Skill 详情

plaud-embedded-transcription-api-skill

实现 Plaud Embedded 的 Transcription API,用于上传和转录音频文件,支持语言检测、降噪和说话人分离。

匹配类型直接匹配已针对 音频与语音 审核
来源plaud-ai/​plaud-embedded-skills外部来源
报告安装量64仅表示受欢迎程度

使用前先检查

自动化审核只检查相关性,不代表安全审查或推荐。使用前请阅读来源中的说明。

已保存的来源预览

SKILL.md

这段内容是审核时保存的快照。外部来源才是完整且最新的版本。

---
name: plaud-embedded-transcription-api-skill
description: Skill for users to implement Plaud Embedded's Transcription API. Use this skill when the user wants to transcribe audio files from their Plaud device or mobile app.
---

# Plaud Embedded Transcription API Skill
This skill provides context and instructions on how to upload audio files to Plaud for transcription.

## When To Use This Skill
Use this skill when a user wants to transcribe audio files using Plaud's Transcription Pipeline (the pipeline performs language detection, 
noise reduction, ASR, Speech-to-text, etc.)

### Prerequisites
- [ ] User has audio files on either: 
    1. Their mobile app synced from a Plaud Device
    2. An audio file available for download on a public URL (i.e. S3 download URL)

## How Plaud's Transcription API works

The **Transcription API** is an AI API with two main endpoints.

1. Triggers an async transcription job given a `file_url` to download an audio file
2. Get transcription task status for polling; On task finish, the conversation transcription will be available

## How to Start Transcribing

The Transcription API itself requires a public download URL to download the audio file. 

**If the developer does not have a way to upload files and generate public download URLs already, use the File Upload API**.

Else, skip the File Upload API and use the Transcription API directly.

### File Upload API \[ONLY FOR USERS WHO DO NOT HAVE A DOWNLOAD URL TO ACCESS THEIR AUDIO FILE\]

The File Upload API is an API to use Plaud's cloud storage to upload audio files. 
This can be performed either from your mobile app or from your backend.

Read the [File Upload Overview](https://docs.plaud.ai/plaud-embedded/file-api-overview.md) for the available File Upload API endpoints and data flow.

It is a **3-step multipart upload**:

1. `POST /generate-presigned-urls` with `filesize` and `filetype` → returns `FileId`, `UploadId`, `ChunkSize`, and a `Parts` array of presigned S3 URLs
2. `PUT` up to `ChunkSize` of raw bytes to each `PresignedUrl` (no auth — these are presigned). **Keep the `ETag` response header from every `PUT`** — the next step needs them
3. `POST /complete-upload` with `file_id`, `upload_id`, the `part_list` of `PartNumber`/`ETag` pairs, `filetype`, and `file_md5` → returns the `DownloadUrl`

**IMPORTANT**: The returned `DownloadUrl` is valid for **24 hours**. Pass it as `file_url` to the Transcription API.

#### API Reference (The Upload Step is not included as it's directly to S3)
* [Generating presigned upload URLs](https://docs.plaud.ai/api-reference/file-upload-api/generate-presigned-upload-urls.md) 
* [Completing the upload](https://docs.plaud.ai/api-reference/file-upload-api/complete-multipart-upload.md) 

### Transcription API
After an audio file has been uploaded to a public download API (either via the File Upload API or through a user's unique cloud storage), the user can use the Transcription API on the file URL.

The [Transcription API Overview](https://docs.plaud.ai/plaud-embedded/transcription-api-overview.md) goes through how Plaud's transcription flow works.

**IMPORTANT**: The Transcription API authenticates with your `X-Client-Id` and `X-Client-Api-Key` headers (the `api_key` is NOT your `client_secret` — grab it from the developer portal under App Settings > API Keys). 

Supported audio formats for `file_url` are **M4A, MP3, and WAV**. Recordings **exceeding 5 hours** should be broken into chunks and transcribed in parts.

`POST` accepts an optional `params` object to tune the pipeline:

| Param | Default | Purpose |
| ---- | ---- | ---- |
| `transcribe.language` | `auto` | BCP-47 code (`en-US`, `zh-CN`) or `auto` |
| `transcribe.detection_level` | `segment` | Language identification level (`segment` or `chapter`) |
| `vad.decode_silence` | `false` | Whether to decode silent regions |
| `diarization.enabled` | `false` | Identify and label speakers |
| `diarization.return_embedding` | `false` | Return speaker embed
在 GitHub 阅读完整来源 (打开外部页面)
相关上下文

相关工作