Claude Skills for Audio and Voice
Audio and voice skills cover work that runs in both directions: speech to text and text to speech. On the speech-to-text side, that includes transcription of audio and video, speaker diarization, word-level timestamps, and summarization of recordings such as interviews or meetings. On the text-to-speech side, it includes narration, voiceovers, accessibility reads, audio prompts, batch speech generation, and dubbing or translation. Some skills also handle general audio processing. The list below is ranked by observed search demand and reviewed for relevance to this topic. It is not an endorsement, and inclusion does not guarantee quality or fit. Use it as a starting point: read each skill's own description and requirements, check what credentials or API keys it expects, and confirm its scope before adding it to your agent.
8 to start with
45 reviewed options
Start with these.
Audio work runs both directions: speech to text (transcribe, summarize) and text to speech (voiceover, dubbing).
These are the closest matches for this work. Start with the first one and open it to install.
- Rank 01View Agent SkillDirect
text-to-speech
Converts text to speech using ElevenLabs voice AI, supporting 70+ languages, multiple models, voice settings, and voiceovers.
Needs (stated): API key · Network access
11,935 · popularity - Rank 02View Agent SkillDirect
speech-to-text
Transcribes audio to text using ElevenLabs Scribe v2 with 90+ languages, speaker diarization, word-level timestamps, and keyterm prompting.
Needs (stated): API key · Network access
8,398 · popularity - Rank 03View Agent SkillDirect
speech
OfficialUse when the user asks for text-to-speech narration or voiceover, accessibility reads, audio prompts, or batch speech generation via the OpenAI Audio API; run the bundled CLI (`scripts/text_to_speech.py`) with built-in voices and require `OPENAI_API_KEY` for live calls. Custom voice creation is out of scope.
- Rank 04View Agent SkillDirect
transcribe
OfficialTranscribe audio files to text with optional diarization and known-speaker hints. Use when a user asks to transcribe speech from audio/video, extract text from recordings, or label speakers in interviews or meetings.
- Rank 05View Agent SkillDirect
qianwen-audio-tts
Synthesizes speech from text with Qwen TTS models, supporting voiceovers, narration, and TTS applications via HTTP API and CosyVoice.
Needs (stated): Python
4,758 · popularity - Rank 06View Agent SkillDirect
qwencloud-audio-tts
Synthesizes speech from text with Qwen TTS models, supporting voiceovers, narration, and TTS applications via HTTP and WebSocket.
Needs (stated): Python
1,439 · popularity - Rank 07View Agent SkillDirect
azure-speech-to-text
Transcribes audio to text using Azure AI Speech Fast Transcription REST API with word-level timestamps, speaker diarization, and multi-language identification.
Needs (stated): Network access
479 · popularity - Rank 08View Agent SkillDirect
azure-text-to-speech
Generates neural narration audio using Azure AI Speech REST text-to-speech with multilingual voices and SSML prosody control.
Needs (stated): Network access
299 · popularity
Explore 37 more Agent Skills
- Rank 09View Agent SkillDirect
byted-text-to-speech
Synthesizes text to speech using the Volcengine Doubao TTS API, supporting streaming, multiple voices, and speed/pitch/volume control.
165 · popularity - Rank 10View Agent SkillDirect
deepgram-js-speech-to-text
Uses Deepgram Speech-to-Text v1 in JavaScript/TypeScript for prerecorded and live audio transcription via REST and WebSocket.
Needs (stated): Node.js
69 · popularity - Rank 11View Agent SkillDirect
fish-audio
Generates AI text-to-speech audio, uses saved voices, or creates one-shot voice clones via the AceDataCloud Fish Audio API.
4,092 · popularity - Rank 12View Agent SkillDirect
fish-audio-tts
Produces speech and clones voices with Fish Audio, covering its hosted TTS API, open-weight models, streaming, and expressive delivery markers.
72 · popularity - Rank 13View Agent SkillDirect
blog-audio
Generates audio narration of blog posts using Google Gemini TTS, supporting summary, full read-aloud, and two-speaker podcast modes with 30 voices.
Needs (stated): API key
2,256 · popularity - Rank 14View Agent SkillDirect
audio-transcription
Transcribes local or remote audio into durable text and timestamp artifacts through PostPlus, with optional timestamps and subtitle-ready outputs.
1,293 · popularity - Rank 15View Agent SkillDirect
eachlabs-voice-audio
Provides text-to-speech, speech-to-text, voice conversion, and audio processing via EachLabs AI models including ElevenLabs TTS and Whisper.
289 · popularity - Rank 16View Agent SkillDirect
venice-audio-transcription
Transcribes audio files to text via Venice's OpenAI-compatible /audio/transcriptions endpoint, covering models, formats, timestamps, and language hints.
156 · popularity - Rank 17View Agent SkillDirect
gemini-tts
Generates speech from text using Google Gemini TTS models, supporting multiple voices, multi-speaker conversations, and streaming.
85 · popularity - Rank 18View Agent SkillDirect
script-to-voiceover
Turns a plain-text script into an AI-narrated voiceover, choosing delivery through voice selection and script craft.
104 · popularity - Rank 19View Agent SkillDirect
plaud-embedded-transcription-api-skill
Implements Plaud Embedded's Transcription API to upload and transcribe audio files, with language detection, noise reduction, and diarization.
64 · popularity - Rank 20View Agent SkillDirect
modelslab-audio-generation
Generates speech, music, and sound effects using ModelsLab's v7 Voice API, supporting TTS, STT, voice conversion, and dubbing.
63 · popularity - Rank 21View Agent SkillDirect
asr
Implements speech-to-text (ASR) using the z-ai-web-dev-sdk, transcribing audio files and supporting voice input features.
204 · popularity - Rank 22View Agent SkillDirect
moss-tts-nano-speech
Expert skill for MOSS-TTS-Nano, a 0.1B multilingual real-time TTS model that runs on CPU with voice cloning and streaming support.
465 · popularity - Rank 23View Agent SkillDirect
voice-audio-engineer
Expert in voice synthesis, TTS, voice cloning, podcast production, and speech processing via ElevenLabs, covering loudness standards and dialogue mixing.
356 · popularity - Rank 24View Agent SkillDirect
ai-voiceover
AI narration and voiceover mini-skill using ElevenLabs to pick voices, write for the ear, and direct delivery for video content.
309 · popularity - Rank 25View Agent SkillDirect
voice-design
Selects and creates AI voices for content using ElevenLabs, Qwen3-TTS, and other platforms, matching voice characteristics to brand and audience.
188 · popularity - Rank 26View Agent SkillDirect
framevideo-voiceover-ssml
Creates narration from project subtitles or transcripts and authors SSML-style voice scripts with phoneme, break, and number markup for voiceovers.
156 · popularity - Rank 27View Agent SkillDirect
azure-speech-to-text-rest-py
Transcribes short audio files (up to 60 seconds) using the Azure Speech to Text REST API in Python without the Speech SDK.
79 · popularity - Rank 28View Agent SkillPossible
azure-ai
OfficialUse for Azure AI: Search, Speech, OpenAI, Document Intelligence. Helps with search, vector/hybrid search, speech-to-text, text-to-speech, transcription, OCR. WHEN: AI Search, query search, vector search, hybrid search, semantic search, speech-to-text, text-to-speech, transcribe, OCR, convert text to speech.
- Rank 29View Agent SkillPossible
azure-ai-voicelive-py
OfficialBuild real-time voice AI applications using Azure AI Voice Live SDK (azure-ai-voicelive). Use this skill when creating Python applications that need real-time bidirectional audio communication with Azure AI, including voice assistants, voice-enabled chatbots, real-time speech-to-speech translation, voice-driven avatars, or any WebSocket-based audio streaming with AI models. Supports Server VAD (Voice Activity Detection), turn-based conversation, function calling, MCP tools, avatar integration, and transcription.
Needs (stated): API key
- Rank 30View Agent SkillPossible
azure-ai-voicelive-dotnet
Official| Azure AI Voice Live SDK for .NET. Build real-time voice AI applications with bidirectional WebSocket communication. Use for voice assistants, conversational AI, real-time speech-to-speech, and voice-enabled chatbots. Triggers: "voice live", "real-time voice", "VoiceLiveClient", "VoiceLiveSession", "voice assistant .NET", "bidirectional audio", "speech-to-speech".
Needs (stated): API key
- Rank 31View Agent SkillPossible
azure-ai-voicelive-java
Official| Azure AI VoiceLive SDK for Java. Real-time bidirectional voice conversations with AI assistants using WebSocket. Triggers: "VoiceLiveClient java", "voice assistant java", "real-time voice java", "audio streaming java", "voice activity detection java".
Needs (stated): API key
- Rank 32View Agent SkillPossible
azure-ai-contentunderstanding-py
Official| Azure AI Content Understanding SDK for Python. Use for multimodal content extraction from documents, images, audio, and video. Triggers: "azure-ai-contentunderstanding", "ContentUnderstandingClient", "multimodal analysis", "document extraction", "video analysis", "audio transcription".
- Rank 33View Agent SkillPossible
azure-ai-openai-dotnet
Official| Azure OpenAI SDK for .NET. Client library for Azure OpenAI and OpenAI services. Use for chat completions, embeddings, image generation, audio transcription, and assistants. Triggers: "Azure OpenAI", "AzureOpenAIClient", "ChatClient", "chat completions .NET", "GPT-4", "embeddings", "DALL-E", "Whisper", "OpenAI .NET".
Needs (stated): API key
- Rank 34View Agent SkillPossible
syncfusion-react-speech-to-text
Implements the Syncfusion React SpeechToText component for real-time speech recognition, microphone input, and voice-enabled forms.
627 · popularity - Rank 35View Agent SkillPossible
syncfusion-blazor-speech-to-text
Implements speech-to-text voice input in Blazor applications using the Syncfusion SpeechToText component with real-time recognition and multi-language support.
324 · popularity - Rank 36View Agent SkillPossible
syncfusion-angular-speech-to-text
Implements the Syncfusion Angular SpeechToText component for real-time speech-to-text conversion with localization and error handling.
268 · popularity - Rank 37View Agent SkillPossible
audio-voice-recovery
Audio forensics and voice recovery guidelines for enhancing degraded recordings, noise reduction, voice isolation, and difficult transcription.
298 · popularity - Rank 38View Agent SkillPossible
voiceover-direction
Directs voice talent to deliver performances matching brand vision, covering VO briefs, casting, speakable scripts, and session direction.
275 · popularity - Rank 39View Agent SkillPossible
precision-voiceover-sync
Aligns narration timing to on-screen actions, running automated alignment then hand-tuning sync points.
102 · popularity - Rank 40View Agent SkillPossible
media-use
NotableAgent Media OS, the single skill for every media need in a HyperFrames project. Resolve BGM, SFX, image, icon, brand logo, voice, color grade, or LUT into a frozen local file or paste-ready block + ledger record (one verb, `resolve`); generate via TTS / music / image models when the catalog misses; produce voiceover, transcription, captions, and background removal through one shared audio engine; operate on media (cut / reframe / transform); and reuse assets across projects. Also use for vague feedback that real footage looks dark, flat, boring, should feel retro/camcorder/print/ASCII, needs p
447,090 · popularity - Rank 41View Agent SkillPossible
ai-avatar-video
NotableCreate AI avatar and talking head videos via inference.sh CLI. Recommended: P-Video-Avatar (fastest, cheapest, built-in TTS). Also: OmniHuman, Fabric, PixVerse. Audio: Inworld TTS-2 (100+ languages, emotion steering for characters), ElevenLabs, Kokoro. Capabilities: audio-driven avatars, text-to-avatar, lipsync videos, talking head generation, virtual presenters, UGC content. Use for: AI presenters, explainer videos, virtual influencers, dubbing, marketing videos, UGC ads, gaming avatars, NPC dialogue. Triggers: ai avatar, talking head, lipsync, avatar video, virtual presenter, ai spokesperson
142,016 · popularity - Rank 42View Agent SkillPossible
wonda-cli
NotableUsing the Wonda CLI to generate images, videos, music, and audio from the terminal — plus LinkedIn, Reddit, and X/Twitter research and automation
74,605 · popularity - Rank 43View Agent SkillPossible
video-producer-agent
Orchestrates complete videos with voiceover, music, and visuals, combining script, Gemini TTS voiceover, and final assembly.
58 · popularity - Rank 44View Agent SkillPossible
fish-audio-sdk
Writes code with official Fish Audio SDKs for text-to-speech, speech-to-text, voice cloning, and realtime WebSocket TTS in Python and JavaScript.
873 · popularity - Rank 45View Agent SkillPossible
byted-mediakit-voiceover-editing
Volcano Engine MediaKit voiceover-editing skill for cutting talking-head content, covering ASR, audio processing, and video export.
301 · popularity
Frequently asked questions
- What kinds of tasks fall under audio and voice skills?
- They cover transcription, text to speech, summarization, dubbing and translation, and audio processing. Audio work runs both directions: speech to text for transcribe and summarize tasks, and text to speech for voiceover and dubbing tasks.
- How is this list ordered?
- It is ranked by observed search demand and reviewed for relevance to the topic. It is not an endorsement of any skill.
- Do these skills need API keys or accounts?
- It depends on the skill. Some text-to-speech skills require an API key for live calls, and others may create custom voices as part of their scope while some explicitly leave that out of scope.
- Can these skills handle multiple languages and speakers?
- Some do. Certain transcription skills support 90+ languages with speaker diarization and word-level timestamps, and some text-to-speech skills support 70+ languages with multiple models and voice settings.
How to install an Agent Skill
Start with the first skill on this page. Claude Code reads a SKILL.md once it is in the skills directory:
- Claude Code — copy the SKILL.md for text-to-speech into ~/.claude/skills/text-to-speech/ for yourself, or .claude/skills/text-to-speech/ to share it with a project.
- From the registry — run npx skills add elevenlabs/skills@text-to-speech.
- Or download text-to-speech from its source and place that SKILL.md in the directory above.
How to choose Check the source and the skill’s own instructions before you install. Official and Notable describe provenance, not safety.