Claude Skills for Audio and Voice

Audio and voice skills cover work that runs in both directions: speech to text and text to speech. On the speech-to-text side, that includes transcription of audio and video, speaker diarization, word-level timestamps, and summarization of recordings such as interviews or meetings. On the text-to-speech side, it includes narration, voiceovers, accessibility reads, audio prompts, batch speech generation, and dubbing or translation. Some skills also handle general audio processing. The list below is ranked by observed search demand and reviewed for relevance to this topic. It is not an endorsement, and inclusion does not guarantee quality or fit. Use it as a starting point: read each skill's own description and requirements, check what credentials or API keys it expects, and confirm its scope before adding it to your agent.

8 to start with
45 reviewed options

Start with these.

Audio work runs both directions: speech to text (transcribe, summarize) and text to speech (voiceover, dubbing).

These are the closest matches for this work. Start with the first one and open it to install.

45 reviewed options

  1. Rank 04
    Direct

    transcribe

    Official

    Transcribe audio files to text with optional diarization and known-speaker hints. Use when a user asks to transcribe speech from audio/video, extract text from recordings, or label speakers in interviews or meetings.

    • Transcription
    View Agent Skill
  2. Rank 05
    Direct

    qianwen-audio-tts

    Synthesizes speech from text with Qwen TTS models, supporting voiceovers, narration, and TTS applications via HTTP API and CosyVoice.

    • Text to speech

    Needs (stated): Python

    4,758 · popularity
    View Agent Skill
  3. Rank 06
    Direct

    qwencloud-audio-tts

    Synthesizes speech from text with Qwen TTS models, supporting voiceovers, narration, and TTS applications via HTTP and WebSocket.

    • Text to speech

    Needs (stated): Python

    1,439 · popularity
    View Agent Skill
  4. Rank 07
    Direct

    azure-speech-to-text

    Transcribes audio to text using Azure AI Speech Fast Transcription REST API with word-level timestamps, speaker diarization, and multi-language identification.

    • Transcription
    • Text to speech
    • Dubbing and translation

    Needs (stated): Network access

    479 · popularity
    View Agent Skill
  5. Rank 08
    Direct

    azure-text-to-speech

    Generates neural narration audio using Azure AI Speech REST text-to-speech with multilingual voices and SSML prosody control.

    • Text to speech

    Needs (stated): Network access

    299 · popularity
    View Agent Skill
Explore 37 more Agent Skills
  1. Rank 09
    Direct

    byted-text-to-speech

    Synthesizes text to speech using the Volcengine Doubao TTS API, supporting streaming, multiple voices, and speed/pitch/volume control.

    • Text to speech
    165 · popularity
    View Agent Skill
  2. Rank 10
    Direct

    deepgram-js-speech-to-text

    Uses Deepgram Speech-to-Text v1 in JavaScript/TypeScript for prerecorded and live audio transcription via REST and WebSocket.

    • Transcription
    • Text to speech

    Needs (stated): Node.js

    69 · popularity
    View Agent Skill
  3. Rank 11
    Direct

    fish-audio

    Generates AI text-to-speech audio, uses saved voices, or creates one-shot voice clones via the AceDataCloud Fish Audio API.

    • Text to speech
    • Audio processing
    4,092 · popularity
    View Agent Skill
  4. Rank 12
    Direct

    fish-audio-tts

    Produces speech and clones voices with Fish Audio, covering its hosted TTS API, open-weight models, streaming, and expressive delivery markers.

    • Text to speech
    72 · popularity
    View Agent Skill
  5. Rank 13
    Direct

    blog-audio

    Generates audio narration of blog posts using Google Gemini TTS, supporting summary, full read-aloud, and two-speaker podcast modes with 30 voices.

    • Text to speech
    • Summarization
    • Audio processing

    Needs (stated): API key

    2,256 · popularity
    View Agent Skill
  6. Rank 14
    Direct

    audio-transcription

    Transcribes local or remote audio into durable text and timestamp artifacts through PostPlus, with optional timestamps and subtitle-ready outputs.

    • Transcription
    • Dubbing and translation
    1,293 · popularity
    View Agent Skill
  7. Rank 15
    Direct

    eachlabs-voice-audio

    Provides text-to-speech, speech-to-text, voice conversion, and audio processing via EachLabs AI models including ElevenLabs TTS and Whisper.

    • Transcription
    • Text to speech
    • Dubbing and translation
    289 · popularity
    View Agent Skill
  8. Rank 16
    Direct

    venice-audio-transcription

    Transcribes audio files to text via Venice's OpenAI-compatible /audio/transcriptions endpoint, covering models, formats, timestamps, and language hints.

    • Transcription
    • Dubbing and translation
    156 · popularity
    View Agent Skill
  9. Rank 17
    Direct

    gemini-tts

    Generates speech from text using Google Gemini TTS models, supporting multiple voices, multi-speaker conversations, and streaming.

    • Text to speech
    • Audio processing
    85 · popularity
    View Agent Skill
  10. Rank 18
    Direct

    script-to-voiceover

    Turns a plain-text script into an AI-narrated voiceover, choosing delivery through voice selection and script craft.

    • Text to speech
    104 · popularity
    View Agent Skill
  11. Rank 19
    Direct

    plaud-embedded-transcription-api-skill

    Implements Plaud Embedded's Transcription API to upload and transcribe audio files, with language detection, noise reduction, and diarization.

    • Transcription
    64 · popularity
    View Agent Skill
  12. Rank 20
    Direct

    modelslab-audio-generation

    Generates speech, music, and sound effects using ModelsLab's v7 Voice API, supporting TTS, STT, voice conversion, and dubbing.

    • Transcription
    • Text to speech
    • Dubbing and translation
    63 · popularity
    View Agent Skill
  13. Rank 21
    Direct

    asr

    Implements speech-to-text (ASR) using the z-ai-web-dev-sdk, transcribing audio files and supporting voice input features.

    • Transcription
    204 · popularity
    View Agent Skill
  14. Rank 22
    Direct

    moss-tts-nano-speech

    Expert skill for MOSS-TTS-Nano, a 0.1B multilingual real-time TTS model that runs on CPU with voice cloning and streaming support.

    • Text to speech
    465 · popularity
    View Agent Skill
  15. Rank 23
    Direct

    voice-audio-engineer

    Expert in voice synthesis, TTS, voice cloning, podcast production, and speech processing via ElevenLabs, covering loudness standards and dialogue mixing.

    • Transcription
    • Text to speech
    • Audio processing
    356 · popularity
    View Agent Skill
  16. Rank 24
    Direct

    ai-voiceover

    AI narration and voiceover mini-skill using ElevenLabs to pick voices, write for the ear, and direct delivery for video content.

    • Text to speech
    • Dubbing and translation
    309 · popularity
    View Agent Skill
  17. Rank 25
    Direct

    voice-design

    Selects and creates AI voices for content using ElevenLabs, Qwen3-TTS, and other platforms, matching voice characteristics to brand and audience.

    • Text to speech
    • Summarization
    • Dubbing and translation
    188 · popularity
    View Agent Skill
  18. Rank 26
    Direct

    framevideo-voiceover-ssml

    Creates narration from project subtitles or transcripts and authors SSML-style voice scripts with phoneme, break, and number markup for voiceovers.

    • Text to speech
    • Dubbing and translation
    • Audio processing
    156 · popularity
    View Agent Skill
  19. Rank 27
    Direct

    azure-speech-to-text-rest-py

    Transcribes short audio files (up to 60 seconds) using the Azure Speech to Text REST API in Python without the Speech SDK.

    • Transcription
    79 · popularity
    View Agent Skill
  20. Rank 28
    Possible

    azure-ai

    Official

    Use for Azure AI: Search, Speech, OpenAI, Document Intelligence. Helps with search, vector/hybrid search, speech-to-text, text-to-speech, transcription, OCR. WHEN: AI Search, query search, vector search, hybrid search, semantic search, speech-to-text, text-to-speech, transcribe, OCR, convert text to speech.

    • Transcription
    • Text to speech
    • Dubbing and translation
    View Agent Skill
  21. Rank 29
    Possible

    azure-ai-voicelive-py

    Official

    Build real-time voice AI applications using Azure AI Voice Live SDK (azure-ai-voicelive). Use this skill when creating Python applications that need real-time bidirectional audio communication with Azure AI, including voice assistants, voice-enabled chatbots, real-time speech-to-speech translation, voice-driven avatars, or any WebSocket-based audio streaming with AI models. Supports Server VAD (Voice Activity Detection), turn-based conversation, function calling, MCP tools, avatar integration, and transcription.

    • Transcription
    • Dubbing and translation

    Needs (stated): API key

    View Agent Skill
  22. Rank 30
    Possible

    azure-ai-voicelive-dotnet

    Official

    | Azure AI Voice Live SDK for .NET. Build real-time voice AI applications with bidirectional WebSocket communication. Use for voice assistants, conversational AI, real-time speech-to-speech, and voice-enabled chatbots. Triggers: "voice live", "real-time voice", "VoiceLiveClient", "VoiceLiveSession", "voice assistant .NET", "bidirectional audio", "speech-to-speech".

    Needs (stated): API key

    View Agent Skill
  23. Rank 31
    Possible

    azure-ai-voicelive-java

    Official

    | Azure AI VoiceLive SDK for Java. Real-time bidirectional voice conversations with AI assistants using WebSocket. Triggers: "VoiceLiveClient java", "voice assistant java", "real-time voice java", "audio streaming java", "voice activity detection java".

    • Transcription

    Needs (stated): API key

    View Agent Skill
  24. Rank 32
    Possible

    azure-ai-contentunderstanding-py

    Official

    | Azure AI Content Understanding SDK for Python. Use for multimodal content extraction from documents, images, audio, and video. Triggers: "azure-ai-contentunderstanding", "ContentUnderstandingClient", "multimodal analysis", "document extraction", "video analysis", "audio transcription".

    • Transcription
    View Agent Skill
  25. Rank 33
    Possible

    azure-ai-openai-dotnet

    Official

    | Azure OpenAI SDK for .NET. Client library for Azure OpenAI and OpenAI services. Use for chat completions, embeddings, image generation, audio transcription, and assistants. Triggers: "Azure OpenAI", "AzureOpenAIClient", "ChatClient", "chat completions .NET", "GPT-4", "embeddings", "DALL-E", "Whisper", "OpenAI .NET".

    • Transcription

    Needs (stated): API key

    View Agent Skill
  26. Rank 34
    Possible

    syncfusion-react-speech-to-text

    Implements the Syncfusion React SpeechToText component for real-time speech recognition, microphone input, and voice-enabled forms.

    • Transcription
    • Dubbing and translation
    627 · popularity
    View Agent Skill
  27. Rank 35
    Possible

    syncfusion-blazor-speech-to-text

    Implements speech-to-text voice input in Blazor applications using the Syncfusion SpeechToText component with real-time recognition and multi-language support.

    • Transcription
    324 · popularity
    View Agent Skill
  28. Rank 36
    Possible

    syncfusion-angular-speech-to-text

    Implements the Syncfusion Angular SpeechToText component for real-time speech-to-text conversion with localization and error handling.

    • Transcription
    • Dubbing and translation
    268 · popularity
    View Agent Skill
  29. Rank 37
    Possible

    audio-voice-recovery

    Audio forensics and voice recovery guidelines for enhancing degraded recordings, noise reduction, voice isolation, and difficult transcription.

    • Transcription
    298 · popularity
    View Agent Skill
  30. Rank 38
    Possible

    voiceover-direction

    Directs voice talent to deliver performances matching brand vision, covering VO briefs, casting, speakable scripts, and session direction.

    • Text to speech
    275 · popularity
    View Agent Skill
  31. Rank 39
    Possible

    precision-voiceover-sync

    Aligns narration timing to on-screen actions, running automated alignment then hand-tuning sync points.

    • Text to speech
    102 · popularity
    View Agent Skill
  32. Rank 40
    Possible

    media-use

    Notable

    Agent Media OS, the single skill for every media need in a HyperFrames project. Resolve BGM, SFX, image, icon, brand logo, voice, color grade, or LUT into a frozen local file or paste-ready block + ledger record (one verb, `resolve`); generate via TTS / music / image models when the catalog misses; produce voiceover, transcription, captions, and background removal through one shared audio engine; operate on media (cut / reframe / transform); and reuse assets across projects. Also use for vague feedback that real footage looks dark, flat, boring, should feel retro/camcorder/print/ASCII, needs p

    • Transcription
    • Text to speech
    447,090 · popularity
    View Agent Skill
  33. Rank 41
    Possible

    ai-avatar-video

    Notable

    Create AI avatar and talking head videos via inference.sh CLI. Recommended: P-Video-Avatar (fastest, cheapest, built-in TTS). Also: OmniHuman, Fabric, PixVerse. Audio: Inworld TTS-2 (100+ languages, emotion steering for characters), ElevenLabs, Kokoro. Capabilities: audio-driven avatars, text-to-avatar, lipsync videos, talking head generation, virtual presenters, UGC content. Use for: AI presenters, explainer videos, virtual influencers, dubbing, marketing videos, UGC ads, gaming avatars, NPC dialogue. Triggers: ai avatar, talking head, lipsync, avatar video, virtual presenter, ai spokesperson

    • Text to speech
    • Dubbing and translation
    142,016 · popularity
    View Agent Skill
  34. Rank 42
    Possible

    wonda-cli

    Notable

    Using the Wonda CLI to generate images, videos, music, and audio from the terminal — plus LinkedIn, Reddit, and X/Twitter research and automation

    • Summarization
    74,605 · popularity
    View Agent Skill
  35. Rank 43
    Possible

    video-producer-agent

    Orchestrates complete videos with voiceover, music, and visuals, combining script, Gemini TTS voiceover, and final assembly.

    • Text to speech
    • Summarization
    58 · popularity
    View Agent Skill
  36. Rank 44
    Possible

    fish-audio-sdk

    Writes code with official Fish Audio SDKs for text-to-speech, speech-to-text, voice cloning, and realtime WebSocket TTS in Python and JavaScript.

    • Transcription
    • Text to speech
    873 · popularity
    View Agent Skill
  37. Rank 45
    Possible

    byted-mediakit-voiceover-editing

    Volcano Engine MediaKit voiceover-editing skill for cutting talking-head content, covering ASR, audio processing, and video export.

    • Text to speech
    301 · popularity
    View Agent Skill
Frequently asked questions

Frequently asked questions

What kinds of tasks fall under audio and voice skills?
They cover transcription, text to speech, summarization, dubbing and translation, and audio processing. Audio work runs both directions: speech to text for transcribe and summarize tasks, and text to speech for voiceover and dubbing tasks.
How is this list ordered?
It is ranked by observed search demand and reviewed for relevance to the topic. It is not an endorsement of any skill.
Do these skills need API keys or accounts?
It depends on the skill. Some text-to-speech skills require an API key for live calls, and others may create custom voices as part of their scope while some explicitly leave that out of scope.
Can these skills handle multiple languages and speakers?
Some do. Certain transcription skills support 90+ languages with speaker diarization and word-level timestamps, and some text-to-speech skills support 70+ languages with multiple models and voice settings.

How to install an Agent Skill

Start with the first skill on this page. Claude Code reads a SKILL.md once it is in the skills directory:

  • Claude Code — copy the SKILL.md for text-to-speech into ~/.claude/skills/text-to-speech/ for yourself, or .claude/skills/text-to-speech/ to share it with a project.
  • From the registry — run npx skills add elevenlabs/skills@text-to-speech.
  • Or download text-to-speech from its source and place that SKILL.md in the directory above.

How to choose Check the source and the skill’s own instructions before you install. Official and Notable describe provenance, not safety.