Skill 詳細

voice-audio-engineer

ElevenLabs による音声合成、TTS、声質クローン、ポッドキャスト制作、音声処理のエキスパート。ラウドネス基準と対話ミックスを網羅。

一致度直接一致オーディオと音声 向けにレビュー済み
出典curiositech/​some_claude_skills外部ソース
報告インストール数356人気度の参考値

使用前に確認

自動レビューは関連性のみを確認し、安全性や推奨を保証しません。使用前に出典の説明を読んでください。

保存された出典プレビュー

SKILL.md

これはレビュー時に保存された抜粋です。完全で最新の内容は外部ソースを確認してください。

---
name: voice-audio-engineer
description: Expert in voice synthesis, TTS, voice cloning, podcast production, speech processing, and voice UI design via ElevenLabs integration. Specializes in vocal clarity, loudness standards (LUFS),
  de-essing, dialogue mixing, and voice transformation. Activate on 'TTS', 'text-to-speech', 'voice clone', 'voice synthesis', 'ElevenLabs', 'podcast', 'voice recording', 'speech-to-speech', 'voice UI',
  'audiobook', 'dialogue'. NOT for spatial audio (use sound-engineer), music production (use DAW tools), game audio middleware (use sound-engineer), sound effects generation (use sound-engineer with ElevenLabs
  SFX), or live concert audio.
allowed-tools: Read,Write,Edit,Bash,mcp__firecrawl__firecrawl_search,WebFetch,mcp__ElevenLabs__text_to_speech,mcp__ElevenLabs__speech_to_speech,mcp__ElevenLabs__voice_clone,mcp__ElevenLabs__search_voices,mcp__ElevenLabs__speech_to_text,mcp__ElevenLabs__isolate_audio,mcp__ElevenLabs__create_agent
metadata:
  category: Design & Creative
  pairs-with:
  - skill: sound-engineer
    reason: Full audio production pipeline
  - skill: speech-pathology-ai
    reason: Clinical voice applications
  tags:
  - voice
  - tts
  - elevenlabs
  - podcast
  - synthesis
---

# Voice & Audio Engineer: Voice Synthesis, TTS & Speech Processing

Expert in voice synthesis, speech processing, and vocal production using ElevenLabs and professional audio techniques. Specializes in TTS, voice cloning, podcast production, and voice UI design.

## When to Use This Skill

✅ **Use for:**
- Text-to-speech (TTS) generation
- Voice cloning and voice design
- Speech-to-speech voice transformation
- Podcast production and editing
- Audiobook production
- Voice UI/conversational AI audio
- Dialogue mixing and processing
- Loudness normalization (LUFS)
- Voice quality enhancement (de-essing, compression)
- Transcription and speech-to-text

❌ **Do NOT use for:**
- Spatial audio (HRTF, Ambisonics) → **sound-engineer**
- Sound effects generation → **sound-engineer** (ElevenLabs SFX)
- Game audio middleware (Wwise, FMOD) → **sound-engineer**
- Music composition/production → DAW tools
- Live concert/event audio → specialized domain

## MCP Integrations

| MCP Tool | Purpose |
|----------|---------|
| `text_to_speech` | Generate speech from text with voice selection |
| `speech_to_speech` | Transform voice recordings to different voices |
| `voice_clone` | Create instant voice clones from audio samples |
| `search_voices` | Find voices in ElevenLabs library |
| `speech_to_text` | Transcribe audio with speaker diarization |
| `isolate_audio` | Separate voice from background noise |
| `create_agent` | Build conversational AI agents with voice |

## Expert vs Novice Shibboleths

| Topic | Novice | Expert |
|-------|--------|--------|
| **TTS quality** | "Any voice works" | Matches voice to brand; considers emotion, pace, style |
| **Voice cloning** | "Upload any audio" | Knows 30s-3min of clean, varied speech needed; single speaker |
| **Loudness** | "Make it loud" | Targets -16 to -19 LUFS for podcasts; -14 for streaming |
| **De-essing** | "Doesn't matter" | Knows sibilance lives at 5-8kHz; frequency-selective compression |
| **Compression** | "Squash it" | Uses 3:1-4:1 for dialogue; slow attack (10-20ms) to preserve transients |
| **High-pass** | "Never use it" | Always HPF at 80-100Hz for voice; removes rumble, plosives |
| **True peak** | "Peak is peak" | Knows intersample peaks exceed 0dBFS; targets -1 dBTP |
| **ElevenLabs models** | "Use default" | `eleven_multilingual_v2` for quality; `eleven_flash_v2_5` for speed |

## Common Anti-Patterns

### Anti-Pattern: Uploading Noisy Audio for Voice Cloning
**What it looks like**: Voice clone from phone recording with background noise, echo
**Why it's wrong**: Clone learns the noise; output has artifacts
**What to do instead**: Use `isolate_audio` first; record in quiet space; provide 1-3 min of varied speech

### Anti-Pattern: Ignoring Loudness Standards
**
GitHub で全文を読む (外部ページ)
関連情報

関連する仕事