Skill 详情
framevideo-voiceover-ssml
从项目字幕或转录稿创建旁白,并编写 SSML 风格语音脚本,包含音素、停顿和数字标记用于配音。
使用前先检查
自动化审核只检查相关性,不代表安全审查或推荐。使用前请阅读来源中的说明。
SKILL.md
这段内容是审核时保存的快照。外部来源才是完整且最新的版本。
--- name: framevideo-voiceover-ssml description: Create narration from project subtitles or AI analysis output, then author SSML-style voice scripts with phoneme, break, and ttnumber markup for FrameVideo voiceovers. Use when the user wants subtitles linked to voice generation, automatic narration from transcript/content, pronunciation fixes, pause insertion, or text normalization before TTS. --- # FrameVideo Voiceover SSML Generate narration scripts from existing project content (subtitles, transcripts, AI analysis) and enhance them with SSML-style markup for pronunciation fixes, pauses, and number reading control. ## When To Use - User wants narration generated from project subtitles or transcript - Need to fix pronunciation of brand names, technical terms, or foreign words - Need to control pause timing for dramatic effect or pacing - Need to normalize how numbers, dates, or prices are spoken - Expanding terse subtitles into natural narration - Creating voiceover that stays aligned with visible text ## Do NOT Use - For simple text-to-speech without markup → use `framevideo-media` (tts) directly - For digital human video synthesis → use `chanjing-digital-human` - For choosing TTS voices or providers → use `framevideo-media` - When the user provides a complete narration script ready for TTS --- ## Quick Start **Minimal workflow:** ```bash # 1. Find source text (subtitle file, transcript, or analysis) cat subtitles.txt # Output: "Visit chanjing.ai for more info" # 2. Add SSML markup for pronunciation echo '<phoneme alphabet="ipa" ph="tʃæn.dʒɪŋ">Chanjing</phoneme> <break time="0.3s"/> Visit chanjing dot AI for more info.' > script-marked.txt # 3. Generate TTS with markup (if provider supports) npx framevideo tts script-marked.txt --voice af_heart --output narration.wav # 4. Transcribe back to get timestamps npx framevideo transcribe narration.wav --output transcript.json # 5. Reference in composition # <audio src="narration.wav" data-start="0" data-duration="5.2"></audio> ``` --- ## Core Concepts ### Source Text Priority Find the source text in this order: 1. **Project subtitles/transcript** — the primary source, already aligned with visible content 2. **AI analysis text** — scene descriptions or product copy extracted during storyboard 3. **Existing narration draft** — if the user provided one 4. **Manual authoring** — last resort when no source exists **Why subtitles first?** They're already timed to the visual beats and match what viewers see on screen. ### SSML-Style Markup Three tags for delivery control: | Tag | Purpose | Example | |-----|---------|---------| | `<phoneme>` | Fix pronunciation | `<phoneme alphabet="ipa" ph="tʃæn.dʒɪŋ">Chanjing</phoneme>` | | `<break>` | Insert pause | `<break time="0.5s"/>` | | `<ttnumber>` | Control number reading | `<ttnumber pronounce="twenty twenty-six">2026</ttnumber>` | **Markup is authoring-layer only.** Keep it in your script file, but strip it before sending to TTS providers that don't support SSML. ### Fallback Strategy Not all TTS providers support SSML tags: - **Kokoro (local):** Does NOT support SSML — strip tags before generation - **Chanjing TTS:** Supports `phoneme`, `break`, `ttnumber` — pass through unchanged - **ElevenLabs:** Supports subset of SSML — check their docs **Always save both versions:** - `script-marked.txt` — original with SSML tags - `script-fallback.txt` — clean text for non-SSML providers --- ## Workflow ### Step 1: Locate Source Text ```bash # Check for existing subtitles ls subtitles.txt captions.json transcript.json # Check AI analysis output (if using website-to-framevideo) cat STORYBOARD.md SCRIPT.md # Check existing narration drafts ls narration-draft.txt script.txt ``` ### Step 2: Normalize Into Speakable Script **Expand terse subtitles:** ``` Source: "New feature: AI search" Narration: "Introducing our new feature: AI-powered search that understands your intent." ``` **Keep alignment with visible text:** - If在 GitHub 阅读完整来源 (打开外部页面)