Skill 详情

framevideo-voiceover-ssml

从项目字幕或转录稿创建旁白,并编写 SSML 风格语音脚本,包含音素、停顿和数字标记用于配音。

匹配类型直接匹配已针对 音频与语音 审核
来源chanjing-ai/​framevideo外部来源
报告安装量156仅表示受欢迎程度

使用前先检查

自动化审核只检查相关性,不代表安全审查或推荐。使用前请阅读来源中的说明。

已保存的来源预览

SKILL.md

这段内容是审核时保存的快照。外部来源才是完整且最新的版本。

---
name: framevideo-voiceover-ssml
description: Create narration from project subtitles or AI analysis output, then author SSML-style voice scripts with phoneme, break, and ttnumber markup for FrameVideo voiceovers. Use when the user wants subtitles linked to voice generation, automatic narration from transcript/content, pronunciation fixes, pause insertion, or text normalization before TTS.
---

# FrameVideo Voiceover SSML

Generate narration scripts from existing project content (subtitles, transcripts, AI analysis) and enhance them with SSML-style markup for pronunciation fixes, pauses, and number reading control.

## When To Use

- User wants narration generated from project subtitles or transcript
- Need to fix pronunciation of brand names, technical terms, or foreign words
- Need to control pause timing for dramatic effect or pacing
- Need to normalize how numbers, dates, or prices are spoken
- Expanding terse subtitles into natural narration
- Creating voiceover that stays aligned with visible text

## Do NOT Use

- For simple text-to-speech without markup → use `framevideo-media` (tts) directly
- For digital human video synthesis → use `chanjing-digital-human`
- For choosing TTS voices or providers → use `framevideo-media`
- When the user provides a complete narration script ready for TTS

---

## Quick Start

**Minimal workflow:**

```bash
# 1. Find source text (subtitle file, transcript, or analysis)
cat subtitles.txt
# Output: "Visit chanjing.ai for more info"

# 2. Add SSML markup for pronunciation
echo '<phoneme alphabet="ipa" ph="tʃæn.dʒɪŋ">Chanjing</phoneme> <break time="0.3s"/> Visit chanjing dot AI for more info.' > script-marked.txt

# 3. Generate TTS with markup (if provider supports)
npx framevideo tts script-marked.txt --voice af_heart --output narration.wav

# 4. Transcribe back to get timestamps
npx framevideo transcribe narration.wav --output transcript.json

# 5. Reference in composition
# <audio src="narration.wav" data-start="0" data-duration="5.2"></audio>
```

---

## Core Concepts

### Source Text Priority

Find the source text in this order:

1. **Project subtitles/transcript** — the primary source, already aligned with visible content
2. **AI analysis text** — scene descriptions or product copy extracted during storyboard
3. **Existing narration draft** — if the user provided one
4. **Manual authoring** — last resort when no source exists

**Why subtitles first?** They're already timed to the visual beats and match what viewers see on screen.

### SSML-Style Markup

Three tags for delivery control:

| Tag | Purpose | Example |
|-----|---------|---------|
| `<phoneme>` | Fix pronunciation | `<phoneme alphabet="ipa" ph="tʃæn.dʒɪŋ">Chanjing</phoneme>` |
| `<break>` | Insert pause | `<break time="0.5s"/>` |
| `<ttnumber>` | Control number reading | `<ttnumber pronounce="twenty twenty-six">2026</ttnumber>` |

**Markup is authoring-layer only.** Keep it in your script file, but strip it before sending to TTS providers that don't support SSML.

### Fallback Strategy

Not all TTS providers support SSML tags:

- **Kokoro (local):** Does NOT support SSML — strip tags before generation
- **Chanjing TTS:** Supports `phoneme`, `break`, `ttnumber` — pass through unchanged
- **ElevenLabs:** Supports subset of SSML — check their docs

**Always save both versions:**
- `script-marked.txt` — original with SSML tags
- `script-fallback.txt` — clean text for non-SSML providers

---

## Workflow

### Step 1: Locate Source Text

```bash
# Check for existing subtitles
ls subtitles.txt captions.json transcript.json

# Check AI analysis output (if using website-to-framevideo)
cat STORYBOARD.md SCRIPT.md

# Check existing narration drafts
ls narration-draft.txt script.txt
```

### Step 2: Normalize Into Speakable Script

**Expand terse subtitles:**

```
Source: "New feature: AI search"
Narration: "Introducing our new feature: AI-powered search that understands your intent."
```

**Keep alignment with visible text:**
- If 
在 GitHub 阅读完整来源 (打开外部页面)
相关上下文

相关工作