Skill detail

framevideo-voiceover-ssml

Creates narration from project subtitles or transcripts and authors SSML-style voice scripts with phoneme, break, and number markup for voiceovers.

MatchDirectReviewed for Audio and Voice
Sourcechanjing-ai/​framevideoExternal source
Reported installs156Popularity signal only

Inspect before use

Automated review checks relevance, not safety or endorsement. Read the source instructions before using this skill.

Saved source preview

SKILL.md

The saved excerpt is a snapshot from review. The external source remains the complete and most current version.

---
name: framevideo-voiceover-ssml
description: Create narration from project subtitles or AI analysis output, then author SSML-style voice scripts with phoneme, break, and ttnumber markup for FrameVideo voiceovers. Use when the user wants subtitles linked to voice generation, automatic narration from transcript/content, pronunciation fixes, pause insertion, or text normalization before TTS.
---

# FrameVideo Voiceover SSML

Generate narration scripts from existing project content (subtitles, transcripts, AI analysis) and enhance them with SSML-style markup for pronunciation fixes, pauses, and number reading control.

## When To Use

- User wants narration generated from project subtitles or transcript
- Need to fix pronunciation of brand names, technical terms, or foreign words
- Need to control pause timing for dramatic effect or pacing
- Need to normalize how numbers, dates, or prices are spoken
- Expanding terse subtitles into natural narration
- Creating voiceover that stays aligned with visible text

## Do NOT Use

- For simple text-to-speech without markup → use `framevideo-media` (tts) directly
- For digital human video synthesis → use `chanjing-digital-human`
- For choosing TTS voices or providers → use `framevideo-media`
- When the user provides a complete narration script ready for TTS

---

## Quick Start

**Minimal workflow:**

```bash
# 1. Find source text (subtitle file, transcript, or analysis)
cat subtitles.txt
# Output: "Visit chanjing.ai for more info"

# 2. Add SSML markup for pronunciation
echo '<phoneme alphabet="ipa" ph="tʃæn.dʒɪŋ">Chanjing</phoneme> <break time="0.3s"/> Visit chanjing dot AI for more info.' > script-marked.txt

# 3. Generate TTS with markup (if provider supports)
npx framevideo tts script-marked.txt --voice af_heart --output narration.wav

# 4. Transcribe back to get timestamps
npx framevideo transcribe narration.wav --output transcript.json

# 5. Reference in composition
# <audio src="narration.wav" data-start="0" data-duration="5.2"></audio>
```

---

## Core Concepts

### Source Text Priority

Find the source text in this order:

1. **Project subtitles/transcript** — the primary source, already aligned with visible content
2. **AI analysis text** — scene descriptions or product copy extracted during storyboard
3. **Existing narration draft** — if the user provided one
4. **Manual authoring** — last resort when no source exists

**Why subtitles first?** They're already timed to the visual beats and match what viewers see on screen.

### SSML-Style Markup

Three tags for delivery control:

| Tag | Purpose | Example |
|-----|---------|---------|
| `<phoneme>` | Fix pronunciation | `<phoneme alphabet="ipa" ph="tʃæn.dʒɪŋ">Chanjing</phoneme>` |
| `<break>` | Insert pause | `<break time="0.5s"/>` |
| `<ttnumber>` | Control number reading | `<ttnumber pronounce="twenty twenty-six">2026</ttnumber>` |

**Markup is authoring-layer only.** Keep it in your script file, but strip it before sending to TTS providers that don't support SSML.

### Fallback Strategy

Not all TTS providers support SSML tags:

- **Kokoro (local):** Does NOT support SSML — strip tags before generation
- **Chanjing TTS:** Supports `phoneme`, `break`, `ttnumber` — pass through unchanged
- **ElevenLabs:** Supports subset of SSML — check their docs

**Always save both versions:**
- `script-marked.txt` — original with SSML tags
- `script-fallback.txt` — clean text for non-SSML providers

---

## Workflow

### Step 1: Locate Source Text

```bash
# Check for existing subtitles
ls subtitles.txt captions.json transcript.json

# Check AI analysis output (if using website-to-framevideo)
cat STORYBOARD.md SCRIPT.md

# Check existing narration drafts
ls narration-draft.txt script.txt
```

### Step 2: Normalize Into Speakable Script

**Expand terse subtitles:**

```
Source: "New feature: AI search"
Narration: "Introducing our new feature: AI-powered search that understands your intent."
```

**Keep alignment with visible text:**
- If 
Read the full source on GitHub (opens external page)
Context

Related work