Skill 詳細

framevideo-voiceover-ssml

プロジェクトの字幕やトランスクリプトからナレーションを作成し、音素、ブレーク、数字マークアップを備えたSSML風の音声スクリプトをボイスオーバー向けに執筆します。

一致度直接一致オーディオと音声 向けにレビュー済み
出典chanjing-ai/​framevideo外部ソース
報告インストール数156人気度の参考値

使用前に確認

自動レビューは関連性のみを確認し、安全性や推奨を保証しません。使用前に出典の説明を読んでください。

保存された出典プレビュー

SKILL.md

これはレビュー時に保存された抜粋です。完全で最新の内容は外部ソースを確認してください。

---
name: framevideo-voiceover-ssml
description: Create narration from project subtitles or AI analysis output, then author SSML-style voice scripts with phoneme, break, and ttnumber markup for FrameVideo voiceovers. Use when the user wants subtitles linked to voice generation, automatic narration from transcript/content, pronunciation fixes, pause insertion, or text normalization before TTS.
---

# FrameVideo Voiceover SSML

Generate narration scripts from existing project content (subtitles, transcripts, AI analysis) and enhance them with SSML-style markup for pronunciation fixes, pauses, and number reading control.

## When To Use

- User wants narration generated from project subtitles or transcript
- Need to fix pronunciation of brand names, technical terms, or foreign words
- Need to control pause timing for dramatic effect or pacing
- Need to normalize how numbers, dates, or prices are spoken
- Expanding terse subtitles into natural narration
- Creating voiceover that stays aligned with visible text

## Do NOT Use

- For simple text-to-speech without markup → use `framevideo-media` (tts) directly
- For digital human video synthesis → use `chanjing-digital-human`
- For choosing TTS voices or providers → use `framevideo-media`
- When the user provides a complete narration script ready for TTS

---

## Quick Start

**Minimal workflow:**

```bash
# 1. Find source text (subtitle file, transcript, or analysis)
cat subtitles.txt
# Output: "Visit chanjing.ai for more info"

# 2. Add SSML markup for pronunciation
echo '<phoneme alphabet="ipa" ph="tʃæn.dʒɪŋ">Chanjing</phoneme> <break time="0.3s"/> Visit chanjing dot AI for more info.' > script-marked.txt

# 3. Generate TTS with markup (if provider supports)
npx framevideo tts script-marked.txt --voice af_heart --output narration.wav

# 4. Transcribe back to get timestamps
npx framevideo transcribe narration.wav --output transcript.json

# 5. Reference in composition
# <audio src="narration.wav" data-start="0" data-duration="5.2"></audio>
```

---

## Core Concepts

### Source Text Priority

Find the source text in this order:

1. **Project subtitles/transcript** — the primary source, already aligned with visible content
2. **AI analysis text** — scene descriptions or product copy extracted during storyboard
3. **Existing narration draft** — if the user provided one
4. **Manual authoring** — last resort when no source exists

**Why subtitles first?** They're already timed to the visual beats and match what viewers see on screen.

### SSML-Style Markup

Three tags for delivery control:

| Tag | Purpose | Example |
|-----|---------|---------|
| `<phoneme>` | Fix pronunciation | `<phoneme alphabet="ipa" ph="tʃæn.dʒɪŋ">Chanjing</phoneme>` |
| `<break>` | Insert pause | `<break time="0.5s"/>` |
| `<ttnumber>` | Control number reading | `<ttnumber pronounce="twenty twenty-six">2026</ttnumber>` |

**Markup is authoring-layer only.** Keep it in your script file, but strip it before sending to TTS providers that don't support SSML.

### Fallback Strategy

Not all TTS providers support SSML tags:

- **Kokoro (local):** Does NOT support SSML — strip tags before generation
- **Chanjing TTS:** Supports `phoneme`, `break`, `ttnumber` — pass through unchanged
- **ElevenLabs:** Supports subset of SSML — check their docs

**Always save both versions:**
- `script-marked.txt` — original with SSML tags
- `script-fallback.txt` — clean text for non-SSML providers

---

## Workflow

### Step 1: Locate Source Text

```bash
# Check for existing subtitles
ls subtitles.txt captions.json transcript.json

# Check AI analysis output (if using website-to-framevideo)
cat STORYBOARD.md SCRIPT.md

# Check existing narration drafts
ls narration-draft.txt script.txt
```

### Step 2: Normalize Into Speakable Script

**Expand terse subtitles:**

```
Source: "New feature: AI search"
Narration: "Introducing our new feature: AI-powered search that understands your intent."
```

**Keep alignment with visible text:**
- If 
GitHub で全文を読む (外部ページ)
関連情報

関連する仕事