Skill 详情

azure-text-to-speech

使用 Azure AI Speech REST 文本转语音生成神经旁白音频,支持多语言语音和 SSML 韵律控制。

声明的前提(自述): Network access

匹配类型直接匹配已针对 音频与语音 审核
来源calesthio/​openmontage外部来源
报告安装量299仅表示受欢迎程度

使用前先检查

自动化审核只检查相关性,不代表安全审查或推荐。使用前请阅读来源中的说明。

已保存的来源预览

SKILL.md

这段内容是审核时保存的快照。外部来源才是完整且最新的版本。

---
name: azure-text-to-speech
description: Generate neural narration audio using Azure AI Speech (REST text-to-speech). Use when synthesizing voiceovers or narration in OpenMontage. Optional cloud TTS provider — preferred when AZURE_SPEECH_KEY is configured; the local piper_tts remains the default offline path. Shares one Speech resource with azure_stt.
license: MIT
compatibility: Requires internet access and an Azure AI Speech resource (AZURE_SPEECH_KEY + AZURE_SPEECH_REGION).
metadata: {"openclaw": {"requires": {"env": ["AZURE_SPEECH_KEY", "AZURE_SPEECH_REGION"]}, "primaryEnv": "AZURE_SPEECH_KEY"}}
---

# Azure AI Speech — Text-to-Speech

Generate narration with **Azure neural TTS** — high-quality multilingual voices,
SSML prosody control, and express-as styles, served synchronously by the REST
`/cognitiveservices/v1` endpoint (no token exchange, Blob storage, or job
polling). In OpenMontage this is exposed through the `azure_tts` tool
(`capability=tts`, `provider=azure`). It is an **optional cloud TTS provider** —
when `AZURE_SPEECH_KEY` is configured, prefer it for high-quality cloud
narration. The local `piper_tts` remains the **default offline path** and the
fallback when Azure is unavailable; `elevenlabs_tts` remains the choice for
voice cloning.

> Docs: [REST text to speech](https://learn.microsoft.com/azure/ai-services/speech-service/rest-text-to-speech) · [Voice gallery](https://speech.microsoft.com/portal/voicegallery)

## Setup

Same Speech resource as `azure_stt` — **one key/region unlocks both directions**
(STT and TTS). Create a **Speech** resource in the
[Azure portal](https://portal.azure.com); copy the key and region from its
**Keys and Endpoint** page.

```bash
export AZURE_SPEECH_KEY=your_speech_resource_key
export AZURE_SPEECH_REGION=eastus        # your resource's region
# export AZURE_TTS_ENDPOINT=https://...  # optional: full custom TTS host
#   (the TTS host is https://<region>.tts.speech.microsoft.com — a different
#    subdomain than the STT endpoint, hence the separate override var)
```

`azure_tts` reports `AVAILABLE` once `AZURE_SPEECH_KEY` plus either
`AZURE_SPEECH_REGION` or `AZURE_TTS_ENDPOINT` are set.

## Using it in a pipeline

Route through `tts_selector` as usual (it auto-discovers `azure_tts`), or call
the provider tool directly when the user has approved Azure:

```python
from tools.tool_registry import registry
registry.discover()
tts = registry._tools["azure_tts"]

result = tts.execute({
    "text": "Every design decision in this dashboard has a reason.",
    "voice": "andrew",                 # alias or full Azure short name
    "rate": "-4%",                     # slightly slower for narration
    # "style": "narration-professional",  # for voices that support styles
    "output_path": "projects/my-video/assets/audio/seg_001.mp3",
    "output_format": "mp3",            # or "wav" (48kHz PCM) for mixing
})
```

If `azure_tts` is unavailable (no key) or errors, fall back per its declared
chain: `elevenlabs_tts` → `openai_tts` → `piper_tts`.

## Voice selection

Curated shortlist (aliases accepted by the `voice` param):

| Alias | Voice | Character |
|-------|-------|-----------|
| `andrew` | en-US-AndrewMultilingualNeural | warm, confident, conversational — the default; founder/explainer register |
| `brandon` | en-US-BrandonMultilingualNeural | deeper, measured |
| `ava` | en-US-AvaMultilingualNeural | confident, bright female |
| `guy` | en-US-GuyNeural | authoritative |
| `jenny` | en-US-JennyNeural | friendly, clear |

Any valid Azure voice short name may be passed verbatim (e.g.
`de-DE-KatjaNeural`); the *Multilingual* voices handle non-English text well —
set `locale` to match the text's language for correct SSML.

## Parameters that matter

- **`rate` / `pitch`** — SSML prosody. Narration usually reads best slightly
  slowed (`"-4%"` to `"-8%"`); leave pitch at `"0%"` unless correcting a voice.
- **`style`** — express-as style for voices that support it
  (`narration-professio
在 GitHub 阅读完整来源 (打开外部页面)
相关上下文

相关工作