Skill 详情
fish-audio-tts
使用 Fish Audio 生成语音并克隆声音,涵盖其托管 TTS API、开放权重模型、流式传输和富有表现力的表达标记。
使用前先检查
自动化审核只检查相关性,不代表安全审查或推荐。使用前请阅读来源中的说明。
SKILL.md
这段内容是审核时保存的快照。外部来源才是完整且最新的版本。
--- name: fish-audio-tts description: >- Produce speech and clone voices with Fish Audio — its hosted TTS API (S2.1-Pro / S2-Pro / S1 model lineup, REST + WebSocket streaming, instant and persistent voice cloning) and its open-weight OpenAudio S1-mini / Fish-Speech models for self-hosting. Use this skill when an agent must generate narration or dialogue through Fish Audio, choose between Fish Audio's hosted models and open weights, clone a voice from reference audio, author emotion/tone/special markers for expressive delivery, estimate cost from UTF-8 bytes, wire real-time streaming for a voice agent, decide whether self-hosting beats the API, or review Fish Audio TTS output for production. Do not use it to pick a different provider — it covers Fish Audio specifically. --- # Fish Audio text-to-speech and voice cloning Fish Audio is a TTS and voice-cloning provider with two distinct product surfaces that share a lineage but differ in licensing and operation: 1. **Hosted API** (`api.fish.audio`) — the commercial service. Current lineup is branded **S2.1-Pro / S2-Pro / S1**. Paid per UTF-8 byte, streaming, cloning, and a free tier. 2. **Open weights** — released under the **OpenAudio** brand (and earlier as **Fish-Speech**). Downloadable from Hugging Face and GitHub for self-hosting. The single most important fact to get right: **the hosted API and the open weights are not the same models, and their licenses differ**. The full flagship weights are hosted-only; only smaller distilled weights are openly published, and those carry a **non-commercial** license. Never assume "Fish is open source, so I can self-host it commercially" — verify which artifact and which license apply. See *Open weights and self-hosting* below. All model names, prices, endpoints, and limits below are volatile. **Verification date for every dated claim in this document: 2026-07-10.** Re-verify against `docs.fish.audio` before quoting these to a user as current. Labels used throughout: - **[Doc]** — stated in official Fish Audio / OpenAudio documentation, model cards, or their published technical reports. - **[Claim]** — a first-party marketing or benchmark claim from Fish Audio/OpenAudio; treat as their assertion, not independently established. - **[Heuristic]** — a production judgment from practice, not a documented guarantee. --- ## When to use Fish Audio (and when not to) **Reach for Fish Audio when:** - You need low-cost, expressive multilingual TTS with strong voice cloning from short reference audio. - You want fine-grained emotional/paralinguistic control through inline markers, not just SSML prosody. - You need a real-time streaming voice for an agent (WebSocket, sub-second time-to-first-audio). **[Claim]** - You want the *option* to self-host later, or to prototype locally on the open weights before committing to the API. **Prefer something else when:** - You need certified, contractual voice likeness rights for a specific celebrity/brand voice — Fish Audio does not pre-clear individual use cases and puts the consent burden on you. **[Doc]** - Your deployment is commercial and you specifically want to run *self-hosted open weights* — the published weights are non-commercial-licensed, so commercial self-hosting needs the hosted API or a separate commercial license. **[Doc]** - You need a provider with a formal, audited enterprise compliance posture (e.g., signed BAA/DPA guarantees on the free tier) — the free tier explicitly carries no such guarantees. **[Doc]** - The languages you need fall outside the S1 supported set (see *Languages*) and you need documented, tested quality there. This skill does not choose *between providers*. If the user is still deciding whether to use Fish Audio at all versus ElevenLabs, Cartesia, etc., surface the tradeoffs but recognize a provider-comparison skill governs that decision. --- ## Hosted API ### Endpoints [Doc, 2026-07-10] - **REST (batch/stream):** `POST https://api.fish.audio/v1/tts` - **WebSoc在 GitHub 阅读完整来源 (打开外部页面)