Skill-Details
fish-audio-tts
Erzeugt Sprache und klont Stimmen mit Fish Audio, einschließlich der gehosteten TTS-API, Open-Weight-Modelle, Streaming und ausdrucksstarker Sprechmarker.
Vor Nutzung prüfen
Die automatische Prüfung bewertet Relevanz, nicht Sicherheit oder Empfehlung. Lies vor der Nutzung die Quellanweisungen.
SKILL.md
Dieser Auszug wurde bei der Prüfung gespeichert. Die externe Quelle enthält die vollständige und aktuelle Version.
--- name: fish-audio-tts description: >- Produce speech and clone voices with Fish Audio — its hosted TTS API (S2.1-Pro / S2-Pro / S1 model lineup, REST + WebSocket streaming, instant and persistent voice cloning) and its open-weight OpenAudio S1-mini / Fish-Speech models for self-hosting. Use this skill when an agent must generate narration or dialogue through Fish Audio, choose between Fish Audio's hosted models and open weights, clone a voice from reference audio, author emotion/tone/special markers for expressive delivery, estimate cost from UTF-8 bytes, wire real-time streaming for a voice agent, decide whether self-hosting beats the API, or review Fish Audio TTS output for production. Do not use it to pick a different provider — it covers Fish Audio specifically. --- # Fish Audio text-to-speech and voice cloning Fish Audio is a TTS and voice-cloning provider with two distinct product surfaces that share a lineage but differ in licensing and operation: 1. **Hosted API** (`api.fish.audio`) — the commercial service. Current lineup is branded **S2.1-Pro / S2-Pro / S1**. Paid per UTF-8 byte, streaming, cloning, and a free tier. 2. **Open weights** — released under the **OpenAudio** brand (and earlier as **Fish-Speech**). Downloadable from Hugging Face and GitHub for self-hosting. The single most important fact to get right: **the hosted API and the open weights are not the same models, and their licenses differ**. The full flagship weights are hosted-only; only smaller distilled weights are openly published, and those carry a **non-commercial** license. Never assume "Fish is open source, so I can self-host it commercially" — verify which artifact and which license apply. See *Open weights and self-hosting* below. All model names, prices, endpoints, and limits below are volatile. **Verification date for every dated claim in this document: 2026-07-10.** Re-verify against `docs.fish.audio` before quoting these to a user as current. Labels used throughout: - **[Doc]** — stated in official Fish Audio / OpenAudio documentation, model cards, or their published technical reports. - **[Claim]** — a first-party marketing or benchmark claim from Fish Audio/OpenAudio; treat as their assertion, not independently established. - **[Heuristic]** — a production judgment from practice, not a documented guarantee. --- ## When to use Fish Audio (and when not to) **Reach for Fish Audio when:** - You need low-cost, expressive multilingual TTS with strong voice cloning from short reference audio. - You want fine-grained emotional/paralinguistic control through inline markers, not just SSML prosody. - You need a real-time streaming voice for an agent (WebSocket, sub-second time-to-first-audio). **[Claim]** - You want the *option* to self-host later, or to prototype locally on the open weights before committing to the API. **Prefer something else when:** - You need certified, contractual voice likeness rights for a specific celebrity/brand voice — Fish Audio does not pre-clear individual use cases and puts the consent burden on you. **[Doc]** - Your deployment is commercial and you specifically want to run *self-hosted open weights* — the published weights are non-commercial-licensed, so commercial self-hosting needs the hosted API or a separate commercial license. **[Doc]** - You need a provider with a formal, audited enterprise compliance posture (e.g., signed BAA/DPA guarantees on the free tier) — the free tier explicitly carries no such guarantees. **[Doc]** - The languages you need fall outside the S1 supported set (see *Languages*) and you need documented, tested quality there. This skill does not choose *between providers*. If the user is still deciding whether to use Fish Audio at all versus ElevenLabs, Cartesia, etc., surface the tradeoffs but recognize a provider-comparison skill governs that decision. --- ## Hosted API ### Endpoints [Doc, 2026-07-10] - **REST (batch/stream):** `POST https://api.fish.audio/v1/tts` - **WebSocVollständige Quelle auf GitHub lesen (öffnet externe Seite)