Skill detail

fish-audio-tts

Produces speech and clones voices with Fish Audio, covering its hosted TTS API, open-weight models, streaming, and expressive delivery markers.

MatchDirectReviewed for Audio and Voice
Sourcecalesthio/​generative-media-skillsExternal source
Reported installs72Popularity signal only

Inspect before use

Automated review checks relevance, not safety or endorsement. Read the source instructions before using this skill.

Saved source preview

SKILL.md

The saved excerpt is a snapshot from review. The external source remains the complete and most current version.

---
name: fish-audio-tts
description: >-
  Produce speech and clone voices with Fish Audio — its hosted TTS API (S2.1-Pro / S2-Pro / S1 model lineup, REST + WebSocket streaming, instant and persistent voice cloning) and its open-weight OpenAudio S1-mini / Fish-Speech models for self-hosting. Use this skill when an agent must generate narration or dialogue through Fish Audio, choose between Fish Audio's hosted models and open weights, clone a voice from reference audio, author emotion/tone/special markers for expressive delivery, estimate cost from UTF-8 bytes, wire real-time streaming for a voice agent, decide whether self-hosting beats the API, or review Fish Audio TTS output for production. Do not use it to pick a different provider — it covers Fish Audio specifically.
---

# Fish Audio text-to-speech and voice cloning

Fish Audio is a TTS and voice-cloning provider with two distinct product surfaces that share a lineage but differ in licensing and operation:

1. **Hosted API** (`api.fish.audio`) — the commercial service. Current lineup is branded **S2.1-Pro / S2-Pro / S1**. Paid per UTF-8 byte, streaming, cloning, and a free tier.
2. **Open weights** — released under the **OpenAudio** brand (and earlier as **Fish-Speech**). Downloadable from Hugging Face and GitHub for self-hosting.

The single most important fact to get right: **the hosted API and the open weights are not the same models, and their licenses differ**. The full flagship weights are hosted-only; only smaller distilled weights are openly published, and those carry a **non-commercial** license. Never assume "Fish is open source, so I can self-host it commercially" — verify which artifact and which license apply. See *Open weights and self-hosting* below.

All model names, prices, endpoints, and limits below are volatile. **Verification date for every dated claim in this document: 2026-07-10.** Re-verify against `docs.fish.audio` before quoting these to a user as current.

Labels used throughout:
- **[Doc]** — stated in official Fish Audio / OpenAudio documentation, model cards, or their published technical reports.
- **[Claim]** — a first-party marketing or benchmark claim from Fish Audio/OpenAudio; treat as their assertion, not independently established.
- **[Heuristic]** — a production judgment from practice, not a documented guarantee.

---

## When to use Fish Audio (and when not to)

**Reach for Fish Audio when:**
- You need low-cost, expressive multilingual TTS with strong voice cloning from short reference audio.
- You want fine-grained emotional/paralinguistic control through inline markers, not just SSML prosody.
- You need a real-time streaming voice for an agent (WebSocket, sub-second time-to-first-audio). **[Claim]**
- You want the *option* to self-host later, or to prototype locally on the open weights before committing to the API.

**Prefer something else when:**
- You need certified, contractual voice likeness rights for a specific celebrity/brand voice — Fish Audio does not pre-clear individual use cases and puts the consent burden on you. **[Doc]**
- Your deployment is commercial and you specifically want to run *self-hosted open weights* — the published weights are non-commercial-licensed, so commercial self-hosting needs the hosted API or a separate commercial license. **[Doc]**
- You need a provider with a formal, audited enterprise compliance posture (e.g., signed BAA/DPA guarantees on the free tier) — the free tier explicitly carries no such guarantees. **[Doc]**
- The languages you need fall outside the S1 supported set (see *Languages*) and you need documented, tested quality there.

This skill does not choose *between providers*. If the user is still deciding whether to use Fish Audio at all versus ElevenLabs, Cartesia, etc., surface the tradeoffs but recognize a provider-comparison skill governs that decision.

---

## Hosted API

### Endpoints [Doc, 2026-07-10]
- **REST (batch/stream):** `POST https://api.fish.audio/v1/tts`
- **WebSoc
Read the full source on GitHub (opens external page)
Context

Related work