Skill 詳細

fish-audio-tts

Fish Audioで音声生成とボイスクローンを実行。ホスト型TTS API、オープンウェイトモデル、ストリーミング、表現力豊かなデリバリーマーカーを網羅。

一致度直接一致オーディオと音声 向けにレビュー済み
出典calesthio/​generative-media-skills外部ソース
報告インストール数72人気度の参考値

使用前に確認

自動レビューは関連性のみを確認し、安全性や推奨を保証しません。使用前に出典の説明を読んでください。

保存された出典プレビュー

SKILL.md

これはレビュー時に保存された抜粋です。完全で最新の内容は外部ソースを確認してください。

---
name: fish-audio-tts
description: >-
  Produce speech and clone voices with Fish Audio — its hosted TTS API (S2.1-Pro / S2-Pro / S1 model lineup, REST + WebSocket streaming, instant and persistent voice cloning) and its open-weight OpenAudio S1-mini / Fish-Speech models for self-hosting. Use this skill when an agent must generate narration or dialogue through Fish Audio, choose between Fish Audio's hosted models and open weights, clone a voice from reference audio, author emotion/tone/special markers for expressive delivery, estimate cost from UTF-8 bytes, wire real-time streaming for a voice agent, decide whether self-hosting beats the API, or review Fish Audio TTS output for production. Do not use it to pick a different provider — it covers Fish Audio specifically.
---

# Fish Audio text-to-speech and voice cloning

Fish Audio is a TTS and voice-cloning provider with two distinct product surfaces that share a lineage but differ in licensing and operation:

1. **Hosted API** (`api.fish.audio`) — the commercial service. Current lineup is branded **S2.1-Pro / S2-Pro / S1**. Paid per UTF-8 byte, streaming, cloning, and a free tier.
2. **Open weights** — released under the **OpenAudio** brand (and earlier as **Fish-Speech**). Downloadable from Hugging Face and GitHub for self-hosting.

The single most important fact to get right: **the hosted API and the open weights are not the same models, and their licenses differ**. The full flagship weights are hosted-only; only smaller distilled weights are openly published, and those carry a **non-commercial** license. Never assume "Fish is open source, so I can self-host it commercially" — verify which artifact and which license apply. See *Open weights and self-hosting* below.

All model names, prices, endpoints, and limits below are volatile. **Verification date for every dated claim in this document: 2026-07-10.** Re-verify against `docs.fish.audio` before quoting these to a user as current.

Labels used throughout:
- **[Doc]** — stated in official Fish Audio / OpenAudio documentation, model cards, or their published technical reports.
- **[Claim]** — a first-party marketing or benchmark claim from Fish Audio/OpenAudio; treat as their assertion, not independently established.
- **[Heuristic]** — a production judgment from practice, not a documented guarantee.

---

## When to use Fish Audio (and when not to)

**Reach for Fish Audio when:**
- You need low-cost, expressive multilingual TTS with strong voice cloning from short reference audio.
- You want fine-grained emotional/paralinguistic control through inline markers, not just SSML prosody.
- You need a real-time streaming voice for an agent (WebSocket, sub-second time-to-first-audio). **[Claim]**
- You want the *option* to self-host later, or to prototype locally on the open weights before committing to the API.

**Prefer something else when:**
- You need certified, contractual voice likeness rights for a specific celebrity/brand voice — Fish Audio does not pre-clear individual use cases and puts the consent burden on you. **[Doc]**
- Your deployment is commercial and you specifically want to run *self-hosted open weights* — the published weights are non-commercial-licensed, so commercial self-hosting needs the hosted API or a separate commercial license. **[Doc]**
- You need a provider with a formal, audited enterprise compliance posture (e.g., signed BAA/DPA guarantees on the free tier) — the free tier explicitly carries no such guarantees. **[Doc]**
- The languages you need fall outside the S1 supported set (see *Languages*) and you need documented, tested quality there.

This skill does not choose *between providers*. If the user is still deciding whether to use Fish Audio at all versus ElevenLabs, Cartesia, etc., surface the tradeoffs but recognize a provider-comparison skill governs that decision.

---

## Hosted API

### Endpoints [Doc, 2026-07-10]
- **REST (batch/stream):** `POST https://api.fish.audio/v1/tts`
- **WebSoc
GitHub で全文を読む (外部ページ)
関連情報

関連する仕事