Skill-Details
fish-audio-sdk
Schreibt Code mit offiziellen Fish Audio SDKs für Text-to-Speech, Speech-to-Text, Voice Cloning und Echtzeit-WebSocket-TTS in Python und JavaScript.
Vor Nutzung prüfen
Die automatische Prüfung bewertet Relevanz, nicht Sicherheit oder Empfehlung. Lies vor der Nutzung die Quellanweisungen.
SKILL.md
Dieser Auszug wurde bei der Prüfung gespeichert. Die externe Quelle enthält die vollständige und aktuelle Version.
---
name: fish-audio-sdk
description: Write code with the official Fish Audio SDKs, Python (`fishaudio`, PyPI `fish-audio-sdk`) and JavaScript/TypeScript (`fish-audio`). Use when the user wants text-to-speech, speech-to-text, voice cloning / voice-model management, or realtime WebSocket TTS through the installed SDK rather than raw HTTP. Covers install and auth, sync + async Python, the TypeScript client, exact method signatures and defaults, model selection (including the S2.1 typing caveat), the real exception types, and the Python↔JavaScript naming differences. For raw REST/WebSocket calls without an SDK (curl, unsupported languages, edge runtimes), use the `fish-audio-api` skill instead.
---
# Fish Audio SDK Skill
Use this skill to generate correct, runnable code with the **official Fish Audio SDKs**:
- **Python**: package `fish-audio-sdk` on PyPI, imported as `fishaudio`. (The same wheel still ships a separate legacy `fish_audio_sdk` package. Do **not** mix them; everything here is the modern `fishaudio` package.)
- **JavaScript / TypeScript**: package `fish-audio` on npm, imported as `FishAudioClient`.
If the user wants raw `curl` / HTTP / WebSocket without installing an SDK, use the **`fish-audio-api`** skill instead.
> This file is the index. Deeper, task-specific rules and full examples live in [`references/`](references/). Read the reference for the task you're doing before writing code.
## Global facts
- **Auth:** both SDKs read the API key from the `FISH_API_KEY` environment variable automatically. Get keys at `https://fish.audio/app/api-keys`. Never hardcode a key.
- **Base URL:** `https://api.fish.audio` (override with `base_url=` in Python / `baseUrl:` in JS).
- **Models:** the API supports `s1`, `s2-pro`, `s2.1-pro` (recommended for production), and `s2.1-pro-free` (free tier), but the SDK type definitions currently list only `s1` and `s2-pro` (`s2-pro` = SDK default). Both SDKs forward the model value without runtime validation, so `"s2.1-pro"` works over the wire. Static type checkers will flag it, so add `# type: ignore` (Python) / an `as` cast (TS), or use the `fish-audio-api` skill for raw calls. `speech-1.5` / `speech-1.6` are **deprecated**. In Python pass `model="s2-pro"` (keyword); in JS pass the **positional** `backend` argument.
- **Audio formats:** `mp3` (default), `wav`, `pcm`, `opus`.
- **Playback in examples:** `play()` shells out to a system audio tool: Python uses **ffmpeg/ffplay** (or `mpv`), JS uses **ffplay**. It is for local/desktop use; in a server, `save()` to a file or stream the bytes instead. See [references/installation.md](references/installation.md).
## Quick start: Python
```python
from fishaudio import FishAudio
from fishaudio.utils import play, save
client = FishAudio() # reads FISH_API_KEY
# Generate speech (returns the full audio as bytes)
audio = client.tts.convert(text="Hello from Fish Audio!")
save(audio, "output.mp3") # write to a file
# play(audio) # or play locally (needs ffmpeg)
```
Async: identical resource tree on `AsyncFishAudio`, used as a context manager:
```python
import asyncio
from fishaudio import AsyncFishAudio
from fishaudio.utils import save
async def main():
async with AsyncFishAudio() as client:
audio = await client.tts.convert(text="Hello from Fish Audio!")
save(audio, "output.mp3")
asyncio.run(main())
```
## Quick start: JavaScript / TypeScript
```ts
import { FishAudioClient, play } from "fish-audio";
const client = new FishAudioClient({ apiKey: process.env.FISH_API_KEY });
// convert() returns audio you can play or pipe to a file
const audio = await client.textToSpeech.convert({
text: "Hello from Fish Audio!",
}); // defaults to model "s2-pro"
await play(audio); // local playback (needs ffplay)
```
To pick a model in JS, pass `backend` as the **positional** argument (not a named option):
```ts
const audio = await client.textToSpeech.convert({ text: "Hi" }, "s1");
```
## Capabilities → referenc