Skill 詳細

fish-audio-sdk

公式Fish Audio SDKを使い、PythonとJavaScriptでテキスト読み上げ、音声認識、音声クローン、リアルタイムWebSocket TTSのコードを作成。

一致度一致の可能性オーディオと音声 向けにレビュー済み
出典docs.fish.audio外部ソース
報告インストール数873人気度の参考値

使用前に確認

自動レビューは関連性のみを確認し、安全性や推奨を保証しません。使用前に出典の説明を読んでください。

保存された出典プレビュー

SKILL.md

これはレビュー時に保存された抜粋です。完全で最新の内容は外部ソースを確認してください。

---
name: fish-audio-sdk
description: Write code with the official Fish Audio SDKs, Python (`fishaudio`, PyPI `fish-audio-sdk`) and JavaScript/TypeScript (`fish-audio`). Use when the user wants text-to-speech, speech-to-text, voice cloning / voice-model management, or realtime WebSocket TTS through the installed SDK rather than raw HTTP. Covers install and auth, sync + async Python, the TypeScript client, exact method signatures and defaults, model selection (including the S2.1 typing caveat), the real exception types, and the Python↔JavaScript naming differences. For raw REST/WebSocket calls without an SDK (curl, unsupported languages, edge runtimes), use the `fish-audio-api` skill instead.
---

# Fish Audio SDK Skill

Use this skill to generate correct, runnable code with the **official Fish Audio SDKs**:

- **Python**: package `fish-audio-sdk` on PyPI, imported as `fishaudio`. (The same wheel still ships a separate legacy `fish_audio_sdk` package. Do **not** mix them; everything here is the modern `fishaudio` package.)
- **JavaScript / TypeScript**: package `fish-audio` on npm, imported as `FishAudioClient`.

If the user wants raw `curl` / HTTP / WebSocket without installing an SDK, use the **`fish-audio-api`** skill instead.

> This file is the index. Deeper, task-specific rules and full examples live in [`references/`](references/). Read the reference for the task you're doing before writing code.

## Global facts

- **Auth:** both SDKs read the API key from the `FISH_API_KEY` environment variable automatically. Get keys at `https://fish.audio/app/api-keys`. Never hardcode a key.
- **Base URL:** `https://api.fish.audio` (override with `base_url=` in Python / `baseUrl:` in JS).
- **Models:** the API supports `s1`, `s2-pro`, `s2.1-pro` (recommended for production), and `s2.1-pro-free` (free tier), but the SDK type definitions currently list only `s1` and `s2-pro` (`s2-pro` = SDK default). Both SDKs forward the model value without runtime validation, so `"s2.1-pro"` works over the wire. Static type checkers will flag it, so add `# type: ignore` (Python) / an `as` cast (TS), or use the `fish-audio-api` skill for raw calls. `speech-1.5` / `speech-1.6` are **deprecated**. In Python pass `model="s2-pro"` (keyword); in JS pass the **positional** `backend` argument.
- **Audio formats:** `mp3` (default), `wav`, `pcm`, `opus`.
- **Playback in examples:** `play()` shells out to a system audio tool: Python uses **ffmpeg/ffplay** (or `mpv`), JS uses **ffplay**. It is for local/desktop use; in a server, `save()` to a file or stream the bytes instead. See [references/installation.md](references/installation.md).

## Quick start: Python

```python
from fishaudio import FishAudio
from fishaudio.utils import play, save

client = FishAudio()  # reads FISH_API_KEY

# Generate speech (returns the full audio as bytes)
audio = client.tts.convert(text="Hello from Fish Audio!")

save(audio, "output.mp3")   # write to a file
# play(audio)               # or play locally (needs ffmpeg)
```

Async: identical resource tree on `AsyncFishAudio`, used as a context manager:

```python
import asyncio
from fishaudio import AsyncFishAudio
from fishaudio.utils import save

async def main():
    async with AsyncFishAudio() as client:
        audio = await client.tts.convert(text="Hello from Fish Audio!")
        save(audio, "output.mp3")

asyncio.run(main())
```

## Quick start: JavaScript / TypeScript

```ts
import { FishAudioClient, play } from "fish-audio";

const client = new FishAudioClient({ apiKey: process.env.FISH_API_KEY });

// convert() returns audio you can play or pipe to a file
const audio = await client.textToSpeech.convert({
  text: "Hello from Fish Audio!",
}); // defaults to model "s2-pro"
await play(audio); // local playback (needs ffplay)
```

To pick a model in JS, pass `backend` as the **positional** argument (not a named option):

```ts
const audio = await client.textToSpeech.convert({ text: "Hi" }, "s1");
```

## Capabilities → referenc
関連情報

関連する仕事