fish-audio-sdk
Write code with the official Fish Audio SDKs, Python (`fishaudio`, PyPI `fish-audio-sdk`) and JavaScript/TypeScript (`fish-audio`). Use when the user wants text-to-speech, speech-to-text, voice cloning / voice-model management, or realtime WebSocket TTS through the installed SDK rather than raw HTTP. Covers install and auth, sync + async Python, the TypeScript client, exact method signatures and defaults, model selection (including the S2.1 typing caveat), the real exception types, and the Python↔JavaScript naming differences. For raw REST/WebSocket calls without an SDK (curl, unsupported languages, edge runtimes), use the `fish-audio-api` skill instead.
How do I install this agent skill?
npx skills add https://docs.fish.audio --skill fish-audio-sdkIs this agent skill safe to install?
No partner audit is available yet. Read the source before installing.
What does this agent skill do?
Fish Audio SDK Skill
Use this skill to generate correct, runnable code with the official Fish Audio SDKs:
- Python: package
fish-audio-sdkon PyPI, imported asfishaudio. (The same wheel still ships a separate legacyfish_audio_sdkpackage. Do not mix them; everything here is the modernfishaudiopackage.) - JavaScript / TypeScript: package
fish-audioon npm, imported asFishAudioClient.
If the user wants raw curl / HTTP / WebSocket without installing an SDK, use the fish-audio-api skill instead.
This file is the index. Deeper, task-specific rules and full examples live in
references/. Read the reference for the task you're doing before writing code.
Global facts
- Auth: both SDKs read the API key from the
FISH_API_KEYenvironment variable automatically. Get keys athttps://fish.audio/app/api-keys. Never hardcode a key. - Base URL:
https://api.fish.audio(override withbase_url=in Python /baseUrl:in JS). - TTS models: the API supports
s1,s2-pro,s2.1-pro(recommended for production), ands2.1-pro-free(free tier), but the SDK type definitions currently list onlys1ands2-pro(s2-pro= SDK default). Both SDKs forward the model value without runtime validation, so"s2.1-pro"works over the wire. Static type checkers will flag it, so add# type: ignore(Python) / anascast (TS), or use thefish-audio-apiskill for raw calls.speech-1.5/speech-1.6are deprecated. In Python passmodel="s2-pro"(keyword); in JS pass the positionalbackendargument. - ASR models: use
transcribe-1-pro(recommended: speaker turns, long recordings, emotion cues). Neither SDK has an ASRmodelargument: send themodel: transcribe-1-proHTTP header on every request (PythonRequestOptions(additional_headers=...), JSrequestOptions.headers). A request without it is served and billed astranscribe-1, the model for short recordings. See references/speech-to-text.md. - Audio formats:
mp3(default),wav,pcm,opus. - Playback in examples:
play()shells out to a system audio tool: Python uses ffmpeg/ffplay (ormpv), JS uses ffplay. It is for local/desktop use; in a server,save()to a file or stream the bytes instead. See references/installation.md.
Quick start: Python
from fishaudio import FishAudio
from fishaudio.utils import play, save
client = FishAudio() # reads FISH_API_KEY
# Generate speech (returns the full audio as bytes)
audio = client.tts.convert(text="Hello from Fish Audio!")
save(audio, "output.mp3") # write to a file
# play(audio) # or play locally (needs ffmpeg)
Async: identical resource tree on AsyncFishAudio, used as a context manager:
import asyncio
from fishaudio import AsyncFishAudio
from fishaudio.utils import save
async def main():
async with AsyncFishAudio() as client:
audio = await client.tts.convert(text="Hello from Fish Audio!")
save(audio, "output.mp3")
asyncio.run(main())
Quick start: JavaScript / TypeScript
import { FishAudioClient, play } from "fish-audio";
const client = new FishAudioClient({ apiKey: process.env.FISH_API_KEY });
// convert() returns audio you can play or pipe to a file
const audio = await client.textToSpeech.convert({
text: "Hello from Fish Audio!",
}); // defaults to model "s2-pro"
await play(audio); // local playback (needs ffplay)
To pick a model in JS, pass backend as the positional argument (not a named option):
const audio = await client.textToSpeech.convert({ text: "Hi" }, "s1");
Capabilities → references
| Task | Reference |
|---|---|
| Install, auth, playback deps, verify a key | references/installation.md |
| Text-to-Speech (convert, stream, formats, prosody, model select) | references/text-to-speech.md |
| Voice cloning (instant references + persistent voice models) | references/voice-cloning.md |
| Speech-to-Text (models, timestamps, fields the SDK drops) | references/speech-to-text.md |
| Realtime WebSocket TTS (stream text → audio) | references/websocket.md |
| Errors, retries, and timeouts (the real exception types) | references/errors.md |
Python ↔ JavaScript name map
The two SDKs do not use the same names. Use this map when porting code between them.
| Concept | Python (fishaudio) | JavaScript (fish-audio) |
|---|---|---|
| Client | FishAudio() / AsyncFishAudio() | new FishAudioClient({ apiKey }) |
| Text-to-Speech | client.tts.convert(text=...) → bytes | client.textToSpeech.convert({ text }) |
| TTS HTTP stream | client.tts.stream(...) → AudioStream | (use convert; realtime streaming is convertRealtime) |
| Realtime WebSocket | client.tts.stream_websocket(text_stream) | client.textToSpeech.convertRealtime(request, textStream) |
| Speech-to-Text | client.asr.transcribe(audio=...) | client.speechToText.convert({ audio }) |
| List voice models | client.voices.list() | client.voices.search() |
| Get voice model | client.voices.get(id) | client.voices.get(id) |
| Create voice (clone) | client.voices.create(title=..., voices=[bytes]) | client.voices.ivc.create({ title, voices: [File] }) |
| Update / delete voice | client.voices.update(id, ...) / delete(id) | client.voices.update(id, ...) / delete(id) |
| Credit balance | client.account.get_credits() | client.user.get_api_credit() |
| Subscription package | client.account.get_package() | client.user.get_package() |
| Choose model | model="s2-pro" keyword arg | positional backend arg, e.g. convert(req, "s2-pro") |
| ASR model (header) | RequestOptions(additional_headers=...) | convert(req, { headers: { model: "transcribe-1-pro" } }) |
Decision shortcuts
- Audio from text →
tts.convert(Python) /textToSpeech.convert(JS). - Reuse a saved voice → pass
reference_id(the voice modelid). - Clone a voice instantly from a clip → pass
references=[ReferenceAudio(audio=..., text=...)](Python) /references: [{ audio, text }](JS). See voice-cloning. - Persistent custom voice to reuse → create a voice model, then use its
idasreference_id. - Stream tokens from an LLM and play speech as it arrives →
tts.stream_websocket(Python) /textToSpeech.convertRealtime(JS). See websocket. - Transcribe audio →
asr.transcribe(Python) /speechToText.convert(JS) with themodel: transcribe-1-proheader (recommended; without it, the request runs ontranscribe-1). Pro request fields,speaker_turns,request_id, and the language fields need raw HTTP in Python (fish-audio-sdk1.3.0). See speech-to-text.
Gotchas (verified against the SDK source)
- Python
latencyaccepts only"normal"or"balanced"(default"balanced"); there is no"low". - The Python client has no
max_retriesand does not auto-retry; the JS client does auto-retry (configurable via per-callrequestOptions.maxRetries). See errors. - Python defines a
ValidationErrorclass but never raises it, so don't catch it expecting validation failures; a 422 surfaces asAPIError. The JS SDK throwsUnprocessableEntityErroron 422. - ASR segment
start/endanddurationare all in seconds. The Python SDK'sASRResponsedocstring says milliseconds; that is wrong. See speech-to-text.
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/docs.fish.audio/fish-audio-sdk">View fish-audio-sdk on skillZs</a>