skillZs
★ LIVE SKILL TAGS ★
>>> LIVE SKILLS INDEX <<<
* OPEN SOURCE *
NO LOGIN, NO TRACKING
※ REAL INSTALL DATA ※
← back to all skills
docs.fish.audio1k installs

fish-audio-sdk

Write code with the official Fish Audio SDKs, Python (`fishaudio`, PyPI `fish-audio-sdk`) and JavaScript/TypeScript (`fish-audio`). Use when the user wants text-to-speech, speech-to-text, voice cloning / voice-model management, or realtime WebSocket TTS through the installed SDK rather than raw HTTP. Covers install and auth, sync + async Python, the TypeScript client, exact method signatures and defaults, model selection (including the S2.1 typing caveat), the real exception types, and the Python↔JavaScript naming differences. For raw REST/WebSocket calls without an SDK (curl, unsupported languages, edge runtimes), use the `fish-audio-api` skill instead.

How do I install this agent skill?

npx skills add https://docs.fish.audio --skill fish-audio-sdk
view source ↗

Is this agent skill safe to install?

No partner audit is available yet. Read the source before installing.

What does this agent skill do?

Fish Audio SDK Skill

Use this skill to generate correct, runnable code with the official Fish Audio SDKs:

  • Python: package fish-audio-sdk on PyPI, imported as fishaudio. (The same wheel still ships a separate legacy fish_audio_sdk package. Do not mix them; everything here is the modern fishaudio package.)
  • JavaScript / TypeScript: package fish-audio on npm, imported as FishAudioClient.

If the user wants raw curl / HTTP / WebSocket without installing an SDK, use the fish-audio-api skill instead.

This file is the index. Deeper, task-specific rules and full examples live in references/. Read the reference for the task you're doing before writing code.

Global facts

  • Auth: both SDKs read the API key from the FISH_API_KEY environment variable automatically. Get keys at https://fish.audio/app/api-keys. Never hardcode a key.
  • Base URL: https://api.fish.audio (override with base_url= in Python / baseUrl: in JS).
  • TTS models: the API supports s1, s2-pro, s2.1-pro (recommended for production), and s2.1-pro-free (free tier), but the SDK type definitions currently list only s1 and s2-pro (s2-pro = SDK default). Both SDKs forward the model value without runtime validation, so "s2.1-pro" works over the wire. Static type checkers will flag it, so add # type: ignore (Python) / an as cast (TS), or use the fish-audio-api skill for raw calls. speech-1.5 / speech-1.6 are deprecated. In Python pass model="s2-pro" (keyword); in JS pass the positional backend argument.
  • ASR models: use transcribe-1-pro (recommended: speaker turns, long recordings, emotion cues). Neither SDK has an ASR model argument: send the model: transcribe-1-pro HTTP header on every request (Python RequestOptions(additional_headers=...), JS requestOptions.headers). A request without it is served and billed as transcribe-1, the model for short recordings. See references/speech-to-text.md.
  • Audio formats: mp3 (default), wav, pcm, opus.
  • Playback in examples: play() shells out to a system audio tool: Python uses ffmpeg/ffplay (or mpv), JS uses ffplay. It is for local/desktop use; in a server, save() to a file or stream the bytes instead. See references/installation.md.

Quick start: Python

from fishaudio import FishAudio
from fishaudio.utils import play, save

client = FishAudio()  # reads FISH_API_KEY

# Generate speech (returns the full audio as bytes)
audio = client.tts.convert(text="Hello from Fish Audio!")

save(audio, "output.mp3")   # write to a file
# play(audio)               # or play locally (needs ffmpeg)

Async: identical resource tree on AsyncFishAudio, used as a context manager:

import asyncio
from fishaudio import AsyncFishAudio
from fishaudio.utils import save

async def main():
    async with AsyncFishAudio() as client:
        audio = await client.tts.convert(text="Hello from Fish Audio!")
        save(audio, "output.mp3")

asyncio.run(main())

Quick start: JavaScript / TypeScript

import { FishAudioClient, play } from "fish-audio";

const client = new FishAudioClient({ apiKey: process.env.FISH_API_KEY });

// convert() returns audio you can play or pipe to a file
const audio = await client.textToSpeech.convert({
  text: "Hello from Fish Audio!",
}); // defaults to model "s2-pro"
await play(audio); // local playback (needs ffplay)

To pick a model in JS, pass backend as the positional argument (not a named option):

const audio = await client.textToSpeech.convert({ text: "Hi" }, "s1");

Capabilities → references

TaskReference
Install, auth, playback deps, verify a keyreferences/installation.md
Text-to-Speech (convert, stream, formats, prosody, model select)references/text-to-speech.md
Voice cloning (instant references + persistent voice models)references/voice-cloning.md
Speech-to-Text (models, timestamps, fields the SDK drops)references/speech-to-text.md
Realtime WebSocket TTS (stream text → audio)references/websocket.md
Errors, retries, and timeouts (the real exception types)references/errors.md

Python ↔ JavaScript name map

The two SDKs do not use the same names. Use this map when porting code between them.

ConceptPython (fishaudio)JavaScript (fish-audio)
ClientFishAudio() / AsyncFishAudio()new FishAudioClient({ apiKey })
Text-to-Speechclient.tts.convert(text=...) → bytesclient.textToSpeech.convert({ text })
TTS HTTP streamclient.tts.stream(...) → AudioStream(use convert; realtime streaming is convertRealtime)
Realtime WebSocketclient.tts.stream_websocket(text_stream)client.textToSpeech.convertRealtime(request, textStream)
Speech-to-Textclient.asr.transcribe(audio=...)client.speechToText.convert({ audio })
List voice modelsclient.voices.list()client.voices.search()
Get voice modelclient.voices.get(id)client.voices.get(id)
Create voice (clone)client.voices.create(title=..., voices=[bytes])client.voices.ivc.create({ title, voices: [File] })
Update / delete voiceclient.voices.update(id, ...) / delete(id)client.voices.update(id, ...) / delete(id)
Credit balanceclient.account.get_credits()client.user.get_api_credit()
Subscription packageclient.account.get_package()client.user.get_package()
Choose modelmodel="s2-pro" keyword argpositional backend arg, e.g. convert(req, "s2-pro")
ASR model (header)RequestOptions(additional_headers=...)convert(req, { headers: { model: "transcribe-1-pro" } })

Decision shortcuts

  • Audio from text → tts.convert (Python) / textToSpeech.convert (JS).
  • Reuse a saved voice → pass reference_id (the voice model id).
  • Clone a voice instantly from a clip → pass references=[ReferenceAudio(audio=..., text=...)] (Python) / references: [{ audio, text }] (JS). See voice-cloning.
  • Persistent custom voice to reuse → create a voice model, then use its id as reference_id.
  • Stream tokens from an LLM and play speech as it arrives → tts.stream_websocket (Python) / textToSpeech.convertRealtime (JS). See websocket.
  • Transcribe audio → asr.transcribe (Python) / speechToText.convert (JS) with the model: transcribe-1-pro header (recommended; without it, the request runs on transcribe-1). Pro request fields, speaker_turns, request_id, and the language fields need raw HTTP in Python (fish-audio-sdk 1.3.0). See speech-to-text.

Gotchas (verified against the SDK source)

  • Python latency accepts only "normal" or "balanced" (default "balanced"); there is no "low".
  • The Python client has no max_retries and does not auto-retry; the JS client does auto-retry (configurable via per-call requestOptions.maxRetries). See errors.
  • Python defines a ValidationError class but never raises it, so don't catch it expecting validation failures; a 422 surfaces as APIError. The JS SDK throws UnprocessableEntityError on 422.
  • ASR segment start / end and duration are all in seconds. The Python SDK's ASRResponse docstring says milliseconds; that is wrong. See speech-to-text.

Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.

<a href="https://skillzs.dev/skills/docs.fish.audio/fish-audio-sdk">View fish-audio-sdk on skillZs</a>