Fish Audio Docs
API ReferenceText to SpeechSynchronous Text to Speech

Synchronous Text to Speech v3

Generate audio synchronously or create recoverable asynchronous jobs with the provider-neutral HTTP v3 contract.

HTTP TTS v3 is the recommended contract for new integrations. The platform owns provider routing; clients use a public voiceId, an optional modelId, and generic controls.

If you omit modelId, the platform selects the default engine for that voice's provider. If you send modelId, the voice's modelIds must contain it, because supported controls, text limits, and output formats vary by model.

Base URL

https://fishaudio.org/api/open/v3

Every endpoint except capabilities and the OpenAPI document requires a server-side Bearer API key. Never embed the key in browser code, mobile binaries, or public repositories.

Choose synchronous or asynchronous delivery

Use caseEndpointWhy
Short text and the caller can keep the HTTP request openPOST /speech/ttsReturns audio bytes directly
Long text, batches, or durable result lookupPOST /speech/tts/jobsReturns a job ID that can be polled and downloaded again
The output must be recoverable after a lost responseAsynchronous JobsThe synchronous endpoint does not replay completed audio

Step 1: discover model capabilities

Capabilities is public. Read it before constructing a generation request instead of hard-coding model behavior.

curl "https://fishaudio.org/api/open/v3/speech/tts/capabilities"
{
  "contractVersion": "v3",
  "route": "/api/open/v3/speech/tts",
  "models": [
    {
      "id": "fishaudio-s21pro-flash",
      "available": true,
      "maxTextLength": 10000,
      "outputFormats": ["mp3", "wav", "ogg"],
      "controls": {
        "speed": true,
        "volume": true,
        "stability": true,
        "similarity": true
      },
      "supportedOperations": {
        "synchronous": true,
        "asynchronous": true
      }
    }
  ]
}

Choose only a model with available=true. Treat its maxTextLength, outputFormats, and controls as authoritative.

Step 2: find a compatible voice

curl "https://fishaudio.org/api/open/v3/voices?page=1&pageSize=20&includePersonal=false" \
  -H "Authorization: Bearer $FISHAUDIO_API_KEY"
{
  "total": 1,
  "page": 1,
  "pageSize": 20,
  "totalPages": 1,
  "items": [
    {
      "voiceId": "00a1b221-6137-4b73-ad62-b0cbce134167",
      "name": "System test voice",
      "isPersonal": false,
      "primaryLanguage": "en",
      "languages": ["en"],
      "modelIds": ["fishaudio-s21pro-flash"]
    }
  ],
  "requestId": "req_xxx"
}

modelIds is the list of public models currently supported by that voice. Generation requires a combination where the model is available and the voice supports it. Use GET /voices/{voiceId} to inspect one voice.

Step 3: generate synchronously

POST /api/open/v3/speech/tts
curl "https://fishaudio.org/api/open/v3/speech/tts" \
  -H "Authorization: Bearer $FISHAUDIO_API_KEY" \
  -H "Content-Type: application/json" \
  -H "X-Request-Id: trace-123-segment-1" \
  -H "Idempotency-Key: tts-order-123-segment-1" \
  -d '{
    "text": "Hello from HTTP TTS v3.",
    "voiceId": "00a1b221-6137-4b73-ad62-b0cbce134167",
    "modelId": "fishaudio-s21pro-flash",
    "format": "mp3",
    "speed": 1,
    "volume": 0,
    "stability": 1,
    "similarity": 1,
    "language": "en",
    "textNormalization": true
  }' \
  --output speech.mp3

Request fields

FieldTypeRequiredConstraints and defaults
textstringYes1–10,000 characters and no more than the model's maxTextLength
voiceIdstringYesReturned by /voices; if modelId is sent, the voice must support it
modelIdstringNoPublic model ID from capabilities; omitted uses the voice provider default
formatstringNomp3, wav, or ogg; default mp3; must be in the model's outputFormats
speednumberNo0.5–2; default 1
volumenumberNo-20–20; default 0
pitchnumberNo-12–12; send only when capabilities declares support
stabilitynumberNo0.5–1.5; send only when capabilities declares support
similaritynumberNo0.5–1.5; send only when capabilities declares support
languagestringNoLanguage hint, 1–64 characters
emotionstringNoGlobal emotion ID, 1–64 characters
instructionstringNoStyle, dialect, role, or emotion instruction, up to 1,600 characters
textNormalizationbooleanNoEnables structured-text pronunciation normalization

The request contract is strict. provider, provider API keys, provider-native model fields, and other unknown properties return 400. Valid generic controls unsupported by the selected model are ignored and listed in X-OpenAPI-Ignored-Parameters.

filterEmoji (experimental, off by default)

Synchronous and asynchronous creation requests can include:

{ "filterEmoji": true }

Only boolean true enables it. Omission or false preserves existing behavior; strings such as "false" are rejected. filterEmoji is a top-level boolean, not nested under experimental. Despite its general name, it currently only filters kaomoji, not graphical emoji such as 😀. This is independent of textNormalization. Only selected complete decorative kaomoji are removed; emoji, ambiguous ASCII faces, code, URLs, quoted literals, TTS tags and explanatory contexts are preserved conservatively. Coverage is incomplete and zero false positives are not guaranteed.

Billing uses filtered text; input limits are checked before filtering. A whitespace-only result returns 400 ERR_INVALID_REQUEST before job creation or charging. Async jobs store filtered synthesis text; historical records are not rewritten. Use a new idempotency key when changing the flag or original input. Legacy HTTP synchronous/asynchronous TTS endpoints also support this flag; Studio, mobile and WebSocket do not.

Successful response

HTTP/1.1 200 OK
Content-Type: audio/mpeg
X-Request-Id: trace-123-segment-1
X-OpenAPI-Quota-Remaining: 99994
X-OpenAPI-Credits-Used: 6
X-OpenAPI-Ignored-Parameters: pitch

<binary audio data>

Check the HTTP status first, then branch on Content-Type. A successful response contains audio bytes; an error contains JSON. X-OpenAPI-Ignored-Parameters is present only when a control was ignored.

Streaming audio with word timestamps

For a Fish Audio model and voice, POST /speech/tts/aligned accepts the same JSON body as synchronous v3 TTS and returns text/event-stream rather than an MP3 file. Query /speech/tts/capabilities first and select a model with supportedOperations.alignedStreaming: true.

curl -N "https://fishaudio.org/api/open/v3/speech/tts/aligned" \
  -H "Authorization: Bearer $FISHAUDIO_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: aligned-example-001" \
  -d '{"text":"Hello, welcome to captions.","voiceId":"00a1b221-6137-4b73-ad62-b0cbce134167","modelId":"fishaudio-s21pro-flash","format":"mp3"}'

Process the SSE events in order:

EventDataClient action
audioaudioBase64Decode and concatenate audio bytes in arrival order. Do not keep only the last packet.
alignment_snapshotchunkSeq, units: "milliseconds", words: [{text,startMs,endMs}]A newer snapshot for the same chunkSeq replaces the old one; times are absolute milliseconds from the start of the full audio.
alignment_invalidatedchunkSeqDiscard prior timing for that chunk; its audio remains valid.
completedrequestId, generationId, alignmentStatus, quotaRemaining, creditsUsedGeneration and billing settlement finished. Treat word timing as complete only when alignmentStatus is complete.
errorrequestId, codeThe stream failed after it started; do not treat partial audio as a complete result.

Pre-stream HTTP errors are JSON. After HTTP 200, keep reading until completed or error. A disconnected stream cannot be resumed; use asynchronous Jobs when you need durable retries and downloads. Native word timing is not guaranteed for every generation; if alignmentStatus is partial, your application can fall back to forced alignment.

Asynchronous jobs

Use Jobs for long text, batches, or workflows that require recoverable output. The create body is the same as synchronous v3.

curl "https://fishaudio.org/api/open/v3/speech/tts/jobs" \
  -H "Authorization: Bearer $FISHAUDIO_API_KEY" \
  -H "Content-Type: application/json" \
  -H "X-Request-Id: trace-job-001" \
  -H "Idempotency-Key: tts-job-001" \
  -d '{
    "text": "Text that needs durable generation and download.",
    "voiceId": "00a1b221-6137-4b73-ad62-b0cbce134167",
    "modelId": "fishaudio-s21pro-flash",
    "format": "mp3"
  }'

Creation returns 202 Accepted. Location points to the job status resource and Retry-After provides the suggested polling delay. Store task.taskId immediately:

{
  "task": {
    "taskId": "task_123",
    "status": "pending",
    "voiceId": "00a1b221-6137-4b73-ad62-b0cbce134167",
    "modelId": "fishaudio-s21pro-flash"
  },
  "requestId": "trace-job-001",
  "quotaRemaining": 99994,
  "creditsUsed": 6,
  "ignoredParameters": []
}
GET /speech/tts/jobs/{jobId}
GET /speech/tts/jobs?page=1&limit=20&status=success
GET /speech/tts/jobs/{jobId}/audio?download=1
GET /speech/tts/jobs/{jobId}/segments/{segmentIndex}
Authorization: Bearer FISHAUDIO_API_KEY

pending and processing are non-terminal. Only success, partial_fail, and fail are terminal. Keep Bearer authentication when downloading audio and allow the client to follow redirects.

Error format and handling

{
  "code": "ERR_REQUEST_ID_CONFLICT",
  "message": "Request id was reused with a different request",
  "requestId": "tts-order-123-segment-1"
}
StatusCommon causeClient action
400Invalid JSON, field, model, voice, or formatFix the request; do not retry unchanged
401Missing or invalid API keyReplace the server-side credential
402Insufficient API quotaStop the workflow and surface a billing state
404Voice or asynchronous job not foundCheck the ID and owning account
409Request active, completed, refunded, or reused with different inputBranch on code; do not blindly use a new ID
413HTTP request body exceeds the limitReduce the request; use Jobs for long text
429Rate limit exceededHonor Retry-After and use exponential backoff
500Platform or provider generation failedPreserve requestId; confirm settlement before retrying
503The selected model is currently unavailableRefresh capabilities and select an available model

Idempotency, billing, and retries

X-Request-Id is a trace id and is echoed in the response; the server generates it when omitted. Idempotency-Key is the business idempotency key and is at most 128 characters. Create one stable idempotency key per logical generation and reuse it only for an identical retry. Never send concurrent requests with the same key.

Existing clients may temporarily continue using X-Request-Id as a compatibility idempotency key. New clients must send the two headers separately. When Idempotency-Key is omitted, the request still runs, but the client cannot rely on an idempotent retry for that call.

  • Synchronous generation has at-most-once semantics. Reusing a completed idempotency key returns 409 ERR_REQUEST_ALREADY_COMPLETED; it does not replay audio and does not charge again.
  • Asynchronous creation replays the accepted task for an identical request, so it remains queryable and is not billed twice.
  • Reusing an idempotency key with different parameters returns 409 ERR_REQUEST_ID_CONFLICT.
  • A client timeout does not mean the server stopped. Use Jobs whenever output must be recovered after a lost response.
  • Validation, authentication, and quota failures are not charged. Generation or settlement failures enter the refund path.

Machine-readable contract and migration

The complete OpenAPI 3.1 document is public:

GET /api/open/v3/openapi.json

OpenAPI defines stable fields and response shapes; capabilities describes live model availability. Read both together. Existing v1 and v2 clients remain compatible, while new integrations should use v3. See the migration guide for field and path changes.