Skip to main content
Model
Speech endpoint (POST)
Need a model that listens and responds across turns? Use Higgs Realtime. Higgs TTS 3 renders speech from text; it does not manage a conversation.

Features

  • Chat-native, low-latency streaming — begin speaking before the full input is finalized.
  • 100 languages — single-digit WER/CER coverage. See Languages.
  • Instant voice cloning — zero-shot from a short reference clip and its transcript. See Voices.
  • Inline control tags — shape emotion, style, prosody, and sound effects with <|emotion:…|>, <|style:…|>, <|prosody:…|>, and <|sfx:…|>. See Tags.

Try it in the playground

The fastest way to hear the model is the playground. Pick a voice, paste text, and press play.

Generate speech with the API

You need a Boson API key stored in BOSON_API_KEY, and available credit — new accounts must claim their free trial credit before their first API call. Set the key in your shell for the current session:
A minimal request needs Authorization, model, and input. Everything else is optional.

Use a preset voice

Use voice to choose a preset speaker.
See Voices for more preset speakers and samples.

Use reference audio

Use ref_audio to clone a voice from a short reference clip. Passing the audio transcript through ref_text can often improve generated audio quality.
To clone from a local file, either encode local file as base64 string or send as `multipart/form-data. Below code shows the latter.
You must own the right to clone the voice.
See Voices for best practices and reusable custom voices.

Fine-grained control

Inline tags control emotion, style, prosody, and sound effects in the generated audio. Add them to input, and the model adjusts the surrounding speech. For example:
See Tags for the complete list and sample audio.

Text normalization

Normalization rewrites the input so the model speaks it the way a person would read it aloud, expanding numbers, dates, times, currency, units, URLs, email addresses, phone numbers, abbreviations, and symbols into spoken form. It is on by default and leaves inline tags and ordinary prose untouched. Set enable_tn: false to synthesize the text exactly as written.
Normalization auto-detects the input language for Chinese, English, French, German, Hungarian, Italian, Japanese, Korean, Portuguese, Russian, Spanish, Swedish, Thai, and Vietnamese. For any other language, and for short or mixed-language text where detection is least reliable, set tn_language to an ISO 639-1 code. Normalization does not run when word-level timestamps are requested.

Word-level timestamps

Set timestamps: true to get the start and end time of every word alongside the audio.
timestamps: true changes the response from audio bytes to JSON, with the audio base64-encoded inside it. Writing the response body straight to a file produces an unplayable file. Decode audio first.
response_format still selects the audio format as usual; the encoded audio is returned in that format inside the envelope’s audio field.
timestamps is null when alignment is unavailable, for example on very short inputs.
Word timestamps are currently available for English, Chinese, and Spanish. Other languages return "timestamps": null.
Timestamps require the full clip, so stream must be false. Setting both returns a 400 timestamps_streaming_unsupported.

Streaming response

When stream: true, set response_format: "pcm".

Limits

We recommend keeping input to roughly 300 characters. Longer text is not yet reliably supported, and may come back truncated or with garbled words as a normal 200 response rather than an error. Support for long inputs is in progress; in the meantime, split long scripts on sentence boundaries and synthesize each chunk separately.

API reference

Full request body:
See the API reference for field details and additional options.

Alternative ways to use the model

Beyond the hosted API, you can run the model yourself: