Model
Speech endpoint (POST)
Need a model that listens and responds across turns? Use Higgs Realtime. Higgs TTS 3 renders speech from text; it does not manage a conversation.
Features
- Chat-native, low-latency streaming — begin speaking before the full input is finalized.
- 100 languages — single-digit WER/CER coverage. See Languages.
- Instant voice cloning — zero-shot from a short reference clip and its transcript. See Voices.
- Inline control tags — shape emotion, style, prosody, and sound effects with
<|emotion:…|>,<|style:…|>,<|prosody:…|>, and<|sfx:…|>. See Tags.
Try it in the playground
The fastest way to hear the model is the playground. Pick a voice, paste text, and press play.Generate speech with the API
You need a Boson API key stored inBOSON_API_KEY, and available credit — new accounts must claim their free trial credit before their first API call. Set the key in your shell for the current session:
Authorization, model, and input. Everything else is optional.
Use a preset voice
Usevoice to choose a preset speaker.
Use reference audio
Useref_audio to clone a voice from a short reference clip. Passing the audio transcript through ref_text can often improve generated audio quality.
Fine-grained control
Inline tags control emotion, style, prosody, and sound effects in the generated audio. Add them toinput, and the model adjusts the surrounding speech. For example:
Text normalization
Normalization rewrites the input so the model speaks it the way a person would read it aloud, expanding numbers, dates, times, currency, units, URLs, email addresses, phone numbers, abbreviations, and symbols into spoken form. It is on by default and leaves inline tags and ordinary prose untouched. Setenable_tn: false to synthesize the text exactly as written.
tn_language to an ISO 639-1 code. Normalization does not run when word-level timestamps are requested.
Word-level timestamps
Settimestamps: true to get the start and end time of every word alongside the audio.
response_format still selects the audio format as usual; the encoded audio is returned in that format inside the envelope’s audio field.
timestamps is null when alignment is unavailable, for example on very short inputs.
Word timestamps are currently available for English, Chinese, and Spanish. Other languages return
"timestamps": null.stream must be false. Setting both returns a 400 timestamps_streaming_unsupported.
Streaming response
Whenstream: true, set response_format: "pcm".
Limits
We recommend keepinginput to roughly 300 characters. Longer text is not yet reliably supported, and may come back truncated or with garbled words as a normal 200 response rather than an error. Support for long inputs is in progress; in the meantime, split long scripts on sentence boundaries and synthesize each chunk separately.
API reference
Full request body:Alternative ways to use the model
Beyond the hosted API, you can run the model yourself:- Hugging Face — open model weights at bosonai/higgs-tts-v3-4b.
- SGLang — serve the model locally for high-throughput inference. See the Higgs TTS cookbook.