Skip to main content
POST
Create a speech

Authorizations

Authorization
string
header
required

Your Boson API key, sent as Authorization: Bearer $BOSON_API_KEY.

Body

input
string
required

Text to convert to speech. May contain inline tags. Inputs longer than 5000 characters return a 400 input_too_long. This is a validation limit; we recommend keeping each request to roughly 300 characters, as longer text is not yet reliably supported.

Required string length: 1 - 5000
Example:

"Hello, this is a test."

model
enum<string>
default:higgs-tts-3

TTS model ID / public alias. Resolved to the served model server-side.

Available options:
higgs-tts-3
voice
string
default:default

Preset voice name or custom voice ID. Mutually exclusive with ref_audio / ref_text when explicitly provided.

response_format
enum<string>
default:mp3

Output audio format. Streaming requires pcm.

Available options:
mp3,
opus,
pcm,
wav,
aac,
flac
stream
boolean
default:false

If true, stream raw PCM chunks as they are decoded. Requires response_format to be pcm. Speed adjustment is not supported when streaming.

ref_audio
string | null

Inline reference audio for one-off cloning: an http(s) URL, data URI, or base64-encoded raw audio bytes. Supported formats: AAC, WAV, MP3, FLAC, OPUS. Inline (base64 / data-URI) payloads: max 10 MB.

ref_text
string | null

Recommended transcript of ref_audio.

enable_tn
boolean
default:true

Expand non-standard words into spoken form before synthesis, including numbers, dates, times, currency, units, URLs, email addresses, phone numbers, abbreviations, and symbols. Enabled by default; set to false to synthesize the text exactly as written. Text with nothing to normalize is passed through unchanged. Inline tags are preserved.

tn_language
string | null

Normalization language as an ISO 639-1 code. Defaults to null, which auto-detects Chinese, English, French, German, Hungarian, Italian, Japanese, Korean, Portuguese, Russian, Spanish, Swedish, Thai, and Vietnamese. Set it explicitly for any other language, and for short or mixed-language text where detection is least reliable. An unsupported code returns 400 invalid_field_value, including when normalization is disabled.

timestamps
boolean
default:false

Attach word-level timestamps to the response. The response becomes a JSON envelope carrying the word list and the audio base64-encoded in the requested response_format, instead of raw audio bytes. Requires stream to be false; combining the two returns 400 timestamps_streaming_unsupported. Available for English, Chinese, and Spanish; other languages return a null word list. Overrides enable_tn, which does not run when timestamps are requested.

Response

Generated audio. The content type depends on response_format. When timestamps is true the response is application/json instead, carrying the audio base64-encoded alongside the word list.

The response is of type file.