Model
WebSocket endpoint
Features
- Full-duplex voice — one persistent session streams audio and text in both directions; the model listens and speaks natively.
- Turn detection and barge-in —
server_vadandsemantic_vaddetect end of turn and reply automatically; interruptions always stop the model mid-utterance. See Turn detection. - Tool calling —
functiontools with JSON-schema parameters, called mid-conversation. See Tool use. - Voices — the
defaultvoice, Boson AI presets, or your own clonedvoice_<id>. See Audio & voices. - Input transcription — live user-speech transcripts via
higgs-stt-3.1. - Text mode — set
output_modalitiesto["text"]to stream text instead of spoken audio. - OpenAI-compatible protocol — the event protocol is compatible with OpenAI’s GA Realtime API; most integrations move with config changes. See Migrate an existing integration.
- Browser and mobile clients — trusted servers mint short-lived client secrets so keys never ship to devices.
Start building
Quickstart
Every step from a fresh account to a working voice assistant.
API reference
Every client and server event, close codes, and session fields.
LiveKit & Pipecat integrations
Production voice-agent frameworks with Higgs Realtime built in.
Migrate from OpenAI
Move an existing OpenAI Realtime integration over.
Higgs Realtime API — a tutorial
Six parts in TypeScript and React, from zero to a working browser voice assistant — with a git checkpoint after every part.