Skip to main content
The Realtime API enables real-time bidirectional voice communication over WebSocket (WebRTC support is in progress). Stream audio and text both ways for voice assistants, phone agents, and interactive voice systems: the model listens and speaks natively in a single full-duplex session — no separate transcription or synthesis steps to manage.
Model
WebSocket endpoint

Features

  • Full-duplex voice — one persistent session streams audio and text in both directions; the model listens and speaks natively.
  • Turn detection and barge-in — server_vad and semantic_vad detect end of turn and reply automatically; interruptions always stop the model mid-utterance. See Turn detection.
  • Tool calling — function tools with JSON-schema parameters, called mid-conversation. See Tool use.
  • Voices — the default voice, Boson AI presets, or your own cloned voice_<id>. See Audio & voices.
  • Input transcription — live user-speech transcripts via higgs-stt-3.1.
  • Text mode — set output_modalities to ["text"] to stream text instead of spoken audio.
  • OpenAI-compatible protocol — the event protocol is compatible with OpenAI’s GA Realtime API; most integrations move with config changes. See Migrate an existing integration.
  • Browser and mobile clients — trusted servers mint short-lived client secrets so keys never ship to devices.

Start building

Quickstart

Every step from a fresh account to a working voice assistant.

API reference

Every client and server event, close codes, and session fields.

LiveKit & Pipecat integrations

Production voice-agent frameworks with Higgs Realtime built in.

Migrate from OpenAI

Move an existing OpenAI Realtime integration over.

Higgs Realtime API — a tutorial

Six parts in TypeScript and React, from zero to a working browser voice assistant — with a git checkpoint after every part.