GPT-Live
New guides for building conversational voice with GPT-Live:
- Getting started with GPT-Live — backends and first browser voice session
- Managing GPT-Live sessions — history, context, transcripts, mic controls, long conversations, and recovery
- Delegation and tools in GPT-Live — client or Responses delegation, context, and functions
- Prompting GPT-Live — style, backchannels, interruptions, and delegation
- Migrate to GPT-Live — paths from Realtime and existing text agents
- GPT-Live partner integrations — LiveKit, Twilio, Telnyx, and Daily/Pipecat
Docs now also list the gpt-live model ID.
Voice and Realtime
Voice transport and operations docs were reorganized under clearer voice-focused guides:
- WebRTC
- WebSockets
- Telephony and SIP
- Server-side controls
- Prompting Realtime models
- Cost optimization
These replace the older Realtime-specific pages:
- Realtime API with WebRTC
- Realtime API with WebSocket
- Realtime API with SIP
- Webhooks and server-side controls
- Using realtime models
- Managing costs
Related hub pages were retitled and reframed:
- Audio and voice (was “Audio and speech”) — start with GPT-Live or pick another audio API
- Getting started with the Realtime API (was “Realtime and audio”) — browser voice agent with the Realtime API and Agents SDK
- Voice agents — updated description for choosing a voice architecture and connecting spoken conversations to agent workflows
Audio and speech
- Audio in Chat Completions — add audio input and output to an existing Chat Completions app
- Custom voices — create an approved custom voice and use it for speech generation and voice agents