Key capabilities
- Realtime speech - Text and audio input and output
- Configurable reasoning - Adjust reasoning effort for voice workflows
- Tool use - Supports function calls in realtime sessions
- Multimodal input - Accepts text, audio, and image input
Quick example
Node.js WebSocket
Install thews package, then connect with the model query parameter and a Bearer token.