Providers and models
Vida supports speech-to-speech through OpenAI GPT Live, Google Gemini Live, and the older OpenAI Realtime integration. Available engines and compatible voices can vary by account. GPT Live uses a separately selected delegation model for reasoning and tool decisions while handling the real-time audio conversation itself.When to use it
Speech-to-speech is a good fit when natural turn-taking, low latency, and expressive delivery matter more than selecting separate speech-recognition and voice providers. Standard mode is a better fit when you need the broader set of speech-recognition controls, voice providers, or tuning options.Configure and test
- Open the Agent editor and select Voice.
- Choose speech-to-speech mode and an available provider.
- Select a compatible voice.
- Test realistic calls, including interruptions, names, numbers, and background noise.
- Publish only after the staged Agent behaves as expected.
Configuring speech mode through the API? See Language, voice, and speech.