Skip to main content
Vida supports standard voice processing and speech-to-speech modes. In speech-to-speech mode, one real-time model listens and responds with audio, which can make conversations feel faster and more expressive.

Providers and models

Vida supports speech-to-speech through OpenAI GPT Live, Google Gemini Live, and the older OpenAI Realtime integration. Available engines and compatible voices can vary by account. GPT Live uses a separately selected delegation model for reasoning and tool decisions while handling the real-time audio conversation itself.

When to use it

Speech-to-speech is a good fit when natural turn-taking, low latency, and expressive delivery matter more than selecting separate speech-recognition and voice providers. Standard mode is a better fit when you need the broader set of speech-recognition controls, voice providers, or tuning options.

Configure and test

  1. Open the Agent editor and select Voice.
  2. Choose speech-to-speech mode and an available provider.
  3. Select a compatible voice.
  4. Test realistic calls, including interruptions, names, numbers, and background noise.
  5. Publish only after the staged Agent behaves as expected.
Changing speech mode can change which voices and controls are available. Review the complete Voice and speech mode guide before switching a live Agent.
Configuring speech mode through the API? See Language, voice, and speech.