Google Launches New Voice AI Models for Building Real-Time Conversational Apps
This week, Google released two new AI models, Gemini 3.8 Live and Gemini 3.5 Transcribe, via the Gemini Live API. These models aim to help developers build low-latency conversational voice agents with multilingual support, visual context understanding, and more.
First, Gemini 3.8 Live is Google's built-in model for fast verbal conversations. The model can run API and tool calls in the background while streaming audio responses without interruptions. It analyzes live images and videos at up to 1 frame per second to connect users' speech with what they see.
This model supports 97 languages, with what Google claims are realistic accents and automatic mid-conversation language switching. Sessions last up to 15 minutes for audio-only and 2 minutes for audio and video. Pricing starts at $0.005 per minute for audio input and $0.018 per minute for audio output.
Gemini 3.5 Transcribe is Google's dedicated speech-to-text model, designed for transcription. The model...
Copyright of this story solely belongs to www.extremetech.com. To see the full text click HERE