Build real-time voice applications with Gemini 3.8 Live and 3.5 Transcribe
Google has released Gemini 3.8 Live, Gemini 3.8 Live Extended Thinking, and Gemini 3.5 Transcribe for developer integration. These models provide new capabilities for real-time voice application development and audio processing.
Verified State Diff
Impact & Verification Analysis
Software developers, voice application engineers, and enterprise customers building real-time AI agents.
This release provides developers with native, high-performance tools for building low-latency voice interfaces, reducing reliance on external transcription services and improving the reasoning capabilities of conversational AI agents.
Full Fact Overview
The release introduces three distinct models: Gemini 3.8 Live, which focuses on low-latency conversational interaction; Gemini 3.8 Live Extended Thinking, which likely incorporates chain-of-thought processing for complex voice queries; and Gemini 3.5 Transcribe, a specialized model for speech-to-text tasks. This expansion signals Google's move to provide modular, specialized audio-processing capabilities within the Gemini ecosystem, moving beyond general-purpose multimodal models to specific, high-performance voice-centric APIs.