Google introduced two new audio models that enable enterprises to deploy intelligent AI agents with reasoning depth and more interactivity and conversational speed.
Launched on September 15, Gemini 3.8 Live and Live Extended Thinking let developers build voice agents that can reason and execute tasks while maintaining the flow of the conversation, Google said.
The 3.8 Live Extended Thinking version supports configurable thinking, a feature that lets users turn an AI model’s internal reasoning on or off.
The release of the audio models comes two weeks after Google initially introduced the Gemini 3.8 Flash and 3.8 Cyber foundation models.
Still, Google tech takes it to the next level by changing how users interact with AI, he said.
Google introduced two new audio models that enable enterprises to deploy intelligent AI agents with reasoning depth and more interactivity and conversational speed.
Launched on September 15, Gemini 3.8 Live and Live Extended Thinking let developers build voice agents that can reason and execute tasks while maintaining the flow of the conversation, Google said. The models’ key capabilities include performing API and tool calls while continuing to stream audio responses to users, providing live visual inputs that help agents understand what users say and see, and supporting more than 97 languages. The 3.8 Live Extended Thinking version supports configurable thinking, a feature that lets users turn an AI model’s internal reasoning on or off.
The release of the audio models comes two weeks after Google initially introduced the Gemini 3.8 Flash and 3.8 Cyber foundation models. With 3.8 Live’s capabilities, Google appears to be catching up to vendors such as Apple, which already offers real-time transcription in productivity apps such as Notes and Voice Memos, said Bradley Shimmin, an analyst at Futurum Group. Still, Google tech takes it to the next level by changing how users interact with AI, he said.