Google is giving Gemini Live a major upgrade with two new models designed to make real-time AI conversations more capable, natural and useful. The company has introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its latest live dialogue models for voice agents and complex, voice-driven tasks.
Announced on September 15, the new models are built around near-real-time reasoning, allowing Gemini to do more than simply respond to spoken prompts. Gemini 3.8 Live is designed for scalable, cost-efficient conversational experiences, while Gemini 3.8 Live Extended Thinking is aimed at more complicated tasks that require deeper, multi-step reasoning.
The models are also being integrated across Google’s ecosystem, including Gemini Live, Search Live and selected Google Workspace experiences. Developers can access both models through the Gemini API and Google AI Studio.
Gemini can keep talking while it works
One of the biggest changes with Gemini 3.8 Live is the ability to execute tools and API calls in the background without stopping the conversation.
That means a voice agent doesn’t necessarily have to become silent while it waits for an external action to finish. Gemini can acknowledge what the user asked, continue the conversation and handle the underlying task asynchronously.
For example, an agent could start an API request or another tool-based operation while continuing to communicate with the user. Google says this is intended to make interactions feel more like a natural conversation rather than a sequence of prompts followed by long pauses.
Gemini 3.8 Live can also incorporate visual information in near real time. This gives the model additional context about what the user is seeing, opening the door to voice interactions where Gemini can simultaneously listen, look and respond.
Google highlights applications such as real-time employee onboarding and troubleshooting, along with demonstrations involving chess and other visually grounded tasks.
The model can also automatically detect and switch between 97 supported languages during a conversation, making it possible to move between languages without manually changing settings.
Extended Thinking brings deeper reasoning to voice
Gemini 3.8 Live Extended Thinking takes the same basic idea further by adding more intensive reasoning for complex workflows.
Rather than forcing users to wait silently while an AI works through a complicated request, the model can reason in the background while continuing to speak. Google says it can provide early verbal cues such as “Let me check that…” and narrate progress as a multi-step task moves forward.
That combination could be particularly useful for voice agents that need to coordinate several actions. Instead of treating reasoning and conversation as separate phases, the model is designed to handle both simultaneously.
Google describes the model as being intended for enterprise-grade task completion and complex workflows, with support for configurable thinking. Developers can therefore use deeper reasoning when a task requires it rather than treating every voice interaction as a simple question-and-answer exchange.
Google points to strong voice-agent benchmark results
Google is positioning Gemini 3.8 Live Extended Thinking as one of its strongest models for real-time speech-to-speech interaction.
According to Google’s reported benchmarks, the model scored 82.6 on Artificial Analysis’ Speech to Speech Quality Index, where Google says it took the overall top position. It also reported scores of 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking benchmark for agentic task completion. On Big Bench Audio, the model scored 97.7%.
Gemini 3.8 Live, meanwhile, placed second in Google’s cited Speech Agent Arena results.
These are Google’s reported benchmark results, so they should be viewed in the context of the particular tests, configurations and evaluation environments used. Google also says both models performed well on ServiceNow’s EVA-Bench, where the company evaluated the balance between task completion and conversational quality.
Built for developers and voice agents
The new models aren’t limited to Google’s own consumer applications. Google is also making them available as part of its developer-focused Gemini Live API.
Developers can use Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking to build voice-driven applications that combine speech, reasoning, visual context and tool execution.
Google lists asynchronous function calling, visual context, multilingual support, alphanumeric precision and incremental content updates among the capabilities available to developers. The company says the models can handle things such as confirmation codes, claim numbers and other technical information where accurate handling of spoken alphanumeric data matters.
The models are also supported through a growing collection of development platforms, including Agora, Fishjam, LiveKit, LangChain, Pipecat, Vercel and Vision Agents. These platforms can handle parts of the real-time media infrastructure, allowing developers to focus more heavily on the actual voice-agent experience.
For developers using the Live API, Google lists pricing of $0.005 per minute for audio input and $0.018 per minute for audio output.
Gemini 3.8 Live is coming to more Google products
The new models are also being pushed into Google’s own products.
Gemini 3.8 Live is rolling out to developers through the Gemini API and Google AI Studio, while enterprise customers can access it through a private preview in Gemini Enterprise. Google says it is also available to everyone through Search Live.
Gemini 3.8 Live Extended Thinking is similarly rolling out through the Gemini API and Google AI Studio. On the consumer side, Google says it is available in Gemini Live, while Google AI Pro and Ultra subscribers can use it in Docs through Workspace. Gmail and Keep are also getting the model for Google AI subscribers.
Google is effectively turning voice interaction into another interface for getting work done rather than treating it purely as a conversational feature. The combination of background tool execution, visual understanding and simultaneous reasoning could make that distinction increasingly important.
Google is also watermarking AI-generated audio
Google says audio generated by its AI products includes SynthID, the company’s imperceptible watermarking technology.
The watermark is embedded into generated audio so that AI-created content can be identified, according to Google. The company says this is intended to improve transparency and help address the potential spread of misinformation involving synthetic audio.
That becomes increasingly relevant as voice models become more convincing and capable of carrying on longer, more natural conversations.
With Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, Google is moving its live AI models toward a model of interaction where the assistant doesn’t simply wait for a prompt, generate a response and stop. Instead, it can listen, see, reason, talk and work on tasks at the same time.
For users, that could make Gemini Live feel considerably less like a voice chatbot and more like an always-available voice interface for getting things done.
Discover more from GadgetBond
Subscribe to get the latest posts sent to your email.
