Voice-based AI agents are getting a much more visible presence. Google Cloud has made Gemini 3.8 Live with Live Avatar generally available, allowing enterprise developers to build conversational agents that can not only listen and speak in real time, but also appear on screen as responsive video avatars.
The technology is now available in Gemini Enterprise after first being previewed at Google Cloud Next 2026. Google says the combination of Gemini 3.8 Live‘s native speech-to-speech capabilities and Live Avatar is designed to make interactions with enterprise AI agents feel more natural across websites, mobile applications, and interactive kiosks.
The most obvious change is the addition of a visual presence. Live Avatar can generate video avatars with synchronized lip movements, giving a voice agent an on-screen persona that can respond while the conversation is happening. Google provides a library of curated, pre-built avatars, while custom avatar creation is currently restricted to customers that receive enterprise allowlisting and verification.
Underneath that visual layer is Gemini 3.8 Live’s real-time speech-to-speech architecture. Instead of treating a voice conversation as a sequence of separate speech recognition, language-model and text-to-speech steps, the system is designed around native spoken interaction. Google says this allows the agent to handle interruptions more naturally without losing the surrounding conversation context or backend transactions.
That becomes particularly interesting when the agent needs to do something while talking. Gemini 3.8 Live supports asynchronous tool calling, allowing API, CRM, ERP and other backend operations to run in the background while the agent continues speaking. Rather than going silent while a transaction completes, the agent can acknowledge the request and continue the conversation before the underlying task finishes.
Google is also emphasizing the system’s multimodal capabilities. Gemini 3.8 Live can process live camera feeds and screen shares alongside audio, allowing an agent to understand what a customer is seeing while simultaneously listening to what they are saying.
That opens up use cases that would be difficult for a conventional voice bot. In one Google demo, an insurance claims agent listens to a customer describe damage while the customer points a camera at it. The information can then be incorporated into the claims workflow while other agents and backend systems handle policy checks and intake rules.
Another demo shows the technology being used for vehicle shopping. Cox Automotive has built an AI shopping assistant for Autotrader that uses conversational interaction, live screen highlighting and tool calling to help customers search for vehicles, compare options and work through financing.
Language support is another major part of the release. Gemini 3.8 Live can understand and speak 97 languages, with automatic language detection allowing conversations to move between supported languages without requiring users to manually switch settings. Google says this is already being used by Equal AI, whose AI handles more than one million live calls per day across nine Indian languages.
For developers, Google is positioning Gemini 3.8 Live as a foundation for real-time, bidirectional applications rather than simply another chatbot model. Developers can connect the Gemini Live API with Google’s Agent Development Kit, allowing them to define agents, manage session memory and stream audio directly to the model without building a traditional speech-to-text pipeline themselves.
The enterprise controls around the system are just as notable as the avatar itself. Gemini 3.8 Live with Live Avatar is available through US and EU endpoints with provisioned throughput, enterprise compliance and data-governance controls. Google also says generated audio and video streams include imperceptible SynthID watermarks to help identify AI-generated media.
Google is keeping custom avatars behind additional safeguards. While companies can use Google’s curated avatar library, creating a custom avatar requires allowlisting and an identity-verification process intended to reduce impersonation and other misuse.
The timing is significant because enterprise AI agents are increasingly moving beyond text interfaces. A voice agent that can see a customer’s screen, understand a camera feed, call business systems and maintain a natural conversation is much closer to a digital service representative than the phone-tree-style bots that have traditionally handled customer calls.
Gemini 3.8 Live with Live Avatar is now available to enterprise customers, with Google offering access through Gemini Enterprise as well as the Gemini Live API for developers building their own applications. Custom avatar deployments and provisioned throughput require working with Google Cloud.
Google is also continuing to develop the Gemini 3.8 Live family. Gemini 3.8 Live Extended Thinking remains in private preview, meaning the Live Avatar release is currently focused on real-time conversational interaction rather than exposing the extended-thinking variant to general enterprise production users.
The bigger shift is that enterprise voice AI no longer has to be voice-only. With Gemini 3.8 Live and Live Avatar, Google is combining speech, video, vision, multilingual conversation and backend actions into one continuous interaction — giving businesses a way to build AI agents that can both talk to customers and actively work alongside them.
Discover more from GadgetBond
Subscribe to get the latest posts sent to your email.