Gemini 3.8 Live Avatar: 97 Languages, Enterprise Only
Google's Live Avatar gives Gemini 3.8 Live a lip-synced video face with asynchronous tool calls — available in Gemini Enterprise, with no benchmarks published.

Google shipped Live Avatar for Gemini 3.8 Live on September 24, giving enterprise voice agents a lip-synced video face that it says switches across 97 languages mid-conversation.
The announcement comes from the Gemini Audio team — Shuo-yiin Chang, a research scientist, and CJ Zheng, a software engineer — and it lands, in Google's words, "building on the momentum of last week's Gemini 3.8 Live launch." Every capability below is Google's own description; the post publishes no benchmarks.
What Live Avatar adds
Google describes the feature as pairing near real-time video generation with speech, producing "an experience that listens, sees, and speaks with a dynamic visual persona." The claimed qualities are precise lip-syncing, natural expressions and fluid turn-taking. Google positions it for customer service and interactive walkthroughs rather than consumer chat.
Availability is narrow and stated plainly: "Starting today, Gemini 3.8 Live with Live Avatar is available in Gemini Enterprise." There is no consumer tier and no free tier in the post.
The part that matters to engineers
The most consequential line is not about the face. Google says Live Avatar supports asynchronous tool calling, so it "can trigger tool calls and fetch data in the background while continuing active dialogue, handling complex tasks while ensuring an uninterrupted conversational flow."
If you have built a voice agent, you know why that sentence exists. Synchronous tool calls are where voice agents go quiet — the model stops talking, the user hears nothing, and the silence reads as a dropped call. Moving the fetch off the speaking path is the difference between a demo and something a support queue can use. Google's illustrative scenario is checking a guest into a hotel while the dialogue continues uninterrupted.
Languages and custom avatars
On multilingual behaviour, Google says the feature "dynamically adapts its lip-sync and expressions and can seamlessly transition across 97 languages without degrading video fidelity or introducing visual drift." That is an unusually strong claim for generated video and the post offers no measurement behind it, so treat 97 languages as a supported-language count rather than a verified quality figure.
Organisations get a library of preset avatars, and Google says developers can generate "a fully animated, responsive avatar" from a high-quality reference image while preserving likeness, brand styling or character identity. That path is gated: custom avatar creation is currently available only through enterprise allowlisting.
Provenance
Google states that all output from its AI products carries SynthID, describing it as an imperceptible watermark "woven directly into the audio and video output" to keep AI-generated content detectable. For a product whose entire purpose is putting a synthetic human face in front of customers, that is the load-bearing safety claim, and it is the one most worth independent testing.
What to watch
Two things will tell you whether this is production-ready. First, latency numbers: Google repeatedly says "near real-time" and never quantifies it, and for a talking face the tolerance is far tighter than for text. Second, whether custom avatars leave allowlisting — as long as likeness generation stays gated, most teams will be shipping a preset face, which changes the buying case from brand presence to plain call handling. The API documentation is the place to check both.
More from DangMua