Skip to content

Product7 publishers3 min readPublished

Google gives Gemini's live voice agents a lip-syncing face for enterprise customer service

Google added lip-syncing video avatars to its Gemini 3.8 Live voice agents, available now to Gemini Enterprise customers in 97 languages. Early reviews called the faces creepy and the launch came with no customer evidence, so support teams will have to measure trust in their own queues.

The Product Desk · Product desk

Illustration accompanying Google gives Gemini's live voice agents a lip-syncing face for enterprise customer service

What happened

  • Google then added Live Avatar, which pairs a near real-time generated face that lip-syncs, shows expressions and takes turns with those live dialogue models.
  • While the avatar keeps talking, the agent behind it can trigger tool calls and fetch data in the background.
  • The avatar works across 97 languages, adjusting its speech, lip movements and expressions as it goes.
  • Live Avatar is available now, but only to companies on Gemini Enterprise.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • decision A support team can ship the voice agent first and add the face later as its own experiment, since the avatar changes how the agent looks and leaves what it can do alone.
  • exposure Because Google designed its SynthID watermark to be imperceptible, telling a customer on screen that the agent is software falls to the company that deploys it.
  • constraint Without allowlisting, a company's first test uses one of Google's preset faces, so an agent built on the brand's own likeness waits on Google's approval.

A customer with a wrong charge on a bill opens the support page on a phone. A generated person answers, wearing a uniform with the company's logo on it, one of the custom touches TechRadar describes [15]. TechRadar lists the intended jobs as customer service, receptionists and training [13]. Digital Trends adds interactive walkthroughs [14]. Research scientist Shuo-yiin Chang and software engineer CJ Zheng said it now feels like AI has a much more "visual presence," TechRadar reported [12].

Take the face away and the work underneath is the same. Digital Trends describes Live Avatar as the September 15 foundation with a visual layer added [4], and background tool execution was already one of the things those models focused on [5]. The lookup or the refund happens in a tool call whether or not a face is on screen.

The story teams will tell themselves is the vendor's. In TechRadar's summary, the avatar makes it feel more like you're speaking to a human [19]. The first outside reactions came from reviewers, and they went the other way. Engadget and TechRadar both put "creepy" in their headlines [17][18], and Engadget wrote that, given people's distaste for AI slop, the avatars "may not be universally popular" [20]. Neither is customer data, and none of the coverage reports a pilot result or a price for the avatar layer. The one cost claim, Google's via TechRadar, is that Live Extended Thinking is cheaper per hour of input audio than two rival voice models [21].

On disclosure, Google points to SynthID. "This imperceptible watermark is woven directly into the audio and video output, helping to ensure AI-generated content remains detectable to help minimise misinformation and misattribution," the company wrote [11]. Google's videos show both realistic and cartoon-like avatars [16]. Of the two, only the cartoon tells a customer it is software before it says anything.

I'd start with voice alone and add the face only in flows where it shows the customer something, such as a walkthrough or a training session. The cost is the brand presence Google is selling. If the face does help some group of customers, a team that held it back finds out later than one that tried it everywhere.

For the decision itself, use a 2x2. One axis is whether the face carries information the voice cannot, as in a walkthrough, or only presence, as in a billing dispute. The other is whether the customer chose a video channel or landed in one by default. Information and choice together is where a face gets tested first. Presence by default is where voice alone should stay. The two mixed quadrants get the same agent run with and without the avatar. Minutes spent watching the face will move in that test and prove nothing about the bill. The scorecard is repeat contacts within seven days and handoffs to a human.

What to watch

  • Whether Google publishes a price for the Live Avatar video layer separate from the voice model's per-hour audio cost.
  • Whether any Gemini Enterprise customer reports repeat-contact or human-handoff numbers from an avatar pilot against a voice-only agent.
  • Whether Google opens custom avatars beyond allowlisted enterprises or offers Live Avatar outside Gemini Enterprise.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories