Skip to content

Product1 publisher2 min readPublished

Gemini 3.8 Live keeps talking while it runs the tool call in the background

Google's two new voice models process speech and reasoning at the same time and can fire tool calls without pausing the conversation. Published scores put the stronger one top of a speech-to-speech index and lowest on banking tasks.

The Product Desk · Product desk

What happened

  • Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, pitching both at the latency problem that has kept voice agents scripted and shallow.
  • Both models execute third-party tool and API calls in the background, so the agent keeps conversing at a normal pace while it works on the task it was just given.
  • On Google's own numbers, Extended Thinking scored 82.6 on the Artificial Analysis Speech to Speech Quality Index, ahead of GPT-Live-1-Astra and Grok Voice Think Fast 2.0.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • cost Output audio bills at 3.6 times input, so a script that makes the agent talk more costs more than one that makes it listen more, and the difference lands on whoever owns the call minutes.
  • constraint Because reasoning tokens are billed on top for Extended Thinking, a per-minute rate cannot be used to forecast what the higher-scoring model will cost; that number has to come from measured calls.
  • decision Teams picking between the cheap predictable model and the top scorer now have to make that call per use case, because the same model's scores move by tens of points across domains.
  • precedent Verbal filler shipped as a documented feature sets the expectation that agents cover their own thinking time with speech, which QA has to start scoring alongside accuracy.

"Let me check that" is now a listed product feature. Google says the models use early verbal cues like it to acknowledge a prompt in a more lifelike way [4]. Contact centres have been running that line for decades under a plainer name. SiliconANGLE's account of Google's post describes near-real-time reasoning without a latency measurement, so how much of the wait actually shrank is not in the published material [21].

The scores are Google's own, and SiliconANGLE published them with the aside "if Google is to be believed" [8]. Inside a single model the range is wide: 97.7% on Big Bench Audio, 68.6% on T-Voice, and 35.1% on T-Voice-banking [7]. Banking sits 33.5 points below general T-Voice and 62.6 points below Big Bench Audio [19]. On that set, 64.9% of attempts were not successes [20]. The standard model placed second on Speech Agent Arena and first on ServiceNow's EVA-Bench [6].

A ten-minute call split evenly between the two sides works out at 5 x $0.005 plus 5 x $0.018, or about 11.5 cents of audio [18]. Extended Thinking bills reasoning tokens on top of that, along with inputs such as video and documents [10].

For the person who has to deploy this, the two models arrive in different places. Both are live in the Gemini API and Google AI Studio, with an enterprise private preview in Gemini Enterprise and Search Live [11]. Extended Thinking also reaches Workspace through Docs, Gmail and Keep for subscribers, plus the Gemini Live apps [12]. Existing voice stacks connect through partners including LiveKit, Pipecat, Agora, Vercel, Fishjam and Vision Agents [13]. The models handle 97 languages, detect the language automatically and can switch mid-conversation [14], and generated audio carries an invisible SynthID watermark that Google said can be used to detect misinformation [15].

Google AI's launch post said the models "let you speak, collaborate, and execute tasks seamlessly, meaning conversing with AI just got a lot more natural" [16]. Two things sort whether that holds for a given use case. The first is whether the tool call can finish while the caller keeps talking, or whether its answer gates the agent's next sentence; background execution only pays in the first case, and in the second the caller hears the filler and then waits anyway. The second is what a wrong answer at that step costs the person on the line, because a mis-stated account balance delivered fluently is a different failure from a mis-stated opening hour. For regulated, transactional work of the kind T-Voice-banking is meant to stand in for, completion rate on last week's real calls is a better guide than any index score.

What to watch

  • Whether Google or an independent tester publishes end-to-end latency figures for either model.
  • Whether the Gemini Enterprise and Search Live private preview opens up with a published per-minute rate for Extended Thinking.
  • Whether an independent run of T-Voice-banking reproduces the score Google reported.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories