Skip to content

BuildReports disagree8 publishers2 min readPublished Updated

Gemini 3.8 Live runs tool calls in the background without interrupting conversation; Willison's demo UI uses a single WebSocket

Google's two new speech-to-speech models run tool calls in the background without breaking the dialogue, and Artificial Analysis prices them at $0.84 and $3.50 per hour of input audio on one benchmark subset.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying Gemini 3.8 Live runs tool calls in the background without interrupting conversation; Willison's demo UI uses a single WebSocket
Generated illustration

What happened

  • Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026, two speech-to-speech models that Simon Willison described as a similar shape to OpenAI's GPT-Live family.
  • Both models detect and switch between 97 languages mid-conversation, execute tool calls and API requests in the background while they keep talking, and process visual input in near real time.
  • On Sierra's banking leaderboard for customer-service tasks, Extended Thinking resolves 35.1%, against 32.0% for GPT-Live-1 Astra and 16.5% for xAI-Realtime.
  • Artificial Analysis's cost per hour of input audio on the Big Bench Audio subset puts 3.8 Live at $0.84 and Extended Thinking at $3.50, against $4.80 for Grok Voice Think Fast 2.0 and $5.83 for GPT-Live-1 Astra.
  • Simon Willison had OpenAI's GPT-6 Astra Extra High read Google's documentation and build him a browser client for holding voice conversations with the two models.

Why it matters

  • constraint Budgeting from the $0.84 and $3.50 figures imports Big Bench Audio's ratio of model reasoning to caller speech. A support queue where the agent thinks more per minute of audio pays above that average.
  • decision The two tiers sit 4.2 times apart in hourly cost and 6.6 points apart in speech quality. Anyone running both has to decide which turns are worth routing to the reasoning tier.
  • capability A client that is one WebSocket and an AudioContext lets a team instrument its own capture, playback and interrupt timings instead of waiting for an SDK release to expose them.
  • exposure Every second of audio these models emit carries a SynthID watermark, so recordings a business republishes or archives are detectable as model-generated.

Google's own phrasing for the new capability is real-time reasoning, visual-context handling and background task execution without interrupting the dialogue [1]. Neither launch account carries a latency number, so a team rebuilding its voice budget around this has to get the numbers from its own traces.

The transport is documented. Willison's client connects to the BidiGenerateContent WebSocket endpoint and uses a Web Audio API AudioContext for both capture and playback [12]. "The implementation uses no libraries," Willison wrote [14]. Interrupting the model mid-sentence works in it [15]. Capture, playback scheduling and barge-in therefore sit in application code.

Now the two tiers. Extended Thinking scores 82.6% on Artificial Analysis's Speech to Speech Quality Index and the standard model 76.0% [5], a gap of 6.6 points [18]. Per hour of input audio, Extended Thinking costs 4.2 times the standard tier [17]. Measured against GPT-Live-1 Astra's hourly figure, the cheap tier is 6.9 times cheaper [25]; officechai.com called it roughly a sixth of the price [22].

Against GPT-Live-1 Astra (Medium) the quality lead is 1.1 points on the speech index [19] and 0.7 points on Artificial Analysis's agentic benchmark [20], while Extended Thinking's hourly cost is 40% lower [21].

For an hourly rate to transfer, the traffic has to resemble the subset it was measured on. The denominator is hours of input audio and the numerator covers whatever the model thinks and says in between, so a queue that reasons more per minute of caller speech will not land on the benchmark average. Extended Thinking is the tier sold on multi-step reasoning [9]. dev.to reports that the announcement material it saw did not include exact pricing or a full regional rollout schedule [26].

The closest thing to a deployment number is Sierra's banking leaderboard, and it leaves 64.9% of tasks unresolved [24]. dev.to arrives at the same point in general terms: the practical value will depend on how a business designs its workflows, integrations and escalation paths [23].

Availability splits by tier. 3.8 Live goes out through the Gemini API, Google AI Studio and Search Live, with enterprise access through Gemini Enterprise, and Extended Thinking adds Gemini Live plus the Workspace apps Docs, Gmail and Keep for AI Pro and Ultra subscribers [6]. dev.to reports the Gemini Enterprise for Customer Experience route as still forthcoming, with enterprise access running through private previews [8]. Google names LiveKit, Pipecat and Agora as platform partners and Salesforce and ServiceNow among the enterprise customers building on the models [10].

What to watch

  • A published Google rate card for the Live API, which would replace Artificial Analysis's per-hour estimate with billable per-token terms.
  • Whether the Gemini Enterprise for Customer Experience route ships, and what escalation hooks it exposes to a support stack.
  • Independent latency figures: time to first audio, and tool-call turnaround while background reasoning holds the floor.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence60
Adoption22
Hype gap+30
Incentives72
Confidence64

Perspective Coverage

8 publishers
Builder
Builder 48%
Operator
Operator 29%
Investor
Investor 23%
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Google describes the models as capable of real-time reasoning, visual-context handling and background task execution without interrupting the dialogue.

  2. [2]

    Both Gemini 3.8 Live models can detect and switch between 97 languages mid-conversation, execute tool calls and API requests in the background while continuing to talk, and process visual input in near real time.

  3. [3]

    Google released Gemini 3.8 Live and 3.8 Live Extended Thinking on 15 September 2026, two speech-to-speech models that are a similar shape to OpenAI's GPT-Live family.

Sources

8 independent publishers whose own reporting we read for this story.

  1. blog.google

    3 articles · September 15, 2026

    Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking
  2. dev.to

    2 articles · September 18, 2026

    Build real-time voice applications with Gemini 3.8 Live and 3.5 Transcribe
  3. officechai.com

    1 article · September 17, 2026

    Google Releases Gemini 3.8 Live-Extended Conversational Model, Claims Better Performance Than Rivals At Lower Price
  4. runtimewire.com

    1 article · September 15, 2026

    Google ships Gemini 3.8 Live to keep voice agents talking while tools work
  5. simonwillison.net

    2 articles · September 15, 2026

    Gemini Live audio
  6. testingcatalog.com

    2 articles · September 15, 2026

    Google rolled out Gemini 3.8 Live and Extended Thinking
  7. the-decoder.com

    4 articles · September 15, 2026

    Apple brings a fully revamped Siri built on Google's Gemini, but not to the EU
  8. thenewstack.io

    2 articles · September 15, 2026

    OpenAI’s voice model doesn’t think. That’s the point.

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Loading related stories