BuildReports disagree8 publishers2 min readPublished Updated
Gemini 3.8 Live runs tool calls in the background without interrupting conversation; Willison's demo UI uses a single WebSocket
Google's two new speech-to-speech models run tool calls in the background without breaking the dialogue, and Artificial Analysis prices them at $0.84 and $3.50 per hour of input audio on one benchmark subset.
The Engineer · Build desk

What happened
- Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026, two speech-to-speech models that Simon Willison described as a similar shape to OpenAI's GPT-Live family.
- Both models detect and switch between 97 languages mid-conversation, execute tool calls and API requests in the background while they keep talking, and process visual input in near real time.
- On Sierra's banking leaderboard for customer-service tasks, Extended Thinking resolves 35.1%, against 32.0% for GPT-Live-1 Astra and 16.5% for xAI-Realtime.
- Artificial Analysis's cost per hour of input audio on the Big Bench Audio subset puts 3.8 Live at $0.84 and Extended Thinking at $3.50, against $4.80 for Grok Voice Think Fast 2.0 and $5.83 for GPT-Live-1 Astra.
- Simon Willison had OpenAI's GPT-6 Astra Extra High read Google's documentation and build him a browser client for holding voice conversations with the two models.
Why it matters
- constraint Budgeting from the $0.84 and $3.50 figures imports Big Bench Audio's ratio of model reasoning to caller speech. A support queue where the agent thinks more per minute of audio pays above that average.
- decision The two tiers sit 4.2 times apart in hourly cost and 6.6 points apart in speech quality. Anyone running both has to decide which turns are worth routing to the reasoning tier.
- capability A client that is one WebSocket and an AudioContext lets a team instrument its own capture, playback and interrupt timings instead of waiting for an SDK release to expose them.
- exposure Every second of audio these models emit carries a SynthID watermark, so recordings a business republishes or archives are detectable as model-generated.
Google's own phrasing for the new capability is real-time reasoning, visual-context handling and background task execution without interrupting the dialogue [1]. Neither launch account carries a latency number, so a team rebuilding its voice budget around this has to get the numbers from its own traces.
The transport is documented. Willison's client connects to the BidiGenerateContent WebSocket endpoint and uses a Web Audio API AudioContext for both capture and playback [12]. "The implementation uses no libraries," Willison wrote [14]. Interrupting the model mid-sentence works in it [15]. Capture, playback scheduling and barge-in therefore sit in application code.
Now the two tiers. Extended Thinking scores 82.6% on Artificial Analysis's Speech to Speech Quality Index and the standard model 76.0% [5], a gap of 6.6 points [18]. Per hour of input audio, Extended Thinking costs 4.2 times the standard tier [17]. Measured against GPT-Live-1 Astra's hourly figure, the cheap tier is 6.9 times cheaper [25]; officechai.com called it roughly a sixth of the price [22].
Against GPT-Live-1 Astra (Medium) the quality lead is 1.1 points on the speech index [19] and 0.7 points on Artificial Analysis's agentic benchmark [20], while Extended Thinking's hourly cost is 40% lower [21].
For an hourly rate to transfer, the traffic has to resemble the subset it was measured on. The denominator is hours of input audio and the numerator covers whatever the model thinks and says in between, so a queue that reasons more per minute of caller speech will not land on the benchmark average. Extended Thinking is the tier sold on multi-step reasoning [9]. dev.to reports that the announcement material it saw did not include exact pricing or a full regional rollout schedule [26].
The closest thing to a deployment number is Sierra's banking leaderboard, and it leaves 64.9% of tasks unresolved [24]. dev.to arrives at the same point in general terms: the practical value will depend on how a business designs its workflows, integrations and escalation paths [23].
Availability splits by tier. 3.8 Live goes out through the Gemini API, Google AI Studio and Search Live, with enterprise access through Gemini Enterprise, and Extended Thinking adds Gemini Live plus the Workspace apps Docs, Gmail and Keep for AI Pro and Ultra subscribers [6]. dev.to reports the Gemini Enterprise for Customer Experience route as still forthcoming, with enterprise access running through private previews [8]. Google names LiveKit, Pipecat and Agora as platform partners and Salesforce and ServiceNow among the enterprise customers building on the models [10].
What to watch
- A published Google rate card for the Live API, which would replace Artificial Analysis's per-hour estimate with billable per-token terms.
- Whether the Gemini Enterprise for Customer Experience route ships, and what escalation hooks it exposes to a support stack.
- Independent latency figures: time to first audio, and tool-call turnaround while background reasoning holds the floor.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence60
- Adoption22
- Hype gap+30
- Incentives72
- Confidence64
Perspective Coverage
8 publishers- Builder
- Builder 48%
- Operator
- Operator 29%
- Investor
- Investor 23%
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Google describes the models as capable of real-time reasoning, visual-context handling and background task execution without interrupting the dialogue.
- [2]
Both Gemini 3.8 Live models can detect and switch between 97 languages mid-conversation, execute tool calls and API requests in the background while continuing to talk, and process visual input in near real time.
- [3]
Google released Gemini 3.8 Live and 3.8 Live Extended Thinking on 15 September 2026, two speech-to-speech models that are a similar shape to OpenAI's GPT-Live family.
- [4]
On Artificial Analysis's cost-per-hour-of-input-audio measure using the Big Bench Audio subset, Gemini 3.8 Live costs $0.84 an hour, 3.8 Live Extended Thinking $3.50, Grok Voice Think Fast 2.0 $4.80 and GPT-Live-1 Astra $5.83.
- [5]
On Artificial Analysis's Speech to Speech Quality Index, Gemini 3.8 Live Extended Thinking scores 82.6%, GPT-Live-1 Astra (Medium) 81.5%, Grok Voice Think Fast 2.0 (High) 81.3% and the standard Gemini 3.8 Live 76.0%.
- [6]
3.8 Live rolls out through the Gemini API, Google AI Studio and Search Live with enterprise access via Gemini Enterprise; Extended Thinking gets the same developer and enterprise rollout plus a consumer release inside Gemini Live and Google Workspace Docs, Gmail and Keep for Google AI Pro and Ultra subscribers.
- [7]
On Artificial Analysis's tau-Voice benchmark, 3.8 Live Extended Thinking scores 68.6%, GPT-Live-1 Astra 67.9% and Grok Voice Think Fast 2.0 56.5%.
- [8]
dev.to reports that Gemini Enterprise customers get private previews and that the models coming to Gemini Enterprise for Customer Experience is explicitly described as forthcoming.
- [9]
Google positions the standard model as a scalable, cost-efficient option for live dialogue and the Extended Thinking version for more complex work that benefits from increased intelligence and multi-step reasoning.
- [10]
Google says it is partnering with platforms including LiveKit, Pipecat and Agora, along with enterprise customers including Salesforce and ServiceNow, to build voice agents on the new models.
- [11]
Simon Willison pointed GPT-6 Astra Extra High at Google's documentation and had it build a web UI for trying the new models, with model and voice preset selection, an optional system prompt, and a voice conversation in the browser.
- [12]
The UI connects to the wss://generativelanguage.googleapis.com BidiGenerateContent WebSocket endpoint and uses a Web Audio API AudioContext for both capture and playback.
- [13]
All audio output from the models carries Google's SynthID watermark, its technology for helping detect AI-generated content.
- [14]
"The implementation uses no libraries," Willison wrote.
- [15]
The browser UI includes the ability to interrupt the model while it is talking.
- [16]
On Sierra's tau-cubed Banking Leaderboard, which tests how well a voice agent resolves customer-service style banking tasks, 3.8 Live Extended Thinking leads with 35.1%, against 32.0% for GPT-Live-1 Astra and 16.5% for xAI-Realtime.
- [17]
Extended Thinking costs about 4.2 times the standard tier per hour of input audio.
- [18]
The Speech to Speech Quality Index gap between Extended Thinking and the standard model is 6.6 points.
- [19]
Extended Thinking's lead over GPT-Live-1 Astra (Medium) on the Speech to Speech Quality Index is 1.1 points.
- [20]
Extended Thinking's lead over GPT-Live-1 Astra on the tau-Voice benchmark is 0.7 points.
- [21]
Extended Thinking's hourly input-audio cost is 40% lower than GPT-Live-1 Astra's.
- [22]
officechai.com wrote that the standard Gemini 3.8 Live model is roughly a sixth of the price of GPT-Live-1 Astra for a good chunk of the same conversational ability.
- [23]
dev.to writes that the practical value will depend on how a business designs its workflows, integrations and escalation paths.
- [24]
Extended Thinking's 35.1% score on Sierra's banking leaderboard leaves 64.9% of tasks unresolved.
- [25]
Gemini 3.8 Live's hourly cost is 6.9 times lower than GPT-Live-1 Astra's, closer to a seventh than a sixth.
- [26]
dev.to reports that Google has not provided exact pricing or a full regional rollout schedule in the material it saw.
Sources
8 independent publishers whose own reporting we read for this story.
- blog.googleIntroducing Gemini 3.8 Live and 3.8 Live Extended Thinking
3 articles · September 15, 2026
- dev.toBuild real-time voice applications with Gemini 3.8 Live and 3.5 Transcribe
2 articles · September 18, 2026
- Google Releases Gemini 3.8 Live-Extended Conversational Model, Claims Better Performance Than Rivals At Lower Price
officechai.com
1 article · September 17, 2026
- runtimewire.comGoogle ships Gemini 3.8 Live to keep voice agents talking while tools work
1 article · September 15, 2026
- simonwillison.netGemini Live audio
2 articles · September 15, 2026
- testingcatalog.comGoogle rolled out Gemini 3.8 Live and Extended Thinking
2 articles · September 15, 2026
- the-decoder.comApple brings a fully revamped Siri built on Google's Gemini, but not to the EU
4 articles · September 15, 2026
- thenewstack.ioOpenAI’s voice model doesn’t think. That’s the point.
2 articles · September 15, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- Voice agents and live captioningFollow
- Speech-to-speech modelsFollow
- Model pricing and inference economicsFollow
Entities
- OpenAIFollow
- GoogleFollow
- Gemini 3.8 Live Extended ThinkingFollow
- Grok Voice Think Fast 2.0Follow
- SynthIDFollow
- GPT-Live-1Follow
- Gemini 3.8 LiveFollow
- Google DeepMindFollow
- Artificial AnalysisFollow
- Gemini Live APIFollow
- Simon WillisonFollow
- SierraFollow