BuildReports disagree6 publishers3 min readPublished Updated
Google splits transcription in two, and quietly absorbs your cleanup layer
Gemini 3.5 Transcribe ships as a sub-second streaming endpoint and a batch one, with filler-word removal and formatting inside the model. The catalog and the pricing page have not caught up.
The Engineer · Build desk

What happened
- Google's new Gemini 3.5 Transcribe serves interactive use through the Live API with bidirectional streaming at sub-second latency, under the identifier gemini-3.5-transcribe-live.
- Filler-word removal, self-correction handling and text formatting are described as part of the model's output rather than a downstream step.
- Runtimewire reports the Gemini model catalog and pricing page still list Flash and Live Translate but not Transcribe, and carry no rate for it.
Why it matters
- capability A hand-maintained cleanup pass can leave the stack, and the formatted string now arrives at the agent boundary instead of being assembled behind it.
- constraint Anything that needs the better error rate has to run off the batch endpoint, which means live agents and their own call records stop being the same artefact.
- cost With no published per-minute rate, the engineering hours saved on post-processing cannot yet be netted against what the model bills.
- contradiction One publisher prints the identifiers and error rates while the other says the announcement supplied neither, so a procurement review has to decide which artefact it is buying against.
The layer coming out of the stack is usually small and annoying to own: a filler-word list, a rule for collapsing repeated tokens, often a second model call that rewrites the raw string into something a downstream parser will accept. Google now describes that work as part of transcription itself, including self-corrections like "let's meet Tuesday, no, Wednesday" and automatic formatting [4]. The maintenance goes away, and so does the unedited string. Runtimewire reports that Google's audio transcription documentation does not establish separate verbatim and smart modes for the model [18]. If you retain transcripts for dispute resolution, clinical notes or QA scoring, settle that before the migration rather than after.
Then the accuracy split, which is the part the two endpoints make unavoidable. Google cites Artificial Analysis at 4.0% word error rate streaming and 2.6% non-streaming [5]. That is 1.4 points, or about 54% more errors on the interactive path [14]: roughly one wrong word in 25 live, against one in 38 for recorded audio [16]. The batch endpoint is also the one carrying speaker attribution and word-level timestamps [3]. Post-call analytics gets the better model and the richer output; the live agent gets neither, which argues for transcribing twice when the transcript is the record of the conversation and not just its input.
Multilingual deployments should size against a different figure again. On FLEURS the model is reported at 5.04% non-streaming and 5.50% streaming [10], and 5.04 is roughly 1.9 times the headline 2.6 [15]. Google says the model auto-detects and transcribes more than 85 languages [6]; runtimewire says the published documentation does not establish that 85-language count [18]. Two publishers, one product, and the gap sits in the exact place a procurement review looks.
That gap is wider than language counts. Runtimewire, writing about the August 26 announcement, says it disclosed no pricing, no benchmarks and no public model identifier, and that the Gemini catalog lists Gemini 3.5 Flash and Gemini 3.5 Live Translate but not Transcribe, with no price on the pricing page [17][8]. Google's own post names both identifiers, gemini-3.5-transcribe-live on the Live API and gemini-3.5-transcribe on the Interactions API [2][3], and publishes the error rates [5]. You can write the call. You cannot yet write the invoice, and Google already runs Cloud Speech-to-Text while general-purpose Gemini models accept audio [11], so this is a choice among its own paths as much as against a specialist vendor.
The feature that actually decides pipeline viability is the narrowest documented one. Wrong order IDs and phone numbers make a transcript useless to downstream automation even when the prose around them is clean [13], and Google's claim to handle alphanumeric entities in noise [12] rests on custom vocabulary that runtimewire says was demonstrated in a simulated interface carrying a warning that compatibility and availability vary [19].
What to watch
- A Gemini 3.5 Transcribe line appearing on Google's pricing page with a per-minute or per-token rate, which is the first point a migration can be costed.
- Documentation of a verbatim mode, or its absence, for teams that must retain an unedited transcript.
- An independent reproduction of the 4.0% streaming WER on noisy audio containing order IDs and postal codes.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence55
- Adoption35
- Hype gap+25
- Incentives70
- Confidence60
Perspective Coverage
6 publishers- Builder
- Builder 48%
- Operator
- Operator 36%
- Investor
- Investor 16%
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Google introduced Gemini 3.5 Transcribe, described as its most precise speech-to-text model yet, converting raw audio directly into accurate, polished, formatted text.
- [2]
Real-time use is served by continuous bidirectional streaming with sub-second latency via the Live API using the model identifier gemini-3.5-transcribe-live.
- [3]
Pre-recorded audio, meetings and call logs are handled with speaker attribution and word-level timestamps via the Interactions API using the identifier gemini-3.5-transcribe.
- [4]
The model handles self-corrections (Google's example: "let's meet Tuesday, no, Wednesday"), removes filler words such as ums and ahs, and auto-formats text.
- [5]
As measured by Artificial Analysis, the model achieves an average word error rate of 4.0% for streaming and 2.6% for non-streaming use cases.
- [6]
Google says the model automatically detects and transcribes over 85 languages, and attributes speech in pre-recorded audio with timestamps for up to three speakers, with support for more than three speakers experimental.
- [7]
Developers can access Gemini 3.5 Transcribe in the Gemini API in Google AI Studio and in Gemini Enterprise Agent Platform.
- [8]
Runtimewire reports Google's current Gemini model catalog lists Gemini 3.5 Flash and Gemini 3.5 Live Translate but not Gemini 3.5 Transcribe, and that Google's pricing page does not list a Gemini 3.5 Transcribe price.
- [9]
Compared with Google's previous transcription model Chirp 3, time to final transcription improves by 70%, as measured by Artificial Analysis.
- [10]
On the FLEURS benchmark across a set of top languages and locales, the model achieves 5.50% WER in streaming mode and 5.04% WER non-streaming, improving over Chirp 3.
- [11]
Google also has an established Cloud Speech-to-Text service, and general-purpose Gemini models already accept audio.
- [12]
Google says the model shows strong performance in noisy real-world environments and accurately captures alphanumeric entities such as postal codes and order IDs.
- [13]
Errors in phone numbers and order IDs can make a transcript useless for downstream automation even when the surrounding conversation is accurate.
- [14]
The streaming path carries 1.4 percentage points more word error than the non-streaming path, about 54% more errors in relative terms.
- [15]
The FLEURS multilingual non-streaming result of 5.04% is about 1.9 times the headline non-streaming average of 2.6%.
- [16]
A 4.0% word error rate is roughly one wrong word in 25; 2.6% is roughly one in 38.
- [17]
Runtimewire reports the model was announced on Aug. 26 and that the post did not disclose pricing, benchmarks or a public model identifier.
- [18]
Runtimewire reports Google's audio transcription documentation does not establish limits and modes attributed to Gemini 3.5 Transcribe, including an 85-language count, code-switching, separate verbatim and smart modes, or vocabulary limits of 1,000 and 100 terms.
- [19]
Runtimewire reports Google showed custom vocabulary in a simulated interface whose image warned that compatibility and availability vary, and that the announcement does not establish whether custom vocabulary is exposed through an API or another surface.
Sources
6 independent publishers whose own reporting we read for this story.
- blog.googleIntelligent transcription with Gemini 3.5 Transcribe
2 articles · August 26, 2026
- blog.vercel.comGemini 3.5 Transcribe now available on AI Gateway
1 article · August 25, 2026
- dev.toGoogle Expands Gemini With Transcription, Video Controls and a Voice-First App Experience
1 article · August 28, 2026
- mezha.netGoogle оновила Gemini Audio: нові моделі покращують транскрипцію та голосові діалоги
1 article · August 27, 2026
- runtimewire.comGoogle announces Gemini 3.5 Transcribe, leaves its API identity unclear
2 articles · August 26, 2026
- the-decoder.comGoogle's Gemini 3.5 Transcribe turns speech to text in 85 languages while auto-correcting your verbal stumbles
2 articles · August 27, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.