Skip to content

BuildReports disagree6 publishers3 min readPublished Updated

Google splits transcription in two, and quietly absorbs your cleanup layer

Gemini 3.5 Transcribe ships as a sub-second streaming endpoint and a batch one, with filler-word removal and formatting inside the model. The catalog and the pricing page have not caught up.

The Engineer · Build desk

How we use AISend a correction

Photograph accompanying Google splits transcription in two, and quietly absorbs your cleanup layer
Photo: runtimewire.com

What happened

  • Google's new Gemini 3.5 Transcribe serves interactive use through the Live API with bidirectional streaming at sub-second latency, under the identifier gemini-3.5-transcribe-live.
  • Filler-word removal, self-correction handling and text formatting are described as part of the model's output rather than a downstream step.
  • Runtimewire reports the Gemini model catalog and pricing page still list Flash and Live Translate but not Transcribe, and carry no rate for it.

Why it matters

  • capability A hand-maintained cleanup pass can leave the stack, and the formatted string now arrives at the agent boundary instead of being assembled behind it.
  • constraint Anything that needs the better error rate has to run off the batch endpoint, which means live agents and their own call records stop being the same artefact.
  • cost With no published per-minute rate, the engineering hours saved on post-processing cannot yet be netted against what the model bills.
  • contradiction One publisher prints the identifiers and error rates while the other says the announcement supplied neither, so a procurement review has to decide which artefact it is buying against.

The layer coming out of the stack is usually small and annoying to own: a filler-word list, a rule for collapsing repeated tokens, often a second model call that rewrites the raw string into something a downstream parser will accept. Google now describes that work as part of transcription itself, including self-corrections like "let's meet Tuesday, no, Wednesday" and automatic formatting [4]. The maintenance goes away, and so does the unedited string. Runtimewire reports that Google's audio transcription documentation does not establish separate verbatim and smart modes for the model [18]. If you retain transcripts for dispute resolution, clinical notes or QA scoring, settle that before the migration rather than after.

Then the accuracy split, which is the part the two endpoints make unavoidable. Google cites Artificial Analysis at 4.0% word error rate streaming and 2.6% non-streaming [5]. That is 1.4 points, or about 54% more errors on the interactive path [14]: roughly one wrong word in 25 live, against one in 38 for recorded audio [16]. The batch endpoint is also the one carrying speaker attribution and word-level timestamps [3]. Post-call analytics gets the better model and the richer output; the live agent gets neither, which argues for transcribing twice when the transcript is the record of the conversation and not just its input.

Multilingual deployments should size against a different figure again. On FLEURS the model is reported at 5.04% non-streaming and 5.50% streaming [10], and 5.04 is roughly 1.9 times the headline 2.6 [15]. Google says the model auto-detects and transcribes more than 85 languages [6]; runtimewire says the published documentation does not establish that 85-language count [18]. Two publishers, one product, and the gap sits in the exact place a procurement review looks.

That gap is wider than language counts. Runtimewire, writing about the August 26 announcement, says it disclosed no pricing, no benchmarks and no public model identifier, and that the Gemini catalog lists Gemini 3.5 Flash and Gemini 3.5 Live Translate but not Transcribe, with no price on the pricing page [17][8]. Google's own post names both identifiers, gemini-3.5-transcribe-live on the Live API and gemini-3.5-transcribe on the Interactions API [2][3], and publishes the error rates [5]. You can write the call. You cannot yet write the invoice, and Google already runs Cloud Speech-to-Text while general-purpose Gemini models accept audio [11], so this is a choice among its own paths as much as against a specialist vendor.

The feature that actually decides pipeline viability is the narrowest documented one. Wrong order IDs and phone numbers make a transcript useless to downstream automation even when the prose around them is clean [13], and Google's claim to handle alphanumeric entities in noise [12] rests on custom vocabulary that runtimewire says was demonstrated in a simulated interface carrying a warning that compatibility and availability vary [19].

What to watch

  • A Gemini 3.5 Transcribe line appearing on Google's pricing page with a per-minute or per-token rate, which is the first point a migration can be costed.
  • Documentation of a verbatim mode, or its absence, for teams that must retain an unedited transcript.
  • An independent reproduction of the 4.0% streaming WER on noisy audio containing order IDs and postal codes.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence55
Adoption35
Hype gap+25
Incentives70
Confidence60

Perspective Coverage

6 publishers
Builder
Builder 48%
Operator
Operator 36%
Investor
Investor 16%
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Google introduced Gemini 3.5 Transcribe, described as its most precise speech-to-text model yet, converting raw audio directly into accurate, polished, formatted text.

  2. [2]

    Real-time use is served by continuous bidirectional streaming with sub-second latency via the Live API using the model identifier gemini-3.5-transcribe-live.

  3. [3]

    Pre-recorded audio, meetings and call logs are handled with speaker attribution and word-level timestamps via the Interactions API using the identifier gemini-3.5-transcribe.

Sources

6 independent publishers whose own reporting we read for this story.

  1. blog.google

    2 articles · August 26, 2026

    Intelligent transcription with Gemini 3.5 Transcribe
  2. blog.vercel.com

    1 article · August 25, 2026

    Gemini 3.5 Transcribe now available on AI Gateway
  3. dev.to

    1 article · August 28, 2026

    Google Expands Gemini With Transcription, Video Controls and a Voice-First App Experience
  4. mezha.net

    1 article · August 27, 2026

    Google оновила Gemini Audio: нові моделі покращують транскрипцію та голосові діалоги
  5. runtimewire.com

    2 articles · August 26, 2026

    Google announces Gemini 3.5 Transcribe, leaves its API identity unclear
  6. the-decoder.com

    2 articles · August 27, 2026

    Google's Gemini 3.5 Transcribe turns speech to text in 85 languages while auto-correcting your verbal stumbles

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Loading related stories