Build2 distinct publishers3 min readPublished
Gemini 3.5 Transcribe ships as a sub-second streaming endpoint and a batch one, with filler-word removal and formatting inside the model. The catalog and the pricing page have not caught up.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
The layer coming out of the stack is usually small and annoying to own: a filler-word list, a rule for collapsing repeated tokens, often a second model call that rewrites the raw string into something a downstream parser will accept. Google now describes that work as part of transcription itself, including self-corrections like "let's meet Tuesday, no, Wednesday" and automatic formatting [4]. The maintenance goes away, and so does the unedited string. Runtimewire reports that Google's audio transcription documentation does not establish separate verbatim and smart modes for the model [12]. If you retain transcripts for dispute resolution, clinical notes or QA scoring, settle that before the migration rather than after.
Then the accuracy split, which is the part the two endpoints make unavoidable. Google cites Artificial Analysis at 4.0% word error rate streaming and 2.6% non-streaming [5]. That is 1.4 points, or about 54% more errors on the interactive path [1]: roughly one wrong word in 25 live, against one in 38 for recorded audio [3]. The batch endpoint is also the one carrying speaker attribution and word-level timestamps [3]. Post-call analytics gets the better model and the richer output; the live agent gets neither, which argues for transcribing twice when the transcript is the record of the conversation and not just its input.
Multilingual deployments should size against a different figure again. On FLEURS the model is reported at 5.04% non-streaming and 5.50% streaming [7], and 5.04 is roughly 1.9 times the headline 2.6 [2]. Google says the model auto-detects and transcribes more than 85 languages [8]; runtimewire says the published documentation does not establish that 85-language count [12]. Two publishers, one product, and the gap sits in the exact place a procurement review looks.
That gap is wider than language counts. Runtimewire, writing about the August 26 announcement, says it disclosed no pricing, no benchmarks and no public model identifier, and that the Gemini catalog lists Gemini 3.5 Flash and Gemini 3.5 Live Translate but not Transcribe, with no price on the pricing page [10][11]. Google's own post names both identifiers, gemini-3.5-transcribe-live on the Live API and gemini-3.5-transcribe on the Interactions API [2][3], and publishes the error rates [5]. You can write the call. You cannot yet write the invoice, and Google already runs Cloud Speech-to-Text while general-purpose Gemini models accept audio [14], so this is a choice among its own paths as much as against a specialist vendor.
The feature that actually decides pipeline viability is the narrowest documented one. Wrong order IDs and phone numbers make a transcript useless to downstream automation even when the prose around them is clean [16], and Google's claim to handle alphanumeric entities in noise [15] rests on custom vocabulary that runtimewire says was demonstrated in a simulated interface carrying a warning that compatibility and availability vary [13].
Ranked by verification strength, evidence, and original report placement.
Google introduced Gemini 3.5 Transcribe, described as its most precise speech-to-text model yet, converting raw audio directly into accurate, polished, formatted text.
Real-time use is served by continuous bidirectional streaming with sub-second latency via the Live API using the model identifier gemini-3.5-transcribe-live.
Pre-recorded audio, meetings and call logs are handled with speaker attribution and word-level timestamps via the Interactions API using the identifier gemini-3.5-transcribe.
The model handles self-corrections (Google's example: "let's meet Tuesday, no, Wednesday"), removes filler words such as ums and ahs, and auto-formats text.
As measured by Artificial Analysis, the model achieves an average word error rate of 4.0% for streaming and 2.6% for non-streaming use cases.
Compared with Google's previous transcription model Chirp 3, time to final transcription improves by 70%, as measured by Artificial Analysis.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 26, 2026
1 article · August 26, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
invest
Google Ships Flash Instead of Pro While OpenAI Loses Its Two Best Operators1 distinct publisher
leadership
Data center opposition is now a siting cost, and the industry is pricing it as a PR line1 distinct publisher
build
Gemini 3.7 Flash goes GA on one model layer, and that is the actual news1 distinct publisher
build
Armenian ASR leaderboard: closed models take the top eight, then lose the domains that matter1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed vendor numbers, no independent verification
The cluster contains unusually specific quantitative claims (4.0%/2.6% WER, 5.50%/5.04% FLEURS, 70% time-to-final improvement) plus named model identifiers and access channels, all from the vendor's own post and attributed to Artificial Analysis without a locatable report, test set or methodology. The only independent artifact is a second publisher's audit of Google's catalog, pricing page and documentation, which finds those surfaces silent or inconsistent with the announcement. That supports the existence and shape of the release strongly, and the performance claims only at vendor word.
Shipped in Google surfaces, preview for everyone else
The model is already powering shipped Google products (Rambler on Android, the Gemini app on macOS, Antigravity, AI Studio Build mode), which is real deployment at consumer scale, and Google names seven voice-infrastructure platforms plus three customers offering feedback. Against that, third-party access is public preview only, no usage volumes, seat counts or traffic figures appear, and the model is absent from Google's own catalog and pricing page, which blocks the procurement step that converts preview interest into production adoption.
Superlatives and capability list outrun the documented surface
The launch frames the model as Google's most precise speech-to-text yet and lists custom vocabulary, code-switching-adjacent language coverage and multi-speaker attribution, while the buying surface lags: no price anywhere, no catalog entry, documentation that does not establish the 85-language count or vocabulary limits, and a custom-vocabulary demo shown in a simulated interface warning that compatibility and availability vary. The headline 2.6% average is also roughly half the 5.04% multilingual FLEURS figure, and the 70% latency gain is measured against Google's own predecessor. The overstatement is one of scope and readiness rather than fabrication, since concrete numbers and shipped surfaces do exist.
Vendor launch post plus gap-audit framing
The dominant source is Google's own product announcement, which exists to drive developer and enterprise preview signups and to position the model against conventional speech recognition and specialist transcription APIs; it selects benchmarks, the comparison baseline (its own Chirp 3) and the customer testimonials it publishes. The second publisher's incentive runs the other way, toward an omissions-and-unresolved-identity frame, and its claim that no benchmarks or identifiers were disclosed is contradicted by the announcement text, indicating framing pressure on that side too. No disinterested third-party measurement appears in the cluster.
Two sources, one vendor, one auditor, no third-party test
Only two publishers are present and they disagree on a basic question of what was disclosed, so the assessment rests on the vendor's text for capability and performance and on a single independent check for the documentation and pricing gaps. The structural facts (two endpoints, in-model cleanup, preview availability, no published price) are corroborated or directly readable in the sources and are held with high confidence; the accuracy and latency figures and the language-count claim are held with lower confidence pending independent measurement or updated Google documentation.