Skip to content

BuildIndependently confirmed2 publishers2 min readPublished

IBM's 20x speech gain came from deleting the decoder, not adding parameters

Granite Speech 5.0 Turbo CTC drops the language model and three quarters of its output steps. The speed looks real; the accuracy figures are IBM's own.

The Engineer · Build desk

How we use AISend a correction

What happened

  • IBM researchers including Brian Kingsbury and George Saon released two speech models on 25 August, built by removing the language-model decoder used in earlier Granite Speech releases.
  • The encoder-only build is 470 million parameters and, IBM says, runs more than 20 times the throughput of the previous Granite Speech models.
  • IBM measures over 12,600 RTFx on one NVIDIA H200 with batched inference, or more than 3.5 hours of speech transcribed in a second.
  • Aggregate word error rate is 5.00% for the Apache 2.0 model and 4.85% for the noncommercial one on OpenASR public English short-form sets, both reported as unofficial.
  • The Apache model trained on about 60,000 hours of public English audio plus synthetic material; the noncommercial one adds GigaSpeech and SPGI Speech for roughly 75,000 hours.

Why it matters

  • decision Fleet sizing changes shape when a single GPU-day covers roughly 302,400 hours of recorded audio: the binding question stops being how many cards and becomes whether the workload needs anything...
  • constraint Speech translation and keyword biasing went out with the decoder, so anyone who needs vocabulary steering or non-English output stays on the LM-equipped line or builds a second pass around this one.
  • cost Commercial deployments pay about 3% relative word error for staying inside the permissive licence, which is cheap unless transcripts feed something that compounds errors downstream.
  • exposure Capacity plans built on the 12,600 RTFx figure rest on IBM's own benchmark harness rather than an outside production run, and that risk sits with whoever signs the hardware order.

The throughput came from removing sequential work, not from removing weights. Earlier Granite encoders emitted 50 characters per second of audio; these emit 12.5 tokens per second [11], one quarter as many output steps for the same recording [19]. Getting to that rate takes three stages of 2x subsampling from a 100 frames-per-second log Mel front end, two of them built into the first two Conformer blocks as strided convolutions [15]. Decoding is non-autoregressive greedy [16], so nothing waits on the token before it, and chunkwise attention keeps the 16-block encoder off the quadratic curve as audio gets longer [12]. Dropping the projector-and-LM arrangement of the previous models [5] is what let that design exist: 470 million parameters is about a seventeenth of the 8-billion-parameter Granite LM that an earlier version of this stack fed [6][2][23].

The licence split is the more interesting arithmetic. The noncommercial weights see roughly 15,000 more hours of natural audio [22] and come back with 0.15 of a point less word error [20], a 3% relative reduction [21]. That gap is also not a clean measurement of what the extra data bought, because the two builds use different tokenizers, SentencePiece for the noncommercial one and BPE for the Apache one [17]. Anyone trying to read the value of GigaSpeech and SPGI Speech hours out of 0.15 points is reading through a second variable.

One set of numbers in the release is not IBM's own scoring of its own run. On the FFASR far-field leaderboard as of 25 August 2026, IBM's post says the Apache model ranked ninth on accuracy and the noncommercial model fifth, and that both were the fastest two entries [9]. Far-field is where small encoders usually come apart, so that is the more informative result. It is still a leaderboard.

Where this lands is not the H200 figure. IBM says the models suit speech-to-text on edge devices [14], and the release ships a WebGPU streaming demo that runs in Chrome or Edge [18]. A 470M encoder with a 12.5 token-per-second output rate fits places an audio-plus-LM stack does not, and that is what deleting the decoder actually bought. The trade only reads as free if the capability that went with it was never on your critical path.

What to watch

  • Official OpenASR leaderboard results including the private test sets, which would either confirm or dent the 5.00% and 4.85% figures IBM published as unofficial.
  • Whether IBM reintroduces keyword biasing as an external pass for the encoder-only models, or leaves it as a reason to stay on the LM-equipped Granite Speech line.
  • Independent RTFx measurements on hardware other than an H200, particularly the edge devices IBM names as the target.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence63
Adoption18
Hype gap+16
Incentives72
Confidence68
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Brian Kingsbury, George Saon and six other IBM researchers released two compact speech-recognition models on August 25 after removing the language-model decoder that gave earlier Granite Speech releases broader capabilities.

  2. [2]

    The Granite Speech 5.0 Turbo CTC models contain 470 million parameters each and focus solely on turning spoken English into text.

  3. [3]

    The encoder-only design provides strong transcription performance, a small memory footprint of only 470M parameters, and over 20x faster throughput than previous Granite Speech models.

Sources

2 independent publishers whose own reporting we read for this story.

  1. huggingface.co

    1 article · August 25, 2026

    Extremely Fast and Accurate Transcription with Granite Speech 5.0 Turbo CTC
  2. runtimewire.com

    1 article · August 25, 2026

    IBM cuts the decoder from Granite Speech, says transcription runs 20x faster

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

Entities

Loading related stories