BuildIndependently confirmed2 publishers2 min readPublished
IBM's 20x speech gain came from deleting the decoder, not adding parameters
Granite Speech 5.0 Turbo CTC drops the language model and three quarters of its output steps. The speed looks real; the accuracy figures are IBM's own.
The Engineer · Build desk
What happened
- IBM researchers including Brian Kingsbury and George Saon released two speech models on 25 August, built by removing the language-model decoder used in earlier Granite Speech releases.
- The encoder-only build is 470 million parameters and, IBM says, runs more than 20 times the throughput of the previous Granite Speech models.
- IBM measures over 12,600 RTFx on one NVIDIA H200 with batched inference, or more than 3.5 hours of speech transcribed in a second.
- Aggregate word error rate is 5.00% for the Apache 2.0 model and 4.85% for the noncommercial one on OpenASR public English short-form sets, both reported as unofficial.
- The Apache model trained on about 60,000 hours of public English audio plus synthetic material; the noncommercial one adds GigaSpeech and SPGI Speech for roughly 75,000 hours.
Why it matters
- decision Fleet sizing changes shape when a single GPU-day covers roughly 302,400 hours of recorded audio: the binding question stops being how many cards and becomes whether the workload needs anything...
- constraint Speech translation and keyword biasing went out with the decoder, so anyone who needs vocabulary steering or non-English output stays on the LM-equipped line or builds a second pass around this one.
- cost Commercial deployments pay about 3% relative word error for staying inside the permissive licence, which is cheap unless transcripts feed something that compounds errors downstream.
- exposure Capacity plans built on the 12,600 RTFx figure rest on IBM's own benchmark harness rather than an outside production run, and that risk sits with whoever signs the hardware order.
The throughput came from removing sequential work, not from removing weights. Earlier Granite encoders emitted 50 characters per second of audio; these emit 12.5 tokens per second [11], one quarter as many output steps for the same recording [19]. Getting to that rate takes three stages of 2x subsampling from a 100 frames-per-second log Mel front end, two of them built into the first two Conformer blocks as strided convolutions [15]. Decoding is non-autoregressive greedy [16], so nothing waits on the token before it, and chunkwise attention keeps the 16-block encoder off the quadratic curve as audio gets longer [12]. Dropping the projector-and-LM arrangement of the previous models [5] is what let that design exist: 470 million parameters is about a seventeenth of the 8-billion-parameter Granite LM that an earlier version of this stack fed [6][2][23].
The licence split is the more interesting arithmetic. The noncommercial weights see roughly 15,000 more hours of natural audio [22] and come back with 0.15 of a point less word error [20], a 3% relative reduction [21]. That gap is also not a clean measurement of what the extra data bought, because the two builds use different tokenizers, SentencePiece for the noncommercial one and BPE for the Apache one [17]. Anyone trying to read the value of GigaSpeech and SPGI Speech hours out of 0.15 points is reading through a second variable.
One set of numbers in the release is not IBM's own scoring of its own run. On the FFASR far-field leaderboard as of 25 August 2026, IBM's post says the Apache model ranked ninth on accuracy and the noncommercial model fifth, and that both were the fastest two entries [9]. Far-field is where small encoders usually come apart, so that is the more informative result. It is still a leaderboard.
Where this lands is not the H200 figure. IBM says the models suit speech-to-text on edge devices [14], and the release ships a WebGPU streaming demo that runs in Chrome or Edge [18]. A 470M encoder with a 12.5 token-per-second output rate fits places an audio-plus-LM stack does not, and that is what deleting the decoder actually bought. The trade only reads as free if the capability that went with it was never on your critical path.
What to watch
- Official OpenASR leaderboard results including the private test sets, which would either confirm or dent the 5.00% and 4.85% figures IBM published as unofficial.
- Whether IBM reintroduces keyword biasing as an external pass for the encoder-only models, or leaves it as a reason to stay on the LM-equipped Granite Speech line.
- Independent RTFx measurements on hardware other than an H200, particularly the edge devices IBM names as the target.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence63
- Adoption18
- Hype gap+16
- Incentives72
- Confidence68
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Brian Kingsbury, George Saon and six other IBM researchers released two compact speech-recognition models on August 25 after removing the language-model decoder that gave earlier Granite Speech releases broader capabilities.
- [2]
The Granite Speech 5.0 Turbo CTC models contain 470 million parameters each and focus solely on turning spoken English into text.
- [3]
The encoder-only design provides strong transcription performance, a small memory footprint of only 470M parameters, and over 20x faster throughput than previous Granite Speech models.
- [4]
The models reach over 12,600 RTFx on an NVIDIA H200 GPU, meaning they can transcribe more than 3.5 hours of speech in one second using batched inference.
- [5]
The new models are encoder-only, unlike prior Granite Speech models which comprise an acoustic encoder, projector, and Granite LM with LoRA adapters.
- [6]
IBM's earlier Granite Speech research connected a Conformer acoustic encoder to 2-billion- and 8-billion-parameter Granite language models for transcription and speech translation.
- [7]
The noncommercial model scores an aggregate 4.85% WER and the Apache 2.0 model 5.00% WER on the public English short-form test sets used by the OpenASR Leaderboard, reported as unofficial results as of 21 August 2026.
- [8]
The throughput number comes from IBM's benchmark setup rather than a third-party production test; IBM labels the OpenASR charts unofficial and says they were generated with Hugging Face Jobs and the leaderboard's scoring tools.
- [9]
On the FFASR Leaderboard as of 25 August 2026, granite-speech-5.0-470m-turboctc ranked ninth in accuracy and granite-speech-5.0-470m-turboctc-nc ranked fifth, while both were the fastest two models; IBM notes these are official results.
- [10]
The new models give up some capabilities of the LM-equipped models, such as speech translation and keyword biasing.
- [11]
Previous Granite encoders generated 50 characters per second; the Granite 5.0 models generate 12.5 tokens per second.
- [12]
Each model comprises 16 Conformer blocks with self-conditioning at the output of the 8th block, employs chunkwise attention to avoid quadratic scaling with sequence length, and optimizes the CTC loss during training.
- [13]
The Apache 2.0 model is intended for commercial applications and was trained on about 60,000 hours of public English audio plus synthetic material; the noncommercial version adds GigaSpeech and SPGI Speech data for roughly 75,000 hours of natural audio and carries a CC-BY-NC-SA-4.0 licence.
- [14]
IBM says the models are ideal for speech-to-text tasks on edge devices.
- [15]
To get from the 100 frames per second log Mel spectrogram front end to 12.5 tokens per second, IBM uses three stages of 2x subsampling; the second and third are built into the first two Conformer blocks using strided convolutions.
- [17]
The noncommercial model uses SentencePiece tokenization and the Apache 2.0 model uses BPE tokenization, with both tokenizers trained on speech transcripts.
- [18]
The release includes a WebGPU demo of streaming speech recognition that only runs on Chrome or Edge browsers.
- [19]
The new output rate is one quarter of the old one, a 75% reduction in output steps per second of audio.
- [20]
The accuracy gap between the noncommercial and Apache models is 0.15 percentage points of aggregate word error rate.
- [21]
That gap is a 3% relative reduction in word error rate.
- [22]
The noncommercial training set holds about 15,000 more hours of natural audio than the Apache set.
- [23]
At 470M parameters, each new model is roughly one seventeenth the size of the 8-billion-parameter Granite LM used in earlier Granite Speech work.
- [24]
At more than 12,600 RTFx, one H200 processes about 302,400 hours of recorded audio per GPU-day.
Sources
2 independent publishers whose own reporting we read for this story.
- huggingface.coExtremely Fast and Accurate Transcription with Granite Speech 5.0 Turbo CTC
1 article · August 25, 2026
- runtimewire.comIBM cuts the decoder from Granite Speech, says transcription runs 20x faster
1 article · August 25, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- Benchmark Provenance and Self-GradingFollow
- Automatic speech recognitionFollow
- Open model licensingFollow
- Inference throughput and efficiencyFollow
- Small models and edge deploymentFollow
Entities
- CC-BY-NC-SA-4.0Follow
- OpenASR LeaderboardFollow
- FFASR LeaderboardFollow
- ConformerFollow
- Connectionist temporal classificationFollow
- Apache License 2.0Follow
- granite-speech-5.0-470m-turboctcFollow
- George SaonFollow
- RuntimeWireFollow
- granite-speech-5.0-470m-turboctc-ncFollow
- PyTorchFollow
- Nvidia H200Follow
- Granite SpeechFollow
- Hugging FaceFollow
- WebGPUFollow
- SPGI SpeechFollow
- GigaSpeechFollow
- Brian KingsburyFollow
- IBMFollow