Build2 distinct publishers2 min readPublished
Granite Speech 5.0 Turbo CTC drops the language model and three quarters of its output steps. The speed looks real; the accuracy figures are IBM's own.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
The throughput came from removing sequential work, not from removing weights. Earlier Granite encoders emitted 50 characters per second of audio; these emit 12.5 tokens per second [12], one quarter as many output steps for the same recording [2]. Getting to that rate takes three stages of 2x subsampling from a 100 frames-per-second log Mel front end, two of them built into the first two Conformer blocks as strided convolutions [13]. Decoding is non-autoregressive greedy [14], so nothing waits on the token before it, and chunkwise attention keeps the 16-block encoder off the quadratic curve as audio gets longer [15]. Dropping the projector-and-LM arrangement of the previous models [5] is what let that design exist: 470 million parameters is about a seventeenth of the 8-billion-parameter Granite LM that an earlier version of this stack fed [6][2][6].
The licence split is the more interesting arithmetic. The noncommercial weights see roughly 15,000 more hours of natural audio [5] and come back with 0.15 of a point less word error [3], a 3% relative reduction [4]. That gap is also not a clean measurement of what the extra data bought, because the two builds use different tokenizers, SentencePiece for the noncommercial one and BPE for the Apache one [17]. Anyone trying to read the value of GigaSpeech and SPGI Speech hours out of 0.15 points is reading through a second variable.
One set of numbers in the release is not IBM's own scoring of its own run. On the FFASR far-field leaderboard as of 25 August 2026, IBM's post says the Apache model ranked ninth on accuracy and the noncommercial model fifth, and that both were the fastest two entries [9]. Far-field is where small encoders usually come apart, so that is the more informative result. It is still a leaderboard.
Where this lands is not the H200 figure. IBM says the models suit speech-to-text on edge devices [11], and the release ships a WebGPU streaming demo that runs in Chrome or Edge [18]. A 470M encoder with a 12.5 token-per-second output rate fits places an audio-plus-LM stack does not, and that is what deleting the decoder actually bought. The trade only reads as free if the capability that went with it was never on your critical path.
Ranked by verification strength, evidence, and original report placement.
Brian Kingsbury, George Saon and six other IBM researchers released two compact speech-recognition models on August 25 after removing the language-model decoder that gave earlier Granite Speech releases broader capabilities.
The Granite Speech 5.0 Turbo CTC models contain 470 million parameters each and focus solely on turning spoken English into text.
The encoder-only design provides strong transcription performance, a small memory footprint of only 470M parameters, and over 20x faster throughput than previous Granite Speech models.
The models reach over 12,600 RTFx on an NVIDIA H200 GPU, meaning they can transcribe more than 3.5 hours of speech in one second using batched inference.
The new models are encoder-only, unlike prior Granite Speech models which comprise an acoustic encoder, projector, and Granite LM with LoRA adapters.
IBM's earlier Granite Speech research connected a Conformer acoustic encoder to 2-billion- and 8-billion-parameter Granite language models for transcription and speech translation.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed vendor disclosure, one official third-party leaderboard
The architectural account is unusually specific and internally consistent across both sources: parameter count, 16 Conformer blocks, self-conditioning placement, three-stage 2x subsampling, tokenizers, decoding scheme and licence/training-data split. Quantitative performance evidence is weaker: the headline throughput and the 4.85%/5.00% WERs are IBM's own runs, self-labelled unofficial, with official OpenASR results deferred to the leaderboard. The FFASR placements are the only figures described as official third-party results, which lifts evidence above pure vendor assertion but stops short of independent production measurement.
Release-stage, no deployment evidence
Adoption signals stop at publication: two models on Hugging Face, a browser demo, self-run public benchmarks and official far-field leaderboard entries. The supplied sources contain no download counts, integrations, customer deployments or production usage disclosures, so measured adoption reflects availability and leaderboard participation only.
Mildly overstated: vendor speed record, undemonstrated in production
The '20x' and 'unprecedented speed' framing rests on IBM's own batched-inference benchmark on an H200, a configuration that flatters throughput and says nothing about single-stream latency or real-world audio. Accuracy claims are self-labelled unofficial. Against that, IBM discloses the trade-offs plainly (lost speech translation and keyword biasing, licence restriction on the stronger model), the FFASR rankings are official, and the runtimewire coverage explicitly discounts the numbers, so the gap is modest rather than severe.
Vendor announcement with commercial licensing steer
The primary source is IBM's own launch post on a distribution platform, publishing its own benchmark charts and superlative framing; the secondary outlet's reporting derives from that newsroom item. The licence design adds a further commercial incentive: the more accurate model is research-only while production users are pointed to the Apache 2.0 variant trained on narrower data. Incentives are disclosed rather than hidden, which is why this sits high but not at the ceiling.
Consistent facts, single originating source
The two sources agree on every material figure and the technical detail is specific enough to be checkable, so factual confidence is solid. It is capped by the cluster resting on one originating announcement, by the key performance figures being vendor-measured and unofficial, and by the complete absence of adoption or third-party workload data.
build
Hugging Face's $13B process puts most teams' model pipeline under a single owner2 distinct publishers
build
Base Compute hands kernel tuning to agents; the carryover claim is the unmeasured part1 distinct publisher
build
A 27B model reportedly beat a license check in 30 minutes. Nobody has seen the binary.1 distinct publisher
build
Qwen 3.8's Apache-licensed 27B is the one you can actually own, and its KV cache is why1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 25, 2026
1 article · August 25, 2026