Build1 publisher3 min readPublished
Whisper's 300ms floor is architecture, not a bug: when live voice needs streaming ASR
A dev.to writeup argues Whisper is a batch model retrofitted for streaming. If your response budget is under a second on mid-range phones, that is a build decision, not a tuning exercise.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Voice-first apps became mainstream in 2024-2026, including voice agents, real-time captions, live meeting transcription and voice notes with instant feedback.
- Developers building live voice apps typically look at Whisper first (whisper.cpp, faster-whisper); it is the de facto default for on-device speech-to-text.
- For batch transcription (upload audio, wait, get transcript) Whisper handles interview, meeting and lecture audio accurately and reliably.
- In live scenarios where a user speaks and expects an instant response, Whisper shows a minimum delay of 300-500 ms, often up to 1-2 seconds; words get sliced at chunk boundaries and the transcript arrives in jerks.
- The post states this is not a Whisper bug but a design mismatch: Whisper is a batch model retrofitted for streaming, while real-time scenarios need a different architecture.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A writeup on dev.to puts a number on something a lot of voice teams have been treating as a tuning problem: used live, Whisper shows a minimum delay of 300-500 ms and often 1-2 seconds, with words sliced at chunk boundaries and transcript arriving in jerks [4]. The post's argument is that this is not a bug but a design mismatch, because Whisper is a batch model retrofitted for streaming [5]. That matters because Whisper, via whisper.cpp and faster-whisper, is the de facto default for on-device speech-to-text [2], and defaults are how latency budgets get spent without anyone deciding to spend them.
The useful part of the post is a taxonomy by lookahead, meaning how much future audio a model sees before it emits text [7]. Infinite lookahead is batch: the whole file is known upfront, which is where Whisper sits [8]. Lookahead of 1000 ms or more is pseudo-streaming, which covers whisper.cpp streaming mode and VAD-based faster-whisper setups [8]. Native streaming lives at 80-320 ms and includes NeMo FastConformer streaming, streaming Conformer variants and Parakeet streaming [8]. Zero lookahead is causal streaming, fastest but usually strictly worse on accuracy [8].
Whisper was trained on 30-second chunks [9]. To make it stream, whisper.cpp runs a rolling window of 500 ms to 3 s and concatenates the outputs of each mini-batch [10]. The side effects are structural: words at chunk borders get cut, repeated or misrecognised; punctuation drifts between chunks; VAD and endpointing need separate tuning because Whisper has no native VAD; and perceived latency equals chunk size plus inference time [11]. On a Snapdragon 662 with base.en at 500 ms chunks, the post measures 700-1500 ms perceived latency [12]. Subtract the chunk and inference alone is eating roughly 200-1000 ms [17], which is the part you cannot shrink by lowering the window.
The VAD wrapper route (Whisper-Streaming, WhisperLive) produces cleaner output with no boundary artifacts, but latency goes up, because you wait for end of speech before transcribing [13]. So the two available workarounds trade opposite failure modes: choppy and early, or clean and late.
Native streaming models are built differently. They are trained knowing they will see only 80-320 ms of future audio, and they cache internal hidden states between chunks so previous audio is not reprocessed [14]. Tokens come out incrementally, so there are no perceived chunk boundaries [15]. The stated cost is accuracy: native streaming models are usually slightly less accurate than their batch counterparts [16]. The post calls that trade-off minimal, but the text available stops mid-sentence before producing a WER figure [16], so treat "minimal" as an assertion, not a measurement.
The decision this sets up is measurable. Native streaming's worst case lookahead, 320 ms, sits at or below Whisper's best case live delay of 300-500 ms [19], and its best case is more than ten times lower than the 1000 ms-plus of pseudo-streaming [18]. If your product can absorb 700-1500 ms on mid-range Android [12], stay on the default and spend the engineering elsewhere. If it cannot, chunk-size tuning will not get you there, because the inference term does not move with it [11].
What to watch: whether anyone publishes accuracy deltas for streaming versus batch on the same audio, on device, rather than the qualitative "slightly less accurate" the post offers [16]. Also worth checking against your own hardware floor: the 700-1500 ms figure is one SoC, one model size, one chunk length [12].