Science1 publisher2 min readPublished
SuperWhisper's open-weight S1-mini cleans up raw speech transcripts offline on a laptop CPU
SuperWhisper's S1-mini, a 596-million-parameter open-weight model, turns raw ASR transcripts into clean text on a laptop CPU with no network calls. Keeping the audio itself off the cloud still depends on the speech recognizer that runs in front of it.
The Scientist · Science desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- SuperWhisper launched S1-mini on August 19, 2026, alongside two cloud-hosted siblings: S1-Voice for speech recognition and S1-Language for heavier formatting.
- On 7,519 held-out English cases from 104 transcripts, SuperWhisper's model card reports 94.8% token accuracy and an 11.6% text-edit error rate.
- A control line added to each input sets styling, structure and context, and the model will only format a list when it finds at least three real items.
- Because it is fine-tuned from Qwen3-0.6B, S1-mini outputs an empty think block and no text unless the chat template sets enable_thinking=False.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
- capability For the cleanup step, a team can verify offline behaviour from the published weights itself, and does not have to take SuperWhisper's word that user data stays out of training.
- exposure Pipelines that test only for crashes or exceptions will pass S1-mini's silent empty output, so a misconfigured deployment can ship blank transcripts to users.
- decision Teams weighing a switch from a prompted 7B or 8B chat model will have to score both on their own transcripts to test the reported parity.
- contradiction KDnuggets headlines S1-mini as built 'just for transcription', but the model card describes a formatter whose input is a transcript some other recognizer already produced.
The design goal is deliberately narrow. The model card calls S1-mini "ruthlessly obedient" and says it will never add content the speaker didn't say, correct a fact, soften profanity or rewrite dialect [5]. Most small models are tuned to be broadly helpful. This one is tuned to format what was said and nothing more [4]. I think that is the right target for a transcript formatter, because a cleanup pass that corrects a speaker's facts adds errors the recognizer never made.
The best test in the evaluation is the one with nothing in it. Given input made only of filler noise, the model returned an empty string 98.6% of the time, according to the model card [10]. That test checks directly for the failure that does most damage to a transcript: invented speech filling a silence. Fewer than 1% of generations looped or cut off early [9].
The accuracy figures need a denominator. The 7,519 held-out cases work out to about 72 per transcript [1]. Cases cut from one transcript share a speaker and a topic, so the independent sample is closer to 104 than to 7,519 [7]. The two figures are also not complements. 100 minus 94.8 is 5.2, while the reported text-edit error rate is 11.6% [2], so the two metrics count errors in different ways. Email output got its own checks: the greeting line was identified correctly 99.3% of the time and the sign-off 97.9% [8].
On size, the model card counts 596 million unique parameters. The 0.8 billion shown in the Hugging Face sidebar counts the tied embedding weights twice [6]. At that size it runs on a laptop CPU with no GPU [6]. KDnuggets reports that on this task S1-mini matches or beats what general 7B or 8B chat models produce [14]. The thing the summary doesn't tell you is the size of that margin, or how long a CPU takes per transcript.
What to watch
- Independent head-to-head scores for S1-mini against a prompted 7B or 8B chat model on the same held-out transcripts.
- Measured CPU latency per transcript on ordinary laptops, which would settle whether local cleanup keeps pace with dictation.
- Any SuperWhisper evaluation beyond English, or a named local speech recognizer paired with S1-mini for a fully offline pipeline.