Build1 publisher3 min readPublished
Lokutor's Oído transcribes any English sentence on an emulated $5 ESP32-S3
Lokutor released Oído, open-source transcription of any English sentence on a $5 ESP32-S3, with every figure so far computed off the board. Whether hardware built for command lists can take open-ended speech now depends on how the real boards measure.
The Engineer · Build desk

What happened
- Oído runs NVIDIA's openly licensed Conformer-CTC Small, a 13-million-parameter model that Lokutor did not retrain or distill.
- Lokutor says Oído makes 2.3 to 2.6 times fewer LibriSpeech errors than Espressif's own recognizer, which matches speech against up to 200 predefined commands.
- Across 14 noisy test conditions, Oído averages an 8.4% word error rate against 12.1% for Whisper tiny running at full precision on a laptop.
- Speed is estimated from instruction counts at 0.7 to 0.95 times real time, and Lokutor says it will publish measured numbers once boards arrive this week.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- exposure With recognition on the chip, a device keeps understanding speech when its connection drops and stops sending users' voices to a server, the two problems besides cost that Lokutor attributes to cloud streaming.
- constraint The int8 model fills 87.5% of the board's 16 MB of flash, so firmware with substantial application code needs the int4 build and the 6 MB it leaves free.
- constraint Text arrives about 3 seconds after a short command ends, a long wait for an interface where the user expects the device to react at once.
- cost A product with closed firmware has to meet GPLv3 terms or buy a commercial licence from Lokutor, and the int4 model adds CC-BY-SA-4.0 share-alike terms.
A 14 MB model on a chip with 512 KB of fast internal memory is 28 times larger than the memory where compute is cheap [4][1]. The weights stay in flash. The ESP32-S3 reads flash through a small cache at a few tens of megabytes per second [5]. "The obstacle isn't arithmetic, it's memory bandwidth," Lokutor wrote [6].
Lokutor's new C inference engine for the chip's vector instructions is designed around that limit [7]. Its integer matrix kernels do 16 multiply-adds per instruction. Attention runs in integers too, softmax included, so the whole model stays in int8 [7]. The schedule is the clever part. Audio goes through in blocks of 64 frames, and each weight is read from flash once per block instead of once per frame [8]. Weight traffic fell from 18 to 7.5 MB per second of audio, a 58% cut, with both cores working in parallel [8][2][7].
Quantisation cost less than the post's own "about a tenth of a point". The int8 model scores 3.70% on LibriSpeech test-clean against 3.68% at full precision, a 0.02-point gap [11][5]. The post's text does not give an error rate for the smaller int4 build. An optional on-chip language model brings the reported error rates down to 3.3% and 7.2% [13].
The benchmarks describe Lokutor's test audio, so whether they hold for a product depends on how close that product's audio is. The noise result, about 31% fewer errors than Whisper tiny [4], applies to rooms that sound like the car, kitchen and cafeteria recordings, background chatter and echo Lokutor built its conditions from [10]. The Espressif comparison is harder to read. A command matcher is built for its own list, and LibriSpeech is the standard English benchmark [9][20]. A product that only ever needs its command list should test both recognisers on its own phrases.
Lokutor argues that a device streaming audio to a GPU "pays for inference for as long as it exists" [21]. Its verification is careful. Every transcript and accuracy figure comes from the exact arithmetic of the on-chip engine, and the firmware reproduces the same transcripts word for word in Espressif's QEMU emulator [14]. A host build runs the chip's integer math on a laptop, with a live microphone demo [18]. At the slow end of the speed estimate, about 5% margin remains before processing a second of audio takes a full second [15][6]. Oído is what cooks in a Spanish kitchen call out to confirm an order [19]. So far its firmware has done its listening inside QEMU [14].
I think the host build is enough to start a prototype. Hardware needs an ESP32-S3-DevKitC-1 N16R8 board and an INMP441 microphone [18], and the model handles only English for now [16].
What to watch
- Measured speed on a physical ESP32-S3 board; Lokutor says boards arrive this week and it will publish the numbers.
- An error rate for the int4 build, the version that leaves room for application code in flash.
- The technical paper Lokutor says it is writing, including how the Espressif command recognizer was scored on LibriSpeech.