Build1 publisher3 min readPublished
The $8 chip did not change. The memory layout did.
A 28.9-million-parameter model now runs on an ESP32-S3 that was topping out at 260K two years ago. The gain came from moving most of the weights into flash and reading 450 bytes per token.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- On July 25, 2026, a repository called esp32-ai hit the front page of Hacker News with a 28.9-million-parameter language model running entirely on an ESP32-S3 microcontroller that costs about eight dollars and has 512KB of SRAM, 8MB of PSRAM and 16MB of flash, generating text at 9.88 tokens per second, roughly reading speed. Tom's Hardware, The Register, Adafruit and CNX Software covered it within a week.
- The previous record for a language model on an ESP32 was about 260 thousand parameters.
- A 28.9M-parameter model quantized to 4 bits needs about 14.9MB; it fits in flash, but flash is far too slow to stream weights through on every token.
- A 14.9MB quantized model in 16MB of flash leaves roughly 1.1MB spare.
- The developer, who goes by slvDev, borrowed the idea of per-layer embeddings from Google's Gemma 3n.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
On July 25, 2026, a repository called esp32-ai reached the front page of Hacker News running a 28.9-million-parameter language model entirely on an ESP32-S3, a microcontroller that costs about eight dollars and carries 512KB of SRAM, 8MB of PSRAM and 16MB of flash [1]. The previous record for a language model on that chip family was about 260 thousand parameters [2], so the question worth asking is not whether it works but what changed, and the answer is where the weights sit.
The arithmetic that used to end this conversation still holds. Quantized to 4 bits, 28.9M parameters need roughly 14.9MB [3], which fits in the 16MB of flash with around 1.1MB spare [4] but is far too slow to stream through on every token [3]. The developer, who goes by slvDev, borrowed per-layer embeddings from Google's Gemma 3n [5]. Embeddings are lookups, not matrix math, and most of a small model's parameters live in them [6]. So the model splits: a dense reasoning core of roughly 560 thousand parameters, about 2 percent of the total [12], runs in RAM every step, while 25 million parameters, roughly 86 percent of the model [13], sit in flash as a memory-mapped table [7].
Generating a token touches about six rows of that table, roughly 450 bytes of flash reads [8]. At the reported 9.88 tokens per second [1] that is about 4.4KB per second of flash traffic [14]. SRAM holds activations and normalization weights, PSRAM holds the core and the output head, flash holds the embedding table [9]. The processor is the same processor. The hierarchy is new.
The lineage is worth naming, because it shows the gain is cumulative engineering rather than a jump in parts. Most microcontroller LLM ports descend from Andrej Karpathy's llama2.c and from DaveBben's esp32-llm, which in 2024 ran a 260K model at 19 tokens per second on the same chip family [10]. Two years later the parameter count is about 111 times higher [11] on hardware that costs the same [1], at a lower token rate. That trade, throughput for capacity, is the whole story.
Sort the field honestly and it splits into tiers where weights and compute both live on the chip, versus where the chip only does the math [15]. The most interesting near-miss belongs to the second category. On August 5, 2026, a developer going by cyfrit posted p-for-llm, a 180.9-million-parameter mixture-of-experts model on an ESP32-P4 with 12 layers, 29 experts per layer, top-1 routing, ternary BitNet-style quantization and a vocabulary pruned from Qwen, on a board costing six to ten dollars [16]. It runs at roughly 9 tokens per second, follows simple ChatML instructions, and has a tool-calling demo the author himself describes as frequently going off the rails [17]. The asterisk: about 44MiB of weights are pushed over USB at startup, with SD-card autonomy still future work [18]. That is more than the 24MB of combined flash and PSRAM on the S3 [19], which is exactly why it needs the umbilical. It was trained on about 12 billion tokens on a single consumer GPU [20], and it got two points on Hacker News [16].
Watch whether p-for-llm's SD-card path lands, since that is the line between a demo with a host attached and a device. Watch also whether per-layer embeddings show up in other ports, because an architectural trick that anyone can copy sets a new floor rather than a record.