Build1 publisher3 min readPublished
The 810-cell sweep that never read a sentence: an author audits his own 100% claim
appgen's README says an exhaustive sweep is why its 100% verification numbers are trustworthy. Its author now points out that none of the 810 cells contains English.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- appgen takes a typed sentence such as "I want a support desk system with priorities, comments, search and closing tickets", writes a single dependency-free Python file, starts it on a private port, exercises every feature requested over real HTTP, and only then shows it to the user; no language model does the generating.
- The tool's README says the verification sweep covers "all 810 domain x feature x family cells" and that this "is why the 100% claims above are trustworthy".
- A generation run takes about 2 ms of CPU and no GPU.
- The code lives in a private research repo, so the post contains no link to it.
- The sweep lives in experiments/028-relations/exhaustive.py and turns on a six-line cells() generator.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
The author of appgen, a deterministic sentence-to-application generator, published a post dismantling his own README, which states that a verification sweep covers "all 810 domain x feature x family cells" and that this "is why the 100% claims above are trustworthy" [2]. The arithmetic holds, but according to the author the sweep touches the user's sentence zero times, and the sentence is the whole interface [11][13].
What appgen does, on the author's account: you type something like "I want a support desk system with priorities, comments, search and closing tickets" and it writes a single dependency-free Python file, starts it on a private port, exercises every requested feature over real HTTP, and only then shows it to you [1]. No language model generates the code, and the run costs about 2 ms of CPU and no GPU [1][3]. The code sits in a private research repo, so there is nothing to inspect independently [4].
The sweep lives in experiments/028-relations/exhaustive.py and hangs on a six-line generator [5]. For each domain it yields an empty feature set, then each single feature, then the full set: nine configurations [6][7][8]. Ten domains gives 90 cells [6][7][1], and nine emitter families multiply that to 810 [9][2]. The author says he checked the arithmetic rather than trusting the README, and 810 is right [10].
The defect is not the number. A cell is a tuple of a dictionary key and a frozen set of strings; nothing in the generator produces English and nothing downstream of it consumes English, because the emitters are handed a schema and a feature set directly [11][12]. The sweep is candid about this at line 96 of sweep_project, where a comment reads "build_project parses a request; drive the emitters directly instead" [13]. That is defensible test design, and the author defends it: a sweep should test one thing [14]. What is not defensible is a README sentence that borrows the composer's 100% to vouch for the tool a user actually drives.
The component the user hits first is a hand-written keyword planner, measured separately in experiment 041, which was built to ask whether a learned model should replace it [15]. The harness is stricter than a random split: eight phrasing families are held out whole, and slot vocabularies are disjoint across the split, so nothing scores by memorising a phrase [16][17]. That is 146 training clauses and 32 test clauses per fold, eight folds, seven operations [18]. Re-running compare.py, the rules score 0.70, nearest-neighbour retrieval 0.48, and a hybrid 0.68, which is why the planner stayed hand-written [19][21][22][23].
Read the winning column as a user and 0.70 means roughly three requests in ten are read wrong on a phrasing the rules were not written for [3][24]. One of the seven operations is kind, which decides whether you get an HTML app, a JSON API or a CLI [25]. A misread there does not produce a failing test; it produces a passing test for the wrong artefact.
One wrinkle: the post describes the Naive Bayes score of 0.24 as "barely above" a most-common-class floor of 0.29, which is 0.05 below it [20][21][4]. Worth watching whether the per-operation breakdown, cut off in the material as published, shows kind among the weakest operations, and whether the README claim is rescoped or a sweep is added that starts at English rather than at a frozenset.