Build1 publisher3 min readPublished
Model's plan rescued zero of the requests lost by olo's empty planner seat
Removing the qwen3-coder:30b planner from the olo app stack lost no working build across 88 requests, its developer reports. The planner is the stack's only neural part, and on this corpus it showed no measurable gain, in figures from a private repository.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- The model's job is a seven-field build plan, with six fields picked from a closed vocabulary and the seventh being the request passed through unchanged.
- A new --planner defer option declines every field and sends the bare sentence through the same validator and command-line assembly the model's plan uses.
- Bare appgen and olo with the planner seat empty returned the same verdict on every one of the 88 requests.
- Requests appgen cannot build are counted as refusals, because the tool declines anything outside its grammar instead of guessing.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision On this corpus, an olo operator could drop the 30B model from the planner seat and lose no working app the model was producing.
- constraint The result covers only requests inside appgen's closed grammar, so it leaves untested the open-domain proposal job the README still reserves for a model.
- precedent Passing an explicit default can disable inference in the layer below, so a harness that fills in flags can end up measuring a different program from the one it names.
The branch that makes the comparison valid is small. A declined field is now left off appgen's command line instead of being filled with appgen's default [9]. appgen reads the sentence for its program kind and language only when those flags are absent, so passing --kind web switches that reading off [9]. A harness that filled in web as a courtesy would have measured appgen with its own router disabled and reported the result as appgen [9].
Earlier measurements could not separate the planner from the code around it. olo puts a local language model in front of appgen [1]. It also wraps appgen in a router, a plan validator, an entity read-back and the code that assembles appgen's command line, and both prior rounds had a model on both sides of the comparison [16]. Any margin from either round belonged to the planner plus all of that.
The author calls the planner the only place in the architecture where a neural network does real work [3]. This round holds every line of code downstream of the plan constant and changes only who fills the seat [5]. "Everything olo does that is not the planner costs nothing and buys nothing once the seat is empty," the author wrote [15]. With a local qwen3-coder:30b planning [12], the model's plan rescued zero requests that the empty seat lost [5]. The author reports that this holds under all four ways of scoring the run, and that it is the only result of the round that does [6]. The available text of the post stops before giving the count in the other direction.
An older figure points the same way. A previous round ran bare appgen over the same corpus and recorded 21 of 88, against 15 for the full stack [7]. That is six fewer builds with the model in place [1]. The older run predates the repaired seam, so it is weaker evidence than the paired result.
For the zero to carry over to another pipeline, the plan has to look like this one. Six of the seven fields are picks from a closed vocabulary [2], and for two of them, kind and language, appgen already infers the answer from the sentence [9]. A planner choosing from a menu the builder can already read off the request has little left to add. The repository's README names the one job it says still has no measured non-neural replacement: proposing a program in the open domain [4]. I think that job is the remaining case for keeping a model, and a closed-vocabulary plan does not test it. Every figure comes from the author's results files and commands in a private research repository [17].
The guard tests deserve credit. Three tests protect the absent-flag branch, and the author proved each one by planting the defect it names and checking that the test went red [10]. The first plant showed the tests had been appended below sys.exit(main()) and had never run [10]. "A file of assertions that cannot execute looks exactly like a file of assertions that pass," the author wrote [11].
What to watch
- The count in the other direction: how many requests the qwen3-coder:30b plan lost that the empty seat built.
- A run on open-domain program proposals, the one job olo's README says has no measured non-neural replacement.
- Whether the author publishes the code or results files from the currently private repository so the 88-request run can be reproduced.