A dev.to field test of 157 planning traces argues teams are hardening tools, memory and orchestration while single-pass decomposition ships plans missing one ordering constraint.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+45
- Incentives60
- Confidence50
One developer's 170-goal sweep sorts almost every planning blocker into three families a linter can name, two of them plain graph properties. Upgrading the planner to GPT-4o left the same pattern behind.
Reality
- Evidence44
- Adoption10
- Hype gap+22
- Incentives50
- Confidence45
Measured five times on the same plan, one safety critic returned five different verdicts and still never approved a seeded defect. It holds because the code downstream never reads the model's severity label.
Reality
- Evidence40
- Adoption12
- Hype gap+8
- Incentives45
- Confidence52
Debashish Ghosal built two LLM verification systems on opposite bets, code and structure, and both did what their designs promised. The blind spot they shared sat in whichever layer each one had decided to trust.
Reality
- Evidence34
- Adoption10
- Hype gap−6
- Incentives45
- Confidence33
PlannerCritic's author tried to inject his own engine. The blocks arrived as feasibility verdicts rather than safety strings, which is a design worth copying and a limit worth reading closely.
Reality
- Evidence44
- Adoption12
- Hype gap+24
- Incentives72
- Confidence41
PlannerCritic's sweeps are cheap enough to gate every release. The harder problem is telling a clean run from a rig that quietly counted nothing.
Reality
- Evidence34
- Adoption11
- Hype gap+24
- Incentives68
- Confidence33
A 157-goal field test of an LLM plan-and-critique loop found failures clustered in three structural families. Swapping in gpt-4o changed the writing, not the dependency graph.
Reality
- Evidence44
- Adoption14
- Hype gap+16
- Incentives55
- Confidence37