Relipa's AI development workflow wraps a coding agent in six stages and fixes the tests before any code is written. Most of its published detail covers the workspace plumbing that runs before a model is ever called.
Reality
- Evidence35
- Adoption8
- Hype gap+5
- Incentives
- Insufficient
- Confidence40
System Design One's newsletter says AI agents on big codebases need written specs, citing a change that passed tests but emailed an opted-out user. Most of its sample spec just records what the old system already does, and a team can write that document before buying any agent platform.
Publishers:newsletter.systemdesign.one
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+25
- Incentives
- Insufficient
- Confidence45
One operator ran spec-driven development at full BMAD weight for four weeks on an autonomous coding repo. The planning stage produced 1,782 files and 16 MB of artifacts, and a single story burned around thirty million tokens.
Reality
- Evidence34
- Adoption20
- Hype gap+20
- Incentives45
- Confidence38
A dev.to writeup wires CPU and RAM telemetry from a Manifest V3 extension down to a standalone Windows binary. The first of its seven written rules: nothing may print to stdout, where Chrome expects a 4-byte length prefix.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+20
- Incentives55
- Confidence45
A dev.to post groups spec-driven development frameworks by what they do with the markdown once the LLM has written the code, and argues for deleting it on the grounds that code is the less ambiguous description.
Reality
- Evidence34
- Adoption20
- Hype gap+22
- Incentives28
- Confidence42
For the first ninety minutes nobody opens an IDE. A mentor has to approve the team's SPEC.md, and that same file is what Antigravity or Cursor reads once the coding starts. It is worth 35 of the 100 points.
Reality
- Evidence42
- Adoption22
- Hype gap+18
- Incentives62
- Confidence55
Offset paging passes a static test suite and then skips rows the moment a merchant inserts a product mid-read. A dev.to case study answers that by freezing the sort key, tie-breaker and error codes in a spec file the agent cannot edit.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+25
- Incentives20
- Confidence50
A month and a half of prototyping an experimental functional language with Codex ended with eight artifacts around the repository and a documentation flow in which ChatGPT writes the prompts and Codex writes the spec.
Reality
- Evidence26
- Adoption12
- Hype gap+22
- Incentives30
- Confidence52
Signadot's argument is that the bill for an unattended coding loop is the iterations it needs times the cost of each one, and that only feedback precise enough to localize a fault shrinks both of those terms at once.
Reality
- Evidence24
- Adoption
- Insufficient
- Hype gap+34
- Incentives80
- Confidence60
Anthropic's playbook says code is no longer the bottleneck, and the spec-driven tools around it each prescribe one fixed sequence of stages. The New Stack argues the process should be data an organization defines itself.
Reality
- Evidence38
- Adoption18
- Hype gap+24
- Incentives55
- Confidence44
SWC's creator fanned a single instruction out to 45 concurrent Labor0 tasks on the ES minifier, then kept review and merge for himself. He says he does not know how long the run took; he was looking at his phone.
Publishers:kdy1.dev
Reality
- Evidence42
- Adoption28
- Hype gap+20
- Incentives82
- Confidence45
A study accepted at GAISS 2026 reports that a specification baseline increased the attributability of reviewers' bug finds without increasing their count, and that on easy tasks the credit for spec-first prompting mostly belongs to reasoning effort.
Reality
- Evidence36
- Adoption
- Insufficient
- Hype gap+8
- Incentives55
- Confidence40
Enforcement is a bash script that counts files against the merge-base with the parent branch, so it bounds one diff at a time, and only if the agent chooses to run it. The 100 is an environment variable's default.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+35
- Incentives
- Insufficient
- Confidence62
The enforcement is a bootstrap prompt the agent checks before every task, so its durability depends on per-harness plumbing. On Hermes Agent, compaction can lose it.
Reality
- Evidence34
- Adoption
- Insufficient
- Hype gap+30
- Incentives
- Insufficient
- Confidence41
Daniel Vaughn's experimental editor treats pseudocode as the durable artifact and generated source as output. The public build sends the whole spec to a model and shows read-only code.
Reality
- Evidence54
- Adoption7
- Hype gap+28
- Incentives44
- Confidence52
Salesforce says builds, tests and reviews validate the diff, not the requirement. Its answer was a specification gate upstream of the code, plus a rule about which questions agents may answer.
Reality
- Evidence45
- Adoption20
- Hype gap+18
- Incentives58
- Confidence48
A dev.to writeup argues long sessions degrade structurally, not linguistically, and proposes research/plan/implement phases with a hard context clear between each.
Reality
- Evidence28
- Adoption12
- Hype gap+32
- Incentives22
- Confidence42
A dev.to comparison of OpenSpec and GitHub Spec Kit argues chat degradation is structural. The interesting part is not the diagnosis but how differently the two tools file the cure.
Reality
- Evidence22
- Adoption
- Insufficient
- Hype gap+48
- Incentives66
- Confidence30