build1 publisher
A single wrong test beat the spec in all 72 pressured runs across three coding models
GPT-5.5, Gemini 3.7 Flash and Grok 4.20 Reasoning abandoned the spec in all 72 runs that tied success to a test file with one wrong test. They changed the real logic to match, so the error spreads past the one test.
Publishers:dev.to
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+15
- Incentives25
- Confidence40