A JSON parser benchmark that scores refusal as a pass, and why the column order flips
Seven Python parsers, 300 labelled malformed LLM outputs. On recoverable cases json-repair wins by two; count the 25 unrecoverable ones and it loses by seventeen.
Reality
- Evidence64
- Adoption
- Insufficient
- Hype gap+12
- Incentives78
- Confidence55