Skip to content

Build1 publisher3 min readPublished

A JSON parser benchmark that scores refusal as a pass, and why the column order flips

Seven Python parsers, 300 labelled malformed LLM outputs. On recoverable cases json-repair wins by two; count the 25 unrecoverable ones and it loses by seventeen.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • jsonshim is a small stdlib tool, maintained by the article's author, that pulls JSON out of a language model's reply.
  • The author built a 300-case conformance suite for that job, named MALFORMED-300.
  • The author had previously published results for his own parser on the suite: 18 failures, five of them values invented where the model had produced nothing.
  • The benchmark covered seven parsers and 300 labelled cases on Python 3.12.3, one run each, with nothing tuned afterwards.
  • The author reports two columns because the libraries do not share a design goal, and one column would quietly punish some of them for a decision their authors made on purpose.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

The maintainer of jsonshim, a small stdlib tool that pulls JSON out of a language model's reply [1], built a 300-case conformance suite called MALFORMED-300 and ran seven Python parsers through it once each, on Python 3.12.3, with nothing tuned afterwards [2][4]. The transferable part is not the ranking but the scoring rule: 25 of the 300 cases are unrecoverable, the model having emitted only `{` or a bare `<redacted>` placeholder, and a library passes those only by refusing, with `{}` graded a failure rather than a near miss [6]. The author reports two columns, on the grounds that these libraries do not share a design goal and a single number would quietly punish some of them for a deliberate choice [5]. The narrow column drops the unrecoverable cases and asks only what a parser returns when there is something to return [7]. There, json-repair beats jsonshim 264 to 262 [8]. Widen it and the order inverts: jsonshim declines 20 of the 25 unrecoverable cases, json-repair declines 1, because json-repair is built to always hand you something [9]. That is 265 of 300 against 282 [1][23], a 17-case swing produced entirely by refusal policy [2] on top of a two-case deficit in recovery [8]. The author's own framing is that always returning is a legitimate design, arguably the right one when you would rather render an empty card than an error, and the wrong one when the parsed object goes into a database or a tool call [10]. The mechanism is that an invented value does not raise and does not log; it arrives looking exactly like data [11]. The worst case in the suite is `{"user_id": <redacted>}` parsed into `{"user_id": "<redacted>"}`, a redaction turned into a user id [12]. jsonshim also fails that one, and the author left it failing on the stated grounds that patching it after seeing the score would convert 94.0% from a measurement into a claim [13][3]. Four of the seven are doing a different job. json5, pyjson5, demjson3 and dirtyjson are dialect parsers reading a looser grammar than strict JSON [14], so they clear grammar categories - unquoted keys 25/25 for all four, trailing commas 20/25, single quotes 16/25, comments 16/25 - and score a flat zero on code fences, prose wrappers and truncation [15][16]. A reply that wraps a fenced object in reasoning text has not produced malformed JSON; locating the payload and parsing the payload are separate problems, and only two of the seven do both [17]. On that basis the author argues the standing advice to "just use json5" answers a question most LLM output does not ask [18]. The per-category grid also splits the two extractors: json-repair takes wrappers 25 to 17, jsonshim takes Python literals 25 to 21 [19]. The harness is the reason to take any of this seriously. Each library is called through its own documented entry point and lenient mode, in an adapters file kept to about a hundred lines so it can be audited [20]; no input is pre-cleaned and every parser sees the identical raw string [21]; the only post-processing normalises library-specific container and sentinel types to plain `dict`, `list` and `None` [22]. Before any third-party row was scored, the harness had to reproduce a previously published jsonshim result exactly, 282/300 with 5 invented values and 4 false refusals, and it did [23]. Ground truth is produced by construction, with a checkpointing renderer emitting the expected value before the malformed text exists, so no parser was ever consulted about the right answer [24]. Two things to watch. The arithmetic is tight enough to check: 18 total failures, of which 5 are unrecoverable cases not declined, leaving 13 among the recoverable set [4].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories