Skip to content

BuildNot yet confirmed elsewhere1 publisher3 min readPublished Updated

format="json" is steering, not a contract: what it takes to put a local model behind an API

Constrained decoding, a repair parser, validation and feedback retries do four different jobs. A dev.to walkthrough shows where each one stops working, and where a 30-second retry budget runs out.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying format="json" is steering, not a contract: what it takes to put a local model behind an API
Generated illustration

What happened

  • Passing a real JSON Schema object instead switches on constrained decoding, which zeroes the probability of any token that would break the schema.
  • For backends that do not constrain decoding, the fallback is json_repair, a drop-in replacement for json.loads that also strips stray prose.
  • Parsed objects are then validated, with failures retried by feeding the error back, capped at one to three attempts before failing loudly and keeping the raw output.

Why it matters

  • exposure A wrong-keyed object that is still valid JSON is discovered by the agent step or ETL row consuming it, not at the boundary where anyone is looking.
  • constraint Marking a field required obliges the model to fill it, and the fabrication clears both the decoder and the validator, so a strict schema can buy confident wrong data.
  • cost Three calls inside a 30-second total budget leaves about ten seconds each on the slow local hardware the retries exist for, so the budget decides whether retry two ever runs.
  • decision Enabling skip_json_loads commits you to knowing in advance which payloads are broken, since the repair parser will reshape the ones that were fine.

The trap is in the API surface. The same `format` parameter takes either a string or a schema object, and the two do entirely different work. The string `"json"` is steering: the model is pushed toward emitting a valid object, and nothing enforces field names, types or required keys, so `{"result": "..."}` is a legal answer to a request for `severity` and `summary` [3]. The schema object is enforcement [4]. Both calls return without complaint, and the response gives you no way to tell which of the two you got.

Enforcement here is mechanical, which is why it is worth the migration. Ollama zeroes the probability of any token that would violate the schema at every generation step, so a markdown fence is not discouraged, it is unreachable [4]. No prompt wording does that. On Ollama 0.3.0 or newer the Python client passes the schema through with Pydantic's `model_json_schema()` [5], and the dev.to walkthrough reports that for straightforward schemas on a 7B-or-larger model this alone ends the parse failures [19].

The mask has a narrow remit, though. The same piece lists three things that still bite, and only one of them sits inside that remit: small models fumbling nested and optional fields [20]. Of the other two, one is a deployment fact, since an older Ollama or any endpoint offering only loose JSON mode never applies the mask at all [6]. The other is a job no decoder constraint can do. Where the text says the salary is competitive and your schema demands an integer, the model supplies an integer [8]. Two of the three named failure modes are therefore outside what constrained decoding can reach [18].

That second one is the expensive failure, because it survives every later layer. A fabricated integer satisfies the grammar that produced it and satisfies the validator that inspects it, since both only ever asked what type the value was [17]. The feedback retry does not catch it either: the retry is triggered by a validation error, and there is no validation error to feed back [1].

The retry ladder is where the budget breaks. `instructor` gives it to you in configuration, with `max_retries=2` and a `timeout` the source flags as total across retries rather than per call, precisely because local models are slow [9]. At the example's 30 seconds, that is three model calls in 30 seconds, roughly ten apiece [16], on hardware chosen because you own it. Rung two of the ladder is theoretical unless that number goes up.

On the defensive side, `json_repair` earns its place as the `json.loads` replacement, in preference to the regex that strips fences until the night it does not [7][2]. Its default ordering is the safe one: stdlib parse first, repair parser only on failure, so valid input passes through untouched [10]. `skip_json_loads=True` removes that fast path, and the library's own docs warn against pointing it at input you expect to be valid, because the repair parser can reshape valid JSON [11]. Where you cannot vouch for the provenance of the bytes, `strict=True`, which raises instead of repairing, is the setting that tells you the truth [14].

And `instructor` does not replace the mask. It talks to the OpenAI-compatible endpoint, which still depends on the backend honouring the schema [12]. Constrain, defend, validate, retry are four jobs, not four names for one [13]. The retries are what you run when the constraint is not there.

What to watch

  • Whether json_repair's schema-guided repair path holds up on the nested and optional fields small models are said to fumble.
  • Whether backends behind instructor's OpenAI-compatible route actually apply the schema, or only claim to accept it.
  • A published per-call latency figure for local 7B extraction would settle whether a 30-second total retry budget is usable.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence41
Adoption
Insufficient
Hype gap+14
Incentives32
Confidence54
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    The parsed object should always be validated against the contract, and on failure the model should be retried with the error fed back rather than blind re-rolled; one to three attempts is the right ceiling, after which the pipeline should fail loudly and keep the raw output for debugging.

  2. [2]

    A local model asked for JSON commonly returns a code fence, prose such as "Here is your result:" and a trailing comma, so json.loads() raises JSONDecodeError; the naive fix is a regex that strips code fences, which works until it does not.

    ReportedSupportedView cited source
  3. [3]

    Ollama's format="json" means JSON mode: the model is steered to emit a valid JSON object, but it does NOT enforce field names, types or required keys, so you can still get {"result":"..."} when you wanted {"severity":"high","summary":"..."}.

    ReportedSupportedView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. dev.to

    1 article · August 26, 2026

    Your LLM Returns JSON That Isn't JSON: A Robust Structured-Output Pipeline for Local Models

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

Loading related stories