BuildNot yet confirmed elsewhere1 publisher3 min readPublished Updated
format="json" is steering, not a contract: what it takes to put a local model behind an API
Constrained decoding, a repair parser, validation and feedback retries do four different jobs. A dev.to walkthrough shows where each one stops working, and where a 30-second retry budget runs out.
The Engineer · Build desk

What happened
- Passing a real JSON Schema object instead switches on constrained decoding, which zeroes the probability of any token that would break the schema.
- For backends that do not constrain decoding, the fallback is json_repair, a drop-in replacement for json.loads that also strips stray prose.
- Parsed objects are then validated, with failures retried by feeding the error back, capped at one to three attempts before failing loudly and keeping the raw output.
Why it matters
- exposure A wrong-keyed object that is still valid JSON is discovered by the agent step or ETL row consuming it, not at the boundary where anyone is looking.
- constraint Marking a field required obliges the model to fill it, and the fabrication clears both the decoder and the validator, so a strict schema can buy confident wrong data.
- cost Three calls inside a 30-second total budget leaves about ten seconds each on the slow local hardware the retries exist for, so the budget decides whether retry two ever runs.
- decision Enabling skip_json_loads commits you to knowing in advance which payloads are broken, since the repair parser will reshape the ones that were fine.
The trap is in the API surface. The same `format` parameter takes either a string or a schema object, and the two do entirely different work. The string `"json"` is steering: the model is pushed toward emitting a valid object, and nothing enforces field names, types or required keys, so `{"result": "..."}` is a legal answer to a request for `severity` and `summary` [3]. The schema object is enforcement [4]. Both calls return without complaint, and the response gives you no way to tell which of the two you got.
Enforcement here is mechanical, which is why it is worth the migration. Ollama zeroes the probability of any token that would violate the schema at every generation step, so a markdown fence is not discouraged, it is unreachable [4]. No prompt wording does that. On Ollama 0.3.0 or newer the Python client passes the schema through with Pydantic's `model_json_schema()` [5], and the dev.to walkthrough reports that for straightforward schemas on a 7B-or-larger model this alone ends the parse failures [19].
The mask has a narrow remit, though. The same piece lists three things that still bite, and only one of them sits inside that remit: small models fumbling nested and optional fields [20]. Of the other two, one is a deployment fact, since an older Ollama or any endpoint offering only loose JSON mode never applies the mask at all [6]. The other is a job no decoder constraint can do. Where the text says the salary is competitive and your schema demands an integer, the model supplies an integer [8]. Two of the three named failure modes are therefore outside what constrained decoding can reach [18].
That second one is the expensive failure, because it survives every later layer. A fabricated integer satisfies the grammar that produced it and satisfies the validator that inspects it, since both only ever asked what type the value was [17]. The feedback retry does not catch it either: the retry is triggered by a validation error, and there is no validation error to feed back [1].
The retry ladder is where the budget breaks. `instructor` gives it to you in configuration, with `max_retries=2` and a `timeout` the source flags as total across retries rather than per call, precisely because local models are slow [9]. At the example's 30 seconds, that is three model calls in 30 seconds, roughly ten apiece [16], on hardware chosen because you own it. Rung two of the ladder is theoretical unless that number goes up.
On the defensive side, `json_repair` earns its place as the `json.loads` replacement, in preference to the regex that strips fences until the night it does not [7][2]. Its default ordering is the safe one: stdlib parse first, repair parser only on failure, so valid input passes through untouched [10]. `skip_json_loads=True` removes that fast path, and the library's own docs warn against pointing it at input you expect to be valid, because the repair parser can reshape valid JSON [11]. Where you cannot vouch for the provenance of the bytes, `strict=True`, which raises instead of repairing, is the setting that tells you the truth [14].
And `instructor` does not replace the mask. It talks to the OpenAI-compatible endpoint, which still depends on the backend honouring the schema [12]. Constrain, defend, validate, retry are four jobs, not four names for one [13]. The retries are what you run when the constraint is not there.
What to watch
- Whether json_repair's schema-guided repair path holds up on the nested and optional fields small models are said to fumble.
- Whether backends behind instructor's OpenAI-compatible route actually apply the schema, or only claim to accept it.
- A published per-call latency figure for local 7B extraction would settle whether a 30-second total retry budget is usable.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence41
- Adoption
- Insufficient
- Hype gap+14
- Incentives32
- Confidence54
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
The parsed object should always be validated against the contract, and on failure the model should be retried with the error fed back rather than blind re-rolled; one to three attempts is the right ceiling, after which the pipeline should fail loudly and keep the raw output for debugging.
- [2]
A local model asked for JSON commonly returns a code fence, prose such as "Here is your result:" and a trailing comma, so json.loads() raises JSONDecodeError; the naive fix is a regex that strips code fences, which works until it does not.
- [3]
Ollama's format="json" means JSON mode: the model is steered to emit a valid JSON object, but it does NOT enforce field names, types or required keys, so you can still get {"result":"..."} when you wanted {"severity":"high","summary":"..."}.
- [4]
When format is set to a JSON Schema object, Ollama applies constrained decoding: at every generation step it sets the probability of any token that would violate the schema to zero, so the model physically cannot emit a markdown fence, surrounding prose, or a structurally invalid object.
- [5]
The official recommended pattern is to pass a real JSON Schema via Pydantic's model_json_schema() to the Python Ollama client, and it works on Ollama 0.3.0 or newer.
- [6]
A different local server, an older Ollama, or any endpoint that only offers loose JSON mode will not constrain decoding, returning fences and partial objects.
- [7]
json_repair (PyPI json-repair) is a drop-in upgrade for json.loads() that fixes missing quotes, trailing commas and truncated values, and strips stray prose.
- [8]
Constrained decoding constrains structure, not truth: it guarantees the shape but not that values are correct, and if the text says "salary is competitive" while the schema demands an integer, the model will hallucinate a number to fill it.
- [9]
The instructor example uses max_retries=2 and timeout=30.0, with the timeout described as TOTAL across retries, which the source flags as important for slow local models.
- [10]
By default json_repair tries stdlib json.loads first and only falls back to the repair parser on failure, so feeding it valid JSON is safe.
- [11]
Passing skip_json_loads=True skips json_repair's fast path, and the library docs say not to use that flag on input you expect to be valid because the repair parser can reshape valid JSON.
- [12]
instructor leans on the OpenAI-compatible endpoint, which still depends on the backend honouring the schema.
- [13]
The robust design stated by the source is: constrain when you can, defend when you can't, validate always, retry with feedback.
- [14]
json_repair also supports schema/pydantic-guided repair and a strict=True mode that raises instead of repairing.
- [15]
The failure is described as hitting developers who have wired a local LLM into an agent, an ETL job, or a backend endpoint.
- [16]
A 30-second timeout that is total across retries, with max_retries=2, budgets three model calls and therefore about 10 seconds per call.
- [17]
A value the model fabricated to satisfy a required typed field passes both the decoding constraint and the schema validator, because the constraint only enforces structure and the validator only checks the contract's shape, so neither layer flags it.
- [18]
Of the three residual failure modes the source names, only one (small models on nested or optional fields) is inside constrained decoding's remit; the other two, unconstraining backends and fabricated values, are not.
- [19]
For most straightforward schemas on a 7B-plus model, schema-constrained decoding alone kills the parse failures.
- [20]
Small models still misbehave on nested or optional fields.
Sources
1 independent publisher whose own reporting we read for this story.
- dev.toYour LLM Returns JSON That Isn't JSON: A Robust Structured-Output Pipeline for Local Models
1 article · August 26, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
Entities
- OllamaFollow
- PydanticFollow
- json-repairFollow
- instructorFollow
- Qwen2.5-7BFollow
- JSON SchemaFollow