Skip to content

Build1 publisher3 min readPublished

The ignored 402 is a lint error, not a judgment call

A new open-source linter reads agent traces after the run and exits non-zero on structural defects. The interesting part is the exit code contract, not the rules.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • tracelint is described as a linter for agent runs: it reads the execution trace and flags structural bugs deterministically, with the exact trace lines as evidence and a CI exit code, running after the run on the trace rather than on the code, with no second model judging it.
  • The opening example: the charge_card tool returned a 402 and the agent kept going and told the customer their order shipped.
  • The author characterises an ignored tool error as a structural defect in the run rather than a hallucination, and says structural defects are decidable by looking at the trace.
  • The author states that published trace-error benchmarks show LLM judges have low localization accuracy, saying something seems off without reliably pointing at which step, and that judges are non-deterministic, cost money per trace, and cannot gate CI.
  • One structurally decidable defect: a tool call whose arguments violate the tool's JSON Schema, resolved by running the schema validator.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A developer has published tracelint, a linter that reads an agent's execution trace after the run, flags structural bugs deterministically with the offending trace lines as evidence, and returns a CI exit code, with no second model judging the output [1]. The bug it opens with is the one worth building around: a `charge_card` tool returned a 402 and the agent kept going and told the customer their order shipped [2].

That is not a made-up fact. It is a structural defect in the run, and per the write-up it is decidable by looking at the trace [3]. The author's argument against using an LLM judge here is a fitness argument rather than a purity one: published trace-error benchmarks show judges have low localization accuracy, telling you something seems off without reliably pointing at which step, and they are also non-deterministic, cost money per trace, and cannot gate CI [4].

The decidable set, as described, is small and boring in the right way. Tool call arguments that violate the tool's JSON Schema, which you settle by running the validator [5]. A tool that returned an error followed by the agent proceeding as if it had not [6]. The same tool called five times with identical arguments and identical results [7]. Arguments that appear nowhere in what the agent observed, as a candidate hallucinated value [8].

The part operators should read closely is the exit code contract. `tracelint check ./trace.json --tools ./tools.json` returns 0 for clean, 2 for a structurally-provable defect, and 3 for an input error, and heuristic findings never fail CI on their own [9][10]. Only provable things such as a schema violation or malformed JSON are hard defects; loops, redundant calls and suspicious arguments surface as candidates with evidence for a human, on the stated reasoning that a retry loop and a stuck loop look structurally similar [11]. Two sample findings from an OpenInference run illustrate the split: an R4 loop on `search` at steps 2, 4 and 6, and an R5 redundant_call at steps 2 and 6, both tagged candidate [12]. Under the documented contract that run does not fail a build [13].

The other design decision that earns its place: if a trace is missing a field a rule needs, such as tool schemas or result payloads, the rule does not silently pass. It suppresses with a stated reason [14]. Silent passes are how linters lose trust.

Ingest is where the distribution case sits. tracelint reads OpenInference, the OpenTelemetry semantic convention for AI, plus Langfuse and OpenAI message formats, so the claim is that you point it at spans you already emit into Phoenix, Langfuse or an OTel collector rather than learning a trace format [15][16]. There is a Python entry point, `lint_otel_trace`, that takes Phoenix dataframe records and returns a report with an exit code [17]. The author says validation ran against real OpenInference exports rather than only hand-built fixtures, and reports that on one real Phoenix trace the linter localized a genuine failure: rule R2a, step 9, `add_spans_to_dataset` returned a GraphQL error [18][19]. `pip install tracelint`, then `tracelint demo --html demo.html` runs a keyless suite with one planted instance of every defect plus clean controls [20][21].

Worth watching: whether the hard-defect tier stays narrow enough that a 2 always means a real bug, and whether the candidate tier stays useful or becomes noise teams filter out. The suppression list is the honest metric here. If most rules suppress on your traces for want of schemas or payloads, the linter is telling you your instrumentation is thin, which is a different repair job than fixing the agent.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories