Skip to content

Build1 publisher2 min readPublished

JetBrains and Lund researchers recast AI code review as a trust-calibration problem

JetBrains and Lund University propose review tools that flag per-segment risk in AI-written code, a design shaped with 17 practitioners and a 43-person survey. They argue diff review cannot keep up with agent-sized changes unless tools show reviewers where to read closely.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • The paper describing the framework is due to be presented at Empirical Software Engineering International Week 2026 in October.
  • The authors define trust calibration as matching review effort to segment-level risk when the code's author cannot be questioned about its confidence or reasoning.
  • Their premise is that a language model presents every generated line with the same apparent confidence, however uncertain it actually was when producing it.
  • The proposed three-level workflow follows Shneiderman's ordering of overview first, then zoom and filter, then details on demand.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint The design is only as useful as its per-segment risk score; a score that misjudges a segment steers reviewers away from the code that most needed a close read.
  • cost Until tools can rank segments by risk, teams adopting coding agents pay for review in reviewer hours that rise as change sets get larger.
  • decision Review-tool builders now have a target they can instrument: effort spent per segment, set against the risk of that segment.

Review between colleagues depends on context the diff never shows. According to the post, you know roughly how senior the author is, which parts of the codebase they are confident in and which parts they rush, and you can ping them on Slack for a rationale in thirty seconds [6]. None of that exists when the author is an LLM [6].

The post opens with agents writing thousands of lines across a dozen files in the time it takes to make coffee [14]. Reading them takes longer. Without the social cues, the safe policy is to read every line, because any line could be the wrong one. The authors say that audit scales terribly as agents produce larger change sets [13].

The authors' target is the diff viewer. A diff viewer assumes the reviewer's job is to understand what changed [8]. With an LLM author, understanding the change is the easy part. The hard part is knowing whether to trust it, and where [8]. "The diff viewer was the right tool for reviewing what your colleague wrote. It may not be the right tool for reviewing what your agent wrote," the researchers wrote [12].

The three-level workflow is the strongest part of the proposal, and I think it is careful design work. It is built on what expert developers say about how they read unfamiliar code [9]. It also restores the step a line-by-line diff makes reviewers skip: forming hypotheses about the whole change before checking details [10].

The weak point is the signal. The framework calls for risk and confidence to be shown at the granularity where developers put their attention [3]. By the paper's own premise, the model's output cannot supply that score, because every line comes out looking equally sure [7]. The post's description of the framework does not specify where the score would come from. In my view, producing a score reviewers can trust is harder than designing the screen that shows it.

The evidence is participatory design sessions with practitioners and a follow-up survey of software professionals [4]. Those methods establish what reviewers want from a tool. Showing that reviews get faster or catch more defects would take a controlled comparison against a plain diff. For the result to transfer to a given team, two things have to hold. Its reviewers have to read unfamiliar code the way the study's experts described [9]. Its tooling has to produce segment-level risk scores that match where the defects actually are. "Until the tools catch up, developers will continue to fly blind," the JetBrains team wrote [11].

What to watch

  • Whether the full paper specifies how per-segment risk and confidence scores are computed, and from which inputs.
  • Any controlled study comparing review time and defects caught with segment-level risk signals against a plain diff view.
  • Whether JetBrains builds the three-level review workflow into its own IDEs or review tooling.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories