Skip to content

Build1 publisher3 min readPublished

Your reviewing model is reading the diff when it should be reading the session

A vendor writeup argues cross-family AI review fails on inputs, not model choice: the diff shows what changed, the agent session shows whether it was asked for. The evidence is one anecdote.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Your reviewing model is reading the diff when it should be reading the session
Generated illustration

What happened

  • Having one model review code written by a model from a different family is increasingly popular: for example building a feature with Codex then asking Claude Code to review it, on the theory that each provider trains its models differently so the reviewing model may catch problems the original missed.
  • A typical AI code review might include the final code, the diff, and the commit history.
  • Those inputs tell the reviewer what changed, but not why it changed.
  • The author built an Arcade feature with Codex, then asked Claude Code to review the result.
  • The author reports the review produced reasonable feedback but felt incomplete and superficial, because Claude could inspect the implementation but did not have the conversation that produced it and could not compare the code against the original intent.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A developer writing on dev.to has put a specific complaint to the now-common practice of building with one model family and reviewing with another: the reviewing model is handed the final code, the diff, and the commit history, which describe what changed but not why [1][2][3]. That relocates the interesting failure from model selection to input selection, which is a cheaper thing for an operator to fix.

The pattern under examination is familiar. Build a feature with Codex, then ask Claude Code to review it, on the theory that different providers train differently and the second model catches what the first missed [1]. The author tried exactly that on an Arcade feature and reports the output was reasonable but incomplete and superficial, because Claude Code could inspect the implementation without having the conversation that produced it [4][5].

The argument is that in agent-assisted work the reasoning lives in the session: original prompts, agent responses, tool calls and commands, files the agent inspected, constraints supplied by the developer, and alternatives considered along the way [6]. Absent that, according to the post, a reviewer can still flag bugs, questionable patterns, and missing tests, but cannot reliably say whether the implementation matches what was requested [7].

The demonstration is small and, to its credit, unambiguous. The author asked for two ad cards, one per game, and the agent produced three [8][9] - one card, or fifty percent, more than specified [10]. Three cards are not broken code, and a conventional review would not necessarily flag clean, tested, technically correct work [11]. With the original prompt loaded into the review criteria, the consolidated report classified the extra card as intent drift [12].

The mechanism is a product. Entire stores prompts and agent responses and connects that context to Git through checkpoints [13]. Running `entire review` walks you through a review profile that defines what gets checked, which agents review, and which agent consolidates [14]. Selected reviewers run in parallel and a judge merges their findings, resolving contradictions, removing duplicates, and prioritising issues supported by evidence [15]. `entire review --edit` is how the intent comparison gets added as a check, and `entire review general` runs the configured profile [16][17].

Two things to hold at arm's length. First, this is Entire's own post arguing for Entire, and the entire evidentiary base is one feature and one miscounted card [18]. There are no numbers on judge accuracy, false positives, review latency, or token cost [19], which are the figures that decide whether a parallel-reviewer-plus-judge topology survives contact with a busy repository. Second, feeding the prompt into the review makes the prompt the specification of record. That is defensible when the prompt was right and the agent wandered. It is less defensible when the developer changed their mind three turns in, or asked for the wrong thing precisely, and the reviewer now has a documented reason to call correct work drift.

Watch whether session transcripts become portable across agent vendors, since a review that depends on captured prompts is only as good as its access to them. Watch whether consolidated verdicts stay auditable back to the individual reviewer that raised each finding [15]. And watch how these tools handle intent that legitimately moved during a session, because that is the common case, not the miscounted card.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories