Skip to content

Build1 publisher2 min readPublished Updated

LiveNerf measures Claude Opus 5.5 drift every day on 78 questions the model sometimes misses

LiveNerf reruns 78 calibrated questions against Claude Opus 5.5 every day, with frozen prompts and a pinned Claude Code CLI. That gives teams building on the model a dated launch baseline to test a suspected regression against.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying LiveNerf measures Claude Opus 5.5 drift every day on 78 questions the model sometimes misses
Generated illustration

What happened

  • Opus 5.5 shipped on September 22, 2026, and LiveNerf's daily runs began two days later, on September 24.
  • The 78 questions are what survived a screen of 2,336 candidates from GPQA Diamond, MMLU-Pro, competition math and AIME 2025-26, each sampled four times.
  • The harness runs on Inspect, the UK AI Security Institute's open evaluation framework, with statistics taken from Anthropic's published work on standard errors in evaluations.
  • By September 29 the project had logged six runs, all on one harness hash, 461391b6fce64167.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Teams watching Opus 5.5 in production have a case for logging output tokens per task beside quality scores, because the project's validation puts the token signal ahead of the accuracy one.
  • constraint The series measures Opus 5.5 as reached through Claude Code 2.1.280 with frozen prompts; a team on another client or effort setting can apply the result to itself only by assumption.
  • constraint The panel is graduate science and contest math, so a flat LiveNerf line says little about coding-agent or tool-use workloads unless those degrade the same way.

The calibration step is the part of the design I would copy. Opus 5.5 answered about 93% of the 2,336 candidates correctly on the first attempt, and 97% of them came back the same way on all four samples [15]. A question answered right four times out of four says little about a small change in the model. The 78 left in the middle are 3.3% of the pool [1]. They are the questions Opus 5.5 sometimes gets right and sometimes misses [16], and a small loss of reasoning has room to show there.

A reasoning model cannot be pinned the usual way. There are no sampling parameters to fix and no switch for the internal thinking [9]. LiveNerf freezes what surrounds the model instead: the prompts, the Claude Code CLI at version 2.1.280, and the harness hash [10]. The project says that with those fixed, any remaining difference has one explanation: the model changed [12].

The claim holds for the client side. From the user's seat, the causes the write-up lists (quantization, routing to a cheaper variant, lower reasoning effort) all count as the model changing [4]. Sampling noise also survives the freeze, and separating it from real change is the job of the statistics. Each question is compared against its own baseline, with errors clustered by question, so item difficulty cancels instead of mixing into the signal [13].

The effort sweep needs a careful reading. According to the project's published validation, low effort cut output tokens 62% while the score fell 8.3 +/- 4.5 points, and medium effort cut tokens 26% while the score fell 4.2 +/- 3.9 points [6]. Read as 95% intervals, both drops exclude zero, the medium one by 0.3 points [2]. Read as one standard error each, neither drop is two standard errors from zero: the low-effort drop is 1.8 and the medium one 1.1 [5][6]. If each question is scored once per run, one question is worth about 1.28 points [3], and the medium-effort drop is roughly three questions [4]. The write-up's case for tokens as the early warning rests on this gap: compute spend moves before accuracy becomes statistically significant [7].

When a model feels worse weeks after launch, users usually have screenshots and complaints and no baseline, and a screenshot comes without an error bar [3]. The 'nerf' label comes from the accusation that Anthropic degrades its models days or weeks after release; the write-up says it could be true, statistical noise, or a mix of both [5]. It does not report whether Opus 5.5 has moved against its launch baseline in the runs logged so far [11].

What to watch

  • A daily run where panel accuracy falls outside the launch baseline's error bars, or where output tokens drop with no announced version change.
  • A new Claude Code pin or harness hash in the repository, since runs after it would no longer compare directly with the ones logged since September 24.
  • Whether Anthropic responds to any drop the series records.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories