Build1 distinct publisher3 min readPublished
Ai2 fit a multidimensional item response model to 100 models across 16 benchmarks. The per-question estimates say parts of its own safety suite are grading general reasoning rather than refusal behaviour.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Per question, the fit estimates two things: how difficult the item is, and how sharply it separates models that are strong or weak on each latent capability. Per model, it estimates strength on those same capabilities [8]. A benchmark average is a sum over items that keeps neither. Every question gets weight 1, whether it discriminates or not, and whichever dimension it loads on.
The part that makes this more than a relabelling exercise is that the dimensions were not supplied. The authors did not tell the model which benchmark was measuring what, and the fit recovered two dominant dimensions on its own [11]. Reruns from scratch produced the same two [12]. That is the check I care about, because a two-factor solution you specified in advance will always find two factors.
It is also the limit. The corpus was six reasoning benchmarks and ten from the Olmo 3 safety suite [10]. Two dominant dimensions is what is identifiable from a corpus built out of two families. Add a code or multilingual family and the honest expectation is a third dimension, with the per-item loadings moving underneath it. Ai2's stability result rules out one random seed as the explanation for the two-factor fit, though it stops short of showing that the basis is complete.
The WMDP sign is the case that should change a scorecard. The benchmark counts refusing or failing to produce dangerous dual-use knowledge as the desired answer [16], so the metric is partly an ignorance test, and it is one that appears in safety tables. Read as a safety composite, an average that includes WMDP subtracts some fraction of reasoning ability from the total. Read as a capability probe, the same score means something coherent. The number does not tell you which reading its publisher intended.
Now the adoption bill. The parameters come from 100 models scored on more than 34,000 questions [9], which is at least 3.4 million model-question observations [18]. Discrimination is defined against that population of models. For the published loadings to transfer to your selection, two things have to hold: your candidate has to behave like something in that pool, and you have to be scoring the same items. If you are evaluating one fine-tune on your own prompt set, you have no discrimination estimates at all, and getting them means running many models over every prompt, not one.
So the practical version of this is narrower than "audit per prompt". Where Ai2 has published loadings for a benchmark you already run, stop reporting the composite and report the subsets, because BBQ aligning with reasoning [14] and the two halves of WildJailbreak [5] mean the composite is an average of different measurements. Where it has not, you are back to writing your own prompt-level splits by hand and defending them with an argument rather than a fit.
In my context, picking a model where over-refusal costs revenue and harmful compliance costs more, those two WildJailbreak halves are two separate purchases [4]. The number I want printed next to a safety average is the discrimination of its items on the safety dimension.
Ranked by verification strength, evidence, and original report placement.
Ai2 introduced BenchMIRT, a method for auditing LLM benchmarks at the level of individual prompts, meaning the questions and tasks a model is scored on.
A benchmark is usually designed to measure a particular ability such as safety, general reasoning, or instruction following, but the individual tasks inside it may depend on more than that stated goal.
BBQ is designed to test whether models rely on social stereotypes; one question about a grandson and grandfather trying to book an Uber probes age bias but also requires the model to track who is who and reason from the evidence provided.
WildJailbreak includes harmful jailbreak prompts alongside benign prompts designed to test whether a model refuses harmless requests too often.
In BenchMIRT's analysis, WildJailbreak's harmful prompts are more closely associated with safety while its benign prompts are more closely associated with general reasoning, and averaging them into a single benchmark score can obscure that difference.
BenchMIRT takes cues from Item Response Theory, a psychometrics technique that starts from the idea that not every question tells you the same amount: some are harder, and some do a better job of distinguishing stronger from weaker performers.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 1, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
DeepSeek paid for Engram's lookup table by deleting 17 of its 72 routed experts1 distinct publisher
build
A refusal-stripped 27B model now ships as a 17.9 GB llama.cpp pull1 distinct publisher
invest
OpenAI's own timeline: twelve days from agent attack to knowing it was them1 distinct publisher
security
OpenAI's evaluation agents turned a package registry into their messaging bus1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed, self-graded, unchecked
Ai2 shows more of its work than a launch post usually does — 100 open-weight models, per-item difficulty and discrimination, a correlation chart on a −1 to 1 scale with significance marks, repeated fits from scratch. All of it is first-party. The lab built the method, supplied ten of the sixteen benchmarks under examination, and is the only one reporting the result; our coverage carries no paper, no code, no outside replication, and the piece stops mid-sentence on the question-ranking finding it was about to deliver.
Announcement day only
Running your own method over your own benchmark suite is a demonstration, not uptake. There is no second lab reporting item-level audits, no leaderboard folding the dimensions in, no download or integration figure — nothing that would let us put a number on who is using this.
Written smaller than it reads
The post talks itself down: these findings 'don't necessarily mean the benchmarks are flawed', they merely 'make the score easier to interpret'. Then it says a low BBQ score may reflect reasoning difficulty, and that better reasoners score worse on WMDP by design. Read together, that means safety scorecards built on those benchmarks are partly reporting capability — a sharper conclusion than the prose allows itself.
Auditing its own scorecard
Ten of the sixteen benchmarks in the fit are Ai2's Olmo 3 safety suite, and the method is framed as the successor to Ai2's Fluid Benchmarking — so the lab both stakes a claim on evaluation methodology and publishes an unflattering result about its own safety set. Those pressures point in opposite directions, which is why this sits mid-scale rather than high: there is a launch to promote, but the finding it promotes costs the promoter something.
Solid on what, provisional on so-what
What BenchMIRT is and what it was fit on can be stated plainly. Whether BBQ genuinely belongs in the reasoning column is one unreplicated model away from becoming received wisdom, and repeating an analysis from scratch guards against a bad run, not against the dimensionality and scoring choices shared by every run. Until someone outside Ai2 refits this, the direction is more trustworthy than the labels.