Skip to content

Build1 publisher3 min readPublished

Five frontier LLMs gave split fact-check verdicts on 63% of 997 real user claims

Five frontier LLMs with web search disagreed on 63% of 997 claims users sent to Lenz.io for checking. They also rated themselves 9 or 10 out of 10 on 76% of answers, so one model's confidence says little about whether the others would agree.

The Engineer · Build desk

Photograph accompanying Five frontier LLMs gave split fact-check verdicts on 63% of 997 real user claims
Photo: lesswrong.com

What happened

  • The study's authors took the 1,000 most recent claims users submitted to fact-checking platform Lenz.io between May 1 and July 18, 2026, and gave five frontier models one prompt.
  • On the 997 claims where all five models returned a usable verdict, at least one model disagreed with the others on 63%.
  • Across all answers, the models rated their own confidence 9 or 10 out of 10 76% of the time, for a mean of 8.99.
  • Agreement between pairs of models ran from 76% for Grok 4.5 and Gemini 3.1 Pro down to 49% for Gemini and Sonar Deep Research.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Swapping Gemini for Sonar Deep Research in a single-model checker would have changed the label on 51% of these claims, so picking a vendor is an editorial decision.
  • constraint An auto-publish rule keyed to confidence of 9 or above would pass three answers in four even though the models disagree on most claims.
  • cost Human review is the price of a unanimity rule: about 628 of the 997 claims would have needed a reviewer or a tie-break.

The five-point scale sets the size of the 63% figure [4]. The authors ranked verdicts from True at 0 through Mostly True, Mixed and Mostly False to False at 4. Then they took the widest gap on each claim [5]. A split between True and Mostly True counts as disagreement. On about 40% of the 997 claims, the spread never went past that single step [1]. The harder cases are the 23% where the two furthest verdicts sat two or more steps apart, for example True against Mixed [5].

Counting votes gives a finer picture. All five models agreed on 37% of claims [7]. No verdict won three votes on 11% [7]. On the other 52%, a majority of three or four held against one or two dissenters [2]. In my view the useful design is a router keyed to the vote. Unanimous verdicts get labelled automatically. Majority verdicts get a quick human check, and claims with no majority get a full review.

Self-reported confidence mostly tracks how often a model picks one end of the scale. Most answers from all five landed on True or False [16]. Gemini gave a True or False verdict on 83% of claims and a confidence of 10 on 70% of its answers, and the authors link the two [11]. Fable 5 was the least confident model in the panel [11]. "High confidence from an individual model was not enough to show that the other models would agree with its verdict," the authors wrote [15].

Agreeing with the majority is a different thing from being right. Sonar Deep Research sided with the strict majority less often than Fable 5, Grok 4.5 or Gemini did. It also spread its verdicts most evenly across the five categories [12]. On a claim that really is mixed, the model that answers Mixed could be the correct outlier. This study cannot settle that. The authors did no human labelling, and they wrote that they "didn't assert that any model was more accurate than the others" [10].

For these figures to carry over to another pipeline, the inputs have to look alike. The claims were user submissions to Lenz.io. Near-duplicates were collapsed by cosine distance on text-embedding-3-small embeddings [1][13]. Every model had web retrieval and deep thinking switched on and got one prompt with five defined categories [2][3]. The panel also contains a substitute. Claude Fable 5 refused 145 claims and Opus 4.8 answered in its place [9], so on 142 of the analysed claims the column labelled Fable 5 is Opus 4.8 [3]. The post opens on the premise that frontier models post similar results on public benchmarks, though it does not list scores for these five [14].

What to watch

  • Human labels on the 997 claims would turn these disagreement figures into per-model accuracy.
  • A rerun on a binary True/False scale would show how much of the 63% comes from the five-point grading.
  • A repeat on a later window of Lenz.io submissions would test whether the 49% to 76% pairwise agreement range holds.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories