Published · 5d agoScience2 min read
Gemini 3.1 Pro's 77.1% on ARC-AGI-2 measures the variable that was easiest to move
Google's headline score is more than double its predecessor's. On Artificial Analysis's real-world agentic evaluation, the same model gained ground and still finished behind four rivals.
Written for builders.See today for builders

What happened
- On ARC-AGI-2, described by Google as a benchmark that evaluates a model's ability to solve entirely new logic patterns, Gemini 3.1 Pro achieved a verified score of 77.1%.
- Google states the ARC-AGI-2 result is more than double the reasoning performance of Gemini 3 Pro.
- Artificial Analysis reports Gemini 3.1 Pro Preview improved on GDPval-AA, its agentic evaluation focused on real-world tasks, increasing its ELO score over 100 points to 1316 from Gemini 3 Pro Preview, but still sitting behind Claude Sonnet 4.6, Opus 4.6, GPT-5.2 (xhigh) and GLM-5.
- Four models are named as ranking ahead of Gemini 3.1 Pro Preview on GDPval-AA, placing it fifth among the models cited.
- Artificial Analysis reports Gemini 3.1 Pro Preview scoring 18% on CritPt, a set of unpublished research-level physics reasoning problems, over 5 percentage points above the next best model.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
Google's launch post for Gemini 3.1 Pro leads with one figure: a verified 77.1% on ARC-AGI-2, a benchmark that Google describes as evaluating a model's ability to solve entirely new logic patterns [1]. The company says that is more than double the reasoning performance of Gemini 3 Pro [2]. The doubling is large and the score is verified. It is also the variable in the release that was easiest to move.
What ARC-AGI-2 scores is pattern induction with a checkable answer. What operators buy is task completion. On the same model, Artificial Analysis reports GDPval-AA, its agentic evaluation of real-world tasks, rising more than 100 ELO points to 1316 while still sitting behind Claude Sonnet 4.6, Opus 4.6, GPT-5.2 (xhigh) and GLM-5 [3] - fifth among the models named [4]. And on CritPt, unpublished research-level physics problems, the leading score is 18%, which Artificial Analysis notes is over 5 percentage points ahead of the next best model [5]. Frontier performance on problems nobody has published an answer to is 18%, not 77%.
This gap between the measured variable and the intended one is a documented feature of evaluation design, not a Google problem. In multimodal multiple-choice benchmarks, the MMEvalPro authors found language models with no visual perception scoring non-trivially, with the average gap between best LLM and best LMM at 14.64% [6], and a prevalent Type-I error in which models produce correct answers without comprehension [7]. A separate paper reports that answer choices themselves act as priors, steering models to plausible options the figure does not support [8]. Those findings are about figure QA, not ARC-AGI-2; the transferable lesson is that a headline score reports what was cheap to grade.
The numbers likelier to change a procurement decision sit lower in the same coverage: $892 to run the Artificial Analysis Intelligence Index, less than half the cost of Opus 4.6 (max) and about twice GLM-5's $547 [9], and a 38 percentage point cut in the AA-Omniscience hallucination rate versus Gemini 3 Pro Preview [10].
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
On ARC-AGI-2, described by Google as a benchmark that evaluates a model's ability to solve entirely new logic patterns, Gemini 3.1 Pro achieved a verified score of 77.1%.
- [2]
Google states the ARC-AGI-2 result is more than double the reasoning performance of Gemini 3 Pro.
- [3]
Artificial Analysis reports Gemini 3.1 Pro Preview improved on GDPval-AA, its agentic evaluation focused on real-world tasks, increasing its ELO score over 100 points to 1316 from Gemini 3 Pro Preview, but still sitting behind Claude Sonnet 4.6, Opus 4.6, GPT-5.2 (xhigh) and GLM-5.
- [5]
Artificial Analysis reports Gemini 3.1 Pro Preview scoring 18% on CritPt, a set of unpublished research-level physics reasoning problems, over 5 percentage points above the next best model.
- [6]
MMEvalPro reports that large language models without any visual perception capabilities achieve non-trivial performance on popular multimodal multiple-choice benchmarks, with the average performance gap between the best LLM and the best LMM at 14.64%, smaller than gaps within LMMs themselves.
- [7]
MMEvalPro's answer consistency test found a prevalent Type-I error in multiple-choice evaluation, where models output correct answers without actual comprehension.
Sources & coverage · 4 publishers
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- blog.google5d agoAnnouncing our latest Gemini AI model
Cited in this coverage: Google (blog.google launch post)
- artificialanalysis.ai5d agoGemini 3.1 Pro Preview
- artificialintelligencemadesimple.com5d agoChatGPT vs Gemini vs Claude: The Best LLM Subscription You Should Buy (2026 Edition)


