Build1 distinct publisher3 min readUpdated
MMEvalPro reports the best text-only LLM trailing the best multimodal model by only 14.64% on three popular benchmarks. That is a smaller gap than the spread among the vision models themselves.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
A group of researchers has published MMEvalPro, a benchmark that rebuilds multimodal multiple-choice tests after finding that large language models with no visual perception at all score non-trivially on them [1]. In a preliminary experiment the authors call the Seeing-or-Not Comparison, the average gap between the best text-only LLM and the best large multimodal model across the tested benchmarks was 14.64%, which the paper notes is smaller than the gap among the multimodal models themselves [3][13].
That second clause is the part that should change how you run an internal eval. If the entire measurable contribution of seeing the image is narrower than the spread between two vision models on the same leaderboard, then a ranking difference between those two models sits inside the band that a model with no eyes can already reach [14]. The number you report as "perception" is partly answer-option prior.
The paper studies MMMU, ScienceQA and MathVista [4], and attributes the text-only scores to three ordinary causes: possible data leakage, questions whose visual information is not actually needed to answer, and plain guessing [5]. None of these require a clever model. They are properties of the item pool. The authors also note that some models have reached or surpassed human scores on certain benchmarks [12], and cite FastV's finding that a multimodal model can do better on some benchmarks using only part of the visual tokens [11]. Earlier work including PCA-Bench, MMStar and MathVerse had already flagged that multimodal MCQ evaluation gives LLMs a shortcut [10].
MMEvalPro's fix is cheap enough to copy. For each original question, human annotators add one perception question and one knowledge anchor question, producing triplets [2]. Two-thirds of the questions are labelled by human experts, with the remainder taken from the source benchmarks [9]. The metric is Genuine Accuracy rather than raw accuracy [8]. The point of the triplet is an error class the authors name directly: an Answer Consistency Test showed a prevalent Type-I error, where a model produces the correct answer without comprehension [6]. Their example is a model that computes the degree of an angle but cannot identify the name of that angle in the figure, which is a prerequisite for the computation [7].
For an operator, the transferable practice is three-part. Run every multimodal suite through a text-only baseline before you trust a single delta, because that baseline tells you the ceiling that requires no perception [1][5]. Add a perception item and a knowledge item beside each reasoning item you care about, then score only the cases where all three land [2][8]. Treat any two-model comparison narrower than your own measured text-only ceiling as unresolved rather than as a win [14].
One caution on the paper itself: the version of the text available to us omits the numeric values for MMEvalPro's headline comparisons, including the gap to human performance and the widened LLM-to-LMM gap [15]. The 14.64% figure from the preliminary experiment is the concrete number to work from, and it is the authors' own measurement rather than an independent replication [3].
What to watch: whether model vendors start publishing text-only baselines beside their multimodal scores, and whether the triplet structure survives contact with annotation cost, since two-thirds human labelling [9] is the expensive part of the recipe and the first thing a team under deadline will drop.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Large language models without any visual perception capabilities achieve non-trivial performance on many multimodal multiple-choice benchmarks, undermining the credibility of those evaluations; MMEvalPro is proposed to address this.
For each original question from existing benchmarks, human annotators augment it by creating one perception question and one knowledge anchor question, so MMEvalPro comprises question triplets.
In the preliminary experiment, the average performance gap between the best LLM and the best LMM is just 14.64%, which is even smaller than the gap within the LMMs themselves.
The paper mainly studies three popular multimodal benchmarks: MMMU, ScienceQA and MathVista.
The authors attribute high LLM scores without visual data to possible data leakage, visual information problems not being related to answering the question, or simply guessing the answer.
The Answer Consistency Test reveals a prevalent Type-I error in multiple-choice evaluation conclusions, where models output correct answers without actual comprehension.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Primary preprint, self-reported, unreplicated
The core findings come directly from the primary document, with a named experimental design (Seeing-or-Not Comparison, Answer Consistency Test) and one concrete quantitative anchor (14.64%). Against that: a single publisher, no peer review or third-party replication in the cluster, and several of the abstract's headline numbers, including the human-performance lag and the MMEvalPro blind-model gap, are absent from the available text, so the strongest 'more trustworthy' claim cannot be checked.
No adoption signal beyond the authors' own release
The only observations available are the authors' own benchmark construction and their own experimental results. The supplied material contains no downstream usage, no external leaderboard adoption of Genuine Accuracy, no model-vendor reporting against MMEvalPro, and no download or citation figures, so adoption cannot be scored.
Mildly overstated relative to what the text shows
The underlying finding is concrete and well framed, and the story's headline restates the paper's own 14.64% figure faithfully. The overstatement is in reach rather than direction: the claim that MMEvalPro is 'more trustworthy' and 'more challenging' rests on numbers missing from the available text, comes from the benchmark's authors, and has zero external adoption or replication attached, so the certainty implied by the framing exceeds what the single source establishes.
Authors diagnose the flaw and sell the replacement
The same team that reports existing multimodal MCQ benchmarks as untrustworthy is proposing MMEvalPro, the Genuine Accuracy metric and the triplet annotation pipeline as the fix, and stakes the benchmark's value on the size of the flaw. That is a clear self-interest in a strong negative finding about MMMU, ScienceQA and MathVista. Mitigating it: the mechanism is documented, and independent prior work is cited in the same direction, but no disinterested source is present in the cluster.
Directionally credible, quantitatively thin
One publisher, one document, an interested author, and missing headline numbers hold confidence below the midpoint. What supports the direction is that the mechanism is specific and testable by any reader with a text-only model, and that the paper cites four independent prior efforts reaching similar conclusions about MCQ shortcuts.
build
Four months of A100 bills say self-hosting is a utilization bet, not a cost saving1 distinct publisher
leadership
Re-baseline AI procurement on cost per completed task, not dollars per million tokens1 distinct publisher
build
Deferred tool schemas cut cost 21% on average, and made one task type 12.3% dearer1 distinct publisher
build
GPT-5.6 ships as three models, and that makes model choice a deployment decision1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 18, 2026