Build1 publisherNot yet confirmed elsewhere3 min readPublished
A blind model scores on your vision benchmark, which means the benchmark grades priors
MMEvalPro reports the best text-only LLM trailing the best multimodal model by only 14.64% on three popular benchmarks. That is a smaller gap than the spread among the vision models themselves.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Large language models without any visual perception capabilities achieve non-trivial performance on many multimodal multiple-choice benchmarks, undermining the credibility of those evaluations; MMEvalPro is proposed to address this.
- For each original question from existing benchmarks, human annotators augment it by creating one perception question and one knowledge anchor question, so MMEvalPro comprises question triplets.
- In the preliminary experiment, the average performance gap between the best LLM and the best LMM is just 14.64%, which is even smaller than the gap within the LMMs themselves.
- The paper mainly studies three popular multimodal benchmarks: MMMU, ScienceQA and MathVista.
- The authors attribute high LLM scores without visual data to possible data leakage, visual information problems not being related to answering the question, or simply guessing the answer.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A group of researchers has published MMEvalPro, a benchmark that rebuilds multimodal multiple-choice tests after finding that large language models with no visual perception at all score non-trivially on them [1]. In a preliminary experiment the authors call the Seeing-or-Not Comparison, the average gap between the best text-only LLM and the best large multimodal model across the tested benchmarks was 14.64%, which the paper notes is smaller than the gap among the multimodal models themselves [3][13].
That second clause is the part that should change how you run an internal eval. If the entire measurable contribution of seeing the image is narrower than the spread between two vision models on the same leaderboard, then a ranking difference between those two models sits inside the band that a model with no eyes can already reach [15]. The number you report as "perception" is partly answer-option prior.
The paper studies MMMU, ScienceQA and MathVista [4], and attributes the text-only scores to three ordinary causes: possible data leakage, questions whose visual information is not actually needed to answer, and plain guessing [5]. None of these require a clever model. They are properties of the item pool. The authors also note that some models have reached or surpassed human scores on certain benchmarks [12], and cite FastV's finding that a multimodal model can do better on some benchmarks using only part of the visual tokens [11]. Earlier work including PCA-Bench, MMStar and MathVerse had already flagged that multimodal MCQ evaluation gives LLMs a shortcut [10].
MMEvalPro's fix is cheap enough to copy. For each original question, human annotators add one perception question and one knowledge anchor question, producing triplets [2]. Two-thirds of the questions are labelled by human experts, with the remainder taken from the source benchmarks [9]. The metric is Genuine Accuracy rather than raw accuracy [8]. The point of the triplet is an error class the authors name directly: an Answer Consistency Test showed a prevalent Type-I error, where a model produces the correct answer without comprehension [6]. Their example is a model that computes the degree of an angle but cannot identify the name of that angle in the figure, which is a prerequisite for the computation [7].
For an operator, the transferable practice is three-part. Run every multimodal suite through a text-only baseline before you trust a single delta, because that baseline tells you the ceiling that requires no perception [1][5]. Add a perception item and a knowledge item beside each reasoning item you care about, then score only the cases where all three land [2][8]. Treat any two-model comparison narrower than your own measured text-only ceiling as unresolved rather than as a win [15].
One caution on the paper itself: the version of the text available to us omits the numeric values for MMEvalPro's headline comparisons, including the gap to human performance and the widened LLM-to-LMM gap [14]. The 14.64% figure from the preliminary experiment is the concrete number to work from, and it is the authors' own measurement rather than an independent replication [3].
What to watch: whether model vendors start publishing text-only baselines beside their multimodal scores, and whether the triplet structure survives contact with annotation cost, since two-thirds human labelling [9] is the expensive part of the recipe and the first thing a team under deadline will drop.