ScienceNot yet confirmed elsewhere1 publisher3 min readPublished
At ICML, authors ranking their own papers beat reviewer scores at spotting future citations
A year-long experiment covering 2,592 ICML 2023 submissions found top self-ranked papers drew about twice the citations of bottom-ranked ones. The signal survived rejection, which is the awkward part.
The Scientist · Science desk

What happened
- ICML organizers approved an experiment that asked authors holding multiple 2023 submissions to rank their own papers by perceived quality.
- Rankings came from 1,342 researchers and covered 2,592 submissions, matched to official review scores and final decisions.
- Measured over 16 months of citations, top self-ranked papers drew about twice as many as bottom-ranked ones, among accepted and rejected papers alike.
- Of the 22 papers that passed 150 citations, 17 had been ranked highest by at least one of their authors.
Why it matters
- capability Programme chairs gain a ranking input that consumes no reviewer time, because it is answered by people who have already read the paper closely and cannot promote all of their entries at once.
- constraint The signal exists only inside one author's portfolio, so any pool dominated by single-submission authors cannot be triaged this way, and coverage shrinks as submissions per author falls.
- exposure Because the pattern held among rejected papers, a decision that overrides a top self-ranking is now checkable after the fact against a published predictor of citations.
- contradiction The yardstick is citations, which the paper itself calls imperfect, while review is also charged with factual correctness, so 'better than review scores' is a claim about one of review's two jobs.
The elicitation is the part worth copying. Authors asked for an absolute quality score will inflate it, so the experiment never asked for one. It asked which of an author's own submissions is stronger, a comparison in which nobody can place every paper first, and which the paper grounds in game-theoretic reasoning about when truthful ranking is an author's best available move [9]. It is also an easier question than assigning a number [9].
Divide the submissions covered by the researchers who ranked them and the ratio is 1.93 [16]. Since eligibility required multiple submissions, every participant ranked at least two, so arithmetic forces at least 92 ranking slots onto submissions that somebody else also ranked [16]. That matters for the tail result, where 77 percent of the papers above 150 citations were top-ranked by at least one author [19]. "At least one" is a softer bar than agreement among co-authors, and the text as published does not give the agreement rate [6].
The incumbent it beat is not a precise instrument either. In a NeurIPS 2021 experiment, roughly half the accepted papers would have been rejected under a second independent review [13]. Outperforming review scores as a citation predictor [7] is therefore a low bar cleared, not a high one; the interesting property is that it was cleared by a question authors answer about work they already know, rather than by more reviewer hours. The robustness checks help here: the associations held after accounting for final decisions, review-score ranges, preprint posting dates and self-citations [8].
The axis is narrower than the framing suggests. Peer review is charged with establishing factual correctness as well as identifying scientific potential [14], and citation counts test neither correctness nor the reasons a paper was rejected. The paper concedes citations are an imperfect measure of impact [15]. So the honest statement is that authors predict one downstream proxy better than reviewers do, on a proxy reviewers were never asked to forecast.
Volume is what makes a free signal worth arguing about. ICML submissions went from 1,676 in 2017 to 12,107 in 2025, a factor of 7.2 [10][17]; NeurIPS went from 3,240 to 21,575 across the same years, a factor of 6.7 [11][18]. That is 33,682 submissions between two conferences in one year [20], against a reviewer pool that did not grow in step, which is why graduate and undergraduate reviewers without prior publications at these venues, and language models, are now doing the work [12].
Read against that, self-rankings are a prior, not a verdict. The defensible use is ordering attention: which papers get a third reviewer, which rejections get an area chair's second look, which of an author's submissions absorbs the scarce senior reviewer. The moment a ranking carries weight in the decision, the incentive that made it informative changes, and the experiment measured the version where it did not.
What to watch
- Whether any conference makes self-rankings consequential in accept/reject decisions, and whether the predictive edge survives once authors know the ranking is scored.
- Replication at a second venue such as NeurIPS, and whether the gap over review scores holds at citation windows longer than 16 months.
- Whether the full paper reports co-author agreement rates, rather than crediting a paper when at least one author ranked it top.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence74
- Adoption24
- Hype gap+12
- Incentives
- Insufficient
- Confidence56
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
The experiment collected self-rankings from 1,342 researchers covering 2,592 submissions, together with official review scores and final decisions.
- [2]
A paper published on nature.com, 'Self-rankings as a predictor of scientific impact beyond peer review', reports a large-scale experiment at a leading AI conference testing whether authors' rankings of their own submissions predict scientific impact.
- [3]
With the conference organizers' approval, the researchers asked authors with multiple ICML 2023 submissions to rank those submissions by perceived quality.
- [4]
The study evaluated whether the rankings predicted citations accumulated over 16 months.
- [5]
Papers ranked highest by their authors received about twice as many citations as those ranked lowest, among both accepted and rejected submissions.
- [6]
Among the 22 papers that received more than 150 citations, 17 were ranked highest by at least one author.
- [7]
Self-rankings were more predictive of future citation counts than peer-review scores.
- [8]
The associations remained statistically significant after checks involving final decisions, review-score ranges, preprint posting dates and self-citations.
- [9]
The authors elicited comparative rankings rather than absolute scores because authors may inflate absolute evaluations; authors cannot label every submission as their best, and under certain conditions the design encourages truthful rankings. The paper describes the approach as grounded in game-theoretic reasoning and notes comparative judgments ask an easier question than an absolute quality score.
- [10]
Submissions to ICML rose from 1,676 in 2017 to 12,107 in 2025.
- [11]
Submissions to NeurIPS rose from 3,240 in 2017 to 21,575 in 2025.
- [12]
The pool of experienced reviewers has not expanded at the same rate as submissions, and conferences have relied more on graduate and undergraduate reviewers, many without prior publications at these venues, and on language models for reviewing.
- [13]
In a NeurIPS 2021 experiment, approximately half of the accepted papers would have been rejected under a second independent review.
- [14]
Peer review in academic research aims to ensure factual correctness and to identify work of high scientific potential.
- [15]
The paper states that citations are widely used although imperfect measures of scientific impact.
- [16]
The experiment covered 1.93 distinct submissions per participating researcher, below the two-submission minimum implied by eligibility, so at least 92 ranking slots fell on submissions ranked by more than one participant.
- [17]
ICML submissions grew by a factor of about 7.2 between 2017 and 2025.
- [18]
NeurIPS submissions grew by a factor of about 6.7 between 2017 and 2025.
- [19]
77 percent of the papers exceeding 150 citations were ranked highest by at least one author.
- [20]
ICML and NeurIPS together took 33,682 submissions in 2025.
- [21]
Because only authors with more than one submission to the same conference can be asked to rank, authors submitting a single paper produce no self-ranking signal at all.
Sources
1 independent publisher whose own reporting we read for this story.
- nature.comSelf-rankings as a predictor of scientific impact beyond peer review
1 article · August 23, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.