Skip to content

ScienceNot yet confirmed elsewhere1 publisher3 min readPublished

At ICML, authors ranking their own papers beat reviewer scores at spotting future citations

A year-long experiment covering 2,592 ICML 2023 submissions found top self-ranked papers drew about twice the citations of bottom-ranked ones. The signal survived rejection, which is the awkward part.

The Scientist · Science desk

How we use AISend a correction

Photograph accompanying At ICML, authors ranking their own papers beat reviewer scores at spotting future citations
Photo: nature.com

What happened

  • ICML organizers approved an experiment that asked authors holding multiple 2023 submissions to rank their own papers by perceived quality.
  • Rankings came from 1,342 researchers and covered 2,592 submissions, matched to official review scores and final decisions.
  • Measured over 16 months of citations, top self-ranked papers drew about twice as many as bottom-ranked ones, among accepted and rejected papers alike.
  • Of the 22 papers that passed 150 citations, 17 had been ranked highest by at least one of their authors.

Why it matters

  • capability Programme chairs gain a ranking input that consumes no reviewer time, because it is answered by people who have already read the paper closely and cannot promote all of their entries at once.
  • constraint The signal exists only inside one author's portfolio, so any pool dominated by single-submission authors cannot be triaged this way, and coverage shrinks as submissions per author falls.
  • exposure Because the pattern held among rejected papers, a decision that overrides a top self-ranking is now checkable after the fact against a published predictor of citations.
  • contradiction The yardstick is citations, which the paper itself calls imperfect, while review is also charged with factual correctness, so 'better than review scores' is a claim about one of review's two jobs.

The elicitation is the part worth copying. Authors asked for an absolute quality score will inflate it, so the experiment never asked for one. It asked which of an author's own submissions is stronger, a comparison in which nobody can place every paper first, and which the paper grounds in game-theoretic reasoning about when truthful ranking is an author's best available move [9]. It is also an easier question than assigning a number [9].

Divide the submissions covered by the researchers who ranked them and the ratio is 1.93 [16]. Since eligibility required multiple submissions, every participant ranked at least two, so arithmetic forces at least 92 ranking slots onto submissions that somebody else also ranked [16]. That matters for the tail result, where 77 percent of the papers above 150 citations were top-ranked by at least one author [19]. "At least one" is a softer bar than agreement among co-authors, and the text as published does not give the agreement rate [6].

The incumbent it beat is not a precise instrument either. In a NeurIPS 2021 experiment, roughly half the accepted papers would have been rejected under a second independent review [13]. Outperforming review scores as a citation predictor [7] is therefore a low bar cleared, not a high one; the interesting property is that it was cleared by a question authors answer about work they already know, rather than by more reviewer hours. The robustness checks help here: the associations held after accounting for final decisions, review-score ranges, preprint posting dates and self-citations [8].

The axis is narrower than the framing suggests. Peer review is charged with establishing factual correctness as well as identifying scientific potential [14], and citation counts test neither correctness nor the reasons a paper was rejected. The paper concedes citations are an imperfect measure of impact [15]. So the honest statement is that authors predict one downstream proxy better than reviewers do, on a proxy reviewers were never asked to forecast.

Volume is what makes a free signal worth arguing about. ICML submissions went from 1,676 in 2017 to 12,107 in 2025, a factor of 7.2 [10][17]; NeurIPS went from 3,240 to 21,575 across the same years, a factor of 6.7 [11][18]. That is 33,682 submissions between two conferences in one year [20], against a reviewer pool that did not grow in step, which is why graduate and undergraduate reviewers without prior publications at these venues, and language models, are now doing the work [12].

Read against that, self-rankings are a prior, not a verdict. The defensible use is ordering attention: which papers get a third reviewer, which rejections get an area chair's second look, which of an author's submissions absorbs the scarce senior reviewer. The moment a ranking carries weight in the decision, the incentive that made it informative changes, and the experiment measured the version where it did not.

What to watch

  • Whether any conference makes self-rankings consequential in accept/reject decisions, and whether the predictive edge survives once authors know the ranking is scored.
  • Replication at a second venue such as NeurIPS, and whether the gap over review scores holds at citation windows longer than 16 months.
  • Whether the full paper reports co-author agreement rates, rather than crediting a paper when at least one author ranked it top.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence74
Adoption24
Hype gap+12
Incentives
Insufficient
Confidence56
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    The experiment collected self-rankings from 1,342 researchers covering 2,592 submissions, together with official review scores and final decisions.

  2. [2]

    A paper published on nature.com, 'Self-rankings as a predictor of scientific impact beyond peer review', reports a large-scale experiment at a leading AI conference testing whether authors' rankings of their own submissions predict scientific impact.

    ReportedSupportedView cited source
  3. [3]

    With the conference organizers' approval, the researchers asked authors with multiple ICML 2023 submissions to rank those submissions by perceived quality.

    ReportedSupportedView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. nature.com

    1 article · August 23, 2026

    Self-rankings as a predictor of scientific impact beyond peer review

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

  • Peer review at scaleFollow
  • Mechanism design for elicitationFollow
  • AI conference ecosystemFollow
  • Research impact measurementFollow
Loading related stories