Science1 distinct publisher3 min readUpdated
A year-long experiment covering 2,592 ICML 2023 submissions found top self-ranked papers drew about twice the citations of bottom-ranked ones. The signal survived rejection, which is the awkward part.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
The elicitation is the part worth copying. Authors asked for an absolute quality score will inflate it, so the experiment never asked for one. It asked which of an author's own submissions is stronger, a comparison in which nobody can place every paper first, and which the paper grounds in game-theoretic reasoning about when truthful ranking is an author's best available move [9]. It is also an easier question than assigning a number [9].
Divide the submissions covered by the researchers who ranked them and the ratio is 1.93 [4]. Since eligibility required multiple submissions, every participant ranked at least two, so arithmetic forces at least 92 ranking slots onto submissions that somebody else also ranked [4]. That matters for the tail result, where 77 percent of the papers above 150 citations were top-ranked by at least one author [3]. "At least one" is a softer bar than agreement among co-authors, and the text as published does not give the agreement rate [6].
The incumbent it beat is not a precise instrument either. In a NeurIPS 2021 experiment, roughly half the accepted papers would have been rejected under a second independent review [13]. Outperforming review scores as a citation predictor [7] is therefore a low bar cleared, not a high one; the interesting property is that it was cleared by a question authors answer about work they already know, rather than by more reviewer hours. The robustness checks help here: the associations held after accounting for final decisions, review-score ranges, preprint posting dates and self-citations [8].
The axis is narrower than the framing suggests. Peer review is charged with establishing factual correctness as well as identifying scientific potential [14], and citation counts test neither correctness nor the reasons a paper was rejected. The paper concedes citations are an imperfect measure of impact [15]. So the honest statement is that authors predict one downstream proxy better than reviewers do, on a proxy reviewers were never asked to forecast.
Volume is what makes a free signal worth arguing about. ICML submissions went from 1,676 in 2017 to 12,107 in 2025, a factor of 7.2 [10][1]; NeurIPS went from 3,240 to 21,575 across the same years, a factor of 6.7 [11][2]. That is 33,682 submissions between two conferences in one year [5], against a reviewer pool that did not grow in step, which is why graduate and undergraduate reviewers without prior publications at these venues, and language models, are now doing the work [12].
Read against that, self-rankings are a prior, not a verdict. The defensible use is ordering attention: which papers get a third reviewer, which rejections get an area chair's second look, which of an author's submissions absorbs the scarce senior reviewer. The moment a ranking carries weight in the decision, the incentive that made it informative changes, and the experiment measured the version where it did not.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
The experiment collected self-rankings from 1,342 researchers covering 2,592 submissions, together with official review scores and final decisions.
A paper published on nature.com, 'Self-rankings as a predictor of scientific impact beyond peer review', reports a large-scale experiment at a leading AI conference testing whether authors' rankings of their own submissions predict scientific impact.
With the conference organizers' approval, the researchers asked authors with multiple ICML 2023 submissions to rank those submissions by perceived quality.
The study evaluated whether the rankings predicted citations accumulated over 16 months.
Papers ranked highest by their authors received about twice as many citations as those ranked lowest, among both accepted and rejected submissions.
Among the 22 papers that received more than 150 citations, 17 were ranked highest by at least one author.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Strong single-study evidence, one venue-year
The core claims rest on a preregistered-style, organizer-approved field experiment with a large sample (1,342 ranking researchers, 2,592 submissions, 39.6% of the conference), outcomes linked to official review scores and decisions, a stated 16-month citation window, and robustness checks across decisions, score ranges, preprint dates, self-citations, Google Scholar citations and GitHub stars. It is nonetheless one conference, one year, one publisher, with no independent replication in the supplied material, and the outcome is an association with an admittedly imperfect impact proxy.
One sanctioned pilot, no process integration
Adoption evidence is limited to a single organizer-approved prereview survey at ICML 2023 with voluntary participation: 30.4% of authors responded and 39.6% of submissions were ranked. The supplied source shows no conference incorporating self-rankings into review or decision workflows, no repeat run at a later venue, and by construction the mechanism cannot cover single-submission authors.
Slightly overstated by the 'beats peer review' framing
The paper's own language is hedged - self-rankings 'complement' peer review, citations do not fully measure quality - and its numbers are reported with robustness checks, so the gap is small. It leans positive because a single-venue, single-year association over 16 months is easily read as a general claim that authors outrank reviewers at spotting impact, while adoption evidence is one voluntary pilot and the truthfulness argument holds only in the consequence-free prereview setting actually tested.
No disclosure in supplied source
The supplied excerpt contains no funding statement, competing-interest declaration, or description of any commercial or institutional relationship between the authors and the OpenRank.cc survey platform or the conference, so no incentive structure can be scored without inference.
Solid primary source, single-source cluster
The claim set is drawn directly from a peer-reviewed primary paper with explicit numbers, which supports high confidence in what was found at ICML 2023. Confidence is held down because the cluster contains exactly one source and one publisher, there is no independent verification or critical commentary, generalization beyond one venue-year is untested, and incentive disclosure is missing.
science
A parameter-free tree replaces UMAP, and turns up an NK cell subtype from the myeloid lineage1 distinct publisher
science
A compact Fanzor2 editor beats Cas12f 2.6-fold, which puts the class average near 13%1 distinct publisher
science
Nature Perspective: patching one fact into a model leaves the reasoning around it broken1 distinct publisher
science
Nature paper pins each gas in lithium-metal cells to an electrode, then buys 10x cycles for free1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 23, 2026