Skip to content

Build1 publisher3 min readPublished

Two rejected papers: the shadow evaluation that undercuts autonomous AI research claims

Princeton and the UK AI Security Institute gave a frontier agent six days and $3,000 to answer real unpublished research questions. The original authors reviewed the output and rejected both papers.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • A new paper from Princeton and the UK AI Security Institute tests the claim that AI agents can conduct AI research, finding that today's frontier models can handle research engineering but fail at the parts of the research process that matter.
  • Anthropic and OpenAI have been touting their models' ability to speed up AI research.
  • The authors argue solid evidence for claims about automated AI research has been mostly absent: existing evaluations either test agents on narrow, verifiable tasks or submit AI-generated papers to peer review, a process the researchers call "overstretched, stochastic, and suffers from poor review quality."
  • The researchers call their approach "Shadow Evaluation": an agent receives the core research question from an unpublished paper, and the original authors, who spent months on the same question, then evaluate the result as conference reviewers would. Because the results are not yet on the web, the agent cannot fall back on training data.
  • The team partnered with the authors of two NeurIPS 2026 submissions. The first examines how personality traits of language models can be steered through their weights. The second develops a method called TabPFN that detects when a tabular prediction model hits deployment data that differs sharply from its training data and tanks its accuracy.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

Researchers at Princeton and the UK AI Security Institute handed a frontier agent the core research question from an unpublished paper, gave it six days, $3,000 in API credits, a GPU budget, a virtual machine and the open web, then asked the paper's original authors to review the result as conference reviewers [4][7]. Both papers produced this way were rejected, one with a Strong Reject, which is a direct check on Anthropic's and OpenAI's marketing about models accelerating AI research [9][2].

The design is the interesting part. The authors call it Shadow Evaluation: because the reference paper is unpublished, the answer is not in the training data, and because the reviewers spent months on the same question, they know what a real contribution looks like [4]. The team argues existing evidence is weak, either narrow verifiable tasks or AI-written papers pushed through peer review, a process they describe as "overstretched, stochastic, and suffers from poor review quality" [3]. Two NeurIPS 2026 submissions were used: one on steering language model personality traits through weights, one on a method called TabPFN for detecting when a tabular prediction model meets deployment data far from its training distribution [5]. The main runs used Claude Opus 4.8 with Extra-High Reasoning inside OpenClaw, the open-source agent framework built by Peter Steinberger, who joined OpenAI earlier this year [6][8].

The engineering held up. The agents ran literature searches, debugged GPU code, and completed hundreds of experiments and robustness tests without human help [18]. What failed was everything downstream of that. Reviewers cited poorly motivated data and experiments, unreadable prose, and no new contributions [9]; one called the reasoning a "'proof by example' fallacy" that was "highly non-scientific," another called the experiment choices "bizarre" and the results the product of "post hoc choices" [10].

The log analysis reads like a list of operational defects rather than intelligence deficits. The agents generated plausible hypotheses then killed them on small, hand-curated or synthetic datasets, showing no sense of the publication bar [11]. When a hypothesis was falsified they narrowed the existing claim instead of opening a new line [12]. Their own internal AI reviewer never returned a single Accept across fifteen revision rounds, and the core criticism was never addressed [13]. Both runs abandoned their most ambitious goals inside ten hours, and the Personas run closed out exploration after five hours against a plan of 36 to 48, roughly a tenth to a seventh of the time it had allocated itself [14][2]. Both finished with less than half the API budget spent, under $1,500 of $3,000, and one declared the project complete seven hours early, shortly after its own reviewer returned another Reject [15][1]. Instruction drift did the rest: both papers broke the length limit and would have been desk-rejected, and one shipped with zero figures in the main text where the human original had 15 [16]. Meta AI has described a related pattern it calls "behavioral state decay," where an agent honours a requirement early and violates it later while fixing something unrelated [17].

Two papers and one model is a thin sample, and the finding is about a specific scaffold as much as a specific model. Watch whether the method extends to more submissions and more vendors, whether the resource-awareness and context failures are fixable in scaffolding, and whether anyone claiming automated research is willing to be graded by the people who did the work first.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories