Published · 5d agoProduct3 min read
Princeton-led study: AI agents can do the engineering of AI research, not the research
Claude Opus 4.8 got six days, $3,000 in credits and a GPU budget to answer two unpublished NeurIPS questions. The papers' original authors graded the output and rejected both.
Not a builder's beat, but builders have a standing stake in it.See today for builders
What happened
- A multi-institution group of researchers led by Peter Kirgis and Sayash Kapoor at Princeton University conducted the study.
- The researchers found that AI agents could solve the engineering problems necessary to do AI research but lacked the judgment and creativity to produce original research at the caliber of papers accepted by a top machine-learning conference.
- The gap suggests that some of the hyped-up timelines for automating AI research may be running ahead of the evidence; the study suggests recursive self-improvement may take a while.
- The researchers proposed a new evaluation method called "shadow evaluation," which requires the AI to answer a research question from a high-quality unpublished paper.
- The researchers asked Anthropic's Claude Opus 4.8, running on open-source software called OpenClaw, to tackle the research questions.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
A multi-institution group led by Peter Kirgis and Sayash Kapoor at Princeton University gave AI agents the money, compute and time to produce a top-conference machine learning paper, and the human authors of the source papers rejected what came back [1][11]. That matters because the case for near-term recursive self-improvement rests on agents handling the open-ended part of AI research, not just the engineering scaffolding around it [2][3].
The industry pitch is that AI will soon improve itself with almost no human oversight, and LLMs already write code, generate synthetic training data and optimize the chips they run on, which is why forecasts of explosive progress treat recursive self-improvement as imminent [24]. Most existing evaluations in this area test narrow tasks with checkable answers, such as solving engineering problems or post-training small language models against a benchmark [22]. Actual research progress also requires picking a set of hypotheses, deciding what evidence would settle a question, and knowing when to start over [23].
The study's method, which the authors call shadow evaluation, hands the agent a research question drawn from a high-quality unpublished paper [4]. The two questions came from papers submitted to NeurIPS 2026: whether a language model's personas can be controlled by editing the model's weights, and how to design a detector that flags when a model predicting from spreadsheet data has become unreliable [6][7][8]. Because the papers were not public, the agents could not memorize the answers from training data or find them online [9].
The setup was Anthropic's Claude Opus 4.8 running on the open-source OpenClaw software, with six days, $3,000 in Anthropic API credits, a GPU budget, dedicated virtual machines and open web access [5][10]. That is roughly $500 a day in model credits before compute [1]. The original authors then graded the resulting papers as they would a conference submission, and rejected both [11].
The engineering held. The agents reviewed the literature, ran hundreds of experiments and compiled the results [12]. "On the other hand, the agents were unambiguously bad at carrying out the research itself," Kapoor says [13]. They ran odd experiments, in some cases testing hypotheses on tiny synthetic datasets, struggled to write intelligibly, and made no novel contribution to either field [14]. According to Kapoor, the papers "were nowhere close to the mark when it came to being at the quality of a top AI conference" [15].
The failure pattern is the operationally interesting part. The agents produced novel and ambitious hypotheses resembling those the human authors themselves began with, then abandoned them on the basis of very limited data [17]. They under-explored alternatives and committed to unpromising approaches too quickly [16]. They could make small pivots but could not rethink an approach or restart from scratch [18]. Given feedback from subagents and external AI reviewing tools, they narrowed their claims and added caveats instead of revising methodology [19]. They also used tokens, compute and time poorly, and did not follow instructions about how long to spend on each phase or how long the paper could be [20][21].
Worth keeping in proportion: this is two papers, one model, one harness [2]. Several of the failures listed read like scaffolding problems rather than intelligence problems, specifically budget adherence, instruction following and the inability to backtrack [18][20][21]. Watch whether the next round of agent frameworks moves any of those three, and whether shadow evaluation gets rerun across more unpublished papers and more models, since a method that depends on unpublished work has a narrow window before the papers go public [4][9].
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
A multi-institution group of researchers led by Peter Kirgis and Sayash Kapoor at Princeton University conducted the study.
ReportedView cited source - [2]
The researchers found that AI agents could solve the engineering problems necessary to do AI research but lacked the judgment and creativity to produce original research at the caliber of papers accepted by a top machine-learning conference.
ReportedView cited source - [3]
The gap suggests that some of the hyped-up timelines for automating AI research may be running ahead of the evidence; the study suggests recursive self-improvement may take a while.
ReportedView cited source - [4]
The researchers proposed a new evaluation method called "shadow evaluation," which requires the AI to answer a research question from a high-quality unpublished paper.
ReportedView cited source - [5]
The researchers asked Anthropic's Claude Opus 4.8, running on open-source software called OpenClaw, to tackle the research questions.
ReportedView cited source - [6]
The questions came from two papers submitted to the machine-learning conference NeurIPS 2026.
ReportedView cited source
Sources & coverage · 1 publisher
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- technologyreview.comMichelle Kim5d agoAI’s recursive self-improvement might not come so quickly after all
- technologyreview.comThomas Macaulay4d agoThe Download: AI’s self-improvement problem, and what’s driving the heat
Additional citations
- Sayash Kapoor, quoted by MIT Technology Review



