Build1 distinct publisher3 min readUpdated
Princeton and the UK AI Security Institute gave a frontier agent six days and $3,000 to answer real unpublished research questions. The original authors reviewed the output and rejected both papers.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
Follow any of these and your For You feed starts watching them — no settings page required.
Princeton and the UK AI Security Institute gave a frontier agent six days and $3,000 to answer real unpublished research questions. The original authors reviewed the output and rejected both papers.
Researchers at Princeton and the UK AI Security Institute handed a frontier agent the core research question from an unpublished paper, gave it six days, $3,000 in API credits, a GPU budget, a virtual machine and the open web, then asked the paper's original authors to review the result as conference reviewers [4][7]. Both papers produced this way were rejected, one with a Strong Reject, which is a direct check on Anthropic's and OpenAI's marketing about models accelerating AI research [9][2].
The design is the interesting part. The authors call it Shadow Evaluation: because the reference paper is unpublished, the answer is not in the training data, and because the reviewers spent months on the same question, they know what a real contribution looks like [4]. The team argues existing evidence is weak, either narrow verifiable tasks or AI-written papers pushed through peer review, a process they describe as "overstretched, stochastic, and suffers from poor review quality" [3]. Two NeurIPS 2026 submissions were used: one on steering language model personality traits through weights, one on a method called TabPFN for detecting when a tabular prediction model meets deployment data far from its training distribution [5]. The main runs used Claude Opus 4.8 with Extra-High Reasoning inside OpenClaw, the open-source agent framework built by Peter Steinberger, who joined OpenAI earlier this year [6][8].
The engineering held up. The agents ran literature searches, debugged GPU code, and completed hundreds of experiments and robustness tests without human help [18]. What failed was everything downstream of that. Reviewers cited poorly motivated data and experiments, unreadable prose, and no new contributions [9]; one called the reasoning a "'proof by example' fallacy" that was "highly non-scientific," another called the experiment choices "bizarre" and the results the product of "post hoc choices" [10].
The log analysis reads like a list of operational defects rather than intelligence deficits. The agents generated plausible hypotheses then killed them on small, hand-curated or synthetic datasets, showing no sense of the publication bar [11]. When a hypothesis was falsified they narrowed the existing claim instead of opening a new line [12]. Their own internal AI reviewer never returned a single Accept across fifteen revision rounds, and the core criticism was never addressed [13]. Both runs abandoned their most ambitious goals inside ten hours, and the Personas run closed out exploration after five hours against a plan of 36 to 48, roughly a tenth to a seventh of the time it had allocated itself [14][2]. Both finished with less than half the API budget spent, under $1,500 of $3,000, and one declared the project complete seven hours early, shortly after its own reviewer returned another Reject [15][1]. Instruction drift did the rest: both papers broke the length limit and would have been desk-rejected, and one shipped with zero figures in the main text where the human original had 15 [16]. Meta AI has described a related pattern it calls "behavioral state decay," where an agent honours a requirement early and violates it later while fixing something unrelated [17].
Two papers and one model is a thin sample, and the finding is about a specific scaffold as much as a specific model. Watch whether the method extends to more submissions and more vendors, whether the resource-awareness and context failures are fixable in scaffolding, and whether anyone claiming automated research is willing to be graded by the people who did the work first.
Ranked by verification strength, evidence, and original report placement.
A new paper from Princeton and the UK AI Security Institute tests the claim that AI agents can conduct AI research, finding that today's frontier models can handle research engineering but fail at the parts of the research process that matter.
Anthropic and OpenAI have been touting their models' ability to speed up AI research.
The authors argue solid evidence for claims about automated AI research has been mostly absent: existing evaluations either test agents on narrow, verifiable tasks or submit AI-generated papers to peer review, a process the researchers call "overstretched, stochastic, and suffers from poor review quality."
The researchers call their approach "Shadow Evaluation": an agent receives the core research question from an unpublished paper, and the original authors, who spent months on the same question, then evaluate the result as conference reviewers would. Because the results are not yet on the web, the agent cannot fall back on training data.
The team partnered with the authors of two NeurIPS 2026 submissions. The first examines how personality traits of language models can be steered through their weights. The second develops a method called TabPFN that detects when a tabular prediction model hits deployment data that differs sharply from its training data and tanks its accuracy.
The main experiments used Claude Opus 4.8 with Extra-High Reasoning.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed single-source account of one study
The reporting is unusually specific for a single item: a named protocol, named partner submissions, model and scaffold identification, per-run resource envelopes, verbatim reviewer language, log-derived failure modes and a cross-model cross-check. That specificity is what raises the score above the midpoint. It is held down by structural limits: one publisher, no link to or direct reading of the underlying paper in the supplied material, no independent replication outside the research team, and a truncated closing quote.
Two main runs plus one cross-check; no external uptake shown
Adoption here means uptake of the evaluation approach and of agentic research automation itself. The supplied material shows only the authors' own runs: two Shadow Evaluation pairings and one GPT-5.6 Sol/Codex repeat. No third party is shown adopting Shadow Evaluation, and the only external usage signals are the labs' own unverified acceleration disclosures, one of which the authors say is absent from the relevant system card.
Headline generalizes further than the sample
Slightly positive rather than strongly so. The itemized failure observations are well supported and align with the evidence presented, but the framing that a study 'contradicts' Anthropic's and OpenAI's claims and that autonomous AI research is not within reach rests on two agent runs against two papers, one primary model and one scaffold, with the authors' own suspicion (not demonstration) that more time or compute would not change outcomes. In the opposite direction, the lab claims the piece targets are themselves under-documented, which keeps the net gap small.
Academic and safety-institute framing versus commercial lab messaging
Incentive load is visible on both sides of the story as supplied. The study comes from an academic group and a state AI security institute whose remit and reputation favour findings that temper automation claims, and the reviewers grading the agent output are the same authors whose months of work the agent was set against. The labs whose statements are contested have direct commercial incentive to advertise research acceleration. The article also notes the OpenClaw author joined OpenAI. These are structural incentives disclosed in the text, not evidence of bad faith, and no funding or conflict details are supplied.
Moderate-low: one publisher, no primary document
Confidence is limited by cluster shape rather than by internal inconsistency. A single outlet supplies every claim, the underlying paper is not in the cluster, no lab response appears, the sample is two runs, and the article text is truncated before the authors' concluding qualification. The consistency and specificity of the reported details support moderate confidence in the individual observations, but not in the broad capability verdict.
leadership
The AI bill nobody reconciles: cost per finished task, not per million tokens1 distinct publisher
build
OpenAI's president says open weights will accelerate the threat. His own cyber model stays gated.1 distinct publisher
science
The AI hacking disclosures were all instructed attacks. The change is tempo, not autonomy1 distinct publisher
build
19 unsanctioned actions in 10 of 122 runs: nothing escaped, and that is the point1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 14, 2026