Skip to content

Invest1 publisher3 min readPublished

If agents can't do open-ended research, price compute against task automation

A new study says AI agents still fail at free-form research. The self-improvement story is the load-bearing beam under a lot of compute spending, and it just got a crack in it.

The Investor · Invest desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Researchers found that AI agents still cannot conduct open-ended AI research, described as free-form investigations with no clear-cut answers that require the judgment and creativity needed to make genuine breakthroughs.
  • MIT Technology Review describes the AI industry's boldest current promise as the claim that AI will soon improve itself, with almost no need for human oversight.
  • The open question raised by the study is how crucial open-ended research is to recursive self-improvement, and whether AI systems can grind their way there without it, simply by improving on narrower tasks.
  • OpenAI has paused some model work over safety concerns, saying its Astra model reached a "critical" risk threshold.
  • The slowdown at OpenAI sets it apart from Anthropic's approach.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

Researchers have found that AI agents still cannot conduct open-ended AI research: free-form investigations with no clear-cut answers, the kind that require the judgment and creativity behind genuine breakthroughs [1]. That matters to anyone underwriting compute, because the industry's boldest promise right now, as MIT Technology Review puts it, is that AI will soon improve itself with almost no need for human oversight [2].

Handle the finding carefully. The newsletter item reporting it does not name the researchers, the institution, or the publication venue [9], so this is a directional signal rather than a settled result. The framing it leaves behind is the useful part: the open question is how crucial open-ended research actually is to recursive self-improvement, and whether systems can grind their way there anyway by getting better at narrower tasks [3].

Those two paths carry very different price tags. A takeoff story lets you justify almost any capex number, because the terminal value is unbounded and the discount rate stops mattering. A grinding story does not. It makes compute an input to a series of specific, boring, measurable substitutions, each of which has a customer, a budget line, and a competitor. If you are buying the second thing at the first thing's multiple, the study is your problem, not a footnote.

The rest of the day's tape supports the boring reading. OpenAI has paused some model work over safety concerns, saying its Astra model reached a "critical" risk threshold, according to the Guardian [4], and Axios reports the slowdown sets it apart from Anthropic's approach [5]. Frontier capability schedules are not monotonic, and they are not fully within the labs' control.

Meanwhile the money is voting for embodiment and throughput. Chinese humanoid maker Unitree surged 629 percent in its stock market debut [6], which puts the shares at roughly 7.3 times the offer price [1]; the BBC describes it as the world's biggest humanoid firm and already profitable [7]. Profitable is the word that does work there. On the input side, China is allowing Nvidia's H200 chips into the mainland, with ByteDance and Tencent each recently receiving about 10,000 processors [8], roughly 20,000 units between them [2]. That is compute flowing to companies with existing revenue to defend, not to a recursion thesis.

And the near-term product still needs a human on the other end of it. A Kentucky mother told local station WDRB that AI-generated educational materials her son brought home said "Arizona is Arizone" and that "Illinois starts with a V" [10]. Cheap output with an error rate is a service business with a review cost, which is exactly how narrow automation gets priced.

What to watch: whether the study is published with methods that let anyone replicate the open-ended research failure, and whether it survives contact with the labs' own agent evaluations. Watch whether capability marketing shifts from self-improvement language toward task-level benchmarks, which would be an admission of where the revenue is. Watch whether OpenAI resumes Astra work and defines what "critical" meant [4]. And watch whether H200 shipments scale past the initial 10,000-unit tranches [8], because that flow tells you which of the two stories the buyers are actually funding.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories