Skip to content

Build1 publisher3 min readPublished

Discovery Loop bets the research bottleneck sits in taste and evaluation

Oriol Vinyals has left Google DeepMind to build Discovery Loop with Jeff Dean, Sanjay Ghemawat and Quoc Le, on the argument that AI already writes the code and runs the experiments while idea generation and judging results lag.

The Engineer · Build desk

Photograph accompanying Discovery Loop bets the research bottleneck sits in taste and evaluation
Photo: radical.vc

What happened

  • Oriol Vinyals, until recently VP of Research at Google DeepMind, spoke at the Agentic AI Summit 2026 days after leaving, arguing that recursive self-improvement is coming but that a sudden intelligence explosion is unlikely.
  • Direct tests hand a system a metric and a compute budget and measure how much it improves itself, but each evaluation runs an agent for hours on tasks far removed from what ultimately matters.
  • Vinyals has co-founded a startup, Discovery Loop, with Jeff Dean, Sanjay Ghemawat and Quoc Le, with the stated aim of automating the scientific research process end to end.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost The measurement that would settle whether a system improves itself costs hours of agent compute per data point, so the cheap proxy scores keep getting quoted in its place.
  • constraint There is no studied training recipe for research taste, so the step Vinyals names as the bottleneck cannot be optimised the way a SWE-Bench score can.
  • decision Anyone reading SWE-Bench-style numbers as evidence of self-improvement is buying a measure of the two steps that already work, and needs a different test to see the other two.
  • capability A tenfold speedup confined to coding and experiments leaves cycle time governed by the unautomated steps, which on a half-and-half split works out at roughly 1.8x end to end.

Point an agent at its own training and the first job is naming the target. Vinyals lists what an AI system could change about itself: its weights, its training data, its training methods, the instructions it receives with every query, external tools such as database access and code execution, and the metrics it uses to track its own progress [1]. Each of those carries different technical and regulatory problems, according to Vinyals [2].

The cycle he describes has four steps: a promising idea, code that implements it, experiments that test it, and a reliable way to judge whether the change helped [3]. AI is already making progress on the middle two. Idea generation and evaluation are where systems still fall short [4].

That split is awkward for the way labs currently keep score. Vinyals says most of them track self-improvement indirectly, through capability benchmarks like SWE-Bench Pro or ML-Bench, climbing the leaderboard and hoping self-improvement emerges as a side effect [5]. Those tests are cheap and well defined, and they mainly cover implementation and experimentation, the steps that already work [6].

A direct test looks different. The system gets a metric and a compute budget, and researchers measure how much it improves itself [7]. Each evaluation runs an agent for hours on tasks far removed from what ultimately matters, and that is expensive [8]. Vinyals' own example has the agent optimizing Tetris while the real goal is to automate an entire research lab and build the world's best model [9]. From years of building game-playing agents, he says, systems exploit objectives in unexpected ways, beating the scoring system instead of playing the game [10].

Grading ideas is harder still. Vinyals calls the missing piece "research taste", the instinct for which ideas are worth pursuing, and says nobody in LLM training has really studied how to teach it [11]. He expects future evaluations to grade how a system got its result, on the criteria conference reviewers use: originality, elegance, efficiency, and whether a technique stands the test of time [12]. Some of that can be written down as rules, checked through reward models and trained on with reinforcement learning, though he says doing it is very hard and will take more time [13]. Human review is expensive as well, and by his account not particularly good at spotting strong ideas either [14]. Anyone who has had a paper reviewed will recognise the complaint.

The ten-times claim and the no-explosion claim are compatible. Vinyals expects AI to speed up certain research and engineering tasks by a factor of ten or more, and considers a sudden, self-accelerating intelligence explosion unlikely [15]. Take an assumption he does not make: implementation and experimentation are half the elapsed work in a research cycle. A tenfold speedup on that half leaves 0.5/10 + 0.5 = 0.55 of the original time, about 1.8x overall [16]. He also points to a floor underneath all of it, that chips cannot compute faster than their design and the speed of light allow [22].

Discovery Loop is co-founded by Vinyals with Jeff Dean, Sanjay Ghemawat and Quoc Le, and its stated aim is to automate the scientific research process end to end [17]. Vinyals spoke at the Agentic AI Summit 2026 days after leaving Google DeepMind, where he was VP of Research and worked on AlphaStar, AlphaCode and Gemini [18]. The-decoder's account does not say whether Dean, Ghemawat or Le have left Google [20]. In the early phase, according to that account, humans and machines will form hypotheses together [19].

What to watch

  • Publication of a direct self-improvement benchmark that states its compute budget and per-run cost.
  • Confirmation of whether Dean, Ghemawat and Le have left Google, which the-decoder's report does not address.
  • Whether Discovery Loop releases a rubric or reward model for research taste that other labs can run.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories