Skip to content

Build1 publisher3 min readPublished

Persuading a frontier lab now takes 100B-parameter scaling laws, Geodesic says

Geodesic Research has published what six months of buying GPU hours taught it about scaling an independent safety agenda, including the model size its frontier lab advisors say a persuasive result now needs.

The Engineer · Build desk

Illustration accompanying Persuading a frontier lab now takes 100B-parameter scaling laws, Geodesic says

What happened

  • Geodesic Research says six months of procuring compute for its safety agenda surfaced non-obvious bottlenecks that can stop independent non-profits from scaling their research quickly.
  • Its post argues that philanthropic money can scale compute and staff time directly, but that some of the bottlenecks it encountered cannot be resolved by funding alone.
  • An appendix sets out a large multi-year compute deal Geodesic recently finalised, together with the strategy it followed through the procurement campaign.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost If persuasive safety work has to be shaped like frontier training, philanthropy is paying for capacity that produces no product and no revenue at the end of the run.
  • constraint Runs that occupy GPUs for multiple weeks cannot be bought at the moment the experiment is ready, so capacity has to be contracted ahead of the research calendar.
  • decision Handing infrastructure and empirical work to agent teams raises GPU hours per researcher, so hiring plans and compute contracts have to be sized against each other.
  • exposure An organisation that stays at 7B parameters with SFT-only post-training risks having its results discounted by the lab safety teams it is writing for.

The post defines intelligence as the time staff spend scoping ideas, engaging advisors, designing and running experiments, and disseminating results, and it gives that time one target: saturating compute with productive, impactful research [9]. Read the other way, it is a purchasing spec. When compute is constrained, Geodesic writes, it cannot productively scale intelligence and is left with "an overhang of research ideas we cannot execute" [8].

Anyone underwriting a spec like that wants to know where the ceiling sits. Geodesic says it has generated detailed forecasting documents on how much further compute it could saturate, and invites readers who would find them helpful to get in touch [6].

That ceiling moves while the deal is being negotiated. Day-to-day research increasingly hands low-level tasks to teams of agents that handle both infrastructure and empirics, multiplying effective headcount. The post expects a growing share of intelligence to come from AI-provided labour in the near term [10]. A researcher supervising several agent-run sweeps consumes more GPU hours per week than the same researcher writing their own training scripts.

The scale requirement is sourced to Geodesic's frontier lab advisors, and it is a claim about one audience. Independent research often focuses on models of 7B parameters or below with simple post-training such as SFT-only, and findings there "do not always scale to the frontier" [11]. To persuade frontier lab researchers, the post says, independent organisations increasingly need compute-intensive research with realistic pipelines: scaling laws with 100B+ parameter models, and multi-stage, agentic post-training [12]. Divide 100 by 7 and the floor of that jump is about fourteen times the parameters [13], before token budgets, optimiser states, or the several runs a scaling law needs rather than one.

The bar binds on your own agenda only if your intended reader is a lab safety team and the effect you are measuring changes with scale. A result written for a regulator, or one that saturates below 7B, can be correct without the 100B run.

Most of Geodesic's compute goes to training runs that range from single-day post-training to multi-week pretraining [7], and the deal described in the appendix is multi-year [4]. In my view the term to negotiate hardest in an agreement shaped like that is the commitment period, because a run that occupies GPUs for three weeks has to sit inside a reservation that already exists. The supplied text breaks off mid-sentence while describing the scale Geodesic works at; price, GPU-hour volume and counterparty are all undisclosed [15].

The post's own claim is narrower than the framing it invites. Geodesic says philanthropic funding can scale both intelligence and compute directly, and that what it hit were "non-obvious bottlenecks that funding alone can't resolve" [3]. A bottleneck that funding alone can't resolve is a different claim from grant size having stopped binding, and the text here does not test the stronger version. What it does argue is that reducing the resource gap, preparing organisations to ride potential funding waves, and forecasting compute supply crunches are old ideas that need more public discourse unified with first-hand decision-making [5].

What to watch

  • Whether Geodesic publishes the term length, price or GPU-hour volume of the multi-year deal described in its appendix.
  • Whether other independent safety organisations take up the offer of the saturation forecasts, or publish forecasts of their own.
  • Whether philanthropic funders start asking applicants for a saturable-compute ceiling as part of a compute grant.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories