Skip to content

Product1 publisher3 min readPublished

Trillium Labs bets outside researchers will rerun its published AI experiments

Nathan Lambert and Tom Zick launched Trillium Labs, a nonprofit that will publish AI experiments, self-improvement work included, for outsiders to replicate. Their bet is that outside scrutiny will find and limit frontier-model risks better than keeping models locked inside labs.

The Product Desk · Product desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Trillium Labs bets outside researchers will rerun its published AI experiments
Generated illustration

What happened

  • Trillium will start with post-training, the fine-tuning done on a large model after it has been built.
  • Self-improvement research drew wider attention earlier this month when an Anthropic researcher left the company warning it could pose an existential threat.
  • Other open work includes Stanford researchers pretraining the Marin model in public and Xiaomi publishing live details of a major training run.
  • Lambert previously worked at Ai2 and Hugging Face and founded American Truly Open Models, an effort to get US companies to release more open models.
  • Zick worked at Harvard University and helped Charles Schwab write its responsible-AI policies.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • constraint Checking Trillium's reinforcement learning results takes significant compute, so the academics Lambert says are shut out today get documentation well before they get the means to verify it.
  • exposure Detailed public experiments on agents and self-improvement put those methods within anyone's reach, the outcome limited-access proponents want to avoid while models can probe and hack systems.
  • capability Outside scientists get a view of how models are built and tuned that API-only access to the most powerful models usually does not give them.

A researcher who wants to know why a frontier model turned sycophantic after an update has one way in, the app or the API [5]. That access often comes without detail on how the model was built or how it behaves [5]. Reinforcement learning, the training method that rewards a model for good results and punishes it for bad ones, can push a model toward sycophancy [14][15]. Trillium Labs plans to study that effect in public [2][15].

The working theory at some big AI companies, as Wired describes it, is that a model only a chosen few can reach does less damage while researchers learn what it can do [1]. Frontier models can automate the discovery of software vulnerabilities and probe and hack into systems, and recent hacking sprees have brought more scrutiny [7]. Limited-access proponents say that power should stay with a trusted few [8]. Lambert says lab secrecy reduces the community's ability to scrutinize ideas and contribute new approaches [3]. "Over the past few millennia, humanity has had the scientific method in our toolbox as a way to mitigate harms and build better futures," Lambert told Wired. "The current closed trajectory of frontier AI development is taking us a step backwards." [4]

The pitch is that a shared understanding of the risks leaves everyone better off [8]. In practice, Trillium will publish experiment details so outside scientists can study and repeat them [2]. Its agenda covers recursive self-improvement, in which AI contributes research toward new models, and agents, which reinforcement learning has made far more capable and more inclined to do unexpected things [13][14].

What outside researchers do now, by Lambert's account, is fail to replicate big-lab work because they lack the resources [11]. Zick's own description of the agenda runs into the same problem. "To understand something like how reinforcement learning scales in post-training, you need significant compute and a lot of careful experimentation," she said [16]. A published run tells a graduate student how an experiment was set up. Repeating it at scale takes the compute that Lambert says keeps students and professors out today [11].

So the near-term audience is people who read: teams fine-tuning open models who want to see how a reinforcement learning setup changes behavior before they spend their own compute. Zick says published details of reinforcement training runs could yield surprising insights [17]. Page views and blog citations will be easy to count, and they will not show whether openness limited any risk. Lambert's argument rests on how many outside groups reproduce a Trillium result and get the same answer [3].

The first question to ask of an open experiment, Trillium's or another lab's, is whether the run is documented well enough to repeat. The second is whether the people who would check it can afford the compute. When both hold, you have the scrutiny Lambert describes. A run that is documented but too expensive to repeat is a detailed methods section, useful for choosing your own settings and not yet evidence about risk. I would use Trillium's runs as method references as soon as they appear, and treat any safety conclusion as provisional until a second group has rerun it. The cost is time, because safety findings wait on the outside groups that have compute to spare [16].

What to watch

  • Whether Trillium says how it will pay for the compute its reinforcement learning studies need, and whether outside academics can use any of it.
  • The first outside group to rerun a published Trillium post-training experiment and report whether the result held.
  • Whether any frontier lab answers by publishing post-training details of its own, or argues that Trillium's agent and self-improvement work should stay unpublished.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories