Skip to content

Build1 publisher3 min readPublished

PhAI Labs' JEPA-Anything outpredicts a standard JEPA in most tests across seven fields

PhAI Labs' JEPA-Anything splits a world model's prediction into several modules, cutting error 35% against a standard JEPA on a Pong task. Its reported comparisons use that one baseline, and its liver cancer pick has only lab and mouse data.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Between fields the team changed only how the data was prepared, and the modules were left to find their own roles during training.
  • On the Burgers fluid-dynamics benchmark, error fell by nearly half in one evaluation, but the advantage shrank to about 3 percent over 50 prediction steps.
  • Image tasks showed only a small difference, and in simulated robot locomotion the standard JEPA beat it in one of three environments.
  • Trained on simulated orbits with no physical quantities supplied, the model recovered Kepler's third-law exponent as -1.4991 against the true -1.5.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint With the same network minus the split as the only baseline, the results support the split as a design change and leave open whether it can displace any field's specialist model.
  • decision Teams that need long rollouts would have to benchmark at their own horizon before adopting it, given how much of the Burgers gain was gone by step 50.
  • capability Unlabeled modules that still yielded a testable drug pairing and a physical exponent mean the partial predictions can be inspected after training as a source of hypotheses.

A JEPA predicts an abstract summary of a missing or future state. It does not try to reconstruct raw data such as pixels [4]. The authors object to the standard design because everything goes into one prediction, so easy patterns drown out hard ones [5]. JEPA-Anything cuts the predicted state into several parts, and each part gets its own prediction module [6]. An added constraint pushes the modules toward different aspects of the state. The model then reassembles their partial predictions into a full one [6].

I like where this puts the generality. One architecture and one constraint carry across fields, and the per-field work stays in data preparation [7]. Leaving the module roles unassigned also spares everyone a review meeting about what module three is for.

The case for a universal world model starts from the claim that each domain has typically needed its own model [3]. The team, led by PhAI Labs with collaborators from Stanford, Oxford and Princeton [2], compared JEPA-Anything with a standard JEPA of the same architecture, trained on the same data under identical conditions [8]. As an ablation it is clean. It isolates what the split buys. The article does not report results against the specialist models each field already runs. For the replacement claim to transfer, the split model would have to beat each field's incumbent on that field's own metric.

Against its unsplit twin, the authors say it won consistently across ten tasks in physics, robotics and weather forecasting [21]. On Pong it also cut error by 13 percent on combinations of interventions it never saw in training [9]. In the simulations of water, quartz, acetaminophen and benzene it still scored best after 100 steps [11]. It assigned single-cell types more reliably, and it did slightly better at predicting more than 1,000 possible clinical disease events [12]. The Burgers decay is the result I would want explained before trusting this on long rollouts.

The liver cancer candidate came from inspecting the modules. The team analyzed the partial predictions a model learned from gene activity, protein levels and CRISPR screens [14]. Its top pick paired IL-18, a signaling molecule that activates immune cells, with blockade of CD73, an enzyme tumors use to suppress nearby immune responses [14]. Testing covered liver cancer cells co-cultured with immune cells, organoids and tumor tissue from three patients each, and mice [15]. T cells and natural killer cells showed stronger activation under the combination [16]. The article's summary says the kill result held in mice too [22]. Its detailed account attributes that result only to the organoids and tissue samples [16]. The study does not establish whether the pairing could become a therapy [17]. What the candidate shows about the method is narrower: the partial predictions produced a hypothesis that survived a first round of lab work, on samples from three patients per assay [15].

On orbits, the learned exponent sits 0.0009 from Kepler's -1.5, an error of about 0.06 percent [20]. The team evaluated one training run [19].

What to watch

  • A head-to-head against a field's own specialist model, such as a dedicated fluid or weather solver, on that field's benchmark.
  • Additional training runs for the orbit experiment showing whether the -1.4991 exponent holds across runs.
  • Whether the IL-18 plus CD73 blockade pairing moves past samples from three patients per assay and mouse work into larger studies.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories