Build1 distinct publisher3 min readUpdated
A new paper says Sora, Genie 3, JEPA and Marble carry no mental state at all. Its own benchmark, 420 of 448 scenes without motion, tests language models instead.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
Follow any of these and your For You feed starts watching them — no settings page required.
The systems the paper indicts are generative world models: Sora, Genie 3, JEPA, Marble [1]. The systems it measures are eight language models from OpenAI and Anthropic [9]. That substitution deserves attention before the headline numbers do. Menti-Bench consists of 320 text descriptions, 100 picture stories and 28 sound-video clips [7], which puts 420 of its 448 scenes, 93.8 percent, outside anything a video model would consume [1]. The demonstrated result, then, is that language models choose better actions when forced to write a belief state down first. Whether a rendering model gains the same from a bolted-on mental layer is a different experiment.
Inside the ablations sits a number the framing plays down. Removing the mental channel costs 12.1 F1 points on average; removing the physical channel costs 16.5 [12]. Physics is carrying 4.4 more points than mind on the authors' own benchmark [2], which is the correct result for a paper arguing that mental variables are a missing addition rather than a replacement. The figure with architectural teeth is 6.4, the loss when the physical and mental transitions are predicted independently instead of jointly [12]. The obvious build, a scene predictor running beside an intent classifier with a merge step at the end, is precisely the configuration that gives those points away.
The compute comparison rewards a close read. Direct answering scores 63.3, six-sample self-consistency 77.9, the full pipeline 87.9 [10], so the scaffold buys 10.0 points past what majority voting already extracts [4]. GPT-4.1 inside the pipeline reaches 84.9 and beats GPT-5.6-Sol answering directly with self-consistency at 83.6 [11], a margin of 1.3 [6]. Thin, but the ordering holds, and weaker base models gain more from the explicit structure than stronger ones [13], which is what you would expect if the scaffold supplies something the model was otherwise improvising.
The gain also concentrates where the benchmark is thickest: 26.4 points on interpersonal scenes against 14.0 on object-focused ones [13], nearly double [7], with 78 percent of scenes containing at least two characters [8]. The benchmark was built to have the property the framework exploits. That is normal for a paper introducing both, and a reason to want someone else's scenes.
Humans score 98.5 against the pipeline's 87.9 [10], a gap of 10.6 points [5] on six-option questions with a human-written reference solution [7]. For a service robot that gap is the whole job. The authors' own illustration is a cup slid across a table, which can be an apology, a deception or an act of care, and only the world model holds the variables that tell them apart [4]. To their credit they do not claim to simulate anything: mental states here are hypotheses inferred from behaviour and context, not measurements, and they ask that systems built this way keep their uncertainty and assumptions visible [5]. The pipeline at least writes a machine-readable intermediate at every stage, so a wrong answer can be traced to the step that produced it [6].
Ranked by verification strength, evidence, and original report placement.
Menti-Bench contains 448 decision scenes: 320 text descriptions, 100 picture stories and 28 sound-video clips. Each scene has six response options and a human-created reference solution documenting the correct action plus the underlying mental and physical states.
78 percent of Menti-Bench scenes involve at least two characters.
The paper argues that existing world models including Sora, Genie 3, JEPA and Marble model only the physical layer of the world (objects, positions, motion, occlusion), and that what people believe, want or consider socially appropriate never appears in their state space.
The authors' example: if someone's cup is moved into a cabinet while they are not looking, the scene looks correct to a purely physical world model but it still predicts the wrong next action; only a model tracking the person's belief about the cup's location explains what they will do.
The framework, called Mental World Modeling (MWM) and published on GitHub, extends classic world models with mental variables including beliefs, attention, goals, intentions, emotions, norms and social relationships; the target agent sees only an egocentric partial view while the world model holds the complete state.
Every action splits into a physical carrier (speaking, pointing, grasping) and a mental payload (comforting, deceiving, rejecting); the same gesture of sliding a cup across a table can be an apology, a deception or an act of care, and only the world model holds the variables that distinguish them.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed but single-sourced and author-reported
The quantitative record is unusually specific for a preprint write-up: a full F1 ladder, three ablations, scene-type breakdowns and oracle-substitution diagnostics. But all of it reaches the cluster through one outlet relaying the paper's own authors, with no peer-review status, author affiliation, independent replication or third-party run of Menti-Bench, and the systems the paper indicts (Sora, Genie 3, JEPA, Marble) are never actually measured.
No adoption evidence beyond the authors' own release
The cluster shows a GitHub publication and the authors' own benchmark run, but no third-party usage, downloads, stars, forks, integrations, deployments or vendor uptake. Nothing supports an adoption reading, and inferring one from a repository announcement would be guessing.
Verdict on video world models outruns a static-scene LLM test
The framing claim is that physics-only world models cannot predict people, naming Sora, Genie 3, JEPA and Marble - yet the evidence offered is eight language models answering six-option questions on a benchmark where 420 of 448 scenes (93.8 percent) have no motion at all, and none of the named systems is evaluated. The article does carry deflating detail (humans still 10.6 points ahead, roughly 80 percent of the residual gap in transition simulation, an explicit disclaimer about not simulating consciousness), which keeps the gap moderate rather than severe.
Authors grade their own framework on their own benchmark
The framework, the reference implementation and the evaluation set all come from the same team, and the scoring protocol and human reference solutions are theirs, so favourable results are self-adjudicated. The outlet's incentive layer is visible too: the piece closes on Hassabis's 'ChatGPT moment' expectation and hundreds of millions flowing into world-model startups, framing the paper against an active funding narrative. Nothing in the cluster discloses funding, affiliation or competitive relationships, so this is a structural reading rather than a documented conflict.
Internally consistent, externally unverified
Confidence is limited by a single publisher, a single self-reporting research team and zero external adoption or replication signal. It is not lower because the reported figures are specific, internally coherent (the arithmetic in the derived claims checks out), accompanied by ablations and failure-mode diagnostics, and hedged with an explicit no-consciousness caveat.
leadership
The AI bill nobody reconciles: cost per finished task, not per million tokens1 distinct publisher
build
OpenAI puts latency on the price list: 750 tokens/sec, gated by workload fit3 distinct publishers
invest
Speed becomes a SKU: OpenAI and Google put a separate price on latency3 distinct publishers
build
Grok 4.6 lands in Copilot two days after launch, and the model picker becomes a procurement problem1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 22, 2026