Build1 distinct publisher3 min readPublished
The Neuron gave it a browser, a Mac, a broken Blender install and an hour. What changed was how rarely it stopped to ask permission, which makes the next piece of work a harness problem rather than a prompt problem.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
A permission prompt does two jobs. It authorises the next action, and it gives you a cheap place to read the plan and kill the run. The Neuron's report that Fable 5.1 is more willing to keep working without asking every few minutes [3] takes away both at once. Longer stretches between returns mean a wrong turn costs forty commands instead of one, so the artifact you now have to get right is the harness: the allowlist, the working directory, the sandbox boundary, and the definition of done.
The Blender episode is the clearest exhibit. Their MCP setup was broken, Corey proposed letting the model repair it through computer use, and the team granted permission [9]. At 25:30 the model had found the ZIP and was installing it [10]. Read that as an agent editing its own toolset mid-task. The approval covered one install; the install changes what every later step can reach. If your permissions are scoped per action, that is contained. If they are granted once per session, you have handed over a standing capability.
Cat Doom is the team's own repeated house test, which they call a wonderfully unscientific benchmark, and their description of earlier attempts as code that technically ran and visually resembled regret is the most honest benchmark note I have read this month [6]. For that result to transfer to your queue, your task has to look like theirs: greenfield, no schema to respect, no credentials, and verifiable by watching it run. The extension is the part worth copying. They asked for 12 levels of rising difficulty and told the model to use subagents if needed [8]. That hands fan-out and spend decisions to the model, which is where a long run's bill actually gets set.
Which brings up what this material does not contain. The cost claim is the load-bearing one, and the writeup carries no token counts; pricing, safeguards and benchmarks sit in a separate launch piece, and one of this article's own section headings concedes that the cost story may matter more than the benchmark story [15][18]. The timestamps do bound the walltime. The best-yet verdict landed at 17:58 [7] and the ZIP install at 25:30 [10], seven minutes and 32 seconds later [16]. Grant's callout about computer-use speed came at 26:44, inside the first 45 percent of the roughly one hour they had allotted [11][17]. That is quick for an environment repair, and walltime is not a cost model.
Two other reads point the same way with different emphasis. Every's Dan Shipper ran a week of testing, Kieran Klaassen rebuilt Every's Proof editor from a single prompt and ran multi-day jobs, and Shipper says a computer-use Mac app called Hands was one-shotted after other models failed [12]. Klaassen's own framing was Fable-class depth plus a collaborator he can trust [13]. All of it is one publisher's account of its own and its colleagues' sessions.
In my context the review question changes shape. It stops being whether the prompt is specific enough and becomes what the agent can reach between checkpoints and what reversing that costs. Permissions are configuration you can inspect before the run. Judgement is something you can only assess after it.
Ranked by verification strength, evidence, and original report placement.
Fable 5.1 produced a playable ray-casting browser game within minutes; the weapon was a spray bottle, cats went to sleep instead of dying, the minimap worked, and the exit changed state after the enemies were cleared.
The team has made versions of Cat Doom with frontier models for a while, calling it a wonderfully unscientific benchmark, and says older versions often produced something that technically ran and visually resembled regret.
By the 17:58 mark of the stream, Corey and Grant were both calling it the best Cat Doom yet.
The team then asked Claude to design 12 levels with rising difficulty and harder cats, using subagents if needed, and later added catnip for distraction, yarn balls as grenade-like weapons, and boxes and bags as hiding mechanics.
The Neuron tested Claude Fable 5.1 live, giving it a browser, their computer, Blender, several coding tasks and roughly one hour.
In the session the model built the best version of Cat Doom the team had made, installed its own Blender MCP connection through computer use, turned a viewer's solar-system theory into an interactive 3D visualization, and made a Floppy Bird clone with a "flamingo speed" superpower.
Distinct publishers with included, body-backed reporting in this cluster.
2 articles · September 1, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
leadership
Anthropic ships a price dial with its new model, and that is now the buying decision1 distinct publisher
leadership
Anthropic cuts Fable 5.1 prices by 25% and launches two-tier safeguard system with Mythos 5.11 distinct publisher
leadership
Every's 30 people, four products and one cloned editor: the self-driving company in practice1 distinct publisher
invest
AIR Security banks $50 million six months in on a census of 17,800 untrusted AI add-ons1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One livestream, told twice
Strip the duplicate posting away and the whole story is one hour of one team's stream. The strongest details are the most verifiable — 17:58, 25:30, 26:44, a working minimap, a ZIP found and installed — but they are all events The Neuron staged, narrated and timestamped itself, on a test it happily calls unscientific. Every's week and Mollick's early access would be the corroboration; both arrive as The Neuron's summary of what other people said, and nothing here separates 'the model did this once, live' from 'the model does this reliably'.
Launch-week hands, nothing at scale
Three sets of hands total: The Neuron for an hour, Every for a week, Mollick under early access. That is enough to say the model is shipped and being driven hard by people close to the launch, and not nearly enough to say anything about production use. No customer counts, no deployment, no cost per run — and the piece's own cost argument is left for a different write-up, so even the economics claim has no adoption footing.
Restrained framing, vibes underneath
The Neuron does more self-limiting than most launch-day coverage: it says outright that none of the game features prove Fable 5.1 is the smartest model, and it reframes its own result as the shrinking distance between a messy idea and a working artifact. The overshoot is in the load the anecdotes are asked to carry. 'Best Cat Doom yet' and 'easier to delegate' are feelings recorded on a stream the team ran; 'more economical on long agent runs' is an economic claim with no number attached anywhere in the piece.
Everyone here got the model early
Follow the access. Mollick's game was made under early access, Every had the model for a week before the public did, and The Neuron built a launch-day livestream plus two cross-linked articles out of it — the piece routes readers to its own breakdown and its own video. None of that makes the observations false, but every enthusiastic voice in the story is someone the vendor's launch cycle put in front of the model first, and Klaassen's 'try it again' pitch to earlier skeptics is marketing shaped like advice.
Believable, unconfirmed
We are fairly confident about what happened on that stream and much less confident about what it means. The timestamps and artifacts are hard to fake in a live recording; the durable claims — faster, cheaper, more autonomous, safer to leave running — depend on repetition nobody has supplied yet, and there is no second publisher in this story to check any of it.