Security1 distinct publisher2 min readPublished
The 50% time horizon measures replaceable serial human labour, not autonomous runtime, and the published interval on the frontier measurement is about as wide as a year of the curve it sits on.
The Watch · Security desk

Compiled by The WatchSomething wrong?How this is made
Put the note's own two numbers on one axis. Horizons have risen about 6x in the nine months since the paper [17], which works out to a doubling roughly every 3.5 months [18]. The published 95% interval on the current frontier measurement spans about 11x end to end [19], which is three and a half doublings, so at the trend's own rate the interval around a single model is as wide as roughly a year of progress [20]. A plan that says capability crosses some hour threshold in a named quarter is quoting a measurement that cannot resolve the quarter.
The domain spread is the part that should reach security planning. Translated onto the same doubling rate, the visual computer use gap is five to seven doublings, or roughly 18 to 23 months behind the software and research figure [21]. Attack chains that need a model to drive a screen, whether that is a console session or an operator UI, sit on the lagging branch. The physical case needs no benchmark: METR's author puts Claude 4.5 Sonnet's real-world coffee-making horizon at about two minutes [9].
The 50% is per task attempt, and when a model does succeed it is usually much faster than the human baseline [3]. The hours on the y-axis are human hours priced at a coin flip.
The soft spots in the level mostly point one way. Aggregating multiple successful baselines with the geometric mean rather than the arithmetic mean holds averages about 25% lower [16], and failed baselines were excluded rather than read as evidence that a task takes longer than the time the baseliner burned, which the author says would increase measured task lengths [15]. Baseline conventions alone could move the figure by more than 1.25x [13]. Pushing the other direction, the baseliner pool was skilled professionals in software engineering, machine learning and cybersecurity, and the author notes top engineers such as lab employees would be faster [14].
The rate survives all of this better than any single level does. Roughly 3.5 months per doubling is what the evidence carries; the position of any one model on that line is good to a factor of a few, and a factor of a few is the entire distance between one planning assumption and its alternative.
Ranked by verification strength, evidence, and original report placement.
A METR note by one of the main authors of the time horizon paper says that while the author still believes the core results, many people overstate the precision of the time horizon measurements and draw conclusions the evidence does not fully support.
Time horizon is not the length of time AIs can work independently; it is the amount of serial human labour they can replace with a 50% success rate.
METR reports Claude Opus 4.5 with a 50% time horizon of around 4 hours 49 minutes, with a 95% confidence interval of 1 hour 49 minutes to 20 hours 25 minutes, generated via bootstrapping.
The author writes: 'I really have no idea whether Claude's "true" time horizon is 3.5h or 6.5h.'
Error bars have historically been a factor of about 2 in each direction, and are worse with current models such as Opus 4.5 as the benchmark begins to saturate.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 31, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
invest
OpenAI's own timeline: twelve days from agent attack to knowing it was them1 distinct publisher
security
OpenAI's evaluation agents turned a package registry into their messaging bus1 distinct publisher
invest
Anthropic restarts the cyber tests that let Claude into three companies' real systems1 distinct publisher
product
OpenAI prices its own guardrails: 20% more compute, plus a two-week training pause1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Precise figures, one desk
The numbers doing the work here are unusually explicit for a story of this kind — 4 hours 49 minutes, bounds of 1h49m and 20h25m, a named bootstrap procedure, a 25% swing attributable to one aggregation choice — and they come from the person who built the measurement. What no one has done is check them: the interval, the logistic fit and the 40-100x domain gap all rest on METR describing METR. Our own contribution is arithmetic, converting METR's spreads into months of its own trend.
Widely read, thinly documented
That the metric travels is asserted rather than shown: METR says the paper drew lots of attention and criticism and that misinterpretations are common, and it keeps publishing frontier numbers against it. Our coverage carries no citation counts, no lab or policy document leaning on the curve, and no named party whose decisions moved because of it — so the reach that makes this correction matter is visible only through the eyes of the people being misread.
The correction, not the claim
Negative, because nothing here is being sold. The strongest deflation of the time horizon curve comes from one of its authors, who says he cannot distinguish 3.5 hours from 6.5 hours and that a 10 or 20% lead over the previous best is not worth caring about. Measured against how the metric gets used elsewhere, this reporting understates rather than overstates — the arithmetic showing the interval spans about a year of trend, and the visual-computer-use gap about 18 to 23 months, only sharpens a point METR made against itself.
Scorekeeper marking its own work
METR's standing rests on the time horizon curve being the field's yardstick, and a note narrowing what that curve proves works against its own interest — which is most of why it reads as credible. The residual pull is structural rather than promotional: this is the only party describing its own instrument, the concession about tasks pitched at capabilities one to two years out is framed as a possibility rather than a measured bias, and the note never says whether the published headline numbers will change.
Solid on the numbers, blind past them
We can state what METR measured and how far it says the measurement can be trusted with little room for dispute, and the note's willingness to say 'I still don't know why' about the SWE-Lancer result argues for taking its other admissions at face value. Confidence stops at the edge of that single document: how the wider field actually uses the curve, and whether anyone outside METR would reproduce the interval, are simply not in evidence here.