BuildNot yet confirmed elsewhere1 publisher3 min readPublished
Berkeley statisticians find METR's time-horizon axis nearly flat between 2 and 30 minutes
Berkeley statisticians Nguyen and Fithian refit METR's time-horizon data and found AI task difficulty barely changes between 2 and 30 minutes of human time. Equal 10x steps on the plot therefore mean unequal capability gains, depending on where on the axis they fall.
The Engineer · Build desk
What happened
- Each AI's 50% horizon is solved from a logistic fit with log2 of human task time as the only covariate, and the dated trend is drawn from those fitted values.
- Runs were far more correlated than independence predicts, with 83% of run sets unanimous against an expected 60% or so.
- Under 5-fold cross-validation split by task family, the authors report their horizon estimates scored better than the baselines.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Adoption or budget cases built on a horizon gain inside the 2-to-30-minute band need task-level evidence before anyone accepts the multiplier as stated.
- constraint A steady doubling rate on the dated plot does not show steady capability growth per year, because each dated point inherits the linear-in-log-time assumption.
- capability The audit needs no new benchmark runs, so any internal eval that logs repeated runs, human times and task families can be refit the same way to test its own axis.
Every point on the METR plot is solved from a fitted curve [4]. Each task in the suite carries a human time, the geometric mean of a skilled engineer's successful timed attempts [3]. Each AI attempts every task several times [3]. For each AI, the standard recipe fits logit p(t) = alpha - beta * log2(t) and solves for the t where p reaches one half [4]. A 2025 METR research note moved to one slope shared across AIs, and the Berkeley paper takes that version as its baseline [5]. The fit never sees a release date. The exponential trend shows up only later, when the fitted horizons are plotted by date [6].
One covariate means one strong assumption. AI difficulty has to rise linearly with the log of human time, so every 10x step in human minutes costs the same success probability [7]. According to a dev.to summary of the paper, nothing in the data collection guarantees that [7].
Nguyen and Fithian kept METR's data, 228 tasks in 79 families across 26 AIs, and changed the statistics [8]. The authors call the statistical fit a "model" and the language model an "AI" [16]. I follow them here. It spares everyone the phrase "each model's model."
The first fit replaces the linear term with a monotone spline so the curve can bend [9]. The second is an explanatory item-response model, the method behind standardized test scoring [10]. Each task gets a latent difficulty. The model predicts that difficulty from human time through the spline, then adds a task-family effect and a per-task residual [10]. That family effect replaces the square-root reweighting METR used to stop large families dominating [12].
The item-response fit also handles correlated runs. If runs were independent given the fitted probabilities, about 60% of an AI's run sets on a task would be all successes or all failures; in the data, 83% are [11]. The gap is 23 percentage points [17]. A correlation parameter downweights task-AI pairs with many redundant runs [11].
I think the validation design is the part worth copying. Folds are split by task family, so held-out tasks come from families the fit never saw [2]. Scoring uses proper rules, including an elementary score at q = 0.5 and 0.8 that grades the horizon estimate itself [2]. The authors report their estimates beat the baselines across that suite [2].
The result turns on the shape of the spline. The conversion from human minutes to AI difficulty is close to linear outside roughly 2 to 30 minutes and nearly flat inside that band [13]. A horizon that moves from 3 to 30 minutes stays inside the flat region the whole way [19]. The authors say that jump is much easier than 30 minutes to 5 hours, though both are 10x [14]. Their conditional success plot suggests a 4-to-15-minute move, a 3.75x multiplier, may be quite small in difficulty terms [15][18].
The paper does not dispute the trend [1]. For a horizon figure to transfer to a team's own backlog, those tasks need to map human minutes to AI difficulty the way METR's suite does. The refit shows that map bends at about 2 and 30 minutes [13]. The dev.to author offers one explanation, labelled as interpretation and not as a finding of the paper. Short tasks mostly test whether an AI follows instructions and calls tools without slipping, while longer ones add planning, error recovery and context management [20].
What to watch
- Whether METR adopts a spline or item-response fit in a future horizon update, as it adopted a shared slope in a 2025 research note.
- Whether the flat 2-to-30-minute region appears when the same fits run on task suites outside METR's 228 tasks.
- Whether refit horizons for the 26 AIs change the doubling rate shown on the dated METR plot.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap0
- Incentives
- Insufficient
- Confidence40
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Berkeley statisticians Nguyen and Fithian published 'On the estimation and validity of AI time horizons' in October 2026; the paper does not dispute the exponential trend in METR's time-horizon plot.
ReportedSupportedSource: dev.to summary of the paper2 sources— create a free account to open themView cited source - [2]
The authors use 5-fold cross-validation split by task family and score with proper scoring rules (log score, Brier score, and an elementary score at q = 0.5 and 0.8 that directly judges the horizon estimate) crossed with three weighting schemes; they report their time-horizon estimates perform better than the baselines.
ReportedSupportedSource: Nguyen and Fithian, per dev.to summary2 sources— create a free account to open themView cited source - [3]
Each task gets a human time, how long a skilled engineer takes, using the geometric mean of successful timed attempts. Each model attempts the tasks several times. The 50% time horizon is the human time at which the model's success probability crosses one half.
- [4]
The 50% horizon is not measured directly; it is read off a fitted per-model logistic regression with log2(human time) as the only covariate, logit p_j(t) = alpha_j - beta_j * log2(t), inverted at p = 0.5.
- [5]
A 2025 METR research note moved to a shared slope across models; the new paper treats this shared-slope version as its baseline.
- [6]
Release dates are not in the fit; the exponential trend comes from plotting the fitted horizons against dates afterward.
- [7]
A logistic regression on log time assumes task difficulty for an AI grows linearly with the logarithm of human time, so every 10x step in human time costs the same amount of success probability. Nothing in the data collection guarantees it.
- [8]
The paper keeps the same data, 228 tasks in 79 task families and 26 models, and fits two richer statistical models.
- [9]
Model 1 replaces the linear term with a monotone spline so the link between human time and difficulty can bend.
- [10]
Model 2 is an explanatory item-response theory model, the machinery behind standardized tests. Each task gets a latent difficulty regressed on human time through the spline, plus a task-family effect and a per-task residual.
- [11]
If runs were independent given the fitted probabilities, about 60% of a model's run sets on a task would be unanimous; in the data, 83% are. Model 2 includes a correlation parameter that downweights task-model pairs with many redundant runs.
- [12]
The family effect replaces the heuristic square-root reweighting METR used so large families did not dominate.
- [13]
The fitted conversion function from human time to AI difficulty is nearly flat from roughly 2 to 30 minutes of human time and close to linear outside that range.
- [14]
The authors' example: a horizon jump from 3 minutes to 30 minutes is much easier than one from 30 minutes to 5 hours, despite both being 10x.
- [15]
The paper's conditional success trajectory plot suggests a move from 4 to 15 minutes may be quite small in difficulty terms.
- [16]
The authors use 'model' for the statistical fit and 'AI' for the language models.
- [17]
Observed unanimous run sets exceed the independence expectation by 23 percentage points.
- [18]
A move from 4 to 15 minutes is a 3.75x multiplier in human time.
- [19]
The 3-to-30-minute horizon jump lies entirely inside the roughly 2-to-30-minute band where the conversion function is nearly flat.
- [20]
Offered as interpretation rather than a finding from the paper: short tasks mostly test whether a model can follow instructions and call tools without slipping, while longer tasks add planning, recovery from errors, and context management.
Sources
1 independent publisher whose own reporting we read for this story.
- dev.toIs a 10x Jump Always a 10x Jump? A Statistical Audit of the METR Time-Horizon Plot
1 article · October 9, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- Simulation Evaluation and Accuracy ClaimsFollow
- AI forecastingFollow
- Item response theoryFollow
Entities
- METRFollow
- METR time horizonFollow
- University of California, BerkeleyFollow