Build1 distinct publisher2 min readUpdated
OmnisBench's author rebuilt his split from date-stamped LiveCodeBench problems. The cheap tier fell from 90% to 60%, and a 4,096-token cap had been marking the frontier model absent.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
Eleven blank answers at 4,096 tokens apiece is 45,056 output tokens billed for nothing a grader could read [14]. That was the entire failure set for the frontier model on the hard split, according to the author: empty responses rather than wrong ones, each the model reasoning up to the cap and then getting cut off before it wrote a line of code [6].
The size of that harness error is worth stating plainly. The truncated run put the frontier model at 60%; with the output budget turned into a config setting and raised, the same model on the same problems came back at 86.7% [5][8][9]. That is 26.7 points of apparent incompetence that belonged to a config value [15]. A single global output ceiling is fine while the tasks are short functions [7], and it quietly converts a reasoning model into an absentee once they are not.
Now the number the post does not compute. The cheap model scores 60% on the fresh split, and routing is reported as buying 33 points there, which puts the routed policy near 93%, roughly six points above the frontier model it escalates to on a fifth of calls [10][11][16][9]. On the old tasks the same policy escalated never and bought ten points [11]. So the contaminated suite was not only flattering the cheap tier, it was grading the escalation logic as surplus to requirements, and escalation logic is the product.
Two things keep 93% from being a headline. The samples are small and the fresh split is competitive programming, so difficulty and recency move together, which the author says himself [12]. And his own figures for the cheap model on old tasks do not agree: 94.5% in the first post, 90% in the re-grade, a 4.5-point gap left unexplained [3][10][18]. Read the 30-point fresh-set drop [13] as a ceiling on memorisation rather than a measurement of it.
The useful part is the direction of the two errors. Contamination lifts the small model; a tight output cap pushes down the model that thinks longest before answering. A suite carrying both does not merely get the magnitudes wrong, it can invert the ordering, and cheap-is-close-enough is exactly the conclusion it produces. All of this was recoverable only because the raw responses were sitting on disk to be opened [1].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
The author concludes the contamination point stands: with the models given all the output room they wanted, the cheap model is thirty points worse on problems it cannot have memorised.
OmnisBench is an open benchmark for LLM routing, built around the author's OmnisRouter, and it publishes the actual model responses so that every number can be re-graded by a reader.
Two commenters, deanlee and jugeni, argued that HumanEval and GSM8K are old enough that the models have almost certainly read the answers, undercutting the benchmark's routing result.
The author built a fresh split from LiveCodeBench, which stamps every problem with a release date, keeping only problems published after the models could plausibly have trained on them, and graded the same routing policies on old and new sets side by side.
In the first fresh run the cheap model's score collapsed and the frontier model came in far below expectation; the truncated run reported the frontier model at 60%.
Every one of the frontier model's failures on the hard problems was an empty answer rather than a wrong one: eleven problems, eleven blanks, each exactly 4,096 tokens of reasoning before the response was cut off without any code.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single-source, self-graded, but with inspectable artifacts
All findings come from one vendor-authored post about its own benchmark and router. The evidentiary strengths are concrete and unusual for this genre: the failure mode was diagnosed from published raw responses (eleven blanks each at exactly the 4,096-token ceiling), the harness change is specific, the split-construction method is reproducible via LiveCodeBench release dates, and results are said to be re-gradable offline. The weaknesses are equally concrete: no independent replication, no disclosed sample sizes or model versions, an acknowledged difficulty/freshness confound, and an unreconciled 94.5%-versus-90% figure for the same cheap tier on the old set. That supports the harness/truncation finding well and the contamination magnitude only weakly.
No third-party adoption disclosed
The supplied material documents only the vendor's own benchmark runs and a harness configuration change. There is no evidence of external users, downstream deployments, independent re-grading, stars/forks, integrations, or any party other than the author and two commenters engaging with OmnisBench or OmnisRouter. Uptake cannot be measured without inferring facts the source does not provide.
Mildly overstated headline, unusually well hedged
The framing that routing is worth roughly three times more on uncontaminated problems rests on one small, self-graded run in which difficulty and freshness are admittedly confounded — a vendor-favourable conclusion drawn from evidence that cannot separate 'fresh' from 'hard'. That pushes the gap positive. It is held close to zero by strong self-correction: the post retracts its own earlier 94.5% narrative, publishes the bug that produced a wrong 60% figure, states plainly that this is 'a first signal, not a verdict', and discloses the cost of its own mistakes. The residual overstatement is the causal attribution to contamination, not the harness finding, which is if anything under-sold.
Vendor grading its own product, partially offset by disclosure
The benchmark is built by Fortitude Omnis explicitly around its own OmnisRouter, and the reported result — that routing earns far more once contamination is removed — is directly commercially favourable to the router being sold. The post also converts the episode into a sales argument against competitors ('if a routing benchmark, or a routing vendor, won't show you the responses'). Offsetting factors are meaningful but do not remove the conflict: publication of raw responses, an offline verify path, retraction of its own earlier flattering number, and disclosure of the cost of the failed runs.
Moderate on method, low on magnitudes
Confidence is asymmetric. The mechanism — a fixed 4,096-token output cap turning extended reasoning into empty, zero-scored answers, and the value of publishing raw responses to catch it — is specific, internally consistent, and independently plausible, so it can be relied on as a practice warning. The quantitative claims (30-point contamination penalty, 33-point routing gain, ~93% routed fresh score) rest on a single self-graded run with undisclosed sample sizes, a confounded design, no replication, and one unreconciled internal discrepancy, so their magnitudes should be treated as provisional.
security
Mandiant found 100 high-severity bugs in two days. Plan for the other side doing the same.1 distinct publisher
science
NIST's own logs show agents looking up the answers, making public benchmark scores soft evidence1 distinct publisher
product
A 27B laptop model scores like a rented one, and thinks three times as hard to do it1 distinct publisher
build
Inco AI's DFlash 2: 21% longer accepted drafts for 1.3% latency and 18.5M parameters1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 21, 2026