Skip to content

Topic

Benchmark Methodology and Reproducibility

Sensitivity of benchmark scores to prompts, temperature and agent scaffolding, and the methodological weaknesses of published benchmarks.

Current stories

build1 publisher

Cache lookahead brings a kaniko fork within 1.6x of BuildKit on GitLab runners

Kaniko's community fork, with cross-stage cache lookahead, ran 1.6x slower than BuildKit on GitLab.com runners in a dev.to benchmark. The 15x gap still quoted against kaniko dates from 2018, so teams choosing an in-cluster builder need figures from their own Dockerfiles.

Publishers:dev.to

Reality

Evidence40
Adoption
Insufficient
Hype gap+30
Incentives
Insufficient
Confidence35
build1 publisher

Confidential inference on Blackwell retains 96-98 percent of throughput with CC-aware adaptations, NVIDIA reports

NVIDIA measured TensorRT LLM holding 96.1 to 98.2 percent of its non-confidential output throughput on Blackwell, and it got there by unpinning host memory on the affected paths, moving decode readback off the scheduler thread, and timing kernel tactics with the GPU's global timer instead of CUDA events.

Reality

Evidence58
Adoption22
Hype gap+10
Incentives80
Confidence55
product1 publisher

Fireworks' own DeepSWE numbers put four coding models inside the noise band

The vendor selling the cheapest model in the comparison reports a 0.7-point quality spread across four frontier models against run-to-run variation of 1.4 to 3.2 points. That leaves price per task, $0.43 against an implied $6.45 for GPT-6 Astra.

Publishers:fireworks.ai

Reality

Evidence42
Adoption18
Hype gap+28
Incentives88
Confidence58

Earlier coverage

  1. Postgres walks all 110,659 pages of a 1.2 GB table to answer OFFSET 9999980 LIMIT 20

    Build · September 15, 2026 · 1 publisher

  2. Trail of Bits says 1Password's 26% AI patch score reflects flawed prompts, no-code-execution trials, and grading errors, not true AI performance

    Security · September 15, 2026 · 1 publisher

  3. Rerunning the same eval suite three times in ten minutes moved its score by two cases

    Build · September 14, 2026 · 1 publisher

  4. AllSpark ran each Iris benchmark twice to separate the model from its scaffolding

    Build · September 13, 2026 · 1 publisher

  5. Dropping a 6.6-second trace phase saved more build time than Turbopack's faster compile

    Build · September 11, 2026 · 1 publisher

  6. One predicate removed 48.5% of a test operator's steady-state reconciles without weakening repair

    Build · September 7, 2026 · 1 publisher

  7. A token-overlap matcher shrugged at more than half of 1,538 agent rule candidates

    Build · September 7, 2026 · 1 publisher

  8. Speed and archive size rank the same nine backup plugins in opposing orders

    Build · September 7, 2026 · 1 publisher

  9. Measuring each patch against the human fix on the same bug leaves 13 of 14 models messier

    Build · September 6, 2026 · 1 publisher

  10. A .NET 10 lab benchmarks four publish modes against ten metrics on one laptop

    Build · September 4, 2026 · 1 publisher

  11. Twenty passing runs move the coding-API decision onto time-to-first-token

    Build · September 2, 2026 · 1 publisher

  12. Reordering two loop indices beat every cache-tiled version on the same machine

    Build · September 2, 2026 · 1 publisher

  13. Meta prices streaming transcription at a fifth of Google Cloud's standard rate

    Leadership · September 1, 2026 · 1 publisher

  14. PgCache invalidates cached aggregates off Postgres's logical replication stream

    Build · September 1, 2026 · 1 publisher

  15. fal reports 35x MiniMax's own H3 endpoint after tuning the weights to its runtime

    Build · September 1, 2026 · 1 publisher

  16. The choice of speaker encoder moved equal error rate five-fold on an identical trial list

    Build · August 31, 2026 · 1 publisher

  17. Farm.js compiles a state update into a direct DOM write when it can prove the target

    Build · August 31, 2026 · 1 publisher

  18. LatticeDB's traversal speed margin over SQLite narrows by three hops, then widens sharply by depth ten

    Build · August 30, 2026 · 1 publisher

  19. The repo's own control run deleted the 5-10x WASM claim from vizcrush's launch copy

    Build · August 29, 2026 · 1 publisher

  20. Eleven agent sessions on one machine settled CPU contention by writing to each other

    Build · August 28, 2026 · 1 publisher

  21. WebKit charges 39 MB for the tab Kestrel's simulator priced at 32 KB

    Build · August 28, 2026 · 1 publisher

  22. Harness choice moved token use 83-fold with the model held constant

    Build · August 27, 2026 · 1 publisher

  23. The one Pylint check Ruff cannot take, and the case for pulling it out of CI

    Build · August 24, 2026 · 1 publisher

  24. Optuna killed 60% of the trials and gave back 28% of the time. The gap is in the config.

    Build · August 24, 2026 · 1 publisher

  25. Harness-Bench makes the scaffold a measured variable, and model-only scores lose meaning

    Build · August 23, 2026 · 1 publisher

  26. A 284B model at 25 tokens a second on one 5090, and 192 GiB of DDR5 doing the quiet part

    Build · August 22, 2026 · 1 publisher

  27. Uniform INT4 beat NVFP4 in 85 of 90 real gradient tests, and the rotation barely moved them

    Build · August 22, 2026 · 1 publisher

  28. AgentCL: if the task stream is not controlled, agent memory gains prove nothing

    Build · August 21, 2026 · 1 publisher

  29. A five-way graph database benchmark spent its first four hours measuring undersea cable

    Build · August 21, 2026 · 1 publisher

  30. Tokenize-then-compress works, but the win is window coverage, not the 45% smaller stream

    Build · August 20, 2026 · 1 publisher

  31. The 21-cent model bake-off that inverted when the judge got audited

    Build · August 20, 2026 · 1 publisher

  32. A JSON parser benchmark that scores refusal as a pass, and why the column order flips

    Build · August 19, 2026 · 1 publisher

  33. Next.js 16.3's memory claim didn't reproduce; its TypeScript handoff cut a build by two thirds

    Build · August 18, 2026 · 1 publisher

  34. Your GPU reports 24GB. Only 7.9GB of it loads a model, and half of that is already gone

    Build · August 18, 2026 · 1 publisher

  35. A harness gain is not a leaderboard win: reading the J-Space DeepSeek report properly

    Build · August 17, 2026 · 1 publisher

  36. Artificial Analysis moves eval onto your data, and turns model choice into procurement

    Build · August 15, 2026 · 1 publisher