Skip to content

Topic

LLM Evaluation

The practice of testing and benchmarking language models to measure accuracy, reasoning, and performance drift over time.

Current stories

build1 publisher

Crutches built from measured failures lift a local Qwen 3B from 33% to 52% on post-cutoff facts

Qwen2.5-3B, wired to a local Wikipedia index, scored 52% on 150 post-cutoff questions it answers none of unaided, up from 33%, in a dev.to author's tests. Each fix targets a measured 3B failure, so a zero-shot 7B gained only 9 points from them, and the two readers' confidence intervals overlap.

Publishers:dev.to

Reality

Evidence35
Adoption
Insufficient
Hype gap+25
Incentives
Insufficient
Confidence40
build1 publisher

Claude Code's plugin eval caught skills firing 55.6% of the time once 89 were loaded

Claude Code's new plugin eval showed one developer's skills firing in 5 of 9 relevant runs once all 89 were loaded, down from every run with one skill. The test covers three prompts in one project, but it gives teams that keep rules in skills a way to measure how often those rules get consulted.

Publishers:dev.to

Reality

Evidence35
Adoption
Insufficient
Hype gap+20
Incentives
Insufficient
Confidence30
build1 publisher

Simon Willison credits two November model releases with making coding agents reliable for daily use

Simon Willison says Claude Opus 4.5 and GPT-5.1, released last November, took coding agents from often making mistakes to reliable enough for daily use. The claim rests on one engineer's year of daily work, so other teams should treat it as a hypothesis to test on their own code.

Reality

Evidence30
Adoption
Insufficient
Hype gap+20
Incentives
Insufficient
Confidence35
build1 publisher

Per-property pass bars expose failures that a 92 percent eval average hides

One developer's support-agent eval suite fails two of its five property bars even though its pooled score is 0.919. The author argues checks like these are what OpenAI lacked when a sycophantic GPT-4o update shipped in April 2025 and was pulled in four days.

Publishers:dev.to

Reality

Evidence40
Adoption
Insufficient
Hype gap+5
Incentives
Insufficient
Confidence50

Earlier coverage

  1. A 4-bit Gemma 4 26B on one L4 trails TypeSafe's Jev by 2.1 points overall

    Build · September 23, 2026 · 1 publisher

  2. A small labelled holdout turns a miscalibrated chat LLM into a working abstention gate

    Build · September 23, 2026 · 1 publisher

  3. Fine-tuning an 8B model on under 250 security examples cost it severity scoring and scope control

    Build · September 22, 2026 · 1 publisher

  4. 4.6 KB of instruction cut Claude Code's answers from 524 words to 258

    Build · September 22, 2026 · 1 publisher

  5. Foundry's automated grading needs the trace to carry the content it grades

    Build · September 21, 2026 · 1 publisher

  6. Adding both the guard and sandbox moved measured attack success from 16% to 20%

    Build · September 21, 2026 · 1 publisher

  7. A six-label scorer keeps a 429 out of the model's error column

    Build · September 21, 2026 · 1 publisher

  8. A fifteen-line abstention rule removed more correct answers than confident wrong ones

    Build · September 21, 2026 · 1 publisher

  9. A candidate model's 1.6-point accuracy gain lands inside both confidence intervals

    Build · September 20, 2026 · 1 publisher

  10. A popularity fallback put off-genre recommendations on 1,724 of 16,311 game pages

    Build · September 20, 2026 · 1 publisher

  11. Bespoke Nimble alone closed 24 of the 27-point gap to Jev within two days of launch

    Invest · September 20, 2026 · 1 publisher

  12. Stamping a trap chunk id into every negative test turned two false passes red

    Build · September 19, 2026 · 1 publisher

  13. A three-entry allowlist test fails the build when this refund agent's rules import a framework

    Build · September 19, 2026 · 1 publisher

  14. Proving an LLM feature still works costs engineer-months after the ten-minute build

    Product · September 19, 2026 · 1 publisher

  15. Raw logs caught an LLM review debate replaying pre-generated text

    Build · September 19, 2026 · 1 publisher

  16. oh-my-agent admits a promoted fixture only when the failing run's output still fails it

    Build · September 18, 2026 · 1 publisher

  17. Databricks calls model switching the biggest lever on its AI coding bill

    Product · September 18, 2026 · 1 publisher

  18. Raindrop replays real production traffic against every pull request to test agent changes

    Product · September 18, 2026 · 1 publisher

  19. Monitoring and rollback account for £45,000 of a £75,000 agent build

    Build · September 17, 2026 · 1 publisher

  20. Catching a model-swap regression takes twenty labeled documents run twenty times each

    Build · September 17, 2026 · 1 publisher

  21. Every unanswerable question cleared the 0.35 refusal threshold by at least 0.09

    Build · September 15, 2026 · 1 publisher

  22. Fine-tuning requests usually mean one of two things: missing knowledge or wrong style

    Build · September 15, 2026 · 1 publisher

  23. Changing only the name on a CV moved its rank across three million model comparisons

    Build · September 14, 2026 · 1 publisher

  24. Claude Code's plugin eval spends six agent runs per case to measure a plugin's lift

    Build · September 14, 2026 · 1 publisher

  25. Open-weight moderation models land inside the proprietary error band on Bluesky posts

    Build · September 13, 2026 · 1 publisher

  26. A stale "unpublished" flag outlived four hours, four documents and two reviewers

    Build · August 17, 2026 · 1 publisher

  27. Orchestra builds its cheaper-model switch gate from the buyer's own corrected traces

    Build · September 13, 2026 · 1 publisher

  28. A support agent's faithfulness check passed on documents from 2024 and 2023

    Build · September 13, 2026 · 1 publisher

  29. Slicing FrontierMath by category leaves about 13 problems behind each error bar

    Build · September 11, 2026 · 1 publisher

  30. Mistral Small 3.2 cuts the same text into 547 Polish tokens and 377 English ones

    Build · September 10, 2026 · 1 publisher

  31. ADR-0001 demotes Langfuse to a projection of Kept's in-process trace

    Build · September 10, 2026 · 1 publisher

  32. Faros telemetry shows pull requests nearly quadrupling across 22,000 developers in two years

    Product · September 9, 2026 · 1 publisher

  33. Deleting an example beat banning it across three rebuilds of a 680-line prompt

    Build · September 7, 2026 · 1 publisher

  34. Integer-cents Python clears the invoice before the local 3B model's draft reaches anyone

    Build · September 6, 2026 · 1 publisher

  35. OpenAI's hosted Evals platform stops accepting writes on October 31, 2026

    Build · September 3, 2026 · 1 publisher

  36. LLM evals are a Cartesian sweep, and the sweep layer was solved in 2015

    Build · August 26, 2026 · 1 publisher

  37. Benchmarks are prototyping tools: how GitHub secret scanning ranked LLM configurations

    Build · August 25, 2026 · 1 publisher

  38. The cheapest LLM eval starts with one question: which mistake cannot be undone?

    Build · August 24, 2026 · 1 publisher

  39. The smallest serious coding agent has a 443-line edit tool, and that is the point

    Build · August 23, 2026 · 1 publisher

  40. Bedrock's evaluation modes grade what they can see, and the dataset outlives both

    Build · August 22, 2026 · 1 publisher