Skip to content

Topic

LLM-as-Judge Pipelines

Use of language models as nondeterministic reviewers within otherwise mechanical validation pipelines.

Current stories

invest1 publisher

MIT and Sakana AI's SIFT finishes a coding agent's self-improvement search on about $34 of API calls

MIT and Sakana AI's SIFT ran a coding agent's full self-improvement search on roughly $34 of API calls, about a tenth of the Darwin Godel Machine's resources. The saving comes from a language-model judge screening patches, so it holds only while that judge picks correctly.

Reality

Evidence40
Adoption
Insufficient
Hype gap+25
Incentives
Insufficient
Confidence45
build1 publisher

A noisy judge drags an adaptive agent stop rule below a fixed six-step budget

CDV's adaptive stop rule truly reached its quality bar in 71% of runs under judge noise of 0.10, against 93% for a fixed six-step budget. Requiring two passing scores in a row lifts that to 97% for 1.4 extra steps, so a one-line change beats both the Bayesian policy and the budget.

Publishers:dev.to

Reality

Evidence58
Adoption
Insufficient
Hype gap+8
Incentives30
Confidence55

Earlier coverage

  1. AWS's turn-level metric separates the one broken turn from the three that inherited it

    Build · September 10, 2026 · 1 publisher

  2. Anthropic broke an agent ceiling by making "is this design good?" a gradable question

    Leadership · September 10, 2026 · 1 publisher

  3. Judge model choice swings AI-Infra-Guard's false positive rate fifteenfold

    Security · September 9, 2026 · 1 publisher

  4. Six calls, eleven evaluators, and the pass condition that still lets dead air through

    Build · August 26, 2026 · 1 publisher

  5. Three agents, one spec, and a blind judging round that mostly proved they vote for themselves

    Build · August 22, 2026 · 1 publisher

  6. Palomar registers Lean proofs against a commit, and refuses to referee them

    Build · August 18, 2026 · 1 publisher