Skip to content

Topic

LLM-as-Judge Evaluation

A technique that uses large language models to score or grade the outputs of other AI systems, often replacing or supplementing human evaluation.

Current stories

invest1 publisherOne report

Coinbase cuts its 90-case support bot test from up to two weeks to under an hour

Coinbase says its Autopilot system now runs about 90 support-bot test cases in 30 to 45 minutes, work that took one to two weeks by hand. A person still approves each procedure release, so the faster cycle gives Coinbase more checks per reviewer while its effect on customers and staffing costs stays unmeasured.

Publishers:cryptoslate.com

Reality

Evidence40
Adoption25
Hype gap+10
Incentives60
Confidence45
build1 publisherOne report

Ten human-labelled prompts calibrate the judges in this abliteration study

Abliteration needs no gradient updates, so a projection at inference time is a fair test of whether a data recipe actually diffused refusal behaviour. The judge you pick to score it changes the answer, and this protocol calibrates its judges on ten prompts.

Publishers:arxiv.org

Reality

Evidence55
Adoption
Insufficient
Hype gap+12
Incentives30
Confidence45