Skip to content

Topic

Medical AI benchmarks

Evaluation datasets and tasks that measure how well AI models perform clinical reasoning, diagnosis and medical record work.

Current clusters

build1 publisher

Agent workflows often made models worse on CMU's Synthetic Hospital chart benchmark

Carnegie Mellon's Synthetic Hospital benchmark found agent wrappers often lowered scores for 10 models reading 1,268 synthetic multi-visit patient charts. The models found facts in the charts but struggled to combine them, so a scaffold that adds retrieval steps may be working on the wrong weakness.

Publishers:theneuron.ai

Reality

Evidence50
Adoption
Insufficient
Hype gap+8
Incentives
Insufficient
Confidence45