Skip to content

model

GPT 5.3

A version of the GPT series of large language models.

Current clusters

build1 publisher

Agent workflows often made models worse on CMU's Synthetic Hospital chart benchmark

Carnegie Mellon's Synthetic Hospital benchmark found agent wrappers often lowered scores for 10 models reading 1,268 synthetic multi-visit patient charts. The models found facts in the charts but struggled to combine them, so a scaffold that adds retrieval steps may be working on the wrong weakness.

Publishers:theneuron.ai

Reality

Evidence50
Adoption
Insufficient
Hype gap+8
Incentives
Insufficient
Confidence45