Skip to content

benchmark

KernelBench

A benchmark testing language models on writing efficient GPU kernels.

Current clusters

build1 publisher

Asking GPT-5.6 Luna to name an amphibian flags benchmark transcripts with black-box access

GPT-5.6 Luna says "frog" 70-95% of the time when asked for an amphibian after capability benchmarks, against 12-38% after real use, a LessWrong post reports. Anyone with black-box access can run the check, though its authors cannot yet say whether it detects evaluation awareness or lexical cues.

Publishers:lesswrong.com

Reality

Evidence45
Adoption
Insufficient
Hype gap+10
Incentives
Insufficient
Confidence40