Skip to content

Topic

Code generation benchmarks

Datasets and scoring harnesses used to measure how well language models write working code, usually by executing tests against generated functions.

Current clusters