Skip to content

Topic

AI-generated code benchmarks

Evaluations that score model- or agent-written programs on measured performance, such as runtime against established library implementations, rather than on test pass rates alone.

Current clusters