Skip to content

benchmark

EDIT-Bench

Benchmark that scores language models on realistic code-editing tasks using pass@1.

Current clusters