Skip to content

benchmark

code800

Dataset of matched answerable and structurally impossible code prompts across 8 categories, used to test abstention on formally invalid input.

Current clusters