Skip to content

benchmark

nocot-bench

A suite of programmatically generated reasoning tasks, each parameterized by a difficulty setting, used to measure how much reasoning a model can complete without emitting chain-of-thought tokens.

Current clusters