Skip to content

benchmark

PinchBench

Benchmark suite that scores how well a language model performs as the reasoning core of an OpenClaw agent across a set of agentic tasks.

Current clusters