Published Build3 min read
A benchmark for the questions nobody can check: inside the Conceptual Reasoning Index
Researchers working with Anthropic built three evals for arguments that have no verifiable answer. The ground truth is human raters, which is both the point and the limit.
Written for builders.See today for builders

What happened
- A LessWrong post titled "Introducing the Conceptual Reasoning Index" announces a suite of three conceptual reasoning benchmarks developed to evaluate reasoning on tasks lacking practical empirical feedback loops.
- The authors state the work was done in collaboration with Anthropic.
- The three benchmarks are LMCA, ACCoRD, and DTBench capabilities.
- The benchmarks are aggregated into the Conceptual Reasoning Index (CRI), available at conceptualreasoning.ai along with methodology details, and the authors say the website will be kept up to date as new models and benchmarks are released.
- Access to the primary conceptual dataset, LMCA, can be requested through a form.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A group of researchers has published the Conceptual Reasoning Index, an aggregate of three new benchmarks meant to score models on argumentation about questions that have no empirically verifiable answer [1][3][4], in work the authors say was done in collaboration with Anthropic [2]. It matters because the authors are targeting the one capability class that current training methods are structurally bad at: they note that AI training leans heavily on abundant data and reliable feedback, so models tend to be worse at tasks that cannot be empirically or mathematically verified [7].
The argument for building this is a list of tasks with broken feedback loops. Much safety work involves reasoning about AIs more generally capable than any human, for which the authors say there is no obvious reference class and no clear way to model the subject [8]. Some decisions have to be right the first time, since a mistake that leads to AGI takeover or an AI-assisted coup might not surface until too late, and choices such as which research agendas to prioritise or which governance interventions to pursue play out over timescales where evidence arrives after it is useful [9]. Some questions, such as which values AIs should have, may have no ground truth at all, though the authors hold that argument can still make progress on them [10]. They define this capability as conceptual reasoning: reasoning where evidence is limited, no verifiable answer exists, and you rely on argumentation [6].
The interesting design choice is in LMCA, or Language Model Conceptual Argumentation, which sidesteps the problem of verifying bottom-line answers by grading arguments instead of conclusions [12]. It holds 560 position texts and 1,461 arguments against them [13], with 2,140 ratings in total, nearly all produced by conceptual researcher Emery Cooper and some independently by at least one other researcher [14]. Models are scored by comparing their ratings to the researchers' [15]. That works out to roughly 2.6 arguments per position text [1] and about 1.46 ratings per argument [2], which means most arguments in the set carry one rater's application of the rubric. The ground truth here is a small expert pool, not a fact.
The authors know this is the load-bearing joint. On arguments rated by at least two people, they report that inter-rater agreement is high compared with agreement between humans and models [16]. Their validation set is roughly 50 arguments, each rated independently by four to six people and then discussed for a total of seven to eight hours [17], which is about eight to ten minutes of discussion per argument [3]. If humans disagreed with each other as much as models disagree with humans, the index would be measuring noise; that comparison is the claim to interrogate first.
LMCA also runs a generation loop: one model writes a fresh argument against a position text, and a second model, given the rubric and few-shot prompted with the existing rated arguments, scores it, which the authors say yields fairly accurate ratings [18]. That is a model grading a model against a rubric distilled from a handful of human judgements, and the calibration claim is the authors' own.
Three things to watch. The index and methodology live at conceptualreasoning.ai, which the authors say will be updated as new models and benchmarks appear [4], so the spread between models is the number to track against the human-human agreement floor. LMCA is access-by-request through a form [5], which constrains outside replication as much as it constrains contamination. And the other two components, ACCoRD and DTBench capabilities [3], are named but not yet documented at the level LMCA is; the authors have promised later posts making the case for the work [19].
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
A LessWrong post titled "Introducing the Conceptual Reasoning Index" announces a suite of three conceptual reasoning benchmarks developed to evaluate reasoning on tasks lacking practical empirical feedback loops.
ReportedView cited source - [2]
The authors state the work was done in collaboration with Anthropic.
- [4]
The benchmarks are aggregated into the Conceptual Reasoning Index (CRI), available at conceptualreasoning.ai along with methodology details, and the authors say the website will be kept up to date as new models and benchmarks are released.
ReportedView cited source - [5]
Access to the primary conceptual dataset, LMCA, can be requested through a form.
ReportedView cited source - [6]
The authors define conceptual reasoning as reasoning about questions where empirical evidence is limited, there is no practically verifiable answer, and one must rely heavily on argumentation.
ReportedView cited source
Sources & coverage · 1 publisher
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- lesswrong.comChi NguyenAug 12Introducing the Conceptual Reasoning Index
Additional citations
- Introducing the Conceptual Reasoning Index, LessWrong
