On August 12, 2026 Emery Cooper, Caspar Oesterheld, Chi Nguyen and Alex Kastner of Redwood Research, with Joe Benton and Ethan Perez of Anthropic, introduced the Conceptual Reasoning Index. The motivation is that current training depends on abundant data and reliable feedback, so models are systematically weaker at tasks that cannot be empirically or mathematically verified. Much of the work of reducing risk from advanced AI falls in exactly that category: reasoning about systems more capable than any human, choosing research agendas whose payoff arrives years later, and arguing about which values AI systems should have. The authors call this capability conceptual reasoning and built three benchmarks to measure it.
LMCA, the Language Model Conceptual Argumentation dataset, contains 560 position texts with 1,461 arguments against them, drawn from decision theory, philosophy and risks from advanced AI. Nearly all arguments were rated against a detailed rubric by a conceptual researcher, with some rated independently by others for 2,140 ratings in total, including a validation set of roughly 50 arguments each rated by four to six people and then discussed for seven to eight hours. Models are scored on how closely their ratings match the human ones. ACCoRD measures whether a model’s stated beliefs and preferences are logically consistent, for example whether its reported probability for A is at least its reported probability for A and B; it holds close to 14,000 generated constraints across 18 types, of which 567 were human-approved and used in the index. DTBench capabilities is 407 handcrafted multiple-choice questions on decision-theoretic situations involving predictions of a model’s own behavior or interactions with near-copies.
The index is currently a weighted average of LMCA at 60 percent, ACCoRD at 20 percent and DTBench capabilities at 20 percent, with scores running from 0 for random guessing to 100. As of August 10, 2026 the highest scorer was Anthropic’s Opus 5 at 73.6, with a 95 percent confidence interval of plus or minus 2.1. Because human ratings are noisy, the authors estimate the practical ceiling of the index at around 91, so the frontier is well short of it. Scores have risen roughly linearly since late 2024 with no sign of flattening. The sub-benchmarks are diverging: DTBench capabilities is close to saturated, with Fable 5 answering 98 percent of questions correctly, LMCA is loosely projected to begin saturating about a year out, and the authors are very uncertain about ACCoRD.
Two methodological details are worth noting because they will show up in other evaluations. All models were run at maximum token limits and effort levels, and where Fable 5 refused to answer a question the authors substituted Opus 5 as a fallback. Refusals are becoming a measurement problem in their own right: the reported ACCoRD figure for GPT-4 is partial because that model refused to fully answer 18 percent of items.
The reason this matters beyond alignment research is that most consequential business judgment looks like conceptual reasoning, not like a graded exam. Strategy calls, regulatory bets and organizational design decisions have no verifiable answer key and are settled by argument. The CRI is the first public attempt to put a trend line on that ability, and it says two things at once: frontier models are meaningfully below the human-expert ceiling on unverifiable reasoning, and the gap has been closing at a steady rate for nearly two years. Live scores are published at conceptualreasoning.ai.