Anthropic launches Conceptual Reasoning Index to measure AI's philosophical thinking
Anthropic: Introducing The Conceptual Reasoning Index

Anthropic and Redwood Research introduce the Conceptual Reasoning Index (CRI), a suite of three benchmarks—LMCA, ACCoRD, and DTBench—to evaluate AI's ability to reason about conceptual questions where empirical feedback is limited. Initial results show top models score 73.6 out of an estimated ceiling of 91, with progress roughly linear and no signs of flattening.
A core hope for managing AI risks is that AIs will help us understand our situation, plan for what lies ahead, and develop risk mitigations.