Anthropic's Claude Fable 5.1 Tops New AI Index with Tougher, Anti-Gaming Tests

Artificial Analysis Intelligence Index v4.2

Anthropic's Claude Fable 5.1 Tops New AI Index with Tougher, Anti-Gaming Tests

Artificial Analysis has released Intelligence Index v4.2, an interim update that adds more complex and realistic benchmarks while increasing the weight of private, held-out test sets to prevent gaming. New evaluations include AA-Briefcase, an agentic knowledge-work test with a private set, and GDP.pdf, a long-context document reasoning benchmark spanning 4,592 pages. The update also retires the saturated GPQA Diamond. Results show Anthropic's Claude Fable 5.1 leading, followed by OpenAI's GPT-6 Astra, with Meta, SpaceXAI, Moonshot/Kimi, Z.AI, and Google trailing. GPT-6 Astra dominates the output token frontier.

We have deliberately held back updates to keep the Index stable through recent major model launches. However, with the frontier moving so quickly in the past weeks, we feel it is important to deliver an immediate interim update to ensure our Index remains as relevant and useful as ever to users.
  1. jjcm

    IMO one of the biggest losses of the OpenAI/Cursor breakup will be the loss of OAI models on CursorBench [1]. Their bench has always been one that most-fit my mental model of how good each of these models are. I find AA’s Intelligence index to often be out of alignment with my own subjective evals.

    [1] https://cursor.com/evals

  2. throwaway13337

    I have no idea how artificial analysis got to be something anyone took seriously. This is their new benchmark set?

    A glance at their new index shows that whatever they're measuring, it isn't useful.

    Spend an hour with gemini 3.8 and tell me that model belongs in 2026. It feels like the model has Alzheimer's. It gets confused about whether what it reads is what it did. Just crazy bad.

    I haven't tried muse spark 1.3. But it must have been a miracle since 1.2 to hit that rank.

    Video game journalism vibes all over this.

  3. redox99

    They realized Astra having the same score as Sol was silly so they rushed to update the index so it fits what people expect.

    The old index was clearly bad (Astra is way better than Sol) but it's also unscientific to tweak it like this.

  4. jascha_eng

    Imo the omniscience index they have has the highest correlation to actual usefulness of the models.

    https://artificialanalysis.ai/evaluations/omniscience

    > measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer.

    This is so useful because it makes you actually trust a models output. A high score on benchmarks is not as useful because a model overtrained to always answer will give confidently wrong responses. But this index measures how often it is correct while penalizing wrong responses so that a high score means you can trust this model more and when it doesn't know it is more likely to tell you that it really doesn't know rather than making shit up.

    Fable also performs a lot better than opus 5 here which correlates very strongly with perceived strength despite the models performing similarly on e.g. DeepSWE

    Astra is a big jump from sol and performs the same or slightly better than fable here.

  5. sanxiyn

    It is very unfortunate they upweighted SciCode from 8% to 10%. SciCode is a broken benchmark: see https://arxiv.org/abs/2608.04975.

More from this day

2026-09-05