ARC-AGI-3 Leaderboard: Measuring True AI Efficiency and Adaptability

ARC-AGI Leaderboard

ARC-AGI-3 Leaderboard: Measuring True AI Efficiency and Adaptability

The ARC-AGI-3 leaderboard tracks how AI agents adapt to novel interactive environments, moving beyond passive intelligence. I highlight the critical balance between performance and cost, comparing top models like Claude Opus 5, Grok 4.5, and GPT-5.6 Sol. This data reveals that true intelligence requires solving problems efficiently with minimal resources, distinguishing advanced reasoning systems from standard base LLMs.

True intelligence isn't just about solving problems, but solving them efficiently with minimal resources.
  1. KaoruAoiShiho

    Appears to be benchmaxxing

    https://x.com/quietnning/status/2080786711861407883

  2. dinp

    The last time I checked, for the arc agi 3 leaderboard, the models are given a simple prompt and the game input and asked to play the game, no harness/tools. If harnesses were allowed, I would expect the benchmark to be saturated. There were a few harness attempts, but they could only be evaluated on the public set, so it's not an apples to apples comparison.

    My guess is, the large score jump for Opus 5 is mainly because of getting the right RL envs for training.

    It's becoming harder and more expensive to build and run meaningful benchmarks, it would be interesting to see what they do with arc agi 4, maybe just give it gameboy/steam games and see how they compare vs a human baseline? The latency requirements and very long horizons in games could be an interesting challenge for llms.

  3. throwaw12

    Why Anthropic models are always leapfrogging these benchmarks, but in real life work I do feel like after 3 weeks I am back to Claude Opus 4.5? (regardless of the model I use, Fable was exception for 1 day when it was released)

  4. MaskNinja

    Not possible. I don't get how Opus 5 gets so high. Have they run it against the private and held-out games?

  5. codedokode

    Why is there no Kimi 3, and GLM5.2 didn't run the third benchmark? I am more interested in knowing the abilities of open weight models.

  6. albatross79

    ARC-AGI is a beauty contest for pigs where the pig's owners compete to see who can apply the lipstick most convincingly.

  7. NooneAtAll3

    games are great (as for a human)

    but I kinda wish I could select level... I accidentally pressed redirect button and when I came back I was once again shown level 1, all progress lost :(

  8. tudelo

    > Only systems which required less than $10,000 to run are shown. (Notes[1])

    Am I lost or are their many models on this ranking (Opus 5 included) that clear this?

More from this day

2026-07-25