AI Benchmarks Plateau: Study Finds Nearly Half Saturated

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

AI Benchmarks Plateau: Study Finds Nearly Half Saturated

A systematic study of 60 language model benchmarks reveals that nearly half exhibit saturation, with rates increasing with age. The research identifies 14 properties related to saturation and finds that expert curation, not public test data, impacts resilience. These insights suggest design choices can extend benchmark longevity and inform more durable evaluation approaches.

We find that nearly half of our benchmarks exhibit saturation, with rates increasing with age.

More from this day

2026-08-04