Expert re-grading shows frontier models nearly ace physics benchmarks
How good are frontier models at physics?
Low scores on physics benchmarks suggest frontier language models struggle with advanced physics, but experts who audited six widely used benchmarks found most errors were not the models' fault. After correcting flawed reference solutions and ambiguous questions, GPT-5.6-Sol's mean@4 on HLE-Physics jumped from 47.3% to 78.7%, and its corrected pass@4 reached 94.4% on retained CritPt challenges. The findings indicate current benchmarks substantially understate model ability and are nearing saturation.
Most audited cases initially evaluated as incorrect reflect these benchmarking issues rather than errors in the models' physics reasoning.