Two MMLU Scores, One Benchmark Name: Why the Numbers Don't Add Up
The Two MMLU Scores: What a Benchmark Name Does Not Fix

Two builds of the same model family report MMLU accuracies of 0.781 and 0.79 under the same benchmark label. The shared name fixes a dataset family, not a measurement procedure: the split, implementation, prompt format, grader, and runner's network access all remain open. Published evidence shows these choices can move scores by whole points, not thousandths. Comparability attaches to the reference the results are traceable to, not to the number itself.
Comparability is a property of the reference the results are traceable to, not of the number.