Recreating Minecraft Is Not a Benchmark

Recreating Minecraft Is Not a Benchmark

The author argues that viral demos like recreating Minecraft in one prompt have become 'demo-benchmarks'—easy to overfit and measure preparation rather than capability. They point to Thinking Machines' Inkling Small scoring within a point of its flagship on the Artificial Analysis Intelligence Index with a third of the parameters, showing that public static tests leak into training data. They suggest holdout evals like LiveBench or ARC-AGI as better alternatives, but acknowledge demos' appeal for their instant clarity.

A test you can perfect on a schedule measures preparation instead of capability, to me that’s anti the very definition of a benchmark, it should be a hard test, something very hard to perfect.

More from this day

2026-09-06