Why a Loop That Does 6% of the Work Takes the Same Time
Gallery of Processor Cache Effects
Igor Ostrovsky uses C# code samples to reveal how processor caches shape real-world performance. A loop touching every 16th array element runs as fast as one touching every element because both fetch the same 64-byte cache lines. He also shows how L1/L2 cache sizes create performance cliffs, how instruction-level parallelism can double loop speed, and how cache associativity causes mysterious slowdowns for certain step sizes.
The second loop only does about 6% of the work of the first loop, but on modern machines, the two for-loops take about the same time: 80 and 78ms respectively on my machine.