Why Your C++ Code Is Slow: The Hidden Cost of Cache Misses

Writing Efficient C++ Code

Why Your C++ Code Is Slow: The Hidden Cost of Cache Misses

C++ gives you direct access to hardware, but that alone doesn't guarantee speed. The real bottleneck is often memory: a single cache miss can cost hundreds of cycles. This article argues for Data-Oriented Design—laying out data contiguously and avoiding pointer-chasing—to write code that's both simple and fast. It also warns against over-abstraction and shows how cache-friendly structures like std::vector can outperform linked lists and trees.

If we try to express as directly as possible what data a program must store and what operations it must perform on that data, the code will be simple, elegant, readable, and efficient at the same time. This contradicts the popular view that optimization means making code complicated and unreadable.
  1. asveikau

    This article reminds me of performance advice I was starting to see in the 2000s decade. Basically it was to not introduce a bunch of pointer heavy data structures to get lower algorithmic complexity. Stuff it all into a vector. You will use some algorithms that the computer science textbook will say it's slower, but if it fits all in cache it doesn't matter. The cache misses following pointers all over town hurts you more.

  2. Jeaye

    While we're here, has anyone seen any resources related to data-oriented design when GCs are involved? So much of data-oriented design is arena-focused, but that's not always possible, when the lifetime model of the code requires a GC (for whatever reason).

    I feel like the DoD movement is a slow-moving, but big, change through how systems programming is done, but that there's still insufficient material for how to do this in different scenarios. I would really like to apply this more to my areas of work, which are also in C++, but there seems to be a gap between what they're presenting and how it can be applied.

    More specifically, I'm using C++ to build a dynamic programming language runtime for a Clojure dialect. That runtime is required to be garbage collected, type-erased, and highly polymorphic. So I surely can't just SoA or AoS everything. Yes, I can pack my data, and I can avoid the GC whenever possible, both in compiler/runtime code and in generated code via escape analysis. But what about everything else, which is the 80% or more of the system? It could be that this runtime is too far at odds with DoD, but I generally see things as a gradient rather than black and white.

  3. hn_submit

    I write in C++ almost every day but never have the need to optimize for speed. Even when you write straightforward code it's already blazingly fast.

  4. MaxBarraclough

    There's no mention of branch prediction, or context switching, or synchronisation. Depending on what you're doing, they could be very consequential. There's only very brief mention of parallelisation with threads and with SIMD.

    High-performance programming is a big topic. The scope is far too broad for a single blog post, which naturally gives only cursory discussion of C++ and computer architecture. The article isn't bad considering, but I do think it's the wrong format. A blog series, or even a book, would be more fitting.

  5. 112233

    "This article was originally published in Polish in issue 4/2013" — a lot of excellent advice. Sad to see C++ have moved in last decade in a direction that makes writing efficient, simple low level code harder and harder :(

More from this day

2026-09-27