42x Faster Prompt Lookup Drafting in llama.cpp

42x Faster Prompt Lookup Drafting in llama.cpp

I optimized prompt lookup decoding in llama.cpp, achieving up to 42x faster drafting and 2.6x lower memory usage. By eliminating unnecessary map copies, replacing std::unordered_map with ankerl::unordered_dense::segmented_map, and using sorted vectors for inner maps, I reduced drafting latency from 165 µs to under 4 µs for a 541 MB corpus. The changes maintain identical acceptance rates and are detailed in my GitHub repository.

This one is almost more of a bug fix than an optimization. I found that the inner maps were being copied unnecessarily in multiple places on every drafting step.
  1. jadidbourbaki

    Btw, if anyone has experience with the open source community in general and llama.cpp in specific, I would greatly appreciate some advice. I’m facing a bit of an interpersonal issue that I really hope is resolved without any ill will. Here is the context:

    https://www.reddit.com/r/LocalLLaMA/comments/1wr5ylm/comment...

    Any advice for what I can do? Due to this, I cannot create a PR or issue in the llama.cpp repository. However, I am worried about bothering the maintainers on other channels in case it aggravates them further. Thank you for your help!

  2. jadidbourbaki

    Fun update to this: Daniel Lemire added another optimization to make this even faster. https://github.com/jadidbourbaki/llama.cpp/pull/12

    I’ll benchmark his change and add it to the article, crediting him for this improvement.

  3. PicardsFlute

    So, instead of resolving it privately like you were asked, you come here and complain? Was the bot posts on Reddit not enough for you? Pro tip: There is a right way and a wrong way to approach these things and you are MOST certainly approaching this the wrong way. And before you go accusing me of anything: I am NOT associated with that project, but I DID see what happened. You are acting like a damned child and should be ashamed of yourself.

More from this day

2026-09-27