42x Faster Prompt Lookup Drafting in llama.cpp

I optimized prompt lookup decoding in llama.cpp, achieving up to 42x faster drafting and 2.6x lower memory usage. By eliminating unnecessary map copies, replacing std::unordered_map with ankerl::unordered_dense::segmented_map, and using sorted vectors for inner maps, I reduced drafting latency from 165 µs to under 4 µs for a 541 MB corpus. The changes maintain identical acceptance rates and are detailed in my GitHub repository.
This one is almost more of a bug fix than an optimization. I found that the inner maps were being copied unnecessarily in multiple places on every drafting step.
- jadidbourbaki
Btw, if anyone has experience with the open source community in general and llama.cpp in specific, I would greatly appreciate some advice. I’m facing a bit of an interpersonal issue that I really hope is resolved without any ill will. Here is the context:
https://www.reddit.com/r/LocalLLaMA/comments/1wr5ylm/comment...
Any advice for what I can do? Due to this, I cannot create a PR or issue in the llama.cpp repository. However, I am worried about bothering the maintainers on other channels in case it aggravates them further. Thank you for your help!
- jadidbourbaki
Fun update to this: Daniel Lemire added another optimization to make this even faster. https://github.com/jadidbourbaki/llama.cpp/pull/12
I’ll benchmark his change and add it to the article, crediting him for this improvement.
- PicardsFlute
So, instead of resolving it privately like you were asked, you come here and complain? Was the bot posts on Reddit not enough for you? Pro tip: There is a right way and a wrong way to approach these things and you are MOST certainly approaching this the wrong way. And before you go accusing me of anything: I am NOT associated with that project, but I DID see what happened. You are acting like a damned child and should be ashamed of yourself.