LRU is harder to beat than the KV-cache papers suggest
A block-granular simulator replayed 68,266 requests from 393 Claude Code sessions and 23,608 Mooncake requests to test whether smarter eviction policies can beat LRU. Three attempts—hazard-based prediction, physical recompute cost, and session-granular eviction—all failed. The reason: under capacity pressure, most recomputation comes from tight tool loops seconds apart, not idle sessions, so the 5-minute TTL never fires.
Requests arriving after a gap longer than 5 minutes account for 17.5% of recompute. Requests arriving within 10 seconds account for 33.1%.
- chaboud
I've been building latency-sensitive LLM systems for a while, and I've come to rely heavily on pre-fill-considerate mechanics like ping-pong overlapped async context construction. For interactive mechanics, the worst case, even if rare, is problematic.
A toy/simplified version lives here:
https://github.com/chaboud/goulash
Consideration of mutation rate (a sort of temporal Shannon-ish coding/ordering) lives in there (with some RoPE-friendly structuring). Note: That was a vacation project, not the day job, but similar principles apply even with larger models.
- augment_me
What I feel a bit annoyed by, and what I feel obviously LLM-run ablations like this fail to capture, is any kind of reflection around previous research or any kind of proof that this is the best you can do. You don't know, you pulled the lever and you got something, is the best? Can you do better? What is the constraint?
As an individual researcher, you do not have 22M$ to run a massive brute-force search for your problem. You are constrained to your little subscription and you will barely dip your toe in the sea of possible solutions to a problem. So letting Claude run an autoresearch loop on your problem and then having it summarize it for you brings 0 value because you dont know what the downsides and trade-offs of LRU caches were, and how you would possible solve it.
- eru
> It didn't work, and why it didn't work turned out to be more interesting than the policy would have been.
Spoken like a true Claude.
Snarking aside, I am glad that our AI agents make it cheap enough to do these experiments and publish these write-ups that people finally bother to publish null findings. Very useful!
- talolard
I work on inference at a neocloud, but opinions are my own .
The economics and thus tools you can throw at inference change at various scales .
As a “blunt” contrived example , on a gb300 the GPUs communicate super fast over nvlink, and the cards can offload kv cache to dram and then disk, “fast enough “ for these tool heavy agentic workloads.
Which come together to mean that at high enough scale and in the right scenario, we can work
with wild ttls on the kv cache and still comfortably hit SLAs and tokenomics.
- bob1029
LRU seems like the ideal strategy for most things LLM-related. Everything in this realm is about recency bias. I think it is a feature in this context, not a problem.
When I give an agent a piece of corrected information regarding a long running task, the last thing I want it to do is try and statistically compensate for the fact that it is new information. I want this new information to dominate the old information.