DeepSeek's Free Gift Shrinks AI Memory by 437x—and Western Labs Are Quietly Cashing In

The AI Race Just Got Awkward

DeepSeek's Free Gift Shrinks AI Memory by 437x—and Western Labs Are Quietly Cashing In

The AI race has taken an awkward turn: Western labs like Anthropic and OpenAI are quietly adopting DeepSeek's KV cache optimizations, which cut memory use by up to 437x. DeepSeek shared these breakthroughs freely, and the recent silent releases of Claude Opus 5.5 and GPT-6.1 Sol show sharp drops in cache-read pricing—60% and 80%, respectively. The author questions why Chinese labs would hand such a lifeline to their loss-making Western rivals.

So the Chinese labs have thrown a lifeline to the Western loss-making labs, and I just have no clue as to why.
  1. cmiles8

    Why is it a problem that the Chinese labs are just distilling down Anthropic’s models? Aren’t Anthropic’s models not just distilling down other people’s work?

    Feels like Anthropic crying do as I say not as I do.

  2. slowin

    I'm also grateful to the Chinese labs for providing workarounds for the walled gardens that the US based AI companies are attempting to create.

    Does anyone know if there are any distillation datasets available? I'd love to see these distributed on BitTorrent. I think it's critical that AI be democratized and not isolated in the hands of a few private companies.

  3. bwest87

    The best explanation is that it's a goal of the CCP to generally commodotize LLMs, because LLMs will ultimately be a compliment to manufacturing (which China dominates), and you always want to "commodotize your compliments".

    I think this explains why they are open sourcing broadly. It's not to be nice. It's a strategic play by the Chinese government to help ensure there are many players in this race and not too much power accumulates to American labs (even if American labs benefit in the process)

  4. listless

    I'm beyond thankful that Chinese AI models are so good. I desperately want us to cure the myriad of maladies that humans suffer needlessly with on a daily basis. We're going to need more powerful models than we have now if we're gonna do that and the Chinese are providing the competition needed to push this thing as fast as we can.

    I realize "going as fast as we can" is not the most popular position atm. But I'm far more interested in what good we can do than 10% apocalypse scenarios. I volunteer with a charity for childhood brain cancer and I do not want to see another 4 year old die. I'm willing to risk anything to stop this.

  5. reedf1

    I've been running Qwen 3.8 27b (an opus 4.6 tier model), locally on a 5090 for just over two weeks @ 170 tokens/s. That's a frontier model from 9 months ago running on consumer hardware. Who knows where distillation and pruning gets us in another year.

  6. reticulates

    “So the Chinese labs have thrown a lifeline to the Western loss-making labs, and I just have no clue as to why.”

    I don’t think it is intentional but this is actually quite bad for the western labs.

    The entire booster narrative has been “look at how their revenue is growing! $10bn to $100bn ARR in under a year! This’ll be a multi-trillion IPO!” and the extrapolated future growth from $100bn to $500bn and $500bn to $1tn justified future investment… but that revenue was just because inference was expensive.

    The revenue growth story is all that matters pre-IPO. If revenue falls from $100bn to $50bn that’s very very bad optics for OpenAI and Anthropic even if they are now profitable, it completely destroys the growth narrative.

  7. eggbrain

    Performance optimizations don't just help the western labs, they also help with running more powerful/useful LLMs locally.

    If local LLMs get "good" enough, people will soon paying for subscriptions to ChatGPT and Claude, which hurts their revenue.

  8. user43928

    > All this must mean the Western AI companies are now extremely inference-margin positive.

    > So the Chinese labs have thrown a lifeline to the Western loss-making labs, and I just have no clue as to why.

    That inference wasn't profitable is a widespread myth.

    Analysis based on Kimi K3 suggests that OpenAI and Anthropic have margins well north of 95%: https://inferencex.semianalysis.com/run/kimi-k3-on-b200

    Over the last months I have seen news that OpenAI made breakthroughs in inference efficiency multiple times.

    I have no reason to believe that the leading US labs don't have their own optimizations, or that they learned of this particular optimization from DeepSeek.

  9. rglover

    > So the Chinese labs have thrown a lifeline to the Western loss-making labs, and I just have no clue as to why.

    "Thus the expert in battle moves the enemy, and is not moved by him."

    They figured out a clever method for avoiding excessive training costs via distillation. That forces the hand of frontier labs to move faster, produce better models, etc. (to avoid embarrassment and 'falling behind'—all the while shouldering most of the cost), which they can just keep distilling—or applying other techniques against—much to the dismay of said frontier labs.

    Checkmate.

  10. wren6991

    The doublethink required to simultaneously believe "our safeguards prevent our models from doing unsanctioned cybersecurity tasks" and "distillation is why Chinese models are getting better at cybersecurity tasks" is genuinely quite funny.

More from this day

2026-09-30