Kimi K3 (2.8T) runs at 1 token/s on a MacBook Pro, streamed from four SSDs
Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs
A new fork of Deltafin, ARGODRIVE, streams the full, unpruned 2.8-trillion-parameter Kimi K3 model from SSDs on Apple Silicon, achieving 1 token/s on a MacBook Pro. The project prioritizes exact model quality, using speculative decoding with small draft models but verifying every token with K3. Benchmarks on an M1 Max show steady improvements, and the code is open-source under MIT.
We choose to run the full 2.8-trillion-parameter model locally, and do the other things, not because they are easy, but because they are hard.
- lukeduff
Reminds me of Deep Thought from Hitchhiker's Guide to the Galaxy
- mandeepj
You currently can't run a 2.8T locally; there's just no way. So, it's a good start.
- dusted
A medium prompt in only 11 days.
- bluechair
I missed the explanation for how the SSDs are connected.
Maybe a dumb question.
- amelius
The SSDs are necessary because Apple's architecture doesn't allow RAM upgrades. Reminds me of someone who said "640KB ought to be enough for anybody".
- pjdesno
I wonder if faster SSDs would help?
In particular you can still get used Optane SSDs on eBay, although they’re fairly pricey. (not the bogus m.2 ones that are slower than a halfway decent consumer NVMe)
- alex7o
I think this is cool not for kimi but for sth like glm flash
- walrus01
Now imagine the token/s rate decline after context fill at 200,000+ context.