Running Kimi K3, a 2.8T LLM, on a Single Apple Silicon Mac
Running Kimi K3 on a M1 Mac
I built Deltafin to run Kimi K3, a massive 2.8 trillion parameter Mixture-of-Experts model, on a standard 64 GB Apple Silicon Mac. By streaming MXFP4 experts on demand and using fused NEON kernels, the system achieves exact, reproducible decoding despite the hardware limits. While inference is slow at roughly 16 seconds per token on an M1 Max, it proves that local execution of frontier models is possible without cloud dependency.
It is not fast — about 16 seconds per token on our M1 Max — but it is exact, reproducible, and it works on a 64 GB laptop.
- antirez
SSD streaming on an M5 Max 128GB: https://x.com/antirez/status/2082136334160818528
Soon decent speed across two Mac Studios with 512GB of RAM.
- Azantys
0.01 tk/s is unusable for anything, you would wait a whole day for just 1000 token of output, what is the point of projects like this?
- mannyv
Going to try this on my M1 Ultra 128gb.
The point of these engineering tricks is to see the envelope of what's possible. You can use these tricks to both run a bigger model on smaller hardware or run a smaller model on smaller hardware.
- acmnrs
The title should probably be edited to specify "M1 Max" instead of "M1 Mac". You aren't running K3 on a base M1 anytime soon. Either way, still a very impressive project.
- ALLTaken
Exactly my machine 64GB M1 Max
So happy about this! ♡
idk how people access (soldout) and even afford 512GB RAM MacStudio's. Isn't it $40k or so?