Maple-Preview - On-device ternary 20B MoE reasoning LLM
Show HN: Maple-Preview – ternary 20B MoE running at 120 tok/s on a iPhone
Maple-Preview is an open-source, 20B-A1B ternary-weight reasoning model that runs at blazing speeds on everyday devices. It achieves 200+ tokens/s on a Mac mini M4 and 127 tokens/s on an iPhone—13x faster than comparable 1-bit models. Despite its small footprint, it delivers state-of-the-art reasoning, solving IMO-level problems and competing with larger models. Designed for on-device adaptation, it can learn user preferences through weight updates, enabling personalized, privacy-preserving AI assistants. Experience the future of efficient, always-on AI with Maple-Preview.
We believe that the precision a model runs at should be the precision it learns at.
- walrus01
I wish that "small" LLMs would stop being confidently very incorrect. Admittedly this is a bit of an intentionally esoteric test, but the confident way in which it presents a totally incorrect answer is a bit concerning.
"please write 250 words on the etymology and history of the word schlong"
The actual origin of the word is from middle high German and Yiddish-speaking Ashkenazi Jewish communities.
For comparison qwen 3.6 35B A3B does perfect on this and will give a solid description of the word's real origins and how it has made it into casual profanity/vulgarity as used in US English, and even mentions specific stand-up comedians and famous public figures of specific ethnic/religious origin in the US NE who introduced it into wider use.
Ask it for something that's not a narrow niche scientific or technical field, but something that would be less common to make it into a 20B size model, and see just how it does.
chat test link: https://chat.deepgrove.ai/
- beautiful_apple
A benchmark table comparing to Qwen 3.5 35B-A3B seems strange when Qwen 3.6 35B-A3B has been out for some time and is significantly better.
I didn't notice the version difference when first reading the article! So this is a heads up to people like me.
- arjie
For small models like this, it’s super important that it works well at tool calling etc. imho because it can’t memorize facts and isn’t big enough to tell when it doesn’t know. I could use it for high quality tool routing or a backup fast model for smaller task set. E.g. I use GPT-5.6 for voice channels at home. I’d prefer to be able to have this do basic tool calls and stuff because of the local speed.
Will give it a crack as a quick model in my clawlike.
- kamranjon
“Current approaches to low precision primarily focus on converting models trained in full precision to lower bitwidths. We view this as fundamentally the wrong approach…”
Very excited to see how it performs, I’ve been a bit skeptical of the efficacy of converting existing models - really cool to see one trained from scratch in the ternary format.
- momojo
At this point I think Apple just needs to simply not do anything stupid and these small model makers are going to hand them models