Muse Glimmer: A Memory Hierarchy Disguised as a 30B Transformer
Muse Glimmer is a memory hierarchy disguised as a 30B Transformer

Meta's Muse Glimmer packs a 30B-class multimodal agent into consumer hardware by rethinking memory. A 55 GiB checkpoint shrinks to under 20 GB via 4-bit quantization, while a hybrid attention schedule—local windows in most layers, global content-based attention every fourth layer—keeps the KV cache tiny. With only two KV heads and a 131K context, the model prioritizes semantic retrieval over precise positioning, making long-horizon agentic tasks feasible on-device.
The design favors robust semantic retrieval over precise global coordinate matching.