Neutrino-1 8B: A Single Model for Datacenter GPUs and Laptops

Neutrino-1 8B: A Single Model for Datacenter GPUs and Laptops

We introduce Neutrino-1 8B, a transformer model using a proprietary ternary format to shrink its footprint to just 3.88 GB. This design allows the same artifact to run efficiently on everything from H100 datacenter GPUs to 16 GB MacBook laptops without conversion. By decoding weights directly inside matrix kernels, we achieve high throughput and enable speculative decoding on consumer hardware while maintaining bit-exact performance across all platforms.

The bits of fp16 72.1 MMLU 763 tok/s Spec decode, H100

More from this day

2026-07-28