Samsung's LittleBit squeezes LLMs down to 0.1 bits per weight

Sub-1-Bit LLM Compression via Latent Factorization

SamsungLabs has released LittleBit, a method that compresses large language models into the sub-1-bit regime by factorizing weight matrices into low-rank latent factors, binarizing them, and restoring magnitude with learned scales. It supports 1.0 to 0.1 bits per weight and keeps the original architecture at inference. The follow-up LittleBit-2, accepted at ICML 2026, adds Joint-ITQ initialization to align latent geometry, available via an opt-in flag with no inference overhead.

LittleBit compresses large language models into the sub-1-bit regime by factorizing each dense weight matrix into low-rank latent factors, binarizing those factors, and restoring magnitude information through lightweight learned scales.
  1. hgoel

    I understand that the interest in these extreme quants stems from wanting to maximize the capability the average user can get from a local LLM in this era of ludicrously expensive memory, but I wonder if this is maybe targeting the wrong axis?

    We've been seeing various optimizations towards streaming, that have been much more impactful in the local AI space, e.g. MoE models where the less busy experts are offloaded to slow RAM or even pruned entirely, engram tables that can be read from NVMe instead of sitting around in RAM etc.

    Maybe the trick with these extreme quants would be to increase the total parameter count while quanting individual weights, such that maybe the active parameter count comes down, or streaming weights from RAM or disk becomes more efficient, or cache behavior improves? Say, replacing a single 4bit/weight matmul with 3 1bit/weight operations that produce a much closer result than a single 1bit/weight matmul would.

  2. augment_me

    Perf goes from 80% to 47% on Wikitext-2. Also no comparisons to FP4 solutions that are able to maintain or exceed perf on the same dataset 80% perf with a 4.25-4.5 big budget.

    I think more meaningful thing here would be a hybrid solution that went down to sub-bit representations when the informational representation does not need it (for example later layers) that still maintains task performance

  3. cpldcpu

    I understand the obsession with low bit quantization, but it is empirically quite evident that it is not possible to compress models to less than 4 bit per weight without severe loss of capabilities².

    It may be nice as an experiment, but it is obviously a very inefficient route for model training: spending all the flops on a saturated model only to prune its capabilties.

    ²As to why, I have seen few explanations. But the empirical evidence is there.

More from this day

2026-10-08