Bonsai 2 27B compresses a 27B model 9x while keeping 98.2% of its performance

Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

Bonsai 2 27B compresses a 27B model 9x while keeping 98.2% of its performance

Prism ML has released Ternary Bonsai 2 27B, a compressed version of Qwen3.8 27B that uses ternary weights to achieve 1.76 effective bits per weight and a 5.9GB footprint. It retains 98.2% of the full-precision model's aggregate benchmark performance, with strong results in coding, vision, and agentic tool use. The model supports a 262K-token context window, runs on NVIDIA GPUs and Apple devices, and is available under the Apache 2.0 license.

At this level of retention, compression becomes a deployment unlock: nearly the same capability, in a footprint that can run in far more places.
  1. simonw

    If you want to try out out the GGUFs from https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf#th... be aware that you need Prism's llama.cpp fork to get them to work, from https://github.com/PrismML-Eng/llama.cpp/releases/tag/prism-...

    This should work:

    cd /tmp

    # Get the Prism macOS runtime

    curl -fL https://github.com/PrismML-Eng/llama.cpp/releases/download/prism-b10685-7dffb15/llama-prism-b10685-7dffb15-bin-macos-arm64.tar.gz -o bonsai-runtime.tar.gz

    tar -xzf bonsai-runtime.tar.gz

    # Get the ~5.95 GB GGUF model:

    curl -fL https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf/resolve/main/Ternary-Bonsai-2-27B-PTQ1_0.gguf -o Ternary-Bonsai-2-27B-PTQ1_0.gguf

    # Run the server, I used port 8331

    ./llama-prism-b10685-7dffb15/llama-server \

    -m Ternary-Bonsai-2-27B-PTQ1_0.gguf \

    --port 8331 -ngl 99 -fa on -c 32768

    Then open http://localhost:8331 for the (very good) baked in llama-server web UI... or run a prompt via the API like this:

    uvx llm openai endpoint http://127.0.0.1:8331/v1 \

    --model bonsai-2-27b --responses hi

    That's running at ~20 token/second for me on an M5 Pro (after a server restart I got 44 token/second, not sure why), but I'm pretty sure something isn't working right, on startup the server said "ggml_metal_device_init: - the tensor API is not supported in this environment - disabling".

  2. miffy900

    I really wish people would stop saying N times smaller than something when making a comparison; that makes no sense - it's 1/9th (11.11%) the size. You don't get a smaller quantity by multiplying by a number greater than 1.0. You could instead reverse the subjects being compared - "the original model is 9x bigger than this new smaller, efficient model" or some such. That makes sense.

    I keep seeing this being used when people talk about efficiency or performance gains and it's just very unintuitive language.

  3. Aurornis

    These are small enough that you can run them entirely in the browser https://huggingface.co/spaces/webml-community/ternary-bonsai...

    Remember to clear the downloaded weights afterward.

    Like the last model, it's amazing they work as well as they do. Use it for any longer task and they fall apart spectacularly and in interesting ways.

  4. adrian17

    > Ternary Bonsai 2 27B uses ternary {−1, 0, +1} weights with FP16 group-wise scaling, for 1.76 effective bits per weight

    If I recall correctly, a recent post [1] has shown that Q2 quants (with like 2.6 bpw) of the same base Qwen model sit at the edge between "noticeably worse" and Q1's "useless". I took a quick glance at Bonsai's blog posts, and don't really see them comparing themselves to "typical" quants or explaining what's the special sauce that makes them better?

    https://news.ycombinator.com/item?id=49611128

  5. huseyinkeles

    Testing on a MBP m4 pro 24gb

    ~100t/s prefill, ~15t/s, dropping to ~10t/s later with 64k context.

    The issue is I have yet to find a useful agentic local llm that I can run on this machine.

    Just given a relatively simple task on a swift app, took 25 minutes, brainstorming like crazy but can not decide on what to do. Eventually I killed it. GPT 5.6 sol-medium took 3 minutes to complete the same task for reference.

  6. danbrooks

    Nice! Does anyone know how this compares to the Unsloth quantizations of this model?

    https://unsloth.ai/docs/models/qwen3.8#run-qwen3.8-guide

  7. zhiyan

    Awesome results. Opens up doors for a lot of people.

  8. jedbrooke

    Running at about 7-8 tok/s (~60 tok/s prefill) on a Mac Mini M2 16GB.

    So far feels smarter than Bonsai 1 27B, it’s slightly larger than the Q1_0 quant. Super exciting stuff :)

More from this day

2026-09-17