ShapeLearn's 13.1 GB Qwen 3.8 27B Quant Hits 99.63% of BF16 Quality

Shapelearn Qwen 3.8 27B (13.1 GB VRAM)

ShapeLearn's 13.1 GB Qwen 3.8 27B Quant Hits 99.63% of BF16 Quality

ByteShape has released full ShapeLearn quantizations for Qwen 3.8 27B, following an earlier Lite run. Across six GPUs, all five models sit on the measured quality-speed frontier. The default GPU-5 (IQ4_XS, 3.84 bpw) reaches 99.63% of BF16's aggregate benchmark score at 13.1 GB VRAM, while the smaller GPU-4 hits 98.72% at 11.0 GB and runs faster. Speculative decoding with MTP or DFlash2 boosts throughput further.

By “frontier,” we mean that no other plotted model is both faster and more accurate.
  1. _ache_

    From my own test.

    It's not faster than the unsloth model.

    Disclarer: I'm unsing Vulkan on an AMD GC.

  2. syntaxing

    I’m on a strix halo @ GPU-5 with MTP and I get 600 prefill and 30 TG which pushes it into a very usable range. The odd thing is that Dflash2 is really slow for me, like sub 10 TG.

  3. Schlagbohrer

    Absolute treasure of a website with these graphs, thank you for sharing this. Huge help for me to find a faster model (smaller quantization) for my VRAM.

  4. npodbielski

    Well I tested it on 7900XTX with the same prompts and their draft model gave me about 30t/s. Their own snippet of code with regular MTP model gave me 60t/s.

    Also model with their draft answered incorrectly.

    With MTP it answered correctly.

    Question was: "Does MikroTik CRS312-4C+8XG-RM have combo ports?". The answer is Yes.

  5. kristianp

    What's GPU-5?

More from this day

2026-09-18