ShapeLearn's 13.1 GB Qwen 3.8 27B Quant Hits 99.63% of BF16 Quality
Shapelearn Qwen 3.8 27B (13.1 GB VRAM)
ByteShape has released full ShapeLearn quantizations for Qwen 3.8 27B, following an earlier Lite run. Across six GPUs, all five models sit on the measured quality-speed frontier. The default GPU-5 (IQ4_XS, 3.84 bpw) reaches 99.63% of BF16's aggregate benchmark score at 13.1 GB VRAM, while the smaller GPU-4 hits 98.72% at 11.0 GB and runs faster. Speculative decoding with MTP or DFlash2 boosts throughput further.
By “frontier,” we mean that no other plotted model is both faster and more accurate.
- _ache_
From my own test.
It's not faster than the unsloth model.
Disclarer: I'm unsing Vulkan on an AMD GC.
- syntaxing
I’m on a strix halo @ GPU-5 with MTP and I get 600 prefill and 30 TG which pushes it into a very usable range. The odd thing is that Dflash2 is really slow for me, like sub 10 TG.
- Schlagbohrer
Absolute treasure of a website with these graphs, thank you for sharing this. Huge help for me to find a faster model (smaller quantization) for my VRAM.
- npodbielski
Well I tested it on 7900XTX with the same prompts and their draft model gave me about 30t/s. Their own snippet of code with regular MTP model gave me 60t/s.
Also model with their draft answered incorrectly.
With MTP it answered correctly.
Question was: "Does MikroTik CRS312-4C+8XG-RM have combo ports?". The answer is Yes.
- kristianp
What's GPU-5?