50 tok/s on a 24GB GPU: How I squeezed Qwen3.8 27B to 256K context
Qwen3.8-27B at 256K on a 24GB RTX PRO 4000 SFF (432 GB/s): 50 tok/s with MTP

A developer details how they achieved 50 tokens per second with Qwen3.8 27B at 256K context on a 24GB RTX PRO 4000 SFF GPU. The key was a custom quantized model that used NVFP4 for bulk matrices and higher precision for sensitive layers, plus an embedded MTP drafter and a patched llama.cpp build. The winning setup was not about any single best component but the fit between quant, drafter, kernels, and memory layout.
The winning setup came from the fit between the quant, drafter, CUDA kernels, memory layout, and workload. No component won on its own.