DeepSeek V4 Flash runs on a single AMD MI300X at 168 tok/s

DeepSeek V4 Flash on a Single AMD MI300X

A GitHub repository details how to run DeepSeek V4 Flash on a single AMD MI300X GPU in production, achieving 168.6 tok/s single-stream decode without quantization or offload. It includes patches for FP8 format, MoE routing, and kernel tuning, plus a Docker Compose stack. The setup uses a 20 GB GPU KV cache and 96 GiB CPU offload, handling up to 64 concurrent streams.

A kernel that assumes OCP semantics on MI300X can be wrong by a factor of two in the scale domain.

More from this day

2026-08-04