Run a 125B AI model on your gaming PC at 100 tokens per second
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

Strata, an open-source inference engine, runs Qwen3.8-Flash-Next on consumer GPUs with 12GB VRAM. It splits the model's 24,576 experts across GPU, RAM, and SSD, achieving up to 94 tokens/s on an RTX 5070 and an estimated 100-140 tokens/s on an RTX 3090. One-click install for Windows and Linux, with OpenAI/Anthropic-compatible API on localhost and optional image input.
Strata makes the model fit by sharing the work across your whole PC.