Slotstream - Run 104GB Qwen3.8-Flash-Next on any Mac by streaming experts from SSD

Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s

Slotstream lets you run the massive Qwen3.8-Flash-Next model (125B MoE, 104GB at 4-bit) on Macs with far less RAM by streaming experts from SSD. It uses MLX and Swift, and provides an Ollama-compatible API. On a 48GB Mac, it achieves ~12 tok/s warm decode, with cold start to first token in ~3s, and peak memory of 32GB (auto-sized, cappable). The tool requires ~110GB free disk, works on Apple Silicon with macOS 14+, and includes a doctor command to check compatibility before downloading. It supports Ollama and OpenAI chat/generate endpoints, streaming, and various sampling options. Memory is elastic, resizing between requests. Installation is a simple curl script, and the binary is small; the weights are a one-time 104GB download. Slotstream is ideal for developers and AI enthusiasts who want to run large models locally on modest hardware.

Cache size changes speed, never output. Greedy decoding is byte-identical between a 4 GB cache and a 24 GB one, and that equivalence is a standing test.

More from this day

2026-09-01