Qwen3.8-Flash-Next: A New Architecture, Towards Ultimate Cost-Efficiency

Qwen3.8-Flash-Next: A New Architecture, Towards Ultimate Cost-Efficiency

Qwen3.8-Flash-Next is an open-weights multimodal MoE model that previews the architecture for Qwen4. It introduces a hybrid Gated DeltaNet + Qwen Sparse Attention (QSA) for efficient memory and retrieval, a Gated Residual stream with four branches, N-gram Embedding to scale capacity cheaply, and the Muon optimizer. With 125B total parameters (6B active) plus 51B N-gram embeddings, it cuts training cost to about 1/9 of Qwen3.7-Plus while improving coding and office tasks. It natively supports 262K context, extendable to 1M, and the production version is priced at $0.16/M input and $0.47/M output tokens.

Put simply: GDN efficiently “remembers,” while QSA precisely “retrieves.”

More from this day

2026-08-26