Qwen3.8: A 2.4T-Parameter Open Model That Rivals Opus 4.8 on Coding Benchmarks

Qwen3.8-2.4T

Qwen3.8: A 2.4T-Parameter Open Model That Rivals Opus 4.8 on Coding Benchmarks

Qwen releases Qwen3.8-2.4T-A95B, a massive mixture-of-experts model with 2.4 trillion total parameters and 95 billion activated per token. It introduces a hybrid architecture combining gated DeltaNet and attention layers, supports a native 262K context (extendable to 1M), and features adjustable reasoning effort. Benchmark results show it competing with top proprietary models like Claude Opus 4.8 and GPT-5.6 Sol on coding and agentic tasks, with FP8 quantized weights available for efficient deployment via vLLM, SGLang, and TokenSpeed.

Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date.
  1. UncleOxidant

    We're all really waiting for the 3.8-27B which is due out in 2 days.

  2. WalterGR

    Submitted 3 hours ago with 63 comments so far: https://news.ycombinator.com/item?id=49273478

  3. xlayn

    I wonder besides big labs and REALLY big corporations who can run the 2.6TB q8...

    I mean, the 1 bit one is 508GB.

    Assuming 256k context size as irrelevant at those sizes, I would need 22 AMD 7900XTX (24GB vram) to run this for the 1 bit one, and 113 for the 2.6TB

    113 and assumming close to constant load and the gpus taking turns as it does on my machine with two gpus is around 113gpus*113W = 11KW... just to hold and run that one instance... and only god knows the cooling requirements...

    I guess just someone on Meta/Googly/Claudy will say... oh nice, let's download it and run it...

    How did the unsloth guy did to process a file that size?

    https://huggingface.co/unsloth/Qwen3.8-2.4T-A95B-GGUF

More from this day

2026-08-12