DeepSeek-V4 Latent Reasoning - Thinking in latent space

Show HN: DeepSeek-V4 Latent Reasoning – moving "thinking" into latent space

DeepSeek-V4-Flash-0731-Latent-Reasoning is a self-contained model that performs reasoning in latent space, quantized to NVFP4 for efficient serving. It integrates a trained latent reasoning head and a learned stop head into the DeepSeek-V4-Flash backbone, enabling multi-step state tracking with a compression factor of 6. The model achieves an aggregate BBH score of 0.880 and is served via a custom vLLM fork with a production-grade runtime. Ideal for complex reasoning tasks, it offers a complete package with weights and serving code in one repository.

The work is now a complete, self-contained model. Every weight needed to serve it ships in one repository. The backbone is quantized down to NVFP4 so it fits on real silicon. And the latent loop is driven by a proper, benchmarked serving runtime.
  1. vikramkr

    Shout-out to anthropic for having their models have such a strongly distinct writing style and personality that you can recognize their work instantly! It's quite nice to have such an immediate signal that if I were to proceed, I would spend orders of magnitude more time and effort reading the the text than the person claiming author credit spent writing or even reading it themselves.

  2. GodelNumbering

    Am I missing something or the evals do not compare it to the baseline deepseek-v4-flash? Without a baseline comparison, it is hard to tell what works well and what doesn't

  3. dtj1123

    I really appreciate all the honesty here.

More from this day

2026-08-09