DeepSeek-V4 Latent Reasoning - Thinking in latent space

Show HN: DeepSeek-V4 Latent Reasoning – moving "thinking" into latent space

DeepSeek-V4-Flash-0731-Latent-Reasoning is a self-contained model that performs reasoning in latent space, quantized to NVFP4 for efficient serving. It integrates a trained latent reasoning head and a learned stop head into the DeepSeek-V4-Flash backbone, enabling multi-step state tracking with a compression factor of 6. The model achieves an aggregate BBH score of 0.880 and is served via a custom vLLM fork with a production-grade runtime. Ideal for complex reasoning tasks, it offers a complete package with weights and serving code in one repository.

The work is now a complete, self-contained model. Every weight needed to serve it ships in one repository. The backbone is quantized down to NVFP4 so it fits on real silicon. And the latent loop is driven by a proper, benchmarked serving runtime.

More from this day

2026-08-09