Inception Labs unveils Mercury 2.5, a diffusion LLM rivaling frontier models at 1,107 tokens per second

Inception Labs has released Mercury 2.5, its most capable production model and, to its knowledge, the largest diffusion language model ever trained. The model delivers a 40% increase in intelligence over Mercury 2, with speeds up to 1,107 tokens per second on NVIDIA GPUs and a 260K-token context window. Priced at $0.20 per million input tokens and $0.75 per million output tokens (with an 80% launch discount), it targets latency-sensitive workloads like search, voice, and coding. Early adopters report dramatic latency and cost improvements: OpenCall cut P99 response times from minutes to one second, and Augment Code reduced context-compaction latency by 82% and cost by 90%. The release also previews Mercury Voice and Mercury Router, a dLLM-based routing system.
In voice, latency isn't an infrastructure detail. It is the pause a caller hears.
- Sphax
Got my hopes up when it said widely available GPUs that it would be open weights but it doesn’t seem like it sadly
- mring33621
I like the model.
FYI:
"If you do not want us to use your User Submissions to train our models, you can opt-out by setting the ‘Improve the model for everyone’ option under User Settings in the API Platform to OFF."
- gertlabs
Inception is one of the most interesting neolabs with their diffusion-based architectures. My understanding is that their primary business is low latency voice applications but they are seriously pursuing coding.
We tested Mercury 2.5 Preview, which is nowhere close to the frontier (and not advertised as such), but it's actually usable as a general-purpose chatbot. It's comparable in problem solving ability to some last-gen open weights models, and the price and cost make it compelling. However, they have not figured out general purpose tool use and agentic coding (their model performs worse on our problems when given a custom harness). If they do, I see a lot of real-time applications that the speed and cost will enable.