Nari Qwen3-TTS and Qwen3-ASR - High-accuracy, low-latency voice AI on the Pareto frontier

Show HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost

Nari Qwen3-TTS and Qwen3-ASR - High-accuracy, low-latency voice AI on the Pareto frontier

Nari Labs' Qwen3-TTS and Qwen3-ASR models lead Coval's voice AI benchmarks, delivering exceptional accuracy and low latency at unbeatable prices. Qwen3-ASR Fast achieves #1 latency (44 ms TTFS) and #2 WER (3.6%), while Qwen3-TTS Fast ranks #2 latency (63 ms TTFA) and #1 WER (3.8%). Both are the cheapest in their categories, making high-performance voice AI accessible for real-time applications like voice agents. Try them free during the public beta and get $20 in credits when you sign up.

Nari Labs leads Coval’s voice AI benchmark by sitting on the quality-latency Pareto Frontier for both Text-to-Speech and Speech-to-Text. We also lead the latency-cost and quality-cost Pareto Frontier out of all publicly available models on the benchmark.
  1. apimade

    https://apimade.com/audio-compare.html

    Added it to my blind TTS model comparison leaderboard. So far Darwin TTS is the open model leading the pack, ElevenLabs is at the lead.

  2. asaiacai

    This is really cool work! I'm curious like what do you see as the biggest lever for speeding up TTS models or from a technical perspective that this was a promising direction in the first place to push on. If I were to guess, some distillation but I'm certain there are probably TTS model aware architectural changes that just make inference wayyyy faster?

  3. recentlypostedj

    Question: How do you plan to differentiate, because there are so many TTS and its constantly changing every month who would become better

  4. karimf

    This is awesome. Thanks for pushing the audio pareto frontier forward.

    Probably far fetched for now, but I think the next big evolution is building the pareto/much cheaper alternative to GPT-Live-1.

    The STT/TTS market is quite saturated, while today, there's almost no cheap/open source alternative to GPT-Live-1.

  5. rahimnathwani

    For some reason it switched voices half way through a 33 second clip.

    For OP the clip name is nari-nina-01a0a12f-980a-765e-8029-fa56bd23210d.wav

More from this day

2026-09-14