Google's Gemini 3.5 Transcribe cuts word error rate to 2.6%
Gemini-3.5-Transcribe

Google has unveiled Gemini 3.5 Transcribe, its most precise speech-to-text model yet, designed for intelligent voice interactions. The model converts raw audio directly into polished, formatted text, handling background noise, jargon, and disfluencies. It achieves a 4.0% word error rate (WER) in streaming mode and 2.6% in non-streaming, a 70% improvement in time-to-final-transcription over its predecessor Chirp 3. Available via the Gemini API and Enterprise Agent Platform, it supports over 85 languages, custom vocabulary, and function calling, and is already integrated into products like Gboard's Rambler and the Gemini macOS app.
Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text.