Grok Voice Transcribe 2.0 doubles accuracy at the same price

xAI released Grok Voice Transcribe 2.0, a speech-to-text model that is twice as accurate as its predecessor at the same price. It ranks first among 32 streaming models on the Artificial Analysis leaderboard and cuts word error rate on short multilingual phrases from 20.6% to 6.8%. The model supports batch and streaming, word-level timestamps, speaker diarization, multichannel transcription, and key term biasing. Atlassian Loom now uses it to transcribe every video.
We've always believed the best way to move work forward is to capture context once and let it flow everywhere. With Grok powering Loom's speech-to-text and Cursor turning that into code, we're closing the loop from context to code: record what you mean, and the work gets done. It's a glimpse of where AI-assisted development is headed.