xAI released Grok Voice Transcribe 2.0 for speech recognition in real recordings with noise, accents, overlapping speakers, and phone-quality audio. xAI says it improves on version 1.0 in its production-derived tests and ranks first for accuracy among 32 streaming models on the public Artificial Analysis leaderboard. These are vendor evaluations and an external leaderboard result, each tied to its own test conditions. Read xAI’s feature announcement.

Audio workflows it may support
The model can transcribe support calls, interviews, video narration, and short voice commands. The API accepts up to eight channels, supports a vocabulary of up to 100 domain terms to help with names and jargon, and returns speaker labels and timestamps. It can detect multiple languages and handle language switches during a recording. Short phrases have little context, so numbers, email addresses, and proper nouns still deserve special review.
Pricing and migration
xAI lists batch transcription at $0.10 per hour of audio and streaming at $0.20 per hour. The company says 2.0 becomes the default while 1.0 is phased out over the following weeks. The two modes have different rates; estimating from source-audio duration is more useful than estimating from transcript length. After switching, check whether error types and downstream timestamps changed even if the endpoint still returns text.
Run a meaningful comparison
- Prepare human-verified recordings that represent phone noise, multiple speakers, accents, short phrases, and spoken numbers.
- Include a domain vocabulary in some requests and compare its effect on recognition.
- Track word error rate, speaker attribution, timestamp drift, and human correction time—not just whether the transcript reads smoothly.
- Check for personal or protected information and follow organizational rules for upload, retention, and access.