Published September 19, 2026 in Technology

Grok Transcribe 2.0 is twice as accurate at the same price

TMRW Editorial
By TMRW Editorial
Editorial desk
Grok Transcribe 2.0 is twice as accurate at the same price
4 min read
Share this post

Cover: AI-generated editorial composition by TMRW, based on SpaceXAI’s 18 September launch post. Independent streaming ranks are from Artificial Analysis.

Most speech-to-text upgrades are a press release and a slightly nicer demo. SpaceXAI’s Grok Voice Transcribe 2.0, shipped 18 September, is narrower and more useful. Same price as version 1.0. Roughly twice the accuracy on the company’s own tests. First place among 32 streaming models on Artificial Analysis’s public leaderboard.

If you record meetings, voice notes, or Loom videos and then need the words, this is the kind of change you can feel. It is not a new assistant. It is a better ear.

The numbers, and who measured them

According to the launch post, Transcribe 2.0 is built on the same audio foundation model as Grok Voice, the system SpaceXAI says already handles tens of thousands of support calls a day and the Grok assistant in Tesla cars. Batch transcription stays $0.10 per hour of audio. Streaming stays $0.20. Diarization, timestamps, and key-term biasing are included. Version 1.0 will be deprecated in the coming weeks; pin grok-voice-transcribe-1.0 if you need it.

Public streaming numbers, as reported by Artificial Analysis: 2.7 percent word error on the first final transcript, at 0.49 seconds after the end of speech, down from 3.9 percent for 1.0. That won the streaming accuracy ranking. It did not win on speed. Muse Voice Transcribe and ElevenLabs Scribe v2 Realtime are faster to a first transcript, with slightly worse error rates. In batch, Transcribe 2.0 ranked fifth of 59 at 2.3 percent word error, behind Alibaba, StepFun, Microsoft, and ElevenLabs Scribe v2.

The company’s own claim worth testing is multilingual short phrases — the in-car, “navigate home” class of utterance with almost no context. SpaceXAI says word error on that set fell from 20.6 percent to 6.8 percent across 19 languages. If you take voice notes in more than one language, that is the number that matters more than a clean-English leaderboard.

Where it sits next to ElevenLabs

ElevenLabs is in our catalog for generation: voices, and now music. Scribe is their transcription product, and on batch accuracy it still sits ahead of Grok 2.0 in Artificial Analysis’s table. Grok’s pitch is live and messy audio — phones, overlapping talk, account numbers read aloud — at a price that undercuts Scribe’s realtime tier.

Atlassian is the named customer. SpaceXAI says Loom found 2.0 more accurate than its previous stack, and quotes SVP Sanchan Saxena on a loop from Loom recording to Cursor. That is a developer workflow. The everyday version is smaller: record the meeting, get a transcript you can skim, paste the action items into whatever you already use.

We have not run a bake-off on our own recordings. SpaceXAI’s internal sets (telephony, Grok chats, spoken credentials, short multilingual commands) are production traffic, which is the right kind of test, and they are still the vendor’s tests. If you already have a speech-to-text bill, the low-risk move is to pin 2.0 on a week of real calls and count names, numbers, and “um”s yourself. The list price did not go up, so the experiment is cheap.

Who should switch, who should wait

Switch, or at least trial, if you stream audio — support, live notes, in-car commands — and you are paying something like ElevenLabs or Deepgram rates for realtime. $0.20 an hour is the argument. Keep your current batch vendor if your files are clean studio speech and you already like Scribe or a specialist system that won the 59-model table.

Skip it as a consumer “app” story. This is an API default change. There is no new Grok button in your meeting tool unless that tool’s vendor adopts it, as Loom says it has. Grok’s separate consumer Voice Mode, rolling out on desktop and mobile, is a different product: talking to the chatbot, not turning a file into text.

The honest limitation is the usual one for speech models. Leaderboards average over datasets. Your conference room, your accent, and your product names are not those datasets. Key-term biasing — up to 100 custom words per request — is there for that reason. Use it. A model that is first on a public board and still misspells your company name is not first for you.