Product1 publisher3 min readPublished
SpaceXAI holds transcription at ten cents an audio hour while claiming twice the accuracy
Version 2.0 arrives at the same price as version 1.0. The accuracy gain SpaceXAI reports is concentrated on short multilingual commands. The evaluation sets behind the claim are drawn from production support calls.
The Product Desk · Product desk

What happened
- Grok Voice Transcribe 2.0 was released on 18 September at $0.10 per hour of audio for batch work and $0.20 per hour for streaming, both the same prices as the version it replaces.
- Speaker labelling, word-level timestamps and key term biasing are now included at no extra cost, having been billable features until recently.
- SpaceXAI says the model ranks first for accuracy among 32 streaming models on the public Artificial Analysis leaderboard, and calls itself one of the most accurate.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
- exposure The caller on a support line has no contract with the AI vendor, so the buyer that routes those calls is the party holding the consent question for the person who read out an account number.
- cost A team running 1,000 hours of batch audio a month spends $100 on it. That puts the unit price below the threshold worth negotiating, and the leverage moves to data handling and latency terms.
- decision Anyone on version 1.0 has to schedule a migration within weeks, because pinning only covers the transition.
- contradiction The doubling is measured against the company's own predecessor while its multilingual set moves about threefold, so what a buyer gains depends on whether their audio resembles short commands or clean speech.
By SpaceXAI's own description, one of the four sets it measures word error rate against is people reading out account codes, phone numbers, email addresses and addresses aloud. The audio comes from production traffic. The company also says where the volume comes from: Grok Voice powers tens of thousands of customer-support calls a day, transcribes millions of hours of video narration, and runs the assistant in Tesla vehicles. It calls its training data "a unique dataset of live, noisy, multilingual audio recorded across a diverse set of environments". SpaceXAI did not say how the spoken-credentials audio is obtained, retained or consented to.
The doubling is measured against Grok Voice Transcribe 1.0, the company's own predecessor. No rival vendor's model is in that figure. On the 19-language voice-assistant set the improvement is larger than that: 6.8 divided by 20.6 is 0.33, so about two thirds of the errors are gone, roughly three times fewer. Short utterances leave a model almost no context for working out which language it is hearing, and the same problem shows up in the car. SpaceXAI has that traffic through Tesla.
What a buyer can check is thinner than what is claimed. The comparison charts are run against ElevenLabs Scribe v2 and Deepgram Nova-3, comparators the vendor picked. The more useful evidence is a buyer who moved: Atlassian says it found the model more accurate than its existing supplier and now transcribes every Loom video with it. The audio Atlassian moved over is video narration.
Everyone selling transcription is pushing the same way. OpenAI has released new voice API models, and Microsoft benchmarked its in-house transcription model against Whisper, Gemini and ElevenLabs across 25 languages. European vendors have gone at multilingual specifically, with DeepL running real-time voice translation across more than 40 languages and ElevenLabs listing transcription among five services on the UK government's cloud framework. Thenextweb.com argues that transcription is now priced as a commodity input and the sellers compete on what they can bundle. Streaming still costs double batch, so anything you need live is priced at twice the batch line.
Leaderboard placing is beside the point for a team. Two things decide it. The first is whether your audio looks like the hard cases: according to Thenextweb.com, clean-speech accuracy is solved and what is left is telephony, accents, crosstalk and in-car commands. The second is how many of your recordings contain someone other than your user.
That gives four boxes. Messy audio with no outside party, which is internal meetings and narration, is where a doubling actually shows up in your output and nothing else changes; switch on accuracy. Messy audio with an outside party, which is support telephony, is where the gain is biggest and where the negotiation is getting retention and purpose limitation in writing. Clean audio decides on latency and price in both boxes, and migrating only makes sense if you can measure the difference in your own transcripts.
One piece of paperwork before anyone signs. The announcement is published under SpaceXAI, the name xAI has carried since July after SpaceX acquired it in February. Coverage still calls the company xAI; its own footer does not.
What to watch
- An evaluation on telephony and crosstalk audio run by someone other than the vendors, which would show how much of the doubling a buyer actually gets.
- Any European regulator asking about retention and purpose limitation for the spoken-credentials corpus.
- Whether the batch price holds once rival vendors match the bundled speaker labels and timestamps.