SpaceXAI launches a $0.20-an-hour transcriber that tops streaming accuracy

Grok Voice Transcribe 2.0 reached 2.7% final word error at 0.49 seconds, while its batch result placed fifth in Artificial Analysis testing.

By · Published

Primary source: Artificial Analysis

Why it matters

SpaceXAI is using low prices and benchmark-leading streaming accuracy to sell Grok as voice infrastructure, although faster rivals and stronger batch models keep the market open.

A close-up of a sleek digital screen showing abstract sound waves transforming into perfectly accurate digital text on a high-tech console.

SpaceXAI released Grok Voice Transcribe 2.0 on September 18th, pairing the lowest word error rate in Artificial Analysis' streaming benchmark with a price of $0.20 per hour of audio.

Artificial Analysis reported in a six-post thread that the model produced a 2.7% word error rate, or WER, on its first final transcript, down from 3.9% for Grok Voice Transcribe 1.0. The result put version 2.0 first for final-transcript accuracy among the 33 models in the evaluator's latest test.

SpaceXAI is the AI operation inside Elon Musk's SpaceX after SpaceX acquired xAI on February 2nd, 2026. The release extends Musk's Grok operation further into the infrastructure behind customer-service systems, voice agents and media transcription, where a dedicated model can be sold by the hour without requiring customers to adopt the Grok chatbot.

Accuracy comes with a latency trade-off

Grok Voice Transcribe 2.0 returned its first final transcript 0.49 seconds after the end of speech in Artificial Analysis' testing. That was slower than Muse Voice Transcribe, which recorded a 3.1% WER at 0.16 seconds, and Cartesia's Ink-2, which recorded a 3.3% WER at 0.21 seconds using semantic endpoint detection.

The pattern held for the first partial transcript, the earliest text a streaming system can act on. Grok Voice Transcribe 2.0 posted a 3.4% WER at 0.49 seconds, narrowly ahead of Muse Voice Transcribe and ElevenLabs Scribe v2 Realtime on accuracy. Both competitors reached 3.6%, while returning text sooner.

That distinction matters for voice agents. An accurate transcript gives the reasoning model cleaner input, but every fraction of a second spent waiting for text reduces the time available for reasoning, tool calls and speech generation before a conversation begins to feel sluggish.

Artificial Analysis' streaming methodology uses about eight hours of audio drawn from a private voice-agent dataset, European parliamentary proceedings and corporate earnings calls. Half of the combined score comes from the private agent-focused dataset, with the two public datasets receiving 25% each. The evaluator measures latency from the point at which its voice-activity detector identifies the end of speech.

Grok Voice Transcribe 2.0 did not lead the separate batch benchmark. It scored a 2.3% WER and ranked fifth among 59 models in the results cited by Artificial Analysis, improving from version 1.0's 4.0%. Alibaba's Fun-Realtime-ASR-preview and StepFun's StepAudio 3 ASR each reached 1.7%, while Microsoft's MAI-Transcribe-2 scored 2.0% and ElevenLabs Scribe v2 recorded 2.2%.

The split result defines the product clearly: SpaceXAI has optimized version 2.0 for developers building live systems, while the strongest offline transcription models still produce fewer errors when immediate output is less important.

SpaceXAI keeps the old price

SpaceXAI charges $0.20 per hour for streaming transcription, equivalent to about $3.33 per 1,000 minutes. Batch processing costs $0.10 per hour, or about $1.67 per 1,000 minutes. Those rates are unchanged from Grok Voice Transcribe 1.0, according to SpaceXAI's September 18th launch post.

The model supports recorded files and live streams, word-level timestamps, speaker labels, as many as eight independent audio channels and biasing toward as many as 100 customer-supplied terms. SpaceXAI also includes filler-word removal, automatic text formatting and turn detection for voice agents.

SpaceXAI's release notes listed grok-voice-transcribe-2.0 as available on September 17th. Version 1.0 remains the default when developers omit a model name, although SpaceXAI says version 2.0 will take over as the default and its predecessor will be deprecated in the coming weeks. Existing integrations can select the new model through the same speech-to-text API.

SpaceXAI says the underlying audio model already handles tens of thousands of customer-support calls each day and millions of hours of video narration. Those usage figures are self-reported. SpaceXAI also named Atlassian's Loom as a customer using version 2.0 to transcribe videos.

The API release turns the Grok voice stack into a direct challenge to ElevenLabs, Cartesia, Deepgram, Microsoft and a growing field of specialist transcription providers. SpaceXAI's pitch rests on a measurable combination: the best streaming accuracy in Artificial Analysis' current test at a price below several close competitors. Developers still have to decide whether the extra wait for Grok's transcript fits the response-time budget of a live agent.

Reader comments

Conversation for this story loads after sign-in.