SpaceXAI is giving its speech-to-text technology a major upgrade with Grok Voice Transcribe 2.0, a new model designed to handle the kind of audio that tends to trip up transcription systems: noisy phone calls, overlapping speakers, accents, spoken email addresses and phone numbers, and short voice commands.
Announced on September 18, Grok Voice Transcribe 2.0 is the successor to Grok Voice Transcribe 1.0 and is available through SpaceXAI’s Speech-to-Text API. SpaceXAI says the new model is twice as accurate as its predecessor in its real-world evaluations while keeping the same pricing.
The model is built on the audio foundation model behind Grok Voice, which SpaceXAI says already handles tens of thousands of customer-support calls each day, millions of hours of video narration, and voice agents in physical products, including the Grok assistant in Tesla vehicles. Its training data includes live, noisy, multilingual audio recorded across different environments, followed by additional post-training.
Built for messy audio
SpaceXAI’s focus with Grok Voice Transcribe 2.0 is less about perfectly clean recordings and more about what happens when audio gets difficult.
The company says the model ranks first for accuracy among 32 streaming speech-to-text models on the public Artificial Analysis leaderboard. Its internal testing also covers four different types of production audio: 8kHz telephony recordings from customer-support calls, conversations with Grok, spoken credentials such as phone numbers and email addresses, and short voice-assistant commands in 19 languages.
Multilingual speech is another area where SpaceXAI says the new model has made substantial progress. Grok Voice Transcribe 2.0 can automatically detect languages and handle switches between languages during the same recording.
On SpaceXAI’s internal short-phrase evaluation, the word error rate fell from 20.6% with the first-generation model to 6.8% with Grok Voice Transcribe 2.0. That is particularly relevant for voice assistants, where commands can be extremely short and provide little contextual information for identifying the language correctly.
A full set of transcription controls
The model supports both batch and real-time transcription. Developers can submit recorded audio files or URLs for transcription, or stream audio for low-latency transcription.
It also includes word-level timestamps with confidence scores, speaker diarization, and multichannel transcription for up to eight channels. Developers can provide up to 100 key terms per request to improve recognition of specialized vocabulary, product names, or other domain-specific terminology.
There are also formatting controls for numbers, dates, currencies, phone numbers, and email addresses. Filler words such as “um” and “uh” can be removed automatically, while Smart Turn detection can help voice agents determine when a speaker has finished talking.
The Speech-to-Text API supports a range of common audio formats, including WAV, MP3, WebM, OGG, and M4A. SpaceXAI’s documentation also lists support for multiple languages, including English, Spanish, French, German, Japanese, Korean, Hindi, Arabic, and several others.
Loom is already using it
SpaceXAI says Atlassian’s Loom is using Grok Voice Transcribe 2.0 to transcribe its videos.
According to SpaceXAI, Atlassian found the new model more accurate than its previous transcription solution and now uses it across Loom videos. The company also points to a workflow where a user can record an action plan in Loom, send the resulting transcript into Cursor, and use it to drive code changes.
That example illustrates where better speech recognition can become more than a simple accessibility or transcription feature. Accurate transcripts can serve as structured input for other AI systems, allowing spoken instructions to become part of larger automated workflows.
Same price as the previous model
Perhaps the most notable part of the announcement for developers is that SpaceXAI isn’t increasing the price.
Grok Voice Transcribe 2.0 costs $0.10 per hour for batch transcription and $0.20 per hour for streaming transcription. SpaceXAI says diarization, timestamps, and key-term support are included in those prices.
The model is available now, and SpaceXAI says it will soon become the default model for its Speech-to-Text API. Grok Voice Transcribe 1.0 will be deprecated in the coming weeks, although developers can temporarily continue using it by explicitly pinning the grok-voice-transcribe-1.0 model.
The current API documentation lists Grok Voice Transcribe 2.0 as the default model for new Speech-to-Text requests, reflecting the transition already underway.
With Grok Voice Transcribe 2.0, SpaceXAI is effectively pushing its voice technology further down the AI stack. Instead of being limited to the conversational Grok experience, the company’s audio models can now provide a transcription layer for customer support, video, voice assistants, multilingual applications, and other AI-powered workflows.
Discover more from GadgetBond
Subscribe to get the latest posts sent to your email.
