Gradium ships a TTS beta it says starts speaking in under 50ms

The cited benchmark pages still show Gradium's existing model at 217ms to 236ms, leaving the beta's test conditions to be reconciled.

By · Published

Primary source: Gradium on X

Why it matters

Sub-50ms TTS would return most of a voice agent's latency budget to transcription and language-model reasoning, but only if the result survives full-stack, independent measurement.

A person wears a sleek wireless earbud, from which a subtle blue ripple of sound emanates, set against a softly blurred professional background.

Gradium, the voice AI developer co-founded by Neil Zeghidour (@neilzegh), released a new text-to-speech beta on September 22nd that it says produces its first audio in less than 50 milliseconds without sacrificing quality.

https://x.com/GradiumAI/status/2102438831668551687

poster=/api/storage/public-objects/tweet-videos/gradium-sub-50ms-tts-beta-voice-agents-poster-3fe4ecec.jpg|Video from @GradiumAI on X

In a two-post thread on X, Gradium said most competing models take more than 100ms and directed developers to test the beta in Gradium Studio. The launch places another experimental model behind gradium-tts-beta, the alias Gradium previously used for a public beta focused on accurately reading phone numbers, email addresses, financial identifiers and other structured text.

The sub-50ms figure is Gradium's claim. The short announcement does not define the percentile, testing region, network path, connection state or whether the clock includes leading silence. Those details decide whether a latency result describes model inference, the provider's internal stack or the delay an application user actually experiences.

The public benchmarks still measure the older model

The independent boards Gradium cited do not yet present an apples-to-apples confirmation of the new number. Coval's Gradium provider page says it last measured the company's default TTS configuration on September 22nd at a 236ms mean time to first audio and a 5.3% word error rate. That put Gradium eighth among 28 TTS models for latency in Coval's 30-day comparison.

Speko's text-to-speech board lists Gradium TTS at a 217ms median and 351ms p90 from US East. Speko also gives the measured Gradium voice a naturalness Elo score of roughly 1,585, placing it close to the top of that board. The published Gradium rows on both services identify the existing production model, rather than the newly announced beta.

That distinction matters because Gradium's claim would represent a substantial step down from its existing public results. A move from roughly 217ms to below 50ms would cut first-audio latency by more than three-quarters, assuming equivalent network locations, connection handling and audio-onset detection.

Gradium has already shown how benchmark definitions can reshape product development. In a September 9th technical post, Gradium and Coval explained that the company's earlier model returned audio quickly but placed a median 225ms of silence before the first phoneme. Coval began including that silence in perceived TTFA, pushing Gradium's reported result from 171.9ms to 429.6ms even though its infrastructure had not changed.

Gradium then trained a replacement that emitted audible sound from its first returned frame. That production model reached about 214ms on the board captured for the September post. The episode established the measurement standard against which the new beta will be judged: first playable, audible speech across the full network path, rather than the first byte or an internal inference timer.

Gradium is using the beta alias as its fast lane

The company first opened gradium-tts-beta on July 30th for an accuracy-focused model trained to handle the strings that frequently break automated calls. Gradium promoted the resulting model to production on August 31st, reporting a 216ms median on Coval and an 81% pass rate across a company-designed, human-rated set of 500 difficult sentences in English, French, German, Spanish and Portuguese.

That test covered spelling, phone numbers, emails, dates, booking references and other content where a fluent-sounding voice can still make a call unusable by dropping one character. Gradium published the evaluation set, but the reported pass rates came from its own comparison. Tuesday's claim that the faster beta maintains the "same quality" does not specify whether Gradium means the production model's naturalness, its hard-case accuracy or another internal evaluation.

The rapid release cadence follows a large financing push. Zeghidour founded Gradium with Olivier Teboul, Laurent Mazare and Alexandre Defossez, a group drawn from Kyutai and earlier roles at Meta, Google DeepMind, Google Brain and Jane Street. Gradium launched publicly in December 2025 with a $70 million seed round led by FirstMark Capital and Eurazeo. In July, Gradium said it had extended total funding to $100 million, added Nvidia as an investor and began establishing a San Francisco Bay Area office.

Gradium is spending that capital on a straightforward competitive bet: voice infrastructure wins when applications can respond quickly without mangling the information that completes the transaction. The new beta is available to test now. Establishing the claimed advantage requires the public boards to measure gradium-tts-beta under the same full-stack conditions used for the rest of the field.

Reader comments

Conversation for this story loads after sign-in.