Google announces Gemini 3.5 Transcribe, leaves its API identity unclear
Google says its latest speech-to-text model handles structured strings, filler removal, formatting and custom vocabulary, while pricing and the access path remain undisclosed.
By RuntimeWire Staff · Published
Primary source: Google DeepMind
Why it matters
Gemini 3.5 Transcribe puts Google's audio work into a specialist speech-to-text pitch, but developers still lack a documented model identifier, access path, benchmark and price. Those omissions make the announcement difficult to compare with established transcription APIs.

Google DeepMind, the combined DeepMind and Google Brain unit formed in April 2023, announced Gemini 3.5 Transcribe on Aug. 26th as its latest speech-to-text model. Demis Hassabis co-founded the original DeepMind in 2010. The post did not disclose pricing, benchmarks or a public model identifier.
Google says the model handles complex phone numbers, postal codes and order IDs more accurately, including in noisy recordings. It can also remove filler words, format transcribed text and recognize custom vocabulary for names and product titles. The announcement provides no word-error-rate results, test set or latency comparison to support a broader accuracy claim.
Google DeepMind's feature announcement
Structured strings are the pitch
Google DeepMind is focusing its announcement on the details that often break transcription workflows. A system can produce readable prose and still fail when it encounters a customer's name, an unfamiliar drug or an internal product code. Errors in phone numbers and order IDs can make a transcript useless for downstream automation even when the surrounding conversation is accurate.
Custom vocabulary could help developers adapt transcription to those narrow domains. Google showed the feature in a simulated interface, however, and the image warned that compatibility and availability vary. The announcement does not establish whether custom vocabulary is exposed through an API, Google AI Studio or another product surface.
Google's audio transcription documentation describes transcription with Gemini models. It does not establish the additional limits and modes previously attributed to Gemini 3.5 Transcribe, including an 85-language count, code-switching, separate verbatim and smart modes, or vocabulary limits of 1,000 and 100 terms.
The model announcement is ahead of the catalog
Google's current Gemini model catalog lists Gemini 3.5 Flash and Gemini 3.5 Live Translate, but it does not list Gemini 3.5 Transcribe. That leaves the Transcribe model's public API name, endpoint, release channel and relationship to Google's existing Gemini audio models unresolved.
The announcement also omits rate limits, latency, supported languages and availability regions. Google's pricing page does not list a Gemini 3.5 Transcribe price. Gemini 3.5 Flash and Gemini 3.5 Live Translate have published prices, but those rates do not establish how Transcribe will be billed. Without a product-specific rate, developers cannot make a direct cost comparison with specialist transcription services.
Google also has an established Cloud Speech-to-Text service, while general-purpose Gemini models already accept audio. Developers evaluating Transcribe will need to know whether it is a distinct endpoint, a configuration of an existing Gemini model or a product layer built around capabilities documented elsewhere.
The release lands during a leadership handoff
The announcement comes as Hassabis steps away from daily management. Google recently said Hassabis would hand over day-to-day operational responsibilities and become chair of Google DeepMind and chief scientist of Alphabet. He will continue advising Google DeepMind's model and research teams while leading Isomorphic Labs.
Koray Kavukcuoglu, a DeepMind researcher for 13 years, is becoming senior vice president of Google DeepMind while retaining his role as Google's chief AI architect. Google says he will oversee Gemini model development, frontier AI research, the Gemini app and developer teams. His earlier work included WaveNet and DQN.
DeepMind was founded in 2010 by Hassabis, Shane Legg and Mustafa Suleyman. It later became part of Google DeepMind when Google combined DeepMind and Google Brain in 2023.
Transcription is a narrower assignment, with immediate operational stakes for call centers, meeting products, medical software and voice agents. Google's announcement establishes the intended product and a short list of capabilities. Its documentation has yet to establish the model identity, access path, performance or price that developers need for production decisions.