Google Launches Gemini 3.5 Transcribe Speech Model with 4.0% Word Error Rate, 85 Languages
Google has released Gemini 3.5 Transcribe, a new speech-to-text model for developers, enterprises, and consumers, achieving a 4.0% word error rate in streaming scenarios and 2.6% in non-streaming mode. The model supports 85 languages, improves accuracy in noisy environments, handles specialized terminology, and can attribute speech to up to three speakers in pre-recorded audio. It is available in public preview through the Gemini API in Google AI Studio and Google Antigravity, and has been integrated into the macOS Gemini app and Android Gboard's Rambler feature.
Google has released Gemini 3.5 Transcribe, the latest speech-to-text model in its Gemini family, now open to developers, enterprises, and general users. The model focuses on improving transcription in noisy environments, with complex specialized terminology, and across natural spoken expression. It has already been integrated into the Gemini app on macOS and the Rambler feature of the Gboard keyboard on Android. Developers can use the model to build voice agents, real-time captioning tools, and post-call analysis workflows.
The new model adds several optimization features, including automatic self-correction, removal of spoken filler words such as "um" and "ah," and automatic formatting of output text. It can also delegate complex tasks such as file analysis and image generation to other Gemini models, creating a multi-model collaborative experience. On accuracy, Gemini 3.5 Transcribe achieves an average word error rate of 4.0% in streaming speech recognition scenarios, dropping to 2.6% in non-streaming scenarios. The model delivers better transcription performance in noisy environments and can more accurately recognize key alphanumeric information such as order numbers and postal codes.
In terms of language support, Gemini 3.5 Transcribe covers 85 languages, accommodating regional accents and dialects. For pre-recorded audio, it can perform timestamped speaker attribution for up to three speakers. The model also adapts to user-defined vocabularies, recognizing specialized terminology, unusual spellings, and industry-specific names. Developers can access the model in public preview through the Gemini API in Google AI Studio and Google Antigravity. General users can experience the features on Android and macOS, and Google plans to bring the functionality to the Chrome browser, where users will be able to dictate text directly into input boxes on any web page.
Google has recently accelerated the rollout of its Gemini product line. The company previously introduced Gemini 3.7 Flash, stepping up competition on AI model pricing. Gemini in Chrome for Android is now available to all users in the United States, and the Gemini assistant has also entered Waymo's next generation of autonomous robotaxis.
Why this event matters
The event has a measured impact on 2 industrys. The strongest current signal is positive for Artificial Intelligence, with intensity 50/100 and 70% confidence over a short term horizon.
Artificial Intelligence
- Direction
- positive
- Intensity
- 50
- Confidence
- 70%
- Horizon
- Short term
Cloud Services & Data Centres
- Direction
- positive
- Intensity
- 40
- Confidence
- 65%
- Horizon
- Short term
Impact figures are analytical estimates that combine direction, intensity, confidence and event importance. They are not investment advice.