Google rolls out Gemini 3.5 Transcribe with real-time transcription and 85+ language support


Google has announced Gemini 3.5 Transcribe, a new speech-to-text model designed for real-time and pre-recorded audio transcription. It handles background noise, specialized terminology and natural speech while converting audio into formatted text. The model is designed for use cases such as voice agents, real-time captioning and post-call analytics.

Gemini 3.5 Transcribe features

Gemini 3.5 Transcribe supports two transcription modes through separate APIs:

  • Real-time streaming: Provides continuous, bidirectional streaming with sub-second latency through the Live API using gemini-3.5-transcribe-live.
  • Pre-recorded audio: Transcribes recorded audio, meetings, call logs and other content through the Interactions API using gemini-3.5-transcribe, with speaker attribution and word-level timestamps.

The model can handle self-corrections in natural speech, remove filler words such as “um” and “ah,” and automatically format the resulting text. It is also designed to understand a speaker’s intent and recognize custom vocabulary.

Other capabilities include:

  • Custom vocabulary: Recognizes specialized jargon and unique spellings provided by developers.
  • Language support: Automatically detects and transcribes more than 85 languages, including regional accents and different dialects.
  • Multi-speaker identification: Attributes speech to up to three speakers in pre-recorded audio with timestamps. Support for more than three speakers is currently experimental.
  • Function calling: Can delegate tasks such as image generation and file analysis to other Gemini models. This is currently available in the Gemini app on macOS.
Accuracy and latency

According to measurements by Artificial Analysis, Gemini 3.5 Transcribe has an average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming use cases. It can also recognize alphanumeric information such as postal codes and order IDs in noisy, real-world environments.

Artificial Analysis reports a 70% improvement in time to final transcription compared with Google’s previous Chirp 3 transcription model. On the FLEURS benchmark, Gemini 3.5 Transcribe recorded a WER of 5.50% in streaming mode and 5.04% in non-streaming use cases across a selection of languages and locales.

Gemini 3.5 Transcribe across Google products

On Gboard for Android, the new Rambler feature uses Gemini 3.5 Transcribe to convert spoken thoughts into formatted text while filtering out filler words. Users can also use voice commands to make edits, correct misspellings and change the writing style. In Google Antigravity, the model can use screen context and chat history, with user permission, to improve transcription of file names, agent thoughts and active documents.

In Google AI Studio, the model is available in Build mode for creating apps using voice input. The Gemini app on macOS uses it to convert natural speech into formatted text and supports voice commands that can use screen context. It can call other Gemini models for tasks such as summarizing local files, repurposing text across apps and generating images at the cursor.

Google is also working on support for Chrome, which will allow users to dictate text into web fields, including replies, posts and prompts.

Developer platforms and feedback

Developer platforms including Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel and Vision Agents are using the Gemini Live API for voice-driven interfaces. These platforms manage the real-time media streaming infrastructure used by these applications.

Vivo, Intellitek Health and Lingopal have also shared feedback on Gemini 3.5 Transcribe, covering its latency, accuracy and language support.

Availability

Gemini 3.5 Transcribe is currently available in public preview for developers and enterprises, while consumer availability is limited to specific Google products, languages and countries.

  • Developers: Available through the Gemini API in Google AI Studio and Google Antigravity.
  • Enterprises: Available through the Gemini Enterprise Agent Platform, with Gemini Enterprise for Customer Experience coming soon.

For consumers, the model is available in the Gemini app on macOS in English and through Rambler on Android in select countries and languages. Support for Chrome is coming soon.