Intelligent transcription with Gemini 3.5 Transcribe
Article image or reusable cover for Google DeepMind
Google DeepMind introduces Gemini 3.5 Transcribe, a new speech-to-text model that converts voice directly into formatted and corrected text.
The model handles background noise, complex technical terminology, and removes filler words such as "um" and "ah". It comes in two versions: one for live streaming speech with less than one second latency, and one for recorded material that can identify and distinguish up to three speakers. Developers can use it via the Gemini API and Google AI Studio to build voice applications, real-time captioning, or conversation analysis. The model supports over 85 languages and can be customized for specialized vocabulary and spellings.
Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text.
Vibekollen prepared this summary with AI from the original publication. The content belongs to Google DeepMind.