Skip to content
VibekollenBETAVibekollen
BlogGoogle DeepMind

Intelligent transcription with Gemini 3.5 Transcribe

Article image or reusable cover for Google DeepMind

Google DeepMind introduces Gemini 3.5 Transcribe, a new speech-to-text model that converts voice directly into formatted and corrected text.

The model handles background noise, complex technical terminology, and removes filler words such as "um" and "ah". It comes in two versions: one for live streaming speech with less than one second latency, and one for recorded material that can identify and distinguish up to three speakers. Developers can use it via the Gemini API and Google AI Studio to build voice applications, real-time captioning, or conversation analysis. The model supports over 85 languages and can be customized for specialized vocabulary and spellings.

Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text.
Verbatim from the article at Google DeepMind
Read the full story at Google DeepMind →

Vibekollen prepared this summary with AI from the original publication. The content belongs to Google DeepMind.

More to read