Intelligent transcription with Gemini 3.5 Transcribe
Article image or reusable cover for Google Gemini
Google introduces Gemini 3.5 Transcribe, a new speech-to-text model that converts voice directly into accurate, formatted text.
It handles background noise, specialized terminology, and natural speech errors better than previous models — and can, for example, remove filler words like "uhm", fix self-corrections, and recognize over 85 languages. Developers can use it through two APIs: one for streaming speech and one for recorded audio, and it is already integrated into Google apps such as Gemini and Android Rambler.
Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text.
Read the full story at Google Gemini →
Vibekollen prepared this summary with AI from the original publication. The content belongs to Google Gemini.