Skip to content
VibekollenBETAVibekollen
VideoAI Engineer

Speech-to-Speech Model Research at Google DeepMind — Valeria Wu Fon & Tom Ouyang, Google DeepMind

Valeria Wu Fon and Tom Ouyang from Google DeepMind present research on speech-to-speech models — systems that listen to speech and respond with speech directly, without converting to text.

Before 2018, recognizing speech required multiple separate components; modern models do it end-to-end. Their model trains on audio, video, and text together, allowing it to handle things nobody explicitly programmed — for example, keeping English terms that a Spanish speaker actually uses. They identify three tensions: the system must be fast for natural conversation, intelligent to understand instructions, and multimodal to accept both speech and visual material. Increasing one capability often weakens the others, and they demonstrate this with demos of live translation in meetings, roadside assistance for vehicle registration, and the system's ability to know when it's not its turn to speak.

Open on YouTube →

Vibekollen prepared this summary with AI from the original publication. The content belongs to AI Engineer.

More to read