Speech-to-Speech Model Research at Google DeepMind — Valeria Wu Fon & Tom Ouyang, Google DeepMind
Valeria Wu Fon and Tom Ouyang from Google DeepMind present research on speech-to-speech models — systems that listen to speech and respond with speech directly, without converting to text.
Before 2018, recognizing speech required multiple separate components; modern models do it end-to-end. Their model trains on audio, video, and text together, allowing it to handle things nobody explicitly programmed — for example, keeping English terms that a Spanish speaker actually uses. They identify three tensions: the system must be fast for natural conversation, intelligent to understand instructions, and multimodal to accept both speech and visual material. Increasing one capability often weakens the others, and they demonstrate this with demos of live translation in meetings, roadside assistance for vehicle registration, and the system's ability to know when it's not its turn to speak.
Vibekollen prepared this summary with AI from the original publication. The content belongs to AI Engineer.