From Voice Agents to AI Avatars with Alexander Smola - #777
Article image or reusable cover for TWIML AI
The podcast explores the evolution from today's voice assistants to AI agents that can both speak and see—so-called audiovisual agents and AI avatars.
Alexander Smola, founder of Boson AI and professor at Carnegie Mellon University, walks through the technical challenges: small delays, awkward interruptions, and wrong tone can quickly break the illusion. He explains technical tradeoffs such as audio tokenization (a way to split sound into units for AI), latency, and computational cost. A key insight is that when systems can both see and be seen, the demands become even higher—and emotional intelligence plus the ability to learn from human interaction become essential to move beyond impressive demos and achieve natural conversation.
Vibekollen prepared this summary with AI from the original publication. The content belongs to TWIML AI.
More to read
Server-Side Code Execution Tools for AI Agents, Compared
OpenRouter 9 h ago
v0.40.0
Ollama 9 h ago
Google froze its open source bug bounty program due to a ‘significant rise’ in AI submissions
TechCrunch AI 13 h ago
Can ‘super intelligence’ and a non-binding safety pact solve AI’s image problem?
TechCrunch AI 13 h ago