Skip to content
VibekollenBETAVibekollen
VideoAI Engineer

Voice agents with Realtime Video — Sidney Primas, LemonSlice

Sidney Primas from LemonSlice demonstrates how to build AI agents with avatars and video that can run continuously for hours.

The main challenge is that video is generated frame by frame—each new image is based on previous ones, and errors accumulate over time. LemonSlice solves this by training the model to only look backward (since future frames don't exist yet) and drastically reducing computational steps. A surprising issue is audio: expressions and facial movements depend on how sound is analyzed, but most audio models are trained on monotone audiobooks, so they built their own. The service costs roughly the same to run as a voice model, but according to Primas, the biggest future value lies in orchestrating computations across GPUs and CPUs without video stuttering.

Open on YouTube →

Vibekollen prepared this summary with AI from the original publication. The content belongs to AI Engineer.

More to read