From VLM/VLA's to Embodied Agents — Armen Aghajanyan, Perceptron AI
Feed a model one hour of video and roughly a million visual tokens go in, yet the loss lands on about two percent of them, since the only ground truth is a transcript or a few labeled frames.
Armen Aghajanyan calls that a humongous waste, and predicting every pixel treats a background pixel with the same weight as a gripper tip or a contact point. Perceptron's answer is a perceptive objective that learns which percepts will matter, rather than hardcoding the gripper. The second problem is context bloat from cameras that never switch off. Patch averaging buys ten times compression, but his fix is data sparse mixture of experts, a router that decides per layer which tokens to read and which to skip. Left alone, the model zooms into the graph in a figure and spends more tokens on fruit when asked to segment fruit. Put together, that produced the model his team released a few weeks earlier, trained on a petabyte spanning text, images, video, and trajectories from desktop use to video games, which he says beats a frontier lab's embodied reasoning model at a fraction of the cost.
The text is the source's own description of its publication. The content belongs to AI Engineer.