Skip to content
VibekollenBETAVibekollen
VideoAI Engineer

Modality Misalignment and Originality Attribution in Short-Form Video — Aditya Gautam, Meta

A video is about sports for its first six seconds, then turns political for half a second.

Catching that across a hundred million plus videos is one of two problems Aditya Gautam works on at Meta; the other is unoriginal content, trivial to make with AI tools and corrosive to attribution. Both sit on messy data: adversarial uploads, multilingual on screen text, drift, and no ground truth. Mismatch across modalities is a solved embedding problem; mismatch within one modality over time is not, so three agents share it. A perceiver splits the video where temporal change happens, not at a fixed frame rate, and emits clip level embeddings, tags, and OCR. A reviewer runs the temporal analysis over that JSON and folds in live comments and sentiment. A retriever indexes topics, embeddings, and entities into inverted, vector, and graph stores for pulling similar clips and authors at inference time. Every agent runs on a small specialized VLM, not a frontier model, since he does not care whether it can code.

Open on YouTube →

The text is the source's own description of its publication. The content belongs to AI Engineer.

More to read