Physical AI's Next Bottleneck Is Finding the Right Video — Rafael Levi, Bright Data
AI without data is just a box." Rafael Levi of Bright Data points out that LLMs train on trillions of words, but robotics has only about a million videos of robots doing things.
Paying people to record staged actions produces biased data, because nobody opens a door naturally when told to. Meanwhile, the web holds billions of hours of real people handling objects and real physics, with gravity, motion and cause and effect. The catch is noise. Levi notes that NVIDIA discards about 96% of the video it downloads to train Cosmos, and Stable Video Diffusion discards 74%, which wastes compute, bandwidth and storage. Meta, on the other hand, trained on about a million hours of public video and needed only 62 hours of real robot data to control a robot. Bright Data's approach is "search first, collect second." It indexes more than a billion videos by the actions in them, not their titles, and returns trimmed clips with timestamps, match scores and frame counts through an API. Levi demos searches for dishwashing and clothes-folding clips and covers uses in self-driving and brand discovery.
Texten är källans egen beskrivning av publiceringen. Innehållet tillhör AI Engineer.
Mer att läsa
Building advertising for the way people use AI
OpenAI för 2 tim sedan
Server-Side Code Execution Tools for AI Agents, Compared
OpenRouter för 12 tim sedan
v0.40.0
Ollama för 12 tim sedan
Google froze its open source bug bounty program due to a ‘significant rise’ in AI submissions
TechCrunch AI för 16 tim sedan