World Models Need Causality, Not Pretty Pixels — Christopher Manning, Moonlake AI
In 2007, Google already had a language model trained on 2 trillion tokens, the same order of magnitude as today's frontier models.
Christopher Manning, Stanford professor, former director of the Stanford AI Lab and now at Moonlake AI, uses that as one stop in a tour of AI history. It runs from Dartmouth in 1956 and Shakey the robot to the rise of LLMs, and it ends at his argument for what comes next: embodied intelligence built on simulation. Manning argues that generative video like Genie 3 simulates observations but has no semantics underneath, so it can't support planning. Moonlake takes a single photo or short video and builds an action-conditioned world in code, with objects you can pick up, open and move. It even researches objects on the web to fill in what the camera can't see, like the tea bags inside a closed box. A loop inspired by Claude Code compares renders against reality to shrink the sim-to-real gap, with the goal of replacing 10,000 hours of teleoperation with simulation. The talk ends with Q&A on gaming, ontologies, physics and discovery in latent space.
The text is the source's own description of its publication. The content belongs to AI Engineer.