VideoAI Engineer
The Best Models Still Reason Like Toddlers — Andrew Dai, Elorian
Show a frontier model part of a chessboard, ask how many white squares are visible, and it answers 32. Andrew Dai's diagnosis: it saw a chessboard, knew chessboards have 32 white squares, and hallucinated the rest. The pattern matching that makes these models superb at naming flowers hurts them once a question needs counting or spatial grounding: they miscount a Catan player's roads from the pieces left off the board, and they miss a robot arm lifting a lid because they cannot hold state across a long video. His test for understanding versus reasoning: if a person can answer in one second, so can the model; if it takes longer, the model falls apart. The benchmarks hide this: one popular reasoning suite uses 32 by 32 pixel images, and a multimodal science exam can mostly be answered without the image. What is missing, he says, is visual thinking. Video generators produce cartoonish explosions because their training data is Hollywood and game engines; detection models are robust but passive.
23 sep. youtube.com