The Best Models Still Reason Like Toddlers — Andrew Dai, Elorian
Show a frontier model part of a chessboard, ask how many white squares are visible, and it answers 32.
Andrew Dai's diagnosis: it saw a chessboard, knew chessboards have 32 white squares, and hallucinated the rest. The pattern matching that makes these models superb at naming flowers hurts them once a question needs counting or spatial grounding: they miscount a Catan player's roads from the pieces left off the board, and they miss a robot arm lifting a lid because they cannot hold state across a long video. His test for understanding versus reasoning: if a person can answer in one second, so can the model; if it takes longer, the model falls apart. The benchmarks hide this: one popular reasoning suite uses 32 by 32 pixel images, and a multimodal science exam can mostly be answered without the image. What is missing, he says, is visual thinking. Video generators produce cartoonish explosions because their training data is Hollywood and game engines; detection models are robust but passive.
Texten är källans egen beskrivning av publiceringen. Innehållet tillhör AI Engineer.
Mer att läsa
Building advertising for the way people use AI
OpenAI för 3 tim sedan
Server-Side Code Execution Tools for AI Agents, Compared
OpenRouter för 13 tim sedan
v0.40.0
Ollama för 13 tim sedan
Google froze its open source bug bounty program due to a ‘significant rise’ in AI submissions
TechCrunch AI för 16 tim sedan