Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents
Article image or reusable cover for NVIDIA
NVIDIA presents its new Vera Rubin NVL72 processor, which handles AI agents 30 times more efficiently than previous models.
Agents consume much more data than simple chat — they search through databases, compare alternatives, and perform reasoning step by step, causing the volume of processed data to grow enormously. Vera Rubin also reduces cost per million processed tokens by 35 times. This matters for companies running AI factories, as it enables running more agents on the same power budget and earning more money per processed request.
According to OpenRouter data, agentic AI workloads consume 15x more tokens than a simple chat request.
Read the full story at NVIDIA →
Vibekollen prepared this summary with AI from the original publication. The content belongs to NVIDIA.