Skip to content
VibekollenBETAVibekollen
VideoAI Engineer

Computer Use at the Edge of the Statistical Precipice — Pierluca D'Oro, Programma Labs

Pierluca D'Oro demonstrates how AI agents can achieve high scores on standard tests by simply replaying recorded sequences of button presses—without actually looking at the screen or understanding what they're doing.

This works on many common benchmarks because the tests are too predictable. He presents the PRISM principles for making environments less exploitable (random variations, sandboxes, multiple starting states), and introduces DIGIWORLD with 15 mobile apps and 3.2 million verified test configurations. He also shows that current uncertainty measures are poor: a 95% confidence interval covers actual performance only 20% of the time, masking large differences between models.

Open on YouTube →

Vibekollen prepared this summary with AI from the original publication. The content belongs to AI Engineer.

More to read