Infra behind Krea 2: How to train and serve at scale — Gabriel Jorge Menezes, Krea.ai
Gabriel Jorge Menezes from Krea.ai explains how to train large AI models across thousands of GPUs.
He shows that many standard metrics are misleading—for example, GPU usage displayed 100 percent constantly even when the cluster wasn't running efficiently. Instead, he tracked tensor core utilization. When training ran on larger images (from 128 to 1024 pixels), this metric increased noticeably. To monitor network communication between servers—a common source of errors—they built custom tools since no off-the-shelf solutions existed. A GPU exceeding 78 degrees is immediately removed from the system to prevent slowing down the entire training. The Krea 2 model was frequently saved with checkpoints to very fast storage that could write a terabyte in under 30 seconds, enabling training to recover quickly after crashes.
Vibekollen prepared this summary with AI from the original publication. The content belongs to AI Engineer.