Taking Reinforcement Learning Cross Datacenter — Nan Jiang, Modal
Nan Jiang from Modal presents a technique for running reinforcement learning (AI training through trial-and-error) on GPU servers distributed across multiple continents.
The problem is that checkpoints — saved versions of a trained AI model — weigh around 500 GB and take hours to transfer between data centers, which prevents rapid model updates. His solution is to send only around 500 MB instead: a small patch containing only the most important changes. This works because less than 1% of the weights actually change between versions — not because gradients are sparse, but because updates are very small, often a thousand times smaller than the precision allows. Modal calls the implementation Stitch.
Vibekollen prepared this summary with AI from the original publication. The content belongs to AI Engineer.