Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL
Article image or reusable cover for Hugging Face
Hugging Face has updated its AsyncGRPOTrainer tool (version 1.14) to train adapters—small add-ons to AI models—instead of training the full model.
These new adapters are only a few megabytes and can sync through a storage bucket instead of between machines. A training job and two inference servers (vLLM) can now run on completely separate machines. A small proxy server routes each request to the right server and broadcasts adapter updates to all servers. In practice, training became five times faster: from 3 hours 27 minutes down to 53 minutes for 500 steps.
AsyncGRPOTrainer can now train a LoRA adapter and sync only that adapter to vLLM (TRL v1.14).
Read the full story at Hugging Face →
Vibekollen prepared this summary with AI from the original publication. The content belongs to Hugging Face.