AERIOXFLUX
← Frontier Labs
Frontier Labs · nvidia

Nvidia Cuts Trillion-Parameter RL Weight Sync to 150 Seconds

NeMo-DCR ships only the roughly 1% of weights that change per training step, turning an 87.5-minute cross-region transfer into a two-and-a-half-minute refit.

Flux Desk·2026-10-09·5 min read

Moving a full one-trillion-parameter checkpoint between two AWS regions takes 87.5 minutes. In reinforcement learning, where the training cluster has to push fresh weights to the rollout cluster after every optimizer step, that is not a transfer. It is a lunch break, repeated hundreds of times. A paper posted to arXiv on October 6 by researchers at Nvidia and Aalto University describes a system that does the same update in 150 seconds.

The system is called NeMo-DCR, short for delta-compressed refit, and the paper's title states the pitch plainly: bit-exact delta-compressed refit for scalable agentic RL at trillion-parameter scale. The authors are Songlin Jiang, Zhiyu Li, Terry Kong, Yu Yao, Youngeun Kwon, Bernard Nguyen, Ashwath Aithal and Mario Di Francesco. Jiang did the work during an Nvidia internship; Jiang and Di Francesco are at Aalto.

Why refit became the bottleneck

Agentic RL has a lopsided cost profile. The paper opens with the observation that it spends over 70% of its wall-clock time in rollout, the phase where the model generates trajectories by acting in environments. To keep up, labs increasingly split rollout from training and run it on separate hardware, sometimes in a different cluster or a different region entirely, wherever GPUs are available.

That split creates a synchronization problem, which the field calls refit. Every time the trainer updates the policy, the rollout workers need the new weights before they can generate the next batch. At small scale this is a footnote. At a trillion parameters, across a wide-area link, it can dominate the step.

Most of the model does not change

NeMo-DCR rests on a measurement rather than a new algorithm. The authors found that across six Qwen3 and Nemotron models, only 0.6 to 1.2% of BF16 weight elements change per step, averaged over the first five GRPO steps. Their real training runs averaged 1.2%. Sending the full checkpoint every time means shipping roughly 99% of the bytes the receiver already has.

So the system sends only the difference. Changed weights are encoded as XOR masks against the version already sitting on the rollout node, with absolute overwrite values used for tensors where XOR is not safe. The packed values and locations are compressed with zstd at level 1. On Qwen3 delta payloads, the mixed XOR encoding cut bytes by 38 to 40% compared with overwrites alone, according to the paper, and the compressed payloads for a 120-billion-parameter model came to 2.4 to 3.8% of the full checkpoint.

The part that makes this usable rather than merely clever is the guarantee. The authors claim, with a formal proof, that receivers end up with parameter and buffer bits identical to a dense refit. That matters because RL is sensitive to silent drift: a lossy update scheme that slowly corrupts the rollout policy would show up as noisy rewards long before anyone traced it to the transport layer. The paper notes bit-exactness is verified in the BF16 receiver configurations it evaluated, and that unsupported loader paths are rejected up front.

The numbers, and what they compare

The headline figures need a little care. The 150-second result is a one-trillion-parameter refit at a 3% change rate using a relay tree, in which a payload goes from the producer to one root and then fans out through a binary tree of receivers. The 87.5-minute baseline is a transport-only full-checkpoint transfer between regions. The authors say that reference excludes checkpoint save and load time, which flatters the baseline rather than the new method.

The 3% and 5% change rates are deliberate stress tests, higher than any rate the team observed in practice. At those rates, the paper reports 12 to 40 times faster refits across models from 30 billion to one trillion parameters. For a 120-billion-parameter model whose checkpoint weighs 247.2 GB, the full transfer took 750 seconds. NeMo-DCR brought that to 22.6 seconds over the relay tree and 24.5 seconds through object storage at 3%, and to 41.6 and 49.7 seconds respectively at 5%. Nvidia's NeMo RL design documentation separately reports 20.2 seconds for the same 120B case over S3 at 3% density.

The authors also killed receivers mid-refit to test recovery, and report that both transports tracked the mean reward and KL trajectories of a standard dense NCCL refit.

What is left to optimize is the network

One finding reframes where future gains come from. In the stress cases, the transport lower bound accounts for 77 to 94% of refit latency, against 18 to 45% for building the deltas. The encoding is no longer the slow part. Faster links between clusters are.

There are trade-offs between the two transports. The relay tree needs direct connectivity between clusters. Object storage works without it, but every rollout node downloads across the boundary once. The source side also keeps a baseline copy in host memory at roughly 1.04 times the checkpoint size.

This is not a lab curiosity waiting for a product. The reference implementation, about 7,000 lines of Python built on Megatron Bridge and vLLM's native weight loader, landed in the open-source NeMo RL repository as pull request 2444, which GitHub shows merged in July. Nvidia's documentation lists the current limits: it requires a non-colocated Megatron policy with a vLLM backend, the same starting checkpoint on both sides, unquantized BF16 or FP16 rollout weights, and GRPO as the algorithm. FP8 rollout is rejected.

Fast weight sync inside a single cluster, over a high-bandwidth fabric, is a problem other teams have attacked too. NeMo-DCR's target is the messier case where training and rollout sit on separate clusters, joined by ordinary networks. As RL post-training spreads across whatever GPUs a lab can rent, that is increasingly the normal case, and a 35-fold cut in the time spent waiting for weights is the kind of saving that compounds over every step of a run.

#nvidia#reinforcement-learning#nemo-rl#distributed-training#weight-sync

The state of AI, in flux.

The directory + magazine for AI tools and the workflows people use to make money with them.

🔥 The Sauce Drop

The week's highest-earning AI workflows, in your inbox.

Some outbound links are affiliate links — Flux may earn a commission at no cost to you; this never affects rankings. Earnings figures are self-reported and not guarantees of income; most people earn less, some earn nothing.