Skip to content

[optim] run plan/metadata coordination over a gloo group - #62

Open
yueming-yuan wants to merge 26 commits into
miles-mainfrom
yueming/dist-ckpt-gloo-coordination
Open

[optim] run plan/metadata coordination over a gloo group#62
yueming-yuan wants to merge 26 commits into
miles-mainfrom
yueming/dist-ckpt-gloo-coordination

Conversation

@yueming-yuan

Copy link
Copy Markdown

Problem

During torch_dist checkpoint load, torch DCP's plan/metadata coordination (_DistWrapper gather/scatter of LoadPlans, coordinator = global rank 0) runs on the default process group. Megatron initializes that group as NCCL-only, so these object collectives are staged through GPU tensors and go over NCCL.

On a 64-rank GLM-5.2 744B run (16×GB300, NVL72), this makes rank 0 open P2P channels to all 63 peers (~35 channels each, 10MB+2MB shareable buffers per channel):

  • rank 0 NCCL allocations: 26.25 GB (4,562 allocs) vs ~3 GB on every other rank
  • the buffers live on the world communicator and are never freed, so rank 0 permanently carries ~24 GB less usable GPU memory — including through the colocate rollout window, where it becomes the fleet-wide binding constraint on sglang mem_fraction_static

The payload transported is only pickled plans/shard descriptors (names, offsets, lengths, storage keys — a few hundred MB at rank 0); tensor data never crosses ranks in this path (each rank reads its own shards from disk).

Fix

Route the coordination through a lazily-created gloo group, for both the load and save planning paths. Tensor I/O is untouched; only the transport of plan objects changes (CPU/TCP instead of GPU/NCCL), so checkpoint bytes are identical.

Verification (A/B, identical 16×GB300 jobs, only this patch differs)

metric (rank-0 GPU, physical) before (job 773) after (job 778)
NCCL alloc total during ckpt load 26.25 GB (4,562 allocs) 2.62 GB (447 allocs)
actor at first update start 46.0 GiB 22.4 GiB
whole-GPU peak during update 188.7 GiB 164.8 GiB
actor residual during rollout 27.7 GiB 4.1 GiB
update_weights wall time 67.7–75.2 s 63.3 s (no regression)

Load correctness covered by the run itself: checkpoint loads, step-0 metrics normal, and the run has per-tensor SHA256 weight verification enabled (--check-rematerialize-param-from-master-weight).

🤖 Generated with Claude Code

yueming-yuan and others added 26 commits February 25, 2026 19:22
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>

Co-authored-by: Yueming Yuan <yym022502@gmail.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
- Detach output layer params to prevent MTP gradient flowing to output layer
- Add mtp_kwargs interface for flexible MTP label/loss_mask passing
- Roll mtp_labels and loss_mask for RL training compatibility

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Yueming Yuan <yym022502@gmail.com>
…yers (#10)

- Add is_mtp flag to MoE layers and multi_token_prediction module
- Bypass routing replay for MTP layers (MTP uses fresh routing)
- Replace rdxa/dev's built-in RouterReplay with miles.utils.routing_replay:
  - moe_utils.py: use get_routing_replay_compute_topk() wrapper
  - router.py: use register_routing_replay() for initialization

Co-authored-by: Yueming Yuan <yym022502@gmail.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Yueming Yuan <yym022502@gmail.com>
After bumping Megatron (rdxa/dev), colocated IPC weight update fails with
torch.AcceleratorError: CUDA error: invalid argument during
torch.multiprocessing serialization of CUDA tensors.

Root cause: Megatron's new TMS hook (PR NVIDIA#3048) alters allocator behavior
in training flow, causing allocations via cuMemCreate/cuMemMap which are
incompatible with CUDA IPC (_share_cuda_() fails).

Fix: resolve mapping.py and dynamic_context.py conflicts to isolate
hook side effects so TMS/allocator state remains IPC-compatible during
the weight update phase.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- merge(): truncate dp_reshardable padding on optimizer/param_state path
- load_parameter_state_from_dp_reshardable: tolerate missing 'padding' key
- ShardedTensor: relax flattened_range to deprecation warning

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This PR rebases from radixark Megatron fork [miles-20260218](https://github.com/radixark/Megatron-LM/tree/miles-20260218) and resolve conflicts.

Upgrade Megatron from Dec 17 (3714d81) to Feb 13 (1dcf0da)

PR link: #13

Co-authored-by: Yueming Yuan <yym022502@gmail.com>
Made-with: Cursor
…se `--disable-weight-backuper` in miles (#18)

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: fzyzcjy <5236035+fzyzcjy@users.noreply.github.com>
Co-authored-by: fzyzcjy <5236035+fzyzcjy@users.noreply.github.com>
Squash merge of the dense true-on-policy Megatron branch.

Co-authored-by: zju-stu-lizheng <lizheng.cs@zju.edu.cn>
Co-authored-by: zyxiyy02 <282300612+zyxiyy02@users.noreply.github.com>
Co-authored-by: Yi Zhang <1109276519@qq.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: fzyzcjy <5236035+fzyzcjy@users.noreply.github.com>
Co-authored-by: Zhiyao Jiang <jessicajiang324@gmail.com>
Co-authored-by: zyzshishui <82826991+zyzshishui@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@yueming-yuan
yueming-yuan marked this pull request as ready for review July 7, 2026 05:00
@yueming-yuan yueming-yuan changed the title dist-ckpt: run plan/metadata coordination over a gloo group [optim] run plan/metadata coordination over a gloo group Jul 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants