Handle step key correctly in dp_reshardable checkpoint save with --optimizer-cpu-offload - #69
Handle step key correctly in dp_reshardable checkpoint save with --optimizer-cpu-offload#69artkorenev wants to merge 27 commits into
step key correctly in dp_reshardable checkpoint save with --optimizer-cpu-offload#69Conversation
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
…rk#4) Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> Co-authored-by: Yueming Yuan <yym022502@gmail.com>
…adixark#5) Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
- Detach output layer params to prevent MTP gradient flowing to output layer - Add mtp_kwargs interface for flexible MTP label/loss_mask passing - Roll mtp_labels and loss_mask for RL training compatibility Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> Co-authored-by: Yueming Yuan <yym022502@gmail.com>
…yers (radixark#10) - Add is_mtp flag to MoE layers and multi_token_prediction module - Bypass routing replay for MTP layers (MTP uses fresh routing) - Replace rdxa/dev's built-in RouterReplay with miles.utils.routing_replay: - moe_utils.py: use get_routing_replay_compute_topk() wrapper - router.py: use register_routing_replay() for initialization Co-authored-by: Yueming Yuan <yym022502@gmail.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> Co-authored-by: Yueming Yuan <yym022502@gmail.com>
After bumping Megatron (rdxa/dev), colocated IPC weight update fails with torch.AcceleratorError: CUDA error: invalid argument during torch.multiprocessing serialization of CUDA tensors. Root cause: Megatron's new TMS hook (PR NVIDIA#3048) alters allocator behavior in training flow, causing allocations via cuMemCreate/cuMemMap which are incompatible with CUDA IPC (_share_cuda_() fails). Fix: resolve mapping.py and dynamic_context.py conflicts to isolate hook side effects so TMS/allocator state remains IPC-compatible during the weight update phase. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- merge(): truncate dp_reshardable padding on optimizer/param_state path - load_parameter_state_from_dp_reshardable: tolerate missing 'padding' key - ShardedTensor: relax flattened_range to deprecation warning Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This PR rebases from radixark Megatron fork [miles-20260218](https://github.com/radixark/Megatron-LM/tree/miles-20260218) and resolve conflicts. Upgrade Megatron from Dec 17 (3714d81) to Feb 13 (1dcf0da) PR link: radixark#13 Co-authored-by: Yueming Yuan <yym022502@gmail.com> Made-with: Cursor
…se `--disable-weight-backuper` in miles (radixark#18) Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: fzyzcjy <5236035+fzyzcjy@users.noreply.github.com>
…adixark#20) Co-authored-by: fzyzcjy <5236035+fzyzcjy@users.noreply.github.com>
Squash merge of the dense true-on-policy Megatron branch. Co-authored-by: zju-stu-lizheng <lizheng.cs@zju.edu.cn> Co-authored-by: zyxiyy02 <282300612+zyxiyy02@users.noreply.github.com> Co-authored-by: Yi Zhang <1109276519@qq.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: fzyzcjy <5236035+fzyzcjy@users.noreply.github.com>
…xark#58) Co-authored-by: Zhiyao Jiang <jessicajiang324@gmail.com>
…xark#60) Co-authored-by: zyzshishui <82826991+zyzshishui@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…-optimizer-cpu-offload` Graft of NVIDIA#2874 (f4502eb): wrap the optimizer `step` in LocalNonpersistentObject in sharded_param_state_dp_reshardable. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
The graft matches upstream exactly, but I traced the load path with
For 2 the easy fix is skipping (Only cpu-offload / torch AdamW runs are affected — FusedAdam keeps step in param_groups, which is also why no existing test catches this.) |
2e64d78 to
235952d
Compare
Grafts NVIDIA#2874 (
f4502eb1c9).Wraps the optimizer
stepinLocalNonpersistentObjectinsharded_param_state_dp_reshardableso it is not persisted into the rank-0common state.
🤖 Generated with Claude Code