Fix recompute checkpointing + training CGs - #3919
Conversation
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
|
/claude review |
| # Skip checkpointing during CUDA graph warmup and capture, matching the behavior of | ||
| # CheckpointWithoutOutput. The graph captures all ops directly; recomputation cannot | ||
| # run inside a captured graph. | ||
| if is_graph_warmup() or is_graph_capturing(): |
There was a problem hiding this comment.
Doesnt this just disable recomputation during graph capture? is_graph_capturing() is true when capturing. Why cant recomputation work with graphs?
There was a problem hiding this comment.
Right, this only disables recomputation during graph capture. That's where I'm seeing errors in main right now: the first time we run through a training pass, we error out because we are capturing the graph while also attempting recompute.
Subsequent training passes seem to run fine.
jiemingz
left a comment
There was a problem hiding this comment.
There might be a way to get recompute working with per layer graphs but I am okay with explicitly disabling it like this unless needed.
|
🔄 Merge queue validation started! You can track the progress here: https://github.com/NVIDIA/Megatron-LM/actions/runs/25839862145 |
|
🔄 Merge queue validation started! You can track the progress here: https://github.com/NVIDIA/Megatron-LM/actions/runs/25858675752 |
* origin/main: (138 commits) Refactor CUDA graph API: decompose cuda_graph_scope into full_iteration impl, inference scope, and per-layer capture modules (NVIDIA#4292) Add high-priority A2A stream and HybridEP preprocessing SMs (NVIDIA#4694) add is_torch_min_version in fsdp src (NVIDIA#4812) [Main][feat] Support A2A Overlap for Megatron-FSDP (NVIDIA#3797) Reorder mtp_post_process after attention backward in 1F1B schedule plan (NVIDIA#4695) [fix] Use MSC for checking checkpoint existence (NVIDIA#4251) Combine GEMM + SwiGLU fused MLP PRs (3890, 4071, 4095, 4219, 4311, 4324) → main (NVIDIA#4636) Strengthen test_checkpoint to verify distributed checkpoint behavior (NVIDIA#4711) Disable MSC by default; opt in via --enable-msc (NVIDIA#4629) additional tests for nvrx (NVIDIA#4522) Update copy-pr-bot.yaml [skip ci] ci: tolerate git-gc race in /home/runner chown after checkout (NVIDIA#4808) Inference: Optimize Prefill Engine Steps for Nemotron (NVIDIA#4764) chore: Update nightly tests golden values (NVIDIA#4805) ci: Update workflow to use same commit for building docker image and running tests (NVIDIA#4787) Update owners (NVIDIA#4794) Bump nvidia-modelopt>=0.44.0 (NVIDIA#4803) fix tokenizers in respect to newer transformers (NVIDIA#4608) Use Protocols to type-check linear_proj submodules of Attention (NVIDIA#3434) Fix recompute checkpointing + training CGs (NVIDIA#3919) ... # Conflicts: # megatron/core/transformer/moe/moe_utils.py
Signed-off-by: yhgalaxy <yhgalaxy@outlook.com>
Signed-off-by: Jon Barker <jbarker@aws-cmh-slurm-1-vscode-02.cm.cluster>
Signed-off-by: Dmytro Pykhtar <dpykhtar@nvidia.com>
What does this PR do ?
Contribution process
Pre-checks
Code review
Feel free to message or comment the @mcore-oncall to help accelerate your merge into main. The less complex your PR is, the faster it will be approved and merged!
All PRs start as draft. If you open a non-draft PR, it will be automatically converted to draft.
Step 1: Mark PR as "Ready for Review"
.github/CODEOWNERS.Final Review might get declined if these requirements are not fulfilled.
Step 2: Final Review
For PRs that change
megatron/core, once all expert reviewers have approved, theFinal Reviewlabel is applied automatically and final reviewers are assigned.For PRs outside
megatron/core, this step is skipped.Step 3: Approved
Once all required reviewers have approved, the
Approvedlabel is applied automatically.Merge
Any member of mcore-engineers will be able to merge your PR.
For MRs into `dev` branch
The proposed review process for `dev` branch is under active discussion.MRs are mergable after one approval by either
eharper@nvidia.comorzijiey@nvidia.com.