Skip to content

[Dev] Fix mis-set decoupled gradient for Megatron-FSDP. - #4426

Merged
cspades merged 2 commits into
NVIDIA:devfrom
cspades:cye/mfsdp-decgrad-segfault-dev
Apr 25, 2026
Merged

[Dev] Fix mis-set decoupled gradient for Megatron-FSDP.#4426
cspades merged 2 commits into
NVIDIA:devfrom
cspades:cye/mfsdp-decgrad-segfault-dev

Conversation

@cspades

@cspades cspades commented Apr 22, 2026

Copy link
Copy Markdown
Member

What does this PR do ?

Main Branch: #4427

Related Consequences

  • Megatron-Bridge logic also needs to be updated to have use_precision_aware_optimizer activate megatron_fsdp_use_decoupled_grad. This caused seg-faults with clip_grad_norm when MLPerf upgraded MCore + MBridge and param.decoupled_grad = None while clip_grad_norm(use_decoupled_grad=True).

Testing

  • While FP8-Delayed does converge even without this fix, now MXFP8 (which falls outside of the scope of use_precision_aware_optimizer_no_fp8_or_ds_fp8 originally responsible for setting use_decoupled_grad=True) now also converges:
# MXFP8 + FusedAdam(use_decoupled_grad=True)
[2026-04-23 19:40:32.890713] iteration      100/15258789 | consumed samples:        12800 | elapsed time per iteration (ms): 10401.4 | throughput per GPU (TFLOP/s/GPU): 1297.3 | learning rate: 4.915263E-05 | global batch size:   128 | lm loss: 1.168019E-02 | loss scale: 1.0 | grad norm: 0.294 | num zeros: 0 | number of skipped iterations:   0 | number of nan iterations:   0 |

TODO

  • Update the TransformerEngine E2E test to print out loss for every training step, and update the MCore commit hash to point to this PR's fix.

Contribution process

Pre-checks

  • I have added relevant unit tests
  • I have added relevant functional tests
  • I have added proper typing to my code Typing guidelines
  • I have added relevant documentation
  • I have run the autoformatter.sh on my PR

Code review

Feel free to message or comment the @mcore-oncall to help accelerate your merge into main. The less complex your PR is, the faster it will be approved and merged!

All PRs start as draft. If you open a non-draft PR, it will be automatically converted to draft.

Step 1: Mark PR as "Ready for Review"

  1. When your PR is ready, click Ready for Review.
  2. An oncall reviewer is auto-assigned and expert reviewers are notified based on your changes.
    • Some PRs may jump straight to step 2. This is determined by .github/CODEOWNERS.

⚠️ Only mark as ready once merge-conflicts are resolved and the CI is passing.
Final Review might get declined if these requirements are not fulfilled.

Step 2: Final Review

For PRs that change megatron/core, once all expert reviewers have approved, the Final Review label is applied automatically and final reviewers are assigned.

For PRs outside megatron/core, this step is skipped.

Step 3: Approved

Once all required reviewers have approved, the Approved label is applied automatically.

Merge

Any member of mcore-engineers will be able to merge your PR.

For MRs into `dev` branch The proposed review process for `dev` branch is under active discussion.

MRs are mergable after one approval by either eharper@nvidia.com or zijiey@nvidia.com.

@cspades cspades self-assigned this Apr 22, 2026
@cspades
cspades requested review from a team as code owners April 22, 2026 15:23
@svcnvidia-nemo-ci svcnvidia-nemo-ci added this to the Core 0.16 milestone Apr 22, 2026
@cspades
cspades force-pushed the cye/mfsdp-decgrad-segfault-dev branch from c271cfd to fdb9849 Compare April 22, 2026 16:23
@cspades
cspades marked this pull request as draft April 22, 2026 16:52
@copy-pr-bot

copy-pr-bot Bot commented Apr 22, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@cspades cspades changed the title [Dev] Fix segfault caused by not using decoupled gradient for Megatron-FSDP. [Dev] Fix invisible decoupled gradient for Megatron-FSDP. Apr 22, 2026
@cspades cspades changed the title [Dev] Fix invisible decoupled gradient for Megatron-FSDP. [Dev] Fix invisible issues related to decoupled gradient for Megatron-FSDP. Apr 22, 2026
cspades added 2 commits April 23, 2026 20:22
…oupled_grad=False).

Signed-off-by: Cory Ye <cye@nvidia.com>
@cspades
cspades force-pushed the cye/mfsdp-decgrad-segfault-dev branch from fdb9849 to aba018e Compare April 24, 2026 03:24
@cspades
cspades marked this pull request as ready for review April 24, 2026 03:24
@cspades cspades changed the title [Dev] Fix invisible issues related to decoupled gradient for Megatron-FSDP. [Dev] Fix mis-set decoupled gradient for Megatron-FSDP. Apr 24, 2026
@cspades cspades added Final Review PR is in the "final review" stage module: megatron-fsdp labels Apr 24, 2026
@cspades
cspades added this pull request to the merge queue Apr 25, 2026
@svcnvidia-nemo-ci

Copy link
Copy Markdown
Contributor

🔄 Merge queue validation started!

You can track the progress here: https://github.com/NVIDIA/Megatron-LM/actions/runs/24920387376

Merged via the queue into NVIDIA:dev with commit 78858b2 Apr 25, 2026
63 of 64 checks passed
@cspades
cspades deleted the cye/mfsdp-decgrad-segfault-dev branch April 25, 2026 02:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants