Skip to content

Cherry-pick: Fix mtp_num_layers and clip_qk issues (#2581, #2776) - #3075

Merged
BestJuly merged 3 commits into
NVIDIA:mainfrom
BestJuly:lit/pick_pr2581_to_main
Feb 5, 2026
Merged

Cherry-pick: Fix mtp_num_layers and clip_qk issues (#2581, #2776)#3075
BestJuly merged 3 commits into
NVIDIA:mainfrom
BestJuly:lit/pick_pr2581_to_main

Conversation

@BestJuly

@BestJuly BestJuly commented Jan 26, 2026

Copy link
Copy Markdown
Contributor

What does this PR do ?

Cherry-pick two merged PRs from dev branch to main.

The followings are from original PR description.

PR#2581 Fix: don't enter branch if mtp_num_layers == 0

The condition if self.config.mtp_num_layers is not None evaluates to True when mtp_num_layers=0, causing the code to enter the MTP branch and crash on labels.clone() when labels=None.

Since range(0) produces an empty loop and torch.chunk(..., 1) is essentially an identity operation, entering this branch with mtp_num_layers=0 does nothing useful but requires labels to be passed.

PR#2776 Fix clip_qk for virtual pipeline size > 1

Add a None check for current_max_attn_logits in clip_qk.

When virtual pipeline is enabled with size > 1, in the first iteration step, only some model chunks have initialized the current_max_attn_logits, other chunks remain with current_max_attn_logits = None, which would lead to a TypeError when all_reduce is called in the following step.

⚠️ For major changes (either in lines of code or in its impact), please make sure to first share a design doc with the team. If you're unsure what's the best way to do so, contact the @mcore-oncall.

Contribution process

flowchart LR
    A[Pre-checks] --> B[PR Tests]
    subgraph Code Review/Approval
        C1[Expert Review] --> C2[Final Review]
    end
    B --> C1
    C2 --> D[Merge]
Loading

Pre-checks

  • I want this PR in a versioned release and have added the appropriate Milestone (e.g., Core 0.8)
  • I have added relevant unit tests
  • I have added relevant functional tests
  • I have added proper typing to my code Typing guidelines
  • I have added relevant documentation
  • I have run the autoformatter.sh on my PR

Code review

The following process is enforced via the CODEOWNERS file for changes into megatron/core. For changes outside of megatron/core, it is up to the PR author whether or not to tag the Final Reviewer team.

For MRs into `main` branch

Feel free to message or comment the @mcore-oncall to help accelerate your merge into main. The less complex your PR is, the faster it will be approved and merged!

(Step 1): Add PR label Expert Review

(Step 2): Collect the expert reviewers reviews

  1. Attach the Expert Review label when your PR is ready for review.
  2. GitHub auto-assigns expert reviewers based on your changes. They will get notified and pick up your PR soon.

⚠️ Only proceed to the next step once all reviewers have approved, merge-conflict are resolved and the CI is passing.
Final Review might get declined if these requirements are not fulfilled.

(Step 3): Final Review

  1. Add Final Review label
  2. GitHub auto-assigns final reviewers based on your changes. They will get notified and pick up your PR soon.

(Optional Step 4): Cherry-pick into release branch

If this PR also needs to be merged into core_r* release branches, after this PR has been merged, select Cherry-pick to open a new PR into the release branch.

For MRs into `dev` branch The proposed review process for `dev` branch is under active discussion.

MRs are mergable after one approval by either eharper@nvidia.com or zijiey@nvidia.com.

Merging your PR

Any member of core-adlr and core-nemo will be able to merge your PR.

Co-authored-by: Xin Yao <xiny@nvidia.com>
@BestJuly BestJuly changed the title Fix: don't enter branch if mtp_num_layers == 0 Cherry-pick PR#2581 and PR#2776 to main Jan 26, 2026
@maanug-nv maanug-nv added Expert Review [deprecated] Apply this label to indicate that your PR is ready for expert review. complexity: low module: optimizer labels Jan 27, 2026
@maanug-nv

Copy link
Copy Markdown
Contributor

Since these are both bug fixes, it seems a good idea to add unit tests to this PR to ensure same things don't happen again. For example, test model creation with mtp_num_layers=0 doesn't crash.

@jaredcasper

Copy link
Copy Markdown
Contributor

Please use descriptive names for PRs, not referring to PR #s.

@Phlip79
Phlip79 requested a review from ericharper January 28, 2026 16:57
@Phlip79 Phlip79 added the Final Review PR is in the "final review" stage label Jan 28, 2026
@BestJuly BestJuly changed the title Cherry-pick PR#2581 and PR#2776 to main Cherry-pick: Fix mtp_num_layers and clip_qk issues (#2581, #2776) Feb 5, 2026
@BestJuly
BestJuly enabled auto-merge February 5, 2026 01:53
@BestJuly
BestJuly added this pull request to the merge queue Feb 5, 2026
Merged via the queue into NVIDIA:main with commit 1b11076 Feb 5, 2026
45 checks passed
@BestJuly
BestJuly deleted the lit/pick_pr2581_to_main branch February 5, 2026 06:43
daiyaanarfeen pushed a commit to daiyaanarfeen/Megatron-LM that referenced this pull request Feb 23, 2026
…IA#2776) (NVIDIA#3075)

Co-authored-by: rj42 <lbkzman@gmail.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>
Co-authored-by: Juntao Wang <juntaow@nvidia.com>
yangbofun pushed a commit to xlm-research/Megatron-LM that referenced this pull request May 22, 2026
…IA#2776) (NVIDIA#3075)

Co-authored-by: rj42 <lbkzman@gmail.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>
Co-authored-by: Juntao Wang <juntaow@nvidia.com>
terminator123 pushed a commit to 021ai/Megatron-LM that referenced this pull request Aug 3, 2026
…IA#2776) (NVIDIA#3075)

Co-authored-by: rj42 <lbkzman@gmail.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>
Co-authored-by: Juntao Wang <juntaow@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

complexity: low Expert Review [deprecated] Apply this label to indicate that your PR is ready for expert review. Final Review PR is in the "final review" stage module: optimizer

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants