Skip to content

Port DeepSeek Sparse Attention to MambaModel - #3553

Merged
Phlip79 merged 12 commits into
NVIDIA:mainfrom
janEbert:mamba-dsa
Apr 17, 2026
Merged

Port DeepSeek Sparse Attention to MambaModel#3553
Phlip79 merged 12 commits into
NVIDIA:mainfrom
janEbert:mamba-dsa

Conversation

@janEbert

@janEbert janEbert commented Feb 24, 2026

Copy link
Copy Markdown
Contributor

What does this PR do ?

Make experimental DeepSeek Sparse Attention (DSA) available to MambaModel. For now, we error out when a user tries to use both standard Attention and DSA. It's not supported right now to have different Attention flavors in the same hybrid model, this requires broader restructuring.

Pre-checks

  • I want this PR in a versioned release and have added the appropriate Milestone (e.g., Core 0.8)
  • I have added relevant unit tests
  • I have added relevant functional tests
  • I have added proper typing to my code Typing guidelines
  • I have added relevant documentation
  • I have run the autoformatter.sh on my PR

Code review

The following process is enforced via the CODEOWNERS file for changes into megatron/core. For changes outside of megatron/core, it is up to the PR author whether or not to tag the Final Reviewer team.

For MRs into `main` branch

Feel free to message or comment the @mcore-oncall to help accelerate your merge into main. The less complex your PR is, the faster it will be approved and merged!

(Step 1): Add PR label Expert Review

(Step 2): Collect the expert reviewers reviews

  1. Attach the Expert Review label when your PR is ready for review.
  2. GitHub auto-assigns expert reviewers based on your changes. They will get notified and pick up your PR soon.

⚠️ Only proceed to the next step once all reviewers have approved, merge-conflict are resolved and the CI is passing.
Final Review might get declined if these requirements are not fulfilled.

(Step 3): Final Review

  1. Add Final Review label
  2. GitHub auto-assigns final reviewers based on your changes. They will get notified and pick up your PR soon.

(Optional Step 4): Cherry-pick into release branch

If this PR also needs to be merged into core_r* release branches, after this PR has been merged, select Cherry-pick to open a new PR into the release branch.

For MRs into `dev` branch The proposed review process for `dev` branch is under active discussion.

MRs are mergable after one approval by either eharper@nvidia.com or zijiey@nvidia.com.

Merging your PR

Any member of core-adlr and core-nemo will be able to merge your PR.

@janEbert
janEbert requested review from a team as code owners February 24, 2026 00:15
@svcnvidia-nemo-ci
svcnvidia-nemo-ci requested a review from a team February 24, 2026 00:15
@svcnvidia-nemo-ci svcnvidia-nemo-ci added this to the Core 0.16 milestone Feb 24, 2026
@janEbert
janEbert marked this pull request as draft February 24, 2026 00:16
@copy-pr-bot

copy-pr-bot Bot commented Feb 24, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@duncanriach duncanriach left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Quick review.

I want this to merge after 3377. This will need to be adjusted to accommodate the changes in that PR

Comment thread megatron/core/ssm/mamba_hybrid_layer_allocation.py Outdated
Comment thread megatron/core/ssm/mamba_mixer.py Outdated
Comment thread tests/unit_tests/models/test_mamba_model.py Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ugh, yeah. Checkpoint compatibility is an issue

@janEbert

Copy link
Copy Markdown
Contributor Author

Thank you for the quick review, Duncan! Should I maybe already start rebasing this on top of #3377?

@duncanriach

Copy link
Copy Markdown
Contributor

Thank you for the quick review, Duncan! Should I maybe already start rebasing this on top of #3377?

Maybe good to hold off a bit since there might be some more changes in the PR now. Quite a lot of feedback just came in, including from you :-)

@janEbert
janEbert force-pushed the mamba-dsa branch 11 times, most recently from de61e61 to bcb6b4f Compare March 3, 2026 14:40
@janEbert

janEbert commented Mar 3, 2026

Copy link
Copy Markdown
Contributor Author

/ok to test bcb6b4f

@janEbert
janEbert force-pushed the mamba-dsa branch 3 times, most recently from df1bfb7 to 27f7c8a Compare March 4, 2026 17:46
@Phlip79 Phlip79 removed the Expert Review [deprecated] Apply this label to indicate that your PR is ready for expert review. label Apr 9, 2026
@Phlip79
Phlip79 removed the request for review from a team April 15, 2026 19:05
Comment thread megatron/core/models/mamba/mamba_model.py
@svcnvidia-nemo-ci svcnvidia-nemo-ci added Approved All necessary approvals have been made and removed Final Review PR is in the "final review" stage labels Apr 16, 2026
@janEbert
janEbert requested a review from a team as a code owner April 16, 2026 19:09
@janEbert
janEbert enabled auto-merge April 16, 2026 19:10
@Phlip79

Phlip79 commented Apr 16, 2026

Copy link
Copy Markdown
Member

/ok to test 769704d

@svcnvidia-nemo-ci

Copy link
Copy Markdown
Contributor

🔄 Merge queue validation started!

You can track the progress here: https://github.com/NVIDIA/Megatron-LM/actions/runs/24533263668

@Phlip79

Phlip79 commented Apr 16, 2026

Copy link
Copy Markdown
Member

/ok to test 931ca7c

@ko3n1g

ko3n1g commented Apr 17, 2026

Copy link
Copy Markdown
Contributor

/ok to test bb32183

janEbert added 12 commits April 17, 2026 14:28
Also prefer modern type hint specification using native types.
And add corresponding tests.

DSA = DeepSeek Sparse Attention
By using a generalized dictionary with the same sorting as in the
hardcoded return dict, we get future-compatibilty in case other layers
are added.
By using a dictionary with the same sorting as in the hardcoded list
used to create the returned list (the specified return type of tuple is
incorrect), we get cleaner code.
Concerns the `get_layer_maps_from_layer_type_list` method. This seems
cleaner.
DSA = DeepSeek Sparse Attention
- New pytest test `test_dsa_gpt_mamba_equivalence.py` builds both a
  GPTModel (DSA, 4 layers) and a MambaModel (pattern S-S-S-S-, 8 layers)
  in-memory, remaps weights GPT→Mamba, and asserts logprob equivalence
  across TP=1/PP=1, TP=2/PP=1, and TP=1/PP=2 distributed configs.

- New checkpoint conversion utility
  `tools/checkpoint/remap_gpt_dsa_to_mamba.py` applies the same
  layer-key remapping (decoder.layers.{N} → {2N}/{2N+1},
  decoder.final_layernorm → decoder.final_norm) to DCP checkpoints.
Extend the DSA GPT/Mamba logprob equivalence suite to cover mixed
dense+MoE architectures, mirroring the real DeepSeek-V3 layout where the
first N layers are dense and the remaining layers use MoE.

Key changes:

- Add `pre_mlp_layernorm.*` routing in `_remap_gpt_to_mamba_state_dict`
  and `_remap_key` (checkpoint tool): MoE layers expose a real TENorm
  for `pre_mlp_layernorm` (not fused), which maps to MoETransformerLayer
  2N+1. Dense layers use IdentityOp and produce no keys, so existing
  tests are unaffected.

- Add `_make_dsa_moe_config` with `moe_layer_freq=[0,0,1,1]` (first 2
  GPT layers dense, last 2 MoE) and proxy MoE params matching the
  DeepSeek-V3 style (4 experts, grouped-gemm, allgather dispatcher,
  shared experts).

- Add `_MOE_MAMBA_PATTERN = "S-S-SESE"` and
  `TestDSAMoEGPTMambaEquivalence` with the same three parametrized tests
  as the dense suite
  (tp=1/2 pp=1/2): logprob match, strict weight loading, and
  golden-value recording/comparison.
@janEbert

Copy link
Copy Markdown
Contributor Author

/ok to test 6fc1c7c

@svcnvidia-nemo-ci

Copy link
Copy Markdown
Contributor

🔄 Merge queue validation started!

You can track the progress here: https://github.com/NVIDIA/Megatron-LM/actions/runs/24567828712

@svcnvidia-nemo-ci

Copy link
Copy Markdown
Contributor

🔄 Merge queue validation started!

You can track the progress here: https://github.com/NVIDIA/Megatron-LM/actions/runs/24568144051

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Approved All necessary approvals have been made complexity: high

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants