Port DeepSeek Sparse Attention to MambaModel - #3553
Conversation
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
duncanriach
left a comment
There was a problem hiding this comment.
Quick review.
I want this to merge after 3377. This will need to be adjusted to accommodate the changes in that PR
There was a problem hiding this comment.
Ugh, yeah. Checkpoint compatibility is an issue
|
Thank you for the quick review, Duncan! Should I maybe already start rebasing this on top of #3377? |
Maybe good to hold off a bit since there might be some more changes in the PR now. Quite a lot of feedback just came in, including from you :-) |
de61e61 to
bcb6b4f
Compare
|
/ok to test bcb6b4f |
df1bfb7 to
27f7c8a
Compare
|
/ok to test 769704d |
|
🔄 Merge queue validation started! You can track the progress here: https://github.com/NVIDIA/Megatron-LM/actions/runs/24533263668 |
|
/ok to test 931ca7c |
|
/ok to test bb32183 |
Also prefer modern type hint specification using native types.
And add corresponding tests. DSA = DeepSeek Sparse Attention
By using a generalized dictionary with the same sorting as in the hardcoded return dict, we get future-compatibilty in case other layers are added.
By using a dictionary with the same sorting as in the hardcoded list used to create the returned list (the specified return type of tuple is incorrect), we get cleaner code.
Concerns the `get_layer_maps_from_layer_type_list` method. This seems cleaner.
DSA = DeepSeek Sparse Attention
- New pytest test `test_dsa_gpt_mamba_equivalence.py` builds both a
GPTModel (DSA, 4 layers) and a MambaModel (pattern S-S-S-S-, 8 layers)
in-memory, remaps weights GPT→Mamba, and asserts logprob equivalence
across TP=1/PP=1, TP=2/PP=1, and TP=1/PP=2 distributed configs.
- New checkpoint conversion utility
`tools/checkpoint/remap_gpt_dsa_to_mamba.py` applies the same
layer-key remapping (decoder.layers.{N} → {2N}/{2N+1},
decoder.final_layernorm → decoder.final_norm) to DCP checkpoints.
Extend the DSA GPT/Mamba logprob equivalence suite to cover mixed dense+MoE architectures, mirroring the real DeepSeek-V3 layout where the first N layers are dense and the remaining layers use MoE. Key changes: - Add `pre_mlp_layernorm.*` routing in `_remap_gpt_to_mamba_state_dict` and `_remap_key` (checkpoint tool): MoE layers expose a real TENorm for `pre_mlp_layernorm` (not fused), which maps to MoETransformerLayer 2N+1. Dense layers use IdentityOp and produce no keys, so existing tests are unaffected. - Add `_make_dsa_moe_config` with `moe_layer_freq=[0,0,1,1]` (first 2 GPT layers dense, last 2 MoE) and proxy MoE params matching the DeepSeek-V3 style (4 experts, grouped-gemm, allgather dispatcher, shared experts). - Add `_MOE_MAMBA_PATTERN = "S-S-SESE"` and `TestDSAMoEGPTMambaEquivalence` with the same three parametrized tests as the dense suite (tp=1/2 pp=1/2): logprob match, strict weight loading, and golden-value recording/comparison.
|
/ok to test 6fc1c7c |
|
🔄 Merge queue validation started! You can track the progress here: https://github.com/NVIDIA/Megatron-LM/actions/runs/24567828712 |
|
🔄 Merge queue validation started! You can track the progress here: https://github.com/NVIDIA/Megatron-LM/actions/runs/24568144051 |
What does this PR do ?
Make experimental DeepSeek Sparse Attention (DSA) available to
MambaModel. For now, we error out when a user tries to use both standard Attention and DSA. It's not supported right now to have different Attention flavors in the same hybrid model, this requires broader restructuring.Pre-checks
Core 0.8)Code review
The following process is enforced via the CODEOWNERS file for changes into
megatron/core. For changes outside ofmegatron/core, it is up to the PR author whether or not to tag the Final Reviewer team.For MRs into `main` branch
Feel free to message or comment the @mcore-oncall to help accelerate your merge into main. The less complex your PR is, the faster it will be approved and merged!
(Step 1): Add PR label
Expert Review(Step 2): Collect the expert reviewers reviews
Expert Reviewlabel when your PR is ready for review.Final Review might get declined if these requirements are not fulfilled.
(Step 3): Final Review
Final Reviewlabel(Optional Step 4): Cherry-pick into release branch
If this PR also needs to be merged into
core_r*release branches, after this PR has been merged, selectCherry-pickto open a new PR into the release branch.For MRs into `dev` branch
The proposed review process for `dev` branch is under active discussion.MRs are mergable after one approval by either
eharper@nvidia.comorzijiey@nvidia.com.Merging your PR
Any member of core-adlr and
core-nemowill be able to merge your PR.