Skip to content

Add vLLM grouped gemm backend for MoE inference - #4566

Merged
santhnm2 merged 101 commits into
NVIDIA:mainfrom
santhnm2:vllm-grouped-gemm-cuda-graphable
May 4, 2026
Merged

Add vLLM grouped gemm backend for MoE inference#4566
santhnm2 merged 101 commits into
NVIDIA:mainfrom
santhnm2:vllm-grouped-gemm-cuda-graphable

Conversation

@santhnm2

Copy link
Copy Markdown
Contributor

What does this PR do ?

Ports the vLLM grouped gemm kernel to the inference-optimized MoE backend.

Issue tracking

For PRs from open-source community contributors:

  • New features: a linked issue is required. Please open a feature request and reference it here before submitting the PR.
  • Small updates (bug fixes, minor improvements): a linked issue is recommended and will accelerate the PR review process.

Linked issue:

Contribution process

Pre-checks

  • I have added relevant unit tests
  • I have added relevant functional tests
  • I have added proper typing to my code Typing guidelines
  • I have added relevant documentation
  • I have run the autoformatter.sh on my PR

Code review

Feel free to message or comment the @mcore-oncall to help accelerate your merge into main. The less complex your PR is, the faster it will be approved and merged!

All PRs start as draft. If you open a non-draft PR, it will be automatically converted to draft.

Step 1: Mark PR as "Ready for Review"

  1. When your PR is ready, click Ready for Review.
  2. An oncall reviewer is auto-assigned and expert reviewers are notified based on your changes.
    • Some PRs may jump straight to step 2. This is determined by .github/CODEOWNERS.

⚠️ Only mark as ready once merge-conflicts are resolved and the CI is passing.
Final Review might get declined if these requirements are not fulfilled.

Step 2: Final Review

For PRs that change megatron/core, once all expert reviewers have approved, the Final Review label is applied automatically and final reviewers are assigned.

For PRs outside megatron/core, this step is skipped.

Step 3: Approved

Once all required reviewers have approved, the Approved label is applied automatically.

Merge

Any member of mcore-engineers will be able to merge your PR.

For MRs into `dev` branch The proposed review process for `dev` branch is under active discussion.

MRs are mergable after one approval by either eharper@nvidia.com or zijiey@nvidia.com.

- Fix NCCLAllGatherDispatcher.set_step_metadata writing _valid_tokens_tensor
  to the subclass instead of InferenceAllGatherDispatcherBase, causing the
  Triton kernel to receive None and crash with 'constexpr has no attr is_ptr'
- Add use_allgather_v support to NCCLAllGatherDispatcher: non-CG steps
  (prefill) all-gather actual per-rank token counts, pad to max, AllGather,
  compact on dispatch; expand, ReduceScatter, truncate on combine
- Context passes use_allgather_v=not using_cuda_graph_this_step() for NCCL;
  both context and wrapper now all-gather actual per-rank counts rather than
  trivially filling (dummy forward always eager → use_allgather_v=True)
- Fix using_cuda_graph_this_step missing parentheses (property → method call)
@santhnm2

santhnm2 commented May 1, 2026

Copy link
Copy Markdown
Contributor Author

/ok to test 4a034a8

Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>
@santhnm2

santhnm2 commented May 1, 2026

Copy link
Copy Markdown
Contributor Author

/ok to test e7840c9

Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>
@santhnm2

santhnm2 commented May 3, 2026

Copy link
Copy Markdown
Contributor Author

/ok to test e5225d5

Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>
@santhnm2

santhnm2 commented May 4, 2026

Copy link
Copy Markdown
Contributor Author

/ok to test f3f2aca

Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>
@santhnm2

santhnm2 commented May 4, 2026

Copy link
Copy Markdown
Contributor Author

/ok to test a883044

@santhnm2
santhnm2 added this pull request to the merge queue May 4, 2026
@svcnvidia-nemo-ci

Copy link
Copy Markdown
Contributor

🔄 Merge queue validation started!

You can track the progress here: https://github.com/NVIDIA/Megatron-LM/actions/runs/25302945651

Merged via the queue into NVIDIA:main with commit bb979dd May 4, 2026
187 of 190 checks passed
@santhnm2
santhnm2 deleted the vllm-grouped-gemm-cuda-graphable branch May 4, 2026 06:08
Andron00e pushed a commit to Andron00e/Megatron-LM that referenced this pull request May 28, 2026
…ync-2026-05

Upstream tip: 932d9ee (Add periodic GPU sniff tests, 2026-05-07).

Conflict resolution:
- 13 modify/delete conflicts in .github/* and skills/*: kept deletion
  (matches policy from c037427 removing upstream-NVIDIA CI machinery
  not relevant to this fork).
- AGENTS.md content conflict: kept fork version (skills/contributing
  sections from upstream don't apply since the directories were removed).
- New upstream additions matching the same removal policy were dropped:
  .claude/settings.json, .github/workflows/nightly-sync-main-to-dev.yml,
  skills/{cicd,linting-and-formatting,nightly-sync,run-on-slurm,testing,
  update-golden-values}/SKILL.md.
- SECURITY.md (new from upstream): kept (generic security policy).

Verified post-merge:
- pretrain_gpt.py logging-patch hook intact (lines 405-408).
- megatron/core/optimizer/__init__.py emerging_optimizers + muon logic intact.
- megatron/training/arguments.py custom muon flags intact (lines 2297-2309).

Notable upstream changes pulled in:
- Removed legacy transformer + legacy GPT modules (NVIDIA#4207, NVIDIA#4322).
- Docker base image bump to 26.04-py3 (NVIDIA#4611).
- Inference fixes (vLLM grouped GEMM NVIDIA#4566, FlashInfer sampling NVIDIA#2456,
  EP sync NVIDIA#4607, MoE dispatcher fixes NVIDIA#4576).
- Gradient corruption fix with layerwise param all-gather overlap (NVIDIA#4609).
- Hybrid model + Flextron + GPU sniff tests + named validation sets +
  fault injection + InJob restart.
yhgalaxy pushed a commit to yhgalaxy/Megatron-LM that referenced this pull request Jun 17, 2026
Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>
Co-authored-by: Siddharth Singh <sidsingh@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Signed-off-by: yhgalaxy <yhgalaxy@outlook.com>
mathemakitten pushed a commit to mathemakitten/Megatron-LM that referenced this pull request Jul 1, 2026
Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>
Co-authored-by: Siddharth Singh <sidsingh@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
jon-barker pushed a commit to jon-barker/Megatron-LM that referenced this pull request Jul 10, 2026
Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>
Co-authored-by: Siddharth Singh <sidsingh@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Signed-off-by: Jon Barker <jbarker@aws-cmh-slurm-1-vscode-02.cm.cluster>
terminator123 pushed a commit to 021ai/Megatron-LM that referenced this pull request Aug 3, 2026
Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>
Co-authored-by: Siddharth Singh <sidsingh@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
svcnvidia-nemo-ci pushed a commit to dimapihtar/Megatron-LM that referenced this pull request Aug 4, 2026
Signed-off-by: Keshav Santhanam <ksanthanam@nvidia.com>
Co-authored-by: Siddharth Singh <sidsingh@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Signed-off-by: Dmytro Pykhtar <dpykhtar@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Approved All necessary approvals have been made complexity: high Run functional tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants