Skip to content

[Bugfix][Kernel] Bump DeepGEMM pin for SM107 JIT compatibility - #59503

Open
lisawym wants to merge 5 commits into
vllm-project:mainfrom
CentML:liswang/fix-deepgemm-sm107-jit
Open

lisawym wants to merge 5 commits into
vllm-project:mainfrom
CentML:liswang/fix-deepgemm-sm107-jit

Conversation

@lisawym

@lisawym lisawym commented Sep 30, 2026 •

Copy link
Copy Markdown

Overview

Bump both DeepGEMM pins to 550036ad, which includes the SM107 runtime JIT workaround merged in vllm-project/DeepGEMM#23. This addresses Rubin startup failures tracked in deepseek-ai/DeepGEMM#461.

Claims

  • Use the merged workaround through both CMake vendoring and the standalone installer, with matching dependency pins.
  • On physical SM107, DeepGEMM selects the compatible 100f JIT target while preserving device capability reporting and other architectures' compiler targets.

Validation

Check Result
Pre-commit on both changed files, shell syntax, and diff checks Passed
Fresh CMake FetchContent check with CUDA compilation disabled Fetched exactly 550036ad and initialized the expected CUTLASS/DeepJIT submodules
C++17 host harness compiling the fetched JIT header with mocked CUDA calls Only SM107 changed to 100f; SM90, SM100, SM103, SM120, SM121 and a synthetic SM108 were preserved

Commands run for the changed files:

git diff upstream/main --check
.venv/bin/python -m pre_commit run --files \
  cmake/external_projects/deepgemm.cmake tools/install_deepgemm.sh
bash -n tools/install_deepgemm.sh
bash tools/install_deepgemm.sh --help

Prior Rubin startup and serving checks used a custom build with the same JIT override at an earlier dependency revision, as described in DeepGEMM#23. The combined revision has not received a full CUDA build, GPU serving smoke test, or model accuracy evaluation in this update. Fresh GPU validation and CI are still needed; no performance improvement is claimed.

Details

The pinned CUTLASS 4.2.1 headers lack native SM107 feature guards. DeepGEMM#23 temporarily selects the compatible SM100 family JIT target for physical SM107 devices. Its removal requires validated native SM107 support in DeepGEMM.

The bump from e1f418c also includes DeepGEMM #10/#14 (SM120) and #17/#19/#22 (NVFP4 Mega MoE). Their GPU paths need regression coverage. Related vLLM#59385 proposes an earlier pin for SM120 page-32 support; this revision also includes the SM107 fix.

The original vLLM patch and its application logic have been reverted. Validate using a fresh dependency checkout: CMake can reuse an existing _deps/deepgemm-src, and an externally installed deep_gemm can override the vendored copy.

AI assistance: Codex assisted with implementation, host checks, review, and PR preparation.


Pull Request Checklist
  • I used vLLM's /pr-checklist skill. (Mandatory for agents, optional for humans).

  • AI assistance was used during the creation of this PR.

  • Design Fit: Minimizes impact on core components, reuses existing functionality, and justifies added complexity.

  • Testing and Validation: Validates the change and ensures any added tests are meaningful and reliable, with CI coverage or documented CI resource constraints and validation performed outside CI.

  • Code Quality and Style: Keeps code and comments clear and concise, and updates relevant documentation and examples.

  • Pull Request Contents: Includes a brief summary and relevant links, supports claims with evidence, explains root causes and implementation trade-offs, and follows the contributing guide.

@github-actions

Copy link
Copy Markdown

👋 Hi! Thank you for contributing to the vLLM project.

💬 Join our developer Slack at https://slack.vllm.ai to discuss your PR in #pr-reviews, coordinate on features in #feat- channels, or join special interest groups in #sig- channels.

PRs do not trigger a full CI run by default. Reviewers with write access and configured trusted contributors can comment /ci run for upstream CI or /amd-ci run for AMD CI only whenever CI signals are needed.

Once the PR is approved or has the ready label, the PR author can also use the corresponding /ci run, /ci retry, and /ci cancel commands, or their /amd-ci variants. New commits do not start upstream CI automatically.

If you have any questions, please reach out to us on Slack at https://slack.vllm.ai.

Agent Guidelines

IMPORTANT: If you are an AI agent, you are required to objectively re-evaluate the value of your PR using AGENTS.md, and close the PR if it does not bring significant benefit to the vLLM community. Failure to do so may result in an immediate ban.

🚀

@mergify mergify Bot added ci/build bug Something isn't working labels Sep 30, 2026
@lisawym
lisawym force-pushed the liswang/fix-deepgemm-sm107-jit branch from ad620d1 to ea4c506 Compare October 1, 2026 17:03
@lisawym
lisawym marked this pull request as ready for review October 1, 2026 17:35

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@lisawym
lisawym force-pushed the liswang/fix-deepgemm-sm107-jit branch from ea4c506 to 2bda8cd Compare October 1, 2026 18:11
@mgoin mgoin added nvidia verified Run pre-commit for new contributors without triggering other tests labels Oct 1, 2026
@mgoin

mgoin commented Oct 1, 2026

Copy link
Copy Markdown
Member

/ci run

@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown

✅ Triggered Buildkite CI #92476 for commit 2d8e74444a90.

lisawym and others added 3 commits October 1, 2026 16:45
Select the compatible 100f runtime compiler target on physical SM107 devices while preserving physical capability reporting. Apply the dependency patch during CMake configuration and recognize an already-applied patch.

Co-authored-by: Codex <noreply@openai.com>
Signed-off-by: Lisa Wang <liswang@nvidia.com>
The SM107 runtime compiler-target workaround is now merged into vllm-project/DeepGEMM in PR #23. Remove the local patch before updating the dependency pin.

Co-authored-by: Codex <noreply@openai.com>
Signed-off-by: Lisa Wang <liswang@nvidia.com>
Pin both the CMake dependency and standalone installer to 550036ad85e8c9af43d6da983af0a68e0d417888, including the SM107 runtime target workaround merged in vllm-project/DeepGEMM#23. The original vLLM patch was reverted before this update.

Co-authored-by: Codex <noreply@openai.com>
Signed-off-by: Lisa Wang <liswang@nvidia.com>
@lisawym
lisawym force-pushed the liswang/fix-deepgemm-sm107-jit branch from 2d8e744 to 6665c90 Compare October 1, 2026 23:51
@lisawym lisawym changed the title [Bugfix][Kernel] Fix DeepGEMM JIT target on SM107 (Rubin) [Bugfix][Kernel] Bump DeepGEMM pin for SM107 JIT compatibility Oct 1, 2026
Preserve the existing fork feature descriptions and document SM120 page_kv=32 support, NVFP4 Mega MoE updates, and the temporary SM107 JIT target workaround with its upstream references.

Co-authored-by: Codex <noreply@openai.com>
Signed-off-by: Lisa Wang <liswang@nvidia.com>
@lisawym

lisawym commented Oct 2, 2026

Copy link
Copy Markdown
Author
  • Reverted the old patch and rebased onto current upstream main.
  • Updated both DeepGEMM pins to 550036ad.

Comment thread tools/install_deepgemm.sh
# activations, plus the CUDA 12.x layout header fix from vllm-project/DeepGEMM#12,
# SM120 page_kv=32 support, NVFP4 Mega MoE updates, and the temporary SM107
# JIT target workaround from vllm-project/DeepGEMM#23 (see deepseek-ai/DeepGEMM#461).
DEEPGEMM_GIT_REF="550036ad85e8c9af43d6da983af0a68e0d417888"

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

not an issue with this PR but we should make this a single-source of truth so these don't need to stay in sync

@tlrmchlsmth tlrmchlsmth added the ready ONLY add when PR is ready to merge/full CI is needed label Oct 2, 2026
@tlrmchlsmth

Copy link
Copy Markdown
Member

/ci run

@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

✅ @lisawym, CI is now available for this PR.

  • /ci run starts upstream CI; /amd-ci run starts AMD CI only.
  • Your branch must contain every commit currently on its upstream target branch. Merge or rebase onto the latest target branch, then rerun the command. Append --allow-stale to a run command to test an outdated branch at your own risk.
  • /ci retry retries failed jobs in the CI build for the current PR head. If the current head has no CI build, it starts a new CI build for the current head containing only jobs that failed in the latest earlier CI build for this PR.
  • /amd-ci retry retries failed jobs in AMD CI for the current PR head. Use /amd-ci run when the current head has no AMD CI build.
  • /ci cancel cancels scheduled or running CI builds for this PR branch; /amd-ci cancel does the same for AMD CI only.

@tlrmchlsmth
tlrmchlsmth enabled auto-merge (squash) October 2, 2026 13:27
@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

✅ Triggered Buildkite CI #92627 for commit be2405a79716.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working ci/build nvidia ready ONLY add when PR is ready to merge/full CI is needed verified Run pre-commit for new contributors without triggering other tests

Projects

Status: Ready

Development

Successfully merging this pull request may close these issues.

3 participants