Skip to content

[Bugfix] Exclude SM107 (Rubin) from support_deep_gemm() - #58022

Closed
meena-at-work wants to merge 1 commit into
vllm-project:mainfrom
meena-at-work:fix/rubin-sm107-deep-gemm-gate
Closed

meena-at-work wants to merge 1 commit into
vllm-project:mainfrom
meena-at-work:fix/rubin-sm107-deep-gemm-gate

Conversation

@meena-at-work

Copy link
Copy Markdown
Contributor

Summary

support_deep_gemm() uses is_device_capability_family(100), which matches
any 10.x compute capability. That silently admits SM107 into the
DeepGEMM/CuTe-gated code paths alongside SM100/SM103 -- but the vendored
DeepGEMM/CuTe kernels aren't built for native sm_107, so any call behind
is_deep_gemm_supported() (e.g. the MHC TileLang prenorm-GEMM path) hits a
cuModuleLoadData -> CUDA_ERROR_ASSERT at kernel load time on that hardware.

This replaces the family-wide check with explicit is_device_capability(100)
/ is_device_capability(103) entries, matching how Hopper (90) and the
120 family are already enumerated individually in this same function. When
DeepGEMM ships kernels built for native sm_107, it can be added back the
same way -- this isn't a permanent hardware exclusion, just a "not built
yet" gap.

Why this isn't a duplicate

Searched open PRs referencing support_deep_gemm/DeepGEMM SM gating. The
closest is #53055 ("Guard DeepGEMM in mhc_pre_broadcast_tilelang with a
torch fallback"), which is a different bug: it adds a missing
is_deep_gemm_supported() check at one call site for platforms where
DeepGEMM isn't installed at all (SM121/GB10). This PR instead fixes the
shared platform-level gate itself, which currently wrongly reports support
on hardware where DeepGEMM is installed but not built for the native arch.

Test plan

Validated live on real SM107 hardware, using the public
vllm/vllm-openai:cu134-nightly image (no NVIDIA-internal patches):

>>> from vllm.platforms import current_platform
>>> current_platform.get_device_capability()
DeviceCapability(major=10, minor=7)
>>> current_platform.support_deep_gemm()
True      # before this change

After applying this diff to the installed vllm/platforms/cuda.py in the
same container:

>>> current_platform.support_deep_gemm()
False     # after this change

Confirms the gate now correctly excludes SM107 while leaving SM100/SM103/
Hopper/120-family behavior unchanged (traced statically; no regression
hardware available to re-verify those paths, but the change is additive/
narrowing only for SM107 and doesn't touch their branches).

AI assistance disclosure: this change was drafted with Claude Code
assistance; I reviewed the diff and reasoning end-to-end, and validated the
behavior change live on SM107 hardware myself before submitting.

is_device_capability_family(100) matches any 10.x capability, so it
silently admits SM107 (Rubin) into the DeepGEMM/CuTe-gated code paths
alongside SM100/SM103 (Blackwell). The vendored DeepGEMM/CuTe kernels
are not built for native sm_107, so this causes a
cuModuleLoadData -> CUDA_ERROR_ASSERT at kernel load time -- observed
via the MHC TileLang path (mhc_pre_big_fuse_with_norm_tilelang) on
SM107 hardware, but the same gate also feeds every other DeepGEMM
call site behind is_deep_gemm_supported().

Replace the family-wide check with explicit SM100/SM103 entries,
matching how Hopper (90) and the 120 family are already enumerated
individually in this function. When DeepGEMM ships kernels built for
native sm_107, add it back the same way.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

Signed-off-by: Meenakshi Venkataraman <meenakshiv@nvidia.com>

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@mergify mergify Bot added nvidia bug Something isn't working labels Sep 21, 2026
@meena-at-work

Copy link
Copy Markdown
Contributor Author

@tlrmchlsmth -- please review this minor SM107 related change, this affects GLM-5.3-flash.

@meena-at-work

Copy link
Copy Markdown
Contributor Author

@wangshangsam @xinli-sw please review and merge.

@meena-at-work

Copy link
Copy Markdown
Contributor Author

Closing this bugfix in favour of: #59503, since the current failure mode is only on the JIT path.

@github-project-automation github-project-automation Bot moved this to Done in NVIDIA Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working nvidia

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

2 participants