Repository navigation
[Bugfix] Exclude SM107 (Rubin) from support_deep_gemm() - #58022
Closed
meena-at-work wants to merge 1 commit into
Closed
meena-at-work wants to merge 1 commit into
meena-at-work wants to merge 1 commit into
Conversation
is_device_capability_family(100) matches any 10.x capability, so it silently admits SM107 (Rubin) into the DeepGEMM/CuTe-gated code paths alongside SM100/SM103 (Blackwell). The vendored DeepGEMM/CuTe kernels are not built for native sm_107, so this causes a cuModuleLoadData -> CUDA_ERROR_ASSERT at kernel load time -- observed via the MHC TileLang path (mhc_pre_big_fuse_with_norm_tilelang) on SM107 hardware, but the same gate also feeds every other DeepGEMM call site behind is_deep_gemm_supported(). Replace the family-wide check with explicit SM100/SM103 entries, matching how Hopper (90) and the 120 family are already enumerated individually in this function. When DeepGEMM ships kernels built for native sm_107, add it back the same way. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Signed-off-by: Meenakshi Venkataraman <meenakshiv@nvidia.com>
Contributor
Author
|
@tlrmchlsmth -- please review this minor SM107 related change, this affects GLM-5.3-flash. |
Contributor
Author
|
@wangshangsam @xinli-sw please review and merge. |
wangshangsam
approved these changes
Sep 22, 2026
5 of 6 tasks
Contributor
Author
|
Closing this bugfix in favour of: #59503, since the current failure mode is only on the JIT path. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
support_deep_gemm()usesis_device_capability_family(100), which matchesany 10.x compute capability. That silently admits SM107 into the
DeepGEMM/CuTe-gated code paths alongside SM100/SM103 -- but the vendored
DeepGEMM/CuTe kernels aren't built for native
sm_107, so any call behindis_deep_gemm_supported()(e.g. the MHC TileLang prenorm-GEMM path) hits acuModuleLoadData -> CUDA_ERROR_ASSERTat kernel load time on that hardware.This replaces the family-wide check with explicit
is_device_capability(100)/
is_device_capability(103)entries, matching how Hopper (90) and the120 family are already enumerated individually in this same function. When
DeepGEMM ships kernels built for native
sm_107, it can be added back thesame way -- this isn't a permanent hardware exclusion, just a "not built
yet" gap.
Why this isn't a duplicate
Searched open PRs referencing
support_deep_gemm/DeepGEMM SM gating. Theclosest is #53055 ("Guard DeepGEMM in mhc_pre_broadcast_tilelang with a
torch fallback"), which is a different bug: it adds a missing
is_deep_gemm_supported()check at one call site for platforms whereDeepGEMM isn't installed at all (SM121/GB10). This PR instead fixes the
shared platform-level gate itself, which currently wrongly reports support
on hardware where DeepGEMM is installed but not built for the native arch.
Test plan
Validated live on real SM107 hardware, using the public
vllm/vllm-openai:cu134-nightlyimage (no NVIDIA-internal patches):After applying this diff to the installed
vllm/platforms/cuda.pyin thesame container:
Confirms the gate now correctly excludes SM107 while leaving SM100/SM103/
Hopper/120-family behavior unchanged (traced statically; no regression
hardware available to re-verify those paths, but the change is additive/
narrowing only for SM107 and doesn't touch their branches).
AI assistance disclosure: this change was drafted with Claude Code
assistance; I reviewed the diff and reasoning end-to-end, and validated the
behavior change live on SM107 hardware myself before submitting.