fix(trtllm): restrict routed-MoE backends to supported architectures (release port for #4107) - #4230
Conversation
…(release) Port of flashinfer-ai#4177 onto release-v0.6.16 for flashinfer-ai#4107. SM12x (Spark, RTX Pro 6000) was mis-dispatching TRTLLM routed-MoE cubins built for sm100f/sm103a and crashing at runtime. - Add isArchCompatible() to batched-GEMM and GEMM runners; reject unknown cubin families and guard Sm107a behind TLLM_RUBIN_FEATURES (Rubin pin only) - Tighten Trtllm*Config.supported() from arch >= 100 to explicit allowlists including sm107 on release - Gate integration tests on config_cls.supported(arch); add CPU contract tests Fixes flashinfer-ai#4107
|
Caution The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased. |
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Plus Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
<!-- .github/pull_request_template.md --> ## 📌 Description This PR relands SM 107 support to main branch (reverted in #4171) as well as some other release fixes. #### Cherry Picks - #4191 - #4189 - #4200 - #4215 - #4225 - #4230 - #4235 - #4226 - #4257 - #4258 - #4261 #### Other Changes - Rubin guards from #4252's conflict resolution (`TLLM_RUBIN_FEATURES`: SiTuGlu static_asserts + tile-192 advertisement, compiled out for the Rubin BMM pin) - Test-contract update: `test_unified_moe.py` arch assertions written post-revert (#4159) flipped to the restored contract (FP4/BF16 claim 107; FP8 stays 100/103) <!-- What does this PR do? Briefly describe the changes and why they’re needed. --> ## 🔍 Related Issues <!-- Link any related issues here --> #4107, #4164, reverts #4171 ## 🚀 Pull Request Checklist Thank you for contributing to FlashInfer! Before we review your pull request, please make sure the following items are complete. ### ✅ Pre-commit Checks - [x] I have installed `pre-commit` by running `pip install pre-commit` (or used your preferred method). - [x] I have installed the hooks with `pre-commit install`. - [x] I have run the hooks manually with `pre-commit run --all-files` and fixed any reported issues. > If you are unsure about how to set up `pre-commit`, see [the pre-commit documentation](https://pre-commit.com/). ## 🧪 Tests - [ ] Tests have been added or updated as needed. - [ ] All tests are passing (`unittest`, etc.). ## Reviewer Notes <!-- Optional: anything you'd like reviewers to focus on, concerns, etc. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added support for Rubin/SM107 GPUs across GEMM, MoE, attention, quantization, sampling, and DeepGEMM workflows. * Added architecture-aware kernel selection, memory sizing, compilation, and artifact handling. * **Bug Fixes** * Improved validation and error messages for incompatible GPU architectures and invalid kernel configurations. * Clearly rejects unsupported NVFP4 KV-cache operations on SM107. * **Documentation** * Updated installation guidance with the SM107 architecture target. * **Tests** * Expanded architecture coverage and compatibility checks across GPU test suites. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Vinnie6167 <Vinnie6167@users.noreply.github.com> Co-authored-by: Ka-Hyun Nam <knam@nvidia.com> Co-authored-by: Alex Yang <aleyang@nvidia.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Jimmy Zhou <79552142+jimmyzho@users.noreply.github.com>
|
#4177 started efforts to addresss the issue in restricting TRT-LLM routed-MoE and GEMM backends to supported architectures. |
Description
Release-specific port of #4177 onto
release-v0.6.16for #4107.On SM12x (Spark, RTX Pro 6000), TRTLLM routed-MoE backends were incorrectly
claiming support (
arch >= 100) and then dispatching sm100f/sm103a cubins,causing
RuntimeError: Error occurred when running GEMM!or segfaults intest_split_fused_moe_kernel_vs_reference.Changes
csrc/trtllm_batched_gemm_runner.cu: Replace per-SM if-chains withisArchCompatible(); reject unknown cubin families; guardSm107abehind#ifdef TLLM_RUBIN_FEATURES(only exists in the Rubin cubin pin's headers).csrc/trtllm_gemm_runner.cu: Same arch filter for the plain GEMM runner(previously had no arch filtering at all).
flashinfer/fused_moe/api.py: TightenTrtllm*Config.supported()fromarch >= 100to explicit allowlists_TRTLLM_ROUTED_ARCHS = (100, 103, 107)and
_TRTLLM_ROUTED_FP8_ARCHS = (100, 103).tests/moe_ep/test_split_fused_moe_kernel_vs_reference.py: Gate GPU testson
config_cls.supported(arch); add CPU contract tests + SM120 regression guard.Release-specific notes
main, whichwill drop 107 after the Adds SM107 support #4122 revert).
Sm107aenum case is gated behindTLLM_RUBIN_FEATURES, matching the existingrelease pattern from Adds SM107 support #4122/Split trtllm-gen cubin pins per arch #4191. Verified against the actual pinned headers:
the default BMM/GEMM pins do not define
Sm107a; only the Rubin pins do.Verification
test_split_fused_moe_kernel_vs_reference.py)isArchCompatible()builds cleanly against both default and Rubin BMMexport headers; ported batched runner compiles against default pin
Related
main, has merge conflicts)main: not yet mergedPre-existing issue (not in scope)
#4213on release referencesoptions.mDtypeSfC, which does not exist in thedefault BMM cubin pin's headers (only in the Rubin pin). This is a separate
release-only compile issue on the non-Rubin module, predating this port.