Skip to content

deep_gemm: arch-based split for DEEPGEMM_RUBIN cubins - #4261

Merged
jimmyzho merged 1 commit into
flashinfer-ai:release-v0.6.16from
jimmyzho:jimmzhou/deepgemm-rubin-arch-split
Jul 30, 2026
Merged

jimmyzho merged 1 commit into
flashinfer-ai:release-v0.6.16from
jimmyzho:jimmzhou/deepgemm-rubin-arch-split

Conversation

@jimmyzho

Copy link
Copy Markdown
Contributor

Summary

Follow-up to #4258 (which added ArtifactPath.DEEPGEMM_RUBIN and its checksum). That PR merged the path + checksum only; this PR adds the arch-based split that actually routes deep-gemm cubin loading to the Rubin directory on sm_107, mirroring the enable_rubin pattern from #4191.

flashinfer/deep_gemm.py

  • New is_rubin_arch() / get_deepgemm_artifact_path() helpers selecting DEEPGEMM_RUBIN on arch 107a and DEEPGEMM otherwise.
  • Used in load_all(), load(), and KernelMap.init_indices() so both the kernel cubins and kernel_map.json are fetched from the correct per-arch directory.
  • Added KernelMap.KERNEL_MAP_HASH_RUBIN = f8bf2b1b…a36e27 for the Rubin kernel_map.json (a separate manifest from the default one), selected by arch.

flashinfer/artifacts.py

  • Added DEEPGEMM_RUBIN to get_subdir_file_list()'s download list so flashinfer artifacts download fetches it, and updated the DEEPGEMM comment to reference the Rubin map hash.

Both Rubin hashes were verified by downloading the artifacts directly from the cubin repository.

Targets the release-v0.6.16 release branch.

🤖 Generated with Claude Code

Select the Rubin (sm_107 / cc 10.7) deep-gemm artifact directory at
runtime, mirroring the enable_rubin split added for the trtllm-gen
GEMM/BMM/MoE modules in PR flashinfer-ai#4191.

- Add is_rubin_arch() / get_deepgemm_artifact_path() helpers and use them
  in load_all(), load(), and KernelMap.init_indices() so cubins and
  kernel_map.json are fetched from DEEPGEMM_RUBIN on Rubin.
- Add KernelMap.KERNEL_MAP_HASH_RUBIN for the Rubin kernel_map.json.
- Register DEEPGEMM_RUBIN in get_subdir_file_list()'s download list.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@coderabbitai

coderabbitai Bot commented Jul 30, 2026

Copy link
Copy Markdown
Contributor

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: d2726c67-0adb-4cb8-b379-5271782b0c76

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@jimmyzho
jimmyzho merged commit d4c14f1 into flashinfer-ai:release-v0.6.16 Jul 30, 2026
4 checks passed
aleozlx added a commit that referenced this pull request Jul 30, 2026
<!-- .github/pull_request_template.md -->

## 📌 Description

This PR relands SM 107 support to main branch (reverted in #4171) as
well as some other release fixes.

#### Cherry Picks
- #4191
- #4189
- #4200
- #4215
- #4225
- #4230
- #4235
- #4226
- #4257
- #4258 
- #4261

#### Other Changes
- Rubin guards from #4252's conflict resolution (`TLLM_RUBIN_FEATURES`:
SiTuGlu
static_asserts + tile-192 advertisement, compiled out for the Rubin BMM
pin)
- Test-contract update: `test_unified_moe.py` arch assertions written
post-revert
(#4159) flipped to the restored contract (FP4/BF16 claim 107; FP8 stays
100/103)

<!-- What does this PR do? Briefly describe the changes and why they’re
needed. -->

## 🔍 Related Issues

<!-- Link any related issues here -->

#4107, #4164, reverts #4171

## 🚀 Pull Request Checklist

Thank you for contributing to FlashInfer! Before we review your pull
request, please make sure the following items are complete.

### ✅ Pre-commit Checks

- [x] I have installed `pre-commit` by running `pip install pre-commit`
(or used your preferred method).
- [x] I have installed the hooks with `pre-commit install`.
- [x] I have run the hooks manually with `pre-commit run --all-files`
and fixed any reported issues.

> If you are unsure about how to set up `pre-commit`, see [the
pre-commit documentation](https://pre-commit.com/).

## 🧪 Tests

- [ ] Tests have been added or updated as needed.
- [ ] All tests are passing (`unittest`, etc.).

## Reviewer Notes

<!-- Optional: anything you'd like reviewers to focus on, concerns, etc.
-->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added support for Rubin/SM107 GPUs across GEMM, MoE, attention,
quantization, sampling, and DeepGEMM workflows.
* Added architecture-aware kernel selection, memory sizing, compilation,
and artifact handling.
* **Bug Fixes**
* Improved validation and error messages for incompatible GPU
architectures and invalid kernel configurations.
  * Clearly rejects unsupported NVFP4 KV-cache operations on SM107.
* **Documentation**
  * Updated installation guidance with the SM107 architecture target.
* **Tests**
* Expanded architecture coverage and compatibility checks across GPU
test suites.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Vinnie6167 <Vinnie6167@users.noreply.github.com>
Co-authored-by: Ka-Hyun Nam <knam@nvidia.com>
Co-authored-by: Alex Yang <aleyang@nvidia.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Jimmy Zhou <79552142+jimmyzho@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant