[None][perf] Custom decode kernels for MinimaxM3 - rebased - #17268
Merged
brb-nv merged 1 commit intoAug 5, 2026
Conversation
brb-nv
requested review from
PerkzZheng,
cascade812,
dc3671,
eopXD,
hyukn,
mikeiovine,
rosong11,
yizhang-nv and
yuxianq
August 4, 2026 23:48
brb-nv
requested review from
QiJune,
jieli-matrix,
schetlur-nv,
tburt-nv and
yuanjingx87
August 4, 2026 23:49
1 task
brb-nv
requested review from
pcicotti,
peihu-nv and
zheyuf
and removed request for
PerkzZheng,
hyukn,
jieli-matrix,
rosong11,
schetlur-nv,
tburt-nv,
yizhang-nv,
yuanjingx87 and
yuxianq
August 4, 2026 23:49
Collaborator
Author
|
/bot run --disable-fail-fast |
Collaborator
|
PR_Github #63868 [ run ] triggered by Bot. Commit: |
brb-nv
force-pushed
the
user/brb/port-vllm-kernels-rebase
branch
from
August 5, 2026 00:29
0332348 to
9fdbcfd
Compare
Signed-off-by: Balaram Buddharaju <169953907+brb-nv@users.noreply.github.com>
brb-nv
force-pushed
the
user/brb/port-vllm-kernels-rebase
branch
from
August 5, 2026 00:33
9fdbcfd to
09893da
Compare
Collaborator
|
PR_Github #63868 [ run ] completed with state |
1 task
1 task
brb-nv
added a commit
to brb-nv/TensorRT-LLM
that referenced
this pull request
Aug 20, 2026
…wiring The MiniMax-M3 decode path currently runs its generation rows through fmha_sm100, which schedules a generation row like a context row. Three kernels ported from vLLM replace that: a CuTe DSL indexer scoring kernel, a Triton sparse block decode kernel and a trtllm-gen dense decode kernel. This lands the kernels, their custom-op registration and their correctness tests. Nothing dispatches to them yet: the MSA backend, indexer and cache manager are untouched, so the decode path is byte-for-byte what it was and the kernels are reachable only from the tests and the microbenchmark. The dispatch is a follow-up, since it rests on the device-side length patching and the fused per-layer cache writes that are still landing on feat/m3_with_msa. Split out of NVIDIA#17268 on feat/m3_with_msa, which carries the same kernels plus that wiring. Two additions differ from it. msa_indexer gains only cutedsl_score_runner and _cutedsl_score, the self-contained entry points the scorer test drives, and not the run_indexer dispatch that calls them. The tests reach fmha_sm100 through a local _flat_page_table helper, because build_kv_page_indices does not take a block table until NVIDIA#16875; the helper feeds it the slot map that block table implies, so the A/B comparisons still run against the production page-table builder rather than a test-local copy. The CuTe DSL indexer decode kernel and its tests were originally contributed to vLLM by Thien Tran (vllm-project/vllm#48582), as were the CuTe utilities (vllm-project/vllm#43273). The Triton sparse decode kernel and its tests were originally contributed to vLLM by Kaichao You (vllm-project/vllm#45381). Thanks to both. No test-list change: l0_b300 already collects unittest/_torch/attention wholesale, so the three new files are picked up there, and each skips itself off SM100/SM103. (cherry picked from commit 727c683) Signed-off-by: Balaram Buddharaju <169953907+brb-nv@users.noreply.github.com>
brb-nv
added a commit
to brb-nv/TensorRT-LLM
that referenced
this pull request
Aug 20, 2026
…wiring The MiniMax-M3 decode path currently runs its generation rows through fmha_sm100, which schedules a generation row like a context row. Three kernels ported from vLLM replace that: a CuTe DSL indexer scoring kernel, a Triton sparse block decode kernel and a trtllm-gen dense decode kernel. This lands the kernels, their custom-op registration and their correctness tests. Nothing dispatches to them yet: the MSA backend, indexer and cache manager are untouched, so the decode path is byte-for-byte what it was and the kernels are reachable only from the tests and the microbenchmark. The dispatch is a follow-up, since it rests on the device-side length patching and the fused per-layer cache writes that are still landing on feat/m3_with_msa. Split out of NVIDIA#17268 on feat/m3_with_msa, which carries the same kernels plus that wiring. Two additions differ from it. msa_indexer gains only cutedsl_score_runner and _cutedsl_score, the self-contained entry points the scorer test drives, and not the run_indexer dispatch that calls them. The tests reach fmha_sm100 through a local _flat_page_table helper, because build_kv_page_indices does not take a block table until NVIDIA#16875; the helper feeds it the slot map that block table implies, so the A/B comparisons still run against the production page-table builder rather than a test-local copy. The CuTe DSL indexer decode kernel and its tests were originally contributed to vLLM by Thien Tran (vllm-project/vllm#48582), as were the CuTe utilities (vllm-project/vllm#43273). The Triton sparse decode kernel and its tests were originally contributed to vLLM by Kaichao You (vllm-project/vllm#45381). Thanks to both. No test-list change: l0_b300 already collects unittest/_torch/attention wholesale, so the three new files are picked up there, and each skips itself off SM100/SM103. (cherry picked from commit 727c683) Signed-off-by: Balaram Buddharaju <169953907+brb-nv@users.noreply.github.com>
brb-nv
added a commit
to brb-nv/TensorRT-LLM
that referenced
this pull request
Aug 21, 2026
…wiring The MiniMax-M3 decode path currently runs its generation rows through fmha_sm100, which schedules a generation row like a context row. Three kernels ported from vLLM replace that: a CuTe DSL indexer scoring kernel, a Triton sparse block decode kernel and a trtllm-gen dense decode kernel. This lands the kernels, their custom-op registration and their correctness tests. Nothing dispatches to them yet: the MSA backend, indexer and cache manager are untouched, so the decode path is byte-for-byte what it was and the kernels are reachable only from the tests and the microbenchmark. The dispatch is a follow-up, since it rests on the device-side length patching and the fused per-layer cache writes that are still landing on feat/m3_with_msa. Split out of NVIDIA#17268 on feat/m3_with_msa, which carries the same kernels plus that wiring. Two additions differ from it. msa_indexer gains only cutedsl_score_runner and _cutedsl_score, the self-contained entry points the scorer test drives, and not the run_indexer dispatch that calls them. The tests reach fmha_sm100 through a local _flat_page_table helper, because build_kv_page_indices does not take a block table until NVIDIA#16875; the helper feeds it the slot map that block table implies, so the A/B comparisons still run against the production page-table builder rather than a test-local copy. The CuTe DSL indexer decode kernel and its tests were originally contributed to vLLM by Thien Tran (vllm-project/vllm#48582), as were the CuTe utilities (vllm-project/vllm#43273). The Triton sparse decode kernel and its tests were originally contributed to vLLM by Kaichao You (vllm-project/vllm#45381). Thanks to both. No test-list change: l0_b300 already collects unittest/_torch/attention wholesale, so the three new files are picked up there, and each skips itself off SM100/SM103. (cherry picked from commit 727c683) Signed-off-by: Balaram Buddharaju <169953907+brb-nv@users.noreply.github.com>
brb-nv
added a commit
to brb-nv/TensorRT-LLM
that referenced
this pull request
Aug 24, 2026
…wiring The MiniMax-M3 decode path currently runs its generation rows through fmha_sm100, which schedules a generation row like a context row. Three kernels ported from vLLM replace that: a CuTe DSL indexer scoring kernel, a Triton sparse block decode kernel and a trtllm-gen dense decode kernel. This lands the kernels, their custom-op registration and their correctness tests. Nothing dispatches to them yet: the MSA backend, indexer and cache manager are untouched, so the decode path is byte-for-byte what it was and the kernels are reachable only from the tests and the microbenchmark. The dispatch is a follow-up, since it rests on the device-side length patching and the fused per-layer cache writes that are still landing on feat/m3_with_msa. Split out of NVIDIA#17268 on feat/m3_with_msa, which carries the same kernels plus that wiring. Two additions differ from it. msa_indexer gains only cutedsl_score_runner and _cutedsl_score, the self-contained entry points the scorer test drives, and not the run_indexer dispatch that calls them. The tests reach fmha_sm100 through a local _flat_page_table helper, because build_kv_page_indices does not take a block table until NVIDIA#16875; the helper feeds it the slot map that block table implies, so the A/B comparisons still run against the production page-table builder rather than a test-local copy. The CuTe DSL indexer decode kernel and its tests were originally contributed to vLLM by Thien Tran (vllm-project/vllm#48582), as were the CuTe utilities (vllm-project/vllm#43273). The Triton sparse decode kernel and its tests were originally contributed to vLLM by Kaichao You (vllm-project/vllm#45381). Thanks to both. No test-list change: l0_b300 already collects unittest/_torch/attention wholesale, so the three new files are picked up there, and each skips itself off SM100/SM103. (cherry picked from commit 727c683) Signed-off-by: Balaram Buddharaju <169953907+brb-nv@users.noreply.github.com>
brb-nv
added a commit
to brb-nv/TensorRT-LLM
that referenced
this pull request
Aug 24, 2026
…wiring The MiniMax-M3 decode path currently runs its generation rows through fmha_sm100, which schedules a generation row like a context row. Three kernels ported from vLLM replace that: a CuTe DSL indexer scoring kernel, a Triton sparse block decode kernel and a trtllm-gen dense decode kernel. This lands the kernels, their custom-op registration and their correctness tests. Nothing dispatches to them yet: the MSA backend, indexer and cache manager are untouched, so the decode path is byte-for-byte what it was and the kernels are reachable only from the tests and the microbenchmark. The dispatch is a follow-up, since it rests on the device-side length patching and the fused per-layer cache writes that are still landing on feat/m3_with_msa. Split out of NVIDIA#17268 on feat/m3_with_msa, which carries the same kernels plus that wiring. Two additions differ from it. msa_indexer gains only cutedsl_score_runner and _cutedsl_score, the self-contained entry points the scorer test drives, and not the run_indexer dispatch that calls them. The tests reach fmha_sm100 through a local _flat_page_table helper, because build_kv_page_indices does not take a block table until NVIDIA#16875; the helper feeds it the slot map that block table implies, so the A/B comparisons still run against the production page-table builder rather than a test-local copy. The CuTe DSL indexer decode kernel and its tests were originally contributed to vLLM by Thien Tran (vllm-project/vllm#48582), as were the CuTe utilities (vllm-project/vllm#43273). The Triton sparse decode kernel and its tests were originally contributed to vLLM by Kaichao You (vllm-project/vllm#45381). Thanks to both. No test-list change: l0_b300 already collects unittest/_torch/attention wholesale, so the three new files are picked up there, and each skips itself off SM100/SM103. (cherry picked from commit 727c683) Signed-off-by: Balaram Buddharaju <169953907+brb-nv@users.noreply.github.com>
1 task
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
This MR adds a few custom decode kernels for MinimaxM3 ported from vLLM.
[GDN] GDN Prefill kernel for SM100 vllm-project/vllm#43273
[M3] Improve indexer for long-context decode (sm100) vllm-project/vllm#48582
Special thanks to the original contributors!
Test Coverage
PR Checklist
Please review the following before submitting your PR:
PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.
PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.
Test cases are provided for new code paths (see test instructions)
If PR introduces API changes, an appropriate PR label is added - either
api-compatibleorapi-breaking. Forapi-breaking, includeBREAKINGin the PR title.Any new dependencies have been scanned for license and vulnerabilities
CODEOWNERS updated if ownership changes
Documentation updated as needed
Update tava architecture diagram if there is a significant design change in PR.
The reviewers assigned automatically/manually are appropriate for the PR.
Please check this after reviewing the above items as appropriate for this PR.
GitHub Bot Help
To see a list of available CI bot commands, please comment
/bot help.