Skip to content

[None][feat] Update the logic of FMHA JIT path - #14291

Merged
lfr-0531 merged 2 commits into
NVIDIA:mainfrom
heyuhhh:user/yuhangh/update_JIT_filter
May 20, 2026
Merged

[None][feat] Update the logic of FMHA JIT path#14291
lfr-0531 merged 2 commits into
NVIDIA:mainfrom
heyuhhh:user/yuhangh/update_JIT_filter

Conversation

@heyuhhh

@heyuhhh heyuhhh commented May 19, 2026

Copy link
Copy Markdown
Collaborator

@coderabbitai summary

Description

In this PR:

  • Update all the cubins/headers/libs from trtllm-gen ToT main with full dynamic sparse kernels and zero sparse MQA/GQA cubins:

    • Elimate the cubins of sparse MQA/GQA because there is no model really uses it.
    • Use the cubins of dynamic sparse kernels but not NVRTC because there is a compiler issue in CUDA13.1.
  • Update the logic of shouldUseNvrtc

Test Coverage

PR Checklist

Please review the following before submitting your PR:

  • PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.

  • PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.

  • Test cases are provided for new code paths (see test instructions)

  • If PR introduces API changes, an appropriate PR label is added - either api-compatible or api-breaking. For api-breaking, include BREAKING in the PR title.

  • Any new dependencies have been scanned for license and vulnerabilities

  • CODEOWNERS updated if ownership changes

  • Documentation updated as needed

  • Update tava architecture diagram if there is a significant design change in PR.

  • The reviewers assigned automatically/manually are appropriate for the PR.

  • Please check this after reviewing the above items as appropriate for this PR.

GitHub Bot Help

To see a list of available CI bot commands, please comment /bot help.

@heyuhhh

heyuhhh commented May 19, 2026

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #49109 [ run ] triggered by Bot. Commit: dd43341 Link to invocation

heyuhhh added 2 commits May 19, 2026 11:58
Signed-off-by: yuhangh <58161490+heyuhhh@users.noreply.github.com>
Signed-off-by: yuhangh <58161490+heyuhhh@users.noreply.github.com>
@heyuhhh
heyuhhh force-pushed the user/yuhangh/update_JIT_filter branch from dd43341 to dd46263 Compare May 19, 2026 12:35
@heyuhhh

heyuhhh commented May 19, 2026

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #49189 [ run ] triggered by Bot. Commit: dd46263 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #49109 [ run ] completed with state ABORTED. Commit: dd43341

Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #49189 [ run ] completed with state FAILURE. Commit: dd46263
/LLM/main/L0_MergeRequest_PR pipeline #38867 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@heyuhhh

heyuhhh commented May 20, 2026

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #49301 [ run ] triggered by Bot. Commit: dd46263 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #49301 [ run ] completed with state FAILURE. Commit: dd46263
/LLM/main/L0_MergeRequest_PR pipeline #38963 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@heyuhhh

heyuhhh commented May 20, 2026

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #49338 [ run ] triggered by Bot. Commit: dd46263 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #49338 [ run ] completed with state FAILURE. Commit: dd46263
/LLM/main/L0_MergeRequest_PR pipeline #38996 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@heyuhhh

heyuhhh commented May 20, 2026

Copy link
Copy Markdown
Collaborator Author

/bot run

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #49384 [ run ] triggered by Bot. Commit: dd46263 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #49384 [ run ] completed with state SUCCESS. Commit: dd46263
/LLM/main/L0_MergeRequest_PR pipeline #39036 completed with status: 'SUCCESS'

CI Report

Link to invocation

@lfr-0531
lfr-0531 merged commit a173761 into NVIDIA:main May 20, 2026
7 checks passed
@github-actions

Copy link
Copy Markdown

LFS objects already in storage (2982 files) — no sync needed.

These LFS-tracked files are already present in this repository's LFS storage:

  • cpp/tensorrt_llm/kernels/trtllmGenKernels/fmha/cubin/FmhaSm100aKernel_QE4m3KvE2m1OE4m3H128PagedKvCausalP32VarSeqQ128Kv128PersistentContext.cubin.tar.zst
  • cpp/tensorrt_llm/kernels/trtllmGenKernels/fmha/cubin/FmhaSm100aKernel_QE4m3KvE2m1OE4m3H128PagedKvCausalP32VarSeqQ128Kv128StaticContext.cubin.tar.zst
  • cpp/tensorrt_llm/kernels/trtllmGenKernels/fmha/cubin/FmhaSm100aKernel_QE4m3KvE2m1OE4m3H128PagedKvCustomP32MultiCtasKvCgaVarSeqQ128Kv128StaticKeepsAbForGen.cubin.tar.zst
  • cpp/tensorrt_llm/kernels/trtllmGenKernels/fmha/cubin/FmhaSm100aKernel_QE4m3KvE2m1OE4m3H128PagedKvCustomP32MultiCtasKvCgaVarSeqQ16Kv128StaticSwapsAbForGen.cubin.tar.zst
  • cpp/tensorrt_llm/kernels/trtllmGenKernels/fmha/cubin/FmhaSm100aKernel_QE4m3KvE2m1OE4m3H128PagedKvCustomP32MultiCtasKvCgaVarSeqQ32Kv128StaticSwapsAbForGen.cubin.tar.zst
  • ...and 2977 more

xxi-nv pushed a commit to xxi-nv/TensorRT-LLM that referenced this pull request May 22, 2026
Signed-off-by: yuhangh <58161490+heyuhhh@users.noreply.github.com>
farazkh80 added a commit to farazkh80/TensorRT-LLM that referenced this pull request May 24, 2026
Rebases the TokenSpeed K2.6 evaluation onto upstream/main (PR NVIDIA#14291,
FMHA JIT namespace fix) and reports the resulting clean TS vs TRTLLM
A/B numbers. Pre-rebase, the TRTLLM baseline crashed during warmup
with NVRTC compilation failures on the fmhaSm103aKernel ...HQk576HV512
...ForGen family; the patch in PR NVIDIA#14291 (cutlass:: -> trtllm::dev::)
restores it.

Results (B300, BF16-KV patched K2.6):
- TP4 1k/1k conc=1: TRTLLM 158.5 / TS 152.8 tok/s (-3.6%)
- TP8 1k/1k conc=1: TRTLLM 182.2 / TS 169.3 tok/s (-7.1%)
- TP4 8k/1k conc=1: TRTLLM 152.0 / TS 146.9 tok/s (-3.4%)
- TP4 1k/1k conc=16: TRTLLM 1239.3 / TS 1246.8 tok/s (+0.6%, tied)

Files:
- new: phase4-rebased-ab.md, nvrtc-rebase-verify.md (verification +
  writeup), bench-config_base.yml (TRTLLM sidecar), scripts/run_bench_v3.sh
  (A/B driver targeting the rebased build)
- updated: phase4-summary.md (rebased context + NVBugs link),
  nvbug-draft-nvrtc-baseline.md (archived banner, fix landed upstream)
- removed: unused bench-60k1k_*.yml, bench-1k1k_tp4_conc16_attndp.yml,
  bench-1k1k_tp4_conc16_mtp3.yml, bench-8k1k_tp4_conc1_mtp3.yml and
  early prototype code/{tokenspeed_mla,test_tokenspeed_mla}.py
  (superseded by the in-tree backend at
  tensorrt_llm/_torch/attention_backend/tokenspeed_mla*.py)

Signed-off-by: Faraz Khoubsirat <58580514+farazkh80@users.noreply.github.com>
bmarimuthu-nv pushed a commit to nv-auto-deploy/TensorRT-LLM that referenced this pull request May 28, 2026
Signed-off-by: yuhangh <58161490+heyuhhh@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants