Skip to content
Closed
Show file tree
Hide file tree
Changes from 1 commit
Commits
Show all changes
204 commits
Select commit Hold shift + click to select a range
63aac74
[None][chore] Bump version to 1.2.0 (#10686)
yiqingy0 Jan 15, 2026
bc5094c
[None][chore] Setup the code review rule on release/1.2 branch (#10694)
yiqingy0 Jan 15, 2026
fe65847
[None][infra] Waive failed cases for release branch on 01/16 (#10748)
EmmaQiaoCh Jan 16, 2026
22b5f0c
[None][fix] Disable short profile for tunable ops with MERGE strategy…
hyukn Jan 16, 2026
06f7970
[https://nvbugs/5800521][ci] Move test_openai_chat_guided_decoding to…
syuoni Jan 16, 2026
3f98409
[None][doc] 1.2 Release Notes Headers (#10722)
pcastonguay Jan 16, 2026
e87a406
[https://nvbugs/5669671][fix] Support GuidedDecoder with sharded logi…
syuoni Jan 16, 2026
c772f3a
[https://nvbugs/5803813][fix] Fix llama 4 min latency (#10724)
mikeiovine Jan 16, 2026
18fe915
[TRTLLM-5366][chore] Add dgx-spark beta notes (#10766)
farazkh80 Jan 17, 2026
a48466f
[None][infra] Update upgrade related docs for release 1.2 (#10760)
EmmaQiaoCh Jan 17, 2026
6a8f18e
[None][infra] Waive failed cases for release on 10/18 (#10781)
EmmaQiaoCh Jan 18, 2026
368cb15
[None][doc] update doc (add minimax model) (#10749)
jmydurant Jan 19, 2026
83be9bb
[None][fix] Fix tmp dir being deleted too early in unit test. (#10741)
hyukn Jan 19, 2026
17f419f
[https://nvbugs/5811697][fix] Fix buffer reuse. (#10716)
yuxianq Jan 19, 2026
bc712a0
[https://nvbugs/5782112][fix] Cherry-pick #10633: Fix hanging issue f…
hyukn Jan 19, 2026
6185464
[None][infra] Waive failed case for release branch on 01/19 (#10795)
EmmaQiaoCh Jan 19, 2026
460b8a5
[None][test] modify ctx config in 128k8k disagg cases (#10779)
ruodil Jan 19, 2026
e5ade20
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jan 19, 2026
c51c75c
[None][test] Update test case for release (#10763)
crazydemo Jan 19, 2026
b979a02
[https://nvbugs/5748664][fix] Increasing disagg acc test timeout (#10…
pcastonguay Jan 19, 2026
6c0b080
[https://nvbugs/5791242][fix] workaround for flashinfer.sampling.samp…
ixlmar Jan 20, 2026
ed6df35
[https://nvbugs/5636916][fix] Fix accuracy issue of TWOSHOT AllReduce…
hyukn Jan 20, 2026
307ea14
[None][infra] Waive failed cases for release branch on 01/20 (#10828)
EmmaQiaoCh Jan 20, 2026
87ff718
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jan 20, 2026
89bb6f3
[None][chore] Reduce tedious logs (#10819)
chzblych Jan 20, 2026
837579c
[None][test] Update case for release (#10811)
crazydemo Jan 21, 2026
226274a
[None][fix] Fix the potential access issue of cache operations betwee…
hyukn Jan 21, 2026
4791ee7
[None][chore] Revert NVIDIA/TensorRT-LLM#10819 (#10870)
chzblych Jan 21, 2026
07a2878
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jan 21, 2026
7a3f264
[https://nvbugs/5814203][fix] Fix port 8000 being used issue in stres…
dominicshanshan Jan 21, 2026
8bf736a
[https://nvbugs/5747938][infra] Unwaive trtllm serve example test (#1…
LinPoly Jan 21, 2026
4e38b99
[https://nvbugs/5754977][fix] Use free port for serve test (#10878)
JunyiXu-nv Jan 21, 2026
0362a42
[https://nvbugs/5740377][fix] Prevent out-of-bounds read (#10868)
HuiGao-NV Jan 22, 2026
8ecfc29
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jan 22, 2026
42c2a41
[https://nvbugs/5769425][fix] add syncthreads for tinygemm to resolve…
dc3671 Jan 22, 2026
9488cbc
[https://nvbugs/5821433][fix] fix test_auto_scaling for 2 GPUs (#10866)
reasonsolo Jan 22, 2026
8cc475c
[https://nvbugs/5779536][fix] Unwaive Llama 3.3 related multi GPU tes…
pengbowang-nv Jan 22, 2026
7539452
[https://nvbugs/5826962][fix] Fix PD disaggregation for VLMs that use…
2ez4bz Jan 22, 2026
b7011a8
[https://nvbugs/5691730][fix] Have LoRa bf16 ckpts work with Llama 3.…
moraxu Jan 22, 2026
d36bd45
[https://nvbugs/5814409][fix] fix pp loop hang because of i-sending n…
reasonsolo Jan 22, 2026
40a8b2a
[None][fix] Always reset drafting states for GuidedDecoder (#10899)
syuoni Jan 22, 2026
6433764
[https://nvbugs/5814247][fix] AutoDeploy: skip mxfp4_moe test unless …
lucaslie Jan 22, 2026
13103af
[https://nvbugs/5769712][fix] fix timeout in AutoDeploy llama accurac…
lucaslie Jan 22, 2026
7f11072
[https://nvbugs/5784543][chore] unwaive test. (#10906)
yuxianq Jan 23, 2026
8cf294a
[https://nvbugs/5701445][chore] unwaive tests. (#10913)
yuxianq Jan 23, 2026
e1d61d1
[https://nvbugs/5748600][ci] Update guided decoding waive list (#10904)
syuoni Jan 23, 2026
00856de
[https://nvbugs/5819021][fix] unwaive some tests (#10836)
byshiue Jan 23, 2026
4689f83
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jan 23, 2026
27feb0c
[https://nvbugs/5680911][fix] Remove @cache decorator to enhance CI s…
zheyuf Jan 23, 2026
9205a4d
[https://nvbugs/5814309][fix] Use NCCL as fallback to avoid crash due…
hyukn Jan 23, 2026
ea4b8c9
[None][test] Update test list (#10883)
crazydemo Jan 23, 2026
133076a
[https://nvbugs/5833795][chore] Waive test test_e2e.py::test_ptp_quic…
yihwang-nv Jan 23, 2026
09e4a94
[https://nvbugs/5741304][chore] Update flashinfer-python to 0.6.1 (#1…
yihwang-nv Jan 23, 2026
cab453f
[https://nvbugs/5814914][fix] Fix llama sm120 spec dec (#10765)
mikeiovine Jan 23, 2026
571521b
[None][fix] Fix MTP 1-model sampler (#10369)
mikeiovine Jan 23, 2026
8159925
[None][ci] Remove long-running sanity check tests on GH200 (#10924)
chzblych Jan 24, 2026
02cb227
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jan 24, 2026
01ba1a1
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jan 24, 2026
fe27e9a
[https://nvbugs/5804146][fix] Enable responses tests and remove ds to…
JunyiXu-nv Jan 24, 2026
da7d1d3
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jan 25, 2026
6d6d0fb
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jan 25, 2026
829170a
[https://nvbugs/5779536][fix] Unwaive DeepSeekR1 nvfp4 pp4 mtp test c…
pengbowang-nv Jan 25, 2026
f048ade
[https://nvbugs/5829097][fix] Re-init TRTLLM sampler to use sample st…
yuxianq Jan 26, 2026
8f4b8f2
[https://nvbugs/5769890][fix] enable system memory to transfer active…
yuxianq Jan 26, 2026
453b7d1
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jan 26, 2026
11f2081
[None][feat] Cherry-pick #10335: Use XQA JIT impl by default and miti…
pengbowang-nv Jan 26, 2026
57a4fb9
[#10614][fix] gpt_oss first iteration streaming in trtllm-serve (#10884)
LinPoly Jan 26, 2026
4b319f0
[TRTLLM-9581][infra] Use /home/scratch.trt_llm_data_ci in computelab …
ZhanruiSunCh Jan 26, 2026
15cb430
[None][infra] Waive failed cases for release branch on 01/26 (#10999)
EmmaQiaoCh Jan 26, 2026
cfddd21
[https://nvbugs/5826689][fix] replace etcd3 with etcd-sdk-python (#10…
reasonsolo Jan 26, 2026
59c8225
[https://nvbugs/5769815][fix] Fix offset calculation in _are_stop_wor…
stnie Jan 26, 2026
fa0b317
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jan 27, 2026
0f0787b
[None][test] Fix missing test cases (#10881)
yufeiwu-nv Jan 27, 2026
bb56001
[None][feat] support Lyris GB200 and increase disagg test timeout (#1…
yingguo-trt Jan 27, 2026
8802eaf
[https://nvbugs/5800646][fix] Fix hang issue by avoid exposing UB buf…
liji-nv Jan 27, 2026
e2957d6
[None][chore] Unwaive helix tests (#11008)
brb-nv Jan 27, 2026
ce36c8a
[https://nvbugs/5835925][fix] Add EPD disagg support for Qwen3 VL MoE…
2ez4bz Jan 28, 2026
6dcac08
[https://nvbugs/5819452][ci] Unwaive LLaMA2 7B FP8 case (#10997)
syuoni Jan 28, 2026
b8318b2
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jan 28, 2026
96c0ae0
[https://nvbugs/5811087][chore] Unwaive Gemma3 27B multimodal test (#…
brb-nv Jan 28, 2026
219e5db
[https://nvbugs/5821433][fix] WAR for popen in QA env (#10989)
reasonsolo Jan 28, 2026
dde9545
[None][infra] Waive failed case for release on 1/28 (#11055)
EmmaQiaoCh Jan 28, 2026
4ebe56f
[https://nvbugs/5839569][test] update test constraint (#11054)
crazydemo Jan 28, 2026
18c4ff6
[None][infra] cherry pick lock file fix to release/1.2 (#10975)
yuanjingx87 Jan 28, 2026
bbc6462
[TRTLLM-10669][fix] Fix Eagle3 draft model weight loading for through…
cascade812 Jan 28, 2026
4f5c9ce
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jan 29, 2026
1bb0ead
[None][infra] Fix TRT-LLM data scratch mount point for gb10x (#10880)…
EmmaQiaoCh Jan 29, 2026
ffaa62b
[None][chore] unwaive qwen3 235B accuracy test (#11058)
kris1025 Jan 29, 2026
1519191
[None][doc] Hardware support update (#10719)
pcastonguay Jan 29, 2026
8e3d1be
[https://nvbugs/5829830][fix] Declare the var in the correct scope (#…
ziyixiong-nv Jan 30, 2026
569947a
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jan 30, 2026
6f7ff1b
[https://nvbugs/5815136][fix] Cherry-pick #11042: nccl symmetric with…
hyukn Jan 30, 2026
fe1ea30
[None][feat] Add documentation on configuring CPU affinity in TRT-LLM…
dhansen-nvidia Jan 30, 2026
3d3b30c
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jan 31, 2026
66c017f
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 1, 2026
8d6bb0a
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 2, 2026
0fa4d4e
[https://nvbugs/5823465][fix] Add CUTEDSL moe backend for deepseek r1…
dominicshanshan Feb 2, 2026
a459360
[https://nvbugs/5787904][fix] update mig tests (#11014)
xinhe-nv Feb 2, 2026
272fee9
[https://nvbugs/5819444][fix] Unwaive gpt-oss test (#10927)
LinPoly Feb 2, 2026
59e30a6
[None][infra] Waive failed cases for release branch on 02/02 (#11182)
EmmaQiaoCh Feb 2, 2026
343b608
[https://nvbugs/5739981][fix] unwaive tests using opt-125M (#11099)
ixlmar Feb 2, 2026
c4b2bbc
[None][chore] Add warning about 2-model MTP deprecation (#11043)
mikeiovine Feb 2, 2026
ed3b831
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 3, 2026
41656da
[https://nvbugs/5854419][fix] Fix Qwen3-VL-Dense/MoE accuracy drop (#…
yechank-nvidia Feb 3, 2026
dcff503
[https://nvbugs/5761391][fix] Cherry-pick #10471: Include triton-kern…
anish-shanbhag Feb 3, 2026
cf67667
[TRTLLM-10803][fix] Cherry-pick of #11200: Fix mocking of HuggingFace…
anish-shanbhag Feb 4, 2026
5a36856
[https://nvbugs/5815025][fix] Fix spec-dec mode flag and related cpp …
pengbowang-nv Feb 4, 2026
a3a293b
[TRTLLM-8425][doc] Update sampling documentation (#10083) (#11270)
stnie Feb 4, 2026
adb133b
[https://nvbugs/5821433][fix] complete WAR for popen in QA env (#11214)
crazydemo Feb 5, 2026
2ffc068
[None][chore] Pass without_comm to cutlass and deepgemm (#11245)
xxi-nv Feb 5, 2026
b60949f
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 5, 2026
7f4a4cd
[None][chore] Fix slurm job name (#11265)
yingguo-trt Feb 5, 2026
baa2abf
[https://nvbugs/5845769][fix] B300(sm103) support on VLMs (#11274)
yechank-nvidia Feb 5, 2026
46b890a
[https://nvbugs/5830877][fix] Use the best (correct) config for GPTOS…
dongfengy Feb 5, 2026
adb8bcb
[https://nvbugs/5826890][fix] Warm-up before disagg benchmarking (Che…
bo-nv Feb 5, 2026
1d875d8
[https://nvbugs/5863443][fix] Fix message truncation in Helix CP cach…
brb-nv Feb 5, 2026
8e01450
[https://nvbugs/5688721][fix] AutoDeploy: unwaive fixed NemotronH acc…
lucaslie Feb 5, 2026
b6f2582
[TRTLLM-10118][fix] Fix vulnerabilities urllib3 nbconvert jaraco-cont…
yiqingy0 Feb 6, 2026
60d53fb
[None][chore] Resolve a conflict in the md file (#11255)
ziyixiong-nv Feb 6, 2026
2027b55
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 7, 2026
a22f63a
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 8, 2026
82bfc5f
[https://nvbugs/5831976][chore] Move test_trtllm_flashinfer_symbol_co…
yihwang-nv Feb 10, 2026
78544d6
[https://nvbugs/5624818][fix] Fix GPT-OSS with non-paged_context_fmha…
pengbowang-nv Feb 10, 2026
daabeab
[https://nvbugs/5814504][fix] Add skip_pre_hopper flag on NVILA & Nan…
yechank-nvidia Feb 10, 2026
7167f2b
[None][infra] Disable release spark stage due to migration of spark c…
EmmaQiaoCh Feb 10, 2026
5038e96
[None][infra] Enable spark stage for release since the spark cloud mi…
EmmaQiaoCh Feb 10, 2026
57c1ecf
[https://nvbugs/5820922][perf] Improve TorchSampler performance by re…
stnie Feb 11, 2026
c679cb5
[None][infra] Pin the version for torchao (#11446)
EmmaQiaoCh Feb 11, 2026
d3b26dc
[https://nvbugs/5889564][fix] fix kwargs name (#11496)
reasonsolo Feb 13, 2026
3cddddd
[https://nvbugs/5833795][fix] Remove test waive and try CI (#11464)
dongfengy Feb 13, 2026
de294fc
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 14, 2026
3cc9e47
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 15, 2026
d1e5956
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 15, 2026
0149e89
[https://nvbugs/5860137][fix] Adjust deepgemm tuning buckets to cover…
dc3671 Feb 15, 2026
1c207c1
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 16, 2026
4bcc441
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 16, 2026
aa4c677
[https://nvbugs/5875296][fix] Fix TritonMOE test for Qwen3_30B_A3B (#…
dongfengy Feb 16, 2026
fbda477
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 17, 2026
cf1b00f
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 17, 2026
4a110dd
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 17, 2026
261627c
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 18, 2026
c426c49
[None][infra] Cherry pick plc pipeline for 1.2 (#11546)
yuanjingx87 Feb 18, 2026
82f6878
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 19, 2026
7d08591
[https://nvbugs/839137][fix] Unwaive disagg unexpected ucx error (#11…
pcastonguay Feb 19, 2026
5006b5f
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 20, 2026
5bd8661
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 21, 2026
1d229e7
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 21, 2026
73db5c0
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 22, 2026
8e9c39a
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 22, 2026
275f615
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 23, 2026
fe91481
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 23, 2026
c5127f6
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 23, 2026
4e96ed1
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 23, 2026
48c03d3
[https://nvbugs/5823783][fix] Fix multi-node trust_remote_code hang i…
JunyiXu-nv Feb 23, 2026
c42ba80
[None][infra] Waive failures on release 1.2 (#11639)
jieli-matrix Feb 23, 2026
4f6acbb
[None][chore] Fix gpu memory requirement in stress test (#11404)
dominicshanshan Feb 23, 2026
1510a14
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 24, 2026
c9a6df9
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 24, 2026
08318ab
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 25, 2026
172ab2a
[https://nvbugs/5839155][test] Unwaive DeepSeekR1 fp8_blockscale thro…
kaiyux Feb 26, 2026
ea5d0b5
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 26, 2026
00de565
[https://nvbugs/5809169][unwaive] Unwaive TestGPTOSS test (#11416)
peaceh-nv Feb 26, 2026
dd89617
[https://nvbugs/5859881][fix] Unwaive test (#11716)
hyukn Feb 26, 2026
6c5f1da
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 26, 2026
aa5e1d9
[https://nvbugs/5799917][fix] Recover from CUTLASS MoE doActivation p…
rosenrodt Feb 26, 2026
11dba1b
[None][feat] add sanity tests for release1.2 version (#11738)
yingguo-trt Feb 26, 2026
6225e34
[https://nvbugs/5889841][fix] Add custom option class to allow subcom…
FrankD412 Feb 26, 2026
95e3682
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 27, 2026
98c7afd
[None][chore]: Add waives for nvbug 5936273 and 5936322 (#11775)
jieli-matrix Feb 27, 2026
1648ce6
[https://nvbugs/5875522][docs] Add known issue for disaggregated serv…
Tabrizian Feb 27, 2026
76f011e
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Feb 28, 2026
c693d41
[https://nvbugs/5756028][fix] Fix VSWA initialization with spec-dec a…
cascade812 Feb 28, 2026
06e6ef6
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Mar 1, 2026
ae32812
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Mar 1, 2026
7a6d551
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Mar 2, 2026
db3af40
[https://nvbugs/5775256] [fix] Reopen fp8_dsl_fused_moe ut. (#11779)
limin2021 Mar 2, 2026
9b0c020
[TRTLLM-11135][fix] Fix vulnerabilities protobuf (#11702)
yiqingy0 Mar 2, 2026
a4ee00f
[https://nvbugs/5762822][chore] Unwaive longbenchV2 test (#11647)
heyuhhh Mar 2, 2026
6c542e9
[https://nvbugs/5823212][fix] Warmup maybe_compiled_cat in forward_co…
yuantailing Mar 2, 2026
76dd900
[https://nvbugs/5747920][bug] Cherry pick 11296 from main (#11771)
yechank-nvidia Mar 3, 2026
527ce26
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Mar 3, 2026
96c09fb
[TRTLLM-11176][fix] Security Issue Fix cherry pick (#11683)
yibinl-nvidia Mar 4, 2026
c8f3331
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Mar 4, 2026
3768b35
[TRTLLM-11135][fix] Fix vulnerability aiohttp (#11778)
yiqingy0 Mar 4, 2026
9c53176
[https://nvbugs/5936273][fix] Fix bugs of Mistral Large3 (#11885)
byshiue Mar 4, 2026
f06eaaa
[https://nvbugs/5949098][doc] Fixing docs links (#11912)
pcastonguay Mar 4, 2026
9f32f48
[None][doc] Replace the TensorRT-LLM with TensorRT LLM (#11914)
nv-guomingz Mar 5, 2026
3d97567
[None][chore] Fix/disagg perf failure detection (#11904)
yingguo-trt Mar 5, 2026
36e34ee
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Mar 5, 2026
14bf96b
[None][test] cherry-pick: add concurrency override and fix for 128k8k…
ruodil Mar 6, 2026
df5d831
[None][infra] Waive 3 failed cases for release/1.2 in post-merge 40 (…
ZhanruiSunCh Mar 6, 2026
e8b1956
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Mar 6, 2026
3429ad6
[None][infra] Waive 2 failed cases for release/1.2 in post-merge 42 (…
ZhanruiSunCh Mar 6, 2026
5e23368
[https://nvbugs/5948878][fix] Fix ClientPayloadError (#11973)
yingguo-trt Mar 6, 2026
8e77388
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Mar 7, 2026
e691be9
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Mar 8, 2026
a19f0f0
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Mar 9, 2026
02e5f84
[None][doc] Release notes for 1.2 release (#11955)
pcastonguay Mar 9, 2026
4988ba0
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Mar 9, 2026
d9cbb34
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Mar 10, 2026
c15f57d
[None][test] Fix disagg test sku for release 1.2 (#12066)
fredricz-20070104 Mar 10, 2026
484f4fe
[https://nvbugs/5924136][fix] Fix bug by add env var (#11974)
benzh-2025 Mar 10, 2026
51f5ef3
[None][test] Fix disagg test gpu (#12078)
fredricz-20070104 Mar 10, 2026
71acaf6
beam-aware logit_row tracking in _execute_logit_post_processors and b…
kyurious-george May 15, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Prev Previous commit
beam-aware logit_row tracking in _execute_logit_post_processors and b…
…uild coherent per-beam token histories by tracing parentIds
  • Loading branch information
kyurious-george committed May 15, 2026
commit 71acaf6e7a15681fc96d856735da4cb4fcc6e9ee
87 changes: 67 additions & 20 deletions tensorrt_llm/_torch/pyexecutor/model_engine.py
Original file line number Diff line number Diff line change
Expand Up @@ -3497,7 +3497,17 @@ def load_weights_from_target_model(self,
def _execute_logit_post_processors(self,
scheduled_requests: ScheduledRequests,
outputs: dict):
"""Apply logit post processors (in-place modify outputs Tensors) if any."""
"""Apply logit post processors (in-place modify outputs Tensors) if any.

For beam search (beam_width > 1), iterates over all beams for each
request and passes per-beam token histories to the callback.
The logits tensor has ``beam_width`` rows per generation request, laid
out as ``[ctx_0, ..., gen_0_beam_0, gen_0_beam_1, ..., gen_1_beam_0, ...]``.

On 1.2.x this must run BEFORE make_decoding_batch_input so the
decoder sees the modified logits (forward_async reads from
decoding_input, not decoder_input_buffers).
"""

if not (self.mapping.is_last_pp_rank()):
return
Expand All @@ -3509,29 +3519,66 @@ def _execute_logit_post_processors(self,
num_ctx_req = len(scheduled_requests.context_requests)
logits_tensor = outputs["logits"]

# Track running offset into logits_tensor because during beam search,
# generation requests contribute beam_width rows each while context
# requests contribute 1 row.
logit_row = 0
for idx, request in enumerate(scheduled_requests.all_requests()):
is_ctx = idx < num_ctx_req
beam_width = 1 if is_ctx else request.sampling_config.beam_width

logits_processors = getattr(request, "py_logits_post_processors",
None)
if not logits_processors:
logit_row += beam_width
continue

token_ids = request.get_tokens(0)
if idx < num_ctx_req and request.py_orig_prompt_len < len(
token_ids):
# Skip as we only need to apply logit processor on the last context request
continue

logits_row = logits_tensor[idx]
# Reshape to align w/ the shape used in the TRT backend,
# so the same logit processors can be used across both backends.
logits_row = logits_row.view(1, 1, -1)
token_ids = [token_ids]
for lp in logits_processors:
lp_params = inspect.signature(lp).parameters

assert 4 <= len(lp_params) <= 5, (
"Logit post processor signature must match the `LogitsProcessor` interface "
"defined in `tensorrtllm.sampling_params`.")
lp(request.py_request_id, logits_row, token_ids, None, None)
if is_ctx:
token_ids = request.get_tokens(0)
if request.py_orig_prompt_len < len(token_ids):
logit_row += 1
continue

logits_tensor[idx] = logits_row.view(-1)
if beam_width == 1:
# Single beam: original path.
token_ids = request.get_tokens(0)
logits_row = logits_tensor[logit_row]
logits_row = logits_row.view(1, 1, -1)
token_ids = [token_ids]
for lp in logits_processors:
lp_params = inspect.signature(lp).parameters
assert 4 <= len(lp_params) <= 5, (
"Logit post processor signature must match the `LogitsProcessor` interface "
"defined in `tensorrtllm.sampling_params`.")
lp(request.py_request_id, logits_row, token_ids, None, None)
logits_tensor[logit_row] = logits_row.view(-1)
logit_row += 1
else:
# Beam search: process all beams together.
vocab_size = logits_tensor.shape[-1]
logits_block = logits_tensor[logit_row:logit_row + beam_width]
# Shape: [1, beam_width, vocab] to match TRT backend convention.
logits_block_3d = logits_block.view(1, beam_width, vocab_size)

# Use coherent per-beam token histories if available.
# Built by PyExecutor._update_requests using parentIds
# tracing from the previous sampling step.
# Falls back to raw slot histories for the first step
# (before any beam reassignment has occurred).
coherent = getattr(request, 'py_coherent_beam_tokens', None)
if coherent is not None and len(coherent) == beam_width:
all_beam_token_ids = coherent
else:
all_beam_token_ids = [
request.get_tokens(beam) for beam in range(beam_width)
]

for lp in logits_processors:
lp_params = inspect.signature(lp).parameters
assert 4 <= len(lp_params) <= 5, (
"Logit post processor signature must match the `LogitsProcessor` interface "
"defined in `tensorrtllm.sampling_params`.")
lp(request.py_request_id, logits_block_3d,
all_beam_token_ids, None, None)
logits_tensor[logit_row:logit_row + beam_width] = logits_block_3d.view(beam_width, vocab_size)
logit_row += beam_width
45 changes: 45 additions & 0 deletions tensorrt_llm/_torch/pyexecutor/py_executor.py
Original file line number Diff line number Diff line change
Expand Up @@ -2448,6 +2448,51 @@ def _update_requests(self,
error_msg = str(e)
logger.error(f"Encountered an error in sampling: {error_msg}")
self._handle_errors(error_msg)
return

# Build coherent per-beam token histories for logits processors.
# request.get_tokens(beam) returns incoherent slot-accumulated
# histories after beam reassignment. We trace parentIds from
# DecoderState to recover the true ancestral token path per beam.
# These are stored on the request and used by model_engine's
# _execute_logit_post_processors on the NEXT iteration.
if isinstance(self.sampler, TRTLLMSampler):
try:
decoder_state = self.sampler.store["decoder_state"]
# parentIds shape: [maxNumSequences, maxBeamWidth, maxSequenceLength]
parent_ids_gpu = decoder_state.parent_ids
for req in (sample_state.scheduled_requests.generation_requests
if sample_state.scheduled_requests else []):
bw = req.sampling_config.beam_width
if bw <= 1 or req.py_seq_slot is None:
continue
slot = req.py_seq_slot
prompt_len = req.py_orig_prompt_len
tokens = [req.get_tokens(b) for b in range(bw)]
num_gen = len(tokens[0]) - prompt_len
if num_gen <= 1:
# First step — no beam reassignment yet.
continue
# D2H copy of this slot's parentIds: [beamWidth, maxSeqLen]
pid_host = parent_ids_gpu[slot, :bw, :].cpu()
# Trace backward per beam to build coherent histories.
coherent = []
for b in range(bw):
beam_gen = len(tokens[b]) - prompt_len
slot_at_step = [0] * beam_gen
s = b
for g in range(beam_gen - 1, -1, -1):
slot_at_step[g] = s
if g > 0:
s = int(pid_host[s, prompt_len + g].item())
path = list(tokens[0][:prompt_len]) # shared prompt
for g in range(beam_gen):
path.append(tokens[slot_at_step[g]][prompt_len + g])
coherent.append(path)
req.py_coherent_beam_tokens = coherent
except Exception:
# Non-fatal: fall back to incoherent histories.
pass

def _handle_errors(self,
error_msg: Optional[str] = None,
Expand Down