Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
1129 commits
Select commit Hold shift + click to select a range
59d4369
[https://nvbugs/6115560][fix] catch OSError in config_file_lock for N…
sara4dev May 28, 2026
89c90eb
[None][fix] Fix OSRB source header and provenance issues in AutoDeplo…
bmarimuthu-nv May 28, 2026
fc3b69e
[https://nvbugs/6185173][fix] Set mamba ssm cache to fp32 for Nemotro…
tensorrt-cicd May 28, 2026
83ec591
[https://nvbugs/5800725][fix] Restore Mistral Large 3 text-only proce…
byshiue May 28, 2026
9feaa76
[None][test] Unwaive fp8 blockscale baseline mtp1 (#14666)
sunnyqgg May 28, 2026
82f69a0
[None][docs] fix incorrect auto sampler behavior description for beam…
fuergaosi233 May 28, 2026
fe07957
[None][feat] Expose host/GPU per-iter time and clarify iter labeling …
eopXD May 28, 2026
f391155
[None][chore] Add test lists with multi-gpu test to CI multi-gpu test…
pengbowang-nv May 28, 2026
09bad73
[None][refactor] Add derived properties for the thop.attention call s…
yuxianq May 28, 2026
e73d068
[https://nvbugs/6160085][fix] At `tensorrt_llm/tokenizer/tokenizer.py…
tensorrt-cicd May 28, 2026
89bcd80
[None][infra] Waive 1 failed cases for main in post-merge 2740 (#14688)
ZhanruiSunCh May 28, 2026
6484b71
[https://nvbugs/6221483][fix] Revert auto_deploy _mamba_ssm_prepare_m…
greg-kwasniewski1 May 28, 2026
82679b5
[TRTLLM-12762][fix] Enable multi-node TP for MiniMax-M2 (#14314)
pcicotti May 28, 2026
0432a81
[https://nvbugs/6043248][fix] Validate tensor payload size on deseria…
yibinl-nvidia May 28, 2026
56be625
[None][perf] Add AutoDeploy NVFP4 RMSNorm quant fusion (#14361)
tcherckez-nvidia May 28, 2026
13ca44a
[None][feat] Support Gemma4 multi-head_dim pools and host-side slicin…
eopXD May 28, 2026
0715e15
[TRTLLM-13960][test] Offline equivalence test for sharding IR (#13963)
greg-kwasniewski1 May 28, 2026
5c896cd
[TRTLLM-11410][feat] MoT World Model Support (#14012)
NVShreyas May 28, 2026
7443f1d
[https://nvbugs/6115036][fix] Fix NVFP4 engine size estimation and at…
hyukn May 28, 2026
50ca49f
[https://nvbugs/5972776][fix] Pass IPC HMAC key through file descript…
yibinl-nvidia May 28, 2026
e8a42a1
[https://nvbugs/5911594][fix] Restrict HTTP cluster storage to loopba…
yibinl-nvidia May 28, 2026
e173446
[None][fix] Exclude post-merge stages from CBTS force-keep filters (#…
achartier May 28, 2026
776bebc
[TRTLLMINF-67][infra] use pre-configured idle GPU exemption (#14587)
tburt-nv May 28, 2026
e96b710
[https://nvbugs/6207749][fix] Replace the spec with `onnx>=1.21.0` in…
tensorrt-cicd May 28, 2026
22c7956
[https://nvbugs/6185480][fix] Autodeploy skip the GLM accuracy test f…
nvchenghaoz May 28, 2026
da9db3a
[https://nvbugs/6165866][infra] Waive 1 failed cases for main in pre-…
taylor-yb-lee May 28, 2026
56df200
[https://nvbugs/6187185][fix] Apply the existing `low_memory_override…
tensorrt-cicd May 28, 2026
e6784d8
[https://nvbugs/6192201][fix] AutoDeploy: unwaive llama perf test and…
MrGeva May 28, 2026
f6ba936
[TRTLLM-12982][feat] improve attention backend selection (#14635)
ixlmar May 28, 2026
e35f0f4
[None][test] Enable test for kv_cache_manager_v2 for A10 (#12885)
lowsfer May 29, 2026
fed47f1
[None][infra] Generate json with cmake fetched contents in build stag…
yuanjingx87 May 29, 2026
1bb1a02
[TRTLLM-12436][feat] visual_gen: add CuTe DSL attention via exported …
xrq-phys May 29, 2026
a471435
[https://nvbugs/6229221][fix] Add a reasoning parser for qwen3_5 (#14…
moraxu May 29, 2026
8cdde83
[None][fix] Reuse batch_indices_cuda across CUDA graph captures in EA…
achartier May 29, 2026
dedd826
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd May 29, 2026
4826f6a
[https://nvbugs/6084447][fix] Fix MoE DeepGEMM workspace size with at…
tensorrt-cicd May 29, 2026
018c432
[https://nvbugs/6189416][fix] Add a Blackwell-specific reference entr…
tensorrt-cicd May 29, 2026
566c226
[#14619][perf] AutoDeploy: tune Llama-3.1-8B-Instruct-FP8 TP=2/4 conf…
MrGeva May 29, 2026
4504fd7
[#13561][feat] AutoDeploy: enable MLIR elementwise fusion and trtllm…
MrGeva May 29, 2026
8c830c9
[https://nvbugs/6221841][fix] Detect via the raw config_dict whether …
tensorrt-cicd May 29, 2026
7bdd835
[None][perf] Replace Parakeet audio encoder with native trtllm layers…
aswinvisva May 29, 2026
6b126ca
[https://nvbugs/6194552][fix] stabilize Triton Mamba softplus (#14652)
hnover-nv May 29, 2026
8d8a259
[TRTLLM-13043][chroe] add VisualGen context to AGENTS.md (#14732)
zhenhuaw-me May 29, 2026
5421ef9
[None][feature] Add thinking token budget control (#14665)
tijyojwad May 29, 2026
ecb1b44
[TRTLLM-13050][test] Remove two-model eagle3 spec-decoding tests (#14…
QiJune May 29, 2026
b1dfd30
[TRTLLM-12653][feat] LTX-2 Ulysses cross-attention for v2a with audio…
luyiyun1021 May 29, 2026
0a19205
[None][refactor] Flatten thop.attention sequence kwargs + rename rota…
yuxianq May 29, 2026
3f21a48
[https://nvbugs/6162857][fix] Use generation metrics for VisualGen pe…
taianz-nv May 29, 2026
3e80e33
[https://nvbugs/6045177][fix] resolve mypy error (#14689)
ixlmar May 29, 2026
c7683f2
[None][feat] add Poolside Laguna tool parser (#14638)
DomBrown May 29, 2026
3d56a4e
[TRTLLM-10004][chore] Enable NCCL symmetric zero-copy by default (#14…
nv-lschneider May 29, 2026
cf3c1d9
[None][infra] Fix cbts tokenmacro b64 (#14718)
crazydemo May 29, 2026
208ed9b
[TRTLLM-12982][perf] remove sync after FlashInfer attention plan() (#…
ixlmar May 29, 2026
27b4e56
[https://nvbugs/6156233][test] unwaive GPT-OSS dflash test since bug …
dongfengy May 29, 2026
29971b4
[TRTLLM-12901][fix] cap per-rank max_num_active_requests by max_num_t…
xwang233 May 29, 2026
59de15d
[None][feat] Enable NVFP4 KV cache support in trtllm-gen attention (#…
yihwang-nv May 29, 2026
027eb72
[https://nvbugs/6185480][fix] autodeploy unwaive the test (#14716)
nvchenghaoz May 29, 2026
6f2055f
[https://nvbugs/6136737][fix] Propagate external SWA window to FMHA k…
tensorrt-cicd May 29, 2026
91371dd
[https://nvbugs/5996024][fix] Enforce trust_remote_code flag (#13527)
yibinl-nvidia May 29, 2026
562d683
[https://nvbugs/5979710][fix] Bound transfer destinations (#13525)
yibinl-nvidia May 29, 2026
ebbbec4
[None][fix] Resolve NVML device index mismatch in get_numa_aware_cpu_…
YPxHolic May 29, 2026
74d7c3a
[None][infra] revert #13607 (#14757)
tburt-nv May 29, 2026
097dab1
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd May 30, 2026
1f02e96
[TRTLLM-12535][chore] Refactor fast path (token ID space) preprocessi…
moraxu May 30, 2026
15bb791
[None][infra] Waive 2 failed cases for main in post-merge 2741 (#14737)
ZhanruiSunCh May 30, 2026
c27da75
[None][infra] Waive 1 failed cases for main in pre-merge 40562 (#14776)
ZhanruiSunCh May 30, 2026
4ec2942
[TRTLLM-12440][feat] Add GMS-only weight sharing support (#13926)
chienchunhung May 30, 2026
32eda52
[None][test] Unwaive some Perf Tests (#14664)
chenfeiz0326 May 30, 2026
f20858c
[https://nvbugs/6204488][fix] Replace fixed disagg fill throttle with…
chienchunhung May 30, 2026
255a54b
[None][chore] Waive failing multi-gpu test (#14788)
brb-nv May 30, 2026
5f106df
[TRTLLM-11408][feat] Add VisualGen TP Support (#13614)
belgarten-nv May 30, 2026
26c099f
[None][test] Add TLLM_SPEC_DECODE_FORCE_NUM_ACCEPTED_TOKENS in Spec D…
chenfeiz0326 May 31, 2026
46bbdf5
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd May 31, 2026
0fb9a36
[https://nvbugs/6165866][infra] Waive 1 failed cases for main in pre-…
taylor-yb-lee May 31, 2026
54259ed
[https://nvbugs/6196391][fix] Carryover disagg TTFT improvements (#14…
brb-nv May 31, 2026
a422420
[https://nvbugs/6244695][fix] Revert Pass IPC HMAC key through file d…
chenfeiz0326 Jun 1, 2026
5f5b772
[None][test] Update datasets path (#14671)
JennyLiu-nv Jun 1, 2026
9ed1669
[None][infra] Update new .test_durations (#14661)
EmmaQiaoCh Jun 1, 2026
4e12ff7
[TRTLLM-13015][feat] drop complex visual_gen CLI example scripts (#14…
zhenhuaw-me Jun 1, 2026
f402178
[https://nvbugs/6117811][fix] Fix XQA IMA for invalid pages with slid…
pengbowang-nv Jun 1, 2026
6b36b0e
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 1, 2026
7a9c186
[None][feat] Tune mamba config by env variables (#14730)
Wanli-Jiang Jun 1, 2026
cde9963
[None][test] Update moe backend for ctx and acceptance length env (#1…
fredricz-20070104 Jun 1, 2026
35cce74
[None][test] Update precision of previous device step time (#14809)
fredricz-20070104 Jun 1, 2026
0b96e3a
[None][infra] Waive 12 failed cases for main in post-merge 2749 (#14802)
ZhanruiSunCh Jun 1, 2026
d2fd17b
[TRTLLM-12971][infra] Fix parse classname logic in timeout result (#1…
yiqingy0 Jun 1, 2026
3ec9e9b
[https://nvbugs/6038228][fix] Propagate event loop errors to await_re…
JunyiXu-nv Jun 1, 2026
71a188c
[TRTLLM-12288][feat] Support Nemotron-H nvfp4 ckpt on Hopper (#14775)
JadoTu Jun 1, 2026
70ab8fb
[TRTLLM-12596][feat] Support simple logprob format (#13972)
tongyuantongyu Jun 1, 2026
2e6f602
[None][fix] Stabilize Mamba replay state update (#14509)
sunnyqgg Jun 1, 2026
2f9b85a
[None][feat] Upgrade NIXL to v1.0.1 and UCX to 1.21 (#14436)
chuangz0 Jun 1, 2026
441eaae
[None][feat] Refactor DWDP from CUDA IPC to CUDA VMM + MNNVL composit…
tianyuz-nv Jun 1, 2026
06456e1
[TRTLLM-10947][perf] eagle3: use cudaMemcpy2DAsync custom op for hidd…
pcicotti Jun 1, 2026
059de9c
[None][fix] PyExecutor Hang in Disagg TP Prefill (#14020)
jthomson04 Jun 1, 2026
d5b19bd
[https://nvbugs/6240561][fix] Autodeploy fix the deepseek accuracy dr…
nvchenghaoz Jun 1, 2026
02a65b5
[#12702][feat] Autodeploy deprecate the legacy triton attention (#14194)
nvchenghaoz Jun 1, 2026
12ea8ef
[None][test] Waive 5 failed cases for main in QA CI (#14789)
tensorrt-cicd Jun 2, 2026
b9453fc
[None][test] Waive 7 failed cases for main in QA CI (#14791)
tensorrt-cicd Jun 2, 2026
c482b81
[https://nvbugs/6240561][fix] Fix AutoDeploy DeepSeek-R1 accuracy dro…
taylor-yb-lee Jun 2, 2026
6222112
[#14588][fix] [AutoDeploy] Fix OOM of DeepSeek-R1 NVFP4 for tp=4 (#14…
taylor-yb-lee Jun 2, 2026
5db9414
[https://nvbugs/6179761][fix] Save LTX-2 BF16 weights to speed up per…
yibinl-nvidia Jun 2, 2026
4ba59c0
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 2, 2026
c3f6d98
[TRTLLM-13028][doc] Add VisualGen API walkthrough example and docs pa…
zhenhuaw-me Jun 2, 2026
702e39d
[None][chore] Update flashinfer-python from 0.6.12rc2 to 0.6.12 (#14805)
yihwang-nv Jun 2, 2026
260b80f
[None][fix] AutoDeploy: Unwaive llmc standalone tests (#14700)
bmarimuthu-nv Jun 2, 2026
ca00411
[TRTLLM-35882][feat] Add cute dsl gvr top-k decode kernel (#14602)
limin2021 Jun 2, 2026
f57f4aa
[https://nvbugs/6222480][test] fix stress test issue on H100 (#14721)
xinhe-nv Jun 2, 2026
06e3a77
[None][test] Waive 6 failed cases for main in QA CI (#14787)
tensorrt-cicd Jun 2, 2026
125c0da
[None][test] Waive 1 failed cases for main in QA CI (#14783)
tensorrt-cicd Jun 2, 2026
18724f7
[None][fix] synchronize MLA cache reuse fallback metadata (#14049)
DhineshPonnarasan Jun 2, 2026
460adc7
[None][feat] Add KV cache prefetch (#14748)
lowsfer Jun 2, 2026
209f371
[https://nvbugs/6191524][fix] In MLA.forward_context, also call the w…
tensorrt-cicd Jun 2, 2026
a2996ae
[None][test] Waive 2 failed cases for main in QA CI (#14839)
tensorrt-cicd Jun 2, 2026
efb71c7
[None][fix] Cherry-pick kv_cache_manager_v2 fixes to main (#14725)
lowsfer Jun 2, 2026
4cc2d8a
[None][test] Waive 11 failed cases for main in post-merge (#14854)
tensorrt-cicd Jun 2, 2026
33e0ee3
[None][feat] Enable flashifner gdn decoding kernel for qwen3.5 (#13645)
nv-guomingz Jun 2, 2026
e58e758
[https://nvbugs/5940460][fix] Harden FP8 quant fusion matching after …
pcicotti Jun 2, 2026
6ef7b38
[https://nvbugs/6221450][fix] AutoDeploy: Qwen3.5 400B NVFP4 accuracy…
taylor-yb-lee Jun 2, 2026
66262cf
[TRTLLM-12648][test] implement disagg cancel stress metrics_thread (#…
chienchunhung Jun 2, 2026
9653450
[None][chore] Update AD model list (#14686)
tcherckez-nvidia Jun 2, 2026
cd38dfb
[https://nvbugs/6226933][fix] canonicalize multimodal cache-key seria…
venkywonka Jun 2, 2026
178c8f4
[https://nvbugs/6240561][fix] Unwaive DeepSeek R1 accuracy test (#14870)
taylor-yb-lee Jun 2, 2026
180eedb
[None][feat] Add Qwen image support (#13449)
pst2154 Jun 3, 2026
6824bd8
[TRTLLM-12507][feat] Per-expert lora support with Cutlass backend (#1…
brb-nv Jun 3, 2026
74ec2a1
[None][chore] Make submit.py can run single GPU test and accept custo…
HuiGao-NV Jun 3, 2026
e1ed5e0
[None][test] Waive 9 failed cases for main in QA CI (#14792)
tensorrt-cicd Jun 3, 2026
aa4276d
[None][test] Update DSV32 32k4k config to avoid timeout issue (#14856)
chenfeiz0326 Jun 3, 2026
bcdf418
[None][chore] Bump version to 1.3.0rc18 (#14872)
yuanjingx87 Jun 3, 2026
66577ac
[None][infra] Waive 5 failed cases for main in post-merge 2755 (#14883)
ZhanruiSunCh Jun 3, 2026
06388ec
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 3, 2026
328ef0b
[None][fix] LTX-2 audio PE pad: use token-axis seq_dim=1 for token-ma…
luyiyun1021 Jun 3, 2026
c798fd9
[#13082][fix] Fix-multimodal embedding mismatch (#13240)
aashirvad08 Jun 3, 2026
e94830c
[None][fix] Pipe stderr separately in subprocess calls to improve err…
yufeiwu-nv Jun 3, 2026
f7f92f4
[None][fix] Use renamed get_param_count_and_checkpoint_size in hybrid…
yufeiwu-nv Jun 3, 2026
8edd72e
[None][test] Remove duplicate test cases in llm_perf_core file (#14749)
yufeiwu-nv Jun 3, 2026
d0cfcde
[None][test] Remove 28 closed-bug waive entries for main (#14545)
tensorrt-cicd Jun 3, 2026
514afc8
[TRTLLM-13022][test] remove deprecated models from tests (#14660)
xinhe-nv Jun 3, 2026
17ccf33
[None][feat] Reserve one more slots for attention_dp in mixed mamba c…
Wanli-Jiang Jun 3, 2026
c938efa
[https://nvbugs/6195110][fix] Restore DeepSeek shared-weights vanilla…
zhaoyangwang-nvidia Jun 3, 2026
a24c3d4
[#12359][feat] AutoDeploy: Support SSM replay kernel for MTP with Fla…
galagam Jun 3, 2026
abc6ba2
[None][test] Waive 1 failed cases for main in QA CI (#14857)
tensorrt-cicd Jun 3, 2026
af9568a
[None][chore] add attention module owner for VisualGen (#14814)
zhenhuaw-me Jun 3, 2026
a336495
[None][fix] release v1 KV blocks on MAX_UTILIZATION pause (#14723)
eopXD Jun 3, 2026
6ab5005
[None][perf] Reduce OpenAI stream postprocess overhead (#14708)
2ez4bz Jun 3, 2026
3959914
[None][fix] propagate chat prompt token ids (#14420) (#14859)
reasonsolo Jun 3, 2026
3630e16
[https://nvbugs/6211193][fix] etcd listen all interfaces (#14863)
reasonsolo Jun 3, 2026
7e8082e
[#5247][fix] auto-detect local cnn_dailymail dataset by directory lay…
guan404ming Jun 3, 2026
e5b8094
[https://nvbugs/6248987][fix] Made the slow-tokenizer swap lazy and i…
tensorrt-cicd Jun 3, 2026
a163d74
[None][chore] redact internal NVIDIA URLs from exec-slurm-compile ski…
ssam18 Jun 3, 2026
b2bb0ad
[None][feat] VisualGen: Attention2D + Ulysses & Multi-GPU LPIPS Evals…
juney-nvidia Jun 3, 2026
fbf66e9
[TRTLLM-13077][feat] Decompose post_load_weights() (#14770)
chienchunhung Jun 3, 2026
7374d1f
[None][fix] Fix config sharing issue for Qwen3-VL (#14766)
2ez4bz Jun 3, 2026
5f64e7d
[https://nvbugs/6104831][fix] Enforce request and buffer index lifecy…
chienchunhung Jun 3, 2026
c16ce24
[None][feat] Add encoder CUDA graph support to llm.encode() (#14326)
tingyangk Jun 4, 2026
6dc60cb
[None][feat] Support Step-3.7-Flash model (#14711)
kaiyux Jun 4, 2026
d64b217
[https://nvbugs/6050489][chore] unwaive tests (#14866)
bo-nv Jun 4, 2026
e1212ad
[None][test] Waive 1 failed cases for main in QA CI (#14896)
tensorrt-cicd Jun 4, 2026
86d08f4
[None][infra] Source code and container vulnerability fix (#14025)
yuanjingx87 Jun 4, 2026
33efef2
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 4, 2026
846d0f4
[None][infra] Waive 11 failed cases for main in post-merge 2757 (#14925)
ZhanruiSunCh Jun 4, 2026
b7ca2a6
[None][test] update rtx6k test list (#14929)
xinhe-nv Jun 4, 2026
00187c0
[None][fix] Update dataset identifier for cnn_dailymail to use namesp…
yufeiwu-nv Jun 4, 2026
864240b
[https://nvbugs/5979673][fix] Unwaive test_agent_multi_backends.py::t…
Shixiaowei02 Jun 4, 2026
4437cb9
[None][fix] Add nemotron-v3 as the proper nemotron-h reasoning parser…
Wanli-Jiang Jun 4, 2026
27af2e5
[https://nvbugs/6193836][test] Use EP=8 + attention DP for minimax_m2…
ruodil Jun 4, 2026
023ac82
[TRTLLM-8236][infra] fix platform tag for public wheel (#14616)
niukuo Jun 4, 2026
0f5dc5a
[None][test] update bug ids in waives (#14946)
xinhe-nv Jun 4, 2026
8b0eba9
[https://nvbugs/6244474][fix] AutoDeploy: skip explicit shape-prop af…
tensorrt-cicd Jun 4, 2026
8c39de8
[None][infra] fix cbts json decode (#14928)
crazydemo Jun 4, 2026
2a934fc
[https://nvbugs/6222480][fix] Fix stress (#14949)
xinhe-nv Jun 4, 2026
941c778
[None][test] Decrease P1 models number and merge sanity test list int…
yufeiwu-nv Jun 4, 2026
8361d42
[None][perf] Use a Triton kernel for Cpp mamba hybrid state update (#…
VALLIS-NERIA Jun 4, 2026
c17611c
[None][chore] Autodeploy unwaive 5888827, 6200112 (#14894)
galagam Jun 4, 2026
222d9e8
[NVBUG-6248780][fix] Add --decoupled flag to benchmark_core_model in …
karljang Jun 4, 2026
0718049
[TRTLLM-12870][feat] Support num_images_per_prompt for FLUX pipelines…
karljang Jun 4, 2026
33b0a32
[TRTLLM-12214][perf] DeepGemmFusedMoE: fuse masked gather + finalize-…
xwang233 Jun 4, 2026
a8c4007
[None][fix] Fix AutoDeploy accuracy tests (#13925)
bmarimuthu-nv Jun 4, 2026
910826b
[TRTLLMINF-69][infra] Migrate A100X-FMHA-Post-Merge-1 and A100X-Trito…
mlefeb01 Jun 4, 2026
8e5d9e2
[TRTLLM-11508][refactor] Merge Eagle3 and MTP-eagle one-model workers…
zhaoyangwang-nvidia Jun 4, 2026
a50b5e2
[https://nvbugs/6143787][fix] Add `kv_cache_config = KvCacheConfig(fr…
tensorrt-cicd Jun 4, 2026
df2d5b9
[https://nvbugs/6248764][fix] Normalize non-sliding KV windows to ful…
eopXD Jun 5, 2026
bd17d1b
[https://nvbugs/6240420][fix] Clamp KV pool window sizes to max_seq_l…
eopXD Jun 5, 2026
81e86a5
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 5, 2026
45d4d54
[None][test] Fix the ci disagg perf local submit test scope too large…
fredricz-20070104 Jun 5, 2026
f0ca418
[None][test] remove outdated model in perf test (#14992)
ruodil Jun 5, 2026
4574851
[TRTLLM-12648][test] implement disagg cancellation injector thread (#…
chienchunhung Jun 5, 2026
21ffdc7
[None][feat] Add AutoDeploy support for StepFun Step-3.7-Flash (#14759)
bmarimuthu-nv Jun 5, 2026
316430f
[None] [waive] Waive the failed step3p7 test case due to ckpt update …
kaiyux Jun 5, 2026
d5de55e
[https://nvbugs/6210714][fix] Fix mamba block calculation (#14524)
VALLIS-NERIA Jun 5, 2026
6818233
[https://nvbugs/5546507][https://nvbugs/5612313][test] Remove obsolet…
xinhe-nv Jun 5, 2026
fdcdcb3
[None][fix] AutoDeploy: Move hf_id_to_local_model_dir to function for…
bmarimuthu-nv Jun 5, 2026
6387eac
[TRTLLM-12893][infra] Parallelize post stages: Rerun Report, Test Cov…
ZhanruiSunCh Jun 5, 2026
2336e47
[None][infra] Waive 11 failed cases for main in post-merge 2760 (#15003)
ZhanruiSunCh Jun 5, 2026
58da60a
[None][fix] Uncomment Qwen3.5 and DSR1 from model registry so that th…
taylor-yb-lee Jun 5, 2026
73c824c
[TRTLLM-11410][feat] Cosmos3 Support (#14824)
NVShreyas Jun 5, 2026
fb5bd44
[https://nvbugs/5859886][fix] Remove the waiver (#14948)
ziyixiong-nv Jun 5, 2026
501b5c2
[https://nvbugs/6248744][fix] Added `trust_remote_code=True` to the `…
tensorrt-cicd Jun 5, 2026
37ece3f
[https://nvbugs/6160629][fix] AutoDeploy: Fix manual seed setting for…
galagam Jun 5, 2026
86f9602
[TRTLLM-12714][feat] KVCacheManagerV2: wire PyExecutor rebalance hook…
thorjohnsen Jun 5, 2026
3e17560
[None][feat] add Wan I2V generation example (#14981)
o-stoner Jun 5, 2026
52ba2bb
[TRTLLM-12527][feat] Parallelize multi-shard visual-gen checkpoint lo…
yibinl-nvidia Jun 5, 2026
d639c57
[https://nvbugs/6250866][fix] Fix deep ep partial warp sync for gptos…
dongfengy Jun 6, 2026
3b21093
[https://nvbugs/6272668][infra] Unwaive DSR1 and Qwen3.5 again (#15010)
taylor-yb-lee Jun 6, 2026
d7a5872
[None][feat] Afmoe trinity support (#13148)
alyosha-swamy Jun 6, 2026
4279e5b
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 6, 2026
e47f26e
[TRTLLM-13027][ci] Relocate under-using tests to right-sized stages (…
QiJune Jun 6, 2026
520262d
[None][feat] Add LTX-2 visual generation example (#14976)
yibinl-nvidia Jun 6, 2026
ec6b284
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 7, 2026
bedad85
[None][feat] AutoDeploy: Fix hardcoded configs (#14943)
taylor-yb-lee Jun 7, 2026
47666de
[#13718][feat] AutoDeploy MoE all-to-all: cache + runtime max-tokens …
greg-kwasniewski1 Jun 7, 2026
428cc3e
[TRTLLM-13177][doc] Add Nemotron 3 Ultra doc (#14964)
nv-guomingz Jun 7, 2026
dcd4e90
[#10710][feat] Make explicit CLI flags take precedence over --config …
marinayanov Jun 7, 2026
8be182d
[https://nvbugs/6260907][fix] unwaive test (#15058)
bo-nv Jun 8, 2026
b8d17d7
[None][chore] Increase GB200-4_GPUs-PyTorch shards (#14836)
tburt-nv Jun 8, 2026
71debd5
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 8, 2026
0e0ee27
[TRTLLM-12648][test] implement disagg cancellation canary thread (#15…
chienchunhung Jun 8, 2026
98a88f7
[TRTLLM-12507][feat] Cudagraph support for routed-expert MoE LoRA wit…
brb-nv Jun 8, 2026
86f33e6
[https://nvbugs/6245317][test] set Harmony tiktoken env for GPT-OSS d…
dongfengy Jun 8, 2026
b4d44d3
[https://nvbugs/6153955][test] unwaive GPT-OSS w4 DP4 CUTLASS (#14884)
dongfengy Jun 8, 2026
ca2bc5e
[None][perf] kv_cache_manager_v2: batch block-key SHA-256 hashing (#1…
lancelly Jun 8, 2026
2cad6db
[TRTLLM-13259][ci] Merge DGX_H100 DeepSeek and GptOss stages (#15035)
QiJune Jun 8, 2026
5fa68a4
[None][infra] Waive 11 failed cases for main in post-merge 2765 (#15080)
ZhanruiSunCh Jun 8, 2026
2632530
[None][infra] Waive 3 failed cases for main in post-merge 2765 (#15082)
ZhanruiSunCh Jun 8, 2026
7e49baa
[None][test] waive weekly qa ci failure cases (#15077)
crazydemo Jun 8, 2026
02f6b2f
[None][feat] AutoDeploy: propagate layer_type hint across pattern-mat…
greg-kwasniewski1 Jun 8, 2026
28dc25e
[None][test] Waive 15 failed cases for main in QA CI (#15056)
tensorrt-cicd Jun 8, 2026
2febb37
[None][infra] Waive 1 failed cases for main in pre-merge 41894 (#15089)
ZhanruiSunCh Jun 8, 2026
09c21b6
[TRTLLM-13262][ci] Move non-default-feature tests to post merge (#15038)
QiJune Jun 8, 2026
c93c63d
[None][feat] Enable disk cache config for KV cache v2 (#14845)
reasonsolo Jun 8, 2026
6dee167
[https://nvbugs/6185446][fix] Add warmup for trtllm-gen fmha JIT kern…
pengbowang-nv Jun 8, 2026
b14794c
[https://nvbugs/6162940][chore] Unwaive fixed test (#15078)
longlee0622 Jun 8, 2026
2bf4d3d
[None][perf] Support Gemma RMSNorm + interleaved mRoPE in fused_qk_no…
nv-guomingz Jun 8, 2026
9eaa468
[None][test] Half K25 Agg Multi Round to Solve Timeout Issue (#15083)
chenfeiz0326 Jun 8, 2026
9af8a16
[None][infra] Reduce Docker image layer count in release stage (#14972)
tburt-nv Jun 8, 2026
cb01607
[#14828][feat] AutoDeploy: support multi KV cache memory pool in trtl…
MrGeva Jun 8, 2026
15d06c0
[None][doc] Refine Nemotron Ultra doc (#15113)
nv-guomingz Jun 8, 2026
8036cde
[None][infra] Waive TestQwen3NextInstruct nvfp4 cases (#15086)
mzweilz Jun 8, 2026
1998324
[https://nvbugs/6248757][fix] Avoid running all reduce in aux stream …
tensorrt-cicd Jun 8, 2026
900d069
[https://nvbugs/6221483][fix] AutoDeploy: Fix Eagle metadata host syn…
govind-ramnarayan Jun 8, 2026
9827c21
[None][feat] add FLUX visual generation examples (#14987)
karljang Jun 8, 2026
b222246
[https://nvbugs/6261164][fix] In the kvcache insert transform (`_Inse…
tensorrt-cicd Jun 8, 2026
c1e9b00
[https://nvbugs/6211189][fix] Lower the reference to 46.5 (matching c…
tensorrt-cicd Jun 9, 2026
bfb4537
[None][refactor] split VisualGen pipeline and model configs (#14956)
bobboli Jun 9, 2026
5e3af40
[TRTLLM-11457][feat] Async Ulysses pipeline (Enabled for LTX-2 + WAN)…
luyiyun1021 Jun 9, 2026
09ebc59
[TRTLLM-11548][doc] Add Qwen3.5 deployment guide doc (#15111)
nv-guomingz Jun 9, 2026
a33dec7
[https://nvbugs/6181383][fix] Build inner text/vision/audio sub-confi…
tensorrt-cicd Jun 9, 2026
041ed83
[None][infra] Check in most recent lock file from nightly pipeline
tensorrt-cicd Jun 9, 2026
2490441
[https://nvbugs/6273850][chore] waive TestQwen3_5_4B::test_bf16 for a…
tburt-nv Jun 9, 2026
64497e2
[None][doc] Add docs for AutoDeploy transforms (#15122)
bmarimuthu-nv Jun 9, 2026
5115366
[TRTLLM-12721][feat] Add disagg transfer state consensus
chienchunhung Jun 9, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
The diff you're trying to view is too large. We only load the first 3000 changed files.
6 changes: 2 additions & 4 deletions .claude/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -43,7 +43,7 @@ There are two ways to trigger skills and agents:
sub-agent automatically.

2. **Manual invoke** — type `/<skill-name>` (e.g. `/perf-analysis`,
`/serve-config-guide`) to explicitly run a skill. For sub-agents, type
`/trtllm-serve-config-guide`) to explicitly run a skill. For sub-agents, type
`@"<agent-name>" (agent)` (e.g. `@"exec-compile-specialist (agent)"`) to
delegate a task directly. This is useful when you know exactly which workflow you want.

Expand All @@ -63,12 +63,10 @@ short and not repeat it.
| Prefix | Domain | Definition |
|---|---|---|
| `ad-` | AutoDeploy | Model onboarding, pipeline debugging, and execution for the AutoDeploy backend |
| `ci-` | CI/CD | CI failure retrieval, test diagnostics, and pipeline workflows |
| `exec-` | Execution infra | Environment setup and job execution (compile, run, container) |
| `kernel-` | Kernel development | Kernel writing, generation, and kernel-specific transforms |
| `perf-` | Performance work | Profiling, analysis, and tuning above the kernel layer (kernel modifications belong under `kernel-`) |
| `serve-` | Serving | Serving configuration, deployment, and runtime workflows |
| `trtllm-` | TRT-LLM dev workflows | Workflows for reading, modifying, and contributing to the codebase (static subsystem knowledge belongs in repo docs) |
| `trtllm-` | TRT-LLM project workflows | Project-specific workflows: codebase exploration, contribution, dependency upgrades, and serving configuration (static subsystem knowledge belongs in repo docs) |

Guidelines:

Expand Down
30 changes: 30 additions & 0 deletions .claude/agent-tests/perf-test-sync/build_prompt.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
"""Promptfoo prompt builder.

Loads perf-test-sync.md, strips Claude Code specific sections, and appends the
user request.

Stripped sections:
- YAML frontmatter between leading `---` markers
- `# Persistent Agent Memory` section and everything after it (Claude Code
memory infrastructure, not relevant to prompt-quality evaluation)
"""

import os
import re

_SCRIPT_DIR = os.path.dirname(os.path.abspath(__file__))
_AGENT_MD = os.path.normpath(os.path.join(_SCRIPT_DIR, "..", "..", "agents", "perf-test-sync.md"))


def _load_agent_body() -> str:
with open(_AGENT_MD, "r", encoding="utf-8") as f:
text = f.read()
text = re.sub(r"\A---\n.*?\n---\n", "", text, count=1, flags=re.DOTALL)
text = text.split("# Persistent Agent Memory", 1)[0].rstrip()
return text


def build(context: dict) -> str:
user_prompt = context["vars"]["prompt"]
agent_body = _load_agent_body()
return f"{agent_body}\n\n## User request\n\n{user_prompt}\n"
Loading