Skip to content

[None][chore] Weekly mass integration of release/1.2 - #11572

Merged
chzblych merged 6 commits into
NVIDIA:mainfrom
mikeiovine:mi-release-1.2
Feb 24, 2026
Merged

[None][chore] Weekly mass integration of release/1.2#11572
chzblych merged 6 commits into
NVIDIA:mainfrom
mikeiovine:mi-release-1.2

Conversation

@mikeiovine

@mikeiovine mikeiovine commented Feb 18, 2026

Copy link
Copy Markdown
Collaborator

Description

Weekly round of mass integration from release branch.

Skipped:

pick d3b26dc4f [https://nvbugs/5889564][fix] fix kwargs name (#11496)
pick 3cdddddcf [https://nvbugs/5833795][fix] Remove test waive and try CI (#11464)
drop de294fcce [None][infra] Check in most recent lock file from nightly pipeline
drop 3cc9e4780 [None][infra] Check in most recent lock file from nightly pipeline
drop d1e5956bf [None][infra] Check in most recent lock file from nightly pipeline
pick 0149e89e9 [https://nvbugs/5860137][fix] Adjust deepgemm tuning buckets to cover larger num_tokens's scope (#11494)
drop 1c207c19d [None][infra] Check in most recent lock file from nightly pipeline
drop 4bcc4418f [None][infra] Check in most recent lock file from nightly pipeline
pick aa4c6779e [https://nvbugs/5875296][fix] Fix TritonMOE test for Qwen3_30B_A3B (#11495)
drop fbda4771d [None][infra] Check in most recent lock file from nightly pipeline
drop cf1b00f7f [None][infra] Check in most recent lock file from nightly pipeline
drop 4a110dd60 [None][infra] Check in most recent lock file from nightly pipeline
drop 261627c55 [None][infra] Check in most recent lock file from nightly pipeline
pick c426c4910 [None][infra] Cherry pick plc pipeline for 1.2 (#11546)
drop 82f687870 [None][infra] Check in most recent lock file from nightly pipeline
pick 7d08591ca [https://nvbugs/839137][fix] Unwaive disagg unexpected ucx error (#11543)
drop 5006b5f66 [None][infra] Check in most recent lock file from nightly pipeline
drop 5bd86618a [None][infra] Check in most recent lock file from nightly pipeline
drop 1d229e7c2 [None][infra] Check in most recent lock file from nightly pipeline
drop 73db5c0ae [None][infra] Check in most recent lock file from nightly pipeline
drop 8e9c39a19 [None][infra] Check in most recent lock file from nightly pipeline
drop 275f61576 [None][infra] Check in most recent lock file from nightly pipeline
drop fe914812d [None][infra] Check in most recent lock file from nightly pipeline
drop c5127f61e [None][infra] Check in most recent lock file from nightly pipeline
drop 4e96ed13b [None][infra] Check in most recent lock file from nightly pipeline
pick 48c03d356 [https://nvbugs/5823783][fix] Fix multi-node trust_remote_code hang i… (#11383)
drop c42ba80f1 [None][infra] Waive failures on release 1.2 (#11639)
pick 4f6acbbde [None][chore] Fix gpu memory requirement in stress test (#11404)

Test Coverage

N/A

PR Checklist

Please review the following before submitting your PR:

  • PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.

  • PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.

  • Test cases are provided for new code paths (see test instructions)

  • Any new dependencies have been scanned for license and vulnerabilities

  • CODEOWNERS updated if ownership changes

  • Documentation updated as needed

  • Update tava architecture diagram if there is a significant design change in PR.

  • The reviewers assigned automatically/manually are appropriate for the PR.

  • Please check this after reviewing the above items as appropriate for this PR.

GitHub Bot Help

/bot [-h] ['run', 'kill', 'skip', 'reuse-pipeline'] ...

Provide a user friendly way for developers to interact with a Jenkins server.

Run /bot [-h|--help] to print this help message.

See details below for each supported subcommand.

Details

run [--reuse-test (optional)pipeline-id --disable-fail-fast --skip-test --stage-list "A10-PyTorch-1, xxx" --gpu-type "A30, H100_PCIe" --test-backend "pytorch, cpp" --add-multi-gpu-test --only-multi-gpu-test --disable-multi-gpu-test --post-merge --extra-stage "H100_PCIe-TensorRT-Post-Merge-1, xxx" --detailed-log --debug(experimental)]

Launch build/test pipelines. All previously running jobs will be killed.

--reuse-test (optional)pipeline-id (OPTIONAL) : Allow the new pipeline to reuse build artifacts and skip successful test stages from a specified pipeline or the last pipeline if no pipeline-id is indicated. If the Git commit ID has changed, this option will be always ignored. The DEFAULT behavior of the bot is to reuse build artifacts and successful test results from the last pipeline.

--disable-reuse-test (OPTIONAL) : Explicitly prevent the pipeline from reusing build artifacts and skipping successful test stages from a previous pipeline. Ensure that all builds and tests are run regardless of previous successes.

--disable-fail-fast (OPTIONAL) : Disable fail fast on build/tests/infra failures.

--skip-test (OPTIONAL) : Skip all test stages, but still run build stages, package stages and sanity check stages. Note: Does NOT update GitHub check status.

--stage-list "A10-PyTorch-1, xxx" (OPTIONAL) : Only run the specified test stages. Examples: "A10-PyTorch-1, xxx". Note: Does NOT update GitHub check status.

--gpu-type "A30, H100_PCIe" (OPTIONAL) : Only run the test stages on the specified GPU types. Examples: "A30, H100_PCIe". Note: Does NOT update GitHub check status.

--test-backend "pytorch, cpp" (OPTIONAL) : Skip test stages which don't match the specified backends. Only support [pytorch, cpp, tensorrt, triton]. Examples: "pytorch, cpp" (does not run test stages with tensorrt or triton backend). Note: Does NOT update GitHub pipeline status.

--only-multi-gpu-test (OPTIONAL) : Only run the multi-GPU tests. Note: Does NOT update GitHub check status.

--disable-multi-gpu-test (OPTIONAL) : Disable the multi-GPU tests. Note: Does NOT update GitHub check status.

--add-multi-gpu-test (OPTIONAL) : Force run the multi-GPU tests in addition to running L0 pre-merge pipeline.

--post-merge (OPTIONAL) : Run the L0 post-merge pipeline instead of the ordinary L0 pre-merge pipeline.

--extra-stage "H100_PCIe-TensorRT-Post-Merge-1, xxx" (OPTIONAL) : Run the ordinary L0 pre-merge pipeline and specified test stages. Examples: --extra-stage "H100_PCIe-TensorRT-Post-Merge-1, xxx".

--detailed-log (OPTIONAL) : Enable flushing out all logs to the Jenkins console. This will significantly increase the log volume and may slow down the job.

--debug (OPTIONAL) : Experimental feature. Enable access to the CI container for debugging purpose. Note: Specify exactly one stage in the stage-list parameter to access the appropriate container environment. Note: Does NOT update GitHub check status.

For guidance on mapping tests to stage names, see docs/source/reference/ci-overview.md
and the scripts/test_to_stage_mapping.py helper.

kill

kill

Kill all running builds associated with pull request.

skip

skip --comment COMMENT

Skip testing for latest commit on pull request. --comment "Reason for skipping build/test" is required. IMPORTANT NOTE: This is dangerous since lack of user care and validation can cause top of tree to break.

reuse-pipeline

reuse-pipeline

Reuse a previous pipeline to validate current commit. This action will also kill all currently running builds associated with the pull request. IMPORTANT NOTE: This is dangerous since lack of user care and validation can cause top of tree to break.

Summary by CodeRabbit

  • Tests
    • Updated test configurations to focus on CUTLASS and TRTLLM backends
    • Consolidated MXFP4 latency test coverage across multiple test suites
    • Simplified test parameterizations for improved maintainability
    • Removed obsolete test skip entries

@mikeiovine

Copy link
Copy Markdown
Collaborator Author

/bot run

@coderabbitai

coderabbitai Bot commented Feb 18, 2026

Copy link
Copy Markdown
Contributor
📝 Walkthrough

Walkthrough

This pull request removes the TRITON MOE backend option from test parameterizations and updates test configurations to use W4A16 MXFP4 variants instead of W4A8 MXFP4, while also renaming a test utility parameter and removing an obsolete test waiver entry.

Changes

Cohort / File(s) Summary
MOE Backend Test Parameterization Updates
tests/integration/defs/accuracy/test_llm_api_pytorch.py
Removed TRITON from MOE backend parameterizations in multiple test functions, simplifying test decorators to use only CUTLASS and TRTLLM backends. Updated test_w4a8_mxfp4, test_w4a16_mxfp4, and test_nvfp4 decorator blocks accordingly.
Test Configuration List Updates
tests/integration/test_lists/qa/llm_function_core.txt, tests/integration/test_lists/qa/llm_function_core_sanity.txt, tests/integration/test_lists/test-db/l0_b200.yml
Replaced W4A8 MXFP4 TRITON latency test entries with W4A16 MXFP4 latency test entries across multiple test list configurations, consolidating latency test coverage to use the updated MXFP4 variant.
Test Waiver Cleanup
tests/integration/test_lists/waives.txt
Removed a single test skip entry for test_ptp_quickstart_advanced[GPT-OSS-120B-gpt_oss/gpt-oss-120b].
Test Utility Parameter Rename
tests/unittest/llmapi/apps/_test_disagg_serving_multi_nodes.py
Renamed keyword argument from time_leniency_seconds to time_tolerance_seconds in validate_timing_metrics call within test_completion function.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly identifies this as a weekly mass integration from the release branch with [chore] type designation, directly matching the PR's purpose.
Description check ✅ Passed The PR description adequately explains the purpose (weekly mass integration from release branch), lists picked/dropped commits with justification, and completes the template sections.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
  • 📝 Generate docstrings (stacked PR)
  • 📝 Generate docstrings (commit on current branch)
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment

Comment @coderabbitai help to get the list of available commands and usage tips.

@mikeiovine
mikeiovine requested a review from dongfengy February 18, 2026 20:38
@mikeiovine

Copy link
Copy Markdown
Collaborator Author

/bot run

@dongfengy dongfengy left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Duplicate comments:
In `@tests/integration/test_lists/qa/llm_function_core.txt`:
- Line 149: Confirm that the test entry
accuracy/test_llm_api_pytorch.py::TestQwen3_30B_A3B::test_w4a16_mxfp4[latency-TRITON]
in tests/integration/test_lists/qa/llm_function_core.txt exactly matches the
validated entry in llm_function_core_sanity.txt: verify the test function name
test_w4a16_mxfp4 and the test class TestQwen3_30B_A3B exist in
accuracy/test_llm_api_pytorch.py and that the "[latency-TRITON]" marker is a
valid parameterization; if anything differs, update the entry in
llm_function_core.txt to match the canonical line in
llm_function_core_sanity.txt (or add the missing test/function to
accuracy/test_llm_api_pytorch.py if the test is absent).

@tburt-nv

Copy link
Copy Markdown
Collaborator

/bot run

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #36182 [ run ] triggered by Bot. Commit: 514411d Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #36182 [ run ] completed with state FAILURE. Commit: 514411d
/LLM/main/L0_MergeRequest_PR pipeline #27965 completed with status: 'FAILURE'

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

Link to invocation

@mikeiovine

Copy link
Copy Markdown
Collaborator Author

/bot run

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #36251 [ run ] triggered by Bot. Commit: 5321a82 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #36251 [ run ] completed with state FAILURE. Commit: 5321a82
/LLM/main/L0_MergeRequest_PR pipeline #28028 completed with status: 'FAILURE'

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

Link to invocation

@mikeiovine
mikeiovine force-pushed the mi-release-1.2 branch 2 times, most recently from 4f23f3b to 15f494a Compare February 19, 2026 16:29
@mikeiovine

Copy link
Copy Markdown
Collaborator Author

/bot run

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #36258 [ run ] triggered by Bot. Commit: 15f494a Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #36258 [ run ] completed with state SUCCESS. Commit: 15f494a
/LLM/main/L0_MergeRequest_PR pipeline #28034 completed with status: 'FAILURE'

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

Link to invocation

@mikeiovine

Copy link
Copy Markdown
Collaborator Author

/bot run

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #36263 [ run ] triggered by Bot. Commit: 77390e3 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #36263 [ run ] completed with state SUCCESS. Commit: 77390e3
/LLM/main/L0_MergeRequest_PR pipeline #28037 completed with status: 'FAILURE'

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

Link to invocation

@mikeiovine

Copy link
Copy Markdown
Collaborator Author

/bot run

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #36281 [ run ] triggered by Bot. Commit: 77390e3 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #36281 [ run ] completed with state SUCCESS. Commit: 77390e3
/LLM/main/L0_MergeRequest_PR pipeline #28054 completed with status: 'FAILURE'

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

Link to invocation

@mikeiovine

Copy link
Copy Markdown
Collaborator Author

/bot run

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #36340 [ run ] triggered by Bot. Commit: 2837103 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #36340 [ run ] completed with state SUCCESS. Commit: 2837103
/LLM/main/L0_MergeRequest_PR pipeline #28110 completed with status: 'SUCCESS'

Link to invocation

@mikeiovine

Copy link
Copy Markdown
Collaborator Author

/bot skip --comment "CI already passed"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #36383 [ skip ] triggered by Bot. Commit: 167e8aa Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #36383 [ skip ] completed with state SUCCESS. Commit: 167e8aa
Skipping testing for commit 167e8aa

Link to invocation

@juney-nvidia
juney-nvidia enabled auto-merge (squash) February 22, 2026 00:23
@mikeiovine

Copy link
Copy Markdown
Collaborator Author

Added more PRs per offline discussion.

@mikeiovine

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #36540 [ run ] triggered by Bot. Commit: eedb211 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #36540 [ run ] completed with state SUCCESS. Commit: eedb211
/LLM/main/L0_MergeRequest_PR pipeline #28276 completed with status: 'FAILURE'

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

Link to invocation

reasonsolo and others added 6 commits February 23, 2026 13:13
Signed-off-by: Lizhi Zhou <1432185+reasonsolo@users.noreply.github.com>
Signed-off-by: Mike Iovine <miovine@nvidia.com>
)

Signed-off-by: Dongfeng Yu <dongfengy@nvidia.com>
Signed-off-by: dongfengy <99041270+dongfengy@users.noreply.github.com>
Signed-off-by: Mike Iovine <miovine@nvidia.com>
…VIDIA#11495)

Signed-off-by: Dongfeng Yu <dongfengy@nvidia.com>
Signed-off-by: Mike Iovine <miovine@nvidia.com>
…DIA#11543)

Signed-off-by: Patrice Castonguay <55748270+pcastonguay@users.noreply.github.com>
Signed-off-by: Mike Iovine <miovine@nvidia.com>
NVIDIA#11383)

Signed-off-by: Junyi Xu <219237550+JunyiXu-nv@users.noreply.github.com>
Signed-off-by: Mike Iovine <miovine@nvidia.com>
Signed-off-by: Wangshanshan <30051912+dominicshanshan@users.noreply.github.com>
Signed-off-by: Mike Iovine <miovine@nvidia.com>
@mikeiovine

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #36550 [ run ] triggered by Bot. Commit: 78e49d1 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #36550 [ run ] completed with state SUCCESS. Commit: 78e49d1
/LLM/main/L0_MergeRequest_PR pipeline #28286 completed with status: 'SUCCESS'
Pipeline passed with automatic retried tests. Check the rerun report for details.

Link to invocation

@chzblych
chzblych merged commit 951a467 into NVIDIA:main Feb 24, 2026
9 checks passed
@mikeiovine
mikeiovine deleted the mi-release-1.2 branch February 24, 2026 05:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.