Skip to content

[CI/Build] Diff-aware source_file_dependencies for CUDA/NPU pipelines - #6597

Merged
Gaohan123 merged 15 commits into
vllm-project:mainfrom
yenuo26:trigger
Sep 10, 2026
Merged

Gaohan123 merged 15 commits into
vllm-project:mainfrom
yenuo26:trigger

Conversation

@yenuo26

@yenuo26 yenuo26 commented Aug 25, 2026 •

Copy link
Copy Markdown
Collaborator

PLEASE FILL IN THE PR DESCRIPTION HERE.

Purpose

Make CUDA L2–L5 and NPU L4 jobs diff-aware on PR labels, so a nightly-test / weekly-test / ready / merge-test push no longer uploads the full pipeline. Scheduled main uploads (NIGHTLY=1, WEEKLY=1 / NON_CRITICAL=1, post-merge L3) still run every job.

Fixes #6507.

source_file_dependencies is a preset key (or a list of keys) from .buildkite/common/ci_source_file_dependencies.yml. upload_pipeline.py expands the key, keeps the step only when a changed file matches a prefix, then strips the field before Buildkite sees the YAML. Jobs with no key are always kept. --all / --e2e and BUILDKITE_BRANCH=main disable filtering.

Registry layout:

  • Feature keys stay cross-cutting (diffusion_distributed_attention, diffusion_tiny_model, …).
  • YAML anchors hold model business-code paths. Job keys (*_function / _perf / _accuracy / _reliability / _doc / _cov) alias those anchors and add the scripts that job actually runs. Pytest paths are listed on the key; they are not inferred from commands.
  • Cross-model invalid-param reliability jobs use reliability_invalid_param_h100 / reliability_invalid_param_l4 (model anchors + tests/dfx/reliability/invalid_param_test/), not *_function.

Job wiring that follows from that:

  • L2/L3 E2E, L4 CUDA/NPU, and weekly Reliability/Perf jobs point at the matching kind key. Weekly E2E (NON_CRITICAL=1 only) has no source filter.
  • Shared H100 diffusion sweeps that timed out are split by model (list below). LingBot single-GPU coverage moved from slow to full_model and is one nightly job, Diffusion X2V · LingBot Function Test (diffusion_lingbot_function). World still skips unless VLLM_OMNI_LINGBOT_WORLD_V2_IMAGE_PATH and VLLM_OMNI_LINGBOT_WORLD_V2_ACTION_DIR are set; CI does not set them.
  • X2V Doc Test uses diffusion_image_to_video_doc for the image-to-video README snippets that job owns: Wan2.2 TI2V/I2V, HunyuanVideo-1.5 I2V, LTX-2, SANA-Video, plus LingBot-Video (README command is skipped on /path/to/, but its model code still selects the job). Python API, Prerequisites, Advanced Features, FAQ, and Wan2.1 VACE stay out of this key.

Unused PR labels omni-test, tts-test, diffusion-x2iat-test, and diffusion-x2v-test no longer trigger CI. nightly-test remains the L4 label.

Test Plan

UT

Uploader unit tests construct synthetic pipelines instead of asserting live job labels. They cover key expansion, unrelated-file drops, jobs with no source key, composed doc paths, and stripping source_file_dependencies / expanding mirror_hardwares.

python -m pytest -sv tests/buildkite/test_upload_pipeline.py"

Registry keys referenced by CUDA/NPU YAML are checked by test_pipeline_source_file_dependency_keys_are_registered in the same file.

Weekly H100 diffusion jobs split from the old shared sweeps

The old Diffusion · H100 · Single-GPU and Diffusion · H100 · 2-GPU sweeps were split by model so one pytest process no longer collects every slow+cards_N file. Each job below must collect at least one test (--collect-only should print node ids, not no tests collected). Collection does not need GPU weights; it does need the repo test imports to succeed.

LingBot is not in this weekly list: those cases moved to nightly Diffusion X2V · LingBot Function Test.

Single-GPU (-m "slow and diffusion and H100 and cards_1" --run-level full_model):

pytest --collect-only -q tests/e2e/online_serving/test_qwen_image_edit.py tests/e2e/online_serving/test_qwen_image_edit_expansion.py tests/e2e/online_serving/test_qwen_image_layered.py tests/e2e/online_serving/test_qwen_image_layered_expansion.py tests/e2e/accuracy/test_qwen_image_edit.py tests/e2e/accuracy/test_qwen_image_layered.py -m "slow and diffusion and H100 and cards_1" --run-level full_model
pytest --collect-only -q tests/e2e/online_serving/test_wan_2_1_vace_expansion.py -m "slow and diffusion and H100 and cards_1" --run-level full_model
pytest --collect-only -q tests/e2e/online_serving/test_longcat_image_expansion.py tests/e2e/online_serving/test_longcat_image_edit_expansion.py -m "slow and diffusion and H100 and cards_1" --run-level full_model
pytest --collect-only -q tests/e2e/online_serving/test_bagel.py tests/e2e/online_serving/test_bagel_expansion.py tests/e2e/offline_inference/test_bagel.py tests/e2e/offline_inference/test_bagel_expansion.py -m "slow and diffusion and H100 and cards_1" --run-level full_model
pytest --collect-only -q tests/e2e/online_serving/test_boogu_image.py tests/e2e/online_serving/test_boogu_image_edit.py tests/e2e/online_serving/test_boogu_image_expansion.py tests/e2e/online_serving/test_boogu_image_edit_expansion.py -m "slow and diffusion and H100 and cards_1" --run-level full_model
pytest --collect-only -q tests/e2e/online_serving/test_flux_2_dev_expansion.py tests/e2e/offline_inference/test_flux1_schnell_expansion.py -m "slow and diffusion and H100 and cards_1" --run-level full_model
pytest --collect-only -q tests/e2e/online_serving/test_ltx.py tests/e2e/online_serving/test_ltx25.py tests/e2e/accuracy/ltx/test_ltx_official_similarity.py tests/e2e/accuracy/ltx/test_ltx25_official_similarity.py -m "slow and diffusion and H100 and cards_1" --run-level full_model
pytest --collect-only -q tests/e2e/online_serving/test_krea2_expansion.py tests/e2e/offline_inference/test_krea2_expansion.py -m "slow and diffusion and H100 and cards_1" --run-level full_model
pytest --collect-only -q tests/e2e/online_serving/test_hidream_i1_full_expansion.py -m "slow and diffusion and H100 and cards_1" --run-level full_model
pytest --collect-only -q tests/e2e/online_serving/test_sana_video_expansion.py -m "slow and diffusion and H100 and cards_1" --run-level full_model
pytest --collect-only -q tests/e2e/offline_inference/test_sensenova_u1_expansion.py tests/e2e/offline_inference/test_sensenova_u1_text2img_expansion.py tests/e2e/offline_inference/test_sensenova_u1_img2img_expansion.py -m "slow and diffusion and H100 and cards_1" --run-level full_model
pytest --collect-only -q tests/e2e/offline_inference/test_mammoth_moda2_expansion.py -m "slow and diffusion and H100 and cards_1" --run-level full_model
pytest --collect-only -q tests/e2e/accuracy/test_video_streaming_output_similarity.py -m "slow and diffusion and H100 and cards_1" --run-level full_model

2-GPU (-m "slow and diffusion and H100 and cards_2" --run-level full_model):

pytest --collect-only -q tests/e2e/online_serving/test_qwen_image_edit_expansion.py tests/e2e/online_serving/test_qwen_image_layered_expansion.py -m "slow and diffusion and H100 and cards_2" --run-level full_model
pytest --collect-only -q tests/e2e/online_serving/test_wan_2_1_vace_expansion.py -m "slow and diffusion and H100 and cards_2" --run-level full_model
pytest --collect-only -q tests/e2e/online_serving/test_longcat_image_expansion.py tests/e2e/online_serving/test_longcat_image_edit_expansion.py -m "slow and diffusion and H100 and cards_2" --run-level full_model
pytest --collect-only -q tests/e2e/online_serving/test_bagel_expansion.py -m "slow and diffusion and H100 and cards_2" --run-level full_model
pytest --collect-only -q tests/e2e/online_serving/test_boogu_image_expansion.py tests/e2e/online_serving/test_boogu_image_edit_expansion.py -m "slow and diffusion and H100 and cards_2" --run-level full_model
pytest --collect-only -q tests/e2e/online_serving/test_flux_2_dev_expansion.py tests/e2e/offline_inference/test_flux1_kontext_expansion.py -m "slow and diffusion and H100 and cards_2" --run-level full_model
pytest --collect-only -q tests/e2e/online_serving/test_glm_image_expansion.py tests/e2e/offline_inference/test_glm_image_autoround_w4a16_expansion.py -m "slow and diffusion and H100 and cards_2" --run-level full_model
pytest --collect-only -q tests/e2e/online_serving/test_ltx.py -m "slow and diffusion and H100 and cards_2" --run-level full_model
pytest --collect-only -q tests/e2e/offline_inference/test_krea2_expansion.py -m "slow and diffusion and H100 and cards_2" --run-level full_model
pytest --collect-only -q tests/e2e/online_serving/test_dreamzero_expansion.py tests/e2e/accuracy/test_dreamzero.py -m "slow and diffusion and H100 and cards_2" --run-level full_model
pytest --collect-only -q tests/e2e/accuracy/test_hidream_o1_image.py -m "slow and diffusion and cuda and H100 and cards_2" --run-level full_model

3&4-GPU (-m "slow and diffusion and H100 and cards_2" --run-level full_model):

pytest --collect-only -q tests/e2e/online_serving/test_boogu_image_edit_expansion.py -k "test_double_guidance_cfg_parallel and cfg_parallel_3_full" -m "slow and diffusion and cuda and H100 and cards_3" --run-level full_model
pytest --collect-only -q tests/e2e/online_serving/test_magi2.py -m "slow and diffusion and cuda and H100 and cards_4" --run-level full_model

LingBot nightly command (H100, 1 GPU; World skips without the two asset env vars):

pytest -s -v tests/e2e/online_serving/test_lingbot_video.py tests/e2e/online_serving/test_lingbot_video_moe.py tests/e2e/offline_inference/test_lingbot_world_v2.py -m "full_model and diffusion and H100 and cards_1" --run-level "full_model"

Full ready/merge/nightly/weekly execution is expected from Buildkite after this PR is labeled. This change is CI YAML and the uploader only.

vLLM Version: N/A (no runtime / vLLM pin change)

vLLM-Omni Commit: 99ede1b

Test Result

  • UT
bdd4c3f9-cc8e-43ac-893c-827c7a131a28
  • Weekly H100 diffusion jobs split from the old shared sweeps
    single GPU
94bd5aa5-a9a4-4c45-a574-f3757dc8a2ca ecfb27aa-2d80-4855-8f4f-8fc9fb143982 fafa620b-e402-496d-8bd0-33ceb9bdc649 1ec2376b-60de-46f0-bb1c-a691e0faeaea 32d4526a-86f2-40b9-9e62-36370bc22954 93778729-da90-40be-8027-33d99aef2d01 89f868b0-ba3a-4d8e-b8ea-800d77040900 40c9a7dc-bbba-4965-a6de-5fa86fdc5251 ec98d1cc-768b-407d-8ee8-62a3eef63c00 2f896621-584c-4a31-8ab2-3965990e1af0 60d082ea-0ce5-4eec-9a6b-1235faf1974d ee3dd059-eb2f-4e4e-9ecc-c6091f7b313e 70e43a9f-30f5-472e-8852-1142b207ba93

2-GPU
cb51e1bf-2a41-4606-86b8-207aa8aa9c43
68820292-9441-4623-9176-6c6a48181068
611f8a36-0102-4dff-8ec9-37ff05957c3c
804d512b-0c24-4a26-a8fe-b2aaa6794009
4c705538-98dd-4522-977b-a1f802bab989
45be15fd-3f5d-48d0-aef8-375a5fce7731
1acb99e3-dd73-4bbb-8c83-87a90413c65a
27f4af52-74db-4514-85d9-7ccebf6c4578
b1ac0a0c-efe4-4f3d-a242-fa8773932583
87251229-5cc7-4e5c-8380-4787a82bb5d4
422dcd8f-cf54-470d-98e6-211fe8999267

3&4-GPU
2ccf3a2f-a276-4373-b327-7509e5a9b32d
da214084-2f2a-40ef-922d-c8ed1769a564

  • LingBot nightly
da3d8ad1-2688-44b6-a208-aa3dcaa3790c
  • Trigger
    when i modify lingbot and minimax testcase + nightly-test label
    GPU
a81149da-6315-4cb2-a6bc-a68f33788581

NPU
d13ded58-4f32-48d3-a2b2-b91b2aa24f75

when i modify lingbot and minimax testcase + NIGHTLY=1
GPU
94e8810d-e461-4c17-8dfb-747457472626

NPU
57a22f7f-44e3-4fdc-883f-57af667d2ee6

BEFORE SUBMITTING: read CONTRIBUTING.md and run the precheck-pr skill with the code agent for a self-check against project conventions.
(anything written below this line will be removed by GitHub Actions)

yenuo26 and others added 2 commits August 25, 2026 11:44
- Refactored `tests/helpers/runtime.py` into `tests/helpers/client.py` and `tests/helpers/clean.py`.
- Updated CI configuration to use new source file dependencies for various tests.
- Enhanced test coverage and consolidated TTS coverage, including moving specific test cases.
- Improved error handling in `OmniServer` and `OmniRunner` teardown processes.
- Adjusted e2e tests to utilize `get_open_port()` and maintain consistent CUDA device mapping.

This refactor aims to improve code organization and maintainability while ensuring comprehensive test coverage across different models and configurations.

Signed-off-by: wangyu <410167048@qq.com>
Signed-off-by: wangyu <410167048@qq.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@yenuo26 yenuo26 changed the title [CI/Build] Diff-aware source_file_dependencies for CUDA/NPU pipelines [CI/Build][WIP] Diff-aware source_file_dependencies for CUDA/NPU pipelines Aug 25, 2026
@yenuo26 yenuo26 added the CI/CD codes related to changes to CI/CD label Aug 25, 2026
@vllm-omni-review-bot

Copy link
Copy Markdown

This PR was classified as CI work.

CI owner: @yenuo26

@yenuo26, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@yenuo26 yenuo26 linked an issue Aug 26, 2026 that may be closed by this pull request
1 task done
Keep source_file_dependencies preset keys; take main's GPU-split job labels and extra path prefixes.

Signed-off-by: wangyu <410167048@qq.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
@yenuo26 yenuo26 changed the title [CI/Build][WIP] Diff-aware source_file_dependencies for CUDA/NPU pipelines [CI/Build] Diff-aware source_file_dependencies for CUDA/NPU pipelines Sep 2, 2026
Keep source_file_dependencies preset keys and label-only filtering; take main's MIRROR_HW inference and GPU-split job labels.

Signed-off-by: wangyu <410167048@qq.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 7, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Resolved as of 0d6baf288dcf: the high-priority or low-quality signal noted on an earlier commit no longer applies.

@yenuo26 yenuo26 added nightly-test label to trigger buildkite nightly test CI cuda-test Used to trigger vllm-omni cuda CI separately. and removed nightly-test label to trigger buildkite nightly test CI cuda-test Used to trigger vllm-omni cuda CI separately. labels Sep 7, 2026
…date test configurations in YAML files to reflect new command structures and improve test accuracy.

Signed-off-by: wangyu <410167048@qq.com>
@yenuo26 yenuo26 added nightly-test label to trigger buildkite nightly test CI cuda-test Used to trigger vllm-omni cuda CI separately. labels Sep 7, 2026
…s; refactor upload_pipeline.py to streamline dependency resolution and improve test configurations. Adjust test cases to align with new dependency structures.

Signed-off-by: wangyu <410167048@qq.com>
@yenuo26 yenuo26 removed nightly-test label to trigger buildkite nightly test CI cuda-test Used to trigger vllm-omni cuda CI separately. labels Sep 8, 2026
Comment thread .buildkite/common/ci_source_file_dependencies.yml Outdated
Comment thread .buildkite/common/ci_source_file_dependencies.yml
Comment thread .buildkite/common/ci_source_file_dependencies.yml Outdated
…xes for clarity; update test configurations in ci_settings.md and test_upload_pipeline.py to align with new dependency structures. Remove obsolete model keys from ci_source_file_dependencies.yml.

Signed-off-by: wangyu <410167048@qq.com>
Signed-off-by: wangyu <410167048@qq.com>
@yenuo26 yenuo26 added nightly-test label to trigger buildkite nightly test CI cuda-test Used to trigger vllm-omni cuda CI separately. npu-test and removed nightly-test label to trigger buildkite nightly test CI cuda-test Used to trigger vllm-omni cuda CI separately. labels Sep 8, 2026
Signed-off-by: wangyu <410167048@qq.com>
@yenuo26 yenuo26 removed nightly-test label to trigger buildkite nightly test CI npu-test labels Sep 8, 2026
Signed-off-by: wangyu <410167048@qq.com>
@Gaohan123 Gaohan123 added this to the v0.30.0 milestone Sep 9, 2026
(NIGHTLY/WEEKLY/post-merge) run the full pipeline. ``--all`` / ``--e2e``
also disable filtering.
"""
if force_all or e2e_only:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Keep affected jobs when the pipeline configuration changes

If a PR only changes .buildkite/npu/test-npu-nightly.yml and is triggered with nightly-test, the YAML path matches none of the registered source dependencies, so every test job is filtered out. The equivalent CUDA change leaves only the email aggregation step. This prevents changes to test commands, environments, or hardware settings from being exercised before merge. Please bypass source filtering for the affected pipeline when its YAML changes, and handle shared uploader/registry changes similarly.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

fixed, Added fallback logic for common modules. If no source code matches are found, it will additionally check whether the changes involve common uploaders, helpers functions, or Buildkite configuration files. If any of these are matched, source filtering will be skipped.

diffusion_text_to_image_doc:
- *qwen_image
- *z_image
- tests/examples/offline_inference/test_text_to_image.py

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Include the example scripts executed by the text-to-image doc tests

This preset includes the test wrappers but omits the example scripts they execute. For example, changing only examples/online_serving/text_to_image/openai_chat_client.py or examples/offline_inference/text_to_image/text_to_image.py does not select the Doc Test job, even though those changes can break its tests. Please include examples/online_serving/text_to_image/ and examples/offline_inference/text_to_image/ in this preset so example-only changes receive the corresponding validation.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

fixed

- tests/e2e/online_serving/run_minicpmo_realtime_duplex_multi_session.py
- tests/e2e/online_serving/run_minicpmo_realtime_duplex_server_vad.py
- tests/e2e/online_serving/run_minicpmo_realtime_duplex_soft_interrupt.py
- tests/e2e/online_serving/helpers/minicpmo_4_5_duplex.py

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Preserve the duplex audio fixtures in the dependency preset

The previous ready/merge dependency lists included tests/assets/minicpmo_4_5/response_required_16k.wav and tests/assets/minicpmo_4_5/soft_interrupt_16k.wav, but this preset drops both. The current duplex helpers still read these files and validate their SHA256 and audio format, so changing either fixture alone can break the tests without selecting the corresponding duplex jobs. Please restore both paths, or include their containing asset directory.

This remains reproducible at 0d6baf288dcf, despite the related earlier discussion being resolved.

@yenuo26 yenuo26 Sep 9, 2026 •

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

When deleting, the consideration was that media files are generally not modified in terms of content, but rather replaced along with test files. However, it might be more appropriate to include tests/assets/minicpmo_4_5/, and this has now been added.

…lter_fallback` key in `ci_source_file_dependencies.yml` to retain all jobs when no job-key prefix matches. Update `upload_pipeline.py` to handle this fallback logic and adjust related test cases in `test_upload_pipeline.py` to ensure correct behavior. Revise documentation in `ci_settings.md` to clarify the filtering process and fallback conditions.

Signed-off-by: wangyu <410167048@qq.com>
…lter fallback behavior. Update assertions to ensure correct handling of job-key matches and pipeline YAML dependencies. Enhance test coverage for scenarios involving matching prefixes and fallback conditions.

Signed-off-by: wangyu <410167048@qq.com>
return doc
if e2e_only:
steps = _select_e2e_group_steps(steps)
# Bypass is a fallback: only when no job-key prefix matched.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Apply the shared-path bypass even when a model dependency matches

The fallback fixes configuration-only changes, but mixed changes still lose required coverage. At 63114228f2a05, changing only .buildkite/npu/test-npu-nightly.yml retains all 15 jobs; adding a Qwen3-Omni model change to the same diff reduces that to its two performance jobs. Other jobs whose commands, environments, or hardware settings may have changed in that YAML are filtered out. The same reduction occurs when combining tests/helpers/runtime.py with a Qwen3-Omni model change.

Please check the pipeline/shared-path bypass independently of _any_source_dependency_match, so an additional model change cannot suppress validation required by the shared change. The mixed-change test should expect all affected jobs to remain selected.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On this point, I lean toward keeping the current design where only the corresponding jobs are triggered based on changes to business code. My reasoning is as follows:
For example, when a user modifies the relevant business code and needs to make changes to the test cases for that model, they would modify the Buildkite configuration to add or adjust the corresponding job, or modify the relevant code in the helper functions. But if we go down the path of mixed-change judgment, then all jobs would be triggered in that case. And I think this kind of situation is likely the majority, because users typically only want to trigger nightly tests in a PR when the nightly test cases have changed. This seems to run counter to our goal of reducing the number of jobs triggered in a PR.

@yenuo26 yenuo26 added nightly-test label to trigger buildkite nightly test CI cuda-test Used to trigger vllm-omni cuda CI separately. and removed nightly-test label to trigger buildkite nightly test CI cuda-test Used to trigger vllm-omni cuda CI separately. labels Sep 10, 2026
yenuo26 and others added 2 commits September 10, 2026 15:26
Resolve conflicts by keeping named MiniCPM source keys, adding the
native chat template to the model anchor, and combining source-filter
fallback with CUDA HF token injection.

Co-authored-by: Cursor <cursoragent@cursor.com>
@Gaohan123 Gaohan123 added the ready label to trigger buildkite CI label Sep 10, 2026

@Gaohan123 Gaohan123 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Thanks

@Gaohan123
Gaohan123 enabled auto-merge (squash) September 10, 2026 08:01
@Gaohan123
Gaohan123 merged commit 5b927b7 into vllm-project:main Sep 10, 2026
6 of 9 checks passed
JoseCarlosGarcia95 added a commit to valendra-tech/vllm-omni that referenced this pull request Sep 16, 2026
* [Bugfix][Examples] Use --profiler-config flag in offline TTS examples (vllm-project#6763)

Signed-off-by: Asthenia <asthenia0412@gmail.com>
Co-authored-by: Asthenia <asthenia0412@gmail.com>

* [Bugfix] Skip HWR store-size scans when no limit is configured (vllm-project#7131)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [CI][ROCm] Route LTX2 Ulysses parity to two-GPU lane (vllm-project#7234)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Bugfix][Model] GR00T-N1.7: honor the per-request seed for flow-matching noise (vllm-project#7253)

Signed-off-by: liangmengh <liangmengh@nvidia.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* Add vLLM-Omni library info to Hugging Face Hub requests (vllm-project#5381)

Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* [Bugfix][NPU] Limit MiniMax H3 modulation grid size (vllm-project#6794)

Signed-off-by: KrystalRay <keeleiray@gmail.com>
Co-authored-by: KrystalRay <keeleiray@gmail.com>

* [Bugfix] Build the forced-aligner prompt without a chat template (word timestamps one bin late) (vllm-project#7240)

Signed-off-by: Tianyao Wu <rayroy31@gmail.com>

* [Refactor][Diffusion] Resolve offload topology through one plan resolver (vllm-project#7209)

Signed-off-by: specture724 <specture724@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [Doc] Add AI usage policy for contributions (vllm-project#7305)

Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>

* [Bugfix][MiMo-Audio] Align code2wav decode with tokenizer device (vllm-project#6539)

Signed-off-by: chaosansui <zzc15560846421@163.com>
Signed-off-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com>

* [Bugfix][MiniCPM-o] Fix the audio_embeds input path (vllm-project#5730)

Signed-off-by: eval-dev <0xe5bca0@gmail.com>
Signed-off-by: eval <74645252+eval-dev@users.noreply.github.com>

* [Feat][OmniVoice]Support Varlen Attn,  Request-Batch and Step-Execution (vllm-project#6408)

Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>

* [Model] Add Audio8 TTS Preview 0.6B (DualAR, 44.1 kHz codec) (vllm-project#6157)

Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com>
Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com>

* [Bugfix][Frontend] Accept the msgpack-numpy package's numpy markers on the OpenPI endpoint (vllm-project#6051)

Signed-off-by: zjli2013 <leezhengjiang@126.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Frontend] Opt-in WebSocket TTS split_granularity and session seed (vllm-project#7046)

Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu>
Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [Bugfix][Frontend] Clear the P0 multimodal cache through the renderer (vllm-project#7003)

Signed-off-by: ZenAlexa <zimingwang945@gmail.com>

* [Bugfix][Frontend] Enforce image pixel limits for video input references (vllm-project#6963)

Signed-off-by: BANANASJIM <bananasjim1@gmail.com>

* [Bugfix][TTS] Isolate shared Higgs v3 reference encode from request cancellation (vllm-project#7076)

Signed-off-by: Allen Wu <allenwu2795@gmail.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

* [Bugfix][CosyVoice3] Resolve hash snapshot pipeline (vllm-project#6896)

Signed-off-by: xutianle <xutianle@fudan.edu.cn>

* [CI] Skip Qwen3-Omni Server VAD multi-turn realtime test (vllm-project#7279) (vllm-project#7314)

Signed-off-by: wangyu <410167048@qq.com>

* [Bugfix][Magi2] Allow import without an active Triton driver (vllm-project#7239)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Core] Split Omni connector model runner mixin (vllm-project#6903)

Signed-off-by: natureofnature <wzliu@connect.hku.hk>

* [Bugfix] Make LTX vocoder decoding deterministic (vllm-project#7231)

Signed-off-by: mglyn <1203789601@qq.com>

* [Doc] [Recipe] Add FLUX.1-schnell recipe for RTX 5090 32GB (vllm-project#7299)

Signed-off-by: Sparks-M <41097544+Sparks-M@users.noreply.github.com>

* [Doc] Qwen3-TTS: add 0.6B on 1x A100 40GB (vllm-project#7289)

Signed-off-by: chi030303 <106855944+chi030303@users.noreply.github.com>

* [Perf][Model] Add optimized LTX-2.5 DiffVAE operators (vllm-project#7308)

Signed-off-by: mglyn <1203789601@qq.com>

* [2/N] Add a minimal temporal chunk callback for MiniMax-H3 (vllm-project#7017)

Signed-off-by: specture724 <specture724@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* [Feature][Diffusion] Expose detailed pipeline timings (vllm-project#6822)

Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com>

* [Bugfix] Resolve vllm-project#6931 hub FA3 on torch 2.13 via kernels 0.16.1 (vllm-project#7185)

Signed-off-by: NumberWan <wantszkin2003@gmail.com>

* [Bugfix][Ascend] fix npu 310/a5 bugs (vllm-project#6685)

Signed-off-by: zouyizhou <zouyizhou@huawei.com>

* [Bugfix][Engine] Group overlapping device stages into one sequential init component (vllm-project#7328)

Signed-off-by: ZhengWG <zwg0606@gmail.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* fix: reserve Qwen3-Omni NVFP4 backend fix (vllm-project#7200)

Signed-off-by: kunkunblueberry <1833921874@qq.com>

* [BugFix] Add field validators for /v1/audio/generate request (vllm-project#4741)

Signed-off-by: Shaun Walsh <shaunwalsh24@gmail.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Nick Cao <ncao@redhat.com>

* [CI][ROCm] Match CUDA/NPU L2/L3 label routing (vllm-project#6966)

Signed-off-by: andyluo7 <andy.luo@amd.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [CI/Build] Avoid duplicate stage CLI deploy config (vllm-project#7007)

Signed-off-by: mershi <mershi@tencent.com>
Co-authored-by: mershi <mershi@tencent.com>

* [CI/Build][ROCm] Normalize SenseNova paged-decode hardware markers (vllm-project#6935)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Model] Skip unused frame packing in Wan2.2 S2V (vllm-project#7155)

Signed-off-by: hyw <yuweih205@gmail.com>

* [Doc] Add dual DGX Spark MiniMax-H3 results (vllm-project#7343)

Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com>

* [Model] Optimize MOSS-TTS Local batched execution and streaming codec (vllm-project#7202)

Signed-off-by: Sy03 <1370724210@qq.com>

* [Bugfix][XPU] Restore N-D output shape for W8A16 FP8 linear (vllm-project#7301)

Signed-off-by: Joshna Medisetty <joshna.medisetty@intel.com>
Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com>

* [Doc] Document num_outputs_per_prompt for /v1/videos (vllm-project#7341)

Signed-off-by: Guangjian <hiro20833@gmail.com>

* [Skills] Add perf-evidence isolation, stage-attribution, and realtime-contract requirements (vllm-project#6820)

Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Co-authored-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>

* [Bugfix] Allow LLM replicas on different GPUs to initialize concurrently (vllm-project#7292)

Signed-off-by: Gao Han <hgaoaf@connect.ust.hk>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* [CI/Build] Stabilize LTX2 vocoder autocast test on ROCm (vllm-project#7336)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [NPU][CI] Add A5 and 310P CI support (vllm-project#6875)

Signed-off-by: Weiming Liao <liaowm5@gmail.com>
Co-authored-by: wangyu <53896905+yenuo26@users.noreply.github.com>

* [Kernel] Enable LTX DiffVAE fusions on SM100 and SM103 (vllm-project#7350)

Signed-off-by: mglyn <1203789601@qq.com>

* [Bugfix][MiniCPM-o] Align structured chat content with native omni rendering (vllm-project#7344)

Signed-off-by: Sy03 <1370724210@qq.com>

* [Rebase] Rebase to vLLM 0.29.0 (vllm-project#7230)

Signed-off-by: tzhouam <tzhouam@connect.ust.hk>
Signed-off-by: Zhou Taichang <tzhouam@connect.ust.hk>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* [Refactor] P0.2: Migrate API server helpers out of api_server (vllm-project#5453)

Signed-off-by: herotai214 <herotai214@gmail.com>

* [CI] Stabilize Qwen3-Omni Server VAD E2E (vllm-project#7356)

Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* [CI/Build] Diff-aware source_file_dependencies for CUDA/NPU pipelines (vllm-project#6597)

Signed-off-by: wangyu <410167048@qq.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [Core][Diffusion] Add a typed pre-D2H video media contract (vllm-project#6615)

Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com>
Signed-off-by: Samit <285365963@qq.com>
Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com>
Co-authored-by: Samit <285365963@qq.com>

* [Bugfix] Bound HWR domain initialization lock waits (vllm-project#7128)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Escalate diffusion worker shutdown and retain survivors (vllm-project#7126)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Misc] Add standalone safetensors retention diagnostic (vllm-project#7145)

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [CI] Isolate layerwise offload memory measurements (vllm-project#6938)

Signed-off-by: andyluo7 <andy.luo@amd.com>

* [Model] Add Cosmos3 mixed W8A8/W8A16 and W4A4/W4A16 denoising (vllm-project#6560)

Signed-off-by: Rahul Steiger <rsteiger@aws-cmh-slurm-1-vscode-04.cm.cluster>
Signed-off-by: Wojciech Kutak <wkutak@nvidia.com>
Co-authored-by: Rahul Steiger <rsteiger@nvidia.com>

* [Test] Use public render_jinja_template in MiniCPM-o native template test (vllm-project#7362)

Signed-off-by: tly <2200895168@qq.com>

* [Bugfix] Fix video prewarm cache retention and cancel-restart delay (vllm-project#7363)

Signed-off-by: psv666 <2693925048@qq.com>

* Cosmos3 action policy improvements (vllm-project#6460)

Signed-off-by: Maciej Bala <mbala@nvidia.com>
Signed-off-by: MaciejBalaNV <mbala@nvidia.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [BugFix][CI] Restore diff-aware source filtering for post-merge L3 (vllm-project#7371)

Signed-off-by: wangyu <410167048@qq.com>

* [Bugfix] Fail when a diffusion LoRA adapter binds no layer (vllm-project#7349)

Signed-off-by: Guangjian <hiro20833@gmail.com>

* [Bugfix] Fix host-memory leak on aborted /v1/images/generations (vllm-project#6462) (vllm-project#6561)

Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Refactor] Declare model-local KV held outside the paged manager (vllm-project#6171)

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>

* [Realtime] Emit current (non-beta) OpenAI audio/transcript event names (vllm-project#7339)

Signed-off-by: Nick Cao <ncao@redhat.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* [Bugfix][Core] Clean up failed HWR atomic metadata writes (vllm-project#6956)

Signed-off-by: BANANASJIM <bananasjim1@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Keep MiniMax-H3 reference audio budgets separate (vllm-project#7281)

Signed-off-by: david6666666 <530634352@qq.com>

* [Bugfix] Fix Helios USP: per-component split for correct sequence parallelism (vllm-project#6930)

Signed-off-by: yancaocn <yancaochn@163.com>
Co-authored-by: yancaocn <yancaochn@163.com>

* [Perf][Diffusion] Optimize HSDP startup via Rank-0 shared weight loading and accelerated LoRA delta computation (vllm-project#7005)

Signed-off-by: samithuang <285365963@qq.com>

* [Example] Migrate HunyuanImage-3.0 to model_extras + shared task examples (vllm-project#5559)

Signed-off-by: suyanli220 <suyanli220@gmail.com>
Signed-off-by: suyan.li <suyan.li@bytedance.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: suyan.li <suyan.li@bytedance.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Model] Avoid scalar synchronizations in GLM-Image preparation (vllm-project#7172)

Signed-off-by: hyw <yuweih205@gmail.com>

* [Model][ERNIE-Image] Delay AdaLN modulation broadcast (vllm-project#7171)

Signed-off-by: hyw <yuweih205@gmail.com>

* [Kernel][MiniMax-H3] Run Q/K RMSNorm-RoPE in one launch (vllm-project#7167)

Signed-off-by: hyw <yuweih205@gmail.com>

* [CI][ROCm] Align AMD image with vLLM 0.29 (vllm-project#7395)

Signed-off-by: andyluo7 <andy.luo@amd.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Add embed_multimodal to MiniCPM-o 4.5 omni LLM class (vllm-project#7384)

Signed-off-by: Guangjian <hiro20833@gmail.com>

* [Model] Add LingBot World Ulysses sequence parallelism (vllm-project#6841)

Signed-off-by: wtz2333 <2955110911@qq.com>
Co-authored-by: Zhou Taichang <tzhouam@connect.ust.hk>

* [Feature][TTS] Add Speech API streaming metrics (vllm-project#6853)

Signed-off-by: XIN GAO <1037396230@qq.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix][Model] Fix FLUX.2 Klein multi-image edit metadata (vllm-project#7430)

Signed-off-by: QI JIA <qi.jia@shengshu.ai>
Co-authored-by: QI JIA <qi.jia@shengshu.ai>
Co-authored-by: Cursor <cursoragent@cursor.com>

* [BugFix] Fix leftovers of the legacy OpenAI realtime API event names (vllm-project#7426)

Signed-off-by: Nick Cao <ncao@redhat.com>
Co-authored-by: Codex <noreply@openai.com>

* [Model] Add Tencent AuK speech generation and editing (encoder + diffusion pipeline) (vllm-project#7385)

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
Co-authored-by: Sy03 <1370724210@qq.com>

* [XPU][Docker] Align XPU image and CI with vLLM v0.29.0 (vllm-project#7441)

Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com>

* [Bugfix] Add explicit error when using CFGP with distilled Cosmos3 models (vllm-project#7427)

Signed-off-by: Maciej Bala <mbala@nvidia.com>

* [Perf][Diffusion] Run MammothModa2 DiT attention through the shared attention layer (vllm-project#7094)

Signed-off-by: MrlixiangWE <mrdanaer@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Bugfix] Give model CLI flags typed owners in the Omni config (vllm-project#7390)

Signed-off-by: Guangjian <hiro20833@gmail.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

* [Bugfix] Require a model for `vllm serve --omni` (fixes vllm-project#4158) (vllm-project#4167)

Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com>

* [Bugfix] Send a downstream terminal chunk when a parked stage ends (vllm-project#6889)

Signed-off-by: psv666 <2693925048@qq.com>

* [NPU] upgrade to v0.29.0 (vllm-project#7433)

Signed-off-by: Weiming Liao <liaowm5@gmail.com>

* [Bugfix][Model][Lance] Support decoded video frames in video editing (vllm-project#5128)

Signed-off-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>
Co-authored-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>

* [Refactor][Diffusion] Remove model-specific names from LoRA and ModelOpt loader defaults (vllm-project#5907)

Signed-off-by: Alicia <115451386+congw729@users.noreply.github.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* Optimize CosyVoice3 Stage1 flow batching (vllm-project#4876)

Signed-off-by: gerayking <399geray@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [3/N] Encode streamed video on the worker with bounded batching (vllm-project#7018)

Signed-off-by: specture724 <specture724@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>

* [Kernel][Boogu-Image] Fuse Q/K RMSNorm + interleaved RoPE via fused_qk_norm_rope (vllm-project#6982)

Signed-off-by: Qihan Kang <rollykanggg@gmail.com>

* [Bugfix][Frontend] Honor output_compression on the image generations route (vllm-project#7447)

Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>

---------

Signed-off-by: Asthenia <asthenia0412@gmail.com>
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: andyluo7 <andy.luo@amd.com>
Signed-off-by: liangmengh <liangmengh@nvidia.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Signed-off-by: KrystalRay <keeleiray@gmail.com>
Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
Signed-off-by: specture724 <specture724@gmail.com>
Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com>
Signed-off-by: chaosansui <zzc15560846421@163.com>
Signed-off-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com>
Signed-off-by: eval-dev <0xe5bca0@gmail.com>
Signed-off-by: eval <74645252+eval-dev@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com>
Signed-off-by: zjli2013 <leezhengjiang@126.com>
Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu>
Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu>
Signed-off-by: ZenAlexa <zimingwang945@gmail.com>
Signed-off-by: BANANASJIM <bananasjim1@gmail.com>
Signed-off-by: Allen Wu <allenwu2795@gmail.com>
Signed-off-by: xutianle <xutianle@fudan.edu.cn>
Signed-off-by: wangyu <410167048@qq.com>
Signed-off-by: natureofnature <wzliu@connect.hku.hk>
Signed-off-by: mglyn <1203789601@qq.com>
Signed-off-by: Sparks-M <41097544+Sparks-M@users.noreply.github.com>
Signed-off-by: chi030303 <106855944+chi030303@users.noreply.github.com>
Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com>
Signed-off-by: NumberWan <wantszkin2003@gmail.com>
Signed-off-by: zouyizhou <zouyizhou@huawei.com>
Signed-off-by: ZhengWG <zwg0606@gmail.com>
Signed-off-by: kunkunblueberry <1833921874@qq.com>
Signed-off-by: Shaun Walsh <shaunwalsh24@gmail.com>
Signed-off-by: mershi <mershi@tencent.com>
Signed-off-by: hyw <yuweih205@gmail.com>
Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com>
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Joshna Medisetty <joshna.medisetty@intel.com>
Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com>
Signed-off-by: Guangjian <hiro20833@gmail.com>
Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Signed-off-by: Gao Han <hgaoaf@connect.ust.hk>
Signed-off-by: Weiming Liao <liaowm5@gmail.com>
Signed-off-by: tzhouam <tzhouam@connect.ust.hk>
Signed-off-by: Zhou Taichang <tzhouam@connect.ust.hk>
Signed-off-by: herotai214 <herotai214@gmail.com>
Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com>
Signed-off-by: Samit <285365963@qq.com>
Signed-off-by: Rahul Steiger <rsteiger@aws-cmh-slurm-1-vscode-04.cm.cluster>
Signed-off-by: Wojciech Kutak <wkutak@nvidia.com>
Signed-off-by: tly <2200895168@qq.com>
Signed-off-by: psv666 <2693925048@qq.com>
Signed-off-by: Maciej Bala <mbala@nvidia.com>
Signed-off-by: MaciejBalaNV <mbala@nvidia.com>
Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com>
Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
Signed-off-by: Nick Cao <ncao@redhat.com>
Signed-off-by: david6666666 <530634352@qq.com>
Signed-off-by: yancaocn <yancaochn@163.com>
Signed-off-by: samithuang <285365963@qq.com>
Signed-off-by: suyanli220 <suyanli220@gmail.com>
Signed-off-by: suyan.li <suyan.li@bytedance.com>
Signed-off-by: wtz2333 <2955110911@qq.com>
Signed-off-by: XIN GAO <1037396230@qq.com>
Signed-off-by: QI JIA <qi.jia@shengshu.ai>
Signed-off-by: MrlixiangWE <mrdanaer@gmail.com>
Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com>
Signed-off-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>
Signed-off-by: Alicia <115451386+congw729@users.noreply.github.com>
Signed-off-by: gerayking <399geray@gmail.com>
Signed-off-by: Qihan Kang <rollykanggg@gmail.com>
Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
Signed-off-by: José Carlos <jose@valendra.tech>
Co-authored-by: Yancy <138764723+Asthenia0412@users.noreply.github.com>
Co-authored-by: Asthenia <asthenia0412@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Co-authored-by: andyluo7 <43718156+andyluo7@users.noreply.github.com>
Co-authored-by: liangmenghuang <liangmengh@nvidia.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Lei Ke <1141466880@qq.com>
Co-authored-by: KrystalRay <keeleiray@gmail.com>
Co-authored-by: Tianyao Wu <54675599+twu3202@users.noreply.github.com>
Co-authored-by: Anjie Hou <149605198+specture724@users.noreply.github.com>
Co-authored-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com>
Co-authored-by: eval <74645252+eval-dev@users.noreply.github.com>
Co-authored-by: boatman <1930807094@qq.com>
Co-authored-by: NancyFyong <88076188+NancyFyong@users.noreply.github.com>
Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com>
Co-authored-by: zhengjia <ZJLi2013@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Rakesh Kariya <83279947+rk9595@users.noreply.github.com>
Co-authored-by: Ziming Wang <125807850+ZenAlexa@users.noreply.github.com>
Co-authored-by: Jim Ban <77719403+BANANASJIM@users.noreply.github.com>
Co-authored-by: Allen Wu <85376543+EchoHayate@users.noreply.github.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: xutianle <24210290017@m.fudan.edu.cn>
Co-authored-by: wangyu <53896905+yenuo26@users.noreply.github.com>
Co-authored-by: NATURE <wzliu@connect.hku.hk>
Co-authored-by: Mu GuanLin <1203789601@qq.com>
Co-authored-by: Sparks <41097544+Sparks-M@users.noreply.github.com>
Co-authored-by: chi030303 <106855944+chi030303@users.noreply.github.com>
Co-authored-by: Bo Li <22713281+bobboli@users.noreply.github.com>
Co-authored-by: NumberWan <wantszkin2003@gmail.com>
Co-authored-by: zyz111222 <zouyizhou@huawei.com>
Co-authored-by: Zheng Wengang <zwg0606@gmail.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>
Co-authored-by: kunkun <72174834+kunkunblueberry@users.noreply.github.com>
Co-authored-by: Shaun Walsh <153730091+Shaun-Walsh@users.noreply.github.com>
Co-authored-by: Nick Cao <ncao@redhat.com>
Co-authored-by: shiyichuan <93317314+CarrotSwordsman@users.noreply.github.com>
Co-authored-by: mershi <mershi@tencent.com>
Co-authored-by: hyw <109567717+yuweih205@users.noreply.github.com>
Co-authored-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com>
Co-authored-by: Sy03 <1370724210@qq.com>
Co-authored-by: Joshna-Medisetty <joshna.medisetty@intel.com>
Co-authored-by: Guangjian Dong <163994576+Hiro208@users.noreply.github.com>
Co-authored-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com>
Co-authored-by: Gao Han <hgaoaf@connect.ust.hk>
Co-authored-by: Weiming Liao <liaowm5@gmail.com>
Co-authored-by: Zhou Taichang <tzhouam@connect.ust.hk>
Co-authored-by: herotai214 <68222888+herotai214@users.noreply.github.com>
Co-authored-by: LHXuuu <xulianhao.xlh@antgroup.com>
Co-authored-by: Samit <285365963@qq.com>
Co-authored-by: wkutak <wkutak@nvidia.com>
Co-authored-by: Rahul Steiger <rsteiger@nvidia.com>
Co-authored-by: tlysanhuo <166924864+tlysanhuo@users.noreply.github.com>
Co-authored-by: psv666 <150513104+psv666@users.noreply.github.com>
Co-authored-by: MaciejBalaNV <mbala@nvidia.com>
Co-authored-by: summer <128961079+zhang-keliang@users.noreply.github.com>
Co-authored-by: Yueqian Lin <70319226+linyueqian@users.noreply.github.com>
Co-authored-by: WeiQing Chen <40507679+david6666666@users.noreply.github.com>
Co-authored-by: Yan Cao <31481315+yancaocn@users.noreply.github.com>
Co-authored-by: yancaocn <yancaochn@163.com>
Co-authored-by: SuyanLi <126558907+suyanli220@users.noreply.github.com>
Co-authored-by: suyan.li <suyan.li@bytedance.com>
Co-authored-by: wtz2333 <2955110911@qq.com>
Co-authored-by: GXIN <37653830+gxxx-hum@users.noreply.github.com>
Co-authored-by: Qi Jia <kuafou@gmail.com>
Co-authored-by: QI JIA <qi.jia@shengshu.ai>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: DanaerLee <mrdanaer@gmail.com>
Co-authored-by: longguo <107740309+abinggo@users.noreply.github.com>
Co-authored-by: junpengw67-max <junpengw67@gmail.com>
Co-authored-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com>
Co-authored-by: Alicia <115451386+congw729@users.noreply.github.com>
Co-authored-by: geray <48796550+gerayking@users.noreply.github.com>
Co-authored-by: KANG Qihan <3149604185@qq.com>
@yenuo26
yenuo26 deleted the trigger branch September 20, 2026 07:54
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
…vllm-project#6597)

Signed-off-by: wangyu <410167048@qq.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CI/CD codes related to changes to CI/CD ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: PR nightly-test label reruns the full nightly pipeline on every push

3 participants