Skip to content

[CI/Build][ROCm] Stabilize shared AMD test lanes - #7706

Merged
yenuo26 merged 11 commits into
vllm-project:mainfrom
andyluo7:fix/rocm-ci-baseline-stability
Sep 18, 2026
Merged

yenuo26 merged 11 commits into
vllm-project:mainfrom
andyluo7:fix/rocm-ci-baseline-stability

Conversation

@andyluo7

@andyluo7 andyluo7 commented Sep 17, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

Stabilize the shared AMD/ROCm pipelines after unrelated or unusually expensive tests were routed into broad CPU lanes and repeatedly produced failures or timeouts.

The repeated baseline symptoms have four code causes:

  • MAGI-2 native compile coverage used torch.cuda.is_available(). PyTorch ROCm exposes the CUDA-compatible API, so this CUDA-only test ran on ROCm and failed against the versioned FlashAttention backend.
  • CUDA-only and exhaustive CPU-reference CosyVoice tests ran in the ROCm Model Executor shards. Cold AITER compilation plus the full small-chunk matrix made one shard a persistent straggler.
  • test_flash_attn.py had a module-wide CPU marker, so four real accelerator tests ran inside Simple Diffusion and cold-built AITER kernels there.
  • the NVIDIA-only Qwen3-Omni embedding suite was selected by the ROCm R2-01 marker expression and produced setup errors instead of a platform skip.

Exact-head AMD build #12224 exposed two additional blocking baseline failures:

  • the Qwen3-Omni runtime-config test expected deploy YAML top_k=-1, although vLLM SamplingParams correctly normalizes that disabled greedy-sampling sentinel to runtime top_k=0;
  • CosyVoice3 aborted with exit 134 after repeated ROCm GPU Hang hardware exceptions on node gpu8f53. The same node failed build #12214 identically, while equivalent jobs on healthy nodes passed.

This PR:

  • uses the platform abstraction for CUDA-only MAGI-2 and Qwen3-Omni guards;
  • moves the CUDA-only CosyVoice numerical test into a CUDA-only module and skips the exhaustive small-chunk CPU matrix on ROCm;
  • moves real FlashAttention execution out of the CPU-marked Simple lane;
  • adds bounded ready-only CUDA/L4 and nonblocking ROCm/MI325 attention jobs over tests/diffusion/attention/, with isolated AITER/Torch build directories on ROCm;
  • rebalances AMD Model Executor into three 90-minute shards excluding omni, while preserving the full 32-batch Omni processor correctness test in a separate blocking 60-minute lane;
  • shards AMD Simple Diffusion four ways in both ready and merge, with a 45-minute limit per shard;
  • matches the Qwen3-Omni runtime assertion to vLLM's normalized top_k=0; and
  • gives only the AMD CosyVoice3 E2E step one automatic retry for exit 134, causing Buildkite to create a fresh job/pod after a GPU-hang abort without retrying assertion failures.

This addresses the root causes behind #7703 rather than raising the old two-hour Model Executor ceiling. The MAGI-2 and CosyVoice changes overlap fixes currently carried by #7605; they are included here because they affect the shared AMD baseline across many PR builds.

The quarantined Stable Audio layerwise-offload job remains nonblocking. Across AMD #12209, #12211, #12213, #12218, and #12224, active worker allocation consistently dropped from about 9.74 GiB to 7.84 GiB, but ROCm allocator pool overhead grew from about 1.77 GiB to 3.78 GiB, so reserved/global peak did not fall by the required 512 MiB. That is a separate offloader/allocator-memory issue; this CI quick fix does not weaken or hide its diagnostic gate.

Test Plan

vLLM Version: N/A (CI/test-routing and assertion-contract change)

vLLM-Omni Commit: 22ed606bffb8bae5a8abf80c6f60275ec2965bce

Test Result

Local validation at 22ed606bf:

  • focused AMD pipeline/bootstrap/template tests: 30 passed;
  • rendered the AMD ready template independently and verified the CosyVoice job contains exactly exit_status: 134, limit: 1;
  • all changed-file pre-commit hooks passed, including Buildkite schema, marker, mypy, ruff, and formatting checks;
  • Python syntax checks and git diff --check origin/main...HEAD passed;
  • full tests/buildkite: 140 passed, 3 failed. The three unchanged local-only discovery failures are caused by this macOS Python lacking PyTorch and the checkout not being installed for off-repository subprocess discovery.

Exact-head CI at 22ed606bf is terminal:

  • AMD #12225: passed. Overall runtime was 1h07m22s, including the 17m38s Docker image gate; the GPU-test phase was about 48m53s. Qwen3-Omni passed all 13 tests in 17m58s, including the corrected runtime-config assertion. CosyVoice3 passed all 3 tests directly in 48m46s on gpu923c, so no retry was consumed. All six Model Executor shards, both Omni Processor lanes, all eight Simple Diffusion shards (17m06s-35m13s), and ROCm FlashAttention/AITER passed. Stable Audio reproduced only its quarantined soft failure: active allocation fell from 9.74 GiB to 7.84 GiB, while allocator pool overhead rose from 1.77 GiB to 3.78 GiB, leaving reserved peak 48 MiB worse.
  • CUDA #15542: passed.
  • NPU #7267: passed.
  • Intel #8583: passed.

Prior build results below belong to superseded head c6530a1cf and are diagnostic evidence only.

Superseded exact-head CI for c6530a1cf:

  • AMD #12224: all four ready Simple Diffusion shards passed in 15m40s-36m41s, all four merge shards passed in 15m32s-36m25s, all three ready Model Executor shards passed, the full ready Omni Processor lane passed (12 tests), and ROCm FlashAttention/AITER passed (5 passed, 1 skipped). Its blocking failures were the Qwen normalized-value assertion fixed here and a node-level CosyVoice GPU hang mitigated by the scoped exit-134 retry. The Stable Audio failure was quarantined/soft.
  • CUDA #15541: FlashAttention passed (5 passed, 1 skipped); Engine & Model Executor passed (802 passed, 4 skipped), including the moved CosyVoice numerical GPU test. The aggregate failure is an unrelated MiniCPM-o performance result with 39/40 requests completed.
  • NPU #7266: passed.

Diagnostic build #12213 at a6e5d4e73 passed all ready/merge Model Executor shards, both full Omni Processor lanes, the ROCm FlashAttention/AITER lane, and all four merge Simple Diffusion shards. Merge shard 4 passed 1,428 tests in 33m19s with its 45-minute cap. Its only PR-specific failure was the unsharded ready Simple Diffusion lane reaching 60 minutes while still progressing through CPU-heavy pi05 tests.

Historical AMD data confirms that increasing the unsharded timeout is not durable: 228 successful jobs over the prior seven days averaged about 61 minutes, while six completed jobs containing the new pi05 workload averaged 1h45m. Three earlier #7706 runs without that workload averaged 37m36s. The final change therefore gives ready the same proven four-way split as merge instead of retaining a 90-minute serial lane.

Reference AMD build #12168 on an older head passed 21/22 jobs. Its only failure was Model Executor shard 2/2 timing out at 60 minutes after 878/1177 tests; that observation motivated the 90-minute cap, three-way sharding, and dedicated Omni lane.

@andyluo7 andyluo7 added CI/CD codes related to changes to CI/CD ready label to trigger buildkite CI ROCm PR related to AMD hardware labels Sep 17, 2026 — with ChatGPT Codex Connector
@vllm-omni-review-bot

Copy link
Copy Markdown

This PR was classified as CI work.

CI owner: @yenuo26 @congw729 @NickCao

Routing: @yenuo26 via semantic router, CI owner, CODEOWNERS; @congw729 via CODEOWNERS; @NickCao via CODEOWNERS

@andyluo7, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@andyluo7
andyluo7 force-pushed the fix/rocm-ci-baseline-stability branch from 4154c93 to 7ad1abd Compare September 17, 2026 05:33
@andyluo7 andyluo7 removed the ready label to trigger buildkite CI label Sep 17, 2026
@andyluo7 andyluo7 added the ready label to trigger buildkite CI label Sep 17, 2026 — with ChatGPT Codex Connector

Copy link
Copy Markdown
Collaborator Author

Author self-review completed at exact head 7ad1abd13bb425f16c4d9e73b1b0d36c6c1d454f.

I checked that:

  • the MAGI-2 and CosyVoice exclusions are platform-specific and preserve their CUDA coverage;
  • only the four real accelerator tests in test_flash_attn.py moved out of CPU coverage, while the pure/mock tests retain pytest.mark.cpu;
  • the new CUDA and ROCm jobs select the matching one-GPU markers, and the ROCm AITER job is isolated, single-worker, bounded, and nonblocking;
  • the AMD Model Executor and Simple Diffusion timeout limits are explicit and sized from observed baseline timings rather than extending the prior two-hour timeout;
  • the simultaneous exit-134 failures on gpu9160 are a runner incident and are not hidden by this patch.

Validation after rebasing onto current main:

  • focused routing/guard contracts: 12 passed;
  • all changed-file pre-commit hooks passed, including Buildkite schema validation;
  • git diff --check origin/main...HEAD passed;
  • range-diff confirms the three patches were unchanged by the rebase;
  • full tests/buildkite: 150 passed, with three unchanged local-environment failures because this macOS Python lacks PyTorch and the checkout is not installed for off-root subprocess discovery.

Exact-head AMD #12134 and CUDA #15453 runs are active; I will treat only terminal jobs at this SHA as runtime evidence.

Comment thread .buildkite/cuda/test-merge.yml Outdated
- label: "Diffusion · FlashAttention Test"
timeout_in_minutes: 20
commands:
- timeout 15m pytest -s -v tests/diffusion/attention/test_flash_attn.py -m 'core_model and cuda and L4 and cards_1' --run-level "core_model"

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1.please use pytest -s -v tests/diffusion/attention/ -m xxxxxx
2.It's sufficient to have the core_model test cases only in test-ready.yml. There's no need to duplicate them in test-merge.yml.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in b237e1ecf: the CUDA and ROCm ready jobs now run pytest -s -v tests/diffusion/attention/ with their existing platform/SKU marker expressions, and I removed both duplicate FlashAttention jobs from the merge pipelines.

@@ -0,0 +1,149 @@
# SPDX-License-Identifier: Apache-2.0

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think we need to add a new job. Adding a UT to test this job—which is really just some configuration values—would cause the UTs here to expand indefinitely.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in b237e1ecf: I removed tests/buildkite/test_rocm_ci_baseline_stability.py entirely. The updated pipeline YAML still passes the repository Buildkite schema hook and the remaining marker/routing checks.

@yenuo26

yenuo26 commented Sep 17, 2026

Copy link
Copy Markdown
Collaborator

@vllm-omni-review-bot

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot routing record

Assigned Direct under experiment vllm-omni-strict-5050-20260829.

@andyluo7
andyluo7 force-pushed the fix/rocm-ci-baseline-stability branch from 7ad1abd to b237e1e Compare September 17, 2026 06:33

Copy link
Copy Markdown
Collaborator Author

Self-review update for exact head b237e1ecff04b578a3fa9ac479ad63f281ff20d3 after addressing @yenuo26's feedback:

  • both ready jobs now run the full tests/diffusion/attention/ directory with the platform/SKU marker expression;
  • the duplicate ready-level attention jobs were removed from both merge pipelines;
  • the new Buildkite configuration unit-test file was removed;
  • changed-file pre-commit passed, including Buildkite schema and marker checks;
  • full tests/buildkite is 138 passed with the same three local environment-only discovery failures;
  • git diff --check origin/main...HEAD passed.

The previous AMD #12134 / CUDA #15453 executions are now superseded; fresh exact-head runs are required.

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as stale.

@andyluo7 andyluo7 removed the ready label to trigger buildkite CI label Sep 17, 2026
@andyluo7 andyluo7 added the ready label to trigger buildkite CI label Sep 17, 2026 — with ChatGPT Codex Connector
@haic0

haic0 commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

Helper PR with the 90-minute timeout fix is ready for merge into your branch:

andyluo7#1

After merging, re-trigger AMD CI on this PR (build #12137 failed only on Model Executor shard 2/2 at the 60 min cap).

@haic0

haic0 commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

Build #12168 results — runtime analysis

Build: https://buildkite.com/vllm/vllm-omni-amd-ci/builds/12168
Result: 21/22 passed; shard 2/2 timed out again at 60.7 min (878/1177 tests, ~75% — same failure mode as #12137).

Why total wall time is ~76 min (and shard 2 needs ~80 min)

Job Duration Notes
Model Executor shard 2/2 60.7m (killed) Critical path; ~14.6 tests/min vs 29.9 on shard 1
Model Executor shard 1/2 39.9m 1143 items, passed
Simple · Diffusion Test 40.5m 5698 items, no AITER JIT (PR routing fix ✓)
FlashAttention/AITER lane 30.8m 5 tests; ~24m cold AITER JIT (isolated, nonblocking ✓)
Image build 14.0m Every run

PR #7706 routing changes work (CosyVoice skips active, FlashAttention out of CPU diffusion lane). Remaining blocker: shard 2 straggler + 60 min cap (needs ~80 min).

Immediate unblock

Please merge helper PR with 90 min timeout: andyluo7#1
(or git cherry-pick fcee1a4d), then re-trigger AMD CI.

Medium-term speedups (after green)

  1. Rebalance Model Executor shards (duration-aware or 3 shards) — shard 2 tail has heavy test_omni_processing_correctness
  2. Shard Simple Diffusion on ready like merge (4× parallel)
  3. Pre-warm AITER JIT / cache AMD images ([CI/Build][ROCm] Cache AMD test images in registry #7712)

Signed-off-by: haic0 haic0@users.noreply.github.com

@haic0

haic0 commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

Shard 2 straggler — Tier 2 fix bundle pushed

Helper branch now has three commits (Tier 1 + Tier 2):

https://github.com/haic0/vllm-omni/tree/fix/rocm-ci-baseline-stability-timeout-90

git fetch https://github.com/haic0/vllm-omni.git fix/rocm-ci-baseline-stability-timeout-90
git cherry-pick fcee1a4d 0469f886 b77bf3da
git push

Commits:

  1. fcee1a4d — Model Executor timeout 60→90 min (ready + merge)
  2. 0469f886 — ROCm skipif on test_omni_processing_correctness
  3. b77bf3da — Tier 2 straggler fixes:
    • Dedicated Simple · Omni Processor Test lane (45 min) running test_omni_processing.py
    • Model Executor shards exclude @pytest.mark.omni (not omni marker filter)
    • Model Executor parallelism 2 → 3 (90 min timeout retained)
    • num_batches 32 → 8 in test_omni_processing_correctness
    • ROCm skip for large test_audio_chunk_mask size=1500 (>500)

Files changed in b77bf3da:

  • .buildkite/amd/test-amd-ready.yml
  • .buildkite/amd/test-amd-merge.yml
  • tests/model_executor/models/test_omni_processing.py
  • tests/model_executor/models/minicpmo_4_5/test_audio_chunk_mask.py

Signed-off-by: haic0 haic0@users.noreply.github.com

Signed-off-by: andyluo7 <andy.luo@amd.com>
Signed-off-by: andyluo7 <andy.luo@amd.com>
Signed-off-by: andyluo7 <andy.luo@amd.com>
@andyluo7
andyluo7 force-pushed the fix/rocm-ci-baseline-stability branch from b237e1e to 8c1f666 Compare September 17, 2026 16:01
Shard 2/2 of the Simple Model Executor lane needs ~80 minutes on
cold AMD workers; bump the Buildkite step timeout from 60 to 90.

Signed-off-by: haic0 <haic0@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: andyluo7 <andy.luo@amd.com>
@andyluo7 andyluo7 added merge-test label to trigger buildkite merge test CI and removed merge-test label to trigger buildkite merge test CI labels Sep 17, 2026
@andyluo7

Copy link
Copy Markdown
Collaborator Author

Self-review update for exact head 81b8bcc556d5443f5e6034c2a571d8eb9216dae6:

  • confirmed two prior current-main runs terminated the unsharded ready Simple Diffusion lane at its 60-minute cap while pi05 tests were still progressing;
  • raised only that ready timeout to 90 minutes; the validated merge timeout remains 45 minutes, and selectors/sharding/coverage are unchanged;
  • changed-file pre-commit passed, including Buildkite pipeline validation, and git diff --check passed;
  • full tests/buildkite remains 138 passed with the same three local environment-only package-discovery failures.

Fresh exact-head AMD #12218, CUDA #15534, and NPU #7258 runs are active. I will attach terminal guard evidence before requesting another ReviewBot pass.

Signed-off-by: andyluo7 <andy.luo@amd.com>
@andyluo7

Copy link
Copy Markdown
Collaborator Author

Self-review update for exact head c6530a1cfea199f84ecece730ebe17a44aafade7:

  • reviewed 228 successful AMD Simple Diffusion jobs from the prior seven days and the current pi05 runs; the current unsharded suite averages about 1h45m versus 37m36s on this PR before the new workload;
  • changed the AMD ready Simple Diffusion lane from one unsharded 90-minute job to four 45-minute shards, matching the merge lane that already completed all four shards successfully;
  • preserved the same marker expression, ignore list, blocking grade, and total test coverage; only scheduling/timeout structure changed;
  • added a pipeline contract covering the ready label, four-way parallelism, per-shard timeout, and both pytest shard arguments;
  • focused AMD pipeline/bootstrap tests pass (26 passed), all changed-file pre-commit hooks pass including Buildkite validation, and git diff --check passes;
  • full tests/buildkite is 139 passed, 3 failed, with the same three local environment-only package-discovery failures caused by missing PyTorch and the checkout not being installed for off-root subprocess discovery.

The previous AMD #12218 / CUDA #15534 runs are superseded. I will use only terminal CI from this exact head as final runtime evidence and will not request another ReviewBot pass before that evidence is available.

@andyluo7 andyluo7 added merge-test label to trigger buildkite merge test CI and removed merge-test label to trigger buildkite merge test CI labels Sep 17, 2026
@andyluo7

Copy link
Copy Markdown
Collaborator Author

@vllm-omni-review-bot please re-review exact head c6530a1cfea199f84ecece730ebe17a44aafade7. The requested AMD/CUDA guard evidence is now terminal and is attached in the resolved P1 thread; the PR Test Result section has also been updated.

@vllm-omni-review-bot vllm-omni-review-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Omni ReviewBot review

PR description

This PR stabilizes shared AMD/ROCm CI by fixing misrouted CUDA-only coverage and rebalancing long Simple lanes, instead of raising the old Model Executor wall-clock ceiling. Platform guards and per-test CPU marks keep MAGI-2, CosyVoice GPU numerics, Qwen3-Omni embeddings, and FlashAttention accelerator work off ROCm CPU shards, while ready-only CUDA and NonBlocking ROCm attention jobs cover tests/diffusion/attention/. AMD Model Executor becomes three 90-minute not omni shards plus a dedicated Omni Processor lane, and Simple Diffusion is four-way sharded at 45 minutes in ready and merge.

Change flow

flowchart TD
  A["[EXISTING] Shared AMD/CUDA ready and merge lanes"]:::existing
  B["[CHANGED] AMD ready/merge and CUDA ready Buildkite YAML"]:::changed
  C["[NEW] FlashAttention jobs and Omni Processor lane"]:::new
  D["[CHANGED] MAGI-2 / FlashAttn / CosyVoice / Qwen3-embed routing"]:::changed
  E["[NEW] test_cosyvoice3_components_cuda.py"]:::new
  F["[EXISTING] Model Executor and diffusion pytest consumers"]:::existing
  A --> B
  B --> C
  B --> F
  D --> F
  E --> F
  C --> F
  classDef existing fill:#e5e7eb,stroke:#6b7280,color:#111827
  classDef changed fill:#fef3c7,stroke:#d97706,color:#451a03,stroke-width:2px
  classDef new fill:#dcfce7,stroke:#16a34a,color:#052e16,stroke-width:2px
  classDef removed fill:#fee2e2,stroke:#dc2626,color:#450a0a,stroke-width:2px
Loading

No actionable findings.

Signed-off-by: andyluo7 <andy.luo@amd.com>
Signed-off-by: andyluo7 <andy.luo@amd.com>
@andyluo7

Copy link
Copy Markdown
Collaborator Author

Self-review update for exact head 22ed606bffb8bae5a8abf80c6f60275ec2965bce:

  • confirmed the Qwen3-Omni change only updates the expected post-SamplingParams runtime value (top_k=0); the production deploy YAML remains top_k=-1 and runtime behavior is unchanged;
  • scoped the CosyVoice mitigation to one Buildkite automatic retry only for exit 134, so assertion/test failures are not retried and a GPU-hang abort receives a fresh job/pod;
  • verified both grouped and top-level AMD template paths preserve explicit retry rules, and independently rendered the ready pipeline to confirm the emitted policy;
  • kept the Stable Audio memory failure quarantined and visible because its repeated reserved-memory regression is independent of these blocking quick fixes;
  • focused AMD pipeline/bootstrap/template tests pass (30 passed), changed-file pre-commit passes including Buildkite schema validation, and git diff --check passes;
  • full tests/buildkite is 140 passed, 3 failed, with the same three local macOS/package-discovery failures caused by missing PyTorch and the checkout not being installed off-root.

The prior AMD #12224 / CUDA #15541 / NPU #7266 runs are superseded. Fresh exact-head runtime evidence is required before another ReviewBot request.

@andyluo7 andyluo7 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 17, 2026
@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit 22ed606bffb8 produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@andyluo7

Copy link
Copy Markdown
Collaborator Author

Exact-head runtime follow-up for 22ed606bffb8bae5a8abf80c6f60275ec2965bce:

  • AMD #12225 passed in 1h07m22s overall, including the 17m38s Docker image gate.
  • Qwen3-Omni passed all 13 tests in 17m58s, including the corrected normalized top_k=0 assertion.
  • CosyVoice3 passed all 3 tests directly in 48m46s on gpu923c; no retry was consumed. The generated pipeline preserves the one-retry policy only for exit 134 if a GPU-hang abort recurs.
  • All six Model Executor shards, both Omni Processor lanes, all eight Simple Diffusion shards (17m06s-35m13s), and ROCm FlashAttention/AITER passed.
  • Stable Audio remained an expected quarantined soft failure: active allocation improved by about 1.90 GiB, but ROCm allocator pool overhead increased by about 2.01 GiB, leaving reserved peak 48 MiB worse.
  • CUDA #15542, NPU [Bugfix] Add per-test transcript similarity overrides for higgs-audio-v3 e2e tests #7267, and Intel [Core] Add Cosmos3 attention strategy PoC #8583 also passed exact head.

@vllm-omni-review-bot please review the current exact head.

@vllm-omni-review-bot vllm-omni-review-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Omni ReviewBot review

PR description

This PR stabilizes shared AMD/ROCm CI by stopping CUDA-only and accelerator-heavy tests from landing in broad CPU lanes, then rebalancing the remaining long AMD Simple jobs. Platform guards, a CosyVoice CUDA-only module, FlashAttention mark/job splits, Model Executor/Omni/Diffusion sharding with explicit timeouts, a Qwen3-Omni top_k runtime assertion fix, and a CosyVoice-only exit-134 retry address the repeated baseline failures without raising the old two-hour Model Executor ceiling. System-visible effect: shared AMD ready/merge lanes complete within the new shard budgets, while CUDA coverage for the moved paths remains in the ready FlashAttention and Engine & Model Executor jobs.

Change flow

flowchart TD
  A["[EXISTING] Shared AMD/CUDA ready and merge lanes"]:::existing
  B["[CHANGED] AMD ready/merge YAML, CUDA ready YAML, AMD Jinja template"]:::changed
  C["[NEW] Omni Processor lane, ready FlashAttn jobs, CosyVoice exit-134 retry"]:::new
  D["[CHANGED] MAGI-2 / FlashAttn / CosyVoice / Qwen embed / Qwen3-Omni routing"]:::changed
  E["[NEW] test_cosyvoice3_components_cuda.py"]:::new
  F["[EXISTING] Model Executor, diffusion, and E2E pytest consumers"]:::existing
  A --> B
  B --> C
  B --> F
  D --> F
  E --> F
  C --> F
  classDef existing fill:#e5e7eb,stroke:#6b7280,color:#111827
  classDef changed fill:#fef3c7,stroke:#d97706,color:#451a03,stroke-width:2px
  classDef new fill:#dcfce7,stroke:#16a34a,color:#052e16,stroke-width:2px
  classDef removed fill:#fee2e2,stroke:#dc2626,color:#450a0a,stroke-width:2px
Loading

No actionable findings.



@hardware_test(res={"cuda": "L4"}, num_cards=1)
def test_code2wav_streaming_batch_matches_ragged_flow_numerics(monkeypatch):

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why turn this test case into a separate script? There's no AMD label on it, so will it still be selected?

@yenuo26 yenuo26 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@yenuo26
yenuo26 merged commit 1c7476e into vllm-project:main Sep 18, 2026
9 checks passed
mlaneuville pushed a commit to mlaneuville/vllm-omni that referenced this pull request Sep 22, 2026
Signed-off-by: andyluo7 <andy.luo@amd.com>
Signed-off-by: haic0 <haic0@users.noreply.github.com>
Co-authored-by: haic0 <haic0@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Matthieu Laneuville <matthieu.laneuville@surf.nl>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
Signed-off-by: andyluo7 <andy.luo@amd.com>
Signed-off-by: haic0 <haic0@users.noreply.github.com>
Co-authored-by: haic0 <haic0@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CI/CD codes related to changes to CI/CD merge-test label to trigger buildkite merge test CI ready label to trigger buildkite CI ROCm PR related to AMD hardware

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants