Repository navigation
[CI/Build][ROCm] Stabilize shared AMD test lanes - #7706
Conversation
|
This PR was classified as CI work. CI owner: @yenuo26 @congw729 @NickCao Routing: @yenuo26 via semantic router, CI owner, CODEOWNERS; @congw729 via CODEOWNERS; @NickCao via CODEOWNERS @andyluo7, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
4154c93 to
7ad1abd
Compare
|
Author self-review completed at exact head I checked that:
Validation after rebasing onto current
Exact-head AMD #12134 and CUDA #15453 runs are active; I will treat only terminal jobs at this SHA as runtime evidence. |
| - label: "Diffusion · FlashAttention Test" | ||
| timeout_in_minutes: 20 | ||
| commands: | ||
| - timeout 15m pytest -s -v tests/diffusion/attention/test_flash_attn.py -m 'core_model and cuda and L4 and cards_1' --run-level "core_model" |
There was a problem hiding this comment.
1.please use pytest -s -v tests/diffusion/attention/ -m xxxxxx
2.It's sufficient to have the core_model test cases only in test-ready.yml. There's no need to duplicate them in test-merge.yml.
There was a problem hiding this comment.
Addressed in b237e1ecf: the CUDA and ROCm ready jobs now run pytest -s -v tests/diffusion/attention/ with their existing platform/SKU marker expressions, and I removed both duplicate FlashAttention jobs from the merge pipelines.
| @@ -0,0 +1,149 @@ | |||
| # SPDX-License-Identifier: Apache-2.0 | |||
There was a problem hiding this comment.
I don't think we need to add a new job. Adding a UT to test this job—which is really just some configuration values—would cause the UTs here to expand indefinitely.
There was a problem hiding this comment.
Addressed in b237e1ecf: I removed tests/buildkite/test_rocm_ci_baseline_stability.py entirely. The updated pipeline YAML still passes the repository Buildkite schema hook and the remaining marker/routing checks.
Omni ReviewBot routing recordAssigned Direct under experiment |
7ad1abd to
b237e1e
Compare
|
Self-review update for exact head
The previous AMD #12134 / CUDA #15453 executions are now superseded; fresh exact-head runs are required. |
Omni ReviewBot attempt recordReview attempt ended as stale. |
|
Helper PR with the 90-minute timeout fix is ready for merge into your branch: After merging, re-trigger AMD CI on this PR (build #12137 failed only on Model Executor shard 2/2 at the 60 min cap). |
Build #12168 results — runtime analysisBuild: https://buildkite.com/vllm/vllm-omni-amd-ci/builds/12168 Why total wall time is ~76 min (and shard 2 needs ~80 min)
PR #7706 routing changes work (CosyVoice skips active, FlashAttention out of CPU diffusion lane). Remaining blocker: shard 2 straggler + 60 min cap (needs ~80 min). Immediate unblockPlease merge helper PR with 90 min timeout: andyluo7#1 Medium-term speedups (after green)
Signed-off-by: haic0 haic0@users.noreply.github.com |
Shard 2 straggler — Tier 2 fix bundle pushedHelper branch now has three commits (Tier 1 + Tier 2): git fetch https://github.com/haic0/vllm-omni.git fix/rocm-ci-baseline-stability-timeout-90
git cherry-pick fcee1a4d 0469f886 b77bf3da
git pushCommits:
Files changed in
Signed-off-by: haic0 haic0@users.noreply.github.com |
Signed-off-by: andyluo7 <andy.luo@amd.com>
Signed-off-by: andyluo7 <andy.luo@amd.com>
Signed-off-by: andyluo7 <andy.luo@amd.com>
b237e1e to
8c1f666
Compare
Shard 2/2 of the Simple Model Executor lane needs ~80 minutes on cold AMD workers; bump the Buildkite step timeout from 60 to 90. Signed-off-by: haic0 <haic0@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: andyluo7 <andy.luo@amd.com>
|
Self-review update for exact head
Fresh exact-head AMD #12218, CUDA #15534, and NPU #7258 runs are active. I will attach terminal guard evidence before requesting another ReviewBot pass. |
Signed-off-by: andyluo7 <andy.luo@amd.com>
|
Self-review update for exact head
The previous AMD #12218 / CUDA #15534 runs are superseded. I will use only terminal CI from this exact head as final runtime evidence and will not request another ReviewBot pass before that evidence is available. |
|
@vllm-omni-review-bot please re-review exact head |
vllm-omni-review-bot
left a comment
There was a problem hiding this comment.
Omni ReviewBot review
PR description
This PR stabilizes shared AMD/ROCm CI by fixing misrouted CUDA-only coverage and rebalancing long Simple lanes, instead of raising the old Model Executor wall-clock ceiling. Platform guards and per-test CPU marks keep MAGI-2, CosyVoice GPU numerics, Qwen3-Omni embeddings, and FlashAttention accelerator work off ROCm CPU shards, while ready-only CUDA and NonBlocking ROCm attention jobs cover tests/diffusion/attention/. AMD Model Executor becomes three 90-minute not omni shards plus a dedicated Omni Processor lane, and Simple Diffusion is four-way sharded at 45 minutes in ready and merge.
Change flow
flowchart TD
A["[EXISTING] Shared AMD/CUDA ready and merge lanes"]:::existing
B["[CHANGED] AMD ready/merge and CUDA ready Buildkite YAML"]:::changed
C["[NEW] FlashAttention jobs and Omni Processor lane"]:::new
D["[CHANGED] MAGI-2 / FlashAttn / CosyVoice / Qwen3-embed routing"]:::changed
E["[NEW] test_cosyvoice3_components_cuda.py"]:::new
F["[EXISTING] Model Executor and diffusion pytest consumers"]:::existing
A --> B
B --> C
B --> F
D --> F
E --> F
C --> F
classDef existing fill:#e5e7eb,stroke:#6b7280,color:#111827
classDef changed fill:#fef3c7,stroke:#d97706,color:#451a03,stroke-width:2px
classDef new fill:#dcfce7,stroke:#16a34a,color:#052e16,stroke-width:2px
classDef removed fill:#fee2e2,stroke:#dc2626,color:#450a0a,stroke-width:2px
No actionable findings.
Signed-off-by: andyluo7 <andy.luo@amd.com>
Signed-off-by: andyluo7 <andy.luo@amd.com>
|
Self-review update for exact head
The prior AMD #12224 / CUDA #15541 / NPU #7266 runs are superseded. Fresh exact-head runtime evidence is required before another ReviewBot request. |
Omni ReviewBot triage noteAutomated triage of commit
These are automated triage suggestions only — the final decision belongs to the maintainers. |
|
Exact-head runtime follow-up for
@vllm-omni-review-bot please review the current exact head. |
vllm-omni-review-bot
left a comment
There was a problem hiding this comment.
Omni ReviewBot review
PR description
This PR stabilizes shared AMD/ROCm CI by stopping CUDA-only and accelerator-heavy tests from landing in broad CPU lanes, then rebalancing the remaining long AMD Simple jobs. Platform guards, a CosyVoice CUDA-only module, FlashAttention mark/job splits, Model Executor/Omni/Diffusion sharding with explicit timeouts, a Qwen3-Omni top_k runtime assertion fix, and a CosyVoice-only exit-134 retry address the repeated baseline failures without raising the old two-hour Model Executor ceiling. System-visible effect: shared AMD ready/merge lanes complete within the new shard budgets, while CUDA coverage for the moved paths remains in the ready FlashAttention and Engine & Model Executor jobs.
Change flow
flowchart TD
A["[EXISTING] Shared AMD/CUDA ready and merge lanes"]:::existing
B["[CHANGED] AMD ready/merge YAML, CUDA ready YAML, AMD Jinja template"]:::changed
C["[NEW] Omni Processor lane, ready FlashAttn jobs, CosyVoice exit-134 retry"]:::new
D["[CHANGED] MAGI-2 / FlashAttn / CosyVoice / Qwen embed / Qwen3-Omni routing"]:::changed
E["[NEW] test_cosyvoice3_components_cuda.py"]:::new
F["[EXISTING] Model Executor, diffusion, and E2E pytest consumers"]:::existing
A --> B
B --> C
B --> F
D --> F
E --> F
C --> F
classDef existing fill:#e5e7eb,stroke:#6b7280,color:#111827
classDef changed fill:#fef3c7,stroke:#d97706,color:#451a03,stroke-width:2px
classDef new fill:#dcfce7,stroke:#16a34a,color:#052e16,stroke-width:2px
classDef removed fill:#fee2e2,stroke:#dc2626,color:#450a0a,stroke-width:2px
No actionable findings.
|
|
||
|
|
||
| @hardware_test(res={"cuda": "L4"}, num_cards=1) | ||
| def test_code2wav_streaming_batch_matches_ragged_flow_numerics(monkeypatch): |
There was a problem hiding this comment.
Why turn this test case into a separate script? There's no AMD label on it, so will it still be selected?
Signed-off-by: andyluo7 <andy.luo@amd.com> Signed-off-by: haic0 <haic0@users.noreply.github.com> Co-authored-by: haic0 <haic0@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Matthieu Laneuville <matthieu.laneuville@surf.nl>
Signed-off-by: andyluo7 <andy.luo@amd.com> Signed-off-by: haic0 <haic0@users.noreply.github.com> Co-authored-by: haic0 <haic0@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
Purpose
Stabilize the shared AMD/ROCm pipelines after unrelated or unusually expensive tests were routed into broad CPU lanes and repeatedly produced failures or timeouts.
The repeated baseline symptoms have four code causes:
torch.cuda.is_available(). PyTorch ROCm exposes the CUDA-compatible API, so this CUDA-only test ran on ROCm and failed against the versioned FlashAttention backend.test_flash_attn.pyhad a module-wide CPU marker, so four real accelerator tests ran inside Simple Diffusion and cold-built AITER kernels there.Exact-head AMD build #12224 exposed two additional blocking baseline failures:
top_k=-1, although vLLMSamplingParamscorrectly normalizes that disabled greedy-sampling sentinel to runtimetop_k=0;GPU Hanghardware exceptions on nodegpu8f53. The same node failed build #12214 identically, while equivalent jobs on healthy nodes passed.This PR:
tests/diffusion/attention/, with isolated AITER/Torch build directories on ROCm;omni, while preserving the full 32-batch Omni processor correctness test in a separate blocking 60-minute lane;top_k=0; andThis addresses the root causes behind #7703 rather than raising the old two-hour Model Executor ceiling. The MAGI-2 and CosyVoice changes overlap fixes currently carried by #7605; they are included here because they affect the shared AMD baseline across many PR builds.
The quarantined Stable Audio layerwise-offload job remains nonblocking. Across AMD #12209, #12211, #12213, #12218, and #12224, active worker allocation consistently dropped from about 9.74 GiB to 7.84 GiB, but ROCm allocator pool overhead grew from about 1.77 GiB to 3.78 GiB, so reserved/global peak did not fall by the required 512 MiB. That is a separate offloader/allocator-memory issue; this CI quick fix does not weaken or hide its diagnostic gate.
Test Plan
vLLM Version: N/A (CI/test-routing and assertion-contract change)
vLLM-Omni Commit:
22ed606bffb8bae5a8abf80c6f60275ec2965bceTest Result
Local validation at
22ed606bf:exit_status: 134,limit: 1;git diff --check origin/main...HEADpassed;tests/buildkite: 140 passed, 3 failed. The three unchanged local-only discovery failures are caused by this macOS Python lacking PyTorch and the checkout not being installed for off-repository subprocess discovery.Exact-head CI at
22ed606bfis terminal:gpu923c, so no retry was consumed. All six Model Executor shards, both Omni Processor lanes, all eight Simple Diffusion shards (17m06s-35m13s), and ROCm FlashAttention/AITER passed. Stable Audio reproduced only its quarantined soft failure: active allocation fell from 9.74 GiB to 7.84 GiB, while allocator pool overhead rose from 1.77 GiB to 3.78 GiB, leaving reserved peak 48 MiB worse.Prior build results below belong to superseded head
c6530a1cfand are diagnostic evidence only.Superseded exact-head CI for
c6530a1cf:Diagnostic build #12213 at
a6e5d4e73passed all ready/merge Model Executor shards, both full Omni Processor lanes, the ROCm FlashAttention/AITER lane, and all four merge Simple Diffusion shards. Merge shard 4 passed 1,428 tests in 33m19s with its 45-minute cap. Its only PR-specific failure was the unsharded ready Simple Diffusion lane reaching 60 minutes while still progressing through CPU-heavypi05tests.Historical AMD data confirms that increasing the unsharded timeout is not durable: 228 successful jobs over the prior seven days averaged about 61 minutes, while six completed jobs containing the new
pi05workload averaged 1h45m. Three earlier #7706 runs without that workload averaged 37m36s. The final change therefore gives ready the same proven four-way split as merge instead of retaining a 90-minute serial lane.Reference AMD build #12168 on an older head passed 21/22 jobs. Its only failure was Model Executor shard 2/2 timing out at 60 minutes after 878/1177 tests; that observation motivated the 90-minute cap, three-way sharding, and dedicated Omni lane.