Repository navigation
[CI/Build][ROCm] Stabilize shared AMD test lanes #7706
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Changes from all commits
1a7cf30
118684d
2afb562
6aece08
62d0dad
f52e6e7
a6e5d4e
81b8bcc
c6530a1
cd93f26
22ed606
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -9,26 +9,39 @@ steps: | |
| steps: | ||
| - label: "Simple · Model Executor Test · Shard %N/%t" | ||
| agent_pool: mi300_1 | ||
| parallelism: 2 | ||
| parallelism: 3 | ||
| depends_on: amd-build | ||
| mirror_hardwares: [amdproduction] | ||
| grade: Blocking | ||
| timeout_in_minutes: 90 | ||
|
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. [P1] Exact-head AMD/CUDA guard evidence missing from Test Result The PR body’s Test Result only reports local For this CI-routing change, that does not cover the live guard commands that now own the fixed lanes, including:
Prior AMD #12127/#12168 evidence is on other SHAs / pre-split budgets and cannot substitute. Please attach terminal exact-head results for those selectors before merge.
Collaborator
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Resolved at exact head
Evidence: AMD #12224, CUDA #15541, NPU #7266 (passed). The aggregate AMD/CUDA failures are outside these changed selectors: AMD reproduced the known Qwen3-Omni assertion plus a CosyVoice3 E2E exit 134; CUDA MiniCPM-o perf completed 39/40 requests. |
||
| commands: | ||
| - export VLLM_ROCM_USE_AITER=0 | ||
| # ignore test_teacache_extractors.py because it use rocm gemm kernel from vLLM | ||
| # that is not supported on CPU | ||
| - "pytest -sv tests/model_executor -m 'core_model and cpu' --num-shards=$$BUILDKITE_PARALLEL_JOB_COUNT --shard-id=$$BUILDKITE_PARALLEL_JOB" | ||
| - "pytest -sv tests/model_executor -m 'core_model and cpu and not omni' --num-shards=$$BUILDKITE_PARALLEL_JOB_COUNT --shard-id=$$BUILDKITE_PARALLEL_JOB" | ||
|
|
||
| - label: "Simple · Diffusion Test" | ||
| - label: "Simple · Omni Processor Test" | ||
| agent_pool: mi300_1 | ||
| depends_on: amd-build | ||
| mirror_hardwares: [amdproduction] | ||
| grade: Blocking | ||
| timeout_in_minutes: 60 | ||
| commands: | ||
| - export VLLM_ROCM_USE_AITER=0 | ||
| - "pytest -sv tests/model_executor/models/test_omni_processing.py -m 'core_model and cpu and omni'" | ||
|
|
||
| - label: "Simple · Diffusion Test · Shard %N/%t" | ||
| agent_pool: mi300_1 | ||
| parallelism: 4 | ||
| depends_on: amd-build | ||
| mirror_hardwares: [amdproduction] | ||
| grade: Blocking | ||
| timeout_in_minutes: 45 | ||
| commands: | ||
| - export VLLM_ROCM_USE_AITER=0 | ||
| # ignore test_teacache_extractors.py because it use rocm gemm kernel from vLLM | ||
| # that is not supported on CPU | ||
| - "pytest -sv tests/diffusion -m 'core_model and cpu' --ignore=tests/diffusion/cache/test_teacache_extractors.py" | ||
| - "pytest -sv tests/diffusion -m 'core_model and cpu' --ignore=tests/diffusion/cache/test_teacache_extractors.py --num-shards=$$BUILDKITE_PARALLEL_JOB_COUNT --shard-id=$$BUILDKITE_PARALLEL_JOB" | ||
|
|
||
| - label: "Simple · Engine&Entrypoints Test" | ||
| agent_pool: mi300_1 | ||
|
|
@@ -69,6 +82,27 @@ steps: | |
|
|
||
| - group: ":card_index_dividers: Diffusion Test" | ||
| steps: | ||
| # Keep real FlashAttention execution out of the CPU-marked Simple lane. | ||
| # AITER may need to build several kernels on a cold worker, so run one | ||
| # pytest worker with job-local build directories and a bounded timeout. | ||
| - label: "ROCm · FlashAttention/AITER GPU Coverage" | ||
| agent_pool: mi300_1 | ||
| depends_on: amd-build | ||
| mirror_hardwares: [amdproduction] | ||
| grade: NonBlocking | ||
| timeout_in_minutes: 60 | ||
| commands: | ||
| - export VLLM_ROCM_USE_AITER=1 | ||
| - export AITER_JIT_DIR="/tmp/vllm-omni-aiter-$$BUILDKITE_JOB_ID" | ||
| - export TORCH_EXTENSIONS_DIR="/tmp/vllm-omni-torch-extensions-$$BUILDKITE_JOB_ID" | ||
| - mkdir -p "$$AITER_JIT_DIR" "$$TORCH_EXTENSIONS_DIR" | ||
| - >- | ||
| timeout --signal=TERM --kill-after=2m 55m | ||
| pytest -s -v -n 1 | ||
| tests/diffusion/attention/ | ||
| -m 'core_model and rocm and MI325 and cards_1' | ||
| --run-level "core_model" | ||
|
|
||
| - label: "Diffusion · Batch Test" | ||
| agent_pool: mi300_1 | ||
| depends_on: amd-build | ||
|
|
@@ -299,6 +333,12 @@ steps: | |
| depends_on: amd-build | ||
| mirror_hardwares: [amdproduction] | ||
| grade: Blocking | ||
| # A ROCm hardware hang aborts the interpreter with 134. Let Buildkite | ||
| # create one fresh pod instead of retrying on the poisoned GPU process. | ||
| retry: | ||
| automatic: | ||
| - exit_status: 134 | ||
| limit: 1 | ||
| commands: | ||
| - pytest -s -v tests/e2e/online_serving/test_cosyvoice3_tts_expansion.py -m "slow" --run-level "core_model" | ||
|
|
||
|
|
||
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
[P1] Exact-head AMD/CUDA guard evidence missing for new lane selectors
At this head the CI contract that must pass is the updated AMD ready selectors:
Simple · Model Executor(core_model and cpu and not omni, parallelism 3, 90m),Simple · Omni Processor Test(... and omni),ROCm · FlashAttention/AITER GPU Coverage(tests/diffusion/attention/ -m 'core_model and rocm and MI325 and cards_1'), plus CUDA readyDiffusion · FlashAttention Test(... -m 'core_model and cuda and L4 and cards_1'). The PR Test Result only reports localtests/buildkite/ pre-commit and cites AMD #12127 at a different SHA, while the body itself says exact-head validation is still required; the author head self-review likewise does not show terminal green for those job selectors. Without that evidence, the 90m/3-shard/Omni-split timeout claim is unverified against the failure mode that previously killed shard 2 at ~60m.