Skip to content

[CI/Build][ROCm] Add tiny diffusion multi-GPU and quantization coverage - #7606

Open
andyluo7 wants to merge 7 commits into
vllm-project:mainfrom
andyluo7:ci/rocm-diffusion-multigpu-quant
Open

andyluo7 wants to merge 7 commits into
vllm-project:mainfrom
andyluo7:ci/rocm-diffusion-multigpu-quant

Conversation

@andyluo7

@andyluo7 andyluo7 commented Sep 15, 2026 •

Copy link
Copy Markdown
Collaborator

Latest coverage repair evidence - September 30, 2026, 23:54 UTC

Current head: 20f02a3d8f2659c790b1904a698f4f1b6f999b74.

FLUX.2 quality on MI300 had no device-profile threshold: the original table contains H100=0.15/B200=0.17 only. Executing the actual resolver confirms MI300 resolves to default and raises. The ROCm quality helper now enforces the strictest existing configured limit (0.15 for this case), preserving CUDA profile selection, scalar thresholds, both eager BF16/FP8 arms and generation parameters. The actual regression test/helper functions pass CUDA H100/B200 and ROCm controls with a stub Torch-version field; GPU quality execution remains pending. All changed-file hooks pass.

At superseded70ca2f1ce, AMD#13067 2-GPU and4-GPU lanes finish exit0nonsoft. All8artifacts in each lane are SHA256verified. JUnit35/26selected pass respectively,zero skips/failures/errors;expected2/4GPUsandimageSHAverified;cleanupPASS withzero leakedprocesses. These are prior-head observations, not current20f head qualification.

Diffusion quality startup repair - September 30, 2026, 23:49 UTC

Current source head: 796429951a0444d675e222afb4a1e04d1e00f77b.

AMD #13066 at superseded f5c7324 exposes a separate FLUX.2-dev quality initialization timeout: its BF16 text encoder/model loading takes 526.11s, then warmup exceeds the default 600s orchestrator budget. Warmup subsequently completes after the timeout. The later agent/container communication loss (exit -7, null agent finish) is a separate incident; the log does not prove OOM.

The quality helper now accepts explicitly configured test initialization limits for both BF16 and FP8 arms. Only the AMD H100-class marker lane exports 1800s overall/1200s stage limits, within its existing 90-minute pytest budget. Ordinary runs retain Omni defaults, both arms retain eager mode and the same attention backend, and all LPIPS thresholds, sampling parameters and marker scopes are unchanged. The lane records its effective limits in environment evidence.

Local validation: 21 focused pipeline/evidence tests and all changed-file pre-commit hooks pass. A probe executing the actual source helper and its regression test verifies default and configured kwargs for both quality arms; this is argument/control-flow evidence, not Torch/GPU runtime. New current-head AMD qualification remains required.

Earlier #13067 at70ca2f1ce passes repaired Simple Other: 4591 tests plus3subtests,68skipped,exit0nonsoft. Its AuK/Breeze failures are covered by#8342/#8201; Engine startup fails after a Hugging Face metadata HTTP disconnect, not a diffusion assertion. #13066 legacy-L4 quantization passes6/2skipped with7SHA256verified artifacts,exactimage1/1GPUandcleanprocess teardown; this is superseded-head evidence.

Current status - September 30, 2026, 23:24 UTC

Head: 70ca2f1ce3dee78811d35ba7578b4de7cc58f655. Rebased onto main 27a6321a6311870fe68c2c9c13b694f2612ff779. All four added lanes remain NonBlocking and require fresh runtime qualification at this head.

Two new repairs follow AMD #13066, which ran superseded f5c7324f05819a59b7ff502b6b7dd364a4aa292d:

  • 175e7a3b2: each lane now validates JUnit reports and retains pytest-result.txt with selected, passed, skipped, failed and executed counts. Missing, malformed, empty, skipped-only, failed and inconsistent reports fail the job. This closes the explicit execution-evidence gap; the native runner also retains pytest failure exit codes.
  • The offline helper comment accidentally matched the marker checker's textual pattern, causing its negative isolation test to fail. The comment is corrected. Loader call contracts now cover default 600s/300s and overridden 900s/600s initialization limits, preserving native checkpoint path and model-class assertions.

Local validation: 21 focused pipeline/evidence tests passed; 3 marker-checker tests passed; all changed-file pre-commit hooks, including mypy and Buildkite schema, passed; compilation and git diff --check passed. A bounded probe executing the actual offline/online runner and contract function bodies verifies all four model/timeout combinations using dependency stubs. That probe checks arguments and control flow, not Torch or GPU runtime. Full loader tests and shard-order execution require CI.

The prior #13066 generic job has three attributable failures repaired above. Its AuK graph failures match #8342 and Breeze RNG failures match #8201; these baseline repairs are tracked separately. Multi-GPU and quantization target results at the old head remain diagnostic and cannot qualify this new head.

Historical purpose and evidence

Everything below describes superseded revisions unless stated otherwise.

Purpose

Advance #5731 items 4, 6, and 7 with non-blocking ROCm nightly coverage for tiny diffusion multi-GPU paths and the current diffusion quantization scopes.

This PR is unstacked and rebased directly onto main at c32aaeb2d. The current exact head is 7e4cc88d50851c633380cbee45ad6a0934a1d15f.

What changed

This PR adds four AMD nightly jobs:

  • tiny diffusion on two MI300-family GPUs;
  • tiny diffusion on four MI300-family GPUs;
  • the current H100-class quantization marker scope on one MI300-family GPU; and
  • the current legacy-L4 quantization marker scope on one MI300-family GPU, including pinned vllm-gguf-plugin==0.0.4 setup.

The CUDA H100, B200, and L4 marks are preserved as test-selection scopes. They do not request those NVIDIA devices, and the legacy-L4 row is not mapped to MI250X.

Each job validates the ROCm build and visible GPU count, runs pytest once, fails closed on zero collection, and retains environment, JUnit, log, summary, runtime, and cleanup evidence. All four jobs remain NonBlocking during burn-in.

The diagnostic AMD runs exposed portability and CI-runtime issues that are fixed here:

  • cleanup evidence ignores the evidence helper's own process;
  • the LTX-2 base vocoder disables the cuDNN/MIOpen path and runs in FP32 on ROCm, restoring its original dtype afterward;
  • tiny multi-GPU initialization limits are configurable and raised to 900s/600s for these lanes;
  • the MAGI-2 compile test now requires the actual CUDA platform rather than the shared torch.cuda namespace;
  • CUDA-only Z-Image and LTX-2 peak-memory comparisons skip on ROCm, while FP8 generation and quantization-quality gates remain active;
  • GGUF generation and cosine-similarity checks remain active on ROCm, with only the CUDA-specific memory-reduction assertion omitted; and
  • the H100-class quality lane uses Torch SDPA on ROCm so first-use AITER JIT cannot consume Omni's 600-second startup budget. Its bounded pytest/job limits are 90/100 minutes, leaving ten minutes for evidence and cleanup.

The attention-backend override is limited to the ROCm quality lane. Each quality case still compares BF16 and FP8 under the same backend, and CUDA behavior and marker selection are unchanged.

Local validation

pytest -o addopts='' --confcutdir=tests/buildkite -q \
  tests/buildkite/test_amd_nightly_diffusion_parity.py
python3 -m py_compile \
  .buildkite/amd/scripts/rocm_ci_evidence.py \
  tests/buildkite/test_amd_nightly_diffusion_parity.py \
  tests/diffusion/models/ltx2/test_ltx2_vae.py \
  tests/diffusion/models/magi2/test_native_compile_cuda.py \
  tests/diffusion/quantization/test_gguf_diffusion.py \
  tests/diffusion/quantization/test_quantization_fp8.py \
  tests/model_tests/diffusion/test_common_offline.py \
  vllm_omni/diffusion/models/ltx2/ltx2_runtime.py
pre-commit run --files \
  .buildkite/amd/test-amd-nightly.yml \
  .buildkite/amd/scripts/rocm_ci_evidence.py \
  tests/buildkite/test_amd_nightly_diffusion_parity.py \
  tests/diffusion/models/ltx2/test_ltx2_vae.py \
  tests/diffusion/models/magi2/test_native_compile_cuda.py \
  tests/diffusion/quantization/test_gguf_diffusion.py \
  tests/diffusion/quantization/test_quantization_fp8.py \
  tests/model_tests/diffusion/test_common_offline.py \
  vllm_omni/diffusion/models/ltx2/ltx2_runtime.py
git diff --check origin/main..HEAD

Results:

  • focused AMD nightly contract: 11 passed;
  • all changed Python files compiled successfully;
  • all changed-file pre-commit hooks passed, including Buildkite schema validation and mypy;
  • git diff --check passed; and
  • git range-diff shows both original PR patches unchanged by the rebase; the latest commit contains only the H100-class lane repair and its contract assertion.

The local macOS environment does not provide ROCm/PyTorch GPU runtime validation.

Hardware evidence

AMD build #11942 is diagnostic only because it ran the superseded stacked head dde475e60.

AMD build #12102 ran the superseded unstacked head fdc4d05a1 and established:

  • 2-GPU tiny diffusion passed 35 tests with zero failures/errors/skips in 1,211 seconds, exact-image and 2/2-GPU evidence, and clean teardown;
  • 4-GPU tiny diffusion passed 20 tests with zero failures/errors/skips in 1,091 seconds, exact-image and 4/4-GPU evidence, and clean teardown;
  • legacy-L4 quantization collected 8 tests, passed 6, and made 2 expected platform skips in 2,538 seconds, with zero failures/errors, exact-image and 1/1-GPU evidence, all seven artifacts, and status=PASS leaked_processes=0;
  • H100-class quantization selected 10 tests and passed both Bagel generation cases, but Z-Image initialization spent the remainder of Omni's 600-second startup budget in first-use AITER attention JIT. Its deferred shutdown left model-parallel state live, the next case failed closed, and the job reached the 60-minute pytest timeout with exit 124, no JUnit, and clean outer teardown.

That failure is not counted as qualification. The current head moves only this ROCm quality lane to Torch SDPA and a 90/100-minute bounded budget. Fresh exact-head AMD and CUDA validation is required before promotion.

The acceptance gate for each new AMD lane remains: exit 0, no soft failure, nonzero execution, readable log/JUnit/environment/summary/cleanup artifacts, exact image SHA and expected visible GPU count, and status=PASS leaked_processes=0.

AI assistance: Used Codex to diagnose the superseded AMD run, unstack and rebase the branch, implement the narrow ROCm portability and runtime fixes, run local validation, and maintain this description. I reviewed the changes and validation results.

@andyluo7 andyluo7 added ROCm PR related to AMD hardware CI/CD codes related to changes to CI/CD amd-test Used to trigger AMD CI separately. nightly-test label to trigger buildkite nightly test CI labels Sep 15, 2026
@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/quantization.md, docs/design/module/diffusion/index.md.

Module owners: @david6666666 @Isotr0py @lishunyang12

Routing: @david6666666 via module named in the PR description; @Isotr0py via module named in the PR description; @lishunyang12 via module named in the PR description

@andyluo7, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@andyluo7

andyluo7 commented Sep 15, 2026 •

Copy link
Copy Markdown
Collaborator Author

Author self-review completed at exact head a2816228a83b6040d73c38ca0bee6c3c31445b80.

I reviewed commit a2816228 separately from its #7604 parent and checked that:

  • all four new jobs remain NonBlocking and use the intended 2-GPU, 4-GPU, and 1-GPU MI300 queues;
  • their pytest paths, CUDA marker scopes, and run levels match the current CUDA nightly definitions;
  • each job executes pytest once, fails on zero collection, validates the visible ROCm device count, and retains environment, runtime, JUnit, log, summary, and cleanup evidence;
  • the legacy-L4 scope pins vllm-gguf-plugin==0.0.4; and
  • no CUDA pipeline or runtime implementation is changed.

I reran the full stacked contract set: 18 passed, the evidence helper compiled, and git diff --check passed.

First exact-head AMD burn-in: build #11942 is still running. The four-GPU job collected 9 tests and completed with 7 passed plus failures in Flux2Klein startup and LTX-2 ROCm execution; its artifacts and cleanup verdict were retained. The H100-class quantization job collected 10 tests but lost its Kubernetes container while starting the LTX-2 FP8 memory test, with no artifacts. These are non-blocking runtime findings, not passing qualification evidence; no threshold or test-selection weakening has been made.

@andyluo7 andyluo7 added cuda-test Used to trigger vllm-omni cuda CI separately. nightly-test label to trigger buildkite nightly test CI and removed amd-test Used to trigger AMD CI separately. nightly-test label to trigger buildkite nightly test CI labels Sep 16, 2026
@andyluo7
andyluo7 force-pushed the ci/rocm-diffusion-multigpu-quant branch from a281622 to dde475e Compare September 16, 2026 05:28
@andyluo7 andyluo7 added nightly-test label to trigger buildkite nightly test CI and removed nightly-test label to trigger buildkite nightly test CI labels Sep 16, 2026
@andyluo7

Copy link
Copy Markdown
Collaborator Author

Stack refresh self-review completed at exact head dde475e60b4648c508981e48ad531f284d984c65.

I replayed the isolated #7606 commit unchanged onto #7604 repaired head 2ea3ef0c. git range-diff shows the #7606 patch is unchanged. The refreshed stacked contract suite passes (21 passed), the evidence helper compiles, all changed-file pre-commit hooks pass, and git diff --check passes.

AMD build #11942 remains diagnostic only: all four new jobs failed and soft-failed, so its aggregate green status is not qualification evidence. CUDA build #15322 belongs to the superseded a2816228 head and is also not qualifying evidence. A fresh exact-head CUDA nightly run is being triggered for dde475e60; the PR remains stacked until #7604 merges.

@andyluo7

Copy link
Copy Markdown
Collaborator Author

Exact-head CUDA update for dde475e60b4648c508981e48ad531f284d984c65:

  • CUDA nightly #15323 passed.
  • All 7 rendered script jobs finished with exit status 0 and soft_failed=false.
  • H100 single-GPU: 20 passed, 3 skipped, 693 deselected.
  • H100 two-GPU: 27 passed, 9 skipped, 680 deselected.
  • 4xH100 multi-replica startup: 3 passed.
  • Build wheels, pre-commit, DCO, and Read the Docs are green; GitHub reports the branch mergeable with no unresolved review threads.

Scope note: the diff-aware CUDA nightly exercised the inherited Qwen3-Omni stack rather than the four new AMD-only diffusion multi-GPU/quantization jobs. After #7604 merges, this stacked PR still needs its planned rebase onto current main and fresh exact-head validation. No merge has been performed.

@andyluo7
andyluo7 force-pushed the ci/rocm-diffusion-multigpu-quant branch from dde475e to fdc4d05 Compare September 16, 2026 21:04
@andyluo7
andyluo7 requested a review from wtomin as a code owner September 16, 2026 21:04
@andyluo7 andyluo7 added ready label to trigger buildkite CI amd-test Used to trigger AMD CI separately. and removed ready label to trigger buildkite CI nightly-test label to trigger buildkite nightly test CI amd-test Used to trigger AMD CI separately. cuda-test Used to trigger vllm-omni cuda CI separately. labels Sep 16, 2026
@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: no human activity for 7 days

@andyluo7 this pull request has had no human commit, comment or review since 2026-09-17. Please confirm the current plan and next step. The author or a maintainer decides whether to change the PR state.

To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline.

@hsliuustc0106 hsliuustc0106 added the quantization Code related to quantization label Sep 25, 2026
@hsliuustc0106 hsliuustc0106 added the high priority high priority issue, needs to be done asap label Sep 30, 2026 — with ChatGPT Codex Connector
Signed-off-by: andyluo7 <andy.luo@amd.com>
Signed-off-by: andyluo7 <andy.luo@amd.com>
@andyluo7
andyluo7 force-pushed the ci/rocm-diffusion-multigpu-quant branch from 7e4cc88 to f5c7324 Compare September 30, 2026 19:04
@andyluo7 andyluo7 added ready label to trigger buildkite CI and removed cuda-test Used to trigger vllm-omni cuda CI separately. ready label to trigger buildkite CI labels Sep 30, 2026
@andyluo7 andyluo7 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 30, 2026
Signed-off-by: andyluo7 <andy.luo@amd.com>
@andyluo7 andyluo7 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 30, 2026
Signed-off-by: andyluo7 <andy.luo@amd.com>
@andyluo7 andyluo7 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 30, 2026
@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit 20f02a3d8f26 produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@hsliuustc0106

Copy link
Copy Markdown
Collaborator

@andyluo7 this PR is labeled ready + high priority, but CI is failing on the latest commit:

Could you please take a look and push an update to get CI green? Once the checks pass we can proceed with review/merge. Thanks!

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot routing record

Assigned Strict on zcode (GLM-5.3-Flash) under experiment fleet-strict-cursor-grok46-zcode-glm53flash-5050-c5-z10-20261002.

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: e851ae95-d6f0-4047-935d-8292ccbc68f1) — check zcode login and the CLI version; retrying strict/zcode/GLM-5.3-Flash in 120s (try ).

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: f21e5af7-f938-4363-8f6c-833b55b3d397) — check zcode login and the CLI version; retrying strict/zcode/GLM-5.3-Flash in 600s (try ).

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: cd979b58-d53f-4eef-85a9-f2053c58e875) — check zcode login and the CLI version; falling back to direct/cursor/auto).

@vllm-omni-review-bot vllm-omni-review-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Omni ReviewBot review

0 actionable finding(s).

CI at 20f02a3d8f26 (2026-10-10T09:58:40.057591+00:00): verification incomplete; required-check status is unknown. Observed Buildkite: buildkite/vllm-omni-npu-ci (failed), buildkite/vllm-omni (passed), buildkite/vllm-omni-amd-ci (failed), and 1 more.

Note: The assigned review arm strict/zcode/GLM-5.3-Flash could not complete this review, so it was produced by the fallback arm direct/cursor/auto. It is excluded from the routing experiment.

Full review analysis

PR description

This PR adds four non-blocking AMD nightly jobs that run the existing CUDA tiny-diffusion multi-GPU and diffusion-quantization marker scopes on MI300-family GPUs. A new evidence helper records the ROCm environment, checks that pytest produced a non-empty JUnit result, and fails the lane when leftover processes remain after teardown. The same change raises startup limits for those lanes, pins the H100-class quality lane to Torch SDPA with a 90/100-minute budget, runs the LTX-2 base vocoder in FP32 with MIOpen disabled on ROCm, and keeps CUDA peak-memory assertions off the ROCm path while image, cosine, and LPIPS checks stay active.

Change flow

flowchart LR
  cudaScopes["[EXISTING] CUDA nightly marker scopes"]:::existing
  amdJobs["[CHANGED] AMD nightly diffusion jobs"]:::changed
  evidence["[NEW] ROCm evidence helper"]:::new
  runtime["[CHANGED] LTX vocoder and quality gates"]:::changed
  lanes["[NEW] NonBlocking MI300 burn-in lanes"]:::new
  cudaScopes --> amdJobs
  runtime --> amdJobs
  amdJobs --> evidence
  evidence --> lanes
  amdJobs --> lanes
  classDef existing fill:#e5e7eb,stroke:#6b7280,color:#111827
  classDef changed fill:#fef3c7,stroke:#d97706,color:#451a03,stroke-width:2px
  classDef new fill:#dcfce7,stroke:#16a34a,color:#052e16,stroke-width:2px
  classDef removed fill:#fee2e2,stroke:#dc2626,color:#450a0a,stroke-width:2px
Loading

No actionable findings.


🤖 This review was generated by InferMatrix Copilot, an open-source repo-maintenance agent for PR review, CI debugging and issue triage. Try it on your own repo, and ⭐ star it if it helped!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CI/CD codes related to changes to CI/CD high priority high priority issue, needs to be done asap nightly-test label to trigger buildkite nightly test CI quantization Code related to quantization ready label to trigger buildkite CI ROCm PR related to AMD hardware

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants