Repository navigation
Conversation
|
This PR appears to belong to: docs/design/module/quantization.md, docs/design/module/diffusion/index.md. Module owners: @david6666666 @Isotr0py @lishunyang12 Routing: @david6666666 via module named in the PR description; @Isotr0py via module named in the PR description; @lishunyang12 via module named in the PR description @andyluo7, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
|
Author self-review completed at exact head I reviewed commit
I reran the full stacked contract set: 18 passed, the evidence helper compiled, and First exact-head AMD burn-in: build #11942 is still running. The four-GPU job collected 9 tests and completed with 7 passed plus failures in Flux2Klein startup and LTX-2 ROCm execution; its artifacts and cleanup verdict were retained. The H100-class quantization job collected 10 tests but lost its Kubernetes container while starting the LTX-2 FP8 memory test, with no artifacts. These are non-blocking runtime findings, not passing qualification evidence; no threshold or test-selection weakening has been made. |
a281622 to
dde475e
Compare
|
Stack refresh self-review completed at exact head I replayed the isolated #7606 commit unchanged onto #7604 repaired head AMD build #11942 remains diagnostic only: all four new jobs failed and soft-failed, so its aggregate green status is not qualification evidence. CUDA build #15322 belongs to the superseded |
|
Exact-head CUDA update for
Scope note: the diff-aware CUDA nightly exercised the inherited Qwen3-Omni stack rather than the four new AMD-only diffusion multi-GPU/quantization jobs. After #7604 merges, this stacked PR still needs its planned rebase onto current |
dde475e to
fdc4d05
Compare
Omni ReviewBot: no human activity for 7 days@andyluo7 this pull request has had no human commit, comment or review since 2026-09-17. Please confirm the current plan and next step. The author or a maintainer decides whether to change the PR state. To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline. |
Signed-off-by: andyluo7 <andy.luo@amd.com>
Signed-off-by: andyluo7 <andy.luo@amd.com>
Signed-off-by: andyluo7 <andy.luo@amd.com>
7e4cc88 to
f5c7324
Compare
Signed-off-by: andyluo7 <andy.luo@amd.com>
Signed-off-by: andyluo7 <andy.luo@amd.com>
Signed-off-by: andyluo7 <andy.luo@amd.com>
Signed-off-by: andyluo7 <andy.luo@amd.com>
Omni ReviewBot triage noteAutomated triage of commit
These are automated triage suggestions only — the final decision belongs to the maintainers. |
|
@andyluo7 this PR is labeled Could you please take a look and push an update to get CI green? Once the checks pass we can proceed with review/merge. Thanks! |
Omni ReviewBot routing recordAssigned Strict on zcode (GLM-5.3-Flash) under experiment |
Omni ReviewBot attempt recordReview attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: e851ae95-d6f0-4047-935d-8292ccbc68f1) — check |
Omni ReviewBot attempt recordReview attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: f21e5af7-f938-4363-8f6c-833b55b3d397) — check |
Omni ReviewBot attempt recordReview attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: cd979b58-d53f-4eef-85a9-f2053c58e875) — check |
vllm-omni-review-bot
left a comment
There was a problem hiding this comment.
Omni ReviewBot review
0 actionable finding(s).
CI at
20f02a3d8f26(2026-10-10T09:58:40.057591+00:00): verification incomplete; required-check status is unknown. Observed Buildkite:buildkite/vllm-omni-npu-ci(failed),buildkite/vllm-omni(passed),buildkite/vllm-omni-amd-ci(failed), and 1 more.
Note: The assigned review arm
strict/zcode/GLM-5.3-Flashcould not complete this review, so it was produced by the fallback armdirect/cursor/auto. It is excluded from the routing experiment.
Full review analysis
PR description
This PR adds four non-blocking AMD nightly jobs that run the existing CUDA tiny-diffusion multi-GPU and diffusion-quantization marker scopes on MI300-family GPUs. A new evidence helper records the ROCm environment, checks that pytest produced a non-empty JUnit result, and fails the lane when leftover processes remain after teardown. The same change raises startup limits for those lanes, pins the H100-class quality lane to Torch SDPA with a 90/100-minute budget, runs the LTX-2 base vocoder in FP32 with MIOpen disabled on ROCm, and keeps CUDA peak-memory assertions off the ROCm path while image, cosine, and LPIPS checks stay active.
Change flow
flowchart LR
cudaScopes["[EXISTING] CUDA nightly marker scopes"]:::existing
amdJobs["[CHANGED] AMD nightly diffusion jobs"]:::changed
evidence["[NEW] ROCm evidence helper"]:::new
runtime["[CHANGED] LTX vocoder and quality gates"]:::changed
lanes["[NEW] NonBlocking MI300 burn-in lanes"]:::new
cudaScopes --> amdJobs
runtime --> amdJobs
amdJobs --> evidence
evidence --> lanes
amdJobs --> lanes
classDef existing fill:#e5e7eb,stroke:#6b7280,color:#111827
classDef changed fill:#fef3c7,stroke:#d97706,color:#451a03,stroke-width:2px
classDef new fill:#dcfce7,stroke:#16a34a,color:#052e16,stroke-width:2px
classDef removed fill:#fee2e2,stroke:#dc2626,color:#450a0a,stroke-width:2px
No actionable findings.
🤖 This review was generated by InferMatrix Copilot, an open-source repo-maintenance agent for PR review, CI debugging and issue triage. Try it on your own repo, and ⭐ star it if it helped!
Latest coverage repair evidence - September 30, 2026, 23:54 UTC
Current head:
20f02a3d8f2659c790b1904a698f4f1b6f999b74.FLUX.2 quality on MI300 had no device-profile threshold: the original table contains H100=0.15/B200=0.17 only. Executing the actual resolver confirms MI300 resolves to default and raises. The ROCm quality helper now enforces the strictest existing configured limit (0.15 for this case), preserving CUDA profile selection, scalar thresholds, both eager BF16/FP8 arms and generation parameters. The actual regression test/helper functions pass CUDA H100/B200 and ROCm controls with a stub Torch-version field; GPU quality execution remains pending. All changed-file hooks pass.
At superseded70ca2f1ce, AMD#13067 2-GPU and4-GPU lanes finish exit0nonsoft. All8artifacts in each lane are SHA256verified. JUnit35/26selected pass respectively,zero skips/failures/errors;expected2/4GPUsandimageSHAverified;cleanupPASS withzero leakedprocesses. These are prior-head observations, not current20f head qualification.
Diffusion quality startup repair - September 30, 2026, 23:49 UTC
Current source head:
796429951a0444d675e222afb4a1e04d1e00f77b.AMD #13066 at superseded f5c7324 exposes a separate FLUX.2-dev quality initialization timeout: its BF16 text encoder/model loading takes 526.11s, then warmup exceeds the default 600s orchestrator budget. Warmup subsequently completes after the timeout. The later agent/container communication loss (exit -7, null agent finish) is a separate incident; the log does not prove OOM.
The quality helper now accepts explicitly configured test initialization limits for both BF16 and FP8 arms. Only the AMD H100-class marker lane exports 1800s overall/1200s stage limits, within its existing 90-minute pytest budget. Ordinary runs retain Omni defaults, both arms retain eager mode and the same attention backend, and all LPIPS thresholds, sampling parameters and marker scopes are unchanged. The lane records its effective limits in environment evidence.
Local validation: 21 focused pipeline/evidence tests and all changed-file pre-commit hooks pass. A probe executing the actual source helper and its regression test verifies default and configured kwargs for both quality arms; this is argument/control-flow evidence, not Torch/GPU runtime. New current-head AMD qualification remains required.
Earlier #13067 at70ca2f1ce passes repaired Simple Other: 4591 tests plus3subtests,68skipped,exit0nonsoft. Its AuK/Breeze failures are covered by#8342/#8201; Engine startup fails after a Hugging Face metadata HTTP disconnect, not a diffusion assertion. #13066 legacy-L4 quantization passes6/2skipped with7SHA256verified artifacts,exactimage1/1GPUandcleanprocess teardown; this is superseded-head evidence.
Current status - September 30, 2026, 23:24 UTC
Head:
70ca2f1ce3dee78811d35ba7578b4de7cc58f655. Rebased onto main27a6321a6311870fe68c2c9c13b694f2612ff779. All four added lanes remain NonBlocking and require fresh runtime qualification at this head.Two new repairs follow AMD #13066, which ran superseded
f5c7324f05819a59b7ff502b6b7dd364a4aa292d:175e7a3b2: each lane now validates JUnit reports and retainspytest-result.txtwith selected, passed, skipped, failed and executed counts. Missing, malformed, empty, skipped-only, failed and inconsistent reports fail the job. This closes the explicit execution-evidence gap; the native runner also retains pytest failure exit codes.Local validation: 21 focused pipeline/evidence tests passed; 3 marker-checker tests passed; all changed-file pre-commit hooks, including mypy and Buildkite schema, passed; compilation and
git diff --checkpassed. A bounded probe executing the actual offline/online runner and contract function bodies verifies all four model/timeout combinations using dependency stubs. That probe checks arguments and control flow, not Torch or GPU runtime. Full loader tests and shard-order execution require CI.The prior #13066 generic job has three attributable failures repaired above. Its AuK graph failures match #8342 and Breeze RNG failures match #8201; these baseline repairs are tracked separately. Multi-GPU and quantization target results at the old head remain diagnostic and cannot qualify this new head.
Historical purpose and evidence
Everything below describes superseded revisions unless stated otherwise.
Purpose
Advance #5731 items 4, 6, and 7 with non-blocking ROCm nightly coverage for tiny diffusion multi-GPU paths and the current diffusion quantization scopes.
This PR is unstacked and rebased directly onto
mainatc32aaeb2d. The current exact head is7e4cc88d50851c633380cbee45ad6a0934a1d15f.What changed
This PR adds four AMD nightly jobs:
vllm-gguf-plugin==0.0.4setup.The CUDA
H100,B200, andL4marks are preserved as test-selection scopes. They do not request those NVIDIA devices, and the legacy-L4 row is not mapped to MI250X.Each job validates the ROCm build and visible GPU count, runs pytest once, fails closed on zero collection, and retains environment, JUnit, log, summary, runtime, and cleanup evidence. All four jobs remain
NonBlockingduring burn-in.The diagnostic AMD runs exposed portability and CI-runtime issues that are fixed here:
torch.cudanamespace;The attention-backend override is limited to the ROCm quality lane. Each quality case still compares BF16 and FP8 under the same backend, and CUDA behavior and marker selection are unchanged.
Local validation
pytest -o addopts='' --confcutdir=tests/buildkite -q \ tests/buildkite/test_amd_nightly_diffusion_parity.py python3 -m py_compile \ .buildkite/amd/scripts/rocm_ci_evidence.py \ tests/buildkite/test_amd_nightly_diffusion_parity.py \ tests/diffusion/models/ltx2/test_ltx2_vae.py \ tests/diffusion/models/magi2/test_native_compile_cuda.py \ tests/diffusion/quantization/test_gguf_diffusion.py \ tests/diffusion/quantization/test_quantization_fp8.py \ tests/model_tests/diffusion/test_common_offline.py \ vllm_omni/diffusion/models/ltx2/ltx2_runtime.py pre-commit run --files \ .buildkite/amd/test-amd-nightly.yml \ .buildkite/amd/scripts/rocm_ci_evidence.py \ tests/buildkite/test_amd_nightly_diffusion_parity.py \ tests/diffusion/models/ltx2/test_ltx2_vae.py \ tests/diffusion/models/magi2/test_native_compile_cuda.py \ tests/diffusion/quantization/test_gguf_diffusion.py \ tests/diffusion/quantization/test_quantization_fp8.py \ tests/model_tests/diffusion/test_common_offline.py \ vllm_omni/diffusion/models/ltx2/ltx2_runtime.py git diff --check origin/main..HEADResults:
11 passed;git diff --checkpassed; andgit range-diffshows both original PR patches unchanged by the rebase; the latest commit contains only the H100-class lane repair and its contract assertion.The local macOS environment does not provide ROCm/PyTorch GPU runtime validation.
Hardware evidence
AMD build #11942 is diagnostic only because it ran the superseded stacked head
dde475e60.AMD build #12102 ran the superseded unstacked head
fdc4d05a1and established:status=PASS leaked_processes=0;That failure is not counted as qualification. The current head moves only this ROCm quality lane to Torch SDPA and a 90/100-minute bounded budget. Fresh exact-head AMD and CUDA validation is required before promotion.
The acceptance gate for each new AMD lane remains: exit 0, no soft failure, nonzero execution, readable log/JUnit/environment/summary/cleanup artifacts, exact image SHA and expected visible GPU count, and
status=PASS leaked_processes=0.AI assistance: Used Codex to diagnose the superseded AMD run, unstack and rebase the branch, implement the narrow ROCm portability and runtime fixes, run local validation, and maintain this description. I reviewed the changes and validation results.