Repository navigation
[CI][ROCm] Align AMD image with vLLM 0.29 - #7395
Conversation
Signed-off-by: andyluo7 <andy.luo@amd.com>
|
Stable Audio follow-up and quarantine disposition are tracked in #7396. The |
|
Author self-review for exact head
Local Docker execution is unavailable on this Mac. The remaining acceptance gate is an exact-head AMD image build followed by a test job confirming vLLM 0.29 and successful diffusion import/startup. |
|
@yenuo26 @hsliuustc0106, no AMD Buildkite context has registered yet for exact head Please start one AMD merge validation for this exact head, preferably through the authorized manual route rather than a broad
|
|
This PR was classified as CI work. Routing: @yenuo26 via semantic router, CI owner, CODEOWNERS; @NickCao via CODEOWNERS @andyluo7, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
|
@vllm-omni-review-bot please review the ROCm image-version alignment and fail-fast compatibility check on the current exact head. |
Omni ReviewBot routing recordAssigned Strict under experiment |
Omni ReviewBot triage noteAutomated triage of commit
These are automated triage suggestions only — the final decision belongs to the maintainers. |
vllm-omni-review-bot
left a comment
There was a problem hiding this comment.
Omni ReviewBot review
Scan:
| Category | Result |
|---|---|
| Tests / verification | 3 finding(s) below |
| Security | no finding reported |
| Docs / comments | 5 finding(s) below |
| Behavior / compatibility | no finding reported |
| Correctness | no finding reported |
Validated:
- [resolved] AMD exact-head validation ask already owned in PR discussion; not re-filed. Residual: static tests ≠ image-build/runtime gate.
- [resolved] Stable Audio/#7396 disposition already tracked; thresholds/grades unchanged here.
- [claim-verified] Docker Hub tag vllm/vllm-openai-rocm:v0.29.0 exists (linux/amd64, active).
- [claim-verified] Scope claim: only Dockerfile.rocm + tests/buildkite/test_rocm_dockerfile.py changed; CUDA CI Dockerfile.ci untouched at v0.29.0.
- [claim-verified] test_rocm_base_tracks_cuda_vllm_release asserts tag equality vs Dockerfile.ci VLLM_BASE_TAG (CI alignment), not Dockerfile.cuda.
- [validated] Producer of required API: vllm_omni/diffusion/diffusion_kv/layout.py:23 imports compute_layout_strides — explains AMD runtime ImportError on 0.28 bases.
Kept four threads: tighten the fail-fast test to the RUN canary (test integrity), resolve Dockerfile paths from __file__ like sibling Buildkite tests, ask whether non-AMD 0.28 pins are intentional deferral after the ROCm-only 0.29 bump (blast-radius), and rename CUDA_DOCKERFILE to CI taxonomy as a nit. Duplicate pairs 0/3 and 2/5 were collapsed onto the surviving indices.
Verdict: COMMENT
Findings
- **[P2] This diff adds
test_rocm_image_fails_fast_on_missing_vllm_apiand insertsRU…** —tests/buildkite/test_rocm_dockerfile.pyThis diff addstest_rocm_image_fails_fast_on_missing_vllm_apiand insertsRUN python3 -c "import vllm; from vllm.v1.kv_cache_interface import compute_layout_strides; ..."atdocker/Dockerfile.rocm:46, but the new test only asserts that the import substring appears anywhere in the file (tests/buildkite/test_rocm_dockerfile.py:35). A comment or echo of the same text would keep the named fail-fast gate green while the executing RUN layer could disappear. Pin the assertion to the concreteRUN python3 -c "import vllm; from vllm.v1.kv_cache_interface import compute_layout_strides;` snippet so the test matches the image-build canary this PR claims to guard.
Evidence: tests/buildkite/test_rocm_dockerfile.py:35 assert "from vllm.v1.kv_cache_interface import compute_layout_strides" in dockerfile — bare substring only; docker/Dockerfile.rocm:46 RUN python3 -c "import vllm; from vllm.v1.kv_cache_interface import compute_layout_strides; print(f'Validated vLLM {vllm.__version__}')" — the executing fail-fast layer the test name claims to guard.
Suggestion: assert (
'RUN python3 -c "import vllm; from vllm.v1.kv_cache_interface import compute_layout_strides;'
in dockerfile
)
- **[P2] ROCm no longer stuck on 0.28 after #7230: docker/Dockerfile.rocm:4
ARG BASE_IM…** —Dockerfile.rocm:4ROCm no longer stuck on 0.28 after #7230: docker/Dockerfile.rocm:4ARG BASE_IMAGE=vllm/vllm-openai-rocm:v0.29.0` (also VLLM_VERSION_OR_COMMIT_HASH=v0.29.0). Residual: unchanged sibling pins still on v0.28.0 — docker/Dockerfile.cuda:1, docker/Dockerfile.xpu:21, docker/Dockerfile.npu:2. AMD runtime CI greenness not verified here.
Evidence: docker/Dockerfile.rocm:4 ARG BASE_IMAGE=vllm/vllm-openai-rocm:v0.29.0 — ROCm pin is v0.29.0. Unchanged by this diff, present in the PR-time tree: docker/Dockerfile.cuda:1 ARG BASE_IMAGE=vllm/vllm-openai:v0.28.0; docker/Dockerfile.xpu:21 ARG VLLM_VERSION=v0.28.0; docker/Dockerfile.npu:2 ARG VLLM_ASCEND_VERSION=v0.28.0.
- [P2] Residual only: API canary at docker/Dockerfile.rocm:46 runs before TorchCodec a… — ``
Residual only: API canary at docker/Dockerfile.rocm:46 runs before TorchCodec and COPY/uv omni install; no post-install re-check of vllm. Acceptable because requirements/*.txt do not declare a vllm dependency that would replace the validated package.
Evidence: docker/Dockerfile.rocm:46 RUN python3 -c "import vllm; from vllm.v1.kv_cache_interface import compute_layout_strides; print(f'Validated vLLM {vllm.__version__}')" — canary; docker/Dockerfile.rocm:48 # Build TorchCodec after any optional nightly vLLM reinstall... then COPY/install TorchCodec; docker/Dockerfile.rocm:70 COPY . ${COMMON_WORKDIR}/vllm-omni and :93–94 uv pip install ... ".[dev]" — omni install after canary with no later Validated vLLM check; unchanged by this diff, present in the PR-time tree: requirements/*.txt have no vllm package pin (grep ^vllm|vllm==|vllm>=|vllm~ → no matches).
- [P2] Fail-fast at docker/Dockerfile.rocm:46 imports compute_layout_strides before om… —
Dockerfile.rocm:46
Fail-fast at docker/Dockerfile.rocm:46 imports compute_layout_strides before omni install; no re-validation afteruv pip install ".[dev]"(lines 92–99). Residual only: vllm is not in omni requirements, so install is unlikely to replace it—minor.
Evidence: docker/Dockerfile.rocm:46 RUN python3 -c "import vllm; from vllm.v1.kv_cache_interface import compute_layout_strides; print(f'Validated vLLM {vllm.__version__}')" — only fail-fast import of compute_layout_strides. Later, unchanged ordering in the same file: docker/Dockerfile.rocm:92–94 RUN cd ${COMMON_WORKDIR}/vllm-omni && \ / uv pip install ... ".[dev]" ... installs omni with no subsequent compute_layout_strides check. Unchanged by this diff, present in the PR-time tree: pyproject.toml:11 dynamic = ["version", "dependencies"] and requirements/common.txt has no vllm package pin (only a comment referencing a vllm_omni path).
- [P2] Residual (out of this PR’s ROCm↔CUDA-CI scope): unchanged docker/Dockerfile.cud… — ``
Residual (out of this PR’s ROCm↔CUDA-CI scope): unchanged docker/Dockerfile.cuda still pinsARG BASE_IMAGE=vllm/vllm-openai:v0.28.0while Dockerfile.ci `VLLM_BASE_TAG` and Dockerfile.rocm defaults are v0.29.0.
Evidence: unchanged by this diff, present in the PR-time tree: docker/Dockerfile.cuda:1 ARG BASE_IMAGE=vllm/vllm-openai:v0.28.0 — contrast docker/Dockerfile.ci:7 ARG VLLM_BASE_TAG=v0.29.0 and docker/Dockerfile.rocm:4 ARG BASE_IMAGE=vllm/vllm-openai-rocm:v0.29.0
- **[P2] This diff lifts only
docker/Dockerfile.rocmtovllm/vllm-openai-rocm:v0.29.0…** —docker/Dockerfile.xpuThis diff lifts onlydocker/Dockerfile.rocmtovllm/vllm-openai-rocm:v0.29.0/VLLM_VERSION_OR_COMMIT_HASH=v0.29.0so AMD CI can importcompute_layout_strides. Same-class pins remain on v0.28.0 indocker/Dockerfile.cuda:1,docker/Dockerfile.xpu:21,docker/Dockerfile.npu/Dockerfile.npu.ci:2, and.buildkite/intel/pipeline-intel.yml:13, while unchangedvllm_omni/diffusion/diffusion_kv/layout.py:23` already imports that 0.29 API. Is the lag intentional unpublished-artifact deferral for non-AMD platforms, or should a follow-up track those pins so non-AMD image builds do not keep the old runtime ImportError?
Evidence: diff-changed: docker/Dockerfile.rocm:4 ARG BASE_IMAGE=vllm/vllm-openai-rocm:v0.29.0; docker/Dockerfile.rocm:16 ARG VLLM_VERSION_OR_COMMIT_HASH=v0.29.0. unchanged by this diff, present in the PR-time tree: docker/Dockerfile.xpu:21 ARG VLLM_VERSION=v0.28.0; docker/Dockerfile.cuda:1 ARG BASE_IMAGE=vllm/vllm-openai:v0.28.0; docker/Dockerfile.npu.ci:2 ARG VLLM_ASCEND_VERSION=v0.28.0; docker/Dockerfile.npu:2 ARG VLLM_ASCEND_VERSION=v0.28.0; .buildkite/intel/pipeline-intel.yml:13 VLLM_VERSION: "v0.28.0"; vllm_omni/diffusion/diffusion_kv/layout.py:23 from vllm.v1.kv_cache_interface import KVCacheSpec, KVCacheTensor, compute_layout_strides
Signed-off-by: andyluo7 <andy.luo@amd.com>
|
Updated self-review for exact head 1e2ce96: Review feedback addressed:
Scope disposition:
Validation:
The focused pytest invocations disabled the repository conftest/addopts because this Mac environment lacks torch and pytest-xdist; no local Docker or AMD runtime result is claimed. @vllm-omni-review-bot please re-review this exact head. @yenuo26 @NickCao please start one AMD merge validation for exact head 1e2ce96. Required evidence remains:
|
vllm-omni-review-bot
left a comment
There was a problem hiding this comment.
Omni ReviewBot review
Scan:
| Category | Result |
|---|---|
| Tests / verification | 7 finding(s) below |
| Security | no finding reported |
| Docs / comments | 1 finding(s) below |
| Behavior / compatibility | no finding reported |
| Correctness | no finding reported |
Validated:
- [resolved] cwd-independent Path(file).resolve().parents[2] — quote tests/buildkite/test_rocm_dockerfile.py:10; residual: still no docker build execution in-test.
- [resolved] CUDA/XPU/NPU still on v0.28 (Dockerfile.cuda:1, Dockerfile.xpu:21, Dockerfile.npu:2) — author deferred; residual: other platforms can still drift vs Dockerfile.ci.
- [resolved] canary startswith(RUN python3 -c ...) — quote lines 35–37; residual: string presence ≠ image build/runtime.
- [resolved] tracks_cuda→tracks_ci naming: CI_DOCKERFILE→Dockerfile.ci — residual: Dockerfile.ci is CUDA CI image, not platform-neutral.
- [resolved] Canary startswith pin: test requires RUN line prefix — residual: proves Dockerfile text, not image-build behavior.
- [resolved] AMD exact-head runtime gate owned in discussion — residual: static tests ≠ image/runtime validation.
Primary mergeable findings center on the new ROCm↔CI alignment contract and fail-fast canary: keep tightening the BASE_IMAGE assert beyond tag-only equality (i=0; dups i=2/i=5), and extend the Dockerfile RUN canary to also import KVCacheLayout as used by layout.py (i=3). Keep the misleading fail-fast test rename as a nit (i=1; dup i=4), and keep gating the source-ref assert on USE_NIGHTLY_BUILD==0 as an optional nit (i=6). No blockers.
Verdict: COMMENT
Findings
- **[P2] test_rocm_base_tracks_ci_vllm_release only checks the tag:
assert rocm_base.rs…** —tests/buildkite/test_rocm_dockerfile.pytest_rocm_base_tracks_ci_vllm_release only checks the tag:assert rocm_base.rsplit(":", 1)[-1] == ci_release(tests/buildkite/test_rocm_dockerfile.py:29). A wrong CUDA-family default likevllm/vllm-openai:v0.29.0would still pass because Dockerfile.ci’sVLLM_BASE_TAGis alsov0.29.0while its image isvllm/vllm-openai(docker/Dockerfile.ci:6-7). Also assert the ROCm image name (allowing a registry prefix), e.g.image_ref.endswith("vllm-openai-rocm")` plus tag equality, matching the default this PR sets at docker/Dockerfile.rocm:4 and that AMD builds without BASE_IMAGE override (.buildkite/amd/test-template-amd-omni.j2:77).
Evidence: tests/buildkite/test_rocm_dockerfile.py:29 assert rocm_base.rsplit(":", 1)[-1] == ci_release — tag only, no image-name check; docker/Dockerfile.rocm:4 ARG BASE_IMAGE=vllm/vllm-openai-rocm:v0.29.0; docker/Dockerfile.ci:6-7 ARG VLLM_BASE_IMAGE=vllm/vllm-openai / ARG VLLM_BASE_TAG=v0.29.0 (same tag, different family); unchanged by this diff, present in the PR-time tree: .buildkite/amd/test-template-amd-omni.j2:77 docker build -f docker/Dockerfile.rocm ... --build-arg ARG_PYTORCH_ROCM_ARCH=gfx942 — no BASE_IMAGE override.
Suggestion: image_ref, _, image_tag = rocm_base.rpartition(":")
assert image_tag == ci_release
assert image_ref.endswith("vllm-openai-rocm"), image_ref
assert rocm_source_ref == ci_release
- **[P2] tests/buildkite/test_rocm_dockerfile.py:29 —
test_rocm_base_tracks_ci_vllm_rel…** — `` tests/buildkite/test_rocm_dockerfile.py:29 —test_rocm_base_tracks_ci_vllm_releaseonly assertsrocm_base.rsplit(":", 1)[-1] == ci_release(androcm_source_ref == ci_release) against Dockerfile.ciVLLM_BASE_TAG; it does not couple to CI image family/VLLM_BASE_IMAGE`. Residual coverage gap only (naming: CI_DOCKERFILE / tracks_ci at L11/L24).
Evidence: tests/buildkite/test_rocm_dockerfile.py:25-30 ci_release = _docker_arg(CI_DOCKERFILE, "VLLM_BASE_TAG") / rocm_base = _docker_arg(ROCM_DOCKERFILE, "BASE_IMAGE") / assert rocm_base.rsplit(":", 1)[-1] == ci_release / assert rocm_source_ref == ci_release — tag-only vs CI; unchanged by this diff, present in PR-time tree: docker/Dockerfile.ci:6-7 ARG VLLM_BASE_IMAGE=vllm/vllm-openai / ARG VLLM_BASE_TAG=v0.29.0 vs docker/Dockerfile.rocm:4 ARG BASE_IMAGE=vllm/vllm-openai-rocm:v0.29.0 (different image family, shared tag).
- [P2] tests/buildkite/test_rocm_dockerfile.py:35–37 only checks that a Dockerfile lin… — ``
tests/buildkite/test_rocm_dockerfile.py:35–37 only checks that a Dockerfile line startswith the canaryRUN python3 -c "import vllm; from vllm.v1.kv_cache_interface import compute_layout_strides;. That is static string presence, not an image build or runtime execution of the fail-fast import (the real RUN lives in docker/Dockerfile.rocm:46).
Evidence: tests/buildkite/test_rocm_dockerfile.py:35 canary = 'RUN python3 -c "import vllm; from vllm.v1.kv_cache_interface import compute_layout_strides;'
tests/buildkite/test_rocm_dockerfile.py:37 assert any(line.startswith(canary) for line in dockerfile.splitlines()) — CPU test reads Dockerfile text only; does not build or run the image. Unchanged residual vs docker/Dockerfile.rocm:46 which does contain the real RUN python3 -c ... layer.
- [P2] Residual drift: only ROCm was bumped to v0.29.0 —
Dockerfile.cuda:1
Residual drift: only ROCm was bumped to v0.29.0; unchanged by this diff, CUDA/XPU/NPU still pin v0.28.0 while docker/Dockerfile.ci:7 has VLLM_BASE_TAG=v0.29.0 (docker/Dockerfile.cuda:1, docker/Dockerfile.xpu:21, docker/Dockerfile.npu:2).
Evidence: unchanged by this diff, present in the PR-time tree: docker/Dockerfile.cuda:1 ARG BASE_IMAGE=vllm/vllm-openai:v0.28.0; docker/Dockerfile.xpu:21 ARG VLLM_VERSION=v0.28.0; docker/Dockerfile.npu:2 ARG VLLM_ASCEND_VERSION=v0.28.0; contrast docker/Dockerfile.ci:7 ARG VLLM_BASE_TAG=v0.29.0 (this PR only updates docker/Dockerfile.rocm to v0.29.0).
Signed-off-by: andyluo7 <andy.luo@amd.com>
|
Follow-up self-review for exact head 5dc9929:
Validation:
@vllm-omni-review-bot please re-review exact head 5dc9929. @yenuo26 @NickCao, this supersedes the prior exact-head request. Please run AMD merge validation against 5dc9929 and retain the same image-version, API-import, diffusion-startup, and #7396 Stable Audio evidence. |
vllm-omni-review-bot
left a comment
There was a problem hiding this comment.
Omni ReviewBot review
Scan:
| Category | Result |
|---|---|
| Tests / verification | 7 finding(s) below |
| Security | no finding reported |
| Docs / comments | 1 finding(s) below |
| Behavior / compatibility | no finding reported |
| Correctness | no finding reported |
Validated:
- [resolved] cwd-independent REPO_ROOT via Path(file).resolve().parents[2] — quote tests/buildkite/test_rocm_dockerfile.py:10; residual: still no docker build execution in-test
- [resolved] CUDA/XPU/NPU still on v0.28 (Dockerfile.cuda:1, Dockerfile.xpu:21, Dockerfile.npu:2) — author deferred; residual: Dockerfile.cuda is also the release/docs build path (.buildkite/release/release-pipeline.yml:66,86)
- [resolved] AMD exact-head validation ask already owned in PR discussion; residual: gate_report clean ≠ proof of AMD
Validated vLLM 0.29.0log for 5dc9929 - [resolved] Stable Audio/#7396 disposition already tracked; thresholds/grades unchanged here
- [resolved] cwd-independent Path(file).resolve().parents[2] (test L10); residual: still no docker build execution in-test.
- [resolved] canary startswith(RUN python3 -c ...) (test L49); residual: order vs nightly block not asserted; USE_NIGHTLY=1 mis-order would still pass the static test.
Keep two concrete follow-ups on the PR’s new ROCm alignment contract: lock USE_NIGHTLY_BUILD=0 in the new regression file (i=0, minor), and fix the Dockerfile.rocm rebase REMINDER so it no longer invites custom ARG pins without updating the new hash lock (i=2, minor; i=1 and i=4 are dups). Also keep a nit rename so the nightly-named source-ref test matches what it actually asserts (i=3). No blockers; primary version bump and API canary are not contested.
Verdict: COMMENT
Findings
- [P2] This diff adds test_rocm_nightly_source_ref_tracks_ci_vllm_release asserting VL… —
docker/Dockerfile.rocm
This diff adds test_rocm_nightly_source_ref_tracks_ci_vllm_release asserting VLLM_VERSION_OR_COMMIT_HASH == Dockerfile.ci VLLM_BASE_TAG on every core_model/cpu run (tests/buildkite/test_rocm_dockerfile.py:34-38). The unchanged rebase REMINDER in docker/Dockerfile.rocm:10-11 still tells maintainers they may change those two ARGs for custom commits, which now fails that CPU gate unless the test is edited in the same change. Update the REMINDER to state the enforced alignment and that intentional commit pins must update the test alongside the ARG.
Evidence: unchanged by this diff, present in the PR-time tree: docker/Dockerfile.rocm:10-11 # 1. If vLLM-omni maintainer is forced to use custom commits / # during rebasing, they can change the two variables; added by this diff: tests/buildkite/test_rocm_dockerfile.py:36-38 rocm_source_ref = _docker_arg(ROCM_DOCKERFILE, "VLLM_VERSION_OR_COMMIT_HASH") then assert rocm_source_ref == ci_release (ci_release from Dockerfile.ci VLLM_BASE_TAG).
Suggestion: # NOTE: REMINDER to vLLM-Omni rebase maintainer
Keep BASE_IMAGE's tag and VLLM_VERSION_OR_COMMIT_HASH equal to
docker/Dockerfile.ci's VLLM_BASE_TAG (enforced by
tests/buildkite/test_rocm_dockerfile.py). An intentional commit pin
must update that test in the same change.
Remember to set USE_NIGHTLY_BUILD to 0 when switching back to the
stable vLLM docker image.
- [P2] CUDA/XPU/NPU remain on v0.28 while CI/ROCm moved to v0.29 (Dockerfile.cuda:1, D… —
Dockerfile.cuda:1
CUDA/XPU/NPU remain on v0.28 while CI/ROCm moved to v0.29 (Dockerfile.cuda:1, Dockerfile.xpu:21, Dockerfile.npu:2). Residual risk: Dockerfile.cuda is the release/docs image path (.buildkite/release/release-pipeline.yml:66,86).
Evidence: unchanged by this diff, present in the PR-time tree: docker/Dockerfile.cuda:1 ARG BASE_IMAGE=vllm/vllm-openai:v0.28.0; docker/Dockerfile.xpu:21 ARG VLLM_VERSION=v0.28.0; docker/Dockerfile.npu:2 ARG VLLM_ASCEND_VERSION=v0.28.0; contrast PR-time docker/Dockerfile.ci:7 ARG VLLM_BASE_TAG=v0.29.0 and docker/Dockerfile.rocm:4 ARG BASE_IMAGE=vllm/vllm-openai-rocm:v0.29.0; release consumer unchanged by this diff: .buildkite/release/release-pipeline.yml:66 -f docker/Dockerfile.cuda . and :86 -f docker/Dockerfile.cuda .
- [P2] Residual only: cpu test
test_rocm_dockerfile_contains_vllm_api_canaryproves… — ``
Residual only: cpu testtest_rocm_dockerfile_contains_vllm_api_canaryproves canary text exists, not that an AMD image build for 5dc9929 printed `Validated vLLM 0.29.0`. gate_report clean ≠ that log.
Evidence: tests/buildkite/test_rocm_dockerfile.py:8 pytestmark = [pytest.mark.core_model, pytest.mark.cpu] — canary test is CPU static; tests/buildkite/test_rocm_dockerfile.py:49 assert any(line.startswith(canary) for line in dockerfile.splitlines()) — asserts Dockerfile text only; docker/Dockerfile.rocm:46 RUN python3 -c "import vllm; from vllm.v1.kv_cache_interface import compute_layout_strides; from vllm.v1.kv_cache_layout import KVCacheLayout; print(f'Validated vLLM {vllm.__version__}')" — log only emitted during ROCm image construction, absent from PR-time tree as an AMD build artifact for 5dc9929.
- **[P2] tests/buildkite/test_rocm_dockerfile.py:49 only checks canary presence via
any…** — `` tests/buildkite/test_rocm_dockerfile.py:49 only checks canary presence viaany(line.startswith(canary)...)`; it does not assert the canary RUN comes after the USE_NIGHTLY_BUILD reinstall block (docker/Dockerfile.rocm:32–46). A canary placed before that block would still pass this static test while leaving USE_NIGHTLY=1 builds unchecked after reinstall.
Evidence: tests/buildkite/test_rocm_dockerfile.py:49 assert any(line.startswith(canary) for line in dockerfile.splitlines()) — presence-only, no order vs nightly. Unchanged ordering fact in PR-time tree: docker/Dockerfile.rocm:32 RUN if [ "${USE_NIGHTLY_BUILD}" = "1" ]; then \ then docker/Dockerfile.rocm:46 RUN python3 -c "import vllm; from vllm.v1.kv_cache_interface import compute_layout_strides; from vllm.v1.kv_cache_layout import KVCacheLayout; print(f'Validated vLLM {vllm.__version__}')" — canary after nightly today, but test does not lock that order.
Signed-off-by: andyluo7 <andy.luo@amd.com>
|
Final follow-up for exact head 177ac10:
The API canary remains a blocking amd-build step intentionally: the shared AMD image must support the repository diffusion path, and failing once during image construction prevents misleading partial greens and repeated downstream runtime failures. The CPU regression still claims only Dockerfile structure; the exact-head AMD build remains execution proof. Validation:
@vllm-omni-review-bot please re-review exact head 177ac10. @yenuo26 @NickCao, this supersedes earlier SHA-specific requests. Please run AMD merge validation against exact head 177ac10 and retain the image-version, API-import, diffusion-startup, and #7396 Stable Audio evidence. |
vllm-omni-review-bot
left a comment
There was a problem hiding this comment.
Omni ReviewBot review
Scan:
| Category | Result |
|---|---|
| Tests / verification | 6 finding(s) below |
| Security | no finding reported |
| Docs / comments | 2 finding(s) below |
| Behavior / compatibility | no finding reported |
| Correctness | no finding reported |
Validated:
- [resolved] CUDA/XPU/NPU still on v0.28 (Dockerfile.cuda:1, Dockerfile.xpu:21, Dockerfile.npu:2) — author deferred; residual: those platforms can still drift vs Dockerfile.ci.
- [resolved] cwd-independent REPO_ROOT Path(file).resolve().parents[2] — quote test_rocm_dockerfile.py:10; residual: no docker build in-test
- [resolved] canary startswith RUN python3 -c … — quote test:55-60; residual: string presence ≠ import execution
- [resolved] USE_NIGHTLY_BUILD=0 lock — quote test:40-41 + Dockerfile:14; residual: CI --build-arg override not locked
- [resolved] USE_NIGHTLY_BUILD=0 locked by test_rocm_defaults_to_prebuilt_base_image — residual: nightly=1 path still unexecuted in CI defaults.
- [resolved] cwd-independent REPO_ROOT via Path(file).resolve().parents[2] — residual: tests assert Dockerfile text only, not a real docker build.
Kept two minors on the primary v0.29 ROCm rebase: docs/users inherit the new BASE_IMAGE default without a CUDA-style pin (blast-radius), and TorchCodec’s pinned branch still needs amd-build confirmation after the base bump (adversary+behavior). Dropped speculative amd-template test locks and a rename-only nit; collapsed the duplicate TorchCodec asks into one.
Verdict: COMMENT
Findings
- [P2] Residual: CUDA/XPU/NPU Dockerfiles still pin v0.28.0 while Dockerfile.ci is on… —
Dockerfile.cuda:1
Residual: CUDA/XPU/NPU Dockerfiles still pin v0.28.0 while Dockerfile.ci is on v0.29.0 (Dockerfile.cuda:1vllm/vllm-openai:v0.28.0, Dockerfile.xpu:21VLLM_VERSION=v0.28.0, Dockerfile.npu:2VLLM_ASCEND_VERSION=v0.28.0). This PR only aligns ROCm + tests/buildkite/test_rocm_dockerfile.py; those platforms can still drift vs CI.
Evidence: unchanged by this diff, present in the PR-time tree: docker/Dockerfile.cuda:1 ARG BASE_IMAGE=vllm/vllm-openai:v0.28.0; docker/Dockerfile.xpu:21 ARG VLLM_VERSION=v0.28.0; docker/Dockerfile.npu:2 ARG VLLM_ASCEND_VERSION=v0.28.0; docker/Dockerfile.ci:7 ARG VLLM_BASE_TAG=v0.29.0 (CI pin this PR aligns ROCm to).
- **[P2] test_rocm_dockerfile_contains_vllm_api_canary only checks Dockerfile text:
_li…** — `` test_rocm_dockerfile_contains_vllm_api_canary only checks Dockerfile text:_line_indexmatches a line startswith theRUN python3 -c "import vllm; …prefix and that it follows the nightlyfi`. That is string presence/order, not import execution (execution happens only if the image is built).
Evidence: tests/buildkite/test_rocm_dockerfile.py:24-27 _line_index uses line.startswith(prefix); tests/buildkite/test_rocm_dockerfile.py:55-60 canary = ('RUN python3 -c "import vllm; ' …); canary_index = _line_index(lines, canary) — no import/runtime of that snippet. Unchanged-by-claim but present in PR-time tree: docker/Dockerfile.rocm:45 RUN python3 -c "import vllm; from vllm.v1.kv_cache_interface import compute_layout_strides; from vllm.v1.kv_cache_layout import KVCacheLayout; print(f'Validated vLLM {vllm.__version__}')" — build-time execution is outside this CPU test.
- **[P2] Add a consumer lock in
tests/buildkite/test_rocm_dockerfile.py(mirroringte…** —.buildkite/amd/test-template-amd-omni.j2:77Add a consumer lock intests/buildkite/test_rocm_dockerfile.py(mirroringtests/buildkite/test_amd_pipeline.py's live pipeline-command assertions) so amd-build at.buildkite/amd/test-template-amd-omni.j2:77stays pinned to-f docker/Dockerfile.rocmwith no--build-arg BASE_IMAGE=` override.
Evidence: unchanged by this diff, present in the PR-time tree: .buildkite/amd/test-template-amd-omni.j2:77 - "docker build -f docker/Dockerfile.rocm -t {{ docker_image_amd }} --target test --build-arg ARG_PYTORCH_ROCM_ARCH=gfx942 --progress plain ." (key amd-build at :79). PR-added tests/buildkite/test_rocm_dockerfile.py only asserts Dockerfile ARG tags/canary (e.g. :30-37 test_rocm_base_tracks_ci_vllm_release / _docker_arg(..., "BASE_IMAGE")) and never reads the AMD template or amd-build command. Sibling contrast, unchanged by this diff: tests/buildkite/test_amd_pipeline.py:31-42 locks live Buildkite step commands (test_qwen3_tts_base_preserves_advanced_model_arguments).
- [P2] Update
docs/getting_started/installation/gpu/rocm.inc.md:135(and:154) sti… —docs/getting_started/installation/gpu/rocm.inc.md:135
Updatedocs/getting_started/installation/gpu/rocm.inc.md:135(and:154) still showing prebuiltvllm/vllm-omni-rocm:v0.28.0, or track that published-omni tag as an explicit follow-up outside this PR's stated CI-only scope.
Evidence: unchanged by this diff, present in the PR-time tree: docs/getting_started/installation/gpu/rocm.inc.md:135 vllm/vllm-omni-rocm:v0.28.0 \ — also docs/getting_started/installation/gpu/rocm.inc.md:154 vllm/vllm-omni-rocm:v0.28.0 (published omni image tag, not the CI vllm/vllm-openai-rocm base bumped in docker/Dockerfile.rocm).
- **[P2] This diff sets
docker/Dockerfile.rocmARG BASE_IMAGE=vllm/vllm-openai-rocm:v…** —docs/getting_started/installation/gpu/rocm.inc.mdThis diff setsdocker/Dockerfile.rocmARG BASE_IMAGE=vllm/vllm-openai-rocm:v0.29.0, so the bare user build indocs/getting_started/installation/gpu/rocm.inc.md(DOCKER_BUILDKIT=1 docker build -f docker/Dockerfile.rocm ...) inherits that default with no version callout, whiledocs/getting_started/installation/gpu/cuda.inc.mdalready shows an optional--build-arg BASE_IMAGE=...:v0.29.0`. Widen the PR scope claim beyond “ROCm CI image only” and mirror that explicit pin (or a one-line note that the default base is now v0.29.0).
Evidence: Changed by this diff: docker/Dockerfile.rocm:4 ARG BASE_IMAGE=vllm/vllm-openai-rocm:v0.29.0. Unchanged by this diff, present in the PR-time tree: docs/getting_started/installation/gpu/rocm.inc.md:75 DOCKER_BUILDKIT=1 docker build -f docker/Dockerfile.rocm -t vllm-omni-rocm .; docs/getting_started/installation/gpu/cuda.inc.md:131-136 If you want to specify the base vLLM version: / DOCKER_BUILDKIT=1 docker build \ / -f docker/Dockerfile.cuda \ / --build-arg BASE_IMAGE=vllm/vllm-openai:v0.29.0 \
Suggestion: DOCKER_BUILDKIT=1 docker build -f docker/Dockerfile.rocm -t vllm-omni-rocm .
Optional: pin the base explicitly (should match docker/Dockerfile.ci VLLM_BASE_TAG)
DOCKER_BUILDKIT=1 docker build
-f docker/Dockerfile.rocm
--build-arg BASE_IMAGE=vllm/vllm-openai-rocm:v0.29.0
-t vllm-omni-rocm .
| # Keep the ROCm runtime aligned with the vLLM release targeted by | ||
| # docker/Dockerfile.ci. A stale base image can build successfully while failing | ||
| # at runtime when vLLM-Omni imports APIs added by the newer vLLM release. | ||
| ARG BASE_IMAGE=vllm/vllm-openai-rocm:v0.29.0 |
There was a problem hiding this comment.
[P2] This diff rebases docker/Dockerfile.rocm ARG BASE_IMAGE/`VLLM_VERSION_OR_CO…
This diff rebases docker/Dockerfile.rocm ARG BASE_IMAGE/VLLM_VERSION_OR_COMMIT_HASH to v0.29.0 and adds an import-only vLLM API canary before the unchanged TorchCodec layer that still runs tools/install_torchcodec_rocm.sh (TORCHCODEC_BRANCH=v0.10.0), whose comment requires revalidation on ROCm base/vLLM rebases. PR validation covers static ARG tests and the canary, not TorchCodec/ABI — please confirm the AMD image-build log for this head (e.g. post-177ac10) shows the TorchCodec/VideoDecoder step succeeding against the v0.29.0 base, or update that pin if it does not.
Evidence: docker/Dockerfile.rocm:4 ARG BASE_IMAGE=vllm/vllm-openai-rocm:v0.29.0; docker/Dockerfile.rocm:15 ARG VLLM_VERSION_OR_COMMIT_HASH=v0.29.0; docker/Dockerfile.rocm:45 RUN python3 -c "import vllm; from vllm.v1.kv_cache_interface import compute_layout_strides; from vllm.v1.kv_cache_layout import KVCacheLayout; print(f'Validated vLLM {vllm.__version__}')"; docker/Dockerfile.rocm:51-59 still COPY tools/install_torchcodec_rocm.sh then bash /tmp/install_torchcodec_rocm.sh after the canary; unchanged by this diff, present in the PR-time tree: tools/install_torchcodec_rocm.sh:11-13 # Keep this pinned for ROCm/PyTorch ABI compatibility. When rebasing the ROCm / base image or vLLM, revalidate this version against the active torch build. / TORCHCODEC_BRANCH="${TORCHCODEC_BRANCH:-v0.10.0}"; tests/buildkite/test_rocm_dockerfile.py has no TorchCodec assertions (grep torchcodec: no matches).
There was a problem hiding this comment.
This remains an exact-head AMD evidence gate rather than a speculative pin change. The requested build for 2e850d4 must show tools/install_torchcodec_rocm.sh completing and the final VideoDecoder import succeeding against the v0.29.0 base. I am leaving this thread open until that artifact exists; if the step fails, the TorchCodec pin will be adjusted from the captured failure.
There was a problem hiding this comment.
Updated for exact head f8af7e2d146cfa78629c36bd9ffd463f4f518b39. This remains the canonical open runtime-evidence thread: the AMD image log must show tools/install_torchcodec_rocm.sh completing and the final from torchcodec.decoders import VideoDecoder import succeeding against the vLLM 0.29 base. If it fails, the pin will be changed from the captured failure rather than speculatively.
There was a problem hiding this comment.
Runtime evidence is now available from AMD Buildkite #11682 on commit f8af7e2: the image used vllm/vllm-openai-rocm:v0.29.0, printed Validated vLLM 0.29.0, completed the TorchCodec v0.10.0 build and installation, and completed Docker step 16 containing the final VideoDecoder import. A Z-Image diffusion worker also initialized and began loading model weights. Commit 15d3c10 changes only documentation and Dockerfile comments; I am leaving this canonical thread open until the final exact-head AMD rerun confirms the new head.
Signed-off-by: andyluo7 <andy.luo@amd.com>
|
Follow-up for exact head 2e850d4:
Validation:
The TorchCodec review thread remains intentionally open pending exact-head hardware evidence. The AMD image log must show install_torchcodec_rocm.sh and the final VideoDecoder import succeeding before that thread is closed. @vllm-omni-review-bot please re-review exact head 2e850d4. @yenuo26 @NickCao, this supersedes earlier SHA-specific requests. Please run AMD merge validation against exact head 2e850d4 and retain the vLLM version/API, TorchCodec/VideoDecoder, diffusion-startup, and #7396 Stable Audio evidence. |
vllm-omni-review-bot
left a comment
There was a problem hiding this comment.
Omni ReviewBot review
Scan:
| Category | Result |
|---|---|
| Tests / verification | 5 finding(s) below |
| Security | no finding reported |
| Docs / comments | 3 finding(s) below |
| Behavior / compatibility | no finding reported |
| Correctness | no finding reported |
Validated:
- [resolved] ReviewBot cwd / CI naming / canary startswith / USE_NIGHTLY=0 / template default locks addressed at head 2e850d4; residual: no in-test docker build execution.
- [resolved] CUDA/XPU/NPU 0.28 drift vs Dockerfile.ci deferred by author; residual siblings still unaligned
- [claim-verified] Early canary imports compute_layout_strides and KVCacheLayout after the nightly block (Dockerfile.rocm:31-45), matching vllm_omni/diffusion/diffusion_kv/layout.py:23-24 consumers.
- [claim-verified] CUDA/XPU/NPU pins unchanged at head: Dockerfile.cuda:1 v0.28.0, Dockerfile.xpu VLLM_VERSION=v0.28.0, Dockerfile.npu VLLM_ASCEND_VERSION=v0.28.0 (author-deferred).
- [validated] requirements/setup do not install/replace vllm; canary-before-omni-install cannot be invalidated by uv pip install ".[dev]" (requirements/common.txt, rocm.txt; setup.py get_install_requires)
- [validated] pytestmark core_model+cpu present (test_rocm_dockerfile.py:8); collected by AMD/CUDA catch-all
tests/ -m 'core_model and cpu'(.buildkite/amd/test-amd-merge.yml:52, .buildkite/cuda/test-ready.yml:34).
Keep three items centered on the ROCm 0.29 base alignment: TorchCodec pin/log confirmation after the ABI rebase (minor), tighter --build-arg matching in the new AMD template lock (nit), and ROCm install-doc note/override for the new default (minor). Drop the uncertain amd-build canary log ask as already covered by PR acceptance criteria, and drop the test-file relocation nit as intentional co-location with the Dockerfile-enforced module.
Verdict: COMMENT
Findings
- [P2] cwd-independent REPO_ROOT via Path(file).resolve().parents[2] (tests/buildk… — ``
cwd-independent REPO_ROOT via Path(file).resolve().parents[2] (tests/buildkite/test_rocm_dockerfile.py:10); residual: siblings like test_amd_pipeline.py still use cwd-relative Path — out of this PR scope.
Evidence: tests/buildkite/test_rocm_dockerfile.py:10 REPO_ROOT = Path(__file__).resolve().parents[2] — cwd-independent. Unchanged by this diff, present in the PR-time tree: tests/buildkite/test_amd_pipeline.py:12 AMD_MERGE_PIPELINE = Path(".buildkite/amd/test-amd-merge.yml") — cwd-relative sibling.
- [P2] Prior ReviewBot locks are present at this head (REPO_ROOT via parents[2], VLLM_… — ``
Prior ReviewBot locks are present at this head (REPO_ROOT via parents[2], VLLM_BASE_TAG tag alignment, canary startswith, USE_NIGHTLY_BUILD=0, AMD template omitting BASE_IMAGE/USE_NIGHTLY overrides). Residual: tests/buildkite/test_rocm_dockerfile.py only read_text-parses Dockerfile/template strings and never executes docker build.
Evidence: tests/buildkite/test_rocm_dockerfile.py:46-53 template = AMD_TEMPLATE.read_text(encoding="utf-8") / build_commands = [line.strip() for line in template.splitlines() if '"docker build ' in line] / assert "BASE_IMAGE" not in build_command / assert "USE_NIGHTLY_BUILD" not in build_command — inspects the build command string only; grep of this file finds no subprocess/docker/os.system/run invocation. Unchanged-by-claim but supporting locks in the same file: :10 REPO_ROOT = Path(__file__).resolve().parents[2]; :42 assert _docker_arg(ROCM_DOCKERFILE, "USE_NIGHTLY_BUILD") == "0"; :72 _line_index(lines, canary) (startswith). Unchanged by this residual, present in PR-time tree: docker/Dockerfile.ci:7 ARG VLLM_BASE_TAG=v0.29.0; .buildkite/amd/test-template-amd-omni.j2:77 docker build uses Dockerfile.rocm defaults.
- [P2] REPO_ROOT is cwd-independent via Path(file).resolve().parents[2] (tests/bui… —
test_rocm_dockerfile.py:10
REPO_ROOT is cwd-independent via Path(file).resolve().parents[2] (tests/buildkite/test_rocm_dockerfile.py:10). Residual: these tests only statically parse Dockerfile/AMD-template text and never execute docker build.
Evidence: tests/buildkite/test_rocm_dockerfile.py:10 REPO_ROOT = Path(__file__).resolve().parents[2]; tests/buildkite/test_rocm_dockerfile.py:47 build_commands = [line.strip() for line in template.splitlines() if '"docker build ' in line] — template string filter only; no subprocess/docker build invocation in this file
- **[P2] Canary is on the shared ROCm image (AMD suite +
docker build -f docker/Dockerf…** —Dockerfile.rocm:45Canary is on the shared ROCm image (AMD suite +docker build -f docker/Dockerfile.rocm` source builds) and fails loud by design — see docker/Dockerfile.rocm:43-45. Residual: that image is not AR-split, so AR rebuilds still run the diffusion KV import canary.
Evidence: docker/Dockerfile.rocm:43-45 # Fail during image construction instead of letting every AMD test fail at / # runtime when the base image and vLLM-Omni target incompatible vLLM APIs. / RUN python3 -c "import vllm; from vllm.v1.kv_cache_interface import compute_layout_strides; from vllm.v1.kv_cache_layout import KVCacheLayout; print(f'Validated vLLM {vllm.__version__}')" — fail-loud canary on the Dockerfile.rocm base stage. Unchanged by this diff beyond the new canary itself, present in the PR-time tree: .buildkite/amd/test-template-amd-omni.j2:77 "docker build -f docker/Dockerfile.rocm -t {{ docker_image_amd }} --target test --build-arg ARG_PYTORCH_ROCM_ARCH=gfx942 --progress plain ." — single shared AMD image build. Unchanged by this diff, present in the PR-time tree: vllm_omni/diffusion/diffusion_kv/layout.py:23-24 from vllm.v1.kv_cache_interface import KVCacheSpec, KVCacheTensor, compute_layout_strides / from vllm.v1.kv_cache_layout import KVCacheLayout — canary symbols are diffusion-KV consumers, so AR-shared rebuilds still pay that import tax.
- [P2] Land or track sibling pin bumps on
Dockerfile.cuda:1,Dockerfile.xpu:21, an… —docker/Dockerfile.cuda:1
Land or track sibling pin bumps onDockerfile.cuda:1,Dockerfile.xpu:21, andDockerfile.npu:2(still v0.28.0) now thatDockerfile.ciandDockerfile.rocmare on v0.29.0 — residual cross-platform drift should not stay silent.
Evidence: unchanged by this diff, present in the PR-time tree: docker/Dockerfile.cuda:1 ARG BASE_IMAGE=vllm/vllm-openai:v0.28.0; docker/Dockerfile.xpu:21 ARG VLLM_VERSION=v0.28.0; docker/Dockerfile.npu:2 ARG VLLM_ASCEND_VERSION=v0.28.0. Contrast with pinned tree: docker/Dockerfile.ci:7 ARG VLLM_BASE_TAG=v0.29.0; docker/Dockerfile.rocm:4 ARG BASE_IMAGE=vllm/vllm-openai-rocm:v0.29.0.
- **[P2] This PR flips
docker/Dockerfile.rocmdefaultBASE_IMAGEtovllm/vllm-opena…** —docs/getting_started/installation/gpu/rocm.inc.mdThis PR flipsdocker/Dockerfile.rocmdefaultBASE_IMAGEtovllm/vllm-openai-rocm:v0.29.0and adds an import canary that fails older bases, so the documented bare build indocs/getting_started/installation/gpu/rocm.inc.md(line 75) now silently produces a 0.29 stack while the adjacent pre-built section still showsvllm/vllm-omni-rocm:v0.28.0. CUDA’s sibling already documents--build-arg BASE_IMAGE=…:v0.29.0; add a one-line note that the source build defaults to v0.29.0 (Hub retag remains #7405) plus a CUDA-style optional--build-arg BASE_IMAGE=` example so rebase/debug overrides of older ROCm bases fail with the explained canary rather than a surprise.
Evidence: unchanged by this diff, present in the PR-time tree: rocm.inc.md:75 DOCKER_BUILDKIT=1 docker build -f docker/Dockerfile.rocm -t vllm-omni-rocm . — bare build, no BASE_IMAGE note; Dockerfile.rocm:4 ARG BASE_IMAGE=vllm/vllm-openai-rocm:v0.29.0; Dockerfile.rocm:45 RUN python3 -c "import vllm; from vllm.v1.kv_cache_interface import compute_layout_strides; from vllm.v1.kv_cache_layout import KVCacheLayout; print(f'Validated vLLM {vllm.__version__}')"; unchanged by this diff: rocm.inc.md:135 vllm/vllm-omni-rocm:v0.28.0 \; unchanged by this diff: cuda.inc.md:136 --build-arg BASE_IMAGE=vllm/vllm-openai:v0.29.0 \
Suggestion: #### Build docker image
The default base is vllm/vllm-openai-rocm:v0.29.0 (aligned with docker/Dockerfile.ci). Older bases fail the image-build API canary. Published vllm/vllm-omni-rocm Hub tags remain on 0.28 until #7405.
DOCKER_BUILDKIT=1 docker build -f docker/Dockerfile.rocm -t vllm-omni-rocm .If you want to specify the base vLLM version:
DOCKER_BUILDKIT=1 docker build \
-f docker/Dockerfile.rocm \
--build-arg BASE_IMAGE=vllm/vllm-openai-rocm:v0.29.0 \
-t vllm-omni-rocm .- [P3] test_amd_build_uses_rocm_dockerfile_defaults (lines 52–53) locks AMD builds to… —
tests/buildkite/test_rocm_dockerfile.py
test_amd_build_uses_rocm_dockerfile_defaults (lines 52–53) locks AMD builds to Dockerfile.rocm defaults by asserting raw substringsBASE_IMAGEandUSE_NIGHTLY_BUILDare absent from the build line. Preferassert "--build-arg BASE_IMAGE" not in build_commandandassert "--build-arg USE_NIGHTLY_BUILD" not in build_commandso a future image/tag/path that merely contains those tokens cannot false-fail the lock.
Evidence: tests/buildkite/test_rocm_dockerfile.py:52-53 assert "BASE_IMAGE" not in build_command / assert "USE_NIGHTLY_BUILD" not in build_command — bare token lock. Unchanged by this diff, present in the PR-time tree: .buildkite/amd/test-template-amd-omni.j2:77 - "docker build -f docker/Dockerfile.rocm -t {{ docker_image_amd }} --target test --build-arg ARG_PYTORCH_ROCM_ARCH=gfx942 --progress plain ." — no --build-arg BASE_IMAGE / --build-arg USE_NIGHTLY_BUILD; intended forbidden overrides are those --build-arg forms.
Suggestion: assert "--build-arg BASE_IMAGE" not in build_command
assert "--build-arg USE_NIGHTLY_BUILD" not in build_command
hsliuustc0106
left a comment
There was a problem hiding this comment.
Findings from a frozen snapshot at 2e850d4 (base main @ 5a386d78). Read-only review — no code from the head was executed.
[P2] TorchCodec ABI pin is not revalidated on the base-image rebase — docker/Dockerfile.rocm:47-65, tools/install_torchcodec_rocm.sh:11-13
This diff moves the ROCm base from v0.28.0 to v0.29.0, which changes the torch/ROCm ABI under the TorchCodec layer, but that layer and its pin are unchanged. The script states its own contract:
# Keep this pinned for ROCm/PyTorch ABI compatibility. When rebasing the ROCm base image or vLLM, revalidate this version against the active torch build.—TORCHCODEC_BRANCH="${TORCHCODEC_BRANCH:-v0.10.0}"
This PR is exactly the "rebasing the ROCm base image" trigger, and no revalidation evidence is attached. Smallest fix: attach the amd-build result for 2e850d4 showing tools/install_torchcodec_rocm.sh completing and the final from torchcodec.decoders import VideoDecoder import succeeding against the v0.29.0 base — or bump TORCHCODEC_BRANCH from the captured failure.
[P3] The alignment contract is enforced for ROCm only, and binds a user-facing image to a CI-only image's tag — tests/buildkite/test_rocm_dockerfile.py:24-38, docker/Dockerfile.cuda:1
test_rocm_base_tracks_ci_vllm_release pins Dockerfile.rocm's BASE_IMAGE to docker/Dockerfile.ci's VLLM_BASE_TAG, and the new comment codifies that as the invariant. Two consequences need an explicit decision:
Dockerfile.rocmis user-facing (per the PR body: users building it without an override), yet it is now bound to a CI-only image's tag. IfDockerfile.cimoves to an unreleased tag for CI reasons, this test forces the user-facing ROCm default to follow.- The sibling user-facing image
docker/Dockerfile.cuda:1still defaults tovllm/vllm-openai:v0.28.0whileDockerfile.ciisv0.29.0. Sincevllm_omni/diffusion/diffusion_kv/layout.py:23-24importscompute_layout_strides/KVCacheLayout(v0.29 APIs), the same import failure fixed here for AMD remains reachable for anyone buildingDockerfile.cuda.
This is pre-existing rather than introduced here, so treating it as a scope question: bump Dockerfile.cuda in this change, or state in the PR that CUDA is tracked separately and open a follow-up. If the intent is "the Dockerfile.ci tag is the project-wide vLLM release of record," say so and the coupling is fine.
Verified sound (checked, no issue)
- The canary targets the real APIs.
docker/Dockerfile.rocm:45imports exactly what production imports —vllm_omni/diffusion/diffusion_kv/layout.py:23(compute_layout_stridesfromvllm.v1.kv_cache_interface) and:24(KVCacheLayoutfromvllm.v1.kv_cache_layout). Not a guessed path. - The new tests do run in CI.
tests/buildkite/is not in the ignore list of the CPU sweep —.buildkite/cuda/test-ready.yml:34and the AMDSimple · Other Teststep both runpytest -sv tests/ -m 'core_model and cpu' --ignore={diffusion,model_executor,entrypoints,engine}, sotest_rocm_dockerfile.pyis collected on every PR. - Markers are legal.
core_modelandcpuare registered inpyproject.toml:250,266, satisfying--strict-markers(pyproject.toml:244). - Every test assumption holds at the frozen head.
Dockerfile.cideclaresARG VLLM_BASE_TAG=exactly once (:7,v0.29.0); ROCm hasBASE_IMAGE=…:v0.29.0,USE_NIGHTLY_BUILD=0,VLLM_VERSION_OR_COMMIT_HASH=v0.29.0; the AMD template has exactly onedocker buildline, it uses-f docker/Dockerfile.rocm, and passes noBASE_IMAGE/USE_NIGHTLY_BUILDoverride (.buildkite/amd/test-template-amd-omni.j2:77). - The AMD build is a real pre-merge gate. The template defines
key: "amd-build"withsoft_fail: false(:73-93), and every AMD test step carriesdepends_on: amd-build. REPO_ROOTis cwd-independent (parents[2]), and.buildkite/is not excluded byDockerfile.ci.dockerignore(shipped viaCOPY --link . .), so the file the tests read exists in CI.
Validation gaps
- The PR's own stated gate has not reported. No
amd-buildcheck-run or status exists on2e850d4. I could not independently confirm thatvllm/vllm-openai-rocm:v0.29.0is published (Docker Hub was unreachable from my environment), so the canary's actual import success is unproven until the AMD build runs. A skipped hardware gate is a gap, not a pass. - The local command quoted in the PR body uses
-o addopts='', which disables--strict-markers. Harmless here since the markers are registered, but it means that run did not exercise the repo's realaddopts.
Signed-off-by: andyluo7 <andy.luo@amd.com>
|
Follow-up for exact head
Validation:
@vllm-omni-review-bot please re-review exact head @yenuo26 @hsliuustc0106 @NickCao, please run one AMD merge validation against this exact head. Required evidence is:
|
vllm-omni-review-bot
left a comment
There was a problem hiding this comment.
Omni ReviewBot review
Scan:
| Category | Result |
|---|---|
| Tests / verification | 6 finding(s) below |
| Security | no finding reported |
| Docs / comments | 4 finding(s) below |
| Behavior / compatibility | no finding reported |
| Correctness | no finding reported |
Validated:
- [resolved] CUDA/XPU/NPU still on v0.28 (Dockerfile.cuda:1, Dockerfile.xpu:21, Dockerfile.npu:2) — deferred to #7410; residual: those platforms can still drift vs Dockerfile.ci.
- [resolved] Prior ReviewBot cwd/CI naming/canary startswith/USE_NIGHTLY=0/template override locks present at head f8af7e2.
- [resolved] USE_NIGHTLY_BUILD=0 lock + template no-override locks — test:41-54; residual: rendered minijinja output not re-checked beyond j2 source
- [resolved] cwd-independent REPO_ROOT / CI naming / canary startswith / USE_NIGHTLY=0 / template locks — present at head; residual: tests are string/content locks, not docker execution.
- [claim-verified] Canary RUN imports compute_layout_strides and KVCacheLayout after the nightly if/fi (docker/Dockerfile.rocm:31-45), matching vllm_omni/diffusion/diffusion_kv/layout.py:23-24.
- [claim-verified] requirements/*.txt and setup.py do not declare vllm, so later uv pip install ".[dev]" does not replace the base-image vLLM the canary checked.
Keep all six with centrality on the ROCm 0.29 BASE_IMAGE flip: fleet-wide inheritance plus exact-head AMD runtime proof (4), canary order locked before TorchCodec not only after nightly fi (0), and docs for coexisting 0.28/0.29 plus #7405 on prebuilt examples (1, 5). Demote the canary version-prefix assert to nit as optional beyond the import canary and static tag lock (2); keep the misleading nightly-default test rename as nit (3). No drops or dups.
Verdict: COMMENT
Findings
- [P2] CUDA/XPU/NPU images remain on v0.28 while CI/ROCm are on v0.29 (docker/Dockerfi… —
Dockerfile.cuda:1
CUDA/XPU/NPU images remain on v0.28 while CI/ROCm are on v0.29 (docker/Dockerfile.cuda:1, docker/Dockerfile.xpu:21, docker/Dockerfile.npu:2). Unchanged by this diff; deferred to #7410. Residual: those platforms can still drift vs docker/Dockerfile.ci because tests/buildkite/test_rocm_dockerfile.py only gates ROCm.
Evidence: unchanged by this diff, present in the PR-time tree: docker/Dockerfile.cuda:1 ARG BASE_IMAGE=vllm/vllm-openai:v0.28.0; docker/Dockerfile.xpu:21 ARG VLLM_VERSION=v0.28.0; docker/Dockerfile.npu:2 ARG VLLM_ASCEND_VERSION=v0.28.0; contrast docker/Dockerfile.ci:7 ARG VLLM_BASE_TAG=v0.29.0 and (in this diff) docker/Dockerfile.rocm:4 ARG BASE_IMAGE=vllm/vllm-openai-rocm:v0.29.0
- [P2] cwd-independent REPO_ROOT via Path(file).resolve().parents[2] at tests/buil… —
test_rocm_dockerfile.py:10
cwd-independent REPO_ROOT via Path(file).resolve().parents[2] at tests/buildkite/test_rocm_dockerfile.py:10; residual: tests only statically parse Dockerfile/AMD template text and never execute docker build.
Evidence: tests/buildkite/test_rocm_dockerfile.py:10 REPO_ROOT = Path(__file__).resolve().parents[2] — cwd-independent root. tests/buildkite/test_rocm_dockerfile.py:45-54 test_amd_build_uses_rocm_dockerfile_defaults only does AMD_TEMPLATE.read_text(...) and filters lines with "docker build " then asserts substrings; no subprocess/docker invocation anywhere in the file.
- [P2] Published
vllm/vllm-omni-rocm:v0.28.0examples are retained (rocm.inc.md:160,… —rocm.inc.md:84
Publishedvllm/vllm-omni-rocm:v0.28.0examples are retained (rocm.inc.md:160,181) with an explicit #7405 published-tag-availability callout (rocm.inc.md:84-86); residual: no v0.29 omni-rocm published tag in these examples.
Evidence: docs/getting_started/installation/gpu/rocm.inc.md:84-86 image-build canary. This upstream base is distinct from the prebuilt / ``vllm/vllm-omni-rocm images discussed below; their published-tag availability / `is tracked in #7405.`; docs/getting_started/installation/gpu/rocm.inc.md:160 ` vllm/vllm-omni-rocm:v0.28.0 `; also line 181 ` vllm/vllm-omni-rocm:v0.28.0`
- [P2] Locks hold in tests/buildkite/test_rocm_dockerfile.py:41-54 (USE_NIGHTLY_BUILD=… — ``
Locks hold in tests/buildkite/test_rocm_dockerfile.py:41-54 (USE_NIGHTLY_BUILD=0 + AMD j2 build command must not pass --build-arg BASE_IMAGE/USE_NIGHTLY_BUILD). Residual: that test only reads .buildkite/amd/test-template-amd-omni.j2 source; live AMD CI renders via minijinja-cli in bootstrap-amd-omni.sh and that rendered output is not re-checked.
Evidence: tests/buildkite/test_rocm_dockerfile.py:41-42 def test_rocm_defaults_to_prebuilt_base_image() -> None: / assert _docker_arg(ROCM_DOCKERFILE, "USE_NIGHTLY_BUILD") == "0"; tests/buildkite/test_rocm_dockerfile.py:46-54 template = AMD_TEMPLATE.read_text(encoding="utf-8") … for arg_name in ("BASE_IMAGE", "USE_NIGHTLY_BUILD"): / assert f"--build-arg {arg_name}" not in build_command; unchanged by this diff, present in PR-time tree: .buildkite/amd/test-template-amd-omni.j2:77 "docker build -f docker/Dockerfile.rocm -t {{ docker_image_amd }} --target test --build-arg ARG_PYTORCH_ROCM_ARCH=gfx942 --progress plain ."; unchanged by this diff, present in PR-time tree: .buildkite/amd/scripts/bootstrap-amd-omni.sh:186 minijinja-cli test-template.j2 "$TEST_YAML" \ — render path exists but no test asserts rendered output for these locks
- [P2] This PR adds the fail-fast canary RUN immediately before the TorchCodec build l… —
tests/buildkite/test_rocm_dockerfile.py
This PR adds the fail-fast canary RUN immediately before the TorchCodec build layer (COPY tools/install_torchcodec_rocm.sh), buttest_rocm_dockerfile_contains_vllm_api_canaryonly requires an intervening nightlyfibeforecanary_index(tests/buildkite/test_rocm_dockerfile.py:75-77). Extend the assertion socanary_indexis also strictly before that TorchCodec COPY line; otherwise the canary can drift below the expensive layer (or pastCOPY .) and the regression still passes while the early-fail claim is lost.
Evidence: tests/buildkite/test_rocm_dockerfile.py:75-77 assert any(line.strip() == "fi" for line in lines[nightly_start + 1 : canary_index]), ( — only lower-bound after nightly; no upper bound vs TorchCodec. docker/Dockerfile.rocm:45 RUN python3 -c "import vllm; from vllm.v1.kv_cache_interface import compute_layout_strides; from vllm.v1.kv_cache_layout import KVCacheLayout; print(f'Validated vLLM {vllm.__version__}')" then docker/Dockerfile.rocm:51 COPY tools/install_torchcodec_rocm.sh /tmp/install_torchcodec_rocm.sh.
Suggestion: assert any(line.strip() == "fi" for line in lines[nightly_start + 1 : canary_index]), (
"vLLM API canary must follow the optional nightly source reinstall block"
)
torchcodec_copy = _line_index(lines, "COPY tools/install_torchcodec_rocm.sh")
assert canary_index < torchcodec_copy, (
"vLLM API canary must run before the TorchCodec build layer"
)
Signed-off-by: andyluo7 <andy.luo@amd.com>
|
Follow-up for exact head
Validation:
AMD #11682 validated parent @vllm-omni-review-bot please re-review this exact head. |
vllm-omni-review-bot
left a comment
There was a problem hiding this comment.
Omni ReviewBot review
Scan:
| Category | Result |
|---|---|
| Tests / verification | 4 finding(s) below |
| Security | no finding reported |
| Docs / comments | 7 finding(s) below |
| Behavior / compatibility | no finding reported |
| Correctness | 1 finding(s) below |
Validated:
- [resolved] TorchCodec ABI revalidation (hsliuustc0106): not re-raised; residual is still the script’s early-exit if base already imports VideoDecoder — author cited AMD #11682 TorchCodec v0.10.0 build on parent f8af7e2
- [resolved] CUDA/XPU/NPU still on v0.28 (Dockerfile.cuda:1, Dockerfile.xpu:21, Dockerfile.npu:2) — deferred by author to #7410; residual siblings remain unaligned with Dockerfile.ci
- [resolved] Prior review items on Path parents[2], CI_DOCKERFILE naming, canary startswith, USE_NIGHTLY=0 lock, template --build-arg forms present at head
- [resolved] CUDA/XPU/NPU v0.28 drift — author #7410; residual siblings still unaligned by design.
- [resolved] USE_NIGHTLY_BUILD=0 lock + unquoted template override bans + canary after nightly — present at head (test:41-54, Dockerfile:14,45).
- [resolved] Prebuilt vllm/vllm-omni-rocm:v0.28.0 left in place — intentional per #7405/PR body; residual is missing release-only disclaimer (comment below), not a tag bump ask.
Keep four items: the unrendered set-up-using-python version contract (major), the pre-built-images missing until-#7405 / backwards “below” note (minor, absorbing 0 and 1), the AMD exact-head canary verification ask (minor question, absorbing 5), and the BASE_IMAGE override test-integrity hole (minor). Drop the test_rocm_* rename nit as pure taxonomy polish.
Verdict: REQUEST CHANGES
Findings
- [P2] Residual (not blocking): tools/install_torchcodec_rocm.sh still early-exits if… — ``
Residual (not blocking): tools/install_torchcodec_rocm.sh still early-exits if the base image already imports VideoDecoder, so a preinstalled TorchCodec can skip the ROCm ABI rebuild the Dockerfile intends. Pin remains TORCHCODEC_BRANCH=v0.10.0. Script unchanged by this PR.
Evidence: unchanged by this diff, present in the PR-time tree: tools/install_torchcodec_rocm.sh:13 TORCHCODEC_BRANCH="${TORCHCODEC_BRANCH:-v0.10.0}"; tools/install_torchcodec_rocm.sh:22-25 if verify_torchcodec >/dev/null 2>&1; then / echo "TorchCodec is already installed and working." / exit 0 / fi — early-exit skips rebuild when base already imports VideoDecoder; Dockerfile.rocm:59 invokes bash /tmp/install_torchcodec_rocm.sh after copying that script.
- [P2] Residual siblings remain on v0.28 while CI is v0.29: docker/Dockerfile.cuda:1,… —
Dockerfile.cuda:1
Residual siblings remain on v0.28 while CI is v0.29: docker/Dockerfile.cuda:1, docker/Dockerfile.xpu:21, docker/Dockerfile.npu:2 vs docker/Dockerfile.ci:7 — ROCm-only in this PR; tracked under #7410.
Evidence: unchanged by this diff, present in the PR-time tree: docker/Dockerfile.cuda:1 ARG BASE_IMAGE=vllm/vllm-openai:v0.28.0; docker/Dockerfile.xpu:21 ARG VLLM_VERSION=v0.28.0; docker/Dockerfile.npu:2 ARG VLLM_ASCEND_VERSION=v0.28.0; docker/Dockerfile.ci:7 ARG VLLM_BASE_TAG=v0.29.0 (ROCm updated in-diff: docker/Dockerfile.rocm:4 ARG BASE_IMAGE=vllm/vllm-openai-rocm:v0.29.0)
- [P2] Leaving
vllm/vllm-omni-rocm:v0.28.0is intentional (#7405 / source-build note) — ``
Leavingvllm/vllm-omni-rocm:v0.28.0is intentional (#7405 / source-build note). Residual: `pre-built-images` still has no release-only disclaimer next to the v0.28.0 examples (unlike wheels at line 26)—docs clarity only, not a tag-bump ask.
Evidence: docs/getting_started/installation/gpu/rocm.inc.md:169 vllm/vllm-omni-rocm:v0.28.0 \ and :190 vllm/vllm-omni-rocm:v0.28.0 — prebuilt examples still on 0.28.0. Contrast wheels disclaimer at :26 These pre-built wheel instructions install the published vLLM-Omni 0.28.0 release. For the 0.29 development line, use the source-install instructions below. — no equivalent beside the prebuilt-images commands (:149–191). Intentional-not-bump: :93–95 This upstream base is distinct from the prebuilt vllm/vllm-omni-rocm images discussed below; their published-tag availability is tracked in [#7405](...). and docker/Dockerfile.rocm:4 ARG BASE_IMAGE=vllm/vllm-openai-rocm:v0.29.0 (upstream base only).
| development line uses vLLM 0.29.x. Published 0.28.0 wheels and images use vLLM | ||
| 0.28.x. | ||
|
|
||
| The Dockerfile's `BASE_IMAGE` pin applies only to Docker builds. The |
There was a problem hiding this comment.
[P1] This diff places the new BASE_IMAGE-vs-non-Docker / “vllm-omni does not insta…
This diff places the new BASE_IMAGE-vs-non-Docker / “vllm-omni does not install vLLM” version-contract prose under # --8<-- [start:set-up-using-python] in rocm.inc.md (lines 11–20). Cited gpu.md includes only requirements, pre-built-wheels, build-wheel-from-source, pre-built-images, and build-docker — not set-up-using-python — so that unique contract never appears on the published GPU install page (README’s short 0.28/0.29 note does not replace the BASE_IMAGE/no-dep sentences). Move the prose into a rendered snippet (pre-built-wheels and/or build-wheel-from-source), or add an explicit --8<-- …:set-up-using-python include in gpu.md.
Evidence: rocm.inc.md:9 # --8<-- [start:set-up-using-python] then 17–20 BASE_IMAGE/no-dep prose; gpu.md:47/67/87/105 include only pre-built-wheels / build-wheel-from-source / pre-built-images / build-docker — no set-up-using-python include.
Suggestion: The Dockerfile's BASE_IMAGE pin applies only to Docker builds. The
vllm-omni package does not install vLLM as a dependency, so non-Docker source
installs must install the matching ROCm vLLM release explicitly before
installing vLLM-Omni, as shown below.
Installation of vLLM
|
|
||
| # Fail during image construction instead of letting every AMD test fail at | ||
| # runtime when the base image and vLLM-Omni target incompatible vLLM APIs. | ||
| RUN python3 -c "import vllm; from vllm.v1.kv_cache_interface import compute_layout_strides; from vllm.v1.kv_cache_layout import KVCacheLayout; print(f'Validated vLLM {vllm.__version__}')" |
There was a problem hiding this comment.
[P2] The fail-fast canary at docker/Dockerfile.rocm:45 (Validated vLLM …) only r…
The fail-fast canary at docker/Dockerfile.rocm:45 (Validated vLLM …) only runs during image construction. L1 tests/buildkite/test_rocm_dockerfile.py only asserts Dockerfile/template text (including canary string presence). Unchanged amd-build at .buildkite/amd/test-template-amd-omni.j2:77 builds -f docker/Dockerfile.rocm without overriding BASE_IMAGE/USE_NIGHTLY_BUILD, so that job is the real gate for vllm/vllm-openai-rocm:v0.29.0. Please link a green amd-build for 15d3c106799d3a90db64df8bdf8a51f961094dec showing Validated vLLM 0.29.0, or explicitly accept parent #11682 plus this docs-only tip before merge.
Evidence: docker/Dockerfile.rocm:4 ARG BASE_IMAGE=vllm/vllm-openai-rocm:v0.29.0; docker/Dockerfile.rocm:45 RUN python3 -c "import vllm; from vllm.v1.kv_cache_interface import compute_layout_strides; from vllm.v1.kv_cache_layout import KVCacheLayout; print(f'Validated vLLM {vllm.__version__}')"; tests/buildkite/test_rocm_dockerfile.py:64-76 def test_rocm_dockerfile_contains_vllm_api_canary() only _line_index/assert on Dockerfile text for the canary import string (no image build); unchanged by this diff, present in the PR-time tree: .buildkite/amd/test-template-amd-omni.j2:77 "docker build -f docker/Dockerfile.rocm -t {{ docker_image_amd }} --target test --build-arg ARG_PYTORCH_ROCM_ARCH=gfx942 --progress plain ." (no --build-arg BASE_IMAGE / USE_NIGHTLY_BUILD).
| pytestmark = [pytest.mark.core_model, pytest.mark.cpu] | ||
|
|
||
| REPO_ROOT = Path(__file__).resolve().parents[2] | ||
| AMD_TEMPLATE = REPO_ROOT / ".buildkite/amd/test-template-amd-omni.j2" |
There was a problem hiding this comment.
[P2] Rename tests/buildkite/test_rocm_dockerfile.py to an amd-prefixed sibling (e.g
Rename tests/buildkite/test_rocm_dockerfile.py to an amd-prefixed sibling (e.g. test_amd_rocm_dockerfile.py alongside test_amd_{bootstrap,pipeline,suite_selection}.py). This file locks the AMD CI template .buildkite/amd/test-template-amd-omni.j2 but breaks the existing test_amd_* AMD-contract naming.
Evidence: tests/buildkite/test_rocm_dockerfile.py:11 AMD_TEMPLATE = REPO_ROOT / ".buildkite/amd/test-template-amd-omni.j2" — new file path is test_rocm_dockerfile.py while locking the AMD template. Unchanged by this diff, present in the PR-time tree: siblings tests/buildkite/test_amd_bootstrap.py, tests/buildkite/test_amd_pipeline.py, tests/buildkite/test_amd_suite_selection.py use the test_amd_* AMD-contract naming.
|
|
||
| # Fail during image construction instead of letting every AMD test fail at | ||
| # runtime when the base image and vLLM-Omni target incompatible vLLM APIs. | ||
| RUN python3 -c "import vllm; from vllm.v1.kv_cache_interface import compute_layout_strides; from vllm.v1.kv_cache_layout import KVCacheLayout; print(f'Validated vLLM {vllm.__version__}')" |
There was a problem hiding this comment.
[P2] The image canary only imports compute_layout_strides and KVCacheLayout (sam…
The image canary only imports compute_layout_strides and KVCacheLayout (same 0.29 pair as vllm_omni/diffusion/diffusion_kv/layout.py:23-24); a green canary still leaves other missing 0.29 APIs beyond those two imports unguarded. Is that two-symbol surface the accepted fail-fast contract for this bump, or should the RUN gate cover additional 0.29 call sites?
Evidence: docker/Dockerfile.rocm:45 RUN python3 -c "import vllm; from vllm.v1.kv_cache_interface import compute_layout_strides; from vllm.v1.kv_cache_layout import KVCacheLayout; print(f'Validated vLLM {vllm.__version__}')" — canary gates only those two symbols. Unchanged by this diff beyond alignment (present in PR-time tree): vllm_omni/diffusion/diffusion_kv/layout.py:23-24 from vllm.v1.kv_cache_interface import KVCacheSpec, KVCacheTensor, compute_layout_strides / from vllm.v1.kv_cache_layout import KVCacheLayout — same compute_layout_strides + KVCacheLayout pair the canary mirrors.
| # Keep the ROCm runtime aligned with the vLLM release targeted by | ||
| # docker/Dockerfile.ci. A stale base image can build successfully while failing | ||
| # at runtime when vLLM-Omni imports APIs added by the newer vLLM release. | ||
| ARG BASE_IMAGE=vllm/vllm-openai-rocm:v0.29.0 |
There was a problem hiding this comment.
[P2] Default BASE_IMAGE (and paired VLLM_VERSION_OR_COMMIT_HASH) bump is inherit…
Default BASE_IMAGE (and paired VLLM_VERSION_OR_COMMIT_HASH) bump is inherited on rebuild by every consumer of docker/Dockerfile.rocm — AMD amd-build (.buildkite/amd/test-template-amd-omni.j2 builds without --build-arg BASE_IMAGE/USE_NIGHTLY_BUILD; merge suite steps depends_on: amd-build), user docker build -f docker/Dockerfile.rocm, and recipes that cite that Dockerfile (Ming-TTS, SenseNova, Qwen3-TTS, Stable-Audio, OmniVoice, MammothModa2). Accept that shared default for those recipe paths, or call out any that must stay pinned off the new base.
Evidence: docker/Dockerfile.rocm:4 ARG BASE_IMAGE=vllm/vllm-openai-rocm:v0.29.0; docker/Dockerfile.rocm:15 ARG VLLM_VERSION_OR_COMMIT_HASH=v0.29.0. Unchanged by this diff, present in the PR-time tree: .buildkite/amd/test-template-amd-omni.j2:77 "docker build -f docker/Dockerfile.rocm -t {{ docker_image_amd }} --target test --build-arg ARG_PYTORCH_ROCM_ARCH=gfx942 --progress plain ." (no BASE_IMAGE/USE_NIGHTLY_BUILD override). PR adds tests/buildkite/test_rocm_dockerfile.py:52-54 asserting AMD build must not pass --build-arg BASE_IMAGE/USE_NIGHTLY_BUILD. Unchanged by this diff: recipes cite docker/Dockerfile.rocm (e.g. recipes/inclusionAI/Ming-omni-tts.md:197, recipes/SenseNova/SenseNova-U1.md:186, recipes/Qwen/Qwen3-TTS.md:359, recipes/StabilityAI/Stable-Audio-Open.md:174, recipes/k2-fsa/OmniVoice.md:47, recipes/MammothModa2/MammothModa2.md:140).
hsliuustc0106
left a comment
There was a problem hiding this comment.
Reviewed 15d3c10. No actionable findings. The five focused static regression tests and diff hygiene checks pass. AMD image-build and diffusion startup validation remain outstanding; Docker is unavailable in the local review environment.
Merge main 01a2f93, including AMD vllm-project#7395, XPU vllm-project#7441 and NPU vllm-project#7433. Preserve both NIXL environment controls and adapt the upstream terminal-chunk test mock to connector metadata forwarding. CPU validation: 1069 passed, 2 skipped, 3 subtests passed; three additional failures reproduced identically on pinned main under the same no-device/offline environment. Hardware CI and H3 E2E remain unverified. Signed-off-by: yuanwu <yuan.wu@intel.com>
* [Bugfix][Examples] Use --profiler-config flag in offline TTS examples (vllm-project#6763) Signed-off-by: Asthenia <asthenia0412@gmail.com> Co-authored-by: Asthenia <asthenia0412@gmail.com> * [Bugfix] Skip HWR store-size scans when no limit is configured (vllm-project#7131) Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [CI][ROCm] Route LTX2 Ulysses parity to two-GPU lane (vllm-project#7234) Signed-off-by: andyluo7 <andy.luo@amd.com> * [Bugfix][Model] GR00T-N1.7: honor the per-request seed for flow-matching noise (vllm-project#7253) Signed-off-by: liangmengh <liangmengh@nvidia.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> * Add vLLM-Omni library info to Hugging Face Hub requests (vllm-project#5381) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * [Bugfix][NPU] Limit MiniMax H3 modulation grid size (vllm-project#6794) Signed-off-by: KrystalRay <keeleiray@gmail.com> Co-authored-by: KrystalRay <keeleiray@gmail.com> * [Bugfix] Build the forced-aligner prompt without a chat template (word timestamps one bin late) (vllm-project#7240) Signed-off-by: Tianyao Wu <rayroy31@gmail.com> * [Refactor][Diffusion] Resolve offload topology through one plan resolver (vllm-project#7209) Signed-off-by: specture724 <specture724@gmail.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * [Doc] Add AI usage policy for contributions (vllm-project#7305) Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com> * [Bugfix][MiMo-Audio] Align code2wav decode with tokenizer device (vllm-project#6539) Signed-off-by: chaosansui <zzc15560846421@163.com> Signed-off-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com> * [Bugfix][MiniCPM-o] Fix the audio_embeds input path (vllm-project#5730) Signed-off-by: eval-dev <0xe5bca0@gmail.com> Signed-off-by: eval <74645252+eval-dev@users.noreply.github.com> * [Feat][OmniVoice]Support Varlen Attn, Request-Batch and Step-Execution (vllm-project#6408) Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com> * [Model] Add Audio8 TTS Preview 0.6B (DualAR, 44.1 kHz codec) (vllm-project#6157) Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com> Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com> * [Bugfix][Frontend] Accept the msgpack-numpy package's numpy markers on the OpenPI endpoint (vllm-project#6051) Signed-off-by: zjli2013 <leezhengjiang@126.com> Co-authored-by: Cursor <cursoragent@cursor.com> * [Frontend] Opt-in WebSocket TTS split_granularity and session seed (vllm-project#7046) Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu> Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * [Bugfix][Frontend] Clear the P0 multimodal cache through the renderer (vllm-project#7003) Signed-off-by: ZenAlexa <zimingwang945@gmail.com> * [Bugfix][Frontend] Enforce image pixel limits for video input references (vllm-project#6963) Signed-off-by: BANANASJIM <bananasjim1@gmail.com> * [Bugfix][TTS] Isolate shared Higgs v3 reference encode from request cancellation (vllm-project#7076) Signed-off-by: Allen Wu <allenwu2795@gmail.com> Co-authored-by: TRAE CLI <traecli@bytedance.com> * [Bugfix][CosyVoice3] Resolve hash snapshot pipeline (vllm-project#6896) Signed-off-by: xutianle <xutianle@fudan.edu.cn> * [CI] Skip Qwen3-Omni Server VAD multi-turn realtime test (vllm-project#7279) (vllm-project#7314) Signed-off-by: wangyu <410167048@qq.com> * [Bugfix][Magi2] Allow import without an active Triton driver (vllm-project#7239) Signed-off-by: andyluo7 <andy.luo@amd.com> * [Core] Split Omni connector model runner mixin (vllm-project#6903) Signed-off-by: natureofnature <wzliu@connect.hku.hk> * [Bugfix] Make LTX vocoder decoding deterministic (vllm-project#7231) Signed-off-by: mglyn <1203789601@qq.com> * [Doc] [Recipe] Add FLUX.1-schnell recipe for RTX 5090 32GB (vllm-project#7299) Signed-off-by: Sparks-M <41097544+Sparks-M@users.noreply.github.com> * [Doc] Qwen3-TTS: add 0.6B on 1x A100 40GB (vllm-project#7289) Signed-off-by: chi030303 <106855944+chi030303@users.noreply.github.com> * [Perf][Model] Add optimized LTX-2.5 DiffVAE operators (vllm-project#7308) Signed-off-by: mglyn <1203789601@qq.com> * [2/N] Add a minimal temporal chunk callback for MiniMax-H3 (vllm-project#7017) Signed-off-by: specture724 <specture724@gmail.com> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * [Feature][Diffusion] Expose detailed pipeline timings (vllm-project#6822) Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com> * [Bugfix] Resolve vllm-project#6931 hub FA3 on torch 2.13 via kernels 0.16.1 (vllm-project#7185) Signed-off-by: NumberWan <wantszkin2003@gmail.com> * [Bugfix][Ascend] fix npu 310/a5 bugs (vllm-project#6685) Signed-off-by: zouyizhou <zouyizhou@huawei.com> * [Bugfix][Engine] Group overlapping device stages into one sequential init component (vllm-project#7328) Signed-off-by: ZhengWG <zwg0606@gmail.com> Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com> * fix: reserve Qwen3-Omni NVFP4 backend fix (vllm-project#7200) Signed-off-by: kunkunblueberry <1833921874@qq.com> * [BugFix] Add field validators for /v1/audio/generate request (vllm-project#4741) Signed-off-by: Shaun Walsh <shaunwalsh24@gmail.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: Nick Cao <ncao@redhat.com> * [CI][ROCm] Match CUDA/NPU L2/L3 label routing (vllm-project#6966) Signed-off-by: andyluo7 <andy.luo@amd.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [CI/Build] Avoid duplicate stage CLI deploy config (vllm-project#7007) Signed-off-by: mershi <mershi@tencent.com> Co-authored-by: mershi <mershi@tencent.com> * [CI/Build][ROCm] Normalize SenseNova paged-decode hardware markers (vllm-project#6935) Signed-off-by: andyluo7 <andy.luo@amd.com> * [Model] Skip unused frame packing in Wan2.2 S2V (vllm-project#7155) Signed-off-by: hyw <yuweih205@gmail.com> * [Doc] Add dual DGX Spark MiniMax-H3 results (vllm-project#7343) Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com> * [Model] Optimize MOSS-TTS Local batched execution and streaming codec (vllm-project#7202) Signed-off-by: Sy03 <1370724210@qq.com> * [Bugfix][XPU] Restore N-D output shape for W8A16 FP8 linear (vllm-project#7301) Signed-off-by: Joshna Medisetty <joshna.medisetty@intel.com> Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com> * [Doc] Document num_outputs_per_prompt for /v1/videos (vllm-project#7341) Signed-off-by: Guangjian <hiro20833@gmail.com> * [Skills] Add perf-evidence isolation, stage-attribution, and realtime-contract requirements (vllm-project#6820) Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com> Co-authored-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com> * [Bugfix] Allow LLM replicas on different GPUs to initialize concurrently (vllm-project#7292) Signed-off-by: Gao Han <hgaoaf@connect.ust.hk> Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com> * [CI/Build] Stabilize LTX2 vocoder autocast test on ROCm (vllm-project#7336) Signed-off-by: andyluo7 <andy.luo@amd.com> * [NPU][CI] Add A5 and 310P CI support (vllm-project#6875) Signed-off-by: Weiming Liao <liaowm5@gmail.com> Co-authored-by: wangyu <53896905+yenuo26@users.noreply.github.com> * [Kernel] Enable LTX DiffVAE fusions on SM100 and SM103 (vllm-project#7350) Signed-off-by: mglyn <1203789601@qq.com> * [Bugfix][MiniCPM-o] Align structured chat content with native omni rendering (vllm-project#7344) Signed-off-by: Sy03 <1370724210@qq.com> * [Rebase] Rebase to vLLM 0.29.0 (vllm-project#7230) Signed-off-by: tzhouam <tzhouam@connect.ust.hk> Signed-off-by: Zhou Taichang <tzhouam@connect.ust.hk> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * [Refactor] P0.2: Migrate API server helpers out of api_server (vllm-project#5453) Signed-off-by: herotai214 <herotai214@gmail.com> * [CI] Stabilize Qwen3-Omni Server VAD E2E (vllm-project#7356) Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com> Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com> * [CI/Build] Diff-aware source_file_dependencies for CUDA/NPU pipelines (vllm-project#6597) Signed-off-by: wangyu <410167048@qq.com> Co-authored-by: Cursor <cursoragent@cursor.com> * [Core][Diffusion] Add a typed pre-D2H video media contract (vllm-project#6615) Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com> Signed-off-by: Samit <285365963@qq.com> Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com> Co-authored-by: Samit <285365963@qq.com> * [Bugfix] Bound HWR domain initialization lock waits (vllm-project#7128) Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Bugfix] Escalate diffusion worker shutdown and retain survivors (vllm-project#7126) Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Misc] Add standalone safetensors retention diagnostic (vllm-project#7145) Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [CI] Isolate layerwise offload memory measurements (vllm-project#6938) Signed-off-by: andyluo7 <andy.luo@amd.com> * [Model] Add Cosmos3 mixed W8A8/W8A16 and W4A4/W4A16 denoising (vllm-project#6560) Signed-off-by: Rahul Steiger <rsteiger@aws-cmh-slurm-1-vscode-04.cm.cluster> Signed-off-by: Wojciech Kutak <wkutak@nvidia.com> Co-authored-by: Rahul Steiger <rsteiger@nvidia.com> * [Test] Use public render_jinja_template in MiniCPM-o native template test (vllm-project#7362) Signed-off-by: tly <2200895168@qq.com> * [Bugfix] Fix video prewarm cache retention and cancel-restart delay (vllm-project#7363) Signed-off-by: psv666 <2693925048@qq.com> * Cosmos3 action policy improvements (vllm-project#6460) Signed-off-by: Maciej Bala <mbala@nvidia.com> Signed-off-by: MaciejBalaNV <mbala@nvidia.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [BugFix][CI] Restore diff-aware source filtering for post-merge L3 (vllm-project#7371) Signed-off-by: wangyu <410167048@qq.com> * [Bugfix] Fail when a diffusion LoRA adapter binds no layer (vllm-project#7349) Signed-off-by: Guangjian <hiro20833@gmail.com> * [Bugfix] Fix host-memory leak on aborted /v1/images/generations (vllm-project#6462) (vllm-project#6561) Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Refactor] Declare model-local KV held outside the paged manager (vllm-project#6171) Signed-off-by: Yueqian Lin <linyueqian@outlook.com> * [Realtime] Emit current (non-beta) OpenAI audio/transcript event names (vllm-project#7339) Signed-off-by: Nick Cao <ncao@redhat.com> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * [Bugfix][Core] Clean up failed HWR atomic metadata writes (vllm-project#6956) Signed-off-by: BANANASJIM <bananasjim1@gmail.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Bugfix] Keep MiniMax-H3 reference audio budgets separate (vllm-project#7281) Signed-off-by: david6666666 <530634352@qq.com> * [Bugfix] Fix Helios USP: per-component split for correct sequence parallelism (vllm-project#6930) Signed-off-by: yancaocn <yancaochn@163.com> Co-authored-by: yancaocn <yancaochn@163.com> * [Perf][Diffusion] Optimize HSDP startup via Rank-0 shared weight loading and accelerated LoRA delta computation (vllm-project#7005) Signed-off-by: samithuang <285365963@qq.com> * [Example] Migrate HunyuanImage-3.0 to model_extras + shared task examples (vllm-project#5559) Signed-off-by: suyanli220 <suyanli220@gmail.com> Signed-off-by: suyan.li <suyan.li@bytedance.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: suyan.li <suyan.li@bytedance.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Model] Avoid scalar synchronizations in GLM-Image preparation (vllm-project#7172) Signed-off-by: hyw <yuweih205@gmail.com> * [Model][ERNIE-Image] Delay AdaLN modulation broadcast (vllm-project#7171) Signed-off-by: hyw <yuweih205@gmail.com> * [Kernel][MiniMax-H3] Run Q/K RMSNorm-RoPE in one launch (vllm-project#7167) Signed-off-by: hyw <yuweih205@gmail.com> * [CI][ROCm] Align AMD image with vLLM 0.29 (vllm-project#7395) Signed-off-by: andyluo7 <andy.luo@amd.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Bugfix] Add embed_multimodal to MiniCPM-o 4.5 omni LLM class (vllm-project#7384) Signed-off-by: Guangjian <hiro20833@gmail.com> * [Model] Add LingBot World Ulysses sequence parallelism (vllm-project#6841) Signed-off-by: wtz2333 <2955110911@qq.com> Co-authored-by: Zhou Taichang <tzhouam@connect.ust.hk> * [Feature][TTS] Add Speech API streaming metrics (vllm-project#6853) Signed-off-by: XIN GAO <1037396230@qq.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Bugfix][Model] Fix FLUX.2 Klein multi-image edit metadata (vllm-project#7430) Signed-off-by: QI JIA <qi.jia@shengshu.ai> Co-authored-by: QI JIA <qi.jia@shengshu.ai> Co-authored-by: Cursor <cursoragent@cursor.com> * [BugFix] Fix leftovers of the legacy OpenAI realtime API event names (vllm-project#7426) Signed-off-by: Nick Cao <ncao@redhat.com> Co-authored-by: Codex <noreply@openai.com> * [Model] Add Tencent AuK speech generation and editing (encoder + diffusion pipeline) (vllm-project#7385) Signed-off-by: Yueqian Lin <linyueqian@outlook.com> Co-authored-by: Sy03 <1370724210@qq.com> * [XPU][Docker] Align XPU image and CI with vLLM v0.29.0 (vllm-project#7441) Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com> * [Bugfix] Add explicit error when using CFGP with distilled Cosmos3 models (vllm-project#7427) Signed-off-by: Maciej Bala <mbala@nvidia.com> * [Perf][Diffusion] Run MammothModa2 DiT attention through the shared attention layer (vllm-project#7094) Signed-off-by: MrlixiangWE <mrdanaer@gmail.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Bugfix] Give model CLI flags typed owners in the Omni config (vllm-project#7390) Signed-off-by: Guangjian <hiro20833@gmail.com> Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com> * [Bugfix] Require a model for `vllm serve --omni` (fixes vllm-project#4158) (vllm-project#4167) Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com> * [Bugfix] Send a downstream terminal chunk when a parked stage ends (vllm-project#6889) Signed-off-by: psv666 <2693925048@qq.com> * [NPU] upgrade to v0.29.0 (vllm-project#7433) Signed-off-by: Weiming Liao <liaowm5@gmail.com> * [Bugfix][Model][Lance] Support decoded video frames in video editing (vllm-project#5128) Signed-off-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com> Co-authored-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com> * [Refactor][Diffusion] Remove model-specific names from LoRA and ModelOpt loader defaults (vllm-project#5907) Signed-off-by: Alicia <115451386+congw729@users.noreply.github.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * Optimize CosyVoice3 Stage1 flow batching (vllm-project#4876) Signed-off-by: gerayking <399geray@gmail.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [3/N] Encode streamed video on the worker with bounded batching (vllm-project#7018) Signed-off-by: specture724 <specture724@gmail.com> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Kernel][Boogu-Image] Fuse Q/K RMSNorm + interleaved RoPE via fused_qk_norm_rope (vllm-project#6982) Signed-off-by: Qihan Kang <rollykanggg@gmail.com> * [Bugfix][Frontend] Honor output_compression on the image generations route (vllm-project#7447) Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com> Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com> Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com> --------- Signed-off-by: Asthenia <asthenia0412@gmail.com> Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com> Signed-off-by: andyluo7 <andy.luo@amd.com> Signed-off-by: liangmengh <liangmengh@nvidia.com> Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> Signed-off-by: KrystalRay <keeleiray@gmail.com> Signed-off-by: Tianyao Wu <rayroy31@gmail.com> Signed-off-by: specture724 <specture724@gmail.com> Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com> Signed-off-by: chaosansui <zzc15560846421@163.com> Signed-off-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com> Signed-off-by: eval-dev <0xe5bca0@gmail.com> Signed-off-by: eval <74645252+eval-dev@users.noreply.github.com> Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com> Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com> Signed-off-by: zjli2013 <leezhengjiang@126.com> Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu> Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu> Signed-off-by: ZenAlexa <zimingwang945@gmail.com> Signed-off-by: BANANASJIM <bananasjim1@gmail.com> Signed-off-by: Allen Wu <allenwu2795@gmail.com> Signed-off-by: xutianle <xutianle@fudan.edu.cn> Signed-off-by: wangyu <410167048@qq.com> Signed-off-by: natureofnature <wzliu@connect.hku.hk> Signed-off-by: mglyn <1203789601@qq.com> Signed-off-by: Sparks-M <41097544+Sparks-M@users.noreply.github.com> Signed-off-by: chi030303 <106855944+chi030303@users.noreply.github.com> Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com> Signed-off-by: NumberWan <wantszkin2003@gmail.com> Signed-off-by: zouyizhou <zouyizhou@huawei.com> Signed-off-by: ZhengWG <zwg0606@gmail.com> Signed-off-by: kunkunblueberry <1833921874@qq.com> Signed-off-by: Shaun Walsh <shaunwalsh24@gmail.com> Signed-off-by: mershi <mershi@tencent.com> Signed-off-by: hyw <yuweih205@gmail.com> Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com> Signed-off-by: Sy03 <1370724210@qq.com> Signed-off-by: Joshna Medisetty <joshna.medisetty@intel.com> Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com> Signed-off-by: Guangjian <hiro20833@gmail.com> Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com> Signed-off-by: Gao Han <hgaoaf@connect.ust.hk> Signed-off-by: Weiming Liao <liaowm5@gmail.com> Signed-off-by: tzhouam <tzhouam@connect.ust.hk> Signed-off-by: Zhou Taichang <tzhouam@connect.ust.hk> Signed-off-by: herotai214 <herotai214@gmail.com> Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com> Signed-off-by: Samit <285365963@qq.com> Signed-off-by: Rahul Steiger <rsteiger@aws-cmh-slurm-1-vscode-04.cm.cluster> Signed-off-by: Wojciech Kutak <wkutak@nvidia.com> Signed-off-by: tly <2200895168@qq.com> Signed-off-by: psv666 <2693925048@qq.com> Signed-off-by: Maciej Bala <mbala@nvidia.com> Signed-off-by: MaciejBalaNV <mbala@nvidia.com> Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com> Signed-off-by: Yueqian Lin <linyueqian@outlook.com> Signed-off-by: Nick Cao <ncao@redhat.com> Signed-off-by: david6666666 <530634352@qq.com> Signed-off-by: yancaocn <yancaochn@163.com> Signed-off-by: samithuang <285365963@qq.com> Signed-off-by: suyanli220 <suyanli220@gmail.com> Signed-off-by: suyan.li <suyan.li@bytedance.com> Signed-off-by: wtz2333 <2955110911@qq.com> Signed-off-by: XIN GAO <1037396230@qq.com> Signed-off-by: QI JIA <qi.jia@shengshu.ai> Signed-off-by: MrlixiangWE <mrdanaer@gmail.com> Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com> Signed-off-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com> Signed-off-by: Alicia <115451386+congw729@users.noreply.github.com> Signed-off-by: gerayking <399geray@gmail.com> Signed-off-by: Qihan Kang <rollykanggg@gmail.com> Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com> Signed-off-by: José Carlos <jose@valendra.tech> Co-authored-by: Yancy <138764723+Asthenia0412@users.noreply.github.com> Co-authored-by: Asthenia <asthenia0412@gmail.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> Co-authored-by: andyluo7 <43718156+andyluo7@users.noreply.github.com> Co-authored-by: liangmenghuang <liangmengh@nvidia.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> Co-authored-by: Lei Ke <1141466880@qq.com> Co-authored-by: KrystalRay <keeleiray@gmail.com> Co-authored-by: Tianyao Wu <54675599+twu3202@users.noreply.github.com> Co-authored-by: Anjie Hou <149605198+specture724@users.noreply.github.com> Co-authored-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com> Co-authored-by: eval <74645252+eval-dev@users.noreply.github.com> Co-authored-by: boatman <1930807094@qq.com> Co-authored-by: NancyFyong <88076188+NancyFyong@users.noreply.github.com> Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com> Co-authored-by: zhengjia <ZJLi2013@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Rakesh Kariya <83279947+rk9595@users.noreply.github.com> Co-authored-by: Ziming Wang <125807850+ZenAlexa@users.noreply.github.com> Co-authored-by: Jim Ban <77719403+BANANASJIM@users.noreply.github.com> Co-authored-by: Allen Wu <85376543+EchoHayate@users.noreply.github.com> Co-authored-by: TRAE CLI <traecli@bytedance.com> Co-authored-by: xutianle <24210290017@m.fudan.edu.cn> Co-authored-by: wangyu <53896905+yenuo26@users.noreply.github.com> Co-authored-by: NATURE <wzliu@connect.hku.hk> Co-authored-by: Mu GuanLin <1203789601@qq.com> Co-authored-by: Sparks <41097544+Sparks-M@users.noreply.github.com> Co-authored-by: chi030303 <106855944+chi030303@users.noreply.github.com> Co-authored-by: Bo Li <22713281+bobboli@users.noreply.github.com> Co-authored-by: NumberWan <wantszkin2003@gmail.com> Co-authored-by: zyz111222 <zouyizhou@huawei.com> Co-authored-by: Zheng Wengang <zwg0606@gmail.com> Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com> Co-authored-by: kunkun <72174834+kunkunblueberry@users.noreply.github.com> Co-authored-by: Shaun Walsh <153730091+Shaun-Walsh@users.noreply.github.com> Co-authored-by: Nick Cao <ncao@redhat.com> Co-authored-by: shiyichuan <93317314+CarrotSwordsman@users.noreply.github.com> Co-authored-by: mershi <mershi@tencent.com> Co-authored-by: hyw <109567717+yuweih205@users.noreply.github.com> Co-authored-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com> Co-authored-by: Sy03 <1370724210@qq.com> Co-authored-by: Joshna-Medisetty <joshna.medisetty@intel.com> Co-authored-by: Guangjian Dong <163994576+Hiro208@users.noreply.github.com> Co-authored-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com> Co-authored-by: Gao Han <hgaoaf@connect.ust.hk> Co-authored-by: Weiming Liao <liaowm5@gmail.com> Co-authored-by: Zhou Taichang <tzhouam@connect.ust.hk> Co-authored-by: herotai214 <68222888+herotai214@users.noreply.github.com> Co-authored-by: LHXuuu <xulianhao.xlh@antgroup.com> Co-authored-by: Samit <285365963@qq.com> Co-authored-by: wkutak <wkutak@nvidia.com> Co-authored-by: Rahul Steiger <rsteiger@nvidia.com> Co-authored-by: tlysanhuo <166924864+tlysanhuo@users.noreply.github.com> Co-authored-by: psv666 <150513104+psv666@users.noreply.github.com> Co-authored-by: MaciejBalaNV <mbala@nvidia.com> Co-authored-by: summer <128961079+zhang-keliang@users.noreply.github.com> Co-authored-by: Yueqian Lin <70319226+linyueqian@users.noreply.github.com> Co-authored-by: WeiQing Chen <40507679+david6666666@users.noreply.github.com> Co-authored-by: Yan Cao <31481315+yancaocn@users.noreply.github.com> Co-authored-by: yancaocn <yancaochn@163.com> Co-authored-by: SuyanLi <126558907+suyanli220@users.noreply.github.com> Co-authored-by: suyan.li <suyan.li@bytedance.com> Co-authored-by: wtz2333 <2955110911@qq.com> Co-authored-by: GXIN <37653830+gxxx-hum@users.noreply.github.com> Co-authored-by: Qi Jia <kuafou@gmail.com> Co-authored-by: QI JIA <qi.jia@shengshu.ai> Co-authored-by: Codex <noreply@openai.com> Co-authored-by: DanaerLee <mrdanaer@gmail.com> Co-authored-by: longguo <107740309+abinggo@users.noreply.github.com> Co-authored-by: junpengw67-max <junpengw67@gmail.com> Co-authored-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com> Co-authored-by: Alicia <115451386+congw729@users.noreply.github.com> Co-authored-by: geray <48796550+gerayking@users.noreply.github.com> Co-authored-by: KANG Qihan <3149604185@qq.com>
Signed-off-by: andyluo7 <andy.luo@amd.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Summary
v0.28.0tov0.29.0BASE_IMAGEoverrideWhy
vLLM-Omni moved to the vLLM 0.29 API in #7230, including
compute_layout_strides, butdocker/Dockerfile.rocmstill installs vLLM 0.28. Current AMD merge builds therefore build successfully and then fail at runtime before model inference:d8d40625ffcaaa94Both report vLLM-Omni 0.29 with vLLM 0.28 and fail importing
compute_layout_stridesfromvllm.v1.kv_cache_interface. This affects the AMD merge suite generally; Stable Audio is only one visible quarantined leaf.Validation
Current exact head:
15d3c106799d3a90db64df8bdf8a51f961094dec--strict-markers:5 passed/tmpwith--strict-markers:5 passedgit diff --check: passedvllm/vllm-openai-rocm:v0.29.0forlinux/amd64v0.29.0contains the requiredcompute_layout_stridesandKVCacheLayoutAPIsAMD #11682 validated parent commit
f8af7e2d146cfa78629c36bd9ffd463f4f518b39before the final documentation/comment-only clarification:vllm/vllm-openai-rocm:v0.29.0and printedValidated vLLM 0.29.0;tools/install_torchcodec_rocm.shbuilt and installed TorchCodec v0.10.0;from torchcodec.decoders import VideoDecodercompleted;The final exact-head AMD rerun remains required before merge because the review clarification changed the PR head after #11682 started.
Scope
This changes the default base used by AMD CI and by users building
docker/Dockerfile.rocmwithout an override, plus static regression coverage and source-build documentation for that contract.The
BASE_IMAGEpin governs Docker builds. Non-Docker installs must install the matching platform-specific vLLM release explicitly; this PR does not add a global vLLM package requirement.It intentionally does not change the CUDA/XPU/NPU image pins in this PR. Their supported release contracts and remaining 0.28 defaults are tracked separately in #7410 rather than being bumped without backend-specific image and runtime validation.
It also does not change Stable Audio thresholds, offloader behavior, or blocking grades. Stable Audio follow-up remains #7396. The separate
vllm/vllm-omni-rocmpublished-image and prebuilt-installation documentation gap remains #7405.@yenuo26 @hsliuustc0106, please review the AMD image update and final documentation clarification at exact head
15d3c106799d3a90db64df8bdf8a51f961094dec.