Skip to content

[CI/Build][ROCm] Cache AMD test images in registry - #7712

Merged
yenuo26 merged 3 commits into
vllm-project:mainfrom
andyluo7:ci/rocm-image-registry-cache
Sep 20, 2026
Merged

yenuo26 merged 3 commits into
vllm-project:mainfrom
andyluo7:ci/rocm-image-registry-cache

Conversation

@andyluo7

@andyluo7 andyluo7 commented Sep 17, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Rebased onto current main after #7771 merged. Current head: b3ec52293884ceae949b900df118df0f45ba56f5; rebase base: fa506e0fe80365c90d75701e44b4308ee8eacf2a.

  • build ROCm CI images with a Buildx docker-container builder
  • import and export a registry-backed BuildKit cache in the existing rocm/vllm-omni repository
  • key dependency caches on Dockerfile.rocm, its Dockerfile-specific ignore file, Python dependency metadata, all requirement files, the TorchCodec ROCm installer, and the ROCm architecture
  • install dependency-heavy layers in a source-independent deps stage, then add the checkout as one final COPY --link layer
  • build and publish the commit-tagged runtime image and cache in one buildx build --push invocation

Motivation

The AMD image job runs on ephemeral amd-cpu workers, so its former plain docker build had no reusable remote cache. In AMD #12134, the image job took 20m36.599s. Its BuildKit stage timings attribute 340.1s to pulling/extracting the ROCm base (including 7.25 GB and 2.82 GB layers), 274.2s to OS packages, 239.3s to the TorchCodec build, and 145.1s to the project/ONNX Runtime installation, before the final image push.

This change lets ordinary source-only PR updates reuse those dependency layers without downloading and extracting the large cached parents. Cache export is best-effort (ignore-error=true) so a transient cache-export failure does not prevent publication of the runtime image.

AMD runtime validation

Workload for every row is the gfx942 AMD CI image build on the amd-cpu queue using vllm/vllm-openai-rocm:v0.29.0. All listed image jobs exited 0 and were not soft-failed.

Run Configuration Source head Worker Image-job time
#12134 pre-change plain Docker build 7ad1abd13 smci250-ccs-aus-c12-43 20m36.599s
#12143 registry cache, cold seed 57d7b686c smci250-ccs-aus-c19-18 14m17.972s
#12147 registry cache, warm a98cde1c2 smci250-ccs-aus-c15-13 6m49.464s
#12150 dependency/source split, cold seed 8a4f239bf smci250-ccs-aus-c19-18 9m22.747s
#12152 dependency/source split, warm, identical tree to #12150 1e5517673 smci250-ccs-aus-c16-28 53.530s
#12155 dependency/source split, warm, source-only change 85594af63 sc-hw-smc-acc-19 1m02.731s

The two cold/warm pairs use tree-identical commits: 57d7b686c/a98cde1c2 for registry caching alone and 8a4f239bf/1e5517673 for the dependency/source split. The pre-rebase exact-head #12155 run changes only tests/buildkite/test_rocm_dockerfile.py; its dependency key remained rocm-deps-cache-e0d4b25c8c5daffe. Stages #8 through #21 were CACHED, while COPY --link . . executed in 0.4s and the new image manifest was pushed successfully. It did not transfer the baseline's 7.25 GB or 2.82 GB parent layers.

For the normal source-change workload, the observed image-job latency fell from 20m36.599s to 1m02.731s (about 95%, or 19.7x). Registry caching alone reduced a warm run to 6m49.464s; splitting dependency installation from the linked source layer reduced the warm path further. These are single observations on different physical workers, not a distribution, so this PR does not claim a mean, percentile, or spread.

As a correctness signal, #12155 published rocm/vllm-omni:85594af63b0825c4c7f730dd0142b2281fb878f1, and dependent MI300 jobs started successfully from the exact-head pipeline. The cold #12150 build also executed (rather than restored) the vLLM API, ONNX Runtime ROCm-provider, and TorchCodec import validation steps before publishing its cache.

Local validation

  • python3 -m pytest --confcutdir=tests/buildkite -o 'addopts=' -q tests/buildkite/test_rocm_dockerfile.py tests/buildkite/test_amd_bootstrap.py — 30 passed
  • changed-file pre-commit hooks — passed, including Ruff, mypy, SPDX, and CI-marker checks
  • bash -n .buildkite/amd/scripts/build-ci-image.sh — passed
  • git diff --check — passed
  • full tests/buildkite suite after rebase — 144 passed; 3 workstation-only failures in test_vllm_omni_package_discovery.py because this macOS Python lacks Torch and the checkout is not installed for off-root subprocess discovery
  • exact-head AMD #12326: all 27 command jobs passed with no soft failures or retries; the image job completed in 1m01.353s; the dependency cache key remains rocm-deps-cache-e0d4b25c8c5daffe, matching the prior warm run; the READY GPU phase was 42m03s, with the long tail in Simple Diffusion shard 4 and the pre-[CI][ROCm] Shorten the blocking CosyVoice test path #7748 full CosyVoice lane
  • exact-head CUDA #15649, Intel #8673, NPU #7378, wheel builds, pre-commit, DCO, and docs passed
  • the prior Omni ReviewBot review found no actionable findings on patch-equivalent head 85594af63; current human review remains requested

Coordination

PR #7311 moves the image-build step from the Jinja template into .buildkite/amd/bootstrap-upload-steps.yml. This PR targets current main; if #7311 lands first, its new image-build step should call this script instead of retaining the inline docker build / docker push pair.

@andyluo7

Copy link
Copy Markdown
Collaborator Author

Author self-review: I checked the Buildx command and cache-key inputs, confirmed the runtime image remains rocm/vllm-omni:$BUILDKITE_COMMIT, and verified that base/nightly Dockerfile defaults are not overridden. I also checked the #7311 overlap and documented the required integration point. Focused Buildkite tests (28), changed-file pre-commit/ShellCheck, shell syntax, and git diff --check pass. The remaining gate is two AMD image runs: one to seed the cache and one to confirm imported cache records, cached dependency stages, and warm-build timing.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR was classified as CI work.

CI owner: @yenuo26 @congw729 @NickCao

Routing: @yenuo26 via semantic router, CI owner, CODEOWNERS; @congw729 via CODEOWNERS; @NickCao via CODEOWNERS

@andyluo7, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@andyluo7 andyluo7 added ready label to trigger buildkite CI ROCm PR related to AMD hardware CI/CD codes related to changes to CI/CD and removed ready label to trigger buildkite CI labels Sep 17, 2026
@andyluo7
andyluo7 force-pushed the ci/rocm-image-registry-cache branch from 57d7b68 to a98cde1 Compare September 17, 2026 07:23
@andyluo7 andyluo7 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 17, 2026
@andyluo7

Copy link
Copy Markdown
Collaborator Author

Self-review update for exact head a98cde1c25b0b5fbf8c34d6a746dd376da7415ca:

  • verified the warm AMD run imported rocm-deps-cache-c9ef10d03b726273;
  • confirmed the expensive OS dependency, vLLM canary, and TorchCodec build stages were CACHED;
  • confirmed the commit-tagged runtime image and refreshed cache manifest were both pushed;
  • image job #12147 completed in 6m49s with exit_status=0 and soft_failed=false, versus the 20m36s baseline;
  • reviewed the residual cost: the current post-COPY . project install still forces cached-layer materialization and should be handled in a separate Dockerfile-layering follow-up.

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: three questions on the performance claim

@andyluo7 this PR reads as a performance or value claim:

  • claim: The warm image job is about 67% faster than the 20m36s baseline.

Before the full evidence checklist, three short questions:

  1. Bottleneck — what is the current bottleneck, and which profile, trace or per-stage measurement shows it?
  2. Value — what does the change buy the user or the system (latency, throughput, memory, cost), and at which workload?
  3. A/B or ablation — is there a same-workload, same-head/config comparison that isolates each main claim on its own? For stacked optimizations, one number per item rather than a blended delta.

When you answer, the evidence that settles it is: base and head SHA, hardware, model, workload, warm-up and repeat count, mean or percentiles with their spread, and a correctness/quality-equivalence signal; an end-to-end claim also needs stage attribution.

@andyluo7 andyluo7 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 17, 2026
@andyluo7
andyluo7 force-pushed the ci/rocm-image-registry-cache branch from 8a4f239 to 1e55176 Compare September 17, 2026 07:59
@andyluo7 andyluo7 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 17, 2026
@andyluo7
andyluo7 marked this pull request as draft September 17, 2026 08:04
@andyluo7
andyluo7 marked this pull request as ready for review September 17, 2026 08:04
@andyluo7 andyluo7 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 17, 2026
@andyluo7

Copy link
Copy Markdown
Collaborator Author

@vllm-omni-review-bot Thanks. Scoped evidence for PR base 8c3d3407d1bc9b1b538d2196f4ed2d802a15cfc3 and current head 85594af63b0825c4c7f730dd0142b2281fb878f1:

  1. Bottleneck. This is a container-image build workload (no model or GPU inference workload). The pre-change plain-Docker image job in AMD #12134, head 7ad1abd13, took 20m36.599s on the amd-cpu queue. BuildKit's per-stage timings show 340.1s pulling/extracting the ROCm base (including 7.25 GB and 2.82 GB layers), 274.2s installing OS packages, 239.3s building TorchCodec, and 145.1s installing the project/ONNX Runtime. The bottleneck was therefore repeated dependency work and large parent-layer materialization on ephemeral workers, rather than the small source update itself.

  2. Value. On the normal PR case where source changes but dependency inputs do not, exact-head AMD #12155 completed the image job in 1m02.731s on sc-hw-smc-acc-19, versus the observed 20m36.599s pre-change run: about 95% lower latency (19.7x). It also avoids repeating the associated amd-cpu, network, and local-storage work. This is one observation per condition on different physical workers, so I am not claiming a mean, percentile, or spread.

  3. Ablation. I used one cold seed and one warm repeat for each architecture, with tree-identical commits inside each cold/warm pair:

There is not a strict same-head before/after comparison because changing the Dockerfile/build script is the optimization itself; the tree-identical cold/warm pairs above isolate cache reuse, and the final source-only run isolates the intended PR workload.

Correctness signal: every listed image job exited 0 with soft_failed=false; #12155 pushed rocm/vllm-omni:85594af63b0825c4c7f730dd0142b2281fb878f1, and dependent MI300 jobs started from that exact-head pipeline. In the cold #12150 build, the vLLM API, ONNX Runtime ROCm-provider, and TorchCodec import checks executed rather than being restored from cache.

@andyluo7

Copy link
Copy Markdown
Collaborator Author

Author self-review for exact head 85594af63b0825c4c7f730dd0142b2281fb878f1:

  • checked that the registry cache key covers the ROCm Dockerfile, Dockerfile-specific ignore rules, package/dependency metadata, TorchCodec installer, and gfx architecture, while ordinary source changes leave that key stable;
  • checked the Dockerfile boundary: dependency installation and ROCm import canaries precede the final COPY --link, with no filesystem-dependent RUN/COPY/ADD afterward;
  • verified exact-head AMD #12155 imported rocm-deps-cache-e0d4b25c8c5daffe, restored dependency stages Add PR and issue templates from vLLM project #8-Code review documentation for PR #17: OmniGPUModelRunner and OmniModelInputForGPU - ✅ All Issues Resolved #21, rebuilt the changed source layer in 0.4s, pushed the commit-tagged image, and finished in 1m02.731s with exit 0 and no soft failure;
  • confirmed dependent MI300 jobs started from the exact-head pipeline;
  • reran focused Buildkite contracts (30 passed), changed-file pre-commit, shell syntax, and diff checks. The full local Buildkite suite is 142 passed with the same 3 workstation-only package-discovery failures caused by missing Torch/off-root installation.

The timing data is one run per condition on different physical amd-cpu workers; I have limited the PR claim accordingly.

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot routing record

Assigned Direct under experiment vllm-omni-strict-5050-20260829.

@andyluo7

Copy link
Copy Markdown
Collaborator Author

Exact-head CI note: AMD #12155's image job is green; the downstream Diffusion · Model Test red is the pre-existing ROCm routing failure in test_inductor_regions_match_eager_in_deterministic_mode (vLLM-bundled versioned FlashAttention requires CUDA). The same test/error occurred on the preceding #12152 image head. PR #7706 changes that module's skip predicate from torch.cuda.is_available() to current_omni_platform.is_cuda() and is the scoped fix. Other #12155 MI300 lanes are already completing successfully, so I am keeping that unrelated test-routing fix out of this image-cache PR.

@vllm-omni-review-bot vllm-omni-review-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Omni ReviewBot review

PR description

AMD CI image builds on ephemeral amd-cpu workers now use a Buildx docker-container builder with a registry-backed BuildKit cache in rocm/vllm-omni, instead of a plain docker build/docker push with no remote cache. Dockerfile.rocm splits dependency install into a deps stage and applies the checkout as a final COPY --link layer so source-only PR changes can reuse cached ROCm dependency layers. The Jinja AMD template calls the new build script; warm AMD runs in the PR body show image-job time dropping from tens of minutes to about one minute when the dependency cache key is unchanged.

Change flow

flowchart LR
  A["[EXISTING] AMD amd-cpu image job"]:::existing --> B["[CHANGED] test-template-amd-omni.j2"]:::changed
  B --> C["[NEW] build-ci-image.sh<br/>buildx registry cache"]:::new
  C --> D["[CHANGED] Dockerfile.rocm<br/>deps then COPY --link"]:::changed
  E["[NEW] Dockerfile.rocm.dockerignore"]:::new --> D
  D --> F["[EXISTING] MI300 pulls rocm/vllm-omni:$COMMIT"]:::existing
  classDef existing fill:#e5e7eb,stroke:#6b7280,color:#111827
  classDef changed fill:#fef3c7,stroke:#d97706,color:#451a03,stroke-width:2px
  classDef new fill:#dcfce7,stroke:#16a34a,color:#052e16,stroke-width:2px
  classDef removed fill:#fee2e2,stroke:#dc2626,color:#450a0a,stroke-width:2px
Loading

No actionable findings.

@hsliuustc0106 hsliuustc0106 added the high priority high priority issue, needs to be done asap label Sep 17, 2026
@andyluo7
andyluo7 force-pushed the ci/rocm-image-registry-cache branch from 85594af to 45660e4 Compare September 18, 2026 03:22
@andyluo7

Copy link
Copy Markdown
Collaborator Author

Self-review update for exact head 45660e4e484975345d385e560f9a16387b199a10 after rebasing onto merged #7706:

  • git range-diff confirms all three cache commits are patch-equivalent to the previously reviewed branch;
  • verified the shared AMD template retains [CI/Build][ROCm] Stabilize shared AMD test lanes #7706's per-step retry rendering while delegating the image gate to build-ci-image.sh;
  • focused Dockerfile/bootstrap contracts pass (30 passed), bash -n, all changed-file pre-commit hooks, and git diff --check pass;
  • the full local tests/buildkite suite is 144 passed with the same three workstation-only package-discovery failures caused by missing Torch and the checkout not being installed off-root.

Fresh exact-head CI is now pending.

@andyluo7 andyluo7 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 18, 2026
Signed-off-by: andyluo7 <andy.luo@amd.com>
Signed-off-by: andyluo7 <andy.luo@amd.com>
Signed-off-by: andyluo7 <andy.luo@amd.com>
@andyluo7
andyluo7 force-pushed the ci/rocm-image-registry-cache branch from 45660e4 to b3ec522 Compare September 19, 2026 17:24
@andyluo7 andyluo7 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 19, 2026
@andyluo7

Copy link
Copy Markdown
Collaborator Author

Self-review update for exact head b3ec52293884ceae949b900df118df0f45ba56f5:

  • rebased the three [CI/Build][ROCm] Cache AMD test images in registry #7712 commits onto upstream fa506e0fe; merged [CI/Build] Restrict ERNIE fused RoPE tests to NVIDIA CUDA #7771 is in the base ancestry, and git range-diff reports all three commits patch-equivalent;
  • verified all three commits retain their Signed-off-by trailers and the worktree is clean;
  • focused Docker/bootstrap contracts pass (30 passed), along with bash -n, changed-file pre-commit hooks, and git diff --check;
  • the full local tests/buildkite suite is 144 passed with only the same three workstation-only package-discovery failures caused by missing local Torch/off-root installation;
  • exact-head AMD #12326 passed all 27 command jobs with no soft failures or retries. The image job completed in 1m01.353s on amd-cpu; its dependency key remains rocm-deps-cache-e0d4b25c8c5daffe, the same key used by the prior warm-cache validation;
  • exact-head CUDA #15649, Intel #8673, NPU #7378, wheel, pre-commit, DCO, and docs checks all passed.

Upstream main advanced by one unrelated #7525 commit after this exact-head run completed; it does not touch the Docker/cache-key inputs changed here.

@yenuo26 @congw729 @NickCao, please review the current head when you have a chance. @vllm-omni-review-bot please review the current exact head.

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (Cursor could not run: [Errno 32] Broken pipe).

@andyluo7

Copy link
Copy Markdown
Collaborator Author

@vllm-omni-review-bot please retry the review for current exact head b3ec52293884ceae949b900df118df0f45ba56f5. The previous attempt ended before reviewing code because the Cursor backend returned [Errno 32] Broken pipe.

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (Cursor could not run: [Errno 32] Broken pipe).

@yenuo26 yenuo26 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@yenuo26
yenuo26 merged commit fa422a3 into vllm-project:main Sep 20, 2026
9 checks passed
mlaneuville pushed a commit to mlaneuville/vllm-omni that referenced this pull request Sep 22, 2026
Signed-off-by: andyluo7 <andy.luo@amd.com>
Signed-off-by: Matthieu Laneuville <matthieu.laneuville@surf.nl>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CI/CD codes related to changes to CI/CD high priority high priority issue, needs to be done asap ready label to trigger buildkite CI ROCm PR related to AMD hardware

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants