Repository navigation
[CI/Build][ROCm] Cache AMD test images in registry - #7712
Conversation
|
Author self-review: I checked the Buildx command and cache-key inputs, confirmed the runtime image remains |
|
This PR was classified as CI work. CI owner: @yenuo26 @congw729 @NickCao Routing: @yenuo26 via semantic router, CI owner, CODEOWNERS; @congw729 via CODEOWNERS; @NickCao via CODEOWNERS @andyluo7, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
57d7b68 to
a98cde1
Compare
|
Self-review update for exact head
|
Omni ReviewBot: three questions on the performance claim@andyluo7 this PR reads as a performance or value claim:
Before the full evidence checklist, three short questions:
When you answer, the evidence that settles it is: base and head SHA, hardware, model, workload, warm-up and repeat count, mean or percentiles with their spread, and a correctness/quality-equivalence signal; an end-to-end claim also needs stage attribution. |
8a4f239 to
1e55176
Compare
|
@vllm-omni-review-bot Thanks. Scoped evidence for PR base
There is not a strict same-head before/after comparison because changing the Dockerfile/build script is the optimization itself; the tree-identical cold/warm pairs above isolate cache reuse, and the final source-only run isolates the intended PR workload. Correctness signal: every listed image job exited 0 with |
|
Author self-review for exact head
The timing data is one run per condition on different physical |
Omni ReviewBot routing recordAssigned Direct under experiment |
|
Exact-head CI note: AMD #12155's image job is green; the downstream |
vllm-omni-review-bot
left a comment
There was a problem hiding this comment.
Omni ReviewBot review
PR description
AMD CI image builds on ephemeral amd-cpu workers now use a Buildx docker-container builder with a registry-backed BuildKit cache in rocm/vllm-omni, instead of a plain docker build/docker push with no remote cache. Dockerfile.rocm splits dependency install into a deps stage and applies the checkout as a final COPY --link layer so source-only PR changes can reuse cached ROCm dependency layers. The Jinja AMD template calls the new build script; warm AMD runs in the PR body show image-job time dropping from tens of minutes to about one minute when the dependency cache key is unchanged.
Change flow
flowchart LR
A["[EXISTING] AMD amd-cpu image job"]:::existing --> B["[CHANGED] test-template-amd-omni.j2"]:::changed
B --> C["[NEW] build-ci-image.sh<br/>buildx registry cache"]:::new
C --> D["[CHANGED] Dockerfile.rocm<br/>deps then COPY --link"]:::changed
E["[NEW] Dockerfile.rocm.dockerignore"]:::new --> D
D --> F["[EXISTING] MI300 pulls rocm/vllm-omni:$COMMIT"]:::existing
classDef existing fill:#e5e7eb,stroke:#6b7280,color:#111827
classDef changed fill:#fef3c7,stroke:#d97706,color:#451a03,stroke-width:2px
classDef new fill:#dcfce7,stroke:#16a34a,color:#052e16,stroke-width:2px
classDef removed fill:#fee2e2,stroke:#dc2626,color:#450a0a,stroke-width:2px
No actionable findings.
85594af to
45660e4
Compare
|
Self-review update for exact head
Fresh exact-head CI is now pending. |
Signed-off-by: andyluo7 <andy.luo@amd.com>
Signed-off-by: andyluo7 <andy.luo@amd.com>
Signed-off-by: andyluo7 <andy.luo@amd.com>
45660e4 to
b3ec522
Compare
|
Self-review update for exact head
Upstream @yenuo26 @congw729 @NickCao, please review the current head when you have a chance. @vllm-omni-review-bot please review the current exact head. |
Omni ReviewBot attempt recordReview attempt ended as failed (Cursor could not run: [Errno 32] Broken pipe). |
|
@vllm-omni-review-bot please retry the review for current exact head |
Omni ReviewBot attempt recordReview attempt ended as failed (Cursor could not run: [Errno 32] Broken pipe). |
Signed-off-by: andyluo7 <andy.luo@amd.com> Signed-off-by: Matthieu Laneuville <matthieu.laneuville@surf.nl>
Signed-off-by: andyluo7 <andy.luo@amd.com>
Summary
Rebased onto current
mainafter #7771 merged. Current head:b3ec52293884ceae949b900df118df0f45ba56f5; rebase base:fa506e0fe80365c90d75701e44b4308ee8eacf2a.rocm/vllm-omnirepositoryDockerfile.rocm, its Dockerfile-specific ignore file, Python dependency metadata, all requirement files, the TorchCodec ROCm installer, and the ROCm architecturedepsstage, then add the checkout as one finalCOPY --linklayerbuildx build --pushinvocationMotivation
The AMD image job runs on ephemeral
amd-cpuworkers, so its former plaindocker buildhad no reusable remote cache. In AMD #12134, the image job took 20m36.599s. Its BuildKit stage timings attribute 340.1s to pulling/extracting the ROCm base (including 7.25 GB and 2.82 GB layers), 274.2s to OS packages, 239.3s to the TorchCodec build, and 145.1s to the project/ONNX Runtime installation, before the final image push.This change lets ordinary source-only PR updates reuse those dependency layers without downloading and extracting the large cached parents. Cache export is best-effort (
ignore-error=true) so a transient cache-export failure does not prevent publication of the runtime image.AMD runtime validation
Workload for every row is the gfx942 AMD CI image build on the
amd-cpuqueue usingvllm/vllm-openai-rocm:v0.29.0. All listed image jobs exited 0 and were not soft-failed.7ad1abd13smci250-ccs-aus-c12-4357d7b686csmci250-ccs-aus-c19-18a98cde1c2smci250-ccs-aus-c15-138a4f239bfsmci250-ccs-aus-c19-181e5517673smci250-ccs-aus-c16-2885594af63sc-hw-smc-acc-19The two cold/warm pairs use tree-identical commits:
57d7b686c/a98cde1c2for registry caching alone and8a4f239bf/1e5517673for the dependency/source split. The pre-rebase exact-head #12155 run changes onlytests/buildkite/test_rocm_dockerfile.py; its dependency key remainedrocm-deps-cache-e0d4b25c8c5daffe. Stages #8 through #21 wereCACHED, whileCOPY --link . .executed in 0.4s and the new image manifest was pushed successfully. It did not transfer the baseline's 7.25 GB or 2.82 GB parent layers.For the normal source-change workload, the observed image-job latency fell from 20m36.599s to 1m02.731s (about 95%, or 19.7x). Registry caching alone reduced a warm run to 6m49.464s; splitting dependency installation from the linked source layer reduced the warm path further. These are single observations on different physical workers, not a distribution, so this PR does not claim a mean, percentile, or spread.
As a correctness signal, #12155 published
rocm/vllm-omni:85594af63b0825c4c7f730dd0142b2281fb878f1, and dependent MI300 jobs started successfully from the exact-head pipeline. The cold #12150 build also executed (rather than restored) the vLLM API, ONNX Runtime ROCm-provider, and TorchCodec import validation steps before publishing its cache.Local validation
python3 -m pytest --confcutdir=tests/buildkite -o 'addopts=' -q tests/buildkite/test_rocm_dockerfile.py tests/buildkite/test_amd_bootstrap.py— 30 passedbash -n .buildkite/amd/scripts/build-ci-image.sh— passedgit diff --check— passedtests/buildkitesuite after rebase — 144 passed; 3 workstation-only failures intest_vllm_omni_package_discovery.pybecause this macOS Python lacks Torch and the checkout is not installed for off-root subprocess discoveryrocm-deps-cache-e0d4b25c8c5daffe, matching the prior warm run; the READY GPU phase was 42m03s, with the long tail in Simple Diffusion shard 4 and the pre-[CI][ROCm] Shorten the blocking CosyVoice test path #7748 full CosyVoice lane85594af63; current human review remains requestedCoordination
PR #7311 moves the image-build step from the Jinja template into
.buildkite/amd/bootstrap-upload-steps.yml. This PR targets currentmain; if #7311 lands first, its new image-build step should call this script instead of retaining the inlinedocker build/docker pushpair.