Repository navigation
fix(ci): skip gha cache-from on docker-build retries - #2008
Conversation
Root cause found via #2006's own CI run: all three attempts (build, retry 1, retry 2) failed identically within the same job, each time on the exact same buildx ref hash across multiple unrelated COPY paths (packages/shared/src/generated, packages/shared/dist, packages/bot/dist, prisma). cache-from is scoped identically (type=gha,scope=<service>) across every attempt in one job, so a retry that reimports a stale or corrupted gha cache entry fails identically instead of getting a fresh build - this is why the retry-count fix in #2007 did not save #2006's run. Add a use-cache input to docker-build-service (default true) and set it to false on both retry steps, so retries skip cache-from and build from scratch instead of re-hitting the same poisoned cache entry. cache-to stays on so a successful retry still writes a fresh cache for the next run.
|
Warning Review limit reached
Next review available in: 14 minutes You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (2)
Warning Billing warning: we have not been able to collect payment for this subscription for more than 72 hours. Please update the payment method or pay any pending invoices in Billing to avoid service interruption. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Failed to generate code suggestions for PR |
There was a problem hiding this comment.
No issues found across 2 files
Requires human review: Adjusts CI caching on retries (skipping cache-from), an operational tradeoff between build time and reliability that merits human judgment.
Re-trigger cubic
There was a problem hiding this comment.
Graphify reviewed this change.
Looks safe to merge — no coupling regressions and no blocking issues, checked against the code graph (not a self-assessment).
Graphify review — findings
This PR modifies the CI Docker build setup to control layer cache usage on retries. It adds a new use-cache input to the docker-build-service composite action, which conditionally includes the cache-from gha cache import based on the input value. In ci.yml, both retry steps now pass use-cache: 'false' so retries build without importing the layer cache, and the accompanying comments are updated to explain the reasoning.
No blocking issues surfaced. 2 lower-confidence candidates did not survive cross-model review.
Analysis details — impact, health, verification
Impact & health
Graphify review
Impact — 0 functions depend on the 0 functions this change touches.
Health — grade A; no new coupling hotspots.
Verification — 0 functions in the blast radius were not formally verified this run (proofs are advisory here).
Gate & verification
graphify gate
PASS — objectively clean (no health regressions, tests not run — proofs not run this pass (advisory)). Grounded, not self-assessed.
|
cache-from alone disabled (confirmed via #2008) did not fix the persistent Build - backend flake - reruns still fail deterministically even with a completely fresh, no-cache-from build. cache-to with mode=max exports every intermediate layer progressively as the build runs; a later stage's COPY --from can plausibly race against that same layer still being finalized for export. Disable cache-to too on retries so they run as a plain uncached build with no export pressure.
🤖 I have created a release *beep* *boop* --- <details><summary>2.39.3</summary> ## [2.39.3](v2.39.2...v2.39.3) (2026-08-12) ### Bug Fixes * **bot:** expose degraded-extractor signal and honest play error ([#1999](#1999)) ([8348bd6](8348bd6)) * **bot:** unlink dead Last.fm sessions on error 9 and notify once ([#1946](#1946)) ([d22b8bd](d22b8bd)) * **ci:** add a second retry for docker-build ([#2007](#2007)) ([9347c55](9347c55)) * **ci:** extract node_modules before per-workspace build steps run ([#2006](#2006)) ([ae577d1](ae577d1)) * **ci:** free disk space before Docker builds to prevent BuildKit GC race ([#2003](#2003)) ([5bcd388](5bcd388)) * **ci:** isolate docker cache scope by branch ([#2012](#2012)) ([5d25ec8](5d25ec8)) * **ci:** pin node:24-alpine base image by digest ([#2004](#2004)) ([f154bfe](f154bfe)) * **ci:** skip gha cache-from on docker-build retries ([#2008](#2008)) ([89dbb90](89dbb90)) * **ci:** split deps-production per production target ([#2005](#2005)) ([43aacce](43aacce)) </details> --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please).



Summary
Why
#2006's own CI run gave the missing piece: all three attempts (build, retry 1, retry 2) failed identically within the same job, each time on the exact same buildx ref hash across multiple unrelated COPY paths (packages/shared/src/generated, packages/shared/dist, packages/bot/dist, prisma). cache-from is scoped identically (type=gha,scope=) across every attempt in a job, so a retry that reimports a stale or corrupted gha cache entry fails identically instead of getting a fresh build. That's why the retry-count fix in #2007 did not save #2006's run.
Adds a
use-cacheinput todocker-build-service(default true) and sets it to false on both retry steps in ci.yml, so retries skip cache-from and build from scratch. cache-to stays on so a successful retry still writes a fresh cache for the next run.Test plan
Summary by cubic
Skip importing the GHA layer cache on docker-build retries to prevent repeated BuildKit failures. Retries now build fresh while still writing a new cache on success.
use-cacheinput to./.github/actions/docker-build-service(defaulttrue); set tofalseon retry steps inci.ymlto omitcache-from.cache-to: type=gha,mode=maxso successful retries repopulate the cache and resolves the "failed to calculate checksum of ref ...: not found" errors from staletype=gha,scope=<service>entries.Written for commit aee532a. Summary will update on new commits.