Repository navigation
fix(ci): disable cache-to (not just cache-from) on docker-build retries - #2013
Conversation
Retries already set use-cache: 'false' to skip importing the gha layer cache after a "failed to calculate checksum of ref ...: not found" failure, but cache-to (mode=max export) stayed active regardless. Evidence from PR #1956/#1952 (2026-08-14): retry 1 and retry 2 both hit the identical error with cache-from already disabled, meaning the export side alone can trigger the same checksum-resolution failure — mode=max has to solve a cache key for every intermediate stage, including the branching installed-deps checkpoint two downstream stages read from concurrently. Gate cache-to by the same flag so retries run fully cache-free.
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (1)
📝 WalkthroughWalkthroughThe Docker build action documents cache import and export behavior during retries. The ChangesDocker cache behavior
Estimated code review effort: 2 (Simple) | ~10 minutes Merge Risk: ⚪ Minimal · up to This PR changes retry behavior so Docker builds skip both cache import and export; no actionable merge-blocking risk remains beyond normal checks and review. Possibly related PRs
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Warning Billing warning: we have not been able to collect payment for this subscription for more than 72 hours. Please update the payment method or pay any pending invoices in Billing to avoid service interruption. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Failed to generate code suggestions for PR |
There was a problem hiding this comment.
No issues found across 1 file
Auto-approved: CI-only fix gating cache-to by use-cache so retries are fully cache-free; bounded development config change with supporting evidence in the description.
Re-trigger cubic
There was a problem hiding this comment.
Graphify reviewed this change.
Worth a look — the grounded gate found no coupling regressions or blocking issues, but 1 advisory finding(s) below merit a look before merge.
Graphify review — findings
This PR modifies the docker-build-service composite GitHub Action so that the use-cache input now controls both cache import (cache-from) and cache export (cache-to), whereas previously cache-to always ran regardless of the input. It also expands the use-cache input description to document the reasoning, referencing specific PRs and BuildKit checksum failures observed on retries. The surface area is limited to one action definition file.
Worth a look
- use-cache=false no longer exports build cache —
.github/actions/docker-build-service/action.yml:74· Escalate · medium- agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review Execution auto-disposal is off for this run; enable it (with sandbox isolation) to have Graphify try to confirm or refute this automatically.
Analysis details — impact, health, verification
Impact & health
Graphify review
Impact — 0 functions depend on the 0 functions this change touches.
Health — grade A; no new coupling hotspots.
Verification — 0 functions in the blast radius were not formally verified this run (proofs are advisory here).
Gate & verification
graphify gate
PASS — objectively clean (no health regressions, tests not run — proofs not run this pass (advisory)). Grounded, not self-assessed.
|
) ## Context Follow-up to #2013. That PR gated \`cache-to\` by \`use-cache\` on retries so they run fully cache-free — this definitively ruled out the gha cache as the cause (see #2002's latest comments): the identical \`failed to calculate checksum of ref ...: not found\` error reproduced 6/6 across PR #1956 and #1952 with cache completely disabled. Real cause is a local BuildKit race (filed as #2015) — not something to fix with another CI-config tweak, needs a Dockerfile restructure. ## Interim mitigation Since it's a race (not deterministic — some runs the same day succeeded), bump retries from 2 to 4 (5 attempts total) to raise the odds one attempt lands clean while #2015 is worked. ## Test plan - [ ] This PR's own docker-build check passes <!-- This is an auto-generated description by cubic. --> --- ## Summary by cubic Increase CI docker-build retries from 2 to 4 (5 attempts total) to mitigate a non-deterministic BuildKit race that triggers "failed to calculate checksum of ref ...: not found". Previously we retried twice with cache disabled on retries; now we retry four times, still disabling cache on retries, trading longer worst-case time for higher pass rates until the Dockerfile fix in issue #2015. - Adds retry steps 2–4 and uses `continue-on-error: true` for the middle attempts to allow progressing to the next retry; the final attempt determines job failure. - Leaves the first attempt unchanged and keeps `use-cache: 'false'` on all retries. - No other workflow logic changes; service-specific `load` behavior remains the same. <sup>Written for commit 1c267f7. Summary will update on new commits.</sup> <a href="https://cubic.dev/pr/LucasSantana-Dev/Lucky/pull/2016?utm_source=github" target="_blank" rel="noopener noreferrer" data-no-image-dialog="true"><picture><source media="(prefers-color-scheme: dark)" srcset="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"><source media="(prefers-color-scheme: light)" srcset="https://www.cubic.dev/buttons/review-in-cubic-light.svg"><img alt="Review in cubic" src="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"></picture></a> <!-- End of auto-generated description by cubic. -->
Fixes the real root cause behind issue #2015 (and #2002): installed-deps is concurrently extended by a second live branch (source-copied, via FROM inheritance) for the rest of the build. deps-production-base's three COPY --from=installed-deps steps were reading that stage while it was still being written to elsewhere in the same build DAG, which hit a BuildKit race resolving the cross-stage ref -- "failed to calculate checksum of ref ...: not found" -- non-deterministically but at a very high rate on GH Actions runners (10/10 recent attempts across #1956 and #1952, even fully cache-free). Retries and cache tuning (#2013, #2016) didn't fix it because it isn't a caching issue. Running npm ci directly in deps-production-base (which already branches straight from node:\${NODE_VERSION}, never touching installed-deps or build) fully decouples this lineage: no more cross-stage read of a stage still being concurrently written to. Cost is one extra @discordjs/opus native compile, once per build (this stage is shared by both deps-production-bot and deps-production-backend), paid once instead of the 3-5 min routinely wasted on exhausted retries. Not locally build-verified end-to-end: native C compilation under QEMU cross-arch emulation (amd64 on this arm64 Mac, via colima) segfaults GCC independent of this change (cc: internal compiler error: Segmentation fault signal terminated program cc1) -- a known QEMU/cross-compile instability class, not present on real amd64 hardware. Relying on CI (real amd64 runners) to validate.
## Root cause (issue #2015) \`installed-deps\` (an empty checkpoint = \`FROM build AS installed-deps\`) is consumed by two different lineages within a single \`docker buildx build\` invocation for any \`production-*\` target: 1. \`source-copied\` (\`FROM installed-deps\`) → \`build-shared\` → \`build-backend\`/\`build-bot\`/\`build-frontend\` — inherits and keeps extending \`installed-deps\`'s filesystem with more COPY/RUN instructions. 2. \`deps-production-base\` did \`COPY --from=installed-deps /app/node_modules ...\` and \`COPY --from=installed-deps /app/packages/shared/node_modules ...\`; \`deps-production-bot\`/\`deps-production-backend\` each did one more \`COPY --from=installed-deps\` for their own package's node_modules. Both lineages read/extend \`installed-deps\` concurrently — BuildKit parallelizes independent branches of the stage DAG by default. The \`COPY --from\` reads in branch 2 race against branch 1 still actively extending the same stage, and intermittently fail: \`\`\` ERROR: failed to calculate checksum of ref ...: "/app/packages/backend/node_modules": not found \`\`\` Confirmed **not** a caching issue: reproduced 10/10 recent attempts across #1956 and #1952 with the gha cache fully disabled (#2013), and retries alone don't reliably dodge it even with 5 attempts (#2016). ## Fix \`deps-production-base\` already branches straight from \`node:${NODE_VERSION}\` (never touches \`installed-deps\` or \`build\`). Give it its own \`npm ci\` instead of copying node_modules out of \`installed-deps\` — this fully removes the cross-stage read of a stage still being concurrently written to. \`deps-production-bot\`/\`deps-production-backend\` no longer need any \`COPY --from=installed-deps\` at all (npm workspaces already installs every workspace's node_modules from the root \`npm ci\`). Cost: one extra \`@discordjs/opus\` native compile, paid once per build (this stage is shared by both bot and backend production targets) — versus the 3-5 min routinely wasted on exhausted retries recently. ## Verification Not locally build-verified end-to-end — native C compilation under QEMU cross-arch emulation (amd64 on an arm64 Mac via colima) segfaults GCC independent of this change (\`cc: internal compiler error: Segmentation fault signal terminated program cc1\`), a known QEMU/cross-compile instability class not present on real amd64 hardware. Relying on this PR's own CI (real amd64 runners, matching production) to validate. ## Test plan - [ ] This PR's own \`docker-build\` matrix (bot/backend/frontend/nginx) passes, ideally on the very first attempt (no retries needed) — confirms the race is actually gone, not just dodged - [ ] \`Verify native modules load — bot\` step still passes (opus still loads correctly from the independently-installed node_modules) - [ ] Once merged, rebase #1956 and #1952 and confirm their Docker build check goes green cleanly <!-- This is an auto-generated description by cubic. --> --- ## Summary by cubic Decouples `deps-production-base` from `installed-deps` by running `npm ci` instead of copying `node_modules`. This removes a BuildKit race that intermittently failed production builds with checksum-not-found errors. - Replaces `COPY --from=installed-deps` with `npm ci` (with cache mount) in `deps-production-base`; removes the per-package `COPY --from=installed-deps` in `deps-production-bot` and `deps-production-backend`; keeps `npm prune --omit=dev`. - Fixes the root cause: `installed-deps` was read while another branch still extended it, triggering parallel BuildKit races (“failed to calculate checksum of ref …: not found”). - Impact: one extra `@discordjs/opus` native compile once per build; no runtime changes; workspace `npm ci` still installs all package `node_modules`. - Rollout: no migrations; CI should pass without retries. Rebase PRs previously failing on issue #2015 after merge. <sup>Written for commit 1b1a877. Summary will update on new commits.</sup> <a href="https://cubic.dev/pr/LucasSantana-Dev/Lucky/pull/2017?utm_source=github" target="_blank" rel="noopener noreferrer" data-no-image-dialog="true"><picture><source media="(prefers-color-scheme: dark)" srcset="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"><source media="(prefers-color-scheme: light)" srcset="https://www.cubic.dev/buttons/review-in-cubic-light.svg"><img alt="Review in cubic" src="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"></picture></a> <!-- End of auto-generated description by cubic. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Improved production container dependency installation and caching. * Streamlined separate production builds for bot and backend services. * Reduced reliance on shared dependency copies between build stages. <!-- end of auto-generated comment: release notes by coderabbit.ai -->
Even fully isolated (no concurrent PR builds, no cache at all), the docker-build backend leg failed 5/5 attempts on both #1956 and #1952 with the same error the Dockerfile fix in #2017 partially addressed: failed to calculate checksum of ref ...: not found That fix (decoupling deps-production-base from installed-deps) validated clean once but didn't fully eliminate the race -- the same signature reappeared on a different COPY --from step (production-backend <- deps-production-backend), with the unrelated build stage's own npm ci still running concurrently in the log at the moment of failure. This is a BuildKit solver-level race under its default concurrent stage scheduling, confirmed independent of caching (#2013) and independent of cross-job/cross-PR concurrency (reproduces running fully alone). buildkitd-config-inline sets max-parallelism=1 on the ephemeral docker-container builder, forcing one stage operation at a time. Trades some build wall-clock for eliminating the race outright instead of retrying around it.
…#2018) ## Context Follow-up to #2017. That fix (decoupling `deps-production-base` from `installed-deps`) validated clean once (4/4 services, attempt 1, no retries) but didn't fully eliminate the race — rebuilding #1956 and #1952 afterward, both failed 5/5 attempts again with the identical error, now on a different `COPY --from` step (`production-backend` reading `deps-production-backend`), with the `build` stage's own unrelated `npm ci` still running concurrently in the log at the moment of failure. Ruled out as causes: gha cache (#2013 — fails identically fully cache-free), cross-job/cross-PR concurrency (fails identically running fully alone on an isolated rerun). This is a BuildKit solver-level race under its default concurrent stage scheduling — confirmed via Docker's own docs that `max-parallelism` is exactly the documented knob for "particularly useful for low-powered machines" style solver concurrency issues. ## Fix `buildkitd-config-inline` on the `docker/setup-buildx-action` step sets `max-parallelism = 1` (both `[worker.oci]` and `[worker.containerd]`, covering whichever worker the image uses), forcing BuildKit to execute one stage operation at a time instead of racing multiple branches concurrently. Trade-off: slower builds (no more parallel stage execution) in exchange for actually eliminating the race instead of retrying around it. Can tune back up (e.g. `max-parallelism = 2`) later if this proves too conservative once we have clean data. ## Test plan - [ ] This PR's own docker-build matrix passes on attempt 1 (no retries needed) — the real signal that the race is gone - [ ] Note build duration vs previous runs to gauge the serialization cost <!-- This is an auto-generated description by cubic. --> --- ## Summary by cubic Serializes Docker BuildKit stage scheduling in CI and production publishing to eliminate intermittent "failed to calculate checksum of ref ...: not found" during COPY --from. Previously stages ran concurrently; now both workflows set max-parallelism=1, trading some build speed for reliability. **Details** - Configure `docker/setup-buildx-action` with `buildkitd-config-inline` setting `[worker.oci]` and `[worker.containerd]` `max-parallelism = 1` in `.github/actions/docker-build-service/action.yml` and `.github/workflows/docker-publish.yml`. - Update `.github/workflows/ci.yml` path filter to include `.github/actions/docker-build-service/` so docker-build runs when its composite action changes. - Clarify Dockerfile header on reproducing vs suppressing the race; no functional Dockerfile changes. <sup>Written for commit 5c23bea. Summary will update on new commits.</sup> <a href="https://cubic.dev/pr/LucasSantana-Dev/Lucky/pull/2018?utm_source=github" target="_blank" rel="noopener noreferrer" data-no-image-dialog="true"><picture><source media="(prefers-color-scheme: dark)" srcset="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"><source media="(prefers-color-scheme: light)" srcset="https://www.cubic.dev/buttons/review-in-cubic-light.svg"><img alt="Review in cubic" src="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"></picture></a> <!-- End of auto-generated description by cubic. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved Docker image build reliability by serializing build stages, helping prevent intermittent checksum-related failures. <!-- end of auto-generated comment: release notes by coderabbit.ai -->
🤖 I have created a release *beep* *boop* --- <details><summary>2.39.4</summary> ## [2.39.4](v2.39.3...v2.39.4) (2026-08-14) ### Bug Fixes * **ci:** bump docker-build retries from 2 to 4 (5 attempts total) ([#2016](#2016)) ([25af082](25af082)) * **ci:** disable cache-to (not just cache-from) on docker-build retries ([#2013](#2013)) ([608f1d5](608f1d5)) * **ci:** serialize buildkit stage execution to kill the checksum race ([#2018](#2018)) ([ec4b5fd](ec4b5fd)) * **docker:** decouple deps-production-base from installed-deps ([#2017](#2017)) ([cfd3a5e](cfd3a5e)) </details> --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please). <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Release** * Updated the application version to **2.39.4**. * Documented improvements to CI retries, Docker caching, BuildKit handling, and production dependency installation. <!-- end of auto-generated comment: release notes by coderabbit.ai -->



Problem
Retries in
.github/workflows/ci.yml'sdocker-buildjob already setuse-cache: 'false'to skip importing the gha layer cache after afailed to calculate checksum of ref ...: not foundfailure — butcache-to(mode=max export) stayed active regardless ofuse-cache.Evidence
PR #1956 and #1952 (2026-08-14) both exhausted all 3 attempts (build + 2 retries) with the identical error, even though retry 1 and retry 2 both had
cache-fromdisabled:Cache import being off but the failure persisting identically means the export side alone can trigger the same checksum-resolution failure —
mode=maxhas to solve a cache key for every intermediate stage, including the branchinginstalled-depscheckpoint that two downstream stages (source-copiedanddeps-production-base) read from concurrently.Fix
Gate
cache-toby the sameuse-cacheflag so retries run a genuinely cache-free build (no import, no export) instead of only skipping import.Test plan
docker-buildmatrix job (backend/bot/frontend/nginx) passes, validating the composite action change worksBuild — Docker imagesrequired check goes greenSummary by cubic
Disable
cache-toondocker-buildretries whenuse-cacheis false to run a fully cache-free build and avoid the BuildKit “failed to calculate checksum of ref …: not found” error. Previously, retries disabled onlycache-fromwhile still exporting withmode=max, which could reproduce the same failure during export..github/actions/docker-build-service/action.yml, gatecache-tobyinputs.use-cache(mirrorscache-from). First attempts unchanged; retries now skip both import and export.docker-buildmatrix should pass and show no cache writes on retries.Build — Docker imagespasses.Written for commit 33bdce6. Summary will update on new commits.
Summary by CodeRabbit