Repository navigation
fix(ci): bump docker-build retries from 2 to 4 (5 attempts total) - #2016
Conversation
Interim mitigation for issue #2015: definitively confirmed the "failed to calculate checksum of ref ...: not found" failure is not a gha-cache problem (reproduces 6/6 with cache fully disabled, see #2013 and #2002). It's a local BuildKit race resolving a COPY --from reference while the shared installed-deps stage is concurrently extended by a second branch in the same build. Not deterministic — some same-day runs succeeded — so more attempts raise the odds one lands clean while the real fix (Dockerfile restructure, #2015) is pending.
|
Warning Review limit reached
Next review available in: 45 minutes You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (1)
Warning Billing warning: we have not been able to collect payment for this subscription for more than 72 hours. Please update the payment method or pay any pending invoices in Billing to avoid service interruption. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Failed to generate code suggestions for PR |
There was a problem hiding this comment.
No issues found across 1 file
Requires human review: Bumps CI build retries and changes failure semantics (continue-on-error) to mask a BuildKit race — an operational tradeoff of time vs. pass rate that a human should weigh.
Re-trigger cubic
There was a problem hiding this comment.
Graphify reviewed this change.
Looks safe to merge — no coupling regressions and no blocking issues, checked against the code graph (not a self-assessment).
Graphify review — findings
This PR modifies the Docker build retry logic in the CI workflow. It increases the number of retry attempts for the service build step from two retries (three attempts total) to four retries (five attempts total), adding two new retry steps that chain on the failure of all preceding attempts. It also rewrites the explanatory comment above the retry steps, changing the stated root cause of the transient failure from a stale/corrupted gha cache entry to a local BuildKit race condition (referencing issues #2015, #2013, #2002), and reframes the mitigation rationale from disabling cache to increasing attempt count.
No blocking issues surfaced. 3 lower-confidence candidates did not survive cross-model review.
Analysis details — impact, health, verification
Impact & health
Graphify review
Impact — 0 functions depend on the 0 functions this change touches.
Health — grade A; no new coupling hotspots.
Verification — 0 functions in the blast radius were not formally verified this run (proofs are advisory here).
Gate & verification
graphify gate
PASS — objectively clean (no health regressions, tests not run — proofs not run this pass (advisory)). Grounded, not self-assessed.
|
Fixes the real root cause behind issue #2015 (and #2002): installed-deps is concurrently extended by a second live branch (source-copied, via FROM inheritance) for the rest of the build. deps-production-base's three COPY --from=installed-deps steps were reading that stage while it was still being written to elsewhere in the same build DAG, which hit a BuildKit race resolving the cross-stage ref -- "failed to calculate checksum of ref ...: not found" -- non-deterministically but at a very high rate on GH Actions runners (10/10 recent attempts across #1956 and #1952, even fully cache-free). Retries and cache tuning (#2013, #2016) didn't fix it because it isn't a caching issue. Running npm ci directly in deps-production-base (which already branches straight from node:\${NODE_VERSION}, never touching installed-deps or build) fully decouples this lineage: no more cross-stage read of a stage still being concurrently written to. Cost is one extra @discordjs/opus native compile, once per build (this stage is shared by both deps-production-bot and deps-production-backend), paid once instead of the 3-5 min routinely wasted on exhausted retries. Not locally build-verified end-to-end: native C compilation under QEMU cross-arch emulation (amd64 on this arm64 Mac, via colima) segfaults GCC independent of this change (cc: internal compiler error: Segmentation fault signal terminated program cc1) -- a known QEMU/cross-compile instability class, not present on real amd64 hardware. Relying on CI (real amd64 runners) to validate.
## Root cause (issue #2015) \`installed-deps\` (an empty checkpoint = \`FROM build AS installed-deps\`) is consumed by two different lineages within a single \`docker buildx build\` invocation for any \`production-*\` target: 1. \`source-copied\` (\`FROM installed-deps\`) → \`build-shared\` → \`build-backend\`/\`build-bot\`/\`build-frontend\` — inherits and keeps extending \`installed-deps\`'s filesystem with more COPY/RUN instructions. 2. \`deps-production-base\` did \`COPY --from=installed-deps /app/node_modules ...\` and \`COPY --from=installed-deps /app/packages/shared/node_modules ...\`; \`deps-production-bot\`/\`deps-production-backend\` each did one more \`COPY --from=installed-deps\` for their own package's node_modules. Both lineages read/extend \`installed-deps\` concurrently — BuildKit parallelizes independent branches of the stage DAG by default. The \`COPY --from\` reads in branch 2 race against branch 1 still actively extending the same stage, and intermittently fail: \`\`\` ERROR: failed to calculate checksum of ref ...: "/app/packages/backend/node_modules": not found \`\`\` Confirmed **not** a caching issue: reproduced 10/10 recent attempts across #1956 and #1952 with the gha cache fully disabled (#2013), and retries alone don't reliably dodge it even with 5 attempts (#2016). ## Fix \`deps-production-base\` already branches straight from \`node:${NODE_VERSION}\` (never touches \`installed-deps\` or \`build\`). Give it its own \`npm ci\` instead of copying node_modules out of \`installed-deps\` — this fully removes the cross-stage read of a stage still being concurrently written to. \`deps-production-bot\`/\`deps-production-backend\` no longer need any \`COPY --from=installed-deps\` at all (npm workspaces already installs every workspace's node_modules from the root \`npm ci\`). Cost: one extra \`@discordjs/opus\` native compile, paid once per build (this stage is shared by both bot and backend production targets) — versus the 3-5 min routinely wasted on exhausted retries recently. ## Verification Not locally build-verified end-to-end — native C compilation under QEMU cross-arch emulation (amd64 on an arm64 Mac via colima) segfaults GCC independent of this change (\`cc: internal compiler error: Segmentation fault signal terminated program cc1\`), a known QEMU/cross-compile instability class not present on real amd64 hardware. Relying on this PR's own CI (real amd64 runners, matching production) to validate. ## Test plan - [ ] This PR's own \`docker-build\` matrix (bot/backend/frontend/nginx) passes, ideally on the very first attempt (no retries needed) — confirms the race is actually gone, not just dodged - [ ] \`Verify native modules load — bot\` step still passes (opus still loads correctly from the independently-installed node_modules) - [ ] Once merged, rebase #1956 and #1952 and confirm their Docker build check goes green cleanly <!-- This is an auto-generated description by cubic. --> --- ## Summary by cubic Decouples `deps-production-base` from `installed-deps` by running `npm ci` instead of copying `node_modules`. This removes a BuildKit race that intermittently failed production builds with checksum-not-found errors. - Replaces `COPY --from=installed-deps` with `npm ci` (with cache mount) in `deps-production-base`; removes the per-package `COPY --from=installed-deps` in `deps-production-bot` and `deps-production-backend`; keeps `npm prune --omit=dev`. - Fixes the root cause: `installed-deps` was read while another branch still extended it, triggering parallel BuildKit races (“failed to calculate checksum of ref …: not found”). - Impact: one extra `@discordjs/opus` native compile once per build; no runtime changes; workspace `npm ci` still installs all package `node_modules`. - Rollout: no migrations; CI should pass without retries. Rebase PRs previously failing on issue #2015 after merge. <sup>Written for commit 1b1a877. Summary will update on new commits.</sup> <a href="https://cubic.dev/pr/LucasSantana-Dev/Lucky/pull/2017?utm_source=github" target="_blank" rel="noopener noreferrer" data-no-image-dialog="true"><picture><source media="(prefers-color-scheme: dark)" srcset="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"><source media="(prefers-color-scheme: light)" srcset="https://www.cubic.dev/buttons/review-in-cubic-light.svg"><img alt="Review in cubic" src="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"></picture></a> <!-- End of auto-generated description by cubic. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Improved production container dependency installation and caching. * Streamlined separate production builds for bot and backend services. * Reduced reliance on shared dependency copies between build stages. <!-- end of auto-generated comment: release notes by coderabbit.ai -->
🤖 I have created a release *beep* *boop* --- <details><summary>2.39.4</summary> ## [2.39.4](v2.39.3...v2.39.4) (2026-08-14) ### Bug Fixes * **ci:** bump docker-build retries from 2 to 4 (5 attempts total) ([#2016](#2016)) ([25af082](25af082)) * **ci:** disable cache-to (not just cache-from) on docker-build retries ([#2013](#2013)) ([608f1d5](608f1d5)) * **ci:** serialize buildkit stage execution to kill the checksum race ([#2018](#2018)) ([ec4b5fd](ec4b5fd)) * **docker:** decouple deps-production-base from installed-deps ([#2017](#2017)) ([cfd3a5e](cfd3a5e)) </details> --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please). <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Release** * Updated the application version to **2.39.4**. * Documented improvements to CI retries, Docker caching, BuildKit handling, and production dependency installation. <!-- end of auto-generated comment: release notes by coderabbit.ai -->



Context
Follow-up to #2013. That PR gated `cache-to` by `use-cache` on retries so they run fully cache-free — this definitively ruled out the gha cache as the cause (see #2002's latest comments): the identical `failed to calculate checksum of ref ...: not found` error reproduced 6/6 across PR #1956 and #1952 with cache completely disabled.
Real cause is a local BuildKit race (filed as #2015) — not something to fix with another CI-config tweak, needs a Dockerfile restructure.
Interim mitigation
Since it's a race (not deterministic — some runs the same day succeeded), bump retries from 2 to 4 (5 attempts total) to raise the odds one attempt lands clean while #2015 is worked.
Test plan
Summary by cubic
Increase CI docker-build retries from 2 to 4 (5 attempts total) to mitigate a non-deterministic BuildKit race that triggers "failed to calculate checksum of ref ...: not found". Previously we retried twice with cache disabled on retries; now we retry four times, still disabling cache on retries, trading longer worst-case time for higher pass rates until the Dockerfile fix in issue #2015.
continue-on-error: truefor the middle attempts to allow progressing to the next retry; the final attempt determines job failure.use-cache: 'false'on all retries.loadbehavior remains the same.Written for commit 1c267f7. Summary will update on new commits.