Repository navigation
fix(OMN-13687): bounded retry on docker push to ECR — absorb transient runner→ECR EOF (emergency prod recovery) - #2134
Merged
Conversation
…t runner→ECR EOF (emergency prod recovery)
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Context
EMERGENCY PROD RECOVERY — OMN-13687 (operator-authorized).
The runtime image build (
build-and-push-runtime.yml) builds and Trivy-scans cleanly on main HEAD2cd720817(the OMN-13670 crypto-floor hotfix #2131), but the finaldocker pushto ECR repeatedly fails with a transient network EOF on a single blob HEAD request from the.201self-hosted runner fleet to272493677981.dkr.ecr.us-east-1.amazonaws.com. Two freshworkflow_dispatchruns on main today (runs 28315044412, 28315257828) both pushed several blobs successfully, then died on a blob HEAD with:Because the EOF recurs across different runs/runners, it is not one bad runner. No fresh
omninode-runtimeimage has reached ECR since 2026-05-25 (sha256:de9a7a4c), which blocks the OMN-13418-gated prod re-pin.Change
Wrap the
docker pushinbuild-and-push-runtime.ymlin a bounded retry loop (5 attempts, exponential backoff 10s, 20s, 40s, 80s, cap 120s), mirroring the existing exponential-backoff resilience idiom already used by.github/actions/resolve-ecr-digest.docker pushis resumable -- already-pushed blobs are skipped on retry -- so retrying absorbs the transient blob-HEAD EOF.Additive only. The build, the image tag/digest computation, the lineage-guard, and the Trivy step are all unchanged. Only the network push is retried.
This PR does NOT deploy, restart, retag, or touch any cluster or
.201host config. The prod re-pin follows separately under OMN-13418 gating.dod_evidence
Evidence-Ticket: OMN-13687
Evidence-Source: OCC#3241
Evidence-Class: hotfix
Active-Hotfix-PR: this PR
hotfix-evidence: OCC-3241
backmerge: #2131
Local gate results (2026-06-28):
pre-commit run --files .github/workflows/build-and-push-runtime.yml-- all applicable hooks Passed (yamlfmt, skip-token rejector, root-cleanliness, etc.)yaml.safe_load(...)on the workflow -- YAML validPush image to Amazon ECRstepBackmerge note
The workflow retry loop should also land on
devto prevent re-divergence on the next dev to main promotion; tracked under OMN-13687.backmerge: #2131references the prior emergency-recovery hotfix in the same OMN-13670/OMN-13687 recovery chain (same pattern #2131 used referencing #2128).check-refresh: 2026-06-28T08:16:54Z