Skip to content

fix(OMN-15249): make the Integration Silent-Skip Guard die AT the Postgres image pull, not on a downstream exit 127 - #2501

Merged
jonahgabriel merged 2 commits into
devfrom
jonah/omn-15249-guard-pull-fatality
Jul 27, 2026
Merged

jonahgabriel merged 2 commits into
devfrom
jonah/omn-15249-guard-pull-fatality

Conversation

@jonahgabriel

@jonahgabriel jonahgabriel commented Jul 27, 2026 •

Copy link
Copy Markdown
Collaborator

Evidence-Source: OCC#5149
Evidence-Ticket: OMN-15249

OMN-15249 — the Integration Silent-Skip Guard must die AT the Postgres image pull

Ticket: OMN-15249 (child of the OMN-14172 guard family). Source-only: workflow YAML, its config comment, and one new test. No lane touched, no runtime mutated.

The live defect (omnibase_infra#2492, job 89995662337)

The integration-guard job provisioned Postgres through a GitHub-managed services: block. That image pull happens inside GitHub's own Initialize containers step, which this repo cannot name, cannot bound, and — critically — which steps carrying if: always() still run past. Verbatim from the job log:

##[command]/usr/bin/docker pull postgres:16-alpine
Error response from daemon: Get "https://registry-1.docker.io/v2/": context deadline exceeded (Client.Timeout exceeded while awaiting headers)
##[warning]Docker pull failed with exit code 1, back off 8.077 seconds before retry.
... (two more attempts) ...
##[error]Docker pull failed with exit code 1

Step conclusions for that job, read live from the API:

# Step Conclusion
2 Initialize containers failure
3–8 checkout → setup-python-uv → resolve host → migrations → export env → curated proofs skipped
9 Enforce no missing-service silent-skips failure
10 Upload integration-guard results success (warned No files were found)

Step 9 carried a bare if: always(), so it fired past the failed container init with no toolchain installed:

/home/runner/work/_temp/....sh: line 1: uv: command not found
##[error]Process completed with exit code 127.

That exit 127 is the last error in the log and therefore what triage reads — a phantom missing-binary problem several steps removed from a registry timeout. It also invites exactly the "transient" mislabeling the operating rules reject.

One correction to the ticket's framing, stated because it changes the fix: the pull failure was already fatal and already bounded — GitHub retries 3 times with backoff and the final attempt is ##[error], not a warning; the two warnings are per-attempt. The real defects are (1) the guard emitting a verdict for a run whose container never materialized, and (2) the misattributed terminal signal. Note that had uv survived, check_integration_skips.py would have exited 2 on the missing JUnit report — still a verdict about silent skips from a run that never provisioned Postgres. The guard was committing its own false-green failure class against itself.

Fix

Provisioning moves out of services: and into steps this repo owns, which is what makes every DoD item mechanically enforceable rather than dependent on GitHub-internal behavior:

  • pull_postgres (first step by design). Bounded retry — GUARD_PG_PULL_ATTEMPTS=3, each attempt wrapped in timeout ${GUARD_PG_PULL_TIMEOUT_SECONDS} (120s), linear backoff — then fails closed with an ::error title=Postgres image pull failed (OMN-15249):: annotation naming the image, the registry, and the per-attempt timeout, and saying in words that this is a registry/network failure, not a missing binary and not a test failure. It runs before checkout and before toolchain setup: nothing after a failed pull can be trusted, so nothing after it is even attempted.
  • start_postgres. Explicit docker run with the same pg_isready health probe and the same ephemeral published port the services: block used, bounded health wait, fail-closed with a named annotation (plus docker logs) if the container never reports healthy, and docker port → $GITHUB_OUTPUT for the downstream host/port resolution.
  • Verdict gating. enforce_no_silent_skips and upload_guard_results move from bare if: always() to always() && steps.run_curated_proofs.conclusion != 'skipped'. The verdict fires when the proofs ran (success or failure) and never when they were not reached. This deliberately preserves the guard's whole purpose: on a genuine integration failure it still runs.
  • stop_postgres. Unconditional docker rm --force ... || true, so a container started before a later step failed is never leaked onto a self-hosted runner and teardown never invents a second, misattributed failure.

scripts/ci/integration_skip_guard.yaml provisioned_by / header comments updated to match (the file is the gate's SYNC anchor).

Proof

tests/ci/test_integration_guard_pull_fatality.py — 8/8 RED at origin/dev, 8/8 GREEN on this branch. Two of the three sections are behavioral, not source-text greps:

Test Proves
test_simulated_registry_timeout_fails_the_pull_step_with_a_named_error Extracts the pull step's real run: body and executes it as a bash subprocess against a stubbed docker on PATH that always fails. Asserts non-zero exit, exactly GUARD_PG_PULL_ATTEMPTS invocations (recorded by the stub, so retry is proven bounded and fail-closed on exhaustion), and that image + registry + timeout all appear in the ::error annotation
test_pull_step_succeeds_and_stops_retrying_when_the_registry_answers Same script, succeeding stub → exit 0 after one attempt. A fail-closed loop that can only ever be red is a disabled check
test_no_toolchain_step_is_reachable_after_a_failed_pull Replays the step graph through a GitHub-if-semantics simulator with the pull marked failed: the verdict and upload steps resolve to skipped, cleanup still runs, and no executed step's shell body contains uv/pytest/python — i.e. the exit-127 path is unreachable, asserted over the graph rather than assumed
test_verdict_still_runs_when_the_curated_proofs_actually_fail The same simulator with the proofs failing: the verdict does run. Guards against trading one false-green for another
test_verdict_runs_on_the_fully_healthy_path Verdict runs, all four owned steps execute
test_guard_does_not_delegate_the_pull_to_a_services_block No services: key — the structural precondition without which none of the above is enforceable
test_pull_is_the_first_step_and_cannot_be_softened Pull is step 0; neither owned step sets continue-on-error (the defect class is literally "container failure downgraded")
test_pull_retry_budget_is_bounded_and_declared Attempts and per-attempt timeout are bounded and declared as job env

The simulator fails closed: any if: form it does not model raises pytest.fail rather than being guessed at, so a future edit cannot silently re-open the warn-and-continue path.

Live end-to-end on real Docker (Linux, .201, ephemeral standalone container omn15249-e2e-proof, no compose project and no lane touched, removed afterward — verified 0 containers remaining). Neither host available for gates has Docker, so this is where the real-daemon behavior was proven:

  • Happy path: pull_postgres → attempt 1/3 success; start_postgres → healthy, published port 33844, port=33844 written to $GITHUB_OUTPUT; TCP connect to 127.0.0.1:33844 OK and psql -c 'select 1' → 1 (so --publish 5432 and the existing host-resolution step still line up); stop_postgres → container gone.
  • Unresolvable registry: 3 bounded attempts, exit 1, ::error annotation naming image/registry/timeout.
  • Blackholed registry (10.255.255.1:5000, the ticket's actual failure mode — a hang, not a DNS error): 3 attempts × 5s timeout = 15s wall clock, exit 1, named annotation. The timeout wrapper is what bounds it.

Gates (rule 11a — run on .200, ssh stickybeatz-studio)

Patch-transfer discipline: edited locally, git diff --cached → scp → git apply --index on .200, then per-file sha256 verified identical on both hosts before any gate ran; ruff format was applied on .200 and the formatted file copied back so the hosts stayed identical. Commit and push executed on .200.

  • ruff check src/ tests/ — All checks passed
  • ruff format --check — clean (one reformat applied and synced back)
  • mypy tests/ci/test_integration_guard_pull_fatality.py — Success, no issues
  • Governed selector (scripts/ci/detect_test_paths.py --feature-flag on) → {"selected_paths":["tests/ci/"],"is_full_suite":false} → uv run pytest tests/ci/ = 980 passed, 1 skipped (the skip is the pre-existing omninode_infra sibling-not-found guard, unrelated)
  • pre-commit run --files <all 3 changed> — 45 passed, 0 failed, no file rewritten by a hook. Includes Integration silent-skip guard selftest (OMN-14172), which passes against the edited config

Deviation to state: the .201 Docker end-to-end could not run on .200 or this Mac — neither has a Docker daemon (docker: command not found on both). It used a standalone container with a unique name, touched no compose project, mutated no lane, and was torn down and verified gone.

Scope note

migration-integration still uses a services: block. It is deliberately untouched: it has no if: always() step, so a pull failure there terminates cleanly at Initialize containers with no misattributed downstream signal. Widening the change would add blast radius without closing a live defect.

Safety

No merge performed by an agent. No lane mutation, no container recreate on any compose project, no prod surface touched.

…tgres image pull, not on a downstream exit 127

The `integration-guard` job provisioned Postgres via a GitHub-managed
`services:` block. On #2492 that image pull timed out against
registry-1.docker.io inside GitHub's own "Initialize containers" step; every
normal step was skipped, but the verdict step carried a bare `if: always()`,
ran with no toolchain installed, and terminated the job on
`uv: command not found` / exit 127 — several steps removed from the registry
timeout that actually caused it.

- Own the pull: explicit first step with bounded retry (3 attempts, 120s
  per-attempt `timeout`) that fails closed with an `::error::` naming the
  image, the registry, and the timeout.
- Own the container: explicit `docker run` with the same health probe and an
  ephemeral published port, fail-closed on an unhealthy container, plus an
  unconditional `docker rm --force` teardown.
- Gate the verdict on `steps.run_curated_proofs.conclusion != 'skipped'` instead
  of bare `always()`, so the guard cannot report on a run whose Postgres never
  materialized, while still firing on genuine test failures.

Proof: tests/ci/test_integration_guard_pull_fatality.py — 8/8 RED at
origin/dev, 8/8 GREEN here. Executes the workflow's real pull `run:` body as
a bash subprocess against a stubbed docker, and replays the step graph through
a GitHub-`if`-semantics simulator that fails closed on unmodelled conditions.
Live end-to-end on real docker (.201, ephemeral standalone container, no lane
touched): happy path pull/start/port/teardown green; blackholed registry
exhausted 3 bounded attempts in 15s and exited 1 with the named annotation.
@coderabbitai

coderabbitai Bot commented Jul 27, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

You’ve reached a temporary PR review limit under our Fair Usage Limits Policy.

Your recent review volume is higher than typical usage, so adaptive limits are currently applied.

Next review available in: 39 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: dd2fd943-bf35-403a-805f-d9f9ec91efc8

📥 Commits

Reviewing files that changed from the base of the PR and between b0e8ca6 and fde67b7.

📒 Files selected for processing (3)
  • .github/workflows/ci.yml
  • scripts/ci/integration_skip_guard.yaml
  • tests/ci/test_integration_guard_pull_fatality.py
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch jonah/omn-15249-guard-pull-fatality

Comment @coderabbitai help to get the list of available commands.

jonahgabriel added a commit to OmniNode-ai/onex_change_control that referenced this pull request Jul 27, 2026
#5149)

* evidence: OCC companion pass 1 for OmniNode-ai/omnibase_infra#2501

* evidence: OCC companion self-bind for #5149

---------

Co-authored-by: node-occ-companion-effect <occ-companion-effect@omninode.ai>
@github-actions

Copy link
Copy Markdown
Contributor

⚠️ Hostile Reviewer — DEGRADED (informational)

Blocking findings (critical): 0
Total findings: 0
Models succeeded: none

Note: All reviewer models failed or were unavailable. Degraded results are informational during the pilot phase (OMN-8468/OMN-8524) and do not block merge. Error: all review endpoints [192.168.86.201:8000 192.168.86.201:8001 ] unreachable — preflight short-circuit (no models available)


Gate semantics (pilot phase)

Verdict Meaning Blocks merge?
passed No critical findings No
blocked CRITICAL findings found Yes
degraded All models unavailable (infra) No (pilot)

Powered by omniintelligence.review_pairing.cli_review — node-based adversarial review via HandlerLlmCliSubprocess (OMN-8468/OMN-8524)

@jonahgabriel
jonahgabriel merged commit 6bc1e7f into dev Jul 27, 2026
94 checks passed
@jonahgabriel
jonahgabriel deleted the jonah/omn-15249-guard-pull-fatality branch July 27, 2026 17:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant