Repository navigation
fix(OMN-12433): github.com egress healthcheck for CI runner image - #1792
Conversation
|
Warning Review limit reached
More reviews will be available in 39 minutes and 32 seconds. Learn how PR review limits work. Your organization has run out of usage credits. Purchase more in the billing tab. ⌛ How to resolve this issue?After more reviews become available, a review can be triggered using the We recommend that you space out your commits to avoid hitting the rate limit. 🚦 How do rate limits work?CodeRabbit enforces hourly rate limits for each developer per organization. Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available. Please see our Fair Usage Limits Policy for further information. ℹ️ Review info⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (4)
📝 WalkthroughWalkthroughAdds a repository CodeQL config and inlines CodeQL execution in the security-scan workflow; implements an egress-capable runner healthcheck deployed via Dockerfile/docker-compose; updates the setup-python-uv composite action for authenticated git fetches and higher retries, and adds tests covering these changes. ChangesCodeQL Configuration and Workflow
Runner Health Check Enhancement
UV composite action and Workflow Wiring
Estimated code review effort🎯 4 (Complex) | ⏱️ ~45 minutes Poem
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
|
Codex update: pushed |
|
Updated #1792 for the repeated CodeQL failure class. The rerun failed with GitHub malformed-request annotations on Evidence:
|
|
Follow-up CodeQL hardening: carried over the supported Evidence:
|
|
CodeQL follow-up: GitHub code-scanning upload/processing is still returning malformed Evidence:
|
There was a problem hiding this comment.
Actionable comments posted: 3
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In @.github/workflows/security-scan.yml:
- Around line 34-35: The checkout step currently uses actions/checkout@v6
without disabling credential persistence; update the Checkout repository step
(uses: actions/checkout@v6) to set persist-credentials: false so the
GITHUB_TOKEN is not written into .git/config during the workflow; modify the
step that defines "name: Checkout repository" to add the persist-credentials:
false input under that action.
- Around line 34-35: Replace floating action tags with pinned commit SHAs for
the GitHub Actions used in this workflow: change uses: actions/checkout@v6,
github/codeql-action/init@v4, github/codeql-action/autobuild@v4, and
github/codeql-action/analyze@v4 to the corresponding full commit SHAs (while
keeping the original tag as a trailing comment for readability); update the four
uses entries so each points to its exact SHA instead of the v-tag to prevent
supply-chain drift and match how other workflows (e.g., attest-source-hash.yml)
are pinned.
- Around line 47-52: The workflow step "Perform CodeQL Analysis" currently sets
the CodeQL action (github/codeql-action/analyze@v4) with upload: never and
wait-for-processing: false which prevents any SARIF/results from being sent to
GitHub and suppresses alerts; update this step to use upload: failure-only or
remove/enable the upload option so results are uploaded (or keep upload but make
the corresponding branch protection/check non-required) and remove the redundant
wait-for-processing: false when upload is enabled; ensure you edit the job step
that references github/codeql-action/analyze@v4 and replace upload: never with
the chosen upload mode.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro
Run ID: 440a7c61-7d3b-4f06-bacc-c5923fd43c3d
📒 Files selected for processing (7)
.github/codeql/codeql-config.yml.github/workflows/security-scan.ymldocker/docker-compose.runners.ymldocker/runners/Dockerfiledocker/runners/healthcheck.shtests/ci/test_ci_workflow_resilience.pytests/unit/observability/runner_health/test_runner_fleet_config.py
|
CI triage for OMN-12433 at 2026-05-31T01:17:51Z: inspected current |
|
Triage update for OMN-12433 current failed check:
Local validation on the ticket worktree at
Rerun attempt for job 78685565504 was rejected by GitHub with |
|
Foreground fix for the unresolved CodeQL review thread in Change: replaced the raw Verification on actual PR head
|
|
Foreground #1792 refresh (2026-05-31T17:08Z) Post-fix #1792 still shows stale cancelled Architecture Handshake / Env Parity contexts and a Docker Build Runtime Image failure. Raw Docker job log for Build Runtime Image returned GitHub Actions:
|
|
Foreground controller update: Docker Build run 26715924945 still exposes a stale failed Build Runtime Image context while the workflow remains queued on its summary job. The raw job log endpoint for job 78738271894 currently returns BlobNotFound, so there is no source failure to patch from yet. I attempted a failed-job rerun path again and will keep polling until GitHub releases the queued parent or provides logs. |
|
Foreground action: Docker Build parent 26715924945 is now terminal failed. Build Runtime Image 78738271894 and Build Summary 78741024303 both fail, but both raw log endpoints return GitHub BlobNotFound, so there is no source-actionable Docker assertion available. Reran failed jobs for 26715924945 now that the parent is terminal. CI parent 26715925012 still has queued CI Tests Gate and stale cancelled contexts, so CI remains active-parent-blocked. |
639630e to
93453d5
Compare
Pull request was converted to draft
93453d5 to
0c60019
Compare
|
Worker AA update for merge-sweep backlog. Decision: repair by rebuilding the PR delta, not retire. The old branch was too stale to merge as-is: carrying it wholesale would have reverted newer Head: Current state after push: draft PR, GitHub reports mergeable |
…er fleet (OmniNode-ai#1833) The live .201 runner fleet was hand-scaled to 48 always-on runners but never reconciled back to the repo source of truth. The repo dev compose defined only 20 runner services and config/runner_fleet.yaml said expected_count:14 / burst_count:20. scripts/deploy-runners.sh rsyncs the REPO compose over the host copy then runs `docker compose up -d --build --force-recreate --remove-orphans`. Running deploy today would orphan-remove live runners 21-48 (48->20 org-CI outage). Reconcile the repo to the proven live fleet (repo follows reality): - config/runner_fleet.yaml: expected_count 14->48, burst_count 20->48. The live fleet runs all 48 as steady-state (no burst tier), so burst_count==expected. - docker/docker-compose.runners.yml: 20->48 steady services. Each service block is byte-identical to the live fleet except for the per-runner OMN-12433 egress healthcheck.sh mount, which the repo intentionally adds (the live compose still carries the older pgrep-only healthcheck and lacks the mount). Header comments updated to reflect the 48-runner resource reality (limits, not reservations). - scripts/deploy-runners.sh: add docker/runners/healthcheck.sh to SYNC_PATHS and the rsync invocation. The compose has bind-mounted ./runners/healthcheck.sh since OMN-12433 but deploy never shipped the artifact, so the mount would resolve to an empty host path. This fixes that latent gap. healthcheck.sh already exists in the repo (added by OMN-12433 OmniNode-ai#1792) and is COPYed into the image by the Dockerfile; no new artifact is introduced. Image ref stays omninode-runner:latest — the OMN-12567 versioned-image bump is a separate concern. dod_evidence: - deploy-runners.sh --dry-run reports `Runner count: 48`, confirming the config drives the deploy to target 48 runners. - No-orphan proof: the reconciled compose defines exactly the 48 service names that match the 48 live running containers (zero orphan-removal). The old 20-service compose would have orphan-removed runners 21-48. - New regression tests assert expected_count==48, 48 steady services each with the healthcheck mount, and that deploy ships healthcheck.sh. They fail on the old 20/14 state (verified TDD). Evidence-Source: <pending-occ> Evidence-Ticket: OMN-12582 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
Summary
Durable fix for the 2026-05-29 CI wedge. The runner Docker healthcheck was
pgrep -f Runner.Listeneronly — it passed even when a runner had silently lost its connection to github.com. That let ~9 of 20 runners sit "Up (healthy)" in Docker while OFFLINE in the GitHub pool, starving the merge queue and blocking #1789/#1781/#1782.docker/runners/healthcheck.sh: requires BOTH theRunner.Listenerprocess AND a short-timeout (--max-time 8) github.com reachability probe. A runner that loses egress now goes unhealthy and is removed from rotation instead of accepting jobs it will fail.Dockerfile(baked) and mounted into every runner service indocker-compose.runners.yml(so already-deployed runners pick it up on recreate, no rebuild required). Anchor + 20 services = 21 mounts.tests/unit/observability/runner_health/test_runner_fleet_config.pyassert the script probes github.com and every runner service uses the egress healthcheck (not bare pgrep).Proof
bash -n+shellcheckclean on healthcheck.sh.tests/unit/observability/runner_health/test_runner_fleet_config.py: 6 passed (incl. 2 new). Broadertests/unit/observability/+ compose tests: 151 passed.healthcheck.test: [CMD-SHELL, /usr/local/bin/healthcheck.sh]and mount the script.Context: this is the durable companion to OMN-12432 (the uv git-auth fix). Both address the same .201 runner-egress fault — OMN-12432 stops it failing git fetches, OMN-12433 stops a degraded runner from silently staying in the pool.
Test plan
Evidence-Source: OCC#1888
Evidence-Ticket: OMN-12433
OMN-12433
Summary by CodeRabbit
New Features
Improvements
Tests