Repository navigation
fix(OMN-12432): authenticate uv git+https fetches in CI (Empty reply from server) - #1789
Conversation
Self-hosted runners ran repeated unauthenticated, uncached git-by-SHA fetches of git+https deps (omnibase-spi, omnibase-core, onex_change_control) on every uv sync. Anonymous github.com requests are rate-limited (60/hr); parallel --no-cache syncs from one runner egress IP tripped "Empty reply from server", persistently failing Type Safety / type-union / deploy-gate and blocking PRs #1781/#1782. Authenticate uv's internal git fetches via a process-scoped insteadOf rewrite (GIT_CONFIG_* env vars — never persisted to disk on the runner) using github.token (override: CROSS_REPO_PAT): - setup-python-uv composite action: new github-token input (defaults to github.token) wires the rewrite before uv sync — fixes every caller. - omni-standards-compliance type-safety / type-union-check: route through the composite action instead of inlining unauthenticated uv sync, so they also gain retries + HTTP/1.1. - handler-contract-compliance: authenticate its pinned onex_change_control git+https install the same way. Proven locally: the exact uv by-SHA fetch of c01f70cd (v0.22.0) and a full `uv sync --no-cache --all-extras` both succeed with the env-var rewrite. Regression guards added in tests/ci/test_ci_workflow_resilience.py.
|
Warning Review limit reached
More reviews will be available in 8 minutes and 45 seconds. Learn how PR review limits work. Your organization has run out of usage credits. Purchase more in the billing tab. ⌛ How to resolve this issue?After more reviews become available, a review can be triggered using the We recommend that you space out your commits to avoid hitting the rate limit. 🚦 How do rate limits work?CodeRabbit enforces hourly rate limits for each developer per organization. Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available. Please see our Fair Usage Limits Policy for further information. ℹ️ Review info⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (4)
📝 WalkthroughWalkthroughAdds a GitHub token input and in-process git auth rewrite to the setup-python-uv composite action, integrates that action into Omni Standards jobs, introduces a repo-local CodeQL config and inlined CodeQL workflow steps, and adds tests covering tokenized git fetches and CodeQL config usage. ChangesCI Authentication and CodeQL Infrastructure
Estimated code review effort🎯 4 (Complex) | ⏱️ ~45 minutes Poem
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
|
Codex update: pushed |
|
Update pushed for the repeated CodeQL failure on current head. Evidence:
|
There was a problem hiding this comment.
🧹 Nitpick comments (1)
.github/workflows/security-scan.yml (1)
34-35: ⚡ Quick winConsider
persist-credentials: falseon checkout.CodeQL
init/analyzeuse their own job token (viasecurity-events: write), so the checkout-persisted GITHUB_TOKEN in.git/configisn't needed by the analysis steps. Disabling it reduces credential exposure duringautobuildon this security workflow.🛡️ Proposed hardening
- name: Checkout repository uses: actions/checkout@v6 + with: + persist-credentials: false🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In @.github/workflows/security-scan.yml around lines 34 - 35, Update the GitHub Actions checkout step that uses actions/checkout@v6 to disable persisting the workflow GITHUB_TOKEN by adding persist-credentials: false; specifically modify the checkout step (the block referencing uses: actions/checkout@v6) to include persist-credentials: false so the repo checkout does not write the default token into .git/config during this CodeQL security workflow.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Nitpick comments:
In @.github/workflows/security-scan.yml:
- Around line 34-35: Update the GitHub Actions checkout step that uses
actions/checkout@v6 to disable persisting the workflow GITHUB_TOKEN by adding
persist-credentials: false; specifically modify the checkout step (the block
referencing uses: actions/checkout@v6) to include persist-credentials: false so
the repo checkout does not write the default token into .git/config during this
CodeQL security workflow.
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro
Run ID: b4b2660c-0f45-401d-96e5-77dded6f7455
📒 Files selected for processing (5)
.github/actions/setup-python-uv/action.yml.github/codeql/codeql-config.yml.github/workflows/omni-standards-compliance.yml.github/workflows/security-scan.ymltests/ci/test_ci_workflow_resilience.py
|
Follow-up on #1789 CodeQL: the repo-local config was loaded, but the job still failed during GitHub code-scanning processing after SARIF upload. I updated the analyze step to submit SARIF without waiting for server-side processing, using the supported Evidence:
|
|
CodeQL follow-up: GitHub code-scanning upload/processing is still returning malformed Evidence:
|
Reformat multiline assert messages flagged by the Lint job's `ruff format --check` step. Format-only; no behavior change. OMN-12432
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In @.github/workflows/security-scan.yml:
- Around line 51-52: The workflow currently disables SARIF upload with the
"upload: never" setting (and also sets "wait-for-processing: false"); instead,
revert "upload: never" to allow CodeQL SARIF uploads and investigate why
server-side processing failed by confirming the CodeQL configuration file that
defines "paths"/"paths-ignore", the scan "languages"/"mode", and whether the
autobuild/build step is causing unintended inclusions; try alternatives
first—ensure the github/codeql-action is up-to-date, apply exclusions in the
build/autobuild or post-process SARIF if needed—and add a short tracking issue
to revert any temporary "upload: never" change and escalate to GitHub support if
exclusions cannot prevent the processing failure.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro
Run ID: 1f0307e8-87d0-4753-9954-c8d647967711
📒 Files selected for processing (3)
.github/actions/setup-python-uv/action.yml.github/workflows/security-scan.ymltests/ci/test_ci_workflow_resilience.py
🚧 Files skipped from review as they are similar to previous changes (1)
- tests/ci/test_ci_workflow_resilience.py
…egration Test Coverage Previous run 26672891204 failed only on 'Failed to set up job. The runner has received a shutdown signal.' before any test ran. Empty commit to re-fire the workflow on a healthy runner. gh run rerun is unavailable (PR-branch workflow-file diff). No code change. OMN-12432
…sted) Prior CI workflow run 26674832759 hung at 'queued' for ~2h and never emitted the required 'CI Summary' status, so the merge queue rejected enqueue with 'Required status check CI Summary is expected'. Empty commit to get one clean CI run that posts the context. No code change. OMN-12432
Prior CI run's Lint job was cancelled (empty step log, no ruff output) — runner cancellation, not a real format violation (ruff format --check is clean locally). Empty commit to get a clean CI run so CI Summary posts. OMN-12432
|
Manual follow-up for OMN-12432 / head
No code change or push from this pass. |
|
Foreground stale-context refresh (2026-05-31T17:01Z) Manual PR-list review found #1789 still blocked by stale cancelled fanout and aggregate CI Summary contexts; no source assertion is visible in the current rollup. Rerunning terminal workflow parents now:
Keeping this in the foreground queue with the CI-substrate set. |
|
Foreground triage: current cancelled contexts are stale inside active/replacement parents. Deploy Gate run 26692494315 is terminal cancelled, but replacement deploy run 26694741061 is already queued on the PR. CI run 26694741054 still has CI Summary queued after rerun, so its cancelled Lint/Version Pin contexts are not safe to rerun independently yet. No source failure indicated; continuing to poll active replacements. |
|
Foreground tick update for #1789:
|
|
Foreground tick update for #1789:
|
Pull request was closed
Rebased onto dev (picks up the git-auth fix #1789, OMN-12432). Enables the persistent uv cache via the setup-python-uv composite action across CI jobs, while keeping cache-enabled: false on the cross-repo git+https fetch jobs that dev deliberately protected (topic-enum-drift, type-safety, type-union-check) so a stale cache restore cannot reintroduce the anonymous-rate-limit flake. Updates test_required_ci_jobs_use_uv_cache_by_default to honor those exemptions and fixes a pre-existing SPDX header year.
Durable fix for the 2026-05-29 CI wedge. The runner Docker healthcheck was pgrep -f Runner.Listener only - it passed even when a runner had silently lost its connection to github.com. That let ~9 of 20 runners sit Up (healthy) in Docker while OFFLINE in the GitHub pool, starving the merge queue. - New docker/runners/healthcheck.sh: requires BOTH the Runner.Listener process AND a short-timeout (--max-time 8) github.com reachability probe. A runner that loses egress now goes unhealthy and is removed from rotation. - Wired into the runner Dockerfile (baked) and mounted into every runner service in docker-compose.runners.yml so already-deployed runners pick it up on recreate. - Dockerfile.runtime plugin external-dep install now retries on transient fetch failures (UV_HTTP_TIMEOUT + retry loop) so a degraded egress does not abort the runner image build. - omni-standards-compliance.yml OCC git+https install hardened with UV_HTTP_TIMEOUT, HTTP/1.1 pin, and a retry loop on top of the OMN-12432 authenticated fetch already on dev. - Regression tests assert the script probes github.com, every runner service uses the egress healthcheck, and the OCC fetch retries. Rebuilt on top of origin/dev (which already carries the OMN-12432 auth fixes from #1789) to scope this PR to the egress-healthcheck delta and clear the merge conflict. Evidence-Source: OCC#1888 Evidence-Ticket: OMN-12433 OMN-12433
Summary
CI jobs that run
uv syncfailed persistently while fetching git-pinned deps (omnibase-spi,omnibase-core,onex_change_control) from GitHub on the self-hosted runners, blocking #1781/#1782 and causing all-day "flakiness" across CodeQL, Type Safety, Migration tests, and deploy-gate (all runuv syncfirst).Root cause (confirmed): the runners performed repeated unauthenticated, uncached git-by-SHA fetches on every job.
--no-cacheforces a fresh fetch each run; anonymous github.com requests are rate-limited (60/hr) and, under many parallel syncs from one egress IP, returnEmpty reply from server. The pin is valid —git ls-remoteconfirmsc01f70cd=refs/tags/v0.22.0^{}. Not a bad pin, not a flake.Fix: authenticate uv's internal
git fetchusing the org-canonicalx-access-tokentoken (secrets.CROSS_REPO_PAT || github.token), applied as a process-scopedinsteadOfrewrite viaGIT_CONFIG_*env vars — so the token is never written to a persistent gitconfig on the self-hosted runner. Authenticated requests get the 5000/hr limit.setup-python-uvcomposite action: newgithub-tokeninput (defaults togithub.token) configures the rewrite beforeuv sync. Fixes every caller (ci.yml, env-parity, check-sibling-compat) with zero per-caller edits.omni-standards-compliance.ymltype-safety/type-union-check: routed through the composite action instead of inlining unauthenticateduv sync --no-cache --all-extras— these are the checks directly blocking fix(OMN-12416): type-scope multi-handler dispatch + per-handler result application #1781/fix(OMN-12421): advance stale OMNIBASE_COMPAT_REF pin so clean redeploy works #1782. They now also gain the retry loop + HTTP/1.1.handler-contract-compliance: its pinnedonex_change_controlgit+httpsinstall now authenticates the same way.Proof (run locally against the exact failing SHA)
git fetch --force --update-head-ok origin '+c01f70cd...:refs/commit/c01f70cd...'→ succeeds, resolves the commit.uv sync --no-cache --all-extraswith the env-var rewrite → built + installedomnibase-spi==0.22.0 (@c01f70cd),omnibase-core,onex-change-control. NoEmpty reply from server.tests/ci/test_ci_workflow_resilience.py: 9 passed (2 new regression guards for the auth wiring); unit CI suite 124 passed.Closes OMN-12432.
Test plan
uv syncfetches git deps withoutEmpty reply from serverOMN-12432
Evidence-Source: 0a670cfbf8bb416ae12d005dd5e5a5c00e4ae336
Evidence-Ticket: OMN-12432
Summary by CodeRabbit
Chores
Tests