Repository navigation
feat(OMN-14873): health-gated, rollback-on-failure stability-lane refresh mechanism - #2370
Conversation
…resh mechanism Wraps the proven manual .201 stability-test refresh recipe (workspace-mode build scoped to the 4 known-good core services + targeted recreate) into scripts/runtime_build/refresh_stability_lane.sh, with pre-state capture, a forward-progress ancestry assertion, a health-gate (verify_stability_refresh.py: digest changed, manifest contract-count floor, /health, rpk cluster health, declared consumer groups Stable, image-revision readback), automatic rollback-and-re-verify on failure, and a durable JSON receipt under ~/.omnibase/state/stability_lane_refresh/. deploy-runtime.sh gets one small additive change (RUNTIME_BUILD_SERVICES_OVERRIDE, defaulting to current behavior) so the build/restart can be scoped to a service subset -- this routes the 4 known-broken release-only services (BUILD_SOURCE selector-mismatch, OMN-14262 residual) around as a controlled decision instead of relying on a partial-build-failure side effect. Cadence (cron/GHA) is deliberately NOT wired -- design only, per feedback_no_landing_automation_before_process_measured, gated on >=5 consecutive clean manual runs. Child of OMN-14263.
|
Warning Review limit reachedYou’ve reached a temporary PR review limit under our Fair Usage Limits Policy. Next review available in: 58 minutes Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (7)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
…rom set -e Canary run surfaced two real bugs in refresh_stability_lane.sh: 1. The rollback's direct `docker compose up --force-recreate` call did not source docker/runtime-policy.env + ~/.omnibase/.env the way deploy-runtime.sh's own compose calls do, so it failed at config interpolation time (e.g. BIFROST_VERIFY_ENDPOINTS unset). 2. That failure was unguarded under `set -e`, so the whole script aborted with NO receipt written at all -- worse than the intended FAILED/ FAILED_ROLLED_BACK receipt outcomes. Fixes: source the same env files at script top (preserving operator OMNI_HOME like deploy-runtime.sh's header does); guard the rollback recreate call and both health-gate JSON reads so a crash anywhere in this path still reaches the final receipt-writing step with an honest FAILED/INFRA_ERROR record instead of a bare script crash.
…egex Real canary findings against the live .201 stability-test lane: 1. `rpk group describe` has NO `-f json` output mode on this rpk version (rejects -f as unknown) -- parse the plain-text STATE line instead. 2. Declared consumer-group names in consumer_groups_stability.yaml were bare prefixes; the real rpk group name carries a `.__i.<instance>.__t.<topic>` per-topic suffix. The bare prefix resolves to a distinct, nonexistent group (STATE=Dead) -- fixed to use the full names copied from `rpk group list`. 3. `rpk cluster health`'s "Healthy:" line padding varies by rpk version (observed much wider than the hardcoded exact-spacing match) -- switched to a regex match. 4. "Empty" (group registered, zero current members) is the NORMAL idle state for a demand-driven consumer and was producing false-negative rollbacks even though the refresh itself was healthy -- only "Dead"/absent (the actual "never registered a group" wiring-death signal) now fails the check. Added a bounded retry (mirrors deploy-runtime.sh's assert_broker_reachable retry shape) since a freshly-recreated container needs a few seconds to reconnect and rejoin its group. 19 -> 23 unit tests (added Empty-is-healthy, Dead-is-not, retry-recovers, retry-exhausts-and-reports coverage).
Release-train canary (OMN-14889) found a real .201 failure: this script's
step 3 ("Refresh the omnibase_infra ambient clone itself to --ref") ran
git checkout dev
git reset --hard "${REF}"
which hard-fails whenever another worktree on the same host (e.g.
deploy-agent's runtime-sync-worktrees/OMN-12618) already has the local
`dev` branch checked out -- git refuses to check the same branch out
into two worktrees at once. This made the deploy step environment-
fragile: it only passed once because the collision happened not to be
present that run.
Fix: resolve --ref to a commit SHA and check it out DETACHED
(`git checkout --force --detach <sha>`). A detached HEAD carries no
branch identity, so it can never collide with a sibling worktree
regardless of what branch that worktree holds. --force preserves the
previous checkout+reset --hard pair's force-discard-local-changes
behavior. The forward-progress ancestry assertion and everything else
is untouched.
Adds tests/unit/scripts/test_refresh_stability_lane_checkout.py:
real-git RED/GREEN reproduction of the exact collision (branch
checkout fails while `dev` is busy in a sibling worktree; detached-
by-SHA checkout succeeds under the identical condition, and leaves the
sibling worktree's own `dev` checkout completely undisturbed), plus a
static regression guard on the shipped script so the collision-prone
bare-branch-checkout pattern cannot silently come back.
Ticket: OMN-14889 (child of OMN-13674)
dod_evidence: 3/3 new tests pass (RED confirmed against the pre-fix
script content read via `git show`); full runtime_build + scripts unit
suite (799 tests) green, no regressions; shellcheck clean; pre-commit
--files clean.
…ty lab lanes (re-cut onto dev) (#2388) * feat(OMN-14889): git-tag-driven release-train trigger for dev+stability lab lanes Re-cut of closed PR #2377 (commit c2ec0d2) onto current dev after #2370 (OMN-14873) merged as 54938fd. File-scoped restore of the 6 paths only; dev's newer refresh_stability_lane.sh (detached-checkout fix) is left untouched. - .github/workflows/release-train-lab.yml (new) - docker/docker-compose.runners.yml (+74, additive) - docs/runbooks/release-train-lab.md (new) - scripts/runtime_build/cut_release_train_tag.sh (new) - scripts/runtime_build/refresh_dev_lane.sh (new) - scripts/runtime_build/verify_dev_refresh.py (new) * docs(OMN-14889): correct false "runner unprovisioned" safety claims The workflow header, runbook, and runners compose block all asserted the omnibase-deploy runner was unprovisioned and that the deploy job "will queue but never pick up." That is FALSE and safety-relevant: a reviewer would conclude a lab/** tag push is inert when it in fact triggers a real lane refresh in execute mode. Verified live 2026-07-21: gh api repos/OmniNode-ai/omnibase_infra/actions/runners -> omninode-deploy-runner status=online, labels [self-hosted, Linux, X64, omnibase-deploy] That is the exact label set both jobs target, and registration is at the REPOSITORY level (not the org level the runbook recipe assumed). Corrected all three sites to state: the runner is online; a lab/dev/** or lab/stability/** tag push DOES refresh the lane in execute mode (intended, since dev+stability are pre-authorized, but not a no-op); and the remaining blocker is host-side git access to the ambient $OMNI_HOME clone, not registration. Five runs on 2026-07-21T03:59-04:19Z all failed in the "Refresh stability lane" step, in three distinct forms as intermediate fixes were attempted: dubious ownership (128), .git/FETCH_HEAD Permission denied (255), and 'dev' already checked out in a sibling worktree (128). All died before any container action -- zero docker compose / up -d / --force-recreate lines in any of the five logs -- so no lane was mutated. Prod boundary statements preserved and reinforced: lab/** and v* remain disjoint namespaces and prod promotion still routes through node_redeploy_orchestrator's grant gate. No functional change; comments and docs only. Closes OMN-14889 * fix(OMN-14889): harden dev release-train refresh --------- Co-authored-by: Jonah Gray <jonah.g.gray@gmail.com>
…ator (count roots, not checks) (#2391) Plan section 3.C5 (docs/plans/2026-07-21-ci-capacity-recovery-plan.md); root-cause class CI-01 from the 2026-07-16 CI remediation plan - designed, not shipped. On omnibase_infra#2370 a single defect (missing Evidence-Source -> occ-preflight eligibility red) amplifies into 36 red check-runs (35 occ-preflight/eligibility, one per calling workflow, + 1 CI Summary) plus 74 skipped needs:occ-preflight dependents - ~37x amplification. Because each occ red is its own workflow run, each is independently rerunnable, which provokes the rerun reflex (a wall of red invites broad reruns, each re-spending the 40-way matrix). Ships scripts/ci/ci_cascade_graph.py, which consumes the head SHA's full cross-workflow check-run list (GET /commits/{sha}/check-runs) and collapses it into: - exactly ONE typed root cause - root election reuses the deterministic, unit-tested product_reason_graph.build_reason_graph classifier, preserving the OCC-independence property (a real product defect roots as PRODUCT_FAILED, not EVIDENCE_MISSING); - every other red/skipped check marked BLOCKED_UPSTREAM, content-addressed to the single root_receipt_id; and - EXACTLY one rerunnable unit (the root) - every dependent carries rerunnable=False, the anti-cosmetic guarantee. Wired as an additive, report-only render in the CI Summary job: if always() + continue-on-error true, writes only to GITHUB_STEP_SUMMARY, ends every command with || true. It NEVER changes the verdict - the poll step remains the sole pass/fail authority. Gate not weakened: a seeded occ-preflight red still yields ci_summary_gate exit 1 (CI Summary = FAILURE). Proof: tests/ci/test_ci_cascade_graph_omn14909.py replays #2370's real captured blocked head (tests/ci/fixtures/omn14909_2370_cascade_checkruns.json) - 1 root (EVIDENCE_MISSING), 35 red + 74 skipped dependents all BLOCKED_UPSTREAM, 1 rerunnable unit, no dependent independently rerunnable. Plus OCC-independence, determinism, green-head base case, and report-only-exit-0 CLI tests. dod_evidence: - uv run pytest tests/ci/test_ci_cascade_graph_omn14909.py -> 7 passed - regression test_product_reason_graph_omn14707.py + test_ci_summary_gate.py -> 58 passed - uv run mypy scripts/ci/ci_cascade_graph.py -> clean; ruff clean; pre-commit 46 passed / 0 failed - gate-not-weakened: seeded occ-preflight-red jobs.json -> ci_summary_gate.py exit 1 (FAILURE) - actionlint: 0 new findings from the added CI Summary step Report-only; no gate weakened; no branch-protection writes; single-repo (fan-out to core/omnimarket/omniclaude/onex_change_control is a follow-up). Closes OMN-14909 Co-authored-by: Jonah Gray <jonah.g.gray@gmail.com>
Summary
Implements + canaries the design in child ticket OMN-14873 (child of OMN-14263: "stability-test proof lane has no refresh cadence").
scripts/runtime_build/refresh_stability_lane.sh— new. Health-gated, rollback-on-failure refresh of the.201stability-test lane. Hardcoded to that lane only (no--laneflag, mirrorscut-lab-ref.sh's prod/judge refusal). Captures pre-state (image IDs + apreflight-<UTC>rollback tag) before any mutation, refreshes theomnibase_infraambient clone + the 4DEPLOY_REF-pinned siblings to--ref(defaultorigin/dev), builds+restarts SCOPED to the 4 known-good core services (omninode-runtime,runtime-effects,runtime-worker,projection-api), asserts forward progress (git merge-base --is-ancestor) for every tracked repo, runs the health-gate, and on failure rolls back + re-verifies. Emits a durable JSON receipt under~/.omnibase/state/stability_lane_refresh/.scripts/runtime_build/verify_stability_refresh.py— new. The health-gate: digest-changed, manifest contract-count floor,/health,rpk cluster health, declared consumer groups Stable (consumer_groups_stability.yaml), and image-revision readback.--no-require-digest-changesupports the post-rollback re-verification pass (where digest is deliberately unchanged).scripts/runtime_build/consumer_groups_stability.yaml— new. Declared consumer-group health-check list (contract-as-data), spanning registration + the delegation pipeline +node_ticket_classify_compute(the def-B canary this refresh also proves live).scripts/runtime_build/tests/test_verify_stability_refresh.py— new. 19 unit tests (mocked docker/HTTP, no live lane needed): PASS/FAIL boundary at exactlymin_contracts, digest-unchanged, revision-mismatch (exists-but-wrong, not silent pass), and the rollback-re-verification receipt logic including the "rollback ALSO fails" branch stayingFAILED(never masked as success).docs/runbooks/stability-lane-refresh.md— new canonical runbook.scripts/deploy-runtime.sh— one small additive change:RUNTIME_BUILD_SERVICES_OVERRIDEenv var, defaulting to currentRUNTIME_SERVICESbehavior (byte-for-byte unchanged for every existing caller — prod, dev,cut-lab-ref.sh,--cold). When set, scopesbuild_images()/restart_services()to a subset.refresh_stability_lane.shsets it to the 4 known-good core services, routing around the still-open BUILD_SOURCE selector-mismatch defect on the 4 release-only services (agent-actions-consumer,skill-lifecycle-consumer,intelligence-api,omninode-contract-resolver; OMN-14262 residual) as a controlled decision instead of relying on a partial-build-failure side effect.Cadence (cron/GHA) is deliberately NOT wired — design only, per
feedback_no_landing_automation_before_process_measured: the bar is >=5 consecutive clean manual runs before a follow-up ticket to wire cadence.Canary proof (this session, against the live
.201stability-test lane)Ran the new mechanism ONCE, manually, against the live lane (refreshing to omnimarket dev HEAD, which includes 60b188ed / OMN-14855's judge-verdict DLQ fix — the prior ad hoc manual refresh at bf0aa219 predated that fix). See the session's final report for the receipt (old→new digest + ancestry proof), health-gate PASS, and the rollback-path proof.
Test plan
uv run pytest scripts/runtime_build/tests/test_verify_stability_refresh.py -v— 19/19 passeduv run pytest tests/scripts/ -q— 338/338 passed (existing deploy-runtime.sh test suite, confirms the additive patch is non-regressing)uv run ruff format --check/uv run ruff check— cleanuv run mypy scripts/runtime_build/verify_stability_refresh.py— cleanshellcheck -xon both shell scripts — cleanpre-commit run --files <changed files>— all hooks greentests/unit/) — 21833 passed, 36 skipped.201stability-test lane — health-gate PASS, lane confirmed healthy (see session report)test_receipt_rollback_reverified_success,test_receipt_rollback_still_unhealthy_is_failed_not_masked)Closes OMN-14873 (child of OMN-14263).
Evidence-Ticket: OMN-14873
Evidence-Source: OCC#4548