Skip to content

feat(OMN-14873): health-gated, rollback-on-failure stability-lane refresh mechanism - #2370

Merged
jonahgabriel merged 4 commits into
devfrom
jonah/omn-14873-stability-lane-refresh-mechanism
Jul 21, 2026
Merged

jonahgabriel merged 4 commits into
devfrom
jonah/omn-14873-stability-lane-refresh-mechanism

Conversation

@jonahgabriel

@jonahgabriel jonahgabriel commented Jul 20, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Implements + canaries the design in child ticket OMN-14873 (child of OMN-14263: "stability-test proof lane has no refresh cadence").

  • scripts/runtime_build/refresh_stability_lane.sh — new. Health-gated, rollback-on-failure refresh of the .201 stability-test lane. Hardcoded to that lane only (no --lane flag, mirrors cut-lab-ref.sh's prod/judge refusal). Captures pre-state (image IDs + a preflight-<UTC> rollback tag) before any mutation, refreshes the omnibase_infra ambient clone + the 4 DEPLOY_REF-pinned siblings to --ref (default origin/dev), builds+restarts SCOPED to the 4 known-good core services (omninode-runtime, runtime-effects, runtime-worker, projection-api), asserts forward progress (git merge-base --is-ancestor) for every tracked repo, runs the health-gate, and on failure rolls back + re-verifies. Emits a durable JSON receipt under ~/.omnibase/state/stability_lane_refresh/.
  • scripts/runtime_build/verify_stability_refresh.py — new. The health-gate: digest-changed, manifest contract-count floor, /health, rpk cluster health, declared consumer groups Stable (consumer_groups_stability.yaml), and image-revision readback. --no-require-digest-change supports the post-rollback re-verification pass (where digest is deliberately unchanged).
  • scripts/runtime_build/consumer_groups_stability.yaml — new. Declared consumer-group health-check list (contract-as-data), spanning registration + the delegation pipeline + node_ticket_classify_compute (the def-B canary this refresh also proves live).
  • scripts/runtime_build/tests/test_verify_stability_refresh.py — new. 19 unit tests (mocked docker/HTTP, no live lane needed): PASS/FAIL boundary at exactly min_contracts, digest-unchanged, revision-mismatch (exists-but-wrong, not silent pass), and the rollback-re-verification receipt logic including the "rollback ALSO fails" branch staying FAILED (never masked as success).
  • docs/runbooks/stability-lane-refresh.md — new canonical runbook.
  • scripts/deploy-runtime.sh — one small additive change: RUNTIME_BUILD_SERVICES_OVERRIDE env var, defaulting to current RUNTIME_SERVICES behavior (byte-for-byte unchanged for every existing caller — prod, dev, cut-lab-ref.sh, --cold). When set, scopes build_images()/restart_services() to a subset. refresh_stability_lane.sh sets it to the 4 known-good core services, routing around the still-open BUILD_SOURCE selector-mismatch defect on the 4 release-only services (agent-actions-consumer, skill-lifecycle-consumer, intelligence-api, omninode-contract-resolver; OMN-14262 residual) as a controlled decision instead of relying on a partial-build-failure side effect.

Cadence (cron/GHA) is deliberately NOT wired — design only, per feedback_no_landing_automation_before_process_measured: the bar is >=5 consecutive clean manual runs before a follow-up ticket to wire cadence.

Canary proof (this session, against the live .201 stability-test lane)

Ran the new mechanism ONCE, manually, against the live lane (refreshing to omnimarket dev HEAD, which includes 60b188ed / OMN-14855's judge-verdict DLQ fix — the prior ad hoc manual refresh at bf0aa219 predated that fix). See the session's final report for the receipt (old→new digest + ancestry proof), health-gate PASS, and the rollback-path proof.

Test plan

  • uv run pytest scripts/runtime_build/tests/test_verify_stability_refresh.py -v — 19/19 passed
  • uv run pytest tests/scripts/ -q — 338/338 passed (existing deploy-runtime.sh test suite, confirms the additive patch is non-regressing)
  • uv run ruff format --check / uv run ruff check — clean
  • uv run mypy scripts/runtime_build/verify_stability_refresh.py — clean
  • shellcheck -x on both shell scripts — clean
  • pre-commit run --files <changed files> — all hooks green
  • Full governed pre-push selector (escalated to tests/unit/) — 21833 passed, 36 skipped
  • Manual canary run against the live .201 stability-test lane — health-gate PASS, lane confirmed healthy (see session report)
  • Rollback path — proven via unit tests (test_receipt_rollback_reverified_success, test_receipt_rollback_still_unhealthy_is_failed_not_masked)

Closes OMN-14873 (child of OMN-14263).

Evidence-Ticket: OMN-14873
Evidence-Source: OCC#4548

…resh mechanism

Wraps the proven manual .201 stability-test refresh recipe (workspace-mode
build scoped to the 4 known-good core services + targeted recreate) into
scripts/runtime_build/refresh_stability_lane.sh, with pre-state capture,
a forward-progress ancestry assertion, a health-gate
(verify_stability_refresh.py: digest changed, manifest contract-count floor,
/health, rpk cluster health, declared consumer groups Stable, image-revision
readback), automatic rollback-and-re-verify on failure, and a durable JSON
receipt under ~/.omnibase/state/stability_lane_refresh/.

deploy-runtime.sh gets one small additive change (RUNTIME_BUILD_SERVICES_OVERRIDE,
defaulting to current behavior) so the build/restart can be scoped to a
service subset -- this routes the 4 known-broken release-only services
(BUILD_SOURCE selector-mismatch, OMN-14262 residual) around as a controlled
decision instead of relying on a partial-build-failure side effect.

Cadence (cron/GHA) is deliberately NOT wired -- design only, per
feedback_no_landing_automation_before_process_measured, gated on >=5
consecutive clean manual runs.

Child of OMN-14263.
@coderabbitai

coderabbitai Bot commented Jul 20, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

You’ve reached a temporary PR review limit under our Fair Usage Limits Policy.

Your recent review volume is higher than typical usage, so adaptive limits are currently applied.

Next review available in: 58 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 55eb50f8-76fc-4cb2-aea1-fe7dd3a58579

📥 Commits

Reviewing files that changed from the base of the PR and between 43935c8 and b30de03.

📒 Files selected for processing (7)
  • docs/runbooks/stability-lane-refresh.md
  • scripts/deploy-runtime.sh
  • scripts/runtime_build/consumer_groups_stability.yaml
  • scripts/runtime_build/refresh_stability_lane.sh
  • scripts/runtime_build/tests/test_verify_stability_refresh.py
  • scripts/runtime_build/verify_stability_refresh.py
  • tests/unit/scripts/test_refresh_stability_lane_checkout.py
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch jonah/omn-14873-stability-lane-refresh-mechanism

Comment @coderabbitai help to get the list of available commands.

Jonah Gray added 2 commits July 20, 2026 16:52
…rom set -e

Canary run surfaced two real bugs in refresh_stability_lane.sh:

1. The rollback's direct `docker compose up --force-recreate` call did not
   source docker/runtime-policy.env + ~/.omnibase/.env the way
   deploy-runtime.sh's own compose calls do, so it failed at config
   interpolation time (e.g. BIFROST_VERIFY_ENDPOINTS unset).
2. That failure was unguarded under `set -e`, so the whole script aborted
   with NO receipt written at all -- worse than the intended FAILED/
   FAILED_ROLLED_BACK receipt outcomes.

Fixes: source the same env files at script top (preserving operator OMNI_HOME
like deploy-runtime.sh's header does); guard the rollback recreate call and
both health-gate JSON reads so a crash anywhere in this path still reaches
the final receipt-writing step with an honest FAILED/INFRA_ERROR record
instead of a bare script crash.
…egex

Real canary findings against the live .201 stability-test lane:

1. `rpk group describe` has NO `-f json` output mode on this rpk version
   (rejects -f as unknown) -- parse the plain-text STATE line instead.
2. Declared consumer-group names in consumer_groups_stability.yaml were bare
   prefixes; the real rpk group name carries a `.__i.<instance>.__t.<topic>`
   per-topic suffix. The bare prefix resolves to a distinct, nonexistent
   group (STATE=Dead) -- fixed to use the full names copied from `rpk group
   list`.
3. `rpk cluster health`'s "Healthy:" line padding varies by rpk version
   (observed much wider than the hardcoded exact-spacing match) -- switched
   to a regex match.
4. "Empty" (group registered, zero current members) is the NORMAL idle state
   for a demand-driven consumer and was producing false-negative rollbacks
   even though the refresh itself was healthy -- only "Dead"/absent (the
   actual "never registered a group" wiring-death signal) now fails the
   check. Added a bounded retry (mirrors deploy-runtime.sh's
   assert_broker_reachable retry shape) since a freshly-recreated container
   needs a few seconds to reconnect and rejoin its group.

19 -> 23 unit tests (added Empty-is-healthy, Dead-is-not, retry-recovers,
retry-exhausts-and-reports coverage).
Release-train canary (OMN-14889) found a real .201 failure: this script's
step 3 ("Refresh the omnibase_infra ambient clone itself to --ref") ran

    git checkout dev
    git reset --hard "${REF}"

which hard-fails whenever another worktree on the same host (e.g.
deploy-agent's runtime-sync-worktrees/OMN-12618) already has the local
`dev` branch checked out -- git refuses to check the same branch out
into two worktrees at once. This made the deploy step environment-
fragile: it only passed once because the collision happened not to be
present that run.

Fix: resolve --ref to a commit SHA and check it out DETACHED
(`git checkout --force --detach <sha>`). A detached HEAD carries no
branch identity, so it can never collide with a sibling worktree
regardless of what branch that worktree holds. --force preserves the
previous checkout+reset --hard pair's force-discard-local-changes
behavior. The forward-progress ancestry assertion and everything else
is untouched.

Adds tests/unit/scripts/test_refresh_stability_lane_checkout.py:
real-git RED/GREEN reproduction of the exact collision (branch
checkout fails while `dev` is busy in a sibling worktree; detached-
by-SHA checkout succeeds under the identical condition, and leaves the
sibling worktree's own `dev` checkout completely undisturbed), plus a
static regression guard on the shipped script so the collision-prone
bare-branch-checkout pattern cannot silently come back.

Ticket: OMN-14889 (child of OMN-13674)
dod_evidence: 3/3 new tests pass (RED confirmed against the pre-fix
script content read via `git show`); full runtime_build + scripts unit
suite (799 tests) green, no regressions; shellcheck clean; pre-commit
--files clean.
@jonahgabriel
jonahgabriel merged commit 54938fd into dev Jul 21, 2026
218 of 259 checks passed
@jonahgabriel
jonahgabriel deleted the jonah/omn-14873-stability-lane-refresh-mechanism branch July 21, 2026 19:35
jonahgabriel added a commit that referenced this pull request Jul 21, 2026
…ty lab lanes (re-cut onto dev) (#2388)

* feat(OMN-14889): git-tag-driven release-train trigger for dev+stability lab lanes

Re-cut of closed PR #2377 (commit c2ec0d2) onto current dev after #2370
(OMN-14873) merged as 54938fd. File-scoped restore of the 6 paths only; dev's
newer refresh_stability_lane.sh (detached-checkout fix) is left untouched.

- .github/workflows/release-train-lab.yml (new)
- docker/docker-compose.runners.yml (+74, additive)
- docs/runbooks/release-train-lab.md (new)
- scripts/runtime_build/cut_release_train_tag.sh (new)
- scripts/runtime_build/refresh_dev_lane.sh (new)
- scripts/runtime_build/verify_dev_refresh.py (new)

* docs(OMN-14889): correct false "runner unprovisioned" safety claims

The workflow header, runbook, and runners compose block all asserted the
omnibase-deploy runner was unprovisioned and that the deploy job "will queue
but never pick up." That is FALSE and safety-relevant: a reviewer would
conclude a lab/** tag push is inert when it in fact triggers a real lane
refresh in execute mode.

Verified live 2026-07-21:
  gh api repos/OmniNode-ai/omnibase_infra/actions/runners
  -> omninode-deploy-runner status=online,
     labels [self-hosted, Linux, X64, omnibase-deploy]

That is the exact label set both jobs target, and registration is at the
REPOSITORY level (not the org level the runbook recipe assumed).

Corrected all three sites to state: the runner is online; a lab/dev/** or
lab/stability/** tag push DOES refresh the lane in execute mode (intended,
since dev+stability are pre-authorized, but not a no-op); and the remaining
blocker is host-side git access to the ambient $OMNI_HOME clone, not
registration.

Five runs on 2026-07-21T03:59-04:19Z all failed in the "Refresh stability
lane" step, in three distinct forms as intermediate fixes were attempted:
dubious ownership (128), .git/FETCH_HEAD Permission denied (255), and 'dev'
already checked out in a sibling worktree (128). All died before any
container action -- zero docker compose / up -d / --force-recreate lines in
any of the five logs -- so no lane was mutated.

Prod boundary statements preserved and reinforced: lab/** and v* remain
disjoint namespaces and prod promotion still routes through
node_redeploy_orchestrator's grant gate.

No functional change; comments and docs only.

Closes OMN-14889

* fix(OMN-14889): harden dev release-train refresh

---------

Co-authored-by: Jonah Gray <jonah.g.gray@gmail.com>
jonahgabriel added a commit that referenced this pull request Jul 22, 2026
…ator (count roots, not checks) (#2391)

Plan section 3.C5 (docs/plans/2026-07-21-ci-capacity-recovery-plan.md); root-cause
class CI-01 from the 2026-07-16 CI remediation plan - designed, not shipped.

On omnibase_infra#2370 a single defect (missing Evidence-Source -> occ-preflight
eligibility red) amplifies into 36 red check-runs (35 occ-preflight/eligibility, one
per calling workflow, + 1 CI Summary) plus 74 skipped needs:occ-preflight dependents
- ~37x amplification. Because each occ red is its own workflow run, each is
independently rerunnable, which provokes the rerun reflex (a wall of red invites
broad reruns, each re-spending the 40-way matrix).

Ships scripts/ci/ci_cascade_graph.py, which consumes the head SHA's full
cross-workflow check-run list (GET /commits/{sha}/check-runs) and collapses it into:
  - exactly ONE typed root cause - root election reuses the deterministic,
    unit-tested product_reason_graph.build_reason_graph classifier, preserving the
    OCC-independence property (a real product defect roots as PRODUCT_FAILED, not
    EVIDENCE_MISSING);
  - every other red/skipped check marked BLOCKED_UPSTREAM, content-addressed to the
    single root_receipt_id; and
  - EXACTLY one rerunnable unit (the root) - every dependent carries
    rerunnable=False, the anti-cosmetic guarantee.

Wired as an additive, report-only render in the CI Summary job: if always() +
continue-on-error true, writes only to GITHUB_STEP_SUMMARY, ends every command with
|| true. It NEVER changes the verdict - the poll step remains the sole pass/fail
authority. Gate not weakened: a seeded occ-preflight red still yields ci_summary_gate
exit 1 (CI Summary = FAILURE).

Proof: tests/ci/test_ci_cascade_graph_omn14909.py replays #2370's real captured
blocked head (tests/ci/fixtures/omn14909_2370_cascade_checkruns.json) - 1 root
(EVIDENCE_MISSING), 35 red + 74 skipped dependents all BLOCKED_UPSTREAM, 1 rerunnable
unit, no dependent independently rerunnable. Plus OCC-independence, determinism,
green-head base case, and report-only-exit-0 CLI tests.

dod_evidence:
- uv run pytest tests/ci/test_ci_cascade_graph_omn14909.py -> 7 passed
- regression test_product_reason_graph_omn14707.py + test_ci_summary_gate.py -> 58 passed
- uv run mypy scripts/ci/ci_cascade_graph.py -> clean; ruff clean; pre-commit 46 passed / 0 failed
- gate-not-weakened: seeded occ-preflight-red jobs.json -> ci_summary_gate.py exit 1 (FAILURE)
- actionlint: 0 new findings from the added CI Summary step

Report-only; no gate weakened; no branch-protection writes; single-repo (fan-out to
core/omnimarket/omniclaude/onex_change_control is a follow-up).

Closes OMN-14909

Co-authored-by: Jonah Gray <jonah.g.gray@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant