Repository navigation
fix(OMN-16030): stop failing the runner-fleet canary on reconnect-gap status flap - #2749
Conversation
… status flap
The canary treated the org REST `status` field as "the AUTHORITATIVE view of
whether runners are serving jobs" and failed whenever offline+missing > 5. On a
72-runner fleet that produces a persistent false red. Measured 2026-08-14 over
7136 jobs (02:00-10:30Z).
Evidence that `offline` is not liveness:
- runner-51 completed a job at 10:24:22Z and runner-67 at 10:25:20Z while both
were labelled offline; runner-30 held an in-progress job while labelled
offline; runner-23 showed Runner.Worker running `uv sync` while the registry
reported it offline.
- The 13 persistently-offline-labelled runners served 153 jobs over 2h40m
(mean 11.8/runner vs 14.7 online) -- ~80% of nominal, not zero.
- `missing` was 0 in every sample; Docker RestartCount was 0 on all 72
containers. Nothing de-registered, nothing crash-looped.
- Offline count correlates POSITIVELY with concurrent job count (r=+0.55, n=7):
it reads worst when the fleet is busiest, the opposite of a liveness signal.
Mechanism -- this is not generic "staleness", it is the OMN-15776 reconnect gap
observed from the other side. Every job completion triggers a retry storm on the
listener's broker long-poll (5-12s backoff); during that gap there is no active
broker session, so the registry reports `offline`, and that is the same window
in which OMN-15776 dispatches are dropped. Of the 13 runners labelled offline at
10:17Z, 10 had an OMN-15776 dispatch-wedge hit in the same window -- expected
3.1 if independent, P(>=10 by chance) = 7.5e-06.
This aligns the canary with what OMN-15255 already concluded ("`ready_count` --
usable capacity. This, not `online_count`") and OMN-14057 recorded as status-lag
corroboration; layer 4 was simply never updated to match.
The gate now fails only on signals a reconnect gap cannot manufacture:
1. `missing > 0` -- a lost registration is unambiguous real fleet loss.
2. offline-and-not-busy >= 50% of fleet -- mass listener death. The 2026-07-03
incident this canary was built for was 37/48 = 77%; observed flap has never
exceeded ~22%, so the bands do not overlap.
A runner offline-but-busy counts ALIVE: it is provably executing a job. The band
between the advisory threshold and 50% now WARNs on a green run.
Verified against real data: today's registry snapshots yield PASS+WARN
(unreachable=10, fail threshold 36); a replay of the 2026-07-03 mode (37/48
offline-idle) still yields FAIL -- detection power preserved.
Cost of the false red: it is indistinguishable from a real outage. It halted two
landing sweeps, and the proposed "recovery" would have force-recreated 12
runners that were actively serving jobs -- killing in-flight work and wiping the
warm tool cache the C2 mirror pre-seed depends on. The claimed dead core
(runners 4/18/19/24/27/59/61) was verified online AND busy at that moment.
Runbook: appends a "Reconnect-gap churn" section to the existing
docs/runbooks/runner-fleet-listener-liveness.md (all 363 prior lines preserved;
this is additive) covering the measurement, the shared mechanism with OMN-15776,
a throughput-based triage recipe, and the ruled-out hypotheses (DNS -- 60
concurrent lookups in 5ms, systemd-resolved already caching, so the OMN-15736
premise is falsified as stated; egress saturation -- fixed by the C2 mirror,
checkout failures 0.41% -> 0.00%; crash-looping -- RestartCount 0 fleet-wide).
Adds step 0 to the operator response: check throughput before bouncing anything.
NOT fixed here, and still real: the OMN-15776 wedge itself. 18 jobs matched its
exact fingerprint (zero steps, 600-601s) in the 8h sample, 13 of them after the
mirror went live, across 17 distinct runners with almost no repeats -- a
fleet-wide GitHub-side race. Layer 5 reruns them so they do not block, but that
is remediation, not prevention (~2 wasted job slots/hour).
|
Warning Review limit reachedYou’ve reached a temporary PR review limit under our Fair Usage Limits Policy. Next review available in: 59 minutes Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (2)
Comment |
|
OCC autobind did not mint a companion for this PR: no changed-file candidate could be proven RED against the merge base, and emitting a PR-existence probe instead would be non-falsifiable evidence (OMN-15247). Hand-authored evidence is required. |
#6514) * evidence(OMN-16030): OCC companion for OmniNode-ai/omnibase_infra#2749 Binds the evidence for the runner-fleet canary fix: the gate now fails only on signals a reconnect gap cannot manufacture (lost registration, mass offline-and-not-busy), an offline-but-busy runner counts alive, and the listener-liveness runbook gains the reconnect-gap churn section. Every content probe is a readback at the pinned product head 175c3809c and is RED at the pre-fix dev base c1032d5ae. * evidence(OMN-16030): self-bind the companion to OCC PR #6514
❓ Hostile Reviewer — UNKNOWNBlocking findings (critical): 0 Gate semantics (pilot phase)
Powered by omniintelligence.review_pairing.cli_review — node-based adversarial review via HandlerLlmCliSubprocess (OMN-8468/OMN-8524) |
Summary
The runner-fleet canary treats the org REST
statusfield as authoritative liveness and fails whenever offline+missing > 5. On a 72-runner fleet that produces a persistent false red: measured over 7136 jobs (02:00-10:30Z on 2026-08-14), runners labelledofflinewere completing jobs at the moment of the label (runner-51 at 10:24:22Z, runner-67 at 10:25:20Z, runner-30 holding an in-progress job), the 13 persistently-offline-labelled runners served 153 jobs over 2h40m (~80% of nominal throughput),missingwas 0 in every sample, Docker RestartCount was 0 on all 72 containers, and the offline count correlates positively with concurrent job load (r=+0.55) — the opposite of a liveness signal. The mechanism is the known reconnect gap: every job completion triggers a retry storm on the listener's broker long-poll, and during that 5-12s backoff window there is no active broker session, so the registry reportsoffline.Changes
scripts/ci/runner_fleet_canary.sh: the gate now fails only on signals a reconnect gap cannot manufacture: (1)missing > 0— a lost registration is unambiguous fleet loss; (2) offline-and-not-busy >= 50% of fleet — mass listener death. A runner that is offline-but-busy counts ALIVE, since it is provably executing a job. The band between the advisory threshold and 50% WARNs instead of failing.docs/runbooks/runner-fleet-listener-liveness.md: additive "Reconnect-gap churn" section (all 363 prior lines preserved) covering the measurement, a throughput-based triage recipe, the ruled-out hypotheses (DNS, egress saturation, crash-looping), and a new step 0 in the operator response: check throughput before bouncing anything.Detection power preserved
Replaying the 2026-07-03 incident this canary was built for (37/48 runners offline-idle = 77%) still yields FAIL; observed flap has never exceeded ~22%, so the bands do not overlap. Today's real registry snapshots yield PASS+WARN under the new gate (unreachable=10 vs fail threshold 36).
Why the false red is expensive
It is indistinguishable from a real outage. It halted two landing sweeps, and the mechanical "recovery" it invites would have force-recreated 12 runners that were actively serving jobs — killing in-flight work and wiping the warm tool cache the egress mirror pre-seed depends on.
Verification
bash -nclean; canary run against live registry snapshots (PASS+WARN) and against the replayed 2026-07-03 incident data (FAIL)Ticket: OMN-16030
Evidence-Source: OCC#6514
Evidence-Ticket: OMN-16030
Change-control companion
onex_change_controlPR #6514 carriescontracts/OMN-16030.yamlwith PASS receipts binding this fix at the pinned head175c3809cc2e0a2f8729b7748f22fb685c2c06bc; merged to OCC dev.