Skip to content

fix(OMN-16030): stop failing the runner-fleet canary on reconnect-gap status flap - #2749

Merged
jonahgabriel merged 1 commit into
devfrom
jonah/omn-16030-fleet-canary-liveness
Aug 15, 2026
Merged

jonahgabriel merged 1 commit into
devfrom
jonah/omn-16030-fleet-canary-liveness

Conversation

@jonahgabriel

@jonahgabriel jonahgabriel commented Aug 15, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

The runner-fleet canary treats the org REST status field as authoritative liveness and fails whenever offline+missing > 5. On a 72-runner fleet that produces a persistent false red: measured over 7136 jobs (02:00-10:30Z on 2026-08-14), runners labelled offline were completing jobs at the moment of the label (runner-51 at 10:24:22Z, runner-67 at 10:25:20Z, runner-30 holding an in-progress job), the 13 persistently-offline-labelled runners served 153 jobs over 2h40m (~80% of nominal throughput), missing was 0 in every sample, Docker RestartCount was 0 on all 72 containers, and the offline count correlates positively with concurrent job load (r=+0.55) — the opposite of a liveness signal. The mechanism is the known reconnect gap: every job completion triggers a retry storm on the listener's broker long-poll, and during that 5-12s backoff window there is no active broker session, so the registry reports offline.

Changes

  • scripts/ci/runner_fleet_canary.sh: the gate now fails only on signals a reconnect gap cannot manufacture: (1) missing > 0 — a lost registration is unambiguous fleet loss; (2) offline-and-not-busy >= 50% of fleet — mass listener death. A runner that is offline-but-busy counts ALIVE, since it is provably executing a job. The band between the advisory threshold and 50% WARNs instead of failing.
  • docs/runbooks/runner-fleet-listener-liveness.md: additive "Reconnect-gap churn" section (all 363 prior lines preserved) covering the measurement, a throughput-based triage recipe, the ruled-out hypotheses (DNS, egress saturation, crash-looping), and a new step 0 in the operator response: check throughput before bouncing anything.

Detection power preserved

Replaying the 2026-07-03 incident this canary was built for (37/48 runners offline-idle = 77%) still yields FAIL; observed flap has never exceeded ~22%, so the bands do not overlap. Today's real registry snapshots yield PASS+WARN under the new gate (unreachable=10 vs fail threshold 36).

Why the false red is expensive

It is indistinguishable from a real outage. It halted two landing sweeps, and the mechanical "recovery" it invites would have force-recreated 12 runners that were actively serving jobs — killing in-flight work and wiping the warm tool cache the egress mirror pre-seed depends on.

Verification

  • bash -n clean; canary run against live registry snapshots (PASS+WARN) and against the replayed 2026-07-03 incident data (FAIL)
  • pre-push governed selector passed honestly on push (no bypass, no skip token)

Ticket: OMN-16030

Evidence-Source: OCC#6514
Evidence-Ticket: OMN-16030

Change-control companion

onex_change_control PR #6514 carries contracts/OMN-16030.yaml with PASS receipts binding this fix at the pinned head 175c3809cc2e0a2f8729b7748f22fb685c2c06bc; merged to OCC dev.

… status flap

The canary treated the org REST `status` field as "the AUTHORITATIVE view of
whether runners are serving jobs" and failed whenever offline+missing > 5. On a
72-runner fleet that produces a persistent false red. Measured 2026-08-14 over
7136 jobs (02:00-10:30Z).

Evidence that `offline` is not liveness:
- runner-51 completed a job at 10:24:22Z and runner-67 at 10:25:20Z while both
  were labelled offline; runner-30 held an in-progress job while labelled
  offline; runner-23 showed Runner.Worker running `uv sync` while the registry
  reported it offline.
- The 13 persistently-offline-labelled runners served 153 jobs over 2h40m
  (mean 11.8/runner vs 14.7 online) -- ~80% of nominal, not zero.
- `missing` was 0 in every sample; Docker RestartCount was 0 on all 72
  containers. Nothing de-registered, nothing crash-looped.
- Offline count correlates POSITIVELY with concurrent job count (r=+0.55, n=7):
  it reads worst when the fleet is busiest, the opposite of a liveness signal.

Mechanism -- this is not generic "staleness", it is the OMN-15776 reconnect gap
observed from the other side. Every job completion triggers a retry storm on the
listener's broker long-poll (5-12s backoff); during that gap there is no active
broker session, so the registry reports `offline`, and that is the same window
in which OMN-15776 dispatches are dropped. Of the 13 runners labelled offline at
10:17Z, 10 had an OMN-15776 dispatch-wedge hit in the same window -- expected
3.1 if independent, P(>=10 by chance) = 7.5e-06.

This aligns the canary with what OMN-15255 already concluded ("`ready_count` --
usable capacity. This, not `online_count`") and OMN-14057 recorded as status-lag
corroboration; layer 4 was simply never updated to match.

The gate now fails only on signals a reconnect gap cannot manufacture:
1. `missing > 0` -- a lost registration is unambiguous real fleet loss.
2. offline-and-not-busy >= 50% of fleet -- mass listener death. The 2026-07-03
   incident this canary was built for was 37/48 = 77%; observed flap has never
   exceeded ~22%, so the bands do not overlap.
A runner offline-but-busy counts ALIVE: it is provably executing a job. The band
between the advisory threshold and 50% now WARNs on a green run.

Verified against real data: today's registry snapshots yield PASS+WARN
(unreachable=10, fail threshold 36); a replay of the 2026-07-03 mode (37/48
offline-idle) still yields FAIL -- detection power preserved.

Cost of the false red: it is indistinguishable from a real outage. It halted two
landing sweeps, and the proposed "recovery" would have force-recreated 12
runners that were actively serving jobs -- killing in-flight work and wiping the
warm tool cache the C2 mirror pre-seed depends on. The claimed dead core
(runners 4/18/19/24/27/59/61) was verified online AND busy at that moment.

Runbook: appends a "Reconnect-gap churn" section to the existing
docs/runbooks/runner-fleet-listener-liveness.md (all 363 prior lines preserved;
this is additive) covering the measurement, the shared mechanism with OMN-15776,
a throughput-based triage recipe, and the ruled-out hypotheses (DNS -- 60
concurrent lookups in 5ms, systemd-resolved already caching, so the OMN-15736
premise is falsified as stated; egress saturation -- fixed by the C2 mirror,
checkout failures 0.41% -> 0.00%; crash-looping -- RestartCount 0 fleet-wide).
Adds step 0 to the operator response: check throughput before bouncing anything.

NOT fixed here, and still real: the OMN-15776 wedge itself. 18 jobs matched its
exact fingerprint (zero steps, 600-601s) in the 8h sample, 13 of them after the
mirror went live, across 17 distinct runners with almost no repeats -- a
fleet-wide GitHub-side race. Layer 5 reruns them so they do not block, but that
is remediation, not prevention (~2 wasted job slots/hour).
@coderabbitai

coderabbitai Bot commented Aug 15, 2026

Copy link
Copy Markdown

Warning

Review limit reached

You’ve reached a temporary PR review limit under our Fair Usage Limits Policy.

Your recent review volume is higher than typical usage, so adaptive limits are currently applied.

Next review available in: 59 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: d5e4c228-a81b-4fb4-b617-e6021130f72c

📥 Commits

Reviewing files that changed from the base of the PR and between c1032d5 and 175c380.

📒 Files selected for processing (2)
  • docs/runbooks/runner-fleet-listener-liveness.md
  • scripts/ci/runner_fleet_canary.sh

Comment @coderabbitai help to get the list of available commands.

@jonahgabriel
jonahgabriel enabled auto-merge (squash) August 15, 2026 03:40
@onexbot-occ-writer

Copy link
Copy Markdown
Contributor

OCC autobind did not mint a companion for this PR: no changed-file candidate could be proven RED against the merge base, and emitting a PR-existence probe instead would be non-falsifiable evidence (OMN-15247). Hand-authored evidence is required.

jonahgabriel added a commit to OmniNode-ai/onex_change_control that referenced this pull request Aug 15, 2026
#6514)

* evidence(OMN-16030): OCC companion for OmniNode-ai/omnibase_infra#2749

Binds the evidence for the runner-fleet canary fix: the gate now fails
only on signals a reconnect gap cannot manufacture (lost registration,
mass offline-and-not-busy), an offline-but-busy runner counts alive, and
the listener-liveness runbook gains the reconnect-gap churn section.
Every content probe is a readback at the pinned product head 175c3809c
and is RED at the pre-fix dev base c1032d5ae.

* evidence(OMN-16030): self-bind the companion to OCC PR #6514
@github-actions

Copy link
Copy Markdown
Contributor

❓ Hostile Reviewer — UNKNOWN

Blocking findings (critical): 0
Total findings: 0
Models succeeded: none


Gate semantics (pilot phase)

Verdict Meaning Blocks merge?
passed No critical findings No
blocked CRITICAL findings found Yes
degraded All models unavailable (infra) No (pilot)

Powered by omniintelligence.review_pairing.cli_review — node-based adversarial review via HandlerLlmCliSubprocess (OMN-8468/OMN-8524)

@jonahgabriel
jonahgabriel merged commit 0f38322 into dev Aug 15, 2026
326 of 392 checks passed
@jonahgabriel
jonahgabriel deleted the jonah/omn-16030-fleet-canary-liveness branch August 15, 2026 08:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant