Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
95 changes: 95 additions & 0 deletions docs/runbooks/runner-fleet-listener-liveness.md
Original file line number Diff line number Diff line change
Expand Up @@ -327,8 +327,103 @@ remediation; all three affected runners cleared on restart.
Coverage: `tests/ci/test_runner_listener_liveness.py` —
`TestHealthcheckBrokerSessionState`.

## Reconnect-gap churn — why `offline` flaps, and why it is NOT capacity loss (OMN-16030, measured 2026-08-14)

The `offline` count in the org registry is **not** a liveness signal on this
fleet, and a red `runner-fleet-canary` is not by itself evidence of an outage.
This section supersedes any reading of layer 4 as authoritative-on-its-own; it
is the same conclusion OMN-15255 already reached ("`ready_count` — usable
capacity. **This, not `online_count`**") and OMN-14057 recorded as status-lag
corroboration, now with a measured mechanism.

**Measurement (72-runner fleet, 2026-08-14, 02:00–10:30Z, 7136 jobs):**

| Observation | Value |
|---|---|
| Runners reporting `offline` at any instant | 11–16 of 72, membership rotating between 20s polls |
| `missing` (lost registrations), every sample | **0** |
| Docker `RestartCount`, all 72 containers | **0** |
| Jobs served by the 13 persistently-offline-labelled runners (2h40m) | **153** (mean 11.8/runner vs 14.7 online — ~80% of nominal) |
| Correlation of offline count with concurrent job count | Pearson **r = +0.55** (n=7) |

Direct disproof of "offline = dead": `runner-51` completed a job at 10:24:22Z
and `runner-67` at 10:25:20Z **while both were labelled offline**; `runner-30`
held an in-progress job while labelled offline; `ps` inside `runner-23` showed
`Runner.Worker` running `uv sync` while the registry reported it offline.

**Mechanism — the same reconnect gap as OMN-15776.** Every job completion
triggers a `TaskCanceledException`/`SocketException(125)` retry storm on the
listener's broker long-poll, with 5–12s backoff. During that gap the runner has
no active broker session, so the registry reports it `offline`. This is the same
window in which OMN-15776 dispatches are dropped — the two are one phenomenon
observed from two sides:

> Of the 13 runners labelled offline at 10:17Z, **10** had an OMN-15776
> dispatch-wedge hit in the same window (expected 3.1 if independent;
> **P(≥10 by chance) = 7.5e-06**).

So `offline` tracks **reconnect churn**, which rises with job turnover — hence
the *positive* correlation with fleet load. The runners that flap offline most
are the ones cycling jobs fastest, i.e. the healthiest-utilised ones.

**Operational consequence: do not bounce on this signal.** A force-recreate
forces *more* reconnects, kills in-flight jobs, and wipes the warm tool cache
the C2 git-mirror pre-seed depends on. On 2026-08-14 a proposed fleet
"recovery" targeted 12 runners that were actively serving jobs, on the strength
of a red canary alone; the claimed dead core (runners 4/18/19/24/27/59/61) was
verified `online` **and** `busy` at that moment. Two landing sweeps were halted
for nothing.

**Residual real cost (this is the part worth fixing).** The wedge itself is
still occurring and is *not* fixed by the git mirror: 18 jobs matched the exact
OMN-15776 fingerprint (zero steps, 600–601s) in the 8h sample, 13 of them after
the mirror went live at 07:30Z, spread across 17 distinct runners with almost no
repeats — a fleet-wide GitHub-side race, not a per-runner defect. Example
citations: `omnibase_infra` job 94704437780 (runner-19, 07:32:28Z),
`onex_change_control` job 94729303340 (runner-28, 09:30:38Z), `omnibase_infra`
job 94726673643 (runner-55, 09:20:47Z). Layer-5
(`runner-broker-dispatch-wedge-rerun`, running ~every 30 min, all green) reruns
them so they do not block, but it is remediation, not prevention — the standing
cost is ~2 wasted job slots/hour plus rerun latency.

### Triage: check throughput, not status labels

```bash
# THE DECIDING CHECK — are jobs completing on self-hosted runners right now?
gh api "repos/OmniNode-ai/omnibase_infra/actions/runs?per_page=20" --jq '.workflow_runs[].id' \
| while read -r id; do
gh api "repos/OmniNode-ai/omnibase_infra/actions/runs/$id/jobs?per_page=100" \
--jq '.jobs[] | select(.runner_name|startswith("omninode-runner"))
| "\(.completed_at // "RUNNING")\t\(.runner_name)\t\(.conclusion // .status)"'
done | sort -r | head -30
```

Healthy baseline (2026-08-14, post-mirror): ~77 jobs completed per 30 min, ~40
in progress at any instant, ~44 distinct runners active per 30 min — *while
11–16 runners were labelled offline.* Declare a real outage only if job
completions have collapsed **and** `missing > 0`.

### Ruled out (do not re-litigate without new evidence)

- **DNS.** OMN-15736 proposes a local DNS cache on the premise of "no local
resolver cache" and a single-upstream chokepoint. Measured on the host
2026-08-14: 60 concurrent lookups of `files.pythonhosted.org` in **5ms**;
`pypi.org` 1ms and `files.pythonhosted.org` 0ms (cache hits — `systemd-resolved`
is already caching); zero resolution failures under burst. The premise is
falsified as stated, and DNS is upstream of nothing in the reconnect-gap path.
- **Egress bandwidth saturation** as a checkout-failure cause. The C2 git-mirror
pre-seed took checkout failures from 21/5111 jobs (0.41%) pre-07:30Z to
**0/1873 (0.00%)** after; queue-wait median 166s → 130s, p90 1009s → 859s.
- **Container crash-looping.** `RestartCount == 0` fleet-wide, `missing == 0` in
every sample.

## Operator response to a canary failure

0. **First: is this a real outage?** Run the throughput check in the
reconnect-gap section above. A red canary with jobs still completing is a
status-flap, not an outage — do not bounce anything. Since OMN-16030 the
canary only *fails* on `missing > 0` or offline-and-idle ≥ 50% of fleet; the
old advisory band now WARNs on a green run.
1. Read the failed `runner-fleet-canary` run summary — it lists offline runner names.
2. Do **NOT** `docker restart` runners (crash-loops: cached creds + expired baked token — OMN-13109).
3. Safe bounce, named services only, fresh token, detached:
Expand Down
119 changes: 101 additions & 18 deletions scripts/ci/runner_fleet_canary.sh
Original file line number Diff line number Diff line change
@@ -1,19 +1,61 @@
#!/usr/bin/env bash
# SPDX-FileCopyrightText: 2025 OmniNode.ai Inc.
# SPDX-License-Identifier: MIT
# runner_fleet_canary.sh — scheduled fleet-status canary (OMN-13915)
# runner_fleet_canary.sh — scheduled fleet-status canary (OMN-13915, OMN-16030)
#
# Compares the GitHub org self-hosted runner registry (the AUTHORITATIVE view
# of whether runners are serving jobs) against the expected fleet size declared
# in config/runner_fleet.yaml, and FAILS LOUDLY when the offline count crosses
# a threshold — BEFORE queued CI runs pile up.
# Compares the GitHub org self-hosted runner registry against the expected fleet
# size declared in config/runner_fleet.yaml, and FAILS LOUDLY on the signals that
# actually prove the fleet stopped serving jobs — BEFORE queued CI runs pile up.
#
# Why this exists: on 2026-07-03 the org API showed 37/48 runners offline while
# every runner container on .201 reported "Up (healthy)". Docker-side checks
# (healthcheck, runner-monitor cron on the host itself) share fate with the
# host; this canary runs on GitHub-hosted compute so it stays alive when the
# fleet — or the whole .201 host — is dead.
#
# OMN-16030 — WHY `status == "offline"` ALONE IS NOT A LIVENESS SIGNAL.
# This script previously treated the org REST `status` field as "the
# AUTHORITATIVE view of whether runners are serving jobs" and failed whenever
# more than RUNNER_CANARY_MAX_OFFLINE runners reported offline. On 2026-08-14
# that assumption was measured and falsified against the 72-runner fleet:
#
# * Runners reporting `offline` were concurrently executing jobs. runner-51
# completed a job at 10:24:22Z and runner-67 at 10:25:20Z while both were
# labelled offline; runner-30 held an in-progress job while labelled
# offline; `ps` inside runner-23 showed Runner.Worker running `uv sync`
# while the registry reported it offline+busy.
# * Over a 2h40m window the 13 persistently-offline-labelled runners served
# 153 jobs (mean 11.8/runner) versus 14.7/runner for online-labelled ones —
# ~80% of nominal throughput, not zero.
# * `missing` was 0 in every sample all day and RestartCount was 0 on all 72
# containers: nothing ever actually de-registered or crashed.
# * The offline count correlates POSITIVELY with concurrent job count
# (Pearson r=+0.55, n=7) — it reads worst exactly when the fleet is
# busiest, which is the opposite of a liveness signal.
#
# Mechanism: these runners use the Actions V2 broker flow (`useV2Flow: true`,
# serverUrlV2 = broker.actions.githubusercontent.com). The org REST `status`
# field reflects broker-session bookkeeping that goes stale under load; it is
# not a heartbeat. A transient stale read is not a dead listener.
#
# Cost of getting this wrong: a persistently-red canary is indistinguishable
# from a real outage, so it trains operators to ignore it AND it halts landing
# sweeps on a false alarm (observed 2026-08-14: two sweeps halted, and a
# proposed "recovery" would have force-recreated healthy runners, killing
# in-flight jobs and wiping the warm tool cache).
#
# What this canary fails on now — signals that cannot be produced by a stale
# read, in descending order of certainty:
# 1. missing > 0 — a runner lost its REGISTRATION. Unambiguous.
# 2. offline fraction >= RUNNER_CANARY_MASS_OFFLINE_PCT — mass listener death
# (the 2026-07-03 mode was 77%). Load-induced staleness has never exceeded
# ~22% in measurement, so this band separates the two cleanly.
# A `busy` runner is counted ALIVE regardless of `status`: it is demonstrably
# executing a job, which is the thing the canary exists to protect.
# Anything between the advisory threshold and the mass-outage threshold is a
# WARNING (visible in the step summary + Slack) on a GREEN run — loud enough to
# investigate, not loud enough to block landing.
#
# Enforcement surface: .github/workflows/runner-fleet-canary.yml runs this on a
# 15-minute schedule on ubuntu-latest. A threshold breach fails the workflow
# run (red X + owner notification). This is not an opt-in script.
Expand All @@ -23,7 +65,10 @@
# GET /orgs/{org}/actions/runners (classic PAT: admin:org read;
# fine-grained: org "Self-hosted runners" read).
# Optional env:
# RUNNER_CANARY_MAX_OFFLINE max offline runners tolerated (default 5)
# RUNNER_CANARY_MAX_OFFLINE offline count above which the run WARNS (default 5).
# Advisory only — see OMN-16030 note above.
# RUNNER_CANARY_MASS_OFFLINE_PCT percent of the fleet reporting offline-and-not-busy
# at which the run FAILS (default 50).
# RUNNER_FLEET_CONFIG_PATH path to runner_fleet.yaml (default config/runner_fleet.yaml)
# GITHUB_API_URL API base (set by Actions; default https://api.github.com)
# GITHUB_STEP_SUMMARY if set, a markdown summary is appended
Expand All @@ -33,6 +78,7 @@ set -euo pipefail

RUNNER_FLEET_CONFIG_PATH="${RUNNER_FLEET_CONFIG_PATH:-config/runner_fleet.yaml}"
RUNNER_CANARY_MAX_OFFLINE="${RUNNER_CANARY_MAX_OFFLINE:-5}"
RUNNER_CANARY_MASS_OFFLINE_PCT="${RUNNER_CANARY_MASS_OFFLINE_PCT:-50}"
GITHUB_API_URL="${GITHUB_API_URL:-https://api.github.com}"

log() { echo "[fleet-canary] $*"; }
Expand Down Expand Up @@ -107,28 +153,45 @@ online_count=$(jq '[ .[] | select(.status == "online") ] | length' <<< "${fleet}
offline_count=$(jq '[ .[] | select(.status != "online") ] | length' <<< "${fleet}")
missing_count=$(( EXPECTED_RUNNERS - total_registered ))
[[ "${missing_count}" -lt 0 ]] && missing_count=0
# A runner that dropped its registration entirely is offline in every way that
# matters — count it against the same threshold.
effective_offline=$(( offline_count + missing_count ))

# OMN-16030: a runner reporting offline while `busy` is demonstrably executing a
# job — the registry read is stale, the listener is not dead. Only offline AND
# not-busy runners are candidates for "actually unreachable".
offline_idle_count=$(jq '[ .[] | select(.status != "online") | select(.busy != true) ] | length' <<< "${fleet}")
offline_busy_count=$(( offline_count - offline_idle_count ))

# Mass-outage fraction is computed against offline-and-not-busy plus lost
# registrations — the two states a stale read cannot manufacture.
unreachable=$(( offline_idle_count + missing_count ))
mass_threshold=$(( EXPECTED_RUNNERS * RUNNER_CANARY_MASS_OFFLINE_PCT / 100 ))

offline_names=$(jq -r '[ .[] | select(.status != "online") | .name ] | join(", ")' <<< "${fleet}")

log "expected=${EXPECTED_RUNNERS} registered=${total_registered} online=${online_count} offline=${offline_count} missing=${missing_count} threshold=${RUNNER_CANARY_MAX_OFFLINE}"
log "expected=${EXPECTED_RUNNERS} registered=${total_registered} online=${online_count} offline=${offline_count} offline_but_busy=${offline_busy_count} offline_idle=${offline_idle_count} missing=${missing_count} unreachable=${unreachable} warn_threshold=${RUNNER_CANARY_MAX_OFFLINE} fail_threshold=${mass_threshold}"

summary() {
cat <<EOF
## Runner fleet canary (OMN-13915)
## Runner fleet canary (OMN-13915, OMN-16030)

| Metric | Value |
|--------|-------|
| Expected fleet size | ${EXPECTED_RUNNERS} |
| Registered (org API) | ${total_registered} |
| Online | ${online_count} |
| Offline | ${offline_count} |
| Offline (reported) | ${offline_count} |
| — of which busy (alive, stale read) | ${offline_busy_count} |
| — of which idle (unreachable candidates) | ${offline_idle_count} |
| Missing registrations | ${missing_count} |
| Offline threshold | ${RUNNER_CANARY_MAX_OFFLINE} |
| **Unreachable (offline-idle + missing)** | **${unreachable}** |
| Warn above | ${RUNNER_CANARY_MAX_OFFLINE} |
| **Fail at or above** | **${mass_threshold}** (${RUNNER_CANARY_MASS_OFFLINE_PCT}% of fleet) |

Offline-reporting runners: ${offline_names:-none}

Offline runners: ${offline_names:-none}
> A runner reporting \`offline\` is NOT proof it is dead. These runners use the
> Actions V2 broker flow, whose REST \`status\` goes stale under load — measured
> 2026-08-14, offline-labelled runners served ~80% of nominal job throughput
> (OMN-16030). Only lost registrations and mass offline-idle fail this gate.
EOF
}

Expand All @@ -137,19 +200,39 @@ if [[ -n "${GITHUB_STEP_SUMMARY:-}" ]]; then
fi

slack_alert() {
local severity="${1}"
local detail="${2}"
[[ -n "${SLACK_BOT_TOKEN:-}" && -n "${SLACK_CHANNEL_ID:-}" ]] || return 0
curl -s -X POST https://slack.com/api/chat.postMessage \
-H "Authorization: Bearer ${SLACK_BOT_TOKEN}" \
-H "Content-Type: application/json" \
-d "$(jq -n \
--arg channel "${SLACK_CHANNEL_ID}" \
--arg text "*[RUNNER FLEET CANARY]* ${effective_offline}/${EXPECTED_RUNNERS} runners offline-or-missing (online=${online_count}, threshold=${RUNNER_CANARY_MAX_OFFLINE}). Offline: ${offline_names:-none}. Docker 'Up (healthy)' is NOT sufficient evidence — see OMN-13915 runbook." \
--arg text "*[RUNNER FLEET CANARY — ${severity}]* ${detail} (expected=${EXPECTED_RUNNERS} online=${online_count} offline=${offline_count} offline_but_busy=${offline_busy_count} offline_idle=${offline_idle_count} missing=${missing_count}). Offline: ${offline_names:-none}. See docs/runbooks/runner-fleet-listener-liveness.md" \
'{channel: $channel, text: $text}')" > /dev/null 2>&1 || true
}

if [[ "${effective_offline}" -gt "${RUNNER_CANARY_MAX_OFFLINE}" ]]; then
slack_alert
fail "${effective_offline}/${EXPECTED_RUNNERS} runners offline-or-missing (> ${RUNNER_CANARY_MAX_OFFLINE}). Offline: ${offline_names:-<registrations missing>}. The fleet is degrading silently — do NOT trust Docker 'Up (healthy)'. See docs/runbooks/runner-fleet-listener-liveness.md"
# --- FAIL 1: lost registrations. A stale status read cannot remove a runner
# from the registry, so this is unambiguous evidence of real fleet loss.
if [[ "${missing_count}" -gt 0 ]]; then
slack_alert "FAIL" "${missing_count} runner registration(s) LOST (registered=${total_registered}/${EXPECTED_RUNNERS})"
fail "${missing_count} runner registration(s) missing (registered=${total_registered}, expected=${EXPECTED_RUNNERS}). A runner dropped its registration entirely — this is real fleet loss, not a stale status read. See docs/runbooks/runner-fleet-listener-liveness.md"
fi

# --- FAIL 2: mass listener death (the 2026-07-03 mode, 77% of fleet).
if [[ "${unreachable}" -ge "${mass_threshold}" ]]; then
slack_alert "FAIL" "${unreachable}/${EXPECTED_RUNNERS} runners offline-and-idle (>= ${mass_threshold})"
fail "${unreachable}/${EXPECTED_RUNNERS} runners offline-and-not-busy (>= ${mass_threshold} = ${RUNNER_CANARY_MASS_OFFLINE_PCT}% of fleet). At this scale it is no longer explainable as broker-status staleness — treat as mass listener death. Offline: ${offline_names:-none}. Do NOT trust Docker 'Up (healthy)'. See docs/runbooks/runner-fleet-listener-liveness.md"
fi

# --- WARN: elevated but within the band that measurement attributes to V2
# broker-status staleness. Visible, not blocking. See OMN-16030.
if [[ "${unreachable}" -gt "${RUNNER_CANARY_MAX_OFFLINE}" ]]; then
slack_alert "WARN" "${unreachable}/${EXPECTED_RUNNERS} runners offline-and-idle (> ${RUNNER_CANARY_MAX_OFFLINE}, below fail threshold ${mass_threshold})"
log "WARN: ${unreachable}/${EXPECTED_RUNNERS} offline-and-idle — above the advisory threshold (${RUNNER_CANARY_MAX_OFFLINE}) but below the mass-outage threshold (${mass_threshold})."
log "WARN: this band is attributed to V2 broker-status staleness (OMN-16030). Confirm with actual job throughput before any restart — offline-labelled runners are usually still serving jobs."
log "OK: fleet serving; not failing on a status-staleness signal."
exit 0
fi

log "OK: fleet within threshold."
Loading