Skip to content

fix(OMN-15509): system-health Slack alert must probe every lane's main runtime endpoint - #2572

Merged
jonahgabriel merged 1 commit into
devfrom
jonah/omn-15509-system-health-alert-lane-runtime-endpoints
Jul 30, 2026
Merged

jonahgabriel merged 1 commit into
devfrom
jonah/omn-15509-system-health-alert-lane-runtime-endpoints

Conversation

@jonahgabriel

@jonahgabriel jonahgabriel commented Jul 30, 2026 •

Copy link
Copy Markdown
Collaborator

Closes OMN-15509.

AC1 — the producing alert code, named, with the evidence that identifies it

Producer: /data/maintenance/bin/omninode-system-slack-report.sh on .201 (omninode-pc), executed as root by /etc/cron.d/omninode-system-slack-report:

5 8  * * *   root  /data/maintenance/bin/omninode-system-slack-report.sh --mode digest
*/15 * * * * root  /data/maintenance/bin/omninode-system-slack-report.sh --mode alert

Confirmed live 2026-07-30T17:00:01-04:00 in /var/log/syslog:
CRON[3258383]: (root) CMD (/data/maintenance/bin/omninode-system-slack-report.sh --mode alert).

The script was untracked by every repository — that is why the defect survived: there was no diff to review, no test to fail, and no grep in any repo that could find it. grep -rn "OmniNode system alert" across the whole workspace returns nothing.

Message matched to the emitting call site. The Slack MCP was down, so the message was read back read-only through the bot token already provisioned on .201 (conversations.history). The 2026-07-30T16:30:02Z message — inside the outage window — is verbatim:

*OmniNode system alert*
Host: omninode-pc
Issues: *1 critical*, *0 warning*
...
*Runtime endpoints*
- `runtime-18085`: HTTP 200 (OK)
- `runtime-28085`: HTTP 200 (OK)
- `projection-api-13002`: HTTP 200 (OK)
- `deploy-agent-8099`: HTTP 200 (OK)
- `web-3003`: HTTP 200 (OK)

That endpoint list is produced character-for-character by collect() + format_digest() in the as-deployed script (tests/fixtures/omn15509/omninode-system-slack-report.as-deployed-20260730.sh.captured:129-133, :151). No runtime-8085. The *OmniNode morning system digest* and *[OmniNode alert resolved]* strings in the same channel come from the same file's case "$MODE" block.

Candidate #1 from the ticket is REFUTED, not merely unconfirmed

omnimarket/src/omnimarket/nodes/node_baseline_capture/handlers/probes/probe_system_health.py is not the producer:

  1. It feeds node_baseline_capture, whose only publish topic is onex.evt.omnimarket.baseline-captured.v1. That topic is recorded as ORPHANED_PRODUCER and the node as DISCONNECTED_SUBGRAPH in the frozen contract graph (omnimarket/src/omnimarket/validators/data/contract_topic_graph_baseline.yaml:14,345); validators/contract_topic_graph.py:187 names it as one of "2 genuinely dead nodes". No Slack module imports it.
  2. Live bus: onex.cmd.omnimarket.slack-publish.v1 and onex.evt.omnimarket.slack-published.v1 are at HW=0 on all three lanes (dev/stability-test/prod), as is onex.cmd.omniclaw.slack-outbound.v1. Nothing reached Slack over the bus.
  3. No container on .201 carries any SLACK_* env var (checked across all 105 running containers via docker inspect), so no containerized service is a Slack producer at all.

Also ruled out along the way: scripts/system_health_check.sh has zero Slack references and its dev-lane-liveness.yml caller was correctly RED all day (14 consecutive failure runs, 06:02Z-16:54Z); monitor_logs.py is not running (onex-log-monitor.service = inactive/not-found); runner-monitor.sh and runner_fleet_canary.sh emit runner-fleet messages, not endpoint lists.

AC2 — every lane's main runtime endpoint, from config not per-call-site

RUNTIME_LANE_SPECS maps dev|stability-test|prod to DEV_RUNTIME_MAIN_PORT / STABILITY_TEST_RUNTIME_MAIN_PORT / PROD_RUNTIME_MAIN_PORT, resolved out of docker/runtime-policy.env by the same targeted-key-extraction idiom scripts/system_health_check.sh uses (deliberately not source, for the reason documented there). Prod is a plain GET /health and nothing else; test_prod_lane_is_probed_with_a_plain_get_only fails the build if any mutating verb ever appears in the file.

AC3 — RED on non-200 or a 200 whose body is not healthy

The old check was grep -Ei 'healthy|ok' against the response body, which matches {"healthy": false} — the substring is present. runtime_body_verdict() now resolves .healthy / .details.healthy / .status as real JSON and returns healthy / unhealthy / unresolvable; unhealthy and unresolvable are both CRITICAL.

AC4 — health=starting alarms

starting_past_start_period() reads each health=starting container's own Config.Healthcheck.StartPeriod and State.StartedAt and goes CRITICAL once it is past its grace (finite default when the image declares none). A container genuinely still inside its start_period does not page.

AC5 — fail closed; Exit-0 one-shots stay quiet

Unreachable (000) is CRITICAL, never omitted. A failed docker query is CRITICAL, not "nothing wrong". Exit accounting was previously reported but excluded from the CRITICAL condition: non-zero exits are now CRITICAL, Exited (0) migration/init one-shots are not.

AC6 — RED-before demonstrated, not asserted

tests/fixtures/omn15509/omninode-system-slack-report.as-deployed-20260730.sh.captured is the byte-for-byte copy of what was live on .201 (sha256 5fe6e5a6...c209da, asserted by test_as_deployed_fixture_is_the_unmodified_201_capture). Both artifacts are driven by the same harness through the same replayed 16:19-16:45Z state (dev 8085 -> 503 healthy=false, dev 8086 refused, stability 18085 -> 200, prod 28085 -> 200, omninode-runtime at health: starting past a 120s start_period, infra containers healthy):

  • test_as_deployed_reports_green_on_the_replayed_outage -> the pre-fix artifact reports Issues: *0 critical*, *0 warning*, - No active warning/critical checks, no dev port in the endpoint section.
  • test_fixed_reports_red_and_names_the_dev_runtime_on_the_same_state -> runtime-dev-8085: HTTP 503 (CRITICAL)`.

AC7 — the omission cannot silently reappear

test_every_lane_main_runtime_port_is_in_the_probe_set iterates the lanes and fails if any lane's main runtime endpoint is missing from the probe set. Mutation-verified: deleting the dev row from RUNTIME_LANE_SPECS turns 6 tests red (AC7 guard, the GREEN-after, the config-source assertion, and all three body/reachability cases) — the guard is load-bearing, not decorative.

dod_evidence

  • 14/14 tests/unit/scripts/test_omninode_system_slack_report.py pass on .200 (stickybeatz-studio), 13.49s.
  • ruff format --check, ruff check, mypy clean on .200.
  • pre-commit run --files <the 4 changed files> clean on .200 (SPDX + OMN-10741 no-env-fallbacks findings fixed in-branch, not bypassed).
  • shellcheck -S warning clean on the new script.
  • Patch-transfer parity: identical sha256 for all four files on this Mac and on .200 before the gate run (rule 11a edit-locality trap closed).
  • Live surfaces read read-only only. No container on .201 was restarted, recreated, or reconfigured; prod was GET-only.

Follow-up, stated not hidden

Landing this PR version-controls and fixes the script; it does not by itself replace the root-owned copy at .201:/data/maintenance/bin/. Deploying it requires a root write on .201, which is out of scope for this lane (the .201 dev-lane containers are owned by a concurrent lane, and this session's write allowlist excludes that path). Until the deployed copy is replaced, the live alert keeps its blind spot. Recommend a follow-up ticket for the deploy step plus an install/sync check so the on-host copy cannot drift from this file again.

Evidence-Ticket: OMN-15509
Evidence-Source: OCC#5634

…runtime endpoint

The .201 system-health Slack alert reported green off an endpoint list that
never included the dev lane runtime. On 2026-07-30 the dev runtime sat at
Docker health=starting with :8085 returning 503 for 26+ minutes while the
16:30:02Z alert showed every listed runtime endpoint as HTTP 200 (OK).

Producer (identified from code, Slack MCP was down):
/data/maintenance/bin/omninode-system-slack-report.sh on .201, run as root by
/etc/cron.d/omninode-system-slack-report. It was untracked by any repo, which
is why the omission was never reviewable. Version-controlled here as
deploy/maintenance/omninode-system-slack-report.sh with the cron unit.

Fixes: lane->port map read from docker/runtime-policy.env (dev/stability-test/
prod main runtime /health, GET-only, prod read-only); RED on non-200 OR a 200
whose health body is not healthy (substring grep matched {"healthy":false});
RED on health=starting past start_period; RED on non-zero exit while Exit(0)
one-shots stay quiet; fail-closed on unreachable endpoints and failed docker
queries.

RED-before/GREEN-after is demonstrated against the artifacts themselves:
tests/fixtures/omn15509/*.sh.captured is the byte-for-byte .201 capture and is
driven through the same replayed outage state, where it reports green.
@coderabbitai

coderabbitai Bot commented Jul 30, 2026

Copy link
Copy Markdown

Warning

Review limit reached

You’ve reached a temporary PR review limit under our Fair Usage Limits Policy.

Your recent review volume is higher than typical usage, so adaptive limits are currently applied.

Next review available in: 28 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 2f76591f-a79f-4081-a73a-ed2c380b68e8

📥 Commits

Reviewing files that changed from the base of the PR and between 5dc6819 and 89de342.

📒 Files selected for processing (4)
  • deploy/maintenance/cron.d/omninode-system-slack-report
  • deploy/maintenance/omninode-system-slack-report.sh
  • tests/fixtures/omn15509/omninode-system-slack-report.as-deployed-20260730.sh.captured
  • tests/unit/scripts/test_omninode_system_slack_report.py

Comment @coderabbitai help to get the list of available commands.

jonahgabriel added a commit to OmniNode-ai/onex_change_control that referenced this pull request Jul 30, 2026
#5634)

* evidence: OCC companion pass 1 for OmniNode-ai/omnibase_infra#2572

* evidence: OCC companion self-bind for #5634

---------

Co-authored-by: node-occ-companion-effect <occ-companion-effect@omninode.ai>
@github-actions

Copy link
Copy Markdown
Contributor

⚠️ Hostile Reviewer — DEGRADED (informational)

Blocking findings (critical): 0
Total findings: 0
Models succeeded: none

Note: All reviewer models failed or were unavailable. Degraded results are informational during the pilot phase (OMN-8468/OMN-8524) and do not block merge. Error: all review endpoints [192.168.86.201:8000 192.168.86.201:8001 ] unreachable — preflight short-circuit (no models available)


Gate semantics (pilot phase)

Verdict Meaning Blocks merge?
passed No critical findings No
blocked CRITICAL findings found Yes
degraded All models unavailable (infra) No (pilot)

Powered by omniintelligence.review_pairing.cli_review — node-based adversarial review via HandlerLlmCliSubprocess (OMN-8468/OMN-8524)

@jonahgabriel
jonahgabriel merged commit 8a3fc36 into dev Jul 30, 2026
154 of 162 checks passed
@jonahgabriel
jonahgabriel deleted the jonah/omn-15509-system-health-alert-lane-runtime-endpoints branch July 30, 2026 18:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant