Repository navigation
fix(OMN-15509): system-health Slack alert must probe every lane's main runtime endpoint - #2572
Conversation
…runtime endpoint
The .201 system-health Slack alert reported green off an endpoint list that
never included the dev lane runtime. On 2026-07-30 the dev runtime sat at
Docker health=starting with :8085 returning 503 for 26+ minutes while the
16:30:02Z alert showed every listed runtime endpoint as HTTP 200 (OK).
Producer (identified from code, Slack MCP was down):
/data/maintenance/bin/omninode-system-slack-report.sh on .201, run as root by
/etc/cron.d/omninode-system-slack-report. It was untracked by any repo, which
is why the omission was never reviewable. Version-controlled here as
deploy/maintenance/omninode-system-slack-report.sh with the cron unit.
Fixes: lane->port map read from docker/runtime-policy.env (dev/stability-test/
prod main runtime /health, GET-only, prod read-only); RED on non-200 OR a 200
whose health body is not healthy (substring grep matched {"healthy":false});
RED on health=starting past start_period; RED on non-zero exit while Exit(0)
one-shots stay quiet; fail-closed on unreachable endpoints and failed docker
queries.
RED-before/GREEN-after is demonstrated against the artifacts themselves:
tests/fixtures/omn15509/*.sh.captured is the byte-for-byte .201 capture and is
driven through the same replayed outage state, where it reports green.
|
Warning Review limit reachedYou’ve reached a temporary PR review limit under our Fair Usage Limits Policy. Next review available in: 28 minutes Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (4)
Comment |
#5634) * evidence: OCC companion pass 1 for OmniNode-ai/omnibase_infra#2572 * evidence: OCC companion self-bind for #5634 --------- Co-authored-by: node-occ-companion-effect <occ-companion-effect@omninode.ai>
|
| Verdict | Meaning | Blocks merge? |
|---|---|---|
passed |
No critical findings | No |
blocked |
CRITICAL findings found | Yes |
degraded |
All models unavailable (infra) | No (pilot) |
Powered by omniintelligence.review_pairing.cli_review — node-based adversarial review via HandlerLlmCliSubprocess (OMN-8468/OMN-8524)
Closes OMN-15509.
AC1 — the producing alert code, named, with the evidence that identifies it
Producer:
/data/maintenance/bin/omninode-system-slack-report.shon.201(omninode-pc), executed asrootby/etc/cron.d/omninode-system-slack-report:Confirmed live 2026-07-30T17:00:01-04:00 in
/var/log/syslog:CRON[3258383]: (root) CMD (/data/maintenance/bin/omninode-system-slack-report.sh --mode alert).The script was untracked by every repository — that is why the defect survived: there was no diff to review, no test to fail, and no grep in any repo that could find it.
grep -rn "OmniNode system alert"across the whole workspace returns nothing.Message matched to the emitting call site. The Slack MCP was down, so the message was read back read-only through the bot token already provisioned on
.201(conversations.history). The 2026-07-30T16:30:02Z message — inside the outage window — is verbatim:That endpoint list is produced character-for-character by
collect()+format_digest()in the as-deployed script (tests/fixtures/omn15509/omninode-system-slack-report.as-deployed-20260730.sh.captured:129-133,:151). Noruntime-8085. The*OmniNode morning system digest*and*[OmniNode alert resolved]*strings in the same channel come from the same file'scase "$MODE"block.Candidate #1 from the ticket is REFUTED, not merely unconfirmed
omnimarket/src/omnimarket/nodes/node_baseline_capture/handlers/probes/probe_system_health.pyis not the producer:node_baseline_capture, whose only publish topic isonex.evt.omnimarket.baseline-captured.v1. That topic is recorded asORPHANED_PRODUCERand the node asDISCONNECTED_SUBGRAPHin the frozen contract graph (omnimarket/src/omnimarket/validators/data/contract_topic_graph_baseline.yaml:14,345);validators/contract_topic_graph.py:187names it as one of "2 genuinely dead nodes". No Slack module imports it.onex.cmd.omnimarket.slack-publish.v1andonex.evt.omnimarket.slack-published.v1are at HW=0 on all three lanes (dev/stability-test/prod), as isonex.cmd.omniclaw.slack-outbound.v1. Nothing reached Slack over the bus..201carries anySLACK_*env var (checked across all 105 running containers viadocker inspect), so no containerized service is a Slack producer at all.Also ruled out along the way:
scripts/system_health_check.shhas zero Slack references and itsdev-lane-liveness.ymlcaller was correctly RED all day (14 consecutivefailureruns, 06:02Z-16:54Z);monitor_logs.pyis not running (onex-log-monitor.service= inactive/not-found);runner-monitor.shandrunner_fleet_canary.shemit runner-fleet messages, not endpoint lists.AC2 — every lane's main runtime endpoint, from config not per-call-site
RUNTIME_LANE_SPECSmapsdev|stability-test|prodtoDEV_RUNTIME_MAIN_PORT/STABILITY_TEST_RUNTIME_MAIN_PORT/PROD_RUNTIME_MAIN_PORT, resolved out ofdocker/runtime-policy.envby the same targeted-key-extraction idiomscripts/system_health_check.shuses (deliberately notsource, for the reason documented there). Prod is a plainGET /healthand nothing else;test_prod_lane_is_probed_with_a_plain_get_onlyfails the build if any mutating verb ever appears in the file.AC3 — RED on non-200 or a 200 whose body is not healthy
The old check was
grep -Ei 'healthy|ok'against the response body, which matches{"healthy": false}— the substring is present.runtime_body_verdict()now resolves.healthy/.details.healthy/.statusas real JSON and returns healthy / unhealthy / unresolvable; unhealthy and unresolvable are both CRITICAL.AC4 —
health=startingalarmsstarting_past_start_period()reads eachhealth=startingcontainer's ownConfig.Healthcheck.StartPeriodandState.StartedAtand goes CRITICAL once it is past its grace (finite default when the image declares none). A container genuinely still inside itsstart_perioddoes not page.AC5 — fail closed; Exit-0 one-shots stay quiet
Unreachable (
000) is CRITICAL, never omitted. A faileddockerquery is CRITICAL, not "nothing wrong". Exit accounting was previously reported but excluded from the CRITICAL condition: non-zero exits are now CRITICAL,Exited (0)migration/init one-shots are not.AC6 — RED-before demonstrated, not asserted
tests/fixtures/omn15509/omninode-system-slack-report.as-deployed-20260730.sh.capturedis the byte-for-byte copy of what was live on.201(sha2565fe6e5a6...c209da, asserted bytest_as_deployed_fixture_is_the_unmodified_201_capture). Both artifacts are driven by the same harness through the same replayed 16:19-16:45Z state (dev 8085 -> 503healthy=false, dev 8086 refused, stability 18085 -> 200, prod 28085 -> 200,omninode-runtimeathealth: startingpast a 120s start_period, infra containers healthy):test_as_deployed_reports_green_on_the_replayed_outage-> the pre-fix artifact reportsIssues: *0 critical*, *0 warning*,- No active warning/critical checks, no dev port in the endpoint section.test_fixed_reports_red_and_names_the_dev_runtime_on_the_same_state->runtime-dev-8085: HTTP 503 (CRITICAL)`.AC7 — the omission cannot silently reappear
test_every_lane_main_runtime_port_is_in_the_probe_setiterates the lanes and fails if any lane's main runtime endpoint is missing from the probe set. Mutation-verified: deleting thedevrow fromRUNTIME_LANE_SPECSturns 6 tests red (AC7 guard, the GREEN-after, the config-source assertion, and all three body/reachability cases) — the guard is load-bearing, not decorative.dod_evidence
tests/unit/scripts/test_omninode_system_slack_report.pypass on.200(stickybeatz-studio), 13.49s.ruff format --check,ruff check,mypyclean on.200.pre-commit run --files <the 4 changed files>clean on.200(SPDX + OMN-10741no-env-fallbacksfindings fixed in-branch, not bypassed).shellcheck -S warningclean on the new script..200before the gate run (rule 11a edit-locality trap closed)..201was restarted, recreated, or reconfigured; prod was GET-only.Follow-up, stated not hidden
Landing this PR version-controls and fixes the script; it does not by itself replace the root-owned copy at
.201:/data/maintenance/bin/. Deploying it requires a root write on.201, which is out of scope for this lane (the.201dev-lane containers are owned by a concurrent lane, and this session's write allowlist excludes that path). Until the deployed copy is replaced, the live alert keeps its blind spot. Recommend a follow-up ticket for the deploy step plus an install/sync check so the on-host copy cannot drift from this file again.Evidence-Ticket: OMN-15509
Evidence-Source: OCC#5634