Repository navigation
fix(OMN-15567): resolve nightly e2e runner-topology connectivity after SIGILL confound clears - #2623
Conversation
|
Warning Review limit reachedYou’ve reached a temporary PR review limit under our Fair Usage Limits Policy. Next review available in: 45 seconds Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (7)
Comment |
#5933) * evidence: OCC companion pass 1 for OmniNode-ai/omnibase_infra#2623 * evidence: OCC companion self-bind for #5933 --------- Co-authored-by: node-occ-companion-effect <occ-companion-effect@omninode.ai>
|
Fresh
Full AC1-4 write-up posted to OMN-15567. |
|
| Verdict | Meaning | Blocks merge? |
|---|---|---|
passed |
No critical findings | No |
blocked |
CRITICAL findings found | Yes |
degraded |
All models unavailable (infra) | No (pilot) |
Powered by omniintelligence.review_pairing.cli_review — node-based adversarial review via HandlerLlmCliSubprocess (OMN-8468/OMN-8524)
|
CI note (not caused by this diff): This is not plausibly caused by this PR's diff: the diff adds one new unit test file (11 cases, ~10s total locally and on .200) plus two one-line env-var reads in existing test files — nothing that would push an unrelated test bucket over a 15-minute budget. A concurrent, unrelated PR on this repo (omninode-infra OMN-15604, run 30735242910) completed its own 3-way split cleanly in the same time window, so this doesn't look like org-wide capacity exhaustion either — more likely this PR's specific split-2 bucket composition is unlucky under the no-cache fallback. Queued another |
|
Remediation round pushed (55ec3ab), all 6 flagged defects fixed and verified — see PR body. CI is currently RED on Confirmed this is live on This traces to the OMN-15628/OMN-15645 delegation-routing-tiers-path binding work (both currently "In Review") — Zombie note: the first CI attempt on this head also showed as |
a5465c8 to
e317e07
Compare
…r redpanda SIGILL confound clears The nightly e2e stack never reached pytest for 8/8 nights (redpanda exited 132/SIGILL). Re-observed post-OMN-15565 on workflow_dispatch run 30681782952 and scheduled run 30685774570: redpanda now starts Healthy in ~5s and pytest runs to completion (2471 passed/444 skipped/1 xfailed/6 failed/13 errors in 848.45s) -- the SIGILL was the OMN-15565 lane-collision confound, not a real defect, and is gone now that the stack is genuinely isolated. The 13 errors are the second defect this ticket targets: the self-hosted omnibase-ci runner does not share a network namespace with the Docker host (same topology reusable-runtime-boot.yml:250-270 already documents), so "localhost:<host-published-port>" is unreachable from the runner process even when the port is the freshly-derived, correctly-read dynamic port. Mirrors reusable-runtime-boot.yml's existing, proven detection/probe pattern: - "Detect runner network topology": identify a containerized runner and resolve the Docker host gateway, so redpanda's external listener can advertise it instead of always defaulting to "localhost". - "Resolve reachable e2e connectivity host": after the stack is Healthy, probe Docker DNS -> localhost -> Docker host gateway in that order and export KAFKA_BOOTSTRAP_SERVERS/OMNIBASE_INFRA_DB_URL/INTEGRATION_POSTGRES_HOST etc. from whichever actually works; fail closed with diagnostics if none do. Both new steps carry their own live-.201 guard, since they compute values the workflow-level static guard never sees. Also fixes two integration tests that hardcoded BOOTSTRAP_SERVERS = "localhost:19092", silently ignoring the CI-derived KAFKA_BOOTSTRAP_SERVERS regardless of topology -- matching the pattern already used correctly by test_runtime_health_monitor.py. OMN-15567
…y, Docker-DNS host cross-wiring, network-leak teardown Adversarial verifier found six defects in the prior push (head 4ef8477): 1. CI Summary FAILURE (required check on omnibase_infra/dev) -- Tests (Split 2/2) was cancelled at the 15-min job timeout 3x on this branch, misdiagnosed in the PR comment as an unlucky pytest-split bucket. Root cause: tests/integration/test_monitor_alert_emitter_integration.py is @pytest.mark.integration-only (missing @pytest.mark.kafka), so CI regular splits `-m "not slow and not chaos and not kafka and not performance"` never deselect it, unlike its sibling files (test_consumer_health_pipeline.py, test_runtime_log_bridge_pipeline.py, test_runtime_health_monitor.py) which already carry the kafka marker for exactly this reason. Added pytestmark = pytest.mark.kafka at module level, matching the established convention. 2. tests/unit/docker/test_nightly_e2e_runner_connectivity.py:171 was permanently RED on the mandated gate host (.200): nightly-integration.yml used `mapfile`, a bash 4+ builtin, and macOS ships bash 3.2.57 at /bin/bash (no newer bash for GPLv3 licensing reasons). Rewrote "Detect runner network topology" to use a `while read` loop instead -- bash-3-compatible, identical behavior. New regression test pins /bin/bash explicitly so a newer bash earlier on PATH cannot hide a future regression back to a bash-4-only construct. 3. PR body's claimed pre-push governed selector result ("1491 passed, 3 skipped") was not reproducible via raw non-login ssh (matches the verifier's invocation): 5 of the 6 failures are pre-existing PATH artifacts of a non-login shell lacking /opt/homebrew/bin (uv/timeout not on PATH) -- confirmed identical under origin/dev's own tree, not introduced by this diff. The 6th (mapfile) is fixed by #2 above and now passes under both a login-shell PATH and a PATH forced to exclude /opt/homebrew/bin (i.e. raw /bin/bash 3.2.57). 4/5. "Resolve reachable e2e connectivity host"'s Docker DNS branch (the ONLY branch live traffic actually took, per the run the PR cited as its own success evidence) cross-wired REDPANDA_ADVERTISE_HOST and OMNI_INFRA_HOST to the *postgres* container's Docker-DNS name instead of redpanda's -- the only branch where postgres and redpanda have DIFFERENT hostnames rather than one shared host disambiguated by port. Split resolved_host (postgres) from a new resolved_kafka_host (redpanda) and wired REDPANDA_ADVERTISE_HOST/OMNI_INFRA_HOST from the latter. New test drives the previously-untested Docker DNS branch via stubbed docker network connect/docker inspect. 6. "Resolve reachable e2e connectivity host" attaches the long-lived self-hosted runner container to the run's compose network via docker network connect and nothing ever disconnected it -- proven live: run 30733477609's teardown logged "Network ... Resource is still in use" without failing. Added a docker network disconnect in "Tear down e2e stack", before the destructive down -v, guarded to no-op when the runner was never attached. Two new tests cover both branches. Gates on .200 (patch-transfer, sha256-verified): targeted suite 33 passed; full tests/ci/+tests/unit/docker/ selector 1495 passed/3 skipped (login-shell PATH with /opt/homebrew/bin); ruff format/check clean; mypy clean except 2 pre-existing unused type:ignore comments at test_monitor_alert_emitter_integration.py (confirmed pre-existing via git stash comparison, unrelated to this diff); actionlint 60 pre-existing findings, zero new; pre-commit run --files clean. OMN-15567
…N_ROUTING_TIERS_PATH opt-out, revert false kafka marker, POSTGRES_PORT re-emit, observable network-disconnect verification Fixes 4 defects found by an independent adversarial verifier on head 4ef8477/55ec3ab: 1. BLOCKER: required CI Summary was FAILURE at PR head due to tests/integration/infra/test_judge_compose_render.py::test_judge_lane_delegation_routing_tiers_path_binding. Root cause confirmed dev-inherited (byte-identical compose/test files between HEAD and origin/dev, reproduces on a clean origin/dev checkout): OMN-15628 PR #2621 added DELEGATION_ROUTING_TIERS_PATH to docker-compose.judge.yml x-judge-runtime-env without the compensating opt-out that docker-compose.infra.yml carries on its own projection-api. Fixed by adding the same DELEGATION_ROUTING_TIERS_PATH: "" override to judge.yml projection-api, mirroring infra.yml. Filed nothing new for this -- OMN-15628 already tracks the area, cited in the compose comment. 2. tests/integration/test_monitor_alert_emitter_integration.py: reverted the pytestmark = pytest.mark.kafka added last round on a disproven claim (every MonitorAlertEmitter(...) construction in the file is wrapped in patch("confluent_kafka.Producer", ...), and the one unpatched construction deletes KAFKA_BOOTSTRAP_SERVERS first so the emitter self-disables before constructing anything -- 8 passed in 0.15s standalone, 8 passed under the CI regular-split marker filter). Live full-suite repro on .200 against a clean origin/dev checkout reproduced the same worker-death signature ("node down: Not properly terminated") on an unrelated Postgres concurrency test with zero Kafka involvement, disproving the Kafka-specific attribution. Filed OMN-15658 to track the real (still-unidentified) root cause; out of OMN-15567 scope. 3. .github/workflows/nightly-integration.yml: POSTGRES_PORT was set once to the ephemeral host-published port and never re-emitted in the Docker-DNS/gateway host-resolution rewrite, leaving it paired with the wrong host. Added the missing echo alongside the other resolved_* keys. 4. .github/workflows/nightly-integration.yml: the network-disconnect teardown fix from last round swallowed its own exit code silently (2>/dev/null || true), reproducing the exact silence class the fix existed to close. Added a post-disconnect docker network inspect verification that prints a non-fatal ::error:: annotation if the runner container is still attached, with 2 new stub-driven tests covering both the surfaced-error and verified-clean paths. [OMN-15567]
e317e07 to
2b4447c
Compare
OMN-15567: nightly e2e runner-topology connectivity (post-SIGILL-confound)
Fixes OMN-15567.
What this proves before touching code
The ticket's headline symptom —
omnibase-infra-redpanda exited (132)(SIGILL), 8/8 nights — is already gone as of two live post-OMN-15565 runs, before any change in this PR:workflow_dispatchrun 30681782952 (2026-08-01T03:20:01Z, headaa2231fc9a44b6738345f6457aa51824f0e8d0c1):Container ... redpanda Healthyat03:21:08.647Z, 5.5s afterStarting. pytest then ran to completion: 2471 passed / 444 skipped / 1 xfailed / 6 failed / 13 errors in 848.45s (0:14:08).Healthyagain in ~5.5s; pytest ran through 84%+ before hitting its own 300s per-test timeout on an unrelated ACL test.This confirms the ticket's own confound hypothesis: the SIGILL was OMN-15565's lane-collision defect (e2e redpanda recreated on the lab lane's data volume under incompatible startup flags), not a genuine CPU/instruction-set incompatibility. AC1/AC2 close on their own — no root-cause-132 write-up is needed because it did not recur on either post-fix run.
What this PR fixes (AC3/AC4)
The 13 errors on run 30681782952 are the ticket's second, independent suspicion:
reusable-runtime-boot.yml:250-270documents that the self-hostedomnibase-cirunner executes inside a container that does not share a network namespace with the Docker host, so host-published ports are unreachable atlocalhostfrom the runner process. Confirmed live — the runner's connection failed even when the port was the freshly-derived, correctly-read dynamic port:This was never a mismatched-port bug; the host string itself was unreachable.
nightly-integration.ymlnever detected or handled this topology;reusable-runtime-boot.yml(Tier-1/Tier-2 smoke) already does, so this PR ports that proven pattern rather than inventing a new one (net-negative-surface rule).Two new steps in
integration-tests:Spin up e2e stack) — samerunner_container_idheuristic and/proc/net/routedefault-gateway read asreusable-runtime-boot.yml. SetsE2E_REDPANDA_ADVERTISE_HOSTso redpanda's external listener (docker-compose.e2e.yml's pre-existing${E2E_REDPANDA_ADVERTISE_HOST:-localhost}in--advertise-kafka-addr/--advertise-pandaproxy-addr) advertises something reachable instead of always defaulting tolocalhost.Wait for health checks, beforeRun integration tests) — probes Docker DNS → localhost → Docker host gateway, in that order, and exportsKAFKA_BOOTSTRAP_SERVERS/OMNIBASE_INFRA_DB_URL/INTEGRATION_POSTGRES_HOST/INTEGRATION_POSTGRES_PORT/OMNI_INFRA_HOST/POSTGRES_HOST/REDPANDA_ADVERTISE_HOSTfrom whichever host actually works. Fails closed with diagnostics (exit 1+ printed probe state) if none are reachable — never lets pytest fail on opaque connection-refused noise.Both new steps carry their own live-
.201guard (E2E_REDPANDA_ADVERTISE_HOST/resolved host(s)== 192.168.86.201→exit 1), since they compute dynamic values the workflow's existing staticguard-live-hostjob never sees (that job only checks the top-levellocalhost-pinned env vars).Also fixed:
tests/integration/test_consumer_health_pipeline.pyandtests/integration/test_runtime_log_bridge_pipeline.pyhardcodedBOOTSTRAP_SERVERS = "localhost:19092", silently ignoring the workflow's run-scopedKAFKA_BOOTSTRAP_SERVERSregardless of topology — the exact same class of defect this ticket targets. Fixed toos.environ.get("KAFKA_BOOTSTRAP_SERVERS", "localhost:19092"), matching the patterntest_runtime_health_monitor.pyalready used correctly.Remediation round (2026-08-02) — six defects found by an independent adversarial verifier on head
4ef8477c, fixed on this same branchCI Summaryrequired check was FAILURE.Tests (Split 2/2)was cancelled at the 15-min job timeout 3x on this head. Root cause (not "flake"):tests/integration/test_monitor_alert_emitter_integration.pyconstructsMonitorAlertEmitter(which builds a realconfluent_kafka.Producer) withKAFKA_BOOTSTRAP_SERVERS=localhost:19092, and was@pytest.mark.integration-only — missing the@pytest.mark.kafkamarker its sibling files (test_consumer_health_pipeline.py,test_runtime_log_bridge_pipeline.py,test_runtime_health_monitor.py) already carry for exactly this reason. CI's regular split selector (-m "not slow and not chaos and not kafka and not performance") never deselected it. Fix: addedpytestmark = pytest.mark.kafkaat module level, matching the established convention. Confirmed live:-m "not slow and not chaos and not kafka and not performance"now deselects all 8 tests in the file (8 deselected);-m kafkaruns them standalone in 0.09s with no hang.test_nightly_e2e_runner_connectivity.py:171permanently RED on the mandated.200gate host. "Detect runner network topology" usedmapfile, a bash 4+ builtin; macOS ships bash 3.2.57 at/bin/bash(Apple stops at GPLv3). Fix: rewrote the array read as awhile IFS= read -rloop (bash-3-compatible, identical behavior), with an explicit${#arr[@]}guard before"${arr[@]}"expansion — bash <4.4 raises "unbound variable" on that expansion for an explicitly-declared-but-empty array underset -u, confirmed live on.200's real/bin/bash. New regression testtest_detect_topology_is_compatible_with_bin_bash_specificallypins/bin/bashexplicitly (skipped if absent) so a newer Homebrewbashearlier on PATH cannot hide a future regression back to a bash-4-only construct.sshinvocation to.200. Root cause: 5 of the reproduced 6 failures (test_integration_guard_pull_fatality.pyx2,test_prepush_hook_host_identity_guard.py,test_verify_pypi_pin_resolvability.pyx2) are PATH artifacts of a non-login shell that lacks/opt/homebrew/bin(timeout/uvnot resolvable) — confirmed identical onorigin/dev's own unmodified tree via the same non-login invocation, i.e. pre-existing and unrelated to this diff, not something this PR introduced or should paper over. The 6th (test_detect_topology_defaults_to_localhost_off_a_bare_runner, themapfiledefect) is this diff's fault and is fixed by Add Claude Code GitHub Workflow #2. Re-verified below under both PATH regimes.REDPANDA_ADVERTISE_HOST/OMNI_INFRA_HOSTto the postgres container.resolved_host="${OMNIBASE_INFRA_POSTGRES_CONTAINER}"was reused for every exported host var including the redpanda-specific ones — proven live on the run this PR cites as its own success evidence (30733477609):REDPANDA_ADVERTISE_HOSTandOMNI_INFRA_HOSTwere both the postgres container name. This is the only branch where postgres and redpanda have different hostnames (Docker DNS container names) rather than one shared host disambiguated by port. Fix: split intoresolved_host(postgres) and a newresolved_kafka_host(redpanda);REDPANDA_ADVERTISE_HOST/OMNI_INFRA_HOSTnow come fromresolved_kafka_host. Also extended the live-.201refusal check to coverresolved_kafka_host._BASE_RESOLVE_ENVcomment said no test droverunner_on_compose_network=true. Fix: new testtest_resolve_connectivity_uses_container_specific_hosts_for_docker_dnsstubsdocker network connect/docker inspectto force that branch and asserts each exported var uses the correct container's name — this is the test that would have caught defect feat: Complete infrastructure containers operational with Docker secrets #4.docker network connectto attach the long-lived self-hosted runner container to the run's network, and nothing ever disconnected it — proven live on run 30733477609's teardown log:Network omnibase-infra-e2e-30733477609-1--network Resource is still in use, non-fatal (silent). Fix: "Tear down e2e stack" now runsdocker network disconnectbefore the destructivedown -v, guarded to no-op when the runner was never attached (OMNIBASE_INFRA_RUNNER_CONTAINER_IDempty). Two new tests cover both branches (test_teardown_disconnects_runner_container_before_compose_down,test_teardown_skips_disconnect_when_runner_was_never_attached); confirmed the pre-existingtest_e2e_compose_lane_isolation.py::test_nightly_teardown_downs_exact_run_project_once(OMN-15565's teardown ratchet, unmodified by this PR) still passes unchanged — the disconnect is additive and only fires when a runner container id is actually set.Remediation round 2 (2026-08-02) — four defects found by an independent adversarial verifier on head
55ec3ab1, fixed on this same branch (commita5465c8f4d)BLOCKER — required
CI Summarywas FAILURE at PR head, not disclosed in this body.Tests (Split 2/3)failed ontests/integration/infra/test_judge_compose_render.py::test_judge_lane_delegation_routing_tiers_path_binding. Root cause confirmed dev-inherited, not introduced by this diff:docker/docker-compose.judge.yml,docker/docker-compose.infra.yml, and the test file are byte-identical between this PR's HEAD andorigin/dev(git diff origin/dev -- <those 3 files>is empty); running the same test on a cleanorigin/devcheckout (0d51fa72b) reproduces the identical failure.OMN-15628PR fix(OMN-15628): bind DELEGATION_ROUTING_TIERS_PATH in every runtime compose lane + bidirectional env parity gate #2621 addedDELEGATION_ROUTING_TIERS_PATHtodocker-compose.judge.yml'sx-judge-runtime-envanchor without the compensating opt-out thatdocker-compose.infra.yml's ownprojection-apicarries (DELEGATION_ROUTING_TIERS_PATH: "", added underOMN-15645). Fix: added the same opt-out to judge.yml'sprojection-api, mirroring infra.yml exactly. Verified the specific test now passes (1 passed in 0.48s), plus the fullx-runtime-env/env-parity/delegation-routing test surface (tests/ci/test_runtime_env_anchor.py,tests/ci/test_env_parity.py, all fourtests/integration/infra/test_*_compose_render.pyfiles,tests/unit/docker/test_delegation_routing_tiers_path_binding_omn15645.py,tests/unit/docker/test_runtime_entrypoint_delegation_tiers_self_heal.py) — 57 passed.codex/omn-15567-delegation-path-ci-fix, OPEN, created 2026-08-02T08:55Z — essentially simultaneous with this remediation round) independently fixes the identicaldocker-compose.judge.ymlline with the same content, and additionally adds 4 missing keys (ONEX_BOUNDARY_DLQ_ENABLED,ONEX_SINGLE_OWNER_COMMAND_TOPICS,ONEX_TOPIC_ENFORCEMENT_MODE,ONEX_WIRING_STRICT_MODE) todocker-compose.infra.yml'sx-runtime-envanchor that this PR'sEnv Parity (docker-compose vs k8s ConfigMap)job separately flagged as missing (cross-repo drift againstomninode_infra's onex-dev k8s manifests — unrelated to OMN-15567, not touched here to avoid a duplicate/conflicting diff against fix(OMN-15567): keep judge projection API routing opt-out #2628's in-flight fix). This PR does not attempt to merge, close, or race fix(OMN-15567): keep judge projection API routing opt-out #2628; whichever lands first, the other needs a trivial rebase (the judge.yml hunks are byte-identical). Codex owns merge sequencing.tests/integration/test_monitor_alert_emitter_integration.py: reverted the falsepytestmark = pytest.mark.kafkaadded in remediation round 1 on a disproven claim. EveryMonitorAlertEmitter(...)construction in the file is wrapped inpatch("confluent_kafka.Producer", ...); the one unpatched construction (test_emitter_disabled_when_kafka_env_missing) deletesKAFKA_BOOTSTRAP_SERVERSfirst so the emitter self-disables before_init_clients()ever callsconfluent_kafka.Producer(...). Verified:8 passed in 0.15sstandalone,8 passedunder CI's exact regular-split marker filter (-m "not slow and not chaos and not kafka and not performance"). What actually caused the 15-min Split-2/2 timeout is still unknown. A live full-suite repro on.200against a cleanorigin/devcheckout, using the identical selector flags, reproduced the same[gwN] node down: Not properly terminatedworker-death signature at ~13% progress ontests/integration/services/snapshot/test_store_postgres_integration.py::TestConcurrentSequenceGeneration::test_high_concurrency_sequence_uniqueness— a Postgres concurrency test with zero Kafka involvement — which rules out this file (and Kafka generally) as the mechanism. Filed OMN-15658 to track the real root cause; out of this ticket's scope (nightly redpanda SIGILL / runner topology, not the regular CITests (Split N)job's full-suite reliability). Also ruled out and documented in OMN-15658 so a future investigator doesn't re-walk the dead end:tests/integration/verification/test_registration_contract_verify.py(skips before reaching its unmockedconfluent_kafkaconstruction, becauseTests (Split N)has no live Postgres service and itsdb_query_fnkwarg is evaluated — and skips — first) andtests/integration/runtime/test_steel_dispatch_golden_chain_live_runtime.py(gated behind a live-reachabilityskipifagainst an unreachable.201host)..github/workflows/nightly-integration.yml:POSTGRES_PORTwas never re-emitted in "Resolve reachable e2e connectivity host." It was set once, in the earlier "Derive isolated e2e namespace" step, to the ephemeral host-published port, and stayed paired with that original value even afterPOSTGRES_HOSTwas overridden toresolved_host— an invalid host/port pair under the Docker-DNS and gateway branches (onlyPOSTGRES_HOSTwas corrected;INTEGRATION_POSTGRES_PORTwas,POSTGRES_PORTwas not). Confirmed live on run 30733477609 (this PR's own cited success evidence):POSTGRES_HOST=omnibase-infra-e2e-30733477609-1--postgresalongsidePOSTGRES_PORT=37919(the original ephemeral host port) — invalid under every topology. Latent that run only because its consumers (tests/integration/test_dispatch_roundtrip.py,tests/integration/injection_effectiveness/conftest.py,tests/integration/migrations/test_node_migration_shape_drift_omn15376.py) were separately gated out. Fix: addedecho "POSTGRES_PORT=${resolved_pg_port}"alongside the otherresolved_*exports..github/workflows/nightly-integration.yml: the network-disconnect teardown fix from remediation round 1 was unobservable in the direction it matters.docker network disconnect ... 2>/dev/null || trueswallows every failure, so if the disconnect doesn't take, the run leaks the attachment exactly as before and the log stays silent — the same silent-failure shape the fix existed to remove. Fix: added a post-disconnectdocker network inspect ... --format '{{json .Containers}}' | grep -q "$OMNIBASE_INFRA_RUNNER_CONTAINER_ID"check that prints a non-fatal::error::annotation naming the network and container if the runner is still attached — the teardown step still completes and still runsdocker compose downeither way, but a real failure to detach is now visible in the log instead of silent. Two new stub-driven tests:test_teardown_surfaces_error_when_disconnect_does_not_take(still-attached path prints::error::) andtest_teardown_silent_when_disconnect_verified_gone(verified-clean path prints nothing) — the existingtest_teardown_disconnects_runner_container_before_compose_down/test_teardown_skips_disconnect_when_runner_was_never_attachedstill pass unmodified.No live
workflow_dispatchrun exists yet reflecting this remediation round's changes (the mapfile→while-read rewrite, the resolved_host/resolved_kafka_host split, and the network-disconnect teardown fix from round 1 were also never exercised live before this round — the onlyworkflow_dispatchrun on this branch, 30733477609, predates all of round 1). A fresh dispatch after this push is still owed before AC1/AC3/AC4 can be called live-proven rather than unit-proven; see the ticket comment for the live run link once available.Gates (remediation round 2)
gates_host:stickybeatz-studio(.200), patch-transfer + sha256-verified per rule 11a/the standing invariant.ruff format --check/ruff check: clean on both touched Python files.mypy: clean ontest_nightly_e2e_runner_connectivity.py;test_monitor_alert_emitter_integration.pyhas the same 2 pre-existingUnused "type: ignore" commentfindings confirmed present onorigin/dev's own unmodified copy of the file (line 256/257, verified via a direct mypy run againstgit show origin/dev:...), unchanged by this diff.tests/integration/test_monitor_alert_emitter_integration.py(8 passed standalone, 8 passed under the CI marker filter),tests/unit/docker/test_nightly_e2e_runner_connectivity.py(17 passed, up from 15 — 2 net new),tests/unit/docker/test_e2e_compose_lane_isolation.py+ connectivity file together (35 passed), the judge-compose + env-parity + delegation-routing surface listed in item 1 above (57 passed),tests/integration/infra/+tests/unit/infra/+tests/unit/docker/(369 passed, 3 skipped, skips are pre-existing Docker-daemon-required).pre-commit run --all-files: clean on all 4 changed files (both hooks that failed repo-wide —URL Authority Gateon 5 unrelated files,ONEX SPDX Header Requirementon 2 unrelated files with a 2025-vs-2026 year mismatch — are pre-existing repo debt in files this diff never touches, confirmed by grep against the changed-file list); the actualgit commitpre-commit run (scoped to the 4 changed files) was 100% Passed/Skipped, zero failures.prepush-smart-tests, OMN-13973), the realgit pushpre-push hook that gated this commit:tests/ci/ tests/unit/docker/— 1497 passed, 3 skipped (same 3 pre-existing Docker-daemon skips; +2 net vs. round 1's 1495 from the 2 net-new teardown-observability tests).docker/docker-compose.judge.ymlchange touches a deploy-gate "runtime path," andEvidence-Source: OCC#5933(this PR body's original citation) resolved to a stale OCC contract snapshot (commit05c85ec8) that predates the falsifiabledod-infra-pr-2623-535f8a74evidence item added by the later, already-merged OCC companion PR #5939 (commit3b30ef918, confirmed viagh api repos/OmniNode-ai/onex_change_control/contents/contracts/OMN-15567.yaml?ref=<sha>at both refs). Updated the citation below toOCC#5939so deploy-gate resolves a contract snapshot that already declares a falsifiablegh api ...probe.Seam definition (every field/key this PR reads or writes across a boundary)
$GITHUB_ENVwithin theintegration-testsjob (GitHub Actions' cross-step boundary — each>>write is visible to every later step in the job):OMNIBASE_INFRA_RUNNER_CONTAINER_IDE2E_REDPANDA_ADVERTISE_HOST"localhost")docker-compose.e2e.ymlredpanda service (--advertise-kafka-addr/--advertise-pandaproxy-addr), "Resolve reachable e2e connectivity host"KAFKA_BOOTSTRAP_SERVERShost:portlocalhost:$KAFKA_PORT) → overridden by "Resolve reachable e2e connectivity host" (resolved_kafka_host:port)Run integration tests(pytest env),test_consumer_health_pipeline.py/test_runtime_log_bridge_pipeline.py/test_runtime_health_monitor.pyOMNIBASE_INFRA_DB_URLresolved_host:pg_port)OMNIBASE_INFRA_DB_URLINTEGRATION_POSTGRES_HOST/INTEGRATION_POSTGRES_PORT/POSTGRES_HOSTresolved_host(postgres-specific host)OMNI_INFRA_HOST/REDPANDA_ADVERTISE_HOSTresolved_kafka_host(redpanda-specific host — NEW in the remediation round, previously incorrectly sourced fromresolved_host/postgres)test_runtime_consumes_build_loop_terminal_event.pybuilds its own DSN from host+port)docker-compose.e2e.ymlboundary (pre-existing key, previously always defaulted):${E2E_REDPANDA_ADVERTISE_HOST:-localhost}— now actually populated by "Detect runner network topology" instead of silently defaulting every run.DOCKER_CALL_LOG-adjacent (test-only, not a runtime boundary): the new teardown tests capturedockerCLI invocations to a log file for ordering assertions, mirroringtest_e2e_compose_lane_isolation.py's pre-existing_run_teardown_with_stubspattern rather than inventing a new stub mechanism.Acceptance criteria → proof
workflow_dispatchrun after OMN-15565, freshSpin up e2e stacklog2471 passed, 444 skipped, 1 xfailed, 6 failed, 13 errors in 848.45s.test_resolve_connectivity_runs_between_health_checks_and_pytestlocks the step ordering that makes this possible.test_detect_topology_defaults_to_localhost_off_a_bare_runner/test_detect_topology_is_compatible_with_bin_bash_specifically(negative default, bash-3-pinned),test_resolve_connectivity_fails_closed_when_nothing_is_reachable(fail-closed),test_resolve_connectivity_prefers_localhost_when_reachable,test_resolve_connectivity_falls_back_to_docker_host_gateway(the exact confirmed-live failure mode),test_resolve_connectivity_uses_container_specific_hosts_for_docker_dns(the exact confirmed-live Docker-DNS branch, remediation round),test_resolve_connectivity_refuses_to_resolve_to_live_201_host(safety),test_teardown_disconnects_runner_container_before_compose_down/test_teardown_skips_disconnect_when_runner_was_never_attached(network-leak fix), plustest_kafka_test_files_read_bootstrap_servers_from_env/test_python_can_import_the_fixed_kafka_test_modulesfor the two hardcoded test files.Tests
tests/unit/docker/test_nightly_e2e_runner_connectivity.py: 16 cases (11 original + 5 added in the remediation round:test_detect_topology_is_compatible_with_bin_bash_specifically,test_resolve_connectivity_uses_container_specific_hosts_for_docker_dns,test_teardown_disconnects_runner_container_before_compose_down,test_teardown_skips_disconnect_when_runner_was_never_attached, plus the parametrized kafka-file case count is unchanged). RED-before/GREEN-after against the realrun:shell of each step via subprocess + stubdocker/python3/uvbinaries — mirrorstest_e2e_compose_lane_isolation.py's existing pattern, not a structural grep.tests/integration/test_monitor_alert_emitter_integration.py: no new test cases (marker-only fix); all 8 existing cases still pass standalone (-m kafka, 0.09s) and are now correctly deselected from the CI regular split (-m "not kafka",8 deselected).Unrelated pre-existing ratchet suite (
test_e2e_compose_lane_isolation.py, 18 cases, OMN-15565) still passes unmodified — confirms the teardown network-disconnect addition is additive, not a regression to that file's project/manifest-guard assertions.Gates
gates_host:stickybeatz-studio(.200), patch-transfer + sha256-verified per rule 11a/the standing invariant, for both the original push and this remediation round.ssh stickybeatz-studio 'zsh -lc "..."'(login shell, canonical.200PATH including/opt/homebrew/binwhereuv/timeout/GNU coreutils/Homebrew bash actually live). A raw non-loginssh stickybeatz-studio 'cmd'(no-l) lacks that PATH and produces 5 spurious "command not found" failures unrelated to code correctness — reproduced and confirmed identical againstorigin/dev's own unmodified tree.ruff format --check/ruff check: clean on all 3 touched Python/YAML-adjacent files (python files only — the workflow YAML is not a ruff target; confirmed by the tool erroring trying to parse YAML as Python when mistakenly included).mypy: clean onnightly-integration.yml's Python test surface andtest_consumer_health_pipeline.py/test_runtime_log_bridge_pipeline.py(from the original push);test_monitor_alert_emitter_integration.py(untouched by this PR except the marker line) has 2 pre-existingUnused "type: ignore" commentfindings at lines that shifted with my diff (256→270, 257→271) — confirmed pre-existing viagit stash/mypy-before comparison, not introduced by this PR.actionlint: 60 shellcheck findings both before and after this remediation round's edit (git stashcomparison) — zero new findings introduced by themapfile→while readrewrite or the disconnect/host-split changes.pre-commit run --files <3 changed files>: all hooks Passed or Skipped (no files to check), zero failures; sha256 of all 3 files unchanged across the pre-commit run (no hook silently mutated content).tests/unit/docker/test_nightly_e2e_runner_connectivity.py+test_e2e_compose_lane_isolation.py): 33 passed in 8.20s.test_nightly_e2e_runner_connectivity.pyalone, PATH forced to exclude/opt/homebrew/bin(i.e.bashresolves to the real/bin/bash3.2.57, matching the verifier's exact reproduction environment): 15 passed in 7.65s — proves themapfilefix holds under the harsher PATH too, not just the login-shell one.test_monitor_alert_emitter_integration.pyunder-m "not slow and not chaos and not kafka and not performance"(CI's regular split selection): 8 deselected. Same file under-m kafka: 8 passed in 0.09s (no hang).prepush-smart-tests, OMN-13973), run twice — once ad hoc on.200(login shell) and once as the actualgit pushpre-push hook that gated this remediation commit's push:tests/ci/ tests/unit/docker/— 1495 passed, 3 skipped (same 3 pre-existing Docker-daemon-required skips as before; +4 net vs. the originally-claimed 1491 from the 4 net-new tests, after subtracting the 1 fixedmapfilefailure that is no longer a failure).Non-closure note
Per the ticket: not closing/muting the schedule here. Once this branch's fresh dispatch run and the AC1-4 write-up land in the ticket, the nightly's stability going forward is an operator call, not this PR's.
Live post-fix
workflow_dispatchconfirmation (run 30741864371)The first run reflecting every remediation-round-2 fix, dispatched after the push above:
exited (132)/ SIGILL — AC1/AC2 stay closed.26 failed, 2678 passed, 237 skipped, 1 xfailed, 25 warnings, 49 errors in 1057.90s (0:17:37)— AC3 (an executed suite, not a green one, is what's required).POSTGRES_HOST/OMNI_INFRA_HOST/REDPANDA_ADVERTISE_HOSTresolved to Docker-DNS container names;POSTGRES_PORTcorrectly moved from the ephemeral47407to the resolved5432at "Run integration tests" (item 3's fix, live-confirmed for the first time) — AC4.Network omnibase-infra-e2e-30741864371-1--network Removedcleanly, no "Resource is still in use," no::error::annotation fired (verified-clean disconnect, not just assumed) — network-leak fix (round 1) and its observability check (round 2, item 4) both live-confirmed on the same run.Full write-up posted to the OMN-15567 ticket.
Evidence-Source: OCC#5939
Evidence-Ticket: OMN-15567