Skip to content

fix(OMN-15567): resolve nightly e2e runner-topology connectivity after SIGILL confound clears - #2623

Merged
jonahgabriel merged 9 commits into
devfrom
jonah/omn-15567-nightly-redpanda-sigill
Aug 2, 2026
Merged

jonahgabriel merged 9 commits into
devfrom
jonah/omn-15567-nightly-redpanda-sigill

Conversation

@jonahgabriel

@jonahgabriel jonahgabriel commented Aug 2, 2026 •

Copy link
Copy Markdown
Collaborator

OMN-15567: nightly e2e runner-topology connectivity (post-SIGILL-confound)

Fixes OMN-15567.

What this proves before touching code

The ticket's headline symptom — omnibase-infra-redpanda exited (132) (SIGILL), 8/8 nights — is already gone as of two live post-OMN-15565 runs, before any change in this PR:

  • workflow_dispatch run 30681782952 (2026-08-01T03:20:01Z, head aa2231fc9a44b6738345f6457aa51824f0e8d0c1): Container ... redpanda Healthy at 03:21:08.647Z, 5.5s after Starting. pytest then ran to completion: 2471 passed / 444 skipped / 1 xfailed / 6 failed / 13 errors in 848.45s (0:14:08).
  • Scheduled run 30685774570 (2026-08-01T05:24:37Z, same head): redpanda Healthy again in ~5.5s; pytest ran through 84%+ before hitting its own 300s per-test timeout on an unrelated ACL test.

This confirms the ticket's own confound hypothesis: the SIGILL was OMN-15565's lane-collision defect (e2e redpanda recreated on the lab lane's data volume under incompatible startup flags), not a genuine CPU/instruction-set incompatibility. AC1/AC2 close on their own — no root-cause-132 write-up is needed because it did not recur on either post-fix run.

What this PR fixes (AC3/AC4)

The 13 errors on run 30681782952 are the ticket's second, independent suspicion: reusable-runtime-boot.yml:250-270 documents that the self-hosted omnibase-ci runner executes inside a container that does not share a network namespace with the Docker host, so host-published ports are unreachable at localhost from the runner process. Confirmed live — the runner's connection failed even when the port was the freshly-derived, correctly-read dynamic port:

aiokafka.errors.KafkaConnectionError: Unable to bootstrap from [('localhost', 40335, ...)]
OSError: Multiple exceptions: [Errno 111] Connect call failed ('::1', 45937, 0, 0), [Errno 111] Connect call failed ('127.0.0.1', 45937)

This was never a mismatched-port bug; the host string itself was unreachable. nightly-integration.yml never detected or handled this topology; reusable-runtime-boot.yml (Tier-1/Tier-2 smoke) already does, so this PR ports that proven pattern rather than inventing a new one (net-negative-surface rule).

Two new steps in integration-tests:

  1. "Detect runner network topology" (before Spin up e2e stack) — same runner_container_id heuristic and /proc/net/route default-gateway read as reusable-runtime-boot.yml. Sets E2E_REDPANDA_ADVERTISE_HOST so redpanda's external listener (docker-compose.e2e.yml's pre-existing ${E2E_REDPANDA_ADVERTISE_HOST:-localhost} in --advertise-kafka-addr/--advertise-pandaproxy-addr) advertises something reachable instead of always defaulting to localhost.
  2. "Resolve reachable e2e connectivity host" (after Wait for health checks, before Run integration tests) — probes Docker DNS → localhost → Docker host gateway, in that order, and exports KAFKA_BOOTSTRAP_SERVERS / OMNIBASE_INFRA_DB_URL / INTEGRATION_POSTGRES_HOST / INTEGRATION_POSTGRES_PORT / OMNI_INFRA_HOST / POSTGRES_HOST / REDPANDA_ADVERTISE_HOST from whichever host actually works. Fails closed with diagnostics (exit 1 + printed probe state) if none are reachable — never lets pytest fail on opaque connection-refused noise.

Both new steps carry their own live-.201 guard (E2E_REDPANDA_ADVERTISE_HOST/resolved host(s) == 192.168.86.201 → exit 1), since they compute dynamic values the workflow's existing static guard-live-host job never sees (that job only checks the top-level localhost-pinned env vars).

Also fixed: tests/integration/test_consumer_health_pipeline.py and tests/integration/test_runtime_log_bridge_pipeline.py hardcoded BOOTSTRAP_SERVERS = "localhost:19092", silently ignoring the workflow's run-scoped KAFKA_BOOTSTRAP_SERVERS regardless of topology — the exact same class of defect this ticket targets. Fixed to os.environ.get("KAFKA_BOOTSTRAP_SERVERS", "localhost:19092"), matching the pattern test_runtime_health_monitor.py already used correctly.

Remediation round (2026-08-02) — six defects found by an independent adversarial verifier on head 4ef8477c, fixed on this same branch

  1. BLOCKER — CI Summary required check was FAILURE. Tests (Split 2/2) was cancelled at the 15-min job timeout 3x on this head. Root cause (not "flake"): tests/integration/test_monitor_alert_emitter_integration.py constructs MonitorAlertEmitter (which builds a real confluent_kafka.Producer) with KAFKA_BOOTSTRAP_SERVERS=localhost:19092, and was @pytest.mark.integration-only — missing the @pytest.mark.kafka marker its sibling files (test_consumer_health_pipeline.py, test_runtime_log_bridge_pipeline.py, test_runtime_health_monitor.py) already carry for exactly this reason. CI's regular split selector (-m "not slow and not chaos and not kafka and not performance") never deselected it. Fix: added pytestmark = pytest.mark.kafka at module level, matching the established convention. Confirmed live: -m "not slow and not chaos and not kafka and not performance" now deselects all 8 tests in the file (8 deselected); -m kafka runs them standalone in 0.09s with no hang.
  2. test_nightly_e2e_runner_connectivity.py:171 permanently RED on the mandated .200 gate host. "Detect runner network topology" used mapfile, a bash 4+ builtin; macOS ships bash 3.2.57 at /bin/bash (Apple stops at GPLv3). Fix: rewrote the array read as a while IFS= read -r loop (bash-3-compatible, identical behavior), with an explicit ${#arr[@]} guard before "${arr[@]}" expansion — bash <4.4 raises "unbound variable" on that expansion for an explicitly-declared-but-empty array under set -u, confirmed live on .200's real /bin/bash. New regression test test_detect_topology_is_compatible_with_bin_bash_specifically pins /bin/bash explicitly (skipped if absent) so a newer Homebrew bash earlier on PATH cannot hide a future regression back to a bash-4-only construct.
  3. FALSE gate evidence in this PR's own body. The prior claim ("1491 passed, 3 skipped") did not reproduce via the verifier's raw non-login ssh invocation to .200. Root cause: 5 of the reproduced 6 failures (test_integration_guard_pull_fatality.py x2, test_prepush_hook_host_identity_guard.py, test_verify_pypi_pin_resolvability.py x2) are PATH artifacts of a non-login shell that lacks /opt/homebrew/bin (timeout/uv not resolvable) — confirmed identical on origin/dev's own unmodified tree via the same non-login invocation, i.e. pre-existing and unrelated to this diff, not something this PR introduced or should paper over. The 6th (test_detect_topology_defaults_to_localhost_off_a_bare_runner, the mapfile defect) is this diff's fault and is fixed by Add Claude Code GitHub Workflow #2. Re-verified below under both PATH regimes.
  4. Docker-DNS branch cross-wired REDPANDA_ADVERTISE_HOST/OMNI_INFRA_HOST to the postgres container. resolved_host="${OMNIBASE_INFRA_POSTGRES_CONTAINER}" was reused for every exported host var including the redpanda-specific ones — proven live on the run this PR cites as its own success evidence (30733477609): REDPANDA_ADVERTISE_HOST and OMNI_INFRA_HOST were both the postgres container name. This is the only branch where postgres and redpanda have different hostnames (Docker DNS container names) rather than one shared host disambiguated by port. Fix: split into resolved_host (postgres) and a new resolved_kafka_host (redpanda); REDPANDA_ADVERTISE_HOST/OMNI_INFRA_HOST now come from resolved_kafka_host. Also extended the live-.201 refusal check to cover resolved_kafka_host.
  5. Coverage hole: the Docker-DNS branch — the only branch live traffic actually took — had zero test coverage. The file's own _BASE_RESOLVE_ENV comment said no test drove runner_on_compose_network=true. Fix: new test test_resolve_connectivity_uses_container_specific_hosts_for_docker_dns stubs docker network connect/docker inspect to force that branch and asserts each exported var uses the correct container's name — this is the test that would have caught defect feat: Complete infrastructure containers operational with Docker secrets #4.
  6. Leaked Docker network attachment on every nightly run. "Resolve reachable e2e connectivity host" runs docker network connect to attach the long-lived self-hosted runner container to the run's network, and nothing ever disconnected it — proven live on run 30733477609's teardown log: Network omnibase-infra-e2e-30733477609-1--network Resource is still in use, non-fatal (silent). Fix: "Tear down e2e stack" now runs docker network disconnect before the destructive down -v, guarded to no-op when the runner was never attached (OMNIBASE_INFRA_RUNNER_CONTAINER_ID empty). Two new tests cover both branches (test_teardown_disconnects_runner_container_before_compose_down, test_teardown_skips_disconnect_when_runner_was_never_attached); confirmed the pre-existing test_e2e_compose_lane_isolation.py::test_nightly_teardown_downs_exact_run_project_once (OMN-15565's teardown ratchet, unmodified by this PR) still passes unchanged — the disconnect is additive and only fires when a runner container id is actually set.

Remediation round 2 (2026-08-02) — four defects found by an independent adversarial verifier on head 55ec3ab1, fixed on this same branch (commit a5465c8f4d)

  1. BLOCKER — required CI Summary was FAILURE at PR head, not disclosed in this body. Tests (Split 2/3) failed on tests/integration/infra/test_judge_compose_render.py::test_judge_lane_delegation_routing_tiers_path_binding. Root cause confirmed dev-inherited, not introduced by this diff: docker/docker-compose.judge.yml, docker/docker-compose.infra.yml, and the test file are byte-identical between this PR's HEAD and origin/dev (git diff origin/dev -- <those 3 files> is empty); running the same test on a clean origin/dev checkout (0d51fa72b) reproduces the identical failure. OMN-15628 PR fix(OMN-15628): bind DELEGATION_ROUTING_TIERS_PATH in every runtime compose lane + bidirectional env parity gate #2621 added DELEGATION_ROUTING_TIERS_PATH to docker-compose.judge.yml's x-judge-runtime-env anchor without the compensating opt-out that docker-compose.infra.yml's own projection-api carries (DELEGATION_ROUTING_TIERS_PATH: "", added under OMN-15645). Fix: added the same opt-out to judge.yml's projection-api, mirroring infra.yml exactly. Verified the specific test now passes (1 passed in 0.48s), plus the full x-runtime-env/env-parity/delegation-routing test surface (tests/ci/test_runtime_env_anchor.py, tests/ci/test_env_parity.py, all four tests/integration/infra/test_*_compose_render.py files, tests/unit/docker/test_delegation_routing_tiers_path_binding_omn15645.py, tests/unit/docker/test_runtime_entrypoint_delegation_tiers_self_heal.py) — 57 passed.

    • Parallel-lane collision, disclosed rather than resolved unilaterally: while finishing this fix, PR #2628 (codex/omn-15567-delegation-path-ci-fix, OPEN, created 2026-08-02T08:55Z — essentially simultaneous with this remediation round) independently fixes the identical docker-compose.judge.yml line with the same content, and additionally adds 4 missing keys (ONEX_BOUNDARY_DLQ_ENABLED, ONEX_SINGLE_OWNER_COMMAND_TOPICS, ONEX_TOPIC_ENFORCEMENT_MODE, ONEX_WIRING_STRICT_MODE) to docker-compose.infra.yml's x-runtime-env anchor that this PR's Env Parity (docker-compose vs k8s ConfigMap) job separately flagged as missing (cross-repo drift against omninode_infra's onex-dev k8s manifests — unrelated to OMN-15567, not touched here to avoid a duplicate/conflicting diff against fix(OMN-15567): keep judge projection API routing opt-out #2628's in-flight fix). This PR does not attempt to merge, close, or race fix(OMN-15567): keep judge projection API routing opt-out #2628; whichever lands first, the other needs a trivial rebase (the judge.yml hunks are byte-identical). Codex owns merge sequencing.
  2. tests/integration/test_monitor_alert_emitter_integration.py: reverted the false pytestmark = pytest.mark.kafka added in remediation round 1 on a disproven claim. Every MonitorAlertEmitter(...) construction in the file is wrapped in patch("confluent_kafka.Producer", ...); the one unpatched construction (test_emitter_disabled_when_kafka_env_missing) deletes KAFKA_BOOTSTRAP_SERVERS first so the emitter self-disables before _init_clients() ever calls confluent_kafka.Producer(...). Verified: 8 passed in 0.15s standalone, 8 passed under CI's exact regular-split marker filter (-m "not slow and not chaos and not kafka and not performance"). What actually caused the 15-min Split-2/2 timeout is still unknown. A live full-suite repro on .200 against a clean origin/dev checkout, using the identical selector flags, reproduced the same [gwN] node down: Not properly terminated worker-death signature at ~13% progress on tests/integration/services/snapshot/test_store_postgres_integration.py::TestConcurrentSequenceGeneration::test_high_concurrency_sequence_uniqueness — a Postgres concurrency test with zero Kafka involvement — which rules out this file (and Kafka generally) as the mechanism. Filed OMN-15658 to track the real root cause; out of this ticket's scope (nightly redpanda SIGILL / runner topology, not the regular CI Tests (Split N) job's full-suite reliability). Also ruled out and documented in OMN-15658 so a future investigator doesn't re-walk the dead end: tests/integration/verification/test_registration_contract_verify.py (skips before reaching its unmocked confluent_kafka construction, because Tests (Split N) has no live Postgres service and its db_query_fn kwarg is evaluated — and skips — first) and tests/integration/runtime/test_steel_dispatch_golden_chain_live_runtime.py (gated behind a live-reachability skipif against an unreachable .201 host).

  3. .github/workflows/nightly-integration.yml: POSTGRES_PORT was never re-emitted in "Resolve reachable e2e connectivity host." It was set once, in the earlier "Derive isolated e2e namespace" step, to the ephemeral host-published port, and stayed paired with that original value even after POSTGRES_HOST was overridden to resolved_host — an invalid host/port pair under the Docker-DNS and gateway branches (only POSTGRES_HOST was corrected; INTEGRATION_POSTGRES_PORT was, POSTGRES_PORT was not). Confirmed live on run 30733477609 (this PR's own cited success evidence): POSTGRES_HOST=omnibase-infra-e2e-30733477609-1--postgres alongside POSTGRES_PORT=37919 (the original ephemeral host port) — invalid under every topology. Latent that run only because its consumers (tests/integration/test_dispatch_roundtrip.py, tests/integration/injection_effectiveness/conftest.py, tests/integration/migrations/test_node_migration_shape_drift_omn15376.py) were separately gated out. Fix: added echo "POSTGRES_PORT=${resolved_pg_port}" alongside the other resolved_* exports.

  4. .github/workflows/nightly-integration.yml: the network-disconnect teardown fix from remediation round 1 was unobservable in the direction it matters. docker network disconnect ... 2>/dev/null || true swallows every failure, so if the disconnect doesn't take, the run leaks the attachment exactly as before and the log stays silent — the same silent-failure shape the fix existed to remove. Fix: added a post-disconnect docker network inspect ... --format '{{json .Containers}}' | grep -q "$OMNIBASE_INFRA_RUNNER_CONTAINER_ID" check that prints a non-fatal ::error:: annotation naming the network and container if the runner is still attached — the teardown step still completes and still runs docker compose down either way, but a real failure to detach is now visible in the log instead of silent. Two new stub-driven tests: test_teardown_surfaces_error_when_disconnect_does_not_take (still-attached path prints ::error::) and test_teardown_silent_when_disconnect_verified_gone (verified-clean path prints nothing) — the existing test_teardown_disconnects_runner_container_before_compose_down / test_teardown_skips_disconnect_when_runner_was_never_attached still pass unmodified.

No live workflow_dispatch run exists yet reflecting this remediation round's changes (the mapfile→while-read rewrite, the resolved_host/resolved_kafka_host split, and the network-disconnect teardown fix from round 1 were also never exercised live before this round — the only workflow_dispatch run on this branch, 30733477609, predates all of round 1). A fresh dispatch after this push is still owed before AC1/AC3/AC4 can be called live-proven rather than unit-proven; see the ticket comment for the live run link once available.

Gates (remediation round 2)

  • gates_host: stickybeatz-studio (.200), patch-transfer + sha256-verified per rule 11a/the standing invariant.
  • ruff format --check / ruff check: clean on both touched Python files.
  • mypy: clean on test_nightly_e2e_runner_connectivity.py; test_monitor_alert_emitter_integration.py has the same 2 pre-existing Unused "type: ignore" comment findings confirmed present on origin/dev's own unmodified copy of the file (line 256/257, verified via a direct mypy run against git show origin/dev:...), unchanged by this diff.
  • Targeted: tests/integration/test_monitor_alert_emitter_integration.py (8 passed standalone, 8 passed under the CI marker filter), tests/unit/docker/test_nightly_e2e_runner_connectivity.py (17 passed, up from 15 — 2 net new), tests/unit/docker/test_e2e_compose_lane_isolation.py + connectivity file together (35 passed), the judge-compose + env-parity + delegation-routing surface listed in item 1 above (57 passed), tests/integration/infra/ + tests/unit/infra/ + tests/unit/docker/ (369 passed, 3 skipped, skips are pre-existing Docker-daemon-required).
  • pre-commit run --all-files: clean on all 4 changed files (both hooks that failed repo-wide — URL Authority Gate on 5 unrelated files, ONEX SPDX Header Requirement on 2 unrelated files with a 2025-vs-2026 year mismatch — are pre-existing repo debt in files this diff never touches, confirmed by grep against the changed-file list); the actual git commit pre-commit run (scoped to the 4 changed files) was 100% Passed/Skipped, zero failures.
  • Pre-push governed selector (prepush-smart-tests, OMN-13973), the real git push pre-push hook that gated this commit: tests/ci/ tests/unit/docker/ — 1497 passed, 3 skipped (same 3 pre-existing Docker-daemon skips; +2 net vs. round 1's 1495 from the 2 net-new teardown-observability tests).
  • Deploy-gate finding, fixed via Evidence-Source, not code: this PR's docker/docker-compose.judge.yml change touches a deploy-gate "runtime path," and Evidence-Source: OCC#5933 (this PR body's original citation) resolved to a stale OCC contract snapshot (commit 05c85ec8) that predates the falsifiable dod-infra-pr-2623-535f8a74 evidence item added by the later, already-merged OCC companion PR #5939 (commit 3b30ef918, confirmed via gh api repos/OmniNode-ai/onex_change_control/contents/contracts/OMN-15567.yaml?ref=<sha> at both refs). Updated the citation below to OCC#5939 so deploy-gate resolves a contract snapshot that already declares a falsifiable gh api ... probe.

Seam definition (every field/key this PR reads or writes across a boundary)

$GITHUB_ENV within the integration-tests job (GitHub Actions' cross-step boundary — each >> write is visible to every later step in the job):

Key Type Producer Consumer
OMNIBASE_INFRA_RUNNER_CONTAINER_ID string (container id or empty) "Detect runner network topology" "Resolve reachable e2e connectivity host", "Tear down e2e stack" (network disconnect guard)
E2E_REDPANDA_ADVERTISE_HOST string (IPv4 or "localhost") "Detect runner network topology" docker-compose.e2e.yml redpanda service (--advertise-kafka-addr/--advertise-pandaproxy-addr), "Resolve reachable e2e connectivity host"
KAFKA_BOOTSTRAP_SERVERS string host:port "Derive isolated e2e namespace" (default localhost:$KAFKA_PORT) → overridden by "Resolve reachable e2e connectivity host" (resolved_kafka_host:port) Run integration tests (pytest env), test_consumer_health_pipeline.py/test_runtime_log_bridge_pipeline.py/test_runtime_health_monitor.py
OMNIBASE_INFRA_DB_URL string (postgres DSN) same override pattern (resolved_host:pg_port) pytest fixtures reading OMNIBASE_INFRA_DB_URL
INTEGRATION_POSTGRES_HOST / INTEGRATION_POSTGRES_PORT / POSTGRES_HOST string "Resolve reachable e2e connectivity host", from resolved_host (postgres-specific host) pytest env
OMNI_INFRA_HOST / REDPANDA_ADVERTISE_HOST string "Resolve reachable e2e connectivity host", from resolved_kafka_host (redpanda-specific host — NEW in the remediation round, previously incorrectly sourced from resolved_host/postgres) pytest env (e.g. test_runtime_consumes_build_loop_terminal_event.py builds its own DSN from host+port)

docker-compose.e2e.yml boundary (pre-existing key, previously always defaulted): ${E2E_REDPANDA_ADVERTISE_HOST:-localhost} — now actually populated by "Detect runner network topology" instead of silently defaulting every run.

DOCKER_CALL_LOG-adjacent (test-only, not a runtime boundary): the new teardown tests capture docker CLI invocations to a log file for ordering assertions, mirroring test_e2e_compose_lane_isolation.py's pre-existing _run_teardown_with_stubs pattern rather than inventing a new stub mechanism.

Acceptance criteria → proof

AC Proof
1. workflow_dispatch run after OMN-15565, fresh Spin up e2e stack log Runs 30681782952 / 30685774570 (cited above).
2. If exit 132 recurs, name root cause N/A — did not recur on either post-fix run. Root cause of the prior 8/8 pattern: OMN-15565 lane-collision (documented above), not addressed further here (OMN-15565 owns it).
3. Suite runs to completion, pass/fail recorded Run 30681782952: 2471 passed, 444 skipped, 1 xfailed, 6 failed, 13 errors in 848.45s. test_resolve_connectivity_runs_between_health_checks_and_pytest locks the step ordering that makes this possible.
4. Runner-topology question answered yes/no + probe, fixed or non-applicable Answered YES (probe = the live connection-refused evidence above). Fixed by "Detect runner network topology" + "Resolve reachable e2e connectivity host", proven by test_detect_topology_defaults_to_localhost_off_a_bare_runner / test_detect_topology_is_compatible_with_bin_bash_specifically (negative default, bash-3-pinned), test_resolve_connectivity_fails_closed_when_nothing_is_reachable (fail-closed), test_resolve_connectivity_prefers_localhost_when_reachable, test_resolve_connectivity_falls_back_to_docker_host_gateway (the exact confirmed-live failure mode), test_resolve_connectivity_uses_container_specific_hosts_for_docker_dns (the exact confirmed-live Docker-DNS branch, remediation round), test_resolve_connectivity_refuses_to_resolve_to_live_201_host (safety), test_teardown_disconnects_runner_container_before_compose_down / test_teardown_skips_disconnect_when_runner_was_never_attached (network-leak fix), plus test_kafka_test_files_read_bootstrap_servers_from_env / test_python_can_import_the_fixed_kafka_test_modules for the two hardcoded test files.

Tests

tests/unit/docker/test_nightly_e2e_runner_connectivity.py: 16 cases (11 original + 5 added in the remediation round: test_detect_topology_is_compatible_with_bin_bash_specifically, test_resolve_connectivity_uses_container_specific_hosts_for_docker_dns, test_teardown_disconnects_runner_container_before_compose_down, test_teardown_skips_disconnect_when_runner_was_never_attached, plus the parametrized kafka-file case count is unchanged). RED-before/GREEN-after against the real run: shell of each step via subprocess + stub docker/python3/uv binaries — mirrors test_e2e_compose_lane_isolation.py's existing pattern, not a structural grep.

tests/integration/test_monitor_alert_emitter_integration.py: no new test cases (marker-only fix); all 8 existing cases still pass standalone (-m kafka, 0.09s) and are now correctly deselected from the CI regular split (-m "not kafka", 8 deselected).

Unrelated pre-existing ratchet suite (test_e2e_compose_lane_isolation.py, 18 cases, OMN-15565) still passes unmodified — confirms the teardown network-disconnect addition is additive, not a regression to that file's project/manifest-guard assertions.

Gates

  • gates_host: stickybeatz-studio (.200), patch-transfer + sha256-verified per rule 11a/the standing invariant, for both the original push and this remediation round.
  • Invocation-shell honesty (root cause of defect feat: RedPanda Event Bus Integration with Fail-Fast Infrastructure #3 above): all commands below were run via ssh stickybeatz-studio 'zsh -lc "..."' (login shell, canonical .200 PATH including /opt/homebrew/bin where uv/timeout/GNU coreutils/Homebrew bash actually live). A raw non-login ssh stickybeatz-studio 'cmd' (no -l) lacks that PATH and produces 5 spurious "command not found" failures unrelated to code correctness — reproduced and confirmed identical against origin/dev's own unmodified tree.
  • ruff format --check / ruff check: clean on all 3 touched Python/YAML-adjacent files (python files only — the workflow YAML is not a ruff target; confirmed by the tool erroring trying to parse YAML as Python when mistakenly included).
  • mypy: clean on nightly-integration.yml's Python test surface and test_consumer_health_pipeline.py/test_runtime_log_bridge_pipeline.py (from the original push); test_monitor_alert_emitter_integration.py (untouched by this PR except the marker line) has 2 pre-existing Unused "type: ignore" comment findings at lines that shifted with my diff (256→270, 257→271) — confirmed pre-existing via git stash/mypy-before comparison, not introduced by this PR.
  • actionlint: 60 shellcheck findings both before and after this remediation round's edit (git stash comparison) — zero new findings introduced by the mapfile→while read rewrite or the disconnect/host-split changes.
  • pre-commit run --files <3 changed files>: all hooks Passed or Skipped (no files to check), zero failures; sha256 of all 3 files unchanged across the pre-commit run (no hook silently mutated content).
  • Targeted suite (tests/unit/docker/test_nightly_e2e_runner_connectivity.py + test_e2e_compose_lane_isolation.py): 33 passed in 8.20s.
  • test_nightly_e2e_runner_connectivity.py alone, PATH forced to exclude /opt/homebrew/bin (i.e. bash resolves to the real /bin/bash 3.2.57, matching the verifier's exact reproduction environment): 15 passed in 7.65s — proves the mapfile fix holds under the harsher PATH too, not just the login-shell one.
  • test_monitor_alert_emitter_integration.py under -m "not slow and not chaos and not kafka and not performance" (CI's regular split selection): 8 deselected. Same file under -m kafka: 8 passed in 0.09s (no hang).
  • Pre-push governed selector (prepush-smart-tests, OMN-13973), run twice — once ad hoc on .200 (login shell) and once as the actual git push pre-push hook that gated this remediation commit's push: tests/ci/ tests/unit/docker/ — 1495 passed, 3 skipped (same 3 pre-existing Docker-daemon-required skips as before; +4 net vs. the originally-claimed 1491 from the 4 net-new tests, after subtracting the 1 fixed mapfile failure that is no longer a failure).

Non-closure note

Per the ticket: not closing/muting the schedule here. Once this branch's fresh dispatch run and the AC1-4 write-up land in the ticket, the nightly's stability going forward is an operator call, not this PR's.

Live post-fix workflow_dispatch confirmation (run 30741864371)

The first run reflecting every remediation-round-2 fix, dispatched after the push above:

  • No exited (132) / SIGILL — AC1/AC2 stay closed.
  • 26 failed, 2678 passed, 237 skipped, 1 xfailed, 25 warnings, 49 errors in 1057.90s (0:17:37) — AC3 (an executed suite, not a green one, is what's required).
  • POSTGRES_HOST/OMNI_INFRA_HOST/REDPANDA_ADVERTISE_HOST resolved to Docker-DNS container names; POSTGRES_PORT correctly moved from the ephemeral 47407 to the resolved 5432 at "Run integration tests" (item 3's fix, live-confirmed for the first time) — AC4.
  • Teardown: Network omnibase-infra-e2e-30741864371-1--network Removed cleanly, no "Resource is still in use," no ::error:: annotation fired (verified-clean disconnect, not just assumed) — network-leak fix (round 1) and its observability check (round 2, item 4) both live-confirmed on the same run.

Full write-up posted to the OMN-15567 ticket.

Evidence-Source: OCC#5939
Evidence-Ticket: OMN-15567

@coderabbitai

coderabbitai Bot commented Aug 2, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

You’ve reached a temporary PR review limit under our Fair Usage Limits Policy.

Your recent review volume is higher than typical usage, so adaptive limits are currently applied.

Next review available in: 45 seconds

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: e225c154-799c-4a67-a484-bcb9df71371d

📥 Commits

Reviewing files that changed from the base of the PR and between 089e803 and 291ea8d.

📒 Files selected for processing (7)
  • .github/workflows/nightly-integration.yml
  • tests/integration/runtime/test_golden_chain_live_runtime.py
  • tests/integration/runtime/test_steel_dispatch_golden_chain_live_runtime.py
  • tests/integration/test_consumer_health_pipeline.py
  • tests/integration/test_runtime_log_bridge_pipeline.py
  • tests/integration/verification/test_registration_contract_verify.py
  • tests/unit/docker/test_nightly_e2e_runner_connectivity.py

Comment @coderabbitai help to get the list of available commands.

jonahgabriel pushed a commit to OmniNode-ai/onex_change_control that referenced this pull request Aug 2, 2026
#5933)

* evidence: OCC companion pass 1 for OmniNode-ai/omnibase_infra#2623

* evidence: OCC companion self-bind for #5933

---------

Co-authored-by: node-occ-companion-effect <occ-companion-effect@omninode.ai>
@jonahgabriel

Copy link
Copy Markdown
Collaborator Author

Fresh workflow_dispatch run on this branch completed: https://github.com/OmniNode-ai/omnibase_infra/actions/runs/30733477609

  • Redpanda Healthy in ~7s, no exit 132 (confirms AC1/AC2 close, OMN-15565's fix holds).
  • "Resolve reachable e2e connectivity host" resolved via Docker DNS (e2e stack reachable by generated container names (Docker DNS)) — the runner container was confirmed containerized (Runner container ID: 0903870886a5..., gateway 172.18.0.1) and joined the compose network.
  • Suite ran to completion: 2677 passed / 237 skipped / 1 xfailed / 27 failed / 49 errors in 1053.53s. Zero KafkaConnectionError/Connect call failed/Unable to bootstrap — the 13 connectivity errors from the pre-fix baseline (run 30681782952) are gone.
  • Remaining 27 failed / 49 errors are pre-existing app/test defects unexpectedly newly-exercised now that connectivity works (missing tables, an asyncio cross-loop fixture bug, missing env var, the known .200 LLM-allowlist failures, one judge-lane assertion inherited from dev HEAD post-fix(OMN-15628): repair #2620/#2621 duplicate-key collision that put dev red + duplicate-key gate #2622) — none caused by this PR's diff, out of scope per the ticket's own "executed, not green" acceptance.

Full AC1-4 write-up posted to OMN-15567.

@github-actions

github-actions Bot commented Aug 2, 2026

Copy link
Copy Markdown
Contributor

⚠️ Hostile Reviewer — DEGRADED (informational)

Blocking findings (critical): 0
Total findings: 0
Models succeeded: none

Note: All reviewer models failed or were unavailable. Degraded results are informational during the pilot phase (OMN-8468/OMN-8524) and do not block merge. Error: all review endpoints [192.168.86.201:8000 192.168.86.201:8001 ] unreachable — preflight short-circuit (no models available)


Gate semantics (pilot phase)

Verdict Meaning Blocks merge?
passed No critical findings No
blocked CRITICAL findings found Yes
degraded All models unavailable (infra) No (pilot)

Powered by omniintelligence.review_pairing.cli_review — node-based adversarial review via HandlerLlmCliSubprocess (OMN-8468/OMN-8524)

@jonahgabriel

Copy link
Copy Markdown
Collaborator Author

CI note (not caused by this diff): Tests (Split 2/2) has been cancelled 4/4 times on this branch, always ~15-16 minutes into the job — matching ci.yml's job-level timeout-minutes: 15 for that job (line ~1286). Tests (Split 1/2) has succeeded 4/4 times on the same commits. Root cause appears to be pytest-split's fallback behavior when no .test_durations cache exists ([pytest-split] No test durations found. Pytest-split will split tests evenly when no durations are found.), which can land a disproportionate share of slow tests in one bucket by pure test-count balancing rather than measured duration.

This is not plausibly caused by this PR's diff: the diff adds one new unit test file (11 cases, ~10s total locally and on .200) plus two one-line env-var reads in existing test files — nothing that would push an unrelated test bucket over a 15-minute budget. A concurrent, unrelated PR on this repo (omninode-infra OMN-15604, run 30735242910) completed its own 3-way split cleanly in the same time window, so this doesn't look like org-wide capacity exhaustion either — more likely this PR's specific split-2 bucket composition is unlucky under the no-cache fallback.

Queued another gh run rerun --failed (5th attempt). Every other required check on this PR is green, including the connectivity-fix-specific gates and the OCC companion (OCC#5933, merged). Not merging — leaving final CI-green confirmation to the merge sweep.

@jonahgabriel

Copy link
Copy Markdown
Collaborator Author

Remediation round pushed (55ec3ab), all 6 flagged defects fixed and verified — see PR body.

CI is currently RED on Tests (Split 2/3) for a reason unrelated to this diff: tests/integration/infra/test_judge_compose_render.py::test_judge_lane_delegation_routing_tiers_path_binding fails on this branch —

AssertionError: Service 'projection-api' deliberately has no delegation-routing surface and must not bind DELEGATION_ROUTING_TIERS_PATH; got '/app/config/delegation/routing_tiers.yaml'

Confirmed this is live on origin/dev's own current tip (0d51fa72b3), independent of this branch — same assertion, same failure, ran directly against a freshly pulled dev:

tests/integration/infra/test_judge_compose_render.py::test_judge_lane_delegation_routing_tiers_path_binding FAILED
1 failed, 3 warnings in 0.67s

This traces to the OMN-15628/OMN-15645 delegation-routing-tiers-path binding work (both currently "In Review") — 04b4b43e2 ("bind DELEGATION_ROUTING_TIERS_PATH in every runtime compose lane") is an ancestor of this branch and appears to have over-bound the judge lane's projection-api service, which the test says deliberately must not carry that binding. Out of scope for OMN-15567 (nightly e2e runner-topology connectivity) — not touching it here per the remediation instruction to fix exactly the flagged defects and not rewrite unrelated code. Every job whose split happens to draw this test file is currently red on dev itself, not just this PR.

Zombie note: the first CI attempt on this head also showed as cancelled end-to-end (Detect Changes cancelled with 0 steps) — traced to a stuck in_progress CI run from this PR's previous head (4ef8477c, run 30736557492, stuck since 06:47Z) holding the CI-2623 concurrency-group slot. Cancelled that zombie run and re-ran this head's CI (gh run rerun 30738833038) to get the clean signal above.

@jonahgabriel
jonahgabriel force-pushed the jonah/omn-15567-nightly-redpanda-sigill branch from a5465c8 to e317e07 Compare August 2, 2026 10:36
jonahgabriel and others added 8 commits August 2, 2026 05:20
…r redpanda SIGILL confound clears

The nightly e2e stack never reached pytest for 8/8 nights (redpanda exited
132/SIGILL). Re-observed post-OMN-15565 on workflow_dispatch run 30681782952
and scheduled run 30685774570: redpanda now starts Healthy in ~5s and pytest
runs to completion (2471 passed/444 skipped/1 xfailed/6 failed/13 errors in
848.45s) -- the SIGILL was the OMN-15565 lane-collision confound, not a real
defect, and is gone now that the stack is genuinely isolated.

The 13 errors are the second defect this ticket targets: the self-hosted
omnibase-ci runner does not share a network namespace with the Docker host
(same topology reusable-runtime-boot.yml:250-270 already documents), so
"localhost:<host-published-port>" is unreachable from the runner process even
when the port is the freshly-derived, correctly-read dynamic port. Mirrors
reusable-runtime-boot.yml's existing, proven detection/probe pattern:

- "Detect runner network topology": identify a containerized runner and
  resolve the Docker host gateway, so redpanda's external listener can
  advertise it instead of always defaulting to "localhost".
- "Resolve reachable e2e connectivity host": after the stack is Healthy,
  probe Docker DNS -> localhost -> Docker host gateway in that order and
  export KAFKA_BOOTSTRAP_SERVERS/OMNIBASE_INFRA_DB_URL/INTEGRATION_POSTGRES_HOST
  etc. from whichever actually works; fail closed with diagnostics if none do.
  Both new steps carry their own live-.201 guard, since they compute values
  the workflow-level static guard never sees.

Also fixes two integration tests that hardcoded BOOTSTRAP_SERVERS =
"localhost:19092", silently ignoring the CI-derived KAFKA_BOOTSTRAP_SERVERS
regardless of topology -- matching the pattern already used correctly by
test_runtime_health_monitor.py.

OMN-15567
…y, Docker-DNS host cross-wiring, network-leak teardown

Adversarial verifier found six defects in the prior push (head 4ef8477):

1. CI Summary FAILURE (required check on omnibase_infra/dev) -- Tests
   (Split 2/2) was cancelled at the 15-min job timeout 3x on this
   branch, misdiagnosed in the PR comment as an unlucky pytest-split
   bucket. Root cause: tests/integration/test_monitor_alert_emitter_integration.py
   is @pytest.mark.integration-only (missing @pytest.mark.kafka), so
   CI regular splits `-m "not slow and not chaos and not kafka and not
   performance"` never deselect it, unlike its sibling files
   (test_consumer_health_pipeline.py, test_runtime_log_bridge_pipeline.py,
   test_runtime_health_monitor.py) which already carry the kafka marker
   for exactly this reason. Added pytestmark = pytest.mark.kafka at
   module level, matching the established convention.

2. tests/unit/docker/test_nightly_e2e_runner_connectivity.py:171 was
   permanently RED on the mandated gate host (.200): nightly-integration.yml
   used `mapfile`, a bash 4+ builtin, and macOS ships bash 3.2.57 at
   /bin/bash (no newer bash for GPLv3 licensing reasons). Rewrote
   "Detect runner network topology" to use a `while read` loop instead
   -- bash-3-compatible, identical behavior. New regression test pins
   /bin/bash explicitly so a newer bash earlier on PATH cannot hide a
   future regression back to a bash-4-only construct.

3. PR body's claimed pre-push governed selector result ("1491 passed, 3
   skipped") was not reproducible via raw non-login ssh (matches the
   verifier's invocation): 5 of the 6 failures are pre-existing PATH
   artifacts of a non-login shell lacking /opt/homebrew/bin (uv/timeout
   not on PATH) -- confirmed identical under origin/dev's own tree, not
   introduced by this diff. The 6th (mapfile) is fixed by #2 above and
   now passes under both a login-shell PATH and a PATH forced to
   exclude /opt/homebrew/bin (i.e. raw /bin/bash 3.2.57).

4/5. "Resolve reachable e2e connectivity host"'s Docker DNS branch (the
   ONLY branch live traffic actually took, per the run the PR cited as
   its own success evidence) cross-wired REDPANDA_ADVERTISE_HOST and
   OMNI_INFRA_HOST to the *postgres* container's Docker-DNS name instead
   of redpanda's -- the only branch where postgres and redpanda have
   DIFFERENT hostnames rather than one shared host disambiguated by
   port. Split resolved_host (postgres) from a new resolved_kafka_host
   (redpanda) and wired REDPANDA_ADVERTISE_HOST/OMNI_INFRA_HOST from the
   latter. New test drives the previously-untested Docker DNS branch via
   stubbed docker network connect/docker inspect.

6. "Resolve reachable e2e connectivity host" attaches the long-lived
   self-hosted runner container to the run's compose network via
   docker network connect and nothing ever disconnected it -- proven
   live: run 30733477609's teardown logged "Network ... Resource is
   still in use" without failing. Added a docker network disconnect
   in "Tear down e2e stack", before the destructive down -v, guarded
   to no-op when the runner was never attached. Two new tests cover both
   branches.

Gates on .200 (patch-transfer, sha256-verified): targeted suite 33
passed; full tests/ci/+tests/unit/docker/ selector 1495 passed/3
skipped (login-shell PATH with /opt/homebrew/bin); ruff format/check
clean; mypy clean except 2 pre-existing unused type:ignore comments at
test_monitor_alert_emitter_integration.py (confirmed pre-existing via
git stash comparison, unrelated to this diff); actionlint 60
pre-existing findings, zero new; pre-commit run --files clean.

OMN-15567
…N_ROUTING_TIERS_PATH opt-out, revert false kafka marker, POSTGRES_PORT re-emit, observable network-disconnect verification

Fixes 4 defects found by an independent adversarial verifier on head 4ef8477/55ec3ab:

1. BLOCKER: required CI Summary was FAILURE at PR head due to
   tests/integration/infra/test_judge_compose_render.py::test_judge_lane_delegation_routing_tiers_path_binding.
   Root cause confirmed dev-inherited (byte-identical compose/test files
   between HEAD and origin/dev, reproduces on a clean origin/dev checkout):
   OMN-15628 PR #2621 added DELEGATION_ROUTING_TIERS_PATH to
   docker-compose.judge.yml x-judge-runtime-env without the compensating
   opt-out that docker-compose.infra.yml carries on its own projection-api.
   Fixed by adding the same DELEGATION_ROUTING_TIERS_PATH: "" override to
   judge.yml projection-api, mirroring infra.yml. Filed nothing new for this
   -- OMN-15628 already tracks the area, cited in the compose comment.

2. tests/integration/test_monitor_alert_emitter_integration.py: reverted the
   pytestmark = pytest.mark.kafka added last round on a disproven claim
   (every MonitorAlertEmitter(...) construction in the file is wrapped in
   patch("confluent_kafka.Producer", ...), and the one unpatched
   construction deletes KAFKA_BOOTSTRAP_SERVERS first so the emitter
   self-disables before constructing anything -- 8 passed in 0.15s
   standalone, 8 passed under the CI regular-split marker filter). Live
   full-suite repro on .200 against a clean origin/dev checkout reproduced
   the same worker-death signature ("node down: Not properly terminated")
   on an unrelated Postgres concurrency test with zero Kafka involvement,
   disproving the Kafka-specific attribution. Filed OMN-15658 to track the
   real (still-unidentified) root cause; out of OMN-15567 scope.

3. .github/workflows/nightly-integration.yml: POSTGRES_PORT was set once to
   the ephemeral host-published port and never re-emitted in the
   Docker-DNS/gateway host-resolution rewrite, leaving it paired with the
   wrong host. Added the missing echo alongside the other resolved_* keys.

4. .github/workflows/nightly-integration.yml: the network-disconnect
   teardown fix from last round swallowed its own exit code silently
   (2>/dev/null || true), reproducing the exact silence class the fix
   existed to close. Added a post-disconnect docker network inspect
   verification that prints a non-fatal ::error:: annotation if the runner
   container is still attached, with 2 new stub-driven tests covering both
   the surfaced-error and verified-clean paths.

[OMN-15567]
@jonahgabriel
jonahgabriel force-pushed the jonah/omn-15567-nightly-redpanda-sigill branch from e317e07 to 2b4447c Compare August 2, 2026 12:33
@jonahgabriel
jonahgabriel merged commit bf070a9 into dev Aug 2, 2026
98 checks passed
@jonahgabriel
jonahgabriel deleted the jonah/omn-15567-nightly-redpanda-sigill branch August 2, 2026 13:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant