Repository navigation
Add Claude Code GitHub Workflow - #2
Merged
jonahgabriel merged 2 commits intoSep 11, 2025
Merged
jonahgabriel merged 2 commits into
jonahgabriel merged 2 commits into
Conversation
This was referenced Dec 4, 2025
This was referenced Dec 15, 2025
Merged
3 tasks done
3 of 4 tasks
This was referenced Dec 24, 2025
7 tasks done
This was referenced Jan 13, 2026
docs: comprehensive documentation rewrite with Quick Start and architecture overview [OMN-1375]
#172
Merged
Merged
jonahgabriel
added a commit
that referenced
this pull request
Apr 20, 2026
…cripts Four findings from the CodeRabbit review on PR #1352, all legitimate correctness improvements to pre-existing behavior that's now in-scope because we're already touching these files. - CR #1, #4: yaml.safe_load may return None or a scalar; guard with isinstance check and fail fast with type-of-value in the message. - CR #2 (MAJOR): missing top-level subscription arrays (READ_MODEL_TOPICS, EXPECTED_TOPICS) were a warning + silent pass. A rename or deletion of either array would silently succeed — exactly the breakage this gate exists to catch. Add required=True kwarg on top-level calls; recursive spread lookups still fall back to topics.ts with a warning. - CR #3 (MAJOR): the parity check only walked consumer -> registry. A newly-declared registry topic that was never wired into READ_MODEL_TOPICS or EXPECTED_TOPICS passed the gate. Add a reverse check that every registry omniclaude evt topic is covered by both consumer arrays. Tests: four new unit tests cover required-array failure, non-dict registry rejection (both scripts), and reverse-parity failure. All 10 tests pass.
github-merge-queue Bot
pushed a commit
that referenced
this pull request
Apr 20, 2026
…6] (#1352) * chore(scripts): relocate topic-parity scripts from omni_home [OMN-9286] omni_home/scripts/ is blocked by the no-functional-code pre-commit hook, which rejects any .py/.sh file in that directory. Two pre-existing scripts (check-topic-parity.py, sync-topic-registry.py — PRs #50/#51, 2026-03-13) violated this and were blocking unrelated docs-only PRs. Relocating to omnibase_infra/scripts/ per the OMN-4922 pattern (pull-all.sh). Changes: * Copy both scripts to omnibase_infra/scripts/ preserving exec bits * Replace module-level global state with OMNI_HOME env var + ModelTopicParityPaths * Add SPDX headers and satisfy mypy --strict + ruff (5 pre-existing PLW0603 + 7 missing-type-arg violations fixed in the move) * Add tests/scripts/test_topic_parity_scripts.py covering shebang, SPDX, argparse surface, and OMNI_HOME resolution Companion omni_home PR will delete the originals and repoint the CI workflow (.github/workflows/topic-parity.yml) at the new location. * fix(scripts): address CodeRabbit findings on relocated topic-parity scripts Four findings from the CodeRabbit review on PR #1352, all legitimate correctness improvements to pre-existing behavior that's now in-scope because we're already touching these files. - CR #1, #4: yaml.safe_load may return None or a scalar; guard with isinstance check and fail fast with type-of-value in the message. - CR #2 (MAJOR): missing top-level subscription arrays (READ_MODEL_TOPICS, EXPECTED_TOPICS) were a warning + silent pass. A rename or deletion of either array would silently succeed — exactly the breakage this gate exists to catch. Add required=True kwarg on top-level calls; recursive spread lookups still fall back to topics.ts with a warning. - CR #3 (MAJOR): the parity check only walked consumer -> registry. A newly-declared registry topic that was never wired into READ_MODEL_TOPICS or EXPECTED_TOPICS passed the gate. Add a reverse check that every registry omniclaude evt topic is covered by both consumer arrays. Tests: four new unit tests cover required-array failure, non-dict registry rejection (both scripts), and reverse-parity failure. All 10 tests pass. * fix(sync-topic-registry): per-entry validation + JSDoc escape Two follow-up CodeRabbit findings on the first fix commit: - CR-minor: load_registry accepted any shape for topics entries; a dict missing 'topic' or both 'event_type'/'topic_base_constant' would raise a raw KeyError downstream instead of a structured exit-2 error with the offending index. Validate each entry's shape on load. - CR-major: descriptions were injected verbatim into /** ... */ JSDoc. A description containing '*/' or a newline would break the generated TypeScript. Escape '*/' to '*\\/' and collapse newlines to spaces. Tests: two new unit tests cover each case. All 12 tests pass. * test(topic-parity): strengthen JSDoc-escape assertion per CR feedback CodeRabbit flagged that the previous test only filtered lines starting with /** and never inspected the full /** ... */ block body, making the */ check vacuous. Parse complete JSDoc blocks with a regex so the assertion actually verifies the escape (and that newlines are collapsed). --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
4 tasks done
jonahgabriel
added a commit
that referenced
this pull request
May 10, 2026
…p config Two CodeRabbit findings + Integration Test Coverage gate: 1. ModelDelegationRequest: add @model_validator(mode='after') rejecting output_schema_key set without compliance_budget. The compliance loop's evaluator requires both — catching this at model construction prevents a downstream assertion crash in the workflow handler when the orchestrator first sees the inference response. (CodeRabbit #2) 2. tests/integration/delegation/test_compliance_loop_wiring_integration.py: exercise HandlerDelegationWorkflow's compliance-loop path against unmocked omnimarket schema-registry / schema-repair / budget-policy handlers and prove the cross-package contract holds end-to-end (3 tests). Closes the Integration Test Coverage gap. Note: CodeRabbit finding #1 (don't require routing_decision for every ROUTED workflow in handle_inference_response) is already addressed by the existing 'if workflow.routing_decision is None: return []' guard at line 389. Evidence-Source: OCC#907 Evidence-Ticket: OMN-10794
jonahgabriel
added a commit
that referenced
this pull request
May 10, 2026
…kflow (#1560) * feat(OMN-10794): wire HandlerComplianceLoop into HandlerDelegationWorkflow Wave 3 / Task 5 of the tokens-to-compliance epic. Adds optional ``output_schema_key`` and ``compliance_budget`` fields to ModelDelegationRequest. When set, HandlerDelegationWorkflow.handle_inference_response invokes HandlerComplianceLoop per attempt: * compliant or budget ABORT → record the attempt, transition ROUTED → INFERENCE_COMPLETED, forward to the quality gate (terminal event carries the running tokens_to_compliance + compliance_attempts) * non-compliant + budget CONTINUE → emit a fresh ModelInferenceIntent carrying the repair prompt and stay in ROUTED (self-loop), incrementing compliance_attempts for the next iteration DelegationWorkflowState gains ``compliance_attempts`` and ``accumulated_tokens``; the FSM transition table allows ROUTED → ROUTED for the repair re-prompt path (both in code _VALID_TRANSITIONS and in contract.yaml). Bumps node contract version to 0.3.0 and updates the FSM transition documentation. The compliance counters are populated onto the terminal ModelDelegationResult and ModelTaskDelegatedEvent so the omnimarket projection (OMN-10793) and the omniclaude sqlite_adapter (OMN-10789) can write them to delegation_events. Backwards compatibility: legacy callers that omit ``output_schema_key`` get the existing single-attempt path unchanged. Their terminal event carries ``compliance_attempts=1`` and ``tokens_to_compliance`` equal to that single attempt's total_tokens — semantically equivalent to first-try success. Tests: 9 new unit tests prove first-try compliance, repair-on-failure with self-loop, two-attempt token accumulation, budget ABORT path, and FSM self-loop legality. 32 existing orchestrator tests still pass. mypy strict clean. pre-commit clean. Evidence-Source: OCC#905 Evidence-Ticket: OMN-10794 * fix(OMN-10794): integration test + model_validator for compliance-loop config Two CodeRabbit findings + Integration Test Coverage gate: 1. ModelDelegationRequest: add @model_validator(mode='after') rejecting output_schema_key set without compliance_budget. The compliance loop's evaluator requires both — catching this at model construction prevents a downstream assertion crash in the workflow handler when the orchestrator first sees the inference response. (CodeRabbit #2) 2. tests/integration/delegation/test_compliance_loop_wiring_integration.py: exercise HandlerDelegationWorkflow's compliance-loop path against unmocked omnimarket schema-registry / schema-repair / budget-policy handlers and prove the cross-package contract holds end-to-end (3 tests). Closes the Integration Test Coverage gap. Note: CodeRabbit finding #1 (don't require routing_decision for every ROUTED workflow in handle_inference_response) is already addressed by the existing 'if workflow.routing_decision is None: return []' guard at line 389. Evidence-Source: OCC#907 Evidence-Ticket: OMN-10794
jonahgabriel
added a commit
that referenced
this pull request
Jul 27, 2026
… + wire live headroom check (#2496) Root cause of the recurring partition-cap regression: docker-compose.stability-test.yml overrides both the redpanda service command AND the redpanda-partition-cap init service with a hardcoded topic_partitions_per_shard value. deploy-runtime.sh's warm_broker_topic_provisioning() force-recreates the redpanda-partition-cap one-shot on EVERY redeploy of the lane (including a --no-deps restart-only deploy that never touches the volume), re-running whatever value is hardcoded there -- silently overwriting any live rpk cluster config set bump made outside the committed config. The lane also silently dropped the base redpanda.yaml startup flag entirely (belt #2 of the documented 3-belt cap design was never applied to this lane). (a) Persist the cap: pin topic_partitions_per_shard=15000 in both belts declared in docker-compose.stability-test.yml (the redpanda service's restored --set flag and the redpanda-partition-cap service's rpk cluster config set call), so every redeploy of this lane carries the raised cap instead of resetting to the 7000 default. 15000 gives durable headroom over the last observed live usage (7046-7047 partitions across the two recorded regressions) without requiring the topic-retirement audit up front. (b) Headroom visibility: add a partition-headroom check (check_partition_headroom) to the EXISTING stability-lane health gate (verify_stability_refresh.py, already invoked by refresh_stability_lane.sh on every lane refresh) -- no new standalone script/dashboard. Queries live topic_partitions_per_shard + summed live partition count; at/over cap is a real, checked FAIL (rpk cluster health alone never sees this); crossing an 80% warn threshold is visible in the report but does not block an otherwise-healthy refresh. (c) Topic-retirement audit and root-causing the lane's disproportionate topic accumulation (OMN-14013 DoD items 1/4) are explicitly OUT of scope here -- documented follow-up, tracked on the ticket itself. Tests: RED proven against the un-pinned 7000 value (regex-based value extraction, not substring matching -- a plain substring check silently passes against this same diff's own forensic comments mentioning superseded values), GREEN after the fix. 32 unit tests for the new headroom check including the exact live incident numbers (7046/7000) as a named regression test. Live verify deferred until post-acceptance -- this PR does not touch the live stability-test lane, restart any broker, or run rpk against 100.109.203.94:39092. Closes OMN-14013 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
jonahgabriel
added a commit
that referenced
this pull request
Aug 2, 2026
…y, Docker-DNS host cross-wiring, network-leak teardown Adversarial verifier found six defects in the prior push (head 4ef8477): 1. CI Summary FAILURE (required check on omnibase_infra/dev) -- Tests (Split 2/2) was cancelled at the 15-min job timeout 3x on this branch, misdiagnosed in the PR comment as an unlucky pytest-split bucket. Root cause: tests/integration/test_monitor_alert_emitter_integration.py is @pytest.mark.integration-only (missing @pytest.mark.kafka), so CI regular splits `-m "not slow and not chaos and not kafka and not performance"` never deselect it, unlike its sibling files (test_consumer_health_pipeline.py, test_runtime_log_bridge_pipeline.py, test_runtime_health_monitor.py) which already carry the kafka marker for exactly this reason. Added pytestmark = pytest.mark.kafka at module level, matching the established convention. 2. tests/unit/docker/test_nightly_e2e_runner_connectivity.py:171 was permanently RED on the mandated gate host (.200): nightly-integration.yml used `mapfile`, a bash 4+ builtin, and macOS ships bash 3.2.57 at /bin/bash (no newer bash for GPLv3 licensing reasons). Rewrote "Detect runner network topology" to use a `while read` loop instead -- bash-3-compatible, identical behavior. New regression test pins /bin/bash explicitly so a newer bash earlier on PATH cannot hide a future regression back to a bash-4-only construct. 3. PR body's claimed pre-push governed selector result ("1491 passed, 3 skipped") was not reproducible via raw non-login ssh (matches the verifier's invocation): 5 of the 6 failures are pre-existing PATH artifacts of a non-login shell lacking /opt/homebrew/bin (uv/timeout not on PATH) -- confirmed identical under origin/dev's own tree, not introduced by this diff. The 6th (mapfile) is fixed by #2 above and now passes under both a login-shell PATH and a PATH forced to exclude /opt/homebrew/bin (i.e. raw /bin/bash 3.2.57). 4/5. "Resolve reachable e2e connectivity host"'s Docker DNS branch (the ONLY branch live traffic actually took, per the run the PR cited as its own success evidence) cross-wired REDPANDA_ADVERTISE_HOST and OMNI_INFRA_HOST to the *postgres* container's Docker-DNS name instead of redpanda's -- the only branch where postgres and redpanda have DIFFERENT hostnames rather than one shared host disambiguated by port. Split resolved_host (postgres) from a new resolved_kafka_host (redpanda) and wired REDPANDA_ADVERTISE_HOST/OMNI_INFRA_HOST from the latter. New test drives the previously-untested Docker DNS branch via stubbed docker network connect/docker inspect. 6. "Resolve reachable e2e connectivity host" attaches the long-lived self-hosted runner container to the run's compose network via docker network connect and nothing ever disconnected it -- proven live: run 30733477609's teardown logged "Network ... Resource is still in use" without failing. Added a docker network disconnect in "Tear down e2e stack", before the destructive down -v, guarded to no-op when the runner was never attached. Two new tests cover both branches. Gates on .200 (patch-transfer, sha256-verified): targeted suite 33 passed; full tests/ci/+tests/unit/docker/ selector 1495 passed/3 skipped (login-shell PATH with /opt/homebrew/bin); ruff format/check clean; mypy clean except 2 pre-existing unused type:ignore comments at test_monitor_alert_emitter_integration.py (confirmed pre-existing via git stash comparison, unrelated to this diff); actionlint 60 pre-existing findings, zero new; pre-commit run --files clean. OMN-15567
jonahgabriel
added a commit
that referenced
this pull request
Aug 2, 2026
…y, Docker-DNS host cross-wiring, network-leak teardown Adversarial verifier found six defects in the prior push (head 4ef8477): 1. CI Summary FAILURE (required check on omnibase_infra/dev) -- Tests (Split 2/2) was cancelled at the 15-min job timeout 3x on this branch, misdiagnosed in the PR comment as an unlucky pytest-split bucket. Root cause: tests/integration/test_monitor_alert_emitter_integration.py is @pytest.mark.integration-only (missing @pytest.mark.kafka), so CI regular splits `-m "not slow and not chaos and not kafka and not performance"` never deselect it, unlike its sibling files (test_consumer_health_pipeline.py, test_runtime_log_bridge_pipeline.py, test_runtime_health_monitor.py) which already carry the kafka marker for exactly this reason. Added pytestmark = pytest.mark.kafka at module level, matching the established convention. 2. tests/unit/docker/test_nightly_e2e_runner_connectivity.py:171 was permanently RED on the mandated gate host (.200): nightly-integration.yml used `mapfile`, a bash 4+ builtin, and macOS ships bash 3.2.57 at /bin/bash (no newer bash for GPLv3 licensing reasons). Rewrote "Detect runner network topology" to use a `while read` loop instead -- bash-3-compatible, identical behavior. New regression test pins /bin/bash explicitly so a newer bash earlier on PATH cannot hide a future regression back to a bash-4-only construct. 3. PR body's claimed pre-push governed selector result ("1491 passed, 3 skipped") was not reproducible via raw non-login ssh (matches the verifier's invocation): 5 of the 6 failures are pre-existing PATH artifacts of a non-login shell lacking /opt/homebrew/bin (uv/timeout not on PATH) -- confirmed identical under origin/dev's own tree, not introduced by this diff. The 6th (mapfile) is fixed by #2 above and now passes under both a login-shell PATH and a PATH forced to exclude /opt/homebrew/bin (i.e. raw /bin/bash 3.2.57). 4/5. "Resolve reachable e2e connectivity host"'s Docker DNS branch (the ONLY branch live traffic actually took, per the run the PR cited as its own success evidence) cross-wired REDPANDA_ADVERTISE_HOST and OMNI_INFRA_HOST to the *postgres* container's Docker-DNS name instead of redpanda's -- the only branch where postgres and redpanda have DIFFERENT hostnames rather than one shared host disambiguated by port. Split resolved_host (postgres) from a new resolved_kafka_host (redpanda) and wired REDPANDA_ADVERTISE_HOST/OMNI_INFRA_HOST from the latter. New test drives the previously-untested Docker DNS branch via stubbed docker network connect/docker inspect. 6. "Resolve reachable e2e connectivity host" attaches the long-lived self-hosted runner container to the run's compose network via docker network connect and nothing ever disconnected it -- proven live: run 30733477609's teardown logged "Network ... Resource is still in use" without failing. Added a docker network disconnect in "Tear down e2e stack", before the destructive down -v, guarded to no-op when the runner was never attached. Two new tests cover both branches. Gates on .200 (patch-transfer, sha256-verified): targeted suite 33 passed; full tests/ci/+tests/unit/docker/ selector 1495 passed/3 skipped (login-shell PATH with /opt/homebrew/bin); ruff format/check clean; mypy clean except 2 pre-existing unused type:ignore comments at test_monitor_alert_emitter_integration.py (confirmed pre-existing via git stash comparison, unrelated to this diff); actionlint 60 pre-existing findings, zero new; pre-commit run --files clean. OMN-15567
jonahgabriel
added a commit
that referenced
this pull request
Aug 2, 2026
…y, Docker-DNS host cross-wiring, network-leak teardown Adversarial verifier found six defects in the prior push (head 4ef8477): 1. CI Summary FAILURE (required check on omnibase_infra/dev) -- Tests (Split 2/2) was cancelled at the 15-min job timeout 3x on this branch, misdiagnosed in the PR comment as an unlucky pytest-split bucket. Root cause: tests/integration/test_monitor_alert_emitter_integration.py is @pytest.mark.integration-only (missing @pytest.mark.kafka), so CI regular splits `-m "not slow and not chaos and not kafka and not performance"` never deselect it, unlike its sibling files (test_consumer_health_pipeline.py, test_runtime_log_bridge_pipeline.py, test_runtime_health_monitor.py) which already carry the kafka marker for exactly this reason. Added pytestmark = pytest.mark.kafka at module level, matching the established convention. 2. tests/unit/docker/test_nightly_e2e_runner_connectivity.py:171 was permanently RED on the mandated gate host (.200): nightly-integration.yml used `mapfile`, a bash 4+ builtin, and macOS ships bash 3.2.57 at /bin/bash (no newer bash for GPLv3 licensing reasons). Rewrote "Detect runner network topology" to use a `while read` loop instead -- bash-3-compatible, identical behavior. New regression test pins /bin/bash explicitly so a newer bash earlier on PATH cannot hide a future regression back to a bash-4-only construct. 3. PR body's claimed pre-push governed selector result ("1491 passed, 3 skipped") was not reproducible via raw non-login ssh (matches the verifier's invocation): 5 of the 6 failures are pre-existing PATH artifacts of a non-login shell lacking /opt/homebrew/bin (uv/timeout not on PATH) -- confirmed identical under origin/dev's own tree, not introduced by this diff. The 6th (mapfile) is fixed by #2 above and now passes under both a login-shell PATH and a PATH forced to exclude /opt/homebrew/bin (i.e. raw /bin/bash 3.2.57). 4/5. "Resolve reachable e2e connectivity host"'s Docker DNS branch (the ONLY branch live traffic actually took, per the run the PR cited as its own success evidence) cross-wired REDPANDA_ADVERTISE_HOST and OMNI_INFRA_HOST to the *postgres* container's Docker-DNS name instead of redpanda's -- the only branch where postgres and redpanda have DIFFERENT hostnames rather than one shared host disambiguated by port. Split resolved_host (postgres) from a new resolved_kafka_host (redpanda) and wired REDPANDA_ADVERTISE_HOST/OMNI_INFRA_HOST from the latter. New test drives the previously-untested Docker DNS branch via stubbed docker network connect/docker inspect. 6. "Resolve reachable e2e connectivity host" attaches the long-lived self-hosted runner container to the run's compose network via docker network connect and nothing ever disconnected it -- proven live: run 30733477609's teardown logged "Network ... Resource is still in use" without failing. Added a docker network disconnect in "Tear down e2e stack", before the destructive down -v, guarded to no-op when the runner was never attached. Two new tests cover both branches. Gates on .200 (patch-transfer, sha256-verified): targeted suite 33 passed; full tests/ci/+tests/unit/docker/ selector 1495 passed/3 skipped (login-shell PATH with /opt/homebrew/bin); ruff format/check clean; mypy clean except 2 pre-existing unused type:ignore comments at test_monitor_alert_emitter_integration.py (confirmed pre-existing via git stash comparison, unrelated to this diff); actionlint 60 pre-existing findings, zero new; pre-commit run --files clean. OMN-15567
jonahgabriel
added a commit
that referenced
this pull request
Aug 2, 2026
…r SIGILL confound clears (#2623) * fix(OMN-15567): resolve nightly e2e runner-topology connectivity after redpanda SIGILL confound clears The nightly e2e stack never reached pytest for 8/8 nights (redpanda exited 132/SIGILL). Re-observed post-OMN-15565 on workflow_dispatch run 30681782952 and scheduled run 30685774570: redpanda now starts Healthy in ~5s and pytest runs to completion (2471 passed/444 skipped/1 xfailed/6 failed/13 errors in 848.45s) -- the SIGILL was the OMN-15565 lane-collision confound, not a real defect, and is gone now that the stack is genuinely isolated. The 13 errors are the second defect this ticket targets: the self-hosted omnibase-ci runner does not share a network namespace with the Docker host (same topology reusable-runtime-boot.yml:250-270 already documents), so "localhost:<host-published-port>" is unreachable from the runner process even when the port is the freshly-derived, correctly-read dynamic port. Mirrors reusable-runtime-boot.yml's existing, proven detection/probe pattern: - "Detect runner network topology": identify a containerized runner and resolve the Docker host gateway, so redpanda's external listener can advertise it instead of always defaulting to "localhost". - "Resolve reachable e2e connectivity host": after the stack is Healthy, probe Docker DNS -> localhost -> Docker host gateway in that order and export KAFKA_BOOTSTRAP_SERVERS/OMNIBASE_INFRA_DB_URL/INTEGRATION_POSTGRES_HOST etc. from whichever actually works; fail closed with diagnostics if none do. Both new steps carry their own live-.201 guard, since they compute values the workflow-level static guard never sees. Also fixes two integration tests that hardcoded BOOTSTRAP_SERVERS = "localhost:19092", silently ignoring the CI-derived KAFKA_BOOTSTRAP_SERVERS regardless of topology -- matching the pattern already used correctly by test_runtime_health_monitor.py. OMN-15567 * chore(OMN-15567): retrigger CI after OCC companion OCC#5933 merged * chore(OMN-15567): retrigger CI with complete Evidence-Source/Evidence-Ticket PR body * test(OMN-15567): bound subprocess.run calls with explicit timeouts (hardening) * fix(OMN-15567): remediation round -- kafka marker, mapfile portability, Docker-DNS host cross-wiring, network-leak teardown Adversarial verifier found six defects in the prior push (head 4ef8477): 1. CI Summary FAILURE (required check on omnibase_infra/dev) -- Tests (Split 2/2) was cancelled at the 15-min job timeout 3x on this branch, misdiagnosed in the PR comment as an unlucky pytest-split bucket. Root cause: tests/integration/test_monitor_alert_emitter_integration.py is @pytest.mark.integration-only (missing @pytest.mark.kafka), so CI regular splits `-m "not slow and not chaos and not kafka and not performance"` never deselect it, unlike its sibling files (test_consumer_health_pipeline.py, test_runtime_log_bridge_pipeline.py, test_runtime_health_monitor.py) which already carry the kafka marker for exactly this reason. Added pytestmark = pytest.mark.kafka at module level, matching the established convention. 2. tests/unit/docker/test_nightly_e2e_runner_connectivity.py:171 was permanently RED on the mandated gate host (.200): nightly-integration.yml used `mapfile`, a bash 4+ builtin, and macOS ships bash 3.2.57 at /bin/bash (no newer bash for GPLv3 licensing reasons). Rewrote "Detect runner network topology" to use a `while read` loop instead -- bash-3-compatible, identical behavior. New regression test pins /bin/bash explicitly so a newer bash earlier on PATH cannot hide a future regression back to a bash-4-only construct. 3. PR body's claimed pre-push governed selector result ("1491 passed, 3 skipped") was not reproducible via raw non-login ssh (matches the verifier's invocation): 5 of the 6 failures are pre-existing PATH artifacts of a non-login shell lacking /opt/homebrew/bin (uv/timeout not on PATH) -- confirmed identical under origin/dev's own tree, not introduced by this diff. The 6th (mapfile) is fixed by #2 above and now passes under both a login-shell PATH and a PATH forced to exclude /opt/homebrew/bin (i.e. raw /bin/bash 3.2.57). 4/5. "Resolve reachable e2e connectivity host"'s Docker DNS branch (the ONLY branch live traffic actually took, per the run the PR cited as its own success evidence) cross-wired REDPANDA_ADVERTISE_HOST and OMNI_INFRA_HOST to the *postgres* container's Docker-DNS name instead of redpanda's -- the only branch where postgres and redpanda have DIFFERENT hostnames rather than one shared host disambiguated by port. Split resolved_host (postgres) from a new resolved_kafka_host (redpanda) and wired REDPANDA_ADVERTISE_HOST/OMNI_INFRA_HOST from the latter. New test drives the previously-untested Docker DNS branch via stubbed docker network connect/docker inspect. 6. "Resolve reachable e2e connectivity host" attaches the long-lived self-hosted runner container to the run's compose network via docker network connect and nothing ever disconnected it -- proven live: run 30733477609's teardown logged "Network ... Resource is still in use" without failing. Added a docker network disconnect in "Tear down e2e stack", before the destructive down -v, guarded to no-op when the runner was never attached. Two new tests cover both branches. Gates on .200 (patch-transfer, sha256-verified): targeted suite 33 passed; full tests/ci/+tests/unit/docker/ selector 1495 passed/3 skipped (login-shell PATH with /opt/homebrew/bin); ruff format/check clean; mypy clean except 2 pre-existing unused type:ignore comments at test_monitor_alert_emitter_integration.py (confirmed pre-existing via git stash comparison, unrelated to this diff); actionlint 60 pre-existing findings, zero new; pre-commit run --files clean. OMN-15567 * fix(OMN-15567): remediation round 2 -- judge projection-api DELEGATION_ROUTING_TIERS_PATH opt-out, revert false kafka marker, POSTGRES_PORT re-emit, observable network-disconnect verification Fixes 4 defects found by an independent adversarial verifier on head 4ef8477/55ec3ab: 1. BLOCKER: required CI Summary was FAILURE at PR head due to tests/integration/infra/test_judge_compose_render.py::test_judge_lane_delegation_routing_tiers_path_binding. Root cause confirmed dev-inherited (byte-identical compose/test files between HEAD and origin/dev, reproduces on a clean origin/dev checkout): OMN-15628 PR #2621 added DELEGATION_ROUTING_TIERS_PATH to docker-compose.judge.yml x-judge-runtime-env without the compensating opt-out that docker-compose.infra.yml carries on its own projection-api. Fixed by adding the same DELEGATION_ROUTING_TIERS_PATH: "" override to judge.yml projection-api, mirroring infra.yml. Filed nothing new for this -- OMN-15628 already tracks the area, cited in the compose comment. 2. tests/integration/test_monitor_alert_emitter_integration.py: reverted the pytestmark = pytest.mark.kafka added last round on a disproven claim (every MonitorAlertEmitter(...) construction in the file is wrapped in patch("confluent_kafka.Producer", ...), and the one unpatched construction deletes KAFKA_BOOTSTRAP_SERVERS first so the emitter self-disables before constructing anything -- 8 passed in 0.15s standalone, 8 passed under the CI regular-split marker filter). Live full-suite repro on .200 against a clean origin/dev checkout reproduced the same worker-death signature ("node down: Not properly terminated") on an unrelated Postgres concurrency test with zero Kafka involvement, disproving the Kafka-specific attribution. Filed OMN-15658 to track the real (still-unidentified) root cause; out of OMN-15567 scope. 3. .github/workflows/nightly-integration.yml: POSTGRES_PORT was set once to the ephemeral host-published port and never re-emitted in the Docker-DNS/gateway host-resolution rewrite, leaving it paired with the wrong host. Added the missing echo alongside the other resolved_* keys. 4. .github/workflows/nightly-integration.yml: the network-disconnect teardown fix from last round swallowed its own exit code silently (2>/dev/null || true), reproducing the exact silence class the fix existed to close. Added a post-disconnect docker network inspect verification that prints a non-fatal ::error:: annotation if the runner container is still attached, with 2 new stub-driven tests covering both the surfaced-error and verified-clean paths. [OMN-15567] * test(OMN-15567): assert postgres port re-emits in docker DNS path * fix(OMN-15567): keep live kafka tests out of standard CI shard * fix(OMN-15567): exclude live verification kafka tests from standard shard
This was referenced Aug 22, 2026
Merged
github-actions Bot
pushed a commit
that referenced
this pull request
Aug 27, 2026
…ial capability from the sweep runner (#2938) The evidence-autoclose sweep records only per-ticket COUNTS. When CI and a local dod_verify disagree on the same ticket and the same contract, the run log cannot say why, because the per-check detail goes to capture files this job never publishes. That is residual #2 of the 2026-08-27 beta DoD-verification status doc, and "which verifier is authoritative" gates `--apply`. Adds an opt-in `diagnose_tickets` workflow_dispatch input. When empty — the value every scheduled run uses — the step is skipped entirely and costs nothing. When set, it runs BEFORE the sweep (so a sweep timeout cannot swallow it) and prints: 1. what this run's GitHub credential can actually READ, per repo the bound contracts cite. Capability only; no token value is printed or derived. 2. the full per-check table for each named ticket — id, status, message — plus the OCC governance ref, refresh outcome and resolved sha. The ticket list is passed through env and format-validated rather than interpolated into the shell (OMN-16323). Refs: OMN-16788
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
🤖 Installing Claude Code GitHub App
This PR adds a GitHub Actions workflow that enables Claude Code integration in our repository.
What is Claude Code?
Claude Code is an AI coding agent that can help with:
How it works
Once this PR is merged, we'll be able to interact with Claude by mentioning @claude in a pull request or issue comment.
Once the workflow is triggered, Claude will analyze the comment and surrounding context, and execute on the request in a GitHub action.
Important Notes
Security
There's more information in the Claude Code action repo.
After merging this PR, let's try mentioning @claude in a comment on any PR to get started!