Skip to content

Add Claude Code GitHub Workflow - #2

Merged
jonahgabriel merged 2 commits into
feature/postgres-adapter-nodefrom
add-claude-github-actions-1757596353634
Sep 11, 2025
Merged

jonahgabriel merged 2 commits into
feature/postgres-adapter-nodefrom
add-claude-github-actions-1757596353634

Conversation

@jonahgabriel

Copy link
Copy Markdown
Collaborator

🤖 Installing Claude Code GitHub App

This PR adds a GitHub Actions workflow that enables Claude Code integration in our repository.

What is Claude Code?

Claude Code is an AI coding agent that can help with:

  • Bug fixes and improvements
  • Documentation updates
  • Implementing new features
  • Code reviews and suggestions
  • Writing tests
  • And more!

How it works

Once this PR is merged, we'll be able to interact with Claude by mentioning @claude in a pull request or issue comment.
Once the workflow is triggered, Claude will analyze the comment and surrounding context, and execute on the request in a GitHub action.

Important Notes

  • This workflow won't take effect until this PR is merged
  • @claude mentions won't work until after the merge is complete
  • The workflow runs automatically whenever Claude is mentioned in PR or issue comments
  • Claude gets access to the entire PR or issue context including files, diffs, and previous comments

Security

  • Our Anthropic API key is securely stored as a GitHub Actions secret
  • Only users with write access to the repository can trigger the workflow
  • All Claude runs are stored in the GitHub Actions run history
  • Claude's default tools are limited to reading/writing files and interacting with our repo by creating comments, branches, and commits.
  • We can add more allowed tools by adding them to the workflow file like:
allowed_tools: Bash(npm install),Bash(npm run build),Bash(npm run lint),Bash(npm run test)

There's more information in the Claude Code action repo.

After merging this PR, let's try mentioning @claude in a comment on any PR to get started!

@jonahgabriel
jonahgabriel merged commit 1500d98 into feature/postgres-adapter-node Sep 11, 2025
@jonahgabriel
jonahgabriel deleted the add-claude-github-actions-1757596353634 branch September 11, 2025 13:12
jonahgabriel added a commit that referenced this pull request Apr 20, 2026
…cripts

Four findings from the CodeRabbit review on PR #1352, all legitimate
correctness improvements to pre-existing behavior that's now in-scope
because we're already touching these files.

- CR #1, #4: yaml.safe_load may return None or a scalar; guard with
  isinstance check and fail fast with type-of-value in the message.
- CR #2 (MAJOR): missing top-level subscription arrays (READ_MODEL_TOPICS,
  EXPECTED_TOPICS) were a warning + silent pass. A rename or deletion of
  either array would silently succeed — exactly the breakage this gate
  exists to catch. Add required=True kwarg on top-level calls; recursive
  spread lookups still fall back to topics.ts with a warning.
- CR #3 (MAJOR): the parity check only walked consumer -> registry. A
  newly-declared registry topic that was never wired into READ_MODEL_TOPICS
  or EXPECTED_TOPICS passed the gate. Add a reverse check that every
  registry omniclaude evt topic is covered by both consumer arrays.

Tests: four new unit tests cover required-array failure, non-dict registry
rejection (both scripts), and reverse-parity failure. All 10 tests pass.
github-merge-queue Bot pushed a commit that referenced this pull request Apr 20, 2026
…6] (#1352)

* chore(scripts): relocate topic-parity scripts from omni_home [OMN-9286]

omni_home/scripts/ is blocked by the no-functional-code pre-commit hook,
which rejects any .py/.sh file in that directory. Two pre-existing scripts
(check-topic-parity.py, sync-topic-registry.py — PRs #50/#51, 2026-03-13)
violated this and were blocking unrelated docs-only PRs. Relocating to
omnibase_infra/scripts/ per the OMN-4922 pattern (pull-all.sh).

Changes:
* Copy both scripts to omnibase_infra/scripts/ preserving exec bits
* Replace module-level global state with OMNI_HOME env var + ModelTopicParityPaths
* Add SPDX headers and satisfy mypy --strict + ruff (5 pre-existing PLW0603
  + 7 missing-type-arg violations fixed in the move)
* Add tests/scripts/test_topic_parity_scripts.py covering shebang, SPDX,
  argparse surface, and OMNI_HOME resolution

Companion omni_home PR will delete the originals and repoint the CI
workflow (.github/workflows/topic-parity.yml) at the new location.

* fix(scripts): address CodeRabbit findings on relocated topic-parity scripts

Four findings from the CodeRabbit review on PR #1352, all legitimate
correctness improvements to pre-existing behavior that's now in-scope
because we're already touching these files.

- CR #1, #4: yaml.safe_load may return None or a scalar; guard with
  isinstance check and fail fast with type-of-value in the message.
- CR #2 (MAJOR): missing top-level subscription arrays (READ_MODEL_TOPICS,
  EXPECTED_TOPICS) were a warning + silent pass. A rename or deletion of
  either array would silently succeed — exactly the breakage this gate
  exists to catch. Add required=True kwarg on top-level calls; recursive
  spread lookups still fall back to topics.ts with a warning.
- CR #3 (MAJOR): the parity check only walked consumer -> registry. A
  newly-declared registry topic that was never wired into READ_MODEL_TOPICS
  or EXPECTED_TOPICS passed the gate. Add a reverse check that every
  registry omniclaude evt topic is covered by both consumer arrays.

Tests: four new unit tests cover required-array failure, non-dict registry
rejection (both scripts), and reverse-parity failure. All 10 tests pass.

* fix(sync-topic-registry): per-entry validation + JSDoc escape

Two follow-up CodeRabbit findings on the first fix commit:

- CR-minor: load_registry accepted any shape for topics entries; a dict
  missing 'topic' or both 'event_type'/'topic_base_constant' would raise
  a raw KeyError downstream instead of a structured exit-2 error with
  the offending index. Validate each entry's shape on load.

- CR-major: descriptions were injected verbatim into /** ... */ JSDoc.
  A description containing '*/' or a newline would break the generated
  TypeScript. Escape '*/' to '*\\/' and collapse newlines to spaces.

Tests: two new unit tests cover each case. All 12 tests pass.

* test(topic-parity): strengthen JSDoc-escape assertion per CR feedback

CodeRabbit flagged that the previous test only filtered lines starting
with /** and never inspected the full /** ... */ block body, making the
*/ check vacuous. Parse complete JSDoc blocks with a regex so the
assertion actually verifies the escape (and that newlines are
collapsed).

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
jonahgabriel added a commit that referenced this pull request May 10, 2026
…p config

Two CodeRabbit findings + Integration Test Coverage gate:

1. ModelDelegationRequest: add @model_validator(mode='after') rejecting
   output_schema_key set without compliance_budget. The compliance loop's
   evaluator requires both — catching this at model construction prevents a
   downstream assertion crash in the workflow handler when the orchestrator
   first sees the inference response. (CodeRabbit #2)

2. tests/integration/delegation/test_compliance_loop_wiring_integration.py:
   exercise HandlerDelegationWorkflow's compliance-loop path against unmocked
   omnimarket schema-registry / schema-repair / budget-policy handlers and
   prove the cross-package contract holds end-to-end (3 tests). Closes the
   Integration Test Coverage gap.

Note: CodeRabbit finding #1 (don't require routing_decision for every ROUTED
workflow in handle_inference_response) is already addressed by the existing
'if workflow.routing_decision is None: return []' guard at line 389.

Evidence-Source: OCC#907
Evidence-Ticket: OMN-10794
jonahgabriel added a commit that referenced this pull request May 10, 2026
…kflow (#1560)

* feat(OMN-10794): wire HandlerComplianceLoop into HandlerDelegationWorkflow

Wave 3 / Task 5 of the tokens-to-compliance epic.

Adds optional ``output_schema_key`` and ``compliance_budget`` fields to
ModelDelegationRequest. When set, HandlerDelegationWorkflow.handle_inference_response
invokes HandlerComplianceLoop per attempt:

* compliant or budget ABORT → record the attempt, transition ROUTED → INFERENCE_COMPLETED, forward to the quality gate (terminal event carries the running tokens_to_compliance + compliance_attempts)
* non-compliant + budget CONTINUE → emit a fresh ModelInferenceIntent carrying the repair prompt and stay in ROUTED (self-loop), incrementing compliance_attempts for the next iteration

DelegationWorkflowState gains ``compliance_attempts`` and ``accumulated_tokens``;
the FSM transition table allows ROUTED → ROUTED for the repair re-prompt path
(both in code _VALID_TRANSITIONS and in contract.yaml). Bumps node contract
version to 0.3.0 and updates the FSM transition documentation.

The compliance counters are populated onto the terminal ModelDelegationResult
and ModelTaskDelegatedEvent so the omnimarket projection (OMN-10793) and the
omniclaude sqlite_adapter (OMN-10789) can write them to delegation_events.

Backwards compatibility: legacy callers that omit ``output_schema_key`` get
the existing single-attempt path unchanged. Their terminal event carries
``compliance_attempts=1`` and ``tokens_to_compliance`` equal to that single
attempt's total_tokens — semantically equivalent to first-try success.

Tests: 9 new unit tests prove first-try compliance, repair-on-failure with
self-loop, two-attempt token accumulation, budget ABORT path, and FSM
self-loop legality. 32 existing orchestrator tests still pass. mypy strict
clean. pre-commit clean.

Evidence-Source: OCC#905
Evidence-Ticket: OMN-10794

* fix(OMN-10794): integration test + model_validator for compliance-loop config

Two CodeRabbit findings + Integration Test Coverage gate:

1. ModelDelegationRequest: add @model_validator(mode='after') rejecting
   output_schema_key set without compliance_budget. The compliance loop's
   evaluator requires both — catching this at model construction prevents a
   downstream assertion crash in the workflow handler when the orchestrator
   first sees the inference response. (CodeRabbit #2)

2. tests/integration/delegation/test_compliance_loop_wiring_integration.py:
   exercise HandlerDelegationWorkflow's compliance-loop path against unmocked
   omnimarket schema-registry / schema-repair / budget-policy handlers and
   prove the cross-package contract holds end-to-end (3 tests). Closes the
   Integration Test Coverage gap.

Note: CodeRabbit finding #1 (don't require routing_decision for every ROUTED
workflow in handle_inference_response) is already addressed by the existing
'if workflow.routing_decision is None: return []' guard at line 389.

Evidence-Source: OCC#907
Evidence-Ticket: OMN-10794
jonahgabriel added a commit that referenced this pull request Jul 27, 2026
… + wire live headroom check (#2496)

Root cause of the recurring partition-cap regression: docker-compose.stability-test.yml
overrides both the redpanda service command AND the redpanda-partition-cap init
service with a hardcoded topic_partitions_per_shard value. deploy-runtime.sh's
warm_broker_topic_provisioning() force-recreates the redpanda-partition-cap one-shot
on EVERY redeploy of the lane (including a --no-deps restart-only deploy that never
touches the volume), re-running whatever value is hardcoded there -- silently
overwriting any live rpk cluster config set bump made outside the committed config.
The lane also silently dropped the base redpanda.yaml startup flag entirely (belt #2
of the documented 3-belt cap design was never applied to this lane).

(a) Persist the cap: pin topic_partitions_per_shard=15000 in both belts declared in
    docker-compose.stability-test.yml (the redpanda service's restored --set flag and
    the redpanda-partition-cap service's rpk cluster config set call), so every
    redeploy of this lane carries the raised cap instead of resetting to the 7000
    default. 15000 gives durable headroom over the last observed live usage
    (7046-7047 partitions across the two recorded regressions) without requiring the
    topic-retirement audit up front.

(b) Headroom visibility: add a partition-headroom check (check_partition_headroom) to
    the EXISTING stability-lane health gate (verify_stability_refresh.py, already
    invoked by refresh_stability_lane.sh on every lane refresh) -- no new standalone
    script/dashboard. Queries live topic_partitions_per_shard + summed live partition
    count; at/over cap is a real, checked FAIL (rpk cluster health alone never sees
    this); crossing an 80% warn threshold is visible in the report but does not block
    an otherwise-healthy refresh.

(c) Topic-retirement audit and root-causing the lane's disproportionate topic
    accumulation (OMN-14013 DoD items 1/4) are explicitly OUT of scope here --
    documented follow-up, tracked on the ticket itself.

Tests: RED proven against the un-pinned 7000 value (regex-based value extraction,
not substring matching -- a plain substring check silently passes against this same
diff's own forensic comments mentioning superseded values), GREEN after the fix.
32 unit tests for the new headroom check including the exact live incident numbers
(7046/7000) as a named regression test. Live verify deferred until post-acceptance --
this PR does not touch the live stability-test lane, restart any broker, or run rpk
against 100.109.203.94:39092.

Closes OMN-14013

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
jonahgabriel added a commit that referenced this pull request Aug 2, 2026
…y, Docker-DNS host cross-wiring, network-leak teardown

Adversarial verifier found six defects in the prior push (head 4ef8477):

1. CI Summary FAILURE (required check on omnibase_infra/dev) -- Tests
   (Split 2/2) was cancelled at the 15-min job timeout 3x on this
   branch, misdiagnosed in the PR comment as an unlucky pytest-split
   bucket. Root cause: tests/integration/test_monitor_alert_emitter_integration.py
   is @pytest.mark.integration-only (missing @pytest.mark.kafka), so
   CI regular splits `-m "not slow and not chaos and not kafka and not
   performance"` never deselect it, unlike its sibling files
   (test_consumer_health_pipeline.py, test_runtime_log_bridge_pipeline.py,
   test_runtime_health_monitor.py) which already carry the kafka marker
   for exactly this reason. Added pytestmark = pytest.mark.kafka at
   module level, matching the established convention.

2. tests/unit/docker/test_nightly_e2e_runner_connectivity.py:171 was
   permanently RED on the mandated gate host (.200): nightly-integration.yml
   used `mapfile`, a bash 4+ builtin, and macOS ships bash 3.2.57 at
   /bin/bash (no newer bash for GPLv3 licensing reasons). Rewrote
   "Detect runner network topology" to use a `while read` loop instead
   -- bash-3-compatible, identical behavior. New regression test pins
   /bin/bash explicitly so a newer bash earlier on PATH cannot hide a
   future regression back to a bash-4-only construct.

3. PR body's claimed pre-push governed selector result ("1491 passed, 3
   skipped") was not reproducible via raw non-login ssh (matches the
   verifier's invocation): 5 of the 6 failures are pre-existing PATH
   artifacts of a non-login shell lacking /opt/homebrew/bin (uv/timeout
   not on PATH) -- confirmed identical under origin/dev's own tree, not
   introduced by this diff. The 6th (mapfile) is fixed by #2 above and
   now passes under both a login-shell PATH and a PATH forced to
   exclude /opt/homebrew/bin (i.e. raw /bin/bash 3.2.57).

4/5. "Resolve reachable e2e connectivity host"'s Docker DNS branch (the
   ONLY branch live traffic actually took, per the run the PR cited as
   its own success evidence) cross-wired REDPANDA_ADVERTISE_HOST and
   OMNI_INFRA_HOST to the *postgres* container's Docker-DNS name instead
   of redpanda's -- the only branch where postgres and redpanda have
   DIFFERENT hostnames rather than one shared host disambiguated by
   port. Split resolved_host (postgres) from a new resolved_kafka_host
   (redpanda) and wired REDPANDA_ADVERTISE_HOST/OMNI_INFRA_HOST from the
   latter. New test drives the previously-untested Docker DNS branch via
   stubbed docker network connect/docker inspect.

6. "Resolve reachable e2e connectivity host" attaches the long-lived
   self-hosted runner container to the run's compose network via
   docker network connect and nothing ever disconnected it -- proven
   live: run 30733477609's teardown logged "Network ... Resource is
   still in use" without failing. Added a docker network disconnect
   in "Tear down e2e stack", before the destructive down -v, guarded
   to no-op when the runner was never attached. Two new tests cover both
   branches.

Gates on .200 (patch-transfer, sha256-verified): targeted suite 33
passed; full tests/ci/+tests/unit/docker/ selector 1495 passed/3
skipped (login-shell PATH with /opt/homebrew/bin); ruff format/check
clean; mypy clean except 2 pre-existing unused type:ignore comments at
test_monitor_alert_emitter_integration.py (confirmed pre-existing via
git stash comparison, unrelated to this diff); actionlint 60
pre-existing findings, zero new; pre-commit run --files clean.

OMN-15567
jonahgabriel added a commit that referenced this pull request Aug 2, 2026
…y, Docker-DNS host cross-wiring, network-leak teardown

Adversarial verifier found six defects in the prior push (head 4ef8477):

1. CI Summary FAILURE (required check on omnibase_infra/dev) -- Tests
   (Split 2/2) was cancelled at the 15-min job timeout 3x on this
   branch, misdiagnosed in the PR comment as an unlucky pytest-split
   bucket. Root cause: tests/integration/test_monitor_alert_emitter_integration.py
   is @pytest.mark.integration-only (missing @pytest.mark.kafka), so
   CI regular splits `-m "not slow and not chaos and not kafka and not
   performance"` never deselect it, unlike its sibling files
   (test_consumer_health_pipeline.py, test_runtime_log_bridge_pipeline.py,
   test_runtime_health_monitor.py) which already carry the kafka marker
   for exactly this reason. Added pytestmark = pytest.mark.kafka at
   module level, matching the established convention.

2. tests/unit/docker/test_nightly_e2e_runner_connectivity.py:171 was
   permanently RED on the mandated gate host (.200): nightly-integration.yml
   used `mapfile`, a bash 4+ builtin, and macOS ships bash 3.2.57 at
   /bin/bash (no newer bash for GPLv3 licensing reasons). Rewrote
   "Detect runner network topology" to use a `while read` loop instead
   -- bash-3-compatible, identical behavior. New regression test pins
   /bin/bash explicitly so a newer bash earlier on PATH cannot hide a
   future regression back to a bash-4-only construct.

3. PR body's claimed pre-push governed selector result ("1491 passed, 3
   skipped") was not reproducible via raw non-login ssh (matches the
   verifier's invocation): 5 of the 6 failures are pre-existing PATH
   artifacts of a non-login shell lacking /opt/homebrew/bin (uv/timeout
   not on PATH) -- confirmed identical under origin/dev's own tree, not
   introduced by this diff. The 6th (mapfile) is fixed by #2 above and
   now passes under both a login-shell PATH and a PATH forced to
   exclude /opt/homebrew/bin (i.e. raw /bin/bash 3.2.57).

4/5. "Resolve reachable e2e connectivity host"'s Docker DNS branch (the
   ONLY branch live traffic actually took, per the run the PR cited as
   its own success evidence) cross-wired REDPANDA_ADVERTISE_HOST and
   OMNI_INFRA_HOST to the *postgres* container's Docker-DNS name instead
   of redpanda's -- the only branch where postgres and redpanda have
   DIFFERENT hostnames rather than one shared host disambiguated by
   port. Split resolved_host (postgres) from a new resolved_kafka_host
   (redpanda) and wired REDPANDA_ADVERTISE_HOST/OMNI_INFRA_HOST from the
   latter. New test drives the previously-untested Docker DNS branch via
   stubbed docker network connect/docker inspect.

6. "Resolve reachable e2e connectivity host" attaches the long-lived
   self-hosted runner container to the run's compose network via
   docker network connect and nothing ever disconnected it -- proven
   live: run 30733477609's teardown logged "Network ... Resource is
   still in use" without failing. Added a docker network disconnect
   in "Tear down e2e stack", before the destructive down -v, guarded
   to no-op when the runner was never attached. Two new tests cover both
   branches.

Gates on .200 (patch-transfer, sha256-verified): targeted suite 33
passed; full tests/ci/+tests/unit/docker/ selector 1495 passed/3
skipped (login-shell PATH with /opt/homebrew/bin); ruff format/check
clean; mypy clean except 2 pre-existing unused type:ignore comments at
test_monitor_alert_emitter_integration.py (confirmed pre-existing via
git stash comparison, unrelated to this diff); actionlint 60
pre-existing findings, zero new; pre-commit run --files clean.

OMN-15567
jonahgabriel added a commit that referenced this pull request Aug 2, 2026
…y, Docker-DNS host cross-wiring, network-leak teardown

Adversarial verifier found six defects in the prior push (head 4ef8477):

1. CI Summary FAILURE (required check on omnibase_infra/dev) -- Tests
   (Split 2/2) was cancelled at the 15-min job timeout 3x on this
   branch, misdiagnosed in the PR comment as an unlucky pytest-split
   bucket. Root cause: tests/integration/test_monitor_alert_emitter_integration.py
   is @pytest.mark.integration-only (missing @pytest.mark.kafka), so
   CI regular splits `-m "not slow and not chaos and not kafka and not
   performance"` never deselect it, unlike its sibling files
   (test_consumer_health_pipeline.py, test_runtime_log_bridge_pipeline.py,
   test_runtime_health_monitor.py) which already carry the kafka marker
   for exactly this reason. Added pytestmark = pytest.mark.kafka at
   module level, matching the established convention.

2. tests/unit/docker/test_nightly_e2e_runner_connectivity.py:171 was
   permanently RED on the mandated gate host (.200): nightly-integration.yml
   used `mapfile`, a bash 4+ builtin, and macOS ships bash 3.2.57 at
   /bin/bash (no newer bash for GPLv3 licensing reasons). Rewrote
   "Detect runner network topology" to use a `while read` loop instead
   -- bash-3-compatible, identical behavior. New regression test pins
   /bin/bash explicitly so a newer bash earlier on PATH cannot hide a
   future regression back to a bash-4-only construct.

3. PR body's claimed pre-push governed selector result ("1491 passed, 3
   skipped") was not reproducible via raw non-login ssh (matches the
   verifier's invocation): 5 of the 6 failures are pre-existing PATH
   artifacts of a non-login shell lacking /opt/homebrew/bin (uv/timeout
   not on PATH) -- confirmed identical under origin/dev's own tree, not
   introduced by this diff. The 6th (mapfile) is fixed by #2 above and
   now passes under both a login-shell PATH and a PATH forced to
   exclude /opt/homebrew/bin (i.e. raw /bin/bash 3.2.57).

4/5. "Resolve reachable e2e connectivity host"'s Docker DNS branch (the
   ONLY branch live traffic actually took, per the run the PR cited as
   its own success evidence) cross-wired REDPANDA_ADVERTISE_HOST and
   OMNI_INFRA_HOST to the *postgres* container's Docker-DNS name instead
   of redpanda's -- the only branch where postgres and redpanda have
   DIFFERENT hostnames rather than one shared host disambiguated by
   port. Split resolved_host (postgres) from a new resolved_kafka_host
   (redpanda) and wired REDPANDA_ADVERTISE_HOST/OMNI_INFRA_HOST from the
   latter. New test drives the previously-untested Docker DNS branch via
   stubbed docker network connect/docker inspect.

6. "Resolve reachable e2e connectivity host" attaches the long-lived
   self-hosted runner container to the run's compose network via
   docker network connect and nothing ever disconnected it -- proven
   live: run 30733477609's teardown logged "Network ... Resource is
   still in use" without failing. Added a docker network disconnect
   in "Tear down e2e stack", before the destructive down -v, guarded
   to no-op when the runner was never attached. Two new tests cover both
   branches.

Gates on .200 (patch-transfer, sha256-verified): targeted suite 33
passed; full tests/ci/+tests/unit/docker/ selector 1495 passed/3
skipped (login-shell PATH with /opt/homebrew/bin); ruff format/check
clean; mypy clean except 2 pre-existing unused type:ignore comments at
test_monitor_alert_emitter_integration.py (confirmed pre-existing via
git stash comparison, unrelated to this diff); actionlint 60
pre-existing findings, zero new; pre-commit run --files clean.

OMN-15567
jonahgabriel added a commit that referenced this pull request Aug 2, 2026
…r SIGILL confound clears (#2623)

* fix(OMN-15567): resolve nightly e2e runner-topology connectivity after redpanda SIGILL confound clears

The nightly e2e stack never reached pytest for 8/8 nights (redpanda exited
132/SIGILL). Re-observed post-OMN-15565 on workflow_dispatch run 30681782952
and scheduled run 30685774570: redpanda now starts Healthy in ~5s and pytest
runs to completion (2471 passed/444 skipped/1 xfailed/6 failed/13 errors in
848.45s) -- the SIGILL was the OMN-15565 lane-collision confound, not a real
defect, and is gone now that the stack is genuinely isolated.

The 13 errors are the second defect this ticket targets: the self-hosted
omnibase-ci runner does not share a network namespace with the Docker host
(same topology reusable-runtime-boot.yml:250-270 already documents), so
"localhost:<host-published-port>" is unreachable from the runner process even
when the port is the freshly-derived, correctly-read dynamic port. Mirrors
reusable-runtime-boot.yml's existing, proven detection/probe pattern:

- "Detect runner network topology": identify a containerized runner and
  resolve the Docker host gateway, so redpanda's external listener can
  advertise it instead of always defaulting to "localhost".
- "Resolve reachable e2e connectivity host": after the stack is Healthy,
  probe Docker DNS -> localhost -> Docker host gateway in that order and
  export KAFKA_BOOTSTRAP_SERVERS/OMNIBASE_INFRA_DB_URL/INTEGRATION_POSTGRES_HOST
  etc. from whichever actually works; fail closed with diagnostics if none do.
  Both new steps carry their own live-.201 guard, since they compute values
  the workflow-level static guard never sees.

Also fixes two integration tests that hardcoded BOOTSTRAP_SERVERS =
"localhost:19092", silently ignoring the CI-derived KAFKA_BOOTSTRAP_SERVERS
regardless of topology -- matching the pattern already used correctly by
test_runtime_health_monitor.py.

OMN-15567

* chore(OMN-15567): retrigger CI after OCC companion OCC#5933 merged

* chore(OMN-15567): retrigger CI with complete Evidence-Source/Evidence-Ticket PR body

* test(OMN-15567): bound subprocess.run calls with explicit timeouts (hardening)

* fix(OMN-15567): remediation round -- kafka marker, mapfile portability, Docker-DNS host cross-wiring, network-leak teardown

Adversarial verifier found six defects in the prior push (head 4ef8477):

1. CI Summary FAILURE (required check on omnibase_infra/dev) -- Tests
   (Split 2/2) was cancelled at the 15-min job timeout 3x on this
   branch, misdiagnosed in the PR comment as an unlucky pytest-split
   bucket. Root cause: tests/integration/test_monitor_alert_emitter_integration.py
   is @pytest.mark.integration-only (missing @pytest.mark.kafka), so
   CI regular splits `-m "not slow and not chaos and not kafka and not
   performance"` never deselect it, unlike its sibling files
   (test_consumer_health_pipeline.py, test_runtime_log_bridge_pipeline.py,
   test_runtime_health_monitor.py) which already carry the kafka marker
   for exactly this reason. Added pytestmark = pytest.mark.kafka at
   module level, matching the established convention.

2. tests/unit/docker/test_nightly_e2e_runner_connectivity.py:171 was
   permanently RED on the mandated gate host (.200): nightly-integration.yml
   used `mapfile`, a bash 4+ builtin, and macOS ships bash 3.2.57 at
   /bin/bash (no newer bash for GPLv3 licensing reasons). Rewrote
   "Detect runner network topology" to use a `while read` loop instead
   -- bash-3-compatible, identical behavior. New regression test pins
   /bin/bash explicitly so a newer bash earlier on PATH cannot hide a
   future regression back to a bash-4-only construct.

3. PR body's claimed pre-push governed selector result ("1491 passed, 3
   skipped") was not reproducible via raw non-login ssh (matches the
   verifier's invocation): 5 of the 6 failures are pre-existing PATH
   artifacts of a non-login shell lacking /opt/homebrew/bin (uv/timeout
   not on PATH) -- confirmed identical under origin/dev's own tree, not
   introduced by this diff. The 6th (mapfile) is fixed by #2 above and
   now passes under both a login-shell PATH and a PATH forced to
   exclude /opt/homebrew/bin (i.e. raw /bin/bash 3.2.57).

4/5. "Resolve reachable e2e connectivity host"'s Docker DNS branch (the
   ONLY branch live traffic actually took, per the run the PR cited as
   its own success evidence) cross-wired REDPANDA_ADVERTISE_HOST and
   OMNI_INFRA_HOST to the *postgres* container's Docker-DNS name instead
   of redpanda's -- the only branch where postgres and redpanda have
   DIFFERENT hostnames rather than one shared host disambiguated by
   port. Split resolved_host (postgres) from a new resolved_kafka_host
   (redpanda) and wired REDPANDA_ADVERTISE_HOST/OMNI_INFRA_HOST from the
   latter. New test drives the previously-untested Docker DNS branch via
   stubbed docker network connect/docker inspect.

6. "Resolve reachable e2e connectivity host" attaches the long-lived
   self-hosted runner container to the run's compose network via
   docker network connect and nothing ever disconnected it -- proven
   live: run 30733477609's teardown logged "Network ... Resource is
   still in use" without failing. Added a docker network disconnect
   in "Tear down e2e stack", before the destructive down -v, guarded
   to no-op when the runner was never attached. Two new tests cover both
   branches.

Gates on .200 (patch-transfer, sha256-verified): targeted suite 33
passed; full tests/ci/+tests/unit/docker/ selector 1495 passed/3
skipped (login-shell PATH with /opt/homebrew/bin); ruff format/check
clean; mypy clean except 2 pre-existing unused type:ignore comments at
test_monitor_alert_emitter_integration.py (confirmed pre-existing via
git stash comparison, unrelated to this diff); actionlint 60
pre-existing findings, zero new; pre-commit run --files clean.

OMN-15567

* fix(OMN-15567): remediation round 2 -- judge projection-api DELEGATION_ROUTING_TIERS_PATH opt-out, revert false kafka marker, POSTGRES_PORT re-emit, observable network-disconnect verification

Fixes 4 defects found by an independent adversarial verifier on head 4ef8477/55ec3ab:

1. BLOCKER: required CI Summary was FAILURE at PR head due to
   tests/integration/infra/test_judge_compose_render.py::test_judge_lane_delegation_routing_tiers_path_binding.
   Root cause confirmed dev-inherited (byte-identical compose/test files
   between HEAD and origin/dev, reproduces on a clean origin/dev checkout):
   OMN-15628 PR #2621 added DELEGATION_ROUTING_TIERS_PATH to
   docker-compose.judge.yml x-judge-runtime-env without the compensating
   opt-out that docker-compose.infra.yml carries on its own projection-api.
   Fixed by adding the same DELEGATION_ROUTING_TIERS_PATH: "" override to
   judge.yml projection-api, mirroring infra.yml. Filed nothing new for this
   -- OMN-15628 already tracks the area, cited in the compose comment.

2. tests/integration/test_monitor_alert_emitter_integration.py: reverted the
   pytestmark = pytest.mark.kafka added last round on a disproven claim
   (every MonitorAlertEmitter(...) construction in the file is wrapped in
   patch("confluent_kafka.Producer", ...), and the one unpatched
   construction deletes KAFKA_BOOTSTRAP_SERVERS first so the emitter
   self-disables before constructing anything -- 8 passed in 0.15s
   standalone, 8 passed under the CI regular-split marker filter). Live
   full-suite repro on .200 against a clean origin/dev checkout reproduced
   the same worker-death signature ("node down: Not properly terminated")
   on an unrelated Postgres concurrency test with zero Kafka involvement,
   disproving the Kafka-specific attribution. Filed OMN-15658 to track the
   real (still-unidentified) root cause; out of OMN-15567 scope.

3. .github/workflows/nightly-integration.yml: POSTGRES_PORT was set once to
   the ephemeral host-published port and never re-emitted in the
   Docker-DNS/gateway host-resolution rewrite, leaving it paired with the
   wrong host. Added the missing echo alongside the other resolved_* keys.

4. .github/workflows/nightly-integration.yml: the network-disconnect
   teardown fix from last round swallowed its own exit code silently
   (2>/dev/null || true), reproducing the exact silence class the fix
   existed to close. Added a post-disconnect docker network inspect
   verification that prints a non-fatal ::error:: annotation if the runner
   container is still attached, with 2 new stub-driven tests covering both
   the surfaced-error and verified-clean paths.

[OMN-15567]

* test(OMN-15567): assert postgres port re-emits in docker DNS path

* fix(OMN-15567): keep live kafka tests out of standard CI shard

* fix(OMN-15567): exclude live verification kafka tests from standard shard
github-actions Bot pushed a commit that referenced this pull request Aug 27, 2026
…ial capability from the sweep runner (#2938)

The evidence-autoclose sweep records only per-ticket COUNTS. When CI and a
local dod_verify disagree on the same ticket and the same contract, the run
log cannot say why, because the per-check detail goes to capture files this
job never publishes. That is residual #2 of the 2026-08-27 beta
DoD-verification status doc, and "which verifier is authoritative" gates
`--apply`.

Adds an opt-in `diagnose_tickets` workflow_dispatch input. When empty — the
value every scheduled run uses — the step is skipped entirely and costs
nothing. When set, it runs BEFORE the sweep (so a sweep timeout cannot
swallow it) and prints:

  1. what this run's GitHub credential can actually READ, per repo the bound
     contracts cite. Capability only; no token value is printed or derived.
  2. the full per-check table for each named ticket — id, status, message —
     plus the OCC governance ref, refresh outcome and resolved sha.

The ticket list is passed through env and format-validated rather than
interpolated into the shell (OMN-16323).

Refs: OMN-16788
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant