Skip to content

chore(OMN-16041): backmerge main lineage into dev alongside the v0.38.5 promotion - #2745

Merged
jonahgabriel merged 22 commits into
devfrom
jonah/omn-16041-backmerge-main-to-dev
Aug 18, 2026
Merged

jonahgabriel merged 22 commits into
devfrom
jonah/omn-16041-backmerge-main-to-dev

Conversation

@jonahgabriel

@jonahgabriel jonahgabriel commented Aug 14, 2026 •

Copy link
Copy Markdown
Collaborator

OMN-16041 — backmerge main lineage into dev

This is the dev-side half of the promotion that landed on main as #2744. It exists so the two branches stop re-conflicting on every promotion, and so dev does not silently regress the fixes that live only on main.


Conflict resolution — 2026-08-18

The branch had gone stale (opened 08-14, main promoted 08-15, dev moved 13 commits since). Current dev was merged in and the conflicts resolved by reading both sides' history per path rather than taking a side wholesale.

Only two files actually conflicted, and both were the same line:

File main lineage dev Decision Why
pyproject.toml 0.38.5 0.38.7 dev main is already at 0.38.6. dev's 0.38.7 is above it, so the release-identity gate is satisfied without any bump here; taking the branch's stale 0.38.5 would have pushed dev below the released version.
uv.lock 0.38.5 0.38.7 dev Same line, same reasoning — the lock's own package entry must track pyproject.toml.

Everything else merged without textual conflict. Each auto-merged path was still checked by hand against both branches, because "no conflict marker" is not the same as "correct":

File Authoritative side Reasoning
.github/workflows/deploy-gate.yml
.github/required-checks.yaml
main These pin the shared deploy-gate workflow by commit. main pins one commit from 2026-07-25 in both files. dev pins two different commits — 2026-07-20 in the workflow and 2026-07-19 in the manifest — so dev was both older and internally inconsistent with itself. Taking main moves to the newest pin and makes the two files agree again.
.github/workflows/env-parity.yml main main widened the sibling-checkout token to try the org-wide token before falling back, because the default token cannot read the private sibling repo. dev never received it and still has the narrower two-option form. Nothing on the dev side competes.
.github/workflows/artifact-reconciliation-webhook.yml main main-only CI change from the previous bridge; dev has not touched this file since the branches diverged.
scripts/validate_handler_contracts.py main The strongest case in the set. dev still carries the original version, which validates a hardcoded list of seven handlers at nodes/handlers/<name>/contract.yaml — a directory that no longer exists in either branch. main replaced that with a glob over the live descriptor location. Verified by running it on the merged tree: 5/5 contracts validated, 0 failed. dev's version cannot pass at all.
src/omnibase_infra/utils/util_runtime_packages.py main main-only; no competing dev change. See the flag below — this one is carried, but it is not clean.
the two runner-health handlers main for the env-read removal, dev for the behaviour main removed direct environment reads from both handlers. dev never got that and still describes its own reads as grandfathered — which they are only against dev's history; measured against main they read as new, which is exactly why the promotion kept conflicting here. The removal is carried. dev's behaviour is preserved, not main's: the staleness threshold stays at dev's 4500s along with the comment explaining why (an idle runner only writes its diagnostic file on its ~50-minute token refresh, so main's 900s classified idle-but-healthy runners as zombies for most of every hour).
src/omnibase_infra/observability/runner_health/model_runner_fleet_config.py both — union The only file where each side had real content. main adds three configuration fields that replace the removed environment reads; dev adds three unrelated optional sub-configs from newer work. Both sets are kept. This is the one file in the merge that is deliberately not byte-identical to main, and that is correct — main simply does not have dev's newer fields yet.

Verification

  • The eight files with no competing dev change are byte-identical to origin/main, so the next promotion has nothing left to conflict on in them.
  • The ninth is the union described above.
  • Targeted suites for the affected surfaces (runner-health observability, runner-fleet maintenance, the environment-read gate, and the runtime-package gate): 245 passed, 4 skipped. The 4 skips are a pre-existing guard for a file-locking tool that is not installed here, unrelated to this change.
  • Handler-contract validator run against the merged tree: 5/5 validated, 0 failed.
  • Full pre-commit hook chain passed on commit; the pre-push gate escalated to the full suite on its own (shared-module rule) and ran to completion before the push was accepted.

Flagged, carried but not endorsed

util_runtime_packages.py does not remove its two environment reads. It hides them: the module-level binding

_ENV_GET = vars(os)["environ"].get

reads exactly the same values, but routes around the gate, which matches on the literal spellings os.environ[, os.environ.get, os.getenv and so on. An indirection like this is invisible to that check by construction.

This is worth naming because the gate's own allowlist explicitly rejects the pattern — one of its entries reads "Formalizing the allowlist entry rather than duplicating that read through a new indirection layer." The correct shape is either a named allowlist entry or resolution through the config overlay, the way the two runner-health handlers in this same merge were fixed.

It is carried here rather than fixed, for one reason: it is already shipped on main. Reverting it on the dev side would re-open the exact promotion conflict this PR exists to close, and would do it in a file where the gate then fires. Fixing it properly means changing both branches, which is its own change with its own test surface — filed as follow-up rather than smuggled into a backmerge.

Known red, disclosed

tests/ci/test_env_parity.py::test_k8s_bound_keys_are_bound_in_compose fails here exactly as it fails on origin/dev. This PR neither introduces nor fixes it; the fix is open as #2740. Nothing here narrows, skips, or reclassifies that test.

Evidence-Source: OCC#6661
Evidence-Ticket: OMN-16041

jonahgabriel and others added 20 commits June 15, 2026 02:20
…1988)

* docs(OMN-12962): contract-store durability audit — cold-runtime census proof (#1937)

Proves cold-start contract census reconstructs from the image-bundled
filesystem manifest (HYBRID-mode bootstrap + PluginLoaderContractSource),
independent of node-registration.v1 retention. The delete-retention topic
feeds only the post-freeze dynamic listener (auto_offset_reset=latest, no
history replay). Live .201 stability-test evidence: filesystem contract_path
in manifest + truncated topic log head with intact census. No store fix
needed; residual dynamic-only gap covered by runtime_sweep sweep check.

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12970): vendor omnimarket projection node migrations into forward/nodes (#1938)

Vendors omnimarket node-owned projection migrations into the namespaced
forward-migration tree so run-forward-migrations.sh materializes them in the
dashboard projection DB (omnidash_analytics) at deploy.

Primary (OMN-12970): creates llm_call_metrics, llm_cost_aggregates, and
capability_scores in the projection DB. These were only ever created in the
omnibase_infra DB by infra migrations 031/060, so the ab-compare,
cost.token_usage, cost.summary, and capability-scores projection topics were
DEGRADED at startup ('table not found') and their dashboard panels rendered
empty.

Also re-syncs three omnimarket node migrations the vendor tree had drifted from
(node_projection_llm_routing, node_projection_overnight, node_projection_savings
/077) — sync-node-migrations.sh --check requires the full vendored tree to match
omnimarket source, and these were missing.

Companion to omnimarket PR for the same ticket (source migrations + projection
table-coverage ratchet test).

Evidence-Ticket: OMN-12970

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12965): stamp runtime image identity (version + revision) in workspace builds (#1943)

* fix(OMN-12965): stamp runtime image identity (version + revision) in workspace builds

The main runtime image stamped org.opencontainers.image.version=0.1.0 with a
blank org.opencontainers.image.revision after workspace rebuilds. A blank
identity degrades every proof packet (runtime SHA + image digest are required
citations in accepted evidence).

Root causes (three build paths under-stamped identity):
- onex up --build (cmd_up) passed only GIT_SHA; the runtime-stage OCI labels
  read VCS_REF (-> blank revision) and RUNTIME_VERSION (-> placeholder 0.1.0).
- deploy-runtime.sh passed VCS_REF but not RUNTIME_VERSION/GIT_SHA.
- Dockerfile silently allowed blank/placeholder identity in workspace mode.

Fix:
- cli._image_identity_build_args() stamps the full quad (GIT_SHA/VCS_REF/
  RUNTIME_VERSION/BUILD_DATE) and fails fast on an unresolved git revision.
- deploy-runtime.sh stamps RUNTIME_VERSION + GIT_SHA and verifies the version
  label is non-placeholder post-deploy.
- Dockerfile.runtime fails workspace builds with blank VCS_REF or placeholder
  RUNTIME_VERSION=0.1.0 (release mode unaffected).

Enforcement ratchet (same PR):
- scripts/check_runtime_image_identity.py static check, wired as pre-commit hook
  + CI gate (ci.yml).
- tests/unit/infra/test_runtime_image_identity_labels.py pins the cli helpers +
  Dockerfile guard; deploy-agent test extended for the quad.

Proven locally via throwaway docker builds: workspace+args -> populated labels;
workspace without args -> guard fails (exit 64); release without args -> 0.1.0
placeholder allowed (no regression).

Evidence-Ticket: OMN-12965

* test(OMN-12965): integration build proof for runtime image identity labels

Builds the real runtime-stage ARG/LABEL/guard block against busybox and asserts
via docker inspect: workspace+args -> populated version/revision; workspace
without args -> guard fails (exit 64); release -> placeholder allowed. Satisfies
the integration-test hard gate and makes the throwaway proof permanent.

Evidence-Ticket: OMN-12965

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12987): workspace-build sibling lock-pin preflight + provenance (#1944)

Recurrence ratchet for the 2026-06-11 stability bootstrap crash. The 11:20Z
--no-cache rebuild vendored omnibase_infra 0.37.0-dev (~2c1d672f) + core 0.42.0
even though omnimarket dev's uv.lock pinned infra 0.38.1 @ e2dbdc95 + core
0.44.0 @ c97c2c9a. The stale sibling predated the OMN-12501 Protocol-quarantine
guard, turning a latent contract defect into a fatal crash that crash-looped the
main runtime.

- check_sibling_lock_pins.py: host-side fail-fast preflight resolving expected
  sibling versions/SHAs from the consuming repo's (omnimarket) uv.lock and
  comparing them against each vendored tree. Mismatch aborts the build.
- stage_workspace.sh: emit a .build-sha marker per staged sibling (rsync drops
  .git) so the preflight and provenance can identify the vendored commit.
- deploy-runtime.sh: run the preflight after staging, before build; abort on
  mismatch. Write the comparison under sibling-repos/.sibling-lock-pins.json.
- compute_workspace_provenance.py + Dockerfile.runtime: fold expected-vs-actual
  lock_pin_comparison into build-provenance.json for deploy verifiers.

Recurrence-guard tests prove a stale infra 0.37.0 vs lock-pinned 0.38.1 fails
the preflight and matched pins pass; deploy script wiring is asserted statically.

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12988): pin stability runtime-worker to replicas 1 + census ratchet (#1942)

The base docker-compose.infra.yml defaults runtime-worker to
replicas 0 (${WORKER_REPLICAS:-0}). The stability-test lane's required
state includes a running worker (4-container census: main, effects,
worker, projection-api), but the override pinned it via an
env-interpolation default (${STABILITY_TEST_WORKER_REPLICAS:-1}) — a
silent-drop surface: a stray exported STABILITY_TEST_WORKER_REPLICAS=0
or removal of the :-1 fallback would scale the worker to 0 with zero
signal on a plain compose up/recreate.

Fix: pin docker-compose.stability-test.yml runtime-worker
deploy.replicas to the literal 1 (no env indirection).

Ratchet (recurrence guards, same PR):
- scripts/deploy-agent/tests/unit/test_runtime_worker_census.py: assert
  runtime-worker stays in the deploy-agent RUNTIME-scope census so a
  missing worker (replicas 0 => absent from docker compose ps) is a
  deploy failure, not silence; assert the override pins a literal 1.
- tests/integration/infra/test_stability_test_runtime_compose_render.py:
  assert the rendered stability worker resolves deploy.replicas == 1.
- tests/unit/infra/test_stability_test_runtime_lane.py: update the
  existing pin assertion to the literal 1.

Evidence-Ticket: OMN-12988
Config-drift family: OMN-12945

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12979): expire-bound topic completeness suppressions (#1940)

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12966): accept community-edition Infisical /api/status in provision readiness gate (#1939)

The fresh-provision readiness gate in provision-infisical.py only accepted the
enterprise {"status": "ok"} payload and rejected the community edition's
{"message": "Ok"}, returning 1 before bootstrap could run. This blocked
provisioning against the Infisical instance deployed on .201 (community edition).

Route the gate through the existing _is_infisical_ready helper (single source of
truth, already used by the already-provisioned path). Add TestMainFreshProvision-
ReadinessGate covering community/enterprise/not-ready cases.

Also adds docker/docker-compose.infisical-stability.yml: an ADDITIVE Infisical
compose project for the stability-test lane (joins the existing network as
external, reuses stability postgres/valkey, no lane mutation), since the lane
overlays disable the in-lane Infisical service via *-disabled profile overrides.

P1.2b-A: Infisical now reachable from the stability runtime/effect containers; a
known secret (OMN_12966_PROBE) seeds and resolves end-to-end via the machine-
identity universal-auth path, verified from inside the effects container.

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-12958): volume-config drift gate + runtime config provenance (#1941)

* feat(OMN-12958): volume-config drift gate + runtime provenance

Compute config provenance (path + sha256) for the runtime-rendered Bifrost
delegation contract; the deployed volume copy survives rebuilds and silently
diverges from packaged source (two competing authorities, OMN-12945).

- runtime/config_provenance.py: ModelConfigProvenance + drift classification,
  sidecar JSON writer (read by sweep + proof packets)
- runtime/health/health_config_provenance.py: drift -> degraded health
- render entrypoint logs provenance line + writes sidecar on every boot
- docs/runbooks/volume-config-drift-and-reseed.md: ledgered re-seed procedure
- validation exemption for config_name (logical identifier, not entity ref)

No live volume mutation: re-seed is an operator deploy step (deploy_pending).

* test(OMN-12958): cover volume config drift reseed flow

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-12957): wire runtime-profile validator + core registry parity guard (#1945)

* feat(OMN-12957): wire runtime-profile validator + core registry parity guard

- Remove the import-time RuntimeError drift raise in runtime_profile.py (a hard
  raise on core/infra version skew would crash the kernel at import); the parity
  invariant is enforced by test_profiles_match_core_registry instead.
- Add tests: _PROFILES keys == omnibase_core REGISTERED_RUNTIME_PROFILES, and
  every CONSUMER_ATTACHED_RUNTIME_PROFILES profile loads as a real lane.
- Wire omnibase_core.validation.validator_runtime_profiles as a pre-commit hook
  + validator-runtime-profiles.yml CI gate on infra contracts.
- Freeze 19 pre-existing violators in validation/runtime_profiles_allowlist.yaml
  (discovered by repo-root walk; drain via OMN-12982). Blocks NEW orphans.

Requires the omnibase_core pin to include OMN-12957's validator (new rules).

Evidence-Ticket: OMN-12957
Evidence-Source: 5463fbaf819409d4fb7f491dd4f276f10d869eba

* ci(OMN-12957): pass runtime profile allowlist to validator

* test(OMN-12957): cover runtime profile registry parity

* fix(OMN-12957): keep runtime profile allowlist under config

* fix(OMN-12957): pin core runtime profile registry

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12973): widen migration-gate healthcheck start_period + ratchet (#1950)

P2.8: classify the prod migration-gate UNHEALTHY-then-self-resolved incident.

Classification: idle-one-shot-mis-modeled = NO. The migration-gate is a
correctly-modeled long-running sentinel (entrypoint `while true; sleep 3600`)
whose healthcheck continuously polls db_metadata.migrations_complete via
check_migrations_complete.sh. The container flipped UNHEALTHY transiently because
its healthcheck start_period (10s) was far shorter than the real cold-volume
migration window (~116s: prod gate started 09:35:22, intelligence-migration
finished 09:37:18). Past the 10s grace window the still-failing probe was
reported UNHEALTHY until migrations completed, then self-resolved — no fault.

Fix: raise migration-gate healthcheck start_period 10s -> 180s in both the
authoritative catalog manifest (docker/catalog/services/migration-gate.yaml,
flows into the generated compose) and the hand-maintained
docker/docker-compose.infra.yml that deploy-runtime.sh applies to .201, so a
still-applying gate stays in `health: starting` instead of flipping UNHEALTHY.

Ratchet (enforcement, not detection): new ValidatorHealthcheckStartPeriod in
omnibase_infra catalog (validator_healthcheck_start_period.py) asserts a 120s
floor for migration-completion gates, wired into `onex validate runtime`
(cmd_validate_runtime) AND backed by unit tests that gate every PR via
pre-commit + CI. A migration-completion gate is identified by BOTH a healthcheck
that polls migration completion AND a service_completed_successfully dependency,
so ordinary app services (e.g. intelligence-api, an HTTP liveness probe with a
legitimately short 40s start_period) are not swept into the floor.

Prod probed read-only only; no prod mutation. Applies to live prod via the
batched stability/prod rebuild (deploy_pending).

Evidence-Ticket: OMN-12973
Evidence-Source: OCC#PENDING

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-12971): deliver Vertex ADC bearer token to runtime effects container (secret-ref discipline) (#1949)

* feat(OMN-12971): deliver Vertex ADC bearer token to runtime effects container (secret-ref discipline)

Wire the Vertex $500-credit ADC path into the runtime, ADDITIVE next to the
Gemini API-key path (provider-agnostic; neither provider removed or forced).

runtime_policy.contract.yaml: add a secret-source mapping llm.vertex.access_token
(source_type=file, /run/secrets/vertex_access_token) to the dev/stability-test/
judge profiles, alongside the existing llm.gemini.api_key env mapping. The ref
name MUST match cloud-vertex-gemini.secret_ref in omnimarket
bifrost_delegation.yaml. The resolved VALUE is a short-lived OAuth bearer token
minted from ADC, refreshed by the operator; the token VALUE is never committed —
only the ref name + in-container path. Add aiplatform.googleapis.com to the
cloud host allowlist.

docker-compose.infra.yml: bind the operator-supplied host token file read-only to
/run/secrets/vertex_access_token on the main and effects runtimes
(VERTEX_ACCESS_TOKEN_HOST_FILE, default /dev/null so lanes without Vertex still
start; Gemini key path unaffected). Pass through BIFROST_VERTEX_GEMINI_ENDPOINT_URL
(overlay supplies the complete Vertex OpenAI-compat URL) and
GOOGLE_CLOUD_PROJECT/LOCATION (default empty).

runtime-policy.env: regenerated from the contract via render_runtime_policy_env
(test_runtime_policy_env_matches_contract_renderer proves contract<->env parity).

test_runtime_policy_contract.py: update host-allowlist assertion for the additive
Vertex host.

Pre-existing platform-wide topic-parity-gate failures (25 unrelated topics) are
identical with and without this change (proven by diff); that gate is pre-commit-
only (not a CI merge gate) and this change adds zero new topic gaps, so that one
hook is SKIP-ped. No deploy/receipt/merge gate is bypassed.

* fix(OMN-12971): make Vertex runtime env contract-owned

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-13008): automated .201 disk maintenance — worktree GC + docker GC + watermark alert ratchet (#1952)

* feat(OMN-13008): automated .201 disk maintenance — worktree GC + docker GC + watermark alert ratchet

Conservative, keep-list-driven disk maintenance to prevent the 2026-06-11
/data ~95% outage that killed all three lanes mid-demo:

- scripts/worktree-gc.sh: drives the canonical omniclaude prune-worktrees.sh
  (merged+clean+pushed safety) on both Mac (merge-sweep tick) and .201
- scripts/disk-gc.sh + disk_gc_plan.py: conservative docker/builder/image GC.
  Pure, testable removal planner honoring a VERSIONED keep-list
  (deploy/disk-gc/keep-list.yaml): never reaps a kept repo, kept tag, in-use
  image, or anything younger than min_age_days; keeps N superseded generations.
- scripts/disk-watermark-check.sh + disk_watermark_event.py: df watermark
  ratchet. >=85% emits a typed disk-watermark bus event (warning) that the
  sweep auto-ticket path turns into a Linear ticket; >=90% emits critical.
  Broker addr is fail-fast from KAFKA_BOOTSTRAP_SERVERS (no localhost default).
- deploy/disk-gc/: systemd USER timer (onex-disk-gc.timer/.service, hourly) +
  install-disk-gc.sh. User units, NOT lane containers.
- tests: 20 unit tests incl. GC plan-safety invariants + dry-run proof that
  default mode issues no destructive op (a wrong-delete GC is worse than none).

Contract: contracts/OMN-13008.yaml

* fix(OMN-13008): pass docker inventory to GC planner via stdin, not env (ARG_MAX)

On a host with many docker images, passing the full image/ps inventory as env
vars to disk_gc_plan.py exceeds ARG_MAX ('Argument list too long', exit 126),
producing an empty plan. Write inventory to per-run scratch files (under the log
dir, never /tmp; cleaned on exit) and hand it to the planner on stdin as a JSON
envelope. Verified the failure live on .201; planner now reads stdin.

* fix(OMN-13008): simplify GC plan stdin pipe (two processes, no nested subprocess)

* fix(OMN-13008): keep-wins reconciliation — never remove an image id with any keep reason

A single image id can surface in multiple 'docker image ls' rows (one per
repo:tag). One tag could route the id to dangling-removal while another routes
it to keep (e.g. tagged 'latest' or within-N-generations). Live .201 plan showed
an id in BOTH remove_image_ids and kept_reasons. Reconcile at the end: keep wins
— any id with a keep reason is dropped from the remove list; remove list deduped.
Adds 2 regression tests. Verified live on .201.

* fix(OMN-13008): timer uses OnCalendar=hourly for reliable re-arm (oneshot service)

OnUnitActiveSec does not reliably re-elapse for a oneshot service once it goes
inactive (observed NextElapseUSecMonotonic=infinity live on .201 after the first
run). Switch to OnCalendar=hourly + RandomizedDelaySec so the timer re-arms every
hour. Keep Persistent=true for missed-run catch-up.

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13005): materialize blocking event_consumer in runtime auto-wiring (runner consume-leg degenerate rows) (#1951)

* fix(OMN-13005): materialize blocking event_consumer in auto-wiring (runner consume-leg degenerate rows)

The runtime auto-wiring materialized event_publisher for handlers that
declare it but had no equivalent for event_consumer. Request/response
EFFECT handlers (HandlerContextRoiRunner) that publish a command then
block on the correlated terminal event fell back to their no-op consumer
default, returning None immediately -> every result row degenerate
(failure_stage=generation, attempt_count=0) while generations succeeded
~1s later.

Adds _make_sync_event_consumer (mirror of _make_sync_event_publisher),
backed by service_terminal_event_consumer.make_terminal_event_consumer:
a sync (topic, correlation_id, timeout) -> dict | None adapter that runs
the proven direct-Kafka correlate-and-wait loop (from RuntimePatternBBroker)
on an isolated event loop in a worker thread, so blocking does not deadlock
the runtime dispatch loop that delivers the awaited terminal.

TDD through the REAL dispatch path: test_event_consumer_injection drives a
trial through _prepare_handler_wiring with a terminal arriving after a delay
and asserts a non-degenerate row; verified RED with injection disabled.

* fix(OMN-13012): two-phase (seek-now/wait-later) terminal event_consumer to close the subscribe-after-publish race (#1954)

The OMN-13005 injected event_consumer is a single callable that does
assign -> seek_to_end -> poll internally, all AFTER the handler has
already published its command. Once OMN-13010 freed the dispatch loop and
generation began completing in ~1s, the correlated terminal lands BEFORE
the single-call consumer's post-publish seek_to_end positions, so
seek_to_end skips PAST the already-emitted terminal and the runner times
out on an offset beyond it (probe3, run_id=20260611T2140Z-probe3 -- both
arms degenerate, zero rebalances).

Splits positioning from waiting so the caller subscribes BEFORE it
publishes:

  session = consumer.open(topic)        # assign + seek_to_end NOW
  publisher(command_topic, payload)     # publish AFTER positioning
  payload = session.wait(cid, timeout)  # block from the captured position

The returned TerminalEventConsumer is still directly callable with the
legacy (topic, cid, timeout) -> dict | None single-call shape for any
consumer that does not need subscribe-before-publish; the runner is the
only consumer today. TerminalConsumerSession owns a dedicated event loop
on a daemon worker thread for the whole open->wait->close lifecycle,
preserving the OMN-13005 loop-isolation discipline so blocking never
deadlocks the runtime dispatch loop.

TDD (real dispatch path, RED-then-GREEN): extends the OMN-13005 injection
test with a terminal emitted IMMEDIATELY after publish. The single-call
(seek-after-publish) consumer MISSES it (degenerate row -- RED test
asserts failure_stage=generation); the two-phase (open-before-publish)
consumer CATCHES it (non-degenerate -- GREEN). The Kafka layer is faked at
the two service seams against a shared in-memory log modeling seek-to-end
semantics. OMN-13005 blocking-correlate behavior preserved. 269/269
auto_wiring unit tests pass; mypy --strict clean.

Sibling to OMN-13010 / OMN-13005 / OMN-13003.

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13005): route terminal consumer through Kafka boundary

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12990): pin worker replicas in ledgered config + fail-fast lane overrides (silent-drop ratchet) (#1948)

* fix(OMN-12990): pin worker replicas in ledgered config + fail-fast lane overrides

The base compose sets runtime-worker deploy replicas to ${WORKER_REPLICAS:-0}
(soft-default ZERO). The stability lane's required state includes a running
worker (GATE_ZERO_PROOF.md: 4 runtime containers), but the worker presence rode
entirely on a compose soft-default — any plain compose up/recreate without the
policy env silently scaled the worker to zero with no error and no signal.

Fix (contract-native + fail-fast):
- Add 'replicas' to ModelRuntimeProcessPolicy; pin replicas: 1 in every lane's
  worker block in runtime_policy.contract.yaml.
- Renderer emits {PROFILE}_WORKER_REPLICAS into the ledgered runtime-policy.env
  for dev/stability-test/judge/prod.
- stability + prod compose overrides reference ${..._WORKER_REPLICAS:?...}
  (fail-fast, NO silent :-1/:-0 default). A recreate that omits the policy env
  now aborts loudly instead of dropping the worker. prod previously had no
  override at all and inherited the dangerous :-0 default.

Ratchet (recurrence guards):
- tests asserting fail-fast override form (no soft default), contract-declared
  replica pin >= 1 per lane, and rendered {PROFILE}_WORKER_REPLICAS=1 in the
  ledgered env.
- runbook deploy/verify procedure adds an expected-container census (worker
  must be present) via verify_container_manifest; a missing worker is a FAILURE,
  not silence.

Note: SKIP=topic-parity-gate — that local-only advisory gate (absent from all
.github/workflows, not a required CI check) fails on 25 pre-existing cross-repo
topic gaps (build-loop/omniclaude/omniweb) identical on pristine base HEAD
8d7da1249; this change adds zero topics. All other hooks ran clean.

Evidence-Ticket: OMN-12990

* test(OMN-12990): cover worker replica policy integration

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-12909): add gateway bus forwarder P0A (#1946)

* feat(OMN-12909): add gateway bus forwarder p0a

* test(OMN-12909): add gateway forwarder integration coverage

* fix(OMN-12909): satisfy gateway forwarder validators

* test(OMN-12909): allow gateway forwarder bus protocol

* fix(OMN-12909): sync gateway forwarder entry point

* fix(OMN-12909): refresh runner image identity lock

* test(OMN-12909): relax JSON normalizer mixed benchmark threshold

* fix(OMN-12909): allow gateway handlers to boot unconfigured

* fix(OMN-12909): declare gateway forwarder runtime profile

* fix(OMN-12909): update runner identity backmerge expectation

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-13011): LANE CENSUS RECONCILIATION ratchet — declared desired-state per lane, drift = auto-ticket (#1955)

The class fix for the recurring lane-drift regression. Nothing reconciled the
declared desired state of a runtime lane against what is actually running, so the
same failure kept recurring with zero signal: volume config drift (OMN-12945),
WORKER_REPLICAS silent zero (OMN-12988/12990), and on 2026-06-11 prod runtime
containers plus the broker network were silently absent for hours during demo prep.

Ships a per-lane DESIRED-STATE census:
- (a) DECLARED in a versioned lane manifest (deploy/lane-census/lane-manifest.yaml):
  container set, network, replicas, image-tag pattern per lane
  (stability-test/prod/judge/dev), derived from the canonical compose lane files.
  A parity ratchet keeps the manifest locked in step with the compose files.
- (b) RECONCILED on a schedule on .201 by SHARING the OMN-13008 systemd timer
  (a drop-in 4th ExecStart on onex-disk-gc.service — never a second timer) and
  on-demand via scripts/lane-census-check.sh / runtime_sweep.
- (c) Drift = typed bus event (onex.evt.infra.lane-census-drift.v1) + Linear
  auto-ticket naming exactly what is missing/extra (container_absent,
  network_detached, replicas_zero, unexpected_container, oneshot_failed/stuck,
  image_tag_mismatch). Fail-fast, no warn-only mode (gates-block policy); exit 30
  on drift; bus publish fail-fast on KAFKA_BOOTSTRAP_SERVERS (no localhost default).

Red fixture reproduces 2026-06-11: prod runtime containers absent + broker network
detached must produce the exact drift findings + a non-zero exit hours before a
human noticed. Pure planner is fully unit-tested; shell driver dry-run-tested.

Builds on the OMN-12988 deploy-agent RUNTIME census (deploy-time) as the
complementary steady-state reconciler; closes the runtime-worker.yaml
container_name: null census gap by sourcing names from the compose lane files.

Evidence-Ticket: OMN-13011
Config-drift family: OMN-12945
Relates-to: OMN-13009, OMN-12988, OMN-13008

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13020): vendor missing node migrations — llm_routing 0000 + context_roi 001 (#1956)

Vendors two omnimarket node-source migrations into the infra forward-migration
tree via scripts/sync-node-migrations.sh (the canonical OMN-12559 mechanism):

- node_projection_llm_routing/0000_create_llm_routing_decisions.sql
  (source: omnimarket #1168 / OMN-12942, merge ed6734f8)
- node_projection_context_roi/001_create_context_roi_scores.sql
  (source: omnimarket #1178 / OMN-12955, merge 5010b1f4)

Without the 0000 base table, node_projection_llm_routing/0001 (CREATE VIEW)
hard-fails against NODE_POSTGRES_DB=omnidash_analytics — exactly the prod
forward-migration exit-3 of 2026-06-11T09:35:52Z, and reproduced by
construction in any clean clones@dev build. Files are byte-identical to the
omnimarket dev blobs (sha256 f8a8b339… / c4126e65…) and to the untracked
hot-patch copies on the .201 stability clone.

Both migrations are self-contained, all-statements-IF-NOT-EXISTS, and 0000
sorts lexically before 0001 within the node's namespaced identity space.

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13058): close TerminalConsumerSession on open() failure (worker thread + event loop leak) (#1957)

TerminalConsumerSession.__init__ starts its dedicated worker loop thread
immediately. TerminalEventConsumer.open did
'session = TerminalConsumerSession(...); return session.open()' with no
cleanup: any failure inside session.open() (consumer start timeout,
partition-assign timeout, broker auth error) propagated out of the raising
expression, the session reference was lost, and the daemon worker thread plus
its never-closed asyncio event loop leaked -- one pair per failed open. The
motivating caller (HandlerContextRoiRunner) opens a session per trial, so a
160-560-trial battery against a degraded broker accumulates hundreds of
leaked threads in the long-lived effects container.

Fix: wrap session.open() in try/except BaseException -> session.close()
(idempotent: stops the loop, joins the thread) -> re-raise. Covers both the
two-phase .open(topic) path and the legacy single-call __call__ path.

Found by the P3.3 doctrinal review of merged #1951 (b9712af9 / 36d98275).

TDD: tests/unit/runtime/test_service_terminal_event_consumer_open_failure.py
injects a real open failure through the production path (event bus without
_bootstrap_servers) and asserts no alive terminal-consumer-* thread after the
raise. Verified RED with the fix stashed (2 failed), GREEN with it (2 passed).

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-13021): non-dev-base guard — fail feature-base PRs absent Stacked-Parent declaration (retro A-6) (#1958)

Any PR whose base is neither dev nor main fails unless the body carries
'Stacked-Parent: #N'. Prevents the feedback_stacked_prs_orphan_from_dev class
(#1185/#1954 auto-merged INTO parent feature branches and stranded off dev).
base=main remains governed by main-target-guard.

Epic OMN-13013 (process enforcement ratchets — June 12 retro).

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-13014): hot-patch ledger rebuild preflight gate (retro B-1) (#1959)

* feat(OMN-13014): hot-patch ledger rebuild preflight gate (retro B-1)

Hot-patches on .201 (.prepatch sibling discipline) silently revert on any
image rebuild/force-recreate — the 2026-06-11 20:58Z rebuild already erased
a live /api/generate patch once. This adds the rebuild-path gate:

- scripts/preflight_hotpatch_ledger.py: given a target container or lane +
  per-repo build refs, hard-fails when any hot-patch ledger row's source PR
  merge commit is not an ancestor of the build ref (git merge-base
  --is-ancestor), plus a .prepatch tripwire of the running container
  (unledgered .prepatch = hard fail; --post-rebuild = zero .prepatch
  expected). Sole bypass: HOTPATCH_PREFLIGHT_BYPASS carrying the Rule-10
  '# skip-token-allowed: <user-approval-receipt-id>' form.
- scripts/deploy-runtime.sh: guard_hotpatch_ledger wired into main() before
  build/preview (both dry-run and execute), lane derived from the compose
  project; skips loudly only when no ledger exists on the host.
- tests/unit/scripts/test_preflight_hotpatch_ledger.py: 17 unit tests
  (ancestor gate, lane scoping, ledger loading, tripwire, bypass forms).
- tests/ci/test_receipt_gate_install_guard.py: repair stale guard — core
  OMN-12565 replaced the OMN-9198 'uv pip uninstall first' install step with
  a cleared workspace venv (uv venv --clear); assert the new contract.

Ledger backfilled from live census (5 .prepatch files / 4 source PRs, all
MERGED to dev) at /data/omninode/hotpatch-ledger/ledger.yaml on .201.

* fix(OMN-13014): scope missing-prepatch tripwire warning to the probed container

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12977): workspace-build sibling-pin ratchet — honor consuming lock, fail-fast on stale vendor (#1936)

* fix(OMN-12977): workspace-build sibling-pin ratchet — honor consuming lock, fail-fast on stale vendor

The workspace-mode image build (BUILD_SOURCE=workspace, used by the stability-test
deploy procedure) vendored sibling/foundation packages from whatever the canonical
OMNI_HOME clones happened to be checked out at, ignoring the consuming repo's
uv.lock. On 2026-06-11 this shipped a 13-day-stale omnibase_infra 0.37.0-dev
(pre-OMN-12501 Protocol-quarantine guard) + core 0.42.0 against an omnimarket dev
lock pinning infra 0.38.1@e2dbdc95 / core 0.44.0@c97c2c9a, dropping the guard and
crashing wire_from_manifest bootstrap fatally (stability lane down on demo day).

Fix + recurrence ratchet (same PR):
- scripts/runtime_build/check_sibling_lock_pins.py: parse the consuming repo's
  uv.lock for expected version+git-rev of each foundation/sibling package
  (scoped to the package's own source line so editable/registry pins are not
  cross-attributed a dependency's rev), resolve the actual clone version+HEAD,
  compare, and classify drift backward/forward/none. Fail-fast (exit 1) on any
  drift; --allow-drift records an explicit operator override in the artifact,
  never silent.
- stage_workspace.sh: runs the preflight against the canonical clones before
  staging; aborts the build (exit 3) on unacknowledged drift and writes
  workspace/sibling-pin-comparison.json.
- compute_workspace_provenance.py: folds the expected-vs-actual comparison into
  build-provenance.json so deploy verifiers can assert the build honored the lock;
  flags unacknowledged drift as a provenance error.
- Dockerfile.runtime: COPY the comparison artifact (committed placeholder so the
  COPY always resolves; overwritten by stage_workspace.sh in workspace mode).
- TDD: 19 unit tests covering lock parsing (git/registry/editable sources),
  drift classification, the exact 0.37.0-vs-0.38.1 stale case, check_pins exit
  codes, and the allow-drift override. Pre-existing mypy-strict bare-dict errors
  in compute_workspace_provenance.py fixed in the same pass.

Evidence-Ticket: OMN-12977
Evidence-Source: pending-occ

* ci(OMN-12977): retry runtime smoke compose port race

* test(OMN-12977): align sibling-pin script tests with current API

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12989): workspace-mode image build must honor sibling lock pins (#1947)

* fix(OMN-12989): workspace build must honor sibling lock pins (fail-fast ratchet)

The 2026-06-11 stability bootstrap crash was caused by a workspace-mode
--no-cache rebuild that vendored omnibase_infra 0.37.0-dev (13-day-stale
worktree) even though omnimarket dev's uv.lock pins infra 0.38.1. The
downgraded sibling predated the OMN-12501 Protocol-quarantine guard and
turned a latent contract defect into a fatal crash.

Fix + ratchet (same PR set):
- scripts/runtime_build/resolve_workspace_pins.py: parse the consuming repo's
  uv.lock for sibling pins (version + git rev); classify each staged/installed
  sibling as exact/ahead/regression/unpinned; fail-fast (WorkspacePinError) on
  any regression BELOW the lock pin. Stdlib version-tuple fallback when
  packaging is absent so the ratchet never fails-open.
- compute_workspace_provenance.py: enforce sibling pins + a host-infra
  self-check (installed omnibase_infra vs lock pin — the exact crash vector,
  since host infra is built from the context, not staged), and emit a
  pin_comparison block into build-provenance.json for deploy verifiers.
- Dockerfile.runtime: COPY resolve_workspace_pins.py beside the provenance
  script so the in-image import resolves.
- TDD: failing tests first (test_resolve_workspace_pins.py, 11 cases) +
  provenance integration tests; fixed a pre-existing stale RUNTIME_VERSION
  assertion to read pyproject dynamically.

Evidence-Ticket: OMN-12989

* fix(OMN-12989): workspace build must honor sibling lock pins (fail-fast ratchet)

The 2026-06-11 stability bootstrap crash was caused by a workspace-mode
--no-cache rebuild that vendored omnibase_infra 0.37.0-dev (13-day-stale
worktree) even though omnimarket dev's uv.lock pins infra 0.38.1. The
downgraded sibling predated the OMN-12501 Protocol-quarantine guard and
turned a latent contract defect into a fatal crash.

Fix + ratchet (same PR set):
- scripts/runtime_build/resolve_workspace_pins.py: parse the consuming repo's
  uv.lock for sibling pins (version + git rev); classify each staged/installed
  sibling as exact/ahead/regression/unpinned; fail-fast (WorkspacePinError) on
  any regression BELOW the lock pin. Stdlib version-tuple fallback when
  packaging is absent so the ratchet never fails-open.
- compute_workspace_provenance.py: enforce sibling pins + a host-infra
  self-check (installed omnibase_infra vs lock pin — the exact crash vector,
  since host infra is built from the context, not staged), and emit a
  pin_comparison block into build-provenance.json for deploy verifiers.
- Dockerfile.runtime: COPY resolve_workspace_pins.py beside the provenance
  script so the in-image import resolves.
- TDD: failing tests first (test_resolve_workspace_pins.py, 11 cases) +
  provenance integration tests; fixed a pre-existing stale RUNTIME_VERSION
  assertion to read pyproject dynamically.

Evidence-Ticket: OMN-12989

* test(OMN-12989): co-locate provenance pin helper in fixture

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13055): absent repos in REPOS list warn and exit 0 instead of failing (#1960)

Missing repos (not cloned locally) now emit a WARN line and are tracked
in a separate WARNED array. Only real fetch/ff failures cause exit 1.
This makes pull-all.sh safe to use on machines with a partial clone set,
while keeping the explicit-list override behavior intact.

Adds three regression tests: absent-only exits 0, absent+present exits 0
with OK for the present repo, present-failed+absent exits 1.

Evidence-Ticket: OMN-13055

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13070): config_prefetcher overlay wins on controlled lanes (#1961)

On infisical_required=True lanes, fetched/overlay config now always wins
over ambient env. Ambient env is retained only as a declared bootstrap
fallback (with an explicit provenance INFO log line) when Infisical
returns None. apply_to_environment also overwrites stale env on controlled
lanes. Uncontrolled lane (infisical_required=False) behaviour is unchanged.

Adds 5 regression tests: controlled-lane Infisical-wins, env-bootstrap-
fallback, apply_to_environment overwrite, missing-from-both-is-error, and
uncontrolled-lane-env-still-wins. Refactors _resolve_key to return a
(outcome, value, error) tuple to satisfy the ≤5-param pattern gate.

Source: docs/audits/2026-06-10-runtime-env-overlay-authority-audit.md

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13062): migration-gate vacuity fix — sentinel discipline, wait-for-postgres, skip-manifest, sync-check nonzero (retro A-10) (#1963)

* fix(OMN-13062): migration-gate vacuity fix — sentinel discipline, wait-for-postgres, skip-manifest, sync-check nonzero

Fixes three bugs identified in retro A-10 (recurrences OMN-12885, OMN-12934):

(1) RUNNER SENTINEL DISCIPLINE
    run-forward-migrations.sh now clears migrations_complete=FALSE at the
    start of every run and sets it TRUE only as its FINAL act after all
    infra and node migrations succeed. Any mid-run failure leaves the gate
    UNHEALTHY. runner_completed_at is stamped at the same final step as
    durable evidence of a successful completion run.

(2) SYNC-NODE-MIGRATIONS VACUOUS GATE
    sync-node-migrations.sh --check now exits 2 (not 0) when the omnimarket
    source tree is unresolvable. Silent exit-0 was hiding drift. The single
    opt-out is SYNC_NODE_MIGRATIONS_SKIP_UNRESOLVABLE=1 for environments
    that intentionally run without the source.

(3) WAIT-FOR-POSTGRES GUARD
    run-forward-migrations.sh now waits up to PG_WAIT_RETRIES (default 30)
    x 2s for Postgres to accept connections before proceeding, guarding
    the first-boot initdb race.

(4) SKIP-MANIFEST
    docker/migrations/skip-manifest.yaml introduced as the sole committed
    escape for intentionally-skipped migrations. The runner reads this at
    startup; listed migrations are recorded in schema_migrations with
    checksum "skip-manifest" without executing the SQL.

(5) MIGRATION 085
    Adds runner_completed_at TIMESTAMPTZ column to db_metadata so the
    runner's final stamp is durable in the schema (idempotent ADD COLUMN IF
    NOT EXISTS). Rollback included.

22 regression tests added covering all five fix surfaces.

* fix(OMN-13062): stamp schema fingerprint for migration 085

Migration 085 (085_add_runner_completed_at_to_db_metadata.sql) was added
in the initial commit but schema_fingerprint.sha256 was not regenerated.
Running `python scripts/check_schema_fingerprint.py stamp` updates the
artifact from the stale hash to match the 71 migration files.

Evidence-Source: OCC#2563
Evidence-Ticket: OMN-13062

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-12864): Bifrost endpoints → committed overlay authority + fail-loud loader (#1964)

* feat(OMN-12864, OMN-12814, OMN-12945): Bifrost endpoints → committed overlay authority + fail-loud loader

OMN-12864 — Committed lane overlay
  - docker/lane-overlays/dev.bifrost.yaml: typed deployment bindings for all
    four BIFROST_LOCAL_*_ENDPOINT_URL values (coder :8000, reasoner :8001,
    embedding :8100, ds4-flash :8101). Previously only available as ephemeral
    shell exports on .201; now committed, auditable, diff-able, and CI-checked.
  - docker/lane-overlays/dev.bifrost.env: generated dotenv sidecar consumed by
    compose via env_file; never edited directly (yaml is authority).
  - docker/docker-compose.infra.yml: wire the env_file block at compose root so
    the four endpoints are injected into the interpolation context on a clean
    shell. Hardcode BIFROST_CONTRACT_PATH (remove :-/empty footgun — OMN-12814).
  - scripts/render_bifrost_lane_overlay_env.py: render script regenerates the
    env sidecar from the YAML source.
  - src/omnibase_infra/runtime/models/model_bifrost_lane_overlay.py:
    ModelBifrostLaneOverlay — typed Pydantic model enforcing URL completeness
    (OMN-12815: every URL must end in /chat/completions).

OMN-12814 — Fail-loud loader
  - render_bifrost_delegation_contract: raises ProtocolConfigurationError on
    FileNotFoundError, YAMLError, ValidationError, and zero-endpoint renders.
    No lru_cache — every restart re-renders from packaged source so a stale
    cache cannot pin a broken result across deploys.

OMN-12945 — Re-seed from packaged source on deploy
  - docker/entrypoint-runtime.sh: set BIFROST_FORCE_RESEED=1 on every container
    restart so the named-volume copy is always rebuilt from the packaged
    bifrost_delegation.yaml merged with committed lane-overlay endpoints.
  - render_bifrost_delegation_contract: honor BIFROST_FORCE_RESEED/force_reseed
    flag to bypass the stale-volume early-return path entirely.

Tests:
  - tests/ci/test_bifrost_lane_overlay.py: CI gate — env sidecar in-sync with
    YAML source; all four BIFROST_LOCAL_* keys present.
  - tests/unit/runtime/models/test_model_bifrost_lane_overlay.py: bare-base URL
    rejection, env dict mapping, extra-field rejection.
  - tests/unit/runtime/test_render_bifrost_delegation_contract.py: fail-loud
    paths, force-reseed, zero-endpoint error, endpoint URL completeness.
  - tests/unit/models/test_model_serialization_roundtrip.py: roundtrip coverage.

* fix(OMN-12864): move bifrost env_file to service level — fix compose schema validation failure

Top-level 'env_file' is rejected by Docker Compose v2 schema validator
('additional properties not allowed'). This caused 10+ compose-render
integration tests to fail in CI.

Fix:
- Remove top-level env_file block from docker-compose.infra.yml
- Add per-service env_file on omninode-runtime, runtime-effects,
  runtime-worker (the three containers that render Bifrost)
- Change BIFROST_LOCAL_*:? to BIFROST_LOCAL_*:- in x-runtime-env
  (compose-level validation removed; Python validates via
  ModelBifrostLaneOverlay + render_bifrost_delegation_contract)
- Add two CI gate tests: compose_env_file_is_service_level_not_top_level
  and runtime_services_have_bifrost_env_file

* fix(OMN-12864): allow BIFROST_LOCAL_* empty defaults in silent-fallback gate

The overlay authority pattern (OMN-12864) passes BIFROST_LOCAL_*_ENDPOINT_URL
via service-level env_file (docker/lane-overlays/dev.bifrost.env), not at
compose config time. Compose-level :? would break CI rendering without the
overlay pre-loaded. Validation at the Python layer (ModelBifrostLaneOverlay
+ render_bifrost_delegation_contract) is the enforcement point.

Add the three failing vars to ALLOWED_EMPTY_DEFAULTS with OMN-12864 citation.

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-12858/OMN-12879/OMN-12880): dispatcher route-coverage CI gate (#1965)

* feat(OMN-12858/OMN-12879/OMN-12880): dispatcher route-coverage CI gate

Static contract analysis: every subscribed command topic (onex.cmd.*)
must declare handler_routing or runtime_dispatch, or the message goes to
DLQ silently. This gate would have caught two recent incidents:

1. June 9 DLQ regression (OMN-12858 post-mortem): node_generation_consumer
   subscribed onex.cmd.omnimarket.node-generation-requested.v1 but a
   sole-handler revert left zero dispatcher routes registered. Messages
   went to DLQ silently with no CI signal.

2. June 12 DEL-01 live finding: onex.cmd.omnimarket.delegate-skill.v1 was
   consumed by dev lane bus but no dispatcher route existed in any deployed
   contract. Discovered via manual rpk consumer-group lag probe (DEL-01
   evidence, docs/evidence/2026-06-12-weekend-pass/).

Deliverables:
- scripts/check_dispatcher_route_coverage.py — static YAML scanner that
  checks both omnibase_infra and omnimarket contract trees; ratchet
  allowlist for known pre-existing violations; --changed-contracts mode
  (OMN-12879) for per-PR scoping; compat publish topics excluded (OMN-12880)
- .github/workflows/dispatcher-route-coverage.yml — CI workflow that
  checks out omnimarket sibling, collects changed contract paths in PR
  mode, and runs the gate; fires on PR, push-to-main, and merge_group
- tests/ci/test_dispatcher_route_coverage_gate.py — 12 unit tests
  covering RED/GREEN/COMPAT/CHANGED-MODE/ALLOWLIST/MULTI-DIR paths plus
  live-contract regression proof against the actual omnibase_infra tree

Allowlist additions:
- onex.cmd.omnibase-infra.pattern-b-dispatch.v1 (RuntimePatternBBroker,
  imperative consumer, OMN-12525 migration target)
- onex.cmd.platform.contract-resolve-requested.v1 (transitional HTTP
  bridge node_contract_resolver_bridge OMN-2756, metadata.transitional=true)

[OMN-12858, OMN-12879, OMN-12880]

* fix(OMN-12858): drop full uv sync from dispatcher-route-coverage workflow

Gate script only needs pyyaml (stdlib + yaml). Using full setup-python-uv
was causing 10+ minute timeout. Replace with direct pip install pyyaml and
invoke python3 directly. Reduces job from 10m timeout to <1m.

[OMN-12858]

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-13026): wire transport-mock-lint pre-commit hook + CI gate (#1962)

* feat(OMN-13026): wire transport-mock-lint pre-commit hook + CI gate (omnibase_infra)

Activates the transport-mock-lint validator (from omnibase_core, OMN-13026)
on omnibase_infra. Ratchet baseline: 218 existing violations across 63 files
frozen in validation/transport_mock_baseline.yaml. New bare AsyncMock/
MagicMock on EventBus/transport surfaces are blocked by pre-commit hook and
CI lint step. Existing violations tracked for drain by per-site tickets
(parent OMN-13026).

Reference incident: PR #1181 bare AsyncMock hid missing EventBusKafka.stop().

Evidence-Ticket: OMN-13026

* fix(OMN-13026): use uv run python for CI lint step + bump omnibase_core pin to include transport_mock_lint

The transport-mock lint CI step previously cloned omnibase_core and ran
`python -m omnibase_core.validators.transport_mock_lint` with PYTHONPATH,
but this failed: `No module named omnibase_core.validators.transport_mock_lint`
because it ran `.venv/bin/python` which uses the locked venv, and the venv
omnibase_core pin (2defabef4) predates the transport_mock_lint module.

Fix: use `uv run python` (removes the clone step) and bump omnibase-core git
pin from 2defabef4 to 309d89fa7 (PR 1231 merge commit on dev) so
transport_mock_lint is available in the locked venv.

* fix(OMN-13026): align transport mock baseline and runner lock

* fix(OMN-13026): sync omnibase_core pin + runner identity lock to dev baseline

Align pyproject.toml omnibase_core rev to 2defabef (required by
test_release_backmerge_preserves_proven_runtime_core_pin) and update
docker/runners/runner-image.lock.json identity_digest/shared_env_digest
to match dev runner image lock (79b08f44 / 90c8b3b9).

Both were stale from the prior session's pin bump that used an older SHA.

* fix(OMN-13026): source transport mock validator from core

* fix(OMN-13026): source transport validator from core dev

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-13094): --output receipt mode on onex node/run — quiet typed receipts with durable capture (#1966)

Phase 2a of the skill-output-suppression slice (epic OMN-13089, plan
docs/plans/2026-06-12-skill-output-suppression-plan.md Phase 2 item 1).

- onex node/onex run gain --output receipt: ALL runtime logging routes to a
  run_id-suffixed capture file under <state-root>/captures/ (no console
  handlers — kills the 25-50-line RuntimeLocal INFO stream at the source);
  stdout carries exactly ONE typed ModelSkillResult JSON with the FULL
  handler result (result_model = concrete handler result type FQN).
- Durable capture: capture log + handler result content-addressed via
  omnibase_core ArtifactStore (OMN-13093); artifact.captured +
  tool.output.captured emitted to the emit daemon socket (--emit-socket,
  default ~/.claude/emit.sock).
- Failure asymmetry: artifact write failure => FULL output printed, no
  receipt (no hidden loss); emission failure => receipt still prints, event
  spooled to <state-root>/emit_spool/ for replay.
- Node failure => status=failed/error with full error + capture log INLINE
  in the receipt (errors are never hidden) and artifact-backed.
- Default output mode unchanged (enforcement is Phase 4).
- RuntimeLocal exposes handler_result (receipt schema identity).
- core pin 2defabef -> ae8793bd (merged OMN-13091/13093 receipt models +
  ArtifactStore); runner-image identity lock regenerated and the OMN-12765
  backmerge identity constants updated for the new pin.

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-13097): onex skill subcommand + declarative skill->node mapping (#1968)

* feat(OMN-13097): onex skill subcommand + declarative skill->node mapping

Phase 4a of the skill-output-suppression slice (epic OMN-13089, plan
docs/plans/2026-06-12-skill-output-suppression-plan.md Phase 4 item 1/2).

A dispatch skill IS one CLI call (user directive 2026-06-12). This adds the
`onex skill <name> [args]` dispatch surface that the 24 omniclaude shim
migrations build on:

- onex.cli entry-point `skill` -> cli_skill.run_skill_by_name. Resolves the
  skill via the declarative skill_mapping.yaml registry, builds the backing
  node's input payload from the skill's CLI args, writes it under the state
  root (.onex_state/tmp/<skill>-<run_id>.json — never /tmp), resolves the
  node's packaged contract exactly like `onex node`, and dispatches through
  the proven receipt-mode path (run_receipt_mode, OMN-13094). stdout is
  exactly one typed ModelSkillResult JSON with the FULL handler result.
- skill_mapping.yaml: declarative DATA mapping all 24 dispatch shims to their
  backing onex.nodes node + typed result-model FQN (verified against each
  live handler handle() return type on origin/dev) + per-arg payload specs +
  static payload + keyword classifiers (delegate task_type as data, not code).
  Adding a skill is a YAML edit + fixture, never a CLI code change (ticket
  deliverable 2/3). Mapping lives beside the node-resolution surface, never
  hardcoded branching in the CLI.
- Typed models split one-per-file (repo convention): ModelSkillArgSpec,
  ModelSkillClassifier, ModelSkillMapping, ModelSkillMappingRegistry,
  EnumSkillArgType. Frozen, extra=forbid, fail-fast coercion/validation.
- validation_exemptions.yaml: Click-callback param-count + literal-identifier
  name-field exemptions mirroring the existing cli_node run_node_by_name
  precedent (OMN-11570) — same pattern, same rationale.

dod_evidence:
- 20 unit tests pass (registry validity, all-24-shims coverage, FQN result
  models, arg parsing/coercion/positional/required, classifiers, payload
  build, receipt-mode dispatch wiring, payload-under-state-root not /tmp).
- uv run mypy src/ --strict: clean (2438 source files).
- ruff format + check: clean. pre-commit run on changed files: pass.
- runner-image identity lock regenerated for the pyproject entry-point add
  (same as OMN-13094).

* test(OMN-13097): rebind OMN-12765 backmerge identity constants for onex skill pyproject change

Adding the `skill` onex.cli entry-point to pyproject.toml changes the
runner-image identity_digest (and shared_env_digest) the lock binds. Update
the hardcoded expected constants in the backmerge-identity test to the
regenerated values — same mechanical rebind OMN-13094 performed for the core
pin bump. Identity + runner-image-identity tests pass (14).

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix+test(OMN-13012): force terminal-topic metadata refresh so both ephemeral consumers assign (runner consume-leg wedge) (#1969)

The runner consume-leg wedged on the live stability battery (image c0505521f1fa,
EXP1-3_RUNNER_CONSUME_LEG_BLOCKER): only the FAILED terminal topic ever subscribed
while the COMPLETED topic never assigned, so the correlated completed terminal was
never read and the 8x2x10 matrix re-fired cell 1 forever, emitting zero
non-degenerate rows. A prior two-strike diagnosis proved the omnimarket handler is
correct (it opens both terminal sessions pre-publish and races them); the defect is
in the omnibase_infra runtime consume leg.

Root cause: _assign_direct_terminal_partitions ignored the metadata future returned
by AIOKafkaClient.set_topics and re-called set_topics([same_topic]) each loop
iteration. aiokafka 0.13.0 set_topics only forces a metadata refresh when the topic
set DIFFERS from the tracked set, so every iteration after the first took the no-op
branch and never re-fetched. An ephemeral group_id=None consumer whose first metadata
fetch had not yet surfaced partitions burned the full 30s assign cap and raised a
bare TimeoutError (the empty-message 'wait failed' seen live).

Fix: register the reply topic once and await that metadata fetch, then on each miss
force a fresh fetch via force_metadata_update (which always fetches) rather than the
no-op set_topics repeat. The assign-cap TimeoutError now carries a diagnostic message
instead of an empty one.

Test: tests/integration/test_terminal_consumer_concurrent_assign_race.py drives the
REAL TerminalEventConsumer (the object wired as event_consumer) through the REAL
open_direct_terminal_consumer/poll path with AIOKafkaConsumer monkeypatched to a fake
that faithfully models aiokafka 0.13.0 set_topics future + metadata-latency semantics.
RED before the fix (bare TimeoutError, the live empty-message signature); GREEN after.
K>=2 multi-trial variant asserts no worker-thread leak across trials.

Evidence-Source: <occ-sha-pending>
Evidence-Ticket: OMN-13012

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-13096): onex delegate single-command subcommand (Phase 2b) (#1967)

* feat(OMN-13096): onex delegate single-command subcommand

Add 'onex delegate "<prompt>" [--task-type X] [--max-tokens N]' as a
subcommand on the existing onex CLI (Phase 2b of the skill-output-suppression
slice, OMN-13089). The command wraps payload construction, node dispatch, and
result extraction internally and prints exactly one
ModelSkillResult[ModelDelegateSkillResponse] to stdout via the OMN-13094
receipt-mode path. RuntimeLocal logs go to the capture file + artifact store,
never to stdout; scratch payloads live under <state-root>/tmp/ with run_id
suffixes (never /tmp).

- cli_delegate.py: classify_task_type (keyword table from legacy skill md),
  payload write, contract resolve, run_receipt_mode dispatch
- register 'delegate' under onex.cli entry points
- exempt delegate_command from the >5-param patterns gate (same Click-callback
  rationale as run_node_by_name)
- 20 unit tests: classification, scratch-under-state-root, single typed
  receipt on stdout, zero INFO log leakage

omnibase_infra does NOT depend on omnimarket; the delegate node is resolved at
runtime via the onex.nodes entry-point group (registered by omnimarket).

* chore(OMN-13096): re-trigger deploy-gate after Evidence-Source set to OCC#2593

No code change — the deploy-gate workflow triggers on synchronize (not edited),
so the PR-body Evidence-Source fix needs a new commit to re-resolve the OCC ref
to the open PR head where contracts/OMN-13096.yaml (with deploy evidence) lives.

* chore(OMN-13096): re-trigger deploy-gate now that OCC#2593 merged to OCC dev

contracts/OMN-13096.yaml (with the dod-deploy-onex-delegate item) is now on
OCC dev, so the deploy-gate OCC-dev checkout resolves the contract + deploy
evidence.

* chore(OMN-13096): regenerate runner-image identity lock for pyproject entry-point add

Adding the 'delegate' onex.cli entry point to pyproject.toml changed the
dependency-manifest digest that scripts/ci/runner_image_identity.py folds into
the runner-image identity lock. Regenerate the lock so
tests/ci/test_runner_image_identity.py matches (was the only CI test failure;
unrelated environmental integration/perf failures excluded).

* test(OMN-13096): update backmerge identity assertions to regenerated lock digests

The runner-image identity lock was regenerated for the pyproject entry-point
add; this test hardcodes the expected identity_digest/shared_env_digest, so
update both to match the new lock (same maintenance OMN-13094 did).

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13118): tolerate partition-less reply topic in terminal consume leg (#1970)

The context-ROI runner opens one ephemeral group_id=None terminal consumer per
terminal topic BEFORE publishing each generation command (subscribe-before-publish,
OMN-13012/13038). The FAILED reply topic is only produced to on contract_passed=False;
in a battery where generations pass it has zero messages, so Redpanda never advertises
a partition for it. _assign_direct_terminal_partitions burned the full 30s assign cap
on every trial then raised a bare TimeoutError (the empty-message 'wait failed'),
stalling each of the 160 battery trials ~30s before the COMPLETED terminal could
correlate -> battery needs >80 min and never completes (verifier-confirmed wedge).

A partition-less reply topic is a valid steady state, not a 30s error:
- _assign_direct_terminal_partitions gives a bounded grace window for a topic that
  exists but is slow to surface metadata, then assigns whatever partitions exist
  (possibly none) and returns promptly instead of burning the cap and raising.
- poll_direct_terminal_consumer treats an empty assignment as 'no terminal will
  arrive here' (sleeps out its timeout, returns None) without calling getone() on
  an unassigned consumer.

Repro: tests/integration/test_terminal_consumer_battery_load_wedge.py drives the
REAL TerminalEventConsumer over K=10 x 2 cells x 2 arms; RED (40/40 trials block a
full assign cap) before the fix, GREEN after.

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13118): pin terminal-consumer read offset synchronously (close lazy seek_to_end publish race) (#1971)

* fix(OMN-13118): pin terminal-consumer read offset synchronously to close lazy seek_to_end publish race

The consume-leg wedge survived PR #1969 (set_topics no-op) and PR #1970
(partition-less assign-cap) because both addressed the assign phase, not the
seek timing. AIOKafkaConsumer.seek_to_end is LAZY: it requests a LATEST offset
reset that only resolves on the first poll — AFTER the caller publishes. With
generation completing in ~1s, the correlated COMPLETED terminal lands in the
open->poll gap, so the lazily-resolved LATEST position is the HWM AFTER the
record and the poll reads past it. The terminal is never read, the trial never
correlates, and the experiment matrix re-fires the same cell forever.

Replace seek_to_end with a synchronous end_offsets() + seek() pin
(_pin_direct_terminal_end_offsets) in both open_direct_terminal_consumer and
RuntimePatternBBroker._dispatch_and_wait_with_direct_kafka_consumer, so the
read position is fixed at open() time, before the publish — the real
subscribe-before-publish guarantee. Empty assignment (partition-less reply
topic, OMN-13118 #1970) is a no-op.

Repro: tests/integration/test_terminal_consumer_seek_reset_race.py drives the
real TerminalEventConsumer.open()/wait() the way HandlerContextRoiRunner does,
publishing the correlated terminal in the open->wait gap across K=10 x 2 cells
x 2 arms; RED with the lazy reset (every cell degenerate), GREEN once the read
offset is pinned. RED verified by git-stashing only the source fix.

Existing consume-leg fakes updated to model end_offsets/seek (they previously
masked the bug by making seek_to_end a synchronous exact snapshot).

* test(OMN-13118): reword assertion (lazy not deferred) for receipt honesty gate

* fix(OMN-13118): bound end_offsets() round-trip with assign-cap timeout (CodeRabbit)

end_offsets() is a broker ListOffsets round-trip aiokafka documents as able to
block indefinitely. Bound it with the same cap as start()/assign so a stalled
broker fails fast instead of hanging the pre-publish positioning.

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13137): _validate_routing strategy-aware event_model checks (#1973)

operation_match routes by the `operation` field and does not use event_model.
The validator was unconditionally requiring event_model.{name,module} for every
handler entry, causing 230/295 omnimarket operation_match contracts (all correct
as authored) to fail routing validation at startup.

Fix: read routing_strategy from the routing map and branch validation:
  - payload_type_match → require event_model.{name, module} (unchanged)
  - operation_match (and any non-payload strategy) → require `operation`;
    skip event_model checks entirely

Updated pre-existing _validate_routing tests to declare routing_strategy:
payload_type_match explicitly (they always tested payload_type_match semantics
but relied on the implicit fallback that is now removed).

Added test_validate_routing_operation_match.py with 4 unit tests:
  1. operation_match without event_model → zero event_model errors
  2. operation_match missing operation field → error
  3. payload_type_match missing event_model → still errors (regression guard)
  4. Real node_integration_sweep_orchestrator routing block → clean (boot gate)

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13118): independent per-terminal-topic consumers in Pattern B direct-Kafka wait (#1972)

The consume-leg wedge survived four merged fixes (#1969 set_topics no-op,
#1970 partition-less assign-cap, #1971 synchronous seek-pin). The STRONG K>=10
multi-cell reprobe on the stability lane still wedged on REBUILD-5
(cdf53d963f7b). Converged diagnosis (strikes 3+4,
docs/evidence/2026-06-12-weekend-pass/experiments/probe4-stability/
reprobe-K10-rebuild5/HALT_K10_WEDGE_PERSISTS.md): the runtime waited for each
trial's terminal across TWO topics (node-generation-completed.v1 +
node-generation-failed.v1) with a SINGLE ephemeral group_id=None consumer
assigned both topics' partitions. One aiokafka consumer holds one manual
subscription; the COMPLETED delivery window collapsed before it surfaced the
correlated record, so the trial never correlated and the matrix re-fired cell 1.

RuntimePatternBBroker._dispatch_and_wait_with_direct_kafka_consumer now opens
ONE independent AIOKafkaConsumer PER terminal topic via
open_direct_terminal_consumer (each started, assigned, and offset-pinned via
end_offsets()+seek() at open() BEFORE publish), then awaits both CONCURRENTLY
via asyncio.wait(FIRST_COMPLETED). The first correlated terminal wins; both are
torn down. No shared consumer, no subscription flip.

Keeps the #1970 partition-less no-op and the #1971 synchronous offset pin (both
live in open_direct_terminal_consumer / poll_direct_terminal_consumer). Removes
the now-dead single-consumer helpers (_assign_terminal_topic_partitions,
method-level _refresh_terminal_topic_metadata, _direct_kafka_* kwargs builders,
_kafka_bootstrap_servers/_kafka_event_bus).

Adds tests/integration/test_terminal_consumer_subscription_flip_wedge.py: a
real-dispatch-path K>=10 x 2-cell x 2-arm repro whose fake models TWO
independent consumers honestly (delivery is faithful only for a single-topic
assignment; a consumer spanning both topics flips and drops the COMPLETED
record). RED genuineness verified by reverting only the source to the
single-consumer shape (test hangs past timeout); GREEN with the fix in 2.3s.

Acceptance is the LIVE K>=10 multi-cell stability-lane reprobe (later phase),
NOT this unit test. A green unit repro is necessary but NOT sufficient.

Refs OMN-13118, OMN-13128.

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13118): long-lived terminal correlator replaces per-trial ephemeral consume leg (#1974)

Tier B canonical redesign (epic OMN-12525). Five offset/subscription patches
(#1969-#1972) tuned the per-trial-ephemeral terminal consumer and all failed
the live K>=10 batt…
…+ OMN-13666 entrypoint tolerance (emergency prod recovery) (#2128)

* fix(OMN-13670): surgical main hotfix — re-apply OMN-13654 dep floors + OMN-13666 entrypoint tolerance

EMERGENCY PROD RECOVERY — OMN-13670 operator-authorized. NOT a normal dev->main promotion.

OMN-13654: Bump pyjwt>=2.13.0, python-multipart>=0.0.30, starlette>=1.3.1 as explicit
security-floor constraints in pyproject.toml. Relocks uv.lock (pyjwt 2.13.0,
python-multipart 0.0.32, starlette 1.3.1). Resolves 5 HIGH CVEs that have been
blocking build-and-push-runtime.yml since 2026-06-07 with Trivy exit-code 1.

OMN-13666: Runtime entrypoint now treats PRIMARY (omnibase_infra) stamp as REQUIRED
(failure aborts boot, exit 1) and SECONDARY (omniintelligence) stamp as BEST-EFFORT
(failure logs WARNING, boot continues). Resolves the prod crash-loop where "permission
denied for table db_metadata" on the omniintelligence DB was bringing all 7 prod
runtime deployments to 0/1.

* fix(OMN-13670): refresh runner image identity lock

* chore(OMN-13670): refresh hotfix check context

* fix(OMN-13670): compare node migrations against main for main hotfix

* fix(OMN-13670): align runner identity backmerge guard

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
… last Trivy HIGH (emergency prod recovery) (#2131)

* fix(OMN-13670): floor cryptography>=48.0.1 in runtime builder — clear last Trivy HIGH

Adds build-time pip floor for cryptography>=48.0.1 in the BUILDER stage of
docker/Dockerfile.runtime, matching the existing protobuf/setuptools floor
pattern. Clears GHSA-537c-gmf6-5ccf (HIGH, cryptography 46.0.7→48.0.1).
This is the final floor completing the clean-main Trivy fix for OMN-13670.

Also includes yamlfmt and ruff format auto-fixes on pre-existing files.

* chore(OMN-13670): refresh hotfix checks

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
…t runner→ECR EOF (emergency prod recovery) (#2134)

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
….46.1 / spi 0.23.0) (#2164)

* docs(OMN-12962): contract-store durability audit — cold-runtime census proof (#1937)

Proves cold-start contract census reconstructs from the image-bundled
filesystem manifest (HYBRID-mode bootstrap + PluginLoaderContractSource),
independent of node-registration.v1 retention. The delete-retention topic
feeds only the post-freeze dynamic listener (auto_offset_reset=latest, no
history replay). Live .201 stability-test evidence: filesystem contract_path
in manifest + truncated topic log head with intact census. No store fix
needed; residual dynamic-only gap covered by runtime_sweep sweep check.

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12970): vendor omnimarket projection node migrations into forward/nodes (#1938)

Vendors omnimarket node-owned projection migrations into the namespaced
forward-migration tree so run-forward-migrations.sh materializes them in the
dashboard projection DB (omnidash_analytics) at deploy.

Primary (OMN-12970): creates llm_call_metrics, llm_cost_aggregates, and
capability_scores in the projection DB. These were only ever created in the
omnibase_infra DB by infra migrations 031/060, so the ab-compare,
cost.token_usage, cost.summary, and capability-scores projection topics were
DEGRADED at startup ('table not found') and their dashboard panels rendered
empty.

Also re-syncs three omnimarket node migrations the vendor tree had drifted from
(node_projection_llm_routing, node_projection_overnight, node_projection_savings
/077) — sync-node-migrations.sh --check requires the full vendored tree to match
omnimarket source, and these were missing.

Companion to omnimarket PR for the same ticket (source migrations + projection
table-coverage ratchet test).

Evidence-Ticket: OMN-12970

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12965): stamp runtime image identity (version + revision) in workspace builds (#1943)

* fix(OMN-12965): stamp runtime image identity (version + revision) in workspace builds

The main runtime image stamped org.opencontainers.image.version=0.1.0 with a
blank org.opencontainers.image.revision after workspace rebuilds. A blank
identity degrades every proof packet (runtime SHA + image digest are required
citations in accepted evidence).

Root causes (three build paths under-stamped identity):
- onex up --build (cmd_up) passed only GIT_SHA; the runtime-stage OCI labels
  read VCS_REF (-> blank revision) and RUNTIME_VERSION (-> placeholder 0.1.0).
- deploy-runtime.sh passed VCS_REF but not RUNTIME_VERSION/GIT_SHA.
- Dockerfile silently allowed blank/placeholder identity in workspace mode.

Fix:
- cli._image_identity_build_args() stamps the full quad (GIT_SHA/VCS_REF/
  RUNTIME_VERSION/BUILD_DATE) and fails fast on an unresolved git revision.
- deploy-runtime.sh stamps RUNTIME_VERSION + GIT_SHA and verifies the version
  label is non-placeholder post-deploy.
- Dockerfile.runtime fails workspace builds with blank VCS_REF or placeholder
  RUNTIME_VERSION=0.1.0 (release mode unaffected).

Enforcement ratchet (same PR):
- scripts/check_runtime_image_identity.py static check, wired as pre-commit hook
  + CI gate (ci.yml).
- tests/unit/infra/test_runtime_image_identity_labels.py pins the cli helpers +
  Dockerfile guard; deploy-agent test extended for the quad.

Proven locally via throwaway docker builds: workspace+args -> populated labels;
workspace without args -> guard fails (exit 64); release without args -> 0.1.0
placeholder allowed (no regression).

Evidence-Ticket: OMN-12965

* test(OMN-12965): integration build proof for runtime image identity labels

Builds the real runtime-stage ARG/LABEL/guard block against busybox and asserts
via docker inspect: workspace+args -> populated version/revision; workspace
without args -> guard fails (exit 64); release -> placeholder allowed. Satisfies
the integration-test hard gate and makes the throwaway proof permanent.

Evidence-Ticket: OMN-12965

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12987): workspace-build sibling lock-pin preflight + provenance (#1944)

Recurrence ratchet for the 2026-06-11 stability bootstrap crash. The 11:20Z
--no-cache rebuild vendored omnibase_infra 0.37.0-dev (~2c1d672f) + core 0.42.0
even though omnimarket dev's uv.lock pinned infra 0.38.1 @ e2dbdc95 + core
0.44.0 @ c97c2c9a. The stale sibling predated the OMN-12501 Protocol-quarantine
guard, turning a latent contract defect into a fatal crash that crash-looped the
main runtime.

- check_sibling_lock_pins.py: host-side fail-fast preflight resolving expected
  sibling versions/SHAs from the consuming repo's (omnimarket) uv.lock and
  comparing them against each vendored tree. Mismatch aborts the build.
- stage_workspace.sh: emit a .build-sha marker per staged sibling (rsync drops
  .git) so the preflight and provenance can identify the vendored commit.
- deploy-runtime.sh: run the preflight after staging, before build; abort on
  mismatch. Write the comparison under sibling-repos/.sibling-lock-pins.json.
- compute_workspace_provenance.py + Dockerfile.runtime: fold expected-vs-actual
  lock_pin_comparison into build-provenance.json for deploy verifiers.

Recurrence-guard tests prove a stale infra 0.37.0 vs lock-pinned 0.38.1 fails
the preflight and matched pins pass; deploy script wiring is asserted statically.

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12988): pin stability runtime-worker to replicas 1 + census ratchet (#1942)

The base docker-compose.infra.yml defaults runtime-worker to
replicas 0 (${WORKER_REPLICAS:-0}). The stability-test lane's required
state includes a running worker (4-container census: main, effects,
worker, projection-api), but the override pinned it via an
env-interpolation default (${STABILITY_TEST_WORKER_REPLICAS:-1}) — a
silent-drop surface: a stray exported STABILITY_TEST_WORKER_REPLICAS=0
or removal of the :-1 fallback would scale the worker to 0 with zero
signal on a plain compose up/recreate.

Fix: pin docker-compose.stability-test.yml runtime-worker
deploy.replicas to the literal 1 (no env indirection).

Ratchet (recurrence guards, same PR):
- scripts/deploy-agent/tests/unit/test_runtime_worker_census.py: assert
  runtime-worker stays in the deploy-agent RUNTIME-scope census so a
  missing worker (replicas 0 => absent from docker compose ps) is a
  deploy failure, not silence; assert the override pins a literal 1.
- tests/integration/infra/test_stability_test_runtime_compose_render.py:
  assert the rendered stability worker resolves deploy.replicas == 1.
- tests/unit/infra/test_stability_test_runtime_lane.py: update the
  existing pin assertion to the literal 1.

Evidence-Ticket: OMN-12988
Config-drift family: OMN-12945

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12979): expire-bound topic completeness suppressions (#1940)

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12966): accept community-edition Infisical /api/status in provision readiness gate (#1939)

The fresh-provision readiness gate in provision-infisical.py only accepted the
enterprise {"status": "ok"} payload and rejected the community edition's
{"message": "Ok"}, returning 1 before bootstrap could run. This blocked
provisioning against the Infisical instance deployed on .201 (community edition).

Route the gate through the existing _is_infisical_ready helper (single source of
truth, already used by the already-provisioned path). Add TestMainFreshProvision-
ReadinessGate covering community/enterprise/not-ready cases.

Also adds docker/docker-compose.infisical-stability.yml: an ADDITIVE Infisical
compose project for the stability-test lane (joins the existing network as
external, reuses stability postgres/valkey, no lane mutation), since the lane
overlays disable the in-lane Infisical service via *-disabled profile overrides.

P1.2b-A: Infisical now reachable from the stability runtime/effect containers; a
known secret (OMN_12966_PROBE) seeds and resolves end-to-end via the machine-
identity universal-auth path, verified from inside the effects container.

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-12958): volume-config drift gate + runtime config provenance (#1941)

* feat(OMN-12958): volume-config drift gate + runtime provenance

Compute config provenance (path + sha256) for the runtime-rendered Bifrost
delegation contract; the deployed volume copy survives rebuilds and silently
diverges from packaged source (two competing authorities, OMN-12945).

- runtime/config_provenance.py: ModelConfigProvenance + drift classification,
  sidecar JSON writer (read by sweep + proof packets)
- runtime/health/health_config_provenance.py: drift -> degraded health
- render entrypoint logs provenance line + writes sidecar on every boot
- docs/runbooks/volume-config-drift-and-reseed.md: ledgered re-seed procedure
- validation exemption for config_name (logical identifier, not entity ref)

No live volume mutation: re-seed is an operator deploy step (deploy_pending).

* test(OMN-12958): cover volume config drift reseed flow

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-12957): wire runtime-profile validator + core registry parity guard (#1945)

* feat(OMN-12957): wire runtime-profile validator + core registry parity guard

- Remove the import-time RuntimeError drift raise in runtime_profile.py (a hard
  raise on core/infra version skew would crash the kernel at import); the parity
  invariant is enforced by test_profiles_match_core_registry instead.
- Add tests: _PROFILES keys == omnibase_core REGISTERED_RUNTIME_PROFILES, and
  every CONSUMER_ATTACHED_RUNTIME_PROFILES profile loads as a real lane.
- Wire omnibase_core.validation.validator_runtime_profiles as a pre-commit hook
  + validator-runtime-profiles.yml CI gate on infra contracts.
- Freeze 19 pre-existing violators in validation/runtime_profiles_allowlist.yaml
  (discovered by repo-root walk; drain via OMN-12982). Blocks NEW orphans.

Requires the omnibase_core pin to include OMN-12957's validator (new rules).

Evidence-Ticket: OMN-12957
Evidence-Source: 5463fbaf819409d4fb7f491dd4f276f10d869eba

* ci(OMN-12957): pass runtime profile allowlist to validator

* test(OMN-12957): cover runtime profile registry parity

* fix(OMN-12957): keep runtime profile allowlist under config

* fix(OMN-12957): pin core runtime profile registry

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12973): widen migration-gate healthcheck start_period + ratchet (#1950)

P2.8: classify the prod migration-gate UNHEALTHY-then-self-resolved incident.

Classification: idle-one-shot-mis-modeled = NO. The migration-gate is a
correctly-modeled long-running sentinel (entrypoint `while true; sleep 3600`)
whose healthcheck continuously polls db_metadata.migrations_complete via
check_migrations_complete.sh. The container flipped UNHEALTHY transiently because
its healthcheck start_period (10s) was far shorter than the real cold-volume
migration window (~116s: prod gate started 09:35:22, intelligence-migration
finished 09:37:18). Past the 10s grace window the still-failing probe was
reported UNHEALTHY until migrations completed, then self-resolved — no fault.

Fix: raise migration-gate healthcheck start_period 10s -> 180s in both the
authoritative catalog manifest (docker/catalog/services/migration-gate.yaml,
flows into the generated compose) and the hand-maintained
docker/docker-compose.infra.yml that deploy-runtime.sh applies to .201, so a
still-applying gate stays in `health: starting` instead of flipping UNHEALTHY.

Ratchet (enforcement, not detection): new ValidatorHealthcheckStartPeriod in
omnibase_infra catalog (validator_healthcheck_start_period.py) asserts a 120s
floor for migration-completion gates, wired into `onex validate runtime`
(cmd_validate_runtime) AND backed by unit tests that gate every PR via
pre-commit + CI. A migration-completion gate is identified by BOTH a healthcheck
that polls migration completion AND a service_completed_successfully dependency,
so ordinary app services (e.g. intelligence-api, an HTTP liveness probe with a
legitimately short 40s start_period) are not swept into the floor.

Prod probed read-only only; no prod mutation. Applies to live prod via the
batched stability/prod rebuild (deploy_pending).

Evidence-Ticket: OMN-12973
Evidence-Source: OCC#PENDING

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-12971): deliver Vertex ADC bearer token to runtime effects container (secret-ref discipline) (#1949)

* feat(OMN-12971): deliver Vertex ADC bearer token to runtime effects container (secret-ref discipline)

Wire the Vertex $500-credit ADC path into the runtime, ADDITIVE next to the
Gemini API-key path (provider-agnostic; neither provider removed or forced).

runtime_policy.contract.yaml: add a secret-source mapping llm.vertex.access_token
(source_type=file, /run/secrets/vertex_access_token) to the dev/stability-test/
judge profiles, alongside the existing llm.gemini.api_key env mapping. The ref
name MUST match cloud-vertex-gemini.secret_ref in omnimarket
bifrost_delegation.yaml. The resolved VALUE is a short-lived OAuth bearer token
minted from ADC, refreshed by the operator; the token VALUE is never committed —
only the ref name + in-container path. Add aiplatform.googleapis.com to the
cloud host allowlist.

docker-compose.infra.yml: bind the operator-supplied host token file read-only to
/run/secrets/vertex_access_token on the main and effects runtimes
(VERTEX_ACCESS_TOKEN_HOST_FILE, default /dev/null so lanes without Vertex still
start; Gemini key path unaffected). Pass through BIFROST_VERTEX_GEMINI_ENDPOINT_URL
(overlay supplies the complete Vertex OpenAI-compat URL) and
GOOGLE_CLOUD_PROJECT/LOCATION (default empty).

runtime-policy.env: regenerated from the contract via render_runtime_policy_env
(test_runtime_policy_env_matches_contract_renderer proves contract<->env parity).

test_runtime_policy_contract.py: update host-allowlist assertion for the additive
Vertex host.

Pre-existing platform-wide topic-parity-gate failures (25 unrelated topics) are
identical with and without this change (proven by diff); that gate is pre-commit-
only (not a CI merge gate) and this change adds zero new topic gaps, so that one
hook is SKIP-ped. No deploy/receipt/merge gate is bypassed.

* fix(OMN-12971): make Vertex runtime env contract-owned

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-13008): automated .201 disk maintenance — worktree GC + docker GC + watermark alert ratchet (#1952)

* feat(OMN-13008): automated .201 disk maintenance — worktree GC + docker GC + watermark alert ratchet

Conservative, keep-list-driven disk maintenance to prevent the 2026-06-11
/data ~95% outage that killed all three lanes mid-demo:

- scripts/worktree-gc.sh: drives the canonical omniclaude prune-worktrees.sh
  (merged+clean+pushed safety) on both Mac (merge-sweep tick) and .201
- scripts/disk-gc.sh + disk_gc_plan.py: conservative docker/builder/image GC.
  Pure, testable removal planner honoring a VERSIONED keep-list
  (deploy/disk-gc/keep-list.yaml): never reaps a kept repo, kept tag, in-use
  image, or anything younger than min_age_days; keeps N superseded generations.
- scripts/disk-watermark-check.sh + disk_watermark_event.py: df watermark
  ratchet. >=85% emits a typed disk-watermark bus event (warning) that the
  sweep auto-ticket path turns into a Linear ticket; >=90% emits critical.
  Broker addr is fail-fast from KAFKA_BOOTSTRAP_SERVERS (no localhost default).
- deploy/disk-gc/: systemd USER timer (onex-disk-gc.timer/.service, hourly) +
  install-disk-gc.sh. User units, NOT lane containers.
- tests: 20 unit tests incl. GC plan-safety invariants + dry-run proof that
  default mode issues no destructive op (a wrong-delete GC is worse than none).

Contract: contracts/OMN-13008.yaml

* fix(OMN-13008): pass docker inventory to GC planner via stdin, not env (ARG_MAX)

On a host with many docker images, passing the full image/ps inventory as env
vars to disk_gc_plan.py exceeds ARG_MAX ('Argument list too long', exit 126),
producing an empty plan. Write inventory to per-run scratch files (under the log
dir, never /tmp; cleaned on exit) and hand it to the planner on stdin as a JSON
envelope. Verified the failure live on .201; planner now reads stdin.

* fix(OMN-13008): simplify GC plan stdin pipe (two processes, no nested subprocess)

* fix(OMN-13008): keep-wins reconciliation — never remove an image id with any keep reason

A single image id can surface in multiple 'docker image ls' rows (one per
repo:tag). One tag could route the id to dangling-removal while another routes
it to keep (e.g. tagged 'latest' or within-N-generations). Live .201 plan showed
an id in BOTH remove_image_ids and kept_reasons. Reconcile at the end: keep wins
— any id with a keep reason is dropped from the remove list; remove list deduped.
Adds 2 regression tests. Verified live on .201.

* fix(OMN-13008): timer uses OnCalendar=hourly for reliable re-arm (oneshot service)

OnUnitActiveSec does not reliably re-elapse for a oneshot service once it goes
inactive (observed NextElapseUSecMonotonic=infinity live on .201 after the first
run). Switch to OnCalendar=hourly + RandomizedDelaySec so the timer re-arms every
hour. Keep Persistent=true for missed-run catch-up.

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13005): materialize blocking event_consumer in runtime auto-wiring (runner consume-leg degenerate rows) (#1951)

* fix(OMN-13005): materialize blocking event_consumer in auto-wiring (runner consume-leg degenerate rows)

The runtime auto-wiring materialized event_publisher for handlers that
declare it but had no equivalent for event_consumer. Request/response
EFFECT handlers (HandlerContextRoiRunner) that publish a command then
block on the correlated terminal event fell back to their no-op consumer
default, returning None immediately -> every result row degenerate
(failure_stage=generation, attempt_count=0) while generations succeeded
~1s later.

Adds _make_sync_event_consumer (mirror of _make_sync_event_publisher),
backed by service_terminal_event_consumer.make_terminal_event_consumer:
a sync (topic, correlation_id, timeout) -> dict | None adapter that runs
the proven direct-Kafka correlate-and-wait loop (from RuntimePatternBBroker)
on an isolated event loop in a worker thread, so blocking does not deadlock
the runtime dispatch loop that delivers the awaited terminal.

TDD through the REAL dispatch path: test_event_consumer_injection drives a
trial through _prepare_handler_wiring with a terminal arriving after a delay
and asserts a non-degenerate row; verified RED with injection disabled.

* fix(OMN-13012): two-phase (seek-now/wait-later) terminal event_consumer to close the subscribe-after-publish race (#1954)

The OMN-13005 injected event_consumer is a single callable that does
assign -> seek_to_end -> poll internally, all AFTER the handler has
already published its command. Once OMN-13010 freed the dispatch loop and
generation began completing in ~1s, the correlated terminal lands BEFORE
the single-call consumer's post-publish seek_to_end positions, so
seek_to_end skips PAST the already-emitted terminal and the runner times
out on an offset beyond it (probe3, run_id=20260611T2140Z-probe3 -- both
arms degenerate, zero rebalances).

Splits positioning from waiting so the caller subscribes BEFORE it
publishes:

  session = consumer.open(topic)        # assign + seek_to_end NOW
  publisher(command_topic, payload)     # publish AFTER positioning
  payload = session.wait(cid, timeout)  # block from the captured position

The returned TerminalEventConsumer is still directly callable with the
legacy (topic, cid, timeout) -> dict | None single-call shape for any
consumer that does not need subscribe-before-publish; the runner is the
only consumer today. TerminalConsumerSession owns a dedicated event loop
on a daemon worker thread for the whole open->wait->close lifecycle,
preserving the OMN-13005 loop-isolation discipline so blocking never
deadlocks the runtime dispatch loop.

TDD (real dispatch path, RED-then-GREEN): extends the OMN-13005 injection
test with a terminal emitted IMMEDIATELY after publish. The single-call
(seek-after-publish) consumer MISSES it (degenerate row -- RED test
asserts failure_stage=generation); the two-phase (open-before-publish)
consumer CATCHES it (non-degenerate -- GREEN). The Kafka layer is faked at
the two service seams against a shared in-memory log modeling seek-to-end
semantics. OMN-13005 blocking-correlate behavior preserved. 269/269
auto_wiring unit tests pass; mypy --strict clean.

Sibling to OMN-13010 / OMN-13005 / OMN-13003.

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13005): route terminal consumer through Kafka boundary

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12990): pin worker replicas in ledgered config + fail-fast lane overrides (silent-drop ratchet) (#1948)

* fix(OMN-12990): pin worker replicas in ledgered config + fail-fast lane overrides

The base compose sets runtime-worker deploy replicas to ${WORKER_REPLICAS:-0}
(soft-default ZERO). The stability lane's required state includes a running
worker (GATE_ZERO_PROOF.md: 4 runtime containers), but the worker presence rode
entirely on a compose soft-default — any plain compose up/recreate without the
policy env silently scaled the worker to zero with no error and no signal.

Fix (contract-native + fail-fast):
- Add 'replicas' to ModelRuntimeProcessPolicy; pin replicas: 1 in every lane's
  worker block in runtime_policy.contract.yaml.
- Renderer emits {PROFILE}_WORKER_REPLICAS into the ledgered runtime-policy.env
  for dev/stability-test/judge/prod.
- stability + prod compose overrides reference ${..._WORKER_REPLICAS:?...}
  (fail-fast, NO silent :-1/:-0 default). A recreate that omits the policy env
  now aborts loudly instead of dropping the worker. prod previously had no
  override at all and inherited the dangerous :-0 default.

Ratchet (recurrence guards):
- tests asserting fail-fast override form (no soft default), contract-declared
  replica pin >= 1 per lane, and rendered {PROFILE}_WORKER_REPLICAS=1 in the
  ledgered env.
- runbook deploy/verify procedure adds an expected-container census (worker
  must be present) via verify_container_manifest; a missing worker is a FAILURE,
  not silence.

Note: SKIP=topic-parity-gate — that local-only advisory gate (absent from all
.github/workflows, not a required CI check) fails on 25 pre-existing cross-repo
topic gaps (build-loop/omniclaude/omniweb) identical on pristine base HEAD
8d7da1249; this change adds zero topics. All other hooks ran clean.

Evidence-Ticket: OMN-12990

* test(OMN-12990): cover worker replica policy integration

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-12909): add gateway bus forwarder P0A (#1946)

* feat(OMN-12909): add gateway bus forwarder p0a

* test(OMN-12909): add gateway forwarder integration coverage

* fix(OMN-12909): satisfy gateway forwarder validators

* test(OMN-12909): allow gateway forwarder bus protocol

* fix(OMN-12909): sync gateway forwarder entry point

* fix(OMN-12909): refresh runner image identity lock

* test(OMN-12909): relax JSON normalizer mixed benchmark threshold

* fix(OMN-12909): allow gateway handlers to boot unconfigured

* fix(OMN-12909): declare gateway forwarder runtime profile

* fix(OMN-12909): update runner identity backmerge expectation

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-13011): LANE CENSUS RECONCILIATION ratchet — declared desired-state per lane, drift = auto-ticket (#1955)

The class fix for the recurring lane-drift regression. Nothing reconciled the
declared desired state of a runtime lane against what is actually running, so the
same failure kept recurring with zero signal: volume config drift (OMN-12945),
WORKER_REPLICAS silent zero (OMN-12988/12990), and on 2026-06-11 prod runtime
containers plus the broker network were silently absent for hours during demo prep.

Ships a per-lane DESIRED-STATE census:
- (a) DECLARED in a versioned lane manifest (deploy/lane-census/lane-manifest.yaml):
  container set, network, replicas, image-tag pattern per lane
  (stability-test/prod/judge/dev), derived from the canonical compose lane files.
  A parity ratchet keeps the manifest locked in step with the compose files.
- (b) RECONCILED on a schedule on .201 by SHARING the OMN-13008 systemd timer
  (a drop-in 4th ExecStart on onex-disk-gc.service — never a second timer) and
  on-demand via scripts/lane-census-check.sh / runtime_sweep.
- (c) Drift = typed bus event (onex.evt.infra.lane-census-drift.v1) + Linear
  auto-ticket naming exactly what is missing/extra (container_absent,
  network_detached, replicas_zero, unexpected_container, oneshot_failed/stuck,
  image_tag_mismatch). Fail-fast, no warn-only mode (gates-block policy); exit 30
  on drift; bus publish fail-fast on KAFKA_BOOTSTRAP_SERVERS (no localhost default).

Red fixture reproduces 2026-06-11: prod runtime containers absent + broker network
detached must produce the exact drift findings + a non-zero exit hours before a
human noticed. Pure planner is fully unit-tested; shell driver dry-run-tested.

Builds on the OMN-12988 deploy-agent RUNTIME census (deploy-time) as the
complementary steady-state reconciler; closes the runtime-worker.yaml
container_name: null census gap by sourcing names from the compose lane files.

Evidence-Ticket: OMN-13011
Config-drift family: OMN-12945
Relates-to: OMN-13009, OMN-12988, OMN-13008

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13020): vendor missing node migrations — llm_routing 0000 + context_roi 001 (#1956)

Vendors two omnimarket node-source migrations into the infra forward-migration
tree via scripts/sync-node-migrations.sh (the canonical OMN-12559 mechanism):

- node_projection_llm_routing/0000_create_llm_routing_decisions.sql
  (source: omnimarket #1168 / OMN-12942, merge ed6734f8)
- node_projection_context_roi/001_create_context_roi_scores.sql
  (source: omnimarket #1178 / OMN-12955, merge 5010b1f4)

Without the 0000 base table, node_projection_llm_routing/0001 (CREATE VIEW)
hard-fails against NODE_POSTGRES_DB=omnidash_analytics — exactly the prod
forward-migration exit-3 of 2026-06-11T09:35:52Z, and reproduced by
construction in any clean clones@dev build. Files are byte-identical to the
omnimarket dev blobs (sha256 f8a8b339… / c4126e65…) and to the untracked
hot-patch copies on the .201 stability clone.

Both migrations are self-contained, all-statements-IF-NOT-EXISTS, and 0000
sorts lexically before 0001 within the node's namespaced identity space.

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13058): close TerminalConsumerSession on open() failure (worker thread + event loop leak) (#1957)

TerminalConsumerSession.__init__ starts its dedicated worker loop thread
immediately. TerminalEventConsumer.open did
'session = TerminalConsumerSession(...); return session.open()' with no
cleanup: any failure inside session.open() (consumer start timeout,
partition-assign timeout, broker auth error) propagated out of the raising
expression, the session reference was lost, and the daemon worker thread plus
its never-closed asyncio event loop leaked -- one pair per failed open. The
motivating caller (HandlerContextRoiRunner) opens a session per trial, so a
160-560-trial battery against a degraded broker accumulates hundreds of
leaked threads in the long-lived effects container.

Fix: wrap session.open() in try/except BaseException -> session.close()
(idempotent: stops the loop, joins the thread) -> re-raise. Covers both the
two-phase .open(topic) path and the legacy single-call __call__ path.

Found by the P3.3 doctrinal review of merged #1951 (b9712af9 / 36d98275).

TDD: tests/unit/runtime/test_service_terminal_event_consumer_open_failure.py
injects a real open failure through the production path (event bus without
_bootstrap_servers) and asserts no alive terminal-consumer-* thread after the
raise. Verified RED with the fix stashed (2 failed), GREEN with it (2 passed).

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-13021): non-dev-base guard — fail feature-base PRs absent Stacked-Parent declaration (retro A-6) (#1958)

Any PR whose base is neither dev nor main fails unless the body carries
'Stacked-Parent: #N'. Prevents the feedback_stacked_prs_orphan_from_dev class
(#1185/#1954 auto-merged INTO parent feature branches and stranded off dev).
base=main remains governed by main-target-guard.

Epic OMN-13013 (process enforcement ratchets — June 12 retro).

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-13014): hot-patch ledger rebuild preflight gate (retro B-1) (#1959)

* feat(OMN-13014): hot-patch ledger rebuild preflight gate (retro B-1)

Hot-patches on .201 (.prepatch sibling discipline) silently revert on any
image rebuild/force-recreate — the 2026-06-11 20:58Z rebuild already erased
a live /api/generate patch once. This adds the rebuild-path gate:

- scripts/preflight_hotpatch_ledger.py: given a target container or lane +
  per-repo build refs, hard-fails when any hot-patch ledger row's source PR
  merge commit is not an ancestor of the build ref (git merge-base
  --is-ancestor), plus a .prepatch tripwire of the running container
  (unledgered .prepatch = hard fail; --post-rebuild = zero .prepatch
  expected). Sole bypass: HOTPATCH_PREFLIGHT_BYPASS carrying the Rule-10
  '# skip-token-allowed: <user-approval-receipt-id>' form.
- scripts/deploy-runtime.sh: guard_hotpatch_ledger wired into main() before
  build/preview (both dry-run and execute), lane derived from the compose
  project; skips loudly only when no ledger exists on the host.
- tests/unit/scripts/test_preflight_hotpatch_ledger.py: 17 unit tests
  (ancestor gate, lane scoping, ledger loading, tripwire, bypass forms).
- tests/ci/test_receipt_gate_install_guard.py: repair stale guard — core
  OMN-12565 replaced the OMN-9198 'uv pip uninstall first' install step with
  a cleared workspace venv (uv venv --clear); assert the new contract.

Ledger backfilled from live census (5 .prepatch files / 4 source PRs, all
MERGED to dev) at /data/omninode/hotpatch-ledger/ledger.yaml on .201.

* fix(OMN-13014): scope missing-prepatch tripwire warning to the probed container

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12977): workspace-build sibling-pin ratchet — honor consuming lock, fail-fast on stale vendor (#1936)

* fix(OMN-12977): workspace-build sibling-pin ratchet — honor consuming lock, fail-fast on stale vendor

The workspace-mode image build (BUILD_SOURCE=workspace, used by the stability-test
deploy procedure) vendored sibling/foundation packages from whatever the canonical
OMNI_HOME clones happened to be checked out at, ignoring the consuming repo's
uv.lock. On 2026-06-11 this shipped a 13-day-stale omnibase_infra 0.37.0-dev
(pre-OMN-12501 Protocol-quarantine guard) + core 0.42.0 against an omnimarket dev
lock pinning infra 0.38.1@e2dbdc95 / core 0.44.0@c97c2c9a, dropping the guard and
crashing wire_from_manifest bootstrap fatally (stability lane down on demo day).

Fix + recurrence ratchet (same PR):
- scripts/runtime_build/check_sibling_lock_pins.py: parse the consuming repo's
  uv.lock for expected version+git-rev of each foundation/sibling package
  (scoped to the package's own source line so editable/registry pins are not
  cross-attributed a dependency's rev), resolve the actual clone version+HEAD,
  compare, and classify drift backward/forward/none. Fail-fast (exit 1) on any
  drift; --allow-drift records an explicit operator override in the artifact,
  never silent.
- stage_workspace.sh: runs the preflight against the canonical clones before
  staging; aborts the build (exit 3) on unacknowledged drift and writes
  workspace/sibling-pin-comparison.json.
- compute_workspace_provenance.py: folds the expected-vs-actual comparison into
  build-provenance.json so deploy verifiers can assert the build honored the lock;
  flags unacknowledged drift as a provenance error.
- Dockerfile.runtime: COPY the comparison artifact (committed placeholder so the
  COPY always resolves; overwritten by stage_workspace.sh in workspace mode).
- TDD: 19 unit tests covering lock parsing (git/registry/editable sources),
  drift classification, the exact 0.37.0-vs-0.38.1 stale case, check_pins exit
  codes, and the allow-drift override. Pre-existing mypy-strict bare-dict errors
  in compute_workspace_provenance.py fixed in the same pass.

Evidence-Ticket: OMN-12977
Evidence-Source: pending-occ

* ci(OMN-12977): retry runtime smoke compose port race

* test(OMN-12977): align sibling-pin script tests with current API

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12989): workspace-mode image build must honor sibling lock pins (#1947)

* fix(OMN-12989): workspace build must honor sibling lock pins (fail-fast ratchet)

The 2026-06-11 stability bootstrap crash was caused by a workspace-mode
--no-cache rebuild that vendored omnibase_infra 0.37.0-dev (13-day-stale
worktree) even though omnimarket dev's uv.lock pins infra 0.38.1. The
downgraded sibling predated the OMN-12501 Protocol-quarantine guard and
turned a latent contract defect into a fatal crash.

Fix + ratchet (same PR set):
- scripts/runtime_build/resolve_workspace_pins.py: parse the consuming repo's
  uv.lock for sibling pins (version + git rev); classify each staged/installed
  sibling as exact/ahead/regression/unpinned; fail-fast (WorkspacePinError) on
  any regression BELOW the lock pin. Stdlib version-tuple fallback when
  packaging is absent so the ratchet never fails-open.
- compute_workspace_provenance.py: enforce sibling pins + a host-infra
  self-check (installed omnibase_infra vs lock pin — the exact crash vector,
  since host infra is built from the context, not staged), and emit a
  pin_comparison block into build-provenance.json for deploy verifiers.
- Dockerfile.runtime: COPY resolve_workspace_pins.py beside the provenance
  script so the in-image import resolves.
- TDD: failing tests first (test_resolve_workspace_pins.py, 11 cases) +
  provenance integration tests; fixed a pre-existing stale RUNTIME_VERSION
  assertion to read pyproject dynamically.

Evidence-Ticket: OMN-12989

* fix(OMN-12989): workspace build must honor sibling lock pins (fail-fast ratchet)

The 2026-06-11 stability bootstrap crash was caused by a workspace-mode
--no-cache rebuild that vendored omnibase_infra 0.37.0-dev (13-day-stale
worktree) even though omnimarket dev's uv.lock pins infra 0.38.1. The
downgraded sibling predated the OMN-12501 Protocol-quarantine guard and
turned a latent contract defect into a fatal crash.

Fix + ratchet (same PR set):
- scripts/runtime_build/resolve_workspace_pins.py: parse the consuming repo's
  uv.lock for sibling pins (version + git rev); classify each staged/installed
  sibling as exact/ahead/regression/unpinned; fail-fast (WorkspacePinError) on
  any regression BELOW the lock pin. Stdlib version-tuple fallback when
  packaging is absent so the ratchet never fails-open.
- compute_workspace_provenance.py: enforce sibling pins + a host-infra
  self-check (installed omnibase_infra vs lock pin — the exact crash vector,
  since host infra is built from the context, not staged), and emit a
  pin_comparison block into build-provenance.json for deploy verifiers.
- Dockerfile.runtime: COPY resolve_workspace_pins.py beside the provenance
  script so the in-image import resolves.
- TDD: failing tests first (test_resolve_workspace_pins.py, 11 cases) +
  provenance integration tests; fixed a pre-existing stale RUNTIME_VERSION
  assertion to read pyproject dynamically.

Evidence-Ticket: OMN-12989

* test(OMN-12989): co-locate provenance pin helper in fixture

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13055): absent repos in REPOS list warn and exit 0 instead of failing (#1960)

Missing repos (not cloned locally) now emit a WARN line and are tracked
in a separate WARNED array. Only real fetch/ff failures cause exit 1.
This makes pull-all.sh safe to use on machines with a partial clone set,
while keeping the explicit-list override behavior intact.

Adds three regression tests: absent-only exits 0, absent+present exits 0
with OK for the present repo, present-failed+absent exits 1.

Evidence-Ticket: OMN-13055

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13070): config_prefetcher overlay wins on controlled lanes (#1961)

On infisical_required=True lanes, fetched/overlay config now always wins
over ambient env. Ambient env is retained only as a declared bootstrap
fallback (with an explicit provenance INFO log line) when Infisical
returns None. apply_to_environment also overwrites stale env on controlled
lanes. Uncontrolled lane (infisical_required=False) behaviour is unchanged.

Adds 5 regression tests: controlled-lane Infisical-wins, env-bootstrap-
fallback, apply_to_environment overwrite, missing-from-both-is-error, and
uncontrolled-lane-env-still-wins. Refactors _resolve_key to return a
(outcome, value, error) tuple to satisfy the ≤5-param pattern gate.

Source: docs/audits/2026-06-10-runtime-env-overlay-authority-audit.md

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13062): migration-gate vacuity fix — sentinel discipline, wait-for-postgres, skip-manifest, sync-check nonzero (retro A-10) (#1963)

* fix(OMN-13062): migration-gate vacuity fix — sentinel discipline, wait-for-postgres, skip-manifest, sync-check nonzero

Fixes three bugs identified in retro A-10 (recurrences OMN-12885, OMN-12934):

(1) RUNNER SENTINEL DISCIPLINE
    run-forward-migrations.sh now clears migrations_complete=FALSE at the
    start of every run and sets it TRUE only as its FINAL act after all
    infra and node migrations succeed. Any mid-run failure leaves the gate
    UNHEALTHY. runner_completed_at is stamped at the same final step as
    durable evidence of a successful completion run.

(2) SYNC-NODE-MIGRATIONS VACUOUS GATE
    sync-node-migrations.sh --check now exits 2 (not 0) when the omnimarket
    source tree is unresolvable. Silent exit-0 was hiding drift. The single
    opt-out is SYNC_NODE_MIGRATIONS_SKIP_UNRESOLVABLE=1 for environments
    that intentionally run without the source.

(3) WAIT-FOR-POSTGRES GUARD
    run-forward-migrations.sh now waits up to PG_WAIT_RETRIES (default 30)
    x 2s for Postgres to accept connections before proceeding, guarding
    the first-boot initdb race.

(4) SKIP-MANIFEST
    docker/migrations/skip-manifest.yaml introduced as the sole committed
    escape for intentionally-skipped migrations. The runner reads this at
    startup; listed migrations are recorded in schema_migrations with
    checksum "skip-manifest" without executing the SQL.

(5) MIGRATION 085
    Adds runner_completed_at TIMESTAMPTZ column to db_metadata so the
    runner's final stamp is durable in the schema (idempotent ADD COLUMN IF
    NOT EXISTS). Rollback included.

22 regression tests added covering all five fix surfaces.

* fix(OMN-13062): stamp schema fingerprint for migration 085

Migration 085 (085_add_runner_completed_at_to_db_metadata.sql) was added
in the initial commit but schema_fingerprint.sha256 was not regenerated.
Running `python scripts/check_schema_fingerprint.py stamp` updates the
artifact from the stale hash to match the 71 migration files.

Evidence-Source: OCC#2563
Evidence-Ticket: OMN-13062

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-12864): Bifrost endpoints → committed overlay authority + fail-loud loader (#1964)

* feat(OMN-12864, OMN-12814, OMN-12945): Bifrost endpoints → committed overlay authority + fail-loud loader

OMN-12864 — Committed lane overlay
  - docker/lane-overlays/dev.bifrost.yaml: typed deployment bindings for all
    four BIFROST_LOCAL_*_ENDPOINT_URL values (coder :8000, reasoner :8001,
    embedding :8100, ds4-flash :8101). Previously only available as ephemeral
    shell exports on .201; now committed, auditable, diff-able, and CI-checked.
  - docker/lane-overlays/dev.bifrost.env: generated dotenv sidecar consumed by
    compose via env_file; never edited directly (yaml is authority).
  - docker/docker-compose.infra.yml: wire the env_file block at compose root so
    the four endpoints are injected into the interpolation context on a clean
    shell. Hardcode BIFROST_CONTRACT_PATH (remove :-/empty footgun — OMN-12814).
  - scripts/render_bifrost_lane_overlay_env.py: render script regenerates the
    env sidecar from the YAML source.
  - src/omnibase_infra/runtime/models/model_bifrost_lane_overlay.py:
    ModelBifrostLaneOverlay — typed Pydantic model enforcing URL completeness
    (OMN-12815: every URL must end in /chat/completions).

OMN-12814 — Fail-loud loader
  - render_bifrost_delegation_contract: raises ProtocolConfigurationError on
    FileNotFoundError, YAMLError, ValidationError, and zero-endpoint renders.
    No lru_cache — every restart re-renders from packaged source so a stale
    cache cannot pin a broken result across deploys.

OMN-12945 — Re-seed from packaged source on deploy
  - docker/entrypoint-runtime.sh: set BIFROST_FORCE_RESEED=1 on every container
    restart so the named-volume copy is always rebuilt from the packaged
    bifrost_delegation.yaml merged with committed lane-overlay endpoints.
  - render_bifrost_delegation_contract: honor BIFROST_FORCE_RESEED/force_reseed
    flag to bypass the stale-volume early-return path entirely.

Tests:
  - tests/ci/test_bifrost_lane_overlay.py: CI gate — env sidecar in-sync with
    YAML source; all four BIFROST_LOCAL_* keys present.
  - tests/unit/runtime/models/test_model_bifrost_lane_overlay.py: bare-base URL
    rejection, env dict mapping, extra-field rejection.
  - tests/unit/runtime/test_render_bifrost_delegation_contract.py: fail-loud
    paths, force-reseed, zero-endpoint error, endpoint URL completeness.
  - tests/unit/models/test_model_serialization_roundtrip.py: roundtrip coverage.

* fix(OMN-12864): move bifrost env_file to service level — fix compose schema validation failure

Top-level 'env_file' is rejected by Docker Compose v2 schema validator
('additional properties not allowed'). This caused 10+ compose-render
integration tests to fail in CI.

Fix:
- Remove top-level env_file block from docker-compose.infra.yml
- Add per-service env_file on omninode-runtime, runtime-effects,
  runtime-worker (the three containers that render Bifrost)
- Change BIFROST_LOCAL_*:? to BIFROST_LOCAL_*:- in x-runtime-env
  (compose-level validation removed; Python validates via
  ModelBifrostLaneOverlay + render_bifrost_delegation_contract)
- Add two CI gate tests: compose_env_file_is_service_level_not_top_level
  and runtime_services_have_bifrost_env_file

* fix(OMN-12864): allow BIFROST_LOCAL_* empty defaults in silent-fallback gate

The overlay authority pattern (OMN-12864) passes BIFROST_LOCAL_*_ENDPOINT_URL
via service-level env_file (docker/lane-overlays/dev.bifrost.env), not at
compose config time. Compose-level :? would break CI rendering without the
overlay pre-loaded. Validation at the Python layer (ModelBifrostLaneOverlay
+ render_bifrost_delegation_contract) is the enforcement point.

Add the three failing vars to ALLOWED_EMPTY_DEFAULTS with OMN-12864 citation.

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-12858/OMN-12879/OMN-12880): dispatcher route-coverage CI gate (#1965)

* feat(OMN-12858/OMN-12879/OMN-12880): dispatcher route-coverage CI gate

Static contract analysis: every subscribed command topic (onex.cmd.*)
must declare handler_routing or runtime_dispatch, or the message goes to
DLQ silently. This gate would have caught two recent incidents:

1. June 9 DLQ regression (OMN-12858 post-mortem): node_generation_consumer
   subscribed onex.cmd.omnimarket.node-generation-requested.v1 but a
   sole-handler revert left zero dispatcher routes registered. Messages
   went to DLQ silently with no CI signal.

2. June 12 DEL-01 live finding: onex.cmd.omnimarket.delegate-skill.v1 was
   consumed by dev lane bus but no dispatcher route existed in any deployed
   contract. Discovered via manual rpk consumer-group lag probe (DEL-01
   evidence, docs/evidence/2026-06-12-weekend-pass/).

Deliverables:
- scripts/check_dispatcher_route_coverage.py — static YAML scanner that
  checks both omnibase_infra and omnimarket contract trees; ratchet
  allowlist for known pre-existing violations; --changed-contracts mode
  (OMN-12879) for per-PR scoping; compat publish topics excluded (OMN-12880)
- .github/workflows/dispatcher-route-coverage.yml — CI workflow that
  checks out omnimarket sibling, collects changed contract paths in PR
  mode, and runs the gate; fires on PR, push-to-main, and merge_group
- tests/ci/test_dispatcher_route_coverage_gate.py — 12 unit tests
  covering RED/GREEN/COMPAT/CHANGED-MODE/ALLOWLIST/MULTI-DIR paths plus
  live-contract regression proof against the actual omnibase_infra tree

Allowlist additions:
- onex.cmd.omnibase-infra.pattern-b-dispatch.v1 (RuntimePatternBBroker,
  imperative consumer, OMN-12525 migration target)
- onex.cmd.platform.contract-resolve-requested.v1 (transitional HTTP
  bridge node_contract_resolver_bridge OMN-2756, metadata.transitional=true)

[OMN-12858, OMN-12879, OMN-12880]

* fix(OMN-12858): drop full uv sync from dispatcher-route-coverage workflow

Gate script only needs pyyaml (stdlib + yaml). Using full setup-python-uv
was causing 10+ minute timeout. Replace with direct pip install pyyaml and
invoke python3 directly. Reduces job from 10m timeout to <1m.

[OMN-12858]

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-13026): wire transport-mock-lint pre-commit hook + CI gate (#1962)

* feat(OMN-13026): wire transport-mock-lint pre-commit hook + CI gate (omnibase_infra)

Activates the transport-mock-lint validator (from omnibase_core, OMN-13026)
on omnibase_infra. Ratchet baseline: 218 existing violations across 63 files
frozen in validation/transport_mock_baseline.yaml. New bare AsyncMock/
MagicMock on EventBus/transport surfaces are blocked by pre-commit hook and
CI lint step. Existing violations tracked for drain by per-site tickets
(parent OMN-13026).

Reference incident: PR #1181 bare AsyncMock hid missing EventBusKafka.stop().

Evidence-Ticket: OMN-13026

* fix(OMN-13026): use uv run python for CI lint step + bump omnibase_core pin to include transport_mock_lint

The transport-mock lint CI step previously cloned omnibase_core and ran
`python -m omnibase_core.validators.transport_mock_lint` with PYTHONPATH,
but this failed: `No module named omnibase_core.validators.transport_mock_lint`
because it ran `.venv/bin/python` which uses the locked venv, and the venv
omnibase_core pin (2defabef4) predates the transport_mock_lint module.

Fix: use `uv run python` (removes the clone step) and bump omnibase-core git
pin from 2defabef4 to 309d89fa7 (PR 1231 merge commit on dev) so
transport_mock_lint is available in the locked venv.

* fix(OMN-13026): align transport mock baseline and runner lock

* fix(OMN-13026): sync omnibase_core pin + runner identity lock to dev baseline

Align pyproject.toml omnibase_core rev to 2defabef (required by
test_release_backmerge_preserves_proven_runtime_core_pin) and update
docker/runners/runner-image.lock.json identity_digest/shared_env_digest
to match dev runner image lock (79b08f44 / 90c8b3b9).

Both were stale from the prior session's pin bump that used an older SHA.

* fix(OMN-13026): source transport mock validator from core

* fix(OMN-13026): source transport validator from core dev

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-13094): --output receipt mode on onex node/run — quiet typed receipts with durable capture (#1966)

Phase 2a of the skill-output-suppression slice (epic OMN-13089, plan
docs/plans/2026-06-12-skill-output-suppression-plan.md Phase 2 item 1).

- onex node/onex run gain --output receipt: ALL runtime logging routes to a
  run_id-suffixed capture file under <state-root>/captures/ (no console
  handlers — kills the 25-50-line RuntimeLocal INFO stream at the source);
  stdout carries exactly ONE typed ModelSkillResult JSON with the FULL
  handler result (result_model = concrete handler result type FQN).
- Durable capture: capture log + handler result content-addressed via
  omnibase_core ArtifactStore (OMN-13093); artifact.captured +
  tool.output.captured emitted to the emit daemon socket (--emit-socket,
  default ~/.claude/emit.sock).
- Failure asymmetry: artifact write failure => FULL output printed, no
  receipt (no hidden loss); emission failure => receipt still prints, event
  spooled to <state-root>/emit_spool/ for replay.
- Node failure => status=failed/error with full error + capture log INLINE
  in the receipt (errors are never hidden) and artifact-backed.
- Default output mode unchanged (enforcement is Phase 4).
- RuntimeLocal exposes handler_result (receipt schema identity).
- core pin 2defabef -> ae8793bd (merged OMN-13091/13093 receipt models +
  ArtifactStore); runner-image identity lock regenerated and the OMN-12765
  backmerge identity constants updated for the new pin.

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-13097): onex skill subcommand + declarative skill->node mapping (#1968)

* feat(OMN-13097): onex skill subcommand + declarative skill->node mapping

Phase 4a of the skill-output-suppression slice (epic OMN-13089, plan
docs/plans/2026-06-12-skill-output-suppression-plan.md Phase 4 item 1/2).

A dispatch skill IS one CLI call (user directive 2026-06-12). This adds the
`onex skill <name> [args]` dispatch surface that the 24 omniclaude shim
migrations build on:

- onex.cli entry-point `skill` -> cli_skill.run_skill_by_name. Resolves the
  skill via the declarative skill_mapping.yaml registry, builds the backing
  node's input payload from the skill's CLI args, writes it under the state
  root (.onex_state/tmp/<skill>-<run_id>.json — never /tmp), resolves the
  node's packaged contract exactly like `onex node`, and dispatches through
  the proven receipt-mode path (run_receipt_mode, OMN-13094). stdout is
  exactly one typed ModelSkillResult JSON with the FULL handler result.
- skill_mapping.yaml: declarative DATA mapping all 24 dispatch shims to their
  backing onex.nodes node + typed result-model FQN (verified against each
  live handler handle() return type on origin/dev) + per-arg payload specs +
  static payload + keyword classifiers (delegate task_type as data, not code).
  Adding a skill is a YAML edit + fixture, never a CLI code change (ticket
  deliverable 2/3). Mapping lives beside the node-resolution surface, never
  hardcoded branching in the CLI.
- Typed models split one-per-file (repo convention): ModelSkillArgSpec,
  ModelSkillClassifier, ModelSkillMapping, ModelSkillMappingRegistry,
  EnumSkillArgType. Frozen, extra=forbid, fail-fast coercion/validation.
- validation_exemptions.yaml: Click-callback param-count + literal-identifier
  name-field exemptions mirroring the existing cli_node run_node_by_name
  precedent (OMN-11570) — same pattern, same rationale.

dod_evidence:
- 20 unit tests pass (registry validity, all-24-shims coverage, FQN result
  models, arg parsing/coercion/positional/required, classifiers, payload
  build, receipt-mode dispatch wiring, payload-under-state-root not /tmp).
- uv run mypy src/ --strict: clean (2438 source files).
- ruff format + check: clean. pre-commit run on changed files: pass.
- runner-image identity lock regenerated for the pyproject entry-point add
  (same as OMN-13094).

* test(OMN-13097): rebind OMN-12765 backmerge identity constants for onex skill pyproject change

Adding the `skill` onex.cli entry-point to pyproject.toml changes the
runner-image identity_digest (and shared_env_digest) the lock binds. Update
the hardcoded expected constants in the backmerge-identity test to the
regenerated values — same mechanical rebind OMN-13094 performed for the core
pin bump. Identity + runner-image-identity tests pass (14).

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix+test(OMN-13012): force terminal-topic metadata refresh so both ephemeral consumers assign (runner consume-leg wedge) (#1969)

The runner consume-leg wedged on the live stability battery (image c0505521f1fa,
EXP1-3_RUNNER_CONSUME_LEG_BLOCKER): only the FAILED terminal topic ever subscribed
while the COMPLETED topic never assigned, so the correlated completed terminal was
never read and the 8x2x10 matrix re-fired cell 1 forever, emitting zero
non-degenerate rows. A prior two-strike diagnosis proved the omnimarket handler is
correct (it opens both terminal sessions pre-publish and races them); the defect is
in the omnibase_infra runtime consume leg.

Root cause: _assign_direct_terminal_partitions ignored the metadata future returned
by AIOKafkaClient.set_topics and re-called set_topics([same_topic]) each loop
iteration. aiokafka 0.13.0 set_topics only forces a metadata refresh when the topic
set DIFFERS from the tracked set, so every iteration after the first took the no-op
branch and never re-fetched. An ephemeral group_id=None consumer whose first metadata
fetch had not yet surfaced partitions burned the full 30s assign cap and raised a
bare TimeoutError (the empty-message 'wait failed' seen live).

Fix: register the reply topic once and await that metadata fetch, then on each miss
force a fresh fetch via force_metadata_update (which always fetches) rather than the
no-op set_topics repeat. The assign-cap TimeoutError now carries a diagnostic message
instead of an empty one.

Test: tests/integration/test_terminal_consumer_concurrent_assign_race.py drives the
REAL TerminalEventConsumer (the object wired as event_consumer) through the REAL
open_direct_terminal_consumer/poll path with AIOKafkaConsumer monkeypatched to a fake
that faithfully models aiokafka 0.13.0 set_topics future + metadata-latency semantics.
RED before the fix (bare TimeoutError, the live empty-message signature); GREEN after.
K>=2 multi-trial variant asserts no worker-thread leak across trials.

Evidence-Source: <occ-sha-pending>
Evidence-Ticket: OMN-13012

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-13096): onex delegate single-command subcommand (Phase 2b) (#1967)

* feat(OMN-13096): onex delegate single-command subcommand

Add 'onex delegate "<prompt>" [--task-type X] [--max-tokens N]' as a
subcommand on the existing onex CLI (Phase 2b of the skill-output-suppression
slice, OMN-13089). The command wraps payload construction, node dispatch, and
result extraction internally and prints exactly one
ModelSkillResult[ModelDelegateSkillResponse] to stdout via the OMN-13094
receipt-mode path. RuntimeLocal logs go to the capture file + artifact store,
never to stdout; scratch payloads live under <state-root>/tmp/ with run_id
suffixes (never /tmp).

- cli_delegate.py: classify_task_type (keyword table from legacy skill md),
  payload write, contract resolve, run_receipt_mode dispatch
- register 'delegate' under onex.cli entry points
- exempt delegate_command from the >5-param patterns gate (same Click-callback
  rationale as run_node_by_name)
- 20 unit tests: classification, scratch-under-state-root, single typed
  receipt on stdout, zero INFO log leakage

omnibase_infra does NOT depend on omnimarket; the delegate node is resolved at
runtime via the onex.nodes entry-point group (registered by omnimarket).

* chore(OMN-13096): re-trigger deploy-gate after Evidence-Source set to OCC#2593

No code change — the deploy-gate workflow triggers on synchronize (not edited),
so the PR-body Evidence-Source fix needs a new commit to re-resolve the OCC ref
to the open PR head where contracts/OMN-13096.yaml (with deploy evidence) lives.

* chore(OMN-13096): re-trigger deploy-gate now that OCC#2593 merged to OCC dev

contracts/OMN-13096.yaml (with the dod-deploy-onex-delegate item) is now on
OCC dev, so the deploy-gate OCC-dev checkout resolves the contract + deploy
evidence.

* chore(OMN-13096): regenerate runner-image identity lock for pyproject entry-point add

Adding the 'delegate' onex.cli entry point to pyproject.toml changed the
dependency-manifest digest that scripts/ci/runner_image_identity.py folds into
the runner-image identity lock. Regenerate the lock so
tests/ci/test_runner_image_identity.py matches (was the only CI test failure;
unrelated environmental integration/perf failures excluded).

* test(OMN-13096): update backmerge identity assertions to regenerated lock digests

The runner-image identity lock was regenerated for the pyproject entry-point
add; this test hardcodes the expected identity_digest/shared_env_digest, so
update both to match the new lock (same maintenance OMN-13094 did).

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13118): tolerate partition-less reply topic in terminal consume leg (#1970)

The context-ROI runner opens one ephemeral group_id=None terminal consumer per
terminal topic BEFORE publishing each generation command (subscribe-before-publish,
OMN-13012/13038). The FAILED reply topic is only produced to on contract_passed=False;
in a battery where generations pass it has zero messages, so Redpanda never advertises
a partition for it. _assign_direct_terminal_partitions burned the full 30s assign cap
on every trial then raised a bare TimeoutError (the empty-message 'wait failed'),
stalling each of the 160 battery trials ~30s before the COMPLETED terminal could
correlate -> battery needs >80 min and never completes (verifier-confirmed wedge).

A partition-less reply topic is a valid steady state, not a 30s error:
- _assign_direct_terminal_partitions gives a bounded grace window for a topic that
  exists but is slow to surface metadata, then assigns whatever partitions exist
  (possibly none) and returns promptly instead of burning the cap and raising.
- poll_direct_terminal_consumer treats an empty assignment as 'no terminal will
  arrive here' (sleeps out its timeout, returns None) without calling getone() on
  an unassigned consumer.

Repro: tests/integration/test_terminal_consumer_battery_load_wedge.py drives the
REAL TerminalEventConsumer over K=10 x 2 cells x 2 arms; RED (40/40 trials block a
full assign cap) before the fix, GREEN after.

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13118): pin terminal-consumer read offset synchronously (close lazy seek_to_end publish race) (#1971)

* fix(OMN-13118): pin terminal-consumer read offset synchronously to close lazy seek_to_end publish race

The consume-leg wedge survived PR #1969 (set_topics no-op) and PR #1970
(partition-less assign-cap) because both addressed the assign phase, not the
seek timing. AIOKafkaConsumer.seek_to_end is LAZY: it requests a LATEST offset
reset that only resolves on the first poll — AFTER the caller publishes. With
generation completing in ~1s, the correlated COMPLETED terminal lands in the
open->poll gap, so the lazily-resolved LATEST position is the HWM AFTER the
record and the poll reads past it. The terminal is never read, the trial never
correlates, and the experiment matrix re-fires the same cell forever.

Replace seek_to_end with a synchronous end_offsets() + seek() pin
(_pin_direct_terminal_end_offsets) in both open_direct_terminal_consumer and
RuntimePatternBBroker._dispatch_and_wait_with_direct_kafka_consumer, so the
read position is fixed at open() time, before the publish — the real
subscribe-before-publish guarantee. Empty assignment (partition-less reply
topic, OMN-13118 #1970) is a no-op.

Repro: tests/integration/test_terminal_consumer_seek_reset_race.py drives the
real TerminalEventConsumer.open()/wait() the way HandlerContextRoiRunner does,
publishing the correlated terminal in the open->wait gap across K=10 x 2 cells
x 2 arms; RED with the lazy reset (every cell degenerate), GREEN once the read
offset is pinned. RED verified by git-stashing only the source fix.

Existing consume-leg fakes updated to model end_offsets/seek (they previously
masked the bug by making seek_to_end a synchronous exact snapshot).

* test(OMN-13118): reword assertion (lazy not deferred) for receipt honesty gate

* fix(OMN-13118): bound end_offsets() round-trip with assign-cap timeout (CodeRabbit)

end_offsets() is a broker ListOffsets round-trip aiokafka documents as able to
block indefinitely. Bound it with the same cap as start()/assign so a stalled
broker fails fast instead of hanging the pre-publish positioning.

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13137): _validate_routing strategy-aware event_model checks (#1973)

operation_match routes by the `operation` field and does not use event_model.
The validator was unconditionally requiring event_model.{name,module} for every
handler entry, causing 230/295 omnimarket operation_match contracts (all correct
as authored) to fail routing validation at startup.

Fix: read routing_strategy from the routing map and branch validation:
  - payload_type_match → require event_model.{name, module} (unchanged)
  - operation_match (and any non-payload strategy) → require `operation`;
    skip event_model checks entirely

Updated pre-existing _validate_routing tests to declare routing_strategy:
payload_type_match explicitly (they always tested payload_type_match semantics
but relied on the implicit fallback that is now removed).

Added test_validate_routing_operation_match.py with 4 unit tests:
  1. operation_match without event_model → zero event_model errors
  2. operation_match missing operation field → error
  3. payload_type_match missing event_model → still errors (regression guard)
  4. Real node_integration_sweep_orchestrator routing block → clean (boot gate)

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13118): independent per-terminal-topic consumers in Pattern B direct-Kafka wait (#1972)

The consume-leg wedge survived four merged fixes (#1969 set_topics no-op,
#1970 partition-less assign-cap, #1971 synchronous seek-pin). The STRONG K>=10
multi-cell reprobe on the stability lane still wedged on REBUILD-5
(cdf53d963f7b). Converged diagnosis (strikes 3+4,
docs/evidence/2026-06-12-weekend-pass/experiments/probe4-stability/
reprobe-K10-rebuild5/HALT_K10_WEDGE_PERSISTS.md): the runtime waited for each
trial's terminal across TWO topics (node-generation-completed.v1 +
node-generation-failed.v1) with a SINGLE ephemeral group_id=None consumer
assigned both topics' partitions. One aiokafka consumer holds one manual
subscription; the COMPLETED delivery window collapsed before it surfaced the
correlated record, so the trial never correlated and the matrix re-fired cell 1.

RuntimePatternBBroker._dispatch_and_wait_with_direct_kafka_consumer now opens
ONE independent AIOKafkaConsumer PER terminal topic via
open_direct_terminal_consumer (each started, assigned, and offset-pinned via
end_offsets()+seek() at open() BEFORE publish), then awaits both CONCURRENTLY
via asyncio.wait(FIRST_COMPLETED). The first correlated terminal wins; both are
torn down. No shared consumer, no subscription flip.

Keeps the #1970 partition-less no-op and the #1971 synchronous offset pin (both
live in open_direct_terminal_consumer / poll_direct_terminal_consumer). Removes
the now-dead single-consumer helpers (_assign_terminal_topic_partitions,
method-level _refresh_terminal_topic_metadata, _direct_kafka_* kwargs builders,
_kafka_bootstrap_servers/_kafka_event_bus).

Adds tests/integration/test_terminal_consumer_subscription_flip_wedge.py: a
real-dispatch-path K>=10 x 2-cell x 2-arm repro whose fake models TWO
independent consumers honestly (delivery is faithful only for a single-topic
assignment; a consumer spanning both topics flips and drops the COMPLETED
record). RED genuineness verified by reverting only the source to the
single-consumer shape (test hangs past timeout); GREEN with the fix in 2.3s.

Acceptance is the LIVE K>=10 multi-cell stability-lane reprobe (later phase),
NOT this unit test. A green unit repro is necessary but NOT sufficient.

Refs OMN-13118, OMN-13128.

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13118): long-lived terminal correlator replaces per-trial ephemeral consume leg (#1974)

Tier B canonical redesign (epic OMN-12525). Five offset/subscription patches
(#1969-#1972) tuned the per-trial-ephemeral terminal consumer and all fa…
…push unblocked) (#2166)

* docs(OMN-12962): contract-store durability audit — cold-runtime census proof (#1937)

Proves cold-start contract census reconstructs from the image-bundled
filesystem manifest (HYBRID-mode bootstrap + PluginLoaderContractSource),
independent of node-registration.v1 retention. The delete-retention topic
feeds only the post-freeze dynamic listener (auto_offset_reset=latest, no
history replay). Live .201 stability-test evidence: filesystem contract_path
in manifest + truncated topic log head with intact census. No store fix
needed; residual dynamic-only gap covered by runtime_sweep sweep check.

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12970): vendor omnimarket projection node migrations into forward/nodes (#1938)

Vendors omnimarket node-owned projection migrations into the namespaced
forward-migration tree so run-forward-migrations.sh materializes them in the
dashboard projection DB (omnidash_analytics) at deploy.

Primary (OMN-12970): creates llm_call_metrics, llm_cost_aggregates, and
capability_scores in the projection DB. These were only ever created in the
omnibase_infra DB by infra migrations 031/060, so the ab-compare,
cost.token_usage, cost.summary, and capability-scores projection topics were
DEGRADED at startup ('table not found') and their dashboard panels rendered
empty.

Also re-syncs three omnimarket node migrations the vendor tree had drifted from
(node_projection_llm_routing, node_projection_overnight, node_projection_savings
/077) — sync-node-migrations.sh --check requires the full vendored tree to match
omnimarket source, and these were missing.

Companion to omnimarket PR for the same ticket (source migrations + projection
table-coverage ratchet test).

Evidence-Ticket: OMN-12970

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12965): stamp runtime image identity (version + revision) in workspace builds (#1943)

* fix(OMN-12965): stamp runtime image identity (version + revision) in workspace builds

The main runtime image stamped org.opencontainers.image.version=0.1.0 with a
blank org.opencontainers.image.revision after workspace rebuilds. A blank
identity degrades every proof packet (runtime SHA + image digest are required
citations in accepted evidence).

Root causes (three build paths under-stamped identity):
- onex up --build (cmd_up) passed only GIT_SHA; the runtime-stage OCI labels
  read VCS_REF (-> blank revision) and RUNTIME_VERSION (-> placeholder 0.1.0).
- deploy-runtime.sh passed VCS_REF but not RUNTIME_VERSION/GIT_SHA.
- Dockerfile silently allowed blank/placeholder identity in workspace mode.

Fix:
- cli._image_identity_build_args() stamps the full quad (GIT_SHA/VCS_REF/
  RUNTIME_VERSION/BUILD_DATE) and fails fast on an unresolved git revision.
- deploy-runtime.sh stamps RUNTIME_VERSION + GIT_SHA and verifies the version
  label is non-placeholder post-deploy.
- Dockerfile.runtime fails workspace builds with blank VCS_REF or placeholder
  RUNTIME_VERSION=0.1.0 (release mode unaffected).

Enforcement ratchet (same PR):
- scripts/check_runtime_image_identity.py static check, wired as pre-commit hook
  + CI gate (ci.yml).
- tests/unit/infra/test_runtime_image_identity_labels.py pins the cli helpers +
  Dockerfile guard; deploy-agent test extended for the quad.

Proven locally via throwaway docker builds: workspace+args -> populated labels;
workspace without args -> guard fails (exit 64); release without args -> 0.1.0
placeholder allowed (no regression).

Evidence-Ticket: OMN-12965

* test(OMN-12965): integration build proof for runtime image identity labels

Builds the real runtime-stage ARG/LABEL/guard block against busybox and asserts
via docker inspect: workspace+args -> populated version/revision; workspace
without args -> guard fails (exit 64); release -> placeholder allowed. Satisfies
the integration-test hard gate and makes the throwaway proof permanent.

Evidence-Ticket: OMN-12965

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12987): workspace-build sibling lock-pin preflight + provenance (#1944)

Recurrence ratchet for the 2026-06-11 stability bootstrap crash. The 11:20Z
--no-cache rebuild vendored omnibase_infra 0.37.0-dev (~2c1d672f) + core 0.42.0
even though omnimarket dev's uv.lock pinned infra 0.38.1 @ e2dbdc95 + core
0.44.0 @ c97c2c9a. The stale sibling predated the OMN-12501 Protocol-quarantine
guard, turning a latent contract defect into a fatal crash that crash-looped the
main runtime.

- check_sibling_lock_pins.py: host-side fail-fast preflight resolving expected
  sibling versions/SHAs from the consuming repo's (omnimarket) uv.lock and
  comparing them against each vendored tree. Mismatch aborts the build.
- stage_workspace.sh: emit a .build-sha marker per staged sibling (rsync drops
  .git) so the preflight and provenance can identify the vendored commit.
- deploy-runtime.sh: run the preflight after staging, before build; abort on
  mismatch. Write the comparison under sibling-repos/.sibling-lock-pins.json.
- compute_workspace_provenance.py + Dockerfile.runtime: fold expected-vs-actual
  lock_pin_comparison into build-provenance.json for deploy verifiers.

Recurrence-guard tests prove a stale infra 0.37.0 vs lock-pinned 0.38.1 fails
the preflight and matched pins pass; deploy script wiring is asserted statically.

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12988): pin stability runtime-worker to replicas 1 + census ratchet (#1942)

The base docker-compose.infra.yml defaults runtime-worker to
replicas 0 (${WORKER_REPLICAS:-0}). The stability-test lane's required
state includes a running worker (4-container census: main, effects,
worker, projection-api), but the override pinned it via an
env-interpolation default (${STABILITY_TEST_WORKER_REPLICAS:-1}) — a
silent-drop surface: a stray exported STABILITY_TEST_WORKER_REPLICAS=0
or removal of the :-1 fallback would scale the worker to 0 with zero
signal on a plain compose up/recreate.

Fix: pin docker-compose.stability-test.yml runtime-worker
deploy.replicas to the literal 1 (no env indirection).

Ratchet (recurrence guards, same PR):
- scripts/deploy-agent/tests/unit/test_runtime_worker_census.py: assert
  runtime-worker stays in the deploy-agent RUNTIME-scope census so a
  missing worker (replicas 0 => absent from docker compose ps) is a
  deploy failure, not silence; assert the override pins a literal 1.
- tests/integration/infra/test_stability_test_runtime_compose_render.py:
  assert the rendered stability worker resolves deploy.replicas == 1.
- tests/unit/infra/test_stability_test_runtime_lane.py: update the
  existing pin assertion to the literal 1.

Evidence-Ticket: OMN-12988
Config-drift family: OMN-12945

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12979): expire-bound topic completeness suppressions (#1940)

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12966): accept community-edition Infisical /api/status in provision readiness gate (#1939)

The fresh-provision readiness gate in provision-infisical.py only accepted the
enterprise {"status": "ok"} payload and rejected the community edition's
{"message": "Ok"}, returning 1 before bootstrap could run. This blocked
provisioning against the Infisical instance deployed on .201 (community edition).

Route the gate through the existing _is_infisical_ready helper (single source of
truth, already used by the already-provisioned path). Add TestMainFreshProvision-
ReadinessGate covering community/enterprise/not-ready cases.

Also adds docker/docker-compose.infisical-stability.yml: an ADDITIVE Infisical
compose project for the stability-test lane (joins the existing network as
external, reuses stability postgres/valkey, no lane mutation), since the lane
overlays disable the in-lane Infisical service via *-disabled profile overrides.

P1.2b-A: Infisical now reachable from the stability runtime/effect containers; a
known secret (OMN_12966_PROBE) seeds and resolves end-to-end via the machine-
identity universal-auth path, verified from inside the effects container.

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-12958): volume-config drift gate + runtime config provenance (#1941)

* feat(OMN-12958): volume-config drift gate + runtime provenance

Compute config provenance (path + sha256) for the runtime-rendered Bifrost
delegation contract; the deployed volume copy survives rebuilds and silently
diverges from packaged source (two competing authorities, OMN-12945).

- runtime/config_provenance.py: ModelConfigProvenance + drift classification,
  sidecar JSON writer (read by sweep + proof packets)
- runtime/health/health_config_provenance.py: drift -> degraded health
- render entrypoint logs provenance line + writes sidecar on every boot
- docs/runbooks/volume-config-drift-and-reseed.md: ledgered re-seed procedure
- validation exemption for config_name (logical identifier, not entity ref)

No live volume mutation: re-seed is an operator deploy step (deploy_pending).

* test(OMN-12958): cover volume config drift reseed flow

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-12957): wire runtime-profile validator + core registry parity guard (#1945)

* feat(OMN-12957): wire runtime-profile validator + core registry parity guard

- Remove the import-time RuntimeError drift raise in runtime_profile.py (a hard
  raise on core/infra version skew would crash the kernel at import); the parity
  invariant is enforced by test_profiles_match_core_registry instead.
- Add tests: _PROFILES keys == omnibase_core REGISTERED_RUNTIME_PROFILES, and
  every CONSUMER_ATTACHED_RUNTIME_PROFILES profile loads as a real lane.
- Wire omnibase_core.validation.validator_runtime_profiles as a pre-commit hook
  + validator-runtime-profiles.yml CI gate on infra contracts.
- Freeze 19 pre-existing violators in validation/runtime_profiles_allowlist.yaml
  (discovered by repo-root walk; drain via OMN-12982). Blocks NEW orphans.

Requires the omnibase_core pin to include OMN-12957's validator (new rules).

Evidence-Ticket: OMN-12957
Evidence-Source: 5463fbaf819409d4fb7f491dd4f276f10d869eba

* ci(OMN-12957): pass runtime profile allowlist to validator

* test(OMN-12957): cover runtime profile registry parity

* fix(OMN-12957): keep runtime profile allowlist under config

* fix(OMN-12957): pin core runtime profile registry

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12973): widen migration-gate healthcheck start_period + ratchet (#1950)

P2.8: classify the prod migration-gate UNHEALTHY-then-self-resolved incident.

Classification: idle-one-shot-mis-modeled = NO. The migration-gate is a
correctly-modeled long-running sentinel (entrypoint `while true; sleep 3600`)
whose healthcheck continuously polls db_metadata.migrations_complete via
check_migrations_complete.sh. The container flipped UNHEALTHY transiently because
its healthcheck start_period (10s) was far shorter than the real cold-volume
migration window (~116s: prod gate started 09:35:22, intelligence-migration
finished 09:37:18). Past the 10s grace window the still-failing probe was
reported UNHEALTHY until migrations completed, then self-resolved — no fault.

Fix: raise migration-gate healthcheck start_period 10s -> 180s in both the
authoritative catalog manifest (docker/catalog/services/migration-gate.yaml,
flows into the generated compose) and the hand-maintained
docker/docker-compose.infra.yml that deploy-runtime.sh applies to .201, so a
still-applying gate stays in `health: starting` instead of flipping UNHEALTHY.

Ratchet (enforcement, not detection): new ValidatorHealthcheckStartPeriod in
omnibase_infra catalog (validator_healthcheck_start_period.py) asserts a 120s
floor for migration-completion gates, wired into `onex validate runtime`
(cmd_validate_runtime) AND backed by unit tests that gate every PR via
pre-commit + CI. A migration-completion gate is identified by BOTH a healthcheck
that polls migration completion AND a service_completed_successfully dependency,
so ordinary app services (e.g. intelligence-api, an HTTP liveness probe with a
legitimately short 40s start_period) are not swept into the floor.

Prod probed read-only only; no prod mutation. Applies to live prod via the
batched stability/prod rebuild (deploy_pending).

Evidence-Ticket: OMN-12973
Evidence-Source: OCC#PENDING

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-12971): deliver Vertex ADC bearer token to runtime effects container (secret-ref discipline) (#1949)

* feat(OMN-12971): deliver Vertex ADC bearer token to runtime effects container (secret-ref discipline)

Wire the Vertex $500-credit ADC path into the runtime, ADDITIVE next to the
Gemini API-key path (provider-agnostic; neither provider removed or forced).

runtime_policy.contract.yaml: add a secret-source mapping llm.vertex.access_token
(source_type=file, /run/secrets/vertex_access_token) to the dev/stability-test/
judge profiles, alongside the existing llm.gemini.api_key env mapping. The ref
name MUST match cloud-vertex-gemini.secret_ref in omnimarket
bifrost_delegation.yaml. The resolved VALUE is a short-lived OAuth bearer token
minted from ADC, refreshed by the operator; the token VALUE is never committed —
only the ref name + in-container path. Add aiplatform.googleapis.com to the
cloud host allowlist.

docker-compose.infra.yml: bind the operator-supplied host token file read-only to
/run/secrets/vertex_access_token on the main and effects runtimes
(VERTEX_ACCESS_TOKEN_HOST_FILE, default /dev/null so lanes without Vertex still
start; Gemini key path unaffected). Pass through BIFROST_VERTEX_GEMINI_ENDPOINT_URL
(overlay supplies the complete Vertex OpenAI-compat URL) and
GOOGLE_CLOUD_PROJECT/LOCATION (default empty).

runtime-policy.env: regenerated from the contract via render_runtime_policy_env
(test_runtime_policy_env_matches_contract_renderer proves contract<->env parity).

test_runtime_policy_contract.py: update host-allowlist assertion for the additive
Vertex host.

Pre-existing platform-wide topic-parity-gate failures (25 unrelated topics) are
identical with and without this change (proven by diff); that gate is pre-commit-
only (not a CI merge gate) and this change adds zero new topic gaps, so that one
hook is SKIP-ped. No deploy/receipt/merge gate is bypassed.

* fix(OMN-12971): make Vertex runtime env contract-owned

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-13008): automated .201 disk maintenance — worktree GC + docker GC + watermark alert ratchet (#1952)

* feat(OMN-13008): automated .201 disk maintenance — worktree GC + docker GC + watermark alert ratchet

Conservative, keep-list-driven disk maintenance to prevent the 2026-06-11
/data ~95% outage that killed all three lanes mid-demo:

- scripts/worktree-gc.sh: drives the canonical omniclaude prune-worktrees.sh
  (merged+clean+pushed safety) on both Mac (merge-sweep tick) and .201
- scripts/disk-gc.sh + disk_gc_plan.py: conservative docker/builder/image GC.
  Pure, testable removal planner honoring a VERSIONED keep-list
  (deploy/disk-gc/keep-list.yaml): never reaps a kept repo, kept tag, in-use
  image, or anything younger than min_age_days; keeps N superseded generations.
- scripts/disk-watermark-check.sh + disk_watermark_event.py: df watermark
  ratchet. >=85% emits a typed disk-watermark bus event (warning) that the
  sweep auto-ticket path turns into a Linear ticket; >=90% emits critical.
  Broker addr is fail-fast from KAFKA_BOOTSTRAP_SERVERS (no localhost default).
- deploy/disk-gc/: systemd USER timer (onex-disk-gc.timer/.service, hourly) +
  install-disk-gc.sh. User units, NOT lane containers.
- tests: 20 unit tests incl. GC plan-safety invariants + dry-run proof that
  default mode issues no destructive op (a wrong-delete GC is worse than none).

Contract: contracts/OMN-13008.yaml

* fix(OMN-13008): pass docker inventory to GC planner via stdin, not env (ARG_MAX)

On a host with many docker images, passing the full image/ps inventory as env
vars to disk_gc_plan.py exceeds ARG_MAX ('Argument list too long', exit 126),
producing an empty plan. Write inventory to per-run scratch files (under the log
dir, never /tmp; cleaned on exit) and hand it to the planner on stdin as a JSON
envelope. Verified the failure live on .201; planner now reads stdin.

* fix(OMN-13008): simplify GC plan stdin pipe (two processes, no nested subprocess)

* fix(OMN-13008): keep-wins reconciliation — never remove an image id with any keep reason

A single image id can surface in multiple 'docker image ls' rows (one per
repo:tag). One tag could route the id to dangling-removal while another routes
it to keep (e.g. tagged 'latest' or within-N-generations). Live .201 plan showed
an id in BOTH remove_image_ids and kept_reasons. Reconcile at the end: keep wins
— any id with a keep reason is dropped from the remove list; remove list deduped.
Adds 2 regression tests. Verified live on .201.

* fix(OMN-13008): timer uses OnCalendar=hourly for reliable re-arm (oneshot service)

OnUnitActiveSec does not reliably re-elapse for a oneshot service once it goes
inactive (observed NextElapseUSecMonotonic=infinity live on .201 after the first
run). Switch to OnCalendar=hourly + RandomizedDelaySec so the timer re-arms every
hour. Keep Persistent=true for missed-run catch-up.

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13005): materialize blocking event_consumer in runtime auto-wiring (runner consume-leg degenerate rows) (#1951)

* fix(OMN-13005): materialize blocking event_consumer in auto-wiring (runner consume-leg degenerate rows)

The runtime auto-wiring materialized event_publisher for handlers that
declare it but had no equivalent for event_consumer. Request/response
EFFECT handlers (HandlerContextRoiRunner) that publish a command then
block on the correlated terminal event fell back to their no-op consumer
default, returning None immediately -> every result row degenerate
(failure_stage=generation, attempt_count=0) while generations succeeded
~1s later.

Adds _make_sync_event_consumer (mirror of _make_sync_event_publisher),
backed by service_terminal_event_consumer.make_terminal_event_consumer:
a sync (topic, correlation_id, timeout) -> dict | None adapter that runs
the proven direct-Kafka correlate-and-wait loop (from RuntimePatternBBroker)
on an isolated event loop in a worker thread, so blocking does not deadlock
the runtime dispatch loop that delivers the awaited terminal.

TDD through the REAL dispatch path: test_event_consumer_injection drives a
trial through _prepare_handler_wiring with a terminal arriving after a delay
and asserts a non-degenerate row; verified RED with injection disabled.

* fix(OMN-13012): two-phase (seek-now/wait-later) terminal event_consumer to close the subscribe-after-publish race (#1954)

The OMN-13005 injected event_consumer is a single callable that does
assign -> seek_to_end -> poll internally, all AFTER the handler has
already published its command. Once OMN-13010 freed the dispatch loop and
generation began completing in ~1s, the correlated terminal lands BEFORE
the single-call consumer's post-publish seek_to_end positions, so
seek_to_end skips PAST the already-emitted terminal and the runner times
out on an offset beyond it (probe3, run_id=20260611T2140Z-probe3 -- both
arms degenerate, zero rebalances).

Splits positioning from waiting so the caller subscribes BEFORE it
publishes:

  session = consumer.open(topic)        # assign + seek_to_end NOW
  publisher(command_topic, payload)     # publish AFTER positioning
  payload = session.wait(cid, timeout)  # block from the captured position

The returned TerminalEventConsumer is still directly callable with the
legacy (topic, cid, timeout) -> dict | None single-call shape for any
consumer that does not need subscribe-before-publish; the runner is the
only consumer today. TerminalConsumerSession owns a dedicated event loop
on a daemon worker thread for the whole open->wait->close lifecycle,
preserving the OMN-13005 loop-isolation discipline so blocking never
deadlocks the runtime dispatch loop.

TDD (real dispatch path, RED-then-GREEN): extends the OMN-13005 injection
test with a terminal emitted IMMEDIATELY after publish. The single-call
(seek-after-publish) consumer MISSES it (degenerate row -- RED test
asserts failure_stage=generation); the two-phase (open-before-publish)
consumer CATCHES it (non-degenerate -- GREEN). The Kafka layer is faked at
the two service seams against a shared in-memory log modeling seek-to-end
semantics. OMN-13005 blocking-correlate behavior preserved. 269/269
auto_wiring unit tests pass; mypy --strict clean.

Sibling to OMN-13010 / OMN-13005 / OMN-13003.

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13005): route terminal consumer through Kafka boundary

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12990): pin worker replicas in ledgered config + fail-fast lane overrides (silent-drop ratchet) (#1948)

* fix(OMN-12990): pin worker replicas in ledgered config + fail-fast lane overrides

The base compose sets runtime-worker deploy replicas to ${WORKER_REPLICAS:-0}
(soft-default ZERO). The stability lane's required state includes a running
worker (GATE_ZERO_PROOF.md: 4 runtime containers), but the worker presence rode
entirely on a compose soft-default — any plain compose up/recreate without the
policy env silently scaled the worker to zero with no error and no signal.

Fix (contract-native + fail-fast):
- Add 'replicas' to ModelRuntimeProcessPolicy; pin replicas: 1 in every lane's
  worker block in runtime_policy.contract.yaml.
- Renderer emits {PROFILE}_WORKER_REPLICAS into the ledgered runtime-policy.env
  for dev/stability-test/judge/prod.
- stability + prod compose overrides reference ${..._WORKER_REPLICAS:?...}
  (fail-fast, NO silent :-1/:-0 default). A recreate that omits the policy env
  now aborts loudly instead of dropping the worker. prod previously had no
  override at all and inherited the dangerous :-0 default.

Ratchet (recurrence guards):
- tests asserting fail-fast override form (no soft default), contract-declared
  replica pin >= 1 per lane, and rendered {PROFILE}_WORKER_REPLICAS=1 in the
  ledgered env.
- runbook deploy/verify procedure adds an expected-container census (worker
  must be present) via verify_container_manifest; a missing worker is a FAILURE,
  not silence.

Note: SKIP=topic-parity-gate — that local-only advisory gate (absent from all
.github/workflows, not a required CI check) fails on 25 pre-existing cross-repo
topic gaps (build-loop/omniclaude/omniweb) identical on pristine base HEAD
8d7da1249; this change adds zero topics. All other hooks ran clean.

Evidence-Ticket: OMN-12990

* test(OMN-12990): cover worker replica policy integration

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-12909): add gateway bus forwarder P0A (#1946)

* feat(OMN-12909): add gateway bus forwarder p0a

* test(OMN-12909): add gateway forwarder integration coverage

* fix(OMN-12909): satisfy gateway forwarder validators

* test(OMN-12909): allow gateway forwarder bus protocol

* fix(OMN-12909): sync gateway forwarder entry point

* fix(OMN-12909): refresh runner image identity lock

* test(OMN-12909): relax JSON normalizer mixed benchmark threshold

* fix(OMN-12909): allow gateway handlers to boot unconfigured

* fix(OMN-12909): declare gateway forwarder runtime profile

* fix(OMN-12909): update runner identity backmerge expectation

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-13011): LANE CENSUS RECONCILIATION ratchet — declared desired-state per lane, drift = auto-ticket (#1955)

The class fix for the recurring lane-drift regression. Nothing reconciled the
declared desired state of a runtime lane against what is actually running, so the
same failure kept recurring with zero signal: volume config drift (OMN-12945),
WORKER_REPLICAS silent zero (OMN-12988/12990), and on 2026-06-11 prod runtime
containers plus the broker network were silently absent for hours during demo prep.

Ships a per-lane DESIRED-STATE census:
- (a) DECLARED in a versioned lane manifest (deploy/lane-census/lane-manifest.yaml):
  container set, network, replicas, image-tag pattern per lane
  (stability-test/prod/judge/dev), derived from the canonical compose lane files.
  A parity ratchet keeps the manifest locked in step with the compose files.
- (b) RECONCILED on a schedule on .201 by SHARING the OMN-13008 systemd timer
  (a drop-in 4th ExecStart on onex-disk-gc.service — never a second timer) and
  on-demand via scripts/lane-census-check.sh / runtime_sweep.
- (c) Drift = typed bus event (onex.evt.infra.lane-census-drift.v1) + Linear
  auto-ticket naming exactly what is missing/extra (container_absent,
  network_detached, replicas_zero, unexpected_container, oneshot_failed/stuck,
  image_tag_mismatch). Fail-fast, no warn-only mode (gates-block policy); exit 30
  on drift; bus publish fail-fast on KAFKA_BOOTSTRAP_SERVERS (no localhost default).

Red fixture reproduces 2026-06-11: prod runtime containers absent + broker network
detached must produce the exact drift findings + a non-zero exit hours before a
human noticed. Pure planner is fully unit-tested; shell driver dry-run-tested.

Builds on the OMN-12988 deploy-agent RUNTIME census (deploy-time) as the
complementary steady-state reconciler; closes the runtime-worker.yaml
container_name: null census gap by sourcing names from the compose lane files.

Evidence-Ticket: OMN-13011
Config-drift family: OMN-12945
Relates-to: OMN-13009, OMN-12988, OMN-13008

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13020): vendor missing node migrations — llm_routing 0000 + context_roi 001 (#1956)

Vendors two omnimarket node-source migrations into the infra forward-migration
tree via scripts/sync-node-migrations.sh (the canonical OMN-12559 mechanism):

- node_projection_llm_routing/0000_create_llm_routing_decisions.sql
  (source: omnimarket #1168 / OMN-12942, merge ed6734f8)
- node_projection_context_roi/001_create_context_roi_scores.sql
  (source: omnimarket #1178 / OMN-12955, merge 5010b1f4)

Without the 0000 base table, node_projection_llm_routing/0001 (CREATE VIEW)
hard-fails against NODE_POSTGRES_DB=omnidash_analytics — exactly the prod
forward-migration exit-3 of 2026-06-11T09:35:52Z, and reproduced by
construction in any clean clones@dev build. Files are byte-identical to the
omnimarket dev blobs (sha256 f8a8b339… / c4126e65…) and to the untracked
hot-patch copies on the .201 stability clone.

Both migrations are self-contained, all-statements-IF-NOT-EXISTS, and 0000
sorts lexically before 0001 within the node's namespaced identity space.

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13058): close TerminalConsumerSession on open() failure (worker thread + event loop leak) (#1957)

TerminalConsumerSession.__init__ starts its dedicated worker loop thread
immediately. TerminalEventConsumer.open did
'session = TerminalConsumerSession(...); return session.open()' with no
cleanup: any failure inside session.open() (consumer start timeout,
partition-assign timeout, broker auth error) propagated out of the raising
expression, the session reference was lost, and the daemon worker thread plus
its never-closed asyncio event loop leaked -- one pair per failed open. The
motivating caller (HandlerContextRoiRunner) opens a session per trial, so a
160-560-trial battery against a degraded broker accumulates hundreds of
leaked threads in the long-lived effects container.

Fix: wrap session.open() in try/except BaseException -> session.close()
(idempotent: stops the loop, joins the thread) -> re-raise. Covers both the
two-phase .open(topic) path and the legacy single-call __call__ path.

Found by the P3.3 doctrinal review of merged #1951 (b9712af9 / 36d98275).

TDD: tests/unit/runtime/test_service_terminal_event_consumer_open_failure.py
injects a real open failure through the production path (event bus without
_bootstrap_servers) and asserts no alive terminal-consumer-* thread after the
raise. Verified RED with the fix stashed (2 failed), GREEN with it (2 passed).

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-13021): non-dev-base guard — fail feature-base PRs absent Stacked-Parent declaration (retro A-6) (#1958)

Any PR whose base is neither dev nor main fails unless the body carries
'Stacked-Parent: #N'. Prevents the feedback_stacked_prs_orphan_from_dev class
(#1185/#1954 auto-merged INTO parent feature branches and stranded off dev).
base=main remains governed by main-target-guard.

Epic OMN-13013 (process enforcement ratchets — June 12 retro).

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-13014): hot-patch ledger rebuild preflight gate (retro B-1) (#1959)

* feat(OMN-13014): hot-patch ledger rebuild preflight gate (retro B-1)

Hot-patches on .201 (.prepatch sibling discipline) silently revert on any
image rebuild/force-recreate — the 2026-06-11 20:58Z rebuild already erased
a live /api/generate patch once. This adds the rebuild-path gate:

- scripts/preflight_hotpatch_ledger.py: given a target container or lane +
  per-repo build refs, hard-fails when any hot-patch ledger row's source PR
  merge commit is not an ancestor of the build ref (git merge-base
  --is-ancestor), plus a .prepatch tripwire of the running container
  (unledgered .prepatch = hard fail; --post-rebuild = zero .prepatch
  expected). Sole bypass: HOTPATCH_PREFLIGHT_BYPASS carrying the Rule-10
  '# skip-token-allowed: <user-approval-receipt-id>' form.
- scripts/deploy-runtime.sh: guard_hotpatch_ledger wired into main() before
  build/preview (both dry-run and execute), lane derived from the compose
  project; skips loudly only when no ledger exists on the host.
- tests/unit/scripts/test_preflight_hotpatch_ledger.py: 17 unit tests
  (ancestor gate, lane scoping, ledger loading, tripwire, bypass forms).
- tests/ci/test_receipt_gate_install_guard.py: repair stale guard — core
  OMN-12565 replaced the OMN-9198 'uv pip uninstall first' install step with
  a cleared workspace venv (uv venv --clear); assert the new contract.

Ledger backfilled from live census (5 .prepatch files / 4 source PRs, all
MERGED to dev) at /data/omninode/hotpatch-ledger/ledger.yaml on .201.

* fix(OMN-13014): scope missing-prepatch tripwire warning to the probed container

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12977): workspace-build sibling-pin ratchet — honor consuming lock, fail-fast on stale vendor (#1936)

* fix(OMN-12977): workspace-build sibling-pin ratchet — honor consuming lock, fail-fast on stale vendor

The workspace-mode image build (BUILD_SOURCE=workspace, used by the stability-test
deploy procedure) vendored sibling/foundation packages from whatever the canonical
OMNI_HOME clones happened to be checked out at, ignoring the consuming repo's
uv.lock. On 2026-06-11 this shipped a 13-day-stale omnibase_infra 0.37.0-dev
(pre-OMN-12501 Protocol-quarantine guard) + core 0.42.0 against an omnimarket dev
lock pinning infra 0.38.1@e2dbdc95 / core 0.44.0@c97c2c9a, dropping the guard and
crashing wire_from_manifest bootstrap fatally (stability lane down on demo day).

Fix + recurrence ratchet (same PR):
- scripts/runtime_build/check_sibling_lock_pins.py: parse the consuming repo's
  uv.lock for expected version+git-rev of each foundation/sibling package
  (scoped to the package's own source line so editable/registry pins are not
  cross-attributed a dependency's rev), resolve the actual clone version+HEAD,
  compare, and classify drift backward/forward/none. Fail-fast (exit 1) on any
  drift; --allow-drift records an explicit operator override in the artifact,
  never silent.
- stage_workspace.sh: runs the preflight against the canonical clones before
  staging; aborts the build (exit 3) on unacknowledged drift and writes
  workspace/sibling-pin-comparison.json.
- compute_workspace_provenance.py: folds the expected-vs-actual comparison into
  build-provenance.json so deploy verifiers can assert the build honored the lock;
  flags unacknowledged drift as a provenance error.
- Dockerfile.runtime: COPY the comparison artifact (committed placeholder so the
  COPY always resolves; overwritten by stage_workspace.sh in workspace mode).
- TDD: 19 unit tests covering lock parsing (git/registry/editable sources),
  drift classification, the exact 0.37.0-vs-0.38.1 stale case, check_pins exit
  codes, and the allow-drift override. Pre-existing mypy-strict bare-dict errors
  in compute_workspace_provenance.py fixed in the same pass.

Evidence-Ticket: OMN-12977
Evidence-Source: pending-occ

* ci(OMN-12977): retry runtime smoke compose port race

* test(OMN-12977): align sibling-pin script tests with current API

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-12989): workspace-mode image build must honor sibling lock pins (#1947)

* fix(OMN-12989): workspace build must honor sibling lock pins (fail-fast ratchet)

The 2026-06-11 stability bootstrap crash was caused by a workspace-mode
--no-cache rebuild that vendored omnibase_infra 0.37.0-dev (13-day-stale
worktree) even though omnimarket dev's uv.lock pins infra 0.38.1. The
downgraded sibling predated the OMN-12501 Protocol-quarantine guard and
turned a latent contract defect into a fatal crash.

Fix + ratchet (same PR set):
- scripts/runtime_build/resolve_workspace_pins.py: parse the consuming repo's
  uv.lock for sibling pins (version + git rev); classify each staged/installed
  sibling as exact/ahead/regression/unpinned; fail-fast (WorkspacePinError) on
  any regression BELOW the lock pin. Stdlib version-tuple fallback when
  packaging is absent so the ratchet never fails-open.
- compute_workspace_provenance.py: enforce sibling pins + a host-infra
  self-check (installed omnibase_infra vs lock pin — the exact crash vector,
  since host infra is built from the context, not staged), and emit a
  pin_comparison block into build-provenance.json for deploy verifiers.
- Dockerfile.runtime: COPY resolve_workspace_pins.py beside the provenance
  script so the in-image import resolves.
- TDD: failing tests first (test_resolve_workspace_pins.py, 11 cases) +
  provenance integration tests; fixed a pre-existing stale RUNTIME_VERSION
  assertion to read pyproject dynamically.

Evidence-Ticket: OMN-12989

* fix(OMN-12989): workspace build must honor sibling lock pins (fail-fast ratchet)

The 2026-06-11 stability bootstrap crash was caused by a workspace-mode
--no-cache rebuild that vendored omnibase_infra 0.37.0-dev (13-day-stale
worktree) even though omnimarket dev's uv.lock pins infra 0.38.1. The
downgraded sibling predated the OMN-12501 Protocol-quarantine guard and
turned a latent contract defect into a fatal crash.

Fix + ratchet (same PR set):
- scripts/runtime_build/resolve_workspace_pins.py: parse the consuming repo's
  uv.lock for sibling pins (version + git rev); classify each staged/installed
  sibling as exact/ahead/regression/unpinned; fail-fast (WorkspacePinError) on
  any regression BELOW the lock pin. Stdlib version-tuple fallback when
  packaging is absent so the ratchet never fails-open.
- compute_workspace_provenance.py: enforce sibling pins + a host-infra
  self-check (installed omnibase_infra vs lock pin — the exact crash vector,
  since host infra is built from the context, not staged), and emit a
  pin_comparison block into build-provenance.json for deploy verifiers.
- Dockerfile.runtime: COPY resolve_workspace_pins.py beside the provenance
  script so the in-image import resolves.
- TDD: failing tests first (test_resolve_workspace_pins.py, 11 cases) +
  provenance integration tests; fixed a pre-existing stale RUNTIME_VERSION
  assertion to read pyproject dynamically.

Evidence-Ticket: OMN-12989

* test(OMN-12989): co-locate provenance pin helper in fixture

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13055): absent repos in REPOS list warn and exit 0 instead of failing (#1960)

Missing repos (not cloned locally) now emit a WARN line and are tracked
in a separate WARNED array. Only real fetch/ff failures cause exit 1.
This makes pull-all.sh safe to use on machines with a partial clone set,
while keeping the explicit-list override behavior intact.

Adds three regression tests: absent-only exits 0, absent+present exits 0
with OK for the present repo, present-failed+absent exits 1.

Evidence-Ticket: OMN-13055

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13070): config_prefetcher overlay wins on controlled lanes (#1961)

On infisical_required=True lanes, fetched/overlay config now always wins
over ambient env. Ambient env is retained only as a declared bootstrap
fallback (with an explicit provenance INFO log line) when Infisical
returns None. apply_to_environment also overwrites stale env on controlled
lanes. Uncontrolled lane (infisical_required=False) behaviour is unchanged.

Adds 5 regression tests: controlled-lane Infisical-wins, env-bootstrap-
fallback, apply_to_environment overwrite, missing-from-both-is-error, and
uncontrolled-lane-env-still-wins. Refactors _resolve_key to return a
(outcome, value, error) tuple to satisfy the ≤5-param pattern gate.

Source: docs/audits/2026-06-10-runtime-env-overlay-authority-audit.md

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13062): migration-gate vacuity fix — sentinel discipline, wait-for-postgres, skip-manifest, sync-check nonzero (retro A-10) (#1963)

* fix(OMN-13062): migration-gate vacuity fix — sentinel discipline, wait-for-postgres, skip-manifest, sync-check nonzero

Fixes three bugs identified in retro A-10 (recurrences OMN-12885, OMN-12934):

(1) RUNNER SENTINEL DISCIPLINE
    run-forward-migrations.sh now clears migrations_complete=FALSE at the
    start of every run and sets it TRUE only as its FINAL act after all
    infra and node migrations succeed. Any mid-run failure leaves the gate
    UNHEALTHY. runner_completed_at is stamped at the same final step as
    durable evidence of a successful completion run.

(2) SYNC-NODE-MIGRATIONS VACUOUS GATE
    sync-node-migrations.sh --check now exits 2 (not 0) when the omnimarket
    source tree is unresolvable. Silent exit-0 was hiding drift. The single
    opt-out is SYNC_NODE_MIGRATIONS_SKIP_UNRESOLVABLE=1 for environments
    that intentionally run without the source.

(3) WAIT-FOR-POSTGRES GUARD
    run-forward-migrations.sh now waits up to PG_WAIT_RETRIES (default 30)
    x 2s for Postgres to accept connections before proceeding, guarding
    the first-boot initdb race.

(4) SKIP-MANIFEST
    docker/migrations/skip-manifest.yaml introduced as the sole committed
    escape for intentionally-skipped migrations. The runner reads this at
    startup; listed migrations are recorded in schema_migrations with
    checksum "skip-manifest" without executing the SQL.

(5) MIGRATION 085
    Adds runner_completed_at TIMESTAMPTZ column to db_metadata so the
    runner's final stamp is durable in the schema (idempotent ADD COLUMN IF
    NOT EXISTS). Rollback included.

22 regression tests added covering all five fix surfaces.

* fix(OMN-13062): stamp schema fingerprint for migration 085

Migration 085 (085_add_runner_completed_at_to_db_metadata.sql) was added
in the initial commit but schema_fingerprint.sha256 was not regenerated.
Running `python scripts/check_schema_fingerprint.py stamp` updates the
artifact from the stale hash to match the 71 migration files.

Evidence-Source: OCC#2563
Evidence-Ticket: OMN-13062

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-12864): Bifrost endpoints → committed overlay authority + fail-loud loader (#1964)

* feat(OMN-12864, OMN-12814, OMN-12945): Bifrost endpoints → committed overlay authority + fail-loud loader

OMN-12864 — Committed lane overlay
  - docker/lane-overlays/dev.bifrost.yaml: typed deployment bindings for all
    four BIFROST_LOCAL_*_ENDPOINT_URL values (coder :8000, reasoner :8001,
    embedding :8100, ds4-flash :8101). Previously only available as ephemeral
    shell exports on .201; now committed, auditable, diff-able, and CI-checked.
  - docker/lane-overlays/dev.bifrost.env: generated dotenv sidecar consumed by
    compose via env_file; never edited directly (yaml is authority).
  - docker/docker-compose.infra.yml: wire the env_file block at compose root so
    the four endpoints are injected into the interpolation context on a clean
    shell. Hardcode BIFROST_CONTRACT_PATH (remove :-/empty footgun — OMN-12814).
  - scripts/render_bifrost_lane_overlay_env.py: render script regenerates the
    env sidecar from the YAML source.
  - src/omnibase_infra/runtime/models/model_bifrost_lane_overlay.py:
    ModelBifrostLaneOverlay — typed Pydantic model enforcing URL completeness
    (OMN-12815: every URL must end in /chat/completions).

OMN-12814 — Fail-loud loader
  - render_bifrost_delegation_contract: raises ProtocolConfigurationError on
    FileNotFoundError, YAMLError, ValidationError, and zero-endpoint renders.
    No lru_cache — every restart re-renders from packaged source so a stale
    cache cannot pin a broken result across deploys.

OMN-12945 — Re-seed from packaged source on deploy
  - docker/entrypoint-runtime.sh: set BIFROST_FORCE_RESEED=1 on every container
    restart so the named-volume copy is always rebuilt from the packaged
    bifrost_delegation.yaml merged with committed lane-overlay endpoints.
  - render_bifrost_delegation_contract: honor BIFROST_FORCE_RESEED/force_reseed
    flag to bypass the stale-volume early-return path entirely.

Tests:
  - tests/ci/test_bifrost_lane_overlay.py: CI gate — env sidecar in-sync with
    YAML source; all four BIFROST_LOCAL_* keys present.
  - tests/unit/runtime/models/test_model_bifrost_lane_overlay.py: bare-base URL
    rejection, env dict mapping, extra-field rejection.
  - tests/unit/runtime/test_render_bifrost_delegation_contract.py: fail-loud
    paths, force-reseed, zero-endpoint error, endpoint URL completeness.
  - tests/unit/models/test_model_serialization_roundtrip.py: roundtrip coverage.

* fix(OMN-12864): move bifrost env_file to service level — fix compose schema validation failure

Top-level 'env_file' is rejected by Docker Compose v2 schema validator
('additional properties not allowed'). This caused 10+ compose-render
integration tests to fail in CI.

Fix:
- Remove top-level env_file block from docker-compose.infra.yml
- Add per-service env_file on omninode-runtime, runtime-effects,
  runtime-worker (the three containers that render Bifrost)
- Change BIFROST_LOCAL_*:? to BIFROST_LOCAL_*:- in x-runtime-env
  (compose-level validation removed; Python validates via
  ModelBifrostLaneOverlay + render_bifrost_delegation_contract)
- Add two CI gate tests: compose_env_file_is_service_level_not_top_level
  and runtime_services_have_bifrost_env_file

* fix(OMN-12864): allow BIFROST_LOCAL_* empty defaults in silent-fallback gate

The overlay authority pattern (OMN-12864) passes BIFROST_LOCAL_*_ENDPOINT_URL
via service-level env_file (docker/lane-overlays/dev.bifrost.env), not at
compose config time. Compose-level :? would break CI rendering without the
overlay pre-loaded. Validation at the Python layer (ModelBifrostLaneOverlay
+ render_bifrost_delegation_contract) is the enforcement point.

Add the three failing vars to ALLOWED_EMPTY_DEFAULTS with OMN-12864 citation.

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-12858/OMN-12879/OMN-12880): dispatcher route-coverage CI gate (#1965)

* feat(OMN-12858/OMN-12879/OMN-12880): dispatcher route-coverage CI gate

Static contract analysis: every subscribed command topic (onex.cmd.*)
must declare handler_routing or runtime_dispatch, or the message goes to
DLQ silently. This gate would have caught two recent incidents:

1. June 9 DLQ regression (OMN-12858 post-mortem): node_generation_consumer
   subscribed onex.cmd.omnimarket.node-generation-requested.v1 but a
   sole-handler revert left zero dispatcher routes registered. Messages
   went to DLQ silently with no CI signal.

2. June 12 DEL-01 live finding: onex.cmd.omnimarket.delegate-skill.v1 was
   consumed by dev lane bus but no dispatcher route existed in any deployed
   contract. Discovered via manual rpk consumer-group lag probe (DEL-01
   evidence, docs/evidence/2026-06-12-weekend-pass/).

Deliverables:
- scripts/check_dispatcher_route_coverage.py — static YAML scanner that
  checks both omnibase_infra and omnimarket contract trees; ratchet
  allowlist for known pre-existing violations; --changed-contracts mode
  (OMN-12879) for per-PR scoping; compat publish topics excluded (OMN-12880)
- .github/workflows/dispatcher-route-coverage.yml — CI workflow that
  checks out omnimarket sibling, collects changed contract paths in PR
  mode, and runs the gate; fires on PR, push-to-main, and merge_group
- tests/ci/test_dispatcher_route_coverage_gate.py — 12 unit tests
  covering RED/GREEN/COMPAT/CHANGED-MODE/ALLOWLIST/MULTI-DIR paths plus
  live-contract regression proof against the actual omnibase_infra tree

Allowlist additions:
- onex.cmd.omnibase-infra.pattern-b-dispatch.v1 (RuntimePatternBBroker,
  imperative consumer, OMN-12525 migration target)
- onex.cmd.platform.contract-resolve-requested.v1 (transitional HTTP
  bridge node_contract_resolver_bridge OMN-2756, metadata.transitional=true)

[OMN-12858, OMN-12879, OMN-12880]

* fix(OMN-12858): drop full uv sync from dispatcher-route-coverage workflow

Gate script only needs pyyaml (stdlib + yaml). Using full setup-python-uv
was causing 10+ minute timeout. Replace with direct pip install pyyaml and
invoke python3 directly. Reduces job from 10m timeout to <1m.

[OMN-12858]

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-13026): wire transport-mock-lint pre-commit hook + CI gate (#1962)

* feat(OMN-13026): wire transport-mock-lint pre-commit hook + CI gate (omnibase_infra)

Activates the transport-mock-lint validator (from omnibase_core, OMN-13026)
on omnibase_infra. Ratchet baseline: 218 existing violations across 63 files
frozen in validation/transport_mock_baseline.yaml. New bare AsyncMock/
MagicMock on EventBus/transport surfaces are blocked by pre-commit hook and
CI lint step. Existing violations tracked for drain by per-site tickets
(parent OMN-13026).

Reference incident: PR #1181 bare AsyncMock hid missing EventBusKafka.stop().

Evidence-Ticket: OMN-13026

* fix(OMN-13026): use uv run python for CI lint step + bump omnibase_core pin to include transport_mock_lint

The transport-mock lint CI step previously cloned omnibase_core and ran
`python -m omnibase_core.validators.transport_mock_lint` with PYTHONPATH,
but this failed: `No module named omnibase_core.validators.transport_mock_lint`
because it ran `.venv/bin/python` which uses the locked venv, and the venv
omnibase_core pin (2defabef4) predates the transport_mock_lint module.

Fix: use `uv run python` (removes the clone step) and bump omnibase-core git
pin from 2defabef4 to 309d89fa7 (PR 1231 merge commit on dev) so
transport_mock_lint is available in the locked venv.

* fix(OMN-13026): align transport mock baseline and runner lock

* fix(OMN-13026): sync omnibase_core pin + runner identity lock to dev baseline

Align pyproject.toml omnibase_core rev to 2defabef (required by
test_release_backmerge_preserves_proven_runtime_core_pin) and update
docker/runners/runner-image.lock.json identity_digest/shared_env_digest
to match dev runner image lock (79b08f44 / 90c8b3b9).

Both were stale from the prior session's pin bump that used an older SHA.

* fix(OMN-13026): source transport mock validator from core

* fix(OMN-13026): source transport validator from core dev

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-13094): --output receipt mode on onex node/run — quiet typed receipts with durable capture (#1966)

Phase 2a of the skill-output-suppression slice (epic OMN-13089, plan
docs/plans/2026-06-12-skill-output-suppression-plan.md Phase 2 item 1).

- onex node/onex run gain --output receipt: ALL runtime logging routes to a
  run_id-suffixed capture file under <state-root>/captures/ (no console
  handlers — kills the 25-50-line RuntimeLocal INFO stream at the source);
  stdout carries exactly ONE typed ModelSkillResult JSON with the FULL
  handler result (result_model = concrete handler result type FQN).
- Durable capture: capture log + handler result content-addressed via
  omnibase_core ArtifactStore (OMN-13093); artifact.captured +
  tool.output.captured emitted to the emit daemon socket (--emit-socket,
  default ~/.claude/emit.sock).
- Failure asymmetry: artifact write failure => FULL output printed, no
  receipt (no hidden loss); emission failure => receipt still prints, event
  spooled to <state-root>/emit_spool/ for replay.
- Node failure => status=failed/error with full error + capture log INLINE
  in the receipt (errors are never hidden) and artifact-backed.
- Default output mode unchanged (enforcement is Phase 4).
- RuntimeLocal exposes handler_result (receipt schema identity).
- core pin 2defabef -> ae8793bd (merged OMN-13091/13093 receipt models +
  ArtifactStore); runner-image identity lock regenerated and the OMN-12765
  backmerge identity constants updated for the new pin.

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-13097): onex skill subcommand + declarative skill->node mapping (#1968)

* feat(OMN-13097): onex skill subcommand + declarative skill->node mapping

Phase 4a of the skill-output-suppression slice (epic OMN-13089, plan
docs/plans/2026-06-12-skill-output-suppression-plan.md Phase 4 item 1/2).

A dispatch skill IS one CLI call (user directive 2026-06-12). This adds the
`onex skill <name> [args]` dispatch surface that the 24 omniclaude shim
migrations build on:

- onex.cli entry-point `skill` -> cli_skill.run_skill_by_name. Resolves the
  skill via the declarative skill_mapping.yaml registry, builds the backing
  node's input payload from the skill's CLI args, writes it under the state
  root (.onex_state/tmp/<skill>-<run_id>.json — never /tmp), resolves the
  node's packaged contract exactly like `onex node`, and dispatches through
  the proven receipt-mode path (run_receipt_mode, OMN-13094). stdout is
  exactly one typed ModelSkillResult JSON with the FULL handler result.
- skill_mapping.yaml: declarative DATA mapping all 24 dispatch shims to their
  backing onex.nodes node + typed result-model FQN (verified against each
  live handler handle() return type on origin/dev) + per-arg payload specs +
  static payload + keyword classifiers (delegate task_type as data, not code).
  Adding a skill is a YAML edit + fixture, never a CLI code change (ticket
  deliverable 2/3). Mapping lives beside the node-resolution surface, never
  hardcoded branching in the CLI.
- Typed models split one-per-file (repo convention): ModelSkillArgSpec,
  ModelSkillClassifier, ModelSkillMapping, ModelSkillMappingRegistry,
  EnumSkillArgType. Frozen, extra=forbid, fail-fast coercion/validation.
- validation_exemptions.yaml: Click-callback param-count + literal-identifier
  name-field exemptions mirroring the existing cli_node run_node_by_name
  precedent (OMN-11570) — same pattern, same rationale.

dod_evidence:
- 20 unit tests pass (registry validity, all-24-shims coverage, FQN result
  models, arg parsing/coercion/positional/required, classifiers, payload
  build, receipt-mode dispatch wiring, payload-under-state-root not /tmp).
- uv run mypy src/ --strict: clean (2438 source files).
- ruff format + check: clean. pre-commit run on changed files: pass.
- runner-image identity lock regenerated for the pyproject entry-point add
  (same as OMN-13094).

* test(OMN-13097): rebind OMN-12765 backmerge identity constants for onex skill pyproject change

Adding the `skill` onex.cli entry-point to pyproject.toml changes the
runner-image identity_digest (and shared_env_digest) the lock binds. Update
the hardcoded expected constants in the backmerge-identity test to the
regenerated values — same mechanical rebind OMN-13094 performed for the core
pin bump. Identity + runner-image-identity tests pass (14).

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix+test(OMN-13012): force terminal-topic metadata refresh so both ephemeral consumers assign (runner consume-leg wedge) (#1969)

The runner consume-leg wedged on the live stability battery (image c0505521f1fa,
EXP1-3_RUNNER_CONSUME_LEG_BLOCKER): only the FAILED terminal topic ever subscribed
while the COMPLETED topic never assigned, so the correlated completed terminal was
never read and the 8x2x10 matrix re-fired cell 1 forever, emitting zero
non-degenerate rows. A prior two-strike diagnosis proved the omnimarket handler is
correct (it opens both terminal sessions pre-publish and races them); the defect is
in the omnibase_infra runtime consume leg.

Root cause: _assign_direct_terminal_partitions ignored the metadata future returned
by AIOKafkaClient.set_topics and re-called set_topics([same_topic]) each loop
iteration. aiokafka 0.13.0 set_topics only forces a metadata refresh when the topic
set DIFFERS from the tracked set, so every iteration after the first took the no-op
branch and never re-fetched. An ephemeral group_id=None consumer whose first metadata
fetch had not yet surfaced partitions burned the full 30s assign cap and raised a
bare TimeoutError (the empty-message 'wait failed' seen live).

Fix: register the reply topic once and await that metadata fetch, then on each miss
force a fresh fetch via force_metadata_update (which always fetches) rather than the
no-op set_topics repeat. The assign-cap TimeoutError now carries a diagnostic message
instead of an empty one.

Test: tests/integration/test_terminal_consumer_concurrent_assign_race.py drives the
REAL TerminalEventConsumer (the object wired as event_consumer) through the REAL
open_direct_terminal_consumer/poll path with AIOKafkaConsumer monkeypatched to a fake
that faithfully models aiokafka 0.13.0 set_topics future + metadata-latency semantics.
RED before the fix (bare TimeoutError, the live empty-message signature); GREEN after.
K>=2 multi-trial variant asserts no worker-thread leak across trials.

Evidence-Source: <occ-sha-pending>
Evidence-Ticket: OMN-13012

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* feat(OMN-13096): onex delegate single-command subcommand (Phase 2b) (#1967)

* feat(OMN-13096): onex delegate single-command subcommand

Add 'onex delegate "<prompt>" [--task-type X] [--max-tokens N]' as a
subcommand on the existing onex CLI (Phase 2b of the skill-output-suppression
slice, OMN-13089). The command wraps payload construction, node dispatch, and
result extraction internally and prints exactly one
ModelSkillResult[ModelDelegateSkillResponse] to stdout via the OMN-13094
receipt-mode path. RuntimeLocal logs go to the capture file + artifact store,
never to stdout; scratch payloads live under <state-root>/tmp/ with run_id
suffixes (never /tmp).

- cli_delegate.py: classify_task_type (keyword table from legacy skill md),
  payload write, contract resolve, run_receipt_mode dispatch
- register 'delegate' under onex.cli entry points
- exempt delegate_command from the >5-param patterns gate (same Click-callback
  rationale as run_node_by_name)
- 20 unit tests: classification, scratch-under-state-root, single typed
  receipt on stdout, zero INFO log leakage

omnibase_infra does NOT depend on omnimarket; the delegate node is resolved at
runtime via the onex.nodes entry-point group (registered by omnimarket).

* chore(OMN-13096): re-trigger deploy-gate after Evidence-Source set to OCC#2593

No code change — the deploy-gate workflow triggers on synchronize (not edited),
so the PR-body Evidence-Source fix needs a new commit to re-resolve the OCC ref
to the open PR head where contracts/OMN-13096.yaml (with deploy evidence) lives.

* chore(OMN-13096): re-trigger deploy-gate now that OCC#2593 merged to OCC dev

contracts/OMN-13096.yaml (with the dod-deploy-onex-delegate item) is now on
OCC dev, so the deploy-gate OCC-dev checkout resolves the contract + deploy
evidence.

* chore(OMN-13096): regenerate runner-image identity lock for pyproject entry-point add

Adding the 'delegate' onex.cli entry point to pyproject.toml changed the
dependency-manifest digest that scripts/ci/runner_image_identity.py folds into
the runner-image identity lock. Regenerate the lock so
tests/ci/test_runner_image_identity.py matches (was the only CI test failure;
unrelated environmental integration/perf failures excluded).

* test(OMN-13096): update backmerge identity assertions to regenerated lock digests

The runner-image identity lock was regenerated for the pyproject entry-point
add; this test hardcodes the expected identity_digest/shared_env_digest, so
update both to match the new lock (same maintenance OMN-13094 did).

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13118): tolerate partition-less reply topic in terminal consume leg (#1970)

The context-ROI runner opens one ephemeral group_id=None terminal consumer per
terminal topic BEFORE publishing each generation command (subscribe-before-publish,
OMN-13012/13038). The FAILED reply topic is only produced to on contract_passed=False;
in a battery where generations pass it has zero messages, so Redpanda never advertises
a partition for it. _assign_direct_terminal_partitions burned the full 30s assign cap
on every trial then raised a bare TimeoutError (the empty-message 'wait failed'),
stalling each of the 160 battery trials ~30s before the COMPLETED terminal could
correlate -> battery needs >80 min and never completes (verifier-confirmed wedge).

A partition-less reply topic is a valid steady state, not a 30s error:
- _assign_direct_terminal_partitions gives a bounded grace window for a topic that
  exists but is slow to surface metadata, then assigns whatever partitions exist
  (possibly none) and returns promptly instead of burning the cap and raising.
- poll_direct_terminal_consumer treats an empty assignment as 'no terminal will
  arrive here' (sleeps out its timeout, returns None) without calling getone() on
  an unassigned consumer.

Repro: tests/integration/test_terminal_consumer_battery_load_wedge.py drives the
REAL TerminalEventConsumer over K=10 x 2 cells x 2 arms; RED (40/40 trials block a
full assign cap) before the fix, GREEN after.

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13118): pin terminal-consumer read offset synchronously (close lazy seek_to_end publish race) (#1971)

* fix(OMN-13118): pin terminal-consumer read offset synchronously to close lazy seek_to_end publish race

The consume-leg wedge survived PR #1969 (set_topics no-op) and PR #1970
(partition-less assign-cap) because both addressed the assign phase, not the
seek timing. AIOKafkaConsumer.seek_to_end is LAZY: it requests a LATEST offset
reset that only resolves on the first poll — AFTER the caller publishes. With
generation completing in ~1s, the correlated COMPLETED terminal lands in the
open->poll gap, so the lazily-resolved LATEST position is the HWM AFTER the
record and the poll reads past it. The terminal is never read, the trial never
correlates, and the experiment matrix re-fires the same cell forever.

Replace seek_to_end with a synchronous end_offsets() + seek() pin
(_pin_direct_terminal_end_offsets) in both open_direct_terminal_consumer and
RuntimePatternBBroker._dispatch_and_wait_with_direct_kafka_consumer, so the
read position is fixed at open() time, before the publish — the real
subscribe-before-publish guarantee. Empty assignment (partition-less reply
topic, OMN-13118 #1970) is a no-op.

Repro: tests/integration/test_terminal_consumer_seek_reset_race.py drives the
real TerminalEventConsumer.open()/wait() the way HandlerContextRoiRunner does,
publishing the correlated terminal in the open->wait gap across K=10 x 2 cells
x 2 arms; RED with the lazy reset (every cell degenerate), GREEN once the read
offset is pinned. RED verified by git-stashing only the source fix.

Existing consume-leg fakes updated to model end_offsets/seek (they previously
masked the bug by making seek_to_end a synchronous exact snapshot).

* test(OMN-13118): reword assertion (lazy not deferred) for receipt honesty gate

* fix(OMN-13118): bound end_offsets() round-trip with assign-cap timeout (CodeRabbit)

end_offsets() is a broker ListOffsets round-trip aiokafka documents as able to
block indefinitely. Bound it with the same cap as start()/assign so a stalled
broker fails fast instead of hanging the pre-publish positioning.

---------

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13137): _validate_routing strategy-aware event_model checks (#1973)

operation_match routes by the `operation` field and does not use event_model.
The validator was unconditionally requiring event_model.{name,module} for every
handler entry, causing 230/295 omnimarket operation_match contracts (all correct
as authored) to fail routing validation at startup.

Fix: read routing_strategy from the routing map and branch validation:
  - payload_type_match → require event_model.{name, module} (unchanged)
  - operation_match (and any non-payload strategy) → require `operation`;
    skip event_model checks entirely

Updated pre-existing _validate_routing tests to declare routing_strategy:
payload_type_match explicitly (they always tested payload_type_match semantics
but relied on the implicit fallback that is now removed).

Added test_validate_routing_operation_match.py with 4 unit tests:
  1. operation_match without event_model → zero event_model errors
  2. operation_match missing operation field → error
  3. payload_type_match missing event_model → still errors (regression guard)
  4. Real node_integration_sweep_orchestrator routing block → clean (boot gate)

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13118): independent per-terminal-topic consumers in Pattern B direct-Kafka wait (#1972)

The consume-leg wedge survived four merged fixes (#1969 set_topics no-op,
#1970 partition-less assign-cap, #1971 synchronous seek-pin). The STRONG K>=10
multi-cell reprobe on the stability lane still wedged on REBUILD-5
(cdf53d963f7b). Converged diagnosis (strikes 3+4,
docs/evidence/2026-06-12-weekend-pass/experiments/probe4-stability/
reprobe-K10-rebuild5/HALT_K10_WEDGE_PERSISTS.md): the runtime waited for each
trial's terminal across TWO topics (node-generation-completed.v1 +
node-generation-failed.v1) with a SINGLE ephemeral group_id=None consumer
assigned both topics' partitions. One aiokafka consumer holds one manual
subscription; the COMPLETED delivery window collapsed before it surfaced the
correlated record, so the trial never correlated and the matrix re-fired cell 1.

RuntimePatternBBroker._dispatch_and_wait_with_direct_kafka_consumer now opens
ONE independent AIOKafkaConsumer PER terminal topic via
open_direct_terminal_consumer (each started, assigned, and offset-pinned via
end_offsets()+seek() at open() BEFORE publish), then awaits both CONCURRENTLY
via asyncio.wait(FIRST_COMPLETED). The first correlated terminal wins; both are
torn down. No shared consumer, no subscription flip.

Keeps the #1970 partition-less no-op and the #1971 synchronous offset pin (both
live in open_direct_terminal_consumer / poll_direct_terminal_consumer). Removes
the now-dead single-consumer helpers (_assign_terminal_topic_partitions,
method-level _refresh_terminal_topic_metadata, _direct_kafka_* kwargs builders,
_kafka_bootstrap_servers/_kafka_event_bus).

Adds tests/integration/test_terminal_consumer_subscription_flip_wedge.py: a
real-dispatch-path K>=10 x 2-cell x 2-arm repro whose fake models TWO
independent consumers honestly (delivery is faithful only for a single-topic
assignment; a consumer spanning both topics flips and drops the COMPLETED
record). RED genuineness verified by reverting only the source to the
single-consumer shape (test hangs past timeout); GREEN with the fix in 2.3s.

Acceptance is the LIVE K>=10 multi-cell stability-lane reprobe (later phase),
NOT this unit test. A green unit repro is necessary but NOT sufficient.

Refs OMN-13118, OMN-13128.

Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>

* fix(OMN-13118): long-lived terminal correlator replaces per-trial ephemeral consume leg (#1974)

Tier B canonical redesign (epic OMN-12525). Five offset/subscription patches
(#1969-#1972) tuned the per-trial-ephemeral terminal consumer and all failed…
* fix(OMN-15115): raise ModelLlmInferenceRequest timeout ceiling + per-model max_retries (#2442)

* fix(OMN-15115): raise ModelLlmInferenceRequest timeout ceiling + add per-model max_retries

qwen3-review-b's hostile-review timeout was pinned at this model's previous
le=600.0 ceiling -- the max the schema allowed. That made the OMN-14176
config-only fix (raising timeout_seconds to 600) structurally incapable of
ever curing the real defect: live-measured throughput on that endpoint is
4.3-4.6 tok/s (about half the ~9 tok/s the 600s value assumed), so a
genuine full-length completion could not finish inside 600s regardless of
contention.

- Raise timeout_seconds ceiling 600.0 -> 1800.0 on
  ModelLlmInferenceRequest (nodes/node_llm_inference_effect -- the model
  HandlerLlmOpenaiCompatible actually consumes; a differently-implemented,
  same-named model at omnibase_infra.models.llm is unrelated, serves the
  CLI-subprocess handler path, and already has its own max_retries field).
- Add a new max_retries field (default 3, matching the transport's
  historical hardcoded behavior) and thread it through
  HandlerLlmOpenaiCompatible._execute_with_auth -> both
  _execute_llm_http_call call sites, so a caller can lower retries for a
  systematically-slow (not transiently-flaky) endpoint without wasting a
  shared single-concurrency-slot's time on doomed retries.
- RED/GREEN proven via git stash: all 8 new/changed assertions fail
  (AttributeError / extra_forbidden / less_than_equal) against the
  pre-fix code, pass after.

OMN-15115

* fix(OMN-15115): sync llm inference contract timeout

* test(OMN-15115): sync llm inference contract runtime expectations

---------

Co-authored-by: Jonah Gray <jonah.g.gray@gmail.com>
(cherry picked from commit 70857e7)

* fix(OMN-14397): remove stale auto_merge skill mapping

* chore(OMN-15115): retrigger deploy gate after reusable workflow fix

---------

Co-authored-by: Jonah Gray <jonah.g.gray@gmail.com>
…me merge-commit exception; heals squash-severed merge-base (OMN-15195)
…2480

Empty commit — no tree change. #2478/#2480 shared the prior commit
SHA (3b67018) with this PR's now-closed
duplicate #2480, which made GitHub's commit->PR resolution in
occ-preflight/receipt-gate ambiguous (kept resolving to the closed #2480).
This commit gives #2479 a unique head SHA while leaving origin/dev's tree
untouched (git diff --quiet against the parent confirms zero diff).
…idge-v2

promote(OMN-15181): dev->main dev-wins bridge — one-time merge-commit exception (OMN-15195)
…-0384-promotion

# Conflicts:
#	src/omnibase_infra/nodes/node_runner_fleet_health_compute/handlers/handler_runner_fleet_health_evaluate.py
#	src/omnibase_infra/nodes/node_runner_health_snapshot_effect/handlers/handler_runner_fleet_snapshot.py
@coderabbitai

coderabbitai Bot commented Aug 14, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

You’ve reached a temporary PR review limit under our Fair Usage Limits Policy.

Your current included review allowance is based on your included PR review attempts over the past 7 days.

Next review available in: 33 minutes

Limit details: You’ve used the included review currently available. Your 112 included PR review attempts over the past 7 days set your current allowance at 1 review per hour.

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits within each organization.

For paid Pro and Pro+ reviews, CodeRabbit uses a developer's included PR review attempts over the past 7 days to set the current hourly allowance. At typical activity levels, the full plan allowance applies. Higher sustained activity can lower the allowance until earlier attempts leave the 7-day window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 7b73627e-7c92-4b09-bff6-93f85d69a8df

📥 Commits

Reviewing files that changed from the base of the PR and between b99e1a9 and 49ff6be.

📒 Files selected for processing (11)
  • .github/required-checks.yaml
  • .github/workflows/artifact-reconciliation-webhook.yml
  • .github/workflows/deploy-gate.yml
  • .github/workflows/env-parity.yml
  • scripts/validate_handler_contracts.py
  • src/omnibase_infra/nodes/node_runner_fleet_health_compute/contract.yaml
  • src/omnibase_infra/nodes/node_runner_fleet_health_compute/handlers/handler_runner_fleet_health_evaluate.py
  • src/omnibase_infra/nodes/node_runner_health_snapshot_effect/contract.yaml
  • src/omnibase_infra/nodes/node_runner_health_snapshot_effect/handlers/handler_runner_fleet_snapshot.py
  • src/omnibase_infra/observability/runner_health/model_runner_fleet_config.py
  • src/omnibase_infra/utils/util_runtime_packages.py

Comment @coderabbitai help to get the list of available commands.

@onexbot-occ-writer

Copy link
Copy Markdown
Contributor

OCC autobind did not mint a companion for this PR: no changed-file candidate could be proven RED against the merge base, and emitting a PR-existence probe instead would be non-falsifiable evidence (OMN-15247). Hand-authored evidence is required.

Resolves the two conflicts (pyproject.toml + uv.lock package version) in
favour of dev's 0.38.7. main is at 0.38.6, so dev's value already satisfies
the release-identity gate and taking the branch's stale 0.38.5 would regress it.

All other main-only surfaces auto-merged cleanly and were verified
file-by-file against origin/main.
@github-actions

Copy link
Copy Markdown
Contributor

✅ Hostile Reviewer — PASSED

Blocking findings (critical): 0
Total findings: 0
Models succeeded: qwen3-review,qwen3-review-b


Gate semantics (pilot phase)

Verdict Meaning Blocks merge?
passed No critical findings No
blocked CRITICAL findings found Yes
degraded All models unavailable (infra) No (pilot)

Powered by omniintelligence.review_pairing.cli_review — node-based adversarial review via HandlerLlmCliSubprocess (OMN-8468/OMN-8524)

jonahgabriel added a commit to OmniNode-ai/onex_change_control that referenced this pull request Aug 18, 2026
#6661)

* evidence(OMN-16041): add content-pinned entry for OmniNode-ai/omnibase_infra#2745

* fix(OMN-16041): drop probe echo variable interpolation for receipt-hardening (OMN-15710)

* evidence(OMN-16041): add self-bind entry for OCC#6661

* evidence(OMN-16041): dod_evidence receipt for OCC#6661 self-bind
…te + node_runner_health_snapshot_effect

Contract Sync Gate (Wave C, OMN-8915) flagged both nodes' handlers as
changed without a matching contract.yaml change on this backmerge.

Root cause: main-side commit b437e48 (OMN-15195, "remove direct env
reads from runner health bridge", 2026-07-26) changed both handlers
without a paired contract.yaml update -- this drift predates the
backmerge and was never caught on main, so there is no "ground truth"
contract state to recover from main; the contracts were simply never
updated when the handlers changed.

node_runner_fleet_health_compute: the contract's description claimed
RUNNER_HEALTH_MAX_DIAG_AGE_SECONDS was "env-overridable". OMN-15195
removed that env-read entirely (confirmed live in the current handler,
which now documents "Nothing constructs this handler with overrides");
the three thresholds are plain module constants (5 / 4500 / 600).
Corrected the description and added a Related Tickets entry.

node_runner_health_snapshot_effect: OMN-15195 replaced this handler's
direct env reads (WEDGE_QUEUE_AGE_SECONDS, RUNNER_CODELOAD_SCAN_LIMIT,
WEDGE_WATCH_REPOS) with the typed ModelRunnerFleetConfig, loaded from
config/runner_fleet.yaml. Added a Related Tickets entry documenting
the config-model migration; no prior description text was factually
wrong here (it didn't mention the env vars), so no description edit.

Both contracts get a patch version bump (documentation-only fix, no
input/output/topic/interface change) and updated metadata.updated /
metadata.ticket.

Verified: scripts/validate-pr-contract-sync.sh passes locally against
both changed files; scripts/validate.py all reports 0 errors for both
contracts (pre-existing runtime_profiles warnings on both files are
unrelated -- present before this change too).
@jonahgabriel
jonahgabriel merged commit bade69e into dev Aug 18, 2026
209 of 220 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant