Repository navigation
fix(OMN-13062): migration-gate vacuity fix — sentinel discipline, wait-for-postgres, skip-manifest, sync-check nonzero (retro A-10) - #1963
Conversation
…t-for-postgres, skip-manifest, sync-check nonzero
Fixes three bugs identified in retro A-10 (recurrences OMN-12885, OMN-12934):
(1) RUNNER SENTINEL DISCIPLINE
run-forward-migrations.sh now clears migrations_complete=FALSE at the
start of every run and sets it TRUE only as its FINAL act after all
infra and node migrations succeed. Any mid-run failure leaves the gate
UNHEALTHY. runner_completed_at is stamped at the same final step as
durable evidence of a successful completion run.
(2) SYNC-NODE-MIGRATIONS VACUOUS GATE
sync-node-migrations.sh --check now exits 2 (not 0) when the omnimarket
source tree is unresolvable. Silent exit-0 was hiding drift. The single
opt-out is SYNC_NODE_MIGRATIONS_SKIP_UNRESOLVABLE=1 for environments
that intentionally run without the source.
(3) WAIT-FOR-POSTGRES GUARD
run-forward-migrations.sh now waits up to PG_WAIT_RETRIES (default 30)
x 2s for Postgres to accept connections before proceeding, guarding
the first-boot initdb race.
(4) SKIP-MANIFEST
docker/migrations/skip-manifest.yaml introduced as the sole committed
escape for intentionally-skipped migrations. The runner reads this at
startup; listed migrations are recorded in schema_migrations with
checksum "skip-manifest" without executing the SQL.
(5) MIGRATION 085
Adds runner_completed_at TIMESTAMPTZ column to db_metadata so the
runner's final stamp is durable in the schema (idempotent ADD COLUMN IF
NOT EXISTS). Rollback included.
22 regression tests added covering all five fix surfaces.
|
Warning Review limit reached
More reviews will be available in 3 hours, 14 minutes, and 58 seconds. Learn how PR review limits work. Your organization has run out of usage credits. Purchase more credits in the billing tab to continue. ⌛ How to resolve this issue?After more reviews become available, a review can be triggered using the We recommend that you space out your commits to avoid hitting the rate limit. 🚦 How do rate limits work?CodeRabbit enforces hourly rate limits for each developer per organization. Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available. Please see our Fair Usage Limits Policy for further information. ℹ️ Review info⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (7)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
Migration 085 (085_add_runner_completed_at_to_db_metadata.sql) was added in the initial commit but schema_fingerprint.sha256 was not regenerated. Running `python scripts/check_schema_fingerprint.py stamp` updates the artifact from the stale hash to match the 71 migration files. Evidence-Source: OCC#2563 Evidence-Ticket: OMN-13062
…1988) * docs(OMN-12962): contract-store durability audit — cold-runtime census proof (#1937) Proves cold-start contract census reconstructs from the image-bundled filesystem manifest (HYBRID-mode bootstrap + PluginLoaderContractSource), independent of node-registration.v1 retention. The delete-retention topic feeds only the post-freeze dynamic listener (auto_offset_reset=latest, no history replay). Live .201 stability-test evidence: filesystem contract_path in manifest + truncated topic log head with intact census. No store fix needed; residual dynamic-only gap covered by runtime_sweep sweep check. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12970): vendor omnimarket projection node migrations into forward/nodes (#1938) Vendors omnimarket node-owned projection migrations into the namespaced forward-migration tree so run-forward-migrations.sh materializes them in the dashboard projection DB (omnidash_analytics) at deploy. Primary (OMN-12970): creates llm_call_metrics, llm_cost_aggregates, and capability_scores in the projection DB. These were only ever created in the omnibase_infra DB by infra migrations 031/060, so the ab-compare, cost.token_usage, cost.summary, and capability-scores projection topics were DEGRADED at startup ('table not found') and their dashboard panels rendered empty. Also re-syncs three omnimarket node migrations the vendor tree had drifted from (node_projection_llm_routing, node_projection_overnight, node_projection_savings /077) — sync-node-migrations.sh --check requires the full vendored tree to match omnimarket source, and these were missing. Companion to omnimarket PR for the same ticket (source migrations + projection table-coverage ratchet test). Evidence-Ticket: OMN-12970 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12965): stamp runtime image identity (version + revision) in workspace builds (#1943) * fix(OMN-12965): stamp runtime image identity (version + revision) in workspace builds The main runtime image stamped org.opencontainers.image.version=0.1.0 with a blank org.opencontainers.image.revision after workspace rebuilds. A blank identity degrades every proof packet (runtime SHA + image digest are required citations in accepted evidence). Root causes (three build paths under-stamped identity): - onex up --build (cmd_up) passed only GIT_SHA; the runtime-stage OCI labels read VCS_REF (-> blank revision) and RUNTIME_VERSION (-> placeholder 0.1.0). - deploy-runtime.sh passed VCS_REF but not RUNTIME_VERSION/GIT_SHA. - Dockerfile silently allowed blank/placeholder identity in workspace mode. Fix: - cli._image_identity_build_args() stamps the full quad (GIT_SHA/VCS_REF/ RUNTIME_VERSION/BUILD_DATE) and fails fast on an unresolved git revision. - deploy-runtime.sh stamps RUNTIME_VERSION + GIT_SHA and verifies the version label is non-placeholder post-deploy. - Dockerfile.runtime fails workspace builds with blank VCS_REF or placeholder RUNTIME_VERSION=0.1.0 (release mode unaffected). Enforcement ratchet (same PR): - scripts/check_runtime_image_identity.py static check, wired as pre-commit hook + CI gate (ci.yml). - tests/unit/infra/test_runtime_image_identity_labels.py pins the cli helpers + Dockerfile guard; deploy-agent test extended for the quad. Proven locally via throwaway docker builds: workspace+args -> populated labels; workspace without args -> guard fails (exit 64); release without args -> 0.1.0 placeholder allowed (no regression). Evidence-Ticket: OMN-12965 * test(OMN-12965): integration build proof for runtime image identity labels Builds the real runtime-stage ARG/LABEL/guard block against busybox and asserts via docker inspect: workspace+args -> populated version/revision; workspace without args -> guard fails (exit 64); release -> placeholder allowed. Satisfies the integration-test hard gate and makes the throwaway proof permanent. Evidence-Ticket: OMN-12965 --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12987): workspace-build sibling lock-pin preflight + provenance (#1944) Recurrence ratchet for the 2026-06-11 stability bootstrap crash. The 11:20Z --no-cache rebuild vendored omnibase_infra 0.37.0-dev (~2c1d672f) + core 0.42.0 even though omnimarket dev's uv.lock pinned infra 0.38.1 @ e2dbdc95 + core 0.44.0 @ c97c2c9a. The stale sibling predated the OMN-12501 Protocol-quarantine guard, turning a latent contract defect into a fatal crash that crash-looped the main runtime. - check_sibling_lock_pins.py: host-side fail-fast preflight resolving expected sibling versions/SHAs from the consuming repo's (omnimarket) uv.lock and comparing them against each vendored tree. Mismatch aborts the build. - stage_workspace.sh: emit a .build-sha marker per staged sibling (rsync drops .git) so the preflight and provenance can identify the vendored commit. - deploy-runtime.sh: run the preflight after staging, before build; abort on mismatch. Write the comparison under sibling-repos/.sibling-lock-pins.json. - compute_workspace_provenance.py + Dockerfile.runtime: fold expected-vs-actual lock_pin_comparison into build-provenance.json for deploy verifiers. Recurrence-guard tests prove a stale infra 0.37.0 vs lock-pinned 0.38.1 fails the preflight and matched pins pass; deploy script wiring is asserted statically. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12988): pin stability runtime-worker to replicas 1 + census ratchet (#1942) The base docker-compose.infra.yml defaults runtime-worker to replicas 0 (${WORKER_REPLICAS:-0}). The stability-test lane's required state includes a running worker (4-container census: main, effects, worker, projection-api), but the override pinned it via an env-interpolation default (${STABILITY_TEST_WORKER_REPLICAS:-1}) — a silent-drop surface: a stray exported STABILITY_TEST_WORKER_REPLICAS=0 or removal of the :-1 fallback would scale the worker to 0 with zero signal on a plain compose up/recreate. Fix: pin docker-compose.stability-test.yml runtime-worker deploy.replicas to the literal 1 (no env indirection). Ratchet (recurrence guards, same PR): - scripts/deploy-agent/tests/unit/test_runtime_worker_census.py: assert runtime-worker stays in the deploy-agent RUNTIME-scope census so a missing worker (replicas 0 => absent from docker compose ps) is a deploy failure, not silence; assert the override pins a literal 1. - tests/integration/infra/test_stability_test_runtime_compose_render.py: assert the rendered stability worker resolves deploy.replicas == 1. - tests/unit/infra/test_stability_test_runtime_lane.py: update the existing pin assertion to the literal 1. Evidence-Ticket: OMN-12988 Config-drift family: OMN-12945 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12979): expire-bound topic completeness suppressions (#1940) Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12966): accept community-edition Infisical /api/status in provision readiness gate (#1939) The fresh-provision readiness gate in provision-infisical.py only accepted the enterprise {"status": "ok"} payload and rejected the community edition's {"message": "Ok"}, returning 1 before bootstrap could run. This blocked provisioning against the Infisical instance deployed on .201 (community edition). Route the gate through the existing _is_infisical_ready helper (single source of truth, already used by the already-provisioned path). Add TestMainFreshProvision- ReadinessGate covering community/enterprise/not-ready cases. Also adds docker/docker-compose.infisical-stability.yml: an ADDITIVE Infisical compose project for the stability-test lane (joins the existing network as external, reuses stability postgres/valkey, no lane mutation), since the lane overlays disable the in-lane Infisical service via *-disabled profile overrides. P1.2b-A: Infisical now reachable from the stability runtime/effect containers; a known secret (OMN_12966_PROBE) seeds and resolves end-to-end via the machine- identity universal-auth path, verified from inside the effects container. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12958): volume-config drift gate + runtime config provenance (#1941) * feat(OMN-12958): volume-config drift gate + runtime provenance Compute config provenance (path + sha256) for the runtime-rendered Bifrost delegation contract; the deployed volume copy survives rebuilds and silently diverges from packaged source (two competing authorities, OMN-12945). - runtime/config_provenance.py: ModelConfigProvenance + drift classification, sidecar JSON writer (read by sweep + proof packets) - runtime/health/health_config_provenance.py: drift -> degraded health - render entrypoint logs provenance line + writes sidecar on every boot - docs/runbooks/volume-config-drift-and-reseed.md: ledgered re-seed procedure - validation exemption for config_name (logical identifier, not entity ref) No live volume mutation: re-seed is an operator deploy step (deploy_pending). * test(OMN-12958): cover volume config drift reseed flow --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12957): wire runtime-profile validator + core registry parity guard (#1945) * feat(OMN-12957): wire runtime-profile validator + core registry parity guard - Remove the import-time RuntimeError drift raise in runtime_profile.py (a hard raise on core/infra version skew would crash the kernel at import); the parity invariant is enforced by test_profiles_match_core_registry instead. - Add tests: _PROFILES keys == omnibase_core REGISTERED_RUNTIME_PROFILES, and every CONSUMER_ATTACHED_RUNTIME_PROFILES profile loads as a real lane. - Wire omnibase_core.validation.validator_runtime_profiles as a pre-commit hook + validator-runtime-profiles.yml CI gate on infra contracts. - Freeze 19 pre-existing violators in validation/runtime_profiles_allowlist.yaml (discovered by repo-root walk; drain via OMN-12982). Blocks NEW orphans. Requires the omnibase_core pin to include OMN-12957's validator (new rules). Evidence-Ticket: OMN-12957 Evidence-Source: 5463fbaf819409d4fb7f491dd4f276f10d869eba * ci(OMN-12957): pass runtime profile allowlist to validator * test(OMN-12957): cover runtime profile registry parity * fix(OMN-12957): keep runtime profile allowlist under config * fix(OMN-12957): pin core runtime profile registry --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12973): widen migration-gate healthcheck start_period + ratchet (#1950) P2.8: classify the prod migration-gate UNHEALTHY-then-self-resolved incident. Classification: idle-one-shot-mis-modeled = NO. The migration-gate is a correctly-modeled long-running sentinel (entrypoint `while true; sleep 3600`) whose healthcheck continuously polls db_metadata.migrations_complete via check_migrations_complete.sh. The container flipped UNHEALTHY transiently because its healthcheck start_period (10s) was far shorter than the real cold-volume migration window (~116s: prod gate started 09:35:22, intelligence-migration finished 09:37:18). Past the 10s grace window the still-failing probe was reported UNHEALTHY until migrations completed, then self-resolved — no fault. Fix: raise migration-gate healthcheck start_period 10s -> 180s in both the authoritative catalog manifest (docker/catalog/services/migration-gate.yaml, flows into the generated compose) and the hand-maintained docker/docker-compose.infra.yml that deploy-runtime.sh applies to .201, so a still-applying gate stays in `health: starting` instead of flipping UNHEALTHY. Ratchet (enforcement, not detection): new ValidatorHealthcheckStartPeriod in omnibase_infra catalog (validator_healthcheck_start_period.py) asserts a 120s floor for migration-completion gates, wired into `onex validate runtime` (cmd_validate_runtime) AND backed by unit tests that gate every PR via pre-commit + CI. A migration-completion gate is identified by BOTH a healthcheck that polls migration completion AND a service_completed_successfully dependency, so ordinary app services (e.g. intelligence-api, an HTTP liveness probe with a legitimately short 40s start_period) are not swept into the floor. Prod probed read-only only; no prod mutation. Applies to live prod via the batched stability/prod rebuild (deploy_pending). Evidence-Ticket: OMN-12973 Evidence-Source: OCC#PENDING Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12971): deliver Vertex ADC bearer token to runtime effects container (secret-ref discipline) (#1949) * feat(OMN-12971): deliver Vertex ADC bearer token to runtime effects container (secret-ref discipline) Wire the Vertex $500-credit ADC path into the runtime, ADDITIVE next to the Gemini API-key path (provider-agnostic; neither provider removed or forced). runtime_policy.contract.yaml: add a secret-source mapping llm.vertex.access_token (source_type=file, /run/secrets/vertex_access_token) to the dev/stability-test/ judge profiles, alongside the existing llm.gemini.api_key env mapping. The ref name MUST match cloud-vertex-gemini.secret_ref in omnimarket bifrost_delegation.yaml. The resolved VALUE is a short-lived OAuth bearer token minted from ADC, refreshed by the operator; the token VALUE is never committed — only the ref name + in-container path. Add aiplatform.googleapis.com to the cloud host allowlist. docker-compose.infra.yml: bind the operator-supplied host token file read-only to /run/secrets/vertex_access_token on the main and effects runtimes (VERTEX_ACCESS_TOKEN_HOST_FILE, default /dev/null so lanes without Vertex still start; Gemini key path unaffected). Pass through BIFROST_VERTEX_GEMINI_ENDPOINT_URL (overlay supplies the complete Vertex OpenAI-compat URL) and GOOGLE_CLOUD_PROJECT/LOCATION (default empty). runtime-policy.env: regenerated from the contract via render_runtime_policy_env (test_runtime_policy_env_matches_contract_renderer proves contract<->env parity). test_runtime_policy_contract.py: update host-allowlist assertion for the additive Vertex host. Pre-existing platform-wide topic-parity-gate failures (25 unrelated topics) are identical with and without this change (proven by diff); that gate is pre-commit- only (not a CI merge gate) and this change adds zero new topic gaps, so that one hook is SKIP-ped. No deploy/receipt/merge gate is bypassed. * fix(OMN-12971): make Vertex runtime env contract-owned --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13008): automated .201 disk maintenance — worktree GC + docker GC + watermark alert ratchet (#1952) * feat(OMN-13008): automated .201 disk maintenance — worktree GC + docker GC + watermark alert ratchet Conservative, keep-list-driven disk maintenance to prevent the 2026-06-11 /data ~95% outage that killed all three lanes mid-demo: - scripts/worktree-gc.sh: drives the canonical omniclaude prune-worktrees.sh (merged+clean+pushed safety) on both Mac (merge-sweep tick) and .201 - scripts/disk-gc.sh + disk_gc_plan.py: conservative docker/builder/image GC. Pure, testable removal planner honoring a VERSIONED keep-list (deploy/disk-gc/keep-list.yaml): never reaps a kept repo, kept tag, in-use image, or anything younger than min_age_days; keeps N superseded generations. - scripts/disk-watermark-check.sh + disk_watermark_event.py: df watermark ratchet. >=85% emits a typed disk-watermark bus event (warning) that the sweep auto-ticket path turns into a Linear ticket; >=90% emits critical. Broker addr is fail-fast from KAFKA_BOOTSTRAP_SERVERS (no localhost default). - deploy/disk-gc/: systemd USER timer (onex-disk-gc.timer/.service, hourly) + install-disk-gc.sh. User units, NOT lane containers. - tests: 20 unit tests incl. GC plan-safety invariants + dry-run proof that default mode issues no destructive op (a wrong-delete GC is worse than none). Contract: contracts/OMN-13008.yaml * fix(OMN-13008): pass docker inventory to GC planner via stdin, not env (ARG_MAX) On a host with many docker images, passing the full image/ps inventory as env vars to disk_gc_plan.py exceeds ARG_MAX ('Argument list too long', exit 126), producing an empty plan. Write inventory to per-run scratch files (under the log dir, never /tmp; cleaned on exit) and hand it to the planner on stdin as a JSON envelope. Verified the failure live on .201; planner now reads stdin. * fix(OMN-13008): simplify GC plan stdin pipe (two processes, no nested subprocess) * fix(OMN-13008): keep-wins reconciliation — never remove an image id with any keep reason A single image id can surface in multiple 'docker image ls' rows (one per repo:tag). One tag could route the id to dangling-removal while another routes it to keep (e.g. tagged 'latest' or within-N-generations). Live .201 plan showed an id in BOTH remove_image_ids and kept_reasons. Reconcile at the end: keep wins — any id with a keep reason is dropped from the remove list; remove list deduped. Adds 2 regression tests. Verified live on .201. * fix(OMN-13008): timer uses OnCalendar=hourly for reliable re-arm (oneshot service) OnUnitActiveSec does not reliably re-elapse for a oneshot service once it goes inactive (observed NextElapseUSecMonotonic=infinity live on .201 after the first run). Switch to OnCalendar=hourly + RandomizedDelaySec so the timer re-arms every hour. Keep Persistent=true for missed-run catch-up. --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13005): materialize blocking event_consumer in runtime auto-wiring (runner consume-leg degenerate rows) (#1951) * fix(OMN-13005): materialize blocking event_consumer in auto-wiring (runner consume-leg degenerate rows) The runtime auto-wiring materialized event_publisher for handlers that declare it but had no equivalent for event_consumer. Request/response EFFECT handlers (HandlerContextRoiRunner) that publish a command then block on the correlated terminal event fell back to their no-op consumer default, returning None immediately -> every result row degenerate (failure_stage=generation, attempt_count=0) while generations succeeded ~1s later. Adds _make_sync_event_consumer (mirror of _make_sync_event_publisher), backed by service_terminal_event_consumer.make_terminal_event_consumer: a sync (topic, correlation_id, timeout) -> dict | None adapter that runs the proven direct-Kafka correlate-and-wait loop (from RuntimePatternBBroker) on an isolated event loop in a worker thread, so blocking does not deadlock the runtime dispatch loop that delivers the awaited terminal. TDD through the REAL dispatch path: test_event_consumer_injection drives a trial through _prepare_handler_wiring with a terminal arriving after a delay and asserts a non-degenerate row; verified RED with injection disabled. * fix(OMN-13012): two-phase (seek-now/wait-later) terminal event_consumer to close the subscribe-after-publish race (#1954) The OMN-13005 injected event_consumer is a single callable that does assign -> seek_to_end -> poll internally, all AFTER the handler has already published its command. Once OMN-13010 freed the dispatch loop and generation began completing in ~1s, the correlated terminal lands BEFORE the single-call consumer's post-publish seek_to_end positions, so seek_to_end skips PAST the already-emitted terminal and the runner times out on an offset beyond it (probe3, run_id=20260611T2140Z-probe3 -- both arms degenerate, zero rebalances). Splits positioning from waiting so the caller subscribes BEFORE it publishes: session = consumer.open(topic) # assign + seek_to_end NOW publisher(command_topic, payload) # publish AFTER positioning payload = session.wait(cid, timeout) # block from the captured position The returned TerminalEventConsumer is still directly callable with the legacy (topic, cid, timeout) -> dict | None single-call shape for any consumer that does not need subscribe-before-publish; the runner is the only consumer today. TerminalConsumerSession owns a dedicated event loop on a daemon worker thread for the whole open->wait->close lifecycle, preserving the OMN-13005 loop-isolation discipline so blocking never deadlocks the runtime dispatch loop. TDD (real dispatch path, RED-then-GREEN): extends the OMN-13005 injection test with a terminal emitted IMMEDIATELY after publish. The single-call (seek-after-publish) consumer MISSES it (degenerate row -- RED test asserts failure_stage=generation); the two-phase (open-before-publish) consumer CATCHES it (non-degenerate -- GREEN). The Kafka layer is faked at the two service seams against a shared in-memory log modeling seek-to-end semantics. OMN-13005 blocking-correlate behavior preserved. 269/269 auto_wiring unit tests pass; mypy --strict clean. Sibling to OMN-13010 / OMN-13005 / OMN-13003. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13005): route terminal consumer through Kafka boundary --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12990): pin worker replicas in ledgered config + fail-fast lane overrides (silent-drop ratchet) (#1948) * fix(OMN-12990): pin worker replicas in ledgered config + fail-fast lane overrides The base compose sets runtime-worker deploy replicas to ${WORKER_REPLICAS:-0} (soft-default ZERO). The stability lane's required state includes a running worker (GATE_ZERO_PROOF.md: 4 runtime containers), but the worker presence rode entirely on a compose soft-default — any plain compose up/recreate without the policy env silently scaled the worker to zero with no error and no signal. Fix (contract-native + fail-fast): - Add 'replicas' to ModelRuntimeProcessPolicy; pin replicas: 1 in every lane's worker block in runtime_policy.contract.yaml. - Renderer emits {PROFILE}_WORKER_REPLICAS into the ledgered runtime-policy.env for dev/stability-test/judge/prod. - stability + prod compose overrides reference ${..._WORKER_REPLICAS:?...} (fail-fast, NO silent :-1/:-0 default). A recreate that omits the policy env now aborts loudly instead of dropping the worker. prod previously had no override at all and inherited the dangerous :-0 default. Ratchet (recurrence guards): - tests asserting fail-fast override form (no soft default), contract-declared replica pin >= 1 per lane, and rendered {PROFILE}_WORKER_REPLICAS=1 in the ledgered env. - runbook deploy/verify procedure adds an expected-container census (worker must be present) via verify_container_manifest; a missing worker is a FAILURE, not silence. Note: SKIP=topic-parity-gate — that local-only advisory gate (absent from all .github/workflows, not a required CI check) fails on 25 pre-existing cross-repo topic gaps (build-loop/omniclaude/omniweb) identical on pristine base HEAD 8d7da1249; this change adds zero topics. All other hooks ran clean. Evidence-Ticket: OMN-12990 * test(OMN-12990): cover worker replica policy integration --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12909): add gateway bus forwarder P0A (#1946) * feat(OMN-12909): add gateway bus forwarder p0a * test(OMN-12909): add gateway forwarder integration coverage * fix(OMN-12909): satisfy gateway forwarder validators * test(OMN-12909): allow gateway forwarder bus protocol * fix(OMN-12909): sync gateway forwarder entry point * fix(OMN-12909): refresh runner image identity lock * test(OMN-12909): relax JSON normalizer mixed benchmark threshold * fix(OMN-12909): allow gateway handlers to boot unconfigured * fix(OMN-12909): declare gateway forwarder runtime profile * fix(OMN-12909): update runner identity backmerge expectation --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13011): LANE CENSUS RECONCILIATION ratchet — declared desired-state per lane, drift = auto-ticket (#1955) The class fix for the recurring lane-drift regression. Nothing reconciled the declared desired state of a runtime lane against what is actually running, so the same failure kept recurring with zero signal: volume config drift (OMN-12945), WORKER_REPLICAS silent zero (OMN-12988/12990), and on 2026-06-11 prod runtime containers plus the broker network were silently absent for hours during demo prep. Ships a per-lane DESIRED-STATE census: - (a) DECLARED in a versioned lane manifest (deploy/lane-census/lane-manifest.yaml): container set, network, replicas, image-tag pattern per lane (stability-test/prod/judge/dev), derived from the canonical compose lane files. A parity ratchet keeps the manifest locked in step with the compose files. - (b) RECONCILED on a schedule on .201 by SHARING the OMN-13008 systemd timer (a drop-in 4th ExecStart on onex-disk-gc.service — never a second timer) and on-demand via scripts/lane-census-check.sh / runtime_sweep. - (c) Drift = typed bus event (onex.evt.infra.lane-census-drift.v1) + Linear auto-ticket naming exactly what is missing/extra (container_absent, network_detached, replicas_zero, unexpected_container, oneshot_failed/stuck, image_tag_mismatch). Fail-fast, no warn-only mode (gates-block policy); exit 30 on drift; bus publish fail-fast on KAFKA_BOOTSTRAP_SERVERS (no localhost default). Red fixture reproduces 2026-06-11: prod runtime containers absent + broker network detached must produce the exact drift findings + a non-zero exit hours before a human noticed. Pure planner is fully unit-tested; shell driver dry-run-tested. Builds on the OMN-12988 deploy-agent RUNTIME census (deploy-time) as the complementary steady-state reconciler; closes the runtime-worker.yaml container_name: null census gap by sourcing names from the compose lane files. Evidence-Ticket: OMN-13011 Config-drift family: OMN-12945 Relates-to: OMN-13009, OMN-12988, OMN-13008 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13020): vendor missing node migrations — llm_routing 0000 + context_roi 001 (#1956) Vendors two omnimarket node-source migrations into the infra forward-migration tree via scripts/sync-node-migrations.sh (the canonical OMN-12559 mechanism): - node_projection_llm_routing/0000_create_llm_routing_decisions.sql (source: omnimarket #1168 / OMN-12942, merge ed6734f8) - node_projection_context_roi/001_create_context_roi_scores.sql (source: omnimarket #1178 / OMN-12955, merge 5010b1f4) Without the 0000 base table, node_projection_llm_routing/0001 (CREATE VIEW) hard-fails against NODE_POSTGRES_DB=omnidash_analytics — exactly the prod forward-migration exit-3 of 2026-06-11T09:35:52Z, and reproduced by construction in any clean clones@dev build. Files are byte-identical to the omnimarket dev blobs (sha256 f8a8b339… / c4126e65…) and to the untracked hot-patch copies on the .201 stability clone. Both migrations are self-contained, all-statements-IF-NOT-EXISTS, and 0000 sorts lexically before 0001 within the node's namespaced identity space. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13058): close TerminalConsumerSession on open() failure (worker thread + event loop leak) (#1957) TerminalConsumerSession.__init__ starts its dedicated worker loop thread immediately. TerminalEventConsumer.open did 'session = TerminalConsumerSession(...); return session.open()' with no cleanup: any failure inside session.open() (consumer start timeout, partition-assign timeout, broker auth error) propagated out of the raising expression, the session reference was lost, and the daemon worker thread plus its never-closed asyncio event loop leaked -- one pair per failed open. The motivating caller (HandlerContextRoiRunner) opens a session per trial, so a 160-560-trial battery against a degraded broker accumulates hundreds of leaked threads in the long-lived effects container. Fix: wrap session.open() in try/except BaseException -> session.close() (idempotent: stops the loop, joins the thread) -> re-raise. Covers both the two-phase .open(topic) path and the legacy single-call __call__ path. Found by the P3.3 doctrinal review of merged #1951 (b9712af9 / 36d98275). TDD: tests/unit/runtime/test_service_terminal_event_consumer_open_failure.py injects a real open failure through the production path (event bus without _bootstrap_servers) and asserts no alive terminal-consumer-* thread after the raise. Verified RED with the fix stashed (2 failed), GREEN with it (2 passed). Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13021): non-dev-base guard — fail feature-base PRs absent Stacked-Parent declaration (retro A-6) (#1958) Any PR whose base is neither dev nor main fails unless the body carries 'Stacked-Parent: #N'. Prevents the feedback_stacked_prs_orphan_from_dev class (#1185/#1954 auto-merged INTO parent feature branches and stranded off dev). base=main remains governed by main-target-guard. Epic OMN-13013 (process enforcement ratchets — June 12 retro). Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13014): hot-patch ledger rebuild preflight gate (retro B-1) (#1959) * feat(OMN-13014): hot-patch ledger rebuild preflight gate (retro B-1) Hot-patches on .201 (.prepatch sibling discipline) silently revert on any image rebuild/force-recreate — the 2026-06-11 20:58Z rebuild already erased a live /api/generate patch once. This adds the rebuild-path gate: - scripts/preflight_hotpatch_ledger.py: given a target container or lane + per-repo build refs, hard-fails when any hot-patch ledger row's source PR merge commit is not an ancestor of the build ref (git merge-base --is-ancestor), plus a .prepatch tripwire of the running container (unledgered .prepatch = hard fail; --post-rebuild = zero .prepatch expected). Sole bypass: HOTPATCH_PREFLIGHT_BYPASS carrying the Rule-10 '# skip-token-allowed: <user-approval-receipt-id>' form. - scripts/deploy-runtime.sh: guard_hotpatch_ledger wired into main() before build/preview (both dry-run and execute), lane derived from the compose project; skips loudly only when no ledger exists on the host. - tests/unit/scripts/test_preflight_hotpatch_ledger.py: 17 unit tests (ancestor gate, lane scoping, ledger loading, tripwire, bypass forms). - tests/ci/test_receipt_gate_install_guard.py: repair stale guard — core OMN-12565 replaced the OMN-9198 'uv pip uninstall first' install step with a cleared workspace venv (uv venv --clear); assert the new contract. Ledger backfilled from live census (5 .prepatch files / 4 source PRs, all MERGED to dev) at /data/omninode/hotpatch-ledger/ledger.yaml on .201. * fix(OMN-13014): scope missing-prepatch tripwire warning to the probed container --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12977): workspace-build sibling-pin ratchet — honor consuming lock, fail-fast on stale vendor (#1936) * fix(OMN-12977): workspace-build sibling-pin ratchet — honor consuming lock, fail-fast on stale vendor The workspace-mode image build (BUILD_SOURCE=workspace, used by the stability-test deploy procedure) vendored sibling/foundation packages from whatever the canonical OMNI_HOME clones happened to be checked out at, ignoring the consuming repo's uv.lock. On 2026-06-11 this shipped a 13-day-stale omnibase_infra 0.37.0-dev (pre-OMN-12501 Protocol-quarantine guard) + core 0.42.0 against an omnimarket dev lock pinning infra 0.38.1@e2dbdc95 / core 0.44.0@c97c2c9a, dropping the guard and crashing wire_from_manifest bootstrap fatally (stability lane down on demo day). Fix + recurrence ratchet (same PR): - scripts/runtime_build/check_sibling_lock_pins.py: parse the consuming repo's uv.lock for expected version+git-rev of each foundation/sibling package (scoped to the package's own source line so editable/registry pins are not cross-attributed a dependency's rev), resolve the actual clone version+HEAD, compare, and classify drift backward/forward/none. Fail-fast (exit 1) on any drift; --allow-drift records an explicit operator override in the artifact, never silent. - stage_workspace.sh: runs the preflight against the canonical clones before staging; aborts the build (exit 3) on unacknowledged drift and writes workspace/sibling-pin-comparison.json. - compute_workspace_provenance.py: folds the expected-vs-actual comparison into build-provenance.json so deploy verifiers can assert the build honored the lock; flags unacknowledged drift as a provenance error. - Dockerfile.runtime: COPY the comparison artifact (committed placeholder so the COPY always resolves; overwritten by stage_workspace.sh in workspace mode). - TDD: 19 unit tests covering lock parsing (git/registry/editable sources), drift classification, the exact 0.37.0-vs-0.38.1 stale case, check_pins exit codes, and the allow-drift override. Pre-existing mypy-strict bare-dict errors in compute_workspace_provenance.py fixed in the same pass. Evidence-Ticket: OMN-12977 Evidence-Source: pending-occ * ci(OMN-12977): retry runtime smoke compose port race * test(OMN-12977): align sibling-pin script tests with current API --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12989): workspace-mode image build must honor sibling lock pins (#1947) * fix(OMN-12989): workspace build must honor sibling lock pins (fail-fast ratchet) The 2026-06-11 stability bootstrap crash was caused by a workspace-mode --no-cache rebuild that vendored omnibase_infra 0.37.0-dev (13-day-stale worktree) even though omnimarket dev's uv.lock pins infra 0.38.1. The downgraded sibling predated the OMN-12501 Protocol-quarantine guard and turned a latent contract defect into a fatal crash. Fix + ratchet (same PR set): - scripts/runtime_build/resolve_workspace_pins.py: parse the consuming repo's uv.lock for sibling pins (version + git rev); classify each staged/installed sibling as exact/ahead/regression/unpinned; fail-fast (WorkspacePinError) on any regression BELOW the lock pin. Stdlib version-tuple fallback when packaging is absent so the ratchet never fails-open. - compute_workspace_provenance.py: enforce sibling pins + a host-infra self-check (installed omnibase_infra vs lock pin — the exact crash vector, since host infra is built from the context, not staged), and emit a pin_comparison block into build-provenance.json for deploy verifiers. - Dockerfile.runtime: COPY resolve_workspace_pins.py beside the provenance script so the in-image import resolves. - TDD: failing tests first (test_resolve_workspace_pins.py, 11 cases) + provenance integration tests; fixed a pre-existing stale RUNTIME_VERSION assertion to read pyproject dynamically. Evidence-Ticket: OMN-12989 * fix(OMN-12989): workspace build must honor sibling lock pins (fail-fast ratchet) The 2026-06-11 stability bootstrap crash was caused by a workspace-mode --no-cache rebuild that vendored omnibase_infra 0.37.0-dev (13-day-stale worktree) even though omnimarket dev's uv.lock pins infra 0.38.1. The downgraded sibling predated the OMN-12501 Protocol-quarantine guard and turned a latent contract defect into a fatal crash. Fix + ratchet (same PR set): - scripts/runtime_build/resolve_workspace_pins.py: parse the consuming repo's uv.lock for sibling pins (version + git rev); classify each staged/installed sibling as exact/ahead/regression/unpinned; fail-fast (WorkspacePinError) on any regression BELOW the lock pin. Stdlib version-tuple fallback when packaging is absent so the ratchet never fails-open. - compute_workspace_provenance.py: enforce sibling pins + a host-infra self-check (installed omnibase_infra vs lock pin — the exact crash vector, since host infra is built from the context, not staged), and emit a pin_comparison block into build-provenance.json for deploy verifiers. - Dockerfile.runtime: COPY resolve_workspace_pins.py beside the provenance script so the in-image import resolves. - TDD: failing tests first (test_resolve_workspace_pins.py, 11 cases) + provenance integration tests; fixed a pre-existing stale RUNTIME_VERSION assertion to read pyproject dynamically. Evidence-Ticket: OMN-12989 * test(OMN-12989): co-locate provenance pin helper in fixture --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13055): absent repos in REPOS list warn and exit 0 instead of failing (#1960) Missing repos (not cloned locally) now emit a WARN line and are tracked in a separate WARNED array. Only real fetch/ff failures cause exit 1. This makes pull-all.sh safe to use on machines with a partial clone set, while keeping the explicit-list override behavior intact. Adds three regression tests: absent-only exits 0, absent+present exits 0 with OK for the present repo, present-failed+absent exits 1. Evidence-Ticket: OMN-13055 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13070): config_prefetcher overlay wins on controlled lanes (#1961) On infisical_required=True lanes, fetched/overlay config now always wins over ambient env. Ambient env is retained only as a declared bootstrap fallback (with an explicit provenance INFO log line) when Infisical returns None. apply_to_environment also overwrites stale env on controlled lanes. Uncontrolled lane (infisical_required=False) behaviour is unchanged. Adds 5 regression tests: controlled-lane Infisical-wins, env-bootstrap- fallback, apply_to_environment overwrite, missing-from-both-is-error, and uncontrolled-lane-env-still-wins. Refactors _resolve_key to return a (outcome, value, error) tuple to satisfy the ≤5-param pattern gate. Source: docs/audits/2026-06-10-runtime-env-overlay-authority-audit.md Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13062): migration-gate vacuity fix — sentinel discipline, wait-for-postgres, skip-manifest, sync-check nonzero (retro A-10) (#1963) * fix(OMN-13062): migration-gate vacuity fix — sentinel discipline, wait-for-postgres, skip-manifest, sync-check nonzero Fixes three bugs identified in retro A-10 (recurrences OMN-12885, OMN-12934): (1) RUNNER SENTINEL DISCIPLINE run-forward-migrations.sh now clears migrations_complete=FALSE at the start of every run and sets it TRUE only as its FINAL act after all infra and node migrations succeed. Any mid-run failure leaves the gate UNHEALTHY. runner_completed_at is stamped at the same final step as durable evidence of a successful completion run. (2) SYNC-NODE-MIGRATIONS VACUOUS GATE sync-node-migrations.sh --check now exits 2 (not 0) when the omnimarket source tree is unresolvable. Silent exit-0 was hiding drift. The single opt-out is SYNC_NODE_MIGRATIONS_SKIP_UNRESOLVABLE=1 for environments that intentionally run without the source. (3) WAIT-FOR-POSTGRES GUARD run-forward-migrations.sh now waits up to PG_WAIT_RETRIES (default 30) x 2s for Postgres to accept connections before proceeding, guarding the first-boot initdb race. (4) SKIP-MANIFEST docker/migrations/skip-manifest.yaml introduced as the sole committed escape for intentionally-skipped migrations. The runner reads this at startup; listed migrations are recorded in schema_migrations with checksum "skip-manifest" without executing the SQL. (5) MIGRATION 085 Adds runner_completed_at TIMESTAMPTZ column to db_metadata so the runner's final stamp is durable in the schema (idempotent ADD COLUMN IF NOT EXISTS). Rollback included. 22 regression tests added covering all five fix surfaces. * fix(OMN-13062): stamp schema fingerprint for migration 085 Migration 085 (085_add_runner_completed_at_to_db_metadata.sql) was added in the initial commit but schema_fingerprint.sha256 was not regenerated. Running `python scripts/check_schema_fingerprint.py stamp` updates the artifact from the stale hash to match the 71 migration files. Evidence-Source: OCC#2563 Evidence-Ticket: OMN-13062 --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12864): Bifrost endpoints → committed overlay authority + fail-loud loader (#1964) * feat(OMN-12864, OMN-12814, OMN-12945): Bifrost endpoints → committed overlay authority + fail-loud loader OMN-12864 — Committed lane overlay - docker/lane-overlays/dev.bifrost.yaml: typed deployment bindings for all four BIFROST_LOCAL_*_ENDPOINT_URL values (coder :8000, reasoner :8001, embedding :8100, ds4-flash :8101). Previously only available as ephemeral shell exports on .201; now committed, auditable, diff-able, and CI-checked. - docker/lane-overlays/dev.bifrost.env: generated dotenv sidecar consumed by compose via env_file; never edited directly (yaml is authority). - docker/docker-compose.infra.yml: wire the env_file block at compose root so the four endpoints are injected into the interpolation context on a clean shell. Hardcode BIFROST_CONTRACT_PATH (remove :-/empty footgun — OMN-12814). - scripts/render_bifrost_lane_overlay_env.py: render script regenerates the env sidecar from the YAML source. - src/omnibase_infra/runtime/models/model_bifrost_lane_overlay.py: ModelBifrostLaneOverlay — typed Pydantic model enforcing URL completeness (OMN-12815: every URL must end in /chat/completions). OMN-12814 — Fail-loud loader - render_bifrost_delegation_contract: raises ProtocolConfigurationError on FileNotFoundError, YAMLError, ValidationError, and zero-endpoint renders. No lru_cache — every restart re-renders from packaged source so a stale cache cannot pin a broken result across deploys. OMN-12945 — Re-seed from packaged source on deploy - docker/entrypoint-runtime.sh: set BIFROST_FORCE_RESEED=1 on every container restart so the named-volume copy is always rebuilt from the packaged bifrost_delegation.yaml merged with committed lane-overlay endpoints. - render_bifrost_delegation_contract: honor BIFROST_FORCE_RESEED/force_reseed flag to bypass the stale-volume early-return path entirely. Tests: - tests/ci/test_bifrost_lane_overlay.py: CI gate — env sidecar in-sync with YAML source; all four BIFROST_LOCAL_* keys present. - tests/unit/runtime/models/test_model_bifrost_lane_overlay.py: bare-base URL rejection, env dict mapping, extra-field rejection. - tests/unit/runtime/test_render_bifrost_delegation_contract.py: fail-loud paths, force-reseed, zero-endpoint error, endpoint URL completeness. - tests/unit/models/test_model_serialization_roundtrip.py: roundtrip coverage. * fix(OMN-12864): move bifrost env_file to service level — fix compose schema validation failure Top-level 'env_file' is rejected by Docker Compose v2 schema validator ('additional properties not allowed'). This caused 10+ compose-render integration tests to fail in CI. Fix: - Remove top-level env_file block from docker-compose.infra.yml - Add per-service env_file on omninode-runtime, runtime-effects, runtime-worker (the three containers that render Bifrost) - Change BIFROST_LOCAL_*:? to BIFROST_LOCAL_*:- in x-runtime-env (compose-level validation removed; Python validates via ModelBifrostLaneOverlay + render_bifrost_delegation_contract) - Add two CI gate tests: compose_env_file_is_service_level_not_top_level and runtime_services_have_bifrost_env_file * fix(OMN-12864): allow BIFROST_LOCAL_* empty defaults in silent-fallback gate The overlay authority pattern (OMN-12864) passes BIFROST_LOCAL_*_ENDPOINT_URL via service-level env_file (docker/lane-overlays/dev.bifrost.env), not at compose config time. Compose-level :? would break CI rendering without the overlay pre-loaded. Validation at the Python layer (ModelBifrostLaneOverlay + render_bifrost_delegation_contract) is the enforcement point. Add the three failing vars to ALLOWED_EMPTY_DEFAULTS with OMN-12864 citation. --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12858/OMN-12879/OMN-12880): dispatcher route-coverage CI gate (#1965) * feat(OMN-12858/OMN-12879/OMN-12880): dispatcher route-coverage CI gate Static contract analysis: every subscribed command topic (onex.cmd.*) must declare handler_routing or runtime_dispatch, or the message goes to DLQ silently. This gate would have caught two recent incidents: 1. June 9 DLQ regression (OMN-12858 post-mortem): node_generation_consumer subscribed onex.cmd.omnimarket.node-generation-requested.v1 but a sole-handler revert left zero dispatcher routes registered. Messages went to DLQ silently with no CI signal. 2. June 12 DEL-01 live finding: onex.cmd.omnimarket.delegate-skill.v1 was consumed by dev lane bus but no dispatcher route existed in any deployed contract. Discovered via manual rpk consumer-group lag probe (DEL-01 evidence, docs/evidence/2026-06-12-weekend-pass/). Deliverables: - scripts/check_dispatcher_route_coverage.py — static YAML scanner that checks both omnibase_infra and omnimarket contract trees; ratchet allowlist for known pre-existing violations; --changed-contracts mode (OMN-12879) for per-PR scoping; compat publish topics excluded (OMN-12880) - .github/workflows/dispatcher-route-coverage.yml — CI workflow that checks out omnimarket sibling, collects changed contract paths in PR mode, and runs the gate; fires on PR, push-to-main, and merge_group - tests/ci/test_dispatcher_route_coverage_gate.py — 12 unit tests covering RED/GREEN/COMPAT/CHANGED-MODE/ALLOWLIST/MULTI-DIR paths plus live-contract regression proof against the actual omnibase_infra tree Allowlist additions: - onex.cmd.omnibase-infra.pattern-b-dispatch.v1 (RuntimePatternBBroker, imperative consumer, OMN-12525 migration target) - onex.cmd.platform.contract-resolve-requested.v1 (transitional HTTP bridge node_contract_resolver_bridge OMN-2756, metadata.transitional=true) [OMN-12858, OMN-12879, OMN-12880] * fix(OMN-12858): drop full uv sync from dispatcher-route-coverage workflow Gate script only needs pyyaml (stdlib + yaml). Using full setup-python-uv was causing 10+ minute timeout. Replace with direct pip install pyyaml and invoke python3 directly. Reduces job from 10m timeout to <1m. [OMN-12858] --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13026): wire transport-mock-lint pre-commit hook + CI gate (#1962) * feat(OMN-13026): wire transport-mock-lint pre-commit hook + CI gate (omnibase_infra) Activates the transport-mock-lint validator (from omnibase_core, OMN-13026) on omnibase_infra. Ratchet baseline: 218 existing violations across 63 files frozen in validation/transport_mock_baseline.yaml. New bare AsyncMock/ MagicMock on EventBus/transport surfaces are blocked by pre-commit hook and CI lint step. Existing violations tracked for drain by per-site tickets (parent OMN-13026). Reference incident: PR #1181 bare AsyncMock hid missing EventBusKafka.stop(). Evidence-Ticket: OMN-13026 * fix(OMN-13026): use uv run python for CI lint step + bump omnibase_core pin to include transport_mock_lint The transport-mock lint CI step previously cloned omnibase_core and ran `python -m omnibase_core.validators.transport_mock_lint` with PYTHONPATH, but this failed: `No module named omnibase_core.validators.transport_mock_lint` because it ran `.venv/bin/python` which uses the locked venv, and the venv omnibase_core pin (2defabef4) predates the transport_mock_lint module. Fix: use `uv run python` (removes the clone step) and bump omnibase-core git pin from 2defabef4 to 309d89fa7 (PR 1231 merge commit on dev) so transport_mock_lint is available in the locked venv. * fix(OMN-13026): align transport mock baseline and runner lock * fix(OMN-13026): sync omnibase_core pin + runner identity lock to dev baseline Align pyproject.toml omnibase_core rev to 2defabef (required by test_release_backmerge_preserves_proven_runtime_core_pin) and update docker/runners/runner-image.lock.json identity_digest/shared_env_digest to match dev runner image lock (79b08f44 / 90c8b3b9). Both were stale from the prior session's pin bump that used an older SHA. * fix(OMN-13026): source transport mock validator from core * fix(OMN-13026): source transport validator from core dev --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13094): --output receipt mode on onex node/run — quiet typed receipts with durable capture (#1966) Phase 2a of the skill-output-suppression slice (epic OMN-13089, plan docs/plans/2026-06-12-skill-output-suppression-plan.md Phase 2 item 1). - onex node/onex run gain --output receipt: ALL runtime logging routes to a run_id-suffixed capture file under <state-root>/captures/ (no console handlers — kills the 25-50-line RuntimeLocal INFO stream at the source); stdout carries exactly ONE typed ModelSkillResult JSON with the FULL handler result (result_model = concrete handler result type FQN). - Durable capture: capture log + handler result content-addressed via omnibase_core ArtifactStore (OMN-13093); artifact.captured + tool.output.captured emitted to the emit daemon socket (--emit-socket, default ~/.claude/emit.sock). - Failure asymmetry: artifact write failure => FULL output printed, no receipt (no hidden loss); emission failure => receipt still prints, event spooled to <state-root>/emit_spool/ for replay. - Node failure => status=failed/error with full error + capture log INLINE in the receipt (errors are never hidden) and artifact-backed. - Default output mode unchanged (enforcement is Phase 4). - RuntimeLocal exposes handler_result (receipt schema identity). - core pin 2defabef -> ae8793bd (merged OMN-13091/13093 receipt models + ArtifactStore); runner-image identity lock regenerated and the OMN-12765 backmerge identity constants updated for the new pin. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13097): onex skill subcommand + declarative skill->node mapping (#1968) * feat(OMN-13097): onex skill subcommand + declarative skill->node mapping Phase 4a of the skill-output-suppression slice (epic OMN-13089, plan docs/plans/2026-06-12-skill-output-suppression-plan.md Phase 4 item 1/2). A dispatch skill IS one CLI call (user directive 2026-06-12). This adds the `onex skill <name> [args]` dispatch surface that the 24 omniclaude shim migrations build on: - onex.cli entry-point `skill` -> cli_skill.run_skill_by_name. Resolves the skill via the declarative skill_mapping.yaml registry, builds the backing node's input payload from the skill's CLI args, writes it under the state root (.onex_state/tmp/<skill>-<run_id>.json — never /tmp), resolves the node's packaged contract exactly like `onex node`, and dispatches through the proven receipt-mode path (run_receipt_mode, OMN-13094). stdout is exactly one typed ModelSkillResult JSON with the FULL handler result. - skill_mapping.yaml: declarative DATA mapping all 24 dispatch shims to their backing onex.nodes node + typed result-model FQN (verified against each live handler handle() return type on origin/dev) + per-arg payload specs + static payload + keyword classifiers (delegate task_type as data, not code). Adding a skill is a YAML edit + fixture, never a CLI code change (ticket deliverable 2/3). Mapping lives beside the node-resolution surface, never hardcoded branching in the CLI. - Typed models split one-per-file (repo convention): ModelSkillArgSpec, ModelSkillClassifier, ModelSkillMapping, ModelSkillMappingRegistry, EnumSkillArgType. Frozen, extra=forbid, fail-fast coercion/validation. - validation_exemptions.yaml: Click-callback param-count + literal-identifier name-field exemptions mirroring the existing cli_node run_node_by_name precedent (OMN-11570) — same pattern, same rationale. dod_evidence: - 20 unit tests pass (registry validity, all-24-shims coverage, FQN result models, arg parsing/coercion/positional/required, classifiers, payload build, receipt-mode dispatch wiring, payload-under-state-root not /tmp). - uv run mypy src/ --strict: clean (2438 source files). - ruff format + check: clean. pre-commit run on changed files: pass. - runner-image identity lock regenerated for the pyproject entry-point add (same as OMN-13094). * test(OMN-13097): rebind OMN-12765 backmerge identity constants for onex skill pyproject change Adding the `skill` onex.cli entry-point to pyproject.toml changes the runner-image identity_digest (and shared_env_digest) the lock binds. Update the hardcoded expected constants in the backmerge-identity test to the regenerated values — same mechanical rebind OMN-13094 performed for the core pin bump. Identity + runner-image-identity tests pass (14). --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix+test(OMN-13012): force terminal-topic metadata refresh so both ephemeral consumers assign (runner consume-leg wedge) (#1969) The runner consume-leg wedged on the live stability battery (image c0505521f1fa, EXP1-3_RUNNER_CONSUME_LEG_BLOCKER): only the FAILED terminal topic ever subscribed while the COMPLETED topic never assigned, so the correlated completed terminal was never read and the 8x2x10 matrix re-fired cell 1 forever, emitting zero non-degenerate rows. A prior two-strike diagnosis proved the omnimarket handler is correct (it opens both terminal sessions pre-publish and races them); the defect is in the omnibase_infra runtime consume leg. Root cause: _assign_direct_terminal_partitions ignored the metadata future returned by AIOKafkaClient.set_topics and re-called set_topics([same_topic]) each loop iteration. aiokafka 0.13.0 set_topics only forces a metadata refresh when the topic set DIFFERS from the tracked set, so every iteration after the first took the no-op branch and never re-fetched. An ephemeral group_id=None consumer whose first metadata fetch had not yet surfaced partitions burned the full 30s assign cap and raised a bare TimeoutError (the empty-message 'wait failed' seen live). Fix: register the reply topic once and await that metadata fetch, then on each miss force a fresh fetch via force_metadata_update (which always fetches) rather than the no-op set_topics repeat. The assign-cap TimeoutError now carries a diagnostic message instead of an empty one. Test: tests/integration/test_terminal_consumer_concurrent_assign_race.py drives the REAL TerminalEventConsumer (the object wired as event_consumer) through the REAL open_direct_terminal_consumer/poll path with AIOKafkaConsumer monkeypatched to a fake that faithfully models aiokafka 0.13.0 set_topics future + metadata-latency semantics. RED before the fix (bare TimeoutError, the live empty-message signature); GREEN after. K>=2 multi-trial variant asserts no worker-thread leak across trials. Evidence-Source: <occ-sha-pending> Evidence-Ticket: OMN-13012 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13096): onex delegate single-command subcommand (Phase 2b) (#1967) * feat(OMN-13096): onex delegate single-command subcommand Add 'onex delegate "<prompt>" [--task-type X] [--max-tokens N]' as a subcommand on the existing onex CLI (Phase 2b of the skill-output-suppression slice, OMN-13089). The command wraps payload construction, node dispatch, and result extraction internally and prints exactly one ModelSkillResult[ModelDelegateSkillResponse] to stdout via the OMN-13094 receipt-mode path. RuntimeLocal logs go to the capture file + artifact store, never to stdout; scratch payloads live under <state-root>/tmp/ with run_id suffixes (never /tmp). - cli_delegate.py: classify_task_type (keyword table from legacy skill md), payload write, contract resolve, run_receipt_mode dispatch - register 'delegate' under onex.cli entry points - exempt delegate_command from the >5-param patterns gate (same Click-callback rationale as run_node_by_name) - 20 unit tests: classification, scratch-under-state-root, single typed receipt on stdout, zero INFO log leakage omnibase_infra does NOT depend on omnimarket; the delegate node is resolved at runtime via the onex.nodes entry-point group (registered by omnimarket). * chore(OMN-13096): re-trigger deploy-gate after Evidence-Source set to OCC#2593 No code change — the deploy-gate workflow triggers on synchronize (not edited), so the PR-body Evidence-Source fix needs a new commit to re-resolve the OCC ref to the open PR head where contracts/OMN-13096.yaml (with deploy evidence) lives. * chore(OMN-13096): re-trigger deploy-gate now that OCC#2593 merged to OCC dev contracts/OMN-13096.yaml (with the dod-deploy-onex-delegate item) is now on OCC dev, so the deploy-gate OCC-dev checkout resolves the contract + deploy evidence. * chore(OMN-13096): regenerate runner-image identity lock for pyproject entry-point add Adding the 'delegate' onex.cli entry point to pyproject.toml changed the dependency-manifest digest that scripts/ci/runner_image_identity.py folds into the runner-image identity lock. Regenerate the lock so tests/ci/test_runner_image_identity.py matches (was the only CI test failure; unrelated environmental integration/perf failures excluded). * test(OMN-13096): update backmerge identity assertions to regenerated lock digests The runner-image identity lock was regenerated for the pyproject entry-point add; this test hardcodes the expected identity_digest/shared_env_digest, so update both to match the new lock (same maintenance OMN-13094 did). --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): tolerate partition-less reply topic in terminal consume leg (#1970) The context-ROI runner opens one ephemeral group_id=None terminal consumer per terminal topic BEFORE publishing each generation command (subscribe-before-publish, OMN-13012/13038). The FAILED reply topic is only produced to on contract_passed=False; in a battery where generations pass it has zero messages, so Redpanda never advertises a partition for it. _assign_direct_terminal_partitions burned the full 30s assign cap on every trial then raised a bare TimeoutError (the empty-message 'wait failed'), stalling each of the 160 battery trials ~30s before the COMPLETED terminal could correlate -> battery needs >80 min and never completes (verifier-confirmed wedge). A partition-less reply topic is a valid steady state, not a 30s error: - _assign_direct_terminal_partitions gives a bounded grace window for a topic that exists but is slow to surface metadata, then assigns whatever partitions exist (possibly none) and returns promptly instead of burning the cap and raising. - poll_direct_terminal_consumer treats an empty assignment as 'no terminal will arrive here' (sleeps out its timeout, returns None) without calling getone() on an unassigned consumer. Repro: tests/integration/test_terminal_consumer_battery_load_wedge.py drives the REAL TerminalEventConsumer over K=10 x 2 cells x 2 arms; RED (40/40 trials block a full assign cap) before the fix, GREEN after. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): pin terminal-consumer read offset synchronously (close lazy seek_to_end publish race) (#1971) * fix(OMN-13118): pin terminal-consumer read offset synchronously to close lazy seek_to_end publish race The consume-leg wedge survived PR #1969 (set_topics no-op) and PR #1970 (partition-less assign-cap) because both addressed the assign phase, not the seek timing. AIOKafkaConsumer.seek_to_end is LAZY: it requests a LATEST offset reset that only resolves on the first poll — AFTER the caller publishes. With generation completing in ~1s, the correlated COMPLETED terminal lands in the open->poll gap, so the lazily-resolved LATEST position is the HWM AFTER the record and the poll reads past it. The terminal is never read, the trial never correlates, and the experiment matrix re-fires the same cell forever. Replace seek_to_end with a synchronous end_offsets() + seek() pin (_pin_direct_terminal_end_offsets) in both open_direct_terminal_consumer and RuntimePatternBBroker._dispatch_and_wait_with_direct_kafka_consumer, so the read position is fixed at open() time, before the publish — the real subscribe-before-publish guarantee. Empty assignment (partition-less reply topic, OMN-13118 #1970) is a no-op. Repro: tests/integration/test_terminal_consumer_seek_reset_race.py drives the real TerminalEventConsumer.open()/wait() the way HandlerContextRoiRunner does, publishing the correlated terminal in the open->wait gap across K=10 x 2 cells x 2 arms; RED with the lazy reset (every cell degenerate), GREEN once the read offset is pinned. RED verified by git-stashing only the source fix. Existing consume-leg fakes updated to model end_offsets/seek (they previously masked the bug by making seek_to_end a synchronous exact snapshot). * test(OMN-13118): reword assertion (lazy not deferred) for receipt honesty gate * fix(OMN-13118): bound end_offsets() round-trip with assign-cap timeout (CodeRabbit) end_offsets() is a broker ListOffsets round-trip aiokafka documents as able to block indefinitely. Bound it with the same cap as start()/assign so a stalled broker fails fast instead of hanging the pre-publish positioning. --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13137): _validate_routing strategy-aware event_model checks (#1973) operation_match routes by the `operation` field and does not use event_model. The validator was unconditionally requiring event_model.{name,module} for every handler entry, causing 230/295 omnimarket operation_match contracts (all correct as authored) to fail routing validation at startup. Fix: read routing_strategy from the routing map and branch validation: - payload_type_match → require event_model.{name, module} (unchanged) - operation_match (and any non-payload strategy) → require `operation`; skip event_model checks entirely Updated pre-existing _validate_routing tests to declare routing_strategy: payload_type_match explicitly (they always tested payload_type_match semantics but relied on the implicit fallback that is now removed). Added test_validate_routing_operation_match.py with 4 unit tests: 1. operation_match without event_model → zero event_model errors 2. operation_match missing operation field → error 3. payload_type_match missing event_model → still errors (regression guard) 4. Real node_integration_sweep_orchestrator routing block → clean (boot gate) Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): independent per-terminal-topic consumers in Pattern B direct-Kafka wait (#1972) The consume-leg wedge survived four merged fixes (#1969 set_topics no-op, #1970 partition-less assign-cap, #1971 synchronous seek-pin). The STRONG K>=10 multi-cell reprobe on the stability lane still wedged on REBUILD-5 (cdf53d963f7b). Converged diagnosis (strikes 3+4, docs/evidence/2026-06-12-weekend-pass/experiments/probe4-stability/ reprobe-K10-rebuild5/HALT_K10_WEDGE_PERSISTS.md): the runtime waited for each trial's terminal across TWO topics (node-generation-completed.v1 + node-generation-failed.v1) with a SINGLE ephemeral group_id=None consumer assigned both topics' partitions. One aiokafka consumer holds one manual subscription; the COMPLETED delivery window collapsed before it surfaced the correlated record, so the trial never correlated and the matrix re-fired cell 1. RuntimePatternBBroker._dispatch_and_wait_with_direct_kafka_consumer now opens ONE independent AIOKafkaConsumer PER terminal topic via open_direct_terminal_consumer (each started, assigned, and offset-pinned via end_offsets()+seek() at open() BEFORE publish), then awaits both CONCURRENTLY via asyncio.wait(FIRST_COMPLETED). The first correlated terminal wins; both are torn down. No shared consumer, no subscription flip. Keeps the #1970 partition-less no-op and the #1971 synchronous offset pin (both live in open_direct_terminal_consumer / poll_direct_terminal_consumer). Removes the now-dead single-consumer helpers (_assign_terminal_topic_partitions, method-level _refresh_terminal_topic_metadata, _direct_kafka_* kwargs builders, _kafka_bootstrap_servers/_kafka_event_bus). Adds tests/integration/test_terminal_consumer_subscription_flip_wedge.py: a real-dispatch-path K>=10 x 2-cell x 2-arm repro whose fake models TWO independent consumers honestly (delivery is faithful only for a single-topic assignment; a consumer spanning both topics flips and drops the COMPLETED record). RED genuineness verified by reverting only the source to the single-consumer shape (test hangs past timeout); GREEN with the fix in 2.3s. Acceptance is the LIVE K>=10 multi-cell stability-lane reprobe (later phase), NOT this unit test. A green unit repro is necessary but NOT sufficient. Refs OMN-13118, OMN-13128. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): long-lived terminal correlator replaces per-trial ephemeral consume leg (#1974) Tier B canonical redesign (epic OMN-12525). Five offset/subscription patches (#1969-#1972) tuned the per-trial-ephemeral terminal consumer and all failed the live K>=10 batt…
….46.1 / spi 0.23.0) (#2164) * docs(OMN-12962): contract-store durability audit — cold-runtime census proof (#1937) Proves cold-start contract census reconstructs from the image-bundled filesystem manifest (HYBRID-mode bootstrap + PluginLoaderContractSource), independent of node-registration.v1 retention. The delete-retention topic feeds only the post-freeze dynamic listener (auto_offset_reset=latest, no history replay). Live .201 stability-test evidence: filesystem contract_path in manifest + truncated topic log head with intact census. No store fix needed; residual dynamic-only gap covered by runtime_sweep sweep check. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12970): vendor omnimarket projection node migrations into forward/nodes (#1938) Vendors omnimarket node-owned projection migrations into the namespaced forward-migration tree so run-forward-migrations.sh materializes them in the dashboard projection DB (omnidash_analytics) at deploy. Primary (OMN-12970): creates llm_call_metrics, llm_cost_aggregates, and capability_scores in the projection DB. These were only ever created in the omnibase_infra DB by infra migrations 031/060, so the ab-compare, cost.token_usage, cost.summary, and capability-scores projection topics were DEGRADED at startup ('table not found') and their dashboard panels rendered empty. Also re-syncs three omnimarket node migrations the vendor tree had drifted from (node_projection_llm_routing, node_projection_overnight, node_projection_savings /077) — sync-node-migrations.sh --check requires the full vendored tree to match omnimarket source, and these were missing. Companion to omnimarket PR for the same ticket (source migrations + projection table-coverage ratchet test). Evidence-Ticket: OMN-12970 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12965): stamp runtime image identity (version + revision) in workspace builds (#1943) * fix(OMN-12965): stamp runtime image identity (version + revision) in workspace builds The main runtime image stamped org.opencontainers.image.version=0.1.0 with a blank org.opencontainers.image.revision after workspace rebuilds. A blank identity degrades every proof packet (runtime SHA + image digest are required citations in accepted evidence). Root causes (three build paths under-stamped identity): - onex up --build (cmd_up) passed only GIT_SHA; the runtime-stage OCI labels read VCS_REF (-> blank revision) and RUNTIME_VERSION (-> placeholder 0.1.0). - deploy-runtime.sh passed VCS_REF but not RUNTIME_VERSION/GIT_SHA. - Dockerfile silently allowed blank/placeholder identity in workspace mode. Fix: - cli._image_identity_build_args() stamps the full quad (GIT_SHA/VCS_REF/ RUNTIME_VERSION/BUILD_DATE) and fails fast on an unresolved git revision. - deploy-runtime.sh stamps RUNTIME_VERSION + GIT_SHA and verifies the version label is non-placeholder post-deploy. - Dockerfile.runtime fails workspace builds with blank VCS_REF or placeholder RUNTIME_VERSION=0.1.0 (release mode unaffected). Enforcement ratchet (same PR): - scripts/check_runtime_image_identity.py static check, wired as pre-commit hook + CI gate (ci.yml). - tests/unit/infra/test_runtime_image_identity_labels.py pins the cli helpers + Dockerfile guard; deploy-agent test extended for the quad. Proven locally via throwaway docker builds: workspace+args -> populated labels; workspace without args -> guard fails (exit 64); release without args -> 0.1.0 placeholder allowed (no regression). Evidence-Ticket: OMN-12965 * test(OMN-12965): integration build proof for runtime image identity labels Builds the real runtime-stage ARG/LABEL/guard block against busybox and asserts via docker inspect: workspace+args -> populated version/revision; workspace without args -> guard fails (exit 64); release -> placeholder allowed. Satisfies the integration-test hard gate and makes the throwaway proof permanent. Evidence-Ticket: OMN-12965 --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12987): workspace-build sibling lock-pin preflight + provenance (#1944) Recurrence ratchet for the 2026-06-11 stability bootstrap crash. The 11:20Z --no-cache rebuild vendored omnibase_infra 0.37.0-dev (~2c1d672f) + core 0.42.0 even though omnimarket dev's uv.lock pinned infra 0.38.1 @ e2dbdc95 + core 0.44.0 @ c97c2c9a. The stale sibling predated the OMN-12501 Protocol-quarantine guard, turning a latent contract defect into a fatal crash that crash-looped the main runtime. - check_sibling_lock_pins.py: host-side fail-fast preflight resolving expected sibling versions/SHAs from the consuming repo's (omnimarket) uv.lock and comparing them against each vendored tree. Mismatch aborts the build. - stage_workspace.sh: emit a .build-sha marker per staged sibling (rsync drops .git) so the preflight and provenance can identify the vendored commit. - deploy-runtime.sh: run the preflight after staging, before build; abort on mismatch. Write the comparison under sibling-repos/.sibling-lock-pins.json. - compute_workspace_provenance.py + Dockerfile.runtime: fold expected-vs-actual lock_pin_comparison into build-provenance.json for deploy verifiers. Recurrence-guard tests prove a stale infra 0.37.0 vs lock-pinned 0.38.1 fails the preflight and matched pins pass; deploy script wiring is asserted statically. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12988): pin stability runtime-worker to replicas 1 + census ratchet (#1942) The base docker-compose.infra.yml defaults runtime-worker to replicas 0 (${WORKER_REPLICAS:-0}). The stability-test lane's required state includes a running worker (4-container census: main, effects, worker, projection-api), but the override pinned it via an env-interpolation default (${STABILITY_TEST_WORKER_REPLICAS:-1}) — a silent-drop surface: a stray exported STABILITY_TEST_WORKER_REPLICAS=0 or removal of the :-1 fallback would scale the worker to 0 with zero signal on a plain compose up/recreate. Fix: pin docker-compose.stability-test.yml runtime-worker deploy.replicas to the literal 1 (no env indirection). Ratchet (recurrence guards, same PR): - scripts/deploy-agent/tests/unit/test_runtime_worker_census.py: assert runtime-worker stays in the deploy-agent RUNTIME-scope census so a missing worker (replicas 0 => absent from docker compose ps) is a deploy failure, not silence; assert the override pins a literal 1. - tests/integration/infra/test_stability_test_runtime_compose_render.py: assert the rendered stability worker resolves deploy.replicas == 1. - tests/unit/infra/test_stability_test_runtime_lane.py: update the existing pin assertion to the literal 1. Evidence-Ticket: OMN-12988 Config-drift family: OMN-12945 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12979): expire-bound topic completeness suppressions (#1940) Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12966): accept community-edition Infisical /api/status in provision readiness gate (#1939) The fresh-provision readiness gate in provision-infisical.py only accepted the enterprise {"status": "ok"} payload and rejected the community edition's {"message": "Ok"}, returning 1 before bootstrap could run. This blocked provisioning against the Infisical instance deployed on .201 (community edition). Route the gate through the existing _is_infisical_ready helper (single source of truth, already used by the already-provisioned path). Add TestMainFreshProvision- ReadinessGate covering community/enterprise/not-ready cases. Also adds docker/docker-compose.infisical-stability.yml: an ADDITIVE Infisical compose project for the stability-test lane (joins the existing network as external, reuses stability postgres/valkey, no lane mutation), since the lane overlays disable the in-lane Infisical service via *-disabled profile overrides. P1.2b-A: Infisical now reachable from the stability runtime/effect containers; a known secret (OMN_12966_PROBE) seeds and resolves end-to-end via the machine- identity universal-auth path, verified from inside the effects container. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12958): volume-config drift gate + runtime config provenance (#1941) * feat(OMN-12958): volume-config drift gate + runtime provenance Compute config provenance (path + sha256) for the runtime-rendered Bifrost delegation contract; the deployed volume copy survives rebuilds and silently diverges from packaged source (two competing authorities, OMN-12945). - runtime/config_provenance.py: ModelConfigProvenance + drift classification, sidecar JSON writer (read by sweep + proof packets) - runtime/health/health_config_provenance.py: drift -> degraded health - render entrypoint logs provenance line + writes sidecar on every boot - docs/runbooks/volume-config-drift-and-reseed.md: ledgered re-seed procedure - validation exemption for config_name (logical identifier, not entity ref) No live volume mutation: re-seed is an operator deploy step (deploy_pending). * test(OMN-12958): cover volume config drift reseed flow --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12957): wire runtime-profile validator + core registry parity guard (#1945) * feat(OMN-12957): wire runtime-profile validator + core registry parity guard - Remove the import-time RuntimeError drift raise in runtime_profile.py (a hard raise on core/infra version skew would crash the kernel at import); the parity invariant is enforced by test_profiles_match_core_registry instead. - Add tests: _PROFILES keys == omnibase_core REGISTERED_RUNTIME_PROFILES, and every CONSUMER_ATTACHED_RUNTIME_PROFILES profile loads as a real lane. - Wire omnibase_core.validation.validator_runtime_profiles as a pre-commit hook + validator-runtime-profiles.yml CI gate on infra contracts. - Freeze 19 pre-existing violators in validation/runtime_profiles_allowlist.yaml (discovered by repo-root walk; drain via OMN-12982). Blocks NEW orphans. Requires the omnibase_core pin to include OMN-12957's validator (new rules). Evidence-Ticket: OMN-12957 Evidence-Source: 5463fbaf819409d4fb7f491dd4f276f10d869eba * ci(OMN-12957): pass runtime profile allowlist to validator * test(OMN-12957): cover runtime profile registry parity * fix(OMN-12957): keep runtime profile allowlist under config * fix(OMN-12957): pin core runtime profile registry --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12973): widen migration-gate healthcheck start_period + ratchet (#1950) P2.8: classify the prod migration-gate UNHEALTHY-then-self-resolved incident. Classification: idle-one-shot-mis-modeled = NO. The migration-gate is a correctly-modeled long-running sentinel (entrypoint `while true; sleep 3600`) whose healthcheck continuously polls db_metadata.migrations_complete via check_migrations_complete.sh. The container flipped UNHEALTHY transiently because its healthcheck start_period (10s) was far shorter than the real cold-volume migration window (~116s: prod gate started 09:35:22, intelligence-migration finished 09:37:18). Past the 10s grace window the still-failing probe was reported UNHEALTHY until migrations completed, then self-resolved — no fault. Fix: raise migration-gate healthcheck start_period 10s -> 180s in both the authoritative catalog manifest (docker/catalog/services/migration-gate.yaml, flows into the generated compose) and the hand-maintained docker/docker-compose.infra.yml that deploy-runtime.sh applies to .201, so a still-applying gate stays in `health: starting` instead of flipping UNHEALTHY. Ratchet (enforcement, not detection): new ValidatorHealthcheckStartPeriod in omnibase_infra catalog (validator_healthcheck_start_period.py) asserts a 120s floor for migration-completion gates, wired into `onex validate runtime` (cmd_validate_runtime) AND backed by unit tests that gate every PR via pre-commit + CI. A migration-completion gate is identified by BOTH a healthcheck that polls migration completion AND a service_completed_successfully dependency, so ordinary app services (e.g. intelligence-api, an HTTP liveness probe with a legitimately short 40s start_period) are not swept into the floor. Prod probed read-only only; no prod mutation. Applies to live prod via the batched stability/prod rebuild (deploy_pending). Evidence-Ticket: OMN-12973 Evidence-Source: OCC#PENDING Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12971): deliver Vertex ADC bearer token to runtime effects container (secret-ref discipline) (#1949) * feat(OMN-12971): deliver Vertex ADC bearer token to runtime effects container (secret-ref discipline) Wire the Vertex $500-credit ADC path into the runtime, ADDITIVE next to the Gemini API-key path (provider-agnostic; neither provider removed or forced). runtime_policy.contract.yaml: add a secret-source mapping llm.vertex.access_token (source_type=file, /run/secrets/vertex_access_token) to the dev/stability-test/ judge profiles, alongside the existing llm.gemini.api_key env mapping. The ref name MUST match cloud-vertex-gemini.secret_ref in omnimarket bifrost_delegation.yaml. The resolved VALUE is a short-lived OAuth bearer token minted from ADC, refreshed by the operator; the token VALUE is never committed — only the ref name + in-container path. Add aiplatform.googleapis.com to the cloud host allowlist. docker-compose.infra.yml: bind the operator-supplied host token file read-only to /run/secrets/vertex_access_token on the main and effects runtimes (VERTEX_ACCESS_TOKEN_HOST_FILE, default /dev/null so lanes without Vertex still start; Gemini key path unaffected). Pass through BIFROST_VERTEX_GEMINI_ENDPOINT_URL (overlay supplies the complete Vertex OpenAI-compat URL) and GOOGLE_CLOUD_PROJECT/LOCATION (default empty). runtime-policy.env: regenerated from the contract via render_runtime_policy_env (test_runtime_policy_env_matches_contract_renderer proves contract<->env parity). test_runtime_policy_contract.py: update host-allowlist assertion for the additive Vertex host. Pre-existing platform-wide topic-parity-gate failures (25 unrelated topics) are identical with and without this change (proven by diff); that gate is pre-commit- only (not a CI merge gate) and this change adds zero new topic gaps, so that one hook is SKIP-ped. No deploy/receipt/merge gate is bypassed. * fix(OMN-12971): make Vertex runtime env contract-owned --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13008): automated .201 disk maintenance — worktree GC + docker GC + watermark alert ratchet (#1952) * feat(OMN-13008): automated .201 disk maintenance — worktree GC + docker GC + watermark alert ratchet Conservative, keep-list-driven disk maintenance to prevent the 2026-06-11 /data ~95% outage that killed all three lanes mid-demo: - scripts/worktree-gc.sh: drives the canonical omniclaude prune-worktrees.sh (merged+clean+pushed safety) on both Mac (merge-sweep tick) and .201 - scripts/disk-gc.sh + disk_gc_plan.py: conservative docker/builder/image GC. Pure, testable removal planner honoring a VERSIONED keep-list (deploy/disk-gc/keep-list.yaml): never reaps a kept repo, kept tag, in-use image, or anything younger than min_age_days; keeps N superseded generations. - scripts/disk-watermark-check.sh + disk_watermark_event.py: df watermark ratchet. >=85% emits a typed disk-watermark bus event (warning) that the sweep auto-ticket path turns into a Linear ticket; >=90% emits critical. Broker addr is fail-fast from KAFKA_BOOTSTRAP_SERVERS (no localhost default). - deploy/disk-gc/: systemd USER timer (onex-disk-gc.timer/.service, hourly) + install-disk-gc.sh. User units, NOT lane containers. - tests: 20 unit tests incl. GC plan-safety invariants + dry-run proof that default mode issues no destructive op (a wrong-delete GC is worse than none). Contract: contracts/OMN-13008.yaml * fix(OMN-13008): pass docker inventory to GC planner via stdin, not env (ARG_MAX) On a host with many docker images, passing the full image/ps inventory as env vars to disk_gc_plan.py exceeds ARG_MAX ('Argument list too long', exit 126), producing an empty plan. Write inventory to per-run scratch files (under the log dir, never /tmp; cleaned on exit) and hand it to the planner on stdin as a JSON envelope. Verified the failure live on .201; planner now reads stdin. * fix(OMN-13008): simplify GC plan stdin pipe (two processes, no nested subprocess) * fix(OMN-13008): keep-wins reconciliation — never remove an image id with any keep reason A single image id can surface in multiple 'docker image ls' rows (one per repo:tag). One tag could route the id to dangling-removal while another routes it to keep (e.g. tagged 'latest' or within-N-generations). Live .201 plan showed an id in BOTH remove_image_ids and kept_reasons. Reconcile at the end: keep wins — any id with a keep reason is dropped from the remove list; remove list deduped. Adds 2 regression tests. Verified live on .201. * fix(OMN-13008): timer uses OnCalendar=hourly for reliable re-arm (oneshot service) OnUnitActiveSec does not reliably re-elapse for a oneshot service once it goes inactive (observed NextElapseUSecMonotonic=infinity live on .201 after the first run). Switch to OnCalendar=hourly + RandomizedDelaySec so the timer re-arms every hour. Keep Persistent=true for missed-run catch-up. --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13005): materialize blocking event_consumer in runtime auto-wiring (runner consume-leg degenerate rows) (#1951) * fix(OMN-13005): materialize blocking event_consumer in auto-wiring (runner consume-leg degenerate rows) The runtime auto-wiring materialized event_publisher for handlers that declare it but had no equivalent for event_consumer. Request/response EFFECT handlers (HandlerContextRoiRunner) that publish a command then block on the correlated terminal event fell back to their no-op consumer default, returning None immediately -> every result row degenerate (failure_stage=generation, attempt_count=0) while generations succeeded ~1s later. Adds _make_sync_event_consumer (mirror of _make_sync_event_publisher), backed by service_terminal_event_consumer.make_terminal_event_consumer: a sync (topic, correlation_id, timeout) -> dict | None adapter that runs the proven direct-Kafka correlate-and-wait loop (from RuntimePatternBBroker) on an isolated event loop in a worker thread, so blocking does not deadlock the runtime dispatch loop that delivers the awaited terminal. TDD through the REAL dispatch path: test_event_consumer_injection drives a trial through _prepare_handler_wiring with a terminal arriving after a delay and asserts a non-degenerate row; verified RED with injection disabled. * fix(OMN-13012): two-phase (seek-now/wait-later) terminal event_consumer to close the subscribe-after-publish race (#1954) The OMN-13005 injected event_consumer is a single callable that does assign -> seek_to_end -> poll internally, all AFTER the handler has already published its command. Once OMN-13010 freed the dispatch loop and generation began completing in ~1s, the correlated terminal lands BEFORE the single-call consumer's post-publish seek_to_end positions, so seek_to_end skips PAST the already-emitted terminal and the runner times out on an offset beyond it (probe3, run_id=20260611T2140Z-probe3 -- both arms degenerate, zero rebalances). Splits positioning from waiting so the caller subscribes BEFORE it publishes: session = consumer.open(topic) # assign + seek_to_end NOW publisher(command_topic, payload) # publish AFTER positioning payload = session.wait(cid, timeout) # block from the captured position The returned TerminalEventConsumer is still directly callable with the legacy (topic, cid, timeout) -> dict | None single-call shape for any consumer that does not need subscribe-before-publish; the runner is the only consumer today. TerminalConsumerSession owns a dedicated event loop on a daemon worker thread for the whole open->wait->close lifecycle, preserving the OMN-13005 loop-isolation discipline so blocking never deadlocks the runtime dispatch loop. TDD (real dispatch path, RED-then-GREEN): extends the OMN-13005 injection test with a terminal emitted IMMEDIATELY after publish. The single-call (seek-after-publish) consumer MISSES it (degenerate row -- RED test asserts failure_stage=generation); the two-phase (open-before-publish) consumer CATCHES it (non-degenerate -- GREEN). The Kafka layer is faked at the two service seams against a shared in-memory log modeling seek-to-end semantics. OMN-13005 blocking-correlate behavior preserved. 269/269 auto_wiring unit tests pass; mypy --strict clean. Sibling to OMN-13010 / OMN-13005 / OMN-13003. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13005): route terminal consumer through Kafka boundary --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12990): pin worker replicas in ledgered config + fail-fast lane overrides (silent-drop ratchet) (#1948) * fix(OMN-12990): pin worker replicas in ledgered config + fail-fast lane overrides The base compose sets runtime-worker deploy replicas to ${WORKER_REPLICAS:-0} (soft-default ZERO). The stability lane's required state includes a running worker (GATE_ZERO_PROOF.md: 4 runtime containers), but the worker presence rode entirely on a compose soft-default — any plain compose up/recreate without the policy env silently scaled the worker to zero with no error and no signal. Fix (contract-native + fail-fast): - Add 'replicas' to ModelRuntimeProcessPolicy; pin replicas: 1 in every lane's worker block in runtime_policy.contract.yaml. - Renderer emits {PROFILE}_WORKER_REPLICAS into the ledgered runtime-policy.env for dev/stability-test/judge/prod. - stability + prod compose overrides reference ${..._WORKER_REPLICAS:?...} (fail-fast, NO silent :-1/:-0 default). A recreate that omits the policy env now aborts loudly instead of dropping the worker. prod previously had no override at all and inherited the dangerous :-0 default. Ratchet (recurrence guards): - tests asserting fail-fast override form (no soft default), contract-declared replica pin >= 1 per lane, and rendered {PROFILE}_WORKER_REPLICAS=1 in the ledgered env. - runbook deploy/verify procedure adds an expected-container census (worker must be present) via verify_container_manifest; a missing worker is a FAILURE, not silence. Note: SKIP=topic-parity-gate — that local-only advisory gate (absent from all .github/workflows, not a required CI check) fails on 25 pre-existing cross-repo topic gaps (build-loop/omniclaude/omniweb) identical on pristine base HEAD 8d7da1249; this change adds zero topics. All other hooks ran clean. Evidence-Ticket: OMN-12990 * test(OMN-12990): cover worker replica policy integration --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12909): add gateway bus forwarder P0A (#1946) * feat(OMN-12909): add gateway bus forwarder p0a * test(OMN-12909): add gateway forwarder integration coverage * fix(OMN-12909): satisfy gateway forwarder validators * test(OMN-12909): allow gateway forwarder bus protocol * fix(OMN-12909): sync gateway forwarder entry point * fix(OMN-12909): refresh runner image identity lock * test(OMN-12909): relax JSON normalizer mixed benchmark threshold * fix(OMN-12909): allow gateway handlers to boot unconfigured * fix(OMN-12909): declare gateway forwarder runtime profile * fix(OMN-12909): update runner identity backmerge expectation --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13011): LANE CENSUS RECONCILIATION ratchet — declared desired-state per lane, drift = auto-ticket (#1955) The class fix for the recurring lane-drift regression. Nothing reconciled the declared desired state of a runtime lane against what is actually running, so the same failure kept recurring with zero signal: volume config drift (OMN-12945), WORKER_REPLICAS silent zero (OMN-12988/12990), and on 2026-06-11 prod runtime containers plus the broker network were silently absent for hours during demo prep. Ships a per-lane DESIRED-STATE census: - (a) DECLARED in a versioned lane manifest (deploy/lane-census/lane-manifest.yaml): container set, network, replicas, image-tag pattern per lane (stability-test/prod/judge/dev), derived from the canonical compose lane files. A parity ratchet keeps the manifest locked in step with the compose files. - (b) RECONCILED on a schedule on .201 by SHARING the OMN-13008 systemd timer (a drop-in 4th ExecStart on onex-disk-gc.service — never a second timer) and on-demand via scripts/lane-census-check.sh / runtime_sweep. - (c) Drift = typed bus event (onex.evt.infra.lane-census-drift.v1) + Linear auto-ticket naming exactly what is missing/extra (container_absent, network_detached, replicas_zero, unexpected_container, oneshot_failed/stuck, image_tag_mismatch). Fail-fast, no warn-only mode (gates-block policy); exit 30 on drift; bus publish fail-fast on KAFKA_BOOTSTRAP_SERVERS (no localhost default). Red fixture reproduces 2026-06-11: prod runtime containers absent + broker network detached must produce the exact drift findings + a non-zero exit hours before a human noticed. Pure planner is fully unit-tested; shell driver dry-run-tested. Builds on the OMN-12988 deploy-agent RUNTIME census (deploy-time) as the complementary steady-state reconciler; closes the runtime-worker.yaml container_name: null census gap by sourcing names from the compose lane files. Evidence-Ticket: OMN-13011 Config-drift family: OMN-12945 Relates-to: OMN-13009, OMN-12988, OMN-13008 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13020): vendor missing node migrations — llm_routing 0000 + context_roi 001 (#1956) Vendors two omnimarket node-source migrations into the infra forward-migration tree via scripts/sync-node-migrations.sh (the canonical OMN-12559 mechanism): - node_projection_llm_routing/0000_create_llm_routing_decisions.sql (source: omnimarket #1168 / OMN-12942, merge ed6734f8) - node_projection_context_roi/001_create_context_roi_scores.sql (source: omnimarket #1178 / OMN-12955, merge 5010b1f4) Without the 0000 base table, node_projection_llm_routing/0001 (CREATE VIEW) hard-fails against NODE_POSTGRES_DB=omnidash_analytics — exactly the prod forward-migration exit-3 of 2026-06-11T09:35:52Z, and reproduced by construction in any clean clones@dev build. Files are byte-identical to the omnimarket dev blobs (sha256 f8a8b339… / c4126e65…) and to the untracked hot-patch copies on the .201 stability clone. Both migrations are self-contained, all-statements-IF-NOT-EXISTS, and 0000 sorts lexically before 0001 within the node's namespaced identity space. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13058): close TerminalConsumerSession on open() failure (worker thread + event loop leak) (#1957) TerminalConsumerSession.__init__ starts its dedicated worker loop thread immediately. TerminalEventConsumer.open did 'session = TerminalConsumerSession(...); return session.open()' with no cleanup: any failure inside session.open() (consumer start timeout, partition-assign timeout, broker auth error) propagated out of the raising expression, the session reference was lost, and the daemon worker thread plus its never-closed asyncio event loop leaked -- one pair per failed open. The motivating caller (HandlerContextRoiRunner) opens a session per trial, so a 160-560-trial battery against a degraded broker accumulates hundreds of leaked threads in the long-lived effects container. Fix: wrap session.open() in try/except BaseException -> session.close() (idempotent: stops the loop, joins the thread) -> re-raise. Covers both the two-phase .open(topic) path and the legacy single-call __call__ path. Found by the P3.3 doctrinal review of merged #1951 (b9712af9 / 36d98275). TDD: tests/unit/runtime/test_service_terminal_event_consumer_open_failure.py injects a real open failure through the production path (event bus without _bootstrap_servers) and asserts no alive terminal-consumer-* thread after the raise. Verified RED with the fix stashed (2 failed), GREEN with it (2 passed). Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13021): non-dev-base guard — fail feature-base PRs absent Stacked-Parent declaration (retro A-6) (#1958) Any PR whose base is neither dev nor main fails unless the body carries 'Stacked-Parent: #N'. Prevents the feedback_stacked_prs_orphan_from_dev class (#1185/#1954 auto-merged INTO parent feature branches and stranded off dev). base=main remains governed by main-target-guard. Epic OMN-13013 (process enforcement ratchets — June 12 retro). Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13014): hot-patch ledger rebuild preflight gate (retro B-1) (#1959) * feat(OMN-13014): hot-patch ledger rebuild preflight gate (retro B-1) Hot-patches on .201 (.prepatch sibling discipline) silently revert on any image rebuild/force-recreate — the 2026-06-11 20:58Z rebuild already erased a live /api/generate patch once. This adds the rebuild-path gate: - scripts/preflight_hotpatch_ledger.py: given a target container or lane + per-repo build refs, hard-fails when any hot-patch ledger row's source PR merge commit is not an ancestor of the build ref (git merge-base --is-ancestor), plus a .prepatch tripwire of the running container (unledgered .prepatch = hard fail; --post-rebuild = zero .prepatch expected). Sole bypass: HOTPATCH_PREFLIGHT_BYPASS carrying the Rule-10 '# skip-token-allowed: <user-approval-receipt-id>' form. - scripts/deploy-runtime.sh: guard_hotpatch_ledger wired into main() before build/preview (both dry-run and execute), lane derived from the compose project; skips loudly only when no ledger exists on the host. - tests/unit/scripts/test_preflight_hotpatch_ledger.py: 17 unit tests (ancestor gate, lane scoping, ledger loading, tripwire, bypass forms). - tests/ci/test_receipt_gate_install_guard.py: repair stale guard — core OMN-12565 replaced the OMN-9198 'uv pip uninstall first' install step with a cleared workspace venv (uv venv --clear); assert the new contract. Ledger backfilled from live census (5 .prepatch files / 4 source PRs, all MERGED to dev) at /data/omninode/hotpatch-ledger/ledger.yaml on .201. * fix(OMN-13014): scope missing-prepatch tripwire warning to the probed container --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12977): workspace-build sibling-pin ratchet — honor consuming lock, fail-fast on stale vendor (#1936) * fix(OMN-12977): workspace-build sibling-pin ratchet — honor consuming lock, fail-fast on stale vendor The workspace-mode image build (BUILD_SOURCE=workspace, used by the stability-test deploy procedure) vendored sibling/foundation packages from whatever the canonical OMNI_HOME clones happened to be checked out at, ignoring the consuming repo's uv.lock. On 2026-06-11 this shipped a 13-day-stale omnibase_infra 0.37.0-dev (pre-OMN-12501 Protocol-quarantine guard) + core 0.42.0 against an omnimarket dev lock pinning infra 0.38.1@e2dbdc95 / core 0.44.0@c97c2c9a, dropping the guard and crashing wire_from_manifest bootstrap fatally (stability lane down on demo day). Fix + recurrence ratchet (same PR): - scripts/runtime_build/check_sibling_lock_pins.py: parse the consuming repo's uv.lock for expected version+git-rev of each foundation/sibling package (scoped to the package's own source line so editable/registry pins are not cross-attributed a dependency's rev), resolve the actual clone version+HEAD, compare, and classify drift backward/forward/none. Fail-fast (exit 1) on any drift; --allow-drift records an explicit operator override in the artifact, never silent. - stage_workspace.sh: runs the preflight against the canonical clones before staging; aborts the build (exit 3) on unacknowledged drift and writes workspace/sibling-pin-comparison.json. - compute_workspace_provenance.py: folds the expected-vs-actual comparison into build-provenance.json so deploy verifiers can assert the build honored the lock; flags unacknowledged drift as a provenance error. - Dockerfile.runtime: COPY the comparison artifact (committed placeholder so the COPY always resolves; overwritten by stage_workspace.sh in workspace mode). - TDD: 19 unit tests covering lock parsing (git/registry/editable sources), drift classification, the exact 0.37.0-vs-0.38.1 stale case, check_pins exit codes, and the allow-drift override. Pre-existing mypy-strict bare-dict errors in compute_workspace_provenance.py fixed in the same pass. Evidence-Ticket: OMN-12977 Evidence-Source: pending-occ * ci(OMN-12977): retry runtime smoke compose port race * test(OMN-12977): align sibling-pin script tests with current API --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12989): workspace-mode image build must honor sibling lock pins (#1947) * fix(OMN-12989): workspace build must honor sibling lock pins (fail-fast ratchet) The 2026-06-11 stability bootstrap crash was caused by a workspace-mode --no-cache rebuild that vendored omnibase_infra 0.37.0-dev (13-day-stale worktree) even though omnimarket dev's uv.lock pins infra 0.38.1. The downgraded sibling predated the OMN-12501 Protocol-quarantine guard and turned a latent contract defect into a fatal crash. Fix + ratchet (same PR set): - scripts/runtime_build/resolve_workspace_pins.py: parse the consuming repo's uv.lock for sibling pins (version + git rev); classify each staged/installed sibling as exact/ahead/regression/unpinned; fail-fast (WorkspacePinError) on any regression BELOW the lock pin. Stdlib version-tuple fallback when packaging is absent so the ratchet never fails-open. - compute_workspace_provenance.py: enforce sibling pins + a host-infra self-check (installed omnibase_infra vs lock pin — the exact crash vector, since host infra is built from the context, not staged), and emit a pin_comparison block into build-provenance.json for deploy verifiers. - Dockerfile.runtime: COPY resolve_workspace_pins.py beside the provenance script so the in-image import resolves. - TDD: failing tests first (test_resolve_workspace_pins.py, 11 cases) + provenance integration tests; fixed a pre-existing stale RUNTIME_VERSION assertion to read pyproject dynamically. Evidence-Ticket: OMN-12989 * fix(OMN-12989): workspace build must honor sibling lock pins (fail-fast ratchet) The 2026-06-11 stability bootstrap crash was caused by a workspace-mode --no-cache rebuild that vendored omnibase_infra 0.37.0-dev (13-day-stale worktree) even though omnimarket dev's uv.lock pins infra 0.38.1. The downgraded sibling predated the OMN-12501 Protocol-quarantine guard and turned a latent contract defect into a fatal crash. Fix + ratchet (same PR set): - scripts/runtime_build/resolve_workspace_pins.py: parse the consuming repo's uv.lock for sibling pins (version + git rev); classify each staged/installed sibling as exact/ahead/regression/unpinned; fail-fast (WorkspacePinError) on any regression BELOW the lock pin. Stdlib version-tuple fallback when packaging is absent so the ratchet never fails-open. - compute_workspace_provenance.py: enforce sibling pins + a host-infra self-check (installed omnibase_infra vs lock pin — the exact crash vector, since host infra is built from the context, not staged), and emit a pin_comparison block into build-provenance.json for deploy verifiers. - Dockerfile.runtime: COPY resolve_workspace_pins.py beside the provenance script so the in-image import resolves. - TDD: failing tests first (test_resolve_workspace_pins.py, 11 cases) + provenance integration tests; fixed a pre-existing stale RUNTIME_VERSION assertion to read pyproject dynamically. Evidence-Ticket: OMN-12989 * test(OMN-12989): co-locate provenance pin helper in fixture --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13055): absent repos in REPOS list warn and exit 0 instead of failing (#1960) Missing repos (not cloned locally) now emit a WARN line and are tracked in a separate WARNED array. Only real fetch/ff failures cause exit 1. This makes pull-all.sh safe to use on machines with a partial clone set, while keeping the explicit-list override behavior intact. Adds three regression tests: absent-only exits 0, absent+present exits 0 with OK for the present repo, present-failed+absent exits 1. Evidence-Ticket: OMN-13055 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13070): config_prefetcher overlay wins on controlled lanes (#1961) On infisical_required=True lanes, fetched/overlay config now always wins over ambient env. Ambient env is retained only as a declared bootstrap fallback (with an explicit provenance INFO log line) when Infisical returns None. apply_to_environment also overwrites stale env on controlled lanes. Uncontrolled lane (infisical_required=False) behaviour is unchanged. Adds 5 regression tests: controlled-lane Infisical-wins, env-bootstrap- fallback, apply_to_environment overwrite, missing-from-both-is-error, and uncontrolled-lane-env-still-wins. Refactors _resolve_key to return a (outcome, value, error) tuple to satisfy the ≤5-param pattern gate. Source: docs/audits/2026-06-10-runtime-env-overlay-authority-audit.md Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13062): migration-gate vacuity fix — sentinel discipline, wait-for-postgres, skip-manifest, sync-check nonzero (retro A-10) (#1963) * fix(OMN-13062): migration-gate vacuity fix — sentinel discipline, wait-for-postgres, skip-manifest, sync-check nonzero Fixes three bugs identified in retro A-10 (recurrences OMN-12885, OMN-12934): (1) RUNNER SENTINEL DISCIPLINE run-forward-migrations.sh now clears migrations_complete=FALSE at the start of every run and sets it TRUE only as its FINAL act after all infra and node migrations succeed. Any mid-run failure leaves the gate UNHEALTHY. runner_completed_at is stamped at the same final step as durable evidence of a successful completion run. (2) SYNC-NODE-MIGRATIONS VACUOUS GATE sync-node-migrations.sh --check now exits 2 (not 0) when the omnimarket source tree is unresolvable. Silent exit-0 was hiding drift. The single opt-out is SYNC_NODE_MIGRATIONS_SKIP_UNRESOLVABLE=1 for environments that intentionally run without the source. (3) WAIT-FOR-POSTGRES GUARD run-forward-migrations.sh now waits up to PG_WAIT_RETRIES (default 30) x 2s for Postgres to accept connections before proceeding, guarding the first-boot initdb race. (4) SKIP-MANIFEST docker/migrations/skip-manifest.yaml introduced as the sole committed escape for intentionally-skipped migrations. The runner reads this at startup; listed migrations are recorded in schema_migrations with checksum "skip-manifest" without executing the SQL. (5) MIGRATION 085 Adds runner_completed_at TIMESTAMPTZ column to db_metadata so the runner's final stamp is durable in the schema (idempotent ADD COLUMN IF NOT EXISTS). Rollback included. 22 regression tests added covering all five fix surfaces. * fix(OMN-13062): stamp schema fingerprint for migration 085 Migration 085 (085_add_runner_completed_at_to_db_metadata.sql) was added in the initial commit but schema_fingerprint.sha256 was not regenerated. Running `python scripts/check_schema_fingerprint.py stamp` updates the artifact from the stale hash to match the 71 migration files. Evidence-Source: OCC#2563 Evidence-Ticket: OMN-13062 --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12864): Bifrost endpoints → committed overlay authority + fail-loud loader (#1964) * feat(OMN-12864, OMN-12814, OMN-12945): Bifrost endpoints → committed overlay authority + fail-loud loader OMN-12864 — Committed lane overlay - docker/lane-overlays/dev.bifrost.yaml: typed deployment bindings for all four BIFROST_LOCAL_*_ENDPOINT_URL values (coder :8000, reasoner :8001, embedding :8100, ds4-flash :8101). Previously only available as ephemeral shell exports on .201; now committed, auditable, diff-able, and CI-checked. - docker/lane-overlays/dev.bifrost.env: generated dotenv sidecar consumed by compose via env_file; never edited directly (yaml is authority). - docker/docker-compose.infra.yml: wire the env_file block at compose root so the four endpoints are injected into the interpolation context on a clean shell. Hardcode BIFROST_CONTRACT_PATH (remove :-/empty footgun — OMN-12814). - scripts/render_bifrost_lane_overlay_env.py: render script regenerates the env sidecar from the YAML source. - src/omnibase_infra/runtime/models/model_bifrost_lane_overlay.py: ModelBifrostLaneOverlay — typed Pydantic model enforcing URL completeness (OMN-12815: every URL must end in /chat/completions). OMN-12814 — Fail-loud loader - render_bifrost_delegation_contract: raises ProtocolConfigurationError on FileNotFoundError, YAMLError, ValidationError, and zero-endpoint renders. No lru_cache — every restart re-renders from packaged source so a stale cache cannot pin a broken result across deploys. OMN-12945 — Re-seed from packaged source on deploy - docker/entrypoint-runtime.sh: set BIFROST_FORCE_RESEED=1 on every container restart so the named-volume copy is always rebuilt from the packaged bifrost_delegation.yaml merged with committed lane-overlay endpoints. - render_bifrost_delegation_contract: honor BIFROST_FORCE_RESEED/force_reseed flag to bypass the stale-volume early-return path entirely. Tests: - tests/ci/test_bifrost_lane_overlay.py: CI gate — env sidecar in-sync with YAML source; all four BIFROST_LOCAL_* keys present. - tests/unit/runtime/models/test_model_bifrost_lane_overlay.py: bare-base URL rejection, env dict mapping, extra-field rejection. - tests/unit/runtime/test_render_bifrost_delegation_contract.py: fail-loud paths, force-reseed, zero-endpoint error, endpoint URL completeness. - tests/unit/models/test_model_serialization_roundtrip.py: roundtrip coverage. * fix(OMN-12864): move bifrost env_file to service level — fix compose schema validation failure Top-level 'env_file' is rejected by Docker Compose v2 schema validator ('additional properties not allowed'). This caused 10+ compose-render integration tests to fail in CI. Fix: - Remove top-level env_file block from docker-compose.infra.yml - Add per-service env_file on omninode-runtime, runtime-effects, runtime-worker (the three containers that render Bifrost) - Change BIFROST_LOCAL_*:? to BIFROST_LOCAL_*:- in x-runtime-env (compose-level validation removed; Python validates via ModelBifrostLaneOverlay + render_bifrost_delegation_contract) - Add two CI gate tests: compose_env_file_is_service_level_not_top_level and runtime_services_have_bifrost_env_file * fix(OMN-12864): allow BIFROST_LOCAL_* empty defaults in silent-fallback gate The overlay authority pattern (OMN-12864) passes BIFROST_LOCAL_*_ENDPOINT_URL via service-level env_file (docker/lane-overlays/dev.bifrost.env), not at compose config time. Compose-level :? would break CI rendering without the overlay pre-loaded. Validation at the Python layer (ModelBifrostLaneOverlay + render_bifrost_delegation_contract) is the enforcement point. Add the three failing vars to ALLOWED_EMPTY_DEFAULTS with OMN-12864 citation. --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12858/OMN-12879/OMN-12880): dispatcher route-coverage CI gate (#1965) * feat(OMN-12858/OMN-12879/OMN-12880): dispatcher route-coverage CI gate Static contract analysis: every subscribed command topic (onex.cmd.*) must declare handler_routing or runtime_dispatch, or the message goes to DLQ silently. This gate would have caught two recent incidents: 1. June 9 DLQ regression (OMN-12858 post-mortem): node_generation_consumer subscribed onex.cmd.omnimarket.node-generation-requested.v1 but a sole-handler revert left zero dispatcher routes registered. Messages went to DLQ silently with no CI signal. 2. June 12 DEL-01 live finding: onex.cmd.omnimarket.delegate-skill.v1 was consumed by dev lane bus but no dispatcher route existed in any deployed contract. Discovered via manual rpk consumer-group lag probe (DEL-01 evidence, docs/evidence/2026-06-12-weekend-pass/). Deliverables: - scripts/check_dispatcher_route_coverage.py — static YAML scanner that checks both omnibase_infra and omnimarket contract trees; ratchet allowlist for known pre-existing violations; --changed-contracts mode (OMN-12879) for per-PR scoping; compat publish topics excluded (OMN-12880) - .github/workflows/dispatcher-route-coverage.yml — CI workflow that checks out omnimarket sibling, collects changed contract paths in PR mode, and runs the gate; fires on PR, push-to-main, and merge_group - tests/ci/test_dispatcher_route_coverage_gate.py — 12 unit tests covering RED/GREEN/COMPAT/CHANGED-MODE/ALLOWLIST/MULTI-DIR paths plus live-contract regression proof against the actual omnibase_infra tree Allowlist additions: - onex.cmd.omnibase-infra.pattern-b-dispatch.v1 (RuntimePatternBBroker, imperative consumer, OMN-12525 migration target) - onex.cmd.platform.contract-resolve-requested.v1 (transitional HTTP bridge node_contract_resolver_bridge OMN-2756, metadata.transitional=true) [OMN-12858, OMN-12879, OMN-12880] * fix(OMN-12858): drop full uv sync from dispatcher-route-coverage workflow Gate script only needs pyyaml (stdlib + yaml). Using full setup-python-uv was causing 10+ minute timeout. Replace with direct pip install pyyaml and invoke python3 directly. Reduces job from 10m timeout to <1m. [OMN-12858] --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13026): wire transport-mock-lint pre-commit hook + CI gate (#1962) * feat(OMN-13026): wire transport-mock-lint pre-commit hook + CI gate (omnibase_infra) Activates the transport-mock-lint validator (from omnibase_core, OMN-13026) on omnibase_infra. Ratchet baseline: 218 existing violations across 63 files frozen in validation/transport_mock_baseline.yaml. New bare AsyncMock/ MagicMock on EventBus/transport surfaces are blocked by pre-commit hook and CI lint step. Existing violations tracked for drain by per-site tickets (parent OMN-13026). Reference incident: PR #1181 bare AsyncMock hid missing EventBusKafka.stop(). Evidence-Ticket: OMN-13026 * fix(OMN-13026): use uv run python for CI lint step + bump omnibase_core pin to include transport_mock_lint The transport-mock lint CI step previously cloned omnibase_core and ran `python -m omnibase_core.validators.transport_mock_lint` with PYTHONPATH, but this failed: `No module named omnibase_core.validators.transport_mock_lint` because it ran `.venv/bin/python` which uses the locked venv, and the venv omnibase_core pin (2defabef4) predates the transport_mock_lint module. Fix: use `uv run python` (removes the clone step) and bump omnibase-core git pin from 2defabef4 to 309d89fa7 (PR 1231 merge commit on dev) so transport_mock_lint is available in the locked venv. * fix(OMN-13026): align transport mock baseline and runner lock * fix(OMN-13026): sync omnibase_core pin + runner identity lock to dev baseline Align pyproject.toml omnibase_core rev to 2defabef (required by test_release_backmerge_preserves_proven_runtime_core_pin) and update docker/runners/runner-image.lock.json identity_digest/shared_env_digest to match dev runner image lock (79b08f44 / 90c8b3b9). Both were stale from the prior session's pin bump that used an older SHA. * fix(OMN-13026): source transport mock validator from core * fix(OMN-13026): source transport validator from core dev --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13094): --output receipt mode on onex node/run — quiet typed receipts with durable capture (#1966) Phase 2a of the skill-output-suppression slice (epic OMN-13089, plan docs/plans/2026-06-12-skill-output-suppression-plan.md Phase 2 item 1). - onex node/onex run gain --output receipt: ALL runtime logging routes to a run_id-suffixed capture file under <state-root>/captures/ (no console handlers — kills the 25-50-line RuntimeLocal INFO stream at the source); stdout carries exactly ONE typed ModelSkillResult JSON with the FULL handler result (result_model = concrete handler result type FQN). - Durable capture: capture log + handler result content-addressed via omnibase_core ArtifactStore (OMN-13093); artifact.captured + tool.output.captured emitted to the emit daemon socket (--emit-socket, default ~/.claude/emit.sock). - Failure asymmetry: artifact write failure => FULL output printed, no receipt (no hidden loss); emission failure => receipt still prints, event spooled to <state-root>/emit_spool/ for replay. - Node failure => status=failed/error with full error + capture log INLINE in the receipt (errors are never hidden) and artifact-backed. - Default output mode unchanged (enforcement is Phase 4). - RuntimeLocal exposes handler_result (receipt schema identity). - core pin 2defabef -> ae8793bd (merged OMN-13091/13093 receipt models + ArtifactStore); runner-image identity lock regenerated and the OMN-12765 backmerge identity constants updated for the new pin. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13097): onex skill subcommand + declarative skill->node mapping (#1968) * feat(OMN-13097): onex skill subcommand + declarative skill->node mapping Phase 4a of the skill-output-suppression slice (epic OMN-13089, plan docs/plans/2026-06-12-skill-output-suppression-plan.md Phase 4 item 1/2). A dispatch skill IS one CLI call (user directive 2026-06-12). This adds the `onex skill <name> [args]` dispatch surface that the 24 omniclaude shim migrations build on: - onex.cli entry-point `skill` -> cli_skill.run_skill_by_name. Resolves the skill via the declarative skill_mapping.yaml registry, builds the backing node's input payload from the skill's CLI args, writes it under the state root (.onex_state/tmp/<skill>-<run_id>.json — never /tmp), resolves the node's packaged contract exactly like `onex node`, and dispatches through the proven receipt-mode path (run_receipt_mode, OMN-13094). stdout is exactly one typed ModelSkillResult JSON with the FULL handler result. - skill_mapping.yaml: declarative DATA mapping all 24 dispatch shims to their backing onex.nodes node + typed result-model FQN (verified against each live handler handle() return type on origin/dev) + per-arg payload specs + static payload + keyword classifiers (delegate task_type as data, not code). Adding a skill is a YAML edit + fixture, never a CLI code change (ticket deliverable 2/3). Mapping lives beside the node-resolution surface, never hardcoded branching in the CLI. - Typed models split one-per-file (repo convention): ModelSkillArgSpec, ModelSkillClassifier, ModelSkillMapping, ModelSkillMappingRegistry, EnumSkillArgType. Frozen, extra=forbid, fail-fast coercion/validation. - validation_exemptions.yaml: Click-callback param-count + literal-identifier name-field exemptions mirroring the existing cli_node run_node_by_name precedent (OMN-11570) — same pattern, same rationale. dod_evidence: - 20 unit tests pass (registry validity, all-24-shims coverage, FQN result models, arg parsing/coercion/positional/required, classifiers, payload build, receipt-mode dispatch wiring, payload-under-state-root not /tmp). - uv run mypy src/ --strict: clean (2438 source files). - ruff format + check: clean. pre-commit run on changed files: pass. - runner-image identity lock regenerated for the pyproject entry-point add (same as OMN-13094). * test(OMN-13097): rebind OMN-12765 backmerge identity constants for onex skill pyproject change Adding the `skill` onex.cli entry-point to pyproject.toml changes the runner-image identity_digest (and shared_env_digest) the lock binds. Update the hardcoded expected constants in the backmerge-identity test to the regenerated values — same mechanical rebind OMN-13094 performed for the core pin bump. Identity + runner-image-identity tests pass (14). --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix+test(OMN-13012): force terminal-topic metadata refresh so both ephemeral consumers assign (runner consume-leg wedge) (#1969) The runner consume-leg wedged on the live stability battery (image c0505521f1fa, EXP1-3_RUNNER_CONSUME_LEG_BLOCKER): only the FAILED terminal topic ever subscribed while the COMPLETED topic never assigned, so the correlated completed terminal was never read and the 8x2x10 matrix re-fired cell 1 forever, emitting zero non-degenerate rows. A prior two-strike diagnosis proved the omnimarket handler is correct (it opens both terminal sessions pre-publish and races them); the defect is in the omnibase_infra runtime consume leg. Root cause: _assign_direct_terminal_partitions ignored the metadata future returned by AIOKafkaClient.set_topics and re-called set_topics([same_topic]) each loop iteration. aiokafka 0.13.0 set_topics only forces a metadata refresh when the topic set DIFFERS from the tracked set, so every iteration after the first took the no-op branch and never re-fetched. An ephemeral group_id=None consumer whose first metadata fetch had not yet surfaced partitions burned the full 30s assign cap and raised a bare TimeoutError (the empty-message 'wait failed' seen live). Fix: register the reply topic once and await that metadata fetch, then on each miss force a fresh fetch via force_metadata_update (which always fetches) rather than the no-op set_topics repeat. The assign-cap TimeoutError now carries a diagnostic message instead of an empty one. Test: tests/integration/test_terminal_consumer_concurrent_assign_race.py drives the REAL TerminalEventConsumer (the object wired as event_consumer) through the REAL open_direct_terminal_consumer/poll path with AIOKafkaConsumer monkeypatched to a fake that faithfully models aiokafka 0.13.0 set_topics future + metadata-latency semantics. RED before the fix (bare TimeoutError, the live empty-message signature); GREEN after. K>=2 multi-trial variant asserts no worker-thread leak across trials. Evidence-Source: <occ-sha-pending> Evidence-Ticket: OMN-13012 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13096): onex delegate single-command subcommand (Phase 2b) (#1967) * feat(OMN-13096): onex delegate single-command subcommand Add 'onex delegate "<prompt>" [--task-type X] [--max-tokens N]' as a subcommand on the existing onex CLI (Phase 2b of the skill-output-suppression slice, OMN-13089). The command wraps payload construction, node dispatch, and result extraction internally and prints exactly one ModelSkillResult[ModelDelegateSkillResponse] to stdout via the OMN-13094 receipt-mode path. RuntimeLocal logs go to the capture file + artifact store, never to stdout; scratch payloads live under <state-root>/tmp/ with run_id suffixes (never /tmp). - cli_delegate.py: classify_task_type (keyword table from legacy skill md), payload write, contract resolve, run_receipt_mode dispatch - register 'delegate' under onex.cli entry points - exempt delegate_command from the >5-param patterns gate (same Click-callback rationale as run_node_by_name) - 20 unit tests: classification, scratch-under-state-root, single typed receipt on stdout, zero INFO log leakage omnibase_infra does NOT depend on omnimarket; the delegate node is resolved at runtime via the onex.nodes entry-point group (registered by omnimarket). * chore(OMN-13096): re-trigger deploy-gate after Evidence-Source set to OCC#2593 No code change — the deploy-gate workflow triggers on synchronize (not edited), so the PR-body Evidence-Source fix needs a new commit to re-resolve the OCC ref to the open PR head where contracts/OMN-13096.yaml (with deploy evidence) lives. * chore(OMN-13096): re-trigger deploy-gate now that OCC#2593 merged to OCC dev contracts/OMN-13096.yaml (with the dod-deploy-onex-delegate item) is now on OCC dev, so the deploy-gate OCC-dev checkout resolves the contract + deploy evidence. * chore(OMN-13096): regenerate runner-image identity lock for pyproject entry-point add Adding the 'delegate' onex.cli entry point to pyproject.toml changed the dependency-manifest digest that scripts/ci/runner_image_identity.py folds into the runner-image identity lock. Regenerate the lock so tests/ci/test_runner_image_identity.py matches (was the only CI test failure; unrelated environmental integration/perf failures excluded). * test(OMN-13096): update backmerge identity assertions to regenerated lock digests The runner-image identity lock was regenerated for the pyproject entry-point add; this test hardcodes the expected identity_digest/shared_env_digest, so update both to match the new lock (same maintenance OMN-13094 did). --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): tolerate partition-less reply topic in terminal consume leg (#1970) The context-ROI runner opens one ephemeral group_id=None terminal consumer per terminal topic BEFORE publishing each generation command (subscribe-before-publish, OMN-13012/13038). The FAILED reply topic is only produced to on contract_passed=False; in a battery where generations pass it has zero messages, so Redpanda never advertises a partition for it. _assign_direct_terminal_partitions burned the full 30s assign cap on every trial then raised a bare TimeoutError (the empty-message 'wait failed'), stalling each of the 160 battery trials ~30s before the COMPLETED terminal could correlate -> battery needs >80 min and never completes (verifier-confirmed wedge). A partition-less reply topic is a valid steady state, not a 30s error: - _assign_direct_terminal_partitions gives a bounded grace window for a topic that exists but is slow to surface metadata, then assigns whatever partitions exist (possibly none) and returns promptly instead of burning the cap and raising. - poll_direct_terminal_consumer treats an empty assignment as 'no terminal will arrive here' (sleeps out its timeout, returns None) without calling getone() on an unassigned consumer. Repro: tests/integration/test_terminal_consumer_battery_load_wedge.py drives the REAL TerminalEventConsumer over K=10 x 2 cells x 2 arms; RED (40/40 trials block a full assign cap) before the fix, GREEN after. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): pin terminal-consumer read offset synchronously (close lazy seek_to_end publish race) (#1971) * fix(OMN-13118): pin terminal-consumer read offset synchronously to close lazy seek_to_end publish race The consume-leg wedge survived PR #1969 (set_topics no-op) and PR #1970 (partition-less assign-cap) because both addressed the assign phase, not the seek timing. AIOKafkaConsumer.seek_to_end is LAZY: it requests a LATEST offset reset that only resolves on the first poll — AFTER the caller publishes. With generation completing in ~1s, the correlated COMPLETED terminal lands in the open->poll gap, so the lazily-resolved LATEST position is the HWM AFTER the record and the poll reads past it. The terminal is never read, the trial never correlates, and the experiment matrix re-fires the same cell forever. Replace seek_to_end with a synchronous end_offsets() + seek() pin (_pin_direct_terminal_end_offsets) in both open_direct_terminal_consumer and RuntimePatternBBroker._dispatch_and_wait_with_direct_kafka_consumer, so the read position is fixed at open() time, before the publish — the real subscribe-before-publish guarantee. Empty assignment (partition-less reply topic, OMN-13118 #1970) is a no-op. Repro: tests/integration/test_terminal_consumer_seek_reset_race.py drives the real TerminalEventConsumer.open()/wait() the way HandlerContextRoiRunner does, publishing the correlated terminal in the open->wait gap across K=10 x 2 cells x 2 arms; RED with the lazy reset (every cell degenerate), GREEN once the read offset is pinned. RED verified by git-stashing only the source fix. Existing consume-leg fakes updated to model end_offsets/seek (they previously masked the bug by making seek_to_end a synchronous exact snapshot). * test(OMN-13118): reword assertion (lazy not deferred) for receipt honesty gate * fix(OMN-13118): bound end_offsets() round-trip with assign-cap timeout (CodeRabbit) end_offsets() is a broker ListOffsets round-trip aiokafka documents as able to block indefinitely. Bound it with the same cap as start()/assign so a stalled broker fails fast instead of hanging the pre-publish positioning. --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13137): _validate_routing strategy-aware event_model checks (#1973) operation_match routes by the `operation` field and does not use event_model. The validator was unconditionally requiring event_model.{name,module} for every handler entry, causing 230/295 omnimarket operation_match contracts (all correct as authored) to fail routing validation at startup. Fix: read routing_strategy from the routing map and branch validation: - payload_type_match → require event_model.{name, module} (unchanged) - operation_match (and any non-payload strategy) → require `operation`; skip event_model checks entirely Updated pre-existing _validate_routing tests to declare routing_strategy: payload_type_match explicitly (they always tested payload_type_match semantics but relied on the implicit fallback that is now removed). Added test_validate_routing_operation_match.py with 4 unit tests: 1. operation_match without event_model → zero event_model errors 2. operation_match missing operation field → error 3. payload_type_match missing event_model → still errors (regression guard) 4. Real node_integration_sweep_orchestrator routing block → clean (boot gate) Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): independent per-terminal-topic consumers in Pattern B direct-Kafka wait (#1972) The consume-leg wedge survived four merged fixes (#1969 set_topics no-op, #1970 partition-less assign-cap, #1971 synchronous seek-pin). The STRONG K>=10 multi-cell reprobe on the stability lane still wedged on REBUILD-5 (cdf53d963f7b). Converged diagnosis (strikes 3+4, docs/evidence/2026-06-12-weekend-pass/experiments/probe4-stability/ reprobe-K10-rebuild5/HALT_K10_WEDGE_PERSISTS.md): the runtime waited for each trial's terminal across TWO topics (node-generation-completed.v1 + node-generation-failed.v1) with a SINGLE ephemeral group_id=None consumer assigned both topics' partitions. One aiokafka consumer holds one manual subscription; the COMPLETED delivery window collapsed before it surfaced the correlated record, so the trial never correlated and the matrix re-fired cell 1. RuntimePatternBBroker._dispatch_and_wait_with_direct_kafka_consumer now opens ONE independent AIOKafkaConsumer PER terminal topic via open_direct_terminal_consumer (each started, assigned, and offset-pinned via end_offsets()+seek() at open() BEFORE publish), then awaits both CONCURRENTLY via asyncio.wait(FIRST_COMPLETED). The first correlated terminal wins; both are torn down. No shared consumer, no subscription flip. Keeps the #1970 partition-less no-op and the #1971 synchronous offset pin (both live in open_direct_terminal_consumer / poll_direct_terminal_consumer). Removes the now-dead single-consumer helpers (_assign_terminal_topic_partitions, method-level _refresh_terminal_topic_metadata, _direct_kafka_* kwargs builders, _kafka_bootstrap_servers/_kafka_event_bus). Adds tests/integration/test_terminal_consumer_subscription_flip_wedge.py: a real-dispatch-path K>=10 x 2-cell x 2-arm repro whose fake models TWO independent consumers honestly (delivery is faithful only for a single-topic assignment; a consumer spanning both topics flips and drops the COMPLETED record). RED genuineness verified by reverting only the source to the single-consumer shape (test hangs past timeout); GREEN with the fix in 2.3s. Acceptance is the LIVE K>=10 multi-cell stability-lane reprobe (later phase), NOT this unit test. A green unit repro is necessary but NOT sufficient. Refs OMN-13118, OMN-13128. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): long-lived terminal correlator replaces per-trial ephemeral consume leg (#1974) Tier B canonical redesign (epic OMN-12525). Five offset/subscription patches (#1969-#1972) tuned the per-trial-ephemeral terminal consumer and all fa…
…push unblocked) (#2166) * docs(OMN-12962): contract-store durability audit — cold-runtime census proof (#1937) Proves cold-start contract census reconstructs from the image-bundled filesystem manifest (HYBRID-mode bootstrap + PluginLoaderContractSource), independent of node-registration.v1 retention. The delete-retention topic feeds only the post-freeze dynamic listener (auto_offset_reset=latest, no history replay). Live .201 stability-test evidence: filesystem contract_path in manifest + truncated topic log head with intact census. No store fix needed; residual dynamic-only gap covered by runtime_sweep sweep check. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12970): vendor omnimarket projection node migrations into forward/nodes (#1938) Vendors omnimarket node-owned projection migrations into the namespaced forward-migration tree so run-forward-migrations.sh materializes them in the dashboard projection DB (omnidash_analytics) at deploy. Primary (OMN-12970): creates llm_call_metrics, llm_cost_aggregates, and capability_scores in the projection DB. These were only ever created in the omnibase_infra DB by infra migrations 031/060, so the ab-compare, cost.token_usage, cost.summary, and capability-scores projection topics were DEGRADED at startup ('table not found') and their dashboard panels rendered empty. Also re-syncs three omnimarket node migrations the vendor tree had drifted from (node_projection_llm_routing, node_projection_overnight, node_projection_savings /077) — sync-node-migrations.sh --check requires the full vendored tree to match omnimarket source, and these were missing. Companion to omnimarket PR for the same ticket (source migrations + projection table-coverage ratchet test). Evidence-Ticket: OMN-12970 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12965): stamp runtime image identity (version + revision) in workspace builds (#1943) * fix(OMN-12965): stamp runtime image identity (version + revision) in workspace builds The main runtime image stamped org.opencontainers.image.version=0.1.0 with a blank org.opencontainers.image.revision after workspace rebuilds. A blank identity degrades every proof packet (runtime SHA + image digest are required citations in accepted evidence). Root causes (three build paths under-stamped identity): - onex up --build (cmd_up) passed only GIT_SHA; the runtime-stage OCI labels read VCS_REF (-> blank revision) and RUNTIME_VERSION (-> placeholder 0.1.0). - deploy-runtime.sh passed VCS_REF but not RUNTIME_VERSION/GIT_SHA. - Dockerfile silently allowed blank/placeholder identity in workspace mode. Fix: - cli._image_identity_build_args() stamps the full quad (GIT_SHA/VCS_REF/ RUNTIME_VERSION/BUILD_DATE) and fails fast on an unresolved git revision. - deploy-runtime.sh stamps RUNTIME_VERSION + GIT_SHA and verifies the version label is non-placeholder post-deploy. - Dockerfile.runtime fails workspace builds with blank VCS_REF or placeholder RUNTIME_VERSION=0.1.0 (release mode unaffected). Enforcement ratchet (same PR): - scripts/check_runtime_image_identity.py static check, wired as pre-commit hook + CI gate (ci.yml). - tests/unit/infra/test_runtime_image_identity_labels.py pins the cli helpers + Dockerfile guard; deploy-agent test extended for the quad. Proven locally via throwaway docker builds: workspace+args -> populated labels; workspace without args -> guard fails (exit 64); release without args -> 0.1.0 placeholder allowed (no regression). Evidence-Ticket: OMN-12965 * test(OMN-12965): integration build proof for runtime image identity labels Builds the real runtime-stage ARG/LABEL/guard block against busybox and asserts via docker inspect: workspace+args -> populated version/revision; workspace without args -> guard fails (exit 64); release -> placeholder allowed. Satisfies the integration-test hard gate and makes the throwaway proof permanent. Evidence-Ticket: OMN-12965 --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12987): workspace-build sibling lock-pin preflight + provenance (#1944) Recurrence ratchet for the 2026-06-11 stability bootstrap crash. The 11:20Z --no-cache rebuild vendored omnibase_infra 0.37.0-dev (~2c1d672f) + core 0.42.0 even though omnimarket dev's uv.lock pinned infra 0.38.1 @ e2dbdc95 + core 0.44.0 @ c97c2c9a. The stale sibling predated the OMN-12501 Protocol-quarantine guard, turning a latent contract defect into a fatal crash that crash-looped the main runtime. - check_sibling_lock_pins.py: host-side fail-fast preflight resolving expected sibling versions/SHAs from the consuming repo's (omnimarket) uv.lock and comparing them against each vendored tree. Mismatch aborts the build. - stage_workspace.sh: emit a .build-sha marker per staged sibling (rsync drops .git) so the preflight and provenance can identify the vendored commit. - deploy-runtime.sh: run the preflight after staging, before build; abort on mismatch. Write the comparison under sibling-repos/.sibling-lock-pins.json. - compute_workspace_provenance.py + Dockerfile.runtime: fold expected-vs-actual lock_pin_comparison into build-provenance.json for deploy verifiers. Recurrence-guard tests prove a stale infra 0.37.0 vs lock-pinned 0.38.1 fails the preflight and matched pins pass; deploy script wiring is asserted statically. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12988): pin stability runtime-worker to replicas 1 + census ratchet (#1942) The base docker-compose.infra.yml defaults runtime-worker to replicas 0 (${WORKER_REPLICAS:-0}). The stability-test lane's required state includes a running worker (4-container census: main, effects, worker, projection-api), but the override pinned it via an env-interpolation default (${STABILITY_TEST_WORKER_REPLICAS:-1}) — a silent-drop surface: a stray exported STABILITY_TEST_WORKER_REPLICAS=0 or removal of the :-1 fallback would scale the worker to 0 with zero signal on a plain compose up/recreate. Fix: pin docker-compose.stability-test.yml runtime-worker deploy.replicas to the literal 1 (no env indirection). Ratchet (recurrence guards, same PR): - scripts/deploy-agent/tests/unit/test_runtime_worker_census.py: assert runtime-worker stays in the deploy-agent RUNTIME-scope census so a missing worker (replicas 0 => absent from docker compose ps) is a deploy failure, not silence; assert the override pins a literal 1. - tests/integration/infra/test_stability_test_runtime_compose_render.py: assert the rendered stability worker resolves deploy.replicas == 1. - tests/unit/infra/test_stability_test_runtime_lane.py: update the existing pin assertion to the literal 1. Evidence-Ticket: OMN-12988 Config-drift family: OMN-12945 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12979): expire-bound topic completeness suppressions (#1940) Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12966): accept community-edition Infisical /api/status in provision readiness gate (#1939) The fresh-provision readiness gate in provision-infisical.py only accepted the enterprise {"status": "ok"} payload and rejected the community edition's {"message": "Ok"}, returning 1 before bootstrap could run. This blocked provisioning against the Infisical instance deployed on .201 (community edition). Route the gate through the existing _is_infisical_ready helper (single source of truth, already used by the already-provisioned path). Add TestMainFreshProvision- ReadinessGate covering community/enterprise/not-ready cases. Also adds docker/docker-compose.infisical-stability.yml: an ADDITIVE Infisical compose project for the stability-test lane (joins the existing network as external, reuses stability postgres/valkey, no lane mutation), since the lane overlays disable the in-lane Infisical service via *-disabled profile overrides. P1.2b-A: Infisical now reachable from the stability runtime/effect containers; a known secret (OMN_12966_PROBE) seeds and resolves end-to-end via the machine- identity universal-auth path, verified from inside the effects container. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12958): volume-config drift gate + runtime config provenance (#1941) * feat(OMN-12958): volume-config drift gate + runtime provenance Compute config provenance (path + sha256) for the runtime-rendered Bifrost delegation contract; the deployed volume copy survives rebuilds and silently diverges from packaged source (two competing authorities, OMN-12945). - runtime/config_provenance.py: ModelConfigProvenance + drift classification, sidecar JSON writer (read by sweep + proof packets) - runtime/health/health_config_provenance.py: drift -> degraded health - render entrypoint logs provenance line + writes sidecar on every boot - docs/runbooks/volume-config-drift-and-reseed.md: ledgered re-seed procedure - validation exemption for config_name (logical identifier, not entity ref) No live volume mutation: re-seed is an operator deploy step (deploy_pending). * test(OMN-12958): cover volume config drift reseed flow --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12957): wire runtime-profile validator + core registry parity guard (#1945) * feat(OMN-12957): wire runtime-profile validator + core registry parity guard - Remove the import-time RuntimeError drift raise in runtime_profile.py (a hard raise on core/infra version skew would crash the kernel at import); the parity invariant is enforced by test_profiles_match_core_registry instead. - Add tests: _PROFILES keys == omnibase_core REGISTERED_RUNTIME_PROFILES, and every CONSUMER_ATTACHED_RUNTIME_PROFILES profile loads as a real lane. - Wire omnibase_core.validation.validator_runtime_profiles as a pre-commit hook + validator-runtime-profiles.yml CI gate on infra contracts. - Freeze 19 pre-existing violators in validation/runtime_profiles_allowlist.yaml (discovered by repo-root walk; drain via OMN-12982). Blocks NEW orphans. Requires the omnibase_core pin to include OMN-12957's validator (new rules). Evidence-Ticket: OMN-12957 Evidence-Source: 5463fbaf819409d4fb7f491dd4f276f10d869eba * ci(OMN-12957): pass runtime profile allowlist to validator * test(OMN-12957): cover runtime profile registry parity * fix(OMN-12957): keep runtime profile allowlist under config * fix(OMN-12957): pin core runtime profile registry --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12973): widen migration-gate healthcheck start_period + ratchet (#1950) P2.8: classify the prod migration-gate UNHEALTHY-then-self-resolved incident. Classification: idle-one-shot-mis-modeled = NO. The migration-gate is a correctly-modeled long-running sentinel (entrypoint `while true; sleep 3600`) whose healthcheck continuously polls db_metadata.migrations_complete via check_migrations_complete.sh. The container flipped UNHEALTHY transiently because its healthcheck start_period (10s) was far shorter than the real cold-volume migration window (~116s: prod gate started 09:35:22, intelligence-migration finished 09:37:18). Past the 10s grace window the still-failing probe was reported UNHEALTHY until migrations completed, then self-resolved — no fault. Fix: raise migration-gate healthcheck start_period 10s -> 180s in both the authoritative catalog manifest (docker/catalog/services/migration-gate.yaml, flows into the generated compose) and the hand-maintained docker/docker-compose.infra.yml that deploy-runtime.sh applies to .201, so a still-applying gate stays in `health: starting` instead of flipping UNHEALTHY. Ratchet (enforcement, not detection): new ValidatorHealthcheckStartPeriod in omnibase_infra catalog (validator_healthcheck_start_period.py) asserts a 120s floor for migration-completion gates, wired into `onex validate runtime` (cmd_validate_runtime) AND backed by unit tests that gate every PR via pre-commit + CI. A migration-completion gate is identified by BOTH a healthcheck that polls migration completion AND a service_completed_successfully dependency, so ordinary app services (e.g. intelligence-api, an HTTP liveness probe with a legitimately short 40s start_period) are not swept into the floor. Prod probed read-only only; no prod mutation. Applies to live prod via the batched stability/prod rebuild (deploy_pending). Evidence-Ticket: OMN-12973 Evidence-Source: OCC#PENDING Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12971): deliver Vertex ADC bearer token to runtime effects container (secret-ref discipline) (#1949) * feat(OMN-12971): deliver Vertex ADC bearer token to runtime effects container (secret-ref discipline) Wire the Vertex $500-credit ADC path into the runtime, ADDITIVE next to the Gemini API-key path (provider-agnostic; neither provider removed or forced). runtime_policy.contract.yaml: add a secret-source mapping llm.vertex.access_token (source_type=file, /run/secrets/vertex_access_token) to the dev/stability-test/ judge profiles, alongside the existing llm.gemini.api_key env mapping. The ref name MUST match cloud-vertex-gemini.secret_ref in omnimarket bifrost_delegation.yaml. The resolved VALUE is a short-lived OAuth bearer token minted from ADC, refreshed by the operator; the token VALUE is never committed — only the ref name + in-container path. Add aiplatform.googleapis.com to the cloud host allowlist. docker-compose.infra.yml: bind the operator-supplied host token file read-only to /run/secrets/vertex_access_token on the main and effects runtimes (VERTEX_ACCESS_TOKEN_HOST_FILE, default /dev/null so lanes without Vertex still start; Gemini key path unaffected). Pass through BIFROST_VERTEX_GEMINI_ENDPOINT_URL (overlay supplies the complete Vertex OpenAI-compat URL) and GOOGLE_CLOUD_PROJECT/LOCATION (default empty). runtime-policy.env: regenerated from the contract via render_runtime_policy_env (test_runtime_policy_env_matches_contract_renderer proves contract<->env parity). test_runtime_policy_contract.py: update host-allowlist assertion for the additive Vertex host. Pre-existing platform-wide topic-parity-gate failures (25 unrelated topics) are identical with and without this change (proven by diff); that gate is pre-commit- only (not a CI merge gate) and this change adds zero new topic gaps, so that one hook is SKIP-ped. No deploy/receipt/merge gate is bypassed. * fix(OMN-12971): make Vertex runtime env contract-owned --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13008): automated .201 disk maintenance — worktree GC + docker GC + watermark alert ratchet (#1952) * feat(OMN-13008): automated .201 disk maintenance — worktree GC + docker GC + watermark alert ratchet Conservative, keep-list-driven disk maintenance to prevent the 2026-06-11 /data ~95% outage that killed all three lanes mid-demo: - scripts/worktree-gc.sh: drives the canonical omniclaude prune-worktrees.sh (merged+clean+pushed safety) on both Mac (merge-sweep tick) and .201 - scripts/disk-gc.sh + disk_gc_plan.py: conservative docker/builder/image GC. Pure, testable removal planner honoring a VERSIONED keep-list (deploy/disk-gc/keep-list.yaml): never reaps a kept repo, kept tag, in-use image, or anything younger than min_age_days; keeps N superseded generations. - scripts/disk-watermark-check.sh + disk_watermark_event.py: df watermark ratchet. >=85% emits a typed disk-watermark bus event (warning) that the sweep auto-ticket path turns into a Linear ticket; >=90% emits critical. Broker addr is fail-fast from KAFKA_BOOTSTRAP_SERVERS (no localhost default). - deploy/disk-gc/: systemd USER timer (onex-disk-gc.timer/.service, hourly) + install-disk-gc.sh. User units, NOT lane containers. - tests: 20 unit tests incl. GC plan-safety invariants + dry-run proof that default mode issues no destructive op (a wrong-delete GC is worse than none). Contract: contracts/OMN-13008.yaml * fix(OMN-13008): pass docker inventory to GC planner via stdin, not env (ARG_MAX) On a host with many docker images, passing the full image/ps inventory as env vars to disk_gc_plan.py exceeds ARG_MAX ('Argument list too long', exit 126), producing an empty plan. Write inventory to per-run scratch files (under the log dir, never /tmp; cleaned on exit) and hand it to the planner on stdin as a JSON envelope. Verified the failure live on .201; planner now reads stdin. * fix(OMN-13008): simplify GC plan stdin pipe (two processes, no nested subprocess) * fix(OMN-13008): keep-wins reconciliation — never remove an image id with any keep reason A single image id can surface in multiple 'docker image ls' rows (one per repo:tag). One tag could route the id to dangling-removal while another routes it to keep (e.g. tagged 'latest' or within-N-generations). Live .201 plan showed an id in BOTH remove_image_ids and kept_reasons. Reconcile at the end: keep wins — any id with a keep reason is dropped from the remove list; remove list deduped. Adds 2 regression tests. Verified live on .201. * fix(OMN-13008): timer uses OnCalendar=hourly for reliable re-arm (oneshot service) OnUnitActiveSec does not reliably re-elapse for a oneshot service once it goes inactive (observed NextElapseUSecMonotonic=infinity live on .201 after the first run). Switch to OnCalendar=hourly + RandomizedDelaySec so the timer re-arms every hour. Keep Persistent=true for missed-run catch-up. --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13005): materialize blocking event_consumer in runtime auto-wiring (runner consume-leg degenerate rows) (#1951) * fix(OMN-13005): materialize blocking event_consumer in auto-wiring (runner consume-leg degenerate rows) The runtime auto-wiring materialized event_publisher for handlers that declare it but had no equivalent for event_consumer. Request/response EFFECT handlers (HandlerContextRoiRunner) that publish a command then block on the correlated terminal event fell back to their no-op consumer default, returning None immediately -> every result row degenerate (failure_stage=generation, attempt_count=0) while generations succeeded ~1s later. Adds _make_sync_event_consumer (mirror of _make_sync_event_publisher), backed by service_terminal_event_consumer.make_terminal_event_consumer: a sync (topic, correlation_id, timeout) -> dict | None adapter that runs the proven direct-Kafka correlate-and-wait loop (from RuntimePatternBBroker) on an isolated event loop in a worker thread, so blocking does not deadlock the runtime dispatch loop that delivers the awaited terminal. TDD through the REAL dispatch path: test_event_consumer_injection drives a trial through _prepare_handler_wiring with a terminal arriving after a delay and asserts a non-degenerate row; verified RED with injection disabled. * fix(OMN-13012): two-phase (seek-now/wait-later) terminal event_consumer to close the subscribe-after-publish race (#1954) The OMN-13005 injected event_consumer is a single callable that does assign -> seek_to_end -> poll internally, all AFTER the handler has already published its command. Once OMN-13010 freed the dispatch loop and generation began completing in ~1s, the correlated terminal lands BEFORE the single-call consumer's post-publish seek_to_end positions, so seek_to_end skips PAST the already-emitted terminal and the runner times out on an offset beyond it (probe3, run_id=20260611T2140Z-probe3 -- both arms degenerate, zero rebalances). Splits positioning from waiting so the caller subscribes BEFORE it publishes: session = consumer.open(topic) # assign + seek_to_end NOW publisher(command_topic, payload) # publish AFTER positioning payload = session.wait(cid, timeout) # block from the captured position The returned TerminalEventConsumer is still directly callable with the legacy (topic, cid, timeout) -> dict | None single-call shape for any consumer that does not need subscribe-before-publish; the runner is the only consumer today. TerminalConsumerSession owns a dedicated event loop on a daemon worker thread for the whole open->wait->close lifecycle, preserving the OMN-13005 loop-isolation discipline so blocking never deadlocks the runtime dispatch loop. TDD (real dispatch path, RED-then-GREEN): extends the OMN-13005 injection test with a terminal emitted IMMEDIATELY after publish. The single-call (seek-after-publish) consumer MISSES it (degenerate row -- RED test asserts failure_stage=generation); the two-phase (open-before-publish) consumer CATCHES it (non-degenerate -- GREEN). The Kafka layer is faked at the two service seams against a shared in-memory log modeling seek-to-end semantics. OMN-13005 blocking-correlate behavior preserved. 269/269 auto_wiring unit tests pass; mypy --strict clean. Sibling to OMN-13010 / OMN-13005 / OMN-13003. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13005): route terminal consumer through Kafka boundary --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12990): pin worker replicas in ledgered config + fail-fast lane overrides (silent-drop ratchet) (#1948) * fix(OMN-12990): pin worker replicas in ledgered config + fail-fast lane overrides The base compose sets runtime-worker deploy replicas to ${WORKER_REPLICAS:-0} (soft-default ZERO). The stability lane's required state includes a running worker (GATE_ZERO_PROOF.md: 4 runtime containers), but the worker presence rode entirely on a compose soft-default — any plain compose up/recreate without the policy env silently scaled the worker to zero with no error and no signal. Fix (contract-native + fail-fast): - Add 'replicas' to ModelRuntimeProcessPolicy; pin replicas: 1 in every lane's worker block in runtime_policy.contract.yaml. - Renderer emits {PROFILE}_WORKER_REPLICAS into the ledgered runtime-policy.env for dev/stability-test/judge/prod. - stability + prod compose overrides reference ${..._WORKER_REPLICAS:?...} (fail-fast, NO silent :-1/:-0 default). A recreate that omits the policy env now aborts loudly instead of dropping the worker. prod previously had no override at all and inherited the dangerous :-0 default. Ratchet (recurrence guards): - tests asserting fail-fast override form (no soft default), contract-declared replica pin >= 1 per lane, and rendered {PROFILE}_WORKER_REPLICAS=1 in the ledgered env. - runbook deploy/verify procedure adds an expected-container census (worker must be present) via verify_container_manifest; a missing worker is a FAILURE, not silence. Note: SKIP=topic-parity-gate — that local-only advisory gate (absent from all .github/workflows, not a required CI check) fails on 25 pre-existing cross-repo topic gaps (build-loop/omniclaude/omniweb) identical on pristine base HEAD 8d7da1249; this change adds zero topics. All other hooks ran clean. Evidence-Ticket: OMN-12990 * test(OMN-12990): cover worker replica policy integration --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12909): add gateway bus forwarder P0A (#1946) * feat(OMN-12909): add gateway bus forwarder p0a * test(OMN-12909): add gateway forwarder integration coverage * fix(OMN-12909): satisfy gateway forwarder validators * test(OMN-12909): allow gateway forwarder bus protocol * fix(OMN-12909): sync gateway forwarder entry point * fix(OMN-12909): refresh runner image identity lock * test(OMN-12909): relax JSON normalizer mixed benchmark threshold * fix(OMN-12909): allow gateway handlers to boot unconfigured * fix(OMN-12909): declare gateway forwarder runtime profile * fix(OMN-12909): update runner identity backmerge expectation --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13011): LANE CENSUS RECONCILIATION ratchet — declared desired-state per lane, drift = auto-ticket (#1955) The class fix for the recurring lane-drift regression. Nothing reconciled the declared desired state of a runtime lane against what is actually running, so the same failure kept recurring with zero signal: volume config drift (OMN-12945), WORKER_REPLICAS silent zero (OMN-12988/12990), and on 2026-06-11 prod runtime containers plus the broker network were silently absent for hours during demo prep. Ships a per-lane DESIRED-STATE census: - (a) DECLARED in a versioned lane manifest (deploy/lane-census/lane-manifest.yaml): container set, network, replicas, image-tag pattern per lane (stability-test/prod/judge/dev), derived from the canonical compose lane files. A parity ratchet keeps the manifest locked in step with the compose files. - (b) RECONCILED on a schedule on .201 by SHARING the OMN-13008 systemd timer (a drop-in 4th ExecStart on onex-disk-gc.service — never a second timer) and on-demand via scripts/lane-census-check.sh / runtime_sweep. - (c) Drift = typed bus event (onex.evt.infra.lane-census-drift.v1) + Linear auto-ticket naming exactly what is missing/extra (container_absent, network_detached, replicas_zero, unexpected_container, oneshot_failed/stuck, image_tag_mismatch). Fail-fast, no warn-only mode (gates-block policy); exit 30 on drift; bus publish fail-fast on KAFKA_BOOTSTRAP_SERVERS (no localhost default). Red fixture reproduces 2026-06-11: prod runtime containers absent + broker network detached must produce the exact drift findings + a non-zero exit hours before a human noticed. Pure planner is fully unit-tested; shell driver dry-run-tested. Builds on the OMN-12988 deploy-agent RUNTIME census (deploy-time) as the complementary steady-state reconciler; closes the runtime-worker.yaml container_name: null census gap by sourcing names from the compose lane files. Evidence-Ticket: OMN-13011 Config-drift family: OMN-12945 Relates-to: OMN-13009, OMN-12988, OMN-13008 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13020): vendor missing node migrations — llm_routing 0000 + context_roi 001 (#1956) Vendors two omnimarket node-source migrations into the infra forward-migration tree via scripts/sync-node-migrations.sh (the canonical OMN-12559 mechanism): - node_projection_llm_routing/0000_create_llm_routing_decisions.sql (source: omnimarket #1168 / OMN-12942, merge ed6734f8) - node_projection_context_roi/001_create_context_roi_scores.sql (source: omnimarket #1178 / OMN-12955, merge 5010b1f4) Without the 0000 base table, node_projection_llm_routing/0001 (CREATE VIEW) hard-fails against NODE_POSTGRES_DB=omnidash_analytics — exactly the prod forward-migration exit-3 of 2026-06-11T09:35:52Z, and reproduced by construction in any clean clones@dev build. Files are byte-identical to the omnimarket dev blobs (sha256 f8a8b339… / c4126e65…) and to the untracked hot-patch copies on the .201 stability clone. Both migrations are self-contained, all-statements-IF-NOT-EXISTS, and 0000 sorts lexically before 0001 within the node's namespaced identity space. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13058): close TerminalConsumerSession on open() failure (worker thread + event loop leak) (#1957) TerminalConsumerSession.__init__ starts its dedicated worker loop thread immediately. TerminalEventConsumer.open did 'session = TerminalConsumerSession(...); return session.open()' with no cleanup: any failure inside session.open() (consumer start timeout, partition-assign timeout, broker auth error) propagated out of the raising expression, the session reference was lost, and the daemon worker thread plus its never-closed asyncio event loop leaked -- one pair per failed open. The motivating caller (HandlerContextRoiRunner) opens a session per trial, so a 160-560-trial battery against a degraded broker accumulates hundreds of leaked threads in the long-lived effects container. Fix: wrap session.open() in try/except BaseException -> session.close() (idempotent: stops the loop, joins the thread) -> re-raise. Covers both the two-phase .open(topic) path and the legacy single-call __call__ path. Found by the P3.3 doctrinal review of merged #1951 (b9712af9 / 36d98275). TDD: tests/unit/runtime/test_service_terminal_event_consumer_open_failure.py injects a real open failure through the production path (event bus without _bootstrap_servers) and asserts no alive terminal-consumer-* thread after the raise. Verified RED with the fix stashed (2 failed), GREEN with it (2 passed). Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13021): non-dev-base guard — fail feature-base PRs absent Stacked-Parent declaration (retro A-6) (#1958) Any PR whose base is neither dev nor main fails unless the body carries 'Stacked-Parent: #N'. Prevents the feedback_stacked_prs_orphan_from_dev class (#1185/#1954 auto-merged INTO parent feature branches and stranded off dev). base=main remains governed by main-target-guard. Epic OMN-13013 (process enforcement ratchets — June 12 retro). Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13014): hot-patch ledger rebuild preflight gate (retro B-1) (#1959) * feat(OMN-13014): hot-patch ledger rebuild preflight gate (retro B-1) Hot-patches on .201 (.prepatch sibling discipline) silently revert on any image rebuild/force-recreate — the 2026-06-11 20:58Z rebuild already erased a live /api/generate patch once. This adds the rebuild-path gate: - scripts/preflight_hotpatch_ledger.py: given a target container or lane + per-repo build refs, hard-fails when any hot-patch ledger row's source PR merge commit is not an ancestor of the build ref (git merge-base --is-ancestor), plus a .prepatch tripwire of the running container (unledgered .prepatch = hard fail; --post-rebuild = zero .prepatch expected). Sole bypass: HOTPATCH_PREFLIGHT_BYPASS carrying the Rule-10 '# skip-token-allowed: <user-approval-receipt-id>' form. - scripts/deploy-runtime.sh: guard_hotpatch_ledger wired into main() before build/preview (both dry-run and execute), lane derived from the compose project; skips loudly only when no ledger exists on the host. - tests/unit/scripts/test_preflight_hotpatch_ledger.py: 17 unit tests (ancestor gate, lane scoping, ledger loading, tripwire, bypass forms). - tests/ci/test_receipt_gate_install_guard.py: repair stale guard — core OMN-12565 replaced the OMN-9198 'uv pip uninstall first' install step with a cleared workspace venv (uv venv --clear); assert the new contract. Ledger backfilled from live census (5 .prepatch files / 4 source PRs, all MERGED to dev) at /data/omninode/hotpatch-ledger/ledger.yaml on .201. * fix(OMN-13014): scope missing-prepatch tripwire warning to the probed container --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12977): workspace-build sibling-pin ratchet — honor consuming lock, fail-fast on stale vendor (#1936) * fix(OMN-12977): workspace-build sibling-pin ratchet — honor consuming lock, fail-fast on stale vendor The workspace-mode image build (BUILD_SOURCE=workspace, used by the stability-test deploy procedure) vendored sibling/foundation packages from whatever the canonical OMNI_HOME clones happened to be checked out at, ignoring the consuming repo's uv.lock. On 2026-06-11 this shipped a 13-day-stale omnibase_infra 0.37.0-dev (pre-OMN-12501 Protocol-quarantine guard) + core 0.42.0 against an omnimarket dev lock pinning infra 0.38.1@e2dbdc95 / core 0.44.0@c97c2c9a, dropping the guard and crashing wire_from_manifest bootstrap fatally (stability lane down on demo day). Fix + recurrence ratchet (same PR): - scripts/runtime_build/check_sibling_lock_pins.py: parse the consuming repo's uv.lock for expected version+git-rev of each foundation/sibling package (scoped to the package's own source line so editable/registry pins are not cross-attributed a dependency's rev), resolve the actual clone version+HEAD, compare, and classify drift backward/forward/none. Fail-fast (exit 1) on any drift; --allow-drift records an explicit operator override in the artifact, never silent. - stage_workspace.sh: runs the preflight against the canonical clones before staging; aborts the build (exit 3) on unacknowledged drift and writes workspace/sibling-pin-comparison.json. - compute_workspace_provenance.py: folds the expected-vs-actual comparison into build-provenance.json so deploy verifiers can assert the build honored the lock; flags unacknowledged drift as a provenance error. - Dockerfile.runtime: COPY the comparison artifact (committed placeholder so the COPY always resolves; overwritten by stage_workspace.sh in workspace mode). - TDD: 19 unit tests covering lock parsing (git/registry/editable sources), drift classification, the exact 0.37.0-vs-0.38.1 stale case, check_pins exit codes, and the allow-drift override. Pre-existing mypy-strict bare-dict errors in compute_workspace_provenance.py fixed in the same pass. Evidence-Ticket: OMN-12977 Evidence-Source: pending-occ * ci(OMN-12977): retry runtime smoke compose port race * test(OMN-12977): align sibling-pin script tests with current API --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12989): workspace-mode image build must honor sibling lock pins (#1947) * fix(OMN-12989): workspace build must honor sibling lock pins (fail-fast ratchet) The 2026-06-11 stability bootstrap crash was caused by a workspace-mode --no-cache rebuild that vendored omnibase_infra 0.37.0-dev (13-day-stale worktree) even though omnimarket dev's uv.lock pins infra 0.38.1. The downgraded sibling predated the OMN-12501 Protocol-quarantine guard and turned a latent contract defect into a fatal crash. Fix + ratchet (same PR set): - scripts/runtime_build/resolve_workspace_pins.py: parse the consuming repo's uv.lock for sibling pins (version + git rev); classify each staged/installed sibling as exact/ahead/regression/unpinned; fail-fast (WorkspacePinError) on any regression BELOW the lock pin. Stdlib version-tuple fallback when packaging is absent so the ratchet never fails-open. - compute_workspace_provenance.py: enforce sibling pins + a host-infra self-check (installed omnibase_infra vs lock pin — the exact crash vector, since host infra is built from the context, not staged), and emit a pin_comparison block into build-provenance.json for deploy verifiers. - Dockerfile.runtime: COPY resolve_workspace_pins.py beside the provenance script so the in-image import resolves. - TDD: failing tests first (test_resolve_workspace_pins.py, 11 cases) + provenance integration tests; fixed a pre-existing stale RUNTIME_VERSION assertion to read pyproject dynamically. Evidence-Ticket: OMN-12989 * fix(OMN-12989): workspace build must honor sibling lock pins (fail-fast ratchet) The 2026-06-11 stability bootstrap crash was caused by a workspace-mode --no-cache rebuild that vendored omnibase_infra 0.37.0-dev (13-day-stale worktree) even though omnimarket dev's uv.lock pins infra 0.38.1. The downgraded sibling predated the OMN-12501 Protocol-quarantine guard and turned a latent contract defect into a fatal crash. Fix + ratchet (same PR set): - scripts/runtime_build/resolve_workspace_pins.py: parse the consuming repo's uv.lock for sibling pins (version + git rev); classify each staged/installed sibling as exact/ahead/regression/unpinned; fail-fast (WorkspacePinError) on any regression BELOW the lock pin. Stdlib version-tuple fallback when packaging is absent so the ratchet never fails-open. - compute_workspace_provenance.py: enforce sibling pins + a host-infra self-check (installed omnibase_infra vs lock pin — the exact crash vector, since host infra is built from the context, not staged), and emit a pin_comparison block into build-provenance.json for deploy verifiers. - Dockerfile.runtime: COPY resolve_workspace_pins.py beside the provenance script so the in-image import resolves. - TDD: failing tests first (test_resolve_workspace_pins.py, 11 cases) + provenance integration tests; fixed a pre-existing stale RUNTIME_VERSION assertion to read pyproject dynamically. Evidence-Ticket: OMN-12989 * test(OMN-12989): co-locate provenance pin helper in fixture --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13055): absent repos in REPOS list warn and exit 0 instead of failing (#1960) Missing repos (not cloned locally) now emit a WARN line and are tracked in a separate WARNED array. Only real fetch/ff failures cause exit 1. This makes pull-all.sh safe to use on machines with a partial clone set, while keeping the explicit-list override behavior intact. Adds three regression tests: absent-only exits 0, absent+present exits 0 with OK for the present repo, present-failed+absent exits 1. Evidence-Ticket: OMN-13055 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13070): config_prefetcher overlay wins on controlled lanes (#1961) On infisical_required=True lanes, fetched/overlay config now always wins over ambient env. Ambient env is retained only as a declared bootstrap fallback (with an explicit provenance INFO log line) when Infisical returns None. apply_to_environment also overwrites stale env on controlled lanes. Uncontrolled lane (infisical_required=False) behaviour is unchanged. Adds 5 regression tests: controlled-lane Infisical-wins, env-bootstrap- fallback, apply_to_environment overwrite, missing-from-both-is-error, and uncontrolled-lane-env-still-wins. Refactors _resolve_key to return a (outcome, value, error) tuple to satisfy the ≤5-param pattern gate. Source: docs/audits/2026-06-10-runtime-env-overlay-authority-audit.md Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13062): migration-gate vacuity fix — sentinel discipline, wait-for-postgres, skip-manifest, sync-check nonzero (retro A-10) (#1963) * fix(OMN-13062): migration-gate vacuity fix — sentinel discipline, wait-for-postgres, skip-manifest, sync-check nonzero Fixes three bugs identified in retro A-10 (recurrences OMN-12885, OMN-12934): (1) RUNNER SENTINEL DISCIPLINE run-forward-migrations.sh now clears migrations_complete=FALSE at the start of every run and sets it TRUE only as its FINAL act after all infra and node migrations succeed. Any mid-run failure leaves the gate UNHEALTHY. runner_completed_at is stamped at the same final step as durable evidence of a successful completion run. (2) SYNC-NODE-MIGRATIONS VACUOUS GATE sync-node-migrations.sh --check now exits 2 (not 0) when the omnimarket source tree is unresolvable. Silent exit-0 was hiding drift. The single opt-out is SYNC_NODE_MIGRATIONS_SKIP_UNRESOLVABLE=1 for environments that intentionally run without the source. (3) WAIT-FOR-POSTGRES GUARD run-forward-migrations.sh now waits up to PG_WAIT_RETRIES (default 30) x 2s for Postgres to accept connections before proceeding, guarding the first-boot initdb race. (4) SKIP-MANIFEST docker/migrations/skip-manifest.yaml introduced as the sole committed escape for intentionally-skipped migrations. The runner reads this at startup; listed migrations are recorded in schema_migrations with checksum "skip-manifest" without executing the SQL. (5) MIGRATION 085 Adds runner_completed_at TIMESTAMPTZ column to db_metadata so the runner's final stamp is durable in the schema (idempotent ADD COLUMN IF NOT EXISTS). Rollback included. 22 regression tests added covering all five fix surfaces. * fix(OMN-13062): stamp schema fingerprint for migration 085 Migration 085 (085_add_runner_completed_at_to_db_metadata.sql) was added in the initial commit but schema_fingerprint.sha256 was not regenerated. Running `python scripts/check_schema_fingerprint.py stamp` updates the artifact from the stale hash to match the 71 migration files. Evidence-Source: OCC#2563 Evidence-Ticket: OMN-13062 --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12864): Bifrost endpoints → committed overlay authority + fail-loud loader (#1964) * feat(OMN-12864, OMN-12814, OMN-12945): Bifrost endpoints → committed overlay authority + fail-loud loader OMN-12864 — Committed lane overlay - docker/lane-overlays/dev.bifrost.yaml: typed deployment bindings for all four BIFROST_LOCAL_*_ENDPOINT_URL values (coder :8000, reasoner :8001, embedding :8100, ds4-flash :8101). Previously only available as ephemeral shell exports on .201; now committed, auditable, diff-able, and CI-checked. - docker/lane-overlays/dev.bifrost.env: generated dotenv sidecar consumed by compose via env_file; never edited directly (yaml is authority). - docker/docker-compose.infra.yml: wire the env_file block at compose root so the four endpoints are injected into the interpolation context on a clean shell. Hardcode BIFROST_CONTRACT_PATH (remove :-/empty footgun — OMN-12814). - scripts/render_bifrost_lane_overlay_env.py: render script regenerates the env sidecar from the YAML source. - src/omnibase_infra/runtime/models/model_bifrost_lane_overlay.py: ModelBifrostLaneOverlay — typed Pydantic model enforcing URL completeness (OMN-12815: every URL must end in /chat/completions). OMN-12814 — Fail-loud loader - render_bifrost_delegation_contract: raises ProtocolConfigurationError on FileNotFoundError, YAMLError, ValidationError, and zero-endpoint renders. No lru_cache — every restart re-renders from packaged source so a stale cache cannot pin a broken result across deploys. OMN-12945 — Re-seed from packaged source on deploy - docker/entrypoint-runtime.sh: set BIFROST_FORCE_RESEED=1 on every container restart so the named-volume copy is always rebuilt from the packaged bifrost_delegation.yaml merged with committed lane-overlay endpoints. - render_bifrost_delegation_contract: honor BIFROST_FORCE_RESEED/force_reseed flag to bypass the stale-volume early-return path entirely. Tests: - tests/ci/test_bifrost_lane_overlay.py: CI gate — env sidecar in-sync with YAML source; all four BIFROST_LOCAL_* keys present. - tests/unit/runtime/models/test_model_bifrost_lane_overlay.py: bare-base URL rejection, env dict mapping, extra-field rejection. - tests/unit/runtime/test_render_bifrost_delegation_contract.py: fail-loud paths, force-reseed, zero-endpoint error, endpoint URL completeness. - tests/unit/models/test_model_serialization_roundtrip.py: roundtrip coverage. * fix(OMN-12864): move bifrost env_file to service level — fix compose schema validation failure Top-level 'env_file' is rejected by Docker Compose v2 schema validator ('additional properties not allowed'). This caused 10+ compose-render integration tests to fail in CI. Fix: - Remove top-level env_file block from docker-compose.infra.yml - Add per-service env_file on omninode-runtime, runtime-effects, runtime-worker (the three containers that render Bifrost) - Change BIFROST_LOCAL_*:? to BIFROST_LOCAL_*:- in x-runtime-env (compose-level validation removed; Python validates via ModelBifrostLaneOverlay + render_bifrost_delegation_contract) - Add two CI gate tests: compose_env_file_is_service_level_not_top_level and runtime_services_have_bifrost_env_file * fix(OMN-12864): allow BIFROST_LOCAL_* empty defaults in silent-fallback gate The overlay authority pattern (OMN-12864) passes BIFROST_LOCAL_*_ENDPOINT_URL via service-level env_file (docker/lane-overlays/dev.bifrost.env), not at compose config time. Compose-level :? would break CI rendering without the overlay pre-loaded. Validation at the Python layer (ModelBifrostLaneOverlay + render_bifrost_delegation_contract) is the enforcement point. Add the three failing vars to ALLOWED_EMPTY_DEFAULTS with OMN-12864 citation. --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12858/OMN-12879/OMN-12880): dispatcher route-coverage CI gate (#1965) * feat(OMN-12858/OMN-12879/OMN-12880): dispatcher route-coverage CI gate Static contract analysis: every subscribed command topic (onex.cmd.*) must declare handler_routing or runtime_dispatch, or the message goes to DLQ silently. This gate would have caught two recent incidents: 1. June 9 DLQ regression (OMN-12858 post-mortem): node_generation_consumer subscribed onex.cmd.omnimarket.node-generation-requested.v1 but a sole-handler revert left zero dispatcher routes registered. Messages went to DLQ silently with no CI signal. 2. June 12 DEL-01 live finding: onex.cmd.omnimarket.delegate-skill.v1 was consumed by dev lane bus but no dispatcher route existed in any deployed contract. Discovered via manual rpk consumer-group lag probe (DEL-01 evidence, docs/evidence/2026-06-12-weekend-pass/). Deliverables: - scripts/check_dispatcher_route_coverage.py — static YAML scanner that checks both omnibase_infra and omnimarket contract trees; ratchet allowlist for known pre-existing violations; --changed-contracts mode (OMN-12879) for per-PR scoping; compat publish topics excluded (OMN-12880) - .github/workflows/dispatcher-route-coverage.yml — CI workflow that checks out omnimarket sibling, collects changed contract paths in PR mode, and runs the gate; fires on PR, push-to-main, and merge_group - tests/ci/test_dispatcher_route_coverage_gate.py — 12 unit tests covering RED/GREEN/COMPAT/CHANGED-MODE/ALLOWLIST/MULTI-DIR paths plus live-contract regression proof against the actual omnibase_infra tree Allowlist additions: - onex.cmd.omnibase-infra.pattern-b-dispatch.v1 (RuntimePatternBBroker, imperative consumer, OMN-12525 migration target) - onex.cmd.platform.contract-resolve-requested.v1 (transitional HTTP bridge node_contract_resolver_bridge OMN-2756, metadata.transitional=true) [OMN-12858, OMN-12879, OMN-12880] * fix(OMN-12858): drop full uv sync from dispatcher-route-coverage workflow Gate script only needs pyyaml (stdlib + yaml). Using full setup-python-uv was causing 10+ minute timeout. Replace with direct pip install pyyaml and invoke python3 directly. Reduces job from 10m timeout to <1m. [OMN-12858] --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13026): wire transport-mock-lint pre-commit hook + CI gate (#1962) * feat(OMN-13026): wire transport-mock-lint pre-commit hook + CI gate (omnibase_infra) Activates the transport-mock-lint validator (from omnibase_core, OMN-13026) on omnibase_infra. Ratchet baseline: 218 existing violations across 63 files frozen in validation/transport_mock_baseline.yaml. New bare AsyncMock/ MagicMock on EventBus/transport surfaces are blocked by pre-commit hook and CI lint step. Existing violations tracked for drain by per-site tickets (parent OMN-13026). Reference incident: PR #1181 bare AsyncMock hid missing EventBusKafka.stop(). Evidence-Ticket: OMN-13026 * fix(OMN-13026): use uv run python for CI lint step + bump omnibase_core pin to include transport_mock_lint The transport-mock lint CI step previously cloned omnibase_core and ran `python -m omnibase_core.validators.transport_mock_lint` with PYTHONPATH, but this failed: `No module named omnibase_core.validators.transport_mock_lint` because it ran `.venv/bin/python` which uses the locked venv, and the venv omnibase_core pin (2defabef4) predates the transport_mock_lint module. Fix: use `uv run python` (removes the clone step) and bump omnibase-core git pin from 2defabef4 to 309d89fa7 (PR 1231 merge commit on dev) so transport_mock_lint is available in the locked venv. * fix(OMN-13026): align transport mock baseline and runner lock * fix(OMN-13026): sync omnibase_core pin + runner identity lock to dev baseline Align pyproject.toml omnibase_core rev to 2defabef (required by test_release_backmerge_preserves_proven_runtime_core_pin) and update docker/runners/runner-image.lock.json identity_digest/shared_env_digest to match dev runner image lock (79b08f44 / 90c8b3b9). Both were stale from the prior session's pin bump that used an older SHA. * fix(OMN-13026): source transport mock validator from core * fix(OMN-13026): source transport validator from core dev --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13094): --output receipt mode on onex node/run — quiet typed receipts with durable capture (#1966) Phase 2a of the skill-output-suppression slice (epic OMN-13089, plan docs/plans/2026-06-12-skill-output-suppression-plan.md Phase 2 item 1). - onex node/onex run gain --output receipt: ALL runtime logging routes to a run_id-suffixed capture file under <state-root>/captures/ (no console handlers — kills the 25-50-line RuntimeLocal INFO stream at the source); stdout carries exactly ONE typed ModelSkillResult JSON with the FULL handler result (result_model = concrete handler result type FQN). - Durable capture: capture log + handler result content-addressed via omnibase_core ArtifactStore (OMN-13093); artifact.captured + tool.output.captured emitted to the emit daemon socket (--emit-socket, default ~/.claude/emit.sock). - Failure asymmetry: artifact write failure => FULL output printed, no receipt (no hidden loss); emission failure => receipt still prints, event spooled to <state-root>/emit_spool/ for replay. - Node failure => status=failed/error with full error + capture log INLINE in the receipt (errors are never hidden) and artifact-backed. - Default output mode unchanged (enforcement is Phase 4). - RuntimeLocal exposes handler_result (receipt schema identity). - core pin 2defabef -> ae8793bd (merged OMN-13091/13093 receipt models + ArtifactStore); runner-image identity lock regenerated and the OMN-12765 backmerge identity constants updated for the new pin. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13097): onex skill subcommand + declarative skill->node mapping (#1968) * feat(OMN-13097): onex skill subcommand + declarative skill->node mapping Phase 4a of the skill-output-suppression slice (epic OMN-13089, plan docs/plans/2026-06-12-skill-output-suppression-plan.md Phase 4 item 1/2). A dispatch skill IS one CLI call (user directive 2026-06-12). This adds the `onex skill <name> [args]` dispatch surface that the 24 omniclaude shim migrations build on: - onex.cli entry-point `skill` -> cli_skill.run_skill_by_name. Resolves the skill via the declarative skill_mapping.yaml registry, builds the backing node's input payload from the skill's CLI args, writes it under the state root (.onex_state/tmp/<skill>-<run_id>.json — never /tmp), resolves the node's packaged contract exactly like `onex node`, and dispatches through the proven receipt-mode path (run_receipt_mode, OMN-13094). stdout is exactly one typed ModelSkillResult JSON with the FULL handler result. - skill_mapping.yaml: declarative DATA mapping all 24 dispatch shims to their backing onex.nodes node + typed result-model FQN (verified against each live handler handle() return type on origin/dev) + per-arg payload specs + static payload + keyword classifiers (delegate task_type as data, not code). Adding a skill is a YAML edit + fixture, never a CLI code change (ticket deliverable 2/3). Mapping lives beside the node-resolution surface, never hardcoded branching in the CLI. - Typed models split one-per-file (repo convention): ModelSkillArgSpec, ModelSkillClassifier, ModelSkillMapping, ModelSkillMappingRegistry, EnumSkillArgType. Frozen, extra=forbid, fail-fast coercion/validation. - validation_exemptions.yaml: Click-callback param-count + literal-identifier name-field exemptions mirroring the existing cli_node run_node_by_name precedent (OMN-11570) — same pattern, same rationale. dod_evidence: - 20 unit tests pass (registry validity, all-24-shims coverage, FQN result models, arg parsing/coercion/positional/required, classifiers, payload build, receipt-mode dispatch wiring, payload-under-state-root not /tmp). - uv run mypy src/ --strict: clean (2438 source files). - ruff format + check: clean. pre-commit run on changed files: pass. - runner-image identity lock regenerated for the pyproject entry-point add (same as OMN-13094). * test(OMN-13097): rebind OMN-12765 backmerge identity constants for onex skill pyproject change Adding the `skill` onex.cli entry-point to pyproject.toml changes the runner-image identity_digest (and shared_env_digest) the lock binds. Update the hardcoded expected constants in the backmerge-identity test to the regenerated values — same mechanical rebind OMN-13094 performed for the core pin bump. Identity + runner-image-identity tests pass (14). --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix+test(OMN-13012): force terminal-topic metadata refresh so both ephemeral consumers assign (runner consume-leg wedge) (#1969) The runner consume-leg wedged on the live stability battery (image c0505521f1fa, EXP1-3_RUNNER_CONSUME_LEG_BLOCKER): only the FAILED terminal topic ever subscribed while the COMPLETED topic never assigned, so the correlated completed terminal was never read and the 8x2x10 matrix re-fired cell 1 forever, emitting zero non-degenerate rows. A prior two-strike diagnosis proved the omnimarket handler is correct (it opens both terminal sessions pre-publish and races them); the defect is in the omnibase_infra runtime consume leg. Root cause: _assign_direct_terminal_partitions ignored the metadata future returned by AIOKafkaClient.set_topics and re-called set_topics([same_topic]) each loop iteration. aiokafka 0.13.0 set_topics only forces a metadata refresh when the topic set DIFFERS from the tracked set, so every iteration after the first took the no-op branch and never re-fetched. An ephemeral group_id=None consumer whose first metadata fetch had not yet surfaced partitions burned the full 30s assign cap and raised a bare TimeoutError (the empty-message 'wait failed' seen live). Fix: register the reply topic once and await that metadata fetch, then on each miss force a fresh fetch via force_metadata_update (which always fetches) rather than the no-op set_topics repeat. The assign-cap TimeoutError now carries a diagnostic message instead of an empty one. Test: tests/integration/test_terminal_consumer_concurrent_assign_race.py drives the REAL TerminalEventConsumer (the object wired as event_consumer) through the REAL open_direct_terminal_consumer/poll path with AIOKafkaConsumer monkeypatched to a fake that faithfully models aiokafka 0.13.0 set_topics future + metadata-latency semantics. RED before the fix (bare TimeoutError, the live empty-message signature); GREEN after. K>=2 multi-trial variant asserts no worker-thread leak across trials. Evidence-Source: <occ-sha-pending> Evidence-Ticket: OMN-13012 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13096): onex delegate single-command subcommand (Phase 2b) (#1967) * feat(OMN-13096): onex delegate single-command subcommand Add 'onex delegate "<prompt>" [--task-type X] [--max-tokens N]' as a subcommand on the existing onex CLI (Phase 2b of the skill-output-suppression slice, OMN-13089). The command wraps payload construction, node dispatch, and result extraction internally and prints exactly one ModelSkillResult[ModelDelegateSkillResponse] to stdout via the OMN-13094 receipt-mode path. RuntimeLocal logs go to the capture file + artifact store, never to stdout; scratch payloads live under <state-root>/tmp/ with run_id suffixes (never /tmp). - cli_delegate.py: classify_task_type (keyword table from legacy skill md), payload write, contract resolve, run_receipt_mode dispatch - register 'delegate' under onex.cli entry points - exempt delegate_command from the >5-param patterns gate (same Click-callback rationale as run_node_by_name) - 20 unit tests: classification, scratch-under-state-root, single typed receipt on stdout, zero INFO log leakage omnibase_infra does NOT depend on omnimarket; the delegate node is resolved at runtime via the onex.nodes entry-point group (registered by omnimarket). * chore(OMN-13096): re-trigger deploy-gate after Evidence-Source set to OCC#2593 No code change — the deploy-gate workflow triggers on synchronize (not edited), so the PR-body Evidence-Source fix needs a new commit to re-resolve the OCC ref to the open PR head where contracts/OMN-13096.yaml (with deploy evidence) lives. * chore(OMN-13096): re-trigger deploy-gate now that OCC#2593 merged to OCC dev contracts/OMN-13096.yaml (with the dod-deploy-onex-delegate item) is now on OCC dev, so the deploy-gate OCC-dev checkout resolves the contract + deploy evidence. * chore(OMN-13096): regenerate runner-image identity lock for pyproject entry-point add Adding the 'delegate' onex.cli entry point to pyproject.toml changed the dependency-manifest digest that scripts/ci/runner_image_identity.py folds into the runner-image identity lock. Regenerate the lock so tests/ci/test_runner_image_identity.py matches (was the only CI test failure; unrelated environmental integration/perf failures excluded). * test(OMN-13096): update backmerge identity assertions to regenerated lock digests The runner-image identity lock was regenerated for the pyproject entry-point add; this test hardcodes the expected identity_digest/shared_env_digest, so update both to match the new lock (same maintenance OMN-13094 did). --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): tolerate partition-less reply topic in terminal consume leg (#1970) The context-ROI runner opens one ephemeral group_id=None terminal consumer per terminal topic BEFORE publishing each generation command (subscribe-before-publish, OMN-13012/13038). The FAILED reply topic is only produced to on contract_passed=False; in a battery where generations pass it has zero messages, so Redpanda never advertises a partition for it. _assign_direct_terminal_partitions burned the full 30s assign cap on every trial then raised a bare TimeoutError (the empty-message 'wait failed'), stalling each of the 160 battery trials ~30s before the COMPLETED terminal could correlate -> battery needs >80 min and never completes (verifier-confirmed wedge). A partition-less reply topic is a valid steady state, not a 30s error: - _assign_direct_terminal_partitions gives a bounded grace window for a topic that exists but is slow to surface metadata, then assigns whatever partitions exist (possibly none) and returns promptly instead of burning the cap and raising. - poll_direct_terminal_consumer treats an empty assignment as 'no terminal will arrive here' (sleeps out its timeout, returns None) without calling getone() on an unassigned consumer. Repro: tests/integration/test_terminal_consumer_battery_load_wedge.py drives the REAL TerminalEventConsumer over K=10 x 2 cells x 2 arms; RED (40/40 trials block a full assign cap) before the fix, GREEN after. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): pin terminal-consumer read offset synchronously (close lazy seek_to_end publish race) (#1971) * fix(OMN-13118): pin terminal-consumer read offset synchronously to close lazy seek_to_end publish race The consume-leg wedge survived PR #1969 (set_topics no-op) and PR #1970 (partition-less assign-cap) because both addressed the assign phase, not the seek timing. AIOKafkaConsumer.seek_to_end is LAZY: it requests a LATEST offset reset that only resolves on the first poll — AFTER the caller publishes. With generation completing in ~1s, the correlated COMPLETED terminal lands in the open->poll gap, so the lazily-resolved LATEST position is the HWM AFTER the record and the poll reads past it. The terminal is never read, the trial never correlates, and the experiment matrix re-fires the same cell forever. Replace seek_to_end with a synchronous end_offsets() + seek() pin (_pin_direct_terminal_end_offsets) in both open_direct_terminal_consumer and RuntimePatternBBroker._dispatch_and_wait_with_direct_kafka_consumer, so the read position is fixed at open() time, before the publish — the real subscribe-before-publish guarantee. Empty assignment (partition-less reply topic, OMN-13118 #1970) is a no-op. Repro: tests/integration/test_terminal_consumer_seek_reset_race.py drives the real TerminalEventConsumer.open()/wait() the way HandlerContextRoiRunner does, publishing the correlated terminal in the open->wait gap across K=10 x 2 cells x 2 arms; RED with the lazy reset (every cell degenerate), GREEN once the read offset is pinned. RED verified by git-stashing only the source fix. Existing consume-leg fakes updated to model end_offsets/seek (they previously masked the bug by making seek_to_end a synchronous exact snapshot). * test(OMN-13118): reword assertion (lazy not deferred) for receipt honesty gate * fix(OMN-13118): bound end_offsets() round-trip with assign-cap timeout (CodeRabbit) end_offsets() is a broker ListOffsets round-trip aiokafka documents as able to block indefinitely. Bound it with the same cap as start()/assign so a stalled broker fails fast instead of hanging the pre-publish positioning. --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13137): _validate_routing strategy-aware event_model checks (#1973) operation_match routes by the `operation` field and does not use event_model. The validator was unconditionally requiring event_model.{name,module} for every handler entry, causing 230/295 omnimarket operation_match contracts (all correct as authored) to fail routing validation at startup. Fix: read routing_strategy from the routing map and branch validation: - payload_type_match → require event_model.{name, module} (unchanged) - operation_match (and any non-payload strategy) → require `operation`; skip event_model checks entirely Updated pre-existing _validate_routing tests to declare routing_strategy: payload_type_match explicitly (they always tested payload_type_match semantics but relied on the implicit fallback that is now removed). Added test_validate_routing_operation_match.py with 4 unit tests: 1. operation_match without event_model → zero event_model errors 2. operation_match missing operation field → error 3. payload_type_match missing event_model → still errors (regression guard) 4. Real node_integration_sweep_orchestrator routing block → clean (boot gate) Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): independent per-terminal-topic consumers in Pattern B direct-Kafka wait (#1972) The consume-leg wedge survived four merged fixes (#1969 set_topics no-op, #1970 partition-less assign-cap, #1971 synchronous seek-pin). The STRONG K>=10 multi-cell reprobe on the stability lane still wedged on REBUILD-5 (cdf53d963f7b). Converged diagnosis (strikes 3+4, docs/evidence/2026-06-12-weekend-pass/experiments/probe4-stability/ reprobe-K10-rebuild5/HALT_K10_WEDGE_PERSISTS.md): the runtime waited for each trial's terminal across TWO topics (node-generation-completed.v1 + node-generation-failed.v1) with a SINGLE ephemeral group_id=None consumer assigned both topics' partitions. One aiokafka consumer holds one manual subscription; the COMPLETED delivery window collapsed before it surfaced the correlated record, so the trial never correlated and the matrix re-fired cell 1. RuntimePatternBBroker._dispatch_and_wait_with_direct_kafka_consumer now opens ONE independent AIOKafkaConsumer PER terminal topic via open_direct_terminal_consumer (each started, assigned, and offset-pinned via end_offsets()+seek() at open() BEFORE publish), then awaits both CONCURRENTLY via asyncio.wait(FIRST_COMPLETED). The first correlated terminal wins; both are torn down. No shared consumer, no subscription flip. Keeps the #1970 partition-less no-op and the #1971 synchronous offset pin (both live in open_direct_terminal_consumer / poll_direct_terminal_consumer). Removes the now-dead single-consumer helpers (_assign_terminal_topic_partitions, method-level _refresh_terminal_topic_metadata, _direct_kafka_* kwargs builders, _kafka_bootstrap_servers/_kafka_event_bus). Adds tests/integration/test_terminal_consumer_subscription_flip_wedge.py: a real-dispatch-path K>=10 x 2-cell x 2-arm repro whose fake models TWO independent consumers honestly (delivery is faithful only for a single-topic assignment; a consumer spanning both topics flips and drops the COMPLETED record). RED genuineness verified by reverting only the source to the single-consumer shape (test hangs past timeout); GREEN with the fix in 2.3s. Acceptance is the LIVE K>=10 multi-cell stability-lane reprobe (later phase), NOT this unit test. A green unit repro is necessary but NOT sufficient. Refs OMN-13118, OMN-13128. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): long-lived terminal correlator replaces per-trial ephemeral consume leg (#1974) Tier B canonical redesign (epic OMN-12525). Five offset/subscription patches (#1969-#1972) tuned the per-trial-ephemeral terminal consumer and all failed…
* chore(OMN-12934): continue dev to main promotion for prod redeploy (#1988)
* docs(OMN-12962): contract-store durability audit — cold-runtime census proof (#1937)
Proves cold-start contract census reconstructs from the image-bundled
filesystem manifest (HYBRID-mode bootstrap + PluginLoaderContractSource),
independent of node-registration.v1 retention. The delete-retention topic
feeds only the post-freeze dynamic listener (auto_offset_reset=latest, no
history replay). Live .201 stability-test evidence: filesystem contract_path
in manifest + truncated topic log head with intact census. No store fix
needed; residual dynamic-only gap covered by runtime_sweep sweep check.
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* fix(OMN-12970): vendor omnimarket projection node migrations into forward/nodes (#1938)
Vendors omnimarket node-owned projection migrations into the namespaced
forward-migration tree so run-forward-migrations.sh materializes them in the
dashboard projection DB (omnidash_analytics) at deploy.
Primary (OMN-12970): creates llm_call_metrics, llm_cost_aggregates, and
capability_scores in the projection DB. These were only ever created in the
omnibase_infra DB by infra migrations 031/060, so the ab-compare,
cost.token_usage, cost.summary, and capability-scores projection topics were
DEGRADED at startup ('table not found') and their dashboard panels rendered
empty.
Also re-syncs three omnimarket node migrations the vendor tree had drifted from
(node_projection_llm_routing, node_projection_overnight, node_projection_savings
/077) — sync-node-migrations.sh --check requires the full vendored tree to match
omnimarket source, and these were missing.
Companion to omnimarket PR for the same ticket (source migrations + projection
table-coverage ratchet test).
Evidence-Ticket: OMN-12970
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* fix(OMN-12965): stamp runtime image identity (version + revision) in workspace builds (#1943)
* fix(OMN-12965): stamp runtime image identity (version + revision) in workspace builds
The main runtime image stamped org.opencontainers.image.version=0.1.0 with a
blank org.opencontainers.image.revision after workspace rebuilds. A blank
identity degrades every proof packet (runtime SHA + image digest are required
citations in accepted evidence).
Root causes (three build paths under-stamped identity):
- onex up --build (cmd_up) passed only GIT_SHA; the runtime-stage OCI labels
read VCS_REF (-> blank revision) and RUNTIME_VERSION (-> placeholder 0.1.0).
- deploy-runtime.sh passed VCS_REF but not RUNTIME_VERSION/GIT_SHA.
- Dockerfile silently allowed blank/placeholder identity in workspace mode.
Fix:
- cli._image_identity_build_args() stamps the full quad (GIT_SHA/VCS_REF/
RUNTIME_VERSION/BUILD_DATE) and fails fast on an unresolved git revision.
- deploy-runtime.sh stamps RUNTIME_VERSION + GIT_SHA and verifies the version
label is non-placeholder post-deploy.
- Dockerfile.runtime fails workspace builds with blank VCS_REF or placeholder
RUNTIME_VERSION=0.1.0 (release mode unaffected).
Enforcement ratchet (same PR):
- scripts/check_runtime_image_identity.py static check, wired as pre-commit hook
+ CI gate (ci.yml).
- tests/unit/infra/test_runtime_image_identity_labels.py pins the cli helpers +
Dockerfile guard; deploy-agent test extended for the quad.
Proven locally via throwaway docker builds: workspace+args -> populated labels;
workspace without args -> guard fails (exit 64); release without args -> 0.1.0
placeholder allowed (no regression).
Evidence-Ticket: OMN-12965
* test(OMN-12965): integration build proof for runtime image identity labels
Builds the real runtime-stage ARG/LABEL/guard block against busybox and asserts
via docker inspect: workspace+args -> populated version/revision; workspace
without args -> guard fails (exit 64); release -> placeholder allowed. Satisfies
the integration-test hard gate and makes the throwaway proof permanent.
Evidence-Ticket: OMN-12965
---------
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* fix(OMN-12987): workspace-build sibling lock-pin preflight + provenance (#1944)
Recurrence ratchet for the 2026-06-11 stability bootstrap crash. The 11:20Z
--no-cache rebuild vendored omnibase_infra 0.37.0-dev (~2c1d672f) + core 0.42.0
even though omnimarket dev's uv.lock pinned infra 0.38.1 @ e2dbdc95 + core
0.44.0 @ c97c2c9a. The stale sibling predated the OMN-12501 Protocol-quarantine
guard, turning a latent contract defect into a fatal crash that crash-looped the
main runtime.
- check_sibling_lock_pins.py: host-side fail-fast preflight resolving expected
sibling versions/SHAs from the consuming repo's (omnimarket) uv.lock and
comparing them against each vendored tree. Mismatch aborts the build.
- stage_workspace.sh: emit a .build-sha marker per staged sibling (rsync drops
.git) so the preflight and provenance can identify the vendored commit.
- deploy-runtime.sh: run the preflight after staging, before build; abort on
mismatch. Write the comparison under sibling-repos/.sibling-lock-pins.json.
- compute_workspace_provenance.py + Dockerfile.runtime: fold expected-vs-actual
lock_pin_comparison into build-provenance.json for deploy verifiers.
Recurrence-guard tests prove a stale infra 0.37.0 vs lock-pinned 0.38.1 fails
the preflight and matched pins pass; deploy script wiring is asserted statically.
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* fix(OMN-12988): pin stability runtime-worker to replicas 1 + census ratchet (#1942)
The base docker-compose.infra.yml defaults runtime-worker to
replicas 0 (${WORKER_REPLICAS:-0}). The stability-test lane's required
state includes a running worker (4-container census: main, effects,
worker, projection-api), but the override pinned it via an
env-interpolation default (${STABILITY_TEST_WORKER_REPLICAS:-1}) — a
silent-drop surface: a stray exported STABILITY_TEST_WORKER_REPLICAS=0
or removal of the :-1 fallback would scale the worker to 0 with zero
signal on a plain compose up/recreate.
Fix: pin docker-compose.stability-test.yml runtime-worker
deploy.replicas to the literal 1 (no env indirection).
Ratchet (recurrence guards, same PR):
- scripts/deploy-agent/tests/unit/test_runtime_worker_census.py: assert
runtime-worker stays in the deploy-agent RUNTIME-scope census so a
missing worker (replicas 0 => absent from docker compose ps) is a
deploy failure, not silence; assert the override pins a literal 1.
- tests/integration/infra/test_stability_test_runtime_compose_render.py:
assert the rendered stability worker resolves deploy.replicas == 1.
- tests/unit/infra/test_stability_test_runtime_lane.py: update the
existing pin assertion to the literal 1.
Evidence-Ticket: OMN-12988
Config-drift family: OMN-12945
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* fix(OMN-12979): expire-bound topic completeness suppressions (#1940)
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* fix(OMN-12966): accept community-edition Infisical /api/status in provision readiness gate (#1939)
The fresh-provision readiness gate in provision-infisical.py only accepted the
enterprise {"status": "ok"} payload and rejected the community edition's
{"message": "Ok"}, returning 1 before bootstrap could run. This blocked
provisioning against the Infisical instance deployed on .201 (community edition).
Route the gate through the existing _is_infisical_ready helper (single source of
truth, already used by the already-provisioned path). Add TestMainFreshProvision-
ReadinessGate covering community/enterprise/not-ready cases.
Also adds docker/docker-compose.infisical-stability.yml: an ADDITIVE Infisical
compose project for the stability-test lane (joins the existing network as
external, reuses stability postgres/valkey, no lane mutation), since the lane
overlays disable the in-lane Infisical service via *-disabled profile overrides.
P1.2b-A: Infisical now reachable from the stability runtime/effect containers; a
known secret (OMN_12966_PROBE) seeds and resolves end-to-end via the machine-
identity universal-auth path, verified from inside the effects container.
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* feat(OMN-12958): volume-config drift gate + runtime config provenance (#1941)
* feat(OMN-12958): volume-config drift gate + runtime provenance
Compute config provenance (path + sha256) for the runtime-rendered Bifrost
delegation contract; the deployed volume copy survives rebuilds and silently
diverges from packaged source (two competing authorities, OMN-12945).
- runtime/config_provenance.py: ModelConfigProvenance + drift classification,
sidecar JSON writer (read by sweep + proof packets)
- runtime/health/health_config_provenance.py: drift -> degraded health
- render entrypoint logs provenance line + writes sidecar on every boot
- docs/runbooks/volume-config-drift-and-reseed.md: ledgered re-seed procedure
- validation exemption for config_name (logical identifier, not entity ref)
No live volume mutation: re-seed is an operator deploy step (deploy_pending).
* test(OMN-12958): cover volume config drift reseed flow
---------
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* feat(OMN-12957): wire runtime-profile validator + core registry parity guard (#1945)
* feat(OMN-12957): wire runtime-profile validator + core registry parity guard
- Remove the import-time RuntimeError drift raise in runtime_profile.py (a hard
raise on core/infra version skew would crash the kernel at import); the parity
invariant is enforced by test_profiles_match_core_registry instead.
- Add tests: _PROFILES keys == omnibase_core REGISTERED_RUNTIME_PROFILES, and
every CONSUMER_ATTACHED_RUNTIME_PROFILES profile loads as a real lane.
- Wire omnibase_core.validation.validator_runtime_profiles as a pre-commit hook
+ validator-runtime-profiles.yml CI gate on infra contracts.
- Freeze 19 pre-existing violators in validation/runtime_profiles_allowlist.yaml
(discovered by repo-root walk; drain via OMN-12982). Blocks NEW orphans.
Requires the omnibase_core pin to include OMN-12957's validator (new rules).
Evidence-Ticket: OMN-12957
Evidence-Source: 5463fbaf819409d4fb7f491dd4f276f10d869eba
* ci(OMN-12957): pass runtime profile allowlist to validator
* test(OMN-12957): cover runtime profile registry parity
* fix(OMN-12957): keep runtime profile allowlist under config
* fix(OMN-12957): pin core runtime profile registry
---------
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* fix(OMN-12973): widen migration-gate healthcheck start_period + ratchet (#1950)
P2.8: classify the prod migration-gate UNHEALTHY-then-self-resolved incident.
Classification: idle-one-shot-mis-modeled = NO. The migration-gate is a
correctly-modeled long-running sentinel (entrypoint `while true; sleep 3600`)
whose healthcheck continuously polls db_metadata.migrations_complete via
check_migrations_complete.sh. The container flipped UNHEALTHY transiently because
its healthcheck start_period (10s) was far shorter than the real cold-volume
migration window (~116s: prod gate started 09:35:22, intelligence-migration
finished 09:37:18). Past the 10s grace window the still-failing probe was
reported UNHEALTHY until migrations completed, then self-resolved — no fault.
Fix: raise migration-gate healthcheck start_period 10s -> 180s in both the
authoritative catalog manifest (docker/catalog/services/migration-gate.yaml,
flows into the generated compose) and the hand-maintained
docker/docker-compose.infra.yml that deploy-runtime.sh applies to .201, so a
still-applying gate stays in `health: starting` instead of flipping UNHEALTHY.
Ratchet (enforcement, not detection): new ValidatorHealthcheckStartPeriod in
omnibase_infra catalog (validator_healthcheck_start_period.py) asserts a 120s
floor for migration-completion gates, wired into `onex validate runtime`
(cmd_validate_runtime) AND backed by unit tests that gate every PR via
pre-commit + CI. A migration-completion gate is identified by BOTH a healthcheck
that polls migration completion AND a service_completed_successfully dependency,
so ordinary app services (e.g. intelligence-api, an HTTP liveness probe with a
legitimately short 40s start_period) are not swept into the floor.
Prod probed read-only only; no prod mutation. Applies to live prod via the
batched stability/prod rebuild (deploy_pending).
Evidence-Ticket: OMN-12973
Evidence-Source: OCC#PENDING
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* feat(OMN-12971): deliver Vertex ADC bearer token to runtime effects container (secret-ref discipline) (#1949)
* feat(OMN-12971): deliver Vertex ADC bearer token to runtime effects container (secret-ref discipline)
Wire the Vertex $500-credit ADC path into the runtime, ADDITIVE next to the
Gemini API-key path (provider-agnostic; neither provider removed or forced).
runtime_policy.contract.yaml: add a secret-source mapping llm.vertex.access_token
(source_type=file, /run/secrets/vertex_access_token) to the dev/stability-test/
judge profiles, alongside the existing llm.gemini.api_key env mapping. The ref
name MUST match cloud-vertex-gemini.secret_ref in omnimarket
bifrost_delegation.yaml. The resolved VALUE is a short-lived OAuth bearer token
minted from ADC, refreshed by the operator; the token VALUE is never committed —
only the ref name + in-container path. Add aiplatform.googleapis.com to the
cloud host allowlist.
docker-compose.infra.yml: bind the operator-supplied host token file read-only to
/run/secrets/vertex_access_token on the main and effects runtimes
(VERTEX_ACCESS_TOKEN_HOST_FILE, default /dev/null so lanes without Vertex still
start; Gemini key path unaffected). Pass through BIFROST_VERTEX_GEMINI_ENDPOINT_URL
(overlay supplies the complete Vertex OpenAI-compat URL) and
GOOGLE_CLOUD_PROJECT/LOCATION (default empty).
runtime-policy.env: regenerated from the contract via render_runtime_policy_env
(test_runtime_policy_env_matches_contract_renderer proves contract<->env parity).
test_runtime_policy_contract.py: update host-allowlist assertion for the additive
Vertex host.
Pre-existing platform-wide topic-parity-gate failures (25 unrelated topics) are
identical with and without this change (proven by diff); that gate is pre-commit-
only (not a CI merge gate) and this change adds zero new topic gaps, so that one
hook is SKIP-ped. No deploy/receipt/merge gate is bypassed.
* fix(OMN-12971): make Vertex runtime env contract-owned
---------
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* feat(OMN-13008): automated .201 disk maintenance — worktree GC + docker GC + watermark alert ratchet (#1952)
* feat(OMN-13008): automated .201 disk maintenance — worktree GC + docker GC + watermark alert ratchet
Conservative, keep-list-driven disk maintenance to prevent the 2026-06-11
/data ~95% outage that killed all three lanes mid-demo:
- scripts/worktree-gc.sh: drives the canonical omniclaude prune-worktrees.sh
(merged+clean+pushed safety) on both Mac (merge-sweep tick) and .201
- scripts/disk-gc.sh + disk_gc_plan.py: conservative docker/builder/image GC.
Pure, testable removal planner honoring a VERSIONED keep-list
(deploy/disk-gc/keep-list.yaml): never reaps a kept repo, kept tag, in-use
image, or anything younger than min_age_days; keeps N superseded generations.
- scripts/disk-watermark-check.sh + disk_watermark_event.py: df watermark
ratchet. >=85% emits a typed disk-watermark bus event (warning) that the
sweep auto-ticket path turns into a Linear ticket; >=90% emits critical.
Broker addr is fail-fast from KAFKA_BOOTSTRAP_SERVERS (no localhost default).
- deploy/disk-gc/: systemd USER timer (onex-disk-gc.timer/.service, hourly) +
install-disk-gc.sh. User units, NOT lane containers.
- tests: 20 unit tests incl. GC plan-safety invariants + dry-run proof that
default mode issues no destructive op (a wrong-delete GC is worse than none).
Contract: contracts/OMN-13008.yaml
* fix(OMN-13008): pass docker inventory to GC planner via stdin, not env (ARG_MAX)
On a host with many docker images, passing the full image/ps inventory as env
vars to disk_gc_plan.py exceeds ARG_MAX ('Argument list too long', exit 126),
producing an empty plan. Write inventory to per-run scratch files (under the log
dir, never /tmp; cleaned on exit) and hand it to the planner on stdin as a JSON
envelope. Verified the failure live on .201; planner now reads stdin.
* fix(OMN-13008): simplify GC plan stdin pipe (two processes, no nested subprocess)
* fix(OMN-13008): keep-wins reconciliation — never remove an image id with any keep reason
A single image id can surface in multiple 'docker image ls' rows (one per
repo:tag). One tag could route the id to dangling-removal while another routes
it to keep (e.g. tagged 'latest' or within-N-generations). Live .201 plan showed
an id in BOTH remove_image_ids and kept_reasons. Reconcile at the end: keep wins
— any id with a keep reason is dropped from the remove list; remove list deduped.
Adds 2 regression tests. Verified live on .201.
* fix(OMN-13008): timer uses OnCalendar=hourly for reliable re-arm (oneshot service)
OnUnitActiveSec does not reliably re-elapse for a oneshot service once it goes
inactive (observed NextElapseUSecMonotonic=infinity live on .201 after the first
run). Switch to OnCalendar=hourly + RandomizedDelaySec so the timer re-arms every
hour. Keep Persistent=true for missed-run catch-up.
---------
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* fix(OMN-13005): materialize blocking event_consumer in runtime auto-wiring (runner consume-leg degenerate rows) (#1951)
* fix(OMN-13005): materialize blocking event_consumer in auto-wiring (runner consume-leg degenerate rows)
The runtime auto-wiring materialized event_publisher for handlers that
declare it but had no equivalent for event_consumer. Request/response
EFFECT handlers (HandlerContextRoiRunner) that publish a command then
block on the correlated terminal event fell back to their no-op consumer
default, returning None immediately -> every result row degenerate
(failure_stage=generation, attempt_count=0) while generations succeeded
~1s later.
Adds _make_sync_event_consumer (mirror of _make_sync_event_publisher),
backed by service_terminal_event_consumer.make_terminal_event_consumer:
a sync (topic, correlation_id, timeout) -> dict | None adapter that runs
the proven direct-Kafka correlate-and-wait loop (from RuntimePatternBBroker)
on an isolated event loop in a worker thread, so blocking does not deadlock
the runtime dispatch loop that delivers the awaited terminal.
TDD through the REAL dispatch path: test_event_consumer_injection drives a
trial through _prepare_handler_wiring with a terminal arriving after a delay
and asserts a non-degenerate row; verified RED with injection disabled.
* fix(OMN-13012): two-phase (seek-now/wait-later) terminal event_consumer to close the subscribe-after-publish race (#1954)
The OMN-13005 injected event_consumer is a single callable that does
assign -> seek_to_end -> poll internally, all AFTER the handler has
already published its command. Once OMN-13010 freed the dispatch loop and
generation began completing in ~1s, the correlated terminal lands BEFORE
the single-call consumer's post-publish seek_to_end positions, so
seek_to_end skips PAST the already-emitted terminal and the runner times
out on an offset beyond it (probe3, run_id=20260611T2140Z-probe3 -- both
arms degenerate, zero rebalances).
Splits positioning from waiting so the caller subscribes BEFORE it
publishes:
session = consumer.open(topic) # assign + seek_to_end NOW
publisher(command_topic, payload) # publish AFTER positioning
payload = session.wait(cid, timeout) # block from the captured position
The returned TerminalEventConsumer is still directly callable with the
legacy (topic, cid, timeout) -> dict | None single-call shape for any
consumer that does not need subscribe-before-publish; the runner is the
only consumer today. TerminalConsumerSession owns a dedicated event loop
on a daemon worker thread for the whole open->wait->close lifecycle,
preserving the OMN-13005 loop-isolation discipline so blocking never
deadlocks the runtime dispatch loop.
TDD (real dispatch path, RED-then-GREEN): extends the OMN-13005 injection
test with a terminal emitted IMMEDIATELY after publish. The single-call
(seek-after-publish) consumer MISSES it (degenerate row -- RED test
asserts failure_stage=generation); the two-phase (open-before-publish)
consumer CATCHES it (non-degenerate -- GREEN). The Kafka layer is faked at
the two service seams against a shared in-memory log modeling seek-to-end
semantics. OMN-13005 blocking-correlate behavior preserved. 269/269
auto_wiring unit tests pass; mypy --strict clean.
Sibling to OMN-13010 / OMN-13005 / OMN-13003.
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* fix(OMN-13005): route terminal consumer through Kafka boundary
---------
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* fix(OMN-12990): pin worker replicas in ledgered config + fail-fast lane overrides (silent-drop ratchet) (#1948)
* fix(OMN-12990): pin worker replicas in ledgered config + fail-fast lane overrides
The base compose sets runtime-worker deploy replicas to ${WORKER_REPLICAS:-0}
(soft-default ZERO). The stability lane's required state includes a running
worker (GATE_ZERO_PROOF.md: 4 runtime containers), but the worker presence rode
entirely on a compose soft-default — any plain compose up/recreate without the
policy env silently scaled the worker to zero with no error and no signal.
Fix (contract-native + fail-fast):
- Add 'replicas' to ModelRuntimeProcessPolicy; pin replicas: 1 in every lane's
worker block in runtime_policy.contract.yaml.
- Renderer emits {PROFILE}_WORKER_REPLICAS into the ledgered runtime-policy.env
for dev/stability-test/judge/prod.
- stability + prod compose overrides reference ${..._WORKER_REPLICAS:?...}
(fail-fast, NO silent :-1/:-0 default). A recreate that omits the policy env
now aborts loudly instead of dropping the worker. prod previously had no
override at all and inherited the dangerous :-0 default.
Ratchet (recurrence guards):
- tests asserting fail-fast override form (no soft default), contract-declared
replica pin >= 1 per lane, and rendered {PROFILE}_WORKER_REPLICAS=1 in the
ledgered env.
- runbook deploy/verify procedure adds an expected-container census (worker
must be present) via verify_container_manifest; a missing worker is a FAILURE,
not silence.
Note: SKIP=topic-parity-gate — that local-only advisory gate (absent from all
.github/workflows, not a required CI check) fails on 25 pre-existing cross-repo
topic gaps (build-loop/omniclaude/omniweb) identical on pristine base HEAD
8d7da1249; this change adds zero topics. All other hooks ran clean.
Evidence-Ticket: OMN-12990
* test(OMN-12990): cover worker replica policy integration
---------
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* feat(OMN-12909): add gateway bus forwarder P0A (#1946)
* feat(OMN-12909): add gateway bus forwarder p0a
* test(OMN-12909): add gateway forwarder integration coverage
* fix(OMN-12909): satisfy gateway forwarder validators
* test(OMN-12909): allow gateway forwarder bus protocol
* fix(OMN-12909): sync gateway forwarder entry point
* fix(OMN-12909): refresh runner image identity lock
* test(OMN-12909): relax JSON normalizer mixed benchmark threshold
* fix(OMN-12909): allow gateway handlers to boot unconfigured
* fix(OMN-12909): declare gateway forwarder runtime profile
* fix(OMN-12909): update runner identity backmerge expectation
---------
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* feat(OMN-13011): LANE CENSUS RECONCILIATION ratchet — declared desired-state per lane, drift = auto-ticket (#1955)
The class fix for the recurring lane-drift regression. Nothing reconciled the
declared desired state of a runtime lane against what is actually running, so the
same failure kept recurring with zero signal: volume config drift (OMN-12945),
WORKER_REPLICAS silent zero (OMN-12988/12990), and on 2026-06-11 prod runtime
containers plus the broker network were silently absent for hours during demo prep.
Ships a per-lane DESIRED-STATE census:
- (a) DECLARED in a versioned lane manifest (deploy/lane-census/lane-manifest.yaml):
container set, network, replicas, image-tag pattern per lane
(stability-test/prod/judge/dev), derived from the canonical compose lane files.
A parity ratchet keeps the manifest locked in step with the compose files.
- (b) RECONCILED on a schedule on .201 by SHARING the OMN-13008 systemd timer
(a drop-in 4th ExecStart on onex-disk-gc.service — never a second timer) and
on-demand via scripts/lane-census-check.sh / runtime_sweep.
- (c) Drift = typed bus event (onex.evt.infra.lane-census-drift.v1) + Linear
auto-ticket naming exactly what is missing/extra (container_absent,
network_detached, replicas_zero, unexpected_container, oneshot_failed/stuck,
image_tag_mismatch). Fail-fast, no warn-only mode (gates-block policy); exit 30
on drift; bus publish fail-fast on KAFKA_BOOTSTRAP_SERVERS (no localhost default).
Red fixture reproduces 2026-06-11: prod runtime containers absent + broker network
detached must produce the exact drift findings + a non-zero exit hours before a
human noticed. Pure planner is fully unit-tested; shell driver dry-run-tested.
Builds on the OMN-12988 deploy-agent RUNTIME census (deploy-time) as the
complementary steady-state reconciler; closes the runtime-worker.yaml
container_name: null census gap by sourcing names from the compose lane files.
Evidence-Ticket: OMN-13011
Config-drift family: OMN-12945
Relates-to: OMN-13009, OMN-12988, OMN-13008
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* fix(OMN-13020): vendor missing node migrations — llm_routing 0000 + context_roi 001 (#1956)
Vendors two omnimarket node-source migrations into the infra forward-migration
tree via scripts/sync-node-migrations.sh (the canonical OMN-12559 mechanism):
- node_projection_llm_routing/0000_create_llm_routing_decisions.sql
(source: omnimarket #1168 / OMN-12942, merge ed6734f8)
- node_projection_context_roi/001_create_context_roi_scores.sql
(source: omnimarket #1178 / OMN-12955, merge 5010b1f4)
Without the 0000 base table, node_projection_llm_routing/0001 (CREATE VIEW)
hard-fails against NODE_POSTGRES_DB=omnidash_analytics — exactly the prod
forward-migration exit-3 of 2026-06-11T09:35:52Z, and reproduced by
construction in any clean clones@dev build. Files are byte-identical to the
omnimarket dev blobs (sha256 f8a8b339… / c4126e65…) and to the untracked
hot-patch copies on the .201 stability clone.
Both migrations are self-contained, all-statements-IF-NOT-EXISTS, and 0000
sorts lexically before 0001 within the node's namespaced identity space.
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* fix(OMN-13058): close TerminalConsumerSession on open() failure (worker thread + event loop leak) (#1957)
TerminalConsumerSession.__init__ starts its dedicated worker loop thread
immediately. TerminalEventConsumer.open did
'session = TerminalConsumerSession(...); return session.open()' with no
cleanup: any failure inside session.open() (consumer start timeout,
partition-assign timeout, broker auth error) propagated out of the raising
expression, the session reference was lost, and the daemon worker thread plus
its never-closed asyncio event loop leaked -- one pair per failed open. The
motivating caller (HandlerContextRoiRunner) opens a session per trial, so a
160-560-trial battery against a degraded broker accumulates hundreds of
leaked threads in the long-lived effects container.
Fix: wrap session.open() in try/except BaseException -> session.close()
(idempotent: stops the loop, joins the thread) -> re-raise. Covers both the
two-phase .open(topic) path and the legacy single-call __call__ path.
Found by the P3.3 doctrinal review of merged #1951 (b9712af9 / 36d98275).
TDD: tests/unit/runtime/test_service_terminal_event_consumer_open_failure.py
injects a real open failure through the production path (event bus without
_bootstrap_servers) and asserts no alive terminal-consumer-* thread after the
raise. Verified RED with the fix stashed (2 failed), GREEN with it (2 passed).
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* feat(OMN-13021): non-dev-base guard — fail feature-base PRs absent Stacked-Parent declaration (retro A-6) (#1958)
Any PR whose base is neither dev nor main fails unless the body carries
'Stacked-Parent: #N'. Prevents the feedback_stacked_prs_orphan_from_dev class
(#1185/#1954 auto-merged INTO parent feature branches and stranded off dev).
base=main remains governed by main-target-guard.
Epic OMN-13013 (process enforcement ratchets — June 12 retro).
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* feat(OMN-13014): hot-patch ledger rebuild preflight gate (retro B-1) (#1959)
* feat(OMN-13014): hot-patch ledger rebuild preflight gate (retro B-1)
Hot-patches on .201 (.prepatch sibling discipline) silently revert on any
image rebuild/force-recreate — the 2026-06-11 20:58Z rebuild already erased
a live /api/generate patch once. This adds the rebuild-path gate:
- scripts/preflight_hotpatch_ledger.py: given a target container or lane +
per-repo build refs, hard-fails when any hot-patch ledger row's source PR
merge commit is not an ancestor of the build ref (git merge-base
--is-ancestor), plus a .prepatch tripwire of the running container
(unledgered .prepatch = hard fail; --post-rebuild = zero .prepatch
expected). Sole bypass: HOTPATCH_PREFLIGHT_BYPASS carrying the Rule-10
'# skip-token-allowed: <user-approval-receipt-id>' form.
- scripts/deploy-runtime.sh: guard_hotpatch_ledger wired into main() before
build/preview (both dry-run and execute), lane derived from the compose
project; skips loudly only when no ledger exists on the host.
- tests/unit/scripts/test_preflight_hotpatch_ledger.py: 17 unit tests
(ancestor gate, lane scoping, ledger loading, tripwire, bypass forms).
- tests/ci/test_receipt_gate_install_guard.py: repair stale guard — core
OMN-12565 replaced the OMN-9198 'uv pip uninstall first' install step with
a cleared workspace venv (uv venv --clear); assert the new contract.
Ledger backfilled from live census (5 .prepatch files / 4 source PRs, all
MERGED to dev) at /data/omninode/hotpatch-ledger/ledger.yaml on .201.
* fix(OMN-13014): scope missing-prepatch tripwire warning to the probed container
---------
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* fix(OMN-12977): workspace-build sibling-pin ratchet — honor consuming lock, fail-fast on stale vendor (#1936)
* fix(OMN-12977): workspace-build sibling-pin ratchet — honor consuming lock, fail-fast on stale vendor
The workspace-mode image build (BUILD_SOURCE=workspace, used by the stability-test
deploy procedure) vendored sibling/foundation packages from whatever the canonical
OMNI_HOME clones happened to be checked out at, ignoring the consuming repo's
uv.lock. On 2026-06-11 this shipped a 13-day-stale omnibase_infra 0.37.0-dev
(pre-OMN-12501 Protocol-quarantine guard) + core 0.42.0 against an omnimarket dev
lock pinning infra 0.38.1@e2dbdc95 / core 0.44.0@c97c2c9a, dropping the guard and
crashing wire_from_manifest bootstrap fatally (stability lane down on demo day).
Fix + recurrence ratchet (same PR):
- scripts/runtime_build/check_sibling_lock_pins.py: parse the consuming repo's
uv.lock for expected version+git-rev of each foundation/sibling package
(scoped to the package's own source line so editable/registry pins are not
cross-attributed a dependency's rev), resolve the actual clone version+HEAD,
compare, and classify drift backward/forward/none. Fail-fast (exit 1) on any
drift; --allow-drift records an explicit operator override in the artifact,
never silent.
- stage_workspace.sh: runs the preflight against the canonical clones before
staging; aborts the build (exit 3) on unacknowledged drift and writes
workspace/sibling-pin-comparison.json.
- compute_workspace_provenance.py: folds the expected-vs-actual comparison into
build-provenance.json so deploy verifiers can assert the build honored the lock;
flags unacknowledged drift as a provenance error.
- Dockerfile.runtime: COPY the comparison artifact (committed placeholder so the
COPY always resolves; overwritten by stage_workspace.sh in workspace mode).
- TDD: 19 unit tests covering lock parsing (git/registry/editable sources),
drift classification, the exact 0.37.0-vs-0.38.1 stale case, check_pins exit
codes, and the allow-drift override. Pre-existing mypy-strict bare-dict errors
in compute_workspace_provenance.py fixed in the same pass.
Evidence-Ticket: OMN-12977
Evidence-Source: pending-occ
* ci(OMN-12977): retry runtime smoke compose port race
* test(OMN-12977): align sibling-pin script tests with current API
---------
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* fix(OMN-12989): workspace-mode image build must honor sibling lock pins (#1947)
* fix(OMN-12989): workspace build must honor sibling lock pins (fail-fast ratchet)
The 2026-06-11 stability bootstrap crash was caused by a workspace-mode
--no-cache rebuild that vendored omnibase_infra 0.37.0-dev (13-day-stale
worktree) even though omnimarket dev's uv.lock pins infra 0.38.1. The
downgraded sibling predated the OMN-12501 Protocol-quarantine guard and
turned a latent contract defect into a fatal crash.
Fix + ratchet (same PR set):
- scripts/runtime_build/resolve_workspace_pins.py: parse the consuming repo's
uv.lock for sibling pins (version + git rev); classify each staged/installed
sibling as exact/ahead/regression/unpinned; fail-fast (WorkspacePinError) on
any regression BELOW the lock pin. Stdlib version-tuple fallback when
packaging is absent so the ratchet never fails-open.
- compute_workspace_provenance.py: enforce sibling pins + a host-infra
self-check (installed omnibase_infra vs lock pin — the exact crash vector,
since host infra is built from the context, not staged), and emit a
pin_comparison block into build-provenance.json for deploy verifiers.
- Dockerfile.runtime: COPY resolve_workspace_pins.py beside the provenance
script so the in-image import resolves.
- TDD: failing tests first (test_resolve_workspace_pins.py, 11 cases) +
provenance integration tests; fixed a pre-existing stale RUNTIME_VERSION
assertion to read pyproject dynamically.
Evidence-Ticket: OMN-12989
* fix(OMN-12989): workspace build must honor sibling lock pins (fail-fast ratchet)
The 2026-06-11 stability bootstrap crash was caused by a workspace-mode
--no-cache rebuild that vendored omnibase_infra 0.37.0-dev (13-day-stale
worktree) even though omnimarket dev's uv.lock pins infra 0.38.1. The
downgraded sibling predated the OMN-12501 Protocol-quarantine guard and
turned a latent contract defect into a fatal crash.
Fix + ratchet (same PR set):
- scripts/runtime_build/resolve_workspace_pins.py: parse the consuming repo's
uv.lock for sibling pins (version + git rev); classify each staged/installed
sibling as exact/ahead/regression/unpinned; fail-fast (WorkspacePinError) on
any regression BELOW the lock pin. Stdlib version-tuple fallback when
packaging is absent so the ratchet never fails-open.
- compute_workspace_provenance.py: enforce sibling pins + a host-infra
self-check (installed omnibase_infra vs lock pin — the exact crash vector,
since host infra is built from the context, not staged), and emit a
pin_comparison block into build-provenance.json for deploy verifiers.
- Dockerfile.runtime: COPY resolve_workspace_pins.py beside the provenance
script so the in-image import resolves.
- TDD: failing tests first (test_resolve_workspace_pins.py, 11 cases) +
provenance integration tests; fixed a pre-existing stale RUNTIME_VERSION
assertion to read pyproject dynamically.
Evidence-Ticket: OMN-12989
* test(OMN-12989): co-locate provenance pin helper in fixture
---------
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* fix(OMN-13055): absent repos in REPOS list warn and exit 0 instead of failing (#1960)
Missing repos (not cloned locally) now emit a WARN line and are tracked
in a separate WARNED array. Only real fetch/ff failures cause exit 1.
This makes pull-all.sh safe to use on machines with a partial clone set,
while keeping the explicit-list override behavior intact.
Adds three regression tests: absent-only exits 0, absent+present exits 0
with OK for the present repo, present-failed+absent exits 1.
Evidence-Ticket: OMN-13055
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* fix(OMN-13070): config_prefetcher overlay wins on controlled lanes (#1961)
On infisical_required=True lanes, fetched/overlay config now always wins
over ambient env. Ambient env is retained only as a declared bootstrap
fallback (with an explicit provenance INFO log line) when Infisical
returns None. apply_to_environment also overwrites stale env on controlled
lanes. Uncontrolled lane (infisical_required=False) behaviour is unchanged.
Adds 5 regression tests: controlled-lane Infisical-wins, env-bootstrap-
fallback, apply_to_environment overwrite, missing-from-both-is-error, and
uncontrolled-lane-env-still-wins. Refactors _resolve_key to return a
(outcome, value, error) tuple to satisfy the ≤5-param pattern gate.
Source: docs/audits/2026-06-10-runtime-env-overlay-authority-audit.md
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* fix(OMN-13062): migration-gate vacuity fix — sentinel discipline, wait-for-postgres, skip-manifest, sync-check nonzero (retro A-10) (#1963)
* fix(OMN-13062): migration-gate vacuity fix — sentinel discipline, wait-for-postgres, skip-manifest, sync-check nonzero
Fixes three bugs identified in retro A-10 (recurrences OMN-12885, OMN-12934):
(1) RUNNER SENTINEL DISCIPLINE
run-forward-migrations.sh now clears migrations_complete=FALSE at the
start of every run and sets it TRUE only as its FINAL act after all
infra and node migrations succeed. Any mid-run failure leaves the gate
UNHEALTHY. runner_completed_at is stamped at the same final step as
durable evidence of a successful completion run.
(2) SYNC-NODE-MIGRATIONS VACUOUS GATE
sync-node-migrations.sh --check now exits 2 (not 0) when the omnimarket
source tree is unresolvable. Silent exit-0 was hiding drift. The single
opt-out is SYNC_NODE_MIGRATIONS_SKIP_UNRESOLVABLE=1 for environments
that intentionally run without the source.
(3) WAIT-FOR-POSTGRES GUARD
run-forward-migrations.sh now waits up to PG_WAIT_RETRIES (default 30)
x 2s for Postgres to accept connections before proceeding, guarding
the first-boot initdb race.
(4) SKIP-MANIFEST
docker/migrations/skip-manifest.yaml introduced as the sole committed
escape for intentionally-skipped migrations. The runner reads this at
startup; listed migrations are recorded in schema_migrations with
checksum "skip-manifest" without executing the SQL.
(5) MIGRATION 085
Adds runner_completed_at TIMESTAMPTZ column to db_metadata so the
runner's final stamp is durable in the schema (idempotent ADD COLUMN IF
NOT EXISTS). Rollback included.
22 regression tests added covering all five fix surfaces.
* fix(OMN-13062): stamp schema fingerprint for migration 085
Migration 085 (085_add_runner_completed_at_to_db_metadata.sql) was added
in the initial commit but schema_fingerprint.sha256 was not regenerated.
Running `python scripts/check_schema_fingerprint.py stamp` updates the
artifact from the stale hash to match the 71 migration files.
Evidence-Source: OCC#2563
Evidence-Ticket: OMN-13062
---------
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* feat(OMN-12864): Bifrost endpoints → committed overlay authority + fail-loud loader (#1964)
* feat(OMN-12864, OMN-12814, OMN-12945): Bifrost endpoints → committed overlay authority + fail-loud loader
OMN-12864 — Committed lane overlay
- docker/lane-overlays/dev.bifrost.yaml: typed deployment bindings for all
four BIFROST_LOCAL_*_ENDPOINT_URL values (coder :8000, reasoner :8001,
embedding :8100, ds4-flash :8101). Previously only available as ephemeral
shell exports on .201; now committed, auditable, diff-able, and CI-checked.
- docker/lane-overlays/dev.bifrost.env: generated dotenv sidecar consumed by
compose via env_file; never edited directly (yaml is authority).
- docker/docker-compose.infra.yml: wire the env_file block at compose root so
the four endpoints are injected into the interpolation context on a clean
shell. Hardcode BIFROST_CONTRACT_PATH (remove :-/empty footgun — OMN-12814).
- scripts/render_bifrost_lane_overlay_env.py: render script regenerates the
env sidecar from the YAML source.
- src/omnibase_infra/runtime/models/model_bifrost_lane_overlay.py:
ModelBifrostLaneOverlay — typed Pydantic model enforcing URL completeness
(OMN-12815: every URL must end in /chat/completions).
OMN-12814 — Fail-loud loader
- render_bifrost_delegation_contract: raises ProtocolConfigurationError on
FileNotFoundError, YAMLError, ValidationError, and zero-endpoint renders.
No lru_cache — every restart re-renders from packaged source so a stale
cache cannot pin a broken result across deploys.
OMN-12945 — Re-seed from packaged source on deploy
- docker/entrypoint-runtime.sh: set BIFROST_FORCE_RESEED=1 on every container
restart so the named-volume copy is always rebuilt from the packaged
bifrost_delegation.yaml merged with committed lane-overlay endpoints.
- render_bifrost_delegation_contract: honor BIFROST_FORCE_RESEED/force_reseed
flag to bypass the stale-volume early-return path entirely.
Tests:
- tests/ci/test_bifrost_lane_overlay.py: CI gate — env sidecar in-sync with
YAML source; all four BIFROST_LOCAL_* keys present.
- tests/unit/runtime/models/test_model_bifrost_lane_overlay.py: bare-base URL
rejection, env dict mapping, extra-field rejection.
- tests/unit/runtime/test_render_bifrost_delegation_contract.py: fail-loud
paths, force-reseed, zero-endpoint error, endpoint URL completeness.
- tests/unit/models/test_model_serialization_roundtrip.py: roundtrip coverage.
* fix(OMN-12864): move bifrost env_file to service level — fix compose schema validation failure
Top-level 'env_file' is rejected by Docker Compose v2 schema validator
('additional properties not allowed'). This caused 10+ compose-render
integration tests to fail in CI.
Fix:
- Remove top-level env_file block from docker-compose.infra.yml
- Add per-service env_file on omninode-runtime, runtime-effects,
runtime-worker (the three containers that render Bifrost)
- Change BIFROST_LOCAL_*:? to BIFROST_LOCAL_*:- in x-runtime-env
(compose-level validation removed; Python validates via
ModelBifrostLaneOverlay + render_bifrost_delegation_contract)
- Add two CI gate tests: compose_env_file_is_service_level_not_top_level
and runtime_services_have_bifrost_env_file
* fix(OMN-12864): allow BIFROST_LOCAL_* empty defaults in silent-fallback gate
The overlay authority pattern (OMN-12864) passes BIFROST_LOCAL_*_ENDPOINT_URL
via service-level env_file (docker/lane-overlays/dev.bifrost.env), not at
compose config time. Compose-level :? would break CI rendering without the
overlay pre-loaded. Validation at the Python layer (ModelBifrostLaneOverlay
+ render_bifrost_delegation_contract) is the enforcement point.
Add the three failing vars to ALLOWED_EMPTY_DEFAULTS with OMN-12864 citation.
---------
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* feat(OMN-12858/OMN-12879/OMN-12880): dispatcher route-coverage CI gate (#1965)
* feat(OMN-12858/OMN-12879/OMN-12880): dispatcher route-coverage CI gate
Static contract analysis: every subscribed command topic (onex.cmd.*)
must declare handler_routing or runtime_dispatch, or the message goes to
DLQ silently. This gate would have caught two recent incidents:
1. June 9 DLQ regression (OMN-12858 post-mortem): node_generation_consumer
subscribed onex.cmd.omnimarket.node-generation-requested.v1 but a
sole-handler revert left zero dispatcher routes registered. Messages
went to DLQ silently with no CI signal.
2. June 12 DEL-01 live finding: onex.cmd.omnimarket.delegate-skill.v1 was
consumed by dev lane bus but no dispatcher route existed in any deployed
contract. Discovered via manual rpk consumer-group lag probe (DEL-01
evidence, docs/evidence/2026-06-12-weekend-pass/).
Deliverables:
- scripts/check_dispatcher_route_coverage.py — static YAML scanner that
checks both omnibase_infra and omnimarket contract trees; ratchet
allowlist for known pre-existing violations; --changed-contracts mode
(OMN-12879) for per-PR scoping; compat publish topics excluded (OMN-12880)
- .github/workflows/dispatcher-route-coverage.yml — CI workflow that
checks out omnimarket sibling, collects changed contract paths in PR
mode, and runs the gate; fires on PR, push-to-main, and merge_group
- tests/ci/test_dispatcher_route_coverage_gate.py — 12 unit tests
covering RED/GREEN/COMPAT/CHANGED-MODE/ALLOWLIST/MULTI-DIR paths plus
live-contract regression proof against the actual omnibase_infra tree
Allowlist additions:
- onex.cmd.omnibase-infra.pattern-b-dispatch.v1 (RuntimePatternBBroker,
imperative consumer, OMN-12525 migration target)
- onex.cmd.platform.contract-resolve-requested.v1 (transitional HTTP
bridge node_contract_resolver_bridge OMN-2756, metadata.transitional=true)
[OMN-12858, OMN-12879, OMN-12880]
* fix(OMN-12858): drop full uv sync from dispatcher-route-coverage workflow
Gate script only needs pyyaml (stdlib + yaml). Using full setup-python-uv
was causing 10+ minute timeout. Replace with direct pip install pyyaml and
invoke python3 directly. Reduces job from 10m timeout to <1m.
[OMN-12858]
---------
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* feat(OMN-13026): wire transport-mock-lint pre-commit hook + CI gate (#1962)
* feat(OMN-13026): wire transport-mock-lint pre-commit hook + CI gate (omnibase_infra)
Activates the transport-mock-lint validator (from omnibase_core, OMN-13026)
on omnibase_infra. Ratchet baseline: 218 existing violations across 63 files
frozen in validation/transport_mock_baseline.yaml. New bare AsyncMock/
MagicMock on EventBus/transport surfaces are blocked by pre-commit hook and
CI lint step. Existing violations tracked for drain by per-site tickets
(parent OMN-13026).
Reference incident: PR #1181 bare AsyncMock hid missing EventBusKafka.stop().
Evidence-Ticket: OMN-13026
* fix(OMN-13026): use uv run python for CI lint step + bump omnibase_core pin to include transport_mock_lint
The transport-mock lint CI step previously cloned omnibase_core and ran
`python -m omnibase_core.validators.transport_mock_lint` with PYTHONPATH,
but this failed: `No module named omnibase_core.validators.transport_mock_lint`
because it ran `.venv/bin/python` which uses the locked venv, and the venv
omnibase_core pin (2defabef4) predates the transport_mock_lint module.
Fix: use `uv run python` (removes the clone step) and bump omnibase-core git
pin from 2defabef4 to 309d89fa7 (PR 1231 merge commit on dev) so
transport_mock_lint is available in the locked venv.
* fix(OMN-13026): align transport mock baseline and runner lock
* fix(OMN-13026): sync omnibase_core pin + runner identity lock to dev baseline
Align pyproject.toml omnibase_core rev to 2defabef (required by
test_release_backmerge_preserves_proven_runtime_core_pin) and update
docker/runners/runner-image.lock.json identity_digest/shared_env_digest
to match dev runner image lock (79b08f44 / 90c8b3b9).
Both were stale from the prior session's pin bump that used an older SHA.
* fix(OMN-13026): source transport mock validator from core
* fix(OMN-13026): source transport validator from core dev
---------
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* feat(OMN-13094): --output receipt mode on onex node/run — quiet typed receipts with durable capture (#1966)
Phase 2a of the skill-output-suppression slice (epic OMN-13089, plan
docs/plans/2026-06-12-skill-output-suppression-plan.md Phase 2 item 1).
- onex node/onex run gain --output receipt: ALL runtime logging routes to a
run_id-suffixed capture file under <state-root>/captures/ (no console
handlers — kills the 25-50-line RuntimeLocal INFO stream at the source);
stdout carries exactly ONE typed ModelSkillResult JSON with the FULL
handler result (result_model = concrete handler result type FQN).
- Durable capture: capture log + handler result content-addressed via
omnibase_core ArtifactStore (OMN-13093); artifact.captured +
tool.output.captured emitted to the emit daemon socket (--emit-socket,
default ~/.claude/emit.sock).
- Failure asymmetry: artifact write failure => FULL output printed, no
receipt (no hidden loss); emission failure => receipt still prints, event
spooled to <state-root>/emit_spool/ for replay.
- Node failure => status=failed/error with full error + capture log INLINE
in the receipt (errors are never hidden) and artifact-backed.
- Default output mode unchanged (enforcement is Phase 4).
- RuntimeLocal exposes handler_result (receipt schema identity).
- core pin 2defabef -> ae8793bd (merged OMN-13091/13093 receipt models +
ArtifactStore); runner-image identity lock regenerated and the OMN-12765
backmerge identity constants updated for the new pin.
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* feat(OMN-13097): onex skill subcommand + declarative skill->node mapping (#1968)
* feat(OMN-13097): onex skill subcommand + declarative skill->node mapping
Phase 4a of the skill-output-suppression slice (epic OMN-13089, plan
docs/plans/2026-06-12-skill-output-suppression-plan.md Phase 4 item 1/2).
A dispatch skill IS one CLI call (user directive 2026-06-12). This adds the
`onex skill <name> [args]` dispatch surface that the 24 omniclaude shim
migrations build on:
- onex.cli entry-point `skill` -> cli_skill.run_skill_by_name. Resolves the
skill via the declarative skill_mapping.yaml registry, builds the backing
node's input payload from the skill's CLI args, writes it under the state
root (.onex_state/tmp/<skill>-<run_id>.json — never /tmp), resolves the
node's packaged contract exactly like `onex node`, and dispatches through
the proven receipt-mode path (run_receipt_mode, OMN-13094). stdout is
exactly one typed ModelSkillResult JSON with the FULL handler result.
- skill_mapping.yaml: declarative DATA mapping all 24 dispatch shims to their
backing onex.nodes node + typed result-model FQN (verified against each
live handler handle() return type on origin/dev) + per-arg payload specs +
static payload + keyword classifiers (delegate task_type as data, not code).
Adding a skill is a YAML edit + fixture, never a CLI code change (ticket
deliverable 2/3). Mapping lives beside the node-resolution surface, never
hardcoded branching in the CLI.
- Typed models split one-per-file (repo convention): ModelSkillArgSpec,
ModelSkillClassifier, ModelSkillMapping, ModelSkillMappingRegistry,
EnumSkillArgType. Frozen, extra=forbid, fail-fast coercion/validation.
- validation_exemptions.yaml: Click-callback param-count + literal-identifier
name-field exemptions mirroring the existing cli_node run_node_by_name
precedent (OMN-11570) — same pattern, same rationale.
dod_evidence:
- 20 unit tests pass (registry validity, all-24-shims coverage, FQN result
models, arg parsing/coercion/positional/required, classifiers, payload
build, receipt-mode dispatch wiring, payload-under-state-root not /tmp).
- uv run mypy src/ --strict: clean (2438 source files).
- ruff format + check: clean. pre-commit run on changed files: pass.
- runner-image identity lock regenerated for the pyproject entry-point add
(same as OMN-13094).
* test(OMN-13097): rebind OMN-12765 backmerge identity constants for onex skill pyproject change
Adding the `skill` onex.cli entry-point to pyproject.toml changes the
runner-image identity_digest (and shared_env_digest) the lock binds. Update
the hardcoded expected constants in the backmerge-identity test to the
regenerated values — same mechanical rebind OMN-13094 performed for the core
pin bump. Identity + runner-image-identity tests pass (14).
---------
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* fix+test(OMN-13012): force terminal-topic metadata refresh so both ephemeral consumers assign (runner consume-leg wedge) (#1969)
The runner consume-leg wedged on the live stability battery (image c0505521f1fa,
EXP1-3_RUNNER_CONSUME_LEG_BLOCKER): only the FAILED terminal topic ever subscribed
while the COMPLETED topic never assigned, so the correlated completed terminal was
never read and the 8x2x10 matrix re-fired cell 1 forever, emitting zero
non-degenerate rows. A prior two-strike diagnosis proved the omnimarket handler is
correct (it opens both terminal sessions pre-publish and races them); the defect is
in the omnibase_infra runtime consume leg.
Root cause: _assign_direct_terminal_partitions ignored the metadata future returned
by AIOKafkaClient.set_topics and re-called set_topics([same_topic]) each loop
iteration. aiokafka 0.13.0 set_topics only forces a metadata refresh when the topic
set DIFFERS from the tracked set, so every iteration after the first took the no-op
branch and never re-fetched. An ephemeral group_id=None consumer whose first metadata
fetch had not yet surfaced partitions burned the full 30s assign cap and raised a
bare TimeoutError (the empty-message 'wait failed' seen live).
Fix: register the reply topic once and await that metadata fetch, then on each miss
force a fresh fetch via force_metadata_update (which always fetches) rather than the
no-op set_topics repeat. The assign-cap TimeoutError now carries a diagnostic message
instead of an empty one.
Test: tests/integration/test_terminal_consumer_concurrent_assign_race.py drives the
REAL TerminalEventConsumer (the object wired as event_consumer) through the REAL
open_direct_terminal_consumer/poll path with AIOKafkaConsumer monkeypatched to a fake
that faithfully models aiokafka 0.13.0 set_topics future + metadata-latency semantics.
RED before the fix (bare TimeoutError, the live empty-message signature); GREEN after.
K>=2 multi-trial variant asserts no worker-thread leak across trials.
Evidence-Source: <occ-sha-pending>
Evidence-Ticket: OMN-13012
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* feat(OMN-13096): onex delegate single-command subcommand (Phase 2b) (#1967)
* feat(OMN-13096): onex delegate single-command subcommand
Add 'onex delegate "<prompt>" [--task-type X] [--max-tokens N]' as a
subcommand on the existing onex CLI (Phase 2b of the skill-output-suppression
slice, OMN-13089). The command wraps payload construction, node dispatch, and
result extraction internally and prints exactly one
ModelSkillResult[ModelDelegateSkillResponse] to stdout via the OMN-13094
receipt-mode path. RuntimeLocal logs go to the capture file + artifact store,
never to stdout; scratch payloads live under <state-root>/tmp/ with run_id
suffixes (never /tmp).
- cli_delegate.py: classify_task_type (keyword table from legacy skill md),
payload write, contract resolve, run_receipt_mode dispatch
- register 'delegate' under onex.cli entry points
- exempt delegate_command from the >5-param patterns gate (same Click-callback
rationale as run_node_by_name)
- 20 unit tests: classification, scratch-under-state-root, single typed
receipt on stdout, zero INFO log leakage
omnibase_infra does NOT depend on omnimarket; the delegate node is resolved at
runtime via the onex.nodes entry-point group (registered by omnimarket).
* chore(OMN-13096): re-trigger deploy-gate after Evidence-Source set to OCC#2593
No code change — the deploy-gate workflow triggers on synchronize (not edited),
so the PR-body Evidence-Source fix needs a new commit to re-resolve the OCC ref
to the open PR head where contracts/OMN-13096.yaml (with deploy evidence) lives.
* chore(OMN-13096): re-trigger deploy-gate now that OCC#2593 merged to OCC dev
contracts/OMN-13096.yaml (with the dod-deploy-onex-delegate item) is now on
OCC dev, so the deploy-gate OCC-dev checkout resolves the contract + deploy
evidence.
* chore(OMN-13096): regenerate runner-image identity lock for pyproject entry-point add
Adding the 'delegate' onex.cli entry point to pyproject.toml changed the
dependency-manifest digest that scripts/ci/runner_image_identity.py folds into
the runner-image identity lock. Regenerate the lock so
tests/ci/test_runner_image_identity.py matches (was the only CI test failure;
unrelated environmental integration/perf failures excluded).
* test(OMN-13096): update backmerge identity assertions to regenerated lock digests
The runner-image identity lock was regenerated for the pyproject entry-point
add; this test hardcodes the expected identity_digest/shared_env_digest, so
update both to match the new lock (same maintenance OMN-13094 did).
---------
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* fix(OMN-13118): tolerate partition-less reply topic in terminal consume leg (#1970)
The context-ROI runner opens one ephemeral group_id=None terminal consumer per
terminal topic BEFORE publishing each generation command (subscribe-before-publish,
OMN-13012/13038). The FAILED reply topic is only produced to on contract_passed=False;
in a battery where generations pass it has zero messages, so Redpanda never advertises
a partition for it. _assign_direct_terminal_partitions burned the full 30s assign cap
on every trial then raised a bare TimeoutError (the empty-message 'wait failed'),
stalling each of the 160 battery trials ~30s before the COMPLETED terminal could
correlate -> battery needs >80 min and never completes (verifier-confirmed wedge).
A partition-less reply topic is a valid steady state, not a 30s error:
- _assign_direct_terminal_partitions gives a bounded grace window for a topic that
exists but is slow to surface metadata, then assigns whatever partitions exist
(possibly none) and returns promptly instead of burning the cap and raising.
- poll_direct_terminal_consumer treats an empty assignment as 'no terminal will
arrive here' (sleeps out its timeout, returns None) without calling getone() on
an unassigned consumer.
Repro: tests/integration/test_terminal_consumer_battery_load_wedge.py drives the
REAL TerminalEventConsumer over K=10 x 2 cells x 2 arms; RED (40/40 trials block a
full assign cap) before the fix, GREEN after.
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* fix(OMN-13118): pin terminal-consumer read offset synchronously (close lazy seek_to_end publish race) (#1971)
* fix(OMN-13118): pin terminal-consumer read offset synchronously to close lazy seek_to_end publish race
The consume-leg wedge survived PR #1969 (set_topics no-op) and PR #1970
(partition-less assign-cap) because both addressed the assign phase, not the
seek timing. AIOKafkaConsumer.seek_to_end is LAZY: it requests a LATEST offset
reset that only resolves on the first poll — AFTER the caller publishes. With
generation completing in ~1s, the correlated COMPLETED terminal lands in the
open->poll gap, so the lazily-resolved LATEST position is the HWM AFTER the
record and the poll reads past it. The terminal is never read, the trial never
correlates, and the experiment matrix re-fires the same cell forever.
Replace seek_to_end with a synchronous end_offsets() + seek() pin
(_pin_direct_terminal_end_offsets) in both open_direct_terminal_consumer and
RuntimePatternBBroker._dispatch_and_wait_with_direct_kafka_consumer, so the
read position is fixed at open() time, before the publish — the real
subscribe-before-publish guarantee. Empty assignment (partition-less reply
topic, OMN-13118 #1970) is a no-op.
Repro: tests/integration/test_terminal_consumer_seek_reset_race.py drives the
real TerminalEventConsumer.open()/wait() the way HandlerContextRoiRunner does,
publishing the correlated terminal in the open->wait gap across K=10 x 2 cells
x 2 arms; RED with the lazy reset (every cell degenerate), GREEN once the read
offset is pinned. RED verified by git-stashing only the source fix.
Existing consume-leg fakes updated to model end_offsets/seek (they previously
masked the bug by making seek_to_end a synchronous exact snapshot).
* test(OMN-13118): reword assertion (lazy not deferred) for receipt honesty gate
* fix(OMN-13118): bound end_offsets() round-trip with assign-cap timeout (CodeRabbit)
end_offsets() is a broker ListOffsets round-trip aiokafka documents as able to
block indefinitely. Bound it with the same cap as start()/assign so a stalled
broker fails fast instead of hanging the pre-publish positioning.
---------
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* fix(OMN-13137): _validate_routing strategy-aware event_model checks (#1973)
operation_match routes by the `operation` field and does not use event_model.
The validator was unconditionally requiring event_model.{name,module} for every
handler entry, causing 230/295 omnimarket operation_match contracts (all correct
as authored) to fail routing validation at startup.
Fix: read routing_strategy from the routing map and branch validation:
- payload_type_match → require event_model.{name, module} (unchanged)
- operation_match (and any non-payload strategy) → require `operation`;
skip event_model checks entirely
Updated pre-existing _validate_routing tests to declare routing_strategy:
payload_type_match explicitly (they always tested payload_type_match semantics
but relied on the implicit fallback that is now removed).
Added test_validate_routing_operation_match.py with 4 unit tests:
1. operation_match without event_model → zero event_model errors
2. operation_match missing operation field → error
3. payload_type_match missing event_model → still errors (regression guard)
4. Real node_integration_sweep_orchestrator routing block → clean (boot gate)
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* fix(OMN-13118): independent per-terminal-topic consumers in Pattern B direct-Kafka wait (#1972)
The consume-leg wedge survived four merged fixes (#1969 set_topics no-op,
#1970 partition-less assign-cap, #1971 synchronous seek-pin). The STRONG K>=10
multi-cell reprobe on the stability lane still wedged on REBUILD-5
(cdf53d963f7b). Converged diagnosis (strikes 3+4,
docs/evidence/2026-06-12-weekend-pass/experiments/probe4-stability/
reprobe-K10-rebuild5/HALT_K10_WEDGE_PERSISTS.md): the runtime waited for each
trial's terminal across TWO topics (node-generation-completed.v1 +
node-generation-failed.v1) with a SINGLE ephemeral group_id=None consumer
assigned both topics' partitions. One aiokafka consumer holds one manual
subscription; the COMPLETED delivery window collapsed before it surfaced the
correlated record, so the trial never correlated and the matrix re-fired cell 1.
RuntimePatternBBroker._dispatch_and_wait_with_direct_kafka_consumer now opens
ONE independent AIOKafkaConsumer PER terminal topic via
open_direct_terminal_consumer (each started, assigned, and offset-pinned via
end_offsets()+seek() at open() BEFORE publish), then awaits both CONCURRENTLY
via asyncio.wait(FIRST_COMPLETED). The first correlated terminal wins; both are
torn down. No shared consumer, no subscription flip.
Keeps the #1970 partition-less no-op and the #1971 synchronous offset pin (both
live in open_direct_terminal_consumer / poll_direct_terminal_consumer). Removes
the now-dead single-consumer helpers (_assign_terminal_topic_partitions,
method-level _refresh_terminal_topic_metadata, _direct_kafka_* kwargs builders,
_kafka_bootstrap_servers/_kafka_event_bus).
Adds tests/integration/test_terminal_consumer_subscription_flip_wedge.py: a
real-dispatch-path K>=10 x 2-cell x 2-arm repro whose fake models TWO
independent consumers honestly (delivery is faithful only for a single-topic
assignment; a consumer spanning both topics flips and drops the COMPLETED
record). RED genuineness verified by reverting only the source to the
single-consumer shape (test hangs past timeout); GREEN with the fix in 2.3s.
Acceptance is the LIVE K>=10 multi-cell stability-lane reprobe (later phase),
NOT this unit test. A green unit repro is necessary but NOT sufficient.
Refs OMN-13118, OMN-13128.
Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
* fix(OMN-13118): long-lived terminal correlator replaces per-trial ephemeral consume leg (#1974)
Tier B canonical redesign (epic OMN-12525). Five offset/subscription patches
(#1969-#1972) tuned the per-tri…
…I release (0.36.1 → 0.38.6) (#2744) * docs(OMN-15198): slim CLAUDE.md 43049 -> 18136 bytes per context-engineering doctrine (#2482) Canary for the org-wide repo-level CLAUDE.md slim (root omni_home pass proved the method: the high-value cut is drift-prone factual state, not prose). - Move Install Model + Service Catalog walkthrough to docs/patterns/service_catalog.md (new), keep a pointer + the hardcoded_env/required_env trap inline. - Replace drift-prone state with the probe that regenerates it: coverage minimum -> pyproject.toml fail_under; transport types -> EnumInfraTransportType; error tree -> omnibase_infra.errors; CLI entry points -> pyproject [project.scripts]; bundle service lists -> docker/catalog/bundles.yaml; compose start_period literal -> docker/docker-compose.infra.yml; branch-protection narrative (OMN-14288) -> the audit_required_context_parity_cli.py report command + enforcement_parity_manifest.yaml. - Delete stale Agent-Driven Development section (predates Workflow-tool dispatch paradigm governed by the root CLAUDE.md). - All 8 KEEP-contract gotchas survive inline (no-publish handlers, AMBIGUOUS_CONTRACT_ CONFIGURATION fail-fast, dispatcher-owned resilience, hardcoded_env rule, RestartCount==0 compose-sandbox gate, ModelIntentPayloadBase removal, INFISICAL_ADDR opt-in prefetch, VERSION_MATRIX single-source-of-truth). Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * chore(OMN-15181): preserve main ancestry on infra dev (#2478) * chore(OMN-12934): continue dev to main promotion for prod redeploy (#1988) * docs(OMN-12962): contract-store durability audit — cold-runtime census proof (#1937) Proves cold-start contract census reconstructs from the image-bundled filesystem manifest (HYBRID-mode bootstrap + PluginLoaderContractSource), independent of node-registration.v1 retention. The delete-retention topic feeds only the post-freeze dynamic listener (auto_offset_reset=latest, no history replay). Live .201 stability-test evidence: filesystem contract_path in manifest + truncated topic log head with intact census. No store fix needed; residual dynamic-only gap covered by runtime_sweep sweep check. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12970): vendor omnimarket projection node migrations into forward/nodes (#1938) Vendors omnimarket node-owned projection migrations into the namespaced forward-migration tree so run-forward-migrations.sh materializes them in the dashboard projection DB (omnidash_analytics) at deploy. Primary (OMN-12970): creates llm_call_metrics, llm_cost_aggregates, and capability_scores in the projection DB. These were only ever created in the omnibase_infra DB by infra migrations 031/060, so the ab-compare, cost.token_usage, cost.summary, and capability-scores projection topics were DEGRADED at startup ('table not found') and their dashboard panels rendered empty. Also re-syncs three omnimarket node migrations the vendor tree had drifted from (node_projection_llm_routing, node_projection_overnight, node_projection_savings /077) — sync-node-migrations.sh --check requires the full vendored tree to match omnimarket source, and these were missing. Companion to omnimarket PR for the same ticket (source migrations + projection table-coverage ratchet test). Evidence-Ticket: OMN-12970 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12965): stamp runtime image identity (version + revision) in workspace builds (#1943) * fix(OMN-12965): stamp runtime image identity (version + revision) in workspace builds The main runtime image stamped org.opencontainers.image.version=0.1.0 with a blank org.opencontainers.image.revision after workspace rebuilds. A blank identity degrades every proof packet (runtime SHA + image digest are required citations in accepted evidence). Root causes (three build paths under-stamped identity): - onex up --build (cmd_up) passed only GIT_SHA; the runtime-stage OCI labels read VCS_REF (-> blank revision) and RUNTIME_VERSION (-> placeholder 0.1.0). - deploy-runtime.sh passed VCS_REF but not RUNTIME_VERSION/GIT_SHA. - Dockerfile silently allowed blank/placeholder identity in workspace mode. Fix: - cli._image_identity_build_args() stamps the full quad (GIT_SHA/VCS_REF/ RUNTIME_VERSION/BUILD_DATE) and fails fast on an unresolved git revision. - deploy-runtime.sh stamps RUNTIME_VERSION + GIT_SHA and verifies the version label is non-placeholder post-deploy. - Dockerfile.runtime fails workspace builds with blank VCS_REF or placeholder RUNTIME_VERSION=0.1.0 (release mode unaffected). Enforcement ratchet (same PR): - scripts/check_runtime_image_identity.py static check, wired as pre-commit hook + CI gate (ci.yml). - tests/unit/infra/test_runtime_image_identity_labels.py pins the cli helpers + Dockerfile guard; deploy-agent test extended for the quad. Proven locally via throwaway docker builds: workspace+args -> populated labels; workspace without args -> guard fails (exit 64); release without args -> 0.1.0 placeholder allowed (no regression). Evidence-Ticket: OMN-12965 * test(OMN-12965): integration build proof for runtime image identity labels Builds the real runtime-stage ARG/LABEL/guard block against busybox and asserts via docker inspect: workspace+args -> populated version/revision; workspace without args -> guard fails (exit 64); release -> placeholder allowed. Satisfies the integration-test hard gate and makes the throwaway proof permanent. Evidence-Ticket: OMN-12965 --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12987): workspace-build sibling lock-pin preflight + provenance (#1944) Recurrence ratchet for the 2026-06-11 stability bootstrap crash. The 11:20Z --no-cache rebuild vendored omnibase_infra 0.37.0-dev (~2c1d672f) + core 0.42.0 even though omnimarket dev's uv.lock pinned infra 0.38.1 @ e2dbdc95 + core 0.44.0 @ c97c2c9a. The stale sibling predated the OMN-12501 Protocol-quarantine guard, turning a latent contract defect into a fatal crash that crash-looped the main runtime. - check_sibling_lock_pins.py: host-side fail-fast preflight resolving expected sibling versions/SHAs from the consuming repo's (omnimarket) uv.lock and comparing them against each vendored tree. Mismatch aborts the build. - stage_workspace.sh: emit a .build-sha marker per staged sibling (rsync drops .git) so the preflight and provenance can identify the vendored commit. - deploy-runtime.sh: run the preflight after staging, before build; abort on mismatch. Write the comparison under sibling-repos/.sibling-lock-pins.json. - compute_workspace_provenance.py + Dockerfile.runtime: fold expected-vs-actual lock_pin_comparison into build-provenance.json for deploy verifiers. Recurrence-guard tests prove a stale infra 0.37.0 vs lock-pinned 0.38.1 fails the preflight and matched pins pass; deploy script wiring is asserted statically. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12988): pin stability runtime-worker to replicas 1 + census ratchet (#1942) The base docker-compose.infra.yml defaults runtime-worker to replicas 0 (${WORKER_REPLICAS:-0}). The stability-test lane's required state includes a running worker (4-container census: main, effects, worker, projection-api), but the override pinned it via an env-interpolation default (${STABILITY_TEST_WORKER_REPLICAS:-1}) — a silent-drop surface: a stray exported STABILITY_TEST_WORKER_REPLICAS=0 or removal of the :-1 fallback would scale the worker to 0 with zero signal on a plain compose up/recreate. Fix: pin docker-compose.stability-test.yml runtime-worker deploy.replicas to the literal 1 (no env indirection). Ratchet (recurrence guards, same PR): - scripts/deploy-agent/tests/unit/test_runtime_worker_census.py: assert runtime-worker stays in the deploy-agent RUNTIME-scope census so a missing worker (replicas 0 => absent from docker compose ps) is a deploy failure, not silence; assert the override pins a literal 1. - tests/integration/infra/test_stability_test_runtime_compose_render.py: assert the rendered stability worker resolves deploy.replicas == 1. - tests/unit/infra/test_stability_test_runtime_lane.py: update the existing pin assertion to the literal 1. Evidence-Ticket: OMN-12988 Config-drift family: OMN-12945 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12979): expire-bound topic completeness suppressions (#1940) Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12966): accept community-edition Infisical /api/status in provision readiness gate (#1939) The fresh-provision readiness gate in provision-infisical.py only accepted the enterprise {"status": "ok"} payload and rejected the community edition's {"message": "Ok"}, returning 1 before bootstrap could run. This blocked provisioning against the Infisical instance deployed on .201 (community edition). Route the gate through the existing _is_infisical_ready helper (single source of truth, already used by the already-provisioned path). Add TestMainFreshProvision- ReadinessGate covering community/enterprise/not-ready cases. Also adds docker/docker-compose.infisical-stability.yml: an ADDITIVE Infisical compose project for the stability-test lane (joins the existing network as external, reuses stability postgres/valkey, no lane mutation), since the lane overlays disable the in-lane Infisical service via *-disabled profile overrides. P1.2b-A: Infisical now reachable from the stability runtime/effect containers; a known secret (OMN_12966_PROBE) seeds and resolves end-to-end via the machine- identity universal-auth path, verified from inside the effects container. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12958): volume-config drift gate + runtime config provenance (#1941) * feat(OMN-12958): volume-config drift gate + runtime provenance Compute config provenance (path + sha256) for the runtime-rendered Bifrost delegation contract; the deployed volume copy survives rebuilds and silently diverges from packaged source (two competing authorities, OMN-12945). - runtime/config_provenance.py: ModelConfigProvenance + drift classification, sidecar JSON writer (read by sweep + proof packets) - runtime/health/health_config_provenance.py: drift -> degraded health - render entrypoint logs provenance line + writes sidecar on every boot - docs/runbooks/volume-config-drift-and-reseed.md: ledgered re-seed procedure - validation exemption for config_name (logical identifier, not entity ref) No live volume mutation: re-seed is an operator deploy step (deploy_pending). * test(OMN-12958): cover volume config drift reseed flow --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12957): wire runtime-profile validator + core registry parity guard (#1945) * feat(OMN-12957): wire runtime-profile validator + core registry parity guard - Remove the import-time RuntimeError drift raise in runtime_profile.py (a hard raise on core/infra version skew would crash the kernel at import); the parity invariant is enforced by test_profiles_match_core_registry instead. - Add tests: _PROFILES keys == omnibase_core REGISTERED_RUNTIME_PROFILES, and every CONSUMER_ATTACHED_RUNTIME_PROFILES profile loads as a real lane. - Wire omnibase_core.validation.validator_runtime_profiles as a pre-commit hook + validator-runtime-profiles.yml CI gate on infra contracts. - Freeze 19 pre-existing violators in validation/runtime_profiles_allowlist.yaml (discovered by repo-root walk; drain via OMN-12982). Blocks NEW orphans. Requires the omnibase_core pin to include OMN-12957's validator (new rules). Evidence-Ticket: OMN-12957 Evidence-Source: 5463fbaf819409d4fb7f491dd4f276f10d869eba * ci(OMN-12957): pass runtime profile allowlist to validator * test(OMN-12957): cover runtime profile registry parity * fix(OMN-12957): keep runtime profile allowlist under config * fix(OMN-12957): pin core runtime profile registry --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12973): widen migration-gate healthcheck start_period + ratchet (#1950) P2.8: classify the prod migration-gate UNHEALTHY-then-self-resolved incident. Classification: idle-one-shot-mis-modeled = NO. The migration-gate is a correctly-modeled long-running sentinel (entrypoint `while true; sleep 3600`) whose healthcheck continuously polls db_metadata.migrations_complete via check_migrations_complete.sh. The container flipped UNHEALTHY transiently because its healthcheck start_period (10s) was far shorter than the real cold-volume migration window (~116s: prod gate started 09:35:22, intelligence-migration finished 09:37:18). Past the 10s grace window the still-failing probe was reported UNHEALTHY until migrations completed, then self-resolved — no fault. Fix: raise migration-gate healthcheck start_period 10s -> 180s in both the authoritative catalog manifest (docker/catalog/services/migration-gate.yaml, flows into the generated compose) and the hand-maintained docker/docker-compose.infra.yml that deploy-runtime.sh applies to .201, so a still-applying gate stays in `health: starting` instead of flipping UNHEALTHY. Ratchet (enforcement, not detection): new ValidatorHealthcheckStartPeriod in omnibase_infra catalog (validator_healthcheck_start_period.py) asserts a 120s floor for migration-completion gates, wired into `onex validate runtime` (cmd_validate_runtime) AND backed by unit tests that gate every PR via pre-commit + CI. A migration-completion gate is identified by BOTH a healthcheck that polls migration completion AND a service_completed_successfully dependency, so ordinary app services (e.g. intelligence-api, an HTTP liveness probe with a legitimately short 40s start_period) are not swept into the floor. Prod probed read-only only; no prod mutation. Applies to live prod via the batched stability/prod rebuild (deploy_pending). Evidence-Ticket: OMN-12973 Evidence-Source: OCC#PENDING Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12971): deliver Vertex ADC bearer token to runtime effects container (secret-ref discipline) (#1949) * feat(OMN-12971): deliver Vertex ADC bearer token to runtime effects container (secret-ref discipline) Wire the Vertex $500-credit ADC path into the runtime, ADDITIVE next to the Gemini API-key path (provider-agnostic; neither provider removed or forced). runtime_policy.contract.yaml: add a secret-source mapping llm.vertex.access_token (source_type=file, /run/secrets/vertex_access_token) to the dev/stability-test/ judge profiles, alongside the existing llm.gemini.api_key env mapping. The ref name MUST match cloud-vertex-gemini.secret_ref in omnimarket bifrost_delegation.yaml. The resolved VALUE is a short-lived OAuth bearer token minted from ADC, refreshed by the operator; the token VALUE is never committed — only the ref name + in-container path. Add aiplatform.googleapis.com to the cloud host allowlist. docker-compose.infra.yml: bind the operator-supplied host token file read-only to /run/secrets/vertex_access_token on the main and effects runtimes (VERTEX_ACCESS_TOKEN_HOST_FILE, default /dev/null so lanes without Vertex still start; Gemini key path unaffected). Pass through BIFROST_VERTEX_GEMINI_ENDPOINT_URL (overlay supplies the complete Vertex OpenAI-compat URL) and GOOGLE_CLOUD_PROJECT/LOCATION (default empty). runtime-policy.env: regenerated from the contract via render_runtime_policy_env (test_runtime_policy_env_matches_contract_renderer proves contract<->env parity). test_runtime_policy_contract.py: update host-allowlist assertion for the additive Vertex host. Pre-existing platform-wide topic-parity-gate failures (25 unrelated topics) are identical with and without this change (proven by diff); that gate is pre-commit- only (not a CI merge gate) and this change adds zero new topic gaps, so that one hook is SKIP-ped. No deploy/receipt/merge gate is bypassed. * fix(OMN-12971): make Vertex runtime env contract-owned --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13008): automated .201 disk maintenance — worktree GC + docker GC + watermark alert ratchet (#1952) * feat(OMN-13008): automated .201 disk maintenance — worktree GC + docker GC + watermark alert ratchet Conservative, keep-list-driven disk maintenance to prevent the 2026-06-11 /data ~95% outage that killed all three lanes mid-demo: - scripts/worktree-gc.sh: drives the canonical omniclaude prune-worktrees.sh (merged+clean+pushed safety) on both Mac (merge-sweep tick) and .201 - scripts/disk-gc.sh + disk_gc_plan.py: conservative docker/builder/image GC. Pure, testable removal planner honoring a VERSIONED keep-list (deploy/disk-gc/keep-list.yaml): never reaps a kept repo, kept tag, in-use image, or anything younger than min_age_days; keeps N superseded generations. - scripts/disk-watermark-check.sh + disk_watermark_event.py: df watermark ratchet. >=85% emits a typed disk-watermark bus event (warning) that the sweep auto-ticket path turns into a Linear ticket; >=90% emits critical. Broker addr is fail-fast from KAFKA_BOOTSTRAP_SERVERS (no localhost default). - deploy/disk-gc/: systemd USER timer (onex-disk-gc.timer/.service, hourly) + install-disk-gc.sh. User units, NOT lane containers. - tests: 20 unit tests incl. GC plan-safety invariants + dry-run proof that default mode issues no destructive op (a wrong-delete GC is worse than none). Contract: contracts/OMN-13008.yaml * fix(OMN-13008): pass docker inventory to GC planner via stdin, not env (ARG_MAX) On a host with many docker images, passing the full image/ps inventory as env vars to disk_gc_plan.py exceeds ARG_MAX ('Argument list too long', exit 126), producing an empty plan. Write inventory to per-run scratch files (under the log dir, never /tmp; cleaned on exit) and hand it to the planner on stdin as a JSON envelope. Verified the failure live on .201; planner now reads stdin. * fix(OMN-13008): simplify GC plan stdin pipe (two processes, no nested subprocess) * fix(OMN-13008): keep-wins reconciliation — never remove an image id with any keep reason A single image id can surface in multiple 'docker image ls' rows (one per repo:tag). One tag could route the id to dangling-removal while another routes it to keep (e.g. tagged 'latest' or within-N-generations). Live .201 plan showed an id in BOTH remove_image_ids and kept_reasons. Reconcile at the end: keep wins — any id with a keep reason is dropped from the remove list; remove list deduped. Adds 2 regression tests. Verified live on .201. * fix(OMN-13008): timer uses OnCalendar=hourly for reliable re-arm (oneshot service) OnUnitActiveSec does not reliably re-elapse for a oneshot service once it goes inactive (observed NextElapseUSecMonotonic=infinity live on .201 after the first run). Switch to OnCalendar=hourly + RandomizedDelaySec so the timer re-arms every hour. Keep Persistent=true for missed-run catch-up. --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13005): materialize blocking event_consumer in runtime auto-wiring (runner consume-leg degenerate rows) (#1951) * fix(OMN-13005): materialize blocking event_consumer in auto-wiring (runner consume-leg degenerate rows) The runtime auto-wiring materialized event_publisher for handlers that declare it but had no equivalent for event_consumer. Request/response EFFECT handlers (HandlerContextRoiRunner) that publish a command then block on the correlated terminal event fell back to their no-op consumer default, returning None immediately -> every result row degenerate (failure_stage=generation, attempt_count=0) while generations succeeded ~1s later. Adds _make_sync_event_consumer (mirror of _make_sync_event_publisher), backed by service_terminal_event_consumer.make_terminal_event_consumer: a sync (topic, correlation_id, timeout) -> dict | None adapter that runs the proven direct-Kafka correlate-and-wait loop (from RuntimePatternBBroker) on an isolated event loop in a worker thread, so blocking does not deadlock the runtime dispatch loop that delivers the awaited terminal. TDD through the REAL dispatch path: test_event_consumer_injection drives a trial through _prepare_handler_wiring with a terminal arriving after a delay and asserts a non-degenerate row; verified RED with injection disabled. * fix(OMN-13012): two-phase (seek-now/wait-later) terminal event_consumer to close the subscribe-after-publish race (#1954) The OMN-13005 injected event_consumer is a single callable that does assign -> seek_to_end -> poll internally, all AFTER the handler has already published its command. Once OMN-13010 freed the dispatch loop and generation began completing in ~1s, the correlated terminal lands BEFORE the single-call consumer's post-publish seek_to_end positions, so seek_to_end skips PAST the already-emitted terminal and the runner times out on an offset beyond it (probe3, run_id=20260611T2140Z-probe3 -- both arms degenerate, zero rebalances). Splits positioning from waiting so the caller subscribes BEFORE it publishes: session = consumer.open(topic) # assign + seek_to_end NOW publisher(command_topic, payload) # publish AFTER positioning payload = session.wait(cid, timeout) # block from the captured position The returned TerminalEventConsumer is still directly callable with the legacy (topic, cid, timeout) -> dict | None single-call shape for any consumer that does not need subscribe-before-publish; the runner is the only consumer today. TerminalConsumerSession owns a dedicated event loop on a daemon worker thread for the whole open->wait->close lifecycle, preserving the OMN-13005 loop-isolation discipline so blocking never deadlocks the runtime dispatch loop. TDD (real dispatch path, RED-then-GREEN): extends the OMN-13005 injection test with a terminal emitted IMMEDIATELY after publish. The single-call (seek-after-publish) consumer MISSES it (degenerate row -- RED test asserts failure_stage=generation); the two-phase (open-before-publish) consumer CATCHES it (non-degenerate -- GREEN). The Kafka layer is faked at the two service seams against a shared in-memory log modeling seek-to-end semantics. OMN-13005 blocking-correlate behavior preserved. 269/269 auto_wiring unit tests pass; mypy --strict clean. Sibling to OMN-13010 / OMN-13005 / OMN-13003. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13005): route terminal consumer through Kafka boundary --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12990): pin worker replicas in ledgered config + fail-fast lane overrides (silent-drop ratchet) (#1948) * fix(OMN-12990): pin worker replicas in ledgered config + fail-fast lane overrides The base compose sets runtime-worker deploy replicas to ${WORKER_REPLICAS:-0} (soft-default ZERO). The stability lane's required state includes a running worker (GATE_ZERO_PROOF.md: 4 runtime containers), but the worker presence rode entirely on a compose soft-default — any plain compose up/recreate without the policy env silently scaled the worker to zero with no error and no signal. Fix (contract-native + fail-fast): - Add 'replicas' to ModelRuntimeProcessPolicy; pin replicas: 1 in every lane's worker block in runtime_policy.contract.yaml. - Renderer emits {PROFILE}_WORKER_REPLICAS into the ledgered runtime-policy.env for dev/stability-test/judge/prod. - stability + prod compose overrides reference ${..._WORKER_REPLICAS:?...} (fail-fast, NO silent :-1/:-0 default). A recreate that omits the policy env now aborts loudly instead of dropping the worker. prod previously had no override at all and inherited the dangerous :-0 default. Ratchet (recurrence guards): - tests asserting fail-fast override form (no soft default), contract-declared replica pin >= 1 per lane, and rendered {PROFILE}_WORKER_REPLICAS=1 in the ledgered env. - runbook deploy/verify procedure adds an expected-container census (worker must be present) via verify_container_manifest; a missing worker is a FAILURE, not silence. Note: SKIP=topic-parity-gate — that local-only advisory gate (absent from all .github/workflows, not a required CI check) fails on 25 pre-existing cross-repo topic gaps (build-loop/omniclaude/omniweb) identical on pristine base HEAD 8d7da1249; this change adds zero topics. All other hooks ran clean. Evidence-Ticket: OMN-12990 * test(OMN-12990): cover worker replica policy integration --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12909): add gateway bus forwarder P0A (#1946) * feat(OMN-12909): add gateway bus forwarder p0a * test(OMN-12909): add gateway forwarder integration coverage * fix(OMN-12909): satisfy gateway forwarder validators * test(OMN-12909): allow gateway forwarder bus protocol * fix(OMN-12909): sync gateway forwarder entry point * fix(OMN-12909): refresh runner image identity lock * test(OMN-12909): relax JSON normalizer mixed benchmark threshold * fix(OMN-12909): allow gateway handlers to boot unconfigured * fix(OMN-12909): declare gateway forwarder runtime profile * fix(OMN-12909): update runner identity backmerge expectation --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13011): LANE CENSUS RECONCILIATION ratchet — declared desired-state per lane, drift = auto-ticket (#1955) The class fix for the recurring lane-drift regression. Nothing reconciled the declared desired state of a runtime lane against what is actually running, so the same failure kept recurring with zero signal: volume config drift (OMN-12945), WORKER_REPLICAS silent zero (OMN-12988/12990), and on 2026-06-11 prod runtime containers plus the broker network were silently absent for hours during demo prep. Ships a per-lane DESIRED-STATE census: - (a) DECLARED in a versioned lane manifest (deploy/lane-census/lane-manifest.yaml): container set, network, replicas, image-tag pattern per lane (stability-test/prod/judge/dev), derived from the canonical compose lane files. A parity ratchet keeps the manifest locked in step with the compose files. - (b) RECONCILED on a schedule on .201 by SHARING the OMN-13008 systemd timer (a drop-in 4th ExecStart on onex-disk-gc.service — never a second timer) and on-demand via scripts/lane-census-check.sh / runtime_sweep. - (c) Drift = typed bus event (onex.evt.infra.lane-census-drift.v1) + Linear auto-ticket naming exactly what is missing/extra (container_absent, network_detached, replicas_zero, unexpected_container, oneshot_failed/stuck, image_tag_mismatch). Fail-fast, no warn-only mode (gates-block policy); exit 30 on drift; bus publish fail-fast on KAFKA_BOOTSTRAP_SERVERS (no localhost default). Red fixture reproduces 2026-06-11: prod runtime containers absent + broker network detached must produce the exact drift findings + a non-zero exit hours before a human noticed. Pure planner is fully unit-tested; shell driver dry-run-tested. Builds on the OMN-12988 deploy-agent RUNTIME census (deploy-time) as the complementary steady-state reconciler; closes the runtime-worker.yaml container_name: null census gap by sourcing names from the compose lane files. Evidence-Ticket: OMN-13011 Config-drift family: OMN-12945 Relates-to: OMN-13009, OMN-12988, OMN-13008 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13020): vendor missing node migrations — llm_routing 0000 + context_roi 001 (#1956) Vendors two omnimarket node-source migrations into the infra forward-migration tree via scripts/sync-node-migrations.sh (the canonical OMN-12559 mechanism): - node_projection_llm_routing/0000_create_llm_routing_decisions.sql (source: omnimarket #1168 / OMN-12942, merge ed6734f8) - node_projection_context_roi/001_create_context_roi_scores.sql (source: omnimarket #1178 / OMN-12955, merge 5010b1f4) Without the 0000 base table, node_projection_llm_routing/0001 (CREATE VIEW) hard-fails against NODE_POSTGRES_DB=omnidash_analytics — exactly the prod forward-migration exit-3 of 2026-06-11T09:35:52Z, and reproduced by construction in any clean clones@dev build. Files are byte-identical to the omnimarket dev blobs (sha256 f8a8b339… / c4126e65…) and to the untracked hot-patch copies on the .201 stability clone. Both migrations are self-contained, all-statements-IF-NOT-EXISTS, and 0000 sorts lexically before 0001 within the node's namespaced identity space. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13058): close TerminalConsumerSession on open() failure (worker thread + event loop leak) (#1957) TerminalConsumerSession.__init__ starts its dedicated worker loop thread immediately. TerminalEventConsumer.open did 'session = TerminalConsumerSession(...); return session.open()' with no cleanup: any failure inside session.open() (consumer start timeout, partition-assign timeout, broker auth error) propagated out of the raising expression, the session reference was lost, and the daemon worker thread plus its never-closed asyncio event loop leaked -- one pair per failed open. The motivating caller (HandlerContextRoiRunner) opens a session per trial, so a 160-560-trial battery against a degraded broker accumulates hundreds of leaked threads in the long-lived effects container. Fix: wrap session.open() in try/except BaseException -> session.close() (idempotent: stops the loop, joins the thread) -> re-raise. Covers both the two-phase .open(topic) path and the legacy single-call __call__ path. Found by the P3.3 doctrinal review of merged #1951 (b9712af9 / 36d98275). TDD: tests/unit/runtime/test_service_terminal_event_consumer_open_failure.py injects a real open failure through the production path (event bus without _bootstrap_servers) and asserts no alive terminal-consumer-* thread after the raise. Verified RED with the fix stashed (2 failed), GREEN with it (2 passed). Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13021): non-dev-base guard — fail feature-base PRs absent Stacked-Parent declaration (retro A-6) (#1958) Any PR whose base is neither dev nor main fails unless the body carries 'Stacked-Parent: #N'. Prevents the feedback_stacked_prs_orphan_from_dev class (#1185/#1954 auto-merged INTO parent feature branches and stranded off dev). base=main remains governed by main-target-guard. Epic OMN-13013 (process enforcement ratchets — June 12 retro). Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13014): hot-patch ledger rebuild preflight gate (retro B-1) (#1959) * feat(OMN-13014): hot-patch ledger rebuild preflight gate (retro B-1) Hot-patches on .201 (.prepatch sibling discipline) silently revert on any image rebuild/force-recreate — the 2026-06-11 20:58Z rebuild already erased a live /api/generate patch once. This adds the rebuild-path gate: - scripts/preflight_hotpatch_ledger.py: given a target container or lane + per-repo build refs, hard-fails when any hot-patch ledger row's source PR merge commit is not an ancestor of the build ref (git merge-base --is-ancestor), plus a .prepatch tripwire of the running container (unledgered .prepatch = hard fail; --post-rebuild = zero .prepatch expected). Sole bypass: HOTPATCH_PREFLIGHT_BYPASS carrying the Rule-10 '# skip-token-allowed: <user-approval-receipt-id>' form. - scripts/deploy-runtime.sh: guard_hotpatch_ledger wired into main() before build/preview (both dry-run and execute), lane derived from the compose project; skips loudly only when no ledger exists on the host. - tests/unit/scripts/test_preflight_hotpatch_ledger.py: 17 unit tests (ancestor gate, lane scoping, ledger loading, tripwire, bypass forms). - tests/ci/test_receipt_gate_install_guard.py: repair stale guard — core OMN-12565 replaced the OMN-9198 'uv pip uninstall first' install step with a cleared workspace venv (uv venv --clear); assert the new contract. Ledger backfilled from live census (5 .prepatch files / 4 source PRs, all MERGED to dev) at /data/omninode/hotpatch-ledger/ledger.yaml on .201. * fix(OMN-13014): scope missing-prepatch tripwire warning to the probed container --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12977): workspace-build sibling-pin ratchet — honor consuming lock, fail-fast on stale vendor (#1936) * fix(OMN-12977): workspace-build sibling-pin ratchet — honor consuming lock, fail-fast on stale vendor The workspace-mode image build (BUILD_SOURCE=workspace, used by the stability-test deploy procedure) vendored sibling/foundation packages from whatever the canonical OMNI_HOME clones happened to be checked out at, ignoring the consuming repo's uv.lock. On 2026-06-11 this shipped a 13-day-stale omnibase_infra 0.37.0-dev (pre-OMN-12501 Protocol-quarantine guard) + core 0.42.0 against an omnimarket dev lock pinning infra 0.38.1@e2dbdc95 / core 0.44.0@c97c2c9a, dropping the guard and crashing wire_from_manifest bootstrap fatally (stability lane down on demo day). Fix + recurrence ratchet (same PR): - scripts/runtime_build/check_sibling_lock_pins.py: parse the consuming repo's uv.lock for expected version+git-rev of each foundation/sibling package (scoped to the package's own source line so editable/registry pins are not cross-attributed a dependency's rev), resolve the actual clone version+HEAD, compare, and classify drift backward/forward/none. Fail-fast (exit 1) on any drift; --allow-drift records an explicit operator override in the artifact, never silent. - stage_workspace.sh: runs the preflight against the canonical clones before staging; aborts the build (exit 3) on unacknowledged drift and writes workspace/sibling-pin-comparison.json. - compute_workspace_provenance.py: folds the expected-vs-actual comparison into build-provenance.json so deploy verifiers can assert the build honored the lock; flags unacknowledged drift as a provenance error. - Dockerfile.runtime: COPY the comparison artifact (committed placeholder so the COPY always resolves; overwritten by stage_workspace.sh in workspace mode). - TDD: 19 unit tests covering lock parsing (git/registry/editable sources), drift classification, the exact 0.37.0-vs-0.38.1 stale case, check_pins exit codes, and the allow-drift override. Pre-existing mypy-strict bare-dict errors in compute_workspace_provenance.py fixed in the same pass. Evidence-Ticket: OMN-12977 Evidence-Source: pending-occ * ci(OMN-12977): retry runtime smoke compose port race * test(OMN-12977): align sibling-pin script tests with current API --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12989): workspace-mode image build must honor sibling lock pins (#1947) * fix(OMN-12989): workspace build must honor sibling lock pins (fail-fast ratchet) The 2026-06-11 stability bootstrap crash was caused by a workspace-mode --no-cache rebuild that vendored omnibase_infra 0.37.0-dev (13-day-stale worktree) even though omnimarket dev's uv.lock pins infra 0.38.1. The downgraded sibling predated the OMN-12501 Protocol-quarantine guard and turned a latent contract defect into a fatal crash. Fix + ratchet (same PR set): - scripts/runtime_build/resolve_workspace_pins.py: parse the consuming repo's uv.lock for sibling pins (version + git rev); classify each staged/installed sibling as exact/ahead/regression/unpinned; fail-fast (WorkspacePinError) on any regression BELOW the lock pin. Stdlib version-tuple fallback when packaging is absent so the ratchet never fails-open. - compute_workspace_provenance.py: enforce sibling pins + a host-infra self-check (installed omnibase_infra vs lock pin — the exact crash vector, since host infra is built from the context, not staged), and emit a pin_comparison block into build-provenance.json for deploy verifiers. - Dockerfile.runtime: COPY resolve_workspace_pins.py beside the provenance script so the in-image import resolves. - TDD: failing tests first (test_resolve_workspace_pins.py, 11 cases) + provenance integration tests; fixed a pre-existing stale RUNTIME_VERSION assertion to read pyproject dynamically. Evidence-Ticket: OMN-12989 * fix(OMN-12989): workspace build must honor sibling lock pins (fail-fast ratchet) The 2026-06-11 stability bootstrap crash was caused by a workspace-mode --no-cache rebuild that vendored omnibase_infra 0.37.0-dev (13-day-stale worktree) even though omnimarket dev's uv.lock pins infra 0.38.1. The downgraded sibling predated the OMN-12501 Protocol-quarantine guard and turned a latent contract defect into a fatal crash. Fix + ratchet (same PR set): - scripts/runtime_build/resolve_workspace_pins.py: parse the consuming repo's uv.lock for sibling pins (version + git rev); classify each staged/installed sibling as exact/ahead/regression/unpinned; fail-fast (WorkspacePinError) on any regression BELOW the lock pin. Stdlib version-tuple fallback when packaging is absent so the ratchet never fails-open. - compute_workspace_provenance.py: enforce sibling pins + a host-infra self-check (installed omnibase_infra vs lock pin — the exact crash vector, since host infra is built from the context, not staged), and emit a pin_comparison block into build-provenance.json for deploy verifiers. - Dockerfile.runtime: COPY resolve_workspace_pins.py beside the provenance script so the in-image import resolves. - TDD: failing tests first (test_resolve_workspace_pins.py, 11 cases) + provenance integration tests; fixed a pre-existing stale RUNTIME_VERSION assertion to read pyproject dynamically. Evidence-Ticket: OMN-12989 * test(OMN-12989): co-locate provenance pin helper in fixture --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13055): absent repos in REPOS list warn and exit 0 instead of failing (#1960) Missing repos (not cloned locally) now emit a WARN line and are tracked in a separate WARNED array. Only real fetch/ff failures cause exit 1. This makes pull-all.sh safe to use on machines with a partial clone set, while keeping the explicit-list override behavior intact. Adds three regression tests: absent-only exits 0, absent+present exits 0 with OK for the present repo, present-failed+absent exits 1. Evidence-Ticket: OMN-13055 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13070): config_prefetcher overlay wins on controlled lanes (#1961) On infisical_required=True lanes, fetched/overlay config now always wins over ambient env. Ambient env is retained only as a declared bootstrap fallback (with an explicit provenance INFO log line) when Infisical returns None. apply_to_environment also overwrites stale env on controlled lanes. Uncontrolled lane (infisical_required=False) behaviour is unchanged. Adds 5 regression tests: controlled-lane Infisical-wins, env-bootstrap- fallback, apply_to_environment overwrite, missing-from-both-is-error, and uncontrolled-lane-env-still-wins. Refactors _resolve_key to return a (outcome, value, error) tuple to satisfy the ≤5-param pattern gate. Source: docs/audits/2026-06-10-runtime-env-overlay-authority-audit.md Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13062): migration-gate vacuity fix — sentinel discipline, wait-for-postgres, skip-manifest, sync-check nonzero (retro A-10) (#1963) * fix(OMN-13062): migration-gate vacuity fix — sentinel discipline, wait-for-postgres, skip-manifest, sync-check nonzero Fixes three bugs identified in retro A-10 (recurrences OMN-12885, OMN-12934): (1) RUNNER SENTINEL DISCIPLINE run-forward-migrations.sh now clears migrations_complete=FALSE at the start of every run and sets it TRUE only as its FINAL act after all infra and node migrations succeed. Any mid-run failure leaves the gate UNHEALTHY. runner_completed_at is stamped at the same final step as durable evidence of a successful completion run. (2) SYNC-NODE-MIGRATIONS VACUOUS GATE sync-node-migrations.sh --check now exits 2 (not 0) when the omnimarket source tree is unresolvable. Silent exit-0 was hiding drift. The single opt-out is SYNC_NODE_MIGRATIONS_SKIP_UNRESOLVABLE=1 for environments that intentionally run without the source. (3) WAIT-FOR-POSTGRES GUARD run-forward-migrations.sh now waits up to PG_WAIT_RETRIES (default 30) x 2s for Postgres to accept connections before proceeding, guarding the first-boot initdb race. (4) SKIP-MANIFEST docker/migrations/skip-manifest.yaml introduced as the sole committed escape for intentionally-skipped migrations. The runner reads this at startup; listed migrations are recorded in schema_migrations with checksum "skip-manifest" without executing the SQL. (5) MIGRATION 085 Adds runner_completed_at TIMESTAMPTZ column to db_metadata so the runner's final stamp is durable in the schema (idempotent ADD COLUMN IF NOT EXISTS). Rollback included. 22 regression tests added covering all five fix surfaces. * fix(OMN-13062): stamp schema fingerprint for migration 085 Migration 085 (085_add_runner_completed_at_to_db_metadata.sql) was added in the initial commit but schema_fingerprint.sha256 was not regenerated. Running `python scripts/check_schema_fingerprint.py stamp` updates the artifact from the stale hash to match the 71 migration files. Evidence-Source: OCC#2563 Evidence-Ticket: OMN-13062 --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12864): Bifrost endpoints → committed overlay authority + fail-loud loader (#1964) * feat(OMN-12864, OMN-12814, OMN-12945): Bifrost endpoints → committed overlay authority + fail-loud loader OMN-12864 — Committed lane overlay - docker/lane-overlays/dev.bifrost.yaml: typed deployment bindings for all four BIFROST_LOCAL_*_ENDPOINT_URL values (coder :8000, reasoner :8001, embedding :8100, ds4-flash :8101). Previously only available as ephemeral shell exports on .201; now committed, auditable, diff-able, and CI-checked. - docker/lane-overlays/dev.bifrost.env: generated dotenv sidecar consumed by compose via env_file; never edited directly (yaml is authority). - docker/docker-compose.infra.yml: wire the env_file block at compose root so the four endpoints are injected into the interpolation context on a clean shell. Hardcode BIFROST_CONTRACT_PATH (remove :-/empty footgun — OMN-12814). - scripts/render_bifrost_lane_overlay_env.py: render script regenerates the env sidecar from the YAML source. - src/omnibase_infra/runtime/models/model_bifrost_lane_overlay.py: ModelBifrostLaneOverlay — typed Pydantic model enforcing URL completeness (OMN-12815: every URL must end in /chat/completions). OMN-12814 — Fail-loud loader - render_bifrost_delegation_contract: raises ProtocolConfigurationError on FileNotFoundError, YAMLError, ValidationError, and zero-endpoint renders. No lru_cache — every restart re-renders from packaged source so a stale cache cannot pin a broken result across deploys. OMN-12945 — Re-seed from packaged source on deploy - docker/entrypoint-runtime.sh: set BIFROST_FORCE_RESEED=1 on every container restart so the named-volume copy is always rebuilt from the packaged bifrost_delegation.yaml merged with committed lane-overlay endpoints. - render_bifrost_delegation_contract: honor BIFROST_FORCE_RESEED/force_reseed flag to bypass the stale-volume early-return path entirely. Tests: - tests/ci/test_bifrost_lane_overlay.py: CI gate — env sidecar in-sync with YAML source; all four BIFROST_LOCAL_* keys present. - tests/unit/runtime/models/test_model_bifrost_lane_overlay.py: bare-base URL rejection, env dict mapping, extra-field rejection. - tests/unit/runtime/test_render_bifrost_delegation_contract.py: fail-loud paths, force-reseed, zero-endpoint error, endpoint URL completeness. - tests/unit/models/test_model_serialization_roundtrip.py: roundtrip coverage. * fix(OMN-12864): move bifrost env_file to service level — fix compose schema validation failure Top-level 'env_file' is rejected by Docker Compose v2 schema validator ('additional properties not allowed'). This caused 10+ compose-render integration tests to fail in CI. Fix: - Remove top-level env_file block from docker-compose.infra.yml - Add per-service env_file on omninode-runtime, runtime-effects, runtime-worker (the three containers that render Bifrost) - Change BIFROST_LOCAL_*:? to BIFROST_LOCAL_*:- in x-runtime-env (compose-level validation removed; Python validates via ModelBifrostLaneOverlay + render_bifrost_delegation_contract) - Add two CI gate tests: compose_env_file_is_service_level_not_top_level and runtime_services_have_bifrost_env_file * fix(OMN-12864): allow BIFROST_LOCAL_* empty defaults in silent-fallback gate The overlay authority pattern (OMN-12864) passes BIFROST_LOCAL_*_ENDPOINT_URL via service-level env_file (docker/lane-overlays/dev.bifrost.env), not at compose config time. Compose-level :? would break CI rendering without the overlay pre-loaded. Validation at the Python layer (ModelBifrostLaneOverlay + render_bifrost_delegation_contract) is the enforcement point. Add the three failing vars to ALLOWED_EMPTY_DEFAULTS with OMN-12864 citation. --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12858/OMN-12879/OMN-12880): dispatcher route-coverage CI gate (#1965) * feat(OMN-12858/OMN-12879/OMN-12880): dispatcher route-coverage CI gate Static contract analysis: every subscribed command topic (onex.cmd.*) must declare handler_routing or runtime_dispatch, or the message goes to DLQ silently. This gate would have caught two recent incidents: 1. June 9 DLQ regression (OMN-12858 post-mortem): node_generation_consumer subscribed onex.cmd.omnimarket.node-generation-requested.v1 but a sole-handler revert left zero dispatcher routes registered. Messages went to DLQ silently with no CI signal. 2. June 12 DEL-01 live finding: onex.cmd.omnimarket.delegate-skill.v1 was consumed by dev lane bus but no dispatcher route existed in any deployed contract. Discovered via manual rpk consumer-group lag probe (DEL-01 evidence, docs/evidence/2026-06-12-weekend-pass/). Deliverables: - scripts/check_dispatcher_route_coverage.py — static YAML scanner that checks both omnibase_infra and omnimarket contract trees; ratchet allowlist for known pre-existing violations; --changed-contracts mode (OMN-12879) for per-PR scoping; compat publish topics excluded (OMN-12880) - .github/workflows/dispatcher-route-coverage.yml — CI workflow that checks out omnimarket sibling, collects changed contract paths in PR mode, and runs the gate; fires on PR, push-to-main, and merge_group - tests/ci/test_dispatcher_route_coverage_gate.py — 12 unit tests covering RED/GREEN/COMPAT/CHANGED-MODE/ALLOWLIST/MULTI-DIR paths plus live-contract regression proof against the actual omnibase_infra tree Allowlist additions: - onex.cmd.omnibase-infra.pattern-b-dispatch.v1 (RuntimePatternBBroker, imperative consumer, OMN-12525 migration target) - onex.cmd.platform.contract-resolve-requested.v1 (transitional HTTP bridge node_contract_resolver_bridge OMN-2756, metadata.transitional=true) [OMN-12858, OMN-12879, OMN-12880] * fix(OMN-12858): drop full uv sync from dispatcher-route-coverage workflow Gate script only needs pyyaml (stdlib + yaml). Using full setup-python-uv was causing 10+ minute timeout. Replace with direct pip install pyyaml and invoke python3 directly. Reduces job from 10m timeout to <1m. [OMN-12858] --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13026): wire transport-mock-lint pre-commit hook + CI gate (#1962) * feat(OMN-13026): wire transport-mock-lint pre-commit hook + CI gate (omnibase_infra) Activates the transport-mock-lint validator (from omnibase_core, OMN-13026) on omnibase_infra. Ratchet baseline: 218 existing violations across 63 files frozen in validation/transport_mock_baseline.yaml. New bare AsyncMock/ MagicMock on EventBus/transport surfaces are blocked by pre-commit hook and CI lint step. Existing violations tracked for drain by per-site tickets (parent OMN-13026). Reference incident: PR #1181 bare AsyncMock hid missing EventBusKafka.stop(). Evidence-Ticket: OMN-13026 * fix(OMN-13026): use uv run python for CI lint step + bump omnibase_core pin to include transport_mock_lint The transport-mock lint CI step previously cloned omnibase_core and ran `python -m omnibase_core.validators.transport_mock_lint` with PYTHONPATH, but this failed: `No module named omnibase_core.validators.transport_mock_lint` because it ran `.venv/bin/python` which uses the locked venv, and the venv omnibase_core pin (2defabef4) predates the transport_mock_lint module. Fix: use `uv run python` (removes the clone step) and bump omnibase-core git pin from 2defabef4 to 309d89fa7 (PR 1231 merge commit on dev) so transport_mock_lint is available in the locked venv. * fix(OMN-13026): align transport mock baseline and runner lock * fix(OMN-13026): sync omnibase_core pin + runner identity lock to dev baseline Align pyproject.toml omnibase_core rev to 2defabef (required by test_release_backmerge_preserves_proven_runtime_core_pin) and update docker/runners/runner-image.lock.json identity_digest/shared_env_digest to match dev runner image lock (79b08f44 / 90c8b3b9). Both were stale from the prior session's pin bump that used an older SHA. * fix(OMN-13026): source transport mock validator from core * fix(OMN-13026): source transport validator from core dev --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13094): --output receipt mode on onex node/run — quiet typed receipts with durable capture (#1966) Phase 2a of the skill-output-suppression slice (epic OMN-13089, plan docs/plans/2026-06-12-skill-output-suppression-plan.md Phase 2 item 1). - onex node/onex run gain --output receipt: ALL runtime logging routes to a run_id-suffixed capture file under <state-root>/captures/ (no console handlers — kills the 25-50-line RuntimeLocal INFO stream at the source); stdout carries exactly ONE typed ModelSkillResult JSON with the FULL handler result (result_model = concrete handler result type FQN). - Durable capture: capture log + handler result content-addressed via omnibase_core ArtifactStore (OMN-13093); artifact.captured + tool.output.captured emitted to the emit daemon socket (--emit-socket, default ~/.claude/emit.sock). - Failure asymmetry: artifact write failure => FULL output printed, no receipt (no hidden loss); emission failure => receipt still prints, event spooled to <state-root>/emit_spool/ for replay. - Node failure => status=failed/error with full error + capture log INLINE in the receipt (errors are never hidden) and artifact-backed. - Default output mode unchanged (enforcement is Phase 4). - RuntimeLocal exposes handler_result (receipt schema identity). - core pin 2defabef -> ae8793bd (merged OMN-13091/13093 receipt models + ArtifactStore); runner-image identity lock regenerated and the OMN-12765 backmerge identity constants updated for the new pin. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13097): onex skill subcommand + declarative skill->node mapping (#1968) * feat(OMN-13097): onex skill subcommand + declarative skill->node mapping Phase 4a of the skill-output-suppression slice (epic OMN-13089, plan docs/plans/2026-06-12-skill-output-suppression-plan.md Phase 4 item 1/2). A dispatch skill IS one CLI call (user directive 2026-06-12). This adds the `onex skill <name> [args]` dispatch surface that the 24 omniclaude shim migrations build on: - onex.cli entry-point `skill` -> cli_skill.run_skill_by_name. Resolves the skill via the declarative skill_mapping.yaml registry, builds the backing node's input payload from the skill's CLI args, writes it under the state root (.onex_state/tmp/<skill>-<run_id>.json — never /tmp), resolves the node's packaged contract exactly like `onex node`, and dispatches through the proven receipt-mode path (run_receipt_mode, OMN-13094). stdout is exactly one typed ModelSkillResult JSON with the FULL handler result. - skill_mapping.yaml: declarative DATA mapping all 24 dispatch shims to their backing onex.nodes node + typed result-model FQN (verified against each live handler handle() return type on origin/dev) + per-arg payload specs + static payload + keyword classifiers (delegate task_type as data, not code). Adding a skill is a YAML edit + fixture, never a CLI code change (ticket deliverable 2/3). Mapping lives beside the node-resolution surface, never hardcoded branching in the CLI. - Typed models split one-per-file (repo convention): ModelSkillArgSpec, ModelSkillClassifier, ModelSkillMapping, ModelSkillMappingRegistry, EnumSkillArgType. Frozen, extra=forbid, fail-fast coercion/validation. - validation_exemptions.yaml: Click-callback param-count + literal-identifier name-field exemptions mirroring the existing cli_node run_node_by_name precedent (OMN-11570) — same pattern, same rationale. dod_evidence: - 20 unit tests pass (registry validity, all-24-shims coverage, FQN result models, arg parsing/coercion/positional/required, classifiers, payload build, receipt-mode dispatch wiring, payload-under-state-root not /tmp). - uv run mypy src/ --strict: clean (2438 source files). - ruff format + check: clean. pre-commit run on changed files: pass. - runner-image identity lock regenerated for the pyproject entry-point add (same as OMN-13094). * test(OMN-13097): rebind OMN-12765 backmerge identity constants for onex skill pyproject change Adding the `skill` onex.cli entry-point to pyproject.toml changes the runner-image identity_digest (and shared_env_digest) the lock binds. Update the hardcoded expected constants in the backmerge-identity test to the regenerated values — same mechanical rebind OMN-13094 performed for the core pin bump. Identity + runner-image-identity tests pass (14). --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix+test(OMN-13012): force terminal-topic metadata refresh so both ephemeral consumers assign (runner consume-leg wedge) (#1969) The runner consume-leg wedged on the live stability battery (image c0505521f1fa, EXP1-3_RUNNER_CONSUME_LEG_BLOCKER): only the FAILED terminal topic ever subscribed while the COMPLETED topic never assigned, so the correlated completed terminal was never read and the 8x2x10 matrix re-fired cell 1 forever, emitting zero non-degenerate rows. A prior two-strike diagnosis proved the omnimarket handler is correct (it opens both terminal sessions pre-publish and races them); the defect is in the omnibase_infra runtime consume leg. Root cause: _assign_direct_terminal_partitions ignored the metadata future returned by AIOKafkaClient.set_topics and re-called set_topics([same_topic]) each loop iteration. aiokafka 0.13.0 set_topics only forces a metadata refresh when the topic set DIFFERS from the tracked set, so every iteration after the first took the no-op branch and never re-fetched. An ephemeral group_id=None consumer whose first metadata fetch had not yet surfaced partitions burned the full 30s assign cap and raised a bare TimeoutError (the empty-message 'wait failed' seen live). Fix: register the reply topic once and await that metadata fetch, then on each miss force a fresh fetch via force_metadata_update (which always fetches) rather than the no-op set_topics repeat. The assign-cap TimeoutError now carries a diagnostic message instead of an empty one. Test: tests/integration/test_terminal_consumer_concurrent_assign_race.py drives the REAL TerminalEventConsumer (the object wired as event_consumer) through the REAL open_direct_terminal_consumer/poll path with AIOKafkaConsumer monkeypatched to a fake that faithfully models aiokafka 0.13.0 set_topics future + metadata-latency semantics. RED before the fix (bare TimeoutError, the live empty-message signature); GREEN after. K>=2 multi-trial variant asserts no worker-thread leak across trials. Evidence-Source: <occ-sha-pending> Evidence-Ticket: OMN-13012 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13096): onex delegate single-command subcommand (Phase 2b) (#1967) * feat(OMN-13096): onex delegate single-command subcommand Add 'onex delegate "<prompt>" [--task-type X] [--max-tokens N]' as a subcommand on the existing onex CLI (Phase 2b of the skill-output-suppression slice, OMN-13089). The command wraps payload construction, node dispatch, and result extraction internally and prints exactly one ModelSkillResult[ModelDelegateSkillResponse] to stdout via the OMN-13094 receipt-mode path. RuntimeLocal logs go to the capture file + artifact store, never to stdout; scratch payloads live under <state-root>/tmp/ with run_id suffixes (never /tmp). - cli_delegate.py: classify_task_type (keyword table from legacy skill md), payload write, contract resolve, run_receipt_mode dispatch - register 'delegate' under onex.cli entry points - exempt delegate_command from the >5-param patterns gate (same Click-callback rationale as run_node_by_name) - 20 unit tests: classification, scratch-under-state-root, single typed receipt on stdout, zero INFO log leakage omnibase_infra does NOT depend on omnimarket; the delegate node is resolved at runtime via the onex.nodes entry-point group (registered by omnimarket). * chore(OMN-13096): re-trigger deploy-gate after Evidence-Source set to OCC#2593 No code change — the deploy-gate workflow triggers on synchronize (not edited), so the PR-body Evidence-Source fix needs a new commit to re-resolve the OCC ref to the open PR head where contracts/OMN-13096.yaml (with deploy evidence) lives. * chore(OMN-13096): re-trigger deploy-gate now that OCC#2593 merged to OCC dev contracts/OMN-13096.yaml (with the dod-deploy-onex-delegate item) is now on OCC dev, so the deploy-gate OCC-dev checkout resolves the contract + deploy evidence. * chore(OMN-13096): regenerate runner-image identity lock for pyproject entry-point add Adding the 'delegate' onex.cli entry point to pyproject.toml changed the dependency-manifest digest that scripts/ci/runner_image_identity.py folds into the runner-image identity lock. Regenerate the lock so tests/ci/test_runner_image_identity.py matches (was the only CI test failure; unrelated environmental integration/perf failures excluded). * test(OMN-13096): update backmerge identity assertions to regenerated lock digests The runner-image identity lock was regenerated for the pyproject entry-point add; this test hardcodes the expected identity_digest/shared_env_digest, so update both to match the new lock (same maintenance OMN-13094 did). --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): tolerate partition-less reply topic in terminal consume leg (#1970) The context-ROI runner opens one ephemeral group_id=None terminal consumer per terminal topic BEFORE publishing each generation command (subscribe-before-publish, OMN-13012/13038). The FAILED reply topic is only produced to on contract_passed=False; in a battery where generations pass it has zero messages, so Redpanda never advertises a partition for it. _assign_direct_terminal_partitions burned the full 30s assign cap on every trial then raised a bare TimeoutError (the empty-message 'wait failed'), stalling each of the 160 battery trials ~30s before the COMPLETED terminal could correlate -> battery needs >80 min and never completes (verifier-confirmed wedge). A partition-less reply topic is a valid steady state, not a 30s error: - _assign_direct_terminal_partitions gives a bounded grace window for a topic that exists but is slow to surface metadata, then assigns whatever partitions exist (possibly none) and returns promptly instead of burning the cap and raising. - poll_direct_terminal_consumer treats an empty assignment as 'no terminal will arrive here' (sleeps out its timeout, returns None) without calling getone() on an unassigned consumer. Repro: tests/integration/test_terminal_consumer_battery_load_wedge.py drives the REAL TerminalEventConsumer over K=10 x 2 cells x 2 arms; RED (40/40 trials block a full assign cap) before the fix, GREEN after. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): pin terminal-consumer read offset synchronously (close lazy seek_to_end publish race) (#1971) * fix(OMN-13118): pin terminal-consumer read offset synchronously to close lazy seek_to_end publish race The consume-leg wedge survived PR #1969 (set_topics no-op) and PR #1970 (partition-less assign-cap) because both addressed the assign phase, not the seek timing. AIOKafkaConsumer.seek_to_end is LAZY: it requests a LATEST offset reset that only resolves on the first poll — AFTER the caller publishes. With generation completing in ~1s, the correlated COMPLETED terminal lands in the open->poll gap, so the lazily-resolved LATEST position is the HWM AFTER the record and the poll reads past it. The terminal is never read, the trial never correlates, and the experiment matrix re-fires the same cell forever. Replace seek_to_end with a synchronous end_offsets() + seek() pin (_pin_direct_terminal_end_offsets) in both open_direct_terminal_consumer and RuntimePatternBBroker._dispatch_and_wait_with_direct_kafka_consumer, so the read position is fixed at open() time, before the publish — the real subscribe-before-publish guarantee. Empty assignment (partition-less reply topic, OMN-13118 #1970) is a no-op. Repro: tests/integration/test_terminal_consumer_seek_reset_race.py drives the real TerminalEventConsumer.open()/wait() the way HandlerContextRoiRunner does, publishing the correlated terminal in the open->wait gap across K=10 x 2 cells x 2 arms; RED with the lazy reset (every cell degenerate), GREEN once the read offset is pinned. RED verified by git-stashing only the source fix. Existing consume-leg fakes updated to model end_offsets/seek (they previously masked the bug by making seek_to_end a synchronous exact snapshot). * test(OMN-13118): reword assertion (lazy not deferred) for receipt honesty gate * fix(OMN-13118): bound end_offsets() round-trip with assign-cap timeout (CodeRabbit) end_offsets() is a broker ListOffsets round-trip aiokafka documents as able to block indefinitely. Bound it with the same cap as start()/assign so a stalled broker fails fast instead of hanging the pre-publish positioning. --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13137): _validate_routing strategy-aware event_model checks (#1973) operation_match routes by the `operation` field and does not use event_model. The validator was unconditionally requiring event_model.{name,module} for every handler entry, causing 230/295 omnimarket operation_match contracts (all correct as authored) to fail routing validation at startup. Fix: read routing_strategy from the routing map and branch validation: - payload_type_match → require event_model.{name, module} (unchanged) - operation_match (and any non-payload strategy) → require `operation`; skip event_model checks entirely Updated pre-existing _validate_routing tests to declare routing_strategy: payload_type_match explicitly (they always tested payload_type_match semantics but relied on the implicit fallback that is now removed). Added test_validate_routing_operation_match.py with 4 unit tests: 1. operation_match without event_model → zero event_model errors 2. operation_match missing operation field → error 3. payload_type_match missing event_model → still errors (regression guard) 4. Real node_integration_sweep_orchestrator routing block → clean (boot gate) Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): independent per-terminal-topic consumers in Pattern B direct-Kafka wait (#1972) The consume-leg wedge survived four merged fixes (#1969 set_topics no-op, #1970 partition-less assign-cap, #1971 synchronous seek-pin). The STRONG K>=10 multi-cell reprobe on the stability lane still wedged on REBUILD-5 (cdf53d963f7b). Converged diagnosis (strikes 3+4, docs/evidence/2026-06-12-weekend-pass/experiments/probe4-stability/ reprobe-K10-rebuild5/HALT_K10_WEDGE_PERSISTS.md): the runtime waited for each trial's terminal across TWO topics (node-generation-completed.v1 + node-generation-failed.v1) with a SINGLE ephemeral group_id=None consumer assigned both topics' partitions. One aiokafka consumer holds one manual subscription; the COMPLETED delivery window collapsed before it surfaced the correlated record, so the trial never correlated and the matrix re-fired cell 1. RuntimePatternBBroker._dispatch_and_wait_with_direct_kafka_consumer now opens ONE independent AIOKafkaConsumer PER terminal topic via open_direct_terminal…
….5 promotion (#2745) * chore(OMN-12934): continue dev to main promotion for prod redeploy (#1988) * docs(OMN-12962): contract-store durability audit — cold-runtime census proof (#1937) Proves cold-start contract census reconstructs from the image-bundled filesystem manifest (HYBRID-mode bootstrap + PluginLoaderContractSource), independent of node-registration.v1 retention. The delete-retention topic feeds only the post-freeze dynamic listener (auto_offset_reset=latest, no history replay). Live .201 stability-test evidence: filesystem contract_path in manifest + truncated topic log head with intact census. No store fix needed; residual dynamic-only gap covered by runtime_sweep sweep check. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12970): vendor omnimarket projection node migrations into forward/nodes (#1938) Vendors omnimarket node-owned projection migrations into the namespaced forward-migration tree so run-forward-migrations.sh materializes them in the dashboard projection DB (omnidash_analytics) at deploy. Primary (OMN-12970): creates llm_call_metrics, llm_cost_aggregates, and capability_scores in the projection DB. These were only ever created in the omnibase_infra DB by infra migrations 031/060, so the ab-compare, cost.token_usage, cost.summary, and capability-scores projection topics were DEGRADED at startup ('table not found') and their dashboard panels rendered empty. Also re-syncs three omnimarket node migrations the vendor tree had drifted from (node_projection_llm_routing, node_projection_overnight, node_projection_savings /077) — sync-node-migrations.sh --check requires the full vendored tree to match omnimarket source, and these were missing. Companion to omnimarket PR for the same ticket (source migrations + projection table-coverage ratchet test). Evidence-Ticket: OMN-12970 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12965): stamp runtime image identity (version + revision) in workspace builds (#1943) * fix(OMN-12965): stamp runtime image identity (version + revision) in workspace builds The main runtime image stamped org.opencontainers.image.version=0.1.0 with a blank org.opencontainers.image.revision after workspace rebuilds. A blank identity degrades every proof packet (runtime SHA + image digest are required citations in accepted evidence). Root causes (three build paths under-stamped identity): - onex up --build (cmd_up) passed only GIT_SHA; the runtime-stage OCI labels read VCS_REF (-> blank revision) and RUNTIME_VERSION (-> placeholder 0.1.0). - deploy-runtime.sh passed VCS_REF but not RUNTIME_VERSION/GIT_SHA. - Dockerfile silently allowed blank/placeholder identity in workspace mode. Fix: - cli._image_identity_build_args() stamps the full quad (GIT_SHA/VCS_REF/ RUNTIME_VERSION/BUILD_DATE) and fails fast on an unresolved git revision. - deploy-runtime.sh stamps RUNTIME_VERSION + GIT_SHA and verifies the version label is non-placeholder post-deploy. - Dockerfile.runtime fails workspace builds with blank VCS_REF or placeholder RUNTIME_VERSION=0.1.0 (release mode unaffected). Enforcement ratchet (same PR): - scripts/check_runtime_image_identity.py static check, wired as pre-commit hook + CI gate (ci.yml). - tests/unit/infra/test_runtime_image_identity_labels.py pins the cli helpers + Dockerfile guard; deploy-agent test extended for the quad. Proven locally via throwaway docker builds: workspace+args -> populated labels; workspace without args -> guard fails (exit 64); release without args -> 0.1.0 placeholder allowed (no regression). Evidence-Ticket: OMN-12965 * test(OMN-12965): integration build proof for runtime image identity labels Builds the real runtime-stage ARG/LABEL/guard block against busybox and asserts via docker inspect: workspace+args -> populated version/revision; workspace without args -> guard fails (exit 64); release -> placeholder allowed. Satisfies the integration-test hard gate and makes the throwaway proof permanent. Evidence-Ticket: OMN-12965 --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12987): workspace-build sibling lock-pin preflight + provenance (#1944) Recurrence ratchet for the 2026-06-11 stability bootstrap crash. The 11:20Z --no-cache rebuild vendored omnibase_infra 0.37.0-dev (~2c1d672f) + core 0.42.0 even though omnimarket dev's uv.lock pinned infra 0.38.1 @ e2dbdc95 + core 0.44.0 @ c97c2c9a. The stale sibling predated the OMN-12501 Protocol-quarantine guard, turning a latent contract defect into a fatal crash that crash-looped the main runtime. - check_sibling_lock_pins.py: host-side fail-fast preflight resolving expected sibling versions/SHAs from the consuming repo's (omnimarket) uv.lock and comparing them against each vendored tree. Mismatch aborts the build. - stage_workspace.sh: emit a .build-sha marker per staged sibling (rsync drops .git) so the preflight and provenance can identify the vendored commit. - deploy-runtime.sh: run the preflight after staging, before build; abort on mismatch. Write the comparison under sibling-repos/.sibling-lock-pins.json. - compute_workspace_provenance.py + Dockerfile.runtime: fold expected-vs-actual lock_pin_comparison into build-provenance.json for deploy verifiers. Recurrence-guard tests prove a stale infra 0.37.0 vs lock-pinned 0.38.1 fails the preflight and matched pins pass; deploy script wiring is asserted statically. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12988): pin stability runtime-worker to replicas 1 + census ratchet (#1942) The base docker-compose.infra.yml defaults runtime-worker to replicas 0 (${WORKER_REPLICAS:-0}). The stability-test lane's required state includes a running worker (4-container census: main, effects, worker, projection-api), but the override pinned it via an env-interpolation default (${STABILITY_TEST_WORKER_REPLICAS:-1}) — a silent-drop surface: a stray exported STABILITY_TEST_WORKER_REPLICAS=0 or removal of the :-1 fallback would scale the worker to 0 with zero signal on a plain compose up/recreate. Fix: pin docker-compose.stability-test.yml runtime-worker deploy.replicas to the literal 1 (no env indirection). Ratchet (recurrence guards, same PR): - scripts/deploy-agent/tests/unit/test_runtime_worker_census.py: assert runtime-worker stays in the deploy-agent RUNTIME-scope census so a missing worker (replicas 0 => absent from docker compose ps) is a deploy failure, not silence; assert the override pins a literal 1. - tests/integration/infra/test_stability_test_runtime_compose_render.py: assert the rendered stability worker resolves deploy.replicas == 1. - tests/unit/infra/test_stability_test_runtime_lane.py: update the existing pin assertion to the literal 1. Evidence-Ticket: OMN-12988 Config-drift family: OMN-12945 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12979): expire-bound topic completeness suppressions (#1940) Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12966): accept community-edition Infisical /api/status in provision readiness gate (#1939) The fresh-provision readiness gate in provision-infisical.py only accepted the enterprise {"status": "ok"} payload and rejected the community edition's {"message": "Ok"}, returning 1 before bootstrap could run. This blocked provisioning against the Infisical instance deployed on .201 (community edition). Route the gate through the existing _is_infisical_ready helper (single source of truth, already used by the already-provisioned path). Add TestMainFreshProvision- ReadinessGate covering community/enterprise/not-ready cases. Also adds docker/docker-compose.infisical-stability.yml: an ADDITIVE Infisical compose project for the stability-test lane (joins the existing network as external, reuses stability postgres/valkey, no lane mutation), since the lane overlays disable the in-lane Infisical service via *-disabled profile overrides. P1.2b-A: Infisical now reachable from the stability runtime/effect containers; a known secret (OMN_12966_PROBE) seeds and resolves end-to-end via the machine- identity universal-auth path, verified from inside the effects container. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12958): volume-config drift gate + runtime config provenance (#1941) * feat(OMN-12958): volume-config drift gate + runtime provenance Compute config provenance (path + sha256) for the runtime-rendered Bifrost delegation contract; the deployed volume copy survives rebuilds and silently diverges from packaged source (two competing authorities, OMN-12945). - runtime/config_provenance.py: ModelConfigProvenance + drift classification, sidecar JSON writer (read by sweep + proof packets) - runtime/health/health_config_provenance.py: drift -> degraded health - render entrypoint logs provenance line + writes sidecar on every boot - docs/runbooks/volume-config-drift-and-reseed.md: ledgered re-seed procedure - validation exemption for config_name (logical identifier, not entity ref) No live volume mutation: re-seed is an operator deploy step (deploy_pending). * test(OMN-12958): cover volume config drift reseed flow --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12957): wire runtime-profile validator + core registry parity guard (#1945) * feat(OMN-12957): wire runtime-profile validator + core registry parity guard - Remove the import-time RuntimeError drift raise in runtime_profile.py (a hard raise on core/infra version skew would crash the kernel at import); the parity invariant is enforced by test_profiles_match_core_registry instead. - Add tests: _PROFILES keys == omnibase_core REGISTERED_RUNTIME_PROFILES, and every CONSUMER_ATTACHED_RUNTIME_PROFILES profile loads as a real lane. - Wire omnibase_core.validation.validator_runtime_profiles as a pre-commit hook + validator-runtime-profiles.yml CI gate on infra contracts. - Freeze 19 pre-existing violators in validation/runtime_profiles_allowlist.yaml (discovered by repo-root walk; drain via OMN-12982). Blocks NEW orphans. Requires the omnibase_core pin to include OMN-12957's validator (new rules). Evidence-Ticket: OMN-12957 Evidence-Source: 5463fbaf819409d4fb7f491dd4f276f10d869eba * ci(OMN-12957): pass runtime profile allowlist to validator * test(OMN-12957): cover runtime profile registry parity * fix(OMN-12957): keep runtime profile allowlist under config * fix(OMN-12957): pin core runtime profile registry --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12973): widen migration-gate healthcheck start_period + ratchet (#1950) P2.8: classify the prod migration-gate UNHEALTHY-then-self-resolved incident. Classification: idle-one-shot-mis-modeled = NO. The migration-gate is a correctly-modeled long-running sentinel (entrypoint `while true; sleep 3600`) whose healthcheck continuously polls db_metadata.migrations_complete via check_migrations_complete.sh. The container flipped UNHEALTHY transiently because its healthcheck start_period (10s) was far shorter than the real cold-volume migration window (~116s: prod gate started 09:35:22, intelligence-migration finished 09:37:18). Past the 10s grace window the still-failing probe was reported UNHEALTHY until migrations completed, then self-resolved — no fault. Fix: raise migration-gate healthcheck start_period 10s -> 180s in both the authoritative catalog manifest (docker/catalog/services/migration-gate.yaml, flows into the generated compose) and the hand-maintained docker/docker-compose.infra.yml that deploy-runtime.sh applies to .201, so a still-applying gate stays in `health: starting` instead of flipping UNHEALTHY. Ratchet (enforcement, not detection): new ValidatorHealthcheckStartPeriod in omnibase_infra catalog (validator_healthcheck_start_period.py) asserts a 120s floor for migration-completion gates, wired into `onex validate runtime` (cmd_validate_runtime) AND backed by unit tests that gate every PR via pre-commit + CI. A migration-completion gate is identified by BOTH a healthcheck that polls migration completion AND a service_completed_successfully dependency, so ordinary app services (e.g. intelligence-api, an HTTP liveness probe with a legitimately short 40s start_period) are not swept into the floor. Prod probed read-only only; no prod mutation. Applies to live prod via the batched stability/prod rebuild (deploy_pending). Evidence-Ticket: OMN-12973 Evidence-Source: OCC#PENDING Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12971): deliver Vertex ADC bearer token to runtime effects container (secret-ref discipline) (#1949) * feat(OMN-12971): deliver Vertex ADC bearer token to runtime effects container (secret-ref discipline) Wire the Vertex $500-credit ADC path into the runtime, ADDITIVE next to the Gemini API-key path (provider-agnostic; neither provider removed or forced). runtime_policy.contract.yaml: add a secret-source mapping llm.vertex.access_token (source_type=file, /run/secrets/vertex_access_token) to the dev/stability-test/ judge profiles, alongside the existing llm.gemini.api_key env mapping. The ref name MUST match cloud-vertex-gemini.secret_ref in omnimarket bifrost_delegation.yaml. The resolved VALUE is a short-lived OAuth bearer token minted from ADC, refreshed by the operator; the token VALUE is never committed — only the ref name + in-container path. Add aiplatform.googleapis.com to the cloud host allowlist. docker-compose.infra.yml: bind the operator-supplied host token file read-only to /run/secrets/vertex_access_token on the main and effects runtimes (VERTEX_ACCESS_TOKEN_HOST_FILE, default /dev/null so lanes without Vertex still start; Gemini key path unaffected). Pass through BIFROST_VERTEX_GEMINI_ENDPOINT_URL (overlay supplies the complete Vertex OpenAI-compat URL) and GOOGLE_CLOUD_PROJECT/LOCATION (default empty). runtime-policy.env: regenerated from the contract via render_runtime_policy_env (test_runtime_policy_env_matches_contract_renderer proves contract<->env parity). test_runtime_policy_contract.py: update host-allowlist assertion for the additive Vertex host. Pre-existing platform-wide topic-parity-gate failures (25 unrelated topics) are identical with and without this change (proven by diff); that gate is pre-commit- only (not a CI merge gate) and this change adds zero new topic gaps, so that one hook is SKIP-ped. No deploy/receipt/merge gate is bypassed. * fix(OMN-12971): make Vertex runtime env contract-owned --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13008): automated .201 disk maintenance — worktree GC + docker GC + watermark alert ratchet (#1952) * feat(OMN-13008): automated .201 disk maintenance — worktree GC + docker GC + watermark alert ratchet Conservative, keep-list-driven disk maintenance to prevent the 2026-06-11 /data ~95% outage that killed all three lanes mid-demo: - scripts/worktree-gc.sh: drives the canonical omniclaude prune-worktrees.sh (merged+clean+pushed safety) on both Mac (merge-sweep tick) and .201 - scripts/disk-gc.sh + disk_gc_plan.py: conservative docker/builder/image GC. Pure, testable removal planner honoring a VERSIONED keep-list (deploy/disk-gc/keep-list.yaml): never reaps a kept repo, kept tag, in-use image, or anything younger than min_age_days; keeps N superseded generations. - scripts/disk-watermark-check.sh + disk_watermark_event.py: df watermark ratchet. >=85% emits a typed disk-watermark bus event (warning) that the sweep auto-ticket path turns into a Linear ticket; >=90% emits critical. Broker addr is fail-fast from KAFKA_BOOTSTRAP_SERVERS (no localhost default). - deploy/disk-gc/: systemd USER timer (onex-disk-gc.timer/.service, hourly) + install-disk-gc.sh. User units, NOT lane containers. - tests: 20 unit tests incl. GC plan-safety invariants + dry-run proof that default mode issues no destructive op (a wrong-delete GC is worse than none). Contract: contracts/OMN-13008.yaml * fix(OMN-13008): pass docker inventory to GC planner via stdin, not env (ARG_MAX) On a host with many docker images, passing the full image/ps inventory as env vars to disk_gc_plan.py exceeds ARG_MAX ('Argument list too long', exit 126), producing an empty plan. Write inventory to per-run scratch files (under the log dir, never /tmp; cleaned on exit) and hand it to the planner on stdin as a JSON envelope. Verified the failure live on .201; planner now reads stdin. * fix(OMN-13008): simplify GC plan stdin pipe (two processes, no nested subprocess) * fix(OMN-13008): keep-wins reconciliation — never remove an image id with any keep reason A single image id can surface in multiple 'docker image ls' rows (one per repo:tag). One tag could route the id to dangling-removal while another routes it to keep (e.g. tagged 'latest' or within-N-generations). Live .201 plan showed an id in BOTH remove_image_ids and kept_reasons. Reconcile at the end: keep wins — any id with a keep reason is dropped from the remove list; remove list deduped. Adds 2 regression tests. Verified live on .201. * fix(OMN-13008): timer uses OnCalendar=hourly for reliable re-arm (oneshot service) OnUnitActiveSec does not reliably re-elapse for a oneshot service once it goes inactive (observed NextElapseUSecMonotonic=infinity live on .201 after the first run). Switch to OnCalendar=hourly + RandomizedDelaySec so the timer re-arms every hour. Keep Persistent=true for missed-run catch-up. --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13005): materialize blocking event_consumer in runtime auto-wiring (runner consume-leg degenerate rows) (#1951) * fix(OMN-13005): materialize blocking event_consumer in auto-wiring (runner consume-leg degenerate rows) The runtime auto-wiring materialized event_publisher for handlers that declare it but had no equivalent for event_consumer. Request/response EFFECT handlers (HandlerContextRoiRunner) that publish a command then block on the correlated terminal event fell back to their no-op consumer default, returning None immediately -> every result row degenerate (failure_stage=generation, attempt_count=0) while generations succeeded ~1s later. Adds _make_sync_event_consumer (mirror of _make_sync_event_publisher), backed by service_terminal_event_consumer.make_terminal_event_consumer: a sync (topic, correlation_id, timeout) -> dict | None adapter that runs the proven direct-Kafka correlate-and-wait loop (from RuntimePatternBBroker) on an isolated event loop in a worker thread, so blocking does not deadlock the runtime dispatch loop that delivers the awaited terminal. TDD through the REAL dispatch path: test_event_consumer_injection drives a trial through _prepare_handler_wiring with a terminal arriving after a delay and asserts a non-degenerate row; verified RED with injection disabled. * fix(OMN-13012): two-phase (seek-now/wait-later) terminal event_consumer to close the subscribe-after-publish race (#1954) The OMN-13005 injected event_consumer is a single callable that does assign -> seek_to_end -> poll internally, all AFTER the handler has already published its command. Once OMN-13010 freed the dispatch loop and generation began completing in ~1s, the correlated terminal lands BEFORE the single-call consumer's post-publish seek_to_end positions, so seek_to_end skips PAST the already-emitted terminal and the runner times out on an offset beyond it (probe3, run_id=20260611T2140Z-probe3 -- both arms degenerate, zero rebalances). Splits positioning from waiting so the caller subscribes BEFORE it publishes: session = consumer.open(topic) # assign + seek_to_end NOW publisher(command_topic, payload) # publish AFTER positioning payload = session.wait(cid, timeout) # block from the captured position The returned TerminalEventConsumer is still directly callable with the legacy (topic, cid, timeout) -> dict | None single-call shape for any consumer that does not need subscribe-before-publish; the runner is the only consumer today. TerminalConsumerSession owns a dedicated event loop on a daemon worker thread for the whole open->wait->close lifecycle, preserving the OMN-13005 loop-isolation discipline so blocking never deadlocks the runtime dispatch loop. TDD (real dispatch path, RED-then-GREEN): extends the OMN-13005 injection test with a terminal emitted IMMEDIATELY after publish. The single-call (seek-after-publish) consumer MISSES it (degenerate row -- RED test asserts failure_stage=generation); the two-phase (open-before-publish) consumer CATCHES it (non-degenerate -- GREEN). The Kafka layer is faked at the two service seams against a shared in-memory log modeling seek-to-end semantics. OMN-13005 blocking-correlate behavior preserved. 269/269 auto_wiring unit tests pass; mypy --strict clean. Sibling to OMN-13010 / OMN-13005 / OMN-13003. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13005): route terminal consumer through Kafka boundary --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12990): pin worker replicas in ledgered config + fail-fast lane overrides (silent-drop ratchet) (#1948) * fix(OMN-12990): pin worker replicas in ledgered config + fail-fast lane overrides The base compose sets runtime-worker deploy replicas to ${WORKER_REPLICAS:-0} (soft-default ZERO). The stability lane's required state includes a running worker (GATE_ZERO_PROOF.md: 4 runtime containers), but the worker presence rode entirely on a compose soft-default — any plain compose up/recreate without the policy env silently scaled the worker to zero with no error and no signal. Fix (contract-native + fail-fast): - Add 'replicas' to ModelRuntimeProcessPolicy; pin replicas: 1 in every lane's worker block in runtime_policy.contract.yaml. - Renderer emits {PROFILE}_WORKER_REPLICAS into the ledgered runtime-policy.env for dev/stability-test/judge/prod. - stability + prod compose overrides reference ${..._WORKER_REPLICAS:?...} (fail-fast, NO silent :-1/:-0 default). A recreate that omits the policy env now aborts loudly instead of dropping the worker. prod previously had no override at all and inherited the dangerous :-0 default. Ratchet (recurrence guards): - tests asserting fail-fast override form (no soft default), contract-declared replica pin >= 1 per lane, and rendered {PROFILE}_WORKER_REPLICAS=1 in the ledgered env. - runbook deploy/verify procedure adds an expected-container census (worker must be present) via verify_container_manifest; a missing worker is a FAILURE, not silence. Note: SKIP=topic-parity-gate — that local-only advisory gate (absent from all .github/workflows, not a required CI check) fails on 25 pre-existing cross-repo topic gaps (build-loop/omniclaude/omniweb) identical on pristine base HEAD 8d7da1249; this change adds zero topics. All other hooks ran clean. Evidence-Ticket: OMN-12990 * test(OMN-12990): cover worker replica policy integration --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12909): add gateway bus forwarder P0A (#1946) * feat(OMN-12909): add gateway bus forwarder p0a * test(OMN-12909): add gateway forwarder integration coverage * fix(OMN-12909): satisfy gateway forwarder validators * test(OMN-12909): allow gateway forwarder bus protocol * fix(OMN-12909): sync gateway forwarder entry point * fix(OMN-12909): refresh runner image identity lock * test(OMN-12909): relax JSON normalizer mixed benchmark threshold * fix(OMN-12909): allow gateway handlers to boot unconfigured * fix(OMN-12909): declare gateway forwarder runtime profile * fix(OMN-12909): update runner identity backmerge expectation --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13011): LANE CENSUS RECONCILIATION ratchet — declared desired-state per lane, drift = auto-ticket (#1955) The class fix for the recurring lane-drift regression. Nothing reconciled the declared desired state of a runtime lane against what is actually running, so the same failure kept recurring with zero signal: volume config drift (OMN-12945), WORKER_REPLICAS silent zero (OMN-12988/12990), and on 2026-06-11 prod runtime containers plus the broker network were silently absent for hours during demo prep. Ships a per-lane DESIRED-STATE census: - (a) DECLARED in a versioned lane manifest (deploy/lane-census/lane-manifest.yaml): container set, network, replicas, image-tag pattern per lane (stability-test/prod/judge/dev), derived from the canonical compose lane files. A parity ratchet keeps the manifest locked in step with the compose files. - (b) RECONCILED on a schedule on .201 by SHARING the OMN-13008 systemd timer (a drop-in 4th ExecStart on onex-disk-gc.service — never a second timer) and on-demand via scripts/lane-census-check.sh / runtime_sweep. - (c) Drift = typed bus event (onex.evt.infra.lane-census-drift.v1) + Linear auto-ticket naming exactly what is missing/extra (container_absent, network_detached, replicas_zero, unexpected_container, oneshot_failed/stuck, image_tag_mismatch). Fail-fast, no warn-only mode (gates-block policy); exit 30 on drift; bus publish fail-fast on KAFKA_BOOTSTRAP_SERVERS (no localhost default). Red fixture reproduces 2026-06-11: prod runtime containers absent + broker network detached must produce the exact drift findings + a non-zero exit hours before a human noticed. Pure planner is fully unit-tested; shell driver dry-run-tested. Builds on the OMN-12988 deploy-agent RUNTIME census (deploy-time) as the complementary steady-state reconciler; closes the runtime-worker.yaml container_name: null census gap by sourcing names from the compose lane files. Evidence-Ticket: OMN-13011 Config-drift family: OMN-12945 Relates-to: OMN-13009, OMN-12988, OMN-13008 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13020): vendor missing node migrations — llm_routing 0000 + context_roi 001 (#1956) Vendors two omnimarket node-source migrations into the infra forward-migration tree via scripts/sync-node-migrations.sh (the canonical OMN-12559 mechanism): - node_projection_llm_routing/0000_create_llm_routing_decisions.sql (source: omnimarket #1168 / OMN-12942, merge ed6734f8) - node_projection_context_roi/001_create_context_roi_scores.sql (source: omnimarket #1178 / OMN-12955, merge 5010b1f4) Without the 0000 base table, node_projection_llm_routing/0001 (CREATE VIEW) hard-fails against NODE_POSTGRES_DB=omnidash_analytics — exactly the prod forward-migration exit-3 of 2026-06-11T09:35:52Z, and reproduced by construction in any clean clones@dev build. Files are byte-identical to the omnimarket dev blobs (sha256 f8a8b339… / c4126e65…) and to the untracked hot-patch copies on the .201 stability clone. Both migrations are self-contained, all-statements-IF-NOT-EXISTS, and 0000 sorts lexically before 0001 within the node's namespaced identity space. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13058): close TerminalConsumerSession on open() failure (worker thread + event loop leak) (#1957) TerminalConsumerSession.__init__ starts its dedicated worker loop thread immediately. TerminalEventConsumer.open did 'session = TerminalConsumerSession(...); return session.open()' with no cleanup: any failure inside session.open() (consumer start timeout, partition-assign timeout, broker auth error) propagated out of the raising expression, the session reference was lost, and the daemon worker thread plus its never-closed asyncio event loop leaked -- one pair per failed open. The motivating caller (HandlerContextRoiRunner) opens a session per trial, so a 160-560-trial battery against a degraded broker accumulates hundreds of leaked threads in the long-lived effects container. Fix: wrap session.open() in try/except BaseException -> session.close() (idempotent: stops the loop, joins the thread) -> re-raise. Covers both the two-phase .open(topic) path and the legacy single-call __call__ path. Found by the P3.3 doctrinal review of merged #1951 (b9712af9 / 36d98275). TDD: tests/unit/runtime/test_service_terminal_event_consumer_open_failure.py injects a real open failure through the production path (event bus without _bootstrap_servers) and asserts no alive terminal-consumer-* thread after the raise. Verified RED with the fix stashed (2 failed), GREEN with it (2 passed). Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13021): non-dev-base guard — fail feature-base PRs absent Stacked-Parent declaration (retro A-6) (#1958) Any PR whose base is neither dev nor main fails unless the body carries 'Stacked-Parent: #N'. Prevents the feedback_stacked_prs_orphan_from_dev class (#1185/#1954 auto-merged INTO parent feature branches and stranded off dev). base=main remains governed by main-target-guard. Epic OMN-13013 (process enforcement ratchets — June 12 retro). Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13014): hot-patch ledger rebuild preflight gate (retro B-1) (#1959) * feat(OMN-13014): hot-patch ledger rebuild preflight gate (retro B-1) Hot-patches on .201 (.prepatch sibling discipline) silently revert on any image rebuild/force-recreate — the 2026-06-11 20:58Z rebuild already erased a live /api/generate patch once. This adds the rebuild-path gate: - scripts/preflight_hotpatch_ledger.py: given a target container or lane + per-repo build refs, hard-fails when any hot-patch ledger row's source PR merge commit is not an ancestor of the build ref (git merge-base --is-ancestor), plus a .prepatch tripwire of the running container (unledgered .prepatch = hard fail; --post-rebuild = zero .prepatch expected). Sole bypass: HOTPATCH_PREFLIGHT_BYPASS carrying the Rule-10 '# skip-token-allowed: <user-approval-receipt-id>' form. - scripts/deploy-runtime.sh: guard_hotpatch_ledger wired into main() before build/preview (both dry-run and execute), lane derived from the compose project; skips loudly only when no ledger exists on the host. - tests/unit/scripts/test_preflight_hotpatch_ledger.py: 17 unit tests (ancestor gate, lane scoping, ledger loading, tripwire, bypass forms). - tests/ci/test_receipt_gate_install_guard.py: repair stale guard — core OMN-12565 replaced the OMN-9198 'uv pip uninstall first' install step with a cleared workspace venv (uv venv --clear); assert the new contract. Ledger backfilled from live census (5 .prepatch files / 4 source PRs, all MERGED to dev) at /data/omninode/hotpatch-ledger/ledger.yaml on .201. * fix(OMN-13014): scope missing-prepatch tripwire warning to the probed container --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12977): workspace-build sibling-pin ratchet — honor consuming lock, fail-fast on stale vendor (#1936) * fix(OMN-12977): workspace-build sibling-pin ratchet — honor consuming lock, fail-fast on stale vendor The workspace-mode image build (BUILD_SOURCE=workspace, used by the stability-test deploy procedure) vendored sibling/foundation packages from whatever the canonical OMNI_HOME clones happened to be checked out at, ignoring the consuming repo's uv.lock. On 2026-06-11 this shipped a 13-day-stale omnibase_infra 0.37.0-dev (pre-OMN-12501 Protocol-quarantine guard) + core 0.42.0 against an omnimarket dev lock pinning infra 0.38.1@e2dbdc95 / core 0.44.0@c97c2c9a, dropping the guard and crashing wire_from_manifest bootstrap fatally (stability lane down on demo day). Fix + recurrence ratchet (same PR): - scripts/runtime_build/check_sibling_lock_pins.py: parse the consuming repo's uv.lock for expected version+git-rev of each foundation/sibling package (scoped to the package's own source line so editable/registry pins are not cross-attributed a dependency's rev), resolve the actual clone version+HEAD, compare, and classify drift backward/forward/none. Fail-fast (exit 1) on any drift; --allow-drift records an explicit operator override in the artifact, never silent. - stage_workspace.sh: runs the preflight against the canonical clones before staging; aborts the build (exit 3) on unacknowledged drift and writes workspace/sibling-pin-comparison.json. - compute_workspace_provenance.py: folds the expected-vs-actual comparison into build-provenance.json so deploy verifiers can assert the build honored the lock; flags unacknowledged drift as a provenance error. - Dockerfile.runtime: COPY the comparison artifact (committed placeholder so the COPY always resolves; overwritten by stage_workspace.sh in workspace mode). - TDD: 19 unit tests covering lock parsing (git/registry/editable sources), drift classification, the exact 0.37.0-vs-0.38.1 stale case, check_pins exit codes, and the allow-drift override. Pre-existing mypy-strict bare-dict errors in compute_workspace_provenance.py fixed in the same pass. Evidence-Ticket: OMN-12977 Evidence-Source: pending-occ * ci(OMN-12977): retry runtime smoke compose port race * test(OMN-12977): align sibling-pin script tests with current API --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12989): workspace-mode image build must honor sibling lock pins (#1947) * fix(OMN-12989): workspace build must honor sibling lock pins (fail-fast ratchet) The 2026-06-11 stability bootstrap crash was caused by a workspace-mode --no-cache rebuild that vendored omnibase_infra 0.37.0-dev (13-day-stale worktree) even though omnimarket dev's uv.lock pins infra 0.38.1. The downgraded sibling predated the OMN-12501 Protocol-quarantine guard and turned a latent contract defect into a fatal crash. Fix + ratchet (same PR set): - scripts/runtime_build/resolve_workspace_pins.py: parse the consuming repo's uv.lock for sibling pins (version + git rev); classify each staged/installed sibling as exact/ahead/regression/unpinned; fail-fast (WorkspacePinError) on any regression BELOW the lock pin. Stdlib version-tuple fallback when packaging is absent so the ratchet never fails-open. - compute_workspace_provenance.py: enforce sibling pins + a host-infra self-check (installed omnibase_infra vs lock pin — the exact crash vector, since host infra is built from the context, not staged), and emit a pin_comparison block into build-provenance.json for deploy verifiers. - Dockerfile.runtime: COPY resolve_workspace_pins.py beside the provenance script so the in-image import resolves. - TDD: failing tests first (test_resolve_workspace_pins.py, 11 cases) + provenance integration tests; fixed a pre-existing stale RUNTIME_VERSION assertion to read pyproject dynamically. Evidence-Ticket: OMN-12989 * fix(OMN-12989): workspace build must honor sibling lock pins (fail-fast ratchet) The 2026-06-11 stability bootstrap crash was caused by a workspace-mode --no-cache rebuild that vendored omnibase_infra 0.37.0-dev (13-day-stale worktree) even though omnimarket dev's uv.lock pins infra 0.38.1. The downgraded sibling predated the OMN-12501 Protocol-quarantine guard and turned a latent contract defect into a fatal crash. Fix + ratchet (same PR set): - scripts/runtime_build/resolve_workspace_pins.py: parse the consuming repo's uv.lock for sibling pins (version + git rev); classify each staged/installed sibling as exact/ahead/regression/unpinned; fail-fast (WorkspacePinError) on any regression BELOW the lock pin. Stdlib version-tuple fallback when packaging is absent so the ratchet never fails-open. - compute_workspace_provenance.py: enforce sibling pins + a host-infra self-check (installed omnibase_infra vs lock pin — the exact crash vector, since host infra is built from the context, not staged), and emit a pin_comparison block into build-provenance.json for deploy verifiers. - Dockerfile.runtime: COPY resolve_workspace_pins.py beside the provenance script so the in-image import resolves. - TDD: failing tests first (test_resolve_workspace_pins.py, 11 cases) + provenance integration tests; fixed a pre-existing stale RUNTIME_VERSION assertion to read pyproject dynamically. Evidence-Ticket: OMN-12989 * test(OMN-12989): co-locate provenance pin helper in fixture --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13055): absent repos in REPOS list warn and exit 0 instead of failing (#1960) Missing repos (not cloned locally) now emit a WARN line and are tracked in a separate WARNED array. Only real fetch/ff failures cause exit 1. This makes pull-all.sh safe to use on machines with a partial clone set, while keeping the explicit-list override behavior intact. Adds three regression tests: absent-only exits 0, absent+present exits 0 with OK for the present repo, present-failed+absent exits 1. Evidence-Ticket: OMN-13055 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13070): config_prefetcher overlay wins on controlled lanes (#1961) On infisical_required=True lanes, fetched/overlay config now always wins over ambient env. Ambient env is retained only as a declared bootstrap fallback (with an explicit provenance INFO log line) when Infisical returns None. apply_to_environment also overwrites stale env on controlled lanes. Uncontrolled lane (infisical_required=False) behaviour is unchanged. Adds 5 regression tests: controlled-lane Infisical-wins, env-bootstrap- fallback, apply_to_environment overwrite, missing-from-both-is-error, and uncontrolled-lane-env-still-wins. Refactors _resolve_key to return a (outcome, value, error) tuple to satisfy the ≤5-param pattern gate. Source: docs/audits/2026-06-10-runtime-env-overlay-authority-audit.md Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13062): migration-gate vacuity fix — sentinel discipline, wait-for-postgres, skip-manifest, sync-check nonzero (retro A-10) (#1963) * fix(OMN-13062): migration-gate vacuity fix — sentinel discipline, wait-for-postgres, skip-manifest, sync-check nonzero Fixes three bugs identified in retro A-10 (recurrences OMN-12885, OMN-12934): (1) RUNNER SENTINEL DISCIPLINE run-forward-migrations.sh now clears migrations_complete=FALSE at the start of every run and sets it TRUE only as its FINAL act after all infra and node migrations succeed. Any mid-run failure leaves the gate UNHEALTHY. runner_completed_at is stamped at the same final step as durable evidence of a successful completion run. (2) SYNC-NODE-MIGRATIONS VACUOUS GATE sync-node-migrations.sh --check now exits 2 (not 0) when the omnimarket source tree is unresolvable. Silent exit-0 was hiding drift. The single opt-out is SYNC_NODE_MIGRATIONS_SKIP_UNRESOLVABLE=1 for environments that intentionally run without the source. (3) WAIT-FOR-POSTGRES GUARD run-forward-migrations.sh now waits up to PG_WAIT_RETRIES (default 30) x 2s for Postgres to accept connections before proceeding, guarding the first-boot initdb race. (4) SKIP-MANIFEST docker/migrations/skip-manifest.yaml introduced as the sole committed escape for intentionally-skipped migrations. The runner reads this at startup; listed migrations are recorded in schema_migrations with checksum "skip-manifest" without executing the SQL. (5) MIGRATION 085 Adds runner_completed_at TIMESTAMPTZ column to db_metadata so the runner's final stamp is durable in the schema (idempotent ADD COLUMN IF NOT EXISTS). Rollback included. 22 regression tests added covering all five fix surfaces. * fix(OMN-13062): stamp schema fingerprint for migration 085 Migration 085 (085_add_runner_completed_at_to_db_metadata.sql) was added in the initial commit but schema_fingerprint.sha256 was not regenerated. Running `python scripts/check_schema_fingerprint.py stamp` updates the artifact from the stale hash to match the 71 migration files. Evidence-Source: OCC#2563 Evidence-Ticket: OMN-13062 --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12864): Bifrost endpoints → committed overlay authority + fail-loud loader (#1964) * feat(OMN-12864, OMN-12814, OMN-12945): Bifrost endpoints → committed overlay authority + fail-loud loader OMN-12864 — Committed lane overlay - docker/lane-overlays/dev.bifrost.yaml: typed deployment bindings for all four BIFROST_LOCAL_*_ENDPOINT_URL values (coder :8000, reasoner :8001, embedding :8100, ds4-flash :8101). Previously only available as ephemeral shell exports on .201; now committed, auditable, diff-able, and CI-checked. - docker/lane-overlays/dev.bifrost.env: generated dotenv sidecar consumed by compose via env_file; never edited directly (yaml is authority). - docker/docker-compose.infra.yml: wire the env_file block at compose root so the four endpoints are injected into the interpolation context on a clean shell. Hardcode BIFROST_CONTRACT_PATH (remove :-/empty footgun — OMN-12814). - scripts/render_bifrost_lane_overlay_env.py: render script regenerates the env sidecar from the YAML source. - src/omnibase_infra/runtime/models/model_bifrost_lane_overlay.py: ModelBifrostLaneOverlay — typed Pydantic model enforcing URL completeness (OMN-12815: every URL must end in /chat/completions). OMN-12814 — Fail-loud loader - render_bifrost_delegation_contract: raises ProtocolConfigurationError on FileNotFoundError, YAMLError, ValidationError, and zero-endpoint renders. No lru_cache — every restart re-renders from packaged source so a stale cache cannot pin a broken result across deploys. OMN-12945 — Re-seed from packaged source on deploy - docker/entrypoint-runtime.sh: set BIFROST_FORCE_RESEED=1 on every container restart so the named-volume copy is always rebuilt from the packaged bifrost_delegation.yaml merged with committed lane-overlay endpoints. - render_bifrost_delegation_contract: honor BIFROST_FORCE_RESEED/force_reseed flag to bypass the stale-volume early-return path entirely. Tests: - tests/ci/test_bifrost_lane_overlay.py: CI gate — env sidecar in-sync with YAML source; all four BIFROST_LOCAL_* keys present. - tests/unit/runtime/models/test_model_bifrost_lane_overlay.py: bare-base URL rejection, env dict mapping, extra-field rejection. - tests/unit/runtime/test_render_bifrost_delegation_contract.py: fail-loud paths, force-reseed, zero-endpoint error, endpoint URL completeness. - tests/unit/models/test_model_serialization_roundtrip.py: roundtrip coverage. * fix(OMN-12864): move bifrost env_file to service level — fix compose schema validation failure Top-level 'env_file' is rejected by Docker Compose v2 schema validator ('additional properties not allowed'). This caused 10+ compose-render integration tests to fail in CI. Fix: - Remove top-level env_file block from docker-compose.infra.yml - Add per-service env_file on omninode-runtime, runtime-effects, runtime-worker (the three containers that render Bifrost) - Change BIFROST_LOCAL_*:? to BIFROST_LOCAL_*:- in x-runtime-env (compose-level validation removed; Python validates via ModelBifrostLaneOverlay + render_bifrost_delegation_contract) - Add two CI gate tests: compose_env_file_is_service_level_not_top_level and runtime_services_have_bifrost_env_file * fix(OMN-12864): allow BIFROST_LOCAL_* empty defaults in silent-fallback gate The overlay authority pattern (OMN-12864) passes BIFROST_LOCAL_*_ENDPOINT_URL via service-level env_file (docker/lane-overlays/dev.bifrost.env), not at compose config time. Compose-level :? would break CI rendering without the overlay pre-loaded. Validation at the Python layer (ModelBifrostLaneOverlay + render_bifrost_delegation_contract) is the enforcement point. Add the three failing vars to ALLOWED_EMPTY_DEFAULTS with OMN-12864 citation. --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12858/OMN-12879/OMN-12880): dispatcher route-coverage CI gate (#1965) * feat(OMN-12858/OMN-12879/OMN-12880): dispatcher route-coverage CI gate Static contract analysis: every subscribed command topic (onex.cmd.*) must declare handler_routing or runtime_dispatch, or the message goes to DLQ silently. This gate would have caught two recent incidents: 1. June 9 DLQ regression (OMN-12858 post-mortem): node_generation_consumer subscribed onex.cmd.omnimarket.node-generation-requested.v1 but a sole-handler revert left zero dispatcher routes registered. Messages went to DLQ silently with no CI signal. 2. June 12 DEL-01 live finding: onex.cmd.omnimarket.delegate-skill.v1 was consumed by dev lane bus but no dispatcher route existed in any deployed contract. Discovered via manual rpk consumer-group lag probe (DEL-01 evidence, docs/evidence/2026-06-12-weekend-pass/). Deliverables: - scripts/check_dispatcher_route_coverage.py — static YAML scanner that checks both omnibase_infra and omnimarket contract trees; ratchet allowlist for known pre-existing violations; --changed-contracts mode (OMN-12879) for per-PR scoping; compat publish topics excluded (OMN-12880) - .github/workflows/dispatcher-route-coverage.yml — CI workflow that checks out omnimarket sibling, collects changed contract paths in PR mode, and runs the gate; fires on PR, push-to-main, and merge_group - tests/ci/test_dispatcher_route_coverage_gate.py — 12 unit tests covering RED/GREEN/COMPAT/CHANGED-MODE/ALLOWLIST/MULTI-DIR paths plus live-contract regression proof against the actual omnibase_infra tree Allowlist additions: - onex.cmd.omnibase-infra.pattern-b-dispatch.v1 (RuntimePatternBBroker, imperative consumer, OMN-12525 migration target) - onex.cmd.platform.contract-resolve-requested.v1 (transitional HTTP bridge node_contract_resolver_bridge OMN-2756, metadata.transitional=true) [OMN-12858, OMN-12879, OMN-12880] * fix(OMN-12858): drop full uv sync from dispatcher-route-coverage workflow Gate script only needs pyyaml (stdlib + yaml). Using full setup-python-uv was causing 10+ minute timeout. Replace with direct pip install pyyaml and invoke python3 directly. Reduces job from 10m timeout to <1m. [OMN-12858] --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13026): wire transport-mock-lint pre-commit hook + CI gate (#1962) * feat(OMN-13026): wire transport-mock-lint pre-commit hook + CI gate (omnibase_infra) Activates the transport-mock-lint validator (from omnibase_core, OMN-13026) on omnibase_infra. Ratchet baseline: 218 existing violations across 63 files frozen in validation/transport_mock_baseline.yaml. New bare AsyncMock/ MagicMock on EventBus/transport surfaces are blocked by pre-commit hook and CI lint step. Existing violations tracked for drain by per-site tickets (parent OMN-13026). Reference incident: PR #1181 bare AsyncMock hid missing EventBusKafka.stop(). Evidence-Ticket: OMN-13026 * fix(OMN-13026): use uv run python for CI lint step + bump omnibase_core pin to include transport_mock_lint The transport-mock lint CI step previously cloned omnibase_core and ran `python -m omnibase_core.validators.transport_mock_lint` with PYTHONPATH, but this failed: `No module named omnibase_core.validators.transport_mock_lint` because it ran `.venv/bin/python` which uses the locked venv, and the venv omnibase_core pin (2defabef4) predates the transport_mock_lint module. Fix: use `uv run python` (removes the clone step) and bump omnibase-core git pin from 2defabef4 to 309d89fa7 (PR 1231 merge commit on dev) so transport_mock_lint is available in the locked venv. * fix(OMN-13026): align transport mock baseline and runner lock * fix(OMN-13026): sync omnibase_core pin + runner identity lock to dev baseline Align pyproject.toml omnibase_core rev to 2defabef (required by test_release_backmerge_preserves_proven_runtime_core_pin) and update docker/runners/runner-image.lock.json identity_digest/shared_env_digest to match dev runner image lock (79b08f44 / 90c8b3b9). Both were stale from the prior session's pin bump that used an older SHA. * fix(OMN-13026): source transport mock validator from core * fix(OMN-13026): source transport validator from core dev --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13094): --output receipt mode on onex node/run — quiet typed receipts with durable capture (#1966) Phase 2a of the skill-output-suppression slice (epic OMN-13089, plan docs/plans/2026-06-12-skill-output-suppression-plan.md Phase 2 item 1). - onex node/onex run gain --output receipt: ALL runtime logging routes to a run_id-suffixed capture file under <state-root>/captures/ (no console handlers — kills the 25-50-line RuntimeLocal INFO stream at the source); stdout carries exactly ONE typed ModelSkillResult JSON with the FULL handler result (result_model = concrete handler result type FQN). - Durable capture: capture log + handler result content-addressed via omnibase_core ArtifactStore (OMN-13093); artifact.captured + tool.output.captured emitted to the emit daemon socket (--emit-socket, default ~/.claude/emit.sock). - Failure asymmetry: artifact write failure => FULL output printed, no receipt (no hidden loss); emission failure => receipt still prints, event spooled to <state-root>/emit_spool/ for replay. - Node failure => status=failed/error with full error + capture log INLINE in the receipt (errors are never hidden) and artifact-backed. - Default output mode unchanged (enforcement is Phase 4). - RuntimeLocal exposes handler_result (receipt schema identity). - core pin 2defabef -> ae8793bd (merged OMN-13091/13093 receipt models + ArtifactStore); runner-image identity lock regenerated and the OMN-12765 backmerge identity constants updated for the new pin. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13097): onex skill subcommand + declarative skill->node mapping (#1968) * feat(OMN-13097): onex skill subcommand + declarative skill->node mapping Phase 4a of the skill-output-suppression slice (epic OMN-13089, plan docs/plans/2026-06-12-skill-output-suppression-plan.md Phase 4 item 1/2). A dispatch skill IS one CLI call (user directive 2026-06-12). This adds the `onex skill <name> [args]` dispatch surface that the 24 omniclaude shim migrations build on: - onex.cli entry-point `skill` -> cli_skill.run_skill_by_name. Resolves the skill via the declarative skill_mapping.yaml registry, builds the backing node's input payload from the skill's CLI args, writes it under the state root (.onex_state/tmp/<skill>-<run_id>.json — never /tmp), resolves the node's packaged contract exactly like `onex node`, and dispatches through the proven receipt-mode path (run_receipt_mode, OMN-13094). stdout is exactly one typed ModelSkillResult JSON with the FULL handler result. - skill_mapping.yaml: declarative DATA mapping all 24 dispatch shims to their backing onex.nodes node + typed result-model FQN (verified against each live handler handle() return type on origin/dev) + per-arg payload specs + static payload + keyword classifiers (delegate task_type as data, not code). Adding a skill is a YAML edit + fixture, never a CLI code change (ticket deliverable 2/3). Mapping lives beside the node-resolution surface, never hardcoded branching in the CLI. - Typed models split one-per-file (repo convention): ModelSkillArgSpec, ModelSkillClassifier, ModelSkillMapping, ModelSkillMappingRegistry, EnumSkillArgType. Frozen, extra=forbid, fail-fast coercion/validation. - validation_exemptions.yaml: Click-callback param-count + literal-identifier name-field exemptions mirroring the existing cli_node run_node_by_name precedent (OMN-11570) — same pattern, same rationale. dod_evidence: - 20 unit tests pass (registry validity, all-24-shims coverage, FQN result models, arg parsing/coercion/positional/required, classifiers, payload build, receipt-mode dispatch wiring, payload-under-state-root not /tmp). - uv run mypy src/ --strict: clean (2438 source files). - ruff format + check: clean. pre-commit run on changed files: pass. - runner-image identity lock regenerated for the pyproject entry-point add (same as OMN-13094). * test(OMN-13097): rebind OMN-12765 backmerge identity constants for onex skill pyproject change Adding the `skill` onex.cli entry-point to pyproject.toml changes the runner-image identity_digest (and shared_env_digest) the lock binds. Update the hardcoded expected constants in the backmerge-identity test to the regenerated values — same mechanical rebind OMN-13094 performed for the core pin bump. Identity + runner-image-identity tests pass (14). --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix+test(OMN-13012): force terminal-topic metadata refresh so both ephemeral consumers assign (runner consume-leg wedge) (#1969) The runner consume-leg wedged on the live stability battery (image c0505521f1fa, EXP1-3_RUNNER_CONSUME_LEG_BLOCKER): only the FAILED terminal topic ever subscribed while the COMPLETED topic never assigned, so the correlated completed terminal was never read and the 8x2x10 matrix re-fired cell 1 forever, emitting zero non-degenerate rows. A prior two-strike diagnosis proved the omnimarket handler is correct (it opens both terminal sessions pre-publish and races them); the defect is in the omnibase_infra runtime consume leg. Root cause: _assign_direct_terminal_partitions ignored the metadata future returned by AIOKafkaClient.set_topics and re-called set_topics([same_topic]) each loop iteration. aiokafka 0.13.0 set_topics only forces a metadata refresh when the topic set DIFFERS from the tracked set, so every iteration after the first took the no-op branch and never re-fetched. An ephemeral group_id=None consumer whose first metadata fetch had not yet surfaced partitions burned the full 30s assign cap and raised a bare TimeoutError (the empty-message 'wait failed' seen live). Fix: register the reply topic once and await that metadata fetch, then on each miss force a fresh fetch via force_metadata_update (which always fetches) rather than the no-op set_topics repeat. The assign-cap TimeoutError now carries a diagnostic message instead of an empty one. Test: tests/integration/test_terminal_consumer_concurrent_assign_race.py drives the REAL TerminalEventConsumer (the object wired as event_consumer) through the REAL open_direct_terminal_consumer/poll path with AIOKafkaConsumer monkeypatched to a fake that faithfully models aiokafka 0.13.0 set_topics future + metadata-latency semantics. RED before the fix (bare TimeoutError, the live empty-message signature); GREEN after. K>=2 multi-trial variant asserts no worker-thread leak across trials. Evidence-Source: <occ-sha-pending> Evidence-Ticket: OMN-13012 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13096): onex delegate single-command subcommand (Phase 2b) (#1967) * feat(OMN-13096): onex delegate single-command subcommand Add 'onex delegate "<prompt>" [--task-type X] [--max-tokens N]' as a subcommand on the existing onex CLI (Phase 2b of the skill-output-suppression slice, OMN-13089). The command wraps payload construction, node dispatch, and result extraction internally and prints exactly one ModelSkillResult[ModelDelegateSkillResponse] to stdout via the OMN-13094 receipt-mode path. RuntimeLocal logs go to the capture file + artifact store, never to stdout; scratch payloads live under <state-root>/tmp/ with run_id suffixes (never /tmp). - cli_delegate.py: classify_task_type (keyword table from legacy skill md), payload write, contract resolve, run_receipt_mode dispatch - register 'delegate' under onex.cli entry points - exempt delegate_command from the >5-param patterns gate (same Click-callback rationale as run_node_by_name) - 20 unit tests: classification, scratch-under-state-root, single typed receipt on stdout, zero INFO log leakage omnibase_infra does NOT depend on omnimarket; the delegate node is resolved at runtime via the onex.nodes entry-point group (registered by omnimarket). * chore(OMN-13096): re-trigger deploy-gate after Evidence-Source set to OCC#2593 No code change — the deploy-gate workflow triggers on synchronize (not edited), so the PR-body Evidence-Source fix needs a new commit to re-resolve the OCC ref to the open PR head where contracts/OMN-13096.yaml (with deploy evidence) lives. * chore(OMN-13096): re-trigger deploy-gate now that OCC#2593 merged to OCC dev contracts/OMN-13096.yaml (with the dod-deploy-onex-delegate item) is now on OCC dev, so the deploy-gate OCC-dev checkout resolves the contract + deploy evidence. * chore(OMN-13096): regenerate runner-image identity lock for pyproject entry-point add Adding the 'delegate' onex.cli entry point to pyproject.toml changed the dependency-manifest digest that scripts/ci/runner_image_identity.py folds into the runner-image identity lock. Regenerate the lock so tests/ci/test_runner_image_identity.py matches (was the only CI test failure; unrelated environmental integration/perf failures excluded). * test(OMN-13096): update backmerge identity assertions to regenerated lock digests The runner-image identity lock was regenerated for the pyproject entry-point add; this test hardcodes the expected identity_digest/shared_env_digest, so update both to match the new lock (same maintenance OMN-13094 did). --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): tolerate partition-less reply topic in terminal consume leg (#1970) The context-ROI runner opens one ephemeral group_id=None terminal consumer per terminal topic BEFORE publishing each generation command (subscribe-before-publish, OMN-13012/13038). The FAILED reply topic is only produced to on contract_passed=False; in a battery where generations pass it has zero messages, so Redpanda never advertises a partition for it. _assign_direct_terminal_partitions burned the full 30s assign cap on every trial then raised a bare TimeoutError (the empty-message 'wait failed'), stalling each of the 160 battery trials ~30s before the COMPLETED terminal could correlate -> battery needs >80 min and never completes (verifier-confirmed wedge). A partition-less reply topic is a valid steady state, not a 30s error: - _assign_direct_terminal_partitions gives a bounded grace window for a topic that exists but is slow to surface metadata, then assigns whatever partitions exist (possibly none) and returns promptly instead of burning the cap and raising. - poll_direct_terminal_consumer treats an empty assignment as 'no terminal will arrive here' (sleeps out its timeout, returns None) without calling getone() on an unassigned consumer. Repro: tests/integration/test_terminal_consumer_battery_load_wedge.py drives the REAL TerminalEventConsumer over K=10 x 2 cells x 2 arms; RED (40/40 trials block a full assign cap) before the fix, GREEN after. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): pin terminal-consumer read offset synchronously (close lazy seek_to_end publish race) (#1971) * fix(OMN-13118): pin terminal-consumer read offset synchronously to close lazy seek_to_end publish race The consume-leg wedge survived PR #1969 (set_topics no-op) and PR #1970 (partition-less assign-cap) because both addressed the assign phase, not the seek timing. AIOKafkaConsumer.seek_to_end is LAZY: it requests a LATEST offset reset that only resolves on the first poll — AFTER the caller publishes. With generation completing in ~1s, the correlated COMPLETED terminal lands in the open->poll gap, so the lazily-resolved LATEST position is the HWM AFTER the record and the poll reads past it. The terminal is never read, the trial never correlates, and the experiment matrix re-fires the same cell forever. Replace seek_to_end with a synchronous end_offsets() + seek() pin (_pin_direct_terminal_end_offsets) in both open_direct_terminal_consumer and RuntimePatternBBroker._dispatch_and_wait_with_direct_kafka_consumer, so the read position is fixed at open() time, before the publish — the real subscribe-before-publish guarantee. Empty assignment (partition-less reply topic, OMN-13118 #1970) is a no-op. Repro: tests/integration/test_terminal_consumer_seek_reset_race.py drives the real TerminalEventConsumer.open()/wait() the way HandlerContextRoiRunner does, publishing the correlated terminal in the open->wait gap across K=10 x 2 cells x 2 arms; RED with the lazy reset (every cell degenerate), GREEN once the read offset is pinned. RED verified by git-stashing only the source fix. Existing consume-leg fakes updated to model end_offsets/seek (they previously masked the bug by making seek_to_end a synchronous exact snapshot). * test(OMN-13118): reword assertion (lazy not deferred) for receipt honesty gate * fix(OMN-13118): bound end_offsets() round-trip with assign-cap timeout (CodeRabbit) end_offsets() is a broker ListOffsets round-trip aiokafka documents as able to block indefinitely. Bound it with the same cap as start()/assign so a stalled broker fails fast instead of hanging the pre-publish positioning. --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13137): _validate_routing strategy-aware event_model checks (#1973) operation_match routes by the `operation` field and does not use event_model. The validator was unconditionally requiring event_model.{name,module} for every handler entry, causing 230/295 omnimarket operation_match contracts (all correct as authored) to fail routing validation at startup. Fix: read routing_strategy from the routing map and branch validation: - payload_type_match → require event_model.{name, module} (unchanged) - operation_match (and any non-payload strategy) → require `operation`; skip event_model checks entirely Updated pre-existing _validate_routing tests to declare routing_strategy: payload_type_match explicitly (they always tested payload_type_match semantics but relied on the implicit fallback that is now removed). Added test_validate_routing_operation_match.py with 4 unit tests: 1. operation_match without event_model → zero event_model errors 2. operation_match missing operation field → error 3. payload_type_match missing event_model → still errors (regression guard) 4. Real node_integration_sweep_orchestrator routing block → clean (boot gate) Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): independent per-terminal-topic consumers in Pattern B direct-Kafka wait (#1972) The consume-leg wedge survived four merged fixes (#1969 set_topics no-op, #1970 partition-less assign-cap, #1971 synchronous seek-pin). The STRONG K>=10 multi-cell reprobe on the stability lane still wedged on REBUILD-5 (cdf53d963f7b). Converged diagnosis (strikes 3+4, docs/evidence/2026-06-12-weekend-pass/experiments/probe4-stability/ reprobe-K10-rebuild5/HALT_K10_WEDGE_PERSISTS.md): the runtime waited for each trial's terminal across TWO topics (node-generation-completed.v1 + node-generation-failed.v1) with a SINGLE ephemeral group_id=None consumer assigned both topics' partitions. One aiokafka consumer holds one manual subscription; the COMPLETED delivery window collapsed before it surfaced the correlated record, so the trial never correlated and the matrix re-fired cell 1. RuntimePatternBBroker._dispatch_and_wait_with_direct_kafka_consumer now opens ONE independent AIOKafkaConsumer PER terminal topic via open_direct_terminal_consumer (each started, assigned, and offset-pinned via end_offsets()+seek() at open() BEFORE publish), then awaits both CONCURRENTLY via asyncio.wait(FIRST_COMPLETED). The first correlated terminal wins; both are torn down. No shared consumer, no subscription flip. Keeps the #1970 partition-less no-op and the #1971 synchronous offset pin (both live in open_direct_terminal_consumer / poll_direct_terminal_consumer). Removes the now-dead single-consumer helpers (_assign_terminal_topic_partitions, method-level _refresh_terminal_topic_metadata, _direct_kafka_* kwargs builders, _kafka_bootstrap_servers/_kafka_event_bus). Adds tests/integration/test_terminal_consumer_subscription_flip_wedge.py: a real-dispatch-path K>=10 x 2-cell x 2-arm repro whose fake models TWO independent consumers honestly (delivery is faithful only for a single-topic assignment; a consumer spanning both topics flips and drops the COMPLETED record). RED genuineness verified by reverting only the source to the single-consumer shape (test hangs past timeout); GREEN with the fix in 2.3s. Acceptance is the LIVE K>=10 multi-cell stability-lane reprobe (later phase), NOT this unit test. A green unit repro is necessary but NOT sufficient. Refs OMN-13118, OMN-13128. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): long-lived terminal correlator replaces per-trial ephemeral consume leg (#1974) Tier B canonical redesign (epic OMN-12525). Five offset/subscription patches (#1…
…olution) (#2800) * chore(OMN-12934): continue dev to main promotion for prod redeploy (#1988) * docs(OMN-12962): contract-store durability audit — cold-runtime census proof (#1937) Proves cold-start contract census reconstructs from the image-bundled filesystem manifest (HYBRID-mode bootstrap + PluginLoaderContractSource), independent of node-registration.v1 retention. The delete-retention topic feeds only the post-freeze dynamic listener (auto_offset_reset=latest, no history replay). Live .201 stability-test evidence: filesystem contract_path in manifest + truncated topic log head with intact census. No store fix needed; residual dynamic-only gap covered by runtime_sweep sweep check. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12970): vendor omnimarket projection node migrations into forward/nodes (#1938) Vendors omnimarket node-owned projection migrations into the namespaced forward-migration tree so run-forward-migrations.sh materializes them in the dashboard projection DB (omnidash_analytics) at deploy. Primary (OMN-12970): creates llm_call_metrics, llm_cost_aggregates, and capability_scores in the projection DB. These were only ever created in the omnibase_infra DB by infra migrations 031/060, so the ab-compare, cost.token_usage, cost.summary, and capability-scores projection topics were DEGRADED at startup ('table not found') and their dashboard panels rendered empty. Also re-syncs three omnimarket node migrations the vendor tree had drifted from (node_projection_llm_routing, node_projection_overnight, node_projection_savings /077) — sync-node-migrations.sh --check requires the full vendored tree to match omnimarket source, and these were missing. Companion to omnimarket PR for the same ticket (source migrations + projection table-coverage ratchet test). Evidence-Ticket: OMN-12970 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12965): stamp runtime image identity (version + revision) in workspace builds (#1943) * fix(OMN-12965): stamp runtime image identity (version + revision) in workspace builds The main runtime image stamped org.opencontainers.image.version=0.1.0 with a blank org.opencontainers.image.revision after workspace rebuilds. A blank identity degrades every proof packet (runtime SHA + image digest are required citations in accepted evidence). Root causes (three build paths under-stamped identity): - onex up --build (cmd_up) passed only GIT_SHA; the runtime-stage OCI labels read VCS_REF (-> blank revision) and RUNTIME_VERSION (-> placeholder 0.1.0). - deploy-runtime.sh passed VCS_REF but not RUNTIME_VERSION/GIT_SHA. - Dockerfile silently allowed blank/placeholder identity in workspace mode. Fix: - cli._image_identity_build_args() stamps the full quad (GIT_SHA/VCS_REF/ RUNTIME_VERSION/BUILD_DATE) and fails fast on an unresolved git revision. - deploy-runtime.sh stamps RUNTIME_VERSION + GIT_SHA and verifies the version label is non-placeholder post-deploy. - Dockerfile.runtime fails workspace builds with blank VCS_REF or placeholder RUNTIME_VERSION=0.1.0 (release mode unaffected). Enforcement ratchet (same PR): - scripts/check_runtime_image_identity.py static check, wired as pre-commit hook + CI gate (ci.yml). - tests/unit/infra/test_runtime_image_identity_labels.py pins the cli helpers + Dockerfile guard; deploy-agent test extended for the quad. Proven locally via throwaway docker builds: workspace+args -> populated labels; workspace without args -> guard fails (exit 64); release without args -> 0.1.0 placeholder allowed (no regression). Evidence-Ticket: OMN-12965 * test(OMN-12965): integration build proof for runtime image identity labels Builds the real runtime-stage ARG/LABEL/guard block against busybox and asserts via docker inspect: workspace+args -> populated version/revision; workspace without args -> guard fails (exit 64); release -> placeholder allowed. Satisfies the integration-test hard gate and makes the throwaway proof permanent. Evidence-Ticket: OMN-12965 --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12987): workspace-build sibling lock-pin preflight + provenance (#1944) Recurrence ratchet for the 2026-06-11 stability bootstrap crash. The 11:20Z --no-cache rebuild vendored omnibase_infra 0.37.0-dev (~2c1d672f) + core 0.42.0 even though omnimarket dev's uv.lock pinned infra 0.38.1 @ e2dbdc95 + core 0.44.0 @ c97c2c9a. The stale sibling predated the OMN-12501 Protocol-quarantine guard, turning a latent contract defect into a fatal crash that crash-looped the main runtime. - check_sibling_lock_pins.py: host-side fail-fast preflight resolving expected sibling versions/SHAs from the consuming repo's (omnimarket) uv.lock and comparing them against each vendored tree. Mismatch aborts the build. - stage_workspace.sh: emit a .build-sha marker per staged sibling (rsync drops .git) so the preflight and provenance can identify the vendored commit. - deploy-runtime.sh: run the preflight after staging, before build; abort on mismatch. Write the comparison under sibling-repos/.sibling-lock-pins.json. - compute_workspace_provenance.py + Dockerfile.runtime: fold expected-vs-actual lock_pin_comparison into build-provenance.json for deploy verifiers. Recurrence-guard tests prove a stale infra 0.37.0 vs lock-pinned 0.38.1 fails the preflight and matched pins pass; deploy script wiring is asserted statically. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12988): pin stability runtime-worker to replicas 1 + census ratchet (#1942) The base docker-compose.infra.yml defaults runtime-worker to replicas 0 (${WORKER_REPLICAS:-0}). The stability-test lane's required state includes a running worker (4-container census: main, effects, worker, projection-api), but the override pinned it via an env-interpolation default (${STABILITY_TEST_WORKER_REPLICAS:-1}) — a silent-drop surface: a stray exported STABILITY_TEST_WORKER_REPLICAS=0 or removal of the :-1 fallback would scale the worker to 0 with zero signal on a plain compose up/recreate. Fix: pin docker-compose.stability-test.yml runtime-worker deploy.replicas to the literal 1 (no env indirection). Ratchet (recurrence guards, same PR): - scripts/deploy-agent/tests/unit/test_runtime_worker_census.py: assert runtime-worker stays in the deploy-agent RUNTIME-scope census so a missing worker (replicas 0 => absent from docker compose ps) is a deploy failure, not silence; assert the override pins a literal 1. - tests/integration/infra/test_stability_test_runtime_compose_render.py: assert the rendered stability worker resolves deploy.replicas == 1. - tests/unit/infra/test_stability_test_runtime_lane.py: update the existing pin assertion to the literal 1. Evidence-Ticket: OMN-12988 Config-drift family: OMN-12945 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12979): expire-bound topic completeness suppressions (#1940) Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12966): accept community-edition Infisical /api/status in provision readiness gate (#1939) The fresh-provision readiness gate in provision-infisical.py only accepted the enterprise {"status": "ok"} payload and rejected the community edition's {"message": "Ok"}, returning 1 before bootstrap could run. This blocked provisioning against the Infisical instance deployed on .201 (community edition). Route the gate through the existing _is_infisical_ready helper (single source of truth, already used by the already-provisioned path). Add TestMainFreshProvision- ReadinessGate covering community/enterprise/not-ready cases. Also adds docker/docker-compose.infisical-stability.yml: an ADDITIVE Infisical compose project for the stability-test lane (joins the existing network as external, reuses stability postgres/valkey, no lane mutation), since the lane overlays disable the in-lane Infisical service via *-disabled profile overrides. P1.2b-A: Infisical now reachable from the stability runtime/effect containers; a known secret (OMN_12966_PROBE) seeds and resolves end-to-end via the machine- identity universal-auth path, verified from inside the effects container. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12958): volume-config drift gate + runtime config provenance (#1941) * feat(OMN-12958): volume-config drift gate + runtime provenance Compute config provenance (path + sha256) for the runtime-rendered Bifrost delegation contract; the deployed volume copy survives rebuilds and silently diverges from packaged source (two competing authorities, OMN-12945). - runtime/config_provenance.py: ModelConfigProvenance + drift classification, sidecar JSON writer (read by sweep + proof packets) - runtime/health/health_config_provenance.py: drift -> degraded health - render entrypoint logs provenance line + writes sidecar on every boot - docs/runbooks/volume-config-drift-and-reseed.md: ledgered re-seed procedure - validation exemption for config_name (logical identifier, not entity ref) No live volume mutation: re-seed is an operator deploy step (deploy_pending). * test(OMN-12958): cover volume config drift reseed flow --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12957): wire runtime-profile validator + core registry parity guard (#1945) * feat(OMN-12957): wire runtime-profile validator + core registry parity guard - Remove the import-time RuntimeError drift raise in runtime_profile.py (a hard raise on core/infra version skew would crash the kernel at import); the parity invariant is enforced by test_profiles_match_core_registry instead. - Add tests: _PROFILES keys == omnibase_core REGISTERED_RUNTIME_PROFILES, and every CONSUMER_ATTACHED_RUNTIME_PROFILES profile loads as a real lane. - Wire omnibase_core.validation.validator_runtime_profiles as a pre-commit hook + validator-runtime-profiles.yml CI gate on infra contracts. - Freeze 19 pre-existing violators in validation/runtime_profiles_allowlist.yaml (discovered by repo-root walk; drain via OMN-12982). Blocks NEW orphans. Requires the omnibase_core pin to include OMN-12957's validator (new rules). Evidence-Ticket: OMN-12957 Evidence-Source: 5463fbaf819409d4fb7f491dd4f276f10d869eba * ci(OMN-12957): pass runtime profile allowlist to validator * test(OMN-12957): cover runtime profile registry parity * fix(OMN-12957): keep runtime profile allowlist under config * fix(OMN-12957): pin core runtime profile registry --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12973): widen migration-gate healthcheck start_period + ratchet (#1950) P2.8: classify the prod migration-gate UNHEALTHY-then-self-resolved incident. Classification: idle-one-shot-mis-modeled = NO. The migration-gate is a correctly-modeled long-running sentinel (entrypoint `while true; sleep 3600`) whose healthcheck continuously polls db_metadata.migrations_complete via check_migrations_complete.sh. The container flipped UNHEALTHY transiently because its healthcheck start_period (10s) was far shorter than the real cold-volume migration window (~116s: prod gate started 09:35:22, intelligence-migration finished 09:37:18). Past the 10s grace window the still-failing probe was reported UNHEALTHY until migrations completed, then self-resolved — no fault. Fix: raise migration-gate healthcheck start_period 10s -> 180s in both the authoritative catalog manifest (docker/catalog/services/migration-gate.yaml, flows into the generated compose) and the hand-maintained docker/docker-compose.infra.yml that deploy-runtime.sh applies to .201, so a still-applying gate stays in `health: starting` instead of flipping UNHEALTHY. Ratchet (enforcement, not detection): new ValidatorHealthcheckStartPeriod in omnibase_infra catalog (validator_healthcheck_start_period.py) asserts a 120s floor for migration-completion gates, wired into `onex validate runtime` (cmd_validate_runtime) AND backed by unit tests that gate every PR via pre-commit + CI. A migration-completion gate is identified by BOTH a healthcheck that polls migration completion AND a service_completed_successfully dependency, so ordinary app services (e.g. intelligence-api, an HTTP liveness probe with a legitimately short 40s start_period) are not swept into the floor. Prod probed read-only only; no prod mutation. Applies to live prod via the batched stability/prod rebuild (deploy_pending). Evidence-Ticket: OMN-12973 Evidence-Source: OCC#PENDING Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12971): deliver Vertex ADC bearer token to runtime effects container (secret-ref discipline) (#1949) * feat(OMN-12971): deliver Vertex ADC bearer token to runtime effects container (secret-ref discipline) Wire the Vertex $500-credit ADC path into the runtime, ADDITIVE next to the Gemini API-key path (provider-agnostic; neither provider removed or forced). runtime_policy.contract.yaml: add a secret-source mapping llm.vertex.access_token (source_type=file, /run/secrets/vertex_access_token) to the dev/stability-test/ judge profiles, alongside the existing llm.gemini.api_key env mapping. The ref name MUST match cloud-vertex-gemini.secret_ref in omnimarket bifrost_delegation.yaml. The resolved VALUE is a short-lived OAuth bearer token minted from ADC, refreshed by the operator; the token VALUE is never committed — only the ref name + in-container path. Add aiplatform.googleapis.com to the cloud host allowlist. docker-compose.infra.yml: bind the operator-supplied host token file read-only to /run/secrets/vertex_access_token on the main and effects runtimes (VERTEX_ACCESS_TOKEN_HOST_FILE, default /dev/null so lanes without Vertex still start; Gemini key path unaffected). Pass through BIFROST_VERTEX_GEMINI_ENDPOINT_URL (overlay supplies the complete Vertex OpenAI-compat URL) and GOOGLE_CLOUD_PROJECT/LOCATION (default empty). runtime-policy.env: regenerated from the contract via render_runtime_policy_env (test_runtime_policy_env_matches_contract_renderer proves contract<->env parity). test_runtime_policy_contract.py: update host-allowlist assertion for the additive Vertex host. Pre-existing platform-wide topic-parity-gate failures (25 unrelated topics) are identical with and without this change (proven by diff); that gate is pre-commit- only (not a CI merge gate) and this change adds zero new topic gaps, so that one hook is SKIP-ped. No deploy/receipt/merge gate is bypassed. * fix(OMN-12971): make Vertex runtime env contract-owned --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13008): automated .201 disk maintenance — worktree GC + docker GC + watermark alert ratchet (#1952) * feat(OMN-13008): automated .201 disk maintenance — worktree GC + docker GC + watermark alert ratchet Conservative, keep-list-driven disk maintenance to prevent the 2026-06-11 /data ~95% outage that killed all three lanes mid-demo: - scripts/worktree-gc.sh: drives the canonical omniclaude prune-worktrees.sh (merged+clean+pushed safety) on both Mac (merge-sweep tick) and .201 - scripts/disk-gc.sh + disk_gc_plan.py: conservative docker/builder/image GC. Pure, testable removal planner honoring a VERSIONED keep-list (deploy/disk-gc/keep-list.yaml): never reaps a kept repo, kept tag, in-use image, or anything younger than min_age_days; keeps N superseded generations. - scripts/disk-watermark-check.sh + disk_watermark_event.py: df watermark ratchet. >=85% emits a typed disk-watermark bus event (warning) that the sweep auto-ticket path turns into a Linear ticket; >=90% emits critical. Broker addr is fail-fast from KAFKA_BOOTSTRAP_SERVERS (no localhost default). - deploy/disk-gc/: systemd USER timer (onex-disk-gc.timer/.service, hourly) + install-disk-gc.sh. User units, NOT lane containers. - tests: 20 unit tests incl. GC plan-safety invariants + dry-run proof that default mode issues no destructive op (a wrong-delete GC is worse than none). Contract: contracts/OMN-13008.yaml * fix(OMN-13008): pass docker inventory to GC planner via stdin, not env (ARG_MAX) On a host with many docker images, passing the full image/ps inventory as env vars to disk_gc_plan.py exceeds ARG_MAX ('Argument list too long', exit 126), producing an empty plan. Write inventory to per-run scratch files (under the log dir, never /tmp; cleaned on exit) and hand it to the planner on stdin as a JSON envelope. Verified the failure live on .201; planner now reads stdin. * fix(OMN-13008): simplify GC plan stdin pipe (two processes, no nested subprocess) * fix(OMN-13008): keep-wins reconciliation — never remove an image id with any keep reason A single image id can surface in multiple 'docker image ls' rows (one per repo:tag). One tag could route the id to dangling-removal while another routes it to keep (e.g. tagged 'latest' or within-N-generations). Live .201 plan showed an id in BOTH remove_image_ids and kept_reasons. Reconcile at the end: keep wins — any id with a keep reason is dropped from the remove list; remove list deduped. Adds 2 regression tests. Verified live on .201. * fix(OMN-13008): timer uses OnCalendar=hourly for reliable re-arm (oneshot service) OnUnitActiveSec does not reliably re-elapse for a oneshot service once it goes inactive (observed NextElapseUSecMonotonic=infinity live on .201 after the first run). Switch to OnCalendar=hourly + RandomizedDelaySec so the timer re-arms every hour. Keep Persistent=true for missed-run catch-up. --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13005): materialize blocking event_consumer in runtime auto-wiring (runner consume-leg degenerate rows) (#1951) * fix(OMN-13005): materialize blocking event_consumer in auto-wiring (runner consume-leg degenerate rows) The runtime auto-wiring materialized event_publisher for handlers that declare it but had no equivalent for event_consumer. Request/response EFFECT handlers (HandlerContextRoiRunner) that publish a command then block on the correlated terminal event fell back to their no-op consumer default, returning None immediately -> every result row degenerate (failure_stage=generation, attempt_count=0) while generations succeeded ~1s later. Adds _make_sync_event_consumer (mirror of _make_sync_event_publisher), backed by service_terminal_event_consumer.make_terminal_event_consumer: a sync (topic, correlation_id, timeout) -> dict | None adapter that runs the proven direct-Kafka correlate-and-wait loop (from RuntimePatternBBroker) on an isolated event loop in a worker thread, so blocking does not deadlock the runtime dispatch loop that delivers the awaited terminal. TDD through the REAL dispatch path: test_event_consumer_injection drives a trial through _prepare_handler_wiring with a terminal arriving after a delay and asserts a non-degenerate row; verified RED with injection disabled. * fix(OMN-13012): two-phase (seek-now/wait-later) terminal event_consumer to close the subscribe-after-publish race (#1954) The OMN-13005 injected event_consumer is a single callable that does assign -> seek_to_end -> poll internally, all AFTER the handler has already published its command. Once OMN-13010 freed the dispatch loop and generation began completing in ~1s, the correlated terminal lands BEFORE the single-call consumer's post-publish seek_to_end positions, so seek_to_end skips PAST the already-emitted terminal and the runner times out on an offset beyond it (probe3, run_id=20260611T2140Z-probe3 -- both arms degenerate, zero rebalances). Splits positioning from waiting so the caller subscribes BEFORE it publishes: session = consumer.open(topic) # assign + seek_to_end NOW publisher(command_topic, payload) # publish AFTER positioning payload = session.wait(cid, timeout) # block from the captured position The returned TerminalEventConsumer is still directly callable with the legacy (topic, cid, timeout) -> dict | None single-call shape for any consumer that does not need subscribe-before-publish; the runner is the only consumer today. TerminalConsumerSession owns a dedicated event loop on a daemon worker thread for the whole open->wait->close lifecycle, preserving the OMN-13005 loop-isolation discipline so blocking never deadlocks the runtime dispatch loop. TDD (real dispatch path, RED-then-GREEN): extends the OMN-13005 injection test with a terminal emitted IMMEDIATELY after publish. The single-call (seek-after-publish) consumer MISSES it (degenerate row -- RED test asserts failure_stage=generation); the two-phase (open-before-publish) consumer CATCHES it (non-degenerate -- GREEN). The Kafka layer is faked at the two service seams against a shared in-memory log modeling seek-to-end semantics. OMN-13005 blocking-correlate behavior preserved. 269/269 auto_wiring unit tests pass; mypy --strict clean. Sibling to OMN-13010 / OMN-13005 / OMN-13003. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13005): route terminal consumer through Kafka boundary --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12990): pin worker replicas in ledgered config + fail-fast lane overrides (silent-drop ratchet) (#1948) * fix(OMN-12990): pin worker replicas in ledgered config + fail-fast lane overrides The base compose sets runtime-worker deploy replicas to ${WORKER_REPLICAS:-0} (soft-default ZERO). The stability lane's required state includes a running worker (GATE_ZERO_PROOF.md: 4 runtime containers), but the worker presence rode entirely on a compose soft-default — any plain compose up/recreate without the policy env silently scaled the worker to zero with no error and no signal. Fix (contract-native + fail-fast): - Add 'replicas' to ModelRuntimeProcessPolicy; pin replicas: 1 in every lane's worker block in runtime_policy.contract.yaml. - Renderer emits {PROFILE}_WORKER_REPLICAS into the ledgered runtime-policy.env for dev/stability-test/judge/prod. - stability + prod compose overrides reference ${..._WORKER_REPLICAS:?...} (fail-fast, NO silent :-1/:-0 default). A recreate that omits the policy env now aborts loudly instead of dropping the worker. prod previously had no override at all and inherited the dangerous :-0 default. Ratchet (recurrence guards): - tests asserting fail-fast override form (no soft default), contract-declared replica pin >= 1 per lane, and rendered {PROFILE}_WORKER_REPLICAS=1 in the ledgered env. - runbook deploy/verify procedure adds an expected-container census (worker must be present) via verify_container_manifest; a missing worker is a FAILURE, not silence. Note: SKIP=topic-parity-gate — that local-only advisory gate (absent from all .github/workflows, not a required CI check) fails on 25 pre-existing cross-repo topic gaps (build-loop/omniclaude/omniweb) identical on pristine base HEAD 8d7da1249; this change adds zero topics. All other hooks ran clean. Evidence-Ticket: OMN-12990 * test(OMN-12990): cover worker replica policy integration --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12909): add gateway bus forwarder P0A (#1946) * feat(OMN-12909): add gateway bus forwarder p0a * test(OMN-12909): add gateway forwarder integration coverage * fix(OMN-12909): satisfy gateway forwarder validators * test(OMN-12909): allow gateway forwarder bus protocol * fix(OMN-12909): sync gateway forwarder entry point * fix(OMN-12909): refresh runner image identity lock * test(OMN-12909): relax JSON normalizer mixed benchmark threshold * fix(OMN-12909): allow gateway handlers to boot unconfigured * fix(OMN-12909): declare gateway forwarder runtime profile * fix(OMN-12909): update runner identity backmerge expectation --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13011): LANE CENSUS RECONCILIATION ratchet — declared desired-state per lane, drift = auto-ticket (#1955) The class fix for the recurring lane-drift regression. Nothing reconciled the declared desired state of a runtime lane against what is actually running, so the same failure kept recurring with zero signal: volume config drift (OMN-12945), WORKER_REPLICAS silent zero (OMN-12988/12990), and on 2026-06-11 prod runtime containers plus the broker network were silently absent for hours during demo prep. Ships a per-lane DESIRED-STATE census: - (a) DECLARED in a versioned lane manifest (deploy/lane-census/lane-manifest.yaml): container set, network, replicas, image-tag pattern per lane (stability-test/prod/judge/dev), derived from the canonical compose lane files. A parity ratchet keeps the manifest locked in step with the compose files. - (b) RECONCILED on a schedule on .201 by SHARING the OMN-13008 systemd timer (a drop-in 4th ExecStart on onex-disk-gc.service — never a second timer) and on-demand via scripts/lane-census-check.sh / runtime_sweep. - (c) Drift = typed bus event (onex.evt.infra.lane-census-drift.v1) + Linear auto-ticket naming exactly what is missing/extra (container_absent, network_detached, replicas_zero, unexpected_container, oneshot_failed/stuck, image_tag_mismatch). Fail-fast, no warn-only mode (gates-block policy); exit 30 on drift; bus publish fail-fast on KAFKA_BOOTSTRAP_SERVERS (no localhost default). Red fixture reproduces 2026-06-11: prod runtime containers absent + broker network detached must produce the exact drift findings + a non-zero exit hours before a human noticed. Pure planner is fully unit-tested; shell driver dry-run-tested. Builds on the OMN-12988 deploy-agent RUNTIME census (deploy-time) as the complementary steady-state reconciler; closes the runtime-worker.yaml container_name: null census gap by sourcing names from the compose lane files. Evidence-Ticket: OMN-13011 Config-drift family: OMN-12945 Relates-to: OMN-13009, OMN-12988, OMN-13008 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13020): vendor missing node migrations — llm_routing 0000 + context_roi 001 (#1956) Vendors two omnimarket node-source migrations into the infra forward-migration tree via scripts/sync-node-migrations.sh (the canonical OMN-12559 mechanism): - node_projection_llm_routing/0000_create_llm_routing_decisions.sql (source: omnimarket #1168 / OMN-12942, merge ed6734f8) - node_projection_context_roi/001_create_context_roi_scores.sql (source: omnimarket #1178 / OMN-12955, merge 5010b1f4) Without the 0000 base table, node_projection_llm_routing/0001 (CREATE VIEW) hard-fails against NODE_POSTGRES_DB=omnidash_analytics — exactly the prod forward-migration exit-3 of 2026-06-11T09:35:52Z, and reproduced by construction in any clean clones@dev build. Files are byte-identical to the omnimarket dev blobs (sha256 f8a8b339… / c4126e65…) and to the untracked hot-patch copies on the .201 stability clone. Both migrations are self-contained, all-statements-IF-NOT-EXISTS, and 0000 sorts lexically before 0001 within the node's namespaced identity space. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13058): close TerminalConsumerSession on open() failure (worker thread + event loop leak) (#1957) TerminalConsumerSession.__init__ starts its dedicated worker loop thread immediately. TerminalEventConsumer.open did 'session = TerminalConsumerSession(...); return session.open()' with no cleanup: any failure inside session.open() (consumer start timeout, partition-assign timeout, broker auth error) propagated out of the raising expression, the session reference was lost, and the daemon worker thread plus its never-closed asyncio event loop leaked -- one pair per failed open. The motivating caller (HandlerContextRoiRunner) opens a session per trial, so a 160-560-trial battery against a degraded broker accumulates hundreds of leaked threads in the long-lived effects container. Fix: wrap session.open() in try/except BaseException -> session.close() (idempotent: stops the loop, joins the thread) -> re-raise. Covers both the two-phase .open(topic) path and the legacy single-call __call__ path. Found by the P3.3 doctrinal review of merged #1951 (b9712af9 / 36d98275). TDD: tests/unit/runtime/test_service_terminal_event_consumer_open_failure.py injects a real open failure through the production path (event bus without _bootstrap_servers) and asserts no alive terminal-consumer-* thread after the raise. Verified RED with the fix stashed (2 failed), GREEN with it (2 passed). Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13021): non-dev-base guard — fail feature-base PRs absent Stacked-Parent declaration (retro A-6) (#1958) Any PR whose base is neither dev nor main fails unless the body carries 'Stacked-Parent: #N'. Prevents the feedback_stacked_prs_orphan_from_dev class (#1185/#1954 auto-merged INTO parent feature branches and stranded off dev). base=main remains governed by main-target-guard. Epic OMN-13013 (process enforcement ratchets — June 12 retro). Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13014): hot-patch ledger rebuild preflight gate (retro B-1) (#1959) * feat(OMN-13014): hot-patch ledger rebuild preflight gate (retro B-1) Hot-patches on .201 (.prepatch sibling discipline) silently revert on any image rebuild/force-recreate — the 2026-06-11 20:58Z rebuild already erased a live /api/generate patch once. This adds the rebuild-path gate: - scripts/preflight_hotpatch_ledger.py: given a target container or lane + per-repo build refs, hard-fails when any hot-patch ledger row's source PR merge commit is not an ancestor of the build ref (git merge-base --is-ancestor), plus a .prepatch tripwire of the running container (unledgered .prepatch = hard fail; --post-rebuild = zero .prepatch expected). Sole bypass: HOTPATCH_PREFLIGHT_BYPASS carrying the Rule-10 '# skip-token-allowed: <user-approval-receipt-id>' form. - scripts/deploy-runtime.sh: guard_hotpatch_ledger wired into main() before build/preview (both dry-run and execute), lane derived from the compose project; skips loudly only when no ledger exists on the host. - tests/unit/scripts/test_preflight_hotpatch_ledger.py: 17 unit tests (ancestor gate, lane scoping, ledger loading, tripwire, bypass forms). - tests/ci/test_receipt_gate_install_guard.py: repair stale guard — core OMN-12565 replaced the OMN-9198 'uv pip uninstall first' install step with a cleared workspace venv (uv venv --clear); assert the new contract. Ledger backfilled from live census (5 .prepatch files / 4 source PRs, all MERGED to dev) at /data/omninode/hotpatch-ledger/ledger.yaml on .201. * fix(OMN-13014): scope missing-prepatch tripwire warning to the probed container --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12977): workspace-build sibling-pin ratchet — honor consuming lock, fail-fast on stale vendor (#1936) * fix(OMN-12977): workspace-build sibling-pin ratchet — honor consuming lock, fail-fast on stale vendor The workspace-mode image build (BUILD_SOURCE=workspace, used by the stability-test deploy procedure) vendored sibling/foundation packages from whatever the canonical OMNI_HOME clones happened to be checked out at, ignoring the consuming repo's uv.lock. On 2026-06-11 this shipped a 13-day-stale omnibase_infra 0.37.0-dev (pre-OMN-12501 Protocol-quarantine guard) + core 0.42.0 against an omnimarket dev lock pinning infra 0.38.1@e2dbdc95 / core 0.44.0@c97c2c9a, dropping the guard and crashing wire_from_manifest bootstrap fatally (stability lane down on demo day). Fix + recurrence ratchet (same PR): - scripts/runtime_build/check_sibling_lock_pins.py: parse the consuming repo's uv.lock for expected version+git-rev of each foundation/sibling package (scoped to the package's own source line so editable/registry pins are not cross-attributed a dependency's rev), resolve the actual clone version+HEAD, compare, and classify drift backward/forward/none. Fail-fast (exit 1) on any drift; --allow-drift records an explicit operator override in the artifact, never silent. - stage_workspace.sh: runs the preflight against the canonical clones before staging; aborts the build (exit 3) on unacknowledged drift and writes workspace/sibling-pin-comparison.json. - compute_workspace_provenance.py: folds the expected-vs-actual comparison into build-provenance.json so deploy verifiers can assert the build honored the lock; flags unacknowledged drift as a provenance error. - Dockerfile.runtime: COPY the comparison artifact (committed placeholder so the COPY always resolves; overwritten by stage_workspace.sh in workspace mode). - TDD: 19 unit tests covering lock parsing (git/registry/editable sources), drift classification, the exact 0.37.0-vs-0.38.1 stale case, check_pins exit codes, and the allow-drift override. Pre-existing mypy-strict bare-dict errors in compute_workspace_provenance.py fixed in the same pass. Evidence-Ticket: OMN-12977 Evidence-Source: pending-occ * ci(OMN-12977): retry runtime smoke compose port race * test(OMN-12977): align sibling-pin script tests with current API --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12989): workspace-mode image build must honor sibling lock pins (#1947) * fix(OMN-12989): workspace build must honor sibling lock pins (fail-fast ratchet) The 2026-06-11 stability bootstrap crash was caused by a workspace-mode --no-cache rebuild that vendored omnibase_infra 0.37.0-dev (13-day-stale worktree) even though omnimarket dev's uv.lock pins infra 0.38.1. The downgraded sibling predated the OMN-12501 Protocol-quarantine guard and turned a latent contract defect into a fatal crash. Fix + ratchet (same PR set): - scripts/runtime_build/resolve_workspace_pins.py: parse the consuming repo's uv.lock for sibling pins (version + git rev); classify each staged/installed sibling as exact/ahead/regression/unpinned; fail-fast (WorkspacePinError) on any regression BELOW the lock pin. Stdlib version-tuple fallback when packaging is absent so the ratchet never fails-open. - compute_workspace_provenance.py: enforce sibling pins + a host-infra self-check (installed omnibase_infra vs lock pin — the exact crash vector, since host infra is built from the context, not staged), and emit a pin_comparison block into build-provenance.json for deploy verifiers. - Dockerfile.runtime: COPY resolve_workspace_pins.py beside the provenance script so the in-image import resolves. - TDD: failing tests first (test_resolve_workspace_pins.py, 11 cases) + provenance integration tests; fixed a pre-existing stale RUNTIME_VERSION assertion to read pyproject dynamically. Evidence-Ticket: OMN-12989 * fix(OMN-12989): workspace build must honor sibling lock pins (fail-fast ratchet) The 2026-06-11 stability bootstrap crash was caused by a workspace-mode --no-cache rebuild that vendored omnibase_infra 0.37.0-dev (13-day-stale worktree) even though omnimarket dev's uv.lock pins infra 0.38.1. The downgraded sibling predated the OMN-12501 Protocol-quarantine guard and turned a latent contract defect into a fatal crash. Fix + ratchet (same PR set): - scripts/runtime_build/resolve_workspace_pins.py: parse the consuming repo's uv.lock for sibling pins (version + git rev); classify each staged/installed sibling as exact/ahead/regression/unpinned; fail-fast (WorkspacePinError) on any regression BELOW the lock pin. Stdlib version-tuple fallback when packaging is absent so the ratchet never fails-open. - compute_workspace_provenance.py: enforce sibling pins + a host-infra self-check (installed omnibase_infra vs lock pin — the exact crash vector, since host infra is built from the context, not staged), and emit a pin_comparison block into build-provenance.json for deploy verifiers. - Dockerfile.runtime: COPY resolve_workspace_pins.py beside the provenance script so the in-image import resolves. - TDD: failing tests first (test_resolve_workspace_pins.py, 11 cases) + provenance integration tests; fixed a pre-existing stale RUNTIME_VERSION assertion to read pyproject dynamically. Evidence-Ticket: OMN-12989 * test(OMN-12989): co-locate provenance pin helper in fixture --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13055): absent repos in REPOS list warn and exit 0 instead of failing (#1960) Missing repos (not cloned locally) now emit a WARN line and are tracked in a separate WARNED array. Only real fetch/ff failures cause exit 1. This makes pull-all.sh safe to use on machines with a partial clone set, while keeping the explicit-list override behavior intact. Adds three regression tests: absent-only exits 0, absent+present exits 0 with OK for the present repo, present-failed+absent exits 1. Evidence-Ticket: OMN-13055 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13070): config_prefetcher overlay wins on controlled lanes (#1961) On infisical_required=True lanes, fetched/overlay config now always wins over ambient env. Ambient env is retained only as a declared bootstrap fallback (with an explicit provenance INFO log line) when Infisical returns None. apply_to_environment also overwrites stale env on controlled lanes. Uncontrolled lane (infisical_required=False) behaviour is unchanged. Adds 5 regression tests: controlled-lane Infisical-wins, env-bootstrap- fallback, apply_to_environment overwrite, missing-from-both-is-error, and uncontrolled-lane-env-still-wins. Refactors _resolve_key to return a (outcome, value, error) tuple to satisfy the ≤5-param pattern gate. Source: docs/audits/2026-06-10-runtime-env-overlay-authority-audit.md Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13062): migration-gate vacuity fix — sentinel discipline, wait-for-postgres, skip-manifest, sync-check nonzero (retro A-10) (#1963) * fix(OMN-13062): migration-gate vacuity fix — sentinel discipline, wait-for-postgres, skip-manifest, sync-check nonzero Fixes three bugs identified in retro A-10 (recurrences OMN-12885, OMN-12934): (1) RUNNER SENTINEL DISCIPLINE run-forward-migrations.sh now clears migrations_complete=FALSE at the start of every run and sets it TRUE only as its FINAL act after all infra and node migrations succeed. Any mid-run failure leaves the gate UNHEALTHY. runner_completed_at is stamped at the same final step as durable evidence of a successful completion run. (2) SYNC-NODE-MIGRATIONS VACUOUS GATE sync-node-migrations.sh --check now exits 2 (not 0) when the omnimarket source tree is unresolvable. Silent exit-0 was hiding drift. The single opt-out is SYNC_NODE_MIGRATIONS_SKIP_UNRESOLVABLE=1 for environments that intentionally run without the source. (3) WAIT-FOR-POSTGRES GUARD run-forward-migrations.sh now waits up to PG_WAIT_RETRIES (default 30) x 2s for Postgres to accept connections before proceeding, guarding the first-boot initdb race. (4) SKIP-MANIFEST docker/migrations/skip-manifest.yaml introduced as the sole committed escape for intentionally-skipped migrations. The runner reads this at startup; listed migrations are recorded in schema_migrations with checksum "skip-manifest" without executing the SQL. (5) MIGRATION 085 Adds runner_completed_at TIMESTAMPTZ column to db_metadata so the runner's final stamp is durable in the schema (idempotent ADD COLUMN IF NOT EXISTS). Rollback included. 22 regression tests added covering all five fix surfaces. * fix(OMN-13062): stamp schema fingerprint for migration 085 Migration 085 (085_add_runner_completed_at_to_db_metadata.sql) was added in the initial commit but schema_fingerprint.sha256 was not regenerated. Running `python scripts/check_schema_fingerprint.py stamp` updates the artifact from the stale hash to match the 71 migration files. Evidence-Source: OCC#2563 Evidence-Ticket: OMN-13062 --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12864): Bifrost endpoints → committed overlay authority + fail-loud loader (#1964) * feat(OMN-12864, OMN-12814, OMN-12945): Bifrost endpoints → committed overlay authority + fail-loud loader OMN-12864 — Committed lane overlay - docker/lane-overlays/dev.bifrost.yaml: typed deployment bindings for all four BIFROST_LOCAL_*_ENDPOINT_URL values (coder :8000, reasoner :8001, embedding :8100, ds4-flash :8101). Previously only available as ephemeral shell exports on .201; now committed, auditable, diff-able, and CI-checked. - docker/lane-overlays/dev.bifrost.env: generated dotenv sidecar consumed by compose via env_file; never edited directly (yaml is authority). - docker/docker-compose.infra.yml: wire the env_file block at compose root so the four endpoints are injected into the interpolation context on a clean shell. Hardcode BIFROST_CONTRACT_PATH (remove :-/empty footgun — OMN-12814). - scripts/render_bifrost_lane_overlay_env.py: render script regenerates the env sidecar from the YAML source. - src/omnibase_infra/runtime/models/model_bifrost_lane_overlay.py: ModelBifrostLaneOverlay — typed Pydantic model enforcing URL completeness (OMN-12815: every URL must end in /chat/completions). OMN-12814 — Fail-loud loader - render_bifrost_delegation_contract: raises ProtocolConfigurationError on FileNotFoundError, YAMLError, ValidationError, and zero-endpoint renders. No lru_cache — every restart re-renders from packaged source so a stale cache cannot pin a broken result across deploys. OMN-12945 — Re-seed from packaged source on deploy - docker/entrypoint-runtime.sh: set BIFROST_FORCE_RESEED=1 on every container restart so the named-volume copy is always rebuilt from the packaged bifrost_delegation.yaml merged with committed lane-overlay endpoints. - render_bifrost_delegation_contract: honor BIFROST_FORCE_RESEED/force_reseed flag to bypass the stale-volume early-return path entirely. Tests: - tests/ci/test_bifrost_lane_overlay.py: CI gate — env sidecar in-sync with YAML source; all four BIFROST_LOCAL_* keys present. - tests/unit/runtime/models/test_model_bifrost_lane_overlay.py: bare-base URL rejection, env dict mapping, extra-field rejection. - tests/unit/runtime/test_render_bifrost_delegation_contract.py: fail-loud paths, force-reseed, zero-endpoint error, endpoint URL completeness. - tests/unit/models/test_model_serialization_roundtrip.py: roundtrip coverage. * fix(OMN-12864): move bifrost env_file to service level — fix compose schema validation failure Top-level 'env_file' is rejected by Docker Compose v2 schema validator ('additional properties not allowed'). This caused 10+ compose-render integration tests to fail in CI. Fix: - Remove top-level env_file block from docker-compose.infra.yml - Add per-service env_file on omninode-runtime, runtime-effects, runtime-worker (the three containers that render Bifrost) - Change BIFROST_LOCAL_*:? to BIFROST_LOCAL_*:- in x-runtime-env (compose-level validation removed; Python validates via ModelBifrostLaneOverlay + render_bifrost_delegation_contract) - Add two CI gate tests: compose_env_file_is_service_level_not_top_level and runtime_services_have_bifrost_env_file * fix(OMN-12864): allow BIFROST_LOCAL_* empty defaults in silent-fallback gate The overlay authority pattern (OMN-12864) passes BIFROST_LOCAL_*_ENDPOINT_URL via service-level env_file (docker/lane-overlays/dev.bifrost.env), not at compose config time. Compose-level :? would break CI rendering without the overlay pre-loaded. Validation at the Python layer (ModelBifrostLaneOverlay + render_bifrost_delegation_contract) is the enforcement point. Add the three failing vars to ALLOWED_EMPTY_DEFAULTS with OMN-12864 citation. --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12858/OMN-12879/OMN-12880): dispatcher route-coverage CI gate (#1965) * feat(OMN-12858/OMN-12879/OMN-12880): dispatcher route-coverage CI gate Static contract analysis: every subscribed command topic (onex.cmd.*) must declare handler_routing or runtime_dispatch, or the message goes to DLQ silently. This gate would have caught two recent incidents: 1. June 9 DLQ regression (OMN-12858 post-mortem): node_generation_consumer subscribed onex.cmd.omnimarket.node-generation-requested.v1 but a sole-handler revert left zero dispatcher routes registered. Messages went to DLQ silently with no CI signal. 2. June 12 DEL-01 live finding: onex.cmd.omnimarket.delegate-skill.v1 was consumed by dev lane bus but no dispatcher route existed in any deployed contract. Discovered via manual rpk consumer-group lag probe (DEL-01 evidence, docs/evidence/2026-06-12-weekend-pass/). Deliverables: - scripts/check_dispatcher_route_coverage.py — static YAML scanner that checks both omnibase_infra and omnimarket contract trees; ratchet allowlist for known pre-existing violations; --changed-contracts mode (OMN-12879) for per-PR scoping; compat publish topics excluded (OMN-12880) - .github/workflows/dispatcher-route-coverage.yml — CI workflow that checks out omnimarket sibling, collects changed contract paths in PR mode, and runs the gate; fires on PR, push-to-main, and merge_group - tests/ci/test_dispatcher_route_coverage_gate.py — 12 unit tests covering RED/GREEN/COMPAT/CHANGED-MODE/ALLOWLIST/MULTI-DIR paths plus live-contract regression proof against the actual omnibase_infra tree Allowlist additions: - onex.cmd.omnibase-infra.pattern-b-dispatch.v1 (RuntimePatternBBroker, imperative consumer, OMN-12525 migration target) - onex.cmd.platform.contract-resolve-requested.v1 (transitional HTTP bridge node_contract_resolver_bridge OMN-2756, metadata.transitional=true) [OMN-12858, OMN-12879, OMN-12880] * fix(OMN-12858): drop full uv sync from dispatcher-route-coverage workflow Gate script only needs pyyaml (stdlib + yaml). Using full setup-python-uv was causing 10+ minute timeout. Replace with direct pip install pyyaml and invoke python3 directly. Reduces job from 10m timeout to <1m. [OMN-12858] --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13026): wire transport-mock-lint pre-commit hook + CI gate (#1962) * feat(OMN-13026): wire transport-mock-lint pre-commit hook + CI gate (omnibase_infra) Activates the transport-mock-lint validator (from omnibase_core, OMN-13026) on omnibase_infra. Ratchet baseline: 218 existing violations across 63 files frozen in validation/transport_mock_baseline.yaml. New bare AsyncMock/ MagicMock on EventBus/transport surfaces are blocked by pre-commit hook and CI lint step. Existing violations tracked for drain by per-site tickets (parent OMN-13026). Reference incident: PR #1181 bare AsyncMock hid missing EventBusKafka.stop(). Evidence-Ticket: OMN-13026 * fix(OMN-13026): use uv run python for CI lint step + bump omnibase_core pin to include transport_mock_lint The transport-mock lint CI step previously cloned omnibase_core and ran `python -m omnibase_core.validators.transport_mock_lint` with PYTHONPATH, but this failed: `No module named omnibase_core.validators.transport_mock_lint` because it ran `.venv/bin/python` which uses the locked venv, and the venv omnibase_core pin (2defabef4) predates the transport_mock_lint module. Fix: use `uv run python` (removes the clone step) and bump omnibase-core git pin from 2defabef4 to 309d89fa7 (PR 1231 merge commit on dev) so transport_mock_lint is available in the locked venv. * fix(OMN-13026): align transport mock baseline and runner lock * fix(OMN-13026): sync omnibase_core pin + runner identity lock to dev baseline Align pyproject.toml omnibase_core rev to 2defabef (required by test_release_backmerge_preserves_proven_runtime_core_pin) and update docker/runners/runner-image.lock.json identity_digest/shared_env_digest to match dev runner image lock (79b08f44 / 90c8b3b9). Both were stale from the prior session's pin bump that used an older SHA. * fix(OMN-13026): source transport mock validator from core * fix(OMN-13026): source transport validator from core dev --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13094): --output receipt mode on onex node/run — quiet typed receipts with durable capture (#1966) Phase 2a of the skill-output-suppression slice (epic OMN-13089, plan docs/plans/2026-06-12-skill-output-suppression-plan.md Phase 2 item 1). - onex node/onex run gain --output receipt: ALL runtime logging routes to a run_id-suffixed capture file under <state-root>/captures/ (no console handlers — kills the 25-50-line RuntimeLocal INFO stream at the source); stdout carries exactly ONE typed ModelSkillResult JSON with the FULL handler result (result_model = concrete handler result type FQN). - Durable capture: capture log + handler result content-addressed via omnibase_core ArtifactStore (OMN-13093); artifact.captured + tool.output.captured emitted to the emit daemon socket (--emit-socket, default ~/.claude/emit.sock). - Failure asymmetry: artifact write failure => FULL output printed, no receipt (no hidden loss); emission failure => receipt still prints, event spooled to <state-root>/emit_spool/ for replay. - Node failure => status=failed/error with full error + capture log INLINE in the receipt (errors are never hidden) and artifact-backed. - Default output mode unchanged (enforcement is Phase 4). - RuntimeLocal exposes handler_result (receipt schema identity). - core pin 2defabef -> ae8793bd (merged OMN-13091/13093 receipt models + ArtifactStore); runner-image identity lock regenerated and the OMN-12765 backmerge identity constants updated for the new pin. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13097): onex skill subcommand + declarative skill->node mapping (#1968) * feat(OMN-13097): onex skill subcommand + declarative skill->node mapping Phase 4a of the skill-output-suppression slice (epic OMN-13089, plan docs/plans/2026-06-12-skill-output-suppression-plan.md Phase 4 item 1/2). A dispatch skill IS one CLI call (user directive 2026-06-12). This adds the `onex skill <name> [args]` dispatch surface that the 24 omniclaude shim migrations build on: - onex.cli entry-point `skill` -> cli_skill.run_skill_by_name. Resolves the skill via the declarative skill_mapping.yaml registry, builds the backing node's input payload from the skill's CLI args, writes it under the state root (.onex_state/tmp/<skill>-<run_id>.json — never /tmp), resolves the node's packaged contract exactly like `onex node`, and dispatches through the proven receipt-mode path (run_receipt_mode, OMN-13094). stdout is exactly one typed ModelSkillResult JSON with the FULL handler result. - skill_mapping.yaml: declarative DATA mapping all 24 dispatch shims to their backing onex.nodes node + typed result-model FQN (verified against each live handler handle() return type on origin/dev) + per-arg payload specs + static payload + keyword classifiers (delegate task_type as data, not code). Adding a skill is a YAML edit + fixture, never a CLI code change (ticket deliverable 2/3). Mapping lives beside the node-resolution surface, never hardcoded branching in the CLI. - Typed models split one-per-file (repo convention): ModelSkillArgSpec, ModelSkillClassifier, ModelSkillMapping, ModelSkillMappingRegistry, EnumSkillArgType. Frozen, extra=forbid, fail-fast coercion/validation. - validation_exemptions.yaml: Click-callback param-count + literal-identifier name-field exemptions mirroring the existing cli_node run_node_by_name precedent (OMN-11570) — same pattern, same rationale. dod_evidence: - 20 unit tests pass (registry validity, all-24-shims coverage, FQN result models, arg parsing/coercion/positional/required, classifiers, payload build, receipt-mode dispatch wiring, payload-under-state-root not /tmp). - uv run mypy src/ --strict: clean (2438 source files). - ruff format + check: clean. pre-commit run on changed files: pass. - runner-image identity lock regenerated for the pyproject entry-point add (same as OMN-13094). * test(OMN-13097): rebind OMN-12765 backmerge identity constants for onex skill pyproject change Adding the `skill` onex.cli entry-point to pyproject.toml changes the runner-image identity_digest (and shared_env_digest) the lock binds. Update the hardcoded expected constants in the backmerge-identity test to the regenerated values — same mechanical rebind OMN-13094 performed for the core pin bump. Identity + runner-image-identity tests pass (14). --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix+test(OMN-13012): force terminal-topic metadata refresh so both ephemeral consumers assign (runner consume-leg wedge) (#1969) The runner consume-leg wedged on the live stability battery (image c0505521f1fa, EXP1-3_RUNNER_CONSUME_LEG_BLOCKER): only the FAILED terminal topic ever subscribed while the COMPLETED topic never assigned, so the correlated completed terminal was never read and the 8x2x10 matrix re-fired cell 1 forever, emitting zero non-degenerate rows. A prior two-strike diagnosis proved the omnimarket handler is correct (it opens both terminal sessions pre-publish and races them); the defect is in the omnibase_infra runtime consume leg. Root cause: _assign_direct_terminal_partitions ignored the metadata future returned by AIOKafkaClient.set_topics and re-called set_topics([same_topic]) each loop iteration. aiokafka 0.13.0 set_topics only forces a metadata refresh when the topic set DIFFERS from the tracked set, so every iteration after the first took the no-op branch and never re-fetched. An ephemeral group_id=None consumer whose first metadata fetch had not yet surfaced partitions burned the full 30s assign cap and raised a bare TimeoutError (the empty-message 'wait failed' seen live). Fix: register the reply topic once and await that metadata fetch, then on each miss force a fresh fetch via force_metadata_update (which always fetches) rather than the no-op set_topics repeat. The assign-cap TimeoutError now carries a diagnostic message instead of an empty one. Test: tests/integration/test_terminal_consumer_concurrent_assign_race.py drives the REAL TerminalEventConsumer (the object wired as event_consumer) through the REAL open_direct_terminal_consumer/poll path with AIOKafkaConsumer monkeypatched to a fake that faithfully models aiokafka 0.13.0 set_topics future + metadata-latency semantics. RED before the fix (bare TimeoutError, the live empty-message signature); GREEN after. K>=2 multi-trial variant asserts no worker-thread leak across trials. Evidence-Source: <occ-sha-pending> Evidence-Ticket: OMN-13012 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13096): onex delegate single-command subcommand (Phase 2b) (#1967) * feat(OMN-13096): onex delegate single-command subcommand Add 'onex delegate "<prompt>" [--task-type X] [--max-tokens N]' as a subcommand on the existing onex CLI (Phase 2b of the skill-output-suppression slice, OMN-13089). The command wraps payload construction, node dispatch, and result extraction internally and prints exactly one ModelSkillResult[ModelDelegateSkillResponse] to stdout via the OMN-13094 receipt-mode path. RuntimeLocal logs go to the capture file + artifact store, never to stdout; scratch payloads live under <state-root>/tmp/ with run_id suffixes (never /tmp). - cli_delegate.py: classify_task_type (keyword table from legacy skill md), payload write, contract resolve, run_receipt_mode dispatch - register 'delegate' under onex.cli entry points - exempt delegate_command from the >5-param patterns gate (same Click-callback rationale as run_node_by_name) - 20 unit tests: classification, scratch-under-state-root, single typed receipt on stdout, zero INFO log leakage omnibase_infra does NOT depend on omnimarket; the delegate node is resolved at runtime via the onex.nodes entry-point group (registered by omnimarket). * chore(OMN-13096): re-trigger deploy-gate after Evidence-Source set to OCC#2593 No code change — the deploy-gate workflow triggers on synchronize (not edited), so the PR-body Evidence-Source fix needs a new commit to re-resolve the OCC ref to the open PR head where contracts/OMN-13096.yaml (with deploy evidence) lives. * chore(OMN-13096): re-trigger deploy-gate now that OCC#2593 merged to OCC dev contracts/OMN-13096.yaml (with the dod-deploy-onex-delegate item) is now on OCC dev, so the deploy-gate OCC-dev checkout resolves the contract + deploy evidence. * chore(OMN-13096): regenerate runner-image identity lock for pyproject entry-point add Adding the 'delegate' onex.cli entry point to pyproject.toml changed the dependency-manifest digest that scripts/ci/runner_image_identity.py folds into the runner-image identity lock. Regenerate the lock so tests/ci/test_runner_image_identity.py matches (was the only CI test failure; unrelated environmental integration/perf failures excluded). * test(OMN-13096): update backmerge identity assertions to regenerated lock digests The runner-image identity lock was regenerated for the pyproject entry-point add; this test hardcodes the expected identity_digest/shared_env_digest, so update both to match the new lock (same maintenance OMN-13094 did). --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): tolerate partition-less reply topic in terminal consume leg (#1970) The context-ROI runner opens one ephemeral group_id=None terminal consumer per terminal topic BEFORE publishing each generation command (subscribe-before-publish, OMN-13012/13038). The FAILED reply topic is only produced to on contract_passed=False; in a battery where generations pass it has zero messages, so Redpanda never advertises a partition for it. _assign_direct_terminal_partitions burned the full 30s assign cap on every trial then raised a bare TimeoutError (the empty-message 'wait failed'), stalling each of the 160 battery trials ~30s before the COMPLETED terminal could correlate -> battery needs >80 min and never completes (verifier-confirmed wedge). A partition-less reply topic is a valid steady state, not a 30s error: - _assign_direct_terminal_partitions gives a bounded grace window for a topic that exists but is slow to surface metadata, then assigns whatever partitions exist (possibly none) and returns promptly instead of burning the cap and raising. - poll_direct_terminal_consumer treats an empty assignment as 'no terminal will arrive here' (sleeps out its timeout, returns None) without calling getone() on an unassigned consumer. Repro: tests/integration/test_terminal_consumer_battery_load_wedge.py drives the REAL TerminalEventConsumer over K=10 x 2 cells x 2 arms; RED (40/40 trials block a full assign cap) before the fix, GREEN after. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): pin terminal-consumer read offset synchronously (close lazy seek_to_end publish race) (#1971) * fix(OMN-13118): pin terminal-consumer read offset synchronously to close lazy seek_to_end publish race The consume-leg wedge survived PR #1969 (set_topics no-op) and PR #1970 (partition-less assign-cap) because both addressed the assign phase, not the seek timing. AIOKafkaConsumer.seek_to_end is LAZY: it requests a LATEST offset reset that only resolves on the first poll — AFTER the caller publishes. With generation completing in ~1s, the correlated COMPLETED terminal lands in the open->poll gap, so the lazily-resolved LATEST position is the HWM AFTER the record and the poll reads past it. The terminal is never read, the trial never correlates, and the experiment matrix re-fires the same cell forever. Replace seek_to_end with a synchronous end_offsets() + seek() pin (_pin_direct_terminal_end_offsets) in both open_direct_terminal_consumer and RuntimePatternBBroker._dispatch_and_wait_with_direct_kafka_consumer, so the read position is fixed at open() time, before the publish — the real subscribe-before-publish guarantee. Empty assignment (partition-less reply topic, OMN-13118 #1970) is a no-op. Repro: tests/integration/test_terminal_consumer_seek_reset_race.py drives the real TerminalEventConsumer.open()/wait() the way HandlerContextRoiRunner does, publishing the correlated terminal in the open->wait gap across K=10 x 2 cells x 2 arms; RED with the lazy reset (every cell degenerate), GREEN once the read offset is pinned. RED verified by git-stashing only the source fix. Existing consume-leg fakes updated to model end_offsets/seek (they previously masked the bug by making seek_to_end a synchronous exact snapshot). * test(OMN-13118): reword assertion (lazy not deferred) for receipt honesty gate * fix(OMN-13118): bound end_offsets() round-trip with assign-cap timeout (CodeRabbit) end_offsets() is a broker ListOffsets round-trip aiokafka documents as able to block indefinitely. Bound it with the same cap as start()/assign so a stalled broker fails fast instead of hanging the pre-publish positioning. --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13137): _validate_routing strategy-aware event_model checks (#1973) operation_match routes by the `operation` field and does not use event_model. The validator was unconditionally requiring event_model.{name,module} for every handler entry, causing 230/295 omnimarket operation_match contracts (all correct as authored) to fail routing validation at startup. Fix: read routing_strategy from the routing map and branch validation: - payload_type_match → require event_model.{name, module} (unchanged) - operation_match (and any non-payload strategy) → require `operation`; skip event_model checks entirely Updated pre-existing _validate_routing tests to declare routing_strategy: payload_type_match explicitly (they always tested payload_type_match semantics but relied on the implicit fallback that is now removed). Added test_validate_routing_operation_match.py with 4 unit tests: 1. operation_match without event_model → zero event_model errors 2. operation_match missing operation field → error 3. payload_type_match missing event_model → still errors (regression guard) 4. Real node_integration_sweep_orchestrator routing block → clean (boot gate) Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): independent per-terminal-topic consumers in Pattern B direct-Kafka wait (#1972) The consume-leg wedge survived four merged fixes (#1969 set_topics no-op, #1970 partition-less assign-cap, #1971 synchronous seek-pin). The STRONG K>=10 multi-cell reprobe on the stability lane still wedged on REBUILD-5 (cdf53d963f7b). Converged diagnosis (strikes 3+4, docs/evidence/2026-06-12-weekend-pass/experiments/probe4-stability/ reprobe-K10-rebuild5/HALT_K10_WEDGE_PERSISTS.md): the runtime waited for each trial's terminal across TWO topics (node-generation-completed.v1 + node-generation-failed.v1) with a SINGLE ephemeral group_id=None consumer assigned both topics' partitions. One aiokafka consumer holds one manual subscription; the COMPLETED delivery window collapsed before it surfaced the correlated record, so the trial never correlated and the matrix re-fired cell 1. RuntimePatternBBroker._dispatch_and_wait_with_direct_kafka_consumer now opens ONE independent AIOKafkaConsumer PER terminal topic via open_direct_terminal_consumer (each started, assigned, and offset-pinned via end_offsets()+seek() at open() BEFORE publish), then awaits both CONCURRENTLY via asyncio.wait(FIRST_COMPLETED). The first correlated terminal wins; both are torn down. No shared consumer, no subscription flip. Keeps the #1970 partition-less no-op and the #1971 synchronous offset pin (both live in open_direct_terminal_consumer / poll_direct_terminal_consumer). Removes the now-dead single-consumer helpers (_assign_terminal_topic_partitions, method-level _refresh_terminal_topic_metadata, _direct_kafka_* kwargs builders, _kafka_bootstrap_servers/_kafka_event_bus). Adds tests/integration/test_terminal_consumer_subscription_flip_wedge.py: a real-dispatch-path K>=10 x 2-cell x 2-arm repro whose fake models TWO independent consumers honestly (delivery is faithful only for a single-topic assignment; a consumer spanning both topics flips and drops the COMPLETED record). RED genuineness verified by reverting only the source to the single-consumer shape (test hangs past timeout); GREEN with the fix in 2.3s. Acceptance is the LIVE K>=10 multi-cell stability-lane reprobe (later phase), NOT this unit test. A green unit repro is necessary but NOT sufficient. Refs OMN-13118, OMN-13128. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): long-lived terminal correlator replaces per-trial ephemeral consume leg (#1974) Tier B canonical redesign (epic OMN-12525). Five offset/subscription patches (#1969-…
…e-red classes behind it (#3041) OMN-16989 recorded 15 host-coupled `omnibase_infra` failures on `.201` and set `repos_denied=omnibase_infra` on `h201`. Every one of the 15 was measured inside the `.201` GATE-RUNNER CONTAINER via `scripts/ci/run_on_gate_runner.sh` -- a different execution environment from the one the table routes work to, which is the `.201` HOST over the OMN-16991 remote leg. Re-measured on the host over the real leg, receipts in `.onex_state/prepush_distribution/receipts.jsonl`: tests/unit/ + tests/integration/chains/ 24930 collected, pytest_exit=0, 695s tests/ci/ (the tree the 15 came from) 1892 collected, pytest_exit=0, 122s both host_mode=authorizing, chosen_host=omninode-pc. Zero host-coupled failures. The denial was pinning a verdict from an environment nothing dispatches to. Four real defects surfaced on the way and are fixed, not denied around: 1. The remote wrapper's PATH prefix was macOS-only by construction. `/opt/homebrew/bin` has no meaning on a Linux row and h201 is the fleet's only one. Its measured non-interactive PATH omits `~/.local/bin`, where BOTH `uv` and `shellcheck` live -- already covered by the existing entries, so this caused no red on h201, but the list is now Linux-complete. The added entries go AFTER every measured one, so they can only add resolution. 2. The transplanted tree has no remote-tracking refs. `git bundle create <b> HEAD` carries one ref, and this suite subprocesses the hook, which resolves `${PREPUSH_BASE_REF:-origin/dev}` first. On h201 the whole `test_prepush_hook_host_identity_guard.py` behavioral proof reduced to `base ref 'origin/dev' could not be resolved`. The dispatcher now hands the wrapper the base ref and sha and the wrapper `update-ref`s it in. BASE_SHA is a merge-base, so it is always an ancestor of HEAD and already in the bundle. Unresolvable => skipped silently: it may only add resolution, never refuse. 3. `check_deploy_scope_dod.resolve_omni_home` walks EVERY parent of the repo root, so on the transplant it reached `/data/omninode` and imported an omniclaude clone pinned at omniclaude#1600 (the host's current clone, #1963, is not an ancestor). Four tests failed on `has no attribute 'parse_evidence_metadata'` -- a red about another repo's clone freshness on whichever host ran the suite, with nothing saying so. Same cross-repo-hidden-input shape OMN-16989 named on the env-parity tests. `load_canonical_validator` now asserts the surface it calls and raises a NAMED staleness error; the fixture skips with that reason. The hook still hard-errors, but with a remediation instead of an AttributeError five frames down. 4. `test_release_identity_passes_when_version_ahead` asserted an AMBIENT repo-state precondition ("someone bumped the version in this PR") rather than stating it. Between a release tag landing and the post-release bump landing (OMN-13912, PR #3033, itself blocked on a stale runner-image identity lock) it is red on EVERY developer clone and ONLY there: `actions/checkout` does not fetch tags, so in CI `git tag --list` is empty, the gate is exempt and the test passes; locally the tags exist and it fails. Measured 2026-08-30 -- dev at 0.38.14 with `v0.38.14` cut at 13:24Z -- it refused every local governed pre-push in the repo while being invisible in CI. The test now names the precondition and skips with it. The GATE is untouched: `check-release-identity` (pre-commit) and the CI gate both still run the real script against the real diff. Table + tests, per the runbook's promotion procedure: - `scripts/hooks/prepush_hosts.tsv`: h201 `repos_denied` -> `-`, note carries the measured numbers. - `test_201_is_denied_for_omnibase_infra_until_omn16989_closes` -> `test_201_denies_no_repo_since_omn16989_closed`. - `test_a_repo_denied_host_is_never_chosen` now drives a SYNTHETIC table (the `_SYNTHETIC_TABLE` precedent set when h105 left shadow), so the placement RULE is no longer pinned to whichever repo the lab happens to deny today -- pinning it made a capacity-policy edit look like a mechanism regression, which is exactly the failure this run hit. - `test_no_row_denies_a_repo_today_so_the_rule_needs_a_synthetic_fixture` guards that fixture choice: the moment a real row denies a repo again, it fails and says the live table may be pinned again. Not fixed here: `pyproject.toml` version-vs-published is a release-cut condition owned by OMN-13912; this diff touches neither `src/**` nor `pyproject.toml`, so the release-identity gate itself does not apply to it.
Summary
Fixes three vacuity bugs in the migration subsystem identified in retro A-10 (recurrences OMN-12885, OMN-12934).
run-forward-migrations.shnow clearsmigrations_complete=FALSEat startup, sets itTRUE+ stampsrunner_completed_atonly as its FINAL act after ALL infra and node migrations succeed. Any mid-run failure leaves the gate UNHEALTHY.sync-node-migrations.sh --checknow exits 2 (not 0) when omnimarket source is unresolvable. Opt-out:SYNC_NODE_MIGRATIONS_SKIP_UNRESOLVABLE=1.PG_WAIT_RETRIESdefault 30 x 2s) guards first-boot initdb race.docker/migrations/skip-manifest.yamlintroduced as the sole committed escape for intentionally-skipped migrations.runner_completed_at TIMESTAMPTZtodb_metadata(idempotentADD COLUMN IF NOT EXISTS). Rollback included.22 regression tests covering all five fix surfaces. No live deploy required.
Evidence-Source: OCC#2563
Evidence-Ticket: OMN-13062