Repository navigation
chore(OMN-16041): backmerge main lineage into dev alongside the v0.38.5 promotion - #2745
Conversation
…1988) * docs(OMN-12962): contract-store durability audit — cold-runtime census proof (#1937) Proves cold-start contract census reconstructs from the image-bundled filesystem manifest (HYBRID-mode bootstrap + PluginLoaderContractSource), independent of node-registration.v1 retention. The delete-retention topic feeds only the post-freeze dynamic listener (auto_offset_reset=latest, no history replay). Live .201 stability-test evidence: filesystem contract_path in manifest + truncated topic log head with intact census. No store fix needed; residual dynamic-only gap covered by runtime_sweep sweep check. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12970): vendor omnimarket projection node migrations into forward/nodes (#1938) Vendors omnimarket node-owned projection migrations into the namespaced forward-migration tree so run-forward-migrations.sh materializes them in the dashboard projection DB (omnidash_analytics) at deploy. Primary (OMN-12970): creates llm_call_metrics, llm_cost_aggregates, and capability_scores in the projection DB. These were only ever created in the omnibase_infra DB by infra migrations 031/060, so the ab-compare, cost.token_usage, cost.summary, and capability-scores projection topics were DEGRADED at startup ('table not found') and their dashboard panels rendered empty. Also re-syncs three omnimarket node migrations the vendor tree had drifted from (node_projection_llm_routing, node_projection_overnight, node_projection_savings /077) — sync-node-migrations.sh --check requires the full vendored tree to match omnimarket source, and these were missing. Companion to omnimarket PR for the same ticket (source migrations + projection table-coverage ratchet test). Evidence-Ticket: OMN-12970 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12965): stamp runtime image identity (version + revision) in workspace builds (#1943) * fix(OMN-12965): stamp runtime image identity (version + revision) in workspace builds The main runtime image stamped org.opencontainers.image.version=0.1.0 with a blank org.opencontainers.image.revision after workspace rebuilds. A blank identity degrades every proof packet (runtime SHA + image digest are required citations in accepted evidence). Root causes (three build paths under-stamped identity): - onex up --build (cmd_up) passed only GIT_SHA; the runtime-stage OCI labels read VCS_REF (-> blank revision) and RUNTIME_VERSION (-> placeholder 0.1.0). - deploy-runtime.sh passed VCS_REF but not RUNTIME_VERSION/GIT_SHA. - Dockerfile silently allowed blank/placeholder identity in workspace mode. Fix: - cli._image_identity_build_args() stamps the full quad (GIT_SHA/VCS_REF/ RUNTIME_VERSION/BUILD_DATE) and fails fast on an unresolved git revision. - deploy-runtime.sh stamps RUNTIME_VERSION + GIT_SHA and verifies the version label is non-placeholder post-deploy. - Dockerfile.runtime fails workspace builds with blank VCS_REF or placeholder RUNTIME_VERSION=0.1.0 (release mode unaffected). Enforcement ratchet (same PR): - scripts/check_runtime_image_identity.py static check, wired as pre-commit hook + CI gate (ci.yml). - tests/unit/infra/test_runtime_image_identity_labels.py pins the cli helpers + Dockerfile guard; deploy-agent test extended for the quad. Proven locally via throwaway docker builds: workspace+args -> populated labels; workspace without args -> guard fails (exit 64); release without args -> 0.1.0 placeholder allowed (no regression). Evidence-Ticket: OMN-12965 * test(OMN-12965): integration build proof for runtime image identity labels Builds the real runtime-stage ARG/LABEL/guard block against busybox and asserts via docker inspect: workspace+args -> populated version/revision; workspace without args -> guard fails (exit 64); release -> placeholder allowed. Satisfies the integration-test hard gate and makes the throwaway proof permanent. Evidence-Ticket: OMN-12965 --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12987): workspace-build sibling lock-pin preflight + provenance (#1944) Recurrence ratchet for the 2026-06-11 stability bootstrap crash. The 11:20Z --no-cache rebuild vendored omnibase_infra 0.37.0-dev (~2c1d672f) + core 0.42.0 even though omnimarket dev's uv.lock pinned infra 0.38.1 @ e2dbdc95 + core 0.44.0 @ c97c2c9a. The stale sibling predated the OMN-12501 Protocol-quarantine guard, turning a latent contract defect into a fatal crash that crash-looped the main runtime. - check_sibling_lock_pins.py: host-side fail-fast preflight resolving expected sibling versions/SHAs from the consuming repo's (omnimarket) uv.lock and comparing them against each vendored tree. Mismatch aborts the build. - stage_workspace.sh: emit a .build-sha marker per staged sibling (rsync drops .git) so the preflight and provenance can identify the vendored commit. - deploy-runtime.sh: run the preflight after staging, before build; abort on mismatch. Write the comparison under sibling-repos/.sibling-lock-pins.json. - compute_workspace_provenance.py + Dockerfile.runtime: fold expected-vs-actual lock_pin_comparison into build-provenance.json for deploy verifiers. Recurrence-guard tests prove a stale infra 0.37.0 vs lock-pinned 0.38.1 fails the preflight and matched pins pass; deploy script wiring is asserted statically. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12988): pin stability runtime-worker to replicas 1 + census ratchet (#1942) The base docker-compose.infra.yml defaults runtime-worker to replicas 0 (${WORKER_REPLICAS:-0}). The stability-test lane's required state includes a running worker (4-container census: main, effects, worker, projection-api), but the override pinned it via an env-interpolation default (${STABILITY_TEST_WORKER_REPLICAS:-1}) — a silent-drop surface: a stray exported STABILITY_TEST_WORKER_REPLICAS=0 or removal of the :-1 fallback would scale the worker to 0 with zero signal on a plain compose up/recreate. Fix: pin docker-compose.stability-test.yml runtime-worker deploy.replicas to the literal 1 (no env indirection). Ratchet (recurrence guards, same PR): - scripts/deploy-agent/tests/unit/test_runtime_worker_census.py: assert runtime-worker stays in the deploy-agent RUNTIME-scope census so a missing worker (replicas 0 => absent from docker compose ps) is a deploy failure, not silence; assert the override pins a literal 1. - tests/integration/infra/test_stability_test_runtime_compose_render.py: assert the rendered stability worker resolves deploy.replicas == 1. - tests/unit/infra/test_stability_test_runtime_lane.py: update the existing pin assertion to the literal 1. Evidence-Ticket: OMN-12988 Config-drift family: OMN-12945 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12979): expire-bound topic completeness suppressions (#1940) Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12966): accept community-edition Infisical /api/status in provision readiness gate (#1939) The fresh-provision readiness gate in provision-infisical.py only accepted the enterprise {"status": "ok"} payload and rejected the community edition's {"message": "Ok"}, returning 1 before bootstrap could run. This blocked provisioning against the Infisical instance deployed on .201 (community edition). Route the gate through the existing _is_infisical_ready helper (single source of truth, already used by the already-provisioned path). Add TestMainFreshProvision- ReadinessGate covering community/enterprise/not-ready cases. Also adds docker/docker-compose.infisical-stability.yml: an ADDITIVE Infisical compose project for the stability-test lane (joins the existing network as external, reuses stability postgres/valkey, no lane mutation), since the lane overlays disable the in-lane Infisical service via *-disabled profile overrides. P1.2b-A: Infisical now reachable from the stability runtime/effect containers; a known secret (OMN_12966_PROBE) seeds and resolves end-to-end via the machine- identity universal-auth path, verified from inside the effects container. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12958): volume-config drift gate + runtime config provenance (#1941) * feat(OMN-12958): volume-config drift gate + runtime provenance Compute config provenance (path + sha256) for the runtime-rendered Bifrost delegation contract; the deployed volume copy survives rebuilds and silently diverges from packaged source (two competing authorities, OMN-12945). - runtime/config_provenance.py: ModelConfigProvenance + drift classification, sidecar JSON writer (read by sweep + proof packets) - runtime/health/health_config_provenance.py: drift -> degraded health - render entrypoint logs provenance line + writes sidecar on every boot - docs/runbooks/volume-config-drift-and-reseed.md: ledgered re-seed procedure - validation exemption for config_name (logical identifier, not entity ref) No live volume mutation: re-seed is an operator deploy step (deploy_pending). * test(OMN-12958): cover volume config drift reseed flow --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12957): wire runtime-profile validator + core registry parity guard (#1945) * feat(OMN-12957): wire runtime-profile validator + core registry parity guard - Remove the import-time RuntimeError drift raise in runtime_profile.py (a hard raise on core/infra version skew would crash the kernel at import); the parity invariant is enforced by test_profiles_match_core_registry instead. - Add tests: _PROFILES keys == omnibase_core REGISTERED_RUNTIME_PROFILES, and every CONSUMER_ATTACHED_RUNTIME_PROFILES profile loads as a real lane. - Wire omnibase_core.validation.validator_runtime_profiles as a pre-commit hook + validator-runtime-profiles.yml CI gate on infra contracts. - Freeze 19 pre-existing violators in validation/runtime_profiles_allowlist.yaml (discovered by repo-root walk; drain via OMN-12982). Blocks NEW orphans. Requires the omnibase_core pin to include OMN-12957's validator (new rules). Evidence-Ticket: OMN-12957 Evidence-Source: 5463fbaf819409d4fb7f491dd4f276f10d869eba * ci(OMN-12957): pass runtime profile allowlist to validator * test(OMN-12957): cover runtime profile registry parity * fix(OMN-12957): keep runtime profile allowlist under config * fix(OMN-12957): pin core runtime profile registry --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12973): widen migration-gate healthcheck start_period + ratchet (#1950) P2.8: classify the prod migration-gate UNHEALTHY-then-self-resolved incident. Classification: idle-one-shot-mis-modeled = NO. The migration-gate is a correctly-modeled long-running sentinel (entrypoint `while true; sleep 3600`) whose healthcheck continuously polls db_metadata.migrations_complete via check_migrations_complete.sh. The container flipped UNHEALTHY transiently because its healthcheck start_period (10s) was far shorter than the real cold-volume migration window (~116s: prod gate started 09:35:22, intelligence-migration finished 09:37:18). Past the 10s grace window the still-failing probe was reported UNHEALTHY until migrations completed, then self-resolved — no fault. Fix: raise migration-gate healthcheck start_period 10s -> 180s in both the authoritative catalog manifest (docker/catalog/services/migration-gate.yaml, flows into the generated compose) and the hand-maintained docker/docker-compose.infra.yml that deploy-runtime.sh applies to .201, so a still-applying gate stays in `health: starting` instead of flipping UNHEALTHY. Ratchet (enforcement, not detection): new ValidatorHealthcheckStartPeriod in omnibase_infra catalog (validator_healthcheck_start_period.py) asserts a 120s floor for migration-completion gates, wired into `onex validate runtime` (cmd_validate_runtime) AND backed by unit tests that gate every PR via pre-commit + CI. A migration-completion gate is identified by BOTH a healthcheck that polls migration completion AND a service_completed_successfully dependency, so ordinary app services (e.g. intelligence-api, an HTTP liveness probe with a legitimately short 40s start_period) are not swept into the floor. Prod probed read-only only; no prod mutation. Applies to live prod via the batched stability/prod rebuild (deploy_pending). Evidence-Ticket: OMN-12973 Evidence-Source: OCC#PENDING Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12971): deliver Vertex ADC bearer token to runtime effects container (secret-ref discipline) (#1949) * feat(OMN-12971): deliver Vertex ADC bearer token to runtime effects container (secret-ref discipline) Wire the Vertex $500-credit ADC path into the runtime, ADDITIVE next to the Gemini API-key path (provider-agnostic; neither provider removed or forced). runtime_policy.contract.yaml: add a secret-source mapping llm.vertex.access_token (source_type=file, /run/secrets/vertex_access_token) to the dev/stability-test/ judge profiles, alongside the existing llm.gemini.api_key env mapping. The ref name MUST match cloud-vertex-gemini.secret_ref in omnimarket bifrost_delegation.yaml. The resolved VALUE is a short-lived OAuth bearer token minted from ADC, refreshed by the operator; the token VALUE is never committed — only the ref name + in-container path. Add aiplatform.googleapis.com to the cloud host allowlist. docker-compose.infra.yml: bind the operator-supplied host token file read-only to /run/secrets/vertex_access_token on the main and effects runtimes (VERTEX_ACCESS_TOKEN_HOST_FILE, default /dev/null so lanes without Vertex still start; Gemini key path unaffected). Pass through BIFROST_VERTEX_GEMINI_ENDPOINT_URL (overlay supplies the complete Vertex OpenAI-compat URL) and GOOGLE_CLOUD_PROJECT/LOCATION (default empty). runtime-policy.env: regenerated from the contract via render_runtime_policy_env (test_runtime_policy_env_matches_contract_renderer proves contract<->env parity). test_runtime_policy_contract.py: update host-allowlist assertion for the additive Vertex host. Pre-existing platform-wide topic-parity-gate failures (25 unrelated topics) are identical with and without this change (proven by diff); that gate is pre-commit- only (not a CI merge gate) and this change adds zero new topic gaps, so that one hook is SKIP-ped. No deploy/receipt/merge gate is bypassed. * fix(OMN-12971): make Vertex runtime env contract-owned --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13008): automated .201 disk maintenance — worktree GC + docker GC + watermark alert ratchet (#1952) * feat(OMN-13008): automated .201 disk maintenance — worktree GC + docker GC + watermark alert ratchet Conservative, keep-list-driven disk maintenance to prevent the 2026-06-11 /data ~95% outage that killed all three lanes mid-demo: - scripts/worktree-gc.sh: drives the canonical omniclaude prune-worktrees.sh (merged+clean+pushed safety) on both Mac (merge-sweep tick) and .201 - scripts/disk-gc.sh + disk_gc_plan.py: conservative docker/builder/image GC. Pure, testable removal planner honoring a VERSIONED keep-list (deploy/disk-gc/keep-list.yaml): never reaps a kept repo, kept tag, in-use image, or anything younger than min_age_days; keeps N superseded generations. - scripts/disk-watermark-check.sh + disk_watermark_event.py: df watermark ratchet. >=85% emits a typed disk-watermark bus event (warning) that the sweep auto-ticket path turns into a Linear ticket; >=90% emits critical. Broker addr is fail-fast from KAFKA_BOOTSTRAP_SERVERS (no localhost default). - deploy/disk-gc/: systemd USER timer (onex-disk-gc.timer/.service, hourly) + install-disk-gc.sh. User units, NOT lane containers. - tests: 20 unit tests incl. GC plan-safety invariants + dry-run proof that default mode issues no destructive op (a wrong-delete GC is worse than none). Contract: contracts/OMN-13008.yaml * fix(OMN-13008): pass docker inventory to GC planner via stdin, not env (ARG_MAX) On a host with many docker images, passing the full image/ps inventory as env vars to disk_gc_plan.py exceeds ARG_MAX ('Argument list too long', exit 126), producing an empty plan. Write inventory to per-run scratch files (under the log dir, never /tmp; cleaned on exit) and hand it to the planner on stdin as a JSON envelope. Verified the failure live on .201; planner now reads stdin. * fix(OMN-13008): simplify GC plan stdin pipe (two processes, no nested subprocess) * fix(OMN-13008): keep-wins reconciliation — never remove an image id with any keep reason A single image id can surface in multiple 'docker image ls' rows (one per repo:tag). One tag could route the id to dangling-removal while another routes it to keep (e.g. tagged 'latest' or within-N-generations). Live .201 plan showed an id in BOTH remove_image_ids and kept_reasons. Reconcile at the end: keep wins — any id with a keep reason is dropped from the remove list; remove list deduped. Adds 2 regression tests. Verified live on .201. * fix(OMN-13008): timer uses OnCalendar=hourly for reliable re-arm (oneshot service) OnUnitActiveSec does not reliably re-elapse for a oneshot service once it goes inactive (observed NextElapseUSecMonotonic=infinity live on .201 after the first run). Switch to OnCalendar=hourly + RandomizedDelaySec so the timer re-arms every hour. Keep Persistent=true for missed-run catch-up. --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13005): materialize blocking event_consumer in runtime auto-wiring (runner consume-leg degenerate rows) (#1951) * fix(OMN-13005): materialize blocking event_consumer in auto-wiring (runner consume-leg degenerate rows) The runtime auto-wiring materialized event_publisher for handlers that declare it but had no equivalent for event_consumer. Request/response EFFECT handlers (HandlerContextRoiRunner) that publish a command then block on the correlated terminal event fell back to their no-op consumer default, returning None immediately -> every result row degenerate (failure_stage=generation, attempt_count=0) while generations succeeded ~1s later. Adds _make_sync_event_consumer (mirror of _make_sync_event_publisher), backed by service_terminal_event_consumer.make_terminal_event_consumer: a sync (topic, correlation_id, timeout) -> dict | None adapter that runs the proven direct-Kafka correlate-and-wait loop (from RuntimePatternBBroker) on an isolated event loop in a worker thread, so blocking does not deadlock the runtime dispatch loop that delivers the awaited terminal. TDD through the REAL dispatch path: test_event_consumer_injection drives a trial through _prepare_handler_wiring with a terminal arriving after a delay and asserts a non-degenerate row; verified RED with injection disabled. * fix(OMN-13012): two-phase (seek-now/wait-later) terminal event_consumer to close the subscribe-after-publish race (#1954) The OMN-13005 injected event_consumer is a single callable that does assign -> seek_to_end -> poll internally, all AFTER the handler has already published its command. Once OMN-13010 freed the dispatch loop and generation began completing in ~1s, the correlated terminal lands BEFORE the single-call consumer's post-publish seek_to_end positions, so seek_to_end skips PAST the already-emitted terminal and the runner times out on an offset beyond it (probe3, run_id=20260611T2140Z-probe3 -- both arms degenerate, zero rebalances). Splits positioning from waiting so the caller subscribes BEFORE it publishes: session = consumer.open(topic) # assign + seek_to_end NOW publisher(command_topic, payload) # publish AFTER positioning payload = session.wait(cid, timeout) # block from the captured position The returned TerminalEventConsumer is still directly callable with the legacy (topic, cid, timeout) -> dict | None single-call shape for any consumer that does not need subscribe-before-publish; the runner is the only consumer today. TerminalConsumerSession owns a dedicated event loop on a daemon worker thread for the whole open->wait->close lifecycle, preserving the OMN-13005 loop-isolation discipline so blocking never deadlocks the runtime dispatch loop. TDD (real dispatch path, RED-then-GREEN): extends the OMN-13005 injection test with a terminal emitted IMMEDIATELY after publish. The single-call (seek-after-publish) consumer MISSES it (degenerate row -- RED test asserts failure_stage=generation); the two-phase (open-before-publish) consumer CATCHES it (non-degenerate -- GREEN). The Kafka layer is faked at the two service seams against a shared in-memory log modeling seek-to-end semantics. OMN-13005 blocking-correlate behavior preserved. 269/269 auto_wiring unit tests pass; mypy --strict clean. Sibling to OMN-13010 / OMN-13005 / OMN-13003. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13005): route terminal consumer through Kafka boundary --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12990): pin worker replicas in ledgered config + fail-fast lane overrides (silent-drop ratchet) (#1948) * fix(OMN-12990): pin worker replicas in ledgered config + fail-fast lane overrides The base compose sets runtime-worker deploy replicas to ${WORKER_REPLICAS:-0} (soft-default ZERO). The stability lane's required state includes a running worker (GATE_ZERO_PROOF.md: 4 runtime containers), but the worker presence rode entirely on a compose soft-default — any plain compose up/recreate without the policy env silently scaled the worker to zero with no error and no signal. Fix (contract-native + fail-fast): - Add 'replicas' to ModelRuntimeProcessPolicy; pin replicas: 1 in every lane's worker block in runtime_policy.contract.yaml. - Renderer emits {PROFILE}_WORKER_REPLICAS into the ledgered runtime-policy.env for dev/stability-test/judge/prod. - stability + prod compose overrides reference ${..._WORKER_REPLICAS:?...} (fail-fast, NO silent :-1/:-0 default). A recreate that omits the policy env now aborts loudly instead of dropping the worker. prod previously had no override at all and inherited the dangerous :-0 default. Ratchet (recurrence guards): - tests asserting fail-fast override form (no soft default), contract-declared replica pin >= 1 per lane, and rendered {PROFILE}_WORKER_REPLICAS=1 in the ledgered env. - runbook deploy/verify procedure adds an expected-container census (worker must be present) via verify_container_manifest; a missing worker is a FAILURE, not silence. Note: SKIP=topic-parity-gate — that local-only advisory gate (absent from all .github/workflows, not a required CI check) fails on 25 pre-existing cross-repo topic gaps (build-loop/omniclaude/omniweb) identical on pristine base HEAD 8d7da1249; this change adds zero topics. All other hooks ran clean. Evidence-Ticket: OMN-12990 * test(OMN-12990): cover worker replica policy integration --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12909): add gateway bus forwarder P0A (#1946) * feat(OMN-12909): add gateway bus forwarder p0a * test(OMN-12909): add gateway forwarder integration coverage * fix(OMN-12909): satisfy gateway forwarder validators * test(OMN-12909): allow gateway forwarder bus protocol * fix(OMN-12909): sync gateway forwarder entry point * fix(OMN-12909): refresh runner image identity lock * test(OMN-12909): relax JSON normalizer mixed benchmark threshold * fix(OMN-12909): allow gateway handlers to boot unconfigured * fix(OMN-12909): declare gateway forwarder runtime profile * fix(OMN-12909): update runner identity backmerge expectation --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13011): LANE CENSUS RECONCILIATION ratchet — declared desired-state per lane, drift = auto-ticket (#1955) The class fix for the recurring lane-drift regression. Nothing reconciled the declared desired state of a runtime lane against what is actually running, so the same failure kept recurring with zero signal: volume config drift (OMN-12945), WORKER_REPLICAS silent zero (OMN-12988/12990), and on 2026-06-11 prod runtime containers plus the broker network were silently absent for hours during demo prep. Ships a per-lane DESIRED-STATE census: - (a) DECLARED in a versioned lane manifest (deploy/lane-census/lane-manifest.yaml): container set, network, replicas, image-tag pattern per lane (stability-test/prod/judge/dev), derived from the canonical compose lane files. A parity ratchet keeps the manifest locked in step with the compose files. - (b) RECONCILED on a schedule on .201 by SHARING the OMN-13008 systemd timer (a drop-in 4th ExecStart on onex-disk-gc.service — never a second timer) and on-demand via scripts/lane-census-check.sh / runtime_sweep. - (c) Drift = typed bus event (onex.evt.infra.lane-census-drift.v1) + Linear auto-ticket naming exactly what is missing/extra (container_absent, network_detached, replicas_zero, unexpected_container, oneshot_failed/stuck, image_tag_mismatch). Fail-fast, no warn-only mode (gates-block policy); exit 30 on drift; bus publish fail-fast on KAFKA_BOOTSTRAP_SERVERS (no localhost default). Red fixture reproduces 2026-06-11: prod runtime containers absent + broker network detached must produce the exact drift findings + a non-zero exit hours before a human noticed. Pure planner is fully unit-tested; shell driver dry-run-tested. Builds on the OMN-12988 deploy-agent RUNTIME census (deploy-time) as the complementary steady-state reconciler; closes the runtime-worker.yaml container_name: null census gap by sourcing names from the compose lane files. Evidence-Ticket: OMN-13011 Config-drift family: OMN-12945 Relates-to: OMN-13009, OMN-12988, OMN-13008 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13020): vendor missing node migrations — llm_routing 0000 + context_roi 001 (#1956) Vendors two omnimarket node-source migrations into the infra forward-migration tree via scripts/sync-node-migrations.sh (the canonical OMN-12559 mechanism): - node_projection_llm_routing/0000_create_llm_routing_decisions.sql (source: omnimarket #1168 / OMN-12942, merge ed6734f8) - node_projection_context_roi/001_create_context_roi_scores.sql (source: omnimarket #1178 / OMN-12955, merge 5010b1f4) Without the 0000 base table, node_projection_llm_routing/0001 (CREATE VIEW) hard-fails against NODE_POSTGRES_DB=omnidash_analytics — exactly the prod forward-migration exit-3 of 2026-06-11T09:35:52Z, and reproduced by construction in any clean clones@dev build. Files are byte-identical to the omnimarket dev blobs (sha256 f8a8b339… / c4126e65…) and to the untracked hot-patch copies on the .201 stability clone. Both migrations are self-contained, all-statements-IF-NOT-EXISTS, and 0000 sorts lexically before 0001 within the node's namespaced identity space. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13058): close TerminalConsumerSession on open() failure (worker thread + event loop leak) (#1957) TerminalConsumerSession.__init__ starts its dedicated worker loop thread immediately. TerminalEventConsumer.open did 'session = TerminalConsumerSession(...); return session.open()' with no cleanup: any failure inside session.open() (consumer start timeout, partition-assign timeout, broker auth error) propagated out of the raising expression, the session reference was lost, and the daemon worker thread plus its never-closed asyncio event loop leaked -- one pair per failed open. The motivating caller (HandlerContextRoiRunner) opens a session per trial, so a 160-560-trial battery against a degraded broker accumulates hundreds of leaked threads in the long-lived effects container. Fix: wrap session.open() in try/except BaseException -> session.close() (idempotent: stops the loop, joins the thread) -> re-raise. Covers both the two-phase .open(topic) path and the legacy single-call __call__ path. Found by the P3.3 doctrinal review of merged #1951 (b9712af9 / 36d98275). TDD: tests/unit/runtime/test_service_terminal_event_consumer_open_failure.py injects a real open failure through the production path (event bus without _bootstrap_servers) and asserts no alive terminal-consumer-* thread after the raise. Verified RED with the fix stashed (2 failed), GREEN with it (2 passed). Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13021): non-dev-base guard — fail feature-base PRs absent Stacked-Parent declaration (retro A-6) (#1958) Any PR whose base is neither dev nor main fails unless the body carries 'Stacked-Parent: #N'. Prevents the feedback_stacked_prs_orphan_from_dev class (#1185/#1954 auto-merged INTO parent feature branches and stranded off dev). base=main remains governed by main-target-guard. Epic OMN-13013 (process enforcement ratchets — June 12 retro). Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13014): hot-patch ledger rebuild preflight gate (retro B-1) (#1959) * feat(OMN-13014): hot-patch ledger rebuild preflight gate (retro B-1) Hot-patches on .201 (.prepatch sibling discipline) silently revert on any image rebuild/force-recreate — the 2026-06-11 20:58Z rebuild already erased a live /api/generate patch once. This adds the rebuild-path gate: - scripts/preflight_hotpatch_ledger.py: given a target container or lane + per-repo build refs, hard-fails when any hot-patch ledger row's source PR merge commit is not an ancestor of the build ref (git merge-base --is-ancestor), plus a .prepatch tripwire of the running container (unledgered .prepatch = hard fail; --post-rebuild = zero .prepatch expected). Sole bypass: HOTPATCH_PREFLIGHT_BYPASS carrying the Rule-10 '# skip-token-allowed: <user-approval-receipt-id>' form. - scripts/deploy-runtime.sh: guard_hotpatch_ledger wired into main() before build/preview (both dry-run and execute), lane derived from the compose project; skips loudly only when no ledger exists on the host. - tests/unit/scripts/test_preflight_hotpatch_ledger.py: 17 unit tests (ancestor gate, lane scoping, ledger loading, tripwire, bypass forms). - tests/ci/test_receipt_gate_install_guard.py: repair stale guard — core OMN-12565 replaced the OMN-9198 'uv pip uninstall first' install step with a cleared workspace venv (uv venv --clear); assert the new contract. Ledger backfilled from live census (5 .prepatch files / 4 source PRs, all MERGED to dev) at /data/omninode/hotpatch-ledger/ledger.yaml on .201. * fix(OMN-13014): scope missing-prepatch tripwire warning to the probed container --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12977): workspace-build sibling-pin ratchet — honor consuming lock, fail-fast on stale vendor (#1936) * fix(OMN-12977): workspace-build sibling-pin ratchet — honor consuming lock, fail-fast on stale vendor The workspace-mode image build (BUILD_SOURCE=workspace, used by the stability-test deploy procedure) vendored sibling/foundation packages from whatever the canonical OMNI_HOME clones happened to be checked out at, ignoring the consuming repo's uv.lock. On 2026-06-11 this shipped a 13-day-stale omnibase_infra 0.37.0-dev (pre-OMN-12501 Protocol-quarantine guard) + core 0.42.0 against an omnimarket dev lock pinning infra 0.38.1@e2dbdc95 / core 0.44.0@c97c2c9a, dropping the guard and crashing wire_from_manifest bootstrap fatally (stability lane down on demo day). Fix + recurrence ratchet (same PR): - scripts/runtime_build/check_sibling_lock_pins.py: parse the consuming repo's uv.lock for expected version+git-rev of each foundation/sibling package (scoped to the package's own source line so editable/registry pins are not cross-attributed a dependency's rev), resolve the actual clone version+HEAD, compare, and classify drift backward/forward/none. Fail-fast (exit 1) on any drift; --allow-drift records an explicit operator override in the artifact, never silent. - stage_workspace.sh: runs the preflight against the canonical clones before staging; aborts the build (exit 3) on unacknowledged drift and writes workspace/sibling-pin-comparison.json. - compute_workspace_provenance.py: folds the expected-vs-actual comparison into build-provenance.json so deploy verifiers can assert the build honored the lock; flags unacknowledged drift as a provenance error. - Dockerfile.runtime: COPY the comparison artifact (committed placeholder so the COPY always resolves; overwritten by stage_workspace.sh in workspace mode). - TDD: 19 unit tests covering lock parsing (git/registry/editable sources), drift classification, the exact 0.37.0-vs-0.38.1 stale case, check_pins exit codes, and the allow-drift override. Pre-existing mypy-strict bare-dict errors in compute_workspace_provenance.py fixed in the same pass. Evidence-Ticket: OMN-12977 Evidence-Source: pending-occ * ci(OMN-12977): retry runtime smoke compose port race * test(OMN-12977): align sibling-pin script tests with current API --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12989): workspace-mode image build must honor sibling lock pins (#1947) * fix(OMN-12989): workspace build must honor sibling lock pins (fail-fast ratchet) The 2026-06-11 stability bootstrap crash was caused by a workspace-mode --no-cache rebuild that vendored omnibase_infra 0.37.0-dev (13-day-stale worktree) even though omnimarket dev's uv.lock pins infra 0.38.1. The downgraded sibling predated the OMN-12501 Protocol-quarantine guard and turned a latent contract defect into a fatal crash. Fix + ratchet (same PR set): - scripts/runtime_build/resolve_workspace_pins.py: parse the consuming repo's uv.lock for sibling pins (version + git rev); classify each staged/installed sibling as exact/ahead/regression/unpinned; fail-fast (WorkspacePinError) on any regression BELOW the lock pin. Stdlib version-tuple fallback when packaging is absent so the ratchet never fails-open. - compute_workspace_provenance.py: enforce sibling pins + a host-infra self-check (installed omnibase_infra vs lock pin — the exact crash vector, since host infra is built from the context, not staged), and emit a pin_comparison block into build-provenance.json for deploy verifiers. - Dockerfile.runtime: COPY resolve_workspace_pins.py beside the provenance script so the in-image import resolves. - TDD: failing tests first (test_resolve_workspace_pins.py, 11 cases) + provenance integration tests; fixed a pre-existing stale RUNTIME_VERSION assertion to read pyproject dynamically. Evidence-Ticket: OMN-12989 * fix(OMN-12989): workspace build must honor sibling lock pins (fail-fast ratchet) The 2026-06-11 stability bootstrap crash was caused by a workspace-mode --no-cache rebuild that vendored omnibase_infra 0.37.0-dev (13-day-stale worktree) even though omnimarket dev's uv.lock pins infra 0.38.1. The downgraded sibling predated the OMN-12501 Protocol-quarantine guard and turned a latent contract defect into a fatal crash. Fix + ratchet (same PR set): - scripts/runtime_build/resolve_workspace_pins.py: parse the consuming repo's uv.lock for sibling pins (version + git rev); classify each staged/installed sibling as exact/ahead/regression/unpinned; fail-fast (WorkspacePinError) on any regression BELOW the lock pin. Stdlib version-tuple fallback when packaging is absent so the ratchet never fails-open. - compute_workspace_provenance.py: enforce sibling pins + a host-infra self-check (installed omnibase_infra vs lock pin — the exact crash vector, since host infra is built from the context, not staged), and emit a pin_comparison block into build-provenance.json for deploy verifiers. - Dockerfile.runtime: COPY resolve_workspace_pins.py beside the provenance script so the in-image import resolves. - TDD: failing tests first (test_resolve_workspace_pins.py, 11 cases) + provenance integration tests; fixed a pre-existing stale RUNTIME_VERSION assertion to read pyproject dynamically. Evidence-Ticket: OMN-12989 * test(OMN-12989): co-locate provenance pin helper in fixture --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13055): absent repos in REPOS list warn and exit 0 instead of failing (#1960) Missing repos (not cloned locally) now emit a WARN line and are tracked in a separate WARNED array. Only real fetch/ff failures cause exit 1. This makes pull-all.sh safe to use on machines with a partial clone set, while keeping the explicit-list override behavior intact. Adds three regression tests: absent-only exits 0, absent+present exits 0 with OK for the present repo, present-failed+absent exits 1. Evidence-Ticket: OMN-13055 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13070): config_prefetcher overlay wins on controlled lanes (#1961) On infisical_required=True lanes, fetched/overlay config now always wins over ambient env. Ambient env is retained only as a declared bootstrap fallback (with an explicit provenance INFO log line) when Infisical returns None. apply_to_environment also overwrites stale env on controlled lanes. Uncontrolled lane (infisical_required=False) behaviour is unchanged. Adds 5 regression tests: controlled-lane Infisical-wins, env-bootstrap- fallback, apply_to_environment overwrite, missing-from-both-is-error, and uncontrolled-lane-env-still-wins. Refactors _resolve_key to return a (outcome, value, error) tuple to satisfy the ≤5-param pattern gate. Source: docs/audits/2026-06-10-runtime-env-overlay-authority-audit.md Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13062): migration-gate vacuity fix — sentinel discipline, wait-for-postgres, skip-manifest, sync-check nonzero (retro A-10) (#1963) * fix(OMN-13062): migration-gate vacuity fix — sentinel discipline, wait-for-postgres, skip-manifest, sync-check nonzero Fixes three bugs identified in retro A-10 (recurrences OMN-12885, OMN-12934): (1) RUNNER SENTINEL DISCIPLINE run-forward-migrations.sh now clears migrations_complete=FALSE at the start of every run and sets it TRUE only as its FINAL act after all infra and node migrations succeed. Any mid-run failure leaves the gate UNHEALTHY. runner_completed_at is stamped at the same final step as durable evidence of a successful completion run. (2) SYNC-NODE-MIGRATIONS VACUOUS GATE sync-node-migrations.sh --check now exits 2 (not 0) when the omnimarket source tree is unresolvable. Silent exit-0 was hiding drift. The single opt-out is SYNC_NODE_MIGRATIONS_SKIP_UNRESOLVABLE=1 for environments that intentionally run without the source. (3) WAIT-FOR-POSTGRES GUARD run-forward-migrations.sh now waits up to PG_WAIT_RETRIES (default 30) x 2s for Postgres to accept connections before proceeding, guarding the first-boot initdb race. (4) SKIP-MANIFEST docker/migrations/skip-manifest.yaml introduced as the sole committed escape for intentionally-skipped migrations. The runner reads this at startup; listed migrations are recorded in schema_migrations with checksum "skip-manifest" without executing the SQL. (5) MIGRATION 085 Adds runner_completed_at TIMESTAMPTZ column to db_metadata so the runner's final stamp is durable in the schema (idempotent ADD COLUMN IF NOT EXISTS). Rollback included. 22 regression tests added covering all five fix surfaces. * fix(OMN-13062): stamp schema fingerprint for migration 085 Migration 085 (085_add_runner_completed_at_to_db_metadata.sql) was added in the initial commit but schema_fingerprint.sha256 was not regenerated. Running `python scripts/check_schema_fingerprint.py stamp` updates the artifact from the stale hash to match the 71 migration files. Evidence-Source: OCC#2563 Evidence-Ticket: OMN-13062 --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12864): Bifrost endpoints → committed overlay authority + fail-loud loader (#1964) * feat(OMN-12864, OMN-12814, OMN-12945): Bifrost endpoints → committed overlay authority + fail-loud loader OMN-12864 — Committed lane overlay - docker/lane-overlays/dev.bifrost.yaml: typed deployment bindings for all four BIFROST_LOCAL_*_ENDPOINT_URL values (coder :8000, reasoner :8001, embedding :8100, ds4-flash :8101). Previously only available as ephemeral shell exports on .201; now committed, auditable, diff-able, and CI-checked. - docker/lane-overlays/dev.bifrost.env: generated dotenv sidecar consumed by compose via env_file; never edited directly (yaml is authority). - docker/docker-compose.infra.yml: wire the env_file block at compose root so the four endpoints are injected into the interpolation context on a clean shell. Hardcode BIFROST_CONTRACT_PATH (remove :-/empty footgun — OMN-12814). - scripts/render_bifrost_lane_overlay_env.py: render script regenerates the env sidecar from the YAML source. - src/omnibase_infra/runtime/models/model_bifrost_lane_overlay.py: ModelBifrostLaneOverlay — typed Pydantic model enforcing URL completeness (OMN-12815: every URL must end in /chat/completions). OMN-12814 — Fail-loud loader - render_bifrost_delegation_contract: raises ProtocolConfigurationError on FileNotFoundError, YAMLError, ValidationError, and zero-endpoint renders. No lru_cache — every restart re-renders from packaged source so a stale cache cannot pin a broken result across deploys. OMN-12945 — Re-seed from packaged source on deploy - docker/entrypoint-runtime.sh: set BIFROST_FORCE_RESEED=1 on every container restart so the named-volume copy is always rebuilt from the packaged bifrost_delegation.yaml merged with committed lane-overlay endpoints. - render_bifrost_delegation_contract: honor BIFROST_FORCE_RESEED/force_reseed flag to bypass the stale-volume early-return path entirely. Tests: - tests/ci/test_bifrost_lane_overlay.py: CI gate — env sidecar in-sync with YAML source; all four BIFROST_LOCAL_* keys present. - tests/unit/runtime/models/test_model_bifrost_lane_overlay.py: bare-base URL rejection, env dict mapping, extra-field rejection. - tests/unit/runtime/test_render_bifrost_delegation_contract.py: fail-loud paths, force-reseed, zero-endpoint error, endpoint URL completeness. - tests/unit/models/test_model_serialization_roundtrip.py: roundtrip coverage. * fix(OMN-12864): move bifrost env_file to service level — fix compose schema validation failure Top-level 'env_file' is rejected by Docker Compose v2 schema validator ('additional properties not allowed'). This caused 10+ compose-render integration tests to fail in CI. Fix: - Remove top-level env_file block from docker-compose.infra.yml - Add per-service env_file on omninode-runtime, runtime-effects, runtime-worker (the three containers that render Bifrost) - Change BIFROST_LOCAL_*:? to BIFROST_LOCAL_*:- in x-runtime-env (compose-level validation removed; Python validates via ModelBifrostLaneOverlay + render_bifrost_delegation_contract) - Add two CI gate tests: compose_env_file_is_service_level_not_top_level and runtime_services_have_bifrost_env_file * fix(OMN-12864): allow BIFROST_LOCAL_* empty defaults in silent-fallback gate The overlay authority pattern (OMN-12864) passes BIFROST_LOCAL_*_ENDPOINT_URL via service-level env_file (docker/lane-overlays/dev.bifrost.env), not at compose config time. Compose-level :? would break CI rendering without the overlay pre-loaded. Validation at the Python layer (ModelBifrostLaneOverlay + render_bifrost_delegation_contract) is the enforcement point. Add the three failing vars to ALLOWED_EMPTY_DEFAULTS with OMN-12864 citation. --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12858/OMN-12879/OMN-12880): dispatcher route-coverage CI gate (#1965) * feat(OMN-12858/OMN-12879/OMN-12880): dispatcher route-coverage CI gate Static contract analysis: every subscribed command topic (onex.cmd.*) must declare handler_routing or runtime_dispatch, or the message goes to DLQ silently. This gate would have caught two recent incidents: 1. June 9 DLQ regression (OMN-12858 post-mortem): node_generation_consumer subscribed onex.cmd.omnimarket.node-generation-requested.v1 but a sole-handler revert left zero dispatcher routes registered. Messages went to DLQ silently with no CI signal. 2. June 12 DEL-01 live finding: onex.cmd.omnimarket.delegate-skill.v1 was consumed by dev lane bus but no dispatcher route existed in any deployed contract. Discovered via manual rpk consumer-group lag probe (DEL-01 evidence, docs/evidence/2026-06-12-weekend-pass/). Deliverables: - scripts/check_dispatcher_route_coverage.py — static YAML scanner that checks both omnibase_infra and omnimarket contract trees; ratchet allowlist for known pre-existing violations; --changed-contracts mode (OMN-12879) for per-PR scoping; compat publish topics excluded (OMN-12880) - .github/workflows/dispatcher-route-coverage.yml — CI workflow that checks out omnimarket sibling, collects changed contract paths in PR mode, and runs the gate; fires on PR, push-to-main, and merge_group - tests/ci/test_dispatcher_route_coverage_gate.py — 12 unit tests covering RED/GREEN/COMPAT/CHANGED-MODE/ALLOWLIST/MULTI-DIR paths plus live-contract regression proof against the actual omnibase_infra tree Allowlist additions: - onex.cmd.omnibase-infra.pattern-b-dispatch.v1 (RuntimePatternBBroker, imperative consumer, OMN-12525 migration target) - onex.cmd.platform.contract-resolve-requested.v1 (transitional HTTP bridge node_contract_resolver_bridge OMN-2756, metadata.transitional=true) [OMN-12858, OMN-12879, OMN-12880] * fix(OMN-12858): drop full uv sync from dispatcher-route-coverage workflow Gate script only needs pyyaml (stdlib + yaml). Using full setup-python-uv was causing 10+ minute timeout. Replace with direct pip install pyyaml and invoke python3 directly. Reduces job from 10m timeout to <1m. [OMN-12858] --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13026): wire transport-mock-lint pre-commit hook + CI gate (#1962) * feat(OMN-13026): wire transport-mock-lint pre-commit hook + CI gate (omnibase_infra) Activates the transport-mock-lint validator (from omnibase_core, OMN-13026) on omnibase_infra. Ratchet baseline: 218 existing violations across 63 files frozen in validation/transport_mock_baseline.yaml. New bare AsyncMock/ MagicMock on EventBus/transport surfaces are blocked by pre-commit hook and CI lint step. Existing violations tracked for drain by per-site tickets (parent OMN-13026). Reference incident: PR #1181 bare AsyncMock hid missing EventBusKafka.stop(). Evidence-Ticket: OMN-13026 * fix(OMN-13026): use uv run python for CI lint step + bump omnibase_core pin to include transport_mock_lint The transport-mock lint CI step previously cloned omnibase_core and ran `python -m omnibase_core.validators.transport_mock_lint` with PYTHONPATH, but this failed: `No module named omnibase_core.validators.transport_mock_lint` because it ran `.venv/bin/python` which uses the locked venv, and the venv omnibase_core pin (2defabef4) predates the transport_mock_lint module. Fix: use `uv run python` (removes the clone step) and bump omnibase-core git pin from 2defabef4 to 309d89fa7 (PR 1231 merge commit on dev) so transport_mock_lint is available in the locked venv. * fix(OMN-13026): align transport mock baseline and runner lock * fix(OMN-13026): sync omnibase_core pin + runner identity lock to dev baseline Align pyproject.toml omnibase_core rev to 2defabef (required by test_release_backmerge_preserves_proven_runtime_core_pin) and update docker/runners/runner-image.lock.json identity_digest/shared_env_digest to match dev runner image lock (79b08f44 / 90c8b3b9). Both were stale from the prior session's pin bump that used an older SHA. * fix(OMN-13026): source transport mock validator from core * fix(OMN-13026): source transport validator from core dev --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13094): --output receipt mode on onex node/run — quiet typed receipts with durable capture (#1966) Phase 2a of the skill-output-suppression slice (epic OMN-13089, plan docs/plans/2026-06-12-skill-output-suppression-plan.md Phase 2 item 1). - onex node/onex run gain --output receipt: ALL runtime logging routes to a run_id-suffixed capture file under <state-root>/captures/ (no console handlers — kills the 25-50-line RuntimeLocal INFO stream at the source); stdout carries exactly ONE typed ModelSkillResult JSON with the FULL handler result (result_model = concrete handler result type FQN). - Durable capture: capture log + handler result content-addressed via omnibase_core ArtifactStore (OMN-13093); artifact.captured + tool.output.captured emitted to the emit daemon socket (--emit-socket, default ~/.claude/emit.sock). - Failure asymmetry: artifact write failure => FULL output printed, no receipt (no hidden loss); emission failure => receipt still prints, event spooled to <state-root>/emit_spool/ for replay. - Node failure => status=failed/error with full error + capture log INLINE in the receipt (errors are never hidden) and artifact-backed. - Default output mode unchanged (enforcement is Phase 4). - RuntimeLocal exposes handler_result (receipt schema identity). - core pin 2defabef -> ae8793bd (merged OMN-13091/13093 receipt models + ArtifactStore); runner-image identity lock regenerated and the OMN-12765 backmerge identity constants updated for the new pin. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13097): onex skill subcommand + declarative skill->node mapping (#1968) * feat(OMN-13097): onex skill subcommand + declarative skill->node mapping Phase 4a of the skill-output-suppression slice (epic OMN-13089, plan docs/plans/2026-06-12-skill-output-suppression-plan.md Phase 4 item 1/2). A dispatch skill IS one CLI call (user directive 2026-06-12). This adds the `onex skill <name> [args]` dispatch surface that the 24 omniclaude shim migrations build on: - onex.cli entry-point `skill` -> cli_skill.run_skill_by_name. Resolves the skill via the declarative skill_mapping.yaml registry, builds the backing node's input payload from the skill's CLI args, writes it under the state root (.onex_state/tmp/<skill>-<run_id>.json — never /tmp), resolves the node's packaged contract exactly like `onex node`, and dispatches through the proven receipt-mode path (run_receipt_mode, OMN-13094). stdout is exactly one typed ModelSkillResult JSON with the FULL handler result. - skill_mapping.yaml: declarative DATA mapping all 24 dispatch shims to their backing onex.nodes node + typed result-model FQN (verified against each live handler handle() return type on origin/dev) + per-arg payload specs + static payload + keyword classifiers (delegate task_type as data, not code). Adding a skill is a YAML edit + fixture, never a CLI code change (ticket deliverable 2/3). Mapping lives beside the node-resolution surface, never hardcoded branching in the CLI. - Typed models split one-per-file (repo convention): ModelSkillArgSpec, ModelSkillClassifier, ModelSkillMapping, ModelSkillMappingRegistry, EnumSkillArgType. Frozen, extra=forbid, fail-fast coercion/validation. - validation_exemptions.yaml: Click-callback param-count + literal-identifier name-field exemptions mirroring the existing cli_node run_node_by_name precedent (OMN-11570) — same pattern, same rationale. dod_evidence: - 20 unit tests pass (registry validity, all-24-shims coverage, FQN result models, arg parsing/coercion/positional/required, classifiers, payload build, receipt-mode dispatch wiring, payload-under-state-root not /tmp). - uv run mypy src/ --strict: clean (2438 source files). - ruff format + check: clean. pre-commit run on changed files: pass. - runner-image identity lock regenerated for the pyproject entry-point add (same as OMN-13094). * test(OMN-13097): rebind OMN-12765 backmerge identity constants for onex skill pyproject change Adding the `skill` onex.cli entry-point to pyproject.toml changes the runner-image identity_digest (and shared_env_digest) the lock binds. Update the hardcoded expected constants in the backmerge-identity test to the regenerated values — same mechanical rebind OMN-13094 performed for the core pin bump. Identity + runner-image-identity tests pass (14). --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix+test(OMN-13012): force terminal-topic metadata refresh so both ephemeral consumers assign (runner consume-leg wedge) (#1969) The runner consume-leg wedged on the live stability battery (image c0505521f1fa, EXP1-3_RUNNER_CONSUME_LEG_BLOCKER): only the FAILED terminal topic ever subscribed while the COMPLETED topic never assigned, so the correlated completed terminal was never read and the 8x2x10 matrix re-fired cell 1 forever, emitting zero non-degenerate rows. A prior two-strike diagnosis proved the omnimarket handler is correct (it opens both terminal sessions pre-publish and races them); the defect is in the omnibase_infra runtime consume leg. Root cause: _assign_direct_terminal_partitions ignored the metadata future returned by AIOKafkaClient.set_topics and re-called set_topics([same_topic]) each loop iteration. aiokafka 0.13.0 set_topics only forces a metadata refresh when the topic set DIFFERS from the tracked set, so every iteration after the first took the no-op branch and never re-fetched. An ephemeral group_id=None consumer whose first metadata fetch had not yet surfaced partitions burned the full 30s assign cap and raised a bare TimeoutError (the empty-message 'wait failed' seen live). Fix: register the reply topic once and await that metadata fetch, then on each miss force a fresh fetch via force_metadata_update (which always fetches) rather than the no-op set_topics repeat. The assign-cap TimeoutError now carries a diagnostic message instead of an empty one. Test: tests/integration/test_terminal_consumer_concurrent_assign_race.py drives the REAL TerminalEventConsumer (the object wired as event_consumer) through the REAL open_direct_terminal_consumer/poll path with AIOKafkaConsumer monkeypatched to a fake that faithfully models aiokafka 0.13.0 set_topics future + metadata-latency semantics. RED before the fix (bare TimeoutError, the live empty-message signature); GREEN after. K>=2 multi-trial variant asserts no worker-thread leak across trials. Evidence-Source: <occ-sha-pending> Evidence-Ticket: OMN-13012 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13096): onex delegate single-command subcommand (Phase 2b) (#1967) * feat(OMN-13096): onex delegate single-command subcommand Add 'onex delegate "<prompt>" [--task-type X] [--max-tokens N]' as a subcommand on the existing onex CLI (Phase 2b of the skill-output-suppression slice, OMN-13089). The command wraps payload construction, node dispatch, and result extraction internally and prints exactly one ModelSkillResult[ModelDelegateSkillResponse] to stdout via the OMN-13094 receipt-mode path. RuntimeLocal logs go to the capture file + artifact store, never to stdout; scratch payloads live under <state-root>/tmp/ with run_id suffixes (never /tmp). - cli_delegate.py: classify_task_type (keyword table from legacy skill md), payload write, contract resolve, run_receipt_mode dispatch - register 'delegate' under onex.cli entry points - exempt delegate_command from the >5-param patterns gate (same Click-callback rationale as run_node_by_name) - 20 unit tests: classification, scratch-under-state-root, single typed receipt on stdout, zero INFO log leakage omnibase_infra does NOT depend on omnimarket; the delegate node is resolved at runtime via the onex.nodes entry-point group (registered by omnimarket). * chore(OMN-13096): re-trigger deploy-gate after Evidence-Source set to OCC#2593 No code change — the deploy-gate workflow triggers on synchronize (not edited), so the PR-body Evidence-Source fix needs a new commit to re-resolve the OCC ref to the open PR head where contracts/OMN-13096.yaml (with deploy evidence) lives. * chore(OMN-13096): re-trigger deploy-gate now that OCC#2593 merged to OCC dev contracts/OMN-13096.yaml (with the dod-deploy-onex-delegate item) is now on OCC dev, so the deploy-gate OCC-dev checkout resolves the contract + deploy evidence. * chore(OMN-13096): regenerate runner-image identity lock for pyproject entry-point add Adding the 'delegate' onex.cli entry point to pyproject.toml changed the dependency-manifest digest that scripts/ci/runner_image_identity.py folds into the runner-image identity lock. Regenerate the lock so tests/ci/test_runner_image_identity.py matches (was the only CI test failure; unrelated environmental integration/perf failures excluded). * test(OMN-13096): update backmerge identity assertions to regenerated lock digests The runner-image identity lock was regenerated for the pyproject entry-point add; this test hardcodes the expected identity_digest/shared_env_digest, so update both to match the new lock (same maintenance OMN-13094 did). --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): tolerate partition-less reply topic in terminal consume leg (#1970) The context-ROI runner opens one ephemeral group_id=None terminal consumer per terminal topic BEFORE publishing each generation command (subscribe-before-publish, OMN-13012/13038). The FAILED reply topic is only produced to on contract_passed=False; in a battery where generations pass it has zero messages, so Redpanda never advertises a partition for it. _assign_direct_terminal_partitions burned the full 30s assign cap on every trial then raised a bare TimeoutError (the empty-message 'wait failed'), stalling each of the 160 battery trials ~30s before the COMPLETED terminal could correlate -> battery needs >80 min and never completes (verifier-confirmed wedge). A partition-less reply topic is a valid steady state, not a 30s error: - _assign_direct_terminal_partitions gives a bounded grace window for a topic that exists but is slow to surface metadata, then assigns whatever partitions exist (possibly none) and returns promptly instead of burning the cap and raising. - poll_direct_terminal_consumer treats an empty assignment as 'no terminal will arrive here' (sleeps out its timeout, returns None) without calling getone() on an unassigned consumer. Repro: tests/integration/test_terminal_consumer_battery_load_wedge.py drives the REAL TerminalEventConsumer over K=10 x 2 cells x 2 arms; RED (40/40 trials block a full assign cap) before the fix, GREEN after. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): pin terminal-consumer read offset synchronously (close lazy seek_to_end publish race) (#1971) * fix(OMN-13118): pin terminal-consumer read offset synchronously to close lazy seek_to_end publish race The consume-leg wedge survived PR #1969 (set_topics no-op) and PR #1970 (partition-less assign-cap) because both addressed the assign phase, not the seek timing. AIOKafkaConsumer.seek_to_end is LAZY: it requests a LATEST offset reset that only resolves on the first poll — AFTER the caller publishes. With generation completing in ~1s, the correlated COMPLETED terminal lands in the open->poll gap, so the lazily-resolved LATEST position is the HWM AFTER the record and the poll reads past it. The terminal is never read, the trial never correlates, and the experiment matrix re-fires the same cell forever. Replace seek_to_end with a synchronous end_offsets() + seek() pin (_pin_direct_terminal_end_offsets) in both open_direct_terminal_consumer and RuntimePatternBBroker._dispatch_and_wait_with_direct_kafka_consumer, so the read position is fixed at open() time, before the publish — the real subscribe-before-publish guarantee. Empty assignment (partition-less reply topic, OMN-13118 #1970) is a no-op. Repro: tests/integration/test_terminal_consumer_seek_reset_race.py drives the real TerminalEventConsumer.open()/wait() the way HandlerContextRoiRunner does, publishing the correlated terminal in the open->wait gap across K=10 x 2 cells x 2 arms; RED with the lazy reset (every cell degenerate), GREEN once the read offset is pinned. RED verified by git-stashing only the source fix. Existing consume-leg fakes updated to model end_offsets/seek (they previously masked the bug by making seek_to_end a synchronous exact snapshot). * test(OMN-13118): reword assertion (lazy not deferred) for receipt honesty gate * fix(OMN-13118): bound end_offsets() round-trip with assign-cap timeout (CodeRabbit) end_offsets() is a broker ListOffsets round-trip aiokafka documents as able to block indefinitely. Bound it with the same cap as start()/assign so a stalled broker fails fast instead of hanging the pre-publish positioning. --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13137): _validate_routing strategy-aware event_model checks (#1973) operation_match routes by the `operation` field and does not use event_model. The validator was unconditionally requiring event_model.{name,module} for every handler entry, causing 230/295 omnimarket operation_match contracts (all correct as authored) to fail routing validation at startup. Fix: read routing_strategy from the routing map and branch validation: - payload_type_match → require event_model.{name, module} (unchanged) - operation_match (and any non-payload strategy) → require `operation`; skip event_model checks entirely Updated pre-existing _validate_routing tests to declare routing_strategy: payload_type_match explicitly (they always tested payload_type_match semantics but relied on the implicit fallback that is now removed). Added test_validate_routing_operation_match.py with 4 unit tests: 1. operation_match without event_model → zero event_model errors 2. operation_match missing operation field → error 3. payload_type_match missing event_model → still errors (regression guard) 4. Real node_integration_sweep_orchestrator routing block → clean (boot gate) Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): independent per-terminal-topic consumers in Pattern B direct-Kafka wait (#1972) The consume-leg wedge survived four merged fixes (#1969 set_topics no-op, #1970 partition-less assign-cap, #1971 synchronous seek-pin). The STRONG K>=10 multi-cell reprobe on the stability lane still wedged on REBUILD-5 (cdf53d963f7b). Converged diagnosis (strikes 3+4, docs/evidence/2026-06-12-weekend-pass/experiments/probe4-stability/ reprobe-K10-rebuild5/HALT_K10_WEDGE_PERSISTS.md): the runtime waited for each trial's terminal across TWO topics (node-generation-completed.v1 + node-generation-failed.v1) with a SINGLE ephemeral group_id=None consumer assigned both topics' partitions. One aiokafka consumer holds one manual subscription; the COMPLETED delivery window collapsed before it surfaced the correlated record, so the trial never correlated and the matrix re-fired cell 1. RuntimePatternBBroker._dispatch_and_wait_with_direct_kafka_consumer now opens ONE independent AIOKafkaConsumer PER terminal topic via open_direct_terminal_consumer (each started, assigned, and offset-pinned via end_offsets()+seek() at open() BEFORE publish), then awaits both CONCURRENTLY via asyncio.wait(FIRST_COMPLETED). The first correlated terminal wins; both are torn down. No shared consumer, no subscription flip. Keeps the #1970 partition-less no-op and the #1971 synchronous offset pin (both live in open_direct_terminal_consumer / poll_direct_terminal_consumer). Removes the now-dead single-consumer helpers (_assign_terminal_topic_partitions, method-level _refresh_terminal_topic_metadata, _direct_kafka_* kwargs builders, _kafka_bootstrap_servers/_kafka_event_bus). Adds tests/integration/test_terminal_consumer_subscription_flip_wedge.py: a real-dispatch-path K>=10 x 2-cell x 2-arm repro whose fake models TWO independent consumers honestly (delivery is faithful only for a single-topic assignment; a consumer spanning both topics flips and drops the COMPLETED record). RED genuineness verified by reverting only the source to the single-consumer shape (test hangs past timeout); GREEN with the fix in 2.3s. Acceptance is the LIVE K>=10 multi-cell stability-lane reprobe (later phase), NOT this unit test. A green unit repro is necessary but NOT sufficient. Refs OMN-13118, OMN-13128. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): long-lived terminal correlator replaces per-trial ephemeral consume leg (#1974) Tier B canonical redesign (epic OMN-12525). Five offset/subscription patches (#1969-#1972) tuned the per-trial-ephemeral terminal consumer and all failed the live K>=10 batt…
…+ OMN-13666 entrypoint tolerance (emergency prod recovery) (#2128) * fix(OMN-13670): surgical main hotfix — re-apply OMN-13654 dep floors + OMN-13666 entrypoint tolerance EMERGENCY PROD RECOVERY — OMN-13670 operator-authorized. NOT a normal dev->main promotion. OMN-13654: Bump pyjwt>=2.13.0, python-multipart>=0.0.30, starlette>=1.3.1 as explicit security-floor constraints in pyproject.toml. Relocks uv.lock (pyjwt 2.13.0, python-multipart 0.0.32, starlette 1.3.1). Resolves 5 HIGH CVEs that have been blocking build-and-push-runtime.yml since 2026-06-07 with Trivy exit-code 1. OMN-13666: Runtime entrypoint now treats PRIMARY (omnibase_infra) stamp as REQUIRED (failure aborts boot, exit 1) and SECONDARY (omniintelligence) stamp as BEST-EFFORT (failure logs WARNING, boot continues). Resolves the prod crash-loop where "permission denied for table db_metadata" on the omniintelligence DB was bringing all 7 prod runtime deployments to 0/1. * fix(OMN-13670): refresh runner image identity lock * chore(OMN-13670): refresh hotfix check context * fix(OMN-13670): compare node migrations against main for main hotfix * fix(OMN-13670): align runner identity backmerge guard --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
… last Trivy HIGH (emergency prod recovery) (#2131) * fix(OMN-13670): floor cryptography>=48.0.1 in runtime builder — clear last Trivy HIGH Adds build-time pip floor for cryptography>=48.0.1 in the BUILDER stage of docker/Dockerfile.runtime, matching the existing protobuf/setuptools floor pattern. Clears GHSA-537c-gmf6-5ccf (HIGH, cryptography 46.0.7→48.0.1). This is the final floor completing the clean-main Trivy fix for OMN-13670. Also includes yamlfmt and ruff format auto-fixes on pre-existing files. * chore(OMN-13670): refresh hotfix checks --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
…t runner→ECR EOF (emergency prod recovery) (#2134) Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
….46.1 / spi 0.23.0) (#2164) * docs(OMN-12962): contract-store durability audit — cold-runtime census proof (#1937) Proves cold-start contract census reconstructs from the image-bundled filesystem manifest (HYBRID-mode bootstrap + PluginLoaderContractSource), independent of node-registration.v1 retention. The delete-retention topic feeds only the post-freeze dynamic listener (auto_offset_reset=latest, no history replay). Live .201 stability-test evidence: filesystem contract_path in manifest + truncated topic log head with intact census. No store fix needed; residual dynamic-only gap covered by runtime_sweep sweep check. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12970): vendor omnimarket projection node migrations into forward/nodes (#1938) Vendors omnimarket node-owned projection migrations into the namespaced forward-migration tree so run-forward-migrations.sh materializes them in the dashboard projection DB (omnidash_analytics) at deploy. Primary (OMN-12970): creates llm_call_metrics, llm_cost_aggregates, and capability_scores in the projection DB. These were only ever created in the omnibase_infra DB by infra migrations 031/060, so the ab-compare, cost.token_usage, cost.summary, and capability-scores projection topics were DEGRADED at startup ('table not found') and their dashboard panels rendered empty. Also re-syncs three omnimarket node migrations the vendor tree had drifted from (node_projection_llm_routing, node_projection_overnight, node_projection_savings /077) — sync-node-migrations.sh --check requires the full vendored tree to match omnimarket source, and these were missing. Companion to omnimarket PR for the same ticket (source migrations + projection table-coverage ratchet test). Evidence-Ticket: OMN-12970 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12965): stamp runtime image identity (version + revision) in workspace builds (#1943) * fix(OMN-12965): stamp runtime image identity (version + revision) in workspace builds The main runtime image stamped org.opencontainers.image.version=0.1.0 with a blank org.opencontainers.image.revision after workspace rebuilds. A blank identity degrades every proof packet (runtime SHA + image digest are required citations in accepted evidence). Root causes (three build paths under-stamped identity): - onex up --build (cmd_up) passed only GIT_SHA; the runtime-stage OCI labels read VCS_REF (-> blank revision) and RUNTIME_VERSION (-> placeholder 0.1.0). - deploy-runtime.sh passed VCS_REF but not RUNTIME_VERSION/GIT_SHA. - Dockerfile silently allowed blank/placeholder identity in workspace mode. Fix: - cli._image_identity_build_args() stamps the full quad (GIT_SHA/VCS_REF/ RUNTIME_VERSION/BUILD_DATE) and fails fast on an unresolved git revision. - deploy-runtime.sh stamps RUNTIME_VERSION + GIT_SHA and verifies the version label is non-placeholder post-deploy. - Dockerfile.runtime fails workspace builds with blank VCS_REF or placeholder RUNTIME_VERSION=0.1.0 (release mode unaffected). Enforcement ratchet (same PR): - scripts/check_runtime_image_identity.py static check, wired as pre-commit hook + CI gate (ci.yml). - tests/unit/infra/test_runtime_image_identity_labels.py pins the cli helpers + Dockerfile guard; deploy-agent test extended for the quad. Proven locally via throwaway docker builds: workspace+args -> populated labels; workspace without args -> guard fails (exit 64); release without args -> 0.1.0 placeholder allowed (no regression). Evidence-Ticket: OMN-12965 * test(OMN-12965): integration build proof for runtime image identity labels Builds the real runtime-stage ARG/LABEL/guard block against busybox and asserts via docker inspect: workspace+args -> populated version/revision; workspace without args -> guard fails (exit 64); release -> placeholder allowed. Satisfies the integration-test hard gate and makes the throwaway proof permanent. Evidence-Ticket: OMN-12965 --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12987): workspace-build sibling lock-pin preflight + provenance (#1944) Recurrence ratchet for the 2026-06-11 stability bootstrap crash. The 11:20Z --no-cache rebuild vendored omnibase_infra 0.37.0-dev (~2c1d672f) + core 0.42.0 even though omnimarket dev's uv.lock pinned infra 0.38.1 @ e2dbdc95 + core 0.44.0 @ c97c2c9a. The stale sibling predated the OMN-12501 Protocol-quarantine guard, turning a latent contract defect into a fatal crash that crash-looped the main runtime. - check_sibling_lock_pins.py: host-side fail-fast preflight resolving expected sibling versions/SHAs from the consuming repo's (omnimarket) uv.lock and comparing them against each vendored tree. Mismatch aborts the build. - stage_workspace.sh: emit a .build-sha marker per staged sibling (rsync drops .git) so the preflight and provenance can identify the vendored commit. - deploy-runtime.sh: run the preflight after staging, before build; abort on mismatch. Write the comparison under sibling-repos/.sibling-lock-pins.json. - compute_workspace_provenance.py + Dockerfile.runtime: fold expected-vs-actual lock_pin_comparison into build-provenance.json for deploy verifiers. Recurrence-guard tests prove a stale infra 0.37.0 vs lock-pinned 0.38.1 fails the preflight and matched pins pass; deploy script wiring is asserted statically. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12988): pin stability runtime-worker to replicas 1 + census ratchet (#1942) The base docker-compose.infra.yml defaults runtime-worker to replicas 0 (${WORKER_REPLICAS:-0}). The stability-test lane's required state includes a running worker (4-container census: main, effects, worker, projection-api), but the override pinned it via an env-interpolation default (${STABILITY_TEST_WORKER_REPLICAS:-1}) — a silent-drop surface: a stray exported STABILITY_TEST_WORKER_REPLICAS=0 or removal of the :-1 fallback would scale the worker to 0 with zero signal on a plain compose up/recreate. Fix: pin docker-compose.stability-test.yml runtime-worker deploy.replicas to the literal 1 (no env indirection). Ratchet (recurrence guards, same PR): - scripts/deploy-agent/tests/unit/test_runtime_worker_census.py: assert runtime-worker stays in the deploy-agent RUNTIME-scope census so a missing worker (replicas 0 => absent from docker compose ps) is a deploy failure, not silence; assert the override pins a literal 1. - tests/integration/infra/test_stability_test_runtime_compose_render.py: assert the rendered stability worker resolves deploy.replicas == 1. - tests/unit/infra/test_stability_test_runtime_lane.py: update the existing pin assertion to the literal 1. Evidence-Ticket: OMN-12988 Config-drift family: OMN-12945 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12979): expire-bound topic completeness suppressions (#1940) Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12966): accept community-edition Infisical /api/status in provision readiness gate (#1939) The fresh-provision readiness gate in provision-infisical.py only accepted the enterprise {"status": "ok"} payload and rejected the community edition's {"message": "Ok"}, returning 1 before bootstrap could run. This blocked provisioning against the Infisical instance deployed on .201 (community edition). Route the gate through the existing _is_infisical_ready helper (single source of truth, already used by the already-provisioned path). Add TestMainFreshProvision- ReadinessGate covering community/enterprise/not-ready cases. Also adds docker/docker-compose.infisical-stability.yml: an ADDITIVE Infisical compose project for the stability-test lane (joins the existing network as external, reuses stability postgres/valkey, no lane mutation), since the lane overlays disable the in-lane Infisical service via *-disabled profile overrides. P1.2b-A: Infisical now reachable from the stability runtime/effect containers; a known secret (OMN_12966_PROBE) seeds and resolves end-to-end via the machine- identity universal-auth path, verified from inside the effects container. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12958): volume-config drift gate + runtime config provenance (#1941) * feat(OMN-12958): volume-config drift gate + runtime provenance Compute config provenance (path + sha256) for the runtime-rendered Bifrost delegation contract; the deployed volume copy survives rebuilds and silently diverges from packaged source (two competing authorities, OMN-12945). - runtime/config_provenance.py: ModelConfigProvenance + drift classification, sidecar JSON writer (read by sweep + proof packets) - runtime/health/health_config_provenance.py: drift -> degraded health - render entrypoint logs provenance line + writes sidecar on every boot - docs/runbooks/volume-config-drift-and-reseed.md: ledgered re-seed procedure - validation exemption for config_name (logical identifier, not entity ref) No live volume mutation: re-seed is an operator deploy step (deploy_pending). * test(OMN-12958): cover volume config drift reseed flow --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12957): wire runtime-profile validator + core registry parity guard (#1945) * feat(OMN-12957): wire runtime-profile validator + core registry parity guard - Remove the import-time RuntimeError drift raise in runtime_profile.py (a hard raise on core/infra version skew would crash the kernel at import); the parity invariant is enforced by test_profiles_match_core_registry instead. - Add tests: _PROFILES keys == omnibase_core REGISTERED_RUNTIME_PROFILES, and every CONSUMER_ATTACHED_RUNTIME_PROFILES profile loads as a real lane. - Wire omnibase_core.validation.validator_runtime_profiles as a pre-commit hook + validator-runtime-profiles.yml CI gate on infra contracts. - Freeze 19 pre-existing violators in validation/runtime_profiles_allowlist.yaml (discovered by repo-root walk; drain via OMN-12982). Blocks NEW orphans. Requires the omnibase_core pin to include OMN-12957's validator (new rules). Evidence-Ticket: OMN-12957 Evidence-Source: 5463fbaf819409d4fb7f491dd4f276f10d869eba * ci(OMN-12957): pass runtime profile allowlist to validator * test(OMN-12957): cover runtime profile registry parity * fix(OMN-12957): keep runtime profile allowlist under config * fix(OMN-12957): pin core runtime profile registry --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12973): widen migration-gate healthcheck start_period + ratchet (#1950) P2.8: classify the prod migration-gate UNHEALTHY-then-self-resolved incident. Classification: idle-one-shot-mis-modeled = NO. The migration-gate is a correctly-modeled long-running sentinel (entrypoint `while true; sleep 3600`) whose healthcheck continuously polls db_metadata.migrations_complete via check_migrations_complete.sh. The container flipped UNHEALTHY transiently because its healthcheck start_period (10s) was far shorter than the real cold-volume migration window (~116s: prod gate started 09:35:22, intelligence-migration finished 09:37:18). Past the 10s grace window the still-failing probe was reported UNHEALTHY until migrations completed, then self-resolved — no fault. Fix: raise migration-gate healthcheck start_period 10s -> 180s in both the authoritative catalog manifest (docker/catalog/services/migration-gate.yaml, flows into the generated compose) and the hand-maintained docker/docker-compose.infra.yml that deploy-runtime.sh applies to .201, so a still-applying gate stays in `health: starting` instead of flipping UNHEALTHY. Ratchet (enforcement, not detection): new ValidatorHealthcheckStartPeriod in omnibase_infra catalog (validator_healthcheck_start_period.py) asserts a 120s floor for migration-completion gates, wired into `onex validate runtime` (cmd_validate_runtime) AND backed by unit tests that gate every PR via pre-commit + CI. A migration-completion gate is identified by BOTH a healthcheck that polls migration completion AND a service_completed_successfully dependency, so ordinary app services (e.g. intelligence-api, an HTTP liveness probe with a legitimately short 40s start_period) are not swept into the floor. Prod probed read-only only; no prod mutation. Applies to live prod via the batched stability/prod rebuild (deploy_pending). Evidence-Ticket: OMN-12973 Evidence-Source: OCC#PENDING Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12971): deliver Vertex ADC bearer token to runtime effects container (secret-ref discipline) (#1949) * feat(OMN-12971): deliver Vertex ADC bearer token to runtime effects container (secret-ref discipline) Wire the Vertex $500-credit ADC path into the runtime, ADDITIVE next to the Gemini API-key path (provider-agnostic; neither provider removed or forced). runtime_policy.contract.yaml: add a secret-source mapping llm.vertex.access_token (source_type=file, /run/secrets/vertex_access_token) to the dev/stability-test/ judge profiles, alongside the existing llm.gemini.api_key env mapping. The ref name MUST match cloud-vertex-gemini.secret_ref in omnimarket bifrost_delegation.yaml. The resolved VALUE is a short-lived OAuth bearer token minted from ADC, refreshed by the operator; the token VALUE is never committed — only the ref name + in-container path. Add aiplatform.googleapis.com to the cloud host allowlist. docker-compose.infra.yml: bind the operator-supplied host token file read-only to /run/secrets/vertex_access_token on the main and effects runtimes (VERTEX_ACCESS_TOKEN_HOST_FILE, default /dev/null so lanes without Vertex still start; Gemini key path unaffected). Pass through BIFROST_VERTEX_GEMINI_ENDPOINT_URL (overlay supplies the complete Vertex OpenAI-compat URL) and GOOGLE_CLOUD_PROJECT/LOCATION (default empty). runtime-policy.env: regenerated from the contract via render_runtime_policy_env (test_runtime_policy_env_matches_contract_renderer proves contract<->env parity). test_runtime_policy_contract.py: update host-allowlist assertion for the additive Vertex host. Pre-existing platform-wide topic-parity-gate failures (25 unrelated topics) are identical with and without this change (proven by diff); that gate is pre-commit- only (not a CI merge gate) and this change adds zero new topic gaps, so that one hook is SKIP-ped. No deploy/receipt/merge gate is bypassed. * fix(OMN-12971): make Vertex runtime env contract-owned --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13008): automated .201 disk maintenance — worktree GC + docker GC + watermark alert ratchet (#1952) * feat(OMN-13008): automated .201 disk maintenance — worktree GC + docker GC + watermark alert ratchet Conservative, keep-list-driven disk maintenance to prevent the 2026-06-11 /data ~95% outage that killed all three lanes mid-demo: - scripts/worktree-gc.sh: drives the canonical omniclaude prune-worktrees.sh (merged+clean+pushed safety) on both Mac (merge-sweep tick) and .201 - scripts/disk-gc.sh + disk_gc_plan.py: conservative docker/builder/image GC. Pure, testable removal planner honoring a VERSIONED keep-list (deploy/disk-gc/keep-list.yaml): never reaps a kept repo, kept tag, in-use image, or anything younger than min_age_days; keeps N superseded generations. - scripts/disk-watermark-check.sh + disk_watermark_event.py: df watermark ratchet. >=85% emits a typed disk-watermark bus event (warning) that the sweep auto-ticket path turns into a Linear ticket; >=90% emits critical. Broker addr is fail-fast from KAFKA_BOOTSTRAP_SERVERS (no localhost default). - deploy/disk-gc/: systemd USER timer (onex-disk-gc.timer/.service, hourly) + install-disk-gc.sh. User units, NOT lane containers. - tests: 20 unit tests incl. GC plan-safety invariants + dry-run proof that default mode issues no destructive op (a wrong-delete GC is worse than none). Contract: contracts/OMN-13008.yaml * fix(OMN-13008): pass docker inventory to GC planner via stdin, not env (ARG_MAX) On a host with many docker images, passing the full image/ps inventory as env vars to disk_gc_plan.py exceeds ARG_MAX ('Argument list too long', exit 126), producing an empty plan. Write inventory to per-run scratch files (under the log dir, never /tmp; cleaned on exit) and hand it to the planner on stdin as a JSON envelope. Verified the failure live on .201; planner now reads stdin. * fix(OMN-13008): simplify GC plan stdin pipe (two processes, no nested subprocess) * fix(OMN-13008): keep-wins reconciliation — never remove an image id with any keep reason A single image id can surface in multiple 'docker image ls' rows (one per repo:tag). One tag could route the id to dangling-removal while another routes it to keep (e.g. tagged 'latest' or within-N-generations). Live .201 plan showed an id in BOTH remove_image_ids and kept_reasons. Reconcile at the end: keep wins — any id with a keep reason is dropped from the remove list; remove list deduped. Adds 2 regression tests. Verified live on .201. * fix(OMN-13008): timer uses OnCalendar=hourly for reliable re-arm (oneshot service) OnUnitActiveSec does not reliably re-elapse for a oneshot service once it goes inactive (observed NextElapseUSecMonotonic=infinity live on .201 after the first run). Switch to OnCalendar=hourly + RandomizedDelaySec so the timer re-arms every hour. Keep Persistent=true for missed-run catch-up. --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13005): materialize blocking event_consumer in runtime auto-wiring (runner consume-leg degenerate rows) (#1951) * fix(OMN-13005): materialize blocking event_consumer in auto-wiring (runner consume-leg degenerate rows) The runtime auto-wiring materialized event_publisher for handlers that declare it but had no equivalent for event_consumer. Request/response EFFECT handlers (HandlerContextRoiRunner) that publish a command then block on the correlated terminal event fell back to their no-op consumer default, returning None immediately -> every result row degenerate (failure_stage=generation, attempt_count=0) while generations succeeded ~1s later. Adds _make_sync_event_consumer (mirror of _make_sync_event_publisher), backed by service_terminal_event_consumer.make_terminal_event_consumer: a sync (topic, correlation_id, timeout) -> dict | None adapter that runs the proven direct-Kafka correlate-and-wait loop (from RuntimePatternBBroker) on an isolated event loop in a worker thread, so blocking does not deadlock the runtime dispatch loop that delivers the awaited terminal. TDD through the REAL dispatch path: test_event_consumer_injection drives a trial through _prepare_handler_wiring with a terminal arriving after a delay and asserts a non-degenerate row; verified RED with injection disabled. * fix(OMN-13012): two-phase (seek-now/wait-later) terminal event_consumer to close the subscribe-after-publish race (#1954) The OMN-13005 injected event_consumer is a single callable that does assign -> seek_to_end -> poll internally, all AFTER the handler has already published its command. Once OMN-13010 freed the dispatch loop and generation began completing in ~1s, the correlated terminal lands BEFORE the single-call consumer's post-publish seek_to_end positions, so seek_to_end skips PAST the already-emitted terminal and the runner times out on an offset beyond it (probe3, run_id=20260611T2140Z-probe3 -- both arms degenerate, zero rebalances). Splits positioning from waiting so the caller subscribes BEFORE it publishes: session = consumer.open(topic) # assign + seek_to_end NOW publisher(command_topic, payload) # publish AFTER positioning payload = session.wait(cid, timeout) # block from the captured position The returned TerminalEventConsumer is still directly callable with the legacy (topic, cid, timeout) -> dict | None single-call shape for any consumer that does not need subscribe-before-publish; the runner is the only consumer today. TerminalConsumerSession owns a dedicated event loop on a daemon worker thread for the whole open->wait->close lifecycle, preserving the OMN-13005 loop-isolation discipline so blocking never deadlocks the runtime dispatch loop. TDD (real dispatch path, RED-then-GREEN): extends the OMN-13005 injection test with a terminal emitted IMMEDIATELY after publish. The single-call (seek-after-publish) consumer MISSES it (degenerate row -- RED test asserts failure_stage=generation); the two-phase (open-before-publish) consumer CATCHES it (non-degenerate -- GREEN). The Kafka layer is faked at the two service seams against a shared in-memory log modeling seek-to-end semantics. OMN-13005 blocking-correlate behavior preserved. 269/269 auto_wiring unit tests pass; mypy --strict clean. Sibling to OMN-13010 / OMN-13005 / OMN-13003. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13005): route terminal consumer through Kafka boundary --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12990): pin worker replicas in ledgered config + fail-fast lane overrides (silent-drop ratchet) (#1948) * fix(OMN-12990): pin worker replicas in ledgered config + fail-fast lane overrides The base compose sets runtime-worker deploy replicas to ${WORKER_REPLICAS:-0} (soft-default ZERO). The stability lane's required state includes a running worker (GATE_ZERO_PROOF.md: 4 runtime containers), but the worker presence rode entirely on a compose soft-default — any plain compose up/recreate without the policy env silently scaled the worker to zero with no error and no signal. Fix (contract-native + fail-fast): - Add 'replicas' to ModelRuntimeProcessPolicy; pin replicas: 1 in every lane's worker block in runtime_policy.contract.yaml. - Renderer emits {PROFILE}_WORKER_REPLICAS into the ledgered runtime-policy.env for dev/stability-test/judge/prod. - stability + prod compose overrides reference ${..._WORKER_REPLICAS:?...} (fail-fast, NO silent :-1/:-0 default). A recreate that omits the policy env now aborts loudly instead of dropping the worker. prod previously had no override at all and inherited the dangerous :-0 default. Ratchet (recurrence guards): - tests asserting fail-fast override form (no soft default), contract-declared replica pin >= 1 per lane, and rendered {PROFILE}_WORKER_REPLICAS=1 in the ledgered env. - runbook deploy/verify procedure adds an expected-container census (worker must be present) via verify_container_manifest; a missing worker is a FAILURE, not silence. Note: SKIP=topic-parity-gate — that local-only advisory gate (absent from all .github/workflows, not a required CI check) fails on 25 pre-existing cross-repo topic gaps (build-loop/omniclaude/omniweb) identical on pristine base HEAD 8d7da1249; this change adds zero topics. All other hooks ran clean. Evidence-Ticket: OMN-12990 * test(OMN-12990): cover worker replica policy integration --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12909): add gateway bus forwarder P0A (#1946) * feat(OMN-12909): add gateway bus forwarder p0a * test(OMN-12909): add gateway forwarder integration coverage * fix(OMN-12909): satisfy gateway forwarder validators * test(OMN-12909): allow gateway forwarder bus protocol * fix(OMN-12909): sync gateway forwarder entry point * fix(OMN-12909): refresh runner image identity lock * test(OMN-12909): relax JSON normalizer mixed benchmark threshold * fix(OMN-12909): allow gateway handlers to boot unconfigured * fix(OMN-12909): declare gateway forwarder runtime profile * fix(OMN-12909): update runner identity backmerge expectation --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13011): LANE CENSUS RECONCILIATION ratchet — declared desired-state per lane, drift = auto-ticket (#1955) The class fix for the recurring lane-drift regression. Nothing reconciled the declared desired state of a runtime lane against what is actually running, so the same failure kept recurring with zero signal: volume config drift (OMN-12945), WORKER_REPLICAS silent zero (OMN-12988/12990), and on 2026-06-11 prod runtime containers plus the broker network were silently absent for hours during demo prep. Ships a per-lane DESIRED-STATE census: - (a) DECLARED in a versioned lane manifest (deploy/lane-census/lane-manifest.yaml): container set, network, replicas, image-tag pattern per lane (stability-test/prod/judge/dev), derived from the canonical compose lane files. A parity ratchet keeps the manifest locked in step with the compose files. - (b) RECONCILED on a schedule on .201 by SHARING the OMN-13008 systemd timer (a drop-in 4th ExecStart on onex-disk-gc.service — never a second timer) and on-demand via scripts/lane-census-check.sh / runtime_sweep. - (c) Drift = typed bus event (onex.evt.infra.lane-census-drift.v1) + Linear auto-ticket naming exactly what is missing/extra (container_absent, network_detached, replicas_zero, unexpected_container, oneshot_failed/stuck, image_tag_mismatch). Fail-fast, no warn-only mode (gates-block policy); exit 30 on drift; bus publish fail-fast on KAFKA_BOOTSTRAP_SERVERS (no localhost default). Red fixture reproduces 2026-06-11: prod runtime containers absent + broker network detached must produce the exact drift findings + a non-zero exit hours before a human noticed. Pure planner is fully unit-tested; shell driver dry-run-tested. Builds on the OMN-12988 deploy-agent RUNTIME census (deploy-time) as the complementary steady-state reconciler; closes the runtime-worker.yaml container_name: null census gap by sourcing names from the compose lane files. Evidence-Ticket: OMN-13011 Config-drift family: OMN-12945 Relates-to: OMN-13009, OMN-12988, OMN-13008 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13020): vendor missing node migrations — llm_routing 0000 + context_roi 001 (#1956) Vendors two omnimarket node-source migrations into the infra forward-migration tree via scripts/sync-node-migrations.sh (the canonical OMN-12559 mechanism): - node_projection_llm_routing/0000_create_llm_routing_decisions.sql (source: omnimarket #1168 / OMN-12942, merge ed6734f8) - node_projection_context_roi/001_create_context_roi_scores.sql (source: omnimarket #1178 / OMN-12955, merge 5010b1f4) Without the 0000 base table, node_projection_llm_routing/0001 (CREATE VIEW) hard-fails against NODE_POSTGRES_DB=omnidash_analytics — exactly the prod forward-migration exit-3 of 2026-06-11T09:35:52Z, and reproduced by construction in any clean clones@dev build. Files are byte-identical to the omnimarket dev blobs (sha256 f8a8b339… / c4126e65…) and to the untracked hot-patch copies on the .201 stability clone. Both migrations are self-contained, all-statements-IF-NOT-EXISTS, and 0000 sorts lexically before 0001 within the node's namespaced identity space. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13058): close TerminalConsumerSession on open() failure (worker thread + event loop leak) (#1957) TerminalConsumerSession.__init__ starts its dedicated worker loop thread immediately. TerminalEventConsumer.open did 'session = TerminalConsumerSession(...); return session.open()' with no cleanup: any failure inside session.open() (consumer start timeout, partition-assign timeout, broker auth error) propagated out of the raising expression, the session reference was lost, and the daemon worker thread plus its never-closed asyncio event loop leaked -- one pair per failed open. The motivating caller (HandlerContextRoiRunner) opens a session per trial, so a 160-560-trial battery against a degraded broker accumulates hundreds of leaked threads in the long-lived effects container. Fix: wrap session.open() in try/except BaseException -> session.close() (idempotent: stops the loop, joins the thread) -> re-raise. Covers both the two-phase .open(topic) path and the legacy single-call __call__ path. Found by the P3.3 doctrinal review of merged #1951 (b9712af9 / 36d98275). TDD: tests/unit/runtime/test_service_terminal_event_consumer_open_failure.py injects a real open failure through the production path (event bus without _bootstrap_servers) and asserts no alive terminal-consumer-* thread after the raise. Verified RED with the fix stashed (2 failed), GREEN with it (2 passed). Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13021): non-dev-base guard — fail feature-base PRs absent Stacked-Parent declaration (retro A-6) (#1958) Any PR whose base is neither dev nor main fails unless the body carries 'Stacked-Parent: #N'. Prevents the feedback_stacked_prs_orphan_from_dev class (#1185/#1954 auto-merged INTO parent feature branches and stranded off dev). base=main remains governed by main-target-guard. Epic OMN-13013 (process enforcement ratchets — June 12 retro). Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13014): hot-patch ledger rebuild preflight gate (retro B-1) (#1959) * feat(OMN-13014): hot-patch ledger rebuild preflight gate (retro B-1) Hot-patches on .201 (.prepatch sibling discipline) silently revert on any image rebuild/force-recreate — the 2026-06-11 20:58Z rebuild already erased a live /api/generate patch once. This adds the rebuild-path gate: - scripts/preflight_hotpatch_ledger.py: given a target container or lane + per-repo build refs, hard-fails when any hot-patch ledger row's source PR merge commit is not an ancestor of the build ref (git merge-base --is-ancestor), plus a .prepatch tripwire of the running container (unledgered .prepatch = hard fail; --post-rebuild = zero .prepatch expected). Sole bypass: HOTPATCH_PREFLIGHT_BYPASS carrying the Rule-10 '# skip-token-allowed: <user-approval-receipt-id>' form. - scripts/deploy-runtime.sh: guard_hotpatch_ledger wired into main() before build/preview (both dry-run and execute), lane derived from the compose project; skips loudly only when no ledger exists on the host. - tests/unit/scripts/test_preflight_hotpatch_ledger.py: 17 unit tests (ancestor gate, lane scoping, ledger loading, tripwire, bypass forms). - tests/ci/test_receipt_gate_install_guard.py: repair stale guard — core OMN-12565 replaced the OMN-9198 'uv pip uninstall first' install step with a cleared workspace venv (uv venv --clear); assert the new contract. Ledger backfilled from live census (5 .prepatch files / 4 source PRs, all MERGED to dev) at /data/omninode/hotpatch-ledger/ledger.yaml on .201. * fix(OMN-13014): scope missing-prepatch tripwire warning to the probed container --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12977): workspace-build sibling-pin ratchet — honor consuming lock, fail-fast on stale vendor (#1936) * fix(OMN-12977): workspace-build sibling-pin ratchet — honor consuming lock, fail-fast on stale vendor The workspace-mode image build (BUILD_SOURCE=workspace, used by the stability-test deploy procedure) vendored sibling/foundation packages from whatever the canonical OMNI_HOME clones happened to be checked out at, ignoring the consuming repo's uv.lock. On 2026-06-11 this shipped a 13-day-stale omnibase_infra 0.37.0-dev (pre-OMN-12501 Protocol-quarantine guard) + core 0.42.0 against an omnimarket dev lock pinning infra 0.38.1@e2dbdc95 / core 0.44.0@c97c2c9a, dropping the guard and crashing wire_from_manifest bootstrap fatally (stability lane down on demo day). Fix + recurrence ratchet (same PR): - scripts/runtime_build/check_sibling_lock_pins.py: parse the consuming repo's uv.lock for expected version+git-rev of each foundation/sibling package (scoped to the package's own source line so editable/registry pins are not cross-attributed a dependency's rev), resolve the actual clone version+HEAD, compare, and classify drift backward/forward/none. Fail-fast (exit 1) on any drift; --allow-drift records an explicit operator override in the artifact, never silent. - stage_workspace.sh: runs the preflight against the canonical clones before staging; aborts the build (exit 3) on unacknowledged drift and writes workspace/sibling-pin-comparison.json. - compute_workspace_provenance.py: folds the expected-vs-actual comparison into build-provenance.json so deploy verifiers can assert the build honored the lock; flags unacknowledged drift as a provenance error. - Dockerfile.runtime: COPY the comparison artifact (committed placeholder so the COPY always resolves; overwritten by stage_workspace.sh in workspace mode). - TDD: 19 unit tests covering lock parsing (git/registry/editable sources), drift classification, the exact 0.37.0-vs-0.38.1 stale case, check_pins exit codes, and the allow-drift override. Pre-existing mypy-strict bare-dict errors in compute_workspace_provenance.py fixed in the same pass. Evidence-Ticket: OMN-12977 Evidence-Source: pending-occ * ci(OMN-12977): retry runtime smoke compose port race * test(OMN-12977): align sibling-pin script tests with current API --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12989): workspace-mode image build must honor sibling lock pins (#1947) * fix(OMN-12989): workspace build must honor sibling lock pins (fail-fast ratchet) The 2026-06-11 stability bootstrap crash was caused by a workspace-mode --no-cache rebuild that vendored omnibase_infra 0.37.0-dev (13-day-stale worktree) even though omnimarket dev's uv.lock pins infra 0.38.1. The downgraded sibling predated the OMN-12501 Protocol-quarantine guard and turned a latent contract defect into a fatal crash. Fix + ratchet (same PR set): - scripts/runtime_build/resolve_workspace_pins.py: parse the consuming repo's uv.lock for sibling pins (version + git rev); classify each staged/installed sibling as exact/ahead/regression/unpinned; fail-fast (WorkspacePinError) on any regression BELOW the lock pin. Stdlib version-tuple fallback when packaging is absent so the ratchet never fails-open. - compute_workspace_provenance.py: enforce sibling pins + a host-infra self-check (installed omnibase_infra vs lock pin — the exact crash vector, since host infra is built from the context, not staged), and emit a pin_comparison block into build-provenance.json for deploy verifiers. - Dockerfile.runtime: COPY resolve_workspace_pins.py beside the provenance script so the in-image import resolves. - TDD: failing tests first (test_resolve_workspace_pins.py, 11 cases) + provenance integration tests; fixed a pre-existing stale RUNTIME_VERSION assertion to read pyproject dynamically. Evidence-Ticket: OMN-12989 * fix(OMN-12989): workspace build must honor sibling lock pins (fail-fast ratchet) The 2026-06-11 stability bootstrap crash was caused by a workspace-mode --no-cache rebuild that vendored omnibase_infra 0.37.0-dev (13-day-stale worktree) even though omnimarket dev's uv.lock pins infra 0.38.1. The downgraded sibling predated the OMN-12501 Protocol-quarantine guard and turned a latent contract defect into a fatal crash. Fix + ratchet (same PR set): - scripts/runtime_build/resolve_workspace_pins.py: parse the consuming repo's uv.lock for sibling pins (version + git rev); classify each staged/installed sibling as exact/ahead/regression/unpinned; fail-fast (WorkspacePinError) on any regression BELOW the lock pin. Stdlib version-tuple fallback when packaging is absent so the ratchet never fails-open. - compute_workspace_provenance.py: enforce sibling pins + a host-infra self-check (installed omnibase_infra vs lock pin — the exact crash vector, since host infra is built from the context, not staged), and emit a pin_comparison block into build-provenance.json for deploy verifiers. - Dockerfile.runtime: COPY resolve_workspace_pins.py beside the provenance script so the in-image import resolves. - TDD: failing tests first (test_resolve_workspace_pins.py, 11 cases) + provenance integration tests; fixed a pre-existing stale RUNTIME_VERSION assertion to read pyproject dynamically. Evidence-Ticket: OMN-12989 * test(OMN-12989): co-locate provenance pin helper in fixture --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13055): absent repos in REPOS list warn and exit 0 instead of failing (#1960) Missing repos (not cloned locally) now emit a WARN line and are tracked in a separate WARNED array. Only real fetch/ff failures cause exit 1. This makes pull-all.sh safe to use on machines with a partial clone set, while keeping the explicit-list override behavior intact. Adds three regression tests: absent-only exits 0, absent+present exits 0 with OK for the present repo, present-failed+absent exits 1. Evidence-Ticket: OMN-13055 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13070): config_prefetcher overlay wins on controlled lanes (#1961) On infisical_required=True lanes, fetched/overlay config now always wins over ambient env. Ambient env is retained only as a declared bootstrap fallback (with an explicit provenance INFO log line) when Infisical returns None. apply_to_environment also overwrites stale env on controlled lanes. Uncontrolled lane (infisical_required=False) behaviour is unchanged. Adds 5 regression tests: controlled-lane Infisical-wins, env-bootstrap- fallback, apply_to_environment overwrite, missing-from-both-is-error, and uncontrolled-lane-env-still-wins. Refactors _resolve_key to return a (outcome, value, error) tuple to satisfy the ≤5-param pattern gate. Source: docs/audits/2026-06-10-runtime-env-overlay-authority-audit.md Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13062): migration-gate vacuity fix — sentinel discipline, wait-for-postgres, skip-manifest, sync-check nonzero (retro A-10) (#1963) * fix(OMN-13062): migration-gate vacuity fix — sentinel discipline, wait-for-postgres, skip-manifest, sync-check nonzero Fixes three bugs identified in retro A-10 (recurrences OMN-12885, OMN-12934): (1) RUNNER SENTINEL DISCIPLINE run-forward-migrations.sh now clears migrations_complete=FALSE at the start of every run and sets it TRUE only as its FINAL act after all infra and node migrations succeed. Any mid-run failure leaves the gate UNHEALTHY. runner_completed_at is stamped at the same final step as durable evidence of a successful completion run. (2) SYNC-NODE-MIGRATIONS VACUOUS GATE sync-node-migrations.sh --check now exits 2 (not 0) when the omnimarket source tree is unresolvable. Silent exit-0 was hiding drift. The single opt-out is SYNC_NODE_MIGRATIONS_SKIP_UNRESOLVABLE=1 for environments that intentionally run without the source. (3) WAIT-FOR-POSTGRES GUARD run-forward-migrations.sh now waits up to PG_WAIT_RETRIES (default 30) x 2s for Postgres to accept connections before proceeding, guarding the first-boot initdb race. (4) SKIP-MANIFEST docker/migrations/skip-manifest.yaml introduced as the sole committed escape for intentionally-skipped migrations. The runner reads this at startup; listed migrations are recorded in schema_migrations with checksum "skip-manifest" without executing the SQL. (5) MIGRATION 085 Adds runner_completed_at TIMESTAMPTZ column to db_metadata so the runner's final stamp is durable in the schema (idempotent ADD COLUMN IF NOT EXISTS). Rollback included. 22 regression tests added covering all five fix surfaces. * fix(OMN-13062): stamp schema fingerprint for migration 085 Migration 085 (085_add_runner_completed_at_to_db_metadata.sql) was added in the initial commit but schema_fingerprint.sha256 was not regenerated. Running `python scripts/check_schema_fingerprint.py stamp` updates the artifact from the stale hash to match the 71 migration files. Evidence-Source: OCC#2563 Evidence-Ticket: OMN-13062 --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12864): Bifrost endpoints → committed overlay authority + fail-loud loader (#1964) * feat(OMN-12864, OMN-12814, OMN-12945): Bifrost endpoints → committed overlay authority + fail-loud loader OMN-12864 — Committed lane overlay - docker/lane-overlays/dev.bifrost.yaml: typed deployment bindings for all four BIFROST_LOCAL_*_ENDPOINT_URL values (coder :8000, reasoner :8001, embedding :8100, ds4-flash :8101). Previously only available as ephemeral shell exports on .201; now committed, auditable, diff-able, and CI-checked. - docker/lane-overlays/dev.bifrost.env: generated dotenv sidecar consumed by compose via env_file; never edited directly (yaml is authority). - docker/docker-compose.infra.yml: wire the env_file block at compose root so the four endpoints are injected into the interpolation context on a clean shell. Hardcode BIFROST_CONTRACT_PATH (remove :-/empty footgun — OMN-12814). - scripts/render_bifrost_lane_overlay_env.py: render script regenerates the env sidecar from the YAML source. - src/omnibase_infra/runtime/models/model_bifrost_lane_overlay.py: ModelBifrostLaneOverlay — typed Pydantic model enforcing URL completeness (OMN-12815: every URL must end in /chat/completions). OMN-12814 — Fail-loud loader - render_bifrost_delegation_contract: raises ProtocolConfigurationError on FileNotFoundError, YAMLError, ValidationError, and zero-endpoint renders. No lru_cache — every restart re-renders from packaged source so a stale cache cannot pin a broken result across deploys. OMN-12945 — Re-seed from packaged source on deploy - docker/entrypoint-runtime.sh: set BIFROST_FORCE_RESEED=1 on every container restart so the named-volume copy is always rebuilt from the packaged bifrost_delegation.yaml merged with committed lane-overlay endpoints. - render_bifrost_delegation_contract: honor BIFROST_FORCE_RESEED/force_reseed flag to bypass the stale-volume early-return path entirely. Tests: - tests/ci/test_bifrost_lane_overlay.py: CI gate — env sidecar in-sync with YAML source; all four BIFROST_LOCAL_* keys present. - tests/unit/runtime/models/test_model_bifrost_lane_overlay.py: bare-base URL rejection, env dict mapping, extra-field rejection. - tests/unit/runtime/test_render_bifrost_delegation_contract.py: fail-loud paths, force-reseed, zero-endpoint error, endpoint URL completeness. - tests/unit/models/test_model_serialization_roundtrip.py: roundtrip coverage. * fix(OMN-12864): move bifrost env_file to service level — fix compose schema validation failure Top-level 'env_file' is rejected by Docker Compose v2 schema validator ('additional properties not allowed'). This caused 10+ compose-render integration tests to fail in CI. Fix: - Remove top-level env_file block from docker-compose.infra.yml - Add per-service env_file on omninode-runtime, runtime-effects, runtime-worker (the three containers that render Bifrost) - Change BIFROST_LOCAL_*:? to BIFROST_LOCAL_*:- in x-runtime-env (compose-level validation removed; Python validates via ModelBifrostLaneOverlay + render_bifrost_delegation_contract) - Add two CI gate tests: compose_env_file_is_service_level_not_top_level and runtime_services_have_bifrost_env_file * fix(OMN-12864): allow BIFROST_LOCAL_* empty defaults in silent-fallback gate The overlay authority pattern (OMN-12864) passes BIFROST_LOCAL_*_ENDPOINT_URL via service-level env_file (docker/lane-overlays/dev.bifrost.env), not at compose config time. Compose-level :? would break CI rendering without the overlay pre-loaded. Validation at the Python layer (ModelBifrostLaneOverlay + render_bifrost_delegation_contract) is the enforcement point. Add the three failing vars to ALLOWED_EMPTY_DEFAULTS with OMN-12864 citation. --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12858/OMN-12879/OMN-12880): dispatcher route-coverage CI gate (#1965) * feat(OMN-12858/OMN-12879/OMN-12880): dispatcher route-coverage CI gate Static contract analysis: every subscribed command topic (onex.cmd.*) must declare handler_routing or runtime_dispatch, or the message goes to DLQ silently. This gate would have caught two recent incidents: 1. June 9 DLQ regression (OMN-12858 post-mortem): node_generation_consumer subscribed onex.cmd.omnimarket.node-generation-requested.v1 but a sole-handler revert left zero dispatcher routes registered. Messages went to DLQ silently with no CI signal. 2. June 12 DEL-01 live finding: onex.cmd.omnimarket.delegate-skill.v1 was consumed by dev lane bus but no dispatcher route existed in any deployed contract. Discovered via manual rpk consumer-group lag probe (DEL-01 evidence, docs/evidence/2026-06-12-weekend-pass/). Deliverables: - scripts/check_dispatcher_route_coverage.py — static YAML scanner that checks both omnibase_infra and omnimarket contract trees; ratchet allowlist for known pre-existing violations; --changed-contracts mode (OMN-12879) for per-PR scoping; compat publish topics excluded (OMN-12880) - .github/workflows/dispatcher-route-coverage.yml — CI workflow that checks out omnimarket sibling, collects changed contract paths in PR mode, and runs the gate; fires on PR, push-to-main, and merge_group - tests/ci/test_dispatcher_route_coverage_gate.py — 12 unit tests covering RED/GREEN/COMPAT/CHANGED-MODE/ALLOWLIST/MULTI-DIR paths plus live-contract regression proof against the actual omnibase_infra tree Allowlist additions: - onex.cmd.omnibase-infra.pattern-b-dispatch.v1 (RuntimePatternBBroker, imperative consumer, OMN-12525 migration target) - onex.cmd.platform.contract-resolve-requested.v1 (transitional HTTP bridge node_contract_resolver_bridge OMN-2756, metadata.transitional=true) [OMN-12858, OMN-12879, OMN-12880] * fix(OMN-12858): drop full uv sync from dispatcher-route-coverage workflow Gate script only needs pyyaml (stdlib + yaml). Using full setup-python-uv was causing 10+ minute timeout. Replace with direct pip install pyyaml and invoke python3 directly. Reduces job from 10m timeout to <1m. [OMN-12858] --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13026): wire transport-mock-lint pre-commit hook + CI gate (#1962) * feat(OMN-13026): wire transport-mock-lint pre-commit hook + CI gate (omnibase_infra) Activates the transport-mock-lint validator (from omnibase_core, OMN-13026) on omnibase_infra. Ratchet baseline: 218 existing violations across 63 files frozen in validation/transport_mock_baseline.yaml. New bare AsyncMock/ MagicMock on EventBus/transport surfaces are blocked by pre-commit hook and CI lint step. Existing violations tracked for drain by per-site tickets (parent OMN-13026). Reference incident: PR #1181 bare AsyncMock hid missing EventBusKafka.stop(). Evidence-Ticket: OMN-13026 * fix(OMN-13026): use uv run python for CI lint step + bump omnibase_core pin to include transport_mock_lint The transport-mock lint CI step previously cloned omnibase_core and ran `python -m omnibase_core.validators.transport_mock_lint` with PYTHONPATH, but this failed: `No module named omnibase_core.validators.transport_mock_lint` because it ran `.venv/bin/python` which uses the locked venv, and the venv omnibase_core pin (2defabef4) predates the transport_mock_lint module. Fix: use `uv run python` (removes the clone step) and bump omnibase-core git pin from 2defabef4 to 309d89fa7 (PR 1231 merge commit on dev) so transport_mock_lint is available in the locked venv. * fix(OMN-13026): align transport mock baseline and runner lock * fix(OMN-13026): sync omnibase_core pin + runner identity lock to dev baseline Align pyproject.toml omnibase_core rev to 2defabef (required by test_release_backmerge_preserves_proven_runtime_core_pin) and update docker/runners/runner-image.lock.json identity_digest/shared_env_digest to match dev runner image lock (79b08f44 / 90c8b3b9). Both were stale from the prior session's pin bump that used an older SHA. * fix(OMN-13026): source transport mock validator from core * fix(OMN-13026): source transport validator from core dev --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13094): --output receipt mode on onex node/run — quiet typed receipts with durable capture (#1966) Phase 2a of the skill-output-suppression slice (epic OMN-13089, plan docs/plans/2026-06-12-skill-output-suppression-plan.md Phase 2 item 1). - onex node/onex run gain --output receipt: ALL runtime logging routes to a run_id-suffixed capture file under <state-root>/captures/ (no console handlers — kills the 25-50-line RuntimeLocal INFO stream at the source); stdout carries exactly ONE typed ModelSkillResult JSON with the FULL handler result (result_model = concrete handler result type FQN). - Durable capture: capture log + handler result content-addressed via omnibase_core ArtifactStore (OMN-13093); artifact.captured + tool.output.captured emitted to the emit daemon socket (--emit-socket, default ~/.claude/emit.sock). - Failure asymmetry: artifact write failure => FULL output printed, no receipt (no hidden loss); emission failure => receipt still prints, event spooled to <state-root>/emit_spool/ for replay. - Node failure => status=failed/error with full error + capture log INLINE in the receipt (errors are never hidden) and artifact-backed. - Default output mode unchanged (enforcement is Phase 4). - RuntimeLocal exposes handler_result (receipt schema identity). - core pin 2defabef -> ae8793bd (merged OMN-13091/13093 receipt models + ArtifactStore); runner-image identity lock regenerated and the OMN-12765 backmerge identity constants updated for the new pin. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13097): onex skill subcommand + declarative skill->node mapping (#1968) * feat(OMN-13097): onex skill subcommand + declarative skill->node mapping Phase 4a of the skill-output-suppression slice (epic OMN-13089, plan docs/plans/2026-06-12-skill-output-suppression-plan.md Phase 4 item 1/2). A dispatch skill IS one CLI call (user directive 2026-06-12). This adds the `onex skill <name> [args]` dispatch surface that the 24 omniclaude shim migrations build on: - onex.cli entry-point `skill` -> cli_skill.run_skill_by_name. Resolves the skill via the declarative skill_mapping.yaml registry, builds the backing node's input payload from the skill's CLI args, writes it under the state root (.onex_state/tmp/<skill>-<run_id>.json — never /tmp), resolves the node's packaged contract exactly like `onex node`, and dispatches through the proven receipt-mode path (run_receipt_mode, OMN-13094). stdout is exactly one typed ModelSkillResult JSON with the FULL handler result. - skill_mapping.yaml: declarative DATA mapping all 24 dispatch shims to their backing onex.nodes node + typed result-model FQN (verified against each live handler handle() return type on origin/dev) + per-arg payload specs + static payload + keyword classifiers (delegate task_type as data, not code). Adding a skill is a YAML edit + fixture, never a CLI code change (ticket deliverable 2/3). Mapping lives beside the node-resolution surface, never hardcoded branching in the CLI. - Typed models split one-per-file (repo convention): ModelSkillArgSpec, ModelSkillClassifier, ModelSkillMapping, ModelSkillMappingRegistry, EnumSkillArgType. Frozen, extra=forbid, fail-fast coercion/validation. - validation_exemptions.yaml: Click-callback param-count + literal-identifier name-field exemptions mirroring the existing cli_node run_node_by_name precedent (OMN-11570) — same pattern, same rationale. dod_evidence: - 20 unit tests pass (registry validity, all-24-shims coverage, FQN result models, arg parsing/coercion/positional/required, classifiers, payload build, receipt-mode dispatch wiring, payload-under-state-root not /tmp). - uv run mypy src/ --strict: clean (2438 source files). - ruff format + check: clean. pre-commit run on changed files: pass. - runner-image identity lock regenerated for the pyproject entry-point add (same as OMN-13094). * test(OMN-13097): rebind OMN-12765 backmerge identity constants for onex skill pyproject change Adding the `skill` onex.cli entry-point to pyproject.toml changes the runner-image identity_digest (and shared_env_digest) the lock binds. Update the hardcoded expected constants in the backmerge-identity test to the regenerated values — same mechanical rebind OMN-13094 performed for the core pin bump. Identity + runner-image-identity tests pass (14). --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix+test(OMN-13012): force terminal-topic metadata refresh so both ephemeral consumers assign (runner consume-leg wedge) (#1969) The runner consume-leg wedged on the live stability battery (image c0505521f1fa, EXP1-3_RUNNER_CONSUME_LEG_BLOCKER): only the FAILED terminal topic ever subscribed while the COMPLETED topic never assigned, so the correlated completed terminal was never read and the 8x2x10 matrix re-fired cell 1 forever, emitting zero non-degenerate rows. A prior two-strike diagnosis proved the omnimarket handler is correct (it opens both terminal sessions pre-publish and races them); the defect is in the omnibase_infra runtime consume leg. Root cause: _assign_direct_terminal_partitions ignored the metadata future returned by AIOKafkaClient.set_topics and re-called set_topics([same_topic]) each loop iteration. aiokafka 0.13.0 set_topics only forces a metadata refresh when the topic set DIFFERS from the tracked set, so every iteration after the first took the no-op branch and never re-fetched. An ephemeral group_id=None consumer whose first metadata fetch had not yet surfaced partitions burned the full 30s assign cap and raised a bare TimeoutError (the empty-message 'wait failed' seen live). Fix: register the reply topic once and await that metadata fetch, then on each miss force a fresh fetch via force_metadata_update (which always fetches) rather than the no-op set_topics repeat. The assign-cap TimeoutError now carries a diagnostic message instead of an empty one. Test: tests/integration/test_terminal_consumer_concurrent_assign_race.py drives the REAL TerminalEventConsumer (the object wired as event_consumer) through the REAL open_direct_terminal_consumer/poll path with AIOKafkaConsumer monkeypatched to a fake that faithfully models aiokafka 0.13.0 set_topics future + metadata-latency semantics. RED before the fix (bare TimeoutError, the live empty-message signature); GREEN after. K>=2 multi-trial variant asserts no worker-thread leak across trials. Evidence-Source: <occ-sha-pending> Evidence-Ticket: OMN-13012 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13096): onex delegate single-command subcommand (Phase 2b) (#1967) * feat(OMN-13096): onex delegate single-command subcommand Add 'onex delegate "<prompt>" [--task-type X] [--max-tokens N]' as a subcommand on the existing onex CLI (Phase 2b of the skill-output-suppression slice, OMN-13089). The command wraps payload construction, node dispatch, and result extraction internally and prints exactly one ModelSkillResult[ModelDelegateSkillResponse] to stdout via the OMN-13094 receipt-mode path. RuntimeLocal logs go to the capture file + artifact store, never to stdout; scratch payloads live under <state-root>/tmp/ with run_id suffixes (never /tmp). - cli_delegate.py: classify_task_type (keyword table from legacy skill md), payload write, contract resolve, run_receipt_mode dispatch - register 'delegate' under onex.cli entry points - exempt delegate_command from the >5-param patterns gate (same Click-callback rationale as run_node_by_name) - 20 unit tests: classification, scratch-under-state-root, single typed receipt on stdout, zero INFO log leakage omnibase_infra does NOT depend on omnimarket; the delegate node is resolved at runtime via the onex.nodes entry-point group (registered by omnimarket). * chore(OMN-13096): re-trigger deploy-gate after Evidence-Source set to OCC#2593 No code change — the deploy-gate workflow triggers on synchronize (not edited), so the PR-body Evidence-Source fix needs a new commit to re-resolve the OCC ref to the open PR head where contracts/OMN-13096.yaml (with deploy evidence) lives. * chore(OMN-13096): re-trigger deploy-gate now that OCC#2593 merged to OCC dev contracts/OMN-13096.yaml (with the dod-deploy-onex-delegate item) is now on OCC dev, so the deploy-gate OCC-dev checkout resolves the contract + deploy evidence. * chore(OMN-13096): regenerate runner-image identity lock for pyproject entry-point add Adding the 'delegate' onex.cli entry point to pyproject.toml changed the dependency-manifest digest that scripts/ci/runner_image_identity.py folds into the runner-image identity lock. Regenerate the lock so tests/ci/test_runner_image_identity.py matches (was the only CI test failure; unrelated environmental integration/perf failures excluded). * test(OMN-13096): update backmerge identity assertions to regenerated lock digests The runner-image identity lock was regenerated for the pyproject entry-point add; this test hardcodes the expected identity_digest/shared_env_digest, so update both to match the new lock (same maintenance OMN-13094 did). --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): tolerate partition-less reply topic in terminal consume leg (#1970) The context-ROI runner opens one ephemeral group_id=None terminal consumer per terminal topic BEFORE publishing each generation command (subscribe-before-publish, OMN-13012/13038). The FAILED reply topic is only produced to on contract_passed=False; in a battery where generations pass it has zero messages, so Redpanda never advertises a partition for it. _assign_direct_terminal_partitions burned the full 30s assign cap on every trial then raised a bare TimeoutError (the empty-message 'wait failed'), stalling each of the 160 battery trials ~30s before the COMPLETED terminal could correlate -> battery needs >80 min and never completes (verifier-confirmed wedge). A partition-less reply topic is a valid steady state, not a 30s error: - _assign_direct_terminal_partitions gives a bounded grace window for a topic that exists but is slow to surface metadata, then assigns whatever partitions exist (possibly none) and returns promptly instead of burning the cap and raising. - poll_direct_terminal_consumer treats an empty assignment as 'no terminal will arrive here' (sleeps out its timeout, returns None) without calling getone() on an unassigned consumer. Repro: tests/integration/test_terminal_consumer_battery_load_wedge.py drives the REAL TerminalEventConsumer over K=10 x 2 cells x 2 arms; RED (40/40 trials block a full assign cap) before the fix, GREEN after. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): pin terminal-consumer read offset synchronously (close lazy seek_to_end publish race) (#1971) * fix(OMN-13118): pin terminal-consumer read offset synchronously to close lazy seek_to_end publish race The consume-leg wedge survived PR #1969 (set_topics no-op) and PR #1970 (partition-less assign-cap) because both addressed the assign phase, not the seek timing. AIOKafkaConsumer.seek_to_end is LAZY: it requests a LATEST offset reset that only resolves on the first poll — AFTER the caller publishes. With generation completing in ~1s, the correlated COMPLETED terminal lands in the open->poll gap, so the lazily-resolved LATEST position is the HWM AFTER the record and the poll reads past it. The terminal is never read, the trial never correlates, and the experiment matrix re-fires the same cell forever. Replace seek_to_end with a synchronous end_offsets() + seek() pin (_pin_direct_terminal_end_offsets) in both open_direct_terminal_consumer and RuntimePatternBBroker._dispatch_and_wait_with_direct_kafka_consumer, so the read position is fixed at open() time, before the publish — the real subscribe-before-publish guarantee. Empty assignment (partition-less reply topic, OMN-13118 #1970) is a no-op. Repro: tests/integration/test_terminal_consumer_seek_reset_race.py drives the real TerminalEventConsumer.open()/wait() the way HandlerContextRoiRunner does, publishing the correlated terminal in the open->wait gap across K=10 x 2 cells x 2 arms; RED with the lazy reset (every cell degenerate), GREEN once the read offset is pinned. RED verified by git-stashing only the source fix. Existing consume-leg fakes updated to model end_offsets/seek (they previously masked the bug by making seek_to_end a synchronous exact snapshot). * test(OMN-13118): reword assertion (lazy not deferred) for receipt honesty gate * fix(OMN-13118): bound end_offsets() round-trip with assign-cap timeout (CodeRabbit) end_offsets() is a broker ListOffsets round-trip aiokafka documents as able to block indefinitely. Bound it with the same cap as start()/assign so a stalled broker fails fast instead of hanging the pre-publish positioning. --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13137): _validate_routing strategy-aware event_model checks (#1973) operation_match routes by the `operation` field and does not use event_model. The validator was unconditionally requiring event_model.{name,module} for every handler entry, causing 230/295 omnimarket operation_match contracts (all correct as authored) to fail routing validation at startup. Fix: read routing_strategy from the routing map and branch validation: - payload_type_match → require event_model.{name, module} (unchanged) - operation_match (and any non-payload strategy) → require `operation`; skip event_model checks entirely Updated pre-existing _validate_routing tests to declare routing_strategy: payload_type_match explicitly (they always tested payload_type_match semantics but relied on the implicit fallback that is now removed). Added test_validate_routing_operation_match.py with 4 unit tests: 1. operation_match without event_model → zero event_model errors 2. operation_match missing operation field → error 3. payload_type_match missing event_model → still errors (regression guard) 4. Real node_integration_sweep_orchestrator routing block → clean (boot gate) Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): independent per-terminal-topic consumers in Pattern B direct-Kafka wait (#1972) The consume-leg wedge survived four merged fixes (#1969 set_topics no-op, #1970 partition-less assign-cap, #1971 synchronous seek-pin). The STRONG K>=10 multi-cell reprobe on the stability lane still wedged on REBUILD-5 (cdf53d963f7b). Converged diagnosis (strikes 3+4, docs/evidence/2026-06-12-weekend-pass/experiments/probe4-stability/ reprobe-K10-rebuild5/HALT_K10_WEDGE_PERSISTS.md): the runtime waited for each trial's terminal across TWO topics (node-generation-completed.v1 + node-generation-failed.v1) with a SINGLE ephemeral group_id=None consumer assigned both topics' partitions. One aiokafka consumer holds one manual subscription; the COMPLETED delivery window collapsed before it surfaced the correlated record, so the trial never correlated and the matrix re-fired cell 1. RuntimePatternBBroker._dispatch_and_wait_with_direct_kafka_consumer now opens ONE independent AIOKafkaConsumer PER terminal topic via open_direct_terminal_consumer (each started, assigned, and offset-pinned via end_offsets()+seek() at open() BEFORE publish), then awaits both CONCURRENTLY via asyncio.wait(FIRST_COMPLETED). The first correlated terminal wins; both are torn down. No shared consumer, no subscription flip. Keeps the #1970 partition-less no-op and the #1971 synchronous offset pin (both live in open_direct_terminal_consumer / poll_direct_terminal_consumer). Removes the now-dead single-consumer helpers (_assign_terminal_topic_partitions, method-level _refresh_terminal_topic_metadata, _direct_kafka_* kwargs builders, _kafka_bootstrap_servers/_kafka_event_bus). Adds tests/integration/test_terminal_consumer_subscription_flip_wedge.py: a real-dispatch-path K>=10 x 2-cell x 2-arm repro whose fake models TWO independent consumers honestly (delivery is faithful only for a single-topic assignment; a consumer spanning both topics flips and drops the COMPLETED record). RED genuineness verified by reverting only the source to the single-consumer shape (test hangs past timeout); GREEN with the fix in 2.3s. Acceptance is the LIVE K>=10 multi-cell stability-lane reprobe (later phase), NOT this unit test. A green unit repro is necessary but NOT sufficient. Refs OMN-13118, OMN-13128. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): long-lived terminal correlator replaces per-trial ephemeral consume leg (#1974) Tier B canonical redesign (epic OMN-12525). Five offset/subscription patches (#1969-#1972) tuned the per-trial-ephemeral terminal consumer and all fa…
…push unblocked) (#2166) * docs(OMN-12962): contract-store durability audit — cold-runtime census proof (#1937) Proves cold-start contract census reconstructs from the image-bundled filesystem manifest (HYBRID-mode bootstrap + PluginLoaderContractSource), independent of node-registration.v1 retention. The delete-retention topic feeds only the post-freeze dynamic listener (auto_offset_reset=latest, no history replay). Live .201 stability-test evidence: filesystem contract_path in manifest + truncated topic log head with intact census. No store fix needed; residual dynamic-only gap covered by runtime_sweep sweep check. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12970): vendor omnimarket projection node migrations into forward/nodes (#1938) Vendors omnimarket node-owned projection migrations into the namespaced forward-migration tree so run-forward-migrations.sh materializes them in the dashboard projection DB (omnidash_analytics) at deploy. Primary (OMN-12970): creates llm_call_metrics, llm_cost_aggregates, and capability_scores in the projection DB. These were only ever created in the omnibase_infra DB by infra migrations 031/060, so the ab-compare, cost.token_usage, cost.summary, and capability-scores projection topics were DEGRADED at startup ('table not found') and their dashboard panels rendered empty. Also re-syncs three omnimarket node migrations the vendor tree had drifted from (node_projection_llm_routing, node_projection_overnight, node_projection_savings /077) — sync-node-migrations.sh --check requires the full vendored tree to match omnimarket source, and these were missing. Companion to omnimarket PR for the same ticket (source migrations + projection table-coverage ratchet test). Evidence-Ticket: OMN-12970 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12965): stamp runtime image identity (version + revision) in workspace builds (#1943) * fix(OMN-12965): stamp runtime image identity (version + revision) in workspace builds The main runtime image stamped org.opencontainers.image.version=0.1.0 with a blank org.opencontainers.image.revision after workspace rebuilds. A blank identity degrades every proof packet (runtime SHA + image digest are required citations in accepted evidence). Root causes (three build paths under-stamped identity): - onex up --build (cmd_up) passed only GIT_SHA; the runtime-stage OCI labels read VCS_REF (-> blank revision) and RUNTIME_VERSION (-> placeholder 0.1.0). - deploy-runtime.sh passed VCS_REF but not RUNTIME_VERSION/GIT_SHA. - Dockerfile silently allowed blank/placeholder identity in workspace mode. Fix: - cli._image_identity_build_args() stamps the full quad (GIT_SHA/VCS_REF/ RUNTIME_VERSION/BUILD_DATE) and fails fast on an unresolved git revision. - deploy-runtime.sh stamps RUNTIME_VERSION + GIT_SHA and verifies the version label is non-placeholder post-deploy. - Dockerfile.runtime fails workspace builds with blank VCS_REF or placeholder RUNTIME_VERSION=0.1.0 (release mode unaffected). Enforcement ratchet (same PR): - scripts/check_runtime_image_identity.py static check, wired as pre-commit hook + CI gate (ci.yml). - tests/unit/infra/test_runtime_image_identity_labels.py pins the cli helpers + Dockerfile guard; deploy-agent test extended for the quad. Proven locally via throwaway docker builds: workspace+args -> populated labels; workspace without args -> guard fails (exit 64); release without args -> 0.1.0 placeholder allowed (no regression). Evidence-Ticket: OMN-12965 * test(OMN-12965): integration build proof for runtime image identity labels Builds the real runtime-stage ARG/LABEL/guard block against busybox and asserts via docker inspect: workspace+args -> populated version/revision; workspace without args -> guard fails (exit 64); release -> placeholder allowed. Satisfies the integration-test hard gate and makes the throwaway proof permanent. Evidence-Ticket: OMN-12965 --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12987): workspace-build sibling lock-pin preflight + provenance (#1944) Recurrence ratchet for the 2026-06-11 stability bootstrap crash. The 11:20Z --no-cache rebuild vendored omnibase_infra 0.37.0-dev (~2c1d672f) + core 0.42.0 even though omnimarket dev's uv.lock pinned infra 0.38.1 @ e2dbdc95 + core 0.44.0 @ c97c2c9a. The stale sibling predated the OMN-12501 Protocol-quarantine guard, turning a latent contract defect into a fatal crash that crash-looped the main runtime. - check_sibling_lock_pins.py: host-side fail-fast preflight resolving expected sibling versions/SHAs from the consuming repo's (omnimarket) uv.lock and comparing them against each vendored tree. Mismatch aborts the build. - stage_workspace.sh: emit a .build-sha marker per staged sibling (rsync drops .git) so the preflight and provenance can identify the vendored commit. - deploy-runtime.sh: run the preflight after staging, before build; abort on mismatch. Write the comparison under sibling-repos/.sibling-lock-pins.json. - compute_workspace_provenance.py + Dockerfile.runtime: fold expected-vs-actual lock_pin_comparison into build-provenance.json for deploy verifiers. Recurrence-guard tests prove a stale infra 0.37.0 vs lock-pinned 0.38.1 fails the preflight and matched pins pass; deploy script wiring is asserted statically. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12988): pin stability runtime-worker to replicas 1 + census ratchet (#1942) The base docker-compose.infra.yml defaults runtime-worker to replicas 0 (${WORKER_REPLICAS:-0}). The stability-test lane's required state includes a running worker (4-container census: main, effects, worker, projection-api), but the override pinned it via an env-interpolation default (${STABILITY_TEST_WORKER_REPLICAS:-1}) — a silent-drop surface: a stray exported STABILITY_TEST_WORKER_REPLICAS=0 or removal of the :-1 fallback would scale the worker to 0 with zero signal on a plain compose up/recreate. Fix: pin docker-compose.stability-test.yml runtime-worker deploy.replicas to the literal 1 (no env indirection). Ratchet (recurrence guards, same PR): - scripts/deploy-agent/tests/unit/test_runtime_worker_census.py: assert runtime-worker stays in the deploy-agent RUNTIME-scope census so a missing worker (replicas 0 => absent from docker compose ps) is a deploy failure, not silence; assert the override pins a literal 1. - tests/integration/infra/test_stability_test_runtime_compose_render.py: assert the rendered stability worker resolves deploy.replicas == 1. - tests/unit/infra/test_stability_test_runtime_lane.py: update the existing pin assertion to the literal 1. Evidence-Ticket: OMN-12988 Config-drift family: OMN-12945 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12979): expire-bound topic completeness suppressions (#1940) Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12966): accept community-edition Infisical /api/status in provision readiness gate (#1939) The fresh-provision readiness gate in provision-infisical.py only accepted the enterprise {"status": "ok"} payload and rejected the community edition's {"message": "Ok"}, returning 1 before bootstrap could run. This blocked provisioning against the Infisical instance deployed on .201 (community edition). Route the gate through the existing _is_infisical_ready helper (single source of truth, already used by the already-provisioned path). Add TestMainFreshProvision- ReadinessGate covering community/enterprise/not-ready cases. Also adds docker/docker-compose.infisical-stability.yml: an ADDITIVE Infisical compose project for the stability-test lane (joins the existing network as external, reuses stability postgres/valkey, no lane mutation), since the lane overlays disable the in-lane Infisical service via *-disabled profile overrides. P1.2b-A: Infisical now reachable from the stability runtime/effect containers; a known secret (OMN_12966_PROBE) seeds and resolves end-to-end via the machine- identity universal-auth path, verified from inside the effects container. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12958): volume-config drift gate + runtime config provenance (#1941) * feat(OMN-12958): volume-config drift gate + runtime provenance Compute config provenance (path + sha256) for the runtime-rendered Bifrost delegation contract; the deployed volume copy survives rebuilds and silently diverges from packaged source (two competing authorities, OMN-12945). - runtime/config_provenance.py: ModelConfigProvenance + drift classification, sidecar JSON writer (read by sweep + proof packets) - runtime/health/health_config_provenance.py: drift -> degraded health - render entrypoint logs provenance line + writes sidecar on every boot - docs/runbooks/volume-config-drift-and-reseed.md: ledgered re-seed procedure - validation exemption for config_name (logical identifier, not entity ref) No live volume mutation: re-seed is an operator deploy step (deploy_pending). * test(OMN-12958): cover volume config drift reseed flow --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12957): wire runtime-profile validator + core registry parity guard (#1945) * feat(OMN-12957): wire runtime-profile validator + core registry parity guard - Remove the import-time RuntimeError drift raise in runtime_profile.py (a hard raise on core/infra version skew would crash the kernel at import); the parity invariant is enforced by test_profiles_match_core_registry instead. - Add tests: _PROFILES keys == omnibase_core REGISTERED_RUNTIME_PROFILES, and every CONSUMER_ATTACHED_RUNTIME_PROFILES profile loads as a real lane. - Wire omnibase_core.validation.validator_runtime_profiles as a pre-commit hook + validator-runtime-profiles.yml CI gate on infra contracts. - Freeze 19 pre-existing violators in validation/runtime_profiles_allowlist.yaml (discovered by repo-root walk; drain via OMN-12982). Blocks NEW orphans. Requires the omnibase_core pin to include OMN-12957's validator (new rules). Evidence-Ticket: OMN-12957 Evidence-Source: 5463fbaf819409d4fb7f491dd4f276f10d869eba * ci(OMN-12957): pass runtime profile allowlist to validator * test(OMN-12957): cover runtime profile registry parity * fix(OMN-12957): keep runtime profile allowlist under config * fix(OMN-12957): pin core runtime profile registry --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12973): widen migration-gate healthcheck start_period + ratchet (#1950) P2.8: classify the prod migration-gate UNHEALTHY-then-self-resolved incident. Classification: idle-one-shot-mis-modeled = NO. The migration-gate is a correctly-modeled long-running sentinel (entrypoint `while true; sleep 3600`) whose healthcheck continuously polls db_metadata.migrations_complete via check_migrations_complete.sh. The container flipped UNHEALTHY transiently because its healthcheck start_period (10s) was far shorter than the real cold-volume migration window (~116s: prod gate started 09:35:22, intelligence-migration finished 09:37:18). Past the 10s grace window the still-failing probe was reported UNHEALTHY until migrations completed, then self-resolved — no fault. Fix: raise migration-gate healthcheck start_period 10s -> 180s in both the authoritative catalog manifest (docker/catalog/services/migration-gate.yaml, flows into the generated compose) and the hand-maintained docker/docker-compose.infra.yml that deploy-runtime.sh applies to .201, so a still-applying gate stays in `health: starting` instead of flipping UNHEALTHY. Ratchet (enforcement, not detection): new ValidatorHealthcheckStartPeriod in omnibase_infra catalog (validator_healthcheck_start_period.py) asserts a 120s floor for migration-completion gates, wired into `onex validate runtime` (cmd_validate_runtime) AND backed by unit tests that gate every PR via pre-commit + CI. A migration-completion gate is identified by BOTH a healthcheck that polls migration completion AND a service_completed_successfully dependency, so ordinary app services (e.g. intelligence-api, an HTTP liveness probe with a legitimately short 40s start_period) are not swept into the floor. Prod probed read-only only; no prod mutation. Applies to live prod via the batched stability/prod rebuild (deploy_pending). Evidence-Ticket: OMN-12973 Evidence-Source: OCC#PENDING Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12971): deliver Vertex ADC bearer token to runtime effects container (secret-ref discipline) (#1949) * feat(OMN-12971): deliver Vertex ADC bearer token to runtime effects container (secret-ref discipline) Wire the Vertex $500-credit ADC path into the runtime, ADDITIVE next to the Gemini API-key path (provider-agnostic; neither provider removed or forced). runtime_policy.contract.yaml: add a secret-source mapping llm.vertex.access_token (source_type=file, /run/secrets/vertex_access_token) to the dev/stability-test/ judge profiles, alongside the existing llm.gemini.api_key env mapping. The ref name MUST match cloud-vertex-gemini.secret_ref in omnimarket bifrost_delegation.yaml. The resolved VALUE is a short-lived OAuth bearer token minted from ADC, refreshed by the operator; the token VALUE is never committed — only the ref name + in-container path. Add aiplatform.googleapis.com to the cloud host allowlist. docker-compose.infra.yml: bind the operator-supplied host token file read-only to /run/secrets/vertex_access_token on the main and effects runtimes (VERTEX_ACCESS_TOKEN_HOST_FILE, default /dev/null so lanes without Vertex still start; Gemini key path unaffected). Pass through BIFROST_VERTEX_GEMINI_ENDPOINT_URL (overlay supplies the complete Vertex OpenAI-compat URL) and GOOGLE_CLOUD_PROJECT/LOCATION (default empty). runtime-policy.env: regenerated from the contract via render_runtime_policy_env (test_runtime_policy_env_matches_contract_renderer proves contract<->env parity). test_runtime_policy_contract.py: update host-allowlist assertion for the additive Vertex host. Pre-existing platform-wide topic-parity-gate failures (25 unrelated topics) are identical with and without this change (proven by diff); that gate is pre-commit- only (not a CI merge gate) and this change adds zero new topic gaps, so that one hook is SKIP-ped. No deploy/receipt/merge gate is bypassed. * fix(OMN-12971): make Vertex runtime env contract-owned --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13008): automated .201 disk maintenance — worktree GC + docker GC + watermark alert ratchet (#1952) * feat(OMN-13008): automated .201 disk maintenance — worktree GC + docker GC + watermark alert ratchet Conservative, keep-list-driven disk maintenance to prevent the 2026-06-11 /data ~95% outage that killed all three lanes mid-demo: - scripts/worktree-gc.sh: drives the canonical omniclaude prune-worktrees.sh (merged+clean+pushed safety) on both Mac (merge-sweep tick) and .201 - scripts/disk-gc.sh + disk_gc_plan.py: conservative docker/builder/image GC. Pure, testable removal planner honoring a VERSIONED keep-list (deploy/disk-gc/keep-list.yaml): never reaps a kept repo, kept tag, in-use image, or anything younger than min_age_days; keeps N superseded generations. - scripts/disk-watermark-check.sh + disk_watermark_event.py: df watermark ratchet. >=85% emits a typed disk-watermark bus event (warning) that the sweep auto-ticket path turns into a Linear ticket; >=90% emits critical. Broker addr is fail-fast from KAFKA_BOOTSTRAP_SERVERS (no localhost default). - deploy/disk-gc/: systemd USER timer (onex-disk-gc.timer/.service, hourly) + install-disk-gc.sh. User units, NOT lane containers. - tests: 20 unit tests incl. GC plan-safety invariants + dry-run proof that default mode issues no destructive op (a wrong-delete GC is worse than none). Contract: contracts/OMN-13008.yaml * fix(OMN-13008): pass docker inventory to GC planner via stdin, not env (ARG_MAX) On a host with many docker images, passing the full image/ps inventory as env vars to disk_gc_plan.py exceeds ARG_MAX ('Argument list too long', exit 126), producing an empty plan. Write inventory to per-run scratch files (under the log dir, never /tmp; cleaned on exit) and hand it to the planner on stdin as a JSON envelope. Verified the failure live on .201; planner now reads stdin. * fix(OMN-13008): simplify GC plan stdin pipe (two processes, no nested subprocess) * fix(OMN-13008): keep-wins reconciliation — never remove an image id with any keep reason A single image id can surface in multiple 'docker image ls' rows (one per repo:tag). One tag could route the id to dangling-removal while another routes it to keep (e.g. tagged 'latest' or within-N-generations). Live .201 plan showed an id in BOTH remove_image_ids and kept_reasons. Reconcile at the end: keep wins — any id with a keep reason is dropped from the remove list; remove list deduped. Adds 2 regression tests. Verified live on .201. * fix(OMN-13008): timer uses OnCalendar=hourly for reliable re-arm (oneshot service) OnUnitActiveSec does not reliably re-elapse for a oneshot service once it goes inactive (observed NextElapseUSecMonotonic=infinity live on .201 after the first run). Switch to OnCalendar=hourly + RandomizedDelaySec so the timer re-arms every hour. Keep Persistent=true for missed-run catch-up. --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13005): materialize blocking event_consumer in runtime auto-wiring (runner consume-leg degenerate rows) (#1951) * fix(OMN-13005): materialize blocking event_consumer in auto-wiring (runner consume-leg degenerate rows) The runtime auto-wiring materialized event_publisher for handlers that declare it but had no equivalent for event_consumer. Request/response EFFECT handlers (HandlerContextRoiRunner) that publish a command then block on the correlated terminal event fell back to their no-op consumer default, returning None immediately -> every result row degenerate (failure_stage=generation, attempt_count=0) while generations succeeded ~1s later. Adds _make_sync_event_consumer (mirror of _make_sync_event_publisher), backed by service_terminal_event_consumer.make_terminal_event_consumer: a sync (topic, correlation_id, timeout) -> dict | None adapter that runs the proven direct-Kafka correlate-and-wait loop (from RuntimePatternBBroker) on an isolated event loop in a worker thread, so blocking does not deadlock the runtime dispatch loop that delivers the awaited terminal. TDD through the REAL dispatch path: test_event_consumer_injection drives a trial through _prepare_handler_wiring with a terminal arriving after a delay and asserts a non-degenerate row; verified RED with injection disabled. * fix(OMN-13012): two-phase (seek-now/wait-later) terminal event_consumer to close the subscribe-after-publish race (#1954) The OMN-13005 injected event_consumer is a single callable that does assign -> seek_to_end -> poll internally, all AFTER the handler has already published its command. Once OMN-13010 freed the dispatch loop and generation began completing in ~1s, the correlated terminal lands BEFORE the single-call consumer's post-publish seek_to_end positions, so seek_to_end skips PAST the already-emitted terminal and the runner times out on an offset beyond it (probe3, run_id=20260611T2140Z-probe3 -- both arms degenerate, zero rebalances). Splits positioning from waiting so the caller subscribes BEFORE it publishes: session = consumer.open(topic) # assign + seek_to_end NOW publisher(command_topic, payload) # publish AFTER positioning payload = session.wait(cid, timeout) # block from the captured position The returned TerminalEventConsumer is still directly callable with the legacy (topic, cid, timeout) -> dict | None single-call shape for any consumer that does not need subscribe-before-publish; the runner is the only consumer today. TerminalConsumerSession owns a dedicated event loop on a daemon worker thread for the whole open->wait->close lifecycle, preserving the OMN-13005 loop-isolation discipline so blocking never deadlocks the runtime dispatch loop. TDD (real dispatch path, RED-then-GREEN): extends the OMN-13005 injection test with a terminal emitted IMMEDIATELY after publish. The single-call (seek-after-publish) consumer MISSES it (degenerate row -- RED test asserts failure_stage=generation); the two-phase (open-before-publish) consumer CATCHES it (non-degenerate -- GREEN). The Kafka layer is faked at the two service seams against a shared in-memory log modeling seek-to-end semantics. OMN-13005 blocking-correlate behavior preserved. 269/269 auto_wiring unit tests pass; mypy --strict clean. Sibling to OMN-13010 / OMN-13005 / OMN-13003. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13005): route terminal consumer through Kafka boundary --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12990): pin worker replicas in ledgered config + fail-fast lane overrides (silent-drop ratchet) (#1948) * fix(OMN-12990): pin worker replicas in ledgered config + fail-fast lane overrides The base compose sets runtime-worker deploy replicas to ${WORKER_REPLICAS:-0} (soft-default ZERO). The stability lane's required state includes a running worker (GATE_ZERO_PROOF.md: 4 runtime containers), but the worker presence rode entirely on a compose soft-default — any plain compose up/recreate without the policy env silently scaled the worker to zero with no error and no signal. Fix (contract-native + fail-fast): - Add 'replicas' to ModelRuntimeProcessPolicy; pin replicas: 1 in every lane's worker block in runtime_policy.contract.yaml. - Renderer emits {PROFILE}_WORKER_REPLICAS into the ledgered runtime-policy.env for dev/stability-test/judge/prod. - stability + prod compose overrides reference ${..._WORKER_REPLICAS:?...} (fail-fast, NO silent :-1/:-0 default). A recreate that omits the policy env now aborts loudly instead of dropping the worker. prod previously had no override at all and inherited the dangerous :-0 default. Ratchet (recurrence guards): - tests asserting fail-fast override form (no soft default), contract-declared replica pin >= 1 per lane, and rendered {PROFILE}_WORKER_REPLICAS=1 in the ledgered env. - runbook deploy/verify procedure adds an expected-container census (worker must be present) via verify_container_manifest; a missing worker is a FAILURE, not silence. Note: SKIP=topic-parity-gate — that local-only advisory gate (absent from all .github/workflows, not a required CI check) fails on 25 pre-existing cross-repo topic gaps (build-loop/omniclaude/omniweb) identical on pristine base HEAD 8d7da1249; this change adds zero topics. All other hooks ran clean. Evidence-Ticket: OMN-12990 * test(OMN-12990): cover worker replica policy integration --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12909): add gateway bus forwarder P0A (#1946) * feat(OMN-12909): add gateway bus forwarder p0a * test(OMN-12909): add gateway forwarder integration coverage * fix(OMN-12909): satisfy gateway forwarder validators * test(OMN-12909): allow gateway forwarder bus protocol * fix(OMN-12909): sync gateway forwarder entry point * fix(OMN-12909): refresh runner image identity lock * test(OMN-12909): relax JSON normalizer mixed benchmark threshold * fix(OMN-12909): allow gateway handlers to boot unconfigured * fix(OMN-12909): declare gateway forwarder runtime profile * fix(OMN-12909): update runner identity backmerge expectation --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13011): LANE CENSUS RECONCILIATION ratchet — declared desired-state per lane, drift = auto-ticket (#1955) The class fix for the recurring lane-drift regression. Nothing reconciled the declared desired state of a runtime lane against what is actually running, so the same failure kept recurring with zero signal: volume config drift (OMN-12945), WORKER_REPLICAS silent zero (OMN-12988/12990), and on 2026-06-11 prod runtime containers plus the broker network were silently absent for hours during demo prep. Ships a per-lane DESIRED-STATE census: - (a) DECLARED in a versioned lane manifest (deploy/lane-census/lane-manifest.yaml): container set, network, replicas, image-tag pattern per lane (stability-test/prod/judge/dev), derived from the canonical compose lane files. A parity ratchet keeps the manifest locked in step with the compose files. - (b) RECONCILED on a schedule on .201 by SHARING the OMN-13008 systemd timer (a drop-in 4th ExecStart on onex-disk-gc.service — never a second timer) and on-demand via scripts/lane-census-check.sh / runtime_sweep. - (c) Drift = typed bus event (onex.evt.infra.lane-census-drift.v1) + Linear auto-ticket naming exactly what is missing/extra (container_absent, network_detached, replicas_zero, unexpected_container, oneshot_failed/stuck, image_tag_mismatch). Fail-fast, no warn-only mode (gates-block policy); exit 30 on drift; bus publish fail-fast on KAFKA_BOOTSTRAP_SERVERS (no localhost default). Red fixture reproduces 2026-06-11: prod runtime containers absent + broker network detached must produce the exact drift findings + a non-zero exit hours before a human noticed. Pure planner is fully unit-tested; shell driver dry-run-tested. Builds on the OMN-12988 deploy-agent RUNTIME census (deploy-time) as the complementary steady-state reconciler; closes the runtime-worker.yaml container_name: null census gap by sourcing names from the compose lane files. Evidence-Ticket: OMN-13011 Config-drift family: OMN-12945 Relates-to: OMN-13009, OMN-12988, OMN-13008 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13020): vendor missing node migrations — llm_routing 0000 + context_roi 001 (#1956) Vendors two omnimarket node-source migrations into the infra forward-migration tree via scripts/sync-node-migrations.sh (the canonical OMN-12559 mechanism): - node_projection_llm_routing/0000_create_llm_routing_decisions.sql (source: omnimarket #1168 / OMN-12942, merge ed6734f8) - node_projection_context_roi/001_create_context_roi_scores.sql (source: omnimarket #1178 / OMN-12955, merge 5010b1f4) Without the 0000 base table, node_projection_llm_routing/0001 (CREATE VIEW) hard-fails against NODE_POSTGRES_DB=omnidash_analytics — exactly the prod forward-migration exit-3 of 2026-06-11T09:35:52Z, and reproduced by construction in any clean clones@dev build. Files are byte-identical to the omnimarket dev blobs (sha256 f8a8b339… / c4126e65…) and to the untracked hot-patch copies on the .201 stability clone. Both migrations are self-contained, all-statements-IF-NOT-EXISTS, and 0000 sorts lexically before 0001 within the node's namespaced identity space. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13058): close TerminalConsumerSession on open() failure (worker thread + event loop leak) (#1957) TerminalConsumerSession.__init__ starts its dedicated worker loop thread immediately. TerminalEventConsumer.open did 'session = TerminalConsumerSession(...); return session.open()' with no cleanup: any failure inside session.open() (consumer start timeout, partition-assign timeout, broker auth error) propagated out of the raising expression, the session reference was lost, and the daemon worker thread plus its never-closed asyncio event loop leaked -- one pair per failed open. The motivating caller (HandlerContextRoiRunner) opens a session per trial, so a 160-560-trial battery against a degraded broker accumulates hundreds of leaked threads in the long-lived effects container. Fix: wrap session.open() in try/except BaseException -> session.close() (idempotent: stops the loop, joins the thread) -> re-raise. Covers both the two-phase .open(topic) path and the legacy single-call __call__ path. Found by the P3.3 doctrinal review of merged #1951 (b9712af9 / 36d98275). TDD: tests/unit/runtime/test_service_terminal_event_consumer_open_failure.py injects a real open failure through the production path (event bus without _bootstrap_servers) and asserts no alive terminal-consumer-* thread after the raise. Verified RED with the fix stashed (2 failed), GREEN with it (2 passed). Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13021): non-dev-base guard — fail feature-base PRs absent Stacked-Parent declaration (retro A-6) (#1958) Any PR whose base is neither dev nor main fails unless the body carries 'Stacked-Parent: #N'. Prevents the feedback_stacked_prs_orphan_from_dev class (#1185/#1954 auto-merged INTO parent feature branches and stranded off dev). base=main remains governed by main-target-guard. Epic OMN-13013 (process enforcement ratchets — June 12 retro). Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13014): hot-patch ledger rebuild preflight gate (retro B-1) (#1959) * feat(OMN-13014): hot-patch ledger rebuild preflight gate (retro B-1) Hot-patches on .201 (.prepatch sibling discipline) silently revert on any image rebuild/force-recreate — the 2026-06-11 20:58Z rebuild already erased a live /api/generate patch once. This adds the rebuild-path gate: - scripts/preflight_hotpatch_ledger.py: given a target container or lane + per-repo build refs, hard-fails when any hot-patch ledger row's source PR merge commit is not an ancestor of the build ref (git merge-base --is-ancestor), plus a .prepatch tripwire of the running container (unledgered .prepatch = hard fail; --post-rebuild = zero .prepatch expected). Sole bypass: HOTPATCH_PREFLIGHT_BYPASS carrying the Rule-10 '# skip-token-allowed: <user-approval-receipt-id>' form. - scripts/deploy-runtime.sh: guard_hotpatch_ledger wired into main() before build/preview (both dry-run and execute), lane derived from the compose project; skips loudly only when no ledger exists on the host. - tests/unit/scripts/test_preflight_hotpatch_ledger.py: 17 unit tests (ancestor gate, lane scoping, ledger loading, tripwire, bypass forms). - tests/ci/test_receipt_gate_install_guard.py: repair stale guard — core OMN-12565 replaced the OMN-9198 'uv pip uninstall first' install step with a cleared workspace venv (uv venv --clear); assert the new contract. Ledger backfilled from live census (5 .prepatch files / 4 source PRs, all MERGED to dev) at /data/omninode/hotpatch-ledger/ledger.yaml on .201. * fix(OMN-13014): scope missing-prepatch tripwire warning to the probed container --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12977): workspace-build sibling-pin ratchet — honor consuming lock, fail-fast on stale vendor (#1936) * fix(OMN-12977): workspace-build sibling-pin ratchet — honor consuming lock, fail-fast on stale vendor The workspace-mode image build (BUILD_SOURCE=workspace, used by the stability-test deploy procedure) vendored sibling/foundation packages from whatever the canonical OMNI_HOME clones happened to be checked out at, ignoring the consuming repo's uv.lock. On 2026-06-11 this shipped a 13-day-stale omnibase_infra 0.37.0-dev (pre-OMN-12501 Protocol-quarantine guard) + core 0.42.0 against an omnimarket dev lock pinning infra 0.38.1@e2dbdc95 / core 0.44.0@c97c2c9a, dropping the guard and crashing wire_from_manifest bootstrap fatally (stability lane down on demo day). Fix + recurrence ratchet (same PR): - scripts/runtime_build/check_sibling_lock_pins.py: parse the consuming repo's uv.lock for expected version+git-rev of each foundation/sibling package (scoped to the package's own source line so editable/registry pins are not cross-attributed a dependency's rev), resolve the actual clone version+HEAD, compare, and classify drift backward/forward/none. Fail-fast (exit 1) on any drift; --allow-drift records an explicit operator override in the artifact, never silent. - stage_workspace.sh: runs the preflight against the canonical clones before staging; aborts the build (exit 3) on unacknowledged drift and writes workspace/sibling-pin-comparison.json. - compute_workspace_provenance.py: folds the expected-vs-actual comparison into build-provenance.json so deploy verifiers can assert the build honored the lock; flags unacknowledged drift as a provenance error. - Dockerfile.runtime: COPY the comparison artifact (committed placeholder so the COPY always resolves; overwritten by stage_workspace.sh in workspace mode). - TDD: 19 unit tests covering lock parsing (git/registry/editable sources), drift classification, the exact 0.37.0-vs-0.38.1 stale case, check_pins exit codes, and the allow-drift override. Pre-existing mypy-strict bare-dict errors in compute_workspace_provenance.py fixed in the same pass. Evidence-Ticket: OMN-12977 Evidence-Source: pending-occ * ci(OMN-12977): retry runtime smoke compose port race * test(OMN-12977): align sibling-pin script tests with current API --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12989): workspace-mode image build must honor sibling lock pins (#1947) * fix(OMN-12989): workspace build must honor sibling lock pins (fail-fast ratchet) The 2026-06-11 stability bootstrap crash was caused by a workspace-mode --no-cache rebuild that vendored omnibase_infra 0.37.0-dev (13-day-stale worktree) even though omnimarket dev's uv.lock pins infra 0.38.1. The downgraded sibling predated the OMN-12501 Protocol-quarantine guard and turned a latent contract defect into a fatal crash. Fix + ratchet (same PR set): - scripts/runtime_build/resolve_workspace_pins.py: parse the consuming repo's uv.lock for sibling pins (version + git rev); classify each staged/installed sibling as exact/ahead/regression/unpinned; fail-fast (WorkspacePinError) on any regression BELOW the lock pin. Stdlib version-tuple fallback when packaging is absent so the ratchet never fails-open. - compute_workspace_provenance.py: enforce sibling pins + a host-infra self-check (installed omnibase_infra vs lock pin — the exact crash vector, since host infra is built from the context, not staged), and emit a pin_comparison block into build-provenance.json for deploy verifiers. - Dockerfile.runtime: COPY resolve_workspace_pins.py beside the provenance script so the in-image import resolves. - TDD: failing tests first (test_resolve_workspace_pins.py, 11 cases) + provenance integration tests; fixed a pre-existing stale RUNTIME_VERSION assertion to read pyproject dynamically. Evidence-Ticket: OMN-12989 * fix(OMN-12989): workspace build must honor sibling lock pins (fail-fast ratchet) The 2026-06-11 stability bootstrap crash was caused by a workspace-mode --no-cache rebuild that vendored omnibase_infra 0.37.0-dev (13-day-stale worktree) even though omnimarket dev's uv.lock pins infra 0.38.1. The downgraded sibling predated the OMN-12501 Protocol-quarantine guard and turned a latent contract defect into a fatal crash. Fix + ratchet (same PR set): - scripts/runtime_build/resolve_workspace_pins.py: parse the consuming repo's uv.lock for sibling pins (version + git rev); classify each staged/installed sibling as exact/ahead/regression/unpinned; fail-fast (WorkspacePinError) on any regression BELOW the lock pin. Stdlib version-tuple fallback when packaging is absent so the ratchet never fails-open. - compute_workspace_provenance.py: enforce sibling pins + a host-infra self-check (installed omnibase_infra vs lock pin — the exact crash vector, since host infra is built from the context, not staged), and emit a pin_comparison block into build-provenance.json for deploy verifiers. - Dockerfile.runtime: COPY resolve_workspace_pins.py beside the provenance script so the in-image import resolves. - TDD: failing tests first (test_resolve_workspace_pins.py, 11 cases) + provenance integration tests; fixed a pre-existing stale RUNTIME_VERSION assertion to read pyproject dynamically. Evidence-Ticket: OMN-12989 * test(OMN-12989): co-locate provenance pin helper in fixture --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13055): absent repos in REPOS list warn and exit 0 instead of failing (#1960) Missing repos (not cloned locally) now emit a WARN line and are tracked in a separate WARNED array. Only real fetch/ff failures cause exit 1. This makes pull-all.sh safe to use on machines with a partial clone set, while keeping the explicit-list override behavior intact. Adds three regression tests: absent-only exits 0, absent+present exits 0 with OK for the present repo, present-failed+absent exits 1. Evidence-Ticket: OMN-13055 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13070): config_prefetcher overlay wins on controlled lanes (#1961) On infisical_required=True lanes, fetched/overlay config now always wins over ambient env. Ambient env is retained only as a declared bootstrap fallback (with an explicit provenance INFO log line) when Infisical returns None. apply_to_environment also overwrites stale env on controlled lanes. Uncontrolled lane (infisical_required=False) behaviour is unchanged. Adds 5 regression tests: controlled-lane Infisical-wins, env-bootstrap- fallback, apply_to_environment overwrite, missing-from-both-is-error, and uncontrolled-lane-env-still-wins. Refactors _resolve_key to return a (outcome, value, error) tuple to satisfy the ≤5-param pattern gate. Source: docs/audits/2026-06-10-runtime-env-overlay-authority-audit.md Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13062): migration-gate vacuity fix — sentinel discipline, wait-for-postgres, skip-manifest, sync-check nonzero (retro A-10) (#1963) * fix(OMN-13062): migration-gate vacuity fix — sentinel discipline, wait-for-postgres, skip-manifest, sync-check nonzero Fixes three bugs identified in retro A-10 (recurrences OMN-12885, OMN-12934): (1) RUNNER SENTINEL DISCIPLINE run-forward-migrations.sh now clears migrations_complete=FALSE at the start of every run and sets it TRUE only as its FINAL act after all infra and node migrations succeed. Any mid-run failure leaves the gate UNHEALTHY. runner_completed_at is stamped at the same final step as durable evidence of a successful completion run. (2) SYNC-NODE-MIGRATIONS VACUOUS GATE sync-node-migrations.sh --check now exits 2 (not 0) when the omnimarket source tree is unresolvable. Silent exit-0 was hiding drift. The single opt-out is SYNC_NODE_MIGRATIONS_SKIP_UNRESOLVABLE=1 for environments that intentionally run without the source. (3) WAIT-FOR-POSTGRES GUARD run-forward-migrations.sh now waits up to PG_WAIT_RETRIES (default 30) x 2s for Postgres to accept connections before proceeding, guarding the first-boot initdb race. (4) SKIP-MANIFEST docker/migrations/skip-manifest.yaml introduced as the sole committed escape for intentionally-skipped migrations. The runner reads this at startup; listed migrations are recorded in schema_migrations with checksum "skip-manifest" without executing the SQL. (5) MIGRATION 085 Adds runner_completed_at TIMESTAMPTZ column to db_metadata so the runner's final stamp is durable in the schema (idempotent ADD COLUMN IF NOT EXISTS). Rollback included. 22 regression tests added covering all five fix surfaces. * fix(OMN-13062): stamp schema fingerprint for migration 085 Migration 085 (085_add_runner_completed_at_to_db_metadata.sql) was added in the initial commit but schema_fingerprint.sha256 was not regenerated. Running `python scripts/check_schema_fingerprint.py stamp` updates the artifact from the stale hash to match the 71 migration files. Evidence-Source: OCC#2563 Evidence-Ticket: OMN-13062 --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12864): Bifrost endpoints → committed overlay authority + fail-loud loader (#1964) * feat(OMN-12864, OMN-12814, OMN-12945): Bifrost endpoints → committed overlay authority + fail-loud loader OMN-12864 — Committed lane overlay - docker/lane-overlays/dev.bifrost.yaml: typed deployment bindings for all four BIFROST_LOCAL_*_ENDPOINT_URL values (coder :8000, reasoner :8001, embedding :8100, ds4-flash :8101). Previously only available as ephemeral shell exports on .201; now committed, auditable, diff-able, and CI-checked. - docker/lane-overlays/dev.bifrost.env: generated dotenv sidecar consumed by compose via env_file; never edited directly (yaml is authority). - docker/docker-compose.infra.yml: wire the env_file block at compose root so the four endpoints are injected into the interpolation context on a clean shell. Hardcode BIFROST_CONTRACT_PATH (remove :-/empty footgun — OMN-12814). - scripts/render_bifrost_lane_overlay_env.py: render script regenerates the env sidecar from the YAML source. - src/omnibase_infra/runtime/models/model_bifrost_lane_overlay.py: ModelBifrostLaneOverlay — typed Pydantic model enforcing URL completeness (OMN-12815: every URL must end in /chat/completions). OMN-12814 — Fail-loud loader - render_bifrost_delegation_contract: raises ProtocolConfigurationError on FileNotFoundError, YAMLError, ValidationError, and zero-endpoint renders. No lru_cache — every restart re-renders from packaged source so a stale cache cannot pin a broken result across deploys. OMN-12945 — Re-seed from packaged source on deploy - docker/entrypoint-runtime.sh: set BIFROST_FORCE_RESEED=1 on every container restart so the named-volume copy is always rebuilt from the packaged bifrost_delegation.yaml merged with committed lane-overlay endpoints. - render_bifrost_delegation_contract: honor BIFROST_FORCE_RESEED/force_reseed flag to bypass the stale-volume early-return path entirely. Tests: - tests/ci/test_bifrost_lane_overlay.py: CI gate — env sidecar in-sync with YAML source; all four BIFROST_LOCAL_* keys present. - tests/unit/runtime/models/test_model_bifrost_lane_overlay.py: bare-base URL rejection, env dict mapping, extra-field rejection. - tests/unit/runtime/test_render_bifrost_delegation_contract.py: fail-loud paths, force-reseed, zero-endpoint error, endpoint URL completeness. - tests/unit/models/test_model_serialization_roundtrip.py: roundtrip coverage. * fix(OMN-12864): move bifrost env_file to service level — fix compose schema validation failure Top-level 'env_file' is rejected by Docker Compose v2 schema validator ('additional properties not allowed'). This caused 10+ compose-render integration tests to fail in CI. Fix: - Remove top-level env_file block from docker-compose.infra.yml - Add per-service env_file on omninode-runtime, runtime-effects, runtime-worker (the three containers that render Bifrost) - Change BIFROST_LOCAL_*:? to BIFROST_LOCAL_*:- in x-runtime-env (compose-level validation removed; Python validates via ModelBifrostLaneOverlay + render_bifrost_delegation_contract) - Add two CI gate tests: compose_env_file_is_service_level_not_top_level and runtime_services_have_bifrost_env_file * fix(OMN-12864): allow BIFROST_LOCAL_* empty defaults in silent-fallback gate The overlay authority pattern (OMN-12864) passes BIFROST_LOCAL_*_ENDPOINT_URL via service-level env_file (docker/lane-overlays/dev.bifrost.env), not at compose config time. Compose-level :? would break CI rendering without the overlay pre-loaded. Validation at the Python layer (ModelBifrostLaneOverlay + render_bifrost_delegation_contract) is the enforcement point. Add the three failing vars to ALLOWED_EMPTY_DEFAULTS with OMN-12864 citation. --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-12858/OMN-12879/OMN-12880): dispatcher route-coverage CI gate (#1965) * feat(OMN-12858/OMN-12879/OMN-12880): dispatcher route-coverage CI gate Static contract analysis: every subscribed command topic (onex.cmd.*) must declare handler_routing or runtime_dispatch, or the message goes to DLQ silently. This gate would have caught two recent incidents: 1. June 9 DLQ regression (OMN-12858 post-mortem): node_generation_consumer subscribed onex.cmd.omnimarket.node-generation-requested.v1 but a sole-handler revert left zero dispatcher routes registered. Messages went to DLQ silently with no CI signal. 2. June 12 DEL-01 live finding: onex.cmd.omnimarket.delegate-skill.v1 was consumed by dev lane bus but no dispatcher route existed in any deployed contract. Discovered via manual rpk consumer-group lag probe (DEL-01 evidence, docs/evidence/2026-06-12-weekend-pass/). Deliverables: - scripts/check_dispatcher_route_coverage.py — static YAML scanner that checks both omnibase_infra and omnimarket contract trees; ratchet allowlist for known pre-existing violations; --changed-contracts mode (OMN-12879) for per-PR scoping; compat publish topics excluded (OMN-12880) - .github/workflows/dispatcher-route-coverage.yml — CI workflow that checks out omnimarket sibling, collects changed contract paths in PR mode, and runs the gate; fires on PR, push-to-main, and merge_group - tests/ci/test_dispatcher_route_coverage_gate.py — 12 unit tests covering RED/GREEN/COMPAT/CHANGED-MODE/ALLOWLIST/MULTI-DIR paths plus live-contract regression proof against the actual omnibase_infra tree Allowlist additions: - onex.cmd.omnibase-infra.pattern-b-dispatch.v1 (RuntimePatternBBroker, imperative consumer, OMN-12525 migration target) - onex.cmd.platform.contract-resolve-requested.v1 (transitional HTTP bridge node_contract_resolver_bridge OMN-2756, metadata.transitional=true) [OMN-12858, OMN-12879, OMN-12880] * fix(OMN-12858): drop full uv sync from dispatcher-route-coverage workflow Gate script only needs pyyaml (stdlib + yaml). Using full setup-python-uv was causing 10+ minute timeout. Replace with direct pip install pyyaml and invoke python3 directly. Reduces job from 10m timeout to <1m. [OMN-12858] --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13026): wire transport-mock-lint pre-commit hook + CI gate (#1962) * feat(OMN-13026): wire transport-mock-lint pre-commit hook + CI gate (omnibase_infra) Activates the transport-mock-lint validator (from omnibase_core, OMN-13026) on omnibase_infra. Ratchet baseline: 218 existing violations across 63 files frozen in validation/transport_mock_baseline.yaml. New bare AsyncMock/ MagicMock on EventBus/transport surfaces are blocked by pre-commit hook and CI lint step. Existing violations tracked for drain by per-site tickets (parent OMN-13026). Reference incident: PR #1181 bare AsyncMock hid missing EventBusKafka.stop(). Evidence-Ticket: OMN-13026 * fix(OMN-13026): use uv run python for CI lint step + bump omnibase_core pin to include transport_mock_lint The transport-mock lint CI step previously cloned omnibase_core and ran `python -m omnibase_core.validators.transport_mock_lint` with PYTHONPATH, but this failed: `No module named omnibase_core.validators.transport_mock_lint` because it ran `.venv/bin/python` which uses the locked venv, and the venv omnibase_core pin (2defabef4) predates the transport_mock_lint module. Fix: use `uv run python` (removes the clone step) and bump omnibase-core git pin from 2defabef4 to 309d89fa7 (PR 1231 merge commit on dev) so transport_mock_lint is available in the locked venv. * fix(OMN-13026): align transport mock baseline and runner lock * fix(OMN-13026): sync omnibase_core pin + runner identity lock to dev baseline Align pyproject.toml omnibase_core rev to 2defabef (required by test_release_backmerge_preserves_proven_runtime_core_pin) and update docker/runners/runner-image.lock.json identity_digest/shared_env_digest to match dev runner image lock (79b08f44 / 90c8b3b9). Both were stale from the prior session's pin bump that used an older SHA. * fix(OMN-13026): source transport mock validator from core * fix(OMN-13026): source transport validator from core dev --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13094): --output receipt mode on onex node/run — quiet typed receipts with durable capture (#1966) Phase 2a of the skill-output-suppression slice (epic OMN-13089, plan docs/plans/2026-06-12-skill-output-suppression-plan.md Phase 2 item 1). - onex node/onex run gain --output receipt: ALL runtime logging routes to a run_id-suffixed capture file under <state-root>/captures/ (no console handlers — kills the 25-50-line RuntimeLocal INFO stream at the source); stdout carries exactly ONE typed ModelSkillResult JSON with the FULL handler result (result_model = concrete handler result type FQN). - Durable capture: capture log + handler result content-addressed via omnibase_core ArtifactStore (OMN-13093); artifact.captured + tool.output.captured emitted to the emit daemon socket (--emit-socket, default ~/.claude/emit.sock). - Failure asymmetry: artifact write failure => FULL output printed, no receipt (no hidden loss); emission failure => receipt still prints, event spooled to <state-root>/emit_spool/ for replay. - Node failure => status=failed/error with full error + capture log INLINE in the receipt (errors are never hidden) and artifact-backed. - Default output mode unchanged (enforcement is Phase 4). - RuntimeLocal exposes handler_result (receipt schema identity). - core pin 2defabef -> ae8793bd (merged OMN-13091/13093 receipt models + ArtifactStore); runner-image identity lock regenerated and the OMN-12765 backmerge identity constants updated for the new pin. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13097): onex skill subcommand + declarative skill->node mapping (#1968) * feat(OMN-13097): onex skill subcommand + declarative skill->node mapping Phase 4a of the skill-output-suppression slice (epic OMN-13089, plan docs/plans/2026-06-12-skill-output-suppression-plan.md Phase 4 item 1/2). A dispatch skill IS one CLI call (user directive 2026-06-12). This adds the `onex skill <name> [args]` dispatch surface that the 24 omniclaude shim migrations build on: - onex.cli entry-point `skill` -> cli_skill.run_skill_by_name. Resolves the skill via the declarative skill_mapping.yaml registry, builds the backing node's input payload from the skill's CLI args, writes it under the state root (.onex_state/tmp/<skill>-<run_id>.json — never /tmp), resolves the node's packaged contract exactly like `onex node`, and dispatches through the proven receipt-mode path (run_receipt_mode, OMN-13094). stdout is exactly one typed ModelSkillResult JSON with the FULL handler result. - skill_mapping.yaml: declarative DATA mapping all 24 dispatch shims to their backing onex.nodes node + typed result-model FQN (verified against each live handler handle() return type on origin/dev) + per-arg payload specs + static payload + keyword classifiers (delegate task_type as data, not code). Adding a skill is a YAML edit + fixture, never a CLI code change (ticket deliverable 2/3). Mapping lives beside the node-resolution surface, never hardcoded branching in the CLI. - Typed models split one-per-file (repo convention): ModelSkillArgSpec, ModelSkillClassifier, ModelSkillMapping, ModelSkillMappingRegistry, EnumSkillArgType. Frozen, extra=forbid, fail-fast coercion/validation. - validation_exemptions.yaml: Click-callback param-count + literal-identifier name-field exemptions mirroring the existing cli_node run_node_by_name precedent (OMN-11570) — same pattern, same rationale. dod_evidence: - 20 unit tests pass (registry validity, all-24-shims coverage, FQN result models, arg parsing/coercion/positional/required, classifiers, payload build, receipt-mode dispatch wiring, payload-under-state-root not /tmp). - uv run mypy src/ --strict: clean (2438 source files). - ruff format + check: clean. pre-commit run on changed files: pass. - runner-image identity lock regenerated for the pyproject entry-point add (same as OMN-13094). * test(OMN-13097): rebind OMN-12765 backmerge identity constants for onex skill pyproject change Adding the `skill` onex.cli entry-point to pyproject.toml changes the runner-image identity_digest (and shared_env_digest) the lock binds. Update the hardcoded expected constants in the backmerge-identity test to the regenerated values — same mechanical rebind OMN-13094 performed for the core pin bump. Identity + runner-image-identity tests pass (14). --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix+test(OMN-13012): force terminal-topic metadata refresh so both ephemeral consumers assign (runner consume-leg wedge) (#1969) The runner consume-leg wedged on the live stability battery (image c0505521f1fa, EXP1-3_RUNNER_CONSUME_LEG_BLOCKER): only the FAILED terminal topic ever subscribed while the COMPLETED topic never assigned, so the correlated completed terminal was never read and the 8x2x10 matrix re-fired cell 1 forever, emitting zero non-degenerate rows. A prior two-strike diagnosis proved the omnimarket handler is correct (it opens both terminal sessions pre-publish and races them); the defect is in the omnibase_infra runtime consume leg. Root cause: _assign_direct_terminal_partitions ignored the metadata future returned by AIOKafkaClient.set_topics and re-called set_topics([same_topic]) each loop iteration. aiokafka 0.13.0 set_topics only forces a metadata refresh when the topic set DIFFERS from the tracked set, so every iteration after the first took the no-op branch and never re-fetched. An ephemeral group_id=None consumer whose first metadata fetch had not yet surfaced partitions burned the full 30s assign cap and raised a bare TimeoutError (the empty-message 'wait failed' seen live). Fix: register the reply topic once and await that metadata fetch, then on each miss force a fresh fetch via force_metadata_update (which always fetches) rather than the no-op set_topics repeat. The assign-cap TimeoutError now carries a diagnostic message instead of an empty one. Test: tests/integration/test_terminal_consumer_concurrent_assign_race.py drives the REAL TerminalEventConsumer (the object wired as event_consumer) through the REAL open_direct_terminal_consumer/poll path with AIOKafkaConsumer monkeypatched to a fake that faithfully models aiokafka 0.13.0 set_topics future + metadata-latency semantics. RED before the fix (bare TimeoutError, the live empty-message signature); GREEN after. K>=2 multi-trial variant asserts no worker-thread leak across trials. Evidence-Source: <occ-sha-pending> Evidence-Ticket: OMN-13012 Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-13096): onex delegate single-command subcommand (Phase 2b) (#1967) * feat(OMN-13096): onex delegate single-command subcommand Add 'onex delegate "<prompt>" [--task-type X] [--max-tokens N]' as a subcommand on the existing onex CLI (Phase 2b of the skill-output-suppression slice, OMN-13089). The command wraps payload construction, node dispatch, and result extraction internally and prints exactly one ModelSkillResult[ModelDelegateSkillResponse] to stdout via the OMN-13094 receipt-mode path. RuntimeLocal logs go to the capture file + artifact store, never to stdout; scratch payloads live under <state-root>/tmp/ with run_id suffixes (never /tmp). - cli_delegate.py: classify_task_type (keyword table from legacy skill md), payload write, contract resolve, run_receipt_mode dispatch - register 'delegate' under onex.cli entry points - exempt delegate_command from the >5-param patterns gate (same Click-callback rationale as run_node_by_name) - 20 unit tests: classification, scratch-under-state-root, single typed receipt on stdout, zero INFO log leakage omnibase_infra does NOT depend on omnimarket; the delegate node is resolved at runtime via the onex.nodes entry-point group (registered by omnimarket). * chore(OMN-13096): re-trigger deploy-gate after Evidence-Source set to OCC#2593 No code change — the deploy-gate workflow triggers on synchronize (not edited), so the PR-body Evidence-Source fix needs a new commit to re-resolve the OCC ref to the open PR head where contracts/OMN-13096.yaml (with deploy evidence) lives. * chore(OMN-13096): re-trigger deploy-gate now that OCC#2593 merged to OCC dev contracts/OMN-13096.yaml (with the dod-deploy-onex-delegate item) is now on OCC dev, so the deploy-gate OCC-dev checkout resolves the contract + deploy evidence. * chore(OMN-13096): regenerate runner-image identity lock for pyproject entry-point add Adding the 'delegate' onex.cli entry point to pyproject.toml changed the dependency-manifest digest that scripts/ci/runner_image_identity.py folds into the runner-image identity lock. Regenerate the lock so tests/ci/test_runner_image_identity.py matches (was the only CI test failure; unrelated environmental integration/perf failures excluded). * test(OMN-13096): update backmerge identity assertions to regenerated lock digests The runner-image identity lock was regenerated for the pyproject entry-point add; this test hardcodes the expected identity_digest/shared_env_digest, so update both to match the new lock (same maintenance OMN-13094 did). --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): tolerate partition-less reply topic in terminal consume leg (#1970) The context-ROI runner opens one ephemeral group_id=None terminal consumer per terminal topic BEFORE publishing each generation command (subscribe-before-publish, OMN-13012/13038). The FAILED reply topic is only produced to on contract_passed=False; in a battery where generations pass it has zero messages, so Redpanda never advertises a partition for it. _assign_direct_terminal_partitions burned the full 30s assign cap on every trial then raised a bare TimeoutError (the empty-message 'wait failed'), stalling each of the 160 battery trials ~30s before the COMPLETED terminal could correlate -> battery needs >80 min and never completes (verifier-confirmed wedge). A partition-less reply topic is a valid steady state, not a 30s error: - _assign_direct_terminal_partitions gives a bounded grace window for a topic that exists but is slow to surface metadata, then assigns whatever partitions exist (possibly none) and returns promptly instead of burning the cap and raising. - poll_direct_terminal_consumer treats an empty assignment as 'no terminal will arrive here' (sleeps out its timeout, returns None) without calling getone() on an unassigned consumer. Repro: tests/integration/test_terminal_consumer_battery_load_wedge.py drives the REAL TerminalEventConsumer over K=10 x 2 cells x 2 arms; RED (40/40 trials block a full assign cap) before the fix, GREEN after. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): pin terminal-consumer read offset synchronously (close lazy seek_to_end publish race) (#1971) * fix(OMN-13118): pin terminal-consumer read offset synchronously to close lazy seek_to_end publish race The consume-leg wedge survived PR #1969 (set_topics no-op) and PR #1970 (partition-less assign-cap) because both addressed the assign phase, not the seek timing. AIOKafkaConsumer.seek_to_end is LAZY: it requests a LATEST offset reset that only resolves on the first poll — AFTER the caller publishes. With generation completing in ~1s, the correlated COMPLETED terminal lands in the open->poll gap, so the lazily-resolved LATEST position is the HWM AFTER the record and the poll reads past it. The terminal is never read, the trial never correlates, and the experiment matrix re-fires the same cell forever. Replace seek_to_end with a synchronous end_offsets() + seek() pin (_pin_direct_terminal_end_offsets) in both open_direct_terminal_consumer and RuntimePatternBBroker._dispatch_and_wait_with_direct_kafka_consumer, so the read position is fixed at open() time, before the publish — the real subscribe-before-publish guarantee. Empty assignment (partition-less reply topic, OMN-13118 #1970) is a no-op. Repro: tests/integration/test_terminal_consumer_seek_reset_race.py drives the real TerminalEventConsumer.open()/wait() the way HandlerContextRoiRunner does, publishing the correlated terminal in the open->wait gap across K=10 x 2 cells x 2 arms; RED with the lazy reset (every cell degenerate), GREEN once the read offset is pinned. RED verified by git-stashing only the source fix. Existing consume-leg fakes updated to model end_offsets/seek (they previously masked the bug by making seek_to_end a synchronous exact snapshot). * test(OMN-13118): reword assertion (lazy not deferred) for receipt honesty gate * fix(OMN-13118): bound end_offsets() round-trip with assign-cap timeout (CodeRabbit) end_offsets() is a broker ListOffsets round-trip aiokafka documents as able to block indefinitely. Bound it with the same cap as start()/assign so a stalled broker fails fast instead of hanging the pre-publish positioning. --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13137): _validate_routing strategy-aware event_model checks (#1973) operation_match routes by the `operation` field and does not use event_model. The validator was unconditionally requiring event_model.{name,module} for every handler entry, causing 230/295 omnimarket operation_match contracts (all correct as authored) to fail routing validation at startup. Fix: read routing_strategy from the routing map and branch validation: - payload_type_match → require event_model.{name, module} (unchanged) - operation_match (and any non-payload strategy) → require `operation`; skip event_model checks entirely Updated pre-existing _validate_routing tests to declare routing_strategy: payload_type_match explicitly (they always tested payload_type_match semantics but relied on the implicit fallback that is now removed). Added test_validate_routing_operation_match.py with 4 unit tests: 1. operation_match without event_model → zero event_model errors 2. operation_match missing operation field → error 3. payload_type_match missing event_model → still errors (regression guard) 4. Real node_integration_sweep_orchestrator routing block → clean (boot gate) Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): independent per-terminal-topic consumers in Pattern B direct-Kafka wait (#1972) The consume-leg wedge survived four merged fixes (#1969 set_topics no-op, #1970 partition-less assign-cap, #1971 synchronous seek-pin). The STRONG K>=10 multi-cell reprobe on the stability lane still wedged on REBUILD-5 (cdf53d963f7b). Converged diagnosis (strikes 3+4, docs/evidence/2026-06-12-weekend-pass/experiments/probe4-stability/ reprobe-K10-rebuild5/HALT_K10_WEDGE_PERSISTS.md): the runtime waited for each trial's terminal across TWO topics (node-generation-completed.v1 + node-generation-failed.v1) with a SINGLE ephemeral group_id=None consumer assigned both topics' partitions. One aiokafka consumer holds one manual subscription; the COMPLETED delivery window collapsed before it surfaced the correlated record, so the trial never correlated and the matrix re-fired cell 1. RuntimePatternBBroker._dispatch_and_wait_with_direct_kafka_consumer now opens ONE independent AIOKafkaConsumer PER terminal topic via open_direct_terminal_consumer (each started, assigned, and offset-pinned via end_offsets()+seek() at open() BEFORE publish), then awaits both CONCURRENTLY via asyncio.wait(FIRST_COMPLETED). The first correlated terminal wins; both are torn down. No shared consumer, no subscription flip. Keeps the #1970 partition-less no-op and the #1971 synchronous offset pin (both live in open_direct_terminal_consumer / poll_direct_terminal_consumer). Removes the now-dead single-consumer helpers (_assign_terminal_topic_partitions, method-level _refresh_terminal_topic_metadata, _direct_kafka_* kwargs builders, _kafka_bootstrap_servers/_kafka_event_bus). Adds tests/integration/test_terminal_consumer_subscription_flip_wedge.py: a real-dispatch-path K>=10 x 2-cell x 2-arm repro whose fake models TWO independent consumers honestly (delivery is faithful only for a single-topic assignment; a consumer spanning both topics flips and drops the COMPLETED record). RED genuineness verified by reverting only the source to the single-consumer shape (test hangs past timeout); GREEN with the fix in 2.3s. Acceptance is the LIVE K>=10 multi-cell stability-lane reprobe (later phase), NOT this unit test. A green unit repro is necessary but NOT sufficient. Refs OMN-13118, OMN-13128. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-13118): long-lived terminal correlator replaces per-trial ephemeral consume leg (#1974) Tier B canonical redesign (epic OMN-12525). Five offset/subscription patches (#1969-#1972) tuned the per-trial-ephemeral terminal consumer and all failed…
* fix(OMN-15115): raise ModelLlmInferenceRequest timeout ceiling + per-model max_retries (#2442) * fix(OMN-15115): raise ModelLlmInferenceRequest timeout ceiling + add per-model max_retries qwen3-review-b's hostile-review timeout was pinned at this model's previous le=600.0 ceiling -- the max the schema allowed. That made the OMN-14176 config-only fix (raising timeout_seconds to 600) structurally incapable of ever curing the real defect: live-measured throughput on that endpoint is 4.3-4.6 tok/s (about half the ~9 tok/s the 600s value assumed), so a genuine full-length completion could not finish inside 600s regardless of contention. - Raise timeout_seconds ceiling 600.0 -> 1800.0 on ModelLlmInferenceRequest (nodes/node_llm_inference_effect -- the model HandlerLlmOpenaiCompatible actually consumes; a differently-implemented, same-named model at omnibase_infra.models.llm is unrelated, serves the CLI-subprocess handler path, and already has its own max_retries field). - Add a new max_retries field (default 3, matching the transport's historical hardcoded behavior) and thread it through HandlerLlmOpenaiCompatible._execute_with_auth -> both _execute_llm_http_call call sites, so a caller can lower retries for a systematically-slow (not transiently-flaky) endpoint without wasting a shared single-concurrency-slot's time on doomed retries. - RED/GREEN proven via git stash: all 8 new/changed assertions fail (AttributeError / extra_forbidden / less_than_equal) against the pre-fix code, pass after. OMN-15115 * fix(OMN-15115): sync llm inference contract timeout * test(OMN-15115): sync llm inference contract runtime expectations --------- Co-authored-by: Jonah Gray <jonah.g.gray@gmail.com> (cherry picked from commit 70857e7) * fix(OMN-14397): remove stale auto_merge skill mapping * chore(OMN-15115): retrigger deploy gate after reusable workflow fix --------- Co-authored-by: Jonah Gray <jonah.g.gray@gmail.com>
Refs OMN-15147
…me merge-commit exception; heals squash-severed merge-base (OMN-15195)
…2480 Empty commit — no tree change. #2478/#2480 shared the prior commit SHA (3b67018) with this PR's now-closed duplicate #2480, which made GitHub's commit->PR resolution in occ-preflight/receipt-gate ambiguous (kept resolving to the closed #2480). This commit gives #2479 a unique head SHA while leaving origin/dev's tree untouched (git diff --quiet against the parent confirms zero diff).
…idge-v2 promote(OMN-15181): dev->main dev-wins bridge — one-time merge-commit exception (OMN-15195)
…-0384-promotion # Conflicts: # src/omnibase_infra/nodes/node_runner_fleet_health_compute/handlers/handler_runner_fleet_health_evaluate.py # src/omnibase_infra/nodes/node_runner_health_snapshot_effect/handlers/handler_runner_fleet_snapshot.py
|
Warning Review limit reachedYou’ve reached a temporary PR review limit under our Fair Usage Limits Policy. Next review available in: 33 minutes Limit details: You’ve used the included review currently available. Your 112 included PR review attempts over the past 7 days set your current allowance at 1 review per hour. Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits within each organization. For paid Pro and Pro+ reviews, CodeRabbit uses a developer's included PR review attempts over the past 7 days to set the current hourly allowance. At typical activity levels, the full plan allowance applies. Higher sustained activity can lower the allowance until earlier attempts leave the 7-day window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (11)
Comment |
|
OCC autobind did not mint a companion for this PR: no changed-file candidate could be proven RED against the merge base, and emitting a PR-existence probe instead would be non-falsifiable evidence (OMN-15247). Hand-authored evidence is required. |
Resolves the two conflicts (pyproject.toml + uv.lock package version) in favour of dev's 0.38.7. main is at 0.38.6, so dev's value already satisfies the release-identity gate and taking the branch's stale 0.38.5 would regress it. All other main-only surfaces auto-merged cleanly and were verified file-by-file against origin/main.
✅ Hostile Reviewer — PASSEDBlocking findings (critical): 0 Gate semantics (pilot phase)
Powered by omniintelligence.review_pairing.cli_review — node-based adversarial review via HandlerLlmCliSubprocess (OMN-8468/OMN-8524) |
#6661) * evidence(OMN-16041): add content-pinned entry for OmniNode-ai/omnibase_infra#2745 * fix(OMN-16041): drop probe echo variable interpolation for receipt-hardening (OMN-15710) * evidence(OMN-16041): add self-bind entry for OCC#6661 * evidence(OMN-16041): dod_evidence receipt for OCC#6661 self-bind
…te + node_runner_health_snapshot_effect Contract Sync Gate (Wave C, OMN-8915) flagged both nodes' handlers as changed without a matching contract.yaml change on this backmerge. Root cause: main-side commit b437e48 (OMN-15195, "remove direct env reads from runner health bridge", 2026-07-26) changed both handlers without a paired contract.yaml update -- this drift predates the backmerge and was never caught on main, so there is no "ground truth" contract state to recover from main; the contracts were simply never updated when the handlers changed. node_runner_fleet_health_compute: the contract's description claimed RUNNER_HEALTH_MAX_DIAG_AGE_SECONDS was "env-overridable". OMN-15195 removed that env-read entirely (confirmed live in the current handler, which now documents "Nothing constructs this handler with overrides"); the three thresholds are plain module constants (5 / 4500 / 600). Corrected the description and added a Related Tickets entry. node_runner_health_snapshot_effect: OMN-15195 replaced this handler's direct env reads (WEDGE_QUEUE_AGE_SECONDS, RUNNER_CODELOAD_SCAN_LIMIT, WEDGE_WATCH_REPOS) with the typed ModelRunnerFleetConfig, loaded from config/runner_fleet.yaml. Added a Related Tickets entry documenting the config-model migration; no prior description text was factually wrong here (it didn't mention the env vars), so no description edit. Both contracts get a patch version bump (documentation-only fix, no input/output/topic/interface change) and updated metadata.updated / metadata.ticket. Verified: scripts/validate-pr-contract-sync.sh passes locally against both changed files; scripts/validate.py all reports 0 errors for both contracts (pre-existing runtime_profiles warnings on both files are unrelated -- present before this change too).
OMN-16041 — backmerge
mainlineage intodevThis is the
dev-side half of the promotion that landed onmainas #2744. It exists so the two branches stop re-conflicting on every promotion, and sodevdoes not silently regress the fixes that live only onmain.Conflict resolution — 2026-08-18
The branch had gone stale (opened 08-14,
mainpromoted 08-15,devmoved 13 commits since). Currentdevwas merged in and the conflicts resolved by reading both sides' history per path rather than taking a side wholesale.Only two files actually conflicted, and both were the same line:
mainlineagedevpyproject.toml0.38.50.38.7mainis already at0.38.6.dev's0.38.7is above it, so the release-identity gate is satisfied without any bump here; taking the branch's stale0.38.5would have pusheddevbelow the released version.uv.lock0.38.50.38.7pyproject.toml.Everything else merged without textual conflict. Each auto-merged path was still checked by hand against both branches, because "no conflict marker" is not the same as "correct":
.github/workflows/deploy-gate.yml.github/required-checks.yamlmainpins one commit from 2026-07-25 in both files.devpins two different commits — 2026-07-20 in the workflow and 2026-07-19 in the manifest — sodevwas both older and internally inconsistent with itself. Takingmainmoves to the newest pin and makes the two files agree again..github/workflows/env-parity.ymlmainwidened the sibling-checkout token to try the org-wide token before falling back, because the default token cannot read the private sibling repo.devnever received it and still has the narrower two-option form. Nothing on thedevside competes..github/workflows/artifact-reconciliation-webhook.ymlmain-only CI change from the previous bridge;devhas not touched this file since the branches diverged.scripts/validate_handler_contracts.pydevstill carries the original version, which validates a hardcoded list of seven handlers atnodes/handlers/<name>/contract.yaml— a directory that no longer exists in either branch.mainreplaced that with a glob over the live descriptor location. Verified by running it on the merged tree:5/5contracts validated,0failed.dev's version cannot pass at all.src/omnibase_infra/utils/util_runtime_packages.pymain-only; no competingdevchange. See the flag below — this one is carried, but it is not clean.mainremoved direct environment reads from both handlers.devnever got that and still describes its own reads as grandfathered — which they are only againstdev's history; measured againstmainthey read as new, which is exactly why the promotion kept conflicting here. The removal is carried.dev's behaviour is preserved, notmain's: the staleness threshold stays atdev's 4500s along with the comment explaining why (an idle runner only writes its diagnostic file on its ~50-minute token refresh, somain's 900s classified idle-but-healthy runners as zombies for most of every hour).src/omnibase_infra/observability/runner_health/model_runner_fleet_config.pymainadds three configuration fields that replace the removed environment reads;devadds three unrelated optional sub-configs from newer work. Both sets are kept. This is the one file in the merge that is deliberately not byte-identical tomain, and that is correct —mainsimply does not havedev's newer fields yet.Verification
devchange are byte-identical toorigin/main, so the next promotion has nothing left to conflict on in them.Flagged, carried but not endorsed
util_runtime_packages.pydoes not remove its two environment reads. It hides them: the module-level bindingreads exactly the same values, but routes around the gate, which matches on the literal spellings
os.environ[,os.environ.get,os.getenvand so on. An indirection like this is invisible to that check by construction.This is worth naming because the gate's own allowlist explicitly rejects the pattern — one of its entries reads "Formalizing the allowlist entry rather than duplicating that read through a new indirection layer." The correct shape is either a named allowlist entry or resolution through the config overlay, the way the two runner-health handlers in this same merge were fixed.
It is carried here rather than fixed, for one reason: it is already shipped on
main. Reverting it on thedevside would re-open the exact promotion conflict this PR exists to close, and would do it in a file where the gate then fires. Fixing it properly means changing both branches, which is its own change with its own test surface — filed as follow-up rather than smuggled into a backmerge.Known red, disclosed
tests/ci/test_env_parity.py::test_k8s_bound_keys_are_bound_in_composefails here exactly as it fails onorigin/dev. This PR neither introduces nor fixes it; the fix is open as #2740. Nothing here narrows, skips, or reclassifies that test.Evidence-Source: OCC#6661
Evidence-Ticket: OMN-16041