Repository navigation
feat(OMN-7404): shadow-mode routing confidence gate in node_model_router_compute - #2376
Conversation
…ter_compute Adds RoutingGate (audit-only, per the plan's reviewed "Clarification from review" correction: a confidence estimator, never a model selector) as an optional injected dependency on HandlerScoreModels. Defaults to gate=None, which reproduces prior behavior byte-for-byte (proven by the existing OMN-14825 golden-equivalence corpus). No classifier has ever been trained (plan Task 7 unbuilt), so this is the only path exercised in production today. Deviates from the ticket's literal `_CLASSIFIER_PATH.exists()` file-check inside the handler: that pattern is I/O inside a COMPUTE node (forbidden by CLAUDE.md 7a). Classifier loading is pushed to whichever caller constructs the gate (a DI/registry seam), never into the handler itself. The literal `handler_model_router.py` target file no longer exists post OMN-14825 canonical def-B regeneration; this lands on its replacement, handler_score_models.py. RED->GREEN: tests/unit/learning/test_routing_gate.py fails at import before gate.py exists; tests/unit/nodes/test_model_router/test_score_models_routing_gate_omn7404.py fails (TypeError: unexpected keyword argument 'gate') on the pre-wiring handler. Both pass after implementation, and the pre-existing golden equivalence test (test_score_models_defb_omn14825.py) remains green unmodified, proving gate=None is a true no-op.
…n gates Governed pre-push selector (OMN-13973) and protocol-ownership allowlist (INFRA-016) both fail-closed on unregistered new modules/protocols; this satisfies both for the new src/omnibase_infra/learning package added in the prior commit.
|
Warning Review limit reachedYou’ve reached a temporary PR review limit under our Fair Usage Limits Policy. Next review available in: 30 minutes Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (11)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
Summary
OMN-7404 — "Task 8: [Phase 2] Deploy routing classifier gate in node_model_router_compute".
Adds
RoutingGateas an audit-only, shadow-mode confidence gate, injectableinto
HandlerScoreModels(the canonical def-B handler fornode_model_router_compute,OMN-14825). The gate never selects a model or alters the routing decision — it only
computes
{confidence, would_flag}for logging, per the plan's own reviewed"Clarification from review" correction (
docs/plans/2026-04-03-learning-infrastructure.md,Task 8): "The classifier predicts 'is this routing decision likely good?' — a binary
confidence estimator, NOT a model selector... Shadow-mode only in Phase 2."
Scope deviations from the ticket text (and why)
The Linear ticket description (and an earlier, superseded draft of the plan) describes
RoutingGate.recommend()with an override-capableuse_classifierflag, wiring intohandler_model_router.py, and aPath(...).exists()check inside the handler. None ofthat is implemented as literally written, for concrete reasons:
handler_model_router.pyno longer exists. OMN-14825 (merged, PR feat(OMN-14825): regenerate node_model_router_compute to canonical def-B #2360) regeneratednode_model_router_computeto the canonical def-B shape; the current handler ishandler_score_models.py::HandlerScoreModels.handle().review" block explicitly rules out an override-capable
recommend()— audit only,shadow-mode only. This PR implements the corrected
audit()design, not the stalerecommend()draft.RoutingClassifier) was neverbuilt — no pickle artifact, no pandas/sklearn dependency anywhere in this repo.
gate=Noneis therefore the only path ever exercised in production today, which is why default
behavior must reproduce prior output byte-for-byte (see golden-equivalence proof below).
node_model_router_computeis aCOMPUTE_GENERICnode — CLAUDE.md §7a requires it remain I/O-free and deterministic.Checking
_CLASSIFIER_PATH.exists()insidehandle()would violate that. Instead, thegate is an optional constructor-injected dependency (
HandlerScoreModels(gate=...));loading any real classifier artifact is pushed to whichever caller constructs the gate
(a DI/registry seam), never to the handler.
What changed
src/omnibase_infra/learning/routing/gate.py—RoutingGate.audit(), pure/dependency-free(no pandas), graceful degradation on
classifier=Noneor classifier exception.src/omnibase_infra/learning/routing/typed_dict_routing_audit.py—TypedDictRoutingAudit(kept the union-usage ratchet at 154/154 instead of a
dict[str, float | bool | None]3-way-union value type).
handler_score_models.py— optionalgate: RoutingGate | None = Noneconstructor param;when present, logs a
routing_audit: ...line after computing the decision, never mutating it.contract.yaml— patch bump (1.0.2 → 1.0.3) documenting the addition.scripts/ci/test_selection_adjacency.yaml+tests/unit/contracts/test_protocol_ownership.py—registered the new
learningmodule andProtocolRoutingClassifierin their respectivefail-closed allowlists (governed pre-push selector OMN-13973, protocol-ownership gate INFRA-016).
RED → GREEN evidence
tests/unit/learning/test_routing_gate.py— fails at collection (ImportError) beforegate.pyexists; 6/6 pass after.tests/unit/nodes/test_model_router/test_score_models_routing_gate_omn7404.py— verified REDvia
git stashof the handler edit (TypeError: HandlerScoreModels() takes no argumentson4/5 gate-dependent tests); 5/5 pass after restoring the wiring.
tests/unit/nodes/test_model_router/test_score_models_defb_omn14825.py(pre-existinggolden-equivalence corpus, untouched) stays green — proves
gate=None(today's onlyproduction path) is a byte-for-byte no-op.
Test plan
uv run pytest tests/unit/learning/ tests/unit/nodes/test_model_router/ tests/unit/contracts/test_protocol_ownership.py tests/unit/scripts/ci/test_test_selection_loader.py -v— 55/55 passuv run ruff format/uv run ruff check --fixcleanuv run mypy src/omnibase_infra/learning src/omnibase_infra/nodes/node_model_router_compute/handlers/handler_score_models.pycleanpre-commit run --all-files(targeted) cleanuv run pytest tests/ --ignore=tests/integration— 21844 passed, 36 skippedgh pr checksgreen (watching)dod_evidence: RED→GREEN proof above; ticket OMN-7404 cited; PR body documents the corrected/authoritative plan source superseding the stale ticket draft.
Evidence-Ticket: OMN-7404
Evidence-Source: OCC#4593
Evidence-Commit: c7a371a9428394710186c03519f43f522a0b36fd