Skip to content

feat(OMN-7404): shadow-mode routing confidence gate in node_model_router_compute - #2376

Merged
jonahgabriel merged 2 commits into
devfrom
jonah/omn-7404-routing-classifier-gate-phase2
Jul 22, 2026
Merged

jonahgabriel merged 2 commits into
devfrom
jonah/omn-7404-routing-classifier-gate-phase2

Conversation

@jonahgabriel

@jonahgabriel jonahgabriel commented Jul 21, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

OMN-7404 — "Task 8: [Phase 2] Deploy routing classifier gate in node_model_router_compute".

Adds RoutingGate as an audit-only, shadow-mode confidence gate, injectable
into HandlerScoreModels (the canonical def-B handler for node_model_router_compute,
OMN-14825). The gate never selects a model or alters the routing decision — it only
computes {confidence, would_flag} for logging, per the plan's own reviewed
"Clarification from review" correction (docs/plans/2026-04-03-learning-infrastructure.md,
Task 8): "The classifier predicts 'is this routing decision likely good?' — a binary
confidence estimator, NOT a model selector... Shadow-mode only in Phase 2."

Scope deviations from the ticket text (and why)

The Linear ticket description (and an earlier, superseded draft of the plan) describes
RoutingGate.recommend() with an override-capable use_classifier flag, wiring into
handler_model_router.py, and a Path(...).exists() check inside the handler. None of
that is implemented as literally written, for concrete reasons:

  1. handler_model_router.py no longer exists. OMN-14825 (merged, PR feat(OMN-14825): regenerate node_model_router_compute to canonical def-B #2360) regenerated
    node_model_router_compute to the canonical def-B shape; the current handler is
    handler_score_models.py::HandlerScoreModels.handle().
  2. The plan itself was corrected after the ticket was drafted. The "Clarification from
    review" block explicitly rules out an override-capable recommend() — audit only,
    shadow-mode only. This PR implements the corrected audit() design, not the stale
    recommend() draft.
  3. No classifier exists. Plan Task 7 (train + persist RoutingClassifier) was never
    built — no pickle artifact, no pandas/sklearn dependency anywhere in this repo. gate=None
    is therefore the only path ever exercised in production today, which is why default
    behavior must reproduce prior output byte-for-byte (see golden-equivalence proof below).
  4. No file-existence check inside the handler. node_model_router_compute is a
    COMPUTE_GENERIC node — CLAUDE.md §7a requires it remain I/O-free and deterministic.
    Checking _CLASSIFIER_PATH.exists() inside handle() would violate that. Instead, the
    gate is an optional constructor-injected dependency (HandlerScoreModels(gate=...));
    loading any real classifier artifact is pushed to whichever caller constructs the gate
    (a DI/registry seam), never to the handler.

What changed

  • src/omnibase_infra/learning/routing/gate.py — RoutingGate.audit(), pure/dependency-free
    (no pandas), graceful degradation on classifier=None or classifier exception.
  • src/omnibase_infra/learning/routing/typed_dict_routing_audit.py — TypedDictRoutingAudit
    (kept the union-usage ratchet at 154/154 instead of a dict[str, float | bool | None]
    3-way-union value type).
  • handler_score_models.py — optional gate: RoutingGate | None = None constructor param;
    when present, logs a routing_audit: ... line after computing the decision, never mutating it.
  • contract.yaml — patch bump (1.0.2 → 1.0.3) documenting the addition.
  • scripts/ci/test_selection_adjacency.yaml + tests/unit/contracts/test_protocol_ownership.py —
    registered the new learning module and ProtocolRoutingClassifier in their respective
    fail-closed allowlists (governed pre-push selector OMN-13973, protocol-ownership gate INFRA-016).

RED → GREEN evidence

  • tests/unit/learning/test_routing_gate.py — fails at collection (ImportError) before
    gate.py exists; 6/6 pass after.
  • tests/unit/nodes/test_model_router/test_score_models_routing_gate_omn7404.py — verified RED
    via git stash of the handler edit (TypeError: HandlerScoreModels() takes no arguments on
    4/5 gate-dependent tests); 5/5 pass after restoring the wiring.
  • tests/unit/nodes/test_model_router/test_score_models_defb_omn14825.py (pre-existing
    golden-equivalence corpus, untouched) stays green — proves gate=None (today's only
    production path) is a byte-for-byte no-op.

Test plan

  • uv run pytest tests/unit/learning/ tests/unit/nodes/test_model_router/ tests/unit/contracts/test_protocol_ownership.py tests/unit/scripts/ci/test_test_selection_loader.py -v — 55/55 pass
  • uv run ruff format / uv run ruff check --fix clean
  • uv run mypy src/omnibase_infra/learning src/omnibase_infra/nodes/node_model_router_compute/handlers/handler_score_models.py clean
  • pre-commit run --all-files (targeted) clean
  • Full local pre-push suite (governed selector escalated to full suite): uv run pytest tests/ --ignore=tests/integration — 21844 passed, 36 skipped
  • gh pr checks green (watching)

dod_evidence: RED→GREEN proof above; ticket OMN-7404 cited; PR body documents the corrected/authoritative plan source superseding the stale ticket draft.

Evidence-Ticket: OMN-7404
Evidence-Source: OCC#4593
Evidence-Commit: c7a371a9428394710186c03519f43f522a0b36fd

Jonah Gray added 2 commits July 20, 2026 19:25
…ter_compute

Adds RoutingGate (audit-only, per the plan's reviewed "Clarification from
review" correction: a confidence estimator, never a model selector) as an
optional injected dependency on HandlerScoreModels. Defaults to gate=None,
which reproduces prior behavior byte-for-byte (proven by the existing
OMN-14825 golden-equivalence corpus). No classifier has ever been trained
(plan Task 7 unbuilt), so this is the only path exercised in production
today.

Deviates from the ticket's literal `_CLASSIFIER_PATH.exists()` file-check
inside the handler: that pattern is I/O inside a COMPUTE node (forbidden by
CLAUDE.md 7a). Classifier loading is pushed to whichever caller constructs
the gate (a DI/registry seam), never into the handler itself. The literal
`handler_model_router.py` target file no longer exists post OMN-14825
canonical def-B regeneration; this lands on its replacement,
handler_score_models.py.

RED->GREEN: tests/unit/learning/test_routing_gate.py fails at import before
gate.py exists; tests/unit/nodes/test_model_router/test_score_models_routing_gate_omn7404.py
fails (TypeError: unexpected keyword argument 'gate') on the pre-wiring
handler. Both pass after implementation, and the pre-existing golden
equivalence test (test_score_models_defb_omn14825.py) remains green
unmodified, proving gate=None is a true no-op.
…n gates

Governed pre-push selector (OMN-13973) and protocol-ownership allowlist
(INFRA-016) both fail-closed on unregistered new modules/protocols;
this satisfies both for the new src/omnibase_infra/learning package added
in the prior commit.
@coderabbitai

coderabbitai Bot commented Jul 21, 2026

Copy link
Copy Markdown

Warning

Review limit reached

You’ve reached a temporary PR review limit under our Fair Usage Limits Policy.

Your recent review volume is higher than typical usage, so adaptive limits are currently applied.

Next review available in: 30 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 1e6bca14-5241-445b-8e14-17e982463049

📥 Commits

Reviewing files that changed from the base of the PR and between 43935c8 and ce4e601.

📒 Files selected for processing (11)
  • scripts/ci/test_selection_adjacency.yaml
  • src/omnibase_infra/learning/__init__.py
  • src/omnibase_infra/learning/routing/__init__.py
  • src/omnibase_infra/learning/routing/gate.py
  • src/omnibase_infra/learning/routing/typed_dict_routing_audit.py
  • src/omnibase_infra/nodes/node_model_router_compute/contract.yaml
  • src/omnibase_infra/nodes/node_model_router_compute/handlers/handler_score_models.py
  • tests/unit/contracts/test_protocol_ownership.py
  • tests/unit/learning/__init__.py
  • tests/unit/learning/test_routing_gate.py
  • tests/unit/nodes/test_model_router/test_score_models_routing_gate_omn7404.py
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch jonah/omn-7404-routing-classifier-gate-phase2

Comment @coderabbitai help to get the list of available commands.

@jonahgabriel
jonahgabriel merged commit 4380e8b into dev Jul 22, 2026
252 of 312 checks passed
@jonahgabriel
jonahgabriel deleted the jonah/omn-7404-routing-classifier-gate-phase2 branch July 22, 2026 04:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant