Skip to content

feat(complexity_router): gate Jev route-downs on classifier confidence - #42387

Open
n24q02m wants to merge 6 commits into
BerriAI:mainfrom
n24q02m:feat/jev-confidence-gate
Open

n24q02m wants to merge 6 commits into
BerriAI:mainfrom
n24q02m:feat/jev-confidence-gate

Conversation

@n24q02m

@n24q02m n24q02m commented Sep 22, 2026

Copy link
Copy Markdown

Closes #42139 (confidence-gate half; the fail-expensive default is a separate, follow-up change).

Rebase note: this branch is cut on current main. #41886 reshapes several of the same files
(jev_classifier.py, complexity_router.py, config.py); happy to rebase on top of it when it
merges. This change deliberately does not touch jev_classifier.py.

Problem

classifier_type: "jev" sends the request to TypeSafe Jev, then applies answer.choice
unconditionally. The classifier's confidence is recorded in the signals and in the routing
decision, but nothing ever consults it: a verdict the classifier itself rates at 0.2 confidence
still moves the request onto a cheaper tier. Under-routing is the expensive failure mode — the
response quality regresses, and nothing in the config lets an operator say "only degrade when Jev
is sure".

Change

Route-downs (verdicts naming a tier below the strongest configured tier) are now gated in the
apply path. Verdicts naming the strongest tier are never gated — routing up is not a degradation.

Three gates join JevClassifierConfig, all defaulting to None (gate off, current behavior):

Field Gate Refusal reason
confidence_threshold answer.confidence >= confidence_threshold low-confidence
complexity_max complexity score <= complexity_max high-complexity
complexity_confidence_min complexity confidence >= complexity_confidence_min low-complexity-confidence

All configured gates must pass; any failure refuses the route-down.

Semantics of a refusal:

  • The router follows the standard classifier-unusable fallback chain — llm_v2 capable tier,
    fallback_tier, then classifier_fallback. Pairing a threshold with an expensive fallback is
    what makes the router fail expensive on unsure verdicts.
  • The refusal is a policy decision, not a classifier health failure: the circuit breaker records
    success, so sustained low-confidence traffic cannot open the circuit.
  • The JevVerdict and its classifier cost stay on the outcome, so existing provenance logging
    (classifier_probabilities, classifier_confidence) is unchanged and the spend is still
    accounted.

Missing data is treated conservatively:

  • The wire protocol rejects an answer without a valid confidence, so that case already falls
    back today; a test pins it.
  • The built-in Jev request asks only the tier question, so there is no complexity answer yet.
    When complexity_max or complexity_confidence_min is configured, a missing complexity answer
    counts as the hardest case (score 1.0, confidence 0.0) and refuses the route-down.
    complexity_max: 1.0 / complexity_confidence_min: 0.0 are the values at which missing data
    passes. When a complexity answer exists on the wire (follow-up), the same gates read it —
    covered by unit tests on the gate function.

Choosing thresholds

Thresholds are pool-specific, not one-size-fits-all:

  • Pool whose cheap tier is nearly as capable as its top tier (small quality gap): 0.55–0.65
    harvests the savings at low risk.
  • Pool with a large capability gap between cheap and top: 0.85 — a wrong down-route is a real
    quality regression.

These live in the field descriptions so operators see the guidance in-config.

Testing

Mock-first; no network or key involved (JevClassifierClient is injected, mirroring the existing
suite):

  • 14 new tests in tests/unit/router_strategy/complexity_router/test_jev_confidence_gate.py:
    below-threshold refuses and fails over to the expensive destination (with verdict/cost
    provenance preserved), above-threshold and at-threshold apply, gate disabled keeps current
    behavior, strongest-tier verdicts never gated, wire answer missing confidence falls back
    conservatively, refusals record breaker success, each complexity-gate refusal reason, gate
    function boundary arithmetic, config validation of the three fields.
  • Existing test_jev_classifier.py (16 tests) and the router-level jev tests in
    test_complexity_router.py (8 tests) pass unchanged — the default config is behavior-neutral.

Compatibility

  • Default config is unchanged: every gate is None/off.
  • No new dependencies; no changes to jev_classifier.py.
  • extra="forbid" on the config model is unaffected: the fields are real model fields.

Follow-ups (happy to open issues/PRs)

  1. Fail-expensive classifier fallback default (the other half of Jev classifier: confidence-gated route-down + fail-expensive option #42139): classifier_fallback
    currently defaults to the local heuristic, which can also land a cheap tier — an unsure Jev
    verdict plus the heuristic fallback can agree on the wrong tier. capability and llm_v2
    already fail closed to the capable tier; jev should have an equivalent option.
  2. A complexity score answer in the Jev request, making the two complexity gates live instead of
    conservative-by-absence.

Update 2026-09-22

One documentation addition on top of the original patch: the complexity_max / complexity_confidence_min config descriptions now state explicitly that wire scores arrive as weighted level indices (0..n-1) and must be normalized to 0..1 before any configured threshold applies. Today the built-in Jev request asks only the tier question (the schema is choice-only), so an absent complexity answer reads as the hardest score (1.0) and a configured gate is fail-closed — but if the protocol gains a score question later, the normalization requirement is now written where the next implementer will look. Suite: 30 tests for the gate + doc patch, all passing.

@n24q02m
n24q02m requested a review from a team September 22, 2026 01:27
@greptile-apps

greptile-apps Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

The routing behavior appears safe, but the repository's explicit source-comment requirement must be satisfied before merging

Findings

  1. P2 Classifier model provenance missing ▶
  2. P2 Test comments violate convention ▶

Summary

This PR adds optional confidence and complexity gates before applying Jev route-down verdicts, preserving fallback behavior, circuit health, classifier forecasts, and cost attribution. It also adds focused configuration and routing tests.

  • Adds three validated, opt-in Jev gate settings
  • Refuses lower-tier verdicts when any configured gate fails
  • Preserves Jev verdict confidence, probabilities, and classifier cost through fallback
  • Treats policy refusals as circuit-breaker successes
  • Adds boundary, fallback, wire-validation, and breaker coverage

Reviews (1) · Last reviewed commit: "docs(complexity_router): note wire score..."

)
# Keep the verdict and its cost on the refusal for provenance: the request paid
# for the Jev call, and the routing decision still logs what it answered.
return refused._replace(classifier_cost=verdict.cost, jev_verdict=verdict)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Classifier model provenance missing Refused outcomes retain Jev forecasts and cost, but omit the model that produced them from routing records.

Knowledge Base Used: Router selection strategies

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Comment on lines +245 to +254
reason = ComplexityRouter._jev_route_down_refusal_reason

# Not a route-down: the strongest tier is never gated, whatever the confidence.
assert reason(config, "REASONING", "REASONING", 0.0) is None
# Confidence gate: closed below the threshold, open at it (complexity gates follow).
assert reason(config, "MEDIUM", "REASONING", 0.69) == "low-confidence"
assert reason(config, "MEDIUM", "REASONING", 0.7) == "high-complexity" # missing score is 1.0 > 0.5
# Complexity gates pass when the wire carries a qualifying answer.
assert reason(config, "MEDIUM", "REASONING", 0.7, complexity_score=0.3, complexity_confidence=0.6) is None
# Missing complexity data is hard: unknown score 1.0 refuses, unknown confidence 0.0 refuses.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Test comments violate convention These comments restate assertions, violating the repository directive that comments explain only complex logic. This requirement must be satisfied before merging.

Context Used: AGENTS.md (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

@codspeed

codspeed Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing n24q02m:feat/jev-confidence-gate (e510d8b) with main (cc30d93)

Open in CodSpeed

)
# Keep the verdict and its cost on the refusal for provenance: the request paid
# for the Jev call, and the routing decision still logs what it answered.
return refused._replace(classifier_cost=verdict.cost, jev_verdict=verdict)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Low: Refused classifier calls bypass spend accounting

When this refusal uses classifier_fallback='default_model', async_pre_routing_hook takes the direct fallback return at line 4604, whose routing decision omits classifier_cost, classifier_model, and the Jev forecast. A caller can repeatedly send prompts that produce below-threshold verdicts and consume paid Jev calls without those costs reaching the session rollup or savings accounting. Propagate this provenance through the direct fallback response as well as retaining it on ClassificationOutcome.

@veria-ai

veria-ai Bot commented Sep 22, 2026

Copy link
Copy Markdown
Contributor

PR overview

This PR updates the complexity router so Jev route-down decisions are gated on classifier confidence, with below-threshold outcomes using the configured fallback behavior.

One issue remains in the default-model fallback path: refused or below-threshold classifier calls can omit classifier cost and forecast metadata from spend and savings accounting. A caller could repeatedly trigger paid Jev classifications that are not reflected in session rollups, though this depends on the affected fallback configuration.

Open issues (1)

Fixed/addressed: 0 · PR risk: 4/10

@codecov

codecov Bot commented Sep 22, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 75.75758% with 8 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
...er_strategy/complexity_router/complexity_router.py 60.00% 8 Missing ⚠️

📢 Thoughts on this report? Let us know!

@n24q02m

n24q02m commented Sep 24, 2026

Copy link
Copy Markdown
Author

Fixed the codecov/patch failure: the gate tests lived under tests/unit/router_strategy/..., a path no workflow shard runs, so they never executed in CI and the patch lines in complexity_router.py reported zero coverage. Moved the file to tests/test_litellm/router_strategy/complexity_router/ (aff654f), which the enterprise-routing shard in test-unit.yml already runs with coverage upload. Same 14 tests, verified passing from the new path.

@n24q02m

n24q02m commented Sep 25, 2026

Copy link
Copy Markdown
Author

Thanks for the pointer to pydecide and the per-type threshold design in 0.2.0 — the
clipping asymmetry you described (one number gating a wide choice and a binary decision
very differently) is exactly the failure mode we were seeing, so we adopted the same
split in this PR.

The route-down gate now distinguishes the decision question's shape:

  • a question with exactly two options (a two-tier custom tier set) is yes/no-shaped and
    uses confidence_threshold_yes_no;
  • a wider question is choice-shaped and uses confidence_threshold_choice;
  • each per-type value falls back to the single confidence_threshold when unset, so
    existing configs that only set the single threshold keep gating every shape exactly
    as before (commit 112c6dd).

The option count comes from the tier-question criteria the router already builds, so no
wire changes were needed. Feedback welcome — especially on whether the two-option shape
should be detected differently once the protocol grows a native yes/no question type.

@n24q02m
n24q02m force-pushed the feat/jev-confidence-gate branch from 112c6dd to 009f279 Compare September 27, 2026 03:39
n24q02m and others added 6 commits September 28, 2026 08:16
A Jev verdict that routes below the strongest tier is applied only when
every configured gate passes. Three gates join JevClassifierConfig:
confidence_threshold (verdict confidence), complexity_max and
complexity_confidence_min (complexity-answer score and confidence). All
default to None, which keeps today's behavior of applying every valid
verdict; thresholds are pool-specific choices, so none are hard-coded.

A refused route-down follows the standard classifier-unusable fallback
chain (llm_v2 capable tier, fallback_tier, classifier_fallback), so
pairing a threshold with an expensive fallback makes the router fail
expensive. The refusal is a policy decision, not a health failure: the
circuit breaker records success, and the verdict plus its cost stay on
the outcome so provenance logging is unchanged.

The built-in Jev request asks only the tier question today, so missing
complexity data is treated as hard (score 1.0, confidence 0.0): a
configured complexity gate refuses route-downs until the wire carries a
complexity answer.

Refs: BerriAI#42139
complexity_max is validated 0-1; decision-API score answers arrive as
weighted level indices (0..levels-1, routing-spec v1.1). The built-in
request asks only the tier question today, so document the normalization
requirement for any future score question instead of encoding a scale
conversion nothing exercises yet.
…ence_min, complexity_max, confidence_threshold)
…nit shard

tests/unit/** is not referenced by any workflow shard, so the gate tests
never ran in CI and codecov/patch stayed red. The enterprise-routing
shard (test-unit.yml) already runs tests/test_litellm/router_strategy,
so the moved file executes there and its coverage uploads.
…thresholds

One confidence threshold behaves very differently for choice and yes/no
questions because Jev clips scores to 0.01-0.99: confidence spreads across
every option of a wide choice but only two of a binary decision. The gate now
follows the decision question's shape - a question with exactly two options
(built-in tiers never produce one; a two-tier custom tier_definitions set
does) is yes/no-shaped and uses confidence_threshold_yes_no, a wider one is
choice-shaped and uses confidence_threshold_choice. Each per-type value falls
back to the single confidence_threshold when absent, so existing single-
threshold configs keep gating every shape exactly as before (pydecide 0.2.0
adopted the same per-type split).

The option count comes from the criteria the router already builds for the
tier question; the member-tunable whitelist and the dashboard schema are
updated with the two new fields.
@n24q02m
n24q02m force-pushed the feat/jev-confidence-gate branch from 009f279 to e510d8b Compare September 28, 2026 01:18

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Jev classifier: confidence-gated route-down + fail-expensive option

1 participant