Repository navigation
Conversation
|
| ) | ||
| # Keep the verdict and its cost on the refusal for provenance: the request paid | ||
| # for the Jev call, and the routing decision still logs what it answered. | ||
| return refused._replace(classifier_cost=verdict.cost, jev_verdict=verdict) |
There was a problem hiding this comment.
Classifier model provenance missing Refused outcomes retain Jev forecasts and cost, but omit the model that produced them from routing records.
Knowledge Base Used: Router selection strategies
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
| reason = ComplexityRouter._jev_route_down_refusal_reason | ||
|
|
||
| # Not a route-down: the strongest tier is never gated, whatever the confidence. | ||
| assert reason(config, "REASONING", "REASONING", 0.0) is None | ||
| # Confidence gate: closed below the threshold, open at it (complexity gates follow). | ||
| assert reason(config, "MEDIUM", "REASONING", 0.69) == "low-confidence" | ||
| assert reason(config, "MEDIUM", "REASONING", 0.7) == "high-complexity" # missing score is 1.0 > 0.5 | ||
| # Complexity gates pass when the wire carries a qualifying answer. | ||
| assert reason(config, "MEDIUM", "REASONING", 0.7, complexity_score=0.3, complexity_confidence=0.6) is None | ||
| # Missing complexity data is hard: unknown score 1.0 refuses, unknown confidence 0.0 refuses. |
There was a problem hiding this comment.
Test comments violate convention These comments restate assertions, violating the repository directive that comments explain only complex logic. This requirement must be satisfied before merging.
Context Used: AGENTS.md (source)
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
| ) | ||
| # Keep the verdict and its cost on the refusal for provenance: the request paid | ||
| # for the Jev call, and the routing decision still logs what it answered. | ||
| return refused._replace(classifier_cost=verdict.cost, jev_verdict=verdict) |
There was a problem hiding this comment.
Low: Refused classifier calls bypass spend accounting
When this refusal uses classifier_fallback='default_model', async_pre_routing_hook takes the direct fallback return at line 4604, whose routing decision omits classifier_cost, classifier_model, and the Jev forecast. A caller can repeatedly send prompts that produce below-threshold verdicts and consume paid Jev calls without those costs reaching the session rollup or savings accounting. Propagate this provenance through the direct fallback response as well as retaining it on ClassificationOutcome.
PR overviewThis PR updates the complexity router so Jev route-down decisions are gated on classifier confidence, with below-threshold outcomes using the configured fallback behavior. One issue remains in the default-model fallback path: refused or below-threshold classifier calls can omit classifier cost and forecast metadata from spend and savings accounting. A caller could repeatedly trigger paid Jev classifications that are not reflected in session rollups, though this depends on the affected fallback configuration. Open issues (1)
Fixed/addressed: 0 · PR risk: 4/10 |
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
|
Fixed the codecov/patch failure: the gate tests lived under |
|
Thanks for the pointer to pydecide and the per-type threshold design in 0.2.0 — the The route-down gate now distinguishes the decision question's shape:
The option count comes from the tier-question criteria the router already builds, so no |
112c6dd to
009f279
Compare
A Jev verdict that routes below the strongest tier is applied only when every configured gate passes. Three gates join JevClassifierConfig: confidence_threshold (verdict confidence), complexity_max and complexity_confidence_min (complexity-answer score and confidence). All default to None, which keeps today's behavior of applying every valid verdict; thresholds are pool-specific choices, so none are hard-coded. A refused route-down follows the standard classifier-unusable fallback chain (llm_v2 capable tier, fallback_tier, classifier_fallback), so pairing a threshold with an expensive fallback makes the router fail expensive. The refusal is a policy decision, not a health failure: the circuit breaker records success, and the verdict plus its cost stay on the outcome so provenance logging is unchanged. The built-in Jev request asks only the tier question today, so missing complexity data is treated as hard (score 1.0, confidence 0.0): a configured complexity gate refuses route-downs until the wire carries a complexity answer. Refs: BerriAI#42139
complexity_max is validated 0-1; decision-API score answers arrive as weighted level indices (0..levels-1, routing-spec v1.1). The built-in request asks only the tier question today, so document the normalization requirement for any future score question instead of encoding a scale conversion nothing exercises yet.
…lients for request_kwargs
…ence_min, complexity_max, confidence_threshold)
…nit shard tests/unit/** is not referenced by any workflow shard, so the gate tests never ran in CI and codecov/patch stayed red. The enterprise-routing shard (test-unit.yml) already runs tests/test_litellm/router_strategy, so the moved file executes there and its coverage uploads.
…thresholds One confidence threshold behaves very differently for choice and yes/no questions because Jev clips scores to 0.01-0.99: confidence spreads across every option of a wide choice but only two of a binary decision. The gate now follows the decision question's shape - a question with exactly two options (built-in tiers never produce one; a two-tier custom tier_definitions set does) is yes/no-shaped and uses confidence_threshold_yes_no, a wider one is choice-shaped and uses confidence_threshold_choice. Each per-type value falls back to the single confidence_threshold when absent, so existing single- threshold configs keep gating every shape exactly as before (pydecide 0.2.0 adopted the same per-type split). The option count comes from the criteria the router already builds for the tier question; the member-tunable whitelist and the dashboard schema are updated with the two new fields.
009f279 to
e510d8b
Compare
Closes #42139 (confidence-gate half; the fail-expensive default is a separate, follow-up change).
Problem
classifier_type: "jev"sends the request to TypeSafe Jev, then appliesanswer.choiceunconditionally. The classifier's
confidenceis recorded in the signals and in the routingdecision, but nothing ever consults it: a verdict the classifier itself rates at 0.2 confidence
still moves the request onto a cheaper tier. Under-routing is the expensive failure mode — the
response quality regresses, and nothing in the config lets an operator say "only degrade when Jev
is sure".
Change
Route-downs (verdicts naming a tier below the strongest configured tier) are now gated in the
apply path. Verdicts naming the strongest tier are never gated — routing up is not a degradation.
Three gates join
JevClassifierConfig, all defaulting toNone(gate off, current behavior):confidence_thresholdanswer.confidence >= confidence_thresholdlow-confidencecomplexity_max<= complexity_maxhigh-complexitycomplexity_confidence_min>= complexity_confidence_minlow-complexity-confidenceAll configured gates must pass; any failure refuses the route-down.
Semantics of a refusal:
fallback_tier, thenclassifier_fallback. Pairing a threshold with an expensive fallback iswhat makes the router fail expensive on unsure verdicts.
success, so sustained low-confidence traffic cannot open the circuit.
JevVerdictand its classifier cost stay on the outcome, so existing provenance logging(
classifier_probabilities,classifier_confidence) is unchanged and the spend is stillaccounted.
Missing data is treated conservatively:
confidence, so that case already fallsback today; a test pins it.
When
complexity_maxorcomplexity_confidence_minis configured, a missing complexity answercounts as the hardest case (score 1.0, confidence 0.0) and refuses the route-down.
complexity_max: 1.0/complexity_confidence_min: 0.0are the values at which missing datapasses. When a complexity answer exists on the wire (follow-up), the same gates read it —
covered by unit tests on the gate function.
Choosing thresholds
Thresholds are pool-specific, not one-size-fits-all:
0.55–0.65harvests the savings at low risk.
0.85— a wrong down-route is a realquality regression.
These live in the field descriptions so operators see the guidance in-config.
Testing
Mock-first; no network or key involved (
JevClassifierClientis injected, mirroring the existingsuite):
tests/unit/router_strategy/complexity_router/test_jev_confidence_gate.py:below-threshold refuses and fails over to the expensive destination (with verdict/cost
provenance preserved), above-threshold and at-threshold apply, gate disabled keeps current
behavior, strongest-tier verdicts never gated, wire answer missing
confidencefalls backconservatively, refusals record breaker success, each complexity-gate refusal reason, gate
function boundary arithmetic, config validation of the three fields.
test_jev_classifier.py(16 tests) and the router-leveljevtests intest_complexity_router.py(8 tests) pass unchanged — the default config is behavior-neutral.Compatibility
None/off.jev_classifier.py.extra="forbid"on the config model is unaffected: the fields are real model fields.Follow-ups (happy to open issues/PRs)
classifier_fallbackcurrently defaults to the local heuristic, which can also land a cheap tier — an unsure Jev
verdict plus the heuristic fallback can agree on the wrong tier.
capabilityandllm_v2already fail closed to the capable tier;
jevshould have an equivalent option.conservative-by-absence.
Update 2026-09-22
One documentation addition on top of the original patch: the
complexity_max/complexity_confidence_minconfig descriptions now state explicitly that wire scores arrive as weighted level indices (0..n-1) and must be normalized to 0..1 before any configured threshold applies. Today the built-in Jev request asks only the tier question (the schema is choice-only), so an absent complexity answer reads as the hardest score (1.0) and a configured gate is fail-closed — but if the protocol gains a score question later, the normalization requirement is now written where the next implementer will look. Suite: 30 tests for the gate + doc patch, all passing.