Repository navigation
fix(complexity_router): fall back to a live peer when the decided tier model is fully cooled down - #39675
Conversation
|
@greptileai can you reply here with a score and why that score? |
|
bugbot run |
|
Score: 4/5. The fix is well-scoped and the core approach is sound: The tests cover the important paths: pinned sessions, fresh classification, all peers being dead, unreadable cooldown state, unknown deployments, adaptive routing metadata, and modality compatibility. The routing-decision type and generated schema are updated consistently. I’m not giving 5/5 only because the required CI status is still pending, so full integration validation is not yet confirmed. Subject to CI passing, this is strong merge-ready work with no blocking correctness issue identified. |
Greptile SummaryThe PR adds a final health gate to complexity-router decisions so an unavailable model group can be replaced by a serving peer from the same tier.
Confidence Score: 5/5The PR appears safe to merge. No blocking failure remains.
|
| Filename | Overview |
|---|---|
| litellm/router_strategy/complexity_router/complexity_router.py | Adds request-aware same-tier health failover at both routing-hook exits; the previously reported peer-eligibility gaps are addressed by delegating capacity checks to the Router. |
| litellm/types/utils.py | Adds the new health-failover value to the routing-decision cause contract. |
| tests/test_litellm/router_strategy/test_complexity_router.py | Adds focused regression coverage for health, policy, rate-limit, prompt-window, plugin, modality, and session-affinity behavior. |
| ui/litellm-dashboard/src/lib/http/schema.d.ts | Keeps the generated dashboard schema aligned with the new routing-decision cause. |
Reviews (3): Last reviewed commit: "fix(complexity_router): fall back to a l..." | Re-trigger Greptile
PR overviewAll previously flagged issues have been addressed. No open security concerns remain on this pull request. Security reviewNo open security issues remain on this pull request. Fixed/addressed: 1 · PR risk: 0/10 |
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
077d813 to
8e7bee5
Compare
|
@greptileai can you reply here with a score and why that score? |
|
bugbot run |
|
Confidence score: 4/5. The fix now applies one health gate at every routing exit, uses the Router’s authoritative healthy-deployment lookup (including cooldown and admin-paused deployments), and fails open when health information is unavailable. The fallback remains within the decided tier and preserves modality/context-window constraints, routing-plugin filtering, adaptive-choice metadata, and the existing non-pinnable semantics. The added tests cover pinned and fresh routing, all-peer/unreadable-health cases, paused peers, plugin exclusions, adaptive routing, modality, and context-window limits. I’m not giving 5/5 yet because the change adds a fairly broad async gate around several routing paths and the PR’s CI status is still pending. I found no remaining blocking correctness issue in the current diff; assuming the targeted tests and CI pass, this is a strong 4/5 and ready for maintainer review. |
8e7bee5 to
33f78ae
Compare
|
@greptileai can you reply here with a score and why that score? |
|
bugbot run |
|
Score: 4/5 This is a strong, focused fix. The health gate is applied at the routing hook’s exits, covers fresh picks and session-pin replays, re-picks only within the decided tier, and preserves modality/context-window constraints. It also correctly accounts for cooldown and admin-paused deployments, fails open when health data is uncertain, preserves plugin exclusions, updates adaptive-routing metadata, and prevents failed-over decisions from being pinned. The added tests cover the main regression plus the important edge cases, and the generated TypeScript union is updated. I’m holding back the final point because the PR’s required CI status is still pending, and the end-to-end results in the description are reported evidence rather than independently verifiable checks here. |
|
@greptileai CI is now fully green on 33f78ae, including codecov/patch. Given that was the only holdback, can you reply here with an updated score and why that score? |
|
Correction on my previous comment: CI is not fully green yet. 21 checks have passed and 35 are still running, none failing. codecov/patch is green. I will re-ping once everything finishes. |
33f78ae to
8b05890
Compare
|
Both findings had one cause: the gate hand-assembled its own eligibility view, so every filter it did not replicate became a finding. It now delegates to async_get_healthy_deployments. Net -1 helper, 3 new tests. |
|
@greptileai can you reply here with a score and why that score? |
|
bugbot run |
|
Confidence Score: 4/5 The fix is well-contained and the current implementation addresses the previously identified correctness issues:
I’m keeping this at 4 rather than 5 because the PR’s CI/build status is still pending, so the full validation signal is not available yet. |
8b05890 to
890d1ad
Compare
|
Both fixed on 890d1ad. The probe now mirrors the owner's full contract: every negative verdict it can raise, and the same prompt arguments the real routing call passes. 3 new tests, each mutation-checked. |
|
@greptileai can you reply here with a score and why that score? |
|
bugbot run |
|
Score: 4/5 This is a strong, well-scoped fix. The health gate is applied at both routing exits in The added coverage is unusually comprehensive, including pinned and fresh routing, cooldowns, paused/excluded/RPM-limited peers, plugin exclusions, modality and context-window constraints, missing deployments, fail-open behavior, and the Responses |
|
The documentation check failure is inherited from staging: it flags heuristic_v2_router_limit, a router kwarg this diff never touches, whose reference doc is tracked in a separate docs PR. Every branch cut from current staging fails this check. |
|
code-quality is the same inherited failure: it runs the identical router-settings doc validation and trips on the same staging key. |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 890d1ad. Configure here.
|
@greptileai CI has finished on 890d1ad: 75 passing, and the only 2 failures are the inherited router-settings doc validation for a staging key this diff never touches (see the flagged comments above). Bugbot's check-run on this sha concluded success with no findings. Can you reply here with an updated score and why that score? |
|
Updated confidence score: 4/5. The implementation addresses the reported failure mode: it gates every decided tier placement, checks capacity through the router’s existing eligibility path, preserves modality/context-window/plugin constraints, fails open on uncertainty, and avoids pinning health substitutions. The added regression coverage is comprehensive, and CI reports 75 passing tests. The only two failures are inherited router-settings documentation validation failures for an unrelated staging key, and Bugbot completed successfully with no findings. I’m keeping this at 4/5 rather than 5/5 because the overall CI run is not fully green and this changes a core routing path, but there are no remaining findings attributable to this diff. |
|
Linked to Resolves LIT-6930. Bugbot is clean and all product checks passed; the only CI reds are the inherited router-settings documentation validation on staging. |
|
The documentation fix is already merged in BerriAI/litellm-docs#1180. I reran the two stale jobs so their checkout consumes that row. |
|
@greptileai review |
…r model is fully cooled down A complexity tier can name several model groups, but the pool pick and the session-pin replay both returned a group without consulting deployment health, so a group whose every deployment was in cooldown was still routed to and the request died at the router's zero-deployment check while a healthy peer sat in the same tier. Gate the decided response at the pre-routing hook's exits, the seam the modality gate already occupies, so every arm that can place a request is covered by one owner: a fresh classification, a replayed or escalated pin, a plan-mode floor, a context-window escalation, an adaptive pick, and whatever arm is added next. Peers come from the decided tier only. Climbing to a higher tier costs more than the classifier asked for and is left to a follow-up. The gate fails open on every uncertainty: an unreadable cooldown view, a decision carrying no tier, a group the router knows no deployments for, or a tier whose peers are all cooling.
890d1ad to
6defd00
Compare
|
Rebased onto df68edc and resolved the one shared _decision_is_pinnable hunk by retaining both the classifier-circuit and health-failover unpinnable rules. 20 focused tests and make check pass; merge state is now MERGEABLE. |
TLDR
Problem this solves:
How it solves it:
User Flow
Before: a developer calling an auto-router whose tier lists two models sees requests fail while one model's deployments are down, even though the other model is healthy
"model": "tier-health-router"and a short promptNo deployments available for selected model, Try again in 3600 secondsAfter: the same traffic serves from the healthy model in the same tier
"model": "tier-health-router"Relevant issues
Linear ticket
Resolves LIT-6930
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Screenshots / Proof of Fix
Shared setup: a live proxy on localhost with real gateway-backed LLM calls (no mocks, real spend). One auto-router whose every tier is the pool
["dead-tier-model", "live-tier-model"].dead-tier-modelpoints at a closed port so its only deployment enters cooldown on first use (allowed_fails: 0,cooldown_time: 3600,num_retries: 0);live-tier-modelis a realopenai/claude-haiku-4-5-20251001deploymentBoth arms ran against the same rig, same database, same config, same commands; only the checked-out source moved. Each arm asserts its own identity first, so the Before capture cannot accidentally run fixed code.
Before (39a1789)
Fresh picks keep landing on the cooled-down model
curl http://127.0.0.1:4620/v1/chat/completions -H "Authorization: Bearer $KEY" -d '{"model":"tier-health-router","messages":[{"role":"user","content":"say hi"}],"max_tokens":10}'dead-tier-modelwhile still warm and 500s withOpenAIException - Connection error, which cools its deployment down. req4, req5, req6, req10 and req11 then return429 No deployments available for selected model, Try again in 3600 seconds. Passed model=dead-tier-modelwhilelive-tier-modelsits healthy in the same tier. 5 of 12 requests failA session pinned to the cooled-down model is stuck for its whole lifetime
curl http://127.0.0.1:4621/v1/chat/completions ... -d '{"model":"tier-health-router","messages":[{"role":"user","content":"turn N say hi"}],"max_tokens":8,"metadata":{"session_id":"pN"}}'cause=session_affinity_pinand no fallback anywhereAfter (890d1ad)
Fresh picks route around the cooled-down model
cause=health_failover, routed_model=live-tier-model, displaced=dead-tier-modelonce per request that drew the dead group after it cooledThe pinned session serves every turn from the live peer
cause=health_failoverin the log. 15 of 16 turns succeed, against 9 before, and the only failure is the warm first hitNo deployments availableerrorsThe messages surface behaves the same
curl http://127.0.0.1:4621/v1/messages -H "Authorization: Bearer $KEY" -d '{"model":"tier-health-router","max_tokens":10,"messages":[{"role":"user","content":"say hi"}],"metadata":{"session_id":"p1"}}'Type
🐛 Bug Fix
Caveats (if any)
Medium
Low
Final Attestation
Note
Medium Risk
Changes core auto-router placement for multi-model tiers and adds async health probes on every routed request when a tier has peers, though behavior only shifts when the original pick would fail downstream.
Overview
When a complexity tier lists multiple model groups but the chosen group has no healthy deployments (cooldown, RPM limits, admin pause, context window, etc.), requests no longer stick on that group and 429. A new
_gate_response_healthstep runs at the end of the pre-routing hook (after modality gating) on every path that returns a decided model.Eligibility is checked via
async_get_healthy_deploymentson a copy ofrequest_kwargs, forwardingmessages,input, and OTEL span so the probe matches real routing. If the decided group cannot serve, the router picks a live peer in the same tier through_pick_model_for_tier(plugins still apply). Routing logscause=health_failoverwithhealth_displaced:<old_model>in signals;health_failoveris not session-pinnable so substitutes do not outlive the outage.The gate fails open when health is unreadable, the tier has only one model, no peer can serve, or plugins empty the pool; it does not escalate to a higher tier.
Reviewed by Cursor Bugbot for commit 890d1ad. Bugbot is set up for automated code reviews on this repo. Configure here.