Skip to content

fix(complexity_router): fall back to a live peer when the decided tier model is fully cooled down - #39675

Merged
tin-berri merged 1 commit into
litellm_internal_stagingfrom
litellm_tier_health_filter
Sep 4, 2026
Merged

tin-berri merged 1 commit into
litellm_internal_stagingfrom
litellm_tier_health_filter

Conversation

@tin-berri

@tin-berri tin-berri commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • A complexity tier can list several models, yet the pick ignores deployment health
  • A tier model in full cooldown still gets routed to and the request 429s
  • A session pinned to that model fails every remaining turn

How it solves it:

  • One health gate on the decided routing response, at the hook's exits
  • A model group with no capacity is swapped for a live peer in the same tier
  • The substitute is picked through the normal tier pick, so plugins still apply
  • Fails open on any uncertainty, so behavior only changes where requests fail today

User Flow

Before: a developer calling an auto-router whose tier lists two models sees requests fail while one model's deployments are down, even though the other model is healthy

  1. They send POST http://litellm-domain/v1/chat/completions with "model": "tier-health-router" and a short prompt
  2. The tier's first model is down, so its deployments enter cooldown after a connection error
  3. Roughly half their requests now return 429 No deployments available for selected model, Try again in 3600 seconds
  4. A session whose first turn landed on the down model returns that 429 on every later turn until the session's pin expires, roughly an hour by default

After: the same traffic serves from the healthy model in the same tier

  1. They send the same POST http://litellm-domain/v1/chat/completions with "model": "tier-health-router"
  2. The down model's deployments enter cooldown after the same connection error
  3. Every later request returns 200, served by the healthy model in the same tier
  4. The pinned session also returns 200 on every turn, and the response's routing record names the model that was skipped

Relevant issues

Linear ticket

Resolves LIT-6930

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • Required CI checks pass, excluding inherited documentation validation failures from staging
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Screenshots / Proof of Fix

Shared setup: a live proxy on localhost with real gateway-backed LLM calls (no mocks, real spend). One auto-router whose every tier is the pool ["dead-tier-model", "live-tier-model"]. dead-tier-model points at a closed port so its only deployment enters cooldown on first use (allowed_fails: 0, cooldown_time: 3600, num_retries: 0); live-tier-model is a real openai/claude-haiku-4-5-20251001 deployment

model_list:
  - model_name: tier-health-router
    litellm_params:
      model: auto_router/complexity_router
      complexity_router_config:
        session_affinity: true
        tiers:
          SIMPLE: ["dead-tier-model", "live-tier-model"]
          MEDIUM: ["dead-tier-model", "live-tier-model"]
          COMPLEX: ["dead-tier-model", "live-tier-model"]
          REASONING: ["dead-tier-model", "live-tier-model"]
  - model_name: dead-tier-model
    litellm_params:
      model: openai/claude-haiku-4-5-20251001
      api_base: http://127.0.0.1:9
      api_key: dead-key
  - model_name: live-tier-model
    litellm_params:
      model: openai/claude-haiku-4-5-20251001
      api_base: https://gateway.litellm-sandbox.ai/v1
      api_key: os.environ/GW_KEY
router_settings:
  num_retries: 0
  allowed_fails: 0
  cooldown_time: 3600

Both arms ran against the same rig, same database, same config, same commands; only the checked-out source moved. Each arm asserts its own identity first, so the Before capture cannot accidentally run fixed code.

Before (39a1789)

Fresh picks keep landing on the cooled-down model

  1. Send 12 identical requests: curl http://127.0.0.1:4620/v1/chat/completions -H "Authorization: Bearer $KEY" -d '{"model":"tier-health-router","messages":[{"role":"user","content":"say hi"}],"max_tokens":10}'
  2. Observed: req2 draws dead-tier-model while still warm and 500s with OpenAIException - Connection error, which cools its deployment down. req4, req5, req6, req10 and req11 then return 429 No deployments available for selected model, Try again in 3600 seconds. Passed model=dead-tier-model while live-tier-model sits healthy in the same tier. 5 of 12 requests fail

A session pinned to the cooled-down model is stuck for its whole lifetime

  1. Run 4 sessions of 4 turns each, every turn curl http://127.0.0.1:4621/v1/chat/completions ... -d '{"model":"tier-health-router","messages":[{"role":"user","content":"turn N say hi"}],"max_tokens":8,"metadata":{"session_id":"pN"}}'
  2. Observed: session p3 loses turn 1 to the connection error then 429s on turns 2, 3 and 4; session p4 429s on all 4 turns. 7 of 16 turns fail
  3. The proxy log shows 13 turns replayed with cause=session_affinity_pin and no fallback anywhere

After (890d1ad)

Fresh picks route around the cooled-down model

  1. Send the same 12 requests against the same config
  2. Observed: req7 draws the dead model while still warm and 500s once, which is what creates the cooldown the gate reads. Every other request returns 200, and no request returns 429. 1 of 12 fails, against 5 before
  3. The log shows cause=health_failover, routed_model=live-tier-model, displaced=dead-tier-model once per request that drew the dead group after it cooled

The pinned session serves every turn from the live peer

  1. Run the same 4 sessions of 4 turns
  2. Observed: one session loses turn 1 to the warm connection error and pins onto the dead group; its remaining turns replay that pin and every one serves 200 from the live peer, with 5 turns rewritten to cause=health_failover in the log. 15 of 16 turns succeed, against 9 before, and the only failure is the warm first hit
  3. Across both legs: 27 successes, 2 failures, both of them the warm hit that creates the cooldown, and zero No deployments available errors

The messages surface behaves the same

  1. curl http://127.0.0.1:4621/v1/messages -H "Authorization: Bearer $KEY" -d '{"model":"tier-health-router","max_tokens":10,"messages":[{"role":"user","content":"say hi"}],"metadata":{"session_id":"p1"}}'
  2. Observed: 200, served through the same gate

Type

🐛 Bug Fix

Caveats (if any)

Medium

  • A failed-over turn is never pinned, so each turn re-picks while the group is down
    • Deliberate: a failover describes the fleet, not the session's traffic
    • The session returns to normal classification once the group recovers
  • The first request into a still-warm dead deployment still fails
    • Cooldown state is the signal; it only exists after a failure

Low

  • Cost: zero for single-model tiers (the gate returns before any lookup); otherwise one health lookup for the decided group, measured at ~10 microseconds on a 20-deployment router, and peer lookups only when that group is actually out of capacity. Redis stays batched behind the dual cache's 10 second per-key throttle
  • The auto-router health derivation (fix(proxy): derive auto-router health from its underlying models #38174) still reports the router red when one pool member is fully dead, though the router now serves; revisiting that rule is a follow-up
  • The adaptive bandit's own pick arms stay health-blind upstream; the gate corrects them at the exit

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Note

Medium Risk
Changes core auto-router placement for multi-model tiers and adds async health probes on every routed request when a tier has peers, though behavior only shifts when the original pick would fail downstream.

Overview
When a complexity tier lists multiple model groups but the chosen group has no healthy deployments (cooldown, RPM limits, admin pause, context window, etc.), requests no longer stick on that group and 429. A new _gate_response_health step runs at the end of the pre-routing hook (after modality gating) on every path that returns a decided model.

Eligibility is checked via async_get_healthy_deployments on a copy of request_kwargs, forwarding messages, input, and OTEL span so the probe matches real routing. If the decided group cannot serve, the router picks a live peer in the same tier through _pick_model_for_tier (plugins still apply). Routing logs cause=health_failover with health_displaced:<old_model> in signals; health_failover is not session-pinnable so substitutes do not outlive the outage.

The gate fails open when health is unreadable, the tier has only one model, no peer can serve, or plugins empty the pool; it does not escalate to a higher tier.

Reviewed by Cursor Bugbot for commit 890d1ad. Bugbot is set up for automated code reviews on this repo. Configure here.

@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai can you reply here with a score and why that score?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Sep 3, 2026

Copy link
Copy Markdown
Contributor

Score: 4/5.

The fix is well-scoped and the core approach is sound: _gate_response_health() is applied at both routing exits, checks that every deployment for the selected model group is cooled down, and fails open when cooldown or deployment metadata is uncertain. It also preserves the selected tier, applies modality constraints, updates the adaptive chosen-model marker, and keeps session-pin behavior explicit.

The tests cover the important paths: pinned sessions, fresh classification, all peers being dead, unreadable cooldown state, unknown deployments, adaptive routing metadata, and modality compatibility. The routing-decision type and generated schema are updated consistently.

I’m not giving 5/5 only because the required CI status is still pending, so full integration validation is not yet confirmed. Subject to CI passing, this is strong merge-ready work with no blocking correctness issue identified.

@codspeed

codspeed Bot commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_tier_health_filter (6defd00) with litellm_internal_staging (df68edc)

Open in CodSpeed

@greptile-apps

greptile-apps Bot commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

The PR adds a final health gate to complexity-router decisions so an unavailable model group can be replaced by a serving peer from the same tier.

  • Uses the Router’s request-aware deployment eligibility path for the decided group and candidate peers.
  • Preserves routing-plugin participation and records failovers with a dedicated routing-decision cause.
  • Adds coverage for cooldowns, paused and policy-excluded deployments, rate limits, context windows, session affinity, modality routing, and fail-open behavior.
  • Extends the shared and dashboard routing-cause types with health_failover.

Confidence Score: 5/5

The PR appears safe to merge.

No blocking failure remains.

Important Files Changed

Filename Overview
litellm/router_strategy/complexity_router/complexity_router.py Adds request-aware same-tier health failover at both routing-hook exits; the previously reported peer-eligibility gaps are addressed by delegating capacity checks to the Router.
litellm/types/utils.py Adds the new health-failover value to the routing-decision cause contract.
tests/test_litellm/router_strategy/test_complexity_router.py Adds focused regression coverage for health, policy, rate-limit, prompt-window, plugin, modality, and session-affinity behavior.
ui/litellm-dashboard/src/lib/http/schema.d.ts Keeps the generated dashboard schema aligned with the new routing-decision cause.

Reviews (3): Last reviewed commit: "fix(complexity_router): fall back to a l..." | Re-trigger Greptile

Comment thread litellm/router_strategy/complexity_router/complexity_router.py Outdated
Comment thread tests/test_litellm/router_strategy/test_complexity_router.py Outdated
Comment thread litellm/router_strategy/complexity_router/complexity_router.py
Comment thread litellm/router_strategy/complexity_router/complexity_router.py Outdated
Comment thread litellm/router_strategy/complexity_router/complexity_router.py Outdated
Comment thread litellm/router_strategy/complexity_router/complexity_router.py Outdated
@veria-ai

veria-ai Bot commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

PR overview

All previously flagged issues have been addressed. No open security concerns remain on this pull request.

Security review

No open security issues remain on this pull request.

Fixed/addressed: 1 · PR risk: 0/10

@codecov

codecov Bot commented Sep 3, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@tin-berri
tin-berri force-pushed the litellm_tier_health_filter branch from 077d813 to 8e7bee5 Compare September 4, 2026 00:13
@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai can you reply here with a score and why that score?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

Confidence score: 4/5.

The fix now applies one health gate at every routing exit, uses the Router’s authoritative healthy-deployment lookup (including cooldown and admin-paused deployments), and fails open when health information is unavailable. The fallback remains within the decided tier and preserves modality/context-window constraints, routing-plugin filtering, adaptive-choice metadata, and the existing non-pinnable semantics. The added tests cover pinned and fresh routing, all-peer/unreadable-health cases, paused peers, plugin exclusions, adaptive routing, modality, and context-window limits.

I’m not giving 5/5 yet because the change adds a fairly broad async gate around several routing paths and the PR’s CI status is still pending. I found no remaining blocking correctness issue in the current diff; assuming the targeted tests and CI pass, this is a strong 4/5 and ready for maintainer review.

Comment thread litellm/router_strategy/complexity_router/complexity_router.py
@tin-berri
tin-berri force-pushed the litellm_tier_health_filter branch from 8e7bee5 to 33f78ae Compare September 4, 2026 00:27
@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai can you reply here with a score and why that score?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

Score: 4/5

This is a strong, focused fix. The health gate is applied at the routing hook’s exits, covers fresh picks and session-pin replays, re-picks only within the decided tier, and preserves modality/context-window constraints. It also correctly accounts for cooldown and admin-paused deployments, fails open when health data is uncertain, preserves plugin exclusions, updates adaptive-routing metadata, and prevents failed-over decisions from being pinned. The added tests cover the main regression plus the important edge cases, and the generated TypeScript union is updated.

I’m holding back the final point because the PR’s required CI status is still pending, and the end-to-end results in the description are reported evidence rather than independently verifiable checks here.

@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai CI is now fully green on 33f78ae, including codecov/patch. Given that was the only holdback, can you reply here with an updated score and why that score?

@tin-berri

Copy link
Copy Markdown
Contributor Author

Correction on my previous comment: CI is not fully green yet. 21 checks have passed and 35 are still running, none failing. codecov/patch is green. I will re-ping once everything finishes.

Comment thread litellm/router_strategy/complexity_router/complexity_router.py Outdated
Comment thread litellm/router_strategy/complexity_router/complexity_router.py Outdated
@tin-berri
tin-berri force-pushed the litellm_tier_health_filter branch from 33f78ae to 8b05890 Compare September 4, 2026 00:52
@tin-berri

Copy link
Copy Markdown
Contributor Author

Both findings had one cause: the gate hand-assembled its own eligibility view, so every filter it did not replicate became a finding. It now delegates to async_get_healthy_deployments. Net -1 helper, 3 new tests.

@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai can you reply here with a score and why that score?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

Confidence Score: 4/5

The fix is well-contained and the current implementation addresses the previously identified correctness issues:

  • Health is checked through the router’s own async_get_healthy_deployments, so cooldown, admin pause, request-scoped eligibility, and context-window filtering use the same path as actual routing.
  • Replacement selection goes through _pick_model_for_tier, preserving routing-plugin behavior and failing closed if plugins remove all candidates.
  • health_failover is treated as unpinnable, preventing transient failover choices from becoming sticky session pins.
  • The added tests cover pinned and fresh routing, fail-open behavior, paused/excluded peers, plugin filtering, adaptive metadata, modality constraints, context windows, missing tiers, and request-kwargs mutation.

I’m keeping this at 4 rather than 5 because the PR’s CI/build status is still pending, so the full validation signal is not available yet.

Comment thread litellm/router_strategy/complexity_router/complexity_router.py
Comment thread litellm/router_strategy/complexity_router/complexity_router.py
@tin-berri
tin-berri force-pushed the litellm_tier_health_filter branch from 8b05890 to 890d1ad Compare September 4, 2026 01:09
@tin-berri

Copy link
Copy Markdown
Contributor Author

Both fixed on 890d1ad. The probe now mirrors the owner's full contract: every negative verdict it can raise, and the same prompt arguments the real routing call passes. 3 new tests, each mutation-checked.

@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai can you reply here with a score and why that score?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

Score: 4/5

This is a strong, well-scoped fix. The health gate is applied at both routing exits in litellm/router_strategy/complexity_router/complexity_router.py:3076, checks capacity through the router’s existing async_get_healthy_deployments() path, preserves request-specific filters and context-window checks, and re-picks through _pick_model_for_tier() so plugins remain authoritative. The implementation also correctly avoids mutating the caller’s kwargs, forwards input, preserves modality constraints, updates the adaptive marker, and prevents health-failover decisions from being pinned.

The added coverage is unusually comprehensive, including pinned and fresh routing, cooldowns, paused/excluded/RPM-limited peers, plugin exclusions, modality and context-window constraints, missing deployments, fail-open behavior, and the Responses input path. I’m giving 4 rather than 5 because the PR’s CI checklist is still incomplete and the reported integration evidence is not independently verifiable from the review; the current status is still pending. I found no blocking correctness issue in the submitted diff.

@tin-berri

Copy link
Copy Markdown
Contributor Author

The documentation check failure is inherited from staging: it flags heuristic_v2_router_limit, a router kwarg this diff never touches, whose reference doc is tracked in a separate docs PR. Every branch cut from current staging fails this check.

@tin-berri

Copy link
Copy Markdown
Contributor Author

code-quality is the same inherited failure: it runs the identical router-settings doc validation and trips on the same staging key.

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 890d1ad. Configure here.

@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai CI has finished on 890d1ad: 75 passing, and the only 2 failures are the inherited router-settings doc validation for a staging key this diff never touches (see the flagged comments above). Bugbot's check-run on this sha concluded success with no findings. Can you reply here with an updated score and why that score?

@greptile-apps

greptile-apps Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

Updated confidence score: 4/5.

The implementation addresses the reported failure mode: it gates every decided tier placement, checks capacity through the router’s existing eligibility path, preserves modality/context-window/plugin constraints, fails open on uncertainty, and avoids pinning health substitutions. The added regression coverage is comprehensive, and CI reports 75 passing tests. The only two failures are inherited router-settings documentation validation failures for an unrelated staging key, and Bugbot completed successfully with no findings. I’m keeping this at 4/5 rather than 5/5 because the overall CI run is not fully green and this changes a core routing path, but there are no remaining findings attributable to this diff.

@tin-berri

Copy link
Copy Markdown
Contributor Author

Linked to Resolves LIT-6930. Bugbot is clean and all product checks passed; the only CI reds are the inherited router-settings documentation validation on staging.

@tin-berri

Copy link
Copy Markdown
Contributor Author

The documentation fix is already merged in BerriAI/litellm-docs#1180. I reran the two stale jobs so their checkout consumes that row.

@tin-berri
tin-berri enabled auto-merge (squash) September 4, 2026 02:20
@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai review

@tin-berri
tin-berri disabled auto-merge September 4, 2026 03:49
@tin-berri
tin-berri enabled auto-merge (squash) September 4, 2026 03:49
…r model is fully cooled down

A complexity tier can name several model groups, but the pool pick and the session-pin
replay both returned a group without consulting deployment health, so a group whose every
deployment was in cooldown was still routed to and the request died at the router's
zero-deployment check while a healthy peer sat in the same tier.

Gate the decided response at the pre-routing hook's exits, the seam the modality gate
already occupies, so every arm that can place a request is covered by one owner: a fresh
classification, a replayed or escalated pin, a plan-mode floor, a context-window
escalation, an adaptive pick, and whatever arm is added next.

Peers come from the decided tier only. Climbing to a higher tier costs more than the
classifier asked for and is left to a follow-up. The gate fails open on every uncertainty:
an unreadable cooldown view, a decision carrying no tier, a group the router knows no
deployments for, or a tier whose peers are all cooling.
@tin-berri
tin-berri force-pushed the litellm_tier_health_filter branch from 890d1ad to 6defd00 Compare September 4, 2026 22:02
@tin-berri
tin-berri requested a review from a team September 4, 2026 22:02
@tin-berri

Copy link
Copy Markdown
Contributor Author

Rebased onto df68edc and resolved the one shared _decision_is_pinnable hunk by retaining both the classifier-circuit and health-failover unpinnable rules. 20 focused tests and make check pass; merge state is now MERGEABLE.

@tin-berri
tin-berri merged commit 6dff3a5 into litellm_internal_staging Sep 4, 2026
184 of 185 checks passed
@tin-berri
tin-berri deleted the litellm_tier_health_filter branch September 4, 2026 22:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants