Skip to content

fix(proxy): derive auto-router health from its underlying models - #38174

Merged
tin-berri merged 1 commit into
litellm_internal_stagingfrom
litellm_autorouter_health_derive
Aug 26, 2026
Merged

tin-berri merged 1 commit into
litellm_internal_stagingfrom
litellm_autorouter_health_derive

Conversation

@tin-berri

@tin-berri tin-berri commented Aug 25, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Auto-routers always report healthy, whatever is behind them
  • A router with a dead tier shows green and 500s
  • Nothing tells an admin which underlying model broke

How it solves it:

  • Resolve each router's tier, default and classifier models
  • Report the router unhealthy when a dependency cannot serve
  • Name the offending model in the error text

User Flow

Before: an admin running a complexity router whose SIMPLE tier points at a broken deployment sees a healthy router, then fields 500s from users

  1. They call GET http://localhost:4000/health with the master key
  2. The router comes back in healthy_endpoints with a green badge, while the deployment its SIMPLE tier routes to sits in unhealthy_endpoints
  3. They open http://localhost:4000/ui/?page=models, Health Status tab, and the router shows healthy
  4. A user sends POST http://localhost:4000/v1/chat/completions for that router and gets a 500 reading There are no healthy deployments for this model
  5. The admin has no way to tell from the health surface which of the router's models is at fault

After: the same health call reports the router unhealthy and names the model that broke

  1. They call GET http://localhost:4000/health with the master key
  2. The router now comes back in unhealthy_endpoints with tier model 'dead-model' has no healthy deployment
  3. They open http://localhost:4000/ui/?page=models, Health Status tab, and the router shows unhealthy with that same text
  4. GET http://localhost:4000/health?model_id= answers 503 rather than 200, so monitoring pages on it
  5. A router whose models are all serving still reports healthy, so only real faults show red

Relevant issues

Linear ticket

Resolves LIT-6073

Pre-Submission checklist

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review

What makes a router report non-green

Every case below is a router reporting unhealthy. Anything not listed leaves it green, and absent information never reds a router.

  1. A tier model has no healthy deployment. Every deployment behind that name failed its own probe in the same run. Tier pools count per member, because pool selection is a blind random.choice with no health awareness, so one dead member fails that share of the tier's traffic

  2. The default model has no healthy deployment. complexity_router_default_model, auto_router_default_model or quality_router_default_model, whichever the router would actually resolve, since the params field overrides the config one

  3. The classifier model has no healthy deployment, for classifier_type: llm. A dead classifier still serves requests by silently falling back, so the routing decision is gone while the traffic looks fine

  4. A router it routes to is itself red. Verdicts settle over rounds, so a parent whose tier or default names another strategy router inherits that child's fault instead of reading the child's unprobed marker as healthy. Rounds are bounded by the marker count, so two routers pointing at each other terminate green rather than recursing

  5. A referenced model name resolves to no deployment at all. Resolution goes through every channel the request path uses, so a working alias, routing group or wildcard is not reported as missing, while an alias whose target is gone reds the router because a request through it fails the same way

Fail-open cases, which stay green:

  • No router was injected, so there is nothing to resolve names against
  • A dependency the run did not fully judge, whether the whole group is invisible to the caller or only one replica of it opted out via disable_background_health_check. A replica that was never probed can still serve what a dead sibling drops, so a fraction of the evidence never decides the verdict
  • A complexity embedding model while semantic_keyword_matching is off, since the router never calls it
  • A complexity router's complexity_router_config.default_model, which init overwrites with a tier-derived value, so only the litellm_params spelling counts (a quality router does fall back to its config field, so both count there)
  • A semantic router's routes live in an auto_router_config JSON string or an auto_router_config_path file, so only its default and embedding models are checked
  • A malformed config yields no dependencies rather than an exception, so a bad config cannot take the whole health response down

Screenshots / Proof of Fix

Config used for both runs, plus a router whose models all serve as the control. good-model is a real billed gateway deployment, dead-model points at a port with nothing on it

model_list:
  - model_name: good-model
    litellm_params: {model: openai/claude-haiku-4-5-20251001, api_base: <gateway>/v1, api_key: os.environ/GW_KEY}
  - model_name: dead-model
    litellm_params: {model: openai/gpt-4o-mini, api_base: http://127.0.0.1:9/v1, api_key: sk-bogus-dead}
  - model_name: router-tier-down
    litellm_params:
      model: auto_router/complexity_router
      complexity_router_config: {tiers: {SIMPLE: dead-model, MEDIUM: good-model, COMPLEX: good-model}}
      complexity_router_default_model: good-model
  - model_name: router-default-down
    litellm_params:
      model: auto_router/complexity_router
      complexity_router_config: {tiers: {SIMPLE: good-model, MEDIUM: good-model, COMPLEX: good-model}}
      complexity_router_default_model: dead-model
  - model_name: router-classifier-down
    litellm_params:
      model: auto_router/complexity_router
      complexity_router_config:
        tiers: {SIMPLE: good-model, MEDIUM: good-model, COMPLEX: good-model}
        classifier_type: llm
        classifier_llm_config: {model: dead-model}
      complexity_router_default_model: good-model
  - model_name: router-misconfigured
    litellm_params:
      model: auto_router/complexity_router
      complexity_router_config: {tiers: {SIMPLE: ghost-model-not-in-list, MEDIUM: good-model, COMPLEX: good-model}}
      complexity_router_default_model: good-model
  - model_name: router-broken-alias
    litellm_params:
      model: auto_router/complexity_router
      complexity_router_config: {tiers: {SIMPLE: broken-alias, MEDIUM: good-model, COMPLEX: good-model}}
      complexity_router_default_model: good-model
  - model_name: router-healthy
    litellm_params:
      model: auto_router/complexity_router
      complexity_router_config: {tiers: {SIMPLE: good-model, MEDIUM: good-model, COMPLEX: good-model}}
      complexity_router_default_model: good-model

Before (43ae350)

Full health check

  1. curl -s -H "Authorization: Bearer $KEY" http://127.0.0.1:4674/health
healthy_count: 7   unhealthy_count: 1
All 6 routers report healthy, including the four with a dependency that cannot serve
UNHEALTHY: openai/gpt-4o-mini -> OpenAIException - Connection error

Per-deployment check, which is what the Admin UI health table issues

  1. for id in ...; do curl -s -w "%{http_code}" ".../health?model_id=$id"; done
good-model               HTTP 200  1 / 0 | healthy
router-tier-down         HTTP 200  1 / 0 | healthy
router-default-down      HTTP 200  1 / 0 | healthy
router-classifier-down   HTTP 200  1 / 0 | healthy
router-misconfigured     HTTP 200  1 / 0 | healthy
router-broken-alias      HTTP 200  1 / 0 | healthy
router-healthy           HTTP 200  1 / 0 | healthy
dead-model               HTTP 503  0 / 1 | OpenAIException - Connection error

A green router cannot serve

  1. curl -s -w "HTTP %{http_code}" -X POST .../v1/chat/completions -d '{"model":"router-tier-down","messages":[{"role":"user","content":"hi"}]}'
HTTP 500
litellm.InternalServerError: OpenAIException - Connection error.. Received Model Group=router-tier-down
  1. The control router answers on the same call
curl ... -d '{"model":"router-healthy", ...}'
HTTP 200  {"id":"chatcmpl-b56d082d...","model":"router-healthy","choices":[{"message":{"content":"# Hey there!..."}}]}

After (216ddd0)

Full health check

  1. curl -s -H "Authorization: Bearer $KEY" http://127.0.0.1:4673/health
healthy_count: 2   unhealthy_count: 6
RED   -> tier model 'dead-model' has no healthy deployment
RED   -> default model 'dead-model' has no healthy deployment
RED   -> classifier model 'dead-model' has no healthy deployment
RED   -> tier model 'ghost-model-not-in-list' matches no deployment on this proxy
RED   -> tier model 'broken-alias' matches no deployment on this proxy
RED   -> tier model 'router-tier-down' has no healthy deployment   (a router whose tier is another, red, router)
GREEN -> router-healthy, tiers {SIMPLE: good-model, MEDIUM: good-model, COMPLEX: good-model}
GREEN -> router-config-default-unused, whose complexity_router_config.default_model is dead-model
         (a complexity router derives its default from the tiers and never calls that field, so it must not red)

Per-deployment check, which is what the Admin UI health table issues

  1. Same loop, same ids
good-model               HTTP 200  1 / 0 | healthy
router-tier-down         HTTP 503  0 / 1 | tier model 'dead-model' has no healthy deployment
router-default-down      HTTP 503  0 / 1 | default model 'dead-model' has no healthy deployment
router-classifier-down   HTTP 503  0 / 1 | classifier model 'dead-model' has no healthy deployment
router-misconfigured     HTTP 503  0 / 1 | tier model 'ghost-model-not-in-list' matches no deployment on this proxy
router-broken-alias      HTTP 503  0 / 1 | tier model 'broken-alias' matches no deployment on this proxy
router-healthy           HTTP 200  1 / 0 | healthy
dead-model               HTTP 503  0 / 1 | OpenAIException - Connection error
  1. Each targeted call returns exactly its own deployment, 1 / 0 or 0 / 1, so the dependencies pulled in to reach the verdict are not reported back

  2. The targeted path agrees with the full list on every case, including router-nested-parent, whose tier is another router that is itself red. Reaching that verdict needs the child's own models probed, not just the child, so expansion follows routers through routers

A green router cannot serve

  1. router-healthy still answers 200 on a real billed call through the gateway, so only routers with a real fault turned red

Type

🐛 Bug Fix

Caveats (if any)

  • Classifier fallback rate is not measured yet, only the model group
  • A dead classifier that still resolves reads green today
  • Wildcard routes match by literal name, inherited from get_model_list
  • 8 pre-existing failures in tests/litellm_utils_tests/test_health_check.py, identical on staging

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Note

Medium Risk
Changes core /health semantics and can add probes on targeted router checks; logic is complex (nested routers, partial evidence, skip-disabled) but heavily tested and fail-open when no router is injected.

Overview
Strategy-router (auto_router/*) deployments no longer show green on /health when their backing models cannot serve. Markers still skip direct LLM probes, but when a Router is passed in, health runs dependency resolution, optional extra probes, and post-processing that can move a router from healthy to unhealthy with a specific error (e.g. tier model 'dead-group' has no healthy deployment).

strategy_router_dependencies() in auto_router_model_naming.py enumerates tier, default, classifier, and embedding names per router kind (with runtime-conditional fields so dead config keys do not false-red). perform_health_check refactors targeting/eligibility into helpers, expands probes only for targeted router checks (not full-list runs), propagates nested router failures in bounded rounds, and strips dependency-only probes from the response.

Wiring: GET /health, shared Redis health checks, and background paths pass llm_router. _run_direct_health_check_with_instrumentation drops unknown optional kwargs by name instead of a fixed retry ladder.

Reviewed by Cursor Bugbot for commit 086b961. Bugbot is set up for automated code reviews on this repo. Configure here.

@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai can you reply here with a score and why that score?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Aug 25, 2026

Copy link
Copy Markdown
Contributor

Score: 4/5

The implementation is solid. Here's the breakdown:


What works well

Correctness of the core verdict logic (_dependency_failure):

visible: Final = resolved & visible_ids
if not visible or not visible <= unhealthy_ids:
    return None

The condition is right: only red a router when all visible deployments for a dependency are unhealthy. If even one is healthy, the router stays green. The fail-open cases (empty alias, unresolvable name when aliased, no router injected) are all handled correctly.

Resolution parity with the request path — using router.get_model_list() rather than a direct name lookup means the health check resolves aliases, routing groups and wildcards exactly as the request would. This avoids false positives for valid alias-resolved tiers.

Dependency probe expansion — _dependency_deployments_to_probe is a no-op on full-list runs (the common case), and only fires on targeted /health?model_id=<router> calls, which is exactly when the dependencies aren't already in scope. Drawing probes from universe (access-filtered) means a targeted check can't probe deployments the caller couldn't already see.

Defensive config parsing — strategy_router_dependencies never raises on malformed configs (wrong type, null, missing keys). Tests confirm this explicitly.

Backwards compatibility — router=None is the default everywhere. Callers that don't pass it see no behavior change.

Test coverage — 8 new focused tests: tier-down, all-healthy, missing name, caller-visibility gating, no-op on full list, probe expansion. The parametrized test_strategy_router_dependencies covers all four router kinds.


Minor concerns (not blockers)

_resolved_deployment_ids is called at least twice per dependency per run — once during _dependency_deployments_to_probe and again during _strategy_router_verdicts → _dependency_failure. For a large model list with many routers sharing tier names this is redundant work, though since health checks are infrequent it won't matter in practice. A small dict-based memo keyed on model_name inside _finalize_strategy_router_endpoints would clean this up without complexity.

The unhealthy_ids source is probed_unhealthy, not the final unhealthy_endpoints — this is fine because _finalize_strategy_router_endpoints computes verdicts before applying the keep filter, so the IDs are consistent. Just worth noting the ordering dependency is load-bearing.

Semantic router routes are unresolvable from the params dict alone (they live in a JSON string or a file path). The PR documents this and accepts it as a known limitation — the default and embedding models are still checked. This is the right call given the complexity of parsing an arbitrary JSON blob at health-check time.


The implementation is clean, well-tested, and the described before/after behavior matches what the code does. The two minor points above are optimization opportunities, not correctness issues.

@greptile-apps

greptile-apps Bot commented Aug 25, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR derives auto-router health from the deployments used by its tier, default, classifier, and embedding models.

  • Adds strategy-router dependency extraction and resolution through Router model lookup
  • Expands targeted checks with dependency probes, then removes those probes from the returned endpoint set
  • Propagates Router context through live, background, and shared health-check paths
  • Adds focused tests for healthy, unhealthy, missing, and caller-invisible dependencies

Confidence Score: 3/5

The PR should not merge until disabled dependency probes and unresolved alias targets produce health results consistent with their request-time behavior

Targeted checks can probe deployments explicitly excluded from health checks, while aliases with missing targets remain green even though requests through them fail

Files Needing Attention: litellm/proxy/health_check.py

Important Files Changed

Filename Overview
litellm/proxy/health_check.py Adds dependency probing and router verdict finalization, but reintroduces disabled deployments and mishandles aliases with missing targets
litellm/router_utils/auto_router_model_naming.py Adds typed dependency extraction for strategy-router configurations with documented fail-open limits
litellm/proxy/health_check_utils/shared_health_check_manager.py Propagates Router context through shared and fallback health checks
litellm/proxy/health_endpoints/_health_endpoints.py Supplies the shared Router to live health checks while preserving caller model scoping
litellm/proxy/proxy_server.py Supplies Router context to direct and shared background health-check execution
tests/test_litellm/proxy/test_health_check_max_tokens.py Adds focused strategy-router verdict and dependency-expansion coverage but omits disabled dependencies and unresolved alias targets
tests/test_litellm/proxy/test_shared_health_check.py Updates shared-health-check call assertions for Router propagation
tests/test_litellm/router_utils/test_auto_router_model_naming.py Exercises dependency extraction across supported strategy-router configuration shapes

Reviews (1): Last reviewed commit: "fix(proxy): derive auto-router health fr..." | Re-trigger Greptile

Comment thread litellm/proxy/health_check.py Outdated
Comment on lines +767 to +770
dependency_probes: Final = (
_dependency_deployments_to_probe(requested, universe, router) if router is not None else ()
)
model_list = requested + list(dependency_probes) # mutable-ok: _perform_health_check takes a list

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Disabled dependencies are still probed

When skip-disabled mode is enabled, dependency expansion uses the unfiltered universe, contacting disabled deployments and using their results in router health

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed. Eligibility now has one owner, applied to the requested set and the probe pool alike, so a disabled dependency cannot re-enter. Test pins it.

Comment thread litellm/proxy/health_check.py Outdated
Comment on lines +218 to +221
resolved: Final = router.get_model_list(model_name=model_name)
if not resolved and model_name not in router.model_group_alias:
return None
return frozenset(ident for entry in (resolved or ()) if (ident := _deployment_id(entry)))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Broken aliases remain healthy

When an alias target has no deployment, this returns an empty set, causing router health to stay green while requests fail

Suggested change
resolved: Final = router.get_model_list(model_name=model_name)
if not resolved and model_name not in router.model_group_alias:
return None
return frozenset(ident for entry in (resolved or ()) if (ident := _deployment_id(entry)))
resolved: Final = router.get_model_list(model_name=model_name)
if not resolved:
return None
return frozenset(ident for entry in resolved if (ident := _deployment_id(entry)))

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed, took your suggestion. An alias resolving to zero deployments is now unresolvable, not unknown. Test covers a broken alias target.

Comment thread litellm/proxy/health_check.py Outdated
dependency_probes: Final = (
_dependency_deployments_to_probe(requested, universe, router) if router is not None else ()
)
model_list = requested + list(dependency_probes) # mutable-ok: _perform_health_check takes a list

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Function parameter is rebound

Reassigning model_list violates the repository's immutable-data convention and obscures the distinction between requested and additionally probed deployments

Context Used: CLAUDE.md (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed. Narrowing and eligibility are pure helpers now, so requested and probed are separate names and the parameter is never rebound.

Comment thread litellm/proxy/health_check.py Outdated
@veria-ai

veria-ai Bot commented Aug 25, 2026 •

Copy link
Copy Markdown
Contributor

PR overview

All previously flagged issues have been addressed. No open security concerns remain on this pull request.

Security review

No open security issues remain on this pull request.

Fixed/addressed: 1 · PR risk: 0/10

Comment thread litellm/router_utils/auto_router_model_naming.py Outdated
@codecov

codecov Bot commented Aug 25, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 98.50746% with 2 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
litellm/proxy/health_check.py 98.91% 1 Missing ⚠️
litellm/proxy/proxy_server.py 85.71% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

@tin-berri
tin-berri force-pushed the litellm_autorouter_health_derive branch from 9ab1438 to 8e87537 Compare August 25, 2026 04:15
@tin-berri

Copy link
Copy Markdown
Contributor Author

Pushed fixes for all five findings. @greptileai can you reply here with a score and why that score?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Aug 25, 2026

Copy link
Copy Markdown
Contributor

Score: 4/5

The implementation is solid. Here's the full breakdown:


What's working well

  • Core verdict logic is correct. _dependency_failure's visible <= unhealthy_ids condition correctly implements "all visible deployments behind a tier name are dead → red", while a partial failure (one dead, one live) stays green because the tier can still serve through the healthy member. This matches how litellm routing actually works.

  • Fail-open semantics are consistent throughout. A dependency whose resolved IDs don't intersect visible_ids → silent. A malformed config → empty tuple, never an exception. A disabled deployment → can't re-enter through a router that depends on it. Each of these is gated at exactly one place.

  • strategy_router_dependencies correctly mirrors the runtime's conditions. Classifier model only enumerated when classifier_type: llm, embedding model only when semantic_keyword_matching is set, default model prefers the litellm_params field over the config field (matching ComplexityRouter's own override logic). dict.fromkeys dedup is a clean touch when the same name appears in both a tier and the default.

  • Targeted /health?model_id=<router> pull-in is scoped correctly. _dependency_deployments_to_probe draws from the eligible universe (already access-filtered), not the raw model_list, so a targeted check cannot expand to a deployment the caller wasn't already granted.

  • Test coverage is thorough. Tier-down, default-down, classifier-down, unresolvable name, broken alias, disabled dep opt-out, full-list no-extra-probe, and the silent-when-dep-not-visible case are all exercised as unit tests rather than integration fixtures.


Two remaining concerns (minor)

  1. frozenset[str]().union(...) in _dependency_deployments_to_probe is valid Python but unusual — frozenset[str] is a types.GenericAlias at runtime, so the subscript serves no purpose beyond what the Final[frozenset[str]] annotation already provides. frozenset().union(...) is identical at runtime and clearer.

  2. _narrow_to_target has an unguarded index on x["litellm_params"]["model"] in the model-name branch. This was pre-existing and is low-risk (every deployment in a litellm Router has litellm_params.model), but since this is now a named helper used in tests, a defensive .get would make it consistent with everything else in this module.

Neither is a blocker — the fix is correct and the scenarios that mattered before this PR (green routers silently 500ing, admin UI showing healthy for dead-tier routers, GET /health?model_id=<router> returning 200) are all resolved.

Comment thread litellm/router_utils/auto_router_model_naming.py Outdated
@tin-berri
tin-berri force-pushed the litellm_autorouter_health_derive branch from 8e87537 to fa4b563 Compare August 25, 2026 04:29
@tin-berri

Copy link
Copy Markdown
Contributor Author

Pushed fixes for the complexity default derivation plus the two cleanups. @greptileai can you reply here with a score and why that score?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Aug 25, 2026

Copy link
Copy Markdown
Contributor

Score: 4.5 / 5

Here's why:

What's strong:

  1. Fail-open semantics are correct throughout. _dependency_failure returns None (green) whenever the answer is ambiguous — dependency not visible to caller, resolves to nothing the caller can see, or opted-out via disable_background_health_check. The only red paths are ones where the runtime would actually fail. This is the right invariant.

  2. _resolved_deployment_ids delegates to router.get_model_list, which is the same resolution path live traffic uses (aliases, groups, wildcards). A mirror of any single channel would give false positives or false negatives; using the same resolver avoids that.

  3. The checked / dependency_probes split is correct. Full-list runs gain zero extra probes (all dependencies are already in checked), so /health doesn't double its provider spend. Targeted ?model_id=<router> calls pull in only the dependencies that weren't already in scope, then strip them from the response. The test test_dependency_probe_expansion_is_a_no_op_when_every_dependency_is_already_checked pins this explicitly.

  4. Backwards-compatible wiring. router: "Router | None" = None is optional on perform_health_check and perform_shared_health_check, and the entire verdict path is gated on router is not None, so callers that don't pass a router behave identically to before.

  5. Test coverage hits the right edge cases: happy path, tier down, default down, classifier down, unresolvable name, caller-visibility guard, disabled-dep guard, broken alias, probe expansion, narrowing fallback. The parametrized malformed-config test (None, "not-a-dict", {"tiers": "not-a-dict"}, etc.) confirms the health endpoint cannot be taken down by a bad config.

  6. complexity_router_default_model (litellm_params) vs config.default_model handling. The PR correctly identifies that the complexity router derives its runtime default from the tiers and overwrites config.default_model at init, so only the litellm_params spelling is a real dependency. test_complexity_ignores_its_config_default_model_and_quality_does_not exercises this distinction. I couldn't read the full complexity_router.py source (128 KB) to independently confirm the overwrite, but the test is specific enough that it would fail if the claim were wrong.

What keeps it from 5/5:

  • The tier-level visible <= unhealthy_ids condition (all visible members of the tier must be unhealthy before the tier is judged down) means a pool where one member is dead and another is healthy reports green. This is called out in the PR as intentional ("Fail-open cases"), but it's a known gap: a pool with a 50% dead rate silently fails half of requests while the router stays green. The PR description acknowledges this but doesn't quantify the traffic impact. Acceptable for now, but worth a follow-up issue.
  • Classifier fallback isn't measured — a dead classifier that falls back silently stays green. Also called out.

Neither of these is a regression from the current state; both are documented known limitations, not new bugs. The PR is ready.

Comment thread litellm/proxy/health_check.py
@tin-berri
tin-berri force-pushed the litellm_autorouter_health_derive branch from fa4b563 to 95f2908 Compare August 25, 2026 04:58
@tin-berri

Copy link
Copy Markdown
Contributor Author

Fixed nested-router propagation. @greptileai can you reply here with a score and why that score?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Aug 25, 2026

Copy link
Copy Markdown
Contributor

4.5 / 5

The nested-router propagation fix is correct and the overall implementation is solid. Here's the breakdown:

What's working well

  • _settle_router_verdicts terminates correctly in all cases. A cycle (A → B → A) produces no fresh entries in round 1 (neither has a probe failure, and neither is in settled yet), so fresh is empty and the function returns settled unchanged — green, not infinite. The bound of len(markers) is tight.
  • The propagation chain (child red → child's ID enters unhealthy_now in round 2 → parent sees it via _dependency_failure) is exactly right, and test_a_router_routing_to_a_red_router_is_itself_red proves it end-to-end.
  • _dependency_failure's fail-open on absent visibility (not visible → None) is the right default. An opted-out dependency can't re-enter through a router that depends on it because _dependency_deployments_to_probe draws from _health_check_eligible(model_list, skip_disabled).
  • Resolution goes through router.get_model_list, which mirrors every channel the actual request path uses (alias, routing group, wildcard). That's the correct mirror.
  • Test suite is comprehensive — positive, negative, and edge-case classes all present.

One remaining concern

_settle_router_verdicts is recursive, not iterative. It is bounded by len(markers) so it terminates, but each call is a real Python stack frame. A deployment with ~900+ router markers in a single health check run would hit CPython's default recursion limit and raise a RecursionError. That's an extreme case today, but converting to a while loop would eliminate the limit entirely:

def _settle_router_verdicts(
    markers: Mapping[str, Mapping[str, object]],
    router: "Router",
    visible_ids: frozenset[str],
    probe_failures: frozenset[str],
    settled: Mapping[str, str],
    rounds_left: int,
) -> Mapping[str, str]:
    while rounds_left > 0:
        unhealthy_now: Final = probe_failures | frozenset(settled)
        fresh = {
            marker_id: error
            for marker_id, deployment in markers.items()
            if marker_id not in settled
            if (error := _strategy_router_dependency_error(deployment, router, visible_ids, unhealthy_now))
        }
        if not fresh:
            return settled
        settled = MappingProxyType({**settled, **fresh})
        rounds_left -= 1
    return settled

That's the only gap I'd flag. Everything else — the dependency extraction logic, the targeted-check probe expansion, the cycle termination proof, the test coverage — is in good shape.

Comment thread litellm/proxy/health_check.py Outdated
@codspeed

codspeed Bot commented Aug 25, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_autorouter_health_derive (086b961) with litellm_internal_staging (f57e4b8)

Open in CodSpeed

@tin-berri
tin-berri force-pushed the litellm_autorouter_health_derive branch from 95f2908 to 8b6c918 Compare August 26, 2026 19:39
@tin-berri

Copy link
Copy Markdown
Contributor Author

Fixed transitive probe expansion for nested routers. @greptileai can you reply here with a score and why that score?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Aug 26, 2026

Copy link
Copy Markdown
Contributor

Confidence score: 5/5.

The transitive probe expansion fixes the remaining targeted-check gap: _reachable_dependency_ids recursively follows dependencies, so a parent router pulls in the child router and the child’s underlying deployments before verdict settlement. Expansion is bounded by rounds_left and stops when no new IDs are found, so router cycles terminate safely. The added tests cover grandchild probing and cycle termination, while the nested-router verdict test confirms that a red child propagates to its parent. Buildkite e2e status is passing, and I have no remaining blocking concerns.

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 8b6c918. Configure here.

@tin-berri
tin-berri force-pushed the litellm_autorouter_health_derive branch from 8b6c918 to 216ddd0 Compare August 26, 2026 19:49
@tin-berri

Copy link
Copy Markdown
Contributor Author

Rebased onto current staging, conflicts resolved. @greptileai can you reply here with a score and why that score?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Aug 26, 2026

Copy link
Copy Markdown
Contributor

Score: 5/5

The rebase does not introduce any new correctness concerns, and the previously identified issues are addressed:

  • _dependency_failure correctly reds a router only when every visible deployment for a dependency is unhealthy, while preserving fail-open behavior for deployments outside the caller’s scope.
  • Targeted checks now use the same eligibility filter as the requested set, so disable_background_health_check deployments cannot be reintroduced or probed (litellm/proxy/health_check.py:212).
  • Empty Router.get_model_list() results are treated as unresolved dependencies, including broken aliases (litellm/proxy/health_check.py:233).
  • Nested router dependencies are expanded transitively and verdicts settle over bounded rounds, covering parent/child propagation without infinite recursion (litellm/proxy/health_check.py:275, litellm/proxy/health_check.py:330).
  • Dependency extraction now matches runtime behavior for complexity defaults, quality defaults, classifier gating, and conditional embedding use (litellm/router_utils/auto_router_model_naming.py:76).
  • Router context is wired through live, background, and shared health-check paths, while targeted responses omit dependency-only probe results.
  • The added tests cover the key failure modes, visibility and opt-out semantics, aliases, nested routers, cycles, and malformed configurations.

No blocking issues remain from this review. The current CI status is still pending, so this score is based on the implementation and test coverage rather than a completed CI run.

@greptile-apps

greptile-apps Bot commented Aug 26, 2026

Copy link
Copy Markdown
Contributor

Confidence score: 4/5

The rebased commit correctly addresses the reported behavior. The implementation:

  • derives dependencies for the supported strategy-router types and respects runtime gates such as classifier type, semantic matching, and complexity-vs-quality default semantics;
  • uses Router.get_model_list() for aliases, routing groups, and wildcard resolution;
  • preserves fail-open behavior for dependencies outside the caller’s visible scope and for disabled background checks;
  • propagates failures through nested routers with bounded multi-round settlement, while terminating safely for cycles;
  • expands only targeted checks transitively and removes dependency-only probes from the response, without adding probes to full-list checks;
  • wires the router through the live, background, and shared health-check paths.

The added tests cover the important positive, negative, malformed-config, alias, visibility, disabled-dependency, nested-router, cycle, and targeted-probe cases. I’m giving 4/5 rather than 5/5 because this changes core /health semantics and the current CI status is still pending, so final confidence should be raised after the required checks complete.

Comment thread litellm/proxy/health_check.py
@tin-berri
tin-berri force-pushed the litellm_autorouter_health_derive branch from 216ddd0 to 4004325 Compare August 26, 2026 20:52
@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptile can you reply here with a score and why

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Aug 26, 2026

Copy link
Copy Markdown
Contributor

Score: 4/5

This is a strong, well-scoped fix. It derives strategy-router health from the models the routing path actually resolves, distinguishes unhealthy dependencies from missing ones, respects disabled or partially judged deployments, propagates failures through nested routers, and bounds cycles. The tests cover tiers, defaults, classifiers, aliases, malformed configs, partial evidence, nested routers, and cycles. It also correctly avoids extra probes during full-list checks while expanding dependencies for targeted router checks.

I’m not giving 5/5 because the change substantially alters /health semantics, while the most important coverage is concentrated in helper-level tests rather than complete endpoint and shared/background health integration tests. That leaves some regression risk around cached results and unusual router-resolution shapes. The PR description’s proof and test breadth nevertheless provide good evidence that the intended cases are fixed.

if router is not None
else ()
)
checked: Final = requested + list(dependency_probes) # mutable-ok: _perform_health_check takes a list

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Targeted probes skip ID deduplication

Low Severity

requested goes through filter_deployments_by_id, but dependency_probes are taken from the undeduped universe via _deployments_by_id. Duplicate rows that share an id (the alias case that filter already exists for) get probed twice on the targeted /health?model_id= path. Divergent results for the same id can also put it in probe_failures while a sibling probe succeeded, which reds the router.

Additional Locations (2)
Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit 4004325. Configure here.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confirmed and fixed at the primitive rather than the call site.

_deployments_by_id is the one answer to "give me the deployments for these ids", and it was the only path that did not honour the one-row-per-id rule filter_deployments_by_id exists for. It now calls that function instead of mirroring a weaker version of it:

matched: Final = tuple(d for d in universe if (uid := _deployment_id(d)) and uid in ids)
return tuple(filter_deployments_by_id(model_list=matched))

Fixing it there covers both of its callers — the returned probe set and the sweep's frontier — so the fix cannot be reintroduced by a future third caller. Both halves you flagged go with it: the duplicate row is no longer probed twice, and one id can no longer land in probe_failures from one probe while a sibling probe of the same id succeeded.

Test: test_dependency_probes_carry_one_row_per_id builds a universe with dead-1 present twice and asserts the probe set carries it once. It fails on the previous commit.

last_type_error: TypeError | None = None
for extra_kwargs in (
{
"router": llm_router,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Retry fallback drops skip-disabled filter

Low Severity

router was folded into the first perform_health_check kwargs set, but the next fallback is only instrumentation_context. A TypeError on the new router argument now retries without health_check_skip_disabled_background_models, so opted-out deployments get probed on the background fallback path.

Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit 4004325. Configure here.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Correct, and it exposed the real defect: the ladder was a hand-maintained power set of kwarg combinations, so every argument added to perform_health_check needs a new rung or it opens exactly this kind of hole. Guarding it with one more rung would have left the next argument to reintroduce the bug.

The ladder is gone. The TypeError already names the argument the callee rejected, so drop that one and keep the rest:

optional: Mapping[str, object] = MappingProxyType(
    {
        "router": llm_router,
        "instrumentation_context": instrumentation_context,
        **health_check_filter_kwargs_from_general_settings(general_settings),
    }
)
for _ in range(len(optional) + 1):
    try:
        return await perform_health_check(..., **optional)
    except TypeError as e:
        rejected = _UNEXPECTED_KWARG.search(str(e))
        if rejected is None or rejected["name"] not in optional:
            raise
        optional = MappingProxyType({k: v for k, v in optional.items() if k != rejected["name"]})

A callee that predates router now loses router and nothing else, so health_check_skip_disabled_background_models survives and opted-out deployments stay unprobed. The rejected["name"] not in optional test keeps the previous re-raise behaviour for a TypeError that is not about one of these arguments, so _is_unexpected_keyword_argument_type_error was deleted rather than replaced.

Tests: the three existing rung tests (legacy three-arg stub, instrumentation-only, filter-only) all still pass, since one-at-a-time removal reaches each of those shapes. New test_run_direct_health_check_drops_only_the_rejected_kwarg covers the case you found — a stub accepting instrumentation_context and the filter but not router — and asserts the callee receives (instrumentation_context, True). It fails on the previous commit.

An auto_router deployment is a marker, not something a probe can contact, so
`_run_model_health_check` returns `{}` for it and it lands healthy whatever is
behind it. This derives its verdict from the models it actually resolves.

Rules and owners:

- `strategy_router_dependencies` is the single answer to "what does this router
  call": tier, default, classifier and embedding names per router kind, aligned
  with what init and the request path actually use.
- `_health_check_eligible` is the single probe-eligibility gate, applied to the
  requested set and to the pool a router's dependencies are drawn from alike, so
  an opted-out deployment cannot re-enter through a router that depends on it.
- `_resolved_deployment_ids` resolves names through `get_model_list`, the same
  composition of alias, routing-group and wildcard channels a request uses.
- A dependency reds its router only when *every* deployment behind the name is
  known unhealthy. A replica this run never judged, hidden from the caller or
  opted out of health checks, can still serve what the dead one drops, so
  partial evidence leaves the verdict green. Absent information never reds.
- Verdicts settle over rounds, because a marker never fails a probe of its own
  and a parent whose tier is a red router must inherit that fault. Both sweeps
  are bounded loops, so a router cycle terminates green.
- Dependency probes are added only on the targeted `/health?model_id=` path the
  dashboard uses per deployment, and are dropped from the response.

Resolves LIT-6073

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@tin-berri
tin-berri force-pushed the litellm_autorouter_health_derive branch from 4004325 to 086b961 Compare August 26, 2026 21:43
@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptile can you reply here with score & why

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Aug 26, 2026

Copy link
Copy Markdown
Contributor

Score: 4/5 (high confidence).

Why:

  • Correctly derives strategy-router health from tier, default, classifier, embedding, and nested-router dependencies.
  • Uses Router.get_model_list, preserving request-path resolution for aliases, routing groups, and wildcards, while distinguishing missing deployments from unhealthy ones.
  • Preserves fail-open behavior for partial/unjudged evidence and opted-out deployments.
  • Targeted checks expand dependency probes without returning those extra deployments; full-list checks avoid duplicate probes.
  • Bounded transitive expansion and verdict rounds handle nested routers and cycles safely.
  • Tests cover healthy/unhealthy dependencies, missing and broken aliases, malformed configs, default precedence, opt-outs, duplicate IDs, nested routers, cycles, and keyword compatibility.

I’m not giving 5/5 because this changes /health status/error semantics and targeted checks can broaden provider probing, creating meaningful operational blast radius. The dynamic unexpected-keyword compatibility path and deployment-ID-based verdict aggregation also merit maintainer validation. I found no blocking correctness issue in the reviewed diff.

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

2 issues from previous reviews remain unresolved.

Fix All in Cursor

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 086b961. Configure here.

@tin-berri
tin-berri enabled auto-merge (squash) August 26, 2026 23:37
@tin-berri
tin-berri merged commit d8edfb6 into litellm_internal_staging Aug 26, 2026
82 checks passed
@tin-berri
tin-berri deleted the litellm_autorouter_health_derive branch August 26, 2026 23:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants