Skip to content

fix(router): don't cool down parent deployment on advisor sub-call failure - #33792

Merged
yassin-berriai merged 4 commits into
litellm_internal_stagingfrom
litellm_lit_4565_advisor_subcall_cooldown
Jul 25, 2026
Merged

yassin-berriai merged 4 commits into
litellm_internal_stagingfrom
litellm_lit_4565_advisor_subcall_cooldown

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Jul 17, 2026 •

Copy link
Copy Markdown
Contributor

Relevant issues

Linear ticket

Resolves LIT-4565

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • My PR passes all CI/CD checks (e.g., lint, format, unit tests)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Screenshots / Proof of Fix

Live proxy, real Fireworks executor + real Anthropic advisor sub-call (no mocks). The proxy is started with no ANTHROPIC_API_KEY so the advisor sub-call gets a genuine 401, which is the condition from the ticket. claude-sonnet-5 maps to a single non-native deployment (fireworks_ai/accounts/fireworks/models/gpt-oss-120b); the advisor tool resolves to api.anthropic.com directly, the same non-native path Bedrock takes

Config used

model_list:
  - model_name: claude-sonnet-5
    litellm_params:
      model: fireworks_ai/accounts/fireworks/models/gpt-oss-120b
      api_key: os.environ/FIREWORKS_AI_API_KEY
router_settings:
  num_retries: 0
  cooldown_time: 30
  allowed_fails: 0

Both runs send the same two requests back to back: first an advisor request (forced tool_choice so the advisor sub-call always fires), then an unrelated well-formed request to the same model group

# 1) advisor request (advisor sub-call 401s because the proxy has no anthropic key)
curl -s -w "\nHTTP %{http_code}\n" http://localhost:4000/v1/messages -H "Content-Type: application/json" -d '{"model":"claude-sonnet-5","max_tokens":1024,"messages":[{"role":"user","content":"What is 2+2? Ask the advisor first."}],"tools":[{"type":"advisor_20260301","name":"advisor","model":"claude-opus-4-8"}],"tool_choice":{"type":"tool","name":"advisor"}}'

# 2) UNRELATED normal request to the same model group, immediately after
curl -s -w "\nHTTP %{http_code}\n" http://localhost:4000/v1/messages -H "Content-Type: application/json" -d '{"model":"claude-sonnet-5","max_tokens":64,"messages":[{"role":"user","content":"Say hi in one word."}]}'

Before the fix, at base commit c5b4456401

# 1) advisor request
{"error":{"message":"{\"type\":\"error\",\"error\":{\"type\":\"authentication_error\",\"message\":\"x-api-key header is required\"}...}. Received Model Group=claude-sonnet-5 ...","code":"401"}}
HTTP 401

# 2) unrelated request
{"error":{"message":"No deployments available for selected model, Try again in 30 seconds. Passed model=claude-sonnet-5. pre-call-checks=False, cooldown_list=['23237369...39f8e']","code":"429"}}
HTTP 429

The advisor sub-call's 401 cooled down the healthy parent deployment, so the unrelated caller is rejected with a 429 naming the deployment in cooldown_list

After the fix, at commit 28b46ae93c

# 1) advisor request
{"error":{"message":"{\"type\":\"error\",\"error\":{\"type\":\"authentication_error\",\"message\":\"x-api-key header is required\"}...}. Received Model Group=claude-sonnet-5 ...","code":"401"}}
HTTP 401

# 2) unrelated request
{"id":"chatcmpl-ad0cd614...","type":"message","role":"assistant","model":"claude-sonnet-5","content":[...,{"type":"text","text":"Hi"}],"stop_reason":"end_turn"}
HTTP 200

The advisor request still surfaces its own 401 (that is the advisor being misconfigured, unchanged), but the parent deployment is no longer cooled down, so the unrelated caller succeeds with a 200

Type

🐛 Bug Fix

Changes

The advisor orchestration handler issues a sub-call to the advisor model, which resolves to a different provider and credentials than the deployment the router selected for the parent request. When that sub-call fails, for example a 401 because no advisor API key is configured, the exception propagates up to Router.deployment_callback_on_failure, which keys cooldown off the parent deployment's model_info.id and never consulted whether the failure actually came from the deployment. The healthy deployment gets cooled down and unrelated callers to the same model group are rejected

A failure that originates from advisor orchestration now carries a signal to that effect. The sentinel and its mark_advisor_orchestration_failure / is_advisor_orchestration_failure predicate live in litellm/router_utils/cooldown_handlers.py, which is provider-neutral and already imported by the router, so deployment_callback_on_failure reads the tag via a module-top import rather than reaching into an Anthropic module. AdvisorOrchestrationHandler.handle tags two orchestration-level failures: the advisor sub-call (wrapped in a try/except) and AdvisorMaxIterationsError when the loop exhausts max_uses while the executor keeps succeeding. deployment_callback_on_failure returns early without incrementing failures or setting cooldown when the tag is present. The exception is tagged rather than wrapped, so its type, the router's retry and fallback classification, and the client-facing error are all unchanged. Genuine executor or deployment failures are never tagged and keep cooling down exactly as before

Regression tests

tests/test_litellm/llms/anthropic/messages/test_advisor_orchestration.py asserts that a failing advisor sub-call and a max-iterations failure both propagate a tagged exception, while an executor failure does not get tagged

tests/test_litellm/test_router.py::TestAdvisorSubCallCooldown asserts that deployment_callback_on_failure cools the deployment down for an untagged auth error but skips cooldown for a tagged advisor orchestration failure

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Link to Devin session: https://app.devin.ai/sessions/d3f8fab3d4fa460887ced329ee7c1b44
Requested by: @shivamrawat1

…ilure

Advisor orchestration issues a sub-call to a different provider/credentials than the selected deployment. When that sub-call fails (e.g. a 401 because no advisor API key is configured), the exception propagates up and the router's deployment_callback_on_failure attributes it to the healthy parent deployment's model_info.id, cooling it down and rejecting unrelated callers to the same model group.

Tag advisor sub-call failures on the exception and skip cooldown for them in deployment_callback_on_failure. The exception is tagged rather than wrapped so its type is preserved and retry/fallback classification and the client-facing error are unchanged. Genuine executor/deployment failures are untagged and still cool down as before.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@shivamrawat1 shivamrawat1 self-assigned this Jul 17, 2026
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR. Add '(aside)' to your comment to have me ignore it.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@greptile-apps

greptile-apps Bot commented Jul 17, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR prevents the router from cooling down a healthy deployment when an advisor sub-call (which targets a different provider/credentials) fails. The fix tags the exception at the point of failure and short-circuits deployment_callback_on_failure before cooldown is applied.

  • advisor.py exposes mark_advisor_sub_call_failure / is_advisor_sub_call_failure helpers and wraps only the advisor sub-call in a try/except that applies the tag; executor failures propagate untagged.
  • router.py imports the advisor-specific helper directly into deployment_callback_on_failure and returns early without cooldown when the tag is present; the check occurs before any failure counter is incremented.
  • Two new test classes cover both directions of the tagging check and verify the cooldown list membership in the router.

Confidence Score: 4/5

The core fix is correct and well-tested; the main open question is whether AdvisorMaxIterationsError should also be exempt from incrementing the failure counter.

The tagged-exception approach is sound and the two test classes exercise the happy and unhappy paths. The AdvisorMaxIterationsError path goes untagged, which can still increment the failure counter for the healthy deployment over repeated max-iteration errors, albeit not trigger full cooldown unless allowed_fails is reached. Importing an Anthropic-specific helper into the generic router is a minor coupling concern.

litellm/router.py and advisor.py — specifically the untagged AdvisorMaxIterationsError path and the provider-specific import in the router.

Important Files Changed

Filename Overview
litellm/llms/anthropic/experimental_pass_through/messages/interceptors/advisor.py Adds tagging helpers and wraps the advisor sub-call in try/except to mark failures; AdvisorMaxIterationsError is not tagged so the failure counter can still be incremented for that path.
litellm/router.py Adds an early-return guard in deployment_callback_on_failure for tagged advisor sub-call failures; imports an Anthropic-specific advisor helper directly into the generic router.
tests/test_litellm/llms/anthropic/messages/test_advisor_orchestration.py Adds two new async tests verifying that advisor sub-call failures are tagged while executor failures are not; tests use mocks with no real network calls.
tests/test_litellm/test_router.py Adds TestAdvisorSubCallCooldown with two tests covering the tagged and untagged paths through deployment_callback_on_failure; correctly verifies cooldown list membership.

Comments Outside Diff (1)

  1. litellm/llms/anthropic/experimental_pass_through/messages/interceptors/advisor.py, line 145-150 (link)

    P2 AdvisorMaxIterationsError not tagged — failure counter still incremented

    When the advisor loop exceeds max_uses, AdvisorMaxIterationsError propagates without the advisor-sub-call tag. The router's deployment_callback_on_failure will call increment_deployment_failures_for_current_minute for the parent deployment even though the executor succeeded on every iteration. Over multiple such timeouts on the same model group, the failure counter can reach allowed_fails, cooling down the healthy deployment — the same class of bug this PR fixes for the sub-call path.

Reviews (1): Last reviewed commit: "fix(router): don't cool down parent depl..." | Re-trigger Greptile

Comment thread litellm/router.py Outdated
Comment on lines +7005 to +7014
from litellm.llms.anthropic.experimental_pass_through.messages.interceptors.advisor import (
is_advisor_sub_call_failure,
)

if is_advisor_sub_call_failure(exception):
verbose_router_logger.debug(
"Router: Exiting 'deployment_callback_on_failure' without cooldown. "
"Failure originated from an advisor sub-call, not the selected deployment."
)
return False

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Provider-specific import in generic router component

Importing is_advisor_sub_call_failure from a deeply-nested Anthropic advisor module inside deployment_callback_on_failure tightly couples the generic router to a single provider's orchestration feature. Any other interceptor (Bedrock, Vertex, etc.) that wants the same "skip cooldown" guarantee would need a separate import here. A more maintainable pattern would be to check a shared, generic exception attribute (e.g., a utility in litellm/utils.py or a base interceptor sentinel) so the router stays provider-agnostic.

Rule Used: What: Avoid writing provider-specific code outside... (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

Runtime test results (LIT-4565)

Tested on a live litellm proxy at localhost:4000, real Fireworks executor + real Anthropic advisor sub-call, no mocks. Proxy started with ANTHROPIC_API_KEY unset so the advisor sub-call gets a genuine 401 (the ticket condition). Same two requests sent back to back at base and at the fix commit.

  • Advisor sub-call still returns HTTP 401 (x-api-key header is required): PASSED
  • Unrelated request to the same model group returns HTTP 200 after the fix (HTTP 429 before): PASSED
After the fix (commit 28b46ae93c) - live run

Advisor request 401s (unchanged), but the unrelated request now returns 200; the parent deployment is no longer cooled down.

after

Before the fix (base c5b4456401) - captured live

The advisor sub-call 401 cooled down the healthy parent deployment, so the unrelated request was rejected with HTTP 429 No deployments available ... cooldown_list=[...].

before

Executor is Fireworks gpt-oss-120b rather than Bedrock because the test Bedrock IAM role cannot InvokeModel; Fireworks is a non-native provider too, so it exercises the identical AdvisorOrchestrationHandler + deployment_callback_on_failure path.

Full test session: https://app.devin.ai/sessions/d3f8fab3d4fa460887ced329ee7c1b44

@codecov

codecov Bot commented Jul 17, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@codspeed

codspeed Bot commented Jul 17, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_lit_4565_advisor_subcall_cooldown (8f8d5f8) with litellm_internal_staging (010b200)1

Open in CodSpeed

Footnotes

  1. No successful run was found on litellm_internal_staging (0439bcb) during the generation of this report, so 010b200 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

…_subcall_cooldown

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

@yassin-berriai yassin-berriai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the fix end to end. The core change is correct, minimal, and the live-proxy before/after is exactly the right kind of proof. A few notes below

What holds up

The tagging mechanism survives end to end. Logging._failure_handler_helper_fn stores the same exception object via self.model_call_details["exception"] = exception (litellm_logging.py:2698), which is what deployment_callback_on_failure later reads through kwargs.get("exception"), so the setattr tag is preserved by object identity. The 429 to 200 flip in the proof confirms it in practice

Scoping is right. The try/except wraps only the advisor sub-call (advisor.py:136-150); the executor call and its re-calls after injection sit outside it, so genuine deployment failures stay untagged and keep cooling down. test_executor_failure_is_not_tagged locks that in

deployment_callback_on_failure is the only cooldown-on-failure path this scenario hits. async_deployment_callback_on_failure only bumps RPM, and the other _set_cooldown_deployments at router.py:7304 fires solely on a pre-call RateLimitError, so the single fix location is complete for this case

The tests are real regressions with good mutation kill: removing the setattr, widening the try to cover the executor, or dropping the early return each break a test, and the positive test_untagged_auth_error_cools_down_deployment guards that normal cooldown still works

Findings

  1. AdvisorMaxIterationsError is untagged, so the same healthy-deployment attribution can still happen through the max-iterations path. It is raised at advisor.py:127, outside the try/except, even though the executor succeeded on every iteration. deployment_callback_on_failure then calls increment_deployment_failures_for_current_minute for the parent, and over repeated max-iteration errors on the same group the counter can reach allowed_fails and cool the healthy deployment down, which is the same class of bug this PR fixes for the sub-call path. It is an orchestration-level failure rather than a deployment-health one, so it belongs in the same exemption; tagging it before the raise closes the gap. This is the one substantive item, and it lines up with the Greptile 4/5 open question

  2. The generic router now imports an Anthropic-specific helper, and does it in the function body. deployment_callback_on_failure adds a from litellm.llms.anthropic...advisor import is_advisor_sub_call_failure inside the method. Two things here: the repo convention is no new in-function imports (the only sanctioned exception is proxy-only deps, which this is not), and a provider-agnostic router taking a dependency on an anthropic module is the coupling Greptile flagged. Both resolve together by moving the sentinel plus the mark_/is_ predicate into a provider-neutral util that the router imports at module top. If the in-function placement was there to avoid a circular import, that relocation also removes the cycle

  3. Minor test nit: test_untagged_auth_error_cools_down_deployment is marked @pytest.mark.asyncio and declared async but awaits nothing and calls a sync method, while its sibling is a plain sync def doing the same kind of work. Not a problem, just inconsistent

CI

The two red shards, proxy-server and proxy-endpoints, are not from this diff. The failures are ValueError: not enough values to unpack (expected 2, got 0) out of router.get_configured_token_limits via litellm/proxy/utils.py:6139 in test_team_model_name_translation.py, code this PR does not touch. Newer open PRs pass these same shards, so this is staging drift captured at this PR's merge ref; a rebase onto current litellm_internal_staging and a re-run should clear both. Worth flipping the "passes all CI/CD checks" checklist item back until it is green

Net: fix is correct and well-tested. Addressing finding 1 (tag AdvisorMaxIterationsError, with a test covering that path) is what I would gate on; findings 2 and 3 are quality

shivamrawat1 and others added 2 commits July 18, 2026 19:10
…tral util

Address review on LIT-4565: move the cooldown-exemption marker into
litellm/router_utils/cooldown_handlers.py so the router imports it at
module top instead of an in-function anthropic import, and extend the
exemption to AdvisorMaxIterationsError so a max-iterations orchestration
failure no longer cools down the healthy executor deployment.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you all sign our Contributor License Agreement before we can accept your contribution.
1 out of 2 committers have signed the CLA.

✅ shivamrawat1
❌ devin-ai-integration[bot]
You have signed the CLA already but the status is still pending? Let us recheck it.

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

Thanks for the thorough review @yassin-berriai. Addressed in 99ddba0 (plus a merge of current litellm_internal_staging)

Finding 1: AdvisorMaxIterationsError is now tagged. It is built, marked with mark_advisor_orchestration_failure, then raised at the max_uses check, so a loop that exhausts iterations while the executor keeps succeeding no longer feeds the parent's failure counter. Covered by test_advisor_max_iterations_failure_is_tagged, which drives the loop purely via max_uses (executor always asks for the advisor, advisor always succeeds) and asserts the raised error is tagged

Finding 2: the sentinel and the mark_/is_ predicate moved into litellm/router_utils/cooldown_handlers.py, which is provider-neutral and already imported by the router. deployment_callback_on_failure now imports is_advisor_orchestration_failure at module top and the in-function anthropic import is gone; advisor.py imports mark_advisor_orchestration_failure from the same place. I renamed the pair from *_sub_call_* to *_orchestration_* since it now covers both the sub-call and the max-iterations path

Finding 3: I kept test_untagged_auth_error_cools_down_deployment async on purpose. Unlike its sibling, it exercises the real cooldown path, and _set_cooldown_deployments schedules router_cooldown_event_callback via asyncio.create_task (cooldown_handlers.py:303), which needs a running loop; the sibling returns early before reaching that, so it stays sync. Happy to unify them if you'd rather, but the marker is load-bearing here

CI: merged current litellm_internal_staging (8f8d5f8) so the proxy-server / proxy-endpoints drift you flagged should clear on the re-run; I'll confirm once the shards report

@yassin-berriai
yassin-berriai merged commit 96f58fa into litellm_internal_staging Jul 25, 2026
78 checks passed
@yassin-berriai
yassin-berriai deleted the litellm_lit_4565_advisor_subcall_cooldown branch July 25, 2026 17:17
Ericcwang23 pushed a commit to Ericcwang23/litellm that referenced this pull request Jul 27, 2026
…ilure (BerriAI#33792)

* fix(router): don't cool down parent deployment on advisor sub-call failure

Advisor orchestration issues a sub-call to a different provider/credentials than the selected deployment. When that sub-call fails (e.g. a 401 because no advisor API key is configured), the exception propagates up and the router's deployment_callback_on_failure attributes it to the healthy parent deployment's model_info.id, cooling it down and rejecting unrelated callers to the same model group.

Tag advisor sub-call failures on the exception and skip cooldown for them in deployment_callback_on_failure. The exception is tagged rather than wrapped so its type is preserved and retry/fallback classification and the client-facing error are unchanged. Genuine executor/deployment failures are untagged and still cool down as before.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(router): tag advisor orchestration failures via provider-neutral util

Address review on LIT-4565: move the cooldown-exemption marker into
litellm/router_utils/cooldown_handlers.py so the router imports it at
module top instead of an in-function anthropic import, and extend the
exemption to AdvisorMaxIterationsError so a max-iterations orchestration
failure no longer cools down the healthy executor deployment.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants