Skip to content

fix(azure): propagate asyncio.CancelledError instead of raising a 500 - #42295

Merged
mateo-berri merged 9 commits into
mainfrom
litellm_fix_azure_cancellederror_cooldown
Sep 22, 2026
Merged

mateo-berri merged 9 commits into
mainfrom
litellm_fix_azure_cancellederror_cooldown

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Supersedes #35330 at its head 70be37a: the same commits and authors, moved to a litellm_ branch so CircleCI runs the legacy provider suites on it

TLDR

Problem this solves:

  • Azure's async handler turned task cancellation into a fake 500
  • With cancel_on_disconnect, one client hang-up became a deployment failure
  • The router benched the healthy Azure deployment and fell back
  • Logs showed AzureException APIError 500s Azure never returned

How it solves it:

  • Re-raise asyncio.CancelledError instead of AzureOpenAIError(500)
  • Keeps the existing post_call logging for observability
  • Unit regression test that fails without the fix
  • Live e2e test: a client hang-up must not bench the Azure deployment
  • e2e harness gains Transport.abandon, which closes the socket mid-answer
  • The cell retries the hang-up when the model answers inside the window

User Flow

Before: a platform team running cancel_on_disconnect: true in front of an Azure model group watches one impatient client take a healthy Azure deployment out of rotation

  1. Their app sends POST https://litellm-domain/v1/chat/completions with "model": "azure-group" and a prompt whose answer takes about a minute
  2. The app's own timeout fires a few seconds in and it closes the connection before any answer arrives
  3. On https://litellm-domain/ui/?page=logs that request shows as a 500 AzureException APIError with an empty error body, booked against the Azure deployment that was serving it
  4. The next call, POST https://litellm-domain/v1/chat/completions with the same "model": "azure-group", comes back 200 with an x-litellm-model-id header naming a different deployment: the one the client hung up on is in cooldown
  5. Every call for the length of the cooldown lands on the fallback deployment, while Azure's own metrics show no failed request

After: the same hang-up is recorded as the client's doing and the Azure deployment keeps serving

  1. Their app sends POST https://litellm-domain/v1/chat/completions with "model": "azure-group" and a prompt whose answer takes about a minute
  2. The app's own timeout fires a few seconds in and it closes the connection before any answer arrives
  3. On https://litellm-domain/ui/?page=logs that request shows as a 499 Client disconnected the request and is not counted against any deployment
  4. The next call, POST https://litellm-domain/v1/chat/completions with the same "model": "azure-group", comes back 200 with x-litellm-model-id naming the same Azure deployment the client hung up on
  5. Every following call keeps landing on that Azure deployment and nothing is in cooldown

Relevant issues

Fixes #35329
Fixes #42222

Affected release

Linear ticket

Resolves LIT-8248

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/llms/azure/test_azure.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Two local proxies, two uvicorn workers each with no Redis, on random free ports: http://127.0.0.1:43305 booted from the merge base tree (a83773c) and http://127.0.0.1:34701 from this branch (70be37a). Both run this config, each against its own Postgres:

general_settings:
  store_model_in_db: true
  cancel_on_disconnect: true
  proxy_batch_write_at: 5
  database_connection_pool_limit: 5

litellm_settings:
  drop_params: true
  request_timeout: 600
  num_retries: 3

router_settings:
  routing_strategy: simple-shuffle
  num_retries: 3
  allowed_fails: 5
  cooldown_time: 30

model_list:
  - model_name: gpt-5.5
    litellm_params:
      model: openai/gpt-5.5
      api_key: os.environ/OPENAI_API_KEY

Every case registers its own group through the API: a real Azure gpt-5.4-nano deployment holding all the weight and benched on its first failure, plus a zero-weight OpenAI backup, which is needed because a single-deployment group never cools down. The key is limited to one parallel request and a $0.05 budget:

POST /model/new     {"model_name": "<grp>", "litellm_params": {"model": "azure/gpt-5.4-nano", "api_key": "os.environ/AZURE_API_KEY", "api_base": "os.environ/AZURE_API_BASE", "api_version": "2024-10-21", "max_retries": 0, "weight": 1, "cooldown_time": 300}, "model_info": {"allowed_fails": 0}}
POST /model/new     {"model_name": "<grp>", "litellm_params": {"model": "openai/gpt-5.5", "api_key": "os.environ/OPENAI_API_KEY", "weight": 0}}
POST /key/generate  {"models": ["<grp>"], "max_parallel_requests": 1, "max_budget": 0.05, "key_alias": "<label>"}

The hang-up in every case is a non-streaming ask for a long telegraph essay with max_tokens: 16384, abandoned 8s in, followed by six say hi <n> calls with the same key. Real Azure and OpenAI calls, real spend. Proxy log times are local (PDT), spend-row times are UTC, and the spend rows are read straight from LiteLLM_SpendLogs (status | model_id | spend | total_tokens | startTime | endTime | status code | error). The client legs run one script (python lit8248-legs-clients.py <port> <label> <httpx|aiohttp>) with httpx 0.28.1 (issue #35329's client: warm-up at timeout=60, essay at httpx.Timeout(8.0), catching httpx.ReadTimeout) or aiohttp 3.14.3 (issue #42222's client: ClientTimeout(total=60) and ClientTimeout(total=8.0), catching asyncio.TimeoutError), printing follow-up <n>: <status> <azure|backup|none> per call

Before (a83773c)

Live e2e cell

  1. From tests/e2e: LITELLM_PROXY_URL=http://127.0.0.1:43305 pytest router/test_reliability_cancel_on_disconnect_e2e.py -p no:cacheprovider -q --no-header -rA
  2. Output (the bench shows on the second follow-up because the first one landed on the worker that had not booked it):
>           assert model_id_of(resp) == azure, (
E           AssertionError: call 2 after the hang-up should have been served by the Azure deployment 4e2fc630-e040-4673-bbf5-96298b0e4b55, the proxy named 'ce2251b3-8d79-402c-aa06-759ea59adaa2': the cancelled call was booked as a failure and benched it
E           assert 'ce2251b3-8d7...-759ea59adaa2' == '4e2fc630-e04...-96298b0e4b55'

tests/e2e/router/test_reliability_cancel_on_disconnect_e2e.py:136: AssertionError
FAILED tests/e2e/router/test_reliability_cancel_on_disconnect_e2e.py::TestReliabilityCancelOnDisconnect::test_client_hanging_up_never_benches_the_deployment
1 failed in 61.38s (0:01:01)
  1. Proxy log for that hang-up:
14:44:39 - LiteLLM Router:DEBUG: cooldown_handlers.py:463 - Attempting to add 4e2fc630-e040-4673-bbf5-96298b0e4b55 to cooldown list
14:44:39 - LiteLLM Router:INFO: router.py:3699 - litellm.acompletion(model=azure/gpt-5.4-nano) Exception litellm.APIError: AzureException APIError - 
14:44:39 - LiteLLM Proxy:ERROR: common_request_processing.py:1545 - litellm.proxy.proxy_server._handle_llm_api_exception(): Exception occured - litellm_call_id=fceb0983-1a8f-485f-ac88-e598680621d0 - litellm.APIError: AzureException APIError - . Received Model Group=reliability-cooldown-disconnect-24b95e6dada0

curl hang-up, then six follow-ups

  1. Register group azure-legs-base2 and its key: azure=dbaa5ca1-1bf3-401c-a640-2531f24ed0a2 backup=8ab1acdc-5a67-46e7-a6ed-4aac3d027868
  2. Warm-up: curl -s -D - -o /dev/null -X POST "$BASE/chat/completions" -H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' -d '{"model": "azure-legs-base2", "messages": [{"role": "user", "content": "say hi"}]}' | grep -iE '^(HTTP|x-litellm-model-id|x-litellm-call-id)'
HTTP/1.1 200 OK
x-litellm-call-id: 0554e39e-79fa-4e1f-996d-d72655d011c7
x-litellm-model-id: dbaa5ca1-1bf3-401c-a640-2531f24ed0a2
  1. Hang up 8s into the essay: curl -s --max-time 8 -X POST "$BASE/chat/completions" -H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' -d '{"model": "azure-legs-base2", "messages": [{"role": "user", "content": "Write a 10000 word essay on the history of the telegraph, one section per decade."}], "max_tokens": 16384}'; echo "curl exit $?"
curl exit 28 (28 = hung up)
  1. Six say hi <n> calls with the same key. They landed on the worker that had not booked the bench, so they still name the Azure id, and the fourth one was refused by the key's budget check while the retried essay's reservation was in flight:
HTTP/1.1 200 OK    x-litellm-model-id: dbaa5ca1-1bf3-401c-a640-2531f24ed0a2
HTTP/1.1 200 OK    x-litellm-model-id: dbaa5ca1-1bf3-401c-a640-2531f24ed0a2
HTTP/1.1 200 OK    x-litellm-model-id: dbaa5ca1-1bf3-401c-a640-2531f24ed0a2
HTTP/1.1 422 Unprocessable Content
HTTP/1.1 200 OK    x-litellm-model-id: dbaa5ca1-1bf3-401c-a640-2531f24ed0a2
HTTP/1.1 200 OK    x-litellm-model-id: dbaa5ca1-1bf3-401c-a640-2531f24ed0a2
  1. Spend rows: the abandoned essay was retried on the backup and billed to the key even though nobody was waiting for it
azure-legs-base2 | success | dbaa5ca1-1bf3-401c-a640-2531f24ed0a2 | 1.91e-05  | 22   | 20:32:55 | 20:32:56
azure-legs-base2 | success | 8ab1acdc-5a67-46e7-a6ed-4aac3d027868 | 0.1819    | 6085 | 20:33:05 | 20:34:35
azure-legs-base2 | success | dbaa5ca1-1bf3-401c-a640-2531f24ed0a2 | 1.2e-05   | 18   | 20:33:08 | 20:33:09
azure-legs-base2 | success | dbaa5ca1-1bf3-401c-a640-2531f24ed0a2 | 1.825e-05 | 23   | 20:33:09 | 20:33:10
azure-legs-base2 | success | dbaa5ca1-1bf3-401c-a640-2531f24ed0a2 | 1.075e-05 | 17   | 20:33:10 | 20:33:11
azure-legs-base2 | success | dbaa5ca1-1bf3-401c-a640-2531f24ed0a2 | 1.45e-05  | 20   | 20:33:11 | 20:33:12
azure-legs-base2 | success | dbaa5ca1-1bf3-401c-a640-2531f24ed0a2 | 1.325e-05 | 19   | 20:33:13 | 20:33:14
  1. Proxy log for the hang-up:
13:33:05 - LiteLLM Router:DEBUG: cooldown_handlers.py:463 - Attempting to add dbaa5ca1-1bf3-401c-a640-2531f24ed0a2 to cooldown list
13:33:05 - LiteLLM Router:DEBUG: cooldown_handlers.py:514 - retrieve cooldown models: [('dbaa5ca1-1bf3-401c-a640-2531f24ed0a2', {'exception_received': 'litellm.APIError: AzureException APIError - ', 'status_code': '500', 'timestamp': 1790022785.376363, 'cooldown_time': 300})]

httpx client (issue #35329's client)

  1. python lit8248-legs-clients.py 43305 base httpx
  2. Output: every follow-up is refused with the key's budget_exceeded 422 ({"error":{"message":"Budget has been exceeded! Key=ecc0dd37... Current cost: 0.541575, Max budget: 0.05","type":"budget_exceeded",...}}) because the abandoned essay is being retried on the backup and its reservation alone exceeds the $0.05 budget
## azure-httpx-base: azure=5c915dfd-2dcf-440b-b9a5-2a5338bf820b backup=55d6bd17-566a-43f8-ae2f-5b6112baccc3
## httpx (issue #35329's client): timeout=8.0s on a long non-streaming answer
warm-up 200 azure
httpx.ReadTimeout after 8.0s (ReadTimeout): client hung up
follow-up 1: 422 (none)
follow-up 2: 422 (none)
follow-up 3: 422 (none)
follow-up 4: 422 (none)
follow-up 5: 422 (none)
follow-up 6: 422 (none)
  1. Proxy log: 14:40:46 - LiteLLM Router:DEBUG: cooldown_handlers.py:463 - Attempting to add 5c915dfd-2dcf-440b-b9a5-2a5338bf820b to cooldown list
  2. Spend rows: the warm-up, then the retried essay on the backup at $0.21, and no row at all for the six refused follow-ups
azure-httpx-base | success | 5c915dfd-2dcf-440b-b9a5-2a5338bf820b | 1.91e-05 | 22   | 21:40:36 | 21:40:38
azure-httpx-base | success | 55d6bd17-566a-43f8-ae2f-5b6112baccc3 | 0.21262  | 7109 | 21:40:47 | 21:42:28

aiohttp client (issue #42222's client)

  1. python lit8248-legs-clients.py 43305 base3 aiohttp
  2. Output: the follow-ups landed on the worker that had not booked the bench, so they still name the Azure id; the bench is in the log and the spend rows
## azure-aiohttp-base3: azure=0aa28d89-0d1f-400b-8557-ddd17b180eb6 backup=32ea426b-8db2-438c-b73e-eb29a9f07580
## aiohttp (issue #42222's client): ClientTimeout(total=8.0) on a long non-streaming answer
warm-up 200 azure
TimeoutError after 8.1s: client hung up
follow-up 1: 200 azure
follow-up 2: 200 azure
follow-up 3: 200 azure
follow-up 4: 200 azure
follow-up 5: 200 azure
follow-up 6: 200 azure
  1. Proxy log: 14:47:55 - LiteLLM Router:DEBUG: cooldown_handlers.py:463 - Attempting to add 0aa28d89-0d1f-400b-8557-ddd17b180eb6 to cooldown list
  2. Spend rows: the abandoned essay was retried to completion, all 10112 tokens of it, for a client that had already gone
azure-aiohttp-base3 | success | 0aa28d89-0d1f-400b-8557-ddd17b180eb6 | 1.91e-05  | 22    | 21:47:45 | 21:47:46
azure-aiohttp-base3 | success | 0aa28d89-0d1f-400b-8557-ddd17b180eb6 | 0.0125959 | 10112 | 21:47:56 | 21:49:48
azure-aiohttp-base3 | success | 0aa28d89-0d1f-400b-8557-ddd17b180eb6 | 1.2e-05   | 18    | 21:47:58 | 21:48:00
azure-aiohttp-base3 | success | 0aa28d89-0d1f-400b-8557-ddd17b180eb6 | 1.575e-05 | 21    | 21:48:00 | 21:48:01
azure-aiohttp-base3 | success | 0aa28d89-0d1f-400b-8557-ddd17b180eb6 | 1.325e-05 | 19    | 21:48:01 | 21:48:02
azure-aiohttp-base3 | success | 0aa28d89-0d1f-400b-8557-ddd17b180eb6 | 1.325e-05 | 19    | 21:48:02 | 21:48:04
azure-aiohttp-base3 | success | 0aa28d89-0d1f-400b-8557-ddd17b180eb6 | 1.825e-05 | 23    | 21:48:04 | 21:48:05
azure-aiohttp-base3 | success | 0aa28d89-0d1f-400b-8557-ddd17b180eb6 | 1.2e-05   | 18    | 21:48:05 | 21:48:06

Proxy log counts over every case above

  1. grep -c on the before proxy's log:
Attempting to add .* to cooldown list                   5
AzureException APIError                                33
client disconnected, upstream LLM request cancelled     5
Should Not Run Cooldown Logic                           1   (at 13:25:06, before any of the cases above ran)

Buildkite run of the cell

  1. litellm-e2e-pr build 571, the cell against a gateway built from a83773c: FAILED on the same assertion as the local run

After (70be37a; the litellm/ tree is byte-identical to the tip 073260c, since 6b8e988 and 073260c touch only tests/e2e)

Live e2e cell

  1. From tests/e2e: LITELLM_PROXY_URL=http://127.0.0.1:34701 pytest router/test_reliability_cancel_on_disconnect_e2e.py -p no:cacheprovider -q --no-header -rA
  2. Output:
PASSED tests/e2e/router/test_reliability_cancel_on_disconnect_e2e.py::TestReliabilityCancelOnDisconnect::test_client_hanging_up_never_benches_the_deployment
1 passed in 69.37s (0:01:09)
  1. Proxy log for that hang-up:
14:44:39 - LiteLLM Proxy:INFO: common_request_processing.py:1533 - litellm.proxy.proxy_server._handle_llm_api_exception(): client disconnected, upstream LLM request cancelled - litellm_call_id=2b645c34-71d3-48c7-afec-43339ae73e79
14:44:39 - LiteLLM Router:DEBUG: cooldown_handlers.py:320 - Should Not Run Cooldown Logic: _is_cooldown_required returned False

curl hang-up, then six follow-ups

  1. Register group azure-legs-head2 and its key: azure=1466fef8-42b3-4892-ab76-c348918ff661 backup=66a78404-4629-49df-8913-4e2bff446665
  2. Warm-up, same curl against http://127.0.0.1:34701:
HTTP/1.1 200 OK
x-litellm-call-id: 073a34b7-de35-47ae-9cef-8a7b0b812831
x-litellm-model-id: 1466fef8-42b3-4892-ab76-c348918ff661
  1. Hang up 8s into the essay, same curl:
curl exit 28 (28 = hung up)
  1. Six say hi <n> calls with the same key, all served by the Azure deployment the client hung up on:
HTTP/1.1 200 OK    x-litellm-model-id: 1466fef8-42b3-4892-ab76-c348918ff661
HTTP/1.1 200 OK    x-litellm-model-id: 1466fef8-42b3-4892-ab76-c348918ff661
HTTP/1.1 200 OK    x-litellm-model-id: 1466fef8-42b3-4892-ab76-c348918ff661
HTTP/1.1 200 OK    x-litellm-model-id: 1466fef8-42b3-4892-ab76-c348918ff661
HTTP/1.1 200 OK    x-litellm-model-id: 1466fef8-42b3-4892-ab76-c348918ff661
HTTP/1.1 200 OK    x-litellm-model-id: 1466fef8-42b3-4892-ab76-c348918ff661
  1. Spend rows: the hang-up is a spend-0 failure row booked as a 499, no retried essay, no backup row
azure-legs-head2 | success | 1466fef8-42b3-4892-ab76-c348918ff661 | 1.91e-05  | 22 | 20:32:54 | 20:32:55
azure-legs-head2 | failure | 1466fef8-42b3-4892-ab76-c348918ff661 | 0         | 0  | 20:32:57 | 20:33:04 | 499 | 499: Client disconnected the request
azure-legs-head2 | success | 1466fef8-42b3-4892-ab76-c348918ff661 | 1.075e-05 | 17 | 20:33:07 | 20:33:08
azure-legs-head2 | success | 1466fef8-42b3-4892-ab76-c348918ff661 | 1.325e-05 | 19 | 20:33:08 | 20:33:09
azure-legs-head2 | success | 1466fef8-42b3-4892-ab76-c348918ff661 | 1.075e-05 | 17 | 20:33:10 | 20:33:11
azure-legs-head2 | success | 1466fef8-42b3-4892-ab76-c348918ff661 | 1.325e-05 | 19 | 20:33:11 | 20:33:12
azure-legs-head2 | success | 1466fef8-42b3-4892-ab76-c348918ff661 | 1.825e-05 | 23 | 20:33:12 | 20:33:13
azure-legs-head2 | success | 1466fef8-42b3-4892-ab76-c348918ff661 | 1.325e-05 | 19 | 20:33:13 | 20:33:14
  1. Proxy log for the hang-up:
13:33:04 - LiteLLM Proxy:INFO: common_request_processing.py:1533 - litellm.proxy.proxy_server._handle_llm_api_exception(): client disconnected, upstream LLM request cancelled - litellm_call_id=19f04473-42d4-499c-978f-68bd44fd477f
13:33:04 - LiteLLM Router:DEBUG: cooldown_handlers.py:320 - Should Not Run Cooldown Logic: _is_cooldown_required returned False

httpx client (issue #35329's client)

  1. python lit8248-legs-clients.py 34701 head httpx
  2. Output:
## azure-httpx-head: azure=d3b161a5-3506-47a7-92f4-d7d53a271de1 backup=fbb90f27-9cc6-4026-b2fe-c10def8287b6
## httpx (issue #35329's client): timeout=8.0s on a long non-streaming answer
warm-up 200 azure
httpx.ReadTimeout after 8.1s (ReadTimeout): client hung up
follow-up 1: 200 azure
follow-up 2: 200 azure
follow-up 3: 200 azure
follow-up 4: 200 azure
follow-up 5: 200 azure
follow-up 6: 200 azure
  1. Proxy log:
14:40:50 - LiteLLM Proxy:INFO: common_request_processing.py:1533 - litellm.proxy.proxy_server._handle_llm_api_exception(): client disconnected, upstream LLM request cancelled - litellm_call_id=c0a31f9b-7aee-40ea-9ea7-a234469054ab
14:40:50 - LiteLLM Router:DEBUG: cooldown_handlers.py:320 - Should Not Run Cooldown Logic: _is_cooldown_required returned False
  1. Spend rows:
azure-httpx-head | success | d3b161a5-3506-47a7-92f4-d7d53a271de1 | 1.91e-05  | 22 | 21:40:40 | 21:40:42
azure-httpx-head | failure | d3b161a5-3506-47a7-92f4-d7d53a271de1 | 0         | 0  | 21:40:42 | 21:40:50 | 499 | 499: Client disconnected the request
azure-httpx-head | success | d3b161a5-3506-47a7-92f4-d7d53a271de1 | 1.2e-05   | 18 | 21:40:53 | 21:40:55
azure-httpx-head | success | d3b161a5-3506-47a7-92f4-d7d53a271de1 | 1.7e-05   | 22 | 21:40:55 | 21:40:56
azure-httpx-head | success | d3b161a5-3506-47a7-92f4-d7d53a271de1 | 1.2e-05   | 18 | 21:40:56 | 21:40:57
azure-httpx-head | success | d3b161a5-3506-47a7-92f4-d7d53a271de1 | 1.075e-05 | 17 | 21:40:57 | 21:40:58
azure-httpx-head | success | d3b161a5-3506-47a7-92f4-d7d53a271de1 | 1.325e-05 | 19 | 21:40:58 | 21:40:59
azure-httpx-head | success | d3b161a5-3506-47a7-92f4-d7d53a271de1 | 1.325e-05 | 19 | 21:40:59 | 21:41:01

aiohttp client (issue #42222's client)

  1. python lit8248-legs-clients.py 34701 head2 aiohttp
  2. Output:
## azure-aiohttp-head2: azure=c3822bcc-92cd-40dd-b785-4dca1537734e backup=dea1b636-c42c-4e4a-8d27-322ae966e854
## aiohttp (issue #42222's client): ClientTimeout(total=8.0) on a long non-streaming answer
warm-up 200 azure
TimeoutError after 9.0s: client hung up
follow-up 1: 200 azure
follow-up 2: 200 azure
follow-up 3: 200 azure
follow-up 4: 200 azure
follow-up 5: 200 azure
follow-up 6: 200 azure
  1. Proxy log:
14:44:09 - LiteLLM Proxy:INFO: common_request_processing.py:1533 - litellm.proxy.proxy_server._handle_llm_api_exception(): client disconnected, upstream LLM request cancelled - litellm_call_id=27fb9402-360b-4677-bc7f-b90e2bbd99e2
14:44:10 - LiteLLM Router:DEBUG: cooldown_handlers.py:320 - Should Not Run Cooldown Logic: _is_cooldown_required returned False
  1. Spend rows:
azure-aiohttp-head2 | success | c3822bcc-92cd-40dd-b785-4dca1537734e | 1.91e-05  | 22 | 21:43:59 | 21:44:00
azure-aiohttp-head2 | failure | c3822bcc-92cd-40dd-b785-4dca1537734e | 0         | 0  | 21:44:01 | 21:44:10 | 499 | 499: Client disconnected the request
azure-aiohttp-head2 | success | c3822bcc-92cd-40dd-b785-4dca1537734e | 1.075e-05 | 17 | 21:44:13 | 21:44:14
azure-aiohttp-head2 | success | c3822bcc-92cd-40dd-b785-4dca1537734e | 1.2e-05   | 18 | 21:44:14 | 21:44:15
azure-aiohttp-head2 | success | c3822bcc-92cd-40dd-b785-4dca1537734e | 1.2e-05   | 18 | 21:44:15 | 21:44:16
azure-aiohttp-head2 | success | c3822bcc-92cd-40dd-b785-4dca1537734e | 1.325e-05 | 19 | 21:44:16 | 21:44:17
azure-aiohttp-head2 | success | c3822bcc-92cd-40dd-b785-4dca1537734e | 1.075e-05 | 17 | 21:44:17 | 21:44:18
azure-aiohttp-head2 | success | c3822bcc-92cd-40dd-b785-4dca1537734e | 1.075e-05 | 17 | 21:44:18 | 21:44:19

Proxy log counts over every case above

  1. grep -c on the after proxy's log, six hang-ups in:
Attempting to add .* to cooldown list                   0
AzureException APIError                                 0
client disconnected, upstream LLM request cancelled     6
Should Not Run Cooldown Logic                           6

Buildkite run of the cell

  1. litellm-e2e-pr builds 569 and 570 (the cell at 70be37a) and 578 (this tip, 073260c, after BerriAI/project-releaser#255 landed): the cell PASSED in each ([gw2] [ 99%] PASSED router/test_reliability_cancel_on_disconnect_e2e.py::TestReliabilityCancelOnDisconnect::test_client_hanging_up_never_benches_the_deployment); build 578 finished 2 failed, 1271 passed, 57 skipped, 30 warnings, 1 rerun in 1996.50s, its two failures being the pre-existing Mistral OCR 429 and vertex cache-read flakes
  2. GitHub Actions test-e2e-changed run 35659271117 at 073260c: 3/3 jobs SUCCESS

Type

🐛 Bug Fix
✅ Test

Caveats (if any)

Low

  • buildkite/e2e-tests is red at the tip on two known flakes only, and it is not a required check on main
    • Mistral OCR 429 in test_ocr_rust_e2e.py::test_rust_ocr_response[mistral] and the vertex cache-read in test_cache_control.py
    • The new cell passed in that same build (578)
  • CircleCI at the tip carries fleet reds that main's own pipeline after the merge base shows identically
  • SDK callers cancelling an Azure acompletion task now see CancelledError, the asyncio contract every other provider follows, instead of APIError
  • Only the non-streaming Azure chat path caught CancelledError; streaming was never affected, and the other handlers under litellm/llms already re-raise
  • post_call logging still runs before the re-raise, unchanged from before
  • The mislabeled exception surfaced as litellm.APIError, which no per-type allowed-fails policy covers, so the test benches on generic allowed_fails: 0
  • A single-deployment group never cools down, so the test needs the zero-weight backup to observe the bench
  • On a Redis-less multi-worker proxy the cooldown lives only in the worker that booked it
    • Follow-ups on another worker still name the Azure id, which is why two Before cases above show the bench in the log and spend rows rather than in the follow-up headers
  • Before the fix the abandoned call is retried in the background and billed (up to $0.21 per hang-up above), and its reservation can 422 the key's next calls; the 499 ends the call, so the fix removes both

QA runbook

  • tests/e2e/router/test_reliability_cancel_on_disconnect_e2e.py::TestReliabilityCancelOnDisconnect::test_client_hanging_up_never_benches_the_deployment - a client that hangs up mid-answer on an Azure deployment does not get it benched; the next six calls all land on it
    • Start the proxy with general_settings.cancel_on_disconnect: true and store_model_in_db: true, with AZURE_API_KEY, AZURE_API_BASE (a resource serving gpt-5.4-nano) and OPENAI_API_KEY in the environment; GET /config/list with the master key must list cancel_on_disconnect as true, which is the test's precondition
    • POST /model/new with the master key: {"model_name": "grp", "litellm_params": {"model": "azure/gpt-5.4-nano", "api_key": "os.environ/AZURE_API_KEY", "api_base": "os.environ/AZURE_API_BASE", "api_version": "2024-10-21", "max_retries": 0, "weight": 1, "cooldown_time": 300}, "model_info": {"allowed_fails": 0}} and note the returned model_info.id
    • POST /model/new with the master key: {"model_name": "grp", "litellm_params": {"model": "openai/gpt-5.5", "api_key": "os.environ/OPENAI_API_KEY", "weight": 0}}
    • Warm up: curl -s -D - -o /dev/null -X POST http://localhost:4000/chat/completions -H "Authorization: Bearer sk-1234" -d '{"model": "grp", "messages": [{"role": "user", "content": "say hi"}]}' and expect 200 with x-litellm-model-id equal to the Azure id (a cold virtual key can spend a couple of seconds in auth, and a hang-up that lands before the provider call is in flight cancels nothing)
    • curl -m 5 -X POST http://localhost:4000/chat/completions -H "Authorization: Bearer sk-1234" -d '{"model": "grp", "messages": [{"role": "user", "content": "Write an essay on the history of the telegraph with one section per decade from the 1830s to the 2020s, each section at least 300 words."}], "max_tokens": 16384, "router_settings_override": {"num_retries": 0}}' and expect curl to exit 28 with no response; if the proxy answers 200 inside 5s, send the ask again (the test retries it up to three times and fails if all three come back early, since then no call was ever in flight to cancel)
    • Wait 15s, then run the warm-up curl six more times and expect 200 with x-litellm-model-id equal to the Azure id every time; before the fix the first follow-up to reach the worker that booked the bench names the backup id
    • Proxy log: expect client disconnected, upstream LLM request cancelled and Should Not Run Cooldown Logic for the hang-up and no Attempting to add <id> to cooldown list line
    • Sanity check: this test makes sense to add and is not hand-wavey (it asserts the exact deployment id on every follow-up, reads cancel_on_disconnect back from the proxy so it cannot pass vacuously, and fails loudly when the model answers all three asks before the hang-up) or potentially flaky

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Link to Devin session: https://app.devin.ai/sessions/c391a7a94c0148b8bfee7ea1bb629e52


Note

Medium Risk
Touches Azure async error handling and cancellation propagation on a path used with router cooldowns; behavior change for SDK callers (CancelledError vs APIError) is intentional but affects failure classification.

Overview
Azure async chat no longer turns client/task cancellation into a synthetic 500 AzureOpenAIError. On asyncio.CancelledError, the handler still logs via post_call but re-raises the cancellation so cancel_on_disconnect and router cooldown logic treat disconnects as client-driven, not deployment failures.

Tests and e2e harness: a unit test asserts litellm.acompletion against Azure surfaces CancelledError; a live reliability cell hangs up mid–long completion and checks the Azure deployment is not benched. Supporting changes add Transport.abandon / AbandonedRequest, general_setting_enabled via GET /config/list, an Azure “bench on first fail” deployment helper, cancel_on_disconnect: true in the stage mirror config, and a coverage-registry entry for reliability.cooldown.client_disconnect.stays_healthy.

Reviewed by Cursor Bugbot for commit 073260c. Bugbot is set up for automated code reviews on this repo. Configure here.

devin-ai-integration Bot and others added 7 commits July 31, 2026 06:03
…OpenAIError(500)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ancelled

Live proxy with cancel_on_disconnect, two-deployment group, generic allowed_fails=0, red at the pre-fix handler and green with the bare raise

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…hub.com/BerriAI/litellm into devin_ai_fix_azure_cancellederror_35329

# Conflicts:
#	tests/e2e/router/reliability_support.py
#	tests/e2e/router/test_reliability_cancel_on_disconnect_e2e.py
#	tests/test_litellm/llms/azure/test_azure.py
@devin-ai-integration

devin-ai-integration Bot commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you all sign our Contributor License Agreement before we can accept your contribution.
1 out of 2 committers have signed the CLA.

✅ mateo-berri
❌ devin-ai-integration[bot]
You have signed the CLA already but the status is still pending? Let us recheck it.

@codspeed

codspeed Bot commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_fix_azure_cancellederror_cooldown (073260c) with main (ebb4d23)

Open in CodSpeed

@greptile-apps

greptile-apps Bot commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge, with no outstanding correctness, security, or repository-rule violations.

Summary

This PR preserves asyncio cancellation semantics in Azure’s non-streaming async completion path so client disconnects are not converted into synthetic provider failures.

  • Re-raises asyncio.CancelledError after retaining post-call logging.
  • Adds unit coverage for cancellation propagation.
  • Adds an end-to-end disconnect test that verifies the Azure deployment remains healthy.
  • Extends the E2E transport and configuration helpers to support socket abandonment and prerequisite checks.

Reviews (4) · Last reviewed commit: "test(e2e): retry the hang-up when the mo..."

Comment thread tests/e2e/router/test_reliability_cancel_on_disconnect_e2e.py Outdated
Comment thread tests/e2e/router/test_reliability_cancel_on_disconnect_e2e.py
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@codecov

codecov Bot commented Sep 21, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@mateo-berri mateo-berri added run-ci and removed run-ci labels Sep 21, 2026

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 073260c. Configure here.

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@mateo-berri
mateo-berri merged commit 9fad216 into main Sep 22, 2026
140 of 149 checks passed
@mateo-berri
mateo-berri deleted the litellm_fix_azure_cancellederror_cooldown branch September 22, 2026 00:24

This branch was successfully deployed

1 active deployment
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

2 participants