Repository navigation
fix(azure): propagate asyncio.CancelledError instead of raising a 500 - #42295
Conversation
…OpenAIError(500) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ancelled Live proxy with cancel_on_disconnect, two-deployment group, generic allowed_fails=0, red at the pre-fix handler and green with the bare raise Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…e Azure deployment
…ncellederror_35329
…hub.com/BerriAI/litellm into devin_ai_fix_azure_cancellederror_35329 # Conflicts: # tests/e2e/router/reliability_support.py # tests/e2e/router/test_reliability_cancel_on_disconnect_e2e.py # tests/test_litellm/llms/azure/test_azure.py
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
|
|
|
…connect cell's prose
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
|
bugbot run |
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 073260c. Configure here.
Supersedes #35330 at its head 70be37a: the same commits and authors, moved to a
litellm_branch so CircleCI runs the legacy provider suites on itTLDR
Problem this solves:
cancel_on_disconnect, one client hang-up became a deployment failureAzureException APIError500s Azure never returnedHow it solves it:
asyncio.CancelledErrorinstead ofAzureOpenAIError(500)Transport.abandon, which closes the socket mid-answerUser Flow
Before: a platform team running
cancel_on_disconnect: truein front of an Azure model group watches one impatient client take a healthy Azure deployment out of rotation"model": "azure-group"and a prompt whose answer takes about a minuteAzureException APIErrorwith an empty error body, booked against the Azure deployment that was serving it"model": "azure-group", comes back 200 with anx-litellm-model-idheader naming a different deployment: the one the client hung up on is in cooldownAfter: the same hang-up is recorded as the client's doing and the Azure deployment keeps serving
"model": "azure-group"and a prompt whose answer takes about a minuteClient disconnected the requestand is not counted against any deployment"model": "azure-group", comes back 200 withx-litellm-model-idnaming the same Azure deployment the client hung up onRelevant issues
Fixes #35329
Fixes #42222
Affected release
Linear ticket
Resolves LIT-8248
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/llms/azure/test_azure.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Two local proxies, two uvicorn workers each with no Redis, on random free ports: http://127.0.0.1:43305 booted from the merge base tree (a83773c) and http://127.0.0.1:34701 from this branch (70be37a). Both run this config, each against its own Postgres:
Every case registers its own group through the API: a real Azure
gpt-5.4-nanodeployment holding all the weight and benched on its first failure, plus a zero-weight OpenAI backup, which is needed because a single-deployment group never cools down. The key is limited to one parallel request and a $0.05 budget:The hang-up in every case is a non-streaming ask for a long telegraph essay with
max_tokens: 16384, abandoned 8s in, followed by sixsay hi <n>calls with the same key. Real Azure and OpenAI calls, real spend. Proxy log times are local (PDT), spend-row times are UTC, and the spend rows are read straight fromLiteLLM_SpendLogs(status | model_id | spend | total_tokens | startTime | endTime | status code | error). The client legs run one script (python lit8248-legs-clients.py <port> <label> <httpx|aiohttp>) with httpx 0.28.1 (issue #35329's client: warm-up attimeout=60, essay athttpx.Timeout(8.0), catchinghttpx.ReadTimeout) or aiohttp 3.14.3 (issue #42222's client:ClientTimeout(total=60)andClientTimeout(total=8.0), catchingasyncio.TimeoutError), printingfollow-up <n>: <status> <azure|backup|none>per callBefore (a83773c)
Live e2e cell
tests/e2e:LITELLM_PROXY_URL=http://127.0.0.1:43305 pytest router/test_reliability_cancel_on_disconnect_e2e.py -p no:cacheprovider -q --no-header -rAcurl hang-up, then six follow-ups
azure-legs-base2and its key:azure=dbaa5ca1-1bf3-401c-a640-2531f24ed0a2 backup=8ab1acdc-5a67-46e7-a6ed-4aac3d027868curl -s -D - -o /dev/null -X POST "$BASE/chat/completions" -H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' -d '{"model": "azure-legs-base2", "messages": [{"role": "user", "content": "say hi"}]}' | grep -iE '^(HTTP|x-litellm-model-id|x-litellm-call-id)'curl -s --max-time 8 -X POST "$BASE/chat/completions" -H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' -d '{"model": "azure-legs-base2", "messages": [{"role": "user", "content": "Write a 10000 word essay on the history of the telegraph, one section per decade."}], "max_tokens": 16384}'; echo "curl exit $?"say hi <n>calls with the same key. They landed on the worker that had not booked the bench, so they still name the Azure id, and the fourth one was refused by the key's budget check while the retried essay's reservation was in flight:httpx client (issue #35329's client)
python lit8248-legs-clients.py 43305 base httpxbudget_exceeded422 ({"error":{"message":"Budget has been exceeded! Key=ecc0dd37... Current cost: 0.541575, Max budget: 0.05","type":"budget_exceeded",...}}) because the abandoned essay is being retried on the backup and its reservation alone exceeds the $0.05 budget14:40:46 - LiteLLM Router:DEBUG: cooldown_handlers.py:463 - Attempting to add 5c915dfd-2dcf-440b-b9a5-2a5338bf820b to cooldown listaiohttp client (issue #42222's client)
python lit8248-legs-clients.py 43305 base3 aiohttp14:47:55 - LiteLLM Router:DEBUG: cooldown_handlers.py:463 - Attempting to add 0aa28d89-0d1f-400b-8557-ddd17b180eb6 to cooldown listProxy log counts over every case above
grep -con the before proxy's log:Buildkite run of the cell
After (70be37a; the
litellm/tree is byte-identical to the tip 073260c, since 6b8e988 and 073260c touch onlytests/e2e)Live e2e cell
tests/e2e:LITELLM_PROXY_URL=http://127.0.0.1:34701 pytest router/test_reliability_cancel_on_disconnect_e2e.py -p no:cacheprovider -q --no-header -rAcurl hang-up, then six follow-ups
azure-legs-head2and its key:azure=1466fef8-42b3-4892-ab76-c348918ff661 backup=66a78404-4629-49df-8913-4e2bff446665say hi <n>calls with the same key, all served by the Azure deployment the client hung up on:httpx client (issue #35329's client)
python lit8248-legs-clients.py 34701 head httpxaiohttp client (issue #42222's client)
python lit8248-legs-clients.py 34701 head2 aiohttpProxy log counts over every case above
grep -con the after proxy's log, six hang-ups in:Buildkite run of the cell
[gw2] [ 99%] PASSED router/test_reliability_cancel_on_disconnect_e2e.py::TestReliabilityCancelOnDisconnect::test_client_hanging_up_never_benches_the_deployment); build 578 finished2 failed, 1271 passed, 57 skipped, 30 warnings, 1 rerun in 1996.50s, its two failures being the pre-existing Mistral OCR 429 and vertex cache-read flakestest-e2e-changedrun 35659271117 at 073260c: 3/3 jobs SUCCESSType
🐛 Bug Fix
✅ Test
Caveats (if any)
Low
buildkite/e2e-testsis red at the tip on two known flakes only, and it is not a required check onmaintest_ocr_rust_e2e.py::test_rust_ocr_response[mistral]and the vertex cache-read intest_cache_control.pymain's own pipeline after the merge base shows identicallytest_jev_classifiertimeouts (feat(auto-router): add JEV classifier alongside LLM classifier #41886), two bedrocktest_get_model_info, ten fireworks translation tests,test_bad_database_url, seventest_e2e_budgetingacompletiontask now seeCancelledError, the asyncio contract every other provider follows, instead ofAPIErrorCancelledError; streaming was never affected, and the other handlers underlitellm/llmsalready re-raisepost_calllogging still runs before the re-raise, unchanged from beforelitellm.APIError, which no per-type allowed-fails policy covers, so the test benches on genericallowed_fails: 0QA runbook
general_settings.cancel_on_disconnect: trueandstore_model_in_db: true, withAZURE_API_KEY,AZURE_API_BASE(a resource servinggpt-5.4-nano) andOPENAI_API_KEYin the environment;GET /config/listwith the master key must listcancel_on_disconnectastrue, which is the test's precondition{"model_name": "grp", "litellm_params": {"model": "azure/gpt-5.4-nano", "api_key": "os.environ/AZURE_API_KEY", "api_base": "os.environ/AZURE_API_BASE", "api_version": "2024-10-21", "max_retries": 0, "weight": 1, "cooldown_time": 300}, "model_info": {"allowed_fails": 0}}and note the returnedmodel_info.id{"model_name": "grp", "litellm_params": {"model": "openai/gpt-5.5", "api_key": "os.environ/OPENAI_API_KEY", "weight": 0}}curl -s -D - -o /dev/null -X POST http://localhost:4000/chat/completions -H "Authorization: Bearer sk-1234" -d '{"model": "grp", "messages": [{"role": "user", "content": "say hi"}]}'and expect 200 withx-litellm-model-idequal to the Azure id (a cold virtual key can spend a couple of seconds in auth, and a hang-up that lands before the provider call is in flight cancels nothing)curl -m 5 -X POST http://localhost:4000/chat/completions -H "Authorization: Bearer sk-1234" -d '{"model": "grp", "messages": [{"role": "user", "content": "Write an essay on the history of the telegraph with one section per decade from the 1830s to the 2020s, each section at least 300 words."}], "max_tokens": 16384, "router_settings_override": {"num_retries": 0}}'and expect curl to exit 28 with no response; if the proxy answers 200 inside 5s, send the ask again (the test retries it up to three times and fails if all three come back early, since then no call was ever in flight to cancel)x-litellm-model-idequal to the Azure id every time; before the fix the first follow-up to reach the worker that booked the bench names the backup idclient disconnected, upstream LLM request cancelledandShould Not Run Cooldown Logicfor the hang-up and noAttempting to add <id> to cooldown listlinecancel_on_disconnectback from the proxy so it cannot pass vacuously, and fails loudly when the model answers all three asks before the hang-up) or potentially flakyFinal Attestation
Link to Devin session: https://app.devin.ai/sessions/c391a7a94c0148b8bfee7ea1bb629e52
Note
Medium Risk
Touches Azure async error handling and cancellation propagation on a path used with router cooldowns; behavior change for SDK callers (CancelledError vs APIError) is intentional but affects failure classification.
Overview
Azure async chat no longer turns client/task cancellation into a synthetic 500
AzureOpenAIError. Onasyncio.CancelledError, the handler still logs viapost_callbut re-raises the cancellation socancel_on_disconnectand router cooldown logic treat disconnects as client-driven, not deployment failures.Tests and e2e harness: a unit test asserts
litellm.acompletionagainst Azure surfacesCancelledError; a live reliability cell hangs up mid–long completion and checks the Azure deployment is not benched. Supporting changes addTransport.abandon/AbandonedRequest,general_setting_enabledviaGET /config/list, an Azure “bench on first fail” deployment helper,cancel_on_disconnect: truein the stage mirror config, and a coverage-registry entry forreliability.cooldown.client_disconnect.stays_healthy.Reviewed by Cursor Bugbot for commit 073260c. Bugbot is set up for automated code reviews on this repo. Configure here.