Repository navigation
feat(router): reject with 429 when a deployment's max_parallel_requests slots are all in use - #41555
Conversation
…29 on overflow Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ded deployments Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…instead of subclassing it Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
|
🤖 Devin AI EngineerI'll be helping with this pull request! Here's what you should know: ✅ I will automatically:
Note: I can only respond to comments from users who have write access to this repository. ⚙️ Control Options:
|
|
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
…ive integer Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
bugbot run |
|
code-quality and documentation fail only because default_max_parallel_requests_queue_size is undocumented; BerriAI/litellm-docs#1513 adds the row, so they pass once it merges |
…ts slots are all in use Replace the per-deployment asyncio.Semaphore with MaxParallelRequestsLimit, which admits a call synchronously or raises the router's RateLimitError (429) right away. Nothing waits for a slot any more, so the max_parallel_requests_queue_size and default_max_parallel_requests_queue_size settings from the earlier commits are dropped along with their proxy validation, dashboard control and generated schema entries. The rpm/tpm derivation of the cap is unchanged. Every router endpoint family now enters the slot through one _deployment_slot context, and the provider coroutine is only created once the slot is held Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 8d972ee. Configure here.
|
@veria-ai please review 8d972ee: semaphore replaced by synchronous slot admission that raises 429 when full, slot release on every exit path including streaming, fallback and cooldown behavior, rpm/tpm derivation unchanged |
…requests_queue_size Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> # Conflicts: # tests/test_litellm/test_utils.py
|
@greptileai please re-review at 61ce1b4 (origin/main merged in, one test-file conflict resolved, no behavior change) |
|
bugbot run |
Bugbot is paused — on-demand spend limit reachedBugbot uses usage-based billing for this team and has hit its on-demand spend limit. A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue. |
|
@veria-ai please review 61ce1b4: origin/main merged in, same synchronous slot admission returning 429 with no wait queue, check release on exceptions and streams |
TLDR
Problem this solves:
rpm/tpmon a deployment silently become an in-process concurrency semaphoreHow it solves it:
max_parallel_requestsslots and nothing elserpm/tpmderivation of the cap is unchanged; only what happens over the cap changedUser Flow
Before: a developer whose deployment has
max_parallel_requests: 1(or anrpm/tpmthat derives it) sends a burst and every request eventually returns 200, silently serialized"model": "qa-gated"x-litellm-response-duration-msunder 2.2 s andx-litellm-attempted-retries: 0, so the 5 s the last request spent parked inside the proxy shows up nowhereAfter: the same burst gets one 200 and five immediate 429s that name the deployment and its cap
"model": "qa-gated"x-litellm-attempted-retries: 0"type":"throttling_error"and the messageDeployment has all max_parallel_requests slots in use. Deployment model_group=qa-gated, id=... already has max_parallel_requests=1 requests in flight. Raise max_parallel_requests (or the rpm/tpm it is derived from) for this deploymentx-litellm-attempted-retries: 1or2) and the rest surfaces as a 429 after about 5 s;num_retries: 0makes the caller see the 429 at onceRelevant issues
Affected release
Linear ticket
Resolves LIT-7675
Root cause
calculate_max_parallel_requests(litellm/utils.py) derives a per-deployment cap frommax_parallel_requests, elserpm, elseint(tpm / 1000 * 6)(minimum 1), elsedefault_max_parallel_requests, andInitalizeCachedClient.set_max_parallel_requests_clientwrapped it in a plainasyncio.Semaphore. Every router path awaited that semaphore, so a request beyond the cap parked inside the proxy until a slot freed up, with no bound, no error and no signal to the caller. The product decision on the ticket is to reject instead of queue, so the semaphore is replaced rather than bounded. The derivation formula, precedence and default cap are unchangedBehavior changes
This changes the default for every deployment that sets
max_parallel_requests,rpmortpm, and for every router withdefault_max_parallel_requests. Where a burst over the cap used to come back as slow 200s it now comes back as one 200 per slot and a 429 for the rest. The Caveats section calls this out as Severe because an operator who relied on the queueing has to raise the cap, add a fallback deployment or handle the 429 in the callerMaxParallelRequestsLimit(litellm/router_utils/client_initalization_utils.py) replaces the semaphore.acquireis synchronous: it checksin_flight >= max_parallel_requestsand increments in the same step with noawaitin between, so two callers on the same event loop cannot both pass a full check. A rejected caller never touches the provider. An admitted caller holds the slot through the provider call, through a streaming response until the stream is exhausted or closed, and releases it on exceptions too; the router's_deployment_slotasync context is the single place every endpoint family (chat, text completion, messages, embeddings, images, audio, files, batches, vector stores, the generic and passthrough helpers) enters and leaves the slot. The provider coroutine is created inside that context, so a rejection never leaves an unawaited coroutine behind (previously_acompletionbuiltlitellm.acompletion(**input_kwargs)before awaiting the semaphore)The rejection is the existing
litellm.RateLimitErrorwithRateLimitErrorCategory.LITELLM_RATE_LIMITandRateLimitType.CONCURRENT_REQUESTS, raised before any provider call. Cooldown accounting only runs on exceptions from the provider call, so a self-inflicted 429 does not put the deployment in cooldown;test_router_max_parallel_requests_overflow_is_429_without_cooldown_or_provider_callasserts_async_get_cooldown_deployments(...) == []afterwards. Retries and fallbacks treat it like any other 429: a model group with a second deployment fails over to it, and a single deployment is retriednum_retriestimes with backoff before the caller sees the error. The live run below shows both the default-retries shape andnum_retries: 0Slot state is per process, exactly like the semaphore it replaces. With
--num_workers 2a burst can land on two different worker-local limits and be admitted by both; a cross-worker gauge is LIT-7024 and out of scopeThe earlier revisions of this PR added
max_parallel_requests_queue_sizeanddefault_max_parallel_requests_queue_sizeto bound the queue instead. Since nothing queues any more those settings, their proxy validation, the Admin UI Router Settings control, the generatedschema.d.tsentries, the docs rows and their tests are all dropped again; the diff against the merge base is limited tolitellm/router.py,litellm/router_utils/client_initalization_utils.py,litellm/types/router.pyand the four test filesQueue wait used to be excluded from
x-litellm-response-duration-msandx-litellm-overhead-duration-ms; with no queue there is nothing left to excludeDependents of the changed surface: the cached object under
<model_id>_max_parallel_requests_clientis only read by the tworouter.pysites this PR rewrites (_acompletionand_deployment_slot) and bytests/local_testing/test_router_max_parallel_requests.py; nothing inenterprise/, the dashboard, the proxy client or the docs touches it.set_max_parallel_requests_clientkeeps its signature and cache key, andcalculate_max_parallel_requestsis untouched. The paths no unit test drives end to end (text completions, messages, responses, embeddings, streaming, an uncapped deployment, and the default retry policy) were A/B'd on the base and head proxies below against the real providerExploitability verdict
Ordinary bug work. Before this change the semaphore already enforced the concurrency cap, so nobody could do anything they were not entitled to; the defect is that overflow was invisible and unbounded. The change swaps waiting for a rejection and does not touch authorization, budgets, guardrails or tenant boundaries
Config
Nothing new to configure.
max_parallel_requestsper deployment,rpm/tpmper deployment anddefault_max_parallel_requestsunderrouter_settingsbehave as documented;num_retries: 0underrouter_settingsmakes the caller see the 429 without the router's retry backoffDocs: BerriAI/litellm-docs#1513 (routing page section on
max_parallel_requests)Tests
Backend (
LITELLM_LOCAL_MODEL_COST_MAP=True uv run --no-sync pytest):tests/test_litellm/router_utils/test_client_initalization_utils.py(429 with deployment and cap in the message when every slot is held, no waiting queue, a burst of exactly the cap is admitted and the rest rejected, release after a provider-like exception, router-derived cap precedence, no limit when nothing is set),tests/test_litellm/test_router.py(10 concurrentacompletioncalls on a cap of 2 give 2 provider calls, 2 peak in flight and 8 429s with nothing held afterwards, streaming and non-streaming; a streaming response holds the slot until exhausted; a rejected call never reaches the provider; no cooldown entry after the 429;aembeddingrejects the same way; an ordinary provider 429 still takes the fallback path; slot released after exit),tests/test_litellm/test_utils.py(derivation precedence and tpm minimum 1),tests/local_testing/test_router_max_parallel_requests.pyupdated from the semaphore internals to the new classMutation checks against the 34 mapped tests, each restored with
cpand verified withcmp -s:>=to>inacquire(13 fail), dropping the increment (14 fail), droppingrelease(5 fail), raisingTimeoutinstead ofRateLimitError(12 fail), removing the router's slot entry (7 fail), provider-originated instead ofLITELLM_RATE_LIMITcategory (12 fail). On the current tip (61ce1b4,origin/mainmerged in, the only conflict wastests/test_litellm/test_utils.pywhere main had dropped unrelated cost-map pinning tests next to the new precedence test): the mappedtests/test_litellmsuites give 1012 passed, 11 skipped, andmake lintexits 0 with every strict, type-discipline, test-quality, LIT and basedpyright budget gate passingPre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Setup shared by both sides: the same Postgres,
PYTHONPATH="$WT/enterprise:$WT"pointing at the tree under test,.venv/bin/python -m litellm.proxy.proxy_cli --config <config> --detailed_debug,OPENAI_API_KEYread from the QA vault at launch, real calls toopenai/gpt-5.4-nanoandopenai/text-embedding-3-small. Each launch prints the checked out sha,litellm.__file__andhas_fix(whetherMaxParallelRequestsLimitexists in the loadedclient_initalization_utils) so the tree serving each port is on record.fire.sh <port> <model> <n>sendsnsimultaneouscurl -sS -D - -w "HTTP=%{http_code} WALL=%{time_total}s"to/v1/chat/completionswith body{"model":"<model>","messages":[{"role":"user","content":"Reply with the single word ok"}],"max_tokens":16}and prints status, wall time,x-litellm-*headers and the start of the body;fire_ep.sh <port> <model> <n> <completions|messages|responses|embeddings|stream>does the same for the other endpoints. Output is in completion orderConfig (
qa-openhas no cap,qa-gatedis capped at 1 explicitly,qa-tpm-gatedderives a cap of 1 fromtpm: 150,qa-embed-gatedis an embedding model capped at 1):Cases 1 to 6 run with that config. Case 7 runs the same config without the
router_settingsblock, so the proxy's default retries apply. Before is the merge base (origin/mainat the time of the merge), After is the PR tipBefore (8fc9c46)
1. /v1/chat/completions, max_parallel_requests: 1
bash fire.sh 20676 qa-gated 6x-litellm-attempted-retries: 0and ax-litellm-response-duration-msfar below the wall time, so the wait shows up nowhere2. /v1/chat/completions, cap derived from tpm: 150
bash fire.sh 20676 qa-tpm-gated 6tpm-derived cap of 1 queues exactly like the explicit one3. /v1/completions, /v1/messages and /v1/responses on the capped deployment
bash fire_ep.sh 20676 qa-gated 3 completions,bash fire_ep.sh 20676 qa-gated 3 messages,bash fire_ep.sh 20676 qa-gated 3 responses4. /v1/embeddings on a capped embedding deployment
bash fire_ep.sh 20676 qa-embed-gated 3 embeddings5. Streaming /v1/chat/completions on the capped deployment
bash fire_ep.sh 20676 qa-gated 3 stream6. Control: deployment with no cap
bash fire.sh 20676 qa-open 67. /v1/chat/completions, max_parallel_requests: 1, proxy default retries
router_settingsblock on port 20677,bash fire.sh 20677 qa-gated 48. Admin UI on the Before proxy
adminwith the master key, open Logs, Filters, Error Code429 - Rate Limited, ApplyAfter (61ce1b4)
1. /v1/chat/completions, max_parallel_requests: 1
bash fire.sh 20686 qa-gated 6"type":"throttling_error"and a message naming the deployment and its cap; the 429s never reach the provider2. /v1/chat/completions, cap derived from tpm: 150
bash fire.sh 20686 qa-tpm-gated 63. /v1/completions, /v1/messages and /v1/responses on the capped deployment
bash fire_ep.sh 20686 qa-gated 3 completions,bash fire_ep.sh 20686 qa-gated 3 messages,bash fire_ep.sh 20686 qa-gated 3 responses/v1/messagesreturns the 429 in Anthropic error shape4. /v1/embeddings on a capped embedding deployment
bash fire_ep.sh 20686 qa-embed-gated 3 embeddings5. Streaming /v1/chat/completions on the capped deployment
bash fire_ep.sh 20686 qa-gated 3 stream6. Control: deployment with no cap
bash fire.sh 20686 qa-open 67. /v1/chat/completions, max_parallel_requests: 1, proxy default retries
router_settingsblock on port 20687,bash fire.sh 20687 qa-gated 4x-litellm-attempted-retries), the fourth surfaces the 429 after about 5 s once retries are exhausted8. Admin UI on the After proxy
adminwith the master key, open Models + Endpoints, click theqa-gatedrow. The LiteLLM Params block shows the cap the router enforces429 - Rate Limited, Apply. Every overflow request from cases 1 to 7 is a Failure row with a duration under 0.1 s (the 4.84 s row is the retry-exhausted request from case 7); before this change these rows did not exist because the requests waited and then succeededError Code: 429and the message naming the deployment and itsmax_parallel_requests=1Type
🐛 Bug Fix
Caveats (if any)
Severe
rpm/tpm/max_parallel_requestsnow 429 instead of queueingMedium
num_retries: 0Low
MaxParallelRequestsLimit.in_flightis a mutable counter on purpose: check and increment must happen in one synchronous step for the same-loop race guaranteeReview gate status at 61ce1b4
CI: 89 check runs passed and one skipped (the stage-mirror e2e job is conditional), none failed or pending. Greptile: 5/5 with no outstanding findings, last reviewed commit 61ce1b4. CodeQL: passed with zero open review threads. Bugbot: a single
bugbot runwas posted by mateo-berri at 2026-09-18 09:14 UTC for this head and no Cursor check run or review started in the following 86 minutes; sibling PRs triggered in the same window are in the same state and the team's Cursor on-demand spend limit is reported as exhausted, so Bugbot is recorded as unavailable for this head rather than passed. The last Bugbot verdict on this branch is "found no new issues" on ancestor 8d972ee, and the only change since is the merge of origin/main with one test-file conflict resolution. Veria has not posted on this PR at any SHAFinal Attestation
ran /live-pr-risk and found no regressions/backward incompatible risks
Link to Devin session: https://app.devin.ai/sessions/7b097835d9504224a951e5eda64247b7
Open in Devin Desktop: https://app.devin.ai/desktop/session/7b097835d9504224a951e5eda64247b7?variant=devin
Requested by: @yassin-berriai