Skip to content

feat(router): reject with 429 when a deployment's max_parallel_requests slots are all in use - #41555

Merged
yassin-berriai merged 6 commits into
mainfrom
litellm_max_parallel_requests_queue_size
Sep 18, 2026
Merged

yassin-berriai merged 6 commits into
mainfrom
litellm_max_parallel_requests_queue_size

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • rpm / tpm on a deployment silently become an in-process concurrency semaphore
  • requests over the cap wait in the proxy for as long as it takes, then return 200
  • callers see a latency staircase and no 429, so nothing tells them the deployment is saturated

How it solves it:

  • the semaphore is gone; a deployment has max_parallel_requests slots and nothing else
  • a request arriving with every slot in use gets a 429 right away, before any provider call
  • the rpm / tpm derivation of the cap is unchanged; only what happens over the cap changed
  • one shared slot context for every router endpoint family; provider coroutine only created once the slot is held

User Flow

Before: a developer whose deployment has max_parallel_requests: 1 (or an rpm / tpm that derives it) sends a burst and every request eventually returns 200, silently serialized

  1. They send 6 parallel POST http://localhost:4000/v1/chat/completions with "model": "qa-gated"
  2. All 6 come back 200, one after another: wall times 2.3 s, 3.1 s, 3.9 s, 4.6 s, 5.2 s, 5.9 s
  3. Every response carries x-litellm-response-duration-ms under 2.2 s and x-litellm-attempted-retries: 0, so the 5 s the last request spent parked inside the proxy shows up nowhere
  4. The same happens for POST /v1/completions, POST /v1/messages, POST /v1/embeddings and streaming chat completions on that deployment

After: the same burst gets one 200 and five immediate 429s that name the deployment and its cap

  1. They send 6 parallel POST http://localhost:4000/v1/chat/completions with "model": "qa-gated"
  2. One comes back 200 in 1.0 s with x-litellm-attempted-retries: 0
  3. The other 5 come back 429 within 0.4 s with "type":"throttling_error" and the message Deployment has all max_parallel_requests slots in use. Deployment model_group=qa-gated, id=... already has max_parallel_requests=1 requests in flight. Raise max_parallel_requests (or the rpm/tpm it is derived from) for this deployment
  4. The same holds for POST /v1/completions, POST /v1/messages, POST /v1/embeddings and streaming chat completions on that deployment, while a deployment with no cap still serves the whole burst in parallel
  5. With the proxy's default retries left on, the router retries the 429 against the same deployment with backoff, so some of the overflow gets through on a later attempt (x-litellm-attempted-retries: 1 or 2) and the rest surfaces as a 429 after about 5 s; num_retries: 0 makes the caller see the 429 at once

Relevant issues

Affected release

Linear ticket

Resolves LIT-7675

Root cause

calculate_max_parallel_requests (litellm/utils.py) derives a per-deployment cap from max_parallel_requests, else rpm, else int(tpm / 1000 * 6) (minimum 1), else default_max_parallel_requests, and InitalizeCachedClient.set_max_parallel_requests_client wrapped it in a plain asyncio.Semaphore. Every router path awaited that semaphore, so a request beyond the cap parked inside the proxy until a slot freed up, with no bound, no error and no signal to the caller. The product decision on the ticket is to reject instead of queue, so the semaphore is replaced rather than bounded. The derivation formula, precedence and default cap are unchanged

Behavior changes

This changes the default for every deployment that sets max_parallel_requests, rpm or tpm, and for every router with default_max_parallel_requests. Where a burst over the cap used to come back as slow 200s it now comes back as one 200 per slot and a 429 for the rest. The Caveats section calls this out as Severe because an operator who relied on the queueing has to raise the cap, add a fallback deployment or handle the 429 in the caller

MaxParallelRequestsLimit (litellm/router_utils/client_initalization_utils.py) replaces the semaphore. acquire is synchronous: it checks in_flight >= max_parallel_requests and increments in the same step with no await in between, so two callers on the same event loop cannot both pass a full check. A rejected caller never touches the provider. An admitted caller holds the slot through the provider call, through a streaming response until the stream is exhausted or closed, and releases it on exceptions too; the router's _deployment_slot async context is the single place every endpoint family (chat, text completion, messages, embeddings, images, audio, files, batches, vector stores, the generic and passthrough helpers) enters and leaves the slot. The provider coroutine is created inside that context, so a rejection never leaves an unawaited coroutine behind (previously _acompletion built litellm.acompletion(**input_kwargs) before awaiting the semaphore)

The rejection is the existing litellm.RateLimitError with RateLimitErrorCategory.LITELLM_RATE_LIMIT and RateLimitType.CONCURRENT_REQUESTS, raised before any provider call. Cooldown accounting only runs on exceptions from the provider call, so a self-inflicted 429 does not put the deployment in cooldown; test_router_max_parallel_requests_overflow_is_429_without_cooldown_or_provider_call asserts _async_get_cooldown_deployments(...) == [] afterwards. Retries and fallbacks treat it like any other 429: a model group with a second deployment fails over to it, and a single deployment is retried num_retries times with backoff before the caller sees the error. The live run below shows both the default-retries shape and num_retries: 0

Slot state is per process, exactly like the semaphore it replaces. With --num_workers 2 a burst can land on two different worker-local limits and be admitted by both; a cross-worker gauge is LIT-7024 and out of scope

The earlier revisions of this PR added max_parallel_requests_queue_size and default_max_parallel_requests_queue_size to bound the queue instead. Since nothing queues any more those settings, their proxy validation, the Admin UI Router Settings control, the generated schema.d.ts entries, the docs rows and their tests are all dropped again; the diff against the merge base is limited to litellm/router.py, litellm/router_utils/client_initalization_utils.py, litellm/types/router.py and the four test files

Queue wait used to be excluded from x-litellm-response-duration-ms and x-litellm-overhead-duration-ms; with no queue there is nothing left to exclude

Dependents of the changed surface: the cached object under <model_id>_max_parallel_requests_client is only read by the two router.py sites this PR rewrites (_acompletion and _deployment_slot) and by tests/local_testing/test_router_max_parallel_requests.py; nothing in enterprise/, the dashboard, the proxy client or the docs touches it. set_max_parallel_requests_client keeps its signature and cache key, and calculate_max_parallel_requests is untouched. The paths no unit test drives end to end (text completions, messages, responses, embeddings, streaming, an uncapped deployment, and the default retry policy) were A/B'd on the base and head proxies below against the real provider

Exploitability verdict

Ordinary bug work. Before this change the semaphore already enforced the concurrency cap, so nobody could do anything they were not entitled to; the defect is that overflow was invisible and unbounded. The change swaps waiting for a rejection and does not touch authorization, budgets, guardrails or tenant boundaries

Config

Nothing new to configure. max_parallel_requests per deployment, rpm / tpm per deployment and default_max_parallel_requests under router_settings behave as documented; num_retries: 0 under router_settings makes the caller see the 429 without the router's retry backoff

Docs: BerriAI/litellm-docs#1513 (routing page section on max_parallel_requests)

Tests

Backend (LITELLM_LOCAL_MODEL_COST_MAP=True uv run --no-sync pytest): tests/test_litellm/router_utils/test_client_initalization_utils.py (429 with deployment and cap in the message when every slot is held, no waiting queue, a burst of exactly the cap is admitted and the rest rejected, release after a provider-like exception, router-derived cap precedence, no limit when nothing is set), tests/test_litellm/test_router.py (10 concurrent acompletion calls on a cap of 2 give 2 provider calls, 2 peak in flight and 8 429s with nothing held afterwards, streaming and non-streaming; a streaming response holds the slot until exhausted; a rejected call never reaches the provider; no cooldown entry after the 429; aembedding rejects the same way; an ordinary provider 429 still takes the fallback path; slot released after exit), tests/test_litellm/test_utils.py (derivation precedence and tpm minimum 1), tests/local_testing/test_router_max_parallel_requests.py updated from the semaphore internals to the new class

Mutation checks against the 34 mapped tests, each restored with cp and verified with cmp -s: >= to > in acquire (13 fail), dropping the increment (14 fail), dropping release (5 fail), raising Timeout instead of RateLimitError (12 fail), removing the router's slot entry (7 fail), provider-originated instead of LITELLM_RATE_LIMIT category (12 fail). On the current tip (61ce1b4, origin/main merged in, the only conflict was tests/test_litellm/test_utils.py where main had dropped unrelated cost-map pinning tests next to the new precedence test): the mapped tests/test_litellm suites give 1012 passed, 11 skipped, and make lint exits 0 with every strict, type-discipline, test-quality, LIT and basedpyright budget gate passing

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Setup shared by both sides: the same Postgres, PYTHONPATH="$WT/enterprise:$WT" pointing at the tree under test, .venv/bin/python -m litellm.proxy.proxy_cli --config <config> --detailed_debug, OPENAI_API_KEY read from the QA vault at launch, real calls to openai/gpt-5.4-nano and openai/text-embedding-3-small. Each launch prints the checked out sha, litellm.__file__ and has_fix (whether MaxParallelRequestsLimit exists in the loaded client_initalization_utils) so the tree serving each port is on record. fire.sh <port> <model> <n> sends n simultaneous curl -sS -D - -w "HTTP=%{http_code} WALL=%{time_total}s" to /v1/chat/completions with body {"model":"<model>","messages":[{"role":"user","content":"Reply with the single word ok"}],"max_tokens":16} and prints status, wall time, x-litellm-* headers and the start of the body; fire_ep.sh <port> <model> <n> <completions|messages|responses|embeddings|stream> does the same for the other endpoints. Output is in completion order

Config (qa-open has no cap, qa-gated is capped at 1 explicitly, qa-tpm-gated derives a cap of 1 from tpm: 150, qa-embed-gated is an embedding model capped at 1):

model_list:
  - model_name: qa-open
    litellm_params: {model: openai/gpt-5.4-nano, api_key: os.environ/OPENAI_API_KEY}
  - model_name: qa-gated
    litellm_params: {model: openai/gpt-5.4-nano, api_key: os.environ/OPENAI_API_KEY, max_parallel_requests: 1}
  - model_name: qa-tpm-gated
    litellm_params: {model: openai/gpt-5.4-nano, api_key: os.environ/OPENAI_API_KEY, tpm: 150}
  - model_name: qa-embed-gated
    litellm_params: {model: openai/text-embedding-3-small, api_key: os.environ/OPENAI_API_KEY, max_parallel_requests: 1}
router_settings:
  num_retries: 0
general_settings:
  master_key: sk-1234

Cases 1 to 6 run with that config. Case 7 runs the same config without the router_settings block, so the proxy's default retries apply. Before is the merge base (origin/main at the time of the merge), After is the PR tip

Before (8fc9c46)

sha=8fc9c46d1aeda6d009718cd22684a526fa4647fd
litellm.__file__ /home/ubuntu/litfix-7675/base-wt/litellm/__init__.py
has_fix False

1. /v1/chat/completions, max_parallel_requests: 1

  1. bash fire.sh 20676 qa-gated 6
  2. Observed: all 6 return 200 in a staircase, roughly one provider call apart, with x-litellm-attempted-retries: 0 and a x-litellm-response-duration-ms far below the wall time, so the wait shows up nowhere
req1 HTTP=200 WALL=2.306823s x-litellm-response-duration-ms: 2108.44 x-litellm-overhead-duration-ms: 73.224 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOk6fddQ5aGOxprdgHLxHiFJzvFt","created":1789722018,"model":"qa-gated","object":"chat.completi ...
req5 HTTP=200 WALL=3.093339s x-litellm-response-duration-ms: 781.701 x-litellm-overhead-duration-ms: 18.764 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOk7gODGyWlBNXOk9gaXwAyiHsg5","created":1789722019,"model":"qa-gated","object":"chat.completi ...
req3 HTTP=200 WALL=3.883346s x-litellm-response-duration-ms: 784.53 x-litellm-overhead-duration-ms: 14.918 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOk8X55zuX6melckvM8xjD5NKz9z","created":1789722020,"model":"qa-gated","object":"chat.completi ...
req4 HTTP=200 WALL=4.639294s x-litellm-response-duration-ms: 751.004 x-litellm-overhead-duration-ms: 17.17 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOk9s8U6toCHv9Eh1Ma87Hqiqeo5","created":1789722021,"model":"qa-gated","object":"chat.completi ...
req2 HTTP=200 WALL=5.162785s x-litellm-response-duration-ms: 518.24 x-litellm-overhead-duration-ms: 13.256 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOk9UfgUhgtR67QRPMLmv3OuCxhN","created":1789722021,"model":"qa-gated","object":"chat.completi ...
req6 HTTP=200 WALL=5.883939s x-litellm-response-duration-ms: 719.955 x-litellm-overhead-duration-ms: 14.366 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOkAaHOL7PCvV7GY4c00G8W2buGS","created":1789722022,"model":"qa-gated","object":"chat.completi ...

2. /v1/chat/completions, cap derived from tpm: 150

  1. bash fire.sh 20676 qa-tpm-gated 6
  2. Observed: the same staircase, so the tpm-derived cap of 1 queues exactly like the explicit one
req3 HTTP=200 WALL=0.827003s x-litellm-response-duration-ms: 788.213 x-litellm-overhead-duration-ms: 5.823 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOkGMczsQ4WCOll0Xrx1FVIi2Vj6","created":1789722028,"model":"qa-tpm-gated","object":"chat.comp ...
req2 HTTP=200 WALL=1.547626s x-litellm-response-duration-ms: 715.745 x-litellm-overhead-duration-ms: 13.472 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOkG86AxTYL10cZmFOTlBCy1DLbr","created":1789722028,"model":"qa-tpm-gated","object":"chat.comp ...
req5 HTTP=200 WALL=2.328332s x-litellm-response-duration-ms: 776.178 x-litellm-overhead-duration-ms: 14.658 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOkHpsGq6w7IMgu6Ub3QFpN9lsOy","created":1789722029,"model":"qa-tpm-gated","object":"chat.comp ...
req4 HTTP=200 WALL=3.115168s x-litellm-response-duration-ms: 780.973 x-litellm-overhead-duration-ms: 12.885 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOkIZkoxw7YEVq920GnQBObm25Pj","created":1789722030,"model":"qa-tpm-gated","object":"chat.comp ...
req6 HTTP=200 WALL=3.562585s x-litellm-response-duration-ms: 442.539 x-litellm-overhead-duration-ms: 13.087 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOkJSOcTy4QJ0DMuOzUTG0gQOmXy","created":1789722031,"model":"qa-tpm-gated","object":"chat.comp ...
req1 HTTP=200 WALL=4.433212s x-litellm-response-duration-ms: 868.827 x-litellm-overhead-duration-ms: 16.155 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOkJ31oFZjj9YpZvohP2par6fVkb","created":1789722031,"model":"qa-tpm-gated","object":"chat.comp ...

3. /v1/completions, /v1/messages and /v1/responses on the capped deployment

  1. bash fire_ep.sh 20676 qa-gated 3 completions, bash fire_ep.sh 20676 qa-gated 3 messages, bash fire_ep.sh 20676 qa-gated 3 responses
  2. Observed: every request reaches the provider one at a time and returns 200
req1 HTTP=200 WALL=0.483036s x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOkKn6rrISMldq0YTpfn0UkALciX","object":"text_completion","created":1789722032,"model":"qa-gat ...
req2 HTTP=200 WALL=1.386790s x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOkLN0WDMiGdRdRjCZCyLCIokFET","object":"text_completion","created":1789722033,"model":"qa-gat ...
req3 HTTP=200 WALL=1.886922s x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOkLCDfl5YZo6XNXbgSO8rcyrudM","object":"text_completion","created":1789722033,"model":"qa-gat ...
== /v1/messages ==
req2 HTTP=200 WALL=1.324235s x-litellm-attempted-retries: 0  {"id":"resp_bGl0ZWxsbTpjdXN0b21fbGxtX3Byb3ZpZGVyOm9wZW5haTttb2RlbF9pZDozNjg0Mjc2MmIwYWY0OTBkODVkNzdmNmEzMGRiZT ...
req3 HTTP=200 WALL=2.167326s x-litellm-attempted-retries: 0  {"id":"resp_bGl0ZWxsbTpjdXN0b21fbGxtX3Byb3ZpZGVyOm9wZW5haTttb2RlbF9pZDozNjg0Mjc2MmIwYWY0OTBkODVkNzdmNmEzMGRiZT ...
req1 HTTP=200 WALL=2.883347s x-litellm-attempted-retries: 0  {"id":"resp_bGl0ZWxsbTpjdXN0b21fbGxtX3Byb3ZpZGVyOm9wZW5haTttb2RlbF9pZDozNjg0Mjc2MmIwYWY0OTBkODVkNzdmNmEzMGRiZT ...
== /v1/responses ==
req1 HTTP=200 WALL=0.978036s x-litellm-attempted-retries: 0  {"id":"resp_zM-5qYj8RtfvINGENE_2hA9sZt5eayG2ET0Zxq4ZyC8KghVBEZHlQRcw-IVOxUm1VVRaAruWx83gG1EGmfD9mSRFs4kb40neOA ...
req2 HTTP=200 WALL=1.710033s x-litellm-attempted-retries: 0  {"id":"resp_W335fNGx-F5kXp354xacxSW4hOuSjK4Ytc3RLcyJC5LLnOsx1I96_EPFX-5BK205LnGDd-9xlLstzsAonaMZXw-so6skr4wOSf ...
req3 HTTP=200 WALL=2.320309s x-litellm-attempted-retries: 0  {"id":"resp_8mrQWn8TiHJy-kQN7KcXHn2UjNDeprTTgvB8ttGGqWwHFH_hmrV0F92udzvmukwRIrMHyU7szxh34oxV_Y50tEAK9DjhImOSgh ...

4. /v1/embeddings on a capped embedding deployment

  1. bash fire_ep.sh 20676 qa-embed-gated 3 embeddings
  2. Observed: all 3 return 200, serialized
req3 HTTP=200 WALL=0.508199s x-litellm-attempted-retries: 0  {"model":"qa-embed-gated","data":[{"embedding":[-0.0265045166015625,0.0259552001953125,0.0029449462890625,0.04 ...
req1 HTTP=200 WALL=0.763287s x-litellm-attempted-retries: 0  {"model":"qa-embed-gated","data":[{"embedding":[-0.026519775390625,0.025909423828125,0.002902984619140625,0.04 ...
req2 HTTP=200 WALL=0.991217s x-litellm-attempted-retries: 0  {"model":"qa-embed-gated","data":[{"embedding":[-0.0265045166015625,0.0259552001953125,0.0029449462890625,0.04 ...

5. Streaming /v1/chat/completions on the capped deployment

  1. bash fire_ep.sh 20676 qa-gated 3 stream
  2. Observed: all 3 stream 200, serialized
req3 HTTP=200 WALL=0.452714s x-litellm-attempted-retries: 0  data: {"id":"chatcmpl-EPOkSYueHGBgoWadStWzukBco6GEl","object":"chat.completion.chunk","created":1789722040,"mo ...
req1 HTTP=200 WALL=0.991346s x-litellm-attempted-retries: 0  data: {"id":"chatcmpl-EPOkScdeV8AYCy0ULaDJTT85epBnv","object":"chat.completion.chunk","created":1789722040,"mo ...
req2 HTTP=200 WALL=1.515234s x-litellm-attempted-retries: 0  data: {"id":"chatcmpl-EPOkTQzSkPD3BzYx9L9ztkTwyJxxn","object":"chat.completion.chunk","created":1789722041,"mo ...

6. Control: deployment with no cap

  1. bash fire.sh 20676 qa-open 6
  2. Observed: all 6 run in parallel
req4 HTTP=200 WALL=0.538150s x-litellm-response-duration-ms: 509.806 x-litellm-overhead-duration-ms: 55.985 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOkTvJg672giVQa952gIMrbx6csQ","created":1789722041,"model":"qa-open","object":"chat.completio ...
req5 HTTP=200 WALL=0.572033s x-litellm-response-duration-ms: 436.209 x-litellm-overhead-duration-ms: 7.461 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOkTYoWrpSzgisT5ZT8iiuQu5Sx2","created":1789722041,"model":"qa-open","object":"chat.completio ...
req3 HTTP=200 WALL=0.583686s x-litellm-response-duration-ms: 531.093 x-litellm-overhead-duration-ms: 39.189 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOkTBEMHgjzxErG0Pzu3961MK4K3","created":1789722041,"model":"qa-open","object":"chat.completio ...
req2 HTTP=200 WALL=0.594099s x-litellm-response-duration-ms: 518.622 x-litellm-overhead-duration-ms: 58.82 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOkTxwK7t8BY8dsWsDXixRzhWeSS","created":1789722041,"model":"qa-open","object":"chat.completio ...
req1 HTTP=200 WALL=0.641359s x-litellm-response-duration-ms: 547.314 x-litellm-overhead-duration-ms: 44.427 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOkUY0ZJSK6CpP7y9SRGTKRP6ETF","created":1789722042,"model":"qa-open","object":"chat.completio ...
req6 HTTP=200 WALL=0.871296s x-litellm-response-duration-ms: 757.554 x-litellm-overhead-duration-ms: 25.159 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOkUJAPKco687TEcsfrq2Sh6D3R2","created":1789722042,"model":"qa-open","object":"chat.completio ...

7. /v1/chat/completions, max_parallel_requests: 1, proxy default retries

  1. same config without the router_settings block on port 20677, bash fire.sh 20677 qa-gated 4
  2. Observed: the same staircase as case 1; retries never fire because nothing fails
req3 HTTP=200 WALL=0.893473s x-litellm-response-duration-ms: 735.154 x-litellm-overhead-duration-ms: 56.07 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOlYFXyzwkRhWwk1KMQDWDHQKEnl","created":1789722108,"model":"qa-gated","object":"chat.completi ...
req2 HTTP=200 WALL=1.495322s x-litellm-response-duration-ms: 597.823 x-litellm-overhead-duration-ms: 22.571 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOlYO0KfGsbT2B5pTHzstv9HN33O","created":1789722108,"model":"qa-gated","object":"chat.completi ...
req4 HTTP=200 WALL=2.024748s x-litellm-response-duration-ms: 524.404 x-litellm-overhead-duration-ms: 16.15 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOlZhHyaCuqMgqyyMsHQETFeuf0U","created":1789722109,"model":"qa-gated","object":"chat.completi ...
req1 HTTP=200 WALL=2.488958s x-litellm-response-duration-ms: 462.507 x-litellm-overhead-duration-ms: 12.821 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOlZX01FYrDNaptol4ZnHJ7zuvor","created":1789722109,"model":"qa-gated","object":"chat.completi ...

8. Admin UI on the Before proxy

  1. Log in to http://localhost:20676/ui/ as admin with the master key, open Logs, Filters, Error Code 429 - Rate Limited, Apply
  2. Observed: nothing to show, every burst request in cases 1 to 7 above returned 200, so the Logs page only has Success rows for the capped deployments and no 429 row exists to open

After (61ce1b4)

sha=61ce1b46d9842ef19a95f9520d430e72c2a093e7
litellm.__file__ /home/ubuntu/repos/litellm/litellm/__init__.py
has_fix True

1. /v1/chat/completions, max_parallel_requests: 1

  1. bash fire.sh 20686 qa-gated 6
  2. Observed: one 200 in under a second, five 429s in about 0.35 s each with "type":"throttling_error" and a message naming the deployment and its cap; the 429s never reach the provider
req3 HTTP=429 WALL=0.352242s  {"error":{"message":"litellm.RateLimitError: Deployment has all max_parallel_requests slots in use. Deployment ...
req1 HTTP=429 WALL=0.363978s  {"error":{"message":"litellm.RateLimitError: Deployment has all max_parallel_requests slots in use. Deployment ...
req2 HTTP=429 WALL=0.371360s  {"error":{"message":"litellm.RateLimitError: Deployment has all max_parallel_requests slots in use. Deployment ...
req6 HTTP=429 WALL=0.379031s  {"error":{"message":"litellm.RateLimitError: Deployment has all max_parallel_requests slots in use. Deployment ...
req4 HTTP=429 WALL=0.386748s  {"error":{"message":"litellm.RateLimitError: Deployment has all max_parallel_requests slots in use. Deployment ...
req5 HTTP=200 WALL=0.961530s x-litellm-response-duration-ms: 821.626 x-litellm-overhead-duration-ms: 79.273 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOpQ5AtEUAcV58Ap7ZptjUtbpQmz","created":1789722348,"model":"qa-gated","object":"chat.completi ...

2. /v1/chat/completions, cap derived from tpm: 150

  1. bash fire.sh 20686 qa-tpm-gated 6
  2. Observed: one 200, five immediate 429s, so the derived cap rejects the same way
req1 HTTP=429 WALL=0.133791s  {"error":{"message":"litellm.RateLimitError: Deployment has all max_parallel_requests slots in use. Deployment ...
req5 HTTP=429 WALL=0.140322s  {"error":{"message":"litellm.RateLimitError: Deployment has all max_parallel_requests slots in use. Deployment ...
req4 HTTP=429 WALL=0.145098s  {"error":{"message":"litellm.RateLimitError: Deployment has all max_parallel_requests slots in use. Deployment ...
req3 HTTP=429 WALL=0.146670s  {"error":{"message":"litellm.RateLimitError: Deployment has all max_parallel_requests slots in use. Deployment ...
req6 HTTP=429 WALL=0.154072s  {"error":{"message":"litellm.RateLimitError: Deployment has all max_parallel_requests slots in use. Deployment ...
req2 HTTP=200 WALL=0.677955s x-litellm-response-duration-ms: 645.021 x-litellm-overhead-duration-ms: 68.539 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOpQE9tE1MxkufqVdhlfENdZZLmn","created":1789722348,"model":"qa-tpm-gated","object":"chat.comp ...

3. /v1/completions, /v1/messages and /v1/responses on the capped deployment

  1. bash fire_ep.sh 20686 qa-gated 3 completions, bash fire_ep.sh 20686 qa-gated 3 messages, bash fire_ep.sh 20686 qa-gated 3 responses
  2. Observed: each endpoint gives one 200 and two 429s in under 0.1 s; /v1/messages returns the 429 in Anthropic error shape
req2 HTTP=429 WALL=0.085300s  {"error":{"message":"litellm.RateLimitError: Deployment has all max_parallel_requests slots in use. Deployment ...
req3 HTTP=429 WALL=0.089766s  {"error":{"message":"litellm.RateLimitError: Deployment has all max_parallel_requests slots in use. Deployment ...
req1 HTTP=200 WALL=0.667626s x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOpRvp8Rsbn2ZrUjTVxM7doLJ98m","object":"text_completion","created":1789722349,"model":"qa-gat ...
== /v1/messages ==
req1 HTTP=429 WALL=0.098626s  {"type":"error","error":{"type":"rate_limit_error","message":"litellm.RateLimitError: Deployment has all max_p ...
req3 HTTP=429 WALL=0.103521s  {"type":"error","error":{"type":"rate_limit_error","message":"litellm.RateLimitError: Deployment has all max_p ...
req2 HTTP=200 WALL=0.949890s x-litellm-attempted-retries: 0  {"id":"resp_bGl0ZWxsbTpjdXN0b21fbGxtX3Byb3ZpZGVyOm9wZW5haTttb2RlbF9pZDozNjg0Mjc2MmIwYWY0OTBkODVkNzdmNmEzMGRiZT ...
== /v1/responses ==
req3 HTTP=429 WALL=0.090930s  {"error":{"message":"litellm.RateLimitError: Deployment has all max_parallel_requests slots in use. Deployment ...
req2 HTTP=429 WALL=0.090962s  {"error":{"message":"litellm.RateLimitError: Deployment has all max_parallel_requests slots in use. Deployment ...
req1 HTTP=200 WALL=0.681928s x-litellm-attempted-retries: 0  {"id":"resp_KqaIa41UJTupEjECGw5twfF_Bun-S7kdrbaS4VLJ0FT-vCk9ZE3f_urOr5f30v0tWB_8Hq97bSGud0vijHGo1QM56V7uLTyiUS ...

4. /v1/embeddings on a capped embedding deployment

  1. bash fire_ep.sh 20686 qa-embed-gated 3 embeddings
  2. Observed: one 200, two immediate 429s
req3 HTTP=429 WALL=0.084969s  {"error":{"message":"litellm.RateLimitError: Deployment has all max_parallel_requests slots in use. Deployment ...
req2 HTTP=429 WALL=0.085033s  {"error":{"message":"litellm.RateLimitError: Deployment has all max_parallel_requests slots in use. Deployment ...
req1 HTTP=200 WALL=0.471711s x-litellm-attempted-retries: 0  {"model":"qa-embed-gated","data":[{"embedding":[-0.0265045166015625,0.0259552001953125,0.0029449462890625,0.04 ...

5. Streaming /v1/chat/completions on the capped deployment

  1. bash fire_ep.sh 20686 qa-gated 3 stream
  2. Observed: one request streams 200, the other two get an immediate 429 before any chunk
req2 HTTP=429 WALL=0.077870s  {"error":{"message":"litellm.RateLimitError: Deployment has all max_parallel_requests slots in use. Deployment ...
req1 HTTP=429 WALL=0.083270s  {"error":{"message":"litellm.RateLimitError: Deployment has all max_parallel_requests slots in use. Deployment ...
req3 HTTP=200 WALL=0.602744s x-litellm-attempted-retries: 0  data: {"id":"chatcmpl-EPOpTEZM2cP59O2olfyJvl9ZQ9vYI","object":"chat.completion.chunk","created":1789722352,"mo ...

6. Control: deployment with no cap

  1. bash fire.sh 20686 qa-open 6
  2. Observed: all 6 run in parallel, unchanged
req6 HTTP=200 WALL=0.654748s x-litellm-response-duration-ms: 488.518 x-litellm-overhead-duration-ms: 8.33 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOpVZSI6mPK5IqM4PiQqo2NUIQs7","created":1789722353,"model":"qa-open","object":"chat.completio ...
req3 HTTP=200 WALL=0.666200s x-litellm-response-duration-ms: 633.271 x-litellm-overhead-duration-ms: 45.253 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOpVx1Ny6T2bL99wZHLuSI9IDbiK","created":1789722353,"model":"qa-open","object":"chat.completio ...
req4 HTTP=200 WALL=0.679988s x-litellm-response-duration-ms: 622.465 x-litellm-overhead-duration-ms: 49.554 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOpVkPX04SaQgj1hSMjPcNqjq2cv","created":1789722353,"model":"qa-open","object":"chat.completio ...
req5 HTTP=200 WALL=0.780977s x-litellm-response-duration-ms: 640.627 x-litellm-overhead-duration-ms: 32.086 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOpVlnxs1OpZrh3KdknmHtM0iGJb","created":1789722353,"model":"qa-open","object":"chat.completio ...
req2 HTTP=200 WALL=1.029844s x-litellm-response-duration-ms: 916.448 x-litellm-overhead-duration-ms: 57.576 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOpVuXkGG7kJ8Pacoi5boBCJVHiI","created":1789722353,"model":"qa-open","object":"chat.completio ...
req1 HTTP=200 WALL=1.157321s x-litellm-response-duration-ms: 1072.989 x-litellm-overhead-duration-ms: 50.362 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOpVT6vNOAq6iUxRGFJ1JI1hxOsY","created":1789722353,"model":"qa-open","object":"chat.completio ...

7. /v1/chat/completions, max_parallel_requests: 1, proxy default retries

  1. same config without the router_settings block on port 20687, bash fire.sh 20687 qa-gated 4
  2. Observed: the router retries the 429 with backoff against the same deployment: three requests get through on attempt 0, 1 and 2 (x-litellm-attempted-retries), the fourth surfaces the 429 after about 5 s once retries are exhausted
req3 HTTP=200 WALL=0.811026s x-litellm-response-duration-ms: 705.822 x-litellm-overhead-duration-ms: 43.732 x-litellm-attempted-retries: 0  {"id":"chatcmpl-EPOpWLKx3pjqUKyKnoY1PcjdW7J77","created":1789722354,"model":"qa-gated","object":"chat.completi ...
req2 HTTP=200 WALL=1.379543s x-litellm-response-duration-ms: 541.817 x-litellm-overhead-duration-ms: 5.364 x-litellm-attempted-retries: 1  {"id":"chatcmpl-EPOpWkIG3rnr8S3mcZAyDYlricJR5","created":1789722354,"model":"qa-gated","object":"chat.completi ...
req4 HTTP=200 WALL=2.577462s x-litellm-response-duration-ms: 679.363 x-litellm-overhead-duration-ms: 5.572 x-litellm-attempted-retries: 2  {"id":"chatcmpl-EPOpYivWxG5kgHT8gSzXuqrLYTgdd","created":1789722356,"model":"qa-gated","object":"chat.completi ...
req1 HTTP=429 WALL=4.950064s  {"error":{"message":"litellm.RateLimitError: Deployment has all max_parallel_requests slots in use. Deployment ...

8. Admin UI on the After proxy

  1. Log in to http://localhost:20686/ui/ as admin with the master key, open Models + Endpoints, click the qa-gated row. The LiteLLM Params block shows the cap the router enforces

Models page, qa-gated deployment with max_parallel_requests 1

  1. Open Logs, Filters, Error Code 429 - Rate Limited, Apply. Every overflow request from cases 1 to 7 is a Failure row with a duration under 0.1 s (the 4.84 s row is the retry-exhausted request from case 7); before this change these rows did not exist because the requests waited and then succeeded

Logs page filtered to error code 429

  1. Click one of the rows. The detail panel shows Error Code: 429 and the message naming the deployment and its max_parallel_requests=1

Log detail with the 429 message

Type

🐛 Bug Fix

Caveats (if any)

Severe

  • default behavior change: bursts over a deployment's rpm / tpm / max_parallel_requests now 429 instead of queueing
    • operators who relied on the queue need to raise the cap, add a fallback deployment or handle 429s
    • the docs PR calls this out on the routing page

Medium

  • with the proxy's default retries, an overflow 429 is retried with backoff against the same deployment, so callers see it a few seconds late unless num_retries: 0
  • slots are per process: with several workers a burst can exceed the cap by up to one cap per worker (LIT-7024)

Low

  • MaxParallelRequestsLimit.in_flight is a mutable counter on purpose: check and increment must happen in one synchronous step for the same-loop race guarantee

Review gate status at 61ce1b4

CI: 89 check runs passed and one skipped (the stage-mirror e2e job is conditional), none failed or pending. Greptile: 5/5 with no outstanding findings, last reviewed commit 61ce1b4. CodeQL: passed with zero open review threads. Bugbot: a single bugbot run was posted by mateo-berri at 2026-09-18 09:14 UTC for this head and no Cursor check run or review started in the following 86 minutes; sibling PRs triggered in the same window are in the same state and the team's Cursor on-demand spend limit is reported as exhausted, so Bugbot is recorded as unavailable for this head rather than passed. The last Bugbot verdict on this branch is "found no new issues" on ancestor 8d972ee, and the only change since is the merge of origin/main with one test-file conflict resolution. Veria has not posted on this PR at any SHA

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

ran /live-pr-risk and found no regressions/backward incompatible risks

Link to Devin session: https://app.devin.ai/sessions/7b097835d9504224a951e5eda64247b7
Open in Devin Desktop: https://app.devin.ai/desktop/session/7b097835d9504224a951e5eda64247b7?variant=devin
Requested by: @yassin-berriai

yassin-berriai and others added 3 commits September 17, 2026 02:50
…29 on overflow

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ded deployments

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…instead of subclassing it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR. Add '(aside)' to your comment to have me ignore it.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@codspeed

codspeed Bot commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_max_parallel_requests_queue_size (61ce1b4) with main (8fc9c46)

Open in CodSpeed

@greptile-apps

greptile-apps Bot commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge, with no outstanding correctness, security, or repository-rule findings.

Summary

This PR replaces deployment-local concurrency semaphores with a fail-fast slot limiter, returning a 429 when all derived or configured max_parallel_requests slots are occupied.

  • Centralizes slot acquisition and release around provider calls across router endpoint families.
  • Avoids creating provider coroutines before slot admission.
  • Holds streaming slots until streams are exhausted or closed and releases slots on exceptions.
  • Adds coverage for overflow rejection, fallback and cooldown behavior, streaming lifecycle, embeddings, limit derivation, and uncapped deployments.
  • The latest merge from origin/main preserved the PR-owned limit-precedence test without changing this feature’s behavior.

Reviews (6) · Last reviewed commit: "Merge remote-tracking branch 'origin/mai..."

Comment thread litellm/router_utils/client_initalization_utils.py Outdated
Comment thread tests/code_coverage_tests/router_code_coverage.py Outdated
@codecov

codecov Bot commented Sep 17, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 96.42857% with 2 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
litellm/router.py 94.28% 2 Missing ⚠️

📢 Thoughts on this report? Let us know!

…ive integer

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

code-quality and documentation fail only because default_max_parallel_requests_queue_size is undocumented; BerriAI/litellm-docs#1513 adds the row, so they pass once it merges

…ts slots are all in use

Replace the per-deployment asyncio.Semaphore with MaxParallelRequestsLimit, which admits a call synchronously or raises the router's RateLimitError (429) right away. Nothing waits for a slot any more, so the max_parallel_requests_queue_size and default_max_parallel_requests_queue_size settings from the earlier commits are dropped along with their proxy validation, dashboard control and generated schema entries. The rpm/tpm derivation of the cap is unchanged. Every router endpoint family now enters the slot through one _deployment_slot context, and the provider coroutine is only created once the slot is held

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration devin-ai-integration Bot changed the title feat(router): bound the max_parallel_requests wait queue and return 429 on overflow feat(router): reject with 429 when a deployment's max_parallel_requests slots are all in use Sep 17, 2026
@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 8d972ee. Configure here.

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@veria-ai please review 8d972ee: semaphore replaced by synchronous slot admission that raises 429 when full, slot release on every exit path including streaming, fallback and cooldown behavior, rpm/tpm derivation unchanged

Comment thread litellm/router_utils/client_initalization_utils.py
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread litellm/router_utils/client_initalization_utils.py
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

…requests_queue_size

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

# Conflicts:
#	tests/test_litellm/test_utils.py
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai please re-review at 61ce1b4 (origin/main merged in, one test-file conflict resolved, no behavior change)

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor

cursor Bot commented Sep 18, 2026

Copy link
Copy Markdown
Contributor

Bugbot is paused — on-demand spend limit reached

Bugbot uses usage-based billing for this team and has hit its on-demand spend limit.

A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue.

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@veria-ai please review 61ce1b4: origin/main merged in, same synchronous slot admission returning 429 with no wait queue, check release on exceptions and streams

@yassin-berriai
yassin-berriai merged commit 660f3df into main Sep 18, 2026
93 checks passed
@yassin-berriai
yassin-berriai deleted the litellm_max_parallel_requests_queue_size branch September 18, 2026 18:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants