Skip to content

perf(proxy): one request-scoped Redis pipeline for auth, spend, rate-limit and routing reads - #43407

Merged
yassin-berriai merged 1 commit into
mainfrom
litellm_redis_p2_request_batch
Sep 29, 2026
Merged

yassin-berriai merged 1 commit into
mainfrom
litellm_redis_p2_request_batch

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Each request can share a Redis batch across authentication, spend checks, rate limits, and routing reads. Awaiting a result sends queued operations through a pipeline for that Redis backend, and the middleware flushes anything left when the request ends. This PR extends the earlier spend-counter batching work to the rest of the request path

The debug line logged when a routing prefetch fails to arm strips CR and LF from the request's model name and the error text, so a request cannot forge a log entry through it (CodeQL py/log-injection)

Files changed

File Change
litellm/caching/redis_batch.py Adds the request-scoped queue that sends Redis operations through shared pipelines
litellm/proxy/auth/auth_object_prefetch.py Batches auth-cache reads and write-backs, with the existing direct path as fallback
litellm/proxy/common_request_processing.py Arms routing reads before request admission flushes queued work
litellm/proxy/hooks/parallel_request_limiter_v3.py Sends rate-limit scripts through the shared pipeline and handles failed groups
litellm/proxy/middleware/redis_request_batch_middleware.py Opens a batch for each request and flushes remaining operations at the end
litellm/proxy/proxy_server.py Registers the request batch middleware
litellm/proxy/spend_tracking/spend_counter_batch.py Adds unread spend counters to the request pipeline and collects their results
litellm/router.py Arms cooldown and usage reads for the requested model
litellm/router_utils/routing_read_batch.py Uses prefetched cooldown and usage data during routing, with a fallback read

Request flow

flowchart LR
    R[Request opens batch] --> Q[Auth spend limits and routing queue Redis work]
    Q --> F[Await or request end flushes pending work]
    F --> P[Pipeline per Redis backend sends queued operations]
    P --> C[Callers receive results]
Loading

User Flow

Before: a developer whose key has a budget, a team budget, an end-user budget and TPM/RPM limits waits on 15 Redis round trips per request, 9 of them before the provider is called

  1. They send POST https://litellm-domain/v1/chat/completions with "model": "gpt-group" and "user": "perf-enduser"
  2. The proxy reads their identity from Redis, writes it back, reads the spend counters, reserves the budget, runs three separate rate-limit scripts, reads the deployment cooldown state, then checks the response cache: nine waits on Redis, one after the other
  3. The response comes back 200 with the mock reply

After: the same request waits on 5 Redis round trips before the provider is called, gets the same 200 and records the same spend

  1. They send the same POST https://litellm-domain/v1/chat/completions
  2. The proxy reads identity and spend counters in one pipeline, reserves the budget, then sends the identity write-back, the cooldown read and the RPM script in one pipeline, the two TPM scripts in one more, and checks the response cache
  3. The response comes back 200 with the same mock reply

Relevant issues

Stacked on #43369; merge after it. Third of the stacked PRs for one Redis pipeline pre-call and one post-call (#43320, then LIT-8881, LIT-8882, LIT-8883); design and per-request measurements are on LIT-8881

Linear ticket

Resolves LIT-8882

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/unit/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Screenshots / Proof of Fix

Same fixture bytes for both arms, PYTHONPATH=<checkout>, real Redis 6.0.16 at 127.0.0.1:6379, real Postgres, mock deployments (litellm_params.mock_response; the change is on the admission path, not a provider path). Every RedisCache method call and every RedisBatch flush is logged with its caller chain by a sitecustomize tracer loaded through PYTHONPATH; the tables list every call between the request marker and the response marker, minus the jobs that fire on timers (_sync_in_memory_spend_with_redis, daily tag spend flush, config prefetch). Calls are round trips: a pipeline or a Lua script counts once; a PIPELINE row names the ops it carried

Fixture: config_full.yaml with two gpt-group deployments (gpt-dep-a, gpt-dep-b, usage-based-routing-v2 routing group) and two shuffle-group deployments (shuffle-dep-a, shuffle-dep-b, top-level simple-shuffle), Redis response cache on, enable_redis_auth_cache: true. Virtual key with max_budget: 1000, tpm_limit: 10000000, rpm_limit: 100000, in team perf-team (same budget and limits), user with max_budget: 1000, end user perf-enduser on a budget of 1000. Requests are sent 12 s apart so the warm pass hits warm caches; the tables are the warm pass

python litellm/proxy/proxy_cli.py --config ~/perf_rt/config_full.yaml --port <port>
curl -s -H "Authorization: Bearer $KEY" -H "Content-Type: application/json" http://127.0.0.1:<port>/v1/chat/completions -d @chat_warm.req.json
# chat_warm.req.json
{"model":"gpt-group","user":"perf-enduser","messages":[{"role":"user","content":"Reply with the single word: pong B"}],"max_tokens":5}
# shuffle_warm.req.json
{"model":"shuffle-group","user":"perf-enduser","messages":[{"role":"user","content":"Reply with the single word: pong B"}],"max_tokens":5}
# chat_stream_warm.req.json
{"model":"gpt-group","user":"perf-enduser","stream":true,"stream_options":{"include_usage":true},"messages":[{"role":"user","content":"Reply with the single word: pong B"}],"max_tokens":5}
# responses_warm.req.json
{"model":"gpt-group","user":"perf-enduser","input":"Reply with the single word: pong B","max_output_tokens":5}
# messages_warm.req.json
{"model":"gpt-group","max_tokens":5,"messages":[{"role":"user","content":"Reply with the single word: pong B"}],"metadata":{"user_id":"perf-enduser"}}

Before (e8affcb, the #43369 tip this PR is stacked on)

POST /v1/chat/completions, non-streaming, usage-based routing

  1. HTTP 200, "model": "gpt-group", content mock reply from b
  2. 15 Redis round trips on the request path (9 pre-call, 6 post-call). Rows 1 to 9 are nine serial waits on Redis from five different owners
# phase RedisCache op keys caller
1 pre async_batch_get_cache ['<team>_<user>', 'team_membership:837848e9-670e-4 _fill_from_redis <- prefetch_auth_objects
2 pre async_set_cache_pipeline_with_ttls (('<team>_<user>', {'user_id': '837848e9-670e-4b34 _write_back <- _fill_from_db
3 pre async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _fetch <- _load
4 pre async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> _reserve_counters <- _reserve_reservable_counters
5 pre EVALSHA rate-limiter check ['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318 atch_rate_limiter_script <- should_rate_limit
6 pre EVALSHA tpm check-and-increment ['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318 heck_and_increment_by_n <- reserve_tpm_tokens
7 pre EVALSHA tpm check-and-increment ['{team:<team>}:window', '{team:<team>}:tokens'] heck_and_increment_by_n <- reserve_tpm_tokens
8 pre async_batch_get_cache ['deployment:gpt-dep-a:cooldown', 'deployment:gpt-dep-b:cooldown', 'deployment:shuffle-dep-a:cooldown', 'deplo _cooldown_deployments <- async_get_healthy_deployments
9 pre async_get_cache <key-hash> ache <- _retrieve_from_cache
10 post async_set_cache <key-hash> async_set_cache <- async_add_cache
11 post async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _fetch <- _load
12 post async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> ed <- increment_spend_counters
13 post async_increment gpt-dep-b:openai/gpt-4o-mini:tpm:<min> async_increment_cache <- async_log_success_event
14 post async_batch_get_cache ['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0'] async_batch_get_cache <- _read_update_cache_values
15 post EVALSHA rate-limiter check ['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3 _execute_token_increment_script <- async_increment_tokens_with_ttl_preservation

POST /v1/chat/completions, non-streaming, simple-shuffle

  1. HTTP 200, "model": "shuffle-group", content mock reply from s
  2. 15 Redis round trips on the request path (9 pre-call, 6 post-call), the same shape: the cooldown read (row 8) is its own round trip for every routing strategy
# phase RedisCache op keys caller
1 pre async_batch_get_cache ['<team>_<user>', 'team_membership:837848e9-670e-4 _fill_from_redis <- prefetch_auth_objects
2 pre async_set_cache_pipeline_with_ttls (('<team>_<user>', {'user_id': '837848e9-670e-4b34 _write_back <- _fill_from_db
3 pre async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _fetch <- _load
4 pre async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> _reserve_counters <- _reserve_reservable_counters
5 pre EVALSHA rate-limiter check ['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318 atch_rate_limiter_script <- should_rate_limit
6 pre EVALSHA tpm check-and-increment ['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318 heck_and_increment_by_n <- reserve_tpm_tokens
7 pre EVALSHA tpm check-and-increment ['{team:<team>}:window', '{team:<team>}:tokens'] heck_and_increment_by_n <- reserve_tpm_tokens
8 pre async_batch_get_cache ['deployment:gpt-dep-a:cooldown', 'deployment:gpt-dep-b:cooldown', 'deployment:shuffle-dep-a:cooldown', 'deplo _cooldown_deployments <- async_get_healthy_deployments
9 pre async_get_cache <key-hash> ache <- _retrieve_from_cache
10 post async_set_cache <key-hash> async_set_cache <- async_add_cache
11 post async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _fetch <- _load
12 post async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> ed <- increment_spend_counters
13 post async_increment shuffle-dep-a:openai/gpt-4o-mini:tpm:<min> async_increment_cache <- async_log_success_event
14 post async_batch_get_cache ['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0'] async_batch_get_cache <- _read_update_cache_values
15 post EVALSHA rate-limiter check ['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3 _execute_token_increment_script <- async_increment_tokens_with_ttl_preservation

POST /v1/chat/completions, streaming

  1. HTTP 200, 8 SSE chunks, "model": "gpt-group", content mock reply from b
  2. 16 Redis round trips on the request path (9 pre-call, 7 post-call)
# phase RedisCache op keys caller
1 pre async_batch_get_cache ['<team>_<user>', 'team_membership:837848e9-670e-4 _fill_from_redis <- prefetch_auth_objects
2 pre async_set_cache_pipeline_with_ttls (('<team>_<user>', {'user_id': '837848e9-670e-4b34 _write_back <- _fill_from_db
3 pre async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _fetch <- _load
4 pre async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> _reserve_counters <- _reserve_reservable_counters
5 pre EVALSHA rate-limiter check ['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318 atch_rate_limiter_script <- should_rate_limit
6 pre EVALSHA tpm check-and-increment ['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318 heck_and_increment_by_n <- reserve_tpm_tokens
7 pre EVALSHA tpm check-and-increment ['{team:<team>}:window', '{team:<team>}:tokens'] heck_and_increment_by_n <- reserve_tpm_tokens
8 pre async_batch_get_cache ['deployment:gpt-dep-a:cooldown', 'deployment:gpt-dep-b:cooldown', 'deployment:shuffle-dep-a:cooldown', 'deplo _cooldown_deployments <- async_get_healthy_deployments
9 pre async_get_cache <key-hash> ache <- _retrieve_from_cache
10 post async_set_cache <key-hash> async_set_cache <- async_add_cache
11 post async_set_cache <key-hash> async_set_cache <- async_add_cache
12 post async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _fetch <- _load
13 post async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> s <- _update_database_and_spend_counters_in_batch
14 post async_increment gpt-dep-b:None:tpm:<min> async_increment_cache <- async_log_success_event
15 post async_batch_get_cache ['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0'] async_batch_get_cache <- _read_update_cache_values
16 post EVALSHA rate-limiter check ['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3 _execute_token_increment_script <- async_increment_tokens_with_ttl_preservation

POST /v1/responses

  1. HTTP 200, "model": "gpt-group", content mock reply from b
  2. 18 Redis round trips on the request path (10 pre-call, 8 post-call); the SYNC rows are the pre-existing synchronous cache calls of the Responses path (noted on perf(router): fetch cooldown state and usage counters in one Redis round trip #43320)
# phase RedisCache op keys caller
1 pre async_batch_get_cache ['<team>_<user>', 'team_membership:837848e9-670e-4 _fill_from_redis <- prefetch_auth_objects
2 pre async_set_cache_pipeline_with_ttls (('<team>_<user>', {'user_id': '837848e9-670e-4b34 _write_back <- _fill_from_db
3 pre async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _fetch <- _load
4 pre async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> _reserve_counters <- _reserve_reservable_counters
5 pre EVALSHA rate-limiter check ['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318 _batch_rate_limiter_script <- should_rate_limit
6 pre EVALSHA tpm check-and-increment ['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318 _check_and_increment_by_n <- reserve_tpm_tokens
7 pre EVALSHA tpm check-and-increment ['{team:<team>}:window', '{team:<team>}:tokens'] _check_and_increment_by_n <- reserve_tpm_tokens
8 pre async_batch_get_cache ['deployment:gpt-dep-a:cooldown', 'deployment:gpt-dep-b:cooldown', 'deployment:shuffle-dep-a:cooldown', 'deplo _api_call_with_fallbacks_responses_attempt <- make_call
9 pre async_get_cache <key-hash> _ageneric_api_call_with_fallbacks_responses_attempt <- make_call
10 pre get_cache SYNC <key-hash> get_cache <- _sync_get_cache
11 post set_cache SYNC <key-hash> add_cache <- sync_set_cache
12 post increment_cache SYNC gpt-dep-b:openai/gpt-4o-mini:tpm:<min> increment_cache <- log_success_event
13 post async_set_cache <key-hash> async_set_cache <- async_add_cache
14 post async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _fetch <- _load
15 post async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> ed <- increment_spend_counters
16 post async_increment gpt-dep-b:openai/gpt-4o-mini:tpm:<min> async_increment_cache <- async_log_success_event
17 post async_batch_get_cache ['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0'] async_batch_get_cache <- _read_update_cache_values
18 post EVALSHA rate-limiter check ['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3 _execute_token_increment_script <- async_increment_tokens_with_ttl_preservation

POST /v1/messages

  1. HTTP 200, "model": "gpt-group", content mock reply from a
  2. 31 Redis round trips on the request path (25 pre-call, 6 post-call). Rows 1 to 13 and 17 to 19 are the pre-existing per-object GET/SET/DEL of the Messages path (side finding on LIT-8881, out of scope here); rows 14 to 16 and 20 to 25 are the same admission shape as chat
# phase RedisCache op keys caller
1 pre async_get_cache end_user_id:perf-enduser async_get_cache <- async_get_cache
2 pre async_get_cache end_user_restricted_registry async_get_cache <- async_get_cache
3 pre async_get_cache end_user_restricted_registry async_get_cache <- async_get_cache
4 pre async_set_cache end_user_restricted_registry _cache_registry_answer <- _fetch_and_cache_registry
5 pre async_set_cache end_user_id:perf-enduser async_set_cache <- async_set_cache
6 pre async_get_cache <key-hash> async_get_cache <- async_get_cache
7 pre async_get_cache <key-hash> async_get_cache <- async_get_cache
8 pre async_set_cache <key-hash> async_set_cache <- async_set_cache
9 pre async_get_cache team_id:<team> get_team_object <- _get_hierarchical_router_settings
10 pre async_delete_cache team_id:<team> get_team_object <- _get_hierarchical_router_settings
11 pre async_set_cache team_id:<team> get_team_object <- _get_hierarchical_router_settings
12 pre delete_cache SYNC team_alias:perf-team get_team_object <- _get_hierarchical_router_settings
13 pre async_delete_cache team_alias:perf-team get_team_object <- _get_hierarchical_router_settings
14 pre async_batch_get_cache ['<team>_<user>', '837848e9-670e-4b34-936e-156409a _fill_from_redis <- prefetch_auth_objects
15 pre async_set_cache_pipeline_with_ttls (('<user>', {'user_id': '<user>', 'user_alias': No _write_back <- _fill_from_db
16 pre async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _fetch <- _load
17 pre async_get_cache model_access_group_registry cached_registry <- _load_bounded_registry
18 pre async_get_cache model_access_group_registry cached_registry <- _load_bounded_registry
19 pre async_set_cache model_access_group_registry _cache_registry <- _load_bounded_registry
20 pre async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> _reserve_counters <- _reserve_reservable_counters
21 pre EVALSHA rate-limiter check ['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318 h_rate_limiter_script <- should_rate_limit
22 pre EVALSHA tpm check-and-increment ['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318 k_and_increment_by_n <- reserve_tpm_tokens
23 pre EVALSHA tpm check-and-increment ['{team:<team>}:window', '{team:<team>}:tokens'] k_and_increment_by_n <- reserve_tpm_tokens
24 pre async_batch_get_cache ['deployment:gpt-dep-a:cooldown', 'deployment:gpt-dep-b:cooldown', 'deployment:shuffle-dep-a:cooldown', 'deplo pt <- make_call
25 pre async_get_cache <key-hash> ropic_messages_attempt <- make_call
26 post async_set_cache <key-hash> async_add_cache <- _complete_cache_write_despite_cancellation
27 post async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _fetch <- _load
28 post async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> ed <- increment_spend_counters
29 post async_increment gpt-dep-a:None:tpm:<min> async_increment_cache <- async_log_success_event
30 post async_batch_get_cache ['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0'] async_batch_get_cache <- _read_update_cache_values
31 post EVALSHA rate-limiter check ['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3 _execute_token_increment_script <- async_increment_tokens_with_ttl_preservation

After (32c0727; the follow-up commits add suppression reasons, a router coverage exemption and the review fixes: a rejected or failed rate-limit group now refunds every group the pipeline incremented and a Redis denial stands even when another group fails, a reply one op cannot decode fails that op alone, and a cooldown recorded in memory after the prefetch left still wins; the round-trip counts are unchanged)

POST /v1/chat/completions, non-streaming, usage-based routing

  1. HTTP 200, "model": "gpt-group", content mock reply from a (the other deployment of the same group; both served requests in both arms)
  2. 11 Redis round trips on the request path (5 pre-call, 6 post-call). Row 1 is the identity MGET and the spend MGET in one pipeline; row 3 is the identity write-back (two SETs), the cooldown read and the RPM rate-limit script in one pipeline; row 4 is the two TPM scripts in one pipeline. Post-call is unchanged (LIT-8883)
  3. Every request of the run returned 200 and the proxy log has no pipeline failure or traceback
# phase RedisCache op keys caller
1 pre PIPELINE MGET+MGET ['<team>_<user>', 'team_membership:837848e9-670e-4 _read_redis_rows <- _fill_from_redis
2 pre async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> _reserve_counters <- _reserve_reservable_counters
3 pre PIPELINE SET+SET+MGET+EVALSHA ['<team>_<user>', 'team_membership:837848e9-670e-4 atch_rate_limiter_script <- should_rate_limit
4 pre PIPELINE EVALSHA+EVALSHA ['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318 heck_and_increment_by_n <- reserve_tpm_tokens
5 pre async_get_cache <key-hash> ache <- _retrieve_from_cache
6 post async_set_cache <key-hash> async_set_cache <- async_add_cache
7 post PIPELINE MGET ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _collect_inflight <- _load
8 post async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> ed <- increment_spend_counters
9 post async_increment gpt-dep-a:openai/gpt-4o-mini:tpm:<min> async_increment_cache <- async_log_success_event
10 post async_batch_get_cache ['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0'] async_batch_get_cache <- _read_update_cache_values
11 post EVALSHA rate-limiter check ['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3 _execute_token_increment_script <- async_increment_tokens_with_ttl_preservation

POST /v1/chat/completions, non-streaming, simple-shuffle

  1. HTTP 200, "model": "shuffle-group", content mock reply from s
  2. 11 Redis round trips on the request path (5 pre-call, 6 post-call); the cooldown read now rides the rate limiter's pipeline (row 3) for this strategy too
# phase RedisCache op keys caller
1 pre PIPELINE MGET+MGET ['<team>_<user>', 'team_membership:837848e9-670e-4 _read_redis_rows <- _fill_from_redis
2 pre async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> _reserve_counters <- _reserve_reservable_counters
3 pre PIPELINE SET+SET+MGET+EVALSHA ['<team>_<user>', 'team_membership:837848e9-670e-4 atch_rate_limiter_script <- should_rate_limit
4 pre PIPELINE EVALSHA+EVALSHA ['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318 heck_and_increment_by_n <- reserve_tpm_tokens
5 pre async_get_cache <key-hash> ache <- _retrieve_from_cache
6 post async_set_cache <key-hash> async_set_cache <- async_add_cache
7 post PIPELINE MGET ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _collect_inflight <- _load
8 post async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> ed <- increment_spend_counters
9 post async_increment shuffle-dep-a:openai/gpt-4o-mini:tpm:<min> async_increment_cache <- async_log_success_event
10 post async_batch_get_cache ['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0'] async_batch_get_cache <- _read_update_cache_values
11 post EVALSHA rate-limiter check ['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3 _execute_token_increment_script <- async_increment_tokens_with_ttl_preservation

POST /v1/chat/completions, streaming

  1. HTTP 200, 8 SSE chunks, "model": "gpt-group", content mock reply from a
  2. 12 Redis round trips on the request path (5 pre-call, 7 post-call)
# phase RedisCache op keys caller
1 pre PIPELINE MGET+MGET ['<team>_<user>', 'team_membership:837848e9-670e-4 _read_redis_rows <- _fill_from_redis
2 pre async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> _reserve_counters <- _reserve_reservable_counters
3 pre PIPELINE SET+SET+MGET+EVALSHA ['<team>_<user>', 'team_membership:837848e9-670e-4 atch_rate_limiter_script <- should_rate_limit
4 pre PIPELINE EVALSHA+EVALSHA ['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318 heck_and_increment_by_n <- reserve_tpm_tokens
5 pre async_get_cache <key-hash> ache <- _retrieve_from_cache
6 post async_set_cache <key-hash> async_set_cache <- async_add_cache
7 post async_set_cache <key-hash> async_set_cache <- async_add_cache
8 post PIPELINE MGET ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _collect_inflight <- _load
9 post async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> s <- _update_database_and_spend_counters_in_batch
10 post async_increment gpt-dep-a:None:tpm:<min> async_increment_cache <- async_log_success_event
11 post async_batch_get_cache ['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0'] async_batch_get_cache <- _read_update_cache_values
12 post EVALSHA rate-limiter check ['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3 _execute_token_increment_script <- async_increment_tokens_with_ttl_preservation

POST /v1/responses

  1. HTTP 200, "model": "gpt-group", content mock reply from b
  2. 14 Redis round trips on the request path (6 pre-call, 8 post-call)
# phase RedisCache op keys caller
1 pre PIPELINE MGET+MGET ['<team>_<user>', 'team_membership:837848e9-670e-4 _read_redis_rows <- _fill_from_redis
2 pre async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> _reserve_counters <- _reserve_reservable_counters
3 pre PIPELINE SET+SET+MGET+EVALSHA ['<team>_<user>', 'team_membership:837848e9-670e-4 _batch_rate_limiter_script <- should_rate_limit
4 pre PIPELINE EVALSHA+EVALSHA ['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318 _check_and_increment_by_n <- reserve_tpm_tokens
5 pre async_get_cache <key-hash> _ageneric_api_call_with_fallbacks_responses_attempt <- make_call
6 pre get_cache SYNC <key-hash> get_cache <- _sync_get_cache
7 post set_cache SYNC <key-hash> add_cache <- sync_set_cache
8 post increment_cache SYNC gpt-dep-b:openai/gpt-4o-mini:tpm:<min> increment_cache <- log_success_event
9 post async_set_cache <key-hash> async_set_cache <- async_add_cache
10 post PIPELINE MGET ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _collect_inflight <- _load
11 post async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> ed <- increment_spend_counters
12 post async_increment gpt-dep-b:openai/gpt-4o-mini:tpm:<min> async_increment_cache <- async_log_success_event
13 post async_batch_get_cache ['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0'] async_batch_get_cache <- _read_update_cache_values
14 post EVALSHA rate-limiter check ['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3 _execute_token_increment_script <- async_increment_tokens_with_ttl_preservation

POST /v1/messages

  1. HTTP 200, "model": "gpt-group", content mock reply from a
  2. 27 Redis round trips on the request path (21 pre-call, 6 post-call); the admission rows collapse the same way, the per-object GET/SET/DEL rows of the Messages path are untouched
# phase RedisCache op keys caller
1 pre async_get_cache end_user_id:perf-enduser async_get_cache <- async_get_cache
2 pre async_get_cache end_user_restricted_registry async_get_cache <- async_get_cache
3 pre async_get_cache end_user_restricted_registry async_get_cache <- async_get_cache
4 pre async_set_cache end_user_restricted_registry _cache_registry_answer <- _fetch_and_cache_registry
5 pre async_set_cache end_user_id:perf-enduser async_set_cache <- async_set_cache
6 pre async_get_cache <key-hash> async_get_cache <- async_get_cache
7 pre async_get_cache <key-hash> async_get_cache <- async_get_cache
8 pre async_set_cache <key-hash> async_set_cache <- async_set_cache
9 pre async_get_cache team_id:<team> get_team_object <- _get_hierarchical_router_settings
10 pre async_delete_cache team_id:<team> get_team_object <- _get_hierarchical_router_settings
11 pre async_set_cache team_id:<team> get_team_object <- _get_hierarchical_router_settings
12 pre delete_cache SYNC team_alias:perf-team get_team_object <- _get_hierarchical_router_settings
13 pre async_delete_cache team_alias:perf-team get_team_object <- _get_hierarchical_router_settings
14 pre PIPELINE MGET+MGET ['<team>_<user>', '837848e9-670e-4b34-936e-156409a _read_redis_rows <- _fill_from_redis
15 pre async_get_cache model_access_group_registry cached_registry <- _load_bounded_registry
16 pre async_get_cache model_access_group_registry cached_registry <- _load_bounded_registry
17 pre async_set_cache model_access_group_registry _cache_registry <- _load_bounded_registry
18 pre async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> _reserve_counters <- _reserve_reservable_counters
19 pre PIPELINE SET+SET+SET+MGET+EVALSHA ['<user>', '<team>_837848e9-670e-4b34-936e-156409a h_rate_limiter_script <- should_rate_limit
20 pre PIPELINE EVALSHA+EVALSHA ['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318 k_and_increment_by_n <- reserve_tpm_tokens
21 pre async_get_cache <key-hash> ropic_messages_attempt <- make_call
22 post async_set_cache <key-hash> async_add_cache <- _complete_cache_write_despite_cancellation
23 post PIPELINE MGET ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _collect_inflight <- _load
24 post async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> ed <- increment_spend_counters
25 post async_increment gpt-dep-a:None:tpm:<min> async_increment_cache <- async_log_success_event
26 post async_batch_get_cache ['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0'] async_batch_get_cache <- _read_update_cache_values
27 post EVALSHA rate-limiter check ['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3 _execute_token_increment_script <- async_increment_tokens_with_ttl_preservation

The five pre-call round trips left on the chat path are separated by data dependencies, not by ownership: the reservation increments (row 2) need the spend read; the rate limiter (row 3) runs after auth and its result gates the TPM scripts (row 4); the response-cache key (row 5) depends on the deployment routing picked. Getting from 5 to 1 needs the speculative admission mode and the Lua check-and-reserve of LIT-8884, which this PR does not ship

Not measured here: a fleet load regime. These are single-request command counts on one box; the prod traces on LIT-8881 are what motivated the change

The fail-closed follow-up at b763921 makes a pipelined rate-limit or TPM group that Redis could not verify refund every applied group and then raise RateLimitUnverifiableError when fail-closed is on, instead of quietly falling back to in-memory counters. Dropping the _reject_if_rate_limit_unverifiable call turns both cases of test_fail_closed_rejects_when_a_pipelined_lua_group_cannot_be_verified red, and restoring it turns them green. 6c6354a only narrows the fake Redis types and pins the cooldown test timestamps

Re-run at the current head 6c6354a against real Redis, Postgres and Anthropic claude-sonnet-4-6, all requests 200:

$ bash run_arm2.sh head43407final ~/repos/litellm
warm 15 reads, chat_master 9, chat_vk 14, messages_vk 16   (redis total_reads_processed deltas)

Admin UI at 6c6354a

The same virtual key request from the fixture, POST http://localhost:4000/v1/chat/completions with sk-perf-vk-0000000001 against Anthropic claude-sonnet-4-6, shows up on http://localhost:4000/ui/?page=logs as a Success row with its real spend, so the batched Redis path still records spend end to end

Request Logs page after a live request at 6c6354a853

Live re-check at ec4dd90

This head keys the request batch by the cache's namespace as well as its connection settings, so two caches on one server with different prefixes each get their own pipeline and key prefix (test_caches_on_one_server_with_different_namespaces_keep_their_own_key_prefix, red when the namespace is dropped from _backend_key). A read-count harness, run from this checkout with PYTHONPATH=<checkout>, one warm request then one request per case, redis_reads_processed from INFO stats minus the three marker commands

$ bash run_arm2.sh <arm> <checkout>  # redis-cli flushall, proxy on :4000 with config_v2.yaml, real Redis 7, real Postgres, Anthropic claude-sonnet-4-6
key/generate 200
warm 200
warm redis_reads_processed=15
chat_master 200
chat_master redis_reads_processed=8
chat_vk 200
chat_vk redis_reads_processed=14
messages_vk 200
messages_vk redis_reads_processed=16

Every response body was pong from claude-group with usage populated, and the proxy log has no ERROR or Traceback lines

Type

🚄 Infrastructure

Caveats (if any)

Medium

  • Awaiting one result flushes every op declared so far, so a consumer that awaits early splits the pipeline. That is why reservation and the TPM scripts are separate round trips today
  • Auth write-backs now leave with the next flush (the rate limiter's pipeline in this fixture) or at request exit when nothing else flushes. The in-memory copy is updated at once as before; another worker reading Redis inside that window fills from the DB, as it would on a cache miss
  • On a RedisClusterCache every declared op still runs on its own, as before this PR. The cluster pipeline is per node and the existing per-op paths already group by slot

Low

  • A single rate-limit Lua call also rides the request batch now, so a request with one hash-tag group shares the round trip with the routing read. Outside a request scope (background jobs, SDK use) every caller runs exactly as before
  • Non-usage routing strategies only go through RoutingReadBatch when the request armed a cooldown prefetch; otherwise the router's plain cooldown read stays in charge, so callers and tests that replace litellm.router._async_get_cooldown_deployments still take effect
  • The proxy's cache and the router's cache share one pipeline when their connection settings compare equal as strings (the router receives its port as a string). Two caches with different settings get separate pipelines, flushed concurrently
  • redis.exceptions is imported lazily inside redis_batch.py, like redis_cache.py, so import litellm without redis installed keeps working

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Link to Devin session: https://app.devin.ai/sessions/a666ccc35fde4ab799d68f3cc5bad523
Open in Devin Desktop: https://app.devin.ai/desktop/session/a666ccc35fde4ab799d68f3cc5bad523?variant=devin
Requested by: @yassin-berriai

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@codecov

codecov Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 90.94828% with 42 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
...itellm/proxy/spend_tracking/spend_counter_batch.py 29.62% 19 Missing ⚠️
litellm/caching/redis_batch.py 95.70% 11 Missing ⚠️
litellm/proxy/hooks/parallel_request_limiter_v3.py 90.00% 7 Missing ⚠️
litellm/proxy/auth/auth_object_prefetch.py 70.58% 5 Missing ⚠️

📢 Thoughts on this report? Let us know!

@greptile-apps

greptile-apps Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[Critical risk] Batches Redis reads across auth, rate limiting, and routing in one pipeline per request.

The PR appears safe to merge based on the reviewed changes and resolved previous findings.

Summary

The PR adds request-scoped Redis pipelines for authentication, spend, rate-limit, and routing reads.

  • The changes since the previous review separate caches with different namespaces into distinct batches and add a regression test.
  • No new actionable issue was identified.

Reviews (15) · Last reviewed commit: "fix(caching): key the request Redis batc..."

Comment thread litellm/proxy/hooks/parallel_request_limiter_v3.py
Comment thread litellm/caching/redis_batch.py Outdated
Comment thread litellm/router_utils/routing_read_batch.py Outdated
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread litellm/proxy/hooks/parallel_request_limiter_v3.py
Comment thread litellm/router.py Outdated
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread tests/unit/caching/test_request_redis_batch_pre_call.py Outdated
Comment thread tests/unit/caching/test_request_redis_batch_pre_call.py Outdated
Comment thread tests/unit/caching/test_request_redis_batch_pre_call.py Outdated
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

1 similar comment
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@yassin-berriai

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/proxy/hooks/parallel_request_limiter_v3.py
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@yassin-berriai

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit ec4dd90. Configure here.

@devin-ai-integration
devin-ai-integration Bot force-pushed the litellm_redis_p1_spend_batch_request_scope branch from 178bd46 to a51b69e Compare September 29, 2026 21:33
@devin-ai-integration
devin-ai-integration Bot force-pushed the litellm_redis_p2_request_batch branch from ec4dd90 to 13e171c Compare September 29, 2026 21:33
Base automatically changed from litellm_redis_p1_spend_batch_request_scope to main September 29, 2026 22:05
@yassin-berriai
yassin-berriai requested a review from a team September 29, 2026 22:05
@devin-ai-integration
devin-ai-integration Bot force-pushed the litellm_redis_p2_request_batch branch from 13e171c to fada595 Compare September 29, 2026 22:07
Comment thread litellm/router.py Fixed
@codspeed

codspeed Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_redis_p2_request_batch (2de50b6) with main (e7460f1)

Open in CodSpeed

…limit and routing reads

RedisBatch: one pipeline per Redis backend for independently declared operations (MGET, GET, Lua
scripts, INCRBYFLOAT, SET, DEL), a future per operation so each owner keeps its own fallback, Redis
Cluster hash-slot fallback. A request-scoped batch middleware shares that pipeline across the auth
identity reads and write-back, the spend counter MGET, the rate limiter Lua groups and the routing
read. A rate-limit denial stands when another pipelined group fails; every pipelined group is refunded
on rejection; local cooldowns win over the prefetch.

The routing prefetch failure log line strips request line breaks (CodeQL py/log-injection)

Resolves LIT-8882

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

This branch is waiting to be deployed

1 waiting deployment
e2e-changed — 2de50b69 Waiting Sep 29, 2026 by devin-ai-integration[bot] via oauth #1890
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants