Skip to content

perf(proxy): one post-call Redis pipeline per backend for spend, rate-limit, routing and response-cache writes - #43424

Merged
yassin-berriai merged 1 commit into
litellm_redis_p2_request_batchfrom
litellm_redis_p3_post_call_batch
Sep 29, 2026
Merged

yassin-berriai merged 1 commit into
litellm_redis_p2_request_batchfrom
litellm_redis_p3_post_call_batch

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • A governed request still makes several separate Redis trips after the provider responds
  • Spend, rate limits, routing usage, response caching and update_cache do their own work

How it solves it:

  • Each Redis backend gets one post-call batch per request
  • Spend, rate-limit, routing and cache writes share that batch
  • update_cache reads join the existing request read pipeline when possible
  • The batch flushes after callbacks finish, with a one-second fallback timer

Files changed

Files What changed
litellm/caching/caching.py, litellm/caching/dual_cache.py Defer eligible response-cache sets and post-call increments into the shared batch
litellm/caching/redis_batch.py Add settled hooks, per-backend post-call batches, deadline flushing and shutdown draining
litellm/litellm_core_utils/litellm_logging.py Flush post-call batches after success and failure callbacks
litellm/proxy/hooks/parallel_request_limiter_v3.py Batch token updates, slot releases and failure refunds
litellm/proxy/hooks/proxy_track_cost_callback.py Arm the update_cache read and use deferred spend updates
litellm/proxy/proxy_server.py Batch spend-counter writes and drain pending post-call work during shutdown
litellm/router_strategy/lowest_tpm_rpm_v2.py Send deployment TPM updates through the post-call path
flowchart TD
    A[Provider response] --> B[Success or failure callbacks]
    B --> C[Post-call Redis batch]
    C --> D[One pipeline per Redis backend]
    D --> E[Spend, limits, routing and cache writes]
Loading

User Flow

Before: a developer whose key has a budget, a team budget, an end-user budget and TPM/RPM limits waits on 11 Redis round trips per request, 6 of them after the provider answered

  1. They send POST https://litellm-domain/v1/chat/completions with "model": "gpt-group" and "user": "perf-enduser"
  2. The response comes back 200 with the mock reply
  3. Behind it the proxy writes the response cache, re-reads the spend counters, increments them, increments the deployment TPM, re-reads the user, end-user and tag spend, then runs the token Lua: six waits on Redis, one after the other

After: the same request does 2 Redis round trips after the provider answered, gets the same 200 and records the same spend

  1. They send the same POST https://litellm-domain/v1/chat/completions
  2. The response comes back 200 with the same mock reply
  3. Behind it the proxy reads the spend counters and the user, end-user and tag spend in one pipeline, then sends the cache write, every counter increment and the token Lua in one more

Relevant issues

Stacked on #43407; merge after it. Last of the stacked PRs for one Redis pipeline pre-call and one post-call (#43320, then LIT-8881, LIT-8882, LIT-8883); design and per-request measurements are on LIT-8881

Linear ticket

Resolves LIT-8883

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/unit/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Screenshots / Proof of Fix

Same fixture bytes for both arms, PYTHONPATH=<checkout>, real Redis 6.0.16 at 127.0.0.1:6379, real Postgres, mock deployments (litellm_params.mock_response; the change is on the accounting path after the provider call, not a provider path). Every RedisCache method call and every RedisBatch flush is logged with its caller chain by a sitecustomize tracer loaded through PYTHONPATH; the tables list every call between the request marker and the response marker, minus the jobs that fire on timers (_sync_in_memory_spend_with_redis, daily tag spend flush, config prefetch). Calls are round trips: a pipeline or a Lua script counts once; a PIPELINE row names the ops it carried

Fixture: config_full.yaml with two gpt-group deployments (gpt-dep-a, gpt-dep-b, usage-based-routing-v2 routing group) and two shuffle-group deployments (shuffle-dep-a, shuffle-dep-b, top-level simple-shuffle), Redis response cache on, enable_redis_auth_cache: true. Virtual key with max_budget: 1000, tpm_limit: 10000000, rpm_limit: 100000, in team perf-team (same budget and limits), user with max_budget: 1000, end user perf-enduser on a budget of 1000. Requests are sent 12 s apart so the warm pass hits warm caches; the tables are the warm pass, plus one response-cache hit

python litellm/proxy/proxy_cli.py --config ~/perf_rt/config_full.yaml --port <port>
curl -s -H "Authorization: Bearer $KEY" -H "Content-Type: application/json" http://127.0.0.1:<port>/v1/chat/completions -d @chat_warm.req.json
# chat_warm.req.json
{"model":"gpt-group","user":"perf-enduser","messages":[{"role":"user","content":"Reply with the single word: pong B"}],"max_tokens":5}
# shuffle_warm.req.json
{"model":"shuffle-group","user":"perf-enduser","messages":[{"role":"user","content":"Reply with the single word: pong B"}],"max_tokens":5}
# chat_stream_warm.req.json
{"model":"gpt-group","user":"perf-enduser","stream":true,"stream_options":{"include_usage":true},"messages":[{"role":"user","content":"Reply with the single word: pong B"}],"max_tokens":5}
# responses_warm.req.json
{"model":"gpt-group","user":"perf-enduser","input":"Reply with the single word: pong B","max_output_tokens":5}
# messages_warm.req.json
{"model":"gpt-group","max_tokens":5,"messages":[{"role":"user","content":"Reply with the single word: pong B"}],"metadata":{"user_id":"perf-enduser"}}
# chat_cachehit: chat_warm.req.json sent again, served from the response cache

Before (3a92125, the #43407 tip this PR is stacked on)

Post-call rows come from five owners on the same Redis: the response-cache SET, the spend reconcile read and increment pipeline, the deployment TPM INCRBYFLOAT, the update_cache read and the rate-limit token EVALSHA

POST /v1/chat/completions, non-streaming, usage-based routing

  1. HTTP 200, "model": "gpt-group", mock reply
  2. 11 Redis round trips on the request path (5 pre-call, 6 post-call)
# phase RedisCache op keys caller
1 pre PIPELINE MGET+MGET ['<team>_<user>', 'team_membership:837848e9-670e-4 _read_redis_rows <- _fill_from_redis
2 pre async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> _reserve_counters <- _reserve_reservable_counters
3 pre PIPELINE SET+SET+MGET+EVALSHA ['<team>_<user>', 'team_membership:837848e9-670e-4 atch_rate_limiter_script <- should_rate_limit
4 pre PIPELINE EVALSHA+EVALSHA ['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318 heck_and_increment_by_n <- reserve_tpm_tokens
5 pre async_get_cache <key-hash> ache <- _retrieve_from_cache
6 post async_set_cache <key-hash> async_set_cache <- async_add_cache
7 post PIPELINE MGET ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _collect_inflight <- _load
8 post async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> ed <- increment_spend_counters
9 post async_increment gpt-dep-a:openai/gpt-4o-mini:tpm:<min> async_increment_cache <- async_log_success_event
10 post async_batch_get_cache ['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0'] async_batch_get_cache <- _read_update_cache_values
11 post EVALSHA rate-limiter check ['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3 _execute_token_increment_script <- async_increment_tokens_with_ttl_preservation

POST /v1/chat/completions, non-streaming, simple-shuffle

  1. HTTP 200, "model": "shuffle-group", mock reply
  2. 11 Redis round trips on the request path (5 pre-call, 6 post-call)
# phase RedisCache op keys caller
1 pre PIPELINE MGET+MGET ['<team>_<user>', 'team_membership:837848e9-670e-4 _read_redis_rows <- _fill_from_redis
2 pre async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> _reserve_counters <- _reserve_reservable_counters
3 pre PIPELINE SET+SET+MGET+EVALSHA ['<team>_<user>', 'team_membership:837848e9-670e-4 atch_rate_limiter_script <- should_rate_limit
4 pre PIPELINE EVALSHA+EVALSHA ['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318 heck_and_increment_by_n <- reserve_tpm_tokens
5 pre async_get_cache <key-hash> ache <- _retrieve_from_cache
6 post async_set_cache <key-hash> async_set_cache <- async_add_cache
7 post PIPELINE MGET ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _collect_inflight <- _load
8 post async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> ed <- increment_spend_counters
9 post async_increment shuffle-dep-a:openai/gpt-4o-mini:tpm:<min> async_increment_cache <- async_log_success_event
10 post async_batch_get_cache ['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0'] async_batch_get_cache <- _read_update_cache_values
11 post EVALSHA rate-limiter check ['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3 _execute_token_increment_script <- async_increment_tokens_with_ttl_preservation

POST /v1/chat/completions, streaming

  1. HTTP 200, "model": "gpt-group", mock reply
  2. 12 Redis round trips on the request path (5 pre-call, 7 post-call)
# phase RedisCache op keys caller
1 pre PIPELINE MGET+MGET ['<team>_<user>', 'team_membership:837848e9-670e-4 _read_redis_rows <- _fill_from_redis
2 pre async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> _reserve_counters <- _reserve_reservable_counters
3 pre PIPELINE SET+SET+MGET+EVALSHA ['<team>_<user>', 'team_membership:837848e9-670e-4 atch_rate_limiter_script <- should_rate_limit
4 pre PIPELINE EVALSHA+EVALSHA ['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318 heck_and_increment_by_n <- reserve_tpm_tokens
5 pre async_get_cache <key-hash> ache <- _retrieve_from_cache
6 post async_set_cache <key-hash> async_set_cache <- async_add_cache
7 post async_set_cache <key-hash> async_set_cache <- async_add_cache
8 post PIPELINE MGET ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _collect_inflight <- _load
9 post async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> s <- _update_database_and_spend_counters_in_batch
10 post async_increment gpt-dep-a:None:tpm:<min> async_increment_cache <- async_log_success_event
11 post async_batch_get_cache ['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0'] async_batch_get_cache <- _read_update_cache_values
12 post EVALSHA rate-limiter check ['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3 _execute_token_increment_script <- async_increment_tokens_with_ttl_preservation

POST /v1/responses

  1. HTTP 200, "model": "gpt-group", mock reply
  2. 14 Redis round trips on the request path (6 pre-call, 8 post-call)
# phase RedisCache op keys caller
1 pre PIPELINE MGET+MGET ['<team>_<user>', 'team_membership:837848e9-670e-4 _read_redis_rows <- _fill_from_redis
2 pre async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> _reserve_counters <- _reserve_reservable_counters
3 pre PIPELINE SET+SET+MGET+EVALSHA ['<team>_<user>', 'team_membership:837848e9-670e-4 _batch_rate_limiter_script <- should_rate_limit
4 pre PIPELINE EVALSHA+EVALSHA ['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318 _check_and_increment_by_n <- reserve_tpm_tokens
5 pre async_get_cache <key-hash> _ageneric_api_call_with_fallbacks_responses_attempt <- make_call
6 pre get_cache SYNC <key-hash> get_cache <- _sync_get_cache
7 post set_cache SYNC <key-hash> add_cache <- sync_set_cache
8 post increment_cache SYNC gpt-dep-b:openai/gpt-4o-mini:tpm:<min> increment_cache <- log_success_event
9 post async_set_cache <key-hash> async_set_cache <- async_add_cache
10 post PIPELINE MGET ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _collect_inflight <- _load
11 post async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> ed <- increment_spend_counters
12 post async_increment gpt-dep-b:openai/gpt-4o-mini:tpm:<min> async_increment_cache <- async_log_success_event
13 post async_batch_get_cache ['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0'] async_batch_get_cache <- _read_update_cache_values
14 post EVALSHA rate-limiter check ['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3 _execute_token_increment_script <- async_increment_tokens_with_ttl_preservation

POST /v1/messages

  1. HTTP 200, "model": "gpt-group", mock reply
  2. 27 Redis round trips on the request path (21 pre-call, 6 post-call)
  3. The 21 pre-call trips here are not a /v1/messages property: this sample landed on the 60s management-object cache expiry (DEFAULT_MANAGEMENT_OBJECT_IN_MEMORY_CACHE_TTL). 16 of the 21 are the auth refresh path in user_api_key_auth (end user 5, key 3, team 5, model-access-group registry 3: miss, DB read, re-set, run one at a time before the request batch is armed); the remaining 5 are the same batched pre-call as chat. The other /v1/messages sample in the same run (messages_cold) shows 4 pre-call trips, identical to chat, and the same refresh landed on chat_cold and responses_cold. See the refresh-path breakdown below.
# phase RedisCache op keys caller
1 pre async_get_cache end_user_id:perf-enduser async_get_cache <- async_get_cache
2 pre async_get_cache end_user_restricted_registry async_get_cache <- async_get_cache
3 pre async_get_cache end_user_restricted_registry async_get_cache <- async_get_cache
4 pre async_set_cache end_user_restricted_registry _cache_registry_answer <- _fetch_and_cache_registry
5 pre async_set_cache end_user_id:perf-enduser async_set_cache <- async_set_cache
6 pre async_get_cache <key-hash> async_get_cache <- async_get_cache
7 pre async_get_cache <key-hash> async_get_cache <- async_get_cache
8 pre async_set_cache <key-hash> async_set_cache <- async_set_cache
9 pre async_get_cache team_id:<team> get_team_object <- _get_hierarchical_router_settings
10 pre async_delete_cache team_id:<team> get_team_object <- _get_hierarchical_router_settings
11 pre async_set_cache team_id:<team> get_team_object <- _get_hierarchical_router_settings
12 pre delete_cache SYNC team_alias:perf-team get_team_object <- _get_hierarchical_router_settings
13 pre async_delete_cache team_alias:perf-team get_team_object <- _get_hierarchical_router_settings
14 pre PIPELINE MGET+MGET ['<team>_<user>', '837848e9-670e-4b34-936e-156409a _read_redis_rows <- _fill_from_redis
15 pre async_get_cache model_access_group_registry cached_registry <- _load_bounded_registry
16 pre async_get_cache model_access_group_registry cached_registry <- _load_bounded_registry
17 pre async_set_cache model_access_group_registry _cache_registry <- _load_bounded_registry
18 pre async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> _reserve_counters <- _reserve_reservable_counters
19 pre PIPELINE SET+SET+SET+MGET+EVALSHA ['<user>', '<team>_837848e9-670e-4b34-936e-156409a h_rate_limiter_script <- should_rate_limit
20 pre PIPELINE EVALSHA+EVALSHA ['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318 k_and_increment_by_n <- reserve_tpm_tokens
21 pre async_get_cache <key-hash> ropic_messages_attempt <- make_call
22 post async_set_cache <key-hash> async_add_cache <- _complete_cache_write_despite_cancellation
23 post PIPELINE MGET ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _collect_inflight <- _load
24 post async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> ed <- increment_spend_counters
25 post async_increment gpt-dep-a:None:tpm:<min> async_increment_cache <- async_log_success_event
26 post async_batch_get_cache ['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0'] async_batch_get_cache <- _read_update_cache_values
27 post EVALSHA rate-limiter check ['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3 _execute_token_increment_script <- async_increment_tokens_with_ttl_preservation

POST /v1/chat/completions, response-cache hit

  1. HTTP 200, "model": "gpt-group", mock reply
  2. 11 Redis round trips on the request path (6 pre-call, 5 post-call)
# phase RedisCache op keys caller
1 pre PIPELINE MGET+MGET ['<team>_<user>', 'team_membership:837848e9-670e-4 _read_redis_rows <- _fill_from_redis
2 pre async_set_max spend:team:<team> _repair_stale_spend_counter <- get_current_spend
3 pre async_set_max spend:end_user:perf-enduser _repair_stale_spend_counter <- get_current_spend
4 pre async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> _reserve_counters <- _reserve_reservable_counters
5 pre PIPELINE SET+SET+MGET+EVALSHA ['<team>_<user>', 'team_membership:837848e9-670e-4 atch_rate_limiter_script <- should_rate_limit
6 pre PIPELINE EVALSHA+EVALSHA ['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318 heck_and_increment_by_n <- reserve_tpm_tokens
7 post PIPELINE MGET ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _collect_inflight <- _load
8 post async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> ed <- increment_spend_counters
9 post async_increment gpt-dep-a:gpt-4o-mini:tpm:<min> async_increment_cache <- async_log_success_event
10 post async_batch_get_cache ['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0'] async_batch_get_cache <- _read_update_cache_values
11 post EVALSHA rate-limiter check ['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3 _execute_token_increment_script <- async_increment_tokens_with_ttl_preservation

After (68cb506d143d68d25fd0fa8aa27ac8f1efde0b8e)

Post-call is two round trips on every async path: one read pipeline (the spend reconcile MGET and the update_cache MGET), then one write pipeline carrying the response-cache SET, every spend and deployment TPM increment and the token EVALSHA. The pre-call rows are unchanged from #43407; the cache-hit arm differs pre-call only because the before run repaired two stale spend counters (async_set_max) that the after run found already repaired

POST /v1/chat/completions, non-streaming, usage-based routing

  1. HTTP 200, "model": "gpt-group", mock reply
  2. 7 Redis round trips on the request path (5 pre-call, 2 post-call)
# phase RedisCache op keys caller
1 pre PIPELINE MGET+MGET ['<team>_<user>', 'team_membership:837848e9-670e-4 _read_redis_rows <- _fill_from_redis
2 pre async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> _reserve_counters <- _reserve_reservable_counters
3 pre PIPELINE SET+SET+MGET+EVALSHA ['<team>_<user>', 'team_membership:837848e9-670e-4 atch_rate_limiter_script <- should_rate_limit
4 pre PIPELINE EVALSHA+EVALSHA ['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318 _wrap_awaitable <- invoke
5 pre async_get_cache <key-hash> ache <- _retrieve_from_cache
6 post PIPELINE MGET+MGET ['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0', 'spend:end_user:perf-enduser' _collect_inflight <- _load
7 post PIPELINE SET+Increment+Increment+Increment+Increment+Increment+Increment+Increment+Increment+EVALSHA ['<key-hash>', 'spend:key:ee876d0d5cd318717fe4f5a0a848f8 invoke <- invoke

POST /v1/chat/completions, non-streaming, simple-shuffle

  1. HTTP 200, "model": "shuffle-group", mock reply
  2. 7 Redis round trips on the request path (5 pre-call, 2 post-call)
# phase RedisCache op keys caller
1 pre PIPELINE MGET+MGET ['<team>_<user>', 'team_membership:837848e9-670e-4 _read_redis_rows <- _fill_from_redis
2 pre async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> _reserve_counters <- _reserve_reservable_counters
3 pre PIPELINE SET+SET+MGET+EVALSHA ['<team>_<user>', 'team_membership:837848e9-670e-4 atch_rate_limiter_script <- should_rate_limit
4 pre PIPELINE EVALSHA+EVALSHA ['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318 _wrap_awaitable <- invoke
5 pre async_get_cache <key-hash> ache <- _retrieve_from_cache
6 post PIPELINE MGET+MGET ['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0', 'spend:end_user:perf-enduser' _collect_inflight <- _load
7 post PIPELINE SET+Increment+Increment+Increment+Increment+Increment+Increment+Increment+Increment+EVALSHA ['<key-hash>', 'spend:key:ee876d0d5cd318717fe4f5a0a848f8 invoke <- invoke

POST /v1/chat/completions, streaming

  1. HTTP 200, "model": "gpt-group", mock reply
  2. 7 Redis round trips on the request path (5 pre-call, 2 post-call)
# phase RedisCache op keys caller
1 pre PIPELINE MGET+MGET ['<team>_<user>', 'team_membership:837848e9-670e-4 _read_redis_rows <- _fill_from_redis
2 pre async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> _reserve_counters <- _reserve_reservable_counters
3 pre PIPELINE SET+SET+MGET+EVALSHA ['<team>_<user>', 'team_membership:837848e9-670e-4 atch_rate_limiter_script <- should_rate_limit
4 pre PIPELINE EVALSHA+EVALSHA ['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318 _wrap_awaitable <- invoke
5 pre async_get_cache <key-hash> ache <- _retrieve_from_cache
6 post PIPELINE MGET+MGET ['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0', 'spend:end_user:perf-enduser' _collect_inflight <- _load
7 post PIPELINE SET+SET+Increment+Increment+Increment+Increment+Increment+Increment+Increment+Increment+EVALSHA ['<key-hash>', '2deb1366c70583e925626e7b712eb99a3f50fd3d invoke <- invoke

POST /v1/responses

  1. HTTP 200, "model": "gpt-group", mock reply
  2. 10 Redis round trips on the request path (6 pre-call, 4 post-call)
# phase RedisCache op keys caller
1 pre PIPELINE MGET+MGET ['<team>_<user>', 'team_membership:837848e9-670e-4 _read_redis_rows <- _fill_from_redis
2 pre async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> _reserve_counters <- _reserve_reservable_counters
3 pre PIPELINE SET+SET+MGET+EVALSHA ['<team>_<user>', 'team_membership:837848e9-670e-4 _batch_rate_limiter_script <- should_rate_limit
4 pre PIPELINE EVALSHA+EVALSHA ['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318 _wrap_awaitable <- invoke
5 pre async_get_cache <key-hash> _ageneric_api_call_with_fallbacks_responses_attempt <- make_call
6 pre get_cache SYNC <key-hash> get_cache <- _sync_get_cache
7 post set_cache SYNC <key-hash> add_cache <- sync_set_cache
8 post increment_cache SYNC gpt-dep-a:openai/gpt-4o-mini:tpm:<min> increment_cache <- log_success_event
9 post PIPELINE MGET+MGET ['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0', 'spend:end_user:perf-enduser' _collect_inflight <- _load
10 post PIPELINE SET+Increment+Increment+Increment+Increment+Increment+Increment+Increment+Increment+EVALSHA ['<key-hash>', 'spend:key:ee876d0d5cd318717fe4f5a0a848f8 invoke <- invoke

POST /v1/messages

  1. HTTP 200, "model": "gpt-group", mock reply
  2. 23 Redis round trips on the request path (21 pre-call, 2 post-call)
  3. The 21 pre-call trips here are not a /v1/messages property: this sample landed on the 60s management-object cache expiry (DEFAULT_MANAGEMENT_OBJECT_IN_MEMORY_CACHE_TTL). 16 of the 21 are the auth refresh path in user_api_key_auth (end user 5, key 3, team 5, model-access-group registry 3: miss, DB read, re-set, run one at a time before the request batch is armed); the remaining 5 are the same batched pre-call as chat. The other /v1/messages sample in the same run (messages_cold) shows 4 pre-call trips, identical to chat, and the same refresh landed on chat_cold and responses_cold. See the refresh-path breakdown below.
# phase RedisCache op keys caller
1 pre async_get_cache end_user_id:perf-enduser async_get_cache <- async_get_cache
2 pre async_get_cache end_user_restricted_registry async_get_cache <- async_get_cache
3 pre async_get_cache end_user_restricted_registry async_get_cache <- async_get_cache
4 pre async_set_cache end_user_restricted_registry _cache_registry_answer <- _fetch_and_cache_registry
5 pre async_set_cache end_user_id:perf-enduser async_set_cache <- async_set_cache
6 pre async_get_cache <key-hash> async_get_cache <- async_get_cache
7 pre async_get_cache <key-hash> async_get_cache <- async_get_cache
8 pre async_set_cache <key-hash> async_set_cache <- async_set_cache
9 pre async_get_cache team_id:<team> get_team_object <- _get_hierarchical_router_settings
10 pre async_delete_cache team_id:<team> get_team_object <- _get_hierarchical_router_settings
11 pre async_set_cache team_id:<team> get_team_object <- _get_hierarchical_router_settings
12 pre delete_cache SYNC team_alias:perf-team get_team_object <- _get_hierarchical_router_settings
13 pre async_delete_cache team_alias:perf-team get_team_object <- _get_hierarchical_router_settings
14 pre PIPELINE MGET+MGET ['<team>_<user>', '837848e9-670e-4b34-936e-156409a _read_redis_rows <- _fill_from_redis
15 pre async_get_cache model_access_group_registry cached_registry <- _load_bounded_registry
16 pre async_get_cache model_access_group_registry cached_registry <- _load_bounded_registry
17 pre async_set_cache model_access_group_registry _cache_registry <- _load_bounded_registry
18 pre async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> _reserve_counters <- _reserve_reservable_counters
19 pre PIPELINE SET+SET+SET+MGET+EVALSHA ['<user>', '<team>_837848e9-670e-4b34-936e-156409a h_rate_limiter_script <- should_rate_limit
20 pre PIPELINE EVALSHA+EVALSHA ['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318 _wrap_awaitable <- invoke
21 pre async_get_cache <key-hash> ropic_messages_attempt <- make_call
22 post PIPELINE MGET+MGET ['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0', 'spend:end_user:perf-enduser' _collect_inflight <- _load
23 post PIPELINE SET+Increment+Increment+Increment+Increment+EVALSHA ['<key-hash>', 'spend:key:ee876d0d5cd318717fe4f5a0a848f8 invoke <- invoke

POST /v1/chat/completions, response-cache hit

  1. HTTP 200, "model": "gpt-group", mock reply
  2. 6 Redis round trips on the request path (4 pre-call, 2 post-call)
# phase RedisCache op keys caller
1 pre PIPELINE MGET+MGET ['<team>_<user>', 'team_membership:837848e9-670e-4 _read_redis_rows <- _fill_from_redis
2 pre async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> _reserve_counters <- _reserve_reservable_counters
3 pre PIPELINE SET+SET+MGET+EVALSHA ['<team>_<user>', 'team_membership:837848e9-670e-4 atch_rate_limiter_script <- should_rate_limit
4 pre PIPELINE EVALSHA+EVALSHA ['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318 _wrap_awaitable <- invoke
5 post PIPELINE MGET+MGET ['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0', 'spend:end_user:perf-enduser' _collect_inflight <- _load
6 post PIPELINE Increment+Increment+Increment+Increment+EVALSHA ['spend:key:<key-hash>', 'spend:team:6690cc1c-3fd6-418a- invoke <- invoke

Follow-up at c18e51e

The update_cache read is now declared only after the spend write to the database succeeds, so a cached spend another callback writes during that write is read instead of an older copy. Rerunning the live harness at c18e51e against the #43407 tip dc0344f with the same fixtures:

$ bash run_arm2.sh head43407fix ~/repos/litellm   # base
warm 14 reads, chat_master 8, chat_vk 15, messages_vk 16   (redis total_reads_processed deltas, all 200)
$ bash run_arm2.sh head43424fix ~/wt_p3           # head
warm 9 reads, chat_master 6, chat_vk 12, messages_vk 12    (all 200)

Mutation check for the ordering: moving arm_update_cache_read back in front of update_database turns test_the_update_cache_read_sees_a_cached_spend_written_while_the_spend_was_persisted red with the stale spend 1.0 instead of 5.0, and restoring it turns the test green

The follow-up test commits at 799fddb hand the limiter its scripts through the fake cache's async_register_script and drive the post-call deadline off the event loop clock instead of a wall-clock sleep. Making _flush_on_deadline skip its flush turns test_a_post_call_batch_nobody_closes_goes_out_at_the_deadline red, and restoring it turns the test green

The cancellation follow-ups at 868b0c2 and a65ce6d stop a cancelled post-call flush from deleting the shared spend counter, since a cancel does not prove the INCRBYFLOAT missed Redis. The cancelled spend is added to a local counter that already exists and an absent one is left unseeded, so a later Redis outage cannot undercount it. Removing the local increment, or the existence guard, turns test_a_cancelled_post_call_flush_keeps_the_shared_spend_counter_and_counts_the_spend_locally red. a2f12ee reads the existing post-call batch into a Final instead of rebinding it

Re-run at the current head a2f12ee against real Redis, Postgres and Anthropic claude-sonnet-4-6, all requests 200:

$ bash run_arm2.sh head43407final ~/repos/litellm   # base, 6c6354a853
warm 15 reads, chat_master 9, chat_vk 14, messages_vk 16
$ bash run_arm2.sh head43424final ~/wt_p3           # head, a2f12ee143
warm 10 reads, chat_master 6, chat_vk 12, messages_vk 12

Admin UI at a2f12ee

The same virtual key request from the fixture, POST http://localhost:4000/v1/chat/completions with sk-perf-vk-0000000001 against Anthropic claude-sonnet-4-6, shows up on http://localhost:4000/ui/?page=logs as a Success row with its real spend, so the batched Redis path still records spend end to end

Request Logs page after a live request at a2f12ee143

Live re-check at eb4379a

This head merges the #43407 head, which keys the request and post-call batches by namespace as well as connection settings. A read-count harness, run from this checkout with PYTHONPATH=<checkout>, one warm request then one request per case, redis_reads_processed from INFO stats minus the three marker commands

$ bash run_arm2.sh <arm> <checkout>  # redis-cli flushall, proxy on :4000 with config_v2.yaml, real Redis 7, real Postgres, Anthropic claude-sonnet-4-6
key/generate 200
warm 200
warm redis_reads_processed=10
chat_master 200
chat_master redis_reads_processed=6
chat_vk 200
chat_vk redis_reads_processed=12
messages_vk 200
messages_vk redis_reads_processed=12

Every response body was pong from claude-group with usage populated, and the proxy log has no ERROR or Traceback lines

Clean re-run, no cache-expiry samples

The tables above were captured with samples 12 to 16 s apart, so roughly every fourth sample landed on the 60 s management-object TTL and paid a 16-call auth refresh (end user 5, key 3, team 5, model-access registry 3) that has nothing to do with the endpoint sampled; that is where the 21 pre-call trips on the /v1/messages sample came from. The refresh path itself is fixed in #43776 (P4). This re-run, done in a separate session, sends a warm-up request right before every traced request and pauses before every fourth warm-up so the expiry happens outside the sample windows; both arms were checked for refresh signatures inside every window (0 found). Before is the #43407 tip fada595, After is 9dad494, the same tree as the current tip 073f036 minus the #43407 log-line fix it now carries, same harness, same fixtures, real Redis 6.0.16 and Postgres 14, every warm-up and sample HTTP 200, chat_cachehit served the same response id as chat_warm in both arms. The before arm needed one rerun (its first attempt had refresh signatures in messages_cold and shuffle_warm), the after arm none

Request Before pre Before post Before total After pre After post After total
chat_cold 5 7 12 5 3 8
chat_stream_cold 5 8 13 5 3 8
shuffle_cold 5 7 12 5 3 8
messages_cold 5 7 12 5 3 8
responses_cold 6 9 15 6 5 11
chat_warm 5 7 12 5 3 8
chat_stream_warm 5 8 13 5 3 8
shuffle_warm 5 7 12 5 3 8
messages_warm 5 7 12 5 3 8
responses_warm 6 9 15 6 5 11
chat_cachehit 4 6 10 4 3 7

Post-call on /v1/chat/completions (both routing groups, streaming or not) and /v1/messages goes from 7 or 8 trips to 3: the spend counter read pipeline, the update_cache read pipeline and the one write pipeline (response cache SET, spend increments, deployment TPM, token Lua). /v1/responses stays at 5 because its native handler's two synchronous worker-thread calls are outside the request scope. Pre-call is unchanged by this PR, as intended

Type

🚄 Infrastructure

Caveats (if any)

Medium

  • Post-call is 2 round trips, not 1: the spend settle still reads the counters first (reservation delta, and only counters that exist are incremented, so a flushed counter is not resurrected), and the write pipeline follows. Folding that read into an existence-guarded Lua increment is the P4 item on LIT-8884
  • The post-call writes leave when the last success or failure callback has run, or 1 s after the first declaration. Another worker reading a spend counter, a token counter or the response cache inside that window sees the pre-request value, as it did during the callback's own gap before this PR. In-memory copies (spend counters, DualCache response cache, parallel slots) are updated at once
  • A failed increment inside the pipeline invalidates that spend counter locally, so the next request rebuilds it from Redis, as a failed INCRBYFLOAT did before. A failed token Lua group falls back to the plain increment pipeline for that group. A parallel slot is removed from the local gauge the moment it is released, so admission on the same worker sees the capacity before the pipeline goes out; the count Redis returns from the pipelined release is not written over the gauge (a newer acquire on the worker may have mirrored a fresher count by then; the next acquire refreshes it from Redis as before), and a failed release Lua leaves the slot released in memory only
  • /v1/responses native handlers still do two synchronous calls (set_cache, increment_cache) from a worker thread; they run outside the async request scope and are not batched, exactly as before this PR

Low

  • The deadline flush is a task the scope does not await; it holds the batch objects alive until it has sent them. Nothing is declared after it except by a callback still running, which flushes at its own end. Every request whose post-call batch still holds ops is registered in a WeakSet, and proxy_shutdown_event drains them (drain_post_call_redis_batches) before the cache disconnects, so a worker stopping inside the callback or deadline window sends the pending writes rather than dropping them
  • The deferred response-cache SET coerces the caller's ttl exactly as BaseCache.get_ttl does (int(ttl), a non-numeric value means the cache default), so ttl="3600" expires in 3600 s on both paths
  • Response-cache writes with nx=True (and any non-Redis cache) stay direct; the deferred SET uses the same key, TTL and serialised bytes RedisCache.async_set_cache would send
  • Outside a request scope (background jobs, SDK use) every owner runs exactly as before

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Link to Devin session: https://app.devin.ai/sessions/f732394a6aca4bc197a2f4b1d54d0e7e
Open in Devin Desktop: https://app.devin.ai/desktop/session/f732394a6aca4bc197a2f4b1d54d0e7e?variant=devin
Requested by: @yassin-berriai

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@greptile-apps

greptile-apps Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[High risk] Refactors Redis write batching across spend, rate-limit, and cache operations.

The PR appears safe to merge based on the changes since the previous review.

Summary

The PR batches post-call Redis writes for spend accounting, rate limiting, routing usage, and response caching. Since the previous review, it also keys shared Redis batches by namespace and adds a test for distinct namespaces.

Reviews (18) · Last reviewed commit: "Merge branch 'litellm_redis_p2_request_b..."

Comment thread litellm/proxy/hooks/parallel_request_limiter_v3.py Outdated
Comment thread litellm/caching/redis_batch.py
Comment thread litellm/caching/caching.py Outdated
@codecov

codecov Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 94.16667% with 14 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
litellm/caching/caching.py 85.71% 5 Missing ⚠️
litellm/caching/dual_cache.py 88.37% 5 Missing ⚠️
litellm/caching/redis_batch.py 95.23% 3 Missing ⚠️
litellm/proxy/hooks/parallel_request_limiter_v3.py 97.50% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread litellm/proxy/hooks/parallel_request_limiter_v3.py
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread litellm/proxy/hooks/parallel_request_limiter_v3.py Outdated
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread tests/unit/caching/test_request_redis_batch_post_call.py Outdated
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread litellm/proxy/hooks/proxy_track_cost_callback.py Outdated
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread tests/unit/caching/test_request_redis_batch_post_call.py Outdated
Comment thread tests/unit/caching/test_request_redis_batch_post_call.py Outdated
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@yassin-berriai

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/proxy/proxy_server.py
@yassin-berriai

Copy link
Copy Markdown
Contributor

@greptileai

@yassin-berriai

Copy link
Copy Markdown
Contributor

bugbot run

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

bugbot run

1 similar comment
@yassin-berriai

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/proxy/hooks/parallel_request_limiter_v3.py
@yassin-berriai

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@yassin-berriai

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit eb4379a. Configure here.

…-limit, routing and response-cache writes

Post-call owners declare into one request-scoped RedisBatch per Redis backend: spend counter
increments and reservation reconciliation, rate-limit token Lua updates and refunds, parallel-slot
release (freed locally at once), deployment TPM, and compatible async response-cache SETs. The batch
is sent once the success and failure callbacks have run, or on a deadline, and pending batches are
drained at shutdown before Redis disconnects. nx writes, non-Redis caches and calls outside a request
stay direct; numeric string TTLs keep the direct-path coercion.

Resolves LIT-8883

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@yassin-berriai
yassin-berriai merged commit 826849f into litellm_redis_p2_request_batch Sep 29, 2026
94 of 97 checks passed
@yassin-berriai
yassin-berriai deleted the litellm_redis_p3_post_call_batch branch September 29, 2026 23:42

This branch is waiting to be deployed

1 waiting deployment
e2e-changed — 5b7b2e4e Waiting Sep 29, 2026 by devin-ai-integration[bot] via oauth #1891
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants