Repository navigation
perf(proxy): one request-scoped Redis pipeline for auth, spend, rate-limit and routing reads - #43407
Conversation
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
|
|
Codecov Report❌ Patch coverage is 📢 Thoughts on this report? Let us know! |
|
1 similar comment
|
bugbot run |
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit ec4dd90. Configure here.
178bd46 to
a51b69e
Compare
ec4dd90 to
13e171c
Compare
13e171c to
fada595
Compare
1476a1c to
da1e602
Compare
…limit and routing reads RedisBatch: one pipeline per Redis backend for independently declared operations (MGET, GET, Lua scripts, INCRBYFLOAT, SET, DEL), a future per operation so each owner keeps its own fallback, Redis Cluster hash-slot fallback. A request-scoped batch middleware shares that pipeline across the auth identity reads and write-back, the spend counter MGET, the rate limiter Lua groups and the routing read. A rate-limit denial stands when another pipelined group fails; every pipelined group is refunded on rejection; local cooldowns win over the prefetch. The routing prefetch failure log line strips request line breaks (CodeQL py/log-injection) Resolves LIT-8882 Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
da1e602 to
2de50b6
Compare
TLDR
Each request can share a Redis batch across authentication, spend checks, rate limits, and routing reads. Awaiting a result sends queued operations through a pipeline for that Redis backend, and the middleware flushes anything left when the request ends. This PR extends the earlier spend-counter batching work to the rest of the request path
The debug line logged when a routing prefetch fails to arm strips CR and LF from the request's model name and the error text, so a request cannot forge a log entry through it (CodeQL py/log-injection)
Files changed
litellm/caching/redis_batch.pylitellm/proxy/auth/auth_object_prefetch.pylitellm/proxy/common_request_processing.pylitellm/proxy/hooks/parallel_request_limiter_v3.pylitellm/proxy/middleware/redis_request_batch_middleware.pylitellm/proxy/proxy_server.pylitellm/proxy/spend_tracking/spend_counter_batch.pylitellm/router.pylitellm/router_utils/routing_read_batch.pyRequest flow
flowchart LR R[Request opens batch] --> Q[Auth spend limits and routing queue Redis work] Q --> F[Await or request end flushes pending work] F --> P[Pipeline per Redis backend sends queued operations] P --> C[Callers receive results]User Flow
Before: a developer whose key has a budget, a team budget, an end-user budget and TPM/RPM limits waits on 15 Redis round trips per request, 9 of them before the provider is called
"model": "gpt-group"and"user": "perf-enduser"After: the same request waits on 5 Redis round trips before the provider is called, gets the same 200 and records the same spend
Relevant issues
Stacked on #43369; merge after it. Third of the stacked PRs for one Redis pipeline pre-call and one post-call (#43320, then LIT-8881, LIT-8882, LIT-8883); design and per-request measurements are on LIT-8881
Linear ticket
Resolves LIT-8882
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/unit/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Screenshots / Proof of Fix
Same fixture bytes for both arms,
PYTHONPATH=<checkout>, real Redis 6.0.16 at 127.0.0.1:6379, real Postgres, mock deployments (litellm_params.mock_response; the change is on the admission path, not a provider path). EveryRedisCachemethod call and everyRedisBatchflush is logged with its caller chain by asitecustomizetracer loaded throughPYTHONPATH; the tables list every call between the request marker and the response marker, minus the jobs that fire on timers (_sync_in_memory_spend_with_redis, daily tag spend flush, config prefetch). Calls are round trips: a pipeline or a Lua script counts once; aPIPELINErow names the ops it carriedFixture:
config_full.yamlwith twogpt-groupdeployments (gpt-dep-a,gpt-dep-b,usage-based-routing-v2routing group) and twoshuffle-groupdeployments (shuffle-dep-a,shuffle-dep-b, top-levelsimple-shuffle), Redis response cache on,enable_redis_auth_cache: true. Virtual key withmax_budget: 1000,tpm_limit: 10000000,rpm_limit: 100000, in teamperf-team(same budget and limits), user withmax_budget: 1000, end userperf-enduseron a budget of 1000. Requests are sent 12 s apart so the warm pass hits warm caches; the tables are the warm passBefore (e8affcb, the #43369 tip this PR is stacked on)
POST /v1/chat/completions, non-streaming, usage-based routing
HTTP 200,"model": "gpt-group", contentmock reply from basync_batch_get_cache['<team>_<user>', 'team_membership:837848e9-670e-4_fill_from_redis <- prefetch_auth_objectsasync_set_cache_pipeline_with_ttls(('<team>_<user>', {'user_id': '837848e9-670e-4b34_write_back <- _fill_from_dbasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_fetch <- _loadasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>_reserve_counters <- _reserve_reservable_countersEVALSHA rate-limiter check['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318atch_rate_limiter_script <- should_rate_limitEVALSHA tpm check-and-increment['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318heck_and_increment_by_n <- reserve_tpm_tokensEVALSHA tpm check-and-increment['{team:<team>}:window', '{team:<team>}:tokens']heck_and_increment_by_n <- reserve_tpm_tokensasync_batch_get_cache['deployment:gpt-dep-a:cooldown', 'deployment:gpt-dep-b:cooldown', 'deployment:shuffle-dep-a:cooldown', 'deplo_cooldown_deployments <- async_get_healthy_deploymentsasync_get_cache<key-hash>ache <- _retrieve_from_cacheasync_set_cache<key-hash>async_set_cache <- async_add_cacheasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_fetch <- _loadasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>ed <- increment_spend_countersasync_incrementgpt-dep-b:openai/gpt-4o-mini:tpm:<min>async_increment_cache <- async_log_success_eventasync_batch_get_cache['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0']async_batch_get_cache <- _read_update_cache_valuesEVALSHA rate-limiter check['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3_execute_token_increment_script <- async_increment_tokens_with_ttl_preservationPOST /v1/chat/completions, non-streaming, simple-shuffle
HTTP 200,"model": "shuffle-group", contentmock reply from sasync_batch_get_cache['<team>_<user>', 'team_membership:837848e9-670e-4_fill_from_redis <- prefetch_auth_objectsasync_set_cache_pipeline_with_ttls(('<team>_<user>', {'user_id': '837848e9-670e-4b34_write_back <- _fill_from_dbasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_fetch <- _loadasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>_reserve_counters <- _reserve_reservable_countersEVALSHA rate-limiter check['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318atch_rate_limiter_script <- should_rate_limitEVALSHA tpm check-and-increment['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318heck_and_increment_by_n <- reserve_tpm_tokensEVALSHA tpm check-and-increment['{team:<team>}:window', '{team:<team>}:tokens']heck_and_increment_by_n <- reserve_tpm_tokensasync_batch_get_cache['deployment:gpt-dep-a:cooldown', 'deployment:gpt-dep-b:cooldown', 'deployment:shuffle-dep-a:cooldown', 'deplo_cooldown_deployments <- async_get_healthy_deploymentsasync_get_cache<key-hash>ache <- _retrieve_from_cacheasync_set_cache<key-hash>async_set_cache <- async_add_cacheasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_fetch <- _loadasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>ed <- increment_spend_countersasync_incrementshuffle-dep-a:openai/gpt-4o-mini:tpm:<min>async_increment_cache <- async_log_success_eventasync_batch_get_cache['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0']async_batch_get_cache <- _read_update_cache_valuesEVALSHA rate-limiter check['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3_execute_token_increment_script <- async_increment_tokens_with_ttl_preservationPOST /v1/chat/completions, streaming
HTTP 200, 8 SSE chunks,"model": "gpt-group", contentmock reply from basync_batch_get_cache['<team>_<user>', 'team_membership:837848e9-670e-4_fill_from_redis <- prefetch_auth_objectsasync_set_cache_pipeline_with_ttls(('<team>_<user>', {'user_id': '837848e9-670e-4b34_write_back <- _fill_from_dbasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_fetch <- _loadasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>_reserve_counters <- _reserve_reservable_countersEVALSHA rate-limiter check['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318atch_rate_limiter_script <- should_rate_limitEVALSHA tpm check-and-increment['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318heck_and_increment_by_n <- reserve_tpm_tokensEVALSHA tpm check-and-increment['{team:<team>}:window', '{team:<team>}:tokens']heck_and_increment_by_n <- reserve_tpm_tokensasync_batch_get_cache['deployment:gpt-dep-a:cooldown', 'deployment:gpt-dep-b:cooldown', 'deployment:shuffle-dep-a:cooldown', 'deplo_cooldown_deployments <- async_get_healthy_deploymentsasync_get_cache<key-hash>ache <- _retrieve_from_cacheasync_set_cache<key-hash>async_set_cache <- async_add_cacheasync_set_cache<key-hash>async_set_cache <- async_add_cacheasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_fetch <- _loadasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>s <- _update_database_and_spend_counters_in_batchasync_incrementgpt-dep-b:None:tpm:<min>async_increment_cache <- async_log_success_eventasync_batch_get_cache['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0']async_batch_get_cache <- _read_update_cache_valuesEVALSHA rate-limiter check['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3_execute_token_increment_script <- async_increment_tokens_with_ttl_preservationPOST /v1/responses
HTTP 200,"model": "gpt-group", contentmock reply from bSYNCrows are the pre-existing synchronous cache calls of the Responses path (noted on perf(router): fetch cooldown state and usage counters in one Redis round trip #43320)async_batch_get_cache['<team>_<user>', 'team_membership:837848e9-670e-4_fill_from_redis <- prefetch_auth_objectsasync_set_cache_pipeline_with_ttls(('<team>_<user>', {'user_id': '837848e9-670e-4b34_write_back <- _fill_from_dbasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_fetch <- _loadasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>_reserve_counters <- _reserve_reservable_countersEVALSHA rate-limiter check['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318_batch_rate_limiter_script <- should_rate_limitEVALSHA tpm check-and-increment['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318_check_and_increment_by_n <- reserve_tpm_tokensEVALSHA tpm check-and-increment['{team:<team>}:window', '{team:<team>}:tokens']_check_and_increment_by_n <- reserve_tpm_tokensasync_batch_get_cache['deployment:gpt-dep-a:cooldown', 'deployment:gpt-dep-b:cooldown', 'deployment:shuffle-dep-a:cooldown', 'deplo_api_call_with_fallbacks_responses_attempt <- make_callasync_get_cache<key-hash>_ageneric_api_call_with_fallbacks_responses_attempt <- make_callget_cacheSYNC <key-hash>get_cache <- _sync_get_cacheset_cacheSYNC <key-hash>add_cache <- sync_set_cacheincrement_cacheSYNC gpt-dep-b:openai/gpt-4o-mini:tpm:<min>increment_cache <- log_success_eventasync_set_cache<key-hash>async_set_cache <- async_add_cacheasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_fetch <- _loadasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>ed <- increment_spend_countersasync_incrementgpt-dep-b:openai/gpt-4o-mini:tpm:<min>async_increment_cache <- async_log_success_eventasync_batch_get_cache['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0']async_batch_get_cache <- _read_update_cache_valuesEVALSHA rate-limiter check['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3_execute_token_increment_script <- async_increment_tokens_with_ttl_preservationPOST /v1/messages
HTTP 200,"model": "gpt-group", contentmock reply from aasync_get_cacheend_user_id:perf-enduserasync_get_cache <- async_get_cacheasync_get_cacheend_user_restricted_registryasync_get_cache <- async_get_cacheasync_get_cacheend_user_restricted_registryasync_get_cache <- async_get_cacheasync_set_cacheend_user_restricted_registry_cache_registry_answer <- _fetch_and_cache_registryasync_set_cacheend_user_id:perf-enduserasync_set_cache <- async_set_cacheasync_get_cache<key-hash>async_get_cache <- async_get_cacheasync_get_cache<key-hash>async_get_cache <- async_get_cacheasync_set_cache<key-hash>async_set_cache <- async_set_cacheasync_get_cacheteam_id:<team>get_team_object <- _get_hierarchical_router_settingsasync_delete_cacheteam_id:<team>get_team_object <- _get_hierarchical_router_settingsasync_set_cacheteam_id:<team>get_team_object <- _get_hierarchical_router_settingsdelete_cacheSYNC team_alias:perf-teamget_team_object <- _get_hierarchical_router_settingsasync_delete_cacheteam_alias:perf-teamget_team_object <- _get_hierarchical_router_settingsasync_batch_get_cache['<team>_<user>', '837848e9-670e-4b34-936e-156409a_fill_from_redis <- prefetch_auth_objectsasync_set_cache_pipeline_with_ttls(('<user>', {'user_id': '<user>', 'user_alias': No_write_back <- _fill_from_dbasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_fetch <- _loadasync_get_cachemodel_access_group_registrycached_registry <- _load_bounded_registryasync_get_cachemodel_access_group_registrycached_registry <- _load_bounded_registryasync_set_cachemodel_access_group_registry_cache_registry <- _load_bounded_registryasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>_reserve_counters <- _reserve_reservable_countersEVALSHA rate-limiter check['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318h_rate_limiter_script <- should_rate_limitEVALSHA tpm check-and-increment['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318k_and_increment_by_n <- reserve_tpm_tokensEVALSHA tpm check-and-increment['{team:<team>}:window', '{team:<team>}:tokens']k_and_increment_by_n <- reserve_tpm_tokensasync_batch_get_cache['deployment:gpt-dep-a:cooldown', 'deployment:gpt-dep-b:cooldown', 'deployment:shuffle-dep-a:cooldown', 'deplopt <- make_callasync_get_cache<key-hash>ropic_messages_attempt <- make_callasync_set_cache<key-hash>async_add_cache <- _complete_cache_write_despite_cancellationasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_fetch <- _loadasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>ed <- increment_spend_countersasync_incrementgpt-dep-a:None:tpm:<min>async_increment_cache <- async_log_success_eventasync_batch_get_cache['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0']async_batch_get_cache <- _read_update_cache_valuesEVALSHA rate-limiter check['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3_execute_token_increment_script <- async_increment_tokens_with_ttl_preservationAfter (32c0727; the follow-up commits add suppression reasons, a router coverage exemption and the review fixes: a rejected or failed rate-limit group now refunds every group the pipeline incremented and a Redis denial stands even when another group fails, a reply one op cannot decode fails that op alone, and a cooldown recorded in memory after the prefetch left still wins; the round-trip counts are unchanged)
POST /v1/chat/completions, non-streaming, usage-based routing
HTTP 200,"model": "gpt-group", contentmock reply from a(the other deployment of the same group; both served requests in both arms)PIPELINE MGET+MGET['<team>_<user>', 'team_membership:837848e9-670e-4_read_redis_rows <- _fill_from_redisasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>_reserve_counters <- _reserve_reservable_countersPIPELINE SET+SET+MGET+EVALSHA['<team>_<user>', 'team_membership:837848e9-670e-4atch_rate_limiter_script <- should_rate_limitPIPELINE EVALSHA+EVALSHA['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318heck_and_increment_by_n <- reserve_tpm_tokensasync_get_cache<key-hash>ache <- _retrieve_from_cacheasync_set_cache<key-hash>async_set_cache <- async_add_cachePIPELINE MGET['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_collect_inflight <- _loadasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>ed <- increment_spend_countersasync_incrementgpt-dep-a:openai/gpt-4o-mini:tpm:<min>async_increment_cache <- async_log_success_eventasync_batch_get_cache['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0']async_batch_get_cache <- _read_update_cache_valuesEVALSHA rate-limiter check['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3_execute_token_increment_script <- async_increment_tokens_with_ttl_preservationPOST /v1/chat/completions, non-streaming, simple-shuffle
HTTP 200,"model": "shuffle-group", contentmock reply from sPIPELINE MGET+MGET['<team>_<user>', 'team_membership:837848e9-670e-4_read_redis_rows <- _fill_from_redisasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>_reserve_counters <- _reserve_reservable_countersPIPELINE SET+SET+MGET+EVALSHA['<team>_<user>', 'team_membership:837848e9-670e-4atch_rate_limiter_script <- should_rate_limitPIPELINE EVALSHA+EVALSHA['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318heck_and_increment_by_n <- reserve_tpm_tokensasync_get_cache<key-hash>ache <- _retrieve_from_cacheasync_set_cache<key-hash>async_set_cache <- async_add_cachePIPELINE MGET['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_collect_inflight <- _loadasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>ed <- increment_spend_countersasync_incrementshuffle-dep-a:openai/gpt-4o-mini:tpm:<min>async_increment_cache <- async_log_success_eventasync_batch_get_cache['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0']async_batch_get_cache <- _read_update_cache_valuesEVALSHA rate-limiter check['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3_execute_token_increment_script <- async_increment_tokens_with_ttl_preservationPOST /v1/chat/completions, streaming
HTTP 200, 8 SSE chunks,"model": "gpt-group", contentmock reply from aPIPELINE MGET+MGET['<team>_<user>', 'team_membership:837848e9-670e-4_read_redis_rows <- _fill_from_redisasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>_reserve_counters <- _reserve_reservable_countersPIPELINE SET+SET+MGET+EVALSHA['<team>_<user>', 'team_membership:837848e9-670e-4atch_rate_limiter_script <- should_rate_limitPIPELINE EVALSHA+EVALSHA['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318heck_and_increment_by_n <- reserve_tpm_tokensasync_get_cache<key-hash>ache <- _retrieve_from_cacheasync_set_cache<key-hash>async_set_cache <- async_add_cacheasync_set_cache<key-hash>async_set_cache <- async_add_cachePIPELINE MGET['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_collect_inflight <- _loadasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>s <- _update_database_and_spend_counters_in_batchasync_incrementgpt-dep-a:None:tpm:<min>async_increment_cache <- async_log_success_eventasync_batch_get_cache['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0']async_batch_get_cache <- _read_update_cache_valuesEVALSHA rate-limiter check['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3_execute_token_increment_script <- async_increment_tokens_with_ttl_preservationPOST /v1/responses
HTTP 200,"model": "gpt-group", contentmock reply from bPIPELINE MGET+MGET['<team>_<user>', 'team_membership:837848e9-670e-4_read_redis_rows <- _fill_from_redisasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>_reserve_counters <- _reserve_reservable_countersPIPELINE SET+SET+MGET+EVALSHA['<team>_<user>', 'team_membership:837848e9-670e-4_batch_rate_limiter_script <- should_rate_limitPIPELINE EVALSHA+EVALSHA['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318_check_and_increment_by_n <- reserve_tpm_tokensasync_get_cache<key-hash>_ageneric_api_call_with_fallbacks_responses_attempt <- make_callget_cacheSYNC <key-hash>get_cache <- _sync_get_cacheset_cacheSYNC <key-hash>add_cache <- sync_set_cacheincrement_cacheSYNC gpt-dep-b:openai/gpt-4o-mini:tpm:<min>increment_cache <- log_success_eventasync_set_cache<key-hash>async_set_cache <- async_add_cachePIPELINE MGET['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_collect_inflight <- _loadasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>ed <- increment_spend_countersasync_incrementgpt-dep-b:openai/gpt-4o-mini:tpm:<min>async_increment_cache <- async_log_success_eventasync_batch_get_cache['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0']async_batch_get_cache <- _read_update_cache_valuesEVALSHA rate-limiter check['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3_execute_token_increment_script <- async_increment_tokens_with_ttl_preservationPOST /v1/messages
HTTP 200,"model": "gpt-group", contentmock reply from aasync_get_cacheend_user_id:perf-enduserasync_get_cache <- async_get_cacheasync_get_cacheend_user_restricted_registryasync_get_cache <- async_get_cacheasync_get_cacheend_user_restricted_registryasync_get_cache <- async_get_cacheasync_set_cacheend_user_restricted_registry_cache_registry_answer <- _fetch_and_cache_registryasync_set_cacheend_user_id:perf-enduserasync_set_cache <- async_set_cacheasync_get_cache<key-hash>async_get_cache <- async_get_cacheasync_get_cache<key-hash>async_get_cache <- async_get_cacheasync_set_cache<key-hash>async_set_cache <- async_set_cacheasync_get_cacheteam_id:<team>get_team_object <- _get_hierarchical_router_settingsasync_delete_cacheteam_id:<team>get_team_object <- _get_hierarchical_router_settingsasync_set_cacheteam_id:<team>get_team_object <- _get_hierarchical_router_settingsdelete_cacheSYNC team_alias:perf-teamget_team_object <- _get_hierarchical_router_settingsasync_delete_cacheteam_alias:perf-teamget_team_object <- _get_hierarchical_router_settingsPIPELINE MGET+MGET['<team>_<user>', '837848e9-670e-4b34-936e-156409a_read_redis_rows <- _fill_from_redisasync_get_cachemodel_access_group_registrycached_registry <- _load_bounded_registryasync_get_cachemodel_access_group_registrycached_registry <- _load_bounded_registryasync_set_cachemodel_access_group_registry_cache_registry <- _load_bounded_registryasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>_reserve_counters <- _reserve_reservable_countersPIPELINE SET+SET+SET+MGET+EVALSHA['<user>', '<team>_837848e9-670e-4b34-936e-156409ah_rate_limiter_script <- should_rate_limitPIPELINE EVALSHA+EVALSHA['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318k_and_increment_by_n <- reserve_tpm_tokensasync_get_cache<key-hash>ropic_messages_attempt <- make_callasync_set_cache<key-hash>async_add_cache <- _complete_cache_write_despite_cancellationPIPELINE MGET['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_collect_inflight <- _loadasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>ed <- increment_spend_countersasync_incrementgpt-dep-a:None:tpm:<min>async_increment_cache <- async_log_success_eventasync_batch_get_cache['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0']async_batch_get_cache <- _read_update_cache_valuesEVALSHA rate-limiter check['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3_execute_token_increment_script <- async_increment_tokens_with_ttl_preservationThe five pre-call round trips left on the chat path are separated by data dependencies, not by ownership: the reservation increments (row 2) need the spend read; the rate limiter (row 3) runs after auth and its result gates the TPM scripts (row 4); the response-cache key (row 5) depends on the deployment routing picked. Getting from 5 to 1 needs the speculative admission mode and the Lua check-and-reserve of LIT-8884, which this PR does not ship
Not measured here: a fleet load regime. These are single-request command counts on one box; the prod traces on LIT-8881 are what motivated the change
The fail-closed follow-up at b763921 makes a pipelined rate-limit or TPM group that Redis could not verify refund every applied group and then raise
RateLimitUnverifiableErrorwhen fail-closed is on, instead of quietly falling back to in-memory counters. Dropping the_reject_if_rate_limit_unverifiablecall turns both cases oftest_fail_closed_rejects_when_a_pipelined_lua_group_cannot_be_verifiedred, and restoring it turns them green. 6c6354a only narrows the fake Redis types and pins the cooldown test timestampsRe-run at the current head 6c6354a against real Redis, Postgres and Anthropic claude-sonnet-4-6, all requests 200:
Admin UI at 6c6354a
The same virtual key request from the fixture, POST http://localhost:4000/v1/chat/completions with
sk-perf-vk-0000000001against Anthropic claude-sonnet-4-6, shows up on http://localhost:4000/ui/?page=logs as a Success row with its real spend, so the batched Redis path still records spend end to endLive re-check at ec4dd90
This head keys the request batch by the cache's namespace as well as its connection settings, so two caches on one server with different prefixes each get their own pipeline and key prefix (
test_caches_on_one_server_with_different_namespaces_keep_their_own_key_prefix, red when the namespace is dropped from_backend_key). A read-count harness, run from this checkout withPYTHONPATH=<checkout>, one warm request then one request per case,redis_reads_processedfromINFO statsminus the three marker commandsEvery response body was
pongfromclaude-groupwith usage populated, and the proxy log has no ERROR or Traceback linesType
🚄 Infrastructure
Caveats (if any)
Medium
RedisClusterCacheevery declared op still runs on its own, as before this PR. The cluster pipeline is per node and the existing per-op paths already group by slotLow
RoutingReadBatchwhen the request armed a cooldown prefetch; otherwise the router's plain cooldown read stays in charge, so callers and tests that replacelitellm.router._async_get_cooldown_deploymentsstill take effectredis.exceptionsis imported lazily insideredis_batch.py, likeredis_cache.py, soimport litellmwithout redis installed keeps workingFinal Attestation
Link to Devin session: https://app.devin.ai/sessions/a666ccc35fde4ab799d68f3cc5bad523
Open in Devin Desktop: https://app.devin.ai/desktop/session/a666ccc35fde4ab799d68f3cc5bad523?variant=devin
Requested by: @yassin-berriai