Repository navigation
perf(proxy): one post-call Redis pipeline per backend for spend, rate-limit, routing and response-cache writes - #43424
Conversation
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
|
|
|
Codecov Report❌ Patch coverage is 📢 Thoughts on this report? Let us know! |
|
bugbot run |
|
bugbot run |
|
bugbot run |
1 similar comment
|
bugbot run |
|
bugbot run |
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit eb4379a. Configure here.
ec4dd90 to
13e171c
Compare
eb4379a to
9ebdb43
Compare
13e171c to
fada595
Compare
502476b to
9dad494
Compare
1476a1c to
da1e602
Compare
9dad494 to
073f036
Compare
…-limit, routing and response-cache writes Post-call owners declare into one request-scoped RedisBatch per Redis backend: spend counter increments and reservation reconciliation, rate-limit token Lua updates and refunds, parallel-slot release (freed locally at once), deployment TPM, and compatible async response-cache SETs. The batch is sent once the success and failure callbacks have run, or on a deadline, and pending batches are drained at shutdown before Redis disconnects. nx writes, non-Redis caches and calls outside a request stay direct; numeric string TTLs keep the direct-path coercion. Resolves LIT-8883 Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
da1e602 to
2de50b6
Compare
073f036 to
5b7b2e4
Compare
826849f
into
litellm_redis_p2_request_batch
TLDR
Problem this solves:
update_cachedo their own workHow it solves it:
update_cachereads join the existing request read pipeline when possibleFiles changed
litellm/caching/caching.py,litellm/caching/dual_cache.pylitellm/caching/redis_batch.pylitellm/litellm_core_utils/litellm_logging.pylitellm/proxy/hooks/parallel_request_limiter_v3.pylitellm/proxy/hooks/proxy_track_cost_callback.pyupdate_cacheread and use deferred spend updateslitellm/proxy/proxy_server.pylitellm/router_strategy/lowest_tpm_rpm_v2.pyflowchart TD A[Provider response] --> B[Success or failure callbacks] B --> C[Post-call Redis batch] C --> D[One pipeline per Redis backend] D --> E[Spend, limits, routing and cache writes]User Flow
Before: a developer whose key has a budget, a team budget, an end-user budget and TPM/RPM limits waits on 11 Redis round trips per request, 6 of them after the provider answered
"model": "gpt-group"and"user": "perf-enduser"After: the same request does 2 Redis round trips after the provider answered, gets the same 200 and records the same spend
Relevant issues
Stacked on #43407; merge after it. Last of the stacked PRs for one Redis pipeline pre-call and one post-call (#43320, then LIT-8881, LIT-8882, LIT-8883); design and per-request measurements are on LIT-8881
Linear ticket
Resolves LIT-8883
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/unit/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Screenshots / Proof of Fix
Same fixture bytes for both arms,
PYTHONPATH=<checkout>, real Redis 6.0.16 at 127.0.0.1:6379, real Postgres, mock deployments (litellm_params.mock_response; the change is on the accounting path after the provider call, not a provider path). EveryRedisCachemethod call and everyRedisBatchflush is logged with its caller chain by asitecustomizetracer loaded throughPYTHONPATH; the tables list every call between the request marker and the response marker, minus the jobs that fire on timers (_sync_in_memory_spend_with_redis, daily tag spend flush, config prefetch). Calls are round trips: a pipeline or a Lua script counts once; aPIPELINErow names the ops it carriedFixture:
config_full.yamlwith twogpt-groupdeployments (gpt-dep-a,gpt-dep-b,usage-based-routing-v2routing group) and twoshuffle-groupdeployments (shuffle-dep-a,shuffle-dep-b, top-levelsimple-shuffle), Redis response cache on,enable_redis_auth_cache: true. Virtual key withmax_budget: 1000,tpm_limit: 10000000,rpm_limit: 100000, in teamperf-team(same budget and limits), user withmax_budget: 1000, end userperf-enduseron a budget of 1000. Requests are sent 12 s apart so the warm pass hits warm caches; the tables are the warm pass, plus one response-cache hitBefore (3a92125, the #43407 tip this PR is stacked on)
Post-call rows come from five owners on the same Redis: the response-cache
SET, the spend reconcile read and increment pipeline, the deployment TPMINCRBYFLOAT, theupdate_cacheread and the rate-limit tokenEVALSHAPOST /v1/chat/completions, non-streaming, usage-based routing
HTTP 200,"model": "gpt-group", mock replyPIPELINE MGET+MGET['<team>_<user>', 'team_membership:837848e9-670e-4_read_redis_rows <- _fill_from_redisasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>_reserve_counters <- _reserve_reservable_countersPIPELINE SET+SET+MGET+EVALSHA['<team>_<user>', 'team_membership:837848e9-670e-4atch_rate_limiter_script <- should_rate_limitPIPELINE EVALSHA+EVALSHA['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318heck_and_increment_by_n <- reserve_tpm_tokensasync_get_cache<key-hash>ache <- _retrieve_from_cacheasync_set_cache<key-hash>async_set_cache <- async_add_cachePIPELINE MGET['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_collect_inflight <- _loadasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>ed <- increment_spend_countersasync_incrementgpt-dep-a:openai/gpt-4o-mini:tpm:<min>async_increment_cache <- async_log_success_eventasync_batch_get_cache['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0']async_batch_get_cache <- _read_update_cache_valuesEVALSHA rate-limiter check['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3_execute_token_increment_script <- async_increment_tokens_with_ttl_preservationPOST /v1/chat/completions, non-streaming, simple-shuffle
HTTP 200,"model": "shuffle-group", mock replyPIPELINE MGET+MGET['<team>_<user>', 'team_membership:837848e9-670e-4_read_redis_rows <- _fill_from_redisasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>_reserve_counters <- _reserve_reservable_countersPIPELINE SET+SET+MGET+EVALSHA['<team>_<user>', 'team_membership:837848e9-670e-4atch_rate_limiter_script <- should_rate_limitPIPELINE EVALSHA+EVALSHA['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318heck_and_increment_by_n <- reserve_tpm_tokensasync_get_cache<key-hash>ache <- _retrieve_from_cacheasync_set_cache<key-hash>async_set_cache <- async_add_cachePIPELINE MGET['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_collect_inflight <- _loadasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>ed <- increment_spend_countersasync_incrementshuffle-dep-a:openai/gpt-4o-mini:tpm:<min>async_increment_cache <- async_log_success_eventasync_batch_get_cache['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0']async_batch_get_cache <- _read_update_cache_valuesEVALSHA rate-limiter check['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3_execute_token_increment_script <- async_increment_tokens_with_ttl_preservationPOST /v1/chat/completions, streaming
HTTP 200,"model": "gpt-group", mock replyPIPELINE MGET+MGET['<team>_<user>', 'team_membership:837848e9-670e-4_read_redis_rows <- _fill_from_redisasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>_reserve_counters <- _reserve_reservable_countersPIPELINE SET+SET+MGET+EVALSHA['<team>_<user>', 'team_membership:837848e9-670e-4atch_rate_limiter_script <- should_rate_limitPIPELINE EVALSHA+EVALSHA['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318heck_and_increment_by_n <- reserve_tpm_tokensasync_get_cache<key-hash>ache <- _retrieve_from_cacheasync_set_cache<key-hash>async_set_cache <- async_add_cacheasync_set_cache<key-hash>async_set_cache <- async_add_cachePIPELINE MGET['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_collect_inflight <- _loadasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>s <- _update_database_and_spend_counters_in_batchasync_incrementgpt-dep-a:None:tpm:<min>async_increment_cache <- async_log_success_eventasync_batch_get_cache['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0']async_batch_get_cache <- _read_update_cache_valuesEVALSHA rate-limiter check['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3_execute_token_increment_script <- async_increment_tokens_with_ttl_preservationPOST /v1/responses
HTTP 200,"model": "gpt-group", mock replyPIPELINE MGET+MGET['<team>_<user>', 'team_membership:837848e9-670e-4_read_redis_rows <- _fill_from_redisasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>_reserve_counters <- _reserve_reservable_countersPIPELINE SET+SET+MGET+EVALSHA['<team>_<user>', 'team_membership:837848e9-670e-4_batch_rate_limiter_script <- should_rate_limitPIPELINE EVALSHA+EVALSHA['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318_check_and_increment_by_n <- reserve_tpm_tokensasync_get_cache<key-hash>_ageneric_api_call_with_fallbacks_responses_attempt <- make_callget_cacheSYNC <key-hash>get_cache <- _sync_get_cacheset_cacheSYNC <key-hash>add_cache <- sync_set_cacheincrement_cacheSYNC gpt-dep-b:openai/gpt-4o-mini:tpm:<min>increment_cache <- log_success_eventasync_set_cache<key-hash>async_set_cache <- async_add_cachePIPELINE MGET['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_collect_inflight <- _loadasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>ed <- increment_spend_countersasync_incrementgpt-dep-b:openai/gpt-4o-mini:tpm:<min>async_increment_cache <- async_log_success_eventasync_batch_get_cache['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0']async_batch_get_cache <- _read_update_cache_valuesEVALSHA rate-limiter check['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3_execute_token_increment_script <- async_increment_tokens_with_ttl_preservationPOST /v1/messages
HTTP 200,"model": "gpt-group", mock replyDEFAULT_MANAGEMENT_OBJECT_IN_MEMORY_CACHE_TTL). 16 of the 21 are the auth refresh path inuser_api_key_auth(end user 5, key 3, team 5, model-access-group registry 3: miss, DB read, re-set, run one at a time before the request batch is armed); the remaining 5 are the same batched pre-call as chat. The other /v1/messages sample in the same run (messages_cold) shows 4 pre-call trips, identical to chat, and the same refresh landed onchat_coldandresponses_cold. See the refresh-path breakdown below.async_get_cacheend_user_id:perf-enduserasync_get_cache <- async_get_cacheasync_get_cacheend_user_restricted_registryasync_get_cache <- async_get_cacheasync_get_cacheend_user_restricted_registryasync_get_cache <- async_get_cacheasync_set_cacheend_user_restricted_registry_cache_registry_answer <- _fetch_and_cache_registryasync_set_cacheend_user_id:perf-enduserasync_set_cache <- async_set_cacheasync_get_cache<key-hash>async_get_cache <- async_get_cacheasync_get_cache<key-hash>async_get_cache <- async_get_cacheasync_set_cache<key-hash>async_set_cache <- async_set_cacheasync_get_cacheteam_id:<team>get_team_object <- _get_hierarchical_router_settingsasync_delete_cacheteam_id:<team>get_team_object <- _get_hierarchical_router_settingsasync_set_cacheteam_id:<team>get_team_object <- _get_hierarchical_router_settingsdelete_cacheSYNC team_alias:perf-teamget_team_object <- _get_hierarchical_router_settingsasync_delete_cacheteam_alias:perf-teamget_team_object <- _get_hierarchical_router_settingsPIPELINE MGET+MGET['<team>_<user>', '837848e9-670e-4b34-936e-156409a_read_redis_rows <- _fill_from_redisasync_get_cachemodel_access_group_registrycached_registry <- _load_bounded_registryasync_get_cachemodel_access_group_registrycached_registry <- _load_bounded_registryasync_set_cachemodel_access_group_registry_cache_registry <- _load_bounded_registryasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>_reserve_counters <- _reserve_reservable_countersPIPELINE SET+SET+SET+MGET+EVALSHA['<user>', '<team>_837848e9-670e-4b34-936e-156409ah_rate_limiter_script <- should_rate_limitPIPELINE EVALSHA+EVALSHA['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318k_and_increment_by_n <- reserve_tpm_tokensasync_get_cache<key-hash>ropic_messages_attempt <- make_callasync_set_cache<key-hash>async_add_cache <- _complete_cache_write_despite_cancellationPIPELINE MGET['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_collect_inflight <- _loadasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>ed <- increment_spend_countersasync_incrementgpt-dep-a:None:tpm:<min>async_increment_cache <- async_log_success_eventasync_batch_get_cache['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0']async_batch_get_cache <- _read_update_cache_valuesEVALSHA rate-limiter check['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3_execute_token_increment_script <- async_increment_tokens_with_ttl_preservationPOST /v1/chat/completions, response-cache hit
HTTP 200,"model": "gpt-group", mock replyPIPELINE MGET+MGET['<team>_<user>', 'team_membership:837848e9-670e-4_read_redis_rows <- _fill_from_redisasync_set_maxspend:team:<team>_repair_stale_spend_counter <- get_current_spendasync_set_maxspend:end_user:perf-enduser_repair_stale_spend_counter <- get_current_spendasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>_reserve_counters <- _reserve_reservable_countersPIPELINE SET+SET+MGET+EVALSHA['<team>_<user>', 'team_membership:837848e9-670e-4atch_rate_limiter_script <- should_rate_limitPIPELINE EVALSHA+EVALSHA['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318heck_and_increment_by_n <- reserve_tpm_tokensPIPELINE MGET['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_collect_inflight <- _loadasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>ed <- increment_spend_countersasync_incrementgpt-dep-a:gpt-4o-mini:tpm:<min>async_increment_cache <- async_log_success_eventasync_batch_get_cache['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0']async_batch_get_cache <- _read_update_cache_valuesEVALSHA rate-limiter check['{api_key:<key-hash>}:tokens', '{user:837848e9-670e-4b3_execute_token_increment_script <- async_increment_tokens_with_ttl_preservationAfter (68cb506d143d68d25fd0fa8aa27ac8f1efde0b8e)
Post-call is two round trips on every async path: one read pipeline (the spend reconcile
MGETand theupdate_cacheMGET), then one write pipeline carrying the response-cacheSET, every spend and deployment TPM increment and the tokenEVALSHA. The pre-call rows are unchanged from #43407; the cache-hit arm differs pre-call only because the before run repaired two stale spend counters (async_set_max) that the after run found already repairedPOST /v1/chat/completions, non-streaming, usage-based routing
HTTP 200,"model": "gpt-group", mock replyPIPELINE MGET+MGET['<team>_<user>', 'team_membership:837848e9-670e-4_read_redis_rows <- _fill_from_redisasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>_reserve_counters <- _reserve_reservable_countersPIPELINE SET+SET+MGET+EVALSHA['<team>_<user>', 'team_membership:837848e9-670e-4atch_rate_limiter_script <- should_rate_limitPIPELINE EVALSHA+EVALSHA['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318_wrap_awaitable <- invokeasync_get_cache<key-hash>ache <- _retrieve_from_cachePIPELINE MGET+MGET['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0', 'spend:end_user:perf-enduser'_collect_inflight <- _loadPIPELINE SET+Increment+Increment+Increment+Increment+Increment+Increment+Increment+Increment+EVALSHA['<key-hash>', 'spend:key:ee876d0d5cd318717fe4f5a0a848f8invoke <- invokePOST /v1/chat/completions, non-streaming, simple-shuffle
HTTP 200,"model": "shuffle-group", mock replyPIPELINE MGET+MGET['<team>_<user>', 'team_membership:837848e9-670e-4_read_redis_rows <- _fill_from_redisasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>_reserve_counters <- _reserve_reservable_countersPIPELINE SET+SET+MGET+EVALSHA['<team>_<user>', 'team_membership:837848e9-670e-4atch_rate_limiter_script <- should_rate_limitPIPELINE EVALSHA+EVALSHA['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318_wrap_awaitable <- invokeasync_get_cache<key-hash>ache <- _retrieve_from_cachePIPELINE MGET+MGET['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0', 'spend:end_user:perf-enduser'_collect_inflight <- _loadPIPELINE SET+Increment+Increment+Increment+Increment+Increment+Increment+Increment+Increment+EVALSHA['<key-hash>', 'spend:key:ee876d0d5cd318717fe4f5a0a848f8invoke <- invokePOST /v1/chat/completions, streaming
HTTP 200,"model": "gpt-group", mock replyPIPELINE MGET+MGET['<team>_<user>', 'team_membership:837848e9-670e-4_read_redis_rows <- _fill_from_redisasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>_reserve_counters <- _reserve_reservable_countersPIPELINE SET+SET+MGET+EVALSHA['<team>_<user>', 'team_membership:837848e9-670e-4atch_rate_limiter_script <- should_rate_limitPIPELINE EVALSHA+EVALSHA['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318_wrap_awaitable <- invokeasync_get_cache<key-hash>ache <- _retrieve_from_cachePIPELINE MGET+MGET['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0', 'spend:end_user:perf-enduser'_collect_inflight <- _loadPIPELINE SET+SET+Increment+Increment+Increment+Increment+Increment+Increment+Increment+Increment+EVALSHA['<key-hash>', '2deb1366c70583e925626e7b712eb99a3f50fd3dinvoke <- invokePOST /v1/responses
HTTP 200,"model": "gpt-group", mock replyPIPELINE MGET+MGET['<team>_<user>', 'team_membership:837848e9-670e-4_read_redis_rows <- _fill_from_redisasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>_reserve_counters <- _reserve_reservable_countersPIPELINE SET+SET+MGET+EVALSHA['<team>_<user>', 'team_membership:837848e9-670e-4_batch_rate_limiter_script <- should_rate_limitPIPELINE EVALSHA+EVALSHA['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318_wrap_awaitable <- invokeasync_get_cache<key-hash>_ageneric_api_call_with_fallbacks_responses_attempt <- make_callget_cacheSYNC <key-hash>get_cache <- _sync_get_cacheset_cacheSYNC <key-hash>add_cache <- sync_set_cacheincrement_cacheSYNC gpt-dep-a:openai/gpt-4o-mini:tpm:<min>increment_cache <- log_success_eventPIPELINE MGET+MGET['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0', 'spend:end_user:perf-enduser'_collect_inflight <- _loadPIPELINE SET+Increment+Increment+Increment+Increment+Increment+Increment+Increment+Increment+EVALSHA['<key-hash>', 'spend:key:ee876d0d5cd318717fe4f5a0a848f8invoke <- invokePOST /v1/messages
HTTP 200,"model": "gpt-group", mock replyDEFAULT_MANAGEMENT_OBJECT_IN_MEMORY_CACHE_TTL). 16 of the 21 are the auth refresh path inuser_api_key_auth(end user 5, key 3, team 5, model-access-group registry 3: miss, DB read, re-set, run one at a time before the request batch is armed); the remaining 5 are the same batched pre-call as chat. The other /v1/messages sample in the same run (messages_cold) shows 4 pre-call trips, identical to chat, and the same refresh landed onchat_coldandresponses_cold. See the refresh-path breakdown below.async_get_cacheend_user_id:perf-enduserasync_get_cache <- async_get_cacheasync_get_cacheend_user_restricted_registryasync_get_cache <- async_get_cacheasync_get_cacheend_user_restricted_registryasync_get_cache <- async_get_cacheasync_set_cacheend_user_restricted_registry_cache_registry_answer <- _fetch_and_cache_registryasync_set_cacheend_user_id:perf-enduserasync_set_cache <- async_set_cacheasync_get_cache<key-hash>async_get_cache <- async_get_cacheasync_get_cache<key-hash>async_get_cache <- async_get_cacheasync_set_cache<key-hash>async_set_cache <- async_set_cacheasync_get_cacheteam_id:<team>get_team_object <- _get_hierarchical_router_settingsasync_delete_cacheteam_id:<team>get_team_object <- _get_hierarchical_router_settingsasync_set_cacheteam_id:<team>get_team_object <- _get_hierarchical_router_settingsdelete_cacheSYNC team_alias:perf-teamget_team_object <- _get_hierarchical_router_settingsasync_delete_cacheteam_alias:perf-teamget_team_object <- _get_hierarchical_router_settingsPIPELINE MGET+MGET['<team>_<user>', '837848e9-670e-4b34-936e-156409a_read_redis_rows <- _fill_from_redisasync_get_cachemodel_access_group_registrycached_registry <- _load_bounded_registryasync_get_cachemodel_access_group_registrycached_registry <- _load_bounded_registryasync_set_cachemodel_access_group_registry_cache_registry <- _load_bounded_registryasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>_reserve_counters <- _reserve_reservable_countersPIPELINE SET+SET+SET+MGET+EVALSHA['<user>', '<team>_837848e9-670e-4b34-936e-156409ah_rate_limiter_script <- should_rate_limitPIPELINE EVALSHA+EVALSHA['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318_wrap_awaitable <- invokeasync_get_cache<key-hash>ropic_messages_attempt <- make_callPIPELINE MGET+MGET['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0', 'spend:end_user:perf-enduser'_collect_inflight <- _loadPIPELINE SET+Increment+Increment+Increment+Increment+EVALSHA['<key-hash>', 'spend:key:ee876d0d5cd318717fe4f5a0a848f8invoke <- invokePOST /v1/chat/completions, response-cache hit
HTTP 200,"model": "gpt-group", mock replyPIPELINE MGET+MGET['<team>_<user>', 'team_membership:837848e9-670e-4_read_redis_rows <- _fill_from_redisasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>_reserve_counters <- _reserve_reservable_countersPIPELINE SET+SET+MGET+EVALSHA['<team>_<user>', 'team_membership:837848e9-670e-4atch_rate_limiter_script <- should_rate_limitPIPELINE EVALSHA+EVALSHA['{api_key:<key-hash>}:window', '{api_key:ee876d0d5cd318_wrap_awaitable <- invokePIPELINE MGET+MGET['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0', 'spend:end_user:perf-enduser'_collect_inflight <- _loadPIPELINE Increment+Increment+Increment+Increment+EVALSHA['spend:key:<key-hash>', 'spend:team:6690cc1c-3fd6-418a-invoke <- invokeFollow-up at c18e51e
The update_cache read is now declared only after the spend write to the database succeeds, so a cached spend another callback writes during that write is read instead of an older copy. Rerunning the live harness at c18e51e against the #43407 tip dc0344f with the same fixtures:
Mutation check for the ordering: moving
arm_update_cache_readback in front ofupdate_databaseturnstest_the_update_cache_read_sees_a_cached_spend_written_while_the_spend_was_persistedred with the stale spend 1.0 instead of 5.0, and restoring it turns the test greenThe follow-up test commits at 799fddb hand the limiter its scripts through the fake cache's
async_register_scriptand drive the post-call deadline off the event loop clock instead of a wall-clock sleep. Making_flush_on_deadlineskip its flush turnstest_a_post_call_batch_nobody_closes_goes_out_at_the_deadlinered, and restoring it turns the test greenThe cancellation follow-ups at 868b0c2 and a65ce6d stop a cancelled post-call flush from deleting the shared spend counter, since a cancel does not prove the INCRBYFLOAT missed Redis. The cancelled spend is added to a local counter that already exists and an absent one is left unseeded, so a later Redis outage cannot undercount it. Removing the local increment, or the existence guard, turns
test_a_cancelled_post_call_flush_keeps_the_shared_spend_counter_and_counts_the_spend_locallyred. a2f12ee reads the existing post-call batch into aFinalinstead of rebinding itRe-run at the current head a2f12ee against real Redis, Postgres and Anthropic claude-sonnet-4-6, all requests 200:
Admin UI at a2f12ee
The same virtual key request from the fixture, POST http://localhost:4000/v1/chat/completions with
sk-perf-vk-0000000001against Anthropic claude-sonnet-4-6, shows up on http://localhost:4000/ui/?page=logs as a Success row with its real spend, so the batched Redis path still records spend end to endLive re-check at eb4379a
This head merges the #43407 head, which keys the request and post-call batches by namespace as well as connection settings. A read-count harness, run from this checkout with
PYTHONPATH=<checkout>, one warm request then one request per case,redis_reads_processedfromINFO statsminus the three marker commandsEvery response body was
pongfromclaude-groupwith usage populated, and the proxy log has no ERROR or Traceback linesClean re-run, no cache-expiry samples
The tables above were captured with samples 12 to 16 s apart, so roughly every fourth sample landed on the 60 s management-object TTL and paid a 16-call auth refresh (end user 5, key 3, team 5, model-access registry 3) that has nothing to do with the endpoint sampled; that is where the 21 pre-call trips on the
/v1/messagessample came from. The refresh path itself is fixed in #43776 (P4). This re-run, done in a separate session, sends a warm-up request right before every traced request and pauses before every fourth warm-up so the expiry happens outside the sample windows; both arms were checked for refresh signatures inside every window (0 found). Before is the #43407 tip fada595, After is 9dad494, the same tree as the current tip 073f036 minus the #43407 log-line fix it now carries, same harness, same fixtures, real Redis 6.0.16 and Postgres 14, every warm-up and sample HTTP 200,chat_cachehitserved the same response id aschat_warmin both arms. The before arm needed one rerun (its first attempt had refresh signatures inmessages_coldandshuffle_warm), the after arm nonechat_coldchat_stream_coldshuffle_coldmessages_coldresponses_coldchat_warmchat_stream_warmshuffle_warmmessages_warmresponses_warmchat_cachehitPost-call on
/v1/chat/completions(both routing groups, streaming or not) and/v1/messagesgoes from 7 or 8 trips to 3: the spend counter read pipeline, theupdate_cacheread pipeline and the one write pipeline (response cache SET, spend increments, deployment TPM, token Lua)./v1/responsesstays at 5 because its native handler's two synchronous worker-thread calls are outside the request scope. Pre-call is unchanged by this PR, as intendedType
🚄 Infrastructure
Caveats (if any)
Medium
DualCacheresponse cache, parallel slots) are updated at onceINCRBYFLOATdid before. A failed token Lua group falls back to the plain increment pipeline for that group. A parallel slot is removed from the local gauge the moment it is released, so admission on the same worker sees the capacity before the pipeline goes out; the count Redis returns from the pipelined release is not written over the gauge (a newer acquire on the worker may have mirrored a fresher count by then; the next acquire refreshes it from Redis as before), and a failed release Lua leaves the slot released in memory only/v1/responsesnative handlers still do two synchronous calls (set_cache,increment_cache) from a worker thread; they run outside the async request scope and are not batched, exactly as before this PRLow
WeakSet, andproxy_shutdown_eventdrains them (drain_post_call_redis_batches) before the cache disconnects, so a worker stopping inside the callback or deadline window sends the pending writes rather than dropping themSETcoerces the caller'sttlexactly asBaseCache.get_ttldoes (int(ttl), a non-numeric value means the cache default), sottl="3600"expires in 3600 s on both pathsnx=True(and any non-Redis cache) stay direct; the deferredSETuses the same key, TTL and serialised bytesRedisCache.async_set_cachewould sendFinal Attestation
Link to Devin session: https://app.devin.ai/sessions/f732394a6aca4bc197a2f4b1d54d0e7e
Open in Devin Desktop: https://app.devin.ai/desktop/session/f732394a6aca4bc197a2f4b1d54d0e7e?variant=devin
Requested by: @yassin-berriai