Repository navigation
perf(proxy): hold one spend counter batch across admission and across post-call accounting - #43369
Conversation
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
|
|
|
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
|
bugbot run |
|
bugbot run |
|
bugbot run |
|
bugbot run |
1 similar comment
|
bugbot run |
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 178bd46. Configure here.
… post-call accounting Auth's spend counter MGET scope spans common checks, model budget check and reservation; reservation increments go out as one pipeline; post-call reconcile adjustments ride the ordinary increment pipeline and update_cache uses one batched read. Over-budget reservation counters are charged one at a time so a rejection never touches the counters after it; post-call counter keys are derived from ids without validating a UserAPIKeyAuth. Resolves LIT-8881 Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
178bd46 to
a51b69e
Compare
TLDR
Problem this solves:
How it solves it:
Files changed
litellm/caching/dual_cache.pylitellm/proxy/auth/user_api_key_auth.pylitellm/proxy/hooks/proxy_track_cost_callback.pylitellm/proxy/proxy_server.pylitellm/proxy/spend_tracking/budget_reservation.pylitellm/proxy/spend_tracking/spend_counter_batch.pyUser Flow
Before: a developer whose key has a budget, a team budget, an end-user budget and TPM/RPM limits waits on 22 Redis round trips per request
"model": "gpt-group"and"user": "perf-enduser"After: the same request uses fewer Redis round trips, returns the same 200, and records the same spend
Relevant issues
Stacked on #43320; merge after it. Second of the stacked PRs for one Redis pipeline pre-call and one post-call (LIT-8881, then LIT-8882 and LIT-8883); design and per-request measurements are on LIT-8881
Linear ticket
Resolves LIT-8881
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/unit/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Screenshots / Proof of Fix
Same fixture bytes for both arms,
PYTHONPATH=<checkout>, real Redis 6.0.16 at 127.0.0.1:6379, real Postgres, mock deployments (litellm_params.mock_response; the change is on the admission and spend path, not a provider path). EveryRedisCachemethod call is logged with its caller chain by asitecustomizetracer loaded throughPYTHONPATH; the tables below list every call between the request marker and the response marker, minus the background jobs that fire on timers (_sync_in_memory_spend_with_redis, daily tag spend flush, config prefetch). Calls are round trips: a pipeline or a Lua script counts once.Fixture:
config_full.yamlwith twogpt-groupdeployments (gpt-dep-a,gpt-dep-b,usage-based-routing-v2routing group), Redis response cache on,enable_redis_auth_cache: true, top-levelsimple-shuffle. Virtual key withmax_budget: 1000,tpm_limit: 10000000,rpm_limit: 100000, in teamperf-team(same budget and limits), user withmax_budget: 1000, end userperf-enduseron a budget of 1000. Requests are sent 12 s apart so the warm pass hits warm caches; the tables are the warm pass.Before (464b584, the #43320 tip this PR is stacked on)
POST /v1/chat/completions, non-streaming
HTTP 200,"model": "gpt-group", contentmock reply from aasync_batch_get_cache['<team>_<user>', 'team_membership:837848e9-670e-4_fill_from_redis <- prefetch_auth_objectsasync_set_cache_pipeline_with_ttls(('<team>_<user>', {'user_id': '837848e9-670e-4b34_write_back <- _fill_from_dbasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_fetch <- _loadasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',read_spend_counter_cache_value <- _is_spend_counter_cache_warmasync_incrementspend:key:<key-hash>_increment_spend_counter_cache <- _reserve_counterasync_incrementspend:team:<team>_increment_spend_counter_cache <- _reserve_counterasync_incrementspend:end_user:perf-enduser_increment_spend_counter_cache <- _reserve_counterEVALSHA rate-limiter['{api_key:<key-hash>}:window', '{api_key:<key-hash>}:tokens']_batch_rate_limiter_script <- should_rate_limitEVALSHA tpm check-and-increment['{api_key:<key-hash>}:window', '{api_key:<key-hash>}:tokens']_check_and_increment_by_n <- reserve_tpm_tokensEVALSHA tpm check-and-increment['{team:<team>}:window', '{team:<team>}:tokens']_check_and_increment_by_n <- reserve_tpm_tokensasync_batch_get_cache['deployment:gpt-dep-a:cooldown', 'deployment:gpt-dep-b:cooldown', 'deployment:shuffle-dep-a:cooldown', 'deplo_cooldown_deployments <- async_get_healthy_deploymentsasync_get_cache<key-hash>async_get_cache <- _retrieve_from_cacheasync_set_cache<key-hash>async_set_cache <- async_add_cacheasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_fetch <- _loadasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>reconcile_budget_reservation <- _reconcile_budget_reservation_before_db_updateasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_fetch <- _loadasync_increment_pipeline[{'key': 'spend:team_member:<user>:<team>', 'incre_apply_spend_counter_increments <- _increment_spend_counters_batchedasync_incrementgpt-dep-a:openai/gpt-4o-mini:tpm:<min>async_increment_cache <- async_log_success_eventasync_get_cachedefault_user_id:spendasync_get_cache <- async_get_cacheasync_get_cachetag:User-Agent: curlasync_get_cache <- async_get_cacheasync_get_cachetag:User-Agent: curl/7.81.0async_get_cache <- async_get_cacheEVALSHA rate-limiter['{api_key:<key-hash>}:tokens', '{user:<user>}:tokens', ...]_execute_token_increment_script <- async_increment_tokens_with_ttl_preservationasync_batch_get_cache['<team>_<user>', 'team_membership:837848e9-670e-4_fill_from_redis <- prefetch_auth_objectsasync_set_cache_pipeline_with_ttls(('<team>_<user>', {'user_id': '837848e9-670e-4b34_write_back <- _fill_from_dbasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_fetch <- _loadasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',read_spend_counter_cache_value <- _is_spend_counter_cache_warmasync_incrementspend:key:<key-hash>_increment_spend_counter_cache <- _reserve_counterasync_incrementspend:team:<team>_increment_spend_counter_cache <- _reserve_counterasync_incrementspend:end_user:perf-enduser_increment_spend_counter_cache <- _reserve_counterEVALSHA rate-limiter['{api_key:<key-hash>}:window', '{api_key:<key-hash>}:tokens']_batch_rate_limiter_script <- should_rate_limitEVALSHA tpm check-and-increment['{api_key:<key-hash>}:window', '{api_key:<key-hash>}:tokens']_check_and_increment_by_n <- reserve_tpm_tokensEVALSHA tpm check-and-increment['{team:<team>}:window', '{team:<team>}:tokens']_check_and_increment_by_n <- reserve_tpm_tokensasync_batch_get_cache['deployment:gpt-dep-a:cooldown', 'deployment:gpt-dep-b:cooldown', 'deployment:shuffle-dep-a:cooldown', 'deplo_cooldown_deployments <- async_get_healthy_deploymentsasync_get_cache<key-hash>async_get_cache <- _retrieve_from_cacheasync_set_cache<key-hash>async_set_cache <- async_add_cacheasync_set_cache<key-hash>async_set_cache <- async_add_cacheasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_fetch <- _loadasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>_reconcile_budget_reservation_before_db_update <- _update_database_and_spend_countersasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_fetch <- _loadasync_increment_pipeline[{'key': 'spend:team_member:<user>:<team>', 'incre_increment_spend_counters_batched <- increment_spend_countersasync_incrementgpt-dep-a:None:tpm:<min>async_increment_cache <- async_log_success_eventasync_get_cachedefault_user_id:spendasync_get_cache <- async_get_cacheasync_get_cachetag:User-Agent: curlasync_get_cache <- async_get_cacheasync_get_cachetag:User-Agent: curl/7.81.0async_get_cache <- async_get_cacheEVALSHA rate-limiter['{api_key:<key-hash>}:tokens', '{user:<user>}:tokens', ...]_execute_token_increment_script <- async_increment_tokens_with_ttl_preservationPOST /v1/chat/completions, streaming
HTTP 200, 8 SSE chunks,"model": "gpt-group", contentmock reply from aasync_batch_get_cache['<team>_<user>', 'team_membership:837848e9-670e-4_fill_from_redis <- prefetch_auth_objectsasync_set_cache_pipeline_with_ttls(('<team>_<user>', {'user_id': '837848e9-670e-4b34_write_back <- _fill_from_dbasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_fetch <- _loadasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',read_spend_counter_cache_value <- _is_spend_counter_cache_warmasync_incrementspend:key:<key-hash>_increment_spend_counter_cache <- _reserve_counterasync_incrementspend:team:<team>_increment_spend_counter_cache <- _reserve_counterasync_incrementspend:end_user:perf-enduser_increment_spend_counter_cache <- _reserve_counterEVALSHA rate-limiter['{api_key:<key-hash>}:window', '{api_key:<key-hash>}:tokens']_batch_rate_limiter_script <- should_rate_limitEVALSHA tpm check-and-increment['{api_key:<key-hash>}:window', '{api_key:<key-hash>}:tokens']_check_and_increment_by_n <- reserve_tpm_tokensEVALSHA tpm check-and-increment['{team:<team>}:window', '{team:<team>}:tokens']_check_and_increment_by_n <- reserve_tpm_tokensasync_batch_get_cache['deployment:gpt-dep-a:cooldown', 'deployment:gpt-dep-b:cooldown', 'deployment:shuffle-dep-a:cooldown', 'deplo_cooldown_deployments <- async_get_healthy_deploymentsasync_get_cache<key-hash>async_get_cache <- _retrieve_from_cacheasync_set_cache<key-hash>async_set_cache <- async_add_cacheasync_set_cache<key-hash>async_set_cache <- async_add_cacheasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_fetch <- _loadasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>_reconcile_budget_reservation_before_db_update <- _update_database_and_spend_countersasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_fetch <- _loadasync_increment_pipeline[{'key': 'spend:team_member:<user>:<team>', 'incre_increment_spend_counters_batched <- increment_spend_countersasync_incrementgpt-dep-a:None:tpm:<min>async_increment_cache <- async_log_success_eventasync_get_cachedefault_user_id:spendasync_get_cache <- async_get_cacheasync_get_cachetag:User-Agent: curlasync_get_cache <- async_get_cacheasync_get_cachetag:User-Agent: curl/7.81.0async_get_cache <- async_get_cacheEVALSHA rate-limiter['{api_key:<key-hash>}:tokens', '{user:<user>}:tokens', ...]_execute_token_increment_script <- async_increment_tokens_with_ttl_preservationPOST /v1/responses
HTTP 200,"model": "gpt-group", contentmock reply from basync_set_maxrows are the pre-existing stale counter repair (_repair_stale_spend_counter), which fires whenever the DB spend the auth checks loaded reads above the counter and is unchanged by this PR| 40 | post |
EVALSHA rate-limiter|['{api_key:<key-hash>}:tokens', '{user:<user>}:tokens', ...]|_execute_token_increment_script <- async_increment_tokens_with_ttl_preservation|async_batch_get_cache['<team>_<user>', 'team_membership:837848e9-670e-4_fill_from_redis <- prefetch_auth_objectsasync_set_cache_pipeline_with_ttls(('<team>_<user>', {'user_id': '837848e9-670e-4b34_write_back <- _fill_from_dbasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_fetch <- _loadasync_set_maxspend:key:<key-hash>_repair_stale_spend_counter <- get_current_spendasync_batch_get_cache['spend:key:<key-hash>']_fetch <- _loadasync_set_maxspend:team:<team>_repair_stale_spend_counter <- get_current_spendasync_set_maxspend:end_user:perf-enduser_repair_stale_spend_counter <- get_current_spendasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',read_spend_counter_cache_value <- _is_spend_counter_cache_warmasync_incrementspend:key:<key-hash>_increment_spend_counter_cache <- _reserve_counterasync_incrementspend:team:<team>_increment_spend_counter_cache <- _reserve_counterasync_incrementspend:end_user:perf-enduser_increment_spend_counter_cache <- _reserve_counterEVALSHA rate-limiter['{api_key:<key-hash>}:window', '{api_key:<key-hash>}:tokens']_batch_rate_limiter_script <- should_rate_limitEVALSHA tpm check-and-increment['{api_key:<key-hash>}:window', '{api_key:<key-hash>}:tokens']_check_and_increment_by_n <- reserve_tpm_tokensEVALSHA tpm check-and-increment['{team:<team>}:window', '{team:<team>}:tokens']_check_and_increment_by_n <- reserve_tpm_tokensasync_batch_get_cache['deployment:gpt-dep-a:cooldown', 'deployment:gpt-dep-b:cooldown', 'deployment:shuffle-dep-a:cooldown', 'deplo_api_call_with_fallbacks_responses_attempt <- make_callasync_get_cache<key-hash>_ageneric_api_call_with_fallbacks_responses_attempt <- make_callget_cacheSYNC <key-hash>get_cache <- _sync_get_cacheset_cacheSYNC <key-hash>add_cache <- sync_set_cacheincrement_cacheSYNC gpt-dep-b:openai/gpt-4o-mini:tpm:<min>increment_cache <- log_success_eventasync_set_cache<key-hash>async_set_cache <- async_add_cacheasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_fetch <- _loadasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>reconcile_budget_reservation <- _reconcile_budget_reservation_before_db_updateasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_fetch <- _loadasync_increment_pipeline[{'key': 'spend:team_member:<user>:<team>', 'incre_apply_spend_counter_increments <- _increment_spend_counters_batchedasync_incrementgpt-dep-b:openai/gpt-4o-mini:tpm:<min>async_increment_cache <- async_log_success_eventasync_get_cachedefault_user_id:spendasync_get_cache <- async_get_cacheasync_get_cachetag:User-Agent: curlasync_get_cache <- async_get_cacheasync_get_cachetag:User-Agent: curl/7.81.0async_get_cache <- async_get_cacheEVALSHA rate-limiter['{api_key:<key-hash>}:tokens', '{user:<user>}:tokens', ...]_execute_token_increment_script <- async_increment_tokens_with_ttl_preservationAfter (d35ede6; aaeee9d only changes the over-budget reservation path and the update_cache read throttle, neither of which this warm fixture exercises, so the tables hold for the tip)
POST /v1/chat/completions, non-streaming
HTTP 200,"model": "gpt-group", contentmock reply from b(the other deployment of the same group;usage-based-routing-v2picks by the minute's usage, both deployments served requests in both arms)spendcolumns all read 0.0081075, so the merged pipeline records each request's cost exactly onceasync_batch_get_cache['<team>_<user>', 'team_membership:837848e9-670e-4_fill_from_redis <- prefetch_auth_objectsasync_set_cache_pipeline_with_ttls(('<team>_<user>', {'user_id': '837848e9-670e-4b34_write_back <- _fill_from_dbasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_fetch <- _loadasync_set_maxspend:team:<team>_repair_stale_spend_counter <- get_current_spendasync_set_maxspend:end_user:perf-enduser_repair_stale_spend_counter <- get_current_spendasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>_reserve_counters <- reserve_budget_for_requestEVALSHA rate-limiter['{api_key:<key-hash>}:window', '{api_key:<key-hash>}:tokens']_batch_rate_limiter_script <- should_rate_limitEVALSHA tpm check-and-increment['{api_key:<key-hash>}:window', '{api_key:<key-hash>}:tokens']_check_and_increment_by_n <- reserve_tpm_tokensEVALSHA tpm check-and-increment['{team:<team>}:window', '{team:<team>}:tokens']_check_and_increment_by_n <- reserve_tpm_tokensasync_batch_get_cache['deployment:gpt-dep-a:cooldown', 'deployment:gpt-dep-b:cooldown', 'deployment:shuffle-dep-a:cooldown', 'deplo_cooldown_deployments <- async_get_healthy_deploymentsasync_get_cache<key-hash>async_get_cache <- _retrieve_from_cacheasync_set_cache<key-hash>async_set_cache <- async_add_cacheasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_fetch <- _loadasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>_increment_spend_counters_batched <- increment_spend_countersasync_incrementgpt-dep-b:openai/gpt-4o-mini:tpm:<min>async_increment_cache <- async_log_success_eventasync_batch_get_cache['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0']async_batch_get_cache <- _read_update_cache_valuesEVALSHA rate-limiter['{api_key:<key-hash>}:tokens', '{user:<user>}:tokens', ...]_execute_token_increment_script <- async_increment_tokens_with_ttl_preservationPOST /v1/chat/completions, streaming
HTTP 200, 8 SSE chunks,"model": "gpt-group", contentmock reply from basync_batch_get_cache['<team>_<user>', 'team_membership:837848e9-670e-4_fill_from_redis <- prefetch_auth_objectsasync_set_cache_pipeline_with_ttls(('<team>_<user>', {'user_id': '837848e9-670e-4b34_write_back <- _fill_from_dbasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_fetch <- _loadasync_set_maxspend:team:<team>_repair_stale_spend_counter <- get_current_spendasync_set_maxspend:end_user:perf-enduser_repair_stale_spend_counter <- get_current_spendasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>_reserve_counters <- reserve_budget_for_requestEVALSHA rate-limiter['{api_key:<key-hash>}:window', '{api_key:<key-hash>}:tokens']_batch_rate_limiter_script <- should_rate_limitEVALSHA tpm check-and-increment['{api_key:<key-hash>}:window', '{api_key:<key-hash>}:tokens']_check_and_increment_by_n <- reserve_tpm_tokensEVALSHA tpm check-and-increment['{team:<team>}:window', '{team:<team>}:tokens']_check_and_increment_by_n <- reserve_tpm_tokensasync_batch_get_cache['deployment:gpt-dep-a:cooldown', 'deployment:gpt-dep-b:cooldown', 'deployment:shuffle-dep-a:cooldown', 'deplo_cooldown_deployments <- async_get_healthy_deploymentsasync_get_cache<key-hash>async_get_cache <- _retrieve_from_cacheasync_set_cache<key-hash>async_set_cache <- async_add_cacheasync_set_cache<key-hash>async_set_cache <- async_add_cacheasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_fetch <- _loadasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>increment_spend_counters <- _update_database_and_spend_counters_in_batchasync_incrementgpt-dep-b:None:tpm:<min>async_increment_cache <- async_log_success_eventasync_batch_get_cache['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0']async_batch_get_cache <- _read_update_cache_valuesEVALSHA rate-limiter['{api_key:<key-hash>}:tokens', '{user:<user>}:tokens', ...]_execute_token_increment_script <- async_increment_tokens_with_ttl_preservationPOST /v1/responses
HTTP 200,"model": "gpt-group", contentmock reply from aasync_batch_get_cache['<team>_<user>', 'team_membership:837848e9-670e-4_fill_from_redis <- prefetch_auth_objectsasync_set_cache_pipeline_with_ttls(('<team>_<user>', {'user_id': '837848e9-670e-4b34_write_back <- _fill_from_dbasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_fetch <- _loadasync_set_maxspend:key:<key-hash>_repair_stale_spend_counter <- get_current_spendasync_set_maxspend:team:<team>_repair_stale_spend_counter <- get_current_spendasync_set_maxspend:end_user:perf-enduser_repair_stale_spend_counter <- get_current_spendasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>_reserve_counters <- reserve_budget_for_requestEVALSHA rate-limiter['{api_key:<key-hash>}:window', '{api_key:<key-hash>}:tokens']_batch_rate_limiter_script <- should_rate_limitEVALSHA tpm check-and-increment['{api_key:<key-hash>}:window', '{api_key:<key-hash>}:tokens']_check_and_increment_by_n <- reserve_tpm_tokensEVALSHA tpm check-and-increment['{team:<team>}:window', '{team:<team>}:tokens']_check_and_increment_by_n <- reserve_tpm_tokensasync_batch_get_cache['deployment:gpt-dep-a:cooldown', 'deployment:gpt-dep-b:cooldown', 'deployment:shuffle-dep-a:cooldown', 'deplo_api_call_with_fallbacks_responses_attempt <- make_callasync_get_cache<key-hash>_ageneric_api_call_with_fallbacks_responses_attempt <- make_callget_cacheSYNC <key-hash>get_cache <- _sync_get_cacheset_cacheSYNC <key-hash>add_cache <- sync_set_cacheincrement_cacheSYNC gpt-dep-a:openai/gpt-4o-mini:tpm:<min>increment_cache <- log_success_eventasync_set_cache<key-hash>async_set_cache <- async_add_cacheasync_batch_get_cache['spend:end_user:perf-enduser', 'spend:key:<key-hash>',_fetch <- _loadasync_increment_pipeline[{'key': 'spend:key:<key-hash>', 'increment_value': <cost>_increment_spend_counters_batched <- increment_spend_countersasync_incrementgpt-dep-a:openai/gpt-4o-mini:tpm:<min>async_increment_cache <- async_log_success_eventasync_batch_get_cache['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0']async_batch_get_cache <- _read_update_cache_valuesEVALSHA rate-limiter['{api_key:<key-hash>}:tokens', '{user:<user>}:tokens', ...]_execute_token_increment_script <- async_increment_tokens_with_ttl_preservationNot measured here: a fleet load regime. These are single-request command counts on one box; the prod traces on LIT-8881 are what motivated the change.
Admin UI at 2c5b5bf
The same virtual key request from the fixture, POST http://localhost:4000/v1/chat/completions with
sk-perf-vk-0000000001against Anthropic claude-sonnet-4-6, shows up on http://localhost:4000/ui/?page=logs as a Success row with its real spend, so the batched Redis path still records spend end to endLive re-check at a69c50b (fde219b after it only rewrites one auth test so it asserts that admission and reservation share one spend counter MGET; moving the batch release ahead of the reservation turns it red)
This head adds the merge of the #43320 head, plus two commits that load the reservable budget counters through an async generator frozen into a tuple (same order, same fail-closed rejection, no mutable accumulator and no recursion). A read-count harness, run from this checkout with
PYTHONPATH=<checkout>, one warm request then one request per case,redis_reads_processedfromINFO statsminus the three marker commandsEvery response body was
pongfromclaude-groupwith usage populated, and the proxy log has no ERROR or Traceback linesType
🚄 Infrastructure
Caveats (if any)
Medium
tests/test_litellm/proxy/test_budget_reservation.pyandtest_spend_counter_batch.pytest_reserved_counter_deleted_during_spend_write_is_reseeded_instead_of_going_negativepins it, and theredis-cli monitorrerun on the virtual key fixture shows exactly one added MGET (32 -> 33 chat, 33 -> 34 messages)reconcile_budget_reservation(apply_consistent=False)returns the consistent adjustments unwritten; the caller must write them and then callstamp_budget_reservation_actual_cost. The two post-call callers do; a failed accounting pipeline invalidates the reserved counters and propagates, which is whattest_budget_reservation_redis_failure.pynow assertsLow
update_cachereads its cached objects withasync_batch_get_cache(throttle_redis=False), so unlike other batch readers it never skips a key that missed memory withinredis_batch_cache_expiry; that matches the per-key GET it replaces_repair_stale_spend_counterfires on every request of this fixture because the DBspendcolumn (0.006486000000000002) reads above the Redis counter (0.006486, INCRBYFLOAT's 17 significant digit formatting) by 2e-18. Pre-existing on the base arm too (the/v1/responsestable); tracked as a follow-up on LIT-8884The remaining pre-call round trips (identity MGET and write-back, three rate-limiter Lua calls, routing MGET, response-cache GET) and post-call round trips (response-cache SET, deployment TPM INCR, rate-limiter token Lua) are the subject of the next two PRs in the stack (LIT-8882, LIT-8883)
Final Attestation
Link to Devin session: https://app.devin.ai/sessions/f3b64ab771264ecf9b474add1ce663de
Open in Devin Desktop: https://app.devin.ai/desktop/session/f3b64ab771264ecf9b474add1ce663de?variant=devin
Requested by: @yassin-berriai
Note
High Risk
Changes core budget reservation, reconcile, and spend-counter Redis semantics (batch lifetimes, single vs serial reservation, pipeline failure handling), which directly affects enforcement accuracy and spend accounting under concurrency and Redis failures.
Overview
This PR cuts duplicate Redis work on the proxy budget and spend path by sharing one spend-counter batch across admission, reservation, and post-call accounting, and by merging pipelines that used to run separately.
Pre-call:
release_spend_counter_batch()now runs in afinallyblock after model budget checks and_reserve_budget_after_common_checks, so admission and reservation reuse one MGET snapshot. Budget reservation charges counters viarun_spend_counter_pipeline(one INCRBYFLOAT pipeline when every counter still fits the estimate; otherwise one-at-a-time so a rejection never touches later counters). Stale counter repair records the repaired value into the open batch instead of forcing another read.Post-call: The cost callback wraps the DB spend write and counter updates in
spend_counter_batch_scope.reconcile_budget_reservation(apply_consistent=False)returns reconcile deltas asPendingSpendIncremententries that are combined with ordinary increments in a single pipeline;stamp_budget_reservation_actual_costupdates reservation metadata after that write. A failed post-call pipeline invalidates affected counters and propagates the error (no silent “reconciled but not incremented” state).Cache:
DualCache.async_batch_get_cachegainsthrottle_redis=Falsesoupdate_cachecan batch-read user/team/tag objects without skipping Redis for keys that recently missed in memory—matching prior per-key GET behavior.update_cacheuses one batched read instead of multipleasync_get_cachecalls.Other:
post_call_counter_keysignores non-string id placeholders; spend counter pipeline logic is split intorun_spend_counter_pipelinevs invalidation-on-failure at the caller.Reviewed by Cursor Bugbot for commit 178bd46. Bugbot is set up for automated code reviews on this repo. Configure here.
Link to Devin session: https://app.devin.ai/sessions/2ca52cb470de404ca807c7265c6c5f66
Open in Devin Desktop: https://app.devin.ai/desktop/session/2ca52cb470de404ca807c7265c6c5f66?variant=devin