Skip to content

perf(proxy): hold one spend counter batch across admission and across post-call accounting - #43369

Merged
yassin-berriai merged 1 commit into
mainfrom
litellm_redis_p1_spend_batch_request_scope
Sep 29, 2026
Merged

yassin-berriai merged 1 commit into
mainfrom
litellm_redis_p1_spend_batch_request_scope

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Admission and post-call accounting read the same spend counters more than once
  • Budget reservation writes each counter separately
  • Cache refreshes read related objects one at a time

How it solves it:

  • One spend counter batch now spans admission checks and budget reservation
  • Reservation and post-call counter updates use Redis pipelines
  • Related cache objects are refreshed with one batched read

Files changed

Files What changed
litellm/caching/dual_cache.py Let cache refreshes fetch every memory miss from Redis, even after a recent lookup
litellm/proxy/auth/user_api_key_auth.py Kept the spend counter batch open through budget checks and reservation
litellm/proxy/hooks/proxy_track_cost_callback.py Reused the counter batch during post-call accounting
litellm/proxy/proxy_server.py Added batched counter reads, writes, and cache refresh handling
litellm/proxy/spend_tracking/budget_reservation.py Batched reservations and reconciled reserved costs safely
litellm/proxy/spend_tracking/spend_counter_batch.py Shared counter-key handling across admission and post-call paths
flowchart LR
  A[Admission read] --> B[Reservation pipeline]
  B --> C[Provider call]
  C --> D[Post-call read]
  D --> E[DB write and increment pipeline]
  E --> F[Batch cache refresh]
Loading

User Flow

Before: a developer whose key has a budget, a team budget, an end-user budget and TPM/RPM limits waits on 22 Redis round trips per request

  1. They send POST https://litellm-domain/v1/chat/completions with "model": "gpt-group" and "user": "perf-enduser"
  2. Auth reads their spend counters, then the budget reservation reads the same counters again and increments them one at a time
  3. The response comes back 200 with the mock reply
  4. Behind the response the proxy reconciles the reservation in its own pipeline, then increments the counters in another, then refreshes the cached key/user/team/tag objects with one GET each

After: the same request uses fewer Redis round trips, returns the same 200, and records the same spend

  1. They send the same POST https://litellm-domain/v1/chat/completions
  2. Auth reads the spend counters once; the reservation reuses that read and reserves every counter in one pipeline (a counter that no longer fits the estimate is charged on its own, so a rejection never touches the counters after it)
  3. The response comes back 200 with the same mock reply
  4. Behind the response the proxy reconciles the reservation from one MGET before the DB write, reads the post-call counters with one MGET after it, writes the increments in one pipeline, and refreshes the cached objects from one batched read

Relevant issues

Stacked on #43320; merge after it. Second of the stacked PRs for one Redis pipeline pre-call and one post-call (LIT-8881, then LIT-8882 and LIT-8883); design and per-request measurements are on LIT-8881

Linear ticket

Resolves LIT-8881

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/unit/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Screenshots / Proof of Fix

Same fixture bytes for both arms, PYTHONPATH=<checkout>, real Redis 6.0.16 at 127.0.0.1:6379, real Postgres, mock deployments (litellm_params.mock_response; the change is on the admission and spend path, not a provider path). Every RedisCache method call is logged with its caller chain by a sitecustomize tracer loaded through PYTHONPATH; the tables below list every call between the request marker and the response marker, minus the background jobs that fire on timers (_sync_in_memory_spend_with_redis, daily tag spend flush, config prefetch). Calls are round trips: a pipeline or a Lua script counts once.

Fixture: config_full.yaml with two gpt-group deployments (gpt-dep-a, gpt-dep-b, usage-based-routing-v2 routing group), Redis response cache on, enable_redis_auth_cache: true, top-level simple-shuffle. Virtual key with max_budget: 1000, tpm_limit: 10000000, rpm_limit: 100000, in team perf-team (same budget and limits), user with max_budget: 1000, end user perf-enduser on a budget of 1000. Requests are sent 12 s apart so the warm pass hits warm caches; the tables are the warm pass.

python litellm/proxy/proxy_cli.py --config ~/perf_rt/config_full.yaml --port <port>
curl -s -H "Authorization: Bearer $KEY" -H "Content-Type: application/json" http://127.0.0.1:<port>/v1/chat/completions -d @chat_warm.req.json
# chat_warm.req.json
{"model":"gpt-group","user":"perf-enduser","messages":[{"role":"user","content":"Reply with the single word: pong B"}],"max_tokens":5}
# chat_stream_warm.req.json
{"model":"gpt-group","user":"perf-enduser","stream":true,"stream_options":{"include_usage":true},"messages":[{"role":"user","content":"Reply with the single word: pong B"}],"max_tokens":5}
# responses_warm.req.json
{"model":"gpt-group","user":"perf-enduser","input":"Reply with the single word: pong B","max_output_tokens":5}

Before (464b584, the #43320 tip this PR is stacked on)

POST /v1/chat/completions, non-streaming

  1. HTTP 200, "model": "gpt-group", content mock reply from a
  2. 22 Redis round trips on the request path (12 pre-call, 10 post-call). Rows 3 and 4 are the same MGET twice (auth checks, then the reservation warm check); rows 5 to 7 are one INCRBYFLOAT per reserved counter; rows 14 to 17 are two MGET plus two pipelines for what is one accounting write; rows 19 to 21 are one GET per cached object
# phase RedisCache op keys caller
1 pre async_batch_get_cache ['<team>_<user>', 'team_membership:837848e9-670e-4 _fill_from_redis <- prefetch_auth_objects
2 pre async_set_cache_pipeline_with_ttls (('<team>_<user>', {'user_id': '837848e9-670e-4b34 _write_back <- _fill_from_db
3 pre async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _fetch <- _load
4 pre async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', read_spend_counter_cache_value <- _is_spend_counter_cache_warm
5 pre async_increment spend:key:<key-hash> _increment_spend_counter_cache <- _reserve_counter
6 pre async_increment spend:team:<team> _increment_spend_counter_cache <- _reserve_counter
7 pre async_increment spend:end_user:perf-enduser _increment_spend_counter_cache <- _reserve_counter
8 pre EVALSHA rate-limiter ['{api_key:<key-hash>}:window', '{api_key:<key-hash>}:tokens'] _batch_rate_limiter_script <- should_rate_limit
9 pre EVALSHA tpm check-and-increment ['{api_key:<key-hash>}:window', '{api_key:<key-hash>}:tokens'] _check_and_increment_by_n <- reserve_tpm_tokens
10 pre EVALSHA tpm check-and-increment ['{team:<team>}:window', '{team:<team>}:tokens'] _check_and_increment_by_n <- reserve_tpm_tokens
11 pre async_batch_get_cache ['deployment:gpt-dep-a:cooldown', 'deployment:gpt-dep-b:cooldown', 'deployment:shuffle-dep-a:cooldown', 'deplo _cooldown_deployments <- async_get_healthy_deployments
12 pre async_get_cache <key-hash> async_get_cache <- _retrieve_from_cache
13 post async_set_cache <key-hash> async_set_cache <- async_add_cache
14 post async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _fetch <- _load
15 post async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> reconcile_budget_reservation <- _reconcile_budget_reservation_before_db_update
16 post async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _fetch <- _load
17 post async_increment_pipeline [{'key': 'spend:team_member:<user>:<team>', 'incre _apply_spend_counter_increments <- _increment_spend_counters_batched
18 post async_increment gpt-dep-a:openai/gpt-4o-mini:tpm:<min> async_increment_cache <- async_log_success_event
19 post async_get_cache default_user_id:spend async_get_cache <- async_get_cache
20 post async_get_cache tag:User-Agent: curl async_get_cache <- async_get_cache
21 post async_get_cache tag:User-Agent: curl/7.81.0 async_get_cache <- async_get_cache
22 post EVALSHA rate-limiter ['{api_key:<key-hash>}:tokens', '{user:<user>}:tokens', ...] _execute_token_increment_script <- async_increment_tokens_with_ttl_preservation
# phase RedisCache op keys caller
1 pre async_batch_get_cache ['<team>_<user>', 'team_membership:837848e9-670e-4 _fill_from_redis <- prefetch_auth_objects
2 pre async_set_cache_pipeline_with_ttls (('<team>_<user>', {'user_id': '837848e9-670e-4b34 _write_back <- _fill_from_db
3 pre async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _fetch <- _load
4 pre async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', read_spend_counter_cache_value <- _is_spend_counter_cache_warm
5 pre async_increment spend:key:<key-hash> _increment_spend_counter_cache <- _reserve_counter
6 pre async_increment spend:team:<team> _increment_spend_counter_cache <- _reserve_counter
7 pre async_increment spend:end_user:perf-enduser _increment_spend_counter_cache <- _reserve_counter
8 pre EVALSHA rate-limiter ['{api_key:<key-hash>}:window', '{api_key:<key-hash>}:tokens'] _batch_rate_limiter_script <- should_rate_limit
9 pre EVALSHA tpm check-and-increment ['{api_key:<key-hash>}:window', '{api_key:<key-hash>}:tokens'] _check_and_increment_by_n <- reserve_tpm_tokens
10 pre EVALSHA tpm check-and-increment ['{team:<team>}:window', '{team:<team>}:tokens'] _check_and_increment_by_n <- reserve_tpm_tokens
11 pre async_batch_get_cache ['deployment:gpt-dep-a:cooldown', 'deployment:gpt-dep-b:cooldown', 'deployment:shuffle-dep-a:cooldown', 'deplo _cooldown_deployments <- async_get_healthy_deployments
12 pre async_get_cache <key-hash> async_get_cache <- _retrieve_from_cache
13 post async_set_cache <key-hash> async_set_cache <- async_add_cache
14 post async_set_cache <key-hash> async_set_cache <- async_add_cache
15 post async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _fetch <- _load
16 post async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> _reconcile_budget_reservation_before_db_update <- _update_database_and_spend_counters
17 post async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _fetch <- _load
18 post async_increment_pipeline [{'key': 'spend:team_member:<user>:<team>', 'incre _increment_spend_counters_batched <- increment_spend_counters
19 post async_increment gpt-dep-a:None:tpm:<min> async_increment_cache <- async_log_success_event
20 post async_get_cache default_user_id:spend async_get_cache <- async_get_cache
21 post async_get_cache tag:User-Agent: curl async_get_cache <- async_get_cache
22 post async_get_cache tag:User-Agent: curl/7.81.0 async_get_cache <- async_get_cache
23 post EVALSHA rate-limiter ['{api_key:<key-hash>}:tokens', '{user:<user>}:tokens', ...] _execute_token_increment_script <- async_increment_tokens_with_ttl_preservation

POST /v1/chat/completions, streaming

  1. HTTP 200, 8 SSE chunks, "model": "gpt-group", content mock reply from a
  2. 23 Redis round trips on the request path (12 pre-call, 11 post-call), same shape as non-streaming plus the stream's second response-cache write
# phase RedisCache op keys caller
1 pre async_batch_get_cache ['<team>_<user>', 'team_membership:837848e9-670e-4 _fill_from_redis <- prefetch_auth_objects
2 pre async_set_cache_pipeline_with_ttls (('<team>_<user>', {'user_id': '837848e9-670e-4b34 _write_back <- _fill_from_db
3 pre async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _fetch <- _load
4 pre async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', read_spend_counter_cache_value <- _is_spend_counter_cache_warm
5 pre async_increment spend:key:<key-hash> _increment_spend_counter_cache <- _reserve_counter
6 pre async_increment spend:team:<team> _increment_spend_counter_cache <- _reserve_counter
7 pre async_increment spend:end_user:perf-enduser _increment_spend_counter_cache <- _reserve_counter
8 pre EVALSHA rate-limiter ['{api_key:<key-hash>}:window', '{api_key:<key-hash>}:tokens'] _batch_rate_limiter_script <- should_rate_limit
9 pre EVALSHA tpm check-and-increment ['{api_key:<key-hash>}:window', '{api_key:<key-hash>}:tokens'] _check_and_increment_by_n <- reserve_tpm_tokens
10 pre EVALSHA tpm check-and-increment ['{team:<team>}:window', '{team:<team>}:tokens'] _check_and_increment_by_n <- reserve_tpm_tokens
11 pre async_batch_get_cache ['deployment:gpt-dep-a:cooldown', 'deployment:gpt-dep-b:cooldown', 'deployment:shuffle-dep-a:cooldown', 'deplo _cooldown_deployments <- async_get_healthy_deployments
12 pre async_get_cache <key-hash> async_get_cache <- _retrieve_from_cache
13 post async_set_cache <key-hash> async_set_cache <- async_add_cache
14 post async_set_cache <key-hash> async_set_cache <- async_add_cache
15 post async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _fetch <- _load
16 post async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> _reconcile_budget_reservation_before_db_update <- _update_database_and_spend_counters
17 post async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _fetch <- _load
18 post async_increment_pipeline [{'key': 'spend:team_member:<user>:<team>', 'incre _increment_spend_counters_batched <- increment_spend_counters
19 post async_increment gpt-dep-a:None:tpm:<min> async_increment_cache <- async_log_success_event
20 post async_get_cache default_user_id:spend async_get_cache <- async_get_cache
21 post async_get_cache tag:User-Agent: curl async_get_cache <- async_get_cache
22 post async_get_cache tag:User-Agent: curl/7.81.0 async_get_cache <- async_get_cache
23 post EVALSHA rate-limiter ['{api_key:<key-hash>}:tokens', '{user:<user>}:tokens', ...] _execute_token_increment_script <- async_increment_tokens_with_ttl_preservation

POST /v1/responses

  1. HTTP 200, "model": "gpt-group", content mock reply from b
  2. 29 Redis round trips on the request path (17 pre-call, 12 post-call). Same admission shape; the three async_set_max rows are the pre-existing stale counter repair (_repair_stale_spend_counter), which fires whenever the DB spend the auth checks loaded reads above the counter and is unchanged by this PR

| 40 | post | EVALSHA rate-limiter | ['{api_key:<key-hash>}:tokens', '{user:<user>}:tokens', ...] | _execute_token_increment_script <- async_increment_tokens_with_ttl_preservation |

# phase RedisCache op keys caller
1 pre async_batch_get_cache ['<team>_<user>', 'team_membership:837848e9-670e-4 _fill_from_redis <- prefetch_auth_objects
2 pre async_set_cache_pipeline_with_ttls (('<team>_<user>', {'user_id': '837848e9-670e-4b34 _write_back <- _fill_from_db
3 pre async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _fetch <- _load
4 pre async_set_max spend:key:<key-hash> _repair_stale_spend_counter <- get_current_spend
5 pre async_batch_get_cache ['spend:key:<key-hash>'] _fetch <- _load
6 pre async_set_max spend:team:<team> _repair_stale_spend_counter <- get_current_spend
7 pre async_set_max spend:end_user:perf-enduser _repair_stale_spend_counter <- get_current_spend
8 pre async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', read_spend_counter_cache_value <- _is_spend_counter_cache_warm
9 pre async_increment spend:key:<key-hash> _increment_spend_counter_cache <- _reserve_counter
10 pre async_increment spend:team:<team> _increment_spend_counter_cache <- _reserve_counter
11 pre async_increment spend:end_user:perf-enduser _increment_spend_counter_cache <- _reserve_counter
12 pre EVALSHA rate-limiter ['{api_key:<key-hash>}:window', '{api_key:<key-hash>}:tokens'] _batch_rate_limiter_script <- should_rate_limit
13 pre EVALSHA tpm check-and-increment ['{api_key:<key-hash>}:window', '{api_key:<key-hash>}:tokens'] _check_and_increment_by_n <- reserve_tpm_tokens
14 pre EVALSHA tpm check-and-increment ['{team:<team>}:window', '{team:<team>}:tokens'] _check_and_increment_by_n <- reserve_tpm_tokens
15 pre async_batch_get_cache ['deployment:gpt-dep-a:cooldown', 'deployment:gpt-dep-b:cooldown', 'deployment:shuffle-dep-a:cooldown', 'deplo _api_call_with_fallbacks_responses_attempt <- make_call
16 pre async_get_cache <key-hash> _ageneric_api_call_with_fallbacks_responses_attempt <- make_call
17 pre get_cache SYNC <key-hash> get_cache <- _sync_get_cache
18 post set_cache SYNC <key-hash> add_cache <- sync_set_cache
19 post increment_cache SYNC gpt-dep-b:openai/gpt-4o-mini:tpm:<min> increment_cache <- log_success_event
20 post async_set_cache <key-hash> async_set_cache <- async_add_cache
21 post async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _fetch <- _load
22 post async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> reconcile_budget_reservation <- _reconcile_budget_reservation_before_db_update
23 post async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _fetch <- _load
24 post async_increment_pipeline [{'key': 'spend:team_member:<user>:<team>', 'incre _apply_spend_counter_increments <- _increment_spend_counters_batched
25 post async_increment gpt-dep-b:openai/gpt-4o-mini:tpm:<min> async_increment_cache <- async_log_success_event
26 post async_get_cache default_user_id:spend async_get_cache <- async_get_cache
27 post async_get_cache tag:User-Agent: curl async_get_cache <- async_get_cache
28 post async_get_cache tag:User-Agent: curl/7.81.0 async_get_cache <- async_get_cache
29 post EVALSHA rate-limiter ['{api_key:<key-hash>}:tokens', '{user:<user>}:tokens', ...] _execute_token_increment_script <- async_increment_tokens_with_ttl_preservation

After (d35ede6; aaeee9d only changes the over-budget reservation path and the update_cache read throttle, neither of which this warm fixture exercises, so the tables hold for the tip)

POST /v1/chat/completions, non-streaming

  1. HTTP 200, "model": "gpt-group", content mock reply from b (the other deployment of the same group; usage-based-routing-v2 picks by the minute's usage, both deployments served requests in both arms)
  2. 17 Redis round trips on the request path (11 pre-call, 6 post-call). One spend MGET pre-call (row 3) shared by the auth checks, the model budget check and the reservation; one reservation pipeline (row 6); one spend MGET (row 13) and one increment pipeline (row 14) post-call carrying the reconcile adjustments and the increments; one batched read for update_cache (row 16). Rows 4 and 5 are the same stale counter repair as before; it now records the repaired value into the open batch, so no MGET follows it
  3. DB spend after both arms: 46 spend logs sum to 0.0081075 and the key, team, user and end-user spend columns all read 0.0081075, so the merged pipeline records each request's cost exactly once
# phase RedisCache op keys caller
1 pre async_batch_get_cache ['<team>_<user>', 'team_membership:837848e9-670e-4 _fill_from_redis <- prefetch_auth_objects
2 pre async_set_cache_pipeline_with_ttls (('<team>_<user>', {'user_id': '837848e9-670e-4b34 _write_back <- _fill_from_db
3 pre async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _fetch <- _load
4 pre async_set_max spend:team:<team> _repair_stale_spend_counter <- get_current_spend
5 pre async_set_max spend:end_user:perf-enduser _repair_stale_spend_counter <- get_current_spend
6 pre async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> _reserve_counters <- reserve_budget_for_request
7 pre EVALSHA rate-limiter ['{api_key:<key-hash>}:window', '{api_key:<key-hash>}:tokens'] _batch_rate_limiter_script <- should_rate_limit
8 pre EVALSHA tpm check-and-increment ['{api_key:<key-hash>}:window', '{api_key:<key-hash>}:tokens'] _check_and_increment_by_n <- reserve_tpm_tokens
9 pre EVALSHA tpm check-and-increment ['{team:<team>}:window', '{team:<team>}:tokens'] _check_and_increment_by_n <- reserve_tpm_tokens
10 pre async_batch_get_cache ['deployment:gpt-dep-a:cooldown', 'deployment:gpt-dep-b:cooldown', 'deployment:shuffle-dep-a:cooldown', 'deplo _cooldown_deployments <- async_get_healthy_deployments
11 pre async_get_cache <key-hash> async_get_cache <- _retrieve_from_cache
12 post async_set_cache <key-hash> async_set_cache <- async_add_cache
13 post async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _fetch <- _load
14 post async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> _increment_spend_counters_batched <- increment_spend_counters
15 post async_increment gpt-dep-b:openai/gpt-4o-mini:tpm:<min> async_increment_cache <- async_log_success_event
16 post async_batch_get_cache ['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0'] async_batch_get_cache <- _read_update_cache_values
17 post EVALSHA rate-limiter ['{api_key:<key-hash>}:tokens', '{user:<user>}:tokens', ...] _execute_token_increment_script <- async_increment_tokens_with_ttl_preservation

POST /v1/chat/completions, streaming

  1. HTTP 200, 8 SSE chunks, "model": "gpt-group", content mock reply from b
  2. 18 Redis round trips on the request path (11 pre-call, 7 post-call)
# phase RedisCache op keys caller
1 pre async_batch_get_cache ['<team>_<user>', 'team_membership:837848e9-670e-4 _fill_from_redis <- prefetch_auth_objects
2 pre async_set_cache_pipeline_with_ttls (('<team>_<user>', {'user_id': '837848e9-670e-4b34 _write_back <- _fill_from_db
3 pre async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _fetch <- _load
4 pre async_set_max spend:team:<team> _repair_stale_spend_counter <- get_current_spend
5 pre async_set_max spend:end_user:perf-enduser _repair_stale_spend_counter <- get_current_spend
6 pre async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> _reserve_counters <- reserve_budget_for_request
7 pre EVALSHA rate-limiter ['{api_key:<key-hash>}:window', '{api_key:<key-hash>}:tokens'] _batch_rate_limiter_script <- should_rate_limit
8 pre EVALSHA tpm check-and-increment ['{api_key:<key-hash>}:window', '{api_key:<key-hash>}:tokens'] _check_and_increment_by_n <- reserve_tpm_tokens
9 pre EVALSHA tpm check-and-increment ['{team:<team>}:window', '{team:<team>}:tokens'] _check_and_increment_by_n <- reserve_tpm_tokens
10 pre async_batch_get_cache ['deployment:gpt-dep-a:cooldown', 'deployment:gpt-dep-b:cooldown', 'deployment:shuffle-dep-a:cooldown', 'deplo _cooldown_deployments <- async_get_healthy_deployments
11 pre async_get_cache <key-hash> async_get_cache <- _retrieve_from_cache
12 post async_set_cache <key-hash> async_set_cache <- async_add_cache
13 post async_set_cache <key-hash> async_set_cache <- async_add_cache
14 post async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _fetch <- _load
15 post async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> increment_spend_counters <- _update_database_and_spend_counters_in_batch
16 post async_increment gpt-dep-b:None:tpm:<min> async_increment_cache <- async_log_success_event
17 post async_batch_get_cache ['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0'] async_batch_get_cache <- _read_update_cache_values
18 post EVALSHA rate-limiter ['{api_key:<key-hash>}:tokens', '{user:<user>}:tokens', ...] _execute_token_increment_script <- async_increment_tokens_with_ttl_preservation

POST /v1/responses

  1. HTTP 200, "model": "gpt-group", content mock reply from a
  2. 21 Redis round trips on the request path (13 pre-call, 8 post-call)
# phase RedisCache op keys caller
1 pre async_batch_get_cache ['<team>_<user>', 'team_membership:837848e9-670e-4 _fill_from_redis <- prefetch_auth_objects
2 pre async_set_cache_pipeline_with_ttls (('<team>_<user>', {'user_id': '837848e9-670e-4b34 _write_back <- _fill_from_db
3 pre async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _fetch <- _load
4 pre async_set_max spend:key:<key-hash> _repair_stale_spend_counter <- get_current_spend
5 pre async_set_max spend:team:<team> _repair_stale_spend_counter <- get_current_spend
6 pre async_set_max spend:end_user:perf-enduser _repair_stale_spend_counter <- get_current_spend
7 pre async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> _reserve_counters <- reserve_budget_for_request
8 pre EVALSHA rate-limiter ['{api_key:<key-hash>}:window', '{api_key:<key-hash>}:tokens'] _batch_rate_limiter_script <- should_rate_limit
9 pre EVALSHA tpm check-and-increment ['{api_key:<key-hash>}:window', '{api_key:<key-hash>}:tokens'] _check_and_increment_by_n <- reserve_tpm_tokens
10 pre EVALSHA tpm check-and-increment ['{team:<team>}:window', '{team:<team>}:tokens'] _check_and_increment_by_n <- reserve_tpm_tokens
11 pre async_batch_get_cache ['deployment:gpt-dep-a:cooldown', 'deployment:gpt-dep-b:cooldown', 'deployment:shuffle-dep-a:cooldown', 'deplo _api_call_with_fallbacks_responses_attempt <- make_call
12 pre async_get_cache <key-hash> _ageneric_api_call_with_fallbacks_responses_attempt <- make_call
13 pre get_cache SYNC <key-hash> get_cache <- _sync_get_cache
14 post set_cache SYNC <key-hash> add_cache <- sync_set_cache
15 post increment_cache SYNC gpt-dep-a:openai/gpt-4o-mini:tpm:<min> increment_cache <- log_success_event
16 post async_set_cache <key-hash> async_set_cache <- async_add_cache
17 post async_batch_get_cache ['spend:end_user:perf-enduser', 'spend:key:<key-hash>', _fetch <- _load
18 post async_increment_pipeline [{'key': 'spend:key:<key-hash>', 'increment_value': <cost> _increment_spend_counters_batched <- increment_spend_counters
19 post async_increment gpt-dep-a:openai/gpt-4o-mini:tpm:<min> async_increment_cache <- async_log_success_event
20 post async_batch_get_cache ['default_user_id:spend', 'tag:User-Agent: curl', 'tag:User-Agent: curl/7.81.0'] async_batch_get_cache <- _read_update_cache_values
21 post EVALSHA rate-limiter ['{api_key:<key-hash>}:tokens', '{user:<user>}:tokens', ...] _execute_token_increment_script <- async_increment_tokens_with_ttl_preservation

Not measured here: a fleet load regime. These are single-request command counts on one box; the prod traces on LIT-8881 are what motivated the change.

Admin UI at 2c5b5bf

The same virtual key request from the fixture, POST http://localhost:4000/v1/chat/completions with sk-perf-vk-0000000001 against Anthropic claude-sonnet-4-6, shows up on http://localhost:4000/ui/?page=logs as a Success row with its real spend, so the batched Redis path still records spend end to end

Request Logs page after a live request at 2c5b5bf4d9

Live re-check at a69c50b (fde219b after it only rewrites one auth test so it asserts that admission and reservation share one spend counter MGET; moving the batch release ahead of the reservation turns it red)

This head adds the merge of the #43320 head, plus two commits that load the reservable budget counters through an async generator frozen into a tuple (same order, same fail-closed rejection, no mutable accumulator and no recursion). A read-count harness, run from this checkout with PYTHONPATH=<checkout>, one warm request then one request per case, redis_reads_processed from INFO stats minus the three marker commands

$ bash run_arm2.sh <arm> <checkout>  # redis-cli flushall, proxy on :4000 with config_v2.yaml, real Redis 7, real Postgres, Anthropic claude-sonnet-4-6
key/generate 200
warm 200
warm redis_reads_processed=15
chat_master 200
chat_master redis_reads_processed=8
chat_vk 200
chat_vk redis_reads_processed=16
messages_vk 200
messages_vk redis_reads_processed=17

Every response body was pong from claude-group with usage populated, and the proxy log has no ERROR or Traceback lines

Type

🚄 Infrastructure

Caveats (if any)

Medium

  • Reservation reserves every counter in one pipeline only when the admission MGET shows each still fits the estimate; otherwise the counters are charged one at a time in the original order, so the over-budget policy settles each before the next is touched and a rejection charges nothing after it. A concurrent request that lands between the MGET and the pipeline can still push a counter over, and that case is handled the same way as before: the over-budget policy resizes or releases the applied entries afterwards. Covered by tests/test_litellm/proxy/test_budget_reservation.py and test_spend_counter_batch.py
  • When the reservation pipeline fails, every counter in it is invalidated (its increment may or may not have landed), where the serial path only invalidated the counter it was on. A counter that cannot be invalidated is released instead, as before
  • The reservation is reconciled from its own MGET before the spend is persisted, and the post-call counters are read in a second MGET after it (8254973, one extra round trip against d35ede6). Sharing one MGET across the DB write let a counter deleted during the write keep a stale value, so its negative adjustment landed on an absent key. test_reserved_counter_deleted_during_spend_write_is_reseeded_instead_of_going_negative pins it, and the redis-cli monitor rerun on the virtual key fixture shows exactly one added MGET (32 -> 33 chat, 33 -> 34 messages)
  • reconcile_budget_reservation(apply_consistent=False) returns the consistent adjustments unwritten; the caller must write them and then call stamp_budget_reservation_actual_cost. The two post-call callers do; a failed accounting pipeline invalidates the reserved counters and propagates, which is what test_budget_reservation_redis_failure.py now asserts

Low

  • update_cache reads its cached objects with async_batch_get_cache(throttle_redis=False), so unlike other batch readers it never skips a key that missed memory within redis_batch_cache_expiry; that matches the per-key GET it replaces

  • _repair_stale_spend_counter fires on every request of this fixture because the DB spend column (0.006486000000000002) reads above the Redis counter (0.006486, INCRBYFLOAT's 17 significant digit formatting) by 2e-18. Pre-existing on the base arm too (the /v1/responses table); tracked as a follow-up on LIT-8884

  • The remaining pre-call round trips (identity MGET and write-back, three rate-limiter Lua calls, routing MGET, response-cache GET) and post-call round trips (response-cache SET, deployment TPM INCR, rate-limiter token Lua) are the subject of the next two PRs in the stack (LIT-8882, LIT-8883)

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Link to Devin session: https://app.devin.ai/sessions/f3b64ab771264ecf9b474add1ce663de
Open in Devin Desktop: https://app.devin.ai/desktop/session/f3b64ab771264ecf9b474add1ce663de?variant=devin
Requested by: @yassin-berriai


Note

High Risk
Changes core budget reservation, reconcile, and spend-counter Redis semantics (batch lifetimes, single vs serial reservation, pipeline failure handling), which directly affects enforcement accuracy and spend accounting under concurrency and Redis failures.

Overview
This PR cuts duplicate Redis work on the proxy budget and spend path by sharing one spend-counter batch across admission, reservation, and post-call accounting, and by merging pipelines that used to run separately.

Pre-call: release_spend_counter_batch() now runs in a finally block after model budget checks and _reserve_budget_after_common_checks, so admission and reservation reuse one MGET snapshot. Budget reservation charges counters via run_spend_counter_pipeline (one INCRBYFLOAT pipeline when every counter still fits the estimate; otherwise one-at-a-time so a rejection never touches later counters). Stale counter repair records the repaired value into the open batch instead of forcing another read.

Post-call: The cost callback wraps the DB spend write and counter updates in spend_counter_batch_scope. reconcile_budget_reservation(apply_consistent=False) returns reconcile deltas as PendingSpendIncrement entries that are combined with ordinary increments in a single pipeline; stamp_budget_reservation_actual_cost updates reservation metadata after that write. A failed post-call pipeline invalidates affected counters and propagates the error (no silent “reconciled but not incremented” state).

Cache: DualCache.async_batch_get_cache gains throttle_redis=False so update_cache can batch-read user/team/tag objects without skipping Redis for keys that recently missed in memory—matching prior per-key GET behavior. update_cache uses one batched read instead of multiple async_get_cache calls.

Other: post_call_counter_keys ignores non-string id placeholders; spend counter pipeline logic is split into run_spend_counter_pipeline vs invalidation-on-failure at the caller.

Reviewed by Cursor Bugbot for commit 178bd46. Bugbot is set up for automated code reviews on this repo. Configure here.

Link to Devin session: https://app.devin.ai/sessions/2ca52cb470de404ca807c7265c6c5f66
Open in Devin Desktop: https://app.devin.ai/desktop/session/2ca52cb470de404ca807c7265c6c5f66?variant=devin

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@greptile-apps

greptile-apps Bot commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[High risk] Refactors spend counter batching across budget checks and cost accounting.

The PR appears safe to merge based on this review; no new actionable issue or outstanding previous finding remains.

Summary

The PR shares spend-counter reads across admission and reservation, batches reservation and post-call counter writes, and batches cached-object refresh reads. The latest change revises the test for the shared admission/reservation read.

Reviews (12) · Last reviewed commit: "test(proxy): prove admission and budget ..."

Comment thread litellm/proxy/spend_tracking/budget_reservation.py Outdated
Comment thread litellm/proxy/proxy_server.py
@codecov

codecov Bot commented Sep 26, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.05882% with 5 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
litellm/proxy/spend_tracking/budget_reservation.py 96.20% 3 Missing ⚠️
litellm/proxy/proxy_server.py 96.55% 2 Missing ⚠️

📢 Thoughts on this report? Let us know!

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread tests/test_litellm/proxy/spend_tracking/test_spend_counter_batch.py Outdated
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread tests/test_litellm/proxy/auth/test_user_api_key_auth.py
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread litellm/proxy/proxy_server.py Outdated
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread tests/test_litellm/proxy/test_budget_reservation.py Outdated
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread litellm/proxy/hooks/proxy_track_cost_callback.py
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

bugbot run

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

bugbot run

Comment thread litellm/proxy/spend_tracking/budget_reservation.py Outdated
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

bugbot run

Comment thread tests/test_litellm/proxy/auth/test_user_api_key_auth.py Outdated
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

bugbot run

1 similar comment
@yassin-berriai

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Base automatically changed from litellm_router_single_redis_round_trip to main September 29, 2026 20:52
@yassin-berriai
yassin-berriai requested a review from a team September 29, 2026 20:52
@yassin-berriai

Copy link
Copy Markdown
Contributor

@greptileai

@yassin-berriai

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 178bd46. Configure here.

@codspeed

codspeed Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_redis_p1_spend_batch_request_scope (a51b69e) with main (d2a574b)

Open in CodSpeed

… post-call accounting

Auth's spend counter MGET scope spans common checks, model budget check and reservation;
reservation increments go out as one pipeline; post-call reconcile adjustments ride the ordinary
increment pipeline and update_cache uses one batched read. Over-budget reservation counters are
charged one at a time so a rejection never touches the counters after it; post-call counter keys are
derived from ids without validating a UserAPIKeyAuth.

Resolves LIT-8881

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration
devin-ai-integration Bot force-pushed the litellm_redis_p1_spend_batch_request_scope branch from 178bd46 to a51b69e Compare September 29, 2026 21:33
@yassin-berriai
yassin-berriai merged commit 2d034bb into main Sep 29, 2026
100 of 103 checks passed
@yassin-berriai
yassin-berriai deleted the litellm_redis_p1_spend_batch_request_scope branch September 29, 2026 22:05

This branch was successfully deployed

1 active deployment
e2e-changed — a51b69ee Deployed Sep 29, 2026 by devin-ai-integration[bot] via oauth #1872
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants