Skip to content

fix(proxy): give spend_counter_cache its own 10000-entry in-memory cache - #42414

Closed
devin-ai-integration[bot] wants to merge 4 commits into
mainfrom
litellm_spend-counter-cache-eviction
Closed

devin-ai-integration[bot] wants to merge 4 commits into
mainfrom
litellm_spend-counter-cache-eviction

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Spend counters shared the generic 200-entry in-memory cache
  • Over 200 active budget scopes evicted still-valid counters
  • Evicted counters reseeded from lagging DB spend on the next request
  • A key just blocked for budget got admitted again

How it solves it:

  • spend_counter_cache gets its own InMemoryCache sized to 10000 entries
  • New constant SPEND_COUNTER_CACHE_MAX_SIZE; no other cache changes size
  • Unit test and an integration test that fail on the old size

User Flow

Before: a key that is over its budget gets admitted again once more than 200 other keys, teams, or end users have made requests

  1. The admin sends POST http://localhost:4000/key/generate with {"max_budget": 0.000001, "models": ["claude-haiku-4-5"]} and gets back a key sk-...
  2. A developer sends POST http://localhost:4000/v1/chat/completions with that key and gets a 200 with a completion
  3. The developer sends the same POST again and gets a 422 budget_exceeded: "Budget has been exceeded! Key=key (sk-...) Current cost: 1.3e-05, Max budget: 1e-06"
  4. Meanwhile 250 other end users (distinct "user" values) send POST http://localhost:4000/v1/chat/completions through a different key and each get a 200
  5. GET http://localhost:4000/key/info?key=sk-... still shows "spend": 0.0 because the batch writer has not flushed yet
  6. The developer sends the same POST on the over-budget key again and gets a 200 with a completion, spending past the budget

After: the over-budget key stays blocked while hundreds of other scopes take traffic

  1. The admin sends POST http://localhost:4000/key/generate with {"max_budget": 0.000001, "models": ["claude-haiku-4-5"]} and gets back a key sk-...
  2. A developer sends POST http://localhost:4000/v1/chat/completions with that key and gets a 200 with a completion
  3. The developer sends the same POST again and gets a 422 budget_exceeded: "Budget has been exceeded! Key=key (sk-...) Current cost: 1.3e-05, Max budget: 1e-06"
  4. Meanwhile 250 other end users (distinct "user" values) send POST http://localhost:4000/v1/chat/completions through a different key and each get a 200
  5. GET http://localhost:4000/key/info?key=sk-... still shows "spend": 0.0 because the batch writer has not flushed yet
  6. The developer sends the same POST on the over-budget key again and gets the same 422 budget_exceeded

Relevant issues

Fixes #40221

Affected release

Linear ticket

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Setup shared by both runs: one proxy worker on port 20221, Postgres 16, no Redis, real Anthropic claude-haiku-4-5 calls (max_tokens: 1), proxy_batch_write_at: 60 so the DB row lags the in-memory counter long enough to see the reseed

model_list:
  - model_name: claude-haiku-4-5
    litellm_params:
      model: anthropic/claude-haiku-4-5
      api_key: os.environ/ANTHROPIC_API_KEY
general_settings:
  master_key: os.environ/LITELLM_MASTER_KEY
  database_url: os.environ/DATABASE_URL
  proxy_batch_write_at: 60
P=http://localhost:20221
KEY=$(curl -s $P/key/generate -H "Authorization: Bearer $MK" -H 'content-type: application/json' -d '{"max_budget":0.000001,"models":["claude-haiku-4-5"]}' | python3 -c 'import sys,json;print(json.load(sys.stdin)["key"])')
FLOOD=$(curl -s $P/key/generate -H "Authorization: Bearer $MK" -H 'content-type: application/json' -d '{"models":["claude-haiku-4-5"]}' | python3 -c 'import sys,json;print(json.load(sys.stdin)["key"])')
chat() { curl -s -o /tmp/body -w "%{http_code}" $P/v1/chat/completions -H "Authorization: Bearer $1" -H 'content-type: application/json' -d "{\"model\":\"claude-haiku-4-5\",\"max_tokens\":1,\"messages\":[{\"role\":\"user\",\"content\":\"hi\"}],\"user\":\"$2\"}"; }
echo "1st call on budgeted key: $(chat $KEY probe-user)"
echo "2nd call on budgeted key: $(chat $KEY probe-user) $(head -c 160 /tmp/body)"
seq 1 250 | xargs -P 24 -I{} bash -c "$(declare -f chat); P=$P; chat $FLOOD flood-user-{} >/dev/null"
echo "DB spend of budgeted key: $(curl -s "$P/key/info?key=$KEY" -H "Authorization: Bearer $MK" | python3 -c 'import sys,json;print(json.load(sys.stdin)["info"]["spend"])')"
echo "3rd call on budgeted key after flood: $(chat $KEY probe-user) $(head -c 160 /tmp/body)"

Before (dfd8ffc)

  1. Run the script above against a proxy started from the merge base
  2. Observed output:
1st call on budgeted key (admitted, spends past budget): 200
2nd call on budgeted key (expect 400 budget exceeded): 422 {"error":{"message":"Budget has been exceeded! Key=key (sk-...zqCg) Current cost: 1.3000000000000001e-05, Max budget: 1e-06","type":"budget_exceeded","param":nu
flooding 250 distinct end_user scopes on a separate key...
DB spend of budgeted key right now: 0.0
3rd call on budgeted key after flood (expect 400 if counter survived): 200 {"id":"chatcmpl-0b779fa7-6d23-4dcc-8dbd-b2dd286ee530","created":1790050964,"model":"claude-haiku-4-5","object":"chat.completion","choices":[{"finish_reason":"le
  1. The third call is a 200: the over-budget key was admitted again after 250 other scopes took traffic

After (1de5a82)

  1. Run the same script against a proxy started from the PR tip
  2. Observed output:
1st call on budgeted key (admitted, spends past budget): 200
2nd call on budgeted key (expect 400 budget exceeded): 422 {"error":{"message":"Budget has been exceeded! Key=key (sk-...ouWg) Current cost: 1.3000000000000001e-05, Max budget: 1e-06","type":"budget_exceeded","param":nu
flooding 250 distinct end_user scopes on a separate key...
DB spend of budgeted key right now: 0.0
3rd call on budgeted key after flood (expect 400 if counter survived): 422 {"error":{"message":"Budget has been exceeded! Key=key (sk-...ouWg) Current cost: 1.3000000000000001e-05, Max budget: 1e-06","type":"budget_exceeded","param":nu
  1. The third call is the same 422 budget_exceeded as the second one

Regression tests: tests/test_litellm/proxy/test_proxy_server.py::test_over_budget_key_spend_survives_hundreds_of_other_scopes_incrementing drives increment_spend_counters for one key and 300 end users, then reads the key back through get_current_spend. It fails on the merge base (key spend read back as 0.0 after 300 other scopes were charged) and passes on the tip. tests/integration/spend/test_spend_counter_scope_churn.py::test_over_budget_key_stays_blocked_while_hundreds_of_end_users_take_traffic (group accounting, uv run --no-sync python tests/integration/run.py accounting) starts an owned local proxy without Redis and with proxy_batch_write_at: 600 against the scripted upstream, registers a model at 0.001 and 0.002 per token, generates a probe key with max_budget: 0.06 and a churn key, makes one 40-token call on the probe key (200), confirms the second call is a 422 budget_exceeded, sends 250 chat completions on the churn key with 250 distinct user values, then asserts the probe key still gets 422 budget_exceeded and that the upstream saw no request for it. With litellm/constants.py and litellm/proxy/proxy_server.py reverted to the fix commit's parent it fails at that assertion (over-budget key admitted after 250 other scopes took traffic: 200 ...) and the group passes 13 of 13 on the tip. No provider credentials or third party service are involved

Admin UI Playground, before and after

Same rig (one worker, Postgres, no Redis, real claude-haiku-4-5, proxy_batch_write_at: 60). Before run captured at merge base 55e95c0279, after run at 1de5a82e04 (the later b0a174a7 and 675de2f1 commits only touch tests). Reproduce it with:

  1. Open http://localhost:20221/ui/api-keys/ and click Create New Key. Pick claude-haiku-4-5, open Optional Settings, set Max Budget to 0.000001, click Create Key and copy the virtual key
  2. Open http://localhost:20221/ui/playground/. Set Virtual Key Source to Virtual Key, paste the key, select claude-haiku-4-5, open Model Settings, enable Use Advanced Parameters and set Max Tokens to 1
  3. Send hi. The first message gets a completion; send hi again and the Playground shows 422 Budget has been exceeded
  4. From a shell, send 250 chat completions through a second key with a different "user" on each (the seq 1 250 | xargs line above); all return 200 and GET /key/info on the probe key still shows "spend": 0.0
  5. Send hi a third time in the Playground

Before (55e95c0279): blocked after the second message, then admitted again after the churn

Playground shows 422 budget exceeded on the second message, before the churn

Playground admits the same key again after 250 other end users made requests

After (1de5a82e04): blocked after the second message and still blocked after the churn, DB spend still 0.0

Playground shows 422 budget exceeded on the second message, before the churn

Playground still shows 422 budget exceeded on the third message after the churn

4-worker base vs tip, all three provider surfaces

Same Postgres and Anthropic, both trees started with --num_workers 4 and proxy_batch_write_at: 600 so the DB row stays at 0.0 for the whole churn. For each of /v1/chat/completions, /v1/messages and /v1/responses: new key with max_budget: 0.000001, one call (200), 12 concurrent warm calls, then 8 concurrent probes that must all be 422 before the churn counts. Churn is 1200 chat completions on another key with 1200 distinct user values (48 in flight, 13s). Then 12 concurrent probes

Merge base 55e95c0279: probes after the churn were 200 422 422 422 200 422 200 422 422 422 422 200 (chat), 200 200 200 422 422 422 422 422 200 422 422 422 (messages), 200 422 422 200 422 422 422 422 422 422 422 422 (responses), DB spend 0.0 before and after. One admission per worker whose counter was evicted. A settled over-budget key was also admitted 2 of 12 times right after Postgres was paused and unpaused under a 30-request mixed burst

Tip 675de2f1: 12 of 12 probes 422 on all three surfaces with the same pre-check, churn and DB spend 0.0, and 8 of 8 probes 422 while Postgres was paused plus 12 of 12 after unpause. A fresh key got a 200 right after unpause, so the block came from the counter and not from the outage

Also run on both trees at 4 workers: happy paths through raw curl, the OpenAI SDK (sync, and async streaming) and the Anthropic SDK, budget enforcement for key, team, end user, per-model key budget and /key/update raising a budget, malformed input, unauthenticated and unknown-key requests, /key/update during in-flight traffic, a probe after the counter TTL and batch flush, and one worker killed under load. No base vs tip difference outside the fixed behaviour. Two defects showed identically on both trees and are left for follow-up issues: messages: 1 on /v1/messages returns 500 instead of 400, and /key/generate with a team_id that does not exist returns 200

Related: #40233 proposes the same idea against the wrong base with a conflict. This PR was written independently from the issue and does not carry code from it

Type

🐛 Bug Fix

Caveats (if any)

Medium

  • Eviction policy is unchanged: past 10000 live scopes the same reseed can happen. Raising the size is what the issue asks for; an LRU rewrite is out of scope
  • Without Redis each proxy worker still keeps its own counter, so a worker that never saw the key reseeds from the DB exactly as before. Not changed by this PR; on both base and tip the 12 concurrent warm-up calls right after the first request show a few 200s until every worker has seen the key

Low

  • No config knob for the new size. Add one if a deployment needs more than 10000 scopes
  • Redis-backed spend counter path not exercised live (no Redis in the rig); the diff does not touch the Redis member of the cache

QA runbook

  • tests/integration/spend/test_spend_counter_scope_churn.py::test_over_budget_key_stays_blocked_while_hundreds_of_end_users_take_traffic - a key blocked for budget stays blocked after 250 other end users make requests. Runs in the accounting integration group on a local proxy, Postgres and the scripted upstream, with Redis removed from the owned proxy so the in-memory counter is the only thing keeping the key blocked while the DB row lags
    • POST /key/generate with {"max_budget": 0.06, "models": [<model>]} (probe key) and again with {"models": [<model>]} (churn key)
    • POST /v1/chat/completions with the probe key; expect 200 with 40 total tokens
    • Repeat with the probe key; expect 422 with "type": "budget_exceeded"
    • Send 250 POST /v1/chat/completions with the churn key, each with a different "user" value; expect more than 200 of them to return 200
    • POST /v1/chat/completions with the probe key once more; expect 422 budget_exceeded and no upstream request (before this PR: 200)
    • Sanity check: this test makes sense to add and is not hand-wavey (e.g., assert actual expected spend instead of just spend > 0) or potentially flaky

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

ran /live-pr-risk and found no regressions/backward incompatible risks

Link to Devin session: https://app.devin.ai/sessions/ae976f398c084972ab30c6bcc600c415
Open in Devin Desktop: https://app.devin.ai/desktop/session/ae976f398c084972ab30c6bcc600c415?variant=devin


Note

Medium Risk
Touches proxy budget enforcement and in-memory spend counters; wrong sizing or eviction could still admit over-budget traffic, though scope is limited to cache configuration plus tests.

Overview
Fixes budget enforcement bypass when many distinct spend scopes (keys, teams, end users) are active: spend_counter_cache no longer uses the default 200-entry in-memory layer, which could evict a just-blocked key’s counter and reseed it from lagging DB spend.

The proxy now wires spend_counter_cache to a dedicated InMemoryCache capped at SPEND_COUNTER_CACHE_MAX_SIZE (10,000) in constants.py / proxy_server.py; other caches are unchanged.

Regression coverage adds a unit test that floods increment_spend_counters with hundreds of end-user scopes then asserts key spend is still readable, plus an integration test (and contracts.json mapping) that an over-budget key stays 422 budget_exceeded after ~250 other end users take traffic with Redis disabled and delayed batch writes.

Reviewed by Cursor Bugbot for commit ec97ed6. Bugbot is set up for automated code reviews on this repo. Configure here.

The spend counter DualCache used the generic InMemoryCache default of 200
entries, so deployments with more than 200 active budget scopes evicted
still-valid counters and reseeded them from lagging DB spend on the next
request. A key the proxy had just rejected for being over budget could be
admitted again until the batch writer flushed.

Fixes #40221

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration
devin-ai-integration Bot requested a review from a team September 22, 2026 04:48
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔍 Devin Review: 2 flags

Not posted on this PR by your GitHub settings — view them in Devin Review. (Configure)

Devin Review

@greptile-apps

greptile-apps Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge; the dedicated cache addresses the reported eviction path and the replacement tests exercise the intended behavior without external-provider dependencies.

Summary

This PR gives spend counters a dedicated 10,000-entry in-memory cache so active budget counters are not displaced by ordinary 200-entry cache churn.

  • Introduces a centralized SPEND_COUNTER_CACHE_MAX_SIZE constant.
  • Configures spend_counter_cache with its own InMemoryCache.
  • Adds unit and owned-proxy integration coverage for preserving an over-budget key’s counter across hundreds of other scopes.
  • Replaces the unsuitable live E2E coverage with a deterministic local integration contract.

Reviews (4) · Last reviewed commit: "test(spend): move the spend-counter scop..."

Comment thread tests/test_litellm/proxy/test_proxy_server.py Outdated
Comment thread tests/e2e/quota_management/budgets/test_spend_counter_scope_churn_e2e.py Outdated

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

…path and on any e2e stack

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@chilkotiKartik

Copy link
Copy Markdown

Review for PR #42414: Give spend_counter_cache its own 10,000-entry in-memory cache

Cache Isolation & Memory Profiling

  1. LRU Eviction Isolation:
    Isolating spend counters into a dedicated 10k-entry DualCache instance prevents high-frequency model completion caching from evicting organization spend limits during traffic surges.

  2. TTL Synchronization:
    Ensure spend_counter_cache sets short local TTLs (e.g. 5-10s) to guarantee accurate budget limit enforcement across multi-instance proxy clusters.

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@codspeed

codspeed Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_spend-counter-cache-eviction (ec97ed6) with main (4bcdaf3)

Open in CodSpeed

@codecov

codecov Bot commented Sep 22, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

…lass

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

…ation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit ec97ed6. Configure here.

This branch was successfully deployed

1 active deployment
e2e-changed — ec97ed69 Deployed Sep 22, 2026 by devin-ai-integration[bot] via oauth #405
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Spend counter cache evicts active budget counters after 200 entries

2 participants