Skip to content

fix(reset_budget_job): reset end users by budget link, not by user id - #40639

Merged
ryan-crabbe-berri merged 1 commit into
litellm_internal_stagingfrom
litellm_enduser_budget_reset_bind_limit
Sep 11, 2026
Merged

ryan-crabbe-berri merged 1 commit into
litellm_internal_stagingfrom
litellm_enduser_budget_reset_bind_limit

Conversation

@ryan-crabbe-berri

@ryan-crabbe-berri ryan-crabbe-berri commented Sep 11, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • A shared budget past ~32,700 customers never resets
  • Those customers stay blocked forever with no signal
  • Every reset attempt rolls the whole thing back

How it solves it:

  • Reset customers by their budget link, not by id
  • Statement size now tracks budgets, not customer count

User Flow

Before: an operator with more than 32,700 customers sharing one budget finds that budget never resets, so every customer at the cap is blocked for good

  1. The operator puts 33,000 customers on one shared budget capped at $0.01 with a 60 second window
  2. Those customers spend past the cap, so POST http://localhost:4071/v1/chat/completions with "user": "cust-00000000" comes back 429 with ExceededBudget: End User=cust-00000000 over budget. Spend=5.0, Budget=0.01
  3. The window expires and several minutes pass
  4. The customer sends the exact same request and gets the exact same 429
  5. The operator checks the shared budget and finds its reset timestamp still sitting minutes in the past, with every customer's spend untouched
  6. The gateway log repeats too many bind variables in prepared statement, expected maximum of 32767, received 33001 on every tick, and will keep doing so for as long as the population stays that size

After: the same window expiry is uneventful and the customers go back to serving traffic

  1. The operator has the same 33,000 customers on the same shared budget
  2. Those customers spend past the cap, so the same POST comes back with the same 429
  3. The window expires and the next tick resets it
  4. The customer sends the exact same request and gets 200 with a real completion
  5. The operator checks the shared budget and finds its reset timestamp moved into the future, with every customer's spend back to 0
  6. No bind-variable errors in the log

Relevant issues

Fixes #40564

Linear ticket

Resolves LIT-7535

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Screenshots / Proof of Fix

Shared setup, applied identically before each run. One shared budget already due for reset, and 33,000 customers on it, each over the $0.01 cap:

INSERT INTO "LiteLLM_BudgetTable"
    (budget_id, max_budget, budget_duration, budget_reset_at, created_by, updated_by)
VALUES ('repro-40564-shared', 0.01, '60s', now() - interval '5 minutes', 'repro', 'repro')
ON CONFLICT (budget_id) DO UPDATE SET
    max_budget = EXCLUDED.max_budget,
    budget_duration = EXCLUDED.budget_duration,
    budget_reset_at = EXCLUDED.budget_reset_at;

INSERT INTO "LiteLLM_EndUserTable" (user_id, spend, budget_id, blocked)
SELECT 'cust-' || lpad(i::text, 8, '0'), 5.0, 'repro-40564-shared', false
FROM generate_series(0, 32999) AS series(i)
ON CONFLICT (user_id) DO UPDATE SET spend = EXCLUDED.spend, budget_id = EXCLUDED.budget_id;

The gateway runs on port 4071 against a real Postgres and real OpenAI, with the reset job sped up to a 30 second tick so a reviewer does not wait ten minutes per window:

PROXY_BUDGET_RESCHEDULER_MIN_TIME=30 PROXY_BUDGET_RESCHEDULER_MAX_TIME=35 \
  python litellm/proxy/proxy_cli.py --config litellm/proxy/dev_config.yaml \
  --port 4071 --detailed_debug --use_v2_migration_resolver 2>&1 | tee run.log

Before (ac66754)

  1. The customer calls the gateway while over the shared cap
curl -s -w '\nHTTP %{http_code}\n' http://localhost:4071/v1/chat/completions \
  -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' \
  -d '{"model":"gpt-5.5","user":"cust-00000000","messages":[{"role":"user","content":"say hi in 3 words"}]}'
{"error":{"message":"ExceededBudget: End User=cust-00000000 over budget. Spend=5.0, Budget=0.01","type":"budget_exceeded","param":null,"code":"429"}}
HTTP 429
  1. Wait a little over two minutes, so the expired window gets four reset attempts, then read the log
grep -c 'too many bind variables' run.log
grep -m1 -o 'too many bind variables in prepared statement, expected maximum of [0-9]*, received [0-9]*' run.log
grep -m1 -o 'Failed to reset the budget table cascade.*next run' run.log
4
too many bind variables in prepared statement, expected maximum of 32767, received 33001
Failed to reset the budget table cascade (team member, enduser, org, tag and model access group spend, plus budget_reset_at); nothing was committed and the budgets stay due for the next run
  1. Check the shared budget and the customers' spend
psql "$DATABASE_URL" -c "SELECT
  (SELECT count(*) FROM \"LiteLLM_EndUserTable\" WHERE budget_id = 'repro-40564-shared' AND spend > 0) AS customers_still_over_cap,
  (SELECT budget_reset_at FROM \"LiteLLM_BudgetTable\" WHERE budget_id = 'repro-40564-shared') AS budget_reset_at,
  now() AS now;"
 customers_still_over_cap |     budget_reset_at     |              now
--------------------------+-------------------------+-------------------------------
                    33000 | 2026-09-11 00:09:47.261 | 2026-09-11 00:17:03.775022+00

The reset timestamp is still more than seven minutes in the past and not one customer was cleared.

  1. The customer retries the identical call
{"error":{"message":"ExceededBudget: End User=cust-00000000 over budget. Spend=5.0, Budget=0.01","type":"budget_exceeded","param":null,"code":"429"}}
HTTP 429

After (760043b)

  1. The customer calls the gateway while over the shared cap, same command as before
{"error":{"message":"ExceededBudget: End User=cust-00000000 over budget. Spend=5.0, Budget=0.01","type":"budget_exceeded","param":null,"code":"429"}}
HTTP 429
  1. Wait for the expired window to be reset, then read the log with the same greps
0

No bind-variable error, and no cascade failure line.

  1. Check the shared budget and the customers' spend with the same query
 customers_still_over_cap |   budget_reset_at   |              now
--------------------------+---------------------+-------------------------------
                        0 | 2026-09-11 00:19:00 | 2026-09-11 00:18:05.934384+00

All 33,000 customers are cleared and the reset timestamp has moved into the future.

  1. The customer retries the identical call and gets a real completion
{"id":"chatcmpl-EMjFuRSP9xflEVOD1TQnEsIZA0xlk","created":1789085886,"model":"gpt-5.5","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Hi there, friend","role":"assistant","provider_specific_fields":{"refusal":null},"annotations":[]},"provider_specific_fields":{}}],"usage":{"completion_tokens":30,"prompt_tokens":12,"total_tokens":42,"completion_tokens_details":{"accepted_prediction_tokens":0,"audio_tokens":0,"reasoning_tokens":17,"rejected_prediction_tokens":0},"prompt_tokens_details":{"audio_tokens":0,"cached_tokens":0}},"service_tier":"default"}
HTTP 200

Type

🐛 Bug Fix

Caveats (if any)

Medium

  • The reset still reads every customer row on the tier each tick, and still invalidates their cached spend one at a time. That is memory and wall clock at 100k customers, not a failure, and it predates this PR. Worth its own issue to bound the read and batch the invalidation

Low

  • Customers already sitting at zero spend are no longer rewritten on reset, since the filter now skips them the same way the other gated tables already do

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

https://claude.ai/code/session_01Hn5E8Jz1LjGLFyiYxBRcBW

The cascade zeroed end-user spend with a single update_many whose where
clause enumerated every dependent user id. Prisma compiles that IN-list
into one prepared statement carrying one bind variable per customer, and
PostgreSQL caps a statement at 32,767 of them. Once a shared budget had
more dependents than that the statement could not be parsed at all, so
the atomic cascade rolled back, budget_reset_at never advanced, and the
tier stayed due on every later tick forever. Customers sitting at their
cap were blocked indefinitely with only a recurring log line to show for
it.

End users now match on budget_id like every other gated table, plus a
NULL-budget_id branch for the implicitly created rows that carry no link
and ride the default tier. The statement's bind count now tracks the
number of expiring tiers rather than the customer population, so a reset
costs the same whether a budget has ten dependents or a million.

Fixes #40564

Claude-Session: https://claude.ai/code/session_01Hn5E8Jz1LjGLFyiYxBRcBW
@ryan-crabbe-berri
ryan-crabbe-berri requested a review from a team September 11, 2026 00:19
@codspeed

codspeed Bot commented Sep 11, 2026

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_enduser_budget_reset_bind_limit (760043b) with litellm_internal_staging (960fc4b)

Open in CodSpeed

@greptile-apps

greptile-apps Bot commented Sep 11, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR resets end-user spend by expiring budget links instead of enumerating customer IDs, keeping SQL statement size independent of the number of customers sharing a budget

  • Reuses the existing linked-budget reset helper and preserves rollover ordering
  • Handles customers with no explicit budget link when the default tier expires
  • Updates database-boundary assertions and adds a 40,000-customer bind-count regression test

Confidence Score: 5/5

This PR appears safe to merge, with no actionable regressions identified

The changed filters address the customer-count bind limit while retaining expiring-tier eligibility, atomic window advancement, and rollover ordering

Important Files Changed

Filename Overview
litellm/proxy/common_utils/reset_budget_job.py Replaces customer-ID filters with budget-link filters while preserving default-tier and rollover handling
tests/litellm_utils_tests/test_proxy_budget_reset.py Updates end-user reset assertions to expect budget-link selection
tests/test_litellm/proxy/common_utils/test_reset_budget_job.py Adds population-independent bind-count coverage and updates linked, default-tier, and rollover reset assertions

Reviews (1): Last reviewed commit: "fix(reset_budget_job): reset end users b..." | Re-trigger Greptile

@codecov

codecov Bot commented Sep 11, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@ryan-crabbe-berri
ryan-crabbe-berri merged commit 8a4fae0 into litellm_internal_staging Sep 11, 2026
84 checks passed
@ryan-crabbe-berri
ryan-crabbe-berri deleted the litellm_enduser_budget_reset_bind_limit branch September 11, 2026 00:31
@ltoniazzi

ltoniazzi commented Sep 11, 2026 •

Copy link
Copy Markdown

@ryan-crabbe-berri Thank you!
I was also wondering if something like this can be done to invalidate the cache for end-users.

This could allow to not have to load all end-users in-memory in the FastApi job

As currently:

async def _invalidate_budget_cascade_caches(self, cascade: _BudgetCascade) -> None:
        for counter_key, new_spend in cascade.counter_resets:
            await self._invalidate_spend_counter(counter_key, new_spend=new_spend)
        for cache_key in cascade.cache_keys:  # <- 1 str per user
            await self._invalidate_user_api_key_cache_entry(cache_key)

So If we have 1M end-users this loop starts to impact the cpu of the pod, right?

Also if we collect only end-users with spend >0 we can already get a bit more scalability (I think end-users with budget assigned get collected even if spend is already 0)

pull Bot pushed a commit to stnxo2023/litellm that referenced this pull request Sep 17, 2026
The budget-tier reset read every customer linked to an expiring tier into
one result set before the write, then invalidated their caches one key at
a time. Both of those scale with the customer count, so a large enough
deployment can OOM the proxy pod on the read, and the tail of the
population sits on a stale spend counter while the per-key invalidations
drain

PR BerriAI#40639 moved the reset write itself to a link-based UPDATE, so that
pre-commit read no longer feeds the write. It only fed cache invalidation
and the service-logging counts, which means it can move after the commit.
This replaces it with a keyset walk over litellm_endusertable ordered by
user_id, taking RESET_BUDGET_JOB_BATCH_SIZE rows per page, the same shape
_reset_windows_for_source already uses, with no per-run page cap for the
same reason that walk has none: the cursor cannot survive the run, so a
cap would restart at the first customer on every tick and never reach the
tail

Each page's counter and cache keys now go out as one batched delete
through a new DualCache.async_delete_cache_keys, which drops the
in-memory entries and chunks the Redis DELETE at
DEFAULT_MAX_REDIS_BATCH_CACHE_SIZE

num_endusers_found and num_endusers_updated now report the customers
whose caches were invalidated after the commit rather than the rows read
before it, so both read 0 when the cascade write fails
yuneng-berri added a commit that referenced this pull request Sep 23, 2026
…x_bp_40639_1101

chore(release): backport #40639 to stable/1.101.x
yuneng-berri added a commit that referenced this pull request Sep 25, 2026
…udget_0924

chore(release): backport #39631, #39729, #40639 to stable/1.98.x
yuneng-berri added a commit that referenced this pull request Sep 25, 2026
…udget_0924

chore(release): backport #39631, #39729, #40639 to stable/1.99.x and cut 1.99.4
yuneng-berri added a commit that referenced this pull request Sep 25, 2026
…budget_0924

chore(release): backport #39631, #39729, #40639 to stable/1.100.x and cut 1.100.3
achraf-mer pushed a commit to achraf-mer/litellm that referenced this pull request Sep 30, 2026
…BerriAI#40639)

Adapted for stable/1.99.x: this line predates budget rollover (BerriAI#38514), so the fix is applied to _commit_budget_cascade_once directly. End users reset on the budget link plus a NULL budget_id branch for the default tier, which is what upstream's _queue_enduser_resets does with rollover off. The rollover test hunk is dropped.

(cherry picked from commit 8a4fae0)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: End-user budget reset exceeds PostgreSQL's 32,767 bind-variable limit and never completes

3 participants