Repository navigation
fix(reset_budget_job): reset end users by budget link, not by user id - #40639
ryan-crabbe-berri merged 1 commit into
Conversation
The cascade zeroed end-user spend with a single update_many whose where clause enumerated every dependent user id. Prisma compiles that IN-list into one prepared statement carrying one bind variable per customer, and PostgreSQL caps a statement at 32,767 of them. Once a shared budget had more dependents than that the statement could not be parsed at all, so the atomic cascade rolled back, budget_reset_at never advanced, and the tier stayed due on every later tick forever. Customers sitting at their cap were blocked indefinitely with only a recurring log line to show for it. End users now match on budget_id like every other gated table, plus a NULL-budget_id branch for the implicitly created rows that carry no link and ride the default tier. The statement's bind count now tracks the number of expiring tiers rather than the customer population, so a reset costs the same whether a budget has ten dependents or a million. Fixes #40564 Claude-Session: https://claude.ai/code/session_01Hn5E8Jz1LjGLFyiYxBRcBW
Greptile SummaryThis PR resets end-user spend by expiring budget links instead of enumerating customer IDs, keeping SQL statement size independent of the number of customers sharing a budget
Confidence Score: 5/5This PR appears safe to merge, with no actionable regressions identified The changed filters address the customer-count bind limit while retaining expiring-tier eligibility, atomic window advancement, and rollover ordering
|
| Filename | Overview |
|---|---|
| litellm/proxy/common_utils/reset_budget_job.py | Replaces customer-ID filters with budget-link filters while preserving default-tier and rollover handling |
| tests/litellm_utils_tests/test_proxy_budget_reset.py | Updates end-user reset assertions to expect budget-link selection |
| tests/test_litellm/proxy/common_utils/test_reset_budget_job.py | Adds population-independent bind-count coverage and updates linked, default-tier, and rollover reset assertions |
Reviews (1): Last reviewed commit: "fix(reset_budget_job): reset end users b..." | Re-trigger Greptile
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
|
@ryan-crabbe-berri Thank you! This could allow to not have to load all As currently: async def _invalidate_budget_cascade_caches(self, cascade: _BudgetCascade) -> None:
for counter_key, new_spend in cascade.counter_resets:
await self._invalidate_spend_counter(counter_key, new_spend=new_spend)
for cache_key in cascade.cache_keys: # <- 1 str per user
await self._invalidate_user_api_key_cache_entry(cache_key)So If we have 1M Also if we collect only |
The budget-tier reset read every customer linked to an expiring tier into one result set before the write, then invalidated their caches one key at a time. Both of those scale with the customer count, so a large enough deployment can OOM the proxy pod on the read, and the tail of the population sits on a stale spend counter while the per-key invalidations drain PR BerriAI#40639 moved the reset write itself to a link-based UPDATE, so that pre-commit read no longer feeds the write. It only fed cache invalidation and the service-logging counts, which means it can move after the commit. This replaces it with a keyset walk over litellm_endusertable ordered by user_id, taking RESET_BUDGET_JOB_BATCH_SIZE rows per page, the same shape _reset_windows_for_source already uses, with no per-run page cap for the same reason that walk has none: the cursor cannot survive the run, so a cap would restart at the first customer on every tick and never reach the tail Each page's counter and cache keys now go out as one batched delete through a new DualCache.async_delete_cache_keys, which drops the in-memory entries and chunks the Redis DELETE at DEFAULT_MAX_REDIS_BATCH_CACHE_SIZE num_endusers_found and num_endusers_updated now report the customers whose caches were invalidated after the commit rather than the rows read before it, so both read 0 when the cascade write fails
…x_bp_40639_1101 chore(release): backport #40639 to stable/1.101.x
…BerriAI#40639) Adapted for stable/1.99.x: this line predates budget rollover (BerriAI#38514), so the fix is applied to _commit_budget_cascade_once directly. End users reset on the budget link plus a NULL budget_id branch for the default tier, which is what upstream's _queue_enduser_resets does with rollover off. The rollover test hunk is dropped. (cherry picked from commit 8a4fae0)
TLDR
Problem this solves:
How it solves it:
User Flow
Before: an operator with more than 32,700 customers sharing one budget finds that budget never resets, so every customer at the cap is blocked for good
POST http://localhost:4071/v1/chat/completionswith"user": "cust-00000000"comes back429withExceededBudget: End User=cust-00000000 over budget. Spend=5.0, Budget=0.01429too many bind variables in prepared statement, expected maximum of 32767, received 33001on every tick, and will keep doing so for as long as the population stays that sizeAfter: the same window expiry is uneventful and the customers go back to serving traffic
POSTcomes back with the same429200with a real completion0Relevant issues
Fixes #40564
Linear ticket
Resolves LIT-7535
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Screenshots / Proof of Fix
Shared setup, applied identically before each run. One shared budget already due for reset, and 33,000 customers on it, each over the $0.01 cap:
The gateway runs on port 4071 against a real Postgres and real OpenAI, with the reset job sped up to a 30 second tick so a reviewer does not wait ten minutes per window:
Before (ac66754)
The reset timestamp is still more than seven minutes in the past and not one customer was cleared.
After (760043b)
No bind-variable error, and no cascade failure line.
All 33,000 customers are cleared and the reset timestamp has moved into the future.
Type
🐛 Bug Fix
Caveats (if any)
Medium
Low
Final Attestation
https://claude.ai/code/session_01Hn5E8Jz1LjGLFyiYxBRcBW