Skip to content

chore(release): backport #40639 to stable/1.101.x - #42635

Merged
yuneng-berri merged 1 commit into
stable/1.101.xfrom
litellm_backport_stable_1_101_x_bp_40639_1101
Sep 23, 2026
Merged

yuneng-berri merged 1 commit into
stable/1.101.xfrom
litellm_backport_stable_1_101_x_bp_40639_1101

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Budget resets for large shared end-user tiers never complete on 1.101.x
  • The reset enumerated every customer id, past ~32,700 PostgreSQL refuses it
  • The cascade rolls back forever and capped customers stay blocked

How it solves it:

User Flow

Before: an operator with a shared budget tier linked to more than 32,767 customers sees the tier's reset never land

  1. They send POST https://litellm-domain/budget/new with {"budget_id": "shared-tier", "max_budget": 10, "budget_duration": "1d"} and get 200
  2. They create tens of thousands of customers with POST https://litellm-domain/customer/new pointing at "budget_id": "shared-tier", each returning 200
  3. Customers who spend up to 10 start getting 429 Budget has been exceeded on POST https://litellm-domain/v1/chat/completions
  4. A day later they call GET https://litellm-domain/customer/info?end_user_id=cust-1 and spend is still at the cap; the same 429 keeps coming back on every later day, with only a recurring reset error in the proxy log

After: the same tier resets on schedule no matter how many customers share it

  1. They send POST https://litellm-domain/budget/new with {"budget_id": "shared-tier", "max_budget": 10, "budget_duration": "1d"} and get 200
  2. They create tens of thousands of customers with POST https://litellm-domain/customer/new pointing at "budget_id": "shared-tier", each returning 200
  3. Customers who spend up to 10 start getting 429 Budget has been exceeded on POST https://litellm-domain/v1/chat/completions
  4. A day later GET https://litellm-domain/customer/info?end_user_id=cust-1 shows spend: 0 and POST https://litellm-domain/v1/chat/completions returns 200 again

Relevant issues

Backport of #40639 (fixes #40564) onto stable/1.101.x. One pick, cherry-picked with -x from 760043b, the PR's single commit under merge 8a4fae0 on main. Patch-id check against the source is VERBATIM, no adaptation. Every referenced symbol (_SPENT_ROWS_WHERE, the extra parameter on _queue_budget_linked_resets) already exists on the line

What is included:

Not included on purpose: #41488 (paged end-user cache invalidation after a reset) is a later perf follow-up on the same surface. It is not needed for this fix to work and was left out to keep the backport to the one requested change. The existing 1.101.x backport commits on this branch's base are untouched

No version bump. 1.101.1 has not shipped: no Docker Hub or GHCR image, no PyPI release, no tag, no release branch, so this pick rides the pending 1.101.1

Known noise on this line: ruff check on the touched test file reports one new UP006 (Dict vs dict) from the picked test helper, the same pattern as 52 pre-existing hits in that file and identical to what main carries. The line's ruff format --check is clean on the touched module

Affected release

Linear ticket

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Live proxy on localhost:4000 against a local PostgreSQL and real OpenAI calls (gpt-4.1-nano), same script run twice: Before at the merge base 4b432a9 (origin/stable/1.101.x), After at this PR's tip 173f71d. Config: master_key: sk-1234, proxy_budget_rescheduler_min_time: 10, proxy_budget_rescheduler_max_time: 15. The tier has 40,000 customers; one is created through the API, the other 39,999 are seeded with a single SQL insert because that many POST /customer/new calls is the only part a reviewer would not want to curl

$ curl -X POST localhost:4000/budget/new -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' \
  -d '{"budget_id":"shared-tier-replay","max_budget":10,"budget_duration":"30s"}'
{"budget_id":"shared-tier-replay","max_budget":10.0,...,"budget_duration":"30s","budget_reset_at":"2026-09-23T03:40:00Z",...}

$ curl -X POST localhost:4000/customer/new -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' \
  -d '{"user_id":"cust-replay-00000000","budget_id":"shared-tier-replay"}'
{"user_id":"cust-replay-00000000","blocked":false,"spend":0.0,"budget_id":"shared-tier-replay",...}

$ psql "$DATABASE_URL" -c "INSERT INTO \"LiteLLM_EndUserTable\" (user_id, spend, budget_id, blocked) SELECT 'cust-replay-' || lpad(g::text, 8, '0'), 10.0, 'shared-tier-replay', false FROM generate_series(1, 39999) g" \
  -c "UPDATE \"LiteLLM_EndUserTable\" SET spend = 10.0 WHERE user_id = 'cust-replay-00000000'" \
  -c "SELECT count(*) FROM \"LiteLLM_EndUserTable\" WHERE budget_id = 'shared-tier-replay' AND spend > 0"
INSERT 0 39999
UPDATE 1
 count
-------
 40000

$ curl -X POST localhost:4000/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' \
  -d '{"model":"gpt-4.1-nano","user":"cust-replay-00000000","messages":[{"role":"user","content":"say hi"}],"max_tokens":5}'
{"error":{"message":"Budget has been exceeded! EndUser=cust-replay-00000000 Current cost: 10.0000029, Max budget: 10.0","type":"budget_exceeded","param":null,"code":"429"}}
HTTP_STATUS:429

Before (merge base), 60 seconds later, after the 30s window expired and the reset job ran three times:

$ curl -X POST localhost:4000/v1/chat/completions ... (same request)
{"error":{"message":"Budget has been exceeded! EndUser=cust-replay-00000000 Current cost: 10.0000029, Max budget: 10.0","type":"budget_exceeded","param":null,"code":"429"}}
HTTP_STATUS:429

$ curl 'localhost:4000/customer/info?end_user_id=cust-replay-00000000' -H 'Authorization: Bearer sk-1234'
{"user_id":"cust-replay-00000000","blocked":false,"spend":10.0,"budget_id":"shared-tier-replay",...}

$ psql "$DATABASE_URL" -c "SELECT budget_reset_at FROM \"LiteLLM_BudgetTable\" WHERE budget_id='shared-tier-replay'" \
  -c "SELECT count(*) FROM \"LiteLLM_EndUserTable\" WHERE budget_id='shared-tier-replay' AND spend > 0"
   budget_reset_at
---------------------
 2026-09-23 03:40:00
 count
-------
 40000

$ grep -E 'Failed to reset|too many bind' before-proxy.log | head -4
03:40:07 - LiteLLM Proxy:ERROR: reset_budget_job.py:804 - Failed to reset the budget table cascade (team member, enduser, org, tag and model access group spend, plus budget_reset_at); nothing was committed and the budgets stay due for the next run: ...
prisma.errors.DataError: Assertion violation on the database: `too many bind variables in prepared statement, expected maximum of 32767, received 40001`
03:40:18 - LiteLLM Proxy:ERROR: reset_budget_job.py:804 - Failed to reset the budget table cascade ...
prisma.errors.DataError: Assertion violation on the database: `too many bind variables in prepared statement, expected maximum of 32767, received 40001`

After (this PR's tip), same 60 seconds later:

$ curl -X POST localhost:4000/v1/chat/completions ... (same request)
{"id":"chatcmpl-ER8ALNLrXYFrxFFSTj88elh0ndkkO","created":1790134953,"model":"gpt-4.1-nano","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Hi!","role":"assistant",...}}],...}
HTTP_STATUS:200

$ curl 'localhost:4000/customer/info?end_user_id=cust-replay-00000000' -H 'Authorization: Bearer sk-1234'
{"user_id":"cust-replay-00000000","blocked":false,"spend":0.0,"budget_id":"shared-tier-replay",...}

$ psql "$DATABASE_URL" -c "SELECT budget_reset_at FROM \"LiteLLM_BudgetTable\" WHERE budget_id='shared-tier-replay'" \
  -c "SELECT count(*) FROM \"LiteLLM_EndUserTable\" WHERE budget_id='shared-tier-replay' AND spend > 0"
   budget_reset_at
---------------------
 2026-09-23 03:42:30
 count
-------
     0

$ grep -E 'Failed to reset|too many bind' after-proxy.log | wc -l
0

Targeted test delta on the line: baseline at 4b432a9 147 passed across tests/test_litellm/proxy/common_utils/test_reset_budget_job.py and tests/litellm_utils_tests/test_proxy_budget_reset.py; post-pick 150 passed, 0 new failures. The three new tests are the pick's own regression tests for #40564, including the 40,000 customer case. The full tests/test_litellm suite is left to CI: the local venv is synced to main and 11 unrelated modules fail collection on this line either way

Gauntlet-mini over the final tree (five lenses: correctness, dependents, blind black-box, contract and backward compatibility, conventions and tests): SURVIVED, no confirmed findings. The black-box lens is the live replay above. Strongest rejected dissent: the new writes filter on spend > 0 where the old enumeration did not, so linked rows at zero spend are no longer touched; rejected because spend is Float @default(0.0) and non-null in schema.prisma, so zeroing a zero is a no-op and the observable state is identical. The NULL-budget_id branch is gated on the same max_end_user_budget_id in budget_ids condition as the existing NULL-row read in _collect_endusers_to_reset, so the two stay symmetric

Type

🐛 Bug Fix

Caveats (if any)

Low

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Link to Devin session: https://app.devin.ai/sessions/f446cd66c6ef417192eb2922ac02b43a
Open in Devin Desktop: https://app.devin.ai/desktop/session/f446cd66c6ef417192eb2922ac02b43a?variant=devin

The cascade zeroed end-user spend with a single update_many whose where
clause enumerated every dependent user id. Prisma compiles that IN-list
into one prepared statement carrying one bind variable per customer, and
PostgreSQL caps a statement at 32,767 of them. Once a shared budget had
more dependents than that the statement could not be parsed at all, so
the atomic cascade rolled back, budget_reset_at never advanced, and the
tier stayed due on every later tick forever. Customers sitting at their
cap were blocked indefinitely with only a recurring log line to show for
it.

End users now match on budget_id like every other gated table, plus a
NULL-budget_id branch for the implicitly created rows that carry no link
and ride the default tier. The statement's bind count now tracks the
number of expiring tiers rather than the customer population, so a reset
costs the same whether a budget has ten dependents or a million.

Fixes #40564

Claude-Session: https://claude.ai/code/session_01Hn5E8Jz1LjGLFyiYxBRcBW
(cherry picked from commit 760043b)
@devin-ai-integration
devin-ai-integration Bot requested a review from a team September 23, 2026 01:18
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@greptile-apps

greptile-apps Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

The reset implementation appears behaviorally sound, but explicit repository requirements for behavioral tests, strong typing, and concise comments must be satisfied before merging

Findings

  1. P2 Test Checks Mock Structure ▶
  2. P2 Helper Uses Coarse Type ▶
  3. P2 Docstring Is Too Detailed ▶

Summary

This backport changes end-user budget resets to update rows by their shared budget link, preventing statement bind counts from scaling with customer population

  • Explicitly linked end users now use the shared budget-reset helper
  • NULL-linked end users continue to follow the configured default tier
  • Tests cover linked predicates, rollover predicates, and large mocked populations
  • Repository test, typing, and source-comment requirements still need correction

Reviews (1) · Last reviewed commit: "fix(reset_budget_job): reset end users b..."

Comment on lines +592 to +596
assert [_bind_count(write["where"]) for write in writes] == [2], (
f"the cascade must not enumerate {population} user ids: past "
f"{_POSTGRES_MAX_BIND_VARIABLES} binds PostgreSQL refuses the statement, got {writes[:1]}"
)
assert _batch_writes(mock_prisma_client, "budget")[0]["data"]["budget_reset_at"] is not None

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Test Checks Mock Structure

This only inspects recorded predicates, violating the requirement to test behavior. Add execution-based coverage before merging

Context Used: CLAUDE.md (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

_POSTGRES_MAX_BIND_VARIABLES: Final = 32767


def _bind_count(where: Dict[str, Any]) -> int:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Helper Uses Coarse Type

The new parameter uses Dict[str, Any], violating the requirement for specific, fully typed parameters. This must be corrected before merging

Context Used: CLAUDE.md (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Comment on lines +227 to +235
"""End users reset on the budget link like every other gated table, plus a
NULL-budget_id branch: rows created implicitly persist no link and ride the
default tier (litellm.max_end_user_budget_id).

Matching on the link rather than enumerating user ids keeps a statement's
bind count proportional to the expiring tiers instead of the customer
population, which past ~32,700 dependents exceeds PostgreSQL's per-statement
bind ceiling and wedges the cascade permanently (#40564).
"""

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Docstring Is Too Detailed

This adds lengthy incident history, violating the requirement that necessary complex-logic comments remain concise. Shorten it before merging

Suggested change
"""End users reset on the budget link like every other gated table, plus a
NULL-budget_id branch: rows created implicitly persist no link and ride the
default tier (litellm.max_end_user_budget_id).
Matching on the link rather than enumerating user ids keeps a statement's
bind count proportional to the expiring tiers instead of the customer
population, which past ~32,700 dependents exceeds PostgreSQL's per-statement
bind ceiling and wedges the cascade permanently (#40564).
"""

Context Used: CLAUDE.md (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

@yuneng-berri yuneng-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified: patch-id matches upstream 760043b, CI green. Greptile P2s target upstream code, kept verbatim for the backport

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants