Repository navigation
chore(release): backport #39631, #39729, #40639 to stable/1.100.x and cut 1.100.3 - #43132
Conversation
|
|
|
| *(key for row in orgs for key in _org_cache_keys(row)), | ||
| *(key for row in tags for key in _tag_cache_keys(row)), | ||
| *(key for row in model_access_groups for key in _model_access_group_cache_keys(row)), | ||
| *(key for row in endusers for key in _enduser_cache_keys(row)), |
There was a problem hiding this comment.
Stale spend on other replicas
When one replica resets a tier, this change deletes each end-user object from that replica’s cache and Redis, but does not tell other replicas to clear their in-memory copies. Another replica can use the old spend as its budget-check floor and keep returning budget-exceeded errors after the database spend reaches zero. Broadcast these evictions as other management writes do.
Knowledge Base Used: Spend, budgets, and rate limits
| population, which past ~32,700 dependents exceeds PostgreSQL's per-statement | ||
| bind ceiling and wedges the cascade permanently (#40564). | ||
| """ | ||
| _queue_budget_linked_resets(writes, cascade, extra=_SPENT_ROWS_WHERE) |
There was a problem hiding this comment.
Reset misses counter invalidation
If an end user is created, moved onto the tier, or given spend after the job collects end users but before its transaction commits, this new budget-wide update can reset that user’s database spend. The job clears counters and caches only for users in the earlier collection, so the affected user can remain blocked by stale spend. The update and invalidation need to cover the same users.
Knowledge Base Used: Spend, budgets, and rate limits
| if model_name.startswith("gpt-6"): | ||
| return True |
There was a problem hiding this comment.
Hardcoded GPT-6 capability
This prefix check unconditionally treats GPT-6 names as eligible for the GPT-5.4-plus request path; the new reasoning-series marker also selects behavior by name. The repository requires model-specific capabilities to be declared in model_prices_and_context_window.json and read through get_model_info, rather than hardcoded in transformations. That requirement must be met before merging so capabilities can change without a code release.
Rule Used: What: Do not hardcode model-specific flags in the codebase. Instead, put them in model_prices_and_context_window.json and then read them in via get_model_info Why: Prevents need for users to upgrade litellm each time a new model supports this featu... (source)
Knowledge Base Used: Provider adapters and capabilities
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
| @@ -680,6 +696,7 @@ async def _collect_budget_cascade(self, budgets_to_reset: Sequence[LiteLLM_Budge | |||
| *(key for row in orgs for key in _org_cache_keys(row)), | |||
| *(key for row in tags for key in _tag_cache_keys(row)), | |||
| *(key for row in model_access_groups for key in _model_access_group_cache_keys(row)), | |||
| *(key for row in endusers for key in _enduser_cache_keys(row)), | |||
There was a problem hiding this comment.
Serial per-customer invalidation
For a tier with tens of thousands of end users, this adds a counter overwrite and a cache deletion for every customer. Both lists are then processed with sequential awaits after the database commit. The sweep can take minutes and delay later budget resets even though the database update is set-based. Batch or page the invalidation work.
Knowledge Base Used: Spend, budgets, and rate limits
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
| (_model_access_group_counter_key(row), _row_carried_spend(row, rollover_caps)) | ||
| for row in model_access_groups | ||
| ), | ||
| *((_enduser_counter_key(row), _enduser_carried_spend(row, rollover_caps)) for row in endusers), |
There was a problem hiding this comment.
Low: Concurrent end-user spend is overwritten
These snapshot-derived values are written to Redis sequentially after the database commit. An authenticated user can submit requests after the reset commits but before their end-user entry is processed; _invalidate_spend_counter() then replaces those live increments with zero or the stale rollover value. This is especially reachable for large shared tiers because every end user requires awaited cache operations. Reset the counter atomically at the transaction boundary, or use window-specific counter keys so post-reset increments cannot be overwritten by the sweep.
PR overviewThis release PR backports three changes to the stable/1.100.x branch and prepares version 1.100.3. The touched code includes updates to the scheduled budget-reset and spend-counter handling logic. One issue remains open in the budget reset flow: an authenticated user can submit requests during a reset window whose spend increments may then be overwritten by stale or zero values. This can undercount usage and weaken budget enforcement, particularly on large shared tiers where counter invalidation takes longer. No reported issues have yet been addressed in this PR. Open issues (1)
Fixed/addressed: 0 · PR risk: 5/10 |
TLDR
Problem this solves:
How it solves it:
User Flow
Before: a developer calling gpt-6 through the proxy gets errors, and customers on a big shared budget tier stay blocked forever
"model": "gpt-6-astra","max_tokens": 200and"reasoning_effort": "low"UnsupportedParamsError: openai does not support parameters: ['reasoning_effort']"model": "gpt-6-astra","temperature": 0.5and"drop_params": trueUnsupported parameter: 'temperature' is not supported with this modelBudget has been exceededon POST https://litellm-domain/v1/chat/completionsAfter: gpt-6 requests succeed, and the tier resets on schedule and the customer is served again without a restart
"model": "gpt-6-astra","max_tokens": 200and"reasoning_effort": "low""model": "gpt-6-astra","temperature": 0.5and"drop_params": true"status": "completed"Budget has been exceededon POST https://litellm-domain/v1/chat/completionsspend: 0, and the same request returns 200 from the same running proxyRelevant issues
Fixes #40564 and #39726 on stable/1.100.x
Backport of #39631 (merge 025a3ca), #39729 (daced81) and #40639 (merge 8a4fae0), all reachable from main and picked with
cherry-pick -x(-m 1for the two merges). This line has budget rollover, so all three apply with their code lines unchanged: #39631 needed an import-context resolution only, #39729 has one generator reflowed to this line's ruff format, and #40639 adds Final to the test module's typing import, which upstream's test file already had. Each commit message carries its noteNot included on purpose: #41488 (paged end-user cache invalidation after a reset). It is a later perf follow-up on the same surface, the fix works without it, and the stable/1.101.x backport (#42635) made the same call
Dependency bumps, each lock-only inside the existing pyproject range and at or below what main resolves: anyio 4.14.2, gitpython 3.1.60, tornado 6.5.8, pypdf 6.16.1 and soupsieve 2.9. After the bumps grype and OSV report the same residual set as main, with no fixed version available for either
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Live proxy against a local PostgreSQL and real OpenAI calls to gpt-6-astra, same scripts run at the merge base and at this PR's tip. Config:
The budget tier is
shared-tier(max_budget: 10,budget_duration: 1d) with 33,000 customerscust-00000000tocust-00032999seeded by one SQL insert, since 33,000POST /customer/newcalls is not something a reviewer would want to replay. Before each run every customer's spend is set to 10 and the window is set an hour ahead; the replay then movesbudget_reset_atinto the past and waits 40 seconds, which covers several reset job runsBefore (9c1216a)
gpt-6 on /v1/chat/completions
gpt-6 on /v1/responses
33,000 customers on one budget tier
After (04fcee8)
gpt-6 on /v1/chat/completions
gpt-6 on /v1/responses
33,000 customers on one budget tier
Type
Bug Fix
Caveats (if any)