Repository navigation
chore(release): backport #39631, #39729, #40639 to stable/1.98.x - #43130
Conversation
…e configs (#39631) Adapted for stable/1.98.x: requires_max_completion_tokens and its test are not on this line (added upstream by #36857), so that hunk is dropped. The Responses parametrize cases target a test class this line does not have, so the two gpt-6-astra cases are kept as a standalone test in the same file. (cherry picked from commit 025a3ca)
…et (#39726) Adapted for stable/1.98.x: this line predates budget rollover (#38514), so each end user's spend counter is zeroed through the existing counter_keys list instead of counter_resets, and its cache entry is evicted via an inline "end_user_id:" key, matching how this line writes it in auth_checks. Signed-off-by: amasen02 <amasen02@users.noreply.github.com> (cherry picked from commit daced81)
…#40639) Adapted for stable/1.98.x: this line predates budget rollover (#38514), so the fix is applied to _commit_budget_cascade_once directly. End users reset on the budget link plus a NULL budget_id branch for the default tier, which is what upstream's _queue_enduser_resets does with rollover off. The rollover test hunk is dropped. (cherry picked from commit 8a4fae0)
|
|
|
| uow.tags.queue_spend_zero(where=_budget_link_where(cascade.budget_ids, _SPENT_ROWS_WHERE)) | ||
| if enduser_ids: | ||
| uow.endusers.queue_spend_zero(where={"user_id": {"in": list(enduser_ids)}}) | ||
| uow.endusers.queue_spend_zero(where=_budget_link_where(cascade.budget_ids, _SPENT_ROWS_WHERE)) |
There was a problem hiding this comment.
Newly linked users stay blocked If an end user joins a due tier after users are collected, this update clears their database spend but misses their counter. Auth still reads that counter, so the user can remain blocked after the reset
Knowledge Base Used: Spend, budgets, and rate limits
| *(_key_counter_key(row) for row in keys), | ||
| *(_org_counter_key(row) for row in orgs), | ||
| *(_tag_counter_key(row) for row in tags), | ||
| *(_enduser_counter_key(row) for row in endusers), |
There was a problem hiding this comment.
New spend can be erased If a request is charged while a large tier’s counters are being cleared after commit, the later reset can overwrite that charge with zero. Subsequent budget checks then undercount the new window’s spend
Knowledge Base Used: Spend, budgets, and rate limits
| *(_enduser_counter_key(row) for row in endusers), | ||
| ), | ||
| cache_keys=( | ||
| *(key for row in team_memberships for key in _team_membership_cache_keys(row)), | ||
| *(key for row in keys for key in _key_cache_keys(row)), | ||
| *(key for row in orgs for key in _org_cache_keys(row)), | ||
| *(key for row in tags for key in _tag_cache_keys(row)), | ||
| *(key for row in endusers for key in _enduser_cache_keys(row)), |
There was a problem hiding this comment.
Large tiers delay resets The job now writes a counter and deletes a cache entry sequentially for every end user. On a tier with tens of thousands of users, this sweep can delay later budget resets; please batch or bound it
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
| GPT_REASONING_SERIES_MARKERS: Final = ("gpt-5", "gpt-6") | ||
|
|
||
|
|
||
| def is_gpt_reasoning_series_name(model: str) -> bool: | ||
| normalized: Final = model.split("/")[-1] | ||
| return any(marker in model for marker in GPT_REASONING_SERIES_MARKERS) and not normalized.startswith("gpt-5-chat") |
There was a problem hiding this comment.
GPT-6 capabilities are hardcoded The new marker check and GPT-6 prefix check classify model behavior in code. Repository rules require model-specific flags in model metadata, read through get_model_info. This requirement must be met before merging
Rule Used: What: Do not hardcode model-specific flags in the codebase. Instead, put them in model_prices_and_context_window.json and then read them in via get_model_info Why: Prevents need for users to upgrade litellm each time a new model supports this featu... (source)
Knowledge Base Used: Provider adapters and capabilities
TLDR
Problem this solves:
How it solves it:
User Flow
Before: a developer calling gpt-6 through the proxy gets errors, and customers on a big shared budget tier stay blocked forever
"model": "gpt-6-astra","max_tokens": 200and"reasoning_effort": "low"UnsupportedParamsError: openai does not support parameters: ['reasoning_effort']"model": "gpt-6-astra","temperature": 0.5and"drop_params": trueUnsupported parameter: 'temperature' is not supported with this modelBudget has been exceededon POST https://litellm-domain/v1/chat/completionsAfter: gpt-6 requests succeed, and the tier resets on schedule and the customer is served again without a restart
"model": "gpt-6-astra","max_tokens": 200and"reasoning_effort": "low""model": "gpt-6-astra","temperature": 0.5and"drop_params": true"status": "completed"Budget has been exceededon POST https://litellm-domain/v1/chat/completionsspend: 0, and the same request returns 200 from the same running proxyRelevant issues
Fixes #40564 and #39726 on stable/1.98.x
Backport of #39631 (merge 025a3ca), #39729 (daced81) and #40639 (merge 8a4fae0), all reachable from main and picked with
cherry-pick -x(-m 1for the two merges). On this line requires_max_completion_tokens does not exist (it came with #36857), so that hunk of #39631 is dropped; the two gpt-6-astra Responses cases become a standalone test because their upstream test class is not here. The line predates budget rollover, so #39729 zeroes each end user's counter through the existing counter_keys list and evicts the inline end_user_id cache key this line writes, and #40639 is applied to _commit_budget_cascade_once as a budget-link reset plus a NULL budget_id branch for the default tier, which is what upstream does with rollover off. Each commit message carries its adaptation noteNot included on purpose: #41488 (paged end-user cache invalidation after a reset). It is a later perf follow-up on the same surface, the fix works without it, and the stable/1.101.x backport (#42635) made the same call
Dependency bumps, each lock-only inside the existing pyproject range and at or below what main resolves: anyio 4.14.2, gitpython 3.1.60, tornado 6.5.8, restrictedpython 8.4, sqlparse 0.6.0, pypdf 6.16.1 and soupsieve 2.9. After the bumps grype and OSV report the same residual set as main, with no fixed version available for either
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Live proxy against a local PostgreSQL and real OpenAI calls to gpt-6-astra, same scripts run at the merge base and at this PR's tip. Config:
The budget tier is
shared-tier(max_budget: 10,budget_duration: 1d) with 33,000 customerscust-00000000tocust-00032999seeded by one SQL insert, since 33,000POST /customer/newcalls is not something a reviewer would want to replay. Before each run every customer's spend is set to 10 and the window is set an hour ahead; the replay then movesbudget_reset_atinto the past and waits 40 seconds, which covers several reset job runsBefore (86f64b0)
gpt-6 on /v1/chat/completions
gpt-6 on /v1/responses
33,000 customers on one budget tier
After (e70af49)
gpt-6 on /v1/chat/completions
gpt-6 on /v1/responses
33,000 customers on one budget tier
Type
Bug Fix
Caveats (if any)