Skip to content

chore(release): backport #39631, #39729, #40639 to stable/1.98.x - #43130

Merged
yuneng-berri merged 10 commits into
stable/1.98.xfrom
litellm_backport_1_98_x_gpt6_budget_0924
Sep 25, 2026
Merged

yuneng-berri merged 10 commits into
stable/1.98.xfrom
litellm_backport_1_98_x_gpt6_budget_0924

Conversation

@yuneng-berri

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • gpt-6 models fail on chat and Responses with reasoning params
  • Budget tiers with over 32,767 customers never reset
  • Customers stay blocked on a running proxy even after a reset

How it solves it:

User Flow

Before: a developer calling gpt-6 through the proxy gets errors, and customers on a big shared budget tier stay blocked forever

  1. They send POST https://litellm-domain/v1/chat/completions with "model": "gpt-6-astra", "max_tokens": 200 and "reasoning_effort": "low"
  2. They get 400 UnsupportedParamsError: openai does not support parameters: ['reasoning_effort']
  3. They send POST https://litellm-domain/v1/responses with "model": "gpt-6-astra", "temperature": 0.5 and "drop_params": true
  4. They get 400 Unsupported parameter: 'temperature' is not supported with this model
  5. A customer on a 33,000-customer budget tier hits the cap and gets 429 Budget has been exceeded on POST https://litellm-domain/v1/chat/completions
  6. The tier's window passes and GET https://litellm-domain/customer/info still shows the capped spend; the customer keeps getting 429

After: gpt-6 requests succeed, and the tier resets on schedule and the customer is served again without a restart

  1. They send POST https://litellm-domain/v1/chat/completions with "model": "gpt-6-astra", "max_tokens": 200 and "reasoning_effort": "low"
  2. They get 200 with the model's reply
  3. They send POST https://litellm-domain/v1/responses with "model": "gpt-6-astra", "temperature": 0.5 and "drop_params": true
  4. They get 200 with "status": "completed"
  5. A customer on a 33,000-customer budget tier hits the cap and gets 429 Budget has been exceeded on POST https://litellm-domain/v1/chat/completions
  6. The tier's window passes, GET https://litellm-domain/customer/info shows spend: 0, and the same request returns 200 from the same running proxy

Relevant issues

Fixes #40564 and #39726 on stable/1.98.x

Backport of #39631 (merge 025a3ca), #39729 (daced81) and #40639 (merge 8a4fae0), all reachable from main and picked with cherry-pick -x (-m 1 for the two merges). On this line requires_max_completion_tokens does not exist (it came with #36857), so that hunk of #39631 is dropped; the two gpt-6-astra Responses cases become a standalone test because their upstream test class is not here. The line predates budget rollover, so #39729 zeroes each end user's counter through the existing counter_keys list and evicts the inline end_user_id cache key this line writes, and #40639 is applied to _commit_budget_cascade_once as a budget-link reset plus a NULL budget_id branch for the default tier, which is what upstream does with rollover off. Each commit message carries its adaptation note

Not included on purpose: #41488 (paged end-user cache invalidation after a reset). It is a later perf follow-up on the same surface, the fix works without it, and the stable/1.101.x backport (#42635) made the same call

Dependency bumps, each lock-only inside the existing pyproject range and at or below what main resolves: anyio 4.14.2, gitpython 3.1.60, tornado 6.5.8, restrictedpython 8.4, sqlparse 0.6.0, pypdf 6.16.1 and soupsieve 2.9. After the bumps grype and OSV report the same residual set as main, with no fixed version available for either

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Live proxy against a local PostgreSQL and real OpenAI calls to gpt-6-astra, same scripts run at the merge base and at this PR's tip. Config:

model_list:
  - model_name: gpt-6-astra
    litellm_params:
      model: openai/gpt-6-astra
      api_key: os.environ/OPENAI_API_KEY
general_settings:
  master_key: sk-...
  database_url: postgresql://...
  proxy_budget_rescheduler_min_time: 5
  proxy_budget_rescheduler_max_time: 6

The budget tier is shared-tier (max_budget: 10, budget_duration: 1d) with 33,000 customers cust-00000000 to cust-00032999 seeded by one SQL insert, since 33,000 POST /customer/new calls is not something a reviewer would want to replay. Before each run every customer's spend is set to 10 and the window is set an hour ahead; the replay then moves budget_reset_at into the past and waits 40 seconds, which covers several reset job runs

Before (86f64b0)

gpt-6 on /v1/chat/completions

  1. Run the request
$ curl -s localhost:4000/v1/chat/completions -H "Authorization: Bearer sk-..." -d '{"model":"gpt-6-astra","max_tokens":200,"reasoning_effort":"low","messages":[{"role":"user","content":"Reply with the word ok"}]}'
  1. Observe the response
{"message": "litellm.UnsupportedParamsError: openai does not support parameters: ['reasoning_effort'], for model=gpt-6-astra. To drop these, set `litellm.drop_params=True` or for proxy:\n\n`litellm_settings:\n drop_params: true`\n. \n If you want to use these 

gpt-6 on /v1/responses

  1. Run the request
$ curl -s localhost:4000/v1/responses -H "Authorization: Bearer sk-..." -d '{"model":"gpt-6-astra","temperature":0.5,"drop_params":true,"input":"Reply with the word ok"}'
  1. Observe the response
{"message": "litellm.BadRequestError: OpenAIException - {\n  \"error\": {\n    \"message\": \"Unsupported parameter: 'temperature' is not supported with this model.\",\n    \"type\": \"invalid_request_error\",\n    \"param\": \"temperature\",\n    \"code\": nu

33,000 customers on one budget tier

  1. Run the replay (steps and output below come straight from the script)
1. customers on shared-tier at the 10 cap: 33000
2. $ curl -s localhost:4000/v1/chat/completions -H "Authorization: Bearer sk-..." -d '{"model":"gpt-6-astra","user":"cust-00000001","messages":[{"role":"user","content":"Reply with the word ok"}]}'
    {"code": "429", "message": "Budget has been exceeded! EndUser=cust-00000001 Current cost: 10.81932, Max budget: 10.0"}
3. shared-tier window expires (budget_reset_at moved into the past), wait 40s for the reset job
4. proxy log: 'too many bind variables' x12, 'Failed to reset the budget table cascade' x6
   too many bind variables in prepared statement, expected maximum of 32767, received 33001
5. customers still at cap: 33000, zeroed: 0, window still due: t
6. $ curl -s localhost:4000/v1/chat/completions (same request as step 2)
    {"code": "429", "message": "Budget has been exceeded! EndUser=cust-00000001 Current cost: 10.81932, Max budget: 10.0"}

After (e70af49)

gpt-6 on /v1/chat/completions

  1. Run the request
$ curl -s localhost:4000/v1/chat/completions -H "Authorization: Bearer sk-..." -d '{"model":"gpt-6-astra","max_tokens":200,"reasoning_effort":"low","messages":[{"role":"user","content":"Reply with the word ok"}]}'
  1. Observe the response
{"model": "gpt-6-astra", "content": "ok", "usage": 4}

gpt-6 on /v1/responses

  1. Run the request
$ curl -s localhost:4000/v1/responses -H "Authorization: Bearer sk-..." -d '{"model":"gpt-6-astra","temperature":0.5,"drop_params":true,"input":"Reply with the word ok"}'
  1. Observe the response
{"status": "completed", "model": "gpt-6-astra", "temperature": 1.0, "output_text": ["ok"]}

33,000 customers on one budget tier

  1. Run the replay (steps and output below come straight from the script)
1. customers on shared-tier at the 10 cap: 33000
2. $ curl -s localhost:4000/v1/chat/completions -H "Authorization: Bearer sk-..." -d '{"model":"gpt-6-astra","user":"cust-00000001","messages":[{"role":"user","content":"Reply with the word ok"}]}'
    {"code": "429", "message": "Budget has been exceeded! EndUser=cust-00000001 Current cost: 10.81932, Max budget: 10.0"}
3. shared-tier window expires (budget_reset_at moved into the past), wait 40s for the reset job
4. proxy log: 'too many bind variables' x0, 'Failed to reset the budget table cascade' x0
5. customers still at cap: 0, zeroed: 33000, window still due: f
6. $ curl -s localhost:4000/v1/chat/completions (same request as step 2)
    {"model": "gpt-6-astra", "content": "ok"}

Type

Bug Fix

Caveats (if any)

mateo-berri and others added 10 commits September 24, 2026 20:31
…e configs (#39631)

Adapted for stable/1.98.x: requires_max_completion_tokens and its test are not on this line (added upstream by #36857), so that hunk is dropped. The Responses parametrize cases target a test class this line does not have, so the two gpt-6-astra cases are kept as a standalone test in the same file.

(cherry picked from commit 025a3ca)
…et (#39726)

Adapted for stable/1.98.x: this line predates budget rollover (#38514), so each end user's spend counter is zeroed through the existing counter_keys list instead of counter_resets, and its cache entry is evicted via an inline "end_user_id:" key, matching how this line writes it in auth_checks.

Signed-off-by: amasen02 <amasen02@users.noreply.github.com>
(cherry picked from commit daced81)
…#40639)

Adapted for stable/1.98.x: this line predates budget rollover (#38514), so the fix is applied to _commit_budget_cascade_once directly. End users reset on the budget link plus a NULL budget_id branch for the default tier, which is what upstream's _queue_enduser_resets does with rollover off. The rollover test hunk is dropped.

(cherry picked from commit 8a4fae0)
@yuneng-berri
yuneng-berri requested a review from a team September 25, 2026 04:18
@CLAassistant

CLAassistant commented Sep 25, 2026 •

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you all sign our Contributor License Agreement before we can accept your contribution.
3 out of 4 committers have signed the CLA.

✅ mateo-berri
✅ ryan-crabbe-berri
✅ yuneng-berri
❌ amasen02
You have signed the CLA already but the status is still pending? Let us recheck it.

@greptile-apps

greptile-apps Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 3/5

[Medium risk] Backport of model classification and budget reset logic changes.

The PR is not safe to merge until end-user reset invalidation handles concurrent spend and the model-capability instruction is satisfied

Findings

  1. P1 Newly linked users stay blocked ▶
  2. P1 New spend can be erased ▶
  3. P2 Large tiers delay resets ▶
  4. P2 GPT-6 capabilities are hardcoded ▶

Summary

This backport adds GPT-6 reasoning handling to OpenAI and Azure chat and Responses paths, changes budget-tier resets to update end users by budget link, and refreshes seven locked dependencies

  • The budget reset also adds end-user counter and cache invalidation
  • Collection and post-commit invalidation can disagree with concurrent spend, and the new per-user sweep has an unbounded cost

Reviews (1) · Last reviewed commit: "chore(deps): bump soupsieve to 2.9"

uow.tags.queue_spend_zero(where=_budget_link_where(cascade.budget_ids, _SPENT_ROWS_WHERE))
if enduser_ids:
uow.endusers.queue_spend_zero(where={"user_id": {"in": list(enduser_ids)}})
uow.endusers.queue_spend_zero(where=_budget_link_where(cascade.budget_ids, _SPENT_ROWS_WHERE))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Newly linked users stay blocked If an end user joins a due tier after users are collected, this update clears their database spend but misses their counter. Auth still reads that counter, so the user can remain blocked after the reset

Knowledge Base Used: Spend, budgets, and rate limits

*(_key_counter_key(row) for row in keys),
*(_org_counter_key(row) for row in orgs),
*(_tag_counter_key(row) for row in tags),
*(_enduser_counter_key(row) for row in endusers),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 New spend can be erased If a request is charged while a large tier’s counters are being cleared after commit, the later reset can overwrite that charge with zero. Subsequent budget checks then undercount the new window’s spend

Knowledge Base Used: Spend, budgets, and rate limits

Comment on lines +375 to +382
*(_enduser_counter_key(row) for row in endusers),
),
cache_keys=(
*(key for row in team_memberships for key in _team_membership_cache_keys(row)),
*(key for row in keys for key in _key_cache_keys(row)),
*(key for row in orgs for key in _org_cache_keys(row)),
*(key for row in tags for key in _tag_cache_keys(row)),
*(key for row in endusers for key in _enduser_cache_keys(row)),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Large tiers delay resets The job now writes a counter and deletes a cache entry sequentially for every end user. On a tier with tens of thousands of users, this sweep can delay later budget resets; please batch or bound it

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Comment on lines +47 to +52
GPT_REASONING_SERIES_MARKERS: Final = ("gpt-5", "gpt-6")


def is_gpt_reasoning_series_name(model: str) -> bool:
normalized: Final = model.split("/")[-1]
return any(marker in model for marker in GPT_REASONING_SERIES_MARKERS) and not normalized.startswith("gpt-5-chat")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 GPT-6 capabilities are hardcoded The new marker check and GPT-6 prefix check classify model behavior in code. Repository rules require model-specific flags in model metadata, read through get_model_info. This requirement must be met before merging

Rule Used: What: Do not hardcode model-specific flags in the codebase. Instead, put them in model_prices_and_context_window.json and then read them in via get_model_info Why: Prevents need for users to upgrade litellm each time a new model supports this featu... (source)

Knowledge Base Used: Provider adapters and capabilities

@yuneng-berri
yuneng-berri merged commit b521c4c into stable/1.98.x Sep 25, 2026
4 checks passed
@yuneng-berri
yuneng-berri deleted the litellm_backport_1_98_x_gpt6_budget_0924 branch September 25, 2026 05:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants