Skip to content

chore(release): backport #39631, #39729, #40639 to stable/1.100.x and cut 1.100.3 - #43132

Merged
yuneng-berri merged 10 commits into
stable/1.100.xfrom
litellm_backport_1_100_x_gpt6_budget_0924
Sep 25, 2026
Merged

yuneng-berri merged 10 commits into
stable/1.100.xfrom
litellm_backport_1_100_x_gpt6_budget_0924

Conversation

@yuneng-berri

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • gpt-6 models fail on chat and Responses with reasoning params
  • Budget tiers with over 32,767 customers never reset
  • Customers stay blocked on a running proxy even after a reset

How it solves it:

User Flow

Before: a developer calling gpt-6 through the proxy gets errors, and customers on a big shared budget tier stay blocked forever

  1. They send POST https://litellm-domain/v1/chat/completions with "model": "gpt-6-astra", "max_tokens": 200 and "reasoning_effort": "low"
  2. They get 400 UnsupportedParamsError: openai does not support parameters: ['reasoning_effort']
  3. They send POST https://litellm-domain/v1/responses with "model": "gpt-6-astra", "temperature": 0.5 and "drop_params": true
  4. They get 400 Unsupported parameter: 'temperature' is not supported with this model
  5. A customer on a 33,000-customer budget tier hits the cap and gets 429 Budget has been exceeded on POST https://litellm-domain/v1/chat/completions
  6. The tier's window passes and GET https://litellm-domain/customer/info still shows the capped spend; the customer keeps getting 429

After: gpt-6 requests succeed, and the tier resets on schedule and the customer is served again without a restart

  1. They send POST https://litellm-domain/v1/chat/completions with "model": "gpt-6-astra", "max_tokens": 200 and "reasoning_effort": "low"
  2. They get 200 with the model's reply
  3. They send POST https://litellm-domain/v1/responses with "model": "gpt-6-astra", "temperature": 0.5 and "drop_params": true
  4. They get 200 with "status": "completed"
  5. A customer on a 33,000-customer budget tier hits the cap and gets 429 Budget has been exceeded on POST https://litellm-domain/v1/chat/completions
  6. The tier's window passes, GET https://litellm-domain/customer/info shows spend: 0, and the same request returns 200 from the same running proxy

Relevant issues

Fixes #40564 and #39726 on stable/1.100.x

Backport of #39631 (merge 025a3ca), #39729 (daced81) and #40639 (merge 8a4fae0), all reachable from main and picked with cherry-pick -x (-m 1 for the two merges). This line has budget rollover, so all three apply with their code lines unchanged: #39631 needed an import-context resolution only, #39729 has one generator reflowed to this line's ruff format, and #40639 adds Final to the test module's typing import, which upstream's test file already had. Each commit message carries its note

Not included on purpose: #41488 (paged end-user cache invalidation after a reset). It is a later perf follow-up on the same surface, the fix works without it, and the stable/1.101.x backport (#42635) made the same call

Dependency bumps, each lock-only inside the existing pyproject range and at or below what main resolves: anyio 4.14.2, gitpython 3.1.60, tornado 6.5.8, pypdf 6.16.1 and soupsieve 2.9. After the bumps grype and OSV report the same residual set as main, with no fixed version available for either

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Live proxy against a local PostgreSQL and real OpenAI calls to gpt-6-astra, same scripts run at the merge base and at this PR's tip. Config:

model_list:
  - model_name: gpt-6-astra
    litellm_params:
      model: openai/gpt-6-astra
      api_key: os.environ/OPENAI_API_KEY
general_settings:
  master_key: sk-...
  database_url: postgresql://...
  proxy_budget_rescheduler_min_time: 5
  proxy_budget_rescheduler_max_time: 6

The budget tier is shared-tier (max_budget: 10, budget_duration: 1d) with 33,000 customers cust-00000000 to cust-00032999 seeded by one SQL insert, since 33,000 POST /customer/new calls is not something a reviewer would want to replay. Before each run every customer's spend is set to 10 and the window is set an hour ahead; the replay then moves budget_reset_at into the past and waits 40 seconds, which covers several reset job runs

Before (9c1216a)

gpt-6 on /v1/chat/completions

  1. Run the request
$ curl -s localhost:4000/v1/chat/completions -H "Authorization: Bearer sk-..." -d '{"model":"gpt-6-astra","max_tokens":200,"reasoning_effort":"low","messages":[{"role":"user","content":"Reply with the word ok"}]}'
  1. Observe the response
{"message": "litellm.UnsupportedParamsError: openai does not support parameters: ['reasoning_effort'], for model=gpt-6-astra. To drop these, set `litellm.drop_params=True` or for proxy:\n\n`litellm_settings:\n drop_params: true`\n. \n If you want to use these  ...(truncated)

gpt-6 on /v1/responses

  1. Run the request
$ curl -s localhost:4000/v1/responses -H "Authorization: Bearer sk-..." -d '{"model":"gpt-6-astra","temperature":0.5,"drop_params":true,"input":"Reply with the word ok"}'
  1. Observe the response
{"message": "litellm.BadRequestError: OpenAIException - {\n  \"error\": {\n    \"message\": \"Unsupported parameter: 'temperature' is not supported with this model.\",\n    \"type\": \"invalid_request_error\",\n    \"param\": \"temperature\",\n    \"code\": nu ...(truncated)

33,000 customers on one budget tier

  1. Run the replay (steps and output below come straight from the script)
1. customers on shared-tier at the 10 cap: 33000
2. $ curl -s localhost:4000/v1/chat/completions -H "Authorization: Bearer sk-..." -d '{"model":"gpt-6-astra","user":"cust-00000001","messages":[{"role":"user","content":"Reply with the word ok"}]}'
    {"code": "429", "message": "Budget has been exceeded! EndUser=cust-00000001 Current cost: 10.81932, Max budget: 10.0"}
3. shared-tier window expires (budget_reset_at moved into the past), wait 40s for the reset job
4. proxy log: 'too many bind variables' x16, 'Failed to reset the budget table cascade' x8
   too many bind variables in prepared statement, expected maximum of 32767, received 33001
5. customers still at cap: 33000, zeroed: 0, window still due: t
6. $ curl -s localhost:4000/v1/chat/completions (same request as step 2)
    {"code": "429", "message": "Budget has been exceeded! EndUser=cust-00000001 Current cost: 10.81932, Max budget: 10.0"}

After (04fcee8)

gpt-6 on /v1/chat/completions

  1. Run the request
$ curl -s localhost:4000/v1/chat/completions -H "Authorization: Bearer sk-..." -d '{"model":"gpt-6-astra","max_tokens":200,"reasoning_effort":"low","messages":[{"role":"user","content":"Reply with the word ok"}]}'
  1. Observe the response
{"model": "gpt-6-astra", "content": "ok", "usage": 4}

gpt-6 on /v1/responses

  1. Run the request
$ curl -s localhost:4000/v1/responses -H "Authorization: Bearer sk-..." -d '{"model":"gpt-6-astra","temperature":0.5,"drop_params":true,"input":"Reply with the word ok"}'
  1. Observe the response
{"status": "completed", "model": "gpt-6-astra", "temperature": 1.0, "output_text": ["ok"]}

33,000 customers on one budget tier

  1. Run the replay (steps and output below come straight from the script)
1. customers on shared-tier at the 10 cap: 33000
2. $ curl -s localhost:4000/v1/chat/completions -H "Authorization: Bearer sk-..." -d '{"model":"gpt-6-astra","user":"cust-00000001","messages":[{"role":"user","content":"Reply with the word ok"}]}'
    {"code": "429", "message": "Budget has been exceeded! EndUser=cust-00000001 Current cost: 10.81932, Max budget: 10.0"}
3. shared-tier window expires (budget_reset_at moved into the past), wait 40s for the reset job
4. proxy log: 'too many bind variables' x0, 'Failed to reset the budget table cascade' x0
5. customers still at cap: 0, zeroed: 33000, window still due: f
6. $ curl -s localhost:4000/v1/chat/completions (same request as step 2)
    {"model": "gpt-6-astra", "content": "ok"}

Type

Bug Fix

Caveats (if any)

mateo-berri and others added 10 commits September 24, 2026 20:38
…e configs (#39631)

Adapted for stable/1.100.x: import-context conflict only (upstream's neighbouring custom_tools import is not on this line); the added and removed lines are identical to upstream.

(cherry picked from commit 025a3ca)
…et (#39726)

Adapted for stable/1.100.x: reflowed one generator to this line's ruff format (upstream reformatted it in a later style commit).

Signed-off-by: amasen02 <amasen02@users.noreply.github.com>
(cherry picked from commit daced81)
…#40639)

Adapted for stable/1.100.x: added Final to the test module's typing import; upstream's test file already imported it.

(cherry picked from commit 8a4fae0)
@yuneng-berri
yuneng-berri requested a review from a team September 25, 2026 04:18
@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you all sign our Contributor License Agreement before we can accept your contribution.
3 out of 4 committers have signed the CLA.

✅ mateo-berri
✅ ryan-crabbe-berri
✅ yuneng-berri
❌ amasen02
You have signed the CLA already but the status is still pending? Let us recheck it.

@greptile-apps

greptile-apps Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 3/5

[Medium risk] Backport of model routing and budget reset fixes to stable branch.

The PR is not yet safe to merge because budget resets can leave end-user enforcement stale, and the model-capability change must satisfy the repository’s metadata requirement.

Findings

  1. P1 Stale spend on other replicas ▶
  2. P1 Reset misses counter invalidation ▶
  3. P2 Hardcoded GPT-6 capability ▶
  4. P2 Serial per-customer invalidation ▶

Summary

The PR backports GPT-6 request handling and a set-based reset for large end-user budget tiers, adds end-user counter and cache invalidation, refreshes selected locked dependencies, and bumps the release to 1.100.3.

  • The set-based reset avoids the per-customer PostgreSQL bind count.
  • End-user invalidation still has cross-replica and collection-versus-write consistency gaps, and its per-customer sweep remains serial.

Reviews (1) · Last reviewed commit: "chore: refresh uv.lock for 1.100.3"

*(key for row in orgs for key in _org_cache_keys(row)),
*(key for row in tags for key in _tag_cache_keys(row)),
*(key for row in model_access_groups for key in _model_access_group_cache_keys(row)),
*(key for row in endusers for key in _enduser_cache_keys(row)),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Stale spend on other replicas
When one replica resets a tier, this change deletes each end-user object from that replica’s cache and Redis, but does not tell other replicas to clear their in-memory copies. Another replica can use the old spend as its budget-check floor and keep returning budget-exceeded errors after the database spend reaches zero. Broadcast these evictions as other management writes do.

Knowledge Base Used: Spend, budgets, and rate limits

population, which past ~32,700 dependents exceeds PostgreSQL's per-statement
bind ceiling and wedges the cascade permanently (#40564).
"""
_queue_budget_linked_resets(writes, cascade, extra=_SPENT_ROWS_WHERE)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Reset misses counter invalidation
If an end user is created, moved onto the tier, or given spend after the job collects end users but before its transaction commits, this new budget-wide update can reset that user’s database spend. The job clears counters and caches only for users in the earlier collection, so the affected user can remain blocked by stale spend. The update and invalidation need to cover the same users.

Knowledge Base Used: Spend, budgets, and rate limits

Comment on lines +119 to +120
if model_name.startswith("gpt-6"):
return True

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Hardcoded GPT-6 capability
This prefix check unconditionally treats GPT-6 names as eligible for the GPT-5.4-plus request path; the new reasoning-series marker also selects behavior by name. The repository requires model-specific capabilities to be declared in model_prices_and_context_window.json and read through get_model_info, rather than hardcoded in transformations. That requirement must be met before merging so capabilities can change without a code release.

Rule Used: What: Do not hardcode model-specific flags in the codebase. Instead, put them in model_prices_and_context_window.json and then read them in via get_model_info Why: Prevents need for users to upgrade litellm each time a new model supports this featu... (source)

Knowledge Base Used: Provider adapters and capabilities

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Comment on lines 690 to +699
@@ -680,6 +696,7 @@ async def _collect_budget_cascade(self, budgets_to_reset: Sequence[LiteLLM_Budge
*(key for row in orgs for key in _org_cache_keys(row)),
*(key for row in tags for key in _tag_cache_keys(row)),
*(key for row in model_access_groups for key in _model_access_group_cache_keys(row)),
*(key for row in endusers for key in _enduser_cache_keys(row)),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Serial per-customer invalidation
For a tier with tens of thousands of end users, this adds a counter overwrite and a cache deletion for every customer. Both lists are then processed with sequential awaits after the database commit. The sweep can take minutes and delay later budget resets even though the database update is set-based. Batch or page the invalidation work.

Knowledge Base Used: Spend, budgets, and rate limits

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

(_model_access_group_counter_key(row), _row_carried_spend(row, rollover_caps))
for row in model_access_groups
),
*((_enduser_counter_key(row), _enduser_carried_spend(row, rollover_caps)) for row in endusers),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Low: Concurrent end-user spend is overwritten

These snapshot-derived values are written to Redis sequentially after the database commit. An authenticated user can submit requests after the reset commits but before their end-user entry is processed; _invalidate_spend_counter() then replaces those live increments with zero or the stale rollover value. This is especially reachable for large shared tiers because every end user requires awaited cache operations. Reset the counter atomically at the transaction boundary, or use window-specific counter keys so post-reset increments cannot be overwritten by the sweep.

@veria-ai

veria-ai Bot commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

PR overview

This release PR backports three changes to the stable/1.100.x branch and prepares version 1.100.3. The touched code includes updates to the scheduled budget-reset and spend-counter handling logic.

One issue remains open in the budget reset flow: an authenticated user can submit requests during a reset window whose spend increments may then be overwritten by stale or zero values. This can undercount usage and weaken budget enforcement, particularly on large shared tiers where counter invalidation takes longer. No reported issues have yet been addressed in this PR.

Open issues (1)

Fixed/addressed: 0 · PR risk: 5/10

@yuneng-berri
yuneng-berri merged commit 385266c into stable/1.100.x Sep 25, 2026
4 checks passed
@yuneng-berri
yuneng-berri deleted the litellm_backport_1_100_x_gpt6_budget_0924 branch September 25, 2026 05:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants