Skip to content

feat(proxy): opt-in budget rollover carrying overage into the next window - #38514

Merged
yassin-berriai merged 2 commits into
litellm_internal_stagingfrom
litellm_budget_rollover_lit_3085
Aug 27, 2026
Merged

yassin-berriai merged 2 commits into
litellm_internal_stagingfrom
litellm_budget_rollover_lit_3085

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Aug 27, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Budget resets forgive all over-cap spend at every window boundary
  • Customers want overage carried into the next window instead

How it solves it:

  • New opt-in budget_rollover litellm setting (default off, admin UI configurable)
  • At reset, spend above max_budget carries forward: carried = max(0, spend - max_budget)
  • Applies to keys, users, teams, team members, orgs, tags, end users, budget tiers, and per-window counters
  • DB writes use atomic decrement so spend landing mid-reset is never erased

User Flow

Before: a proxy admin caps a key at $100/month; a user who spends $150 gets a clean $100 again next month, so the overage is free

  1. Admin creates a key: POST http://localhost:4000/key/generate with {"max_budget": 100, "budget_duration": "30d"}
  2. The user spends $150 through POST http://localhost:4000/v1/chat/completions (requests keep passing until enforcement catches up)
  3. When the window resets, GET http://localhost:4000/key/info shows "spend": 0.0, so the $50 overage is forgiven
  4. The user can spend another full $100 in the new window

After: with rollover enabled, the overage is deducted from the next window's allowance

  1. Admin turns on budget_rollover in the Admin UI general settings (or litellm_settings.budget_rollover: true)
  2. Admin creates the same key: POST http://localhost:4000/key/generate with {"max_budget": 100, "budget_duration": "30d"}
  3. The user spends $150 through POST http://localhost:4000/v1/chat/completions
  4. When the window resets, GET http://localhost:4000/key/info shows "spend": 50.0, so only $50 of budget remains for the new window
  5. A key that stayed under its cap still resets to "spend": 0.0 exactly as before

Relevant issues

Linear ticket

Resolves LIT-3085

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Screenshots / Proof of Fix

Live proxy on localhost:4000 (Postgres + real OpenAI calls), PROXY_BUDGET_RESCHEDULER_MIN_TIME=10 PROXY_BUDGET_RESCHEDULER_MAX_TIME=15 so resets fire quickly. Fixture is byte-identical across both arms; the arm is identified by whether the budget_rollover field exists in GET /config/list?config_type=general_settings, never by changing the payload. Both arms exercise the direct key reset path (reset_budget_for_litellm_keys); the cascade/linked-row and per-window paths are covered by the unit tests in this PR.

Before (cd63c7e, merge base)

Over-cap spend is forgiven at reset

  1. curl -s "http://localhost:4000/config/list?config_type=general_settings" -H "Authorization: Bearer sk-1234" | grep -c budget_rollover returns 0 (MARKER budget_rollover ABSENT: unfixed build)
  2. curl -s -X POST http://localhost:4000/key/generate -H "Authorization: Bearer sk-1234" -d '{"max_budget": 0.00001, "budget_duration": "30s", "key_alias": "lit3085-before"}' returns 200 with the key
  3. curl -s -X POST http://localhost:4000/v1/chat/completions -H "Authorization: Bearer <key>" -d '{"model": "gpt-4o-mini", "messages": [{"role": "user", "content": "say OK"}]}' returns 200 with a real completion; key spend reaches 3.3e-05, over the 1e-05 cap
  4. After the 30s window resets: curl -s "http://localhost:4000/key/info?key=<key>" -H "Authorization: Bearer sk-1234" shows "spend": 0.0, "budget_reset_at": "2026-08-27T13:00:30+00:00", so the 2.3e-05 overage is fully forgiven

After (9caa257, PR tip)

Over-cap spend carries into the next window

  1. curl -s "http://localhost:4000/config/list?config_type=general_settings" -H "Authorization: Bearer sk-1234" | grep -c budget_rollover returns 1 (MARKER budget_rollover PRESENT: fixed build, setting visible in Admin UI general settings)
  2. curl -s -X POST http://localhost:4000/config/field/update -H "Authorization: Bearer sk-1234" -d '{"field_name": "budget_rollover", "field_value": true, "config_type": "general_settings"}' returns {"message":"Field budget_rollover updated","status":"success"} (persists in DB and hot-reloads)
  3. curl -s -X POST http://localhost:4000/key/generate -H "Authorization: Bearer sk-1234" -d '{"max_budget": 0.000001, "budget_duration": "30s", "key_alias": "lit3085-tip2"}' returns 200
  4. Same completion request returns 200 with a real completion; curl -s "http://localhost:4000/key/info?key=<key>" -H "Authorization: Bearer sk-1234" shows pre-reset "spend": 2.55e-06 against the 1e-06 cap
  5. After two 30s windows elapse: "spend": 5.5e-07, "budget_reset_at": "2026-08-27T13:37:00+00:00", exactly 2.55e-06 - 2 * 1e-06, one cap deducted per window with the remainder carried forward

Under-cap spend still resets to zero

  1. Same fixture with "max_budget": 0.00001; pre-reset "spend": 2.55e-06 (under cap)
  2. After reset: "spend": 0.0, "budget_reset_at": "2026-08-27T13:35:30+00:00", unchanged legacy behavior

Type

🆕 New Feature

Caveats (if any)

Medium

  • Carried overage shows as plain spend on the new window; no separate "carried over" field in API responses
  • Live proof covers the direct key reset path; cascade/linked-row, end-user, and per-window paths are unit-tested only

Low

  • Team aggregate ceiling with member rollover (team 100, members total 150) unchanged: team budget still enforces its own cap independently
  • Spend recorded by async batch reconciliation after a reset lands in the new window (same as today)

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

@greptileai

Link to Devin session: https://app.devin.ai/sessions/327a7b4156a6492eabc012fb228bd913
Open in Devin Desktop: https://app.devin.ai/desktop/session/327a7b4156a6492eabc012fb228bd913?variant=devin
Requested by: @yassin-berriai

…ndow

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR. Add '(aside)' to your comment to have me ignore it.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

PR #38514 is a Devin-authored PR in BerriAI/litellm, but it has no enterprise label — exiting with no changes per scope rules.

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@greptile-apps

greptile-apps Bot commented Aug 27, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR adds opt-in budget rollover across direct, linked, end-user, and per-window budget resets

  • Carries over-cap spend by atomically decrementing each applicable budget cap
  • Orders cascade writes so under-cap rows are zeroed before over-cap rows are decremented
  • Synchronizes post-reset spend counters with carried values and adds sequential-write regression coverage

Confidence Score: 5/5

The PR appears safe to merge

The previously reported cascade ordering issue is fixed by zeroing under-cap rows before decrementing over-cap rows, and no blocking failure remains

Important Files Changed

Filename Overview
litellm/proxy/common_utils/reset_budget_job.py Implements rollover calculation, ordered cascade resets, direct resets, counter synchronization, and per-window rollover without leaving the previously reported ordering defect
litellm/repositories/unit_of_work.py Adds atomic decrement payloads for direct and linked spend resets
tests/test_litellm/proxy/common_utils/test_reset_budget_job.py Covers enabled and disabled rollover, direct and cascade paths, end users, counters, and sequential statement execution
litellm/proxy/proxy_server.py Registers budget rollover in the Admin UI general-settings schema
litellm/constants.py Allows the new rollover flag to be safely loaded from database-backed settings

Reviews (2): Last reviewed commit: "fix(proxy): zero under-cap rows before d..." | Re-trigger Greptile

Comment on lines +172 to +178
for budget_id, cap in cascade.rollover_caps.items():
writes.queue_spend_decrement(
where={"budget_id": budget_id, **extra, "spend": {"gt": cap}}, amount=cap
) # mutable-ok: prisma where filter must be a dict
writes.queue_spend_zero(
where={"budget_id": budget_id, **extra, "spend": {"gt": 0, "lte": cap}}
) # mutable-ok: prisma where filter must be a dict

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Sequential updates erase rollover

When a linked row is above its cap but no more than twice the cap, the first sequential update_many decrements it into the 0 < spend <= cap range and the following update resets it to zero, forgiving the overage instead of carrying it forward. The end-user reset at lines 199-204 has the same overlapping predicates.

Knowledge Base Used: Spend, budgets, and rate limits

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 9caa257: the zero now runs before the decrement, plus a sequential-replay regression test.

@@ -1222,7 +1350,7 @@ async def _reset_budget_common(
still holds the pre-reset value, admitting requests past the cap.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Reset mutates input object

Assigning the calculated rollover directly to item.spend mutates the helper's input parameter, introducing hidden caller-visible state changes and making the reset flow dependent on mutation order. Construct an immutable reset value instead, as required by the repository's immutability guidance.

Context Used: CLAUDE.md (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In-place mutation is this helper's documented contract; callers return the same object. Keeping it avoids a wider refactor here.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You're right — given that _reset_budget_common is documented to mutate the supplied object and callers return that same object, the assignment is consistent with the helper's contract. My immutability concern doesn't apply here, so no change is needed for this comment.

…cascade resets

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@codecov

codecov Bot commented Aug 27, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 92.55319% with 7 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
litellm/proxy/common_utils/reset_budget_job.py 91.35% 7 Missing ⚠️

📢 Thoughts on this report? Let us know!

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai please re-review at 9caa257: cascade write ordering fixed and sequential-replay tests added.

@codspeed

codspeed Bot commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_budget_rollover_lit_3085 (9caa257) with litellm_internal_staging (8ebcb3e)1

Open in CodSpeed

Footnotes

  1. No successful run was found on litellm_internal_staging (cd63c7e) during the generation of this report, so 8ebcb3e was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

@yassin-berriai
yassin-berriai merged commit de53283 into litellm_internal_staging Aug 27, 2026
80 checks passed
@yassin-berriai
yassin-berriai deleted the litellm_budget_rollover_lit_3085 branch August 27, 2026 19:46
achraf-mer pushed a commit to achraf-mer/litellm that referenced this pull request Sep 30, 2026
…et (BerriAI#39726)

Adapted for stable/1.99.x: this line predates budget rollover (BerriAI#38514), so each end user's spend counter is zeroed through the existing counter_keys list instead of counter_resets, and its cache entry is evicted via end_user_cache_key.

Signed-off-by: amasen02 <amasen02@users.noreply.github.com>
(cherry picked from commit daced81)
achraf-mer pushed a commit to achraf-mer/litellm that referenced this pull request Sep 30, 2026
…BerriAI#40639)

Adapted for stable/1.99.x: this line predates budget rollover (BerriAI#38514), so the fix is applied to _commit_budget_cascade_once directly. End users reset on the budget link plus a NULL budget_id branch for the default tier, which is what upstream's _queue_enduser_resets does with rollover off. The rollover test hunk is dropped.

(cherry picked from commit 8a4fae0)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants