Skip to content

fix!: re-check budget on router fallback targets - #41379

Merged
ryan-crabbe-berri merged 3 commits into
BerriAI:mainfrom
runjivu:fix/fallback-budget-check
Sep 16, 2026
Merged

ryan-crabbe-berri merged 3 commits into
BerriAI:mainfrom
runjivu:fix/fallback-budget-check

Conversation

@runjivu

@runjivu runjivu commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Budget is checked against the requested model, not the model that runs
  • A free model with a paid fallback bills with no budget gate
  • An over-budget key can spend without limit through that fallback

How it solves it:

  • Re-check budget per fallback target, where the paid model is chosen
  • Skip targets the caller cannot pay for, keep walking the list
  • Leave the free primary attempt alone, so it is never blocked
  • Mirrors the existing fallback_access_check, on by default

User Flow

Before: a developer on a free model with a paid fallback keeps being billed after their budget is gone

  1. Their admin gives them a key with a $1 cap and a free model that falls back to a paid one
  2. They send POST https://litellm-domain/v1/chat/completions for the free model; it is served free and costs $0
  3. Their free provider goes down, so the same POST now fails upstream
  4. The request still returns 200, silently answered by the paid model and billed in full
  5. They pass the cap and keep going. Every request keeps succeeding and keeps billing
  6. https://litellm-domain/ui/?page=logs shows spend climbing past the cap with no rejections
  7. Anyone holding an over-budget key can keep spending through that fallback with no ceiling

After: the free model still works, and only the paid fallback is refused once they are over budget

  1. Their admin upgrades the proxy. No config change is needed
  2. They send the same POST https://litellm-domain/v1/chat/completions for the free model while over budget
  3. While the free provider is healthy it returns 200 and costs $0, so the cap does not block it
  4. Once the free provider fails, the same POST returns the free provider's own error instead of a billed 200, and the response carries no x-litellm-attempted-fallbacks header
  5. https://litellm-domain/ui/?page=logs shows no new spend
  6. A colleague who is under budget sends the same request and still gets the paid fallback
  7. An over-budget key can no longer spend through the fallback, unless the admin sets enforce_fallback_budget: false to get the old behavior back

Relevant issues

Relates to #41344

Related, and compatible with, #29912 and #38515: both ask for the zero-cost bypass to be extended to _PROXY_MaxBudgetLimiter. That hook is currently the only thing bounding fallback spend for a zero-cost group, so this change is what makes extending it safe. See the note under Proof of Fix, which shows that hook masking the leak on the personal-budget path today.

Affected release

This is a regression, not a long-standing gap. #20249 (v1.81.16) added the zero-cost bypass that waives every auth-time budget check, and that is what lets an over-budget key reach the router at all. Before it, the key's max_budget check refused the request whatever model it named, so the fallback never ran.

Linear ticket

Pre-Submission checklist

  • I have added meaningful tests
  • The handful of test files covering my change pass locally
  • My PR passes all required CI/CD checks. documentation and code-quality cannot pass from this repo; they need docs: document fallback_budget_check and enforce_fallback_budget litellm-docs#1493 merged (details below)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5. Currently 3/5 on one remaining finding that needs a maintainer decision (details below)

Screenshots / Proof of Fix

Shared setup. Real proxy, real Postgres, real router fallback walk, real spend pipeline, and a real paid provider: the fallback target is OpenAI gpt-5.5 on a live API key, so every fallback that gets taken costs real money and lands in the spend table. The zero-cost primary points at a closed port so it always fails. Nothing is mocked.

model_list:
  - model_name: free-model          # zero-cost primary; api_base is a closed port, so it always fails
    litellm_params:
      model: openai/free
      api_base: http://127.0.0.1:8398
      api_key: sk-none
      input_cost_per_token: 0
      output_cost_per_token: 0
    model_info: {input_cost_per_token: 0, output_cost_per_token: 0}
  - model_name: paid-model          # real priced fallback target
    litellm_params:
      model: openai/gpt-5.5
      api_key: os.environ/OPENAI_API_KEY
litellm_settings:
  fallbacks:
    - free-model: ["paid-model"]
general_settings:
  master_key: sk-1234

Two keys, both created through POST /key/generate:

  • over-budget: max_budget: 0.0005, driven past its cap with two direct paid-model calls, then confirmed at $0.00065 through /key/info
  • under-budget: max_budget: 100.0

Cases 1 and 2 run against the config above. Case 3 adds enforce_fallback_budget: false under general_settings and nothing else.

Before (930ec96)

1. Over-budget key requests the zero-cost model

  1. GET /key/info -> spend: 0.0013, max_budget: 0.0005
  2. POST /v1/chat/completions {"model": "free-model"}
HTTP/1.1 200 OK
x-litellm-attempted-fallbacks: 1
model: gpt-5.5-2026-04-23 | content: ok
  1. GET /key/info -> spend: 0.00162. The paid fallback billed while the key was already over its cap

2. Under-budget key requests the zero-cost model

  1. POST /v1/chat/completions {"model": "free-model"} -> HTTP/1.1 200 OK, x-litellm-attempted-fallbacks: 1

3. Over-budget key with enforce_fallback_budget: false

  1. GET /key/info -> spend: 0.00162, max_budget: 0.0005
  2. POST /v1/chat/completions {"model": "free-model"} -> HTTP/1.1 200 OK, x-litellm-attempted-fallbacks: 1
  3. GET /key/info -> spend: 0.00178. The setting is in the config and has no effect, because this build does not know it

After (17844cf)

1. Over-budget key requests the zero-cost model

  1. GET /key/info -> spend: 0.00178, max_budget: 0.0005
  2. POST /v1/chat/completions {"model": "free-model"}
HTTP/1.1 500 Internal Server Error
{"error":{"message":"litellm.InternalServerError: InternalServerError: OpenAIException - Connection error..
  Received Model Group=free-model\nAvailable Model Group Fallbacks=['paid-model']"}}
  1. No x-litellm-attempted-fallbacks header. The paid target was skipped, not attempted
  2. GET /key/info -> spend: 0.00178, unchanged. The caller gets free-model's own error instead of a billed answer

This is the whole point of the change: there is no enforce_fallback_budget anywhere in the config for this run. An unconfigured proxy now refuses the paid fallback.

2. Under-budget key requests the zero-cost model

  1. POST /v1/chat/completions {"model": "free-model"} -> HTTP/1.1 200 OK, x-litellm-attempted-fallbacks: 1

The paid fallback is still taken. Budget state is the only difference between this and case 1, so the 500 above is the skip, not a dead upstream.

3. Over-budget key with enforce_fallback_budget: false

  1. POST /v1/chat/completions {"model": "free-model"} -> HTTP/1.1 200 OK, x-litellm-attempted-fallbacks: 1
  2. GET /key/info -> spend: 0.00178 then 0.00194. The opt-out puts the old unguarded behavior back, exactly as the Before run shows it

Note for #29912 / #38515

An earlier run of this used a personal max_budget and returned 429 Max budget limit reached at auth, never reaching the fallback. _PROXY_MaxBudgetLimiter still blocks zero-cost models when the user's personal budget is exhausted, the exact behavior those two issues ask to remove, so it incidentally masks this leak on that one path. It reads user_max_budget only, so key budgets (used above) are already unguarded today. Removing that hook's blocking without a fallback-time check would open the personal-budget path too.


Implementation notes

The mechanism is the budget sibling of fallback_access_check, whose module docstring describes this same class of problem: auth-time checks run against the requested group, the router picks a different one afterwards.

router
  ├─ attempt "free-model" → zero-cost deployment      ($0)
  │      └─ upstream failure (429, quota exhausted, ...)
  │
  └─ run_async_fallback()
        for each target in fallbacks["free-model"]:
           ├─ _is_fallback_target_authorized(target)     ← existed
           ├─ _is_fallback_target_within_budget(target)  ← added
           └─ execute target
access (existed) budget (added)
Router(fallback_access_check=…) Router(fallback_budget_check=…)
RouterFallbackAccessCheck RouterFallbackBudgetCheck
can_key_call_resolved_model(...) key and user max_budget, evaluated against the target
general_settings.enforce_fallback_model_access (opt-in) general_settings.enforce_fallback_budget (on, opt out with false)

Three deliberate behaviours:

  1. A zero-cost fallback target is always allowed, checked before any spend lookup. Refusing a free target would deny a request on spend some other model accrued, the same reasoning as the auth-time bypass.
  2. A team key does not inherit the key owner's personal budget unless apply_user_budget_to_team_keys is set, matching _PROXY_MaxBudgetLimiter.
  3. Fails closed on lookup error, matching fallback_model_access.py. The tradeoff is that a transient spend-counter failure denies the paid fallback to someone under budget.

Scope and known limitations are stated in the module docstring: this covers the key's and the user's max_budget. Team, team-member, end-user, org, global, per-model and the key's rolling budget_limits windows are not covered. It also reads the spend counter rather than reserving against it, so concurrent fallbacks can cross a cap together; and a request reaching the router without metadata["user_api_key_auth"] is not restricted. The last two are shared with fallback_model_access.py.

Type

🐛 Bug Fix

Caveats (if any)

Severe

  • Changes auth behavior on upgrade, with no config change
  • Over-budget callers stop getting paid fallback targets
  • Operators who rely on that must set enforce_fallback_budget: false

Medium

  • Counter is read, not reserved: concurrent fallbacks can cross a cap
  • Only key and user max_budget; team, org, per-model, windows uncovered
  • Adds a spend-counter read per paid fallback target

Low

  • Async fallback path only, same as fallback_access_check
  • Requests without metadata["user_api_key_auth"] stay unrestricted

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

@CLAassistant

CLAassistant commented Sep 16, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@codspeed

codspeed Bot commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing runjivu:fix/fallback-budget-check (17844cf) with main (930ec96)

Open in CodSpeed

@runjivu
runjivu force-pushed the fix/fallback-budget-check branch from 7718d42 to 039a4d7 Compare September 16, 2026 06:22
@codecov

codecov Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 98.59155% with 1 line in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
litellm/router_utils/fallback_event_handlers.py 90.90% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

@runjivu

runjivu commented Sep 16, 2026

Copy link
Copy Markdown
Contributor Author

The documentation and code-quality failures are tests/documentation_tests/test_router_settings.py, which requires every Router.__init__ kwarg to appear in the router_settings - Reference table. That table lives in BerriAI/litellm-docs (checked out into docs/my-website at CI time), so it can't be fixed from this repo.

Docs PR: BerriAI/litellm-docs#1493 — adds the fallback_budget_check row, the enforce_fallback_budget general_settings entry, and an "Enforce Budget on Fallbacks" section mirroring the access one.

These two jobs will stay red until that merges. Verified locally with the docs branch checked out at docs/my-website: all four documentation_tests pass.

lint was a real ruff format miss on my side and is fixed in the latest push.

@runjivu
runjivu force-pushed the fix/fallback-budget-check branch from 88c0d51 to 2e0aa52 Compare September 16, 2026 06:53
@runjivu
runjivu marked this pull request as ready for review September 16, 2026 07:56
@runjivu
runjivu requested a review from a team September 16, 2026 07:56
@greptile-apps

greptile-apps Bot commented Sep 16, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR adds opt-in key and user budget validation before cross-model fallback attempts and skips fallback targets whose caller is over budget

  • Adds a proxy-injected asynchronous fallback budget predicate
  • Preserves zero-cost targets and the original primary error when every fallback is skipped
  • Adds router wiring, protocol definitions, and focused unit coverage
  • The key lookup still needs authoritative spend verification, and the implementation adds prohibited database-backed work to the fallback request path

Confidence Score: 3/5

This PR is not yet safe to merge because fallback key-budget enforcement can accept stale-low spend state, and explicit repository requirements remain unsatisfied

The key fallback lookup does not supply its budget to get_current_spend, bypassing the authoritative floor check used by existing key enforcement. The new per-target database-backed lookups and structural test also violate repository requirements

Files Needing Attention: litellm/proxy/auth/fallback_budget.py, tests/test_litellm/proxy/test_proxy_server.py

Security Review

A blocking authentication-layer correctness gap remains because the key fallback decision can use stale-low spend state without authoritative floor verification

Important Files Changed
Filename Overview
litellm/proxy/auth/fallback_budget.py Adds fallback budget evaluation, but omits authoritative key-spend verification and introduces prohibited database-backed critical-path lookups
litellm/router_utils/fallback_event_handlers.py Integrates budget rejection into the existing fallback traversal while preserving original-error behavior
litellm/router.py Adds the optional budget-check dependency to Router without changing default behavior
litellm/proxy/proxy_server.py Injects the proxy fallback budget checker into both router construction paths
tests/test_litellm/proxy/auth/test_fallback_budget.py Covers key, user, team-key, zero-cost, metadata, and failure behavior but not stale-low authoritative key-spend handling
tests/test_litellm/proxy/test_proxy_server.py Adds a structural callback identity assertion instead of an observable configuration behavior test

Reviews (1): Last reviewed commit: "fix: re-check budget on router fallback ..." | Re-trigger Greptile

Comment on lines +90 to +93
key_spend: Final = await _counter_spend(
counter_key=f"spend:key:{valid_token.token}",
fallback_spend=valid_token.spend or 0.0,
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 security Stale spend can pass

Omitting max_budget lets stale-low counters allow paid fallbacks.

How this was verified: Existing key checks pass max_budget for authoritative verification.

Rule Used: What: Fail any PR which may contains a security incident on litellm's authentication layer Why: Do not cause security incidents Bad: ```python # Check cache first cache_key = ( f"oidc_userinfo_{token[:20]}" # Use fi... (source)

Comment on lines +68 to +70

return general_settings.get("apply_user_budget_to_team_keys") is True

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Fallbacks add database reads

Each paid target can call database-backed get_current_spend twice, violating the requirement to avoid critical-path database requests and use budget object helpers.

Rule Used: What: Avoid creating new database requests or Router objects in the critical request path. Why: Creating these objects on every request causes performance degradation and unnecessary resource consumption. (source)


router, _, _ = await ProxyConfig().load_config(router=None, config_file_path=str(config_file))

assert router.fallback_budget_check is router_fallback_budget_check

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Test checks wiring only

This identity assertion tests callback wiring, not fallback behavior, violating the requirement that tests verify function rather than code structure before merging.

Context Used: AGENTS.md (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Comment thread litellm/proxy/auth/fallback_budget.py Outdated
if _is_model_cost_zero(model=model, llm_router=llm_router):
return True

key_budget: Final = valid_token.max_budget

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Low: Non-key budget scopes are not enforced

This predicate checks only key and personal-user limits. A caller whose team, team-member, end-user, organization, global, or per-model budget is exhausted can request a free group and still consume a paid fallback because auth skipped all those checks for the original free model. The fallback evaluation needs read-only checks for every budget scope that common_checks bypasses.

if valid_token is None:
return True
try:
return await is_token_within_budget_for_model(model=model, valid_token=valid_token, llm_router=llm_router)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Low: Fallback admission does not reserve budget

This performs a read-only spend check after auth skipped budget reservation for the free model. An attacker can submit many concurrent requests while the counter is just below its limit; each observes the same available budget and proceeds to the paid fallback before cost callbacks update the counter. Atomically reserve the fallback's estimated cost against the applicable counters before attempting it, and reconcile or release that reservation through the existing completion callbacks.

if not self.is_enforced():
return True
valid_token: Final = _user_api_key_auth_from_request(request_kwargs)
if valid_token is None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Low: Missing auth metadata fails open

Authenticated proxy requests do not all carry user_api_key_auth. In particular, /queue/chat/completions manually adds scalar user_api_key_* metadata and then calls schedule_acompletion, so an over-budget caller can trigger a paid fallback through that endpoint and this branch allows it. Ensure every authenticated router call carries the trusted auth object, or distinguish known internal calls and fail closed for proxy requests where it is absent.

@veria-ai

veria-ai Bot commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

PR overview

This pull request adds budget re-checking when the router moves a request from a free model group to a paid fallback target. The changes focus on fallback-budget handling in the proxy authentication flow.

One issue has been addressed, but paid fallbacks can still proceed despite exhausted budgets in several cases, including non-key scopes, rolling key windows, and requests missing trusted authentication metadata. Concurrent requests can also pass read-only checks before spend counters update, allowing aggregate usage beyond configured limits. These gaps permit authenticated callers to generate unintended paid usage, though the impact is limited to budget enforcement and excess spend.

Open issues (4)

Fixed/addressed: 1 · PR risk: 6/10

@runjivu
runjivu force-pushed the fix/fallback-budget-check branch from 2e0aa52 to de7716b Compare September 16, 2026 08:49
@runjivu

runjivu commented Sep 16, 2026

Copy link
Copy Markdown
Contributor Author

Thanks both — pushed de7716b addressing the two actionable findings, and answering the rest below.

Fixed

Stale spend can pass (Greptile P1, fallback_budget.py:93) — correct, and it was a real gap. Both existing enforcement sites pass max_budget= to get_current_spend; mine did not, so the authoritative floor re-check never ran and a counter restored from an older snapshot would keep admitting paid fallbacks. Both counter reads now pass it, with a test that a stale-low counter still refuses a paid target.

Test checks wiring only (Greptile P2, test_proxy_server.py) — fair, CLAUDE.md says "Never test structure of code only function of it". It was a mirror of the neighbouring test_load_config_router_authorizes_fallback_targets_against_the_calling_key, but the rule is explicit, so it now drives the config-loaded router's predicate and asserts an over-budget caller is refused and an under-budget one is not.

Documented as known limitations

No reservation on fallback admission (veria-ai) — accurate. This reads the counter rather than reserving against it, so concurrent requests all observe the same pre-billing figure and a cap can be crossed by roughly the number of in-flight fallbacks times their cost. Auth-time enforcement avoids this via reserve_budget_for_request, which the zero-cost bypass skips. Making it a hard cap means reserving per fallback attempt and reconciling on completion — the same direction discussed on #41344. Noted in the module docstring; it turns unbounded overspend into bounded overspend, which I think is the right first step but is worth being explicit about.

Missing auth metadata fails open (veria-ai) — verified, and correct: only add_litellm_data_to_request writes metadata["user_api_key_auth"], and /queue/chat/completions assembles metadata by hand, so the caller falls through as unauthenticated. Worth flagging that fallback_model_access.py has byte-identical handling, so this is a pre-existing gap in the shipped access check rather than something introduced here. I've documented it instead of diverging from the access check unilaterally — happy to fix both in a follow-up if you'd like them to fail closed together.

Answered, no change

Enforcement disabled by default (veria-ai) — deliberate, and the open question I raised in #41344. Defaulting to true changes behavior for every existing deployment: paid fallbacks start being refused for over-budget callers who are getting them today. That seems like a maintainer call rather than something to flip in this PR. I'll switch it if you'd prefer default-on.

Non-key budget scopes not enforced (veria-ai) — deliberate and disclosed in the PR description and #41344. The auth-path functions for team, team-member, end-user, org, global and per-model budgets raise rather than return a verdict, and reusing them here would fire threshold alerts and take spend reservations for a target that is then skipped. Full coverage wants read-only evaluation paths per dimension; happy to follow up.

Fallbacks add database reads (Greptile P2) — worth noting this pulls directly against the P1 fix: passing max_budget is exactly what enables the authoritative read. Mitigations are that it runs only on the fallback path, only when the feature is enabled, and get_current_spend skips the DB read for healthy counters and caches in-process for a few seconds, so a persistently stale counter drives at most one read per counter per window. I took correctness over the read here; say the word if you'd rather have the cheaper check.

CI note: documentation and code-quality remain red pending BerriAI/litellm-docs#1493, as described above — they can't go green from this repo.

@greptileai

@greptile-apps

greptile-apps Bot commented Sep 16, 2026

Copy link
Copy Markdown
Contributor

Thanks — the stale-counter finding is addressed: both counter reads now pass max_budget, and the regression test verifies that argument is supplied. The config-loaded test also now exercises the predicate’s over/under-budget behavior rather than only asserting callback identity.

The database-path finding remains unresolved under the repository’s request-path rule. litellm/proxy/auth/fallback_budget.py:81 invokes get_current_spend during fallback admission, potentially once for the key and once for the user for every paid fallback target. Passing max_budget is necessary for authoritative verification, but it also means this path can perform database-backed work; a healthy-counter fast path and short cache reduce frequency but do not remove the new request-path dependency. Please route this through an approved existing budget helper/object path, or obtain an explicit maintainer exemption for this rule.

The default-off behavior, limited budget scopes, missing metadata behavior, and lack of reservation are documented design limitations rather than additional actionable findings for this scoped change. I would clear the stale-spend and test-wiring findings, but keep the confidence at 3/5 until the request-path database work is resolved or explicitly accepted.

Tip: You can customize Greptile's behavior for this repo with .greptile/rules.md and .greptile/config.json.

if _is_model_cost_zero(model=model, llm_router=llm_router):
return True

key_budget: Final = valid_token.max_budget

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Low: Rolling key budgets are not enforced

A key that has exhausted a budget_limits window can request a free model group and still trigger its paid fallback because this path checks only max_budget. Evaluate every configured window against its spend:key:{token}:window:{budget_duration} counter using the same window parameters as _virtual_key_multi_budget_check before admitting the fallback.

Budget is enforced once during auth, against the requested model group.
`_is_model_cost_zero` waives every budget check for a zero-cost group, and the
router then picks a fallback target afterwards, inside `run_async_fallback`,
where nothing re-checks budget. A free model with a paid fallback therefore
bills with no budget gate at all.

Add `fallback_budget_check`, the budget sibling of the existing
`fallback_access_check`: a predicate awaited per fallback target that skips
targets the caller cannot pay for. The primary attempt is untouched, so a
zero-cost model is never blocked by budget and only the paid fallback is
refused.

Counter reads pass `max_budget` so `get_current_spend` verifies against
authoritative recorded spend, matching the auth-time key and user checks; a
counter restored from an older snapshot reads as a hit rather than a clean
miss, so without it a stale-low value would keep admitting paid fallbacks.

A zero-cost fallback target is always allowed, and a team key does not inherit
the key owner's personal budget unless `apply_user_budget_to_team_keys` is set,
matching `_PROXY_MaxBudgetLimiter`.

Scope is key and user budgets. Team, team-member, end-user, org, global and
per-model budgets are not covered yet: those auth-path functions enforce rather
than report, so reusing them would fire threshold alerts and take spend
reservations for a target that is then skipped. Two limitations of that scope
are documented in the module docstring: the check reads the spend counter
rather than reserving against it, so concurrent fallbacks can cross a cap
together; and a request reaching the router without
`metadata["user_api_key_auth"]` is not restricted. Both are shared with
`fallback_model_access.py`.

Opt-in via `general_settings.enforce_fallback_budget`.

Relates to BerriAI#41344

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@runjivu
runjivu force-pushed the fix/fallback-budget-check branch from de7716b to 4a70bc3 Compare September 16, 2026 10:06
@runjivu

runjivu commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor Author

@ryan-crabbe-berri — tagging you since this is the direct sibling of fallback_access_check, which you added in 3ea5014. Happy to redirect if someone else owns budgets.

This is out of draft and the Proof of Fix section now has a full e2e run against a real proxy and Postgres, with the leak reproduced at the merge base and closed at the PR tip. Two things need a decision from a maintainer rather than more work from me.

1. A Greptile rule that conflicts with a Veria finding

Greptile is holding at 3/5 on one finding and offers two exits: "resolved or explicitly accepted." It objects that get_current_spend on the fallback path is request-path database work.

The conflict is that Veria's findings on the same PR ask for more budget dimensions — specifically the key's rolling budget_limits windows. Those windows keep their accumulated spend only in spend:key:{token}:window:{budget_duration} counters; there is no auth-loaded field carrying it. So they cannot be enforced without counter reads. The two asks are mutually exclusive, not a matter of effort.

Greptile's own P1 on this PR also asked me to pass max_budget to get_current_spend (fixed, and correct — it is what triggers authoritative verification). That argument is precisely what enables the read it now objects to.

What I'd like from you, whichever you prefer:

  • Accept the reads, on the grounds that they happen only on the fallback path, only when enforce_fallback_budget is on, and get_current_spend already skips the DB for healthy counters and caches in-process; or
  • Tell me to drop them and read the auth-loaded valid_token.spend / user_spend instead. Cheaper and clears the rule mechanically, at the cost of a roughly request-duration-stale figure and permanently foreclosing rolling-window support in this design.

I did not want to trade away a capability to satisfy a bot without you weighing in.

2. The two red checks are not fixable from this repo

documentation and code-quality both run tests/documentation_tests/test_router_settings.py, which requires every Router.__init__ kwarg to appear in the router_settings - Reference table. That table lives in BerriAI/litellm-docs, checked out into docs/my-website at CI time.

BerriAI/litellm-docs#1493 adds the row, plus the enforce_fallback_budget general_settings entry and an "Enforce Budget on Fallbacks" section mirroring the access one. It is 38 additive lines and not a draft. Merging it turns both checks green here. Verified locally with that branch checked out at docs/my-website: all four documentation tests pass.

Also worth your attention

The e2e surfaced something relevant to #29912 and #38515. My first run used a personal max_budget and got 429 Max budget limit reached at auth — _PROXY_MaxBudgetLimiter still blocks zero-cost models there, which is exactly what those issues ask to remove. It reads user_max_budget only, so key budgets are already unguarded today (that is the run in the PR description). Removing that hook's blocking without a fallback-time check would open the personal-budget path as well.

A budget bypass that ships off by default stays open for every deployment
that does not know to look for the flag, so `enforce_fallback_budget` now
defaults to true and `general_settings.enforce_fallback_budget: false` is
the opt-out for anyone who wants the old unguarded behaviour back.

BREAKING CHANGE: a paid fallback target is now refused for callers who are
over their key or user `max_budget`. Deployments relying on fallbacks to
keep serving over-budget callers must set enforce_fallback_budget: false.
@ryan-crabbe-berri ryan-crabbe-berri changed the title fix: re-check budget on router fallback targets fix!: re-check budget on router fallback targets Sep 16, 2026

@ryan-crabbe-berri ryan-crabbe-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM; thanks!

ryan-crabbe-berri added a commit to BerriAI/litellm-docs that referenced this pull request Sep 16, 2026
* docs: document fallback_budget_check and enforce_fallback_budget

Router gained `fallback_budget_check` (BerriAI/litellm#41379), the budget
sibling of `fallback_access_check`: budget is enforced once during auth against
the requested model group, while the fallback target is chosen afterwards
inside the router, so a zero-cost model with a paid fallback bills with no
budget gate.

Adds the `router_settings` reference row, the `general_settings` reference row
and yaml entry for `enforce_fallback_budget`, and an "Enforce Budget on
Fallbacks" section mirroring "Enforce Key Model Access on Fallbacks".

Relates to BerriAI/litellm#41344

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: enforce_fallback_budget is on by default

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants