Skip to content

fix(proxy): remove duplicate user budget hook that 429'd zero-cost models - #41345

Merged
ryan-crabbe-berri merged 3 commits into
mainfrom
litellm_remove_duplicate_user_budget_hook
Sep 16, 2026
Merged

ryan-crabbe-berri merged 3 commits into
mainfrom
litellm_remove_duplicate_user_budget_hook

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Over-budget users got 429 on zero-cost models too
  • Free-model exemption only lived in auth, a second hook re-checked without it

How it solves it:

  • Delete _PROXY_MaxBudgetLimiter, auth already enforces the user budget
  • One place decides, so the zero-cost exemption applies once and for all
  • Regression tests: common_checks lets an over-budget user through on a 0/0 model and still raises BudgetExceededError on a priced one, and the registered hook set no longer re-enforces the user budget

User Flow

Before: a user who has spent their personal budget is locked out of models the admin priced at $0

  1. Admin creates a user with max_budget and gives them a personal key via POST http://localhost:4000/user/new
  2. The user sends POST http://localhost:4000/v1/chat/completions with a paid model until their spend passes the budget
  3. They send POST http://localhost:4000/v1/chat/completions with a model configured with input_cost_per_token: 0 and output_cost_per_token: 0
  4. They get 429 {"message": "Max budget limit reached.", "type": "throttling_error"}
  5. They send the same request with the paid model and get 429 {"message": "ExceededBudget: User=... over budget...", "type": "budget_exceeded"}

After: the same user keeps using $0 models while paid models stay blocked

  1. Admin creates a user with max_budget and gives them a personal key via POST http://localhost:4000/user/new
  2. The user sends POST http://localhost:4000/v1/chat/completions with a paid model until their spend passes the budget
  3. They send POST http://localhost:4000/v1/chat/completions with the $0 model
  4. They get 200 with a normal completion
  5. They send the same request with the paid model and still get 429 {"message": "ExceededBudget: User=... over budget...", "type": "budget_exceeded"}

Relevant issues

Affected release

Linear ticket

Resolves LIT-7464

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Proxy started with uv run --no-sync litellm --config config.yaml --port 4000 against Postgres, real OpenAI calls. Config:

model_list:
  - model_name: free-gpt-5.4-nano
    litellm_params:
      model: openai/gpt-5.4-nano
      api_key: os.environ/OPENAI_API_KEY
    model_info:
      input_cost_per_token: 0
      output_cost_per_token: 0
  - model_name: paid-gpt-5.4-nano
    litellm_params:
      model: openai/gpt-5.4-nano
      api_key: os.environ/OPENAI_API_KEY

general_settings:
  master_key: sk-1234

Setup shared by both runs: create the user and key, then burn the budget with one paid call

curl -s http://localhost:4000/user/new -H "Authorization: Bearer sk-1234" -H "Content-Type: application/json" \
  -d '{"user_id":"lit7464-user","max_budget":0.000001,"models":["free-gpt-5.4-nano","paid-gpt-5.4-nano"]}'
# -> {"user_id": "lit7464-user", "max_budget": 1e-06, "key": "sk-..."}

curl -s http://localhost:4000/v1/chat/completions -H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
  -d '{"model":"paid-gpt-5.4-nano","messages":[{"role":"user","content":"say hi"}]}'
# -> 200, "Hi! How can I help you today?"

curl -s "http://localhost:4000/user/info?user_id=lit7464-user" -H "Authorization: Bearer sk-1234"
# -> {"spend": 1.91e-05, "max_budget": 1e-06}   (user is now over budget)

Then the same loop in both runs:

for m in free-gpt-5.4-nano paid-gpt-5.4-nano; do
  echo "# model=$m"
  curl -s -w '\nHTTP %{http_code}\n' http://localhost:4000/v1/chat/completions \
    -H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
    -d "{\"model\":\"$m\",\"messages\":[{\"role\":\"user\",\"content\":\"say hi\"}]}"
done

Before (0b3e564)

Over-budget user, $0 model

  1. Run the loop above
  2. Observed:
# model=free-gpt-5.4-nano
HTTP 429
{"message": "Max budget limit reached.", "type": "throttling_error", "param": null, "code": "429"}

Over-budget user, paid model

  1. Same loop
  2. Observed:
# model=paid-gpt-5.4-nano
HTTP 429
{"message": "ExceededBudget: User=lit7464-user over budget. Spend=1.91e-05, Budget=1e-06", "type": "budget_exceeded", "param": null, "code": "429"}

After (789b828)

Over-budget user, $0 model

  1. Run the loop above
  2. Observed:
# model=free-gpt-5.4-nano
HTTP 200
{"content": "Hi! How can I help you today?"}

Over-budget user, paid model

  1. Same loop
  2. Observed:
# model=paid-gpt-5.4-nano
HTTP 429
{"message": "ExceededBudget: User=lit7464-user over budget. Spend=1.91e-05, Budget=1e-06", "type": "budget_exceeded", "param": null, "code": "429"}

Type

🐛 Bug Fix

Caveats (if any)

Medium

  • Custom auth deployments without custom_auth_run_common_checks: true skip common_checks, so the deleted hook was the only thing enforcing a user_max_budget that the custom auth returned on the token. Those deployments now need the flag (the proxy already warns at boot that budgets are not enforced without it). Kept out of this PR on purpose: re-adding the check there recreates the drift that caused this bug

Low

  • litellm.proxy.hooks.max_budget_limiter import path is gone, no in-repo users remained
  • Over-budget paid requests now always report budget_exceeded, never throttling_error
  • A free model with a paid fallback is not left unguarded: fix!: re-check budget on router fallback targets #41379 (on main since this branch was rebased) re-checks the key and user budget on each fallback target before it is attempted

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Link to Devin session: https://app.devin.ai/sessions/4f219316eeb043c2a65ecc1c4b46e208
Open in Devin Desktop: https://app.devin.ai/desktop/session/4f219316eeb043c2a65ecc1c4b46e208?variant=devin

@devin-ai-integration
devin-ai-integration Bot requested a review from a team September 16, 2026 01:09
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR. Add '(aside)' to your comment to have me ignore it.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@codspeed

codspeed Bot commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_remove_duplicate_user_budget_hook (789b828) with main (680233f)

Open in CodSpeed

@greptile-apps

greptile-apps Bot commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR removes the duplicate user-budget pre-call hook so zero-cost models remain available to users whose personal budget is exhausted, while the authentication and fallback-budget paths continue rejecting priced models.

  • Deletes _PROXY_MaxBudgetLimiter and removes its hook registration and runtime references.
  • Adds an authentication-layer regression test covering zero-cost allowance and priced-model rejection.
  • Updates hook lifecycle tests, rate-limit tests, architecture documentation, and related references.
  • The rebase incorporates the fallback path’s model-aware budget gate.

Confidence Score: 4/5

The budget behavior appears correct, but the unresolved repository requirement against removing the existing hook compatibility path without a user-controlled flag must be addressed before merging.

The new authentication regression coverage addresses the prior test finding, and the rebased fallback gate preserves model-aware key and user budget enforcement. However, the existing unresolved compatibility finding remains: max_budget_limiter is still removed from its import path and hook registry without a migration flag.

Files Needing Attention: litellm/proxy/hooks/init.py, litellm/proxy/hooks/max_budget_limiter.py

Important Files Changed
Filename Overview
litellm/proxy/hooks/init.py Removes the duplicate budget limiter from the proxy hook registry, but the existing compatibility-rule finding remains outstanding because the public hook path disappears without a flag.
litellm/proxy/hooks/max_budget_limiter.py Deletes the redundant user-budget hook whose model-agnostic check blocked zero-cost requests.
litellm/proxy/utils.py Removes construction and imports of the deleted limiter from proxy logging initialization.
tests/test_litellm/proxy/auth/test_auth_checks.py Verifies that the authentication budget gate permits an exhausted user on a zero-cost model while rejecting the same user on a priced model.
tests/test_litellm/proxy/utils/proxy_logging/test_pre_call_hook.py Confirms the registered pre-call hook set no longer independently enforces personal user budgets.

Reviews (3): Last reviewed commit: "test(auth): cover over-budget user on ze..." | Re-trigger Greptile

Comment on lines 19 to 20
PROXY_HOOKS: Final = {
"max_budget_limiter": _PROXY_MaxBudgetLimiter,
"parallel_request_limiter": _PROXY_MaxParallelRequestsHandler_v3,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Compatibility path removed

Deleting the import and registry key without a flag violates the repository directive to avoid backward-incompatible changes without user-controlled flags.

Rule Used: What: avoid backwards-incompatible changes without user-controlled flags Why: This breaks current behaviour for users using existing functionality Example of BAD: this PR (#22164) introduced run_post_custom... (source)

Comment on lines +403 to +431
@pytest.mark.asyncio
async def test_registered_hooks_do_not_enforce_user_budget(proxy_logging, monkeypatch):
"""
Personal budget is auth's job (`_user_max_budget_check`), which exempts
zero-cost models. A hook re-checking the same counter without that
exemption is what 429'd free models once a user was over budget.
"""
monkeypatch.setattr(litellm, "callbacks", [])
with patch("litellm.proxy.proxy_server.prisma_client", None):
proxy_logging._add_proxy_hooks(llm_router=None)
ProxyLogging._callback_capabilities_cache.clear()

over_budget_user = UserAPIKeyAuth(
api_key="sk-personal",
user_id="user-over-budget",
user_max_budget=1.0,
user_spend=5.0,
team_id=None,
)
data = {"model": "free-model", "messages": [{"role": "user", "content": "hi"}]}

with patch("litellm.proxy.proxy_server.get_current_spend", new=AsyncMock(return_value=5.0)):
out = await proxy_logging.pre_call_hook(
user_api_key_dict=over_budget_user,
data=data,
call_type="completion",
)

assert out == data

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Regression test misses auth

This only calls the hook layer, so it never proves auth allows free models while blocking paid models. Repository test coverage requires this before merging.

Rule Used: What: Flag any modifications to existing tests and verify they don't weaken test coverage or mask regressions. Why: Developers may alter tests to make failing code pass rather than fix the actual bug, hiding regressions. Good: ``` // Test updated t... (source)

Knowledge Base Used: Spend, budgets, and rate limits

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

@codecov

codecov Bot commented Sep 16, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai please re-review: added an auth-level regression test for free vs paid models and documented the custom auth caveat in the description

ryan-crabbe-berri and others added 2 commits September 16, 2026 15:42
…dels

_PROXY_MaxBudgetLimiter re-checked spend:user:{id} against user_max_budget in
async_pre_call_hook without the zero-cost model exemption that
_user_max_budget_check applies in auth, so free models were rejected with
"Max budget limit reached." once a user was over budget. Auth already owns
this check, so the hook is deleted rather than taught the exemption again

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…on_checks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@ryan-crabbe-berri
ryan-crabbe-berri force-pushed the litellm_remove_duplicate_user_budget_hook branch from 6a93a5f to dae1626 Compare September 16, 2026 22:44
@ryan-crabbe-berri

Copy link
Copy Markdown
Contributor

@greptileai rebased on latest main, please re-review. #41379 landed first, so the fallback path now has its own budget gate.

@ryan-crabbe-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit dae1626. Configure here.

@ryan-crabbe-berri

Copy link
Copy Markdown
Contributor

Not adding a shim. _PROXY_MaxBudgetLimiter is underscore-private and registered internally only, so a compat path would preserve the duplicate gate we are deleting.

@ryan-crabbe-berri ryan-crabbe-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM; thanks!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant