Repository navigation
fix(proxy): preserve spend metadata on auth failures - #38493
Open
RealJasonHu wants to merge 12 commits into
Open
RealJasonHu wants to merge 12 commits into
RealJasonHu wants to merge 12 commits into
Conversation
…ticated requests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…stency Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…nd Model Hub UI Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…mutation Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Contributor
Greptile SummaryThe PR preserves valid caller-supplied spend metadata when authentication fails while safely handling malformed headers and requests without a headers scope.
Confidence Score: 5/5The PR appears safe to merge. No blocking failure remains.
|
| Filename | Overview |
|---|---|
| litellm/proxy/auth/auth_exception_handler.py | Safely reads request headers and enriches authentication-failure logging with validated spend metadata and requester IP information. |
| litellm/proxy/litellm_pre_call_utils.py | Exposes a typed parser that accepts JSON-object spend metadata and safely ignores malformed or incompatible header values. |
| tests/test_litellm/proxy/auth/test_auth_exception_handler.py | Adds focused auth-failure regression tests, including the missing-headers-scope case requested in the follow-up. |
Reviews (2): Last reviewed commit: "fix(proxy): safely read auth failure hea..." | Re-trigger Greptile
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
Author
|
@greptileai Please re-review the latest commit, which safely handles requests without a headers scope and adds focused regression coverage |
Contributor
…ow read fails Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…iption feat(model_hub): surface model_info.description in Model Hub
…ip_save_on_failed_read fix(health): skip background health check DB writes when the latest-row read fails
feat(dd_span_tagger): emit litellm.user_email span tag for JWT-authenticated requests
devin-ai-integration
Bot
changed the base branch from
litellm_internal_staging
to
main
September 23, 2026 14:54
devin-ai-integration
Bot
requested review from
ryan-crabbe-berri and
yuneng-berri
as code owners
September 23, 2026 14:54
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
TLDR
Problem this solves:
How it solves it:
User Flow
Before: a gateway operator cannot correlate a budget rejection with the request ID supplied by their client
x-litellm-spend-logs-metadata: {"my_request_id":"req_rejected"}type: "budget_exceeded"spend_logs_metadata: nullAfter: the same budget rejection is searchable by the request ID supplied by their client
x-litellm-spend-logs-metadata: {"my_request_id":"req_rejected"}type: "budget_exceeded"spend_logs_metadata: {"my_request_id":"req_rejected"}Relevant issues
Fixes #37260
Affected release
Linear ticket
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you are seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Ran on September 22, 2026 (Asia/Shanghai), using a live proxy on
http://127.0.0.1:14093, PostgreSQL 17.11,STORE_MODEL_IN_DB=True, andproxy_batch_write_at: 1. Each revision used a separate clean database and the same configuration and date window. The configured model alias wasauth-metadata-repro. These requests fail during authentication, before any provider invocation, so no billable LLM call appliesShared
config.yamlused for both revisions, withDATABASE_URLpointing to that revision's clean databaseThe database was initialized from each checkout's
litellm/proxy/schema.prisma. Port 15993 had no listener; the provider was never needed because authentication rejected every requestBefore each run, created a virtual key with the following command, setting
PHASEtobeforeorafter, and stored the returned key asKEYBoth runs returned HTTP 200 with
"spend": 1.0and"max_budget": 0.01The log query below returned HTTP 200 and three failure rows per run after the background flush. Define this URL once and use the same query for each case
LOGS_URL='http://127.0.0.1:14093/spend/logs/v2?start_date=2026-09-20%2000%3A00%3A00&end_date=2026-09-22%2023%3A59%3A59&page=1&page_size=100'Before (252c71c)
Chat Completions
Send the over-budget request with caller metadata
Observe HTTP 429 with error fields
"type": "budget_exceeded"and"code": "429"Read the persisted failure log
Matching row, showing the relevant fields:
{"status":"failure","call_type":"acompletion","metadata":{"spend_logs_metadata":null}}Responses
Send the over-budget request with caller metadata
Observe HTTP 429 with error fields
"type": "budget_exceeded"and"code": "429"Read the persisted failure log
Matching row, showing the relevant fields:
{"status":"failure","call_type":"aresponses","metadata":{"spend_logs_metadata":null}}Anthropic Messages
Send the over-budget request with caller metadata
Observe HTTP 429 with error fields
"type": "budget_exceeded"and"code": "429"Read the persisted failure log
Matching row, showing the relevant fields:
{"status":"failure","call_type":"anthropic_messages","metadata":{"spend_logs_metadata":null}}After (033fe54)
Chat Completions
Send the over-budget request with caller metadata
Observe HTTP 429 with error fields
"type": "budget_exceeded"and"code": "429"Read the persisted failure log
Matching row, showing the relevant fields:
{"status":"failure","call_type":"acompletion","metadata":{"spend_logs_metadata":{"my_request_id":"req_rejected_after_chat"}}}Responses
Send the over-budget request with caller metadata
Observe HTTP 429 with error fields
"type": "budget_exceeded"and"code": "429"Read the persisted failure log
Matching row, showing the relevant fields:
{"status":"failure","call_type":"aresponses","metadata":{"spend_logs_metadata":{"my_request_id":"req_rejected_after_responses"}}}Anthropic Messages
Send the over-budget request with caller metadata
Observe HTTP 429 with error fields
"type": "budget_exceeded"and"code": "429"Read the persisted failure log
Matching row, showing the relevant fields:
{"status":"failure","call_type":"anthropic_messages","metadata":{"spend_logs_metadata":{"my_request_id":"req_rejected_after_messages"}}}Type
Bug Fix
Validation
At
033fe540, the auth handler suite passes 53 cases and the request setup suite passes 373 cases. Related auth, rate-limit, and invalid-key metrics suites pass 304 cases; their remaining expired-key test passes withTZ=UTC, matching CI. That test also fails on the unchanged base in UTC+8 because it constructs a naive local expiry timestampRepository-wide Ruff checks, formatting of changed production files, strict-rule, type-discipline and test-quality gates, circular-import and import-safety checks pass. The e2e type checker reports zero errors. The canonical full lint check, including the core type-check gate, also passes on this commit. Auth, proxy-endpoint, key-generation, JWT, proxy-core, MCP and proxy-behavior checks have passed in Actions; the remaining jobs and final coverage result are still running
Caveats (if any)
The PR still targets
litellm_internal_staging. Its unrelated documentation parser and dependency scan failures require upstream baseline updates. Vertex's four cost assertions also fail on that base. A timing-sensitive database test has an upstream fix in #40996 that the target branch does not contain. Current-tip CI is still running, so the all-checks and review attestations remain uncheckedFinal Attestation