Skip to content

fix(proxy): preserve spend metadata on auth failures - #38493

Open
RealJasonHu wants to merge 12 commits into
BerriAI:mainfrom
RealJasonHu:bugfix/auth-spend-metadata
Open

RealJasonHu wants to merge 12 commits into
BerriAI:mainfrom
RealJasonHu:bugfix/auth-spend-metadata

Conversation

@RealJasonHu

@RealJasonHu RealJasonHu commented Aug 27, 2026 •

Copy link
Copy Markdown

TLDR

Problem this solves:

  • Budget-rejected requests drop caller-provided spend metadata
  • Operators cannot correlate auth failures with gateway request IDs

How it solves it:

  • Preserve valid spend metadata before failure logging
  • Keep malformed headers and existing request metadata safe
  • Preserve upstream authentication errors and invalid-key logging markers

User Flow

Before: a gateway operator cannot correlate a budget rejection with the request ID supplied by their client

  1. They send POST https://litellm-domain/v1/chat/completions with an over-budget virtual key and x-litellm-spend-logs-metadata: {"my_request_id":"req_rejected"}
  2. The proxy returns HTTP 429 with type: "budget_exceeded"
  3. They send GET https://litellm-domain/spend/logs/v2 and find the failure row with spend_logs_metadata: null

After: the same budget rejection is searchable by the request ID supplied by their client

  1. They send the same POST https://litellm-domain/v1/chat/completions with an over-budget virtual key and x-litellm-spend-logs-metadata: {"my_request_id":"req_rejected"}
  2. The proxy returns the same HTTP 429 with type: "budget_exceeded"
  3. They send the same GET https://litellm-domain/spend/logs/v2 and find the failure row with spend_logs_metadata: {"my_request_id":"req_rejected"}

Relevant issues

Fixes #37260

Affected release

Linear ticket

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you are seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Ran on September 22, 2026 (Asia/Shanghai), using a live proxy on http://127.0.0.1:14093, PostgreSQL 17.11, STORE_MODEL_IN_DB=True, and proxy_batch_write_at: 1. Each revision used a separate clean database and the same configuration and date window. The configured model alias was auth-metadata-repro. These requests fail during authentication, before any provider invocation, so no billable LLM call applies

Shared config.yaml used for both revisions, with DATABASE_URL pointing to that revision's clean database

model_list:
  - model_name: auth-metadata-repro
    litellm_params:
      model: openai/gpt-5
      api_key: sk-local-unused-provider-key
      api_base: http://127.0.0.1:15993/v1
litellm_settings:
  drop_params: true
  set_verbose: false
general_settings:
  master_key: sk-local-auth-metadata-master
  database_url: os.environ/DATABASE_URL
  proxy_batch_write_at: 1
  disable_prisma_schema_update: true

The database was initialized from each checkout's litellm/proxy/schema.prisma. Port 15993 had no listener; the provider was never needed because authentication rejected every request

Before each run, created a virtual key with the following command, setting PHASE to before or after, and stored the returned key as KEY

curl -sS -X POST http://127.0.0.1:14093/key/generate \
  -H 'Authorization: Bearer sk-local-auth-metadata-master' \
  -H 'Content-Type: application/json' \
  -d "{\"key_alias\":\"budget-repro-$PHASE-20260922\",\"models\":[\"auth-metadata-repro\"],\"spend\":1,\"max_budget\":0.01}"

Both runs returned HTTP 200 with "spend": 1.0 and "max_budget": 0.01

The log query below returned HTTP 200 and three failure rows per run after the background flush. Define this URL once and use the same query for each case

LOGS_URL='http://127.0.0.1:14093/spend/logs/v2?start_date=2026-09-20%2000%3A00%3A00&end_date=2026-09-22%2023%3A59%3A59&page=1&page_size=100'

Before (252c71c)

Chat Completions

  1. Send the over-budget request with caller metadata

    curl -sS -i -X POST http://127.0.0.1:14093/v1/chat/completions \
      -H "Authorization: Bearer $KEY" \
      -H 'Content-Type: application/json' \
      -H 'x-litellm-spend-logs-metadata: {"my_request_id":"req_rejected_before_chat"}' \
      -d '{"model":"auth-metadata-repro","messages":[{"role":"user","content":"hi"}]}'
  2. Observe HTTP 429 with error fields "type": "budget_exceeded" and "code": "429"

  3. Read the persisted failure log

    curl -sS "$LOGS_URL" -H 'Authorization: Bearer sk-local-auth-metadata-master'

    Matching row, showing the relevant fields:

    {"status":"failure","call_type":"acompletion","metadata":{"spend_logs_metadata":null}}

Responses

  1. Send the over-budget request with caller metadata

    curl -sS -i -X POST http://127.0.0.1:14093/v1/responses \
      -H "Authorization: Bearer $KEY" \
      -H 'Content-Type: application/json' \
      -H 'x-litellm-spend-logs-metadata: {"my_request_id":"req_rejected_before_responses"}' \
      -d '{"model":"auth-metadata-repro","input":"hi"}'
  2. Observe HTTP 429 with error fields "type": "budget_exceeded" and "code": "429"

  3. Read the persisted failure log

    curl -sS "$LOGS_URL" -H 'Authorization: Bearer sk-local-auth-metadata-master'

    Matching row, showing the relevant fields:

    {"status":"failure","call_type":"aresponses","metadata":{"spend_logs_metadata":null}}

Anthropic Messages

  1. Send the over-budget request with caller metadata

    curl -sS -i -X POST http://127.0.0.1:14093/v1/messages \
      -H "Authorization: Bearer $KEY" \
      -H 'Content-Type: application/json' \
      -H 'x-litellm-spend-logs-metadata: {"my_request_id":"req_rejected_before_messages"}' \
      -d '{"model":"auth-metadata-repro","max_tokens":16,"messages":[{"role":"user","content":"hi"}]}'
  2. Observe HTTP 429 with error fields "type": "budget_exceeded" and "code": "429"

  3. Read the persisted failure log

    curl -sS "$LOGS_URL" -H 'Authorization: Bearer sk-local-auth-metadata-master'

    Matching row, showing the relevant fields:

    {"status":"failure","call_type":"anthropic_messages","metadata":{"spend_logs_metadata":null}}

After (033fe54)

Chat Completions

  1. Send the over-budget request with caller metadata

    curl -sS -i -X POST http://127.0.0.1:14093/v1/chat/completions \
      -H "Authorization: Bearer $KEY" \
      -H 'Content-Type: application/json' \
      -H 'x-litellm-spend-logs-metadata: {"my_request_id":"req_rejected_after_chat"}' \
      -d '{"model":"auth-metadata-repro","messages":[{"role":"user","content":"hi"}]}'
  2. Observe HTTP 429 with error fields "type": "budget_exceeded" and "code": "429"

  3. Read the persisted failure log

    curl -sS "$LOGS_URL" -H 'Authorization: Bearer sk-local-auth-metadata-master'

    Matching row, showing the relevant fields:

    {"status":"failure","call_type":"acompletion","metadata":{"spend_logs_metadata":{"my_request_id":"req_rejected_after_chat"}}}

Responses

  1. Send the over-budget request with caller metadata

    curl -sS -i -X POST http://127.0.0.1:14093/v1/responses \
      -H "Authorization: Bearer $KEY" \
      -H 'Content-Type: application/json' \
      -H 'x-litellm-spend-logs-metadata: {"my_request_id":"req_rejected_after_responses"}' \
      -d '{"model":"auth-metadata-repro","input":"hi"}'
  2. Observe HTTP 429 with error fields "type": "budget_exceeded" and "code": "429"

  3. Read the persisted failure log

    curl -sS "$LOGS_URL" -H 'Authorization: Bearer sk-local-auth-metadata-master'

    Matching row, showing the relevant fields:

    {"status":"failure","call_type":"aresponses","metadata":{"spend_logs_metadata":{"my_request_id":"req_rejected_after_responses"}}}

Anthropic Messages

  1. Send the over-budget request with caller metadata

    curl -sS -i -X POST http://127.0.0.1:14093/v1/messages \
      -H "Authorization: Bearer $KEY" \
      -H 'Content-Type: application/json' \
      -H 'x-litellm-spend-logs-metadata: {"my_request_id":"req_rejected_after_messages"}' \
      -d '{"model":"auth-metadata-repro","max_tokens":16,"messages":[{"role":"user","content":"hi"}]}'
  2. Observe HTTP 429 with error fields "type": "budget_exceeded" and "code": "429"

  3. Read the persisted failure log

    curl -sS "$LOGS_URL" -H 'Authorization: Bearer sk-local-auth-metadata-master'

    Matching row, showing the relevant fields:

    {"status":"failure","call_type":"anthropic_messages","metadata":{"spend_logs_metadata":{"my_request_id":"req_rejected_after_messages"}}}

Type

Bug Fix

Validation

At 033fe540, the auth handler suite passes 53 cases and the request setup suite passes 373 cases. Related auth, rate-limit, and invalid-key metrics suites pass 304 cases; their remaining expired-key test passes with TZ=UTC, matching CI. That test also fails on the unchanged base in UTC+8 because it constructs a naive local expiry timestamp

Repository-wide Ruff checks, formatting of changed production files, strict-rule, type-discipline and test-quality gates, circular-import and import-safety checks pass. The e2e type checker reports zero errors. The canonical full lint check, including the core type-check gate, also passes on this commit. Auth, proxy-endpoint, key-generation, JWT, proxy-core, MCP and proxy-behavior checks have passed in Actions; the remaining jobs and final coverage result are still running

Caveats (if any)

The PR still targets litellm_internal_staging. Its unrelated documentation parser and dependency scan failures require upstream baseline updates. Vertex's four cost assertions also fail on that base. A timing-sensitive database test has an upstream fix in #40996 that the target branch does not contain. Current-tip CI is still running, so the all-checks and review attestations remain unchecked

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

milan-berri and others added 5 commits July 28, 2026 14:14
…ticated requests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…stency

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…nd Model Hub UI

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…mutation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@CLAassistant

CLAassistant commented Aug 27, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@greptile-apps

greptile-apps Bot commented Aug 27, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

The PR preserves valid caller-supplied spend metadata when authentication fails while safely handling malformed headers and requests without a headers scope.

  • Adds spend metadata to auth-failure logging without mutating the original request body.
  • Centralizes validation of the spend-metadata header as a JSON object.
  • Adds focused regression coverage for valid, malformed, null, absent-scope, and existing-body-metadata cases.

Confidence Score: 5/5

The PR appears safe to merge.

No blocking failure remains.

Important Files Changed

Filename Overview
litellm/proxy/auth/auth_exception_handler.py Safely reads request headers and enriches authentication-failure logging with validated spend metadata and requester IP information.
litellm/proxy/litellm_pre_call_utils.py Exposes a typed parser that accepts JSON-object spend metadata and safely ignores malformed or incompatible header values.
tests/test_litellm/proxy/auth/test_auth_exception_handler.py Adds focused auth-failure regression tests, including the missing-headers-scope case requested in the follow-up.

Reviews (2): Last reviewed commit: "fix(proxy): safely read auth failure hea..." | Re-trigger Greptile

@codecov

codecov Bot commented Aug 27, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@RealJasonHu

Copy link
Copy Markdown
Author

@greptileai Please re-review the latest commit, which safely handles requests without a headers scope and adds focused regression coverage

@codspeed

codspeed Bot commented Aug 27, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing RealJasonHu:bugfix/auth-spend-metadata (033fe54) with litellm_internal_staging (252c71c)

Open in CodSpeed

@yuneng-berri
yuneng-berri deleted the branch BerriAI:main September 13, 2026 04:45
@yuneng-berri yuneng-berri reopened this Sep 13, 2026
yassin-berriai and others added 6 commits September 13, 2026 09:44
…ow read fails

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…iption

feat(model_hub): surface model_info.description in Model Hub
…ip_save_on_failed_read

fix(health): skip background health check DB writes when the latest-row read fails
feat(dd_span_tagger): emit litellm.user_email span tag for JWT-authenticated requests
@devin-ai-integration
devin-ai-integration Bot changed the base branch from litellm_internal_staging to main September 23, 2026 14:54
@devin-ai-integration
devin-ai-integration Bot requested a review from a team September 23, 2026 14:54

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Auth-stage failure spend logs (budget_exceeded) drop x-litellm-spend-logs-metadata even though the header is present

6 participants