Skip to content

feat(proxy): record in spend logs whether a request used a client-forwarded Anthropic OAuth token - #43063

Merged
mateo-berri merged 8 commits into
mainfrom
litellm_spend_log_client_oauth_flag
Oct 1, 2026
Merged

mateo-berri merged 8 commits into
mainfrom
litellm_spend_log_client_oauth_flag

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Spend logs cannot tell a Max-seat request from a configured-key one
  • Admins cannot reconcile the Anthropic API bill against the gateway logs
  • Nothing on the Logs page says which credential paid for a row

How it solves it:

  • Every spend log row gets metadata.used_client_oauth_token, true or false
  • Stamped when the proxy forwards the token, resolved at log time against the provider the call went to
  • A caller's own metadata.used_client_oauth_token is overwritten by the proxy's stamp
  • Rows keep the flag when a guardrail adds its own metadata bucket to a chat request
  • /spend/logs/ui (and v2) take used_client_oauth_token=true|false
  • Logs page gains a Credential filter and shows the value in the row drawer
  • The token itself never reaches the log, only the boolean

Intentional product change: the Logs page Filters drawer gains a Credential select (All Credentials, Client OAuth token, Configured key) and the row drawer's Request Details gains a Credential field, so admins can split seat-billed rows from key-billed ones; nothing existing moves or goes away

User Flow

Before: an admin whose developers run Claude Code on Max seats through the gateway cannot tell which spend log rows a seat paid for, so the API bill they reconcile from the logs is wrong

  1. A developer signed into Claude Code with a Max login points it at the gateway (ANTHROPIC_BASE_URL=https://litellm-domain, ANTHROPIC_CUSTOM_HEADERS="x-litellm-api-key: Bearer sk-...") and their prompt gets a reply
  2. The admin sends POST https://litellm-domain/v1/messages with Authorization: Bearer sk-... alone, so the deployment's configured Anthropic key pays, and gets a reply too
  3. The admin opens https://litellm-domain/ui/?page=logs: both rows show list-price spend and nothing on either row or its detail drawer says which credential paid
  4. GET https://litellm-domain/spend/logs/ui?used_client_oauth_token=true returns 200 with both rows, the parameter is ignored

Filters drawer at https://litellm-domain/ui/?page=logs: Cache is followed straight by Key Alias, no Credential filter

ui_before_filters

Row drawer of the Claude Code request: Request Details ends at IP Address, no Credential field

ui_before_row_detail

After: the same two rows differ on metadata.used_client_oauth_token, so the admin can list exactly the seat-billed requests

  1. A developer signed into Claude Code with a Max login points it at the gateway (ANTHROPIC_BASE_URL=https://litellm-domain, ANTHROPIC_CUSTOM_HEADERS="x-litellm-api-key: Bearer sk-...") and their prompt gets a reply
  2. The admin sends POST https://litellm-domain/v1/messages with Authorization: Bearer sk-... alone, so the deployment's configured Anthropic key pays, and gets a reply too
  3. The admin opens https://litellm-domain/ui/?page=logs, picks Credential: Client OAuth token, and only the Claude Code row stays; its detail drawer reads Credential: Client OAuth token, the curl row's reads Configured key, both still at list-price spend
  4. GET https://litellm-domain/spend/logs/ui?used_client_oauth_token=true returns 200 with only the Claude Code row, metadata.used_client_oauth_token: true (false on the curl row even when its body claimed true, absent on rows written before the upgrade), and no part of the token appears in either row

Filters drawer at https://litellm-domain/ui/?page=logs: a Credential select sits between Cache and Key Alias

ui_after_filters

Row drawer of the Claude Code request: Request Details gains Credential: Client OAuth token

ui_after_row_detail_success

Relevant issues

Related to #23755

Linear ticket

Resolves LIT-8589

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Both legs booted python litellm/proxy/proxy_cli.py --config config.yaml --num_workers 2 --use_v2_migration_resolver, two uvicorn workers on their own Postgres database each, with real Anthropic calls. Before is the merge base 317430d on port 31734 and After is this PR's tip 79bda61 on port 58912, each with its own commit's dashboard served by next dev against it. The After leg also turned on the generic_api callback, posting to a local recorder, so the callback payload is checked next to the spend rows. The After config is below, and Before ran the same file without the litellm_settings block

model_list:
  - model_name: claude-sonnet-5-5
    litellm_params:
      model: anthropic/claude-sonnet-5-5
      api_key: os.environ/ANTHROPIC_API_KEY
guardrails:
  - guardrail_name: qa-content-filter
    litellm_params:
      guardrail: litellm_content_filter
      mode: pre_call
      default_on: false
      blocked_words:
        - keyword: qablockedword
          action: BLOCK
litellm_settings:
  callbacks: ["generic_api"]
general_settings:
  master_key: os.environ/LITELLM_MASTER_KEY
  forward_client_headers_to_llm_api: true
  store_prompts_in_spend_logs: true

Two credentials drive every case: $LITELLM_MASTER_KEY alone, so the deployment's configured Anthropic key pays, and $OAUTH_TOKEN, the sk-ant-oat... access token of a Claude Max seat (read from Claude Code's credential store), sent as the bearer with x-litellm-api-key: Bearer $LITELLM_MASTER_KEY alongside, which is how Claude Code itself sends it. The Claude Code case runs the real TUI (v2.1.284, signed into a Max seat) under tmux; the curl cases replay the same prompt on every unified endpoint

Before (317430d)

Merge base proxy on port 31734 with --num_workers 2 and one Postgres database, claude-sonnet-5-5 routed to Anthropic with the configured key

Claude Code on /v1/messages with the Max seat

  1. export ANTHROPIC_BASE_URL=http://127.0.0.1:31734; export ANTHROPIC_CUSTOM_HEADERS="x-litellm-api-key: Bearer $LITELLM_MASTER_KEY"; claude --model claude-sonnet-5-5 (Claude Code v2.1.284, Sonnet 5.5 · Claude Max), then type Reply with the single word pong. and press Enter
  2. The pane shows pong (Crunched for 1s, done 2:55 PM), and the proxy logs the turn as rows msg_011CfWakZsi4AjWqQcmgxeVX and msg_011CfWakb7NjMBqvv23pSCcv, both ua=claude-cli/2.1.284 in the listing below. Row aff94a49 right before them is a request Claude Code sent at startup, before the prompt was typed, that Anthropic answered with 429; the Logs page groups it into the same Claude Code session

before_claude_pane

curl /v1/messages

  1. Configured key:

    curl -s -w '\n%{http_code}' http://127.0.0.1:31734/v1/messages -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
      -H "anthropic-version: 2023-06-01" -H "content-type: application/json" \
      -d '{"model":"claude-sonnet-5-5","max_tokens":16,"messages":[{"role":"user","content":"Reply with the single word pong."}]}' \
      | jq -c '{id, type, text: .content[0].text, stop_reason}'
    HTTP 200
    {"id":"msg_011CfWaneseBciUVz9JDQrs8","type":"message","text":"pong","stop_reason":"end_turn"}
    
  2. Seat token (-H "Authorization: Bearer $OAUTH_TOKEN" -H "x-litellm-api-key: Bearer $LITELLM_MASTER_KEY", same body):

    HTTP 429
    {"type":"error","error":{"type":"rate_limit_error","message":"litellm.RateLimitError: AnthropicException - {\"type\":\"error\",\"error\":{\"type\":\"rate_limit_error\",\"message\":\"Error\"},\"request_id\":\"req_011CfWanvJptmNAuBJGEWZMJ\"}\n\nLiteLLM: model group 'claude-sonnet-5-5' failed with the error above. No fallback was attempted."}}
    

curl /v1/chat/completions

  1. Configured key:

    curl -s -w '\n%{http_code}' http://127.0.0.1:31734/v1/chat/completions -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
      -H "content-type: application/json" \
      -d '{"model":"claude-sonnet-5-5","max_tokens":16,"messages":[{"role":"user","content":"Reply with the single word pong."}]}' \
      | jq -c '{id, object, text: .choices[0].message.content, finish_reason: .choices[0].finish_reason}'
    HTTP 200
    {"id":"chatcmpl-207f402b-c444-4ddd-aff5-2ef82027622d","object":"chat.completion","text":"pong","finish_reason":"stop"}
    
  2. Seat token (same body, bearer $OAUTH_TOKEN plus x-litellm-api-key):

    HTTP 429
    {"error":{"message":"litellm.RateLimitError: AnthropicException - {\"type\":\"error\",\"error\":{\"type\":\"rate_limit_error\",\"message\":\"Error\"},\"request_id\":\"req_011CfWaoSwa1p1Sr7AVvVVZp\"}\n\nLiteLLM: model group 'claude-sonnet-5-5' failed with the error above. No fallback was attempted.","type":"throttling_error","param":null,"code":"429"}}
    

curl /v1/responses

  1. Configured key:

    curl -s -w '\n%{http_code}' http://127.0.0.1:31734/v1/responses -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
      -H "content-type: application/json" \
      -d '{"model":"claude-sonnet-5-5","max_output_tokens":16,"input":"Reply with the single word pong."}' \
      | jq -c '{id: .id[:24], object, text: .output[-1].content[0].text, status}'
    HTTP 200
    {"id":"resp_9OzmfxWXq57qXsZ0dP5","object":"response","text":"pong","status":"completed"}
    
  2. Seat token (same body, bearer $OAUTH_TOKEN plus x-litellm-api-key):

    HTTP 429
    {"error":{"message":"litellm.RateLimitError: AnthropicException - {\"type\":\"error\",\"error\":{\"type\":\"rate_limit_error\",\"message\":\"Error\"},\"request_id\":\"req_011CfWap6hvvHNoSeoAFnHyR\"}\n\nLiteLLM: model group 'claude-sonnet-5-5' failed with the error above. No fallback was attempted.","type":"throttling_error","param":null,"code":"429"}}
    

curl /v1/messages claiming the flag in its own metadata

  1. Configured key, the body carrying "metadata":{"used_client_oauth_token":true} plus a max_tokens Anthropic rejects, so the row is a failure that spent nothing (same merge base proxy on port 31734):

    curl -s -w '\n%{http_code}' http://127.0.0.1:31734/v1/messages -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
      -H "anthropic-version: 2023-06-01" -H "content-type: application/json" \
      -d '{"model":"claude-sonnet-5-5","max_tokens":999999,"metadata":{"used_client_oauth_token":true},"messages":[{"role":"user","content":"Reply with the single word pong."}]}'
    HTTP 400
    {"type":"error","error":{"type":"invalid_request_error","message":"litellm.BadRequestError: AnthropicException - {\"type\":\"error\",\"error\":{\"type\":\"invalid_request_error\",\"message\":\"max_tokens: 999999 > 128000, which is the maximum allowed number of output tokens for claude-sonnet-5-5\"},\"request_id\":\"req_011CfWaq6kC1Py525QMPrsCW\"}\n\nLiteLLM: model group 'claude-sonnet-5-5' failed with the error above. No fallback was attempted."}}
    
  2. The listing command of the next case bounded to this call (start_date=2026-09-28%2021:56:16&end_date=2026-09-28%2021:56:22, port 31734): the row carries no used_client_oauth_token at all, so the body's claim is not stored

    HTTP 200
    rows: 1
    c3993573-e408-4e06-86a2-d0562225adf2 anthropic_messages failure spend=0.0 ua= used_client_oauth_token=null
    0
    0
    
  3. Same listing with &used_client_oauth_token=true: HTTP 200, rows: 1, the same row, the parameter is ignored

GET /spend/logs/ui

  1. All rows of the run window:

    Q="start_date=2026-09-28%2021:54:58&end_date=2026-09-28%2021:56:30&page=1&page_size=50"
    curl -s -w '\nHTTP %{http_code}\n' "http://127.0.0.1:31734/spend/logs/ui?$Q" -H "Authorization: Bearer $LITELLM_MASTER_KEY" > out.json
    tail -1 out.json; sed '$d' out.json | jq -r 'if .data then ("rows: \(.total)", (.data | sort_by(.startTime)[] | "\(.request_id[:36]) \(.call_type) \(.status) spend=\(.spend) ua=\(.metadata.user_agent // "" | .[:18]) used_client_oauth_token=\(.metadata.used_client_oauth_token)")) else . end'
    sed '$d' out.json | grep -c -i 'sk-ant-oat'
    sed '$d' out.json | grep -c -F "${OAUTH_TOKEN: -16}"
    HTTP 200
    rows: 10
    aff94a49-3e95-4455-8a4e-4d8698d05dea anthropic_messages failure spend=0.0 ua= used_client_oauth_token=null
    msg_011CfWakZsi4AjWqQcmgxeVX anthropic_messages success spend=0.00248 ua=claude-cli/2.1.284 used_client_oauth_token=null
    msg_011CfWakb7NjMBqvv23pSCcv anthropic_messages success spend=0.073592 ua=claude-cli/2.1.284 used_client_oauth_token=null
    msg_011CfWaneseBciUVz9JDQrs8 anthropic_messages success spend=0.000076 ua=curl/8.7.1 used_client_oauth_token=null
    c97b03e1-b4b7-4c9b-a4ae-7088beafb2ea anthropic_messages failure spend=0.0 ua= used_client_oauth_token=null
    chatcmpl-207f402b-c444-4ddd-aff5-2ef acompletion success spend=0.000076 ua=curl/8.7.1 used_client_oauth_token=null
    031390fd-6f11-4a7c-b444-73164b0711ed acompletion failure spend=0.0 ua=curl/8.7.1 used_client_oauth_token=null
    chatcmpl-21685ded-f04b-4f6a-99c6-399 aresponses success spend=0.000076 ua=curl/8.7.1 used_client_oauth_token=null
    ab9ba63a-e92e-4743-b118-5b718eb70480 aresponses failure spend=0.0 ua= used_client_oauth_token=null
    c3993573-e408-4e06-86a2-d0562225adf2 anthropic_messages failure spend=0.0 ua= used_client_oauth_token=null
    0
    0
    
  2. Same command with &used_client_oauth_token=true on the query: HTTP 200, rows: 10, the identical ten rows, every one used_client_oauth_token=null, so the parameter is ignored

  3. Same command with &used_client_oauth_token=false: HTTP 200, rows: 10, the identical ten rows

  4. Same command with &used_client_oauth_token=seat: HTTP 200, rows: 10, the identical ten rows, the junk value is not rejected

Logs page

This commit's dashboard served with NEXT_PUBLIC_BASE_URL=http://127.0.0.1:31734 npx next dev -H 127.0.0.1 -p 51779 from ui/litellm-dashboard, the proxy restarted on the same port with PROXY_BASE_URL=http://127.0.0.1:31734 so the dashboard calls it (the rows above live in Postgres and carry over), signed in as admin with the master key

  1. Open http://127.0.0.1:51779/?page=logs and click Filters: the drawer lists Team ID, Span Type, Status, Cache, Key Alias, User ID, End User, Error Code, Error Message, Key Hash, Session ID, Model and Public model / search tool, with no Credential filter (first Before screenshot under User Flow)
  2. Click the Claude Code row msg_011CfWakb7NjMBqvv23pSCcv (Tags User-Agent: claude-cli/2.1.284 (external, cli)): Request Details shows Model, Provider, Call Type, User, Model ID, API Base and IP Address, with no Credential field (second Before screenshot under User Flow)

After (79bda61)

This run checks the After leg at 79bda61, the commit that reads used_client_oauth_token from the bucket the route stamped, so a guardrail that adds litellm_metadata to a /v1/chat/completions request no longer leaves the flag null. Same topology as Before: two workers, a fresh Postgres database, port 58912, PROXY_BASE_URL=http://localhost:58912, and litellm_settings.callbacks: ["generic_api"] with GENERIC_LOGGER_ENDPOINT=http://127.0.0.1:46360/log, a local recorder that appends every JSON array the callback POSTs to recorded.ndjson. The config also has one litellm_content_filter guardrail for the new case at the end of the curl sections

Claude Code on /v1/messages with the Max seat

  1. export ANTHROPIC_BASE_URL=http://127.0.0.1:58912; export ANTHROPIC_CUSTOM_HEADERS="x-litellm-api-key: Bearer $LITELLM_MASTER_KEY"; claude --model claude-sonnet-5-5, then type Reply with the single word pong. and press Enter
  2. The pane shows pong (Churned for 4s, done 5:23 PM). The proxy logs the turn as rows msg_011CfWn2dNHSUfkGcsWmqpZZ and msg_011CfWn2mZQP74SgV6CSAynm, both ua=claude-cli/2.1.284 and used_client_oauth_token=true in the listing below. They come after the 429 row 360fe2aa, which Claude Code sent at launch before the prompt was typed

pr43063-79bda61212-after_claude_pane.png

curl /v1/messages

  1. Configured key (same command as Before, port 58912):

    HTTP 200
    {"id":"msg_011CfWn4LcXCoCEUyrbkHJBQ","type":"message","text":"pong","stop_reason":"end_turn"}
    
  2. Seat token (same command, bearer $OAUTH_TOKEN plus x-litellm-api-key):

    HTTP 429
    {"type":"error","error":{"type":"rate_limit_error","message":"litellm.RateLimitError: AnthropicException - {\"type\":\"error\",\"error\":{\"type\":\"rate_limit_error\",\"message\":\"Error\"},\"request_id\":\"req_011CfWn4nZyMEoChB6oL4Lo5\"}\n\nLiteLLM: model group 'claude-sonnet-5-5' failed with the error above. No fallback was attempted."}}
    

curl /v1/chat/completions

  1. Configured key (same command as Before, port 58912):

    HTTP 200
    {"id":"chatcmpl-197f83f4-af1c-44c5-8f12-617239bfbb7a","object":"chat.completion","text":"pong","finish_reason":"stop"}
    
  2. Seat token:

    HTTP 429
    {"error":{"message":"litellm.RateLimitError: AnthropicException - {\"type\":\"error\",\"error\":{\"type\":\"rate_limit_error\",\"message\":\"Error\"},\"request_id\":\"req_011CfWn5m1Eq3P6F6xqi69Wr\"}\n\nLiteLLM: model group 'claude-sonnet-5-5' failed with the error above. No fallback was attempted.","type":"throttling_error","param":null,"code":"429"}}
    

curl /v1/responses

  1. Configured key (same command as Before, port 58912):

    HTTP 200
    {"id":"resp_gGNrCKbZOvfzTWmpq5j","object":"response","text":"pong","status":"completed"}
    
  2. Seat token:

    HTTP 429
    {"error":{"message":"litellm.RateLimitError: AnthropicException - {\"type\":\"error\",\"error\":{\"type\":\"rate_limit_error\",\"message\":\"Error\"},\"request_id\":\"req_011CfWn6abT9sLLFFKLJtev3\"}\n\nLiteLLM: model group 'claude-sonnet-5-5' failed with the error above. No fallback was attempted.","type":"throttling_error","param":null,"code":"429"}}
    

curl /v1/messages claiming the flag in its own metadata

  1. Same command as Before, port 58912:

    HTTP 400
    {"type":"error","error":{"type":"invalid_request_error","message":"litellm.BadRequestError: AnthropicException - {\"type\":\"error\",\"error\":{\"type\":\"invalid_request_error\",\"message\":\"max_tokens: 999999 > 128000, which is the maximum allowed number of output tokens for claude-sonnet-5-5\"},\"request_id\":\"req_011CfWn7y9J9mVfNgzguorih\"}\n\nLiteLLM: model group 'claude-sonnet-5-5' failed with the error above. No fallback was attempted."}}
    
  2. It is row 202894d0 in the listings below. It reads used_client_oauth_token=false although the body claimed true, because the configured key made the call, and it answers to =false and not to =true

Successful calls claiming the flag, all three endpoints

This is the case Bugbot raised. Each call uses the configured key and puts "metadata":{"used_client_oauth_token":true} in the body:

curl -s -w '\n%{http_code}' http://127.0.0.1:58912/v1/messages \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" -H "anthropic-version: 2023-06-01" -H "content-type: application/json" \
  -d '{"model":"claude-sonnet-5-5","max_tokens":16,"metadata":{"used_client_oauth_token":true},"messages":[{"role":"user","content":"Reply with the single word pong."}]}'

curl -s -w '\n%{http_code}' http://127.0.0.1:58912/v1/responses \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" -H "content-type: application/json" \
  -d '{"model":"claude-sonnet-5-5","max_output_tokens":16,"metadata":{"used_client_oauth_token":true},"input":"Reply with the single word pong."}'

curl -s -w '\n%{http_code}' http://127.0.0.1:58912/v1/chat/completions \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" -H "content-type: application/json" \
  -d '{"model":"claude-sonnet-5-5","max_tokens":16,"metadata":{"used_client_oauth_token":true},"messages":[{"role":"user","content":"Reply with the single word pong."}]}'
HTTP 200
{"id":"msg_011CfWn8SQdnxRnQypRVdxCx","type":"message","text":"pong","stop_reason":"end_turn"}
HTTP 200
{"id":"resp_1nvbjFNnw87CMwEtWK7","object":"response","text":"pong","status":"completed"}
HTTP 200
{"id":"chatcmpl-b385fe01-f324-48c4-9aef-161099ccf177","object":"chat.completion","text":"pong","finish_reason":"stop"}

All three rows read used_client_oauth_token=false in the spend-log listings below (msg_011CfWn8SQ, chatcmpl-2bc9fe2b for the responses call, chatcmpl-b385fe01)

Content filter guardrail on /v1/chat/completions

This is the case the new commit fixes. The config adds one pre_call content filter that a request turns on by naming it:

guardrails:
  - guardrail_name: qa-content-filter
    litellm_params:
      guardrail: litellm_content_filter
      mode: pre_call
      default_on: false
      blocked_words:
        - keyword: qablockedword
          action: BLOCK
  1. Configured key and seat token, both asking for the guardrail:

    curl -s -w '\n%{http_code}' http://127.0.0.1:58912/v1/chat/completions \
      -H "Authorization: Bearer $LITELLM_MASTER_KEY" -H "content-type: application/json" \
      -d '{"model":"claude-sonnet-5-5","max_tokens":16,"guardrails":["qa-content-filter"],"messages":[{"role":"user","content":"Reply with the single word pong."}]}'
    
    curl -s -w '\n%{http_code}' http://127.0.0.1:58912/v1/chat/completions \
      -H "Authorization: Bearer $OAUTH_TOKEN" -H "x-litellm-api-key: Bearer $LITELLM_MASTER_KEY" -H "content-type: application/json" \
      -d '{"model":"claude-sonnet-5-5","max_tokens":16,"guardrails":["qa-content-filter"],"messages":[{"role":"user","content":"Reply with the single word pong."}]}'
    HTTP 200
    {"id":"chatcmpl-95565e95-7565-488b-94b2-eaf5506b4cd6","object":"chat.completion","text":"pong","finish_reason":"stop"}
    HTTP 429
    {"error":{"message":"litellm.RateLimitError: AnthropicException - {\"type\":\"error\",\"error\":{\"type\":\"rate_limit_error\",\"message\":\"Error\"},\"request_id\":\"req_011CfWn9wVdWNQCpDyZyQyJ3\"}\n\nLiteLLM: model group 'claude-sonnet-5-5' failed with the error above. No fallback was attempted.","type":"throttling_error","param":null,"code":"429"}}
    
  2. The same pair with "content":"Reply with the single word qablockedword.", so the guardrail blocks both before any call to Anthropic. Both answer the same way:

    HTTP 400
    {"error":{"message":"Content blocked: keyword 'qablockedword' detected","type":"invalid_request_error","param":null,"code":"400","provider_specific_fields":{"error":"Content blocked: keyword 'qablockedword' detected","keyword":"qablockedword","description":null,"guardrail_name":"qa-content-filter","guardrail_mode":"pre_call"}}}
    
  3. The four rows, from the all-rows listing below:

    sed '$d' out.json | jq -c '.data | sort_by(.startTime)[] | select(.metadata.guardrail_information != null) | {id: .request_id[:18], status, used_client_oauth_token: .metadata.used_client_oauth_token, guardrail: [.metadata.guardrail_information[] | {guardrail_name, guardrail_mode, guardrail_status}]}'
    {"id":"chatcmpl-95565e95-","status":"success","used_client_oauth_token":false,"guardrail":[{"guardrail_name":"qa-content-filter","guardrail_mode":"pre_call","guardrail_status":"success"}]}
    {"id":"981da83e-a52e-4330","status":"failure","used_client_oauth_token":true,"guardrail":[{"guardrail_name":"qa-content-filter","guardrail_mode":"pre_call","guardrail_status":"success"}]}
    {"id":"356ec199-db9e-42e7","status":"failure","used_client_oauth_token":false,"guardrail":[{"guardrail_name":"qa-content-filter","guardrail_mode":"pre_call","guardrail_status":"guardrail_intervened"}]}
    {"id":"5ece92ba-bf51-40be","status":"failure","used_client_oauth_token":true,"guardrail":[{"guardrail_name":"qa-content-filter","guardrail_mode":"pre_call","guardrail_status":"guardrail_intervened"}]}
    

    The guardrail ran on all four requests. The seat rows (981da83e for the pong call, 5ece92ba for the blocked call) read true and both come back under ?used_client_oauth_token=true. The configured-key rows (chatcmpl-95565e95 and 356ec199) read false and come back under =false. None of the four is null, and the generic_api payloads below carry the same values

generic_api callback payloads

Every standard logging payload the callback delivered in the run window, in time order:

jq -c --argjson ws 1790641372 --argjson we 1790641512 '[.[] | select(.startTime >= $ws and .startTime <= $we)] | sort_by(.startTime)[] | {id: .id[:36], call_type, status, used_client_oauth_token: .metadata.used_client_oauth_token, requester_metadata_claim: .metadata.requester_metadata.used_client_oauth_token}' recorded.ndjson
{"id":"360fe2aa-e5f8-4a3f-ae39-947790937f00","call_type":"anthropic_messages","status":"failure","used_client_oauth_token":true,"requester_metadata_claim":null}
{"id":"msg_011CfWn2dNHSUfkGcsWmqpZZ","call_type":"anthropic_messages","status":"success","used_client_oauth_token":true,"requester_metadata_claim":null}
{"id":"msg_011CfWn2mZQP74SgV6CSAynm","call_type":"anthropic_messages","status":"success","used_client_oauth_token":true,"requester_metadata_claim":null}
{"id":"msg_011CfWn4LcXCoCEUyrbkHJBQ","call_type":"anthropic_messages","status":"success","used_client_oauth_token":false,"requester_metadata_claim":null}
{"id":"207cd0f8-ce0b-47e3-b002-c8b23e3d9c51","call_type":"anthropic_messages","status":"failure","used_client_oauth_token":true,"requester_metadata_claim":null}
{"id":"5fd44b3b-10a8-4c0b-b155-27a095a07fab","call_type":"acompletion","status":"failure","used_client_oauth_token":true,"requester_metadata_claim":true}
{"id":"chatcmpl-197f83f4-af1c-44c5-8f12-617","call_type":"acompletion","status":"success","used_client_oauth_token":false,"requester_metadata_claim":false}
{"id":"chatcmpl-13e3cef1-c771-4b54-a48a-47b","call_type":"aresponses","status":"success","used_client_oauth_token":false,"requester_metadata_claim":null}
{"id":"41d6255b-291a-4a6d-b7ff-6c52b3e9b9f6","call_type":"aresponses","status":"failure","used_client_oauth_token":true,"requester_metadata_claim":null}
{"id":"202894d0-56da-4501-9534-965e3ec98dd1","call_type":"anthropic_messages","status":"failure","used_client_oauth_token":false,"requester_metadata_claim":true}
{"id":"msg_011CfWn8SQdnxRnQypRVdxCx","call_type":"anthropic_messages","status":"success","used_client_oauth_token":false,"requester_metadata_claim":true}
{"id":"chatcmpl-2bc9fe2b-5b8a-4f0e-9da1-24e","call_type":"aresponses","status":"success","used_client_oauth_token":false,"requester_metadata_claim":true}
{"id":"chatcmpl-b385fe01-f324-48c4-9aef-161","call_type":"acompletion","status":"success","used_client_oauth_token":false,"requester_metadata_claim":false}
{"id":"chatcmpl-95565e95-7565-488b-94b2-eaf","call_type":"acompletion","status":"success","used_client_oauth_token":false,"requester_metadata_claim":false}
{"id":"981da83e-a52e-4330-9b79-2617f70ce5c1","call_type":"acompletion","status":"failure","used_client_oauth_token":true,"requester_metadata_claim":true}
{"id":"356ec199-db9e-42e7-9e5f-10391c9b1b15","call_type":"acompletion","status":"failure","used_client_oauth_token":false,"requester_metadata_claim":false}
{"id":"5ece92ba-bf51-40be-98c4-f6f55f5ef83a","call_type":"acompletion","status":"failure","used_client_oauth_token":true,"requester_metadata_claim":true}

The four body claims (202894d0 and the three 200s) all read false in .metadata.used_client_oauth_token, the same as the spend-log rows. The eight true payloads are all seat-token requests: the three curl 429s, the guardrail 429 and the guardrail block, Claude Code's launch 429, and Claude Code's two successful calls

metadata.requester_metadata for the claims depends on the endpoint. On /v1/messages and /v1/responses it is {"used_client_oauth_token":true}, the caller's own value kept as sent. On /v1/chat/completions the claim row's requester_metadata is {"used_client_oauth_token":false,"headers":{host, user-agent, accept, content-type, content-length}}, so it holds the proxy stamp instead of the caller's true. On chat the PR writes the stamp into data["metadata"] before requester_metadata is copied from that same dict, so the stamp replaces the caller's key. The chat calls with no caller metadata get the same stamp there: false for chatcmpl-197f83f4 and the two configured-key guardrail rows, true for the seat 429 on chat (5fd44b3b) and the two seat guardrail rows. The headers part is what main already does

GET /spend/logs/ui

  1. All rows of the run window:

    Q="start_date=2026-09-29%2000:22:52&end_date=2026-09-29%2000:25:12&page=1&page_size=50"
    curl -s -w '\nHTTP %{http_code}\n' "http://127.0.0.1:58912/spend/logs/ui?$Q" -H "Authorization: Bearer $LITELLM_MASTER_KEY" > out.json
    tail -1 out.json; sed '$d' out.json | jq -r 'if .data then ("rows: \(.total)", (.data | sort_by(.startTime)[] | "\(.request_id[:36]) \(.call_type) \(.status) spend=\(.spend) ua=\(.metadata.user_agent // "" | .[:18]) used_client_oauth_token=\(.metadata.used_client_oauth_token)")) else . end'
    sed '$d' out.json | grep -c -i 'sk-ant-oat'
    HTTP 200
    rows: 17
    360fe2aa-e5f8-4a3f-ae39-947790937f00 anthropic_messages failure spend=0.0 ua= used_client_oauth_token=true
    msg_011CfWn2dNHSUfkGcsWmqpZZ anthropic_messages success spend=0.00248 ua=claude-cli/2.1.284 used_client_oauth_token=true
    msg_011CfWn2mZQP74SgV6CSAynm anthropic_messages success spend=0.074272 ua=claude-cli/2.1.284 used_client_oauth_token=true
    msg_011CfWn4LcXCoCEUyrbkHJBQ anthropic_messages success spend=0.000076 ua=curl/8.7.1 used_client_oauth_token=false
    207cd0f8-ce0b-47e3-b002-c8b23e3d9c51 anthropic_messages failure spend=0.0 ua= used_client_oauth_token=true
    chatcmpl-197f83f4-af1c-44c5-8f12-617 acompletion success spend=0.000076 ua=curl/8.7.1 used_client_oauth_token=false
    5fd44b3b-10a8-4c0b-b155-27a095a07fab acompletion failure spend=0.0 ua=curl/8.7.1 used_client_oauth_token=true
    chatcmpl-13e3cef1-c771-4b54-a48a-47b aresponses success spend=0.000076 ua=curl/8.7.1 used_client_oauth_token=false
    41d6255b-291a-4a6d-b7ff-6c52b3e9b9f6 aresponses failure spend=0.0 ua= used_client_oauth_token=true
    202894d0-56da-4501-9534-965e3ec98dd1 anthropic_messages failure spend=0.0 ua= used_client_oauth_token=false
    msg_011CfWn8SQdnxRnQypRVdxCx anthropic_messages success spend=0.000076 ua=curl/8.7.1 used_client_oauth_token=false
    chatcmpl-2bc9fe2b-5b8a-4f0e-9da1-24e aresponses success spend=0.000076 ua=curl/8.7.1 used_client_oauth_token=false
    chatcmpl-b385fe01-f324-48c4-9aef-161 acompletion success spend=0.000076 ua=curl/8.7.1 used_client_oauth_token=false
    chatcmpl-95565e95-7565-488b-94b2-eaf acompletion success spend=0.000076 ua= used_client_oauth_token=false
    981da83e-a52e-4330-9b79-2617f70ce5c1 acompletion failure spend=0.0 ua=curl/8.7.1 used_client_oauth_token=true
    356ec199-db9e-42e7-9e5f-10391c9b1b15 acompletion failure spend=0.0 ua=curl/8.7.1 used_client_oauth_token=false
    5ece92ba-bf51-40be-98c4-f6f55f5ef83a acompletion failure spend=0.0 ua=curl/8.7.1 used_client_oauth_token=true
    0
    
  2. Same command with &used_client_oauth_token=true:

    HTTP 200
    rows: 8
    360fe2aa-e5f8-4a3f-ae39-947790937f00 anthropic_messages failure spend=0.0 ua= used_client_oauth_token=true
    msg_011CfWn2dNHSUfkGcsWmqpZZ anthropic_messages success spend=0.00248 ua=claude-cli/2.1.284 used_client_oauth_token=true
    msg_011CfWn2mZQP74SgV6CSAynm anthropic_messages success spend=0.074272 ua=claude-cli/2.1.284 used_client_oauth_token=true
    207cd0f8-ce0b-47e3-b002-c8b23e3d9c51 anthropic_messages failure spend=0.0 ua= used_client_oauth_token=true
    5fd44b3b-10a8-4c0b-b155-27a095a07fab acompletion failure spend=0.0 ua=curl/8.7.1 used_client_oauth_token=true
    41d6255b-291a-4a6d-b7ff-6c52b3e9b9f6 aresponses failure spend=0.0 ua= used_client_oauth_token=true
    981da83e-a52e-4330-9b79-2617f70ce5c1 acompletion failure spend=0.0 ua=curl/8.7.1 used_client_oauth_token=true
    5ece92ba-bf51-40be-98c4-f6f55f5ef83a acompletion failure spend=0.0 ua=curl/8.7.1 used_client_oauth_token=true
    0
    
  3. Same command with &used_client_oauth_token=false:

    HTTP 200
    rows: 9
    msg_011CfWn4LcXCoCEUyrbkHJBQ anthropic_messages success spend=0.000076 ua=curl/8.7.1 used_client_oauth_token=false
    chatcmpl-197f83f4-af1c-44c5-8f12-617 acompletion success spend=0.000076 ua=curl/8.7.1 used_client_oauth_token=false
    chatcmpl-13e3cef1-c771-4b54-a48a-47b aresponses success spend=0.000076 ua=curl/8.7.1 used_client_oauth_token=false
    202894d0-56da-4501-9534-965e3ec98dd1 anthropic_messages failure spend=0.0 ua= used_client_oauth_token=false
    msg_011CfWn8SQdnxRnQypRVdxCx anthropic_messages success spend=0.000076 ua=curl/8.7.1 used_client_oauth_token=false
    chatcmpl-2bc9fe2b-5b8a-4f0e-9da1-24e aresponses success spend=0.000076 ua=curl/8.7.1 used_client_oauth_token=false
    chatcmpl-b385fe01-f324-48c4-9aef-161 acompletion success spend=0.000076 ua=curl/8.7.1 used_client_oauth_token=false
    chatcmpl-95565e95-7565-488b-94b2-eaf acompletion success spend=0.000076 ua= used_client_oauth_token=false
    356ec199-db9e-42e7-9e5f-10391c9b1b15 acompletion failure spend=0.0 ua=curl/8.7.1 used_client_oauth_token=false
    0
    
  4. Same command with &used_client_oauth_token=seat:

    HTTP 422
    {
      "detail": [
        {
          "type": "bool_parsing",
          "loc": [
            "query",
            "used_client_oauth_token"
          ],
          "msg": "Input should be a valid boolean, unable to interpret input"
        }
      ]
    }
    0
    

The true and false sets add up to all 17 rows, and the database holds no other spend row, so no row in this run has a null flag. No row or callback payload carries the seat token, its last 16 characters, the master key or the Anthropic key. The sk-ant-oat, token-tail, master-key and API-key grep counts are 0 in all four listings, every curl output and recorded.ndjson

Logs page

  1. Open http://localhost:42141/logs/ (the branch dashboard against the After proxy) and click Filters. The drawer now has a Credential select (first After screenshot under User Flow)
  2. Open the select, and the options are All Credentials, Client OAuth token and Configured key

pr43063-79bda61212-qa8589_ui_after_credential_options.png

  1. Pick Client OAuth token and Apply Filters. A Credential: Client OAuth token chip sits above the table, and only this run's eight seat rows stay: Claude Code's session of three, the three curl 429s, and the two guardrail rows 981da83e and 5ece92ba. The configured-key rows, the four body claims and the configured-key guardrail rows are gone

pr43063-79bda61212-qa8589_ui_after_filtered_client_oauth.png

  1. Open the Claude Code session and click its msg_011CfWn2mZQP74SgV6CSAynm span. Request Details reads Credential: Client OAuth token at cost $0.07427200 (second After screenshot under User Flow)
  2. Click the failed seat row 207cd0f8 (anthropic_messages, 429), and its drawer also reads Credential: Client OAuth token

pr43063-79bda61212-qa8589_ui_after_row_detail_failure.png

  1. Click the seat guardrail row 981da83e (acompletion, 429). Its drawer shows 1 guardrail evaluated, Guardrail: qa-content-filter and Credential: Client OAuth token

pr43063-79bda61212-qa8589_ui_after_row_detail_guardrail_seat.png

  1. Reload http://localhost:42141/logs/ with no filter and click row 202894d0, the call whose body claimed the flag. Its drawer reads Credential: Configured key

pr43063-79bda61212-qa8589_ui_after_row_detail_spoof.png

  1. Click the configured-key guardrail row chatcmpl-95565e95 (acompletion, 200). Its drawer shows 1 guardrail evaluated, Guardrail: qa-content-filter and Credential: Configured key

pr43063-79bda61212-qa8589_ui_after_row_detail_guardrail_key.png

What surprised the run, each with whether this PR touches it:

  • Seat token through raw curl gets Anthropic 429; untouched
  • Claude Code's launch-time request 429s on the seat; untouched
  • Rejected seat-token rows still read true; by design, see Caveats
  • Failed messages/responses rows lack metadata.user_agent, tags have it; untouched
  • The successful guardrail chat row chatcmpl-95565e95 has no metadata.user_agent, while its tags and the failed guardrail rows have it; untouched, since the PR changes user_agent only in tests
  • The 422 no longer echoes input, from main; untouched
  • Drawer Start Time shows local time with Z suffix; untouched
  • Chat requester_metadata holds the proxy stamp, not the caller's value; this PR
  • The dashboard login expires after ten minutes; untouched

Live risk check (317430d vs 79bda61)

Each leg ran from its own worktree with --num_workers 2 on one shared Postgres, one after the other, with the merge base on port 50132 and this tip on 52454. The only Anthropic upstream was a recording forwarder on 54301 in front of https://api.anthropic.com that logged header names, credential kind, body keys, and send count, never values. callbacks: ["generic_api"] pointed at a sink on 48229 that kept every StandardLoggingPayload. Each leg sent the same 32 requests as the c4f6 run (the original 20, six configured-key spoofs, two seat tokens claiming false, and four /v1/chat/completions requests with a litellm_content_filter guardrail), then 7 more aimed at the new failure-path route rule. Four of those were refused at the API key check (401) while claiming true, in the bucket the route does not read on /v1/chat/completions, /v1/messages, and /v1/responses and in the one it does read on /v1/responses. The other three were /v1/messages and /v1/responses requests the guardrail blocked (400). Rows were read back through GET /spend/logs/ui and GET /spend/logs/v2 and joined to payloads by request id. A merged-tree leg on 41713 sent all 39 again. Row writes lagged well behind the requests under load, so the base leg's first read came back short and was read again from a fresh base proxy once all 39 rows had landed

  • Breaking: none. All 39 requests got the same HTTP status, response body shape, and response header names on the base and head legs, and every outbound request to Anthropic matched on credential kind, header names, body keys, body size, and send count (43 sends on each leg, none of them from the 7 new requests). Neither sk-ant-oat nor the seat token's last 16 characters appears in any spend row, callback payload, recorded request, response body, response header, or proxy log on any leg
  • Backward incompatible: none observed. Callback payloads gain one key, metadata.used_client_oauth_token, which no base payload carries. The merge base ignores the new query param and returns the same rows for true, false, and maybe. This tip returns 9 rows for true, 21 for false, all 32 without the param, and HTTP 422 for maybe on both /spend/logs/ui and /spend/logs/v2, and 2, 2, 7, and 422 for the 7 new requests
  • Regression risk, all driven on the base and head legs:
    • Fixed since c4f6. The four guardrail rows on /v1/chat/completions now agree with their payloads, true for the seat token (429) and false for the configured key (200), the key claiming true (200), and the blocked request (400, nothing sent upstream). The three new guardrail rows agree too, false for the key on /v1/messages and /v1/responses and true for the seat token on /v1/messages, all blocked with nothing sent upstream. This is the only change from the c4f6 results, and it moves the split from 8/18/6 to 9/21/2, where the 2 are the passthrough rows
    • New, Low. Rows resolve the flag the same way the payload does for every authenticated request driven here, but not for a 401 whose own claim sits in the bucket its route does not read. The failure row reads only the route's bucket through metadata_variable_name_for_route, while the payload takes litellm_metadata first and falls back to metadata. So a litellm_metadata claim on /v1/chat/completions and a metadata claim on /v1/messages or /v1/responses each give a null row and a true payload, while a claim in the route's own bucket gives true in both. The analyzer printed no mismatch for the two /v1/responses 401s because failure payloads there carry no cell tag, so those were matched to their requests by start time. Only unauthenticated requests with a forged claim are affected, the row takes the safer value, and nothing is sent upstream
    • The 401 carry-forward from the Caveats is still there. Across the six forged 401 requests the payload reads true all six times and the row three times (the in-bucket ones), so an unauthenticated caller still gets a true row and payload by putting the claim in its route's bucket. No stamp runs on an auth failure, so writing null there for both row and payload would close this and the mismatch above
    • The six configured-key spoofs (metadata and litellm_metadata claims of true on all three unified endpoints) still log false in both row and payload, and the two seat tokens claiming false on /v1/messages and /v1/responses still log true (429). The route rule did not change any of them, and row and payload agree on all 32 of the c4f6 requests
    • Still Low and unchanged from c4f6. On all 16 chat requests, this tip's payload metadata.requester_metadata gains used_client_oauth_token with the proxy's value, overwriting whatever the caller sent. The key spoof, the both-buckets claim, and the guardrail claim go from true to false, and the junk values 7, "", [true], and the 5 KB string all become false. /v1/messages and /v1/responses keep the caller's own value there
    • The original 20 requests match the previous findings (junk values stamped false, streaming false, metadata.user_id still sent upstream, seat 429s true on all three unified endpoints, passthrough rows null)
    • Unchanged from the merge base, so outside this diff. The passthrough seat request's payload carries the caller's x-litellm-api-key header value in metadata.requester_custom_headers (one hit of the master key's last 16 characters per leg on base, head, and merged, none in spend rows)
  • Dependency graph: proxy_stamped_used_client_oauth_token in litellm/litellm_core_utils/core_helpers.py has two callers, the success row in litellm/proxy/spend_tracking/spend_tracking_utils.py and the callback payload in litellm_logging.py. metadata_variable_name_for_route in litellm/proxy/litellm_pre_call_utils.py backs _get_metadata_variable_name, which picks the bucket for the pre-call stamp, and the failure row writer _proxy_stamped_used_client_oauth_token in litellm/proxy/hooks/proxy_track_cost_callback.py, which falls back to get_metadata_variable_name_from_kwargs only when the route is unset. The merged tree was origin/main 3572d35, nine commits past the merge base (the fix(cost-map): add Vertex batch cache prices to vertex_ai/claude-sonnet-5-5 #43602 cost map change, the revert: "feat(usage): search team keys beyond the top-N in the Team usage view (#42857)" #43377 revert of feat(usage): search team keys beyond the top-N in the Team usage view #42857, feat(cache): select Rust caching through explicit cache objects #43601 Rust caching, feat(providers): add Prism provider (internal copy of #40914) #41961 Prism provider, Bedrock prices, credential canary tests, and CI), with no dependency file changes. It merged cleanly and its diff against origin/main is this PR's 24 files. Driven with all 39 requests it matched this tip on every status, send count, row flag, payload flag, and filter count. Two responses differ only because Anthropic chose to think on the merged leg (thinking blocks, finish_reason length, and the reasoning cost header), while the recorded outbound requests were identical and their flags did not change
  • Not verified live: a seat-token call that Anthropic answers with 200 on the unified endpoints, since the Max seat got a 429 (3 sends) on every unified call on every leg, so every true row seen live on those routes is a failure row. The route-unset fallback to get_metadata_variable_name_from_kwargs never ran, because every failure driven here, the 401s included, had its route set. Also not driven were a /v1/responses guardrail request with a seat token, the skills hook and compression interception, the Prometheus and GCS Pub/Sub consumers, and Bedrock and Vertex deployments

Type

🆕 New Feature

Caveats (if any)

Medium

  • spend on a seat-billed row stays at list price; reporting subtracts rows where the flag is true
  • The /anthropic/* passthrough forwards a caller's OAuth bearer next to the configured x-api-key (same on the merge base) and its rows carry no flag, so they match neither filter value (follow-up in LIT-8607)

Low

  • The flag records that the client sent an Anthropic OAuth bearer and the call went to an anthropic deployment, so a request routed to Bedrock or Vertex reads false even if the client sent one
  • A rejected OAuth token still reads true; the row's status shows the failure
  • Rows written before this release carry no value and match neither true nor false
  • Only the Anthropic OAuth forward sets true; the flag says nothing about other forwarded client headers
  • The deprecated /spend/logs route gets no filter; /spend/logs/ui and its v2 form do
  • A request refused at the API key check is logged before the stamp, so that zero-spend 401 row keeps the bool its own body claimed, the way the merge base already keeps requester_ip_address and applied_guardrails from those bodies
    • The row reads only the bucket its route uses while the callback payload reads litellm_metadata first, so a 401 whose claim sits in the other bucket logs null in the row and true in the payload
  • The provider check behind the flag lives in litellm/llms/anthropic/common_utils.py, next to the header forward it mirrors
  • On /v1/chat/completions the row's requester_metadata shows the proxy's stamp rather than the value the caller sent, because chat writes proxy fields into metadata before that snapshot, the way main already does for user_api_key, tags and headers
  • Every non-required CI job red at this tip is red the same way on main's own runs at 431ecd8, the main commit this branch merged in, so none is this PR's
    • misc / Run tests fails the same 4 test_openapi_compliance tests (main run)
    • proxy-endpoints / Run tests fails the same 2 test_create__non_object_metadata_is_400 batches tests (main run, again at 54ae4c5 in 36802988446)
    • proxy-behavior passes all 929 tests and fails only its Codecov upload step, which cannot verify the codecov CLI signature (main run)
    • rust-wheel fails its pytest step against the compiled extension (main run)

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR
  • 79bda61 passes /live-pr-risk

Note

Medium Risk
Changes how spend and standard logging metadata is stamped and resolved across success and failure paths; behavior is additive but affects billing reconciliation and can diverge on unauthenticated 401 rows vs callback payloads per documented caveats.

Overview
Adds metadata.used_client_oauth_token to spend logs and standard logging so admins can tell client-forwarded Anthropic OAuth (Max seat) from deployment API key billing—only a boolean is stored, never the token.

On pre-call, when the proxy forwards an Anthropic OAuth Authorization header it stamps the flag into the route’s metadata bucket (metadata vs litellm_metadata via metadata_variable_name_for_route). At log time resolve_used_client_oauth_token sets the field to true only if the stamp was true and the call went to the anthropic provider; caller-supplied values in their own metadata do not win over the proxy stamp. Failure callbacks stamp the same way from the correct bucket.

/spend/logs/ui (and v2) accept used_client_oauth_token=true|false; the Logs UI adds a Credential filter and shows Client OAuth token vs Configured key in the row drawer (legacy rows stay unset and match neither filter).

Reviewed by Cursor Bugbot for commit 7b791f3. Bugbot is set up for automated code reviews on this repo. Configure here.

…warded Anthropic OAuth token

Stamp metadata.used_client_oauth_token where the proxy decides to forward a
client's Anthropic OAuth token, carry it through StandardLoggingMetadata into
the spend log row, add a used_client_oauth_token filter to /spend/logs/ui, and
surface it on the Logs page as a Credential filter and drawer field. The token
itself never reaches the log
@devin-ai-integration

devin-ai-integration Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@greptile-apps

greptile-apps Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[Medium risk] Adds OAuth credential tracking to spend logs.

The PR appears safe to merge with respect to the reviewed credential-logging changes.

Summary

Adds a proxy-stamped boolean to spend and standard logs identifying requests that forwarded a client Anthropic OAuth token. The spend-log API and dashboard can filter by that credential and display it in request details. The merge since the previous review did not establish a new issue with this feature.

Reviews (7) · Last reviewed commit: "Merge remote-tracking branch 'origin/mai..."

Comment thread litellm/proxy/litellm_pre_call_utils.py
@codspeed

codspeed Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_spend_log_client_oauth_flag (7b791f3) with main (54ae4c5)

Open in CodSpeed

@codecov

codecov Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/proxy/hooks/proxy_track_cost_callback.py Outdated
Comment thread litellm/litellm_core_utils/get_provider_specific_headers.py Outdated
… rows and move the resolver under llms/anthropic
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@mateo-berri
mateo-berri enabled auto-merge (squash) September 25, 2026 01:22
@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/litellm_core_utils/litellm_logging.py
…adata slot

On routes that carry proxy metadata in litellm_metadata, metadata is the
caller's own body field, and merge_litellm_metadata lets it win. Resolve the
flag from litellm_metadata when the proxy stamped it there so a caller cannot
set it in the standard logging payload
@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

Comment thread litellm/proxy/litellm_pre_call_utils.py

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

…te stamped

A guardrail on the unified path adds litellm_metadata to a chat request after
the proxy stamped metadata, so both spend row writers read the new bucket and
stored null. The success row now resolves the flag the same way the callback
payload does, and the failure row picks the bucket from the request route.
@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

…ent_oauth_flag

# Conflicts:
#	litellm/proxy/hooks/proxy_track_cost_callback.py
#	litellm/proxy/spend_tracking/spend_tracking_utils.py
#	tests/test_litellm/proxy/spend_tracking/test_spend_management_endpoints.py
@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 7b791f3. Configure here.

@mateo-berri
mateo-berri merged commit 0c515ed into main Oct 1, 2026
103 of 110 checks passed
@mateo-berri
mateo-berri deleted the litellm_spend_log_client_oauth_flag branch October 1, 2026 02:17
jan-sauer-reef added a commit to jan-sauer-reef/litellm that referenced this pull request Oct 1, 2026
…ject_key_prefix

* upstream/main: (62 commits)
  fix(guardrails): scan Responses API input in Azure Prompt Shield (BerriAI#43786)
  feat(lens): investigate sampled traces and retain batch results (BerriAI#43942)
  fix(proxy): restore pre-config-wins handling of pass-through endpoints (BerriAI#43962)
  fix(cost-map): raise baseten DeepSeek-V4.1-Flash max output to 262144 (BerriAI#43916)
  chore(cost-map): add deprecation date for anthropic claude-sonnet-4-5 (BerriAI#43898)
  chore(cost-map): add fireworks inkling priority prices from the prices api (BerriAI#43949)
  feat(guardrails): honor litellm_params.timeout in every HTTP guardrail (BerriAI#43134)
  test(e2e): typed per-test metadata for the e2e suite (BerriAI#42044)
  fix(caching): write the response-cache SET to Redis at once instead of on the post-call batch (BerriAI#43973)
  feat(ui): filter tags by name and description on the Tag Management page (BerriAI#42949)
  feat(providers): add Cortecs as an OpenAI-compatible provider (BerriAI#43872)
  feat(e2e): record each e2e test's steps, starting with ProxyClient (BerriAI#42393)
  test(ci): repair stale tests and move retired OpenAI text-completion fixtures (BerriAI#43958)
  feat(proxy): record in spend logs whether a request used a client-forwarded Anthropic OAuth token (BerriAI#43063)
  fix(azure_storage): keep the DataLakeServiceClient alive until its TTL elapses (BerriAI#43082)
  chore(deps): bump gitpython and tornado, extend diskcache osv ignore to Nov 1 (BerriAI#43961)
  fix(guardrails): treat an unknown straiker api_version as unset instead of skipping the guardrail (BerriAI#43956)
  fix(azure_storage): name Data Lake objects without base64 padding or slashes (BerriAI#43914)
  fix(grayswan): send request conversation and tool calls to post-call monitor (BerriAI#43770)
  chore(cost-map): sync openrouter prices from the models API (BerriAI#43950)
  ...
yuneng-berri added a commit that referenced this pull request Oct 1, 2026
* test(ci): add used_client_oauth_token to the GCS pub/sub spend-log golden

#43063 stamps used_client_oauth_token into spend-log metadata, so
test_async_gcs_pub_sub_v1 failed on main with an extra metadata key

* test(ui): give the auto-router threshold save wait room for the availability debounce

#42625 keeps Save disabled while a 300ms-debounced availability check runs.
This test waits for Save right after the change, so the whole debounce lands
inside waitFor's 1s default and it times out under CI load. It is the
recurring UI Unit Tests failure on main since #42625 landed

* test(e2e): expect no pricing tier on bills for streamed calls OpenAI served at default

#42870 added both the rule that a served default or standard tier bills at
base pricing and records no service_tier, and streamed tests expecting the
row to record 'default'. They have failed on every scheduled litellm-e2e run
since. The tests now map the served tier to the pricing basis the bill must
record and check input is billed at that basis's rate; the messages case
registers custom rates so the rate check has something to compare against

* test(e2e-ui): wait for the call-id search before hovering the logs row

The row the spec hovers is already on the unfiltered first page, so it was
found before the search request returned. The search response then
re-rendered the table under the mouse, and the Base UI tooltip never opened.
Reproduced with Playwright against a local proxy: hovering right after the
fill never shows the tooltip, hovering after the search response shows the
call id every time

* test(e2e): run the Together structured-output case on the hybrid Qwen with reasoning off

The case picked the cheapest Together row flagged supports_response_schema.
DeepSeek-V4-Flash-0731 hit its cost-map deprecation date on 2026-09-29, so the
pick moved to GLM-5.3-Flash, a reasoning-only model that spends the 1024-token
budget thinking and returns content=None. Qwen3.5-9B is the pinned hybrid model
the reasoning_effort=none case already exercises, and Together lists it with
structured output support

* test(integration): read the agent 365 guardrail status by its own name in spend logs

The MCP shard runs under xdist against one database, and a sibling file creates a
default_on pre_mcp_call content filter there. The owned proxy reloads DB guardrails, so
that filter's 'success' entry could land first in guardrail_information and the test
read it instead of the agent 365 verdict

* test(unit): ignore asyncio's leaked-task records in the budget limiter push-failure log check

gc.collect() inside the caplog window can collect a pending task an earlier test left
on a closed loop, and asyncio logs 'Task was destroyed but it is pending' into this
test's records. The check still counts every LiteLLM logger, and unretrieved task
exceptions on this loop still go through the asserted exception handler

* test(e2e-ui): fill the create-tag fields inside the dialog

#42949 added 'Filter by tag name' and 'Filter by description' inputs to the Tag
Management page, so page-wide getByLabel('Tag Name') and getByLabel('Description')
match two elements and Playwright's strict mode fails the create step

* test(integration): run integration proxies with the CI license

Multi-worker proxies start each uvicorn worker in a fresh process, so every
worker reads the license from its environment. Forward LITELLM_LICENSE into the
proxy and test runner environments

* ci: save GitHub Actions caches only from main and bump codecov-action to 5.5.5

Every pull request saved its own uv, maturin, Rust and Prisma caches, about
4.5 GB per PR, so the repository's 10 GB cache budget evicted main's entries
within minutes. Pull request jobs then missed every cache, downloaded all
dependencies from PyPI and hit the install step timeouts. Pull requests now
restore only, and main keeps the caches warm for them. test-linting and
check-ui-api-types run only on pull requests and keep saving

codecov-action 5.5.4 imports its signing key from the deleted codecovsecurity
keybase account, so every upload failed signature verification. 5.5.5 reads it
from codecovsecops; the key ID matches the one signing the current CLI

* test(unit): join the session-minting thread before collecting the handler

asyncio.to_thread resumes the test as soon as the worker sets its result,
while the pool thread can still hold the work item and through it the
handler. gc.collect() then cannot finalize the handler and the session stays
open. A pool that shuts down before the test continues drops that reference

* test(integration): relaunch owned proxies that lose their port, expire idle gateway connections early

owned_proxy_process released its reserved port and the proxy bound it only
after full startup, so another xdist worker or an outgoing connection could
take it first and the proxy exited with 'address already in use'. The launch
now retries on a fresh port when that happens and stops every failed attempt.

uvicorn closes idle keep-alive connections after 5 seconds and httpx expired
them at the same 5 seconds, so a request sent right at that mark could reuse a
socket the server was closing and get 'Connection reset by peer'. Gateway
clients now drop idle connections after 2 seconds

* ci(circleci): give the base SDK wheel build the same 30 minute no-output window as the Windows build

The release profile builds with fat LTO and one codegen unit, so the final
link of litellm-cache-s3 runs silently for minutes. Successful builds take
711 to 749 seconds, right at the default 10 minute no-output limit, and about
30% of recent runs were killed there

* test(integration): model the budget-reset database outage as 10 seconds instead of 5 refused connections

The proxy retries the database about every 30 seconds and each retry opens
roughly one connection, so a 5-connection outage took 3 to 4 retries to clear
and recovery landed between 60 and 90 seconds, straddling the test's 80 second
reset window. A fixed 10 second outage still refuses the immediate reconnect
and recovers on the next retry

* ci: move the unit-test uv cache split into a composite action

check_workflow_startup_safety sums every setup step's timeout, so the save and
restore variants each counted 5 minutes although only one runs. One composite
step keeps the setup ceiling at 35 minutes

* test(unit): point tiktoken at the bundled cache for every unit test

The rust_bridge tokenizer tests loaded o200k_base before any test in their
xdist worker had imported default_encoding, so tiktoken fell back to the
temp cache and tried to download under pytest-socket. Move the session
fixture from litellm_core_utils/conftest.py to the root unit conftest.

* test(integration): answer model discovery probes in the hosted_vllm wire tests

The router's periodic upstream model info refresh sends GET /v1/models to
hosted_vllm deployments, so a wire server that is live during a refresh
sees an extra request. Answer the probe with an empty model list and leave
it out of the provider-call assertions, matching the responses bridge
tests.

This branch is waiting to be deployed

1 waiting deployment
e2e-changed — 7b791f34 Waiting Oct 1, 2026 by mateo-berri via oauth #2221
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant