Repository navigation
feat(proxy): record in spend logs whether a request used a client-forwarded Anthropic OAuth token - #43063
Conversation
…warded Anthropic OAuth token Stamp metadata.used_client_oauth_token where the proxy decides to forward a client's Anthropic OAuth token, carry it through StandardLoggingMetadata into the spend log row, add a used_client_oauth_token filter to /spend/logs/ui, and surface it on the Logs page as a Credential filter and drawer field. The token itself never reaches the log
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
|
Codecov Report❌ Patch coverage is 📢 Thoughts on this report? Let us know! |
… litellm_metadata routes
|
bugbot run |
… rows and move the resolver under llms/anthropic
|
bugbot run |
|
bugbot run |
…adata slot On routes that carry proxy metadata in litellm_metadata, metadata is the caller's own body field, and merge_litellm_metadata lets it win. Resolve the flag from litellm_metadata when the proxy stamped it there so a caller cannot set it in the standard logging payload
|
bugbot run |
…te stamped A guardrail on the unified path adds litellm_metadata to a chat request after the proxy stamped metadata, so both spend row writers read the new bucket and stored null. The success row now resolves the flag the same way the callback payload does, and the failure row picks the bucket from the request route.
|
bugbot run |
…ent_oauth_flag # Conflicts: # litellm/proxy/hooks/proxy_track_cost_callback.py # litellm/proxy/spend_tracking/spend_tracking_utils.py # tests/test_litellm/proxy/spend_tracking/test_spend_management_endpoints.py
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 7b791f3. Configure here.
…ject_key_prefix * upstream/main: (62 commits) fix(guardrails): scan Responses API input in Azure Prompt Shield (BerriAI#43786) feat(lens): investigate sampled traces and retain batch results (BerriAI#43942) fix(proxy): restore pre-config-wins handling of pass-through endpoints (BerriAI#43962) fix(cost-map): raise baseten DeepSeek-V4.1-Flash max output to 262144 (BerriAI#43916) chore(cost-map): add deprecation date for anthropic claude-sonnet-4-5 (BerriAI#43898) chore(cost-map): add fireworks inkling priority prices from the prices api (BerriAI#43949) feat(guardrails): honor litellm_params.timeout in every HTTP guardrail (BerriAI#43134) test(e2e): typed per-test metadata for the e2e suite (BerriAI#42044) fix(caching): write the response-cache SET to Redis at once instead of on the post-call batch (BerriAI#43973) feat(ui): filter tags by name and description on the Tag Management page (BerriAI#42949) feat(providers): add Cortecs as an OpenAI-compatible provider (BerriAI#43872) feat(e2e): record each e2e test's steps, starting with ProxyClient (BerriAI#42393) test(ci): repair stale tests and move retired OpenAI text-completion fixtures (BerriAI#43958) feat(proxy): record in spend logs whether a request used a client-forwarded Anthropic OAuth token (BerriAI#43063) fix(azure_storage): keep the DataLakeServiceClient alive until its TTL elapses (BerriAI#43082) chore(deps): bump gitpython and tornado, extend diskcache osv ignore to Nov 1 (BerriAI#43961) fix(guardrails): treat an unknown straiker api_version as unset instead of skipping the guardrail (BerriAI#43956) fix(azure_storage): name Data Lake objects without base64 padding or slashes (BerriAI#43914) fix(grayswan): send request conversation and tool calls to post-call monitor (BerriAI#43770) chore(cost-map): sync openrouter prices from the models API (BerriAI#43950) ...
* test(ci): add used_client_oauth_token to the GCS pub/sub spend-log golden #43063 stamps used_client_oauth_token into spend-log metadata, so test_async_gcs_pub_sub_v1 failed on main with an extra metadata key * test(ui): give the auto-router threshold save wait room for the availability debounce #42625 keeps Save disabled while a 300ms-debounced availability check runs. This test waits for Save right after the change, so the whole debounce lands inside waitFor's 1s default and it times out under CI load. It is the recurring UI Unit Tests failure on main since #42625 landed * test(e2e): expect no pricing tier on bills for streamed calls OpenAI served at default #42870 added both the rule that a served default or standard tier bills at base pricing and records no service_tier, and streamed tests expecting the row to record 'default'. They have failed on every scheduled litellm-e2e run since. The tests now map the served tier to the pricing basis the bill must record and check input is billed at that basis's rate; the messages case registers custom rates so the rate check has something to compare against * test(e2e-ui): wait for the call-id search before hovering the logs row The row the spec hovers is already on the unfiltered first page, so it was found before the search request returned. The search response then re-rendered the table under the mouse, and the Base UI tooltip never opened. Reproduced with Playwright against a local proxy: hovering right after the fill never shows the tooltip, hovering after the search response shows the call id every time * test(e2e): run the Together structured-output case on the hybrid Qwen with reasoning off The case picked the cheapest Together row flagged supports_response_schema. DeepSeek-V4-Flash-0731 hit its cost-map deprecation date on 2026-09-29, so the pick moved to GLM-5.3-Flash, a reasoning-only model that spends the 1024-token budget thinking and returns content=None. Qwen3.5-9B is the pinned hybrid model the reasoning_effort=none case already exercises, and Together lists it with structured output support * test(integration): read the agent 365 guardrail status by its own name in spend logs The MCP shard runs under xdist against one database, and a sibling file creates a default_on pre_mcp_call content filter there. The owned proxy reloads DB guardrails, so that filter's 'success' entry could land first in guardrail_information and the test read it instead of the agent 365 verdict * test(unit): ignore asyncio's leaked-task records in the budget limiter push-failure log check gc.collect() inside the caplog window can collect a pending task an earlier test left on a closed loop, and asyncio logs 'Task was destroyed but it is pending' into this test's records. The check still counts every LiteLLM logger, and unretrieved task exceptions on this loop still go through the asserted exception handler * test(e2e-ui): fill the create-tag fields inside the dialog #42949 added 'Filter by tag name' and 'Filter by description' inputs to the Tag Management page, so page-wide getByLabel('Tag Name') and getByLabel('Description') match two elements and Playwright's strict mode fails the create step * test(integration): run integration proxies with the CI license Multi-worker proxies start each uvicorn worker in a fresh process, so every worker reads the license from its environment. Forward LITELLM_LICENSE into the proxy and test runner environments * ci: save GitHub Actions caches only from main and bump codecov-action to 5.5.5 Every pull request saved its own uv, maturin, Rust and Prisma caches, about 4.5 GB per PR, so the repository's 10 GB cache budget evicted main's entries within minutes. Pull request jobs then missed every cache, downloaded all dependencies from PyPI and hit the install step timeouts. Pull requests now restore only, and main keeps the caches warm for them. test-linting and check-ui-api-types run only on pull requests and keep saving codecov-action 5.5.4 imports its signing key from the deleted codecovsecurity keybase account, so every upload failed signature verification. 5.5.5 reads it from codecovsecops; the key ID matches the one signing the current CLI * test(unit): join the session-minting thread before collecting the handler asyncio.to_thread resumes the test as soon as the worker sets its result, while the pool thread can still hold the work item and through it the handler. gc.collect() then cannot finalize the handler and the session stays open. A pool that shuts down before the test continues drops that reference * test(integration): relaunch owned proxies that lose their port, expire idle gateway connections early owned_proxy_process released its reserved port and the proxy bound it only after full startup, so another xdist worker or an outgoing connection could take it first and the proxy exited with 'address already in use'. The launch now retries on a fresh port when that happens and stops every failed attempt. uvicorn closes idle keep-alive connections after 5 seconds and httpx expired them at the same 5 seconds, so a request sent right at that mark could reuse a socket the server was closing and get 'Connection reset by peer'. Gateway clients now drop idle connections after 2 seconds * ci(circleci): give the base SDK wheel build the same 30 minute no-output window as the Windows build The release profile builds with fat LTO and one codegen unit, so the final link of litellm-cache-s3 runs silently for minutes. Successful builds take 711 to 749 seconds, right at the default 10 minute no-output limit, and about 30% of recent runs were killed there * test(integration): model the budget-reset database outage as 10 seconds instead of 5 refused connections The proxy retries the database about every 30 seconds and each retry opens roughly one connection, so a 5-connection outage took 3 to 4 retries to clear and recovery landed between 60 and 90 seconds, straddling the test's 80 second reset window. A fixed 10 second outage still refuses the immediate reconnect and recovers on the next retry * ci: move the unit-test uv cache split into a composite action check_workflow_startup_safety sums every setup step's timeout, so the save and restore variants each counted 5 minutes although only one runs. One composite step keeps the setup ceiling at 35 minutes * test(unit): point tiktoken at the bundled cache for every unit test The rust_bridge tokenizer tests loaded o200k_base before any test in their xdist worker had imported default_encoding, so tiktoken fell back to the temp cache and tried to download under pytest-socket. Move the session fixture from litellm_core_utils/conftest.py to the root unit conftest. * test(integration): answer model discovery probes in the hosted_vllm wire tests The router's periodic upstream model info refresh sends GET /v1/models to hosted_vllm deployments, so a wire server that is live during a refresh sees an extra request. Answer the probe with an empty model list and leave it out of the provider-call assertions, matching the responses bridge tests.
TLDR
Problem this solves:
How it solves it:
metadata.used_client_oauth_token, true or falsemetadata.used_client_oauth_tokenis overwritten by the proxy's stamp/spend/logs/ui(and v2) takeused_client_oauth_token=true|falseIntentional product change: the Logs page Filters drawer gains a Credential select (All Credentials, Client OAuth token, Configured key) and the row drawer's Request Details gains a Credential field, so admins can split seat-billed rows from key-billed ones; nothing existing moves or goes away
User Flow
Before: an admin whose developers run Claude Code on Max seats through the gateway cannot tell which spend log rows a seat paid for, so the API bill they reconcile from the logs is wrong
ANTHROPIC_BASE_URL=https://litellm-domain,ANTHROPIC_CUSTOM_HEADERS="x-litellm-api-key: Bearer sk-...") and their prompt gets a replyAuthorization: Bearer sk-...alone, so the deployment's configured Anthropic key pays, and gets a reply tooFilters drawer at https://litellm-domain/ui/?page=logs: Cache is followed straight by Key Alias, no Credential filter
Row drawer of the Claude Code request: Request Details ends at IP Address, no Credential field
After: the same two rows differ on
metadata.used_client_oauth_token, so the admin can list exactly the seat-billed requestsANTHROPIC_BASE_URL=https://litellm-domain,ANTHROPIC_CUSTOM_HEADERS="x-litellm-api-key: Bearer sk-...") and their prompt gets a replyAuthorization: Bearer sk-...alone, so the deployment's configured Anthropic key pays, and gets a reply tooCredential: Client OAuth token, the curl row's readsConfigured key, both still at list-price spendmetadata.used_client_oauth_token: true(falseon the curl row even when its body claimedtrue, absent on rows written before the upgrade), and no part of the token appears in either rowFilters drawer at https://litellm-domain/ui/?page=logs: a Credential select sits between Cache and Key Alias
Row drawer of the Claude Code request: Request Details gains
Credential: Client OAuth tokenRelevant issues
Related to #23755
Linear ticket
Resolves LIT-8589
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Both legs booted
python litellm/proxy/proxy_cli.py --config config.yaml --num_workers 2 --use_v2_migration_resolver, two uvicorn workers on their own Postgres database each, with real Anthropic calls. Before is the merge base 317430d on port 31734 and After is this PR's tip 79bda61 on port 58912, each with its own commit's dashboard served bynext devagainst it. The After leg also turned on thegeneric_apicallback, posting to a local recorder, so the callback payload is checked next to the spend rows. The After config is below, and Before ran the same file without thelitellm_settingsblockTwo credentials drive every case:
$LITELLM_MASTER_KEYalone, so the deployment's configured Anthropic key pays, and$OAUTH_TOKEN, thesk-ant-oat...access token of a Claude Max seat (read from Claude Code's credential store), sent as the bearer withx-litellm-api-key: Bearer $LITELLM_MASTER_KEYalongside, which is how Claude Code itself sends it. The Claude Code case runs the real TUI (v2.1.284, signed into a Max seat) under tmux; the curl cases replay the same prompt on every unified endpointBefore (317430d)
Merge base proxy on port 31734 with
--num_workers 2and one Postgres database,claude-sonnet-5-5routed to Anthropic with the configured keyClaude Code on /v1/messages with the Max seat
export ANTHROPIC_BASE_URL=http://127.0.0.1:31734; export ANTHROPIC_CUSTOM_HEADERS="x-litellm-api-key: Bearer $LITELLM_MASTER_KEY"; claude --model claude-sonnet-5-5(Claude Code v2.1.284,Sonnet 5.5 · Claude Max), then typeReply with the single word pong.and press Enterpong(Crunched for 1s, done 2:55 PM), and the proxy logs the turn as rows msg_011CfWakZsi4AjWqQcmgxeVX and msg_011CfWakb7NjMBqvv23pSCcv, bothua=claude-cli/2.1.284in the listing below. Row aff94a49 right before them is a request Claude Code sent at startup, before the prompt was typed, that Anthropic answered with 429; the Logs page groups it into the same Claude Code sessioncurl /v1/messages
Configured key:
Seat token (
-H "Authorization: Bearer $OAUTH_TOKEN" -H "x-litellm-api-key: Bearer $LITELLM_MASTER_KEY", same body):curl /v1/chat/completions
Configured key:
Seat token (same body, bearer
$OAUTH_TOKENplusx-litellm-api-key):curl /v1/responses
Configured key:
Seat token (same body, bearer
$OAUTH_TOKENplusx-litellm-api-key):curl /v1/messages claiming the flag in its own metadata
Configured key, the body carrying
"metadata":{"used_client_oauth_token":true}plus amax_tokensAnthropic rejects, so the row is a failure that spent nothing (same merge base proxy on port 31734):The listing command of the next case bounded to this call (
start_date=2026-09-28%2021:56:16&end_date=2026-09-28%2021:56:22, port 31734): the row carries noused_client_oauth_tokenat all, so the body's claim is not storedSame listing with
&used_client_oauth_token=true:HTTP 200,rows: 1, the same row, the parameter is ignoredGET /spend/logs/ui
All rows of the run window:
Same command with
&used_client_oauth_token=trueon the query:HTTP 200,rows: 10, the identical ten rows, every oneused_client_oauth_token=null, so the parameter is ignoredSame command with
&used_client_oauth_token=false:HTTP 200,rows: 10, the identical ten rowsSame command with
&used_client_oauth_token=seat:HTTP 200,rows: 10, the identical ten rows, the junk value is not rejectedLogs page
This commit's dashboard served with
NEXT_PUBLIC_BASE_URL=http://127.0.0.1:31734 npx next dev -H 127.0.0.1 -p 51779fromui/litellm-dashboard, the proxy restarted on the same port withPROXY_BASE_URL=http://127.0.0.1:31734so the dashboard calls it (the rows above live in Postgres and carry over), signed in asadminwith the master keyUser-Agent: claude-cli/2.1.284 (external, cli)): Request Details shows Model, Provider, Call Type, User, Model ID, API Base and IP Address, with no Credential field (second Before screenshot under User Flow)After (79bda61)
This run checks the After leg at 79bda61, the commit that reads
used_client_oauth_tokenfrom the bucket the route stamped, so a guardrail that addslitellm_metadatato a /v1/chat/completions request no longer leaves the flag null. Same topology as Before: two workers, a fresh Postgres database, port 58912,PROXY_BASE_URL=http://localhost:58912, andlitellm_settings.callbacks: ["generic_api"]withGENERIC_LOGGER_ENDPOINT=http://127.0.0.1:46360/log, a local recorder that appends every JSON array the callback POSTs to recorded.ndjson. The config also has onelitellm_content_filterguardrail for the new case at the end of the curl sectionsClaude Code on /v1/messages with the Max seat
export ANTHROPIC_BASE_URL=http://127.0.0.1:58912; export ANTHROPIC_CUSTOM_HEADERS="x-litellm-api-key: Bearer $LITELLM_MASTER_KEY"; claude --model claude-sonnet-5-5, then typeReply with the single word pong.and press Enterpong(Churned for 4s, done 5:23 PM). The proxy logs the turn as rows msg_011CfWn2dNHSUfkGcsWmqpZZ and msg_011CfWn2mZQP74SgV6CSAynm, bothua=claude-cli/2.1.284andused_client_oauth_token=truein the listing below. They come after the 429 row 360fe2aa, which Claude Code sent at launch before the prompt was typedcurl /v1/messages
Configured key (same command as Before, port 58912):
Seat token (same command, bearer
$OAUTH_TOKENplusx-litellm-api-key):curl /v1/chat/completions
Configured key (same command as Before, port 58912):
Seat token:
curl /v1/responses
Configured key (same command as Before, port 58912):
Seat token:
curl /v1/messages claiming the flag in its own metadata
Same command as Before, port 58912:
It is row 202894d0 in the listings below. It reads
used_client_oauth_token=falsealthough the body claimedtrue, because the configured key made the call, and it answers to=falseand not to=trueSuccessful calls claiming the flag, all three endpoints
This is the case Bugbot raised. Each call uses the configured key and puts
"metadata":{"used_client_oauth_token":true}in the body:All three rows read
used_client_oauth_token=falsein the spend-log listings below (msg_011CfWn8SQ, chatcmpl-2bc9fe2b for the responses call, chatcmpl-b385fe01)Content filter guardrail on /v1/chat/completions
This is the case the new commit fixes. The config adds one pre_call content filter that a request turns on by naming it:
Configured key and seat token, both asking for the guardrail:
The same pair with
"content":"Reply with the single word qablockedword.", so the guardrail blocks both before any call to Anthropic. Both answer the same way:The four rows, from the all-rows listing below:
The guardrail ran on all four requests. The seat rows (981da83e for the pong call, 5ece92ba for the blocked call) read
trueand both come back under?used_client_oauth_token=true. The configured-key rows (chatcmpl-95565e95 and 356ec199) readfalseand come back under=false. None of the four is null, and the generic_api payloads below carry the same valuesgeneric_api callback payloads
Every standard logging payload the callback delivered in the run window, in time order:
jq -c --argjson ws 1790641372 --argjson we 1790641512 '[.[] | select(.startTime >= $ws and .startTime <= $we)] | sort_by(.startTime)[] | {id: .id[:36], call_type, status, used_client_oauth_token: .metadata.used_client_oauth_token, requester_metadata_claim: .metadata.requester_metadata.used_client_oauth_token}' recorded.ndjsonThe four body claims (202894d0 and the three 200s) all read
falsein.metadata.used_client_oauth_token, the same as the spend-log rows. The eighttruepayloads are all seat-token requests: the three curl 429s, the guardrail 429 and the guardrail block, Claude Code's launch 429, and Claude Code's two successful callsmetadata.requester_metadatafor the claims depends on the endpoint. On /v1/messages and /v1/responses it is{"used_client_oauth_token":true}, the caller's own value kept as sent. On /v1/chat/completions the claim row's requester_metadata is{"used_client_oauth_token":false,"headers":{host, user-agent, accept, content-type, content-length}}, so it holds the proxy stamp instead of the caller'strue. On chat the PR writes the stamp intodata["metadata"]before requester_metadata is copied from that same dict, so the stamp replaces the caller's key. The chat calls with no caller metadata get the same stamp there:falsefor chatcmpl-197f83f4 and the two configured-key guardrail rows,truefor the seat 429 on chat (5fd44b3b) and the two seat guardrail rows. Theheaderspart is what main already doesGET /spend/logs/ui
All rows of the run window:
Same command with
&used_client_oauth_token=true:Same command with
&used_client_oauth_token=false:Same command with
&used_client_oauth_token=seat:The true and false sets add up to all 17 rows, and the database holds no other spend row, so no row in this run has a null flag. No row or callback payload carries the seat token, its last 16 characters, the master key or the Anthropic key. The
sk-ant-oat, token-tail, master-key and API-key grep counts are 0 in all four listings, every curl output and recorded.ndjsonLogs page
Credential: Client OAuth tokenchip sits above the table, and only this run's eight seat rows stay: Claude Code's session of three, the three curl 429s, and the two guardrail rows 981da83e and 5ece92ba. The configured-key rows, the four body claims and the configured-key guardrail rows are goneCredential: Client OAuth tokenat cost $0.07427200 (second After screenshot under User Flow)Credential: Client OAuth token1 guardrail evaluated,Guardrail: qa-content-filterandCredential: Client OAuth tokenCredential: Configured key1 guardrail evaluated,Guardrail: qa-content-filterandCredential: Configured keyWhat surprised the run, each with whether this PR touches it:
input, from main; untouchedLive risk check (317430d vs 79bda61)
Each leg ran from its own worktree with
--num_workers 2on one shared Postgres, one after the other, with the merge base on port 50132 and this tip on 52454. The only Anthropic upstream was a recording forwarder on 54301 in front ofhttps://api.anthropic.comthat logged header names, credential kind, body keys, and send count, never values.callbacks: ["generic_api"]pointed at a sink on 48229 that kept every StandardLoggingPayload. Each leg sent the same 32 requests as the c4f6 run (the original 20, six configured-key spoofs, two seat tokens claiming false, and four/v1/chat/completionsrequests with alitellm_content_filterguardrail), then 7 more aimed at the new failure-path route rule. Four of those were refused at the API key check (401) while claimingtrue, in the bucket the route does not read on/v1/chat/completions,/v1/messages, and/v1/responsesand in the one it does read on/v1/responses. The other three were/v1/messagesand/v1/responsesrequests the guardrail blocked (400). Rows were read back throughGET /spend/logs/uiandGET /spend/logs/v2and joined to payloads by request id. A merged-tree leg on 41713 sent all 39 again. Row writes lagged well behind the requests under load, so the base leg's first read came back short and was read again from a fresh base proxy once all 39 rows had landedsk-ant-oatnor the seat token's last 16 characters appears in any spend row, callback payload, recorded request, response body, response header, or proxy log on any legmetadata.used_client_oauth_token, which no base payload carries. The merge base ignores the new query param and returns the same rows fortrue,false, andmaybe. This tip returns 9 rows fortrue, 21 forfalse, all 32 without the param, and HTTP 422 formaybeon both/spend/logs/uiand/spend/logs/v2, and 2, 2, 7, and 422 for the 7 new requests/v1/chat/completionsnow agree with their payloads,truefor the seat token (429) andfalsefor the configured key (200), the key claiming true (200), and the blocked request (400, nothing sent upstream). The three new guardrail rows agree too,falsefor the key on/v1/messagesand/v1/responsesandtruefor the seat token on/v1/messages, all blocked with nothing sent upstream. This is the only change from the c4f6 results, and it moves the split from 8/18/6 to 9/21/2, where the 2 are the passthrough rowsmetadata_variable_name_for_route, while the payload takeslitellm_metadatafirst and falls back tometadata. So alitellm_metadataclaim on/v1/chat/completionsand ametadataclaim on/v1/messagesor/v1/responseseach give anullrow and atruepayload, while a claim in the route's own bucket givestruein both. The analyzer printed no mismatch for the two/v1/responses401s because failure payloads there carry no cell tag, so those were matched to their requests by start time. Only unauthenticated requests with a forged claim are affected, the row takes the safer value, and nothing is sent upstreamtrueall six times and the row three times (the in-bucket ones), so an unauthenticated caller still gets atruerow and payload by putting the claim in its route's bucket. No stamp runs on an auth failure, so writingnullthere for both row and payload would close this and the mismatch abovemetadataandlitellm_metadataclaims oftrueon all three unified endpoints) still logfalsein both row and payload, and the two seat tokens claimingfalseon/v1/messagesand/v1/responsesstill logtrue(429). The route rule did not change any of them, and row and payload agree on all 32 of the c4f6 requestsmetadata.requester_metadatagainsused_client_oauth_tokenwith the proxy's value, overwriting whatever the caller sent. The key spoof, the both-buckets claim, and the guardrail claim go fromtruetofalse, and the junk values7,"",[true], and the 5 KB string all becomefalse./v1/messagesand/v1/responseskeep the caller's own value therefalse, streamingfalse,metadata.user_idstill sent upstream, seat 429strueon all three unified endpoints, passthrough rowsnull)x-litellm-api-keyheader value inmetadata.requester_custom_headers(one hit of the master key's last 16 characters per leg on base, head, and merged, none in spend rows)proxy_stamped_used_client_oauth_tokeninlitellm/litellm_core_utils/core_helpers.pyhas two callers, the success row inlitellm/proxy/spend_tracking/spend_tracking_utils.pyand the callback payload inlitellm_logging.py.metadata_variable_name_for_routeinlitellm/proxy/litellm_pre_call_utils.pybacks_get_metadata_variable_name, which picks the bucket for the pre-call stamp, and the failure row writer_proxy_stamped_used_client_oauth_tokeninlitellm/proxy/hooks/proxy_track_cost_callback.py, which falls back toget_metadata_variable_name_from_kwargsonly when the route is unset. The merged tree was origin/main 3572d35, nine commits past the merge base (the fix(cost-map): add Vertex batch cache prices to vertex_ai/claude-sonnet-5-5 #43602 cost map change, the revert: "feat(usage): search team keys beyond the top-N in the Team usage view (#42857)" #43377 revert of feat(usage): search team keys beyond the top-N in the Team usage view #42857, feat(cache): select Rust caching through explicit cache objects #43601 Rust caching, feat(providers): add Prism provider (internal copy of #40914) #41961 Prism provider, Bedrock prices, credential canary tests, and CI), with no dependency file changes. It merged cleanly and its diff against origin/main is this PR's 24 files. Driven with all 39 requests it matched this tip on every status, send count, row flag, payload flag, and filter count. Two responses differ only because Anthropic chose to think on the merged leg (thinking blocks,finish_reasonlength, and the reasoning cost header), while the recorded outbound requests were identical and their flags did not changetruerow seen live on those routes is a failure row. The route-unset fallback toget_metadata_variable_name_from_kwargsnever ran, because every failure driven here, the 401s included, had its route set. Also not driven were a/v1/responsesguardrail request with a seat token, the skills hook and compression interception, the Prometheus and GCS Pub/Sub consumers, and Bedrock and Vertex deploymentsType
🆕 New Feature
Caveats (if any)
Medium
spendon a seat-billed row stays at list price; reporting subtracts rows where the flag is true/anthropic/*passthrough forwards a caller's OAuth bearer next to the configuredx-api-key(same on the merge base) and its rows carry no flag, so they match neither filter value (follow-up in LIT-8607)Low
anthropicdeployment, so a request routed to Bedrock or Vertex reads false even if the client sent onetruenorfalse/spend/logsroute gets no filter;/spend/logs/uiand its v2 form dorequester_ip_addressandapplied_guardrailsfrom those bodieslitellm_metadatafirst, so a 401 whose claim sits in the other bucket logsnullin the row andtruein the payloadlitellm/llms/anthropic/common_utils.py, next to the header forward it mirrors/v1/chat/completionsthe row'srequester_metadatashows the proxy's stamp rather than the value the caller sent, because chat writes proxy fields intometadatabefore that snapshot, the way main already does foruser_api_key, tags and headersmisc / Run testsfails the same 4test_openapi_compliancetests (main run)proxy-endpoints / Run testsfails the same 2test_create__non_object_metadata_is_400batches tests (main run, again at 54ae4c5 in 36802988446)proxy-behaviorpasses all 929 tests and fails only its Codecov upload step, which cannot verify the codecov CLI signature (main run)rust-wheelfails its pytest step against the compiled extension (main run)Final Attestation
Note
Medium Risk
Changes how spend and standard logging metadata is stamped and resolved across success and failure paths; behavior is additive but affects billing reconciliation and can diverge on unauthenticated 401 rows vs callback payloads per documented caveats.
Overview
Adds
metadata.used_client_oauth_tokento spend logs and standard logging so admins can tell client-forwarded Anthropic OAuth (Max seat) from deployment API key billing—only a boolean is stored, never the token.On pre-call, when the proxy forwards an Anthropic OAuth
Authorizationheader it stamps the flag into the route’s metadata bucket (metadatavslitellm_metadataviametadata_variable_name_for_route). At log timeresolve_used_client_oauth_tokensets the field to true only if the stamp was true and the call went to theanthropicprovider; caller-supplied values in their own metadata do not win over the proxy stamp. Failure callbacks stamp the same way from the correct bucket./spend/logs/ui(and v2) acceptused_client_oauth_token=true|false; the Logs UI adds a Credential filter and shows Client OAuth token vs Configured key in the row drawer (legacy rows stay unset and match neither filter).Reviewed by Cursor Bugbot for commit 7b791f3. Bugbot is set up for automated code reviews on this repo. Configure here.