Skip to content

fix(passthrough): stop leaking the caller's virtual key on credential-less Vertex passthrough - #38114

Merged
mateo-berri merged 12 commits into
litellm_internal_stagingfrom
litellm_fix_5997_vertex_pt_key_leak
Aug 24, 2026
Merged

fix(passthrough): stop leaking the caller's virtual key on credential-less Vertex passthrough#38114
mateo-berri merged 12 commits into
litellm_internal_stagingfrom
litellm_fix_5997_vertex_pt_key_leak

Conversation

@mateo-berri

@mateo-berri mateo-berri commented Aug 24, 2026

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Credential-less Vertex passthrough forwarded the caller's proxy auth headers to Google
  • Every header the proxy accepts for caller auth leaked upstream: Authorization, x-litellm-api-key, x-goog-api-key, api-key, x-api-key, Ocp-Apim-Subscription-Key, the mapped-route litellm_user_api_key header, and the operator-configured litellm_key_header_name
  • The LiteLLM virtual key, and any other caller secret in those headers, reached a third-party provider

How it solves it:

  • Derive the name-drop set from the canonical SpecialHeaders.litellm_credential_header_names(), so every proxy-only auth header Google never consumes (x-litellm-api-key, api-key, x-api-key, Ocp-Apim-Subscription-Key) is dropped, plus any operator-configured caller-key header (litellm_key_header_name and each pass_through_endpoints entry's litellm_user_api_key), and future additions are covered automatically
  • Resolve the caller's key by the same precedence user_api_key_auth uses and value-strip exactly that value from Authorization / x-goog-api-key (which can instead hold a real Google credential) and from the operator-configured custom key header, normalizing the value with the auth module's own _get_bearer_token so every scheme it accepts (Bearer / bearer / Basic / AWS4-HMAC-SHA256) is matched
  • Fail with a clean 401 when no real Google credential is present, so nothing is forwarded

User Flow

Before: a developer calls Vertex passthrough on a proxy with no Vertex credential configured, and their LiteLLM key is handed to Google

  1. They POST https://litellm-domain/vertex_ai/v1/projects/my-proj/locations/global/publishers/google/models/gemini-2.5-pro:generateContent with Authorization: Bearer sk-... (their LiteLLM virtual key) and a JSON body
  2. The proxy authenticates them, finds no Vertex credential, and forwards the request to https://aiplatform.googleapis.com/v1/projects/my-proj/...:generateContent with that same Authorization: Bearer sk-... still on it
  3. Google receives the developer's LiteLLM virtual key; sending the key in x-litellm-api-key, x-goog-api-key, api-key, x-api-key, or the operator's configured key header instead leaks it the same way
  4. Another party who can read that upstream request now holds a working LiteLLM key and can call the proxy as that developer

After: the same call fails fast with a clean 401, and the key is never forwarded

  1. They POST the same URL with Authorization: Bearer sk-... and the same body
  2. The proxy authenticates them, finds no Vertex credential and no bring-your-own Google credential, and returns 401 saying no Vertex credential is configured and the virtual key is not forwarded
  3. Nothing is sent to https://aiplatform.googleapis.com; sending the key in x-litellm-api-key or x-goog-api-key gives the same 401, and any api-key / x-api-key / configured custom key header is stripped before forwarding
  4. A developer who brings their own Google credential (an OAuth token in Authorization, or a real Google API key in x-goog-api-key) still has it forwarded, now with the LiteLLM key and the other proxy auth headers stripped out
  5. No other party can obtain a LiteLLM virtual key from this branch anymore

Relevant issues

Linear ticket

Resolves LIT-5997

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Screenshots / Proof of Fix

Shared setup (no real virtual key is ever sent to real Google: a local sink intercepts every egress and the capture is the proof)

  • Proxy config has no Vertex credential: no DEFAULT_VERTEXAI_* env vars, and no use_in_pass_through model, so the passthrough takes the credential-less branch
  • A local mitmproxy sink sits on the proxy's egress. It intercepts every request whose host ends in googleapis.com, records the exact headers that would have gone to Google, and returns a synthetic 599 so nothing reaches Google. A leak shows up as the virtual key (or any caller secret) appearing in that capture
  • Generate a virtual key against the proxy: curl -sX POST http://127.0.0.1:PORT/key/generate -H "<auth header>: Bearer $MASTER_KEY" -d '{"duration":"2h"}' returns sk-…
  • Payload for every call: {"contents":[{"role":"user","parts":[{"text":"hi"}]}]}
  • The custom-key-header cases (Case G) run against a second proxy configured with general_settings.litellm_key_header_name: x-company-key; every other case runs against a default-config proxy

Before (28b433a)

Case A: virtual key in Authorization

  1. curl -sX POST http://127.0.0.1:49346/vertex_ai/v1/projects/my-proj/locations/global/publishers/google/models/gemini-2.5-pro:generateContent -H "Authorization: Bearer sk-…" -d '<payload>'
  2. Sink capture of the outbound to aiplatform.googleapis.com: authorization: Bearer sk-…, the virtual key was on its way to Google

Case B: virtual key in x-litellm-api-key

  1. curl -sX POST http://127.0.0.1:49346/vertex_ai/.../gemini-2.5-pro:generateContent -H "x-litellm-api-key: sk-…" -d '<payload>'
  2. Sink capture of the outbound to aiplatform.googleapis.com: x-litellm-api-key: sk-…, the virtual key was on its way to Google

Case D: virtual key in x-goog-api-key

  1. curl -sX POST http://127.0.0.1:49346/vertex_ai/.../gemini-2.5-pro:generateContent -H "x-litellm-api-key: sk-…" -H "x-goog-api-key: sk-…" -d '<payload>'
  2. Sink capture of the outbound to aiplatform.googleapis.com: x-goog-api-key: sk-… and x-litellm-api-key: sk-…. The virtual key was on its way to Google, this time posing as a Google API key

Case F: distinct caller secrets in api-key and x-api-key

  1. curl -sX POST http://127.0.0.1:49346/vertex_ai/.../gemini-2.5-pro:generateContent -H "x-litellm-api-key: sk-…" -H "Authorization: Bearer ya29.fake-google-oauth" -H "api-key: azure-secret-abc123" -H "x-api-key: anthropic-secret-def456" -d '<payload>'
  2. Sink capture of the outbound to aiplatform.googleapis.com: authorization: Bearer ya29.fake-google-oauth, api-key: azure-secret-abc123, x-api-key: anthropic-secret-def456, and x-litellm-api-key: sk-…. Every proxy auth header, including two unrelated caller secrets and the virtual key, went to Google

Case G: virtual key in the operator-configured custom key header

Against a merge-base proxy configured with general_settings.litellm_key_header_name: x-company-key

  1. curl -sX POST http://127.0.0.1:42826/vertex_ai/.../gemini-2.5-pro:generateContent -H "x-company-key: Bearer sk-…" -H "x-goog-api-key: AIzaSyReal-Google-Api-Key-000" -d '<payload>'
  2. Sink capture of the outbound to aiplatform.googleapis.com: x-company-key: Bearer sk-… and x-goog-api-key: AIzaSyReal-Google-Api-Key-000. The caller authenticated with the custom header, and that same header carrying the virtual key was forwarded to Google alongside the real Google key

Case H: caller Azure APIM secret in Ocp-Apim-Subscription-Key

  1. curl -sX POST http://127.0.0.1:49346/vertex_ai/.../gemini-2.5-pro:generateContent -H "x-litellm-api-key: sk-…" -H "Authorization: Bearer ya29.fake-google-oauth" -H "Ocp-Apim-Subscription-Key: azure-apim-secret-xyz789" -d '<payload>'
  2. Sink capture of the outbound to aiplatform.googleapis.com: authorization: Bearer ya29.fake-google-oauth and Ocp-Apim-Subscription-Key: azure-apim-secret-xyz789. The caller's Azure APIM subscription key, an auth header the proxy accepts but Google never consumes, was forwarded to Google

Case J: virtual key echoed into Authorization with a Basic scheme

  1. curl -sX POST http://127.0.0.1:49346/vertex_ai/.../gemini-2.5-pro:generateContent -H "x-litellm-api-key: sk-…" -H "Authorization: Basic sk-…" -d '<payload>'
  2. Sink capture of the outbound to aiplatform.googleapis.com: authorization: Basic sk-…. The caller authenticated with x-litellm-api-key and echoed the same virtual key into Authorization under a Basic scheme, and that header carrying the key was forwarded to Google

After (16a81c9)

The header-filtering behavior in Cases A-H is stable across the hardening commits and was captured against the default-config proxy on port 41337 and the custom-key-header proxy on port 41779. Case I was run against a fresh default-config proxy on port 40923, and Case J plus the no-regression re-check of Cases C and E against a fresh default-config proxy on port 40611 at this exact tip.

Case A: virtual key in Authorization

  1. curl -sw '%{http_code}' -X POST http://127.0.0.1:41337/vertex_ai/.../gemini-2.5-pro:generateContent -H "Authorization: Bearer sk-…" -d '<payload>'
  2. Response: HTTP 401 with {"detail":"No Vertex AI credential is configured on this proxy and the request carried no upstream Google credential. The LiteLLM virtual key is not forwarded to Google. ..."}. The sink recorded nothing new: no request reached googleapis.com

Case B: virtual key in x-litellm-api-key

  1. curl -sw '%{http_code}' -X POST http://127.0.0.1:41337/vertex_ai/.../gemini-2.5-pro:generateContent -H "x-litellm-api-key: sk-…" -d '<payload>'
  2. Response: HTTP 401 with the same "no Vertex AI credential is configured" detail. The sink recorded nothing new: no request reached googleapis.com

Case C: bring-your-own Google token, virtual key in x-litellm-api-key for proxy auth

  1. curl -sw '%{http_code}' -X POST http://127.0.0.1:41337/vertex_ai/.../gemini-2.5-pro:generateContent -H "x-litellm-api-key: sk-…" -H "Authorization: Bearer ya29.fake-google-oauth" -d '<payload>'
  2. Response: HTTP 599 from the sink (the request was forwarded). Sink capture of the outbound to aiplatform.googleapis.com: authorization: Bearer ya29.fake-google-oauth, x-litellm-api-key: <absent>, and the virtual key is nowhere in the outbound. Real bring-your-own passthrough still works, with the key stripped

Case D: virtual key in x-goog-api-key

  1. curl -sw '%{http_code}' -X POST http://127.0.0.1:41337/vertex_ai/.../gemini-2.5-pro:generateContent -H "x-litellm-api-key: sk-…" -H "x-goog-api-key: sk-…" -d '<payload>'
  2. Response: HTTP 401 with the same "no Vertex AI credential is configured" detail. The sink recorded nothing new: no request reached googleapis.com. The key posing as a Google API key no longer satisfies the gate

Case E: bring-your-own real Google API key in x-goog-api-key, virtual key in x-litellm-api-key for proxy auth

  1. curl -sw '%{http_code}' -X POST http://127.0.0.1:41337/vertex_ai/.../gemini-2.5-pro:generateContent -H "x-litellm-api-key: sk-…" -H "x-goog-api-key: AIzaSyReal-Google-Api-Key-000" -d '<payload>'
  2. Response: HTTP 599 from the sink (the request was forwarded). Sink capture of the outbound to aiplatform.googleapis.com: x-goog-api-key: AIzaSyReal-Google-Api-Key-000, x-litellm-api-key: <absent>, and the virtual key is nowhere in the outbound. A real Google API key that differs from the virtual key still forwards, with the key stripped

Case F: distinct caller secrets in api-key and x-api-key, bring-your-own Google token for the real credential

  1. curl -sw '%{http_code}' -X POST http://127.0.0.1:41337/vertex_ai/.../gemini-2.5-pro:generateContent -H "x-litellm-api-key: sk-…" -H "Authorization: Bearer ya29.fake-google-oauth" -H "api-key: azure-secret-abc123" -H "x-api-key: anthropic-secret-def456" -d '<payload>'
  2. Response: HTTP 599 from the sink (the request was forwarded). Sink capture of the outbound to aiplatform.googleapis.com: authorization: Bearer ya29.fake-google-oauth only. x-litellm-api-key, api-key, and x-api-key are all <absent>, and neither azure-secret-abc123, anthropic-secret-def456, nor the virtual key appears anywhere. The Google token still forwards, every proxy auth header is dropped

Case G: virtual key in the operator-configured custom key header

Against the proxy configured with general_settings.litellm_key_header_name: x-company-key

  1. curl -sw '%{http_code}' -X POST http://127.0.0.1:41779/vertex_ai/.../gemini-2.5-pro:generateContent -H "x-company-key: Bearer sk-…" -H "x-goog-api-key: AIzaSyReal-Google-Api-Key-000" -d '<payload>'
  2. Response: HTTP 599 from the sink (the request was forwarded). Sink capture of the outbound to aiplatform.googleapis.com: x-goog-api-key: AIzaSyReal-Google-Api-Key-000 only, x-company-key: <absent>, and the virtual key is nowhere in the outbound. The caller still authenticates with the custom header, the real Google key still forwards, and the virtual key is stripped
  3. Sending the virtual key in x-company-key with no real Google credential returns HTTP 401 and nothing reaches googleapis.com

Case H: caller Azure APIM secret in Ocp-Apim-Subscription-Key, legitimate Google header preserved

  1. curl -sw '%{http_code}' -X POST http://127.0.0.1:41337/vertex_ai/.../gemini-2.5-pro:generateContent -H "x-litellm-api-key: sk-…" -H "Authorization: Bearer ya29.fake-google-oauth" -H "Ocp-Apim-Subscription-Key: azure-apim-secret-xyz789" -H "X-Goog-User-Project: my-billing-proj" -d '<payload>'
  2. Response: HTTP 599 from the sink (the request was forwarded). Sink capture of the outbound to aiplatform.googleapis.com: authorization: Bearer ya29.fake-google-oauth and X-Goog-User-Project: my-billing-proj, while Ocp-Apim-Subscription-Key is <absent> and azure-apim-secret-xyz789 appears nowhere. The caller's Azure APIM secret is dropped, and the genuine Google X-Goog-User-Project header is preserved so real Vertex requests keep working

Case I: virtual key resolved through the full auth precedence, and no regression to real credentials

The route authenticates through Depends(user_api_key_auth), which accepts the caller key from every header in SpecialHeaders.litellm_credential_header_names(), x-goog-api-key included. The filter now resolves the caller key by that same precedence and value-strips exactly the value that authenticated, so x-goog-api-key is stripped when it carried the key and preserved when it carried a real Google key alongside a higher-precedence virtual key. All against the port 40923 proxy at this tip:

  1. curl -sw '%{http_code}' -X POST http://127.0.0.1:40923/vertex_ai/.../gemini-2.5-pro:generateContent -H "x-goog-api-key: sk-…" -d '<payload>' returns HTTP 401 and the sink records nothing new (the proxy's own auth rejects a virtual key in x-goog-api-key on this route today, and the filter would strip it regardless, closing the boundary that a mocked unit test exercises directly)
  2. Case C re-run (x-litellm-api-key: sk-… + Authorization: Bearer ya29.fake-google-oauth) still returns HTTP 599 with authorization: Bearer ya29.fake-google-oauth forwarded and the virtual key absent
  3. Case E re-run (x-litellm-api-key: sk-… + x-goog-api-key: AIzaSyReal-Google-Api-Key-000) still returns HTTP 599 with x-goog-api-key: AIzaSyReal-Google-Api-Key-000 forwarded and the virtual key absent, so resolving by precedence does not strip a genuine Google key

Case J: virtual key echoed into Authorization with a Basic scheme, and no regression to real credentials

Against a fresh default-config proxy on port 40611 at this tip:

  1. curl -sw '%{http_code}' -X POST http://127.0.0.1:40611/vertex_ai/.../gemini-2.5-pro:generateContent -H "x-litellm-api-key: sk-…" -H "Authorization: Basic sk-…" -d '<payload>' returns HTTP 401 and the sink records nothing new. Normalizing the value with the auth module's own _get_bearer_token strips the Basic scheme just as authentication does, so the echoed virtual key matches the authenticated key and Authorization is dropped, leaving no upstream credential
  2. Case C re-run (x-litellm-api-key: sk-… + Authorization: Bearer ya29.fake-google-oauth) still returns HTTP 599 with authorization: Bearer ya29.fake-google-oauth forwarded and the virtual key absent
  3. Case E re-run (x-litellm-api-key: sk-… + x-goog-api-key: AIzaSyReal-Google-Api-Key-000) still returns HTTP 599 with x-goog-api-key: AIzaSyReal-Google-Api-Key-000 forwarded and the virtual key absent

Case K: virtual key in the mapped-route litellm_user_api_key header, with real credentials preserved

The /vertex_ai prefix is a mapped pass-through route, so user_api_key_auth accepts the caller key from a header literally named litellm_user_api_key and applies it last, overriding every other source. Against fresh default-config proxies on ports 40611 (before this commit) and 40337 (this tip), same request each side: litellm_user_api_key: sk-… (auth) + Authorization: Bearer ya29.fake-google-oauth + x-goog-api-key: AIzaSyReal-Google-Api-Key-000.

  1. Before (port 40611): HTTP 599, and the sink capture to aiplatform.googleapis.com shows litellm_user_api_key: sk-… forwarded with the virtual key, while the real Authorization was dropped. The virtual key reached Google and the bring-your-own token was lost
  2. After (port 40337): HTTP 599, and the sink capture shows authorization: Bearer ya29.fake-google-oauth and x-goog-api-key: AIzaSyReal-Google-Api-Key-000 both forwarded, litellm_user_api_key <absent>, and the virtual key nowhere. The key that authenticated is stripped, and both real credentials are preserved

Type

🐛 Bug Fix

Caveats (if any)

Scope of this PR is the credential-less Vertex passthrough leaking the caller's LiteLLM credential through request headers. One adjacent vector is intentionally left for a separate, focused change: a caller can also send a virtual key in the ?key= URL query parameter (the Google AI Studio auth convention that user_api_key_auth reads on generateContent routes). A client that authenticates only with ?key= is already rejected with a 401 on this branch (no surviving upstream Google credential), so that common case does not leak. The virtual key does still ride the forwarded URL query when the caller both authenticates with a header and brings a real Google credential, but stripping credential query params belongs in the shared pass-through URL-forwarding path (it affects every provider, not just Vertex) and is being tracked as its own follow-up rather than widening this PR's blast radius.

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Live PR risk

/live-pr-risk CHECKED e6eb6a4: SAFE. The graph is tiny and fully local: _forwarded_headers_for_credentialless_vertex_passthrough is called only by _prepare_vertex_auth_headers, which is called only by _base_vertex_proxy_route, which is reached by the two public routes vertex_proxy_route and vertex_discovery_proxy_route. The re-signed _prepare_vertex_auth_headers now returns Mapping[str, str] and its credential-less branch raises HTTPException(401); nothing downstream mutates the returned headers, and create_pass_through_route copies them with dict(...), so the immutable MappingProxyType is safe on every path. vertex_proxy_route was driven live in the proof above; vertex_discovery_proxy_route runs the identical credential-less branch and is covered by unit tests. Direct-caller and related passthrough tests pass on the fixed head (160 passed), including test_vertex_passthrough_load_balancing.py, which unpacks the new tuple directly.

/live-pr-risk CHECKED 4bc0977: SAFE. This commit hardens the same branch to strip the virtual key from every forwarded header by value (normalizing any Bearer prefix) instead of only from Authorization by name, closing the x-goog-api-key vector Greptile flagged. Re-walked the graph for the added code: the new _bearer_stripped is a pure module-private helper referenced only inside _forwarded_headers_for_credentialless_vertex_passthrough, and that function still has exactly one caller, so the blast radius is identical to the commit above. No new dependents, side effects, or raises.

/live-pr-risk CHECKED e7c2ede: SAFE. This commit additionally drops the proxy-only auth headers Google never consumes (x-litellm-api-key, api-key, x-api-key) by name via a new module-private _HEADERS_NEVER_FORWARDED_TO_VERTEX frozenset, keeping the by-value virtual-key strip for Authorization / x-goog-api-key. The frozenset is referenced only inside _forwarded_headers_for_credentialless_vertex_passthrough, which still has exactly one caller, so the blast radius is unchanged; no new dependents, side effects, or raises, and the gate is untouched.

/live-pr-risk CHECKED ee03632: SAFE. This commit closes the last vector Greptile flagged: user_api_key_auth also authenticates a caller from the operator-configured general_settings.litellm_key_header_name, read straight off the request, so a virtual key sent there survived the filter. The new module-private _credentialless_caller_key_values reads general_settings (read-only, via the same lazy import already used elsewhere in the module) and returns the set of accepted key values; the filter now value-strips any header matching one of them. Blast radius unchanged: both the helper and the filter are called only from the single existing caller, no new raises (the 401 gate is untouched), no mutation, no signature change. Verified live on both sides: at the merge-base with litellm_key_header_name: x-company-key set, x-company-key: Bearer <vkey> forwarded to aiplatform.googleapis.com alongside a real Google key; on this tip the same request forwards only the real Google key with the custom header and the virtual key stripped, and the custom header alone returns 401. Cases A-F re-verified unchanged on this tip.

/live-pr-risk CHECKED ab93636: SAFE. This commit replaces the hand-rolled name-drop set with one derived from the canonical SpecialHeaders.litellm_credential_header_names() minus the two headers that double as real Google credentials (Authorization, x-goog-api-key), which are value-stripped instead. This is the same source user_api_key_auth reads the caller's key from, so the drop set now cannot drift out of sync with what authenticates, and it picked up Ocp-Apim-Subscription-Key, which the hand-rolled set missed. SpecialHeaders is a pure enum already imported into the module via its _types star import; the derived frozenset is evaluated once at import with no side effects, and both module-level constants are referenced only inside the single-caller filter, so the blast radius is unchanged. Verified live on both sides: at the merge-base a distinct caller Azure APIM secret in Ocp-Apim-Subscription-Key forwarded to aiplatform.googleapis.com; on this tip it is dropped while the genuine Google X-Goog-User-Project header is preserved, and all of Cases A-G re-verified unchanged.

/live-pr-risk CHECKED f3dc339: SAFE. This commit resolves the caller key by the same precedence get_api_key uses (custom litellm_key_header_name, then x-litellm-api-key, Authorization, api-key, x-api-key, x-goog-api-key, Ocp-Apim-Subscription-Key) and value-strips exactly the one value that authenticated, instead of only the value from x-litellm-api-key / Authorization / the custom header. This closes the structural gap Greptile flagged: x-goog-api-key is an accepted auth source that the old caller-key set omitted while keeping the header, so a virtual key authenticated through it would have been forwarded. The precedence tuple is built once at import from the same SpecialHeaders enum; _authenticated_caller_key_values reads request headers and general_settings read-only and is still called only by the single-caller filter, so the blast radius is unchanged, no new raises, no mutation. Verified live at this tip: a real Google key in x-goog-api-key alongside a higher-precedence virtual key in x-litellm-api-key is still forwarded (Case E, 599), the bring-your-own OAuth token is still forwarded (Case C, 599), and a virtual key sent only in x-goog-api-key is rejected. Note the proxy's own auth currently rejects a virtual key presented in x-goog-api-key on this route before the filter runs, so this commit is defense-in-depth on the forwarding boundary, proven directly by the added unit tests.

/live-pr-risk CHECKED 5d8286c: SAFE. This commit swaps the filter's own Bearer-only stripping for the auth module's _get_bearer_token, so the caller-key comparison normalizes exactly the schemes authentication accepts (Bearer / bearer / Basic / AWS4-HMAC-SHA256), with a raw-value fallback for a bare token. It closes the case Greptile flagged: a virtual key echoed as Authorization: Basic <key> alongside a higher-precedence auth header did not match the caller key under the old normalization and was forwarded. _get_bearer_token is a pure function in user_api_key_auth (already imported into this module for user_api_key_auth), with no side effects; the new _normalize_credential_value wrapper is referenced only by the single-caller resolver and filter, so the blast radius is unchanged. Verified live on both sides: at the merge-base Authorization: Basic <vkey> forwarded to aiplatform.googleapis.com; on this tip it returns 401 with nothing forwarded, while the bring-your-own OAuth token (Case C) and a real Google key (Case E) still forward with the virtual key absent.

/live-pr-risk CHECKED 2fe1e7e: SAFE. user_api_key_auth also accepts the caller key from a pass_through_endpoints entry's headers.litellm_user_api_key, not only litellm_key_header_name. This commit adds _operator_configured_caller_key_header_names, which reads both from general_settings (read-only), drops every configured caller-key header by name, and feeds them as top-precedence caller-key sources into the resolver. The helper is pure and referenced only by the single-caller resolver and filter, so the blast radius is unchanged, no new raises, no mutation. Config values are read defensively with isinstance guards. Covered by unit tests that configure each source through general_settings and assert the configured header is dropped while a real Google key in x-goog-api-key is preserved; the full passthrough test file (160 tests) passes, including the LIT-4761 streaming-classification suite whose fixture now sends the virtual key in x-litellm-api-key, matching a real request.

/live-pr-risk CHECKED fcc047b: SAFE. Corrects the precedence of the operator-configured key headers to match get_api_key exactly: litellm_key_header_name overrides everything so it resolves first, then the built-in headers in get_api_key order, then a pass_through_endpoints litellm_user_api_key header which get_api_key checks last. The prior commit had lumped both configured sources at the top, so a request that authenticated via Authorization while also carrying a pass-through header could pick the wrong value and leave the authenticated Authorization key forwarded. _operator_configured_caller_key_header_names now returns (override, pass_through) and both the resolver ordering and the name-drop consume it; still pure, still called only by the single-caller resolver and filter, no new raises or mutation. Covered by a new unit test where Authorization holds the authenticated key, a pass-through header holds a decoy, and a real x-goog-api-key is present: the Authorization key is stripped, the pass-through header dropped, and the real Google key preserved. Full passthrough test file (161 tests) green.

/live-pr-risk CHECKED 16a81c9: SAFE. Closes a high-severity vector Cursor Bugbot flagged: on mapped pass-through routes (of which /vertex_ai is one), user_api_key_auth accepts the caller key from a header literally named litellm_user_api_key via check_api_key_for_custom_headers_or_pass_through_endpoints, applied last so it overrides every other source. The filter now drops that header by name and resolves it at highest precedence. Adds a module constant plus a prepend to the resolver order and a union into the name-drop set; still pure, still called only by the single-caller resolver and filter, no new raises or mutation. Verified live on both sides: at the previous tip a virtual key in litellm_user_api_key forwarded to aiplatform.googleapis.com while the real Authorization was wrongly stripped; on this tip the virtual key is dropped and both the bring-your-own OAuth token and a real Google key are preserved. Full passthrough test file (163 tests) green.


Note

High Risk
Security-sensitive change to Vertex passthrough auth and header forwarding; wrong filtering could break BYO-Google flows or still leak secrets upstream.

Overview
Fixes LIT-5997: when the proxy has no Vertex credential, the bring-your-own-credentials branch no longer forwards the full incoming header set to Google.

Credential-less Vertex passthrough now builds upstream headers via _forwarded_headers_for_credentialless_vertex_passthrough instead of copying all request headers. Proxy-only auth headers (from SpecialHeaders.litellm_credential_header_names() except Authorization / x-goog-api-key, plus operator litellm_key_header_name and pass-through key headers) are dropped by name. The caller key is resolved with the same precedence as user_api_key_auth and stripped by value from any remaining header (using _get_bearer_token for scheme normalization). If neither a surviving Authorization nor x-goog-api-key remains, the route returns 401 with guidance instead of calling Google.

Bring-your-own Google OAuth or API keys still forward; LiteLLM virtual keys in Authorization, x-litellm-api-key, x-goog-api-key, or custom headers no longer reach upstream. Tests were updated and expanded (TestVertexCredentiallessPassthroughVirtualKeyLeak) for these cases.

Reviewed by Cursor Bugbot for commit 16a81c9. Bugbot is set up for automated code reviews on this repo. Configure here.

…-less Vertex passthrough

When no Vertex credential is configured (no default_vertex_config, no matching
use_in_pass_through deployment, no vector-store credential), the Vertex passthrough
took the bring-your-own-credentials branch and forwarded the entire incoming header
set upstream to Google. That set included whichever header carried the caller's
LiteLLM virtual key: x-litellm-api-key, or Authorization when get_litellm_virtual_key
read the key from there. The proxy's own secret was sent to a third-party provider.

The credential-less branch now drops x-litellm-api-key and the Authorization value
that equals the virtual key, keeping a genuine bring-your-own Google credential
(an OAuth token in Authorization, or x-goog-api-key) so real BYO passthrough still
works. When neither survives, the request fails with a clean 401 telling the operator
no credential is configured, instead of forwarding the virtual key.

Regression coverage in the mapped test path asserts the 401-and-never-forwarded
behavior for both leak vectors and that a real Google credential still passes through
with the virtual key stripped.
@codecov

codecov Bot commented Aug 24, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 62.50000% with 15 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
...ass_through_endpoints/llm_passthrough_endpoints.py 62.50% 15 Missing ⚠️

📢 Thoughts on this report? Let us know!

@greptile-apps

greptile-apps Bot commented Aug 24, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR hardens credential-less Vertex passthrough handling and rejects requests that lack a usable upstream credential

  • Derives proxy-only credential headers from the canonical authentication-header set
  • Matches caller-key precedence across built-in, mapped, and operator-configured headers
  • Preserves distinct bring-your-own Google credentials while removing proxy credentials
  • Adds focused regression coverage for the previously reported header and normalization cases

Confidence Score: 5/5

The PR appears safe to merge

No blocking failure remains

Important Files Changed

Filename Overview
litellm/proxy/pass_through_endpoints/llm_passthrough_endpoints.py Adds credential-aware filtering and an early rejection path; the fixes address all displayed prior findings
tests/test_litellm/proxy/pass_through_endpoints/test_llm_pass_through_endpoints.py Adds regression tests for credential source precedence, configured headers, supported schemes, and bring-your-own credentials

Reviews (11): Last reviewed commit: "fix(vertex-passthrough): cover the mappe..." | Re-trigger Greptile

Comment thread litellm/proxy/pass_through_endpoints/llm_passthrough_endpoints.py Outdated
…ss Vertex forward

The credential-less Vertex passthrough dropped the caller's LiteLLM
virtual key only from Authorization by exact match. A caller who sent
the same key in x-goog-api-key (which doubles as a real Google
credential) had it accepted as a credential and forwarded upstream.

Drop the virtual key by value across every forwarded header, normalizing
any Bearer prefix, so no header name carries it to Google.
@mateo-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread litellm/proxy/pass_through_endpoints/llm_passthrough_endpoints.py
Adds a regression asserting the value-based strip also drops the caller's
virtual key when it is duplicated into the api-key and x-api-key headers,
while a genuine bring-your-own Google credential still forwards.
On the credential-less Vertex passthrough branch, drop every header that
can only carry LiteLLM caller auth (x-litellm-api-key, api-key, x-api-key)
by name, since Google never consumes them, and strip the virtual key by
value from Authorization / x-goog-api-key, which may instead hold a genuine
bring-your-own Google credential. This closes the residual leak where a
distinct caller secret in api-key or x-api-key still reached upstream.
@mateo-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread litellm/proxy/pass_through_endpoints/llm_passthrough_endpoints.py Outdated
@veria-ai

veria-ai Bot commented Aug 24, 2026

Copy link
Copy Markdown
Contributor

PR overview

All previously flagged issues have been addressed. No open security concerns remain on this pull request.

Security review

No open security issues remain on this pull request.

Fixed/addressed: 2 · PR risk: 0/10

Comment thread litellm/proxy/pass_through_endpoints/llm_passthrough_endpoints.py Outdated
user_api_key_auth also authenticates a caller from the operator-configured
general_settings.litellm_key_header_name, reading that header straight off
the request, so a virtual key sent there survived the credential-less Vertex
forwarding filter and reached Google alongside a real bring-your-own
credential. Value-strip every header whose value matches the caller's key
from any accepted source, including that custom header.
@mateo-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread litellm/proxy/pass_through_endpoints/llm_passthrough_endpoints.py Outdated
…alHeaders

The hand-rolled drop set missed Ocp-Apim-Subscription-Key, so a caller
Azure APIM secret in that header was forwarded to Google on the
credential-less branch. Derive the name-drop set from the canonical
SpecialHeaders.litellm_credential_header_names(), minus Authorization and
x-goog-api-key which double as real Google credentials and are value-stripped
instead. New credential headers added there are now dropped automatically.
@mateo-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread litellm/proxy/pass_through_endpoints/llm_passthrough_endpoints.py Outdated
The credential-less filter derived the caller key only from x-litellm-api-key,
Authorization, and the custom header, but the route authenticates through
Depends(user_api_key_auth), which also accepts the key from x-goog-api-key. A
virtual key sent only in x-goog-api-key therefore authenticated yet was kept as
a preserved upstream header and forwarded to Google. Resolve the caller key by
the same precedence get_api_key uses and value-strip exactly that, so a key in
x-goog-api-key is stripped while a real Google key alongside a higher-precedence
virtual key is preserved.
@mateo-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread litellm/proxy/pass_through_endpoints/llm_passthrough_endpoints.py
…er_token

The filter's own Bearer-only stripping missed the other schemes
user_api_key_auth accepts, so a virtual key echoed as `Authorization: Basic
<key>` alongside a higher-precedence auth header did not match the caller key
and was forwarded to Google. Reuse the auth module's _get_bearer_token so the
comparison strips exactly what authentication does (Bearer / bearer / Basic /
AWS4-HMAC-SHA256), falling back to the raw value for a bare token.
@mateo-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

…in streaming tests

The LIT-4761 streaming-classification tests passed only the bring-your-own
Google OAuth token in Authorization and mocked get_litellm_virtual_key, a shape
that cannot authenticate in production. The credential-less filter now resolves
the caller key by auth precedence, so a lone Authorization value reads as the
key and is stripped. Send the virtual key in x-litellm-api-key, matching a real
request, so Authorization is preserved and the classification assertions run.
@mateo-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread litellm/proxy/pass_through_endpoints/llm_passthrough_endpoints.py
…eaders

user_api_key_auth also accepts the caller key from a pass_through_endpoints
entry's headers.litellm_user_api_key, not just litellm_key_header_name. Drop
every operator-configured caller-key header by name and treat them as
top-precedence caller-key sources, so a virtual key sent through one is never
forwarded to Google.
@mateo-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread litellm/proxy/pass_through_endpoints/llm_passthrough_endpoints.py Outdated
Comment thread litellm/proxy/pass_through_endpoints/llm_passthrough_endpoints.py Outdated
…key headers

The resolver placed both operator-configured key headers at the top of its
precedence, but user_api_key_auth only overrides with litellm_key_header_name;
a pass_through_endpoints litellm_user_api_key is checked last. So a request that
authenticated via Authorization while also sending a pass-through header could
have the wrong value chosen, leaving the authenticated Authorization key
forwarded. Order the resolver exactly like get_api_key: override first, built-in
headers next, pass-through header last.
@mateo-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

Comment thread litellm/proxy/pass_through_endpoints/llm_passthrough_endpoints.py
…header

On mapped pass-through routes, of which /vertex_ai is one,
user_api_key_auth accepts the caller key from a header literally named
litellm_user_api_key and applies it last, so it overrides every other source.
The credential-less filter neither dropped it nor resolved the caller key from
it, so a virtual key there reached Google past a real x-goog-api-key, and a
bring-your-own Authorization could be stripped when auth actually came from that
header. Drop it by name and resolve it at highest precedence.
@mateo-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@mateo-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 16a81c9. Configure here.

@tin-berri tin-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving — real leak, right fix. Forwarding _safe_get_request_headers(request).copy() wholesale meant the caller's virtual key went to Google on every credential-less Vertex passthrough, and splitting it into "drop the proxy-only auth headers by name, drop the value that actually authenticated wherever it appears, 401 if nothing usable survives" is the correct decomposition. Reusing _get_bearer_token from user_api_key_auth rather than re-deriving the scheme stripping is exactly right — that comparison has to agree with authentication or it strips the wrong thing.

Traced the cases that decide whether this is correct or just looks correct:

  • virtual key in Authorization only → stripped by value, nothing survives, 401 with the actionable message. Right.
  • virtual key in x-litellm-api-key + real Google token in Authorization → key dropped by name, Authorization value ≠ caller key so it's preserved. Right.
  • virtual key in x-goog-api-key → resolved as the authenticated value and dropped by value, not kept just because Google consumes that header. This is the case a by-name-only fix would have missed.

Also checked the dictMapping return-type change for runtime breakage: create_pass_through_route does dict(param_custom_headers) at the pass_through_request call, and forward_headers_from_request rebinds ({**request_headers, **headers}) rather than mutating in place, so the MappingProxyType never reaches anything that writes to it.

One thing worth a guard. The whole fix rests on forward_headers being False on this route, and nothing here says so. If it's ever True, forward_headers_from_request merges the raw incoming headers back in — and it only pops a request header when that name is already present in headers, so every header this PR dropped by name (x-litellm-api-key, api-key, x-api-key, Ocp-Apim-Subscription-Key, litellm_user_api_key, the operator-configured ones) gets re-added verbatim and the leak is back. Today _base_vertex_proxy_route doesn't pass _forward_headers and the default is False, so it's safe — but that's an invisible dependency holding up a security fix. Either assert it at the call site or note it where _forward_headers_for_credentialless_vertex_passthrough is defined, so someone enabling header forwarding on this route trips over it.

code-quality is the recursive_detector red on llm_request_utils.py — base drift, and #38149 fixes it.

@mateo-berri
mateo-berri merged commit 46c2328 into litellm_internal_staging Aug 24, 2026
80 of 82 checks passed
@mateo-berri
mateo-berri deleted the litellm_fix_5997_vertex_pt_key_leak branch August 24, 2026 21:54
@codspeed-hq

codspeed-hq Bot commented Aug 24, 2026

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_fix_5997_vertex_pt_key_leak (16a81c9) with litellm_internal_staging (a91cac7)1

Open in CodSpeed

Footnotes

  1. No successful run was found on litellm_internal_staging (6147b3c) during the generation of this report, so a91cac7 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report.

mateo-berri added a commit to FelipeRodriguesGare/litellm that referenced this pull request Aug 26, 2026
…ertex passthrough

PR BerriAI#38114 dropped whichever header user_api_key_auth would read the caller's
key from, by precedence. Under custom_auth, JWT auth, or no master key that
header is the caller's own Google token, so the bring-your-own-credentials
Vertex branch answered 401 to every valid request.

A header value is now dropped only when it is the master key or when its
hash is the api_key that authenticated the request, so a Google token that
auth never consumed keeps flowing while a LiteLLM key still never reaches
Google.

test_passthrough_post_call_guardrails.py no longer plants a MagicMock
proxy_server module in sys.modules at import, which poisoned sibling tests
that read module globals at call time.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants