Skip to content

fix(vertex): add the API version to versionless project routes on the Vertex passthrough - #39625

Merged
mateo-berri merged 2 commits into
litellm_internal_stagingfrom
litellm_lit6873_vertex_passthrough_api_version
Sep 3, 2026
Merged

mateo-berri merged 2 commits into
litellm_internal_stagingfrom
litellm_lit6873_vertex_passthrough_api_version

Conversation

@mateo-berri

@mateo-berri mateo-berri commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Claude Code in Vertex mode 404s on every prompt through the Vertex passthrough
  • Its SDK posts /projects/<project>/locations/... routes with no API version
  • The passthrough forwarded those routes unversioned, and Vertex requires /v1
  • The Claude Code gateway docs give the base URL without /v1, so the documented setup fails

How it solves it:

  • A /projects/... route with no version segment gets /v1 prepended before forwarding
  • cachedContent routes get /v1beta1, matching the routes that omit the project
  • Routes that already start with a version are forwarded unchanged

User Flow

Before: a developer running Claude Code in Vertex mode against LiteLLM, with the base URL shape the Claude Code gateway docs give, gets "model not available" on every prompt

  1. They export CLAUDE_CODE_USE_VERTEX=1, CLAUDE_CODE_SKIP_VERTEX_AUTH=1, ANTHROPIC_VERTEX_BASE_URL=https://litellm-domain/vertex_ai, ANTHROPIC_VERTEX_PROJECT_ID=<project>, CLOUD_ML_REGION=global, ANTHROPIC_AUTH_TOKEN=<virtual key> and start claude
  2. They type a prompt, and Claude Code sends POST https://litellm-domain/vertex_ai/projects//locations/global/publishers/anthropic/models/claude-sonnet-4-6:streamRawPredict with the Vertex messages body
  3. The gateway answers 404 with no body; Claude Code falls back to POST https://litellm-domain/vertex_ai/projects//locations/global/publishers/anthropic/models/claude-sonnet-4-6:rawPredict and gets 404 again
  4. The screen shows "The model claude-sonnet-4-6 is not available on your vertex deployment. Try /model to switch..." and every prompt ends the same way
  5. The same request with /vertex_ai/v1/projects/... returns 200, so the only way out is adding /v1 to the base URL by hand, against the docs example

After: the same setup answers the prompt

  1. They export CLAUDE_CODE_USE_VERTEX=1, CLAUDE_CODE_SKIP_VERTEX_AUTH=1, ANTHROPIC_VERTEX_BASE_URL=https://litellm-domain/vertex_ai, ANTHROPIC_VERTEX_PROJECT_ID=<project>, CLOUD_ML_REGION=global, ANTHROPIC_AUTH_TOKEN=<virtual key> and start claude
  2. They type a prompt, and Claude Code sends POST https://litellm-domain/vertex_ai/projects//locations/global/publishers/anthropic/models/claude-sonnet-4-6:streamRawPredict with the Vertex messages body
  3. The gateway answers 200 with the SSE stream (message_start, then text_delta chunks)
  4. The screen shows the model's reply, and the token count call POST https://litellm-domain/vertex_ai/projects//locations/global/publishers/anthropic/models/count-tokens:rawPredict returns 200 with {"input_tokens": N}
  5. A base URL that already ends in /v1 keeps working exactly as before

Relevant issues

Linear ticket

Resolves LIT-6873

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Shared setup: a two-worker proxy with no database, real Vertex AI, claude-sonnet-4-6 in the global location

model_list:
  - model_name: claude-sonnet-4-6
    litellm_params:
      model: vertex_ai/claude-sonnet-4-6
      vertex_project: os.environ/VERTEXAI_PROJECT
      vertex_location: global

default_vertex_config:
  vertex_project: os.environ/VERTEXAI_PROJECT
  vertex_location: global
  vertex_credentials: os.environ/GOOGLE_APPLICATION_CREDENTIALS

general_settings:
  master_key: os.environ/LITELLM_MASTER_KEY
python litellm/proxy/proxy_cli.py --config config.yaml --port 33673 --num_workers 2 --detailed_debug
BASE=http://localhost:33673/vertex-ai/projects/$VERTEXAI_PROJECT/locations/global/publishers/anthropic/models

Control, straight to Vertex with gcloud auth print-access-token, showing Vertex itself needs the version segment:

$ curl -sS -o /dev/null -w '%{http_code}\n' -X POST "https://aiplatform.googleapis.com/v1/projects/$VERTEXAI_PROJECT/locations/global/publishers/anthropic/models/claude-sonnet-4-6:rawPredict" -H "Authorization: Bearer $(gcloud auth print-access-token)" -H 'content-type: application/json' -d '{"anthropic_version":"vertex-2023-10-16","max_tokens":16,"messages":[{"role":"user","content":"Say hi"}]}'
200
$ curl -sS -o /dev/null -w '%{http_code}\n' -X POST "https://aiplatform.googleapis.com/projects/$VERTEXAI_PROJECT/locations/global/publishers/anthropic/models/claude-sonnet-4-6:rawPredict" -H "Authorization: Bearer $(gcloud auth print-access-token)" -H 'content-type: application/json' -d '{"anthropic_version":"vertex-2023-10-16","max_tokens":16,"messages":[{"role":"user","content":"Say hi"}]}'
404

Claude Code setup for case 3 (stock 2.1.251 TUI under tmux, base URL in the shape the Claude Code gateway docs give, no /v1):

export CLAUDE_CODE_USE_VERTEX=1 CLAUDE_CODE_SKIP_VERTEX_AUTH=1
export ANTHROPIC_VERTEX_BASE_URL=http://localhost:33673/vertex-ai
export ANTHROPIC_VERTEX_PROJECT_ID=$VERTEXAI_PROJECT CLOUD_ML_REGION=global
export ANTHROPIC_AUTH_TOKEN=$LITELLM_MASTER_KEY ANTHROPIC_MODEL=claude-sonnet-4-6
claude --debug-file cc.log

Before (9212208)

Case 1: streamRawPredict through the documented base URL shape (no /v1)

  1. Command

    curl -sS -w '\nHTTP %{http_code}\n' -X POST "$BASE/claude-sonnet-4-6:streamRawPredict" -H "Authorization: Bearer $LITELLM_MASTER_KEY" -H 'content-type: application/json' -d '{"anthropic_version":"vertex-2023-10-16","max_tokens":16,"stream":true,"messages":[{"role":"user","content":"Say hi"}]}'
    
  2. Observed output (the proxy log shows it forwarded to https://aiplatform.googleapis.com/projects/<project>/locations/global/publishers/anthropic/models/claude-sonnet-4-6:streamRawPredict, no version)

    HTTP 404
    
  3. Same command with the /vertex_ai/projects/... prefix

    HTTP 404
    

Case 2: count-tokens through the documented base URL shape (no /v1)

  1. Command

    curl -sS -w '\nHTTP %{http_code}\n' -X POST "$BASE/count-tokens:rawPredict" -H "Authorization: Bearer $LITELLM_MASTER_KEY" -H 'content-type: application/json' -d '{"model":"claude-sonnet-4-6","messages":[{"role":"user","content":"hello"}]}'
    
  2. Observed output

    HTTP 404
    

Case 3: Claude Code 2.1.251 in Vertex mode pointed at the documented base URL shape

  1. Type Say hi in exactly three words in the TUI

  2. Observed on screen

    ❯ Say hi in exactly three words
    ⏺ The model claude-sonnet-4-6 is not available on your vertex deployment. Try /model to switch to claude-sonnet-4-5@20250929, or ask your admin to enable this model.
    
  3. Claude Code debug log

    [API REQUEST] /vertex-ai/projects/<project>/locations/global/publishers/anthropic/models/claude-sonnet-4-6:streamRawPredict source=repl_main_thread
    [ERROR] API error (attempt 1/11): 404 404 status code (no body)
    [WARN] Streaming endpoint returned 404, falling back to non-streaming mode
    [ERROR] Non-streaming fallback also failed: 404 status code (no body)
    

Case 4: Gemini generateContent through the documented base URL shape (no /v1)

  1. Command

    curl -sS -w '\nHTTP %{http_code}\n' -X POST "http://localhost:33673/vertex_ai/projects/$VERTEXAI_PROJECT/locations/global/publishers/google/models/gemini-2.5-flash:generateContent" -H "Authorization: Bearer $LITELLM_MASTER_KEY" -H 'content-type: application/json' -d '{"contents":[{"role":"user","parts":[{"text":"Say hi in one word"}]}],"generationConfig":{"maxOutputTokens":16}}'
    
  2. Observed output

    HTTP 404
    

Case 5: cachedContents list through the documented base URL shape (no /v1beta1)

  1. Command

    curl -sS -w '\nHTTP %{http_code}\n' "http://localhost:33673/vertex_ai/projects/$VERTEXAI_PROJECT/locations/global/cachedContents" -H "Authorization: Bearer $LITELLM_MASTER_KEY"
    
  2. Observed output (Google's own HTML 404 page for /projects/<project>/locations/global/cachedContents)

    <p>The requested URL <code>/projects/<project>/locations/global/cachedContents</code> was not found on this server.
    HTTP 404
    

Case 6: discovery dataStores list through the documented base URL shape (no /v1)

  1. Command

    curl -sS -w '\nHTTP %{http_code}\n' "http://localhost:33673/vertex_ai/discovery/projects/$VERTEXAI_PROJECT/locations/global/collections/default_collection/dataStores" -H "Authorization: Bearer $LITELLM_MASTER_KEY"
    
  2. Observed output (Google's own HTML 404 page, the route never reached the Discovery Engine API)

    <p>The requested URL <code>/projects/<project>/locations/global/collections/default_collection/dataStores</code> was not found on this server.
    HTTP 404
    

After (60b725c)

Case 1: streamRawPredict through the documented base URL shape (no /v1)

  1. Command

    curl -sS -w '\nHTTP %{http_code}\n' -X POST "$BASE/claude-sonnet-4-6:streamRawPredict" -H "Authorization: Bearer $LITELLM_MASTER_KEY" -H 'content-type: application/json' -d '{"anthropic_version":"vertex-2023-10-16","max_tokens":16,"stream":true,"messages":[{"role":"user","content":"Say hi"}]}'
    
  2. Observed output (the proxy log now shows https://aiplatform.googleapis.com/v1/projects/<project>/locations/global/publishers/anthropic/models/claude-sonnet-4-6:streamRawPredict)

    event: message_start
    data: {"type":"message_start","message":{"model":"claude-sonnet-4-6","id":"msg_vrtx_011Ceh5uLmMkuJbcFP1Ptb2F","type":"message","role":"assistant",...}}
    data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Hi"}}
    data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":" there"}}
    HTTP 200
    
  3. Same command with the /vertex_ai/projects/... prefix

    HTTP 200
    

Case 2: count-tokens through the documented base URL shape (no /v1)

  1. Command

    curl -sS -w '\nHTTP %{http_code}\n' -X POST "$BASE/count-tokens:rawPredict" -H "Authorization: Bearer $LITELLM_MASTER_KEY" -H 'content-type: application/json' -d '{"model":"claude-sonnet-4-6","messages":[{"role":"user","content":"hello"}]}'
    
  2. Observed output

    {"input_tokens":8}
    HTTP 200
    

Case 3: Claude Code 2.1.251 in Vertex mode pointed at the documented base URL shape

  1. Type Say hi in exactly three words in the TUI

  2. Observed on screen

    ❯ Say hi in exactly three words
    ⏺ Hey there, friend
    
  3. Claude Code debug log (no API errors)

    [API REQUEST] /vertex-ai/projects/<project>/locations/global/publishers/anthropic/models/claude-sonnet-4-6:streamRawPredict source=repl_main_thread
    

Case 4: Gemini generateContent through the documented base URL shape (no /v1)

  1. Command

    curl -sS -w '\nHTTP %{http_code}\n' -X POST "http://localhost:33673/vertex_ai/projects/$VERTEXAI_PROJECT/locations/global/publishers/google/models/gemini-2.5-flash:generateContent" -H "Authorization: Bearer $LITELLM_MASTER_KEY" -H 'content-type: application/json' -d '{"contents":[{"role":"user","parts":[{"text":"Say hi in one word"}]}],"generationConfig":{"maxOutputTokens":16}}'
    
  2. Observed output

    "candidates": [ ... "text": "Hello" ... ]
    HTTP 200
    

Case 5: cachedContents list through the documented base URL shape (no /v1beta1)

  1. Command

    curl -sS -w '\nHTTP %{http_code}\n' "http://localhost:33673/vertex_ai/projects/$VERTEXAI_PROJECT/locations/global/cachedContents" -H "Authorization: Bearer $LITELLM_MASTER_KEY"
    
  2. Observed output (forwarded to /v1beta1/projects/..., an empty list on this project)

    {}
    HTTP 200
    

Case 6: discovery dataStores list through the documented base URL shape (no /v1)

  1. Command

    curl -sS -w '\nHTTP %{http_code}\n' "http://localhost:33673/vertex_ai/discovery/projects/$VERTEXAI_PROJECT/locations/global/collections/default_collection/dataStores" -H "Authorization: Bearer $LITELLM_MASTER_KEY"
    
  2. Observed output (the route now reaches the Discovery Engine API; this service account lacks that permission, and the /v1/ and /v1alpha/ forms of the same route answer the identical 403 on both sides)

    "message": "Permission 'discoveryengine.dataStores.list' denied on resource '//discoveryengine.googleapis.com/projects/<project>/locations/global/collections/default_collection' (or it may not exist). ..."
    HTTP 403
    

Routes that already carry /v1/ returned 200 on both sides ($BASE with /vertex-ai/v1/projects/...), so callers who worked around the bug by adding /v1 keep working

Type

🐛 Bug Fix

Caveats (if any)

Low

  • Only routes whose first segment is projects get a version; anything else is forwarded as before
  • Routes with no /projects/.../locations/... segment still 500 with vertex_location is required
    • pre-existing on the merge base, unrelated to this change, tracked as LIT-6905
  • Versionless routes get v1, or v1beta1 when the route names cached content
    • a beta-only API still needs an explicit /v1beta1/ prefix, which is forwarded verbatim
  • CircleCI local_testing_part1 and local_testing_part2 are red on the three bedrock/cohere.command-r-plus-v1:0 tests
  • The code-quality GitHub Actions job is red on staging itself (get_configured_mode in router.py has no test)
    • not a required check; the router change landed after this branch's merge base

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

  • 60b725c passes /live-pr-risk

@codspeed

codspeed Bot commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_lit6873_vertex_passthrough_api_version (60b725c) with litellm_internal_staging (e4b8cae)

Open in CodSpeed

@greptile-apps

greptile-apps Bot commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

The PR updates Vertex AI passthrough URL construction so versionless /projects/... routes receive the API version required by Vertex while explicitly versioned routes remain unchanged.

  • Uses /v1beta1 for cached-content routes and /v1 for other versionless project routes.
  • Adds typed parameterized coverage for version insertion, project/location replacement, cached content, and already-versioned routes.

Confidence Score: 5/5

The PR appears safe to merge.

No blocking failure remains.

Important Files Changed

Filename Overview
litellm/llms/vertex_ai/common_utils.py Centralizes Vertex route-version selection and prepends a version only to versionless project routes.
tests/test_litellm/llms/vertex_ai/test_vertex_ai_common_utils.py Adds typed parameterized regression coverage; the previously reported missing parameter annotations are now present.

Reviews (2): Last reviewed commit: "test(vertex): type the parametrized vers..." | Re-trigger Greptile

Comment thread tests/test_litellm/llms/vertex_ai/test_vertex_ai_common_utils.py
@codecov

codecov Bot commented Sep 3, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@mateo-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri mateo-berri added run-ci and removed run-ci labels Sep 3, 2026
@mateo-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 60b725c. Configure here.

@tin-berri tin-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@mateo-berri
mateo-berri merged commit 8699998 into litellm_internal_staging Sep 3, 2026
124 of 128 checks passed
@mateo-berri
mateo-berri deleted the litellm_lit6873_vertex_passthrough_api_version branch September 3, 2026 21:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants