Repository navigation
fix(vertex): add the API version to versionless project routes on the Vertex passthrough - #39625
Merged
mateo-berri merged 2 commits intoSep 3, 2026
Conversation
… Vertex passthrough
Contributor
Contributor
Greptile SummaryThe PR updates Vertex AI passthrough URL construction so versionless
Confidence Score: 5/5The PR appears safe to merge. No blocking failure remains.
|
| Filename | Overview |
|---|---|
| litellm/llms/vertex_ai/common_utils.py | Centralizes Vertex route-version selection and prepends a version only to versionless project routes. |
| tests/test_litellm/llms/vertex_ai/test_vertex_ai_common_utils.py | Adds typed parameterized regression coverage; the previously reported missing parameter annotations are now present. |
Reviews (2): Last reviewed commit: "test(vertex): type the parametrized vers..." | Re-trigger Greptile
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
Contributor
Author
Contributor
Author
|
bugbot run |
Contributor
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 60b725c. Configure here.
mateo-berri
enabled auto-merge
September 3, 2026 20:45
mateo-berri
merged commit Sep 3, 2026
8699998
into
litellm_internal_staging
124 of 128 checks passed
mateo-berri
deleted the
litellm_lit6873_vertex_passthrough_api_version
branch
September 3, 2026 21:16
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
TLDR
Problem this solves:
/projects/<project>/locations/...routes with no API version/v1/v1, so the documented setup failsHow it solves it:
/projects/...route with no version segment gets/v1prepended before forwardingcachedContentroutes get/v1beta1, matching the routes that omit the projectUser Flow
Before: a developer running Claude Code in Vertex mode against LiteLLM, with the base URL shape the Claude Code gateway docs give, gets "model not available" on every prompt
CLAUDE_CODE_USE_VERTEX=1,CLAUDE_CODE_SKIP_VERTEX_AUTH=1,ANTHROPIC_VERTEX_BASE_URL=https://litellm-domain/vertex_ai,ANTHROPIC_VERTEX_PROJECT_ID=<project>,CLOUD_ML_REGION=global,ANTHROPIC_AUTH_TOKEN=<virtual key>and startclaude/vertex_ai/v1/projects/...returns 200, so the only way out is adding/v1to the base URL by hand, against the docs exampleAfter: the same setup answers the prompt
CLAUDE_CODE_USE_VERTEX=1,CLAUDE_CODE_SKIP_VERTEX_AUTH=1,ANTHROPIC_VERTEX_BASE_URL=https://litellm-domain/vertex_ai,ANTHROPIC_VERTEX_PROJECT_ID=<project>,CLOUD_ML_REGION=global,ANTHROPIC_AUTH_TOKEN=<virtual key>and startclaudemessage_start, thentext_deltachunks){"input_tokens": N}/v1keeps working exactly as beforeRelevant issues
Linear ticket
Resolves LIT-6873
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Shared setup: a two-worker proxy with no database, real Vertex AI,
claude-sonnet-4-6in thegloballocationControl, straight to Vertex with
gcloud auth print-access-token, showing Vertex itself needs the version segment:Claude Code setup for case 3 (stock 2.1.251 TUI under tmux, base URL in the shape the Claude Code gateway docs give, no
/v1):Before (9212208)
Case 1: streamRawPredict through the documented base URL shape (no /v1)
Command
Observed output (the proxy log shows it forwarded to
https://aiplatform.googleapis.com/projects/<project>/locations/global/publishers/anthropic/models/claude-sonnet-4-6:streamRawPredict, no version)Same command with the
/vertex_ai/projects/...prefixCase 2: count-tokens through the documented base URL shape (no /v1)
Command
Observed output
Case 3: Claude Code 2.1.251 in Vertex mode pointed at the documented base URL shape
Type
Say hi in exactly three wordsin the TUIObserved on screen
Claude Code debug log
Case 4: Gemini generateContent through the documented base URL shape (no /v1)
Command
Observed output
Case 5: cachedContents list through the documented base URL shape (no /v1beta1)
Command
Observed output (Google's own HTML 404 page for
/projects/<project>/locations/global/cachedContents)Case 6: discovery dataStores list through the documented base URL shape (no /v1)
Command
Observed output (Google's own HTML 404 page, the route never reached the Discovery Engine API)
After (60b725c)
Case 1: streamRawPredict through the documented base URL shape (no /v1)
Command
Observed output (the proxy log now shows
https://aiplatform.googleapis.com/v1/projects/<project>/locations/global/publishers/anthropic/models/claude-sonnet-4-6:streamRawPredict)Same command with the
/vertex_ai/projects/...prefixCase 2: count-tokens through the documented base URL shape (no /v1)
Command
Observed output
Case 3: Claude Code 2.1.251 in Vertex mode pointed at the documented base URL shape
Type
Say hi in exactly three wordsin the TUIObserved on screen
Claude Code debug log (no API errors)
Case 4: Gemini generateContent through the documented base URL shape (no /v1)
Command
Observed output
Case 5: cachedContents list through the documented base URL shape (no /v1beta1)
Command
Observed output (forwarded to
/v1beta1/projects/..., an empty list on this project)Case 6: discovery dataStores list through the documented base URL shape (no /v1)
Command
Observed output (the route now reaches the Discovery Engine API; this service account lacks that permission, and the
/v1/and/v1alpha/forms of the same route answer the identical 403 on both sides)Routes that already carry
/v1/returned 200 on both sides ($BASEwith/vertex-ai/v1/projects/...), so callers who worked around the bug by adding/v1keep workingType
🐛 Bug Fix
Caveats (if any)
Low
projectsget a version; anything else is forwarded as before/projects/.../locations/...segment still 500 withvertex_location is requiredv1, orv1beta1when the route names cached content/v1beta1/prefix, which is forwarded verbatimlocal_testing_part1andlocal_testing_part2are red on the threebedrock/cohere.command-r-plus-v1:0testscode-qualityGitHub Actions job is red on staging itself (get_configured_modein router.py has no test)Final Attestation
The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR
60b725c passes /live-pr-risk