Repository navigation
feat(vertex-openai): add 13 partner-model mappings - #2343
Conversation
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository UI Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (10)
🚧 Files skipped from review as they are similar to previous changes (10)
WalkthroughAdds region-aware Vertex OpenAI support: new ChangesVertex OpenAI Provider Support
Sequence Diagram(s)sequenceDiagram
participant Client
participant getProviderEndpoint
participant regionConfig
participant ProviderKeyOptions
participant VertexAPI
Client->>getProviderEndpoint: call(region?, providerKeyOptions?)
getProviderEndpoint->>regionConfig: lookup endpointMap[region]
regionConfig-->>getProviderEndpoint: regionBaseUrl (if present)
alt regionBaseUrl present
getProviderEndpoint->>getProviderEndpoint: derive vertexOpenaiDefaultHost from regionBaseUrl
else
getProviderEndpoint->>getProviderEndpoint: use https://aiplatform.googleapis.com as defaultHost
end
getProviderEndpoint->>ProviderKeyOptions: read vertex_openai_region from providerKeyOptions
getProviderEndpoint-->>Client: return endpoint with resolved baseUrl and vertexRegion
Client->>VertexAPI: request to resolved baseUrl/region
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~25 minutes Possibly related PRs
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Pull request overview
Adds Vertex AI OpenAI-compatible partner model coverage and region support for additional MaaS endpoints.
Changes:
- Adds new
vertex-openaimappings for MiniMax, Moonshot/Kimi, ZAI/GLM, OpenAI GPT OSS, and Qwen models. - Expands
vertex-openairegional configuration and provider key option typing. - Updates Grok 4.20 Vertex pricing from tiered pricing to current flat Vertex rates.
Reviewed changes
Copilot reviewed 8 out of 8 changed files in this pull request and generated 1 comment.
Show a summary per file
| File | Description |
|---|---|
packages/models/src/providers.ts |
Adds us-central1 and us-south1 to Vertex OpenAI region config. |
packages/models/src/models/zai.ts |
Adds Vertex OpenAI mappings for GLM-5 and GLM-4.7. |
packages/models/src/models/xai.ts |
Updates Grok 4.20 Vertex pricing. |
packages/models/src/models/openai.ts |
Adds Vertex OpenAI mappings for GPT OSS 120B and 20B. |
packages/models/src/models/moonshot.ts |
Adds Vertex OpenAI mapping for Kimi K2 Thinking. |
packages/models/src/models/minimax.ts |
Adds skipped-test Vertex OpenAI mapping for MiniMax M2. |
packages/models/src/models/alibaba.ts |
Adds Vertex OpenAI mappings for Qwen3 models. |
packages/db/src/schema.ts |
Widens vertex_openai_region provider key option type. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| { | ||
| providerId: "vertex-openai", | ||
| modelName: "openai/gpt-oss-20b-maas", | ||
| region: "us-central1", |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 0001b4ae19
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| providerId: "vertex-openai", | ||
| modelName: "openai/gpt-oss-120b-maas", | ||
| region: "us-central1", |
There was a problem hiding this comment.
Avoid requiring per-region keys for new Vertex mappings
These new vertex-openai mappings are pinned to non-default regions (for example us-central1), but in credits mode the gateway drops non-default region mappings unless a region-suffixed API key env var exists (hasRegionSpecificEnvKey in apps/gateway/src/chat/chat.ts only accepts base key for the provider default region). With a common setup that only provides LLM_VERTEX_OPENAI_SERVICE_ACCOUNT_JSON, these models are filtered out and never routed, so the added mappings are effectively unavailable unless operators add ...__US_CENTRAL1/...__US_SOUTH1 keys.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: fd4c4f4dc0
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| providerId: "vertex-openai", | ||
| modelName: "openai/gpt-oss-20b-maas", | ||
| region: "us-central1", | ||
| inputPrice: "0.07e-6", | ||
| cachedInputPrice: "0.007e-6", |
There was a problem hiding this comment.
Avoid selecting a non-global Vertex model for key validation
Adding openai/gpt-oss-20b-maas as a vertex-openai mapping in us-central1 makes it the cheapest candidate, so validateProviderKey() now tends to pick it as the validation model; that selector only filters by providerMapping.regions (not singular region) and then builds the endpoint with the provider default region (global). In practice, validating a Vertex OpenAI key can fail with a region/model mismatch even when the credentials are valid, because the request uses a global endpoint for a us-central1-only model.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
packages/actions/src/get-provider-endpoint.ts (1)
96-105:⚠️ Potential issue | 🟠 Major | ⚡ Quick winUnify Vertex region resolution for both host and location path.
vertex-openaicurrently derives host fromregion(function arg) but derives location fromregion ?? providerKeyOptions?.vertex_openai_region ?? env. Whenregionis unset and only provider options/env define region, requests still default to global host, which breaks consistent regional routing.Suggested fix
- let regionBaseUrl: string | undefined; - if (region) { - const providerDef = providers.find((p) => p.id === provider) as - | ProviderDefinition - | undefined; - const endpointMap = providerDef?.regionConfig?.endpointMap as - | Record<string, string> - | undefined; - regionBaseUrl = endpointMap?.[region]; - } + let regionBaseUrl: string | undefined; + let resolvedRegionForRouting = region; + if (provider === "vertex-openai" && !resolvedRegionForRouting) { + resolvedRegionForRouting = + providerKeyOptions?.vertex_openai_region ?? + getProviderEnvValue("vertex-openai", "region", configIndex, "global") ?? + "global"; + } + if (resolvedRegionForRouting) { + const providerDef = providers.find((p) => p.id === provider) as + | ProviderDefinition + | undefined; + const endpointMap = providerDef?.regionConfig?.endpointMap as + | Record<string, string> + | undefined; + regionBaseUrl = endpointMap?.[resolvedRegionForRouting]; + } ... - const vertexRegion = - region ?? - providerKeyOptions?.vertex_openai_region ?? - getProviderEnvValue("vertex-openai", "region", configIndex, "global") ?? - "global"; + const vertexRegion = + resolvedRegionForRouting ?? + providerKeyOptions?.vertex_openai_region ?? + getProviderEnvValue("vertex-openai", "region", configIndex, "global") ?? + "global";Also applies to: 151-161, 379-383
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@packages/actions/src/get-provider-endpoint.ts` around lines 96 - 105, Compute a single resolvedRegion variable (e.g., const resolvedRegion = region ?? providerKeyOptions?.vertex_openai_region ?? env) and use that instead of the raw function arg `region` when looking up the provider host (the `regionBaseUrl` code that uses `providerDef`, `endpointMap`, and `providers`) and anywhere else the code derives a Vertex location/path; replace occurrences where host uses `region` but location uses `region ?? providerKeyOptions?.vertex_openai_region ?? env` so both host and location use `resolvedRegion` (apply the same change to the other similar blocks around the `regionBaseUrl` logic).
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Outside diff comments:
In `@packages/actions/src/get-provider-endpoint.ts`:
- Around line 96-105: Compute a single resolvedRegion variable (e.g., const
resolvedRegion = region ?? providerKeyOptions?.vertex_openai_region ?? env) and
use that instead of the raw function arg `region` when looking up the provider
host (the `regionBaseUrl` code that uses `providerDef`, `endpointMap`, and
`providers`) and anywhere else the code derives a Vertex location/path; replace
occurrences where host uses `region` but location uses `region ??
providerKeyOptions?.vertex_openai_region ?? env` so both host and location use
`resolvedRegion` (apply the same change to the other similar blocks around the
`regionBaseUrl` logic).
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository UI
Review profile: CHILL
Plan: Pro
Run ID: dd098a55-00ab-4efa-92f4-f8995ccbb9ba
📒 Files selected for processing (10)
packages/actions/src/get-provider-endpoint.tspackages/db/src/schema.tspackages/models/src/models/alibaba.tspackages/models/src/models/deepseek.tspackages/models/src/models/minimax.tspackages/models/src/models/moonshot.tspackages/models/src/models/openai.tspackages/models/src/models/xai.tspackages/models/src/models/zai.tspackages/models/src/providers.ts
0fac4a2 to
68a912f
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 68a912f264
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| announcement: null, | ||
| priority: 0.9, | ||
| regionConfig: { | ||
| optionsKey: "vertex_openai_region", |
There was a problem hiding this comment.
Accept vertex_openai_region in key options
This change adds regionConfig.optionsKey = "vertex_openai_region" for vertex-openai, so clients now submit that option when a region is selected, but the provider-key create validator still whitelists only vertex_openai_project_id in apps/api/src/routes/keys-provider.ts and does not include vertex_openai_region. Because Zod object parsing strips unknown keys by default, the selected region is dropped on save, so BYOK/hybrid keys silently fall back to global and the new regional routing cannot be configured through the API.
Useful? React with 👍 / 👎.
68a912f to
64bee15
Compare
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
Actionable comments posted: 0 |
64bee15 to
6dd740c
Compare
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
Actionable comments posted: 0 |
Adds DeepSeek-V3.1 (us-west2) and DeepSeek-V3.2 (global) on Vertex AI MaaS via the OpenAI-compatible chat completions endpoint. Introduces vertex-openai regionConfig so per-model regions select the correct regional host. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Adds Vertex AI MaaS mappings via the OpenAI-compatible endpoint for: MiniMax M2, Kimi K2 Thinking, GLM 4.7, GLM 5, gpt-oss 120b/20b, Qwen3 235B/Coder 480B/Next 80B Thinking/Instruct. Also refreshes Grok 4.20 pricing to the new flat Vertex rates and expands the vertex-openai regionConfig with us-central1 and us-south1. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Vertex's gpt-oss MaaS endpoints support function calling but reject
forced tool_choice. Mirror the qwen-thinking pattern: flip tools back
on and drop tool_choice from supportedParameters. jsonOutput stays
disabled until the {"final": "..."} envelope unwrap lands.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Verified via e2e and direct curls that Vertex's gpt-oss-120b and
gpt-oss-20b MaaS endpoints return clean JSON with response_format:
json_object — no {"final": ...} envelope on the Vertex deployment.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
6dd740c to
a37b662
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: a37b66273f
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| regionConfig: { | ||
| optionsKey: "vertex_openai_region", | ||
| defaultRegion: "global", |
There was a problem hiding this comment.
Keep regionless Vertex mappings selectable
Adding regionConfig makes chat.ts fill usedRegion with the provider default after endpoint construction (apps/gateway/src/chat/chat.ts:5187-5195). The existing Vertex OpenAI Grok 4.20 mappings in packages/models/src/models/xai.ts are still regionless but define supportedParameters; once usedRegion becomes "global", later exact lookups like p.region === usedRegion (apps/gateway/src/chat/chat.ts:5882-5885) no longer find those mappings, so unsupported parameters such as frequency_penalty/presence_penalty stop being stripped and regress to upstream 400s for vertex-openai/grok-4.20-* requests. Mark the existing mappings as region: "global" or use the fallback mapping selector for these checks.
Useful? React with 👍 / 👎.
These three Vertex MaaS mappings have genuine output defects: gpt-oss
120b/20b intermittently leak the harmony-channel {"final": ...}
envelope into JSON output, and minimax-m2 emits <think> prose instead
of JSON with unreliable tool calls. Drop them; the remaining 11
partner-model mappings return clean chat/tools/JSON.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 10ae27d027
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Only deepseek-v3.1 needs an explicit region (404s on the global endpoint, served from us-west2). Drop the redundant region:"global" from the other mappings so they default to global like grok, route qwen3-coder via global (now available there), and trim the unused us-central1/us-south1 entries from the region config. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: e080c8b7a4
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| } | ||
| const vertexRegion = | ||
| region ?? | ||
| providerKeyOptions?.vertex_openai_region ?? |
There was a problem hiding this comment.
Honor the routed Vertex region before key defaults
When a BYOK/hybrid vertex-openai key is configured with vertex_openai_region: "us-west2" to use the new DeepSeek V3.1 mapping, chat.ts intentionally leaves usedRegion undefined for regionless/global mappings such as the existing Grok/Qwen/Kimi entries, but this fallback still applies the key option inside endpoint construction. Those requests are then treated as regionless for routing/logging/validation while the URL path is built as /locations/us-west2 (with the host already chosen from the undefined region), so global-only Vertex OpenAI mappings can be sent to the wrong regional endpoint. Only apply the key option after the selected mapping has been resolved to that compatible region, or promote and validate the region consistently before building the URL.
Useful? React with 👍 / 👎.
Removes the deepseek-v3.1 vertex-openai MaaS mapping (kept the bytedance mapping). It was the only model needing a non-global region, so vertex-openai is now global-only: trim us-west2 from the region config and schema union. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 4804132fe7
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| outputPrice: "0.88e-6", | ||
| requestPrice: "0", | ||
| contextSize: 262144, | ||
| maxOutput: 32768, |
There was a problem hiding this comment.
Lower the Vertex Qwen3 235B max output cap
When callers send max_tokens between 16,385 and 32,768 to vertex-openai/qwen3-235b-a22b-instruct-2507, the gateway will accept this mapping because of this value and forward a request that Vertex rejects; Google's MaaS model card lists qwen3-235b-a22b-instruct-2507-maas with max output 16,384 for both global and us-south1 (https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/maas/qwen/qwen3-235b). Set this cap to the upstream limit so the gateway rejects or reroutes before the provider call.
Useful? React with 👍 / 👎.
| contextSize: 131072, | ||
| maxOutput: 32768, |
There was a problem hiding this comment.
Raise the Vertex Qwen3 Next limits
For vertex-openai/qwen3-next-80b-a3b-thinking-maas, Vertex documents both context length and max output as 262,144 (https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/maas/qwen/qwen3-next-thinking), but this mapping advertises 131,072 context and 32,768 output. model: auto filters candidates by contextSize, and explicit/provider-specific routes enforce maxOutput, so valid long-context or long-generation requests are routed away or fail with a gateway 400 even though Vertex supports them; the instruct sibling below has the same underreported limits.
Useful? React with 👍 / 👎.
Summary
Adds 13 Vertex AI MaaS partner-model mappings via the OpenAI-compatible
vertex-openaiendpoint and expands itsregionConfigwith the regions those partner models require.This PR includes the DeepSeek V3.1/V3.2 mappings originally proposed in #2341 plus 11 additional partner models, refreshes Grok 4.20 pricing to current Vertex rates, and wires per-mapping region into the endpoint construction. #2341 is now redundant and can be closed.
New mappings:
deepseek-ai/deepseek-v3.1-maasdeepseek-ai/deepseek-v3.2-maasminimaxai/minimax-m2-maasmoonshotai/kimi-k2-thinking-maaszai-org/glm-4.7-maaszai-org/glm-5-maasopenai/gpt-oss-120b-maasopenai/gpt-oss-20b-maasqwen/qwen3-235b-a22b-instruct-2507-maasqwen/qwen3-coder-480b-a35b-instruct-maasqwen/qwen3-next-80b-a3b-thinking-maasqwen/qwen3-next-80b-a3b-instruct-maasgrok-4.20-reasoning/grok-4.20-non-reasoningInfrastructure changes:
vertex-openaiprovider gains aregionConfigwithglobal,us-central1,us-south1,us-west2and matchingendpointMapentries.getProviderEndpointreads the per-mappingregionand uses the matching regional host + path location forvertex-openai.ProviderKeyOptions.vertex_openai_regionwidened to"global" | "us-central1" | "us-south1" | "us-west2".Caveats discovered during testing:
gpt-oss-120b-maasandgpt-oss-20b-maason Vertex don't reliably emit OpenAI-styletool_callsand wrap JSON output in a{"final": …}envelope, so those mappings aretools: false/jsonOutput: false. Other providers for the same model keep tools/JSON on.minimax-m2-maason Vertex emits reasoning-only chunks for short prompts (nocontentdeltas), so its mapping istest: "skip"to avoid flaky CI. Model still works for production traffic.globalendpoint shows transient HTML 404 throttle responses under bursty traffic; this is a Vertex-side rate limit, not a gateway issue.Test plan
pnpm formatpnpm build🤖 Generated with Claude Code
Summary by CodeRabbit
New Features
Updates