Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 8 additions & 3 deletions packages/actions/src/get-provider-endpoint.ts
Original file line number Diff line number Diff line change
Expand Up @@ -201,14 +201,17 @@ export function getProviderEndpoint(
"https://aiplatform.googleapis.com",
) ?? "https://aiplatform.googleapis.com";
break;
case "vertex-openai":
case "vertex-openai": {
const vertexOpenaiDefaultHost =
regionBaseUrl ?? "https://aiplatform.googleapis.com";
url =
envValueOrDefault(
"vertex-openai",
"baseUrl",
"https://aiplatform.googleapis.com",
) ?? "https://aiplatform.googleapis.com";
vertexOpenaiDefaultHost,
) ?? vertexOpenaiDefaultHost;
break;
}
case "vertex-anthropic": {
const vaDefaultRegion =
providerKeyOptions?.vertex_anthropic_region ??
Expand Down Expand Up @@ -451,6 +454,8 @@ export function getProviderEndpoint(
);
}
const vertexRegion =
region ??
providerKeyOptions?.vertex_openai_region ??

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Honor the routed Vertex region before key defaults

When a BYOK/hybrid vertex-openai key is configured with vertex_openai_region: "us-west2" to use the new DeepSeek V3.1 mapping, chat.ts intentionally leaves usedRegion undefined for regionless/global mappings such as the existing Grok/Qwen/Kimi entries, but this fallback still applies the key option inside endpoint construction. Those requests are then treated as regionless for routing/logging/validation while the URL path is built as /locations/us-west2 (with the host already chosen from the undefined region), so global-only Vertex OpenAI mappings can be sent to the wrong regional endpoint. Only apply the key option after the selected mapping has been resolved to that compatible region, or promote and validate the region consistently before building the URL.

Useful? React with 👍 / 👎.

getProviderEnvValue("vertex-openai", "region", configIndex, "global") ??
"global";
return `${url}/v1/projects/${projectId}/locations/${vertexRegion}/endpoints/openapi/chat/completions`;
Expand Down
1 change: 1 addition & 0 deletions packages/db/src/schema.ts
Original file line number Diff line number Diff line change
Expand Up @@ -1090,6 +1090,7 @@ export interface ProviderKeyOptions {
alibaba_region?: "singapore" | "us-virginia" | "cn-beijing";
google_vertex_project_id?: string;
vertex_openai_project_id?: string;
vertex_openai_region?: "global";
vertex_anthropic_region?: string;
}

Expand Down
67 changes: 67 additions & 0 deletions packages/models/src/models/alibaba.ts
Original file line number Diff line number Diff line change
Expand Up @@ -511,6 +511,20 @@ export const alibabaModels = [
tools: true,
jsonOutput: true,
},
{
providerId: "vertex-openai",
externalId: "qwen/qwen3-235b-a22b-instruct-2507-maas",
inputPrice: "0.22e-6",
outputPrice: "0.88e-6",
requestPrice: "0",
contextSize: 262144,
maxOutput: 32768,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Lower the Vertex Qwen3 235B max output cap

When callers send max_tokens between 16,385 and 32,768 to vertex-openai/qwen3-235b-a22b-instruct-2507, the gateway will accept this mapping because of this value and forward a request that Vertex rejects; Google's MaaS model card lists qwen3-235b-a22b-instruct-2507-maas with max output 16,384 for both global and us-south1 (https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/maas/qwen/qwen3-235b). Set this cap to the upstream limit so the gateway rejects or reroutes before the provider call.

Useful? React with 👍 / 👎.

reasoning: false,
streaming: true,
vision: false,
tools: true,
jsonOutput: true,
},
],
},
{
Expand Down Expand Up @@ -897,6 +911,20 @@ export const alibabaModels = [
tools: true,
jsonOutput: true,
},
{
providerId: "vertex-openai",
externalId: "qwen/qwen3-coder-480b-a35b-instruct-maas",
inputPrice: "0.22e-6",
cachedInputPrice: "0.022e-6",
outputPrice: "1.8e-6",
requestPrice: "0",
contextSize: 262144,
maxOutput: 65536,
streaming: true,
vision: false,
tools: true,
jsonOutput: true,
},
],
},
{
Expand Down Expand Up @@ -1138,6 +1166,32 @@ export const alibabaModels = [
"tools",
],
},
{
providerId: "vertex-openai",
externalId: "qwen/qwen3-next-80b-a3b-thinking-maas",
inputPrice: "0.15e-6",
outputPrice: "1.2e-6",
requestPrice: "0",
contextSize: 131072,
maxOutput: 32768,
Comment on lines +1175 to +1176

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Raise the Vertex Qwen3 Next limits

For vertex-openai/qwen3-next-80b-a3b-thinking-maas, Vertex documents both context length and max output as 262,144 (https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/maas/qwen/qwen3-next-thinking), but this mapping advertises 131,072 context and 32,768 output. model: auto filters candidates by contextSize, and explicit/provider-specific routes enforce maxOutput, so valid long-context or long-generation requests are routed away or fail with a gateway 400 even though Vertex supports them; the instruct sibling below has the same underreported limits.

Useful? React with 👍 / 👎.

reasoning: true,
reasoningOutput: "omit",
streaming: true,
vision: false,
tools: true,
jsonOutput: false,
supportedParameters: [
"temperature",
"max_tokens",
"top_p",
"frequency_penalty",
"presence_penalty",
"stop",
"stream",
"response_format",
"tools",
],
},
],
},
{
Expand Down Expand Up @@ -1174,6 +1228,19 @@ export const alibabaModels = [
tools: true,
jsonOutput: true,
},
{
providerId: "vertex-openai",
externalId: "qwen/qwen3-next-80b-a3b-instruct-maas",
inputPrice: "0.15e-6",
outputPrice: "1.2e-6",
requestPrice: "0",
contextSize: 131072,
maxOutput: 32768,
streaming: true,
vision: false,
tools: true,
jsonOutput: true,
},
],
},
{
Expand Down
14 changes: 14 additions & 0 deletions packages/models/src/models/deepseek.ts
Original file line number Diff line number Diff line change
Expand Up @@ -206,6 +206,20 @@ export const deepseekModels = [
tools: true,
jsonOutput: true,
},
{
providerId: "vertex-openai",
externalId: "deepseek-ai/deepseek-v3.2-maas",
inputPrice: "0.56e-6",
cachedInputPrice: "0.056e-6",
outputPrice: "1.68e-6",
requestPrice: "0",
contextSize: 163840,
maxOutput: 65536,
streaming: true,
vision: false,
tools: true,
jsonOutput: true,
},
],
},
{
Expand Down
15 changes: 15 additions & 0 deletions packages/models/src/models/moonshot.ts
Original file line number Diff line number Diff line change
Expand Up @@ -163,6 +163,21 @@ export const moonshotModels = [
jsonOutput: true,
deactivatedAt: new Date("2026-07-08"),
},
{
providerId: "vertex-openai",
externalId: "moonshotai/kimi-k2-thinking-maas",
inputPrice: "0.6e-6",
cachedInputPrice: "0.06e-6",
outputPrice: "2.5e-6",
requestPrice: "0",
contextSize: 262144,
maxOutput: 32768,
reasoning: true,
streaming: true,
vision: false,
tools: true,
jsonOutput: true,
},
],
},
{
Expand Down
40 changes: 4 additions & 36 deletions packages/models/src/models/xai.ts
Original file line number Diff line number Diff line change
Expand Up @@ -627,24 +627,8 @@ export const xaiModels = [
providerId: "vertex-openai",
stability: "unstable" as const,
externalId: "grok-4.20-reasoning",
inputPrice: "2.0e-6",
outputPrice: "6.0e-6",
pricingTiers: [
{
name: "Up to 200K",
upToTokens: 200000,
inputPrice: "2.0e-6",
outputPrice: "6.0e-6",
cachedInputPrice: "0.2e-6",
},
{
name: "Over 200K",
upToTokens: Infinity,
inputPrice: "4.0e-6",
outputPrice: "12.0e-6",
cachedInputPrice: "0.4e-6",
},
],
inputPrice: "1.25e-6",
outputPrice: "2.5e-6",
cachedInputPrice: "0.2e-6",
requestPrice: "0",
contextSize: 2_000_000,
Expand All @@ -670,24 +654,8 @@ export const xaiModels = [
providerId: "vertex-openai",
stability: "unstable" as const,
externalId: "grok-4.20-non-reasoning",
inputPrice: "2.0e-6",
outputPrice: "6.0e-6",
pricingTiers: [
{
name: "Up to 200K",
upToTokens: 200000,
inputPrice: "2.0e-6",
outputPrice: "6.0e-6",
cachedInputPrice: "0.2e-6",
},
{
name: "Over 200K",
upToTokens: Infinity,
inputPrice: "4.0e-6",
outputPrice: "12.0e-6",
cachedInputPrice: "0.4e-6",
},
],
inputPrice: "1.25e-6",
outputPrice: "2.5e-6",
cachedInputPrice: "0.2e-6",
requestPrice: "0",
contextSize: 2_000_000,
Expand Down
29 changes: 29 additions & 0 deletions packages/models/src/models/zai.ts
Original file line number Diff line number Diff line change
Expand Up @@ -316,6 +316,21 @@ export const zaiModels = [
tools: true,
jsonOutput: true,
},
{
providerId: "vertex-openai",
externalId: "zai-org/glm-5-maas",
inputPrice: "1e-6",
cachedInputPrice: "0.1e-6",
outputPrice: "3.2e-6",
requestPrice: "0",
contextSize: 202752,
maxOutput: 32768,
streaming: true,
reasoning: true,
vision: false,
tools: true,
jsonOutput: true,
},
],
},
{
Expand Down Expand Up @@ -712,6 +727,20 @@ export const zaiModels = [
"reasoning_effort",
],
},
{
providerId: "vertex-openai",
externalId: "zai-org/glm-4.7-maas",
inputPrice: "0.6e-6",
outputPrice: "2.2e-6",
requestPrice: "0",
contextSize: 202752,
maxOutput: 128000,
streaming: true,
reasoning: true,
vision: false,
tools: true,
jsonOutput: true,
},
],
},
{
Expand Down
8 changes: 8 additions & 0 deletions packages/models/src/providers.ts
Original file line number Diff line number Diff line change
Expand Up @@ -417,6 +417,14 @@ export const providers: ProviderDefinition[] = [
statusPageUrl: "https://status.cloud.google.com",
announcement: null,
priority: 0.2,
regionConfig: {
Comment thread
steebchen marked this conversation as resolved.
optionsKey: "vertex_openai_region",

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Accept vertex_openai_region in key options

This change adds regionConfig.optionsKey = "vertex_openai_region" for vertex-openai, so clients now submit that option when a region is selected, but the provider-key create validator still whitelists only vertex_openai_project_id in apps/api/src/routes/keys-provider.ts and does not include vertex_openai_region. Because Zod object parsing strips unknown keys by default, the selected region is dropped on save, so BYOK/hybrid keys silently fall back to global and the new regional routing cannot be configured through the API.

Useful? React with 👍 / 👎.

defaultRegion: "global",
Comment on lines +420 to +422

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Keep regionless Vertex mappings selectable

Adding regionConfig makes chat.ts fill usedRegion with the provider default after endpoint construction (apps/gateway/src/chat/chat.ts:5187-5195). The existing Vertex OpenAI Grok 4.20 mappings in packages/models/src/models/xai.ts are still regionless but define supportedParameters; once usedRegion becomes "global", later exact lookups like p.region === usedRegion (apps/gateway/src/chat/chat.ts:5882-5885) no longer find those mappings, so unsupported parameters such as frequency_penalty/presence_penalty stop being stripped and regress to upstream 400s for vertex-openai/grok-4.20-* requests. Mark the existing mappings as region: "global" or use the fallback mapping selector for these checks.

Useful? React with 👍 / 👎.

regions: [{ id: "global", label: "Global (default)" }],
endpointMap: {
global: "https://aiplatform.googleapis.com",
},
},
termsUrl: "https://cloud.google.com/terms/service-terms",
privacyPolicyUrl: "https://cloud.google.com/terms/data-processing-addendum",
headquarters: "US",
Expand Down
Loading