Skip to content

fix(models): drop vertex flex tier on gemini-3.6 - #3170

Merged
steebchen merged 1 commit into
mainfrom
fix/gemini-3-6-flash-vertex-flex
Jul 21, 2026
Merged

steebchen merged 1 commit into
mainfrom
fix/gemini-3-6-flash-vertex-flex

Conversation

@smakosh

@smakosh smakosh commented Jul 21, 2026 •

Copy link
Copy Markdown
Member

Problem

Flex was declared on both provider mappings for the new gemini-3.6-flash (#3164), but Google Vertex AI rejects it upstream:

400 INVALID_ARGUMENT: Flex API is not supported for model: gemini-3.6-flash.

Since routing deprioritizes AI Studio (priority 0.8), an unpinned gemini-3.6-flash + service_tier: "flex" request landed on Vertex first and hard-failed with that 400 (fatal with x-no-fallback, wasted roundtrip otherwise).

Fix

Remove "flex" from the google-vertex mapping of gemini-3.6-flash (capability-flag correction, same pattern as gemini-2.5-pro's priority-only Vertex mapping). Priority stays: verified served as ON_DEMAND_PRIORITY on Vertex.

Also adds .claude/** to the vitest exclude list — stale .claude/worktrees/ copies were being collected and failing pnpm test:unit with import errors, same class of noise as the existing .conductor/** exclude.

Verified live (through the local gateway, real upstreams)

Request Before After
gemini-3.6-flash + flex (unpinned, no-fallback) 400 from Vertex routes to AI Studio, used_service_tier: flex, billed 0.5x
google-ai-studio/gemini-3.6-flash + flex ✓ served flex ✓ served flex (also streaming)
google-vertex/gemini-3.6-flash + flex proxied Vertex 400 clean unsupported_service_tier validation error
google-ai-studio/gemini-3.5-flash-lite + flex ✓ served flex ✓ unchanged
google-vertex/gemini-3.5-flash-lite + flex ✓ ON_DEMAND_FLEX ✓ unchanged
google-vertex/{3.6-flash,3.5-flash-lite} + priority ✓ ON_DEMAND_PRIORITY ✓ unchanged

pnpm test:unit green (172 files / 2866 tests), full pnpm build green.

https://claude.ai/code/session_01WNfVNjJoNbVvrzurgZ5J78

Summary by CodeRabbit

  • Bug Fixes

    • Updated Google Vertex configuration to use the supported Priority service tier for the Gemini 3.6 Flash model.
  • Tests

    • Excluded Claude configuration files from automated test discovery.

Google Vertex AI rejects Flex requests for gemini-3.6-flash with
"Flex API is not supported for model: gemini-3.6-flash", so an
unpinned flex request that routed to Vertex hard-failed with a 400.
Remove "flex" from the google-vertex mapping (priority verified
working, ON_DEMAND_PRIORITY) so flex requests route to Google AI
Studio, where the tier is served and billed correctly. Verified
live: AI Studio serves flex for both gemini-3.6-flash and
gemini-3.5-flash-lite; Vertex serves flex for 3.5-flash-lite.

Also exclude .claude/** from vitest so stale worktree copies no
longer fail pnpm test:unit collection (same as .conductor/**).

Claude-Session: https://claude.ai/code/session_01WNfVNjJoNbVvrzurgZ5J78
@coderabbitai

coderabbitai Bot commented Jul 21, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro

Run ID: ac9af6f1-4377-4c86-8c35-1f36f2ff352f

📥 Commits

Reviewing files that changed from the base of the PR and between 57ad9fc and eb75d41.

📒 Files selected for processing (2)
  • packages/models/src/models/google.ts
  • vitest.config.mts

Walkthrough

The Vertex configuration for gemini-3.6-flash now supports only the priority service tier, and Vitest excludes .claude/** from test discovery.

Changes

Vertex service tier update

Layer / File(s) Summary
Gemini Vertex service tier contract
packages/models/src/models/google.ts
The gemini-3.6-flash Vertex provider removes Flex support and documents Vertex’s rejection of that tier.

Vitest discovery update

Layer / File(s) Summary
Test exclusion configuration
vitest.config.mts
Vitest excludes the .claude/** path from test discovery.

Estimated code review effort: 1 (Trivial) | ~3 minutes

Possibly related PRs

Suggested reviewers: steebchen

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly summarizes the main change: removing Vertex Flex support for the Gemini 3.6 model mapping.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/gemini-3-6-flash-vertex-flex

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@steebchen
steebchen merged commit 68d46a2 into main Jul 21, 2026
22 of 23 checks passed
@steebchen
steebchen deleted the fix/gemini-3-6-flash-vertex-flex branch July 21, 2026 21:08
SeyedHashtag pushed a commit to SeyedHashtag/llmgateway that referenced this pull request Jul 22, 2026
## Summary

Reverts the service-tier change from theopenco#3170:
`google-vertex/gemini-3.6-flash` supports Flex PayGo again
(`serviceTiers: ["flex", "priority"]`).

Google's Flex PayGo documentation lists `gemini-3.6-flash` as supported
on the global Vertex endpoint, and live probes confirm it. The "Flex API
is not supported for model: gemini-3.6-flash" rejection that motivated
theopenco#3170 no longer reproduces — most likely a day-one rollout gap, since
the model launched the same day the tier was dropped.

## Verification

- **Direct Vertex probe** (global endpoint,
`X-Vertex-AI-LLM-Request-Type: shared` +
`X-Vertex-AI-LLM-Shared-Request-Type: flex`): HTTP 200 with
`usageMetadata.trafficType: ON_DEMAND_FLEX`. Priority still returns
`ON_DEMAND_PRIORITY`.
- **Through the gateway** (local build, `service_tier: "flex"`, pinned
with `x-no-fallback`): non-streaming and streaming both return
`metadata.requested_service_tier: "flex"` and `used_service_tier:
"flex"`, billed at the 0.5× Flex multiplier (e.g. input 12 tokens →
$9e-06 = 12 × 1.5e-6 × 0.5).
- **E2E**: `TEST_MODELS="google-vertex/gemini-3.6-flash" pnpm test:e2e`
— 26 files passed, 91 tests passed, 0 failures.
- `pnpm format` and full `pnpm build` pass.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added Flex service tier support for the Gemini 3.6 Flash Google Vertex
provider.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
steebchen added a commit that referenced this pull request Aug 17, 2026
## Problem

`google-vertex/gemini-3.7-flash` rejects `service_tier: "flex"` with
`unsupported_service_tier`, even though every other Gemini Vertex
mapping
in the catalogue offers Flex and Google's Flex PayGo docs list the model
on the global endpoint.

The mapping was added on the model's launch day (#3592), where a live
probe returned `400 "Flex API is not supported for model"`, so only
`priority` was declared. That is the same day-one rollout gap that
produced #3170 → #3176 for `gemini-3.6-flash`.

## Fix

Declare `serviceTiers: ["flex", "priority"]` on the `google-vertex`
mapping and drop the now-stale comment. Catalogue-only change; the tier
headers, pricing multipliers and e2e cases are all derived from it.

## Verification

- **Direct Vertex probe** (global endpoint, OAuth,
  `X-Vertex-AI-LLM-Request-Type: shared` +
  `X-Vertex-AI-LLM-Shared-Request-Type: flex`): HTTP 200 with
  `usageMetadata.trafficType: ON_DEMAND_FLEX`, three consecutive runs.
  Priority still returns `ON_DEMAND_PRIORITY`.
- **Through the gateway** (local build, `x-no-fallback`): `flex` →
  `requested_service_tier: flex`, `used_service_tier: flex`, input cost
  `9 × 0.75e-6 × 0.5 = 3.375e-6` (Flex 0.5× multiplier). `priority` →
  `9 × 0.75e-6 × 1.8 = 1.215e-5`.
- **E2E**: `TEST_MODELS="google-vertex/gemini-3.7-flash" FULL_MODE=true
  pnpm test:e2e` — the two new Flex cases (non-streaming + streaming,
  including the multiplier-scaled cost assertion) pass alongside the
  existing Priority ones. The only failures in the scoped run are in
  `keys-provider.e2e.ts` for unrelated providers whose keys are missing
  from this machine's local env.
- **Unit**: 299/301 files pass. The two failures
  (`onboarding-sponsorship.spec.ts`, and a `logs.spec.ts` case that
  hardcodes the default gateway port) reproduce on the unmodified
  baseline and on the isolated-stack ports respectively.
- `pnpm format` and full `pnpm build` pass.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants