Repository navigation
fix(models): drop vertex flex tier on gemini-3.6 - #3170
Merged
Merged
Conversation
Google Vertex AI rejects Flex requests for gemini-3.6-flash with "Flex API is not supported for model: gemini-3.6-flash", so an unpinned flex request that routed to Vertex hard-failed with a 400. Remove "flex" from the google-vertex mapping (priority verified working, ON_DEMAND_PRIORITY) so flex requests route to Google AI Studio, where the tier is served and billed correctly. Verified live: AI Studio serves flex for both gemini-3.6-flash and gemini-3.5-flash-lite; Vertex serves flex for 3.5-flash-lite. Also exclude .claude/** from vitest so stale worktree copies no longer fail pnpm test:unit collection (same as .conductor/**). Claude-Session: https://claude.ai/code/session_01WNfVNjJoNbVvrzurgZ5J78
Contributor
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository UI Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (2)
WalkthroughThe Vertex configuration for ChangesVertex service tier update
Vitest discovery update
Estimated code review effort: 1 (Trivial) | ~3 minutes Possibly related PRs
Suggested reviewers: 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
SeyedHashtag
pushed a commit
to SeyedHashtag/llmgateway
that referenced
this pull request
Jul 22, 2026
## Summary Reverts the service-tier change from theopenco#3170: `google-vertex/gemini-3.6-flash` supports Flex PayGo again (`serviceTiers: ["flex", "priority"]`). Google's Flex PayGo documentation lists `gemini-3.6-flash` as supported on the global Vertex endpoint, and live probes confirm it. The "Flex API is not supported for model: gemini-3.6-flash" rejection that motivated theopenco#3170 no longer reproduces — most likely a day-one rollout gap, since the model launched the same day the tier was dropped. ## Verification - **Direct Vertex probe** (global endpoint, `X-Vertex-AI-LLM-Request-Type: shared` + `X-Vertex-AI-LLM-Shared-Request-Type: flex`): HTTP 200 with `usageMetadata.trafficType: ON_DEMAND_FLEX`. Priority still returns `ON_DEMAND_PRIORITY`. - **Through the gateway** (local build, `service_tier: "flex"`, pinned with `x-no-fallback`): non-streaming and streaming both return `metadata.requested_service_tier: "flex"` and `used_service_tier: "flex"`, billed at the 0.5× Flex multiplier (e.g. input 12 tokens → $9e-06 = 12 × 1.5e-6 × 0.5). - **E2E**: `TEST_MODELS="google-vertex/gemini-3.6-flash" pnpm test:e2e` — 26 files passed, 91 tests passed, 0 failures. - `pnpm format` and full `pnpm build` pass. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added Flex service tier support for the Gemini 3.6 Flash Google Vertex provider. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This was referenced Aug 13, 2026
steebchen
added a commit
that referenced
this pull request
Aug 17, 2026
## Problem `google-vertex/gemini-3.7-flash` rejects `service_tier: "flex"` with `unsupported_service_tier`, even though every other Gemini Vertex mapping in the catalogue offers Flex and Google's Flex PayGo docs list the model on the global endpoint. The mapping was added on the model's launch day (#3592), where a live probe returned `400 "Flex API is not supported for model"`, so only `priority` was declared. That is the same day-one rollout gap that produced #3170 → #3176 for `gemini-3.6-flash`. ## Fix Declare `serviceTiers: ["flex", "priority"]` on the `google-vertex` mapping and drop the now-stale comment. Catalogue-only change; the tier headers, pricing multipliers and e2e cases are all derived from it. ## Verification - **Direct Vertex probe** (global endpoint, OAuth, `X-Vertex-AI-LLM-Request-Type: shared` + `X-Vertex-AI-LLM-Shared-Request-Type: flex`): HTTP 200 with `usageMetadata.trafficType: ON_DEMAND_FLEX`, three consecutive runs. Priority still returns `ON_DEMAND_PRIORITY`. - **Through the gateway** (local build, `x-no-fallback`): `flex` → `requested_service_tier: flex`, `used_service_tier: flex`, input cost `9 × 0.75e-6 × 0.5 = 3.375e-6` (Flex 0.5× multiplier). `priority` → `9 × 0.75e-6 × 1.8 = 1.215e-5`. - **E2E**: `TEST_MODELS="google-vertex/gemini-3.7-flash" FULL_MODE=true pnpm test:e2e` — the two new Flex cases (non-streaming + streaming, including the multiplier-scaled cost assertion) pass alongside the existing Priority ones. The only failures in the scoped run are in `keys-provider.e2e.ts` for unrelated providers whose keys are missing from this machine's local env. - **Unit**: 299/301 files pass. The two failures (`onboarding-sponsorship.spec.ts`, and a `logs.spec.ts` case that hardcodes the default gateway port) reproduce on the unmodified baseline and on the isolated-stack ports respectively. - `pnpm format` and full `pnpm build` pass. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Flex was declared on both provider mappings for the new
gemini-3.6-flash(#3164), but Google Vertex AI rejects it upstream:Since routing deprioritizes AI Studio (priority 0.8), an unpinned
gemini-3.6-flash+service_tier: "flex"request landed on Vertex first and hard-failed with that 400 (fatal withx-no-fallback, wasted roundtrip otherwise).Fix
Remove
"flex"from thegoogle-vertexmapping ofgemini-3.6-flash(capability-flag correction, same pattern asgemini-2.5-pro's priority-only Vertex mapping). Priority stays: verified served asON_DEMAND_PRIORITYon Vertex.Also adds
.claude/**to the vitest exclude list — stale.claude/worktrees/copies were being collected and failingpnpm test:unitwith import errors, same class of noise as the existing.conductor/**exclude.Verified live (through the local gateway, real upstreams)
gemini-3.6-flash+ flex (unpinned, no-fallback)used_service_tier: flex, billed 0.5xgoogle-ai-studio/gemini-3.6-flash+ flexgoogle-vertex/gemini-3.6-flash+ flexunsupported_service_tiervalidation errorgoogle-ai-studio/gemini-3.5-flash-lite+ flexgoogle-vertex/gemini-3.5-flash-lite+ flexON_DEMAND_FLEXgoogle-vertex/{3.6-flash,3.5-flash-lite}+ priorityON_DEMAND_PRIORITYpnpm test:unitgreen (172 files / 2866 tests), fullpnpm buildgreen.https://claude.ai/code/session_01WNfVNjJoNbVvrzurgZ5J78
Summary by CodeRabbit
Bug Fixes
Tests