Skip to content

feat(models): add Gemini TTS on Vertex+AI Studio - #3143

Merged
steebchen merged 1 commit into
mainfrom
add-gemini-tts-models
Jul 19, 2026
Merged

steebchen merged 1 commit into
mainfrom
add-gemini-tts-models

Conversation

@steebchen

Copy link
Copy Markdown
Member

Summary

Adds the two Gemini text-to-speech models via both Google Vertex AI and Google AI Studio:

  • Gemini 3.1 Flash TTS Preview (gemini-3.1-flash-tts-preview, released 2026-04-15): new model entry with google-ai-studio and google-vertex mappings — $1.00/M text input, $20.00/M audio output (per Gemini API pricing), 8k context / 16k max output, same 30-voice catalog as the 2.5 TTS family.
  • Gemini 2.5 Pro TTS: adds a google-vertex mapping to the existing gemini-2.5-pro-preview-tts model using the GA Vertex model id gemini-2.5-pro-tts (the preview id is AI-Studio-only, per the Gemini-TTS docs). Same pricing as the existing AI Studio mapping ($1/M in, $20/M audio out). Both models are served on Vertex's global location, matching the gateway's default Vertex region.

Gateway: google-vertex support for /v1/audio/speech

The speech endpoint previously only supported google-ai-studio, openai, and elevenlabs — a Vertex TTS mapping would have been rejected with 400 unsupported_provider. This PR adds a google-vertex branch mirroring the embeddings Vertex path:

  • URL: https://aiplatform.googleapis.com/v1/projects/{project}/locations/{region}/publishers/google/models/{model}:generateContent
  • Project id from google_vertex_project_id provider-key option or LLM_GOOGLE_CLOUD_PROJECT; region from LLM_GOOGLE_VERTEX_REGION defaulting to global
  • Token type via resolveVertexTokenType so the Authorization: Bearer header and the ?key= query param always agree; the resolved type is passed to getProviderHeaders
  • Response parsing is unchanged — Vertex returns the same inline-PCM generateContent shape as AI Studio

Tests

  • New speech.e2e.ts driven by a new speechModels list in chat-helpers.e2e.ts (same TEST_MODELS/deactivation/stability filtering as embeddingModels), asserting a valid RIFF/WAVE payload. The speech endpoint previously had no e2e coverage.
  • Unit specs: two new speech.spec.ts tests covering the Vertex WAV path (including usedModelMapping/cost assertions) and the 3.1 Flash PCM path; the mock server's Vertex generateContent route now handles the AUDIO response modality.
  • Live e2e run (real providers): all four mappings return real audio (215–290 KB WAV):
    TEST_MODELS="google-ai-studio/gemini-3.1-flash-tts-preview,google-vertex/gemini-3.1-flash-tts-preview,google-vertex/gemini-2.5-pro-preview-tts,google-ai-studio/gemini-2.5-pro-preview-tts" pnpm test:e2e
    ✓ speech google-ai-studio/gemini-2.5-pro-preview-tts
    ✓ speech google-vertex/gemini-2.5-pro-preview-tts
    ✓ speech google-ai-studio/gemini-3.1-flash-tts-preview
    ✓ speech google-vertex/gemini-3.1-flash-tts-preview
    
  • Full pnpm test:unit (171 files, 2852 tests) and pnpm build pass; pnpm format run.

🤖 Generated with Claude Code

Add gemini-3.1-flash-tts-preview (AI Studio + Vertex) and a
google-vertex mapping (GA id gemini-2.5-pro-tts) for
gemini-2.5-pro-preview-tts, both at $1/M text input and $20/M
audio output.

Implement google-vertex support in the /v1/audio/speech handler:
project/region/token-type resolution mirroring the embeddings
Vertex path, with OAuth-vs-API-key auth agreement between the
Authorization header and the ?key= query param.

Add a speech.e2e.ts suite driven by a new speechModels list so
TEST_MODELS can exercise TTS mappings end-to-end, plus unit specs
and mock-server AUDIO support for the Vertex generateContent route.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings July 19, 2026 13:18
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@coderabbitai

coderabbitai Bot commented Jul 19, 2026

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@steebchen, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 13 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro

Run ID: 08c61cb8-b687-48b7-b53f-b2e13f86fdc3

📥 Commits

Reviewing files that changed from the base of the PR and between 135a701 and b52ee34.

📒 Files selected for processing (6)
  • apps/gateway/src/chat-helpers.e2e.ts
  • apps/gateway/src/speech.e2e.ts
  • apps/gateway/src/speech/speech.spec.ts
  • apps/gateway/src/speech/speech.ts
  • apps/gateway/src/test-utils/mock-openai-server.ts
  • packages/models/src/models/google.ts
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch add-gemini-tts-models

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@steebchen steebchen changed the title feat(models): add Gemini TTS via Vertex + AI Studio feat(models): add Gemini TTS on Vertex+AI Studio Jul 19, 2026
@steebchen
steebchen enabled auto-merge July 19, 2026 13:25
@steebchen
steebchen added this pull request to the merge queue Jul 19, 2026
Merged via the queue into main with commit 54699d5 Jul 19, 2026
22 checks passed
@steebchen
steebchen deleted the add-gemini-tts-models branch July 19, 2026 13:41
pull Bot pushed a commit to soitun/llmgateway that referenced this pull request Jul 23, 2026
## Summary

Three related marketing-surface updates:

**SCX.ai "Up to 4x faster" badge on model cards**
(https://llmgateway.io/providers/scx-ai)
- New optional `modelCardBadge` field on `ProviderDefinition`
(data-driven — any provider can carry one), set to `"Up to 4x faster"`
on `scx-ai`.
- Rendered in the shared model card's provider header as an amber Zap
badge, next to the provider name. Plumbed through the shared
`ApiProvider` type and the provider detail page's catalogue conversion.

**Theme-aware SCX logo**
- The SCX wordmark SVG hardcoded `fill="#262626"`, making it
near-invisible in dark mode. All four fills now use `currentColor`, so
it renders black in light mode and white in dark mode (matching the
pattern used by the other monochrome provider icons).

**Changelog: July roundup**
(`/changelog/upgrade-rollover-new-providers`)
- New entry (id 67) covering everything user-facing since the Jul 16
Reset Passes entry: DevPass upgrade rollover + upgrade-timing choice
(theopenco#3147), SCX.ai (theopenco#2958) and Gonka24 (theopenco#3189) providers, Nebius mappings
(theopenco#3187), Gemini 3.6 Flash / 3.5 Flash Lite (theopenco#3164), Gemini TTS on two
providers (theopenco#3143), the Empryo coding agent (theopenco#3161), and the new
request-timeouts docs page (theopenco#3190).
- OG image generated with gpt-image-2 in the house circuit-board style.

## Testing

- Full `pnpm build` passes (validates the changelog frontmatter against
the content-collections schema and typechecks the badge plumbing).
- Verified live on the dev server: badge renders on every SCX model card
in both themes, and the logo is black-on-light / white-on-dark
(screenshots shared in session).

https://claude.ai/code/session_01UGktbNs7a35FU1k2ooKfaB

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added provider badges to model cards, including an “Up to 4x faster”
badge for SCX.ai.
* Added new inference providers and expanded model availability,
including Gemini models and Gemini TTS.
* Upgrade credits can roll over, with immediate or next-renewal upgrade
options.
* Added Empryo coding-agent usage attribution and request-timeout
documentation.
* **Style**
  * Improved provider icon coloring for better theme compatibility.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Luca Steeb <contact@luca-steeb.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants