Skip to content

feat(api): add Google AI Studio Gemini TTS - #11315

Merged
diegosouzapw merged 1 commit into
diegosouzapw:release/v3.8.50from
RaviTharuma:fix/10590-gemini-tts
Aug 24, 2026
Merged

diegosouzapw merged 1 commit into
diegosouzapw:release/v3.8.50from
RaviTharuma:fix/10590-gemini-tts

Conversation

@RaviTharuma

@RaviTharuma RaviTharuma commented Aug 24, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • register public google/gemini-*-tts speech models while reusing stored gemini credentials
  • translate OpenAI-compatible /v1/audio/speech requests to the Google AI Studio generateContent audio contract
  • reuse the Vertex inline-audio parser, PCM sample-rate parser, and WAV wrapper instead of duplicating conversion logic
  • preserve AI Studio upstream status/messages and reject successful responses that contain no audio

This intentionally covers batch TTS only. Gemini Live / bidirectional sessions remain a separate protocol and are not implemented here.

Request contract

POST https://generativelanguage.googleapis.com/v1beta/models/{model}:generateContent

  • authentication: x-goog-api-key for Gemini API keys (Bearer fallback for OAuth credentials)
  • responseModalities: ["AUDIO"]
  • speechConfig.voiceConfig.prebuiltVoiceConfig.voiceName
  • inline audio/L16;...;rate=N is decoded and returned as audio/wav

Tests Added Or Updated

  • tests/unit/gemini-tts.test.ts covers public parsing, credential remap, exact URL/header/body, API-key header precedence, PCM sample-rate-to-WAV conversion, default voice, missing audio, and upstream errors
  • tests/unit/vertex-media.test.ts covers the reused Vertex parser/conversion sibling path
  • tests/unit/audio-speech-handler.test.ts covers the complete speech-handler sibling surface

Verification

  • Focused tests and category gates from the golden path
  • rebased onto release/v3.8.50 at ac02c5b42
  • npx tsx --test tests/unit/gemini-tts.test.ts tests/unit/vertex-media.test.ts tests/unit/audio-speech-handler.test.ts (41/41)
  • focused ESLint on all changed TypeScript files
  • npm run typecheck:core
  • npm run typecheck:noimplicit:core
  • changed-file formatting verified (the pre-existing CRLF vertexMedia.ts is intentionally kept as a 3-line semantic diff)
  • changelog fragment at changelog.d/features/10590-google-ai-studio-tts.md
  • upstream GitHub checks (running)

Closes #10590
Related: #10591

@RaviTharuma

Copy link
Copy Markdown
Contributor Author

Implementation evidence:

  • Focused + sibling tests: 41/41 pass
  • Focused ESLint: pass
  • typecheck:core: pass
  • typecheck:noimplicit:core: pass
  • Prettier: pass
  • Branch is mergeable; four upstream fast-path CI shards are now running.

I also attempted squash auto-merge as requested, but GitHub rejected it because this fork account does not have upstream merge permission.

@RaviTharuma
RaviTharuma force-pushed the fix/10590-gemini-tts branch from c52cf6c to 5713b74 Compare August 24, 2026 02:36
@RaviTharuma

Copy link
Copy Markdown
Contributor Author

Rebased onto current release/v3.8.50 (ac02c5b42), reran the 41 focused/sibling tests plus ESLint and both core typechecks, moved the changelog fragment into changelog.d/features/, and force-pushed with lease. Current head: 5713b746b.

@RaviTharuma
RaviTharuma force-pushed the fix/10590-gemini-tts branch from 5713b74 to a0b367e Compare August 24, 2026 03:26
@RaviTharuma

Copy link
Copy Markdown
Contributor Author

Follow-up CI hardening pushed as a0b367eec:

  • extracted Google speech response handling from the frozen speech handler, reducing open-sse/handlers/audioSpeech.ts to 997 lines
  • check:file-size -- --base-ref ac02c5b42: pass
  • check:open-sse-typecheck: pass (only 5 frozen pre-existing errors remain)
  • focused/sibling tests: 41/41 pass
  • focused ESLint: pass

The prior DAST failure is unrelated to this audio-only diff (Schemathesis probing /v1/models returned 500); the force-push triggered fresh checks.

@RaviTharuma

Copy link
Copy Markdown
Contributor Author

Merge automation was attempted again with both gh pr merge --squash --auto and direct --squash. GitHub rejected both because RaviTharuma lacks upstream MergePullRequest permission. The branch is clean, pushed, and GitHub reports it mergeable; fresh CI is running on a0b367eec.

@RaviTharuma
RaviTharuma force-pushed the fix/10590-gemini-tts branch from a0b367e to a693631 Compare August 24, 2026 04:08
@RaviTharuma

Copy link
Copy Markdown
Contributor Author

Fixed the shard-3 failure exposed by CI: after extracting the Gemini adapter for the file-size gate, non-upstream adapter errors escaped the parent handler instead of preserving its JSON 500 contract. handleGeminiTtsSpeech now returns the same sanitized Speech request failed: ... response.

Current head a693631dd:

  • focused/sibling tests 41/41 pass
  • file-size gate passes (audioSpeech.ts: 997 lines)
  • open-sse typecheck gate passes
  • focused ESLint passes

Fresh checks were triggered.

@RaviTharuma

Copy link
Copy Markdown
Contributor Author

CI follow-up on a693631dd: Fast Quality Gates and shards 1/2/3 now pass. Shard 4 failed only on unrelated timing flake tests/unit/stream-timing.test.ts:78 (assert.ok(total >= 15)); this PR does not touch stream timing. Attempting a job rerun returned HTTP 403 because the fork account lacks upstream Actions admin rights. The separate ESLint job still fails on stale upstream suppressions (There are suppressions left that do not occur anymore), while focused changed-file ESLint passes.

@RaviTharuma

Copy link
Copy Markdown
Contributor Author

Additional unrelated CI evidence: DAST failed on DELETE /api/keys/{id} returning 500 for Schemathesis-generated id �, outside this audio-only diff. The feature-specific and impacted quality gates pass. Merge/auto-merge remain permission-blocked for this fork account.

@diegosouzapw
diegosouzapw merged commit 9b14896 into diegosouzapw:release/v3.8.50 Aug 24, 2026
13 of 16 checks passed
@RaviTharuma
RaviTharuma deleted the fix/10590-gemini-tts branch September 23, 2026 19:37
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
Validated on a 17-PR combined board: gemini-tts + vertex-media + audio-speech-handler (41/41) within the board's 287/287, typecheck:core clean. Registers public google/gemini-*-tts speech models and translates OpenAI-compatible /v1/audio/speech to the AI Studio generateContent audio contract, reusing the Vertex inline-audio/PCM/WAV conversion path. Batch TTS only, Gemini Live is out of scope. Thank you @RaviTharuma!
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(providers): Google AI Studio Gemini TTS + Gemini Live are not proxied

2 participants