feat(api): add Google AI Studio Gemini TTS - #11315
Conversation
|
Implementation evidence:
I also attempted squash auto-merge as requested, but GitHub rejected it because this fork account does not have upstream merge permission. |
c52cf6c to
5713b74
Compare
|
Rebased onto current |
5713b74 to
a0b367e
Compare
|
Follow-up CI hardening pushed as
The prior DAST failure is unrelated to this audio-only diff (Schemathesis probing |
|
Merge automation was attempted again with both |
a0b367e to
a693631
Compare
|
Fixed the shard-3 failure exposed by CI: after extracting the Gemini adapter for the file-size gate, non-upstream adapter errors escaped the parent handler instead of preserving its JSON 500 contract. Current head
Fresh checks were triggered. |
|
CI follow-up on |
|
Additional unrelated CI evidence: DAST failed on |
9b14896
into
diegosouzapw:release/v3.8.50
Validated on a 17-PR combined board: gemini-tts + vertex-media + audio-speech-handler (41/41) within the board's 287/287, typecheck:core clean. Registers public google/gemini-*-tts speech models and translates OpenAI-compatible /v1/audio/speech to the AI Studio generateContent audio contract, reusing the Vertex inline-audio/PCM/WAV conversion path. Batch TTS only, Gemini Live is out of scope. Thank you @RaviTharuma!
Summary
google/gemini-*-ttsspeech models while reusing storedgeminicredentials/v1/audio/speechrequests to the Google AI StudiogenerateContentaudio contractThis intentionally covers batch TTS only. Gemini Live / bidirectional sessions remain a separate protocol and are not implemented here.
Request contract
POST https://generativelanguage.googleapis.com/v1beta/models/{model}:generateContentx-goog-api-keyfor Gemini API keys (Bearer fallback for OAuth credentials)responseModalities: ["AUDIO"]speechConfig.voiceConfig.prebuiltVoiceConfig.voiceNameaudio/L16;...;rate=Nis decoded and returned asaudio/wavTests Added Or Updated
tests/unit/gemini-tts.test.tscovers public parsing, credential remap, exact URL/header/body, API-key header precedence, PCM sample-rate-to-WAV conversion, default voice, missing audio, and upstream errorstests/unit/vertex-media.test.tscovers the reused Vertex parser/conversion sibling pathtests/unit/audio-speech-handler.test.tscovers the complete speech-handler sibling surfaceVerification
release/v3.8.50atac02c5b42npx tsx --test tests/unit/gemini-tts.test.ts tests/unit/vertex-media.test.ts tests/unit/audio-speech-handler.test.ts(41/41)npm run typecheck:corenpm run typecheck:noimplicit:corevertexMedia.tsis intentionally kept as a 3-line semantic diff)changelog.d/features/10590-google-ai-studio-tts.mdCloses #10590
Related: #10591