Skip to content

feat(media): Vertex AI (Google) speech, transcription, music & video generation - #3929

Merged
diegosouzapw merged 3 commits into
diegosouzapw:release/v3.8.26from
artickc:feat/vertex-media-generation
Jun 15, 2026
Merged

diegosouzapw merged 3 commits into
diegosouzapw:release/v3.8.26from
artickc:feat/vertex-media-generation

Conversation

@artickc

@artickc artickc commented Jun 15, 2026

Copy link
Copy Markdown
Contributor

Summary

The dedicated media-generation endpoints only recognized third-party providers in their
registries:

  • /v1/videos/generations → kie, haiper, leonardo, pollinations, minimax, together, replicate, comfyui, sdwebui, veoaifree-web, runwayml
  • /v1/music/generations → kie, suno, udio, minimax, comfyui
  • /v1/audio/speech & /v1/audio/transcriptions → openai, deepgram, groq, elevenlabs, …

Google/Vertex was absent from all of them. So any vertex/* media model (Veo, Lyria,
Gemini TTS / native-audio) was rejected with Invalid <kind> model: vertex/... Use format: provider/model — even though the format was already correct. The provider simply wasn't
supported for those media types, while it is fully supported through Vertex AI.

Changes

  • open-sse/executors/vertexMedia.ts (new): a shared Vertex media client that reuses the
    Vertex chat executor's auth (parseSAFromApiKey / getAccessToken for Service Account JSON →
    OAuth bearer, or Express API key via ?key=) and implements the verified contracts:
    • Speech — {model}:generateContent with responseModalities:["AUDIO"] + speechConfig,
      returning PCM L16 which is wrapped into a WAV container.
    • Transcription — {model}:generateContent with inline audio + a transcribe instruction.
    • Music (Lyria) — {model}:predict → predictions[0].bytesBase64Encoded (WAV).
    • Video (Veo) — {model}:predictLongRunning → poll {model}:fetchPredictOperation until
      done → response.videos[0].bytesBase64Encoded (MP4).
  • Registries: added a vertex provider to videoRegistry (Veo 3.0/2.0), musicRegistry
    (Lyria 2), and audioRegistry speech (Gemini 2.5 TTS) + transcription (Gemini 2.5/2.0).
  • Handlers: added vertex-veo, vertex-lyria, vertex-gemini-tts, and vertex-gemini
    format branches to the respective handlers, returning the standard OpenAI-like shapes
    ({created, data:[{b64_json, format}]} for video/music; audio stream for speech; {text}
    for transcription).

Because all four reuse getProviderCredentials("vertex"), an existing Vertex connection (SA
JSON or Express key) now serves media generation with no extra configuration.

Verification

All four contracts were validated live against Vertex AI with a real Service Account:

Endpoint Result
Speech (gemini-2.5-flash-preview-tts) PCM L16 24kHz → WAV ✓
Transcription (gemini-2.5-flash) round-tripped "the quick brown fox…" ✓
Music (lyria-002:predict) WAV base64 ✓
Video (veo-3.0-fast-generate-001:predictLongRunning) operation → poll → MP4 ✓

Unit tests in tests/unit/vertex-media.test.ts (7 tests, all pass) cover WAV wrapping, the
SA-project vs Express ?key= URL building, request payload shapes for each modality, and the
Veo submit→poll→result flow (fetch mocked). The existing media handler/registry suites
(registry-utils, video-generation-handler, music-generation-handler,
audio-speech-handler, new-content-providers, registry-direct-exports) still pass (104/104).

Notes

  • Express-mode keys use the project-less global publisher endpoint (?key=) — best-effort;
    the Service Account JSON path is fully verified.
  • veo-3.1-fast-generate-preview is not GA in all projects; veo-3.0-* are used as the
    default registry entries.

…deo generation

The dedicated media endpoints (/v1/audio/speech, /v1/audio/transcriptions, /v1/music/generations, /v1/videos/generations) only supported third-party providers (kie/suno/minimax/openai/deepgram/...). Google/Vertex was absent, so any vertex/* media model was rejected with 'Invalid model. Use format: provider/model' even though the format was correct.

Adds a 'vertex' provider to the video, music, speech and transcription registries plus a shared open-sse/executors/vertexMedia.ts client that reuses the Vertex chat auth (Service Account JSON -> OAuth bearer, or Express API key) and implements the verified Vertex contracts: Gemini TTS (generateContent + responseModalities:[AUDIO] -> PCM wrapped as WAV), Gemini transcription (generateContent with inline audio), Lyria music (predict), and Veo video (predictLongRunning + fetchPredictOperation poll).

All four contracts verified live against Vertex AI. Unit tests (tests/unit/vertex-media.test.ts) cover URL building (SA vs Express), WAV wrapping, request payload shapes, and the Veo submit->poll flow.
@artickc
artickc requested a review from diegosouzapw as a code owner June 15, 2026 21:43
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

@diegosouzapw
diegosouzapw changed the base branch from main to release/v3.8.26 June 15, 2026 22:39
diegosouzapw and others added 2 commits June 15, 2026 19:51
…anches

PR diegosouzapw#3929 own growth: audioSpeech.ts 952->965, videoGeneration.ts 1026->1078
(vertex/* media branches). Core logic lives in vertexMedia.ts (under cap).

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
@diegosouzapw

Copy link
Copy Markdown
Owner

Thanks, @artickc! 🙏 Solid feature — Vertex AI media across all 4 endpoints (Gemini 2.5 TTS → WAV, transcription, Lyria 2 music, Veo 3.x video via predictLongRunning poll), reusing the existing Vertex SA/Express auth. The shared vertexMedia.ts executor (under the 800 cap) + per-handler vertex/* branches follow the existing executor/registry/handler pattern cleanly, and the 7 unit tests cover WAV RIFF wrapping, SA-vs-Express URL building, and the Veo poll loop. Green locally (7/7). Small baseline bump for the two handler branches. Merging into release/v3.8.26. 🚀

@diegosouzapw
diegosouzapw merged commit 2b88690 into diegosouzapw:release/v3.8.26 Jun 15, 2026
2 checks passed
@diegosouzapw diegosouzapw mentioned this pull request Jun 16, 2026
diegosouzapw added a commit that referenced this pull request Jun 16, 2026
OmniRoute v3.8.26 — see CHANGELOG.md [3.8.26] for the full notes.

Highlights: Vertex AI media generation (#3929), GLM-5.2 effort-tier routing (#3885),
sticky round-robin combos (#3846), OpenRouter connection presets (#3878), compression
prompt-cache fix (#3936/#3890), and a security pass (form-data/vite + workflow hardening, #3949).

Co-authored-by: artickc <artickc@users.noreply.github.com>
Co-authored-by: rdself <rdself@users.noreply.github.com>
Co-authored-by: herjarsa <herjarsa@users.noreply.github.com>
Co-authored-by: Jack Smith <16862258+YunyunZhai@users.noreply.github.com>
Co-authored-by: dhaern <dhaern@users.noreply.github.com>
Co-authored-by: adivekar-utexas <adivekar-utexas@users.noreply.github.com>
Co-authored-by: megamen32 <megamen32@users.noreply.github.com>
Co-authored-by: zhiru <zhiru@users.noreply.github.com>
Co-authored-by: insoln <insoln@users.noreply.github.com>
Co-authored-by: diego-anselmo <diego-anselmo@users.noreply.github.com>
HouMinXi pushed a commit to HouMinXi/OmniRoute that referenced this pull request Aug 2, 2026
OmniRoute v3.8.26 — see CHANGELOG.md [3.8.26] for the full notes.

Highlights: Vertex AI media generation (diegosouzapw#3929), GLM-5.2 effort-tier routing (diegosouzapw#3885),
sticky round-robin combos (diegosouzapw#3846), OpenRouter connection presets (diegosouzapw#3878), compression
prompt-cache fix (diegosouzapw#3936/diegosouzapw#3890), and a security pass (form-data/vite + workflow hardening, diegosouzapw#3949).

Co-authored-by: artickc <artickc@users.noreply.github.com>
Co-authored-by: rdself <rdself@users.noreply.github.com>
Co-authored-by: herjarsa <herjarsa@users.noreply.github.com>
Co-authored-by: Jack Smith <16862258+YunyunZhai@users.noreply.github.com>
Co-authored-by: dhaern <dhaern@users.noreply.github.com>
Co-authored-by: adivekar-utexas <adivekar-utexas@users.noreply.github.com>
Co-authored-by: megamen32 <megamen32@users.noreply.github.com>
Co-authored-by: zhiru <zhiru@users.noreply.github.com>
Co-authored-by: insoln <insoln@users.noreply.github.com>
Co-authored-by: diego-anselmo <diego-anselmo@users.noreply.github.com>
Poid-ZA pushed a commit to Poid-ZA/OmniRoute that referenced this pull request Aug 5, 2026
OmniRoute v3.8.26 — see CHANGELOG.md [3.8.26] for the full notes.

Highlights: Vertex AI media generation (diegosouzapw#3929), GLM-5.2 effort-tier routing (diegosouzapw#3885),
sticky round-robin combos (diegosouzapw#3846), OpenRouter connection presets (diegosouzapw#3878), compression
prompt-cache fix (diegosouzapw#3936/diegosouzapw#3890), and a security pass (form-data/vite + workflow hardening, diegosouzapw#3949).

Co-authored-by: artickc <artickc@users.noreply.github.com>
Co-authored-by: rdself <rdself@users.noreply.github.com>
Co-authored-by: herjarsa <herjarsa@users.noreply.github.com>
Co-authored-by: Jack Smith <16862258+YunyunZhai@users.noreply.github.com>
Co-authored-by: dhaern <dhaern@users.noreply.github.com>
Co-authored-by: adivekar-utexas <adivekar-utexas@users.noreply.github.com>
Co-authored-by: megamen32 <megamen32@users.noreply.github.com>
Co-authored-by: zhiru <zhiru@users.noreply.github.com>
Co-authored-by: insoln <insoln@users.noreply.github.com>
Co-authored-by: diego-anselmo <diego-anselmo@users.noreply.github.com>
tkgo11 pushed a commit to tkgo11/OmniRoute that referenced this pull request Sep 23, 2026
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
OmniRoute v3.8.26 — see CHANGELOG.md [3.8.26] for the full notes.

Highlights: Vertex AI media generation (diegosouzapw#3929), GLM-5.2 effort-tier routing (diegosouzapw#3885),
sticky round-robin combos (diegosouzapw#3846), OpenRouter connection presets (diegosouzapw#3878), compression
prompt-cache fix (diegosouzapw#3936/diegosouzapw#3890), and a security pass (form-data/vite + workflow hardening, diegosouzapw#3949).

Co-authored-by: artickc <artickc@users.noreply.github.com>
Co-authored-by: rdself <rdself@users.noreply.github.com>
Co-authored-by: herjarsa <herjarsa@users.noreply.github.com>
Co-authored-by: Jack Smith <16862258+YunyunZhai@users.noreply.github.com>
Co-authored-by: dhaern <dhaern@users.noreply.github.com>
Co-authored-by: adivekar-utexas <adivekar-utexas@users.noreply.github.com>
Co-authored-by: megamen32 <megamen32@users.noreply.github.com>
Co-authored-by: zhiru <zhiru@users.noreply.github.com>
Co-authored-by: insoln <insoln@users.noreply.github.com>
Co-authored-by: diego-anselmo <diego-anselmo@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants