Skip to content

feat(realtime): add Gemini Live voice support - #3315

Merged
steebchen merged 4 commits into
theopenco:mainfrom
RATCHAW:RATCHAW/google-live-realtime-voice
Jul 31, 2026
Merged

steebchen merged 4 commits into
theopenco:mainfrom
RATCHAW:RATCHAW/google-live-realtime-voice

Conversation

@RATCHAW

@RATCHAW RATCHAW commented Jul 29, 2026 •

Copy link
Copy Markdown
Contributor

Adds Gemini Live as a second realtime provider alongside OpenAI, proxying Google's native BidiGenerateContent protocol end to end rather than translating it: a sibling GeminiRealtimeProxySession handles setup validation, media/hosted-tool rejection, per-generation-stage billing from usageMetadata, and post-billing account/credit/rate-limit gates, while connectUpstream and the account gate are extracted into shared modules and the OpenAI path is left unchanged (session.spec.ts is untouched as the regression proof). Two models are added to the catalogue — gemini-2.5-flash-native-audio-preview-12-2025 and gemini-3.1-flash-live-preview — with prices taken from Google's published pricing page, and the playground gains a Gemini call hook plus a provider-selecting facade so the browser speaks whichever protocol serves the selected model.

Usage normalization was derived from live traces rather than assumption: Gemini excludes thoughtsTokenCount from totalTokenCount, and 3.1 leaves ~18 prompt tokens per turn unattributed to any modality, so the normalizer treats the reported total as an upper bound and bills unattributed input at the text rate — still failing closed on over-attribution, unpriceable modalities, or a block whose details explain less than half the total. Verified with pnpm build, pnpm lint, and 305 targeted unit tests, plus live calls through the local gateway against both models where a billed turn matched the persisted billing_cost exactly; no database migration is required.

One assumption remains unvalidated: a toolCall arriving without buffered usageMetadata fails the session closed, which no exercised path covers yet, so it is worth a look during review.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added Gemini Live realtime voice conversations with streamed audio, transcripts, tool interactions, and usage tracking.
    • Added new Gemini realtime models, provider-pinned model selection, native transcription, and saved-call continuation support.
    • Added provider-specific realtime connections with secure credential handling.
  • Bug Fixes

    • Improved account authorization, credit checks, billing safeguards, session draining, and backpressure handling.
    • Hardened malformed usage and connection error handling, including a Gemini availability switch.
  • Tests

    • Expanded coverage for Gemini pricing, model selection, connections, and realtime sessions.

RATCHAW and others added 2 commits July 29, 2026 17:32
Serve Google's Gemini Live models over /v1/realtime using Gemini's own
BidiGenerateContent protocol, alongside the existing OpenAI realtime path.
Shared connect/preflight/account-gate logic is factored out of session.ts
so both providers run the same authorization and billing gates.

Adds gemini-2.5-flash-native-audio-preview-12-2025 and
gemini-3.1-flash-live-preview to the catalogue, a strict usage normalizer
for Gemini's usageMetadata blocks, and the playground call UI.

Co-Authored-By: Claude <noreply@anthropic.com>
…-realtime-voice

# Conflicts:
#	apps/playground/src/components/playground/realtime-page-client.tsx
@coderabbitai

coderabbitai Bot commented Jul 29, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 826f6710-9b90-4cf0-95a4-e9b285e499d3

📥 Commits

Reviewing files that changed from the base of the PR and between 67f8bb3 and 972682c.

📒 Files selected for processing (2)
  • apps/playground/src/components/playground/realtime-page-client.tsx
  • apps/playground/src/hooks/use-voice-call.ts

Walkthrough

Adds Gemini realtime model definitions, provider-specific gateway proxying, usage normalization and billing, account gating, a Gemini WebSocket session, and provider-aware playground voice-call support with validation and lifecycle tests.

Changes

Gemini realtime voice integration

Layer / File(s) Summary
Model catalog and provider selection
packages/models/src/models/google.ts, packages/models/src/realtime-models.spec.ts, apps/gateway/src/realtime/catalog.spec.ts, apps/playground/src/lib/realtime-model-value.*, apps/playground/src/app/api/realtime/session/route.ts
Adds two Gemini realtime mappings, provider-pinned model selection, native transcription handling, and related catalog and selector tests.
Gateway authorization and provider wiring
apps/gateway/src/realtime/account-gate.ts, apps/gateway/src/realtime/connect-upstream.*, apps/gateway/src/realtime/preflight.*, apps/gateway/src/realtime/server.*, apps/gateway/src/realtime/session.ts
Centralizes account and credit gates, adds Gemini upstream targeting and handshake handling, supports the Gemini kill switch, recognizes provider-neutral subprotocols, and creates provider-specific sessions.
Gemini session protocol and billing
apps/gateway/src/realtime/gemini-pricing.*, apps/gateway/src/realtime/gemini-session.*
Adds Gemini message validation, upstream frame forwarding, usage normalization, per-stage billing, post-billing gates, draining, shutdown, backpressure, and lifecycle coverage.
Provider-aware playground voice calls
apps/playground/src/hooks/use-gemini-realtime-call.ts, apps/playground/src/hooks/use-voice-call.ts, apps/playground/src/components/playground/realtime-page-client.tsx
Adds Gemini audio capture and playback, transcription and usage tracking, barge-in handling, provider selection, and mapping-aware call startup.

Estimated code review effort: 5 (Critical) | ~120 minutes

Sequence Diagram(s)

sequenceDiagram
  participant VoiceCall as useVoiceCall
  participant SessionRoute as realtime session route
  participant Gateway as realtime server
  participant GeminiSession as GeminiRealtimeProxySession
  participant Google as Gemini Live
  participant Billing as account and billing services

  VoiceCall->>SessionRoute: request realtime session
  SessionRoute->>Gateway: open WebSocket with provider and model
  Gateway->>GeminiSession: create provider-specific session
  GeminiSession->>Google: connect using BidiGenerateContent target
  VoiceCall->>GeminiSession: send setup and audio input
  GeminiSession->>Google: forward validated Gemini frames
  Google-->>GeminiSession: content and usageMetadata frames
  GeminiSession->>Billing: normalize usage and persist stage billing
  GeminiSession-->>VoiceCall: forward content, transcripts, and terminal events
Loading

Possibly related PRs

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 51.85% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the primary change: adding Gemini Live voice support to realtime functionality.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: bf8eca8a39

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

if (this.finalized) {
return;
}
this.finalized = true;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Drain pending usage before finalizing locally

When a local shutdown occurs after a usageMetadata snapshot but before its terminal event—for example, the duration timer, ping timeout, backpressure timeout, or server closeAll—setting finalized here causes handleUpstreamClose to return without billing pendingUsage. Google may therefore charge the generation while the gateway records no billing row; locally initiated shutdowns need to drain or explicitly bill any buffered stage before suppressing the close handler.

Useful? React with 👍 / 👎.

Comment on lines +358 to +362
if (turnClosedRef.current || !userTurnIdRef.current) {
turnCounterRef.current += 1;
userTurnIdRef.current = `user-${turnCounterRef.current}`;
assistantTurnIdRef.current = null;
turnClosedRef.current = false;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Keep late input transcription on the completed turn

When Gemini delivers inputTranscription after turnComplete—an ordering the surrounding comments explicitly intend to support—turnClosedRef.current is already true, so this condition allocates a new user ID and clears the previous assistant ID. The late text is consequently saved as a new, incorrectly paired conversation turn instead of being appended to the completed user bubble.

Useful? React with 👍 / 👎.

Comment on lines +584 to +588
const response = await fetch("/api/realtime/session", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ model: currentModel }),
});

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Move session minting behind the typed internal API

Replace this new raw request to the Next.js /api/realtime/session route with the generated internal-API client and host the minting operation in apps/api. As written, the Gemini call path has no generated request/response contract, so changes to the session payload or authentication can compile successfully and fail only at runtime, contrary to the repository's explicit frontend and backend-operation boundary.

AGENTS.md reference: AGENTS.md:L212-L213

Useful? React with 👍 / 👎.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (3)
apps/gateway/src/realtime/account-gate.ts (1)

67-80: 🚀 Performance & Scalability | 🔵 Trivial | ⚡ Quick win

Parallelize the independent account lookups.

findApiKeyByToken, findOrganizationById, and findProjectById don't depend on each other but are awaited sequentially. This function runs on every response.create and every completed transcription in a live voice session (see runGenerationGatesInner/enforceLimitsAfterTranscription in session.ts), so serializing three round trips adds avoidable latency to a user-facing low-latency path.

⚡ Proposed fix
-	let freshKey;
-	let freshOrg;
-	let freshProject;
 	try {
-		freshKey = await findApiKeyByToken(input.gatewayToken);
-		freshOrg = await findOrganizationById(preflight.project.organizationId);
-		freshProject = await findProjectById(preflight.project.id);
+		[freshKey, freshOrg, freshProject] = await Promise.all([
+			findApiKeyByToken(input.gatewayToken),
+			findOrganizationById(preflight.project.organizationId),
+			findProjectById(preflight.project.id),
+		]);
 	} catch (error) {
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@apps/gateway/src/realtime/account-gate.ts` around lines 67 - 80, Update the
account lookup block in the realtime authorization flow to start
findApiKeyByToken, findOrganizationById, and findProjectById concurrently and
await their results together, while preserving the existing error handling and
assignments to freshKey, freshOrg, and freshProject.
apps/playground/src/hooks/use-gemini-realtime-call.ts (1)

446-455: 🩺 Stability & Availability | 🔵 Trivial | 💤 Low value

Clear any existing elapsed timer before starting a new one.

A second setupComplete frame (or a re-handshake) would overwrite elapsedTimerRef.current and leak the previous interval until cleanup() runs.

♻️ Proposed guard
 			if (message.setupComplete !== undefined) {
 				updateStatus("live");
 				setElapsedSeconds(0);
+				if (elapsedTimerRef.current) {
+					clearInterval(elapsedTimerRef.current);
+				}
 				const startedAt = Date.now();
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@apps/playground/src/hooks/use-gemini-realtime-call.ts` around lines 446 -
455, In the setupComplete handling block, clear any existing elapsed timer
before assigning a new interval to elapsedTimerRef.current. Preserve the current
timer initialization and elapsed-time updates while preventing repeated
setupComplete frames or re-handshakes from leaving the prior interval active.
apps/playground/src/app/api/realtime/session/route.ts (1)

166-166: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Consider sharing the Gemini provider id constant.

"google-ai-studio" is hardcoded here and again as GEMINI_PROVIDER_ID in apps/playground/src/hooks/use-voice-call.ts (Line 17). A single exported constant keeps the mint-side and client-side protocol decisions from drifting.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@apps/playground/src/app/api/realtime/session/route.ts` at line 166, Replace
the hardcoded provider ID in the usesNativeTranscription logic with the shared
GEMINI_PROVIDER_ID constant already used by use-voice-call.ts, importing or
exporting it as needed so mint-side and client-side checks use one source of
truth.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@apps/gateway/src/realtime/gemini-session.ts`:
- Around line 1171-1180: Update the drain-timeout callback around
this.drainTimer to bill any buffered usage before shutting down: invoke
billStage("cancelled", { requireUsage: false }) using the same handling as the
upstream-close path, then retain the existing shutdown behavior and diagnostics
logging.

---

Nitpick comments:
In `@apps/gateway/src/realtime/account-gate.ts`:
- Around line 67-80: Update the account lookup block in the realtime
authorization flow to start findApiKeyByToken, findOrganizationById, and
findProjectById concurrently and await their results together, while preserving
the existing error handling and assignments to freshKey, freshOrg, and
freshProject.

In `@apps/playground/src/app/api/realtime/session/route.ts`:
- Line 166: Replace the hardcoded provider ID in the usesNativeTranscription
logic with the shared GEMINI_PROVIDER_ID constant already used by
use-voice-call.ts, importing or exporting it as needed so mint-side and
client-side checks use one source of truth.

In `@apps/playground/src/hooks/use-gemini-realtime-call.ts`:
- Around line 446-455: In the setupComplete handling block, clear any existing
elapsed timer before assigning a new interval to elapsedTimerRef.current.
Preserve the current timer initialization and elapsed-time updates while
preventing repeated setupComplete frames or re-handshakes from leaving the prior
interval active.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 79772690-d20f-44a8-aa52-60bc7474fd1b

📥 Commits

Reviewing files that changed from the base of the PR and between f4e89b3 and bf8eca8.

📒 Files selected for processing (21)
  • apps/gateway/src/realtime/account-gate.ts
  • apps/gateway/src/realtime/catalog.spec.ts
  • apps/gateway/src/realtime/connect-upstream.spec.ts
  • apps/gateway/src/realtime/connect-upstream.ts
  • apps/gateway/src/realtime/gemini-pricing.spec.ts
  • apps/gateway/src/realtime/gemini-pricing.ts
  • apps/gateway/src/realtime/gemini-session.spec.ts
  • apps/gateway/src/realtime/gemini-session.ts
  • apps/gateway/src/realtime/preflight.spec.ts
  • apps/gateway/src/realtime/preflight.ts
  • apps/gateway/src/realtime/server.spec.ts
  • apps/gateway/src/realtime/server.ts
  • apps/gateway/src/realtime/session.ts
  • apps/playground/src/app/api/realtime/session/route.ts
  • apps/playground/src/components/playground/realtime-page-client.tsx
  • apps/playground/src/hooks/use-gemini-realtime-call.ts
  • apps/playground/src/hooks/use-voice-call.ts
  • apps/playground/src/lib/realtime-model-value.spec.ts
  • apps/playground/src/lib/realtime-model-value.ts
  • packages/models/src/models/google.ts
  • packages/models/src/realtime-models.spec.ts

Comment thread apps/gateway/src/realtime/gemini-session.ts
RATCHAW and others added 2 commits July 29, 2026 18:15
…ealtime-voice

# Conflicts:
#	apps/playground/src/components/playground/realtime-page-client.tsx
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@steebchen
steebchen added this pull request to the merge queue Jul 31, 2026
Merged via the queue into theopenco:main with commit 0484221 Jul 31, 2026
9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants