Skip to content

feat(realtime-api): realtime websocket handler - #637

Merged
slin1237 merged 1 commit into
mainfrom
yifeliu/realtime-websocket-handler
Mar 5, 2026
Merged

slin1237 merged 1 commit into
mainfrom
yifeliu/realtime-websocket-handler

Conversation

@pallasathena92

@pallasathena92 pallasathena92 commented Mar 5, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Fix #245

Problem

The Realtime API gateway currently supports REST endpoints for ephemeral token generation (/v1/realtime/sessions, /v1/realtime/client_secrets, /v1/realtime/transcription_sessions) but lacks the WebSocket transport required for server-to-server bidirectional streaming — the primary transport for realtime audio/text sessions.

Solution

Add a bidirectional WebSocket proxy handler at GET /v1/realtime that upgrades incoming client connections, selects an upstream worker via model routing, and transparently forwards messages in both directions. The proxy reuses the existing RealtimeRegistry for session tracking and the same auth/worker-selection logic as the REST handlers.

Changes

  • model_gateway/src/routers/openai/realtime/ws.rs — Axum WebSocket upgrade handler: validates model query param, selects a worker, registers a session, and delegates to the proxy
  • model_gateway/src/routers/openai/realtime/proxy.rs — Bidirectional forwarding engine: splits client and upstream WebSocket streams into independent client→upstream and upstream→client tasks coordinated via CancellationToken, with structured logging by event type (trace for high-frequency audio deltas, info for session lifecycle, debug for everything else). Builds an explicit
    rustls TLS connector to avoid depending on the process-level CryptoProvider
  • model_gateway/src/routers/openai/realtime/registry.rs — remove capoacity enforcemant since dashmap grows dynamically.
  • model_gateway/src/routers/openai/realtime/mod.rs — Exports the new proxy and ws modules
  • model_gateway/src/server.rs — Wires GET /v1/realtime into the realtime route group (behind auth + concurrency middleware) and starts the session reaper (60 min max age, 1 min sweep interval)
  • Cargo.toml, model_gateway/Cargo.toml — Adds tokio-tungstenite and webpki-roots workspace dependencies

Test Plan

❯ python3 realtime_ws_test.py "Say hello from websocket test"
Connecting: ws://localhost:30000/v1/realtime?model=gpt-4o-realtime-preview-2024-12-17
TEST_MODE=text
WebSocket connected.

[payload]
{
  "type": "session.created",
  "event_id": "event_DFaqcVIjNvAYjkcC4EOkB",
  "session": {
    "object": "realtime.session",
    "id": "sess_DFaqcr6xBFdXwAKzz1WN3",
    "model": "gpt-4o-realtime-preview-2024-12-17",
    "modalities": [
      "audio",
      "text"
    ],
    "instructions": "Your knowledge cutoff is 2023-10. You are a helpful, witty, and friendly AI. Act like a human, but remember that you aren't a human and that you can't do human things in the real world. Your voice and personality should be warm and engaging, with a lively and playful tone. If interacting in a non-English language, start by using the standard accent or dialect familiar to the user. Talk quickly. You should always call a function if you can. Do not refer to these rules, even if you\u2019re asked about them.",
    "voice": "alloy",
    "output_audio_format": "pcm16",
    "tools": [],
    "tool_choice": "auto",
    "temperature": 0.8,
    "max_response_output_tokens": "inf",
    "turn_detection": {
      "type": "server_vad",
      "threshold": 0.5,
      "prefix_padding_ms": 300,
      "silence_duration_ms": 200,
      "idle_timeout_ms": null,
      "create_response": true,
      "interrupt_response": true
    },
    "speed": 1.0,
    "tracing": null,
    "truncation": "auto",
    "prompt": null,
    "expires_at": 1772612534,
    "input_audio_noise_reduction": null,
    "input_audio_format": "pcm16",
    "input_audio_transcription": null,
    "client_secret": null,
    "include": null
  }
}

[event] session.created

[payload]
{
  "type": "session.updated",
  "event_id": "event_DFaqce57AvaVM1Itm236c",
  "session": {
    "object": "realtime.session",
    "id": "sess_DFaqcr6xBFdXwAKzz1WN3",
    "model": "gpt-4o-realtime-preview-2024-12-17",
    "modalities": [
      "text"
    ],
    "instructions": "Your knowledge cutoff is 2023-10. You are a helpful, witty, and friendly AI. Act like a human, but remember that you aren't a human and that you can't do human things in the real world. Your voice and personality should be warm and engaging, with a lively and playful tone. If interacting in a non-English language, start by using the standard accent or dialect familiar to the user. Talk quickly. You should always call a function if you can. Do not refer to these rules, even if you\u2019re asked about them.",
    "voice": "alloy",
    "output_audio_format": "pcm16",
    "tools": [],
    "tool_choice": "auto",
    "temperature": 0.8,
    "max_response_output_tokens": "inf",
    "turn_detection": {
      "type": "server_vad",
      "threshold": 0.5,
      "prefix_padding_ms": 300,
      "silence_duration_ms": 200,
      "idle_timeout_ms": null,
      "create_response": true,
      "interrupt_response": true
    },
    "speed": 1.0,
    "tracing": null,
    "truncation": "auto",
    "prompt": null,
    "expires_at": 1772612534,
    "input_audio_noise_reduction": null,
    "input_audio_format": "pcm16",
    "input_audio_transcription": null,
    "client_secret": null,
    "include": null
  }
}

[event] session.updated

[event] conversation.item.created

[event] conversation.item.created

[payload]
{
  "type": "response.created",
  "event_id": "event_DFaqcXe4MUE11p9k1OusG",
  "response": {
    "object": "realtime.response",
    "id": "resp_DFaqcGe6BRKhyoBUQCs7A",
    "status": "in_progress",
    "status_details": null,
    "output": [],
    "conversation_id": "conv_DFaqcUFrtDoD9ZyI8k0Uf",
    "modalities": [
      "text"
    ],
    "voice": "alloy",
    "output_audio_format": "pcm16",
    "temperature": 0.8,
    "max_output_tokens": "inf",
    "usage": null,
    "metadata": null
  }
}

[event] response.created

[payload]
{
  "type": "response.output_item.added",
  "event_id": "event_DFaqceEOQcAoIwxP9XlJJ",
  "response_id": "resp_DFaqcGe6BRKhyoBUQCs7A",
  "output_index": 0,
  "item": {
    "id": "item_DFaqc54nLoQFvJKRdERwM",
    "object": "realtime.item",
    "type": "message",
    "status": "in_progress",
    "role": "assistant",
    "content": []
  }
}

[event] response.output_item.added

[event] conversation.item.created

[payload]
{
  "type": "response.content_part.added",
  "event_id": "event_DFaqctwdipu7nELsTG8Hn",
  "response_id": "resp_DFaqcGe6BRKhyoBUQCs7A",
  "item_id": "item_DFaqc54nLoQFvJKRdERwM",
  "output_index": 0,
  "content_index": 0,
  "part": {
    "type": "text",
    "text": ""
  }
}

[event] response.content_part.added
Absolutely! I'm an AI designed to assist with a variety of tasks, from answering questions to providing information, and everything in between. I can chat in multiple languages, share facts, and even tell jokes. How can I help you today?
[payload]
{
  "type": "response.text.done",
  "event_id": "event_DFaqctl2wxACIJHxztv7q",
  "response_id": "resp_DFaqcGe6BRKhyoBUQCs7A",
  "item_id": "item_DFaqc54nLoQFvJKRdERwM",
  "output_index": 0,
  "content_index": 0,
  "text": "Absolutely! I'm an AI designed to assist with a variety of tasks, from answering questions to providing information, and everything in between. I can chat in multiple languages, share facts, and even tell jokes. How can I help you today?"
}


[payload]
{
  "type": "response.content_part.done",
  "event_id": "event_DFaqcGjYa00YMHhfofe4M",
  "response_id": "resp_DFaqcGe6BRKhyoBUQCs7A",
  "item_id": "item_DFaqc54nLoQFvJKRdERwM",
  "output_index": 0,
  "content_index": 0,
  "part": {
    "type": "text",
    "text": "Absolutely! I'm an AI designed to assist with a variety of tasks, from answering questions to providing information, and everything in between. I can chat in multiple languages, share facts, and even tell jokes. How can I help you today?"
  }
}

[event] response.content_part.done

[payload]
{
  "type": "response.output_item.done",
  "event_id": "event_DFaqcHeYoIHajTtOctZ1o",
  "response_id": "resp_DFaqcGe6BRKhyoBUQCs7A",
  "output_index": 0,
  "item": {
    "id": "item_DFaqc54nLoQFvJKRdERwM",
    "object": "realtime.item",
    "type": "message",
    "status": "completed",
    "role": "assistant",
    "content": [
      {
        "type": "text",
        "text": "Absolutely! I'm an AI designed to assist with a variety of tasks, from answering questions to providing information, and everything in between. I can chat in multiple languages, share facts, and even tell jokes. How can I help you today?"
      }
    ]
  }
}

[event] response.output_item.done

[payload]
{
  "type": "response.done",
  "event_id": "event_DFaqcQJ4VlwUd913mwdT4",
  "response": {
    "object": "realtime.response",
    "id": "resp_DFaqcGe6BRKhyoBUQCs7A",
    "status": "completed",
    "status_details": null,
    "output": [
      {
        "id": "item_DFaqc54nLoQFvJKRdERwM",
        "object": "realtime.item",
        "type": "message",
        "status": "completed",
        "role": "assistant",
        "content": [
          {
            "type": "text",
            "text": "Absolutely! I'm an AI designed to assist with a variety of tasks, from answering questions to providing information, and everything in between. I can chat in multiple languages, share facts, and even tell jokes. How can I help you today?"
          }
        ]
      }
    ],
    "conversation_id": "conv_DFaqcUFrtDoD9ZyI8k0Uf",
    "modalities": [
      "text"
    ],
    "voice": "alloy",
    "output_audio_format": "pcm16",
    "temperature": 0.8,
    "max_output_tokens": "inf",
    "usage": {
      "total_tokens": 181,
      "input_tokens": 131,
      "output_tokens": 50,
      "input_token_details": {
        "text_tokens": 131,
        "audio_tokens": 0,
        "image_tokens": 0,
        "cached_tokens": 0,
        "cached_tokens_details": {
          "text_tokens": 0,
          "audio_tokens": 0,
          "image_tokens": 0
        }
      },
      "output_token_details": {
        "text_tokens": 50,
        "audio_tokens": 0
      }
    },
    "metadata": null
  }
}

---
Response complete.
Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated

Summary by CodeRabbit

  • New Features

    • WebSocket realtime endpoint at /v1/realtime for bidirectional model interactions
    • Model selection via query parameter; connections require valid upstream authorization
  • Behavior Changes

    • Session registry no longer enforces fixed capacity limits; Connected sessions persist
    • Reaper cleanup now ages Pending and Disconnected entries differently; Connected entries are retained
    • Enhanced TLS-backed WebSocket proxying and improved session lifecycle handling

@github-actions github-actions Bot added dependencies Dependency updates model-gateway Model gateway crate changes realtime-api Realtime API related changes openai OpenAI router changes labels Mar 5, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly enhances the Realtime API gateway by introducing a WebSocket transport layer. Previously, the gateway only supported REST endpoints for token generation, lacking the necessary bidirectional streaming for interactive realtime sessions. The new implementation provides a dedicated WebSocket endpoint that upgrades client connections, intelligently routes them to appropriate upstream workers, and transparently proxies messages, thereby enabling full server-to-server realtime communication.

Highlights

  • Realtime WebSocket Handler: Introduced a new WebSocket upgrade handler at GET /v1/realtime to enable server-to-server bidirectional streaming for realtime audio/text sessions.
  • Bidirectional WebSocket Proxy: Implemented a robust bidirectional forwarding engine that manages client and upstream WebSocket streams, coordinating message flow and handling session lifecycle events with structured logging.
  • Dependency Updates: Added tokio-tungstenite and webpki-roots dependencies to support WebSocket communication and explicit TLS connector building.
  • Session Management Integration: Integrated the new WebSocket route into the server, reusing the existing RealtimeRegistry for session tracking and incorporating authentication and worker-selection logic. A session reaper was also started to manage session lifetimes.
Changelog
  • Cargo.toml
    • Added webpki-roots and multer dependencies.
  • model_gateway/Cargo.toml
    • Added tokio-tungstenite and webpki-roots workspace dependencies.
  • model_gateway/src/routers/openai/realtime/mod.rs
    • Exported new proxy and ws modules to expose WebSocket-related functionalities.
  • model_gateway/src/routers/openai/realtime/proxy.rs
    • Implemented the core logic for bidirectional WebSocket message forwarding between client and upstream services.
  • model_gateway/src/routers/openai/realtime/ws.rs
    • Created the Axum WebSocket upgrade handler responsible for validating requests, selecting workers, and initiating the proxy.
  • model_gateway/src/server.rs
    • Wired the new GET /v1/realtime route into the application's router.
    • Started a background task for the realtime session reaper to clean up expired sessions.
Activity
  • The author provided a detailed test plan demonstrating successful WebSocket connection, session creation, updates, and streaming of responses, including text output from an AI model.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@coderabbitai

coderabbitai Bot commented Mar 5, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

Adds workspace deps and WebSocket realtime support: new bidirectional TLS WebSocket proxy, a /v1/realtime WebSocket handler that registers sessions with a realtime registry, registry reaper integration in server startup, and registry API changes (dynamic sizing, get/set state).

Changes

Cohort / File(s) Summary
Workspace Dependencies
Cargo.toml, model_gateway/Cargo.toml
Added workspace deps: webpki-roots, multer; added crate-scoped tokio-tungstenite and webpki-roots in model_gateway/Cargo.toml.
Realtime Module Surface
model_gateway/src/routers/openai/realtime/mod.rs
Exposed new public submodules proxy and ws.
WebSocket Proxy
model_gateway/src/routers/openai/realtime/proxy.rs
New bidirectional proxy: establishes TLS upstream connection (explicit rustls + webpki roots), spawns paired forward tasks with CancellationToken, updates session state, and implements TLS connector builder.
WebSocket Handler
model_gateway/src/routers/openai/realtime/ws.rs
New ws_handler and RealtimeQueryParams: selects worker, validates auth, registers session (gets cancel token), upgrades client to WebSocket, and delegates to proxy for message forwarding.
Registry changes
model_gateway/src/routers/openai/realtime/registry.rs
Removed capacity reservation and atomics; register_session/register_call now return entries directly; added get_*/set_* helpers; reaper accepts pending_max_age and uses state-aware staleness rules; counts derived from container sizes.
Server integration
model_gateway/src/server.rs
Starts realtime registry reaper with explicit TTLs and registers /v1/realtime route to the new ws handler.

Sequence Diagram

sequenceDiagram
    participant Client
    participant WsHandler as "ws_handler (HTTP)"
    participant Registry as "RealtimeRegistry"
    participant Proxy as "WebSocket Proxy"
    participant Upstream as "Upstream Realtime"

    Client->>WsHandler: GET /v1/realtime?model=...
    WsHandler->>WsHandler: Authenticate & select worker
    WsHandler->>Registry: Register session -> cancel_token
    WsHandler->>WsHandler: Build upstream WebSocket URL
    WsHandler->>Client: Upgrade to WebSocket
    WsHandler->>Proxy: hand off WebSocket + session info

    Proxy->>Upstream: Establish TLS WebSocket connection
    Proxy->>Registry: set session state = Connected

    par Bidirectional forwarding
        loop client -> upstream
            Client->>Proxy: WS message (Text/Binary/Ping/Close)
            Proxy->>Proxy: parse/log ClientEvent
            Proxy->>Upstream: forward message
        end
    and
        loop upstream -> client
            Upstream->>Proxy: WS message
            Proxy->>Proxy: parse/log ServerEvent
            Proxy->>Client: forward message
        end
    end

    Client-->>Proxy: Close or cancel_token triggered
    Proxy->>Upstream: Close upstream connection
    Proxy->>Registry: set session state = Disconnected
Loading

Estimated Code Review Effort

🎯 4 (Complex) | ⏱️ ~50 minutes

Possibly Related Issues

Possibly Related PRs

Suggested Reviewers

  • CatherineSue
  • key4ng
  • slin1237

Poem

🐰 I hopped a tiny TLS stream,

WebSockets hummed like a dream,
Messages danced both ways with glee,
Sessions kept and reaped for me,
Hooray for realtime — hop, whee! 🎉

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'feat(realtime-api): realtime websocket handler' clearly and concisely summarizes the main change: adding a WebSocket handler for the realtime API, which aligns with the primary objective of implementing the realtime WebSocket proxy handler.
Linked Issues check ✅ Passed The PR implements all key objectives from issue #245: ws_handler for model validation and worker selection [#245], proxy.rs with bidirectional forwarding via forward_client_to_upstream and forward_upstream_to_client tasks [#245], registry modifications enabling session lifecycle management [#245], and proper error handling and logging throughout [#245].
Out of Scope Changes check ✅ Passed All changes are within scope: workspace dependencies (webpki-roots, multer, tokio-tungstenite) directly support TLS and WebSocket functionality [#245]; registry refactoring (capacity removal, added get/set methods) enables session state management [#245]; server.rs wiring and mod.rs exports complete the WebSocket integration [#245].
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
  • 📝 Generate docstrings (stacked PR)
  • 📝 Generate docstrings (commit on current branch)
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch yifeliu/realtime-websocket-handler

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a WebSocket handler for the OpenAI Realtime API, utilizing good patterns for stream management and session handling. However, a critical Resource Exhaustion (DoS) vulnerability was identified: sessions are registered before the WebSocket upgrade is completed, potentially filling the registry's capacity with stale 'Pending' sessions. Beyond this, consider optimizing performance by avoiding unnecessary data clones in the message forwarding hot path, and enhancing robustness through URL encoding for query parameters and improved authentication header validation.

Comment thread model_gateway/src/routers/openai/realtime/ws.rs
Comment thread model_gateway/src/routers/openai/realtime/ws.rs Outdated
Comment thread model_gateway/src/routers/openai/realtime/proxy.rs
Comment thread model_gateway/src/routers/openai/realtime/proxy.rs
Comment thread model_gateway/src/routers/openai/realtime/ws.rs

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 3426ebe914

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread model_gateway/src/routers/openai/realtime/ws.rs
Comment thread model_gateway/src/routers/openai/realtime/proxy.rs

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 5

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@model_gateway/src/routers/openai/realtime/proxy.rs`:
- Around line 66-103: When one forwarder join handle finishes in the
tokio::select! (client_to_upstream or upstream_to_client), explicitly abort the
sibling join handle (call .abort() on the other JoinHandle), then await that
handle to ensure it has terminated and log any JoinError; do this after
signalling cancel_token.cancel() so both forward_client_to_upstream and
forward_upstream_to_client are deterministically stopped and cleaned up instead
of being dropped and left running.
- Around line 47-49: Wrap the call to
tokio_tungstenite::connect_async_tls_with_config(request, None, false,
Some(connector)).await in a tokio::time::timeout(...) with the standard
connection timeout duration used elsewhere, and handle a timeout by mapping the
timeout error into the same error type via .map_err(...) so the dial failure
returns a clear timeout error instead of hanging; keep the existing variable
names (request, connector, upstream_ws) and ensure the map_err path produces the
same error shape used by the surrounding function.

In `@model_gateway/src/routers/openai/realtime/ws.rs`:
- Around line 115-117: The build_upstream_ws_url function currently interpolates
model directly into the query string which allows reserved chars to break URL
semantics; update build_upstream_ws_url(worker_url: &str, model: &str) to
percent-encode the model value before insertion (e.g., using the
percent-encoding or url crate / Url::parse_with_params) and then format the
final URL from the trimmed base and the encoded model so characters like &, =, #
are safely escaped in the query component.
- Around line 63-66: The code currently coerces non-UTF8 Authorization headers
into an empty string (auth_str) which then becomes a valid empty header
upstream; instead, in the websocket handler where extract_auth_header(...) and
auth_header_value are used, treat a to_str() failure as an invalid header and
reject the request with a 401 rather than converting to "". Replace the
unwrap_or("") behavior for auth_header_value -> auth_str so that when
HeaderValue::to_str() returns Err you return an immediate 401 Unauthorized
(mirroring rest.rs behavior) and only forward a valid UTF-8 auth_str upstream;
ensure the error path references extract_auth_header, auth_header_value,
auth_str and worker.api_key() to locate the change.

In `@model_gateway/src/server.rs`:
- Around line 566-570: start_reaper currently spawns an unguarded background
task (called from realtime_registry.start_reaper) every time build_app/test
helpers run, causing duplicate reapers and losing the CancellationToken; make it
idempotent by adding an internal guard or stored token: either use a Once or
OnceLock inside realtime_registry to run start_reaper only once, or have
start_reaper store and return a CancellationToken field on the registry (check
if it's already Some and return early) so repeated calls do nothing and the
token can be used for cleanup; update callers (build_app/test helpers) to rely
on the registry’s idempotent behavior instead of spawning new tasks.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: ca462f6e-d123-425b-9071-15dc3255c104

📥 Commits

Reviewing files that changed from the base of the PR and between 636fe1d and 3426ebe.

📒 Files selected for processing (6)
  • Cargo.toml
  • model_gateway/Cargo.toml
  • model_gateway/src/routers/openai/realtime/mod.rs
  • model_gateway/src/routers/openai/realtime/proxy.rs
  • model_gateway/src/routers/openai/realtime/ws.rs
  • model_gateway/src/server.rs

Comment thread model_gateway/src/routers/openai/realtime/proxy.rs Outdated
Comment thread model_gateway/src/routers/openai/realtime/proxy.rs
Comment thread model_gateway/src/routers/openai/realtime/ws.rs
Comment thread model_gateway/src/routers/openai/realtime/ws.rs Outdated
Comment thread model_gateway/src/server.rs Outdated
@pallasathena92
pallasathena92 force-pushed the yifeliu/realtime-websocket-handler branch from 3426ebe to 1586b8a Compare March 5, 2026 03:07

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

♻️ Duplicate comments (1)
model_gateway/src/routers/openai/realtime/ws.rs (1)

119-121: ⚠️ Potential issue | 🟡 Minor

URL-encode model before interpolating into the query string.

Model names containing reserved characters (&, =, #, ?, spaces) would corrupt the URL. The url crate is already available (transitive dependency).

🔧 Proposed fix
 fn build_upstream_ws_url(worker_url: &str, model: &str) -> String {
     let base = worker_url.trim_end_matches('/');
-    format!("{base}/v1/realtime?model={model}")
+    let encoded_model: String =
+        url::form_urlencoded::byte_serialize(model.as_bytes()).collect();
+    format!("{base}/v1/realtime?model={encoded_model}")
 }
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@model_gateway/src/routers/openai/realtime/ws.rs` around lines 119 - 121, The
build_upstream_ws_url function currently interpolates model raw into the query
string; percent-encode the model value before formatting to avoid breaking URLs
when it contains reserved characters. Use the url crate (e.g.,
url::form_urlencoded::byte_serialize or
url::percent_encoding::utf8_percent_encode) to encode the model string, then
call format!("{base}/v1/realtime?model={encoded_model}") in
build_upstream_ws_url so the query parameter is safely encoded.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Duplicate comments:
In `@model_gateway/src/routers/openai/realtime/ws.rs`:
- Around line 119-121: The build_upstream_ws_url function currently interpolates
model raw into the query string; percent-encode the model value before
formatting to avoid breaking URLs when it contains reserved characters. Use the
url crate (e.g., url::form_urlencoded::byte_serialize or
url::percent_encoding::utf8_percent_encode) to encode the model string, then
call format!("{base}/v1/realtime?model={encoded_model}") in
build_upstream_ws_url so the query parameter is safely encoded.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 8d073e4a-4d2f-4f64-b039-7a6aa9502f47

📥 Commits

Reviewing files that changed from the base of the PR and between 3426ebe and 1586b8a.

📒 Files selected for processing (7)
  • Cargo.toml
  • model_gateway/Cargo.toml
  • model_gateway/src/routers/openai/realtime/mod.rs
  • model_gateway/src/routers/openai/realtime/proxy.rs
  • model_gateway/src/routers/openai/realtime/registry.rs
  • model_gateway/src/routers/openai/realtime/ws.rs
  • model_gateway/src/server.rs

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 1586b8a606

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment on lines +104 to +105
error!(session_id, error = %e, "Realtime WebSocket proxy error");
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Record WebSocket proxy failures on the selected worker

ws_handler uses select_worker, which filters by circuit_breaker().can_execute(), but this error branch only logs and never calls worker.record_outcome(false). In a deployment where a worker’s realtime endpoint is failing (bad upstream config, repeated handshake/runtime errors), /v1/realtime traffic will keep selecting that worker because realtime failures never feed into the circuit breaker, causing persistent user-facing connection failures instead of tripping and shedding to healthier workers.

Useful? React with 👍 / 👎.

Comment thread model_gateway/src/routers/openai/realtime/registry.rs
Signed-off-by: yifeliu <yifengliu9@gmail.com>
@pallasathena92
pallasathena92 force-pushed the yifeliu/realtime-websocket-handler branch from 1586b8a to cd80e62 Compare March 5, 2026 03:22

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@model_gateway/src/routers/openai/realtime/proxy.rs`:
- Around line 87-106: The select block drops the surviving task's JoinHandle
(client_to_upstream or upstream_to_client) after calling cancel_token.cancel(),
detaching it instead of awaiting completion; update the logic so after detecting
one branch finished you call await on the other task's JoinHandle (the surviving
client_to_upstream or upstream_to_client handle) to ensure deterministic cleanup
once cancel_token.cancel() is issued, while still keeping the
cancel_token.cancelled() branch; locate the select around those symbols and add
a follow-up await for the other handle (handling/ignoring its Result) after the
select completes to avoid leaving the task detached.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: fdd747dc-8df0-4ebd-bfcb-af79480bc36b

📥 Commits

Reviewing files that changed from the base of the PR and between 1586b8a and cd80e62.

📒 Files selected for processing (7)
  • Cargo.toml
  • model_gateway/Cargo.toml
  • model_gateway/src/routers/openai/realtime/mod.rs
  • model_gateway/src/routers/openai/realtime/proxy.rs
  • model_gateway/src/routers/openai/realtime/registry.rs
  • model_gateway/src/routers/openai/realtime/ws.rs
  • model_gateway/src/server.rs

Comment on lines +87 to +106
// Wait for either task to finish (or cancellation)
tokio::select! {
result = client_to_upstream => {
cancel_token.cancel();
debug!(session_id, "Client→upstream task ended");
if let Err(e) = result {
error!(session_id, error = %e, "Client→upstream task panicked");
}
}
result = upstream_to_client => {
cancel_token.cancel();
debug!(session_id, "Upstream→client task ended");
if let Err(e) = result {
error!(session_id, error = %e, "Upstream→client task panicked");
}
}
() = cancel_token.cancelled() => {
debug!(session_id, "Session cancelled via token");
}
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick | 🔵 Trivial

Consider awaiting the surviving task for deterministic cleanup.

When one forwarding task completes, the other's JoinHandle is dropped without awaiting. While cancel_token.cancel() signals the surviving task to exit, dropping the handle detaches it rather than ensuring completion.

Per the past discussion, the author notes that the surviving task exits within one event loop tick via its cancellation branch, making this acceptable in practice. If you want structural guarantees, you could await the surviving handle after cancellation:

// After the select! block, in each branch:
let _ = other_handle.await;

This is low-severity given the task's short lifetime and fast cancellation response.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@model_gateway/src/routers/openai/realtime/proxy.rs` around lines 87 - 106,
The select block drops the surviving task's JoinHandle (client_to_upstream or
upstream_to_client) after calling cancel_token.cancel(), detaching it instead of
awaiting completion; update the logic so after detecting one branch finished you
call await on the other task's JoinHandle (the surviving client_to_upstream or
upstream_to_client handle) to ensure deterministic cleanup once
cancel_token.cancel() is issued, while still keeping the
cancel_token.cancelled() branch; locate the select around those symbols and add
a follow-up await for the other handle (handling/ignoring its Result) after the
select completes to avoid leaving the task detached.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: cd80e62da4

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

));

let realtime_routes = Router::new()
.route("/v1/realtime", get(realtime_ws::ws_handler))

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Add connection-lifetime limiting for /v1/realtime

Mounting GET /v1/realtime inside realtime_routes means it only gets the existing concurrency_limit_middleware, but that middleware accounts tokens against the HTTP response body lifetime (TokenGuardBody) rather than the upgraded socket lifetime. For WebSocket upgrades, the 101 response body is dropped immediately, so the token is released while the session remains open; in environments that rely on this limiter to cap concurrent load, clients can open many long-lived realtime sockets and bypass the configured concurrency guard.

Useful? React with 👍 / 👎.

@slin1237
slin1237 merged commit 1190b5a into main Mar 5, 2026
24 of 27 checks passed
@slin1237
slin1237 deleted the yifeliu/realtime-websocket-handler branch March 5, 2026 14:46
slin1237 added a commit that referenced this pull request Mar 5, 2026
proxy.rs:
- Replace full ClientEvent/ServerEvent deserialization with lightweight
  EventTypeOnly struct for event logging. Avoids parsing large boxed
  variants (SessionConfig 624B, ResponseCreateParams 384B) on every
  audio frame in the hot path.
- Cache TLS ClientConfig via OnceLock instead of rebuilding it
  (including cloning all webpki root certificates) per connection.
- Remove unused openai_protocol::realtime_events import.

registry.rs:
- Unify SessionEntry and CallEntry (identical except for id field name)
  into a single ConnectionEntry type.
- Extract shared CRUD and reap logic into ConnectionMap, eliminating
  ~30 lines of duplicated code across session/call methods and the
  reaper task.

Refs: #637
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
slin1237 added a commit that referenced this pull request Mar 6, 2026
proxy.rs:
- Use Cow<str> instead of &str for EventTypeOnly::event_type so serde
  can handle JSON escape sequences (e.g. \u002E) that require allocation,
  while preserving zero-copy for the common unescaped case.
- Change "Safety:" comment to "INVARIANT:" per repo convention.

registry.rs:
- Replace two-pass collect-then-remove in reap_stale with single-pass
  DashMap::retain, eliminating intermediate Vec allocation.

Refs: #637
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
slin1237 added a commit that referenced this pull request Mar 7, 2026
proxy.rs:
- Replace full ClientEvent/ServerEvent deserialization with lightweight
  EventTypeOnly struct for event logging. Avoids parsing large boxed
  variants (SessionConfig 624B, ResponseCreateParams 384B) on every
  audio frame in the hot path.
- Cache TLS ClientConfig via OnceLock instead of rebuilding it
  (including cloning all webpki root certificates) per connection.
- Remove unused openai_protocol::realtime_events import.

registry.rs:
- Unify SessionEntry and CallEntry (identical except for id field name)
  into a single ConnectionEntry type.
- Extract shared CRUD and reap logic into ConnectionMap, eliminating
  ~30 lines of duplicated code across session/call methods and the
  reaper task.

Refs: #637
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
slin1237 added a commit that referenced this pull request Mar 7, 2026
proxy.rs:
- Use Cow<str> instead of &str for EventTypeOnly::event_type so serde
  can handle JSON escape sequences (e.g. \u002E) that require allocation,
  while preserving zero-copy for the common unescaped case.
- Change "Safety:" comment to "INVARIANT:" per repo convention.

registry.rs:
- Replace two-pass collect-then-remove in reap_stale with single-pass
  DashMap::retain, eliminating intermediate Vec allocation.

Refs: #637
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

dependencies Dependency updates model-gateway Model gateway crate changes openai OpenAI router changes realtime-api Realtime API related changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Realtime Api] gateway-realtime-websocket-handler

2 participants