Skip to content

fix(realtime-api): worker health tracking in websocket session - #725

Merged
slin1237 merged 1 commit into
mainfrom
yifeliu/realtime-ws-cleanup
Mar 11, 2026
Merged

slin1237 merged 1 commit into
mainfrom
yifeliu/realtime-ws-cleanup

Conversation

@pallasathena92

@pallasathena92 pallasathena92 commented Mar 11, 2026 •

Copy link
Copy Markdown
Collaborator

Ref: #637

Description

Problem

The realtime WebSocket handler does not call worker.record_outcome() after the proxy session completes. This means worker health tracking has no signal from realtime WebSocket sessions — unlike the REST path which already records outcomes. Without this, the load balancer cannot distinguish healthy workers from unhealthy ones based on realtime session results.

Solution

Capture the proxy result as a success boolean and call worker.record_outcome(success) before cleaning up the session. This brings the WS handler in line with existing patterns in the REST path.

Changes

  • model_gateway/src/routers/openai/realtime/ws.rs
    • Refactor if let Err(e) = run_ws_proxy(...) into match that captures a success: bool
    • Call worker.record_outcome(success) after the proxy completes
    • Import proxy module directly instead of proxy::run_ws_proxy (aligns with module-level import convention)
    • Use explicit tracing::{debug, error} imports instead of tracing::error! / tracing::debug! via path

Test Plan

Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated

Summary by CodeRabbit

  • Refactor

    • Improved error handling and logging for real-time WebSocket connections with enhanced session tracking and outcome recording.
  • Chores

    • Updated error handling semantics and adjusted logging levels for better diagnostics and cleanup operations.

Signed-off-by: yifeliu <yifengliu9@gmail.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Repo admins can enable using credits for code reviews in their settings.

@github-actions github-actions Bot added model-gateway Model gateway crate changes realtime-api Realtime API related changes openai OpenAI router changes labels Mar 11, 2026
@coderabbitai

coderabbitai Bot commented Mar 11, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

Modified the OpenAI realtime WebSocket handler to replace direct error handling with a guard-based approach that returns a success indicator, enabling outcome recording via worker.record_outcome() and restructuring error logging semantics.

Changes

Cohort / File(s) Summary
OpenAI Realtime WebSocket Handler
model_gateway/src/routers/openai/realtime/ws.rs
Refactored proxy invocation from if let Err(e) to a match pattern; added boolean success tracking; introduced worker.record_outcome(success) call; adjusted logging from tracing::error! to error! for failures and changed cleanup log to debug! level; updated imports to reference proxy as a module.

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~10 minutes

Possibly related issues

  • Issue #245: Directly addresses changes to the same WebSocket handler implementation and proxy invocation pattern in model_gateway/src/routers/openai/realtime/ws.rs.

Possibly related PRs

  • PR #476: Modifies usage patterns of worker.record_outcome() with similar outcome recording logic.
  • PR #637: Touches the same realtime WebSocket handler file and adjusts proxy call behavior and error handling semantics.

Suggested labels

model-gateway, realtime-api, openai

Suggested reviewers

  • CatherineSue
  • key4ng
  • slin1237

Poem

🐰 A whisker-twitch of proxy calls,
Guard gates the errors, success enthralls,
The outcome recorded, so neat and clean,
Logs now whisper what once was seen,
Debug hops softly through session halls! 🌙

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Title check ✅ Passed The title directly and accurately summarizes the main change: adding worker health tracking to the WebSocket session handler, which is the core objective of the PR.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
  • 📝 Generate docstrings (stacked PR)
  • 📝 Generate docstrings (commit on current branch)
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch yifeliu/realtime-ws-cleanup

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request addresses a critical gap in worker health tracking for real-time WebSocket sessions. Previously, the system lacked a mechanism to report the outcome of these sessions, preventing the load balancer from accurately assessing worker health. The changes introduce a robust way to capture the success or failure of WebSocket proxy operations and record these outcomes, bringing the real-time path in line with existing patterns in the REST API and ensuring more reliable load balancing.

Highlights

  • Worker Health Tracking: Implemented worker.record_outcome() for real-time WebSocket sessions, ensuring the load balancer can now accurately track worker health based on these sessions.
  • Proxy Result Capture: Refactored the WebSocket proxy call from an if let Err statement to a match expression to explicitly capture and utilize the success or failure outcome of the proxy operation.
  • Tracing Imports: Updated tracing macro usage to explicitly import debug and error from the tracing crate, aligning with module-level import conventions.
  • Module Import Consistency: Changed the proxy module import from a specific function (run_ws_proxy) to a module-level import for improved consistency.
Changelog
  • model_gateway/src/routers/openai/realtime/ws.rs
    • Refactored the run_ws_proxy call from an if let Err statement to a match expression to capture the success boolean.
    • Added a call to worker.record_outcome(success) after the WebSocket proxy completes to track worker health.
    • Updated tracing imports to explicitly use debug and error macros.
    • Modified the proxy module import to be module-level instead of importing a specific function.
Activity
  • No human activity (comments, reviews, etc.) has been recorded on this pull request yet.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

The pull request refactors the real-time WebSocket proxy handling to explicitly capture and record the success or failure of the proxy operation. This involves updating tracing imports, adjusting the proxy module import, and introducing a worker.record_outcome call to track the session's outcome before cleanup.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (2)
model_gateway/src/routers/openai/realtime/ws.rs (2)

120-137: ⚠️ Potential issue | 🟠 Major

success still over-reports healthy WS sessions.

Lines 120-137 treat any Ok(()) from proxy::run_ws_proxy as success, but model_gateway/src/routers/openai/realtime/proxy.rs currently returns Ok(()) for all post-connect exits and only logs task failures before falling through. That means a session can terminate abnormally after connect and still call worker.record_outcome(true), so the health signal remains misleading. Please propagate a richer proxy outcome here, or make abnormal forwarding/session termination return Err instead of Ok(()).

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@model_gateway/src/routers/openai/realtime/ws.rs` around lines 120 - 137, The
code treats any Ok(()) from proxy::run_ws_proxy as a healthy session, but
proxy::run_ws_proxy currently returns Ok(()) for abnormal post-connect exits,
causing worker.record_outcome(true) to over-report health; change the proxy API
or caller to propagate a richer outcome and only count true for genuinely
successful sessions. Specifically, update proxy::run_ws_proxy to return a
Result<ProxyOutcome, E> (or Result<bool, E>) that distinguishes normal
completion vs abnormal termination, modify the match here to treat abnormal
outcomes as Err/false (e.g., match proxy::run_ws_proxy(...) .await {
Ok(ProxyOutcome::Completed) => true, Ok(ProxyOutcome::Aborted) | Err(_) => {
error!(...); false } }), and then pass that boolean into worker.record_outcome
so only true reflects a truly healthy forwarded session; ensure references to
proxy::run_ws_proxy and worker.record_outcome are updated accordingly.

119-139: 🧹 Nitpick | 🔵 Trivial

Add a regression test for the new health accounting.

This branch changes worker health state, but there is no coverage here to lock down the true/false mapping. A focused test with a fake Worker that asserts record_outcome(true) on a clean proxy completion and record_outcome(false) on proxy failure would make this much safer to maintain.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@model_gateway/src/routers/openai/realtime/ws.rs` around lines 119 - 139, Add
a regression test that asserts the worker health accounting mapping by creating
a fake/mock Worker and exercising the ws upgrade path: call the code that
invokes proxy::run_ws_proxy (or directly invoke the closure used in
ws.on_upgrade) with two scenarios—one where proxy::run_ws_proxy returns Ok(())
and one where it returns Err(...). For each scenario, verify
FakeWorker.record_outcome(true) is called on the clean completion case and
FakeWorker.record_outcome(false) is called on the failure case; use the same
symbols from the diff (proxy::run_ws_proxy, Worker.record_outcome,
realtime_registry.remove_session, session_id) to locate and wire the test, and
inject/mocking run_ws_proxy or the WebSocket input so the test controls success
vs failure deterministically. Ensure the test runs under tokio async context and
cleans up realtime_registry entries as the production code does.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Outside diff comments:
In `@model_gateway/src/routers/openai/realtime/ws.rs`:
- Around line 120-137: The code treats any Ok(()) from proxy::run_ws_proxy as a
healthy session, but proxy::run_ws_proxy currently returns Ok(()) for abnormal
post-connect exits, causing worker.record_outcome(true) to over-report health;
change the proxy API or caller to propagate a richer outcome and only count true
for genuinely successful sessions. Specifically, update proxy::run_ws_proxy to
return a Result<ProxyOutcome, E> (or Result<bool, E>) that distinguishes normal
completion vs abnormal termination, modify the match here to treat abnormal
outcomes as Err/false (e.g., match proxy::run_ws_proxy(...) .await {
Ok(ProxyOutcome::Completed) => true, Ok(ProxyOutcome::Aborted) | Err(_) => {
error!(...); false } }), and then pass that boolean into worker.record_outcome
so only true reflects a truly healthy forwarded session; ensure references to
proxy::run_ws_proxy and worker.record_outcome are updated accordingly.
- Around line 119-139: Add a regression test that asserts the worker health
accounting mapping by creating a fake/mock Worker and exercising the ws upgrade
path: call the code that invokes proxy::run_ws_proxy (or directly invoke the
closure used in ws.on_upgrade) with two scenarios—one where proxy::run_ws_proxy
returns Ok(()) and one where it returns Err(...). For each scenario, verify
FakeWorker.record_outcome(true) is called on the clean completion case and
FakeWorker.record_outcome(false) is called on the failure case; use the same
symbols from the diff (proxy::run_ws_proxy, Worker.record_outcome,
realtime_registry.remove_session, session_id) to locate and wire the test, and
inject/mocking run_ws_proxy or the WebSocket input so the test controls success
vs failure deterministically. Ensure the test runs under tokio async context and
cleans up realtime_registry entries as the production code does.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: e5b47a4a-bd70-4c4d-8603-3a6c6916afac

📥 Commits

Reviewing files that changed from the base of the PR and between 79dc7e4 and e5430e2.

📒 Files selected for processing (1)
  • model_gateway/src/routers/openai/realtime/ws.rs

@slin1237
slin1237 merged commit cae4c4d into main Mar 11, 2026
62 of 65 checks passed
@slin1237
slin1237 deleted the yifeliu/realtime-ws-cleanup branch March 11, 2026 18:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

model-gateway Model gateway crate changes openai OpenAI router changes realtime-api Realtime API related changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants