Skip to content

refactor(gateway): split OpenAI router.rs into chat and health modules - #726

Merged
slin1237 merged 2 commits into
mainfrom
slin/oai-refactor-3
Mar 11, 2026
Merged

slin1237 merged 2 commits into
mainfrom
slin/oai-refactor-3

Conversation

@slin1237

@slin1237 slin1237 commented Mar 11, 2026 •

Copy link
Copy Markdown
Member

Summary

  • Extract route_chat logic (~250 lines) from the monolithic router.rs (~988 lines) into openai/chat.rs as a free function with a RouterContext struct for shared infrastructure
  • Extract health_generate and get_server_info (~60 lines) into openai/health.rs
  • router.rs becomes a thinner dispatcher (~680 lines), with resolve_provider extracted as a pub(super) free function shared by chat.rs and route_responses

Refs: #721

What changed

File Action
model_gateway/src/routers/openai/chat.rs New — RouterContext struct + route_chat free function with full chat completion routing (worker selection, provider transform, retry loop, streaming/non-streaming)
model_gateway/src/routers/openai/health.rs New — health_generate + get_server_info free functions with external_workers helper inlined
model_gateway/src/routers/openai/mod.rs Added mod chat; and mod health;
model_gateway/src/routers/openai/router.rs Removed moved bodies, RouterTrait methods now delegate to extracted modules. Extracted resolve_provider as pub(super). Removed unused shared_components() accessor.

Why

router.rs mixed chat dispatch, health endpoints, responses orchestration, storage queries, and realtime methods in a single ~988-line file. Splitting improves navigability and sets up further extractions (responses in a follow-up PR, realtime after that).

How

Uses free functions with a RouterContext struct rather than splitting impl blocks across files, since OpenAIRouter fields are private. This matches the convention used in the Anthropic router (RouterContext = shared infrastructure, RequestContext = per-request input) and Gemini router (streaming::execute(&router_ctx, req_ctx)).

Pure refactor — no behavior changes.

Test plan

  • cargo build -p smg — clean, no warnings
  • cargo clippy -p smg --all-targets -- -D warnings — clean
  • cargo test -p smg --lib — 409 tests pass

Summary by CodeRabbit

  • New Features
    • Chat completion routing that forwards requests to upstream workers, with full streaming and non‑streaming support and retry/circuit‑breaker handling.
    • Health monitoring for backend workers, reporting overall availability and unhealthy worker details.
    • Server status endpoint exposing router metrics and worker/service summary for diagnostics.

Extract route_chat and health/server-info endpoints from the monolithic
router.rs (~988 lines) into focused modules, reducing it to a thin
dispatcher (~680 lines).

What changed:
- model_gateway/src/routers/openai/chat.rs: new module containing
  ChatDeps struct and route_chat free function, with the full chat
  completion routing logic (worker selection, provider transform,
  retry loop with streaming/non-streaming handling)
- model_gateway/src/routers/openai/health.rs: new module containing
  health_generate and get_server_info free functions, with the
  external_workers helper inlined
- model_gateway/src/routers/openai/mod.rs: added mod chat and mod health
- model_gateway/src/routers/openai/router.rs: removed route_chat body
  (~250 lines) and health/server-info bodies (~60 lines), RouterTrait
  methods now delegate to the extracted modules. Extracted
  resolve_provider as pub(super) free function shared by chat.rs and
  route_responses. Removed unused shared_components() accessor.

Why:
router.rs mixed chat dispatch, health endpoints, responses orchestration,
storage queries, and realtime methods. Splitting improves navigability
and sets up further extractions (responses in PR 4, realtime in PR 5).

How:
Uses free functions with dependency structs (ChatDeps) rather than
splitting impl blocks across files, since OpenAIRouter fields are
private. This matches the pattern used in Anthropic/Gemini routers
(e.g. streaming::execute(&router_ctx, req_ctx)).

Pure refactor — no behavior changes.

Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Repo admins can enable using credits for code reviews in their settings.

@github-actions github-actions Bot added model-gateway Model gateway crate changes openai OpenAI router changes labels Mar 11, 2026
@coderabbitai

coderabbitai Bot commented Mar 11, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

Adds a new OpenAI chat routing module that forwards chat completion requests to upstream workers with retries, provider-specific transformations, streaming and non-streaming response handling, extensive error metrics, and a health module that reports worker/registry status. Router refactoring delegates provider resolution and chat/health flows to new modules.

Changes

Cohort / File(s) Summary
Chat routing
model_gateway/src/routers/openai/chat.rs
New route_chat() and RouterContext implemented. Selects worker, resolves provider, transforms payload, builds upstream request, executes retryable requests, handles streaming vs non-streaming responses, updates circuit breaker and metrics, and returns proxied responses.
Health endpoints
model_gateway/src/routers/openai/health.rs
New health helpers: health_generate() and get_server_info() that enumerate external workers, partition healthy/unhealthy, return service_unavailable or OK, and expose JSON server info.
Module wiring
model_gateway/src/routers/openai/mod.rs
Added mod chat; and mod health; to include new submodules.
Router refactor / provider resolution
model_gateway/src/routers/openai/router.rs
Introduces resolve_provider() helper and delegates chat routing and health/server-info logic to the new modules; removes large inlined logic from router.

Sequence Diagram

sequenceDiagram
    participant Client as Client
    participant Router as Router (route_chat)
    participant Selector as WorkerSelector
    participant Provider as Provider
    participant Worker as Upstream Worker
    participant Retry as RetryExecutor

    Client->>Router: ChatCompletionRequest
    Router->>Router: record incoming metrics
    Router->>Selector: select worker
    Selector-->>Router: selected worker
    Router->>Provider: resolve & apply transformations
    Provider-->>Router: transformed payload
    Router->>Retry: prepare RequestContext & start retries
    loop retry attempts
        Retry->>Worker: POST /chat/completions (with headers)
        Worker-->>Retry: Response (stream or body)
    end
    alt streaming
        Retry->>Client: relay event stream (preserve content-type/status)
    else non-streaming
        Retry->>Client: forward body (preserve content-type/status)
    end
    Router->>Router: update circuit breaker & record final metrics
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

Suggested reviewers

  • CatherineSue
  • key4ng

Poem

🐰
A hop, a route, a stream set free,
I patch and proxy merrily,
Health beacons blink, retries take flight,
Providers shape the model's light,
The gateway hums into the night.

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 38.46% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main refactoring effort: extracting chat and health routing logic from router.rs into dedicated modules, which aligns with the substantial code reorganization across four files.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
  • 📝 Generate docstrings (stacked PR)
  • 📝 Generate docstrings (commit on current branch)
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch slin/oai-refactor-3

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request refactors the OpenAI router by splitting its monolithic router.rs file into more manageable and focused modules. The primary goal is to improve code navigability and maintainability by separating concerns, specifically moving chat completion routing and health check functionalities into their own dedicated files. This is a pure refactor, ensuring no changes in external behavior.

Highlights

  • Chat Routing Extraction: The route_chat logic, including worker selection, provider transformation, and retry mechanisms, was moved from router.rs to a new openai/chat.rs module.
  • Health Endpoint Extraction: The health_generate and get_server_info functions were extracted from router.rs into a new openai/health.rs module.
  • Router Simplification: The main router.rs file was significantly reduced in size and now delegates chat and health responsibilities to the new modules.
  • Shared Provider Resolution: The resolve_provider function was extracted as a pub(super) free function, allowing it to be shared by the new chat.rs module and other parts of the router.
  • RouterContext Introduction: A RouterContext struct was introduced to provide shared infrastructure (worker registry, provider registry, client, retry config) to the extracted free functions.
Changelog
  • model_gateway/src/routers/openai/chat.rs
    • Created new module for OpenAI chat completion routing.
    • Implemented RouterContext struct for shared dependencies.
    • Implemented route_chat function, encapsulating chat routing logic.
  • model_gateway/src/routers/openai/health.rs
    • Created new module for OpenAI health check and server info endpoints.
    • Implemented external_workers helper function.
    • Implemented health_generate and get_server_info functions.
  • model_gateway/src/routers/openai/mod.rs
    • Added mod chat; to include the new chat routing module.
    • Added mod health; to include the new health endpoint module.
  • model_gateway/src/routers/openai/router.rs
    • Removed route_chat implementation, delegating to chat::route_chat.
    • Removed health_generate and get_server_info implementations, delegating to health module functions.
    • Extracted resolve_provider into a pub(super) free function.
    • Removed unused external_workers and shared_components accessor methods.
    • Removed unused mpsc and UnboundedReceiverStream imports.
    • Updated RouterTrait methods to call the new module functions.
Activity
  • The author, slin1237, performed self-testing, confirming that cargo build -p smg, cargo clippy -p smg --all-targets -D warnings, and cargo test -p smg --lib all passed successfully.
  • No other review comments or activities have been recorded yet.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request refactors router.rs into chat.rs and health.rs, improving maintainability and readability. However, it highlights existing security concerns: unvalidated user input for model names in metrics can lead to Denial of Service, and internal worker URLs are exposed in health and server info endpoints, causing information leakage. Additionally, there are opportunities to improve efficiency in the new health.rs module by reducing redundant iterations and allocations.

Comment thread model_gateway/src/routers/openai/chat.rs
Comment thread model_gateway/src/routers/openai/health.rs
Comment thread model_gateway/src/routers/openai/health.rs

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@model_gateway/src/routers/openai/chat.rs`:
- Around line 48-50: The code resolves an override via model_id into the local
variable model but never rewrites the outgoing chat request or the
RequestContext, causing mismatched metrics/provider selection vs the upstream
payload; update the chat path to mutate or build a new request_body where
body.model is set to model (i.e., apply the model_id override into the
serialized payload used for routing/worker lookup) before any provider
resolution and before serializing the body, and pass request_body.clone() into
RequestContext::for_chat so downstream state and metrics use the same overridden
model; ensure all branches that inspect body.model or serialize body (the
selection logic around model usage and the serialization near where
streaming/worker lookup occurs) use this rewritten request_body.

In `@model_gateway/src/routers/openai/health.rs`:
- Around line 60-68: The response currently returns JSON as plain text; change
the return to use axum's JSON wrapper so the Content-Type is application/json.
Replace the final line that returns (StatusCode::OK,
info.to_string()).into_response() with (StatusCode::OK,
axum::Json(info)).into_response() (and add use axum::Json if not already
imported); ensure the `info` value remains a serde_json::Value (from json!) so
axum::Json can serialize it.

In `@model_gateway/src/routers/openai/router.rs`:
- Around line 56-70: The helper resolve_provider currently only checks the
optional override (model_id) and skips provider resolution when model_id is
None, causing non-default models to fall back to default_provider_arc; update
the call sites (this file where resolve_provider is invoked and
model_gateway/src/routers/openai/chat.rs around the noted lines) to pass the
already-computed effective model string instead of the raw optional override,
and adjust resolve_provider (and its signature if needed) to use that effective
model (so its logic still uses worker.provider_for_model(...) and
ProviderType::from_model_name(...) on the effective model) rather than relying
on an Option that may be None. Ensure resolve_provider no longer ignores the
effective request model and that both callers supply that resolved model value.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: ce473558-e432-42f5-ab8a-bf3e08abb9eb

📥 Commits

Reviewing files that changed from the base of the PR and between 79dc7e4 and 11fc8dd.

📒 Files selected for processing (4)
  • model_gateway/src/routers/openai/chat.rs
  • model_gateway/src/routers/openai/health.rs
  • model_gateway/src/routers/openai/mod.rs
  • model_gateway/src/routers/openai/router.rs

Comment on lines +48 to +50
let start = Instant::now();
let model = model_id.unwrap_or(body.model.as_str());
let streaming = body.stream;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Rewrite the chat request with the override model before routing it.

The function resolves model from model_id, but Line 64 still selects on body.model and Line 85 serializes the original body unchanged. If model_id is set, metrics/provider resolution point at one model while worker lookup and the upstream payload use another.

🐛 Proposed fix
-    let start = Instant::now();
-    let model = model_id.unwrap_or(body.model.as_str());
+    let start = Instant::now();
+    let mut request_body = body.clone();
+    if let Some(model_id) = model_id {
+        request_body.model = model_id.to_string();
+    }
+    let model = request_body.model.as_str();
     let streaming = body.stream;
@@
         .select_worker(&SelectWorkerRequest {
-            model_id: body.model.as_str(),
+            model_id: request_body.model.as_str(),
             headers,
             provider: Some(ProviderType::OpenAI),
             ..Default::default()
         })
@@
-    let mut payload = match to_value(body) {
+    let mut payload = match to_value(&request_body) {
         Ok(v) => v,

I'd also pass request_body.clone() into RequestContext::for_chat so downstream state stays aligned.

Also applies to: 61-68, 85-120

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@model_gateway/src/routers/openai/chat.rs` around lines 48 - 50, The code
resolves an override via model_id into the local variable model but never
rewrites the outgoing chat request or the RequestContext, causing mismatched
metrics/provider selection vs the upstream payload; update the chat path to
mutate or build a new request_body where body.model is set to model (i.e., apply
the model_id override into the serialized payload used for routing/worker
lookup) before any provider resolution and before serializing the body, and pass
request_body.clone() into RequestContext::for_chat so downstream state and
metrics use the same overridden model; ensure all branches that inspect
body.model or serialize body (the selection logic around model usage and the
serialization near where streaming/worker lookup occurs) use this rewritten
request_body.

Comment thread model_gateway/src/routers/openai/health.rs Outdated
Comment thread model_gateway/src/routers/openai/router.rs
What changed:
- model_gateway/src/routers/openai/router.rs: resolve_provider now takes
  &str instead of Option<&str>, so provider resolution always uses the
  effective model (not just the URL override). route_responses call site
  updated to pass the computed model.
- model_gateway/src/routers/openai/chat.rs: worker selection now uses the
  effective model (model_id override or body.model) instead of always
  using body.model. When model_id overrides the body model, the
  serialized payload is patched to match. resolve_provider call updated
  to pass &str.
- model_gateway/src/routers/openai/health.rs: get_server_info now returns
  application/json Content-Type via axum::Json instead of text/plain.

Why:
- resolve_provider with Option<&str> skipped provider lookup entirely for
  body-driven requests (model_id=None), always returning default_provider.
  Passing the effective model ensures correct provider selection.
- Worker selection using body.model while metrics/provider used model_id
  caused inconsistent routing when a URL path override was provided.
- get_server_info returned JSON content with text/plain Content-Type,
  breaking clients that check Content-Type headers.

Signed-off-by: Simo Lin <linsimo.mark@gmail.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@model_gateway/src/routers/openai/chat.rs`:
- Around line 191-220: The code currently unconditionally sets CONTENT_TYPE to
"text/event-stream" for streaming paths; change this so we preserve the upstream
Content-Type for non-success responses and only force "text/event-stream" for
successful SSE responses: in the is_streaming branch (where resp.bytes_stream()
is used and you build Response::new(...)), check status.is_success() (or check
resp.headers().get(CONTENT_TYPE)) before calling
response.headers_mut().insert(CONTENT_TYPE,
HeaderValue::from_static("text/event-stream")); only insert/override the header
when the upstream status is success (or when no Content-Type exists) so 4xx/5xx
JSON error payloads keep their original Content-Type.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 13439dd3-6ce4-4c68-84e4-cb9d98d70bc8

📥 Commits

Reviewing files that changed from the base of the PR and between 11fc8dd and 3fff1cc.

📒 Files selected for processing (3)
  • model_gateway/src/routers/openai/chat.rs
  • model_gateway/src/routers/openai/health.rs
  • model_gateway/src/routers/openai/router.rs

Comment thread model_gateway/src/routers/openai/chat.rs
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

model-gateway Model gateway crate changes openai OpenAI router changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants