Skip to content

refactor(gateway): extract route_responses orchestration into responses/route.rs - #735

Merged
slin1237 merged 1 commit into
mainfrom
slin/oai-refactor-6
Mar 11, 2026
Merged

slin1237 merged 1 commit into
mainfrom
slin/oai-refactor-6

Conversation

@slin1237

@slin1237 slin1237 commented Mar 11, 2026 •

Copy link
Copy Markdown
Member

Summary

Extract ~145 lines of inline route_responses orchestration logic from router.rs into a dedicated responses/route.rs module, matching the existing chat.rs delegation pattern. Continues the OpenAI router cleanup series (PRs #730, #732).

Closes: N/A (refactor)
Refs: #730, #732

What changed

  • responses/route.rs (new): ResponsesRouterContext struct + route_responses() free function with all orchestration logic — metrics recording, worker selection, mutual exclusivity validation, history loading, payload serialization, provider transform, RequestContext setup, streaming/non-streaming dispatch, duration metrics
  • responses/mod.rs: added pub(crate) mod route; declaration
  • router.rs: replaced 145-line inline route_responses with 8-line delegation (pack deps into ResponsesRouterContext, delegate to responses_route::route_responses()). Removed dead responses_components() helper. Cleaned up unused imports (Instant, to_value, ResponseInput, ResponseInputOutputItem, ComponentRefs, PayloadState, RequestContext, WorkerSelection, handle_non_streaming_response, handle_streaming_response, Endpoint, bool_to_static_str, error)
  • chat.rs: renamed RouterContext → ChatRouterContext for naming consistency with ResponsesRouterContext

Why

route_chat was already 8 lines of delegation while route_responses was 145 lines inline — inconsistent patterns in the same file. Now both follow identical structure: pack borrowed refs into a context struct, delegate to a module-level free function. Reduces router.rs from ~422 to ~278 lines.

How

Pure mechanical extraction — orchestration logic moved verbatim with self.* replaced by deps.* on the new context struct. Worker selection inlined via WorkerSelector::new() matching the chat.rs pattern.

Test plan

  • cargo build -p smg — compiles clean
  • cargo clippy -p smg --all-targets -- -D warnings — no warnings
  • cargo test -p smg — all 93+ tests pass across all test suites
  • Pure refactor, no behavior changes — RouterTrait interface unchanged

Summary by CodeRabbit

  • Refactor
    • Reorganized OpenAI router architecture by centralizing responses routing logic into a dedicated module and introducing specialized routing contexts for chat and responses handling.

…es/route.rs

Move ~145 lines of inline orchestration logic from the RouterTrait
impl in router.rs into a dedicated responses/route.rs module,
mirroring the existing chat.rs delegation pattern.

What changed:
- model_gateway/src/routers/openai/responses/route.rs: new file with
  ResponsesRouterContext struct and route_responses() free function
  containing all responses orchestration logic (metrics, worker
  selection, validation, history loading, payload prep, provider
  transform, RequestContext setup, streaming dispatch, duration metrics)
- model_gateway/src/routers/openai/responses/mod.rs: add pub(crate)
  mod route declaration
- model_gateway/src/routers/openai/router.rs: replace 145-line inline
  route_responses with 8-line delegation (pack deps + delegate),
  remove dead responses_components() helper, clean up unused imports
  (Instant, to_value, ResponseInput, ResponseInputOutputItem,
  ComponentRefs, PayloadState, RequestContext, WorkerSelection,
  handle_non_streaming_response, handle_streaming_response, Endpoint,
  bool_to_static_str, error)
- model_gateway/src/routers/openai/chat.rs: rename RouterContext to
  ChatRouterContext for consistency with ResponsesRouterContext

Why: router.rs now has both route_chat and route_responses following
the same delegation pattern — pack borrowed refs into a context struct,
delegate to a module-level free function. This reduces router.rs from
~422 to ~278 lines of pure routing dispatch.

How: pure mechanical extraction — the orchestration logic is moved
verbatim with self.* references replaced by deps.* accesses on the
new context struct. Worker selection is inlined via WorkerSelector::new()
matching the chat.rs pattern.

Signed-off-by: Simon Lin <simon@seekerlabs.ai>
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
@github-actions github-actions Bot added model-gateway Model gateway crate changes openai OpenAI router changes labels Mar 11, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request focuses on improving the modularity and organization of the OpenAI router by refactoring the route_responses logic. It extracts a substantial block of code into a new, dedicated module, establishing a consistent delegation pattern similar to existing chat routing. This change reduces the complexity of the main router file and enhances the maintainability of the codebase without altering any external behavior.

Highlights

  • Code Extraction: Extracted approximately 145 lines of route_responses orchestration logic from router.rs into a new, dedicated module: responses/route.rs.
  • New Module Creation: Introduced responses/route.rs which now houses the ResponsesRouterContext struct and the route_responses() free function, centralizing the responses API routing logic.
  • Router Simplification: The router.rs file was significantly reduced in size (from ~422 to ~278 lines) by replacing the inline route_responses implementation with a concise 8-line delegation call to the new module.
  • Naming Consistency: Renamed RouterContext in chat.rs to ChatRouterContext to align with the newly introduced ResponsesRouterContext, improving naming consistency across the router components.
  • Code Cleanup: Removed a dead responses_components() helper function and several unused imports from router.rs, enhancing code cleanliness and maintainability.
Changelog
  • model_gateway/src/routers/openai/chat.rs
    • Renamed RouterContext to ChatRouterContext for better naming consistency.
    • Updated the route_chat function signature to use the new ChatRouterContext.
  • model_gateway/src/routers/openai/responses/mod.rs
    • Added a new module declaration pub(crate) mod route; to expose the new responses routing logic.
  • model_gateway/src/routers/openai/responses/route.rs
    • Created a new file containing the ResponsesRouterContext struct and the route_responses function.
    • Implemented the full orchestration logic for responses API requests, including metrics, worker selection, validation, history loading, payload serialization, provider transformation, and request dispatch.
  • model_gateway/src/routers/openai/router.rs
    • Removed the extensive inline implementation of route_responses and replaced it with a delegation call to responses_route::route_responses().
    • Deleted the responses_components() helper function as it is no longer needed.
    • Cleaned up numerous unused imports, including Instant, to_value, ResponseInput, RequestContext, and others.
Activity
  • No human activity has been recorded on this pull request yet.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@coderabbitai

coderabbitai Bot commented Mar 11, 2026 •

Copy link
Copy Markdown

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 7a82c426-310b-47dd-a4e6-e3d80ea43c80

📥 Commits

Reviewing files that changed from the base of the PR and between 8d1b7b2 and 97447de.

📒 Files selected for processing (4)
  • model_gateway/src/routers/openai/chat.rs
  • model_gateway/src/routers/openai/responses/mod.rs
  • model_gateway/src/routers/openai/responses/route.rs
  • model_gateway/src/routers/openai/router.rs

📝 Walkthrough

Walkthrough

This PR refactors OpenAI routing by extracting response routing logic into a dedicated module. It renames RouterContext to ChatRouterContext, introduces a new ResponsesRouterContext, and centralizes response routing orchestration in responses/route.rs with comprehensive metrics, validation, and provider handling.

Changes

Cohort / File(s) Summary
Chat Routing Context
model_gateway/src/routers/openai/chat.rs
Renamed struct from RouterContext to ChatRouterContext and updated function signature; added public fields for HTTP client and retry configuration.
Response Routing Module Setup
model_gateway/src/routers/openai/responses/mod.rs
Added new crate-internal route module export to enable response routing delegation.
Response Routing Implementation
model_gateway/src/routers/openai/responses/route.rs (NEW)
Introduced ResponsesRouterContext and comprehensive route_responses function orchestrating model extraction, worker selection, mutual exclusivity validation, request transformation, provider resolution, and response handling with integrated metrics.
Main Router Refactoring
model_gateway/src/routers/openai/router.rs
Replaced inline response routing logic with delegation to responses_route::route_responses; updated chat routing to use ChatRouterContext; simplified overall router implementation by removing 143 lines of complex routing logic.

Sequence Diagram

sequenceDiagram
    participant Client
    participant Router as OpenAI Router
    participant RouteResp as route_responses
    participant WorkerSel as WorkerSelector
    participant WRegistry as WorkerRegistry
    participant PRegistry as ProviderRegistry
    participant RespHandler as Response Handler
    participant Upstream as Upstream Service

    Client->>Router: POST /responses
    Router->>RouteResp: route_responses(context, headers, body)
    
    RouteResp->>RouteResp: Extract model & streaming<br/>Record router metric
    RouteResp->>WorkerSel: Select worker with model
    WorkerSel->>WRegistry: Lookup worker
    WRegistry-->>WorkerSel: Worker result
    WorkerSel-->>RouteResp: Worker or error
    
    alt Worker Selection Failed
        RouteResp->>RouteResp: Record error metric
        RouteResp-->>Router: Return error response
    else Success
        RouteResp->>RouteResp: Validate conversation vs<br/>previous_response_id
        
        alt Validation Failed
            RouteResp->>RouteResp: Record validation error
            RouteResp-->>Router: Return 400 error
        else Validation Passed
            RouteResp->>RouteResp: Load input history
            RouteResp->>RouteResp: Strip reasoning items<br/>if store disabled
            RouteResp->>PRegistry: Resolve provider
            PRegistry-->>RouteResp: Provider result
            
            alt Provider Resolution Failed
                RouteResp->>RouteResp: Record validation error
                RouteResp-->>Router: Return 400 error
            else Success
                RouteResp->>RouteResp: Apply provider transform
                RouteResp->>RespHandler: Handle response<br/>(streaming or non-streaming)
                RespHandler->>Upstream: Forward request
                Upstream-->>RespHandler: Response
                RespHandler-->>RouteResp: Return response
                RouteResp->>RouteResp: Record route duration metric
                RouteResp-->>Router: Return response
            end
        end
    end
    
    Router-->>Client: HTTP Response
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~50 minutes

Possibly related PRs

Suggested reviewers

  • CatherineSue
  • key4ng

🐰 A router hops with glee,
Routes now flow more clearly,
Chat and responses split—
Each with context fit,
The refactor's done with care! ✨

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 33.33% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately reflects the main change: extracting the route_responses orchestration logic from router.rs into a new responses/route.rs module, which is the primary objective of this refactoring PR.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
  • 📝 Generate docstrings (stacked PR)
  • 📝 Generate docstrings (commit on current branch)
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch slin/oai-refactor-6

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request is a solid refactoring that extracts the route_responses logic into its own module, improving code organization and consistency with the route_chat pattern. The changes are clean and well-executed.

I've identified a minor improvement opportunity to reduce redundancy in the newly created ResponsesRouterContext. The client field is redundant as it can be accessed through responses_components. Removing it would simplify the context struct. A similar pattern exists in ChatRouterContext, which could also be simplified in a follow-up change.

pub worker_registry: &'a WorkerRegistry,
pub provider_registry: &'a ProviderRegistry,
pub responses_components: &'a Arc<ResponsesComponents>,
pub client: &'a reqwest::Client,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The client field in ResponsesRouterContext is redundant, as the reqwest::Client can be accessed via deps.responses_components.shared.client. Removing this field will simplify the context struct.

bool_to_static_str(streaming),
);

let worker = match WorkerSelector::new(deps.worker_registry, deps.client)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

To go along with the removal of the redundant client field from ResponsesRouterContext, please update this to access the client via deps.responses_components.shared.client.

Suggested change
let worker = match WorkerSelector::new(deps.worker_registry, deps.client)
let worker = match WorkerSelector::new(deps.worker_registry, &deps.responses_components.shared.client)

Comment on lines +165 to 170
let deps = ResponsesRouterContext {
worker_registry: &self.worker_registry,
provider_registry: &self.provider_registry,
responses_components: &self.responses_components,
client: &self.responses_components.shared.client,
};

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Following the removal of the redundant client field from ResponsesRouterContext, this field should be removed from the struct initialization.

        let deps = ResponsesRouterContext {
            worker_registry: &self.worker_registry,
            provider_registry: &self.provider_registry,
            responses_components: &self.responses_components,
        };

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

model-gateway Model gateway crate changes openai OpenAI router changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant