Skip to content

refactor(responses): retrieval to use data layer directly - #938

Merged
slin1237 merged 1 commit into
mainfrom
refactor/responses-data-layer-retrieval
Mar 27, 2026
Merged

slin1237 merged 1 commit into
mainfrom
refactor/responses-data-layer-retrieval

Conversation

@zhaowenzi

@zhaowenzi zhaowenzi commented Mar 27, 2026 •

Copy link
Copy Markdown
Contributor

Description

Problem

When running with --enable-igw, response retrieval APIs such as GET /v1/responses/{response_id} will return not implemented.
The root cause is that these APIs do not carry a model_id, so IGW cannot reliably choose the correct router. Before this change, these endpoints were routed through the generic router-selection path, which works for inference requests but is not a good fit for storage-backed response APIs. Depending on which router was selected, the request could land on a router that does not implement the endpoint and return 501 Not Implemented.
This is different from /v1/conversations/*, which already operates directly on the shared data layer instead of relying on router dispatch.

Solution

Move response retrieval-style APIs to the data layer directly, following the same pattern as conversations.
This PR makes GET /v1/responses/{response_id}, GET /v1/responses/{response_id}/input_items, and DELETE /v1/responses/{response_id} use shared response storage directly from the server layer instead of dispatching through RouterTrait / RouterManager.
It also removes the now-unused router trait methods and related plumbing for response retrieval, so these endpoints are no longer coupled to IGW router selection.
Known limitation: HTTP regular router response persistence is still incomplete. Responses created through that path may not be persisted to shared response_storage, so this PR does not attempt to change that behavior. That issue is intentionally left out of scope for this change.

Changes

  • Route GET /v1/responses/{response_id} directly to shared response storage
  • Route GET /v1/responses/{response_id}/input_items directly to shared response storage
  • Implement DELETE /v1/responses/{response_id} against shared response storage
  • Remove response retrieval / deletion methods from RouterTrait and router implementations
  • Add a dedicated shared responses module aligned with the existing conversations structure
  • Remove unused ResponsesGetParams
  • Update tests to validate data-layer-authoritative behavior instead of router fanout behavior

Test Plan

Before

1. cargo run --bin smg -- --enable-igw --port 9999
2. curl -X POST http://localhost:9999/workers \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://api.openai.com",
    "api_key": "......",
    "runtime": "external",
    "disable_health_check": true
  }'
3. curl http://localhost:9999/v1/responses \
  -H "Authorization: Bearer ......" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.4",
    "input": "hello"
}'  -> return responses
4. curl http://localhost:9999/v1/responses/<response_id from the above request>
-> *Get response not implemented*

After

1. cargo run --bin smg -- --enable-igw --port 9999
2. curl -X POST http://localhost:9999/workers \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://api.openai.com",
    "api_key": "......",
    "runtime": "external",
    "disable_health_check": true
  }'
3. curl http://localhost:9999/v1/responses \
  -H "Authorization: Bearer ......" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.4",
    "input": "hello"
}'  -> return responses
4. curl http://localhost:9999/v1/responses/<response_id from the above request>
-> *Return the right response*
Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated
  • (Optional) Please join us on Slack #sig-smg to discuss, review, and merge PRs

Summary by CodeRabbit

  • Refactor
    • Refactored internal routing architecture for response operations to improve system performance and organization.
    • Consolidated response handler implementations for cleaner code structure.
    • All response API endpoints (retrieve, delete, list) continue functioning as before.

Signed-off-by: Ziwen Zhao <zzw.mose@gmail.com>
@github-actions github-actions Bot added grpc gRPC client and router changes tests Test changes model-gateway Model gateway crate changes openai OpenAI router changes labels Mar 27, 2026
@coderabbitai

coderabbitai Bot commented Mar 27, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

The PR refactors response retrieval endpoints (GET, DELETE, list input items) by removing protocol-specific router implementations and consolidating shared handler logic into a new dedicated module. Endpoints now directly call handlers backed by ResponseStorage instead of delegating through router trait methods.

Changes

Cohort / File(s) Summary
Protocol Definitions
crates/protocols/src/responses.rs
Removed public ResponsesGetParams struct and all associated serde field defaults.
Router Trait
model_gateway/src/routers/mod.rs
Removed three trait methods from RouterTrait: get_response, delete_response, and list_response_input_items; added module export for responses.
Router Implementations
model_gateway/src/routers/grpc/router.rs, model_gateway/src/routers/http/router.rs, model_gateway/src/routers/openai/router.rs, model_gateway/src/routers/grpc/common/responses/handlers.rs, model_gateway/src/routers/openai/responses/mod.rs, model_gateway/src/routers/router_manager.rs
Removed protocol-specific implementations of get_response, delete_response, and list_response_input_items methods and their supporting handler functions; removed ResponsesGetParams imports.
Shared Response Handlers
model_gateway/src/routers/responses/mod.rs, model_gateway/src/routers/responses/handlers.rs
Added new shared response handler module with three async functions (get_response, delete_response, list_response_input_items) that directly interface with ResponseStorage to provide response retrieval, deletion, and input item listing operations.
Server Routing
model_gateway/src/server.rs
Rewired three response endpoints (/v1/responses/{response_id} GET/DELETE, /v1/responses/{response_id}/input_items GET) to directly call shared response handlers instead of delegating through router trait methods; removed HeaderMap and ResponsesGetParams parameters from endpoint signatures.
Tests
model_gateway/tests/api/api_endpoints_test.rs, model_gateway/tests/routing/test_openai_routing.rs
Updated tests to prepopulate response_storage directly with StoredResponse objects instead of issuing HTTP POST requests; removed ResponsesGetParams usage; adjusted test assertions to validate storage state and endpoint responses (DELETE now expects 200 OK success instead of NOT_IMPLEMENTED).

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

Suggested labels

model-gateway, tests, grpc, openai, data-connector

Suggested reviewers

  • CatherineSue
  • key4ng
  • slin1237

Poem

🐰 The routers once danced with every request,
But handlers, they said, deserved rest.
To storage they bounded, direct and free,
No traits in between—just ResponseStorage!
The code now flows cleaner, all agree. ✨

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 38.89% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'refactor(responses): retrieval to use data layer directly' accurately and specifically describes the main change—moving response retrieval APIs to operate on the shared data layer instead of through routers.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch refactor/responses-data-layer-retrieval

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request refactors the response management system by centralizing the logic for retrieving, deleting, and listing response input items into a new responses module. It removes these methods from the RouterTrait and its various implementations (gRPC, HTTP, OpenAI), allowing the server to handle these requests directly via the shared response_storage. This change simplifies the routing architecture and ensures consistent behavior across different router types. I have no feedback to provide as there were no review comments to assess.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: b1b6ea7260

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread model_gateway/src/server.rs

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
model_gateway/tests/api/api_endpoints_test.rs (1)

999-1007: ⚠️ Potential issue | 🟡 Minor

Strengthen shared-storage retrieval assertion beyond status code.

Line 1006 only checks 200 OK, so this test can pass even if the wrong response object is returned.

✅ Suggested test hardening
         let resp = app.clone().oneshot(req).await.unwrap();
         assert_eq!(resp.status(), StatusCode::OK);
+        let body = axum::body::to_bytes(resp.into_body(), usize::MAX)
+            .await
+            .unwrap();
+        let get_json: serde_json::Value = serde_json::from_slice(&body).unwrap();
+        assert_eq!(get_json["id"], rid);
+        assert_eq!(get_json["object"], "response");
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@model_gateway/tests/api/api_endpoints_test.rs` around lines 999 - 1007, The
test currently only asserts the HTTP status for the GET /v1/responses/{rid}
call; update the assertion to parse the response body (from the Response
returned by app.clone().oneshot) as JSON and validate that it contains the
expected response object fields (at minimum that the "id" matches rid, and
preferably all fields or an exact JSON equality against the originally stored
response); modify the test around the Request/resp handling in
api_endpoints_test.rs so after asserting StatusCode::OK you read the body bytes,
deserialize to the same response struct or serde_json::Value, and assert the id
and other fields match the expected values to ensure the correct object was
returned.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@model_gateway/src/routers/responses/handlers.rs`:
- Around line 36-37: The current code does a get_response(&id) followed by
delete_response(&id), which has a TOCTOU race if another caller deletes between
those calls; change logic to avoid the separate existence check: call
response_storage.delete_response(&id).await directly and handle its result
idempotently (treat “not found” as success) or make delete_response itself
return success when the record is already absent. Update handlers that reference
get_response and delete_response to rely on delete_response’s idempotent
behavior and remove the pre-check to eliminate the race.
- Around line 72-82: The current items_with_ids creation in
list_response_input_items injects a random ID via generate_id("msg") at read
time causing non-deterministic IDs across requests; to fix, replace the
ephemeral random generation with a deterministic ID (for example compute a
stable hash of the item's JSON plus its index or other stable fields) or persist
the generated ID back into the stored response so subsequent reads return the
same id; locate the items_with_ids mapping and change the id assignment logic
(currently calling generate_id("msg")) to either (a) compute a deterministic id
from item content/position or (b) write the generated id back into the response
storage so future calls see the same id.

---

Outside diff comments:
In `@model_gateway/tests/api/api_endpoints_test.rs`:
- Around line 999-1007: The test currently only asserts the HTTP status for the
GET /v1/responses/{rid} call; update the assertion to parse the response body
(from the Response returned by app.clone().oneshot) as JSON and validate that it
contains the expected response object fields (at minimum that the "id" matches
rid, and preferably all fields or an exact JSON equality against the originally
stored response); modify the test around the Request/resp handling in
api_endpoints_test.rs so after asserting StatusCode::OK you read the body bytes,
deserialize to the same response struct or serde_json::Value, and assert the id
and other fields match the expected values to ensure the correct object was
returned.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 54e8b2e6-abd1-4d1a-85ed-4a7d558201cd

📥 Commits

Reviewing files that changed from the base of the PR and between dcc2225 and b1b6ea7.

📒 Files selected for processing (13)
  • crates/protocols/src/responses.rs
  • model_gateway/src/routers/grpc/common/responses/handlers.rs
  • model_gateway/src/routers/grpc/router.rs
  • model_gateway/src/routers/http/router.rs
  • model_gateway/src/routers/mod.rs
  • model_gateway/src/routers/openai/responses/mod.rs
  • model_gateway/src/routers/openai/router.rs
  • model_gateway/src/routers/responses/handlers.rs
  • model_gateway/src/routers/responses/mod.rs
  • model_gateway/src/routers/router_manager.rs
  • model_gateway/src/server.rs
  • model_gateway/tests/api/api_endpoints_test.rs
  • model_gateway/tests/routing/test_openai_routing.rs
💤 Files with no reviewable changes (1)
  • crates/protocols/src/responses.rs

Comment thread model_gateway/src/routers/responses/handlers.rs
Comment thread model_gateway/src/routers/responses/handlers.rs
@slin1237
slin1237 merged commit 1a999fc into main Mar 27, 2026
61 of 65 checks passed
@slin1237
slin1237 deleted the refactor/responses-data-layer-retrieval branch March 27, 2026 01:31
smfirmin pushed a commit to smfirmin/smg that referenced this pull request Apr 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

grpc gRPC client and router changes model-gateway Model gateway crate changes openai OpenAI router changes tests Test changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants