feat(openai): wire image_generation streaming events in Responses router (R6.2) - #1356
Conversation
…ter (R6.2) R6.1 (#1355) landed the shared `ResponseFormat::ImageGenerationCall` plumbing and left the OpenAI Responses router with placeholder marker comments (no real per-site wiring in `responses/streaming.rs`). R6.2 finishes the per-router wiring so `image_generation` MCP tool calls emit the full `response.image_generation_call.*` event sequence through the OpenAI Responses streaming pipeline, matching the shape used for web_search / code_interpreter / file_search. Changes in `model_gateway/src/routers/openai/responses/streaming.rs`: - import `ImageGenerationCallEvent` from `openai_protocol::event_types`. - `apply_event_transformations_inplace`: extend the `ResponseFormat` match so upstream `function_call` items whose session `response_format` is `ImageGenerationCall` are rewritten to `ItemType::IMAGE_GENERATION_CALL` with the `ig_` id prefix, instead of falling through to the generic `mcp_call` arm. This mirrors the existing `WebSearchCall` arm. - `maybe_inject_tool_in_progress`: add an `ItemType::IMAGE_GENERATION_CALL` arm that emits `ImageGenerationCallEvent::IN_PROGRESS`, so the in_progress event is injected after the rewritten `output_item.added` (parallel to the existing `WEB_SEARCH_CALL` / `CODE_INTERPRETER_CALL` / `FILE_SEARCH_CALL` arms). Changes in `model_gateway/src/routers/openai/mcp/tool_loop.rs`: - Drop the temporary "stubbed" marker comments left by R6.1 on the three `ImageGenerationCall` arms (`send_tool_call_intermediate_event`, `stable_streaming_tool_item_id`, `non_streaming_tool_item_id_source`). The underlying values (`ImageGenerationCallEvent::GENERATING`, `ig_` prefix, shared `fc_`/`call_` strip behavior) already matched the other three hosted-tool formats and now stand as the real wiring. - Replace the R6.2-TODO comment on the intermediate event with a short design note: `generating` is the coarse intermediate event emitted by the tool_loop; `partial_image` preview events are the responsibility of the underlying tool when it streams chunks, not the tool_loop itself. End-to-end flow for `image_generation`: 1. upstream emits `output_item.added` with a `function_call` item → rewritten to `image_generation_call` with an `ig_*` id. 2. `response.image_generation_call.in_progress` is injected right after. 3. `send_tool_call_intermediate_event` emits `response.image_generation_call.generating` while the tool runs. 4. `send_tool_call_completion_events` emits `response.image_generation_call.completed` and `response.output_item.done` with the fully-transformed item from the shared `to_image_generation_call` transformer. Scope: protocol-only router wiring under `model_gateway/src/routers/openai/`. Non-OpenAI transports (gRPC, HTTP workers) are handled by R6.3 / R6.4 and are intentionally not touched here. Gates: - `cargo check -p smg --lib` clean - `cargo test -p smg --lib` 616 passed, 0 failed - `cargo fmt --all` - `cargo clippy -p smg --lib --tests -- -D warnings` clean Refs: R6.1 #1355 (prerequisite) Signed-off-by: Simo Lin <linsimo.mark@gmail.com> Co-authored-by: Tingting Zhou <zhoutt96@gmail.com> Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (2)
📝 WalkthroughWalkthroughAdds streaming support for ImageGenerationCall: maps ResponseFormat::ImageGenerationCall to ItemType::IMAGE_GENERATION_CALL with an "ig_" ID prefix, updates in-progress injection to emit ImageGenerationCallEvent::IN_PROGRESS, and clarifies related tool-loop documentation and naming comments. Changes
Sequence Diagram(s)sequenceDiagram
participant Client as Client
participant Router as OpenAI Router
participant Tool as Image Tool Stream
participant MCP as MCP Tool Loop
Client->>Router: request with image_generation_call
Router->>MCP: create tool-call item (type: IMAGE_GENERATION_CALL, id:"ig_*")
MCP->>Tool: start image generation stream
Tool-->>MCP: partial_image events (streamed)
MCP-->>Router: send intermediate 'generating' event
Router-->>Client: emit ImageGenerationCallEvent::IN_PROGRESS (with ig_ id)
Tool-->>MCP: final image result
MCP-->>Router: completion events
Router-->>Client: final completion events
Estimated code review effort🎯 2 (Simple) | ⏱️ ~10 minutes Possibly related PRs
Suggested labels
Suggested reviewers
Poem
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
There was a problem hiding this comment.
Code Review
This pull request integrates ImageGenerationCall into the OpenAI response streaming and tool loop logic, handling event transformations, ID prefixing, and intermediate event emission. A review comment suggests extending the transformation logic to explicitly handle CodeInterpreterCall and FileSearchCall for consistency across all hosted tool formats.
| ResponseFormat::ImageGenerationCall => { | ||
| (ItemType::IMAGE_GENERATION_CALL, "ig_") | ||
| } |
There was a problem hiding this comment.
The apply_event_transformations_inplace function is missing transformation logic for the CodeInterpreterCall and FileSearchCall response formats. While WebSearchCall and the newly added ImageGenerationCall are handled, these other built-in tool formats will currently fall through to the generic MCP_CALL type with an mcp_ ID prefix. To ensure consistency across all hosted tool formats as intended by this PR, they should be explicitly mapped to their respective item types and ID prefixes (ci_ and fs_). This ensures that the transformation logic correctly identifies the item variants present in the final response output.
ResponseFormat::CodeInterpreterCall => {
(ItemType::CODE_INTERPRETER_CALL, "ci_")
}
ResponseFormat::FileSearchCall => (ItemType::FILE_SEARCH_CALL, "fs_"),
ResponseFormat::ImageGenerationCall => {
(ItemType::IMAGE_GENERATION_CALL, "ig_")
}References
- When processing response items, ensure the logic targets the specific item variants that are actually present in the final response output.
There was a problem hiding this comment.
Good catch on the gap — you're right that CodeInterpreterCall and FileSearchCall currently fall through to (ItemType::MCP_CALL, "mcp_") in this match. That predates R6.2; the existing WebSearchCall arm is the only explicit hosted-tool branch today. R6.1 / R6.2 are scoped to adding the image_generation_call path to match web_search_call, not to retrofit the other two formats.
Extending this match to also cover CodeInterpreterCall and FileSearchCall is a straightforward cleanup but touches behavior for two unrelated tool types, so it belongs in a follow-up rather than in the R6.x series (which is staged around the image_generation tool specifically). I'd rather keep this PR's blast radius narrow and file a separate PR that handles all the remaining hosted-tool formats uniformly.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 452e49572c
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| ResponseFormat::ImageGenerationCall => { | ||
| (ItemType::IMAGE_GENERATION_CALL, "ig_") | ||
| } |
There was a problem hiding this comment.
Align image-generation argument event IDs
Rewriting ResponseFormat::ImageGenerationCall to image_generation_call with an ig_ item ID here causes a protocol mismatch because argument events in this same module are still emitted as response.mcp_call_arguments.* with mcp_* IDs (FunctionCallEvent::ARGUMENTS_DONE and send_buffered_arguments). For image-generation tool calls, clients now receive one call under two different item_id namespaces (ig_* vs mcp_*), which breaks correlation logic that joins argument deltas/done to the corresponding output item.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Thanks for the P1 flag. To unpack it: the response.mcp_call_arguments.* events in this module (via FunctionCallEvent::ARGUMENTS_DONE and send_buffered_arguments) always rewrite item_id via mcp_response_item_id and stamp the type as McpEvent::CALL_ARGUMENTS_DONE regardless of response_format. That behavior predates R6.2 and applies uniformly to all three existing hosted-tool formats (web_search_call with ws_*, code_interpreter_call with ci_*, file_search_call with fs_*) — each of them already surfaces the same split-namespace pattern between the output_item's <prefix>_* id and the mcp_* argument-event item_id.
The R6.2 task is scoped to wiring image_generation_call through the same path as the other three hosted formats (the task instructions call out "stay consistent with the other 3 tools' shape; don't invent new patterns"). Fixing the argument-event namespace mismatch consistently for all four hosted formats — either by per-format id preservation or (more likely correct per OpenAI's spec) by suppressing mcp_call_arguments.* altogether for non-Passthrough formats — is a cross-cutting cleanup that belongs in a follow-up PR on top of R6.2/R6.3/R6.4 rather than a localized image-gen-only change that would itself diverge from the existing three-format shape.
Happy to pick that up as a follow-up once the R6.x series lands; leaving this PR narrowly scoped to image_generation wiring.
| ItemType::WEB_SEARCH_CALL => WebSearchCallEvent::IN_PROGRESS, | ||
| ItemType::CODE_INTERPRETER_CALL => CodeInterpreterCallEvent::IN_PROGRESS, | ||
| ItemType::FILE_SEARCH_CALL => FileSearchCallEvent::IN_PROGRESS, | ||
| ItemType::IMAGE_GENERATION_CALL => ImageGenerationCallEvent::IN_PROGRESS, |
There was a problem hiding this comment.
🟡 Nit: The docstring for this function (line 419) enumerates the handled item types but doesn't include image_generation_call after this addition. Consider updating it to stay in sync:
/// Handles mcp_call, web_search_call, code_interpreter_call, file_search_call, and image_generation_call items.
There was a problem hiding this comment.
Addressed in c0dfc00 — updated the docstring to include image_generation_call in the enumerated handled types, plus two parallel docstrings in tool_loop.rs (send_tool_call_completion_events and build_transformed_mcp_call_item) that had the same stale enumeration.
There was a problem hiding this comment.
Clean, well-scoped PR. The image_generation_call wiring in the Responses streaming router follows the established pattern (matches web_search_call). Comment updates in tool_loop.rs accurately replace the stale R6.2 stubs. One minor nit on a docstring enumeration — no blocking issues.
Address Claude PR Review nit on #1356: three helper docstrings enumerated the handled tool-call output item types (`mcp_call`, `web_search_call`, `code_interpreter_call`, `file_search_call`) but were not updated when R6.2 added `image_generation_call`. Add `image_generation_call` to the enumeration in: - `model_gateway/src/routers/openai/responses/streaming.rs` (`maybe_inject_tool_in_progress`) - `model_gateway/src/routers/openai/mcp/tool_loop.rs` (`send_tool_call_completion_events`, `build_transformed_mcp_call_item`) Docs-only change; no behavior difference. Signed-off-by: Simo Lin <linsimo.mark@gmail.com> Co-authored-by: Tingting Zhou <zhoutt96@gmail.com> Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: c0dfc0024b
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| ResponseFormat::ImageGenerationCall => { | ||
| (ItemType::IMAGE_GENERATION_CALL, "ig_") | ||
| } |
There was a problem hiding this comment.
Keep image_generation item IDs consistent across tool events
Fresh evidence in this commit: apply_event_transformations_inplace now rewrites image-generation response.output_item.* entries to image_generation_call with an ig_ prefix, but this module still emits argument events as response.mcp_call_arguments.* and normalizes item_id with mcp_response_item_id(...) (mcp_*). For image-generation tool calls, clients that correlate argument and lifecycle events by item_id will now see two namespaces for the same call (ig_* vs mcp_*), which breaks joining argument delta/done to the corresponding added/in_progress/completed events.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
This is a re-post of the same P1 from your previous review. See my reply on the original thread: #1356 (comment) — the response.mcp_call_arguments.* + mcp_response_item_id behavior is pre-existing across all three other hosted-tool formats (web_search_call, code_interpreter_call, file_search_call), not a regression introduced here. Fixing it consistently across all four formats is a separate cross-cutting cleanup tracked as a follow-up to the R6.x series.
Respond to the coordinator directive on PR #1365: remove the R6.3/R6.4 ``skip_for_runtime`` decorators, add local-gRPC fixtures for both engines, and harden the assertions on what the gateway emits. R6.1-R6.4 all merged. Real lane failures surfaced by this suite will be filed as R6.6/R6.7 follow-ups rather than patched from R6.5, so the CI signal is genuine rather than papered over. Engine matrix ------------- Three test classes, sharing a ``_ImageGenerationAssertions`` mix-in so the cloud + gRPC lanes stay strictly in lock-step: * ``TestImageGenerationCloud`` (vendor=openai, gpu=0) — OpenAI cloud backend with gpt-5-nano, mock MCP server standing in for the image backend. Exercises the OpenAI-compat router (R6.2). * ``TestImageGenerationGrpcSglang`` (engine=sglang, gpu=1, model=openai/gpt-oss-20b) — local SGLang worker via harmony. Exercises R6.3 gRPC-harmony wiring. * ``TestImageGenerationGrpcVllm`` (engine=vllm, gpu=1, model=meta-llama/Llama-3.1-8B-Instruct) — local vLLM worker via regular. Exercises R6.4 gRPC-regular wiring. Each lane's fixture wires its gateway at the same shared in-process ``MockMcpServer`` so deterministic base64-roundtrip / size-override assertions stay valid across engines. Fixture additions (``e2e_test/responses/conftest.py``) ------------------------------------------------------ * ``gateway_with_mock_mcp_cloud`` replaces the old ``gateway_with_mock_mcp``. Yields ``(gateway, client, mock, model)`` with ``model="gpt-5-nano"``. The old name is kept as a backward-compat alias so any unmerged branch referencing it still works. * ``_start_local_grpc_gateway_with_mcp`` helper: launches one gRPC worker for a given engine + model_id, spins up a gateway with ``--mcp-config-path`` pointing at the shared mock MCP config, wraps client instantiation in try/except so a failed init doesn't leak the gateway or the worker. * ``gateway_with_mock_mcp_grpc_sglang`` / ``gateway_with_mock_mcp_grpc_vllm``: class-scoped fixtures built on the helper. Each skips (not fails) when the gRPC worker can't start — CI lanes without GPUs otherwise poison every engine-parametrized suite. Harder assertions (``test_image_generation.py``) ------------------------------------------------ * ``_assert_image_generation_call_item`` now asserts every documented field (``type``, ``id`` ig-prefix, ``status``, ``result``, ``revised_prompt``) rather than just a subset. * ``_assert_streaming_envelope`` validates the full emitted envelope: * ``response.created`` → ``response.output_item.added`` → ``response.image_generation_call.{in_progress, generating, [partial_image], completed}`` → ``response.output_item.done`` → ``response.completed`` * each event's first occurrence strictly precedes the next required event's first occurrence * exactly one ``response.created`` / ``response.completed`` / ``output_item.added`` / ``output_item.done`` per image_gen call * ``sequence_number`` strictly monotonically increasing with no gaps > 1 (only checked when every event has one) * optional ``partial_image`` sits between ``generating`` and ``completed`` when present * Streaming and non-streaming paths share the same mix-in body, so the gRPC lanes exercise the full assertion surface without duplication. Local gates (ruff check/format, mypy, pytest --collect-only) clean; 12 tests collected (3 classes × 4 tests). Refs: R6.1 #1355, R6.2 #1356, R6.3 #1359, R6.4 #1358 (all merged) Co-authored-by: Tingting Zhou <zhoutt96@gmail.com> Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
Introduce e2e_test/infra/mock_mcp_server.py — a thin wrapper around the official MCP SDK's FastMCP base class that exposes streamable-HTTP transport on a local port. The existing MCP e2e test (e2e_test/messages/test_mcp_tool.py) depends on a live Brave MCP server which is flaky, slow, and non-deterministic; routing-layer tests for the gateway's built-in tool plumbing need an in-process, deterministic counterpart. This mock supplies one. Files added: * e2e_test/infra/mock_mcp_server.py — MockMcpServer class. Runs FastMCP via uvicorn on a background thread; auto-allocates a free port when the caller passes port=0. Registers an image_generation tool that returns a hard-coded 1x1 transparent PNG (92-char base64) plus the prompt echoed as revised_prompt, so tests can make byte-for-byte assertions. Records every call into call_log and surfaces the last call via last_call_args for override-verification tests. TODO stubs for web_search, file_search, code_interpreter sit inline with the same shape so future R6.x PRs just uncomment and adjust. * e2e_test/infra/mock_mcp.py — mock_mcp_server session-scoped pytest fixture plus IMAGE_GENERATION_PNG_BASE64 re-export. Files modified: * e2e_test/infra/constants.py — new MOCK_MCP_HOST constant (default 127.0.0.1) paralleling the existing BRAVE_MCP_HOST. * e2e_test/infra/__init__.py — export MockMcpServer, mock_mcp_server, IMAGE_GENERATION_PNG_BASE64, MOCK_MCP_HOST. Why streamable HTTP: matches the protocol the gateway's MCP client already speaks against Brave in production (see crates/mcp/src/core/config.rs::McpTransport::Streamable). The streamable_http_app() FastMCP method yields a /mcp Starlette route mountable directly under uvicorn. How to extend: FastMCP exposes a decorator-driven registration API; add a new @fastmcp.tool in MockMcpServer._register_tools, append to self._call_log inside the body for introspection, and you're done. See the module docstring for the extension recipe. Refs: R6.1 #1355, R6.2 #1356, R6.3 #1359, R6.4 #1358 (all merged) Co-authored-by: Tingting Zhou <zhoutt96@gmail.com> Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
Add the first end-to-end tests for the image_generation built-in tool,
using the in-process MockMcpServer introduced in the previous commit.
R6.1-R6.4 wired image_generation across the openai, gRPC-regular, and
gRPC-harmony routers end-to-end; this PR provides the first automated
coverage that exercises those paths without external-service
dependencies.
Files added:
* e2e_test/responses/conftest.py — fixture module for the Responses
suite. Provides:
- mock_mcp_config_file: writes a session-scoped YAML config to a
tempdir that matches crates/mcp/src/core/config.rs::McpConfig,
registering the mock server with builtin_type: image_generation
and response_format: image_generation_call.
- gateway_with_mock_mcp: launches an OpenAI cloud gateway with
--mcp-config-path pointing at that tempfile; yields
(gateway, client, mock_mcp_server). Skips if OPENAI_API_KEY is
absent.
- image_gen_tool_args: canonical tool payload shared by all tests
so size/quality overrides have a clean starting point.
* e2e_test/responses/test_image_generation.py — TestImageGeneration
class with four tests on the OpenAI cloud backend:
1. test_image_generation_non_streaming: verifies the response
carries an ImageGenerationCall output item with the id ig_
prefix, status=completed, result=<mock base64>, and
revised_prompt echoing the input (R6.2 output-item emission).
2. test_image_generation_streaming: verifies the
response.image_generation_call.{in_progress,generating,
completed} events fire in the documented order; partial_image
is asserted optional-but-ordered-correctly (R6.2/R6.3/R6.4
streaming contract).
3. test_image_generation_tool_overrides_size: pins
size=512x512 and quality=high on the tool payload and asserts
the mock observed those exact values via last_call_args. This
catches compactor regressions where the override pipeline
would silently drop or rewrite user-provided arguments.
4. test_image_generation_compactor_strips_base64: creates a
stored conversation, fetches /v1/conversations/{id}/items, and
asserts the raw base64 bytes do NOT appear in persisted state
— so multi-turn replay never re-ships the image to the model
(R6.1 compactor behavior).
Engine matrix: openai only for now. skip_for_runtime("sglang") and
skip_for_runtime("vllm") guard the gRPC lanes until R6.3 and R6.4
stabilise in CI; a follow-up PR can drop those skip decorators once
the local-worker lanes are green.
Other changes:
* e2e_test/pyproject.toml — add mcp>=1.0 and uvicorn to the e2e test
deps. The mcp SDK is required by the mock server; uvicorn was
already pulled transitively but is called directly from the mock
server module and so we declare it explicitly.
Manual verification:
* ruff check e2e_test/responses/test_image_generation.py \
e2e_test/infra/mock_mcp_server.py e2e_test/infra/mock_mcp.py \
e2e_test/responses/conftest.py — clean.
* mypy on the same files with --ignore-missing-imports — clean.
* pytest e2e_test/responses/test_image_generation.py --collect-only —
4 tests discovered, no import errors.
* In-process smoke test of MockMcpServer via the mcp streamable_http
client confirmed the tool registration, deterministic payload, and
last_call_args introspection work end-to-end.
Full gated run (live gateway + OpenAI key) is deferred to CI.
Refs: R6.1 #1355, R6.2 #1356, R6.3 #1359, R6.4 #1358 (all merged)
Co-authored-by: Tingting Zhou <zhoutt96@gmail.com>
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
Respond to the coordinator directive on PR #1365: remove the R6.3/R6.4 ``skip_for_runtime`` decorators, add local-gRPC fixtures for both engines, and harden the assertions on what the gateway emits. R6.1-R6.4 all merged. Real lane failures surfaced by this suite will be filed as R6.6/R6.7 follow-ups rather than patched from R6.5, so the CI signal is genuine rather than papered over. Engine matrix ------------- Three test classes, sharing a ``_ImageGenerationAssertions`` mix-in so the cloud + gRPC lanes stay strictly in lock-step: * ``TestImageGenerationCloud`` (vendor=openai, gpu=0) — OpenAI cloud backend with gpt-5-nano, mock MCP server standing in for the image backend. Exercises the OpenAI-compat router (R6.2). * ``TestImageGenerationGrpcSglang`` (engine=sglang, gpu=1, model=openai/gpt-oss-20b) — local SGLang worker via harmony. Exercises R6.3 gRPC-harmony wiring. * ``TestImageGenerationGrpcVllm`` (engine=vllm, gpu=1, model=meta-llama/Llama-3.1-8B-Instruct) — local vLLM worker via regular. Exercises R6.4 gRPC-regular wiring. Each lane's fixture wires its gateway at the same shared in-process ``MockMcpServer`` so deterministic base64-roundtrip / size-override assertions stay valid across engines. Fixture additions (``e2e_test/responses/conftest.py``) ------------------------------------------------------ * ``gateway_with_mock_mcp_cloud`` replaces the old ``gateway_with_mock_mcp``. Yields ``(gateway, client, mock, model)`` with ``model="gpt-5-nano"``. The old name is kept as a backward-compat alias so any unmerged branch referencing it still works. * ``_start_local_grpc_gateway_with_mcp`` helper: launches one gRPC worker for a given engine + model_id, spins up a gateway with ``--mcp-config-path`` pointing at the shared mock MCP config, wraps client instantiation in try/except so a failed init doesn't leak the gateway or the worker. * ``gateway_with_mock_mcp_grpc_sglang`` / ``gateway_with_mock_mcp_grpc_vllm``: class-scoped fixtures built on the helper. Each skips (not fails) when the gRPC worker can't start — CI lanes without GPUs otherwise poison every engine-parametrized suite. Harder assertions (``test_image_generation.py``) ------------------------------------------------ * ``_assert_image_generation_call_item`` now asserts every documented field (``type``, ``id`` ig-prefix, ``status``, ``result``, ``revised_prompt``) rather than just a subset. * ``_assert_streaming_envelope`` validates the full emitted envelope: * ``response.created`` → ``response.output_item.added`` → ``response.image_generation_call.{in_progress, generating, [partial_image], completed}`` → ``response.output_item.done`` → ``response.completed`` * each event's first occurrence strictly precedes the next required event's first occurrence * exactly one ``response.created`` / ``response.completed`` / ``output_item.added`` / ``output_item.done`` per image_gen call * ``sequence_number`` strictly monotonically increasing with no gaps > 1 (only checked when every event has one) * optional ``partial_image`` sits between ``generating`` and ``completed`` when present * Streaming and non-streaming paths share the same mix-in body, so the gRPC lanes exercise the full assertion surface without duplication. Local gates (ruff check/format, mypy, pytest --collect-only) clean; 12 tests collected (3 classes × 4 tests). Refs: R6.1 #1355, R6.2 #1356, R6.3 #1359, R6.4 #1358 (all merged) Co-authored-by: Tingting Zhou <zhoutt96@gmail.com> Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
Introduce e2e_test/infra/mock_mcp_server.py — a thin wrapper around the official MCP SDK's FastMCP base class that exposes streamable-HTTP transport on a local port. The existing MCP e2e test (e2e_test/messages/test_mcp_tool.py) depends on a live Brave MCP server which is flaky, slow, and non-deterministic; routing-layer tests for the gateway's built-in tool plumbing need an in-process, deterministic counterpart. This mock supplies one. Files added: * e2e_test/infra/mock_mcp_server.py — MockMcpServer class. Runs FastMCP via uvicorn on a background thread; auto-allocates a free port when the caller passes port=0. Registers an image_generation tool that returns a hard-coded 1x1 transparent PNG (92-char base64) plus the prompt echoed as revised_prompt, so tests can make byte-for-byte assertions. Records every call into call_log and surfaces the last call via last_call_args for override-verification tests. TODO stubs for web_search, file_search, code_interpreter sit inline with the same shape so future R6.x PRs just uncomment and adjust. * e2e_test/infra/mock_mcp.py — mock_mcp_server session-scoped pytest fixture plus IMAGE_GENERATION_PNG_BASE64 re-export. Files modified: * e2e_test/infra/constants.py — new MOCK_MCP_HOST constant (default 127.0.0.1) paralleling the existing BRAVE_MCP_HOST. * e2e_test/infra/__init__.py — export MockMcpServer, mock_mcp_server, IMAGE_GENERATION_PNG_BASE64, MOCK_MCP_HOST. Why streamable HTTP: matches the protocol the gateway's MCP client already speaks against Brave in production (see crates/mcp/src/core/config.rs::McpTransport::Streamable). The streamable_http_app() FastMCP method yields a /mcp Starlette route mountable directly under uvicorn. How to extend: FastMCP exposes a decorator-driven registration API; add a new @fastmcp.tool in MockMcpServer._register_tools, append to self._call_log inside the body for introspection, and you're done. See the module docstring for the extension recipe. Refs: R6.1 #1355, R6.2 #1356, R6.3 #1359, R6.4 #1358 (all merged) Co-authored-by: Tingting Zhou <zhoutt96@gmail.com> Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
Add the first end-to-end tests for the image_generation built-in tool,
using the in-process MockMcpServer introduced in the previous commit.
R6.1-R6.4 wired image_generation across the openai, gRPC-regular, and
gRPC-harmony routers end-to-end; this PR provides the first automated
coverage that exercises those paths without external-service
dependencies.
Files added:
* e2e_test/responses/conftest.py — fixture module for the Responses
suite. Provides:
- mock_mcp_config_file: writes a session-scoped YAML config to a
tempdir that matches crates/mcp/src/core/config.rs::McpConfig,
registering the mock server with builtin_type: image_generation
and response_format: image_generation_call.
- gateway_with_mock_mcp: launches an OpenAI cloud gateway with
--mcp-config-path pointing at that tempfile; yields
(gateway, client, mock_mcp_server). Skips if OPENAI_API_KEY is
absent.
- image_gen_tool_args: canonical tool payload shared by all tests
so size/quality overrides have a clean starting point.
* e2e_test/responses/test_image_generation.py — TestImageGeneration
class with four tests on the OpenAI cloud backend:
1. test_image_generation_non_streaming: verifies the response
carries an ImageGenerationCall output item with the id ig_
prefix, status=completed, result=<mock base64>, and
revised_prompt echoing the input (R6.2 output-item emission).
2. test_image_generation_streaming: verifies the
response.image_generation_call.{in_progress,generating,
completed} events fire in the documented order; partial_image
is asserted optional-but-ordered-correctly (R6.2/R6.3/R6.4
streaming contract).
3. test_image_generation_tool_overrides_size: pins
size=512x512 and quality=high on the tool payload and asserts
the mock observed those exact values via last_call_args. This
catches compactor regressions where the override pipeline
would silently drop or rewrite user-provided arguments.
4. test_image_generation_compactor_strips_base64: creates a
stored conversation, fetches /v1/conversations/{id}/items, and
asserts the raw base64 bytes do NOT appear in persisted state
— so multi-turn replay never re-ships the image to the model
(R6.1 compactor behavior).
Engine matrix: openai only for now. skip_for_runtime("sglang") and
skip_for_runtime("vllm") guard the gRPC lanes until R6.3 and R6.4
stabilise in CI; a follow-up PR can drop those skip decorators once
the local-worker lanes are green.
Other changes:
* e2e_test/pyproject.toml — add mcp>=1.0 and uvicorn to the e2e test
deps. The mcp SDK is required by the mock server; uvicorn was
already pulled transitively but is called directly from the mock
server module and so we declare it explicitly.
Manual verification:
* ruff check e2e_test/responses/test_image_generation.py \
e2e_test/infra/mock_mcp_server.py e2e_test/infra/mock_mcp.py \
e2e_test/responses/conftest.py — clean.
* mypy on the same files with --ignore-missing-imports — clean.
* pytest e2e_test/responses/test_image_generation.py --collect-only —
4 tests discovered, no import errors.
* In-process smoke test of MockMcpServer via the mcp streamable_http
client confirmed the tool registration, deterministic payload, and
last_call_args introspection work end-to-end.
Full gated run (live gateway + OpenAI key) is deferred to CI.
Refs: R6.1 #1355, R6.2 #1356, R6.3 #1359, R6.4 #1358 (all merged)
Co-authored-by: Tingting Zhou <zhoutt96@gmail.com>
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
Respond to the coordinator directive on PR #1365: remove the R6.3/R6.4 ``skip_for_runtime`` decorators, add local-gRPC fixtures for both engines, and harden the assertions on what the gateway emits. R6.1-R6.4 all merged. Real lane failures surfaced by this suite will be filed as R6.6/R6.7 follow-ups rather than patched from R6.5, so the CI signal is genuine rather than papered over. Engine matrix ------------- Three test classes, sharing a ``_ImageGenerationAssertions`` mix-in so the cloud + gRPC lanes stay strictly in lock-step: * ``TestImageGenerationCloud`` (vendor=openai, gpu=0) — OpenAI cloud backend with gpt-5-nano, mock MCP server standing in for the image backend. Exercises the OpenAI-compat router (R6.2). * ``TestImageGenerationGrpcSglang`` (engine=sglang, gpu=1, model=openai/gpt-oss-20b) — local SGLang worker via harmony. Exercises R6.3 gRPC-harmony wiring. * ``TestImageGenerationGrpcVllm`` (engine=vllm, gpu=1, model=meta-llama/Llama-3.1-8B-Instruct) — local vLLM worker via regular. Exercises R6.4 gRPC-regular wiring. Each lane's fixture wires its gateway at the same shared in-process ``MockMcpServer`` so deterministic base64-roundtrip / size-override assertions stay valid across engines. Fixture additions (``e2e_test/responses/conftest.py``) ------------------------------------------------------ * ``gateway_with_mock_mcp_cloud`` replaces the old ``gateway_with_mock_mcp``. Yields ``(gateway, client, mock, model)`` with ``model="gpt-5-nano"``. The old name is kept as a backward-compat alias so any unmerged branch referencing it still works. * ``_start_local_grpc_gateway_with_mcp`` helper: launches one gRPC worker for a given engine + model_id, spins up a gateway with ``--mcp-config-path`` pointing at the shared mock MCP config, wraps client instantiation in try/except so a failed init doesn't leak the gateway or the worker. * ``gateway_with_mock_mcp_grpc_sglang`` / ``gateway_with_mock_mcp_grpc_vllm``: class-scoped fixtures built on the helper. Each skips (not fails) when the gRPC worker can't start — CI lanes without GPUs otherwise poison every engine-parametrized suite. Harder assertions (``test_image_generation.py``) ------------------------------------------------ * ``_assert_image_generation_call_item`` now asserts every documented field (``type``, ``id`` ig-prefix, ``status``, ``result``, ``revised_prompt``) rather than just a subset. * ``_assert_streaming_envelope`` validates the full emitted envelope: * ``response.created`` → ``response.output_item.added`` → ``response.image_generation_call.{in_progress, generating, [partial_image], completed}`` → ``response.output_item.done`` → ``response.completed`` * each event's first occurrence strictly precedes the next required event's first occurrence * exactly one ``response.created`` / ``response.completed`` / ``output_item.added`` / ``output_item.done`` per image_gen call * ``sequence_number`` strictly monotonically increasing with no gaps > 1 (only checked when every event has one) * optional ``partial_image`` sits between ``generating`` and ``completed`` when present * Streaming and non-streaming paths share the same mix-in body, so the gRPC lanes exercise the full assertion surface without duplication. Local gates (ruff check/format, mypy, pytest --collect-only) clean; 12 tests collected (3 classes × 4 tests). Refs: R6.1 #1355, R6.2 #1356, R6.3 #1359, R6.4 #1358 (all merged) Co-authored-by: Tingting Zhou <zhoutt96@gmail.com> Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
Introduce e2e_test/infra/mock_mcp_server.py — a thin wrapper around the official MCP SDK's FastMCP base class that exposes streamable-HTTP transport on a local port. The existing MCP e2e test (e2e_test/messages/test_mcp_tool.py) depends on a live Brave MCP server which is flaky, slow, and non-deterministic; routing-layer tests for the gateway's built-in tool plumbing need an in-process, deterministic counterpart. This mock supplies one. Files added: * e2e_test/infra/mock_mcp_server.py — MockMcpServer class. Runs FastMCP via uvicorn on a background thread; auto-allocates a free port when the caller passes port=0. Registers an image_generation tool that returns a hard-coded 1x1 transparent PNG (92-char base64) plus the prompt echoed as revised_prompt, so tests can make byte-for-byte assertions. Records every call into call_log and surfaces the last call via last_call_args for override-verification tests. TODO stubs for web_search, file_search, code_interpreter sit inline with the same shape so future R6.x PRs just uncomment and adjust. * e2e_test/infra/mock_mcp.py — mock_mcp_server session-scoped pytest fixture plus IMAGE_GENERATION_PNG_BASE64 re-export. Files modified: * e2e_test/infra/constants.py — new MOCK_MCP_HOST constant (default 127.0.0.1) paralleling the existing BRAVE_MCP_HOST. * e2e_test/infra/__init__.py — export MockMcpServer, mock_mcp_server, IMAGE_GENERATION_PNG_BASE64, MOCK_MCP_HOST. Why streamable HTTP: matches the protocol the gateway's MCP client already speaks against Brave in production (see crates/mcp/src/core/config.rs::McpTransport::Streamable). The streamable_http_app() FastMCP method yields a /mcp Starlette route mountable directly under uvicorn. How to extend: FastMCP exposes a decorator-driven registration API; add a new @fastmcp.tool in MockMcpServer._register_tools, append to self._call_log inside the body for introspection, and you're done. See the module docstring for the extension recipe. Refs: R6.1 #1355, R6.2 #1356, R6.3 #1359, R6.4 #1358 (all merged) Co-authored-by: Tingting Zhou <zhoutt96@gmail.com> Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
Add the first end-to-end tests for the image_generation built-in tool,
using the in-process MockMcpServer introduced in the previous commit.
R6.1-R6.4 wired image_generation across the openai, gRPC-regular, and
gRPC-harmony routers end-to-end; this PR provides the first automated
coverage that exercises those paths without external-service
dependencies.
Files added:
* e2e_test/responses/conftest.py — fixture module for the Responses
suite. Provides:
- mock_mcp_config_file: writes a session-scoped YAML config to a
tempdir that matches crates/mcp/src/core/config.rs::McpConfig,
registering the mock server with builtin_type: image_generation
and response_format: image_generation_call.
- gateway_with_mock_mcp: launches an OpenAI cloud gateway with
--mcp-config-path pointing at that tempfile; yields
(gateway, client, mock_mcp_server). Skips if OPENAI_API_KEY is
absent.
- image_gen_tool_args: canonical tool payload shared by all tests
so size/quality overrides have a clean starting point.
* e2e_test/responses/test_image_generation.py — TestImageGeneration
class with four tests on the OpenAI cloud backend:
1. test_image_generation_non_streaming: verifies the response
carries an ImageGenerationCall output item with the id ig_
prefix, status=completed, result=<mock base64>, and
revised_prompt echoing the input (R6.2 output-item emission).
2. test_image_generation_streaming: verifies the
response.image_generation_call.{in_progress,generating,
completed} events fire in the documented order; partial_image
is asserted optional-but-ordered-correctly (R6.2/R6.3/R6.4
streaming contract).
3. test_image_generation_tool_overrides_size: pins
size=512x512 and quality=high on the tool payload and asserts
the mock observed those exact values via last_call_args. This
catches compactor regressions where the override pipeline
would silently drop or rewrite user-provided arguments.
4. test_image_generation_compactor_strips_base64: creates a
stored conversation, fetches /v1/conversations/{id}/items, and
asserts the raw base64 bytes do NOT appear in persisted state
— so multi-turn replay never re-ships the image to the model
(R6.1 compactor behavior).
Engine matrix: openai only for now. skip_for_runtime("sglang") and
skip_for_runtime("vllm") guard the gRPC lanes until R6.3 and R6.4
stabilise in CI; a follow-up PR can drop those skip decorators once
the local-worker lanes are green.
Other changes:
* e2e_test/pyproject.toml — add mcp>=1.0 and uvicorn to the e2e test
deps. The mcp SDK is required by the mock server; uvicorn was
already pulled transitively but is called directly from the mock
server module and so we declare it explicitly.
Manual verification:
* ruff check e2e_test/responses/test_image_generation.py \
e2e_test/infra/mock_mcp_server.py e2e_test/infra/mock_mcp.py \
e2e_test/responses/conftest.py — clean.
* mypy on the same files with --ignore-missing-imports — clean.
* pytest e2e_test/responses/test_image_generation.py --collect-only —
4 tests discovered, no import errors.
* In-process smoke test of MockMcpServer via the mcp streamable_http
client confirmed the tool registration, deterministic payload, and
last_call_args introspection work end-to-end.
Full gated run (live gateway + OpenAI key) is deferred to CI.
Refs: R6.1 #1355, R6.2 #1356, R6.3 #1359, R6.4 #1358 (all merged)
Co-authored-by: Tingting Zhou <zhoutt96@gmail.com>
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
Respond to the coordinator directive on PR #1365: remove the R6.3/R6.4 ``skip_for_runtime`` decorators, add local-gRPC fixtures for both engines, and harden the assertions on what the gateway emits. R6.1-R6.4 all merged. Real lane failures surfaced by this suite will be filed as R6.6/R6.7 follow-ups rather than patched from R6.5, so the CI signal is genuine rather than papered over. Engine matrix ------------- Three test classes, sharing a ``_ImageGenerationAssertions`` mix-in so the cloud + gRPC lanes stay strictly in lock-step: * ``TestImageGenerationCloud`` (vendor=openai, gpu=0) — OpenAI cloud backend with gpt-5-nano, mock MCP server standing in for the image backend. Exercises the OpenAI-compat router (R6.2). * ``TestImageGenerationGrpcSglang`` (engine=sglang, gpu=1, model=openai/gpt-oss-20b) — local SGLang worker via harmony. Exercises R6.3 gRPC-harmony wiring. * ``TestImageGenerationGrpcVllm`` (engine=vllm, gpu=1, model=meta-llama/Llama-3.1-8B-Instruct) — local vLLM worker via regular. Exercises R6.4 gRPC-regular wiring. Each lane's fixture wires its gateway at the same shared in-process ``MockMcpServer`` so deterministic base64-roundtrip / size-override assertions stay valid across engines. Fixture additions (``e2e_test/responses/conftest.py``) ------------------------------------------------------ * ``gateway_with_mock_mcp_cloud`` replaces the old ``gateway_with_mock_mcp``. Yields ``(gateway, client, mock, model)`` with ``model="gpt-5-nano"``. The old name is kept as a backward-compat alias so any unmerged branch referencing it still works. * ``_start_local_grpc_gateway_with_mcp`` helper: launches one gRPC worker for a given engine + model_id, spins up a gateway with ``--mcp-config-path`` pointing at the shared mock MCP config, wraps client instantiation in try/except so a failed init doesn't leak the gateway or the worker. * ``gateway_with_mock_mcp_grpc_sglang`` / ``gateway_with_mock_mcp_grpc_vllm``: class-scoped fixtures built on the helper. Each skips (not fails) when the gRPC worker can't start — CI lanes without GPUs otherwise poison every engine-parametrized suite. Harder assertions (``test_image_generation.py``) ------------------------------------------------ * ``_assert_image_generation_call_item`` now asserts every documented field (``type``, ``id`` ig-prefix, ``status``, ``result``, ``revised_prompt``) rather than just a subset. * ``_assert_streaming_envelope`` validates the full emitted envelope: * ``response.created`` → ``response.output_item.added`` → ``response.image_generation_call.{in_progress, generating, [partial_image], completed}`` → ``response.output_item.done`` → ``response.completed`` * each event's first occurrence strictly precedes the next required event's first occurrence * exactly one ``response.created`` / ``response.completed`` / ``output_item.added`` / ``output_item.done`` per image_gen call * ``sequence_number`` strictly monotonically increasing with no gaps > 1 (only checked when every event has one) * optional ``partial_image`` sits between ``generating`` and ``completed`` when present * Streaming and non-streaming paths share the same mix-in body, so the gRPC lanes exercise the full assertion surface without duplication. Local gates (ruff check/format, mypy, pytest --collect-only) clean; 12 tests collected (3 classes × 4 tests). Refs: R6.1 #1355, R6.2 #1356, R6.3 #1359, R6.4 #1358 (all merged) Co-authored-by: Tingting Zhou <zhoutt96@gmail.com> Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
Introduce e2e_test/infra/mock_mcp_server.py — a thin wrapper around the official MCP SDK's FastMCP base class that exposes streamable-HTTP transport on a local port. The existing MCP e2e test (e2e_test/messages/test_mcp_tool.py) depends on a live Brave MCP server which is flaky, slow, and non-deterministic; routing-layer tests for the gateway's built-in tool plumbing need an in-process, deterministic counterpart. This mock supplies one. Files added: * e2e_test/infra/mock_mcp_server.py — MockMcpServer class. Runs FastMCP via uvicorn on a background thread; auto-allocates a free port when the caller passes port=0. Registers an image_generation tool that returns a hard-coded 1x1 transparent PNG (92-char base64) plus the prompt echoed as revised_prompt, so tests can make byte-for-byte assertions. Records every call into call_log and surfaces the last call via last_call_args for override-verification tests. TODO stubs for web_search, file_search, code_interpreter sit inline with the same shape so future R6.x PRs just uncomment and adjust. * e2e_test/infra/mock_mcp.py — mock_mcp_server session-scoped pytest fixture plus IMAGE_GENERATION_PNG_BASE64 re-export. Files modified: * e2e_test/infra/constants.py — new MOCK_MCP_HOST constant (default 127.0.0.1) paralleling the existing BRAVE_MCP_HOST. * e2e_test/infra/__init__.py — export MockMcpServer, mock_mcp_server, IMAGE_GENERATION_PNG_BASE64, MOCK_MCP_HOST. Why streamable HTTP: matches the protocol the gateway's MCP client already speaks against Brave in production (see crates/mcp/src/core/config.rs::McpTransport::Streamable). The streamable_http_app() FastMCP method yields a /mcp Starlette route mountable directly under uvicorn. How to extend: FastMCP exposes a decorator-driven registration API; add a new @fastmcp.tool in MockMcpServer._register_tools, append to self._call_log inside the body for introspection, and you're done. See the module docstring for the extension recipe. Refs: R6.1 #1355, R6.2 #1356, R6.3 #1359, R6.4 #1358 (all merged) Co-authored-by: Tingting Zhou <zhoutt96@gmail.com> Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
Add the first end-to-end tests for the image_generation built-in tool,
using the in-process MockMcpServer introduced in the previous commit.
R6.1-R6.4 wired image_generation across the openai, gRPC-regular, and
gRPC-harmony routers end-to-end; this PR provides the first automated
coverage that exercises those paths without external-service
dependencies.
Files added:
* e2e_test/responses/conftest.py — fixture module for the Responses
suite. Provides:
- mock_mcp_config_file: writes a session-scoped YAML config to a
tempdir that matches crates/mcp/src/core/config.rs::McpConfig,
registering the mock server with builtin_type: image_generation
and response_format: image_generation_call.
- gateway_with_mock_mcp: launches an OpenAI cloud gateway with
--mcp-config-path pointing at that tempfile; yields
(gateway, client, mock_mcp_server). Skips if OPENAI_API_KEY is
absent.
- image_gen_tool_args: canonical tool payload shared by all tests
so size/quality overrides have a clean starting point.
* e2e_test/responses/test_image_generation.py — TestImageGeneration
class with four tests on the OpenAI cloud backend:
1. test_image_generation_non_streaming: verifies the response
carries an ImageGenerationCall output item with the id ig_
prefix, status=completed, result=<mock base64>, and
revised_prompt echoing the input (R6.2 output-item emission).
2. test_image_generation_streaming: verifies the
response.image_generation_call.{in_progress,generating,
completed} events fire in the documented order; partial_image
is asserted optional-but-ordered-correctly (R6.2/R6.3/R6.4
streaming contract).
3. test_image_generation_tool_overrides_size: pins
size=512x512 and quality=high on the tool payload and asserts
the mock observed those exact values via last_call_args. This
catches compactor regressions where the override pipeline
would silently drop or rewrite user-provided arguments.
4. test_image_generation_compactor_strips_base64: creates a
stored conversation, fetches /v1/conversations/{id}/items, and
asserts the raw base64 bytes do NOT appear in persisted state
— so multi-turn replay never re-ships the image to the model
(R6.1 compactor behavior).
Engine matrix: openai only for now. skip_for_runtime("sglang") and
skip_for_runtime("vllm") guard the gRPC lanes until R6.3 and R6.4
stabilise in CI; a follow-up PR can drop those skip decorators once
the local-worker lanes are green.
Other changes:
* e2e_test/pyproject.toml — add mcp>=1.0 and uvicorn to the e2e test
deps. The mcp SDK is required by the mock server; uvicorn was
already pulled transitively but is called directly from the mock
server module and so we declare it explicitly.
Manual verification:
* ruff check e2e_test/responses/test_image_generation.py \
e2e_test/infra/mock_mcp_server.py e2e_test/infra/mock_mcp.py \
e2e_test/responses/conftest.py — clean.
* mypy on the same files with --ignore-missing-imports — clean.
* pytest e2e_test/responses/test_image_generation.py --collect-only —
4 tests discovered, no import errors.
* In-process smoke test of MockMcpServer via the mcp streamable_http
client confirmed the tool registration, deterministic payload, and
last_call_args introspection work end-to-end.
Full gated run (live gateway + OpenAI key) is deferred to CI.
Refs: R6.1 #1355, R6.2 #1356, R6.3 #1359, R6.4 #1358 (all merged)
Co-authored-by: Tingting Zhou <zhoutt96@gmail.com>
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
Respond to the coordinator directive on PR #1365: remove the R6.3/R6.4 ``skip_for_runtime`` decorators, add local-gRPC fixtures for both engines, and harden the assertions on what the gateway emits. R6.1-R6.4 all merged. Real lane failures surfaced by this suite will be filed as R6.6/R6.7 follow-ups rather than patched from R6.5, so the CI signal is genuine rather than papered over. Engine matrix ------------- Three test classes, sharing a ``_ImageGenerationAssertions`` mix-in so the cloud + gRPC lanes stay strictly in lock-step: * ``TestImageGenerationCloud`` (vendor=openai, gpu=0) — OpenAI cloud backend with gpt-5-nano, mock MCP server standing in for the image backend. Exercises the OpenAI-compat router (R6.2). * ``TestImageGenerationGrpcSglang`` (engine=sglang, gpu=1, model=openai/gpt-oss-20b) — local SGLang worker via harmony. Exercises R6.3 gRPC-harmony wiring. * ``TestImageGenerationGrpcVllm`` (engine=vllm, gpu=1, model=meta-llama/Llama-3.1-8B-Instruct) — local vLLM worker via regular. Exercises R6.4 gRPC-regular wiring. Each lane's fixture wires its gateway at the same shared in-process ``MockMcpServer`` so deterministic base64-roundtrip / size-override assertions stay valid across engines. Fixture additions (``e2e_test/responses/conftest.py``) ------------------------------------------------------ * ``gateway_with_mock_mcp_cloud`` replaces the old ``gateway_with_mock_mcp``. Yields ``(gateway, client, mock, model)`` with ``model="gpt-5-nano"``. The old name is kept as a backward-compat alias so any unmerged branch referencing it still works. * ``_start_local_grpc_gateway_with_mcp`` helper: launches one gRPC worker for a given engine + model_id, spins up a gateway with ``--mcp-config-path`` pointing at the shared mock MCP config, wraps client instantiation in try/except so a failed init doesn't leak the gateway or the worker. * ``gateway_with_mock_mcp_grpc_sglang`` / ``gateway_with_mock_mcp_grpc_vllm``: class-scoped fixtures built on the helper. Each skips (not fails) when the gRPC worker can't start — CI lanes without GPUs otherwise poison every engine-parametrized suite. Harder assertions (``test_image_generation.py``) ------------------------------------------------ * ``_assert_image_generation_call_item`` now asserts every documented field (``type``, ``id`` ig-prefix, ``status``, ``result``, ``revised_prompt``) rather than just a subset. * ``_assert_streaming_envelope`` validates the full emitted envelope: * ``response.created`` → ``response.output_item.added`` → ``response.image_generation_call.{in_progress, generating, [partial_image], completed}`` → ``response.output_item.done`` → ``response.completed`` * each event's first occurrence strictly precedes the next required event's first occurrence * exactly one ``response.created`` / ``response.completed`` / ``output_item.added`` / ``output_item.done`` per image_gen call * ``sequence_number`` strictly monotonically increasing with no gaps > 1 (only checked when every event has one) * optional ``partial_image`` sits between ``generating`` and ``completed`` when present * Streaming and non-streaming paths share the same mix-in body, so the gRPC lanes exercise the full assertion surface without duplication. Local gates (ruff check/format, mypy, pytest --collect-only) clean; 12 tests collected (3 classes × 4 tests). Refs: R6.1 #1355, R6.2 #1356, R6.3 #1359, R6.4 #1358 (all merged) Co-authored-by: Tingting Zhou <zhoutt96@gmail.com> Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
Description
Problem
R6.1 (#1355) landed the shared
ResponseFormat::ImageGenerationCallplumbing — the variant, theBuiltinToolType::ImageGenerationconfig, theImageGenerationCallEventprotocol constants, theto_image_generation_callMCP transformer, and the base64-stripping compactor — but left the OpenAI Responses router itself without real per-site wiring.image_generationMCP tool calls therefore fell through to the genericmcp_callpath inmodel_gateway/src/routers/openai/responses/streaming.rsand emitted the wrong event type to clients.The three sites in
model_gateway/src/routers/openai/mcp/tool_loop.rsthat R6.1 introduced were marked with// stubbed:comments deferring final wiring to R6.2, even though the underlying values (ImageGenerationCallEvent::GENERATING, theig_id prefix, and the sharedfc_/call_strip behavior) already matched the pattern used by the other three hosted-tool formats.Solution
Finish the per-router wiring so
image_generationMCP tool calls emit the fullresponse.image_generation_call.*event sequence through the OpenAI Responses streaming pipeline, matching the shape already in place forweb_search/code_interpreter/file_search.Changes
Protocol-only router wiring under
model_gateway/src/routers/openai/:model_gateway/src/routers/openai/responses/streaming.rsImageGenerationCallEventfromopenai_protocol::event_types.apply_event_transformations_inplace: extend theResponseFormatmatch so upstreamfunction_callitems whose sessionresponse_formatisImageGenerationCallare rewritten toItemType::IMAGE_GENERATION_CALLwith theig_id prefix, instead of falling through to the genericmcp_callarm. Mirrors the existingWebSearchCallarm.maybe_inject_tool_in_progress: add anItemType::IMAGE_GENERATION_CALLarm that emitsImageGenerationCallEvent::IN_PROGRESSso the in_progress event is injected right after the rewrittenoutput_item.addedevent. Parallel to the existingWEB_SEARCH_CALL/CODE_INTERPRETER_CALL/FILE_SEARCH_CALLarms.model_gateway/src/routers/openai/mcp/tool_loop.rs// stubbed:marker comments left by R6.1 on theImageGenerationCallarms insend_tool_call_intermediate_event,stable_streaming_tool_item_id, andnon_streaming_tool_item_id_source. The underlying values already matched the other three hosted-tool formats and now stand as the real wiring.generatingis the coarse intermediate event emitted bytool_loop, on par withsearchingfor web/file search andinterpretingfor code;partial_imagepreview events are the responsibility of the underlying tool when it streams chunks, not thetool_looppath.End-to-end event sequence for
image_generationresponse.output_item.addedwith afunction_callitem → rewritten toimage_generation_callwith anig_*id.response.image_generation_call.in_progressis injected right after.send_tool_call_intermediate_eventemitsresponse.image_generation_call.generatingwhile the tool runs.send_tool_call_completion_eventsemitsresponse.image_generation_call.completedplusresponse.output_item.donewith the fully-transformed item from the sharedto_image_generation_calltransformer.Matches the 4-event sequence documented in
.claude/_audit/openai-responses-api-spec.md§tools (image_generation) and mirrored by OpenAI Python SDK v2.8.1.Scope
partial_imagepreview events. The tool_loop path emits the coarsein_progress → generating → completedsequence; inline preview chunks are produced by the underlying tool when it streams them.Test plan
cargo check -p smg --lib— clean.cargo test -p smg --lib— 616 passed, 0 failed, 4 ignored.cargo fmt --all— no changes.cargo clippy -p smg --lib --tests -- -D warnings— clean.ResponseFormat::ImageGenerationCallsite undermodel_gateway/src/routers/openai/is now real wiring (no remaining// stubbed:markers or R6.2 TODOs).Refs: R6.1 #1355 (prerequisite)
Checklist
Summary by CodeRabbit
New Features
Documentation