Skip to content

feat(openai): wire image_generation streaming events in Responses router (R6.2) - #1356

Merged
slin1237 merged 2 commits into
mainfrom
feat/r6-02-openai-image-generation
Apr 23, 2026
Merged

slin1237 merged 2 commits into
mainfrom
feat/r6-02-openai-image-generation

Conversation

@slin1237

@slin1237 slin1237 commented Apr 23, 2026 •

Copy link
Copy Markdown
Member

Description

Problem

R6.1 (#1355) landed the shared ResponseFormat::ImageGenerationCall plumbing — the variant, the BuiltinToolType::ImageGeneration config, the ImageGenerationCallEvent protocol constants, the to_image_generation_call MCP transformer, and the base64-stripping compactor — but left the OpenAI Responses router itself without real per-site wiring. image_generation MCP tool calls therefore fell through to the generic mcp_call path in model_gateway/src/routers/openai/responses/streaming.rs and emitted the wrong event type to clients.

The three sites in model_gateway/src/routers/openai/mcp/tool_loop.rs that R6.1 introduced were marked with // stubbed: comments deferring final wiring to R6.2, even though the underlying values (ImageGenerationCallEvent::GENERATING, the ig_ id prefix, and the shared fc_/call_ strip behavior) already matched the pattern used by the other three hosted-tool formats.

Solution

Finish the per-router wiring so image_generation MCP tool calls emit the full response.image_generation_call.* event sequence through the OpenAI Responses streaming pipeline, matching the shape already in place for web_search / code_interpreter / file_search.

Changes

Protocol-only router wiring under model_gateway/src/routers/openai/:

model_gateway/src/routers/openai/responses/streaming.rs

  • Import ImageGenerationCallEvent from openai_protocol::event_types.
  • apply_event_transformations_inplace: extend the ResponseFormat match so upstream function_call items whose session response_format is ImageGenerationCall are rewritten to ItemType::IMAGE_GENERATION_CALL with the ig_ id prefix, instead of falling through to the generic mcp_call arm. Mirrors the existing WebSearchCall arm.
  • maybe_inject_tool_in_progress: add an ItemType::IMAGE_GENERATION_CALL arm that emits ImageGenerationCallEvent::IN_PROGRESS so the in_progress event is injected right after the rewritten output_item.added event. Parallel to the existing WEB_SEARCH_CALL / CODE_INTERPRETER_CALL / FILE_SEARCH_CALL arms.

model_gateway/src/routers/openai/mcp/tool_loop.rs

  • Drop the three // stubbed: marker comments left by R6.1 on the ImageGenerationCall arms in send_tool_call_intermediate_event, stable_streaming_tool_item_id, and non_streaming_tool_item_id_source. The underlying values already matched the other three hosted-tool formats and now stand as the real wiring.
  • Replace the R6.2-TODO comment on the intermediate event with a short design note: generating is the coarse intermediate event emitted by tool_loop, on par with searching for web/file search and interpreting for code; partial_image preview events are the responsibility of the underlying tool when it streams chunks, not the tool_loop path.

End-to-end event sequence for image_generation

  1. Upstream emits response.output_item.added with a function_call item → rewritten to image_generation_call with an ig_* id.
  2. response.image_generation_call.in_progress is injected right after.
  3. send_tool_call_intermediate_event emits response.image_generation_call.generating while the tool runs.
  4. send_tool_call_completion_events emits response.image_generation_call.completed plus response.output_item.done with the fully-transformed item from the shared to_image_generation_call transformer.

Matches the 4-event sequence documented in .claude/_audit/openai-responses-api-spec.md §tools (image_generation) and mirrored by OpenAI Python SDK v2.8.1.

Scope

  • In scope: OpenAI Responses streaming router wiring only.
  • Out of scope: non-OpenAI transports — the gRPC and HTTP worker paths are owned by R6.3 / R6.4 and are deliberately not touched here.
  • Out of scope: partial_image preview events. The tool_loop path emits the coarse in_progress → generating → completed sequence; inline preview chunks are produced by the underlying tool when it streams them.

Test plan

  • cargo check -p smg --lib — clean.
  • cargo test -p smg --lib — 616 passed, 0 failed, 4 ignored.
  • cargo fmt --all — no changes.
  • cargo clippy -p smg --lib --tests -- -D warnings — clean.
  • Verified every ResponseFormat::ImageGenerationCall site under model_gateway/src/routers/openai/ is now real wiring (no remaining // stubbed: markers or R6.2 TODOs).

Refs: R6.1 #1355 (prerequisite)

Checklist
  • Documentation updated
  • (Optional) Please join us on Slack #sig-smg to discuss, review, and merge PRs

Summary by CodeRabbit

  • New Features

    • Improved streaming event support for image-generation tool calls, including in-progress events, consistent event sequencing, and standardized ID prefix handling.
  • Documentation

    • Clarified tool-call event definitions to treat image generation as a first-class streamed format and removed outdated references.

…ter (R6.2)

R6.1 (#1355) landed the shared `ResponseFormat::ImageGenerationCall`
plumbing and left the OpenAI Responses router with placeholder marker
comments (no real per-site wiring in `responses/streaming.rs`). R6.2
finishes the per-router wiring so `image_generation` MCP tool calls
emit the full `response.image_generation_call.*` event sequence through
the OpenAI Responses streaming pipeline, matching the shape used for
web_search / code_interpreter / file_search.

Changes in `model_gateway/src/routers/openai/responses/streaming.rs`:
- import `ImageGenerationCallEvent` from `openai_protocol::event_types`.
- `apply_event_transformations_inplace`: extend the `ResponseFormat`
  match so upstream `function_call` items whose session `response_format`
  is `ImageGenerationCall` are rewritten to `ItemType::IMAGE_GENERATION_CALL`
  with the `ig_` id prefix, instead of falling through to the generic
  `mcp_call` arm. This mirrors the existing `WebSearchCall` arm.
- `maybe_inject_tool_in_progress`: add an `ItemType::IMAGE_GENERATION_CALL`
  arm that emits `ImageGenerationCallEvent::IN_PROGRESS`, so the in_progress
  event is injected after the rewritten `output_item.added` (parallel to the
  existing `WEB_SEARCH_CALL` / `CODE_INTERPRETER_CALL` / `FILE_SEARCH_CALL`
  arms).

Changes in `model_gateway/src/routers/openai/mcp/tool_loop.rs`:
- Drop the temporary "stubbed" marker comments left by R6.1 on the three
  `ImageGenerationCall` arms (`send_tool_call_intermediate_event`,
  `stable_streaming_tool_item_id`, `non_streaming_tool_item_id_source`).
  The underlying values (`ImageGenerationCallEvent::GENERATING`, `ig_`
  prefix, shared `fc_`/`call_` strip behavior) already matched the other
  three hosted-tool formats and now stand as the real wiring.
- Replace the R6.2-TODO comment on the intermediate event with a short
  design note: `generating` is the coarse intermediate event emitted by
  the tool_loop; `partial_image` preview events are the responsibility of
  the underlying tool when it streams chunks, not the tool_loop itself.

End-to-end flow for `image_generation`:
1. upstream emits `output_item.added` with a `function_call` item →
   rewritten to `image_generation_call` with an `ig_*` id.
2. `response.image_generation_call.in_progress` is injected right after.
3. `send_tool_call_intermediate_event` emits
   `response.image_generation_call.generating` while the tool runs.
4. `send_tool_call_completion_events` emits
   `response.image_generation_call.completed` and
   `response.output_item.done` with the fully-transformed item from the
   shared `to_image_generation_call` transformer.

Scope: protocol-only router wiring under `model_gateway/src/routers/openai/`.
Non-OpenAI transports (gRPC, HTTP workers) are handled by R6.3 / R6.4 and
are intentionally not touched here.

Gates:
- `cargo check -p smg --lib` clean
- `cargo test -p smg --lib` 616 passed, 0 failed
- `cargo fmt --all`
- `cargo clippy -p smg --lib --tests -- -D warnings` clean

Refs: R6.1 #1355 (prerequisite)

Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
Co-authored-by: Tingting Zhou <zhoutt96@gmail.com>
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
@coderabbitai

coderabbitai Bot commented Apr 23, 2026 •

Copy link
Copy Markdown

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 927f95eb-f091-41fd-8df1-05b7af2de56b

📥 Commits

Reviewing files that changed from the base of the PR and between 452e495 and c0dfc00.

📒 Files selected for processing (2)
  • model_gateway/src/routers/openai/mcp/tool_loop.rs
  • model_gateway/src/routers/openai/responses/streaming.rs

📝 Walkthrough

Walkthrough

Adds streaming support for ImageGenerationCall: maps ResponseFormat::ImageGenerationCall to ItemType::IMAGE_GENERATION_CALL with an "ig_" ID prefix, updates in-progress injection to emit ImageGenerationCallEvent::IN_PROGRESS, and clarifies related tool-loop documentation and naming comments.

Changes

Cohort / File(s) Summary
Tool Loop Documentation
model_gateway/src/routers/openai/mcp/tool_loop.rs
Clarifies that image_generation_call is a first-class streamed format; documents generating as the image-generation intermediate event and that partial_image events originate from the underlying tool stream; updates comments on ig_ prefix and removes outdated references.
Streaming Event Support
model_gateway/src/routers/openai/responses/streaming.rs
Maps ResponseFormat::ImageGenerationCall → ItemType::IMAGE_GENERATION_CALL, rewrites IDs with "ig_" prefix, and extends maybe_inject_tool_in_progress to emit ImageGenerationCallEvent::IN_PROGRESS for image-generation items; imports updated accordingly.

Sequence Diagram(s)

sequenceDiagram
    participant Client as Client
    participant Router as OpenAI Router
    participant Tool as Image Tool Stream
    participant MCP as MCP Tool Loop

    Client->>Router: request with image_generation_call
    Router->>MCP: create tool-call item (type: IMAGE_GENERATION_CALL, id:"ig_*")
    MCP->>Tool: start image generation stream
    Tool-->>MCP: partial_image events (streamed)
    MCP-->>Router: send intermediate 'generating' event
    Router-->>Client: emit ImageGenerationCallEvent::IN_PROGRESS (with ig_ id)
    Tool-->>MCP: final image result
    MCP-->>Router: completion events
    Router-->>Client: final completion events
Loading

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~10 minutes

Possibly related PRs

Suggested labels

mcp

Suggested reviewers

  • CatherineSue
  • key4ng
  • zhaowenzi

Poem

🐰 I nibble prefixes, "ig_" so neat,
I hop through streams where partial images meet,
generating hums as pixels take flight,
A rabbit cheers for events sent right! 📸✨

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately describes the main change: wiring image_generation streaming events in the Responses router, a key missing piece from R6.1.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/r6-02-openai-image-generation

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions github-actions Bot added model-gateway Model gateway crate changes openai OpenAI router changes labels Apr 23, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request integrates ImageGenerationCall into the OpenAI response streaming and tool loop logic, handling event transformations, ID prefixing, and intermediate event emission. A review comment suggests extending the transformation logic to explicitly handle CodeInterpreterCall and FileSearchCall for consistency across all hosted tool formats.

Comment on lines +153 to +155
ResponseFormat::ImageGenerationCall => {
(ItemType::IMAGE_GENERATION_CALL, "ig_")
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The apply_event_transformations_inplace function is missing transformation logic for the CodeInterpreterCall and FileSearchCall response formats. While WebSearchCall and the newly added ImageGenerationCall are handled, these other built-in tool formats will currently fall through to the generic MCP_CALL type with an mcp_ ID prefix. To ensure consistency across all hosted tool formats as intended by this PR, they should be explicitly mapped to their respective item types and ID prefixes (ci_ and fs_). This ensures that the transformation logic correctly identifies the item variants present in the final response output.

                                ResponseFormat::CodeInterpreterCall => {
                                    (ItemType::CODE_INTERPRETER_CALL, "ci_")
                                }
                                ResponseFormat::FileSearchCall => (ItemType::FILE_SEARCH_CALL, "fs_"),
                                ResponseFormat::ImageGenerationCall => {
                                    (ItemType::IMAGE_GENERATION_CALL, "ig_")
                                }
References
  1. When processing response items, ensure the logic targets the specific item variants that are actually present in the final response output.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch on the gap — you're right that CodeInterpreterCall and FileSearchCall currently fall through to (ItemType::MCP_CALL, "mcp_") in this match. That predates R6.2; the existing WebSearchCall arm is the only explicit hosted-tool branch today. R6.1 / R6.2 are scoped to adding the image_generation_call path to match web_search_call, not to retrofit the other two formats.

Extending this match to also cover CodeInterpreterCall and FileSearchCall is a straightforward cleanup but touches behavior for two unrelated tool types, so it belongs in a follow-up rather than in the R6.x series (which is staged around the image_generation tool specifically). I'd rather keep this PR's blast radius narrow and file a separate PR that handles all the remaining hosted-tool formats uniformly.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 452e49572c

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment on lines +153 to +155
ResponseFormat::ImageGenerationCall => {
(ItemType::IMAGE_GENERATION_CALL, "ig_")
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Align image-generation argument event IDs

Rewriting ResponseFormat::ImageGenerationCall to image_generation_call with an ig_ item ID here causes a protocol mismatch because argument events in this same module are still emitted as response.mcp_call_arguments.* with mcp_* IDs (FunctionCallEvent::ARGUMENTS_DONE and send_buffered_arguments). For image-generation tool calls, clients now receive one call under two different item_id namespaces (ig_* vs mcp_*), which breaks correlation logic that joins argument deltas/done to the corresponding output item.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the P1 flag. To unpack it: the response.mcp_call_arguments.* events in this module (via FunctionCallEvent::ARGUMENTS_DONE and send_buffered_arguments) always rewrite item_id via mcp_response_item_id and stamp the type as McpEvent::CALL_ARGUMENTS_DONE regardless of response_format. That behavior predates R6.2 and applies uniformly to all three existing hosted-tool formats (web_search_call with ws_*, code_interpreter_call with ci_*, file_search_call with fs_*) — each of them already surfaces the same split-namespace pattern between the output_item's <prefix>_* id and the mcp_* argument-event item_id.

The R6.2 task is scoped to wiring image_generation_call through the same path as the other three hosted formats (the task instructions call out "stay consistent with the other 3 tools' shape; don't invent new patterns"). Fixing the argument-event namespace mismatch consistently for all four hosted formats — either by per-format id preservation or (more likely correct per OpenAI's spec) by suppressing mcp_call_arguments.* altogether for non-Passthrough formats — is a cross-cutting cleanup that belongs in a follow-up PR on top of R6.2/R6.3/R6.4 rather than a localized image-gen-only change that would itself diverge from the existing three-format shape.

Happy to pick that up as a follow-up once the R6.x series lands; leaving this PR narrowly scoped to image_generation wiring.

ItemType::WEB_SEARCH_CALL => WebSearchCallEvent::IN_PROGRESS,
ItemType::CODE_INTERPRETER_CALL => CodeInterpreterCallEvent::IN_PROGRESS,
ItemType::FILE_SEARCH_CALL => FileSearchCallEvent::IN_PROGRESS,
ItemType::IMAGE_GENERATION_CALL => ImageGenerationCallEvent::IN_PROGRESS,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Nit: The docstring for this function (line 419) enumerates the handled item types but doesn't include image_generation_call after this addition. Consider updating it to stay in sync:

/// Handles mcp_call, web_search_call, code_interpreter_call, file_search_call, and image_generation_call items.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in c0dfc00 — updated the docstring to include image_generation_call in the enumerated handled types, plus two parallel docstrings in tool_loop.rs (send_tool_call_completion_events and build_transformed_mcp_call_item) that had the same stale enumeration.

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Clean, well-scoped PR. The image_generation_call wiring in the Responses streaming router follows the established pattern (matches web_search_call). Comment updates in tool_loop.rs accurately replace the stale R6.2 stubs. One minor nit on a docstring enumeration — no blocking issues.

Address Claude PR Review nit on #1356: three helper docstrings
enumerated the handled tool-call output item types
(`mcp_call`, `web_search_call`, `code_interpreter_call`,
`file_search_call`) but were not updated when R6.2 added
`image_generation_call`. Add `image_generation_call` to the enumeration
in:

- `model_gateway/src/routers/openai/responses/streaming.rs`
  (`maybe_inject_tool_in_progress`)
- `model_gateway/src/routers/openai/mcp/tool_loop.rs`
  (`send_tool_call_completion_events`, `build_transformed_mcp_call_item`)

Docs-only change; no behavior difference.

Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
Co-authored-by: Tingting Zhou <zhoutt96@gmail.com>
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: c0dfc0024b

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment on lines +153 to +155
ResponseFormat::ImageGenerationCall => {
(ItemType::IMAGE_GENERATION_CALL, "ig_")
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Keep image_generation item IDs consistent across tool events

Fresh evidence in this commit: apply_event_transformations_inplace now rewrites image-generation response.output_item.* entries to image_generation_call with an ig_ prefix, but this module still emits argument events as response.mcp_call_arguments.* and normalizes item_id with mcp_response_item_id(...) (mcp_*). For image-generation tool calls, clients that correlate argument and lifecycle events by item_id will now see two namespaces for the same call (ig_* vs mcp_*), which breaks joining argument delta/done to the corresponding added/in_progress/completed events.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a re-post of the same P1 from your previous review. See my reply on the original thread: #1356 (comment) — the response.mcp_call_arguments.* + mcp_response_item_id behavior is pre-existing across all three other hosted-tool formats (web_search_call, code_interpreter_call, file_search_call), not a regression introduced here. Fixing it consistently across all four formats is a separate cross-cutting cleanup tracked as a follow-up to the R6.x series.

@slin1237
slin1237 merged commit 4ac08a7 into main Apr 23, 2026
57 checks passed
@slin1237
slin1237 deleted the feat/r6-02-openai-image-generation branch April 23, 2026 19:26
slin1237 added a commit that referenced this pull request Apr 23, 2026
Respond to the coordinator directive on PR #1365: remove the R6.3/R6.4
``skip_for_runtime`` decorators, add local-gRPC fixtures for both
engines, and harden the assertions on what the gateway emits.

R6.1-R6.4 all merged. Real lane failures surfaced by this suite will be
filed as R6.6/R6.7 follow-ups rather than patched from R6.5, so the CI
signal is genuine rather than papered over.

Engine matrix
-------------

Three test classes, sharing a ``_ImageGenerationAssertions`` mix-in so
the cloud + gRPC lanes stay strictly in lock-step:

* ``TestImageGenerationCloud`` (vendor=openai, gpu=0) — OpenAI cloud
  backend with gpt-5-nano, mock MCP server standing in for the image
  backend. Exercises the OpenAI-compat router (R6.2).
* ``TestImageGenerationGrpcSglang`` (engine=sglang, gpu=1,
  model=openai/gpt-oss-20b) — local SGLang worker via harmony. Exercises
  R6.3 gRPC-harmony wiring.
* ``TestImageGenerationGrpcVllm`` (engine=vllm, gpu=1,
  model=meta-llama/Llama-3.1-8B-Instruct) — local vLLM worker via regular.
  Exercises R6.4 gRPC-regular wiring.

Each lane's fixture wires its gateway at the same shared in-process
``MockMcpServer`` so deterministic base64-roundtrip / size-override
assertions stay valid across engines.

Fixture additions (``e2e_test/responses/conftest.py``)
------------------------------------------------------

* ``gateway_with_mock_mcp_cloud`` replaces the old
  ``gateway_with_mock_mcp``. Yields ``(gateway, client, mock, model)``
  with ``model="gpt-5-nano"``. The old name is kept as a backward-compat
  alias so any unmerged branch referencing it still works.
* ``_start_local_grpc_gateway_with_mcp`` helper: launches one gRPC worker
  for a given engine + model_id, spins up a gateway with
  ``--mcp-config-path`` pointing at the shared mock MCP config, wraps
  client instantiation in try/except so a failed init doesn't leak the
  gateway or the worker.
* ``gateway_with_mock_mcp_grpc_sglang`` / ``gateway_with_mock_mcp_grpc_vllm``:
  class-scoped fixtures built on the helper. Each skips (not fails) when
  the gRPC worker can't start — CI lanes without GPUs otherwise poison
  every engine-parametrized suite.

Harder assertions (``test_image_generation.py``)
------------------------------------------------

* ``_assert_image_generation_call_item`` now asserts every documented
  field (``type``, ``id`` ig-prefix, ``status``, ``result``,
  ``revised_prompt``) rather than just a subset.
* ``_assert_streaming_envelope`` validates the full emitted envelope:
    * ``response.created`` → ``response.output_item.added`` →
      ``response.image_generation_call.{in_progress, generating,
      [partial_image], completed}`` → ``response.output_item.done`` →
      ``response.completed``
    * each event's first occurrence strictly precedes the next required
      event's first occurrence
    * exactly one ``response.created`` / ``response.completed`` /
      ``output_item.added`` / ``output_item.done`` per image_gen call
    * ``sequence_number`` strictly monotonically increasing with no
      gaps > 1 (only checked when every event has one)
    * optional ``partial_image`` sits between ``generating`` and
      ``completed`` when present
* Streaming and non-streaming paths share the same mix-in body, so the
  gRPC lanes exercise the full assertion surface without duplication.

Local gates (ruff check/format, mypy, pytest --collect-only) clean;
12 tests collected (3 classes × 4 tests).

Refs: R6.1 #1355, R6.2 #1356, R6.3 #1359, R6.4 #1358 (all merged)

Co-authored-by: Tingting Zhou <zhoutt96@gmail.com>
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
slin1237 added a commit that referenced this pull request Apr 24, 2026
Introduce e2e_test/infra/mock_mcp_server.py — a thin wrapper around the
official MCP SDK's FastMCP base class that exposes streamable-HTTP
transport on a local port. The existing MCP e2e test
(e2e_test/messages/test_mcp_tool.py) depends on a live Brave MCP server
which is flaky, slow, and non-deterministic; routing-layer tests for
the gateway's built-in tool plumbing need an in-process, deterministic
counterpart. This mock supplies one.

Files added:

* e2e_test/infra/mock_mcp_server.py — MockMcpServer class. Runs FastMCP
  via uvicorn on a background thread; auto-allocates a free port when
  the caller passes port=0. Registers an image_generation tool that
  returns a hard-coded 1x1 transparent PNG (92-char base64) plus the
  prompt echoed as revised_prompt, so tests can make byte-for-byte
  assertions. Records every call into call_log and surfaces the last
  call via last_call_args for override-verification tests. TODO stubs
  for web_search, file_search, code_interpreter sit inline with the
  same shape so future R6.x PRs just uncomment and adjust.

* e2e_test/infra/mock_mcp.py — mock_mcp_server session-scoped pytest
  fixture plus IMAGE_GENERATION_PNG_BASE64 re-export.

Files modified:

* e2e_test/infra/constants.py — new MOCK_MCP_HOST constant (default
  127.0.0.1) paralleling the existing BRAVE_MCP_HOST.

* e2e_test/infra/__init__.py — export MockMcpServer, mock_mcp_server,
  IMAGE_GENERATION_PNG_BASE64, MOCK_MCP_HOST.

Why streamable HTTP: matches the protocol the gateway's MCP client
already speaks against Brave in production (see
crates/mcp/src/core/config.rs::McpTransport::Streamable). The
streamable_http_app() FastMCP method yields a /mcp Starlette route
mountable directly under uvicorn.

How to extend: FastMCP exposes a decorator-driven registration API;
add a new @fastmcp.tool in MockMcpServer._register_tools, append to
self._call_log inside the body for introspection, and you're done.
See the module docstring for the extension recipe.

Refs: R6.1 #1355, R6.2 #1356, R6.3 #1359, R6.4 #1358 (all merged)

Co-authored-by: Tingting Zhou <zhoutt96@gmail.com>
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
slin1237 added a commit that referenced this pull request Apr 24, 2026
Add the first end-to-end tests for the image_generation built-in tool,
using the in-process MockMcpServer introduced in the previous commit.
R6.1-R6.4 wired image_generation across the openai, gRPC-regular, and
gRPC-harmony routers end-to-end; this PR provides the first automated
coverage that exercises those paths without external-service
dependencies.

Files added:

* e2e_test/responses/conftest.py — fixture module for the Responses
  suite. Provides:
    - mock_mcp_config_file: writes a session-scoped YAML config to a
      tempdir that matches crates/mcp/src/core/config.rs::McpConfig,
      registering the mock server with builtin_type: image_generation
      and response_format: image_generation_call.
    - gateway_with_mock_mcp: launches an OpenAI cloud gateway with
      --mcp-config-path pointing at that tempfile; yields
      (gateway, client, mock_mcp_server). Skips if OPENAI_API_KEY is
      absent.
    - image_gen_tool_args: canonical tool payload shared by all tests
      so size/quality overrides have a clean starting point.

* e2e_test/responses/test_image_generation.py — TestImageGeneration
  class with four tests on the OpenAI cloud backend:

    1. test_image_generation_non_streaming: verifies the response
       carries an ImageGenerationCall output item with the id ig_
       prefix, status=completed, result=<mock base64>, and
       revised_prompt echoing the input (R6.2 output-item emission).

    2. test_image_generation_streaming: verifies the
       response.image_generation_call.{in_progress,generating,
       completed} events fire in the documented order; partial_image
       is asserted optional-but-ordered-correctly (R6.2/R6.3/R6.4
       streaming contract).

    3. test_image_generation_tool_overrides_size: pins
       size=512x512 and quality=high on the tool payload and asserts
       the mock observed those exact values via last_call_args. This
       catches compactor regressions where the override pipeline
       would silently drop or rewrite user-provided arguments.

    4. test_image_generation_compactor_strips_base64: creates a
       stored conversation, fetches /v1/conversations/{id}/items, and
       asserts the raw base64 bytes do NOT appear in persisted state
       — so multi-turn replay never re-ships the image to the model
       (R6.1 compactor behavior).

Engine matrix: openai only for now. skip_for_runtime("sglang") and
skip_for_runtime("vllm") guard the gRPC lanes until R6.3 and R6.4
stabilise in CI; a follow-up PR can drop those skip decorators once
the local-worker lanes are green.

Other changes:

* e2e_test/pyproject.toml — add mcp>=1.0 and uvicorn to the e2e test
  deps. The mcp SDK is required by the mock server; uvicorn was
  already pulled transitively but is called directly from the mock
  server module and so we declare it explicitly.

Manual verification:

* ruff check e2e_test/responses/test_image_generation.py \
    e2e_test/infra/mock_mcp_server.py e2e_test/infra/mock_mcp.py \
    e2e_test/responses/conftest.py — clean.
* mypy on the same files with --ignore-missing-imports — clean.
* pytest e2e_test/responses/test_image_generation.py --collect-only —
  4 tests discovered, no import errors.
* In-process smoke test of MockMcpServer via the mcp streamable_http
  client confirmed the tool registration, deterministic payload, and
  last_call_args introspection work end-to-end.

Full gated run (live gateway + OpenAI key) is deferred to CI.

Refs: R6.1 #1355, R6.2 #1356, R6.3 #1359, R6.4 #1358 (all merged)

Co-authored-by: Tingting Zhou <zhoutt96@gmail.com>
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
slin1237 added a commit that referenced this pull request Apr 24, 2026
Respond to the coordinator directive on PR #1365: remove the R6.3/R6.4
``skip_for_runtime`` decorators, add local-gRPC fixtures for both
engines, and harden the assertions on what the gateway emits.

R6.1-R6.4 all merged. Real lane failures surfaced by this suite will be
filed as R6.6/R6.7 follow-ups rather than patched from R6.5, so the CI
signal is genuine rather than papered over.

Engine matrix
-------------

Three test classes, sharing a ``_ImageGenerationAssertions`` mix-in so
the cloud + gRPC lanes stay strictly in lock-step:

* ``TestImageGenerationCloud`` (vendor=openai, gpu=0) — OpenAI cloud
  backend with gpt-5-nano, mock MCP server standing in for the image
  backend. Exercises the OpenAI-compat router (R6.2).
* ``TestImageGenerationGrpcSglang`` (engine=sglang, gpu=1,
  model=openai/gpt-oss-20b) — local SGLang worker via harmony. Exercises
  R6.3 gRPC-harmony wiring.
* ``TestImageGenerationGrpcVllm`` (engine=vllm, gpu=1,
  model=meta-llama/Llama-3.1-8B-Instruct) — local vLLM worker via regular.
  Exercises R6.4 gRPC-regular wiring.

Each lane's fixture wires its gateway at the same shared in-process
``MockMcpServer`` so deterministic base64-roundtrip / size-override
assertions stay valid across engines.

Fixture additions (``e2e_test/responses/conftest.py``)
------------------------------------------------------

* ``gateway_with_mock_mcp_cloud`` replaces the old
  ``gateway_with_mock_mcp``. Yields ``(gateway, client, mock, model)``
  with ``model="gpt-5-nano"``. The old name is kept as a backward-compat
  alias so any unmerged branch referencing it still works.
* ``_start_local_grpc_gateway_with_mcp`` helper: launches one gRPC worker
  for a given engine + model_id, spins up a gateway with
  ``--mcp-config-path`` pointing at the shared mock MCP config, wraps
  client instantiation in try/except so a failed init doesn't leak the
  gateway or the worker.
* ``gateway_with_mock_mcp_grpc_sglang`` / ``gateway_with_mock_mcp_grpc_vllm``:
  class-scoped fixtures built on the helper. Each skips (not fails) when
  the gRPC worker can't start — CI lanes without GPUs otherwise poison
  every engine-parametrized suite.

Harder assertions (``test_image_generation.py``)
------------------------------------------------

* ``_assert_image_generation_call_item`` now asserts every documented
  field (``type``, ``id`` ig-prefix, ``status``, ``result``,
  ``revised_prompt``) rather than just a subset.
* ``_assert_streaming_envelope`` validates the full emitted envelope:
    * ``response.created`` → ``response.output_item.added`` →
      ``response.image_generation_call.{in_progress, generating,
      [partial_image], completed}`` → ``response.output_item.done`` →
      ``response.completed``
    * each event's first occurrence strictly precedes the next required
      event's first occurrence
    * exactly one ``response.created`` / ``response.completed`` /
      ``output_item.added`` / ``output_item.done`` per image_gen call
    * ``sequence_number`` strictly monotonically increasing with no
      gaps > 1 (only checked when every event has one)
    * optional ``partial_image`` sits between ``generating`` and
      ``completed`` when present
* Streaming and non-streaming paths share the same mix-in body, so the
  gRPC lanes exercise the full assertion surface without duplication.

Local gates (ruff check/format, mypy, pytest --collect-only) clean;
12 tests collected (3 classes × 4 tests).

Refs: R6.1 #1355, R6.2 #1356, R6.3 #1359, R6.4 #1358 (all merged)

Co-authored-by: Tingting Zhou <zhoutt96@gmail.com>
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
slin1237 added a commit that referenced this pull request Apr 24, 2026
Introduce e2e_test/infra/mock_mcp_server.py — a thin wrapper around the
official MCP SDK's FastMCP base class that exposes streamable-HTTP
transport on a local port. The existing MCP e2e test
(e2e_test/messages/test_mcp_tool.py) depends on a live Brave MCP server
which is flaky, slow, and non-deterministic; routing-layer tests for
the gateway's built-in tool plumbing need an in-process, deterministic
counterpart. This mock supplies one.

Files added:

* e2e_test/infra/mock_mcp_server.py — MockMcpServer class. Runs FastMCP
  via uvicorn on a background thread; auto-allocates a free port when
  the caller passes port=0. Registers an image_generation tool that
  returns a hard-coded 1x1 transparent PNG (92-char base64) plus the
  prompt echoed as revised_prompt, so tests can make byte-for-byte
  assertions. Records every call into call_log and surfaces the last
  call via last_call_args for override-verification tests. TODO stubs
  for web_search, file_search, code_interpreter sit inline with the
  same shape so future R6.x PRs just uncomment and adjust.

* e2e_test/infra/mock_mcp.py — mock_mcp_server session-scoped pytest
  fixture plus IMAGE_GENERATION_PNG_BASE64 re-export.

Files modified:

* e2e_test/infra/constants.py — new MOCK_MCP_HOST constant (default
  127.0.0.1) paralleling the existing BRAVE_MCP_HOST.

* e2e_test/infra/__init__.py — export MockMcpServer, mock_mcp_server,
  IMAGE_GENERATION_PNG_BASE64, MOCK_MCP_HOST.

Why streamable HTTP: matches the protocol the gateway's MCP client
already speaks against Brave in production (see
crates/mcp/src/core/config.rs::McpTransport::Streamable). The
streamable_http_app() FastMCP method yields a /mcp Starlette route
mountable directly under uvicorn.

How to extend: FastMCP exposes a decorator-driven registration API;
add a new @fastmcp.tool in MockMcpServer._register_tools, append to
self._call_log inside the body for introspection, and you're done.
See the module docstring for the extension recipe.

Refs: R6.1 #1355, R6.2 #1356, R6.3 #1359, R6.4 #1358 (all merged)

Co-authored-by: Tingting Zhou <zhoutt96@gmail.com>
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
slin1237 added a commit that referenced this pull request Apr 24, 2026
Add the first end-to-end tests for the image_generation built-in tool,
using the in-process MockMcpServer introduced in the previous commit.
R6.1-R6.4 wired image_generation across the openai, gRPC-regular, and
gRPC-harmony routers end-to-end; this PR provides the first automated
coverage that exercises those paths without external-service
dependencies.

Files added:

* e2e_test/responses/conftest.py — fixture module for the Responses
  suite. Provides:
    - mock_mcp_config_file: writes a session-scoped YAML config to a
      tempdir that matches crates/mcp/src/core/config.rs::McpConfig,
      registering the mock server with builtin_type: image_generation
      and response_format: image_generation_call.
    - gateway_with_mock_mcp: launches an OpenAI cloud gateway with
      --mcp-config-path pointing at that tempfile; yields
      (gateway, client, mock_mcp_server). Skips if OPENAI_API_KEY is
      absent.
    - image_gen_tool_args: canonical tool payload shared by all tests
      so size/quality overrides have a clean starting point.

* e2e_test/responses/test_image_generation.py — TestImageGeneration
  class with four tests on the OpenAI cloud backend:

    1. test_image_generation_non_streaming: verifies the response
       carries an ImageGenerationCall output item with the id ig_
       prefix, status=completed, result=<mock base64>, and
       revised_prompt echoing the input (R6.2 output-item emission).

    2. test_image_generation_streaming: verifies the
       response.image_generation_call.{in_progress,generating,
       completed} events fire in the documented order; partial_image
       is asserted optional-but-ordered-correctly (R6.2/R6.3/R6.4
       streaming contract).

    3. test_image_generation_tool_overrides_size: pins
       size=512x512 and quality=high on the tool payload and asserts
       the mock observed those exact values via last_call_args. This
       catches compactor regressions where the override pipeline
       would silently drop or rewrite user-provided arguments.

    4. test_image_generation_compactor_strips_base64: creates a
       stored conversation, fetches /v1/conversations/{id}/items, and
       asserts the raw base64 bytes do NOT appear in persisted state
       — so multi-turn replay never re-ships the image to the model
       (R6.1 compactor behavior).

Engine matrix: openai only for now. skip_for_runtime("sglang") and
skip_for_runtime("vllm") guard the gRPC lanes until R6.3 and R6.4
stabilise in CI; a follow-up PR can drop those skip decorators once
the local-worker lanes are green.

Other changes:

* e2e_test/pyproject.toml — add mcp>=1.0 and uvicorn to the e2e test
  deps. The mcp SDK is required by the mock server; uvicorn was
  already pulled transitively but is called directly from the mock
  server module and so we declare it explicitly.

Manual verification:

* ruff check e2e_test/responses/test_image_generation.py \
    e2e_test/infra/mock_mcp_server.py e2e_test/infra/mock_mcp.py \
    e2e_test/responses/conftest.py — clean.
* mypy on the same files with --ignore-missing-imports — clean.
* pytest e2e_test/responses/test_image_generation.py --collect-only —
  4 tests discovered, no import errors.
* In-process smoke test of MockMcpServer via the mcp streamable_http
  client confirmed the tool registration, deterministic payload, and
  last_call_args introspection work end-to-end.

Full gated run (live gateway + OpenAI key) is deferred to CI.

Refs: R6.1 #1355, R6.2 #1356, R6.3 #1359, R6.4 #1358 (all merged)

Co-authored-by: Tingting Zhou <zhoutt96@gmail.com>
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
slin1237 added a commit that referenced this pull request Apr 24, 2026
Respond to the coordinator directive on PR #1365: remove the R6.3/R6.4
``skip_for_runtime`` decorators, add local-gRPC fixtures for both
engines, and harden the assertions on what the gateway emits.

R6.1-R6.4 all merged. Real lane failures surfaced by this suite will be
filed as R6.6/R6.7 follow-ups rather than patched from R6.5, so the CI
signal is genuine rather than papered over.

Engine matrix
-------------

Three test classes, sharing a ``_ImageGenerationAssertions`` mix-in so
the cloud + gRPC lanes stay strictly in lock-step:

* ``TestImageGenerationCloud`` (vendor=openai, gpu=0) — OpenAI cloud
  backend with gpt-5-nano, mock MCP server standing in for the image
  backend. Exercises the OpenAI-compat router (R6.2).
* ``TestImageGenerationGrpcSglang`` (engine=sglang, gpu=1,
  model=openai/gpt-oss-20b) — local SGLang worker via harmony. Exercises
  R6.3 gRPC-harmony wiring.
* ``TestImageGenerationGrpcVllm`` (engine=vllm, gpu=1,
  model=meta-llama/Llama-3.1-8B-Instruct) — local vLLM worker via regular.
  Exercises R6.4 gRPC-regular wiring.

Each lane's fixture wires its gateway at the same shared in-process
``MockMcpServer`` so deterministic base64-roundtrip / size-override
assertions stay valid across engines.

Fixture additions (``e2e_test/responses/conftest.py``)
------------------------------------------------------

* ``gateway_with_mock_mcp_cloud`` replaces the old
  ``gateway_with_mock_mcp``. Yields ``(gateway, client, mock, model)``
  with ``model="gpt-5-nano"``. The old name is kept as a backward-compat
  alias so any unmerged branch referencing it still works.
* ``_start_local_grpc_gateway_with_mcp`` helper: launches one gRPC worker
  for a given engine + model_id, spins up a gateway with
  ``--mcp-config-path`` pointing at the shared mock MCP config, wraps
  client instantiation in try/except so a failed init doesn't leak the
  gateway or the worker.
* ``gateway_with_mock_mcp_grpc_sglang`` / ``gateway_with_mock_mcp_grpc_vllm``:
  class-scoped fixtures built on the helper. Each skips (not fails) when
  the gRPC worker can't start — CI lanes without GPUs otherwise poison
  every engine-parametrized suite.

Harder assertions (``test_image_generation.py``)
------------------------------------------------

* ``_assert_image_generation_call_item`` now asserts every documented
  field (``type``, ``id`` ig-prefix, ``status``, ``result``,
  ``revised_prompt``) rather than just a subset.
* ``_assert_streaming_envelope`` validates the full emitted envelope:
    * ``response.created`` → ``response.output_item.added`` →
      ``response.image_generation_call.{in_progress, generating,
      [partial_image], completed}`` → ``response.output_item.done`` →
      ``response.completed``
    * each event's first occurrence strictly precedes the next required
      event's first occurrence
    * exactly one ``response.created`` / ``response.completed`` /
      ``output_item.added`` / ``output_item.done`` per image_gen call
    * ``sequence_number`` strictly monotonically increasing with no
      gaps > 1 (only checked when every event has one)
    * optional ``partial_image`` sits between ``generating`` and
      ``completed`` when present
* Streaming and non-streaming paths share the same mix-in body, so the
  gRPC lanes exercise the full assertion surface without duplication.

Local gates (ruff check/format, mypy, pytest --collect-only) clean;
12 tests collected (3 classes × 4 tests).

Refs: R6.1 #1355, R6.2 #1356, R6.3 #1359, R6.4 #1358 (all merged)

Co-authored-by: Tingting Zhou <zhoutt96@gmail.com>
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
slin1237 added a commit that referenced this pull request Apr 24, 2026
Introduce e2e_test/infra/mock_mcp_server.py — a thin wrapper around the
official MCP SDK's FastMCP base class that exposes streamable-HTTP
transport on a local port. The existing MCP e2e test
(e2e_test/messages/test_mcp_tool.py) depends on a live Brave MCP server
which is flaky, slow, and non-deterministic; routing-layer tests for
the gateway's built-in tool plumbing need an in-process, deterministic
counterpart. This mock supplies one.

Files added:

* e2e_test/infra/mock_mcp_server.py — MockMcpServer class. Runs FastMCP
  via uvicorn on a background thread; auto-allocates a free port when
  the caller passes port=0. Registers an image_generation tool that
  returns a hard-coded 1x1 transparent PNG (92-char base64) plus the
  prompt echoed as revised_prompt, so tests can make byte-for-byte
  assertions. Records every call into call_log and surfaces the last
  call via last_call_args for override-verification tests. TODO stubs
  for web_search, file_search, code_interpreter sit inline with the
  same shape so future R6.x PRs just uncomment and adjust.

* e2e_test/infra/mock_mcp.py — mock_mcp_server session-scoped pytest
  fixture plus IMAGE_GENERATION_PNG_BASE64 re-export.

Files modified:

* e2e_test/infra/constants.py — new MOCK_MCP_HOST constant (default
  127.0.0.1) paralleling the existing BRAVE_MCP_HOST.

* e2e_test/infra/__init__.py — export MockMcpServer, mock_mcp_server,
  IMAGE_GENERATION_PNG_BASE64, MOCK_MCP_HOST.

Why streamable HTTP: matches the protocol the gateway's MCP client
already speaks against Brave in production (see
crates/mcp/src/core/config.rs::McpTransport::Streamable). The
streamable_http_app() FastMCP method yields a /mcp Starlette route
mountable directly under uvicorn.

How to extend: FastMCP exposes a decorator-driven registration API;
add a new @fastmcp.tool in MockMcpServer._register_tools, append to
self._call_log inside the body for introspection, and you're done.
See the module docstring for the extension recipe.

Refs: R6.1 #1355, R6.2 #1356, R6.3 #1359, R6.4 #1358 (all merged)

Co-authored-by: Tingting Zhou <zhoutt96@gmail.com>
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
slin1237 added a commit that referenced this pull request Apr 24, 2026
Add the first end-to-end tests for the image_generation built-in tool,
using the in-process MockMcpServer introduced in the previous commit.
R6.1-R6.4 wired image_generation across the openai, gRPC-regular, and
gRPC-harmony routers end-to-end; this PR provides the first automated
coverage that exercises those paths without external-service
dependencies.

Files added:

* e2e_test/responses/conftest.py — fixture module for the Responses
  suite. Provides:
    - mock_mcp_config_file: writes a session-scoped YAML config to a
      tempdir that matches crates/mcp/src/core/config.rs::McpConfig,
      registering the mock server with builtin_type: image_generation
      and response_format: image_generation_call.
    - gateway_with_mock_mcp: launches an OpenAI cloud gateway with
      --mcp-config-path pointing at that tempfile; yields
      (gateway, client, mock_mcp_server). Skips if OPENAI_API_KEY is
      absent.
    - image_gen_tool_args: canonical tool payload shared by all tests
      so size/quality overrides have a clean starting point.

* e2e_test/responses/test_image_generation.py — TestImageGeneration
  class with four tests on the OpenAI cloud backend:

    1. test_image_generation_non_streaming: verifies the response
       carries an ImageGenerationCall output item with the id ig_
       prefix, status=completed, result=<mock base64>, and
       revised_prompt echoing the input (R6.2 output-item emission).

    2. test_image_generation_streaming: verifies the
       response.image_generation_call.{in_progress,generating,
       completed} events fire in the documented order; partial_image
       is asserted optional-but-ordered-correctly (R6.2/R6.3/R6.4
       streaming contract).

    3. test_image_generation_tool_overrides_size: pins
       size=512x512 and quality=high on the tool payload and asserts
       the mock observed those exact values via last_call_args. This
       catches compactor regressions where the override pipeline
       would silently drop or rewrite user-provided arguments.

    4. test_image_generation_compactor_strips_base64: creates a
       stored conversation, fetches /v1/conversations/{id}/items, and
       asserts the raw base64 bytes do NOT appear in persisted state
       — so multi-turn replay never re-ships the image to the model
       (R6.1 compactor behavior).

Engine matrix: openai only for now. skip_for_runtime("sglang") and
skip_for_runtime("vllm") guard the gRPC lanes until R6.3 and R6.4
stabilise in CI; a follow-up PR can drop those skip decorators once
the local-worker lanes are green.

Other changes:

* e2e_test/pyproject.toml — add mcp>=1.0 and uvicorn to the e2e test
  deps. The mcp SDK is required by the mock server; uvicorn was
  already pulled transitively but is called directly from the mock
  server module and so we declare it explicitly.

Manual verification:

* ruff check e2e_test/responses/test_image_generation.py \
    e2e_test/infra/mock_mcp_server.py e2e_test/infra/mock_mcp.py \
    e2e_test/responses/conftest.py — clean.
* mypy on the same files with --ignore-missing-imports — clean.
* pytest e2e_test/responses/test_image_generation.py --collect-only —
  4 tests discovered, no import errors.
* In-process smoke test of MockMcpServer via the mcp streamable_http
  client confirmed the tool registration, deterministic payload, and
  last_call_args introspection work end-to-end.

Full gated run (live gateway + OpenAI key) is deferred to CI.

Refs: R6.1 #1355, R6.2 #1356, R6.3 #1359, R6.4 #1358 (all merged)

Co-authored-by: Tingting Zhou <zhoutt96@gmail.com>
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
slin1237 added a commit that referenced this pull request Apr 24, 2026
Respond to the coordinator directive on PR #1365: remove the R6.3/R6.4
``skip_for_runtime`` decorators, add local-gRPC fixtures for both
engines, and harden the assertions on what the gateway emits.

R6.1-R6.4 all merged. Real lane failures surfaced by this suite will be
filed as R6.6/R6.7 follow-ups rather than patched from R6.5, so the CI
signal is genuine rather than papered over.

Engine matrix
-------------

Three test classes, sharing a ``_ImageGenerationAssertions`` mix-in so
the cloud + gRPC lanes stay strictly in lock-step:

* ``TestImageGenerationCloud`` (vendor=openai, gpu=0) — OpenAI cloud
  backend with gpt-5-nano, mock MCP server standing in for the image
  backend. Exercises the OpenAI-compat router (R6.2).
* ``TestImageGenerationGrpcSglang`` (engine=sglang, gpu=1,
  model=openai/gpt-oss-20b) — local SGLang worker via harmony. Exercises
  R6.3 gRPC-harmony wiring.
* ``TestImageGenerationGrpcVllm`` (engine=vllm, gpu=1,
  model=meta-llama/Llama-3.1-8B-Instruct) — local vLLM worker via regular.
  Exercises R6.4 gRPC-regular wiring.

Each lane's fixture wires its gateway at the same shared in-process
``MockMcpServer`` so deterministic base64-roundtrip / size-override
assertions stay valid across engines.

Fixture additions (``e2e_test/responses/conftest.py``)
------------------------------------------------------

* ``gateway_with_mock_mcp_cloud`` replaces the old
  ``gateway_with_mock_mcp``. Yields ``(gateway, client, mock, model)``
  with ``model="gpt-5-nano"``. The old name is kept as a backward-compat
  alias so any unmerged branch referencing it still works.
* ``_start_local_grpc_gateway_with_mcp`` helper: launches one gRPC worker
  for a given engine + model_id, spins up a gateway with
  ``--mcp-config-path`` pointing at the shared mock MCP config, wraps
  client instantiation in try/except so a failed init doesn't leak the
  gateway or the worker.
* ``gateway_with_mock_mcp_grpc_sglang`` / ``gateway_with_mock_mcp_grpc_vllm``:
  class-scoped fixtures built on the helper. Each skips (not fails) when
  the gRPC worker can't start — CI lanes without GPUs otherwise poison
  every engine-parametrized suite.

Harder assertions (``test_image_generation.py``)
------------------------------------------------

* ``_assert_image_generation_call_item`` now asserts every documented
  field (``type``, ``id`` ig-prefix, ``status``, ``result``,
  ``revised_prompt``) rather than just a subset.
* ``_assert_streaming_envelope`` validates the full emitted envelope:
    * ``response.created`` → ``response.output_item.added`` →
      ``response.image_generation_call.{in_progress, generating,
      [partial_image], completed}`` → ``response.output_item.done`` →
      ``response.completed``
    * each event's first occurrence strictly precedes the next required
      event's first occurrence
    * exactly one ``response.created`` / ``response.completed`` /
      ``output_item.added`` / ``output_item.done`` per image_gen call
    * ``sequence_number`` strictly monotonically increasing with no
      gaps > 1 (only checked when every event has one)
    * optional ``partial_image`` sits between ``generating`` and
      ``completed`` when present
* Streaming and non-streaming paths share the same mix-in body, so the
  gRPC lanes exercise the full assertion surface without duplication.

Local gates (ruff check/format, mypy, pytest --collect-only) clean;
12 tests collected (3 classes × 4 tests).

Refs: R6.1 #1355, R6.2 #1356, R6.3 #1359, R6.4 #1358 (all merged)

Co-authored-by: Tingting Zhou <zhoutt96@gmail.com>
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
slin1237 added a commit that referenced this pull request Apr 24, 2026
Introduce e2e_test/infra/mock_mcp_server.py — a thin wrapper around the
official MCP SDK's FastMCP base class that exposes streamable-HTTP
transport on a local port. The existing MCP e2e test
(e2e_test/messages/test_mcp_tool.py) depends on a live Brave MCP server
which is flaky, slow, and non-deterministic; routing-layer tests for
the gateway's built-in tool plumbing need an in-process, deterministic
counterpart. This mock supplies one.

Files added:

* e2e_test/infra/mock_mcp_server.py — MockMcpServer class. Runs FastMCP
  via uvicorn on a background thread; auto-allocates a free port when
  the caller passes port=0. Registers an image_generation tool that
  returns a hard-coded 1x1 transparent PNG (92-char base64) plus the
  prompt echoed as revised_prompt, so tests can make byte-for-byte
  assertions. Records every call into call_log and surfaces the last
  call via last_call_args for override-verification tests. TODO stubs
  for web_search, file_search, code_interpreter sit inline with the
  same shape so future R6.x PRs just uncomment and adjust.

* e2e_test/infra/mock_mcp.py — mock_mcp_server session-scoped pytest
  fixture plus IMAGE_GENERATION_PNG_BASE64 re-export.

Files modified:

* e2e_test/infra/constants.py — new MOCK_MCP_HOST constant (default
  127.0.0.1) paralleling the existing BRAVE_MCP_HOST.

* e2e_test/infra/__init__.py — export MockMcpServer, mock_mcp_server,
  IMAGE_GENERATION_PNG_BASE64, MOCK_MCP_HOST.

Why streamable HTTP: matches the protocol the gateway's MCP client
already speaks against Brave in production (see
crates/mcp/src/core/config.rs::McpTransport::Streamable). The
streamable_http_app() FastMCP method yields a /mcp Starlette route
mountable directly under uvicorn.

How to extend: FastMCP exposes a decorator-driven registration API;
add a new @fastmcp.tool in MockMcpServer._register_tools, append to
self._call_log inside the body for introspection, and you're done.
See the module docstring for the extension recipe.

Refs: R6.1 #1355, R6.2 #1356, R6.3 #1359, R6.4 #1358 (all merged)

Co-authored-by: Tingting Zhou <zhoutt96@gmail.com>
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
slin1237 added a commit that referenced this pull request Apr 24, 2026
Add the first end-to-end tests for the image_generation built-in tool,
using the in-process MockMcpServer introduced in the previous commit.
R6.1-R6.4 wired image_generation across the openai, gRPC-regular, and
gRPC-harmony routers end-to-end; this PR provides the first automated
coverage that exercises those paths without external-service
dependencies.

Files added:

* e2e_test/responses/conftest.py — fixture module for the Responses
  suite. Provides:
    - mock_mcp_config_file: writes a session-scoped YAML config to a
      tempdir that matches crates/mcp/src/core/config.rs::McpConfig,
      registering the mock server with builtin_type: image_generation
      and response_format: image_generation_call.
    - gateway_with_mock_mcp: launches an OpenAI cloud gateway with
      --mcp-config-path pointing at that tempfile; yields
      (gateway, client, mock_mcp_server). Skips if OPENAI_API_KEY is
      absent.
    - image_gen_tool_args: canonical tool payload shared by all tests
      so size/quality overrides have a clean starting point.

* e2e_test/responses/test_image_generation.py — TestImageGeneration
  class with four tests on the OpenAI cloud backend:

    1. test_image_generation_non_streaming: verifies the response
       carries an ImageGenerationCall output item with the id ig_
       prefix, status=completed, result=<mock base64>, and
       revised_prompt echoing the input (R6.2 output-item emission).

    2. test_image_generation_streaming: verifies the
       response.image_generation_call.{in_progress,generating,
       completed} events fire in the documented order; partial_image
       is asserted optional-but-ordered-correctly (R6.2/R6.3/R6.4
       streaming contract).

    3. test_image_generation_tool_overrides_size: pins
       size=512x512 and quality=high on the tool payload and asserts
       the mock observed those exact values via last_call_args. This
       catches compactor regressions where the override pipeline
       would silently drop or rewrite user-provided arguments.

    4. test_image_generation_compactor_strips_base64: creates a
       stored conversation, fetches /v1/conversations/{id}/items, and
       asserts the raw base64 bytes do NOT appear in persisted state
       — so multi-turn replay never re-ships the image to the model
       (R6.1 compactor behavior).

Engine matrix: openai only for now. skip_for_runtime("sglang") and
skip_for_runtime("vllm") guard the gRPC lanes until R6.3 and R6.4
stabilise in CI; a follow-up PR can drop those skip decorators once
the local-worker lanes are green.

Other changes:

* e2e_test/pyproject.toml — add mcp>=1.0 and uvicorn to the e2e test
  deps. The mcp SDK is required by the mock server; uvicorn was
  already pulled transitively but is called directly from the mock
  server module and so we declare it explicitly.

Manual verification:

* ruff check e2e_test/responses/test_image_generation.py \
    e2e_test/infra/mock_mcp_server.py e2e_test/infra/mock_mcp.py \
    e2e_test/responses/conftest.py — clean.
* mypy on the same files with --ignore-missing-imports — clean.
* pytest e2e_test/responses/test_image_generation.py --collect-only —
  4 tests discovered, no import errors.
* In-process smoke test of MockMcpServer via the mcp streamable_http
  client confirmed the tool registration, deterministic payload, and
  last_call_args introspection work end-to-end.

Full gated run (live gateway + OpenAI key) is deferred to CI.

Refs: R6.1 #1355, R6.2 #1356, R6.3 #1359, R6.4 #1358 (all merged)

Co-authored-by: Tingting Zhou <zhoutt96@gmail.com>
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
slin1237 added a commit that referenced this pull request Apr 24, 2026
Respond to the coordinator directive on PR #1365: remove the R6.3/R6.4
``skip_for_runtime`` decorators, add local-gRPC fixtures for both
engines, and harden the assertions on what the gateway emits.

R6.1-R6.4 all merged. Real lane failures surfaced by this suite will be
filed as R6.6/R6.7 follow-ups rather than patched from R6.5, so the CI
signal is genuine rather than papered over.

Engine matrix
-------------

Three test classes, sharing a ``_ImageGenerationAssertions`` mix-in so
the cloud + gRPC lanes stay strictly in lock-step:

* ``TestImageGenerationCloud`` (vendor=openai, gpu=0) — OpenAI cloud
  backend with gpt-5-nano, mock MCP server standing in for the image
  backend. Exercises the OpenAI-compat router (R6.2).
* ``TestImageGenerationGrpcSglang`` (engine=sglang, gpu=1,
  model=openai/gpt-oss-20b) — local SGLang worker via harmony. Exercises
  R6.3 gRPC-harmony wiring.
* ``TestImageGenerationGrpcVllm`` (engine=vllm, gpu=1,
  model=meta-llama/Llama-3.1-8B-Instruct) — local vLLM worker via regular.
  Exercises R6.4 gRPC-regular wiring.

Each lane's fixture wires its gateway at the same shared in-process
``MockMcpServer`` so deterministic base64-roundtrip / size-override
assertions stay valid across engines.

Fixture additions (``e2e_test/responses/conftest.py``)
------------------------------------------------------

* ``gateway_with_mock_mcp_cloud`` replaces the old
  ``gateway_with_mock_mcp``. Yields ``(gateway, client, mock, model)``
  with ``model="gpt-5-nano"``. The old name is kept as a backward-compat
  alias so any unmerged branch referencing it still works.
* ``_start_local_grpc_gateway_with_mcp`` helper: launches one gRPC worker
  for a given engine + model_id, spins up a gateway with
  ``--mcp-config-path`` pointing at the shared mock MCP config, wraps
  client instantiation in try/except so a failed init doesn't leak the
  gateway or the worker.
* ``gateway_with_mock_mcp_grpc_sglang`` / ``gateway_with_mock_mcp_grpc_vllm``:
  class-scoped fixtures built on the helper. Each skips (not fails) when
  the gRPC worker can't start — CI lanes without GPUs otherwise poison
  every engine-parametrized suite.

Harder assertions (``test_image_generation.py``)
------------------------------------------------

* ``_assert_image_generation_call_item`` now asserts every documented
  field (``type``, ``id`` ig-prefix, ``status``, ``result``,
  ``revised_prompt``) rather than just a subset.
* ``_assert_streaming_envelope`` validates the full emitted envelope:
    * ``response.created`` → ``response.output_item.added`` →
      ``response.image_generation_call.{in_progress, generating,
      [partial_image], completed}`` → ``response.output_item.done`` →
      ``response.completed``
    * each event's first occurrence strictly precedes the next required
      event's first occurrence
    * exactly one ``response.created`` / ``response.completed`` /
      ``output_item.added`` / ``output_item.done`` per image_gen call
    * ``sequence_number`` strictly monotonically increasing with no
      gaps > 1 (only checked when every event has one)
    * optional ``partial_image`` sits between ``generating`` and
      ``completed`` when present
* Streaming and non-streaming paths share the same mix-in body, so the
  gRPC lanes exercise the full assertion surface without duplication.

Local gates (ruff check/format, mypy, pytest --collect-only) clean;
12 tests collected (3 classes × 4 tests).

Refs: R6.1 #1355, R6.2 #1356, R6.3 #1359, R6.4 #1358 (all merged)

Co-authored-by: Tingting Zhou <zhoutt96@gmail.com>
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

model-gateway Model gateway crate changes openai OpenAI router changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant