feat(gateway): add Messages API type scaffolding to gRPC router - #739
Conversation
Add the type-level foundation for Anthropic Messages API (/v1/messages) support in the gRPC router pipeline. This PR introduces the RequestType::Messages and FinalResponse::Messages variants plus all supporting context helpers, without wiring up the actual pipeline stages. What changed: - context.rs: Add RequestType::Messages(Arc<CreateMessageRequest>) variant, FinalResponse::Messages(Message) variant, for_messages() factory, messages_request()/messages_request_arc() typed accessors, and is_streaming() support for Messages requests - metrics.rs: Add ENDPOINT_MESSAGES constant for metrics labeling - dispatch_metadata.rs: Handle Messages variant for model extraction - pipeline.rs: Add Messages to all FinalResponse match arms in execute_chat, execute_generate, execute_chat_for_responses - regular/stages/request_building.rs: Add Messages to wrong-pipeline error arm alongside Responses - regular/stages/response_processing.rs: Add Messages to wrong-pipeline error arm alongside Responses - harmony/stages/request_building.rs: Add Messages not-supported error for Harmony models - harmony/stages/response_processing.rs: Add Messages to not-supported arm for Harmony pipeline Why: The Messages API will be a first-class pipeline in the gRPC router, with its own endpoint-specific stages (1, 4, 7) sharing common stages (2, 3, 5, 6) — same architecture as chat. This scaffolding PR establishes the type system so follow-up PRs can add the actual message-specific stages using git cp from chat stage files. How: Each existing exhaustive match on RequestType and FinalResponse was updated to handle the new Messages variant. New context factory and accessor methods follow the same patterns as chat/responses/embedding. The new variants are annotated with #[expect(dead_code)] where appropriate since no pipeline constructs them yet. Signed-off-by: Si Lin <silin@meta.com> Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request establishes the essential type system and initial integration points for the Anthropic Messages API within the gRPC router. It lays the groundwork for a dedicated Messages API pipeline, mirroring the architecture of the existing chat pipeline, without yet implementing the full pipeline stages. This preparatory work ensures that subsequent development can build upon a solid, type-safe foundation. Highlights
Changelog
Activity
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension. Footnotes
|
|
Warning Gemini encountered an error creating the review. You can try again by commenting |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (2)
📝 WalkthroughWalkthroughAdds Messages support across the gRPC pipeline: new RequestType::Messages, RequestContext constructors/accessors, FinalResponse::Messages, a metrics label for the messages endpoint, and pipeline stage updates to route or reject Messages where unsupported. Changes
Sequence Diagram(s)sequenceDiagram
participant Client
participant RequestContext
participant Pipeline
participant HarmonyHandler
participant RegularHandler
participant ResponseProcessor
Client->>RequestContext: for_messages(Arc<CreateMessageRequest), headers,...
RequestContext->>Pipeline: submit(RequestType::Messages)
Pipeline->>Pipeline: route by RequestType
alt Routed to Harmony
Pipeline->>HarmonyHandler: forward Messages
HarmonyHandler->>Pipeline: return bad_request (not_supported_in_harmony)
else Routed to Regular
Pipeline->>RegularHandler: forward Messages
RegularHandler->>ResponseProcessor: process Messages
ResponseProcessor->>Pipeline: return FinalResponse::Messages
end
Pipeline->>Client: respond (error or FinalResponse::Messages)
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~25 minutes Suggested reviewers
Poem
🚥 Pre-merge checks | ✅ 3✅ Passed checks (3 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches
🧪 Generate unit tests (beta)
Comment |
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@model_gateway/src/routers/grpc/regular/stages/request_building.rs`:
- Around line 44-52: Split the combined match arm handling
RequestType::Responses(_) | RequestType::Messages(_) into two distinct arms so
each variant returns a separate wrong-pipeline diagnostic; in
RequestBuildingStage::execute match the RequestType::Responses(_) arm and call
grpc_error::internal_error with a unique code (e.g., "wrong_pipeline_responses")
and an explanatory message, and do the same for RequestType::Messages(_) with
its own code (e.g., "wrong_pipeline_messages"); also update the error! log to
include the specific variant name so logs show which endpoint was misrouted.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: ASSERTIVE
Plan: Pro
Run ID: 64bc3f59-a79b-43ef-a237-f931eab73120
📒 Files selected for processing (8)
model_gateway/src/observability/metrics.rsmodel_gateway/src/routers/grpc/common/stages/dispatch_metadata.rsmodel_gateway/src/routers/grpc/context.rsmodel_gateway/src/routers/grpc/harmony/stages/request_building.rsmodel_gateway/src/routers/grpc/harmony/stages/response_processing.rsmodel_gateway/src/routers/grpc/pipeline.rsmodel_gateway/src/routers/grpc/regular/stages/request_building.rsmodel_gateway/src/routers/grpc/regular/stages/response_processing.rs
CatherineSue
left a comment
There was a problem hiding this comment.
overall LGTM. I agree with coderabbit regarding our error msg. Could be refined. I hate keep editing the error msg.
- Add `impl Display for RequestType` in context.rs so variant names are available in formatted error messages - Update request_building.rs and response_processing.rs to use `request_type @` pattern binding with Display formatting, making wrong-pipeline errors include the specific request type name - Addresses PR review feedback to keep distinct diagnostics per request type without duplicating match arms Signed-off-by: Simon Lin <simon@lightseek.ai> Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
- Collapse 4 separate unsupported-request-type arms into a single `request_type @` pattern in HarmonyRequestBuildingStage and HarmonyResponseProcessingStage - Error messages now dynamically include the request type name via the Display impl added in the previous commit Signed-off-by: Simon Lin <simon@lightseek.ai> Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
…ages API
Add the first message-specific pipeline stage (Stage 1: Preparation) and
the utility functions it needs to convert Anthropic Messages API types
into the internal chat template format.
What changed:
- New message_utils.rs with conversion functions:
- process_messages(): top-level orchestrator parallel to process_chat_messages()
- process_message_content_format(): converts InputMessage to Vec<Value> JSON
- convert_user_message(): handles user messages, splits ToolResult into
separate "tool" role messages
- convert_assistant_message(): extracts text, tool_calls, reasoning_content
- extract_chat_tools(): filters Custom tools and converts to chat::Tool
- convert_message_tool_choice(): maps Messages ToolChoice to chat ToolChoice
- extract_tool_result_text(): helper for ToolResult content extraction
- 7 unit tests covering all major conversion paths
- New MessagePreparationStage (git-cp'd from ChatPreparationStage):
- Same structure as ChatPreparationStage (impl method pattern)
- Resolves tokenizer, converts/filters tools, processes messages,
tokenizes, builds tool constraints, creates stop decoder
- Multimodal processing postponed (marked with async for future .await)
- Made process_tool_call_arguments pub(crate) in chat_utils.rs for reuse
- Updated delegating PreparationStage to use Display-based error messages
- Removed stale #[expect(dead_code)] from messages_request_arc (now used)
Why:
This is PR 2 in the Messages API gRPC pipeline series. PR 1 (#739) added
type scaffolding. This PR adds the preparation stage that converts
Messages API requests into the shared internal format, enabling the
existing request building and response processing stages to work with
Messages API requests in follow-up PRs.
How:
Follows the same architecture as chat: reuses shared utilities
(resolve_tokenizer, filter_tools_by_tool_choice, generate_tool_constraints,
create_stop_decoder, process_tool_call_arguments) and only replaces the
message-specific conversion layer (process_content_format → process_message_content_format).
Uses git-cp to preserve file history from chat/preparation.rs for reviewability.
Refs: #738
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
…ages API
Add the first message-specific pipeline stage (Stage 1: Preparation) and
the utility functions it needs to convert Anthropic Messages API types
into the internal chat template format.
What changed:
- New message_utils.rs with conversion functions:
- process_messages(): top-level orchestrator parallel to process_chat_messages()
- process_message_content_format(): converts InputMessage to Vec<Value> JSON
- convert_user_message(): handles user messages, splits ToolResult into
separate "tool" role messages
- convert_assistant_message(): extracts text, tool_calls, reasoning_content
- extract_chat_tools(): filters Custom tools and converts to chat::Tool
- convert_message_tool_choice(): maps Messages ToolChoice to chat ToolChoice
- extract_tool_result_text(): helper for ToolResult content extraction
- 7 unit tests covering all major conversion paths
- New MessagePreparationStage (git-cp'd from ChatPreparationStage):
- Same structure as ChatPreparationStage (impl method pattern)
- Resolves tokenizer, converts/filters tools, processes messages,
tokenizes, builds tool constraints, creates stop decoder
- Multimodal processing postponed (marked with async for future .await)
- Made process_tool_call_arguments pub(crate) in chat_utils.rs for reuse
- Updated delegating PreparationStage to use Display-based error messages
- Removed stale #[expect(dead_code)] from messages_request_arc (now used)
Why:
This is PR 2 in the Messages API gRPC pipeline series. PR 1 (#739) added
type scaffolding. This PR adds the preparation stage that converts
Messages API requests into the shared internal format, enabling the
existing request building and response processing stages to work with
Messages API requests in follow-up PRs.
How:
Follows the same architecture as chat: reuses shared utilities
(resolve_tokenizer, filter_tools_by_tool_choice, generate_tool_constraints,
create_stop_decoder, process_tool_call_arguments) and only replaces the
message-specific conversion layer (process_content_format → process_message_content_format).
Uses git-cp to preserve file history from chat/preparation.rs for reviewability.
Refs: #738
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
…ages API
Add the first message-specific pipeline stage (Stage 1: Preparation) and
the utility functions it needs to convert Anthropic Messages API types
into the internal chat template format.
What changed:
- New message_utils.rs with conversion functions:
- process_messages(): top-level orchestrator parallel to process_chat_messages()
- process_message_content_format(): converts InputMessage to Vec<Value> JSON
- convert_user_message(): handles user messages, splits ToolResult into
separate "tool" role messages
- convert_assistant_message(): extracts text, tool_calls, reasoning_content
- extract_chat_tools(): filters Custom tools and converts to chat::Tool
- convert_message_tool_choice(): maps Messages ToolChoice to chat ToolChoice
- extract_tool_result_text(): helper for ToolResult content extraction
- 7 unit tests covering all major conversion paths
- New MessagePreparationStage (parallel to ChatPreparationStage):
- Same structure as ChatPreparationStage (impl method pattern)
- Resolves tokenizer, converts/filters tools, processes messages,
tokenizes, builds tool constraints, creates stop decoder
- Multimodal processing postponed (marked with async for future .await)
- Made process_tool_call_arguments pub(crate) in chat_utils.rs for reuse
- Updated delegating PreparationStage to use Display-based error messages
- Removed stale #[expect(dead_code)] from messages_request_arc (now used)
Why:
This is PR 2 in the Messages API gRPC pipeline series. PR 1 (#739) added
type scaffolding. This PR adds the preparation stage that converts
Messages API requests into the shared internal format, enabling the
existing request building and response processing stages to work with
Messages API requests in follow-up PRs.
How:
Follows the same architecture as chat: reuses shared utilities
(resolve_tokenizer, filter_tools_by_tool_choice, generate_tool_constraints,
create_stop_decoder, process_tool_call_arguments) and only replaces the
message-specific conversion layer (process_content_format → process_message_content_format).
Refs: #738
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
…ages API
Add the first message-specific pipeline stage (Stage 1: Preparation) and
the utility functions it needs to convert Anthropic Messages API types
into the internal chat template format.
What changed:
- New message_utils.rs with conversion functions:
- process_messages(): top-level orchestrator parallel to process_chat_messages()
- process_message_content_format(): converts InputMessage to Vec<Value> JSON
- convert_user_message(): handles user messages, splits ToolResult into
separate "tool" role messages
- convert_assistant_message(): extracts text, tool_calls, reasoning_content
- extract_chat_tools(): filters Custom tools and converts to chat::Tool
- convert_message_tool_choice(): maps Messages ToolChoice to chat ToolChoice
- extract_tool_result_text(): helper for ToolResult content extraction
- 7 unit tests covering all major conversion paths
- New MessagePreparationStage (parallel to ChatPreparationStage):
- Same structure as ChatPreparationStage (impl method pattern)
- Resolves tokenizer, converts/filters tools, processes messages,
tokenizes, builds tool constraints, creates stop decoder
- Multimodal processing postponed (marked with async for future .await)
- Made process_tool_call_arguments pub(crate) in chat_utils.rs for reuse
- Updated delegating PreparationStage to use Display-based error messages
- Removed stale #[expect(dead_code)] from messages_request_arc (now used)
Why:
This is PR 2 in the Messages API gRPC pipeline series. PR 1 (#739) added
type scaffolding. This PR adds the preparation stage that converts
Messages API requests into the shared internal format, enabling the
existing request building and response processing stages to work with
Messages API requests in follow-up PRs.
How:
Follows the same architecture as chat: reuses shared utilities
(resolve_tokenizer, filter_tools_by_tool_choice, generate_tool_constraints,
create_stop_decoder, process_tool_call_arguments) and only replaces the
message-specific conversion layer (process_content_format → process_message_content_format).
Refs: #738
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
…arams for Messages API
Add Stage 4 (request building) for the Messages API gRPC pipeline,
converting PreparationOutput + CreateMessageRequest sampling parameters
into backend-specific proto GenerateRequest.
What changed:
- model_gateway/src/routers/grpc/regular/stages/messages/request_building.rs:
New MessageRequestBuildingStage (copied from chat, adapted for Messages).
Uses msg_{uuid} request ID prefix, calls build_messages_request(),
skips multimodal (postponed), no filtered_request pattern.
- model_gateway/src/routers/grpc/client.rs:
Add build_messages_request() dispatcher on GrpcClient enum, dispatching
to each backend's build_generate_request_from_messages().
- crates/grpc_client/src/sglang_scheduler.rs:
Add build_generate_request_from_messages() and
build_grpc_sampling_params_from_messages(). Maps CreateMessageRequest
fields (max_tokens, temperature, top_p, top_k, stop_sequences) to
sglang proto SamplingParams with sensible defaults for missing fields.
- crates/grpc_client/src/vllm_engine.rs:
Same pattern for vLLM backend. Handles vLLM-specific differences
(top_k=0 for disabled, Optional<f32> temperature).
- crates/grpc_client/src/trtllm_service.rs:
Same pattern for TRT-LLM backend using SamplingConfig, OutputConfig,
and GuidedDecodingParams proto types.
- model_gateway/src/routers/grpc/regular/stages/messages/mod.rs:
Wire request_building module and re-export MessageRequestBuildingStage.
Why: This is PR 3 in the Messages API gRPC series. Stage 4 bridges
the gap between preparation (Stage 1, PR #741) and response processing
(Stage 7, future PR), enabling the pipeline to build backend-specific
proto requests from Messages API parameters.
Refs: #739, #741
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
…on-streaming) Add Stage 7 (response processing) for the Messages API gRPC pipeline. This converts backend ProtoGenerateComplete responses into Anthropic Message format with proper ContentBlock construction and StopReason mapping. Non-streaming only; streaming deferred to follow-up PR. What changed: - processor.rs: add process_non_streaming_messages_response() to ResponseProcessor — full pipeline: token decoding, reasoning parsing, tool call parsing, content block construction (Thinking → Text → ToolUse), StopReason mapping (EndTurn/MaxTokens/StopSequence/ToolUse), and messages::Usage building - messages/response_processing.rs: new MessageResponseProcessingStage that extracts execution result, dispatch metadata, tokenizer, and stop decoder from RequestContext, delegates to ResponseProcessor, and stores FinalResponse::Messages - message_utils.rs: add get_history_tool_calls_count_messages() for counting tool use blocks in Messages API request history (needed for KimiK2-style tool call ID generation) - messages/mod.rs: wire response_processing module with unused_imports expect (wired in pipeline factory PR) Why: This is the fourth PR in the Messages API gRPC support series. With preparation (PR #741), request building (PR #744), and now response processing, three of the four endpoint-specific pipeline stages are complete. The shared stages (worker selection, client acquisition, dispatch, execution) are reused from the existing pipeline. How: Follows the same architecture as chat's response processing but adapted for Anthropic Message types: - Reuses existing convert_message_tool_choice() from message_utils to bridge Messages ToolChoice → Chat ToolChoice for parse_json_schema_response - Reuses ResponseProcessor's parse_tool_calls() for model-predicted path - Content blocks ordered per Anthropic convention: Thinking first, Text, then ToolUse blocks - Tool calls parsed as OpenAI ToolCall (via existing parsers) then converted to ContentBlock::ToolUse with JSON input - Messages always n=1, no logprobs - ThinkingConfig::Enabled check replaces separate_reasoning bool Refs: #739, #741, #744 Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
Summary
Add the type-level foundation for Anthropic Messages API (
/v1/messages) support in the gRPC router pipeline. IntroducesRequestType::MessagesandFinalResponse::Messagesvariants plus all supporting context helpers, without wiring up actual pipeline stages.This is the first in a series of PRs to add first-class Messages API support to the gRPC router, following the same 7-stage pipeline architecture as chat.
What changed
RequestType::Messages(Arc<CreateMessageRequest>)variant,FinalResponse::Messages(Message)variant,for_messages()factory,messages_request()/messages_request_arc()typed accessors,is_streaming()supportENDPOINT_MESSAGESconstantFinalResponsematch armsWhy
The Messages API will be a first-class pipeline in the gRPC router with its own endpoint-specific stages (1, 4, 7) sharing common stages (2, 3, 5, 6) — same architecture as chat. This scaffolding PR establishes the type system so follow-up PRs can add the actual message-specific stages using
git cpfrom chat stage files.How
Each existing exhaustive match on
RequestTypeandFinalResponsewas updated to handle the newMessagesvariant. New context factory and accessor methods follow the same patterns as chat/responses/embedding. New variants are annotated with#[expect(dead_code)]where appropriate since no pipeline constructs them yet.Test plan
cargo fmt— passescargo clippy --all-targets --all-features -- -D warnings— passes, zero errorsSummary by CodeRabbit
New Features
Bug Fixes / Validation