Skip to content

feat(gateway): add Messages API type scaffolding to gRPC router - #739

Merged
slin1237 merged 3 commits into
mainfrom
slin/msg-1
Mar 12, 2026
Merged

slin1237 merged 3 commits into
mainfrom
slin/msg-1

Conversation

@slin1237

@slin1237 slin1237 commented Mar 12, 2026 •

Copy link
Copy Markdown
Member

Summary

Add the type-level foundation for Anthropic Messages API (/v1/messages) support in the gRPC router pipeline. Introduces RequestType::Messages and FinalResponse::Messages variants plus all supporting context helpers, without wiring up actual pipeline stages.

This is the first in a series of PRs to add first-class Messages API support to the gRPC router, following the same 7-stage pipeline architecture as chat.

What changed

  • context.rs: RequestType::Messages(Arc<CreateMessageRequest>) variant, FinalResponse::Messages(Message) variant, for_messages() factory, messages_request()/messages_request_arc() typed accessors, is_streaming() support
  • metrics.rs: ENDPOINT_MESSAGES constant
  • dispatch_metadata.rs: Messages variant match arm for model extraction
  • pipeline.rs: Messages added to all FinalResponse match arms
  • regular/stages/request_building.rs: Messages added to wrong-pipeline error arm
  • regular/stages/response_processing.rs: Messages added to wrong-pipeline error arm
  • harmony/stages/request_building.rs: Messages not-supported error for Harmony models
  • harmony/stages/response_processing.rs: Messages added to not-supported arm

Why

The Messages API will be a first-class pipeline in the gRPC router with its own endpoint-specific stages (1, 4, 7) sharing common stages (2, 3, 5, 6) — same architecture as chat. This scaffolding PR establishes the type system so follow-up PRs can add the actual message-specific stages using git cp from chat stage files.

How

Each existing exhaustive match on RequestType and FinalResponse was updated to handle the new Messages variant. New context factory and accessor methods follow the same patterns as chat/responses/embedding. New variants are annotated with #[expect(dead_code)] where appropriate since no pipeline constructs them yet.

Test plan

  • cargo fmt — passes
  • cargo clippy --all-targets --all-features -- -D warnings — passes, zero errors
  • No new files created — purely additive type changes to existing files
  • All existing tests unaffected (new variants are unreachable in current pipelines)

Summary by CodeRabbit

  • New Features

    • Added Messages API support across gRPC routing, request context, and final responses.
    • Integrated Messages into metrics labeling.
  • Bug Fixes / Validation

    • Prevented Messages requests from using unsupported pipelines/models with clearer, unified error responses.
    • Consolidated and unified error messages for Responses and Messages to improve diagnostics.

Add the type-level foundation for Anthropic Messages API (/v1/messages)
support in the gRPC router pipeline. This PR introduces the
RequestType::Messages and FinalResponse::Messages variants plus all
supporting context helpers, without wiring up the actual pipeline stages.

What changed:
- context.rs: Add RequestType::Messages(Arc<CreateMessageRequest>)
  variant, FinalResponse::Messages(Message) variant, for_messages()
  factory, messages_request()/messages_request_arc() typed accessors,
  and is_streaming() support for Messages requests
- metrics.rs: Add ENDPOINT_MESSAGES constant for metrics labeling
- dispatch_metadata.rs: Handle Messages variant for model extraction
- pipeline.rs: Add Messages to all FinalResponse match arms in
  execute_chat, execute_generate, execute_chat_for_responses
- regular/stages/request_building.rs: Add Messages to wrong-pipeline
  error arm alongside Responses
- regular/stages/response_processing.rs: Add Messages to wrong-pipeline
  error arm alongside Responses
- harmony/stages/request_building.rs: Add Messages not-supported error
  for Harmony models
- harmony/stages/response_processing.rs: Add Messages to not-supported
  arm for Harmony pipeline

Why:
The Messages API will be a first-class pipeline in the gRPC router,
with its own endpoint-specific stages (1, 4, 7) sharing common stages
(2, 3, 5, 6) — same architecture as chat. This scaffolding PR
establishes the type system so follow-up PRs can add the actual
message-specific stages using git cp from chat stage files.

How:
Each existing exhaustive match on RequestType and FinalResponse was
updated to handle the new Messages variant. New context factory and
accessor methods follow the same patterns as chat/responses/embedding.
The new variants are annotated with #[expect(dead_code)] where
appropriate since no pipeline constructs them yet.

Signed-off-by: Si Lin <silin@meta.com>
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
@github-actions github-actions Bot added grpc gRPC client and router changes model-gateway Model gateway crate changes labels Mar 12, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request establishes the essential type system and initial integration points for the Anthropic Messages API within the gRPC router. It lays the groundwork for a dedicated Messages API pipeline, mirroring the architecture of the existing chat pipeline, without yet implementing the full pipeline stages. This preparatory work ensures that subsequent development can build upon a solid, type-safe foundation.

Highlights

  • Messages API Scaffolding: Introduced foundational type-level support for the Anthropic Messages API (/v1/messages) within the gRPC router pipeline.
  • New Request and Response Types: Added RequestType::Messages and FinalResponse::Messages variants, along with associated context helpers like for_messages(), messages_request(), messages_request_arc(), and is_streaming() support.
  • Pipeline Integration Points: Updated various existing pipeline stages to recognize and handle the new Messages variant, primarily for model extraction, error handling for unsupported pipelines (Harmony, regular), and correct response type checks.
  • Metrics: Added a new ENDPOINT_MESSAGES constant for observability.
Changelog
  • model_gateway/src/observability/metrics.rs
    • Added ENDPOINT_MESSAGES constant.
  • model_gateway/src/routers/grpc/common/stages/dispatch_metadata.rs
    • Added Messages variant match arm for model extraction.
  • model_gateway/src/routers/grpc/context.rs
    • Added RequestType::Messages(Arc<CreateMessageRequest>) variant.
    • Added FinalResponse::Messages(Message) variant.
    • Introduced for_messages() factory.
    • Added messages_request() and messages_request_arc() typed accessors.
    • Extended is_streaming() support for Messages requests.
  • model_gateway/src/routers/grpc/harmony/stages/request_building.rs
    • Added not-supported error for Messages requests with Harmony models.
  • model_gateway/src/routers/grpc/harmony/stages/response_processing.rs
    • Added Messages to the list of request types not supported in the Harmony pipeline.
  • model_gateway/src/routers/grpc/pipeline.rs
    • Added Messages to FinalResponse match arms in error logging for execute_chat, execute_generate, and execute_chat_for_responses.
  • model_gateway/src/routers/grpc/regular/stages/request_building.rs
    • Added Messages to the wrong-pipeline error arm, indicating it should use a dedicated pipeline.
  • model_gateway/src/routers/grpc/regular/stages/response_processing.rs
    • Added Messages to the wrong-pipeline error arm, indicating it should use a dedicated pipeline.
Activity
  • The author confirmed cargo fmt passed.
  • The author confirmed cargo clippy --all-targets --all-features -- -D warnings passed with zero errors.
  • The author noted that no new files were created, only additive type changes to existing files.
  • The author confirmed all existing tests were unaffected, as new variants are currently unreachable in the pipelines.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

Gemini encountered an error creating the review. You can try again by commenting /gemini review.

@coderabbitai

coderabbitai Bot commented Mar 12, 2026 •

Copy link
Copy Markdown

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 96551544-f4f0-452d-9c7b-f072b53159a5

📥 Commits

Reviewing files that changed from the base of the PR and between 352c744 and 110e7b2.

📒 Files selected for processing (2)
  • model_gateway/src/routers/grpc/harmony/stages/request_building.rs
  • model_gateway/src/routers/grpc/harmony/stages/response_processing.rs

📝 Walkthrough

Walkthrough

Adds Messages support across the gRPC pipeline: new RequestType::Messages, RequestContext constructors/accessors, FinalResponse::Messages, a metrics label for the messages endpoint, and pipeline stage updates to route or reject Messages where unsupported.

Changes

Cohort / File(s) Summary
Metrics Configuration
model_gateway/src/observability/metrics.rs
Added ENDPOINT_MESSAGES constant ("messages") to metrics_labels.
Core Request Context
model_gateway/src/routers/grpc/context.rs
Added RequestType::Messages(Arc<CreateMessageRequest>), for_messages() constructor, messages_request() / messages_request_arc() accessors, updated is_streaming to include Messages, and added FinalResponse::Messages(Message).
Dispatch Metadata
model_gateway/src/routers/grpc/common/stages/dispatch_metadata.rs
Model extraction now handles RequestType::Messages by cloning its model field like other request types.
Harmony Pipeline
model_gateway/src/routers/grpc/harmony/stages/request_building.rs, model_gateway/src/routers/grpc/harmony/stages/response_processing.rs
Added Messages to unsupported-types paths; unified unsupported error path returning a standardized not_supported_in_harmony response for Messages.
Regular Pipeline
model_gateway/src/routers/grpc/regular/stages/request_building.rs, model_gateway/src/routers/grpc/regular/stages/response_processing.rs
Treats RequestType::Messages alongside RequestType::Responses with unified wrong_pipeline error handling and updated logs/messages.
Final Response Handling
model_gateway/src/routers/grpc/pipeline.rs
Endpoints updated to accept FinalResponse::Messages as a final response variant and updated error/log text where applicable.

Sequence Diagram(s)

sequenceDiagram
    participant Client
    participant RequestContext
    participant Pipeline
    participant HarmonyHandler
    participant RegularHandler
    participant ResponseProcessor

    Client->>RequestContext: for_messages(Arc<CreateMessageRequest), headers,...
    RequestContext->>Pipeline: submit(RequestType::Messages)
    Pipeline->>Pipeline: route by RequestType

    alt Routed to Harmony
        Pipeline->>HarmonyHandler: forward Messages
        HarmonyHandler->>Pipeline: return bad_request (not_supported_in_harmony)
    else Routed to Regular
        Pipeline->>RegularHandler: forward Messages
        RegularHandler->>ResponseProcessor: process Messages
        ResponseProcessor->>Pipeline: return FinalResponse::Messages
    end

    Pipeline->>Client: respond (error or FinalResponse::Messages)
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Suggested reviewers

  • key4ng

Poem

🐰 I hopped through code with nimble paws,
A "messages" label now joins the laws,
Context learns the request and streams,
Harmony may say no, Regular redeems,
Metrics and responses clap their jaws — hooray!

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately describes the main change: adding Messages API type scaffolding to the gRPC router. It is concise, specific, and clearly summarizes the primary objective without unnecessary detail.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
  • 📝 Generate docstrings (stacked PR)
  • 📝 Generate docstrings (commit on current branch)
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch slin/msg-1

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@model_gateway/src/routers/grpc/regular/stages/request_building.rs`:
- Around line 44-52: Split the combined match arm handling
RequestType::Responses(_) | RequestType::Messages(_) into two distinct arms so
each variant returns a separate wrong-pipeline diagnostic; in
RequestBuildingStage::execute match the RequestType::Responses(_) arm and call
grpc_error::internal_error with a unique code (e.g., "wrong_pipeline_responses")
and an explanatory message, and do the same for RequestType::Messages(_) with
its own code (e.g., "wrong_pipeline_messages"); also update the error! log to
include the specific variant name so logs show which endpoint was misrouted.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 64bc3f59-a79b-43ef-a237-f931eab73120

📥 Commits

Reviewing files that changed from the base of the PR and between 8d6689b and 733557d.

📒 Files selected for processing (8)
  • model_gateway/src/observability/metrics.rs
  • model_gateway/src/routers/grpc/common/stages/dispatch_metadata.rs
  • model_gateway/src/routers/grpc/context.rs
  • model_gateway/src/routers/grpc/harmony/stages/request_building.rs
  • model_gateway/src/routers/grpc/harmony/stages/response_processing.rs
  • model_gateway/src/routers/grpc/pipeline.rs
  • model_gateway/src/routers/grpc/regular/stages/request_building.rs
  • model_gateway/src/routers/grpc/regular/stages/response_processing.rs

Comment thread model_gateway/src/routers/grpc/regular/stages/request_building.rs Outdated
Comment thread model_gateway/src/routers/grpc/harmony/stages/response_processing.rs Outdated

@CatherineSue CatherineSue left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

overall LGTM. I agree with coderabbit regarding our error msg. Could be refined. I hate keep editing the error msg.

- Add `impl Display for RequestType` in context.rs so variant names
  are available in formatted error messages
- Update request_building.rs and response_processing.rs to use
  `request_type @` pattern binding with Display formatting, making
  wrong-pipeline errors include the specific request type name
- Addresses PR review feedback to keep distinct diagnostics per
  request type without duplicating match arms

Signed-off-by: Simon Lin <simon@lightseek.ai>
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
- Collapse 4 separate unsupported-request-type arms into a single
  `request_type @` pattern in HarmonyRequestBuildingStage and
  HarmonyResponseProcessingStage
- Error messages now dynamically include the request type name via
  the Display impl added in the previous commit

Signed-off-by: Simon Lin <simon@lightseek.ai>
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
@slin1237
slin1237 merged commit e1d2183 into main Mar 12, 2026
26 of 29 checks passed
@slin1237
slin1237 deleted the slin/msg-1 branch March 12, 2026 05:38
slin1237 added a commit that referenced this pull request Mar 12, 2026
…ages API

Add the first message-specific pipeline stage (Stage 1: Preparation) and
the utility functions it needs to convert Anthropic Messages API types
into the internal chat template format.

What changed:
- New message_utils.rs with conversion functions:
  - process_messages(): top-level orchestrator parallel to process_chat_messages()
  - process_message_content_format(): converts InputMessage to Vec<Value> JSON
  - convert_user_message(): handles user messages, splits ToolResult into
    separate "tool" role messages
  - convert_assistant_message(): extracts text, tool_calls, reasoning_content
  - extract_chat_tools(): filters Custom tools and converts to chat::Tool
  - convert_message_tool_choice(): maps Messages ToolChoice to chat ToolChoice
  - extract_tool_result_text(): helper for ToolResult content extraction
  - 7 unit tests covering all major conversion paths
- New MessagePreparationStage (git-cp'd from ChatPreparationStage):
  - Same structure as ChatPreparationStage (impl method pattern)
  - Resolves tokenizer, converts/filters tools, processes messages,
    tokenizes, builds tool constraints, creates stop decoder
  - Multimodal processing postponed (marked with async for future .await)
- Made process_tool_call_arguments pub(crate) in chat_utils.rs for reuse
- Updated delegating PreparationStage to use Display-based error messages
- Removed stale #[expect(dead_code)] from messages_request_arc (now used)

Why:
This is PR 2 in the Messages API gRPC pipeline series. PR 1 (#739) added
type scaffolding. This PR adds the preparation stage that converts
Messages API requests into the shared internal format, enabling the
existing request building and response processing stages to work with
Messages API requests in follow-up PRs.

How:
Follows the same architecture as chat: reuses shared utilities
(resolve_tokenizer, filter_tools_by_tool_choice, generate_tool_constraints,
create_stop_decoder, process_tool_call_arguments) and only replaces the
message-specific conversion layer (process_content_format → process_message_content_format).
Uses git-cp to preserve file history from chat/preparation.rs for reviewability.

Refs: #738
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
slin1237 added a commit that referenced this pull request Mar 12, 2026
…ages API

Add the first message-specific pipeline stage (Stage 1: Preparation) and
the utility functions it needs to convert Anthropic Messages API types
into the internal chat template format.

What changed:
- New message_utils.rs with conversion functions:
  - process_messages(): top-level orchestrator parallel to process_chat_messages()
  - process_message_content_format(): converts InputMessage to Vec<Value> JSON
  - convert_user_message(): handles user messages, splits ToolResult into
    separate "tool" role messages
  - convert_assistant_message(): extracts text, tool_calls, reasoning_content
  - extract_chat_tools(): filters Custom tools and converts to chat::Tool
  - convert_message_tool_choice(): maps Messages ToolChoice to chat ToolChoice
  - extract_tool_result_text(): helper for ToolResult content extraction
  - 7 unit tests covering all major conversion paths
- New MessagePreparationStage (git-cp'd from ChatPreparationStage):
  - Same structure as ChatPreparationStage (impl method pattern)
  - Resolves tokenizer, converts/filters tools, processes messages,
    tokenizes, builds tool constraints, creates stop decoder
  - Multimodal processing postponed (marked with async for future .await)
- Made process_tool_call_arguments pub(crate) in chat_utils.rs for reuse
- Updated delegating PreparationStage to use Display-based error messages
- Removed stale #[expect(dead_code)] from messages_request_arc (now used)

Why:
This is PR 2 in the Messages API gRPC pipeline series. PR 1 (#739) added
type scaffolding. This PR adds the preparation stage that converts
Messages API requests into the shared internal format, enabling the
existing request building and response processing stages to work with
Messages API requests in follow-up PRs.

How:
Follows the same architecture as chat: reuses shared utilities
(resolve_tokenizer, filter_tools_by_tool_choice, generate_tool_constraints,
create_stop_decoder, process_tool_call_arguments) and only replaces the
message-specific conversion layer (process_content_format → process_message_content_format).
Uses git-cp to preserve file history from chat/preparation.rs for reviewability.

Refs: #738
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
slin1237 added a commit that referenced this pull request Mar 12, 2026
…ages API

Add the first message-specific pipeline stage (Stage 1: Preparation) and
the utility functions it needs to convert Anthropic Messages API types
into the internal chat template format.

What changed:
- New message_utils.rs with conversion functions:
  - process_messages(): top-level orchestrator parallel to process_chat_messages()
  - process_message_content_format(): converts InputMessage to Vec<Value> JSON
  - convert_user_message(): handles user messages, splits ToolResult into
    separate "tool" role messages
  - convert_assistant_message(): extracts text, tool_calls, reasoning_content
  - extract_chat_tools(): filters Custom tools and converts to chat::Tool
  - convert_message_tool_choice(): maps Messages ToolChoice to chat ToolChoice
  - extract_tool_result_text(): helper for ToolResult content extraction
  - 7 unit tests covering all major conversion paths
- New MessagePreparationStage (parallel to ChatPreparationStage):
  - Same structure as ChatPreparationStage (impl method pattern)
  - Resolves tokenizer, converts/filters tools, processes messages,
    tokenizes, builds tool constraints, creates stop decoder
  - Multimodal processing postponed (marked with async for future .await)
- Made process_tool_call_arguments pub(crate) in chat_utils.rs for reuse
- Updated delegating PreparationStage to use Display-based error messages
- Removed stale #[expect(dead_code)] from messages_request_arc (now used)

Why:
This is PR 2 in the Messages API gRPC pipeline series. PR 1 (#739) added
type scaffolding. This PR adds the preparation stage that converts
Messages API requests into the shared internal format, enabling the
existing request building and response processing stages to work with
Messages API requests in follow-up PRs.

How:
Follows the same architecture as chat: reuses shared utilities
(resolve_tokenizer, filter_tools_by_tool_choice, generate_tool_constraints,
create_stop_decoder, process_tool_call_arguments) and only replaces the
message-specific conversion layer (process_content_format → process_message_content_format).

Refs: #738
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
slin1237 added a commit that referenced this pull request Mar 12, 2026
…ages API

Add the first message-specific pipeline stage (Stage 1: Preparation) and
the utility functions it needs to convert Anthropic Messages API types
into the internal chat template format.

What changed:
- New message_utils.rs with conversion functions:
  - process_messages(): top-level orchestrator parallel to process_chat_messages()
  - process_message_content_format(): converts InputMessage to Vec<Value> JSON
  - convert_user_message(): handles user messages, splits ToolResult into
    separate "tool" role messages
  - convert_assistant_message(): extracts text, tool_calls, reasoning_content
  - extract_chat_tools(): filters Custom tools and converts to chat::Tool
  - convert_message_tool_choice(): maps Messages ToolChoice to chat ToolChoice
  - extract_tool_result_text(): helper for ToolResult content extraction
  - 7 unit tests covering all major conversion paths
- New MessagePreparationStage (parallel to ChatPreparationStage):
  - Same structure as ChatPreparationStage (impl method pattern)
  - Resolves tokenizer, converts/filters tools, processes messages,
    tokenizes, builds tool constraints, creates stop decoder
  - Multimodal processing postponed (marked with async for future .await)
- Made process_tool_call_arguments pub(crate) in chat_utils.rs for reuse
- Updated delegating PreparationStage to use Display-based error messages
- Removed stale #[expect(dead_code)] from messages_request_arc (now used)

Why:
This is PR 2 in the Messages API gRPC pipeline series. PR 1 (#739) added
type scaffolding. This PR adds the preparation stage that converts
Messages API requests into the shared internal format, enabling the
existing request building and response processing stages to work with
Messages API requests in follow-up PRs.

How:
Follows the same architecture as chat: reuses shared utilities
(resolve_tokenizer, filter_tools_by_tool_choice, generate_tool_constraints,
create_stop_decoder, process_tool_call_arguments) and only replaces the
message-specific conversion layer (process_content_format → process_message_content_format).

Refs: #738
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
slin1237 added a commit that referenced this pull request Mar 12, 2026
…arams for Messages API

Add Stage 4 (request building) for the Messages API gRPC pipeline,
converting PreparationOutput + CreateMessageRequest sampling parameters
into backend-specific proto GenerateRequest.

What changed:
- model_gateway/src/routers/grpc/regular/stages/messages/request_building.rs:
  New MessageRequestBuildingStage (copied from chat, adapted for Messages).
  Uses msg_{uuid} request ID prefix, calls build_messages_request(),
  skips multimodal (postponed), no filtered_request pattern.
- model_gateway/src/routers/grpc/client.rs:
  Add build_messages_request() dispatcher on GrpcClient enum, dispatching
  to each backend's build_generate_request_from_messages().
- crates/grpc_client/src/sglang_scheduler.rs:
  Add build_generate_request_from_messages() and
  build_grpc_sampling_params_from_messages(). Maps CreateMessageRequest
  fields (max_tokens, temperature, top_p, top_k, stop_sequences) to
  sglang proto SamplingParams with sensible defaults for missing fields.
- crates/grpc_client/src/vllm_engine.rs:
  Same pattern for vLLM backend. Handles vLLM-specific differences
  (top_k=0 for disabled, Optional<f32> temperature).
- crates/grpc_client/src/trtllm_service.rs:
  Same pattern for TRT-LLM backend using SamplingConfig, OutputConfig,
  and GuidedDecodingParams proto types.
- model_gateway/src/routers/grpc/regular/stages/messages/mod.rs:
  Wire request_building module and re-export MessageRequestBuildingStage.

Why: This is PR 3 in the Messages API gRPC series. Stage 4 bridges
the gap between preparation (Stage 1, PR #741) and response processing
(Stage 7, future PR), enabling the pipeline to build backend-specific
proto requests from Messages API parameters.

Refs: #739, #741
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
slin1237 added a commit that referenced this pull request Mar 12, 2026
…on-streaming)

Add Stage 7 (response processing) for the Messages API gRPC pipeline.
This converts backend ProtoGenerateComplete responses into Anthropic
Message format with proper ContentBlock construction and StopReason
mapping. Non-streaming only; streaming deferred to follow-up PR.

What changed:
- processor.rs: add process_non_streaming_messages_response() to
  ResponseProcessor — full pipeline: token decoding, reasoning parsing,
  tool call parsing, content block construction (Thinking → Text →
  ToolUse), StopReason mapping (EndTurn/MaxTokens/StopSequence/ToolUse),
  and messages::Usage building
- messages/response_processing.rs: new MessageResponseProcessingStage
  that extracts execution result, dispatch metadata, tokenizer, and
  stop decoder from RequestContext, delegates to ResponseProcessor,
  and stores FinalResponse::Messages
- message_utils.rs: add get_history_tool_calls_count_messages() for
  counting tool use blocks in Messages API request history (needed for
  KimiK2-style tool call ID generation)
- messages/mod.rs: wire response_processing module with unused_imports
  expect (wired in pipeline factory PR)

Why:
This is the fourth PR in the Messages API gRPC support series. With
preparation (PR #741), request building (PR #744), and now response
processing, three of the four endpoint-specific pipeline stages are
complete. The shared stages (worker selection, client acquisition,
dispatch, execution) are reused from the existing pipeline.

How:
Follows the same architecture as chat's response processing but adapted
for Anthropic Message types:
- Reuses existing convert_message_tool_choice() from message_utils to
  bridge Messages ToolChoice → Chat ToolChoice for parse_json_schema_response
- Reuses ResponseProcessor's parse_tool_calls() for model-predicted path
- Content blocks ordered per Anthropic convention: Thinking first, Text,
  then ToolUse blocks
- Tool calls parsed as OpenAI ToolCall (via existing parsers) then
  converted to ContentBlock::ToolUse with JSON input
- Messages always n=1, no logprobs
- ThinkingConfig::Enabled check replaces separate_reasoning bool

Refs: #739, #741, #744
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

grpc gRPC client and router changes model-gateway Model gateway crate changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants