Skip to content

feat(gateway): add MessageRequestBuildingStage for Messages API - #744

Merged
slin1237 merged 1 commit into
mainfrom
slin/msg-3
Mar 12, 2026
Merged

slin1237 merged 1 commit into
mainfrom
slin/msg-3

Conversation

@slin1237

@slin1237 slin1237 commented Mar 12, 2026 •

Copy link
Copy Markdown
Member

Summary

  • Add Stage 4 (MessageRequestBuildingStage) for the Messages API gRPC pipeline
  • Add build_generate_request_from_messages + sampling params builders to all 3 backends (sglang, vLLM, TRT-LLM)
  • Add build_messages_request dispatcher to GrpcClient

What changed

New file:

  • model_gateway/src/routers/grpc/regular/stages/messages/request_building.rs — MessageRequestBuildingStage, copied from chat's request_building.rs and adapted: msg_{uuid} request ID prefix, ctx.messages_request_arc(), build_messages_request(), no filtered_request pattern, multimodal postponed

Backend sampling params (3 files in crates/grpc_client/src/):

  • sglang_scheduler.rs — build_generate_request_from_messages() + build_grpc_sampling_params_from_messages(). Maps CreateMessageRequest fields to sglang proto SamplingParams (top_k=-1 for disabled)
  • vllm_engine.rs — Same pattern for vLLM (top_k=0 for disabled, Option<f32> temperature)
  • trtllm_service.rs — Same pattern using TRT-LLM's SamplingConfig + OutputConfig + GuidedDecodingParams proto types

GrpcClient dispatcher:

  • model_gateway/src/routers/grpc/client.rs — build_messages_request() dispatches to each backend's build_generate_request_from_messages()

Module wiring:

  • model_gateway/src/routers/grpc/regular/stages/messages/mod.rs — Wire request_building module + re-export

Why

PR 3 in the Messages API gRPC series. Bridges preparation (Stage 1, #741) and response processing (Stage 7, future PR). CreateMessageRequest has fewer sampling knobs than ChatCompletionRequest — no min_p, frequency_penalty, presence_penalty, repetition_penalty, n, logprobs, response_format/ebnf/regex. Constraints are limited to tool_call_constraint from the preparation stage (same pattern as Responses API).

Test plan

  • cargo clippy -p smg --all-targets --all-features -- -D warnings — passes
  • cargo clippy -p smg-grpc-client --all-targets --all-features -- -D warnings — passes
  • cargo fmt --check — passes
  • Stage is not yet wired into pipeline factory (follow-up PR), so #![allow(dead_code)] is used

Refs: #739, #741

Summary by CodeRabbit

  • New Features
    • Extended gRPC backend support to handle Anthropic Messages API format
    • All backends (SGLang, TRTLlm, vLLM) now process message-based requests with full sampling configuration and constraint handling
    • Improved routing and pipeline integration for message-based generation requests

…arams for Messages API

Add Stage 4 (request building) for the Messages API gRPC pipeline,
converting PreparationOutput + CreateMessageRequest sampling parameters
into backend-specific proto GenerateRequest.

What changed:
- model_gateway/src/routers/grpc/regular/stages/messages/request_building.rs:
  New MessageRequestBuildingStage (copied from chat, adapted for Messages).
  Uses msg_{uuid} request ID prefix, calls build_messages_request(),
  skips multimodal (postponed), no filtered_request pattern.
- model_gateway/src/routers/grpc/client.rs:
  Add build_messages_request() dispatcher on GrpcClient enum, dispatching
  to each backend's build_generate_request_from_messages().
- crates/grpc_client/src/sglang_scheduler.rs:
  Add build_generate_request_from_messages() and
  build_grpc_sampling_params_from_messages(). Maps CreateMessageRequest
  fields (max_tokens, temperature, top_p, top_k, stop_sequences) to
  sglang proto SamplingParams with sensible defaults for missing fields.
- crates/grpc_client/src/vllm_engine.rs:
  Same pattern for vLLM backend. Handles vLLM-specific differences
  (top_k=0 for disabled, Optional<f32> temperature).
- crates/grpc_client/src/trtllm_service.rs:
  Same pattern for TRT-LLM backend using SamplingConfig, OutputConfig,
  and GuidedDecodingParams proto types.
- model_gateway/src/routers/grpc/regular/stages/messages/mod.rs:
  Wire request_building module and re-export MessageRequestBuildingStage.

Why: This is PR 3 in the Messages API gRPC series. Stage 4 bridges
the gap between preparation (Stage 1, PR #741) and response processing
(Stage 7, future PR), enabling the pipeline to build backend-specific
proto requests from Messages API parameters.

Refs: #739, #741
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly advances the Messages API gRPC series by integrating the request building logic. It establishes the necessary infrastructure to translate incoming CreateMessageRequest objects into backend-specific gRPC GenerateRequest formats, bridging the gap between initial request preparation and subsequent response processing. This work is crucial for enabling the full functionality of the new Messages API within the gateway.

Highlights

  • Messages API Request Building Stage: Implemented MessageRequestBuildingStage as Stage 4 for the Messages API gRPC pipeline, responsible for constructing backend-specific GenerateRequest objects.
  • Backend Integration for Messages API: Introduced build_generate_request_from_messages and associated sampling parameter builders across Sglang, vLLM, and TRT-LLM backends to support the new Messages API requests.
  • gRPC Client Dispatcher Update: Added a build_messages_request dispatcher to the GrpcClient to route Messages API requests to the appropriate backend client.
Changelog
  • crates/grpc_client/src/sglang_scheduler.rs
    • Added CreateMessageRequest import.
    • Implemented build_generate_request_from_messages to construct Sglang-specific GenerateRequest from CreateMessageRequest.
    • Implemented build_grpc_sampling_params_from_messages to map CreateMessageRequest sampling parameters to Sglang's SamplingParams.
  • crates/grpc_client/src/trtllm_service.rs
    • Added CreateMessageRequest import.
    • Implemented build_generate_request_from_messages to construct TRT-LLM-specific GenerateRequest from CreateMessageRequest.
    • Implemented build_sampling_config_from_messages to map CreateMessageRequest sampling parameters to TRT-LLM's SamplingConfig.
  • crates/grpc_client/src/vllm_engine.rs
    • Added CreateMessageRequest import.
    • Implemented build_generate_request_from_messages to construct vLLM-specific GenerateRequest from CreateMessageRequest.
    • Implemented build_grpc_sampling_params_from_messages to map CreateMessageRequest sampling parameters to vLLM's SamplingParams.
  • model_gateway/src/routers/grpc/client.rs
    • Added CreateMessageRequest import.
    • Added build_messages_request method to dispatch CreateMessageRequest to the appropriate backend client (Sglang, vLLM, or TRT-LLM).
  • model_gateway/src/routers/grpc/regular/stages/messages/mod.rs
    • Updated module to include and re-export the new request_building module and MessageRequestBuildingStage.
  • model_gateway/src/routers/grpc/regular/stages/messages/request_building.rs
    • Created a new file defining MessageRequestBuildingStage.
    • Implemented the execute method for MessageRequestBuildingStage to build backend-specific GenerateRequest from CreateMessageRequest and preparation output.
Activity
  • The cargo clippy -p smg --all-targets --all-features -- -D warnings command passed.
  • The cargo clippy -p smg-grpc-client --all-targets --all-features -- -D warnings command passed.
  • The cargo fmt --check command passed.
  • The newly added stage is currently marked with #![allow(dead_code)] as it is not yet wired into the pipeline factory, pending a follow-up PR.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@coderabbitai

coderabbitai Bot commented Mar 12, 2026 •

Copy link
Copy Markdown

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 9920c93f-0fd5-4224-a0d1-b8d1ee0ebf71

📥 Commits

Reviewing files that changed from the base of the PR and between 57a9db1 and 6cc1356.

📒 Files selected for processing (6)
  • crates/grpc_client/src/sglang_scheduler.rs
  • crates/grpc_client/src/trtllm_service.rs
  • crates/grpc_client/src/vllm_engine.rs
  • model_gateway/src/routers/grpc/client.rs
  • model_gateway/src/routers/grpc/regular/stages/messages/mod.rs
  • model_gateway/src/routers/grpc/regular/stages/messages/request_building.rs

📝 Walkthrough

Walkthrough

This PR introduces support for the Anthropic Messages API path in the gRPC request-building pipeline by adding build_generate_request_from_messages methods to three backend clients (SGLang, TRT-LLM, vLLM), a router coordination method, and a new MessageRequestBuildingStage pipeline component that transforms prepared message requests into backend-specific GenerateRequest objects.

Changes

Cohort / File(s) Summary
Backend GenerateRequest Builders
crates/grpc_client/src/sglang_scheduler.rs, crates/grpc_client/src/trtllm_service.rs, crates/grpc_client/src/vllm_engine.rs
Added public build_generate_request_from_messages method and private helper for building sampling parameters from CreateMessageRequest. Each method assembles tokenized input, sampling config, stop sequences, multimodal inputs, and constraint handling following existing chat builder patterns.
Router Request Coordination
model_gateway/src/routers/grpc/client.rs
Added build_messages_request method to dispatch CreateMessageRequest to the appropriate backend variant (Sglang, Vllm, Trtllm) and delegate to respective build_generate_request_from_messages methods, mirroring existing chat request logic.
Messages Pipeline Stage
model_gateway/src/routers/grpc/regular/stages/messages/mod.rs, model_gateway/src/routers/grpc/regular/stages/messages/request_building.rs
Introduced new MessageRequestBuildingStage that validates preparation state, selects appropriate gRPC client, generates request IDs, builds proto GenerateRequest via router coordination, handles build errors, optionally injects PD metadata, and stores result in context.

Sequence Diagram

sequenceDiagram
    participant Client as HTTP Client
    participant Router as GrpcClient Router
    participant Stage as MessageRequestBuildingStage
    participant Backend as Backend Client<br/>(SGLang/TRT-LLM/vLLM)
    participant GrpcService as gRPC Service

    Client->>Router: CreateMessageRequest
    Router->>Stage: execute(RequestContext)
    
    Stage->>Stage: Validate preparation & clients
    Stage->>Stage: Select backend client
    Stage->>Stage: Generate request_id
    
    Stage->>Router: build_messages_request()
    Router->>Backend: build_generate_request_from_messages()
    
    Backend->>Backend: Build sampling params<br/>Handle constraints<br/>Include multimodal inputs
    Backend-->>Router: ProtoGenerateRequest
    
    Router-->>Stage: GenerateRequest
    Stage->>Stage: Optionally inject PD metadata
    Stage->>Stage: Store in context.proto_request
    
    Stage-->>Client: Continue pipeline<br/>with generated request
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

  • PR #741: Introduces MessagePreparationStage that produces the processed_messages consumed by this PR's new request-building functionality.
  • PR #739: Also implements Anthropic Messages API support and overlaps in RequestType::Messages handling and CreateMessageRequest integration points.
  • PR #606: Modifies gRPC request-building and GenerateRequest construction logic including stop/bad-word handling that relates to the constraint integration in this PR.

Suggested labels

grpc, model-gateway

Suggested reviewers

  • key4ng
  • CatherineSue

Poem

🐰 A rabbit hops through gRPC gates,
Where Messages now find their mates!
From request prep to backend's door,
The pipeline flows like never before! ✨

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'feat(gateway): add MessageRequestBuildingStage for Messages API' clearly and specifically summarizes the main change: adding a new MessageRequestBuildingStage to the gateway for Messages API support.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
  • 📝 Generate docstrings (stacked PR)
  • 📝 Generate docstrings (commit on current branch)
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch slin/msg-3
📝 Coding Plan
  • Generate coding plan for human review comments

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions github-actions Bot added grpc gRPC client and router changes model-gateway Model gateway crate changes labels Mar 12, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces the MessageRequestBuildingStage for the Messages API, adding the necessary request-building logic for Sglang, vLLM, and TRT-LLM backends. The changes are well-structured and follow the existing patterns in the codebase for other API endpoints. The new stage correctly prepares the GenerateRequest for each backend based on the CreateMessageRequest.

My main feedback is regarding code duplication in the GrpcClient dispatcher, where the logic for building requests for different backends is highly repetitive. I've left a suggestion to refactor this to improve maintainability.

Overall, this is a solid contribution that moves the Messages API implementation forward.

Comment thread model_gateway/src/routers/grpc/client.rs

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 6cc1356666

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment on lines +655 to +656
let skip_special_tokens =
tool_call_constraint.is_none() && request.tools.as_ref().is_none_or(|t| t.is_empty());

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Honor tool_choice none when setting skip_special_tokens

This logic treats any non-empty request.tools as a signal to keep special tokens, even when tool use is effectively disabled (for example tool_choice: none, which yields no tool constraint) or when tools were filtered out earlier for gRPC use. In those cases generation is plain text, but skip_special_tokens is forced to false, which can leak backend control/special tokens into user-visible output; the same pattern is also present in the new vLLM Messages builder.

Useful? React with 👍 / 👎.

Ok(proto::SamplingParams {
temperature: request.temperature.unwrap_or(1.0) as f32,
top_p: request.top_p.unwrap_or(1.0) as f32,
top_k: request.top_k.map(|v| v as i32).unwrap_or(-1),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Validate top_k before casting to signed backend fields

CreateMessageRequest.top_k is an unsigned integer, but this cast writes it into a signed i32 protobuf field with as; values above i32::MAX will wrap to negative numbers and silently change sampling behavior (including accidental disablement/invalid values). This should be range-checked before conversion; the same lossy cast is also introduced in the new TRT-LLM Messages sampling config path.

Useful? React with 👍 / 👎.

@slin1237
slin1237 merged commit 4d4d1d2 into main Mar 12, 2026
32 checks passed
@slin1237
slin1237 deleted the slin/msg-3 branch March 12, 2026 20:35
slin1237 added a commit that referenced this pull request Mar 12, 2026
…on-streaming)

Add Stage 7 (response processing) for the Messages API gRPC pipeline.
This converts backend ProtoGenerateComplete responses into Anthropic
Message format with proper ContentBlock construction and StopReason
mapping. Non-streaming only; streaming deferred to follow-up PR.

What changed:
- processor.rs: add process_non_streaming_messages_response() to
  ResponseProcessor — full pipeline: token decoding, reasoning parsing,
  tool call parsing, content block construction (Thinking → Text →
  ToolUse), StopReason mapping (EndTurn/MaxTokens/StopSequence/ToolUse),
  and messages::Usage building
- messages/response_processing.rs: new MessageResponseProcessingStage
  that extracts execution result, dispatch metadata, tokenizer, and
  stop decoder from RequestContext, delegates to ResponseProcessor,
  and stores FinalResponse::Messages
- message_utils.rs: add get_history_tool_calls_count_messages() for
  counting tool use blocks in Messages API request history (needed for
  KimiK2-style tool call ID generation)
- messages/mod.rs: wire response_processing module with unused_imports
  expect (wired in pipeline factory PR)

Why:
This is the fourth PR in the Messages API gRPC support series. With
preparation (PR #741), request building (PR #744), and now response
processing, three of the four endpoint-specific pipeline stages are
complete. The shared stages (worker selection, client acquisition,
dispatch, execution) are reused from the existing pipeline.

How:
Follows the same architecture as chat's response processing but adapted
for Anthropic Message types:
- Reuses existing convert_message_tool_choice() from message_utils to
  bridge Messages ToolChoice → Chat ToolChoice for parse_json_schema_response
- Reuses ResponseProcessor's parse_tool_calls() for model-predicted path
- Content blocks ordered per Anthropic convention: Thinking first, Text,
  then ToolUse blocks
- Tool calls parsed as OpenAI ToolCall (via existing parsers) then
  converted to ContentBlock::ToolUse with JSON input
- Messages always n=1, no logprobs
- ThinkingConfig::Enabled check replaces separate_reasoning bool

Refs: #739, #741, #744
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
@coderabbitai coderabbitai Bot mentioned this pull request Jun 16, 2026
2 of 4 tasks
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

grpc gRPC client and router changes model-gateway Model gateway crate changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant