feat(gateway): add MessageRequestBuildingStage for Messages API - #744
Conversation
…arams for Messages API
Add Stage 4 (request building) for the Messages API gRPC pipeline,
converting PreparationOutput + CreateMessageRequest sampling parameters
into backend-specific proto GenerateRequest.
What changed:
- model_gateway/src/routers/grpc/regular/stages/messages/request_building.rs:
New MessageRequestBuildingStage (copied from chat, adapted for Messages).
Uses msg_{uuid} request ID prefix, calls build_messages_request(),
skips multimodal (postponed), no filtered_request pattern.
- model_gateway/src/routers/grpc/client.rs:
Add build_messages_request() dispatcher on GrpcClient enum, dispatching
to each backend's build_generate_request_from_messages().
- crates/grpc_client/src/sglang_scheduler.rs:
Add build_generate_request_from_messages() and
build_grpc_sampling_params_from_messages(). Maps CreateMessageRequest
fields (max_tokens, temperature, top_p, top_k, stop_sequences) to
sglang proto SamplingParams with sensible defaults for missing fields.
- crates/grpc_client/src/vllm_engine.rs:
Same pattern for vLLM backend. Handles vLLM-specific differences
(top_k=0 for disabled, Optional<f32> temperature).
- crates/grpc_client/src/trtllm_service.rs:
Same pattern for TRT-LLM backend using SamplingConfig, OutputConfig,
and GuidedDecodingParams proto types.
- model_gateway/src/routers/grpc/regular/stages/messages/mod.rs:
Wire request_building module and re-export MessageRequestBuildingStage.
Why: This is PR 3 in the Messages API gRPC series. Stage 4 bridges
the gap between preparation (Stage 1, PR #741) and response processing
(Stage 7, future PR), enabling the pipeline to build backend-specific
proto requests from Messages API parameters.
Refs: #739, #741
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request significantly advances the Messages API gRPC series by integrating the request building logic. It establishes the necessary infrastructure to translate incoming Highlights
Changelog
Activity
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension. Footnotes
|
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (6)
📝 WalkthroughWalkthroughThis PR introduces support for the Anthropic Messages API path in the gRPC request-building pipeline by adding Changes
Sequence DiagramsequenceDiagram
participant Client as HTTP Client
participant Router as GrpcClient Router
participant Stage as MessageRequestBuildingStage
participant Backend as Backend Client<br/>(SGLang/TRT-LLM/vLLM)
participant GrpcService as gRPC Service
Client->>Router: CreateMessageRequest
Router->>Stage: execute(RequestContext)
Stage->>Stage: Validate preparation & clients
Stage->>Stage: Select backend client
Stage->>Stage: Generate request_id
Stage->>Router: build_messages_request()
Router->>Backend: build_generate_request_from_messages()
Backend->>Backend: Build sampling params<br/>Handle constraints<br/>Include multimodal inputs
Backend-->>Router: ProtoGenerateRequest
Router-->>Stage: GenerateRequest
Stage->>Stage: Optionally inject PD metadata
Stage->>Stage: Store in context.proto_request
Stage-->>Client: Continue pipeline<br/>with generated request
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~25 minutes Possibly related PRs
Suggested labels
Suggested reviewers
Poem
🚥 Pre-merge checks | ✅ 3✅ Passed checks (3 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches
🧪 Generate unit tests (beta)
📝 Coding Plan
Comment |
There was a problem hiding this comment.
Code Review
This pull request introduces the MessageRequestBuildingStage for the Messages API, adding the necessary request-building logic for Sglang, vLLM, and TRT-LLM backends. The changes are well-structured and follow the existing patterns in the codebase for other API endpoints. The new stage correctly prepares the GenerateRequest for each backend based on the CreateMessageRequest.
My main feedback is regarding code duplication in the GrpcClient dispatcher, where the logic for building requests for different backends is highly repetitive. I've left a suggestion to refactor this to improve maintainability.
Overall, this is a solid contribution that moves the Messages API implementation forward.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 6cc1356666
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| let skip_special_tokens = | ||
| tool_call_constraint.is_none() && request.tools.as_ref().is_none_or(|t| t.is_empty()); |
There was a problem hiding this comment.
Honor tool_choice none when setting skip_special_tokens
This logic treats any non-empty request.tools as a signal to keep special tokens, even when tool use is effectively disabled (for example tool_choice: none, which yields no tool constraint) or when tools were filtered out earlier for gRPC use. In those cases generation is plain text, but skip_special_tokens is forced to false, which can leak backend control/special tokens into user-visible output; the same pattern is also present in the new vLLM Messages builder.
Useful? React with 👍 / 👎.
| Ok(proto::SamplingParams { | ||
| temperature: request.temperature.unwrap_or(1.0) as f32, | ||
| top_p: request.top_p.unwrap_or(1.0) as f32, | ||
| top_k: request.top_k.map(|v| v as i32).unwrap_or(-1), |
There was a problem hiding this comment.
Validate top_k before casting to signed backend fields
CreateMessageRequest.top_k is an unsigned integer, but this cast writes it into a signed i32 protobuf field with as; values above i32::MAX will wrap to negative numbers and silently change sampling behavior (including accidental disablement/invalid values). This should be range-checked before conversion; the same lossy cast is also introduced in the new TRT-LLM Messages sampling config path.
Useful? React with 👍 / 👎.
…on-streaming) Add Stage 7 (response processing) for the Messages API gRPC pipeline. This converts backend ProtoGenerateComplete responses into Anthropic Message format with proper ContentBlock construction and StopReason mapping. Non-streaming only; streaming deferred to follow-up PR. What changed: - processor.rs: add process_non_streaming_messages_response() to ResponseProcessor — full pipeline: token decoding, reasoning parsing, tool call parsing, content block construction (Thinking → Text → ToolUse), StopReason mapping (EndTurn/MaxTokens/StopSequence/ToolUse), and messages::Usage building - messages/response_processing.rs: new MessageResponseProcessingStage that extracts execution result, dispatch metadata, tokenizer, and stop decoder from RequestContext, delegates to ResponseProcessor, and stores FinalResponse::Messages - message_utils.rs: add get_history_tool_calls_count_messages() for counting tool use blocks in Messages API request history (needed for KimiK2-style tool call ID generation) - messages/mod.rs: wire response_processing module with unused_imports expect (wired in pipeline factory PR) Why: This is the fourth PR in the Messages API gRPC support series. With preparation (PR #741), request building (PR #744), and now response processing, three of the four endpoint-specific pipeline stages are complete. The shared stages (worker selection, client acquisition, dispatch, execution) are reused from the existing pipeline. How: Follows the same architecture as chat's response processing but adapted for Anthropic Message types: - Reuses existing convert_message_tool_choice() from message_utils to bridge Messages ToolChoice → Chat ToolChoice for parse_json_schema_response - Reuses ResponseProcessor's parse_tool_calls() for model-predicted path - Content blocks ordered per Anthropic convention: Thinking first, Text, then ToolUse blocks - Tool calls parsed as OpenAI ToolCall (via existing parsers) then converted to ContentBlock::ToolUse with JSON input - Messages always n=1, no logprobs - ThinkingConfig::Enabled check replaces separate_reasoning bool Refs: #739, #741, #744 Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
Summary
build_generate_request_from_messages+ sampling params builders to all 3 backends (sglang, vLLM, TRT-LLM)build_messages_requestdispatcher toGrpcClientWhat changed
New file:
model_gateway/src/routers/grpc/regular/stages/messages/request_building.rs— MessageRequestBuildingStage, copied from chat's request_building.rs and adapted:msg_{uuid}request ID prefix,ctx.messages_request_arc(),build_messages_request(), nofiltered_requestpattern, multimodal postponedBackend sampling params (3 files in
crates/grpc_client/src/):sglang_scheduler.rs—build_generate_request_from_messages()+build_grpc_sampling_params_from_messages(). MapsCreateMessageRequestfields to sglang protoSamplingParams(top_k=-1 for disabled)vllm_engine.rs— Same pattern for vLLM (top_k=0 for disabled,Option<f32>temperature)trtllm_service.rs— Same pattern using TRT-LLM'sSamplingConfig+OutputConfig+GuidedDecodingParamsproto typesGrpcClient dispatcher:
model_gateway/src/routers/grpc/client.rs—build_messages_request()dispatches to each backend'sbuild_generate_request_from_messages()Module wiring:
model_gateway/src/routers/grpc/regular/stages/messages/mod.rs— Wirerequest_buildingmodule + re-exportWhy
PR 3 in the Messages API gRPC series. Bridges preparation (Stage 1, #741) and response processing (Stage 7, future PR).
CreateMessageRequesthas fewer sampling knobs thanChatCompletionRequest— nomin_p,frequency_penalty,presence_penalty,repetition_penalty,n,logprobs,response_format/ebnf/regex. Constraints are limited totool_call_constraintfrom the preparation stage (same pattern as Responses API).Test plan
cargo clippy -p smg --all-targets --all-features -- -D warnings— passescargo clippy -p smg-grpc-client --all-targets --all-features -- -D warnings— passescargo fmt --check— passes#![allow(dead_code)]is usedRefs: #739, #741
Summary by CodeRabbit