Skip to content

feat(completions): add CompletionRequestBuildingStage and backend sampling params - #915

Merged
slin1237 merged 1 commit into
mainfrom
mourya/cmp-3
Mar 27, 2026
Merged

slin1237 merged 1 commit into
mainfrom
mourya/cmp-3

Conversation

@vschandramourya

@vschandramourya vschandramourya commented Mar 26, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Add Stage 4 (CompletionRequestBuildingStage) for the /v1/completions gRPC pipeline
  • Add build_generate_request_from_completion + sampling params builders to all 3 backends (SGLang, vLLM, TRT-LLM)
  • Add build_completion_request dispatcher to GrpcClient

PR 3 in the Completions API gRPC pipeline series. PR 1 was #840 (type scaffolding), PR 2 was #907 (preparation).

What changed

New file

  • model_gateway/src/routers/grpc/regular/stages/completion/request_building.rs — CompletionRequestBuildingStage, parallel to MessageRequestBuildingStage. Uses cmpl_{uuid} request ID prefix, calls build_completion_request(), no multimodal, no tools.

Backend sampling params (3 files in crates/grpc_client/src/)

  • sglang_scheduler.rs — build_generate_request_from_completion() + build_grpc_sampling_params_from_completion(). Maps all CompletionRequest sampling fields (temperature, top_p, top_k, min_p, frequency_penalty, presence_penalty, repetition_penalty, n, logprobs, stop, stop_token_ids, ignore_eos, no_stop_trim) + structured output constraints (json_schema, regex, ebnf).
  • vllm_engine.rs — Same pattern for vLLM. Handles vLLM-specific differences (top_k=0 for disabled, Option<f32> temperature).
  • trtllm_service.rs — Same pattern using TRT-LLM's SamplingConfig + OutputConfig + GuidedDecodingParams proto types. Includes seed, min_tokens, min_p mapping.

GrpcClient dispatcher

  • model_gateway/src/routers/grpc/client.rs — build_completion_request() dispatches to each backend's build_generate_request_from_completion().

Module wiring

  • model_gateway/src/routers/grpc/regular/stages/completion/mod.rs — Wire request_building module + re-export CompletionRequestBuildingStage.

How

Follows the same architecture as Messages request building (#744):

Key difference from Messages: CompletionRequest has richer sampling knobs (frequency_penalty, presence_penalty, repetition_penalty, min_p, n, logprobs, ignore_eos, no_stop_trim) and direct structured output constraints (regex, ebnf, json_schema), but no tools and no multimodal.

Test plan

  • cargo clippy -p smg --all-targets --all-features -- -D warnings — passes
  • cargo clippy -p smg-grpc-client --all-targets --all-features -- -D warnings — passes
  • cargo fmt --check — passes
  • Stage is not yet wired into pipeline factory (follow-up PR), so #![allow(dead_code)] is used

Refs: #840, #907

Summary by CodeRabbit

  • New Features
    • Added end-to-end support for OpenAI-style completions: backend request construction, sampling options, streaming, stop-sequence handling, logprob/hidden-state options, and EOS/stop-token behavior.
    • Single-constraint guided decoding enforced (returns an error for conflicting constraints).
    • Pipeline stage added to build and validate completion requests, with optional metadata injection into backend requests.

@github-actions github-actions Bot added grpc gRPC client and router changes model-gateway Model gateway crate changes labels Mar 26, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request is the third in a series aimed at fully implementing the /v1/completions gRPC pipeline. It establishes the core logic for processing completion requests by introducing a new request building stage and integrating it across various backend clients (SGLang, vLLM, TRT-LLM). The changes enable the system to correctly parse and translate completion-specific parameters, including advanced sampling and structured output constraints, into the backend's native GenerateRequest format.

Highlights

  • Completions API Pipeline Stage: Introduced CompletionRequestBuildingStage as Stage 4 for the /v1/completions gRPC pipeline, continuing the implementation of the native completions endpoint.
  • Backend Request Builders: Added build_generate_request_from_completion and associated sampling parameter builders to the SGLang, vLLM, and TRT-LLM clients, enabling them to process completion-specific request parameters.
  • gRPC Client Dispatcher: Implemented a build_completion_request dispatcher within the GrpcClient to correctly route and process completion requests to the appropriate backend client.
  • New Completion Stages: Created new modules and stages, CompletionPreparationStage and CompletionRequestBuildingStage, to specifically handle the tokenization, stop decoder creation, and protobuf request building for /v1/completions requests.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@mergify

mergify Bot commented Mar 26, 2026

Copy link
Copy Markdown
Contributor

Hi @vschandramourya, this PR has merge conflicts that must be resolved before it can be merged. Please rebase your branch:

git fetch origin main
git rebase origin/main
# resolve any conflicts, then:
git push --force-with-lease

@mergify mergify Bot added the needs-rebase PR has merge conflicts that need to be resolved label Mar 26, 2026
@coderabbitai

coderabbitai Bot commented Mar 26, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

Adds /v1/completions request-building: backend-specific builders that convert OpenAI CompletionRequest into proto GenerateRequest for Sglang, vLLM, and TRTLLM, plus a pipeline CompletionRequestBuildingStage and a GrpcClient dispatcher to produce ProtoGenerateRequest instances.

Changes

Cohort / File(s) Summary
Backend Completion Builders
crates/grpc_client/src/sglang_scheduler.rs, crates/grpc_client/src/vllm_engine.rs, crates/grpc_client/src/trtllm_service.rs
Added build_generate_request_from_completion(...) on each client plus helpers to translate sampling params, stop sequences, logprobs, constraint selection (json_schema/regex/ebnf mutual exclusion), and other completion-specific flags into backend proto types.
GrpcClient Completion Dispatcher
model_gateway/src/routers/grpc/client.rs
Added build_completion_request(...) that dispatches CompletionRequest -> backend-specific proto builder and wraps result into ProtoGenerateRequest.
Completion Pipeline Stage
model_gateway/src/routers/grpc/regular/stages/completion/mod.rs, model_gateway/src/routers/grpc/regular/stages/completion/request_building.rs
New CompletionRequestBuildingStage pipeline stage: retrieves preparation and clients from context, selects builder (Single/Dual), generates cmpl_<uuid-v7> request id, calls backend builder, optionally injects PD metadata, stores ProtoRequest::Generate, and maps builder errors to bad_request.
Imports / Minor Exports
crates/grpc_client/*, model_gateway/src/routers/grpc/regular/stages/completion/mod.rs
Added openai_protocol::completion::CompletionRequest imports and re-exported CompletionRequestBuildingStage for pipeline wiring.

Sequence Diagram(s)

sequenceDiagram
  participant Client
  participant PipelineStage as CompletionRequestBuildingStage
  participant GrpcClient
  participant Backend as BackendBuilder
  participant Worker as Worker/Engine

  Client->>PipelineStage: incoming /v1/completions request
  PipelineStage->>PipelineStage: extract preparation (original_text, token_ids)
  PipelineStage->>GrpcClient: build_completion_request(request_id, CompletionRequest, original_text, token_ids)
  GrpcClient->>Backend: call build_generate_request_from_completion(...)
  Backend->>Worker: produce proto GenerateRequest -> forward to backend worker/engine
  Worker-->>GrpcClient: ack / response stream
  GrpcClient-->>PipelineStage: proto request created
  PipelineStage-->>Client: continue pipeline (proto stored)
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

Suggested reviewers

  • CatherineSue
  • key4ng
  • slin1237

Poem

🐰 twitches whiskers
I hop from token to token bright,
Spinning completions into light,
Three backends hum, the pipeline sings,
Proto dreams on tiny wings 🥕✨

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 47.62% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately captures the main changes: adding CompletionRequestBuildingStage and backend sampling parameters for completions processing.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch mourya/cmp-3

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces comprehensive support for the /v1/completions API endpoint across Sglang, Trtllm, and Vllm backends. It includes adding CompletionRequest handling to the gRPC client implementations, mapping completion parameters to internal GenerateRequest structures and sampling configurations. New pipeline stages, CompletionPreparationStage and CompletionRequestBuildingStage, have been integrated into the model gateway to process these requests, managing tokenization, stop decoder creation, and constructing backend-specific gRPC requests. Feedback suggests refactoring for improved code conciseness, efficiency, and clarity, particularly by utilizing StringOrArray::to_vec() for stop sequence initialization and optimizing constraint-building functions to reduce allocations and cloning.

Comment thread crates/grpc_client/src/sglang_scheduler.rs
Comment thread crates/grpc_client/src/sglang_scheduler.rs
Comment thread crates/grpc_client/src/trtllm_service.rs
Comment thread crates/grpc_client/src/trtllm_service.rs
Comment thread crates/grpc_client/src/vllm_engine.rs
Comment thread crates/grpc_client/src/vllm_engine.rs
@mergify mergify Bot removed the needs-rebase PR has merge conflicts that need to be resolved label Mar 26, 2026
@vschandramourya vschandramourya changed the title Mourya/cmp 3 feat(completions): add CompletionRequestBuildingStage and backend sampling params Mar 26, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@crates/grpc_client/src/sglang_scheduler.rs`:
- Around line 717-735: CompletionRequest.min_tokens is not being propagated into
proto::SamplingParams, so completions can end before the caller requested;
update the SamplingParams construction in sglang_scheduler to map
request.min_tokens into the sampling params (set
proto::SamplingParams.min_new_tokens to request.min_tokens.unwrap_or_default()
or an explicit 0 where appropriate) alongside the existing max_new_tokens
mapping, using the same pattern as other optional fields so min_tokens is
honored by the backend.

In `@crates/grpc_client/src/trtllm_service.rs`:
- Around line 750-777: The GenerateRequest currently hardcodes
include_stop_token_in_output to false so CompletionRequest.no_stop_trim is
ignored; update the construction of proto::GenerateRequest in trtllm_service.rs
to read the no_stop_trim flag from the incoming CompletionRequest (e.g.,
body.no_stop_trim) and set include_stop_token_in_output accordingly (true when
no_stop_trim is true, fallback to false when absent) so the
include_stop_token_in_output field is threaded through to TensorRT-LLM.

In `@crates/grpc_client/src/vllm_engine.rs`:
- Around line 648-666: The proto::SamplingParams construction is missing mapping
for CompletionRequest.min_tokens and CompletionRequest.no_stop_trim, so add
sampling.min_tokens and sampling.include_stop_str_in_output to the returned
proto::SamplingParams: set min_tokens from the request's min_tokens (use the
request default when absent) and set include_stop_str_in_output to the logical
inverse of no_stop_trim (i.e., include_stop_str_in_output = !no_stop_trim,
handling the Option properly), ensuring the field names min_tokens and
include_stop_str_in_output appear in the struct literal alongside the existing
fields like temperature, top_p, stop, and ignore_eos.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: df880bfe-fa72-4d6a-8ff0-3880903832da

📥 Commits

Reviewing files that changed from the base of the PR and between 6baac44 and 2674efa.

📒 Files selected for processing (6)
  • crates/grpc_client/src/sglang_scheduler.rs
  • crates/grpc_client/src/trtllm_service.rs
  • crates/grpc_client/src/vllm_engine.rs
  • model_gateway/src/routers/grpc/client.rs
  • model_gateway/src/routers/grpc/regular/stages/completion/mod.rs
  • model_gateway/src/routers/grpc/regular/stages/completion/request_building.rs

Comment thread crates/grpc_client/src/sglang_scheduler.rs
Comment thread crates/grpc_client/src/trtllm_service.rs
Comment thread crates/grpc_client/src/vllm_engine.rs
…pling params

Signed-off-by: VS Chandra Mourya <msrinivasa@together.ai>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@model_gateway/src/routers/grpc/client.rs`:
- Around line 387-423: The three backend builders disagree on how to handle
CompletionRequest.max_tokens == None, so in build_completion_request you must
normalize max_tokens once before calling
client.build_generate_request_from_completion for Sglang, Vllm and Trtllm:
compute an effective_max_tokens (e.g., let normalized =
body.max_tokens.or(Some(<agreed-default>)) or explicitly keep None if that is
the agreed behavior) and pass a modified/temporary CompletionRequest (or the
normalized value) into each client's build_generate_request_from_completion call
so all branches (Sglang, Vllm, Trtllm) receive the same effective max_tokens
value when constructing the ProtoGenerateRequest.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 9f013ae6-9d72-4240-889a-5b786448c7d6

📥 Commits

Reviewing files that changed from the base of the PR and between 2674efa and 3b8ac76.

📒 Files selected for processing (6)
  • crates/grpc_client/src/sglang_scheduler.rs
  • crates/grpc_client/src/trtllm_service.rs
  • crates/grpc_client/src/vllm_engine.rs
  • model_gateway/src/routers/grpc/client.rs
  • model_gateway/src/routers/grpc/regular/stages/completion/mod.rs
  • model_gateway/src/routers/grpc/regular/stages/completion/request_building.rs

Comment thread model_gateway/src/routers/grpc/client.rs
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

grpc gRPC client and router changes model-gateway Model gateway crate changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants