Skip to content

fix(grpc): add matched_stop support for vLLM and TensorRT-LLM - #602

Merged
CatherineSue merged 1 commit into
mainfrom
chang/fix-vllm-stop-reason
Mar 3, 2026
Merged

CatherineSue merged 1 commit into
mainfrom
chang/fix-vllm-stop-reason

Conversation

@CatherineSue

@CatherineSue CatherineSue commented Mar 3, 2026 •

Copy link
Copy Markdown
Member

Description

Problem

vLLM gRPC responses were missing stop_reason (which specific stop string/token triggered a stop). vLLM HTTP returns it (e.g., stop_reason=','), but the gRPC proto never defined the field. TensorRT-LLM had optional string stop_reason in its proto but the Rust wrapper ignored it. Only SGLang's oneof matched_stop was wired through.

Solution

  • Add oneof matched_stop { matched_token_id, matched_stop_str } to vLLM's GenerateComplete proto, matching SGLang's pattern to preserve int | str type distinction from vLLM's CompletionOutput.stop_reason
  • Wire both vLLM and TensorRT-LLM through matched_stop_json() in the proto wrapper
  • Remove the now-unused SGLang-only matched_stop() helper

Changes

  • grpc_client/proto/vllm_engine.proto: Add oneof matched_stop (fields 10-11) to GenerateComplete
  • model_gateway/src/routers/grpc/proto_wrapper.rs: Rewrite matched_stop_json() to handle all three backends (SGLang oneof, vLLM oneof, TensorRT-LLM string), remove dead matched_stop() method

Note: The corresponding vLLM grpc_server.py change to populate the new field is in the vLLM fork.

Test Plan

  • cargo build -p smg (proto re-generation + compilation)
  • cargo clippy -p smg -- -D warnings — zero warnings
  • cargo test -p smg — all tests pass
  • pre-commit run --all-files — all hooks pass
  • E2E: verify stop_reason appears in vLLM gRPC chat completion responses with stop sequences
Screenshot 2026-03-03 at 2 09 39 PM
Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated

Summary by CodeRabbit

  • New Features
    • Generation responses now include stop-sequence matching information indicating either a matched token ID or a matched stop string.
    • The prior free-form stop_reason has been replaced with this explicit matched-stop representation in generation results.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Repo admins can enable using credits for code reviews in their settings.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request enhances the gRPC responses for vLLM and TensorRT-LLM by introducing comprehensive support for matched_stop reasons. Previously, vLLM gRPC lacked specific stop reason details, and TensorRT-LLM's stop_reason was not fully utilized. The changes standardize how stop reasons are communicated across different backend engines, improving the clarity and completeness of API responses.

Highlights

  • vLLM gRPC stop_reason support: Add oneof matched_stop to vLLM's GenerateComplete proto to provide detailed stop reason information, aligning with SGLang's existing pattern.
  • Unified matched_stop_json() handling: Wire both vLLM and TensorRT-LLM stop reasons through the matched_stop_json() function in the proto wrapper, enabling consistent handling across different backends.
  • Code cleanup: Remove the now-unused matched_stop() helper method, which was previously only applicable to SGLang.
Changelog
  • grpc_client/proto/vllm_engine.proto
    • Added a oneof matched_stop field to the GenerateComplete message, including matched_token_id and matched_stop_str.
  • model_gateway/src/routers/grpc/proto_wrapper.rs
    • Updated use statements to alias MatchedStop for SGLang and vLLM protos.
    • Removed the matched_stop() method.
    • Modified the matched_stop_json() method to process matched_stop from SGLang and vLLM, and stop_reason from TensorRT-LLM.
Activity
  • Successfully passed cargo build -p smg for proto re-generation and compilation.
  • Achieved zero warnings with cargo clippy -p smg -- -D warnings.
  • All tests passed with cargo test -p smg.
  • All hooks passed with pre-commit run --all-files.
  • Passed cargo +nightly fmt.
  • E2E verification for stop_reason in vLLM gRPC chat completion responses with stop sequences is pending.
  • Documentation was not updated.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@coderabbitai

coderabbitai Bot commented Mar 3, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

Protos changed to represent matched stop data as a oneof (token id or string) in GenerateComplete for multiple backends; the gRPC wrapper removed direct matched_stop access and added a backend-aware matched_stop_json() serializer.

Changes

Cohort / File(s) Summary
vLLM proto
grpc_client/proto/vllm_engine.proto
Added oneof matched_stop to GenerateComplete with uint32 matched_token_id = 10 and string matched_stop_str = 11.
TRT-LLM proto
grpc_client/proto/trtllm_service.proto
Replaced stop_reason with oneof matched_stop containing matched_stop_str (tag 4) and matched_token_id (tag 13); stop_reason removed.
gRPC wrapper
model_gateway/src/routers/grpc/proto_wrapper.rs
Removed matched_stop() accessor; added matched_stop_json() that serializes backend-specific MatchedStop variants (SGLang, vLLM, TRT-LLM) into serde_json::Value (number for token id, string for stop str), using a macro to unify conversions.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Suggested reviewers

  • key4ng
  • slin1237
  • XinyueZhang369

Poem

🐰 A tiny proto hop, a change so neat,
Tokens or strings now tell when we beat,
Wrappers tuck answers into JSON light,
Backends whisper what ended the write,
Hooray — I nibble bugs and dance in delight!

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly identifies the main change: adding matched_stop support for vLLM and TensorRT-LLM in gRPC. It is specific, concise, and directly reflects the core objective of the PR.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
  • 📝 Generate docstrings (stacked PR)
  • 📝 Generate docstrings (commit on current branch)
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch chang/fix-vllm-stop-reason

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions github-actions Bot added grpc gRPC client and router changes model-gateway Model gateway crate changes labels Mar 3, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request adds support for matched_stop in vLLM and TensorRT-LLM gRPC backends, aligning them with the existing SGLang implementation. The changes involve updating the vLLM protobuf definition and modifying the Rust wrapper to handle the new fields. The implementation is clean and correct. The suggestion to refactor a small piece of duplicated code in proto_wrapper.rs to improve maintainability is valid and aligns with best practices.

Comment thread model_gateway/src/routers/grpc/proto_wrapper.rs
vLLM gRPC was missing stop_reason entirely (proto had no field), and
TensorRT-LLM had optional string stop_reason which lost the int type
for token IDs. Add oneof matched_stop { matched_token_id, matched_stop_str }
to both protos (wire-compatible for TRT-LLM) and wire through the Rust
wrapper using a local macro to deduplicate the three backend arms.

Signed-off-by: Chang Su <chang.s.su@oracle.com>
@CatherineSue
CatherineSue force-pushed the chang/fix-vllm-stop-reason branch from d3c536f to b82bb23 Compare March 3, 2026 21:25
@CatherineSue
CatherineSue merged commit 829c9d3 into main Mar 3, 2026
28 checks passed
@CatherineSue
CatherineSue deleted the chang/fix-vllm-stop-reason branch March 3, 2026 22:20
key4ng pushed a commit that referenced this pull request May 1, 2026
…ixes

The previous pin (fd080fc7) is the commit immediately before
lightseekorg/tokenspeed#578, which adds defensive Finished-state
handlers to the scheduler FSM. Without #578 the engine crashes under
retract pressure with:

    RuntimeError: FSM transition invalid:
      event=tokenspeed::fsm::ExtendResultEvent;
      state=tokenspeed::fsm::Finished

Reproduced on the nightly Qwen3-30B-A3B bench: when the host KV cache
fills up and a retract fails, AbortEvent terminalizes the request →
Finished, but overlap scheduling has already dispatched a forward
batch including it, and the late ExtendResultEvent commit hits a
strict FSM handler that throws and kills the scheduler event loop.

Bump to current lightseekorg/tokenspeed main (eabeb106) so we also
pick up #602 (release scheduler slot + cancel non-stream handlers on
client disconnect), which removes the long pre-crash stream of
``Received output for rid=... but the state was deleted in AsyncLLM``
warnings caused by aborted requests still occupying engine slots.

Signed-off-by: Chang Su <chang.s.su@oracle.com>
key4ng added a commit that referenced this pull request May 1, 2026
…ixes

The previous pin (fd080fc7) is the commit immediately before
lightseekorg/tokenspeed#578, which adds defensive Finished-state
handlers to the scheduler FSM. Without #578 the engine crashes under
retract pressure with:

    RuntimeError: FSM transition invalid:
      event=tokenspeed::fsm::ExtendResultEvent;
      state=tokenspeed::fsm::Finished

Reproduced on the nightly Qwen3-30B-A3B bench: when the host KV cache
fills up and a retract fails, AbortEvent terminalizes the request →
Finished, but overlap scheduling has already dispatched a forward
batch including it, and the late ExtendResultEvent commit hits a
strict FSM handler that throws and kills the scheduler event loop.

Bump to current lightseekorg/tokenspeed main (eabeb106) so we also
pick up #602 (release scheduler slot + cancel non-stream handlers on
client disconnect), which removes the long pre-crash stream of
``Received output for rid=... but the state was deleted in AsyncLLM``
warnings caused by aborted requests still occupying engine slots.

Signed-off-by: key4ng <rukeyang@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

grpc gRPC client and router changes model-gateway Model gateway crate changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant