Skip to content

fix(trtllm): tokenize and inject user stop sequences for TRT-LLM requests - #346

Merged
CatherineSue merged 3 commits into
mainfrom
str
Feb 11, 2026
Merged

CatherineSue merged 3 commits into
mainfrom
str

Conversation

@ppraneth

@ppraneth ppraneth commented Feb 6, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Problem

closes #221

TRT-LLM requires stop words as tokenized TokenSequence entries (token IDs), unlike SGLang/vLLM which handle stop strings at the client level. Previously, the extract_stop_words method in TrtllmServiceClient was a no-op that always returned an empty vec, meaning user-provided stop sequences were silently ignored.

Solution Implementation

Stop sequences are now tokenized and injected during the Request Building Stage, right after the proto request is constructed and before it is dispatched. A single helper function handles both tokenization and injection in one step, keeping the logic encapsulated and avoiding leaky abstractions.

Key Changes:

  1. Removed dead code in grpc_client

    • Deleted the no-op extract_stop_words method from TrtllmServiceClient — it always returned vec![] with a comment saying "the router should handle this".
    • The call site now uses vec![] directly with a comment explaining that stop words are injected by the router post-build.
  2. Single-responsibility helper: inject_trtllm_stop_words

    • Added utils::inject_trtllm_stop_words(&mut req, tokenizer, stop) that takes a mutable reference to the TRT-LLM request, tokenizes each stop string, and pushes TokenSequence entries directly — no intermediate allocations or conversions needed at the call site.
  3. Request Building Stages

    • Both Chat and Generate request-building stages call inject_trtllm_stop_words after constructing the proto request, gated on ProtoGenerateRequest::Trtllm and the presence of a tokenizer + stop sequences.
  4. E2E Tests

    • Added test_stop_sequences and test_stop_sequences_stream to the chat completions e2e test suite, verifying that stop sequences cause generation to halt (both non-streaming and streaming) and that the stop token does not appear in the output.

Verification

  • cargo clippy --all-targets --all-features -- -D warnings passes
  • cargo +nightly fmt passes
  • E2E tests added for stop sequences (non-streaming and streaming)
Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated

@github-actions github-actions Bot added the model-gateway Model gateway crate changes label Feb 6, 2026
@coderabbitai

coderabbitai Bot commented Feb 6, 2026 •

Copy link
Copy Markdown

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

This PR implements stop sequence tokenization and injection for TRT-LLM requests across chat and generate request builders. A utility function tokenizes stop sequences into token IDs, which are then injected into the proto request's stop_words field. Tests and utility logic for reasoning parser handling were also added.

Changes

Cohort / File(s) Summary
Test Coverage
model_gateway/src/routers/grpc/regular/stages/generate/preparation.rs
Added test module with MockTokenizer implementation and tokio test for GeneratePreparationStage stop sequences validation.
Stop Sequence Injection
model_gateway/src/routers/grpc/regular/stages/chat/request_building.rs, model_gateway/src/routers/grpc/regular/stages/generate/request_building.rs
Injected on-the-fly tokenization of stop sequences into TRT-LLM proto requests. Tokenizes stop sequences from the original request and appends as TokenSequence.token_ids to stop_words. Logs warning if tokenizer unavailable.
Utility Functions
model_gateway/src/routers/grpc/utils.rs
Added public helper function tokenize_stop_sequences() to convert stop sequences to token ID vectors. Updated create_reasoning_parser() return type from ReasoningPooledParser to Option<Box<dyn ReasoningParser>> with fallback registry logic. Expanded test coverage for new tokenization path.

Estimated Code Review Effort

🎯 3 (Moderate) | ⏱️ ~22 minutes

Possibly Related Issues

Poem

🐰 Hopping through stop words with glee,
Tokenizing sequences, one, two, three,
TRT-LLM requests now know when to cease,
A rabbit's gift of completion peace! ✨

🚥 Pre-merge checks | ✅ 4 | ❌ 1
❌ Failed checks (1 inconclusive)
Check name Status Explanation Resolution
Out of Scope Changes check ❓ Inconclusive All changes are scoped to stop sequence tokenization and injection for TRT-LLM. The updates to create_reasoning_parser signature appear necessary for compilation compatibility but represent a minor signature change outside core objectives. Clarify whether the create_reasoning_parser signature change is required for this PR or if it should be addressed in a separate commit for better focus.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title accurately summarizes the main change: tokenizing and injecting stop sequences for TRT-LLM requests, which is the primary objective of the PR.
Linked Issues check ✅ Passed The PR implements all key objectives from issue #221: tokenizing stop sequences using the router-level tokenizer, filtering empty sequences, propagating them through contexts, and injecting into TRT-LLM requests.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch str

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello @ppraneth, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request enhances the TRT-LLM request handling within the Harmony pipeline by introducing robust support for user-defined stop sequences. It ensures that custom stop strings provided by users are correctly tokenized upstream and then seamlessly injected into the TRT-LLM generation requests, providing more precise control over model output termination.

Highlights

  • User Stop Sequence Tokenization: Implemented logic within the Harmony pipeline to tokenize user-defined stop strings using the model's tokenizer, including filtering out any empty token sequences.
  • Pipeline Propagation: The tokenized stop sequences are now propagated through the Harmony build output and the preparation context.
  • TRT-LLM Injection: These pre-tokenized sequences are injected into the stop_words field of the TRT-LLM generate request during final request construction.
  • Helper Function & Unit Tests: A new helper function was added to extract and tokenize stop sequences, and comprehensive unit tests cover single string, array, and empty stop inputs.
  • Expanded Data Structures: The PreparationOutput and HarmonyBuildOutput structs were extended to accommodate the new harmony_stop_sequences field.
Changelog
  • model_gateway/src/routers/grpc/context.rs
    • Added harmony_stop_sequences: Option<Vec<Vec<u32>>> to PreparationOutput struct.
  • model_gateway/src/routers/grpc/harmony/builder.rs
    • Introduced extract_and_tokenize_stop_sequences private helper function to handle tokenization of StringOrArray stop inputs.
    • Modified the chat-based HarmonyBuilder to utilize the new helper for additional_stop_sequences.
    • Added a placeholder additional_stop_sequences: vec![] for response-based builders, noting current lack of support.
    • Included unit tests for extract_and_tokenize_stop_sequences covering various input types.
  • model_gateway/src/routers/grpc/harmony/stages/preparation.rs
    • Updated HarmonyPreparationStage to pass additional_stop_sequences from HarmonyBuildOutput to PreparationOutput.
  • model_gateway/src/routers/grpc/harmony/stages/request_building.rs
    • Enhanced TRT-LLM request building to iterate through and inject harmony_stop_sequences into the req.stop_words field.
    • Updated debug logging to reflect the total count of injected stop words.
  • model_gateway/src/routers/grpc/harmony/types.rs
    • Added additional_stop_sequences: Vec<Vec<u32>> to HarmonyBuildOutput struct.
  • model_gateway/src/routers/grpc/regular/stages/chat/preparation.rs
    • Initialized the new harmony_stop_sequences field to None in the ChatPreparationStage.
  • model_gateway/src/routers/grpc/regular/stages/embedding/preparation.rs
    • Initialized the new harmony_stop_sequences field to None in the EmbeddingPreparationStage.
  • model_gateway/src/routers/grpc/regular/stages/generate/preparation.rs
    • Initialized the new harmony_stop_sequences field to None in the GeneratePreparationStage.
Activity
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request correctly implements the tokenization and injection of user-defined stop sequences for TRT-LLM requests. The changes are well-structured, propagating the tokenized sequences through the Harmony pipeline. I've included a couple of suggestions to improve code clarity and conciseness by using more idiomatic Rust iterator patterns. Overall, this is a solid implementation that addresses the issue effectively.

Comment thread model_gateway/src/routers/grpc/harmony/builder.rs Outdated
Comment thread model_gateway/src/routers/grpc/harmony/stages/request_building.rs

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Fix all issues with AI agents
In `@model_gateway/src/routers/grpc/harmony/builder.rs`:
- Around line 253-254: build_from_responses() is silently dropping user-supplied
stop sequences by hardcoding additional_stop_sequences: vec![]; replicate the
chat-path behavior by extracting and tokenizing the stop sequences from
request.stop (the ResponsesRequest.stop: Option<StringOrArray>) using
extract_and_tokenize_stop_sequences(request.stop.as_ref()) and set
additional_stop_sequences to the result, or if Responses cannot support stops,
validate request.stop and return an explicit error/log; update the function
build_from_responses() to reference extract_and_tokenize_stop_sequences and
populate additional_stop_sequences accordingly (or validate and fail fast).

In `@model_gateway/src/routers/grpc/harmony/stages/request_building.rs`:
- Around line 251-274: The user stop sequences are only added when the internal
Harmony stops exist because the user-stop injection is inside the if let
Some(harmony_stops) guard; move the user stop injection out so it always runs:
keep the loop that pushes internal Harmony stop tokens inside the if let
Some(harmony_stops) block (iterating over harmony_stops and pushing
TokenSequence into req.stop_words), then after that block (outside it) check
prep.harmony_stop_sequences and extend req.stop_words with mapped TokenSequence
entries from each seq, and update the debug logging to report the final
total_stop_count and user_stop_count where appropriate; reference symbols:
harmony_stops, prep.harmony_stop_sequences, req.stop_words, TokenSequence.

Comment thread model_gateway/src/routers/grpc/harmony/builder.rs Outdated
Comment thread model_gateway/src/routers/grpc/harmony/stages/request_building.rs Outdated

@CatherineSue CatherineSue left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think the original issues means the regular gRPC. This PR focuses on Harmony pipeline.

Comment thread model_gateway/src/routers/grpc/harmony/stages/preparation.rs Outdated

@CatherineSue CatherineSue left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A few notes:

  1. Harmony doesn't support user-provided stop strings — this is intentional for now, even for SGLang/vLLM. The Harmony pipeline only injects its own internal stop tokens (<|return|>, <|call|>). We can revisit later if needed.

  2. Please continue working on this PR to make it consistent with issue #221. The fix should handle stop sequences at the router level (post-build injection into the proto request) rather than threading a new field through pipeline structs. See how the regular gRPC path and Harmony path already do post-build injection for reference.

  3. After making relevant changes, we should turn on relevant unit tests if possible in the CI, or add them. Make sure the stop sequence tokenization and injection logic is covered by tests.

@ppraneth
ppraneth marked this pull request as draft February 7, 2026 04:00
@ppraneth
ppraneth marked this pull request as ready for review February 7, 2026 06:13
@ppraneth
ppraneth requested a review from CatherineSue February 7, 2026 06:14
Comment thread model_gateway/src/routers/grpc/context.rs Outdated

@CatherineSue CatherineSue left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A few issues with the current state:

  1. Dead code in trtllm_service.rs: extract_stop_words (line 482) takes a request, ignores it, and returns vec![]. Now that stop sequences are tokenized and injected in
    the request building stage, this dead code should be cleaned up — remove the method and replace the call site with let stop_words = vec![]; directly.

  2. Silent error swallowing: tokenize_stop_sequences drops encoding failures silently via .filter_map(Result::ok). Please add a warn! per failed encoding so we have
    visibility in production.

  3. Remove tests: The 142-line test in generate/preparation.rs is vestigial — it no longer tests stop sequences after the logic moved to request building. The
    MockTokenizer is also duplicated across two files. Please drop both.

Comment thread model_gateway/src/routers/grpc/regular/stages/chat/request_building.rs Outdated
Comment thread model_gateway/src/routers/grpc/regular/stages/chat/request_building.rs Outdated
Comment thread model_gateway/src/routers/grpc/regular/stages/generate/request_building.rs Outdated
@github-actions github-actions Bot added the grpc gRPC client and router changes label Feb 11, 2026
@CatherineSue

CatherineSue commented Feb 11, 2026 •

Copy link
Copy Markdown
Member

Hey @ppraneth, thanks for working on this — I appreciate the effort you put in to tackle #221.

After reviewing the changes, I have a few concerns I want to share:

1. Leaky abstraction with tokenize_stop_sequences

The current approach splits the operation into two parts: tokenize_stop_sequences() returns raw Option<Vec<Vec<u32>>>, and then every call site has to manually convert those into TokenSequence structs:

req.stop_words
    .extend(user_stops.iter().map(|seq| TokenSequence {
        token_ids: seq.clone(),
    }));

This forces the caller to know about TokenSequence internals and duplicates the conversion logic in both chat and generate request-building stages. A single helper that takes &mut TrtllmGenerateRequest directly (tokenize + inject in one step) would be cleaner — the caller shouldn't need to care about the intermediate representation.

2. Unnecessary allocation

.iter().map(|seq| seq.clone()) copies every Vec<u32> for no reason. If the helper owned the injection end-to-end, it could just push each TokenSequence directly from the encoding result without the extra clone.

3. The resolve_tokenizer caching early-return

if let Some(tokenizer) = &ctx.state.tokenizer {
    return Ok(tokenizer.clone());
}

This is grafted onto resolve_tokenizer() as a cache check, but the function already sets ctx.state.tokenizer at the end of its normal flow. The simpler approach is to just access ctx.state.tokenizer directly at the injection site (it's already resolved by the preparation stage). No need to modify the shared utility.


I know this has taken a fair amount of your time already, and I really do appreciate you digging into this. I had actually been working on the same fix on a separate branch (chang/trt-stop) with a slightly different approach that addresses the points above, so I'm going to go ahead and land that version on this PR to keep things moving. Hope that's okay — happy to walk through the changes together if you'd like.

@CatherineSue
CatherineSue merged commit dcb06ee into main Feb 11, 2026
55 of 59 checks passed
@CatherineSue
CatherineSue deleted the str branch February 11, 2026 23:28
ppraneth added a commit that referenced this pull request Feb 18, 2026
…ests (#346)

Co-authored-by: Chang Su <chang.s.su@oracle.com>
Signed-off-by: ppraneth <pranethparuchuri@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci CI/CD configuration changes grpc gRPC client and router changes model-gateway Model gateway crate changes tests Test changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

TRT-LLM: stop sequences (stop strings) not mapped in chat completions

2 participants