Skip to content

feat(gprc): add TensorRT-LLM multimodal support - #504

Merged
slin1237 merged 1 commit into
mainfrom
chang/mm-6
Feb 22, 2026
Merged

slin1237 merged 1 commit into
mainfrom
chang/mm-6

Conversation

@CatherineSue

@CatherineSue CatherineSue commented Feb 22, 2026 •

Copy link
Copy Markdown
Member

Description

Problem

The gRPC pipeline doesn't support multimodal (image) requests for the TensorRT-LLM backend.

TRT-LLM's LLM API already supports VLM inference when given prompt_token_ids + multi_modal_data.

Solution

Wire TensorRT-LLM multimodal support through the gRPC pipeline using the same approach as vLLM: send unexpanded token IDs + raw image bytes. TRT-LLM's input processor handles hashing, position tracking, and vision encoding server-side.

Changes

  • Proto (grpc_client/proto/trtllm_service.proto): Simplify MultimodalInput to carry only repeated bytes image_data — remove unused multimodal_hashes, multimodal_positions, multimodal_lengths fields
  • Proto wrapper (proto_wrapper.rs): Add into_trtllm_proto() conversion method on MultimodalData
  • gRPC client (trtllm_service.rs): Accept multimodal_input: Option<proto::MultimodalInput> in build_generate_request_from_chat() and wire it into the proto request
  • Client dispatch (client.rs): Convert and pass multimodal data in the TRT-LLM branch
  • Harmony pipeline (harmony/stages/request_building.rs): Pass None for the new multimodal param (Harmony doesn't support multimodal)
  • Doc comment (multimodal.rs): Update process_for_backend() doc to reflect TRT-LLM now uses the vLLM (raw bytes) branch

Test Plan

  • E2E testing with TRT-LLM VLM model (e.g., Qwen2-VL) via chat completions with image content
  • Verify non-multimodal TRT-LLM requests are unaffected (multimodal_input = None)
  • Verify Harmony pipeline compiles and works (passes None for multimodal)
Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated

Summary by CodeRabbit

  • Changes
    • Multimodal inputs now carry raw image bytes instead of separate client-side metadata.
    • Server now performs image hashing, position tracking, and vision encoding.
    • Client submission simplified for non-SGLang backends; unified image handling across backends.
    • Introduces a breaking change to the public multimodal request surface (compatibility impact).

@github-actions github-actions Bot added grpc gRPC client and router changes model-gateway Model gateway crate changes labels Feb 22, 2026
@coderabbitai

coderabbitai Bot commented Feb 22, 2026 •

Copy link
Copy Markdown

No actionable comments were generated in the recent review. 🎉


📝 Walkthrough

Walkthrough

Replaces per-image metadata fields in the TRT-LLM proto with a repeated raw image_data bytes field, threads an optional MultimodalInput through request builders, adds a into_trtllm_proto() converter, and moves image hashing/encoding responsibilities to the server-side.

Changes

Cohort / File(s) Summary
Proto Schema
grpc_client/proto/trtllm_service.proto
Removed multimodal_hashes, multimodal_positions, multimodal_lengths from MultimodalInput; added image_data (repeated bytes). Breaking change to proto binary surface.
TRT-LLM Client
grpc_client/src/trtllm_service.rs
build_generate_request_from_chat(...) gains multimodal_input: Option<proto::MultimodalInput> and populates the request with it instead of None.
Gateway: Proto Converters
model_gateway/src/routers/grpc/proto_wrapper.rs
Added pub fn into_trtllm_proto(self) -> trtllm::MultimodalInput converting MultimodalData into proto image_data.
Gateway: GRPC Client Wiring
model_gateway/src/routers/grpc/client.rs
For Trtllm variant, map multimodal inputs with into_trtllm_proto() and pass resulting Option into build_generate_request_from_chat.
Gateway: Harmony Request Building
model_gateway/src/routers/grpc/harmony/stages/request_building.rs
Vllm and Trtllm Harmony paths now pass None for the multimodal argument, keeping Harmony pipeline without multimodal payloads.
Gateway: Multimodal Processing
model_gateway/src/routers/grpc/multimodal.rs
Renamed build_vllm_multimodal_data→build_raw_multimodal_data; process_for_backend signature extended and non-SGLang path now builds raw image bytes for backends.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Suggested labels

multimodal

Suggested reviewers

  • key4ng
  • slin1237

Poem

🐰 I pack pixels into tidy bytes,

off they hop to server lights,
No more scattered hashes, please—
The gateway sends raw image cheese,
Hooray for cleaner multimodal flights! 🥕✨

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly summarizes the main change: adding TensorRT-LLM multimodal support to the gRPC pipeline, which aligns with the primary objective and all code changes in the changeset.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
  • 📝 Generate docstrings (stacked PR)
  • 📝 Generate docstrings (commit on current branch)
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch chang/mm-6

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello @CatherineSue, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces multimodal (image) support for the TensorRT-LLM backend within the gRPC pipeline, addressing the previous rejection of such requests. The solution streamlines the MultimodalInput protobuf to carry only raw image bytes, consistent with the approach used for vLLM. This raw data is then wired through the gRPC client, allowing TensorRT-LLM's server-side input processor to handle the complex tasks of hashing, position tracking, and vision encoding, thereby extending VLM inference capabilities.

Highlights

  • Protobuf Definition Simplification: The MultimodalInput protobuf definition for TensorRT-LLM was simplified to exclusively use raw image bytes, removing previous fields for hashes, positions, and lengths.
  • New Multimodal Data Conversion: A new into_trtllm_proto() conversion method was introduced on MultimodalData to adapt it for the TensorRT-LLM protobuf structure.
  • TensorRT-LLM gRPC Client Integration: The TensorRT-LLM gRPC client (trtllm_service.rs) was updated to accept and correctly wire multimodal input into generate requests.
  • Client Dispatch Logic Update: The gRPC client dispatch logic (client.rs) was modified to convert and pass multimodal data specifically for the TensorRT-LLM backend.
  • Harmony Pipeline Compatibility: The Harmony pipeline (request_building.rs) was updated to explicitly pass None for multimodal input, ensuring compatibility as it does not support this feature.
  • Documentation Update: Relevant documentation (multimodal.rs) was revised to reflect that TensorRT-LLM now processes raw image bytes for multimodal input, aligning its behavior with vLLM.
Changelog
  • grpc_client/proto/trtllm_service.proto
    • Removed multimodal_hashes, multimodal_positions, and multimodal_lengths fields from MultimodalInput.
    • Added a repeated bytes image_data field to MultimodalInput.
  • grpc_client/src/trtllm_service.rs
    • Modified build_generate_request_from_chat to accept an Option<proto::MultimodalInput> parameter.
    • Passed the new multimodal_input parameter directly into the GenerateRequest construction.
  • model_gateway/src/routers/grpc/client.rs
    • Added a call to multimodal_inputs.map(|mm| mm.into_trtllm_proto()) to convert multimodal data for TRT-LLM.
    • Passed the converted trtllm_mm to client.build_generate_request_from_chat.
  • model_gateway/src/routers/grpc/harmony/stages/request_building.rs
    • Inserted None as the multimodal_input argument when calling build_generate_request_from_chat for the Harmony pipeline.
  • model_gateway/src/routers/grpc/multimodal.rs
    • Updated the doc comment for process_for_backend to indicate that both vLLM and TRT-LLM handle raw image bytes and internal preprocessing.
  • model_gateway/src/routers/grpc/proto_wrapper.rs
    • Implemented into_trtllm_proto for MultimodalData to convert it into the TensorRT-LLM specific protobuf format.
Activity
  • The code has been formatted using cargo +nightly fmt.
  • All clippy warnings have been addressed and cargo clippy --all-targets --all-features -D warnings passes.
  • Documentation updates are marked as optional and not yet completed.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

The pull request successfully implements multimodal support for the TensorRT-LLM backend by wiring it through the gRPC pipeline. The approach correctly mirrors the vLLM implementation by sending raw image bytes, which is appropriate as TRT-LLM handles preprocessing server-side. The changes are consistent across the proto definitions, client wrappers, and the gateway pipeline. I have provided a few suggestions to improve maintainability, reduce code duplication, and adhere to Protobuf best practices regarding field type changes.

Comment thread grpc_client/proto/trtllm_service.proto
Comment thread model_gateway/src/routers/grpc/multimodal.rs
@CatherineSue CatherineSue changed the title feat(multimodal): add TensorRT-LLM multimodal support in gRPC pipeline feat(gprc): add TensorRT-LLM multimodal support Feb 22, 2026
Wire multimodal image data through the gRPC pipeline for TRT-LLM:
- Simplify proto MultimodalInput to carry raw image bytes only
- Add into_trtllm_proto() conversion on MultimodalData
- Accept and pass multimodal_input in TRT-LLM client build methods
- Route multimodal data through to TRT-LLM in client dispatch

TRT-LLM handles hashing, token expansion, and vision encoding
server-side via its input processor, so the router only sends
unexpanded token IDs and raw JPEG/PNG bytes (same as vLLM path).

Signed-off-by: Chang Su <chang.s.su@oracle.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

grpc gRPC client and router changes model-gateway Model gateway crate changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants