feat(gprc): add TensorRT-LLM multimodal support - #504
Conversation
|
No actionable comments were generated in the recent review. 🎉 📝 WalkthroughWalkthroughReplaces per-image metadata fields in the TRT-LLM proto with a repeated raw Changes
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~20 minutes Possibly related PRs
Suggested labels
Suggested reviewers
Poem
🚥 Pre-merge checks | ✅ 3✅ Passed checks (3 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches
🧪 Generate unit tests (beta)
Comment |
Summary of ChangesHello @CatherineSue, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request introduces multimodal (image) support for the TensorRT-LLM backend within the gRPC pipeline, addressing the previous rejection of such requests. The solution streamlines the Highlights
Changelog
Activity
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here. You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension. Footnotes
|
There was a problem hiding this comment.
Code Review
The pull request successfully implements multimodal support for the TensorRT-LLM backend by wiring it through the gRPC pipeline. The approach correctly mirrors the vLLM implementation by sending raw image bytes, which is appropriate as TRT-LLM handles preprocessing server-side. The changes are consistent across the proto definitions, client wrappers, and the gateway pipeline. I have provided a few suggestions to improve maintainability, reduce code duplication, and adhere to Protobuf best practices regarding field type changes.
Wire multimodal image data through the gRPC pipeline for TRT-LLM: - Simplify proto MultimodalInput to carry raw image bytes only - Add into_trtllm_proto() conversion on MultimodalData - Accept and pass multimodal_input in TRT-LLM client build methods - Route multimodal data through to TRT-LLM in client dispatch TRT-LLM handles hashing, token expansion, and vision encoding server-side via its input processor, so the router only sends unexpanded token IDs and raw JPEG/PNG bytes (same as vLLM path). Signed-off-by: Chang Su <chang.s.su@oracle.com>
2a52e7e to
eeec140
Compare
Description
Problem
The gRPC pipeline doesn't support multimodal (image) requests for the TensorRT-LLM backend.
TRT-LLM's LLM API already supports VLM inference when given
prompt_token_ids+multi_modal_data.Solution
Wire TensorRT-LLM multimodal support through the gRPC pipeline using the same approach as vLLM: send unexpanded token IDs + raw image bytes. TRT-LLM's input processor handles hashing, position tracking, and vision encoding server-side.
Changes
grpc_client/proto/trtllm_service.proto): SimplifyMultimodalInputto carry onlyrepeated bytes image_data— remove unusedmultimodal_hashes,multimodal_positions,multimodal_lengthsfieldsproto_wrapper.rs): Addinto_trtllm_proto()conversion method onMultimodalDatatrtllm_service.rs): Acceptmultimodal_input: Option<proto::MultimodalInput>inbuild_generate_request_from_chat()and wire it into the proto requestclient.rs): Convert and pass multimodal data in the TRT-LLM branchharmony/stages/request_building.rs): PassNonefor the new multimodal param (Harmony doesn't support multimodal)multimodal.rs): Updateprocess_for_backend()doc to reflect TRT-LLM now uses the vLLM (raw bytes) branchTest Plan
Checklist
cargo +nightly fmtpassescargo clippy --all-targets --all-features -- -D warningspassesSummary by CodeRabbit