Skip to content

refactor(grpc): split MultimodalData into backend-specific variants - #588

Merged
CatherineSue merged 5 commits into
mainfrom
chang/grpc-mm-backend-specific
Mar 3, 2026
Merged

CatherineSue merged 5 commits into
mainfrom
chang/grpc-mm-backend-specific

Conversation

@CatherineSue

@CatherineSue CatherineSue commented Mar 3, 2026 •

Copy link
Copy Markdown
Member

Description

Problem

The unified MultimodalData struct carries fields for all three backends (vLLM, SGLang, TRT-LLM), but each backend only uses a subset:

  • TRT-LLM: only needs raw image bytes — wastes compute serializing megabytes of pixel_values it never sends
  • SGLang: needs pixel_values, model_specific_tensors, patch_offsets, im_token_id — never uses mm_hashes, batched_keys, flat_keys
  • vLLM: needs pixel_values, model_specific_tensors, mm_hashes, batched_keys, flat_keys — never uses sglang_patch_offsets or image_data

As backends diverge (e.g. vLLM's keep_on_cpu field configs, TRT-LLM's upcoming preprocessed tensor support), the god struct keeps growing with fields that most backends ignore.

Solution

Replace the unified MultimodalData struct with an enum of three backend-specific structs. Introduce MultimodalIntermediate as a lightweight preparation output that holds preprocessing results without serializing tensors. The assembly into backend-specific data is deferred to request_building, where the target backend is known after worker selection.

Preparation → MultimodalIntermediate (lightweight, backend-agnostic)
                        ↓ (worker selection)
Request Building → assemble_multimodal_data(intermediate, client)
                        ↓
              MultimodalData::Sglang(...)
              MultimodalData::Vllm(...)
              MultimodalData::Trtllm(...)

Changes

  • proto_wrapper.rs: Replace unified MultimodalData struct with MultimodalData enum + SglangMultimodalData, VllmMultimodalData, TrtllmMultimodalData structs. Move into_*_proto() to per-struct into_proto().
  • multimodal.rs: Add MultimodalIntermediate struct. Extract serialize_pixel_values() and serialize_model_specific() helpers. Add assemble_sglang(), assemble_vllm(), assemble_trtllm(), assemble_multimodal_data(). Update process_multimodal() to return intermediate. Remove build_multimodal_data().
  • mod.rs: ProcessedMessages.multimodal_data → multimodal_intermediate.
  • preparation.rs: Store intermediate instead of MultimodalData.
  • request_building.rs: Call assemble_multimodal_data(intermediate, builder_client) after worker selection.
  • client.rs: Update build_chat_request() to match on enum variant.

Test Plan

  • E2E tested with both vLLM and SGLang backends
  • No functional change — same data reaches each backend, just assembled at a later stage
Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated

Summary by CodeRabbit

  • Refactor

    • Multimodal handling reorganized into a two-stage flow: intermediate preprocessing followed by backend-specific assembly.
    • Direct multimodal payload exposure replaced by an intermediate container and stricter backend-variant alignment to prevent mismatches.
  • New Features

    • Added backend-specific multimodal variants with tailored assembly so each backend receives properly formatted multimodal payloads.
    • Public API expanded to expose a field-layout type for multimodal layouts.

Replace the unified MultimodalData struct with an enum of three
backend-specific structs (SglangMultimodalData, VllmMultimodalData,
TrtllmMultimodalData). Each variant carries only the fields its
backend needs:

- SGLang: pixel_values + model_specific + patch-only placeholders
- vLLM: pixel_values + model_specific + structural placeholders +
  hashes + field keys
- TRT-LLM: raw image bytes only (no tensor serialization)

Introduce MultimodalIntermediate as a lightweight preparation output
that holds preprocessing results without serializing tensors. The
assembly into backend-specific data is deferred to request_building,
where the target backend is known after worker selection. This avoids
wasting work (e.g. TRT-LLM no longer serializes megabytes of
pixel_values it never sends).

Signed-off-by: Chang Su <chang.s.su@oracle.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Repo admins can enable using credits for code reviews in their settings.

@github-actions github-actions Bot added grpc gRPC client and router changes multimodal Multimodal crate changes model-gateway Model gateway crate changes labels Mar 3, 2026
@coderabbitai

coderabbitai Bot commented Mar 3, 2026 •

Copy link
Copy Markdown

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 8832e59 and b4f40e5.

📒 Files selected for processing (1)
  • model_gateway/src/routers/grpc/multimodal.rs

📝 Walkthrough

Walkthrough

Refactors multimodal handling to produce a MultimodalIntermediate in preparation, adds assemble_multimodal_data(...) to build backend-specific MultimodalData variants (Sglang/Vllm/Trtllm) at request-build time, converts MultimodalData into an enum with per-backend types, and updates call sites with variant-aware matching and unreachable guards.

Changes

Cohort / File(s) Summary
Proto wrapper & backend variants
model_gateway/src/routers/grpc/proto_wrapper.rs
Replace old single MultimodalData struct with a MultimodalData enum and backend-specific structs (SglangMultimodalData, VllmMultimodalData, TrtllmMultimodalData); add into_proto() on each variant and adjust per-backend fields (mm_placeholders, pixel_values, model tensors, hashes, keys, etc.).
Multimodal pipeline & assembly
model_gateway/src/routers/grpc/multimodal.rs
Introduce MultimodalIntermediate and make outputs carry it; add assemble_multimodal_data(intermediate, client) plus backend assemblers (assemble_sglang, assemble_vllm, assemble_trtllm) and serialization helpers (pixel/model-specific tensor serialization).
Client request building / variant handling
model_gateway/src/routers/grpc/client.rs
Update build_chat_request to pattern-match MultimodalData variants and call appropriate into_proto() per variant; add unreachable! guards and clippy expectations enforcing backend/variant invariants.
ProcessedMessages field & flow updates
model_gateway/src/routers/grpc/mod.rs, model_gateway/src/routers/grpc/utils.rs, model_gateway/src/routers/grpc/regular/stages/chat/preparation.rs
Rename public ProcessedMessages.multimodal_data → pub(crate) multimodal_intermediate: Option<multimodal::MultimodalIntermediate> and propagate the rename through preparation and utils.
Request-building integration
model_gateway/src/routers/grpc/regular/stages/chat/request_building.rs
Import and invoke assemble_multimodal_data() to convert multimodal_intermediate into a backend-specific MultimodalData using the active GrpcClient before constructing the proto request.
Public exports
multimodal/src/lib.rs
Add FieldLayout to public re-exports from the multimodal types module.

Sequence Diagram

sequenceDiagram
    participant Client as Client/Request
    participant Prep as PreparationStage
    participant Intermediate as MultimodalIntermediate
    participant Assembler as AssemblyLayer
    participant Variant as MultimodalDataVariant
    participant Proto as ProtoConversion

    Client->>Prep: submit message (may include images)
    Prep->>Intermediate: produce MultimodalIntermediate (preprocessed, images, placeholders, layouts)
    Prep-->>Client: return ProcessedMessages with intermediate

    Client->>Assembler: request_building invokes assemble_multimodal_data(intermediate, GrpcClient)
    Assembler->>Assembler: select backend assembler (Sglang/Vllm/Trtllm)
    Assembler->>Variant: assemble backend-specific MultimodalData
    Variant->>Proto: call into_proto()
    Proto-->>Client: return backend-specific proto payload
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~55 minutes

Possibly related issues

  • Issue #566: Implements the MultimodalIntermediate → assemble_multimodal_data flow and backend-specific MultimodalData variants described by the issue.

Possibly related PRs

  • PR #495: Modifies the gRPC multimodal pipeline and types; overlaps with assembly and proto-variant changes.
  • PR #504: Adds TRT-LLM proto conversion and client wiring; closely related to the new Trtllm variant handling.
  • PR #497: Changes MultimodalData flow and build_chat_request handling; directly connected to this client's variant-aware conversions.

Suggested reviewers

  • key4ng
  • slin1237

Poem

🐰 I tucked the pixels safe and neat,

Held placeholders beneath my feet,
I stitch per-backend, variant tight,
Hop—each request now flies just right,
Tiny paws assemble bytes by light.

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately captures the main refactoring change: splitting a unified MultimodalData struct into backend-specific variants (enum with Sglang, Vllm, and Trtllm variants). This directly reflects the core architectural change across proto_wrapper.rs and related modules.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
  • 📝 Generate docstrings (stacked PR)
  • 📝 Generate docstrings (commit on current branch)
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch chang/grpc-mm-backend-specific

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly refactors the handling of multimodal data within the system's gRPC layer. By replacing a monolithic MultimodalData struct with an enum of backend-specific variants and introducing an intermediate, lightweight data structure, the system now defers the final assembly of multimodal data until the specific backend is identified. This architectural change optimizes data serialization, ensuring that each backend receives only the data it requires, thereby enhancing efficiency and maintainability as backend requirements evolve.

Highlights

  • Refactored MultimodalData: The monolithic MultimodalData struct has been refactored into an enum with distinct backend-specific variants (SGLang, vLLM, TRT-LLM), each containing only the data relevant to its target backend.
  • Introduced MultimodalIntermediate: A new MultimodalIntermediate struct was introduced to store lightweight, backend-agnostic preprocessing results, deferring heavy serialization until the target backend is known.
  • Deferred Data Assembly: The assembly of backend-specific multimodal data has been moved to the request_building stage, occurring after worker selection, which ensures optimal data handling.
  • Improved Efficiency: This refactoring improves efficiency by avoiding the serialization and transmission of unnecessary data to each backend, reducing wasted compute and enhancing maintainability.
Changelog
  • model_gateway/src/routers/grpc/client.rs
    • Updated build_chat_request to use a match statement on the new MultimodalData enum, ensuring correct data conversion for each backend.
    • Added #[expect(clippy::unreachable)] annotations to handle expected unreachable match arms.
  • model_gateway/src/routers/grpc/mod.rs
    • Replaced multimodal_data with multimodal_intermediate in ProcessedMessages and updated its documentation.
  • model_gateway/src/routers/grpc/multimodal.rs
    • Defined the new MultimodalIntermediate struct to hold preprocessed data.
    • Modified MultimodalOutput to return MultimodalIntermediate.
    • Removed the build_multimodal_data function.
    • Added assemble_multimodal_data and backend-specific assembly functions (assemble_sglang, assemble_vllm, assemble_trtllm) to convert MultimodalIntermediate into the appropriate MultimodalData enum variant.
    • Extracted serialize_pixel_values and serialize_model_specific into dedicated helper functions.
  • model_gateway/src/routers/grpc/proto_wrapper.rs
    • Transformed MultimodalData from a struct into an enum with SglangMultimodalData, VllmMultimodalData, and TrtllmMultimodalData variants.
    • Defined the new backend-specific structs, each containing only the relevant fields.
    • Moved the into_*_proto conversion methods from the old MultimodalData struct to the new backend-specific structs and renamed them to into_proto.
  • model_gateway/src/routers/grpc/regular/stages/chat/preparation.rs
    • Updated the process_multimodal call to store the returned MultimodalIntermediate instead of MultimodalData.
    • Adjusted the assignment to processed_messages.multimodal_intermediate.
  • model_gateway/src/routers/grpc/regular/stages/chat/request_building.rs
    • Imported the assemble_multimodal_data function.
    • Modified the request building logic to call assemble_multimodal_data with the multimodal_intermediate and the selected builder_client to obtain the correct MultimodalData variant.
  • model_gateway/src/routers/grpc/utils.rs
    • Adjusted process_chat_messages to initialize and return multimodal_intermediate as None in relevant paths.
  • multimodal/src/lib.rs
    • Exported the FieldLayout type from the types module.
Activity
  • No specific human activity (comments, reviews, or progress updates) has been recorded for this pull request.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request is a well-executed refactoring that improves the handling of multimodal data. By splitting the monolithic MultimodalData struct into backend-specific variants within an enum, the code is now cleaner and more efficient, as each backend only deals with the data it requires. The introduction of MultimodalIntermediate to defer data assembly until after worker selection is a smart design choice that enhances modularity.

I have a couple of suggestions for further improvement:

  • An optimization to reduce memory allocations and copies when collecting raw image data, aligning with the principle of avoiding unnecessary intermediate allocations.
  • A small refactoring to make a data serialization function more idiomatic and concise.

Comment thread model_gateway/src/routers/grpc/multimodal.rs
Comment thread model_gateway/src/routers/grpc/multimodal.rs Outdated
…pecific

Replace mutable HashMap loop with idiomatic filter_map().collect() pattern.

Signed-off-by: Chang Su <chang.s.su@oracle.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
model_gateway/src/routers/grpc/multimodal.rs (1)

453-551: 🧹 Nitpick | 🔵 Trivial

Add unit tests for backend assembly invariants.

This PR adds critical backend-specific assembly paths (assemble_sglang, assemble_vllm, assemble_trtllm) but current tests in this file do not cover them. Please add focused tests for placeholder selection, hash/image mapping, and serialization shapes/dtypes to prevent silent payload regressions.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@model_gateway/src/routers/grpc/multimodal.rs` around lines 453 - 551, Add
unit tests that exercise assemble_sglang, assemble_vllm, and assemble_trtllm to
lock-in assembly invariants: (1) For assemble_sglang, test that when
MultimodalIntermediate.patch_offsets is Some the returned
SglangMultimodalData.mm_placeholders uses those patch-only offsets and when None
it falls back to MultimodalIntermediate.placeholders (check offset/length types
and ordering); (2) For assemble_vllm, assert mm_hashes matches
intermediate.images[*].hash and image-derived fields (pixel_values,
pixel_values_shape, model_specific_tensors) are populated, and validate
batched_keys and flat_keys come from PreprocessedImages::batched_keys and
::flat_keys respectively; (3) For assemble_trtllm, verify image_data is a Vec of
the images' raw_bytes; and (4) directly test serialize_pixel_values and
serialize_model_specific produce correct byte-layout (little-endian f32 bytes
and shape Vec<u32>) and that model_specific_to_tensor_bytes entries become keys
in the HashMap of TensorBytes. Use small deterministic PreprocessedImages and
MultimodalIntermediate fixtures to assert types, lengths, and exact bytes where
practical.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@model_gateway/src/routers/grpc/multimodal.rs`:
- Around line 461-468: The fallback currently only triggers when
intermediate.patch_offsets is None, but when patch_offsets is Some(vec![]) we
must treat it as unavailable; update the mm_placeholders selection to use
patch_offsets only if it is Some and non-empty (e.g., check .as_ref().filter(|v|
!v.is_empty()) or match and treat empty vec as None) and otherwise build
mm_placeholders from intermediate.placeholders (referencing mm_placeholders,
intermediate.patch_offsets, and intermediate.placeholders) so empty
patch_offsets will not produce an empty placeholder list.

---

Outside diff comments:
In `@model_gateway/src/routers/grpc/multimodal.rs`:
- Around line 453-551: Add unit tests that exercise assemble_sglang,
assemble_vllm, and assemble_trtllm to lock-in assembly invariants: (1) For
assemble_sglang, test that when MultimodalIntermediate.patch_offsets is Some the
returned SglangMultimodalData.mm_placeholders uses those patch-only offsets and
when None it falls back to MultimodalIntermediate.placeholders (check
offset/length types and ordering); (2) For assemble_vllm, assert mm_hashes
matches intermediate.images[*].hash and image-derived fields (pixel_values,
pixel_values_shape, model_specific_tensors) are populated, and validate
batched_keys and flat_keys come from PreprocessedImages::batched_keys and
::flat_keys respectively; (3) For assemble_trtllm, verify image_data is a Vec of
the images' raw_bytes; and (4) directly test serialize_pixel_values and
serialize_model_specific produce correct byte-layout (little-endian f32 bytes
and shape Vec<u32>) and that model_specific_to_tensor_bytes entries become keys
in the HashMap of TensorBytes. Use small deterministic PreprocessedImages and
MultimodalIntermediate fixtures to assert types, lengths, and exact bytes where
practical.

ℹ️ Review info

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between d6420f5 and 1978e84.

📒 Files selected for processing (1)
  • model_gateway/src/routers/grpc/multimodal.rs

Comment thread model_gateway/src/routers/grpc/multimodal.rs Outdated
…ssembly

Fall back to full structural placeholder ranges when patch_offsets is
Some(vec![]) rather than only on None, preventing empty mm_placeholders
from being sent to SGLang.

Signed-off-by: Chang Su <chang.s.su@oracle.com>
Signed-off-by: Chang Su <chang.s.su@oracle.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@model_gateway/src/routers/grpc/multimodal.rs`:
- Around line 453-514: Add unit tests that exercise assemble_sglang,
assemble_vllm, and assemble_trtllm using minimal MultimodalIntermediate
fixtures: verify assemble_sglang returns SglangMultimodalData with correct
image_data, pixel_values, pixel_values_shape, model_specific_tensors,
im_token_id, and that mm_placeholders uses patch_offsets when present and
non-empty and falls back to placeholders mapping (offset,length) when
patch_offsets is None or empty; verify assemble_vllm returns VllmMultimodalData
with expected pixel_values, pixel_values_shape, model_specific_tensors,
im_token_id, mm_placeholders (from placeholders), mm_hashes (from images.hash),
and batched_keys/flat_keys (use PreprocessedImages::batched_keys/flat_keys
deterministically); verify assemble_trtllm returns TrtllmMultimodalData with
image_data from images.raw_bytes; create small helper builders for
MultimodalIntermediate/PreprocessedImages to set images,
preprocessed.model_specific, im_token_id, placeholders and patch_offsets to
assert both code paths.
- Around line 545-553: The serializer currently drops unsupported
ModelSpecificValue entries silently in serialize_model_specific by using
filter_map with model_specific_to_tensor_bytes; change this so dropped keys are
surfaced: either (preferred) make serialize_model_specific return
Result<HashMap<String,TensorBytes>, SerializeError> and return an Err listing
the unsupported keys (using model_specific_to_tensor_bytes failures), or (if
non-breaking) collect the dropped keys and emit a warning (e.g., tracing::warn!
or your crate logger) that includes the key names and their value types before
returning the partial map; apply the same fix to the analogous map-to-tensor
conversion at the other site referenced (lines ~590-592) so unsupported entries
are consistently reported.

ℹ️ Review info

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 174dd8c and 8832e59.

📒 Files selected for processing (1)
  • model_gateway/src/routers/grpc/multimodal.rs

Comment thread model_gateway/src/routers/grpc/multimodal.rs
Comment thread model_gateway/src/routers/grpc/multimodal.rs
Signed-off-by: Chang Su <chang.s.su@oracle.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

grpc gRPC client and router changes model-gateway Model gateway crate changes multimodal Multimodal crate changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant