feat(grpc): add vLLM multimodal support and split pipeline into fetch + preprocess - #497
Conversation
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
📝 WalkthroughWalkthroughIntroduces backend-agnostic multimodal types and a two‑phase multimodal pipeline in the model gateway, adds mm_inputs to vLLM proto and threads multimodal through vllm client, and disables multimodal propagation in the Go bindings by passing None. Changes
Sequence Diagram(s)sequenceDiagram
participant Client
participant ModelGateway
participant MultimodalModule
participant BackendBuilder
participant gRPC
Client->>ModelGateway: Send chat request (may include images)
ModelGateway->>MultimodalModule: fetch_images(messages) -- Phase 1
MultimodalModule-->>ModelGateway: Vec<Arc<ImageFrame>>
alt SGLang backend
ModelGateway->>MultimodalModule: preprocess_for_sglang(images, token_ids)
MultimodalModule-->>ModelGateway: (expanded_token_ids, MultimodalData)
else non-SGLang backend
ModelGateway->>MultimodalModule: process_for_backend(images, is_sglang=false)
MultimodalModule-->>ModelGateway: (token_ids, MultimodalData)
end
ModelGateway->>BackendBuilder: build_chat_request(token_ids, MultimodalData)
BackendBuilder->>gRPC: build proto request (mm_inputs set via conversion)
gRPC-->>BackendBuilder: proto GenerateRequest
BackendBuilder-->>Client: stream/response
Estimated code review effort🎯 4 (Complex) | ⏱️ ~60 minutes Possibly related PRs
Suggested labels
Suggested reviewers
Poem
🚥 Pre-merge checks | ✅ 3✅ Passed checks (3 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
Summary of ChangesHello @CatherineSue, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request significantly enhances the gRPC multimodal pipeline by introducing support for vLLM backends and decoupling the image processing workflow. Previously, the system was tightly coupled to SGLang's preprocessing requirements, leading to inefficiencies when interacting with vLLM. The changes streamline the process by fetching raw images first, then applying backend-specific preprocessing and token expansion only when the target backend is identified, thereby optimizing resource usage and improving compatibility. Highlights
Changelog
Activity
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here. You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension. Footnotes
|
There was a problem hiding this comment.
Code Review
This pull request introduces vLLM multimodal support and refactors the gRPC multimodal pipeline. A high-severity Server-Side Request Forgery (SSRF) vulnerability was identified due to a lack of URL validation in the image fetching phase, which should be addressed in a dedicated security PR. A critical data serialization bug was found where u32 values are incorrectly serialized as 64-bit integers, leading to potential data corruption; explicit endianness conversion is recommended. Additionally, there's an opportunity to reduce code duplication for improved maintainability.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 6ffe53b0b0
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
There was a problem hiding this comment.
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (4)
bindings/golang/src/client.rs (1)
199-207:⚠️ Potential issue | 🟠 MajorDon’t silently drop multimodal inputs.
Passing
Nonehere will ignore images if the request includes them. Please return an explicit “multimodal not supported in Go bindings” error when multimodal content is present (e.g., viaprocessed_messages.multimodal_imagesor inspectingchat_request.messages) to avoid silent data loss.🔧 Proposed fix
// Build GenerateRequest let request_id = format!("chatcmpl-{}", Uuid::new_v4()); + if processed_messages.multimodal_images.is_some() { + set_error_message(error_out, "Multimodal inputs are not supported in Go bindings"); + return SglErrorCode::InvalidArgument; + } let proto_request = match client.build_generate_request_from_chat( request_id.clone(), &chat_request, processed_messages.text, token_ids, None, // multimodal not supported in golang bindings tool_constraint, ) {🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@bindings/golang/src/client.rs` around lines 199 - 207, Before calling client.build_generate_request_from_chat, detect whether the incoming request contains multimodal content (e.g., check processed_messages.multimodal_images and inspect chat_request.messages for image/attachment entries) and if any multimodal content exists return an explicit error like "multimodal not supported in Go bindings" instead of proceeding with the call that passes None; modify the code around the request construction (the call to build_generate_request_from_chat and the variables request_id, processed_messages) to perform this check and early-return the error so images are not silently dropped.bindings/golang/src/policy.rs (1)
547-555:⚠️ Potential issue | 🟠 MajorReturn an error when multimodal inputs are provided.
This path also drops multimodal inputs silently. Please add an explicit “multimodal not supported” error when the request includes images to prevent confusing output.
🔧 Proposed fix
// Build GenerateRequest let request_id = format!("chatcmpl-{}", Uuid::new_v4()); + if processed_messages.multimodal_images.is_some() { + set_error_message(error_out, "Multimodal inputs are not supported in Go bindings"); + return SglErrorCode::InvalidArgument; + } let proto_request = match client.build_generate_request_from_chat( request_id.clone(), &chat_request, processed_messages.text, token_ids, None, // multimodal not supported in golang bindings tool_constraint, ) {🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@bindings/golang/src/policy.rs` around lines 547 - 555, The code currently calls client.build_generate_request_from_chat and passes None for multimodal, silently dropping images; change it to detect multimodal inputs in processed_messages (e.g., check processed_messages.images or any image/multimodal field) before building the request and return an explicit error like "multimodal not supported" instead of proceeding. Update the code path around build_generate_request_from_chat (where request_id and proto_request are created) to return Err(...) / propagate an error when images are present so callers receive a clear failure rather than silently losing multimodal data.model_gateway/src/routers/grpc/multimodal.rs (1)
527-568:⚠️ Potential issue | 🟠 MajorFix uint32 tensor serialization width mismatch.
UintTensorandUintVecstoreVec<u32>(confirmed via type definitions), but are serialized using(*v as i64).to_le_bytes()(8 bytes) while declaringdtype: "uint32"(4 bytes). This mismatch corrupts tensor payloads for backends expecting 4‑byte elements. Serialize using the native width by callingv.to_le_bytes()directly.🛠️ Suggested fix
ModelSpecificValue::UintTensor { data, shape } => Some(super::TensorBytes { data: data .iter() - .flat_map(|v| (*v as i64).to_le_bytes()) + .flat_map(|v| v.to_le_bytes()) .collect(), shape: shape.iter().map(|&d| d as u32).collect(), dtype: "uint32".to_string(), }), ModelSpecificValue::UintVec(v) => Some(super::TensorBytes { data: v .iter() - .flat_map(|val| (*val as i64).to_le_bytes()) + .flat_map(|val| val.to_le_bytes()) .collect(), shape: vec![v.len() as u32], dtype: "uint32".to_string(), }),🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@model_gateway/src/routers/grpc/multimodal.rs` around lines 527 - 568, The serialization for uint32 tensors in model_specific_to_tensor_bytes is using (*v as i64).to_le_bytes() which writes 8-byte values while dtype is "uint32"; update the ModelSpecificValue::UintTensor and ModelSpecificValue::UintVec branches to serialize each u32 with v.to_le_bytes() (native 4-byte little-endian) and keep dtype "uint32" and shape logic unchanged so payload width matches the declared dtype.model_gateway/src/routers/grpc/regular/stages/chat/preparation.rs (1)
120-128: 🧹 Nitpick | 🔵 TrivialAdd a short TODO documenting the temporary 400‑mapping for multimodal fetch errors.
All fetch failures currently map to
bad_request; a brief note keeps the planned 4xx/5xx split visible until error taxonomy improves.📝 Suggested inline note
- return Err(error::bad_request( + // TODO: Multimodal fetch errors currently map to 400; + // refine to 4xx/5xx once multimodal error taxonomy is available. + return Err(error::bad_request(Based on learnings: In model_gateway/src/routers/grpc/regular/stages/chat/preparation.rs, multimodal processing failures now return 400 Bad Request due to the generic anyhow::Error lacking distinguished error types; implement explicit error categorization and, until then, document the current behavior and anticipated extension.
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@model_gateway/src/routers/grpc/regular/stages/chat/preparation.rs` around lines 120 - 128, Add a short TODO comment right above the error! log and Err(error::bad_request(...)) return in ChatPreparationStage::execute explaining that multimodal image fetch failures are currently mapped to 400 Bad Request because the underlying anyhow::Error lacks discriminated error types, and note the intended future change to split client (4xx) vs server (5xx) errors once explicit error taxonomy is implemented; reference the multimodal processing branch (the error! call and the error::bad_request(...) return) so reviewers can find and later replace the temporary mapping.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Outside diff comments:
In `@bindings/golang/src/client.rs`:
- Around line 199-207: Before calling client.build_generate_request_from_chat,
detect whether the incoming request contains multimodal content (e.g., check
processed_messages.multimodal_images and inspect chat_request.messages for
image/attachment entries) and if any multimodal content exists return an
explicit error like "multimodal not supported in Go bindings" instead of
proceeding with the call that passes None; modify the code around the request
construction (the call to build_generate_request_from_chat and the variables
request_id, processed_messages) to perform this check and early-return the error
so images are not silently dropped.
In `@bindings/golang/src/policy.rs`:
- Around line 547-555: The code currently calls
client.build_generate_request_from_chat and passes None for multimodal, silently
dropping images; change it to detect multimodal inputs in processed_messages
(e.g., check processed_messages.images or any image/multimodal field) before
building the request and return an explicit error like "multimodal not
supported" instead of proceeding. Update the code path around
build_generate_request_from_chat (where request_id and proto_request are
created) to return Err(...) / propagate an error when images are present so
callers receive a clear failure rather than silently losing multimodal data.
In `@model_gateway/src/routers/grpc/multimodal.rs`:
- Around line 527-568: The serialization for uint32 tensors in
model_specific_to_tensor_bytes is using (*v as i64).to_le_bytes() which writes
8-byte values while dtype is "uint32"; update the ModelSpecificValue::UintTensor
and ModelSpecificValue::UintVec branches to serialize each u32 with
v.to_le_bytes() (native 4-byte little-endian) and keep dtype "uint32" and shape
logic unchanged so payload width matches the declared dtype.
In `@model_gateway/src/routers/grpc/regular/stages/chat/preparation.rs`:
- Around line 120-128: Add a short TODO comment right above the error! log and
Err(error::bad_request(...)) return in ChatPreparationStage::execute explaining
that multimodal image fetch failures are currently mapped to 400 Bad Request
because the underlying anyhow::Error lacks discriminated error types, and note
the intended future change to split client (4xx) vs server (5xx) errors once
explicit error taxonomy is implemented; reference the multimodal processing
branch (the error! call and the error::bad_request(...) return) so reviewers can
find and later replace the temporary mapping.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 9b51a186ac
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: bcb5cca89e
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
|
Hi @CatherineSue, this PR has merge conflicts that must be resolved before it can be merged. Please rebase your branch: git fetch origin main
git rebase origin/main
# resolve any conflicts, then:
git push --force-with-lease |
bcb5cca to
2251d12
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 2251d120d7
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@model_gateway/src/routers/grpc/regular/stages/chat/request_building.rs`:
- Around line 108-113: Add a short inline comment documenting the fallback
behavior where tokenizer_source is set: explain that
ctx.components.tokenizer_registry.get_by_name(model_id) may return None and in
that case tokenizer_source is intentionally set to an empty string via
unwrap_or_default(), so downstream code should expect an empty source when no
tokenizer is registered for model_id; reference tokenizer_source,
tokenizer_registry.get_by_name, model_id and unwrap_or_default in the comment.
There was a problem hiding this comment.
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
model_gateway/src/routers/grpc/multimodal.rs (1)
426-427:⚠️ Potential issue | 🟡 MinorMinor:
as u32cast could mask invalid negative token IDs.If
PromptReplacement.tokensever contained a negative value (e.g., due to upstream bug),*&t as u32would silently wrap to a large u32. Consider usingtry_into()with an error, or document that negative token IDs are invalid.🛡️ Optional defensive fix
- expanded.extend(repl.tokens.iter().map(|&t| t as u32)); + expanded.extend(repl.tokens.iter().map(|&t| { + debug_assert!(t >= 0, "Negative token ID in prompt replacement"); + t as u32 + }));🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@model_gateway/src/routers/grpc/multimodal.rs` around lines 426 - 427, The current expansion call expanded.extend(repl.tokens.iter().map(|&t| t as u32)) silently wraps negative PromptReplacement.tokens into large u32 values; change the conversion to check for negativity using TryInto (or an explicit check) on each token from repl.tokens in the function that builds expanded (e.g., where PromptReplacement.tokens is iterated) and return or propagate an error if a token is negative (or otherwise invalid) instead of using as u32 so invalid token IDs are caught and handled.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Outside diff comments:
In `@model_gateway/src/routers/grpc/multimodal.rs`:
- Around line 426-427: The current expansion call
expanded.extend(repl.tokens.iter().map(|&t| t as u32)) silently wraps negative
PromptReplacement.tokens into large u32 values; change the conversion to check
for negativity using TryInto (or an explicit check) on each token from
repl.tokens in the function that builds expanded (e.g., where
PromptReplacement.tokens is iterated) and return or propagate an error if a
token is negative (or otherwise invalid) instead of using as u32 so invalid
token IDs are caught and handled.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: f081e3b3db
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
This comment was marked as resolved.
This comment was marked as resolved.
…dalData Introduce MultimodalData wrapper in proto_wrapper.rs to decouple the multimodal pipeline from backend-specific proto types. Each backend converts via into_sglang_proto() / into_vllm_proto() (consuming self to avoid expensive clones). For vLLM, raw image bytes are sent via the new proto MultimodalInputs message. vLLM handles preprocessing internally while preserving our pre-expanded token IDs. Signed-off-by: Chang Su <chang.s.su@oracle.com> # Conflicts: # model_gateway/src/routers/grpc/multimodal.rs
vLLM handles its own multimodal placeholder expansion internally. The router was sending already-expanded token IDs (40→183 tokens), causing vLLM to fail with "Failed to apply prompt replacement" because placeholder tokens were already replaced. Store original (unexpanded) token IDs in PreparationOutput and use them when building vLLM requests, while SGLang continues to receive expanded tokens. Signed-off-by: Chang Su <chang.s.su@oracle.com>
…sing Split the monolithic process_multimodal() into two phases: - Phase 1 (preparation stage): fetch images via tracker, enforce modality limits, store raw ImageFrames on ProcessedMessages - Phase 2 (request building stage): backend-specific processing where SGLang gets full pixel preprocessing + token expansion, while vLLM gets raw image bytes only (handles preprocessing internally) This removes the original_token_ids workaround and avoids wasting CPU on SGLang-specific preprocessing when the backend is vLLM. Signed-off-by: Chang Su <chang.s.su@oracle.com>
…ed main Adapt the cherry-picked multimodal gRPC code to the stricter clippy configuration on main: inline format args, replace #[allow] with #[cfg(test)], use exhaustive match arms, restore async fn for prepare_chat. Signed-off-by: Chang Su <chang.s.su@oracle.com>
Guard against as_slice()/as_slice_memory_order() returning None for non-contiguous ndarray tensors. Falls back to element-wise iteration instead of silently producing empty bytes via unwrap_or_default(). Signed-off-by: Chang Su <chang.s.su@oracle.com>
Signed-off-by: Chang Su <chang.s.su@oracle.com>
- Fix UintTensor/UintVec serialization: was casting u32 to i64 (8 bytes) while declaring dtype "uint32" (4 bytes), causing byte-width mismatch. Now serializes as native u32 little-endian bytes. - Import MultimodalData and TensorBytes at module top to eliminate repeated super:: references. Signed-off-by: Chang Su <chang.s.su@oracle.com>
…eview feedback - Extract duplicated config+spec lookup into resolve_model_spec() - Use crate:: imports instead of super:: for MultimodalData/TensorBytes - Add TODO for 4xx/5xx error distinction in multimodal fetch - Add TODO for token expansion vs routing policy interaction Signed-off-by: Chang Su <chang.s.su@oracle.com>
…g from model registry TrackerConfig required resolving the model spec (via ModelRegistry) at tracker construction time to obtain two pieces of information: 1. placeholder_tokens — the model-specific placeholder token string (e.g. "<image>", "<|image|>") used to build ConversationSegment tracking and PlaceholderMap/PlaceholderHandle bookkeeping within the tracker. 2. modality_limits — per-modality item count limits from the model spec, enforced via the ModalityLimit error variant. Both are dead in the current architecture: - ConversationSegment tracking (text-position-aware placeholder insertion) was designed for reconstructing prompt text with placeholders, but token expansion now operates directly on token_ids in Phase 2 (preprocess_for_sglang). The tracker never needs to track text positions or know the placeholder token string. - Modality limits were never consumed downstream and only rejected requests early. If limit enforcement is needed, it belongs at the model spec level in Phase 2 where the model is already resolved. The practical problem: TrackerConfig forced ChatPreparationStage (Phase 1) to resolve the model spec via ModelRegistry just to fetch images. This created an unnecessary dependency — whether a model has token expansion configs (like llava, phi3v, llama4_vision) should not matter when all Phase 1 does is download image bytes. By removing TrackerConfig, Phase 1 only needs a MediaConnector, and the ModelRegistry dependency is confined to Phase 2 in grpc/multimodal.rs where it actually belongs. Also removes the now-dead types: ConversationSegment, PlaceholderHandle, PlaceholderMap, DEFAULT_PLACEHOLDERS, and the ModalityLimit error variant. Signed-off-by: Chang Su <chang.s.su@oracle.com>
Signed-off-by: Chang Su <chang.s.su@oracle.com>
f081e3b to
f862344
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: f8623447d9
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
…eline Remove the early-return rejection for multimodal input on the TRT-LLM backend. TRT-LLM can accept pre-tokenized (unexpanded) token IDs plus raw image data — its LLM API will decode and re-process internally. This unblocks multimodal request flow for TRT-LLM in the chat pipeline. Signed-off-by: Chang Su <changsu@nvidia.com> Signed-off-by: Chang Su <chang.s.su@oracle.com>
The tokenizer_source lookup used get_by_name() only, so requests identifying a model by UUID would silently fall back to an empty source path, causing multimodal preprocessing to read configs from the wrong directory. Fall back to get_by_id() when name lookup fails. Signed-off-by: Chang Su <chang.s.su@oracle.com>
d88fcf9 to
1cecc03
Compare
…ration.rs PR #497 split multimodal processing into Phase 1 (fetch in preparation.rs) and Phase 2 (preprocess in request_building.rs) because vLLM needed raw image bytes. Now that vLLM uses the same preprocessed path as SGLang, the split is no longer justified. Collapse back to the clean single-phase pattern from PR #495 (3a05e6b). multimodal.rs: - Add process_multimodal() as single async entry point combining fetch → preprocess → expand tokens → build MultimodalData. - Inline resolve_model_spec() and preprocess() into process_multimodal() since both were single-use helpers. This also avoids computing placeholder_token twice. - Collect mm_hashes inside build_multimodal_data() in the same pass as image_data (single iteration over images). - Remove fetch_images(), process_for_backend(), build_raw_multimodal_data(), and the is_trtllm/is_vllm parameters that were added for the two-phase split. preparation.rs: - Replace Phase 1 fetch-only block with full process_multimodal() call. Store expanded token_ids and MultimodalData directly. - Add tokenizer_source empty-string guard with a clear error message instead of letting it bubble up as a confusing file-not-found from get_or_load_config(). request_building.rs: - Remove entire Phase 2 multimodal block (~55 lines). - Use take() instead of as_ref() on ctx.state.preparation since request_building is the last consumer (worker_selection already ran). This eliminates all clones on token_ids, text, multimodal_data, and tool_constraints — previously cloning MultimodalData copied megabytes of pixel data unnecessarily. Signed-off-by: Chang Su <chang.s.su@oracle.com>
…ration.rs PR #497 split multimodal processing into Phase 1 (fetch in preparation.rs) and Phase 2 (preprocess in request_building.rs) because vLLM needed raw image bytes. Now that vLLM uses the same preprocessed path as SGLang, the split is no longer justified. Collapse back to the clean single-phase pattern from PR #495 (3a05e6b). multimodal.rs: - Add process_multimodal() as single async entry point combining fetch → preprocess → expand tokens → build MultimodalData. - Inline resolve_model_spec() and preprocess() into process_multimodal() since both were single-use helpers. This also avoids computing placeholder_token twice. - Collect mm_hashes inside build_multimodal_data() in the same pass as image_data (single iteration over images). - Remove fetch_images(), process_for_backend(), build_raw_multimodal_data(), and the is_trtllm/is_vllm parameters that were added for the two-phase split. preparation.rs: - Replace Phase 1 fetch-only block with full process_multimodal() call. Store expanded token_ids and MultimodalData directly. - Add tokenizer_source empty-string guard with a clear error message instead of letting it bubble up as a confusing file-not-found from get_or_load_config(). request_building.rs: - Remove entire Phase 2 multimodal block (~55 lines). - Use take() instead of as_ref() on ctx.state.preparation since request_building is the last consumer (worker_selection already ran). This eliminates all clones on token_ids, text, multimodal_data, and tool_constraints — previously cloning MultimodalData copied megabytes of pixel data unnecessarily. Signed-off-by: Chang Su <chang.s.su@oracle.com>
Description
Problem
The gRPC multimodal pipeline only supported SGLang. Adding vLLM support revealed a fundamental architecture issue: SGLang and vLLM have completely different expectations for multimodal data.
SGLang requires:
im_token_idfor itspad_input_tokensfunctionvLLM requires:
The old
process_multimodal()ran the full SGLang-specific pipeline (fetch + preprocess + expand) inChatPreparationStage, before the backend type was known (backend is determined atClientAcquisitionStage). This meant:original_token_idsworkaround was needed to send unexpanded tokens to vLLMSolution
Split the multimodal pipeline into two phases, with the backend-agnostic work in Phase 1 and backend-specific work in Phase 2:
Phase 1 —
ChatPreparationStage(backend-agnostic):fetch_images(): Uses the multimodal tracker to extract image URLs from chat messages and fetch them concurrently as rawImageFramesProcessedMessages.multimodal_imagesasVec<Arc<ImageFrame>>MediaConnectorPhase 2 —
ChatRequestBuildingStage(backend-specific):process_for_backend()routes based onbuilder_client.is_sglang():preprocess_for_sglang(): load model config, run vision preprocessor (pixel normalization, tiling), compute prompt replacements, expand placeholder tokens, build fullMultimodalDatawith pixel tensors + metadatabuild_vllm_multimodal_data(): extract raw JPEG/PNG bytes fromImageFrame, build minimalMultimodalDatawith onlyimage_dataThis also introduces
MultimodalDataandTensorBytesas backend-agnostic types withinto_sglang_proto()andinto_vllm_proto()converters.TrackerConfig removal: The old
TrackerConfigrequired resolving the model spec (viaModelRegistry) at tracker construction to obtain placeholder token strings (forConversationSegmenttracking) and modality limits (for early rejection viaModalityLimit). Both are dead in the new architecture — token expansion operates directly ontoken_idsin Phase 2, andConversationSegment/PlaceholderMapbookkeeping is no longer needed. RemovingTrackerConfigeliminates the model registry dependency from Phase 1, so whether a model has token expansion configs (e.g. llava, phi3v, llama4_vision) is irrelevant when all Phase 1 does is download image bytes.Changes
TensorBytes,MultimodalDatatypes withinto_sglang_proto()/into_vllm_proto()converters inproto_wrapper.rsMultimodalInputsmessage to vLLM proto and engine clientprocess_multimodal()intofetch_images(),preprocess_for_sglang(),build_vllm_multimodal_data(),process_for_backend()ProcessedMessages.multimodal_inputs→multimodal_images: Option<Vec<Arc<ImageFrame>>>ChatPreparationStagetoChatRequestBuildingStageoriginal_token_idsworkaround fromPreparationOutputUintTensor/UintVecserialization: was casting u32→i64 (8 bytes) with dtype "uint32" (4 bytes)build_multimodal_dataunwrap()with proper error handling in request building stageTrackerConfig,ConversationSegment,PlaceholderHandle,PlaceholderMap,DEFAULT_PLACEHOLDERS, andModalityLimiterror variantAsyncMultiModalTracker::new()to only requireMediaConnector(no model registry needed)Test Plan
cargo test -p smg -- grpc::multimodal— all 9 tests passcargo clippy --workspace --all-targets -- -D warnings— cleanChecklist
cargo +nightly fmtpassescargo clippy --all-targets --all-features -- -D warningspassesSummary by CodeRabbit
New Features
Changes
Notes