Skip to content

feat(grpc): send preprocessed multimodal data to vLLM with hashing and structured tokens - #570

Merged
CatherineSue merged 22 commits into
mainfrom
chang/grpc-mm
Mar 1, 2026
Merged

CatherineSue merged 22 commits into
mainfrom
chang/grpc-mm

Conversation

@CatherineSue

@CatherineSue CatherineSue commented Mar 1, 2026 •

Copy link
Copy Markdown
Member

Description

Problem

PR #497 split the gRPC multimodal pipeline into two phases because vLLM
and TRT-LLM could not accept preprocessed pixel tensors — they needed
raw image bytes. This meant smg's Rust-based preprocessing was wasted for
those backends: vLLM re-ran image resize/normalize/crop via HF processor,
and TRT-LLM even decoded token IDs back to text and re-tokenized.

We discovered that vLLM can accept preprocessed tensors through its
MultiModalInputs API (with MultiModalKwargsItems.from_hf_inputs +
MultiModalFieldConfig), bypassing HF processor entirely. This eliminates
redundant CPU work and enables encoder output caching via blake3 hashes.

Additionally, the Llama 4 gRPC path produced incorrect model responses for
multi-tile images because placeholder tokens were flat repeated <|patch|>
IDs instead of the structured token sequences (with <|image_start|>,
tile separators, <|image_end|>) that HF's _prompt_split_image produces.

Solution

1. Send preprocessed data to vLLM (reverts the Phase 1/Phase 2 split from #497)

Route vLLM through the same preprocessing pipeline as SGLang. The two-phase
split is no longer needed — collapse back to a single process_multimodal()
entry point in preparation.rs. Only TRT-LLM remains on raw bytes (reverted
in this PR since TRT-LLM support is not yet ready).

2. Blake3 image hashing for encoder output caching

Compute blake3 hex-digest of raw image bytes at decode time (ImageFrame.hash).
Send per-image hashes via mm_hashes in the proto. This enables vLLM's encoder
output caching and prefix caching for multimodal requests.

3. Field layout metadata for tensor slicing

Add FieldLayout enum (Batched / Flat) to ModelProcessorSpec so the
router explicitly communicates how each tensor maps to images. Send
batched_keys and flat_keys in the proto so the vLLM gRPC server can
construct correct MultiModalFieldConfig without heuristic shape inference.

4. Structured prompt tokens for Llama 4

Change prompt_replacements trait to accept &PreprocessedImages instead of
&[ImageSize], mirroring vLLM's _get_prompt_updates(out_mm_kwargs) pattern.
For Llama 4, build structured token sequences matching HF's
_prompt_split_image format with tile row/column separators. Extract
aspect_ratios from preprocessor output to get correct tile grids (via
get_best_fit, respecting max_patches cap).

Also fixes a tuple order bug where image sizes (h, w) from the preprocessor
were read as (w, h) in the gRPC router.

5. Qwen3-VL model processor spec

Add Qwen3VLVisionSpec with config-driven placeholder token resolution via
id_to_token(image_token_id) and vision start/end token support.

Changes

Proto

  • Expand vLLM MultimodalInputs with pixel_values, model_specific_tensors,
    im_token_id, mm_placeholders, mm_hashes, batched_keys, flat_keys

multimodal crate

  • Add hasher.rs — blake3 hex-digest for per-image cache keys
  • Add FieldLayout enum and field_layouts() to ModelProcessorSpec trait
  • Change prompt_replacements() signature: &[ImageSize] → &PreprocessedImages
  • Add Llama4Spec::extract_aspect_ratios() and structured token generation
  • Add Qwen3VLVisionSpec model processor spec
  • Compute blake3 hash in MediaConnector::decode_image() and store on ImageFrame
  • Fix patches_per_image dtype: uint32 → int64 (torch.uint32 pickle crash)

model_gateway crate

  • Collapse two-phase pipeline back to single process_multimodal() in preparation.rs
  • Remove Phase 2 multimodal block from request_building.rs
  • build_multimodal_data() computes batched_keys/flat_keys from field_layouts()
  • into_vllm_proto() sends full preprocessed data instead of raw image bytes
  • Revert TRT-LLM proto expansion (not ready for this PR)
  • Use take() instead of clone() on PreparationOutput to avoid copying
    megabytes of pixel data

Downstream changes required

vLLM (grpc_server.py):

  • _build_preprocessed_mm_inputs() — deserialize proto tensors, construct
    MultiModalFieldConfig from batched_keys/flat_keys, build is_embed
    mask on PlaceholderRange using im_token_id

SGLang: No downstream changes needed. expand_tokens() now computes
patch-only placeholder offsets (contiguous runs of im_token_id) during
token expansion at zero extra cost, and into_sglang_proto() uses them
instead of the full structural mm_placeholders.

Why sglang needs patch-only offsets but vLLM does not

                    SMG Gateway (Rust)
                         │
                   mm_placeholders = full structural range
                   e.g. offset=179, length=2467
                   covers: <|image_start|> + patches + separators + <|image_end|>
                         │
              ┌──────────┴──────────┐
              │                     │
           vLLM                  sglang
              │                     │
     has is_embed mask        no is_embed concept
              │                     │
     PlaceholderRange         uses mm_placeholders
     offset=179, length=2467  to locate embeddings
              │                     │
     masked_scatter_()        torch.isin(input_ids, pad_value)
     only writes where        counts N pad_value tokens
     is_embed=True            expects N == num_embeddings
     (patch positions)              │
              │               if offsets include structural
     structural tokens        tokens → N = 2467
     are skipped by mask      but encoder only outputs 2448
              │               embeddings (patches only)
     works correctly                │
              ✅              ❌ 2467 != 2448 → crash
                                    │
                              ─── fix ───
                                    │
                              into_sglang_proto()
                              scans within mm_placeholders
                              emits patch-only offsets
                              offset=180, length=2448
                                    │
                              N = 2448 == 2448
                                    │
                              ✅ works correctly

Test Plan

  • cargo test -p llm-multimodal — all tests pass including new Llama 4
    structured token tests and Qwen3-VL spec tests
  • cargo clippy --workspace --all-targets -- -D warnings — clean
  • E2E: vLLM receives preprocessed pixel_values + model_specific_tensors,
    bypasses HF processor, produces correct responses
  • E2E: Llama 4 multi-tile images produce correct structured token sequences
    matching HF _prompt_split_image output

Same request, before vs after in the same screenshot.

Before Structured prompt tokens for Llama 4, the model only sees one image.

Screenshot 2026-02-28 at 9 25 25 PM Screenshot 2026-02-28 at 9 34 41 PM
Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated

Summary by CodeRabbit

Release Notes

  • New Features

    • Enhanced multimodal input handling with support for per-image tensor metadata and hashing
    • Expanded vision model support including improved Qwen3VL and enhanced layout management for Llama4Vision
  • Refactor

    • Consolidated multimodal processing pipeline for more consistent tensor handling across all backends
    • Improved per-image metadata extraction and tensor layout configuration

@github-actions github-actions Bot added dependencies Dependency updates grpc gRPC client and router changes multimodal Multimodal crate changes model-gateway Model gateway crate changes labels Mar 1, 2026
@coderabbitai

coderabbitai Bot commented Mar 1, 2026 •

Copy link
Copy Markdown

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

The changes expand the multimodal input surface by introducing TensorData and PlaceholderRange protobuf messages, consolidating the preprocessing pipeline into a unified process_multimodal function, adding blake3-based image hashing, tracking per-image and per-batch metadata (mm_hashes, batched_keys, flat_keys), and introducing a FieldLayout abstraction for tensor mapping strategies across model specifications.

Changes

Cohort / File(s) Summary
Protobuf multimodal messages
grpc_client/proto/vllm_engine.proto
Added TensorData and PlaceholderRange message types; restructured MultimodalInputs to replace raw image_data with pixel_values (TensorData), model_specific_tensors map, placeholders, hashes, and batched/flat keys.
Multimodal processing pipeline
model_gateway/src/routers/grpc/multimodal.rs
Consolidated fetch_images → process_multimodal function that inlines the complete pipeline: image fetching, model spec resolution, preprocessing, prompt replacement computation, token expansion, and MultimodalData construction; removed phase-specific helpers.
Multimodal data structures
model_gateway/src/routers/grpc/..., multimodal/src/...
Updated ProcessedMessages, MultimodalData, and ImageFrame to carry expanded metadata: sglang_patch_offsets, mm_hashes, batched_keys, flat_keys; added hash field to ImageFrame.
Proto conversion layer
model_gateway/src/routers/grpc/proto_wrapper.rs
Extended into_vllm_proto and into_sglang_proto to populate new MultimodalInputs fields (mm_hashes, batched_keys, flat_keys) and use sglang_patch_offsets when available.
Chat request handling
model_gateway/src/routers/grpc/regular/stages/chat/preparation.rs, request_building.rs, model_gateway/src/routers/grpc/utils.rs
Replaced fetch_images call with process_multimodal invocation; refactored to own preparation state and multimodal_data; updated field references from multimodal_images to multimodal_data; eliminated redundant multimodal processing.
Image hashing utilities
multimodal/src/hasher.rs, multimodal/src/lib.rs
Added new blake3-based hashing module with hash_image and hash_images functions; integrated hasher into media.rs to compute and attach image hashes at decode time.
Model spec registry and field layouts
multimodal/src/registry.rs, multimodal/src/types.rs, multimodal/src/vision/image_processor.rs
Added FieldLayout enum (Batched, Flat) and field_layouts() method to ModelProcessorSpec; updated prompt_replacements signature to consume PreprocessedImages; added Qwen3VLVisionSpec; extended all specs to compute and expose per-image and batched/flat tensor metadata.
Vision processor extensions
multimodal/src/vision/processors/llama4_vision.rs, multimodal/src/vision/image_processor.rs
Added patches_per_image tracking and per-image metadata extraction; introduced batched_keys/flat_keys helpers on PreprocessedImages; added first_dim() and num_images() utility methods.
Dependencies
multimodal/Cargo.toml
Added blake3 as workspace dependency.

Sequence Diagram

sequenceDiagram
    participant Client
    participant ChatPreparation
    participant ProcessMultimodal
    participant ModelRegistry
    participant VisionProcessor
    participant ProtoConverter
    participant vLLMBackend

    Client->>ChatPreparation: Send chat message with images
    ChatPreparation->>ChatPreparation: Compute tokenizer source
    ChatPreparation->>ProcessMultimodal: Invoke with messages, model_id, tokenizer
    
    ProcessMultimodal->>VisionProcessor: Fetch images from URLs/base64
    VisionProcessor->>VisionProcessor: Decode and compute blake3 hash
    ProcessMultimodal->>ModelRegistry: Resolve model spec
    ModelRegistry->>ModelRegistry: Return spec with field_layouts
    
    ProcessMultimodal->>VisionProcessor: Preprocess images
    VisionProcessor->>VisionProcessor: Apply vision processor, compute aspect ratios
    ProcessMultimodal->>ModelRegistry: Build prompt_replacements
    ModelRegistry->>ModelRegistry: Generate tokens from per-image metadata
    
    ProcessMultimodal->>ProcessMultimodal: Expand token_ids with prompt tokens
    ProcessMultimodal->>ProcessMultimodal: Compute patch_offsets and placeholder ranges
    ProcessMultimodal->>ProcessMultimodal: Extract batched_keys and flat_keys
    ProcessMultimodal->>ChatPreparation: Return MultimodalOutput
    
    ChatPreparation->>ChatPreparation: Update token_ids, store multimodal_data
    ChatPreparation->>ProtoConverter: Call into_vllm_proto
    ProtoConverter->>ProtoConverter: Map pixel_values, mm_hashes, batched/flat_keys
    ProtoConverter->>vLLMBackend: Send MultimodalInputs proto
    vLLMBackend->>vLLMBackend: Process multimodal tokens and tensors
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes

Possibly related issues

  • Implements the exact code-level changes specified in #566: adding TensorData/PlaceholderRange protos, consolidating preprocessing into process_multimodal, computing and sending mm_hashes/batched_keys/flat_keys, and introducing field layouts for per-image tensor metadata.

Possibly related PRs

Suggested labels

protocols, multimodal, vision

Suggested reviewers

  • slin1237
  • key4ng

Poem

🐰 Hops of joy through tensor lanes,
Hashes bloom in blake3's chains,
From fetch to spec, one flow unified,
Per-image metadata comes alive,
Batched and flat, now side by side! ✨

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately describes the main change: sending preprocessed multimodal data to vLLM with hashing and structured tokens, which is the core objective of this PR.
Docstring Coverage ✅ Passed Docstring coverage is 87.32% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
  • 📝 Generate docstrings (stacked PR)
  • 📝 Generate docstrings (commit on current branch)
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch chang/grpc-mm

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly enhances the multimodal processing pipeline by enabling vLLM to directly consume preprocessed image data, which was previously handled inefficiently. The changes aim to improve performance by reducing redundant preprocessing and ensure correctness for models like Llama 4 by generating properly structured prompt tokens. Additionally, it introduces image hashing for better caching and provides explicit metadata for tensor slicing, streamlining the interaction between the gateway and vLLM.

Highlights

  • vLLM Preprocessing Integration: Reverted the two-phase multimodal pipeline split from a previous PR, allowing vLLM to directly accept preprocessed pixel tensors and model-specific data, eliminating redundant CPU work.
  • Image Hashing for Caching: Implemented Blake3 hex-digest hashing of raw image bytes to enable vLLM's encoder output caching and prefix caching for multimodal requests.
  • Structured Prompt Tokens for Llama 4: Updated Llama 4's prompt replacement logic to generate structured token sequences, including image start/end and tile separators, matching Hugging Face's _prompt_split_image format, and fixed a tuple order bug for image sizes.
  • Field Layout Metadata: Introduced FieldLayout enum and related logic to explicitly communicate how multimodal tensors map to images, allowing the vLLM gRPC server to correctly construct MultiModalFieldConfig.
  • Qwen3-VL Model Support: Added a new Qwen3VLVisionSpec for Qwen3-VL models, supporting config-driven placeholder token resolution and vision start/end tokens.
Changelog
  • grpc_client/proto/vllm_engine.proto
    • Added TensorData and PlaceholderRange message types.
    • Expanded MultimodalInputs message with fields for preprocessed pixel values, model-specific tensors, image token IDs, placeholder ranges, Blake3 hashes, and tensor field layout keys.
  • model_gateway/src/routers/grpc/mod.rs
    • Removed Arc<ImageFrame> and multimodal_images from ProcessedMessages.
    • Replaced multimodal_images with multimodal_data in ProcessedMessages to hold preprocessed multimodal output.
  • model_gateway/src/routers/grpc/multimodal.rs
    • Removed ImageSize import as it's no longer directly used in this file.
    • Deleted resolve_model_spec function.
    • Refactored fetch_images and preprocess_for_sglang into a single process_multimodal function, handling the complete multimodal pipeline.
    • Removed build_raw_multimodal_data and process_for_backend functions.
    • Updated build_multimodal_data to include Blake3 hashes and field layout keys from the model spec.
  • model_gateway/src/routers/grpc/proto_wrapper.rs
    • Updated MultimodalData struct to include mm_hashes, batched_keys, and flat_keys.
    • Modified into_vllm_proto to populate the newly added multimodal fields for vLLM proto conversion.
  • model_gateway/src/routers/grpc/regular/stages/chat/preparation.rs
    • Modified ChatPreparationStage to perform full multimodal processing (fetch, preprocess, expand tokens, hash) in a single step.
    • Updated to use multimodal::process_multimodal and store its output in multimodal_data.
  • model_gateway/src/routers/grpc/regular/stages/chat/request_building.rs
    • Removed the two-phase multimodal processing logic, as it's now handled in the preparation stage.
    • Updated to directly use multimodal_data from the PreparationOutput and take ownership of relevant fields to avoid cloning.
  • model_gateway/src/routers/grpc/utils.rs
    • Updated ProcessedMessages to use multimodal_data instead of multimodal_images.
  • multimodal/Cargo.toml
    • Added blake3 as a workspace dependency.
  • multimodal/src/hasher.rs
    • Added a new module hasher for Blake3 image hashing utilities.
    • Implemented hash_image to compute a Blake3 hex-digest for raw bytes.
    • Implemented hash_images to compute per-image hashes keyed by modality.
  • multimodal/src/lib.rs
    • Added hasher module to the library exports.
  • multimodal/src/media.rs
    • Updated MediaConnector::decode_image to compute and store a Blake3 hash for each ImageFrame.
  • multimodal/src/registry.rs
    • Imported FieldLayout and PreprocessedImages.
    • Modified ModelProcessorSpec::prompt_replacements trait method to accept &PreprocessedImages instead of &[ImageSize].
    • Added a default field_layouts method to ModelProcessorSpec to declare tensor layout.
    • Added Qwen3VLVisionSpec for Qwen3-VL models, including placeholder token resolution and structured prompt replacements.
    • Updated LlavaSpec, QwenVLVisionSpec, and Phi3VisionSpec to use the new prompt_replacements signature.
    • Enhanced Llama4Spec::prompt_replacements to generate structured token sequences with image start/end and tile separators, and added extract_aspect_ratios helper.
    • Added field_layouts implementation for Qwen3VLVisionSpec, QwenVLVisionSpec, and Llama4Spec.
  • multimodal/src/types.rs
    • Added hash: String field to the ImageFrame struct.
    • Updated ImageFrame::new constructor to accept and store the image hash.
    • Added FieldLayout enum to describe how multimodal tensor dimensions map to images, with Batched and Flat variants.
  • multimodal/src/vision/image_processor.rs
    • Imported FieldLayout.
    • Added first_dim method to ModelSpecificValue to get the first dimension of tensor variants.
    • Added num_images, batched_keys, and flat_keys methods to PreprocessedImages for extracting tensor layout information.
  • multimodal/src/vision/processors/llama4_vision.rs
    • Updated Llama4VisionProcessor::preprocess to store patches_per_image as i64 in model-specific data to avoid torch.uint32 pickle issues.
Activity
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@mergify

This comment was marked as resolved.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request is a significant and well-executed refactoring of the multimodal processing pipeline. It streamlines the process by removing the two-phase approach for vLLM, allowing preprocessed data to be sent directly. This is a major improvement for both performance and code simplicity. The introduction of image hashing for caching, explicit field layout metadata, and structured prompt tokens for Llama 4 are all excellent additions that enhance correctness and efficiency. The code is generally of high quality, with good attention to performance details like avoiding unnecessary clones. I've found one potential issue that could lead to a panic under specific data corruption scenarios, which I've detailed in a specific comment, aligning with the repository's rule on preventing panics. Overall, this is a great contribution.

Comment thread multimodal/src/registry.rs Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 837d3bc5ae

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment on lines +263 to +266
metadata
.tokenizer
.id_to_token(token_id)
.ok_or_else(|| ModelRegistryError::TokenNotFound {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Avoid requiring reverse token lookup for Qwen3 placeholders

The new Qwen3 placeholder resolution now fails hard if id_to_token(image_token_id) is unavailable. That adds a strict reverse-vocab dependency that was not required before, so tokenizer implementations that can encode tokens (token_to_id) but do not provide reverse lookup will reject all Qwen3 multimodal requests with TokenNotFound. This should fall back to a known placeholder token string or config-driven token text instead of treating missing reverse lookup as fatal.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pre-existing issue, not introduced by this PR. The qwen3_vl_includes_end_token test was already failing before these changes. Will address in a separate PR.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 4

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@model_gateway/src/routers/grpc/regular/stages/chat/preparation.rs`:
- Around line 87-97: The code in ChatPreparationStage::execute currently treats
an empty tokenizer_source as an immediate bad_request error and returns Err;
instead, remove the early return and do not convert an empty tokenizer_source
into a 400 so the intended fallback resolution can proceed. Replace the return
Err(error::bad_request(...)) with either no-op or a non-fatal log (e.g.,
trace/warn) and allow the function to continue using an empty tokenizer_source
to trigger registry fallback logic elsewhere; update any surrounding logic that
assumed a guaranteed tokenizer_source accordingly (refer to tokenizer_source and
ChatPreparationStage::execute).

In `@multimodal/src/registry.rs`:
- Around line 515-523: The current parsing of aspect_ratios can panic because it
only checks data.len() >= 2 but then assumes every chunk from chunks(2) has two
elements; change the guard and iteration to ensure even-length pairs and avoid
indexing panics: when matching ModelSpecificValue::IntTensor for
preprocessed.model_specific.get("aspect_ratios"), require data.len() >= 2 &&
data.len() % 2 == 0 (and optionally validate shape vs. data length), and iterate
using chunks_exact(2) or a filter_map that safely converts only complete
2-element slices into (usize, usize) tuples so no chunk[1] indexing can panic.
- Around line 261-269: The placeholder_token method currently treats
metadata.tokenizer.id_to_token(token_id) returning None as an error (raising
ModelRegistryError::TokenNotFound), which breaks Qwen3 when reverse_vocab lacks
that ID; change placeholder_token (and keep using Self::pad_token_id(metadata)?
as u32 and the same token_id formatting) to return a safe fallback string
instead of error — call metadata.tokenizer.id_to_token(token_id) and if it
returns None return a generic placeholder (e.g., empty string "" or "
image_token " as used elsewhere) so the function succeeds even when reverse
lookup is unavailable rather than producing ModelRegistryError::TokenNotFound.
- Around line 133-143: The mapping in image_sizes_hw currently assumes tuples
are (height, width) but producers are inconsistent; standardize on a canonical
(width, height) tuple. Update image_sizes_hw to interpret tuples as (width,
height) and construct ImageSize { width: w, height: h } from (w, h) tuples, then
fix the producers that emit (height, width) (phi4_vision.rs, pixtral.rs,
llama4_vision.rs) to push (width, height) instead of (height, width); leave
llava.rs and phi3_vision.rs as-is since they already produce (width, height).
Ensure all producers and the helper use the same (width, height) convention.

ℹ️ Review info

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between cab2d13 and 837d3bc.

📒 Files selected for processing (15)
  • grpc_client/proto/vllm_engine.proto
  • model_gateway/src/routers/grpc/mod.rs
  • model_gateway/src/routers/grpc/multimodal.rs
  • model_gateway/src/routers/grpc/proto_wrapper.rs
  • model_gateway/src/routers/grpc/regular/stages/chat/preparation.rs
  • model_gateway/src/routers/grpc/regular/stages/chat/request_building.rs
  • model_gateway/src/routers/grpc/utils.rs
  • multimodal/Cargo.toml
  • multimodal/src/hasher.rs
  • multimodal/src/lib.rs
  • multimodal/src/media.rs
  • multimodal/src/registry.rs
  • multimodal/src/types.rs
  • multimodal/src/vision/image_processor.rs
  • multimodal/src/vision/processors/llama4_vision.rs

Comment thread model_gateway/src/routers/grpc/regular/stages/chat/preparation.rs
Comment thread multimodal/src/registry.rs
Comment thread multimodal/src/registry.rs
Comment thread multimodal/src/registry.rs
Add hasher.rs to the multimodal crate for computing per-image
blake3 hex-digest hashes. These hashes are used as cache keys
for vLLM's MultiModalInputs.mm_hashes field, enabling encoder
output caching and prefix caching.

Signed-off-by: Chang Su <chang.s.su@oracle.com>
Add TensorData, PlaceholderRange messages and expand MultimodalInputs
to carry preprocessed pixel_values, model_specific_tensors, im_token_id,
mm_placeholders, and mm_hashes (blake3) for encoder output caching.
This enables the Rust router to send fully preprocessed multimodal data
to vLLM instead of raw image bytes.

Signed-off-by: Chang Su <chang.s.su@oracle.com>
Add mm_hashes field to MultimodalData for per-image blake3 cache keys.
Expand into_vllm_proto() to send full preprocessed data (pixel_values,
model_specific_tensors, mm_placeholders, mm_hashes) instead of raw
image bytes only.

Signed-off-by: Chang Su <chang.s.su@oracle.com>
Route vLLM through the same preprocessing pipeline as SGLang instead
of sending raw image bytes. This computes pixel_values, model_specific
tensors, placeholder expansion, and blake3 image hashes for encoder
output caching. Only TRT-LLM remains on the raw bytes path.

Signed-off-by: Chang Su <chang.s.su@oracle.com>
Compute blake3 hex-digest of raw image bytes in decode_image() and
store it on ImageFrame.hash. This makes the hash available throughout
the pipeline without recomputing, enabling vLLM encoder output caching.

Signed-off-by: Chang Su <chang.s.su@oracle.com>
…imodal_data

Change the multimodal field on ProcessedMessages from
`multimodal_images: Option<Vec<Arc<ImageFrame>>>` to
`multimodal_data: Option<MultimodalData>`, carrying the fully
preprocessed backend-agnostic data through the pipeline instead of
raw image frames.

Update the two construction sites in utils::process_chat_messages to
initialize the new field.

Signed-off-by: Chang Su <chang.s.su@oracle.com>
…ration.rs

PR #497 split multimodal processing into Phase 1 (fetch in
preparation.rs) and Phase 2 (preprocess in request_building.rs) because
vLLM needed raw image bytes. Now that vLLM uses the same preprocessed
path as SGLang, the split is no longer justified. Collapse back to the
clean single-phase pattern from PR #495 (3a05e6b).

multimodal.rs:
- Add process_multimodal() as single async entry point combining
  fetch → preprocess → expand tokens → build MultimodalData.
- Inline resolve_model_spec() and preprocess() into
  process_multimodal() since both were single-use helpers. This also
  avoids computing placeholder_token twice.
- Collect mm_hashes inside build_multimodal_data() in the same pass
  as image_data (single iteration over images).
- Remove fetch_images(), process_for_backend(),
  build_raw_multimodal_data(), and the is_trtllm/is_vllm parameters
  that were added for the two-phase split.

preparation.rs:
- Replace Phase 1 fetch-only block with full process_multimodal()
  call. Store expanded token_ids and MultimodalData directly.
- Add tokenizer_source empty-string guard with a clear error message
  instead of letting it bubble up as a confusing file-not-found from
  get_or_load_config().

request_building.rs:
- Remove entire Phase 2 multimodal block (~55 lines).
- Use take() instead of as_ref() on ctx.state.preparation since
  request_building is the last consumer (worker_selection already ran).
  This eliminates all clones on token_ids, text, multimodal_data, and
  tool_constraints — previously cloning MultimodalData copied megabytes
  of pixel data unnecessarily.

Signed-off-by: Chang Su <chang.s.su@oracle.com>
Signed-off-by: Chang Su <chang.s.su@oracle.com>
Add a `batched_keys` field to the vLLM MultimodalInputs proto message
so the Rust router explicitly communicates which model-specific tensors
are per-image (first dim = num_images) vs shared. This removes the need
for the Python side to infer field configs from tensor shapes.

- proto: `repeated string batched_keys = 7` on MultimodalInputs
- proto_wrapper.rs: add `batched_keys` to MultimodalData, pass through
  in `into_vllm_proto()`
- multimodal.rs: compute batched_keys in `build_multimodal_data()` by
  filtering model_specific_tensors where shape[0] == num_images

Signed-off-by: Chang Su <chang.s.su@oracle.com>
Add TensorData, PlaceholderRange, and expand MultimodalInput in
trtllm_service.proto with preprocessed data fields (pixel_values,
model_specific_tensors, mm_placeholders, mm_hashes, batched_keys).

Update into_trtllm_proto() to send full preprocessed data instead of
only raw image bytes. This enables the TRT-LLM gRPC server to bypass
its HF input processor when preprocessed tensors are present.

Signed-off-by: Chang Su <chang.s.su@oracle.com>
… fields"

This reverts commit fb16c7208515b1cfcd145a004d3a33e43bceb48d.

Signed-off-by: Chang Su <chang.s.su@oracle.com>
Add FieldLayout enum to declare how each tensor's first dimension maps
to images (Batched vs Flat). Replace the heuristic batched_keys detection
with explicit field_layouts() from the model spec. Add flat_keys map to
vLLM proto and MultimodalData for variable-length per-image slicing
(e.g. Llama4 pixel_values split by patches_per_image).

Signed-off-by: Chang Su <chang.s.su@oracle.com>
Add Qwen3VLVisionSpec with config-driven placeholder token resolution
via id_to_token(image_token_id) instead of hardcoded string. Includes
vision_start/end token support and patch grid calculation. Registered
before QwenVL so qwen3 matches first.

Signed-off-by: Chang Su <chang.s.su@oracle.com>
…2 pickle crash

torch.uint32 tensors lack a typed storage class in PyTorch, causing
UntypedStorage.dtype AttributeError when multiproc workers deserialize
via shared memory broadcast. Use i64 to match vLLM's native HF
processor output.

Signed-off-by: Chang Su <chang.s.su@oracle.com>
…cessedImages to prompt_replacements

Change prompt_replacements trait to accept &PreprocessedImages instead of
&[ImageSize], mirroring vLLM's _get_prompt_updates(out_mm_kwargs) pattern.
This lets each model spec extract whatever metadata it needs from the
preprocessor output.

For Llama4, build structured token sequences matching HF's
_prompt_split_image format (<|image_start|>, tile separators,
<|image_end|>) instead of flat repeated <|patch|> tokens. Extract
aspect_ratios from the preprocessor to get correct tile grids.

Also fixes a pre-existing tuple order bug where image_sizes (h,w) from
the preprocessor were read as (w,h) in the gRPC router.

Signed-off-by: Chang Su <chang.s.su@oracle.com>
Signed-off-by: Chang Su <chang.s.su@oracle.com>
vLLM no longer receives raw image bytes — it uses preprocessed
pixel_values exclusively. Reserve field number 1 for backward
compatibility. image_data remains in the SGLang and TRT-LLM protos
where it is still used.

Signed-off-by: Chang Su <chang.s.su@oracle.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 62b4846b14

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread multimodal/src/registry.rs

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

♻️ Duplicate comments (4)
model_gateway/src/routers/grpc/regular/stages/chat/preparation.rs (1)

79-97: ⚠️ Potential issue | 🟠 Major

Do not hard-fail on empty tokenizer_source in the multimodal path.

Line [87]-[97] turns a tokenizer registry miss into a 400, which breaks the intended fallback flow for multimodal config resolution.

💡 Proposed fix
                 let tokenizer_source = ctx
                     .components
                     .tokenizer_registry
                     .get_by_name(model_id)
                     .or_else(|| ctx.components.tokenizer_registry.get_by_id(model_id))
                     .map(|e| e.source)
                     .unwrap_or_default();

-                if tokenizer_source.is_empty() {
-                    error!(
-                        function = "ChatPreparationStage::execute",
-                        model = %model_id,
-                        "Tokenizer source path not found for multimodal processing"
-                    );
-                    return Err(error::bad_request(
-                        "multimodal_config_missing",
-                        format!("Tokenizer source path not found for model: {model_id}"),
-                    ));
-                }
+                let tokenizer_source = if tokenizer_source.is_empty() {
+                    model_id.to_string()
+                } else {
+                    tokenizer_source
+                };

Based on learnings: in repo lightseekorg/smg, when fetching tokenizer_source, an empty string is an intentional valid fallback when the model isn't in the tokenizer registry.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@model_gateway/src/routers/grpc/regular/stages/chat/preparation.rs` around
lines 79 - 97, The current ChatPreparationStage::execute turns an empty
tokenizer_source (from tokenizer_registry.get_by_name/get_by_id -> .map(|e|
e.source).unwrap_or_default()) into a hard 400 error, breaking multimodal
fallback; change this to allow empty tokenizer_source as a valid fallback by
removing the error::bad_request return and replacing it with a non-fatal log
(info/warn/debug) so execution can continue for multimodal config resolution;
update the block that inspects tokenizer_source to log the missing source
(including model = %model_id) but do not Err out from
ChatPreparationStage::execute when tokenizer_source.is_empty().
multimodal/src/registry.rs (3)

515-523: ⚠️ Potential issue | 🔴 Critical

Guard aspect_ratios parsing against malformed tensor lengths.

Line [518]-[523] uses chunks(2) and indexes chunk[1]; odd-length data can panic.

🐛 Proposed hardening
         if let Some(ModelSpecificValue::IntTensor { data, shape }) =
             preprocessed.model_specific.get("aspect_ratios")
         {
-            if shape.len() == 2 && shape[1] == 2 && data.len() >= 2 {
+            if shape.len() == 2 && shape[1] == 2 && data.len() == shape[0] * shape[1] {
                 return data
-                    .chunks(2)
-                    .map(|chunk| (chunk[0] as usize, chunk[1] as usize))
+                    .chunks_exact(2)
+                    .map(|pair| (pair[0] as usize, pair[1] as usize))
                     .collect();
             }
         }
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@multimodal/src/registry.rs` around lines 515 - 523, The parsing of
aspect_ratios in the block using
preprocessed.model_specific.get("aspect_ratios") and matching
ModelSpecificValue::IntTensor { data, shape } can panic because data.chunks(2)
then indexing chunk[1] assumes even-length chunks; change the logic to only
accept well-formed pairs (e.g., require data.len() >= 2 and data.len() % 2 == 0)
or use chunks_exact(2) so you only iterate valid pair slices, and return an
error/empty result if the tensor length is malformed, keeping the existing shape
check (shape.len() == 2 && shape[1] == 2) and the returned type from the
function intact.

133-143: ⚠️ Potential issue | 🟠 Major

Standardize image_sizes tuple orientation before grid/tile math.

Line [133]-[143] and Line [525]-[533] treat stored sizes as (height, width), but PreprocessedImages.image_sizes is documented as (width, height) in multimodal/src/vision/image_processor.rs. This can transpose non-square image calculations.

💡 Proposed fix (align with documented contract)
-/// Convert preprocessor `(height, width)` tuples to `ImageSize` values.
+/// Convert preprocessor `(width, height)` tuples to `ImageSize` values.
 fn image_sizes_hw(preprocessed: &PreprocessedImages) -> Vec<ImageSize> {
     preprocessed
         .image_sizes
         .iter()
-        .map(|&(h, w)| ImageSize {
-            width: w,
-            height: h,
+        .map(|&(w, h)| ImageSize {
+            width: w,
+            height: h,
         })
         .collect()
 }

Also applies to: 525-533

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@multimodal/src/registry.rs` around lines 133 - 143, The code in
image_sizes_hw currently treats tuples as (height, width) but
PreprocessedImages.image_sizes is documented as (width, height); update the
mapping so the first element is width and the second is height (e.g.,
pattern-match the tuple as |&(w, h)| and create ImageSize { width: w, height: h
}). Apply the same correction to the duplicate occurrence that performs the same
tuple-to-ImageSize conversion (the block around the second occurrence referenced
in the review) so all grid/tile math uses the documented (width, height)
orientation.

261-269: ⚠️ Potential issue | 🔴 Critical

Qwen3 placeholder resolution should not hard-fail on missing reverse lookup.

Line [261]-[269] makes multimodal prompt replacement fail whenever id_to_token(image_token_id) returns None, even if image_token_id is valid. That creates a brittle runtime dependency on reverse vocab completeness.

Run this read-only check to confirm the failure path and missing fallback:

#!/bin/bash
# 1) Inspect Qwen3 placeholder_token implementation
rg -nP --type rust 'fn placeholder_token\(&self, metadata: &ModelMetadata\)' multimodal/src/registry.rs -A18 -B3

echo "---"

# 2) Confirm id_to_token is used as a hard requirement in this file
rg -nP --type rust 'id_to_token\(' multimodal/src/registry.rs -A3 -B3

echo "---"

# 3) Confirm test tokenizer reverse lookup behavior in this file
rg -nP --type rust 'fn id_to_token\(&self, _id: u32\) -> Option<String>' multimodal/src/registry.rs -A4 -B2

echo "---"

# 4) Confirm multimodal expansion currently searches by token string-derived ID
rg -nP --type rust 'let search_token_id = tokenizer\.token_to_id\(&placeholder_token\)' model_gateway/src/routers/grpc/multimodal.rs -A4 -B4
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@multimodal/src/registry.rs` around lines 261 - 269, The placeholder_token
function currently hard-fails when metadata.tokenizer.id_to_token(token_id)
returns None; instead, change it to return a sensible fallback string (e.g.
format!("image_token_id:{token_id}")) rather than an Err. Locate fn
placeholder_token(&self, metadata: &ModelMetadata) and replace the ok_or_else
path with a match/map_or_else that returns the found token string if Some,
otherwise returns Ok(format!("image_token_id:{token_id}")), preserving the same
token_id computation via Self::pad_token_id(metadata) and avoiding
ModelRegistryError::TokenNotFound for this case.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Duplicate comments:
In `@model_gateway/src/routers/grpc/regular/stages/chat/preparation.rs`:
- Around line 79-97: The current ChatPreparationStage::execute turns an empty
tokenizer_source (from tokenizer_registry.get_by_name/get_by_id -> .map(|e|
e.source).unwrap_or_default()) into a hard 400 error, breaking multimodal
fallback; change this to allow empty tokenizer_source as a valid fallback by
removing the error::bad_request return and replacing it with a non-fatal log
(info/warn/debug) so execution can continue for multimodal config resolution;
update the block that inspects tokenizer_source to log the missing source
(including model = %model_id) but do not Err out from
ChatPreparationStage::execute when tokenizer_source.is_empty().

In `@multimodal/src/registry.rs`:
- Around line 515-523: The parsing of aspect_ratios in the block using
preprocessed.model_specific.get("aspect_ratios") and matching
ModelSpecificValue::IntTensor { data, shape } can panic because data.chunks(2)
then indexing chunk[1] assumes even-length chunks; change the logic to only
accept well-formed pairs (e.g., require data.len() >= 2 and data.len() % 2 == 0)
or use chunks_exact(2) so you only iterate valid pair slices, and return an
error/empty result if the tensor length is malformed, keeping the existing shape
check (shape.len() == 2 && shape[1] == 2) and the returned type from the
function intact.
- Around line 133-143: The code in image_sizes_hw currently treats tuples as
(height, width) but PreprocessedImages.image_sizes is documented as (width,
height); update the mapping so the first element is width and the second is
height (e.g., pattern-match the tuple as |&(w, h)| and create ImageSize { width:
w, height: h }). Apply the same correction to the duplicate occurrence that
performs the same tuple-to-ImageSize conversion (the block around the second
occurrence referenced in the review) so all grid/tile math uses the documented
(width, height) orientation.
- Around line 261-269: The placeholder_token function currently hard-fails when
metadata.tokenizer.id_to_token(token_id) returns None; instead, change it to
return a sensible fallback string (e.g. format!("image_token_id:{token_id}"))
rather than an Err. Locate fn placeholder_token(&self, metadata: &ModelMetadata)
and replace the ok_or_else path with a match/map_or_else that returns the found
token string if Some, otherwise returns
Ok(format!("image_token_id:{token_id}")), preserving the same token_id
computation via Self::pad_token_id(metadata) and avoiding
ModelRegistryError::TokenNotFound for this case.

ℹ️ Review info

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 837d3bc and 62b4846.

📒 Files selected for processing (15)
  • grpc_client/proto/vllm_engine.proto
  • model_gateway/src/routers/grpc/mod.rs
  • model_gateway/src/routers/grpc/multimodal.rs
  • model_gateway/src/routers/grpc/proto_wrapper.rs
  • model_gateway/src/routers/grpc/regular/stages/chat/preparation.rs
  • model_gateway/src/routers/grpc/regular/stages/chat/request_building.rs
  • model_gateway/src/routers/grpc/utils.rs
  • multimodal/Cargo.toml
  • multimodal/src/hasher.rs
  • multimodal/src/lib.rs
  • multimodal/src/media.rs
  • multimodal/src/registry.rs
  • multimodal/src/types.rs
  • multimodal/src/vision/image_processor.rs
  • multimodal/src/vision/processors/llama4_vision.rs

…aspect_ratios

Prevent potential panic when aspect_ratios tensor has mismatched data
length. Use data.len() == shape[0] * 2 instead of data.len() >= 2,
and chunks_exact(2) instead of chunks(2).

Signed-off-by: Chang Su <chang.s.su@oracle.com>
Comment thread multimodal/src/registry.rs

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

♻️ Duplicate comments (1)
multimodal/src/registry.rs (1)

261-269: ⚠️ Potential issue | 🟠 Major

Test will fail: TestTokenizer::id_to_token always returns None.

The Qwen3VLVisionSpec::placeholder_token method requires id_to_token(token_id) to succeed, but TestTokenizer::id_to_token at line 711-712 unconditionally returns None. This causes prompt_replacements to fail with TokenNotFound, making the test at line 806 (qwen3_vl_includes_end_token) panic on .unwrap().

Either update TestTokenizer to support reverse lookup, or provide a fallback in placeholder_token:

🧪 Option 1: Fix TestTokenizer to support id_to_token
 struct TestTokenizer {
     vocab: HashMap<String, u32>,
+    reverse_vocab: HashMap<u32, String>,
 }

 impl TestTokenizer {
     fn new(pairs: &[(&str, u32)]) -> Self {
         let vocab = pairs
             .iter()
             .map(|(token, id)| ((*token).to_string(), *id))
             .collect();
-        Self { vocab }
+        let reverse_vocab = pairs
+            .iter()
+            .map(|(token, id)| (*id, (*token).to_string()))
+            .collect();
+        Self { vocab, reverse_vocab }
     }
 }

 // ... in TokenizerTrait impl:
 fn id_to_token(&self, id: u32) -> Option<String> {
-    None
+    self.reverse_vocab.get(&id).cloned()
 }

Then update the test to include the image token:

-let tokenizer = TestTokenizer::new(&[("<image>", 999)]);
+let tokenizer = TestTokenizer::new(&[
+    ("<image>", 999),
+    ("<|image_pad|>", 151655),  // image_token_id from config
+]);
🛡️ Option 2: Use fallback in placeholder_token
 fn placeholder_token(&self, metadata: &ModelMetadata) -> RegistryResult<String> {
     let token_id = Self::pad_token_id(metadata)? as u32;
     metadata
         .tokenizer
         .id_to_token(token_id)
-        .ok_or_else(|| ModelRegistryError::TokenNotFound {
-            token: format!("image_token_id:{token_id}"),
-        })
+        .unwrap_or_else(|| format!("<|image_pad:{token_id}|>"))
+        .pipe(Ok)
 }

Also applies to: 711-712

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@multimodal/src/registry.rs` around lines 261 - 269, The test fails because
TestTokenizer::id_to_token always returns None causing
Qwen3VLVisionSpec::placeholder_token to hit ModelRegistryError::TokenNotFound;
fix by making placeholder_token tolerant: call
metadata.tokenizer.id_to_token(token_id) and if it returns None, return a
sensible fallback string (e.g. synthesize the image token text using the
token_id) instead of erroring, so placeholder_token (and prompt_replacements)
won't panic when TestTokenizer::id_to_token is unimplemented; reference:
placeholder_token, metadata.tokenizer.id_to_token,
ModelRegistryError::TokenNotFound, and TestTokenizer::id_to_token.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Duplicate comments:
In `@multimodal/src/registry.rs`:
- Around line 261-269: The test fails because TestTokenizer::id_to_token always
returns None causing Qwen3VLVisionSpec::placeholder_token to hit
ModelRegistryError::TokenNotFound; fix by making placeholder_token tolerant:
call metadata.tokenizer.id_to_token(token_id) and if it returns None, return a
sensible fallback string (e.g. synthesize the image token text using the
token_id) instead of erroring, so placeholder_token (and prompt_replacements)
won't panic when TestTokenizer::id_to_token is unimplemented; reference:
placeholder_token, metadata.tokenizer.id_to_token,
ModelRegistryError::TokenNotFound, and TestTokenizer::id_to_token.

ℹ️ Review info

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 62b4846 and c6b75ff.

📒 Files selected for processing (1)
  • multimodal/src/registry.rs

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: c6b75ff3a4

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment on lines +99 to +103
match multimodal::process_multimodal(
&request.messages,
model_id,
&*tokenizer,
token_ids,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Add raw-byte fallback when multimodal config files are unavailable

ChatPreparationStage now always calls process_multimodal() before backend selection, which forces loading config.json/preprocessor_config.json from tokenizer_source; this is a regression for deployments where tokenizer sources are model IDs or otherwise not local config directories, because requests that previously worked via raw-image forwarding on vLLM/TRT now fail with multimodal_processing_failed. Please keep a backend-aware fallback (or resolve an actual local config path) instead of hard-failing all multimodal requests when local config files are missing.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not a regression — the preprocessed multimodal pipeline requires config files by design. This PR only adds the preprocessed path for gRPC backends that support it. The raw-image forwarding path (used when config files are unavailable) is unchanged and continues to work as before.

dtype: "float32".to_string(),
}),
model_specific_tensors,
im_token_id: self.im_token_id,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Send the token ID that actually marks Qwen vision embeds

into_vllm_proto() now propagates im_token_id, and downstream placeholder embedding logic depends on that ID, but QwenVLVisionSpec builds replacement spans from vision_token_id while placeholder_token_id() still yields image_token_id; when those differ (as in the Qwen2-VL config used in tests), no token inside mm_placeholders matches im_token_id, so image spans are not marked as embeds and multimodal alignment can break. The emitted im_token_id needs to match the repeated vision token used in replacements.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch — this is a real bug. QwenVLVisionSpec::placeholder_token_id() was returning image_token_id (151655) but the expanded token sequence uses vision_token_id (151654) via pad_token_id(). These are different tokens in Qwen2-VL. Fixed placeholder_token_id() to return pad_token_id() (= vision_token_id) so im_token_id matches the actual pad tokens in the sequence.

Signed-off-by: Chang Su <chang.s.su@oracle.com>
mm_placeholders covers the full structural expansion (image_start, tile
separators, image_end, etc.) but sglang's embedding merge expects offsets
aligned 1:1 with vision encoder output (patch tokens only). This causes
both wrong responses and token count crashes for models like Llama-4 that
have structural tokens in their expansion.

into_sglang_proto() now scans within mm_placeholder ranges for contiguous
runs of im_token_id and emits patch-only offsets. into_vllm_proto() keeps
the full structural offsets since vLLM handles this via the is_embed mask.

Signed-off-by: Chang Su <chang.s.su@oracle.com>
QwenVLVisionSpec::placeholder_token_id() was returning image_token_id
(151655) but the expanded token sequence uses vision_token_id (151654)
via pad_token_id(). These differ in Qwen2-VL, causing im_token_id to
not match actual pad tokens — breaking is_embed on vLLM and
pad_input_tokens on SGLang.

Also fix TestTokenizer::id_to_token() to support reverse lookup, and
add the image_token_id mapping to the Qwen3-VL test tokenizer.

Signed-off-by: Chang Su <chang.s.su@oracle.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@model_gateway/src/routers/grpc/proto_wrapper.rs`:
- Around line 91-116: The loop over self.mm_placeholders may index past
token_ids and panic; update the code in the method that iterates
self.mm_placeholders (the block referencing start/end, token_ids, im_id,
run_start, and pushing sglang::PlaceholderRange into offsets) to defensively
check bounds: skip or clamp ranges where start >= token_ids.len(), set end =
end.min(token_ids.len()), and use token_ids.get(i) (or similar) when reading
token_ids[i] so stale/invalid offsets cannot panic; preserve the existing
run_start/PlaceholderRange logic but ensure you only use valid indices when
creating PlaceholderRange entries.

In `@multimodal/src/registry.rs`:
- Around line 512-520: Ensure the returned aspect_ratios are validated against
the image batch size: when reading
preprocessed.model_specific.get("aspect_ratios") and matching
ModelSpecificValue::IntTensor { data, shape }, in addition to the existing shape
checks (shape.len()==2, shape[1]==2, data.len()==shape[0]*2), also verify that
shape[0] == preprocessed.image_sizes.len() (or otherwise handle mismatch by
error/early return/defaulting). Apply the same cardinality check and handling to
the analogous branch around lines 523–531 so aspect ratios cannot drift from
preprocessed.image_sizes and downstream multimodal payloads stay aligned.

In `@multimodal/src/vision/processors/llama4_vision.rs`:
- Around line 463-468: Add a regression assertion to lock the contract around
patches_per_image: after computing patches_per_image from all_outputs (the
Vec<i64> created from all_outputs.iter().map(|o| o.shape()[0] as i64)), assert
that patches_per_image.len() == batch_size and that
patches_per_image.iter().sum::<i64>() == pixel_values.shape()[0] before
inserting into model_specific; additionally add a unit/integration test that
builds a representative pixel_values and all_outputs and verifies the same two
conditions to prevent future regressions.

ℹ️ Review info

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between c6b75ff and 2713a29.

📒 Files selected for processing (4)
  • model_gateway/src/routers/grpc/client.rs
  • model_gateway/src/routers/grpc/proto_wrapper.rs
  • multimodal/src/registry.rs
  • multimodal/src/vision/processors/llama4_vision.rs

Comment thread model_gateway/src/routers/grpc/proto_wrapper.rs Outdated
Comment thread multimodal/src/registry.rs
Comment thread multimodal/src/vision/processors/llama4_vision.rs

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 2713a2923b

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment on lines +99 to +103
match multimodal::process_multimodal(
&request.messages,
model_id,
&*tokenizer,
token_ids,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Keep backend fallback for non-local tokenizer sources

ChatPreparationStage now runs process_multimodal for every multimodal chat before backend selection, which unconditionally requires loading config.json/preprocessor_config.json from tokenizer_source; this makes vLLM/TRT requests fail when tokenizers were registered from non-filesystem sources (for example, HF model IDs), returning multimodal_config_missing or file-read errors instead of taking the previous raw-image path. Fresh evidence versus the earlier thread: this repo explicitly allows non-path tokenizer sources and stores them verbatim (TokenizerConfigRequest.source in model_gateway/src/core/steps/tokenizer_registration.rs and TokenizerEntry.source assignment in tokenizer/src/registry.rs), so this code path is reachable.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same as the earlier comment — not a regression. See #570 (comment)

Move patch-only offset computation from into_sglang_proto() into
expand_tokens(), piggybacking on the existing iteration over replacement
tokens. This eliminates the separate O(placeholder_tokens) scan and
produces sglang_patch_offsets at zero extra cost.

into_sglang_proto() now uses precomputed offsets via unwrap_or fallback,
and no longer needs the expanded token_ids as input.

Signed-off-by: Chang Su <chang.s.su@oracle.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@model_gateway/src/routers/grpc/multimodal.rs`:
- Around line 364-380: Document and enforce the invariant that
PromptReplacement.tokens are non-negative by adding a brief comment before the
loop and inserting debug_assert!(t >= 0) inside the for loop that iterates
repl.tokens (the loop using repl.tokens.iter().enumerate()) before the cast t as
u32 is performed; this ensures im_token_id comparison (t as u32 == im_id) is
safe and will catch violations in development while leaving release behavior
unchanged.

ℹ️ Review info

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 2713a29 and 02fd148.

📒 Files selected for processing (2)
  • model_gateway/src/routers/grpc/multimodal.rs
  • model_gateway/src/routers/grpc/proto_wrapper.rs

Comment thread model_gateway/src/routers/grpc/multimodal.rs

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 02fd148108

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment on lines +76 to +78
fn field_layouts(&self) -> HashMap<String, FieldLayout> {
// Default: pixel_values is batched (most models).
HashMap::from([("pixel_values".to_string(), FieldLayout::Batched)])

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Override Phi3 field layouts for per-image tensors

field_layouts() now defaults to only pixel_values as batched, while unlisted tensors are documented as shared; for Phi3VisionSpec this means image_sizes and num_img_tokens are sent without per-image slicing in vLLM multimodal requests. Fresh evidence versus the earlier thread: multimodal/src/vision/processors/phi3_vision.rs currently inserts both tensors into model_specific (lines 346-366), so this path is active and multi-image Phi3 prompts can get misaligned metadata.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same as the earlier comment — Phi3 image_sizes and num_img_tokens are struct fields on PreprocessedImages, not model_specific tensors, so field_layouts() does not apply to them. See #570 (comment)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

dependencies Dependency updates grpc gRPC client and router changes model-gateway Model gateway crate changes multimodal Multimodal crate changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant