Skip to content

fix(multimodal): fix Phi-3-vision image processing for string-format chat templates - #942

Merged
slin1237 merged 1 commit into
mainfrom
fix/phi3v-multimodal-placeholders
Mar 27, 2026
Merged

slin1237 merged 1 commit into
mainfrom
fix/phi3v-multimodal-placeholders

Conversation

@CatherineSue

@CatherineSue CatherineSue commented Mar 27, 2026 •

Copy link
Copy Markdown
Member

Description

Problem

Phi-3-vision multimodal requests via the OpenAI chat completions API were completely broken — images were either rejected or silently ignored by the model. Four separate issues contributed:

  1. Wrong placeholder token — Phi3VisionSpec used <image> but the Phi-3-vision tokenizer vocabulary uses <|image|> (token id 32044), causing a 400 error on every multimodal request.

  2. Missing placeholder injection — Phi-3-vision's chat template uses string content format (message['content'] as plain text). SMG's transform_content_field stripped all image_url parts for string format, so zero <|image|> tokens appeared in the tokenized prompt. expand_tokens found nothing to replace, and the model received text-only input, hallucinating image descriptions.

  3. Wrong field layout for image_sizes — image_sizes defaulted to SharedField but vLLM's Phi3VForCausalLM._get_mm_fields_config declares it as BatchedField. This caused ValueError: image_sizes dim[0] expected 'bn'=2, got 3 when sending multiple images.

  4. Fixed num_img_tokens — prompt_replacements used the static config value (144, base patch count for 336×336) instead of the actual per-image token count from the HD transform preprocessor. For a 1008×1344 image the real count is 1921, causing a placeholder/embedding size mismatch.

Solution

Mirrors vLLM's approach: resolve the model-specific placeholder token early, thread it through process_chat_messages → transform_content_field, so image parts become placeholder strings instead of being stripped. The multimodal context (model_id, tokenizer_source, placeholder) is computed once and reused for both placeholder injection and process_multimodal, eliminating duplicate lookups.

Changes

  • phi3_v.rs: Fix placeholder token <image> → <|image|>, add image_sizes as Batched field layout, use preprocessed.num_img_tokens instead of fixed config value
  • multimodal.rs: Add resolve_placeholder_token() to look up model placeholder early
  • chat_utils.rs: transform_content_field accepts optional image_placeholder; for string format, image parts become the placeholder string instead of being stripped
  • message_utils.rs: Same change for format_content_parts in the Messages API path
  • chat/preparation.rs, messages/preparation.rs: Resolve multimodal context once (placeholder + model_id + tokenizer_source), pass placeholder to message processing, reuse context for process_multimodal
  • bindings/golang/: Update callers to pass None for the new image_placeholder parameter

Test Plan

Setup:

vllm serve /raid/models/microsoft/Phi-3-vision-128k-instruct \
  --tensor-parallel-size 1 --port 8080 --grpc --trust-remote-code \
  --max-model-len 8192

cargo run --bin smg -- --host 0.0.0.0 --port 3002 --prometheus-port 9321 \
  --worker-urls grpc://127.0.0.1:8080 \
  --model-path /raid/models/microsoft/Phi-3-vision-128k-instruct \
  --log-level debug

Test script:

from openai import OpenAI
client = OpenAI(base_url="http://localhost:3002/v1", api_key="test")
response = client.chat.completions.create(
    model="/raid/models/microsoft/Phi-3-vision-128k-instruct",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "Describe what you see in this image."},
            {"type": "image_url", "image_url": {"url": "https://picsum.photos/id/237/300/200"}},
        ],
    }],
    max_tokens=300,
)
print(response.choices[0].message.content)

Before (main):

# Error 1: placeholder token not found
ERROR smg::routers::grpc::regular::stages::chat::preparation:
  Multimodal processing failed error=Failed to compute prompt replacements:
  token '<image>' not found in tokenizer vocabulary

# Or if that was worked around, model hallucinated:
Choice(content="It doesn't appear that you've sent any images...")

After (this PR):

# Debug log confirms correct token expansion:
Image preprocessing complete num_images=1 total_tokens=1921
Token expansion complete original_len=19 expanded_len=1929
  placeholder_count=1 search_token_id=Some(32044) im_token_id=Some(32044)

# Model correctly describes the image:
Choice(content="The image shows a black Labrador puppy lying on a
  wooden floor, looking directly at the camera with soft brown eyes...")
Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated
  • (Optional) Please join us on Slack #sig-smg to discuss, review, and merge PRs

Summary by CodeRabbit

Release Notes

  • New Features

    • Enhanced multimodal image handling in chat completion messages with placeholder token support, allowing images to be preserved during message processing instead of being stripped.
  • Improvements

    • Updated Phi-3 Vision model support with refined image token representation for improved multimodal request processing.

…ld layouts

Phi-3-vision multimodal requests were failing for four reasons:

1. Wrong placeholder token: The Phi3VisionSpec used `<image>` as the
   placeholder token, but the Phi-3-vision tokenizer uses `<|image|>`
   (token id 32044). This caused a 400 error on every multimodal request.

2. Missing placeholder injection: Phi-3-vision's chat template uses
   string content format, which strips image_url parts during message
   processing. No placeholder tokens were inserted into the text, so
   the tokenized prompt had zero image markers and expand_tokens found
   nothing to replace. This mirrors vLLM's approach of injecting
   model-specific placeholders during content formatting.

   The fix threads an optional image_placeholder through
   process_chat_messages/process_messages -> transform_content_field/
   format_content_parts, resolved early via resolve_placeholder_token().
   Multimodal context (model_id, tokenizer_source, components) is now
   computed once and reused for both placeholder resolution and
   process_multimodal, eliminating duplicate lookups.

3. Wrong field layout for image_sizes: The image_sizes tensor was
   treated as SharedField (default) but vLLM expects it as BatchedField,
   matching its _get_mm_fields_config. This caused a shape validation
   error when processing multiple images.

4. Fixed num_img_tokens: prompt_replacements used the fixed config value
   (144) instead of the actual per-image token count computed by the HD
   transform preprocessor, causing placeholder/embedding size mismatches.

Signed-off-by: Chang Su <chang.s.su@oracle.com>
@github-actions github-actions Bot added grpc gRPC client and router changes model-gateway Model gateway crate changes labels Mar 27, 2026
@coderabbitai

coderabbitai Bot commented Mar 27, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

This PR refactors the multimodal chat message preprocessing pipeline to resolve image placeholders upfront, then thread them through message processing and tokenization instead of stripping image content. It updates three Go bindings to pass a None third argument, modifies Phi-3 vision model spec with new placeholder token format and field layouts, adds a placeholder token resolution helper, and refactors chat/message preparation stages to pre-compute multimodal state before processing.

Changes

Cohort / File(s) Summary
Go Bindings Signature Updates
bindings/golang/src/client.rs, bindings/golang/src/policy.rs, bindings/golang/src/preprocessor.rs
Updated process_chat_messages calls to pass an additional None third argument across three binding entry points.
Phi-3 Vision Model Registry
crates/multimodal/src/registry/phi3_v.rs
Updated placeholder token from "<image>" to `"<
Multimodal Placeholder Resolution
model_gateway/src/routers/grpc/multimodal.rs
Added new crate-visible async helper resolve_placeholder_token() that loads model config, constructs metadata, and returns the placeholder token string or None for unrecognized models.
Chat Preparation Refactoring
model_gateway/src/routers/grpc/regular/stages/chat/preparation.rs
Refactored to pre-resolve multimodal state and image placeholder upfront, pass placeholder into message processing, defer multimodal token expansion to post-tokenization stage gated by resolved mm_context.
Message Preparation Refactoring
model_gateway/src/routers/grpc/regular/stages/messages/preparation.rs
Similar refactoring to pre-resolve multimodal state before message processing, pass image placeholder to message utilities, and conditionally run multimodal expansion after tokenization using cached mm_context.
Message Processing Utilities
model_gateway/src/routers/grpc/utils/chat_utils.rs, model_gateway/src/routers/grpc/utils/message_utils.rs
Updated process_chat_messages and process_messages signatures to accept image_placeholder: Option<&str> parameter; modified content formatting to map image parts to placeholder string when provided, using newline separators instead of spaces.

Sequence Diagram

sequenceDiagram
    participant Client
    participant ChatPrep as Chat Preparation
    participant Registry as Model Registry
    participant MsgProc as Message Processor
    participant Tokenizer
    participant MMProc as Multimodal Processor

    Client->>ChatPrep: Chat request with images
    ChatPrep->>Registry: resolve_placeholder_token()
    Registry-->>ChatPrep: placeholder token or None
    ChatPrep->>MsgProc: process_chat_messages(request, tokenizer, placeholder)
    MsgProc->>MsgProc: map images to placeholder in content
    MsgProc-->>ChatPrep: processed messages with placeholders
    ChatPrep->>Tokenizer: tokenize processed text
    Tokenizer-->>ChatPrep: token ids
    alt multimodal context available
        ChatPrep->>MMProc: process_multimodal(model_id, tokens, images)
        MMProc-->>ChatPrep: expanded token ids + metadata
    end
    ChatPrep-->>Client: prepared request with expanded tokens
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~50 minutes

Possibly related PRs

Suggested labels

multimodal, model-gateway, grpc

Suggested reviewers

  • slin1237
  • key4ng
  • whybeyoung

Poem

🐰 Placeholders bloom where images dwell,
No longer stripped but threaded so well,
Through messages processed with tokens so bright,
Multimodal visions now shine in the light! 🌟

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main change: fixing Phi-3-vision image processing for string-format chat templates, which is the core objective of this multi-file refactoring.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/phi3v-multimodal-placeholders

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request implements model-specific placeholder token injection for multimodal models during chat message processing. It refactors the preparation stages to resolve model-specific placeholders (such as "<|image|>" for Phi-3-vision) early and passes them through to the message transformation utilities. When the chat template format is a string, image parts are now replaced with these placeholders instead of being stripped, and text parts are joined with newlines to ensure correct tokenization for multimodal expansion. Additionally, the Phi-3-vision specification was updated with the correct placeholder token and field layouts. I have no feedback to provide.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
model_gateway/src/routers/grpc/utils/chat_utils.rs (1)

688-876: 🧹 Nitpick | 🔵 Trivial

Add a positive image_placeholder test here.

All updated assertions still pass None, so the Some(image_placeholder) branch that fixes string-format multimodal chats is still unexercised. One mixed text/image case and one image-only case would keep this helper—and the mirrored Messages API helper—protected against future stripping regressions.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@model_gateway/src/routers/grpc/utils/chat_utils.rs` around lines 688 - 876,
Add tests that exercise the Some(image_placeholder) branch of
process_content_format when ChatTemplateContentFormat::String is used: create
one mixed text+image case (mirror test_transform_messages_mixed_content_types)
and one image-only case (mirror test_transform_messages_empty_text_parts) but
call process_content_format with Some("[image]") (or another short placeholder)
and assert that mixed case produces "With image\n[image]" (or placeholder
inserted where images were) and image-only case returns the placeholder string
instead of an array; reference process_content_format,
ChatTemplateContentFormat::String, and the new test names like
test_transform_messages_string_format_with_image_placeholder and
test_transform_messages_image_only_with_placeholder to locate where to add them.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@model_gateway/src/routers/grpc/regular/stages/chat/preparation.rs`:
- Around line 79-87: The call to multimodal::resolve_placeholder_token currently
uses .await.ok().flatten(), which hides real errors by converting Err into None;
change it to preserve and propagate errors instead (e.g., await the Result and
use the ? operator or an explicit match) so that resolve_placeholder_token's Err
is returned immediately rather than being treated like Ok(None). Update the
assignment to let placeholder =
multimodal::resolve_placeholder_token(...).await? (or equivalent match that
returns Err) so failures fail fast while keeping Ok(None) semantics for
genuinely absent placeholders.

In `@model_gateway/src/routers/grpc/regular/stages/messages/preparation.rs`:
- Around line 99-112: The code is currently swallowing errors by calling
.await.ok().flatten() on multimodal::resolve_placeholder_token which hides both
Err and Ok(None); instead, await the result and match it, and if it is Err(_) or
Ok(None) return a 400 Bad Request immediately (don’t continue to tokenization)
so that multimodal failures are surfaced; update the block that assigns
placeholder (and the tuple (placeholder, Some((mm_components, model_id,
tokenizer_source)))) to perform this check and return the appropriate 400
response from the enclosing function when resolution fails.

---

Outside diff comments:
In `@model_gateway/src/routers/grpc/utils/chat_utils.rs`:
- Around line 688-876: Add tests that exercise the Some(image_placeholder)
branch of process_content_format when ChatTemplateContentFormat::String is used:
create one mixed text+image case (mirror
test_transform_messages_mixed_content_types) and one image-only case (mirror
test_transform_messages_empty_text_parts) but call process_content_format with
Some("[image]") (or another short placeholder) and assert that mixed case
produces "With image\n[image]" (or placeholder inserted where images were) and
image-only case returns the placeholder string instead of an array; reference
process_content_format, ChatTemplateContentFormat::String, and the new test
names like test_transform_messages_string_format_with_image_placeholder and
test_transform_messages_image_only_with_placeholder to locate where to add them.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 6d539062-1ebc-4a5d-9137-02d9e7d8baa8

📥 Commits

Reviewing files that changed from the base of the PR and between d582379 and 6dd4dc8.

📒 Files selected for processing (9)
  • bindings/golang/src/client.rs
  • bindings/golang/src/policy.rs
  • bindings/golang/src/preprocessor.rs
  • crates/multimodal/src/registry/phi3_v.rs
  • model_gateway/src/routers/grpc/multimodal.rs
  • model_gateway/src/routers/grpc/regular/stages/chat/preparation.rs
  • model_gateway/src/routers/grpc/regular/stages/messages/preparation.rs
  • model_gateway/src/routers/grpc/utils/chat_utils.rs
  • model_gateway/src/routers/grpc/utils/message_utils.rs

Comment thread model_gateway/src/routers/grpc/regular/stages/chat/preparation.rs
@slin1237
slin1237 merged commit c147d0c into main Mar 27, 2026
38 checks passed
@slin1237
slin1237 deleted the fix/phi3v-multimodal-placeholders branch March 27, 2026 14:58
CatherineSue added a commit that referenced this pull request Mar 27, 2026
…wallowing

resolve_placeholder_token errors were silently converted to None via
.ok().flatten(), making real failures (e.g. corrupt config.json)
indistinguishable from Ok(None). This let requests proceed without
image placeholders, deferring the failure to a vague downstream
mismatch instead of a clear error.

Now propagate the Err as a 500 Internal Server Error immediately,
since this is a gateway-level preparation failure (config loading,
spec lookup), not a backend/worker error. Ok(None) (model not
recognized as multimodal) still proceeds without placeholders,
allowing text-only fallback for unknown models.

Addresses review feedback from #942.

Signed-off-by: Chang Su <chang.s.su@oracle.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

grpc gRPC client and router changes model-gateway Model gateway crate changes multimodal Multimodal crate changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants