Skip to content

fix(multimodal): extract images from tool role messages - #1307

Merged
CatherineSue merged 1 commit into
mainfrom
connorli/fix-multimodal-tool-images
Apr 22, 2026
Merged

CatherineSue merged 1 commit into
mainfrom
connorli/fix-multimodal-tool-images

Conversation

@ConnorLi96

@ConnorLi96 ConnorLi96 commented Apr 22, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

has_multimodal_content and extract_content_parts only match User/System/Developer messages, silently dropping images in Tool role messages. This causes:

  • Tool-only image: model sees <|media_pad|> placeholder but receives no pixel data → replies "I don't see any image"
  • User + tool images: 2 placeholders vs 1 extracted image → TRT-LLM 400 "More media placeholder tokens (2) than media items (1)"

Fix

Add ChatMessage::Tool { content, .. } => Some(content) to both match arms. Tool.content is MessageContent — same type as User/System/Developer. 2-line change.

Note

OpenAI Chat Completions spec limits tool content to text-only, but their newer Responses API supports images in tool output. SMG's protocol layer already deserializes image parts in tool messages, and Kimi's chat template generates placeholders for all roles — this closes the gap in the extraction layer.

Verified

E2E on Kimi K2.5 NVFP4 TP=4 TRT-LLM:

Request Before After
Tool-only image "I don't see any screenshot" extracted 1 multimodal images
User + tool images Only 1 image extracted; prod 400 extracted 2 multimodal images
User-only image (control) Works Still works

Summary by CodeRabbit

  • Bug Fixes
    • Enhanced multimodal content detection and extraction for OpenAI ChatMessage Tool variants, ensuring images and text content within Tool messages are properly identified and processed.

has_multimodal_content and extract_content_parts only matched
User/System/Developer, silently dropping images in Tool messages.
This caused placeholder-vs-image count mismatches on backends
(TRT-LLM 400) or invisible images (model replies "I don't see").

Signed-off-by: Connor Li <ConnorLi96@users.noreply.github.com>
Made-with: Cursor
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

@github-actions github-actions Bot added grpc gRPC client and router changes model-gateway Model gateway crate changes labels Apr 22, 2026
@coderabbitai

coderabbitai Bot commented Apr 22, 2026 •

Copy link
Copy Markdown

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 24147d7a-4261-4276-81e5-eff5a38e55ad

📥 Commits

Reviewing files that changed from the base of the PR and between 97dc848 and 35f71ed.

📒 Files selected for processing (1)
  • model_gateway/src/routers/grpc/multimodal.rs

📝 Walkthrough

Walkthrough

Extended multimodal content detection and extraction in the OpenAI ChatMessage handler to process Tool message variants. The changes enable identification of image URLs within Tool message content and conversion of Tool message content into appropriate media content parts.

Changes

Cohort / File(s) Summary
Multimodal Tool Support
model_gateway/src/routers/grpc/multimodal.rs
Extended has_multimodal_content and extract_content_parts functions to handle ChatMessage::Tool variants, scanning for ContentPart::ImageUrl and converting Tool content to MediaContentParts.

Estimated code review effort

🎯 1 (Trivial) | ⏱️ ~3 minutes

Possibly related PRs

Suggested labels

grpc, model-gateway, multimodal

Suggested reviewers

  • key4ng
  • slin1237

Poem

🐰 A tool message arrives with images bright,
Now detected and extracted with might!
Two tiny lines, yet oh so fine,
Make multimodal content truly shine! ✨

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and specifically describes the main change: adding image extraction support for tool role messages in multimodal handling.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch connorli/fix-multimodal-tool-images

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Clean 2-line fix. ChatMessage::Tool has the same content: MessageContent type as the other variants already matched, so the new arms are type-safe and consistent. The failure modes described in the PR (placeholder/image count mismatch, missing pixel data) are real and this closes the gap correctly.

@CatherineSue
CatherineSue merged commit a92cfb3 into main Apr 22, 2026
49 checks passed
@CatherineSue
CatherineSue deleted the connorli/fix-multimodal-tool-images branch April 22, 2026 17:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

grpc gRPC client and router changes model-gateway Model gateway crate changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants