fix(multimodal): extract images from tool role messages - #1307
Conversation
has_multimodal_content and extract_content_parts only matched User/System/Developer, silently dropping images in Tool messages. This caused placeholder-vs-image count mismatches on backends (TRT-LLM 400) or invisible images (model replies "I don't see"). Signed-off-by: Connor Li <ConnorLi96@users.noreply.github.com> Made-with: Cursor
|
Warning You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again! |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (1)
📝 WalkthroughWalkthroughExtended multimodal content detection and extraction in the OpenAI ChatMessage handler to process Changes
Estimated code review effort🎯 1 (Trivial) | ⏱️ ~3 minutes Possibly related PRs
Suggested labels
Suggested reviewers
Poem
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Clean 2-line fix. ChatMessage::Tool has the same content: MessageContent type as the other variants already matched, so the new arms are type-safe and consistent. The failure modes described in the PR (placeholder/image count mismatch, missing pixel data) are real and this closes the gap correctly.
Problem
has_multimodal_contentandextract_content_partsonly matchUser/System/Developermessages, silently dropping images inToolrole messages. This causes:<|media_pad|>placeholder but receives no pixel data → replies "I don't see any image""More media placeholder tokens (2) than media items (1)"Fix
Add
ChatMessage::Tool { content, .. } => Some(content)to both match arms.Tool.contentisMessageContent— same type as User/System/Developer. 2-line change.Note
OpenAI Chat Completions spec limits tool content to text-only, but their newer Responses API supports images in tool output. SMG's protocol layer already deserializes image parts in tool messages, and Kimi's chat template generates placeholders for all roles — this closes the gap in the extraction layer.
Verified
E2E on Kimi K2.5 NVFP4 TP=4 TRT-LLM:
extracted 1 multimodal imagesextracted 2 multimodal imagesSummary by CodeRabbit