Repository navigation
fix(hosted_vllm): stop rewriting the caller's video file part into video_url in place - #44977
sungbin1015 wants to merge 1 commit into
Conversation
…deo_url in place hosted_vllm converted a video file part into video_url by assigning into the caller's own content list, so a Router fallback to another provider received a part it does not read and the video was dropped without an error. Build a new content list for the vLLM request instead
|
|
|
| video_file_part = {"type": "file", "file": {"file_id": "https://example.com/video.mp4", "format": "video/mp4"}} | ||
| messages = [{"role": "user", "content": [{"type": "text", "text": "Describe this video"}, video_file_part]}] |
There was a problem hiding this comment.
Required annotations are missing
video_file_part, messages, and route lack the guide’s required : Final annotations. Add them before merging
Context Used: AGENTS.md (source)
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
TLDR
Problem this solves:
hosted_vllmrewrites a videofilepart intovideo_urlinside the caller's own messageHow it solves it:
filepartUser Flow
Before: a team that falls back from a vLLM video model to Gemini gets an answer that never saw the video
hosted_vllmmodel with a text part and a videofilepart (file_idURL orfile_datadata URI)After: the fallback sees the same message the client sent
Pre-Submission checklist
python -m pytest tests/unit/llms/hosted_vllm(91 passed, the new test fails onmain)Screenshots / Proof of Fix
I do not have a Gemini key on this machine, so the two upstreams are local HTTP servers: one on 127.0.0.1:18081 answers every vLLM request with 500, one on 127.0.0.1:18082 answers as Gemini and records the request body it got. The Router is real, with
fallbacks=[{"video-model": ["video-backup"]}]andnum_retries=0Before (a5a86cd)
After (this branch)
router.acompletionwith afile_datadata URI shows the same change: Gemini got only the text part before and getsinline_datawithvideo/mp4afterType
Bug Fix
Caveats (if any)
Medium
Low
HostedVLLMChatConfig().transform_request(messages=...)call still reassignscontenton the message dict it is given, as the assistant branch already does.litellm.completionpasses per-call copies of the message dicts, so callers of the public API are not affectedimage_urlstring to object,filecleanup). Those keep the part type, so other providers still read them, and I left them aloneFinal Attestation
Related work
I found no open or closed PR or issue about
hosted_vllmrewriting the caller'sfilepartmap_openai_paramsin the same file. This PR only touches the user-message branch of_transform_messages, so the two do not overlapcache_controlbeing stripped from the caller's messages in the OpenAI transform: different file, different fieldfunction_call_promptmutating messages: different fileSearched:
hosted_vllm mutates messages,hosted_vllm video_url file,hosted_vllm _transform_messages,HostedVLLMChatConfig,mutates caller messages,messages mutated in place fallback,video_url gemini fallback,fallback drops video, plus open PRs withhosted_vllmorvllmin the title