Skip to content

fix(hosted_vllm): stop rewriting the caller's video file part into video_url in place - #44977

Open
sungbin1015 wants to merge 1 commit into
BerriAI:mainfrom
sungbin1015:bugfix/hosted-vllm-video-file-in-place
Open

sungbin1015 wants to merge 1 commit into
BerriAI:mainfrom
sungbin1015:bugfix/hosted-vllm-video-file-in-place

Conversation

@sungbin1015

Copy link
Copy Markdown

TLDR

Problem this solves:

  • hosted_vllm rewrites a video file part into video_url inside the caller's own message
  • A fallback to another provider then gets a part it does not read
  • Gemini silently drops the video and answers from the text alone

How it solves it:

  • Build a new content list for the vLLM request
  • The caller's message keeps its file part

User Flow

Before: a team that falls back from a vLLM video model to Gemini gets an answer that never saw the video

  1. They send POST https://litellm-domain/v1/chat/completions to a hosted_vllm model with a text part and a video file part (file_id URL or file_data data URI)
  2. The vLLM server answers 500, so the configured fallback to a Gemini model runs
  3. Gemini receives only the text part. The response is 200 and nothing says the video was left out

After: the fallback sees the same message the client sent

  1. They send the same POST https://litellm-domain/v1/chat/completions
  2. The vLLM server answers 500, so the fallback to the Gemini model runs
  3. Gemini receives the text part and the video, and the 200 response describes the video

Pre-Submission checklist

  • I have added meaningful tests
  • The handful of test files covering my change pass locally: python -m pytest tests/unit/llms/hosted_vllm (91 passed, the new test fails on main)
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review

Screenshots / Proof of Fix

I do not have a Gemini key on this machine, so the two upstreams are local HTTP servers: one on 127.0.0.1:18081 answers every vLLM request with 500, one on 127.0.0.1:18082 answers as Gemini and records the request body it got. The Router is real, with fallbacks=[{"video-model": ["video-backup"]}] and num_retries=0

router = litellm.Router(
    model_list=[
        {"model_name": "video-model", "litellm_params": {"model": "hosted_vllm/qwen-vl", "api_base": "http://127.0.0.1:18081/v1"}},
        {"model_name": "video-backup", "litellm_params": {"model": "gemini/gemini-2.5-flash", "api_key": "none", "api_base": "http://127.0.0.1:18082"}},
    ],
    fallbacks=[{"video-model": ["video-backup"]}],
    num_retries=0,
)
messages = [{"role": "user", "content": [
    {"type": "text", "text": "Describe this video"},
    {"type": "file", "file": {"file_id": "https://example.com/video.mp4", "format": "video/mp4"}},
]}]
router.completion(model="video-model", messages=messages)

Before (a5a86cd)

  1. Run the script above and print what each upstream received
  2. Output:
vLLM received parts:   [{"type": "text", "text": "Describe this video"}, {"type": "video_url", "video_url": {"url": "https://example.com/video.mp4"}}]
Gemini received parts: [{"text": "Describe this video"}]
caller's messages now: [{"type": "text", "text": "Describe this video"}, {"type": "video_url", "video_url": {"url": "https://example.com/video.mp4"}}]

After (this branch)

  1. Same script
  2. Output:
vLLM received parts:   [{"type": "text", "text": "Describe this video"}, {"type": "video_url", "video_url": {"url": "https://example.com/video.mp4"}}]
Gemini received parts: [{"text": "Describe this video"}, {"file_data": {"mime_type": "video/mp4", "file_uri": "https://example.com/video.mp4"}}]
caller's messages now: [{"type": "text", "text": "Describe this video"}, {"type": "file", "file": {"file_id": "https://example.com/video.mp4", "format": "video/mp4"}}]

router.acompletion with a file_data data URI shows the same change: Gemini got only the text part before and gets inline_data with video/mp4 after

Type

Bug Fix

Caveats (if any)

Medium

  • Proof uses local stand-in upstreams, not a live vLLM server or the real Gemini API

Low

  • A direct HostedVLLMChatConfig().transform_request(messages=...) call still reassigns content on the message dict it is given, as the assistant branch already does. litellm.completion passes per-call copies of the message dicts, so callers of the public API are not affected
  • The base OpenAI transform still normalizes parts of the caller's list in place (image_url string to object, file cleanup). Those keep the part type, so other providers still read them, and I left them alone

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Related work

I found no open or closed PR or issue about hosted_vllm rewriting the caller's file part

Searched: hosted_vllm mutates messages, hosted_vllm video_url file, hosted_vllm _transform_messages, HostedVLLMChatConfig, mutates caller messages, messages mutated in place fallback, video_url gemini fallback, fallback drops video, plus open PRs with hosted_vllm or vllm in the title

…deo_url in place

hosted_vllm converted a video file part into video_url by assigning into the caller's own content list, so a Router fallback to another provider received a part it does not read and the video was dropped without an error. Build a new content list for the vLLM request instead
@CLAassistant

CLAassistant commented Oct 7, 2026 •

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@greptile-apps

greptile-apps Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

[Medium risk] Fixes in-place mutation of caller's message objects.

The caller-preservation fix looks sound, but the required Final annotations must be added before merging

Findings

  1. P2 Required annotations are missing ▶

Summary

hosted_vllm now builds a separate user content list when converting video files to video_url. This keeps the caller’s video file part available for later calls

  • Adds a mocked public-call regression test that checks the outgoing video part and the caller’s original part
  • The new test needs the required Final annotations
  • sungbin1015 explicitly acknowledged direct config message reassignment and existing OpenAI part cleanup as outside this fix. Neither is reported
  • sungbin1015 explicitly acknowledged that the supplied proof uses local stand-in providers rather than live APIs. This limitation is not reported

Reviews (1) · Last reviewed commit: "fix(hosted_vllm): stop rewriting the cal..."

Comment on lines +49 to +50
video_file_part = {"type": "file", "file": {"file_id": "https://example.com/video.mp4", "format": "video/mp4"}}
messages = [{"role": "user", "content": [{"type": "text", "text": "Describe this video"}, video_file_part]}]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Required annotations are missing

video_file_part, messages, and route lack the guide’s required : Final annotations. Add them before merging

Context Used: AGENTS.md (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

@codspeed

codspeed Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing sungbin1015:bugfix/hosted-vllm-video-file-in-place (6c37fbe) with main (a5a86cd)

Open in CodSpeed

@codecov

codecov Bot commented Oct 7, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants