Repository navigation
fix(cost): carry image and video input tokens through the Responses usage bridge (internal copy of #36887) - #41237
Conversation
…sage bridge Realtime cost is computed from *_tokens_details after the usage round-trips through the Responses shape, and the input half of that shape carried audio only, so image and video prompt tokens stopped being billable as themselves. Vertex splits prompt tokens by modality, so a session sending camera frames arrives with image_tokens set. Those were folded into text_tokens and lost their attribution. The amount happens not to move today, because the calculator falls back to input_cost_per_token when no per-modality rate is set, but the tokens have to survive before any such rate can ever apply. InputTokensDetails now declares image_tokens and video_tokens instead of leaning on pydantic extras, the repeated per-field copying is a loop over the modality names so adding a modality no longer adds a branch, and the read-back in ResponseAPILoggingUtils picks up video_tokens, which PromptTokensDetailsWrapper already declared. The output half of the original change is dropped: 449c091 landed the same OutputTokensDetails.audio_tokens fix upstream, with its own coverage in test_gemini_realtime_transformation.py, and it always sets output_tokens_details rather than only when non-empty. That structure is kept as upstream wrote it.
🤖 Devin AI EngineerI'll be helping with this pull request! Here's what you should know: ✅ I will automatically:
Note: I can only respond to comments from users who have write access to this repository. ⚙️ Control Options:
|
|
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit b8c2787. Configure here.
Internal copy of #36887 by @marty-sullivan. Its head
6ab56b8fe5sits on an org fork that maintainers cannot push to, so the commit is cherry-picked here unchanged with the author's commit preserved, the same route #40915 took to land #37075. #36887 closes once this mergesTLDR
Problem this solves:
/v1/responsesdrops image and video input token counts for Vertex Geminitext_tokens, though Gemini reports every modalityHow it solves it:
image_tokensandvideo_tokensoninput_tokens_detailsvideo_tokensback on the reverse bridge for cost and logsUser Flow
Before: a developer sending an image to Gemini over the Responses API gets only the text token count back, so their per-modality usage tracking undercounts
"model": "gemini-3.8-flash"and oneinput_textpart plus oneinput_imagepartusage.input_tokens: 1093, butusage.input_tokens_detailsreads{"text_tokens": 13, "cached_tokens": 0}, so 1080 tokens belong to no modalityinput_filepart pointing at a video and getusage.input_tokens: 5200against{"text_tokens": 13, "audio_tokens": 1425}, with the 3762 video tokens missingusage.prompt_tokens_details: {"text_tokens": 13, "image_tokens": 1080}, so only the Responses route loses the breakdownAfter: the same requests come back with every modality counted, and the counts add up to
input_tokens"model": "gemini-3.8-flash"and oneinput_textpart plus oneinput_imagepartusage.input_tokens: 1093andusage.input_tokens_detailsreads{"text_tokens": 13, "image_tokens": 1080, "cached_tokens": 0}usage.input_tokens: 5200against{"text_tokens": 13, "audio_tokens": 1425, "video_tokens": 3762}{"text_tokens": 13, "image_tokens": 1080}Relevant issues
Affected release
Linear ticket
Resolves LIT-5844
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Both legs boot the same proxy with no database, two uvicorn workers, no
.envin either worktree, and the local cost map, from the same.venv. The Before leg runs the litellm package checked out at the merge base5fc510a6fdonPYTHONPATH, the After leg runs this branch atb8c2787bcf. Each leg gets its own random free portconfig.yaml:
The image and video are Google's public samples, so anyone with Vertex access can rerun this. Payloads:
req_image.json:
{"model":"gemini-3.8-flash","input":[{"role":"user","content":[{"type":"input_text","text":"What is in this image? Answer in at most five words."},{"type":"input_image","image_url":"gs://cloud-samples-data/generative-ai/image/scones.jpg"}]}]}req_video.json:
{"model":"gemini-3.8-flash","input":[{"role":"user","content":[{"type":"input_text","text":"What happens in this video? Answer in at most five words."},{"type":"input_file","file_url":"gs://cloud-samples-data/generative-ai/video/pixel8.mp4"}]}]}req_image_stream.json:
{"model":"gemini-3.8-flash","stream":true,"input":[{"role":"user","content":[{"type":"input_text","text":"What is in this image? Answer in at most five words."},{"type":"input_image","image_url":"gs://cloud-samples-data/generative-ai/image/scones.jpg"}]}]}req_chat_image.json:
{"model":"gemini-3.8-flash","messages":[{"role":"user","content":[{"type":"text","text":"What is in this image? Answer in at most five words."},{"type":"image_url","image_url":{"url":"gs://cloud-samples-data/generative-ai/image/scones.jpg"}}]}]}req_messages_image.json:
{"model":"gemini-3.8-flash","max_tokens":1024,"messages":[{"role":"user","content":[{"type":"text","text":"What is in this image? Answer in at most five words."},{"type":"image","source":{"type":"url","url":"gs://cloud-samples-data/generative-ai/image/scones.jpg"}}]}]}Every curl below carries
-H "Authorization: Bearer sk-lit5844" -H 'Content-Type: application/json', and the jq filter for the non-streaming Responses cases isBefore (5fc510a)
image part on /v1/responses
curl -s http://localhost:$PORT/v1/responses -d @req_image.json -D image.headers -o image.json && jq -c '<filter above>' image.jsoninput_tokens_detailshas noimage_tokenskey at all, so 1080 of the 1093 input tokens belong to no modalitygrep -i '^x-litellm-response-cost-input' image.headersvideo part on /v1/responses
curl -s http://localhost:$PORT/v1/responses -d @req_video.json -D video.headers -o video.json && jq -c '<filter above>' video.jsonThe audio track is counted, the 3762 video tokens are not
grep -i '^x-litellm-response-cost-input' video.headersimage part on /v1/responses with stream: true
curl -s -N http://localhost:$PORT/v1/responses -d @req_image_stream.json -D image_stream.headers -o image_stream.sse && head -1 image_stream.headers && grep '^data:' image_stream.sse | sed 's/^data: //' | grep -v '^\[DONE\]' | jq -c 'select(.type=="response.completed") | {type, input_tokens: .response.usage.input_tokens, input_tokens_details: .response.usage.input_tokens_details}'image part on /v1/chat/completions
curl -s http://localhost:$PORT/v1/chat/completions -d @req_chat_image.json -D chat_image.headers -o chat_image.json && jq -c '{text: .choices[0].message.content, prompt_tokens: .usage.prompt_tokens, prompt_tokens_details: .usage.prompt_tokens_details}' chat_image.jsonThe chat route already reports the image tokens, so the loss is specific to the Responses route
grep -i '^x-litellm-response-cost-input' chat_image.headersimage part on /v1/messages
curl -s http://localhost:$PORT/v1/messages -d @req_messages_image.json -D messages_image.headers -o messages_image.json && jq -c '{text: [.content[] | select(.type=="text") | .text] | join(" "), usage}' messages_image.jsongrep -i '^x-litellm-response-cost-input' messages_image.headersAfter (b8c2787)
image part on /v1/responses
curl -s http://localhost:$PORT/v1/responses -d @req_image.json -D image.headers -o image.json && jq -c '<filter above>' image.json13 text plus 1080 image tokens add up to the 1093 input tokens
grep -i '^x-litellm-response-cost-input' image.headersvideo part on /v1/responses
curl -s http://localhost:$PORT/v1/responses -d @req_video.json -D video.headers -o video.json && jq -c '<filter above>' video.json13 text plus 1425 audio plus 3762 video tokens add up to the 5200 input tokens
grep -i '^x-litellm-response-cost-input' video.headersimage part on /v1/responses with stream: true
curl -s -N http://localhost:$PORT/v1/responses -d @req_image_stream.json -D image_stream.headers -o image_stream.sse && head -1 image_stream.headers && grep '^data:' image_stream.sse | sed 's/^data: //' | grep -v '^\[DONE\]' | jq -c 'select(.type=="response.completed") | {type, input_tokens: .response.usage.input_tokens, input_tokens_details: .response.usage.input_tokens_details}'image part on /v1/chat/completions
curl -s http://localhost:$PORT/v1/chat/completions -d @req_chat_image.json -D chat_image.headers -o chat_image.json && jq -c '{text: .choices[0].message.content, prompt_tokens: .usage.prompt_tokens, prompt_tokens_details: .usage.prompt_tokens_details}' chat_image.jsonUnchanged from Before
grep -i '^x-litellm-response-cost-input' chat_image.headersimage part on /v1/messages
curl -s http://localhost:$PORT/v1/messages -d @req_messages_image.json -D messages_image.headers -o messages_image.json && jq -c '{text: [.content[] | select(.type=="text") | .text] | join(" "), usage}' messages_image.jsonUnchanged from Before, this route does not ride the Responses usage bridge
grep -i '^x-litellm-response-cost-input' messages_image.headersObservations from the run:
/v1/messagesusage has no modality breakdown on either leg, pre-existingType
🐛 Bug Fix
Caveats (if any)
Low
gemini-3.8-flash, it has no per-image or per-video rateinput_cost_per_image_tokenorinput_cost_per_video_tokenbill differently after this, and today no reachable path meets one: the four Bedrock Nova 2 Pro chat entries carry an image rate but Bedrock reports no modality counts, and the eight Gemini Live realtime entries carry an image rate but the proxy's realtime bridge forwards onlyinput_textandinput_audio_buffer.appendto Gemini, so no image or video tokens reach that usage. The bill changes the day a Gemini chat entry gains such a rate or the realtime bridge forwards frames, and then it bills the modality at its own rate with text as the remainder, no double countgemini/realtime/transformation.py) and was not exercised live; it forwards no image or video frames, its unit tests pass at the tip, and its alias helper drops null keys, so the new fields only appear there once Gemini reports them/v1/responsesreplies for every provider now carryimage_tokensandvideo_tokenskeys, null when unreported, next to the nulltext_tokensandaudio_tokensthey already carried; streamingresponse.completedomits null keys, pre-existingvideo_tokensline has no unit assertion; every proxy cost calculation reaches it with the usage object, and a test-only commit would rerun bugbot, Greptile, CI, and both QA legs for one assertion, so it staysci/circleci: litellm_utils_testingis red at this tip ontests/litellm_utils_tests/test_utils.py::test_models_by_provider('transcribe'missing frommodels_by_provider), which this PR does not touch: the same test fails the same way at the merge base5fc510a6fdlocally and on main's own scheduled pipelines 89818 (500e880a40) and 89834 (0a792c0f6b), both ancestors of the merge base, since feat(proxy): add Amazon Transcribe pass-through with completion-time job pricing #41515 (784fe5bfd8) pricedtranscribe/StartTranscriptionJobunder a provider nothing registers. fix(proxy): register transcribe as a known provider for model grants #41926 registers the provider and fixes the test on main; this job is not a required checkci/circleci: proxy_e2e_anthropic_messages_testsis red at this tip on both Bedrock cases oftest_all_beta_headers.py::test_bedrock_invoke_messages_with_all_beta_headers(Bedrock answersinvalid beta flag), which this PR does not touch: the same two cases fail the same way on main's own scheduled pipelines 89818 (500e880a40) and 89834 (0a792c0f6b), both ancestors of the merge base5fc510a6fd, since fix(anthropic): register thinking-binding-controls-2026-08-01 in beta headers config #41203 (a7a3f6802e, also an ancestor) registeredthinking-binding-controls-2026-08-01for Bedrock and the test sends every mapped beta at Opus 4.5 and Sonnet 4.5, which do not implement it. test(e2e): stop sending model-scoped anthropic betas at models without them #41927 fixes the test on main; this job is not a required checkFinal Attestation
Link to Devin session: https://app.devin.ai/sessions/ab05fd2ac04e4ce3aeb46ceceed3269e
Open in Devin Desktop: https://app.devin.ai/desktop/session/ab05fd2ac04e4ce3aeb46ceceed3269e?variant=devin
Requested by: @mateo-berri