Skip to content

fix(cost): carry image and video input tokens through the Responses usage bridge (internal copy of #36887) - #41237

Merged
mateo-berri merged 2 commits into
mainfrom
litellm_internal_copy_36887
Sep 19, 2026
Merged

mateo-berri merged 2 commits into
mainfrom
litellm_internal_copy_36887

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

Internal copy of #36887 by @marty-sullivan. Its head 6ab56b8fe5 sits on an org fork that maintainers cannot push to, so the commit is cherry-picked here unchanged with the author's commit preserved, the same route #40915 took to land #37075. #36887 closes once this merges

TLDR

Problem this solves:

  • /v1/responses drops image and video input token counts for Vertex Gemini
  • The reply shows only text_tokens, though Gemini reports every modality
  • Per-modality input rates never get anything to price

How it solves it:

  • Declare image_tokens and video_tokens on input_tokens_details
  • Carry both from the chat usage into the Responses usage
  • Read video_tokens back on the reverse bridge for cost and logs

User Flow

Before: a developer sending an image to Gemini over the Responses API gets only the text token count back, so their per-modality usage tracking undercounts

  1. They send POST https://litellm-domain/v1/responses with "model": "gemini-3.8-flash" and one input_text part plus one input_image part
  2. The 200 reply carries usage.input_tokens: 1093, but usage.input_tokens_details reads {"text_tokens": 13, "cached_tokens": 0}, so 1080 tokens belong to no modality
  3. They send the same route with an input_file part pointing at a video and get usage.input_tokens: 5200 against {"text_tokens": 13, "audio_tokens": 1425}, with the 3762 video tokens missing
  4. They send the same image to POST https://litellm-domain/v1/chat/completions and get usage.prompt_tokens_details: {"text_tokens": 13, "image_tokens": 1080}, so only the Responses route loses the breakdown

After: the same requests come back with every modality counted, and the counts add up to input_tokens

  1. They send POST https://litellm-domain/v1/responses with "model": "gemini-3.8-flash" and one input_text part plus one input_image part
  2. The 200 reply carries usage.input_tokens: 1093 and usage.input_tokens_details reads {"text_tokens": 13, "image_tokens": 1080, "cached_tokens": 0}
  3. The video request comes back with usage.input_tokens: 5200 against {"text_tokens": 13, "audio_tokens": 1425, "video_tokens": 3762}
  4. POST https://litellm-domain/v1/chat/completions is unchanged and still returns {"text_tokens": 13, "image_tokens": 1080}

Relevant issues

Affected release

Linear ticket

Resolves LIT-5844

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Both legs boot the same proxy with no database, two uvicorn workers, no .env in either worktree, and the local cost map, from the same .venv. The Before leg runs the litellm package checked out at the merge base 5fc510a6fd on PYTHONPATH, the After leg runs this branch at b8c2787bcf. Each leg gets its own random free port

LITELLM_LOCAL_MODEL_COST_MAP=True python litellm/proxy/proxy_cli.py --config config.yaml --port $PORT --num_workers 2

config.yaml:

model_list:
  - model_name: gemini-3.8-flash
    litellm_params:
      model: vertex_ai/gemini-3.8-flash
      vertex_project: os.environ/VERTEXAI_PROJECT
      vertex_location: global

general_settings:
  master_key: sk-lit5844

The image and video are Google's public samples, so anyone with Vertex access can rerun this. Payloads:

req_image.json:

{"model":"gemini-3.8-flash","input":[{"role":"user","content":[{"type":"input_text","text":"What is in this image? Answer in at most five words."},{"type":"input_image","image_url":"gs://cloud-samples-data/generative-ai/image/scones.jpg"}]}]}

req_video.json:

{"model":"gemini-3.8-flash","input":[{"role":"user","content":[{"type":"input_text","text":"What happens in this video? Answer in at most five words."},{"type":"input_file","file_url":"gs://cloud-samples-data/generative-ai/video/pixel8.mp4"}]}]}

req_image_stream.json:

{"model":"gemini-3.8-flash","stream":true,"input":[{"role":"user","content":[{"type":"input_text","text":"What is in this image? Answer in at most five words."},{"type":"input_image","image_url":"gs://cloud-samples-data/generative-ai/image/scones.jpg"}]}]}

req_chat_image.json:

{"model":"gemini-3.8-flash","messages":[{"role":"user","content":[{"type":"text","text":"What is in this image? Answer in at most five words."},{"type":"image_url","image_url":{"url":"gs://cloud-samples-data/generative-ai/image/scones.jpg"}}]}]}

req_messages_image.json:

{"model":"gemini-3.8-flash","max_tokens":1024,"messages":[{"role":"user","content":[{"type":"text","text":"What is in this image? Answer in at most five words."},{"type":"image","source":{"type":"url","url":"gs://cloud-samples-data/generative-ai/image/scones.jpg"}}]}]}

Every curl below carries -H "Authorization: Bearer sk-lit5844" -H 'Content-Type: application/json', and the jq filter for the non-streaming Responses cases is

jq -c '{status, text: [.output[] | select(.type=="message") | .content[].text] | join(" "), input_tokens: .usage.input_tokens, input_tokens_details: .usage.input_tokens_details}'

Before (5fc510a)

image part on /v1/responses

  1. curl -s http://localhost:$PORT/v1/responses -d @req_image.json -D image.headers -o image.json && jq -c '<filter above>' image.json

    {"status":"completed","text":"Blueberry scones, coffee, and flowers.","input_tokens":1093,"input_tokens_details":{"audio_tokens":null,"cached_tokens":0,"cached_tokens_details":null,"text_tokens":13}}
    

    input_tokens_details has no image_tokens key at all, so 1080 of the 1093 input tokens belong to no modality

  2. grep -i '^x-litellm-response-cost-input' image.headers

    x-litellm-response-cost-input: 0.0008197500000000001
    

video part on /v1/responses

  1. curl -s http://localhost:$PORT/v1/responses -d @req_video.json -D video.headers -o video.json && jq -c '<filter above>' video.json

    {"status":"completed","text":"Woman films Tokyo at night.","input_tokens":5200,"input_tokens_details":{"audio_tokens":1425,"cached_tokens":0,"cached_tokens_details":null,"text_tokens":13}}
    

    The audio track is counted, the 3762 video tokens are not

  2. grep -i '^x-litellm-response-cost-input' video.headers

    x-litellm-response-cost-input: 0.0028312500000000004
    

image part on /v1/responses with stream: true

  1. curl -s -N http://localhost:$PORT/v1/responses -d @req_image_stream.json -D image_stream.headers -o image_stream.sse && head -1 image_stream.headers && grep '^data:' image_stream.sse | sed 's/^data: //' | grep -v '^\[DONE\]' | jq -c 'select(.type=="response.completed") | {type, input_tokens: .response.usage.input_tokens, input_tokens_details: .response.usage.input_tokens_details}'

    HTTP/1.1 200 OK
    {"type":"response.completed","input_tokens":1093,"input_tokens_details":{"cached_tokens":0,"text_tokens":13}}
    

image part on /v1/chat/completions

  1. curl -s http://localhost:$PORT/v1/chat/completions -d @req_chat_image.json -D chat_image.headers -o chat_image.json && jq -c '{text: .choices[0].message.content, prompt_tokens: .usage.prompt_tokens, prompt_tokens_details: .usage.prompt_tokens_details}' chat_image.json

    {"text":"Blueberry scones, coffee, and peonies.","prompt_tokens":1093,"prompt_tokens_details":{"text_tokens":13,"image_tokens":1080}}
    

    The chat route already reports the image tokens, so the loss is specific to the Responses route

  2. grep -i '^x-litellm-response-cost-input' chat_image.headers

    x-litellm-response-cost-input: 0.0008197500000000001
    

image part on /v1/messages

  1. curl -s http://localhost:$PORT/v1/messages -d @req_messages_image.json -D messages_image.headers -o messages_image.json && jq -c '{text: [.content[] | select(.type=="text") | .text] | join(" "), usage}' messages_image.json

    {"text":"Blueberry scones, coffee, and flowers.","usage":{"input_tokens":1093,"output_tokens":140}}
    
  2. grep -i '^x-litellm-response-cost-input' messages_image.headers

    x-litellm-response-cost-input: 0.0008197500000000001
    

After (b8c2787)

image part on /v1/responses

  1. curl -s http://localhost:$PORT/v1/responses -d @req_image.json -D image.headers -o image.json && jq -c '<filter above>' image.json

    {"status":"completed","text":"Blueberry scones, coffee, and peonies.","input_tokens":1093,"input_tokens_details":{"audio_tokens":null,"cached_tokens":0,"cached_tokens_details":null,"image_tokens":1080,"text_tokens":13,"video_tokens":null}}
    

    13 text plus 1080 image tokens add up to the 1093 input tokens

  2. grep -i '^x-litellm-response-cost-input' image.headers

    x-litellm-response-cost-input: 0.0008197500000000001
    

video part on /v1/responses

  1. curl -s http://localhost:$PORT/v1/responses -d @req_video.json -D video.headers -o video.json && jq -c '<filter above>' video.json

    {"status":"completed","text":"Woman films Tokyo at night.","input_tokens":5200,"input_tokens_details":{"audio_tokens":1425,"cached_tokens":0,"cached_tokens_details":null,"image_tokens":null,"text_tokens":13,"video_tokens":3762}}
    

    13 text plus 1425 audio plus 3762 video tokens add up to the 5200 input tokens

  2. grep -i '^x-litellm-response-cost-input' video.headers

    x-litellm-response-cost-input: 0.0028312500000000004
    

image part on /v1/responses with stream: true

  1. curl -s -N http://localhost:$PORT/v1/responses -d @req_image_stream.json -D image_stream.headers -o image_stream.sse && head -1 image_stream.headers && grep '^data:' image_stream.sse | sed 's/^data: //' | grep -v '^\[DONE\]' | jq -c 'select(.type=="response.completed") | {type, input_tokens: .response.usage.input_tokens, input_tokens_details: .response.usage.input_tokens_details}'

    HTTP/1.1 200 OK
    {"type":"response.completed","input_tokens":1093,"input_tokens_details":{"cached_tokens":0,"image_tokens":1080,"text_tokens":13}}
    

image part on /v1/chat/completions

  1. curl -s http://localhost:$PORT/v1/chat/completions -d @req_chat_image.json -D chat_image.headers -o chat_image.json && jq -c '{text: .choices[0].message.content, prompt_tokens: .usage.prompt_tokens, prompt_tokens_details: .usage.prompt_tokens_details}' chat_image.json

    {"text":"Blueberry scones, coffee, and flowers.","prompt_tokens":1093,"prompt_tokens_details":{"text_tokens":13,"image_tokens":1080}}
    

    Unchanged from Before

  2. grep -i '^x-litellm-response-cost-input' chat_image.headers

    x-litellm-response-cost-input: 0.0008197500000000001
    

image part on /v1/messages

  1. curl -s http://localhost:$PORT/v1/messages -d @req_messages_image.json -D messages_image.headers -o messages_image.json && jq -c '{text: [.content[] | select(.type=="text") | .text] | join(" "), usage}' messages_image.json

    {"text":"Blueberry scones, coffee, and peonies.","usage":{"input_tokens":1093,"output_tokens":282}}
    

    Unchanged from Before, this route does not ride the Responses usage bridge

  2. grep -i '^x-litellm-response-cost-input' messages_image.headers

    x-litellm-response-cost-input: 0.0008197500000000001
    

Observations from the run:

  • Input cost identical on both legs, this model has no per-modality rate
  • Video audio tokens priced at 0 on both legs, pre-existing (LIT-7791)
  • Streaming usage omits null detail keys, non-streaming keeps them, pre-existing
  • /v1/messages usage has no modality breakdown on either leg, pre-existing

Type

🐛 Bug Fix

Caveats (if any)

Low

  • Billed amount unchanged for gemini-3.8-flash, it has no per-image or per-video rate
    • input cost 0.00081975 for the image and 0.00283125 for the video on both legs
  • Only price entries carrying input_cost_per_image_token or input_cost_per_video_token bill differently after this, and today no reachable path meets one: the four Bedrock Nova 2 Pro chat entries carry an image rate but Bedrock reports no modality counts, and the eight Gemini Live realtime entries carry an image rate but the proxy's realtime bridge forwards only input_text and input_audio_buffer.append to Gemini, so no image or video tokens reach that usage. The bill changes the day a Gemini chat entry gains such a rate or the realtime bridge forwards frames, and then it bills the modality at its own rate with text as the remainder, no double count
  • Video audio tokens billed at 0 on both legs, pre-existing, tracked in LIT-7791
  • The realtime Gemini path rides the same forward bridge (gemini/realtime/transformation.py) and was not exercised live; it forwards no image or video frames, its unit tests pass at the tip, and its alias helper drops null keys, so the new fields only appear there once Gemini reports them
  • Non-streaming /v1/responses replies for every provider now carry image_tokens and video_tokens keys, null when unreported, next to the null text_tokens and audio_tokens they already carried; streaming response.completed omits null keys, pre-existing
  • The new unit test drives the reverse bridge through its dict branch, so the object branch's video_tokens line has no unit assertion; every proxy cost calculation reaches it with the usage object, and a test-only commit would rerun bugbot, Greptile, CI, and both QA legs for one assertion, so it stays
  • ci/circleci: litellm_utils_testing is red at this tip on tests/litellm_utils_tests/test_utils.py::test_models_by_provider ('transcribe' missing from models_by_provider), which this PR does not touch: the same test fails the same way at the merge base 5fc510a6fd locally and on main's own scheduled pipelines 89818 (500e880a40) and 89834 (0a792c0f6b), both ancestors of the merge base, since feat(proxy): add Amazon Transcribe pass-through with completion-time job pricing #41515 (784fe5bfd8) priced transcribe/StartTranscriptionJob under a provider nothing registers. fix(proxy): register transcribe as a known provider for model grants #41926 registers the provider and fixes the test on main; this job is not a required check
  • ci/circleci: proxy_e2e_anthropic_messages_tests is red at this tip on both Bedrock cases of test_all_beta_headers.py::test_bedrock_invoke_messages_with_all_beta_headers (Bedrock answers invalid beta flag), which this PR does not touch: the same two cases fail the same way on main's own scheduled pipelines 89818 (500e880a40) and 89834 (0a792c0f6b), both ancestors of the merge base 5fc510a6fd, since fix(anthropic): register thinking-binding-controls-2026-08-01 in beta headers config #41203 (a7a3f6802e, also an ancestor) registered thinking-binding-controls-2026-08-01 for Bedrock and the test sends every mapped beta at Opus 4.5 and Sonnet 4.5, which do not implement it. test(e2e): stop sending model-scoped anthropic betas at models without them #41927 fixes the test on main; this job is not a required check

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Link to Devin session: https://app.devin.ai/sessions/ab05fd2ac04e4ce3aeb46ceceed3269e
Open in Devin Desktop: https://app.devin.ai/desktop/session/ab05fd2ac04e4ce3aeb46ceceed3269e?variant=devin
Requested by: @mateo-berri

…sage bridge

Realtime cost is computed from *_tokens_details after the usage round-trips
through the Responses shape, and the input half of that shape carried audio
only, so image and video prompt tokens stopped being billable as themselves.

Vertex splits prompt tokens by modality, so a session sending camera frames
arrives with image_tokens set. Those were folded into text_tokens and lost
their attribution. The amount happens not to move today, because the
calculator falls back to input_cost_per_token when no per-modality rate is
set, but the tokens have to survive before any such rate can ever apply.

InputTokensDetails now declares image_tokens and video_tokens instead of
leaning on pydantic extras, the repeated per-field copying is a loop over the
modality names so adding a modality no longer adds a branch, and the read-back
in ResponseAPILoggingUtils picks up video_tokens, which
PromptTokensDetailsWrapper already declared.

The output half of the original change is dropped: 449c091 landed the same
OutputTokensDetails.audio_tokens fix upstream, with its own coverage in
test_gemini_realtime_transformation.py, and it always sets
output_tokens_details rather than only when non-empty. That structure is kept
as upstream wrote it.
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR. Add '(aside)' to your comment to have me ignore it.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@codspeed

codspeed Bot commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_internal_copy_36887 (b8c2787) with main (5fc510a)

Open in CodSpeed

@greptile-apps

greptile-apps Bot commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge because the modality fields are carried through both conversion directions without changing aggregate token counts

Summary

Preserves image and video input-token attribution across the Chat Completions and Responses usage bridge, including the reverse conversion used for cost calculation and logging

  • Extends Responses input-token details with image and video fields
  • Copies modality counts into Responses usage
  • Restores video counts when converting Responses usage back to chat usage
  • Adds a focused round-trip regression test

Reviews (2) · Last reviewed commit: "Merge remote-tracking branch 'origin/mai..."

@codecov

codecov Bot commented Sep 15, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri mateo-berri added run-ci and removed run-ci labels Sep 19, 2026
@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit b8c2787. Configure here.

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@mateo-berri
mateo-berri merged commit 825e287 into main Sep 19, 2026
149 of 152 checks passed
@mateo-berri
mateo-berri deleted the litellm_internal_copy_36887 branch September 19, 2026 10:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants