Skip to content

fix(cost): bill uncached realtime tokens per modality using cached_tokens_details - #38858

Closed
devin-ai-integration[bot] wants to merge 3 commits into
litellm_internal_stagingfrom
litellm_fix_realtime_cached_audio_billing
Closed

devin-ai-integration[bot] wants to merge 3 commits into
litellm_internal_stagingfrom
litellm_fix_realtime_cached_audio_billing

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Aug 30, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Realtime sessions with cached audio input over-bill uncached tokens
  • Uncached budget is filled audio-first, so uncached text bills at audio rate
  • Provider already reports which modalities the cache covered

How it solves it:

  • Parse cached_tokens_details from realtime usage into the spend pipeline
  • Subtract the cache from each modality before pricing uncached tokens
  • Fall back to the old audio-first split when the provider omits the breakdown

User Flow

Before: a developer running realtime voice sessions through the gateway sees spend that is higher than what OpenAI bills them

  1. They connect to wss://litellm-domain/v1/realtime?model=gpt-realtime and hold a voice conversation whose prompt is mostly cached audio
  2. The session's response.done reports 3400 input tokens (1400 text, 2000 audio, 2000 cached of which 1500 audio and 500 text) and 1100 output tokens
  3. They open https://litellm-domain/ui/?page=logs and the session shows $0.1112: all 1400 uncached tokens were priced as audio even though 900 of them were text

After: the same session logs the price OpenAI actually charges

  1. They connect to wss://litellm-domain/v1/realtime?model=gpt-realtime and hold the same conversation
  2. The session's response.done reports the same usage
  3. https://litellm-domain/ui/?page=logs now shows $0.0860: 500 uncached audio at the audio rate, 900 uncached text at the text rate, 2000 cached at the cache rate

Relevant issues

Linear ticket

Resolves LIT-6513

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Live proxy run on localhost:4400 with LITELLM_LOCAL_MODEL_COST_MAP=True, a spend-logging callback writing to a jsonl file, and a realtime backend replaying a fixed GA response.done usage (3400 input: 1400 text + 2000 audio, 2000 cached with cached_tokens_details of 1500 audio + 500 text; 1100 output: 100 text + 1000 audio) so the token counts are identical across both runs. gpt-realtime rates: text $4/M, audio $32/M, cached $0.40/M, text out $16/M, audio out $64/M. Correct price by hand: 500 audio x $32/M + 900 text x $4/M + 2000 cached x $0.40/M + 1000 audio out x $64/M + 100 text out x $16/M = $0.0860

Before (d44d281)

  1. Start the proxy: LITELLM_LOCAL_MODEL_COST_MAP=True uv run python litellm/proxy/proxy_cli.py --config /tmp/rt_repro/config.yaml --port 4400
  2. Run a realtime session: connect to ws://127.0.0.1:4400/v1/realtime?model=gpt-realtime, send {"type": "response.create"}, receive response.done, disconnect
  3. cat /tmp/rt_repro/spend.jsonl shows {"call_type": "_arealtime", "model": "gpt-realtime", "response_cost": 0.11120000000000001, "prompt_tokens": 3400, "completion_tokens": 1100, "spend": 0.11120000000000001}: all 1400 uncached tokens billed as audio ($0.0448 instead of $0.0196)

After (a42f152)

  1. Start the proxy the same way
  2. Run the same realtime session
  3. cat /tmp/rt_repro/spend.jsonl shows {"call_type": "_arealtime", "model": "gpt-realtime", "response_cost": 0.08600000000000001, "prompt_tokens": 3400, "completion_tokens": 1100, "spend": 0.08600000000000001}, matching the hand-computed $0.0860

Type

🐛 Bug Fix

Caveats (if any)

Low

  • Providers that omit cached_tokens_details keep the old audio-first fallback

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Link to Devin session: https://app.devin.ai/sessions/d8222dd829134327a93215f75c4500f9
Open in Devin Desktop: https://app.devin.ai/desktop/session/d8222dd829134327a93215f75c4500f9?variant=devin

…s modality split

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR. Add '(aside)' to your comment to have me ignore it.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

PR #38858 lacks the enterprise label (labels are empty), so no labeling or Linear routing was performed. No changes made.

@codspeed

codspeed Bot commented Aug 30, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_fix_realtime_cached_audio_billing (a42f152) with litellm_internal_staging (d44d281)

Open in CodSpeed

@greptile-apps

greptile-apps Bot commented Aug 30, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR carries realtime cached-token modality details through usage normalization and aggregation, then uses them to price uncached input by modality

  • Adds a typed cached-token modality breakdown to normalized prompt usage
  • Aggregates the breakdown across realtime response completion events
  • Preserves the previous audio-first allocation when the provider omits the breakdown
  • Adds regression coverage for modality-aware pricing and aggregation

Confidence Score: 4/5

The pricing fix appears safe to merge after the non-blocking Python guideline violations are cleaned up

The cached-token metadata follows the realtime normalization, aggregation, and cost-allocation path, with focused regression coverage; the remaining concern is code-quality compliance

Files Needing Attention: litellm/cost_calculator.py, litellm/litellm_core_utils/llm_cost_calc/utils.py, litellm/responses/utils.py

Important Files Changed

Filename Overview
litellm/cost_calculator.py Aggregates cached-token modality details across realtime usage objects, but the helper mutates its parameter and includes an overlong call
litellm/litellm_core_utils/llm_cost_calc/utils.py Allocates the uncached input budget using provider-reported cache modalities while preserving the old fallback; several added lines exceed repository limits
litellm/responses/utils.py Propagates cached_tokens_details from Responses and realtime usage normalization, with one overlong mapping line
litellm/types/utils.py Adds normalized models for text, audio, and image cached-token details
tests/test_litellm/litellm_core_utils/llm_cost_calc/test_llm_cost_calc_utils.py Adds focused regression tests for modality-aware realtime input pricing and fallback behavior
tests/test_litellm/test_cost_calculator.py Verifies per-modality cached-token counts are summed when usage objects are combined

Reviews (1): Last reviewed commit: "Merge remote-tracking branch 'origin/lit..." | Re-trigger Greptile

Comment on lines +2286 to +2289
return
if combined_details.cached_tokens_details is None:
combined_details.cached_tokens_details = CachedTokensDetails()
combined_cached_details: Final = combined_details.cached_tokens_details

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Mutable aggregation and long lines

This helper mutates its parameter without justification, and several added lines exceed 120 characters, preventing the change from meeting required Python standards

Context Used: CLAUDE.md (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@codecov

codecov Bot commented Aug 30, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@yuneng-berri
yuneng-berri deleted the branch litellm_internal_staging September 13, 2026 04:46
@yuneng-berri yuneng-berri reopened this Sep 13, 2026
@mateo-berri

Copy link
Copy Markdown
Contributor

Superseded by #40627, which landed the same per-modality cache subtraction plus the audio cache-read rate, so closing this one

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants