Repository navigation
fix(cost): bill uncached realtime tokens per modality using cached_tokens_details - #38858
devin-ai-integration[bot] wants to merge 3 commits into
Conversation
…s modality split Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…itellm_fix_realtime_cached_audio_billing
🤖 Devin AI EngineerI'll be helping with this pull request! Here's what you should know: ✅ I will automatically:
Note: I can only respond to comments from users who have write access to this repository. ⚙️ Control Options:
|
|
|
|
PR #38858 lacks the |
Merging this PR will not alter performance
Comparing |
Greptile SummaryThis PR carries realtime cached-token modality details through usage normalization and aggregation, then uses them to price uncached input by modality
Confidence Score: 4/5The pricing fix appears safe to merge after the non-blocking Python guideline violations are cleaned up The cached-token metadata follows the realtime normalization, aggregation, and cost-allocation path, with focused regression coverage; the remaining concern is code-quality compliance Files Needing Attention: litellm/cost_calculator.py, litellm/litellm_core_utils/llm_cost_calc/utils.py, litellm/responses/utils.py
|
| Filename | Overview |
|---|---|
| litellm/cost_calculator.py | Aggregates cached-token modality details across realtime usage objects, but the helper mutates its parameter and includes an overlong call |
| litellm/litellm_core_utils/llm_cost_calc/utils.py | Allocates the uncached input budget using provider-reported cache modalities while preserving the old fallback; several added lines exceed repository limits |
| litellm/responses/utils.py | Propagates cached_tokens_details from Responses and realtime usage normalization, with one overlong mapping line |
| litellm/types/utils.py | Adds normalized models for text, audio, and image cached-token details |
| tests/test_litellm/litellm_core_utils/llm_cost_calc/test_llm_cost_calc_utils.py | Adds focused regression tests for modality-aware realtime input pricing and fallback behavior |
| tests/test_litellm/test_cost_calculator.py | Verifies per-modality cached-token counts are summed when usage objects are combined |
Reviews (1): Last reviewed commit: "Merge remote-tracking branch 'origin/lit..." | Re-trigger Greptile
| return | ||
| if combined_details.cached_tokens_details is None: | ||
| combined_details.cached_tokens_details = CachedTokensDetails() | ||
| combined_cached_details: Final = combined_details.cached_tokens_details |
There was a problem hiding this comment.
Mutable aggregation and long lines
This helper mutates its parameter without justification, and several added lines exceed 120 characters, preventing the change from meeting required Python standards
Context Used: CLAUDE.md (source)
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
|
Superseded by #40627, which landed the same per-modality cache subtraction plus the audio cache-read rate, so closing this one |
TLDR
Problem this solves:
How it solves it:
cached_tokens_detailsfrom realtime usage into the spend pipelineUser Flow
Before: a developer running realtime voice sessions through the gateway sees spend that is higher than what OpenAI bills them
response.donereports 3400 input tokens (1400 text, 2000 audio, 2000 cached of which 1500 audio and 500 text) and 1100 output tokensAfter: the same session logs the price OpenAI actually charges
response.donereports the same usageRelevant issues
Linear ticket
Resolves LIT-6513
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Live proxy run on localhost:4400 with
LITELLM_LOCAL_MODEL_COST_MAP=True, a spend-logging callback writing to a jsonl file, and a realtime backend replaying a fixed GAresponse.doneusage (3400 input: 1400 text + 2000 audio, 2000 cached withcached_tokens_detailsof 1500 audio + 500 text; 1100 output: 100 text + 1000 audio) so the token counts are identical across both runs. gpt-realtime rates: text $4/M, audio $32/M, cached $0.40/M, text out $16/M, audio out $64/M. Correct price by hand: 500 audio x $32/M + 900 text x $4/M + 2000 cached x $0.40/M + 1000 audio out x $64/M + 100 text out x $16/M = $0.0860Before (d44d281)
LITELLM_LOCAL_MODEL_COST_MAP=True uv run python litellm/proxy/proxy_cli.py --config /tmp/rt_repro/config.yaml --port 4400ws://127.0.0.1:4400/v1/realtime?model=gpt-realtime, send{"type": "response.create"}, receiveresponse.done, disconnectcat /tmp/rt_repro/spend.jsonlshows{"call_type": "_arealtime", "model": "gpt-realtime", "response_cost": 0.11120000000000001, "prompt_tokens": 3400, "completion_tokens": 1100, "spend": 0.11120000000000001}: all 1400 uncached tokens billed as audio ($0.0448 instead of $0.0196)After (a42f152)
cat /tmp/rt_repro/spend.jsonlshows{"call_type": "_arealtime", "model": "gpt-realtime", "response_cost": 0.08600000000000001, "prompt_tokens": 3400, "completion_tokens": 1100, "spend": 0.08600000000000001}, matching the hand-computed $0.0860Type
🐛 Bug Fix
Caveats (if any)
Low
cached_tokens_detailskeep the old audio-first fallbackFinal Attestation
Link to Devin session: https://app.devin.ai/sessions/d8222dd829134327a93215f75c4500f9
Open in Devin Desktop: https://app.devin.ai/desktop/session/d8222dd829134327a93215f75c4500f9?variant=devin