Repository navigation
fix(cost_calculator): bill ultrafast prompts above 272k at the ultrafast long-context rates - #43764
Conversation
…ast long-context rates Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
|
|
|
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
… pricing fields Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 17e4264. Configure here.
* upstream/main: fix(cost_calculator): bill ultrafast prompts above 272k at the ultrafast long-context rates (BerriAI#43764) chore(model_prices): add Gemini Veo, Mistral and Azure Claude 4.5 deprecation dates (BerriAI#43857)
Narrow the arguments passed to the storage cost helper, add the storage charge to the total outside the additional_costs sum, and type the new test helpers so the changed files add no reportUnknownArgumentType or reportGeneralTypeIssues errors. Raise the reportUnknownArgumentType limit by 97 for the new cache_storage_cost_per_token_per_hour pricing field: like every pricing field, it adds one diagnostic at each existing untyped GenericLiteLLMParams(**kwargs) call site (same approach as BerriAI#43764).
TLDR
Problem this solves:
gpt-6-astraUltrafast prompts above 272K billed at standard long-context rates*_above_272k_tokens_ultrafastprices leaked into the provider requestHow it solves it:
get_model_infoandModelInfoBaseUser Flow
Before: a team sending long prompts on the Ultrafast tier sees roughly a sixth of the real cost
"model": "gpt-6-astra","service_tier": "ultrafast"and a ~280K token input"service_tier": "ultrafast"and 279,788 input tokensx-litellm-response-costheader and the spend log both read $6.99506, the standard above-272K priceAfter: the same request is billed at the Ultrafast above-272K prices
Linear ticket
Resolves LIT-8993
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/unit/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Real OpenAI calls through a local proxy on real Postgres, one proxy per commit, same config:
long_body.jsonis{"model": "gpt-6-astra", "service_tier": "ultrafast", "input": "<279,788 tokens of numbered records>", "max_output_tokens": 16}Expected prices come from https://developers.openai.com/api/docs/pricing (checked 2026-09-29): Ultrafast above 272K is $120 / $450 per 1M input / output, $12 cached read, $150 cache write
Before (27c110c, first run at f4a7c04)
Long Ultrafast prompt, cache write
curl -sD - localhost:4101/v1/responses -H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' --data @long_body.json"service_tier": "ultrafast",input_tokens: 279788,cache_write_tokens: 279785,output_tokens: 5x-litellm-response-cost: 6.9950600000000005, spend log row6.99506. The Ultrafast price is 3 x 1.2e-4 + 279,785 x 1.5e-4 + 5 x 4.5e-4 = 41.97036After (4dfedd8, cache-write run at 89255e909b)
Long Ultrafast prompt, cache write
localhost:4102with a fresh prefix so nothing is cached"service_tier": "ultrafast",input_tokens: 279792,cache_write_tokens: 279789,output_tokens: 5x-litellm-response-cost: 41.97095999999999, spend log row41.97096, which equals 3 x 1.2e-4 + 279,789 x 1.5e-4 + 5 x 4.5e-4Long Ultrafast prompt, cache read
localhost:4102with the original bodycached_tokens: 279785,output_tokens: 5x-litellm-response-cost: 3.3600300000000005, which equals 3 x 1.2e-4 + 279,785 x 1.2e-5 + 5 x 4.5e-4Short Ultrafast prompts (8 in, 5 out) cost $0.00198 on both commits, so ordinary Ultrafast pricing is unchanged
Rerun at the final head (17e4264 vs merge base 27c110c)
The same A/B with a new uncached prompt, sent first to base and then unchanged to head:
"service_tier": "ultrafast",input_tokens: 279792,cache_write_tokens: 279789,output_tokens: 16,x-litellm-response-cost: 6.995985000000001, the same value in the spend log. That is still the standard above-272K pricecached_tokens: 279789,output_tokens: 16,x-litellm-response-cost: 3.365028, the same value in the spend log, which equals 3 x 1.2e-4 + 279,789 x 1.2e-5 + 16 x 4.5e-4The new
tests/integration/pricing/test_service_tier_pricing.pycases run each scenario through both/v1/chat/completionsand/v1/responses. They failed on unfixed main (Obtained: 0.90292, Expected: 15.042, plus the 4 price keys reaching the upstream body) and pass on this branchType
🐛 Bug Fix
✅ Test
Caveats (if any)
Medium
basedpyright-code-budget.jsonraises thereportUnknownArgumentTypelimit by 444 (44,358 to 44,802), with maintainer approval. Each new field adds one diagnostic at each of 111 existing untypedGenericLiteLLMParams(**kwargs)call sitesLow
service_tier: "ultrafast"with 400, only Responses accepts itFinal Attestation
Link to Devin session: https://app.devin.ai/sessions/1594dd34eab14f28a3e07b61dba753fa
Open in Devin Desktop: https://app.devin.ai/desktop/session/1594dd34eab14f28a3e07b61dba753fa?variant=devin
Requested by: @kerry-berri
Note
Medium Risk
Changes token billing and spend tracking for a specific service tier and prompt-length band; impact is mitigated by broad unit, integration, and Rust tests but billing logic is financially sensitive.
Overview
Fixes under-billing for Ultrafast requests with prompts above 272K by adding the four missing
*_above_272k_tokens_ultrafastpricing fields (input, output, cache read, cache creation) to model metadata:ModelInfo/GenericLiteLLMParams,get_model_infomapping, OpenAPIschema.d.ts, and the catalog JSON schema test.generic_cost_per_tokencan now pick Ultrafast long-context rates (same pattern as flex/priority), including deployment overrides via routerlitellm_params, with tests covering catalog values, cost math, spend headers/logs on chat and responses, Rust tier+threshold selection, and router registration isolation.The basedpyright
reportUnknownArgumentTypebudget is raised slightly for the new typed fields.Reviewed by Cursor Bugbot for commit 17e4264. Bugbot is set up for automated code reviews on this repo. Configure here.