Skip to content

fix(cost_calculator): bill ultrafast prompts above 272k at the ultrafast long-context rates - #43764

Merged
kerry-berri merged 8 commits into
mainfrom
litellm_ultrafast_long_context_pricing
Sep 30, 2026
Merged

kerry-berri merged 8 commits into
mainfrom
litellm_ultrafast_long_context_pricing

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • gpt-6-astra Ultrafast prompts above 272K billed at standard long-context rates
  • Deployment-level *_above_272k_tokens_ultrafast prices leaked into the provider request
  • Main's catalog schema test fails on the 8 new Ultrafast keys

How it solves it:

  • Carry the 4 above-272K Ultrafast fields through get_model_info and ModelInfoBase
  • Register them as custom pricing params so they bill and get stripped
  • Add the 8 Ultrafast keys to the catalog schema test

User Flow

Before: a team sending long prompts on the Ultrafast tier sees roughly a sixth of the real cost

  1. They send POST http://localhost:4000/v1/responses with "model": "gpt-6-astra", "service_tier": "ultrafast" and a ~280K token input
  2. OpenAI answers 200 with "service_tier": "ultrafast" and 279,788 input tokens
  3. The x-litellm-response-cost header and the spend log both read $6.99506, the standard above-272K price

After: the same request is billed at the Ultrafast above-272K prices

  1. They send the same POST http://localhost:4000/v1/responses
  2. OpenAI answers the same way
  3. The header and spend log read $41.97096 on a cache write and $3.36003 on a cache read, matching OpenAI's Ultrafast long-context rates to the cent

Linear ticket

Resolves LIT-8993

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/unit/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Real OpenAI calls through a local proxy on real Postgres, one proxy per commit, same config:

model_list:
  - model_name: gpt-6-astra
    litellm_params:
      model: openai/gpt-6-astra
      api_key: os.environ/OPENAI_API_KEY

long_body.json is {"model": "gpt-6-astra", "service_tier": "ultrafast", "input": "<279,788 tokens of numbered records>", "max_output_tokens": 16}

Expected prices come from https://developers.openai.com/api/docs/pricing (checked 2026-09-29): Ultrafast above 272K is $120 / $450 per 1M input / output, $12 cached read, $150 cache write

Before (27c110c, first run at f4a7c04)

Long Ultrafast prompt, cache write

  1. curl -sD - localhost:4101/v1/responses -H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' --data @long_body.json
  2. 200, "service_tier": "ultrafast", input_tokens: 279788, cache_write_tokens: 279785, output_tokens: 5
  3. x-litellm-response-cost: 6.9950600000000005, spend log row 6.99506. The Ultrafast price is 3 x 1.2e-4 + 279,785 x 1.5e-4 + 5 x 4.5e-4 = 41.97036

After (4dfedd8, cache-write run at 89255e909b)

Long Ultrafast prompt, cache write

  1. Same curl on localhost:4102 with a fresh prefix so nothing is cached
  2. 200, "service_tier": "ultrafast", input_tokens: 279792, cache_write_tokens: 279789, output_tokens: 5
  3. x-litellm-response-cost: 41.97095999999999, spend log row 41.97096, which equals 3 x 1.2e-4 + 279,789 x 1.5e-4 + 5 x 4.5e-4

Long Ultrafast prompt, cache read

  1. Same curl on localhost:4102 with the original body
  2. 200, cached_tokens: 279785, output_tokens: 5
  3. x-litellm-response-cost: 3.3600300000000005, which equals 3 x 1.2e-4 + 279,785 x 1.2e-5 + 5 x 4.5e-4

Short Ultrafast prompts (8 in, 5 out) cost $0.00198 on both commits, so ordinary Ultrafast pricing is unchanged

Rerun at the final head (17e4264 vs merge base 27c110c)

The same A/B with a new uncached prompt, sent first to base and then unchanged to head:

  1. Base: 200, "service_tier": "ultrafast", input_tokens: 279792, cache_write_tokens: 279789, output_tokens: 16, x-litellm-response-cost: 6.995985000000001, the same value in the spend log. That is still the standard above-272K price
  2. Head: 200, cached_tokens: 279789, output_tokens: 16, x-litellm-response-cost: 3.365028, the same value in the spend log, which equals 3 x 1.2e-4 + 279,789 x 1.2e-5 + 16 x 4.5e-4
  3. A short prompt (35 in, 16 out) costs $0.0069 on both commits

The new tests/integration/pricing/test_service_tier_pricing.py cases run each scenario through both /v1/chat/completions and /v1/responses. They failed on unfixed main (Obtained: 0.90292, Expected: 15.042, plus the 4 price keys reaching the upstream body) and pass on this branch

Type

🐛 Bug Fix
✅ Test

Caveats (if any)

Medium

  • basedpyright-code-budget.json raises the reportUnknownArgumentType limit by 444 (44,358 to 44,802), with maintainer approval. Each new field adds one diagnostic at each of 111 existing untyped GenericLiteLLMParams(**kwargs) call sites

Low

  • OpenAI Chat Completions rejects service_tier: "ultrafast" with 400, only Responses accepts it
  • Main's 4 interactions OpenAPI tests and the osv-scan findings fail there too

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Link to Devin session: https://app.devin.ai/sessions/1594dd34eab14f28a3e07b61dba753fa
Open in Devin Desktop: https://app.devin.ai/desktop/session/1594dd34eab14f28a3e07b61dba753fa?variant=devin
Requested by: @kerry-berri


Note

Medium Risk
Changes token billing and spend tracking for a specific service tier and prompt-length band; impact is mitigated by broad unit, integration, and Rust tests but billing logic is financially sensitive.

Overview
Fixes under-billing for Ultrafast requests with prompts above 272K by adding the four missing *_above_272k_tokens_ultrafast pricing fields (input, output, cache read, cache creation) to model metadata: ModelInfo / GenericLiteLLMParams, get_model_info mapping, OpenAPI schema.d.ts, and the catalog JSON schema test.

generic_cost_per_token can now pick Ultrafast long-context rates (same pattern as flex/priority), including deployment overrides via router litellm_params, with tests covering catalog values, cost math, spend headers/logs on chat and responses, Rust tier+threshold selection, and router registration isolation.

The basedpyright reportUnknownArgumentType budget is raised slightly for the new typed fields.

Reviewed by Cursor Bugbot for commit 17e4264. Bugbot is set up for automated code reviews on this repo. Configure here.

…ast long-context rates

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration
devin-ai-integration Bot requested a review from a team September 29, 2026 20:15
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@greptile-apps

greptile-apps Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

[Critical risk] Adds pricing fields that affect how customer charges are calculated.

The PR should not merge until the prohibited type-check budget increase is removed

Findings

  1. P1 Type-check budget increased ▶

Summary

The PR wires above-272K ultrafast prices through model information and deployment pricing, updates dashboard types, and adds billing tests for chat and Responses. The remaining concern is the prohibited type-check budget increase

Reviews (5) · Last reviewed commit: "chore(lint): raise reportUnknownArgument..."

Comment thread tests/integration/pricing/test_service_tier_pricing.py

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@codspeed

codspeed Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_ultrafast_long_context_pricing (17e4264) with main (ffb15f9)

Open in CodSpeed

@codecov

codecov Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

kerry and others added 4 commits September 29, 2026 20:45
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread tests/unit/test_openai_service_tier_long_context_pricing.py Outdated
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

… pricing fields

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Comment thread basedpyright-code-budget.json

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 17e4264. Configure here.

@kerry-berri
kerry-berri merged commit 9dda4d8 into main Sep 30, 2026
100 of 104 checks passed
fishingpvalues pushed a commit to fishingpvalues/litellm that referenced this pull request Sep 30, 2026
* upstream/main:
  fix(cost_calculator): bill ultrafast prompts above 272k at the ultrafast long-context rates (BerriAI#43764)
  chore(model_prices): add Gemini Veo, Mistral and Azure Claude 4.5 deprecation dates (BerriAI#43857)
roasted-penguin added a commit to roasted-penguin/litellm that referenced this pull request Oct 8, 2026
Narrow the arguments passed to the storage cost helper, add the storage
charge to the total outside the additional_costs sum, and type the new
test helpers so the changed files add no reportUnknownArgumentType or
reportGeneralTypeIssues errors.

Raise the reportUnknownArgumentType limit by 97 for the new
cache_storage_cost_per_token_per_hour pricing field: like every pricing
field, it adds one diagnostic at each existing untyped
GenericLiteLLMParams(**kwargs) call site (same approach as BerriAI#43764).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant