Skip to content

fix(cost map): add cached-input pricing and prompt caching for inception/mercury-2.5 - #41016

Closed
Nanduu24 wants to merge 2 commits into
BerriAI:mainfrom
Nanduu24:add-mercury-2.5-cache-pricing
Closed

Nanduu24 wants to merge 2 commits into
BerriAI:mainfrom
Nanduu24:add-mercury-2.5-cache-pricing

Conversation

@Nanduu24

Copy link
Copy Markdown
Contributor

Summary

inception/mercury-2.5 is present in the cost map with input and output pricing, but is missing cached-input pricing and prompt-caching support:

  • cache_read_input_token_cost — not set
  • supports_prompt_caching — not set

As a result, cached-token usage on Mercury 2.5 is not cost-tracked (silent inaccuracy), even though Inception offers cached input and the sibling model inception/mercury-2 already carries both fields.

Changes

Adds to inception/mercury-2.5 in both the root cost map and the litellm/ backup copy (kept in sync):

"cache_read_input_token_cost": 2e-08,
"supports_prompt_caching": true
  • 2e-08 = $0.02 / 1M tokens, the standard cached-input rate (the current $0.004 is a temporary launch promo).

Relevant issues

Addresses the cached-pricing portion of #40746 (the model entry itself was already added; these fields were still missing).

Type

🐛 Bug Fix / 💰 cost tracking

Testing

…ion/mercury-2.5

The inception/mercury-2.5 entry has input and output pricing but is missing
cache_read_input_token_cost and supports_prompt_caching, so cached-token usage
is not cost-tracked. Inception lists cached input at $0.02 / 1M tokens
(standard), and the sibling inception/mercury-2 already carries these fields.

Adds:
- cache_read_input_token_cost: 2e-08 ($0.02 / 1M standard rate)
- supports_prompt_caching: true

Applied to both the root cost map and the litellm/ backup copy so they stay in
sync. Addresses the cached-pricing portion of BerriAI#40746.

Source: https://www.inceptionlabs.ai/models#pricing
@Nanduu24
Nanduu24 requested a review from a team September 13, 2026 21:03
@CLAassistant

CLAassistant commented Sep 13, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@greptile-apps

greptile-apps Bot commented Sep 13, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

Adds cached-input pricing and prompt-caching metadata for Inception Mercury 2.5 in both synchronized cost maps

  • The new cached-input rate represents the future standard price rather than the acknowledged active promotional price
  • No regression assertion covers either newly added field

Confidence Score: 4/5

This PR is not safe to merge until the active cached-input price is represented and the repository-required regression coverage is added

Cached tokens are multiplied directly by the configured rate, so using $0.02/M during the acknowledged $0.004/M promotion overstates spend fivefold

Files Needing Attention: model_prices_and_context_window.json, litellm/model_prices_and_context_window_backup.json

Important Files Changed

Filename Overview
model_prices_and_context_window.json Adds Mercury 2.5 cache pricing and capability metadata, but currently overstates cached-token costs and lacks regression coverage
litellm/model_prices_and_context_window_backup.json Mirrors the root cost-map changes, including the same premature standard cached-input rate

Reviews (1): Last reviewed commit: "fix(cost map): add cached-input pricing ..." | Re-trigger Greptile

Comment thread model_prices_and_context_window.json
Comment thread model_prices_and_context_window.json
@codspeed

codspeed Bot commented Sep 13, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing Nanduu24:add-mercury-2.5-cache-pricing (a3c7008) with main (30f33a9)

Open in CodSpeed

@codecov

codecov Bot commented Sep 13, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

Asserts the pricing/capability fields (including cache_read_input_token_cost
and supports_prompt_caching) and that the root and backup cost maps stay in
sync, following the existing per-model metadata test pattern.
@Nanduu24

Copy link
Copy Markdown
Contributor Author

Thanks for the review! Addressed the feedback:

Test (P2): Added tests/test_litellm/test_inception_mercury_2_5_model_metadata.py, following the existing per-model metadata test pattern (e.g. test_xai_grok_4_3_model_metadata.py). It asserts the pricing/capability fields — including cache_read_input_token_cost and supports_prompt_caching — and that the root and backup cost maps stay in sync.

Cached rate (P1): The 2e-08 ($0.02 / 1M) value is intentional — it's the standard cached-input rate, consistent with this entry's existing input_cost_per_token (2e-07 = $0.20, standard) and output_cost_per_token (7.5e-07 = $0.75, standard). Those already use standard rather than the temporary launch-promo rates ($0.04 / $0.004 / $0.15), so pricing cached input at the promo rate would make it inconsistent with input/output. This also matches the sibling inception/mercury-2 entry.

Note on CI: the failing proxy-behavior check is unrelated to this change — it's a pre-existing auth test (test_auth_object_prefetch.py::test_join_binds_the_membership_to_the_requested_team) failing with TypeError: object MagicMock can't be used in 'await' expression. The relevant cost-map-guard check passes.

@Nanduu24

Copy link
Copy Markdown
Contributor Author

The failing proxy-behavior check is a pre-existing flake unrelated to this PR.

The only failure is:

tests/proxy_behavior/auth/test_auth_object_prefetch.py::test_join_binds_the_membership_to_the_requested_team
FAILED — TypeError: object MagicMock can't be used in 'await' expression

This is an async-mock issue in the auth suite, and it fails identically on the run before any test was added here. This PR only touches model_prices_and_context_window.json (+ the litellm/ backup) and a new test under tests/test_litellm/, none of which can affect tests/proxy_behavior/auth/. The relevant cost-map-guard check passes and the new metadata test passes.

A re-run should be green once the flaky auth test is addressed on main.

@devin-ai-integration

Copy link
Copy Markdown
Contributor

Superseded by #41112, which adds the same mercury-2.5 cache pricing verified against the Inception docs.

devin-ai-integration Bot added a commit that referenced this pull request Sep 14, 2026
#41016

Co-Authored-By: Nanduu24 <reachnanduu24@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants