Skip to content

fix(cost-map): add Vertex batch cache prices to vertex_ai/claude-sonnet-5-5 - #43602

Merged
kerry-berri merged 1 commit into
mainfrom
litellm-providers/price-sync-vertex-ai
Sep 28, 2026
Merged

kerry-berri merged 1 commit into
mainfrom
litellm-providers/price-sync-vertex-ai

Conversation

@berriai-litellm-provider-info-sync

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • The Vertex pricing page now lists Claude Sonnet 5.5, including batch cache prices that vertex_ai/claude-sonnet-5-5 and @default on main do not carry: "5m Batch Cache Write $1.25", "Batch Cache Hit $0.10"
  • Without them, cost calculation for Sonnet 5.5 batch jobs on Vertex falls back to the non batch cache rates, overstating batch cache spend by 2x

How it solves it:

  • Adds cache_creation_input_token_cost_batches and cache_read_input_token_cost_batches to both keys (merge mode, every other field untouched)
  • Same edit in the main cost map and its backup copy

User Flow

Before: a developer running Claude Sonnet 5.5 batch jobs with prompt caching on Vertex sees cache tokens billed at the non batch rate

  1. They submit a Vertex batch prediction job for vertex_ai/claude-sonnet-5-5 that reads and writes the prompt cache
  2. Cost calculation finds no batch cache field and uses cache_read_input_token_cost $0.20 and cache_creation_input_token_cost $2.50 per million
  3. https://litellm-domain/ui/?page=logs overstates the batch cache spend by 2x

After: the same job is priced at the published Vertex batch cache rates

  1. They submit the same batch job
  2. Cache hits are priced at $0.10 and 5m cache writes at $1.25 per million tokens
  3. https://litellm-domain/ui/?page=logs matches the Vertex invoice

Rows

vertex-ai sync run, list-only lane. prices vertex-ai reports no structured price source, so both fields are kind page

vertex_ai/claude-sonnet-5-5 and vertex_ai/claude-sonnet-5-5@default, source https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing, Anthropic's Claude models global table, row "Sonnet 5.5":

  • cache_creation_input_token_cost_batches: 1.25e-06: "5m Batch Cache Write $1.25"
  • cache_read_input_token_cost_batches: 1e-07: "Batch Cache Hit $0.10"
  • "1h Batch Cache Write $2.00" has no catalog field; no Vertex Claude row carries one, so it is not represented

Every other price on the row already matches the same table: "Input $2.00", "Output $10.00", "Batch Input $1.00", "Batch Output $5.00", "5m Cache Write $2.50", "1h Cache Write $4.00", "Cache Hit $0.20", and the regional table "Input $2.20" matches regional_endpoint_uplift_multiplier 1.1. Field names follow vertex_ai/claude-opus-5-5, which carries both batch cache fields

Changed pages

Rulings applied

  • vertex-ai-delisted-no-date-keep: the 26 delisted ids in that ruling keep their rows and prices unchanged
  • vertex-ai-no-per-token-price-stay-out, vertex-ai-per-character-skip: textembedding-gecko, chirp-2, chirp-3, gemini-robotics-er-2-preview-info, translate-llm and translategemma stay out
  • vertex-gemini-flash-promo-switch-on-the-day: gemini-3.6/3.7/3.8-flash keep the promo price, no change before 2027-01-01
  • vertex-ai-gemini-2-5-flash-image-drop-priority, vertex-ai-gemini-2-5-flash-image-drop-cache-read: already applied on main

Delisted by the provider

No published date, kept with current prices per vertex-ai-delisted-no-date-keep: deep-research-pro-preview-12-2025, gemini-2.5-flash-lite-preview-09-2025, gemini-2.5-flash-preview-09-2025, gemini-2.5-pro-preview-tts, gemini-3-pro-image-preview, gemini-3.1-flash-lite-preview, gemini-3.1-pro-preview-customtools, gemini-3.8-flash-cyber, gemini-embedding-2-preview, gemini-flash-experimental, gemini-live-2.5-flash-preview-native-audio-09-2025, gemini-robotics-er-1.5-preview, medlm-large, medlm-medium, text-embedding-preview-0409, text-unicorn, vertex_ai/chirp, vertex_ai/claude-3-sonnet, vertex_ai/codestral-2, vertex_ai/codestral-2501, vertex_ai/mistral-large-2411, vertex_ai/mistral-medium-3, vertex_ai/mistral-ocr-2505, vertex_ai/mistral-small-2503, vertex_ai/veo-3.1-fast-generate-preview, vertex_ai/veo-3.1-generate-preview

Already carry a deprecation_date: text-embedding-004, vertex_ai/claude-3-5-haiku, vertex_ai/claude-3-5-sonnet, vertex_ai/claude-3-haiku, vertex_ai/claude-3-opus, vertex_ai/claude-sonnet-4-5, vertex_ai/gemini-2.0-flash, vertex_ai/gemini-2.0-flash-lite, vertex_ai/gemini-2.5-flash-native-audio, vertex_ai/jamba-1.5, vertex_ai/jamba-1.5-large, vertex_ai/jamba-1.5-mini

Naming mismatches with a live list id, not delistings: vertex_ai/chirp_2 and vertex_ai/chirp_3 (list: chirp-2, chirp-3), vertex_ai/gemini-3.5-transcribe and vertex_ai/gemini-3.5-transcribe-live (list: the -preview ids), vertex_ai/gemini-omni-1.1-flash (list: gemini-omni-1.1-flash-preview), vertex_ai/gemma-4-26b-a4b-it (list: gemma-4-26b-a4b-it-maas)

No published date, left for a human: vertex_ai/gemini-2.5-flash-preview-tts (list now carries gemini-2.5-flash-tts)

Coverage gap: vertex_ai/search_api, the publisher models list does not cover the Search API

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests (data only change, covered by the existing schema and cost map guard tests)
  • The handful of test files covering my change pass locally
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review

Type

🐛 Bug Fix

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@cursor

cursor Bot commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

Bugbot is paused — on-demand spend limit reached

Bugbot uses usage-based billing for this team and has hit its on-demand spend limit.

A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue.

@greptile-apps

greptile-apps Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

[Critical risk] Adds pricing data that affects billing calculations.

The PR appears safe to merge; the remaining concern is a non-blocking regression test for the new prices.

Findings

  1. P2 Batch prices lack a regression test ▶

Summary

Adds Vertex Claude Sonnet 5.5 batch cache-write and cache-hit prices to both model identifiers in the main and backup cost maps.

  • The added values use supported batch-pricing fields.
  • A model-specific cost assertion would guard the intended billing result.

Reviews (1) · Last reviewed commit: "fix(cost-map): add Vertex batch cache pr..."

Comment on lines +78413 to +78415
"cache_creation_input_token_cost_batches": 1.25e-06,
"cache_read_input_token_cost": 2e-07,
"cache_read_input_token_cost_batches": 1e-07,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Batch prices lack a regression test The existing tests check that the main and backup catalogs match and that batch cache rates work for a synthetic model. They do not verify the calculated cache-hit and 5-minute cache-write costs for either Vertex Sonnet 5.5 identifier. A later edit could change these prices without failing those tests; please add a model-specific cost test.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That’s fair. Since this is a data-only catalog update, the generic batch-cache cost-path test validates the calculation behavior and the cost-map guard tests validate the catalog entries and backup consistency. A model-specific test would mainly duplicate those checks without adding meaningful coverage here, so no additional test is required.

@codecov

codecov Bot commented Sep 28, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@codspeed

codspeed Bot commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm-providers/price-sync-vertex-ai (f95116d) with main (0f96d09)1

Open in CodSpeed

Footnotes

  1. No successful run was found on main (f20c400) during the generation of this report, so 0f96d09 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant