Repository navigation
fix(cost-map): add Vertex batch cache prices to vertex_ai/claude-sonnet-5-5 - #43602
Conversation
…et-5-5 Price-Sync: litellm-providers
|
|
Bugbot is paused — on-demand spend limit reachedBugbot uses usage-based billing for this team and has hit its on-demand spend limit. A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue. |
|
| "cache_creation_input_token_cost_batches": 1.25e-06, | ||
| "cache_read_input_token_cost": 2e-07, | ||
| "cache_read_input_token_cost_batches": 1e-07, |
There was a problem hiding this comment.
Batch prices lack a regression test The existing tests check that the main and backup catalogs match and that batch cache rates work for a synthetic model. They do not verify the calculated cache-hit and 5-minute cache-write costs for either Vertex Sonnet 5.5 identifier. A later edit could change these prices without failing those tests; please add a model-specific cost test.
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
There was a problem hiding this comment.
That’s fair. Since this is a data-only catalog update, the generic batch-cache cost-path test validates the calculation behavior and the cost-map guard tests validate the catalog entries and backup consistency. A model-specific test would mainly duplicate those checks without adding meaningful coverage here, so no additional test is required.
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
TLDR
Problem this solves:
vertex_ai/claude-sonnet-5-5and@defaulton main do not carry: "5m Batch Cache Write $1.25", "Batch Cache Hit $0.10"How it solves it:
cache_creation_input_token_cost_batchesandcache_read_input_token_cost_batchesto both keys (merge mode, every other field untouched)User Flow
Before: a developer running Claude Sonnet 5.5 batch jobs with prompt caching on Vertex sees cache tokens billed at the non batch rate
vertex_ai/claude-sonnet-5-5that reads and writes the prompt cachecache_read_input_token_cost$0.20 andcache_creation_input_token_cost$2.50 per millionAfter: the same job is priced at the published Vertex batch cache rates
Rows
vertex-ai sync run, list-only lane.
prices vertex-aireports no structured price source, so both fields are kind pagevertex_ai/claude-sonnet-5-5andvertex_ai/claude-sonnet-5-5@default, source https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing, Anthropic's Claude models global table, row "Sonnet 5.5":cache_creation_input_token_cost_batches: 1.25e-06: "5m Batch Cache Write $1.25"cache_read_input_token_cost_batches: 1e-07: "Batch Cache Hit $0.10"Every other price on the row already matches the same table: "Input $2.00", "Output $10.00", "Batch Input $1.00", "Batch Output $5.00", "5m Cache Write $2.50", "1h Cache Write $4.00", "Cache Hit $0.20", and the regional table "Input $2.20" matches
regional_endpoint_uplift_multiplier1.1. Field names followvertex_ai/claude-opus-5-5, which carries both batch cache fieldsChanged pages
Rulings applied
vertex-ai-delisted-no-date-keep: the 26 delisted ids in that ruling keep their rows and prices unchangedvertex-ai-no-per-token-price-stay-out,vertex-ai-per-character-skip: textembedding-gecko, chirp-2, chirp-3, gemini-robotics-er-2-preview-info, translate-llm and translategemma stay outvertex-gemini-flash-promo-switch-on-the-day: gemini-3.6/3.7/3.8-flash keep the promo price, no change before 2027-01-01vertex-ai-gemini-2-5-flash-image-drop-priority,vertex-ai-gemini-2-5-flash-image-drop-cache-read: already applied on mainDelisted by the provider
No published date, kept with current prices per
vertex-ai-delisted-no-date-keep: deep-research-pro-preview-12-2025, gemini-2.5-flash-lite-preview-09-2025, gemini-2.5-flash-preview-09-2025, gemini-2.5-pro-preview-tts, gemini-3-pro-image-preview, gemini-3.1-flash-lite-preview, gemini-3.1-pro-preview-customtools, gemini-3.8-flash-cyber, gemini-embedding-2-preview, gemini-flash-experimental, gemini-live-2.5-flash-preview-native-audio-09-2025, gemini-robotics-er-1.5-preview, medlm-large, medlm-medium, text-embedding-preview-0409, text-unicorn, vertex_ai/chirp, vertex_ai/claude-3-sonnet, vertex_ai/codestral-2, vertex_ai/codestral-2501, vertex_ai/mistral-large-2411, vertex_ai/mistral-medium-3, vertex_ai/mistral-ocr-2505, vertex_ai/mistral-small-2503, vertex_ai/veo-3.1-fast-generate-preview, vertex_ai/veo-3.1-generate-previewAlready carry a
deprecation_date: text-embedding-004, vertex_ai/claude-3-5-haiku, vertex_ai/claude-3-5-sonnet, vertex_ai/claude-3-haiku, vertex_ai/claude-3-opus, vertex_ai/claude-sonnet-4-5, vertex_ai/gemini-2.0-flash, vertex_ai/gemini-2.0-flash-lite, vertex_ai/gemini-2.5-flash-native-audio, vertex_ai/jamba-1.5, vertex_ai/jamba-1.5-large, vertex_ai/jamba-1.5-miniNaming mismatches with a live list id, not delistings: vertex_ai/chirp_2 and vertex_ai/chirp_3 (list: chirp-2, chirp-3), vertex_ai/gemini-3.5-transcribe and vertex_ai/gemini-3.5-transcribe-live (list: the -preview ids), vertex_ai/gemini-omni-1.1-flash (list: gemini-omni-1.1-flash-preview), vertex_ai/gemma-4-26b-a4b-it (list: gemma-4-26b-a4b-it-maas)
No published date, left for a human: vertex_ai/gemini-2.5-flash-preview-tts (list now carries gemini-2.5-flash-tts)
Coverage gap: vertex_ai/search_api, the publisher models list does not cover the Search API
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
Type
🐛 Bug Fix
Final Attestation