Repository navigation
chore(prices): sync Azure, Azure AI, Gemini, OpenAI, Bedrock, Together AI, Fireworks and Vertex prices: 278 models, 59 new, 30 deprecated - #41154
Conversation
|
|
|
Price check at 5dc31fa: 4 models changed, 0 mismatches against models.dev; none listed on OpenRouter. Details in #litellm-provider-info-sync. |
Greptile SummaryThis PR synchronizes provider pricing, model metadata, regional Azure entries, and deprecation dates across the primary and backup model catalogs.
Confidence Score: 5/5The PR appears safe to merge, with no outstanding blocking findings. No new actionable defect was established. The earlier incorrect Gemini per-second billing finding was withdrawn and resolved. devin-ai-integration[bot] accepted the temporary pricing-test coverage gap because the removed assertions pinned changing vendor prices and stated that catalog-independent replacements are planned; that thread was subsequently resolved. Important Files Changed
Reviews (4): Last reviewed commit: "fix(prices): add tpm and rpm to gemini 3..." | Re-trigger Greptile |
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
|
Price check at 956b211: 25 Bedrock entries only gained a source field, 4 Gemini models changed, 0 mismatches. Details in #litellm-provider-info-sync. |
956b211 to
a6bcc44
Compare
|
Price check at a6bcc44: 25 Bedrock entries only gained a source field, 4 Gemini entries changed secondary fields, 0 mismatches. Details in #litellm-provider-info-sync. |
|
Price check at 3986ea1: 18 entries changed only batch or audio fields, 0 mismatches against models.dev and OpenRouter. Details in #litellm-provider-info-sync |
|
Price check at 8c1300b: 18 entries changed only batch, audio or image fields, 0 mismatches against models.dev and OpenRouter. Details in #litellm-provider-info-sync |
|
Price check at 77812c6: 43 entries changed only source, batch, audio, image or per-second fields, 0 mismatches against models.dev and OpenRouter. Details in #litellm-provider-info-sync |
|
Price check at 8e2effb: 39 entries changed only source, batch, audio, image or per-second fields, 0 mismatches on fields this PR changes. Details in #litellm-provider-info-sync |
|
Price check at 6ce0144: 43 entries changed only source, batch, audio, image or per-second fields, 0 mismatches on fields this PR changes. Details in #litellm-provider-info-sync |
gemini-2.5-flash-preview-tts: output_cost_per_audio_token, input_cost_per_token_batches gemini/gemini-2.5-flash-preview-tts: output_cost_per_audio_token, input_cost_per_token_batches gemini-3.1-flash-live-preview: input_cost_per_second gemini/gemini-3.1-flash-live-preview: input_cost_per_second
anthropic.claude-fable-5: anthropic.claude-fable-5-1: anthropic.claude-opus-4-7: anthropic.claude-opus-4-8: anthropic.claude-opus-5: anthropic.claude-sonnet-4-6: anthropic.claude-sonnet-5: global.anthropic.claude-fable-5: global.anthropic.claude-fable-5-1: global.anthropic.claude-opus-4-7: global.anthropic.claude-opus-4-8: global.anthropic.claude-opus-5: global.anthropic.claude-sonnet-4-6: global.anthropic.claude-sonnet-5: us-gov.anthropic.claude-fable-5-1: us-gov.anthropic.claude-opus-4-8: us-gov.anthropic.claude-opus-5: us-gov.anthropic.claude-sonnet-5: us.anthropic.claude-fable-5: us.anthropic.claude-fable-5-1: us.anthropic.claude-opus-4-7: us.anthropic.claude-opus-4-8: us.anthropic.claude-opus-5: us.anthropic.claude-sonnet-4-6: us.anthropic.claude-sonnet-5:
chatgpt-image-latest: input_cost_per_image_token_batches gemini-2.0-flash: input_cost_per_audio_token_batches gemini-2.0-flash-lite: input_cost_per_audio_token_batches gemini-2.5-flash: input_cost_per_audio_token_batches gemini-2.5-flash-lite: input_cost_per_audio_token_batches gemini-3-flash-preview: input_cost_per_audio_token_batches vertex_ai/gemini-3-flash-preview: input_cost_per_audio_token_batches gemini-3.1-flash-lite: input_cost_per_audio_token_batches vertex_ai/gemini-3.1-flash-lite: input_cost_per_audio_token_batches gpt-image-1: input_cost_per_image_token_batches gpt-image-1-mini: input_cost_per_image_token_batches gpt-image-1.5: input_cost_per_image_token_batches gpt-image-1.5-2025-12-16: input_cost_per_image_token_batches gpt-image-2: input_cost_per_image_token_batches
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
us-gov.anthropic.claude-fable-5-1: us-gov.anthropic.claude-opus-4-8: us-gov.anthropic.claude-opus-5: us-gov.anthropic.claude-sonnet-5:
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
Price check at 91c8d1c: 86 models changed, 2 mismatches against models.dev and OpenRouter (azure/gpt-5.6-sol cache prices, Llama-4-Maverick OpenRouter only); 75 Azure eu/us tiers have no direct comparison. Details in #litellm-provider-info-sync. |
… Google Gemini, sync failed: AWS Bedrock, 6 held] gemini/gemini-3.1-pro-preview-customtools: input_cost_per_token_flex, output_cost_per_token_flex, cache_read_input_token_cost_flex
|
Price check at c5a388a: 86 keys with price changes, 75 regional tiers with no direct comparison, 1 mismatch (azure/gpt-5.6-sol cache fields). Details in #litellm-provider-info-sync |
gemini/gemini-robotics-er-2-preview: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, input_cost_per_token_batches, output_cost_per_token_batches
|
Price check at 641dcd5: 87 keys changed, 75 regional tiers with no direct comparison, 2 mismatches (azure/gpt-5.6-sol cache fields, Phi-4-mini-reasoning context). Details in #litellm-provider-info-sync |
together_ai/google/gemma-4-31B-it: deprecation_date together_ai/intfloat/multilingual-e5-large-instruct: deprecation_date together_ai/openai/gpt-oss-20b: deprecation_date together_ai/thinkingmachines/Inkling-Small: deprecation_date
|
Price check at 0443605: 87 keys changed, 75 regional tiers with no direct comparison, 2 mismatches (azure/gpt-5.6-sol cache fields, Phi-4-mini-reasoning context), details in #litellm-provider-info-sync |
|
Price check at d398c91: 87 keys changed, 75 regional tiers with no direct comparison, 1 mismatch (azure/gpt-5.6-sol cache fields), details in #litellm-provider-info-sync |
azure_ai/Codestral-2501: azure_ai/cohere-command-a: azure_ai/deepseek-r1: azure_ai/deepseek-v3: azure_ai/deepseek-v3-0324: azure_ai/deepseek-v3.1: azure_ai/deepseek-v3.2: azure_ai/deepseek-v3.2-speciale: azure_ai/deepseek-v4-flash: azure_ai/DeepSeek-V4-Flash-0731: azure_ai/deepseek-v4-pro: azure_ai/embed-v-4-0: azure_ai/FW-DeepSeek-V3.2: azure_ai/FW-DeepSeek-V4-Pro: azure_ai/FW-GLM-5: azure_ai/FW-GLM-5.1: azure_ai/FW-GLM-5.2: azure_ai/FW-GLM-5.2-Fast: azure_ai/FW-Inkling: azure_ai/FW-Kimi-K2.5: azure_ai/FW-Kimi-K2.6: azure_ai/FW-Kimi-K2.7-Code: azure_ai/FW-Kimi-K3: azure_ai/FW-MiniMax-M2.5: azure_ai/FW-MiniMax-M3: azure_ai/FW-Nemotron-3-Ultra-NVFP4: azure_ai/FW-Nemotron-Lightning-3.5-30B-A3B: azure_ai/gpt-oss-120b: azure_ai/grok-3: azure_ai/global/grok-3: azure_ai/grok-3-mini: azure_ai/global/grok-3-mini: azure_ai/grok-4: azure_ai/grok-4-1-fast-non-reasoning: azure_ai/grok-4-1-fast-reasoning: azure_ai/grok-4-20-non-reasoning: azure_ai/grok-4-20-reasoning: azure_ai/grok-4-fast-non-reasoning: azure_ai/grok-4-fast-reasoning: azure_ai/grok-4.3: azure_ai/grok-4.6: azure_ai/grok-code-fast-1: azure_ai/kimi-k2.5: azure_ai/kimi-k2.6: azure_ai/kimi-k2.7-code: azure_ai/Llama-3.3-70B-Instruct: azure_ai/Llama-4-Maverick-17B-128E-Instruct-FP8: azure_ai/MAI-DS-R1: azure_ai/MAI-Image-2.5: azure_ai/MAI-Image-2.5-Flash: azure_ai/MAI-Image-2e: azure_ai/MAI-Thinking-1: azure_ai/mistral-large-3: azure_ai/Phi-3-medium-128k-instruct: azure_ai/Phi-3-medium-4k-instruct: azure_ai/Phi-3-mini-128k-instruct: azure_ai/Phi-3-mini-4k-instruct: azure_ai/Phi-3-small-128k-instruct: azure_ai/Phi-3-small-8k-instruct: azure_ai/Phi-3.5-mini-instruct:
|
Price check at 8096856: 87 keys changed, 75 regional tiers with no direct comparison, 2 mismatches (azure/gpt-5.6-sol cache fields, Phi-4-mini-reasoning context), details in #litellm-provider-info-sync |
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
Price check at cab6732: 87 keys changed, 75 regional tiers with no direct comparison, 2 mismatches (azure/gpt-5.6-sol cache fields, Phi-4-mini-reasoning context), details in #litellm-provider-info-sync |
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
Price check at b9ab362: 87 keys changed, 75 regional tiers with no direct comparison, 1 mismatch (azure/gpt-5.6-sol cache fields), details in #litellm-provider-info-sync |
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit b9ab362. Configure here.
TLDR
Problem this solves:
How it solves it:
azure/eu/*andazure/us/*rows priced from the Azure retail APIgemini/gemini-3.8-liveandgemini/gemini-3.8-live-extended-thinkingdeprecation_dateon 30 Together AI, Fireworks and Vertex Imagen rowssourceURL on 228 rows so every value is traceableUser Flow
Before: a proxy admin routing
azure/gpt-5.6-solsees spend logged at the old, higher price"model": "azure/gpt-5.6-sol""model": "azure/eu/gpt-5.2"and get a 400 because the model is not in the catalogAfter: the same requests are billed at the current Azure price and the eu tier resolves
"model": "azure/gpt-5.6-sol""model": "azure/eu/gpt-5.2"and it is priced at the eu data-zone rateRelevant issues
Affected release
Linear ticket
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Live proxy proof does not apply to an automated price sync. Each value was read from the provider's own price page, and the
sourcefield on every touched row points at that page. The price-check bot compared every changed row against models.dev and OpenRouter on each push; the two remaining differences are listed under Caveats and were reviewed and accepted. Per-provider sync reports are in the bot comments on this PRType
Chore
What changed
59 new rows, 278 changed rows, nothing removed, in both
model_prices_and_context_window.jsonand the backup copyPrice fields that move on existing rows:
azure/gpt-5.6-soland its eu/us tiers: input 5.00 -> 4.00, output 30.00 -> 20.00, cache read 0.50 -> 0.40, cache write 6.25 -> 5.00 per 1Mazure/gpt-4o-2024-11-20,azure/gpt-4o-mini-2024-07-18,azure/o1-mini: drop the 10% premium that had been applied to the base rowazure/eu|us/gpt-5.1*,gpt-5.4*,gpt-5.1-codex-mini: rounding fixes on input and cache read (1.38 -> 1.375 per 1M and the like)azure_ai/Llama-4-Maverick-17B-128E-Instruct-FP8: input 1.41 -> 0.25, output 0.35 -> 1.00 per 1M per the Azure retail APIazure_ai/FW-GLM-5.2-Fast,FW-Inkling,FW-Nemotron-3-Ultra-NVFP4: up 10%azure_ai/Phi-4-mini-reasoning: input 0.08 -> 0.075, output 0.32 -> 0.30 per 1Mgemini/gemini-robotics-er-2-preview: halved to 1.00 / 5.00 per 1Moutput_cost_per_audio_tokenandinput_cost_per_secondDeprecation dates:
130 of the 278 changed rows change only the
sourcefieldThe two new
gemini/gemini-3.8-live*rows carrytpm: 250000andrpm: 10, copied from the siblinggemini/gemini-3.1-flash-live-previewrow, so everygemini/chat row keeps rate limitsTests removed because they pinned catalog prices and failed once the catalog moved to the provider's current numbers:
test_azure_ai_fw_cost_per_token(Azure AI Foundry FW-* per-token prices)test_sub_input_cache_write_price_is_an_extra_saving(azure/eu/gpt-4o-2024-11-20cache write price)test_azure_gpt_5_6_cache_write_tokens_are_billed(azure/gpt-5.6-lunaliteral rates)test_azure_gpt_5_6_rates_match_azure_price_page(1.25x cache write and 1.1x data zone ratios from a dated price page)A follow up PR replaces the wider set of price pinning tests with catalog independent ones
Caveats (if any)
Medium
azure/gpt-5.6-solcache read 0.40 and cache write 5.00 per 1M; models.dev lists 0.50 and 6.25 for Azure. Azure's page states the lower values, so the PR keeps themazure_ai/Phi-4-mini-reasoningcontext stays 131072; models.dev says 128000. Not changed by this PRo1,eu/o1,us/o1: parser found no input price, left as isgemini-robotics-er-2-streaming-preview: no pricing table on page, left as isLow
azure/eu/*andazure/us/*tiers have no third-party comparison; models.dev and OpenRouter do not list data-zone tiersQA runbook
Only test deletions, no e2e files are touched. Run
uv run pytest tests/test_litellm/test_cost_calculator.py tests/test_litellm/test_utils.py -k "gemini or azure" -qto confirm the remaining cost and model info tests still pass against the new catalogFinal Attestation
Link to Devin session: https://app.devin.ai/sessions/de5c4a28cf3f4b64a6a4b56dc4b46429
Open in Devin Desktop: https://app.devin.ai/desktop/session/de5c4a28cf3f4b64a6a4b56dc4b46429?variant=devin
Note
Medium Risk
Large catalog edits directly change how LiteLLM bills and reports spend for many deployments; wrong values would misstate costs, though runtime billing logic is unchanged.
Overview
Refreshes the built-in model cost map (
model_prices_and_context_window.jsonand its backup) so proxy spend logs andcompletion_costmatch current vendor pricing across Azure, Azure AI, Gemini, OpenAI, Bedrock, Together AI, Fireworks, and Vertex (~278 updated rows, ~59 new).Catalog additions and corrections: Adds Azure EU/US data-zone keys (e.g.
azure/eu/*,azure/us/*) and new Gemini live models; updates token, cache, batch/flex/priority, and modality-specific rates (notably lower azure/gpt-5.6-sol input/output/cache prices, Azure AI Foundry fixes, Gemini TTS/live audio fields). Setsdeprecation_dateon ~30 Together AI, Fireworks, and Vertex Imagen entries. Stamps a traceablesourceURL on many rows.Tests: Removes four unit tests that asserted fixed catalog numbers (Azure AI FW-* pricing, Azure EU cache write, GPT-5.6 cache/ratio checks) so CI does not break on routine price syncs; remaining cost-calculator tests still exercise behavior without pinning vendor list prices.
Reviewed by Cursor Bugbot for commit b9ab362. Bugbot is set up for automated code reviews on this repo. Configure here.