Skip to content

chore(prices): sync prices for 5 providers: 278 models, 34 new - #40944

Closed
berriai-litellm-provider-info-sync[bot] wants to merge 1 commit into
litellm_internal_stagingfrom
litellm-providers/price-sync-litellm_internal_staging
Closed

berriai-litellm-provider-info-sync[bot] wants to merge 1 commit into
litellm_internal_stagingfrom
litellm-providers/price-sync-litellm_internal_staging

Conversation

@berriai-litellm-provider-info-sync

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Catalog prices drifted from what providers now publish
  • 34 listed models had no entry, so requests log $0
  • Several entries pointed at stale or moved pricing pages

How it solves it:

  • Syncs 278 models across 5 providers from provider sources
  • Adds the 34 newly listed models with their published prices
  • Repoints source links at each provider's current pricing endpoint

User Flow

Before: a developer calling fireworks_ai/deepseek-v4-pro gets spend figured from prices the provider no longer charges

  1. They send POST https://litellm-domain/v1/chat/completions with "model": "fireworks_ai/deepseek-v4-pro"
  2. The completion comes back 200 as usual
  3. They open https://litellm-domain/ui/?page=logs and see that request costed at $1.74 per million input tokens and $3.48 per million output tokens
  4. That is higher than the $1.20 in and $1.20 out the provider's own API reports

After: the same request is costed at the provider's current published price

  1. They send the same POST https://litellm-domain/v1/chat/completions with "model": "fireworks_ai/deepseek-v4-pro"
  2. The completion comes back 200 as usual
  3. https://litellm-domain/ui/?page=logs now shows the request at $1.20 per million input tokens and $1.20 per million output tokens
  4. Cached reads on that model now show at $0.60 per million instead of $0.145 per million

Relevant issues

Affected release

Linear ticket

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

This does not apply to an automated price sync, which only edits the price catalog JSON

Type

New Feature

Caveats (if any)

Medium

  • Together AI: 98 models skipped, no serverless token price
  • OpenAI: 7 rows skipped, long-context prices with no stated threshold, Sora size variants, and a data-sharing discount with no litellm key
  • Vertex AI: 3 rows skipped, including Gemini 2.5 Flash audio input above 200K, which has no matching field
  • Skipped rows keep whatever prices they already had

Low

  • Many Anthropic and Together AI entries only change the source link
  • OpenAI additions are mostly batch and flex price fields on existing models

QA runbook

No tests were added or changed, so this does not apply

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR
Sync report

Updates 278 models (34 new) across Anthropic, Fireworks AI, OpenAI, Together AI, Vertex AI, each read from its provider's own published prices.

Anthropic

Updates 13 models from the provider's pricing page (page d79ad2856719).

Changes

Fireworks AI

Updates 38 models (1 new) from the provider's pricing API.

Changes

OpenAI

Updates 120 models (2 new) from the provider's pricing page (page 63878741d55c).

Changes

  • babbage-002: input_cost_per_token_batches $0.2/1M, output_cost_per_token_batches $0.2/1M, source set
  • chat-latest: source https://developers.openai.com/api/docs/models/chat-latest → https://developers.openai.com/api/docs/pricing
  • chatgpt-image-latest: input_cost_per_image_token $10/1M → $8/1M, input_cost_per_token_batches $2.5/1M, output_cost_per_image_token $40/1M → $32/1M, output_cost_per_token $10/1M, output_cost_per_token_batches $5/1M, source set
  • davinci-002: input_cost_per_token_batches $1/1M, output_cost_per_token_batches $1/1M, source set
  • ft:babbage-002: input_cost_per_token_batches $0.2/1M → $0.8/1M, output_cost_per_token_batches $0.2/1M → $0.9/1M, source set
  • ft:davinci-002: input_cost_per_token_batches $1/1M → $6/1M, output_cost_per_token_batches $1/1M → $6/1M, source set
  • ft:gpt-3.5-turbo: source set
  • ft:gpt-4.1-2025-04-14: source set
  • ft:gpt-4.1-mini-2025-04-14: source set
  • ft:gpt-4.1-nano-2025-04-14: source set
  • ft:gpt-4o-2024-08-06: input_cost_per_token_batches $1.875/1M → $2.225/1M, output_cost_per_token_batches $7.5/1M → $12.5/1M, source set
  • ft:gpt-4o-mini-2024-07-18: source set
  • ft:o4-mini-2025-04-16: source set
  • gpt-3.5-turbo: source set
  • gpt-3.5-turbo-0125: input_cost_per_token_batches $0.25/1M, output_cost_per_token_batches $0.75/1M, source set
  • gpt-3.5-turbo-1106: input_cost_per_token_batches $1/1M, output_cost_per_token_batches $2/1M, source set
  • gpt-3.5-turbo-instruct: source set
  • gpt-4-0613: input_cost_per_token_batches $15/1M, output_cost_per_token_batches $30/1M, source set
  • gpt-4-turbo-2024-04-09: input_cost_per_token_batches $5/1M, output_cost_per_token_batches $15/1M, source set
  • gpt-4.1: source set
  • gpt-4.1-2025-04-14: source set
  • gpt-4.1-mini: source set
  • gpt-4.1-mini-2025-04-14: source set
  • gpt-4.1-nano: source set
  • gpt-4.1-nano-2025-04-14: source set
  • gpt-4o: source set
  • gpt-4o-2024-08-06: source set
  • gpt-4o-2024-11-20: source set
  • gpt-4o-2024-05-13: source set
  • gpt-4o-mini: source set
  • gpt-4o-mini-2024-07-18: source set
  • gpt-4o-mini-transcribe: input_cost_per_second $0.00005/s, source set
  • gpt-4o-mini-transcribe-2025-03-20: input_cost_per_second $0.00005/s, source set
  • gpt-4o-mini-transcribe-2025-12-15: input_cost_per_second $0.00005/s, source set
  • gpt-4o-mini-tts: input_cost_per_token $2.5/1M → $0.6/1M, source set
  • gpt-4o-mini-tts-2025-03-20: input_cost_per_token $2.5/1M → $0.6/1M, source set
  • gpt-4o-mini-tts-2025-12-15: input_cost_per_token $2.5/1M → $0.6/1M, source set
  • gpt-4o-transcribe: input_cost_per_second $0.0001/s, source set
  • gpt-4o-transcribe-diarize: input_cost_per_second $0.0001/s, source set
  • gpt-5: input_cost_per_token_batches $0.625/1M, output_cost_per_token_batches $5/1M, source set
  • gpt-5-2025-08-07: input_cost_per_token_batches $0.625/1M, output_cost_per_token_batches $5/1M, source set
  • gpt-5-mini: input_cost_per_token_batches $0.125/1M, output_cost_per_token_batches $1/1M, source set
  • gpt-5-mini-2025-08-07: input_cost_per_token_batches $0.125/1M, output_cost_per_token_batches $1/1M, source set
  • gpt-5-nano: input_cost_per_token_batches $0.025/1M, output_cost_per_token_batches $0.2/1M, source set
  • gpt-5-nano-2025-08-07: input_cost_per_token_batches $0.025/1M, output_cost_per_token_batches $0.2/1M, source set
  • gpt-5-pro: source set
  • gpt-5-pro-2025-10-06: source set
  • gpt-5-search-api: source set
  • gpt-5-search-api-2025-10-14: source set
  • gpt-5.1: cache_read_input_token_cost_flex $0.0625/1M, input_cost_per_token_batches $0.625/1M, input_cost_per_token_flex $0.625/1M, output_cost_per_token_batches $5/1M, output_cost_per_token_flex $5/1M, source set
  • gpt-5.1-2025-11-13: cache_read_input_token_cost_flex $0.0625/1M, input_cost_per_token_batches $0.625/1M, input_cost_per_token_flex $0.625/1M, output_cost_per_token_batches $5/1M, output_cost_per_token_flex $5/1M, source set
  • gpt-5.2: cache_read_input_token_cost_flex $0.0875/1M, input_cost_per_token_batches $0.875/1M, input_cost_per_token_flex $0.875/1M, output_cost_per_token_batches $7/1M, output_cost_per_token_flex $7/1M, source set
  • gpt-5.2-2025-12-11: cache_read_input_token_cost_flex $0.0875/1M, input_cost_per_token_batches $0.875/1M, input_cost_per_token_flex $0.875/1M, output_cost_per_token_batches $7/1M, output_cost_per_token_flex $7/1M, source set
  • gpt-5.2-pro: input_cost_per_token_batches $10.5/1M, output_cost_per_token_batches $84/1M, source set
  • gpt-5.2-pro-2025-12-11: input_cost_per_token_batches $10.5/1M, output_cost_per_token_batches $84/1M, source set
  • gpt-5.3-codex: source set
  • gpt-5.4: source set
  • gpt-5.4-2026-03-05: source set
  • gpt-5.4-mini: source set
  • gpt-5.4-mini-2026-03-17: source set
  • …and 60 more

Skipped by the parser (7)

  • gpt-6-astra: long-context prices present but the row states no threshold
  • gpt-5.6-sol: long-context prices present but the row states no threshold
  • gpt-5.6-terra: long-context prices present but the row states no threshold
  • gpt-5.6-luna: long-context prices present but the row states no threshold
  • sora-2-pro 1024p: size variant has no litellm key stated on the page
  • sora-2-pro 1080p: size variant has no litellm key stated on the page
  • ft:o4-mini-2025-04-16: data-sharing discount has no litellm key

Together AI

Updates 74 models (22 new) from the provider's pricing API.

Changes

Skipped by the parser (98)

  • Prism-ML/Ternary-Bonsai-27B: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • nim/meta/llama-3.2-11b-vision-instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • nim/meta/llama-3.2-90b-vision-instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • nim/mistralai/mixtral-8x22b-instruct-v01: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • nim/meta/llama-3.3-70b-instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • nim/nvidia/llama-3.1-nemotron-70b-instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • nim/meta/llama-3.1-8b-instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • nim/meta/llama-3.1-70b-instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • nim/nv-mistralai/mistral-nemo-12b-instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • nim/mistralai/mixtral-8x7b-instruct-v01: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • nim/nvidia/llama-3.3-nemotron-super-49b-v1: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • google/gemma-3-27b-it: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • meta-llama/Llama-2-7b-chat-hf: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • deepseek-ai/DeepSeek-R1-Distill-Qwen-7B: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • google/gemma-3-1b-it: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • google/gemma-3-4b-it: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • deepcogito/cogito-v1-preview-llama-8B: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • deepcogito/cogito-v1-preview-qwen-32B: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • deepcogito/cogito-v1-preview-qwen-14B: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • deepcogito/cogito-v1-preview-llama-70B: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • deepcogito/cogito-v1-preview-llama-70B-Turbo: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • meta-llama/Llama-3.3-70B-Instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • Qwen/Qwen2.5-32B: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • Qwen/Qwen2.5-72B: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • Qwen/Qwen2.5-3B-Instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • Qwen/Qwen2.5-1.5B-Instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • Qwen/Qwen2.5-14B: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • Qwen/Qwen2.5-7B: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • Qwen/Qwen2.5-1.5B: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • meta-llama/Meta-Llama-3.1-70B: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • meta-llama/Llama-3.2-1B: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • Qwen/Qwen2.5-7B-Instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • Qwen/Qwen2.5-32B-Instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • meta-llama/Llama-3.1-405B: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • agentica-org/DeepCoder-14B-Preview: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • mistralai/Mistral-7B-v0.1: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • mistralai/Devstral-Small-2505: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • mistralai/Mixtral-8x22B-Instruct-v0.1: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • allenai/Molmo-7B-D-0924: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • Qwen/Qwen3-8B: no serverless token price (pricing.input and pricing.output are 0 or absent)
  • …and 58 more

Vertex AI

Updates 33 models (9 new) from the provider's pricing page (page be62e269e1d8).

Changes

Skipped by the parser (3)

  • gemini-2.5-flash: "Audio Input" above 200K is $0.30, not the base price; litellm has no field for it
  • Gemini 2.0 Flash Image Generation: litellm keys only the free experimental model (gemini-2.0-flash-exp-image-generation)
  • Gemini 2.0 Flash Live API: litellm has no vertex_ai/ key for the 2.0 Live API

Opened by the litellm-providers price sync. Every price is read from the provider's own published source and gated before it is applied; the audit trail for each value is in the portal's sync history.

fireworks_ai/accounts/fireworks/models/deepseek-v4-flash-0731: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/deepseek-v4-flash-0731: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/deepseek-v4-flash-vision-exp: 
fireworks_ai/deepseek-v4-flash-vision-exp: 
fireworks_ai/accounts/fireworks/models/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/deepseek-v4-pro-0813: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/deepseek-v4p1-flash: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/deepseek-v4p1-flash: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/glm-5p2: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/glm-5p2: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/glm-5p3: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/glm-5p3-flash: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/gpt-oss-120b: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/gpt-oss-120b: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/kimi-k2p6: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/kimi-k2p6: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/kimi-k2p7-code: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/kimi-k2p7-code: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/kimi-k3: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/kimi-k3: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/minimax-m2p7: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/minimax-m2p7: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/minimax-m3: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/minimax-m3: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/models/muse-glimmer-30b: 
fireworks_ai/muse-glimmer-30b: 
fireworks_ai/accounts/fireworks/models/nemotron-3-ultra-nvfp4: 
fireworks_ai/nemotron-3-ultra-nvfp4: 
fireworks_ai/accounts/fireworks/models/qwen3-embedding-8b: 
fireworks_ai/accounts/fireworks/models/qwen3-reranker-8b: input_cost_per_token
fireworks_ai/accounts/fireworks/models/qwen3p7-plus: 
fireworks_ai/qwen3p7-plus: 
fireworks_ai/accounts/fireworks/models/qwen3p8-max: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/qwen3p8-max: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority
fireworks_ai/accounts/fireworks/routers/glm-5p2-fast: 
fireworks_ai/accounts/fireworks/routers/glm-5p3-fast: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
fireworks_ai/accounts/fireworks/routers/kimi-k3-fast: 
together_ai/arcee-ai/trinity-mini: input_cost_per_token, output_cost_per_token
together_ai/arize-ai/qwen-2-1.5b-instruct: 
babbage-002: input_cost_per_token_batches, output_cost_per_token_batches
chat-latest: 
chatgpt-image-latest: output_cost_per_token, input_cost_per_image_token, output_cost_per_image_token, input_cost_per_token_batches, output_cost_per_token_batches
claude-fable-5: 
claude-fable-5-1: 
claude-haiku-4-5: 
claude-mythos-5: 
claude-mythos-5-1: 
claude-opus-4-5: 
claude-opus-4-6: 
claude-opus-4-7: 
claude-opus-4-8: 
claude-opus-5: 
claude-sonnet-4-5: 
claude-sonnet-4-6: 
claude-sonnet-5: 
davinci-002: input_cost_per_token_batches, output_cost_per_token_batches
deep-research-pro-preview-12-2025: cache_read_input_token_cost
together_ai/deepseek-ai/deepseek-coder-33b-instruct: input_cost_per_token, output_cost_per_token
together_ai/deepseek-ai/DeepSeek-R1-0528:
@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@greptile-apps

greptile-apps Bot commented Sep 13, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR synchronizes model pricing and source metadata across Anthropic, Fireworks AI, OpenAI, Together AI, and Vertex AI

  • Updates pricing metadata for 278 models
  • Adds 34 newly listed models
  • Adds batch, flex, priority, cache, and modality-specific price fields
  • Keeps the canonical and bundled fallback catalogs byte-for-byte synchronized

Confidence Score: 5/5

The catalog-only update appears safe to merge, with no concrete pricing lookup, cost calculation, or synchronization defect identified

Both catalogs remain identical, new entries contain the required lookup metadata, and the investigated pricing-field combinations did not establish an incorrect runtime cost path

Important Files Changed

Filename Overview
model_prices_and_context_window.json Updates the canonical model catalog with provider pricing, pricing tiers, source links, and 34 new model entries
litellm/model_prices_and_context_window_backup.json Mirrors the canonical catalog exactly for bundled and offline fallback loading

Reviews (1): Last reviewed commit: "chore(prices): sync prices for 5 provide..." | Re-trigger Greptile

@codspeed

codspeed Bot commented Sep 13, 2026

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm-providers/price-sync-litellm_internal_staging (3b37355) with litellm_internal_staging (c2c2a62)

Open in CodSpeed

@codecov

codecov Bot commented Sep 13, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@devin-ai-integration

Copy link
Copy Markdown
Contributor

Price check against models.dev and OpenRouter for head 3b37355: 46 models with changed prices, 12 with mismatches, 33 listed in neither catalog. Details in #litellm-provider-info-sync.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants