Repository navigation
chore(prices): sync prices for 5 providers: 278 models, 34 new - #40944
berriai-litellm-provider-info-sync[bot] wants to merge 1 commit into
Conversation
fireworks_ai/accounts/fireworks/models/deepseek-v4-flash-0731: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/deepseek-v4-flash-0731: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/accounts/fireworks/models/deepseek-v4-flash-vision-exp: fireworks_ai/deepseek-v4-flash-vision-exp: fireworks_ai/accounts/fireworks/models/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/accounts/fireworks/models/deepseek-v4-pro-0813: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/accounts/fireworks/models/deepseek-v4p1-flash: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/deepseek-v4p1-flash: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/accounts/fireworks/models/glm-5p2: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/glm-5p2: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/accounts/fireworks/models/glm-5p3: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/accounts/fireworks/models/glm-5p3-flash: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/accounts/fireworks/models/gpt-oss-120b: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/gpt-oss-120b: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/accounts/fireworks/models/kimi-k2p6: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/kimi-k2p6: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/accounts/fireworks/models/kimi-k2p7-code: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/kimi-k2p7-code: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/accounts/fireworks/models/kimi-k3: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/kimi-k3: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/accounts/fireworks/models/minimax-m2p7: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/minimax-m2p7: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/accounts/fireworks/models/minimax-m3: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/minimax-m3: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/accounts/fireworks/models/muse-glimmer-30b: fireworks_ai/muse-glimmer-30b: fireworks_ai/accounts/fireworks/models/nemotron-3-ultra-nvfp4: fireworks_ai/nemotron-3-ultra-nvfp4: fireworks_ai/accounts/fireworks/models/qwen3-embedding-8b: fireworks_ai/accounts/fireworks/models/qwen3-reranker-8b: input_cost_per_token fireworks_ai/accounts/fireworks/models/qwen3p7-plus: fireworks_ai/qwen3p7-plus: fireworks_ai/accounts/fireworks/models/qwen3p8-max: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/qwen3p8-max: input_cost_per_token_priority, output_cost_per_token_priority, cache_read_input_token_cost_priority fireworks_ai/accounts/fireworks/routers/glm-5p2-fast: fireworks_ai/accounts/fireworks/routers/glm-5p3-fast: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost fireworks_ai/accounts/fireworks/routers/kimi-k3-fast: together_ai/arcee-ai/trinity-mini: input_cost_per_token, output_cost_per_token together_ai/arize-ai/qwen-2-1.5b-instruct: babbage-002: input_cost_per_token_batches, output_cost_per_token_batches chat-latest: chatgpt-image-latest: output_cost_per_token, input_cost_per_image_token, output_cost_per_image_token, input_cost_per_token_batches, output_cost_per_token_batches claude-fable-5: claude-fable-5-1: claude-haiku-4-5: claude-mythos-5: claude-mythos-5-1: claude-opus-4-5: claude-opus-4-6: claude-opus-4-7: claude-opus-4-8: claude-opus-5: claude-sonnet-4-5: claude-sonnet-4-6: claude-sonnet-5: davinci-002: input_cost_per_token_batches, output_cost_per_token_batches deep-research-pro-preview-12-2025: cache_read_input_token_cost together_ai/deepseek-ai/deepseek-coder-33b-instruct: input_cost_per_token, output_cost_per_token together_ai/deepseek-ai/DeepSeek-R1-0528:
|
|
Greptile SummaryThis PR synchronizes model pricing and source metadata across Anthropic, Fireworks AI, OpenAI, Together AI, and Vertex AI
Confidence Score: 5/5The catalog-only update appears safe to merge, with no concrete pricing lookup, cost calculation, or synchronization defect identified Both catalogs remain identical, new entries contain the required lookup metadata, and the investigated pricing-field combinations did not establish an incorrect runtime cost path
|
| Filename | Overview |
|---|---|
| model_prices_and_context_window.json | Updates the canonical model catalog with provider pricing, pricing tiers, source links, and 34 new model entries |
| litellm/model_prices_and_context_window_backup.json | Mirrors the canonical catalog exactly for bundled and offline fallback loading |
Reviews (1): Last reviewed commit: "chore(prices): sync prices for 5 provide..." | Re-trigger Greptile
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
|
Price check against models.dev and OpenRouter for head 3b37355: 46 models with changed prices, 12 with mismatches, 33 listed in neither catalog. Details in #litellm-provider-info-sync. |
TLDR
Problem this solves:
How it solves it:
User Flow
Before: a developer calling
fireworks_ai/deepseek-v4-progets spend figured from prices the provider no longer charges"model": "fireworks_ai/deepseek-v4-pro"After: the same request is costed at the provider's current published price
"model": "fireworks_ai/deepseek-v4-pro"Relevant issues
Affected release
Linear ticket
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
This does not apply to an automated price sync, which only edits the price catalog JSON
Type
New Feature
Caveats (if any)
Medium
Low
QA runbook
No tests were added or changed, so this does not apply
Final Attestation
Sync report
Updates 278 models (34 new) across Anthropic, Fireworks AI, OpenAI, Together AI, Vertex AI, each read from its provider's own published prices.
Anthropic
Updates 13 models from the provider's pricing page (page
d79ad2856719).Changes
claude-fable-5: source https://docs.anthropic.com/en/docs/about-claude/models/overview → https://platform.claude.com/docs/en/about-claude/pricingclaude-fable-5-1: source https://platform.claude.com/docs/en/models/fable-5-1/overview → https://platform.claude.com/docs/en/about-claude/pricingclaude-haiku-4-5: source setclaude-mythos-5: source https://platform.claude.com/docs/en/about-claude/models/overview → https://platform.claude.com/docs/en/about-claude/pricingclaude-mythos-5-1: source https://platform.claude.com/docs/en/models/mythos-5-1/overview → https://platform.claude.com/docs/en/about-claude/pricingclaude-opus-4-5: source setclaude-opus-4-6: source setclaude-opus-4-7: source setclaude-opus-4-8: source setclaude-opus-5: source https://docs.anthropic.com/en/docs/about-claude/models/overview → https://platform.claude.com/docs/en/about-claude/pricingclaude-sonnet-4-5: source setclaude-sonnet-4-6: source setclaude-sonnet-5: source https://docs.anthropic.com/en/docs/about-claude/models/overview → https://platform.claude.com/docs/en/about-claude/pricingFireworks AI
Updates 38 models (1 new) from the provider's pricing API.
Changes
fireworks_ai/accounts/fireworks/models/deepseek-v4-flash-0731: cache_read_input_token_cost_priority $0.00875/1M, input_cost_per_token_priority $0.275/1M, output_cost_per_token_priority $0.825/1M, source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/deepseek-v4-flash-0731: cache_read_input_token_cost_priority $0.00875/1M, input_cost_per_token_priority $0.275/1M, output_cost_per_token_priority $0.825/1M, source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/accounts/fireworks/models/deepseek-v4-flash-vision-exp: source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/deepseek-v4-flash-vision-exp: source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/accounts/fireworks/models/deepseek-v4-pro: cache_read_input_token_cost $0.145/1M → $0.6/1M, cache_read_input_token_cost_priority $0.6/1M, input_cost_per_token $1.74/1M → $1.2/1M, input_cost_per_token_priority $1.2/1M, output_cost_per_token $3.48/1M → $1.2/1M, output_cost_per_token_priority $1.2/1M, source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/deepseek-v4-pro: cache_read_input_token_cost $0.145/1M → $0.6/1M, cache_read_input_token_cost_priority $0.6/1M, input_cost_per_token $1.74/1M → $1.2/1M, input_cost_per_token_priority $1.2/1M, output_cost_per_token $3.48/1M → $1.2/1M, output_cost_per_token_priority $1.2/1M, source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/accounts/fireworks/models/deepseek-v4-pro-0813: cache_read_input_token_cost_priority $0.055/1M, input_cost_per_token_priority $1.65/1M, output_cost_per_token_priority $4.95/1M, source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/accounts/fireworks/models/deepseek-v4p1-flash: cache_read_input_token_cost_priority $0.00875/1M, input_cost_per_token_priority $0.275/1M, output_cost_per_token_priority $0.825/1M, source https://fireworks.ai/models/deepseek-ai/deepseek-v4p1-flash → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/deepseek-v4p1-flash: cache_read_input_token_cost_priority $0.00875/1M, input_cost_per_token_priority $0.275/1M, output_cost_per_token_priority $0.825/1M, source https://fireworks.ai/models/deepseek-ai/deepseek-v4p1-flash → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/accounts/fireworks/models/glm-5p2: cache_read_input_token_cost_priority $0.175/1M, input_cost_per_token_priority $1.75/1M, output_cost_per_token_priority $5.5/1M, source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/glm-5p2: cache_read_input_token_cost_priority $0.175/1M, input_cost_per_token_priority $1.75/1M, output_cost_per_token_priority $5.5/1M, source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/accounts/fireworks/models/glm-5p3: cache_read_input_token_cost_priority $0.325/1M, input_cost_per_token_priority $1.75/1M, output_cost_per_token_priority $5.5/1M, source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/accounts/fireworks/models/glm-5p3-flash: cache_read_input_token_cost_priority $0.0375/1M, input_cost_per_token_priority $0.1875/1M, output_cost_per_token_priority $0.625/1M, source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/accounts/fireworks/models/gpt-oss-120b: cache_read_input_token_cost_priority $0.018/1M, input_cost_per_token_priority $0.18/1M, output_cost_per_token_priority $0.72/1M, source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/gpt-oss-120b: cache_read_input_token_cost_priority $0.018/1M, input_cost_per_token_priority $0.18/1M, output_cost_per_token_priority $0.72/1M, source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/accounts/fireworks/models/kimi-k2p6: cache_read_input_token_cost_priority $0.22/1M, input_cost_per_token_priority $1.5/1M, output_cost_per_token_priority $6/1M, source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/kimi-k2p6: cache_read_input_token_cost_priority $0.22/1M, input_cost_per_token_priority $1.5/1M, output_cost_per_token_priority $6/1M, source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/accounts/fireworks/models/kimi-k2p7-code: cache_read_input_token_cost_priority $0.285/1M, input_cost_per_token_priority $1.425/1M, output_cost_per_token_priority $6/1M, source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/kimi-k2p7-code: cache_read_input_token_cost_priority $0.285/1M, input_cost_per_token_priority $1.425/1M, output_cost_per_token_priority $6/1M, source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/accounts/fireworks/models/kimi-k3: cache_read_input_token_cost_priority $0.375/1M, input_cost_per_token_priority $3.75/1M, output_cost_per_token_priority $18.75/1M, source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/kimi-k3: cache_read_input_token_cost_priority $0.375/1M, input_cost_per_token_priority $3.75/1M, output_cost_per_token_priority $18.75/1M, source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/accounts/fireworks/models/minimax-m2p7: cache_read_input_token_cost_priority $0.6/1M, input_cost_per_token_priority $1.2/1M, output_cost_per_token_priority $1.2/1M, source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/minimax-m2p7: cache_read_input_token_cost_priority $0.6/1M, input_cost_per_token_priority $1.2/1M, output_cost_per_token_priority $1.2/1M, source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/accounts/fireworks/models/minimax-m3: cache_read_input_token_cost_priority $0.09/1M, input_cost_per_token_priority $0.45/1M, output_cost_per_token_priority $1.8/1M, source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/minimax-m3: cache_read_input_token_cost_priority $0.09/1M, input_cost_per_token_priority $0.45/1M, output_cost_per_token_priority $1.8/1M, source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/accounts/fireworks/models/muse-glimmer-30b: source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/muse-glimmer-30b: source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/accounts/fireworks/models/nemotron-3-ultra-nvfp4: source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/nemotron-3-ultra-nvfp4: source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/accounts/fireworks/models/qwen3-embedding-8b: source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/accounts/fireworks/models/qwen3-reranker-8b: input_cost_per_token $0/1M → $0.2/1M, source setfireworks_ai/accounts/fireworks/models/qwen3p7-plus: source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/qwen3p7-plus: source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/accounts/fireworks/models/qwen3p8-max: cache_read_input_token_cost_priority $0.375/1M, input_cost_per_token_priority $3/1M, output_cost_per_token_priority $9/1M, source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/qwen3p8-max: cache_read_input_token_cost_priority $0.375/1M, input_cost_per_token_priority $3/1M, output_cost_per_token_priority $9/1M, source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/accounts/fireworks/routers/glm-5p2-fast: source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsfireworks_ai/accounts/fireworks/routers/glm-5p3-fast(new): cache_read_input_token_cost $0.39/1M, input_cost_per_token $2.1/1M, output_cost_per_token $6.6/1Mfireworks_ai/accounts/fireworks/routers/kimi-k3-fast: source https://docs.fireworks.ai/serverless/pricing → https://api.fireworks.ai/v1/serverless/modelsOpenAI
Updates 120 models (2 new) from the provider's pricing page (page
63878741d55c).Changes
babbage-002: input_cost_per_token_batches $0.2/1M, output_cost_per_token_batches $0.2/1M, source setchat-latest: source https://developers.openai.com/api/docs/models/chat-latest → https://developers.openai.com/api/docs/pricingchatgpt-image-latest: input_cost_per_image_token $10/1M → $8/1M, input_cost_per_token_batches $2.5/1M, output_cost_per_image_token $40/1M → $32/1M, output_cost_per_token $10/1M, output_cost_per_token_batches $5/1M, source setdavinci-002: input_cost_per_token_batches $1/1M, output_cost_per_token_batches $1/1M, source setft:babbage-002: input_cost_per_token_batches $0.2/1M → $0.8/1M, output_cost_per_token_batches $0.2/1M → $0.9/1M, source setft:davinci-002: input_cost_per_token_batches $1/1M → $6/1M, output_cost_per_token_batches $1/1M → $6/1M, source setft:gpt-3.5-turbo: source setft:gpt-4.1-2025-04-14: source setft:gpt-4.1-mini-2025-04-14: source setft:gpt-4.1-nano-2025-04-14: source setft:gpt-4o-2024-08-06: input_cost_per_token_batches $1.875/1M → $2.225/1M, output_cost_per_token_batches $7.5/1M → $12.5/1M, source setft:gpt-4o-mini-2024-07-18: source setft:o4-mini-2025-04-16: source setgpt-3.5-turbo: source setgpt-3.5-turbo-0125: input_cost_per_token_batches $0.25/1M, output_cost_per_token_batches $0.75/1M, source setgpt-3.5-turbo-1106: input_cost_per_token_batches $1/1M, output_cost_per_token_batches $2/1M, source setgpt-3.5-turbo-instruct: source setgpt-4-0613: input_cost_per_token_batches $15/1M, output_cost_per_token_batches $30/1M, source setgpt-4-turbo-2024-04-09: input_cost_per_token_batches $5/1M, output_cost_per_token_batches $15/1M, source setgpt-4.1: source setgpt-4.1-2025-04-14: source setgpt-4.1-mini: source setgpt-4.1-mini-2025-04-14: source setgpt-4.1-nano: source setgpt-4.1-nano-2025-04-14: source setgpt-4o: source setgpt-4o-2024-08-06: source setgpt-4o-2024-11-20: source setgpt-4o-2024-05-13: source setgpt-4o-mini: source setgpt-4o-mini-2024-07-18: source setgpt-4o-mini-transcribe: input_cost_per_second $0.00005/s, source setgpt-4o-mini-transcribe-2025-03-20: input_cost_per_second $0.00005/s, source setgpt-4o-mini-transcribe-2025-12-15: input_cost_per_second $0.00005/s, source setgpt-4o-mini-tts: input_cost_per_token $2.5/1M → $0.6/1M, source setgpt-4o-mini-tts-2025-03-20: input_cost_per_token $2.5/1M → $0.6/1M, source setgpt-4o-mini-tts-2025-12-15: input_cost_per_token $2.5/1M → $0.6/1M, source setgpt-4o-transcribe: input_cost_per_second $0.0001/s, source setgpt-4o-transcribe-diarize: input_cost_per_second $0.0001/s, source setgpt-5: input_cost_per_token_batches $0.625/1M, output_cost_per_token_batches $5/1M, source setgpt-5-2025-08-07: input_cost_per_token_batches $0.625/1M, output_cost_per_token_batches $5/1M, source setgpt-5-mini: input_cost_per_token_batches $0.125/1M, output_cost_per_token_batches $1/1M, source setgpt-5-mini-2025-08-07: input_cost_per_token_batches $0.125/1M, output_cost_per_token_batches $1/1M, source setgpt-5-nano: input_cost_per_token_batches $0.025/1M, output_cost_per_token_batches $0.2/1M, source setgpt-5-nano-2025-08-07: input_cost_per_token_batches $0.025/1M, output_cost_per_token_batches $0.2/1M, source setgpt-5-pro: source setgpt-5-pro-2025-10-06: source setgpt-5-search-api: source setgpt-5-search-api-2025-10-14: source setgpt-5.1: cache_read_input_token_cost_flex $0.0625/1M, input_cost_per_token_batches $0.625/1M, input_cost_per_token_flex $0.625/1M, output_cost_per_token_batches $5/1M, output_cost_per_token_flex $5/1M, source setgpt-5.1-2025-11-13: cache_read_input_token_cost_flex $0.0625/1M, input_cost_per_token_batches $0.625/1M, input_cost_per_token_flex $0.625/1M, output_cost_per_token_batches $5/1M, output_cost_per_token_flex $5/1M, source setgpt-5.2: cache_read_input_token_cost_flex $0.0875/1M, input_cost_per_token_batches $0.875/1M, input_cost_per_token_flex $0.875/1M, output_cost_per_token_batches $7/1M, output_cost_per_token_flex $7/1M, source setgpt-5.2-2025-12-11: cache_read_input_token_cost_flex $0.0875/1M, input_cost_per_token_batches $0.875/1M, input_cost_per_token_flex $0.875/1M, output_cost_per_token_batches $7/1M, output_cost_per_token_flex $7/1M, source setgpt-5.2-pro: input_cost_per_token_batches $10.5/1M, output_cost_per_token_batches $84/1M, source setgpt-5.2-pro-2025-12-11: input_cost_per_token_batches $10.5/1M, output_cost_per_token_batches $84/1M, source setgpt-5.3-codex: source setgpt-5.4: source setgpt-5.4-2026-03-05: source setgpt-5.4-mini: source setgpt-5.4-mini-2026-03-17: source setSkipped by the parser (7)
gpt-6-astra: long-context prices present but the row states no thresholdgpt-5.6-sol: long-context prices present but the row states no thresholdgpt-5.6-terra: long-context prices present but the row states no thresholdgpt-5.6-luna: long-context prices present but the row states no thresholdsora-2-pro 1024p: size variant has no litellm key stated on the pagesora-2-pro 1080p: size variant has no litellm key stated on the pageft:o4-mini-2025-04-16: data-sharing discount has no litellm keyTogether AI
Updates 74 models (22 new) from the provider's pricing API.
Changes
together_ai/arcee-ai/trinity-mini(new): input_cost_per_token $0.045/1M, output_cost_per_token $0.15/1Mtogether_ai/arize-ai/qwen-2-1.5b-instruct: source https://docs.together.ai/docs/serverless-models → https://api.together.ai/v1/modelstogether_ai/deepseek-ai/deepseek-coder-33b-instruct(new): input_cost_per_token $0.8/1M, output_cost_per_token $0.8/1Mtogether_ai/deepseek-ai/DeepSeek-R1-0528: source https://api.together.xyz/v1/models → https://api.together.ai/v1/modelstogether_ai/deepseek-ai/DeepSeek-R1-Distill-Llama-70B(new): input_cost_per_token $2/1M, output_cost_per_token $2/1Mtogether_ai/deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B(new): input_cost_per_token $0.18/1M, output_cost_per_token $0.18/1Mtogether_ai/deepseek-ai/DeepSeek-R1-Distill-Qwen-14B(new): input_cost_per_token $1.6/1M, output_cost_per_token $1.6/1Mtogether_ai/deepseek-ai/DeepSeek-V3.1: source https://www.together.ai/models/deepseek-v3-1 → https://api.together.ai/v1/modelstogether_ai/deepseek-ai/DeepSeek-V4-Flash-0731: source https://docs.together.ai/docs/serverless-models → https://api.together.ai/v1/modelstogether_ai/deepseek-ai/DeepSeek-V4-Pro-0813: source https://docs.together.ai/docs/serverless-models → https://api.together.ai/v1/modelstogether_ai/deepseek-ai/DeepSeek-V4.1-Flash(new): cache_read_input_token_cost $0.006/1M, input_cost_per_token $0.3/1M, output_cost_per_token $1.2/1Mtogether_ai/google/gemma-2-27b-it(new): input_cost_per_token $0.8/1M, output_cost_per_token $0.8/1Mtogether_ai/google/gemma-4-31B-it: source https://docs.together.ai/docs/serverless-models → https://api.together.ai/v1/modelstogether_ai/meta-llama/Llama-3-8b-chat-hf(new): input_cost_per_token $0.2/1M, output_cost_per_token $0.2/1Mtogether_ai/meta-llama/Llama-3.1-405B-Instruct(new): input_cost_per_token $3.5/1M, output_cost_per_token $3.5/1Mtogether_ai/meta-llama/Llama-3.2-1B-Instruct(new): input_cost_per_token $0.06/1M, output_cost_per_token $0.06/1Mtogether_ai/meta-llama/Llama-3.2-3B-Instruct(new): input_cost_per_token $0.06/1M, output_cost_per_token $0.06/1Mtogether_ai/meta-llama/Llama-3.3-70B-Instruct-Turbo: source https://docs.together.ai/docs/serverless-models → https://api.together.ai/v1/modelstogether_ai/meta-llama/Llama-4-Scout-17B-16E-Instruct: source settogether_ai/meta-llama/Meta-Llama-3-70B-Instruct-Turbo(new): input_cost_per_token $0.88/1M, output_cost_per_token $0.88/1Mtogether_ai/meta-llama/Meta-Llama-3-8B-Instruct(new): input_cost_per_token $0.2/1M, output_cost_per_token $0.2/1Mtogether_ai/meta-llama/Meta-Llama-3.1-70B-Instruct-Turbo: source settogether_ai/meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo: source settogether_ai/meta-models/Muse-Glimmer-30B: source https://docs.together.ai/docs/serverless-models → https://api.together.ai/v1/modelstogether_ai/MiniMaxAI/MiniMax-M2.7: source https://api.together.xyz/v1/models → https://api.together.ai/v1/modelstogether_ai/MiniMaxAI/MiniMax-M3: source https://docs.together.ai/docs/serverless-models → https://api.together.ai/v1/modelstogether_ai/mistralai/Ministral-3-14B-Instruct-2512: source https://api.together.xyz/v1/models → https://api.together.ai/v1/modelstogether_ai/mistralai/Mistral-7B-Instruct-v0.1: source https://api.together.xyz/v1/models → https://api.together.ai/v1/modelstogether_ai/mistralai/Mistral-7B-Instruct-v0.3: source https://api.together.xyz/v1/models → https://api.together.ai/v1/modelstogether_ai/mistralai/Mistral-Small-24B-Instruct-2501: source https://api.together.xyz/v1/models → https://api.together.ai/v1/modelstogether_ai/mistralai/Mixtral-8x7B-Instruct-v0.1: source settogether_ai/moonshotai/Kimi-K2.5-fp4: source https://api.together.xyz/v1/models → https://api.together.ai/v1/modelstogether_ai/moonshotai/Kimi-K2.6: source https://api.together.xyz/v1/models → https://api.together.ai/v1/modelstogether_ai/moonshotai/Kimi-K2.7-Code: source https://docs.together.ai/docs/serverless-models → https://api.together.ai/v1/modelstogether_ai/moonshotai/Kimi-K3: source https://docs.together.ai/docs/serverless-models → https://api.together.ai/v1/modelstogether_ai/NousResearch/Nous-Hermes-2-Mixtral-8x7B-DPO(new): input_cost_per_token $0.6/1M, output_cost_per_token $0.6/1Mtogether_ai/nvidia/Llama-3.1-Nemotron-70B-Instruct-HF(new): input_cost_per_token $0.88/1M, output_cost_per_token $0.88/1Mtogether_ai/nvidia/nemotron-3-ultra-550b-a55b: source https://docs.together.ai/docs/serverless-models → https://api.together.ai/v1/modelstogether_ai/nvidia/NVIDIA-Nemotron-Nano-9B-v2: source https://api.together.xyz/v1/models → https://api.together.ai/v1/modelstogether_ai/openai/gpt-oss-120b: source https://www.together.ai/models/gpt-oss-120b → https://api.together.ai/v1/modelstogether_ai/openai/gpt-oss-20b: source https://www.together.ai/models/gpt-oss-20b → https://api.together.ai/v1/modelstogether_ai/Qwen/Qwen2-1.5B-Instruct(new): input_cost_per_token $0.02/1M, output_cost_per_token $0.02/1Mtogether_ai/Qwen/Qwen2-72B-Instruct(new): input_cost_per_token $0.9/1M, output_cost_per_token $0.9/1Mtogether_ai/Qwen/Qwen2-VL-72B-Instruct(new): input_cost_per_token $1.2/1M, output_cost_per_token $1.2/1Mtogether_ai/Qwen/Qwen2.5-14B-Instruct(new): input_cost_per_token $0.8/1M, output_cost_per_token $0.8/1Mtogether_ai/Qwen/Qwen2.5-72B-Instruct(new): input_cost_per_token $1.2/1M, output_cost_per_token $1.2/1Mtogether_ai/Qwen/Qwen2.5-72B-Instruct-Turbo: source https://api.together.xyz/v1/models → https://api.together.ai/v1/modelstogether_ai/Qwen/Qwen2.5-7B-Instruct-Turbo: source https://api.together.xyz/v1/models → https://api.together.ai/v1/modelstogether_ai/Qwen/Qwen2.5-Coder-32B-Instruct(new): input_cost_per_token $0.8/1M, output_cost_per_token $0.8/1Mtogether_ai/Qwen/Qwen2.5-VL-72B-Instruct(new): input_cost_per_token $1.95/1M, output_cost_per_token $8/1Mtogether_ai/Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8: source https://www.together.ai/models/qwen3-coder-480b-a35b-instruct → https://api.together.ai/v1/modelstogether_ai/Qwen/Qwen3-Coder-Next-FP8: source https://api.together.xyz/v1/models → https://api.together.ai/v1/modelstogether_ai/Qwen/Qwen3-Next-80B-A3B-Instruct: source https://www.together.ai/models/qwen3-next-80b-a3b-instruct → https://api.together.ai/v1/modelstogether_ai/Qwen/Qwen3-Next-80B-A3B-Thinking: source https://www.together.ai/models/qwen3-next-80b-a3b-thinking → https://api.together.ai/v1/modelstogether_ai/Qwen/Qwen3-VL-32B-Instruct: source https://api.together.xyz/v1/models → https://api.together.ai/v1/modelstogether_ai/Qwen/Qwen3-VL-8B-Instruct: source https://api.together.xyz/v1/models → https://api.together.ai/v1/modelstogether_ai/Qwen/Qwen3.5-397B-A17B: source https://www.together.ai/models/qwen3-5-397b-a17b → https://api.together.ai/v1/modelstogether_ai/Qwen/Qwen3.5-9B: source https://docs.together.ai/docs/serverless-models → https://api.together.ai/v1/modelstogether_ai/Qwen/Qwen3.6-Plus: source https://docs.together.ai/docs/serverless-models → https://api.together.ai/v1/modelstogether_ai/Qwen/Qwen3.7-Max: source https://docs.together.ai/docs/serverless-models → https://api.together.ai/v1/modelsSkipped by the parser (98)
Prism-ML/Ternary-Bonsai-27B: no serverless token price (pricing.input and pricing.output are 0 or absent)nim/meta/llama-3.2-11b-vision-instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)nim/meta/llama-3.2-90b-vision-instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)nim/mistralai/mixtral-8x22b-instruct-v01: no serverless token price (pricing.input and pricing.output are 0 or absent)nim/meta/llama-3.3-70b-instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)nim/nvidia/llama-3.1-nemotron-70b-instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)nim/meta/llama-3.1-8b-instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)nim/meta/llama-3.1-70b-instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)nim/nv-mistralai/mistral-nemo-12b-instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)nim/mistralai/mixtral-8x7b-instruct-v01: no serverless token price (pricing.input and pricing.output are 0 or absent)nim/nvidia/llama-3.3-nemotron-super-49b-v1: no serverless token price (pricing.input and pricing.output are 0 or absent)google/gemma-3-27b-it: no serverless token price (pricing.input and pricing.output are 0 or absent)meta-llama/Llama-2-7b-chat-hf: no serverless token price (pricing.input and pricing.output are 0 or absent)deepseek-ai/DeepSeek-R1-Distill-Qwen-7B: no serverless token price (pricing.input and pricing.output are 0 or absent)google/gemma-3-1b-it: no serverless token price (pricing.input and pricing.output are 0 or absent)google/gemma-3-4b-it: no serverless token price (pricing.input and pricing.output are 0 or absent)deepcogito/cogito-v1-preview-llama-8B: no serverless token price (pricing.input and pricing.output are 0 or absent)deepcogito/cogito-v1-preview-qwen-32B: no serverless token price (pricing.input and pricing.output are 0 or absent)deepcogito/cogito-v1-preview-qwen-14B: no serverless token price (pricing.input and pricing.output are 0 or absent)deepcogito/cogito-v1-preview-llama-70B: no serverless token price (pricing.input and pricing.output are 0 or absent)deepcogito/cogito-v1-preview-llama-70B-Turbo: no serverless token price (pricing.input and pricing.output are 0 or absent)meta-llama/Llama-3.3-70B-Instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)Qwen/Qwen2.5-32B: no serverless token price (pricing.input and pricing.output are 0 or absent)Qwen/Qwen2.5-72B: no serverless token price (pricing.input and pricing.output are 0 or absent)Qwen/Qwen2.5-3B-Instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)Qwen/Qwen2.5-1.5B-Instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)Qwen/Qwen2.5-14B: no serverless token price (pricing.input and pricing.output are 0 or absent)Qwen/Qwen2.5-7B: no serverless token price (pricing.input and pricing.output are 0 or absent)Qwen/Qwen2.5-1.5B: no serverless token price (pricing.input and pricing.output are 0 or absent)meta-llama/Meta-Llama-3.1-70B: no serverless token price (pricing.input and pricing.output are 0 or absent)meta-llama/Llama-3.2-1B: no serverless token price (pricing.input and pricing.output are 0 or absent)Qwen/Qwen2.5-7B-Instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)Qwen/Qwen2.5-32B-Instruct: no serverless token price (pricing.input and pricing.output are 0 or absent)meta-llama/Llama-3.1-405B: no serverless token price (pricing.input and pricing.output are 0 or absent)agentica-org/DeepCoder-14B-Preview: no serverless token price (pricing.input and pricing.output are 0 or absent)mistralai/Mistral-7B-v0.1: no serverless token price (pricing.input and pricing.output are 0 or absent)mistralai/Devstral-Small-2505: no serverless token price (pricing.input and pricing.output are 0 or absent)mistralai/Mixtral-8x22B-Instruct-v0.1: no serverless token price (pricing.input and pricing.output are 0 or absent)allenai/Molmo-7B-D-0924: no serverless token price (pricing.input and pricing.output are 0 or absent)Qwen/Qwen3-8B: no serverless token price (pricing.input and pricing.output are 0 or absent)Vertex AI
Updates 33 models (9 new) from the provider's pricing page (page
be62e269e1d8).Changes
deep-research-pro-preview-12-2025: cache_read_input_token_cost $0.2/1M, source https://ai.google.dev/gemini-api/docs/pricing → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricinggemini-2.0-flash: input_cost_per_audio_token $0.7/1M → $1/1M, input_cost_per_character $0.0375/1M chars, input_cost_per_token $0.1/1M → $0.15/1M, input_cost_per_token_batches $0.075/1M, output_cost_per_token $0.4/1M → $0.6/1M, output_cost_per_token_batches $0.3/1M, source https://ai.google.dev/pricing#2_0flash → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricinggemini-2.0-flash-lite: input_cost_per_character $0.01875/1M chars, input_cost_per_token_batches $0.0375/1M, output_cost_per_token_batches $0.15/1M, source https://cloud.google.com/vertex-ai/generative-ai/docs/learn/models#gemini-2.0-flash → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricinggemini-2.5-computer-use-preview-10-2025: source https://cloud.google.com/vertex-ai/generative-ai/docs/computer-use → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricinggemini-2.5-flash: cache_read_input_token_cost_flex $0.03/1M, cache_read_input_token_cost_priority $0.054/1M, input_cost_per_token_batches $0.15/1M, input_cost_per_token_flex $0.15/1M, input_cost_per_token_priority $0.54/1M, output_cost_per_token_batches $1.25/1M, output_cost_per_token_flex $1.25/1M, output_cost_per_token_priority $4.5/1M, source https://ai.google.dev/gemini-api/docs/models#gemini-2.5-flash-preview → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricinggemini-2.5-flash-image: input_cost_per_token_batches $0.15/1M, input_cost_per_token_flex $0.15/1M, input_cost_per_token_priority $0.54/1M, output_cost_per_token_batches $1.25/1M, output_cost_per_token_flex $1.25/1M, source https://ai.google.dev/gemini-api/docs/pricing#gemini-2.5-flash-image → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricinggemini-2.5-flash-lite: cache_read_input_token_cost_flex $0.01/1M, cache_read_input_token_cost_priority $0.018/1M, input_cost_per_token_batches $0.05/1M, input_cost_per_token_flex $0.05/1M, input_cost_per_token_priority $0.18/1M, output_cost_per_token_batches $0.2/1M, output_cost_per_token_flex $0.2/1M, output_cost_per_token_priority $0.72/1M, source https://ai.google.dev/gemini-api/docs/models#gemini-2.5-flash-preview → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricingvertex_ai/gemini-2.5-flash-native-audio(new): input_cost_per_audio_token $3/1M, input_cost_per_token $0.5/1M, output_cost_per_audio_token $12/1M, output_cost_per_token $2/1Mvertex_ai/gemini-2.5-flash-preview-tts(new): input_cost_per_token $0.5/1M, input_cost_per_token_batches $0.25/1M, output_cost_per_audio_token $10/1Mgemini-2.5-pro: cache_read_input_token_cost_above_200k_tokens_priority $0.45/1M, cache_read_input_token_cost_flex $0.125/1M, cache_read_input_token_cost_priority $0.225/1M, input_cost_per_token_above_200k_tokens_priority $4.5/1M, input_cost_per_token_batches $0.625/1M, input_cost_per_token_flex $0.625/1M, input_cost_per_token_priority $2.25/1M, output_cost_per_token_above_200k_tokens_priority $27/1M, output_cost_per_token_batches $5/1M, output_cost_per_token_flex $5/1M, output_cost_per_token_priority $18/1M, source https://cloud.google.com/vertex-ai/generative-ai/pricing → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricinggemini-2.5-pro-preview-tts: input_cost_per_token_batches $0.5/1M, output_cost_per_audio_token $20/1M, source https://ai.google.dev/gemini-api/docs/pricing → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricinggemini-3-flash-preview: cache_read_input_token_cost_flex $0.05/1M, input_cost_per_token_batches $0.25/1M, input_cost_per_token_flex $0.25/1M, output_cost_per_token_batches $1.5/1M, output_cost_per_token_flex $1.5/1M, source https://ai.google.dev/gemini-api/docs/pricing → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricinggemini-3-pro-image: cache_read_input_token_cost $0.2/1M, cache_read_input_token_cost_above_200k_tokens $0.4/1M, cache_read_input_token_cost_above_200k_tokens_priority $0.72/1M, cache_read_input_token_cost_flex $0.1/1M, cache_read_input_token_cost_priority $0.36/1M, input_cost_per_token_above_200k_tokens $4/1M, input_cost_per_token_above_200k_tokens_priority $7.2/1M, input_cost_per_token_flex $1/1M, input_cost_per_token_priority $3.6/1M, output_cost_per_token_above_200k_tokens $18/1M, output_cost_per_token_above_200k_tokens_priority $32.4/1M, output_cost_per_token_flex $6/1M, output_cost_per_token_priority $21.6/1M, source https://ai.google.dev/gemini-api/docs/pricing#gemini-3-pro-image → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricinggemini-3.1-flash-image: cache_read_input_token_cost $0.05/1M, cache_read_input_token_cost_flex $0.025/1M, input_cost_per_token_batches $0.25/1M, input_cost_per_token_flex $0.25/1M, output_cost_per_token_batches $1.5/1M, output_cost_per_token_flex $1.5/1M, source https://cloud.google.com/vertex-ai/generative-ai/pricing#gemini-models → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricinggemini-3.1-flash-lite: source https://ai.google.dev/gemini-api/docs/pricing#gemini-3.1-flash-lite → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricinggemini-3.1-flash-lite-image: cache_read_input_token_cost_flex $0.0125/1M, input_cost_per_token_flex $0.125/1M, output_cost_per_token_flex $0.75/1Mvertex_ai/gemini-3.1-flash-live-preview(new): input_cost_per_audio_token $3/1M, input_cost_per_second $0.0000833333/s, input_cost_per_token $0.75/1M, output_cost_per_audio_token $12/1M, output_cost_per_token $4.5/1Mvertex_ai/gemini-3.1-flash-tts-preview(new): input_cost_per_token $1/1M, input_cost_per_token_batches $0.5/1M, output_cost_per_audio_token $20/1Mgemini-3.1-pro-preview: cache_read_input_token_cost_flex $0.2/1M, input_cost_per_token_flex $1/1M, output_cost_per_token_flex $6/1M, source https://cloud.google.com/vertex-ai/generative-ai/pricing#gemini-models → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricinggemini-3.5-flash: source https://ai.google.dev/gemini-api/docs/pricing → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricinggemini-3.5-flash-lite: cache_read_input_token_cost_priority $0.05/1M → $0.054/1M, source https://ai.google.dev/gemini-api/docs/pricing → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricingvertex_ai/gemini-3.5-live-translate-preview: input_cost_per_second $0.0000883333/svertex_ai/gemini-3.5-transcribe(new): input_cost_per_audio_token $2/1M, input_cost_per_second $0.00005/s, output_cost_per_token $12/1Mvertex_ai/gemini-3.5-transcribe-live(new): input_cost_per_audio_token $3.5/1M, input_cost_per_second $0.0000833333/s, output_cost_per_token $21/1Mgemini-3.6-flash: source https://ai.google.dev/gemini-api/docs/pricing → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricinggemini-3.7-flash: source https://ai.google.dev/gemini-api/docs/pricing → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricinggemini-3.8-flash: source https://ai.google.dev/gemini-api/docs/pricing → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricinggemini-embedding-2: input_cost_per_audio_token $6.5/1M, input_cost_per_image_token $0.45/1M, input_cost_per_token_batches $0.1/1M, source https://cloud.google.com/vertex-ai/generative-ai/pricing → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricinggemini-live-2.5-flash-preview-native-audio-09-2025: input_cost_per_image_token $3/1M, source https://ai.google.dev/gemini-api/docs/pricing → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricingvertex_ai/gemini-omni-1.1-flash(new): input_cost_per_token $1.5/1M, output_cost_per_token $9/1Mgemini-omni-flash-preview: source https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/omni-flash-preview → https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricingvertex_ai/gemini-robotics-er-2(new): cache_read_input_token_cost $0.1/1M, input_cost_per_token $1/1M, input_cost_per_token_batches $0.5/1M, output_cost_per_token $5/1M, output_cost_per_token_batches $2.5/1Mvertex_ai/gemma-4-26b-a4b-it(new): cache_read_input_token_cost $0.015/1M, input_cost_per_token $0.15/1M, output_cost_per_token $0.6/1MSkipped by the parser (3)
gemini-2.5-flash: "Audio Input" above 200K is $0.30, not the base price; litellm has no field for itGemini 2.0 Flash Image Generation: litellm keys only the free experimental model (gemini-2.0-flash-exp-image-generation)Gemini 2.0 Flash Live API: litellm has no vertex_ai/ key for the 2.0 Live APIOpened by the litellm-providers price sync. Every price is read from the provider's own published source and gated before it is applied; the audit trail for each value is in the portal's sync history.