Repository navigation
fix(cost-map): add perplexity, openrouter, voyage and nebius models and fix registry metadata - #43907
Conversation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
|
|
|
| }, | ||
| "nebius/Qwen/Qwen3.8-27B": { | ||
| "input_cost_per_token": 4.5e-07, | ||
| "litellm_provider": "nebius", |
There was a problem hiding this comment.
Nebius discovery misses Qwen When a proxy uses Nebius wildcard model discovery, this new model is omitted. The registry adds models to the Nebius discovery list only when their provider is
nebius-chat-models, but this entry uses nebius. Direct lookup can price the model, yet nebius/* expansion and provider-based model listings cannot expose it.
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
…y_audit_1790796018
…k, gemini and azure models Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…y_audit_1790796018
…ct gemini, nebius and perplexity metadata Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
TLDR
Problem this solves:
openai/gpt-5.6-solis priced above the published Perplexity ratedeprecation_dateHow it solves it:
User Flow
Before: a proxy user calling the new Perplexity GPT-6.1 Sol model gets no cost, and GPT-5.6 Sol is overcharged
perplexity/openai/gpt-6.1-solthrough POST https://litellm-domain/v1/responsesperplexity/openai/gpt-5.6-solcall is logged at $5/M input and $30/M output tokensAfter: both calls are priced at the rates Perplexity publishes
perplexity/openai/gpt-6.1-solthrough POST https://litellm-domain/v1/responsesperplexity/openai/gpt-5.6-solcall is logged at $4/M input and $20/M output tokensRegistry changes by provider
Perplexity
Source for every value: the raw HTML table at https://docs.perplexity.ai/docs/agent-api/models (Input, Output and Cache read columns, read row by row). Cache rows given as a percentage discount were converted from the input price
Added, all
litellm_providerperplexity,moderesponses:perplexity/anthropic/claude-fable-5-1perplexity/anthropic/claude-opus-5-5perplexity/openai/gpt-6.1-solperplexity/openai/gpt-6-solperplexity/openai/gpt-6-lunaperplexity/google/gemini-3.8-flashperplexity/xai/grok-4.7The table notes that not every third-party model supports every feature, so no
supports_function_callingorsupports_web_searchflags were set on the new rowsChanged:
perplexity/openai/gpt-5.6-solrow says $4/$8 input and $20/$30 output at 272k with a 90% cache discount.input_cost_per_token5e-06 -> 4e-06,output_cost_per_token3e-05 -> 2e-05,cache_read_input_token_cost5e-07 -> 4e-07,input_cost_per_token_above_272k_tokens1e-05 -> 8e-06,output_cost_per_token_above_272k_tokens4.5e-05 -> 3e-05,cache_read_input_token_cost_above_272k_tokens1e-06 -> 8e-07. Every other existing Perplexity row was compared against the same table and matchesOpenRouter
Source: the official OpenRouter models API https://openrouter.ai/api/v1/models
openrouter/apodex/apodex-1.1-mini:freeadded: prompt and completion price 0,context_length262144 ->max_input_tokens,top_provider.max_completion_tokens235929 ->max_output_tokens/max_tokens, text only input,tools,tool_choice,structured_outputsandreasoninginsupported_parametersopenrouter/unbiased/pareto-26.10-previewadded: prompt 0.0000008, completion 0.0000032,input_cache_read0.00000003,context_length1048576, max completion 131072, input modalities text and image,toolsandtool_choiceinsupported_parametersVoyage
Sources: https://docs.voyageai.com/docs/pricing (price rows) and https://docs.voyageai.com/docs/reranker, https://docs.voyageai.com/docs/embeddings (context rows)
voyage/voyage-large-2-instructadded: $0.12/M, 16000 token context, 1024 dimensions.voyage/rerank-1added: $0.05/M, 8000 token context.voyage/rerank-lite-1added: $0.02/M, 4000 token contextGemini API
Source for each flag: the raw HTML of the model page under https://ai.google.dev/gemini-api/docs/models/ (Capabilities table, "Supported" rows)
supports_reasoningtrue added togemini/gemini-2.5-flash-native-audio-preview-12-2025,gemini/gemini-3.1-flash-live-preview,gemini/gemini-3.8-liveandgemini/gemini-3.8-live-extended-thinking(Thinking: Supported)supports_prompt_cachingtrue added togemini/gemini-3.8-flash-ttsandgemini/gemini-3.8-flash-lite-tts(Caching: Supported, and both rows already carry a cache read price from https://ai.google.dev/gemini-api/docs/pricing)supports_url_contexttrue added togemini/gemini-2.5-pro(URL context: Supported)gemini/lyria-3.5max_input_tokens1048576 -> 131072 (Input token limit 131,072 on https://ai.google.dev/gemini-api/docs/models/lyria-3.5)Nebius
Sources: https://tokenfactory.nebius.com/model-catalog.md and the public catalog API https://tokenfactory.nebius.com/api/public/models_info (both official Nebius)
nebius/zai-org/GLM-5.3-Flashgainssupports_visiontrue: the catalog lists it asimage2textwithimagein its use casesnebius/Qwen/Qwen3.8-27Bis new (earlier commit):input_cost_per_token4.5e-07,output_cost_per_token3e-06,max_input_tokens262144 (catalogmax_model_len),supports_function_callingandsupports_reasoningfrom the catalog use cases, source https://tokenfactory.nebius.com/models/catalog/text2text/Qwen%2FQwen3.8-27B. It istext2text, so no vision flagnebius/deepseek-ai/DeepSeek-V4-Pro-0813gainsmax_input_tokens979000.nebius/zai-org/GLM-5.3goes from 1048576 to 1024000 formax_input_tokensandmax_tokens.nebius/deepseek-ai/DeepSeek-V4.1-Flashgoes from 1048576 to 1048000 formax_input_tokens. All match the catalogmax_model_lenGreptile's note that Nebius wildcard discovery misses these rows is real but is a code issue: the discovery list only collects rows whose provider is
nebius-chat-models, while all 60 existing Nebius chat rows usenebius. Changing that is out of scope for a data-only PRDeprecation dates (earlier commit)
DeepSeek, source https://api-docs.deepseek.com/updates:
deprecation_date2026-07-24 added todeepseek-chat,deepseek-reasoner,deepseek/deepseek-chatanddeepseek/deepseek-reasonerGemini API, source https://ai.google.dev/gemini-api/docs/deprecations:
deprecation_date2026-07-28 added togemini/gemini-2.5-computer-use-preview-10-2025Azure OpenAI, source https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-retirement-schedule:
azure/gpt-4.1-nano,azure/gpt-4.1-nano-2025-04-14,azure/us/gpt-4.1-nano,azure/us/gpt-4.1-nano-2025-04-14,azure/eu/gpt-4.1-nanoazure/gpt-4o-transcribeazure/gpt-realtime-2.1,azure/gpt-realtime-2.1-miniOther open registry PRs
This is the rolling registry PR. Open PRs touching the registry were diffed field by field against their own merge base and re-checked against current
mainand official sources. None had a change that could be verified and was not already onmain, so nothing was absorbed and no PR was closedDropped as unverifiable or already covered (data-only PRs):
moonshot/kimi-k2.5already onmain, 8192 output cap not in Moonshot docsdeepseek-v3rowtinkeris not a LiteLLM providermainmain, cache and context edits conflict with https://api.fireworks.ai/v1/serverless/modelslist_foundation_modelsLeft alone because they change code or CI behavior and only touch the JSON as part of that: #35617, #43581, #43390, #43363 and the remaining non data-only PRs
ci_cd/cost_map_pins.jsonis not onmain(it is added by #43390). This PR does not touch the Fireworks MiniMax M3 rows that file pinsChecked and not changed
Groq
allam-2-7b: https://console.groq.com/docs/model/allam-2-7b lists pricing as pending, so it was not added. Together Kimi K3 and media models, Cohere new models (Embed v5, Command A variants, North): no per-token price on https://www.together.ai/pricing or https://cohere.com/pricing for the missing rows. Bedrockluma.ray-v2:0and new regional inference profiles: no per-unit row matched with confidence on https://aws.amazon.com/bedrock/pricing/. Google Antigravity previews andaqa: no per-token price on https://ai.google.dev/gemini-api/docs/pricing. OpenRouteropenrouter/auto-beta,openrouter/fusion,openrouter/pareto-codeare routers with variable pricing, and the Lyria previews are music models, so they were not addedAnthropic models and prices (https://api.anthropic.com/v1/models, https://platform.claude.com/docs/en/about-claude/pricing), OpenAI (https://developers.openai.com/api/docs/pricing), DeepSeek (https://api-docs.deepseek.com/quick_start/pricing), Mistral (https://api.mistral.ai/v1/models) and xAI model lists all matched the registry on the rows checked. Mistral
codestral-2508keeps 128k since the model card wins over the API's 256k. Google image model limits where the API and model page disagree were left aloneFireworks
qwen3p8rows: pricing was not published in https://api.fireworks.ai/inference/v1/models, the raw or r.jina.ai copy of https://fireworks.ai/models, or the rendered page, so nothing was addedPre-Submission checklist
Screenshots / Proof of Fix
Probe:
LITELLM_LOCAL_MODEL_COST_MAP=True python -c "litellm.get_model_info(<model>)"Before (fcf8797)
perplexity/openai/gpt-6.1-sol-> no registry entry, so lookup failsopenrouter/unbiased/pareto-26.10-preview-> no registry entry, so lookup failsvoyage/rerank-1-> no registry entry, so lookup failsgemini/lyria-3.5-> max_input_tokens 1048576perplexity/openai/gpt-5.6-sol-> input 5e-06, output 3e-05After (8ef6abd)
perplexity/openai/gpt-6.1-sol-> input 2e-06, output 1e-05, above 272k 4e-06 and 1.5e-05openrouter/unbiased/pareto-26.10-preview-> max_input_tokens 1048576, input 8e-07, output 3.2e-06voyage/rerank-1-> max_input_tokens 8000, input 5e-08gemini/lyria-3.5-> max_input_tokens 131072perplexity/openai/gpt-5.6-sol-> input 4e-06, output 2e-05python3 ci_cd/check_files_match.pypasses, the schema generator leaves the schema unchanged,tests/unit/test_model_prices_schema.py,tests/unit/test_model_cost_aliases.pyandtests/unit/test_cost_map_guard.pypass (98 passed), and the cost map guard passes againstmainType
🐛 Bug Fix
Link to Devin session: https://app.devin.ai/sessions/d50c257f78224426bbe6e9a4db7c8264
Open in Devin Desktop: https://app.devin.ai/desktop/session/d50c257f78224426bbe6e9a4db7c8264?variant=devin