Skip to content

fix(cost-map): add perplexity, openrouter, voyage and nebius models and fix registry metadata - #43907

Merged
kerry-berri merged 5 commits into
mainfrom
litellm_model_registry_audit_1790796018
Oct 1, 2026
Merged

kerry-berri merged 5 commits into
mainfrom
litellm_model_registry_audit_1790796018

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Seven Perplexity Agent API models and two OpenRouter models have no registry entry
  • Three Voyage models listed on the Voyage pricing page have no entry
  • Perplexity openai/gpt-5.6-sol is priced above the published Perplexity rate
  • Several Gemini rows miss documented reasoning, caching or URL context flags
  • Nebius Qwen3.8-27B is missing and three Nebius rows have wrong context limits
  • Retired DeepSeek and Gemini models and some Azure rows carry wrong deprecation_date

How it solves it:

  • Adds 12 entries from official Perplexity, OpenRouter and Voyage sources
  • Corrects six Perplexity GPT-5.6 Sol price fields to the official table
  • Adds Gemini, Nebius capability flags and fixes the Lyria 3.5 input limit
  • Keeps the earlier Nebius limit and deprecation date fixes from this rolling PR

User Flow

Before: a proxy user calling the new Perplexity GPT-6.1 Sol model gets no cost, and GPT-5.6 Sol is overcharged

  1. They call perplexity/openai/gpt-6.1-sol through POST https://litellm-domain/v1/responses
  2. Cost lookup fails with "This model isn't mapped yet", so spend is logged as $0
  3. A perplexity/openai/gpt-5.6-sol call is logged at $5/M input and $30/M output tokens

After: both calls are priced at the rates Perplexity publishes

  1. They call perplexity/openai/gpt-6.1-sol through POST https://litellm-domain/v1/responses
  2. Spend is logged at $2/M input and $10/M output tokens, $4/$15 above 272k
  3. The perplexity/openai/gpt-5.6-sol call is logged at $4/M input and $20/M output tokens

Registry changes by provider

Perplexity

Source for every value: the raw HTML table at https://docs.perplexity.ai/docs/agent-api/models (Input, Output and Cache read columns, read row by row). Cache rows given as a percentage discount were converted from the input price

Added, all litellm_provider perplexity, mode responses:

Key Input /M Output /M Cache read /M Above tier
perplexity/anthropic/claude-fable-5-1 $10 $50 $0.25 none
perplexity/anthropic/claude-opus-5-5 $4 $20 $0.20 none
perplexity/openai/gpt-6.1-sol $2 $10 $0.10 (95% off) >272k: $4 in, $15 out, $0.20 cache
perplexity/openai/gpt-6-sol $2 $10 $0.20 (90% off) >272k: $4 in, $15 out, $0.40 cache
perplexity/openai/gpt-6-luna $0.10 $0.50 $0.01 (90% off) >272k: $0.20 in, $0.75 out, $0.02 cache
perplexity/google/gemini-3.8-flash $0.75 $3.75 $0.075 none
perplexity/xai/grok-4.7 $2 $6 $0.50 >200k: $4 in, $12 out, $1 cache

The table notes that not every third-party model supports every feature, so no supports_function_calling or supports_web_search flags were set on the new rows

Changed: perplexity/openai/gpt-5.6-sol row says $4/$8 input and $20/$30 output at 272k with a 90% cache discount. input_cost_per_token 5e-06 -> 4e-06, output_cost_per_token 3e-05 -> 2e-05, cache_read_input_token_cost 5e-07 -> 4e-07, input_cost_per_token_above_272k_tokens 1e-05 -> 8e-06, output_cost_per_token_above_272k_tokens 4.5e-05 -> 3e-05, cache_read_input_token_cost_above_272k_tokens 1e-06 -> 8e-07. Every other existing Perplexity row was compared against the same table and matches

OpenRouter

Source: the official OpenRouter models API https://openrouter.ai/api/v1/models

openrouter/apodex/apodex-1.1-mini:free added: prompt and completion price 0, context_length 262144 -> max_input_tokens, top_provider.max_completion_tokens 235929 -> max_output_tokens/max_tokens, text only input, tools, tool_choice, structured_outputs and reasoning in supported_parameters

openrouter/unbiased/pareto-26.10-preview added: prompt 0.0000008, completion 0.0000032, input_cache_read 0.00000003, context_length 1048576, max completion 131072, input modalities text and image, tools and tool_choice in supported_parameters

Voyage

Sources: https://docs.voyageai.com/docs/pricing (price rows) and https://docs.voyageai.com/docs/reranker, https://docs.voyageai.com/docs/embeddings (context rows)

voyage/voyage-large-2-instruct added: $0.12/M, 16000 token context, 1024 dimensions. voyage/rerank-1 added: $0.05/M, 8000 token context. voyage/rerank-lite-1 added: $0.02/M, 4000 token context

Gemini API

Source for each flag: the raw HTML of the model page under https://ai.google.dev/gemini-api/docs/models/ (Capabilities table, "Supported" rows)

supports_reasoning true added to gemini/gemini-2.5-flash-native-audio-preview-12-2025, gemini/gemini-3.1-flash-live-preview, gemini/gemini-3.8-live and gemini/gemini-3.8-live-extended-thinking (Thinking: Supported)

supports_prompt_caching true added to gemini/gemini-3.8-flash-tts and gemini/gemini-3.8-flash-lite-tts (Caching: Supported, and both rows already carry a cache read price from https://ai.google.dev/gemini-api/docs/pricing)

supports_url_context true added to gemini/gemini-2.5-pro (URL context: Supported)

gemini/lyria-3.5 max_input_tokens 1048576 -> 131072 (Input token limit 131,072 on https://ai.google.dev/gemini-api/docs/models/lyria-3.5)

Nebius

Sources: https://tokenfactory.nebius.com/model-catalog.md and the public catalog API https://tokenfactory.nebius.com/api/public/models_info (both official Nebius)

nebius/zai-org/GLM-5.3-Flash gains supports_vision true: the catalog lists it as image2text with image in its use cases

nebius/Qwen/Qwen3.8-27B is new (earlier commit): input_cost_per_token 4.5e-07, output_cost_per_token 3e-06, max_input_tokens 262144 (catalog max_model_len), supports_function_calling and supports_reasoning from the catalog use cases, source https://tokenfactory.nebius.com/models/catalog/text2text/Qwen%2FQwen3.8-27B. It is text2text, so no vision flag

nebius/deepseek-ai/DeepSeek-V4-Pro-0813 gains max_input_tokens 979000. nebius/zai-org/GLM-5.3 goes from 1048576 to 1024000 for max_input_tokens and max_tokens. nebius/deepseek-ai/DeepSeek-V4.1-Flash goes from 1048576 to 1048000 for max_input_tokens. All match the catalog max_model_len

Greptile's note that Nebius wildcard discovery misses these rows is real but is a code issue: the discovery list only collects rows whose provider is nebius-chat-models, while all 60 existing Nebius chat rows use nebius. Changing that is out of scope for a data-only PR

Deprecation dates (earlier commit)

DeepSeek, source https://api-docs.deepseek.com/updates: deprecation_date 2026-07-24 added to deepseek-chat, deepseek-reasoner, deepseek/deepseek-chat and deepseek/deepseek-reasoner

Gemini API, source https://ai.google.dev/gemini-api/docs/deprecations: deprecation_date 2026-07-28 added to gemini/gemini-2.5-computer-use-preview-10-2025

Azure OpenAI, source https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-retirement-schedule:

Rows Old New
azure/gpt-4.1-nano, azure/gpt-4.1-nano-2025-04-14, azure/us/gpt-4.1-nano, azure/us/gpt-4.1-nano-2025-04-14, azure/eu/gpt-4.1-nano 2027-04-14 2026-10-14
azure/gpt-4o-transcribe 2026-12-31 2026-10-15
azure/gpt-realtime-2.1, azure/gpt-realtime-2.1-mini 2027-07-31 2027-06-25

Other open registry PRs

This is the rolling registry PR. Open PRs touching the registry were diffed field by field against their own merge base and re-checked against current main and official sources. None had a change that could be verified and was not already on main, so nothing was absorbed and no PR was closed

Dropped as unverifiable or already covered (data-only PRs):

Left alone because they change code or CI behavior and only touch the JSON as part of that: #35617, #43581, #43390, #43363 and the remaining non data-only PRs

ci_cd/cost_map_pins.json is not on main (it is added by #43390). This PR does not touch the Fireworks MiniMax M3 rows that file pins

Checked and not changed

Groq allam-2-7b: https://console.groq.com/docs/model/allam-2-7b lists pricing as pending, so it was not added. Together Kimi K3 and media models, Cohere new models (Embed v5, Command A variants, North): no per-token price on https://www.together.ai/pricing or https://cohere.com/pricing for the missing rows. Bedrock luma.ray-v2:0 and new regional inference profiles: no per-unit row matched with confidence on https://aws.amazon.com/bedrock/pricing/. Google Antigravity previews and aqa: no per-token price on https://ai.google.dev/gemini-api/docs/pricing. OpenRouter openrouter/auto-beta, openrouter/fusion, openrouter/pareto-code are routers with variable pricing, and the Lyria previews are music models, so they were not added

Anthropic models and prices (https://api.anthropic.com/v1/models, https://platform.claude.com/docs/en/about-claude/pricing), OpenAI (https://developers.openai.com/api/docs/pricing), DeepSeek (https://api-docs.deepseek.com/quick_start/pricing), Mistral (https://api.mistral.ai/v1/models) and xAI model lists all matched the registry on the rows checked. Mistral codestral-2508 keeps 128k since the model card wins over the API's 256k. Google image model limits where the API and model page disagree were left alone

Fireworks qwen3p8 rows: pricing was not published in https://api.fireworks.ai/inference/v1/models, the raw or r.jina.ai copy of https://fireworks.ai/models, or the rendered page, so nothing was added

Pre-Submission checklist

  • The handful of test files covering my change pass locally
  • My PR's scope is as isolated as possible; it only solves 1 specific problem

Screenshots / Proof of Fix

Probe: LITELLM_LOCAL_MODEL_COST_MAP=True python -c "litellm.get_model_info(<model>)"

Before (fcf8797)

  1. perplexity/openai/gpt-6.1-sol -> no registry entry, so lookup fails
  2. openrouter/unbiased/pareto-26.10-preview -> no registry entry, so lookup fails
  3. voyage/rerank-1 -> no registry entry, so lookup fails
  4. gemini/lyria-3.5 -> max_input_tokens 1048576
  5. perplexity/openai/gpt-5.6-sol -> input 5e-06, output 3e-05

After (8ef6abd)

  1. perplexity/openai/gpt-6.1-sol -> input 2e-06, output 1e-05, above 272k 4e-06 and 1.5e-05
  2. openrouter/unbiased/pareto-26.10-preview -> max_input_tokens 1048576, input 8e-07, output 3.2e-06
  3. voyage/rerank-1 -> max_input_tokens 8000, input 5e-08
  4. gemini/lyria-3.5 -> max_input_tokens 131072
  5. perplexity/openai/gpt-5.6-sol -> input 4e-06, output 2e-05

python3 ci_cd/check_files_match.py passes, the schema generator leaves the schema unchanged, tests/unit/test_model_prices_schema.py, tests/unit/test_model_cost_aliases.py and tests/unit/test_cost_map_guard.py pass (98 passed), and the cost map guard passes against main

Type

🐛 Bug Fix

Link to Devin session: https://app.devin.ai/sessions/d50c257f78224426bbe6e9a4db7c8264
Open in Devin Desktop: https://app.devin.ai/desktop/session/d50c257f78224426bbe6e9a4db7c8264?variant=devin

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration
devin-ai-integration Bot requested a review from a team September 30, 2026 19:24
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@greptile-apps

greptile-apps Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

[Medium risk] Updates model pricing and context window metadata.

The PR should not merge until the new Qwen model is available through Nebius wildcard discovery and provider-based listings.

Findings

  1. P1 Nebius discovery misses Qwen ▶

Summary

The PR adds pricing and context metadata for a Nebius Qwen model and updates three Nebius context limits in both registry copies.

  • The new model is not included in Nebius provider-based discovery or wildcard expansion.

Reviews (1) · Last reviewed commit: "fix(cost-map): add nebius qwen3.8-27b an..."

},
"nebius/Qwen/Qwen3.8-27B": {
"input_cost_per_token": 4.5e-07,
"litellm_provider": "nebius",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Nebius discovery misses Qwen When a proxy uses Nebius wildcard model discovery, this new model is omitted. The registry adds models to the Nebius discovery list only when their provider is nebius-chat-models, but this entry uses nebius. Direct lookup can price the model, yet nebius/* expansion and provider-based model listings cannot expose it.

@codecov

codecov Bot commented Sep 30, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@codspeed

codspeed Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_model_registry_audit_1790796018 (8ef6abd) with main (fcf8797)

Open in CodSpeed

…k, gemini and azure models

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration devin-ai-integration Bot changed the title fix(cost-map): add nebius qwen3.8-27b and correct nebius context limits fix(cost-map): add nebius qwen3.8-27b, fix nebius limits and provider deprecation dates Oct 1, 2026
…ct gemini, nebius and perplexity metadata

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration devin-ai-integration Bot changed the title fix(cost-map): add nebius qwen3.8-27b, fix nebius limits and provider deprecation dates fix(cost-map): add perplexity, openrouter, voyage and nebius models and fix registry metadata Oct 1, 2026
@kerry-berri
kerry-berri merged commit ac8c5aa into main Oct 1, 2026
103 of 105 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant