Skip to content

fix(cost-map): retirement dates, chatgpt reasoning flags, bing pricing, bedrock mantle and mythos, azure gpt-5.6 alias, anthropic batch rates, new nebius, openrouter and xai rows - #42951

Merged
kerry-berri merged 23 commits into
mainfrom
litellm_registry_audit_20260924
Sep 26, 2026

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • OpenAI announced shutdown dates for 8 gpt-5.x codex and chat-latest models, none in the cost map
  • Google announced shutdown dates for 3 Gemini 3 preview models, none in the cost map
  • 12 chatgpt/* rows had no reasoning flags while their openai twins did
  • bing_grounding/search charged $35 per 1k searches, Microsoft lists $14
  • 8 Bedrock Mantle open-weight models (DeepSeek V3.1, Kimi K2 Thinking, Qwen3) had no rows
  • 7 Azure AI DeepSeek and Grok deployments and Cohere Aya Expanse 32B had no rows
  • anthropic.claude-mythos-preview base row billed $0 while its regional rows bill $27.50 / $137.50
  • azure/gpt-5.6 alias rows still billed $5 / $30 although gpt-5.6 routes to GPT-5.6 Sol at $4 / $20
  • 19 priced anthropic rows had no Batch API prices
  • Nebius DeepSeek V4.1 Flash, OpenRouter Perceptron Mk1.5 and xAI grok-imagine-image-pro had no rows
  • gemini-2.5-flash-native-audio-latest claimed a 1M input window, the API says 131072

How it solves it:

User Flow

Before: an operator reading /model/info cannot see that these models have a published retirement date, and a chatgpt reasoning model looks like a non-reasoning one

  1. They send GET https://litellm-domain/v1/model/info and look at gpt-5.1-codex and gemini/gemini-3.1-flash-lite-preview
  2. Neither has a deprecation_date, although both providers have published a shutdown date
  3. chatgpt/gpt-5.6-sol comes back with supports_reasoning unset, so their dashboard offers no reasoning effort for it

After: every row carries the provider's published date and the chatgpt twin matches its openai row

  1. They send the same GET https://litellm-domain/v1/model/info
  2. gpt-5.1-codex shows "deprecation_date": "2026-07-23" and gemini/gemini-3.1-flash-lite-preview shows "2026-05-25"
  3. chatgpt/gpt-5.6-sol shows supports_reasoning: true with the same effort flags as gpt-5.6-sol

Relevant issues

Fixes #41394 (absorbed from #42923)
Supersedes #43016, #42727, #42594, #42728, #42704, #43048, #43176 and #43171
Supports #26900 (deprecation metadata in the registry)

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests (test_chatgpt_rows_carry_their_openai_twin_reasoning_annotations fails on main's data with all 38 mismatches and passes here; the date edits are covered by tests/test_litellm/test_model_prices_schema.py and test_cost_map_guard.py)
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.). The two red jobs, misc (tests/unit/test_unit_shard_missing_paths.py, missing .circleci/scripts/unit_selection.sh) and mcp-integration (test_sse_mcp_handler_mock, test_call_tool), fail identically on main's latest runs
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 (5/5 at aa88064, Bugbot found no issues) before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

This is the rolling registry PR for the 2026-09-24 audit run. The previous rolling PR #42947 merged earlier today, so this one is new

Sources for every changed value

Every row below was read from the raw HTML table (curl with a cache buster, <tr>/<td> parsed) on 2026-09-24 and cross-checked with r.jina.ai, which agreed on every row

OpenAI, https://platform.openai.com/docs/deprecations

Registry key deprecation_date Raw row (Shutdown date, Model, Recommended replacement)
gpt-5-codex 2026-07-23 July 23, 2026 | gpt-5-codex | gpt-5.6-sol
gpt-5.1-chat-latest 2026-07-23 July 23, 2026 | gpt-5.1-chat-latest | gpt-5.6-sol
gpt-5.1-codex 2026-07-23 July 23, 2026 | gpt-5.1-codex | gpt-5.6-sol
gpt-5.1-codex-max 2026-07-23 July 23, 2026 | gpt-5.1-codex-max | gpt-5.6-sol
gpt-5.1-codex-mini 2026-07-23 July 23, 2026 | gpt-5.1-codex-mini | gpt-5.6-terra
gpt-5.2-codex 2026-07-23 July 23, 2026 | gpt-5.2-codex | gpt-5.6-sol
gpt-5.2-chat-latest 2026-08-10 Aug 10, 2026 | gpt-5.2-chat-latest | gpt-5.6-sol
gpt-5.3-chat-latest 2026-08-10 Aug 10, 2026 | gpt-5.3-chat-latest | gpt-5.6-sol

Only the bare openai keys change. The azure/* twins keep their dates from the Azure schedule, and the chatgpt/, github_copilot/ and openrouter/ twins are other providers' routes with no published date

Gemini API, https://ai.google.dev/gemini-api/docs/deprecations

Registry key deprecation_date Raw row (Model, Release date, Shutdown date, Recommended replacement)
gemini/gemini-3-pro-image-preview 2026-06-25 gemini-3-pro-image-preview | November 20, 2025 | June 25, 2026 | gemini-3-pro-image
gemini/gemini-3.1-flash-image-preview 2026-06-25 gemini-3.1-flash-image-preview | February 26, 2026 | June 25, 2026 | gemini-3.1-flash-image
gemini/gemini-3.1-flash-lite-preview 2026-05-25 gemini-3.1-flash-lite-preview | March 3, 2026 | May 25, 2026 | gemini-3.1-flash-lite

The vertex_ai/ and bare gemini-3* rows are Vertex routes; the Vertex deprecations page does not list these preview models, so they are left alone

chatgpt twins, absorbed from #42923

Not a vendor value. ChatGPTConfig and ChatGPTResponsesAPIConfig subclass the openai configs, so a chatgpt/<model> row behaves as the bare <model> openai row. The 38 fields are copied verbatim from the bare twin on this branch (supports_reasoning, supports_minimal_reasoning_effort, supports_none_reasoning_effort, supports_xhigh_reasoning_effort, default_reasoning_effort) onto chatgpt/gpt-5.1-codex-max, gpt-5.1-codex-mini, gpt-5.2, gpt-5.2-codex, gpt-5.3-chat-latest, gpt-5.3-codex, gpt-5.4, gpt-5.4-pro, gpt-5.5, gpt-5.6-luna, gpt-5.6-sol, gpt-5.6-terra. chatgpt/gpt-5.3-codex-spark and chatgpt/gpt-5.3-instant have no bare twin and are untouched. The new test fails on main's data (38 mismatches) and passes here. Commit carries a Co-authored-by trailer for the original author

Bing Grounding, https://www.microsoft.com/en-us/bing/apis/grounding-pricing

Raw page and r.jina.ai both read Grounding with Bing Search ... $14 per 1,000 transactions, so bing_grounding/search.input_cost_per_query goes from 0.035 to 0.014 and the metadata.notes text follows. tests/search_tests/test_bing_grounding_search.py asserted the old 0.035 and now asserts 0.014

Bedrock Mantle, https://aws.amazon.com/bedrock/pricing/ and the AWS model cards

Prices are the standard on-demand US East rows, read from the raw pricing page, r.jina.ai and the AWS public price list JSON (per 1K tokens), which all agree. Context window, max output, reasoning, image input and client-side tool calling come from each model card's raw HTML (checkmark icons parsed per cell). Every card marks the Responses API as not supported on the bedrock-mantle endpoint, so supported_endpoints is /v1/chat/completions only

Registry key in / out per 1M context / max out flags model card
bedrock_mantle/deepseek.v3.1 $0.58 / $1.68 128K / 8K tools https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-deepseek-deepseek-v3-1.html
bedrock_mantle/moonshotai.kimi-k2-thinking $0.60 / $2.50 256K / 16K tools, reasoning https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-moonshot-ai-kimi-k2-thinking.html
bedrock_mantle/qwen.qwen3-235b-a22b-2507 $0.22 / $0.88 256K / 8K tools, reasoning https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-qwen-qwen3-235b-a22b-2507.html
bedrock_mantle/qwen.qwen3-32b $0.15 / $0.60 32K / 8K tools, reasoning https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-qwen-qwen3-32b.html
bedrock_mantle/qwen.qwen3-coder-30b-a3b-instruct $0.15 / $0.60 256K / 16K tools https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-qwen-qwen3-coder-30b-a3b-instruct.html
bedrock_mantle/qwen.qwen3-coder-480b-a35b-instruct $0.45 / $1.80 128K / 16K tools https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-qwen-qwen3-coder-480b-a35b-instruct.html
bedrock_mantle/qwen.qwen3-next-80b-a3b-instruct $0.14 / $1.20 256K / 8K tools, reasoning https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-qwen-qwen3-next-80b-a3b.html
bedrock_mantle/qwen.qwen3-vl-235b-a22b-instruct $0.53 / $2.66 256K / 8K tools, vision https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-qwen-qwen3-vl-235b-a22b.html

The prices match #42727. Its token limits did not match the model cards (for example 32B at 131072 context and 81920 or 131072 max output on several rows) and its supports_reasoning on DeepSeek V3.1 and the coder models has no card support, so those fields were taken from the cards instead. Its test file pinned those vendor values as literals and was not carried over

Azure AI, https://prices.azure.com/api/retail/prices?$filter=serviceName%20eq%20'Foundry%20Models'%20and%20armRegionName%20eq%20'eastus'%20and%20priceType%20eq%20'Consumption'

Global (glbl) meters read from the Retail Prices API on 2026-09-24, per 1K tokens, in / out: R1 0.00135 / 0.0054, V3-0324 0.00114 / 0.00456, V3.1 0.00123 / 0.00494, Grok-3 0.003 / 0.015, Grok-3 Mini 0.00025 / 0.00127, Grok4 Fast 0.0002 / 0.0005. New rows azure_ai/deepseek-r1, azure_ai/deepseek-v3-0324, azure_ai/deepseek-v3.1, azure_ai/grok-3, azure_ai/grok-3-mini, azure_ai/grok-4-fast-reasoning, azure_ai/grok-4-fast-non-reasoning carry only price, mode, provider and source. #42594 also proposed context limits, capability flags and retirement dates for them that the API does not publish, so those were not carried, and its azure_ai/MAI-Image-2e row has no meter in the API and was dropped. The rest of #42594's 24 keys already exist on main with matching prices

Cohere, https://docs.cohere.com/docs/models and https://cohere.com/pricing

c4ai-aya-expanse-32b: models table row Text | 128k | 4k | Chat, pricing page Aya Expanse 8B/32B: $0.50 / $1.50 per 1M. c4ai-aya-vision-32b, cohere-transcribe-03-2026 and the embed-*-v3.0-image ids are on the models list but have no per-token price on the pricing page, so no rows

2026-09-25 run

Bedrock, https://pricing.us-east-1.amazonaws.com/offers/v1.0/aws/AmazonBedrockFoundationModels/current/index.json (absorbed from #43048)

Product rows with servicename "Claude Mythos Preview (Amazon Bedrock Edition)" in us-east-1, per 1M tokens: Million Input Tokens Standard 27.50, Million Response Tokens Standard 137.50, Million Cache Read Input Tokens Standard 2.75, Million Cache Write Input Tokens Standard 34.375, Million 1 hour Cache Write Input Tokens Standard 55.00. anthropic.claude-mythos-preview goes from 0 / 0 to those five values with supports_prompt_caching true, matching the existing us. / apac. / au. rows. The PR's test asserts that any priced cross-region bedrock profile has a priced base row, an invariant, so it is carried. Commit carries a Co-authored-by trailer

Azure, https://prices.azure.com/api/retail/prices?$filter=serviceName%20eq%20'Foundry%20Models'%20and%20armRegionName%20eq%20'eastus'%20and%20priceType%20eq%20'Consumption' (absorbed from #43176)

Azure publishes no bare "5.6" meter, only 5.6 sol, 5.6 terra and 5.6 luna; the OpenAI model page https://platform.openai.com/docs/models/gpt-5.6 says "The gpt-5.6 alias routes requests to GPT-5.6 Sol". The Retail Prices API reads 5.6 sol ShortCo Inp Std Gl 4.0, Opt Std Gl 20.0, Cd Inp Std Gl 0.4, Cd Wr Std Gl 5.0, LongCo Inp Std Gl 8.0, LongCo Opt Std Gl 30.0, priority Inp PP Gl 8.0 / Opt PP Gl 40.0, and the DZ meters at 1.1x, which are exactly the azure/gpt-5.6-sol rows already on main. The 16 cost fields on azure/gpt-5.6, azure/us/gpt-5.6 and azure/eu/gpt-5.6 now equal their -sol sibling field for field, and the PR's test pins that invariant. Commit carries a Co-authored-by trailer

Anthropic Batch API, https://platform.claude.com/docs/en/about-claude/pricing (absorbed from #43171)

Raw Batch table rows read today: Claude Fable 5.1 $5 / $25, Claude Opus 5.5 $2 / $10, Claude Sonnet 5 $1 / $5, Claude Haiku 4.5 $0.50 / $2.50, Claude Mythos 5.1 $5 / $25, Claude Fable 5 $5 / $25, Claude Mythos 5 $5 / $25, Claude Opus 5 / 4.8 / 4.7 / 4.6 / 4.5 $2.50 / $12.50, Claude Sonnet 4.6 / 4.5 $1.50 / $7.50. The page also states "These multipliers stack with other pricing modifiers, including the Batch API discount" for cache writes and reads and "Prompt caching and batch processing discounts apply at standard rates across the full context window", so each of the 19 rows gets *_batches at exactly half of every cache and above-200k field it already carries (84 fields, checked programmatically). claude-mythos-preview has no Batch row and is untouched. The PR's source edits on 5 rows were not carried. The schema was regenerated for the four new *_above_200k_tokens_batches keys and the inline schema in tests/unit/test_utils.py, the ModelInfo types in litellm/types/utils.py and the strict Rust catalog struct in litellm-rust/crates/model-catalog/src/model_info.rs list them too. The Rust workflow is path filtered and did not run on the registry-only commits, so cargo test -p litellm-model-catalog --all-features -- --include-ignored was run locally: it also failed on fields already on main (output_cost_per_image_0.5K/1K/2K/4K, computer_use_*, file_search_*, vector_store_cost_per_gb_per_day, fallback_generalizations.rules), so those are declared in the struct in the same commit and the suite is now 26 passed. The same four fields are also copied onto ModelInfo in _get_model_info_helper (the existing 272k batch copy did not cover them, caught by Bugbot) and ui/litellm-dashboard/src/types/schema.d.ts is regenerated for them. _DEPLOYMENT_PRICING_KEYS in litellm_logging.py also lists the four keys so a deployment that overrides only a 200k batch rate is billed at that rate (caught by Greptile). After merging main, which landed three of the four 200k batch fields on its own, this PR's Python wiring is reduced to the one main lacks, cache_creation_input_token_cost_above_200k_tokens_batches

Nebius, https://tokenfactory.nebius.com/endpoints?modals=endpoint-details&model-id=deepseek-ai/DeepSeek-V4.1-Flash

Endpoint page (rendered through r.jina.ai, the catalog is a JS shell): $0.30 / 1M In $1.20 / 1M Out, Modality Vision, Context 1,048K, Tool calling N/A, Reasoning N/A. New row nebius/deepseek-ai/DeepSeek-V4.1-Flash with 3e-7 / 1.2e-6, 1048576 limits and supports_vision, no tool or reasoning flag

OpenRouter, https://openrouter.ai/api/v1/models

perceptron/perceptron-mk1.5: pricing.prompt 0.00000015, pricing.completion 0.0000015, context_length 36864, top_provider.max_completion_tokens 8192, input_modalities text, image, video, audio, supported_parameters include tools, tool_choice, structured_outputs, reasoning. New row openrouter/perceptron/perceptron-mk1.5 with those values and flags

xAI, https://api.x.ai/v1/models

grok-imagine-image-quality lists "aliases": ["grok-imagine-image-quality-20260403", "grok-imagine-image-quality-latest", "grok-imagine-image-pro"] with image_price 500000000 ($0.05). New row xai/grok-imagine-image-pro copies the xai/grok-imagine-image-quality row (deprecation fields excluded, another automation owns them)

Gemini API, https://generativelanguage.googleapis.com/v1beta/models

models/gemini-2.5-flash-native-audio-latest returns inputTokenLimit 131072, outputTokenLimit 8192, the same as both dated previews. gemini-2.5-flash-native-audio-latest and gemini/gemini-2.5-flash-native-audio-latest had max_input_tokens 1048576 and now read 131072

Checked and intentionally not changed

  • Azure, https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-retirement-schedule: the 8 dates from fix(cost-map): update azure gpt-4.1-nano and gpt-4o-2024-05-13 retirement dates #42947 (gpt-4.1-nano family 2026-10-14, gpt-4o-2024-05-13 family 2026-12-09) match the raw rows re-read today. azure/gpt-realtime-mini-2025-10-06 is left alone because the page lists that version twice with conflicting dates (2027-04-06 and 2026-09-21)
  • Anthropic: 14 models listed as Active | N/A | Not sooner than <date> on https://docs.anthropic.com/en/docs/about-claude/model-deprecations. Those are earliest-possible floors, not announced dates, so no deprecation_date is written. claude-mythos-preview keeps 2026-06-09 (added in chore(models): add deprecation_date to claude-mythos-preview from the Anthropic deprecations page #42845); the page not listing it is not evidence the date changed
  • OpenAI gpt-4o-audio-preview-*, gpt-4o-mini-audio-preview-2024-12-17, gpt-4o-mini-realtime-preview-2024-12-17 keep 2027-01-20: the page dates the undated -preview aliases, not these snapshots
  • Bedrock Mantle and Azure GPT-6 Sol/Luna rows from fix(prices): add estimated Bedrock prices for GPT-6 Sol and GPT-6 Luna #42728 and feat(azure): price GPT-6 Sol and GPT-6 Luna on the azure route #42704 are already on main with the same prices. fix(prices): add estimated Bedrock prices for GPT-6 Sol and GPT-6 Luna #42728 additionally set supports_web_search and search_context_cost_per_query on the mantle rows and feat(azure): price GPT-6 Sol and GPT-6 Luna on the azure route #42704 flipped supports_max_reasoning_effort and added supports_prompt_cache_breakpoint, none of which the AWS model cards or Azure docs state, so not carried
  • Voyage legacy voyage-01, voyage-02, rerank-1 are priced on https://docs.voyageai.com/docs/pricing but have no documented context length; voyage-4-nano has no price row. OpenAI dated snapshots (o3-deep-research-2025-06-26, gpt-4o-search-preview-2025-03-11 and 7 more), Gemini antigravity-preview-* and aqa, Fireworks qwen3p8-2p4t-a95b, OpenRouter Lyria 3 and the Together model list have no official per-token price row, so no rows
  • 2026-09-25 breadth pass (provider /models APIs plus raw pricing pages): OpenAI 163 ids, Anthropic 12, Gemini 61, Bedrock 120, xAI 54, DeepSeek 2, Mistral 53, OpenRouter 458, Fireworks 27 and the Groq, Nebius, Voyage, Perplexity and Cohere public tables. Price spot-checks on the newest rows (gpt-5.6 family, gpt-5.3-codex, Claude Fable 5.1 / Opus 5.5 / Opus 5 / Sonnet 5 / Haiku 4.5, Gemini 3.x flash and pro, Bedrock Claude, DeepSeek flash / v4-pro, grok-4.7, Groq gpt-oss and whisper, Voyage 4 family, Nebius catalog, Perplexity sonar) all match
  • Groq llama-3.1-8b-instant, llama-3.3-70b-versatile, minimaxai/minimax-m2.7 are listed on https://console.groq.com/docs/models as Contact Sales, no per-token price, so no rows
  • Cohere command-a-reasoning-08-2025, command-a-vision-07-2025, command-a-translate-08-2025, command-r7b-arabic-02-2025, c4ai-aya-vision-32b, cohere-transcribe-03-2026 and the embed-*-v3.0-image ids are in /v1/models but https://cohere.com/pricing publishes no per-token price for them, so no rows
  • Bedrock amazon.nova-reel-v1:0 / v1:1 are billed per second of video on https://aws.amazon.com/bedrock/pricing/ and the row did not extract from the raw page or the price list in this run, so no rows yet
  • OpenRouter google/lyria-3-*-preview (per-song price in the description, 0 in the API), openrouter/auto-beta, fusion, pareto-code (dynamic -1 pricing) are left alone
  • computer-use-preview and us./global.openai.gpt-5.x rows have mode: chat with only /v1/responses in supported_endpoints, and the snowflake, novita and sarvam rows have max_output_tokens above max_input_tokens; neither could be settled from an official table in this run, so they are unchanged
  • Together lists about 111 chat models with no row, but /v1/models carries no price fields, so they need a dedicated per-model pass
  • Groq, Nebius, Voyage and Jina have no credential in the environment or the shared vault; their public docs and catalog pages were used instead
  • Pricing spot-check against the raw OpenAI, Anthropic, Gemini, Vertex, Bedrock, Azure, Groq, Mistral, Cohere, DeepSeek, xAI and Fireworks tables found no new verifiable mismatch. xAI, Groq, Cohere and Mistral price pages are JS shells with no raw rows; Fireworks has no deprecations page (404)

Open registry PRs reviewed for absorption

Absorbed: #42923 (chatgpt reasoning flags plus test), #43016 (Bing pricing plus test, co-authored), #42727 (Bedrock Mantle prices, limits corrected from the model cards, co-authored), #42594 (7 azure_ai rows at price level, co-authored), #43048 (Mythos Preview base row plus test, co-authored), #43176 (azure gpt-5.6 alias rates plus test, co-authored), #43171 (anthropic batch prices, co-authored, its source edits dropped). Superseded by main: #42728 and #42704

Not absorbed, data PRs that could not be verified from an official source (unchanged from earlier runs): #41376 (xAI grok-4 flags, JS-rendered docs), #40368 (Claude 3 Haiku Bedrock cache prices listed as N/A by AWS), #36274 (Azure gpt-5.6 luna/terra rates contradict the Azure Retail Prices API), #36084, #35617, #32842 (undocumented openai/z-ai/glm-5.2 route), #32117 (Cohere embed v3 light, no published price), #31155 (GitHub Copilot rows, no official model table), #30775, #29920 (qwen3.7-max, no raw Alibaba price row). #41775 is a maintainer-authored together_ai sync left to its owner

Code or behaviour PRs that only incidentally touch the JSON, left alone: #43158 (Databricks gateway behaviour), #42840 (Sail provider), #40755 (external provider addition that also edits litellm/__init__.py, conflicting with main), #42880, #42854, #42829, #42822, #42789, #42726 and the rest of the open registry-touching PRs

Local checks at the PR tip (2026-09-25, 9a4779f)

$ python3 ci_cd/check_files_match.py
Passed! Files model_prices_and_context_window.json and litellm/model_prices_and_context_window_backup.json match.
$ uv run --no-sync python ci_cd/generate_model_prices_schema.py
$ LITELLM_LOCAL_MODEL_COST_MAP=True uv run --no-sync pytest -q tests/unit/test_model_prices_schema.py tests/unit/test_model_cost_aliases.py tests/local_testing/test_get_model_info.py tests/test_litellm/litellm_core_utils/llm_cost_calc/test_llm_cost_calc_utils.py
293 passed

Local checks at the 2026-09-24 tip

$ sha256sum model_prices_and_context_window.json litellm/model_prices_and_context_window_backup.json
65ba66a370fd59fb07ae6428dc0c8d897c69f3db11f350b1d5f353b59827ce2c  model_prices_and_context_window.json
65ba66a370fd59fb07ae6428dc0c8d897c69f3db11f350b1d5f353b59827ce2c  litellm/model_prices_and_context_window_backup.json
$ python3 ci_cd/check_files_match.py
Passed! Files model_prices_and_context_window.json and litellm/model_prices_and_context_window_backup.json match.
$ uv run --no-sync python ci_cd/generate_model_prices_schema.py --check
model_prices_and_context_window.schema.json is in sync and model_prices_and_context_window.json validates against it
$ LITELLM_LOCAL_MODEL_COST_MAP=True uv run --no-sync pytest -q tests/test_litellm/test_model_prices_schema.py tests/test_litellm/test_model_cost_aliases.py tests/search_tests/test_bing_grounding_search.py
79 passed, 3 skipped

With the two JSON files reverted to the state before the chatgpt fields and the new test kept: 1 failed (the 38 mismatches), 85 passed

Link to Devin session: https://app.devin.ai/sessions/efaa7951599c42d59888a499f118b53a
Open in Devin Desktop: https://app.devin.ai/desktop/session/efaa7951599c42d59888a499f118b53a?variant=devin


Note

Medium Risk
Changes directly affect reported model metadata and cost calculation for many deployments; incorrect values would mis-bill or mis-advertise capabilities, though the edits are mostly catalog data and additive pricing keys with tests.

Overview
This is a rolling model cost-map sync that refreshes model_prices_and_context_window.json (and its backup) with provider-sourced pricing, limits, deprecation metadata, and capability flags.

Registry highlights: deprecation_date on retiring OpenAI and Gemini preview models; reasoning annotations copied onto chatgpt/* twins to match their bare OpenAI rows; Bing Grounding per-query cost cut from $0.035 to $0.014; Azure gpt-5.6 alias rows aligned with GPT-5.6 Sol pricing; Anthropic Batch API *_batches rates (including above-200k tiers) on many Claude rows; priced Bedrock Mythos Preview base row and new bedrock_mantle/*, azure_ai/*, Nebius, OpenRouter, and xAI entries; Gemini native audio max_input_tokens corrected to 131072.

Plumbing: Extends the catalog/schema/types (Python, JSON schema, Rust ModelInfo, dashboard schema.d.ts) for new cost fields—especially cache_creation_input_token_cost_above_200k_tokens_batches—and registers that key in deployment pricing so batch/cache overrides bill correctly. Adds regression tests for chatgpt reasoning parity, Azure 5.6 alias pricing, Bedrock cross-region base pricing, Bing cost, and 200k batch deployment overrides.

Reviewed by Cursor Bugbot for commit 9a4779f. Bugbot is set up for automated code reviews on this repo. Configure here.

devin-ai-integration Bot and others added 2 commits September 24, 2026 13:31
…iew retirement dates

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Dmitry Voropaev <dy.voropaev@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@greptile-apps

greptile-apps Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[Medium risk] Updates pricing data and adds new cost fields for batch operations.

The PR appears safe to merge; the remaining batch-override test-coverage concern is non-blocking.

Findings

  1. P2 Batch override coverage is incomplete ▶

Summary

This PR updates model pricing, capabilities, limits, and retirement dates in the cost registry and its backup, with corresponding schema and pricing-plumbing changes.

  • Adds provider rows and corrects Bing, Azure alias, Bedrock Mythos, and Anthropic batch rates.
  • Adds regression tests for registry parity and deployment pricing overrides.

Reviews (14) · Last reviewed commit: "Merge remote-tracking branch 'origin/mai..."

Comment thread tests/test_litellm/test_model_prices_schema.py Outdated
@codspeed

codspeed Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_registry_audit_20260924 (9a4779f) with main (a11a93f)

Open in CodSpeed

@codecov

codecov Bot commented Sep 24, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ent on the openai row

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

Fixed in 77ebdb3 and 0295f42: the twin check now compares presence and value in both directions using a sentinel

devin-ai-integration Bot and others added 5 commits September 24, 2026 19:03
Co-authored-by: Vishnu Nair <nvishnu22@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Mariano Billinghurst <mariano.billinghurst@pmi.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration devin-ai-integration Bot changed the title fix(cost-map): add openai and gemini retirement dates and chatgpt reasoning flags fix(cost-map): retirement dates, chatgpt reasoning flags, bing pricing and bedrock mantle, azure ai and cohere rows Sep 24, 2026
devin-ai-integration Bot and others added 3 commits September 25, 2026 19:03
…aws price list

Co-authored-by: Yuneng Jiang <yuneng@berri.ai>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: krrish-berri-2 <270687000+krrish-berri-2@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

…t_20260924

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

# Conflicts:
#	litellm/litellm_core_utils/litellm_logging.py
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

1 issue from previous review remains unresolved.

Fix All in Cursor

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 9a4779f. Configure here.

@kerry-berri
kerry-berri merged commit 4179860 into main Sep 26, 2026
105 of 107 checks passed
yuneng-berri pushed a commit that referenced this pull request Sep 26, 2026
…g, bedrock mantle and mythos, azure gpt-5.6 alias, anthropic batch rates, new nebius, openrouter and xai rows (#42951)

(cherry picked from commit 4179860)
yuneng-berri added a commit that referenced this pull request Sep 26, 2026
… to rc/1.104.0 (#43366)

* fix(cost-map): retirement dates, chatgpt reasoning flags, bing pricing, bedrock mantle and mythos, azure gpt-5.6 alias, anthropic batch rates, new nebius, openrouter and xai rows (#42951)

(cherry picked from commit 4179860)

* ci: cut CircleCI wall time without loosening test isolation (#43347)

* ci: cut CircleCI wall time without loosening test isolation

* fix(ci): parse integration split files that follow --results

The CircleCI machine image ships Python 3.12.2, whose argparse leaves the
files positional empty when it follows an option and another positional, so
every extensions node exited with 'unrecognized arguments'. Reproduced on
3.12.2; parse_intermixed_args selects the files on 3.12.2, 3.12.13 and 3.13

* test(ci): resolve command references in the Rust toolchain guard

The Windows rustup install moved into the install_windows_toolchain command,
which the guard only recognized for install_rust. It now accepts any command
that installs a pinned rustup and reads the Windows toolchain pin from it

* ci: cache the Windows release cargo build from main

windows_release_wheel rebuilt every dependency with fat LTO on each run. It now
restores the release target and cargo registry saved by main's scheduled run,
drops the workspace crates' fingerprints so they always rebuild from the
checked-out source, and still runs the full LTO link

* ci: run the Windows release wheel build on windows.xlarge

The fat-LTO release build is the slowest job in the pipeline; more cores
speed up the dependency compile ahead of the final link

* ci: skip the Windows fingerprint cleanup when the cargo cache missed

On a cold cache the release fingerprint directory does not exist, and the
CircleCI PowerShell wrapper failed the step on the suppressed not-found error

(cherry picked from commit 635a718)

---------

Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: All chatgpt/* models missing reasoning annotations present on their openai/ and azure/ twins

3 participants