Skip to content

fix(registry): sync Azure/Together deprecation dates, fix computer-use-preview and OpenRouter metadata - #40855

Merged
mateo-berri merged 6 commits into
litellm_internal_stagingfrom
litellm_registry_audit_2026_09_12
Sep 12, 2026
Merged

mateo-berri merged 6 commits into
litellm_internal_stagingfrom
litellm_registry_audit_2026_09_12

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 12, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Azure o-series entries carry stale or missing retirement dates
  • Together AI entries miss provider-announced deprecation dates
  • computer-use-preview is tagged azure and claims structured outputs
  • OpenRouter deepseek-v3.2 output limit and cache-read price are wrong

How it solves it:

  • Set deprecation_date on 11 Azure o1/o3/o4 keys to 2026-11-19
  • Add or correct deprecation_date on 11 Together AI keys
  • Fix computer-use-preview provider, supports_response_schema and supports_pdf_input
  • Fix OpenRouter deepseek-v3.2 limits and cache price, flag structured outputs
  • Root registry and bundled backup stay byte-identical

This is the rolling registry PR for the scheduled registry audit. All values below were read from the raw provider HTML or the provider's public API (curl, table rows or JSON parsed with python), cross-checked with r.jina.ai

Evidence rows

OpenAI

Source: https://platform.openai.com/docs/models/computer-use-preview (raw HTML, feature cells text-gray-400 -> text-tertiary)

8,192 context window | 1,024 max output tokens | Price $3 / $12 per 1M
Streaming -> Not supported
Structured outputs -> Not supported
Fine-tuning -> Not supported
Predicted outputs -> Not supported

Key computer-use-preview (bare key, OpenAI direct): litellm_provider azure -> openai (the bare key was a copy of azure/computer-use-preview; the Azure key is unchanged), supports_response_schema true -> false, supports_pdf_input true added (https://platform.openai.com/docs/guides/file-inputs: PDF input is available on models with vision capabilities, and this model lists Image input; the existing test_openai_models_in_model_info guard requires it for OpenAI vision models), source added. Snapshot computer-use-preview-2025-03-11 is marked Deprecated on the same page and was not added (deprecation work is out of scope here)

OpenRouter

Source: https://openrouter.ai/api/v1/models (public API, fields top_provider.max_completion_tokens, pricing.input_cache_read, supported_parameters)

deepseek/deepseek-v3.2 | context_length 163840 | max_completion_tokens 65536 | prompt 0.000000269 | completion 0.0000004 | input_cache_read 0.0000001345 | supported_parameters includes response_format, structured_outputs, tools, tool_choice
anthropic/claude-opus-4.6 | context_length 1000000 | max_completion_tokens 128000 | prompt 0.000005 | completion 0.000025 | input_cache_read 0.0000005 | input_cache_write 0.00000625 | supported_parameters includes response_format, structured_outputs, tools, tool_choice

Key openrouter/deepseek/deepseek-v3.2: max_output_tokens and max_tokens 163840 -> 65536, input_cost_per_token_cache_hit 2.8e-08 -> 1.345e-07, supports_response_schema true added, source added. Key openrouter/anthropic/claude-opus-4.6: supports_response_schema true added, source added (limits and prices already matched the API)

Azure

https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-retirement-schedule, "Azure OpenAI" table (model | version | status | retirement | replacement). The separate "Fine-tuned models" table has its own dates for o4-mini fine-tunes and is not applied to these base-model keys

o1 | 2024-12-17 | Deprecated | 2026-11-19 | gpt-5.6-sol
o3 | 2025-04-16 | Deprecated | 2026-11-19 | gpt-5.6-sol
o3-deep-research | 2025-06-26 | Deprecated | 2026-11-19 | gpt-5.6-sol
o3-mini | 2025-01-31 | Deprecated | 2026-11-19 | gpt-5.6-terra
o3-pro | 2025-06-10 | Deprecated | 2026-11-19 | gpt-5.6-sol
o4-mini | 2025-04-16 | Deprecated | 2026-11-19 | gpt-5.6-terra

Keys updated: azure/o1, azure/o1-2024-12-17, azure/o3, azure/o3-2025-04-16, azure/o3-deep-research, azure/o3-mini, azure/o3-mini-2025-01-31, azure/o3-pro, azure/o3-pro-2025-06-10, azure/o4-mini, azure/o4-mini-2025-04-16, and the matching data-zone keys azure/us/o1-2024-12-17, azure/us/o3-2025-04-16, azure/us/o3-mini-2025-01-31, azure/us/o4-mini-2025-04-16, azure/eu/o1-2024-12-17, azure/eu/o3-mini-2025-01-31 (the Azure schedule is per model, not per data zone). Previous registry values were 2026-10-01, 2026-10-16, 2026-10-21, 2026-12-17, 2026-12-26 or absent

Together AI

https://docs.together.ai/docs/deprecations. Scheduled deprecations table (date | model | replacement | dedicated available) and inference deprecation history table (date | model | dedicated available)

2026-09-14 | openai/gpt-oss-20b | Qwen/Qwen3.5-9B | Yes
2026-09-14 | google/gemma-4-31B-it | zai-org/GLM-5.3-Flash | Yes
2026-09-14 | thinkingmachines/Inkling-Small | zai-org/GLM-5.3-Flash | Yes
2026-09-14 | intfloat/multilingual-e5-large-instruct |  | No
2026-08-19 | moonshotai/Kimi-K2.6 | Yes
2026-07-10 | zai-org/GLM-5.1 | Yes
2026-06-22 | zai-org/GLM-5 | No
2026-05-14 | Qwen/Qwen3-Coder-Next-FP8 | Yes
2026-04-16 | Qwen/Qwen3-VL-8B-Instruct | Yes
2026-02-25 | Qwen/Qwen3-VL-32B-Instruct | No
2025-11-13 | Qwen/QwQ-32B | No

Keys updated: the matching together_ai/<model> keys. together_ai/Qwen/Qwen3-235B-A22B-Thinking-2507 was left at 2026-04-16 because the page lists it in two tables with different dates (2026-04-16 and 2026-03-06) and the scope of the second table is not clear from the page

Checked, not changed (missing-model and metadata audit, 2026-09-12 run)

Provider model lists were diffed against the registry using official APIs or raw docs. Apparent gaps were first normalized against LiteLLM key conventions (bare vs openai/, us./global./eu. Bedrock prefixes, groq/<org>/<model>), which removed nearly all of them

  • OpenAI (GET /v1/models, 138 IDs): all present. Only computer-use-preview-2025-03-11 is missing and the docs mark it Deprecated
  • Anthropic (GET /v1/models), xAI (/v1/models), Mistral (/v1/models), DeepSeek (/models): no missing IDs
  • Gemini (v1beta/models, 40 generateContent IDs): only antigravity-preview-05-2026, deep-research(-max)-preview-04-2026, aqa, lyria-realtime-exp lack keys; no per-token prices for them on https://ai.google.dev/gemini-api/docs/pricing, so not added
  • Bedrock (ListFoundationModels, us-east-1, 120 IDs): remaining gaps are provisioned-throughput variants (amazon.nova-*:24k/300k, titan-embed-*:8k, cohere.embed-*:512) plus writer.palmyra-vision-7b. The AWS offer file (https://b0.p.awsstatic.com/pricing/2.0/meteredUnitMaps/bedrockfoundationmodels/USD/current/bedrockfoundationmodels.json, resolved against the row hashes in https://aws.amazon.com/bedrock/pricing/) gives Palmyra Vision 7B $0.15 / $0.60 per 1M, but the AWS model docs expose no context or output limit for it, so no entry was added
  • Groq (https://console.groq.com/docs/models raw table): groq/compound and groq/compound-mini have no token price in the table, minimaxai/minimax-m2.7 is Contact Sales; not added
  • Cohere (https://docs.cohere.com/docs/models): newer IDs (command-a-plus-05-2026, command-a-translate-08-2025, command-a-vision-07-2025, Tiny Aya, North, Transcribe) have no per-token API price on https://cohere.com/pricing (Command A+ is listed as a download); not added
  • Fireworks: fireworks.ai/models and the pricing page return no rows via curl and the API returned HTTP 412, so https://docs.fireworks.ai/serverless/pricing raw HTML (32 rows) and r.jina.ai were used. Every priced row already has a fireworks_ai/ key; no change
  • Together (/v1/models, 275 IDs): the 220 unkeyed IDs are zero-priced, dedicated-only, image, video, embedding or rerank types; none carried a serverless per-token price to add
  • OpenRouter (/api/v1/models, 418 priced IDs): 179 unkeyed IDs are upstream aliases or provider routes (~openai/..., inference-net/...); the two metadata fixes above were the only verified corrections on existing keys
  • Perplexity (Router docs), Nebius (pricing page states per-token prices are not published), Voyage and rerank providers: existing keys match the official tables; no change
  • Mechanical scans (mode vs supported_endpoints, *-latest drift, max_output_tokens > max_input_tokens) produced candidates only; none had an official row contradicting the registry

Other open registry PRs

Absorbed at field level: #40840 (Together deprecation dates, verified above; its numeric reformatting of unrelated entries was not carried over)

Closed as already handled on the base branch: #30019 wanted gpt-5.4-mini max_input_tokens 400000, but the registry follows the OpenAI convention where the 400k context window is 272k input plus 128k output (same as gpt-5, gpt-5.1, gpt-5-mini), and e0946cc already pinned these limits with tests/test_litellm/test_gpt_5_4_model_metadata.py. An earlier revision of this PR briefly applied 400000; the tip keeps 272000

Not absorbed, not verifiable from an official numeric source: #40368 (Bedrock Claude 3 Haiku cache read/write rates; the AWS offer file row for Claude 3 Haiku in US East (N. Virginia) resolves to 0.25 | 1.25 | 0.125 | 0.625 | N/A | N/A | N/A for input, output, batch input, batch output, 5m cache write, 1h cache write, cache read, so AWS publishes no cache meters for that model), #40849 (Bedrock Mantle GovCloud gpt-5.6-sol; the Bedrock pricing page has no gpt-5.6 rows and https://aws.amazon.com/bedrock/mantle/ renders no pricing via curl or r.jina.ai), #29920 (Qwen 3.7 Max), #40755 (Phoenix Grove, mixed with code and UI changes), #40832 (about 200 changed entries across Gemini, Together and Fireworks; too broad to re-verify row by row this run, left open for its author)

Left alone as code or mixed PRs whose registry edits are incidental: #40627 (realtime cache-read billing fix), #40018 (Gemini Live and Azure realtime, its Vertex corrections are already on the base branch, its remaining cost-calculator changes are code), #37972 (new provider), #40370, #39861, #39845, #40871 and #40873 (automated provider-info sync against main)

User Flow

Before: an operator loads the registry with LITELLM_LOCAL_MODEL_COST_MAP=True and reads litellm.model_cost["azure/o3-mini"]. There is no deprecation_date, so nothing warns them the deployment retires on 2026-11-19. Same for together_ai/zai-org/GLM-5. litellm.get_llm_provider("computer-use-preview") raises BadRequestError: LLM Provider NOT provided because the registry tags the bare key as azure, and litellm.model_cost["openrouter/deepseek/deepseek-v3.2"]["max_output_tokens"] says 163840, so a max_tokens above 65536 passes local validation and fails at OpenRouter

After: the same lookup returns deprecation_date 2026-11-19 for azure/o3-mini and 2026-06-22 for together_ai/zai-org/GLM-5, ready for the #26900 deprecation handling. litellm.get_llm_provider("computer-use-preview") resolves to openai and supports_response_schema is false; openrouter/deepseek/deepseek-v3.2 reports max_output_tokens 65536 and a cache-read price of 1.345e-07

GET /v1/model/info does not surface deprecation_date yet, so on the proxy endpoint those values show as null on both sides today

Relevant issues

Supports #26900. Supersedes #40840. Closes #30019 as already handled

Linear ticket

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests (data-only change; covered by the existing schema and cost-map guard tests)
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Local checks at the tip (414442c): python3 ci_cd/check_files_match.py passes, uv run --no-sync python ci_cd/generate_model_prices_schema.py leaves the schema unchanged, LITELLM_LOCAL_MODEL_COST_MAP=True uv run --no-sync pytest -q tests/test_litellm/test_model_prices_schema.py tests/test_litellm/test_model_cost_aliases.py tests/test_litellm/test_utils.py -k "model_prices or aliases or test_openai_models_in_model_info" gives 52 passed

Before (9071ca5)

  1. LITELLM_LOCAL_MODEL_COST_MAP=True python3 -c "import litellm; m=litellm.model_cost; print(m['azure/o3-mini'].get('deprecation_date'), m['together_ai/zai-org/GLM-5'].get('deprecation_date'), m['computer-use-preview']['litellm_provider'], m['computer-use-preview']['supports_response_schema'], m['computer-use-preview'].get('supports_pdf_input'), m['openrouter/deepseek/deepseek-v3.2']['max_output_tokens'], m['openrouter/deepseek/deepseek-v3.2']['input_cost_per_token_cache_hit'])"
  2. Output: 2026-10-01 None azure True None 163840 2.8e-08

After (414442c)

  1. Same command
  2. Output: 2026-11-19 2026-06-22 openai False True 65536 1.345e-07

Type

🐛 Bug Fix

Caveats (if any)

Low

  • deprecation_date is stored but not returned by GET /v1/model/info until feat(proxy): proactive model deprecation alerts and /model/deprecations endpoint #26900 lands
  • Azure fine-tuned o4-mini deployments retire later (2027) than the base model date applied here
  • Together Qwen3-235B-A22B-Thinking-2507 kept its existing date; the page lists two conflicting dates
  • azure/computer-use-preview keeps supports_response_schema true; Microsoft's Azure docs were not re-checked for that flag in this run

QA runbook

Run the command from the After section with LITELLM_LOCAL_MODEL_COST_MAP=True; it prints 2026-11-19 2026-06-22 openai False True 65536 1.345e-07

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR
  • 414442c passes /live-pr-risk (data-only diff; dependents are get_model_info, get_llm_provider, supports_response_schema, supports_pdf_input and cost lookups for the three keys, A/B'd at 9071ca5 vs 414442c with LITELLM_LOCAL_MODEL_COST_MAP=True; only the intended values changed, and the bare computer-use-preview key now resolves a provider where the base raised)

Link to Devin session: https://app.devin.ai/sessions/f7c8f7d679564891a3e6218f6b202aeb
Open in Devin Desktop: https://app.devin.ai/desktop/session/f7c8f7d679564891a3e6218f6b202aeb?variant=devin
Requested by: @mateo-berri


Note

Low Risk
Data-only registry updates; behavior shifts are limited to deprecation warnings, provider resolution for computer-use-preview, and local validation/cost for a few model keys.

Overview
This PR refreshes the model cost/registry JSON (model_prices_and_context_window.json and its backup) so LiteLLM’s local metadata matches current provider docs.

Azure OpenAI: deprecation_date on many o1/o3/o4 entries (including EU/US data-zone keys) is set or corrected to 2026-11-19, replacing earlier October/December 2026 values or missing dates.

Together AI: Adds deprecation_date on several serverless models (e.g. gpt-oss-20b, Gemma, Kimi, GLM, Qwen variants) per Together’s deprecation tables.

OpenAI computer-use-preview: litellm_provider changes from azure to openai; supports_response_schema becomes false; supports_pdf_input and a source URL are added.

OpenRouter: openrouter/deepseek/deepseek-v3.2 gets corrected max_output_tokens/max_tokens (65536), input_cost_per_token_cache_hit, and supports_response_schema; openrouter/anthropic/claude-opus-4.6 gains supports_response_schema and source.

Reviewed by Cursor Bugbot for commit 414442c. Bugbot is set up for automated code reviews on this repo. Configure here.

…t-5.4-mini/nano context window

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR. Add '(aside)' to your comment to have me ignore it.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@greptile-apps

greptile-apps Bot commented Sep 12, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR synchronizes model-registry metadata across the primary and bundled backup registries.

  • Updates Azure and Together AI deprecation dates.
  • Corrects the bare computer-use-preview provider and capability metadata.
  • Corrects OpenRouter output-limit, cache-pricing, and structured-output metadata.
  • Keeps both registry files byte-identical.

Confidence Score: 5/5

The PR appears safe to merge with no outstanding correctness or repository-rule violations.

The added PDF capability metadata satisfies the existing OpenAI vision-model registry convention and does not change request-path behavior. The previous input-limit thread was manually resolved without explanation and is not outstanding; the current registry retains the corrected 272K input limit.

Important Files Changed

Filename Overview
model_prices_and_context_window.json Updates provider deprecation dates, model capabilities, token limits, and pricing metadata without an identified regression.
litellm/model_prices_and_context_window_backup.json Mirrors the root registry changes and remains synchronized with the canonical file.

Reviews (5): Last reviewed commit: "fix(registry): mark computer-use-preview..." | Re-trigger Greptile

Comment thread model_prices_and_context_window.json Outdated
@codspeed

codspeed Bot commented Sep 12, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_registry_audit_2026_09_12 (414442c) with litellm_internal_staging (a73454b)1

Open in CodSpeed

Footnotes

  1. No successful run was found on litellm_internal_staging (559247f) during the generation of this report, so a73454b was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

@codecov

codecov Bot commented Sep 12, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration devin-ai-integration Bot changed the title fix(registry): sync Azure o-series and Together deprecation dates, gpt-5.4-mini/nano context window fix(registry): sync Azure o-series and Together AI deprecation dates Sep 12, 2026
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread model_prices_and_context_window.json
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

mateo-berri and others added 2 commits September 12, 2026 19:04
…OpenRouter deepseek-v3.2 / claude-opus-4.6 metadata

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration devin-ai-integration Bot changed the title fix(registry): sync Azure o-series and Together AI deprecation dates fix(registry): sync Azure/Together deprecation dates, fix computer-use-preview and OpenRouter metadata Sep 12, 2026
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 414442c. Configure here.

@mateo-berri mateo-berri left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@mateo-berri
mateo-berri merged commit 9276317 into litellm_internal_staging Sep 12, 2026
87 checks passed
@mateo-berri
mateo-berri deleted the litellm_registry_audit_2026_09_12 branch September 12, 2026 22:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant