Repository navigation
fix(registry): sync Azure/Together deprecation dates, fix computer-use-preview and OpenRouter metadata - #40855
Conversation
…t-5.4-mini/nano context window Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
🤖 Devin AI EngineerI'll be helping with this pull request! Here's what you should know: ✅ I will automatically:
Note: I can only respond to comments from users who have write access to this repository. ⚙️ Control Options:
|
Greptile SummaryThis PR synchronizes model-registry metadata across the primary and bundled backup registries.
Confidence Score: 5/5The PR appears safe to merge with no outstanding correctness or repository-rule violations. The added PDF capability metadata satisfies the existing OpenAI vision-model registry convention and does not change request-path behavior. The previous input-limit thread was manually resolved without explanation and is not outstanding; the current registry retains the corrected 272K input limit.
|
| Filename | Overview |
|---|---|
| model_prices_and_context_window.json | Updates provider deprecation dates, model capabilities, token limits, and pricing metadata without an identified regression. |
| litellm/model_prices_and_context_window_backup.json | Mirrors the root registry changes and remains synchronized with the canonical file. |
Reviews (5): Last reviewed commit: "fix(registry): mark computer-use-preview..." | Re-trigger Greptile
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
bugbot run |
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
bugbot run |
…OpenRouter deepseek-v3.2 / claude-opus-4.6 metadata Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 414442c. Configure here.
TLDR
Problem this solves:
computer-use-previewis taggedazureand claims structured outputsHow it solves it:
deprecation_dateon 11 Azure o1/o3/o4 keys to 2026-11-19deprecation_dateon 11 Together AI keyscomputer-use-previewprovider,supports_response_schemaandsupports_pdf_inputThis is the rolling registry PR for the scheduled registry audit. All values below were read from the raw provider HTML or the provider's public API (curl, table rows or JSON parsed with python), cross-checked with r.jina.ai
Evidence rows
OpenAI
Source: https://platform.openai.com/docs/models/computer-use-preview (raw HTML, feature cells
text-gray-400 -> text-tertiary)Key
computer-use-preview(bare key, OpenAI direct):litellm_providerazure -> openai (the bare key was a copy ofazure/computer-use-preview; the Azure key is unchanged),supports_response_schematrue -> false,supports_pdf_inputtrue added (https://platform.openai.com/docs/guides/file-inputs: PDF input is available on models with vision capabilities, and this model lists Image input; the existingtest_openai_models_in_model_infoguard requires it for OpenAI vision models),sourceadded. Snapshotcomputer-use-preview-2025-03-11is marked Deprecated on the same page and was not added (deprecation work is out of scope here)OpenRouter
Source: https://openrouter.ai/api/v1/models (public API, fields
top_provider.max_completion_tokens,pricing.input_cache_read,supported_parameters)Key
openrouter/deepseek/deepseek-v3.2:max_output_tokensandmax_tokens163840 -> 65536,input_cost_per_token_cache_hit2.8e-08 -> 1.345e-07,supports_response_schematrue added,sourceadded. Keyopenrouter/anthropic/claude-opus-4.6:supports_response_schematrue added,sourceadded (limits and prices already matched the API)Azure
https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-retirement-schedule, "Azure OpenAI" table (model | version | status | retirement | replacement). The separate "Fine-tuned models" table has its own dates for o4-mini fine-tunes and is not applied to these base-model keys
Keys updated:
azure/o1,azure/o1-2024-12-17,azure/o3,azure/o3-2025-04-16,azure/o3-deep-research,azure/o3-mini,azure/o3-mini-2025-01-31,azure/o3-pro,azure/o3-pro-2025-06-10,azure/o4-mini,azure/o4-mini-2025-04-16, and the matching data-zone keysazure/us/o1-2024-12-17,azure/us/o3-2025-04-16,azure/us/o3-mini-2025-01-31,azure/us/o4-mini-2025-04-16,azure/eu/o1-2024-12-17,azure/eu/o3-mini-2025-01-31(the Azure schedule is per model, not per data zone). Previous registry values were 2026-10-01, 2026-10-16, 2026-10-21, 2026-12-17, 2026-12-26 or absentTogether AI
https://docs.together.ai/docs/deprecations. Scheduled deprecations table (date | model | replacement | dedicated available) and inference deprecation history table (date | model | dedicated available)
Keys updated: the matching
together_ai/<model>keys.together_ai/Qwen/Qwen3-235B-A22B-Thinking-2507was left at 2026-04-16 because the page lists it in two tables with different dates (2026-04-16 and 2026-03-06) and the scope of the second table is not clear from the pageChecked, not changed (missing-model and metadata audit, 2026-09-12 run)
Provider model lists were diffed against the registry using official APIs or raw docs. Apparent gaps were first normalized against LiteLLM key conventions (bare vs
openai/,us./global./eu.Bedrock prefixes,groq/<org>/<model>), which removed nearly all of themGET /v1/models, 138 IDs): all present. Onlycomputer-use-preview-2025-03-11is missing and the docs mark it DeprecatedGET /v1/models), xAI (/v1/models), Mistral (/v1/models), DeepSeek (/models): no missing IDsv1beta/models, 40 generateContent IDs): onlyantigravity-preview-05-2026,deep-research(-max)-preview-04-2026,aqa,lyria-realtime-explack keys; no per-token prices for them on https://ai.google.dev/gemini-api/docs/pricing, so not addedListFoundationModels, us-east-1, 120 IDs): remaining gaps are provisioned-throughput variants (amazon.nova-*:24k/300k,titan-embed-*:8k,cohere.embed-*:512) pluswriter.palmyra-vision-7b. The AWS offer file (https://b0.p.awsstatic.com/pricing/2.0/meteredUnitMaps/bedrockfoundationmodels/USD/current/bedrockfoundationmodels.json, resolved against the row hashes in https://aws.amazon.com/bedrock/pricing/) gives Palmyra Vision 7B $0.15 / $0.60 per 1M, but the AWS model docs expose no context or output limit for it, so no entry was addedgroq/compoundandgroq/compound-minihave no token price in the table,minimaxai/minimax-m2.7is Contact Sales; not addedcommand-a-plus-05-2026,command-a-translate-08-2025,command-a-vision-07-2025, Tiny Aya, North, Transcribe) have no per-token API price on https://cohere.com/pricing (Command A+ is listed as a download); not addedfireworks_ai/key; no change/v1/models, 275 IDs): the 220 unkeyed IDs are zero-priced, dedicated-only, image, video, embedding or rerank types; none carried a serverless per-token price to add/api/v1/models, 418 priced IDs): 179 unkeyed IDs are upstream aliases or provider routes (~openai/...,inference-net/...); the two metadata fixes above were the only verified corrections on existing keyssupported_endpoints,*-latestdrift,max_output_tokens>max_input_tokens) produced candidates only; none had an official row contradicting the registryOther open registry PRs
Absorbed at field level: #40840 (Together deprecation dates, verified above; its numeric reformatting of unrelated entries was not carried over)
Closed as already handled on the base branch: #30019 wanted gpt-5.4-mini
max_input_tokens400000, but the registry follows the OpenAI convention where the 400k context window is 272k input plus 128k output (same as gpt-5, gpt-5.1, gpt-5-mini), and e0946cc already pinned these limits withtests/test_litellm/test_gpt_5_4_model_metadata.py. An earlier revision of this PR briefly applied 400000; the tip keeps 272000Not absorbed, not verifiable from an official numeric source: #40368 (Bedrock Claude 3 Haiku cache read/write rates; the AWS offer file row for Claude 3 Haiku in US East (N. Virginia) resolves to
0.25 | 1.25 | 0.125 | 0.625 | N/A | N/A | N/Afor input, output, batch input, batch output, 5m cache write, 1h cache write, cache read, so AWS publishes no cache meters for that model), #40849 (Bedrock Mantle GovCloud gpt-5.6-sol; the Bedrock pricing page has no gpt-5.6 rows and https://aws.amazon.com/bedrock/mantle/ renders no pricing via curl or r.jina.ai), #29920 (Qwen 3.7 Max), #40755 (Phoenix Grove, mixed with code and UI changes), #40832 (about 200 changed entries across Gemini, Together and Fireworks; too broad to re-verify row by row this run, left open for its author)Left alone as code or mixed PRs whose registry edits are incidental: #40627 (realtime cache-read billing fix), #40018 (Gemini Live and Azure realtime, its Vertex corrections are already on the base branch, its remaining cost-calculator changes are code), #37972 (new provider), #40370, #39861, #39845, #40871 and #40873 (automated provider-info sync against
main)User Flow
Before: an operator loads the registry with
LITELLM_LOCAL_MODEL_COST_MAP=Trueand readslitellm.model_cost["azure/o3-mini"]. There is nodeprecation_date, so nothing warns them the deployment retires on 2026-11-19. Same fortogether_ai/zai-org/GLM-5.litellm.get_llm_provider("computer-use-preview")raisesBadRequestError: LLM Provider NOT providedbecause the registry tags the bare key asazure, andlitellm.model_cost["openrouter/deepseek/deepseek-v3.2"]["max_output_tokens"]says 163840, so amax_tokensabove 65536 passes local validation and fails at OpenRouterAfter: the same lookup returns
deprecation_date2026-11-19 forazure/o3-miniand 2026-06-22 fortogether_ai/zai-org/GLM-5, ready for the #26900 deprecation handling.litellm.get_llm_provider("computer-use-preview")resolves toopenaiandsupports_response_schemais false;openrouter/deepseek/deepseek-v3.2reportsmax_output_tokens65536 and a cache-read price of 1.345e-07GET /v1/model/info does not surface
deprecation_dateyet, so on the proxy endpoint those values show asnullon both sides todayRelevant issues
Supports #26900. Supersedes #40840. Closes #30019 as already handled
Linear ticket
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Local checks at the tip (414442c):
python3 ci_cd/check_files_match.pypasses,uv run --no-sync python ci_cd/generate_model_prices_schema.pyleaves the schema unchanged,LITELLM_LOCAL_MODEL_COST_MAP=True uv run --no-sync pytest -q tests/test_litellm/test_model_prices_schema.py tests/test_litellm/test_model_cost_aliases.py tests/test_litellm/test_utils.py -k "model_prices or aliases or test_openai_models_in_model_info"gives 52 passedBefore (9071ca5)
LITELLM_LOCAL_MODEL_COST_MAP=True python3 -c "import litellm; m=litellm.model_cost; print(m['azure/o3-mini'].get('deprecation_date'), m['together_ai/zai-org/GLM-5'].get('deprecation_date'), m['computer-use-preview']['litellm_provider'], m['computer-use-preview']['supports_response_schema'], m['computer-use-preview'].get('supports_pdf_input'), m['openrouter/deepseek/deepseek-v3.2']['max_output_tokens'], m['openrouter/deepseek/deepseek-v3.2']['input_cost_per_token_cache_hit'])"2026-10-01 None azure True None 163840 2.8e-08After (414442c)
2026-11-19 2026-06-22 openai False True 65536 1.345e-07Type
🐛 Bug Fix
Caveats (if any)
Low
deprecation_dateis stored but not returned by GET /v1/model/info until feat(proxy): proactive model deprecation alerts and/model/deprecationsendpoint #26900 landsQwen3-235B-A22B-Thinking-2507kept its existing date; the page lists two conflicting datesazure/computer-use-previewkeepssupports_response_schematrue; Microsoft's Azure docs were not re-checked for that flag in this runQA runbook
Run the command from the After section with
LITELLM_LOCAL_MODEL_COST_MAP=True; it prints2026-11-19 2026-06-22 openai False True 65536 1.345e-07Final Attestation
get_model_info,get_llm_provider,supports_response_schema,supports_pdf_inputand cost lookups for the three keys, A/B'd at 9071ca5 vs 414442c withLITELLM_LOCAL_MODEL_COST_MAP=True; only the intended values changed, and the barecomputer-use-previewkey now resolves a provider where the base raised)Link to Devin session: https://app.devin.ai/sessions/f7c8f7d679564891a3e6218f6b202aeb
Open in Devin Desktop: https://app.devin.ai/desktop/session/f7c8f7d679564891a3e6218f6b202aeb?variant=devin
Requested by: @mateo-berri
Note
Low Risk
Data-only registry updates; behavior shifts are limited to deprecation warnings, provider resolution for
computer-use-preview, and local validation/cost for a few model keys.Overview
This PR refreshes the model cost/registry JSON (
model_prices_and_context_window.jsonand its backup) so LiteLLM’s local metadata matches current provider docs.Azure OpenAI:
deprecation_dateon many o1/o3/o4 entries (including EU/US data-zone keys) is set or corrected to 2026-11-19, replacing earlier October/December 2026 values or missing dates.Together AI: Adds
deprecation_dateon several serverless models (e.g. gpt-oss-20b, Gemma, Kimi, GLM, Qwen variants) per Together’s deprecation tables.OpenAI
computer-use-preview:litellm_providerchanges fromazuretoopenai;supports_response_schemabecomes false;supports_pdf_inputand asourceURL are added.OpenRouter:
openrouter/deepseek/deepseek-v3.2gets correctedmax_output_tokens/max_tokens(65536),input_cost_per_token_cache_hit, andsupports_response_schema;openrouter/anthropic/claude-opus-4.6gainssupports_response_schemaandsource.Reviewed by Cursor Bugbot for commit 414442c. Bugbot is set up for automated code reviews on this repo. Configure here.