Repository navigation
fix(models): rolling registry audit: Azure retirement dates, Bedrock Mantle Grok 4.3 context window - #41597
Conversation
…xt window Azure schedule: https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-retirement-schedule AWS card: https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-xai-grok-4-3.html Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
🤖 Devin AI EngineerI'll be helping with this pull request! Here's what you should know: ✅ I will automatically:
Note: I can only respond to comments from users who have write access to this repository. ⚙️ Control Options:
|
|
|
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
|
…th model card Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…054/M" This reverts commit 14e4b9f. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 86f6257. Configure here.
| }, | ||
| "azure/gpt-image-1.5": { | ||
| "deprecation_date": "2027-06-16", | ||
| "deprecation_date": "2026-12-16", |
There was a problem hiding this comment.
Wrong gpt-image-1.5 retirement date
Medium Severity
deprecation_date on azure/gpt-image-1.5 and azure/gpt-image-1.5-2025-12-16 was moved from 2027-06-16 to 2026-12-16. Microsoft’s Foundry retirement pages still list June 16, 2027 for this GA image model, so deprecation reporting will warn six months early.
Additional Locations (2)
Reviewed by Cursor Bugbot for commit 86f6257. Configure here.
There was a problem hiding this comment.
Re-fetched the raw retirement schedule just now: the gpt-image-1.5 row reads 2025-12-16 | GA | 2026-12-16, so 2026-12-16 stands


TLDR
Problem this solves:
deprecation_datevs the Azure retirement schedulebedrock_mantle/xai.grok-4.3stored a 128k context window; AWS lists 1MHow it solves it:
deprecation_dateon the Azure rows, one field eachmax_input_tokensto 1048576 on the Grok 4.3 Mantle row and itsus-gov-west-1regional rowUser Flow
Before: a proxy admin checking which deployments are about to retire gets an incomplete or wrong list
azure/gpt-4.1-nanoandazure/eu/gpt-5deployments and call GET http://localhost:4000/model/deprecations?warn_within_days=400gpt-4.1-nanocomes back withdeprecation_date: 2027-04-14, six months later than Azure's published 2026-10-14azure/eu/gpt-5is not in the response at all, so nobody is warned about its 2027-02-09 retirementbedrock_mantle/xai.grok-4.3withmax_input_tokens: 131072, so long prompts are rejected or mis-routed before reaching BedrockAfter: the same calls return the dates Azure actually publishes and the context window AWS documents
gpt-4.1-nanocomes back withdeprecation_date: 2026-10-14azure/eu/gpt-5is listed withdeprecation_date: 2027-02-09bedrock_mantle/xai.grok-4.3withmax_input_tokens: 1048576Relevant issues
Supports #26900 (deprecation reporting), which reads
deprecation_datefrom this registryAffected release
Linear ticket
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Shared setup: proxy started with
LITELLM_LOCAL_MODEL_COST_MAP=True python litellm/proxy/proxy_cli.py --config /tmp/dep_config.yaml --port 4011, where the config has three deployments:gpt-4.1-nano->azure/gpt-4.1-nano,gpt-5-eu->azure/eu/gpt-5,grok-4.3->bedrock_mantle/xai.grok-4.3, andmaster_key: sk-1234. No LLM calls are made; these endpoints only read the registryBefore (4b368bf)
curl -s 'http://localhost:4011/model/deprecations?warn_within_days=400' -H 'Authorization: Bearer sk-1234' | python3 -m json.tool | grep -E '"(model_name|litellm_model|deprecation_date|status)"'curl -s http://localhost:4011/model/info -H 'Authorization: Bearer sk-1234' | python3 -c "import sys,json; [print(d['model_name'], 'max_input_tokens=', d['model_info'].get('max_input_tokens')) for d in json.load(sys.stdin)['data'] if d['model_name']=='grok-4.3']"After (f79c3eb)
Type
🐛 Bug Fix
Caveats (if any)
Low
azure/gpt-realtime-mini-2025-10-06left undated: Azure lists two retirement rows for that version (2026-09-21 and 2027-04-06)gpt-realtime-mini,gpt-4o-mini-transcribe,gpt-4o-mini-tts) left alone, as in fix(models): rolling registry audit: Gemini latest aliases, Nova cache pricing, OpenRouter/Together sync, Mistral GLM 5.3, Azure snapshots, Grok caching #41112max_output_tokensleft at 16384: the AWS card publishes no max outputEvidence
Every value below was read from the raw HTML table rows of the official page (curl with a cache buster,
<tr>/<td>parsed with a short Python snippet), with r.jina.ai as a second readerAzure retirement schedule
Source: https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-retirement-schedule (main Azure OpenAI table, columns model | version | lifecycle | retirement date | replacement). The separate fine-tuned model table was not used
Corrected dates (registry had a later date than Azure publishes):
Dated snapshot keys that were missing a date (exact model-version match):
Data zone (
azure/eu/...,azure/us/...) aliases that were missing the date already carried by the matching globalazure/<model>row, where that date is the latest version's row in the schedule:Bedrock Mantle Grok 4.3
Source: https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-xai-grok-4-3.html, raw card row
Context window: 1M tokens(1048576).bedrock_mantle/xai.grok-4.3.max_input_tokens131072 -> 1048576. The same card listsus-gov-west-1(GovCloud West) as a supported region for the same model, sobedrock_mantle/us-gov-west-1/xai.grok-4.3.max_input_tokensgets the same 131072 -> 1048576 (raised by Greptile on the first revision)Gemini 3.5 Flash-Lite priority cache read: reverted (no change)
An earlier revision of this PR set
gemini/gemini-3.5-flash-lite.cache_read_input_token_cost_priority5e-08 -> 5.4e-08 and updated the pinned expectation intests/test_litellm/litellm_core_utils/llm_cost_calc/test_llm_cost_calc_utils.py. That was a misread: the quoted "$0.54 / $4.50 / $0.054" row on https://ai.google.dev/gemini-api/docs/pricing is the Gemini 2.5 Flash priority row. The raw Gemini 3.5 Flash-Lite priority row readsContext caching price | $0.05 (text / image / video), so the existing 5e-08 is correct and both the registry value and the test expectation are restored in the follow-up commit. The Vertex row (vertex_ai/gemini-3.5-flash-lite, 5.4e-08 on main) is unchanged and matches the raw Vertex pricing rowGemini 3.5 Flash-Lite | Input | Global | $0.54 | $0.54 | $0.054 | $0.054; the two platforms publish different priority caching prices for this modelProviders re-read on 2026-09-17 with nothing to change
Model inventories were diffed against the registry via the official model APIs for OpenAI, Anthropic, Gemini, xAI, Mistral, DeepSeek and Together, and via the docs model tables for Cohere, Groq, Perplexity, Voyage and Nebius. Pricing pages were re-read raw for OpenAI, Anthropic, Gemini, Vertex, Together, Voyage, Mistral, DeepSeek and Perplexity. Nothing else verifiably missing or wrong was found. Not added, with the reason: Gemini
deep-research-preview-04-2026,deep-research-max-preview-04-2026andantigravity-preview-05-2026are Interactions API agents that the pricing page bills "at standard Gemini list rates" of the underlying model, so there is no per-model price to record; Coherecommand-a-plus-05-2026,rerank-v4.0-proandrerank-v4.0-fastappear in the models table but https://cohere.com/pricing publishes no row for them; Fireworksqwen3p8-2p4t-a95bis markedServerless: Not supportedon https://fireworks.ai/models/fireworks/qwen3p8-2p4t-a95b so it has no per-token price (the Fireworks models API returned an account-suspended error, the pricing page and per-model page were read instead); the remaining Together inventory gaps have zero serverless pricing in the Together models API (dedicated-only). Bedrock pricing (https://aws.amazon.com/bedrock/pricing/) rendered unresolved{priceOf!...}placeholders in raw HTML and r.jina.ai, and the Azure models page https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/models returned 404 on both readersProviders re-read with nothing to change
Bedrock lifecycle (https://docs.aws.amazon.com/bedrock/latest/userguide/model-lifecycle-legacy.html, all 14 EOL rows already match the registry), OpenAI deprecations and pricing, Anthropic, Gemini, Vertex, Mistral, Cohere, DeepSeek, Groq, Together. Fireworks deprecations (404) and the xAI models page (JS shell) had no parseable rows, so nothing was changed for them
Open registry PRs reviewed this run
Reviewed on 2026-09-17:
gpt-5.6-sol | $4.00 | $0.40 | $5.00 | $20.00) does not prove Bedrock's rates, so not verifiableazure/gpt-5.6-luna-2026-07-09alias): the Azure models page returned 404, so the snapshot name is not verifiableRe-verified against official sources in earlier runs and not absorbed:
serviceName eq 'Foundry Models', eastus) shows luna at $0.20 / $1.20 / $0.02 per M and terra at $2.00 / $12.00 / $0.20 per M for input / output / cached input, which is what main already has. The proposed 5x and 1.25x values match no meterDeepSeek-V3.2 | 32k tokens, matching the current 32768Context window: 256K tokensand128K tokens, matching the current 256000 / 128000 rather than 262144 / 131072. Only the Grok 4.3 context window (1M) was confirmed and is included here; no card publishes a max outputsourcefields; left to that bot's ownersChecks
python3 ci_cd/check_files_match.pypassed,uv run --no-sync python ci_cd/generate_model_prices_schema.pyproduced no schema change,pytest tests/test_litellm/test_model_prices_schema.py tests/test_litellm/test_model_cost_aliases.py53 passed,pytest tests/test_litellm/litellm_core_utils/llm_cost_calc/test_llm_cost_calc_utils.py192 passed (re-run after the Gemini revert: 13 passed for-k "flash_lite or flash-lite or priority")Final Attestation
Link to Devin session: https://app.devin.ai/sessions/6e030a36e04045c9bb7b376ede251746
Open in Devin Desktop: https://app.devin.ai/desktop/session/6e030a36e04045c9bb7b376ede251746?variant=devin