Skip to content

fix(models): rolling registry audit: Azure retirement dates, Bedrock Mantle Grok 4.3 context window - #41597

Merged
kerry-berri merged 5 commits into
mainfrom
litellm_registry_audit_2026_09_17
Sep 18, 2026
Merged

kerry-berri merged 5 commits into
mainfrom
litellm_registry_audit_2026_09_17

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • 53 Azure rows lacked or had wrong deprecation_date vs the Azure retirement schedule
  • bedrock_mantle/xai.grok-4.3 stored a 128k context window; AWS lists 1M

How it solves it:

  • Adds or corrects deprecation_date on the Azure rows, one field each
  • Sets max_input_tokens to 1048576 on the Grok 4.3 Mantle row and its us-gov-west-1 regional row
  • Root and backup registry files kept byte-identical

User Flow

Before: a proxy admin checking which deployments are about to retire gets an incomplete or wrong list

  1. They run the proxy with azure/gpt-4.1-nano and azure/eu/gpt-5 deployments and call GET http://localhost:4000/model/deprecations?warn_within_days=400
  2. gpt-4.1-nano comes back with deprecation_date: 2027-04-14, six months later than Azure's published 2026-10-14
  3. azure/eu/gpt-5 is not in the response at all, so nobody is warned about its 2027-02-09 retirement
  4. GET http://localhost:4000/model/info shows bedrock_mantle/xai.grok-4.3 with max_input_tokens: 131072, so long prompts are rejected or mis-routed before reaching Bedrock

After: the same calls return the dates Azure actually publishes and the context window AWS documents

  1. They run the proxy with the same config and call GET http://localhost:4000/model/deprecations?warn_within_days=400
  2. gpt-4.1-nano comes back with deprecation_date: 2026-10-14
  3. azure/eu/gpt-5 is listed with deprecation_date: 2027-02-09
  4. GET http://localhost:4000/model/info shows bedrock_mantle/xai.grok-4.3 with max_input_tokens: 1048576

Relevant issues

Supports #26900 (deprecation reporting), which reads deprecation_date from this registry

Affected release

Linear ticket

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests (data-only change; the existing schema and alias tests cover it, and vendor dates must not be pinned in tests)
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Shared setup: proxy started with LITELLM_LOCAL_MODEL_COST_MAP=True python litellm/proxy/proxy_cli.py --config /tmp/dep_config.yaml --port 4011, where the config has three deployments: gpt-4.1-nano -> azure/gpt-4.1-nano, gpt-5-eu -> azure/eu/gpt-5, grok-4.3 -> bedrock_mantle/xai.grok-4.3, and master_key: sk-1234. No LLM calls are made; these endpoints only read the registry

Before (4b368bf)

  1. curl -s 'http://localhost:4011/model/deprecations?warn_within_days=400' -H 'Authorization: Bearer sk-1234' | python3 -m json.tool | grep -E '"(model_name|litellm_model|deprecation_date|status)"'
"model_name": "gpt-4.1-nano",
"litellm_model": "azure/gpt-4.1-nano",
"deprecation_date": "2027-04-14",
"status": "imminent",
  1. curl -s http://localhost:4011/model/info -H 'Authorization: Bearer sk-1234' | python3 -c "import sys,json; [print(d['model_name'], 'max_input_tokens=', d['model_info'].get('max_input_tokens')) for d in json.load(sys.stdin)['data'] if d['model_name']=='grok-4.3']"
grok-4.3 max_input_tokens= 131072

After (f79c3eb)

  1. Same deprecations curl
"model_name": "gpt-4.1-nano",
"litellm_model": "azure/gpt-4.1-nano",
"deprecation_date": "2026-10-14",
"status": "imminent",
"model_name": "gpt-5-eu",
"litellm_model": "azure/eu/gpt-5",
"deprecation_date": "2027-02-09",
"status": "imminent",
  1. Same model/info curl
grok-4.3 max_input_tokens= 1048576

Type

🐛 Bug Fix

Caveats (if any)

Low

Evidence

Every value below was read from the raw HTML table rows of the official page (curl with a cache buster, <tr>/<td> parsed with a short Python snippet), with r.jina.ai as a second reader

Azure retirement schedule

Source: https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-retirement-schedule (main Azure OpenAI table, columns model | version | lifecycle | retirement date | replacement). The separate fine-tuned model table was not used

Corrected dates (registry had a later date than Azure publishes):

gpt-4.1-nano | 2025-04-14 | Legacy | 2026-10-14 | -
  azure/gpt-4.1-nano, azure/gpt-4.1-nano-2025-04-14, azure/us/gpt-4.1-nano-2025-04-14: 2027-04-14 -> 2026-10-14
  azure/eu/gpt-4.1-nano, azure/us/gpt-4.1-nano: none -> 2026-10-14
gpt-image-1.5 | 2025-12-16 | GA | 2026-12-16 | -
  azure/gpt-image-1.5, azure/gpt-image-1.5-2025-12-16: 2027-06-16 -> 2026-12-16

Dated snapshot keys that were missing a date (exact model-version match):

gpt-4o | 2024-05-13 | Deprecated | 2026-10-01     -> azure/eu/gpt-4o-2024-05-13, azure/us/gpt-4o-2024-05-13
o3 | 2025-04-16 | Deprecated | 2026-11-19         -> azure/eu/o3-2025-04-16
o4-mini | 2025-04-16 | Deprecated | 2026-11-19    -> azure/eu/o4-mini-2025-04-16

Data zone (azure/eu/..., azure/us/...) aliases that were missing the date already carried by the matching global azure/<model> row, where that date is the latest version's row in the schedule:

codex-mini | 2025-05-16 | Deprecated | 2026-11-15
gpt-4.1 | 2025-04-14 | Legacy | 2027-04-14
gpt-4.1-mini | 2025-04-14 | Legacy | 2027-04-14
gpt-5 | 2025-08-07 | GA | 2027-02-09
gpt-5-mini | 2025-08-07 | GA | 2027-02-09
gpt-5-nano | 2025-08-07 | GA | 2027-02-09
gpt-5-codex | 2025-09-15 | GA | 2027-03-17
gpt-5-pro | 2025-10-06 | GA | 2027-04-07
gpt-5.1-codex-max | 2025-12-04 | GA | 2027-05-18
gpt-5.2 | 2025-12-11 | GA | 2027-06-08
gpt-5.2-chat | 2026-02-10 | Retired | 2026-06-29
gpt-5.2-codex | 2026-01-14 | GA | 2027-07-13
gpt-5.3-chat | 2026-03-03 | Retired | 2026-06-29
gpt-5.3-codex | 2026-02-24 | GA | 2027-08-24
gpt-5.4-mini | 2026-03-17 | GA | 2027-09-21
gpt-5.4-nano | 2026-03-17 | GA | 2027-09-21
gpt-5.4-pro | 2026-03-05 | GA | 2027-09-07
o3-deep-research | 2025-06-26 | Deprecated | 2026-11-19
text-embedding-3-large | 1 | GA | 2028-02-09
text-embedding-3-small | 1 | GA | 2028-02-09
text-embedding-ada-002 | 2 | GA | 2028-02-09

Bedrock Mantle Grok 4.3

Source: https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-xai-grok-4-3.html, raw card row Context window: 1M tokens (1048576). bedrock_mantle/xai.grok-4.3.max_input_tokens 131072 -> 1048576. The same card lists us-gov-west-1 (GovCloud West) as a supported region for the same model, so bedrock_mantle/us-gov-west-1/xai.grok-4.3.max_input_tokens gets the same 131072 -> 1048576 (raised by Greptile on the first revision)

Gemini 3.5 Flash-Lite priority cache read: reverted (no change)

An earlier revision of this PR set gemini/gemini-3.5-flash-lite.cache_read_input_token_cost_priority 5e-08 -> 5.4e-08 and updated the pinned expectation in tests/test_litellm/litellm_core_utils/llm_cost_calc/test_llm_cost_calc_utils.py. That was a misread: the quoted "$0.54 / $4.50 / $0.054" row on https://ai.google.dev/gemini-api/docs/pricing is the Gemini 2.5 Flash priority row. The raw Gemini 3.5 Flash-Lite priority row reads Context caching price | $0.05 (text / image / video), so the existing 5e-08 is correct and both the registry value and the test expectation are restored in the follow-up commit. The Vertex row (vertex_ai/gemini-3.5-flash-lite, 5.4e-08 on main) is unchanged and matches the raw Vertex pricing row Gemini 3.5 Flash-Lite | Input | Global | $0.54 | $0.54 | $0.054 | $0.054; the two platforms publish different priority caching prices for this model

Providers re-read on 2026-09-17 with nothing to change

Model inventories were diffed against the registry via the official model APIs for OpenAI, Anthropic, Gemini, xAI, Mistral, DeepSeek and Together, and via the docs model tables for Cohere, Groq, Perplexity, Voyage and Nebius. Pricing pages were re-read raw for OpenAI, Anthropic, Gemini, Vertex, Together, Voyage, Mistral, DeepSeek and Perplexity. Nothing else verifiably missing or wrong was found. Not added, with the reason: Gemini deep-research-preview-04-2026, deep-research-max-preview-04-2026 and antigravity-preview-05-2026 are Interactions API agents that the pricing page bills "at standard Gemini list rates" of the underlying model, so there is no per-model price to record; Cohere command-a-plus-05-2026, rerank-v4.0-pro and rerank-v4.0-fast appear in the models table but https://cohere.com/pricing publishes no row for them; Fireworks qwen3p8-2p4t-a95b is marked Serverless: Not supported on https://fireworks.ai/models/fireworks/qwen3p8-2p4t-a95b so it has no per-token price (the Fireworks models API returned an account-suspended error, the pricing page and per-model page were read instead); the remaining Together inventory gaps have zero serverless pricing in the Together models API (dedicated-only). Bedrock pricing (https://aws.amazon.com/bedrock/pricing/) rendered unresolved {priceOf!...} placeholders in raw HTML and r.jina.ai, and the Azure models page https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/models returned 404 on both readers

Providers re-read with nothing to change

Bedrock lifecycle (https://docs.aws.amazon.com/bedrock/latest/userguide/model-lifecycle-legacy.html, all 14 EOL rows already match the registry), OpenAI deprecations and pricing, Anthropic, Gemini, Vertex, Mistral, Cohere, DeepSeek, Groq, Together. Fireworks deprecations (404) and the xAI models page (JS shell) had no parseable rows, so nothing was changed for them

Open registry PRs reviewed this run

Reviewed on 2026-09-17:

Re-verified against official sources in earlier runs and not absorbed:

Checks

python3 ci_cd/check_files_match.py passed, uv run --no-sync python ci_cd/generate_model_prices_schema.py produced no schema change, pytest tests/test_litellm/test_model_prices_schema.py tests/test_litellm/test_model_cost_aliases.py 53 passed, pytest tests/test_litellm/litellm_core_utils/llm_cost_calc/test_llm_cost_calc_utils.py 192 passed (re-run after the Gemini revert: 13 passed for -k "flash_lite or flash-lite or priority")

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Link to Devin session: https://app.devin.ai/sessions/6e030a36e04045c9bb7b376ede251746
Open in Devin Desktop: https://app.devin.ai/desktop/session/6e030a36e04045c9bb7b376ede251746?variant=devin

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR. Add '(aside)' to your comment to have me ignore it.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

CLAassistant commented Sep 17, 2026 •

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you all sign our Contributor License Agreement before we can accept your contribution.
0 out of 2 committers have signed the CLA.

❌ devin-ai-integration[bot]
❌ kerry-berri
You have signed the CLA already but the status is still pending? Let us recheck it.

@codspeed

codspeed Bot commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_registry_audit_2026_09_17 (86f6257) with main (c25c098)

Open in CodSpeed

@codecov

codecov Bot commented Sep 17, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@greptile-apps

greptile-apps Bot commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge, with no outstanding correctness or repository-rule issues.

Summary

Corrects model registry metadata across both synchronized registry copies.

  • Updates Azure model retirement dates for global and regional entries.
  • Raises the Bedrock Mantle Grok 4.3 context window to 1,048,576 tokens, including GovCloud.
  • Keeps the primary and backup registry files byte-identical and valid JSON.

Reviews (3) · Last reviewed commit: "Revert "fix(gemini): gemini-3.5-flash-li..."

Comment thread model_prices_and_context_window.json
…th model card

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration devin-ai-integration Bot changed the title fix(models): rolling registry audit: Azure retirement dates, Bedrock Mantle Grok 4.3 context window fix(models): rolling registry audit: Azure retirement dates, Bedrock Mantle Grok 4.3 context window, Gemini 3.5 Flash-Lite priority cache read Sep 17, 2026
…054/M"

This reverts commit 14e4b9f.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration devin-ai-integration Bot changed the title fix(models): rolling registry audit: Azure retirement dates, Bedrock Mantle Grok 4.3 context window, Gemini 3.5 Flash-Lite priority cache read fix(models): rolling registry audit: Azure retirement dates, Bedrock Mantle Grok 4.3 context window Sep 18, 2026

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 86f6257. Configure here.

},
"azure/gpt-image-1.5": {
"deprecation_date": "2027-06-16",
"deprecation_date": "2026-12-16",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Wrong gpt-image-1.5 retirement date

Medium Severity

deprecation_date on azure/gpt-image-1.5 and azure/gpt-image-1.5-2025-12-16 was moved from 2027-06-16 to 2026-12-16. Microsoft’s Foundry retirement pages still list June 16, 2027 for this GA image model, so deprecation reporting will warn six months early.

Additional Locations (2)
Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit 86f6257. Configure here.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-fetched the raw retirement schedule just now: the gpt-image-1.5 row reads 2025-12-16 | GA | 2026-12-16, so 2026-12-16 stands

@kerry-berri
kerry-berri merged commit 0342c75 into main Sep 18, 2026
89 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant