Repository navigation
fix(azure_ai): price seven Foundry catalog names and charge the model router fee once - #40189
Conversation
… router fee once Add cost map entries for azure_ai/gpt-chat-latest, codex-mini, whisper, model-router, cohere-command-a, grok-4-20-reasoning, and grok-4-20-non-reasoning, priced from the live Azure AI Foundry and Azure OpenAI pricing pages and the Azure Retail Prices API. Skip the model router flat fee when the response model is the router entry itself, since the generic cost already priced that fee. Before, azure_ai/model_router charged it twice. Resolves LIT-3157
🤖 Devin AI EngineerI'll be helping with this pull request! Here's what you should know: ✅ I will automatically:
Note: I can only respond to comments from users who have write access to this repository. ⚙️ Control Options:
|
Merging this PR will not alter performance
Comparing |
Greptile SummaryThis PR adds pricing and metadata for seven Azure AI Foundry catalog names, ensures the Model Router fee is charged exactly once, and preserves the Azure AI provider identity for transcription cost calculation.
Confidence Score: 5/5The PR appears safe to merge with no actionable regressions remaining. The current implementation preserves provider attribution for transcription and charges Azure AI Model Router fees once across direct and routed response shapes. Earlier findings about the removed public keyword and test placement were manually resolved without explanation; the provider-boundary and substring findings were withdrawn after confirming they were not introduced by this PR, and the test-coverage concern was conceded after the cached-token case was restored.
|
| Filename | Overview |
|---|---|
| litellm/llms/azure_ai/cost_calculator.py | Refactors Model Router detection and fee calculation so routed and router-shaped responses are charged once. |
| litellm/cost_calculator.py | Avoids applying the separate Azure AI additional-cost path when the priced model already carries the router fee. |
| litellm/llms/azure/audio_transcriptions.py | Preserves the selected Azure or Azure AI provider in transcription response metadata. |
| litellm/main.py | Passes the parsed provider through the shared Azure transcription implementation. |
| model_prices_and_context_window.json | Adds the seven Azure AI Foundry pricing and capability entries. |
| litellm/model_prices_and_context_window_backup.json | Mirrors the new Foundry catalog metadata in the backup cost map. |
| tests/test_litellm/llms/azure_ai/test_azure_ai_cost_calculator.py | Covers direct router names, routed responses, cached prompts, public keyword compatibility, and cost breakdowns. |
| tests/test_litellm/llms/azure/test_audio_transcriptions.py | Verifies Azure AI transcription pricing and unchanged Azure provider attribution. |
| tests/test_litellm/llms/azure_ai/test_azure_ai_foundry_catalog_model_metadata.py | Validates catalog registration, cost-map consistency, and model metadata. |
Reviews (8): Last reviewed commit: "fix(azure): price azure_ai transcription..." | Re-trigger Greptile
…d azure_ai directory The new metadata test sat at the top of tests/test_litellm. The azure_ai metadata tests live in tests/test_litellm/llms/azure_ai next to the cost calculator test, so this moves it there and bumps its repo-root lookup by the two extra directory levels. No test changes.
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
Devin e2e QATested at: 831a2a1 Both legs ran the same script against a two-worker proxy on port 37129 (checked free with a socket bind first) with no database, once at the merge base 9d0c9b9 and once at the PR tip 831a2a1. Every model points at one Azure AI Foundry resource (
model_list:
- model_name: gpt-chat-latest
litellm_params:
model: azure_ai/gpt-chat-latest
api_base: os.environ/AZURE_AI_API_BASE
api_key: os.environ/AZURE_AI_API_KEY
- model_name: cohere-command-a
litellm_params:
model: azure_ai/cohere-command-a
api_base: os.environ/AZURE_AI_API_BASE
api_key: os.environ/AZURE_AI_API_KEY
- model_name: model-router
litellm_params:
model: azure_ai/model-router
api_base: os.environ/AZURE_AI_API_BASE
api_key: os.environ/AZURE_AI_API_KEY
- model_name: model_router
litellm_params:
model: azure_ai/model_router
api_base: os.environ/AZURE_AI_API_BASE
api_key: os.environ/AZURE_AI_API_KEY
- model_name: grok-4-20-reasoning
litellm_params:
model: azure_ai/grok-4-20-reasoning
api_base: os.environ/AZURE_AI_API_BASE
api_key: os.environ/AZURE_AI_API_KEY
- model_name: grok-4.3
litellm_params:
model: azure_ai/grok-4.3
api_base: os.environ/AZURE_AI_API_BASE
api_key: os.environ/AZURE_AI_API_KEY
general_settings:
master_key: sk-1234
litellm_settings:
drop_params: trueProxy start, once per leg from a worktree at that commit (PYTHONPATH pointed at the worktree so the proxy imports that revision's price map): Shell helper used by every POST case: it prints the status line, every BASE=http://127.0.0.1:37129; KEY=sk-1234
call() {
local route="$1" body="$2"
local H="$(mktemp)" B="$(mktemp)"
curl -s -D "$H" -o "$B" "${BASE}${route}" -H "Authorization: Bearer ${KEY}" -H 'Content-Type: application/json' -d "$body"
grep -iE '^HTTP/|^x-litellm-response-cost' "$H" | tr -d '\r'
python3 - "$B" <<'PY'
import json, sys
obj = json.load(open(sys.argv[1]))
def text(o):
for path in (("choices", 0, "message", "content"), ("content", 0, "text"), ("output", -1, "content", 0, "text")):
cur = o
try:
for p in path:
cur = cur[p]
return cur
except Exception:
continue
usage = obj.get("usage")
if isinstance(usage, dict):
usage = {k: v for k, v in usage.items() if not k.endswith("_details")}
keep = {"model": obj.get("model"), "text": text(obj), "usage": usage}
if "error" in obj:
keep["error"] = obj["error"]
print(json.dumps(keep)[:700])
PY
rm -f "$H" "$B"
}Before (9d0c9b9)GET /model/info
POST /v1/chat/completions gpt-chat-latest
POST /v1/chat/completions cohere-command-a
POST /v1/chat/completions grok-4-20-reasoning
POST /v1/chat/completions model-router
POST /v1/chat/completions model_router
POST /v1/chat/completions grok-4.3 (control)
POST /v1/responses gpt-chat-latest
POST /v1/responses grok-4-20-reasoning
POST /v1/responses model-router
POST /v1/messages grok-4-20-reasoning
POST /v1/messages cohere-command-a
POST /v1/messages gpt-chat-latest (max_completion_tokens)
POST /v1/messages model-router
After (831a2a1)GET /model/info
POST /v1/chat/completions gpt-chat-latest
POST /v1/chat/completions cohere-command-a
POST /v1/chat/completions grok-4-20-reasoning
POST /v1/chat/completions model-router
POST /v1/chat/completions model_router
POST /v1/chat/completions grok-4.3 (control)
POST /v1/responses gpt-chat-latest
POST /v1/responses grok-4-20-reasoning
POST /v1/responses model-router
POST /v1/messages grok-4-20-reasoning
POST /v1/messages cohere-command-a
POST /v1/messages gpt-chat-latest (max_completion_tokens)
POST /v1/messages model-router
Observations from the run, none caused by this PR:
Verdict: PASS |
…imits The router fee was folded into azure_ai.cost_per_token and then added again by the additional_costs hook, so every routed request paid it twice. The hook now owns the fee, the entry named by the deployment supplies the price, and a response priced as the router entry itself is not charged again model-router, gpt-chat-latest and cohere-command-a carry the limits from the Foundry models page, and model-router and grok-4-20-* carry their retirement dates. The router tests now run at the completion_cost level with a Logging object, which is the path the proxy takes, and fail at the merge base
Restores the public keyword removed at 415bdbf. A direct caller that names the Model Router as the request model gets the routing fee folded into the prompt cost once; completion_cost never passes it and charges the fee through the additional-costs hook as before
…ce grok-4-20 cache reads Direct litellm.cost_per_token callers that name a Model Router deployment as the model get the routing fee again, as they did before this branch, and the fee is still charged exactly once on every completion_cost path. The grok-4-20 entries bill cached prompt tokens at the input rate, since Azure has no cached-input meter for them, and the model_router twin carries the same limits and retirement date as model-router. The catalog test now exercises the cost calculator and map relations instead of pinning map fields.
…lculator litellm's azure_ai config rejects reasoning_effort for gpt-chat-latest and Azure documents a fixed reasoning level for it, so the entry no longer advertises reasoning_effort_levels. The catalog metadata tests compare cost_per_token and the whisper transcription cost with the entry the calculator read instead of with list-price literals, the pattern #40195 removed
|
|
||
| # Get additional costs from provider (e.g., routing fees, infrastructure costs) | ||
| if custom_llm_provider == "azure_ai": | ||
| if custom_llm_provider == "azure_ai" and not azure_ai_is_model_router_name(model): |
There was a problem hiding this comment.
Provider Logic Escapes Boundary
The generic cost-calculation path now imports an Azure AI router predicate and branches explicitly on custom_llm_provider == "azure_ai". This violates the repository directive to keep provider-specific behavior under llms/, so the provider boundary must be restored before merging.
Rule Used: What: Avoid writing provider-specific code outside... (source)
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
There was a problem hiding this comment.
The azure_ai branch and provider import are the merge base's own hook. This tip only narrows it so the fee is charged once
There was a problem hiding this comment.
You're right. I rechecked the merge-base diff: the existing azure_ai cost-calculator import and custom_llm_provider == "azure_ai" hook are unchanged; this tip only adds the provider-owned router predicate and uses it to skip the existing additional-cost hook when the router fee is already included. My comment treated pre-existing provider integration as introduced by this PR, so it is not actionable for this diff. I withdraw the finding.
| entry: Final = get_model_info(f"azure_ai/{catalog_name}") | ||
| prompt_cost, completion_cost_usd = cost_per_token( | ||
| model=f"azure_ai/{catalog_name}", prompt_tokens=A_MILLION, completion_tokens=A_MILLION | ||
| ) | ||
| assert prompt_cost > 0 | ||
| assert prompt_cost == pytest.approx(A_MILLION * entry["input_cost_per_token"]) | ||
| assert completion_cost_usd == pytest.approx(A_MILLION * entry["output_cost_per_token"]) |
There was a problem hiding this comment.
Pricing Tests Lose Independence
These modified tests derive expected token costs through get_model_info, which reads the same cost map used by cost_per_token; an incorrect map price therefore passes both sides of the assertion. The Whisper test also replaces its fixed $0.36/hour expectation with relative scaling, while the router suite removes the cached-token billing case. This weakens existing regression coverage and violates the repository directive to prevent test modifications from masking regressions, so independent expected prices and equivalent cached-token coverage must be restored before merging.
Rule Used: What: Flag any modifications to existing tests and... (source)
There was a problem hiding this comment.
You're right on both points. Given #40195 removed repo-wide literal price assertions, deriving the expected values from the registered entry is appropriate here: the test verifies that each catalog name is present, has a nonzero input price, and that cost_per_token actually uses that entry. The source/backup comparison separately guards map consistency. And with the cached-token fee case restored in 3cadf2f8f7, my previous coverage objection is resolved. No further issue from this comment.
1 similar comment
|
bugbot run |
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 831a2a1. Configure here.
TLDR
Problem this solves:
azure_ai/priceazure_ai/whisperwas priced with theazurelookup, where the name has no entry, so it logged $0 even with anazure_ai/whisperprice on the mapHow it solves it:
gpt-chat-latestadvertisingreasoning_effortlevels theazure_aiprovider cannot sendazure_ai/whispertranscription against theazure_ai/entry it was configured with, whileazure/whisper-1keeps itsazurepriceUser Flow
Before: a proxy admin adds seven of the customer's Foundry deployments next to one that already prices, and every call to the new names comes back with no cost
azure_ai/gpt-chat-latest,azure_ai/cohere-command-a,azure_ai/model-router,azure_ai/grok-4-20-reasoning, andazure_ai/grok-4-20-non-reasoningto the config next toazure_ai/grok-4.6, all on the same Foundry endpoint, plusazure_ai/whisperon a second Foundry resource that has awhisperdeployment, and starts the proxyinput_cost_per_token: 0,output_cost_per_token: 0, nomax_input_tokens, and nomode(whisperalso with no per-second price), whilegrok-4.6shows $2/M in, $6/M out, 200000 max input, mode chat"model": "gpt-chat-latest"and gets a 200 with the answer and 12 prompt and 9 completion tokens, but nox-litellm-response-costheader, onlyx-litellm-response-cost-input: 0.0andx-litellm-response-cost-output: 0.0"model": "cohere-command-a", with"model": "grok-4-20-reasoning", and with"model": "grok-4-20-non-reasoning"comes back the same way: 200, tokens counted, no cost header"model": "gpt-chat-latest"returns a 200 with"usage": {"input_tokens": 12, "output_tokens": 9, "cost": null}and no cost header"model": "grok-4-20-reasoning"returns a 200 with no cost header"model": "model-router"returns a 200 answered byazure_ai/grok-4-1-fast-reasoningwithx-litellm-response-cost: 0.00012968: the routed model's 0.000128 of reasoning output plus the 6-token router fee of 8.4e-07 twice, once inx-litellm-response-cost-inputand once more that only shows in the total. The samemodel-routerrequest on POST http://localhost:4000/v1/responses returnsx-litellm-response-cost: 0.00020268and on POST http://localhost:4000/v1/messagesx-litellm-response-cost: 4.658e-05, the fee twice eachmodel-routerrequest on POST http://localhost:4000/v1/chat/completions with"stream": trueand"stream_options": {"include_usage": true}streams an answer fromgrok-4-1-fast-reasoningwhose last chunk carries"usage": {..., "cost": 6.418e-05}: 125 reasoning tokens at the routed model's $0.50/M (6.25e-05) plus the 6-token router fee of 8.4e-07 twicemodel=whisperand a 17.6 second wav returns a 200 with the transcript and nox-litellm-response-costheader, while the same file sent asmodel=whisper-1(anazure/whisper-1deployment on the same resource) returnsx-litellm-response-cost: 0.0017577052154195013After: the same config prices every new name, every endpoint returns a real cost, and the router fee lands once
azure_ai/gpt-chat-latest,azure_ai/cohere-command-a,azure_ai/model-router,azure_ai/grok-4-20-reasoning, andazure_ai/grok-4-20-non-reasoningto the config next toazure_ai/grok-4.6, all on the same Foundry endpoint, plusazure_ai/whisperon a second Foundry resource that has awhisperdeployment, and starts the proxygpt-chat-latestat $5/M in, $30/M out, 272000 max input, mode chat;cohere-command-aat $2.50/M in, $10/M out, 131072 max input;model-routerat $0.14/M in, 200000 max input;grok-4-20-reasoningandgrok-4-20-non-reasoningat $1.25/M in, $2.50/M out, 262000 max input;whisperat $0.0001 per second, mode audio_transcription;grok-4.6unchanged"model": "gpt-chat-latest"and gets a 200 withx-litellm-response-cost: 0.00033for the same 12 prompt and 9 completion tokens"model": "cohere-command-a"returnsx-litellm-response-cost: 6.5e-05, with"model": "grok-4-20-reasoning"returnsx-litellm-response-cost: 0.00146, reasoning tokens included, and with"model": "grok-4-20-non-reasoning"returnsx-litellm-response-cost: 2.5e-05for 10 prompt and 5 completion tokens"model": "gpt-chat-latest"returns a 200 withx-litellm-response-cost: 0.00033"model": "grok-4-20-reasoning"returns a 200 withx-litellm-response-cost: 2.25e-05"model": "model-router"returns a 200 answered byazure_ai/grok-4-1-fast-reasoningwithx-litellm-response-cost: 0.00014434: the routed model's 0.0001435 of reasoning output plus the 6-token router fee of 8.4e-07 once, andx-litellm-response-cost-input: 0.0because the routed model's prompt tokens came back as cache reads it prices at zero. The samemodel-routerrequest on POST http://localhost:4000/v1/responses returnsx-litellm-response-cost: 0.00011534and on POST http://localhost:4000/v1/messagesx-litellm-response-cost: 2.504e-05, the fee once eachmodel-routerrequest on POST http://localhost:4000/v1/chat/completions with"stream": trueand"stream_options": {"include_usage": true}streams an answer fromgrok-4-1-fast-reasoningwhose last chunk carries"usage": {..., "cost": 0.00031134}: 621 reasoning tokens at the routed model's $0.50/M (0.0003105) plus the 6-token router fee of 8.4e-07 oncemodel=whisperand the same 17.6 second wav returns a 200 with the transcript andx-litellm-response-cost: 0.0017577052154195013(17.577 seconds at $0.0001 per second), the same numbermodel=whisper-1returned before and still returnsRelevant issues
FLUX.2-flexfrom the same customer list is priced by feat(azure_ai): support FLUX.2 flex images #39424 (per-megapixel image pricing), so it is left out heremax_tokensvsmax_completion_tokens400 thatgpt-chat-latestreturns on /v1/messages in the QA belowLinear ticket
Resolves LIT-3157
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
The proof ran on two rigs, each a two-worker proxy with no database booted once at the merge base and once at the PR tip from the same config. The first rig (port 27233,
lit3157.yaml) drivesgpt-chat-latest,cohere-command-a,model-router,grok-4-20-reasoning, and thegrok-4.6control, all deployed under those names on one Azure AI Foundry resource. The second rig (port 29792,lit3157-r4.yaml) drivesgrok-4-20-non-reasoningand a streamedmodel-routercall on that same resource, pluswhisper(azure_ai/whisper) and thewhisper-1control (azure/whisper-1) on a second Foundry resource where both are deployed.grok-4.6andwhisper-1are the controls: both were already priced before this PR. Theazure_ai/model-routerdeployment answers with whichever model Azure picks, so its token counts differ between legs. Themodel-routercalls on /v1/responses and /v1/messages joined the script after the first Before run, so their Before blocks come from a second run of just those two calls at the merge base with the same helperlit3157.yaml:lit3157-r4.yaml:Proxy start, once per leg after
git checkout <hash>(lit3157-r4.yamlon port 29792 for the second rig):Shell helper used by every POST case: it prints the status line, every
x-litellm-response-cost*header, and the body trimmed tomodel, the first text, andusagewith its*_detailssub-objects folded intocached_tokensandreasoning_tokensBefore (9d0c9b9)
GET /model/info
POST /v1/chat/completions gpt-chat-latest
x-litellm-response-costheaderPOST /v1/chat/completions cohere-command-a
POST /v1/chat/completions grok-4-20-reasoning
total_tokens, no cost headerPOST /v1/chat/completions model-router
azure_ai/grok-4-1-fast-reasoning, total 0.00012968 = 0.000128 (256 reasoning tokens at the routed model's $0.50/M) plus the 6-token routing fee of 8.4e-07 twice, once in the input header and once more that only shows in the totalPOST /v1/chat/completions grok-4.6 (control)
POST /v1/responses gpt-chat-latest
usage.costis null, no cost headerPOST /v1/messages grok-4-20-reasoning
POST /v1/messages gpt-chat-latest
max_completion_tokensfor this model (see fix(proxy): use max_completion_tokens for gpt-chat-latest health #30987)POST /v1/responses model-router
azure_ai/grok-4-1-fast-reasoning, total 0.00020268 = 402 reasoning tokens x $0.50/M (413 total minus 6 in and 5 out) plus the 6-token routing fee of 8.4e-07 twice, once in the input header and once more that only shows in the totalPOST /v1/messages model-router
azure_ai/grok-4-1-fast-reasoning, total 4.658e-05 = 3 output tokens x $0.50/M plus the routing fee over the 161 cache-read tokens this route reports as the prompt twice, 2.254e-05 in the input header and 2.254e-05 more that only shows in the totalGET /model/info (second rig)
POST /v1/chat/completions grok-4-20-non-reasoning
POST /v1/chat/completions model-router, stream=true
grok-4-1-fast-reasoning, and the last chunk's usage carriescost: 6.418e-05= 6.25e-05 (125 reasoning tokens at the routed model's $0.50/M) plus the 6-token routing fee of 8.4e-07 twicePOST /v1/audio/transcriptions whisper
POST /v1/audio/transcriptions whisper-1 (control)
azure/whisper-1deployment on the same resource already bills 17.577 seconds at $0.0001 per secondAfter (3cadf2f)
The one commit since, 831a2a1, changes only how a transcription's provider is recorded for pricing, which none of the first-rig calls below touch, so their proof stands at 3cadf2f; the second rig further down ran at 831a2a1
GET /model/info
POST /v1/chat/completions gpt-chat-latest
x-litellm-response-cost: 0.00033(12 x $5/M + 9 x $30/M)POST /v1/chat/completions cohere-command-a
x-litellm-response-cost: 6.5e-05(6 x $2.50/M + 5 x $10/M)POST /v1/chat/completions grok-4-20-reasoning
x-litellm-response-cost: 0.00146(10 x $1.25/M + 579 reasoning tokens x $2.50/M; the 4 visible tokens go unbilled on this route, LIT-7189)POST /v1/chat/completions model-router
azure_ai/grok-4-1-fast-reasoning, total 0.00014434 = 287 reasoning tokens x $0.50/M plus the 6-token routing fee of 8.4e-07 charged once, as theAzure Model Router Flat Costline; input 0.0 because the routed model reports 161 cache-read tokens against 6 prompt tokens and its entry has no cache-read pricePOST /v1/chat/completions grok-4.6 (control)
POST /v1/responses gpt-chat-latest
x-litellm-response-cost: 0.00033, the same price as the chat completions call for the same tokensPOST /v1/responses model-router
azure_ai/grok-4-1-fast-reasoning, total 0.00011534 = 229 reasoning tokens x $0.50/M (238 total minus 6 in and 3 out) plus the 6-token routing fee of 8.4e-07 once, where the Before block above pays it twicePOST /v1/messages grok-4-20-reasoning
x-litellm-response-cost: 2.25e-05(10 x $1.25/M + 4 x $2.50/M); the reasoning tokens behind those 4 visible tokens are missing from this route's usage block, so they go unbilled here (LIT-7189)POST /v1/messages model-router
azure_ai/grok-4-1-fast-reasoning, total 2.504e-05 = 5 output tokens x $0.50/M plus the routing fee once over the 161 tokens this route reports as the prompt (0 input plus 161 cache reads), where the Before block above pays it twicePOST /v1/messages gpt-chat-latest
After (831a2a1), second rig
GET /model/info (second rig)
whisperper second with the transcription modePOST /v1/chat/completions grok-4-20-non-reasoning
x-litellm-response-cost: 2.5e-05= 10 prompt tokens at $1.25/M plus 5 completion tokens at $2.50/MPOST /v1/chat/completions model-router, stream=true
grok-4-1-fast-reasoning, and the last chunk's usage carriescost: 0.00031134= 0.0003105 (621 reasoning tokens at the routed model's $0.50/M) plus the 6-token routing fee of 8.4e-07 oncePOST /v1/audio/transcriptions whisper
x-litellm-response-cost: 0.0017577052154195013, 17.577 seconds at $0.0001 per secondPOST /v1/audio/transcriptions whisper-1 (control)
Observations from the run, none caused or worsened by this PR:
gpt-chat-latestrejectsmax_tokenson /v1/messages upstream, both legs: left alone (fix(proxy): use max_completion_tokens for gpt-chat-latest health #30987)model-routerprompt priced at zero, both legs: cached tokens outnumber prompt tokensVerdict: PASS
No embedding or rerank name made the cut, so there is no embedding or rerank leg.
codex-minihas no live leg: Azure refuses newcodex-minideployments (ServiceModelDeprecating) and none of the QA subscriptions has one, so it is covered by the metadata test onlyThe customer's list had 33 names with no
azure_ai/pricing. Seven are live Foundry catalog entries and ship here (gpt-chat-latest,codex-mini,whisper,model-router,cohere-command-a,grok-4-20-reasoning,grok-4-20-non-reasoning). The other 26 are left out on purpose:FLUX.2-flex: covered by the open PR feat(azure_ai): support FLUX.2 flex images #39424, which adds the per-megapixel image pricing this PR does not modelada,babbage,curie,davinci: retired on 2024-06-14AI21-Jamba-Instruct: retired on 2025-03-01AI21-Jamba-1.5-Large,AI21-Jamba-1.5-Mini: retired on 2025-08-01Cohere-command-r,Cohere-command-r-plus: retired on 2025-06-30Cohere-command-r-08-2024,Cohere-command-r-plus-08-2024: retired on 2026-05-12Meta-Llama-3-8B-Instruct: retired on 2025-06-30Mistral-Large-2411: retired on 2026-01-30Kimi-K2-Thinking: retired on 2026-03-29DeepSeek-R1-0528: retired on 2026-07-13gpt-5.1-chat,gpt-5.2-chat,gpt-5.3-chat: retired on 2026-06-29 and 2026-05-13, replaced bygpt-chat-latestgpt-35-turbo-16k(0613): retired on 2025-04-30gpt-4-32k: retired on 2025-06-06gpt-35-turbo,o1-mini,sora: absent from the Foundry Models API in every one of the 12 regions checked, shown as retired or archived in the cataloggpt-oss-20b: no per-token Azure retail meter exists for it (onlygpt-oss-120bis metered), so there is no list price to writeqwen3-32b: Azure lists only fine-tuned (FT) meters for it ($0.30/M in, $1.20/M out) and no base-model meter, so there is no list price for the catalog nameType
🐛 Bug Fix
Caveats (if any)
Medium
model-routerresponse the fee shows only as theAzure Model Router Flat Costbreakdown line and in the total, not inx-litellm-response-cost-inputor the OTel input cost attributegpt-chat-lateston /v1/messages returns Azure'smax_tokens400 on both legsLow
LITELLM_LOCAL_MODEL_COST_MAP=Truekeeps an older proxy on its shipped map until it upgradescodex-mini: metadata test only, no live legcodex-minideployments (ServiceModelDeprecating) and no QA subscription has an existing onegrok-4-20-*cached prompt tokens bill at the input price: Azure lists no cached meter for Grok 4.2model-routerprompt prices at zero when the routed model reports more cached tokens than prompt tokens, both legsmodel-router/model_router, with a second copy in the proxy request path, so a deployment alias containing that substring gets the flat feecohere-command-amax_output_tokensis 8182, copied from Microsoft Learn as printedmodel-routerlimits are the router deployment's own (200000 in, 32768 out); the routed model's limits apply per callFinal Attestation
Link to Devin session: https://app.devin.ai/sessions/c438be9d29cc43d7a76e43a46f560c4a
Open in Devin Desktop: https://app.devin.ai/desktop/session/c438be9d29cc43d7a76e43a46f560c4a?variant=devin
Requested by: @mateo-berri
Note
Medium Risk
Changes core spend calculation for Azure AI (catalog pricing and router fee logic); regressions would show up as incorrect totals or headers, though behavior is heavily test-covered.
Overview
Adds seven Azure AI Foundry catalog models to the cost map (
gpt-chat-latest,codex-mini,whisper,model-router,cohere-command-a,grok-4-20-reasoning,grok-4-20-non-reasoning) so proxy/model info and response cost headers stop reporting $0 for those names.Refactors Azure AI Model Router billing so the $0.14/M input routing fee is applied once: response tokens are priced on the actual model, the fee is folded into prompt cost when the priced name is router-shaped, and
completion_costskips the separate “additional costs” path when the response model name is already a router (avoiding the previous double charge). Router fee lookup also supports bothmodel_routerandmodel-routerspellings.Azure audio transcriptions now propagate
custom_llm_providerinto hidden params soazure_ai/whisperis billed against the new per-second entry instead of defaulting toazure.Tests cover catalog metadata, router fee edge cases, cost breakdown lines, and transcription provider tagging.
Reviewed by Cursor Bugbot for commit 831a2a1. Bugbot is set up for automated code reviews on this repo. Configure here.