Skip to content

fix(azure_ai): price seven Foundry catalog names and charge the model router fee once - #40189

Merged
mateo-berri merged 8 commits into
litellm_internal_stagingfrom
litellm_lit_3157_azure_ai_catalog_models
Sep 9, 2026
Merged

mateo-berri merged 8 commits into
litellm_internal_stagingfrom
litellm_lit_3157_azure_ai_catalog_models

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 8, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Seven live Azure AI Foundry catalog names have no azure_ai/ price
  • Calls to them return no cost header and log $0 spend
  • Every request through the model router paid the $0.14/M routing fee twice
  • A transcription through azure_ai/whisper was priced with the azure lookup, where the name has no entry, so it logged $0 even with an azure_ai/whisper price on the map

How it solves it:

  • Adds the seven entries with today's Azure list prices, context limits, and retirement dates
  • Charges the routing fee once per request, whichever spelling of the router name the cost calculator sees
  • Adds tests that fail without the entries or with the fee charged twice
  • Stops gpt-chat-latest advertising reasoning_effort levels the azure_ai provider cannot send
  • Prices an azure_ai/whisper transcription against the azure_ai/ entry it was configured with, while azure/whisper-1 keeps its azure price

User Flow

Before: a proxy admin adds seven of the customer's Foundry deployments next to one that already prices, and every call to the new names comes back with no cost

  1. The admin adds azure_ai/gpt-chat-latest, azure_ai/cohere-command-a, azure_ai/model-router, azure_ai/grok-4-20-reasoning, and azure_ai/grok-4-20-non-reasoning to the config next to azure_ai/grok-4.6, all on the same Foundry endpoint, plus azure_ai/whisper on a second Foundry resource that has a whisper deployment, and starts the proxy
  2. They send GET http://localhost:4000/model/info and the six new rows come back with input_cost_per_token: 0, output_cost_per_token: 0, no max_input_tokens, and no mode (whisper also with no per-second price), while grok-4.6 shows $2/M in, $6/M out, 200000 max input, mode chat
  3. A developer sends POST http://localhost:4000/v1/chat/completions with "model": "gpt-chat-latest" and gets a 200 with the answer and 12 prompt and 9 completion tokens, but no x-litellm-response-cost header, only x-litellm-response-cost-input: 0.0 and x-litellm-response-cost-output: 0.0
  4. The same POST with "model": "cohere-command-a", with "model": "grok-4-20-reasoning", and with "model": "grok-4-20-non-reasoning" comes back the same way: 200, tokens counted, no cost header
  5. POST http://localhost:4000/v1/responses with "model": "gpt-chat-latest" returns a 200 with "usage": {"input_tokens": 12, "output_tokens": 9, "cost": null} and no cost header
  6. POST http://localhost:4000/v1/messages with "model": "grok-4-20-reasoning" returns a 200 with no cost header
  7. POST http://localhost:4000/v1/chat/completions with "model": "model-router" returns a 200 answered by azure_ai/grok-4-1-fast-reasoning with x-litellm-response-cost: 0.00012968: the routed model's 0.000128 of reasoning output plus the 6-token router fee of 8.4e-07 twice, once in x-litellm-response-cost-input and once more that only shows in the total. The same model-router request on POST http://localhost:4000/v1/responses returns x-litellm-response-cost: 0.00020268 and on POST http://localhost:4000/v1/messages x-litellm-response-cost: 4.658e-05, the fee twice each
  8. The same model-router request on POST http://localhost:4000/v1/chat/completions with "stream": true and "stream_options": {"include_usage": true} streams an answer from grok-4-1-fast-reasoning whose last chunk carries "usage": {..., "cost": 6.418e-05}: 125 reasoning tokens at the routed model's $0.50/M (6.25e-05) plus the 6-token router fee of 8.4e-07 twice
  9. POST http://localhost:4000/v1/audio/transcriptions with model=whisper and a 17.6 second wav returns a 200 with the transcript and no x-litellm-response-cost header, while the same file sent as model=whisper-1 (an azure/whisper-1 deployment on the same resource) returns x-litellm-response-cost: 0.0017577052154195013

After: the same config prices every new name, every endpoint returns a real cost, and the router fee lands once

  1. The admin adds azure_ai/gpt-chat-latest, azure_ai/cohere-command-a, azure_ai/model-router, azure_ai/grok-4-20-reasoning, and azure_ai/grok-4-20-non-reasoning to the config next to azure_ai/grok-4.6, all on the same Foundry endpoint, plus azure_ai/whisper on a second Foundry resource that has a whisper deployment, and starts the proxy
  2. They send GET http://localhost:4000/model/info and see gpt-chat-latest at $5/M in, $30/M out, 272000 max input, mode chat; cohere-command-a at $2.50/M in, $10/M out, 131072 max input; model-router at $0.14/M in, 200000 max input; grok-4-20-reasoning and grok-4-20-non-reasoning at $1.25/M in, $2.50/M out, 262000 max input; whisper at $0.0001 per second, mode audio_transcription; grok-4.6 unchanged
  3. The developer sends POST http://localhost:4000/v1/chat/completions with "model": "gpt-chat-latest" and gets a 200 with x-litellm-response-cost: 0.00033 for the same 12 prompt and 9 completion tokens
  4. The same POST with "model": "cohere-command-a" returns x-litellm-response-cost: 6.5e-05, with "model": "grok-4-20-reasoning" returns x-litellm-response-cost: 0.00146, reasoning tokens included, and with "model": "grok-4-20-non-reasoning" returns x-litellm-response-cost: 2.5e-05 for 10 prompt and 5 completion tokens
  5. POST http://localhost:4000/v1/responses with "model": "gpt-chat-latest" returns a 200 with x-litellm-response-cost: 0.00033
  6. POST http://localhost:4000/v1/messages with "model": "grok-4-20-reasoning" returns a 200 with x-litellm-response-cost: 2.25e-05
  7. POST http://localhost:4000/v1/chat/completions with "model": "model-router" returns a 200 answered by azure_ai/grok-4-1-fast-reasoning with x-litellm-response-cost: 0.00014434: the routed model's 0.0001435 of reasoning output plus the 6-token router fee of 8.4e-07 once, and x-litellm-response-cost-input: 0.0 because the routed model's prompt tokens came back as cache reads it prices at zero. The same model-router request on POST http://localhost:4000/v1/responses returns x-litellm-response-cost: 0.00011534 and on POST http://localhost:4000/v1/messages x-litellm-response-cost: 2.504e-05, the fee once each
  8. The same model-router request on POST http://localhost:4000/v1/chat/completions with "stream": true and "stream_options": {"include_usage": true} streams an answer from grok-4-1-fast-reasoning whose last chunk carries "usage": {..., "cost": 0.00031134}: 621 reasoning tokens at the routed model's $0.50/M (0.0003105) plus the 6-token router fee of 8.4e-07 once
  9. POST http://localhost:4000/v1/audio/transcriptions with model=whisper and the same 17.6 second wav returns a 200 with the transcript and x-litellm-response-cost: 0.0017577052154195013 (17.577 seconds at $0.0001 per second), the same number model=whisper-1 returned before and still returns

Relevant issues

Linear ticket

Resolves LIT-3157

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

The proof ran on two rigs, each a two-worker proxy with no database booted once at the merge base and once at the PR tip from the same config. The first rig (port 27233, lit3157.yaml) drives gpt-chat-latest, cohere-command-a, model-router, grok-4-20-reasoning, and the grok-4.6 control, all deployed under those names on one Azure AI Foundry resource. The second rig (port 29792, lit3157-r4.yaml) drives grok-4-20-non-reasoning and a streamed model-router call on that same resource, plus whisper (azure_ai/whisper) and the whisper-1 control (azure/whisper-1) on a second Foundry resource where both are deployed. grok-4.6 and whisper-1 are the controls: both were already priced before this PR. The azure_ai/model-router deployment answers with whichever model Azure picks, so its token counts differ between legs. The model-router calls on /v1/responses and /v1/messages joined the script after the first Before run, so their Before blocks come from a second run of just those two calls at the merge base with the same helper

lit3157.yaml:

model_list:
  - model_name: gpt-chat-latest
    litellm_params:
      model: azure_ai/gpt-chat-latest
      api_base: os.environ/AZURE_AI_API_BASE
      api_key: os.environ/AZURE_AI_API_KEY
  - model_name: cohere-command-a
    litellm_params:
      model: azure_ai/cohere-command-a
      api_base: os.environ/AZURE_AI_API_BASE
      api_key: os.environ/AZURE_AI_API_KEY
  - model_name: model-router
    litellm_params:
      model: azure_ai/model-router
      api_base: os.environ/AZURE_AI_API_BASE
      api_key: os.environ/AZURE_AI_API_KEY
  - model_name: grok-4-20-reasoning
    litellm_params:
      model: azure_ai/grok-4-20-reasoning
      api_base: os.environ/AZURE_AI_API_BASE
      api_key: os.environ/AZURE_AI_API_KEY
  - model_name: grok-4.6
    litellm_params:
      model: azure_ai/grok-4.6
      api_base: os.environ/AZURE_AI_API_BASE
      api_key: os.environ/AZURE_AI_API_KEY
general_settings:
  master_key: sk-1234
litellm_settings:
  drop_params: true

lit3157-r4.yaml:

model_list:
  - model_name: grok-4-20-non-reasoning
    litellm_params:
      model: azure_ai/grok-4-20-non-reasoning
      api_base: os.environ/AZURE_AI_API_BASE
      api_key: os.environ/AZURE_AI_API_KEY
  - model_name: model-router
    litellm_params:
      model: azure_ai/model-router
      api_base: os.environ/AZURE_AI_API_BASE
      api_key: os.environ/AZURE_AI_API_KEY
  - model_name: whisper
    litellm_params:
      model: azure_ai/whisper
      api_base: os.environ/AZURE_AUDIO_API_BASE
      api_key: os.environ/AZURE_AUDIO_API_KEY
      api_version: "2024-06-01"
  - model_name: whisper-1
    litellm_params:
      model: azure/whisper-1
      api_base: os.environ/AZURE_AUDIO_API_BASE
      api_key: os.environ/AZURE_AUDIO_API_KEY
      api_version: "2024-06-01"
general_settings:
  master_key: sk-1234
litellm_settings:
  drop_params: true

Proxy start, once per leg after git checkout <hash> (lit3157-r4.yaml on port 29792 for the second rig):

LITELLM_LOCAL_MODEL_COST_MAP=True python litellm/proxy/proxy_cli.py --config lit3157.yaml --port 27233 --num_workers 2 --detailed_debug

Shell helper used by every POST case: it prints the status line, every x-litellm-response-cost* header, and the body trimmed to model, the first text, and usage with its *_details sub-objects folded into cached_tokens and reasoning_tokens

BASE=http://127.0.0.1:27233; KEY=sk-1234
call() {
  local route="$1" body="$2"
  local H="$(mktemp)" B="$(mktemp)"
  curl -s -D "$H" -o "$B" "${BASE}${route}" -H "Authorization: Bearer ${KEY}" -H 'Content-Type: application/json' -d "$body"
  grep -iE '^HTTP/|^x-litellm-response-cost' "$H" | tr -d '\r'
  python3 - "$B" <<'PY'
import json, sys
obj = json.load(open(sys.argv[1]))
def text(o):
    for path in (("choices", 0, "message", "content"), ("content", 0, "text"), ("output", -1, "content", 0, "text")):
        cur = o
        try:
            for p in path:
                cur = cur[p]
            return cur
        except Exception:
            continue
usage = obj.get("usage")
if isinstance(usage, dict):
    prompt_details = usage.get("prompt_tokens_details") or {}
    completion_details = usage.get("completion_tokens_details") or {}
    usage = {k: v for k, v in usage.items() if not k.endswith("_details")}
    for label, details, key in (("cached_tokens", prompt_details, "cached_tokens"), ("reasoning_tokens", completion_details, "reasoning_tokens")):
        if isinstance(details, dict) and details.get(key) is not None:
            usage[label] = details[key]
keep = {"model": obj.get("model"), "text": text(obj), "usage": usage}
if "error" in obj:
    keep["error"] = obj["error"]
print(json.dumps(keep)[:700])
PY
  rm -f "$H" "$B"
}

Before (9d0c9b9)

GET /model/info

  1. Command:
    curl -s $BASE/model/info -H "Authorization: Bearer $KEY" | python3 -c '
    import json, sys
    for m in json.load(sys.stdin)["data"]:
        mi = m["model_info"]
        print(m["model_name"], "|", m["litellm_params"]["model"], "| in", mi.get("input_cost_per_token"), "| out", mi.get("output_cost_per_token"), "| max_in", mi.get("max_input_tokens"), "| mode", mi.get("mode"))'
    
  2. Output: the four new names carry no price, no limit, and no mode
    gpt-chat-latest | azure_ai/gpt-chat-latest | in 0 | out 0 | max_in None | mode None
    cohere-command-a | azure_ai/cohere-command-a | in 0 | out 0 | max_in None | mode None
    model-router | azure_ai/model-router | in 0 | out 0 | max_in None | mode None
    grok-4-20-reasoning | azure_ai/grok-4-20-reasoning | in 0 | out 0 | max_in None | mode None
    grok-4.6 | azure_ai/grok-4.6 | in 2e-06 | out 6e-06 | max_in 200000 | mode chat
    

POST /v1/chat/completions gpt-chat-latest

  1. Command:
    call /v1/chat/completions '{"model":"gpt-chat-latest","messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200, tokens counted, no x-litellm-response-cost header
    HTTP/1.1 200 OK
    x-litellm-response-cost-original: 0.0
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 0.0
    x-litellm-response-cost-output: 0.0
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "gpt-chat-latest", "text": "Hi there, friend!", "usage": {"completion_tokens": 9, "prompt_tokens": 12, "total_tokens": 21, "latency_checkpoint": {"engine_tbt_ms": 5, "engine_ttft_ms": 127, "engine_ttlt_ms": 169, "pre_inference_ms": 159, "service_tbt_ms": 9, "service_ttft_ms": 823, "service_ttlt_ms": 880, "total_duration_ms": 745, "user_visible_ttft_ms": 664}}}
    

POST /v1/chat/completions cohere-command-a

  1. Command:
    call /v1/chat/completions '{"model":"cohere-command-a","messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200, no cost header
    HTTP/1.1 200 OK
    x-litellm-response-cost-original: 0.0
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 0.0
    x-litellm-response-cost-output: 0.0
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "cohere-command-a", "text": "Hi there, friend!", "usage": {"completion_tokens": 5, "prompt_tokens": 6, "total_tokens": 11, "audio_prompt_tokens": 0}}
    

POST /v1/chat/completions grok-4-20-reasoning

  1. Command:
    call /v1/chat/completions '{"model":"grok-4-20-reasoning","messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200, 407 reasoning tokens counted in total_tokens, no cost header
    HTTP/1.1 200 OK
    x-litellm-response-cost-original: 0.0
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 0.0
    x-litellm-response-cost-output: 0.0
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "grok-4-20-reasoning", "text": "Hello my friend.", "usage": {"completion_tokens": 4, "prompt_tokens": 10, "total_tokens": 421, "audio_prompt_tokens": 0, "num_sources_used": 0}}
    

POST /v1/chat/completions model-router

  1. Command:
    call /v1/chat/completions '{"model":"model-router","messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200 answered by azure_ai/grok-4-1-fast-reasoning, total 0.00012968 = 0.000128 (256 reasoning tokens at the routed model's $0.50/M) plus the 6-token routing fee of 8.4e-07 twice, once in the input header and once more that only shows in the total
    HTTP/1.1 200 OK
    x-litellm-response-cost: 0.00012968
    x-litellm-response-cost-original: 0.00012968
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 8.4e-07
    x-litellm-response-cost-output: 0.000128
    x-litellm-response-cost-reasoning: 0.000128
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "azure_ai/grok-4-1-fast-reasoning", "text": "Hi there, friend!", "usage": {"completion_tokens": 5, "prompt_tokens": 6, "total_tokens": 267, "audio_prompt_tokens": 0, "num_sources_used": 0}}
    

POST /v1/chat/completions grok-4.6 (control)

  1. Command:
    call /v1/chat/completions '{"model":"grok-4.6","messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200 with a real cost, since this name was already priced
    HTTP/1.1 200 OK
    x-litellm-response-cost: 0.002546
    x-litellm-response-cost-original: 0.002546
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 1.9999999999999998e-05
    x-litellm-response-cost-output: 0.002526
    x-litellm-response-cost-reasoning: 0.002526
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "grok-4.6", "text": "Hello there friend.", "usage": {"completion_tokens": 4, "prompt_tokens": 10, "total_tokens": 435, "audio_prompt_tokens": 0, "num_sources_used": 0}}
    

POST /v1/responses gpt-chat-latest

  1. Command:
    call /v1/responses '{"model":"gpt-chat-latest","input":"Say hi in three words."}'
    
  2. Output: 200, usage.cost is null, no cost header
    HTTP/1.1 200 OK
    x-litellm-response-cost-original: 0.0
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 0.0
    x-litellm-response-cost-output: 0.0
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "gpt-chat-latest", "text": "Hi there, friend!", "usage": {"input_tokens": 12, "output_tokens": 9, "total_tokens": 21, "cost": null}, "error": null}
    

POST /v1/messages grok-4-20-reasoning

  1. Command:
    call /v1/messages '{"model":"grok-4-20-reasoning","max_tokens":64,"messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200, no cost header
    HTTP/1.1 200 OK
    x-litellm-response-cost-original: 0.0
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 0.0
    x-litellm-response-cost-output: 0.0
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "grok-4-20-reasoning", "text": "Hello there friend.", "usage": {"input_tokens": 10, "output_tokens": 4}}
    

POST /v1/messages gpt-chat-latest

  1. Command:
    call /v1/messages '{"model":"gpt-chat-latest","max_tokens":64,"messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 400 from Azure, which wants max_completion_tokens for this model (see fix(proxy): use max_completion_tokens for gpt-chat-latest health #30987)
    HTTP/1.1 400 Bad Request
    x-litellm-response-cost: 0
    {"model": null, "text": null, "usage": null, "error": {"type": "invalid_request_error", "message": "litellm.BadRequestError: Azure_aiException - {\n  \"error\": {\n    \"message\": \"Unsupported parameter: 'max_tokens' is not supported with this model. Use 'max_completion_tokens' instead.\",\n    \"type\": \"invalid_request_error\",\n    \"param\": \"max_tokens\",\n    \"code\": \"unsupported_parameter\"\n  }\n}. Received Model Group=gpt-chat-latest\nAvailable Model Group Fallbacks=None"}}
    

POST /v1/responses model-router

  1. Command:
    call /v1/responses '{"model":"model-router","input":"Say hi in three words."}'
    
  2. Output: 200 answered by azure_ai/grok-4-1-fast-reasoning, total 0.00020268 = 402 reasoning tokens x $0.50/M (413 total minus 6 in and 5 out) plus the 6-token routing fee of 8.4e-07 twice, once in the input header and once more that only shows in the total
    HTTP/1.1 200 OK
    x-litellm-response-cost: 0.00020267999999999998
    x-litellm-response-cost-original: 0.00020267999999999998
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 8.4e-07
    x-litellm-response-cost-output: 0.00020099999999999998
    x-litellm-response-cost-reasoning: 0.00020099999999999998
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "azure_ai/grok-4-1-fast-reasoning", "text": "Hi there, friend!", "usage": {"input_tokens": 6, "output_tokens": 5, "total_tokens": 413, "cost": null}, "error": null}
    

POST /v1/messages model-router

  1. Command:
    call /v1/messages '{"model":"model-router","max_tokens":64,"messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200 answered by azure_ai/grok-4-1-fast-reasoning, total 4.658e-05 = 3 output tokens x $0.50/M plus the routing fee over the 161 cache-read tokens this route reports as the prompt twice, 2.254e-05 in the input header and 2.254e-05 more that only shows in the total
    HTTP/1.1 200 OK
    x-litellm-response-cost: 4.6580000000000005e-05
    x-litellm-response-cost-original: 4.6580000000000005e-05
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 2.254e-05
    x-litellm-response-cost-output: 1.5e-06
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "azure_ai/grok-4-1-fast-reasoning", "text": "Hi there!", "usage": {"input_tokens": 0, "output_tokens": 3, "cache_read_input_tokens": 161}}
    

GET /model/info (second rig)

  1. Command:
    curl -s $BASE/model/info -H "Authorization: Bearer $KEY" | python3 -c '
    import json, sys
    for m in json.load(sys.stdin)["data"]:
        mi = m["model_info"]
        print(m["model_name"], "|", m["litellm_params"]["model"], "| in", mi.get("input_cost_per_token"), "| out", mi.get("output_cost_per_token"), "| in/s", mi.get("input_cost_per_second"), "| max_in", mi.get("max_input_tokens"), "| mode", mi.get("mode"))'
    
  2. Output: the three new names carry no price, no limit, and no mode
    grok-4-20-non-reasoning | azure_ai/grok-4-20-non-reasoning | in 0 | out 0 | in/s None | max_in None | mode None
    model-router | azure_ai/model-router | in 0 | out 0 | in/s None | max_in None | mode None
    whisper | azure_ai/whisper | in 0 | out 0 | in/s None | max_in None | mode None
    

POST /v1/chat/completions grok-4-20-non-reasoning

  1. Command:
    call /v1/chat/completions '{"model":"grok-4-20-non-reasoning","messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200 with 10 prompt and 5 completion tokens, no total cost header, zero breakdown headers
    HTTP/1.1 200 OK
    x-litellm-response-cost-original: 0.0
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 0.0
    x-litellm-response-cost-output: 0.0
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "grok-4-20-non-reasoning", "text": "Hi there, friend!", "usage": {"completion_tokens": 5, "prompt_tokens": 10, "total_tokens": 15, "audio_prompt_tokens": 0, "num_sources_used": 0, "cached_tokens": 0}}
    

POST /v1/chat/completions model-router, stream=true

  1. Command:
    curl -s -N $BASE/v1/chat/completions -H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' \
      -d '{"model":"model-router","stream":true,"stream_options":{"include_usage":true},"messages":[{"role":"user","content":"Say hi in three words."}]}' \
      | grep '^data: ' | sed 's/^data: //' | grep -v '^\[DONE\]' | python3 -c '
    import json, sys
    chunks = [json.loads(l) for l in sys.stdin if l.strip()]
    text = "".join((c["choices"][0].get("delta") or {}).get("content") or "" for c in chunks if c.get("choices"))
    usage = next((c["usage"] for c in reversed(chunks) if c.get("usage")), None) or {}
    routed = next((c["model"] for c in chunks if c.get("model") and c["model"] != "model-router"), None)
    print(json.dumps({"chunks": len(chunks), "routed_model": routed, "text": text, "usage": usage}))'
    
  2. Output: streamed answer from grok-4-1-fast-reasoning, and the last chunk's usage carries cost: 6.418e-05 = 6.25e-05 (125 reasoning tokens at the routed model's $0.50/M) plus the 6-token routing fee of 8.4e-07 twice
    {"chunks": 5, "routed_model": "grok-4-1-fast-reasoning", "text": "Hi there!", "usage": {"completion_tokens": 3, "prompt_tokens": 6, "total_tokens": 9, "completion_tokens_details": {"accepted_prediction_tokens": 0, "audio_tokens": 0, "reasoning_tokens": 125, "rejected_prediction_tokens": 0}, "prompt_tokens_details": {"audio_tokens": 0, "cached_tokens": 161, "text_tokens": 6, "image_tokens": 0}, "cost": 6.418e-05}}
    

POST /v1/audio/transcriptions whisper

  1. Command:
    curl -s -D - $BASE/v1/audio/transcriptions -H "Authorization: Bearer $KEY" -F model=whisper -F file=@tests/gettysburg.wav | grep -iE '^HTTP/|^x-litellm-response-cost|^\{'
    
  2. Output: 200 with the transcript and no cost header at all
    HTTP/1.1 200 OK
    {"text":"Four score and seven years ago, our fathers brought forth on this continent a new nation, conceived in liberty and dedicated to the proposition that all men are created equal. Now we are engaged in a great civil war, testing whether that nation, or any nation so conceived and so dedicated, can long endure.","usage":null}
    

POST /v1/audio/transcriptions whisper-1 (control)

  1. Command:
    curl -s -D - $BASE/v1/audio/transcriptions -H "Authorization: Bearer $KEY" -F model=whisper-1 -F file=@tests/gettysburg.wav | grep -iE '^HTTP/|^x-litellm-response-cost|^\{'
    
  2. Output: the azure/whisper-1 deployment on the same resource already bills 17.577 seconds at $0.0001 per second
    HTTP/1.1 200 OK
    x-litellm-response-cost: 0.0017577052154195013
    {"text":"Four score and seven years ago, our fathers brought forth on this continent a new nation, conceived in liberty and dedicated to the proposition that all men are created equal. Now we are engaged in a great civil war, testing whether that nation, or any nation so conceived and so dedicated, can long endure.","usage":null}
    

After (3cadf2f)

The one commit since, 831a2a1, changes only how a transcription's provider is recorded for pricing, which none of the first-rig calls below touch, so their proof stands at 3cadf2f; the second rig further down ran at 831a2a1

GET /model/info

  1. Command:
    curl -s $BASE/model/info -H "Authorization: Bearer $KEY" | python3 -c '
    import json, sys
    for m in json.load(sys.stdin)["data"]:
        mi = m["model_info"]
        print(m["model_name"], "|", m["litellm_params"]["model"], "| in", mi.get("input_cost_per_token"), "| out", mi.get("output_cost_per_token"), "| max_in", mi.get("max_input_tokens"), "| mode", mi.get("mode"))'
    
  2. Output: every new name now carries its list price, context limit, and mode
    gpt-chat-latest | azure_ai/gpt-chat-latest | in 5e-06 | out 3e-05 | max_in 272000 | mode chat
    cohere-command-a | azure_ai/cohere-command-a | in 2.5e-06 | out 1e-05 | max_in 131072 | mode chat
    model-router | azure_ai/model-router | in 1.4e-07 | out 0 | max_in 200000 | mode chat
    grok-4-20-reasoning | azure_ai/grok-4-20-reasoning | in 1.25e-06 | out 2.5e-06 | max_in 262000 | mode chat
    grok-4.6 | azure_ai/grok-4.6 | in 2e-06 | out 6e-06 | max_in 200000 | mode chat
    

POST /v1/chat/completions gpt-chat-latest

  1. Command:
    call /v1/chat/completions '{"model":"gpt-chat-latest","messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200 with x-litellm-response-cost: 0.00033 (12 x $5/M + 9 x $30/M)
    HTTP/1.1 200 OK
    x-litellm-response-cost: 0.00033
    x-litellm-response-cost-original: 0.00033
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 6.000000000000001e-05
    x-litellm-response-cost-output: 0.00027
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "gpt-chat-latest", "text": "Hi there, friend!", "usage": {"completion_tokens": 9, "prompt_tokens": 12, "total_tokens": 21, "latency_checkpoint": {"engine_tbt_ms": 9, "engine_ttft_ms": 93, "engine_ttlt_ms": 172, "pre_inference_ms": 163, "service_tbt_ms": 10, "service_ttft_ms": 846, "service_ttlt_ms": 929, "total_duration_ms": 773, "user_visible_ttft_ms": 683}, "cached_tokens": 0, "reasoning_tokens": 0}}
    

POST /v1/chat/completions cohere-command-a

  1. Command:
    call /v1/chat/completions '{"model":"cohere-command-a","messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200 with x-litellm-response-cost: 6.5e-05 (6 x $2.50/M + 5 x $10/M)
    HTTP/1.1 200 OK
    x-litellm-response-cost: 6.500000000000001e-05
    x-litellm-response-cost-original: 6.500000000000001e-05
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 1.5000000000000002e-05
    x-litellm-response-cost-output: 5e-05
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "cohere-command-a", "text": "Hi there, friend!", "usage": {"completion_tokens": 5, "prompt_tokens": 6, "total_tokens": 11, "audio_prompt_tokens": 0, "cached_tokens": 0, "reasoning_tokens": 0}}
    

POST /v1/chat/completions grok-4-20-reasoning

  1. Command:
    call /v1/chat/completions '{"model":"grok-4-20-reasoning","messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200 with x-litellm-response-cost: 0.00146 (10 x $1.25/M + 579 reasoning tokens x $2.50/M; the 4 visible tokens go unbilled on this route, LIT-7189)
    HTTP/1.1 200 OK
    x-litellm-response-cost: 0.0014600000000000001
    x-litellm-response-cost-original: 0.0014600000000000001
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 1.25e-05
    x-litellm-response-cost-output: 0.0014475000000000002
    x-litellm-response-cost-reasoning: 0.0014475000000000002
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "grok-4-20-reasoning", "text": "Hi there friend.", "usage": {"completion_tokens": 4, "prompt_tokens": 10, "total_tokens": 593, "audio_prompt_tokens": 0, "num_sources_used": 0, "cached_tokens": 0, "reasoning_tokens": 579}}
    

POST /v1/chat/completions model-router

  1. Command:
    call /v1/chat/completions '{"model":"model-router","messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200 answered by azure_ai/grok-4-1-fast-reasoning, total 0.00014434 = 287 reasoning tokens x $0.50/M plus the 6-token routing fee of 8.4e-07 charged once, as the Azure Model Router Flat Cost line; input 0.0 because the routed model reports 161 cache-read tokens against 6 prompt tokens and its entry has no cache-read price
    HTTP/1.1 200 OK
    x-litellm-response-cost: 0.00014434
    x-litellm-response-cost-original: 0.00014434
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 0.0
    x-litellm-response-cost-output: 0.0001435
    x-litellm-response-cost-reasoning: 0.0001435
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "azure_ai/grok-4-1-fast-reasoning", "text": "Hi there!", "usage": {"completion_tokens": 3, "prompt_tokens": 6, "total_tokens": 296, "audio_prompt_tokens": 0, "num_sources_used": 0, "cached_tokens": 161, "reasoning_tokens": 287}}
    

POST /v1/chat/completions grok-4.6 (control)

  1. Command:
    call /v1/chat/completions '{"model":"grok-4.6","messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200 with a real cost, same per-token prices as before, 380 reasoning tokens this run
    HTTP/1.1 200 OK
    x-litellm-response-cost: 0.0023
    x-litellm-response-cost-original: 0.0023
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 1.9999999999999998e-05
    x-litellm-response-cost-output: 0.00228
    x-litellm-response-cost-reasoning: 0.00228
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "grok-4.6", "text": "Hello there friend.", "usage": {"completion_tokens": 4, "prompt_tokens": 10, "total_tokens": 394, "audio_prompt_tokens": 0, "num_sources_used": 0, "cached_tokens": 0, "reasoning_tokens": 380}}
    

POST /v1/responses gpt-chat-latest

  1. Command:
    call /v1/responses '{"model":"gpt-chat-latest","input":"Say hi in three words."}'
    
  2. Output: 200 with x-litellm-response-cost: 0.00033, the same price as the chat completions call for the same tokens
    HTTP/1.1 200 OK
    x-litellm-response-cost: 0.00033
    x-litellm-response-cost-original: 0.00033
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 6.000000000000001e-05
    x-litellm-response-cost-output: 0.00027
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "gpt-chat-latest", "text": "Hi there, friend!", "usage": {"input_tokens": 12, "output_tokens": 9, "total_tokens": 21, "cost": null}, "error": null}
    

POST /v1/responses model-router

  1. Command:
    call /v1/responses '{"model":"model-router","input":"Say hi in three words."}'
    
  2. Output: 200 answered by azure_ai/grok-4-1-fast-reasoning, total 0.00011534 = 229 reasoning tokens x $0.50/M (238 total minus 6 in and 3 out) plus the 6-token routing fee of 8.4e-07 once, where the Before block above pays it twice
    HTTP/1.1 200 OK
    x-litellm-response-cost: 0.00011533999999999999
    x-litellm-response-cost-original: 0.00011533999999999999
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 0.0
    x-litellm-response-cost-output: 0.00011449999999999999
    x-litellm-response-cost-reasoning: 0.00011449999999999999
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "azure_ai/grok-4-1-fast-reasoning", "text": "Hi there!", "usage": {"input_tokens": 6, "output_tokens": 3, "total_tokens": 238, "cost": null}, "error": null}
    

POST /v1/messages grok-4-20-reasoning

  1. Command:
    call /v1/messages '{"model":"grok-4-20-reasoning","max_tokens":64,"messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200 with x-litellm-response-cost: 2.25e-05 (10 x $1.25/M + 4 x $2.50/M); the reasoning tokens behind those 4 visible tokens are missing from this route's usage block, so they go unbilled here (LIT-7189)
    HTTP/1.1 200 OK
    x-litellm-response-cost: 2.25e-05
    x-litellm-response-cost-original: 2.25e-05
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 1.25e-05
    x-litellm-response-cost-output: 1e-05
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "grok-4-20-reasoning", "text": "Hi there friend.", "usage": {"input_tokens": 10, "output_tokens": 4}}
    

POST /v1/messages model-router

  1. Command:
    call /v1/messages '{"model":"model-router","max_tokens":64,"messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200 answered by azure_ai/grok-4-1-fast-reasoning, total 2.504e-05 = 5 output tokens x $0.50/M plus the routing fee once over the 161 tokens this route reports as the prompt (0 input plus 161 cache reads), where the Before block above pays it twice
    HTTP/1.1 200 OK
    x-litellm-response-cost: 2.504e-05
    x-litellm-response-cost-original: 2.504e-05
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 0.0
    x-litellm-response-cost-output: 2.4999999999999998e-06
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "azure_ai/grok-4-1-fast-reasoning", "text": "Hi, what's up?", "usage": {"input_tokens": 0, "output_tokens": 5, "cache_read_input_tokens": 161}}
    

POST /v1/messages gpt-chat-latest

  1. Command:
    call /v1/messages '{"model":"gpt-chat-latest","max_tokens":64,"messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: the same 400 from Azure as before, unrelated to pricing
    HTTP/1.1 400 Bad Request
    x-litellm-response-cost: 0
    {"model": null, "text": null, "usage": null, "error": {"type": "invalid_request_error", "message": "litellm.BadRequestError: Azure_aiException - {\n  \"error\": {\n    \"message\": \"Unsupported parameter: 'max_tokens' is not supported with this model. Use 'max_completion_tokens' instead.\",\n    \"type\": \"invalid_request_error\",\n    \"param\": \"max_tokens\",\n    \"code\": \"unsupported_parameter\"\n  }\n}. Received Model Group=gpt-chat-latest\nAvailable Model Group Fallbacks=None"}}
    

After (831a2a1), second rig

GET /model/info (second rig)

  1. Command: same as the Before block
  2. Output: the three new names priced, whisper per second with the transcription mode
    grok-4-20-non-reasoning | azure_ai/grok-4-20-non-reasoning | in 1.25e-06 | out 2.5e-06 | in/s None | max_in 262000 | mode chat
    model-router | azure_ai/model-router | in 1.4e-07 | out 0 | in/s None | max_in 200000 | mode chat
    whisper | azure_ai/whisper | in 0 | out 0 | in/s 0.0001 | max_in None | mode audio_transcription
    

POST /v1/chat/completions grok-4-20-non-reasoning

  1. Command:
    call /v1/chat/completions '{"model":"grok-4-20-non-reasoning","messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200 with x-litellm-response-cost: 2.5e-05 = 10 prompt tokens at $1.25/M plus 5 completion tokens at $2.50/M
    HTTP/1.1 200 OK
    x-litellm-response-cost: 2.5e-05
    x-litellm-response-cost-original: 2.5e-05
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 1.25e-05
    x-litellm-response-cost-output: 1.25e-05
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "grok-4-20-non-reasoning", "text": "Hi there, friend!", "usage": {"completion_tokens": 5, "prompt_tokens": 10, "total_tokens": 15, "audio_prompt_tokens": 0, "num_sources_used": 0, "cached_tokens": 0}}
    

POST /v1/chat/completions model-router, stream=true

  1. Command: same as the Before block
  2. Output: streamed answer from grok-4-1-fast-reasoning, and the last chunk's usage carries cost: 0.00031134 = 0.0003105 (621 reasoning tokens at the routed model's $0.50/M) plus the 6-token routing fee of 8.4e-07 once
    {"chunks": 7, "routed_model": "grok-4-1-fast-reasoning", "text": "Hi there, friend!", "usage": {"completion_tokens": 5, "prompt_tokens": 6, "total_tokens": 11, "completion_tokens_details": {"accepted_prediction_tokens": 0, "audio_tokens": 0, "reasoning_tokens": 621, "rejected_prediction_tokens": 0}, "prompt_tokens_details": {"audio_tokens": 0, "cached_tokens": 161, "text_tokens": 6, "image_tokens": 0}, "cost": 0.00031134}}
    

POST /v1/audio/transcriptions whisper

  1. Command: same as the Before block
  2. Output: 200 with x-litellm-response-cost: 0.0017577052154195013, 17.577 seconds at $0.0001 per second
    HTTP/1.1 200 OK
    x-litellm-response-cost: 0.0017577052154195013
    {"text":"Four score and seven years ago, our fathers brought forth on this continent a new nation, conceived in liberty and dedicated to the proposition that all men are created equal. Now we are engaged in a great civil war, testing whether that nation, or any nation so conceived and so dedicated, can long endure.","usage":null}
    

POST /v1/audio/transcriptions whisper-1 (control)

  1. Command: same as the Before block
  2. Output: unchanged
    HTTP/1.1 200 OK
    x-litellm-response-cost: 0.0017577052154195013
    {"text":"Four score and seven years ago, our fathers brought forth on this continent a new nation, conceived in liberty and dedicated to the proposition that all men are created equal. Now we are engaged in a great civil war, testing whether that nation, or any nation so conceived and so dedicated, can long endure.","usage":null}
    

Observations from the run, none caused or worsened by this PR:

  • gpt-chat-latest rejects max_tokens on /v1/messages upstream, both legs: left alone (fix(proxy): use max_completion_tokens for gpt-chat-latest health #30987)
  • The routing fee shows only in the total, since the breakdown headers carry no additional-costs line: left alone
  • grok deployments report xAI-style usage, so chat bills reasoning tokens but not the visible completion tokens and /v1/messages bills only the visible ones, both legs: LIT-7189
  • Unpriced names get zero breakdown headers but no total header: left alone
  • Streamed model-router prompt priced at zero, both legs: cached tokens outnumber prompt tokens
  • On /v1/messages the routing fee counts the 161 cache-read tokens the Anthropic-shape usage reports as the prompt, where chat counts the 6 prompt tokens Azure reports next to them; the same usage translation exists at the merge base: left alone

Verdict: PASS

No embedding or rerank name made the cut, so there is no embedding or rerank leg. codex-mini has no live leg: Azure refuses new codex-mini deployments (ServiceModelDeprecating) and none of the QA subscriptions has one, so it is covered by the metadata test only

The customer's list had 33 names with no azure_ai/ pricing. Seven are live Foundry catalog entries and ship here (gpt-chat-latest, codex-mini, whisper, model-router, cohere-command-a, grok-4-20-reasoning, grok-4-20-non-reasoning). The other 26 are left out on purpose:

  • FLUX.2-flex: covered by the open PR feat(azure_ai): support FLUX.2 flex images #39424, which adds the per-megapixel image pricing this PR does not model
  • ada, babbage, curie, davinci: retired on 2024-06-14
  • AI21-Jamba-Instruct: retired on 2025-03-01
  • AI21-Jamba-1.5-Large, AI21-Jamba-1.5-Mini: retired on 2025-08-01
  • Cohere-command-r, Cohere-command-r-plus: retired on 2025-06-30
  • Cohere-command-r-08-2024, Cohere-command-r-plus-08-2024: retired on 2026-05-12
  • Meta-Llama-3-8B-Instruct: retired on 2025-06-30
  • Mistral-Large-2411: retired on 2026-01-30
  • Kimi-K2-Thinking: retired on 2026-03-29
  • DeepSeek-R1-0528: retired on 2026-07-13
  • gpt-5.1-chat, gpt-5.2-chat, gpt-5.3-chat: retired on 2026-06-29 and 2026-05-13, replaced by gpt-chat-latest
  • gpt-35-turbo-16k (0613): retired on 2025-04-30
  • gpt-4-32k: retired on 2025-06-06
  • gpt-35-turbo, o1-mini, sora: absent from the Foundry Models API in every one of the 12 regions checked, shown as retired or archived in the catalog
  • gpt-oss-20b: no per-token Azure retail meter exists for it (only gpt-oss-120b is metered), so there is no list price to write
  • qwen3-32b: Azure lists only fine-tuned (FT) meters for it ($0.30/M in, $1.20/M out) and no base-model meter, so there is no list price for the catalog name

Type

🐛 Bug Fix

Caveats (if any)

Medium

  • grok deployments under-bill: chat skips visible tokens, /v1/messages skips reasoning tokens (LIT-7189)
    • xAI-style usage, on both legs, predates this PR
  • On a routed model-router response the fee shows only as the Azure Model Router Flat Cost breakdown line and in the total, not in x-litellm-response-cost-input or the OTel input cost attribute
    • the merge base put it in both places, which is the double charge; the Admin UI cost breakdown shows the line
  • gpt-chat-latest on /v1/messages returns Azure's max_tokens 400 on both legs

Low

  • Older proxies on the hosted map pay the fee three times on a router-echo response, twice on a routed one
    • LITELLM_LOCAL_MODEL_COST_MAP=True keeps an older proxy on its shipped map until it upgrades
  • 26 of the customer's 33 names stay unpriced on purpose; reasons under the proof
  • codex-mini: metadata test only, no live leg
    • Azure refuses new codex-mini deployments (ServiceModelDeprecating) and no QA subscription has an existing one
  • grok-4-20-* cached prompt tokens bill at the input price: Azure lists no cached meter for Grok 4.2
    • if Azure adds a discounted meter later, the entry over-bills until it is updated
  • A streamed model-router prompt prices at zero when the routed model reports more cached tokens than prompt tokens, both legs
  • Router detection is a substring match on model-router / model_router, with a second copy in the proxy request path, so a deployment alias containing that substring gets the flat fee
  • A router deployment under an alias without that substring still raises when the routed model is unmapped, as before the PR
  • cohere-command-a max_output_tokens is 8182, copied from Microsoft Learn as printed
  • model-router limits are the router deployment's own (200000 in, 32768 out); the routed model's limits apply per call

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Link to Devin session: https://app.devin.ai/sessions/c438be9d29cc43d7a76e43a46f560c4a
Open in Devin Desktop: https://app.devin.ai/desktop/session/c438be9d29cc43d7a76e43a46f560c4a?variant=devin
Requested by: @mateo-berri


Note

Medium Risk
Changes core spend calculation for Azure AI (catalog pricing and router fee logic); regressions would show up as incorrect totals or headers, though behavior is heavily test-covered.

Overview
Adds seven Azure AI Foundry catalog models to the cost map (gpt-chat-latest, codex-mini, whisper, model-router, cohere-command-a, grok-4-20-reasoning, grok-4-20-non-reasoning) so proxy/model info and response cost headers stop reporting $0 for those names.

Refactors Azure AI Model Router billing so the $0.14/M input routing fee is applied once: response tokens are priced on the actual model, the fee is folded into prompt cost when the priced name is router-shaped, and completion_cost skips the separate “additional costs” path when the response model name is already a router (avoiding the previous double charge). Router fee lookup also supports both model_router and model-router spellings.

Azure audio transcriptions now propagate custom_llm_provider into hidden params so azure_ai/whisper is billed against the new per-second entry instead of defaulting to azure.

Tests cover catalog metadata, router fee edge cases, cost breakdown lines, and transcription provider tagging.

Reviewed by Cursor Bugbot for commit 831a2a1. Bugbot is set up for automated code reviews on this repo. Configure here.

… router fee once

Add cost map entries for azure_ai/gpt-chat-latest, codex-mini, whisper,
model-router, cohere-command-a, grok-4-20-reasoning, and
grok-4-20-non-reasoning, priced from the live Azure AI Foundry and Azure
OpenAI pricing pages and the Azure Retail Prices API.

Skip the model router flat fee when the response model is the router
entry itself, since the generic cost already priced that fee. Before,
azure_ai/model_router charged it twice.

Resolves LIT-3157
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR. Add '(aside)' to your comment to have me ignore it.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@codspeed

codspeed Bot commented Sep 8, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_lit_3157_azure_ai_catalog_models (831a2a1) with litellm_internal_staging (075655c)

Open in CodSpeed

@greptile-apps

greptile-apps Bot commented Sep 8, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR adds pricing and metadata for seven Azure AI Foundry catalog names, ensures the Model Router fee is charged exactly once, and preserves the Azure AI provider identity for transcription cost calculation.

  • Adds synchronized primary and backup cost-map entries.
  • Refactors Azure AI router fee calculation and fallback behavior.
  • Propagates transcription provider metadata through synchronous and asynchronous paths.
  • Expands regression coverage for router billing, catalog metadata, and transcription pricing.

Confidence Score: 5/5

The PR appears safe to merge with no actionable regressions remaining.

The current implementation preserves provider attribution for transcription and charges Azure AI Model Router fees once across direct and routed response shapes. Earlier findings about the removed public keyword and test placement were manually resolved without explanation; the provider-boundary and substring findings were withdrawn after confirming they were not introduced by this PR, and the test-coverage concern was conceded after the cached-token case was restored.

Important Files Changed

Filename Overview
litellm/llms/azure_ai/cost_calculator.py Refactors Model Router detection and fee calculation so routed and router-shaped responses are charged once.
litellm/cost_calculator.py Avoids applying the separate Azure AI additional-cost path when the priced model already carries the router fee.
litellm/llms/azure/audio_transcriptions.py Preserves the selected Azure or Azure AI provider in transcription response metadata.
litellm/main.py Passes the parsed provider through the shared Azure transcription implementation.
model_prices_and_context_window.json Adds the seven Azure AI Foundry pricing and capability entries.
litellm/model_prices_and_context_window_backup.json Mirrors the new Foundry catalog metadata in the backup cost map.
tests/test_litellm/llms/azure_ai/test_azure_ai_cost_calculator.py Covers direct router names, routed responses, cached prompts, public keyword compatibility, and cost breakdowns.
tests/test_litellm/llms/azure/test_audio_transcriptions.py Verifies Azure AI transcription pricing and unchanged Azure provider attribution.
tests/test_litellm/llms/azure_ai/test_azure_ai_foundry_catalog_model_metadata.py Validates catalog registration, cost-map consistency, and model metadata.

Reviews (8): Last reviewed commit: "fix(azure): price azure_ai transcription..." | Re-trigger Greptile

Comment thread tests/test_litellm/llms/azure_ai/test_azure_ai_foundry_catalog_model_metadata.py Outdated
…d azure_ai directory

The new metadata test sat at the top of tests/test_litellm. The azure_ai metadata tests live in tests/test_litellm/llms/azure_ai next to the cost calculator test, so this moves it there and bumps its repo-root lookup by the two extra directory levels. No test changes.
@codecov

codecov Bot commented Sep 8, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.05882% with 1 line in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
litellm/llms/azure/audio_transcriptions.py 50.00% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

@devin-ai-integration

devin-ai-integration Bot commented Sep 8, 2026 •

Copy link
Copy Markdown
Contributor Author

Devin e2e QA

Tested at: 831a2a1
Session: https://app.devin.ai/sessions/c438be9d29cc43d7a76e43a46f560c4a

Both legs ran the same script against a two-worker proxy on port 37129 (checked free with a socket bind first) with no database, once at the merge base 9d0c9b9 and once at the PR tip 831a2a1. Every model points at one Azure AI Foundry resource (lit6549-audio-qa, swedencentral) where gpt-chat-latest, cohere-command-a, model-router, model_router, grok-4-20-reasoning, and grok-4.3 are deployed under those names. grok-4.3 is the control (priced before this PR); grok-4.6 could not be deployed because the subscription's grok-4.6 quota is fully used (50/50). model_router is a second deployment of the same Azure model router, added to cover the underscore alias. The model-router deployments answer with whichever model Azure picks, so token counts differ between legs and every call is real Azure spend

lit3157.yaml:

model_list:
  - model_name: gpt-chat-latest
    litellm_params:
      model: azure_ai/gpt-chat-latest
      api_base: os.environ/AZURE_AI_API_BASE
      api_key: os.environ/AZURE_AI_API_KEY
  - model_name: cohere-command-a
    litellm_params:
      model: azure_ai/cohere-command-a
      api_base: os.environ/AZURE_AI_API_BASE
      api_key: os.environ/AZURE_AI_API_KEY
  - model_name: model-router
    litellm_params:
      model: azure_ai/model-router
      api_base: os.environ/AZURE_AI_API_BASE
      api_key: os.environ/AZURE_AI_API_KEY
  - model_name: model_router
    litellm_params:
      model: azure_ai/model_router
      api_base: os.environ/AZURE_AI_API_BASE
      api_key: os.environ/AZURE_AI_API_KEY
  - model_name: grok-4-20-reasoning
    litellm_params:
      model: azure_ai/grok-4-20-reasoning
      api_base: os.environ/AZURE_AI_API_BASE
      api_key: os.environ/AZURE_AI_API_KEY
  - model_name: grok-4.3
    litellm_params:
      model: azure_ai/grok-4.3
      api_base: os.environ/AZURE_AI_API_BASE
      api_key: os.environ/AZURE_AI_API_KEY
general_settings:
  master_key: sk-1234
litellm_settings:
  drop_params: true

Proxy start, once per leg from a worktree at that commit (PYTHONPATH pointed at the worktree so the proxy imports that revision's price map):

LITELLM_LOCAL_MODEL_COST_MAP=True python litellm/proxy/proxy_cli.py --config lit3157.yaml --port 37129 --num_workers 2 --detailed_debug

Shell helper used by every POST case: it prints the status line, every x-litellm-response-cost* header, and the body trimmed to model, the first text, and usage without its *_details sub-objects

BASE=http://127.0.0.1:37129; KEY=sk-1234
call() {
  local route="$1" body="$2"
  local H="$(mktemp)" B="$(mktemp)"
  curl -s -D "$H" -o "$B" "${BASE}${route}" -H "Authorization: Bearer ${KEY}" -H 'Content-Type: application/json' -d "$body"
  grep -iE '^HTTP/|^x-litellm-response-cost' "$H" | tr -d '\r'
  python3 - "$B" <<'PY'
import json, sys
obj = json.load(open(sys.argv[1]))
def text(o):
    for path in (("choices", 0, "message", "content"), ("content", 0, "text"), ("output", -1, "content", 0, "text")):
        cur = o
        try:
            for p in path:
                cur = cur[p]
            return cur
        except Exception:
            continue
usage = obj.get("usage")
if isinstance(usage, dict):
    usage = {k: v for k, v in usage.items() if not k.endswith("_details")}
keep = {"model": obj.get("model"), "text": text(obj), "usage": usage}
if "error" in obj:
    keep["error"] = obj["error"]
print(json.dumps(keep)[:700])
PY
  rm -f "$H" "$B"
}

Before (9d0c9b9)

GET /model/info

  1. Command:
    curl -s $BASE/model/info -H "Authorization: Bearer $KEY" | python3 -c '
    import json, sys
    for m in json.load(sys.stdin)["data"]:
        mi = m["model_info"]
        print(m["model_name"], "|", m["litellm_params"]["model"], "| in", mi.get("input_cost_per_token"), "| out", mi.get("output_cost_per_token"), "| max_in", mi.get("max_input_tokens"), "| mode", mi.get("mode"))'
    
  2. Output: the four new names carry no price, no limit, and no mode; model_router and the grok-4.3 control were already priced
    gpt-chat-latest | azure_ai/gpt-chat-latest | in 0 | out 0 | max_in None | mode None
    cohere-command-a | azure_ai/cohere-command-a | in 0 | out 0 | max_in None | mode None
    model-router | azure_ai/model-router | in 0 | out 0 | max_in None | mode None
    model_router | azure_ai/model_router | in 1.4e-07 | out 0 | max_in None | mode chat
    grok-4-20-reasoning | azure_ai/grok-4-20-reasoning | in 0 | out 0 | max_in None | mode None
    grok-4.3 | azure_ai/grok-4.3 | in 1.25e-06 | out 2.5e-06 | max_in 200000 | mode chat
    

POST /v1/chat/completions gpt-chat-latest

  1. Command:
    call /v1/chat/completions '{"model":"gpt-chat-latest","messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200, tokens counted, no x-litellm-response-cost header
    HTTP/1.1 200 OK
    x-litellm-response-cost-original: 0.0
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 0.0
    x-litellm-response-cost-output: 0.0
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "gpt-chat-latest", "text": "Hi there, friend!", "usage": {"completion_tokens": 9, "prompt_tokens": 12, "total_tokens": 21, "latency_checkpoint": {"engine_tbt_ms": 5, "engine_ttft_ms": 109, "engine_ttlt_ms": 154, "pre_inference_ms": 64, "service_tbt_ms": 7, "service_ttft_ms": 247, "service_ttlt_ms": 299, "total_duration_ms": 242, "user_visible_ttft_ms": 183}}}
    

POST /v1/chat/completions cohere-command-a

  1. Command:
    call /v1/chat/completions '{"model":"cohere-command-a","messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200, no cost header
    HTTP/1.1 200 OK
    x-litellm-response-cost-original: 0.0
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 0.0
    x-litellm-response-cost-output: 0.0
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "cohere-command-a", "text": "Hi there, friend!", "usage": {"completion_tokens": 5, "prompt_tokens": 6, "total_tokens": 11, "audio_prompt_tokens": 0}}
    

POST /v1/chat/completions grok-4-20-reasoning

  1. Command:
    call /v1/chat/completions '{"model":"grok-4-20-reasoning","messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200, reasoning tokens counted in total_tokens, no cost header
    HTTP/1.1 200 OK
    x-litellm-response-cost-original: 0.0
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 0.0
    x-litellm-response-cost-output: 0.0
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "grok-4-20-reasoning", "text": "Hello, my friend.", "usage": {"completion_tokens": 5, "prompt_tokens": 10, "total_tokens": 891, "audio_prompt_tokens": 0, "num_sources_used": 0}}
    

POST /v1/chat/completions model-router

  1. Command:
    call /v1/chat/completions '{"model":"model-router","messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200 answered by azure_ai/grok-4-1-fast-reasoning; -input 8.4e-07 is the fee for 6 prompt tokens, and the total 0.00014118 exceeds input + output (0.00014034) by another 8.4e-07: the fee is charged twice
    HTTP/1.1 200 OK
    x-litellm-response-cost: 0.00014118
    x-litellm-response-cost-original: 0.00014118
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 8.4e-07
    x-litellm-response-cost-output: 0.0001395
    x-litellm-response-cost-reasoning: 0.0001395
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "azure_ai/grok-4-1-fast-reasoning", "text": "Hi there, friend!", "usage": {"completion_tokens": 5, "prompt_tokens": 6, "total_tokens": 290, "audio_prompt_tokens": 0, "num_sources_used": 0}}
    

POST /v1/chat/completions model_router

  1. Command:
    call /v1/chat/completions '{"model":"model_router","messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200, same routed model; total 0.00022718 = input 8.4e-07 + output 0.0002255 + a second 8.4e-07, fee twice on the underscore alias too
    HTTP/1.1 200 OK
    x-litellm-response-cost: 0.00022717999999999998
    x-litellm-response-cost-original: 0.00022717999999999998
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 8.4e-07
    x-litellm-response-cost-output: 0.00022549999999999998
    x-litellm-response-cost-reasoning: 0.00022549999999999998
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "azure_ai/grok-4-1-fast-reasoning", "text": "Hi there friend.", "usage": {"completion_tokens": 4, "prompt_tokens": 6, "total_tokens": 461, "audio_prompt_tokens": 0, "num_sources_used": 0}}
    

POST /v1/chat/completions grok-4.3 (control)

  1. Command:
    call /v1/chat/completions '{"model":"grok-4.3","messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200 with a real cost, this name was already priced
    HTTP/1.1 200 OK
    x-litellm-response-cost: 0.0018075
    x-litellm-response-cost-original: 0.0018075
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 1.25e-05
    x-litellm-response-cost-output: 0.0017950000000000002
    x-litellm-response-cost-reasoning: 0.0017950000000000002
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "grok-4.3", "text": "Hi there friend.", "usage": {"completion_tokens": 4, "prompt_tokens": 10, "total_tokens": 732, "audio_prompt_tokens": 0, "num_sources_used": 0}}
    

POST /v1/responses gpt-chat-latest

  1. Command:
    call /v1/responses '{"model":"gpt-chat-latest","input":"Say hi in three words."}'
    
  2. Output: 200, usage.cost is null, no cost header
    HTTP/1.1 200 OK
    x-litellm-response-cost-original: 0.0
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 0.0
    x-litellm-response-cost-output: 0.0
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "gpt-chat-latest", "text": "Hi there, friend!", "usage": {"input_tokens": 12, "output_tokens": 9, "total_tokens": 21, "cost": null}, "error": null}
    

POST /v1/responses grok-4-20-reasoning

  1. Command:
    call /v1/responses '{"model":"grok-4-20-reasoning","input":"Say hi in three words."}'
    
  2. Output: 200, no cost header
    HTTP/1.1 200 OK
    x-litellm-response-cost-original: 0.0
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 0.0
    x-litellm-response-cost-output: 0.0
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "grok-4-20-reasoning", "text": "Hi there friend.", "usage": {"input_tokens": 10, "output_tokens": 4, "total_tokens": 480, "cost": null}, "error": null}
    

POST /v1/responses model-router

  1. Command:
    call /v1/responses '{"model":"model-router","input":"Say hi in three words."}'
    
  2. Output: 200; total 8.668e-05 = input 8.4e-07 + output 8.5e-05 + a second 8.4e-07, fee twice
    HTTP/1.1 200 OK
    x-litellm-response-cost: 8.667999999999999e-05
    x-litellm-response-cost-original: 8.667999999999999e-05
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 8.4e-07
    x-litellm-response-cost-output: 8.499999999999999e-05
    x-litellm-response-cost-reasoning: 8.499999999999999e-05
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "azure_ai/grok-4-1-fast-reasoning", "text": "Hi there friend!", "usage": {"input_tokens": 6, "output_tokens": 4, "total_tokens": 180, "cost": null}, "error": null}
    

POST /v1/messages grok-4-20-reasoning

  1. Command:
    call /v1/messages '{"model":"grok-4-20-reasoning","max_tokens":64,"messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200, no cost header
    HTTP/1.1 200 OK
    x-litellm-response-cost-original: 0.0
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 0.0
    x-litellm-response-cost-output: 0.0
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "grok-4-20-reasoning", "text": "Hello my friend.", "usage": {"input_tokens": 10, "output_tokens": 4}}
    

POST /v1/messages cohere-command-a

  1. Command:
    call /v1/messages '{"model":"cohere-command-a","max_tokens":64,"messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200, no cost header
    HTTP/1.1 200 OK
    x-litellm-response-cost-original: 0.0
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 0.0
    x-litellm-response-cost-output: 0.0
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "cohere-command-a", "text": "Hi there, friend!", "usage": {"input_tokens": 6, "output_tokens": 5}}
    

POST /v1/messages gpt-chat-latest (max_completion_tokens)

  1. Command:
    call /v1/messages '{"model":"gpt-chat-latest","max_tokens":64,"max_completion_tokens":64,"messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 400 from Azure: the Messages route needs max_tokens to be accepted at all, and the proxy forwards it as max_tokens, which this model rejects (fix(proxy): use max_completion_tokens for gpt-chat-latest health #30987)
    HTTP/1.1 400 Bad Request
    x-litellm-response-cost: 0
    {"model": null, "text": null, "usage": null, "error": {"type": "invalid_request_error", "message": "litellm.BadRequestError: Azure_aiException - {\n  \"error\": {\n    \"message\": \"Unsupported parameter: 'max_tokens' is not supported with this model. Use 'max_completion_tokens' instead.\",\n    \"type\": \"invalid_request_error\",\n    \"param\": \"max_tokens\",\n    \"code\": \"unsupported_parameter\"\n  }\n}. Received Model Group=gpt-chat-latest\nAvailable Model Group Fallbacks=None"}}
    

POST /v1/messages model-router

  1. Command:
    call /v1/messages '{"model":"model-router","max_tokens":64,"messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200; -input 2.254e-05 is the fee over 161 cached prompt tokens, and total 4.708e-05 = 2.254e-05 + output 2e-06 + a second 2.254e-05, fee twice
    HTTP/1.1 200 OK
    x-litellm-response-cost: 4.708e-05
    x-litellm-response-cost-original: 4.708e-05
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 2.254e-05
    x-litellm-response-cost-output: 2e-06
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "azure_ai/grok-4-1-fast-reasoning", "text": "Hi there friend!", "usage": {"input_tokens": 0, "output_tokens": 4, "cache_read_input_tokens": 161}}
    

After (831a2a1)

GET /model/info

  1. Command:
    curl -s $BASE/model/info -H "Authorization: Bearer $KEY" | python3 -c '
    import json, sys
    for m in json.load(sys.stdin)["data"]:
        mi = m["model_info"]
        print(m["model_name"], "|", m["litellm_params"]["model"], "| in", mi.get("input_cost_per_token"), "| out", mi.get("output_cost_per_token"), "| max_in", mi.get("max_input_tokens"), "| mode", mi.get("mode"))'
    
  2. Output: every new name carries its list price, limit, and mode (gpt-chat-latest 272000, model-router 200000 at this tip); model-router and model_router both show the $0.14/M fee
    gpt-chat-latest | azure_ai/gpt-chat-latest | in 5e-06 | out 3e-05 | max_in 272000 | mode chat
    cohere-command-a | azure_ai/cohere-command-a | in 2.5e-06 | out 1e-05 | max_in 131072 | mode chat
    model-router | azure_ai/model-router | in 1.4e-07 | out 0 | max_in 200000 | mode chat
    model_router | azure_ai/model_router | in 1.4e-07 | out 0 | max_in 200000 | mode chat
    grok-4-20-reasoning | azure_ai/grok-4-20-reasoning | in 1.25e-06 | out 2.5e-06 | max_in 262000 | mode chat
    grok-4.3 | azure_ai/grok-4.3 | in 1.25e-06 | out 2.5e-06 | max_in 200000 | mode chat
    

POST /v1/chat/completions gpt-chat-latest

  1. Command:
    call /v1/chat/completions '{"model":"gpt-chat-latest","messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200 with x-litellm-response-cost: 0.00033 (12 x $5/M + 9 x $30/M)
    HTTP/1.1 200 OK
    x-litellm-response-cost: 0.00033
    x-litellm-response-cost-original: 0.00033
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 6.000000000000001e-05
    x-litellm-response-cost-output: 0.00027
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "gpt-chat-latest", "text": "Hi there, friend!", "usage": {"completion_tokens": 9, "prompt_tokens": 12, "total_tokens": 21, "latency_checkpoint": {"engine_tbt_ms": 5, "engine_ttft_ms": 121, "engine_ttlt_ms": 165, "pre_inference_ms": 86, "service_tbt_ms": 6, "service_ttft_ms": 242, "service_ttlt_ms": 292, "total_duration_ms": 212, "user_visible_ttft_ms": 157}}}
    

POST /v1/chat/completions cohere-command-a

  1. Command:
    call /v1/chat/completions '{"model":"cohere-command-a","messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200 with x-litellm-response-cost: 6.5e-05 (6 x $2.50/M + 5 x $10/M)
    HTTP/1.1 200 OK
    x-litellm-response-cost: 6.500000000000001e-05
    x-litellm-response-cost-original: 6.500000000000001e-05
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 1.5000000000000002e-05
    x-litellm-response-cost-output: 5e-05
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "cohere-command-a", "text": "Hi there, friend!", "usage": {"completion_tokens": 5, "prompt_tokens": 6, "total_tokens": 11, "audio_prompt_tokens": 0}}
    

POST /v1/chat/completions grok-4-20-reasoning

  1. Command:
    call /v1/chat/completions '{"model":"grok-4-20-reasoning","messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200 with x-litellm-response-cost: 0.0019025 (10 x $1.25/M + 756 reasoning tokens x $2.50/M)
    HTTP/1.1 200 OK
    x-litellm-response-cost: 0.0019025000000000001
    x-litellm-response-cost-original: 0.0019025000000000001
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 1.25e-05
    x-litellm-response-cost-output: 0.0018900000000000002
    x-litellm-response-cost-reasoning: 0.0018900000000000002
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "grok-4-20-reasoning", "text": "Hey there friend.", "usage": {"completion_tokens": 4, "prompt_tokens": 10, "total_tokens": 770, "audio_prompt_tokens": 0, "num_sources_used": 0}}
    

POST /v1/chat/completions model-router

  1. Command:
    call /v1/chat/completions '{"model":"model-router","messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200 answered by azure_ai/grok-4-1-fast-reasoning; total 0.00023984 = output 0.000239 + 8.4e-07, the fee for 6 prompt tokens charged once. At this tip the fee is booked as the router's additional-cost line, so -input reads 0.0 instead of 8.4e-07
    HTTP/1.1 200 OK
    x-litellm-response-cost: 0.00023983999999999998
    x-litellm-response-cost-original: 0.00023983999999999998
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 0.0
    x-litellm-response-cost-output: 0.00023899999999999998
    x-litellm-response-cost-reasoning: 0.00023899999999999998
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "azure_ai/grok-4-1-fast-reasoning", "text": "Hi there, friend!", "usage": {"completion_tokens": 5, "prompt_tokens": 6, "total_tokens": 489, "audio_prompt_tokens": 0, "num_sources_used": 0}}
    

POST /v1/chat/completions model_router

  1. Command:
    call /v1/chat/completions '{"model":"model_router","messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200; total 0.00018584 = output 0.000185 + 8.4e-07, fee once on the underscore alias too
    HTTP/1.1 200 OK
    x-litellm-response-cost: 0.00018584
    x-litellm-response-cost-original: 0.00018584
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 0.0
    x-litellm-response-cost-output: 0.000185
    x-litellm-response-cost-reasoning: 0.000185
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "azure_ai/grok-4-1-fast-reasoning", "text": "Hi there, friend!", "usage": {"completion_tokens": 5, "prompt_tokens": 6, "total_tokens": 381, "audio_prompt_tokens": 0, "num_sources_used": 0}}
    

POST /v1/chat/completions grok-4.3 (control)

  1. Command:
    call /v1/chat/completions '{"model":"grok-4.3","messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200 with the same per-token prices as before
    HTTP/1.1 200 OK
    x-litellm-response-cost: 0.0017933
    x-litellm-response-cost-original: 0.0017933
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 7.500000000000002e-06
    x-litellm-response-cost-output: 0.0017850000000000001
    x-litellm-response-cost-cache-read: 8e-07
    x-litellm-response-cost-reasoning: 0.0017850000000000001
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "grok-4.3", "text": "Hi there friend", "usage": {"completion_tokens": 3, "prompt_tokens": 10, "total_tokens": 727, "audio_prompt_tokens": 0, "num_sources_used": 0}}
    

POST /v1/responses gpt-chat-latest

  1. Command:
    call /v1/responses '{"model":"gpt-chat-latest","input":"Say hi in three words."}'
    
  2. Output: 200 with x-litellm-response-cost: 0.00033, same as chat completions
    HTTP/1.1 200 OK
    x-litellm-response-cost: 0.00033
    x-litellm-response-cost-original: 0.00033
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 6.000000000000001e-05
    x-litellm-response-cost-output: 0.00027
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "gpt-chat-latest", "text": "Hi there, friend!", "usage": {"input_tokens": 12, "output_tokens": 9, "total_tokens": 21, "cost": null}, "error": null}
    

POST /v1/responses grok-4-20-reasoning

  1. Command:
    call /v1/responses '{"model":"grok-4-20-reasoning","input":"Say hi in three words."}'
    
  2. Output: 200 with x-litellm-response-cost: 0.0015075
    HTTP/1.1 200 OK
    x-litellm-response-cost: 0.0015075
    x-litellm-response-cost-original: 0.0015075
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 1.25e-05
    x-litellm-response-cost-output: 0.001495
    x-litellm-response-cost-reasoning: 0.001495
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "grok-4-20-reasoning", "text": "Hello my friend.", "usage": {"input_tokens": 10, "output_tokens": 4, "total_tokens": 612, "cost": null}, "error": null}
    

POST /v1/responses model-router

  1. Command:
    call /v1/responses '{"model":"model-router","input":"Say hi in three words."}'
    
  2. Output: 200; total 0.00014634 = output 0.0001455 + 8.4e-07, fee once
    HTTP/1.1 200 OK
    x-litellm-response-cost: 0.00014633999999999998
    x-litellm-response-cost-original: 0.00014633999999999998
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 0.0
    x-litellm-response-cost-output: 0.00014549999999999999
    x-litellm-response-cost-reasoning: 0.00014549999999999999
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "azure_ai/grok-4-1-fast-reasoning", "text": "Hi there, friend!", "usage": {"input_tokens": 6, "output_tokens": 5, "total_tokens": 302, "cost": null}, "error": null}
    

POST /v1/messages grok-4-20-reasoning

  1. Command:
    call /v1/messages '{"model":"grok-4-20-reasoning","max_tokens":64,"messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200 priced (-original 0.0012825, input 1.25e-05, reasoning included). The total x-litellm-response-cost header is absent whenever -reasoning is present on /v1/messages; header code is untouched by this PR
    HTTP/1.1 200 OK
    x-litellm-response-cost-original: 0.0012825
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 1.25e-05
    x-litellm-response-cost-output: 0.00127
    x-litellm-response-cost-reasoning: 0.00127
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "grok-4-20-reasoning", "text": "Hello there friend.", "usage": {"input_tokens": 10, "output_tokens": 4}}
    

POST /v1/messages cohere-command-a

  1. Command:
    call /v1/messages '{"model":"cohere-command-a","max_tokens":64,"messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200 with x-litellm-response-cost: 6.5e-05
    HTTP/1.1 200 OK
    x-litellm-response-cost: 6.500000000000001e-05
    x-litellm-response-cost-original: 6.500000000000001e-05
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 1.5000000000000002e-05
    x-litellm-response-cost-output: 5e-05
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "cohere-command-a", "text": "Hi there, friend!", "usage": {"input_tokens": 6, "output_tokens": 5}}
    

POST /v1/messages gpt-chat-latest (max_completion_tokens)

  1. Command:
    call /v1/messages '{"model":"gpt-chat-latest","max_tokens":64,"max_completion_tokens":64,"messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: same 400 as before: the proxy still sends max_tokens upstream even with max_completion_tokens in the body (fix(proxy): use max_completion_tokens for gpt-chat-latest health #30987), unrelated to pricing
    HTTP/1.1 400 Bad Request
    x-litellm-response-cost: 0
    {"model": null, "text": null, "usage": null, "error": {"type": "invalid_request_error", "message": "litellm.BadRequestError: Azure_aiException - {\n  \"error\": {\n    \"message\": \"Unsupported parameter: 'max_tokens' is not supported with this model. Use 'max_completion_tokens' instead.\",\n    \"type\": \"invalid_request_error\",\n    \"param\": \"max_tokens\",\n    \"code\": \"unsupported_parameter\"\n  }\n}. Received Model Group=gpt-chat-latest\nAvailable Model Group Fallbacks=None"}}
    

POST /v1/messages model-router

  1. Command:
    call /v1/messages '{"model":"model-router","max_tokens":64,"messages":[{"role":"user","content":"Say hi in three words."}]}'
    
  2. Output: 200; total 2.404e-05 = output 1.5e-06 + 2.254e-05, the fee over 161 cached prompt tokens charged once
    HTTP/1.1 200 OK
    x-litellm-response-cost: 2.404e-05
    x-litellm-response-cost-original: 2.404e-05
    x-litellm-response-cost-discount-amount: 0.0
    x-litellm-response-cost-margin-amount: 0.0
    x-litellm-response-cost-margin-percent: 0.0
    x-litellm-response-cost-input: 0.0
    x-litellm-response-cost-output: 1.5e-06
    x-litellm-response-cost-tool-usage: 0.0
    {"model": "azure_ai/grok-4-1-fast-reasoning", "text": "Hi there!", "usage": {"input_tokens": 0, "output_tokens": 3, "cache_read_input_tokens": 161}}
    

Observations from the run, none caused by this PR:

  • gpt-chat-latest on /v1/messages returns Azure's 400 on both legs even when the request carries max_completion_tokens: the Anthropic Messages schema requires max_tokens, and the proxy forwards it upstream as max_tokens (fix(proxy): use max_completion_tokens for gpt-chat-latest health #30987). Pricing of that name on /v1/messages therefore could not be observed; cohere-command-a and grok-4-20-reasoning prove the /v1/messages cost path instead
  • On /v1/messages, when x-litellm-response-cost-reasoning is present (grok-4-20-reasoning at the tip, three repeat calls), the total x-litellm-response-cost header is missing while -original, -input, -output are set. This PR does not touch header code; the name was unpriced at the base so the same call could not show it there
  • At the base every router call carries the $0.14/M fee twice: once in -input and once more as the additional-cost line, so the total exceeds input + output by exactly one fee (8.4e-07 for 6 prompt tokens, 2.254e-05 for 161 cached tokens). At the tip the total equals output + one fee on model-router, model_router, /v1/responses, and /v1/messages; the fee now lives only in the additional-cost line, so -input reads 0.0
  • grok-4.6 control replaced by grok-4.3 for quota reasons; its prices are unchanged between legs

Verdict: PASS

…imits

The router fee was folded into azure_ai.cost_per_token and then added again
by the additional_costs hook, so every routed request paid it twice. The hook
now owns the fee, the entry named by the deployment supplies the price, and a
response priced as the router entry itself is not charged again

model-router, gpt-chat-latest and cohere-command-a carry the limits from the
Foundry models page, and model-router and grok-4-20-* carry their retirement
dates. The router tests now run at the completion_cost level with a Logging
object, which is the path the proxy takes, and fail at the merge base
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

Comment thread litellm/cost_calculator.py
Restores the public keyword removed at 415bdbf. A direct caller that
names the Model Router as the request model gets the routing fee folded
into the prompt cost once; completion_cost never passes it and charges
the fee through the additional-costs hook as before
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

…ce grok-4-20 cache reads

Direct litellm.cost_per_token callers that name a Model Router deployment as
the model get the routing fee again, as they did before this branch, and the
fee is still charged exactly once on every completion_cost path. The
grok-4-20 entries bill cached prompt tokens at the input rate, since Azure has
no cached-input meter for them, and the model_router twin carries the same
limits and retirement date as model-router. The catalog test now exercises
the cost calculator and map relations instead of pinning map fields.
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

Comment thread litellm/llms/azure_ai/cost_calculator.py
…lculator

litellm's azure_ai config rejects reasoning_effort for gpt-chat-latest and Azure documents a fixed reasoning level for it, so the entry no longer advertises reasoning_effort_levels. The catalog metadata tests compare cost_per_token and the whisper transcription cost with the entry the calculator read instead of with list-price literals, the pattern #40195 removed
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai


# Get additional costs from provider (e.g., routing fees, infrastructure costs)
if custom_llm_provider == "azure_ai":
if custom_llm_provider == "azure_ai" and not azure_ai_is_model_router_name(model):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Provider Logic Escapes Boundary

The generic cost-calculation path now imports an Azure AI router predicate and branches explicitly on custom_llm_provider == "azure_ai". This violates the repository directive to keep provider-specific behavior under llms/, so the provider boundary must be restored before merging.

Rule Used: What: Avoid writing provider-specific code outside... (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The azure_ai branch and provider import are the merge base's own hook. This tip only narrows it so the fee is charged once

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You're right. I rechecked the merge-base diff: the existing azure_ai cost-calculator import and custom_llm_provider == "azure_ai" hook are unchanged; this tip only adds the provider-owned router predicate and uses it to skip the existing additional-cost hook when the router fee is already included. My comment treated pre-existing provider integration as introduced by this PR, so it is not actionable for this diff. I withdraw the finding.

Comment on lines +59 to +65
entry: Final = get_model_info(f"azure_ai/{catalog_name}")
prompt_cost, completion_cost_usd = cost_per_token(
model=f"azure_ai/{catalog_name}", prompt_tokens=A_MILLION, completion_tokens=A_MILLION
)
assert prompt_cost > 0
assert prompt_cost == pytest.approx(A_MILLION * entry["input_cost_per_token"])
assert completion_cost_usd == pytest.approx(A_MILLION * entry["output_cost_per_token"])

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Pricing Tests Lose Independence

These modified tests derive expected token costs through get_model_info, which reads the same cost map used by cost_per_token; an incorrect map price therefore passes both sides of the assertion. The Whisper test also replaces its fixed $0.36/hour expectation with relative scaling, while the router suite removes the cached-token billing case. This weakens existing regression coverage and violates the repository directive to prevent test modifications from masking regressions, so independent expected prices and equivalent cached-token coverage must be restored before merging.

Rule Used: What: Flag any modifications to existing tests and... (source)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Literal prices are what #40195 just removed repo-wide, and these tests fail without the entries. The cached-token fee case is back in 3cadf2f

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You're right on both points. Given #40195 removed repo-wide literal price assertions, deriving the expected values from the registered entry is appropriate here: the test verifies that each catalog name is present, has a nonzero input price, and that cost_per_token actually uses that entry. The source/backup comparison separately guards map consistency. And with the cached-token fee case restored in 3cadf2f8f7, my previous coverage objection is resolved. No further issue from this comment.

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

1 similar comment
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 831a2a1. Configure here.

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@mateo-berri
mateo-berri merged commit ee7c7e1 into litellm_internal_staging Sep 9, 2026
83 of 84 checks passed
@mateo-berri
mateo-berri deleted the litellm_lit_3157_azure_ai_catalog_models branch September 9, 2026 03:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant