Skip to content

fix(cost): bill cached realtime audio tokens at the audio cache-read rate - #40627

Merged
mateo-berri merged 14 commits into
mainfrom
litellm_fix_realtime_cached_audio_cost
Sep 14, 2026
Merged

mateo-berri merged 14 commits into
mainfrom
litellm_fix_realtime_cached_audio_cost

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Realtime requests with cached audio input are overbilled, about 2x on the ticket's turn
  • OpenAI's cached_tokens_details text/audio split was dropped at every layer
  • Cached audio tokens were then billed at the full $32/M audio rate

How it solves it:

  • Carry cached_tokens_details through realtime usage into prompt_tokens_details
  • Bill cached audio at cache_read_input_audio_token_cost, fallback to cache_read_input_token_cost
  • Split the cost breakdown's cache read cost the same way, so cached audio is priced at the audio cache-read rate there too
  • Carry cache_read_input_audio_token_cost through get_model_info, so proxy and router lookups bill cached audio at that rate (the mini realtime family was falling back to the text cache-read rate)
  • Subtract cached text/audio/image from the full-rate buckets
  • Sum the nested split when combining a session's response.done events
  • Add cache_read_input_audio_token_cost to gpt-realtime, gpt-realtime-1.5, gpt-realtime-2025-08-28
  • Fill the realtime entries that still lacked a cache-read rate: gpt-realtime-mini gains cache_read_input_token_cost $0.06/M (it billed every cached token at $0 before), azure/gpt-realtime-mini and azure/gpt-realtime-mini-2025-10-06 gain cache_read_input_audio_token_cost $0.30/M, and azure/gpt-realtime-2025-08-28 and azure/gpt-realtime-1.5-2026-02-23 gain cache_read_input_audio_token_cost $0.40/M with cache_read_input_token_cost lowered from $4/M (the full input rate) to $0.40/M. The rates are OpenAI's pricing page for the mini family and Azure's retail prices API for the Azure entries (Global meters gpt rt txt 0828 cchd Inp glbl, gpt rt aud 0828 cchd Inp glbl, gpt rt 1.5 txt cd inp Gl, gpt rt 1.5 aud cd inp Gl, gpt rt txt mini cchd Inp glbl, gpt rt aud mini cchd Inp glbl), and each Azure entry now matches the OpenAI entry for the same model

User Flow

Before: a voice app on gpt-realtime-2 sees each turn's spend come out about double what OpenAI charges

  1. They open a websocket to wss://litellm-domain/v1/realtime?model=gpt-realtime-2 and hold a multi-turn audio conversation
  2. A turn's response.done event reports input_tokens: 283, text_tokens: 116, audio_tokens: 167, cached_tokens: 192, with cached_tokens_details: {text_tokens: 64, audio_tokens: 128}
  3. They open https://litellm-domain/ui/?page=logs and the turn's input spend reads $0.0029888, against OpenAI's $0.0015328

After: the same turn is billed at what OpenAI charges

  1. They open a websocket to wss://litellm-domain/v1/realtime?model=gpt-realtime-2 and hold the same conversation
  2. The same response.done event arrives with the same usage
  3. https://litellm-domain/ui/?page=logs shows the turn's input spend at $0.0015328: 52 text at $4/M, 39 audio at $32/M, and all 192 cached tokens at $0.40/M

Relevant issues

Linear ticket

Resolves LIT-4994

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Both legs run the same 4-turn websocket session on three models: gpt-realtime-2 (the customer's model, where the text and audio cache-read rates coincide), gpt-realtime-2.1-mini (where they differ), and gpt-realtime-mini (whose cost map entry had no cache-read rate before this PR), against a live proxy with 2 uvicorn workers, its own Postgres database, and the checkout's own cost map, real OpenAI calls. The before proxy is the merge base c134fb7 on port 32566 and the after proxy is this PR's tip a94c060 on port 42155, each booted with

LITELLM_LOCAL_MODEL_COST_MAP=True LITELLM_MASTER_KEY=$KEY PROXY_BASE_URL=http://localhost:$PORT python litellm/proxy/proxy_cli.py --config config.yaml --port $PORT --num_workers 2 --use_v2_migration_resolver

config.yaml:

model_list:
  - model_name: gpt-realtime-2
    litellm_params:
      model: openai/gpt-realtime-2
      api_key: os.environ/OPENAI_API_KEY
  - model_name: gpt-5-mini
    litellm_params:
      model: openai/gpt-5-mini
      api_key: os.environ/OPENAI_API_KEY
  - model_name: gpt-audio-mini
    litellm_params:
      model: openai/gpt-audio-mini
      api_key: os.environ/OPENAI_API_KEY
  - model_name: claude-haiku-4-5
    litellm_params:
      model: anthropic/claude-haiku-4-5
      api_key: os.environ/ANTHROPIC_API_KEY
  - model_name: gpt-realtime-2.1-mini
    litellm_params:
      model: openai/gpt-realtime-2.1-mini
      api_key: os.environ/OPENAI_API_KEY
  - model_name: gpt-realtime-mini
    litellm_params:
      model: openai/gpt-realtime-mini
      api_key: os.environ/OPENAI_API_KEY

general_settings:
  master_key: os.environ/LITELLM_MASTER_KEY
  database_url: os.environ/DATABASE_URL

The client opens ws://localhost:$PORT/v1/realtime?model=$MODEL with Authorization: Bearer $KEY, sends session.update (text output, pcm16 input, no turn detection), and for each of 4 turns appends about 25s of pcm16 speech (generated once with POST /v1/audio/speech, response_format: pcm), commits, and calls response.create, printing each response.done usage. Once the conversation crosses OpenAI's 1024-token cache threshold, cached_tokens_details reports a text/audio split. The script then adds up the session, prints what it costs at the model's cost map rates ((text - cached text) x input rate + (audio - cached audio) x audio rate + cached text x cache-read rate + cached audio x audio cache-read rate + output text x output rate), and reads the spend row back through GET /spend/logs

for MODEL in gpt-realtime-2 gpt-realtime-2.1-mini gpt-realtime-mini; do python realtime_session.py $PORT $KEY speech.pcm 4 $MODEL; done
realtime_session.py
import asyncio, base64, json, sys, time, urllib.request
from datetime import datetime, timezone

port, key, audio_path, turns, model = sys.argv[1], sys.argv[2], sys.argv[3], int(sys.argv[4]), sys.argv[5]
pcm = open(audio_path, "rb").read()
RATES = {
    "gpt-realtime-2": {"text": 4e-06, "audio": 3.2e-05, "cache_text": 4e-07, "cache_audio": 4e-07, "out_text": 2.4e-05},
    "gpt-realtime-2.1-mini": {"text": 6e-07, "audio": 1e-05, "cache_text": 6e-08, "cache_audio": 3e-07, "out_text": 2.4e-06},
    "gpt-realtime-mini": {"text": 6e-07, "audio": 1e-05, "cache_text": 6e-08, "cache_audio": 3e-07, "out_text": 2.4e-06},
}


async def run():
    import websockets

    url = f"ws://localhost:{port}/v1/realtime?model={model}"
    totals = {"text": 0, "audio": 0, "cached": 0, "cached_text": 0, "cached_audio": 0, "out_text": 0, "out_audio": 0}
    async with websockets.connect(url, additional_headers={"Authorization": f"Bearer {key}"}, max_size=None) as ws:
        await ws.send(json.dumps({"type": "session.update", "session": {"type": "realtime", "output_modalities": ["text"], "audio": {"input": {"format": {"type": "audio/pcm", "rate": 24000}, "turn_detection": None}}}}))
        for turn in range(1, turns + 1):
            for i in range(0, len(pcm), 96000):
                await ws.send(json.dumps({"type": "input_audio_buffer.append", "audio": base64.b64encode(pcm[i:i + 96000]).decode()}))
            await ws.send(json.dumps({"type": "input_audio_buffer.commit"}))
            await ws.send(json.dumps({"type": "response.create"}))
            while True:
                ev = json.loads(await ws.recv())
                if ev["type"] == "error":
                    print("error:", json.dumps(ev)[:400]); return
                if ev["type"] == "response.done":
                    u = ev["response"]["usage"]
                    d, o = u["input_token_details"], u["output_token_details"]
                    ct = d.get("cached_tokens_details") or {}
                    print(f"turn {turn}: input={u['input_tokens']} text={d['text_tokens']} audio={d['audio_tokens']} cached={d['cached_tokens']} cached_tokens_details={ct} output={u['output_tokens']}")
                    totals["text"] += d["text_tokens"]; totals["audio"] += d["audio_tokens"]; totals["cached"] += d["cached_tokens"]
                    totals["cached_text"] += ct.get("text_tokens", 0); totals["cached_audio"] += ct.get("audio_tokens", 0)
                    totals["out_text"] += o.get("text_tokens", 0); totals["out_audio"] += o.get("audio_tokens", 0)
                    break
    print(f"session totals: text={totals['text']} audio={totals['audio']} cached={totals['cached']} (cached text={totals['cached_text']}, cached audio={totals['cached_audio']}) output text={totals['out_text']} audio={totals['out_audio']}")
    return totals


started = datetime.now(timezone.utc)
totals = asyncio.run(run())
if totals is None:
    sys.exit(1)
r = RATES[model]
expected = ((totals["text"] - totals["cached_text"]) * r["text"] + (totals["audio"] - totals["cached_audio"]) * r["audio"] + totals["cached_text"] * r["cache_text"] + totals["cached_audio"] * r["cache_audio"] + totals["out_text"] * r["out_text"])
print(f"expected cost at the cost map rates for {model} (text ${r['text']*1e6:g}/M, audio ${r['audio']*1e6:g}/M, cached text ${r['cache_text']*1e6:g}/M, cached audio ${r['cache_audio']*1e6:g}/M, output text ${r['out_text']*1e6:g}/M): {round(expected, 7)}")
deadline = time.time() + 90
while time.time() < deadline:
    req = urllib.request.Request(f"http://localhost:{port}/spend/logs", headers={"Authorization": f"Bearer {key}"})
    rows = [r_ for r_ in json.load(urllib.request.urlopen(req)) if datetime.fromisoformat(r_["startTime"].replace("Z", "+00:00")) >= started and model in r_["model"]]
    if rows:
        for r_ in rows:
            print(f"GET /spend/logs (rows since this session started) -> model={r_['model']} spend={r_['spend']} prompt_tokens={r_['prompt_tokens']} completion_tokens={r_['completion_tokens']} request_id={r_['request_id']}")
        break
    time.sleep(3)
else:
    print("GET /spend/logs -> no rows for this session after 90s")

To see the same rows in the Admin UI, open http://localhost:$PORT/ui/logs, log in as admin with the master key, and click the gpt-realtime-2, gpt-realtime-2.1-mini, or gpt-realtime-mini row; the detail pane shows Cost and Prompt Cache Read Tokens

gpt-realtime-2, before (c134fb7)

turn 1: input=443 text=126 audio=317 cached=0 cached_tokens_details={'text_tokens': 0, 'audio_tokens': 0, 'image_tokens': 0} output=312
turn 2: input=1009 text=375 audio=634 cached=0 cached_tokens_details={'text_tokens': 0, 'audio_tokens': 0, 'image_tokens': 0} output=252
turn 3: input=1518 text=567 audio=951 cached=1024 cached_tokens_details={'text_tokens': 384, 'audio_tokens': 640, 'image_tokens': 0} output=254
turn 4: input=2030 text=762 audio=1268 cached=1536 cached_tokens_details={'text_tokens': 576, 'audio_tokens': 960, 'image_tokens': 0} output=208
session totals: text=1830 audio=3170 cached=2560 (cached text=960, cached audio=1600) output text=1026 audio=0
expected cost at the cost map rates for gpt-realtime-2 (text $4/M, audio $32/M, cached text $0.4/M, cached audio $0.4/M, output text $24/M): 0.079368
GET /spend/logs (rows since this session started) -> model=openai/gpt-realtime-2 spend=0.103728 prompt_tokens=5000 completion_tokens=1026 request_id=b070841a-9ccb-49d8-bdab-ec97fe54a9bc

Logged spend is $0.103728 against $0.079368 at the cost map rates, $0.02436 (31%) over. With the cached text/audio split dropped, all 2560 cached tokens are subtracted from the 1830 text tokens first and the 730 remainder from audio, so 2440 audio tokens bill at $32/M where only 1570 were uncached

Logs page and the row's detail pane:

pr40627-a94c060b84-shot_before7_r2_logs.png
pr40627-a94c060b84-shot_before7_r2_logs_detail.png

gpt-realtime-2, after (a94c060)

turn 1: input=443 text=126 audio=317 cached=0 cached_tokens_details={'text_tokens': 0, 'audio_tokens': 0, 'image_tokens': 0} output=399
turn 2: input=1064 text=430 audio=634 cached=448 cached_tokens_details={'text_tokens': 128, 'audio_tokens': 320, 'image_tokens': 0} output=255
turn 3: input=1577 text=626 audio=951 cached=1088 cached_tokens_details={'text_tokens': 448, 'audio_tokens': 640, 'image_tokens': 0} output=236
turn 4: input=2107 text=839 audio=1268 cached=1600 cached_tokens_details={'text_tokens': 640, 'audio_tokens': 960, 'image_tokens': 0} output=212
session totals: text=2021 audio=3170 cached=3136 (cached text=1216, cached audio=1920) output text=1102 audio=0
expected cost at the cost map rates for gpt-realtime-2 (text $4/M, audio $32/M, cached text $0.4/M, cached audio $0.4/M, output text $24/M): 0.0709224
GET /spend/logs (rows since this session started) -> model=openai/gpt-realtime-2 spend=0.0709224 prompt_tokens=5191 completion_tokens=1102 request_id=dea5e569-65a4-4279-993e-87896cd906d9

Logged spend equals the cost map rate cost to the last digit: 805 non-cached text at $4/M, 1250 non-cached audio at $32/M, 3136 cached at $0.40/M, 1102 output text at $24/M

pr40627-a94c060b84-shot_after8_r2_logs.png
pr40627-a94c060b84-shot_after8_r2_logs_detail.png

gpt-realtime-2.1-mini, before (c134fb7)

turn 1: input=443 text=126 audio=317 cached=0 cached_tokens_details={'text_tokens': 0, 'audio_tokens': 0, 'image_tokens': 0} output=218
turn 2: input=907 text=273 audio=634 cached=448 cached_tokens_details={'text_tokens': 128, 'audio_tokens': 320, 'image_tokens': 0} output=210
turn 3: input=1390 text=439 audio=951 cached=896 cached_tokens_details={'text_tokens': 256, 'audio_tokens': 640, 'image_tokens': 0} output=179
turn 4: input=1854 text=586 audio=1268 cached=1408 cached_tokens_details={'text_tokens': 448, 'audio_tokens': 960, 'image_tokens': 0} output=222
session totals: text=1424 audio=3170 cached=2752 (cached text=832, cached audio=1920) output text=829 audio=0
expected cost at the cost map rates for gpt-realtime-2.1-mini (text $0.6/M, audio $10/M, cached text $0.06/M, cached audio $0.3/M, output text $2.4/M): 0.0154707
GET /spend/logs (rows since this session started) -> model=openai/gpt-realtime-2.1-mini spend=0.02057472 prompt_tokens=4594 completion_tokens=829 request_id=c9fe4048-bce5-4c4b-82c4-b616042d8504

Logged spend is $0.02057472 against $0.0154707 at the cost map rates, 33% over: the 2752 cached tokens are subtracted from the 1424 text tokens first and the 1328 remainder from audio, so 1842 audio tokens bill at $10/M where only 1250 were uncached, and every cached token bills at the $0.06/M text cache-read rate

pr40627-a94c060b84-shot_before7_21mini_logs_detail.png

gpt-realtime-2.1-mini, after (a94c060)

turn 1: input=443 text=126 audio=317 cached=0 cached_tokens_details={'text_tokens': 0, 'audio_tokens': 0, 'image_tokens': 0} output=448
turn 2: input=1103 text=469 audio=634 cached=0 cached_tokens_details={'text_tokens': 0, 'audio_tokens': 0, 'image_tokens': 0} output=255
turn 3: input=1634 text=683 audio=951 cached=1088 cached_tokens_details={'text_tokens': 448, 'audio_tokens': 640, 'image_tokens': 0} output=173
turn 4: input=2076 text=808 audio=1268 cached=1664 cached_tokens_details={'text_tokens': 704, 'audio_tokens': 960, 'image_tokens': 0} output=236
session totals: text=2086 audio=3170 cached=2752 (cached text=1152, cached audio=1600) output text=1112 audio=0
expected cost at the cost map rates for gpt-realtime-2.1-mini (text $0.6/M, audio $10/M, cached text $0.06/M, cached audio $0.3/M, output text $2.4/M): 0.0194783
GET /spend/logs (rows since this session started) -> model=openai/gpt-realtime-2.1-mini spend=0.01947832 prompt_tokens=5256 completion_tokens=1112 request_id=40f940a0-864e-42a3-ae71-a5a72ad555a0

Logged spend $0.01947832 is the cost map rate cost (the expected line rounds it to 7 places): 934 non-cached text at $0.60/M, 1570 non-cached audio at $10/M, 1152 cached text at $0.06/M, 1600 cached audio at $0.30/M, 1112 output text at $2.40/M. Without the get_model_info change the same session bills cached audio at the $0.06/M text cache-read rate

pr40627-a94c060b84-shot_after8_21mini_logs_detail.png

gpt-realtime-mini, before (c134fb7)

turn 1: input=436 text=119 audio=317 cached=0 cached_tokens_details={'text_tokens': 0, 'audio_tokens': 0, 'image_tokens': 0} output=164
turn 2: input=927 text=293 audio=634 cached=576 cached_tokens_details={'text_tokens': 256, 'audio_tokens': 320, 'image_tokens': 0} output=140
turn 3: input=1394 text=443 audio=951 cached=0 cached_tokens_details={'text_tokens': 0, 'audio_tokens': 0, 'image_tokens': 0} output=131
turn 4: input=1852 text=584 audio=1268 cached=1536 cached_tokens_details={'text_tokens': 576, 'audio_tokens': 960, 'image_tokens': 0} output=177
session totals: text=1439 audio=3170 cached=2112 (cached text=832, cached audio=1280) output text=612 audio=0
expected cost at the cost map rates for gpt-realtime-mini (text $0.6/M, audio $10/M, cached text $0.06/M, cached audio $0.3/M, output text $2.4/M): 0.0211669
GET /spend/logs (rows since this session started) -> model=openai/gpt-realtime-mini spend=0.0264388 prompt_tokens=4609 completion_tokens=612 request_id=656aef21-a4b0-46bb-afeb-da120b66330c

Logged spend is $0.0264388 against $0.0211669 at the cost map rates, 25% over. The old cost map entry had no cache-read rate at all, so all 2112 cached tokens bill at $0, while the split-less subtraction still charges 2497 audio tokens at $10/M where only 1890 were uncached: 2497 x $10/M + 612 output x $2.40/M = $0.0264388 exactly

pr40627-a94c060b84-shot_before7_mini_logs_detail.png

gpt-realtime-mini, after (a94c060)

turn 1: input=436 text=119 audio=317 cached=64 cached_tokens_details={'text_tokens': 64, 'audio_tokens': 0, 'image_tokens': 0} output=138
turn 2: input=901 text=267 audio=634 cached=576 cached_tokens_details={'text_tokens': 256, 'audio_tokens': 320, 'image_tokens': 0} output=108
turn 3: input=1336 text=385 audio=951 cached=1024 cached_tokens_details={'text_tokens': 384, 'audio_tokens': 640, 'image_tokens': 0} output=109
turn 4: input=1772 text=504 audio=1268 cached=1408 cached_tokens_details={'text_tokens': 448, 'audio_tokens': 960, 'image_tokens': 0} output=148
session totals: text=1275 audio=3170 cached=3072 (cached text=1152, cached audio=1920) output text=503 audio=0
expected cost at the cost map rates for gpt-realtime-mini (text $0.6/M, audio $10/M, cached text $0.06/M, cached audio $0.3/M, output text $2.4/M): 0.0144261
GET /spend/logs (rows since this session started) -> model=openai/gpt-realtime-mini spend=0.01442612 prompt_tokens=4445 completion_tokens=503 request_id=9bd46713-5be2-4da4-a31d-44b7918e29ab

Logged spend $0.01442612 is the cost map rate cost with the entry's new $0.06/M cache-read rate: 123 non-cached text at $0.60/M, 1250 non-cached audio at $10/M, 1152 cached text at $0.06/M, 1920 cached audio at $0.30/M, 503 output text at $2.40/M

pr40627-a94c060b84-shot_after8_mini_logs_detail.png

GET /model/info on both proxies (same config, same request) differs only in the touched keys: gpt-realtime-mini reports cache_read_input_token_cost 6e-08 and cache_read_input_audio_token_cost 3e-07 on the after proxy where the before proxy reported neither, and gpt-realtime-2 and gpt-realtime-2.1-mini gain cache_read_input_audio_token_cost (4e-07 and 3e-07); input, audio, and output rates are identical on every model

Observed alongside the fix:

  • OpenAI cache hits vary per run; leg totals differ
  • Some turns report cached=0 mid-session; provider behavior
  • gpt-realtime-mini after turn 1 reused an earlier session's cache
  • UI shows cached total only, no text/audio split; unchanged

Type

🐛 Bug Fix

Caveats (if any)

Medium

  • When a provider's nested cached_tokens_details add up to more than cached_tokens, parse_prompt_tokens_details caps them by clipping audio first, then text, then image. Nothing in the OpenAI usage contract picks that order and the PR does not argue for it; the alternatives are scaling every modality down proportionally, trusting the nested counts over the total, or dropping the split and billing all cached tokens at the text cache-read rate when the two disagree. No response in the QA runs hit this branch, so it only changes billing on a self-inconsistent usage payload; the order is an open decision for the author to confirm or change

Low

  • Cached image tokens bill at the generic cache-read rate; no image-specific cache key exists
  • Non-OpenAI realtime providers do not emit cached_tokens_details, so their billing is unchanged
  • A session where only some response.done events carry cached_tokens_details keeps the split from the events that did and bills the rest of the cached total at the text cache-read rate, which is the pre-fix billing for those tokens; OpenAI sent the split on every response.done in the QA legs (12 of 12) and test_realtime_combine_keeps_cached_split_when_only_one_usage_has_details pins both orders
  • 76ae35d moves CachedTokensDetails into litellm.types.llms.base to clear the two CodeQL cyclic-import alerts, adds the test above, and commits the regenerated schema.d.ts (two soft_budget docstring lines that drifted on main); it cannot change behavior, so the QA legs and the /live-pr-risk pass at a94c060 carry over
  • 6f7882d keeps the BaseLiteLLMOpenAIResponseObject import line in litellm/types/llms/openai.py byte-identical to main and pulls CachedTokensDetails in on its own relative import. CodeQL resolves from openai import Omit in that file to the module itself, so every importer of a name whose definition line is in the diff gets a py/unsafe-cyclic-import alert; the 76ae35d edit to that line produced two such alerts at unchanged files (litellm/types/google_genai/main.py:4 and litellm/proxy/hooks/parallel_request_limiter_v3.py:63), and main itself reports none. The commit cannot change behavior either
  • get_model_info() and GET /model/info now carry cache_read_input_audio_token_cost for every model, null where the cost map has no such rate
  • Every /v1/responses usage now carries input_tokens_details.cached_tokens_details, null when the provider reports no split; additive, the existing fields are unchanged
  • azure/gpt-realtime-2025-08-28 and azure/gpt-realtime-1.5-2026-02-23 cached text drops from $4/M to $0.40/M, so logged spend for cached prompts on those Azure deployments goes down; the old value equalled the full input rate and Azure's retail meter lists $0.40/M
  • The Azure realtime entries were not driven live (no Azure realtime deployment was available to this run); their fills are covered by the parametrized cost test and the retail meters above, and the live legs cover the OpenAI entries only
  • CircleCI ocr_testing is red at this tip because the job collects zero tests: .circleci/config.yml still names tests/test_litellm/ocr/test_rust_bridge.py, which refactor(ocr): complete native lifecycle and preserve Azure auth #40734 deleted and that commit is in this PR's merge base. The base branch's own scheduled pipeline fails the same way and a follow-up tracks the config fix; nothing in this PR touches OCR or the CI config. CircleCI upload-coverage runs only after every test job passes, so it stays not run behind that failure
  • GitHub Actions proxy-behavior (Postgres Tests) is red at this tip on tests/proxy_behavior/auth/test_auth_object_prefetch.py::test_join_binds_the_membership_to_the_requested_team, which awaits a MagicMock. main's own Postgres Tests runs at the merge base fail the same way (34737619378 and 34737299659); nothing in this PR touches auth or that test, and the check is not in main's required set

Final Attestation

  • 5737cab passes /live-pr-risk

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR


Note

Medium Risk
Changes billing and spend logging for realtime and any usage with cached_tokens_details; incorrect caps when nested details exceed cached_tokens could mis-price edge-case payloads.

Overview
Fixes overbilling on OpenAI/Azure realtime when prompt cache hits include audio: the provider’s cached_tokens_details split is now kept end-to-end and used in cost math instead of treating all cached tokens as text.

Adds CachedTokensDetails on prompt_tokens_details / input_tokens_details and threads it through Responses/realtime usage transforms, usage merging (_combine_prompt_tokens_details), and logging. parse_prompt_tokens_details derives capped cached text/audio/image counts and bills non-cached text/audio/image at full input rates. Cache-read cost splits: non-audio cached tokens at cache_read_input_token_cost, cached audio at cache_read_input_audio_token_cost (fallback to text cache-read). BilledTokenRates, get_model_info, and breakdown helpers expose the audio cache-read rate.

Updates model cost maps for gpt-realtime* / azure/gpt-realtime* (new or corrected cache_read_input_* keys, including fixing Azure entries that used the full input rate for cache read).

Reviewed by Cursor Bugbot for commit 6f7882d. Bugbot is set up for automated code reviews on this repo. Configure here.

Link to Devin session: https://app.devin.ai/sessions/02c812fd9fa940c4baa2543d662d6754
Open in Devin Desktop: https://app.devin.ai/desktop/session/02c812fd9fa940c4baa2543d662d6754?variant=devin
Requested by: @shivamrawat1

shivamrawat1 and others added 2 commits September 10, 2026 22:02
…rate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR. Add '(aside)' to your comment to have me ignore it.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@codspeed

codspeed Bot commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_fix_realtime_cached_audio_cost (6f7882d) with main (30f33a9)

Open in CodSpeed

@greptile-apps

greptile-apps Bot commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR preserves cached-token modality details throughout realtime usage processing and applies the appropriate text and audio cache-read rates during cost calculation.

  • Aggregates nested cached-token details across realtime events.
  • Separates cached audio from full-rate audio billing and cost breakdowns.
  • Exposes audio cache-read rates through model metadata.
  • Updates OpenAI and Azure realtime pricing entries and adds regression coverage.
  • The only change since the previous review adjusts an internal import without changing runtime behavior.

Confidence Score: 5/5

The PR appears safe to merge, with no outstanding previous findings or actionable regressions introduced since the previous review.

All previous threads are resolved, including the fully addressed cached-token capping, comment, and post-construction mutation findings; the parameter-mutation finding was correctly withdrawn. The only subsequent change switches one package import from absolute to relative syntax, which resolves to the same module and introduces no loading or cycle issue.

Important Files Changed

Filename Overview
litellm/cost_calculator.py Aggregates cached-token modality details when combining realtime usage.
litellm/litellm_core_utils/llm_cost_calc/utils.py Separates cached audio from uncached audio and bills it at the configured audio cache-read rate.
litellm/responses/litellm_completion_transformation/transformation.py Preserves nested cached-token details while transforming usage into Responses API structures.
litellm/types/llms/openai.py Defines realtime cached-token detail typing; the post-review import adjustment is behaviorally equivalent.
model_prices_and_context_window.json Adds and corrects cache-read pricing for affected OpenAI and Azure realtime models.

Reviews (15): Last reviewed commit: "fix(types): import CachedTokensDetails o..." | Re-trigger Greptile

Comment thread litellm/litellm_core_utils/llm_cost_calc/utils.py Outdated
Comment thread litellm/cost_calculator.py Outdated
@codecov

codecov Bot commented Sep 10, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

…itellm_fix_realtime_cached_audio_cost

# Conflicts:
#	litellm/responses/litellm_completion_transformation/transformation.py
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri mateo-berri added run-ci and removed run-ci labels Sep 12, 2026
Comment thread litellm/cost_calculator.py Outdated
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

…info

Every proxy and router cost lookup goes through get_model_info, which copies
cost map keys explicitly, so the new audio cache-read branch always fell back
to the text cache-read rate there. Copy the key so models whose audio
cache-read rate differs from the text one bill cached audio correctly.
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@devin-ai-integration
devin-ai-integration Bot changed the base branch from litellm_internal_staging to main September 13, 2026 04:24
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

Comment thread litellm/cost_calculator.py Fixed
Comment thread litellm/types/utils.py Fixed
@yuneng-berri
yuneng-berri deleted the branch main September 13, 2026 04:51

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/cost_calculator.py
@mateo-berri mateo-berri reopened this Sep 13, 2026
CodeQL flagged two module-level cyclic imports introduced by defining
CachedTokensDetails in litellm.types.llms.openai and importing it from
litellm.types.utils and litellm.cost_calculator. The class now lives in
litellm.types.llms.base, which imports nothing from litellm, and every
user imports it from there.

Also pins that combining realtime usages where only one response.done
carries cached_tokens_details keeps the earlier modality split in both
orders, and commits the regenerated dashboard API types.
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

CodeQL resolves `from openai import Omit` in litellm/types/llms/openai.py to the
module itself, so every importer of a name whose definition line is in the diff
is reported as an unsafe cyclic import. 76ae35d edited the line that defines
BaseLiteLLMOpenAIResponseObject there and got two alerts at files this PR does
not touch. That line is now byte-identical to main and CachedTokensDetails
arrives through a relative import isort keeps separate.
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor

cursor Bot commented Sep 13, 2026

Copy link
Copy Markdown
Contributor

Bugbot couldn't run

Something went wrong. Try again by commenting "Cursor review" or "bugbot run", or contact support (requestId: serverGenReqId_569613cf-ab77-4823-8948-1ce0bef81c39).

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 6f7882d. Configure here.

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@mateo-berri
mateo-berri merged commit 6e9e084 into main Sep 14, 2026
89 of 91 checks passed
@mateo-berri
mateo-berri deleted the litellm_fix_realtime_cached_audio_cost branch September 14, 2026 17:50
doonga pushed a commit to greyrock-labs/home-ops that referenced this pull request Sep 28, 2026
…103.0) (#290)

This PR contains the following updates:

| Package | Update | Change |
|---|---|---|
| [ghcr.io/berriai/litellm](https://images.chainguard.dev/directory/image/wolfi-base/overview) ([source](https://github.com/BerriAI/litellm)) | minor | `v1.102.1` → `v1.103.0` |

---

### Release Notes

<details>
<summary>BerriAI/litellm (ghcr.io/berriai/litellm)</summary>

### [`v1.103.0`](https://github.com/BerriAI/litellm/releases/tag/v1.103.0)

[Compare Source](https://github.com/BerriAI/litellm/compare/v1.102.1...v1.103.0)

#### Verify Docker Image Signature

All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0).

**Verify using the pinned commit hash (recommended):**

A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \
  ghcr.io/berriai/litellm:v1.103.0
```

**Verify using the release tag (convenience):**

Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/v1.103.0/cosign.pub \
  ghcr.io/berriai/litellm:v1.103.0
```

Expected output:

```
The following checks were performed on each of these signatures:
  - The cosign claims were validated
  - The signatures were verified against the specified public key
```

***

#### What's Changed

- fix(responses): translate the reasoning object into a chat-completion reasoning effort by [@&#8203;joshgarnett](https://github.com/joshgarnett) in [#&#8203;36363](https://github.com/BerriAI/litellm/pull/36363)
- fix(proxy): bound tool and guardrail index create\_many by the spend-log statement budgets by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40561](https://github.com/BerriAI/litellm/pull/40561)
- fix(mcp): require admission for delegated OAuth by [@&#8203;joshua-berri](https://github.com/joshua-berri) in [#&#8203;40923](https://github.com/BerriAI/litellm/pull/40923)
- fix(logging): log one bounded summary for a burst of timed-out LoggingWorker callbacks by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40912](https://github.com/BerriAI/litellm/pull/40912)
- fix(fireworks): resolve short model names to long cost map keys by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40929](https://github.com/BerriAI/litellm/pull/40929)
- ci: remove main branch source guard by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;40172](https://github.com/BerriAI/litellm/pull/40172)
- chore(ci): promote internal staging to main by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;40942](https://github.com/BerriAI/litellm/pull/40942)
- fix(spend\_logs): store litellm\_call\_id and match it in request\_id lookups by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;39068](https://github.com/BerriAI/litellm/pull/39068)
- fix(auth): refresh lite login session token grants from the live user and team rows by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;40657](https://github.com/BerriAI/litellm/pull/40657)
- feat(bedrock): support file delete and list for S3-backed managed files by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;39836](https://github.com/BerriAI/litellm/pull/39836)
- fix(proxy): gate the webhook test alert on proxy admins by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;40814](https://github.com/BerriAI/litellm/pull/40814)
- fix(ui): hide admin write-form tabs on the models page from view-only admins by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;38867](https://github.com/BerriAI/litellm/pull/38867)
- fix(anthropic-adapter): surface mid-stream provider errors as Anthropic error events by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33352](https://github.com/BerriAI/litellm/pull/33352)
- docs(github): add an Affected release section to the PR template by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;40618](https://github.com/BerriAI/litellm/pull/40618)
- docs(e2e): ban unit tests under tests/e2e by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;33852](https://github.com/BerriAI/litellm/pull/33852)
- fix(router): preserve Azure Entra ID params in reusable credentials by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40889](https://github.com/BerriAI/litellm/pull/40889)
- docs(user endpoints): remove unsupported soft\_budget param from user docstrings by [@&#8203;shivamrawat1](https://github.com/shivamrawat1) in [#&#8203;36585](https://github.com/BerriAI/litellm/pull/36585)
- feat(friendli): auto-sync Friendli model metadata into price registry by [@&#8203;Lee-Si-Yoon](https://github.com/Lee-Si-Yoon) in [#&#8203;35918](https://github.com/BerriAI/litellm/pull/35918)
- build(deps): bump smol-toml to 1.8.0 to clear GHSA-7w5x-hrqm-74c2 in osv-scan by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40478](https://github.com/BerriAI/litellm/pull/40478)
- fix(bedrock\_mantle): price GovCloud regions from the regional cost row and accept region-prefixed model names by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;39846](https://github.com/BerriAI/litellm/pull/39846)
- chore(ci): remerge internal staging by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;40943](https://github.com/BerriAI/litellm/pull/40943)
- test(auth): freeze the cache clock in auth prefetch tests by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40996](https://github.com/BerriAI/litellm/pull/40996)
- feat(jwt): allow virtual\_key\_claim\_field per issuer by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40927](https://github.com/BerriAI/litellm/pull/40927)
- fix(cost): bill cached realtime audio tokens at the audio cache-read rate by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40627](https://github.com/BerriAI/litellm/pull/40627)
- perf(logging): skip correlation contextvar stamping when request\_correlation\_in\_logs is off by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41054](https://github.com/BerriAI/litellm/pull/41054)
- feat(pricing): add azure gpt-chat-latest rates and drop retired friendliai llama-3.1 entries by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40976](https://github.com/BerriAI/litellm/pull/40976)
- build(deps): re-suppress GHSA-h7x2-h6g9-p789 in osv-scan on main, mlflow still has no fixed release by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41104](https://github.com/BerriAI/litellm/pull/41104)
- fix(otel): cap per-index OpenInference message attributes span-wide by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40562](https://github.com/BerriAI/litellm/pull/40562)
- feat(proxy): add general\_settings.allowed\_file\_extensions for /v1/files uploads by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41106](https://github.com/BerriAI/litellm/pull/41106)
- fix(proxy): forward provider request id headers on mapped error responses by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40925](https://github.com/BerriAI/litellm/pull/40925)
- fix(router): name the all-deployments-in-cooldown error on 429 responses by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40995](https://github.com/BerriAI/litellm/pull/40995)
- fix(ui): show the team alias on the model info page and in its raw JSON by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40992](https://github.com/BerriAI/litellm/pull/40992)
- feat(proxy): honor LITELLM\_DISABLE\_ACCESS\_LOG\_PATHS to drop noisy uvicorn access log lines by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41096](https://github.com/BerriAI/litellm/pull/41096)
- fix(prometheus): label pre-call rate limit failures with the resolved api\_provider by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41059](https://github.com/BerriAI/litellm/pull/41059)
- perf(proxy): serialize /model/info listing once with orjson by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41114](https://github.com/BerriAI/litellm/pull/41114)
- fix(utils): stop wrapper\_async submitting the sync success handler twice by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41115](https://github.com/BerriAI/litellm/pull/41115)
- fix(redis): log a timeout streak once per interval instead of one line per cache call by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40817](https://github.com/BerriAI/litellm/pull/40817)
- fix(router): record flat retry attempts and cap retries from attempted\_retries by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40930](https://github.com/BerriAI/litellm/pull/40930)
- refactor(prometheus): source PROXY\_LLM\_PROVIDER\_FALLBACK from litellm.constants by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41118](https://github.com/BerriAI/litellm/pull/41118)
- fix(proxy): hide default credentials login hint when UI\_PASSWORD is set by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41107](https://github.com/BerriAI/litellm/pull/41107)
- fix(cli): show routed models and session stats for LLM API keys by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41116](https://github.com/BerriAI/litellm/pull/41116)
- fix(bedrock/realtime): propagate deferred Nova Sonic stream failures to the router by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41064](https://github.com/BerriAI/litellm/pull/41064)
- fix(proxy): keep org admins' own team memberships in other orgs visible on team list by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41086](https://github.com/BerriAI/litellm/pull/41086)
- feat(model\_info): provider-scoped fill\_missing\_for\_providers backfill from fallback generalization rules by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41093](https://github.com/BerriAI/litellm/pull/41093)
- fix(auth): load team membership once per request and skip prisma on an L1 hit by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41102](https://github.com/BerriAI/litellm/pull/41102)
- refactor(harness): expand independent trace coverage by [@&#8203;yujonglee-berri](https://github.com/yujonglee-berri) in [#&#8203;41120](https://github.com/BerriAI/litellm/pull/41120)
- fix(proxy): release max\_parallel\_requests slot when a realtime session ends without LLM callbacks by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41113](https://github.com/BerriAI/litellm/pull/41113)
- fix(router): cool down team deployments on 429 when a sibling serves the same public model by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40991](https://github.com/BerriAI/litellm/pull/40991)
- fix(ui): move tags typed into key metadata JSON into the Tags field by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41023](https://github.com/BerriAI/litellm/pull/41023)
- fix(ui): let team admins grant a team all proxy models by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;40196](https://github.com/BerriAI/litellm/pull/40196)
- chore(lint): graduate 12 rules from the strict-gate ratchet by [@&#8203;HUAHAODIA](https://github.com/HUAHAODIA) in [#&#8203;41048](https://github.com/BerriAI/litellm/pull/41048)
- test: add dedicated CircleCI integration contract foundation by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41066](https://github.com/BerriAI/litellm/pull/41066)
- fix(utils): keep litellm params out of provider request bodies by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41018](https://github.com/BerriAI/litellm/pull/41018)
- fix(openai): keep extra\_headers out of the chat request body on the httpx handler path by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41141](https://github.com/BerriAI/litellm/pull/41141)
- test: cover persisted updates and warmed authorization policies by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41070](https://github.com/BerriAI/litellm/pull/41070)
- fix(proxy): resolve x-litellm-call-id from response metadata when routes omit call\_id by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41056](https://github.com/BerriAI/litellm/pull/41056)
- chore(prices): sync Vertex AI prices: 14 models by [@&#8203;berriai-litellm-provider-info-sync](https://github.com/berriai-litellm-provider-info-sync)\[bot] in [#&#8203;40955](https://github.com/BerriAI/litellm/pull/40955)
- ci(codeql): exclude noisy Python quality queries by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41142](https://github.com/BerriAI/litellm/pull/41142)
- feat(proxy): predict prompt-cache costs across deployments by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;40877](https://github.com/BerriAI/litellm/pull/40877)
- fix(prompt\_security): keep polling file sanitization through non-terminal statuses by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41131](https://github.com/BerriAI/litellm/pull/41131)
- fix(health): skip background health check DB writes when the latest-row read fails by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41145](https://github.com/BerriAI/litellm/pull/41145)
- feat(model\_armor): logging\_only mode scans completed streams after delivery by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40702](https://github.com/BerriAI/litellm/pull/40702)
- fix(bedrock guardrails): derive contextual grounding source and query from plain messages by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41132](https://github.com/BerriAI/litellm/pull/41132)
- fix(cli): drop enum.StrEnum so the CLI imports on Python 3.10 by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41046](https://github.com/BerriAI/litellm/pull/41046)
- fix(responses): route mid-stream error events through exception\_type so content\_policy\_fallbacks fire by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40988](https://github.com/BerriAI/litellm/pull/40988)
- fix(cost): bill gemini-embedding-2 per token and stop double charging audio by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41157](https://github.com/BerriAI/litellm/pull/41157)
- test: bind management E2E callers and isolate JWT actors by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;40892](https://github.com/BerriAI/litellm/pull/40892)
- fix(headroom): protect cache\_control-marked rows anywhere in history by [@&#8203;rad-p44](https://github.com/rad-p44) in [#&#8203;40315](https://github.com/BerriAI/litellm/pull/40315)
- test: add strict stateless provider replay identity by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41149](https://github.com/BerriAI/litellm/pull/41149)
- fix(ci): test checked-out model pricing in unit jobs by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41181](https://github.com/BerriAI/litellm/pull/41181)
- fix(guardrails): write per-message guardrail rewrites back onto Responses input items by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40939](https://github.com/BerriAI/litellm/pull/40939)
- fix(proxy): log the provider usage on deferred /v1/messages calls and price cache writes without a creation rate by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41172](https://github.com/BerriAI/litellm/pull/41172)
- fix(guardrails): record not\_run evaluation when scoping leaves nothing to scan by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;39050](https://github.com/BerriAI/litellm/pull/39050)
- fix(responses): preserve provider affinity by [@&#8203;AaronHowell](https://github.com/AaronHowell) in [#&#8203;40228](https://github.com/BerriAI/litellm/pull/40228)
- fix(sdk): keep body and proxy headers on BadRequestError mapped from a litellm\_proxy 400 by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40994](https://github.com/BerriAI/litellm/pull/40994)
- fix(responses): hoist Codex additional\_tools input items into the chat bridge tools by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40989](https://github.com/BerriAI/litellm/pull/40989)
- fix(router): honor team and key provider weights by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41072](https://github.com/BerriAI/litellm/pull/41072)
- test(e2e): verify streamed answers and tool continuation by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41194](https://github.com/BerriAI/litellm/pull/41194)
- fix(cli): label router costs and simplify the routed-model header by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41186](https://github.com/BerriAI/litellm/pull/41186)
- test(spend): reconcile concurrent requests and daily activity by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41188](https://github.com/BerriAI/litellm/pull/41188)
- fix(guardrails): scan the Anthropic top-level system prompt and tool\_use arguments by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40984](https://github.com/BerriAI/litellm/pull/40984)
- fix(router): count num\_retries\_per\_request across fallback hops by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41191](https://github.com/BerriAI/litellm/pull/41191)
- fix(bedrock): grant rerank, retrieve, agent, and agentcore actions in the web identity session policy by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41168](https://github.com/BerriAI/litellm/pull/41168)
- fix(vertex-live): bill Gemini Live sessions end to end (internal copy of [#&#8203;37075](https://github.com/BerriAI/litellm/issues/37075)) by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40915](https://github.com/BerriAI/litellm/pull/40915)
- fix(health): resolve litellm\_credential\_name in realtime health checks by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41173](https://github.com/BerriAI/litellm/pull/41173)
- feat(proxy): unified custom\_key\_policy hook for key generate, update and regenerate by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40921](https://github.com/BerriAI/litellm/pull/40921)
- fix(proxy): enforce custom\_key\_update policy on /key/regenerate by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40695](https://github.com/BerriAI/litellm/pull/40695)
- fix(router): preserve session model choice within each complexity tier by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41174](https://github.com/BerriAI/litellm/pull/41174)
- test(pricing): let synced GovCloud Bedrock rows cite the AWS price list by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41263](https://github.com/BerriAI/litellm/pull/41263)
- docs(github): ask for interactive coding-tool proof in the PR template by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41257](https://github.com/BerriAI/litellm/pull/41257)
- feat(proxy): add POST /management/v1/users/bulk for batched user and team membership creation by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41028](https://github.com/BerriAI/litellm/pull/41028)
- fix(credentials): answer 409 on a credential name collision, make Terraform adoption opt-in by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;40917](https://github.com/BerriAI/litellm/pull/40917)
- feat(proxy): add POST /management/v1/users/bulk\_delete and POST /management/v1/teams/{team\_id}/members/bulk\_delete by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41039](https://github.com/BerriAI/litellm/pull/41039)
- fix(proxy): list directly assigned team models in model access errors by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41256](https://github.com/BerriAI/litellm/pull/41256)
- feat(auto-router): allow opted-in team members to manage their routers by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41175](https://github.com/BerriAI/litellm/pull/41175)
- build(rust-bridge): add typed \_native stub and validate it with mypy.stubtest by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41180](https://github.com/BerriAI/litellm/pull/41180)
- feat(guardrails): add new upstream presidio pii entities including german set by [@&#8203;MvdB](https://github.com/MvdB) in [#&#8203;36775](https://github.com/BerriAI/litellm/pull/36775)
- fix(responses): filter bridged kwargs like the native Responses path by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41144](https://github.com/BerriAI/litellm/pull/41144)
- test(e2e): cover the reliability retry, cooldown, fallback, and routing-strategy cells by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;39857](https://github.com/BerriAI/litellm/pull/39857)
- fix(anthropic): add the per-turn-control beta when a message carries output\_config by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41189](https://github.com/BerriAI/litellm/pull/41189)
- fix(router): bind per-request routing\_strategy override selectors to the request's callbacks by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41178](https://github.com/BerriAI/litellm/pull/41178)
- feat(proxy): bind JWT claims to registered agents via agent\_id\_jwt\_field by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40904](https://github.com/BerriAI/litellm/pull/40904)
- fix(proxy): enforce organization budgets when max\_budget is 0 by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;41271](https://github.com/BerriAI/litellm/pull/41271)
- fix(alerting): send llm\_exceptions Slack alert for 5xx HTTPException and ProxyException by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41125](https://github.com/BerriAI/litellm/pull/41125)
- fix(headroom): protect the cached prefix through the last cache\_control breakpoint by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41161](https://github.com/BerriAI/litellm/pull/41161)
- fix(utils): cache custom HuggingFace tokenizers across /utils/token\_counter requests by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41216](https://github.com/BerriAI/litellm/pull/41216)
- fix(router): keep weighted routing when a deployment id equals a model\_name by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41156](https://github.com/BerriAI/litellm/pull/41156)
- feat(router): add capability classifier as Fuse foundation by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41270](https://github.com/BerriAI/litellm/pull/41270)
- fix(proxy): keep access-group raw SQL writes on the writer while writer\_unavailable is stale by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41283](https://github.com/BerriAI/litellm/pull/41283)
- fix(prometheus): count 401 auth failures in litellm\_proxy\_failed\_requests\_metric by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41170](https://github.com/BerriAI/litellm/pull/41170)
- test: drop tests that pin vendor facts and add the CLAUDE.md rule by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41269](https://github.com/BerriAI/litellm/pull/41269)
- fix(proxy): run the remaining inline token counts off the event loop by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40262](https://github.com/BerriAI/litellm/pull/40262)
- fix(proxy): log blocked streaming guardrail responses as failures, not success by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40191](https://github.com/BerriAI/litellm/pull/40191)
- feat(proxy): add tpd\_limit (tokens per day) for batch submissions by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40997](https://github.com/BerriAI/litellm/pull/40997)
- fix(proxy): reconcile budget reservation before enqueuing spend to the DB by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40310](https://github.com/BerriAI/litellm/pull/40310)
- fix(xai): stop sending web\_search\_options to xAI's retired Live Search path by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;38278](https://github.com/BerriAI/litellm/pull/38278)
- feat(terraform): add tpm\_limit, rpm\_limit, budget\_duration, allowed\_models to litellm\_team\_member\_add by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;38682](https://github.com/BerriAI/litellm/pull/38682)
- fix(rerank): bill Vertex search\_units from input records and give every rerank response a unique id by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35180](https://github.com/BerriAI/litellm/pull/35180)
- fix(router): stop counting caller-set timeout 408s toward deployment cooldown by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41230](https://github.com/BerriAI/litellm/pull/41230)
- feat(router): add Fuse V2 classifier after capability forecasting by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41272](https://github.com/BerriAI/litellm/pull/41272)
- fix(proxy): keep client User-Agent on auth failure spend logs by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41291](https://github.com/BerriAI/litellm/pull/41291)
- fix(proxy): reset budgets by decrementing pre-reset spend instead of zeroing rows by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41279](https://github.com/BerriAI/litellm/pull/41279)
- fix(xai): honor nested web\_search filters on the xAI Responses API by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;38268](https://github.com/BerriAI/litellm/pull/38268)
- fix(router): stop registering a caller-supplied credential as a router deployment by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;41289](https://github.com/BerriAI/litellm/pull/41289)
- fix(router): accept custom\_provider\_map providers before the first completion call by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41300](https://github.com/BerriAI/litellm/pull/41300)
- fix(proxy): return 400 instead of 500 for lone surrogate escapes in request body by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41297](https://github.com/BerriAI/litellm/pull/41297)
- fix(langsmith): keep events appended during an in-flight flush instead of clearing them by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41288](https://github.com/BerriAI/litellm/pull/41288)
- fix(logging): track spend for streams a deployment hook converted to non-streaming by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41171](https://github.com/BerriAI/litellm/pull/41171)
- fix(bedrock): sanitize client tool\_call ids to Bedrock toolUseId constraints by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40872](https://github.com/BerriAI/litellm/pull/40872)
- fix(passthrough): attribute Vertex passthrough successes to the resolved router deployment by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41307](https://github.com/BerriAI/litellm/pull/41307)
- feat(ui): persist Models table search, filters, sort and page in the URL by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41296](https://github.com/BerriAI/litellm/pull/41296)
- fix(jwt-auth): scope JWT key mappings by issuer to prevent cross-issuer collisions by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;41281](https://github.com/BerriAI/litellm/pull/41281)
- feat(openai): add openai\_system\_messages\_first to put system messages first for prompt caching by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41304](https://github.com/BerriAI/litellm/pull/41304)
- feat(ui): add custom request headers to the API Playground by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41309](https://github.com/BerriAI/litellm/pull/41309)
- feat(cli): sync Codex /model picker from proxy /v1/models in lite codex by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40476](https://github.com/BerriAI/litellm/pull/40476)
- chore: bump litellm-enterprise 0.1.67 -> 0.1.68, litellm-proxy-extras 0.4.97 -> 0.4.98, litellm 1.102.0 -> 1.103.0 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41321](https://github.com/BerriAI/litellm/pull/41321)
- feat: add aihubmix provider pricing entries by [@&#8203;IToSSc](https://github.com/IToSSc) in [#&#8203;41179](https://github.com/BerriAI/litellm/pull/41179)
- feat(auto-router): add per-model Fast mode toggle by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41282](https://github.com/BerriAI/litellm/pull/41282)
- fix(proxy): include litellm\_call\_id in LLM API exception logs by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41205](https://github.com/BerriAI/litellm/pull/41205)
- fix(proxy): keep yaml pass-through endpoints visible to auth after db overlay by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41303](https://github.com/BerriAI/litellm/pull/41303)
- fix(proxy): resolve router\_settings.model\_group\_alias before key/team model auth by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41308](https://github.com/BerriAI/litellm/pull/41308)
- fix(ui): block usage export and flag the range when a spend page fails by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;41294](https://github.com/BerriAI/litellm/pull/41294)
- fix(vertex\_ai): bill Gemini Omni Interactions usage and Veo sampleCount on passthrough by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41322](https://github.com/BerriAI/litellm/pull/41322)
- fix(proxy): honor LITELLM\_LOG for uvicorn and proxy extras loggers by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41306](https://github.com/BerriAI/litellm/pull/41306)
- fix(cost): price native Responses WebSocket turns at their returned service\_tier by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41318](https://github.com/BerriAI/litellm/pull/41318)
- fix(proxy): key model rpm/tpm override takes precedence over team model limit by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41302](https://github.com/BerriAI/litellm/pull/41302)
- fix(proxy): track per-member organization spend by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41255](https://github.com/BerriAI/litellm/pull/41255)
- feat(proxy): add /nvidia\_nim passthrough route for NIM object detection and OCR /v1/infer by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41316](https://github.com/BerriAI/litellm/pull/41316)
- feat(model\_info): add provider-neutral Gemini 2.5+ chat baseline fallback generalization by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41320](https://github.com/BerriAI/litellm/pull/41320)
- fix(spend): sum multi-round session duration in logs UI by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35388](https://github.com/BerriAI/litellm/pull/35388)
- feat(router): limit unlicensed Capability and Fuse v2 routers to one each by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41326](https://github.com/BerriAI/litellm/pull/41326)
- fix(guardrails): resolve caller identity from metadata buckets in custom code guardrail by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41126](https://github.com/BerriAI/litellm/pull/41126)
- fix(e2e): onboard dashboard users through invitations by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41319](https://github.com/BerriAI/litellm/pull/41319)
- feat(guardrails): add Microsoft Agent 365 MCP tool-call guardrail by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;38241](https://github.com/BerriAI/litellm/pull/38241)
- test: drop remaining tests that pin cost-map vendor facts by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41298](https://github.com/BerriAI/litellm/pull/41298)
- feat(ui): show average response time per model in usage model activity by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41313](https://github.com/BerriAI/litellm/pull/41313)
- fix(proxy): preserve Anthropic pricing modifiers in router savings by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41341](https://github.com/BerriAI/litellm/pull/41341)
- feat(guardrails): support pre\_call and during\_call modes for llm\_as\_a\_judge by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41128](https://github.com/BerriAI/litellm/pull/41128)
- fix(gemini): propagate the provider's modelVersion to the response model by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41338](https://github.com/BerriAI/litellm/pull/41338)
- fix(fireworks-ai): bill cache-write, reasoning and audio tokens via the shared cost calculator by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41339](https://github.com/BerriAI/litellm/pull/41339)
- feat(guardrails): singulr v2 API contract with logging\_only, pre\_mcp\_call and post\_mcp\_call by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;41329](https://github.com/BerriAI/litellm/pull/41329)
- ci(image-scan): ignore zlib CVE-2026-85091 until Wolfi ships the fix by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41353](https://github.com/BerriAI/litellm/pull/41353)
- feat(e2e): reuse exact provider responses for 24 hours by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41346](https://github.com/BerriAI/litellm/pull/41346)
- fix(xai): keep 'instructions' on the xAI Responses API so system messages survive web search by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;38254](https://github.com/BerriAI/litellm/pull/38254)
- feat(ui): configure capability and Fuse v2 classifiers by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41315](https://github.com/BerriAI/litellm/pull/41315)
- fix(anthropic): tolerate message\_delta events without usage when streaming by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41336](https://github.com/BerriAI/litellm/pull/41336)
- test(router): ignore deployment-selection logs in the fallback log assertion by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41358](https://github.com/BerriAI/litellm/pull/41358)
- test(proxy): assert budget resets decrement the cleared spend by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41359](https://github.com/BerriAI/litellm/pull/41359)
- fix(e2e): expect models filters to persist after reload by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41348](https://github.com/BerriAI/litellm/pull/41348)
- fix(e2e): record cookie-setting provider responses and keep prompt-caching tests live by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41366](https://github.com/BerriAI/litellm/pull/41366)
- fix(responses): recount tokens when a streamed response completes without usage by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41337](https://github.com/BerriAI/litellm/pull/41337)
- fix(ui): simplify Capability and Fuse advanced routing options by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;41371](https://github.com/BerriAI/litellm/pull/41371)
- fix(mcp): authorize JWT OAuth credential persistence by [@&#8203;joshua-berri](https://github.com/joshua-berri) in [#&#8203;41314](https://github.com/BerriAI/litellm/pull/41314)
- feat(router): stream shadow traffic and fan out silent\_model to multiple targets by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41368](https://github.com/BerriAI/litellm/pull/41368)
- perf(content\_filter): scan a bounded window per streamed chunk by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41407](https://github.com/BerriAI/litellm/pull/41407)
- fix(proxy): hide model allowlist from client-facing model access denied errors by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41310](https://github.com/BerriAI/litellm/pull/41310)
- feat(http): opt-in outbound HTTP/2 for httpx clients by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41268](https://github.com/BerriAI/litellm/pull/41268)
- refactor(rust): remove gateway, config, router, realtime, and Rust trace-parity instrumentation by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41432](https://github.com/BerriAI/litellm/pull/41432)
- fix(guardrails): don't add post\_call output scan for MCP-only Presidio modes by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40571](https://github.com/BerriAI/litellm/pull/40571)
- chore(prices): sync Azure, Azure AI, Gemini, OpenAI, Bedrock, Together AI, Fireworks and Vertex prices: 278 models, 59 new, 30 deprecated by [@&#8203;berriai-litellm-provider-info-sync](https://github.com/berriai-litellm-provider-info-sync)\[bot] in [#&#8203;41154](https://github.com/BerriAI/litellm/pull/41154)
- fix(rag): forward retrieval\_filter from retrieval\_config to vector store search by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;34427](https://github.com/BerriAI/litellm/pull/34427)
- refactor(rust): extract auth and cache crates by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41464](https://github.com/BerriAI/litellm/pull/41464)
- fix(proxy): default litellm\_trace\_id to the OTel server span trace id by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41386](https://github.com/BerriAI/litellm/pull/41386)
- chore(codeowners): add ryan and kerry as owners of the cost map by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41333](https://github.com/BerriAI/litellm/pull/41333)
- fix(responses): guard empty-choices chunks in the Responses API streaming bridge by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;34455](https://github.com/BerriAI/litellm/pull/34455)
- chore(prices): sync Google Gemini prices: 22 models by [@&#8203;berriai-litellm-provider-info-sync](https://github.com/berriai-litellm-provider-info-sync)\[bot] in [#&#8203;41457](https://github.com/BerriAI/litellm/pull/41457)
- fix(bedrock): forward userContext in Knowledge Base Retrieve requests by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41475](https://github.com/BerriAI/litellm/pull/41475)
- ci(rust): split rust jobs, use nextest and Swatinem/rust-cache by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41480](https://github.com/BerriAI/litellm/pull/41480)
- fix(fireworks\_ai): flatten dict-form reasoning\_effort to its effort string by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41335](https://github.com/BerriAI/litellm/pull/41335)
- fix(proxy): never forward the LiteLLM virtual key to Anthropic on the /anthropic passthrough by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41340](https://github.com/BerriAI/litellm/pull/41340)
- fix(proxy): rename AWS Secrets Manager secret when key alias changes by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41468](https://github.com/BerriAI/litellm/pull/41468)
- feat(otel): promote nested request metadata keys to litellm.metadata.\* span attributes by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41462](https://github.com/BerriAI/litellm/pull/41462)
- fix(proxy): sync AWS Secrets Manager on body-less key regenerate and key alias changes by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41458](https://github.com/BerriAI/litellm/pull/41458)
- fix(http\_handler): keep a handler alive while a response it issued is still reading by [@&#8203;max-sixty](https://github.com/max-sixty) in [#&#8203;34829](https://github.com/BerriAI/litellm/pull/34829)
- fix(bedrock): make prompt caching work on the Nova InvokeModel route by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41343](https://github.com/BerriAI/litellm/pull/41343)
- ci(migrations): flag defaulted ADD COLUMN on request-log tables by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41460](https://github.com/BerriAI/litellm/pull/41460)
- feat(prometheus): add customer (end\_user) budget gauges by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41472](https://github.com/BerriAI/litellm/pull/41472)
- fix(otel): drop None metric and event attributes before OTLP export by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36815](https://github.com/BerriAI/litellm/pull/36815)
- fix(anthropic): carry the served model from message\_start onto stream chunks by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41446](https://github.com/BerriAI/litellm/pull/41446)
- fix(models): rolling registry audit: Gemini latest aliases, Nova cache pricing, OpenRouter/Together sync, Mistral GLM 5.3, Azure snapshots, Grok caching by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41112](https://github.com/BerriAI/litellm/pull/41112)
- fix(router): count TPM/RPM usage before building rate-limit headers by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41474](https://github.com/BerriAI/litellm/pull/41474)
- feat(guardrails): release buffered stream chunks after each passing scan by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41425](https://github.com/BerriAI/litellm/pull/41425)
- fix!: re-check budget on router fallback targets by [@&#8203;runjivu](https://github.com/runjivu) in [#&#8203;41379](https://github.com/BerriAI/litellm/pull/41379)
- refactor(ocr): move file preparation from the python bridge into litellm-core by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41489](https://github.com/BerriAI/litellm/pull/41489)
- feat(s3): add s3\_log\_prompts\_only option to log prompts without responses by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41327](https://github.com/BerriAI/litellm/pull/41327)
- feat(team): team-level model\_max\_budget with key-level overrides by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41330](https://github.com/BerriAI/litellm/pull/41330)
- feat(keys): filter /key/list by active, expired, revoked or deleted status and serve deleted keys from /key/info by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41311](https://github.com/BerriAI/litellm/pull/41311)
- feat(proxy): expose lifetime total\_spend on virtual keys by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41403](https://github.com/BerriAI/litellm/pull/41403)
- fix(proxy): release completed max-parallel slots promptly by [@&#8203;elifozdamar](https://github.com/elifozdamar) in [#&#8203;40843](https://github.com/BerriAI/litellm/pull/40843)
- feat(ui): accept ssh clone urls when registering a skill by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35418](https://github.com/BerriAI/litellm/pull/35418)
- fix(prices): dedupe Nova cache\_read\_input\_token\_cost keys left by a text merge by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41496](https://github.com/BerriAI/litellm/pull/41496)
- fix(otel): propagate W3C trace context on HTTP and WebSocket passthrough by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40669](https://github.com/BerriAI/litellm/pull/40669)
- fix(proxy): remove duplicate user budget hook that 429'd zero-cost models by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41345](https://github.com/BerriAI/litellm/pull/41345)
- test(logging): pick this test's own records out of the shared log batch by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41487](https://github.com/BerriAI/litellm/pull/41487)
- test(together\_ai): move request-shape checks to the mapped file, drop the live ones by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41360](https://github.com/BerriAI/litellm/pull/41360)
- feat(ui): shared URL-state layer for tables and tabs by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;41331](https://github.com/BerriAI/litellm/pull/41331)
- feat(e2e): make the provider cache reusable across builds and mount Bedrock behind it by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41402](https://github.com/BerriAI/litellm/pull/41402)
- fix(mcp): fail closed on missing upstream credentials by [@&#8203;joshua-berri](https://github.com/joshua-berri) in [#&#8203;41364](https://github.com/BerriAI/litellm/pull/41364)
- feat(rust): scaffold Redis cache crate by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41501](https://github.com/BerriAI/litellm/pull/41501)
- fix(dashscope): forward reasoning\_effort to the provider by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37506](https://github.com/BerriAI/litellm/pull/37506)
- fix(proxy): carry litellm\_call\_id through endpoint specific error logs and failure responses by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41356](https://github.com/BerriAI/litellm/pull/41356)
- fix(proxy): retry rate-limit fallbacks from a pristine request snapshot by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40596](https://github.com/BerriAI/litellm/pull/40596)
- fix(gemini): map minimal thinking to low for Gemini 3.7 and 3.8 Flash by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41201](https://github.com/BerriAI/litellm/pull/41201)
- fix(proxy): stop forwarding LiteLLM credential headers on Bedrock agent-runtime passthrough by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41504](https://github.com/BerriAI/litellm/pull/41504)
- fix(streaming): estimate interrupted Anthropic stream usage from reasoning\_content by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41503](https://github.com/BerriAI/litellm/pull/41503)
- fix(azure\_ai): route Responses API to native /openai/v1/responses for Foundry Models by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33856](https://github.com/BerriAI/litellm/pull/33856)
- fix(proxy): show all model groups to proxy admins in /model\_group/info by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41094](https://github.com/BerriAI/litellm/pull/41094)
- feat(proxy): let proxy admins choose which team fields team admins may edit by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;39996](https://github.com/BerriAI/litellm/pull/39996)
- fix(bedrock): neutralize orphaned tool blocks instead of raising or injecting a dummy tool (internal copy of [#&#8203;31400](https://github.com/BerriAI/litellm/issues/31400)) by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41513](https://github.com/BerriAI/litellm/pull/41513)
- feat(ui): persist organizations and projects list, detail tab and key table state in the URL by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;41445](https://github.com/BerriAI/litellm/pull/41445)
- fix(bedrock\_mantle): accept and forward verbosity on gpt-5.x chat completions by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41509](https://github.com/BerriAI/litellm/pull/41509)
- test: cover database transactions and persisted accounting by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41073](https://github.com/BerriAI/litellm/pull/41073)
- test: provider wire contracts, streaming and recovery by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41075](https://github.com/BerriAI/litellm/pull/41075)
- fix(mcp): count admin static headers as api\_key credential slots by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41514](https://github.com/BerriAI/litellm/pull/41514)
- ci: auto-merge provider-info-sync PRs when CI, Greptile and Bugbot are clean by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41494](https://github.com/BerriAI/litellm/pull/41494)
- feat(rust): add standalone framing crate by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41500](https://github.com/BerriAI/litellm/pull/41500)
- fix(proxy): enforce tag budgets for tags added by guardrails by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40842](https://github.com/BerriAI/litellm/pull/40842)
- fix(utils): run post-call deployment hook on converted chat streams by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41495](https://github.com/BerriAI/litellm/pull/41495)
- fix(e2e): bind provider-cache recordings to the deployment's test, not the serving process by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41520](https://github.com/BerriAI/litellm/pull/41520)
- test: add extension and browser integration contracts by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41078](https://github.com/BerriAI/litellm/pull/41078)
- fix(logging): scan each log record once and collapse base64 payloads before the secret regex by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40934](https://github.com/BerriAI/litellm/pull/40934)
- fix(spend\_tracking): attribute router-rejected requests to the model group provider by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41507](https://github.com/BerriAI/litellm/pull/41507)
- feat(router): discover token limits for hosted OpenAI-compatible models by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41508](https://github.com/BerriAI/litellm/pull/41508)
- feat(proxy): let team admins edit rpm\_limit and max\_budget when enabled by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;41525](https://github.com/BerriAI/litellm/pull/41525)
- fix(otel): fit per-index OpenInference messages to the span's remaining attribute budget by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41498](https://github.com/BerriAI/litellm/pull/41498)
- test(aws): verify rotated secret value by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41524](https://github.com/BerriAI/litellm/pull/41524)
- test(e2e): read a deleted key back as deleted, not as a 404 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41551](https://github.com/BerriAI/litellm/pull/41551)
- fix(otel v2): map the caller's Langfuse user, session and tags onto the root and generation spans by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41140](https://github.com/BerriAI/litellm/pull/41140)
- test: fix seven tests left stale by [#&#8203;41311](https://github.com/BerriAI/litellm/issues/41311), [#&#8203;41337](https://github.com/BerriAI/litellm/issues/41337), [#&#8203;39996](https://github.com/BerriAI/litellm/issues/39996), [#&#8203;41310](https://github.com/BerriAI/litellm/issues/41310), [#&#8203;41289](https://github.com/BerriAI/litellm/issues/41289) and [#&#8203;41315](https://github.com/BerriAI/litellm/issues/41315) by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41527](https://github.com/BerriAI/litellm/pull/41527)
- test(budgets): cover management null handling by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41563](https://github.com/BerriAI/litellm/pull/41563)
- test(e2e): drop the auto-router select "opens below" spec by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41568](https://github.com/BerriAI/litellm/pull/41568)
- fix(guardrails): stream Prompt Security post\_call redactions in incremental\_diff mode by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41558](https://github.com/BerriAI/litellm/pull/41558)
- fix(guardrails): give post-call scans the scoped request conversation and tools by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41220](https://github.com/BerriAI/litellm/pull/41220)
- feat(openrouter): add stealth/union-alpha to the model cost map by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41576](https://github.com/BerriAI/litellm/pull/41576)
- test(management): cover project authorization lifecycle by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41573](https://github.com/BerriAI/litellm/pull/41573)
- feat(rust): map Anthropic Messages transformations by [@&#8203;yujonglee-berri](https://github.com/yujonglee-berri) in [#&#8203;41531](https://github.com/BerriAI/litellm/pull/41531)
- fix(e2e): clear the three standing errors in the scheduled Buildkite suite by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41616](https://github.com/BerriAI/litellm/pull/41616)
- refactor(rust\_bridge): declarative route catalog and shared runtime selection by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41479](https://github.com/BerriAI/litellm/pull/41479)
- fix(mock\_completion): keep the resolved provider so router custom pricing resolves for azure\_ai deployments by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41623](https://github.com/BerriAI/litellm/pull/41623)
- fix(mcp): restrict health discovery to virtual key grants by [@&#8203;joshua-berri](https://github.com/joshua-berri) in [#&#8203;41609](https://github.com/BerriAI/litellm/pull/41609)
- fix(mcp): preserve request-selected guardrails during tool execution by [@&#8203;joshua-berri](https://github.com/joshua-berri) in [#&#8203;41619](https://github.com/BerriAI/litellm/pull/41619)
- refactor(ocr): mirror Python provider layout and preserve tests by [@&#8203;yujonglee-berri](https://github.com/yujonglee-berri) in [#&#8203;41550](https://github.com/BerriAI/litellm/pull/41550)
- test(fireworks\_ai): stop pinning vision support on minimax-m3 by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41627](https://github.com/BerriAI/litellm/pull/41627)
- perf(spend\_tracking): index LiteLLM\_SpendLogs by (api\_key, startTime) by [@&#8203;etiennechabert](https://github.com/etiennechabert) in [#&#8203;37983](https://github.com/BerriAI/litellm/pull/37983)
- fix(proxy): reject non-string model with 400 and log its spend as unknown-model by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41633](https://github.com/BerriAI/litellm/pull/41633)
- test(together\_ai): stop pinning successor deprecation status by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41635](https://github.com/BerriAI/litellm/pull/41635)
- chore(prices): sync Together AI prices: 6 models, 6 deprecated \[sync failed: Google Gemini] by [@&#8203;berriai-litellm-provider-info-sync](https://github.com/berriai-litellm-provider-info-sync)\[bot] in [#&#8203;41570](https://github.com/BerriAI/litellm/pull/41570)
- fix(budgets): page end-user cache invalidation after a budget reset by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;41488](https://github.com/BerriAI/litellm/pull/41488)
- chore: bump litellm-proxy-extras 0.4.98 -> 0.4.99 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41659](https://github.com/BerriAI/litellm/pull/41659)
- fix(tests): resolve the integration support package without run.py's PYTHONPATH by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41373](https://github.com/BerriAI/litellm/pull/41373)
- fix(ui): keep untimed guardrail entries on the request lifecycle by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;41374](https://github.com/BerriAI/litellm/pull/41374)
- fix(anthropic-bridge): convert mid-conversation system turns to user turns on /v1/messages to chat completions by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41493](https://github.com/BerriAI/litellm/pull/41493)
- fix(bedrock): support aws-sdk-bedrock-runtime 0.10/0.11 in Bedrock Realtime by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41542](https://github.com/BerriAI/litellm/pull/41542)
- feat(cli): deprecate the litellm-proxy entrypoint in favour of lite by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41673](https://github.com/BerriAI/litellm/pull/41673)
- fix(scim): align pagination `count` validation with RFC 7644 by [@&#8203;zachbernstein-sdx](https://github.com/zachbernstein-sdx) in [#&#8203;41444](https://github.com/BerriAI/litellm/pull/41444)
- fix(bedrock): never emit Converse cachePoint for OpenAI-family models by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41419](https://github.com/BerriAI/litellm/pull/41419)
- fix(images): stop forwarding the raw image\[] and mask\[] form keys by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;39512](https://github.com/BerriAI/litellm/pull/39512)
- feat(management\_v1): bulk update team member budgets by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;41632](https://github.com/BerriAI/litellm/pull/41632)
- refactor(rust): extract provider translations by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41690](https://github.com/BerriAI/litellm/pull/41690)
- feat(cli): rename lite autoroute up/down to start/stop, keeping the old names as deprecated aliases by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41672](https://github.com/BerriAI/litellm/pull/41672)
- fix(responses): keep the addressed response id off bridged provider requests by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41689](https://github.com/BerriAI/litellm/pull/41689)
- fix(license): let a wildcard allowed\_features license grant the auto\_router feature by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41684](https://github.com/BerriAI/litellm/pull/41684)
- fix(ui): list every provider in the cache leakage by-model table by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;40875](https://github.com/BerriAI/litellm/pull/40875)
- fix(team): keep a forked member budget's reset window and audit bulk member budget writes by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;41686](https://github.com/BerriAI/litellm/pull/41686)
- feat(proxy): add TypeSafe AI Jev evaluate passthrough with registry-priced spend tracking by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41607](https://github.com/BerriAI/litellm/pull/41607)
- test(e2e): cover bedrock batch file upload and create in the us-gov-west-1 partition by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41536](https://github.com/BerriAI/litellm/pull/41536)
- feat(grafana): add all-metrics dashboard and fix stale dashboard\_v2 gauges by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41578](https://github.com/BerriAI/litellm/pull/41578)
- fix(cost): price Azure PTU spillover requests at standard token rates by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41569](https://github.com/BerriAI/litellm/pull/41569)
- build(deps): bump soupsieve to 2.9.2 to clear the osv-scan advisories by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41703](https://github.com/BerriAI/litellm/pull/41703)
- fix(fireworks\_ai): restore supports\_vision on minimax-m3 in the cost map by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41699](https://github.com/BerriAI/litellm/pull/41699)
- feat(policy\_engine): explicit priority for policy attachment execution order by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41571](https://github.com/BerriAI/litellm/pull/41571)
- feat(router): add TypeSafe Jev as a complexity router classifier by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;41615](https://github.com…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants