Skip to content

feat(router): make routing groups callable as virtual models and list them in /v1/models - #36519

Merged
tin-berri merged 7 commits into
litellm_internal_stagingfrom
litellm_callable_routing_groups
Aug 12, 2026
Merged

tin-berri merged 7 commits into
litellm_internal_stagingfrom
litellm_callable_routing_groups

Conversation

@tin-berri

@tin-berri tin-berri commented Aug 11, 2026

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Routing group names were never callable as models
  • The Create Group modal has promised exactly that since day one
  • Groups never appeared in /v1/models, so Claude Code and Codex discovery cannot surface them
  • A group over single-deployment model names did nothing at all

How it solves it:

  • model=<group_name> now routes across the union of member deployments
  • The group's own strategy picks among them
  • Group names appear in /v1/models and are grantable on keys and teams
  • Rides the model_group_alias machinery instead of a parallel mechanism

User Flow

Before, creating a routing group gave you a name you could not use.

  1. Open Router Settings > Routing Groups, create group claude-quality over claude-fast and claude-smart with cost-based routing
  2. GET /v1/models with your key: claude-quality is not in the list, so Claude Code's model picker (with gateway discovery on) never shows it
  3. POST /v1/chat/completions with "model": "claude-quality": 400 "Invalid model name passed in"
  4. The modal said "Use this name as the model in API calls", so step 3 reads like a gateway bug

After, the group behaves like the modal always claimed.

  1. Create the same group in Router Settings > Routing Groups
  2. GET /v1/models: claude-quality is listed, and Claude Code or Codex discovery shows it in the picker
  3. POST /v1/chat/completions (or /v1/messages, /v1/responses) with "model": "claude-quality": 200, and the response's model field says claude-quality
  4. Repeat the call: the gateway keeps picking the cheapest member deployment, and the spend page attributes rows to claude-quality with the actual backend model recorded per row

Relevant issues

Linear ticket

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • My PR passes all CI/CD checks (e.g., lint, format, unit tests)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Live proxy on port 4020 (proxy_cli.py --config qa_callgroups_config.yaml --detailed_debug --use_v2_migration_resolver, DB-backed), real Bedrock calls. Before captured at b0fac57 (unmodified base), after at e7905f8, identical commands. Config: two single-deployment model names (claude-fast on haiku, claude-smart on opus, per-deployment cost overrides since the lowest-cost selector prices bedrock/-prefixed names at its 10.0 default) and one group:

router_settings:
  routing_strategy: simple-shuffle
  routing_groups:
    - group_name: claude-quality
      models: [claude-fast, claude-smart]
      routing_strategy: cost-based-routing

Before (b0fac57):

$ curl -s localhost:4020/v1/models -H "Authorization: Bearer sk-qa..." | jq -r '[.data[].id]|sort'
["claude-fast", "claude-smart"]
$ curl -s -o /dev/null -w "%{http_code}" localhost:4020/v1/chat/completions ... -d '{"model":"claude-quality",...}'
400            # same 400 on /v1/messages and /v1/responses

After (e7905f8):

$ curl -s localhost:4020/v1/models -H "Authorization: Bearer sk-qa..." | jq -r '[.data[].id]|sort'
["claude-fast", "claude-quality", "claude-smart"]
$ for i in 1 2 3 4; do curl -s localhost:4020/v1/chat/completions ... -d '{"model":"claude-quality","messages":[{"role":"user","content":"say ok"}],"max_tokens":5}'; done
model=claude-quality, 200 x4
$ curl -s localhost:4020/v1/messages ... -d '{"model":"claude-quality",...}'     -> type=message
$ curl -s localhost:4020/v1/responses ... -d '{"model":"claude-quality",...}'    -> status=completed
$ grep "get_available_deployment for model: claude-quality" qa_callgroups.log | grep -o "'model': 'bedrock/[^']*'" | sort | uniq -c
   6 'model': 'bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0'   # cost-based always picked the cheap member across BOTH model names
$ grep "routing_group=" qa_callgroups.log | tail -1
routing_group=claude-quality model=claude-quality strategy=cost-based-routing

Authz is symmetric for a scoped key (/key/generate with {"models": ["claude-quality"]}):

$ curl -s localhost:4020/v1/models -H "Authorization: Bearer <scoped>" | jq -r '[.data[].id]'
["claude-quality"]                       # only the group is listed
$ ... -d '{"model":"claude-quality",...}'   -> 200
$ ... -d '{"model":"claude-fast",...}'      -> 403   # member not granted, member-direct denied

Spend rows attribute the group with the real backend recorded per row:

$ psql ... -c 'SELECT model, model_group, LEFT(model_id,12) FROM "LiteLLM_SpendLogs" WHERE model_group=$$claude-quality$$ LIMIT 3;'
bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0 | claude-quality | 95ac109effb5   (x3)

The Models + Endpoints admin table stays deployment-only (/v2/model/info returns claude-fast, claude-smart; no group row), and member-direct calls behave exactly as before.

UI proof, captured on the live rig above. Router Settings > Routing Groups showing the claude-quality group:

routing groups tab

The Create Group modal, whose "Use this name as the model in API calls" copy this PR makes accurate with no UI change:

create group modal

Models + Endpoints > All Models listing deployments only, with no claude-quality row:

all models table

For picker discovery, point any OpenAI-compatible client at the gateway (Claude Code with CLAUDE_CODE_ENABLE_GATEWAY_MODEL_DISCOVERY=1, or Codex): claude-quality appears via /v1/models as shown in the scoped-key listing above

Type

🆕 New Feature

Caveats (if any)

  • Claude Desktop's client-side validator only accepts Anthropic-shaped names, so name groups like claude-quality for Desktop pickability
  • least-busy, lowest-cost, lowest-latency, and usage-v1 selectors key state by the requested name, so group calls and member-direct calls keep separate stats (same as model_group_alias today)
  • Fallbacks and model_group_retry_policy are keyed by name; a group needs its own entries
  • Access groups do not contain routing groups; grant the group name itself on keys or teams
  • A group name colliding with a model_name or alias warns and stays shadowed instead of failing config load

Changes

Router.get_routing_group resolves a callable group (a real deployment model_name added later shadows a same-named group, and config-time collisions with model names or aliases are rejected at init). _common_checks_available_deployment consults it before the early-resolve step so wildcard or pattern routes cannot hijack a group call, and serves the union of member deployments while model stays the group name; every selector already keys its state by that string. _get_routing_context resolves the group's own strategy for group-name requests. Discovery reuses the model_group_alias materialization: get_model_list_from_routing_groups emits member deployments under the group name via _get_all_deployments' existing model_alias rewrite, which is what surfaces groups in get_model_names, /v1/models, get_model_group_usage, /model_group/info, and the blocked and unhealthy hiding. The proxy gates (route_request and the adapter pass-through endpoint) now share one Router-owned is_recognized_model predicate covering model names, deployment ids, aliases, and routing groups, so the next virtual-model kind cannot be forgotten at one gate; both gates shrank in the process. Cooldown handling keys on the failing request's model group (read through the dual metadata-bucket owner and alias-resolved), so a 429 on a group call moves traffic to the group's alternatives while direct calls to a single-deployment member keep their cooldown exemption. Deleting a deployment now also drops its name from model_names at the removal-repair owner, so a runtime shadow's deletion restores the group, and group rebuilds invalidate the model-group info cache

Authz needs no code: a group is granted by putting its name on a key or team, and listing and enforcement both read the same literal allowlists, so they agree by construction. Member expansion in either direction was rejected on purpose since any-member semantics leak access

Tests extend the mapped files: 12 engine cases in test_router_routing_groups.py (union dispatch, discovery, group strategy selection, collisions, runtime shadowing, alias-to-group composition, empty group, cooldown alternatives, member-direct invariance; 10 of 12 fail on the base commit), a request-gate case in test_route_llm_request.py, and two cooldown cases in test_cooldown_handlers.py

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai can you reply here with a score and why that score?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Aug 11, 2026

Copy link
Copy Markdown
Contributor

Score: 4/5

Here's the breakdown:


What's done well

  • Correct reuse of existing machinery. Routing group discovery is wired through get_model_list_from_routing_groups_get_all_deployments(model_alias=group.group_name), the exact same path model_group_alias uses. Nothing in /v1/models, get_model_names, or get_model_group_usage needed touching separately.
  • Precedence is right. Checking _get_routing_group_deployments before _try_early_resolve_deployments_for_model_not_in_names ensures wildcard/pattern routes cannot hijack a group call. Alias-to-group composition is tested and works because _model_from_alias is already resolved before the group check.
  • Zero DB reads in the critical path. get_routing_group and _get_routing_group_deployments are pure in-memory dict lookups.
  • Cooldown fix is correct. A single-deployment member of a multi-deployment group should not get the single-deployment exemption; deployment_has_routing_group_alternatives fixes this precisely.
  • Init-time collision detection. Rejecting group_name == model_name and group_name == alias_name at startup prevents silent shadowing.
  • Tests. 12 engine cases, 2 cooldown cases, 1 gate case, with 10/12 failing on the base commit — the tests verify the actual new behavior, not just the happy path.

What holds it back from 5/5

  1. _init_routing_groups(None) at line 607 is a hidden contract. The None call is purely to ensure self._routing_groups exists before later uses; a one-liner self._routing_groups: dict[str, RoutingGroup] = {} directly in __init__ would be clearer and safer. The current pattern relies on readers knowing that _init_routing_groups(None) is a no-op loop.

  2. DB-backed group loading path is not tested. When routing_groups is loaded from the database and _init_routing_groups is called after a hot-reload, the collision check runs against self.model_group_alias and known_model_names at that moment in time. There's no test that a DB reload with conflicting names raises rather than silently overwriting.

  3. get_model_list_from_routing_groups return type annotation (Sequence) doesn't match the actual usage pattern. .extend() works on any iterable so this is not a bug, but the annotation is imprecise — callers extend a mutable list with the result, yet the declared return type is the immutable-view Sequence.

  4. No proxy-layer integration test for /v1/models actually listing the group. All tests are at the router engine layer. The gate test confirms the routing group name passes route_request, but there's no test that GET /v1/models with a scoped key returns only the group name. Given the PR's central claim for Claude Code / Codex discovery, a proxy-layer test here would close the loop.

@greptile-apps

greptile-apps Bot commented Aug 11, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

Callable routing groups are added to request routing and model discovery, with group-specific strategy selection and revised cooldown handling.

  • Resolves group names to the union of member deployments.
  • Materializes group names through the existing model-alias discovery machinery.
  • Allows the proxy request gate to dispatch group-name requests.
  • Adjusts cooldown behavior when a member has alternatives elsewhere in its routing group.
  • Adds routing, discovery, collision, request-gate, and cooldown tests.

Confidence Score: 3/5

The PR should not merge until direct-member cooldown behavior and restoration after deleting a runtime group shadow are corrected.

Group membership now causes direct single-deployment requests to lose their cooldown exemption, and model_names-based shadowing can keep a callable group disabled after the shadowing deployment is deleted.

Files Needing Attention: litellm/router.py, litellm/router_utils/cooldown_handlers.py

Important Files Changed

Filename Overview
litellm/router.py Adds callable-group resolution, discovery materialization, strategy selection, and shadowing, but deleted runtime shadows can leave groups disabled.
litellm/router_utils/cooldown_handlers.py Disables the single-deployment cooldown exemption based on group-wide alternatives, which also changes direct-member request behavior.
litellm/proxy/route_llm_request.py Extends the existing model gate to accept configured routing-group names.
tests/test_litellm/router_strategy/test_router_routing_groups.py Adds broad engine coverage but does not cover deleting a runtime deployment that shadows a callable group.
tests/test_litellm/router_utils/test_cooldown_handlers.py Tests the new cooldown decision but not the resulting failure of subsequent direct-member requests.
tests/test_litellm/proxy/test_route_llm_request.py Verifies routing-group names pass the proxy request gate and dispatch through the router.

Reviews (1): Last reviewed commit: "feat(router): make routing groups callab..." | Re-trigger Greptile

Comment on lines 343 to +346
if model_group is not None and len(model_group) == 1:
is_single_deployment_model_group = True
is_single_deployment_model_group = not litellm_router_instance.deployment_has_routing_group_alternatives(
deployment_id=deployment
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Direct-member cooldown behavior breaks

When a direct request to a single-deployment member receives a 429 or crosses the failure threshold, routing-group alternatives disable its single-deployment exemption even though those alternatives are unavailable to the direct request, causing subsequent calls to that member to fail with no healthy deployment until cooldown expires.

Knowledge Base Used: Router: deployment selection, retries, fallbacks, and cooldowns

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 8cbadcc: cooldown now keys on the failing request's model group, so direct member calls keep the exemption

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 8cbadcc: cooldown now keys on the failing request's model group, so direct member calls keep the exemption

Comment thread litellm/router.py
Comment on lines +1099 to +1101
group: Final = self._routing_groups.get(model_name)
if group is None or model_name in self.model_names:
return None

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Deleted shadow remains active

When a runtime deployment whose model name matches a routing group is added and later deleted, its name remains in self.model_names, so get_routing_group continues suppressing the group and requests and discovery cannot use it until the model list is rebuilt.

Knowledge Base Used: Router: deployment selection, retries, fallbacks, and cooldowns

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 8cbadcc: model_names is now maintained on deletion at the removal-repair owner, healing the other stale consumers too

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 8cbadcc: model_names is now maintained on deletion at the removal-repair owner, healing the other stale consumers too

Comment thread litellm/router.py
@veria-ai

veria-ai Bot commented Aug 11, 2026

Copy link
Copy Markdown
Contributor

PR overview

All previously flagged issues have been addressed. No open security concerns remain on this pull request.

Security review

No open security issues remain on this pull request.

Fixed/addressed: 1 · PR risk: 0/10

@codecov

codecov Bot commented Aug 11, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 98.30508% with 1 line in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
...y/pass_through_endpoints/pass_through_endpoints.py 0.00% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

Comment thread litellm/proxy/route_llm_request.py Outdated
@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai can you reply here with a score and why that score?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Aug 11, 2026

Copy link
Copy Markdown
Contributor

Score: 4/5


What's strong:

  1. Correct use of existing machineryget_model_list_from_routing_groups reuses _get_all_deployments's model_alias rewrite identically to how get_model_list_from_model_alias works, so groups surface everywhere that aliases do: get_model_names, /v1/models, get_model_group_usage, blocked/unhealthy hiding. No parallel mechanism.

  2. Precedence is sound — routing group resolution runs inside _common_checks_available_deployment before _try_early_resolve_deployments_for_model_not_in_names, so a wildcard default_deployment or pattern route cannot hijack a group call. The documented order (specific deployment → model id → alias → routing group → model_name → team/pattern/default) is correct.

  3. Collision protection_init_routing_groups rejects group names that collide with existing model_name values or model_group_alias entries at init time, and get_routing_group applies the runtime-shadowing rule (a deployment added later wins) correctly.

  4. Cooldown logic is rightrouting_group_has_alternatives properly distinguishes the two cases: a 429 on a group call (where other members exist as alternatives → cool the member down) vs. a 429 on a direct member call (single-deployment exemption preserved). Passing requested_model_group from the failure callback through _set_cooldown_deployments_should_cooldown_deployment is the right hook.

  5. Tests are thorough — 15 new tests, 10/12 failing on the base commit. Covers union dispatch, discovery, group strategy selection, both collision types, runtime shadowing and restore, alias-to-group composition, empty group, cooldown with/without alternatives, and member-direct invariance.

  6. Live proof — the screenshots + psql output in the description confirm the end-to-end behavior including spend attribution.


What's worth a second look:

  1. _invalidate_model_group_info_cache() in the early init path_init_routing_groups(None) is now called at the very top of __init__ (before model_group_alias and many other attributes exist). _init_routing_groups now calls _invalidate_model_group_info_cache() unconditionally before the if not groups_input: return guard. If that method touches an attribute that hasn't been set yet, it'll raise AttributeError on every Router construction. Worth a quick check that the method is safe when called with a partially-initialized self (e.g., it just does self._model_group_info_cache = None).

  2. Cooldown path: indirect requested_model_group propagation — the failure callback reads (litellm_params.get("metadata") or {}).get("model_group") to recover the requested group name. This is correct if metadata is consistently populated for group calls. The unit tests for cooldown call _should_cooldown_deployment directly with the group name rather than exercising the full failure callback → _set_cooldown_deployments chain, so there's no test confirming the metadata key is always present for group calls.

  3. model_names.discard in _update_deployment_indices_after_removal — this is a real existing-bug fix (without it the shadow/restore test fails), but it's a side-change that touches deployment removal for all router users, not just routing-group users. Low risk, but worth noting.

None of these are blockers — (1) is likely a non-issue (the method probably just assigns None), (2) is a test coverage gap rather than a correctness gap, and (3) is a net improvement. The design is clean and the feature fills a genuine user-facing gap.

Comment thread litellm/router.py
Comment thread litellm/router.py Outdated
@codspeed

codspeed Bot commented Aug 11, 2026

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_callable_routing_groups (b81f318) with litellm_internal_staging (bea3187)

Open in CodSpeed

@tin-berri
tin-berri force-pushed the litellm_callable_routing_groups branch from 8cbadcc to 4bee953 Compare August 11, 2026 22:55
@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai can you reply here with a score and why that score?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Aug 11, 2026

Copy link
Copy Markdown
Contributor

Score: 3.5 / 5

Here's the breakdown:


What's well done

  • is_recognized_model() is the right abstraction — centralizing all virtual-model kinds into one predicate prevents future gate-miss bugs, and the pass-through / route-request refactors are correct and semantically equivalent to the original logic.
  • Reusing _get_all_deployments(model_alias=group.group_name) to materialize group rows mirrors how model_group_alias works, so discovery, /v1/models, get_model_group_usage, and blocked/unhealthy hiding all get groups for free.
  • Config-time collision detection is correct, and the runtime-shadowing precedence rule (real deployment wins, group restores after deletion) is well-specified and tested.
  • The cooldown fix rationale (routing_group_has_alternatives) is logically sound: a 429 on a member with siblings should trigger cooldown so the group selector can move on.
  • Test breadth is good — 15 new cases covering union dispatch, strategy selection, collisions, shadowing, alias composition, empty group, cooldown alternatives, member-direct invariance.

Concerns

1. Cooldown metadata gap (medium risk)
The fix only works if litellm_metadata["model_group"] is already set to the group name when the failure callback fires. test_failure_callback_reads_model_group_from_litellm_metadata_bucket proves the callback reads it correctly, but it sets litellm_metadata: {"model_group": "quality"} by hand. There's no test or diff showing that the real routing path (a live acompletion call via a group name) actually writes the group name into litellm_metadata["model_group"] — or that it was already there from the pre-existing model_group tracking infrastructure. If it isn't, routing_group_has_alternatives is never invoked and the cooldown exemption for single-member model-groups silently suppresses cooldowns even when the call came through a multi-member group.

2. Hot-path overhead
_get_routing_group_deployments is called on every _common_checks_available_deployment invocation, including all non-group requests. For a router with no routing groups this is two no-op dict lookups, but for a busy proxy the early-exit could be tightened to a not self._routing_groups guard.

3. _init_routing_groups(None) calls _invalidate_model_group_info_cache() very early in __init__
This is the first thing set on self after _zero_cost_cache. If _invalidate_model_group_info_cache touches any attribute not yet initialized at that point in __init__, this will AttributeError on every Router construction. The existing test suite passing suggests it's safe in practice, but the ordering is fragile and worth a comment or guard.

4. Missing integration test for alias → group cooldown path
test_routing_group_has_alternatives_resolves_aliases covers the alias-resolution branch of routing_group_has_alternatives in isolation, but there's no end-to-end test that a 429 on a group-aliased call actually cools down the deployment.


The design is clean and the test coverage is above average for a feature of this complexity, but the cooldown metadata gap is a real risk: the fix could silently be a no-op in production if model_group isn't written into litellm_metadata by the existing routing infrastructure for group calls.

Comment thread litellm/router.py
Comment thread litellm/router.py
Comment thread litellm/router.py
…trip member access_groups from group rows, prove cooldown wiring end to end
@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai can you reply here with a score and why that score?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Aug 11, 2026

Copy link
Copy Markdown
Contributor

Score: 4/5

This is a well-engineered, security-conscious feature with strong test coverage. Here's the breakdown:

What's done well:

  • is_recognized_model predicate shared across both proxy gates (route_request, pass_through_endpoints) — prevents future virtual-model kinds from being forgotten at one gate, and the unification is correct (covers model_names, model_id, alias, routing group)
  • _as_routing_group_row correctly strips access_groups from member deployments when emitted under the group name — the access leakage case is explicitly tested
  • Collision detection at init time (_init_routing_groups rejects group names that shadow existing model_names or aliases)
  • Runtime shadowing + deletion restore is covered by test (test_real_model_added_later_shadows_group) and the model_names.discard fix in _update_deployment_indices_after_removal closes the loop
  • Cache invalidation (_invalidate_model_group_info_cache, _invalidate_access_groups_cache) on every group rebuild
  • Cooldown path passes requested_model_group through the metadata bucket so a 429 on a group call cools the member while member-direct calls keep the single-deployment exemption — and the E2E tests cover all three cases

What keeps it from a 5:

  1. requested_model_group metadata dependency: deployment_callback_on_failure reads requested_model_group from litellm_metadata. If that key is absent (e.g., some non-standard call path that doesn't populate litellm_metadata correctly), requested_model_group is None and the single-deployment exemption silently applies even for group calls. The test mocks the metadata directly rather than exercising the full call stack, so this code path isn't end-to-end verified for correctness under missing metadata.

  2. _get_routing_group_deployments on the critical path: It's called for every _common_checks_available_deployment invocation, including non-group requests. The fast path (not self._routing_groups) is O(1), but when groups exist, every non-group request still pays two dict lookups (_routing_groups.get + model_name_to_deployment_indices check). Minor, but worth noting for high-throughput deployments with groups configured.

  3. get_model_list_from_routing_groups materializes member deployments for each group on every get_model_list call: The per-name filtering is efficient, but the all-groups path iterates and expands all member deployments. This is acceptable at current scale but could compound with large group counts.

Overall this is a solid, production-ready feature — the security and authz considerations are correct by construction, the test coverage is unusually thorough, and the proxy gate unification is a genuine improvement over the prior state.

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit b98a7c2. Configure here.

… cache owner and drop the redundant wiring test
@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai can you reply here with a score and why that score?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Aug 11, 2026

Copy link
Copy Markdown
Contributor

Score: 4 / 5


Why not 5:

1. Backwards-incompatible init-time error for existing configs
_init_routing_groups now raises at startup if a group_name matches an existing model_name. Any production config that already had a routing group with the same name as a deployment model (even though that group was previously un-callable) will now refuse to start. The rule against backwards-incompatible breaks without a flag applies; a warning with graceful degradation (or at least a documented migration step) would be safer for existing users.

2. Critical-path complexity in _common_checks_available_deployment
The routing group lookup now runs on every request before the early-resolve step. get_routing_group does two dict lookups and one model_name_to_deployment_indices lookup per call — cheap in isolation, but this function is on the hot path. There's no fast-exit when _routing_groups is empty (the if not self._routing_groups: return None in get_routing_group helps, but _get_routing_group_deployments wraps it without a top-level guard). Minor, but worth noting.

3. requested_model_group threading relies on metadata being set correctly
The cooldown fix reads get_litellm_metadata_from_kwargs(kwargs).get("model_group") inside deployment_callback_on_failure. If anything in the call chain doesn't set model_group in the metadata bucket (e.g., a non-standard call path or a future refactor), the cooldown silently falls back to the old behavior. There's no test that verifies the metadata is actually populated end-to-end for a group call before the cooldown handler reads it.


What's solid:

  • The is_recognized_model predicate unifying all proxy gates is a clean and correct simplification — future virtual-model kinds can't be forgotten at one gate
  • Access group stripping in _as_routing_group_row is correct and tested
  • Cache invalidation (_routing_group_rows = None in _invalidate_model_group_info_cache) is wired to all the right call sites
  • model_names.discard on deployment removal restoring the group after a shadow is deleted is a nice correctness detail
  • Test coverage is strong: 12 engine cases covering dispatch, discovery, collisions, shadowing, alias composition, empty groups, cooldown alternatives, and member-direct invariance

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 9ecc0e8. Configure here.

…t coverage for group helpers, faithful router doubles in a2a and cursor tests
@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai can you reply here with a score and why that score?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Aug 12, 2026

Copy link
Copy Markdown
Contributor

Score: 4/5


What this gets right:

  • Routing precedence is correct. _common_checks_available_deployment consults _get_routing_group_deployments before _try_early_resolve_deployments_for_model_not_in_names, so wildcard/pattern routes cannot hijack a group call. The chain (deployment → model id → alias → routing group → model_name → fallbacks) is well-ordered.

  • is_recognized_model unification is a real improvement. The previous gate in route_llm_request.py had a subtle operator-precedence bug (or ... and ... without parens) that this PR incidentally fixes while collapsing three separate checks into one predicate. Same cleanup in pass_through_endpoints.py and response_api_endpoints/endpoints.py. A new virtual-model kind now only needs one place to register.

  • Security is correctly handled. _as_routing_group_row strips access_groups from the member's model_info so holding a member's access group cannot unlock the whole group. The test test_group_rows_do_not_inherit_member_access_groups covers this.

  • Cooldown logic is sound. routing_group_has_alternatives correctly gates the single-deployment exemption: a 429 on a group call (which has alternatives) cools the member, while a direct call to a single-deployment member keeps its exemption. The alias-to-group path in routing_group_has_alternatives (_get_model_from_alias then get_routing_group) is also correct.

  • Cache invalidation is complete. _routing_group_rows is cleared in _invalidate_model_group_info_cache (called by set_model_list, add_deployment, _init_routing_groups) and in _init_routing_groups itself. The model_names.discard(model_name) fix in _update_deployment_indices_after_removal lets a shadow deployment's deletion restore the group.

  • Test coverage is substantive. 12 engine cases, including union dispatch, strategy selection, collision/shadowing, alias-to-group composition, empty group, cooldown alternatives, and member-direct invariance. The end-to-end cooldown tests use mock_response="litellm.RateLimitError" so they exercise the real failure path.


Where the point comes off:

Cooldown metadata dependency is fragile. In deployment_callback_on_failure:

requested_model_group=(get_litellm_metadata_from_kwargs(kwargs) or {}).get("model_group"),

_should_cooldown_deployment only honors routing_group_has_alternatives when requested_model_group is the group name. If model_group in metadata is absent or set to something else by the time the callback fires (e.g., an intermediate retry that rewrites the metadata bucket, or a fallback chain), the group call silently falls back to the single-deployment exemption and the member is not cooled down. The unit tests construct the metadata correctly, but there is no integration-level assertion that model_group is reliably the group name in the metadata bucket at callback time across retry/fallback paths. A mis-propagation here would be invisible until production traffic.


Minor notes (not scoring):

  • _init_routing_groups(None) is called in __init__ before set_model_list, which is fine — _routing_groups, _model_to_group, and _group_selectors are always reset at the top of the function and the early-return on not groups_input makes the pre-set_model_list call safe.
  • The warning-not-error on collision is a deliberate design call (the group still applies strategy to members, the name just isn't callable). Acceptable.
  • get_routing_group is O(1) and guarded by if not self._routing_groups: return None, so the added call in the hot path is negligible for users without routing groups.

Overall: solid, well-tested feature with a clean architecture. The one deduction is the silent-failure risk in the cooldown metadata propagation path.

Comment thread litellm/router.py
@tin-berri

Copy link
Copy Markdown
Contributor Author

On the cooldown metadata concern: the same callback stamp already drives least-busy and spend accounting for ordinary model groups, and absence degrades to the pre-existing exemption

@tin-berri

Copy link
Copy Markdown
Contributor Author

Fallbacks re-stamp the failing group before each attempt; 64ab436 adds a retry-path cooldown test alongside the existing group and alias e2e tests

@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai can you reply here with a score and why that score?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Aug 12, 2026

Copy link
Copy Markdown
Contributor

Score: 4/5


What's good:

  • Correct resolution precedence — specific deployment > model id > alias > routing group > model_name. _common_checks_available_deployment inserts the group check before early-resolve so wildcard/default-deployment routes can't hijack a group call. This is the right call.
  • is_recognized_model consolidation — three proxy gates (route_llm_request, pass_through_endpoints, response_api_endpoints) now share one predicate instead of duplicating alias+model_names+model_id checks. This is a real maintainability win.
  • Security posture on _as_routing_group_row — stripping access_groups from materialized group rows is correct; inheriting them would let a key holding a member access group call the whole group.
  • model_names.discard after deployment removal — fixes a latent bug where deleting a runtime-shadow deployment left the name in model_names, permanently blocking group restoration.
  • Test coverage — 12 engine cases + gate cases + cooldown cases. The end-to-end cooldown tests (_call_and_get_cooldowns with pinned choice) are particularly well-designed.

Concerns (not blocking, but worth noting):

  1. requested_model_group extraction is an indirect dependency. In deployment_callback_on_failure, requested_model_group is pulled from get_litellm_metadata_from_kwargs(kwargs).get("model_group"). Whether model_group in the metadata bucket holds the routing group name (vs. the underlying member's group) for group calls depends on the metadata-bucket ownership chain — which is not touched by this PR. The end-to-end tests confirm it works today, but it's a somewhat fragile extraction point if metadata tracking changes.

  2. _routing_group_rows cache has a soft race on first population. The all-groups path checks cached = self._routing_group_rows then assigns if None. Under concurrent asyncio, two coroutines can both see None and both materialize, with the second overwriting the first. This is idempotent (same result) but doubles work on cold start under load. A None-check + assign pattern is fine here given asyncio's cooperative scheduling, but it's worth noting.

  3. Per-name queries bypass the all-groups cache. get_model_list_from_routing_groups(model_name="quality") always re-materializes, never reads from or writes to _routing_group_rows. This is intentional (model_name-scoped queries return a subset) but it means /model_group/info or any per-name caller pays re-materialization cost every time.

  4. _init_routing_groups(None) at __init__ time clears caches that were just initialized. It's harmless (the caches are empty), but calling _invalidate_model_group_info_cache and _invalidate_access_groups_cache before any data is loaded is a no-op that adds confusion about ordering invariants.


Summary:

The feature closes a real UX gap that the UI has promised since day one, the implementation is additive and well-contained, access control semantics are correct, and the test suite is thorough. The concerns above are minor and none affect correctness in the happy path. Ready for maintainer review.

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 64ab436. Configure here.

@tin-berri

Copy link
Copy Markdown
Contributor Author
shot1_routing_groups_tab shot2_create_group_modal shot3_all_models_no_group_row UI proof screenshots: Routing Groups tab with claude-quality, the Create Group modal copy this PR makes accurate, and the All Models table showing deployments only

@tin-berri
tin-berri merged commit 06943b6 into litellm_internal_staging Aug 12, 2026
83 checks passed
@tin-berri
tin-berri deleted the litellm_callable_routing_groups branch August 12, 2026 01:41
doonga pushed a commit to greyrock-labs/home-ops that referenced this pull request Aug 24, 2026
…8.0) (#393)

This PR contains the following updates:

| Package | Update | Change |
|---|---|---|
| [ghcr.io/berriai/litellm](https://images.chainguard.dev/directory/image/wolfi-base/overview) ([source](https://github.com/BerriAI/litellm)) | minor | `v1.97.0` → `v1.98.0` |

---

### Release Notes

<details>
<summary>BerriAI/litellm (ghcr.io/berriai/litellm)</summary>

### [`v1.98.0`](https://github.com/BerriAI/litellm/releases/tag/v1.98.0)

[Compare Source](https://github.com/BerriAI/litellm/compare/v1.98.0...v1.98.0)

##### Verify Docker Image Signature

All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0).

**Verify using the pinned commit hash (recommended):**

A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \
  ghcr.io/berriai/litellm:v1.98.0
```

**Verify using the release tag (convenience):**

Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/v1.98.0/cosign.pub \
  ghcr.io/berriai/litellm:v1.98.0
```

Expected output:

```
The following checks were performed on each of these signatures:
  - The cosign claims were validated
  - The signatures were verified against the specified public key
```

***

##### What's Changed

- fix(bedrock): drop toolSpec.strict for Claude Sonnet 5 on Converse by [@&#8203;kr0k](https://github.com/kr0k) in [#&#8203;33196](https://github.com/BerriAI/litellm/pull/33196)
- fix(batches): attribute Vertex passthrough batch cost to key/team/tags by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;34456](https://github.com/BerriAI/litellm/pull/34456)
- docs: rewrite the CLAUDE.md comment rule with explicit exceptions by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36301](https://github.com/BerriAI/litellm/pull/36301)
- fix(proxy): scope file list pagination cursors to the caller by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36093](https://github.com/BerriAI/litellm/pull/36093)
- fix(proxy): skip prisma-dependent hooks when no database is attached by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36273](https://github.com/BerriAI/litellm/pull/36273)
- fix(proxy): report has\_more false on caller-scoped file list pages by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36326](https://github.com/BerriAI/litellm/pull/36326)
- fix(proxy): restore management\_v1 query-param validation under fastapi>=0.140.7 by [@&#8203;HuanQian571](https://github.com/HuanQian571) in [#&#8203;35773](https://github.com/BerriAI/litellm/pull/35773)
- fix(proxy): stop /{provider}/v1/files from capturing /openai\_passthrough by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36092](https://github.com/BerriAI/litellm/pull/36092)
- chore(typing): remove 914 basedpyright Any errors across 16 hotspot files by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36386](https://github.com/BerriAI/litellm/pull/36386)
- fix(router): keep batch fallbacks inside the model group that owns the file by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36181](https://github.com/BerriAI/litellm/pull/36181)
- feat(ptu): configure provisioned-throughput flat cost on a model deployment by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;35341](https://github.com/BerriAI/litellm/pull/35341)
- docs: clarify the CLAUDE.md comment exceptions are any-of by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36421](https://github.com/BerriAI/litellm/pull/36421)
- docs: replace the Changes PR template section with Caveats by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36423](https://github.com/BerriAI/litellm/pull/36423)
- fix(bedrock): enable native structured output for GLM 5 and DeepSeek V3.2 by [@&#8203;alexshtf](https://github.com/alexshtf) in [#&#8203;35669](https://github.com/BerriAI/litellm/pull/35669)
- feat(ptu): daily rollup writes per-model PTU flat cost by active hour by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;35343](https://github.com/BerriAI/litellm/pull/35343)
- feat(logging): add opt-in session\_id and trace\_id correlation to JSON log records via contextvars by [@&#8203;deepanshululla](https://github.com/deepanshululla) in [#&#8203;34418](https://github.com/BerriAI/litellm/pull/34418)
- feat(ptu): surface PTU flat cost on the daily activity read path by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;35391](https://github.com/BerriAI/litellm/pull/35391)
- feat(router): add per-deployment allowed\_fails\_policy and cooldown\_time override support by [@&#8203;deepanshululla](https://github.com/deepanshululla) in [#&#8203;34416](https://github.com/BerriAI/litellm/pull/34416)
- feat(ptu): add PTU inputs to the model form and flat cost to the Usage page by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;35393](https://github.com/BerriAI/litellm/pull/35393)
- fix(cost): price dict-shaped image input token details at the image rate by [@&#8203;vairodp](https://github.com/vairodp) in [#&#8203;33490](https://github.com/BerriAI/litellm/pull/33490)
- fix(model\_prices): refresh deprecation dates, correct xAI pricing and add missing provider models by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36403](https://github.com/BerriAI/litellm/pull/36403)
- feat(ptu): gate PTU flat-cost attribution behind an opt-in env var by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36138](https://github.com/BerriAI/litellm/pull/36138)
- ci: cache Prisma CLI and engine binaries, split test timeout from setup by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36417](https://github.com/BerriAI/litellm/pull/36417)
- feat(rate limiting): configurable estimated output tokens per key, team and model by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36143](https://github.com/BerriAI/litellm/pull/36143)
- fix(ui): hide admin-only Logs tabs from roles that cannot call their endpoints by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36333](https://github.com/BerriAI/litellm/pull/36333)
- test(proxy): guard management\_v1 against fastapi names removed in supported releases by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36336](https://github.com/BerriAI/litellm/pull/36336)
- fix(ui): gate policy and prompt lookups on an admin capability by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36335](https://github.com/BerriAI/litellm/pull/36335)
- build(deps): bump pypdf to 6.15.0 to clear osv-scan by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36350](https://github.com/BerriAI/litellm/pull/36350)
- fix(proxy): isolate guardrail load failures per row by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36432](https://github.com/BerriAI/litellm/pull/36432)
- fix(ui): gate organization and agent usage views behind capabilities by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36334](https://github.com/BerriAI/litellm/pull/36334)
- fix(reset\_budget\_job): atomic budget cascade with chunked reset scans by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36287](https://github.com/BerriAI/litellm/pull/36287)
- feat(proxy): add GET /v1/indexes to list vector store indexes by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36289](https://github.com/BerriAI/litellm/pull/36289)
- feat(ui): show vector store indexes on the Vector Stores page by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36306](https://github.com/BerriAI/litellm/pull/36306)
- fix(proxy): treat SAML as configured in UI SSO detection by [@&#8203;fancybear-dev](https://github.com/fancybear-dev) in [#&#8203;36196](https://github.com/BerriAI/litellm/pull/36196)
- fix(bedrock): reject Anthropic server-side web\_search tool with actionable error by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36473](https://github.com/BerriAI/litellm/pull/36473)
- fix(ui): open the classifier prompt editor above the edit auto-router form by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36438](https://github.com/BerriAI/litellm/pull/36438)
- fix(arize): trace MCP tool calls instead of crashing on CallToolResult by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36453](https://github.com/BerriAI/litellm/pull/36453)
- refactor(ui): make illegal DataTable prop combinations unrepresentable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36470](https://github.com/BerriAI/litellm/pull/36470)
- fix(ui): scope Virtual Keys and Logs team lists to the caller by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36472](https://github.com/BerriAI/litellm/pull/36472)
- fix(ui): gate the Old Usage page behind a proxy-admin capability by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36469](https://github.com/BerriAI/litellm/pull/36469)
- docs(terraform): describe the provider release as automatic by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36467](https://github.com/BerriAI/litellm/pull/36467)
- feat(proxy): add per-deployment keepalive\_seconds SSE heartbeat to prevent load-balancer timeout on long streams by [@&#8203;deepanshululla](https://github.com/deepanshululla) in [#&#8203;34423](https://github.com/BerriAI/litellm/pull/34423)
- fix(router): cool down failed fallback deployments and correct cooldown TTL after Redis backfill by [@&#8203;deepanshululla](https://github.com/deepanshululla) in [#&#8203;35104](https://github.com/BerriAI/litellm/pull/35104)
- perf(spend): write each daily spend batch in one upsert statement by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36448](https://github.com/BerriAI/litellm/pull/36448)
- fix(ui): gate four sidebar pages on the roles their endpoints allow by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36475](https://github.com/BerriAI/litellm/pull/36475)
- fix(ui): restore the Logs Deleted Teams tab for organization admins by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36478](https://github.com/BerriAI/litellm/pull/36478)
- fix(websearch): stop leaking interception control fields to providers by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36480](https://github.com/BerriAI/litellm/pull/36480)
- test(e2e): cover the Anthropic web\_search server tool on Bedrock by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36443](https://github.com/BerriAI/litellm/pull/36443)
- fix(router): warn when a deployment's credentials contradict its provider by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36486](https://github.com/BerriAI/litellm/pull/36486)
- fix: net prompt-caching savings against the cache-write premium by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36452](https://github.com/BerriAI/litellm/pull/36452)
- feat(ui): deployment affinity toggle for the auto-router by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36302](https://github.com/BerriAI/litellm/pull/36302)
- fix(bedrock): use deployment credentials for AWS requests by [@&#8203;daleselaji-dev](https://github.com/daleselaji-dev) in [#&#8203;36160](https://github.com/BerriAI/litellm/pull/36160)
- fix(anthropic): preserve midturn system corrections by [@&#8203;eugene-yao-zocdoc](https://github.com/eugene-yao-zocdoc) in [#&#8203;34290](https://github.com/BerriAI/litellm/pull/34290)
- fix(email): stop duplicate legacy invitation email and fix its onboarding link by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;36455](https://github.com/BerriAI/litellm/pull/36455)
- feat(ui): show models under each tier in routing benchmark chart by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36291](https://github.com/BerriAI/litellm/pull/36291)
- fix(proxy): inject streaming usage cost on openai passthrough streams by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36503](https://github.com/BerriAI/litellm/pull/36503)
- docs: require a user flow and live-proxy proof in bug reports by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36498](https://github.com/BerriAI/litellm/pull/36498)
- fix(proxy): add config\_updated\_at audit timestamp for virtual keys by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36488](https://github.com/BerriAI/litellm/pull/36488)
- docs: require a user flow and a stuck-at proof in feature requests by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36500](https://github.com/BerriAI/litellm/pull/36500)
- feat(router): add required-AND (&) tag prefix and allow\_fail\_open flag by [@&#8203;deepanshululla](https://github.com/deepanshululla) in [#&#8203;36193](https://github.com/BerriAI/litellm/pull/36193)
- feat(proxy): per-key prompt caching toggle via enable\_prompt\_caching by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36466](https://github.com/BerriAI/litellm/pull/36466)
- fix(bedrock): send tool-search beta header for Haiku 4.5 on Invoke /v1/messages by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36502](https://github.com/BerriAI/litellm/pull/36502)
- fix(bedrock): preserve adaptive thinking effort through the /v1/messages bridge by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36507](https://github.com/BerriAI/litellm/pull/36507)
- ci: retry transient network fetch failures in lint workflow by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36563](https://github.com/BerriAI/litellm/pull/36563)
- fix(ui): stub useIsOrgAdmin in UsageTab tests so useCan needs no QueryClient by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36565](https://github.com/BerriAI/litellm/pull/36565)
- fix(alerting): dedupe scheduled Slack spend reports across pods by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36489](https://github.com/BerriAI/litellm/pull/36489)
- chore(typing): clear 1.6k basedpyright Any errors across 56 files by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36543](https://github.com/BerriAI/litellm/pull/36543)
- fix(bedrock): add text block to converse user messages carrying documents by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36499](https://github.com/BerriAI/litellm/pull/36499)
- fix(deps): ship boto3 with the base SDK so bedrock works out of the box by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;36568](https://github.com/BerriAI/litellm/pull/36568)
- fix(model\_prices): add provider-announced deprecation dates for Bedrock, Mistral, Cohere and Gemini models by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36538](https://github.com/BerriAI/litellm/pull/36538)
- chore: bump litellm-enterprise 0.1.54 -> 0.1.55, litellm-proxy-extras 0.4.84 -> 0.4.85, litellm 1.97.0 -> 1.98.0 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36577](https://github.com/BerriAI/litellm/pull/36577)
- fix(bedrock\_guardrails): skip ApplyGuardrail when there is no content to scan by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36441](https://github.com/BerriAI/litellm/pull/36441)
- fix(e2e): assert on the gen-AI span that served the stream, not the span count by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36582](https://github.com/BerriAI/litellm/pull/36582)
- test(e2e): harden vendor API coverage by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;34557](https://github.com/BerriAI/litellm/pull/34557)
- test(e2e): add reproducers for passthrough and model budget gaps by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;34657](https://github.com/BerriAI/litellm/pull/34657)
- test(e2e): cover google-native generateContent framing and prometheus queue time by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;34650](https://github.com/BerriAI/litellm/pull/34650)
- chore(ci): promote internal staging to main by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36560](https://github.com/BerriAI/litellm/pull/36560)
- feat(router): make routing groups callable as virtual models and list them in /v1/models by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36519](https://github.com/BerriAI/litellm/pull/36519)
- fix(xai): bill web\_search from server\_side\_tool\_usage\_details by [@&#8203;geraint0923](https://github.com/geraint0923) in [#&#8203;30817](https://github.com/BerriAI/litellm/pull/30817)
- fix(responses): init completed\_response on bridge streaming iterator ([#&#8203;35411](https://github.com/BerriAI/litellm/issues/35411)) by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35413](https://github.com/BerriAI/litellm/pull/35413)
- fix(batches): attribute Anthropic passthrough batch cost to the creating key, team and tags by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36468](https://github.com/BerriAI/litellm/pull/36468)
- feat(dashscope): add latest Model Studio models to the cost map by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36496](https://github.com/BerriAI/litellm/pull/36496)
- fix(proxy): track streamed passthrough Responses cost by [@&#8203;william-xue](https://github.com/william-xue) in [#&#8203;36529](https://github.com/BerriAI/litellm/pull/36529)
- fix(model\_prices): advertise native structured output on every Bedrock DeepSeek V3.2 and GLM 5 id by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36597](https://github.com/BerriAI/litellm/pull/36597)
- test(bedrock): repoint live Claude tests off the retired Claude 3 Sonnet by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36600](https://github.com/BerriAI/litellm/pull/36600)
- fix(anthropic): preserve speed=fast in usage for /v1/messages and pass-through by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36447](https://github.com/BerriAI/litellm/pull/36447)
- fix(proxy): forward resolved provider and deployment pricing in /cost/estimate by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35880](https://github.com/BerriAI/litellm/pull/35880)
- feat(proxy): global SSE keepalive ping interval for OpenAI-shaped streaming routes by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36154](https://github.com/BerriAI/litellm/pull/36154)
- fix(responses): preserve Codex namespace tool calls by [@&#8203;dcadenas](https://github.com/dcadenas) in [#&#8203;32536](https://github.com/BerriAI/litellm/pull/32536)
- fix(nvidia\_nim): preserve image passages and stop sending top\_k to /v1/ranking by [@&#8203;atomic](https://github.com/atomic) in [#&#8203;34177](https://github.com/BerriAI/litellm/pull/34177)
- fix: refactor HTTP handler initialization with client support by [@&#8203;Praveen11558](https://github.com/Praveen11558) in [#&#8203;30952](https://github.com/BerriAI/litellm/pull/30952)
- feat(lint): gate writable TypedDict fields with LIT012 by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36590](https://github.com/BerriAI/litellm/pull/36590)
- perf(proxy): stagger scheduled background jobs across jobs and pods by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36589](https://github.com/BerriAI/litellm/pull/36589)
- test: remove four mirror test files that exercise none of their module by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;34635](https://github.com/BerriAI/litellm/pull/34635)
- fix(router): stop re-applying router-selecting request tags to the routed tier's deployments by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36628](https://github.com/BerriAI/litellm/pull/36628)
- test: remove tests that never execute by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36681](https://github.com/BerriAI/litellm/pull/36681)
- fix(ui): align spend and budget columns by [@&#8203;daniel-meismer-zocdoc](https://github.com/daniel-meismer-zocdoc) in [#&#8203;35176](https://github.com/BerriAI/litellm/pull/35176)
- test: rename tests that a later definition shadowed by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36685](https://github.com/BerriAI/litellm/pull/36685)
- fix(passthrough): carry the budget reservation into request metadata by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36592](https://github.com/BerriAI/litellm/pull/36592)
- fix(mcp): bound MCP client requests with a session read timeout by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36675](https://github.com/BerriAI/litellm/pull/36675)
- fix(proxy): log requests rejected for an unparsable body in spend logs by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36673](https://github.com/BerriAI/litellm/pull/36673)
- refactor(ui): migrate cost-optimization to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36629](https://github.com/BerriAI/litellm/pull/36629)
- refactor(ui): migrate cost-tracking to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36631](https://github.com/BerriAI/litellm/pull/36631)
- refactor(ui): migrate admin-panel to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36635](https://github.com/BerriAI/litellm/pull/36635)
- refactor(ui): migrate users dashboard to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36642](https://github.com/BerriAI/litellm/pull/36642)
- refactor(ui): migrate prompts to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36643](https://github.com/BerriAI/litellm/pull/36643)
- refactor(ui): migrate team settings to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36641](https://github.com/BerriAI/litellm/pull/36641)
- refactor(ui): migrate models-and-endpoints to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36648](https://github.com/BerriAI/litellm/pull/36648)
- refactor(ui): migrate policy impact popover to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36653](https://github.com/BerriAI/litellm/pull/36653)
- fix(proxy): expand config-defined model access groups when resolving team models for /v2/model/info by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;34211](https://github.com/BerriAI/litellm/pull/34211)
- fix(batches): strip NUL bytes from passthrough batch tags before the managed object write by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36688](https://github.com/BerriAI/litellm/pull/36688)
- test(e2e-ui): verify UI mutations against the API instead of trusting the toast by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36632](https://github.com/BerriAI/litellm/pull/36632)
- fix(proxy): serialize model reconciles so concurrent model writes stop evicting each other by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36687](https://github.com/BerriAI/litellm/pull/36687)
- chore(e2e): port the compat-matrix cron publisher to tests/e2e/claude\_code by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36465](https://github.com/BerriAI/litellm/pull/36465)
- fix(router): never price a strategy-router alias by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36691](https://github.com/BerriAI/litellm/pull/36691)
- feat(model\_prices): add NVIDIA Nemotron 3.5 Lightning on OpenRouter and DeepInfra by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36696](https://github.com/BerriAI/litellm/pull/36696)
- feat(terraform/aws): make VPC, Aurora, and Redis optional by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36676](https://github.com/BerriAI/litellm/pull/36676)
- feat(ui): warn in the Admin UI when no Redis is configured by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36495](https://github.com/BerriAI/litellm/pull/36495)
- fix(ui): show and edit key-level router settings on a virtual key by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36674](https://github.com/BerriAI/litellm/pull/36674)
- fix(router): forward auto-router alias params from the marker entry, not the first same-name deployment by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36626](https://github.com/BerriAI/litellm/pull/36626)
- fix(bedrock\_mantle): 1M context window and long-context pricing for GPT-5.6 Sol/Terra/Luna by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36698](https://github.com/BerriAI/litellm/pull/36698)
- fix(model\_prices): sync the Groq registry with Groq's docs by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36664](https://github.com/BerriAI/litellm/pull/36664)
- fix(router): let untagged requests bypass a tagged pre-routing strategy on shared model names by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36627](https://github.com/BerriAI/litellm/pull/36627)
- fix(spend): stop losing spend log rows when a flush is cancelled by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;34826](https://github.com/BerriAI/litellm/pull/34826)
- docs(claude): drop the @&#8203; prefix from the PR template path by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36726](https://github.com/BerriAI/litellm/pull/36726)
- fix(langfuse): emit otel trace version and release on the keys langfuse v4 reads by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36702](https://github.com/BerriAI/litellm/pull/36702)
- test(interactions): follow Google spec drift replacing Turn with typed steps by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36730](https://github.com/BerriAI/litellm/pull/36730)
- refactor(ui): migrate team detail controls to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36695](https://github.com/BerriAI/litellm/pull/36695)
- refactor(ui): migrate guardrail and duration controls to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36693](https://github.com/BerriAI/litellm/pull/36693)
- refactor(ui): migrate guardrails-monitor, projects, logs to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;34606](https://github.com/BerriAI/litellm/pull/34606)
- refactor(ui): migrate search and user controls to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36694](https://github.com/BerriAI/litellm/pull/36694)
- fix(guardrails): scan and re-emit raw Anthropic SSE streams in the bedrock post-call hook by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36598](https://github.com/BerriAI/litellm/pull/36598)
- fix(helm): render nodeSelector on the migrations job by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36747](https://github.com/BerriAI/litellm/pull/36747)
- fix(langfuse): coerce header-sourced mask and trace-update steering values by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36740](https://github.com/BerriAI/litellm/pull/36740)
- refactor(ui): migrate usage tables to shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36707](https://github.com/BerriAI/litellm/pull/36707)
- refactor(ui): migrate guardrails monitor table to shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36709](https://github.com/BerriAI/litellm/pull/36709)
- refactor(ui): migrate guardrails content tables to shared DataTable by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36708](https://github.com/BerriAI/litellm/pull/36708)
- feat(gemini): day-0 pricing for gemini-3.7-flash by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36792](https://github.com/BerriAI/litellm/pull/36792)
- ci: promote staging to main by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36725](https://github.com/BerriAI/litellm/pull/36725)
- build(deps): bump nanoid to 3.3.18 to clear osv-scan by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36787](https://github.com/BerriAI/litellm/pull/36787)
- fix(router): stop scoring system prompt text for code/technical complexity by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36721](https://github.com/BerriAI/litellm/pull/36721)
- feat(complexity\_router): calibrate the classifier rubric with worked examples, selectable per router by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36578](https://github.com/BerriAI/litellm/pull/36578)
- fix(interactions): map step and turn history to Responses API roles and content types by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36733](https://github.com/BerriAI/litellm/pull/36733)
- fix(ui): restore playground model filtering by endpoint by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;36130](https://github.com/BerriAI/litellm/pull/36130)
- fix(proxy/batches): stop forwarding custom\_llm\_provider twice in list and cancel by [@&#8203;anxkhn](https://github.com/anxkhn) in [#&#8203;32813](https://github.com/BerriAI/litellm/pull/32813)
- refactor(ui): migrate TokenFlow and JsonViewer to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36735](https://github.com/BerriAI/litellm/pull/36735)
- feat: pre-adoption shadow eval for the auto-router (blind pairwise judge, derived state) by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36587](https://github.com/BerriAI/litellm/pull/36587)
- refactor(ui): migrate SimpleMessageBlock and SimpleToolCallBlock to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36737](https://github.com/BerriAI/litellm/pull/36737)
- refactor(ui): migrate HistoryTree and CollapsibleMessage to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36738](https://github.com/BerriAI/litellm/pull/36738)
- refactor: replace Any with precise types across responses, proxy, and llms modules by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36763](https://github.com/BerriAI/litellm/pull/36763)
- refactor(ui): migrate TruncatedValue and OutputCard to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36739](https://github.com/BerriAI/litellm/pull/36739)
- refactor(ui): migrate SectionHeader and ToolsSection to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36793](https://github.com/BerriAI/litellm/pull/36793)
- feat(ui): migrate playground chat controls to shadcn by [@&#8203;mubashir1osmani](https://github.com/mubashir1osmani) in [#&#8203;36129](https://github.com/BerriAI/litellm/pull/36129)
- feat(xai): day-0 pricing for grok-4.6 by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36805](https://github.com/BerriAI/litellm/pull/36805)
- feat(ui): highlight Auto Router in the navbar announcement by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36315](https://github.com/BerriAI/litellm/pull/36315)
- test(e2e): assert the model allow-list permits, not only denies by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36823](https://github.com/BerriAI/litellm/pull/36823)
- fix(proxy): tolerate a concurrent creator when creating spend views by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36824](https://github.com/BerriAI/litellm/pull/36824)
- fix(proxy): honor explicit null budget\_duration on team and key create + clearable UI dropdowns by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;36699](https://github.com/BerriAI/litellm/pull/36699)
- feat(model\_prices): add meta/muse-spark-1.2 and its contributor tier by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36717](https://github.com/BerriAI/litellm/pull/36717)
- fix(auth): carry team grants in lite login session tokens by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36826](https://github.com/BerriAI/litellm/pull/36826)
- feat(ui): show provider prompt cache tokens in chat response metrics by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36827](https://github.com/BerriAI/litellm/pull/36827)
- fix(auth): stop the team fallback from widening model access by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36837](https://github.com/BerriAI/litellm/pull/36837)
- fix(proxy/team): resolve member\_delete cleanup by user id, not the addressed email by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36839](https://github.com/BerriAI/litellm/pull/36839)
- fix(cli): launch agents as a child process on Windows by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36822](https://github.com/BerriAI/litellm/pull/36822)
- feat(ui): shadow evals tab beside auto-router usage by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36588](https://github.com/BerriAI/litellm/pull/36588)
- feat(cli): make the hidden `lite` command list configurable by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36816](https://github.com/BerriAI/litellm/pull/36816)
- feat(azure\_ai): add Fireworks FW model pricing on Azure AI Foundry by [@&#8203;emerzon](https://github.com/emerzon) in [#&#8203;35613](https://github.com/BerriAI/litellm/pull/35613)
- fix: enable xhigh reasoning support for gpt-5.4-mini models by [@&#8203;emerzon](https://github.com/emerzon) in [#&#8203;26909](https://github.com/BerriAI/litellm/pull/26909)
- feat(azure-ai): add Grok 4.3 model metadata by [@&#8203;emerzon](https://github.com/emerzon) in [#&#8203;27932](https://github.com/BerriAI/litellm/pull/27932)
- feat(ui): render request metrics on the /ui/chat surface by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36845](https://github.com/BerriAI/litellm/pull/36845)
- fix(ui): stop a deselected MCP server keeping its grant on a virtual key by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36840](https://github.com/BerriAI/litellm/pull/36840)
- fix(team): sweep dangling team references and cache on team delete by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36819](https://github.com/BerriAI/litellm/pull/36819)
- fix(mcp): resolve admin OAuth sessions from any worker via DB-backed drafts by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36844](https://github.com/BerriAI/litellm/pull/36844)
- refactor(ui): migrate usage to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36834](https://github.com/BerriAI/litellm/pull/36834)
- refactor(ui): migrate guardrails-monitor to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36838](https://github.com/BerriAI/litellm/pull/36838)
- refactor(ui): migrate playground to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36847](https://github.com/BerriAI/litellm/pull/36847)
- refactor(ui): migrate guardrails to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36832](https://github.com/BerriAI/litellm/pull/36832)
- fix(batches): stop uncostable batches from starving the cost poll page by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36714](https://github.com/BerriAI/litellm/pull/36714)
- perf(spend-logs): bound retention cleanup so one run cannot saturate the database by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36594](https://github.com/BerriAI/litellm/pull/36594)
- fix(proxy): fail config load when a callbacks entry is not dispatchable by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36858](https://github.com/BerriAI/litellm/pull/36858)
- fix(bedrock): hoist custom.defer\_loading before dropping custom on invoke tools by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36855](https://github.com/BerriAI/litellm/pull/36855)
- fix(access groups): sync assigned\_key\_ids from the key write paths by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36843](https://github.com/BerriAI/litellm/pull/36843)
- fix(mcp): expose client HTTP headers to logging callbacks and hooks by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36724](https://github.com/BerriAI/litellm/pull/36724)
- fix(ptu): stop per-token billing on a PTU-configured deployment by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36829](https://github.com/BerriAI/litellm/pull/36829)
- fix(ui): add nvidia riva to the model provider list by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36769](https://github.com/BerriAI/litellm/pull/36769)
- fix(scripts): end make check with a ran/skipped summary and verdict by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36864](https://github.com/BerriAI/litellm/pull/36864)
- fix(proxy): track spend for OpenAI passthrough /v1/embeddings by [@&#8203;lostmartian](https://github.com/lostmartian) in [#&#8203;36660](https://github.com/BerriAI/litellm/pull/36660)
- test(proxy): stop monkeypatch.undo re-planting fixture-mocked prisma\_client by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36872](https://github.com/BerriAI/litellm/pull/36872)
- fix(access groups): sync assigned\_team\_ids from the team write paths by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36825](https://github.com/BerriAI/litellm/pull/36825)
- ci: drop the CircleCI ui\_build and ui\_unit\_tests jobs by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36893](https://github.com/BerriAI/litellm/pull/36893)
- fix(langfuse)!: source the emitted metadata blob from StandardLoggingPayload by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36744](https://github.com/BerriAI/litellm/pull/36744)
- refactor(ui): migrate Navbar off antd to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36902](https://github.com/BerriAI/litellm/pull/36902)
- refactor(ui): migrate log details drawer off antd to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36904](https://github.com/BerriAI/litellm/pull/36904)
- refactor(ui): migrate AI Hub off antd and tremor to shadcn by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36908](https://github.com/BerriAI/litellm/pull/36908)
- refactor(ui): move the shared dropdowns and selectors onto shadcn primitives by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36924](https://github.com/BerriAI/litellm/pull/36924)
- refactor(ui): move the root-level dashboard components onto shadcn primitives by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36927](https://github.com/BerriAI/litellm/pull/36927)
- refactor(ui): move the settings page and bulk user invite onto shadcn primitives by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36936](https://github.com/BerriAI/litellm/pull/36936)
- refactor(ui): move the cost tracking components onto shadcn primitives by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36955](https://github.com/BerriAI/litellm/pull/36955)
- ci: drop the duplicate proxy\_unit\_tests letter-shard workflow by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36866](https://github.com/BerriAI/litellm/pull/36866)
- refactor(ui): migrate shared common\_components off antd and tremor by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36910](https://github.com/BerriAI/litellm/pull/36910)
- refactor(ui): migrate key info and permissions views off antd and tremor by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36913](https://github.com/BerriAI/litellm/pull/36913)
- feat(proxy): serve Anthropic-native /v1/models for Claude Code gateway discovery by [@&#8203;Ar-maan05](https://github.com/Ar-maan05) in [#&#8203;35455](https://github.com/BerriAI/litellm/pull/35455)
- refactor(ui): migrate router settings and shared badges off antd and tremor by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36915](https://github.com/BerriAI/litellm/pull/36915)
- refactor(ui): move the model hub and model select onto shadcn primitives by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36918](https://github.com/BerriAI/litellm/pull/36918)
- fix(ui): keep the cost tracking removal confirmation open until it settles by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36960](https://github.com/BerriAI/litellm/pull/36960)
- refactor(ui): declare DateRangePickerValue locally instead of importing it from tremor by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36962](https://github.com/BerriAI/litellm/pull/36962)
- fix(main): an explicit provider outranks a known OpenAI model name by [@&#8203;FahimaGold](https://github.com/FahimaGold) in [#&#8203;36800](https://github.com/BerriAI/litellm/pull/36800)
- fix(exception\_mapping): bare 429 in an error body no longer outranks the status code by [@&#8203;FahimaGold](https://github.com/FahimaGold) in [#&#8203;36705](https://github.com/BerriAI/litellm/pull/36705)
- refactor(ui): move MCP permission panels onto shadcn primitives by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36964](https://github.com/BerriAI/litellm/pull/36964)
- refactor(ui): migrate ten small dashboard files off antd and tremor by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36966](https://github.com/BerriAI/litellm/pull/36966)
- fix(proxy): force prisma recreate on postgres cached-plan error by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36428](https://github.com/BerriAI/litellm/pull/36428)
- fix(transcription): stop a zero output rate from zeroing transcription cost by [@&#8203;hMED22](https://github.com/hMED22) in [#&#8203;36914](https://github.com/BerriAI/litellm/pull/36914)
- fix(langfuse): restrict trace steering keys to real langfuse trace fields by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36862](https://github.com/BerriAI/litellm/pull/36862)
- Revert "fix(auth): stop the team fallback from widening model access" ([#&#8203;36837](https://github.com/BerriAI/litellm/issues/36837)) by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36982](https://github.com/BerriAI/litellm/pull/36982)
- fix(ui): show zeroed auto-router usage stats when a window has no sessions by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36868](https://github.com/BerriAI/litellm/pull/36868)
- fix(mcp): keep admin-entered oauth endpoints in management reads by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36888](https://github.com/BerriAI/litellm/pull/36888)
- fix(ui): distinguish hosted and local vLLM in the provider dropdown by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36974](https://github.com/BerriAI/litellm/pull/36974)
- fix(openai,azure): return a length-truncated 200 when the output budget fits no token by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36859](https://github.com/BerriAI/litellm/pull/36859)
- fix(proxy): always emit the Anthropic /v1/models token limits, null when unknown by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36961](https://github.com/BerriAI/litellm/pull/36961)
- feat(helm): add startupProbe and hpa.behavior to the componentized chart by [@&#8203;Louis-Vauterin](https://github.com/Louis-Vauterin) in [#&#8203;36382](https://github.com/BerriAI/litellm/pull/36382)
- fix(proxy): serve aggregate MCP endpoint on bare /mcp instead of 307-redirecting by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;34845](https://github.com/BerriAI/litellm/pull/34845)
- feat(shadow\_eval): add reverse-direction shadow eval jobs by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36865](https://github.com/BerriAI/litellm/pull/36865)
- fix(proxy): requeue Redis spend buffer transactions when the DB commit fails by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;33881](https://github.com/BerriAI/litellm/pull/33881)
- feat(search): add Nimble as a search provider by [@&#8203;ilchemla](https://github.com/ilchemla) in [#&#8203;36347](https://github.com/BerriAI/litellm/pull/36347)
- fix(mcp): drop caller host and configured upstream headers from logged metadata by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;36901](https://github.com/BerriAI/litellm/pull/36901)
- fix(azure\_ai): recognize real Search doc endpoints so teams can read/write via passthrough by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36798](https://github.com/BerriAI/litellm/pull/36798)
- fix(anthropic): aggregate 5m/1h cache-write split across iterations path by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;34860](https://github.com/BerriAI/litellm/pull/34860)
- fix(anthropic cost): apply regional geo uplift to cached tokens by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;34850](https://github.com/BerriAI/litellm/pull/34850)
- fix(ui): match the MCP servers count badge to its sibling permission badges by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36984](https://github.com/BerriAI/litellm/pull/36984)
- fix(batches): mark terminal batch with no output file as processed in CheckBatchCost by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35360](https://github.com/BerriAI/litellm/pull/35360)
- fix(caching): cache anthropic /v1/messages responses, including streaming by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;34581](https://github.com/BerriAI/litellm/pull/34581)
- fix(anthropic\_messages): make tool\_result images visible to OpenAI-compatible providers by [@&#8203;hMED22](https://github.com/hMED22) in [#&#8203;34462](https://github.com/BerriAI/litellm/pull/34462)
- feat(fireworks\_ai): translate NIM/vLLM extra params to Fireworks-native args by [@&#8203;milesadkins](https://github.com/milesadkins) in [#&#8203;35969](https://github.com/BerriAI/litellm/pull/35969)
- fix(ui): stop the models tab strip from scrolling vertically by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36993](https://github.com/BerriAI/litellm/pull/36993)
- fix(ui): anchor chips-combobox popups to the field instead of the inner input by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36995](https://github.com/BerriAI/litellm/pull/36995)
- feat(proxy): per-component response cost headers by [@&#8203;erensh27](https://github.com/erensh27) in [#&#8203;36965](https://github.com/BerriAI/litellm/pull/36965)
- fix(cost): track OpenAI/Azure web search tool cost per call by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35286](https://github.com/BerriAI/litellm/pull/35286)
- fix(bedrock): resolve aliases in batch file records by [@&#8203;daleselaji-dev](https://github.com/daleselaji-dev) in [#&#8203;36159](https://github.com/BerriAI/litellm/pull/36159)
- fix: report real token usage on guardrail-blocked /v1/responses replies by [@&#8203;guptaishaan](https://github.com/guptaishaan) in [#&#8203;36907](https://github.com/BerriAI/litellm/pull/36907)
- fix(proxy): requeue spend logs when the DB write fails with a transport error by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36716](https://github.com/BerriAI/litellm/pull/36716)
- fix(cost): tiered pricing supports cache creation cost and is all-or-nothing by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36720](https://github.com/BerriAI/litellm/pull/36720)
- fix(vertex\_ai): translate /v1/embeddings batch rows to the Gemini embedding shape by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;35092](https://github.com/BerriAI/litellm/pull/35092)
- docs(claude): require ReadOnly on every TypedDict field (LIT012) by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37005](https://github.com/BerriAI/litellm/pull/37005)
- refactor(ui): migrate access group create modal to RHF + zod + shadcn by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37033](https://github.com/BerriAI/litellm/pull/37033)
- refactor(ui): re-sync badge and skeleton onto the base-vega shadcn style by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;36991](https://github.com/BerriAI/litellm/pull/36991)
- feat(ui): link user detail team names to team pages by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37022](https://github.com/BerriAI/litellm/pull/37022)
- fix(model\_prices): correct DeepSeek V4 max output tokens by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36925](https://github.com/BerriAI/litellm/pull/36925)
- fix(ui): rename models table Status column to Source by [@&#8203;ryan-crabbe-berri](https://github.com/ryan-crabbe-berri) in [#&#8203;37021](https://github.com/BerriAI/litellm/pull/37021)
- chore: bump litellm-enterprise 0.1.55 -> 0.1.56, litellm-proxy-extras 0.4.85 -> 0.4.86 by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37045](https://github.com/BerriAI/litellm/pull/37045)
- feat(proxy): gate the Global Control Plane worker registry on an enterprise license by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;36996](https://github.com/BerriAI/litellm/pull/36996)
- fix(model\_prices): add gemini 3.1 flash tts preview and legacy OpenAI shutdown dates by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36788](https://github.com/BerriAI/litellm/pull/36788)
- fix(panw\_prisma\_airs): surface scan\_id on allowed requests by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37037](https://github.com/BerriAI/litellm/pull/37037)
- fix(model\_map): flag native structured outputs on Anthropic-direct claude-sonnet-5 and claude-haiku-4-5 by [@&#8203;anmolg1997](https://github.com/anmolg1997) in [#&#8203;35930](https://github.com/BerriAI/litellm/pull/35930)
- fix(router): stop get\_router\_model\_info from wiping cached pricing by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36985](https://github.com/BerriAI/litellm/pull/36985)
- fix(redis): unwrap decorated \_\_init\_\_s when deriving the from\_url kwargs allowlist by [@&#8203;anmolg1997](https://github.com/anmolg1997) in [#&#8203;36654](https://github.com/BerriAI/litellm/pull/36654)
- fix(proxy): reserve the larger declared output budget for TPM limits by [@&#8203;yassin-berriai](https://github.com/yassin-berriai) in [#&#8203;37001](https://github.com/BerriAI/litellm/pull/37001)
- fix(databricks): surface provider usage, including prompt-cache counts, in streaming chunks by [@&#8203;pokepoke81](https://github.com/pokepoke81) in [#&#8203;36943](https://github.com/BerriAI/litellm/pull/36943)
- fix(spend): give a batch's cost row a primary key of its own by [@&#8203;marty-sullivan](https://github.com/marty-sullivan) in [#&#8203;36876](https://github.com/BerriAI/litellm/pull/36876)
- feat: shadow eval samples /v1/messages and /v1/responses traffic by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36830](https://github.com/BerriAI/litellm/pull/36830)
- fix(ptu): stop a PTU deployment billing for grounded search by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;37043](https://github.com/BerriAI/litellm/pull/37043)
- fix(fireworks\_ai): support router slugs via routers/ prefix by [@&#8203;heathriel](https://github.com/heathriel) in [#&#8203;34257](https://github.com/BerriAI/litellm/pull/34257)
- fix(bedrock): register managed-batch litellm\_params so they stop leaking to the provider (internal copy of [#&#8203;36633](https://github.com/BerriAI/litellm/issues/36633)) by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37048](https://github.com/BerriAI/litellm/pull/37048)
- fix(bedrock): resolve the managed-batch output bucket on every path that reads it by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37047](https://github.com/BerriAI/litellm/pull/37047)
- fix(bedrock): resolve the managed-batch output bucket on every path that reads it by [@&#8203;marty-sullivan](https://github.com/marty-sullivan) in [#&#8203;36634](https://github.com/BerriAI/litellm/pull/36634)
- feat(scripts): queue heavy gates behind a machine-wide slot lock by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36988](https://github.com/BerriAI/litellm/pull/36988)
- feat(mcp): scope gateway session bearers to the RFC 8707 resource by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;35045](https://github.com/BerriAI/litellm/pull/35045)
- feat(ui): direction picker and reverse-mode display for shadow evals by [@&#8203;tin-berri](https://github.com/tin-berri) in [#&#8203;36994](https://github.com/BerriAI/litellm/pull/36994)
- fix(guardrails): return the full PANW AIRS scan response on blocked requests by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37036](https://github.com/BerriAI/litellm/pull/37036)
- fix(passthrough): stop forwarding client Accept-Encoding upstream by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37058](https://github.com/BerriAI/litellm/pull/37058)
- fix(batches): account a managed batch's cost exactly once by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37050](https://github.com/BerriAI/litellm/pull/37050)
- fix(panw\_prisma\_airs): scan tool call args as plain text, not a tool\_event by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37038](https://github.com/BerriAI/litellm/pull/37038)
- feat(lint): exempt TypedDict-annotated dict literals from LIT002 by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36869](https://github.com/BerriAI/litellm/pull/36869)
- docs(claude): tell agents to let heavy gates queue for machine-wide slots by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;37057](https://github.com/BerriAI/litellm/pull/37057)
- test: unstick the suites CircleCI is failing on by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37059](https://github.com/BerriAI/litellm/pull/37059)
- docs(github): proof-of-fix section shows only the latest run as Before/After with nested cases by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;37063](https://github.com/BerriAI/litellm/pull/37063)
- test(e2e): assert provider error shape instead of pinned prose by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37065](https://github.com/BerriAI/litellm/pull/37065)
- fix(ui): de-duplicate the reset budget option and polish shadcn surfaces by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37010](https://github.com/BerriAI/litellm/pull/37010)
- chore: rebuild Admin UI bundle from litellm\_internal\_staging by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37066](https://github.com/BerriAI/litellm/pull/37066)
- test(e2e/ui): assert the log drawer chevrons by their lucide classes by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37069](https://github.com/BerriAI/litellm/pull/37069)
- chore(ci): promote internal staging to main by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37042](https://github.com/BerriAI/litellm/pull/37042)
- fix(ui): keep completion-mode models in the playground chat dropdown (backport to rc/1.98.0) by [@&#8203;yuneng-berri](https://github.com/yuneng-berri) in [#&#8203;37955](https://github.com/BerriAI/litellm/pull/37955)

##### New Contributors

- [@&#8203;kr0k](https://github.com/kr0k) made their first contribution in [#&#8203;33196](https://github.com/BerriAI/litellm/pull/33196)
- [@&#8203;HuanQian571](https://github.com/HuanQian571) made their first contribution in [#&#8203;35773](https://github.com/BerriAI/litellm/pull/35773)
- [@&#8203;alexshtf](https://github.com/alexshtf) made their first contribution in [#&#8203;35669](https://github.com/BerriAI/litellm/pull/35669)
- [@&#8203;vairodp](https://github.com/vairodp) made their first contribution in [#&#8203;33490](https://github.com/BerriAI/litellm/pull/33490)
- [@&#8203;fancybear-dev](https://github.com/fancybear-dev) made their first contribution in [#&#8203;36196](https://github.com/BerriAI/litellm/pull/36196)
- [@&#8203;daleselaji-dev](https://github.com/daleselaji-dev) made their first contribution in [#&#8203;36160](https://github.com/BerriAI/litellm/pull/36160)
- [@&#8203;eugene-yao-zocdoc](https://github.com/eugene-yao-zocdoc) made their first contribution in [#&#8203;34290](https://github.com/BerriAI/litellm/pull/34290)
- [@&#8203;geraint0923](https://github.com/geraint0923) made their first contribution in [#&#8203;30817](https://github.com/BerriAI/litellm/pull/30817)
- [@&#8203;william-xue](https://github.com/william-xue) made their first contribution in [#&#8203;36529](https://github.com/BerriAI/litellm/pull/36529)
- [@&#8203;dcadenas](https://github.com/dcadenas) made their first contribution in [#&#8203;32536](https://github.com/BerriAI/litellm/pull/32536)
- [@&#8203;atomic](https://github.com/atomic) made their first contribution in [#&#8203;34177](https://github.com/BerriAI/litellm/pull/34177)
- [@&#8203;Praveen11558](https://github.com/Praveen11558) made their first contribution in [#&#8203;30952](https://github.com/BerriAI/litellm/pull/30952)
- [@&#8203;anxkhn](https://github.com/anxkhn) made their first contribution in [#&#8203;32813](https://github.com/BerriAI/litellm/pull/32813)
- [@&#8203;lostmartian](https://github.com/lostmartian) made their first contribution in [#&#8203;36660](https://github.com/BerriAI/litellm/pull/36660)
- [@&#8203;FahimaGold](https://github.com/FahimaGold) made their first contribution in [#&#8203;36800](https://github.com/BerriAI/litellm/pull/36800)
- [@&#8203;Louis-Vauterin](https://github.com/Louis-Vauterin) made their first contribution in [#&#8203;36382](https://github.com/BerriAI/litellm/pull/36382)
- [@&#8203;ilchemla](https://github.com/ilchemla) made their first contribution in [#&#8203;36347](https://github.com/BerriAI/litellm/pull/36347)
- [@&#8203;milesadkins](https://github.com/milesadkins) made their first contribution in [#&#8203;35969](https://github.com/BerriAI/litellm/pull/35969)
- [@&#8203;erensh27](https://github.com/erensh27) made their first contribution in [#&#8203;36965](https://github.com/BerriAI/litellm/pull/36965)
- [@&#8203;guptaishaan](https://github.com/guptaishaan) made their first contribution in [#&#8203;36907](https://github.com/BerriAI/litellm/pull/36907)
- [@&#8203;pokepoke81](https://github.com/pokepoke81) made their first contribution in [#&#8203;36943](https://github.com/BerriAI/litellm/pull/36943)
- [@&#8203;heathriel](https://github.com/heathriel) made their first contribution in [#&#8203;34257](https://github.com/BerriAI/litellm/pull/34257)

**Full Changelog**: <https://github.com/BerriAI/litellm/compare/v1.97.0...v1.98.0>

### [`v1.98.0`](https://github.com/BerriAI/litellm/releases/tag/v1.98.0)

[Compare Source](https://github.com/BerriAI/litellm/compare/v1.97.0...v1.98.0)

##### Verify Docker Image Signature

All LiteLLM Docker images are signed with [cosign](https://docs.sigstore.dev/cosign/overview/). Every release is signed with the same key introduced in [commit `0112e53`](https://github.com/BerriAI/litellm/commit/0112e53046018d726492c814b3644b7d376029d0).

**Verify using the pinned commit hash (recommended):**

A commit hash is cryptographically immutable, so this is the strongest way to ensure you are using the original signing key:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/0112e53046018d726492c814b3644b7d376029d0/cosign.pub \
  ghcr.io/berriai/litellm:v1.98.0
```

**Verify using the release tag (convenience):**

Tags are protected in this repository and resolve to the same key. This option is easier to read but relies on tag protection rules:

```bash
cosign verify \
  --key https://raw.githubusercontent.com/BerriAI/litellm/v1.98.0/cosign.pub \
  ghcr.io/berriai/litellm:v1.98.0
```

Expected output:

```
The following checks were performed on each of these signatures:
  - The cosign claims were validated
  - The signatures were verified against the specified public key
```

***

##### What's Changed

- fix(bedrock): drop toolSpec.strict for Claude Sonnet 5 on Converse by [@&#8203;kr0k](https://github.com/kr0k) in [#&#8203;33196](https://github.com/BerriAI/litellm/pull/33196)
- fix(batches): attribute Vertex passthrough batch cost to key/team/tags by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;34456](https://github.com/BerriAI/litellm/pull/34456)
- docs: rewrite the CLAUDE.md comment rule with explicit exceptions by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36301](https://github.com/BerriAI/litellm/pull/36301)
- fix(proxy): scope file list pagination cursors to the caller by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36093](https://github.com/BerriAI/litellm/pull/36093)
- fix(proxy): skip prisma-dependent hooks when no database is attached by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36273](https://github.com/BerriAI/litellm/pull/36273)
- fix(proxy): report has\_more false on caller-scoped file list pages by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36326](https://github.com/BerriAI/litellm/pull/36326)
- fix(proxy): restore management\_v1 query-param validation under fastapi>=0.140.7 by [@&#8203;HuanQian571](https://github.com/HuanQian571) in [#&#8203;35773](https://github.com/BerriAI/litellm/pull/35773)
- fix(proxy): stop /{provider}/v1/files from capturing /openai\_passthrough by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36092](https://github.com/BerriAI/litellm/pull/36092)
- chore(typing): remove 914 basedpyright Any errors across 16 hotspot files by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36386](https://github.com/BerriAI/litellm/pull/36386)
- fix(router): keep batch fallbacks inside the model group that owns the file by [@&#8203;mateo-berri](https://github.com/mateo-berri) in [#&#8203;36181](https://github.com/BerriAI/litellm/pull/36181)
- feat(ptu): configure provisioned-throughput flat cost on a model deployment by [@&#8203;yucheng-berri](https://github.com/yucheng-berri) in [#&#8203;35341](https://github.com/BerriAI/litellm/pull/35341)
- docs: clarify the CLAUDE.md comment exceptions are any-of by [@&#8203;devin-ai-integration](https://github.com/devin-ai-integration)\[bot] in [#&#8203;36421](https://github.com/BerriAI/litellm/pull/36421)
- docs: replace the Changes …
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants