Skip to content

fix(router): resolve team-scoped auto-routers by their public name - #40432

Merged
tin-berri merged 1 commit into
litellm_internal_stagingfrom
litellm_lit7363_team_autorouter_public_name
Sep 9, 2026
Merged

tin-berri merged 1 commit into
litellm_internal_stagingfrom
litellm_lit7363_team_autorouter_public_name

Conversation

@tin-berri

@tin-berri tin-berri commented Sep 9, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Team-scoped auto-routers 400 "Unmapped LLM provider" on every call
  • The Auto-Routers tab already lets team admins create them, so each one is dead on arrival

How it solves it:

  • Strategy lookup and compression policy resolve the public name the way deployment lookup already does
  • Marker deployments are dropped on both resolution exits, so marker-only names fail clearly

User Flow

Before: a team admin creates a complexity router for their team, sees it listed, and every call to it fails

  1. They send POST https://litellm-domain/model/new with model_name: "team-router", litellm_params.model: "auto_router/complexity_router", tiers over models the team can use, and model_info.team_id set to their team; they get 200
  2. A key on that team sends GET https://litellm-domain/v1/models and sees team-router in the list
  3. The same key sends POST https://litellm-domain/v1/chat/completions with "model": "team-router" and gets 400 Unmapped LLM provider for this endpoint. You passed model=complexity_router, custom_llm_provider=auto_router
  4. The identical router created without model_info.team_id answers with 200 through its cheapest tier

After: the same router routes the team's requests like the unscoped one does

  1. They send the same POST https://litellm-domain/model/new and get 200
  2. The team key sees team-router in GET https://litellm-domain/v1/models
  3. The team key sends the same POST https://litellm-domain/v1/chat/completions and gets 200, with x-litellm-model-name naming the tier model the router picked and a non-zero x-litellm-response-cost
  4. POST https://litellm-domain/v1/messages, POST https://litellm-domain/v1/responses and streaming chat completions behave the same way
  5. A key on another team still cannot reach it, and a proxy admin without a team can call it by its public name like any other team model

Relevant issues

  • Team-scoped strategy routers (complexity, semantic, adaptive, quality) were registered under their internal model_name_{team}_{uuid} while requests arrive with the team public name, so no strategy was ever selected for them
  • The strategy lookup now derives the registry keys from the deployments the request resolves to, through the same team-first resolution the deployment path uses, so the two arms cannot disagree about precedence
  • The team early-resolve exit returned marker deployments without the marker filter; both exits now share one filter
  • The auto-router compression policy resolves its marker through the same request-scoped lookup, so a team router reached by any principal carries its policy
  • Sibling of LIT-4664 (model_group_alias pointing at an auto-router)

Linear ticket

Resolves LIT-7363

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Screenshots / Proof of Fix

Rig: proxy from this worktree on 127.0.0.1:4737, fresh Postgres, store_model_in_db: true, two tier deployments (haiku-direct, sonnet-direct) pointing at a real upstream gateway, so every 200 below is a billed model call. Setup shared by both runs, $SFX is a per-run suffix so prompts are unique upstream:

curl -s -X POST $BASE/team/new -H "Authorization: Bearer $MASTER" -H 'Content-Type: application/json' \
  -d '{"team_alias":"ar-team-'$SFX'","models":["haiku-direct","sonnet-direct"]}'          # -> team_id
curl -s -X POST $BASE/model/new -H "Authorization: Bearer $MASTER" -H 'Content-Type: application/json' \
  -d '{"model_name":"team-router-'$SFX'","litellm_params":{"model":"auto_router/complexity_router","complexity_router_config":{"tiers":{"SIMPLE":"haiku-direct","MEDIUM":"sonnet-direct","COMPLEX":"sonnet-direct"}}},"model_info":{"team_id":"'$TEAM_ID'"}}'
curl -s -X POST $BASE/model/new -H "Authorization: Bearer $MASTER" -H 'Content-Type: application/json' \
  -d '{"model_name":"global-router-'$SFX'","litellm_params":{"model":"auto_router/complexity_router","complexity_router_config":{"tiers":{"SIMPLE":"haiku-direct","MEDIUM":"sonnet-direct","COMPLEX":"sonnet-direct"}}}}'
curl -s -X POST $BASE/key/generate -H "Authorization: Bearer $MASTER" -H 'Content-Type: application/json' \
  -d '{"team_id":"'$TEAM_ID'"}'                                                             # -> TEAM_KEY
curl -s $BASE/v1/models -H "Authorization: Bearer $TEAM_KEY" | jq -c '[.data[].id]'

Both runs listed ["haiku-direct","sonnet-direct","team-router-$SFX"] for the team key and returned 200 to both /model/new calls

Before (ea0851d)

Team key, POST /v1/chat/completions to the team router

  1. curl -s -D - -X POST $BASE/v1/chat/completions -H "Authorization: Bearer $TEAM_KEY" -H 'Content-Type: application/json' -d '{"model":"team-router-5491","messages":[{"role":"user","content":"Say hi in one word. Reference 5491-1"}]}' (sent three times)
  2. All three: HTTP 400, x-litellm-response-cost: 0, body {"error":{"message":"litellm.BadRequestError: Unmapped LLM provider for this endpoint. You passed model=complexity_router, custom_llm_provider=auto_router. ... Received Model Group=team-router-5491\nAvailable Model Group Fallbacks=None","type":"invalid_request_error","code":"400"}}

Control: master key, POST /v1/chat/completions to the unscoped twin

  1. Same body with "model":"global-router-5491" and the master key
  2. HTTP 200, x-litellm-model-name: anthropic/claude-haiku-4-5, x-litellm-response-cost: 8.900000000000001e-05, x-litellm-model-group: global-router-5491, content "Hey\n\n(Reference: 5491-g)"

After (1b9c631)

Team key, POST /v1/chat/completions to the team router

  1. Same command, "model":"team-router-0051", sent three times
  2. All three: HTTP 200, x-litellm-model-name: anthropic/claude-haiku-4-5, x-litellm-model-group: team-router-0051, x-litellm-response-cost: 3.9e-05 / 7.900000000000001e-05 / 3.9e-05, content "Hi", "Hi\n\nReference: 0051-2", "Hi"

Control: master key, POST /v1/chat/completions to the unscoped twin

  1. Same body with "model":"global-router-0051" and the master key
  2. HTTP 200, x-litellm-model-name: anthropic/claude-haiku-4-5, x-litellm-response-cost: 7.900000000000001e-05, content "Hi\n\nReference: 0051-g"

Team key, POST /v1/messages to the team router

  1. curl -s -D - -X POST $BASE/v1/messages -H "Authorization: Bearer $TEAM_KEY" -H 'Content-Type: application/json' -d '{"model":"team-router-0051","max_tokens":32,"messages":[{"role":"user","content":"Say hi in one word. Reference 0051-msg"}]}'
  2. HTTP 200, x-litellm-model-name: anthropic/claude-haiku-4-5, x-litellm-response-cost: 3.9e-05, body {"model":"team-router-0051","id":"msg_011CetbHnyjB2fuA9PDcQGMa","type":"message","role":"assistant","content":[{"type":"text","text":"Hey"}],...}

Team key, POST /v1/responses to the team router

  1. curl -s -D - -X POST $BASE/v1/responses -H "Authorization: Bearer $TEAM_KEY" -H 'Content-Type: application/json' -d '{"model":"team-router-0051","input":"Say hi in one word. Reference 0051-resp"}'
  2. HTTP 200, x-litellm-model-name: anthropic/claude-haiku-4-5, x-litellm-response-cost: 3.9e-05, body {"id":"resp_qmppnSVCrPHHBWBzq_HEIaL7BuNQ...

Team key, streaming POST /v1/chat/completions to the team router

  1. Same chat body with "stream": true
  2. HTTP 200, x-litellm-model-name: anthropic/claude-haiku-4-5, 3 data: chunks, first chunk {"id":"chatcmpl-21cda7e6-7a48-4f1f-9aa5-636690682623","object":"chat.completion.chunk","model":"team-router-0051",...}

Proxy admin without a team, POST /v1/chat/completions to the team router by its public name

  1. Same chat body with the master key and "model":"team-router-0051"
  2. HTTP 200, x-litellm-model-name: anthropic/claude-haiku-4-5, x-litellm-response-cost: 8.900000000000001e-05, content "Hey\n\n(Reference: 0051-admin)"

A second team's key, POST /v1/chat/completions to the team router

  1. POST /team/new for a second team allowed only haiku-direct, POST /key/generate on it, then the same chat body with that key
  2. HTTP 403, {"error":{"message":"team not allowed to access model. This team can only access models=['haiku-direct']. Tried to access team-router-0051","type":"team_model_access_denied","code":"403"}}

Type

🐛 Bug Fix

Caveats (if any)

Low

  • osv-scan is red on GHSA-7w5x-hrqm-74c2 (smol-toml, a dashboard dev dependency this PR does not touch); the advisory published 2026-09-09 18:07 UTC and reds every open PR, fix(ui): bump smol-toml to 1.8.0 to clear GHSA-7w5x-hrqm-74c2 in osv-scan #40442 bumps the lock file
  • The shadow-eval validator _is_configured_pre_routing_strategy still keys on internal names, so it cannot name a team router by its public name; follow-up
  • The sync Router.get_available_deployment path runs no pre-routing hook at all, which predates this PR and only affects SDK-sync callers

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

🤖 Generated with Claude Code

https://claude.ai/code/session_01NU97S7d2FUDDvTk59k53Wp

@tin-berri
tin-berri requested a review from a team September 9, 2026 18:00
@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai can you reply here with a score and why that score?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

Confidence score: 5/5

This directly fixes the root cause: strategy lookup now derives registry keys from the same team-first/global/admin deployment resolution used by deployment selection, so team public names resolve to their internally registered marker deployments without changing the caller-facing model name.

Marker filtering is applied consistently to both early and normal resolution paths, including a clear failure for marker-only model groups. The shared team-id extraction also removes the previous metadata-bucket inconsistency.

The test coverage is strong: all four strategy registries, team isolation, public-name shadowing, tag selection, proxy-admin access, multi-team ambiguity, marker-only rejection, and parity between strategy and deployment resolution are covered. The provided integration evidence additionally validates chat, streaming, Messages, Responses, admin access, and cross-team denial. The documented shadow-eval and sync-SDK caveats are outside this fix’s requested async proxy path.

I found no blocking correctness or security concerns.

@codspeed

codspeed Bot commented Sep 9, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_lit7363_team_autorouter_public_name (1b9c631) with litellm_internal_staging (9315f5d)

Open in CodSpeed

@greptile-apps

greptile-apps Bot commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR resolves team-scoped auto-router strategies by deriving their internal registry names from the deployments visible through a public team model name. It also centralizes team ID extraction and removes strategy markers from early deployment-resolution results.

  • Adds team-aware and proxy-admin-aware strategy resolution
  • Prevents marker-only public names from reaching provider invocation
  • Extends router tests across strategy registries, tags, team isolation, and admin access
  • Leaves teamless proxy-admin compression policy resolution inconsistent with the new strategy path

Confidence Score: 4/5

The PR should not merge until teamless proxy-admin calls apply the same team-scoped compression policy as strategy selection

The new admin-across-teams strategy path successfully selects the team router, but its configured compression policy resolves through a narrower lookup and is silently skipped

Files Needing Attention: litellm/router.py, litellm/proxy/guardrails/auto_router_compression.py

Important Files Changed

Filename Overview
litellm/router.py Adds team-public strategy resolution and shared marker filtering, but the teamless-admin path is not mirrored by compression policy lookup
litellm/proxy/guardrails/auto_router_compression.py Reuses centralized team ID extraction while retaining a deployment resolver that cannot find team-public markers for teamless admins
litellm/router_utils/common_utils.py Centralizes validated team ID extraction and applies it to team-based deployment filtering
tests/test_litellm/test_router.py Adds broad team-public auto-router regression coverage, but does not cover configured compression for a teamless proxy admin

Reviews (1): Last reviewed commit: "fix(router): resolve team-scoped auto-ro..." | Re-trigger Greptile

Comment thread litellm/router.py Outdated
deployments: Final = self._get_all_deployments(model_name=model, team_id=team_id)
if deployments or team_id is not None or not _is_proxy_admin_request(request_kwargs):
return deployments
return self._team_deployments_across_teams(model)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Compression policy is skipped

Teamless proxy admins resolve the team router here, but compression uses a narrower lookup, silently skipping configured routing and model compression.

Knowledge Base Used:

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 2537ee3 by removing the third resolver rather than mirroring the admin arm into it: policy_for_model now takes the request kwargs and resolves its markers through Router.deployments_for_request, the same alias, team-first, global, admin-across-teams lookup strategy selection and marker param forwarding use, and team_id_from_request is gone. Pinned by test_compression_policy_follows_the_same_resolution_for_every_principal (team key and teamless proxy admin get the team marker's policy, another team and a keyless caller get none), which fails on the previous revision

@codecov

codecov Bot commented Sep 9, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.50000% with 1 line in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
litellm/router.py 96.55% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/router.py Outdated
@tin-berri
tin-berri force-pushed the litellm_lit7363_team_autorouter_public_name branch from f03a21e to 2537ee3 Compare September 9, 2026 18:16
@tin-berri

Copy link
Copy Markdown
Contributor Author

Pushed 2537ee3: one round for the shared finding. The compression policy resolver was a third resolver of "which markers does this name mean for this caller", so it now goes through Router.deployments_for_request (made public, alias-resolving, and the single owner for strategy selection, marker param forwarding, and compression policy); team_id_from_request is deleted. The rebase onto current staging also adapts two tests staging added after this branch's base (test_forwardable_alias_marker_params_reads_the_marker_entry_only gains the new request_kwargs argument, and test_model_name_has_plain_deployments_reflects_the_pool is dropped since that helper no longer exists and its two behavioural neighbours in the same class already pin the pool check). Live proof re-run on the new tip is in the description

@greptileai can you reply here with a score and why that score?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

Confidence score: 5/5

The shared finding is fully addressed in 2537ee3601. Compression policy resolution now uses the public Router.deployments_for_request resolver, so strategy selection, marker-parameter forwarding, and compression policy all apply the same alias/team-first/global/admin resolution rules. This correctly covers teamless proxy-admin requests while preserving cross-team isolation.

The follow-up removes the duplicated team_id_from_request path and adds regression coverage for compression policy resolution across team, admin, other-team, and keyless callers. The broader tests cover all four strategy registries, tag selection, marker-only rejection, public-name shadowing, and strategy/deployment parity, and the re-run live proof covers the relevant endpoints and streaming behavior.

I found no remaining blocking correctness or security concerns.

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

A team-scoped auto-router is stored under an internal
model_name_{team_id}_{uuid} with the caller-facing name in
model_info.team_public_model_name, and the four pre-routing strategy
registries key on that internal name. A team key asks for the public name,
so the strategy lookup missed, the team early-resolve exit handed back the
marker deployment itself, and every call 400'd with "Unmapped LLM provider".

The strategy lookup now resolves the requested name through the same
team-first, then global, then admin-across-teams deployment resolution the
deployment path uses, and looks the registries up under the model_name of
whatever that resolves to. Both exits of _common_checks_available_deployment
drop strategy markers through one helper, so a marker-only resolution is
rejected as uncallable on every path. The request team id has one reader.

Resolves LIT-7363

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NU97S7d2FUDDvTk59k53Wp
@tin-berri
tin-berri force-pushed the litellm_lit7363_team_autorouter_public_name branch from 2537ee3 to 1b9c631 Compare September 9, 2026 21:42
@tin-berri

Copy link
Copy Markdown
Contributor Author

Pushed 1b9c631: the one proxy-infra red was test_common_processing_pre_call_logic_arms_auto_router_compression_before_guardrails, whose fake router stubbed get_model_list, which the compression policy no longer reads; the fake now stubs deployments_for_request. No production change. Live proof re-run on this tip is in the description. osv-scan is a new smol-toml advisory (published 2026-09-09 18:07 UTC) that reds every open PR and is fixed by #40442

@greptileai can you reply here with a score and why that score?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

Score: 4/5

The fix is well-supported, and the latest test-only follow-up addresses the single proxy-infra failure without changing production behavior. The implementation now uses one caller-scoped deployment-resolution path for strategy lookup, compression policy, marker parameter forwarding, and deployment selection; marker-only results are also rejected on both resolution exits. The added tests cover all four strategy registries, team isolation, public-name shadowing, tag-scoped siblings, proxy-admin access, compression-policy alignment, and marker-only failures. The live proof validates the affected chat, Messages, Responses, and streaming paths.

I’m holding back the fifth point because required CI is still pending, and the PR documents remaining limitations in shadow-eval validation and the sync SDK path. The osv-scan failure appears to be an unrelated newly published advisory affecting open PRs and does not reduce confidence in this change itself.

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 1b9c631. Configure here.

@tin-berri

Copy link
Copy Markdown
Contributor Author

CI has finished on 1b9c631: 83 checks green, 1 skipped, codecov/patch success, and the only red is osv-scan. That red is the smol-toml advisory GHSA-7w5x-hrqm-74c2 published today at 18:07 UTC: it reds every open PR against staging, #40456 and #40444 merged this evening carrying the same red, this branch never touches the dashboard lock file, and #40442 bumps it. Bugbot reviewed this sha and found no new issues. The two remaining checklist boxes (required CI, Greptile score) are now ticked in the description

@greptileai can you reply here with a score and why that score?

@greptile-apps

greptile-apps Bot commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

Greptile confidence score: 5/5

This is a focused and correct fix. Strategy selection and compression-policy lookup now use the same request-scoped, team-first deployment resolution as the actual deployment path, so team-scoped routers consistently resolve their public names. Shared marker filtering also prevents marker-only deployments from being treated as callable on either resolution path.

The tests cover all four strategy registries, team isolation, public-name shadowing, tag selection, proxy-admin access, marker-only rejection, marker-parameter forwarding, compression policy, and agreement between strategy and deployment resolution. The reported end-to-end results additionally verify chat, streaming, Messages, Responses, and access-control behavior.

The only failing check is the unrelated newly published smol-toml advisory, which this PR does not touch and is already addressed separately. The shadow-eval and sync-path items are clearly identified as pre-existing follow-ups. No issues found that should block merging.

@tin-berri
tin-berri merged commit eb45a08 into litellm_internal_staging Sep 9, 2026
87 of 89 checks passed
@tin-berri
tin-berri deleted the litellm_lit7363_team_autorouter_public_name branch September 9, 2026 22:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants