Skip to content

fix(license): backport #41684 to stable/1.101.x so a wildcard license grants auto_router - #41698

Merged
mateo-berri merged 1 commit into
stable/1.101.xfrom
litellm_cherrypick_1_101_x
Sep 17, 2026
Merged

mateo-berri merged 1 commit into
stable/1.101.xfrom
litellm_cherrypick_1_101_x

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

How it solves it:

Backport notes

  1. Cherry-picked from main: c2fbb11, the single commit of fix(license): let a wildcard allowed_features license grant the auto_router feature #41684 (main lands PRs as merge commits, so there is no squash commit to pick)
  2. tests/test_litellm/proxy/auth/test_litellm_license.py applied clean
  3. litellm/proxy/auth/litellm_license.py conflicted only in the docstring of auto_router_capability_limit, whose wording on this line predates main's; the resolution takes fix(license): let a wildcard allowed_features license grant the auto_router feature #41684's docstring, and the whole region from AUTO_ROUTER_LICENSE_FEATURE down to the method's return 1 is identical to main
  4. No version bump: cutting v1.101.1 is a separate step, the same way fix(responses): backport mid-stream content_policy_violation fallback routing to stable/1.100.x #41208 landed on stable/1.100.x
  5. The design carries over unchanged from fix(license): let a wildcard allowed_features license grant the auto_router feature #41684; every caller of auto_router_capability_limit on this line (the boot-time validation, the two Router constructions, and POST /model/new) is also a caller on main, so the main-side risk check applies here

User Flow

Before: an enterprise operator upgrading to v1.101.0 with two custom auto-routers cannot start the proxy at all, even though their license covers every feature

  1. They set LITELLM_LICENSE to their enterprise key, whose decoded payload reads "allowed_features": ["*"], and keep a config.yaml with two auto_router/complexity_router deployments that each define their own tier_definitions
  2. They start the v1.101.0 proxy (litellm --config config.yaml)
  3. Startup aborts with ValueError: config.yaml model_list: At most 1 auto-router(s) with operator-defined tier_definitions or an operator-written classifier prompt can be registered but this would make 2. ... A LiteLLM license with the 'auto_router' feature lifts the limit. and the process exits
  4. GET http://litellm-domain/health/liveliness gets no answer: the connection fails or times out because no worker ever starts serving

After: the same operator starts the proxy built from this line and both routers serve traffic

  1. They set LITELLM_LICENSE to the same enterprise key ("allowed_features": ["*"]) and keep the same config.yaml with two custom auto-routers
  2. They start the proxy with this fix (litellm --config config.yaml)
  3. The proxy boots and GET http://litellm-domain/health/liveliness returns "I'm alive!"
  4. GET http://litellm-domain/health/license returns 200 with "allowed_features": ["*"]
  5. POST http://litellm-domain/v1/chat/completions with "model": "support-router" and again with "model": "engineering-router" each return 200 with a real completion

Relevant issues

Backport of #41684 (main)

Affected release

regression in v1.101.0 (first in v1.101.0-rc.1, from #39674 on top of #39468); this PR carries the fix onto stable/1.101.x

Linear ticket

Resolves LIT-8019

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Every leg is a live proxy booted from a detached worktree at the named commit with its own venv and 2 uvicorn workers on a random free port, no database. The license is read from the LITELLM_LICENSE env var by each worker on its own, so there is no state shared between workers to cross-check, and each chat request is sent twice. The routers call OpenAI for real (gpt-5.4-mini classifies and answers the easy tier, gpt-5.6 the hard one). The three license keys are offline-signed test keys that differ only in allowed_features; their payloads carry max_users: 100, max_teams: 5, and an expiry of 2027-07-06. Chat responses are trimmed with jq to the model, the content, and the token count

config.yaml, the same file on every leg:

model_list:
  - model_name: gpt-5.4-mini
    litellm_params:
      model: openai/gpt-5.4-mini
      api_key: os.environ/OPENAI_API_KEY
  - model_name: gpt-5.6
    litellm_params:
      model: openai/gpt-5.6
      api_key: os.environ/OPENAI_API_KEY
  - model_name: support-router
    litellm_params:
      model: auto_router/complexity_router
      complexity_router_config:
        classifier_type: llm
        classifier_llm_config:
          model: gpt-5.4-mini
        tier_definitions:
          - name: routine
            description: short factual answers and routine drafting
          - name: hard
            description: multi-step reasoning or analysis
        tiers:
          routine: gpt-5.4-mini
          hard: gpt-5.6
        fallback_tier: routine
  - model_name: engineering-router
    litellm_params:
      model: auto_router/complexity_router
      complexity_router_config:
        classifier_type: llm
        classifier_llm_config:
          model: gpt-5.4-mini
        tier_definitions:
          - name: quick
            description: one-line code questions and lookups
          - name: deep
            description: architecture, debugging across files, or proofs
        tiers:
          quick: gpt-5.4-mini
          deep: gpt-5.6
        fallback_tier: quick

general_settings:
  master_key: sk-1234

Before is the tip of stable/1.101.x (18243cd, the commit tagged v1.101.0); After is this PR's tip (c9ed764). Each row is a separate live proxy:

Checkpoint Proxy commit License Expected Boot support-router engineering-router
Before this PR 18243cd ["*"] blocked blocked, ValueError no answer no answer
Before this PR 18243cd ["auto_router"] boots boots 200 + 200 200 + 200
Before this PR 18243cd ["sso", "audit_logs"] blocked blocked, ValueError no answer no answer
After this PR c9ed764 ["*"] boots boots 200 + 200 200 + 200
After this PR c9ed764 ["auto_router"] boots boots 200 + 200 200 + 200
After this PR c9ed764 ["sso", "audit_logs"] blocked blocked, ValueError no answer no answer

Before (18243cd)

Wildcard license ("allowed_features": ["*"])

  1. Start the proxy; startup aborts in both workers and nothing ever serves

    export LITELLM_LICENSE="$(cat license_wildcard.txt)"
    litellm --config config.yaml --port 50185 --num_workers 2
    ValueError: config.yaml model_list: At most 1 auto-router(s) with operator-defined tier_definitions or an operator-written classifier prompt can be registered but this would make 2. Use the shipped tiers and classifier prompt for this router or remove an existing router with tier_definitions or its own classifier prompt. A LiteLLM license with the 'auto_router' feature lifts the limit.
    ERROR:    Application startup failed. Exiting.
    
  2. GET /health/liveliness gets no answer

    $ curl -sS -m 5 http://localhost:50185/health/liveliness
    curl: (7) Failed to connect to localhost port 50185 after 3957 ms: Couldn't connect to server
    

License naming the feature ("allowed_features": ["auto_router"])

  1. Start the proxy; both workers come up and it serves traffic

    export LITELLM_LICENSE="$(cat license_control_auto_router.txt)"
    litellm --config config.yaml --port 40830 --num_workers 2
    
  2. GET /health/liveliness returns 200

    $ curl -s -w '\nHTTP %{http_code}\n' http://localhost:40830/health/liveliness
    "I'm alive!"
    HTTP 200
    
  3. GET /health/license returns 200

    $ curl -s -w '\nHTTP %{http_code}\n' http://localhost:40830/health/license -H 'Authorization: Bearer sk-1234'
    {"has_license":true,"license_type":"enterprise","expiration_date":"2027-07-06","allowed_features":["auto_router"],"limits":{"max_users":100,"max_teams":5}}
    HTTP 200
    
  4. POST /v1/chat/completions with "model": "support-router", sent twice; both return 200 with a real completion

    $ curl -s -o out.json -w 'HTTP %{http_code}\n' http://localhost:40830/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"model":"support-router","messages":[{"role":"user","content":"What year did the first moon landing happen? One line."}]}'; jq -c '{model, content: .choices[0].message.content, total_tokens: .usage.total_tokens}' out.json
    HTTP 200
    {"model":"support-router","content":"1969.","total_tokens":24}
    $ (same request again)
    HTTP 200
    {"model":"support-router","content":"1969","total_tokens":23}
    
  5. POST /v1/chat/completions with "model": "engineering-router", sent twice; both return 200 with a real completion

    $ curl -s -o out.json -w 'HTTP %{http_code}\n' http://localhost:40830/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"model":"engineering-router","messages":[{"role":"user","content":"In Python, what does dict.setdefault do? One line."}]}'; jq -c '{model, content: .choices[0].message.content, total_tokens: .usage.total_tokens}' out.json
    HTTP 200
    {"model":"engineering-router","content":"`dict.setdefault(key, default)` returns the value for `key` if it exists; otherwise it inserts `key` with `default` and returns `default`.","total_tokens":54}
    $ (same request again)
    HTTP 200
    {"model":"engineering-router","content":"`dict.setdefault(key, default)` returns the value for `key` if it exists; otherwise it inserts `key` with `default` and returns `default`.","total_tokens":54}
    

License with other features only ("allowed_features": ["sso", "audit_logs"])

  1. Start the proxy; startup aborts in both workers and nothing ever serves

    export LITELLM_LICENSE="$(cat license_other_features.txt)"
    litellm --config config.yaml --port 59205 --num_workers 2
    ValueError: config.yaml model_list: At most 1 auto-router(s) with operator-defined tier_definitions or an operator-written classifier prompt can be registered but this would make 2. Use the shipped tiers and classifier prompt for this router or remove an existing router with tier_definitions or its own classifier prompt. A LiteLLM license with the 'auto_router' feature lifts the limit.
    ERROR:    Application startup failed. Exiting.
    
  2. GET /health/liveliness gets no answer

    $ curl -sS -m 5 http://localhost:59205/health/liveliness
    curl: (28) Connection timed out after 5003 milliseconds
    

After (c9ed764)

Wildcard license ("allowed_features": ["*"])

  1. Start the proxy; both workers come up and it serves traffic

    export LITELLM_LICENSE="$(cat license_wildcard.txt)"
    litellm --config config.yaml --port 24204 --num_workers 2
    
  2. GET /health/liveliness returns 200

    $ curl -s -w '\nHTTP %{http_code}\n' http://localhost:24204/health/liveliness
    "I'm alive!"
    HTTP 200
    
  3. GET /health/license returns 200

    $ curl -s -w '\nHTTP %{http_code}\n' http://localhost:24204/health/license -H 'Authorization: Bearer sk-1234'
    {"has_license":true,"license_type":"enterprise","expiration_date":"2027-07-06","allowed_features":["*"],"limits":{"max_users":100,"max_teams":5}}
    HTTP 200
    
  4. POST /v1/chat/completions with "model": "support-router", sent twice; both return 200 with a real completion

    $ curl -s -o out.json -w 'HTTP %{http_code}\n' http://localhost:24204/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"model":"support-router","messages":[{"role":"user","content":"What year did the first moon landing happen? One line."}]}'; jq -c '{model, content: .choices[0].message.content, total_tokens: .usage.total_tokens}' out.json
    HTTP 200
    {"model":"support-router","content":"1969","total_tokens":23}
    $ (same request again)
    HTTP 200
    {"model":"support-router","content":"1969","total_tokens":23}
    
  5. POST /v1/chat/completions with "model": "engineering-router", sent twice; both return 200 with a real completion

    $ curl -s -o out.json -w 'HTTP %{http_code}\n' http://localhost:24204/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"model":"engineering-router","messages":[{"role":"user","content":"In Python, what does dict.setdefault do? One line."}]}'; jq -c '{model, content: .choices[0].message.content, total_tokens: .usage.total_tokens}' out.json
    HTTP 200
    {"model":"engineering-router","content":"`dict.setdefault(key, default)` returns the value for `key` if it exists; otherwise it inserts `key` with `default` and returns `default`.","total_tokens":54}
    $ (same request again)
    HTTP 200
    {"model":"engineering-router","content":"`dict.setdefault(key, default)` returns the value for `key` if it exists; otherwise it inserts `key` with `default` and returns `default`.","total_tokens":54}
    

License naming the feature ("allowed_features": ["auto_router"])

  1. Start the proxy; both workers come up and it serves traffic

    export LITELLM_LICENSE="$(cat license_control_auto_router.txt)"
    litellm --config config.yaml --port 26738 --num_workers 2
    
  2. GET /health/liveliness returns 200

    $ curl -s -w '\nHTTP %{http_code}\n' http://localhost:26738/health/liveliness
    "I'm alive!"
    HTTP 200
    
  3. GET /health/license returns 200

    $ curl -s -w '\nHTTP %{http_code}\n' http://localhost:26738/health/license -H 'Authorization: Bearer sk-1234'
    {"has_license":true,"license_type":"enterprise","expiration_date":"2027-07-06","allowed_features":["auto_router"],"limits":{"max_users":100,"max_teams":5}}
    HTTP 200
    
  4. POST /v1/chat/completions with "model": "support-router", sent twice; both return 200 with a real completion

    $ curl -s -o out.json -w 'HTTP %{http_code}\n' http://localhost:26738/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"model":"support-router","messages":[{"role":"user","content":"What year did the first moon landing happen? One line."}]}'; jq -c '{model, content: .choices[0].message.content, total_tokens: .usage.total_tokens}' out.json
    HTTP 200
    {"model":"support-router","content":"1969.","total_tokens":24}
    $ (same request again)
    HTTP 200
    {"model":"support-router","content":"1969.","total_tokens":24}
    
  5. POST /v1/chat/completions with "model": "engineering-router", sent twice; both return 200 with a real completion

    $ curl -s -o out.json -w 'HTTP %{http_code}\n' http://localhost:26738/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"model":"engineering-router","messages":[{"role":"user","content":"In Python, what does dict.setdefault do? One line."}]}'; jq -c '{model, content: .choices[0].message.content, total_tokens: .usage.total_tokens}' out.json
    HTTP 200
    {"model":"engineering-router","content":"`dict.setdefault(key, default)` returns the value for `key` if it exists; otherwise it inserts `key` with `default` and returns `default`.","total_tokens":54}
    $ (same request again)
    HTTP 200
    {"model":"engineering-router","content":"`dict.setdefault(key, default)` returns the value for `key` if it exists; otherwise it inserts `key` with `default` and returns `default`.","total_tokens":54}
    

License with other features only ("allowed_features": ["sso", "audit_logs"])

  1. Start the proxy; startup aborts in both workers and nothing ever serves

    export LITELLM_LICENSE="$(cat license_other_features.txt)"
    litellm --config config.yaml --port 41876 --num_workers 2
    ValueError: config.yaml model_list: At most 1 auto-router(s) with operator-defined tier_definitions or an operator-written classifier prompt can be registered but this would make 2. Use the shipped tiers and classifier prompt for this router or remove an existing router with tier_definitions or its own classifier prompt. A LiteLLM license with the 'auto_router' feature lifts the limit.
    ERROR:    Application startup failed. Exiting.
    
  2. GET /health/liveliness gets no answer

    $ curl -sS -m 5 http://localhost:41876/health/liveliness
    curl: (7) Failed to connect to localhost port 41876 after 2020 ms: Couldn't connect to server
    

Notes from the run:

Type

🐛 Bug Fix

Caveats (if any)

Low

  • Only the exact * entry is a wildcard; glob patterns like auto_* still match nothing
    • The license generator only ever emits the literal *, so nothing in the field uses patterns
  • GET /health/license keeps reporting allowed_features verbatim, so a wildcard license still shows ["*"] there
  • A bare-string allowed_features is read as a one-item list, the same way GET /health/license reports it
  • A license verified through the API carries no feature list, so it keeps the limit, as before
  • The POST /model/new path of the same gate was proven on fix(license): let a wildcard allowed_features license grant the auto_router feature #41684 (second router 403 before, 200 after) and not re-run on this line; the gate's code region is identical to main's
  • CircleCI local_testing_part1 (job), local_testing_part2 (job), and llm_translation_testing (job) are red only on Together AI tests that call the serverless openai/gpt-oss-20b the provider has since withdrawn (Unable to access non-serverless model); main moved those tests off that model in 1aa2e19, 8c046e1, and b478131 after this line branched, main's latest pipeline passes all three jobs, this PR touches no Together AI path, and CircleCI is not a required check on stable/1.101.x

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Note

Medium Risk
Changes signed-license feature gating used at proxy startup and model registration; mis-handling could block valid enterprise deployments or over-grant features, though scope is limited to auto_router limit logic.

Overview
Backports the #41684 fix so enterprise licenses with allowed_features: ["*"] (the generator default) are treated as granting auto_router, not only licenses that name that feature explicitly.

LicenseCheck.grants_feature is added with a LICENSE_ALL_FEATURES ("*") wildcard; auto_router_capability_limit now delegates to it instead of inlining list checks. Licenses with "*" in the list (or as a bare string) return unlimited auto-router capacity; behavior for named auto_router, missing features, and API-only verification is unchanged.

Tests cover wildcard list/string forms, grants_feature, and signed-license verification paths.

Reviewed by Cursor Bugbot for commit c9ed764. Bugbot is set up for automated code reviews on this repo. Configure here.

…router feature

Backport of #41684 to stable/1.101.x.
Cherry-picked from c2fbb11 (litellm_wildcard_license_auto_router).
@devin-ai-integration

devin-ai-integration Bot commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor Author

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR. Add '(aside)' to your comment to have me ignore it.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@greptile-apps

greptile-apps Bot commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge because wildcard entitlements remain restricted to verified license payloads and have focused regression coverage

Summary

This PR makes wildcard licenses grant the auto_router feature

  • Adds a reusable feature-entitlement check supporting named and wildcard grants
  • Applies the check to the auto-router capability limit
  • Adds direct and signed-license regression coverage for wildcard behavior

Reviews (1) · Last reviewed commit: "fix(license): let a wildcard allowed_fea..."

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit c9ed764. Configure here.

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@mateo-berri
mateo-berri merged commit 751fdde into stable/1.101.x Sep 17, 2026
48 of 52 checks passed
@mateo-berri
mateo-berri deleted the litellm_cherrypick_1_101_x branch September 17, 2026 23:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant