Skip to content

fix(license): backport the wildcard license auto_router grant to rc/1.102.0 (#41684) - #41701

Merged
mateo-berri merged 1 commit into
rc/1.102.0from
litellm_cherrypick_rc_1_102_0
Sep 18, 2026
Merged

mateo-berri merged 1 commit into
rc/1.102.0from
litellm_cherrypick_rc_1_102_0

Conversation

@mateo-berri

@mateo-berri mateo-berri commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

How it solves it:

Backport notes

  1. Cherry-picked from main: c2fbb11, the single commit of fix(license): let a wildcard allowed_features license grant the auto_router feature #41684 (main lands PRs as merge commits, so there is no squash commit to pick)
  2. tests/test_litellm/proxy/auth/test_litellm_license.py applied clean
  3. litellm/proxy/auth/litellm_license.py conflicted only in the docstring of auto_router_capability_limit, whose wording on this line predates main's; the resolution takes fix(license): let a wildcard allowed_features license grant the auto_router feature #41684's docstring, and the whole region from AUTO_ROUTER_LICENSE_FEATURE down to the method's return 1 is identical to main
  4. No version bump: the release cut from this line is a separate step, the same way fix(responses): backport request-param leak fixes to rc/1.102.0 (#41018, #41141, #41144) #41342 landed here and fix(responses): backport mid-stream content_policy_violation fallback routing to stable/1.100.x #41208 landed on stable/1.100.x
  5. The design carries over unchanged from fix(license): let a wildcard allowed_features license grant the auto_router feature #41684; every caller of auto_router_capability_limit on this line (the boot-time validation, the two Router constructions, and POST /model/new) is also a caller on main, so the main-side risk check applies here

User Flow

Before: an enterprise operator upgrading to v1.102.0-rc.2 with two custom auto-routers cannot start the proxy at all, even though their license covers every feature

  1. They set LITELLM_LICENSE to their enterprise key, whose decoded payload reads "allowed_features": ["*"], and keep a config.yaml with two auto_router/complexity_router deployments that each define their own tier_definitions
  2. They start the v1.102.0-rc.2 proxy (litellm --config config.yaml)
  3. Startup aborts with ValueError: config.yaml model_list: At most 1 auto-router(s) with operator-defined tier_definitions or an operator-written classifier prompt can be registered but this would make 2. ... A LiteLLM license with the 'auto_router' feature lifts the limit. and the process exits
  4. GET http://litellm-domain/health/liveliness gets no answer: the connection fails or times out because no worker ever starts serving

After: the same operator starts the proxy built from this line and both routers serve traffic

  1. They set LITELLM_LICENSE to the same enterprise key ("allowed_features": ["*"]) and keep the same config.yaml with two custom auto-routers
  2. They start the proxy with this fix (litellm --config config.yaml)
  3. The proxy boots and GET http://litellm-domain/health/liveliness returns "I'm alive!"
  4. GET http://litellm-domain/health/license returns 200 with "allowed_features": ["*"]
  5. POST http://litellm-domain/v1/chat/completions with "model": "support-router" and again with "model": "engineering-router" each return 200 with a real completion

Relevant issues

Backport of #41684 (main)

Affected release

regression in v1.101.0 (first in v1.101.0-rc.1, from #39674 on top of #39468), still present in v1.102.0-rc.2; this PR carries the fix onto rc/1.102.0 ahead of the v1.102.0 stable cut

Linear ticket

Resolves LIT-8019

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Every leg is a live proxy booted from a detached worktree at the named commit with its own venv and 2 uvicorn workers on a random free port, no database. The license is read from the LITELLM_LICENSE env var by each worker on its own, so there is no state shared between workers to cross-check, and each chat request is sent twice. The routers call OpenAI for real (gpt-5.4-mini classifies and answers the easy tier, gpt-5.6 the hard one). The three license keys are offline-signed test keys that differ only in allowed_features; their payloads carry max_users: 100, max_teams: 5, and an expiry of 2027-07-06. Chat responses are trimmed with jq to the model, the content, and the token count

config.yaml, the same file on every leg:

model_list:
  - model_name: gpt-5.4-mini
    litellm_params:
      model: openai/gpt-5.4-mini
      api_key: os.environ/OPENAI_API_KEY
  - model_name: gpt-5.6
    litellm_params:
      model: openai/gpt-5.6
      api_key: os.environ/OPENAI_API_KEY
  - model_name: support-router
    litellm_params:
      model: auto_router/complexity_router
      complexity_router_config:
        classifier_type: llm
        classifier_llm_config:
          model: gpt-5.4-mini
        tier_definitions:
          - name: routine
            description: short factual answers and routine drafting
          - name: hard
            description: multi-step reasoning or analysis
        tiers:
          routine: gpt-5.4-mini
          hard: gpt-5.6
        fallback_tier: routine
  - model_name: engineering-router
    litellm_params:
      model: auto_router/complexity_router
      complexity_router_config:
        classifier_type: llm
        classifier_llm_config:
          model: gpt-5.4-mini
        tier_definitions:
          - name: quick
            description: one-line code questions and lookups
          - name: deep
            description: architecture, debugging across files, or proofs
        tiers:
          quick: gpt-5.4-mini
          deep: gpt-5.6
        fallback_tier: quick

general_settings:
  master_key: sk-1234

Before is this PR's merge base on rc/1.102.0 (3f96566, the commit tagged v1.102.0-rc.2); After is this PR's tip (e71b78f). Each row is a separate live proxy:

Checkpoint Proxy commit License Expected Boot support-router engineering-router
Before this PR 3f96566 ["*"] blocked blocked, ValueError no answer no answer
Before this PR 3f96566 ["auto_router"] boots boots 200 + 200 200 + 200
Before this PR 3f96566 ["sso", "audit_logs"] blocked blocked, ValueError no answer no answer
After this PR e71b78f ["*"] boots boots 200 + 200 200 + 200
After this PR e71b78f ["auto_router"] boots boots 200 + 200 200 + 200
After this PR e71b78f ["sso", "audit_logs"] blocked blocked, ValueError no answer no answer

Before (3f96566)

Wildcard license ("allowed_features": ["*"])

  1. Start the proxy; startup aborts in both workers and nothing ever serves

    export LITELLM_LICENSE="$(cat license_wildcard.txt)"
    litellm --config config.yaml --port 49708 --num_workers 2
    ValueError: config.yaml model_list: At most 1 auto-router(s) with operator-defined tier_definitions or an operator-written classifier prompt can be registered but this would make 2. Use the shipped tiers and classifier prompt for this router or remove an existing router with tier_definitions or its own classifier prompt. A LiteLLM license with the 'auto_router' feature lifts the limit.
    ERROR:    Application startup failed. Exiting.
    
  2. GET /health/liveliness gets no answer

    $ curl -sS -m 5 http://localhost:49708/health/liveliness
    curl: (7) Failed to connect to localhost port 49708 after 3937 ms: Couldn't connect to server
    

License naming the feature ("allowed_features": ["auto_router"])

  1. Start the proxy; both workers come up and it serves traffic

    export LITELLM_LICENSE="$(cat license_control_auto_router.txt)"
    litellm --config config.yaml --port 25048 --num_workers 2
    
  2. GET /health/liveliness returns 200

    $ curl -s -w '\nHTTP %{http_code}\n' http://localhost:25048/health/liveliness
    "I'm alive!"
    HTTP 200
    
  3. GET /health/license returns 200

    $ curl -s -w '\nHTTP %{http_code}\n' http://localhost:25048/health/license -H 'Authorization: Bearer sk-1234'
    {"has_license":true,"license_type":"enterprise","expiration_date":"2027-07-06","allowed_features":["auto_router"],"limits":{"max_users":100,"max_teams":5}}
    HTTP 200
    
  4. POST /v1/chat/completions with "model": "support-router", sent twice; both return 200 with a real completion

    $ curl -s -o out.json -w 'HTTP %{http_code}\n' http://localhost:25048/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"model":"support-router","messages":[{"role":"user","content":"What year did the first moon landing happen? One line."}]}'; jq -c '{model, content: .choices[0].message.content, total_tokens: .usage.total_tokens}' out.json
    HTTP 200
    {"model":"support-router","content":"1969","total_tokens":23}
    $ (same request again)
    HTTP 200
    {"model":"support-router","content":"1969.","total_tokens":24}
    
  5. POST /v1/chat/completions with "model": "engineering-router", sent twice; both return 200 with a real completion

    $ curl -s -o out.json -w 'HTTP %{http_code}\n' http://localhost:25048/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"model":"engineering-router","messages":[{"role":"user","content":"In Python, what does dict.setdefault do? One line."}]}'; jq -c '{model, content: .choices[0].message.content, total_tokens: .usage.total_tokens}' out.json
    HTTP 200
    {"model":"engineering-router","content":"`dict.setdefault(key, default)` returns the value for `key` if it exists; otherwise it inserts `key` with `default` and returns `default`.","total_tokens":54}
    $ (same request again)
    HTTP 200
    {"model":"engineering-router","content":"`dict.setdefault(key, default)` returns the value for `key` if it exists; otherwise it inserts `key` with `default` and returns that default.","total_tokens":54}
    

License with other features only ("allowed_features": ["sso", "audit_logs"])

  1. Start the proxy; startup aborts in both workers and nothing ever serves

    export LITELLM_LICENSE="$(cat license_other_features.txt)"
    litellm --config config.yaml --port 50257 --num_workers 2
    ValueError: config.yaml model_list: At most 1 auto-router(s) with operator-defined tier_definitions or an operator-written classifier prompt can be registered but this would make 2. Use the shipped tiers and classifier prompt for this router or remove an existing router with tier_definitions or its own classifier prompt. A LiteLLM license with the 'auto_router' feature lifts the limit.
    ERROR:    Application startup failed. Exiting.
    
  2. GET /health/liveliness gets no answer

    $ curl -sS -m 5 http://localhost:50257/health/liveliness
    curl: (28) Connection timed out after 5002 milliseconds
    

After (e71b78f)

Wildcard license ("allowed_features": ["*"])

  1. Start the proxy; both workers come up and it serves traffic

    export LITELLM_LICENSE="$(cat license_wildcard.txt)"
    litellm --config config.yaml --port 40291 --num_workers 2
    
  2. GET /health/liveliness returns 200

    $ curl -s -w '\nHTTP %{http_code}\n' http://localhost:40291/health/liveliness
    "I'm alive!"
    HTTP 200
    
  3. GET /health/license returns 200

    $ curl -s -w '\nHTTP %{http_code}\n' http://localhost:40291/health/license -H 'Authorization: Bearer sk-1234'
    {"has_license":true,"license_type":"enterprise","expiration_date":"2027-07-06","allowed_features":["*"],"limits":{"max_users":100,"max_teams":5}}
    HTTP 200
    
  4. POST /v1/chat/completions with "model": "support-router", sent twice; both return 200 with a real completion

    $ curl -s -o out.json -w 'HTTP %{http_code}\n' http://localhost:40291/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"model":"support-router","messages":[{"role":"user","content":"What year did the first moon landing happen? One line."}]}'; jq -c '{model, content: .choices[0].message.content, total_tokens: .usage.total_tokens}' out.json
    HTTP 200
    {"model":"support-router","content":"1969","total_tokens":23}
    $ (same request again)
    HTTP 200
    {"model":"support-router","content":"1969","total_tokens":23}
    
  5. POST /v1/chat/completions with "model": "engineering-router", sent twice; both return 200 with a real completion

    $ curl -s -o out.json -w 'HTTP %{http_code}\n' http://localhost:40291/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"model":"engineering-router","messages":[{"role":"user","content":"In Python, what does dict.setdefault do? One line."}]}'; jq -c '{model, content: .choices[0].message.content, total_tokens: .usage.total_tokens}' out.json
    HTTP 200
    {"model":"engineering-router","content":"`dict.setdefault(key, default)` returns the value for `key` if it exists; otherwise it inserts `key` with `default` and returns `default`.","total_tokens":54}
    $ (same request again)
    HTTP 200
    {"model":"engineering-router","content":"`dict.setdefault(key, default)` returns `dict[key]` if the key exists; otherwise it inserts `key` with `default` and returns `default`.","total_tokens":54}
    

License naming the feature ("allowed_features": ["auto_router"])

  1. Start the proxy; both workers come up and it serves traffic

    export LITELLM_LICENSE="$(cat license_control_auto_router.txt)"
    litellm --config config.yaml --port 27877 --num_workers 2
    
  2. GET /health/liveliness returns 200

    $ curl -s -w '\nHTTP %{http_code}\n' http://localhost:27877/health/liveliness
    "I'm alive!"
    HTTP 200
    
  3. GET /health/license returns 200

    $ curl -s -w '\nHTTP %{http_code}\n' http://localhost:27877/health/license -H 'Authorization: Bearer sk-1234'
    {"has_license":true,"license_type":"enterprise","expiration_date":"2027-07-06","allowed_features":["auto_router"],"limits":{"max_users":100,"max_teams":5}}
    HTTP 200
    
  4. POST /v1/chat/completions with "model": "support-router", sent twice; both return 200 with a real completion

    $ curl -s -o out.json -w 'HTTP %{http_code}\n' http://localhost:27877/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"model":"support-router","messages":[{"role":"user","content":"What year did the first moon landing happen? One line."}]}'; jq -c '{model, content: .choices[0].message.content, total_tokens: .usage.total_tokens}' out.json
    HTTP 200
    {"model":"support-router","content":"1969.","total_tokens":24}
    $ (same request again)
    HTTP 200
    {"model":"support-router","content":"1969","total_tokens":23}
    
  5. POST /v1/chat/completions with "model": "engineering-router", sent twice; both return 200 with a real completion

    $ curl -s -o out.json -w 'HTTP %{http_code}\n' http://localhost:27877/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"model":"engineering-router","messages":[{"role":"user","content":"In Python, what does dict.setdefault do? One line."}]}'; jq -c '{model, content: .choices[0].message.content, total_tokens: .usage.total_tokens}' out.json
    HTTP 200
    {"model":"engineering-router","content":"`dict.setdefault(key, default)` returns the value for `key` if it exists; otherwise it inserts `key` with `default` and returns `default`.","total_tokens":54}
    $ (same request again)
    HTTP 200
    {"model":"engineering-router","content":"`dict.setdefault(key, default)` returns the value for `key` if it exists; otherwise it inserts `key` with `default` and returns `default`.","total_tokens":54}
    

License with other features only ("allowed_features": ["sso", "audit_logs"])

  1. Start the proxy; startup aborts in both workers and nothing ever serves

    export LITELLM_LICENSE="$(cat license_other_features.txt)"
    litellm --config config.yaml --port 27053 --num_workers 2
    ValueError: config.yaml model_list: At most 1 auto-router(s) with operator-defined tier_definitions or an operator-written classifier prompt can be registered but this would make 2. Use the shipped tiers and classifier prompt for this router or remove an existing router with tier_definitions or its own classifier prompt. A LiteLLM license with the 'auto_router' feature lifts the limit.
    ERROR:    Application startup failed. Exiting.
    
  2. GET /health/liveliness gets no answer

    $ curl -sS -m 5 http://localhost:27053/health/liveliness
    curl: (7) Failed to connect to localhost port 27053 after 2076 ms: Couldn't connect to server
    

Notes from the run:

Type

🐛 Bug Fix

Caveats (if any)

Low

  • Only the exact * entry is a wildcard; glob patterns like auto_* still match nothing
    • The license generator only ever emits the literal *, so nothing in the field uses patterns
  • GET /health/license keeps reporting allowed_features verbatim, so a wildcard license still shows ["*"] there
  • A bare-string allowed_features is read as a one-item list, the same way GET /health/license reports it
  • A license verified through the API carries no feature list, so it keeps the limit, as before
  • The POST /model/new path of the same gate was proven on fix(license): let a wildcard allowed_features license grant the auto_router feature #41684 (second router 403 before, 200 after) and not re-run on this line; the gate's code region is identical to main's
  • CircleCI local_testing_part1 (job), local_testing_part2 (job), and llm_translation_testing (job) are red only on the 22 Together AI tests that call the serverless openai/gpt-oss-20b the provider has since withdrawn (Unable to access non-serverless model); main moved those tests off that model in 1aa2e19, 8c046e1, and b478131, none of which is on rc/1.102.0, main's latest pipeline passes all three jobs, this PR touches no Together AI path, and rc/1.102.0 has no required status checks

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Note

Low Risk
Narrow change to offline license feature parsing for auto-router limits; API-verified licenses without feature lists behave as before.

Overview
Fixes enterprise proxies that fail startup when allowed_features is ["*"] (the default) but the auto-router gate only recognized the literal auto_router entry.

Adds LicenseCheck.grants_feature, which treats either the requested feature or the LICENSE_ALL_FEATURES (*) wildcard in signed (airgapped) license data as a grant, including when allowed_features is a bare string. auto_router_capability_limit now delegates to that helper instead of inlining list checks, so wildcard licenses lift the one-router-per-capability cap the same as explicitly licensed auto_router.

Tests cover list and string wildcards, mixed feature lists, and signed licenses that name only other features (limit stays at 1).

Reviewed by Cursor Bugbot for commit e71b78f. Bugbot is set up for automated code reviews on this repo. Configure here.

…router feature

Backport of #41684 to rc/1.102.0.
Cherry-picked from c2fbb11 (litellm_wildcard_license_auto_router).
@greptile-apps

greptile-apps Bot commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge, with wildcard handling limited to verified signed-license feature data and consistent behavior across capability-limit consumers

Summary

This backport recognizes the signed-license "*" feature claim when evaluating auto-router entitlements

  • Adds a shared feature-grant helper that handles explicit and wildcard grants
  • Uses that helper to lift auto-router capability limits for wildcard offline licenses
  • Adds direct and cryptographically signed regression coverage for wildcard, explicit, unrelated, missing, and null feature claims

Reviews (1) · Last reviewed commit: "fix(license): let a wildcard allowed_fea..."

@mateo-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit e71b78f. Configure here.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant