Skip to content

feat(router): meter auto-router tier and prompt customization against the auto_router license feature - #39674

Merged
tin-berri merged 1 commit into
litellm_internal_stagingfrom
litellm_autorouter_tier_license
Sep 5, 2026
Merged

tin-berri merged 1 commit into
litellm_internal_stagingfrom
litellm_autorouter_tier_license

Conversation

@tin-berri

@tin-berri tin-berri commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

How it solves it:

  • Generalizes feat(router): limit heuristic_v2 auto-routers to one without the auto_router license feature #39468's machinery into a capability table (predicate + SQL + message per capability)
  • Meters operator-defined tier_definitions and any operator-written classifier prompt text as ONE customization slot
  • Without the license: one heuristic_v2 router, and one router customizing its tiers or classifier prompt
  • The shipped default prompt, the rubric presets, tier_labels and tier model choices stay free

User Flow

Before: a proxy admin on a plain license customizes tiers and classifier prompts on as many auto-routers as they like

  1. They boot a config.yaml holding two auto-routers that each define their own tiers via tier_definitions, and both serve
  2. They POST http://localhost:4000/model/new with a third auto-router carrying tier_definitions and get 200
  3. They do the same with auto-routers carrying their own classifier_llm_config.system_prompt and those all get 200 too
  4. Nothing anywhere tells them these are licensed features

After: the same admin gets one customized auto-router per capability; the license lifts the caps

  1. Booting the same two-router config.yaml fails fast with "At most 1 auto-router(s) with operator-defined tier_definitions can be registered but this would make 2. Keep the built-in tiers for this router or remove an existing router with tier_definitions. A LiteLLM license with the 'auto_router' feature lifts the limit."
  2. With one custom-tier router live, POST http://localhost:4000/model/new with a second one returns 403 with the same message; so does a PATCH that adds tier_definitions to an existing router
  3. A second auto-router carrying its own classifier system_prompt returns 403 naming that capability and pointing at the shipped rubric instead
  4. Routers on the shipped default prompt, on a classification_rubric preset, on the built-in tiers, or renaming them with tier_labels all still create with 200
  5. With auto_router in the signed license's allowed_features, the same configs boot and the same creates return 200 with no cap

What is metered, and what stays free

Two capabilities, each with its own count. A single router can only ever claim one of them, because the config validator already forbids every combination.

Capability Claimed when Gated
heuristic_v2 classifier_type: heuristic_v2 yes, shipped in #39468
tier_or_classifier_prompt the config defines its own tier set, OR the operator wrote any part of the classifier prompt (classifier_llm_config.system_prompt, classification_prompt, or classification_examples) on a classifier that calls an LLM yes, new here

Customizing tiers and customizing the classifier prompt share ONE slot, so an operator cannot get a second unlicensed customized router by switching which form of customization they use.

Free, and deliberately so: the shipped default classifier prompt, the classification_rubric presets, tier_labels renames of the built-in tiers, and choosing which model serves each tier.

The three operator prompt fields are covered together because the dashboard prompt editor (#39688) writes opening instructions and calibration examples as their own top-level fields on a built-in-tier router, so gating only system_prompt would leave that editor ungated.

Migration

Router(heuristic_v2_router_limit=...) becomes Router(auto_router_capability_limit=...), same () -> int | None shape. No compatibility shim, because nothing can be depending on the old name yet: it shipped in #39468 earlier today and carries no release tag (git tag --contains df73c623b2 is empty), the proxy is its only caller, and ROUTER_SETTINGS_MANAGED_OUTSIDE_CONFIG blocks it from router_settings in config.yaml. Docs row renamed in litellm-docs#1179.

Relevant issues

Follow-up to #39468

Linear ticket

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Screenshots / Proof of Fix

Shared setup: proxy from source on port 4610 against a dedicated Postgres, master key auth, no auto_router license feature unless the case says so. $CT is a complexity-router config with classifier_type: llm and a two-tier tier_definitions set (routine/hard).

Before (ab0478f)

config.yaml holding two custom-tier auto-routers boots

  1. python litellm/proxy/proxy_cli.py --config /tmp/tierlic/config.yaml --port 4610 with two tier_definitions routers in model_list
  2. curl -s -H "Authorization: Bearer $K" http://127.0.0.1:4610/v1/models -> ['cheap-direct', 'custom-tiers-a', 'custom-tiers-b', 'premium-direct']

/model/new accepts a third and fourth custom-tier router

  1. curl -X POST http://127.0.0.1:4610/model/new ... complexity_router_config: $CT -> HTTP 200
  2. Same call with another name -> HTTP 200; /v1/models now lists four custom-tier routers

control: the heuristic_v2 ceiling on the same rig does fire

  1. Two /model/new calls with classifier_type: heuristic_v2 -> first 200, second 403 "At most 1 auto-router(s) with classifier_type 'heuristic_v2' ..."

After (7970331)

config.yaml holding two custom-tier auto-routers refuses to boot

  1. Same launch as Before
  2. Startup fails: ValueError: config.yaml model_list: At most 1 auto-router(s) with operator-defined tier_definitions can be registered but this would make 2. Keep the built-in tiers for this router or remove an existing router with tier_definitions. A LiteLLM license with the 'auto_router' feature lifts the limit.

tier and prompt customizations share one slot, the shipped prompt and rubric presets stay free

  1. Boot a config with one tier_definitions router and one heuristic_v2 router -> /v1/models serves both
  2. POST /model/new with a router carrying its own classifier_llm_config.system_prompt -> HTTP 403 "At most 1 auto-router(s) with operator-defined tier_definitions or a custom classifier system_prompt ... Use the shipped tiers and classifier prompt for this router ..."
  3. POST /model/new with a second custom-tier router -> HTTP 403 with the same message
  4. POST /model/new with classification_rubric: agentic and built-in tiers -> HTTP 200
  5. POST /model/new with no prompt fields at all (shipped default) -> HTTP 200
  6. Booting a config with two customized routers (any mix of tiers and prompts) fails at startup with the same message

the dashboard prompt editor's own fields (#39688) are covered on built-in-tier routers

  1. Same rig, one custom-tier router live
  2. POST /model/new with classification_prompt: "Grade by data sensitivity" on built-in tiers -> HTTP 403 "... operator-written classifier prompt ..."
  3. POST /model/new with classification_examples: '- "reset my password" -> SIMPLE' on built-in tiers -> HTTP 403 with the same message
  4. PATCH /model/{id}/update adding classification_examples to a plain router -> HTTP 403, stored config still has neither field
  5. POST /model/new with classification_rubric: agentic -> HTTP 200; with no prompt fields at all -> HTTP 200

heuristic_v2 keeps its own slot beside the customization slot

  1. Same rig: one heuristic_v2 router and one custom-tier router serving together under a limit of one each
  2. POST /model/new with a second heuristic_v2 router -> HTTP 403 "At most 1 auto-router(s) with classifier_type 'heuristic_v2' ..."
  3. POST /model/new with tier_labels: {"SIMPLE": "Cheap", "MEDIUM": "Standard"} and built-in tiers -> HTTP 200
  4. PATCH /model/{id}/update adding tier_definitions to a plain router -> HTTP 403, router keeps serving its stored config

non-complexity models cannot claim a customization slot

  1. Create an ordinary openai/gpt-4o-mini deployment
  2. Send model-less PATCH /model/{id}/update and POST /model/update bodies containing a valid custom-tier config -> each HTTP 400 naming the missing auto_router/ prefix
  3. Create a genuine custom complexity router afterwards -> HTTP 200, proving the rejected regular row did not consume the slot
  4. A second genuine customized router -> HTTP 403, proving the normal ceiling still binds

concurrent creates cannot race past the shared slot, even mixing forms

  1. 6 parallel POST /model/new calls, 3 custom-tier and 3 custom-prompt, against a proxy holding no customization
  2. Results 1x200 5x403; customized rows persisted -> 1

the auto_router license feature lifts the cap

  1. Sign a test license with allowed_features: ["auto_router"], boot the two-router config from Before -> serves both
  2. POST /model/new with a third custom-tier router -> HTTP 200

Type

🆕 New Feature

Caveats (if any)

Low

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Note

Medium Risk
Changes license-gated deployment limits across proxy startup, DB writes, and router registration; misclassification could wrongly allow or deny enterprise features, though behavior is heavily tested and mirrors the prior heuristic_v2 pattern.

Overview
Generalizes the heuristic_v2-only license ceiling into per-capability metering for complexity auto-routers. Without the signed auto_router feature, the proxy allows one router per capability: heuristic_v2 (unchanged) and a shared customization slot for operator tier_definitions or any operator-written classifier prompt (system_prompt, classification_prompt, classification_examples). Shipped rubrics, default prompts, and tier_labels stay ungated.

GatedAutoRouterCapability centralizes detection, SQL predicates, and error text; enforcement runs at config.yaml startup, Router registration, and model add/patch/update via the renamed auto_router_capability_limit hook (replacing heuristic_v2_router_limit). DB slot accounting uses capability-specific SQL and decrypts stored models under the existing advisory lock.

Write validation now merges effective model + config (decrypting at-rest model when patches omit it), blocking model-less patches that attach router configs to regular deployments before they could consume a slot.

Reviewed by Cursor Bugbot for commit 7970331. Bugbot is set up for automated code reviews on this repo. Configure here.

@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai can you reply here with a score and why that score?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Sep 3, 2026

Copy link
Copy Markdown
Contributor

Confidence score: 4.5/5

The change is well covered and the enforcement path is coherent: tier_definitions and heuristic_v2 have independent capability predicates, limits, SQL counts, and messages; startup validation prevents silently dropped config deployments; model creates/patches use the transaction-scoped advisory lock to prevent races; and the signed auto_router feature correctly lifts the limit. The tests cover capability classification, cross-capability coexistence, tier-label exemptions, startup failures, edits, rollback behavior, licensing, and the concurrent-write path. The available E2E status is also successful.

I’m not giving 5/5 because the PR is a fairly broad cross-layer change and the full required CI suite is not yet marked complete in the PR checklist; remaining confidence depends on those CI/lint/type/schema checks.

@codspeed

codspeed Bot commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_autorouter_tier_license (e04e96d) with litellm_internal_staging (aea5358)

Open in CodSpeed

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@greptile-apps

greptile-apps Bot commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR generalizes the existing heuristic-v2 license ceiling into per-capability auto-router metering.

  • Retains an independent unlicensed slot for heuristic_v2.
  • Adds one shared customization slot for custom tiers, classifier system prompts, classification_prompt, and classification_examples.
  • Keeps built-in tiers, tier labels, default prompts, and shipped rubric presets free.
  • Applies matching checks during configuration startup, Router registration, and transactional model-management writes.
  • Adds coverage for prompt customization introduced by the rebase, predicate parity, capability exclusivity, concurrent writes, and license-based limit removal.

Confidence Score: 5/5

The PR appears safe to merge; the rebased prompt fields are covered by the same customization limit without introducing a second-slot bypass.

No actionable issue remains. classification_prompt and classification_examples are included in both the runtime predicate and persisted-row SQL predicate, valid configurations cannot claim both gated capabilities, and the write path retains transactional cross-pod serialization.

Important Files Changed

Filename Overview
litellm/router_utils/auto_router_model_naming.py Defines mutually exclusive gated capabilities and matching in-process and SQL predicates, including both newly supported operator prompt fields.
litellm/proxy/management_endpoints/model_management_endpoints.py Enforces capability-specific limits under an advisory transaction lock while excluding the row being updated.
litellm/proxy/proxy_server.py Fails startup when config-defined routers exceed any unlicensed capability limit and injects the live license resolver into Router instances.
litellm/router.py Checks the claimed capability during deployment registration while preserving rollback behavior for previously serving deployments.
litellm/proxy/auth/litellm_license.py Exposes a per-capability limit that becomes unlimited only when the signed license includes the auto_router feature.
tests/test_litellm/router_utils/test_auto_router_model_naming.py Covers classification prompt and example metering, free presets, SQL predicate construction, and capability exclusivity.
tests/test_litellm/proxy/management_endpoints/test_model_management_endpoints.py Exercises shared-slot accounting across customization forms and transactional refusal behavior.
tests/test_litellm/router_strategy/test_complexity_router.py Verifies accepted prompt configurations and validator constraints that prevent overlapping capability claims.

Reviews (3): Last reviewed commit: "feat(router): meter auto-router tier and..." | Re-trigger Greptile

@codecov

codecov Bot commented Sep 3, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 98.88889% with 1 line in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
...management_endpoints/model_management_endpoints.py 97.05% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

@tin-berri
tin-berri force-pushed the litellm_autorouter_tier_license branch from ac3f42e to 7c58bc6 Compare September 4, 2026 00:01
@tin-berri tin-berri changed the title feat(router): meter operator-defined auto-router tier sets against the auto_router license feature feat(router): meter operator-customized auto-router tiers and classifier prompts against the auto_router license feature Sep 4, 2026
@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai can you reply here with a score and why that score?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

Score: 4/5.

This is a strong, well-tested implementation. The capability abstraction keeps the in-process predicate, SQL predicate, limit message, and remedy together; separate per-capability counts correctly allow one heuristic_v2 and one custom-tier router to coexist; and the advisory-locked transaction covers concurrent model writes. The tests cover startup validation, add/update/PATCH paths, license behavior, tier-label exclusions, malformed inputs, rollback, and the race case.

I’m not giving 5/5 because the implementation is broader than the stated scope: it adds a third gated capability (classifier_system_prompt) even though the description presents tier_definitions as the second capability. It also renames the public Router kwarg from heuristic_v2_router_limit to auto_router_capability_limit without a compatibility path. If both changes are intentional, the PR description and migration guidance should explicitly document them; otherwise, narrow them before merge.

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai both intentional: the third capability was requested, and the renamed kwarg is unreleased and proxy-injected only. Body documents both now. Can you re-score with why?

@tin-berri
tin-berri force-pushed the litellm_autorouter_tier_license branch from 7c58bc6 to 5d61a07 Compare September 4, 2026 00:48
@tin-berri tin-berri changed the title feat(router): meter operator-customized auto-router tiers and classifier prompts against the auto_router license feature feat(router): meter auto-router tier and prompt customization against the auto_router license feature Sep 4, 2026
@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai clarified product invariant: custom tier definitions and custom classifier prompts share one unlicensed slot; heuristic_v2 stays separate. Updated single commit at 5d61a07, tests and live proof. Can you reply with a score and why?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

Score: 5/5.

The clarified invariant is implemented consistently: tier_definitions and custom classifier system_prompt share one customization slot, while heuristic_v2 has an independent slot. The capability table keeps the in-process predicate, SQL predicate, counting, limit message, and remedy aligned; startup validation prevents silently omitted config deployments; and the transaction-scoped advisory lock closes the concurrent-write race across both customization forms. The implementation also preserves the intended free cases (built-in tiers, tier_labels, rubric presets, and the shipped prompt), correctly lifts limits for the signed auto_router feature, and covers add/update/PATCH, rollback, malformed inputs, coexistence, licensing, startup, and concurrency. The latest commit has successful E2E status, and the previously raised scope/rename concerns are intentional and documented.

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai #39688 added classification_prompt and classification_examples on built-in-tier routers; the customization predicate now covers both. Rebased, tests and live proof added. Score and why?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/proxy/management_endpoints/model_management_endpoints.py
@veria-ai

veria-ai Bot commented Sep 5, 2026 •

Copy link
Copy Markdown
Contributor

PR overview

All previously flagged issues have been addressed. No open security concerns remain on this pull request.

Security review

No open security issues remain on this pull request.

Fixed/addressed: 1 · PR risk: 0/10

@tin-berri

Copy link
Copy Markdown
Contributor Author

Veria finding fixed in 7970331. The gate now classifies the effective decrypted model at every write-side decision, so only complexity routers claim slots.

@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai Veria's model-scope finding is fixed in 7970331. Please reply with a score and why.

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Sep 5, 2026

Copy link
Copy Markdown
Contributor

Score: 5/5.

The model-scope finding is fixed. _effective_model() now decrypts the stored model when a PATCH/update omits litellm_params.model, and _strategy_router_write_violation() evaluates that effective model before applying complexity-router validation or capability metering. This prevents a regular deployment from attaching a capability-shaped complexity_router_config and consuming a slot, while keeping genuine auto_router/complexity_router rows in scope. The DB candidate count also filters candidates by the decrypted/classified model rather than trusting configuration alone. The added tests cover model-less PATCH and legacy update rejection, encrypted stored router models, and the no-write/no-slot-consumption behavior.

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 7970331. Configure here.

… the auto_router license feature

Generalizes the heuristic_v2 ceiling from #39468 into a capability table whose
records own their in-process predicate, SQL spelling and refusal wording. The
existing heuristic_v2 capability keeps its own one-router ceiling. A single
customization capability combines operator-defined tier definitions with every
operator-written part of the classifier prompt. The prompt half only applies to
classifier types that call an LLM. The shipped default prompt, classification
rubric presets, tier-label renames and tier model choices remain ungated.

Scope every enforcement point to actual complexity routers. A model-less PATCH
or legacy update now decrypts the stored model before accepting strategy-router
settings, so a regular model cannot acquire a router config or spend a license
slot. Under the existing advisory lock, the cross-pod candidate query returns
only model scalars and the count decrypts and classifies them in process; old
non-router rows carrying a capability-shaped config no longer block a real
complexity router. The signed auto_router license feature makes both ceilings
unlimited.
@tin-berri
tin-berri force-pushed the litellm_autorouter_tier_license branch from 7970331 to e04e96d Compare September 5, 2026 05:21
@tin-berri
tin-berri enabled auto-merge (squash) September 5, 2026 09:18
@tin-berri
tin-berri merged commit d0d09e5 into litellm_internal_staging Sep 5, 2026
187 checks passed
@tin-berri
tin-berri deleted the litellm_autorouter_tier_license branch September 5, 2026 16:51
pull Bot pushed a commit to TKaxv-7S/litellm that referenced this pull request Sep 5, 2026
test_no_linear_scans_in_router: BerriAI#39674 renamed heuristic_v2_router_limit_violation
to auto_router_capability_violation, so the allowlist entry stopped matching and the
same admin-only scan tripped the static check. Rename the entry to follow it.

tableScrolling.spec.ts: 9ba6cab (LIT-4738) gave the Tags and Model Hub tables
client-side pagination at 25 rows, so the 40 seeded rows no longer render on one
page. Select 50 rows per page before counting, as the Logs case already does.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants