Repository navigation
feat(router): meter auto-router tier and prompt customization against the auto_router license feature - #39674
Conversation
|
@greptileai can you reply here with a score and why that score? |
|
bugbot run |
|
Confidence score: 4.5/5 The change is well covered and the enforcement path is coherent: I’m not giving 5/5 because the PR is a fairly broad cross-layer change and the full required CI suite is not yet marked complete in the PR checklist; remaining confidence depends on those CI/lint/type/schema checks. |
Greptile SummaryThis PR generalizes the existing heuristic-v2 license ceiling into per-capability auto-router metering.
Confidence Score: 5/5The PR appears safe to merge; the rebased prompt fields are covered by the same customization limit without introducing a second-slot bypass. No actionable issue remains.
|
| Filename | Overview |
|---|---|
| litellm/router_utils/auto_router_model_naming.py | Defines mutually exclusive gated capabilities and matching in-process and SQL predicates, including both newly supported operator prompt fields. |
| litellm/proxy/management_endpoints/model_management_endpoints.py | Enforces capability-specific limits under an advisory transaction lock while excluding the row being updated. |
| litellm/proxy/proxy_server.py | Fails startup when config-defined routers exceed any unlicensed capability limit and injects the live license resolver into Router instances. |
| litellm/router.py | Checks the claimed capability during deployment registration while preserving rollback behavior for previously serving deployments. |
| litellm/proxy/auth/litellm_license.py | Exposes a per-capability limit that becomes unlimited only when the signed license includes the auto_router feature. |
| tests/test_litellm/router_utils/test_auto_router_model_naming.py | Covers classification prompt and example metering, free presets, SQL predicate construction, and capability exclusivity. |
| tests/test_litellm/proxy/management_endpoints/test_model_management_endpoints.py | Exercises shared-slot accounting across customization forms and transactional refusal behavior. |
| tests/test_litellm/router_strategy/test_complexity_router.py | Verifies accepted prompt configurations and validator constraints that prevent overlapping capability claims. |
Reviews (3): Last reviewed commit: "feat(router): meter auto-router tier and..." | Re-trigger Greptile
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
ac3f42e to
7c58bc6
Compare
|
@greptileai can you reply here with a score and why that score? |
|
bugbot run |
|
Score: 4/5. This is a strong, well-tested implementation. The capability abstraction keeps the in-process predicate, SQL predicate, limit message, and remedy together; separate per-capability counts correctly allow one I’m not giving 5/5 because the implementation is broader than the stated scope: it adds a third gated capability ( |
|
@greptileai both intentional: the third capability was requested, and the renamed kwarg is unreleased and proxy-injected only. Body documents both now. Can you re-score with why? |
7c58bc6 to
5d61a07
Compare
|
@greptileai clarified product invariant: custom tier definitions and custom classifier prompts share one unlicensed slot; heuristic_v2 stays separate. Updated single commit at 5d61a07, tests and live proof. Can you reply with a score and why? |
|
bugbot run |
|
Score: 5/5. The clarified invariant is implemented consistently: |
5d61a07 to
8c944ce
Compare
|
@greptileai #39688 added classification_prompt and classification_examples on built-in-tier routers; the customization predicate now covers both. Rebased, tests and live proof added. Score and why? |
|
bugbot run |
PR overviewAll previously flagged issues have been addressed. No open security concerns remain on this pull request. Security reviewNo open security issues remain on this pull request. Fixed/addressed: 1 · PR risk: 0/10 |
8c944ce to
7970331
Compare
|
Veria finding fixed in 7970331. The gate now classifies the effective decrypted model at every write-side decision, so only complexity routers claim slots. |
|
@greptileai Veria's model-scope finding is fixed in 7970331. Please reply with a score and why. |
|
bugbot run |
|
Score: 5/5. The model-scope finding is fixed. |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 7970331. Configure here.
… the auto_router license feature Generalizes the heuristic_v2 ceiling from #39468 into a capability table whose records own their in-process predicate, SQL spelling and refusal wording. The existing heuristic_v2 capability keeps its own one-router ceiling. A single customization capability combines operator-defined tier definitions with every operator-written part of the classifier prompt. The prompt half only applies to classifier types that call an LLM. The shipped default prompt, classification rubric presets, tier-label renames and tier model choices remain ungated. Scope every enforcement point to actual complexity routers. A model-less PATCH or legacy update now decrypts the stored model before accepting strategy-router settings, so a regular model cannot acquire a router config or spend a license slot. Under the existing advisory lock, the cross-pod candidate query returns only model scalars and the count decrypts and classifies them in process; old non-router rows carrying a capability-shaped config no longer block a real complexity router. The signed auto_router license feature makes both ceilings unlimited.
7970331 to
e04e96d
Compare
test_no_linear_scans_in_router: BerriAI#39674 renamed heuristic_v2_router_limit_violation to auto_router_capability_violation, so the allowlist entry stopped matching and the same admin-only scan tripped the static check. Rename the entry to follow it. tableScrolling.spec.ts: 9ba6cab (LIT-4738) gave the Tags and Model Hub tables client-side pagination at 25 rows, so the 40 seeded rows no longer render on one page. Select 50 rows per page before counting, as the Logs case already does.
TLDR
Problem this solves:
auto_routerlicense feature can hold unlimited customized auto-routersHow it solves it:
tier_definitionsand any operator-written classifier prompt text as ONE customization slottier_labelsand tier model choices stay freeUser Flow
Before: a proxy admin on a plain license customizes tiers and classifier prompts on as many auto-routers as they like
tier_definitions, and both servetier_definitionsand get 200classifier_llm_config.system_promptand those all get 200 tooAfter: the same admin gets one customized auto-router per capability; the license lifts the caps
tier_definitionsto an existing routersystem_promptreturns 403 naming that capability and pointing at the shipped rubric insteadclassification_rubricpreset, on the built-in tiers, or renaming them withtier_labelsall still create with 200auto_routerin the signed license'sallowed_features, the same configs boot and the same creates return 200 with no capWhat is metered, and what stays free
Two capabilities, each with its own count. A single router can only ever claim one of them, because the config validator already forbids every combination.
heuristic_v2classifier_type: heuristic_v2tier_or_classifier_promptclassifier_llm_config.system_prompt,classification_prompt, orclassification_examples) on a classifier that calls an LLMCustomizing tiers and customizing the classifier prompt share ONE slot, so an operator cannot get a second unlicensed customized router by switching which form of customization they use.
Free, and deliberately so: the shipped default classifier prompt, the
classification_rubricpresets,tier_labelsrenames of the built-in tiers, and choosing which model serves each tier.The three operator prompt fields are covered together because the dashboard prompt editor (#39688) writes opening instructions and calibration examples as their own top-level fields on a built-in-tier router, so gating only
system_promptwould leave that editor ungated.Migration
Router(heuristic_v2_router_limit=...)becomesRouter(auto_router_capability_limit=...), same() -> int | Noneshape. No compatibility shim, because nothing can be depending on the old name yet: it shipped in #39468 earlier today and carries no release tag (git tag --contains df73c623b2is empty), the proxy is its only caller, andROUTER_SETTINGS_MANAGED_OUTSIDE_CONFIGblocks it fromrouter_settingsin config.yaml. Docs row renamed in litellm-docs#1179.Relevant issues
Follow-up to #39468
Linear ticket
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Screenshots / Proof of Fix
Shared setup: proxy from source on port 4610 against a dedicated Postgres, master key auth, no
auto_routerlicense feature unless the case says so.$CTis a complexity-router config withclassifier_type: llmand a two-tiertier_definitionsset (routine/hard).Before (ab0478f)
config.yaml holding two custom-tier auto-routers boots
python litellm/proxy/proxy_cli.py --config /tmp/tierlic/config.yaml --port 4610with twotier_definitionsrouters in model_listcurl -s -H "Authorization: Bearer $K" http://127.0.0.1:4610/v1/models->['cheap-direct', 'custom-tiers-a', 'custom-tiers-b', 'premium-direct']/model/new accepts a third and fourth custom-tier router
curl -X POST http://127.0.0.1:4610/model/new ... complexity_router_config: $CT-> HTTP 200/v1/modelsnow lists four custom-tier routerscontrol: the heuristic_v2 ceiling on the same rig does fire
/model/newcalls withclassifier_type: heuristic_v2-> first 200, second 403 "At most 1 auto-router(s) with classifier_type 'heuristic_v2' ..."After (7970331)
config.yaml holding two custom-tier auto-routers refuses to boot
ValueError: config.yaml model_list: At most 1 auto-router(s) with operator-defined tier_definitions can be registered but this would make 2. Keep the built-in tiers for this router or remove an existing router with tier_definitions. A LiteLLM license with the 'auto_router' feature lifts the limit.tier and prompt customizations share one slot, the shipped prompt and rubric presets stay free
tier_definitionsrouter and oneheuristic_v2router ->/v1/modelsserves bothPOST /model/newwith a router carrying its ownclassifier_llm_config.system_prompt-> HTTP 403 "At most 1 auto-router(s) with operator-defined tier_definitions or a custom classifier system_prompt ... Use the shipped tiers and classifier prompt for this router ..."POST /model/newwith a second custom-tier router -> HTTP 403 with the same messagePOST /model/newwithclassification_rubric: agenticand built-in tiers -> HTTP 200POST /model/newwith no prompt fields at all (shipped default) -> HTTP 200the dashboard prompt editor's own fields (#39688) are covered on built-in-tier routers
POST /model/newwithclassification_prompt: "Grade by data sensitivity"on built-in tiers -> HTTP 403 "... operator-written classifier prompt ..."POST /model/newwithclassification_examples: '- "reset my password" -> SIMPLE'on built-in tiers -> HTTP 403 with the same messagePATCH /model/{id}/updateaddingclassification_examplesto a plain router -> HTTP 403, stored config still has neither fieldPOST /model/newwithclassification_rubric: agentic-> HTTP 200; with no prompt fields at all -> HTTP 200heuristic_v2 keeps its own slot beside the customization slot
heuristic_v2router and one custom-tier router serving together under a limit of one eachPOST /model/newwith a second heuristic_v2 router -> HTTP 403 "At most 1 auto-router(s) with classifier_type 'heuristic_v2' ..."POST /model/newwithtier_labels: {"SIMPLE": "Cheap", "MEDIUM": "Standard"}and built-in tiers -> HTTP 200PATCH /model/{id}/updateaddingtier_definitionsto a plain router -> HTTP 403, router keeps serving its stored confignon-complexity models cannot claim a customization slot
openai/gpt-4o-minideploymentPATCH /model/{id}/updateandPOST /model/updatebodies containing a valid custom-tier config -> each HTTP 400 naming the missingauto_router/prefixconcurrent creates cannot race past the shared slot, even mixing forms
POST /model/newcalls, 3 custom-tier and 3 custom-prompt, against a proxy holding no customization1x200 5x403; customized rows persisted -> 1the auto_router license feature lifts the cap
allowed_features: ["auto_router"], boot the two-router config from Before -> serves bothPOST /model/newwith a third custom-tier router -> HTTP 200Type
🆕 New Feature
Caveats (if any)
Low
litellm_params.modelin SQL, so it selects only candidate model scalars and decrypts/classifies them under the existing advisory lockFinal Attestation
Note
Medium Risk
Changes license-gated deployment limits across proxy startup, DB writes, and router registration; misclassification could wrongly allow or deny enterprise features, though behavior is heavily tested and mirrors the prior heuristic_v2 pattern.
Overview
Generalizes the heuristic_v2-only license ceiling into per-capability metering for complexity auto-routers. Without the signed
auto_routerfeature, the proxy allows one router per capability:heuristic_v2(unchanged) and a shared customization slot for operatortier_definitionsor any operator-written classifier prompt (system_prompt,classification_prompt,classification_examples). Shipped rubrics, default prompts, andtier_labelsstay ungated.GatedAutoRouterCapabilitycentralizes detection, SQL predicates, and error text; enforcement runs at config.yaml startup, Router registration, and model add/patch/update via the renamedauto_router_capability_limithook (replacingheuristic_v2_router_limit). DB slot accounting uses capability-specific SQL and decrypts stored models under the existing advisory lock.Write validation now merges effective model + config (decrypting at-rest
modelwhen patches omit it), blocking model-less patches that attach router configs to regular deployments before they could consume a slot.Reviewed by Cursor Bugbot for commit 7970331. Bugbot is set up for automated code reviews on this repo. Configure here.