Repository navigation
fix(auto_router): derive tier definitions in prompt editor - #39688
Conversation
|
@greptileai can you reply here with a score and why that score? |
|
bugbot run |
Greptile SummaryThe PR introduces opening-only prompt customization for built-in complexity-router tiers while deriving tier labels, criteria, trust-boundary text, and closing instructions from structured configuration.
Confidence Score: 5/5The PR appears safe to merge, with the live prompt, preview API, configuration validation, and dashboard persistence paths aligned around the derived-tier prompt contract. The new built-in opening path derives active tier labels and criteria from the same router configuration used by classifier output validation, preserves the injection guard, rejects conflicting prompt modes, and is consistently represented by the dashboard builders and preview transport.
|
| Filename | Overview |
|---|---|
| litellm/router_strategy/complexity_router/complexity_router.py | Adds built-in opening-only prompt assembly and uses it before the default full-rubric path while preserving derived labels and trust-boundary text. |
| litellm/router_strategy/complexity_router/config.py | Allows built-in classification_prompt values for LLM classifiers and rejects conflicting legacy whole-prompt configurations. |
| litellm/proxy/management_endpoints/model_management_endpoints.py | Extends prompt preview requests with optional built-in labels and rubric selection while sharing label validation with the existing GET endpoint. |
| ui/litellm-dashboard/src/components/add_model/ClassificationMethodConfig.tsx | Selects the opening-only editor for derived prompts and retains the legacy editor when a whole-prompt override already exists. |
| ui/litellm-dashboard/src/components/add_model/OpeningPromptEditor.tsx | Generalizes the custom-tier editor to preview and save opening-only prompts for both custom and built-in tier sources. |
| ui/litellm-dashboard/src/components/add_model/build_complexity_router_config.ts | Serializes built-in opening prompts only for LLM classifiers without a conflicting legacy system prompt. |
| ui/litellm-dashboard/src/components/edit_auto_router/edit_auto_router_modal.tsx | Keeps classification_prompt under managed edit-state persistence so resets and updates are reflected in saved configuration. |
| ui/litellm-dashboard/src/components/networking.tsx | Replaces the custom-tier-only preview helper with a discriminated source supporting either definitions or built-in labels and rubric. |
Reviews (1): Last reviewed commit: "fix(auto_router): derive tier definition..." | Re-trigger Greptile
|
Confidence score: 4/5. The change correctly centralizes built-in tier prompt assembly in The added backend and dashboard tests cover derived criteria, label validation, legacy behavior, reset behavior, and the create/edit flows. I found no blocking correctness issues in the reviewed diff. I’m stopping at 4 rather than 5 because the PR shows only the focused Buildkite status, while the full required CI/CD validation is still pending. |
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit e93300f. Configure here.
|
@greptileai can you reply here with a score and why that score? |
|
Confidence score: 4/5. The PR correctly centralizes built-in prompt assembly in The dashboard changes also cleanly distinguish derived built-in/custom opening prompts from legacy whole-prompt mode, and the added backend/UI tests cover derived criteria, label validation, reset behavior, and create/edit flows. The focused Buildkite status and coverage report are positive. I’m stopping at 4 rather than 5 because the PR’s full required CI/CD checks are still not complete, and the documentation check is currently blocked by a base-branch dependency. |
d19d84b to
70baeb6
Compare
|
@greptileai can you reply here with a score and why that score? |
|
bugbot run |
|
Confidence score: 4/5. The implementation correctly makes built-in prompt assembly derive tier labels, criteria, the injection guard, and closing instructions from the same structured configuration used by classifier output validation. The live classifier path and preview API are aligned, opening-only prompts are validated for the appropriate LLM flow, conflicting legacy whole-prompt overrides are rejected, and existing legacy configurations remain editable. The dashboard changes consistently support create, edit, preview, reset, and persistence for both derived and legacy modes, with focused backend and UI coverage. I found no blocking correctness issues in the reviewed changes. I’m stopping at 4 rather than 5 because the PR’s full required CI/CD validation is not complete yet: the available Buildkite status is still pending, and the documentation check remains dependent on the base-branch documentation change. |
70baeb6 to
8bc4dfb
Compare
|
@greptileai can you reply here with a score and why that score? |
|
bugbot run |
|
Confidence score: 4/5. The change addresses the reported mismatch at the right boundary: built-in tier labels and criteria are now derived by the proxy, while only operator-owned instructions and examples remain editable. The live classifier and preview share the same assembly path, legacy whole-prompt overrides remain supported, and the backend/dashboard tests cover label renames, legacy splitting, validation, reset flows, rubric selection, and preview parity. The existing trust-boundary guard is also retained after operator-controlled sections. I’m not giving 5/5 because this is a relatively broad backend/dashboard change with load-time migration logic; malformed or unusual legacy prompt layouts could still be a compatibility risk. The focused coverage and passing reported e2e status provide good evidence, but the full required CI suite is not yet confirmed. |
8bc4dfb to
46609cf
Compare
|
@greptileai can you reply here with a score and why that score? |
|
bugbot run |
|
Confidence score: 4/5 The change is well-scoped and addresses the core mismatch: built-in routers now derive tier bullets, labels, rubric criteria, trust-boundary text, and closing instructions from structured configuration, while operator-provided instructions and calibration examples remain editable. The preview and live classifier share the same prompt assembly path, legacy whole-prompt overrides remain supported, and the added tests cover renamed tiers, invalid labels, custom-tier behavior, preview parity, normalization limits, and dashboard flows. The previously identified rubric-cancel and load-time prompt-splitting concerns are also addressed in the latest commit. I’m giving 4 rather than 5 because this is a relatively broad cross-layer change (backend, API validation, and dashboard) with 23 files changed, and the PR checklist still shows required CI/CD checks incomplete. The focused tests and available E2E status are positive, but a full score would require all required CI checks and the remaining integration coverage to be confirmed. |
46609cf to
f380917
Compare
|
@greptileai can you reply here with a score and why that score? |
|
bugbot run |
|
Score: 4/5 This is a strong, well-scoped fix. The live router and preview share the same prompt assembly paths, so tier labels, criteria, calibration sections, the trust boundary, and closing line remain consistent after renames. Legacy whole-prompt overrides are preserved, while the new section-based fields have clear mutual-exclusion, classifier-type, normalization, and length validation. The backend and dashboard tests cover derived prompts, preview parity, legacy behavior, validation failures, and create/edit round trips. I’m not giving 5/5 because this is a behavior-sensitive change across 23 files and the required CI/e2e status is still pending. I did not identify a concrete correctness defect in the reviewed changes; assuming the pending suites remain green, this is close to merge-ready. |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit f380917. Configure here.
|
CI settled: 183 of 186 pass, codecov included. The 3 reds are staging-wide tests this PR does not touch, failing the same way on #39770, #39766, #39763 and #39762 @greptileai can you reply here with a score and why that score? |
|
Score: 4/5 This change directly fixes prompt/tier drift by deriving tier definitions, labels, guards, and closing text from the current router configuration while preserving legacy whole-prompt overrides. The coverage is meaningful across derived prompts, legacy behavior, preview validation, and dashboard create/edit flows. CI is also strong: 183/186 checks pass, with 86.76% diff coverage against a 79.05% target. The three failures are staging-wide, reproduce on multiple unrelated PRs, and do not touch this change, so I do not consider them regressions here. I’m not giving 5/5 because the full required suite is not green and this changes live classifier prompt assembly, which retains some integration risk despite the focused tests. |
… the auto_router license feature Generalizes the heuristic_v2 ceiling from #39468 into a capability table whose records own their in-process predicate, SQL spelling and refusal wording. The existing heuristic_v2 capability keeps its own one-router ceiling. A single customization capability combines operator-defined tier_definitions with every operator-written part of the classifier prompt: a whole replacement system_prompt, replacement opening instructions (classification_prompt) and replacement calibration examples (classification_examples, the fields the dashboard prompt editor from #39688 writes on built-in-tier routers). So an unlicensed proxy can have at most one router that customizes either form. The prompt half only applies to classifier types that call an LLM. The shipped default prompt, classification rubric presets, tier_labels renames and tier model choices remain ungated. The signed auto_router license feature makes both ceilings unlimited.
TLDR
Problem this solves:
How it solves it:
User Flow
Before: an admin can only replace the complete classifier prompt, so a tier rename can leave the prompt describing labels the classifier cannot return
Tiers:section, then rename SIMPLE to CHEAPAfter: the admin edits only the classifier opening and examples, while the proxy always supplies the current tier definitions
Relevant issues
Linear ticket
Pre-Submission checklist
CI status
183 of 186 checks pass on f380917, including codecov/patch at 86.76% of diff hit against a 79.05% target. The 3 reds are two staging-wide tests this PR does not touch:
test_event_loop_stall_timeout_burst_keeps_breaker_closed(redis breaker timing, py3.10 and 3.11) andtest_qualifiers_and_optionality_are_unwrapped(py3.10 only, typing qualifier gap intest_http_parsing_utils.py). Both reproduce on rerun and fail identically on unrelated open PRs #39770, #39766, #39763 and #39762Screenshots / Proof of Fix
Shared setup: an isolated proxy on port 4275 used a cloned local Postgres database and a real provider gateway. The test router used built-in tiers renamed CHEAP, STANDARD, PREMIUM, and DEEP
Before (39a1789)
Built-in opening-only prompt
/model/newwith built-in tiers, an LLM classifier, andclassification_promptclassification_prompt requires tier_definitionsTier-label agreement
- SIMPLE:and rename SIMPLE to CHEAP/v1/chat/completionsthrough that router- SIMPLE:while its response schema enum isCHEAP,STANDARD,PREMIUM,DEEPAfter (e93300f)
Built-in opening-only prompt
/model/newwith the same built-in tiers, LLM classifier, andclassification_promptprompt-router-opening-id/v1/chat/completions,/v1/responses, and/v1/messagesthrough the new routerTier-label agreement
/auto_router/classifier/default_promptwith the opening text, labels CHEAP and DEEP, and the business rubric- CHEAP:and- DEEP:criteria plus the injection guard- CHEAP:bullet and the schema enumCHEAP,STANDARD,PREMIUM,DEEPtier_labelsvalidation errorType
🐛 Bug Fix
Caveats (if any)
Low
Final Attestation
Note
Medium Risk
Changes how classifier system prompts are built and validated for auto-routers, which directly affects routing tier decisions; behavior is heavily tested but misconfiguration could shift classification until operators reset prompts.
Overview
Built-in LLM auto-routers can now customize classifier rubrics in sections instead of pasting a full prompt that freezes tier names. A new
classification_examplesconfig field (with its own length cap and normalizer) sits besideclassification_prompt; the router assembles instructions → derived tier bullets → calibration examples → trust boundary → closing line viabuilt_in_tier_classification_promptand shared_sectioned_prompt. Custom tier sets use the same section model and pass examples through to the live classifier.Validation and preview align with that model:
classification_prompt/classification_examplesare allowed on built-in routers with an LLM classifier but are rejected alongside legacyclassifier_llm_config.system_promptor non-LLM classifiers. The POST/auto_router/classifier/default_promptpreview accepts optionaltier_labels,classification_rubric, and examples, and routes built-in vs custom tier previews through the same assembly functions the router uses.Dashboard create/edit flows replace the standalone rubric dropdown and
CustomTierPromptEditorwithOpeningPromptEditor(instructions + examples + base rubric in one dialog, proxy-assembled preview viagetAutoRouterAssembledPromptCall). Built-in configs can persist opening sections on save; legacy whole-prompt overrides keepClassifierPromptEditorwith clearer warnings about frozen tiers.Reviewed by Cursor Bugbot for commit f380917. Bugbot is set up for automated code reviews on this repo. Configure here.