Skip to content

fix(auto_router): derive tier definitions in prompt editor - #39688

Merged
tin-berri merged 1 commit into
litellm_internal_stagingfrom
litellm_autorouter_prompt_tiers_readonly
Sep 4, 2026
Merged

tin-berri merged 1 commit into
litellm_internal_stagingfrom
litellm_autorouter_prompt_tiers_readonly

Conversation

@tin-berri

@tin-berri tin-berri commented Sep 4, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Built-in prompt edits duplicate tier definitions into free text
  • Tier renames can then disagree with classifier response labels

How it solves it:

  • New built-in edits save opening instructions and calibration examples only
  • The proxy derives tier criteria, labels, guard, and closing line
  • Existing whole-prompt overrides stay editable in legacy mode

User Flow

Before: an admin can only replace the complete classifier prompt, so a tier rename can leave the prompt describing labels the classifier cannot return

  1. They open http://localhost:4000/ui/?page=models and edit a built-in LLM auto-router
  2. They change the classifier prompt, including its Tiers: section, then rename SIMPLE to CHEAP
  3. The saved prompt still says SIMPLE while the router accepts CHEAP

After: the admin edits only the classifier opening and examples, while the proxy always supplies the current tier definitions

  1. They open http://localhost:4000/ui/?page=models and edit the same built-in LLM auto-router
  2. They choose Edit prompt and change the opening instructions or calibration examples
  3. The preview shows the current CHEAP tier definition beneath their text
  4. A legacy whole-prompt override remains editable with a warning and can be reset to the derived mode

Relevant issues

  • Keeps the prompt editor aligned with structured tier configuration
  • Preserves existing legacy whole-prompt configurations unchanged

Linear ticket

Pre-Submission checklist

  • I have added meaningful tests
  • The focused backend and dashboard suites pass locally
  • My PR passes all required CI/CD checks (183 of 186; see CI status below)
  • My PR's scope is isolated to classifier prompt editing
  • I have received a Greptile Confidence Score of at least 4/5 before requesting maintainer review

CI status

183 of 186 checks pass on f380917, including codecov/patch at 86.76% of diff hit against a 79.05% target. The 3 reds are two staging-wide tests this PR does not touch: test_event_loop_stall_timeout_burst_keeps_breaker_closed (redis breaker timing, py3.10 and 3.11) and test_qualifiers_and_optionality_are_unwrapped (py3.10 only, typing qualifier gap in test_http_parsing_utils.py). Both reproduce on rerun and fail identically on unrelated open PRs #39770, #39766, #39763 and #39762

Screenshots / Proof of Fix

image

Shared setup: an isolated proxy on port 4275 used a cloned local Postgres database and a real provider gateway. The test router used built-in tiers renamed CHEAP, STANDARD, PREMIUM, and DEEP

Before (39a1789)

Built-in opening-only prompt

  1. POST /model/new with built-in tiers, an LLM classifier, and classification_prompt
  2. Response: HTTP 400, classification_prompt requires tier_definitions

Tier-label agreement

  1. Create a router with an existing whole-prompt override containing - SIMPLE: and rename SIMPLE to CHEAP
  2. Send POST /v1/chat/completions through that router
  3. The classifier request contains prompt bullet - SIMPLE: while its response schema enum is CHEAP, STANDARD, PREMIUM, DEEP

After (e93300f)

Built-in opening-only prompt

  1. POST /model/new with the same built-in tiers, LLM classifier, and classification_prompt
  2. Response: HTTP 200, model id prompt-router-opening-id
  3. POST /v1/chat/completions, /v1/responses, and /v1/messages through the new router
  4. All three responses return HTTP 200

Tier-label agreement

  1. POST /auto_router/classifier/default_prompt with the opening text, labels CHEAP and DEEP, and the business rubric
  2. Response: HTTP 200; prompt begins with the operator text and contains derived - CHEAP: and - DEEP: criteria plus the injection guard
  3. The real classifier request contains the same derived - CHEAP: bullet and the schema enum CHEAP, STANDARD, PREMIUM, DEEP
  4. POST the same preview with blank SIMPLE label
  5. Response: HTTP 400 with a tier_labels validation error

Type

🐛 Bug Fix

Caveats (if any)

Low

Final Attestation

  • The tests cover derived tier criteria, legacy behavior, preview validation, and the dashboard create and edit flows

Note

Medium Risk
Changes how classifier system prompts are built and validated for auto-routers, which directly affects routing tier decisions; behavior is heavily tested but misconfiguration could shift classification until operators reset prompts.

Overview
Built-in LLM auto-routers can now customize classifier rubrics in sections instead of pasting a full prompt that freezes tier names. A new classification_examples config field (with its own length cap and normalizer) sits beside classification_prompt; the router assembles instructions → derived tier bullets → calibration examples → trust boundary → closing line via built_in_tier_classification_prompt and shared _sectioned_prompt. Custom tier sets use the same section model and pass examples through to the live classifier.

Validation and preview align with that model: classification_prompt / classification_examples are allowed on built-in routers with an LLM classifier but are rejected alongside legacy classifier_llm_config.system_prompt or non-LLM classifiers. The POST /auto_router/classifier/default_prompt preview accepts optional tier_labels, classification_rubric, and examples, and routes built-in vs custom tier previews through the same assembly functions the router uses.

Dashboard create/edit flows replace the standalone rubric dropdown and CustomTierPromptEditor with OpeningPromptEditor (instructions + examples + base rubric in one dialog, proxy-assembled preview via getAutoRouterAssembledPromptCall). Built-in configs can persist opening sections on save; legacy whole-prompt overrides keep ClassifierPromptEditor with clearer warnings about frozen tiers.

Reviewed by Cursor Bugbot for commit f380917. Bugbot is set up for automated code reviews on this repo. Configure here.

@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai can you reply here with a score and why that score?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@codspeed

codspeed Bot commented Sep 4, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_autorouter_prompt_tiers_readonly (f380917) with litellm_internal_staging (2e73400)

Open in CodSpeed

@greptile-apps

greptile-apps Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

The PR introduces opening-only prompt customization for built-in complexity-router tiers while deriving tier labels, criteria, trust-boundary text, and closing instructions from structured configuration.

  • Adds shared built-in prompt assembly and validates conflicts with legacy whole-prompt overrides.
  • Extends the prompt-preview API to support built-in tier labels and rubric presets.
  • Updates dashboard create, edit, reset, preview, and persistence flows for the new prompt shape.
  • Preserves legacy whole-prompt editing for existing overrides and expands focused backend and UI coverage.

Confidence Score: 5/5

The PR appears safe to merge, with the live prompt, preview API, configuration validation, and dashboard persistence paths aligned around the derived-tier prompt contract.

The new built-in opening path derives active tier labels and criteria from the same router configuration used by classifier output validation, preserves the injection guard, rejects conflicting prompt modes, and is consistently represented by the dashboard builders and preview transport.

Important Files Changed

Filename Overview
litellm/router_strategy/complexity_router/complexity_router.py Adds built-in opening-only prompt assembly and uses it before the default full-rubric path while preserving derived labels and trust-boundary text.
litellm/router_strategy/complexity_router/config.py Allows built-in classification_prompt values for LLM classifiers and rejects conflicting legacy whole-prompt configurations.
litellm/proxy/management_endpoints/model_management_endpoints.py Extends prompt preview requests with optional built-in labels and rubric selection while sharing label validation with the existing GET endpoint.
ui/litellm-dashboard/src/components/add_model/ClassificationMethodConfig.tsx Selects the opening-only editor for derived prompts and retains the legacy editor when a whole-prompt override already exists.
ui/litellm-dashboard/src/components/add_model/OpeningPromptEditor.tsx Generalizes the custom-tier editor to preview and save opening-only prompts for both custom and built-in tier sources.
ui/litellm-dashboard/src/components/add_model/build_complexity_router_config.ts Serializes built-in opening prompts only for LLM classifiers without a conflicting legacy system prompt.
ui/litellm-dashboard/src/components/edit_auto_router/edit_auto_router_modal.tsx Keeps classification_prompt under managed edit-state persistence so resets and updates are reflected in saved configuration.
ui/litellm-dashboard/src/components/networking.tsx Replaces the custom-tier-only preview helper with a discriminated source supporting either definitions or built-in labels and rubric.

Reviews (1): Last reviewed commit: "fix(auto_router): derive tier definition..." | Re-trigger Greptile

@greptile-apps

greptile-apps Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

Confidence score: 4/5.

The change correctly centralizes built-in tier prompt assembly in built_in_tier_classification_prompt, so renamed labels, tier criteria, the trust boundary, and the closing line are derived from the same router configuration (litellm/router_strategy/complexity_router/complexity_router.py:212). It also preserves legacy whole-prompt overrides, validates preview labels through the same ComplexityRouterConfig path as writes, and cleanly separates built-in versus custom prompt editing in the dashboard.

The added backend and dashboard tests cover derived criteria, label validation, legacy behavior, reset behavior, and the create/edit flows. I found no blocking correctness issues in the reviewed diff. I’m stopping at 4 rather than 5 because the PR shows only the focused Buildkite status, while the full required CI/CD validation is still pending.

@codecov

codecov Bot commented Sep 4, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit e93300f. Configure here.

@tin-berri
tin-berri enabled auto-merge (squash) September 4, 2026 01:17
@tin-berri

Copy link
Copy Markdown
Contributor Author

The documentation check is blocked by the base branch setting heuristic_v2_router_limit, which is documented on litellm-docs main by PR #1180. The new UI behavior is documented separately in litellm-docs PR #1183.

@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai can you reply here with a score and why that score?

@greptile-apps

greptile-apps Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

Confidence score: 4/5.

The PR correctly centralizes built-in prompt assembly in built_in_tier_classification_prompt, ensuring renamed labels, tier criteria, the injection guard, and the closing line are derived from the same router configuration used by classifier output validation. The live classifier path and preview API are aligned, built-in opening-only prompts are validated for LLM-only use and conflicts with legacy whole-prompt overrides, and existing legacy overrides remain editable.

The dashboard changes also cleanly distinguish derived built-in/custom opening prompts from legacy whole-prompt mode, and the added backend/UI tests cover derived criteria, label validation, reset behavior, and create/edit flows. The focused Buildkite status and coverage report are positive. I’m stopping at 4 rather than 5 because the PR’s full required CI/CD checks are still not complete, and the documentation check is currently blocked by a base-branch dependency.

@tin-berri
tin-berri force-pushed the litellm_autorouter_prompt_tiers_readonly branch 2 times, most recently from d19d84b to 70baeb6 Compare September 4, 2026 03:19
@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai can you reply here with a score and why that score?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

Confidence score: 4/5.

The implementation correctly makes built-in prompt assembly derive tier labels, criteria, the injection guard, and closing instructions from the same structured configuration used by classifier output validation. The live classifier path and preview API are aligned, opening-only prompts are validated for the appropriate LLM flow, conflicting legacy whole-prompt overrides are rejected, and existing legacy configurations remain editable.

The dashboard changes consistently support create, edit, preview, reset, and persistence for both derived and legacy modes, with focused backend and UI coverage. I found no blocking correctness issues in the reviewed changes. I’m stopping at 4 rather than 5 because the PR’s full required CI/CD validation is not complete yet: the available Buildkite status is still pending, and the documentation check remains dependent on the base-branch documentation change.

Comment thread ui/litellm-dashboard/src/components/add_model/OpeningPromptEditor.tsx Outdated
@tin-berri
tin-berri force-pushed the litellm_autorouter_prompt_tiers_readonly branch from 70baeb6 to 8bc4dfb Compare September 4, 2026 03:44
@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai can you reply here with a score and why that score?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

Confidence score: 4/5.

The change addresses the reported mismatch at the right boundary: built-in tier labels and criteria are now derived by the proxy, while only operator-owned instructions and examples remain editable. The live classifier and preview share the same assembly path, legacy whole-prompt overrides remain supported, and the backend/dashboard tests cover label renames, legacy splitting, validation, reset flows, rubric selection, and preview parity. The existing trust-boundary guard is also retained after operator-controlled sections.

I’m not giving 5/5 because this is a relatively broad backend/dashboard change with load-time migration logic; malformed or unusual legacy prompt layouts could still be a compatibility risk. The focused coverage and passing reported e2e status provide good evidence, but the full required CI suite is not yet confirmed.

Comment thread litellm/router_strategy/complexity_router/config.py Outdated
@tin-berri
tin-berri force-pushed the litellm_autorouter_prompt_tiers_readonly branch from 8bc4dfb to 46609cf Compare September 4, 2026 17:18
@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai can you reply here with a score and why that score?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

Confidence score: 4/5

The change is well-scoped and addresses the core mismatch: built-in routers now derive tier bullets, labels, rubric criteria, trust-boundary text, and closing instructions from structured configuration, while operator-provided instructions and calibration examples remain editable. The preview and live classifier share the same prompt assembly path, legacy whole-prompt overrides remain supported, and the added tests cover renamed tiers, invalid labels, custom-tier behavior, preview parity, normalization limits, and dashboard flows. The previously identified rubric-cancel and load-time prompt-splitting concerns are also addressed in the latest commit.

I’m giving 4 rather than 5 because this is a relatively broad cross-layer change (backend, API validation, and dashboard) with 23 files changed, and the PR checklist still shows required CI/CD checks incomplete. The focused tests and available E2E status are positive, but a full score would require all required CI checks and the remaining integration coverage to be confirmed.

@tin-berri
tin-berri force-pushed the litellm_autorouter_prompt_tiers_readonly branch from 46609cf to f380917 Compare September 4, 2026 17:37
@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai can you reply here with a score and why that score?

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

Score: 4/5

This is a strong, well-scoped fix. The live router and preview share the same prompt assembly paths, so tier labels, criteria, calibration sections, the trust boundary, and closing line remain consistent after renames. Legacy whole-prompt overrides are preserved, while the new section-based fields have clear mutual-exclusion, classifier-type, normalization, and length validation. The backend and dashboard tests cover derived prompts, preview parity, legacy behavior, validation failures, and create/edit round trips.

I’m not giving 5/5 because this is a behavior-sensitive change across 23 files and the required CI/e2e status is still pending. I did not identify a concrete correctness defect in the reviewed changes; assuming the pending suites remain green, this is close to merge-ready.

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit f380917. Configure here.

@tin-berri

Copy link
Copy Markdown
Contributor Author

CI settled: 183 of 186 pass, codecov included. The 3 reds are staging-wide tests this PR does not touch, failing the same way on #39770, #39766, #39763 and #39762

@greptileai can you reply here with a score and why that score?

@greptile-apps

greptile-apps Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

Score: 4/5

This change directly fixes prompt/tier drift by deriving tier definitions, labels, guards, and closing text from the current router configuration while preserving legacy whole-prompt overrides. The coverage is meaningful across derived prompts, legacy behavior, preview validation, and dashboard create/edit flows. CI is also strong: 183/186 checks pass, with 86.76% diff coverage against a 79.05% target. The three failures are staging-wide, reproduce on multiple unrelated PRs, and do not touch this change, so I do not consider them regressions here.

I’m not giving 5/5 because the full required suite is not green and this changes live classifier prompt assembly, which retains some integration risk despite the focused tests.

@tin-berri
tin-berri merged commit b3c867c into litellm_internal_staging Sep 4, 2026
266 of 272 checks passed
@tin-berri
tin-berri deleted the litellm_autorouter_prompt_tiers_readonly branch September 4, 2026 23:48
tin-berri added a commit that referenced this pull request Sep 5, 2026
… the auto_router license feature

Generalizes the heuristic_v2 ceiling from #39468 into a capability table whose
records own their in-process predicate, SQL spelling and refusal wording. The
existing heuristic_v2 capability keeps its own one-router ceiling. A single
customization capability combines operator-defined tier_definitions with every
operator-written part of the classifier prompt: a whole replacement
system_prompt, replacement opening instructions (classification_prompt) and
replacement calibration examples (classification_examples, the fields the
dashboard prompt editor from #39688 writes on built-in-tier routers). So an
unlicensed proxy can have at most one router that customizes either form. The
prompt half only applies to classifier types that call an LLM. The shipped
default prompt, classification rubric presets, tier_labels renames and tier
model choices remain ungated. The signed auto_router license feature makes both
ceilings unlimited.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants