Skip to content

chore(backport): add the remaining jev changes to stable/1.99.x - #42669

Open
devin-ai-integration[bot] wants to merge 19 commits into
stable/1.99.xfrom
litellm_jev_backport_1_99_x
Open

devin-ai-integration[bot] wants to merge 19 commits into
stable/1.99.xfrom
litellm_jev_backport_1_99_x

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

How it solves it:

  • Cherry-picks the six follow-up Jev PRs from main onto stable/1.99.x, one commit per PR, original author kept
  • Ports the four main prerequisites those picks build on as their own commits
  • Adapts the dashboard pieces to what this line ships and regenerates the openapi snapshot and api types under Python 3.12
  • No version bump in this PR; the release bump is handled separately

Included PRs:

Backport notes: git cherry-pick -n -m 1 from the main merge commit for cf42b60, 34718f0, 2edda5a, 1e161f5 and a83773c, plain pick for the #42301 squash 1106b16, each committed with the original author, subject and a two-line provenance body. 4bfac88281 (the stable dashboard adaptation from #42595) is picked with -x right after #41886

Prerequisites, each its own commit with a provenance body naming the main sha: 7015bf37bb (apply team model aliases on the JWT auth path), 109ca70f66 (allow opted-in team members to manage their routers: the DatabaseClient protocol, AUTO_ROUTER_MANAGE, the prisma_client kwarg on can_key_call_model / can_team_access_model, the lazy team_membership params on _check_team_member_model_access, the member dry-run plumbing and matching test helpers), d8edfb69c2 (#38174, derive auto-router health from its underlying models, which #41615 and #41879 build on and which this oldest line did not have yet) and 510424c86c (classifier circuit breaker, which the Jev classifier plugs into)

Conflicts resolved, always toward what main has after the pick: this is the oldest line and lacks CustomTierSet, the NON_REASONING tier, heuristic_first / hybrid / heuristic_v2, classification_examples, forecast_classifier_config and AutoRouterClassifierTabs, so shared files took ours plus each pick's delta and the tier and classifier-tab files #41886 brought from main that depend on those features were dropped again in the stable adaptation. schema.d.ts and _lazy_openapi_snapshot.json were regenerated wholesale (see below). Own follow-up commits, all Devin identity: fix(ui): adapt the jev dashboard pieces to stable/1.99.x (classifier fields passed at the serializer call site in add_auto_router_tab.tsx, dead hydratePlanModeMinTier removed, nested ternary replaced by a lookup map in autoRouterRows.ts, stale imports removed and call sites hoisted to named consts in edit_auto_router_modal.tsx), fix(ui): pass the jev classifier fields the serializer requires on stable/1.99.x, fix(ui): clear the eslint errors on the touched dashboard files, fix(lint): satisfy the strict and type-discipline gates on the picked code, style(dashboard): prettier-format the touched dashboard files

Generated artifacts: #41615 and #41757 brought litellm/proxy/_lazy_openapi_snapshot.json and ui/litellm-dashboard/src/lib/http/schema.d.ts verbatim from main. chore(backport): regenerate the openapi snapshot and dashboard api types for stable/1.99.x regenerates both under Python 3.12 (uv sync --python 3.12 --inexact --frozen --extra proxy --group proxy-dev --group e2e-dev, scripts/prisma_generate_if_needed.py, python -m litellm.proxy._lazy_openapi_snapshot, npm run gen:api), so the snapshot drops agent_365 from the unreachable_fallback description and schema.d.ts drops the upstream-only classifier surface and JsonValue and gains /auto_router/manage

Review-loop follow-ups after c2bcf43, both Devin identity with a body: a7bdc0a753 restores main's call shape in the test-routing preview (forward request_kwargs["messages"] and refresh the raw-body snapshot instead of synthesizing one user message from prompt), and 333fb3c951 gates classifier_fallback on usesClassifierContext in the dashboard serializer so a fallback picked for a jev router is saved, matching main

Checks run on the tip c2bcf43 and repeated on 333fb3c. Python: every tests/test_litellm file the picks touch was run in the worktree with LITELLM_LOCAL_MODEL_COST_MAP=True, 563 passed, 14 failed; the 14 are the requires_semantic_router tests in test_complexity_router.py, which fail on the line's tip before these picks too because semantic_router is not installed. Dashboard, under ui/litellm-dashboard: npx vitest run on every touched test file PASS (391 tests), npx eslint on every changed file PASS with 0 errors and no net new local / no-large-inline-object-arg warnings, npx tsc --noEmit on non-test sources PASS, npm run test:types PASS, npm run build PASS. make check at the repo root: FAIL, on exactly one finding, LIT002: total 26877 over limit 26873 (this change added 1) - litellm/integrations/otel/model/payloads.py:104,106. That file is byte-identical to origin/stable/1.99.x and no commit on this branch touches it. The gate measures against merge-base(origin/litellm_internal_staging, HEAD), so it counts drift the stable line picked up since it forked. Running the same gate with the line as base, scripts/type_discipline_gate.py --base origin/stable/1.99.x, prints OK: every LIT rule is within its codebase ceiling. Every other make check gate passed. No budget JSON was edited

User Flow

Before: an admin on stable 1.99.x cannot use Jev anywhere except the raw passthrough

  1. The admin sets TYPESAFE_API_KEY, starts the proxy from stable/1.99.x and opens http://localhost:4000/ui/?page=models to add an auto-router
  2. The classification method dropdown offers only the LLM classifier, there is no Jev option
  3. They add a typesafe_compaction guardrail to config.yaml and the proxy refuses to start with an unknown guardrail error
  4. A developer sends POST http://localhost:4000/auto_router/test_routing with "classifier_type": "jev" and gets 422, jev is not an accepted classifier
  5. A developer sends POST http://localhost:4000/openrouter/api/alpha/decisions and gets 404

After: the same admin configures Jev routing, compaction and budgets from the UI and API

  1. The admin sets TYPESAFE_API_KEY, starts the proxy from this branch and opens http://localhost:4000/ui/?page=models to add an auto-router
  2. The classification method dropdown offers Jev, the connection test returns a tier for the sample prompt, and the router saves
  3. They add the typesafe_compaction guardrail and the proxy starts and trims low-relevance turns before forwarding
  4. The developer sends the same POST http://localhost:4000/auto_router/test_routing and gets 200 with "cause": "jev_classifier" and the chosen tier
  5. The developer sends POST http://localhost:4000/openrouter/api/alpha/decisions and gets the OpenRouter decision back with 200, logged with spend under typesafe/jev-1.13
  6. A virtual key over its budget that sends a Jev test-routing request gets 400 with the budget error instead of a free classification

Relevant issues

Backport of #41615, #41723, #41757, #41879, #41886 and #42301. Follows #42598 on this line and mirrors #42595

Linear ticket

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Two proxies from two worktrees, 2 uvicorn workers each, own Postgres each, real TypeSafe, OpenRouter and OpenAI calls, no mocks. Before is the merge base on port 18402, After is the PR tip on port 18403

Before (1394d33)

POST /openrouter/api/alpha/decisions

$ curl -s -X POST -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"model":"typesafe/jev-1.13","state":"draft","questions":{"q1":{"type":"choice","instructions":"which","options":{"a":"x","b":"y"},"criteria":{"c":"z"}}}}' http://localhost:18402/openrouter/api/alpha/decisions
{"detail":"Not Found"}

POST /model/new with a jev auto-router

$ curl -s -X POST -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"model_name":"qa99-jev-router","litellm_params":{"model":"auto_router/complexity_router","complexity_router_config":{"tiers":{"SIMPLE":"gpt-4o-mini"},"classifier_type":"jev","jev_classifier_config":{"model":"jev-latest"}}}}' http://localhost:18402/model/new
{"error":{"message":"complexity_router_config is invalid at classifier_type: Input should be 'heuristic', 'llm' or 'custom'. The router would drop this deployment at load time, so the write is rejected instead.","type":"validation_error","param":"litellm_params.model","code":"400"}}

POST /auto_router/test_routing with a jev classifier

$ curl -s -X POST -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"prompt":"<long prompt>","complexity_router_config":{"tiers":{"SIMPLE":"gpt-4o-mini","COMPLEX":"gpt-4o"},"classifier_type":"jev","jev_classifier_config":{"model":"jev-latest"}}}' http://localhost:18402/auto_router/test_routing
{"detail":[{"type":"literal_error","loc":["body","complexity_router_config","classifier_type"],"msg":"Input should be 'heuristic', 'llm' or 'custom'","input":"jev", ...}]}

Over-budget member key on /auto_router/test_routing

Not reachable: the jev classifier config is rejected at validation before the budget check runs

After (333fb3c)

POST /openrouter/api/alpha/decisions

$ curl -s -X POST -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '<same body>' http://localhost:18403/openrouter/api/alpha/decisions
{"model":"typesafe/jev-1.13-20260917","answers":{"q1":{"type":"choice","choice":"speed","probabilities":{"speed":1},"confidence":1}},"usage":{"input_tokens":292,"output_tokens":25,"cost":0.000012264},"id":"gen-dec-1790145767-LvEa4ObgOMa6EmBsCBmq","provider":"TypeSafe"}

POST /model/new with a jev auto-router

$ curl -s -X POST -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '<same body>' http://localhost:18403/model/new
{"model_id":"8ab24476-0150-4ba4-b65c-1aae025f5656","model_name":"qa99-jev-router","litellm_params":{"model":"BUJpwholAHxF3aQW_1LloIAgX4_fVZHA_ConLch...
$ curl -s -X POST -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"model":"qa99-jev-router","messages":[{"role":"user","content":"say ok"}]}' http://localhost:18403/v1/chat/completions
200, model qa99-jev-router, content "Ok!" (real OpenAI completion)

POST /auto_router/test_routing with a jev classifier

$ curl -s -X POST -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '<same body>' http://localhost:18403/auto_router/test_routing
{"routed_model":"gpt-4o","routed_model_configured":false,"routing_decision":{"router_model_name":"auto_router_routing_test","router_type":"complexity","routed_model":"gpt-4o","cause":"jev_classifier","tier":"COMPLEX","signals":["jev-classifier:COMPLEX","jev-confidence=0.920000", ...],"classifier_model":"typesafe/jev-1.13.0","classifier_cost":0.000020286,"classifier_probabilities":{"SIMPLE":0.0,"COMPLEX":0.94,"REASONING":0.06,"MEDIUM":0.0},"classifier_confidence":0.92,"conversation_continuing":false}}

Over-budget member key on /auto_router/test_routing

Member key under a team with team_member_permissions=["/auto_router/manage"], max_budget=0.000001, spend pushed over the cap by one real chat completion

$ curl -s -X POST -H 'Authorization: Bearer <member key>' -H 'Content-Type: application/json' -d '{"prompt":"write fizzbuzz","team_id":"5cc9b812-b092-421c-81d0-90090bf507a0","complexity_router_config":{"tiers":{"SIMPLE":"gpt-4o-mini"},"classifier_type":"jev","jev_classifier_config":{"model":"jev-latest"}}}' http://localhost:18403/auto_router/test_routing
{"error":{"message":"Budget has been exceeded! Key=key (sk-...n7bA) Current cost: 6.599999999999265e-06, Max budget: 1e-06","type":"budget_exceeded","param":null,"code":"400"}}

Also on the After side: GET /typesafe/v1/models returns the TypeSafe model list, and a typesafe_compaction guardrail loaded from config shows up in GET /v2/guardrails/list

Type

🆕 New Feature

Caveats (if any)

Medium

  • make check fails on this branch on LIT002 in litellm/integrations/otel/model/payloads.py, a file this branch does not touch; the same gate passes with --base origin/stable/1.99.x

Low

  • fix(ui): adapt the jev dashboard pieces to stable/1.99.x (521ebb4) has no commit body; rewriting it would need a force push
  • The 14 requires_semantic_router tests fail on the line's tip with or without this PR

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR
  • 333fb3c passes /live-pr-risk

Link to Devin session: https://app.devin.ai/sessions/dcedce8035d64a199713bae4b2b90b7f
Open in Devin Desktop: https://app.devin.ai/desktop/session/dcedce8035d64a199713bae4b2b90b7f?variant=devin
Requested by: @mateo-berri

ryan-crabbe-berri and others added 17 commits September 23, 2026 03:15
Prerequisite for #41615 on stable/1.100.x.

Cherry-picked from 7015bf3 (main).

(cherry picked from commit bcbda1466ea40207227362c36a733a62284ecc77)
Prerequisite for #41615 and #41879 on stable/1.99.x.

Cherry-picked from 109ca70 (main).
)

Prerequisite for #41615 and #41879 on stable/1.99.x.

Cherry-picked from d8edfb6 (main).
Prerequisite for #41615 on stable/1.99.x.

Cherry-picked from 510424c (main).
Backport of #41615 to stable/1.99.x.
Cherry-picked from cf42b60 (main).
Backport of #41723 to stable/1.99.x.
Cherry-picked from 34718f0 (main).
Backport of #41757 to stable/1.99.x.
Cherry-picked from 2edda5a (main).
Backport of #41879 to stable/1.99.x.
Cherry-picked from 1e161f5 (main).
Backport of #41886 to stable/1.99.x.
Cherry-picked from a83773c (main).
…pes for stable/1.99.x

The #41615 and #41757 picks brought litellm/proxy/_lazy_openapi_snapshot.json and ui/litellm-dashboard/src/lib/http/schema.d.ts verbatim from main. Both files were regenerated under Python 3.12.
…ions pass-through (#42301)

Backport of #42301 to stable/1.99.x.
Cherry-picked from 1106b16 (main).
…able/1.99.x

The add-auto-router submit path builds BuildComplexityRouterConfigParams, which requires classificationPrompt and classifierContextBudgetChars; both now come from the form config. hydratePlanModeMinTier referenced the CustomTierSet type that only exists on lines with the custom-tier UI, which this line lacks, and nothing on this line calls it, so the dead export is removed instead of stubbing the type.
eslint flagged seven unused imports left over from dropping the upstream-only tier surface, a nested ternary in autoRouterRows, and new local/no-large-inline-object-arg warnings versus the stable base: the ternary is a label map, the inline objects in the edit modal and the transition test are named variables, and the edit modal import block regains the classifier and adaptive types its interfaces reference.
… code

The picks brought upstream imports, signatures and literals that tripped the 1.99.x lint gates: ruff flagged two unused imports and one unsorted import block, the strict gate flagged BLE001 and the same F401s, and the type-discipline gate flagged mutable-annotation (LIT001/2) and **kwargs (LIT008) lines. Imports cleaned and sorted, and each remaining flagged line carries the required suppression with its reason.
make check runs prettier over the dashboard tree; two touched files drifted from its output.
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

CLAassistant commented Sep 23, 2026 •

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you all sign our Contributor License Agreement before we can accept your contribution.
5 out of 7 committers have signed the CLA.

✅ tin-berri
✅ mateo-berri
✅ ryan-crabbe-berri
✅ moe-berri
✅ yuneng-berri
❌ yassin-berriai
❌ devin-ai-integration[bot]
You have signed the CLA already but the status is still pending? Let us recheck it.

@greptile-apps

greptile-apps Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge; the latest changes resolve the remaining preview-history and JEV-fallback gaps without introducing a new actionable failure.

Summary

This backport adds the remaining TypeSafe JEV routing, compaction, budget enforcement, passthrough, logging, and dashboard functionality to the stable 1.99.x line.

  • Adds JEV as a complexity-router classifier with context handling, circuit breaking, cost tracking, and fallback behavior.
  • Adds TypeSafe compaction guardrails and broadens TypeSafe passthrough method support.
  • Adds the OpenRouter decisions passthrough and pricing metadata.
  • Adds member-scoped auto-router authorization and dependency validation.
  • Adds dashboard creation, editing, connection testing, and routing-decision presentation for JEV routers.
  • Regenerates the OpenAPI snapshot and dashboard API types.

Reviews (2) · Last reviewed commit: "fix(dashboard): serialize classifier_fal..."

Comment thread litellm/proxy/management_endpoints/model_management_endpoints.py
Comment thread litellm/proxy/pass_through_endpoints/llm_passthrough_endpoints.py
Comment thread litellm/proxy/management_helpers/auto_router_permissions.py
Comment thread litellm/proxy/auth/team_grants.py
Comment thread litellm/router_strategy/complexity_router/config.py
Comment thread litellm/proxy/guardrails/guardrail_hooks/typesafe/typesafe.py
Comment thread ui/litellm-dashboard/src/components/add_model/build_complexity_router_config.ts Outdated
Comment thread litellm/proxy/pass_through_endpoints/success_handler.py
Comment thread litellm/proxy/auth/team_grants.py
@greptile-apps

This comment has been minimized.

The backport dropped the main-line call shape in preview_auto_router_routing: it synthesized a single user message from data.prompt and skipped refresh_proxy_server_request_body_snapshot, so a dry run carrying messages classified a null-content turn and the raw-body snapshot was never refreshed. Restore the imported helper call and pass request_kwargs["messages"] as main does.
The 1.99.x serializer gated classifier_fallback on classifierType === "llm", so a fallback chosen for a jev classifier was silently dropped on save. Main includes it whenever usesClassifierContext holds, which covers jev. Gate on usesClassifierContext to match.
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 333fb3c. Configure here.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants