Skip to content

fix(autorouter): preserve historical savings and align dashboard totals - #42647

Closed
tin-berri wants to merge 4 commits into
mainfrom
litellm_restore_historical_router_savings
Closed

tin-berri wants to merge 4 commits into
mainfrom
litellm_restore_historical_router_savings

Conversation

@tin-berri

@tin-berri tin-berri commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Historical savings disappeared after switching estimation methods
  • Daily costs included requests outside the selected dates
  • Savings cards mixed comparable costs with unestimated requests

How it solves it:

  • Preserve recorded savings and account by request date
  • Recover historical costs from reconciled retained records
  • Compare actual and baseline spend for eligible requests

Intentional product change: the savings card's actual spend now includes only requests with a baseline estimate, so both sides cover the same requests. The duplicate actual-spend row is removed, and LLM/classification subtotals are hidden when estimates cover only some requests because those subtotals include all usage. A tooltip explains eligibility and coverage remains visible

User Flow

Before: a request without an estimate inflates actual spend beside the savings comparison

  1. Open http://127.0.0.1:48182/cost-optimization/ and select Auto-Router, then Usage
  2. Select review-mixed over the last 30 days
  3. See $1 saved at 50%, but $100 actual and another $1 actual-spend row

Before: actual spend includes the unestimated $99 request

After: actual and baseline spend describe the same eligible requests

  1. Open http://127.0.0.1:48184/cost-optimization/ and select Auto-Router, then Usage
  2. Select review-mixed over the last 30 days
  3. See $1 saved at 50%, $1 actual, $2 baseline, and coverage of one request out of two
  4. Hover the actual-spend information icon to see why the $99 request is excluded

After: actual and baseline spend cover the same request

Tooltip explains estimate eligibility

Relevant issues

Corrects the historical reporting regression introduced by #41177

Affected release

Since v1.103.0-rc.1

Changes

Preserve signed historical savings, add new savings once, and persist request-date costs and estimate coverage consistently across old and new workers. Historical cost recovery reconciles retained logs against daily counts, spend and tokens within a two-second query limit. Missing comparison data stays unavailable instead of becoming zero

Skip installed schema changes on restart and bound retries for required accounting repairs. Daily totals reuse the savings aggregation without fetching unused usage breakdowns or key metadata. Sessions, selected routers and the CLI compare matching estimated requests

The dashboard uses eligible actual spend in the savings card. Total actual spend remains available through existing accounting APIs. Complete comparisons retain the LLM/classification breakdown; partial comparisons show actual spend, baseline spend, and coverage

Validation

Current revision: f96bfdd66dbe3e2643dbbe2da422d0e40c14d961, based on 474ab91c09, which includes the merged CI fix from #43235

All 64 tests across the three affected dashboard files pass, and scoped formatting, lint and budget checks pass. Reverting the display to all-request actual spend makes the mixed-request regression fail: expected $1.00 / $2.00, received $100.00 / $2.00. The correct implementation was restored before committing

This update changes four existing files, with 40 additions and 53 deletions. Across the PR, tests and fixtures add 621 lines against 652 production lines, a 0.95:1 ratio, excluding migrations, schema copies and generated types

An earlier broader local backend run on f996803d0b finished with 867 passed and one failure: test_prometheus_metrics_port_starts_separate_metrics_process expected exit 2 and received exit 1. Its cause is unresolved. Hosted proxy-server / Run tests, misc / Run tests, assert-ci-coverage, and assert-shard-coverage passed on that revision. The latest change is confined to the dashboard; hosted CI is rerunning on the current tip

Pre-Submission checklist

  • I have added meaningful tests
  • The affected dashboard test files pass locally
  • My PR passes all required CI/CD checks
  • My PR's scope is as isolated as possible
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review

Screenshots / Proof of Fix

Both revisions use real localhost proxies, the rendered dashboard and the same synthetic records in Postgres, with no mocked network responses. Request A costs $1 against a $2 baseline; request B costs $99 and has no baseline estimate. This flow reads recorded charges, so provider inference does not apply

Before (474ab91)

  1. Open http://127.0.0.1:48182/cost-optimization/, select Auto-Router, Usage, and review-mixed over the last 30 days
  2. Observe $1 saved at 50%, $100 actual spend, a second $1 actual-spend row, and $2 baseline, as shown in the Before screenshot above

After (f96bfdd)

  1. Open http://127.0.0.1:48184/cost-optimization/, select Auto-Router, Usage, and review-mixed over the last 30 days
  2. Observe $1 saved at 50%, $1 actual spend, $2 baseline, and “Savings based on 1 of 2 requests,” as shown in the After screenshot above
  3. Hover the actual-spend information icon: the tooltip says requests without a recorded baseline estimate are excluded and actual spend includes classification costs

Type

Bug Fix

Caveats (if any)

Medium

  • Missing retained logs can leave historical comparisons unavailable
  • Slow historical recovery preserves savings but skips cost details
  • Selected-router views cover whole overlapping sessions

Low

  • Partial comparisons omit the all-request LLM/classification breakdown
  • Permanent installation failures also receive four bounded attempts

Final Attestation

  • Regression coverage includes history, dates, scope and mixed estimates

Note

Medium Risk
Changes financial reporting paths (daily spend writes, DB triggers, benchmarks API) and requires successful migration install at boot; recovery and partial-coverage logic can leave costs unavailable rather than wrong.

Overview
Adds request-date auto-router cost accounting on LiteLLM_DailyUserSpend (LLM spend, classifier cost, routed/estimated request counts) plus an idempotent SQL migration that installs columns, a baseline-observation trigger, and apply_autorouter_daily_coverage() after Prisma setup. Spend writes, daily upserts, and queue aggregation now populate and roll up those fields; incomplete daily rows can be reconciled from spend logs within bounded queries.

Benchmarks and cost-optimization totals no longer derive savings only from session rollups: headline saved_spend, spend, baseline, and coverage come from the same daily aggregation as overall cost optimization, with optional historical recovery. Per-router groups still describe whole sessions overlapping the window but keep recorded saved_spend and historical baselines. Session and CLI responses split full-session baseline from estimate-only baseline and fix partial-coverage savings display.

The dashboard savings hero compares actual vs baseline only for requests with a baseline estimate (tooltip + “N of M requests”), hides LLM/classification subtotals when coverage is partial, and surfaces cost_coverage / nullable totals fields in the API types.

Reviewed by Cursor Bugbot for commit f96bfdd. Bugbot is set up for automated code reviews on this repo. Configure here.

@tin-berri
tin-berri requested a review from a team September 23, 2026 01:56
@tin-berri tin-berri added the backport-stable P0 regression fix only (Urgent ticket): cherry-pick onto the baking rc line before the stable tag label Sep 23, 2026
@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@greptile-apps

greptile-apps Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

[Critical risk] Adds database migration and schema changes for autorouter cost tracking.

The PR should not merge until mixed-history router groups present a savings comparison over one consistent request population.

Findings

  1. P1 Mixed savings populations ▶

Summary

The PR preserves recorded auto-router savings, adds request-date cost and estimate-coverage accounting with bounded historical recovery, and changes the dashboard to compare eligible actual spend with its baseline. The new mixed-history group comparison still pairs total savings with a percentage calculated for a smaller population.

Reviews (13) · Last reviewed commit: "fix(autorouter): show only eligible spen..."

@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@tin-berri
tin-berri force-pushed the litellm_restore_historical_router_savings branch from ea98fcc to 4ad588d Compare September 23, 2026 02:02
@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai Please review the current commit, including updated key and user dashboard regressions for request-date savings and unavailable costs

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@codspeed

codspeed Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_restore_historical_router_savings (f96bfdd) with main (d08746f)

Open in CodSpeed

@codecov

codecov Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 99.25926% with 1 line in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
...llm/proxy/client/cli/commands/statusline_script.py 92.30% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

@tin-berri
tin-berri force-pushed the litellm_restore_historical_router_savings branch from 4ad588d to 4f014ae Compare September 23, 2026 02:27
@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai Rebased onto the upstream CI fix and added invalid-baseline regressions. Please review the current commit, including preserved historical totals

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@tin-berri
tin-berri force-pushed the litellm_restore_historical_router_savings branch from 4f014ae to 48ef1f4 Compare September 23, 2026 17:29
@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai Please review this updated commit. Unknown-only sessions now retain unavailable savings, with API-to-CLI regression coverage.

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

Comment thread litellm/proxy/management_endpoints/auto_router_endpoints.py Outdated
@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai Please reconsider the finding against the requested no-coverage rule. Both counters default to zero; regression tests preserve all provable comparisons

@greptile-apps

greptile-apps Bot commented Sep 23, 2026

Copy link
Copy Markdown
Contributor

Reconsidered. I’m withdrawing the finding: under the requested no-coverage rule, savings_estimated_turns == 0 means there is no provable savings comparison, and the default saved_spend == 0 cannot distinguish a legacy/unknown session from a genuinely recorded zero. The updated tests cover the relevant comparisons (including estimated zero, signed savings, historical savings, and unknown-only sessions), while keeping unknown-only baseline_spend unavailable. I don’t have a confirmed correctness issue against this PR’s stated semantics.

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/proxy/management_endpoints/auto_router_endpoints.py
@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai The zero-coverage finding has been withdrawn. Please refresh the current-commit review summary to reflect the verified session behavior

@tin-berri
tin-berri force-pushed the litellm_restore_historical_router_savings branch from 48ef1f4 to 4f9156a Compare September 23, 2026 17:45
@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai Please review the updated commit. Session and router-group readers now share the unavailable-baseline rule, with API, CLI and dashboard regressions

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@tin-berri

Copy link
Copy Markdown
Contributor Author

@cursor review

Please review 8215a1b, including historical-query cancellation, restored pooled settings, preserved savings, and date-aligned dashboard costs

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@tin-berri
tin-berri force-pushed the litellm_restore_historical_router_savings branch from 8215a1b to 175ec8f Compare September 24, 2026 05:34
@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai Please review 175ec8f, especially catalog-only restarts, concurrent installation rechecks, trigger repair, and preservation of compatible newer accounting functions

@tin-berri

Copy link
Copy Markdown
Contributor Author

@cursor review

Please review 175ec8f for startup under database contention, schema-scoped installation, repairs, and compatibility with newer accounting functions

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/proxy/db/prisma_client.py
@tin-berri
tin-berri force-pushed the litellm_restore_historical_router_savings branch from 175ec8f to b0b24e7 Compare September 24, 2026 09:18
@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai Please review b0b24e7. Newer installations are protected as a whole, transient installation failures retry, and authenticated database regressions pass

@tin-berri

Copy link
Copy Markdown
Contributor Author

@cursor Please review b0b24e7. Installation retries are fixed; the existing pending_projection producer and writer tests demonstrate why coverage is not double-counted

@cursor

cursor Bot commented Sep 24, 2026

Copy link
Copy Markdown
Contributor

You need to increase your spend limit or enable usage-based billing to run background agents. Go to Cursor

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

Please review b0b24e7 for bounded installation retries, preservation of newer schema contracts, and existing pending_projection coverage ownership

Comment thread litellm-proxy-extras/litellm_proxy_extras/utils.py

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@tin-berri
tin-berri force-pushed the litellm_restore_historical_router_savings branch from 9e23797 to f996803 Compare September 25, 2026 23:45
@tin-berri

Copy link
Copy Markdown
Contributor Author

@greptileai review pls

@tin-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

Comment on lines +716 to +720
comparison_savings: Final = row.savings_estimated_saved_spend if row.savings_estimated_turns else row.saved_spend
comparison_actual: Final = row.savings_estimated_actual_spend if row.savings_estimated_turns else row.spend
baseline_spend: Final = (
comparison_actual + comparison_savings
if row.savings_estimated_turns or not row.turns

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Mixed savings populations

When a router group contains historical turns with recorded savings and newer turns with estimates, this code calculates the baseline and percentage from only the newer turns but returns savings for all turns. The dashboard can show $30 saved beside a $1.50 baseline and a −33% badge, while saying savings are based on 4 of 40 requests. Those figures describe different requests and give users a misleading comparison.

Knowledge Base Used: Dashboard and enterprise UI

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit f96bfdd. Configure here.

@tin-berri tin-berri closed this Sep 27, 2026

This branch is waiting to be deployed

1 waiting deployment
e2e-changed — f96bfdd6 Waiting Sep 26, 2026 by tin-berri via oauth #1280
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backport-stable P0 regression fix only (Urgent ticket): cherry-pick onto the baking rc line before the stable tag

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant