Repository navigation
Conversation
|
bugbot run |
|
ea98fcc to
4ad588d
Compare
|
@greptileai Please review the current commit, including updated key and user dashboard regressions for request-date savings and unavailable costs |
|
bugbot run |
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
4ad588d to
4f014ae
Compare
|
@greptileai Rebased onto the upstream CI fix and added invalid-baseline regressions. Please review the current commit, including preserved historical totals |
|
bugbot run |
4f014ae to
48ef1f4
Compare
|
@greptileai Please review this updated commit. Unknown-only sessions now retain unavailable savings, with API-to-CLI regression coverage. |
|
bugbot run |
|
@greptileai Please reconsider the finding against the requested no-coverage rule. Both counters default to zero; regression tests preserve all provable comparisons |
|
Reconsidered. I’m withdrawing the finding: under the requested no-coverage rule, |
|
@greptileai The zero-coverage finding has been withdrawn. Please refresh the current-commit review summary to reflect the verified session behavior |
48ef1f4 to
4f9156a
Compare
|
@greptileai Please review the updated commit. Session and router-group readers now share the unavailable-baseline rule, with API, CLI and dashboard regressions |
|
bugbot run |
8215a1b to
175ec8f
Compare
|
@greptileai Please review 175ec8f, especially catalog-only restarts, concurrent installation rechecks, trigger repair, and preservation of compatible newer accounting functions |
175ec8f to
b0b24e7
Compare
|
@greptileai Please review b0b24e7. Newer installations are protected as a whole, transient installation failures retry, and authenticated database regressions pass |
|
You need to increase your spend limit or enable usage-based billing to run background agents. Go to Cursor |
|
bugbot run Please review b0b24e7 for bounded installation retries, preservation of newer schema contracts, and existing pending_projection coverage ownership |
9e23797 to
f996803
Compare
|
@greptileai review pls |
|
bugbot run |
| comparison_savings: Final = row.savings_estimated_saved_spend if row.savings_estimated_turns else row.saved_spend | ||
| comparison_actual: Final = row.savings_estimated_actual_spend if row.savings_estimated_turns else row.spend | ||
| baseline_spend: Final = ( | ||
| comparison_actual + comparison_savings | ||
| if row.savings_estimated_turns or not row.turns |
There was a problem hiding this comment.
When a router group contains historical turns with recorded savings and newer turns with estimates, this code calculates the baseline and percentage from only the newer turns but returns savings for all turns. The dashboard can show $30 saved beside a $1.50 baseline and a −33% badge, while saying savings are based on 4 of 40 requests. Those figures describe different requests and give users a misleading comparison.
Knowledge Base Used: Dashboard and enterprise UI
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit f96bfdd. Configure here.
TLDR
Problem this solves:
How it solves it:
Intentional product change: the savings card's actual spend now includes only requests with a baseline estimate, so both sides cover the same requests. The duplicate actual-spend row is removed, and LLM/classification subtotals are hidden when estimates cover only some requests because those subtotals include all usage. A tooltip explains eligibility and coverage remains visible
User Flow
Before: a request without an estimate inflates actual spend beside the savings comparison
http://127.0.0.1:48182/cost-optimization/and select Auto-Router, then Usagereview-mixedover the last 30 daysAfter: actual and baseline spend describe the same eligible requests
http://127.0.0.1:48184/cost-optimization/and select Auto-Router, then Usagereview-mixedover the last 30 daysRelevant issues
Corrects the historical reporting regression introduced by #41177
Affected release
Since v1.103.0-rc.1
Changes
Preserve signed historical savings, add new savings once, and persist request-date costs and estimate coverage consistently across old and new workers. Historical cost recovery reconciles retained logs against daily counts, spend and tokens within a two-second query limit. Missing comparison data stays unavailable instead of becoming zero
Skip installed schema changes on restart and bound retries for required accounting repairs. Daily totals reuse the savings aggregation without fetching unused usage breakdowns or key metadata. Sessions, selected routers and the CLI compare matching estimated requests
The dashboard uses eligible actual spend in the savings card. Total actual spend remains available through existing accounting APIs. Complete comparisons retain the LLM/classification breakdown; partial comparisons show actual spend, baseline spend, and coverage
Validation
Current revision:
f96bfdd66dbe3e2643dbbe2da422d0e40c14d961, based on474ab91c09, which includes the merged CI fix from #43235All 64 tests across the three affected dashboard files pass, and scoped formatting, lint and budget checks pass. Reverting the display to all-request actual spend makes the mixed-request regression fail: expected
$1.00 / $2.00, received$100.00 / $2.00. The correct implementation was restored before committingThis update changes four existing files, with 40 additions and 53 deletions. Across the PR, tests and fixtures add 621 lines against 652 production lines, a 0.95:1 ratio, excluding migrations, schema copies and generated types
An earlier broader local backend run on
f996803d0bfinished with 867 passed and one failure:test_prometheus_metrics_port_starts_separate_metrics_processexpected exit 2 and received exit 1. Its cause is unresolved. Hostedproxy-server / Run tests,misc / Run tests,assert-ci-coverage, andassert-shard-coveragepassed on that revision. The latest change is confined to the dashboard; hosted CI is rerunning on the current tipPre-Submission checklist
Screenshots / Proof of Fix
Both revisions use real localhost proxies, the rendered dashboard and the same synthetic records in Postgres, with no mocked network responses. Request A costs $1 against a $2 baseline; request B costs $99 and has no baseline estimate. This flow reads recorded charges, so provider inference does not apply
Before (474ab91)
http://127.0.0.1:48182/cost-optimization/, select Auto-Router, Usage, andreview-mixedover the last 30 daysAfter (f96bfdd)
http://127.0.0.1:48184/cost-optimization/, select Auto-Router, Usage, andreview-mixedover the last 30 daysType
Bug Fix
Caveats (if any)
Medium
Low
Final Attestation
Note
Medium Risk
Changes financial reporting paths (daily spend writes, DB triggers, benchmarks API) and requires successful migration install at boot; recovery and partial-coverage logic can leave costs unavailable rather than wrong.
Overview
Adds request-date auto-router cost accounting on
LiteLLM_DailyUserSpend(LLM spend, classifier cost, routed/estimated request counts) plus an idempotent SQL migration that installs columns, a baseline-observation trigger, andapply_autorouter_daily_coverage()after Prisma setup. Spend writes, daily upserts, and queue aggregation now populate and roll up those fields; incomplete daily rows can be reconciled from spend logs within bounded queries.Benchmarks and cost-optimization totals no longer derive savings only from session rollups: headline
saved_spend, spend, baseline, and coverage come from the same daily aggregation as overall cost optimization, with optional historical recovery. Per-router groups still describe whole sessions overlapping the window but keep recordedsaved_spendand historical baselines. Session and CLI responses split full-session baseline from estimate-only baseline and fix partial-coverage savings display.The dashboard savings hero compares actual vs baseline only for requests with a baseline estimate (tooltip + “N of M requests”), hides LLM/classification subtotals when coverage is partial, and surfaces
cost_coverage/ nullable totals fields in the API types.Reviewed by Cursor Bugbot for commit f96bfdd. Bugbot is set up for automated code reviews on this repo. Configure here.