Repository navigation
feat(proxy): add LiteLLM_DailyGlobalSpend key-free rollup for the usage dashboard - #41324
Conversation
…ge dashboard Adds a daily spend table without api_key or user_id, written atomically alongside LiteLLM_DailyUserSpend from the batched writer, reconciled from history by a scheduled job that advances a marker in LiteLLM_Config, and read by the key-free arm of the aggregated usage query once the marker covers the requested range. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
🤖 Devin AI EngineerI'll be helping with this pull request! Here's what you should know: ✅ I will automatically:
Note: I can only respond to comments from users who have write access to this repository. ⚙️ Control Options:
|
|
|
|
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
…d split the key-free read at the marker The write path no longer dual-writes the global table. The cron rolls up closed UTC days only, so a pod still flushing the current day can never leave the global table short. The key-free arm reads days through the marker from the global table and later days from LiteLLM_DailyUserSpend in one UNION ALL, and the marker comes from the config cache rather than a per-request database lookup. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
bugbot run |
…lobal days The reconcile now records the database clock of the scan behind the last complete run and, on the next run, rewrites every closed day with per-key rows updated since then, however old the day is. Replaying only the marker day and the one before it missed a delayed flush or retry that landed on an older date, and reads through the marker come from the global table alone, so that spend was never counted. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
bugbot run |
…ily_global_spend_table
…pend Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…_split' into litellm_daily_global_spend_table Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> # Conflicts: # litellm/proxy/management_endpoints/common_daily_activity.py # tests/test_litellm/proxy/management_endpoints/test_common_daily_activity.py
…avior, not SQL text or add_job arguments Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
bugbot run |
|
bugbot run |
…pping reconcile run Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
bugbot run |
…upsert so overlapping runs cannot rewind it Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 3449ae9. Configure here.
… global spend rollup code from rc/1.104.0 (#43385) * Revert "feat(proxy): server-side Team Usage export beyond the top-N key cap (#42996)" This reverts commit 77eccac. * Revert "feat(usage): search team keys beyond the top-N in the Team usage view (#42857)" This reverts commit c2eb549. * Revert "feat(usage): search keys beyond the top-N usage subset (#42827)" This reverts commit 5a8ec13. * Revert "Merge pull request #41324 from BerriAI/litellm_daily_global_spend_table" Removes the daily global spend rollup job and its reads. Keeps the LiteLLM_DailyGlobalSpend model and migration so databases that already applied it are untouched and no new proxy-extras version is needed. This reverts commit 87694c2, reversing changes made to its first parent. * Revert "Merge pull request #41293 from BerriAI/litellm_usage_key_free_aggregate_split" Restores the aggregated usage query without the top-N key cap, so the Usage pages, per-key widgets and exports cover every key again. CacheLeakageCard keeps the date picker removal from #42055. This reverts commit 2e46b10, reversing changes made to its first parent. * chore: update Next.js build artifacts (2026-09-27 00:33 UTC, node v24.19.0)
…r the usage dashboard (#41324)" (#43595) * Revert "Merge pull request #41324 from BerriAI/litellm_daily_global_spend_table" * Revert "Merge pull request #41293 from BerriAI/litellm_usage_key_free_aggregate_split" (#43596) Co-authored-by: yassin <yassin@berri.ai> --------- Co-authored-by: yassin <yassin@berri.ai> Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…)" (#43378) * Revert "feat(usage): search keys beyond the top-N usage subset (#42827)" * revert: "feat(proxy): add LiteLLM_DailyGlobalSpend key-free rollup for the usage dashboard (#41324)" (#43595) * Revert "Merge pull request #41324 from BerriAI/litellm_daily_global_spend_table" * Revert "Merge pull request #41293 from BerriAI/litellm_usage_key_free_aggregate_split" (#43596) Co-authored-by: yassin <yassin@berri.ai> --------- Co-authored-by: yassin <yassin@berri.ai> Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…sage view (#42857)" (#43377) * Revert "feat(usage): search team keys beyond the top-N in the Team usage view (#42857)" * revert: "feat(usage): search keys beyond the top-N usage subset (#42827)" (#43378) * Revert "feat(usage): search keys beyond the top-N usage subset (#42827)" * revert: "feat(proxy): add LiteLLM_DailyGlobalSpend key-free rollup for the usage dashboard (#41324)" (#43595) * Revert "Merge pull request #41324 from BerriAI/litellm_daily_global_spend_table" * Revert "Merge pull request #41293 from BerriAI/litellm_usage_key_free_aggregate_split" (#43596) Co-authored-by: yassin <yassin@berri.ai> --------- Co-authored-by: yassin <yassin@berri.ai> Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>
TLDR
Problem this solves:
How it solves it:
LiteLLM_DailyGlobalSpendtable keyed on (date, model, model_group, provider, mcp_tool, endpoint), no api_key or userBuilt on #41293, which is now in main: this PR is
c8a2d8c,ad8de0e,84c098d,0601d2b,abf530fand3449ae9, plus sync merges (225fc53,b92820d,6a6ae2d) that bring in main.6a6ae2dpicks up main's anyio 4.14.2 lock bump so osv-scan passes; nothing in this PR's own files changed.0601d2bcarries main's newtotal_response_time_msandtimed_requestscolumns through the global table, the rollup SQL and the key-free read arm.abf530fanswers a Bugbot finding: every marker write now starts from the marker stored in the database (read uncached) instead of the snapshot the run scanned from, so two overlapping runs (Redis unreachable, or the pod lock expiring on a long backfill) can only ever advance it. Before that, the slower run put its older prefix back and droppedscanned_at, which sent usage reads for every day in between back to the per-key table until the next run.3449ae9answers the Greptile follow-up on that: the read and the upsert were still two statements, so two runs finishing within milliseconds of each other could interleave between them. The marker is now advanced by oneINSERT ... ON CONFLICT DO UPDATEwhoseSETtakesGREATESTof the stored and the incomingreconciled_throughandscanned_at, so the maximum is enforced by Postgres inside the row lock and no read is involvedUser Flow
Before: a proxy admin on a large deployment opens the Usage page and every widget waits on a query that reads all per-key spend rows for the range
After: the same page gets the same numbers from a table with one row per (day, model, provider, endpoint), so the totals arm no longer grows with the number of keys
metadata.total_spend,total_api_requestsand per-day breakdowns as before; the first request after upgrading may still read the per-key table until the nightly job (which also runs two minutes after boot) has covered the rangeRelevant issues
Affected release
Linear ticket
Resolves LIT-7818
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Shared setup: local Postgres seeded with 270,000
LiteLLM_DailyUserSpendrows (3,000 api keys x 3 models x 30 days, 2026-08-01 to 2026-08-30, total spend 13201.65, 1,350,000 requests). Proxy started withpython litellm/proxy/proxy_cli.py --config litellm/proxy/dev_config.yaml --use_v2_migration_resolver. Before is captured at the tip of #41293, which this PR stacks on; at the merge base with main the same request 500s (shown in #41293).SUMMARYbelow ispython3 -c "import json; d=json.load(open('$F')); m=d['metadata']; print(len(d['results']), round(m['total_spend'],3), m['total_api_requests'], m['api_key_limit'], len(d['results'][0]['breakdown']['api_keys']))"andEXPLAINispsql "$DATABASE_URL" -tAc 'EXPLAIN (ANALYZE, SUMMARY) SELECT date, model, custom_llm_provider, endpoint, SUM(spend), SUM(api_requests) FROM "$T" WHERE date BETWEEN '"'"'2026-08-01'"'"' AND '"'"'2026-08-30'"'"' GROUP BY GROUPING SETS ((), (date), (date, model), (model), (custom_llm_provider), (endpoint))' | grep -E "MixedAggregate|Seq Scan"Before (c3e937b)
Usage totals for a 30-day range
psql "$DATABASE_URL" -tAc 'select count(*) from "LiteLLM_DailyGlobalSpend"; select count(*) from "LiteLLM_Config" where param_name='"'"'daily_global_spend_reconciled_through'"'"''curl -s -o /tmp/before.json -w "status=%{http_code} bytes=%{size_download} time=%{time_total}s\n" "http://localhost:4000/user/daily/activity/aggregated?start_date=2026-08-01&end_date=2026-08-30" -H "Authorization: Bearer sk-1234"Spend that lands on an already rolled-up day
Admin UI Usage page
Admin UI dev server (
npm run devinui/litellm-dashboard, port 3000) pointed at the Before proxy, logged in as admin, open http://localhost:3000/usage, set the date picker to 2026-08-01 through 2026-08-30, Apply. Captured in Chrome at the perf(proxy): split aggregated usage query into key-free rollups and bounded top-N keys #41293 tip3d2913307aagainst the same fixture. Cost tab: Total Spend $13,201.65, Daily Spend bars for the 30 days, request cards at 0 (they read the gateway request log, which the fixture does not seed)Key Activity tab:
Showing 100 of 100 keysandOnly the 100 highest-spend keys of 3,000 are loaded(the perf(proxy): split aggregated usage query into key-free rollups and bounded top-N keys #41293 truncation notice)Response time metrics
total_response_time_msandtimed_requestslanded on main after this capture, so at this commit the response has neither field. On main today they come out of the same 270,000-row per-key scan as every other totalAfter (3449ae9)
The usage, filter, timezone and response time cases below were captured at
84c098d/0601d2b;abf530fand3449ae9only change how the reconcile job writes its marker (git diff 6a6ae2d 3449ae9 --statis the rollup module and its test), so the read path they exercise is byte-identical. The overlapping runs case is a fresh A/B of6a6ae2dagainst3449ae9Overlapping reconcile runs
Same 270,000-row fixture. The marker is deleted so every closed day is pending, then two runs of the reconcile job are started against the real Postgres: run B first, slowed to 0.3s per day, and run A 1.5s later at full speed. Once A finishes, the stored marker is sampled every 50ms until B finishes. Base and head are run by pointing
PYTHONPATHat a worktree of each commit6a6ae2d): A finishes with the marker on 2026-08-30. B then writes over it day by day, so for the rest of B's run the marker reads as far back as 2026-08-03 with noscanned_at3449ae9): same two runs, the marker never moves below 2026-08-30 and keeps ascanned_atin every sample. The finalscanned_atis A's, the later of the two scans, since the upsert keeps the greater value. The global table ends identical, which is why the old behaviour was a read-path regression and never a wrong totalabf530f(same write path as3449ae9when nothing is pending) right after finds nothing pending and only refreshesscanned_at:ReconcileResult(days_reconciled=(), reconciled_through='2026-08-30', failed_day=None), marker beforereconciled_through='2026-08-30' scanned_at='2026-09-18 20:57:29.846435', afterreconciled_through='2026-08-30' scanned_at='2026-09-18 20:58:04.39293'Usage totals for a 30-day range
psql "$DATABASE_URL" -tAc 'delete from "LiteLLM_Config" where param_name='"'"'daily_global_spend_reconciled_through'"'"'; delete from "LiteLLM_DailyGlobalSpend"'psql "$DATABASE_URL" -tAc 'select param_value from "LiteLLM_Config" where param_name='"'"'daily_global_spend_reconciled_through'"'"'; select count(*), count(distinct date), round(sum(spend)::numeric,3), sum(api_requests) from "LiteLLM_DailyGlobalSpend"'curl -s -o /tmp/after.json -w "status=%{http_code} time=%{time_total}s\n" "http://localhost:4000/user/daily/activity/aggregated?start_date=2026-08-01&end_date=2026-08-30" -H "Authorization: Bearer sk-1234"/tmp/before.jsonagainst/tmp/after.jsonwith a 1e-6 float tolerance, skipping theapi_keysandapi_key_breakdownmaps: 0 differences in totals, per-day, per-model, per-model-group, per-provider and per-endpoint values. Inside the key maps two keys at the top-N boundary differ (key-00772,key-00869vskey-01839,key-02130, all at the same spend), where 3,000 seeded keys share exactly 4 spend values (see Low caveat)Spend that lands on an already rolled-up day
psql "$DATABASE_URL" -Atc "INSERT INTO \"LiteLLM_DailyUserSpend\" (id,user_id,date,api_key,model,model_group,custom_llm_provider,mcp_namespaced_tool_name,endpoint,spend,api_requests,successful_requests,updated_at) VALUES ('late-1','user-late','2026-08-03','key-late','gpt-5','gpt-5','openai','','/chat/completions',100.0,7,7,NOW() AT TIME ZONE 'UTC'), ('late-2','user-late','2026-08-10','key-late','gpt-5','gpt-5','openai','','/chat/completions',50.0,3,3,NOW() AT TIME ZONE 'UTC');"psql "$DATABASE_URL" -Atc 'SELECT round(SUM(spend)::numeric,3), SUM(api_requests) FROM "LiteLLM_DailyUserSpend"' -c 'SELECT round(SUM(spend)::numeric,3), SUM(api_requests) FROM "LiteLLM_DailyGlobalSpend"'curl -s -o /tmp/stale.json -w "status=%{http_code} time=%{time_total}s\n" "http://localhost:4000/user/daily/activity/aggregated?start_date=2026-08-01&end_date=2026-08-30" -H "Authorization: Bearer sk-1234"; python3 -c "import json; m=json.load(open('/tmp/stale.json'))['metadata']; print(round(m['total_spend'],3), m['total_api_requests'])"psql "$DATABASE_URL" -Atc 'SELECT param_value FROM "LiteLLM_Config" WHERE param_name=$$daily_global_spend_reconciled_through$$' -c 'SELECT round(SUM(spend)::numeric,3), SUM(api_requests) FROM "LiteLLM_DailyGlobalSpend"' -c 'SELECT date, MAX(updated_at) FROM "LiteLLM_DailyGlobalSpend" GROUP BY date ORDER BY 2 DESC LIMIT 3'curl -s -o /tmp/late.json -w "status=%{http_code} time=%{time_total}s\n" "http://localhost:4000/user/daily/activity/aggregated?start_date=2026-08-01&end_date=2026-08-30" -H "Authorization: Bearer sk-1234"; python3 -c "import json; m=json.load(open('/tmp/late.json'))['metadata']; print(round(m['total_spend'],3), m['total_api_requests'])"Admin UI Usage page
Same dev server pointed at the After proxy, same login, same http://localhost:3000/usage, same 2026-08-01 through 2026-08-30 range, Apply. Captured in Chrome at
f631301cfa;abf530fand3449ae9only change how the reconcile job writes its marker, so the page reads the same code. Cost tab shows the same Total Spend, Total Tokens and Daily Spend bars as Before, and the page load moved the global table'spg_stat_user_tables.idx_scanfrom 181 to 183, so the closed days came from the rollupKey Activity tab still says
Showing 100 of 100 keysandOnly the 100 highest-spend keys of 3,000 are loaded, unchanged by this PR. Browser console and page errors: noneResponse time metrics
20260915000000_add_daily_response_timemigration ran at boot, a late per-key row with response time lands on a closed day:psql "$DATABASE_URL" -Atc "INSERT INTO \"LiteLLM_DailyUserSpend\" (id,user_id,date,api_key,model,model_group,custom_llm_provider,mcp_namespaced_tool_name,endpoint,prompt_tokens,completion_tokens,spend,api_requests,successful_requests,failed_requests,total_response_time_ms,timed_requests,updated_at) VALUES ('late-rt-1','user-0001','2026-08-15','key-00001','gpt-5','gpt-5','openai','','/v1/chat/completions',10,5,25.0,4,4,0,9000,4,NOW() AT TIME ZONE 'UTC')"psql "$DATABASE_URL" -Atc 'select sum(spend), sum(api_requests), sum(total_response_time_ms), sum(timed_requests) from "LiteLLM_DailyGlobalSpend"' -c 'select date, spend, total_response_time_ms, timed_requests from "LiteLLM_DailyGlobalSpend" where date=$$2026-08-15$$ and model=$$gpt-5$$'curl -s "http://localhost:4000/user/daily/activity/aggregated?start_date=2026-08-01&end_date=2026-08-30" -H "Authorization: Bearer sk-1234" | python3 -c "import json,sys; d=json.load(sys.stdin); m=d['metadata']; print(m['total_spend'], m['total_api_requests'], m['total_response_time_ms'], m['total_timed_requests']); r=[x for x in d['results'] if x['date']=='2026-08-15'][0]; print(r['metrics']['total_response_time_ms'], r['metrics']['timed_requests'], r['breakdown']['models']['gpt-5']['metrics']['total_response_time_ms'])"Type
🆕 New Feature
Caveats (if any)
Medium
INSERT ... SELECT ... GROUP BYper day, in a background job (not the migration). Until the marker reaches the requested range the usage query silently stays on the per-key table for those days, so a fresh upgrade on a huge history sees no speedup for a whilead8de0e(noscanned_at) makes the first run on this commit rewrite every closed day once, the same cost as the initial backfillLow
''where the per-key tables store NULL, so the unique index can dedupe them. The read path treats both as "not set" alreadySUM(spend) DESC, api_key; with float sums and exact ties the boundary can move between runs. Pre-existing in perf(proxy): split aggregated usage query into key-free rollups and bounded top-N keys #41293, this PR does not change that arm; fix goes theregen_random_uuid(), core Postgres since 13, so no pgcrypto extension is neededLive base vs head risk check (
/live-pr-riskat3449ae9, base6a6ae2d)Breaking: none observed. The unfiltered aggregate, every filtered read (user, team, key, model, timezone), the current day, late rows on a closed day and two overlapping reconcile runs return the same totals on base and head against the 270,000-row Postgres fixture, with the head answering the unfiltered request where the merge base returned HTTP 500 (the section above)
Backward incompatible, needs a recorded decision from the reviewer: spend that lands on an already rolled-up day (a flush straddling midnight, a retry after an outage) shows on the Usage page only after the next reconcile run (00:30 UTC, or two minutes after a boot), where main shows it on the next page load. LIT-7818 step 1 asked for the global table to be written on the spend write path as well, which would remove this delay at the cost of a seventh table in every flush. This PR ships the reconcile-only design, so a day that got a late row can read low for up to a day. Not fixed in a commit; approve or ask for the write-path arm before merge
Regression risk: none left untested on the dependent graph. The marker format written by
3449ae9is the same jsonb object (reconciled_through,scanned_at) thatread_markerand the read arm consume, verified live by the scheduled run two minutes after boot on this commit and byread_markerin the overlapping runs proofDependency graph:
_advance_markeris called only fromrun_daily_global_spend_reconcileand_reconcile_and_record(both in the rollup module, verified live). The stored marker is read byread_markerandreconciled_through, whose only consumers are_scan_pending(verified live) andcommon_daily_activity's key-free arm (verified live, unfiltered Usage page before and after).LiteLLM_Configupserts elsewhere (ConfigRepository.set_param) never touch thisparam_name. No dashboard or SDK consumer reads the markerNot verified: the failure-mode leg (Postgres paused mid-run) and the packaging leg were not run on
3449ae9. The PR moves no runtime pin, and a failed day is covered by the unit suite (the marker stays on the last good day) rather than a live pauseReview gates at
3449ae9Greptile: confidence 5/5, summary refreshed at 21:20 UTC after the
@greptileaitrigger on this tip, last reviewed commit3449ae9, all seven inline threads are resolved. Bugbot: cursor[bot] review at 21:23 UTC on3449ae9, "found no new issues", triggered once on this tip by the human trigger comment. CodeQL: passed, zero open alerts on the PR ref. Veria: "No security issues found" on3449ae9, check completed 21:42 UTC after the review request on this tip (https://github.com/BerriAI/litellm/runs/105768955877). Reviewer approval from yuneng-berri is on6a6ae2d, beforeabf530fand3449ae9CI: every required check passes on
3449ae9. The only red job is the non-requiredproxy-infra / Run tests, failing ontest_login_throttle_settings_are_not_hot_applied_from_the_database, which this PR does not touch. The same single test fails the same way on the last sevenmaincommits, for example https://github.com/BerriAI/litellm/actions/runs/35395850849/job/105764387519 on this PR and https://github.com/BerriAI/litellm/actions/runs/35396309077/job/105765823358 onmainheada43a4924a6Final Attestation
Link to Devin session: https://app.devin.ai/sessions/31421bb901114ac6abff4a10ac140881
Open in Devin Desktop: https://app.devin.ai/desktop/session/31421bb901114ac6abff4a10ac140881?variant=devin
Requested by: @yassin-berriai
Link to Devin session: https://app.devin.ai/sessions/09ed50d5f86a4716931d706cbd7148f2
Open in Devin Desktop: https://app.devin.ai/desktop/session/09ed50d5f86a4716931d706cbd7148f2?variant=devin
Note
Medium Risk
Changes how global usage totals are computed and introduces async reconciliation; incorrect marker or rollup logic could skew dashboard spend until the next successful run, though reads fall back to the per-key table when the marker is unavailable.
Overview
Adds
LiteLLM_DailyGlobalSpend, a key-free daily rollup of per-user/key spend, so unfiltered usage dashboard totals can scan far fewer rows.A new background reconcile job aggregates closed UTC days from
LiteLLM_DailyUserSpend, stores progress inLiteLLM_Config(daily_global_spend_reconciled_through), and runs on a nightly cron plus a short post-boot catch-up. Marker advances only forward (including under overlapping pods), rewrites days with late per-key updates, and alerts on failure.get_daily_activity_aggregatedchanges the key-free SQL arm only for unfilteredlitellm_dailyuserspendreads: days through the marker come from the global table, newer days still from the per-key table (UNION). Filtered reads, entity scoping, and the per-key breakdown path are unchanged.Reviewed by Cursor Bugbot for commit 3449ae9. Bugbot is set up for automated code reviews on this repo. Configure here.