Skip to content

feat(proxy): add LiteLLM_DailyGlobalSpend key-free rollup for the usage dashboard - #41324

Merged
yassin-berriai merged 14 commits into
mainfrom
litellm_daily_global_spend_table
Sep 18, 2026
Merged

yassin-berriai merged 14 commits into
mainfrom
litellm_daily_global_spend_table

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

How it solves it:

  • New LiteLLM_DailyGlobalSpend table keyed on (date, model, model_group, provider, mcp_tool, endpoint), no api_key or user
  • A scheduled job rolls each closed UTC day into it and records how far it got
  • The key-free arm of the usage query reads days through that marker from the global table, later days from the per-key table, in one statement
  • The current day is never rolled up, so pods still flushing it during a rolling deploy cannot leave the global table short
  • Per-key rows are dated by request start, so spend can land on a day that was already rolled up. The job records the database clock of its scan and the next run rewrites every closed day with rows updated since then, whatever the date
  • Everything else (top-N keys, filters, entity tables, the write path) stays on the per-key tables

Built on #41293, which is now in main: this PR is c8a2d8c, ad8de0e, 84c098d, 0601d2b, abf530f and 3449ae9, plus sync merges (225fc53, b92820d, 6a6ae2d) that bring in main. 6a6ae2d picks up main's anyio 4.14.2 lock bump so osv-scan passes; nothing in this PR's own files changed. 0601d2b carries main's new total_response_time_ms and timed_requests columns through the global table, the rollup SQL and the key-free read arm. abf530f answers a Bugbot finding: every marker write now starts from the marker stored in the database (read uncached) instead of the snapshot the run scanned from, so two overlapping runs (Redis unreachable, or the pod lock expiring on a long backfill) can only ever advance it. Before that, the slower run put its older prefix back and dropped scanned_at, which sent usage reads for every day in between back to the per-key table until the next run. 3449ae9 answers the Greptile follow-up on that: the read and the upsert were still two statements, so two runs finishing within milliseconds of each other could interleave between them. The marker is now advanced by one INSERT ... ON CONFLICT DO UPDATE whose SET takes GREATEST of the stored and the incoming reconciled_through and scanned_at, so the maximum is enforced by Postgres inside the row lock and no read is involved

User Flow

Before: a proxy admin on a large deployment opens the Usage page and every widget waits on a query that reads all per-key spend rows for the range

  1. They open https://litellm-domain/ui/?page=usage for the last 30 days
  2. The page calls GET https://litellm-domain/user/daily/activity/aggregated?start_date=...&end_date=...
  3. The totals, per-model, per-provider and per-endpoint numbers come from a scan of every per-key daily row (270k rows in the run below, tens of millions on the deployments in the linked tickets)
  4. On those deployments the request either returns after a long wait or the database side runs out of memory and the page shows an error

After: the same page gets the same numbers from a table with one row per (day, model, provider, endpoint), so the totals arm no longer grows with the number of keys

  1. They open https://litellm-domain/ui/?page=usage for the last 30 days
  2. The page calls the same GET https://litellm-domain/user/daily/activity/aggregated?start_date=...&end_date=...
  3. The totals, per-model, per-provider and per-endpoint numbers for every closed day come from the global table (90 rows in the run below); today's slice and the top-N key breakdown still read per-key rows
  4. The response has the same metadata.total_spend, total_api_requests and per-day breakdowns as before; the first request after upgrading may still read the per-key table until the nightly job (which also runs two minutes after boot) has covered the range
  5. Spend that shows up later for a day already on the page (a retry, a flush that ran past midnight) is in the totals after the next nightly run

Relevant issues

Affected release

Linear ticket

Resolves LIT-7818

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Shared setup: local Postgres seeded with 270,000 LiteLLM_DailyUserSpend rows (3,000 api keys x 3 models x 30 days, 2026-08-01 to 2026-08-30, total spend 13201.65, 1,350,000 requests). Proxy started with python litellm/proxy/proxy_cli.py --config litellm/proxy/dev_config.yaml --use_v2_migration_resolver. Before is captured at the tip of #41293, which this PR stacks on; at the merge base with main the same request 500s (shown in #41293). SUMMARY below is python3 -c "import json; d=json.load(open('$F')); m=d['metadata']; print(len(d['results']), round(m['total_spend'],3), m['total_api_requests'], m['api_key_limit'], len(d['results'][0]['breakdown']['api_keys']))" and EXPLAIN is psql "$DATABASE_URL" -tAc 'EXPLAIN (ANALYZE, SUMMARY) SELECT date, model, custom_llm_provider, endpoint, SUM(spend), SUM(api_requests) FROM "$T" WHERE date BETWEEN '"'"'2026-08-01'"'"' AND '"'"'2026-08-30'"'"' GROUP BY GROUPING SETS ((), (date), (date, model), (model), (custom_llm_provider), (endpoint))' | grep -E "MixedAggregate|Seq Scan"

Before (c3e937b)

Usage totals for a 30-day range

  1. psql "$DATABASE_URL" -tAc 'select count(*) from "LiteLLM_DailyGlobalSpend"; select count(*) from "LiteLLM_Config" where param_name='"'"'daily_global_spend_reconciled_through'"'"''
    0
    0
    
    (the table exists because the migration ships from the installed proxy extras package; nothing writes or reads it)
  2. curl -s -o /tmp/before.json -w "status=%{http_code} bytes=%{size_download} time=%{time_total}s\n" "http://localhost:4000/user/daily/activity/aggregated?start_date=2026-08-01&end_date=2026-08-30" -H "Authorization: Bearer sk-1234"
    status=200 bytes=13560818 time=3.638277s
    
  3. SUMMARY with F=/tmp/before.json
    30 13201.65 1350000 100 100
    
  4. EXPLAIN with T=LiteLLM_DailyUserSpend
    MixedAggregate  (cost=0.00..25018.58 rows=126 width=88) (actual time=212.474..212.497 rows=126 loops=1)
      ->  Seq Scan on "LiteLLM_DailyUserSpend"  (cost=0.00..12867.00 rows=270000 width=64) (actual time=0.012..67.295 rows=270000 loops=1)
    

Spend that lands on an already rolled-up day

  1. Nothing is rolled up at this commit (step 1 above), so every request reads the per-key table and pays the 270,000-row scan in step 4. There is no rollup for a late row to fall behind

Admin UI Usage page

  1. Admin UI dev server (npm run dev in ui/litellm-dashboard, port 3000) pointed at the Before proxy, logged in as admin, open http://localhost:3000/usage, set the date picker to 2026-08-01 through 2026-08-30, Apply. Captured in Chrome at the perf(proxy): split aggregated usage query into key-free rollups and bounded top-N keys #41293 tip 3d2913307a against the same fixture. Cost tab: Total Spend $13,201.65, Daily Spend bars for the 30 days, request cards at 0 (they read the gateway request log, which the fixture does not seed)

    Usage Cost tab before

  2. Key Activity tab: Showing 100 of 100 keys and Only the 100 highest-spend keys of 3,000 are loaded (the perf(proxy): split aggregated usage query into key-free rollups and bounded top-N keys #41293 truncation notice)

Response time metrics

  1. total_response_time_ms and timed_requests landed on main after this capture, so at this commit the response has neither field. On main today they come out of the same 270,000-row per-key scan as every other total

After (3449ae9)

The usage, filter, timezone and response time cases below were captured at 84c098d/0601d2b; abf530f and 3449ae9 only change how the reconcile job writes its marker (git diff 6a6ae2d 3449ae9 --stat is the rollup module and its test), so the read path they exercise is byte-identical. The overlapping runs case is a fresh A/B of 6a6ae2d against 3449ae9

Overlapping reconcile runs

Same 270,000-row fixture. The marker is deleted so every closed day is pending, then two runs of the reconcile job are started against the real Postgres: run B first, slowed to 0.3s per day, and run A 1.5s later at full speed. Once A finishes, the stored marker is sampled every 50ms until B finishes. Base and head are run by pointing PYTHONPATH at a worktree of each commit

  1. Before (6a6ae2d): A finishes with the marker on 2026-08-30. B then writes over it day by day, so for the rest of B's run the marker reads as far back as 2026-08-03 with no scanned_at
    run A: days=30 through=2026-08-30 failed=None
    run B: days=30 through=2026-08-30 failed=None
    marker right after A finished: {'scanned_at': '2026-09-18 20:57:09.069416', 'reconciled_through': '2026-08-30'}
    samples while B still running: 172; lowest reconciled_through seen: 2026-08-03
    samples with reconciled_through below A's marker: 167
    samples with scanned_at missing: 167
    final marker: {'scanned_at': '2026-09-18 20:57:09.026254', 'reconciled_through': '2026-08-30'}
    global table after both runs: {'n': 90, 'spend': '13201.649999999994', 'req': '1350000'}
    
  2. After (3449ae9): same two runs, the marker never moves below 2026-08-30 and keeps a scanned_at in every sample. The final scanned_at is A's, the later of the two scans, since the upsert keeps the greater value. The global table ends identical, which is why the old behaviour was a read-path regression and never a wrong total
    run A: days=30 through=2026-08-30 failed=None
    run B: days=30 through=2026-08-30 failed=None
    marker right after A finished: {'scanned_at': '2026-09-18 21:12:05.203891', 'reconciled_through': '2026-08-30'}
    samples while B still running: 354; lowest reconciled_through seen: 2026-08-30
    samples with reconciled_through below A's marker: 0
    samples with scanned_at missing: 0
    final marker: {'scanned_at': '2026-09-18 21:12:05.203891', 'reconciled_through': '2026-08-30'}
    global table after both runs: {'n': 90, 'spend': '13201.649999999994', 'req': '1350000'}
    
  3. A third, idle run on abf530f (same write path as 3449ae9 when nothing is pending) right after finds nothing pending and only refreshes scanned_at: ReconcileResult(days_reconciled=(), reconciled_through='2026-08-30', failed_day=None), marker before reconciled_through='2026-08-30' scanned_at='2026-09-18 20:57:29.846435', after reconciled_through='2026-08-30' scanned_at='2026-09-18 20:58:04.39293'

Usage totals for a 30-day range

  1. Clear any earlier rollup so the job starts from nothing, then start the proxy: psql "$DATABASE_URL" -tAc 'delete from "LiteLLM_Config" where param_name='"'"'daily_global_spend_reconciled_through'"'"'; delete from "LiteLLM_DailyGlobalSpend"'
  2. Two minutes after boot the job has run: psql "$DATABASE_URL" -tAc 'select param_value from "LiteLLM_Config" where param_name='"'"'daily_global_spend_reconciled_through'"'"'; select count(*), count(distinct date), round(sum(spend)::numeric,3), sum(api_requests) from "LiteLLM_DailyGlobalSpend"'
    {"scanned_at": "2026-09-16 00:46:22.444768", "reconciled_through": "2026-08-30"}
    90|30|13201.650|1350000
    
    The marker stops at the last closed day that has rows, never today (2026-09-16), and records the database clock of the scan
  3. curl -s -o /tmp/after.json -w "status=%{http_code} time=%{time_total}s\n" "http://localhost:4000/user/daily/activity/aggregated?start_date=2026-08-01&end_date=2026-08-30" -H "Authorization: Bearer sk-1234"
    status=200 time=2.900808s
    
  4. SUMMARY with F=/tmp/after.json
    30 13201.65 1350000 100 100
    
  5. EXPLAIN with T=LiteLLM_DailyGlobalSpend
    MixedAggregate  (cost=0.00..9.22 rows=66 width=88) (actual time=0.123..0.145 rows=126 loops=1)
      ->  Seq Scan on "LiteLLM_DailyGlobalSpend"  (cost=0.00..4.35 rows=90 width=64) (actual time=0.013..0.031 rows=90 loops=1)
    
  6. Diffing /tmp/before.json against /tmp/after.json with a 1e-6 float tolerance, skipping the api_keys and api_key_breakdown maps: 0 differences in totals, per-day, per-model, per-model-group, per-provider and per-endpoint values. Inside the key maps two keys at the top-N boundary differ (key-00772, key-00869 vs key-01839, key-02130, all at the same spend), where 3,000 seeded keys share exactly 4 spend values (see Low caveat)

Spend that lands on an already rolled-up day

  1. With the marker from step 2 above in place, rows for two old days arrive the way a delayed flush or retry lands them: psql "$DATABASE_URL" -Atc "INSERT INTO \"LiteLLM_DailyUserSpend\" (id,user_id,date,api_key,model,model_group,custom_llm_provider,mcp_namespaced_tool_name,endpoint,spend,api_requests,successful_requests,updated_at) VALUES ('late-1','user-late','2026-08-03','key-late','gpt-5','gpt-5','openai','','/chat/completions',100.0,7,7,NOW() AT TIME ZONE 'UTC'), ('late-2','user-late','2026-08-10','key-late','gpt-5','gpt-5','openai','','/chat/completions',50.0,3,3,NOW() AT TIME ZONE 'UTC');"
    INSERT 0 2
    
  2. psql "$DATABASE_URL" -Atc 'SELECT round(SUM(spend)::numeric,3), SUM(api_requests) FROM "LiteLLM_DailyUserSpend"' -c 'SELECT round(SUM(spend)::numeric,3), SUM(api_requests) FROM "LiteLLM_DailyGlobalSpend"'
    13351.650|1350010
    13201.650|1350000
    
  3. Until the next run the dashboard reads the closed days from the global table, so the late rows are not in the totals yet: curl -s -o /tmp/stale.json -w "status=%{http_code} time=%{time_total}s\n" "http://localhost:4000/user/daily/activity/aggregated?start_date=2026-08-01&end_date=2026-08-30" -H "Authorization: Bearer sk-1234"; python3 -c "import json; m=json.load(open('/tmp/stale.json'))['metadata']; print(round(m['total_spend'],3), m['total_api_requests'])"
    status=200 time=2.746843s
    13201.65 1350000
    
  4. Restart the proxy so the job runs again (the nightly 00:30 UTC run does the same). Two minutes later only the two touched days were rewritten, the marker did not move back, and the global table matches the per-key table: psql "$DATABASE_URL" -Atc 'SELECT param_value FROM "LiteLLM_Config" WHERE param_name=$$daily_global_spend_reconciled_through$$' -c 'SELECT round(SUM(spend)::numeric,3), SUM(api_requests) FROM "LiteLLM_DailyGlobalSpend"' -c 'SELECT date, MAX(updated_at) FROM "LiteLLM_DailyGlobalSpend" GROUP BY date ORDER BY 2 DESC LIMIT 3'
    {"scanned_at": "2026-09-16 00:50:48.97129", "reconciled_through": "2026-08-30"}
    13351.650|1350010
    2026-08-10|2026-09-16 00:50:49.066
    2026-08-03|2026-09-16 00:50:49.04
    2026-08-30|2026-09-16 00:46:23.082
    
  5. curl -s -o /tmp/late.json -w "status=%{http_code} time=%{time_total}s\n" "http://localhost:4000/user/daily/activity/aggregated?start_date=2026-08-01&end_date=2026-08-30" -H "Authorization: Bearer sk-1234"; python3 -c "import json; m=json.load(open('/tmp/late.json'))['metadata']; print(round(m['total_spend'],3), m['total_api_requests'])"
    status=200 time=2.794093s
    13351.65 1350010
    

Admin UI Usage page

  1. Same dev server pointed at the After proxy, same login, same http://localhost:3000/usage, same 2026-08-01 through 2026-08-30 range, Apply. Captured in Chrome at f631301cfa; abf530f and 3449ae9 only change how the reconcile job writes its marker, so the page reads the same code. Cost tab shows the same Total Spend, Total Tokens and Daily Spend bars as Before, and the page load moved the global table's pg_stat_user_tables.idx_scan from 181 to 183, so the closed days came from the rollup

    Usage Cost tab after

  2. Key Activity tab still says Showing 100 of 100 keys and Only the 100 highest-spend keys of 3,000 are loaded, unchanged by this PR. Browser console and page errors: none

Response time metrics

  1. After main's 20260915000000_add_daily_response_time migration ran at boot, a late per-key row with response time lands on a closed day: psql "$DATABASE_URL" -Atc "INSERT INTO \"LiteLLM_DailyUserSpend\" (id,user_id,date,api_key,model,model_group,custom_llm_provider,mcp_namespaced_tool_name,endpoint,prompt_tokens,completion_tokens,spend,api_requests,successful_requests,failed_requests,total_response_time_ms,timed_requests,updated_at) VALUES ('late-rt-1','user-0001','2026-08-15','key-00001','gpt-5','gpt-5','openai','','/v1/chat/completions',10,5,25.0,4,4,0,9000,4,NOW() AT TIME ZONE 'UTC')"
    INSERT 0 1
    
  2. Two minutes after boot the job has rewritten that day: psql "$DATABASE_URL" -Atc 'select sum(spend), sum(api_requests), sum(total_response_time_ms), sum(timed_requests) from "LiteLLM_DailyGlobalSpend"' -c 'select date, spend, total_response_time_ms, timed_requests from "LiteLLM_DailyGlobalSpend" where date=$$2026-08-15$$ and model=$$gpt-5$$'
    13376.649999999994|1350014|9000|4
    2026-08-15|171.68499999999992|9000|4
    
  3. The dashboard reads those values from the global table: curl -s "http://localhost:4000/user/daily/activity/aggregated?start_date=2026-08-01&end_date=2026-08-30" -H "Authorization: Bearer sk-1234" | python3 -c "import json,sys; d=json.load(sys.stdin); m=d['metadata']; print(m['total_spend'], m['total_api_requests'], m['total_response_time_ms'], m['total_timed_requests']); r=[x for x in d['results'] if x['date']=='2026-08-15'][0]; print(r['metrics']['total_response_time_ms'], r['metrics']['timed_requests'], r['breakdown']['models']['gpt-5']['metrics']['total_response_time_ms'])"
    13376.649999999992 1350014 9000 4
    9000 4 9000
    

Type

🆕 New Feature

Caveats (if any)

Medium

  • The backfill walks every historical day once, one INSERT ... SELECT ... GROUP BY per day, in a background job (not the migration). Until the marker reaches the requested range the usage query silently stays on the per-key table for those days, so a fresh upgrade on a huge history sees no speedup for a while
  • The current day always reads from the per-key table, so a range ending today still scans today's per-key rows (one day, not the whole range)
  • A marker written by ad8de0e (no scanned_at) makes the first run on this commit rewrite every closed day once, the same cost as the initial backfill
  • Redis down means the job runs on every pod without the lock. The day-level upserts are idempotent and the marker only ever advances, so this is wasted work, not wrong data

Low

Live base vs head risk check (/live-pr-risk at 3449ae9, base 6a6ae2d)

Breaking: none observed. The unfiltered aggregate, every filtered read (user, team, key, model, timezone), the current day, late rows on a closed day and two overlapping reconcile runs return the same totals on base and head against the 270,000-row Postgres fixture, with the head answering the unfiltered request where the merge base returned HTTP 500 (the section above)

Backward incompatible, needs a recorded decision from the reviewer: spend that lands on an already rolled-up day (a flush straddling midnight, a retry after an outage) shows on the Usage page only after the next reconcile run (00:30 UTC, or two minutes after a boot), where main shows it on the next page load. LIT-7818 step 1 asked for the global table to be written on the spend write path as well, which would remove this delay at the cost of a seventh table in every flush. This PR ships the reconcile-only design, so a day that got a late row can read low for up to a day. Not fixed in a commit; approve or ask for the write-path arm before merge

Regression risk: none left untested on the dependent graph. The marker format written by 3449ae9 is the same jsonb object (reconciled_through, scanned_at) that read_marker and the read arm consume, verified live by the scheduled run two minutes after boot on this commit and by read_marker in the overlapping runs proof

Dependency graph: _advance_marker is called only from run_daily_global_spend_reconcile and _reconcile_and_record (both in the rollup module, verified live). The stored marker is read by read_marker and reconciled_through, whose only consumers are _scan_pending (verified live) and common_daily_activity's key-free arm (verified live, unfiltered Usage page before and after). LiteLLM_Config upserts elsewhere (ConfigRepository.set_param) never touch this param_name. No dashboard or SDK consumer reads the marker

Not verified: the failure-mode leg (Postgres paused mid-run) and the packaging leg were not run on 3449ae9. The PR moves no runtime pin, and a failed day is covered by the unit suite (the marker stays on the last good day) rather than a live pause

Review gates at 3449ae9

Greptile: confidence 5/5, summary refreshed at 21:20 UTC after the @greptileai trigger on this tip, last reviewed commit 3449ae9, all seven inline threads are resolved. Bugbot: cursor[bot] review at 21:23 UTC on 3449ae9, "found no new issues", triggered once on this tip by the human trigger comment. CodeQL: passed, zero open alerts on the PR ref. Veria: "No security issues found" on 3449ae9, check completed 21:42 UTC after the review request on this tip (https://github.com/BerriAI/litellm/runs/105768955877). Reviewer approval from yuneng-berri is on 6a6ae2d, before abf530f and 3449ae9

CI: every required check passes on 3449ae9. The only red job is the non-required proxy-infra / Run tests, failing on test_login_throttle_settings_are_not_hot_applied_from_the_database, which this PR does not touch. The same single test fails the same way on the last seven main commits, for example https://github.com/BerriAI/litellm/actions/runs/35395850849/job/105764387519 on this PR and https://github.com/BerriAI/litellm/actions/runs/35396309077/job/105765823358 on main head a43a4924a6

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Link to Devin session: https://app.devin.ai/sessions/31421bb901114ac6abff4a10ac140881
Open in Devin Desktop: https://app.devin.ai/desktop/session/31421bb901114ac6abff4a10ac140881?variant=devin
Requested by: @yassin-berriai

Link to Devin session: https://app.devin.ai/sessions/09ed50d5f86a4716931d706cbd7148f2
Open in Devin Desktop: https://app.devin.ai/desktop/session/09ed50d5f86a4716931d706cbd7148f2?variant=devin


Note

Medium Risk
Changes how global usage totals are computed and introduces async reconciliation; incorrect marker or rollup logic could skew dashboard spend until the next successful run, though reads fall back to the per-key table when the marker is unavailable.

Overview
Adds LiteLLM_DailyGlobalSpend, a key-free daily rollup of per-user/key spend, so unfiltered usage dashboard totals can scan far fewer rows.

A new background reconcile job aggregates closed UTC days from LiteLLM_DailyUserSpend, stores progress in LiteLLM_Config (daily_global_spend_reconciled_through), and runs on a nightly cron plus a short post-boot catch-up. Marker advances only forward (including under overlapping pods), rewrites days with late per-key updates, and alerts on failure.

get_daily_activity_aggregated changes the key-free SQL arm only for unfiltered litellm_dailyuserspend reads: days through the marker come from the global table, newer days still from the per-key table (UNION). Filtered reads, entity scoping, and the per-key breakdown path are unchanged.

Reviewed by Cursor Bugbot for commit 3449ae9. Bugbot is set up for automated code reviews on this repo. Configure here.

…ge dashboard

Adds a daily spend table without api_key or user_id, written atomically alongside
LiteLLM_DailyUserSpend from the batched writer, reconciled from history by a
scheduled job that advances a marker in LiteLLM_Config, and read by the key-free
arm of the aggregated usage query once the marker covers the requested range.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR. Add '(aside)' to your comment to have me ignore it.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@codspeed

codspeed Bot commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_daily_global_spend_table (3449ae9) with main (59c24ab)

Open in CodSpeed

@greptile-apps

greptile-apps Bot commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge; no actionable regression remains in the changes since the previous review.

Summary

This PR adds a key-free daily global-spend rollup for unfiltered usage-dashboard totals.

  • Adds the global rollup schema and migration.
  • Reconciles closed UTC days in a scheduled, cross-pod-aware background job and revisits late updates.
  • Splits unfiltered usage reads between reconciled global rows and newer per-key rows.
  • Atomically advances reconciliation metadata so overlapping runs cannot rewind progress.
  • Adds behavioral PostgreSQL coverage for rollup parity, marker advancement, read routing, and scheduler wiring.

Reviews (11) · Last reviewed commit: "fix(proxy): advance the daily global spe..."

Comment thread litellm/proxy/spend_tracking/daily_global_spend_rollup.py Outdated
Comment thread litellm/proxy/management_endpoints/common_daily_activity.py
@codecov

codecov Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 99.36306% with 1 line in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
.../proxy/spend_tracking/daily_global_spend_rollup.py 99.19% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

…d split the key-free read at the marker

The write path no longer dual-writes the global table. The cron rolls up closed UTC days
only, so a pod still flushing the current day can never leave the global table short. The
key-free arm reads days through the marker from the global table and later days from
LiteLLM_DailyUserSpend in one UNION ALL, and the marker comes from the config cache
rather than a per-request database lookup.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/proxy/spend_tracking/daily_global_spend_rollup.py Outdated
…lobal days

The reconcile now records the database clock of the scan behind the last complete
run and, on the next run, rewrites every closed day with per-key rows updated since
then, however old the day is. Replaying only the marker day and the one before it
missed a delayed flush or retry that landed on an older date, and reads through the
marker come from the global table alone, so that spend was never counted.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

yassin-berriai and others added 3 commits September 16, 2026 01:48
…pend

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…_split' into litellm_daily_global_spend_table

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

# Conflicts:
#	litellm/proxy/management_endpoints/common_daily_activity.py
#	tests/test_litellm/proxy/management_endpoints/test_common_daily_activity.py
@devin-ai-integration
devin-ai-integration Bot changed the base branch from main to litellm_usage_key_free_aggregate_split September 16, 2026 11:30
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread tests/test_litellm/proxy/management_endpoints/test_common_daily_activity.py Outdated
…avior, not SQL text or add_job arguments

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/proxy/spend_tracking/daily_global_spend_rollup.py Outdated
Base automatically changed from litellm_usage_key_free_aggregate_split to main September 18, 2026 20:16
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/proxy/spend_tracking/daily_global_spend_rollup.py
…pping reconcile run

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread litellm/proxy/spend_tracking/daily_global_spend_rollup.py Outdated

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

…upsert so overlapping runs cannot rewind it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 3449ae9. Configure here.

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@veria-ai please review 3449ae9: the marker is now advanced by one conditional upsert in Postgres, everything else is unchanged since f631301

@yassin-berriai
yassin-berriai merged commit 87694c2 into main Sep 18, 2026
100 of 102 checks passed
@yassin-berriai
yassin-berriai deleted the litellm_daily_global_spend_table branch September 18, 2026 21:53
yuneng-berri added a commit that referenced this pull request Sep 26, 2026
…pend rollup code from rc/1.103.0 (#43326)

Reverts the code from #41324 (merge 87694c2) and #41293 (merge 2e46b10).
The LiteLLM_DailyGlobalSpend model and migration stay so databases that applied
rc.1 keep a consistent table and rollup marker when the feature returns.
yuneng-berri added a commit that referenced this pull request Sep 27, 2026
… global spend rollup code from rc/1.104.0 (#43385)

* Revert "feat(proxy): server-side Team Usage export beyond the top-N key cap (#42996)"

This reverts commit 77eccac.

* Revert "feat(usage): search team keys beyond the top-N in the Team usage view (#42857)"

This reverts commit c2eb549.

* Revert "feat(usage): search keys beyond the top-N usage subset (#42827)"

This reverts commit 5a8ec13.

* Revert "Merge pull request #41324 from BerriAI/litellm_daily_global_spend_table"

Removes the daily global spend rollup job and its reads. Keeps the
LiteLLM_DailyGlobalSpend model and migration so databases that already
applied it are untouched and no new proxy-extras version is needed.

This reverts commit 87694c2, reversing changes made to its first parent.

* Revert "Merge pull request #41293 from BerriAI/litellm_usage_key_free_aggregate_split"

Restores the aggregated usage query without the top-N key cap, so the
Usage pages, per-key widgets and exports cover every key again.
CacheLeakageCard keeps the date picker removal from #42055.

This reverts commit 2e46b10, reversing changes made to its first parent.

* chore: update Next.js build artifacts (2026-09-27 00:33 UTC, node v24.19.0)
devin-ai-integration Bot added a commit that referenced this pull request Sep 28, 2026
…pend_table"

This reverts commit 87694c2, reversing
changes made to 47209d3.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
devin-ai-integration Bot pushed a commit that referenced this pull request Sep 28, 2026
devin-ai-integration Bot pushed a commit that referenced this pull request Sep 28, 2026
yuneng-berri pushed a commit that referenced this pull request Sep 28, 2026
…r the usage dashboard (#41324)" (#43595)

* Revert "Merge pull request #41324 from BerriAI/litellm_daily_global_spend_table"

* Revert "Merge pull request #41293 from BerriAI/litellm_usage_key_free_aggregate_split" (#43596)

Co-authored-by: yassin <yassin@berri.ai>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>
yuneng-berri pushed a commit that referenced this pull request Sep 28, 2026
…)" (#43378)

* Revert "feat(usage): search keys beyond the top-N usage subset (#42827)"

* revert: "feat(proxy): add LiteLLM_DailyGlobalSpend key-free rollup for the usage dashboard (#41324)" (#43595)

* Revert "Merge pull request #41324 from BerriAI/litellm_daily_global_spend_table"

* Revert "Merge pull request #41293 from BerriAI/litellm_usage_key_free_aggregate_split" (#43596)

Co-authored-by: yassin <yassin@berri.ai>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>
yuneng-berri pushed a commit that referenced this pull request Sep 28, 2026
…sage view (#42857)" (#43377)

* Revert "feat(usage): search team keys beyond the top-N in the Team usage view (#42857)"

* revert: "feat(usage): search keys beyond the top-N usage subset (#42827)" (#43378)

* Revert "feat(usage): search keys beyond the top-N usage subset (#42827)"

* revert: "feat(proxy): add LiteLLM_DailyGlobalSpend key-free rollup for the usage dashboard (#41324)" (#43595)

* Revert "Merge pull request #41324 from BerriAI/litellm_daily_global_spend_table"

* Revert "Merge pull request #41293 from BerriAI/litellm_usage_key_free_aggregate_split" (#43596)

Co-authored-by: yassin <yassin@berri.ai>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants