Repository navigation
refactor(repositories): daily activity repository with centralized bounded usage queries - #43398
Conversation
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
|
|
Codecov Report❌ Patch coverage is 📢 Thoughts on this report? Let us know! |
|
1 similar comment
b8a9b52 to
52c2990
Compare
f834012 to
1165fe8
Compare
3c4cc18 to
17c22bc
Compare
1165fe8 to
d032e61
Compare
21bf013 to
4df79af
Compare
|
@greptileai review |
|
bugbot run |
4df79af to
a37bb58
Compare
|
@greptileai re review head a37bb58 (rebased on main, search and model top keys now rank by SUM(spend::numeric)) |
|
bugbot run |
a37bb58 to
0e475ac
Compare
|
@greptileai re review head 0e475ac (rebased on main; CodeQL py/mixed-returns fixed: _daily_rows_table and _export_grouping are explicit if chains ending in assert_never) |
|
bugbot run |
…unded usage queries Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
0e475ac to
7b25632
Compare
|
@greptileai review |
|
bugbot run |
|
Live retest on head 7b25632 (this branch on port 4000, Unassigned entity rollup, team rows on 2026-06-17 with
Regression checks on the 2026-06-15 fixture (106 keys plus PTU): head returns 100 of 106 keys with Admin UI usage page on the same head, proof-user on 2026-06-15: |
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 7b25632. Configure here.
| FROM scoped | ||
| {" ".join(joins)} | ||
| WHERE TRUE{cursor_clause} | ||
| GROUP BY {", ".join(grouping_keys)} |
There was a problem hiding this comment.
Model export SQL grouping mismatch
Medium Severity
DAILY_WITH_MODELS export selects NULLIF(COALESCE(scoped.model, ''), '') while grouping only by COALESCE(scoped.model, ''). Postgres requires non-aggregated SELECT expressions to match GROUP BY exactly, so this query is rejected at runtime and export_rows cannot produce a model breakdown.
Additional Locations (1)
Reviewed by Cursor Bugbot for commit 7b25632. Configure here.
There was a problem hiding this comment.
Not a bug. Postgres accepts a select-list expression whose only ungrouped column reference appears inside a subexpression that textually matches a GROUP BY expression, so NULLIF(COALESCE(scoped.model, ''), '') is valid when grouping by COALESCE(scoped.model, '').
Verified on PostgreSQL 14.24:
SELECT NULLIF(COALESCE(m, ''), '') AS model, COUNT(*)
FROM (VALUES ('gpt-4o'), (NULL), ('')) t(m)
GROUP BY COALESCE(m, '') ORDER BY 1;
model | count
--------+-------
gpt-4o | 1
| 2
and the generated query itself: build_export_sql(scope, export_type=ExportType.DAILY_WITH_MODELS, after=None, batch_size=1000) executed against the live LiteLLM_DailyTeamSpend fixture on this head returns rows (2 for the June 2026 team scope), same as the DAILY, DAILY_WITH_KEYS and DAILY_WITH_USERS variants.
* test(proxy): stop the proxy_server app fixture leaking LITELLM_LOG The session app fixture set LITELLM_LOG=ERROR with os.environ.setdefault and never removed it, so later tests on the same xdist worker inherited it. test_drop_params_env_var spawns a subprocess with os.environ and lost the warning it asserts on. Scope the variable to the import with a MonkeyPatch context * test(secret-detection): give the hand-built redaction request an ASGI path Since #43975 _read_request_body checks the route path via request.scope, and a scope without path raised KeyError that was swallowed into an empty body, so chat_completion failed with a missing messages parameter. Real ASGI scopes always carry path * test(integration): isolate litellm callback lists per sdk test usage-based-routing-v2 Routers register their selector in litellm.callbacks and nothing removes it, not even Router.reset(). The counter TTL and Redis service metrics tests left their selectors behind, and the next usage routing test ran their pre-call checks against its own rpm=1 deployments, raising "Deployment over defined rpm limit". An autouse fixture now gives each sdk test copies of the callback lists and restores the originals afterwards * test(integration): keep the owner-lookup fault proxy off the shared read replica The owned proxy points DATABASE_URL at a scratch database but inherited DATABASE_URL_READ_REPLICA from the replica job, so auth read the shared database and rejected the freshly created key with token_not_found_in_db. Drop the replica variable like the other scratch-database owned proxies * test(integration): request every seeded key in the team owner breakdown The aggregated team activity endpoint now caps breakdown.api_keys at the top 100 keys by default (#43398), so the 300 seeded keys came back as 100 rows. The test guarantees each key is reported with its own owner, so ask for an api_key_limit that covers all seeded keys * test(integration): give every owned Redis its own port in the redis-cache container On CircleCI every owned Redis ran on the fixed port 16379 inside the shared redis-cache container. When an earlier server still held that port, the new one failed to bind, readiness pinged the old server, the pidfile read failed and cleanup then reported "Owned Redis still serves after shutdown" Reserve an ephemeral port for the docker-exec path the same way the local binary path already does, and refuse to start when something already serves the chosen port so the failure names the real cause * test(e2e): skip the Vertex Mistral partner case the e2e project cannot reach The e2e Vertex project gets a 404 publisher model not found for vertex_ai/mistral-small-2503, so the case can only fail * test(e2e): skip the Vertex gpt-oss partner case the e2e project never serves vertex_ai/openai/gpt-oss-120b-maas has hit a 60s read timeout with no response headers on every run in the e2e Vertex project since the case was ported, and no other Vertex partner chat model passes there to switch to * test(e2e): check only stored message content for a leaked card number The Presidio spend-log check ran the card-number pattern over the whole serialized response, so a Luhn-valid usage.cost float (0.0003466000000000001) failed the streaming /v1/messages case although the stored content was <CREDIT_CARD>. The check now reads the content and text strings of the stored response, which is where a raw card would land, and still requires the placeholder there * test(e2e): assert the proxy decodes token-array embeddings for titan The port in #44120 carried over a legacy SDK-direct test that expected Bedrock to reject token ids with a 400. Through the proxy, /embeddings decodes token arrays to text for providers that cannot embed tokens, so titan answers 200. The test now sends a token array and its decoded sentence and requires the two vectors to match, which fails if the proxy stops decoding or decodes with the wrong tokenizer * test(e2e): run the Bedrock extended-thinking round trip on a model that honors enabled thinking us.anthropic.claude-sonnet-5-5 is adaptive-only, so litellm sends thinking.type=enabled with a 1024 budget as adaptive with low effort, and Bedrock returned no reasoning blocks on 5 of 5 identical Converse calls (boto3 direct agreed). us.anthropic.claude-sonnet-4-6 accepts the legacy shape verbatim and returned reasoning on 5 of 5. The non-thinking Bedrock case stays on sonnet-5-5 * test(proxy): stop unit modules forcing DEBUG logging into the event-loop lag tests Five tests/unit modules set verbose_proxy_logger to DEBUG at import, so every xdist worker that collected them logged the 2.4MB pass-through response from a worker thread, and secret redaction of that line held the GIL for ~0.8s+ inside the timed window. The lag tests now pin the LiteLLM loggers to WARNING and freeze gc while timing, and the module-level DEBUG overrides are removed * test(e2e): cite the tokenizer and date behind the titan token-array fixture * test(e2e): let migration seed replicas finish their request-log indexes before cloning Since #43948 a serving proxy builds the two LiteLLM_SpendLogs indexes on a background thread after it reports ready. The seed fixtures stopped the replica at readiness, so every cloned legacy database lacked an index no real deployment would be missing, and the v2 baseline diff refused it. Seeds now wait until both indexes exist and are valid in the database's schema * test(passthrough): give the pass-through MockRequest an httpx URL and ASGI scope #43626 made get_request_route read request.scope during pass-through kwarg setup; the MockRequest in tests/unit/passthrough had neither a scope nor a URL object, so both stream-param tests raised before reaching the code they check. Mirrors the repair #43626 made to the tests/pass_through_unit_tests fake * test(integration): ignore foreign allow_all_keys MCP servers in the access matrix tool list test_toolset_gateway_url_serves_a_team_granted_toolset_to_a_key_without_its_own_grant (#43908) registers an allow_all_keys server on the shared gateway, and allow_all_keys servers are listed to every key by design, so a matrix case running on another xdist worker at the same time saw its tools. The matrix now drops tools of allow_all_keys servers it did not create, read from LiteLLM_MCPServerTable before and after listing, and still compares everything else exactly
…) (#44254) * test(proxy): stop the proxy_server app fixture leaking LITELLM_LOG The session app fixture set LITELLM_LOG=ERROR with os.environ.setdefault and never removed it, so later tests on the same xdist worker inherited it. test_drop_params_env_var spawns a subprocess with os.environ and lost the warning it asserts on. Scope the variable to the import with a MonkeyPatch context * test(secret-detection): give the hand-built redaction request an ASGI path Since #43975 _read_request_body checks the route path via request.scope, and a scope without path raised KeyError that was swallowed into an empty body, so chat_completion failed with a missing messages parameter. Real ASGI scopes always carry path * test(integration): isolate litellm callback lists per sdk test usage-based-routing-v2 Routers register their selector in litellm.callbacks and nothing removes it, not even Router.reset(). The counter TTL and Redis service metrics tests left their selectors behind, and the next usage routing test ran their pre-call checks against its own rpm=1 deployments, raising "Deployment over defined rpm limit". An autouse fixture now gives each sdk test copies of the callback lists and restores the originals afterwards * test(integration): keep the owner-lookup fault proxy off the shared read replica The owned proxy points DATABASE_URL at a scratch database but inherited DATABASE_URL_READ_REPLICA from the replica job, so auth read the shared database and rejected the freshly created key with token_not_found_in_db. Drop the replica variable like the other scratch-database owned proxies * test(integration): request every seeded key in the team owner breakdown The aggregated team activity endpoint now caps breakdown.api_keys at the top 100 keys by default (#43398), so the 300 seeded keys came back as 100 rows. The test guarantees each key is reported with its own owner, so ask for an api_key_limit that covers all seeded keys * test(integration): give every owned Redis its own port in the redis-cache container On CircleCI every owned Redis ran on the fixed port 16379 inside the shared redis-cache container. When an earlier server still held that port, the new one failed to bind, readiness pinged the old server, the pidfile read failed and cleanup then reported "Owned Redis still serves after shutdown" Reserve an ephemeral port for the docker-exec path the same way the local binary path already does, and refuse to start when something already serves the chosen port so the failure names the real cause * test(e2e): skip the Vertex Mistral partner case the e2e project cannot reach The e2e Vertex project gets a 404 publisher model not found for vertex_ai/mistral-small-2503, so the case can only fail * test(e2e): skip the Vertex gpt-oss partner case the e2e project never serves vertex_ai/openai/gpt-oss-120b-maas has hit a 60s read timeout with no response headers on every run in the e2e Vertex project since the case was ported, and no other Vertex partner chat model passes there to switch to * test(e2e): check only stored message content for a leaked card number The Presidio spend-log check ran the card-number pattern over the whole serialized response, so a Luhn-valid usage.cost float (0.0003466000000000001) failed the streaming /v1/messages case although the stored content was <CREDIT_CARD>. The check now reads the content and text strings of the stored response, which is where a raw card would land, and still requires the placeholder there * test(e2e): assert the proxy decodes token-array embeddings for titan The port in #44120 carried over a legacy SDK-direct test that expected Bedrock to reject token ids with a 400. Through the proxy, /embeddings decodes token arrays to text for providers that cannot embed tokens, so titan answers 200. The test now sends a token array and its decoded sentence and requires the two vectors to match, which fails if the proxy stops decoding or decodes with the wrong tokenizer * test(e2e): run the Bedrock extended-thinking round trip on a model that honors enabled thinking us.anthropic.claude-sonnet-5-5 is adaptive-only, so litellm sends thinking.type=enabled with a 1024 budget as adaptive with low effort, and Bedrock returned no reasoning blocks on 5 of 5 identical Converse calls (boto3 direct agreed). us.anthropic.claude-sonnet-4-6 accepts the legacy shape verbatim and returned reasoning on 5 of 5. The non-thinking Bedrock case stays on sonnet-5-5 * test(proxy): stop unit modules forcing DEBUG logging into the event-loop lag tests Five tests/unit modules set verbose_proxy_logger to DEBUG at import, so every xdist worker that collected them logged the 2.4MB pass-through response from a worker thread, and secret redaction of that line held the GIL for ~0.8s+ inside the timed window. The lag tests now pin the LiteLLM loggers to WARNING and freeze gc while timing, and the module-level DEBUG overrides are removed * test(e2e): cite the tokenizer and date behind the titan token-array fixture * test(e2e): let migration seed replicas finish their request-log indexes before cloning Since #43948 a serving proxy builds the two LiteLLM_SpendLogs indexes on a background thread after it reports ready. The seed fixtures stopped the replica at readiness, so every cloned legacy database lacked an index no real deployment would be missing, and the v2 baseline diff refused it. Seeds now wait until both indexes exist and are valid in the database's schema * test(passthrough): give the pass-through MockRequest an httpx URL and ASGI scope #43626 made get_request_route read request.scope during pass-through kwarg setup; the MockRequest in tests/unit/passthrough had neither a scope nor a URL object, so both stream-param tests raised before reaching the code they check. Mirrors the repair #43626 made to the tests/pass_through_unit_tests fake * test(integration): ignore foreign allow_all_keys MCP servers in the access matrix tool list test_toolset_gateway_url_serves_a_team_granted_toolset_to_a_key_without_its_own_grant (#43908) registers an allow_all_keys server on the shared gateway, and allow_all_keys servers are listed to every key by design, so a matrix case running on another xdist worker at the same time saw its tools. The matrix now drops tools of allow_all_keys servers it did not create, read from LiteLLM_MCPServerTable before and after listing, and still compares everything else exactly (cherry picked from commit 9b5562f)
) * test: repair stale and polluting tests red on scheduled main CI (#44229) Partial backport to rc/1.104.0: only the owned Redis port, Presidio stored-content check and event-loop lag logging fixes apply here. The other repairs target tests or behavior this line does not have * test(proxy): stop the proxy_server app fixture leaking LITELLM_LOG The session app fixture set LITELLM_LOG=ERROR with os.environ.setdefault and never removed it, so later tests on the same xdist worker inherited it. test_drop_params_env_var spawns a subprocess with os.environ and lost the warning it asserts on. Scope the variable to the import with a MonkeyPatch context * test(secret-detection): give the hand-built redaction request an ASGI path Since #43975 _read_request_body checks the route path via request.scope, and a scope without path raised KeyError that was swallowed into an empty body, so chat_completion failed with a missing messages parameter. Real ASGI scopes always carry path * test(integration): isolate litellm callback lists per sdk test usage-based-routing-v2 Routers register their selector in litellm.callbacks and nothing removes it, not even Router.reset(). The counter TTL and Redis service metrics tests left their selectors behind, and the next usage routing test ran their pre-call checks against its own rpm=1 deployments, raising "Deployment over defined rpm limit". An autouse fixture now gives each sdk test copies of the callback lists and restores the originals afterwards * test(integration): keep the owner-lookup fault proxy off the shared read replica The owned proxy points DATABASE_URL at a scratch database but inherited DATABASE_URL_READ_REPLICA from the replica job, so auth read the shared database and rejected the freshly created key with token_not_found_in_db. Drop the replica variable like the other scratch-database owned proxies * test(integration): request every seeded key in the team owner breakdown The aggregated team activity endpoint now caps breakdown.api_keys at the top 100 keys by default (#43398), so the 300 seeded keys came back as 100 rows. The test guarantees each key is reported with its own owner, so ask for an api_key_limit that covers all seeded keys * test(integration): give every owned Redis its own port in the redis-cache container On CircleCI every owned Redis ran on the fixed port 16379 inside the shared redis-cache container. When an earlier server still held that port, the new one failed to bind, readiness pinged the old server, the pidfile read failed and cleanup then reported "Owned Redis still serves after shutdown" Reserve an ephemeral port for the docker-exec path the same way the local binary path already does, and refuse to start when something already serves the chosen port so the failure names the real cause * test(e2e): skip the Vertex Mistral partner case the e2e project cannot reach The e2e Vertex project gets a 404 publisher model not found for vertex_ai/mistral-small-2503, so the case can only fail * test(e2e): skip the Vertex gpt-oss partner case the e2e project never serves vertex_ai/openai/gpt-oss-120b-maas has hit a 60s read timeout with no response headers on every run in the e2e Vertex project since the case was ported, and no other Vertex partner chat model passes there to switch to * test(e2e): check only stored message content for a leaked card number The Presidio spend-log check ran the card-number pattern over the whole serialized response, so a Luhn-valid usage.cost float (0.0003466000000000001) failed the streaming /v1/messages case although the stored content was <CREDIT_CARD>. The check now reads the content and text strings of the stored response, which is where a raw card would land, and still requires the placeholder there * test(e2e): assert the proxy decodes token-array embeddings for titan The port in #44120 carried over a legacy SDK-direct test that expected Bedrock to reject token ids with a 400. Through the proxy, /embeddings decodes token arrays to text for providers that cannot embed tokens, so titan answers 200. The test now sends a token array and its decoded sentence and requires the two vectors to match, which fails if the proxy stops decoding or decodes with the wrong tokenizer * test(e2e): run the Bedrock extended-thinking round trip on a model that honors enabled thinking us.anthropic.claude-sonnet-5-5 is adaptive-only, so litellm sends thinking.type=enabled with a 1024 budget as adaptive with low effort, and Bedrock returned no reasoning blocks on 5 of 5 identical Converse calls (boto3 direct agreed). us.anthropic.claude-sonnet-4-6 accepts the legacy shape verbatim and returned reasoning on 5 of 5. The non-thinking Bedrock case stays on sonnet-5-5 * test(proxy): stop unit modules forcing DEBUG logging into the event-loop lag tests Five tests/unit modules set verbose_proxy_logger to DEBUG at import, so every xdist worker that collected them logged the 2.4MB pass-through response from a worker thread, and secret redaction of that line held the GIL for ~0.8s+ inside the timed window. The lag tests now pin the LiteLLM loggers to WARNING and freeze gc while timing, and the module-level DEBUG overrides are removed * test(e2e): cite the tokenizer and date behind the titan token-array fixture * test(e2e): let migration seed replicas finish their request-log indexes before cloning Since #43948 a serving proxy builds the two LiteLLM_SpendLogs indexes on a background thread after it reports ready. The seed fixtures stopped the replica at readiness, so every cloned legacy database lacked an index no real deployment would be missing, and the v2 baseline diff refused it. Seeds now wait until both indexes exist and are valid in the database's schema * test(passthrough): give the pass-through MockRequest an httpx URL and ASGI scope * test(integration): ignore foreign allow_all_keys MCP servers in the access matrix tool list test_toolset_gateway_url_serves_a_team_granted_toolset_to_a_key_without_its_own_grant (#43908) registers an allow_all_keys server on the shared gateway, and allow_all_keys servers are listed to every key by design, so a matrix case running on another xdist worker at the same time saw its tools. The matrix now drops tools of allow_all_keys servers it did not create, read from LiteLLM_MCPServerTable before and after listing, and still compares everything else exactly (cherry picked from commit 9b5562f) * fix(proxy-extras): retry P3009 when a peer already recovered the named migration row (#44283) * fix(proxy-extras): retry P3009 when a peer already recovered the named migration row * fix(proxy-extras): pin P3009 recovery locals as Final and cover a recovered sibling row * test(proxy-extras): stop the P3009 ledger fakes shadowing the partition detector cursor * fix(proxy-extras): annotate the migration ledger reader locals as Final --------- Co-authored-by: yuneng <yuneng@berri.ai> (cherry picked from commit 8b11b68) --------- Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>


TLDR
Problem this solves:
breakdown.api_keyswas unbounded, so a user or team with thousands of keys returned every keyHow it solves it:
DailyActivityRepositoryowns every daily activity DB read behind frozen typed scopesdaily_activity_sql.pybind a caller controlledapi_key_limit(default 100, max 1000)total_api_keys,api_key_limitandentity_total_api_keystell callers when the list was cutIntentional product change: aggregated
breakdown.api_keys(and the per-model and per-entity key breakdowns) now return the topapi_key_limitkeys by spend instead of every key. Totals and per-day metrics are unchanged. Callers detect the cut fromtotal_api_keysversuslen(api_keys); the repository acceptsapi_key_limitup to 1000 and exposes paged and searched key reads, which layer 2 (#43408) wires to the routes. No environment variable controls the bound; this is the contract that replaced the reverted #41293.Files changed
litellm/repositories/daily_activity_repository.pyDailyActivityRepository:daily_rows,aggregate,entity_rollups,key_metadata,search_keys,key_page,model_top_keys,cache_leakage_keys,export_batches. The only module that runs daily activityquery_rawand Prisma reads; validates every row with Pydantic adapterslitellm/repositories/daily_activity_sql.pySqlQuery(sql, params). Bound the PTU sentinel andapi_key_limit, rank bySUM(spend::numeric)so ties are deterministic, search unionsLiteLLM_DeletedVerificationToken, exports use keyset cursorslitellm/types/repositories/daily_activity.pylitellm.constantsandlitellm.typesonlylitellm/types/repositories/__init__.pylitellm/constants.pyUSAGE_*default and max bounds andPTU_SENTINEL_API_KEYlitellm/proxy/management_endpoints/common_daily_activity.py_ProxyDailyActivityReads; keeps response shaping, daily-spend owner recovery from main (#43642) and user detail attachmentlitellm/proxy/management_endpoints/internal_user_endpoints.pyDailyActivityRepositoryand a typedDailyActivityScopefor the user daily activity routeslitellm/proxy/management_endpoints/team_endpoints.pylitellm/proxy/management_endpoints/usage_endpoints/ai_usage_chat.pydb_not_connected_errorwhen there is no Prisma clientlitellm/types/proxy/management_endpoints/common_daily_activity.pyapi_key_limit,total_api_keys,entity_total_api_keysui/litellm-dashboard/src/lib/http/schema.d.tslitellm/proxy/_lazy_openapi_snapshot.jsonflowchart LR R[user / team / AI usage route] --> P[_ProxyDailyActivityReads] P --> Repo[DailyActivityRepository] Repo --> SQL[daily_activity_sql builders] SQL --> PG[(Postgres daily spend tables)] Repo --> Meta[key metadata: active tokens, then deleted archive] P --> Own[owner recovery from daily spend, ownerless keys only]User Flow
Before: a proxy admin whose team has 106 keys gets every key back and no way to tell the list is complete or cut
results[0].breakdown.api_keyscarries all 106 keys andmetadatahas onlytotal_spendandtotal_api_requestsAfter: the same request returns the top 100 keys by spend plus the count of all keys, and totals are unchanged
results[0].breakdown.api_keyscarries the 100 highest spending keys,metadata.total_api_keysis 106 andmetadata.api_key_limitis 100,metadata.total_spendis unchangedapi_key_limitargument (default 100, max 1000,USAGE_TOP_API_KEYS_*inconstants.py); the routes in this PR use the default, layer 2 exposes it to callersLinear ticket
Resolves LIT-8897
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/unit/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Shared setup: local Postgres
circle_test, two proxies started fromtests/integration/proxy_config.yamlwithPYTHONPATHpointed at each checkout (base worktree on port 4001, this branch on port 4000), both reading the same database. Fixture inserted once withpsql: userproof-user, teamproof-team, date2026-06-15, 106 keysproof-key-000toproof-key-105with spendi + 1.0on modelgpt-5, plus one__ptu_flat_cost__sentinel row with spend 1000 on the user table only. Expected totals: user 6671.0 over 106 requests, team 5671.0. Build marker: the string"total_api_keys"inGET /openapi.jsonexists only on this branch.Before (ca05eca, the
maincommit the base worktree served)Build marker
curl -s http://127.0.0.1:4001/openapi.json | grep -o '"total_api_keys"'User aggregated
curl -s -H "Authorization: Bearer $MASTER" 'http://127.0.0.1:4001/user/daily/activity/aggregated?user_id=proof-user&start_date=2026-06-15&end_date=2026-06-15'Team aggregated
curl -s -H "Authorization: Bearer $MASTER" 'http://127.0.0.1:4001/team/daily/activity/aggregated?team_ids=proof-team&start_date=2026-06-15&end_date=2026-06-15'After (7b25632)
Build marker
curl -s http://127.0.0.1:4000/openapi.json | grep -o '"total_api_keys"'"total_api_keys"(marker present, this branch is serving)User aggregated
curl -s -H "Authorization: Bearer $MASTER" 'http://127.0.0.1:4000/user/daily/activity/aggregated?user_id=proof-user&start_date=2026-06-15&end_date=2026-06-15'proof-key-000toproof-key-005are the ones cutTeam aggregated
curl -s -H "Authorization: Bearer $MASTER" 'http://127.0.0.1:4000/team/daily/activity/aggregated?team_ids=proof-team&start_date=2026-06-15&end_date=2026-06-15'Unassigned entity rollup (NULL and empty team ids)
Fixture: three team rows on
2026-06-17,team_idNULL with keyqa-null-keyspend 3,team_id''with keyqa-empty-keyspend 7, and a NULL-team__ptu_flat_cost__row with spend 13. Both proxies read the same rows.curl -s -H "Authorization: Bearer $MASTER" 'http://127.0.0.1:4001/team/daily/activity/aggregated?start_date=2026-06-17&end_date=2026-06-17'(base)metadata.total_spend 23.0,entities.Unassigned.metrics.spend 16.0,entity_total_api_keys ABSENT(the NULL and empty rows arrive as two SQL groups and the second overwrites the first underUnassigned)curl -s -H "Authorization: Bearer $MASTER" 'http://127.0.0.1:4000/team/daily/activity/aggregated?start_date=2026-06-17&end_date=2026-06-17'(this branch)metadata.total_spend 23.0,metadata.total_api_keys 2,entities.Unassigned.metrics.spend 23.0,entity_total_api_keys {'Unassigned': 2},Unassigned.api_key_breakdownkeysqa-empty-key,qa-null-key, sentinel excludedbuild_entity_rollup_sqlnow groups byCOALESCE("<entity_id_field>", '')so NULL and empty entity ids are one row before Python maps''toUnassigned; the pre-existing spend overwrite onmainand the undercount in the newentity_total_api_keysfield are both closed by the same change.Test evidence at 7b25632
Local runs (not a substitute for the curls above):
uv run pytest tests/unit/repositories tests/unit/proxy/management_endpoints/test_common_daily_activity.py tests/unit/proxy/management_endpoints/test_internal_user_endpoints.py tests/unit/proxy/management_endpoints/test_team_endpoints.py tests/unit/proxy/test_lazy_openapi_snapshot.py tests/unit/proxy/proxy_server/test_lifecycle.py: 1003 passedpython tests/integration/run.py accounting tests/integration/spend/test_daily_activity_repository.py tests/integration/spend/test_daily_activity_aggregated_breakdowns.pyagainst real Postgres: 16 passedmake lint,scripts/ruff_strict_gate.py --base origin/main,scripts/type_discipline_gate.py --base origin/main,tests/code_coverage_tests/check_unbounded_in_lists.py: all pass, 0 new budget entriescpand verified withcmp -s:USAGE_TOP_API_KEYS_MAXinstead of the caller'sapi_key_limitin the aggregate SQL:test_daily_activity_sql.pyand the aggregated breakdown integration test failLiteLLM_DeletedVerificationTokenunion frombuild_key_search_sql: unit and integration deleted-alias search tests failSUM(spend::numeric)toSUM(spend)in search and model top keys: 2 unit tests failattach_user_details(combined)): 2 owner recovery unit tests failCOALESCE(..., '')grouping from either the entity rollup select or theentity_api_keysCTE inbuild_entity_rollup_sql:test_team_entity_rollups_merge_null_and_empty_entity_idsfailsTaxonomy audit
Confirmed and fixed in this diff:
api_keys, three new metadata fields) declared above as an intentional product change; totals, status codes and defaults for every other field unchangedORDER BY SUM(spend::numeric) DESC, api_keyeverywhere a key list is cut (aggregate, entity rollups, key page, search, model top keys); export pagination is keyset on the grouping columns, neverOFFSETentity_ids=None(no filter) is distinct fromentity_ids=()(match nothing, emitsFALSE);api_keyslikewise;0.0spend rows are keptCOALESCE(vt.*, dvt.*),dvt ON vt.token IS NULL); owner recovery only fills keys that have no owner after the token lookupdb_not_connected_error, not an empty 200MappingProxyType({})as the empty default)tests/moved with its code intotests/unit/repositories/test_daily_activity_sql.pyor was replaced by a stricter one (entity map now also asserts theUnassignedbucket andentity_total_api_keys); no relaxed assertionsOptional, no new baredict/Any, sentinel and bounds inconstants.py,schema.d.tsand the OpenAPI snapshot regenerated, no new source comments beyond docstringspy/mixed-returns:_daily_rows_tableand_export_groupingare explicitifchains ending inassert_neverplusraiseJudged not applicable:
Type
🧹 Refactoring
Caveats (if any)
Medium
breakdown.api_keysis cut atapi_key_limit(default 100, max 1000); dashboard screens that rank, search or export the key list client side see the top 100 until layer 2 (feat(proxy): bounded daily activity routes (aggregated, search, model_top_keys, export, cache_leakage_keys) for all usage entities #43408) and layer 3 (feat(ui): usage pages consume bounded daily activity routes instead of storing all keys client-side #43409) move those to the server-side key page and search reads this repository already exposesILIKEover active and deleted tokens and cannot use an index, same as the pre-existing active-token branchLow
schema-migrationis red on this head fortest_db_push_timeout_hint_names_the_per_command_budget(Lens-safety probe refuses a connection to port 9); the same test fails the same way on recentmainruns and does not touch this difftests/integration/spend/test_daily_activity_aggregated_breakdowns.pyonmainasserted that all 105 fixture keys come back; it now asserts the top 100 by spend plustotal_api_keys == 105, which is the product change above, and adds the one-key and limit-equals-count casestest_internal_user_endpoints.pyandtest_team_endpoints.pyno longer assert mockcall_kwargson the old inline helper; they assert the typedDailyActivityScopethe route builds (table, entity ids, timezone offset)tests/unit/proxy/management_endpoints/test_common_daily_activity.pywere SQL-shape checks on the old inline builder; their replacements live intests/unit/repositories/test_daily_activity_sql.pycodecov/patchreports 76.55% of the diff hit against an 80.76% target; it is not a required check, the uncovered lines areassert_neverfallbacks and typed-row validatorsPostgres Tests / schema-migrationis red on this PR and on the last fivemaincommits (test_db_push_timeout_hint_names_the_per_command_budget, Lens data safety check refusing port 9); this PR does not touchtests/proxy_migration_testsFinal Attestation
Link to Devin session: https://app.devin.ai/sessions/f9176ba9648649df8b51a8375fc7aba9
Open in Devin Desktop: https://app.devin.ai/desktop/session/f9176ba9648649df8b51a8375fc7aba9?variant=devin
Requested by: @yassin-berriai