Skip to content

refactor(repositories): daily activity repository with centralized bounded usage queries - #43398

Merged
yassin-berriai merged 1 commit into
mainfrom
litellm_usage_repository
Oct 1, 2026
Merged

yassin-berriai merged 1 commit into
mainfrom
litellm_usage_repository

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Daily activity routes built SQL and read Prisma inline, duplicated across user, team and AI-usage paths
  • Aggregated breakdown.api_keys was unbounded, so a user or team with thousands of keys returned every key
  • Deleted keys lost their alias and owner in exports and search, and ranking ties were order dependent

How it solves it:

  • New DailyActivityRepository owns every daily activity DB read behind frozen typed scopes
  • Pure SQL builders in daily_activity_sql.py bind a caller controlled api_key_limit (default 100, max 1000)
  • Totals still cover every row; total_api_keys, api_key_limit and entity_total_api_keys tell callers when the list was cut
  • Key metadata and search read active tokens first, then the deleted-token archive

Intentional product change: aggregated breakdown.api_keys (and the per-model and per-entity key breakdowns) now return the top api_key_limit keys by spend instead of every key. Totals and per-day metrics are unchanged. Callers detect the cut from total_api_keys versus len(api_keys); the repository accepts api_key_limit up to 1000 and exposes paged and searched key reads, which layer 2 (#43408) wires to the routes. No environment variable controls the bound; this is the contract that replaced the reverted #41293.

Files changed

File What changed and why
litellm/repositories/daily_activity_repository.py New DailyActivityRepository: daily_rows, aggregate, entity_rollups, key_metadata, search_keys, key_page, model_top_keys, cache_leakage_keys, export_batches. The only module that runs daily activity query_raw and Prisma reads; validates every row with Pydantic adapters
litellm/repositories/daily_activity_sql.py Pure SQL builders returning SqlQuery(sql, params). Bound the PTU sentinel and api_key_limit, rank by SUM(spend::numeric) so ties are deterministic, search unions LiteLLM_DeletedVerificationToken, exports use keyset cursors
litellm/types/repositories/daily_activity.py Frozen, slotted scope, row, result and metadata types shared by the repository and its callers. Imports stdlib, litellm.constants and litellm.types only
litellm/types/repositories/__init__.py Package marker
litellm/constants.py USAGE_* default and max bounds and PTU_SENTINEL_API_KEY
litellm/proxy/management_endpoints/common_daily_activity.py Routes delegate reads to the repository through _ProxyDailyActivityReads; keeps response shaping, daily-spend owner recovery from main (#43642) and user detail attachment
litellm/proxy/management_endpoints/internal_user_endpoints.py Builds DailyActivityRepository and a typed DailyActivityScope for the user daily activity routes
litellm/proxy/management_endpoints/team_endpoints.py Same for the team daily activity routes
litellm/proxy/management_endpoints/usage_endpoints/ai_usage_chat.py Same for AI usage; returns 500 db_not_connected_error when there is no Prisma client
litellm/types/proxy/management_endpoints/common_daily_activity.py Response metadata gains api_key_limit, total_api_keys, entity_total_api_keys
ui/litellm-dashboard/src/lib/http/schema.d.ts Generated types for the three new metadata fields
litellm/proxy/_lazy_openapi_snapshot.json OpenAPI snapshot for the new fields
flowchart LR
    R[user / team / AI usage route] --> P[_ProxyDailyActivityReads]
    P --> Repo[DailyActivityRepository]
    Repo --> SQL[daily_activity_sql builders]
    SQL --> PG[(Postgres daily spend tables)]
    Repo --> Meta[key metadata: active tokens, then deleted archive]
    P --> Own[owner recovery from daily spend, ownerless keys only]
Loading

User Flow

Before: a proxy admin whose team has 106 keys gets every key back and no way to tell the list is complete or cut

  1. They send GET https://litellm-domain/team/daily/activity/aggregated?team_ids=proof-team&start_date=2026-06-15&end_date=2026-06-15
  2. results[0].breakdown.api_keys carries all 106 keys and metadata has only total_spend and total_api_requests
  3. Searching usage by the alias of a key that was deleted last month returns no match

After: the same request returns the top 100 keys by spend plus the count of all keys, and totals are unchanged

  1. They send the same GET https://litellm-domain/team/daily/activity/aggregated?team_ids=proof-team&start_date=2026-06-15&end_date=2026-06-15
  2. results[0].breakdown.api_keys carries the 100 highest spending keys, metadata.total_api_keys is 106 and metadata.api_key_limit is 100, metadata.total_spend is unchanged
  3. The bound is the server-side api_key_limit argument (default 100, max 1000, USAGE_TOP_API_KEYS_* in constants.py); the routes in this PR use the default, layer 2 exposes it to callers
  4. Searching by a deleted key's alias, user id or email finds its historical spend

Linear ticket

Resolves LIT-8897

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/unit/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Shared setup: local Postgres circle_test, two proxies started from tests/integration/proxy_config.yaml with PYTHONPATH pointed at each checkout (base worktree on port 4001, this branch on port 4000), both reading the same database. Fixture inserted once with psql: user proof-user, team proof-team, date 2026-06-15, 106 keys proof-key-000 to proof-key-105 with spend i + 1.0 on model gpt-5, plus one __ptu_flat_cost__ sentinel row with spend 1000 on the user table only. Expected totals: user 6671.0 over 106 requests, team 5671.0. Build marker: the string "total_api_keys" in GET /openapi.json exists only on this branch.

Before (ca05eca, the main commit the base worktree served)

Build marker

  1. curl -s http://127.0.0.1:4001/openapi.json | grep -o '"total_api_keys"'
  2. Observed: no output (marker ABSENT, base code is serving)

User aggregated

  1. curl -s -H "Authorization: Bearer $MASTER" 'http://127.0.0.1:4001/user/daily/activity/aggregated?user_id=proof-user&start_date=2026-06-15&end_date=2026-06-15'
  2. Observed:
metadata.total_spend 6671.0 total_api_requests 106
metadata.total_api_keys ABSENT metadata.api_key_limit ABSENT
len(results[0].breakdown.api_keys) 106 min key proof-key-000 max key proof-key-105
__ptu_flat_cost__ in api_keys False
results[0].metrics.spend 6671.0
models.gpt-5.metrics.spend 6671.0 len(models.gpt-5.api_key_breakdown) 106

Team aggregated

  1. curl -s -H "Authorization: Bearer $MASTER" 'http://127.0.0.1:4001/team/daily/activity/aggregated?team_ids=proof-team&start_date=2026-06-15&end_date=2026-06-15'
  2. Observed:
metadata.total_spend 5671.0 metadata.total_api_keys ABSENT api_key_limit ABSENT
len(api_keys) 106 entity_total_api_keys ABSENT
entities {'proof-team': 106}

After (7b25632)

Build marker

  1. curl -s http://127.0.0.1:4000/openapi.json | grep -o '"total_api_keys"'
  2. Observed: "total_api_keys" (marker present, this branch is serving)

User aggregated

  1. curl -s -H "Authorization: Bearer $MASTER" 'http://127.0.0.1:4000/user/daily/activity/aggregated?user_id=proof-user&start_date=2026-06-15&end_date=2026-06-15'
  2. Observed:
metadata.total_spend 6671.0 total_api_requests 106
metadata.total_api_keys 106 metadata.api_key_limit 100
len(results[0].breakdown.api_keys) 100 min key proof-key-006 max key proof-key-105
__ptu_flat_cost__ in api_keys False
results[0].metrics.spend 6671.0
models.gpt-5.metrics.spend 6671.0 len(models.gpt-5.api_key_breakdown) 100
  1. Totals still include every key and the PTU sentinel row (6671.0); the six lowest spending keys proof-key-000 to proof-key-005 are the ones cut

Team aggregated

  1. curl -s -H "Authorization: Bearer $MASTER" 'http://127.0.0.1:4000/team/daily/activity/aggregated?team_ids=proof-team&start_date=2026-06-15&end_date=2026-06-15'
  2. Observed:
metadata.total_spend 5671.0 metadata.total_api_keys 106 api_key_limit 100
len(api_keys) 100 entity_total_api_keys {'proof-team': 106}
entities {'proof-team': 100}

Unassigned entity rollup (NULL and empty team ids)

Fixture: three team rows on 2026-06-17, team_id NULL with key qa-null-key spend 3, team_id '' with key qa-empty-key spend 7, and a NULL-team __ptu_flat_cost__ row with spend 13. Both proxies read the same rows.

  1. curl -s -H "Authorization: Bearer $MASTER" 'http://127.0.0.1:4001/team/daily/activity/aggregated?start_date=2026-06-17&end_date=2026-06-17' (base)
  2. Observed: metadata.total_spend 23.0, entities.Unassigned.metrics.spend 16.0, entity_total_api_keys ABSENT (the NULL and empty rows arrive as two SQL groups and the second overwrites the first under Unassigned)
  3. curl -s -H "Authorization: Bearer $MASTER" 'http://127.0.0.1:4000/team/daily/activity/aggregated?start_date=2026-06-17&end_date=2026-06-17' (this branch)
  4. Observed: metadata.total_spend 23.0, metadata.total_api_keys 2, entities.Unassigned.metrics.spend 23.0, entity_total_api_keys {'Unassigned': 2}, Unassigned.api_key_breakdown keys qa-empty-key, qa-null-key, sentinel excluded

build_entity_rollup_sql now groups by COALESCE("<entity_id_field>", '') so NULL and empty entity ids are one row before Python maps '' to Unassigned; the pre-existing spend overwrite on main and the undercount in the new entity_total_api_keys field are both closed by the same change.

Test evidence at 7b25632

Local runs (not a substitute for the curls above):

  1. uv run pytest tests/unit/repositories tests/unit/proxy/management_endpoints/test_common_daily_activity.py tests/unit/proxy/management_endpoints/test_internal_user_endpoints.py tests/unit/proxy/management_endpoints/test_team_endpoints.py tests/unit/proxy/test_lazy_openapi_snapshot.py tests/unit/proxy/proxy_server/test_lifecycle.py: 1003 passed
  2. python tests/integration/run.py accounting tests/integration/spend/test_daily_activity_repository.py tests/integration/spend/test_daily_activity_aggregated_breakdowns.py against real Postgres: 16 passed
  3. make lint, scripts/ruff_strict_gate.py --base origin/main, scripts/type_discipline_gate.py --base origin/main, tests/code_coverage_tests/check_unbounded_in_lists.py: all pass, 0 new budget entries
  4. Mutation checks, each restored with cp and verified with cmp -s:
    • bind USAGE_TOP_API_KEYS_MAX instead of the caller's api_key_limit in the aggregate SQL: test_daily_activity_sql.py and the aggregated breakdown integration test fail
    • drop the LiteLLM_DeletedVerificationToken union from build_key_search_sql: unit and integration deleted-alias search tests fail
    • revert SUM(spend::numeric) to SUM(spend) in search and model top keys: 2 unit tests fail
    • run owner recovery before merging rollups (attach_user_details(combined)): 2 owner recovery unit tests fail
    • drop the COALESCE(..., '') grouping from either the entity rollup select or the entity_api_keys CTE in build_entity_rollup_sql: test_team_entity_rollups_merge_null_and_empty_entity_ids fails

Taxonomy audit

Confirmed and fixed in this diff:

  • C1-C7: response shape change (bounded api_keys, three new metadata fields) declared above as an intentional product change; totals, status codes and defaults for every other field unchanged
  • S: ranking is ORDER BY SUM(spend::numeric) DESC, api_key everywhere a key list is cut (aggregate, entity rollups, key page, search, model top keys); export pagination is keyset on the grouping columns, never OFFSET
  • E5, X1: entity_ids=None (no filter) is distinct from entity_ids=() (match nothing, emits FALSE); api_keys likewise; 0.0 spend rows are kept
  • Y: active token metadata wins over the deleted archive (COALESCE(vt.*, dvt.*), dvt ON vt.token IS NULL); owner recovery only fills keys that have no owner after the token lookup
  • Z1, D4: owner recovery returns an empty mapping on any DB error and never widens a key's attribution; a missing Prisma client is a 500 db_not_connected_error, not an empty 200
  • W1, W4: scopes, rows and results are frozen slotted dataclasses or tuples; caller dicts are never mutated (MappingProxyType({}) as the empty default)
  • T: every removed assert in tests/ moved with its code into tests/unit/repositories/test_daily_activity_sql.py or was replaced by a stricter one (entity map now also asserts the Unassigned bucket and entity_total_api_keys); no relaxed assertions
  • H: no Optional, no new bare dict/Any, sentinel and bounds in constants.py, schema.d.ts and the OpenAPI snapshot regenerated, no new source comments beyond docstrings
  • CodeQL py/mixed-returns: _daily_rows_table and _export_grouping are explicit if chains ending in assert_never plus raise

Judged not applicable:

  • F3, O4: no provider, streaming or guardrail surface; the six daily activity entity families all go through the one repository
  • V1: no outbound API key leaves the proxy
  • A1, A2, B1: reads only; no cache writes or read-modify-write columns
  • N, M: no secrets or OAuth

Type

🧹 Refactoring

Caveats (if any)

Medium

Low

  • schema-migration is red on this head for test_db_push_timeout_hint_names_the_per_command_budget (Lens-safety probe refuses a connection to port 9); the same test fails the same way on recent main runs and does not touch this diff
  • tests/integration/spend/test_daily_activity_aggregated_breakdowns.py on main asserted that all 105 fixture keys come back; it now asserts the top 100 by spend plus total_api_keys == 105, which is the product change above, and adds the one-key and limit-equals-count cases
  • Route tests in test_internal_user_endpoints.py and test_team_endpoints.py no longer assert mock call_kwargs on the old inline helper; they assert the typed DailyActivityScope the route builds (table, entity ids, timezone offset)
  • Removed asserts in tests/unit/proxy/management_endpoints/test_common_daily_activity.py were SQL-shape checks on the old inline builder; their replacements live in tests/unit/repositories/test_daily_activity_sql.py
  • codecov/patch reports 76.55% of the diff hit against an 80.76% target; it is not a required check, the uncovered lines are assert_never fallbacks and typed-row validators
  • Postgres Tests / schema-migration is red on this PR and on the last five main commits (test_db_push_timeout_hint_names_the_per_command_budget, Lens data safety check refusing port 9); this PR does not touch tests/proxy_migration_tests

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Link to Devin session: https://app.devin.ai/sessions/f9176ba9648649df8b51a8375fc7aba9
Open in Devin Desktop: https://app.devin.ai/desktop/session/f9176ba9648649df8b51a8375fc7aba9?variant=devin
Requested by: @yassin-berriai

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@codecov

codecov Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

@greptile-apps

greptile-apps Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[High risk] Refactors daily activity queries into a centralized repository layer.

This PR appears safe to merge; no new actionable issue or outstanding previous finding was identified.

Summary

The PR moves daily activity reads into a typed repository and SQL builders. It bounds per-key aggregate lists while retaining full usage totals, and adds metadata for detecting truncated lists.

  • The latest change aligns NULL and empty entity IDs in the “Unassigned” rollup.

Reviews (10) · Last reviewed commit: "refactor(repositories): daily activity r..."

Comment thread litellm/repositories/daily_activity_sql.py Outdated
Comment thread litellm/repositories/daily_activity_repository.py Outdated
Comment thread litellm/repositories/daily_activity_sql.py Outdated
@devin-ai-integration devin-ai-integration Bot changed the title refactor(repositories): DailyActivityRepository with centralized bounded usage queries refactor(repositories): daily activity repository with centralized bounded usage queries Sep 27, 2026
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

1 similar comment
@yassin-berriai

Copy link
Copy Markdown
Contributor

@greptileai

@devin-ai-integration
devin-ai-integration Bot changed the base branch from litellm_revert_42827 to litellm_revert_41293 September 28, 2026 19:38
@devin-ai-integration
devin-ai-integration Bot force-pushed the litellm_usage_repository branch 3 times, most recently from f834012 to 1165fe8 Compare September 28, 2026 21:06
@devin-ai-integration
devin-ai-integration Bot force-pushed the litellm_revert_41293 branch 2 times, most recently from 3c4cc18 to 17c22bc Compare September 28, 2026 21:27
@devin-ai-integration
devin-ai-integration Bot force-pushed the litellm_usage_repository branch from 1165fe8 to d032e61 Compare September 28, 2026 21:38
Base automatically changed from litellm_revert_41293 to litellm_revert_41324 September 28, 2026 21:39
Base automatically changed from litellm_revert_41324 to litellm_revert_42827 September 28, 2026 21:39
@devin-ai-integration
devin-ai-integration Bot force-pushed the litellm_usage_repository branch from 21bf013 to 4df79af Compare October 1, 2026 19:31
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai review

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/repositories/daily_activity_sql.py Outdated
Comment thread litellm/repositories/daily_activity_repository.py Fixed

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/repositories/daily_activity_sql.py Outdated
@devin-ai-integration
devin-ai-integration Bot force-pushed the litellm_usage_repository branch from 4df79af to a37bb58 Compare October 1, 2026 19:48
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai re review head a37bb58 (rebased on main, search and model top keys now rank by SUM(spend::numeric))

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@devin-ai-integration
devin-ai-integration Bot force-pushed the litellm_usage_repository branch from a37bb58 to 0e475ac Compare October 1, 2026 20:27
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai re review head 0e475ac (rebased on main; CodeQL py/mixed-returns fixed: _daily_rows_table and _export_grouping are explicit if chains ending in assert_never)

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

…unded usage queries

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration
devin-ai-integration Bot force-pushed the litellm_usage_repository branch from 0e475ac to 7b25632 Compare October 1, 2026 21:33
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai review

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

Live retest on head 7b25632 (this branch on port 4000, main base worktree on 4001, same Postgres fixture).

Unassigned entity rollup, team rows on 2026-06-17 with team_id NULL (spend 3), team_id '' (spend 7) and a NULL-team PTU sentinel (spend 13):

base head
metadata.total_spend 23.0 23.0
entities.Unassigned.metrics.spend 16.0 23.0
metadata.entity_total_api_keys absent {"Unassigned": 2}
Unassigned.api_key_breakdown qa-empty-key, qa-null-key qa-empty-key, qa-null-key, sentinel excluded

Regression checks on the 2026-06-15 fixture (106 keys plus PTU): head returns 100 of 106 keys with total_api_keys 106 and api_key_limit 100, totals 6671 (user) and 5671 (team) match base, PTU stays in totals and out of every key breakdown. Deleted key keeps its historical spend, alias and owner with key_exists false. Empty user and team ranges return 200 with zero spend. Caller-controlled api_key_limit is not exposed on routes here by design (layer 2, #43408).

Admin UI usage page on the same head, proof-user on 2026-06-15:

Usage page total and daily chart

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 7b25632. Configure here.

FROM scoped
{" ".join(joins)}
WHERE TRUE{cursor_clause}
GROUP BY {", ".join(grouping_keys)}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Model export SQL grouping mismatch

Medium Severity

DAILY_WITH_MODELS export selects NULLIF(COALESCE(scoped.model, ''), '') while grouping only by COALESCE(scoped.model, ''). Postgres requires non-aggregated SELECT expressions to match GROUP BY exactly, so this query is rejected at runtime and export_rows cannot produce a model breakdown.

Additional Locations (1)
Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit 7b25632. Configure here.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not a bug. Postgres accepts a select-list expression whose only ungrouped column reference appears inside a subexpression that textually matches a GROUP BY expression, so NULLIF(COALESCE(scoped.model, ''), '') is valid when grouping by COALESCE(scoped.model, '').

Verified on PostgreSQL 14.24:

SELECT NULLIF(COALESCE(m, ''), '') AS model, COUNT(*)
FROM (VALUES ('gpt-4o'), (NULL), ('')) t(m)
GROUP BY COALESCE(m, '') ORDER BY 1;
 model  | count
--------+-------
 gpt-4o |     1
        |     2

and the generated query itself: build_export_sql(scope, export_type=ExportType.DAILY_WITH_MODELS, after=None, batch_size=1000) executed against the live LiteLLM_DailyTeamSpend fixture on this head returns rows (2 for the June 2026 team scope), same as the DAILY, DAILY_WITH_KEYS and DAILY_WITH_USERS variants.

@yassin-berriai
yassin-berriai merged commit aa601ce into main Oct 1, 2026
106 of 108 checks passed
@yassin-berriai
yassin-berriai deleted the litellm_usage_repository branch October 1, 2026 23:02
yuneng-berri added a commit that referenced this pull request Oct 2, 2026
* test(proxy): stop the proxy_server app fixture leaking LITELLM_LOG

The session app fixture set LITELLM_LOG=ERROR with os.environ.setdefault and never removed it, so later tests on the same xdist worker inherited it. test_drop_params_env_var spawns a subprocess with os.environ and lost the warning it asserts on. Scope the variable to the import with a MonkeyPatch context

* test(secret-detection): give the hand-built redaction request an ASGI path

Since #43975 _read_request_body checks the route path via request.scope, and a scope without path raised KeyError that was swallowed into an empty body, so chat_completion failed with a missing messages parameter. Real ASGI scopes always carry path

* test(integration): isolate litellm callback lists per sdk test

usage-based-routing-v2 Routers register their selector in litellm.callbacks and nothing removes it, not even Router.reset(). The counter TTL and Redis service metrics tests left their selectors behind, and the next usage routing test ran their pre-call checks against its own rpm=1 deployments, raising "Deployment over defined rpm limit". An autouse fixture now gives each sdk test copies of the callback lists and restores the originals afterwards

* test(integration): keep the owner-lookup fault proxy off the shared read replica

The owned proxy points DATABASE_URL at a scratch database but inherited
DATABASE_URL_READ_REPLICA from the replica job, so auth read the shared
database and rejected the freshly created key with token_not_found_in_db.
Drop the replica variable like the other scratch-database owned proxies

* test(integration): request every seeded key in the team owner breakdown

The aggregated team activity endpoint now caps breakdown.api_keys at the top
100 keys by default (#43398), so the 300 seeded keys came back as 100 rows.
The test guarantees each key is reported with its own owner, so ask for an
api_key_limit that covers all seeded keys

* test(integration): give every owned Redis its own port in the redis-cache container

On CircleCI every owned Redis ran on the fixed port 16379 inside the shared
redis-cache container. When an earlier server still held that port, the new
one failed to bind, readiness pinged the old server, the pidfile read failed
and cleanup then reported "Owned Redis still serves after shutdown"

Reserve an ephemeral port for the docker-exec path the same way the local
binary path already does, and refuse to start when something already serves
the chosen port so the failure names the real cause

* test(e2e): skip the Vertex Mistral partner case the e2e project cannot reach

The e2e Vertex project gets a 404 publisher model not found for vertex_ai/mistral-small-2503, so the case can only fail

* test(e2e): skip the Vertex gpt-oss partner case the e2e project never serves

vertex_ai/openai/gpt-oss-120b-maas has hit a 60s read timeout with no response headers on every run in the e2e Vertex project since the case was ported, and no other Vertex partner chat model passes there to switch to

* test(e2e): check only stored message content for a leaked card number

The Presidio spend-log check ran the card-number pattern over the whole serialized response, so a Luhn-valid usage.cost float (0.0003466000000000001) failed the streaming /v1/messages case although the stored content was <CREDIT_CARD>. The check now reads the content and text strings of the stored response, which is where a raw card would land, and still requires the placeholder there

* test(e2e): assert the proxy decodes token-array embeddings for titan

The port in #44120 carried over a legacy SDK-direct test that expected Bedrock to reject token ids with a 400. Through the proxy, /embeddings decodes token arrays to text for providers that cannot embed tokens, so titan answers 200. The test now sends a token array and its decoded sentence and requires the two vectors to match, which fails if the proxy stops decoding or decodes with the wrong tokenizer

* test(e2e): run the Bedrock extended-thinking round trip on a model that honors enabled thinking

us.anthropic.claude-sonnet-5-5 is adaptive-only, so litellm sends thinking.type=enabled with a 1024 budget as adaptive with low effort, and Bedrock returned no reasoning blocks on 5 of 5 identical Converse calls (boto3 direct agreed). us.anthropic.claude-sonnet-4-6 accepts the legacy shape verbatim and returned reasoning on 5 of 5. The non-thinking Bedrock case stays on sonnet-5-5

* test(proxy): stop unit modules forcing DEBUG logging into the event-loop lag tests

Five tests/unit modules set verbose_proxy_logger to DEBUG at import, so every xdist worker that collected them logged the 2.4MB pass-through response from a worker thread, and secret redaction of that line held the GIL for ~0.8s+ inside the timed window. The lag tests now pin the LiteLLM loggers to WARNING and freeze gc while timing, and the module-level DEBUG overrides are removed

* test(e2e): cite the tokenizer and date behind the titan token-array fixture

* test(e2e): let migration seed replicas finish their request-log indexes before cloning

Since #43948 a serving proxy builds the two LiteLLM_SpendLogs indexes on a background thread after it reports ready. The seed fixtures stopped the replica at readiness, so every cloned legacy database lacked an index no real deployment would be missing, and the v2 baseline diff refused it. Seeds now wait until both indexes exist and are valid in the database's schema

* test(passthrough): give the pass-through MockRequest an httpx URL and ASGI scope

#43626 made get_request_route read request.scope during pass-through kwarg setup; the MockRequest in tests/unit/passthrough had neither a scope nor a URL object, so both stream-param tests raised before reaching the code they check. Mirrors the repair #43626 made to the tests/pass_through_unit_tests fake

* test(integration): ignore foreign allow_all_keys MCP servers in the access matrix tool list

test_toolset_gateway_url_serves_a_team_granted_toolset_to_a_key_without_its_own_grant (#43908) registers an allow_all_keys server on the shared gateway, and allow_all_keys servers are listed to every key by design, so a matrix case running on another xdist worker at the same time saw its tools. The matrix now drops tools of allow_all_keys servers it did not create, read from LiteLLM_MCPServerTable before and after listing, and still compares everything else exactly
yuneng-berri added a commit that referenced this pull request Oct 2, 2026
…) (#44254)

* test(proxy): stop the proxy_server app fixture leaking LITELLM_LOG

The session app fixture set LITELLM_LOG=ERROR with os.environ.setdefault and never removed it, so later tests on the same xdist worker inherited it. test_drop_params_env_var spawns a subprocess with os.environ and lost the warning it asserts on. Scope the variable to the import with a MonkeyPatch context

* test(secret-detection): give the hand-built redaction request an ASGI path

Since #43975 _read_request_body checks the route path via request.scope, and a scope without path raised KeyError that was swallowed into an empty body, so chat_completion failed with a missing messages parameter. Real ASGI scopes always carry path

* test(integration): isolate litellm callback lists per sdk test

usage-based-routing-v2 Routers register their selector in litellm.callbacks and nothing removes it, not even Router.reset(). The counter TTL and Redis service metrics tests left their selectors behind, and the next usage routing test ran their pre-call checks against its own rpm=1 deployments, raising "Deployment over defined rpm limit". An autouse fixture now gives each sdk test copies of the callback lists and restores the originals afterwards

* test(integration): keep the owner-lookup fault proxy off the shared read replica

The owned proxy points DATABASE_URL at a scratch database but inherited
DATABASE_URL_READ_REPLICA from the replica job, so auth read the shared
database and rejected the freshly created key with token_not_found_in_db.
Drop the replica variable like the other scratch-database owned proxies

* test(integration): request every seeded key in the team owner breakdown

The aggregated team activity endpoint now caps breakdown.api_keys at the top
100 keys by default (#43398), so the 300 seeded keys came back as 100 rows.
The test guarantees each key is reported with its own owner, so ask for an
api_key_limit that covers all seeded keys

* test(integration): give every owned Redis its own port in the redis-cache container

On CircleCI every owned Redis ran on the fixed port 16379 inside the shared
redis-cache container. When an earlier server still held that port, the new
one failed to bind, readiness pinged the old server, the pidfile read failed
and cleanup then reported "Owned Redis still serves after shutdown"

Reserve an ephemeral port for the docker-exec path the same way the local
binary path already does, and refuse to start when something already serves
the chosen port so the failure names the real cause

* test(e2e): skip the Vertex Mistral partner case the e2e project cannot reach

The e2e Vertex project gets a 404 publisher model not found for vertex_ai/mistral-small-2503, so the case can only fail

* test(e2e): skip the Vertex gpt-oss partner case the e2e project never serves

vertex_ai/openai/gpt-oss-120b-maas has hit a 60s read timeout with no response headers on every run in the e2e Vertex project since the case was ported, and no other Vertex partner chat model passes there to switch to

* test(e2e): check only stored message content for a leaked card number

The Presidio spend-log check ran the card-number pattern over the whole serialized response, so a Luhn-valid usage.cost float (0.0003466000000000001) failed the streaming /v1/messages case although the stored content was <CREDIT_CARD>. The check now reads the content and text strings of the stored response, which is where a raw card would land, and still requires the placeholder there

* test(e2e): assert the proxy decodes token-array embeddings for titan

The port in #44120 carried over a legacy SDK-direct test that expected Bedrock to reject token ids with a 400. Through the proxy, /embeddings decodes token arrays to text for providers that cannot embed tokens, so titan answers 200. The test now sends a token array and its decoded sentence and requires the two vectors to match, which fails if the proxy stops decoding or decodes with the wrong tokenizer

* test(e2e): run the Bedrock extended-thinking round trip on a model that honors enabled thinking

us.anthropic.claude-sonnet-5-5 is adaptive-only, so litellm sends thinking.type=enabled with a 1024 budget as adaptive with low effort, and Bedrock returned no reasoning blocks on 5 of 5 identical Converse calls (boto3 direct agreed). us.anthropic.claude-sonnet-4-6 accepts the legacy shape verbatim and returned reasoning on 5 of 5. The non-thinking Bedrock case stays on sonnet-5-5

* test(proxy): stop unit modules forcing DEBUG logging into the event-loop lag tests

Five tests/unit modules set verbose_proxy_logger to DEBUG at import, so every xdist worker that collected them logged the 2.4MB pass-through response from a worker thread, and secret redaction of that line held the GIL for ~0.8s+ inside the timed window. The lag tests now pin the LiteLLM loggers to WARNING and freeze gc while timing, and the module-level DEBUG overrides are removed

* test(e2e): cite the tokenizer and date behind the titan token-array fixture

* test(e2e): let migration seed replicas finish their request-log indexes before cloning

Since #43948 a serving proxy builds the two LiteLLM_SpendLogs indexes on a background thread after it reports ready. The seed fixtures stopped the replica at readiness, so every cloned legacy database lacked an index no real deployment would be missing, and the v2 baseline diff refused it. Seeds now wait until both indexes exist and are valid in the database's schema

* test(passthrough): give the pass-through MockRequest an httpx URL and ASGI scope

#43626 made get_request_route read request.scope during pass-through kwarg setup; the MockRequest in tests/unit/passthrough had neither a scope nor a URL object, so both stream-param tests raised before reaching the code they check. Mirrors the repair #43626 made to the tests/pass_through_unit_tests fake

* test(integration): ignore foreign allow_all_keys MCP servers in the access matrix tool list

test_toolset_gateway_url_serves_a_team_granted_toolset_to_a_key_without_its_own_grant (#43908) registers an allow_all_keys server on the shared gateway, and allow_all_keys servers are listed to every key by design, so a matrix case running on another xdist worker at the same time saw its tools. The matrix now drops tools of allow_all_keys servers it did not create, read from LiteLLM_MCPServerTable before and after listing, and still compares everything else exactly

(cherry picked from commit 9b5562f)
yuneng-berri added a commit that referenced this pull request Oct 3, 2026
)

* test: repair stale and polluting tests red on scheduled main CI (#44229)

Partial backport to rc/1.104.0: only the owned Redis port, Presidio stored-content check and event-loop lag logging fixes apply here. The other repairs target tests or behavior this line does not have

* test(proxy): stop the proxy_server app fixture leaking LITELLM_LOG

The session app fixture set LITELLM_LOG=ERROR with os.environ.setdefault and never removed it, so later tests on the same xdist worker inherited it. test_drop_params_env_var spawns a subprocess with os.environ and lost the warning it asserts on. Scope the variable to the import with a MonkeyPatch context

* test(secret-detection): give the hand-built redaction request an ASGI path

Since #43975 _read_request_body checks the route path via request.scope, and a scope without path raised KeyError that was swallowed into an empty body, so chat_completion failed with a missing messages parameter. Real ASGI scopes always carry path

* test(integration): isolate litellm callback lists per sdk test

usage-based-routing-v2 Routers register their selector in litellm.callbacks and nothing removes it, not even Router.reset(). The counter TTL and Redis service metrics tests left their selectors behind, and the next usage routing test ran their pre-call checks against its own rpm=1 deployments, raising "Deployment over defined rpm limit". An autouse fixture now gives each sdk test copies of the callback lists and restores the originals afterwards

* test(integration): keep the owner-lookup fault proxy off the shared read replica

The owned proxy points DATABASE_URL at a scratch database but inherited
DATABASE_URL_READ_REPLICA from the replica job, so auth read the shared
database and rejected the freshly created key with token_not_found_in_db.
Drop the replica variable like the other scratch-database owned proxies

* test(integration): request every seeded key in the team owner breakdown

The aggregated team activity endpoint now caps breakdown.api_keys at the top
100 keys by default (#43398), so the 300 seeded keys came back as 100 rows.
The test guarantees each key is reported with its own owner, so ask for an
api_key_limit that covers all seeded keys

* test(integration): give every owned Redis its own port in the redis-cache container

On CircleCI every owned Redis ran on the fixed port 16379 inside the shared
redis-cache container. When an earlier server still held that port, the new
one failed to bind, readiness pinged the old server, the pidfile read failed
and cleanup then reported "Owned Redis still serves after shutdown"

Reserve an ephemeral port for the docker-exec path the same way the local
binary path already does, and refuse to start when something already serves
the chosen port so the failure names the real cause

* test(e2e): skip the Vertex Mistral partner case the e2e project cannot reach

The e2e Vertex project gets a 404 publisher model not found for vertex_ai/mistral-small-2503, so the case can only fail

* test(e2e): skip the Vertex gpt-oss partner case the e2e project never serves

vertex_ai/openai/gpt-oss-120b-maas has hit a 60s read timeout with no response headers on every run in the e2e Vertex project since the case was ported, and no other Vertex partner chat model passes there to switch to

* test(e2e): check only stored message content for a leaked card number

The Presidio spend-log check ran the card-number pattern over the whole serialized response, so a Luhn-valid usage.cost float (0.0003466000000000001) failed the streaming /v1/messages case although the stored content was <CREDIT_CARD>. The check now reads the content and text strings of the stored response, which is where a raw card would land, and still requires the placeholder there

* test(e2e): assert the proxy decodes token-array embeddings for titan

The port in #44120 carried over a legacy SDK-direct test that expected Bedrock to reject token ids with a 400. Through the proxy, /embeddings decodes token arrays to text for providers that cannot embed tokens, so titan answers 200. The test now sends a token array and its decoded sentence and requires the two vectors to match, which fails if the proxy stops decoding or decodes with the wrong tokenizer

* test(e2e): run the Bedrock extended-thinking round trip on a model that honors enabled thinking

us.anthropic.claude-sonnet-5-5 is adaptive-only, so litellm sends thinking.type=enabled with a 1024 budget as adaptive with low effort, and Bedrock returned no reasoning blocks on 5 of 5 identical Converse calls (boto3 direct agreed). us.anthropic.claude-sonnet-4-6 accepts the legacy shape verbatim and returned reasoning on 5 of 5. The non-thinking Bedrock case stays on sonnet-5-5

* test(proxy): stop unit modules forcing DEBUG logging into the event-loop lag tests

Five tests/unit modules set verbose_proxy_logger to DEBUG at import, so every xdist worker that collected them logged the 2.4MB pass-through response from a worker thread, and secret redaction of that line held the GIL for ~0.8s+ inside the timed window. The lag tests now pin the LiteLLM loggers to WARNING and freeze gc while timing, and the module-level DEBUG overrides are removed

* test(e2e): cite the tokenizer and date behind the titan token-array fixture

* test(e2e): let migration seed replicas finish their request-log indexes before cloning

Since #43948 a serving proxy builds the two LiteLLM_SpendLogs indexes on a background thread after it reports ready. The seed fixtures stopped the replica at readiness, so every cloned legacy database lacked an index no real deployment would be missing, and the v2 baseline diff refused it. Seeds now wait until both indexes exist and are valid in the database's schema

* test(passthrough): give the pass-through MockRequest an httpx URL and ASGI scope

* test(integration): ignore foreign allow_all_keys MCP servers in the access matrix tool list

test_toolset_gateway_url_serves_a_team_granted_toolset_to_a_key_without_its_own_grant (#43908) registers an allow_all_keys server on the shared gateway, and allow_all_keys servers are listed to every key by design, so a matrix case running on another xdist worker at the same time saw its tools. The matrix now drops tools of allow_all_keys servers it did not create, read from LiteLLM_MCPServerTable before and after listing, and still compares everything else exactly

(cherry picked from commit 9b5562f)

* fix(proxy-extras): retry P3009 when a peer already recovered the named migration row (#44283)

* fix(proxy-extras): retry P3009 when a peer already recovered the named migration row

* fix(proxy-extras): pin P3009 recovery locals as Final and cover a recovered sibling row

* test(proxy-extras): stop the P3009 ledger fakes shadowing the partition detector cursor

* fix(proxy-extras): annotate the migration ledger reader locals as Final

---------

Co-authored-by: yuneng <yuneng@berri.ai>
(cherry picked from commit 8b11b68)

---------

Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants