Skip to content

feat: improve trace ingestion and trace details - #43975

Merged
yujonglee-berri merged 25 commits into
litellm_tracing_dependenciesfrom
litellm_gateway_traces_boundary
Oct 1, 2026
Merged

yujonglee-berri merged 25 commits into
litellm_tracing_dependenciesfrom
litellm_gateway_traces_boundary

Conversation

@yujonglee-berri

@yujonglee-berri yujonglee-berri commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Stack

Managed with gh stack in stack ID 44039. Depends on #44035. Merge the receiver/access-context refactor first, then rebase this branch and retarget it to main

TLDR

Problem this solves:

  • Resource metadata was copied onto every decoded span
  • Decoding lacked combined allocation and structural bounds
  • Trace ingestion and detail handling needed stronger bounds

How it solves it:

  • Preserve resource sharing through Rust, Python, and insertion
  • Encapsulate ownership behind one Shared<T> storage alias
  • Stream insert encoding without cloning entire batches
  • Bound HTTP bodies and gzip expansion before decoding
  • Parse media types and encode protobuf errors in Rust
  • Expose paginated span errors in trace details

Intentional product change: Gzip without Content-Encoding: gzip returns 400 because the receiver follows HTTP metadata. Incomplete uploads expire after 30 seconds. Authentication errors use safe HTTP reason phrases instead of exception details

User Flow

Before: compressed exports are accepted without declaring their encoding

  1. An exporter sends gzipped OTLP JSON to POST /v1/traces
  2. It omits Content-Encoding: gzip and receives HTTP 200
  3. An unauthenticated JSON export receives an OpenAI-shaped 401 error

After: exporters receive protocol-specific validation and errors

  1. An exporter sends gzipped OTLP JSON to POST /v1/traces
  2. It omits Content-Encoding: gzip and receives HTTP 400
  3. An unauthenticated JSON export receives an OTLP-shaped 401 error

Implementation

Shared<T> keeps the reference-counted storage choice inside litellm-traces. It remains trace-owned because trace decoding and insertion are its only current use cases. Its private Storage<T> alias currently uses Arc<T>; Box<T> passed value-parity tests during development. This alias is distinct from litellm-storage-clickhouse::Storage, which owns generic ClickHouse connections and transport. Clone accounting charges additional allocation only when the selected storage actually copies

litellm-host-python owns reusable ToPythonCache and FromPythonCache helpers. They memoize successful conversions by source identity and retain source lifetimes, without depending on Shared<T> or trace types. Trace fields and the choice of values to share remain in the bridge

The PyO3 bridge preserves resource and scope identity, tenant stamping copies each resource group once, and insert conversion preserves that sharing. Rust validates insert values directly, avoiding Pydantic rebuilding every nested map. The encoder streams into the hash and gzip writers with the existing 64 MiB encoded limit. The generic ClickHouse crate sends the compressed body. Two serialization passes preserve the existing deduplication token and server receive timestamp

The 16 MiB HTTP default and decompressed-body limit remain configurable. The decoded-output cap remains because unique values can expand during conversion, even after resource sharing. Structural limits remain separate

Validation

The earlier 304 focused Python tests pass, including prefixed routes, safe error responses, missing native support, and upload deadline slot release. Python lint and formatting pass. CI exposed invalid synthetic request scopes in existing tests. Those fixtures now supply method and path, and the temporary production fallback was removed. The branch includes #44035's shutdown retry fix at 4921a1c81b. At 801953454c9, all 339 tests in the auth request-flow file pass. The five CI failures reproduced locally before adding the missing path to six fixtures. Assertions and production code are unchanged

Real ClickHouse queries reproduced different preview and diagnostic messages for a duplicated span. The revised queries agree for different start timestamps, different receive timestamps, and fully tied timestamps. Rust integration coverage now exercises all three cases. All three Rust integration cases pass against real ClickHouse. The native rebuild and make check pass, including the latest fixture changes

Live prefixed gzip ingestion returned 200, and prefixed unauthenticated ingestion returned a sanitized 401

A live unfinished upload returned 503 after 30.0 seconds, and the next complete export returned 200. Veria reports no remaining security concerns

The live HTTP checks below use PostgreSQL and ClickHouse with no mocks. The Python proxy was rerun at 801953454c9 with the rebuilt native extension. All seven live HTTP checks pass. Hosted CI and fresh bot reviews remain pending

Review fixes share prefix-aware route detection between auth and error formatting, replace exception text with HTTP reason phrases, allow protobuf errors when native support is absent, and expire body reads inside the shielded ingestion task after 30 seconds. Both diagnostic queries use the same deterministic row order. The unused Axum tracing scaffold was removed

The protobuf Status.code finding was rejected because OTLP/HTTP permits omitting that field. HTTP status governs failure and retry handling

Screenshots / Proof of Fix

Captured against real local PostgreSQL and ClickHouse. The before proxy uses port 49173 and the after proxy uses port 4000. TEAM_AUTH points to a private curl config containing the team key's Authorization header. No LLM call is involved

Create the same compressed fixture for both runs:

cat > export.json <<'JSON'
{"resourceSpans": [{"resource": {"attributes": [{"key": "litellm.team_id", "value": {"stringValue": "spoofed-team"}}]}, "scopeSpans": [{"spans": [{"traceId": "00112233445566778899aabbccddeeff", "spanId": "0011223344556677", "name": "dependency-injection-qa-protobuf", "startTimeUnixNano": "1790812800000000000", "endTimeUnixNano": "1790812800001000000"}]}]}]}
JSON
gzip -c export.json > export.json.gz

Before (4921a1c)

Undeclared gzip

  1. Send the export

    curl -sS --config "$TEAM_AUTH" -H 'Content-Type: application/json' http://127.0.0.1:49173/v1/traces --data-binary @export.json.gz -w '\nHTTP %{http_code}\n'
  2. Observed output: {}, HTTP 200

Declared gzip

  1. Send the export

    curl -sS --config "$TEAM_AUTH" -H 'Content-Type: application/json' -H 'Content-Encoding: gzip' http://127.0.0.1:49173/v1/traces --data-binary @export.json.gz -w '\nHTTP %{http_code}\n'
  2. Observed output: {}, HTTP 200

Missing API key

  1. Send the export

    curl -sS -H 'Content-Type: application/json' http://127.0.0.1:49173/v1/traces -d '{}' -w '\nHTTP %{http_code}\n'
  2. Observed output: {"error":{"message":"Authentication Error, No api key passed in.","type":"auth_error","param":"None","code":"401"}}, HTTP 401

After (8019534)

Undeclared gzip

  1. Send the export

    curl -sS --config "$TEAM_AUTH" -H 'Content-Type: application/json' http://127.0.0.1:4000/v1/traces --data-binary @export.json.gz -w '\nHTTP %{http_code}\n'
  2. Observed output: {"message": "invalid OTLP trace payload"}, HTTP 400

Declared gzip

  1. Send the export

    curl -sS --config "$TEAM_AUTH" -H 'Content-Type: application/json' -H 'Content-Encoding: gzip' http://127.0.0.1:4000/v1/traces --data-binary @export.json.gz -w '\nHTTP %{http_code}\n'
  2. Observed output: {}, HTTP 200

Missing API key

  1. Send the export

    curl -sS -H 'Content-Type: application/json' http://127.0.0.1:4000/v1/traces -d '{}' -w '\nHTTP %{http_code}\n'
  2. Observed output: {"message": "Unauthorized"}, HTTP 401

Measurements

Earlier local debug-build comparisons, three fresh processes per comparison, 8 KiB shared across 1,024 spans. These compare intermediate implementations, not the current merge base:

Concurrent requests Previous peak RSS increase Shared peak RSS increase
1 76.77 MiB 23.25 MiB
2 132.03 MiB 29.66 MiB

Two-request elapsed time was effectively unchanged at about 827 ms. These measurements demonstrate a memory improvement, not production throughput

Temporarily removing the decoded cap let a 15.68 MiB protobuf export add 636.67 MiB at two concurrent requests through string expansion. A 3.96 MiB version added 188.61 MiB. The retained cap rejects both expansion cases, while the previously rejected resource-fanout case succeeds

The memory figures above are prior local profiling observations. The temporary Python profiling harness and environment-driven timing probes have been removed from the PR. Maintained loopback coverage lives in Rust: cargo test --manifest-path litellm-rust/Cargo.toml -p litellm-traces --test insert verifies 1,024 rows sharing a 16 KiB resource through gzip and HTTP, including two concurrent requests, and rejects 64 MiB resource fanout before transport

The resource fanout microbenchmark lives in litellm-rust/crates/traces/benches/resource-fanout.rs. Run cargo bench --manifest-path litellm-rust/Cargo.toml -p litellm-traces --bench resource-fanout to compare owned Box<T> copying with production Shared<T> cloning in one run. Fixtures are generated in code, and benchmark results are not checked in. This benchmark measures cloning and dropping values, not end-to-end process RSS

Pre-Submission checklist

  • I have added meaningful tests
  • Focused Python, native, and trace-detail UI tests pass
  • My PR passes all required CI/CD checks
  • My PR's scope is as isolated as possible
  • I have received a Greptile Confidence Score of at least 4/5

Type

New Feature, Bug Fix, Refactoring

Caveats

Low

  • Ingestion measurements use debug builds on a shared machine
  • The macOS test linker warned about large unwind metadata
  • Required CI and fresh Greptile review remain pending
  • Bugbot review is unavailable on this tip
  • Unrelated auto-router UI cases failed in the previous run
  • Existing team detail authorization still needs follow-up
  • Dashboard screenshots were not refreshed for this rebase

Final Attestation

  • Focused tests cover sharing, isolation, limits, and insertion behavior

Devin Review

@yujonglee-berri yujonglee-berri changed the title refactor: separate OTLP HTTP decoding from trace codec feat: improve trace ingestion and trace details Oct 1, 2026
github-advanced-security[bot]

This comment was marked as resolved.

@codspeed

codspeed Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_gateway_traces_boundary (8019534) with litellm_tracing_dependencies (4921a1c)

Open in CodSpeed

@codecov

codecov Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 93.46405% with 20 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
litellm/proxy/tracing_endpoints.py 73.33% 8 Missing ⚠️
litellm/tracing/decode.py 96.87% 5 Missing ⚠️
litellm/tracing/receiver.py 93.61% 3 Missing ⚠️
litellm/proxy/proxy_server.py 88.88% 2 Missing ⚠️
litellm/rust_bridge/traces.py 87.50% 1 Missing ⚠️
litellm/tracing/store.py 95.65% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

@yujonglee-berri
yujonglee-berri force-pushed the litellm_gateway_traces_boundary branch from cf1a1e1 to 1ad31dc Compare October 1, 2026 17:02
@yujonglee-berri
yujonglee-berri changed the base branch from main to litellm_tracing_dependencies October 1, 2026 17:03
@yujonglee-berri
yujonglee-berri added this pull request to stack #44039 October 1, 2026 17:07
@yujonglee-berri
yujonglee-berri force-pushed the litellm_gateway_traces_boundary branch from 17675f3 to 2629d67 Compare October 1, 2026 17:38
Comment thread litellm/proxy/tracing_endpoints.py
@yujonglee-berri
yujonglee-berri force-pushed the litellm_gateway_traces_boundary branch 2 times, most recently from 11e7c96 to 61ae0c3 Compare October 1, 2026 18:25
Comment thread litellm/tracing/receiver.py Outdated
@veria-ai

veria-ai Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

PR overview

All previously flagged issues have been addressed. No open security concerns remain on this pull request.

Security review

No open security issues remain on this pull request.

Fixed/addressed: 1 · PR risk: 0/10

@yujonglee-berri

Copy link
Copy Markdown
Contributor Author

@greptile-app review Please review the current tip after conflict resolution and report actionable findings on this pull request.

@yujonglee-berri

Copy link
Copy Markdown
Contributor Author

@greptile-app review The branch tip changed after the migration test fix. Please review this latest commit and report actionable findings.

@yujonglee-berri

Copy link
Copy Markdown
Contributor Author

@greptileai Please review the latest tip. Valid findings are fixed, duplicate-span queries passed live ClickHouse checks, and all existing threads have replies

@yujonglee-berri

Copy link
Copy Markdown
Contributor Author

@greptileai Please review 484ce13. The follow-up preserves partial request scopes, and all 150 affected tests pass. Previous review fixes remain included

@yujonglee-berri

Copy link
Copy Markdown
Contributor Author

@greptileai Please review d9ed39a. Invalid test fixtures now use valid HTTP scopes, the production fallback is removed, and all 304 focused tests pass

@yujonglee-berri

Copy link
Copy Markdown
Contributor Author

@greptileai Please review 8019534. Completed auth request fixtures; all 339 file tests, seven live checks, and make check pass.

@yujonglee-berri
yujonglee-berri merged commit ec60582 into main Oct 1, 2026
101 of 103 checks passed
@yujonglee-berri
yujonglee-berri deleted the litellm_gateway_traces_boundary branch October 1, 2026 20:45
yuneng-berri added a commit that referenced this pull request Oct 2, 2026
* test(proxy): stop the proxy_server app fixture leaking LITELLM_LOG

The session app fixture set LITELLM_LOG=ERROR with os.environ.setdefault and never removed it, so later tests on the same xdist worker inherited it. test_drop_params_env_var spawns a subprocess with os.environ and lost the warning it asserts on. Scope the variable to the import with a MonkeyPatch context

* test(secret-detection): give the hand-built redaction request an ASGI path

Since #43975 _read_request_body checks the route path via request.scope, and a scope without path raised KeyError that was swallowed into an empty body, so chat_completion failed with a missing messages parameter. Real ASGI scopes always carry path

* test(integration): isolate litellm callback lists per sdk test

usage-based-routing-v2 Routers register their selector in litellm.callbacks and nothing removes it, not even Router.reset(). The counter TTL and Redis service metrics tests left their selectors behind, and the next usage routing test ran their pre-call checks against its own rpm=1 deployments, raising "Deployment over defined rpm limit". An autouse fixture now gives each sdk test copies of the callback lists and restores the originals afterwards

* test(integration): keep the owner-lookup fault proxy off the shared read replica

The owned proxy points DATABASE_URL at a scratch database but inherited
DATABASE_URL_READ_REPLICA from the replica job, so auth read the shared
database and rejected the freshly created key with token_not_found_in_db.
Drop the replica variable like the other scratch-database owned proxies

* test(integration): request every seeded key in the team owner breakdown

The aggregated team activity endpoint now caps breakdown.api_keys at the top
100 keys by default (#43398), so the 300 seeded keys came back as 100 rows.
The test guarantees each key is reported with its own owner, so ask for an
api_key_limit that covers all seeded keys

* test(integration): give every owned Redis its own port in the redis-cache container

On CircleCI every owned Redis ran on the fixed port 16379 inside the shared
redis-cache container. When an earlier server still held that port, the new
one failed to bind, readiness pinged the old server, the pidfile read failed
and cleanup then reported "Owned Redis still serves after shutdown"

Reserve an ephemeral port for the docker-exec path the same way the local
binary path already does, and refuse to start when something already serves
the chosen port so the failure names the real cause

* test(e2e): skip the Vertex Mistral partner case the e2e project cannot reach

The e2e Vertex project gets a 404 publisher model not found for vertex_ai/mistral-small-2503, so the case can only fail

* test(e2e): skip the Vertex gpt-oss partner case the e2e project never serves

vertex_ai/openai/gpt-oss-120b-maas has hit a 60s read timeout with no response headers on every run in the e2e Vertex project since the case was ported, and no other Vertex partner chat model passes there to switch to

* test(e2e): check only stored message content for a leaked card number

The Presidio spend-log check ran the card-number pattern over the whole serialized response, so a Luhn-valid usage.cost float (0.0003466000000000001) failed the streaming /v1/messages case although the stored content was <CREDIT_CARD>. The check now reads the content and text strings of the stored response, which is where a raw card would land, and still requires the placeholder there

* test(e2e): assert the proxy decodes token-array embeddings for titan

The port in #44120 carried over a legacy SDK-direct test that expected Bedrock to reject token ids with a 400. Through the proxy, /embeddings decodes token arrays to text for providers that cannot embed tokens, so titan answers 200. The test now sends a token array and its decoded sentence and requires the two vectors to match, which fails if the proxy stops decoding or decodes with the wrong tokenizer

* test(e2e): run the Bedrock extended-thinking round trip on a model that honors enabled thinking

us.anthropic.claude-sonnet-5-5 is adaptive-only, so litellm sends thinking.type=enabled with a 1024 budget as adaptive with low effort, and Bedrock returned no reasoning blocks on 5 of 5 identical Converse calls (boto3 direct agreed). us.anthropic.claude-sonnet-4-6 accepts the legacy shape verbatim and returned reasoning on 5 of 5. The non-thinking Bedrock case stays on sonnet-5-5

* test(proxy): stop unit modules forcing DEBUG logging into the event-loop lag tests

Five tests/unit modules set verbose_proxy_logger to DEBUG at import, so every xdist worker that collected them logged the 2.4MB pass-through response from a worker thread, and secret redaction of that line held the GIL for ~0.8s+ inside the timed window. The lag tests now pin the LiteLLM loggers to WARNING and freeze gc while timing, and the module-level DEBUG overrides are removed

* test(e2e): cite the tokenizer and date behind the titan token-array fixture

* test(e2e): let migration seed replicas finish their request-log indexes before cloning

Since #43948 a serving proxy builds the two LiteLLM_SpendLogs indexes on a background thread after it reports ready. The seed fixtures stopped the replica at readiness, so every cloned legacy database lacked an index no real deployment would be missing, and the v2 baseline diff refused it. Seeds now wait until both indexes exist and are valid in the database's schema

* test(passthrough): give the pass-through MockRequest an httpx URL and ASGI scope

#43626 made get_request_route read request.scope during pass-through kwarg setup; the MockRequest in tests/unit/passthrough had neither a scope nor a URL object, so both stream-param tests raised before reaching the code they check. Mirrors the repair #43626 made to the tests/pass_through_unit_tests fake

* test(integration): ignore foreign allow_all_keys MCP servers in the access matrix tool list

test_toolset_gateway_url_serves_a_team_granted_toolset_to_a_key_without_its_own_grant (#43908) registers an allow_all_keys server on the shared gateway, and allow_all_keys servers are listed to every key by design, so a matrix case running on another xdist worker at the same time saw its tools. The matrix now drops tools of allow_all_keys servers it did not create, read from LiteLLM_MCPServerTable before and after listing, and still compares everything else exactly
yuneng-berri added a commit that referenced this pull request Oct 2, 2026
…) (#44254)

* test(proxy): stop the proxy_server app fixture leaking LITELLM_LOG

The session app fixture set LITELLM_LOG=ERROR with os.environ.setdefault and never removed it, so later tests on the same xdist worker inherited it. test_drop_params_env_var spawns a subprocess with os.environ and lost the warning it asserts on. Scope the variable to the import with a MonkeyPatch context

* test(secret-detection): give the hand-built redaction request an ASGI path

Since #43975 _read_request_body checks the route path via request.scope, and a scope without path raised KeyError that was swallowed into an empty body, so chat_completion failed with a missing messages parameter. Real ASGI scopes always carry path

* test(integration): isolate litellm callback lists per sdk test

usage-based-routing-v2 Routers register their selector in litellm.callbacks and nothing removes it, not even Router.reset(). The counter TTL and Redis service metrics tests left their selectors behind, and the next usage routing test ran their pre-call checks against its own rpm=1 deployments, raising "Deployment over defined rpm limit". An autouse fixture now gives each sdk test copies of the callback lists and restores the originals afterwards

* test(integration): keep the owner-lookup fault proxy off the shared read replica

The owned proxy points DATABASE_URL at a scratch database but inherited
DATABASE_URL_READ_REPLICA from the replica job, so auth read the shared
database and rejected the freshly created key with token_not_found_in_db.
Drop the replica variable like the other scratch-database owned proxies

* test(integration): request every seeded key in the team owner breakdown

The aggregated team activity endpoint now caps breakdown.api_keys at the top
100 keys by default (#43398), so the 300 seeded keys came back as 100 rows.
The test guarantees each key is reported with its own owner, so ask for an
api_key_limit that covers all seeded keys

* test(integration): give every owned Redis its own port in the redis-cache container

On CircleCI every owned Redis ran on the fixed port 16379 inside the shared
redis-cache container. When an earlier server still held that port, the new
one failed to bind, readiness pinged the old server, the pidfile read failed
and cleanup then reported "Owned Redis still serves after shutdown"

Reserve an ephemeral port for the docker-exec path the same way the local
binary path already does, and refuse to start when something already serves
the chosen port so the failure names the real cause

* test(e2e): skip the Vertex Mistral partner case the e2e project cannot reach

The e2e Vertex project gets a 404 publisher model not found for vertex_ai/mistral-small-2503, so the case can only fail

* test(e2e): skip the Vertex gpt-oss partner case the e2e project never serves

vertex_ai/openai/gpt-oss-120b-maas has hit a 60s read timeout with no response headers on every run in the e2e Vertex project since the case was ported, and no other Vertex partner chat model passes there to switch to

* test(e2e): check only stored message content for a leaked card number

The Presidio spend-log check ran the card-number pattern over the whole serialized response, so a Luhn-valid usage.cost float (0.0003466000000000001) failed the streaming /v1/messages case although the stored content was <CREDIT_CARD>. The check now reads the content and text strings of the stored response, which is where a raw card would land, and still requires the placeholder there

* test(e2e): assert the proxy decodes token-array embeddings for titan

The port in #44120 carried over a legacy SDK-direct test that expected Bedrock to reject token ids with a 400. Through the proxy, /embeddings decodes token arrays to text for providers that cannot embed tokens, so titan answers 200. The test now sends a token array and its decoded sentence and requires the two vectors to match, which fails if the proxy stops decoding or decodes with the wrong tokenizer

* test(e2e): run the Bedrock extended-thinking round trip on a model that honors enabled thinking

us.anthropic.claude-sonnet-5-5 is adaptive-only, so litellm sends thinking.type=enabled with a 1024 budget as adaptive with low effort, and Bedrock returned no reasoning blocks on 5 of 5 identical Converse calls (boto3 direct agreed). us.anthropic.claude-sonnet-4-6 accepts the legacy shape verbatim and returned reasoning on 5 of 5. The non-thinking Bedrock case stays on sonnet-5-5

* test(proxy): stop unit modules forcing DEBUG logging into the event-loop lag tests

Five tests/unit modules set verbose_proxy_logger to DEBUG at import, so every xdist worker that collected them logged the 2.4MB pass-through response from a worker thread, and secret redaction of that line held the GIL for ~0.8s+ inside the timed window. The lag tests now pin the LiteLLM loggers to WARNING and freeze gc while timing, and the module-level DEBUG overrides are removed

* test(e2e): cite the tokenizer and date behind the titan token-array fixture

* test(e2e): let migration seed replicas finish their request-log indexes before cloning

Since #43948 a serving proxy builds the two LiteLLM_SpendLogs indexes on a background thread after it reports ready. The seed fixtures stopped the replica at readiness, so every cloned legacy database lacked an index no real deployment would be missing, and the v2 baseline diff refused it. Seeds now wait until both indexes exist and are valid in the database's schema

* test(passthrough): give the pass-through MockRequest an httpx URL and ASGI scope

#43626 made get_request_route read request.scope during pass-through kwarg setup; the MockRequest in tests/unit/passthrough had neither a scope nor a URL object, so both stream-param tests raised before reaching the code they check. Mirrors the repair #43626 made to the tests/pass_through_unit_tests fake

* test(integration): ignore foreign allow_all_keys MCP servers in the access matrix tool list

test_toolset_gateway_url_serves_a_team_granted_toolset_to_a_key_without_its_own_grant (#43908) registers an allow_all_keys server on the shared gateway, and allow_all_keys servers are listed to every key by design, so a matrix case running on another xdist worker at the same time saw its tools. The matrix now drops tools of allow_all_keys servers it did not create, read from LiteLLM_MCPServerTable before and after listing, and still compares everything else exactly

(cherry picked from commit 9b5562f)
yuneng-berri added a commit that referenced this pull request Oct 3, 2026
)

* test: repair stale and polluting tests red on scheduled main CI (#44229)

Partial backport to rc/1.104.0: only the owned Redis port, Presidio stored-content check and event-loop lag logging fixes apply here. The other repairs target tests or behavior this line does not have

* test(proxy): stop the proxy_server app fixture leaking LITELLM_LOG

The session app fixture set LITELLM_LOG=ERROR with os.environ.setdefault and never removed it, so later tests on the same xdist worker inherited it. test_drop_params_env_var spawns a subprocess with os.environ and lost the warning it asserts on. Scope the variable to the import with a MonkeyPatch context

* test(secret-detection): give the hand-built redaction request an ASGI path

Since #43975 _read_request_body checks the route path via request.scope, and a scope without path raised KeyError that was swallowed into an empty body, so chat_completion failed with a missing messages parameter. Real ASGI scopes always carry path

* test(integration): isolate litellm callback lists per sdk test

usage-based-routing-v2 Routers register their selector in litellm.callbacks and nothing removes it, not even Router.reset(). The counter TTL and Redis service metrics tests left their selectors behind, and the next usage routing test ran their pre-call checks against its own rpm=1 deployments, raising "Deployment over defined rpm limit". An autouse fixture now gives each sdk test copies of the callback lists and restores the originals afterwards

* test(integration): keep the owner-lookup fault proxy off the shared read replica

The owned proxy points DATABASE_URL at a scratch database but inherited
DATABASE_URL_READ_REPLICA from the replica job, so auth read the shared
database and rejected the freshly created key with token_not_found_in_db.
Drop the replica variable like the other scratch-database owned proxies

* test(integration): request every seeded key in the team owner breakdown

The aggregated team activity endpoint now caps breakdown.api_keys at the top
100 keys by default (#43398), so the 300 seeded keys came back as 100 rows.
The test guarantees each key is reported with its own owner, so ask for an
api_key_limit that covers all seeded keys

* test(integration): give every owned Redis its own port in the redis-cache container

On CircleCI every owned Redis ran on the fixed port 16379 inside the shared
redis-cache container. When an earlier server still held that port, the new
one failed to bind, readiness pinged the old server, the pidfile read failed
and cleanup then reported "Owned Redis still serves after shutdown"

Reserve an ephemeral port for the docker-exec path the same way the local
binary path already does, and refuse to start when something already serves
the chosen port so the failure names the real cause

* test(e2e): skip the Vertex Mistral partner case the e2e project cannot reach

The e2e Vertex project gets a 404 publisher model not found for vertex_ai/mistral-small-2503, so the case can only fail

* test(e2e): skip the Vertex gpt-oss partner case the e2e project never serves

vertex_ai/openai/gpt-oss-120b-maas has hit a 60s read timeout with no response headers on every run in the e2e Vertex project since the case was ported, and no other Vertex partner chat model passes there to switch to

* test(e2e): check only stored message content for a leaked card number

The Presidio spend-log check ran the card-number pattern over the whole serialized response, so a Luhn-valid usage.cost float (0.0003466000000000001) failed the streaming /v1/messages case although the stored content was <CREDIT_CARD>. The check now reads the content and text strings of the stored response, which is where a raw card would land, and still requires the placeholder there

* test(e2e): assert the proxy decodes token-array embeddings for titan

The port in #44120 carried over a legacy SDK-direct test that expected Bedrock to reject token ids with a 400. Through the proxy, /embeddings decodes token arrays to text for providers that cannot embed tokens, so titan answers 200. The test now sends a token array and its decoded sentence and requires the two vectors to match, which fails if the proxy stops decoding or decodes with the wrong tokenizer

* test(e2e): run the Bedrock extended-thinking round trip on a model that honors enabled thinking

us.anthropic.claude-sonnet-5-5 is adaptive-only, so litellm sends thinking.type=enabled with a 1024 budget as adaptive with low effort, and Bedrock returned no reasoning blocks on 5 of 5 identical Converse calls (boto3 direct agreed). us.anthropic.claude-sonnet-4-6 accepts the legacy shape verbatim and returned reasoning on 5 of 5. The non-thinking Bedrock case stays on sonnet-5-5

* test(proxy): stop unit modules forcing DEBUG logging into the event-loop lag tests

Five tests/unit modules set verbose_proxy_logger to DEBUG at import, so every xdist worker that collected them logged the 2.4MB pass-through response from a worker thread, and secret redaction of that line held the GIL for ~0.8s+ inside the timed window. The lag tests now pin the LiteLLM loggers to WARNING and freeze gc while timing, and the module-level DEBUG overrides are removed

* test(e2e): cite the tokenizer and date behind the titan token-array fixture

* test(e2e): let migration seed replicas finish their request-log indexes before cloning

Since #43948 a serving proxy builds the two LiteLLM_SpendLogs indexes on a background thread after it reports ready. The seed fixtures stopped the replica at readiness, so every cloned legacy database lacked an index no real deployment would be missing, and the v2 baseline diff refused it. Seeds now wait until both indexes exist and are valid in the database's schema

* test(passthrough): give the pass-through MockRequest an httpx URL and ASGI scope

* test(integration): ignore foreign allow_all_keys MCP servers in the access matrix tool list

test_toolset_gateway_url_serves_a_team_granted_toolset_to_a_key_without_its_own_grant (#43908) registers an allow_all_keys server on the shared gateway, and allow_all_keys servers are listed to every key by design, so a matrix case running on another xdist worker at the same time saw its tools. The matrix now drops tools of allow_all_keys servers it did not create, read from LiteLLM_MCPServerTable before and after listing, and still compares everything else exactly

(cherry picked from commit 9b5562f)

* fix(proxy-extras): retry P3009 when a peer already recovered the named migration row (#44283)

* fix(proxy-extras): retry P3009 when a peer already recovered the named migration row

* fix(proxy-extras): pin P3009 recovery locals as Final and cover a recovered sibling row

* test(proxy-extras): stop the P3009 ledger fakes shadowing the partition detector cursor

* fix(proxy-extras): annotate the migration ledger reader locals as Final

---------

Co-authored-by: yuneng <yuneng@berri.ai>
(cherry picked from commit 8b11b68)

---------

Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>

This branch is waiting to be deployed

1 waiting deployment
e2e-changed — 80195345 Waiting Oct 1, 2026 by yujonglee-berri via oauth #2411
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants