Skip to content

feat(tracing): add ingestion and UI on Rust foundation - #43864

Closed
yujonglee-berri wants to merge 8 commits into
mainfrom
litellm_agent_tracing_integration
Closed

yujonglee-berri wants to merge 8 commits into
mainfrom
litellm_agent_tracing_integration

Conversation

@yujonglee-berri

@yujonglee-berri yujonglee-berri commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

This replaces #43816, which GitHub marked merged after its head briefly reached the foundation branch

TLDR

Problem this solves:

  • Developers cannot export and inspect agent traces in LiteLLM
  • An accepted export must survive a ClickHouse write failure

How it solves it:

  • Add authenticated OTLP ingestion and scoped trace reads
  • Confirm ClickHouse writes before answering OTLP success
  • Show agent traces, spans, and token counts in Logs
  • Keep parsing and storage in Rust, tenant policy in Python
  • Replace tracing README prose with concise boundary rules
  • Run new tracing and ClickHouse tests in the unit shard

Intentional product change: The Logs page gains Agent traces in All, and failed trace writes return 503 with Retry-After so exporters can retry

User Flow

Before: a developer exports agent spans, but the proxy has no trace view

  1. They enable tracing and send OTLP spans to POST http://localhost:4000/v1/traces
  2. The proxy has no ingestion endpoint to accept the spans
  3. They open http://localhost:4000/ui/?page=logs and cannot inspect the agent run

After: the same spans appear as an agent trace after storage confirms the write

  1. They enable tracing and send OTLP spans to POST http://localhost:4000/v1/traces
  2. The proxy returns OTLP success after storage accepts the spans, or 503 with Retry-After on failure
  3. They open http://localhost:4000/ui/?page=logs and inspect the trace tree, steps, graph, and span details

Pre-Submission checklist

  • I have added meaningful tests
  • Focused Rust, native bridge, and Python tests passed locally on the stack
  • My PR passes all required CI/CD checks
  • My PR's scope is as isolated as possible
  • I have received a Greptile Confidence Score of at least 4/5

Type

New Feature

Caveats (if any)

Medium

Low

  • Live proxy proof, UI screenshots, and reviews have not been rerun at this tip

Final Attestation

  • The tests check the right things, including edge cases and real-world regressions

Devin Review

@CLAassistant

CLAassistant commented Sep 30, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@codecov

codecov Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

@codspeed

codspeed Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_agent_tracing_integration (b7ab950) with main (264b09a)1

Open in CodSpeed

Footnotes

  1. No successful run was found on main (41df8cf) during the generation of this report, so 264b09a was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 3 potential issues.

1 flag not posted on this PR by your GitHub settings — view it in Devin Review. (Configure)

Devin Review

Comment on lines +120 to +122
content_encoding: Option<&str>,
max_decompressed_bytes: usize,
) -> PyResult<Bound<'py, PyAny>> {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Invalid OTLP payloads return server errors

When OTLP decoding rejects malformed or oversized compressed input, trace_decode_otlp raises ValueError. The endpoint does not catch it, so exporters receive 500 instead of a payload error.

Learn more

The Rust decoder distinguishes invalid payloads from decompressed bodies over its size limit, but this binding converts both into a Python ValueError. ingest_otlp_traces only handles RuntimeError and TracingPayloadTooLargeError. A decode failure therefore escapes as HTTP 500, while a compressed body can pass the raw-body size check and fail its decompressed limit later.

Example: A client sends invalid protobuf bytes with Content-Type: application/x-protobuf; decoding fails and the proxy returns 500 rather than 400. A small gzip body expanding past 8 MiB also returns 500 rather than 413.

Recommended fix: Preserve DecodeError::TooLarge separately at the bridge boundary and map invalid payloads to a distinct exception. Translate both in ingest_otlp_traces to 413 and 400 respectively.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Comment on lines +60 to +62
.append_pair("async_insert", "1")
.append_pair("wait_for_async_insert", "1")
.append_pair("date_time_input_format", "best_effort");

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Retried exports duplicate trace totals

When ClickHouse commits an insert but its response is lost, insert_rows returns an error and the exporter retries. Without async-insert deduplication, both writes enter the rollup and double span and token counts.

Learn more

The endpoint acknowledges an export only after the ClickHouse HTTP response succeeds. If ClickHouse commits a batch but the response is lost or times out, the exporter retries that same batch. The materialized view increments span and token totals for every insertion; neither the request nor the table establishes retry deduplication.

Example: ClickHouse stores a 10-span batch, but the 30-second HTTP timeout fires. The exporter retries and gets 200. The list view reports 20 spans even though the trace contains only 10 unique spans.

Recommended fix: Establish a stable batch identity across exporter retries or deduplicate (TeamId, TraceId, SpanId) before rollup; validate the chosen ClickHouse async-insert deduplication settings and their behavior across supported versions and partial batches.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Comment thread litellm/tracing/store.py
Comment on lines +65 to +68
LIMIT {{limit:UInt32}}
) AS t
GROUP BY t.TraceId
ORDER BY start_ms DESC, t.TraceId DESC

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Trace listing discards traces across teams

For admins viewing traces from several teams, LIST_TRACES_SQL limits grouped team rows before regrouping by trace ID. Multiple teams using the same trace ID consume page slots, so the visible page returns fewer traces and cursors skip remaining results.

Learn more

The inner query groups by both TeamId and TraceId and applies the page limit; the outer query groups again by TraceId. Admin scope has an empty team list, so equal trace IDs from different teams produce multiple inner rows but only one response row. The cursor is built from the final response rows in list_traces, and cannot recover team rows discarded by the inner limit.

Example: With page size 2, team A and team B each have trace ID same, and another trace next is older. Both same rows occupy the inner two slots, the response contains one result, and next never appears because there is no next cursor.

Recommended fix: Apply the page limit after aggregating by the public trace identity, or preserve the team identity in the response and cursor so pagination uses the same grouping key throughout.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

@greptile-apps

greptile-apps Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 0/5

[High risk] Adds tracing ingestion and storage on Rust and Python backends.

This PR is not safe to merge until trace isolation, write authorization, bounded intake, and loss-prone write paths are corrected.

Findings

  1. P1 Security Trace summaries cross key boundaries ▶
  2. P1 Security Oversized exports fill proxy memory ▶
  3. P1 Security View-only administrators can ingest traces ▶
  4. P1 Invalid exports return server errors ▶
  5. P1 Cancellation loses pending batch ▶
  6. P1 Large traces fail detail reads ▶
  7. P1 Buffered rows exceed configured cap ▶

Summary

Adds authenticated OTLP ingestion, confirmed ClickHouse writes, scoped trace reads, and a trace explorer in Logs.

  • The read scope and write-role boundaries need correction.
  • Oversized and malformed exports need bounded intake and appropriate error responses.
  • Batch-log durability and large-trace reads need attention.

Reviews (1) · Last reviewed commit: b0c1c56

Comment thread litellm/tracing/store.py
Comment on lines +55 to +58
FROM {AGENT_TRACES_TABLE} AS a
WHERE (empty({{team_ids:Array(String)}}) OR TeamId IN {{team_ids:Array(String)}})
AND ({{api_key_hash:String}} = '' OR TraceId IN (
SELECT TraceId FROM {OTEL_TRACES_TABLE} WHERE {_SCOPE_OTEL}))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 security Trace summaries cross key boundaries

If two team-less keys export spans with the same trace ID, this query checks that the caller owns a span with that ID but reads a summary shared by both keys. The caller can see the other key’s root name and input preview, while the detail queries show only their own spans. The summary must be scoped by key, not trace ID alone.

How this was verified: The materialized view groups spans by team and trace ID without an API-key field, and the list query checks key ownership only through a trace-ID subquery.

Knowledge Base Used: Proxy authentication and authorization

Comment on lines +158 to +159
if _normalize_media_type(content_type) in _BINARY_CONTENT_TYPES:
parsed_body = _parse_binary_body(await request.body())

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 security Oversized exports fill proxy memory

If the optional request-size middleware is disabled, authentication buffers the entire protobuf body here, and the trace endpoint reads the body before checking the 8 MB OTLP limit. An authenticated exporter can therefore send a much larger body and exhaust proxy memory before receiving 413. The limit needs to apply before the body is buffered.

How this was verified: Authentication and ingestion both request the complete exporter-controlled body, while ingestion checks its size only afterward.

Knowledge Base Used: Proxy authentication and authorization

Comment thread litellm/proxy/_types.py
Comment on lines +522 to +525
# agent tracing: OTLP ingest + reads (scoped to the caller's team in the handler)
"/v1/traces",
"/v1/traces/{trace_id}",
"/v1/traces/{trace_id}/spans/{span_id}",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 security View-only administrators can ingest traces

A view-only administrator can POST to /v1/traces: registering it as an LLM API route makes the route checker allow it before reaching the view-only restriction, and the ingestion handler does not check for write access. A read-only credential can consequently insert arbitrary traces.

How this was verified: The LLM-route branch permits the new path before the viewer check, and the POST handler accepts the authenticated identity without checking its role.

Knowledge Base Used: Proxy authentication and authorization

Comment on lines +130 to +133
)
})
.map_err(|error| PyValueError::new_err(error.to_string()))?;
litellm_host_python::Pythonized(spans).into_pyobject(py)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Invalid exports return server errors

Invalid OTLP JSON or protobuf, as well as a gzip body that expands beyond the size limit, becomes a ValueError here. The ingestion handler does not catch it, so the proxy returns 500 instead of a client error or 413. Exporters receive the wrong signal for a request they need to correct.

Comment on lines +70 to +73
while self.log_queue:
batch = self.log_queue[: self.batch_size]
self.log_queue = self.log_queue[len(batch) :]
if not await self._insert(batch):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Cancellation loses pending batch

If a background flush is cancelled while an insert is pending, this code has already removed the batch from the queue. _insert restores rows only for ordinary exceptions, not cancellation, so unconfirmed rows can be lost during shutdown. Keep the batch available until the write is confirmed or cancellation cleanup restores it.

Knowledge Base Used: Preserve Failed Spend Logs Across Pod Restarts

Comment thread litellm/tracing/store.py
Comment on lines +71 to +82
TRACE_SPANS_SQL: Final = f"""
SELECT o.SpanId AS span_id, o.ParentSpanId AS parent_span_id, o.SpanName AS name,
o.ObservationType AS type, o.AgentName AS agent, o.StatusCode AS status,
o.StatusMessage AS status_message,
toUnixTimestamp64Nano(o.Timestamp) AS start_ns, o.Duration AS duration_ns,
o.ServiceName AS service, o.InputPreview AS input_preview, o.Model AS model,
o.InputTokens AS input_tokens, o.OutputTokens AS output_tokens,
o.LiteLLMRequestId AS litellm_request_id
FROM {OTEL_TRACES_TABLE} AS o
WHERE o.TraceId = {{trace_id:String}} AND {_SCOPE_OTEL}
ORDER BY o.Timestamp
LIMIT 1 BY o.SpanId

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Large traces fail detail reads

When an agent run has more than 1,000 spans, this unpaginated detail query exceeds the native read transport’s 1,000-row limit, which is configured to throw on overflow. GET /v1/traces/{trace_id} then fails instead of displaying the trace tree.

Comment on lines +59 to +63
def enqueue(self, rows: list[dict[str, Any]]) -> None:
"""Never awaits ClickHouse. Kicks off an early flush once a full batch is queued."""
self.log_queue.extend(rows)
if len(self.log_queue) >= self.batch_size:
asyncio.get_running_loop().create_task(self.flush_queue())

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Buffered rows exceed configured cap

enqueue adds rows without checking CLICKHOUSE_MAX_BUFFERED_ROWS; is_full() is only advisory. The spend logger in PR #43865 calls enqueue directly, so sustained ClickHouse failures can grow its in-memory queue past the configured cap.

Base automatically changed from litellm_trace_query_api to main September 30, 2026 19:00
@yujonglee-berri
yujonglee-berri force-pushed the litellm_agent_tracing_integration branch from df529d4 to b7ab950 Compare September 30, 2026 19:00
Comment thread litellm/tracing/store.py
groupUniqArrayArray(AgentNames) AS AgentNames
FROM {AGENT_TRACES_TABLE} AS a
WHERE (empty({{team_ids:Array(String)}}) OR TeamId IN {{team_ids:Array(String)}})
AND ({{api_key_hash:String}} = '' OR TraceId IN (

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Medium: Cross-key trace summary disclosure

A team-less key holder who knows another key’s trace ID can ingest a non-root span with that ID, then request /v1/traces to read the other key’s root input preview. The subquery checks ownership of the ID, but the outer query aggregates all rows with that team ID and trace ID; retain the key hash in the rollup and filter it before aggregation.

.unwrap_or(SpanKind::Unspecified)
.as_str_name()
.to_owned(),
resource_attributes: resource_attributes.clone(),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Medium: Small OTLP exports can exhaust worker memory

An authenticated caller can send a roughly 2 KB gzip export containing 20,000 spans and one 64 KiB resource attribute; cloning that attribute into every span allocates about 1.25 GiB before the insert-size limit runs. Bound the span count and aggregate decoded size before making per-span resource copies.

@veria-ai

veria-ai Bot commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

PR overview

The PR adds trace ingestion and a tracing UI on a Rust foundation, including OTLP processing and trace storage and listing.

Two security issues remain open, with none addressed yet. A team-less key holder who knows another key’s trace ID can expose that key’s root input preview through trace listing. Authenticated callers can also submit a small compressed OTLP export that expands into roughly 1.25 GiB of allocations, potentially exhausting worker memory before insert-size limits apply.

Open issues (2)

Fixed/addressed: 0 · PR risk: 7/10

yujonglee-berri and others added 8 commits September 30, 2026 19:27
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration
devin-ai-integration Bot force-pushed the litellm_agent_tracing_integration branch from b7ab950 to 4a5f738 Compare September 30, 2026 19:45
@devin-ai-integration
devin-ai-integration Bot removed this pull request from stack #43866 September 30, 2026 19:45
@yujonglee-berri
yujonglee-berri enabled auto-merge (squash) September 30, 2026 19:45
auto-merge was automatically disabled September 30, 2026 20:00

Pull request was closed

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants