Skip to content

feat(lens): analyze agent activity with a separate worker - #43889

Merged
moe-berri merged 13 commits into
mainfrom
litellm_agent_engine
Sep 30, 2026
Merged

moe-berri merged 13 commits into
mainfrom
litellm_agent_engine

Conversation

@moe-berri

@moe-berri moe-berri commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Reviewing agent activity requires inspecting traces individually

How it solves it:

  • Preview matching runs, then save filters and custom questions
  • Separate workers produce findings linked to original evidence
  • Read clear findings, expand evidence, and track scan progress

Targets main directly. The trace-ingestion foundation is merged; this PR does not depend on #43865 or #43864

Intentional product change: adds Lens Beta under Observability for proxy admins and admin viewers. Setup previews recorded runs and defaults to a single scan. Issues and patterns have separate views; original evidence is expandable. Admins control setup and spending; existing Logs and trace inspection remain available

User Flow

Before: an admin can inspect recorded activity but cannot run Lens analysis

  1. Open http://127.0.0.1:18443/ui/lens/ with recorded release-review traces
  2. The dashboard returns “404, This page could not be found”

Before: Lens route missing

After: the admin configures a lens and reviews findings with source evidence

  1. Open http://127.0.0.1:18443/ui/lens/ with recorded release-review traces
  2. Click New lens, select agent traces, and filter by the swarm’s recorded metadata
  3. Add questions, choose a model and budget, then connect the Docker worker
  4. Start analysis and follow review counts, grouping progress, evidence checks, and elapsed time
  5. Open a finding, then Open trace to inspect the existing trace viewer

After: Lens dashboard with real findings

Live progress during the same scan

Implementation

The worker runs on the proxy host or another server and only needs outbound access to LiteLLM. It has a revocable Lens credential, no provider keys or database credentials. The proxy handles scoped storage reads and model calls

Each scan samples matching executions, screens content in chunks, consolidates observations across batches, then investigates up to ten patterns with a bounded read-only model loop. Both worker and proxy validate exact evidence quotes. Model responses include schema limits and get one repair attempt if validation fails. Findings can be resolved or dismissed, and new supporting executions can reopen resolved findings

PostgreSQL stores lenses, jobs, worker credentials, and findings. ClickHouse reads use scoped queries in the shared Rust storage layer. Jobs have leases and bounded retries; each model call reserves budget before inference. Setup and limits are in deploy/lens/README.md

Pre-Submission checklist

  • Added meaningful tests
  • Focused tests pass locally
  • Required CI/CD checks pass
  • Scope is limited to Lens analysis of recorded activity
  • Greptile confidence is at least 4/5 before maintainer review

Validation

Local make check passes, including Python lint, type budgets, dashboard lint, test quality and generated API types. All 60 focused Python tests pass. All 19 Lens UI tests pass, including duration-unit and analyzer setup regressions. A real ClickHouse test confirms caller-tagged requests remain eligible while server-marked analysis calls are excluded

Greptile’s Linux connectivity, duration conversion and later-page evidence findings are fixed. Veria’s caller-tag exclusion bypass is fixed using server-owned context that overwrites caller metadata. The UI selects a published worker image containing the evidence fix

A fresh paid scan reviewed the known Release-19 error, completed for $0.081326 and produced a finding with four evidence quotes. Greptile’s repeat review scored 5/5 and Veria reports no open security findings. Required Python CI and coverage remain pending. GitHub frontend lint and Dashboard build pass. The documentation check fails on the same three ClickHouse environment variables as main, so that failure is pre-existing. Bugbot could not run because the team’s spending limit was reached

Setup

Open Lens and choose Set up analysis. The URL is your existing LiteLLM deployment. Generate and copy the Docker command, then run it on a machine with Docker. The analyzer polls LiteLLM for manually started or scheduled scans. It needs no public URL, incoming port, provider keys or database credentials

Earlier end-to-end evidence, current-tip verification pending

Shared setup: PostgreSQL, ClickHouse, separate Docker workers, and real paid inference. A controlled three-agent workflow used researcher, verifier, and lead roles across 21 release-review runs, with clean inputs, injected timeouts, conflicting evidence, prompt-injection notes, long content, redacted outputs, and dropped handoffs. The agents made 188 successful model calls and stored 440 spans. Tool faults were deliberately injected; model outputs were not mocked

Before (0841d91)

  1. Serve the parent dashboard and visit http://127.0.0.1:18443/ui/lens/ with the same recorded activity
  2. Observe the 404 shown in the Before screenshot above

After (f330e45)

  1. Open the Lens URL above. Live swarm reliability contains the retained real 21-run scan, with three issue cards and seven pattern cards
  2. Click New lens. Choose Agent runs, service lens-live-swarm, and metadata swarm is release-review. The preview shows 21 named runs with timestamps and links to their original traces
  3. Continue to questions, then Review & run. Choose lens-real, a budget, and a sample limit. New lenses run once unless background monitoring is enabled
  4. Start analysis. A fresh full 21-run scan was launched through this flow and later cancelled for the user's demo deadline. A separate focused scan completed on the actual Release-19 trace with the revised writing instructions, producing “Missing verifier handoff led to wrong release decision”; its all evidence quotes were verified
  5. Return to Live swarm reliability. Needs attention puts high-priority issues first, Patterns separates trends and successful behavior, and each finding groups expandable evidence by run
  6. Open an original step or the Runs tab to inspect the frozen sample. Refreshing preserves the selected lens through its URL
  7. The retained 21-run results were not rewritten: all 72 exact quotes remain valid. The controlled dataset has 188 successful real model calls and 440 spans; it is not a recall benchmark

Setup with matching recorded activity

Type

New Feature

Caveats

Medium

  • Dependency scan flags vulnerabilities in unchanged lockfiles
  • Documentation CI flags undocumented tracing variables inherited from parent
  • ClickHouse and retained content are required for analysis
  • Reviews are bounded samples, not exhaustive session reconstruction
  • Serial model calls can make large scans take minutes
  • Analysis quality needs human review, including grouping decisions
  • Lens budgets are separate from virtual-key budgets
  • Admin-only setup and run controls are intentional in v1

Low

  • No automated fixes or persistent per-execution analysis cache
  • No remaining-time estimate; progress uses counts and elapsed time
  • Historical findings retain original wording and may contain duplicate patterns

Final Attestation

  • Required CI and review must finish before merge

Latest integration validation

Merged main through 52b9fa2 and reconciled named trace reads with Lens queries. Lens preview metadata excludes the authentication attribute server-side, with a regression test for opaque tokens

Added a real Postgres scan lifecycle test covering creation, worker authentication, claiming, progress, budget settlement, completion, reruns, cancellation and credential revocation. The existing Postgres CI job now uploads this coverage. The Rust extension builds locally; current-commit CI and updated combined Codecov coverage are pending

The batch-metadata and interactions OpenAPI failures also occur on main's unit test run. Documentation and code quality both fail on main's undocumented ClickHouse environment variables. These unrelated baseline failures are unchanged

@codecov

codecov Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

@CLAassistant

CLAassistant commented Sep 30, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@moe-berri
moe-berri removed this pull request from stack #43912 September 30, 2026 21:25
@moe-berri
moe-berri changed the base branch from litellm_trace_spend_enrichment to main September 30, 2026 21:25
@moe-berri

Copy link
Copy Markdown
Contributor Author

@greptileai please review this PR for correctness and merge readiness.

@moe-berri

Copy link
Copy Markdown
Contributor Author

@veria-ai please review this PR for correctness and merge readiness.

@moe-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor

cursor Bot commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

Bugbot is paused — on-demand spend limit reached

Bugbot uses usage-based billing for this team and has hit its on-demand spend limit.

A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue.

@codspeed

codspeed Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_agent_engine (fdbf73c) with main (632b69b)

Open in CodSpeed

@moe-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

@moe-berri

Copy link
Copy Markdown
Contributor Author

@veria-ai

@moe-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor

cursor Bot commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

Bugbot is paused — on-demand spend limit reached

Bugbot uses usage-based billing for this team and has hit its on-demand spend limit.

A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue.

@moe-berri
moe-berri marked this pull request as ready for review September 30, 2026 22:07
@moe-berri
moe-berri requested a review from a team September 30, 2026 22:07
Comment thread litellm/proxy/engine/sources.py Outdated
@moe-berri
moe-berri disabled auto-merge September 30, 2026 22:23
@moe-berri
moe-berri enabled auto-merge (squash) September 30, 2026 22:29
@moe-berri
moe-berri merged commit 6fd9334 into main Sep 30, 2026
105 of 116 checks passed
@moe-berri
moe-berri deleted the litellm_agent_engine branch September 30, 2026 22:42

This branch is waiting to be deployed

1 waiting deployment
e2e-changed — fdbf73c1 Waiting Sep 30, 2026 by moe-berri via oauth #2143
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants