Skip to content

fix(otel v2): a callback exports only to the destination it owns, so vendor spans stay off the generic OTEL_* collector - #40881

Open
devin-ai-integration[bot] wants to merge 6 commits into
mainfrom
litellm_otel_v2_langfuse_exporter_isolation
Open

devin-ai-integration[bot] wants to merge 6 commits into
mainfrom
litellm_otel_v2_langfuse_exporter_isolation

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 12, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

How it solves it:

  • a preset lists only the exporter it owns, never the inherited OTEL_* one
  • otel alone serves OTEL_*; a generic collector plus a vendor now needs both in callbacks
  • otel reuses only a logger whose every exporter is the operator's collector (a plain otel logger, or a mapper-only preset like langtrace), and the published global provider prefers that one
  • when otel rides a mapper-only preset built before it, that logger still gets callback_settings.otel and the Phoenix auto-init a lone otel entry performs
  • the Langfuse logger stamps the request root only when its own provider created it
  • Weave v2 reads its config without rewriting the process-wide OTEL_* env
  • the interim langfuse.* prefix filter from the first revision of this PR is gone

User Flow

Before: an operator who ships traces to both a generic OTLP collector and Langfuse sees Langfuse-only attributes and duplicate spans in the generic collector

  1. They set callbacks: ["otel", "langfuse_otel"], OTEL_EXPORTER_OTLP_ENDPOINT to their collector, and LANGFUSE_* to their Langfuse project
  2. A developer sends POST https://litellm-domain/v1/chat/completions with header langfuse_trace_name: private-langfuse-only-name and gets HTTP 200
  3. Langfuse shows a trace named private-langfuse-only-name with a chat proof-model generation, as expected
  4. The generic collector shows the request root span with langfuse.trace.name=private-langfuse-only-name, the standard chat proof-model span, and a second chat proof-model span carrying langfuse.observation.input, langfuse.observation.output and six more langfuse.* keys
  5. If they swap the callback order to ["langfuse_otel", "otel"], the collector instead gets every span with langfuse.* keys, and with callbacks: ["langfuse_otel"] and no OTEL_* env the same spans are printed to the proxy's stdout

After: the generic collector only sees standard spans, Langfuse is unchanged

  1. They set callbacks: ["otel", "langfuse_otel"], OTEL_EXPORTER_OTLP_ENDPOINT to their collector, and LANGFUSE_* to their Langfuse project
  2. A developer sends POST https://litellm-domain/v1/chat/completions with header langfuse_trace_name: private-langfuse-only-name and gets HTTP 200
  3. Langfuse shows a trace named private-langfuse-only-name with a chat proof-model generation, as before
  4. The generic collector shows the request root, the auth span and one chat proof-model span, all with zero langfuse.* keys
  5. Swapping the callback order changes nothing, and with callbacks: ["langfuse_otel"] alone nothing is printed to stdout; an operator who wants the collector as well adds otel to callbacks

Relevant issues

Follow-up to #40793 and #28909

Linear ticket

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Live no-mock audit of every OTel v2 callback (otel, langfuse_otel, arize, arize_phoenix, weave_otel, newrelic, agentops, levo, langtrace). Before is the merge base 27f8d2a9baa1f81b37243878a17f7e55733a64e9, After is this tip b565a8d155f1dc2b19d208295c39bce0cce9f87a. Each leg runs from its own worktree with its own Postgres database and a 2-worker proxy (Before on 21081, After on 21080), started as python litellm/proxy/proxy_cli.py --config cfg/<K>.yaml --port $PORT --num_workers 2 --use_v2_migration_resolver with LITELLM_OTEL_V2=true and LITELLM_DISABLE_NO_REDIS_WARNING=true. Providers are real OpenAI (openai/gpt-5.4-mini as proof-model) and real Anthropic (anthropic/claude-sonnet-4-6 as proof-claude). Destinations are a real Jaeger all-in-one as the generic OTEL_EXPORTER_OTLP_ENDPOINT=http://127.0.0.1:21118 collector (UI on 21186), real Langfuse cloud, real Weave, real Arize and New Relic OTLP ingest endpoints, and a real Arize Phoenix on 21106. Two destinations are disclosed OTLP/HTTP emulators because no account exists: Levo (21119) and the team-level Langfuse callback_vars host (21120). Every request is checked at every destination by its gen_ai.response.id, never by count. Full per-cell tables (146 happy cells, sad, edge and 190 chaos requests) live in the audit report attached to the linked Devin session

Config for the all-callbacks cases (cfg/K2.yaml); the other configurations change only callbacks and the env, as named in each step

model_list:
  - model_name: proof-model
    litellm_params:
      model: openai/gpt-5.4-mini
      api_key: os.environ/OPENAI_API_KEY
  - model_name: proof-claude
    litellm_params:
      model: anthropic/claude-sonnet-4-6
      api_key: os.environ/ANTHROPIC_API_KEY
general_settings:
  master_key: sk-1234
  database_url: os.environ/DATABASE_URL
litellm_settings:
  callbacks: ["otel", "langfuse_otel", "arize", "arize_phoenix", "weave_otel", "newrelic", "levo", "langtrace"]

Request and readback shapes used by every step below. $K is the configuration, $LEG is base or head, every stream is consumed to the end before the readback. gen=N[...] is the number of spans the generic collector holds for that response id and the vocabulary on each, root is the vendor keys on the request root span at the collector

H=(-H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -H "langfuse_trace_name: litaudit-$K-$LEG")
curl -s -D - localhost:$PORT/v1/chat/completions "${H[@]}" \
  -d '{"model":"proof-model","messages":[{"role":"user","content":"Reply with the single word ok"}],"max_tokens":20,"stream":false}'
curl -s -D - localhost:$PORT/v1/messages "${H[@]}" \
  -d '{"model":"proof-claude","messages":[{"role":"user","content":"Reply with the single word ok"}],"max_tokens":20,"stream":true}'
curl -s -D - localhost:$PORT/v1/responses "${H[@]}" \
  -d '{"model":"proof-model","input":"Reply with the single word ok","max_output_tokens":20}'
curl -s "localhost:21186/api/traces?service=litaudit-$LEG&tags=%7B%22gen_ai.response.id%22%3A%22$RID%22%7D"   # generic collector
curl -s -u "$LANGFUSE_PUBLIC_KEY:$LANGFUSE_SECRET_KEY" "$LANGFUSE_HOST/api/public/observations?name=chat%20proof-model&limit=50"
curl -s localhost:21119/dump; curl -s localhost:21120/dump                                                  # Levo and team emulators
psql "$DATABASE_URL" -c "select request_id from \"LiteLLM_SpendLogs\" where request_id = '$RID'"

Before (27f8d2a)

/v1/chat/completions, all eight v2 callbacks

  1. K=K2 chat curl, stream off: HTTP 200, chatcmpl-EO7dw6wP6ogaPNQnoJYXyLVOZ3KL3. Collector: gen=5[generic, langfuse, openinference, openinference, openinference+weave], root langfuse.trace.name. Langfuse 1 observation named litaudit-K2-base, Weave 4 duplicate calls, Phoenix 1, Levo 1, spend log 1
  2. chat curl, stream on: HTTP 200, chatcmpl-EO7dwRvaOrVL0u2BGavbxx5tgjZRV. Collector: gen=5 with the same four vendor duplicates, root leaked
  3. Same request from the dashboard leg (langfuse_trace_name: audit-ui-base): HTTP 200, chatcmpl-EO9HBH6On35YrZDOB0BsN8xh0EGXR, x-litellm-call-id: eaf90459-63ca-4e74-ae03-5e411d53fdf9, Jaeger trace 15b5d6a1ca7992b1af81875fb2003d91 with 9 spans, four chat proof-model, langfuse.observation.*, llm.* plus openinference.span.kind, weave.*, and langfuse.trace.name=audit-ui-base on the root

before: chat trace with 9 spans and vendor duplicates

before: OpenInference and Weave keys on a duplicate chat span

before: request root span carrying langfuse.trace.name

/v1/messages, all eight v2 callbacks

  1. messages curl, stream off: HTTP 200, msg_011Cf3y8wDRXYyhQFMfktRXe. Collector gen=5[langfuse, generic, openinference, openinference, openinference+weave], root leaked
  2. messages curl, stream on: HTTP 200, msg_011Cf3y91yuUY5mw2p6Bnsec. Collector gen=5, root leaked. Langfuse 1 named, Phoenix 1, Levo 1, spend log 1
  3. Dashboard leg: HTTP 200, msg_011Cf477E73eTY9bqL4N5zqE, x-litellm-call-id: 32828cc4-27b3-4f6b-8023-54646d5095da

/v1/responses, all eight v2 callbacks

  1. responses curl, stream off: HTTP 200, resp_iyx7bFqbWthEBEhZxSQzzFFmqWyYNUf1P-OZNZi6_Vw. Collector gen=5, root leaked. Langfuse 1 named, Phoenix 1, Levo 1, spend log 1
  2. responses curl, stream on: HTTP 200, resp_FfmUL5RPf2GlJnYuRfTkF1gsvoRvQBhxd9poyOhZFAj. The streamed client id differs from the logged id and the collector was restarted (chaos C1) before a re-read, so this cell is NOT RUN at the collector; the LiteLLM SDK stream cells below cover streamed responses
  3. Dashboard leg: HTTP 200, x-litellm-call-id: d92d2d21-9025-4c9a-b623-d37eba372b34

OpenAI, Anthropic and LiteLLM SDK clients through the proxy, all eight v2 callbacks

  1. openai sync and async chat.completions.create and responses.create, stream on and off, anthropic sync and async messages.create, stream on and off, base_url=http://localhost:21081: 12 cells, every one HTTP 200, every collector readback gen=5 with vendor duplicates and root leaked
  2. litellm.completion / acompletion / responses / aresponses with model="openai/proof-model", api_base=http://localhost:21081, stream on and off: 8 cells, all HTTP 200, e.g. chatcmpl-EO7rHY8MVgjDTruzGHzG3YnfTY2tQ gen=5[openinference, ...], resp_8tN4ZskAq-zr5s-QUvhc2lILlK0awQW-gGEZNVCqgNf (stream) gen=5[langfuse, ...]

Other callback configurations, one 2-worker boot each, three endpoints each

  1. K1 ["otel"], OTEL_* set: chatcmpl-EO7uuIthLSg0LTjvyqn7GbGojK0yF gen=1[generic], root clean
  2. K3 ["langfuse_otel", "arize", "arize_phoenix", "weave_otel", "newrelic", "levo"] without otel, OTEL_* set: chatcmpl-EO7wOeRQ8tHBaHypZupeVtsWWPTza gen=4[langfuse, openinference, openinference, openinference+weave], root leaked; the collector receives vendor spans it never asked for
  3. K4 ["langfuse_otel", "otel"]: chatcmpl-EO7yGZ7sKjWKSJeHuSKmHwDR0AjBX gen=1[langfuse], root langfuse.trace.name; the otel entry reused the Langfuse logger, Langfuse 1 named litaudit-K4-base
  4. K5 ["weave_otel", "otel"]: chatcmpl-EO7zugaamhNxAkpNlwH0p7Obyt4ht gen=1[openinference+weave], Weave 1
  5. K6 ["langtrace", "otel"] plus PHOENIX_COLLECTOR_HTTP_ENDPOINT: chatcmpl-EO81KwwoiVUjFi5dx395bUKyIH4UB Phoenix 0 spans for all three ids, the auto-init is skipped behind the mapper-only callback
  6. K7 ["otel"] plus PHOENIX_*: chatcmpl-EO82ig93HSNfkJRoExxVwYi4qBAXb Phoenix 3 spans
  7. K8 ["langtrace"]: chatcmpl-EO843K7ztZpHbxpdq5crqlTaeSeze gen=1[langtrace]
  8. K9 ["langfuse_otel"], no OTEL_*: chatcmpl-EO87R7XY5rdcsk4TMwR4TOvpd88hx Langfuse 1 named litaudit-K9-base, and grep -c '"resource":' proxy.log = 12, every span printed as JSON to stdout
  9. K10 ["otel"], OTEL_EXPORTER=console: chatcmpl-EO892WX75MxXwpo1SQiTWs6R5KV1s, 12 span objects on stdout, spend log 1
  10. K11 ["otel", "langfuse_otel"], LANGFUSE_* unset: chatcmpl-EO8ATy5uM1fxVJTEFDfHCMYO4cSn5 HTTP 200, gen=1[generic], Langfuse 0
  11. K12 ["otel", "langfuse_otel"] with a team whose metadata.logging names langfuse_otel and callback_vars.langfuse_host=http://127.0.0.1:21120, keys on that team, on a team without settings, and with no team, then /team/update and a 65 s wait for the metadata cache: Team A chatcmpl-EO8GNH0tzGQZwsx8w2xFBUNW0PJeQ gen=2[generic, langfuse], root leaked on the collector and on the tenant root, team sink 1 generic span; after the update chatcmpl-EO8Ki1obnYCh5FDjO4DzCXnKJfhV0 gen=2[langfuse, generic], root leaked, team sink 1
  12. K13 ["otel", "agentops"], no AgentOps key: chatcmpl-EO8C1PnJZqfD20XAAhHfvvUsraTYE gen=2[generic, generic]
  13. K14 K2 plus callback_settings.otel.message_logging: false: chatcmpl-EO8OIHuLpMYwM4sXV5YOwxDBz8sgN gen=5 vendor duplicates as K2
  14. K15 callbacks: ["nope"] then callbacks: null: each worker exits at boot, ValueError: Empty module name and AttributeError: 'NoneType' object has no attribute 'startswith'
  15. K16 callbacks: [], OTEL_* set: chatcmpl-EO8Tx3v9frNFsCg1seRndO6Lp5qhk gen=1[generic] from the env-built provider

Sad paths, K2 configuration

  1. "model":"no-such-model" on all three routes: HTTP 400 Invalid model name passed in model=no-such-model
  2. proof-model pointed at a bad OPENAI_API_KEY: HTTP 401 AuthenticationError ... Incorrect API key provided, then HTTP 429 No deployments available ... cooldown once the router cooled the deployment down
  3. "max_tokens": 10000000: HTTP 400 on chat and messages, HTTP 200 on responses (provider accepted it)
  4. POST /key/health on a key with Langfuse callback_vars: HTTP 200; GET /health/liveliness: HTTP 200 after every failure above; a valid key on proof-model right after: HTTP 200
  5. A real vendor OTLP endpoint answered 403 to export batches during K2, K3 and K14: callers unaffected, logged as exporter failures only

Edge paths, K2 configuration

  1. Three identical chat requests in a row: each id landed exactly five spans on the collector (one generic plus four vendor duplicates), exactly one spend row
  2. Three concurrent POST /team/update on Team A while 12 Team A requests ran: all HTTP 200, zero new tracebacks in either worker log

Chaos, K2 configuration, 30 to 40 concurrent mixed requests (chat, messages, responses, stream on and off)

  1. C1 docker stop Jaeger for 20 s mid-burst, restart within the exporter retry window: 30/30 callers HTTP 200, Levo got every id once, collector served post-restore traffic. Exactly-once landing of the outage-window spans is not provable, the in-memory Jaeger dropped its store on stop
  2. C2 docker pause Jaeger for 20 s mid-burst, unpause: 30/30 HTTP 200, /health/liveliness 200 every 5 s, after unpause every id landed with gen=5 leaked, spend log 30/30
  3. C3 SIGTERM the proxy master 1.2 s into a 40-request burst, restart: 40/40 accepted requests finished 200, proxy back with 2 workers, spend log 40/40, exporter queue drained for 22/40 ids
  4. C4 kill -KILL one worker 4 s into a 30-request burst: 29/30 HTTP 200 plus 1 connection reset, master respawned a worker, 25/30 spend rows at readback
  5. C5 docker stop a dedicated Postgres for 27 s mid-burst (control proxy on 21083): 29 HTTP 200 plus 1 HTTP 429 (provider cooldown), liveliness 200 throughout, /health/readiness reported db: disconnected then recovered, collector 28/30 ids gen=5 leaked, 26/30 spend rows

Dashboard and destination UI

  1. Open http://localhost:21081/ui, log in as admin with the master key, open Logs (/ui/?page=logs): the audited ids are listed
  2. Click the chatcmpl-EO9HBH6On35YrZDOB0BsN8xh0EGXR row: detail page with request, response, tokens and cost
  3. Open Playground, pick proof-model, send a message: the assistant reply renders
  4. Open http://localhost:21106 and find the trace for the same id: Phoenix trace detail present

before: dashboard logs page

before: dashboard log detail

before: playground

before: Phoenix trace detail

After (b565a8d)

/v1/chat/completions, all eight v2 callbacks

  1. K=K2 chat curl, stream off: HTTP 200, chatcmpl-EO7dwcOxIvfbuK6SjYTm7FJOWnjqt. Collector: gen=2[generic, langtrace], root clean. Langfuse 1 observation named litaudit-K2-head, Weave 1 call, Phoenix 1, Levo 1 (generic vocabulary), spend log 1
  2. chat curl, stream on: HTTP 200, chatcmpl-EO7dwfPqEVuGiZus1Axra5YkbsjUR. Collector: gen=2[generic, langtrace], root clean
  3. Same request from the dashboard leg (langfuse_trace_name: audit-ui-head): HTTP 200, chatcmpl-EO9H7ItlZZ1E3hAcTmHUuOk3cBrxz, x-litellm-call-id: 2f97ce6e-2490-4ca3-95e8-f8c77d3f824d, Jaeger trace 371e55467b8f3d896a7757c811e2b802 with 6 spans, one generic chat proof-model plus the mapper-only Langtrace copy, zero langfuse.*, openinference.* or weave.* keys, root clean

after: chat trace with 6 spans, gen_ai.* only

/v1/messages, all eight v2 callbacks

  1. messages curl, stream off: HTTP 200, msg_011Cf3y8wRpYeGTEaXnKCdr1. Collector gen=2[generic, langtrace], root clean
  2. messages curl, stream on: HTTP 200, msg_011Cf3y91oVTj1xxignTvhgh. Collector gen=2, root clean. Langfuse 1 named, Phoenix 1, Levo 1, spend log 1
  3. Dashboard leg: HTTP 200, msg_011Cf476ujVLz71YzF8UYBJz, x-litellm-call-id: a5eb1638-8b80-4b39-b560-2b1280bd3efc, Jaeger trace 5070fb5d5a591c8977e2d053d9d807a7, 6 spans, 0 vendor keys

after: /v1/messages trace, no vendor keys

/v1/responses, all eight v2 callbacks

  1. responses curl, stream off: HTTP 200, resp_ySsvIjjsBztPCyPKLPl6yZ7U1iPZKto6NjtK69fwDPd. Collector gen=2[generic, langtrace], root clean. Langfuse 1 named, Phoenix 1, Levo 1, spend log 1
  2. responses curl, stream on: HTTP 200, resp_afI0c6LFBbV-RlM5MVy7toZya4xw2veXBGHPjm2IIMd. NOT RUN at the collector for the same reason as Before; covered by the LiteLLM SDK stream cells below
  3. Dashboard leg: HTTP 200, x-litellm-call-id: 7495d2be-7b0b-476a-aeab-ed976039a5a9, Jaeger trace f6a40488b34e91dbdd502b7182bd3008, 6 spans, 0 vendor keys

after: /v1/responses trace, no vendor keys

OpenAI, Anthropic and LiteLLM SDK clients through the proxy, all eight v2 callbacks

  1. Same 12 openai and anthropic cells against http://localhost:21080: every one HTTP 200, every collector readback gen=2[generic, langtrace], root clean
  2. Same 8 litellm.* cells: all HTTP 200, e.g. chatcmpl-EO7rHiRh277TmeYXJThpYnuz2XWo3 gen=2[generic, langtrace], resp_KRxlbk8nfZqSDKJkGyXB5Gxgx0E_WFzlvf0RECWfJBM (stream) gen=2[generic, langtrace], resp_pa9f8XtwUkIa9UrviJNxMV-h_BZevv4GgIc_h0SFoZt (async stream) gen=2

Other callback configurations, one 2-worker boot each, three endpoints each

  1. K1 ["otel"], OTEL_* set: chatcmpl-EO7uuuTvMGDAECG2J1Z6n26KwQaqa gen=1[generic], root clean, unchanged
  2. K3 same six vendor callbacks without otel, OTEL_* set: chatcmpl-EO7wOypCH9WFovofZi6DmRN0UWmcC gen=0, root clean, Phoenix 1 for every id. The collector gets nothing until otel is added
  3. K4 ["langfuse_otel", "otel"]: chatcmpl-EO7yGuXnnAUcq9c4Iy3jwwYzMjm9r gen=1[generic], root clean, Langfuse 1 named litaudit-K4-head; callback order no longer matters
  4. K5 ["weave_otel", "otel"]: chatcmpl-EO7zuPj4fVhsbvhzhN59sCCOSLtab gen=1[generic], Weave 1
  5. K6 ["langtrace", "otel"] plus PHOENIX_*: chatcmpl-EO81KkemrqG0n5ZZOF8aBk4eo7A06 Phoenix 3 spans, 2 and 2 for the messages and responses ids, matching a lone otel
  6. K7 ["otel"] plus PHOENIX_*: chatcmpl-EO82htlBdXhuE6j67TO6p7M9qZGyq Phoenix 3 spans, unchanged
  7. K8 ["langtrace"]: chatcmpl-EO843M89mTACc8nDUoH09qwvdrUaE gen=1[langtrace], unchanged
  8. K9 ["langfuse_otel"], no OTEL_*: chatcmpl-EO87R5haOpiQQOccIM5oDBFgrsDDo Langfuse 1 named litaudit-K9-head, grep -c '"resource":' proxy.log = 0
  9. K10 ["otel"], OTEL_EXPORTER=console: chatcmpl-EO892o401RHwawHIzcNM6cwvgY2az, 12 span objects on stdout, unchanged
  10. K11 ["otel", "langfuse_otel"], LANGFUSE_* unset: chatcmpl-EO8ATCjJFDFsgP817wkhOIAtMsNC7 HTTP 200, gen=1[generic], Langfuse 0, unchanged
  11. K12 same teams, keys and 65 s cache wait: Team A chatcmpl-EO8GB3eKElZefmAJEuEMQWHt8JsDO gen=1[generic], root clean on the collector and the tenant root, team sink 1 generic span, Langfuse 1 named litaudit-K12-TEAM_A-head, spend log 1; team without settings and no-team keys the same; after the update chatcmpl-EO8KiibJNnmkcQLyD2Lqa7ycjceKd gen=1[generic], team sink 1, root clean
  12. K13 ["otel", "agentops"], no AgentOps key: chatcmpl-EO8C19M0dvC13jpEdwkhL3HbHTXUA gen=1[generic], caller 200, AgentOps exporter 401 in the log only
  13. K14 K2 plus callback_settings.otel: chatcmpl-EO8OI8MEDfuqBmdhnVCHLHh8s9PuS gen=2[generic, langtrace], root clean
  14. K15 ["nope"] and null: same boot failures as Before, unchanged
  15. K16 [], OTEL_* set: chatcmpl-EO8Tw5VgnZoN3q91hT09udtICbktG gen=1[generic], unchanged

Sad paths, K2 configuration

  1. Bad model: HTTP 400 Invalid model name passed in model=no-such-model on all three routes
  2. Bad provider key: HTTP 401 AuthenticationError ... Incorrect API key provided, then HTTP 429 cooldown, same as Before
  3. "max_tokens": 10000000: HTTP 400 on chat and messages, HTTP 200 on responses, unchanged
  4. POST /key/health 200, GET /health/liveliness 200 after every failure, valid key right after 200
  5. Vendor OTLP 403 during K2, K3 and K14: callers unaffected

Edge paths, K2 configuration

  1. Three identical chat requests: each id landed exactly one generic span plus one Langtrace-mapped span on the collector, exactly one spend row, no duplicates
  2. Three concurrent POST /team/update during 12 Team A requests: all HTTP 200, zero new tracebacks

Chaos, K2 configuration, 30 to 40 concurrent mixed requests (chat, messages, responses, stream on and off)

  1. C1 stop Jaeger 20 s mid-burst, restart: 30/30 HTTP 200, Levo every id once, collector clean after restore, exactly-once not provable for the outage window (in-memory store)
  2. C2 pause Jaeger 20 s mid-burst, unpause: 30/30 HTTP 200, liveliness 200 throughout, after unpause every id landed exactly once with gen=2 clean, spend log 30/30
  3. C3 SIGTERM 1.2 s into a 40-request burst, restart: 40/40 accepted requests finished 200, proxy back with 2 workers, spend log 40/40, exporter queue drained for 0/40 ids (queued spans lost on SIGTERM, same class of loss as Before)
  4. C4 kill one worker 4 s into a 30-request burst: 30/30 HTTP 200, replacement worker spawned, 29/30 ids on the collector clean, 25/30 spend rows
  5. C5 stop dedicated Postgres 27 s mid-burst (proxy on 21082): 29 HTTP 200 plus 1 HTTP 429, liveliness 200 throughout, readiness db: disconnected then recovered, collector 24/30 ids gen=2 clean, 28/30 spend rows

Dashboard and destination UI

  1. Open http://localhost:21080/ui, log in as admin with the master key, open Logs: the audited ids are listed
  2. Click the chatcmpl-EO9H7ItlZZ1E3hAcTmHUuOk3cBrxz row: detail page with request, response, tokens and cost
  3. Open Playground, pick proof-model, send a message: the assistant reply renders
  4. Open Logging & Alerts (/ui/?page=logging): all eight configured v2 callbacks listed
  5. Open http://localhost:21106 and find the trace for the same id: Phoenix trace detail present

after: dashboard logs page

after: dashboard log detail

after: playground

after: callbacks list

after: Phoenix trace detail

Recording of the UI flow on both legs, 1 fps preview (the full annotated mp4 is attached to the linked Devin session)

UI audit recording preview

Not run on either leg: destination 404 (the organic vendor 403 covers a rejecting destination), Langfuse host on a closed port, direct os.environ inspection of a K5 worker (Weave isolation inferred from the clean generic span), Levo idempotence for the three repeated-request ids, the streamed /v1/responses collector cells for raw curl and the OpenAI SDK (see above), AgentOps happy path (no key), Arize and New Relic readback (ingest-only keys), Langfuse cloud UI screenshot (no UI login)

Type

🐛 Bug Fix

Caveats (if any)

Severe

Medium

  • per-team Langfuse destinations receive the fan-out copy of the generic spans; they never carried langfuse.observation.* and now also lack the root langfuse.trace.name that leaked there before
  • tenant fan-out mirrors the global provider's spans as they are (see the TenantFanOutSpanProcessor docstring). When a vendor preset is the only v2 callback, its provider is the global one, so a team destination on a different vendor receives that vendor's vocabulary, including the Langfuse root attributes. Base does the same; re-mapping fan-out copies per destination owner is a follow-up

Low

  • sync SDK calls emit no OTel v2 spans on either side; pre-existing, not touched here
  • SIGTERM, a worker kill and a Postgres outage lose queued exports and spend-log rows on both legs (chaos C3, C4, C5 above); pre-existing, not touched here
  • legacy (non LITELLM_OTEL_V2) OpenTelemetry and Langfuse loggers are untouched; get_weave_otel_config keeps its v1 side effect of setting OTEL_* env
  • with PHOENIX_* env set, otel listed after a mapper-only preset now also auto-starts the Phoenix logger, as a lone otel entry always did; base skipped it on that order
  • callbacks: ["otel", "langtrace"] (otel first) still builds a second logger for the Langtrace mapper, so the collector gets two chat spans per request, one of them with langtrace.* keys. Base does the same; a mapper-only preset reusing an already built otel logger is a follow-up
  • callbacks: ["langtrace"] alone now honors callback_settings.otel (service name, message logging, metrics, events, baggage, content capture), which base ignored for that config
  • CI proxy-behavior fails on test_join_binds_the_membership_to_the_requested_team (TypeError: object MagicMock can't be used in 'await' expression from auth_checks.py). The same test fails the same way on main itself at 30f33a9 (the proxy-behavior job on that commit), with none of this PR's changes, and this PR touches no auth code

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR
  • merge-ref check: b565a8d merged into current main (30f33a9, which already contains the old litellm_internal_staging history via chore(ci): remerge internal staging #40943) passes the OTel, Weave, and Langfuse OTel suites (767 passed) and the ["otel", "langfuse_otel"] live leg (generic sink 12 spans, 0 langfuse.* keys; Langfuse sink 4 spans, trace names present), run before the audit above
  • b565a8d passes /live-pr-risk

Note

Medium Risk
Changes default observability routing for multi-callback and preset-only setups (breaking if users relied on inherited OTEL_* exporters), but behavior is heavily tested and confined to OTel v2 integration plumbing.

Overview
OTel v2 callbacks no longer fan vendor traces into the operator’s generic OTEL_* collector. Vendor presets (Langfuse, Arize, Phoenix, Weave, New Relic, etc.) now register only their own exporter instead of merging in *base.exporters, so langfuse.* / OpenInference spans stay on the vendor backend. The otel callback alone owns the generic collector; select_global_otel_v2_logger and proxy_server.open_telemetry_logger prefer a logger whose exporters are all operator-owned (serves_generic_collector), and the otel entry reuses mapper-only presets (e.g. Langtrace) that already serve the collector.

Langfuse stamps trace name and root I/O only when the request root belongs to its tracer provider (_owned_recording_root). Weave v2 uses read_weave_otel_config so preset setup does not rewrite process-wide OTEL_* env (v1 get_weave_otel_config keeps the old side effect). Credential-gated exporter layering helpers are removed; _maybe_construct_otel_v2 is simplified accordingly. Operators who want both destinations need callbacks: ["otel", "<vendor>"].

Reviewed by Cursor Bugbot for commit b565a8d. Bugbot is set up for automated code reviews on this repo. Configure here.

Link to Devin session: https://app.devin.ai/sessions/ef5e75844495467e92662e55750e1dd1
Open in Devin Desktop: https://app.devin.ai/desktop/session/ef5e75844495467e92662e55750e1dd1?variant=devin
Requested by: @yucheng-berri

…not Langfuse sinks

The Langfuse preset builds its exporter list as [*base.exporters, langfuse], and both
the Langfuse mapper and the Langfuse logger's log_pre_api_call stamp langfuse.*
attributes onto spans every exporter of the provider shares, so the operator's generic
OTLP collector received Langfuse-only attributes whenever otel and langfuse_otel ran
together. Wrap every exporter that is not a Langfuse sink in a processor that hands it
a span view without langfuse.* attributes, and apply the same rule to per-request
tenant destinations in the fan-out.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR. Add '(aside)' to your comment to have me ignore it.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@codspeed

codspeed Bot commented Sep 12, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_otel_v2_langfuse_exporter_isolation (b565a8d) with main (30f33a9)1

Open in CodSpeed

Footnotes

  1. No successful run was found on main (e240997) during the generation of this report, so 30f33a9 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

@greptile-apps

greptile-apps Bot commented Sep 12, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR separates OpenTelemetry v2 export ownership so vendor callbacks send spans only to their own destinations while the otel callback owns the generic collector

  • Vendor presets no longer inherit generic OTEL_* exporters
  • Global provider selection prefers a generic-collector logger regardless of callback order
  • Mapper-only presets retain generic collector settings and Phoenix initialization
  • Langfuse writes root attributes only when its provider owns the root
  • Weave v2 reads configuration without mutating process-wide OTEL environment variables
  • Regression coverage exercises preset ownership, callback ordering, degraded credentials, tenant fan-out, and Langfuse root behavior

Confidence Score: 5/5

The PR appears safe to merge, with no outstanding previous findings or accepted new issues

The current tip is identical to the previous review SHA, both previous threads are resolved, and the full-PR rule report referenced a file outside this changeset

Important Files Changed

Filename Overview
litellm/integrations/otel/langfuse_logger.py Restricts Langfuse trace-name and root-content attributes to recording roots owned by the Langfuse provider
litellm/integrations/otel/logger.py Adds generic collector ownership classification and makes canonical provider selection independent of callback order
litellm/litellm_core_utils/litellm_logging.py Separates preset construction from generic OTEL export and preserves OTEL callback settings when reusing mapper-only loggers
litellm/integrations/weave/weave_otel.py Splits side-effect-free Weave configuration reads from the legacy helper that publishes OTEL environment variables
tests/test_litellm/integrations/otel/test_otel_v2_destinations.py Adds broad regression coverage for destination ownership, callback order, configuration reuse, and Phoenix initialization
tests/test_litellm/integrations/otel/test_otel_v2_multibackend.py Verifies generic and vendor exporters remain isolated across multibackend tracing flows

Reviews (8): Last reviewed commit: "fix(otel v2): otel entry auto-initialize..." | Re-trigger Greptile

greptile-apps[bot]

This comment was marked as resolved.

@codecov

codecov Bot commented Sep 12, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

…rter filter

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Devin Review: No Issues Found

Devin Review analyzed this PR and found no bugs or issues to report.

Devin Review

Presets (langfuse_otel, arize, phoenix, weave_otel, newrelic, agentops, levo) no longer fold the operator's OTEL_* exporter into their own exporter list, so vendor-mapped spans reach only the vendor backend. The otel callback alone serves OTEL_*, is the only logger a second otel entry reuses, and takes the proxy global slot over a preset whatever the callback order. The Langfuse logger stamps trace name and content on the request root only when its own provider created that root. The langfuse.* attribute filter is removed since no exporter can receive another callback's vocabulary any more. Credentialless presets keep only their header-gated exporter instead of inheriting the console placeholder, and the v2 Weave preset reads its config without publishing it as process-wide OTEL_* env

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration devin-ai-integration Bot changed the title fix(otel v2): keep langfuse.* span attributes off exporters that are not Langfuse sinks fix(otel v2): a callback exports only to the destination it owns, so vendor spans stay off the generic OTEL_* collector Sep 13, 2026
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

Re Greptile fan-out finding: fan-out mirrors the global provider's spans by design and did so on base too. Per-destination remapping is out of scope, noted as a caveat

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@yucheng-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/litellm_core_utils/litellm_logging.py Outdated
… the collector

callbacks: [langtrace, otel] built a second generic logger on the same OTEL_*
collector, so every chat span exported twice. Generic reuse, proxy slot and
global provider selection now key off exporter ownership instead of the
callback name, so Langtrace (collector plus mapper) counts as the generic logger

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

Comment thread litellm/litellm_core_utils/litellm_logging.py
…tings.otel

With callbacks [langtrace, otel], the otel entry reuses the Langtrace logger because it already
serves the operator's OTEL_* collector. That logger was built from the bare preset, so
callback_settings.otel (service_name, message_logging, metrics, events, baggage) was dropped.
A preset whose exporters are all operator-owned now loads the otel settings when it is built,
so whichever entry builds the collector logger first, the operator's settings apply

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/litellm_core_utils/litellm_logging.py
… a mapper-only logger

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit b565a8d. Configure here.

@devin-ai-integration
devin-ai-integration Bot changed the base branch from litellm_internal_staging to main September 13, 2026 04:24
@yuneng-berri
yuneng-berri deleted the branch main September 13, 2026 04:52
@mateo-berri mateo-berri reopened this Sep 13, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants