Repository navigation
fix(otel v2): a callback exports only to the destination it owns, so vendor spans stay off the generic OTEL_* collector - #40881
Conversation
…not Langfuse sinks The Langfuse preset builds its exporter list as [*base.exporters, langfuse], and both the Langfuse mapper and the Langfuse logger's log_pre_api_call stamp langfuse.* attributes onto spans every exporter of the provider shares, so the operator's generic OTLP collector received Langfuse-only attributes whenever otel and langfuse_otel ran together. Wrap every exporter that is not a Langfuse sink in a processor that hands it a span view without langfuse.* attributes, and apply the same rule to per-request tenant destinations in the fan-out. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
🤖 Devin AI EngineerI'll be helping with this pull request! Here's what you should know: ✅ I will automatically:
Note: I can only respond to comments from users who have write access to this repository. ⚙️ Control Options:
|
Greptile SummaryThis PR separates OpenTelemetry v2 export ownership so vendor callbacks send spans only to their own destinations while the
Confidence Score: 5/5The PR appears safe to merge, with no outstanding previous findings or accepted new issues The current tip is identical to the previous review SHA, both previous threads are resolved, and the full-PR rule report referenced a file outside this changeset
|
| Filename | Overview |
|---|---|
| litellm/integrations/otel/langfuse_logger.py | Restricts Langfuse trace-name and root-content attributes to recording roots owned by the Langfuse provider |
| litellm/integrations/otel/logger.py | Adds generic collector ownership classification and makes canonical provider selection independent of callback order |
| litellm/litellm_core_utils/litellm_logging.py | Separates preset construction from generic OTEL export and preserves OTEL callback settings when reusing mapper-only loggers |
| litellm/integrations/weave/weave_otel.py | Splits side-effect-free Weave configuration reads from the legacy helper that publishes OTEL environment variables |
| tests/test_litellm/integrations/otel/test_otel_v2_destinations.py | Adds broad regression coverage for destination ownership, callback order, configuration reuse, and Phoenix initialization |
| tests/test_litellm/integrations/otel/test_otel_v2_multibackend.py | Verifies generic and vendor exporters remain isolated across multibackend tracing flows |
Reviews (8): Last reviewed commit: "fix(otel v2): otel entry auto-initialize..." | Re-trigger Greptile
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
…rter filter Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
bugbot run |
Presets (langfuse_otel, arize, phoenix, weave_otel, newrelic, agentops, levo) no longer fold the operator's OTEL_* exporter into their own exporter list, so vendor-mapped spans reach only the vendor backend. The otel callback alone serves OTEL_*, is the only logger a second otel entry reuses, and takes the proxy global slot over a preset whatever the callback order. The Langfuse logger stamps trace name and content on the request root only when its own provider created that root. The langfuse.* attribute filter is removed since no exporter can receive another callback's vocabulary any more. Credentialless presets keep only their header-gated exporter instead of inheriting the console placeholder, and the v2 Weave preset reads its config without publishing it as process-wide OTEL_* env Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
bugbot run |
|
Re Greptile fan-out finding: fan-out mirrors the global provider's spans by design and did so on base too. Per-destination remapping is out of scope, noted as a caveat |
|
bugbot run |
… the collector callbacks: [langtrace, otel] built a second generic logger on the same OTEL_* collector, so every chat span exported twice. Generic reuse, proxy slot and global provider selection now key off exporter ownership instead of the callback name, so Langtrace (collector plus mapper) counts as the generic logger Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…tings.otel With callbacks [langtrace, otel], the otel entry reuses the Langtrace logger because it already serves the operator's OTEL_* collector. That logger was built from the bare preset, so callback_settings.otel (service_name, message_logging, metrics, events, baggage) was dropped. A preset whose exporters are all operator-owned now loads the otel settings when it is built, so whichever entry builds the collector logger first, the operator's settings apply Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
bugbot run |
… a mapper-only logger Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit b565a8d. Configure here.
TLDR
Problem this solves:
OTEL_*collectorlangfuse.*, OpenInference) lands on a collector that never asked for it["langfuse_otel", "otel"]theotelentry reuses the Langfuse logger, so callback order decides what leaksOTEL_*env prints every span as JSON to proxy stdoutlangfuse_otelits own class, which surfaced it in the SDK pathHow it solves it:
OTEL_*oneotelalone servesOTEL_*; a generic collector plus a vendor now needs both incallbacksotelreuses only a logger whose every exporter is the operator's collector (a plainotellogger, or a mapper-only preset likelangtrace), and the published global provider prefers that oneotelrides a mapper-only preset built before it, that logger still getscallback_settings.oteland the Phoenix auto-init a loneotelentry performsOTEL_*envlangfuse.*prefix filter from the first revision of this PR is goneUser Flow
Before: an operator who ships traces to both a generic OTLP collector and Langfuse sees Langfuse-only attributes and duplicate spans in the generic collector
callbacks: ["otel", "langfuse_otel"],OTEL_EXPORTER_OTLP_ENDPOINTto their collector, andLANGFUSE_*to their Langfuse projectlangfuse_trace_name: private-langfuse-only-nameand gets HTTP 200private-langfuse-only-namewith achat proof-modelgeneration, as expectedlangfuse.trace.name=private-langfuse-only-name, the standardchat proof-modelspan, and a secondchat proof-modelspan carryinglangfuse.observation.input,langfuse.observation.outputand six morelangfuse.*keys["langfuse_otel", "otel"], the collector instead gets every span withlangfuse.*keys, and withcallbacks: ["langfuse_otel"]and noOTEL_*env the same spans are printed to the proxy's stdoutAfter: the generic collector only sees standard spans, Langfuse is unchanged
callbacks: ["otel", "langfuse_otel"],OTEL_EXPORTER_OTLP_ENDPOINTto their collector, andLANGFUSE_*to their Langfuse projectlangfuse_trace_name: private-langfuse-only-nameand gets HTTP 200private-langfuse-only-namewith achat proof-modelgeneration, as beforechat proof-modelspan, all with zerolangfuse.*keyscallbacks: ["langfuse_otel"]alone nothing is printed to stdout; an operator who wants the collector as well addsoteltocallbacksRelevant issues
Follow-up to #40793 and #28909
Linear ticket
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Live no-mock audit of every OTel v2 callback (
otel,langfuse_otel,arize,arize_phoenix,weave_otel,newrelic,agentops,levo,langtrace). Before is the merge base27f8d2a9baa1f81b37243878a17f7e55733a64e9, After is this tipb565a8d155f1dc2b19d208295c39bce0cce9f87a. Each leg runs from its own worktree with its own Postgres database and a 2-worker proxy (Before on 21081, After on 21080), started aspython litellm/proxy/proxy_cli.py --config cfg/<K>.yaml --port $PORT --num_workers 2 --use_v2_migration_resolverwithLITELLM_OTEL_V2=trueandLITELLM_DISABLE_NO_REDIS_WARNING=true. Providers are real OpenAI (openai/gpt-5.4-miniasproof-model) and real Anthropic (anthropic/claude-sonnet-4-6asproof-claude). Destinations are a real Jaeger all-in-one as the genericOTEL_EXPORTER_OTLP_ENDPOINT=http://127.0.0.1:21118collector (UI on 21186), real Langfuse cloud, real Weave, real Arize and New Relic OTLP ingest endpoints, and a real Arize Phoenix on 21106. Two destinations are disclosed OTLP/HTTP emulators because no account exists: Levo (21119) and the team-level Langfusecallback_varshost (21120). Every request is checked at every destination by itsgen_ai.response.id, never by count. Full per-cell tables (146 happy cells, sad, edge and 190 chaos requests) live in the audit report attached to the linked Devin sessionConfig for the all-callbacks cases (
cfg/K2.yaml); the other configurations change onlycallbacksand the env, as named in each stepRequest and readback shapes used by every step below.
$Kis the configuration,$LEGisbaseorhead, every stream is consumed to the end before the readback.gen=N[...]is the number of spans the generic collector holds for that response id and the vocabulary on each,rootis the vendor keys on the request root span at the collectorBefore (27f8d2a)
/v1/chat/completions, all eight v2 callbacks
K=K2chat curl, stream off: HTTP 200,chatcmpl-EO7dw6wP6ogaPNQnoJYXyLVOZ3KL3. Collector: gen=5[generic, langfuse, openinference, openinference, openinference+weave], rootlangfuse.trace.name. Langfuse 1 observation namedlitaudit-K2-base, Weave 4 duplicate calls, Phoenix 1, Levo 1, spend log 1chatcmpl-EO7dwRvaOrVL0u2BGavbxx5tgjZRV. Collector: gen=5 with the same four vendor duplicates, root leakedlangfuse_trace_name: audit-ui-base): HTTP 200,chatcmpl-EO9HBH6On35YrZDOB0BsN8xh0EGXR,x-litellm-call-id: eaf90459-63ca-4e74-ae03-5e411d53fdf9, Jaeger trace15b5d6a1ca7992b1af81875fb2003d91with 9 spans, fourchat proof-model,langfuse.observation.*,llm.*plusopeninference.span.kind,weave.*, andlangfuse.trace.name=audit-ui-baseon the root/v1/messages, all eight v2 callbacks
msg_011Cf3y8wDRXYyhQFMfktRXe. Collector gen=5[langfuse, generic, openinference, openinference, openinference+weave], root leakedmsg_011Cf3y91yuUY5mw2p6Bnsec. Collector gen=5, root leaked. Langfuse 1 named, Phoenix 1, Levo 1, spend log 1msg_011Cf477E73eTY9bqL4N5zqE,x-litellm-call-id: 32828cc4-27b3-4f6b-8023-54646d5095da/v1/responses, all eight v2 callbacks
resp_iyx7bFqbWthEBEhZxSQzzFFmqWyYNUf1P-OZNZi6_Vw. Collector gen=5, root leaked. Langfuse 1 named, Phoenix 1, Levo 1, spend log 1resp_FfmUL5RPf2GlJnYuRfTkF1gsvoRvQBhxd9poyOhZFAj. The streamed client id differs from the logged id and the collector was restarted (chaos C1) before a re-read, so this cell is NOT RUN at the collector; the LiteLLM SDK stream cells below cover streamed responsesx-litellm-call-id: d92d2d21-9025-4c9a-b623-d37eba372b34OpenAI, Anthropic and LiteLLM SDK clients through the proxy, all eight v2 callbacks
openaisync and asyncchat.completions.createandresponses.create, stream on and off,anthropicsync and asyncmessages.create, stream on and off,base_url=http://localhost:21081: 12 cells, every one HTTP 200, every collector readback gen=5 with vendor duplicates and root leakedlitellm.completion/acompletion/responses/aresponseswithmodel="openai/proof-model",api_base=http://localhost:21081, stream on and off: 8 cells, all HTTP 200, e.g.chatcmpl-EO7rHY8MVgjDTruzGHzG3YnfTY2tQgen=5[openinference, ...],resp_8tN4ZskAq-zr5s-QUvhc2lILlK0awQW-gGEZNVCqgNf(stream) gen=5[langfuse, ...]Other callback configurations, one 2-worker boot each, three endpoints each
["otel"],OTEL_*set:chatcmpl-EO7uuIthLSg0LTjvyqn7GbGojK0yFgen=1[generic], root clean["langfuse_otel", "arize", "arize_phoenix", "weave_otel", "newrelic", "levo"]withoutotel,OTEL_*set:chatcmpl-EO7wOeRQ8tHBaHypZupeVtsWWPTzagen=4[langfuse, openinference, openinference, openinference+weave], root leaked; the collector receives vendor spans it never asked for["langfuse_otel", "otel"]:chatcmpl-EO7yGZ7sKjWKSJeHuSKmHwDR0AjBXgen=1[langfuse], rootlangfuse.trace.name; theotelentry reused the Langfuse logger, Langfuse 1 namedlitaudit-K4-base["weave_otel", "otel"]:chatcmpl-EO7zugaamhNxAkpNlwH0p7Obyt4htgen=1[openinference+weave], Weave 1["langtrace", "otel"]plusPHOENIX_COLLECTOR_HTTP_ENDPOINT:chatcmpl-EO81KwwoiVUjFi5dx395bUKyIH4UBPhoenix 0 spans for all three ids, the auto-init is skipped behind the mapper-only callback["otel"]plusPHOENIX_*:chatcmpl-EO82ig93HSNfkJRoExxVwYi4qBAXbPhoenix 3 spans["langtrace"]:chatcmpl-EO843K7ztZpHbxpdq5crqlTaeSezegen=1[langtrace]["langfuse_otel"], noOTEL_*:chatcmpl-EO87R7XY5rdcsk4TMwR4TOvpd88hxLangfuse 1 namedlitaudit-K9-base, andgrep -c '"resource":' proxy.log= 12, every span printed as JSON to stdout["otel"],OTEL_EXPORTER=console:chatcmpl-EO892WX75MxXwpo1SQiTWs6R5KV1s, 12 span objects on stdout, spend log 1["otel", "langfuse_otel"],LANGFUSE_*unset:chatcmpl-EO8ATy5uM1fxVJTEFDfHCMYO4cSn5HTTP 200, gen=1[generic], Langfuse 0["otel", "langfuse_otel"]with a team whosemetadata.loggingnameslangfuse_otelandcallback_vars.langfuse_host=http://127.0.0.1:21120, keys on that team, on a team without settings, and with no team, then/team/updateand a 65 s wait for the metadata cache: Team Achatcmpl-EO8GNH0tzGQZwsx8w2xFBUNW0PJeQgen=2[generic, langfuse], root leaked on the collector and on the tenant root, team sink 1 generic span; after the updatechatcmpl-EO8Ki1obnYCh5FDjO4DzCXnKJfhV0gen=2[langfuse, generic], root leaked, team sink 1["otel", "agentops"], no AgentOps key:chatcmpl-EO8C1PnJZqfD20XAAhHfvvUsraTYEgen=2[generic, generic]callback_settings.otel.message_logging: false:chatcmpl-EO8OIHuLpMYwM4sXV5YOwxDBz8sgNgen=5 vendor duplicates as K2callbacks: ["nope"]thencallbacks: null: each worker exits at boot,ValueError: Empty module nameandAttributeError: 'NoneType' object has no attribute 'startswith'callbacks: [],OTEL_*set:chatcmpl-EO8Tx3v9frNFsCg1seRndO6Lp5qhkgen=1[generic] from the env-built providerSad paths, K2 configuration
"model":"no-such-model"on all three routes: HTTP 400Invalid model name passed in model=no-such-modelproof-modelpointed at a badOPENAI_API_KEY: HTTP 401AuthenticationError ... Incorrect API key provided, then HTTP 429No deployments available ... cooldownonce the router cooled the deployment down"max_tokens": 10000000: HTTP 400 on chat and messages, HTTP 200 on responses (provider accepted it)POST /key/healthon a key with Langfusecallback_vars: HTTP 200;GET /health/liveliness: HTTP 200 after every failure above; a valid key onproof-modelright after: HTTP 200Edge paths, K2 configuration
POST /team/updateon Team A while 12 Team A requests ran: all HTTP 200, zero new tracebacks in either worker logChaos, K2 configuration, 30 to 40 concurrent mixed requests (chat, messages, responses, stream on and off)
docker stopJaeger for 20 s mid-burst, restart within the exporter retry window: 30/30 callers HTTP 200, Levo got every id once, collector served post-restore traffic. Exactly-once landing of the outage-window spans is not provable, the in-memory Jaeger dropped its store on stopdocker pauseJaeger for 20 s mid-burst, unpause: 30/30 HTTP 200,/health/liveliness200 every 5 s, after unpause every id landed with gen=5 leaked, spend log 30/30kill -KILLone worker 4 s into a 30-request burst: 29/30 HTTP 200 plus 1 connection reset, master respawned a worker, 25/30 spend rows at readbackdocker stopa dedicated Postgres for 27 s mid-burst (control proxy on 21083): 29 HTTP 200 plus 1 HTTP 429 (provider cooldown), liveliness 200 throughout,/health/readinessreporteddb: disconnectedthen recovered, collector 28/30 ids gen=5 leaked, 26/30 spend rowsDashboard and destination UI
http://localhost:21081/ui, log in asadminwith the master key, open Logs (/ui/?page=logs): the audited ids are listedchatcmpl-EO9HBH6On35YrZDOB0BsN8xh0EGXRrow: detail page with request, response, tokens and costproof-model, send a message: the assistant reply rendershttp://localhost:21106and find the trace for the same id: Phoenix trace detail presentAfter (b565a8d)
/v1/chat/completions, all eight v2 callbacks
K=K2chat curl, stream off: HTTP 200,chatcmpl-EO7dwcOxIvfbuK6SjYTm7FJOWnjqt. Collector: gen=2[generic, langtrace], root clean. Langfuse 1 observation namedlitaudit-K2-head, Weave 1 call, Phoenix 1, Levo 1 (generic vocabulary), spend log 1chatcmpl-EO7dwfPqEVuGiZus1Axra5YkbsjUR. Collector: gen=2[generic, langtrace], root cleanlangfuse_trace_name: audit-ui-head): HTTP 200,chatcmpl-EO9H7ItlZZ1E3hAcTmHUuOk3cBrxz,x-litellm-call-id: 2f97ce6e-2490-4ca3-95e8-f8c77d3f824d, Jaeger trace371e55467b8f3d896a7757c811e2b802with 6 spans, one genericchat proof-modelplus the mapper-only Langtrace copy, zerolangfuse.*,openinference.*orweave.*keys, root clean/v1/messages, all eight v2 callbacks
msg_011Cf3y8wRpYeGTEaXnKCdr1. Collector gen=2[generic, langtrace], root cleanmsg_011Cf3y91oVTj1xxignTvhgh. Collector gen=2, root clean. Langfuse 1 named, Phoenix 1, Levo 1, spend log 1msg_011Cf476ujVLz71YzF8UYBJz,x-litellm-call-id: a5eb1638-8b80-4b39-b560-2b1280bd3efc, Jaeger trace5070fb5d5a591c8977e2d053d9d807a7, 6 spans, 0 vendor keys/v1/responses, all eight v2 callbacks
resp_ySsvIjjsBztPCyPKLPl6yZ7U1iPZKto6NjtK69fwDPd. Collector gen=2[generic, langtrace], root clean. Langfuse 1 named, Phoenix 1, Levo 1, spend log 1resp_afI0c6LFBbV-RlM5MVy7toZya4xw2veXBGHPjm2IIMd. NOT RUN at the collector for the same reason as Before; covered by the LiteLLM SDK stream cells belowx-litellm-call-id: 7495d2be-7b0b-476a-aeab-ed976039a5a9, Jaeger tracef6a40488b34e91dbdd502b7182bd3008, 6 spans, 0 vendor keysOpenAI, Anthropic and LiteLLM SDK clients through the proxy, all eight v2 callbacks
openaiandanthropiccells againsthttp://localhost:21080: every one HTTP 200, every collector readback gen=2[generic, langtrace], root cleanlitellm.*cells: all HTTP 200, e.g.chatcmpl-EO7rHiRh277TmeYXJThpYnuz2XWo3gen=2[generic, langtrace],resp_KRxlbk8nfZqSDKJkGyXB5Gxgx0E_WFzlvf0RECWfJBM(stream) gen=2[generic, langtrace],resp_pa9f8XtwUkIa9UrviJNxMV-h_BZevv4GgIc_h0SFoZt(async stream) gen=2Other callback configurations, one 2-worker boot each, three endpoints each
["otel"],OTEL_*set:chatcmpl-EO7uuuTvMGDAECG2J1Z6n26KwQaqagen=1[generic], root clean, unchangedotel,OTEL_*set:chatcmpl-EO7wOypCH9WFovofZi6DmRN0UWmcCgen=0, root clean, Phoenix 1 for every id. The collector gets nothing untilotelis added["langfuse_otel", "otel"]:chatcmpl-EO7yGuXnnAUcq9c4Iy3jwwYzMjm9rgen=1[generic], root clean, Langfuse 1 namedlitaudit-K4-head; callback order no longer matters["weave_otel", "otel"]:chatcmpl-EO7zuPj4fVhsbvhzhN59sCCOSLtabgen=1[generic], Weave 1["langtrace", "otel"]plusPHOENIX_*:chatcmpl-EO81KkemrqG0n5ZZOF8aBk4eo7A06Phoenix 3 spans, 2 and 2 for the messages and responses ids, matching a loneotel["otel"]plusPHOENIX_*:chatcmpl-EO82htlBdXhuE6j67TO6p7M9qZGyqPhoenix 3 spans, unchanged["langtrace"]:chatcmpl-EO843M89mTACc8nDUoH09qwvdrUaEgen=1[langtrace], unchanged["langfuse_otel"], noOTEL_*:chatcmpl-EO87R5haOpiQQOccIM5oDBFgrsDDoLangfuse 1 namedlitaudit-K9-head,grep -c '"resource":' proxy.log= 0["otel"],OTEL_EXPORTER=console:chatcmpl-EO892o401RHwawHIzcNM6cwvgY2az, 12 span objects on stdout, unchanged["otel", "langfuse_otel"],LANGFUSE_*unset:chatcmpl-EO8ATCjJFDFsgP817wkhOIAtMsNC7HTTP 200, gen=1[generic], Langfuse 0, unchangedchatcmpl-EO8GB3eKElZefmAJEuEMQWHt8JsDOgen=1[generic], root clean on the collector and the tenant root, team sink 1 generic span, Langfuse 1 namedlitaudit-K12-TEAM_A-head, spend log 1; team without settings and no-team keys the same; after the updatechatcmpl-EO8KiibJNnmkcQLyD2Lqa7ycjceKdgen=1[generic], team sink 1, root clean["otel", "agentops"], no AgentOps key:chatcmpl-EO8C19M0dvC13jpEdwkhL3HbHTXUAgen=1[generic], caller 200, AgentOps exporter 401 in the log onlycallback_settings.otel:chatcmpl-EO8OI8MEDfuqBmdhnVCHLHh8s9PuSgen=2[generic, langtrace], root clean["nope"]andnull: same boot failures as Before, unchanged[],OTEL_*set:chatcmpl-EO8Tw5VgnZoN3q91hT09udtICbktGgen=1[generic], unchangedSad paths, K2 configuration
Invalid model name passed in model=no-such-modelon all three routesAuthenticationError ... Incorrect API key provided, then HTTP 429 cooldown, same as Before"max_tokens": 10000000: HTTP 400 on chat and messages, HTTP 200 on responses, unchangedPOST /key/health200,GET /health/liveliness200 after every failure, valid key right after 200Edge paths, K2 configuration
POST /team/updateduring 12 Team A requests: all HTTP 200, zero new tracebacksChaos, K2 configuration, 30 to 40 concurrent mixed requests (chat, messages, responses, stream on and off)
db: disconnectedthen recovered, collector 24/30 ids gen=2 clean, 28/30 spend rowsDashboard and destination UI
http://localhost:21080/ui, log in asadminwith the master key, open Logs: the audited ids are listedchatcmpl-EO9H7ItlZZ1E3hAcTmHUuOk3cBrxzrow: detail page with request, response, tokens and costproof-model, send a message: the assistant reply renders/ui/?page=logging): all eight configured v2 callbacks listedhttp://localhost:21106and find the trace for the same id: Phoenix trace detail presentRecording of the UI flow on both legs, 1 fps preview (the full annotated mp4 is attached to the linked Devin session)
Not run on either leg: destination 404 (the organic vendor 403 covers a rejecting destination), Langfuse host on a closed port, direct
os.environinspection of a K5 worker (Weave isolation inferred from the clean generic span), Levo idempotence for the three repeated-request ids, the streamed/v1/responsescollector cells for raw curl and the OpenAI SDK (see above), AgentOps happy path (no key), Arize and New Relic readback (ingest-only keys), Langfuse cloud UI screenshot (no UI login)Type
🐛 Bug Fix
Caveats (if any)
Severe
callbacks: ["langfuse_otel"],["arize"], ...) plusOTEL_EXPORTER_OTLP_ENDPOINTno longer also exports to that collector; addoteltocallbacksto keep it (docs: docs(otel v2): presets export only to their own backend, add otel for a generic collector litellm-docs#1447)OTEL_EXPORTER=consoleunder a preset alone no longer prints spans to stdoutMedium
langfuse.observation.*and now also lack the rootlangfuse.trace.namethat leaked there beforeTenantFanOutSpanProcessordocstring). When a vendor preset is the only v2 callback, its provider is the global one, so a team destination on a different vendor receives that vendor's vocabulary, including the Langfuse root attributes. Base does the same; re-mapping fan-out copies per destination owner is a follow-upLow
LITELLM_OTEL_V2) OpenTelemetry and Langfuse loggers are untouched;get_weave_otel_configkeeps its v1 side effect of settingOTEL_*envPHOENIX_*env set,otellisted after a mapper-only preset now also auto-starts the Phoenix logger, as a loneotelentry always did; base skipped it on that ordercallbacks: ["otel", "langtrace"](otel first) still builds a second logger for the Langtrace mapper, so the collector gets twochatspans per request, one of them withlangtrace.*keys. Base does the same; a mapper-only preset reusing an already builtotellogger is a follow-upcallbacks: ["langtrace"]alone now honorscallback_settings.otel(service name, message logging, metrics, events, baggage, content capture), which base ignored for that configproxy-behaviorfails ontest_join_binds_the_membership_to_the_requested_team(TypeError: object MagicMock can't be used in 'await' expressionfromauth_checks.py). The same test fails the same way onmainitself at 30f33a9 (theproxy-behaviorjob on that commit), with none of this PR's changes, and this PR touches no auth codeFinal Attestation
main(30f33a9, which already contains the oldlitellm_internal_staginghistory via chore(ci): remerge internal staging #40943) passes the OTel, Weave, and Langfuse OTel suites (767 passed) and the["otel", "langfuse_otel"]live leg (generic sink 12 spans, 0langfuse.*keys; Langfuse sink 4 spans, trace names present), run before the audit aboveNote
Medium Risk
Changes default observability routing for multi-callback and preset-only setups (breaking if users relied on inherited OTEL_* exporters), but behavior is heavily tested and confined to OTel v2 integration plumbing.
Overview
OTel v2 callbacks no longer fan vendor traces into the operator’s generic
OTEL_*collector. Vendor presets (Langfuse, Arize, Phoenix, Weave, New Relic, etc.) now register only their own exporter instead of merging in*base.exporters, solangfuse.*/ OpenInference spans stay on the vendor backend. Theotelcallback alone owns the generic collector;select_global_otel_v2_loggerandproxy_server.open_telemetry_loggerprefer a logger whose exporters are all operator-owned (serves_generic_collector), and theotelentry reuses mapper-only presets (e.g. Langtrace) that already serve the collector.Langfuse stamps trace name and root I/O only when the request root belongs to its tracer provider (
_owned_recording_root). Weave v2 usesread_weave_otel_configso preset setup does not rewrite process-wideOTEL_*env (v1get_weave_otel_configkeeps the old side effect). Credential-gated exporter layering helpers are removed;_maybe_construct_otel_v2is simplified accordingly. Operators who want both destinations needcallbacks: ["otel", "<vendor>"].Reviewed by Cursor Bugbot for commit b565a8d. Bugbot is set up for automated code reviews on this repo. Configure here.
Link to Devin session: https://app.devin.ai/sessions/ef5e75844495467e92662e55750e1dd1
Open in Devin Desktop: https://app.devin.ai/desktop/session/ef5e75844495467e92662e55750e1dd1?variant=devin
Requested by: @yucheng-berri