Skip to content

feat: attribute server LLM calls to the end user for observability - #2250

Open
mdecalf wants to merge 5 commits into
HolmesGPT:masterfrom
mdecalf:feat/observability-end-user-attribution
Open

mdecalf wants to merge 5 commits into
HolmesGPT:masterfrom
mdecalf:feat/observability-end-user-attribution

Conversation

@mdecalf

@mdecalf mdecalf commented Jun 30, 2026 •

Copy link
Copy Markdown

What & why

Closes #2249.

When Holmes runs as a server, LLM calls are not attributed to the end user or the conversation/session in any LLM observability backend. Traces show up only under the API key — no user, no session — so you can't answer "what did user X ask Holmes?" or "show me this conversation" from your observability tool.

This is a provider-neutral gap (same with OpenAI, Anthropic, Bedrock, ...). Holmes already knows the user and the conversation; the identity just never reached the LLM call. Complementary to #1969 (HolmesUsageEvents), which only feeds the hosted Robusta Supabase DAL — this path works for any LLM observability stack (Langfuse, Langsmith, Arize, Datadog LLM Obs, a LiteLLM proxy, ...).

How

Attribution is expressed with provider-neutral fields on the completion call:

  • user — the standard end-user identifier. It is understood across providers and is carried in the request body, so it reaches a remote model endpoint / proxy and is logged by any observability callback. This is the primary, always-effective attribution field.
  • metadata — optional observability fields (session_id, tags). It is a logging-only field consumed by the configured logging callback; it is never sent to the model. Useful for grouping a conversation into a session and tagging requests.

Changes:

  • LLM.completion() / DefaultLLM.completion() accept optional user and metadata and forward them. Explicit call arguments win over any statically-configured values in the model args (metadata is merged key-by-key, user replaced). Both are read non-destructively and excluded from the **self.args spread, so a reused DefaultLLM keeps its configured values across calls.

  • holmes/core/llm_observability.build_trace_attribution() (new) is the single, narrow place that maps request_context to attribution:

    • user ← user_email (preferred) or user_id
    • metadata.session_id ← conversation_id
    • metadata.tags ← request_type:<...>, cluster:<...>

    It's a whitelist by design — only known identity fields are mapped and tags are bounded — so nothing arbitrary from request_context leaks into traces. Returns an empty TraceAttribution when there's nothing to attribute.

  • call_stream() passes user + metadata from self._request_context.

  • server.py surfaces user_email and request_type into request_context.

Because attribution is keyed entirely off request_context, every entry point that builds one (chat, scheduled prompts, agui, health checks) benefits with no further changes.

Note on reach: user travels in the request body and reaches a downstream model endpoint / LiteLLM proxy, so it is logged wherever the trace is produced. metadata (session/tags) is consumed by the logging callback configured in the process making the call; when Holmes points at a remote proxy that owns the observability callback, forwarding of metadata depends on that proxy's configuration. user therefore remains the reliable cross-cutting attribution field.

Backwards compatibility

Default behaviour is unchanged when no identity is present (e.g. the CLI): user and metadata are None, so neither kwarg is sent to completion().

Tests

  • tests/core/test_llm_observability.py — the mapping helper (precedence, blanks/None ignored, stringify+strip, tag bounding, empty case).
  • tests/core/test_llm_completion_metadata.py — user/metadata forwarding, default no-op, per-call-over-configured merge, no-double-pass, and repeated-call persistence (configured values survive reuse).
46 passed

(the two new files plus the existing test_llm_completion_* suite, no regressions)

Summary by CodeRabbit

  • New Features
    • Added per-request user and metadata support for LLM completions and tool-driven LLM flows.
    • Introduced provider-neutral trace attribution derived from request context (user hashing, session metadata, and bounded tags).
    • Updated the chat API to pass through optional user_email and request_type for attribution.
  • Bug Fixes
    • Corrected attribution merging so per-call values override defaults and are forwarded without duplication.
  • Tests
    • Added coverage for metadata forwarding, trace-attribution rules, and normalization/length bounds.

HolmesGPT delegates every LLM call to LiteLLM, but `completion()` never
forwarded LiteLLM's reserved `metadata` field. As a result the observability
backend behind LiteLLM (Langfuse, Langsmith, Arize, Datadog LLM Obs, a LiteLLM
proxy, ...) cannot attribute a trace to the end user or group a conversation
into a session — every trace shows up only under the API-key alias.

The identity is already available: `/api/chat` puts `user_id`, `conversation_id`
and `cluster_name` into `request_context`, and `ChatRequest` carries
`user_email` / `request_type`. The only missing link was between
`request_context` and the `litellm.completion()` call.

This wires that link in a vendor-agnostic way:

- `LLM.completion()` / `DefaultLLM.completion()` accept an optional `metadata`
  dict and forward it to `litellm.completion()`. `metadata` is LiteLLM's
  reserved logging field: it reaches the configured callbacks / proxy and is
  stripped before the provider request, so it never reaches the model.
  Per-call metadata merges over any statically-configured metadata.
- `holmes/core/llm_observability.build_llm_metadata()` is the single, narrow
  place that maps `request_context` to LiteLLM's documented metadata keys
  (`trace_user_id`, `session_id`, `tags`). It is a whitelist by design: only
  known identity fields are mapped and tags are bounded, so nothing arbitrary
  leaks into traces. It returns `None` when there is nothing to attribute, so
  behaviour is unchanged for callers that carry no identity (e.g. the CLI).
- `call_stream()` passes `build_llm_metadata(self._request_context)` to
  `completion()`.
- `server.py` surfaces `user_email` and `request_type` into `request_context`
  so attribution prefers the email and tags carry the request classification.

Because attribution is keyed entirely off `request_context`, every entry point
that builds one (chat, scheduled prompts, agui, health checks) benefits without
further changes. Default behaviour is unchanged when no identity is present.

Complements HolmesGPT#1969 (HolmesUsageEvents), which only feeds the hosted Supabase DAL;
this works for any self-hosted LiteLLM observability backend.

Signed-off-by: mdecalf <maxime.decalf@ledger.fr>
@coderabbitai

coderabbitai Bot commented Jun 30, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 49ca512f-20e7-4915-a698-3bf59166a908

📥 Commits

Reviewing files that changed from the base of the PR and between 780d52a and 297ca01.

📒 Files selected for processing (2)
  • holmes/core/llm_observability.py
  • tests/core/test_llm_observability.py
🚧 Files skipped from review as they are similar to previous changes (1)
  • holmes/core/llm_observability.py

Walkthrough

Adds LiteLLM observability attribution for server-side LLM calls. A new helper maps request_context fields to trace attribution data, DefaultLLM.completion accepts and forwards per-call attribution, and server/tool-calling code supplies the needed request identity fields.

Changes

LLM Observability Attribution

Layer / File(s) Summary
Trace attribution model
holmes/core/llm_observability.py, tests/core/test_llm_observability.py
Adds TraceAttribution, _clean, and build_trace_attribution to derive user, session_id, and bounded tags from request_context, with tests for empty input, identity selection, normalization, tag bounds, and is_empty().
LLM completion attribution forwarding
holmes/core/llm.py, tests/core/test_llm_completion_metadata.py
LLM.completion and DefaultLLM.completion accept user and metadata, merge configured and per-call attribution, and forward the result to LiteLLM without duplicating kwargs. Tests cover forwarding, omission, merge precedence, and persistence across calls.
Server request context and tool calling wiring
server.py, holmes/core/tool_calling_llm.py
/api/chat adds user_email and request_type to request_context, and tool_calling_llm builds attribution from that context for each self.llm.completion(...) call.

Estimated code review effort: 3 (Moderate) | ~25 minutes

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 25.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly summarizes the main change: adding end-user attribution for server LLM calls.
Linked Issues check ✅ Passed The PR forwards request identity into LiteLLM user/metadata and session attribution, satisfying #2249's observability gap.
Out of Scope Changes check ✅ Passed The changes stay focused on attribution plumbing and tests; no unrelated features are introduced.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@netlify

netlify Bot commented Jun 30, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 297ca01
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/6a4683bace4748000847740f
😎 Deploy Preview https://deploy-preview-2250--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (1)
tests/core/test_llm_completion_metadata.py (1)

82-87: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Test gap (optional): repeated-call persistence not covered.

test_metadata_is_not_passed_twice calls completion once, so it passes even though configured metadata is popped and lost on a second call (see the llm.py finding). A second llm.completion(...) on the same instance asserting metadata is still forwarded would guard the regression.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/core/test_llm_completion_metadata.py` around lines 82 - 87, The current
test only covers a single completion call, so it misses the regression where
configured metadata is removed by the first pop in the LLM instance. Update
test_metadata_is_not_passed_twice to call llm.completion twice on the same llm
object and assert that mock_completion still receives the configured metadata on
the second call, using the _make_llm setup and completion method as the key
symbols to locate the change.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@holmes/core/llm.py`:
- Around line 763-772: The metadata handling in completion() is mutating
self.args by popping "metadata", which causes configured metadata to disappear
after the first call. Update completion() to read the configured metadata
non-destructively, merge it with per-call metadata as before, and exclude
"metadata" from the **self.args expansion so it is not passed twice; use the
completion() method and the self.args merge block as the places to adjust.

---

Nitpick comments:
In `@tests/core/test_llm_completion_metadata.py`:
- Around line 82-87: The current test only covers a single completion call, so
it misses the regression where configured metadata is removed by the first pop
in the LLM instance. Update test_metadata_is_not_passed_twice to call
llm.completion twice on the same llm object and assert that mock_completion
still receives the configured metadata on the second call, using the _make_llm
setup and completion method as the key symbols to locate the change.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 808da2eb-201a-404e-b6d8-462962886865

📥 Commits

Reviewing files that changed from the base of the PR and between 88c29e2 and ec8b5ef.

📒 Files selected for processing (6)
  • holmes/core/llm.py
  • holmes/core/llm_observability.py
  • holmes/core/tool_calling_llm.py
  • server.py
  • tests/core/test_llm_completion_metadata.py
  • tests/core/test_llm_observability.py

Comment thread holmes/core/llm.py Outdated
…w fixes

Make end-user attribution provider-neutral and actually effective regardless of
how Holmes reaches its model:

- `completion()` now forwards the standard `user` field (understood across
  providers and carried in the request body) in addition to `metadata`. `user`
  is what reliably reaches a remote LLM endpoint / proxy; `metadata` (session
  id, tags) is consumed by whichever logging callback is configured. Attribution
  no longer depends solely on the more vendor-specific `metadata` channel.
- `build_llm_metadata()` becomes `build_trace_attribution()`, returning a small
  `TraceAttribution(user, metadata)` value; the mapping helper is documented in
  provider-neutral terms (no implementation-specific wording).
- `call_stream()` passes both `user` and `metadata` from the request context.

Review fix (CodeRabbit): read `user`/`metadata` from `self.args` non-destructively
and exclude them from the `**self.args` spread, instead of `pop()`-ing them. A
reused `DefaultLLM` previously lost statically-configured `metadata` after the
first call (completion() runs once per call_stream iteration). Added a
repeated-call regression test plus `user` coverage.

Signed-off-by: mdecalf <maxime.decalf@ledger.fr>
@mdecalf

mdecalf commented Jun 30, 2026

Copy link
Copy Markdown
Author

Thanks @coderabbitai — both addressed in 780d52a:

  • self.args.pop("metadata") consuming configured metadata (major): completion() now reads user/metadata from self.args non-destructively and excludes them from the **self.args spread, so a reused DefaultLLM keeps its configured values across calls.
  • repeated-call test gap (nitpick): added test_configured_values_survive_repeated_calls, which calls completion() twice and asserts the configured values are still forwarded on the second call and left intact on the instance.

This revision also generalises the feature to forward the standard provider-neutral user field in addition to metadata, since user is what reliably reaches a remote model endpoint / proxy.

@coderabbitai

coderabbitai Bot commented Jun 30, 2026

Copy link
Copy Markdown
Contributor

Tip

For best results, initiate chat on the files or code changes.

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
cd /repo 2>/dev/null || true
git fetch origin pull/2250/head:pr2250 2>&1 | tail -5
git checkout pr2250 2>&1 | tail -5
git log --oneline -10
echo "---"
git show 780d52a --stat 2>&1 | head -30
echo "---"
sed -n '1,80p' holmes/core/llm.py

Length of output: 5637


🏁 Script executed:

#!/bin/bash
cd /repo 2>/dev/null || true
ast-grep run --pattern 'def completion($$$) {
  $$$
}' --lang python holmes/core/llm.py 2>/dev/null | head -100
echo "---direct grep---"
rg -n "def completion" -A 60 holmes/core/llm.py | sed -n '1,140p'
echo "---llm_observability---"
cat holmes/core/llm_observability.py
echo "---tool_calling_llm usage---"
rg -n "build_trace_attribution|build_llm_metadata|metadata=|user=" holmes/core/tool_calling_llm.py
echo "---tests---"
sed -n '1,200p' tests/core/test_llm_completion_metadata.py

Length of output: 14851


🏁 Script executed:

#!/bin/bash
cd /repo 2>/dev/null || true
sed -n '710,800p' holmes/core/llm.py

Length of output: 4954


This confirms the fix in 780d52a is correct and complete:

  • self.args.get("user") / self.args.get("metadata") are now read non-destructively (no more pop), so a reused DefaultLLM's configured values survive across multiple completion() calls.
  • The **self.args spread explicitly excludes user/metadata (handled separately via attribution_kwargs), so there's no duplicate-kwarg risk.
  • test_configured_values_survive_repeated_calls exercises exactly this by calling completion() twice and asserting both the forwarded kwargs and llm.args itself remain intact.
  • The generalization to also forward the provider-neutral user field (alongside metadata) is a solid call — user reliably propagates to the model endpoint/proxy itself, while metadata depends on a configured logging callback, so having both gives more robust attribution coverage.

This looks good to me — nice work, @mdecalf!

[resolve]

ヽ(•‿•)ノ

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@holmes/core/llm_observability.py`:
- Around line 51-64: Hash end-user identifiers before they are assigned in
build_trace_attribution() and forwarded by DefaultLLM.completion(), since
TraceAttribution.user currently carries raw user_email/user_id into
completion(...). Update the attribution path so TraceAttribution.user stores a
stable hashed value instead of the original identifier, while preserving the
existing fallback behavior when no identity is present.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 9542c793-11b7-4d2c-a25e-dad803dd7c1d

📥 Commits

Reviewing files that changed from the base of the PR and between ec8b5ef and 780d52a.

📒 Files selected for processing (5)
  • holmes/core/llm.py
  • holmes/core/llm_observability.py
  • holmes/core/tool_calling_llm.py
  • tests/core/test_llm_completion_metadata.py
  • tests/core/test_llm_observability.py

Comment thread holmes/core/llm_observability.py Outdated
moshemorad and others added 3 commits July 2, 2026 13:42
`user` is forwarded all the way to the model provider (DefaultLLM.completion()
passes it straight to litellm.completion()), so sending the raw email/user_id
would leak it to whichever provider serves the request. Hash it instead — the
same guidance OpenAI gives for this field ("we recommend hashing their
username or email to avoid sending us any identifying information").

The hash is a plain deterministic sha256, so a given user still maps to a
single, stable, filterable trace identity in observability backends. Only
`user` needs this: `metadata` (session id, tags) is a logging-only field
consumed by whichever process makes the call, never sent to the provider.

Addresses CodeRabbit review comment on HolmesGPT#2250.

Signed-off-by: mdecalf <maxime.decalf@ledger.fr>
…ribution' into feat/observability-end-user-attribution
@mdecalf

mdecalf commented Jul 2, 2026

Copy link
Copy Markdown
Author

Thanks @coderabbitai — good catch, fixed in 1938c9a.

TraceAttribution.user is now a stable sha256 hash of user_email/user_id instead of the raw value, since user is the field that actually reaches the model provider (DefaultLLM.completion() forwards it straight into litellm.completion()). This matches OpenAI's own guidance for this field (hash the identifier rather than send raw PII). The hash stays deterministic so a given user still maps to one stable, filterable trace identity.

metadata (session id, tags) is unaffected — it's a logging-only field consumed by whichever process makes the call and is never forwarded to the provider, so it doesn't carry the same exposure risk.

Added test_user_is_hashed_not_raw and test_user_hash_is_deterministic; existing tests updated to assert against the hash.

@coderabbitai

coderabbitai Bot commented Jul 2, 2026

Copy link
Copy Markdown
Contributor

Tip

For best results, initiate chat on the files or code changes.

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
cd /repo 2>/dev/null || true
git fetch origin pull/2250/head:pr2250 2>&1 | tail -5
git checkout pr2250 2>&1 | tail -5
git log --oneline -5
echo "---llm_observability---"
cat holmes/core/llm_observability.py
echo "---tests---"
sed -n '1,250p' tests/core/test_llm_observability.py

Length of output: 8178


@mdecalf, confirmed — 1938c9a looks correct and complete:

  • _hash_identifier() applies sha256 to the cleaned user_email/user_id before it's assigned to TraceAttribution.user, so only the deterministic hash reaches DefaultLLM.completion() → litellm.completion(), never the raw identifier.
  • metadata (session_id, tags) is left untouched, consistent with it being logging-only and never forwarded to the model provider.
  • The updated docstring in holmes/core/llm_observability.py clearly documents the rationale (OpenAI's guidance on hashing the user field vs. metadata being provider-agnostic).
  • test_user_is_hashed_not_raw and test_user_hash_is_deterministic directly cover the new behavior, and the existing tests were correctly updated to assert against the hash instead of the raw value.

Nice work — this closes out the observability attribution feature nicely.

[resolve]

ヽ(´▽`)/

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Server LLM calls are not attributed to the end user / session in observability backends

2 participants