Skip to content

fix(responses): keep provider response headers in streaming logging callbacks - #38131

Merged
yucheng-berri merged 2 commits into
litellm_internal_stagingfrom
litellm_responses_streaming_provider_headers
Sep 3, 2026
Merged

yucheng-berri merged 2 commits into
litellm_internal_stagingfrom
litellm_responses_streaming_provider_headers

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Aug 24, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Streaming /v1/responses logs carry no provider response headers
  • Callbacks cannot read the provider request id for a stream
  • Non-streaming responses already log those headers

How it solves it:

  • Put the provider headers back on the copy sent to logging
  • Prefer headers the provider transform set, else the stream's own
  • Leave the event the caller iterates untouched

User Flow

Before: someone streaming responses cannot find the provider's own request id for that call in their logging backend

  1. They send POST https://litellm-domain/v1/responses with "stream": true
  2. The call returns 200 and the SSE stream reads back normally
  3. They open the log for that call in their logging backend
  4. The log has no provider headers at all, so there is no provider request id to hand to the provider's support
  5. They send the same request with "stream": false and that log does carry the provider headers, so only streaming is affected

After: the same streaming call logs the provider headers, so streaming and non-streaming logs agree

  1. They send the same POST https://litellm-domain/v1/responses with "stream": true
  2. The call returns 200 and the SSE stream reads back unchanged
  3. They open the log for that call in their logging backend
  4. The log now carries the provider headers, including the provider request id and, on Azure, apim-request-id and the serving region
  5. The "stream": false log is unchanged

Relevant issues

Linear ticket

Resolves LIT-6055

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Screenshots / Proof of Fix

Shared setup for every case below. Same config, same commands, only the checked out commit changes.

config.yaml

model_list:
  - model_name: gpt-4.1-mini
    litellm_params:
      model: openai/gpt-4.1-mini
      api_key: os.environ/OPENAI_API_KEY
  - model_name: azure-stub
    litellm_params:
      model: azure/gpt-4.1-mini
      api_base: os.environ/AZURE_STUB_BASE
      api_key: fake-azure-key
      api_version: "2025-04-01-preview"
litellm_settings:
  callbacks: ["datadog"]
general_settings:
  master_key: sk-1234
python litellm/proxy/proxy_cli.py --config config.yaml --port 4000

The gpt-4.1-mini cases are real, paid OpenAI calls. The azure-stub case points at a local Azure-shaped upstream that answers /openai/v1/responses with Azure's apim-request-id and x-ms-region headers, because no Azure OpenAI deployment was reachable with the credentials on hand. Every response is read back out of Datadog through its logs search API, so what you see is what a logging integration actually received.

Before (31ca4dd)

OpenAI streaming, the regression

  1. Send the streaming request and read the log back
$ curl -sS http://localhost:4000/v1/responses -H "Authorization: Bearer sk-1234" -H 'Content-Type: application/json' \
    -d '{"model":"gpt-4.1-mini","input":"lit6055q1788334589streamtrue","max_output_tokens":24,"stream":true}' -o /dev/null -w '%{http_code}\n'
200

$ curl -sS "https://api.$DD_SITE/api/v2/logs/events/search" -H "DD-API-KEY: $DD_API_KEY" -H "DD-APPLICATION-KEY: $DD_APP_KEY" \
    -H 'Content-Type: application/json' -d '{"filter":{"query":"*:*lit6055q1788334589streamtrue*","from":"now-15m","to":"now"}}' \
    | jq '.data[0].attributes.attributes.hidden_params.additional_headers | if . == null then "MISSING" else {header_count: (keys|length), "llm_provider-x-request-id", "llm_provider-openai-processing-ms"} end'
"MISSING"

OpenAI non-streaming, the control

  1. Send the same request without streaming and read the log back
$ curl -sS http://localhost:4000/v1/responses -H "Authorization: Bearer sk-1234" -H 'Content-Type: application/json' \
    -d '{"model":"gpt-4.1-mini","input":"lit6055q1788334589streamfalse","max_output_tokens":24,"stream":false}' -o /dev/null -w '%{http_code}\n'
200

$ curl -sS "https://api.$DD_SITE/api/v2/logs/events/search" -H "DD-API-KEY: $DD_API_KEY" -H "DD-APPLICATION-KEY: $DD_APP_KEY" \
    -H 'Content-Type: application/json' -d '{"filter":{"query":"*:*lit6055q1788334589streamfalse*","from":"now-15m","to":"now"}}' \
    | jq '.data[0].attributes.attributes.hidden_params.additional_headers | if . == null then "MISSING" else {header_count: (keys|length), "llm_provider-x-request-id", "llm_provider-openai-processing-ms"} end'
{
  "header_count": 33,
  "llm_provider-x-request-id": "req_c7e91a849bfb451f9862ce3b87fdbe40",
  "llm_provider-openai-processing-ms": "1091"
}

Azure-shaped upstream, streaming

  1. Send the streaming request and look for apim-request-id
$ curl -sS http://localhost:4000/v1/responses -H "Authorization: Bearer sk-1234" -H 'Content-Type: application/json' \
    -d '{"model":"azure-stub","input":"lit6055az1788334914","stream":true}' -o /dev/null -w '%{http_code}\n'
200

$ curl -sS "https://api.$DD_SITE/api/v2/logs/events/search" -H "DD-API-KEY: $DD_API_KEY" -H "DD-APPLICATION-KEY: $DD_APP_KEY" \
    -H 'Content-Type: application/json' -d '{"filter":{"query":"*:*lit6055az1788334914*","from":"now-15m","to":"now"}}' \
    | jq '.data[0].attributes.attributes.hidden_params.additional_headers | if . == null then "MISSING" else {header_count: (keys|length), "llm_provider-apim-request-id", "llm_provider-x-ms-region"} end'
"MISSING"

After (d0eab54)

OpenAI streaming, the regression

  1. Send the streaming request and read the log back
$ curl -sS http://localhost:4000/v1/responses -H "Authorization: Bearer sk-1234" -H 'Content-Type: application/json' \
    -d '{"model":"gpt-4.1-mini","input":"lit6055v3d0eabstreamtrue","max_output_tokens":24,"stream":true}' -o /dev/null -w '%{http_code}\n'
200

$ curl -sS "https://api.$DD_SITE/api/v2/logs/events/search" -H "DD-API-KEY: $DD_API_KEY" -H "DD-APPLICATION-KEY: $DD_APP_KEY" \
    -H 'Content-Type: application/json' -d '{"filter":{"query":"*:*lit6055v3d0eabstreamtrue*","from":"now-15m","to":"now"}}' \
    | jq '.data[0].attributes.attributes.hidden_params.additional_headers | if . == null then "MISSING" else {header_count: (keys|length), "llm_provider-x-request-id", "llm_provider-openai-processing-ms"} end'
{
  "header_count": 20,
  "llm_provider-x-request-id": "req_886d2ec6faec49838f32b0c1ccdde529",
  "llm_provider-openai-processing-ms": "187"
}

OpenAI non-streaming, the control

  1. Send the same request without streaming and read the log back
$ curl -sS http://localhost:4000/v1/responses -H "Authorization: Bearer sk-1234" -H 'Content-Type: application/json' \
    -d '{"model":"gpt-4.1-mini","input":"lit6055v3d0eabstreamfalse","max_output_tokens":24,"stream":false}' -o /dev/null -w '%{http_code}\n'
200

$ curl -sS "https://api.$DD_SITE/api/v2/logs/events/search" -H "DD-API-KEY: $DD_API_KEY" -H "DD-APPLICATION-KEY: $DD_APP_KEY" \
    -H 'Content-Type: application/json' -d '{"filter":{"query":"*:*lit6055v3d0eabstreamfalse*","from":"now-15m","to":"now"}}' \
    | jq '.data[0].attributes.attributes.hidden_params.additional_headers | if . == null then "MISSING" else {header_count: (keys|length), "llm_provider-x-request-id", "llm_provider-openai-processing-ms"} end'
{
  "header_count": 33,
  "llm_provider-x-request-id": "req_dc9f488bbff64e048fa6d20be14c4925",
  "llm_provider-openai-processing-ms": "1033"
}

Azure-shaped upstream, streaming

  1. Send the streaming request and look for apim-request-id
$ curl -sS http://localhost:4000/v1/responses -H "Authorization: Bearer sk-1234" -H 'Content-Type: application/json' \
    -d '{"model":"azure-stub","input":"lit6055v3azd0eab","stream":true}' -o /dev/null -w '%{http_code}\n'
200

$ curl -sS "https://api.$DD_SITE/api/v2/logs/events/search" -H "DD-API-KEY: $DD_API_KEY" -H "DD-APPLICATION-KEY: $DD_APP_KEY" \
    -H 'Content-Type: application/json' -d '{"filter":{"query":"*:*lit6055v3azd0eab*","from":"now-15m","to":"now"}}' \
    | jq '.data[0].attributes.attributes.hidden_params.additional_headers | if . == null then "MISSING" else {header_count: (keys|length), "llm_provider-apim-request-id", "llm_provider-x-ms-region"} end'
{
  "header_count": 12,
  "llm_provider-apim-request-id": "azure-correlation-abc123",
  "llm_provider-x-ms-region": "East US 2"
}

Datadog, side by side

Five streaming /v1/responses calls on each side, counted in Datadog by whether the log carries the provider headers. Before is 5 logged and none carrying them, after is 5 logged and 5 carrying them

LIT-6055 Datadog before and after

Type

🐛 Bug Fix

Caveats (if any)

Medium

  • Streaming responses now honor an upstream gateway's cost header
    • fires only when the upstream is itself a LiteLLM proxy sending x-litellm-response-cost
    • non-streaming responses and streaming chat completions already behaved this way
    • measured against a stub sending 42.5: the logged cost goes from a locally computed 1.16e-05 to 42.5

Low

  • Prometheus remaining-request gauges now emit for streaming responses, matching the non-streaming path
  • On a cache replay the logged additional_headers is now {} where it used to be null
  • Client facing HTTP response headers are byte for byte identical before and after, names and values
  • The Azure leg ran against a local Azure-shaped upstream, since no Azure OpenAI deployment was reachable with the credentials on hand

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Credit to the original implementation on this branch by Devin, which found the same root cause. This revision rewrites history so every commit is signed for the CLA, restricts the restore to the two header keys so a chained gateway's other hidden params cannot ride along, copies by value so nothing aliases the dicts the proxy uses for the client's own headers, and skips the restore entirely when the logging copy fell back to the caller's event

ran /live-pr-risk and found no regressions/backward incompatible risks

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR. Add '(aside)' to your comment to have me ignore it.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

CLAassistant commented Aug 24, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@greptile-apps

greptile-apps Bot commented Aug 24, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

The PR preserves provider response headers on the logging copy used for streaming Responses API callbacks without modifying the event returned to callers

  • Captures an immutable copy of raw upstream response headers
  • Restores processed and raw headers after the logging response is reconstructed
  • Adds regression coverage for provider headers, transform metadata, copy isolation, and serialization fallback
  • Updates the implementation and tests to satisfy the staging lint requirements

Confidence Score: 5/5

The PR appears safe to merge

No blocking failure remains

Important Files Changed

Filename Overview
litellm/responses/streaming_iterator.py Restores provider headers only onto successfully copied logging responses and leaves the caller event untouched when copying fails
tests/test_litellm/responses/test_streaming_iterator.py Adds focused streaming callback tests covering restored headers, copy independence, transform precedence, and fallback behavior

Reviews (3): Last reviewed commit: "fix: satisfy LIT002 mutable-collection g..." | Re-trigger Greptile

Comment thread litellm/responses/streaming_iterator.py Outdated
copied: Final = self._copy_for_logging(completed)
inner_response: Final = getattr(copied, "response", None)
if isinstance(inner_response, ResponsesAPIResponse):
inner_response._hidden_params = dict(self._hidden_params) # mutable-ok: model_dump drops private attrs

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Fallback mutates returned response

When the logging copy raises during serialization or validation, _copy_for_logging returns the original event and this assignment adds logging-only metadata, including raw provider headers, to the same response object yielded to SDK callers.

submit_times = recording_executor.submit_times_for(logging_obj)
assert len(submit_times) == 1
assert submit_times[0] >= recorder.async_hook_finished

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 New test exceeds line limit

The new test declaration exceeds the repository's 120-character Python line limit, as does the httpx.Response construction on line 172, creating avoidable lint and formatting failures.

Context Used: CLAUDE.md (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

@codecov

codecov Bot commented Aug 24, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@codspeed

codspeed Bot commented Aug 24, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_responses_streaming_provider_headers (d0eab54) with litellm_internal_staging (86ca146)

Open in CodSpeed

@yucheng-berri
yucheng-berri force-pushed the litellm_responses_streaming_provider_headers branch from 9572b60 to 6520b7d Compare September 2, 2026 07:32
@yucheng-berri

Copy link
Copy Markdown
Contributor

@greptileai please review the current head 6520b7d, the branch was rewritten to a single commit

bugbot run

@yucheng-berri

Copy link
Copy Markdown
Contributor

bugbot run

…allbacks

The responses streaming iterator captures the provider's HTTP response headers into
its own _hidden_params, but never puts them on the completed response, and the
model_validate(model_dump()) copy made for logging drops pydantic private attributes.
Success callbacks and StandardLoggingPayload.hidden_params.additional_headers therefore
saw an empty dict for streaming /v1/responses, so Azure's apim-request-id was unreadable
from the callback payload.

Restore the headers on the nested response of the logging copy, preferring any the
provider transform already set (the fake_stream path) and falling back to the ones the
iterator captured from the stream. Skipped when the copy fell back to the original event,
so a serialization failure never leaves logging-only state on the caller's object.
@yucheng-berri
yucheng-berri force-pushed the litellm_responses_streaming_provider_headers branch from 6520b7d to 7a1f2db Compare September 2, 2026 21:26

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 6520b7d. Configure here.

@yucheng-berri

Copy link
Copy Markdown
Contributor

@greptileai please review the current head d0eab54, it fixes the LIT002 lint gate failure from the staging rebase

@yucheng-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit d0eab54. Configure here.

@yucheng-berri
yucheng-berri merged commit c19d49d into litellm_internal_staging Sep 3, 2026
82 checks passed
@yucheng-berri
yucheng-berri deleted the litellm_responses_streaming_provider_headers branch September 3, 2026 00:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants