Skip to content

feat: configure assistant reasoning history for compatible providers - #41050

Open
jibanez-staticduo wants to merge 38 commits into
BerriAI:mainfrom
jibanez-staticduo:litellm_forward_reasoning_content
Open

jibanez-staticduo wants to merge 38 commits into
BerriAI:mainfrom
jibanez-staticduo:litellm_forward_reasoning_content

Conversation

@jibanez-staticduo

@jibanez-staticduo jibanez-staticduo commented Sep 14, 2026 •

Copy link
Copy Markdown

TLDR

Problem this solves:

  • Compatible endpoints can require different historical reasoning field names

How it solves it:

  • Keep upstream string forwarding, allow explicit opt-out
  • Select reasoning explicitly with per-model reasoning_content_field
  • Preserve default behavior, caller messages, and cache isolation

User Flow

Before: compatible requests retain the default field name

  1. Configure a compatible backend with reasoning_content_field: reasoning
  2. Send POST http://127.0.0.1:4000/v1/chat/completions, or /v1/responses with Chat Completions routing, containing assistant reasoning and two completed tools
  3. Inspect the backend request: both providers retain reasoning_content

After: configured requests carry a single canonical reasoning field

  1. Configure a compatible backend with reasoning_content_field: reasoning
  2. Send the same request and tool history
  3. Inspect the backend request: each supplied reasoning marker appears once under reasoning, with tool IDs, arguments, results, and ordering preserved

reasoning_content_field defaults to reasoning_content, meaning no additional normalization. With reasoning selected, a non-null target wins, including an empty string. Otherwise the source value is used when non-null. Only outgoing assistant copies lose reasoning_content. Omitting either transport option from a model update preserves its stored value; an explicit false value still disables forwarding. The model representation also accepts encrypted stored values, so subsequent management updates can read them. Invalid selectors fail with HTTP 400 at the Chat dispatch boundary before reaching the backend

Hosted forwarding follows current upstream: omitting the option or setting it true forwards supplied strings. Explicit false excludes assistant reasoning history. When normalization is selected with false, both field names are removed. Hosted drops non-string values even after normalization. Default and explicit true share a cache key; explicit false is isolated. The OpenAI adapter retains its existing forwarding behavior. Template preserve_thinking, client responses, native Responses, and signed thinking-block handling remain unchanged

The field selector improves compatibility and endpoint consistency. In vLLM revision 2a02f6efe319c885e3ccbcecde402e0028f9ec1e, Chat Completions already normalizes reasoning_content, while /tokenize does not. This extension does not claim demonstrated Chat Completions history loss at that backend layer or a performance improvement

Relevant issues

Product configuration and compatibility documentation: BerriAI/litellm-docs#1456

Current upstream restores default hosted string forwarding. This PR retains that behavior and adds a field selector and explicit exclusion

Affected release

Linear ticket

Pre-Submission checklist

  • I have added meaningful tests
  • The focused test files covering my change pass locally
  • My PR passes all required CI/CD checks
  • My PR's scope includes only its feature and necessary CI repairs
  • Current-tip Greptile confidence is at least 4/5

At 9f6bfb26ac5a3375cc2f1fe9a7ab60d70ed3bdb1, the branch includes upstream 0980f756bd031993329eb0b8b2caa193047e6465 and is mergeable. make check passes against that upstream snapshot, including Python lint and type budgets, test quality, E2E types, dashboard lint, and API-type synchronization. Coverage thresholds, lint budgets, and signature verification remain unchanged

Local validation passes 515 reasoning tests, 122 cache/OpenAI tests, and 32 router tests with eight existing skips. Model construction now uses model_validate with preprocessing preserved, including reserved-key filtering and retry conversion. A mutation that bypasses preprocessing fails the regression test; restoring it passes all 12 API-base tests. Two health-test failures also reproduce on clean upstream in the same environment

Current-tip codecov/patch and Veria pass. The remote proxy-behavior tests pass all 929 cases, then the Lens coverage upload fails because Codecov CLI 11.3.1 cannot import its GPG key and reports No public key, matching codecov-action#1876. An attempted workflow rerun was rejected because it requires repository-admin rights. The rust-test check is cancelled rather than passed. Required CI remains incomplete

Greptile's visible 2/5 belongs to 2a7abe8ca5, not this tip. Current-tip Greptile confirmation is unavailable, and Bugbot has no visible result. Maintainer review has not been requested

Screenshots / Proof of Fix

Fresh before/after captures against the merge base and this exact public tip are unavailable. Local regression checks and live deployment smoke checks do not replace that feature-specific evidence

Type

Bug Fix, New Feature

Caveats

Medium

  • Fresh real-provider before/after validation is unavailable for this tip
  • Required CI and current-tip Greptile/Bugbot validation remain incomplete

Final Attestation

  • All real-world workflows and edge cases are verified

@jibanez-staticduo
jibanez-staticduo requested a review from a team September 14, 2026 08:15
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 14, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-14T09:26:30.738771Z 8a82ea7 Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@codspeed

codspeed Bot commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing jibanez-staticduo:litellm_forward_reasoning_content (9f6bfb2) with main (0980f75)

Open in CodSpeed

@greptile-apps

greptile-apps Bot commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 2/5

[Medium risk] Adds reasoning content forwarding configuration for compatible LLM providers.

The PR does not appear safe to merge while omitted forwarding can disclose hosted reasoning history and reuse responses cached under the old behavior

Findings

  1. P1 Security Hosted reasoning forwards by default ▶
  2. P1 Old cache entries change meaning ▶

Summary

The PR adds configurable assistant-reasoning field normalization and Hosted vLLM forwarding, with transport-option propagation and cache identity changes

  • The current default forwards hosted reasoning even when forwarding was not enabled
  • The cache-key change can mix responses across the old and new defaults during an upgrade

Reviews (17) · Last reviewed commit: "test: remove duplicate catalog fields af..."

Comment thread litellm/llms/hosted_vllm/chat/reasoning_content.md Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 30bee8adc9

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread litellm/types/utils.py Outdated
@codecov

codecov Bot commented Sep 14, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@jibanez-staticduo

Copy link
Copy Markdown
Author

@greptileai Please re-review 8a82ea7: cache isolation is fixed, documentation moved to litellm-docs, and real-provider evidence now matches the current tip

@jibanez-staticduo

Copy link
Copy Markdown
Author

@codex review Please verify the cache-key correction in 8a82ea7, including preserved default keys, semantic scopes, and the new cache regressions

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Delightful!

Reviewed commit: 8a82ea72ac

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@jibanez-staticduo

Copy link
Copy Markdown
Author

The auth failure reproduces on unchanged upstream; #41055 fixes it and passes proxy-behavior. OSV flags unchanged mlflow. Details and evidence are documented above

@jibanez-staticduo jibanez-staticduo changed the title fix(hosted_vllm): opt in to forwarding assistant reasoning history feat: configure assistant reasoning history for compatible providers Sep 15, 2026
@jibanez-staticduo

Copy link
Copy Markdown
Author

@greptileai Please review the new per-model reasoning field selector, forwarding-policy interaction, cache isolation, and outgoing HTTP tests in this revision

@jibanez-staticduo

Copy link
Copy Markdown
Author

@greptileai Please recheck this revision: omitted model-update fields preserve stored reasoning settings, and the generated snapshot now matches the CI Python version

Comment thread litellm/llms/hosted_vllm/chat/transformation.py Outdated
@jibanez-staticduo

Copy link
Copy Markdown
Author

@greptileai Please review the final revision, including immutable hosted filtering, encrypted model-setting roundtrips, and preservation of omitted options during partial updates

Comment thread litellm/types/router.py Outdated
Comment thread litellm/llms/hosted_vllm/chat/transformation.py Outdated
@jibanez-staticduo

Copy link
Copy Markdown
Author

@greptileai Please recheck selector validation before Chat dispatch and typed hosted overrides. Both findings have regression coverage, including encrypted storage and native Responses

@jibanez-staticduo

Copy link
Copy Markdown
Author

@greptileai Please review the current tip 6716494e1f, which adds one commit on top of eac8c95cb9 that you already scored 5/5.

The extra commit touches only tests/unit/test_unit_shard_missing_paths.py: upstream #43186 moved that file with "UNIT_FLAG" repeated inside one env dict literal, which fails F601 under ruff-tests.toml and turns the lint job red for every pull request that touches tests, including main itself. Removing the duplicate leaves the effective environment identical; 3 passed locally and ruff check --config ruff-tests.toml on the file is clean.

CI on this tip is 94 successful and 1 skipped with no failures, and the reasoning-scope files plus the owned-inventory file are 1866 passed locally with 8 skipped.

Upstream moved tests/test_litellm core utils, routing, responses and caching
into tests/unit (BerriAI#43199), which collided with the reasoning-history tests.
The merge keeps both sides: reasoning normalization and forwarding still run
ahead of the configured system-message ordering in the synchronous and
asynchronous paths, and the coverage now lives at the upstream locations.

Verified against upstream/main 797fddf: 1539 passed in the relocated
hosted_vllm, openai chat and litellm_params files, 32 passed and 8 skipped in
the router forward_reasoning_content suite, 31 in tests/unit/caching, 308 in
the Responses-to-Chat bridge, 337 in model management endpoints.
ruff_strict_gate, type_discipline_gate and test_quality_gate pass with
--base upstream/main.
@jibanez-staticduo

Copy link
Copy Markdown
Author

@greptileai Please review 04ea8f4, which merges current upstream main. The upstream test relocation in #43199 moved the affected suites into tests/unit; reasoning normalization still runs ahead of the configured system-message ordering in both the synchronous and asynchronous paths, and the coverage now lives at the upstream locations. Relocated suites pass locally: 1539 passed across hosted_vllm chat, openai chat and litellm_params, 32 passed and 8 skipped in the router suite, 31 in caching, 308 in the Responses-to-Chat bridge, 337 in model management endpoints.

Upstream added 185 commits since the previous sync, including the owned
`Client` pool refactor (BerriAI#43245), the Responses lifecycle-event hold (BerriAI#43238)
and the driver preflight split (BerriAI#43259). One conflict, in the
`get_litellm_params` import block in `litellm/main.py`: upstream added
`InvalidControlOption`, `parse_control_options` and `with_control_options`
where this branch added `REASONING_TRANSPORT_KWARGS_KEYS`. Both stay, in the
order the import sorter wants.

Reasoning normalization still runs ahead of the configured system-message
ordering, and the net diff against upstream is the same 16 paths and 1005
added lines as before the sync.

Verified on this merge: 1574 passed across the relocated hosted_vllm, openai
chat and litellm_params files, 32 passed and 8 skipped in the router suite,
32 in caching, 308 in the Responses-to-Chat bridge, 337 in model management
endpoints. ruff_strict_gate, type_discipline_gate and test_quality_gate pass
with --base upstream/main.
…m models that use them

Declaring `forward_reasoning_content` and `reasoning_content_field` on
`GenericLiteLLMParams` reached every caller that builds those params from an
unannotated `**kwargs`, because basedpyright reports one unknown-argument
diagnostic per declared parameter at every such spread site: 113 sites, so the
two fields added 226 diagnostics and pushed `reportUnknownArgumentType` past its
ceiling without any of those files changing. `LiteLLM_Params` (deployment and
creation) and `updateLiteLLMParams` (model update) are the models that carry the
options, and neither is built by spreading unknown kwargs, so declaring them
there keeps the API, the generated dashboard types and the partial-update
behaviour identical while the shared signature stays as it was.
@jibanez-staticduo

Copy link
Copy Markdown
Author

@greptileai Please re-review 1e7d04a: upstream is merged through the client-pool refactor, and the two reasoning-transport options are now declared on LiteLLM_Params and updateLiteLLMParams instead of GenericLiteLLMParams. Nothing about the transport behaviour, the accepted keys, the generated dashboard types or the partial-update preservation changes — the move only stops 113 unrelated GenericLiteLLMParams(**kwargs) call sites from each reporting two more unknown-argument diagnostics, which was what pushed reportUnknownArgumentType past its ceiling. The locally measured gate, ruff-strict, type-discipline, test-quality and 1,543 focused tests are green.

Comment thread litellm/litellm_core_utils/get_litellm_params.py Outdated
@jibanez-staticduo

Copy link
Copy Markdown
Author

@greptileai Please review 82ef0b9: removed the flagged comments and fixed the shared catalog and Interactions contract tests without skipping validation

@jibanez-staticduo

Copy link
Copy Markdown
Author

@greptileai Please review the updated upstream merge, default forwarding alignment, explicit exclusion, cache identity, and normalized string validation regressions

) -> dict[str, object]: # mutable-ok: provider request contract
request_messages: Final = normalize_reasoning_content(
messages,
forward=litellm_params.get("forward_reasoning_content") is not False,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 security Hosted reasoning forwards by default

When forward_reasoning_content is omitted, this condition now sends supplied assistant reasoning to Hosted vLLM. Existing configurations can therefore forward history without opting in.

How this was verified: An omitted option enables forwarding, and the hosted request transform retains string-valued assistant reasoning_content.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Comment on lines +408 to +409
if forward_reasoning_content is False:
cache_key += "forward_reasoning_content: False"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Old cache entries change meaning

During a rolling upgrade with a shared cache, an unexpired response created when omitted forwarding sent no history can be returned for a request that now forwards history. Both use the same legacy cache key.

- batches _raw_batches_request now provides the ASGI scope real starlette
  requests carry, so the OTLP-trace branch of _read_request_body no longer
  collapses the body and the metadata 400 names the offending field again
- test_outer_deadline_delivers_session_termination gives the initialize/
  tools-call handshake its own 2s budget and keeps the original 0.2s
  fail_after as the real outer cancellation deadline; the in-flight task is
  cancelled and awaited in finally so session-termination DELETE is
  delivered before the caller resumes

Validated: batches 155 passed; test_mcp_client 409 passed; 14 cancel/teardown
variants green under 8-way CPU load.
… native binding

stubtest reports the stub parameter "sql" inconsistent with the pyo3
runtime parameter "query" (python-bridge routes/traces.rs), failing the
rust-wheel job. Python callers pass the argument positionally, so only
the stub name changes; upstream main still carries the stale name.
Resolve the OpenAPI compliance conflict by keeping the branch helpers
(_model_request_schema/_interaction_operation/_resolve_local_ref) with
their stronger required/readOnly/path-param validations, dropping the
now-unused upstream helpers _model_create_request_schema and
_interaction_resource_path, and preserving upstream's auto-merged
_declared_type_value fix. Inherit canonical GitPython/Tornado bumps
from upstream; uv.lock matches the snapshot exactly.
Whitespace-only normalization (blank-line collapse and line wrapping)
required by the ruff format gate after merging upstream 424bfd8; no
semantic or runtime change. Verified: 8 passed/13 skipped identical to
pre-format, ruff-tests and py_compile clean.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants