Skip to content

fix(router): strip encrypted reasoning the pinned deployment cannot decrypt - #43781

Merged
yassin-berriai merged 1 commit into
mainfrom
litellm_encrypted_affinity_mixed_origin_strip
Sep 30, 2026
Merged

yassin-berriai merged 1 commit into
mainfrom
litellm_encrypted_affinity_mixed_origin_strip

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Codex switching OpenAI, then Azure, then OpenAI fails with invalid_encrypted_content
  • Affinity pins to the first marked reasoning item and forwards every other one unchanged
  • The pinned deployment cannot decrypt reasoning minted by the other provider or resource

How it solves it:

  • After pinning, strip only reasoning the pinned deployment cannot decrypt
  • Keep reasoning from the pinned deployment and from deployments on its encryption boundary
  • Same filter on Responses input and on Anthropic /v1/messages history

User Flow

Before: a Codex user switching models mid-session gets a hard 400 on the third turn

  1. Codex (store=false) sends POST http://localhost:4000/v1/responses with model gpt-openai, gets 391
  2. They switch to gpt-azure with /model, ask "add 5", get 396
  3. They switch back to gpt-openai, ask "double it"
  4. The proxy returns 400 The encrypted content for item rs_... could not be verified (invalid_encrypted_content)
  5. https://localhost:4000/ui/?page=logs shows that request as Failure

After: the same session keeps working across both switches

  1. Same first turn, 391
  2. Same switch to gpt-azure, 396
  3. Same switch back to gpt-openai, "double it"
  4. The proxy returns 200 with 792
  5. The Logs page shows every turn as Success

Linear ticket

Resolves LIT-9014

Files changed

File Change
litellm/router_utils/pre_call_checks/encrypted_content_affinity_check.py After the model-id pin and the encryption-boundary pin, strips encrypted reasoning whose origin deployment is neither a target nor on a target's encryption boundary; factors the per-item and per-block origin lookup out of the existing extractors
litellm/responses/utils.py strip_encrypted_reasoning_from_input takes an optional keyword should_strip predicate; unset keeps the existing strip-all behavior
litellm/litellm_core_utils/prompt_templates/common_utils.py strip_encrypted_reasoning_from_messages takes the same optional predicate for Anthropic content blocks
flowchart LR
  A[follow-up request with mixed-origin reasoning] --> B[EncryptedContentAffinityCheck.async_filter_deployments]
  B -->|first marker pins deployment X| C[strip items whose origin is not X and not X's boundary]
  C --> D[upstream call to X]
  B -->|no pin possible| E[existing strip-all path, unchanged]
Loading

The filter runs inside the router pre-call check, only on the two pin branches. Requests with no marker, and the existing fall-through that strips everything when no deployment can serve the origin, are unchanged.

Pre-Submission checklist

  • I have added meaningful tests
  • The handful of test files covering my change pass locally
  • My PR passes all required CI/CD checks
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review

Screenshots / Proof of Fix

Setup: one proxy per arm, Postgres, PYTHONPATH set to that arm's checkout, same config. Real OpenAI and Azure OpenAI calls, no mocks.

model_list:
  - model_name: gpt-openai      # openai/gpt-5.4-mini, id dep-openai-mini
  - model_name: gpt-azure       # azure/gpt-5.4-mini, id dep-azure-mini
  - model_name: gpt-azure-sol   # azure/gpt-5.6-sol, same Azure resource, id dep-azure-sol
router_settings:
  enable_pre_call_checks: true
  optional_pre_call_checks: [encrypted_content_affinity]

The scripted client replays the Codex request shape (store: false, include: ["reasoning.encrypted_content"], full history resent every turn). Prompts are byte-identical across arms: "What is 17x23?", "Now add 5", "Now double it".

Before (e7460f1, ancestor of this PR's merge base; no change to the touched files in between other than #43407)

Codex CLI interactive, OpenAI then Azure then OpenAI

  1. Drove codex interactively in tmux, /model switches between turns
  2. Turn 1 391, turn 2 396, turn 3 fails with invalid_encrypted_content

Codex pane before the fix, turn 3 fails with invalid_encrypted_content

Admin UI Logs

  1. Opened http://localhost:4014/ui/?page=logs after the run
  2. The turn-3 row is a Failure with the OpenAI invalid_encrypted_content error

Logs page before the fix
Failed request detail before the fix

/v1/responses scripted, stream and non-stream

  1. gpt-openai,gpt-azure,gpt-openai stream: turn 3 HTTP 400 OpenAIException ... The encrypted content for item rs_... could not be verified
  2. Same, non-stream: turn 3 HTTP 400, same error
  3. gpt-azure,gpt-openai,gpt-azure stream: turn 3 HTTP 400 AzureException ... could not be verified whenever turn 2 emitted a reasoning item
  4. Controls: same model 3 turns 200; same-resource gpt-azure,gpt-azure-sol,gpt-azure 200

After (08cdc51; 9b13ad2 changes only the test file and rebases onto current main)

Codex CLI interactive, OpenAI then Azure then OpenAI

  1. Same interactive session shape against the fixed proxy
  2. Turn 1 391, turn 2 396, turn 3 792

Codex pane after the fix, all three turns succeed

Admin UI Logs

  1. Ran A-B-A then opened http://127.0.0.1:4013/ui/?page=logs
  2. The three new rows (OpenAI, Azure, OpenAI) are Success; the Failure rows below them are the dashboard's own 401 polling from a stale browser session, not LLM calls

Logs page after the fix
Turn 3 detail after the fix, openai/gpt-5.4-mini on dep-openai-mini, Success

/v1/responses and /v1/messages matrix

Cell Before turn 3 After turn 3
responses stream, OpenAI-Azure-OpenAI 400 invalid_encrypted_content 200 792, TTFD 0.8-1.8s
responses non-stream, OpenAI-Azure-OpenAI 400 invalid_encrypted_content 200 792
responses stream, Azure-OpenAI-Azure 400 when turn 2 emits reasoning 200 792
same-resource Azure-AzureSol-Azure, effort high 200 200 792, Azure marker kept
same model 3 turns 200 200, origin marker kept
messages, OpenAI-Azure-OpenAI 200 200 792

3 reps each of both A-B-A orders on the fixed proxy: 18/18 turns HTTP 200.

Tests

  1. tests/unit/router_utils/pre_call_checks/test_encrypted_content_affinity_check.py: 57 passed, including a real Router test that drives async_get_available_deployment and asserts the selected deployment plus the request input it forwards
  2. tests/unit/litellm_core_utils/prompt_templates/test_litellm_core_utils_prompt_templates_common_utils.py: 183 passed
  3. Mutation check at 08cdc51, 7/7 mutants killed (M1 and M4 re-run against the rewritten real-router test at 9b13ad2, both killed):
    • M1 no strip on the model-id pin: real-router, Anthropic-origin, unknown-origin tests fail
    • M2 no strip on the boundary pin: boundary-pin test fails
    • M3 ignore encryption boundaries: same-boundary and boundary-pin tests fail
    • M4 treat every origin as decryptable: 4 tests fail
    • M5 keep reasoning from unknown origins: unknown-origin test fails
    • M6 Responses helper ignores the predicate: 5 tests fail
    • M7 Anthropic helper ignores the predicate: Anthropic-origin and common-utils predicate tests fail

Taxonomy audit

  • F3 (sibling surfaces): Responses input and Anthropic /v1/messages both filtered; compact routes through the same check
  • F3 not applicable: Responses WebSocket frames after the first skip router selection entirely today, pre-existing and out of scope
  • F3 not applicable: DeploymentAffinityCheck pins without encrypted-content affinity configured never had marker handling; when both are on, this check runs first
  • W1: mutates the request list in place on purpose, matching the existing strip helpers (the router's fallback snapshot shares the list)
  • X1: predicate is None-checked, not falsy-checked
  • H2/H4: no new comments, no Optional, no bare dict/Any; test lines under 120 chars

Type

🐛 Bug Fix
✅ Test

Caveats (if any)

Low

  • Reasoning from an unknown or deleted origin deployment is stripped, its summary kept
  • WebSocket follow-up frames are not re-filtered, same as before this PR

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Link to Devin session: https://app.devin.ai/sessions/fbb7b80872e44f5889056bca094b3660
Open in Devin Desktop: https://app.devin.ai/desktop/session/fbb7b80872e44f5889056bca094b3660?variant=devin
Requested by: @yassin-berriai


Note

Medium Risk
Changes router pre-call request mutation for encrypted reasoning on affinity pin paths; incorrect boundary logic could strip needed blobs or leave undecryptable ones, but behavior is scoped to marked follow-up requests and is heavily tested.

Overview
Fixes multi-provider Codex sessions where encrypted-content affinity pinned a deployment but still forwarded reasoning blobs from other providers, causing invalid_encrypted_content on the next turn.

After a model-id or encryption-boundary pin, EncryptedContentAffinityCheck now runs _strip_reasoning_the_target_cannot_decrypt: it keeps encrypted reasoning only when the pinned target (or a deployment on the same (api_base, api_key) boundary) can decrypt it, and drops encrypted_content from foreign or unknown origins while preserving readable summaries. The same logic applies to Responses input and Anthropic messages history.

The shared strip helpers gain an optional should_strip predicate on strip_encrypted_reasoning_from_input and strip_encrypted_reasoning_from_messages; omitting it preserves the previous strip-all behavior (including the existing fall-through when no deployment can serve the origin). Origin lookup is refactored into _model_id_of_input_item / _model_id_of_anthropic_block.

Reviewed by Cursor Bugbot for commit c3683a3. Bugbot is set up for automated code reviews on this repo. Configure here.

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai review 08cdc51

@greptile-apps

greptile-apps Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[Medium risk] Adds conditional filtering to encrypted reasoning stripping.

The PR appears safe to merge based on the reviewed changes

Summary

The PR filters mixed-origin encrypted reasoning after deployment affinity pins, preserving reasoning the selected deployment can decrypt. The rebase onto main did not change the five PR files since the previous review

Reviews (4) · Last reviewed commit: "fix(router): strip encrypted reasoning t..."

Comment thread tests/unit/router_utils/pre_call_checks/test_encrypted_content_affinity_check.py Outdated
@codecov

codecov Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 96.15385% with 2 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
...re_call_checks/encrypted_content_affinity_check.py 95.55% 2 Missing ⚠️

📢 Thoughts on this report? Let us know!

@codspeed

codspeed Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_encrypted_affinity_mixed_origin_strip (c3683a3) with main (82d8b37)

Open in CodSpeed

@devin-ai-integration
devin-ai-integration Bot force-pushed the litellm_encrypted_affinity_mixed_origin_strip branch from 08cdc51 to 9b13ad2 Compare September 30, 2026 00:31
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai review 9b13ad2

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai full review 9b13ad2

Review 2 was incremental across a rebase, so it counted main commits that landed in between (Redis post-call batching, MCP OAuth metadata, OpenAPI tool listing) as part of this PR. None of those are in this PR. git diff origin/main...9b13ad2 touches only encrypted_content_affinity_check.py, responses/utils.py, prompt_templates/common_utils.py and their two mapped test files. Please judge only that diff

…ecrypt

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration
devin-ai-integration Bot force-pushed the litellm_encrypted_affinity_mixed_origin_strip branch from 9b13ad2 to c3683a3 Compare September 30, 2026 16:34
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai review again please, rebased onto main at c3683a3

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit c3683a3. Configure here.

@yassin-berriai
yassin-berriai merged commit 2ed9761 into main Sep 30, 2026
100 of 103 checks passed
@yassin-berriai
yassin-berriai deleted the litellm_encrypted_affinity_mixed_origin_strip branch September 30, 2026 16:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants