Repository navigation
fix(router): strip encrypted reasoning the pinned deployment cannot decrypt - #43781
Conversation
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
|
|
|
@greptileai review 08cdc51 |
|
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
08cdc51 to
9b13ad2
Compare
|
@greptileai review 9b13ad2 |
|
@greptileai full review 9b13ad2 Review 2 was incremental across a rebase, so it counted main commits that landed in between (Redis post-call batching, MCP OAuth metadata, OpenAPI tool listing) as part of this PR. None of those are in this PR. |
…ecrypt Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
9b13ad2 to
c3683a3
Compare
|
@greptileai review again please, rebased onto main at c3683a3 |
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit c3683a3. Configure here.
TLDR
Problem this solves:
invalid_encrypted_contentHow it solves it:
inputand on Anthropic/v1/messageshistoryUser Flow
Before: a Codex user switching models mid-session gets a hard 400 on the third turn
gpt-openai, gets391gpt-azurewith/model, ask "add 5", get396gpt-openai, ask "double it"The encrypted content for item rs_... could not be verified(invalid_encrypted_content)After: the same session keeps working across both switches
391gpt-azure,396gpt-openai, "double it"792Linear ticket
Resolves LIT-9014
Files changed
litellm/router_utils/pre_call_checks/encrypted_content_affinity_check.pylitellm/responses/utils.pystrip_encrypted_reasoning_from_inputtakes an optional keywordshould_strippredicate; unset keeps the existing strip-all behaviorlitellm/litellm_core_utils/prompt_templates/common_utils.pystrip_encrypted_reasoning_from_messagestakes the same optional predicate for Anthropic content blocksThe filter runs inside the router pre-call check, only on the two pin branches. Requests with no marker, and the existing fall-through that strips everything when no deployment can serve the origin, are unchanged.
Pre-Submission checklist
Screenshots / Proof of Fix
Setup: one proxy per arm, Postgres,
PYTHONPATHset to that arm's checkout, same config. Real OpenAI and Azure OpenAI calls, no mocks.The scripted client replays the Codex request shape (
store: false,include: ["reasoning.encrypted_content"], full history resent every turn). Prompts are byte-identical across arms: "What is 17x23?", "Now add 5", "Now double it".Before (e7460f1, ancestor of this PR's merge base; no change to the touched files in between other than #43407)
Codex CLI interactive, OpenAI then Azure then OpenAI
codexinteractively in tmux,/modelswitches between turns391, turn 2396, turn 3 fails withinvalid_encrypted_contentAdmin UI Logs
invalid_encrypted_contenterror/v1/responses scripted, stream and non-stream
gpt-openai,gpt-azure,gpt-openaistream: turn 3 HTTP 400OpenAIException ... The encrypted content for item rs_... could not be verifiedgpt-azure,gpt-openai,gpt-azurestream: turn 3 HTTP 400AzureException ... could not be verifiedwhenever turn 2 emitted a reasoning itemgpt-azure,gpt-azure-sol,gpt-azure200After (08cdc51; 9b13ad2 changes only the test file and rebases onto current main)
Codex CLI interactive, OpenAI then Azure then OpenAI
391, turn 2396, turn 3792Admin UI Logs
/v1/responses and /v1/messages matrix
792, TTFD 0.8-1.8s792792792, Azure marker kept7923 reps each of both A-B-A orders on the fixed proxy: 18/18 turns HTTP 200.
Tests
tests/unit/router_utils/pre_call_checks/test_encrypted_content_affinity_check.py: 57 passed, including a realRoutertest that drivesasync_get_available_deploymentand asserts the selected deployment plus the requestinputit forwardstests/unit/litellm_core_utils/prompt_templates/test_litellm_core_utils_prompt_templates_common_utils.py: 183 passedTaxonomy audit
inputand Anthropic/v1/messagesboth filtered; compact routes through the same checkDeploymentAffinityCheckpins without encrypted-content affinity configured never had marker handling; when both are on, this check runs firstNone-checked, not falsy-checkedOptional, no baredict/Any; test lines under 120 charsType
🐛 Bug Fix
✅ Test
Caveats (if any)
Low
Final Attestation
Link to Devin session: https://app.devin.ai/sessions/fbb7b80872e44f5889056bca094b3660
Open in Devin Desktop: https://app.devin.ai/desktop/session/fbb7b80872e44f5889056bca094b3660?variant=devin
Requested by: @yassin-berriai
Note
Medium Risk
Changes router pre-call request mutation for encrypted reasoning on affinity pin paths; incorrect boundary logic could strip needed blobs or leave undecryptable ones, but behavior is scoped to marked follow-up requests and is heavily tested.
Overview
Fixes multi-provider Codex sessions where encrypted-content affinity pinned a deployment but still forwarded reasoning blobs from other providers, causing
invalid_encrypted_contenton the next turn.After a model-id or encryption-boundary pin,
EncryptedContentAffinityChecknow runs_strip_reasoning_the_target_cannot_decrypt: it keeps encrypted reasoning only when the pinned target (or a deployment on the same(api_base, api_key)boundary) can decrypt it, and dropsencrypted_contentfrom foreign or unknown origins while preserving readable summaries. The same logic applies to Responsesinputand Anthropicmessageshistory.The shared strip helpers gain an optional
should_strippredicate onstrip_encrypted_reasoning_from_inputandstrip_encrypted_reasoning_from_messages; omitting it preserves the previous strip-all behavior (including the existing fall-through when no deployment can serve the origin). Origin lookup is refactored into_model_id_of_input_item/_model_id_of_anthropic_block.Reviewed by Cursor Bugbot for commit c3683a3. Bugbot is set up for automated code reviews on this repo. Configure here.