Repository navigation
fix(codex): sanitize Responses replay state - #1868
diegosouzapw merged 1 commit into
Conversation
There was a problem hiding this comment.
Code Review
This pull request updates the CLI fingerprint configuration by expanding the bodyFieldOrder and introduces a sanitization mechanism to filter out internal assistant commentary from conversation history. It also refines SSE stream processing to ensure only visible text deltas are accumulated for assistant content in logs. Additionally, the test suite is expanded with new integration and unit tests covering OAuth fingerprinting and commentary filtering. I have no feedback to provide.
|
Thanks @dhaern for this excellent contribution! 🎉 The replay sanitization and tests are incredibly thorough and safe. It has been successfully merged into the |
|
@dhaern @diegosouzapw why this is implemented and merged? Agent forgot what he says and says the same comments again and again: Maybe we should revert it? No other routers does it AFAIK. |
Im gonna check this but this is a simple sanitizer because last fixes added a regression where many internal messages from codex (tested with gpt 5.5 high and xhigh) were escaping and showing during the responses. What platform are you using? And what model? Anyways next time show proofs, logs and more info before accusing without have a single idea. This PR is mandatory because model was escaping useless output text. |
|
I checked this again @diegosouzapw @AveryanAlex The accusation still doesn’t prove that this PR caused the repeated planning text. #1868 is a narrow sanitizer for Codex Responses replay state. It only drops assistant replay items that explicitly carry The original issue was real: internal Codex/OpenCode commentary frames were being stored and sent back upstream in later The snippets you posted look like normal visible assistant/planning output, not evidence that sanitized So no, this is not a reason to revert. If you can provide actual payload evidence, I’ll check it. Otherwise this looks like a separate model/prompt/replay-loop issue, not this sanitizer. |
|
Hello @dhaern, sorry for the delay. I don’t think Evidence:
I also tested this by reverting the commentary stripping and am currently running a patched version with this commit: AveryanAlex@436f2d0. It fixes the issue where the assistant sends near-identical commentary/progress messages repeatedly. So “reverting would reintroduce a leak” is not accurate. The leak would be showing commentary as final visible output, not sending commentary back to the model with its phase. Stripping commentary from replay may actually explain repeated preamble/progress text like reported here: #1868 (comment) |
|
Thanks @AveryanAlex, after re-checking I agree that stripping However, this is already fixed in current The remaining behavior only prevents non- |
Integrated into release/v3.7.8
Integrated into release/v3.7.8
Integrated into release/v3.7.8
Summary
phase: \"commentary\"are not stored and re-injected into later/responsesrequests.response.output_text.deltaas visible assistant content for logs/replay payloads.Why
Recent Codex Responses fixes improved several real edge cases, but together they exposed a gap in replay hygiene:
response.completed.response.outputfrom observed output items./responsesclients./responses/compactJSON-only endpoint.That replay state should contain durable conversation/tool state, not local runtime phases. If a client sends assistant-side internal frames such as
phase: \"commentary\", storing and replaying them can leak hidden working notes back into the next Codex request. This PR keeps the replay behavior from #1750/#1791, but filters non-final assistant phases at the state boundary.While adding a full pipeline regression for this, the Codex CLI fingerprint also showed a real mismatch: the
codexprofile still prioritized chat-completions fields likemessages, while the active Codex OAuth path uses Responses fields such asinput,instructions,store,reasoning,prompt_cache_key, andclient_metadata.Changes
commentary, analysis-like runtime frames, etc.);type;deltavalues still count toward fallback usage estimates;response.output_text.deltais accumulated as visible assistant content.anytest helpers.Validation
node --import tsx/esm --test tests/unit/executor-codex.test.ts tests/unit/stream-utilities.test.ts tests/integration/chat-pipeline.test.ts --test-name-pattern 'Codex|codex|reasoning deltas|passthrough'npm run typecheck:corenpx eslint --quiet open-sse/config/cliFingerprints.ts open-sse/utils/stream.ts open-sse/services/responsesToolCallState.ts tests/unit/executor-codex.test.ts tests/unit/stream-utilities.test.ts tests/integration/chat-pipeline.test.tsgit diff --checkCompatibility
This keeps the replay repairs from #1750 and #1791 in place. The sanitizer only removes assistant messages that explicitly declare a non-final phase, so normal visible assistant messages and tool state continue to replay.
/responses/compactremains JSON-only as handled by #1777; this PR targets the normal Codex/responsesreplay/fingerprint path.