Skip to content

fix(codex): avoid orphaned Responses message IDs - #97445

Closed
salehelsayed wants to merge 1 commit into
NousResearch:mainfrom
salehelsayed:fix/gpt56-reasoning-message-pair
Closed

salehelsayed wants to merge 1 commit into
NousResearch:mainfrom
salehelsayed:fix/gpt56-reasoning-message-pair

Conversation

@salehelsayed

@salehelsayed salehelsayed commented Aug 28, 2026 •

Copy link
Copy Markdown

What does this PR do?

Fixes #97442.
Related to #97427, which appears to report the same GPT-5.6 Responses replay failure.

Hermes intentionally strips replayed reasoning-item IDs when store=false because some Responses-compatible endpoints otherwise try to look them up server-side. A dependent native assistant msg_* ID can nevertheless survive, producing an invalid partial item graph such as:

reasoning{encrypted_content} -> message{id: msg_*}

GPT-5.6 rejects that request because the message identity depends on the stripped rs_* reasoning identity.

This refreshed implementation is rebased onto current main and enforces the invariant at the final provider-call boundary, after request middleware and Responses preflight:

  • if a native reasoning item starts an output group, dependent typed assistant message IDs are removed until a real group boundary;
  • a function_call does not prematurely close the reasoning group;
  • function_call_output and ordinary role messages close the group;
  • if encrypted reasoning replay has been disabled for recovery, all typed assistant message IDs are omitted because their original dependency can no longer be proven;
  • independent message IDs are preserved, so the existing cache optimization remains where it is demonstrably safe;
  • message content, status, phase, and tool-call pairing are unchanged.

Current-main refresh

The original PR head had become conflicted after substantial upstream refactoring. On 2026-09-11 the branch was rebuilt as a single focused commit on current main (ad03f20dd61919ca2135d6904e787a94284aacaf) rather than mechanically resolving the old diff.

Current changes:

  • agent/responses_replay_identity.py
    • Adds the dependency-closure guard for Responses replay identities.
  • agent/turn_api_call.py
    • Applies the guard at the final provider-call boundary after middleware and preflight.
  • tests/agent/test_responses_replay_identity.py
    • Covers linked reasoning/message replay, function-call grouping, tool-output boundaries, independent message IDs, recovery mode, and non-mutation of the original request.
  • contributors/emails/saleh.fekry@gmail.com
    • Adds the standard contributor mapping for salehelsayed.

Type of Change

  • 🐛 Bug fix (non-breaking)
  • ✅ Regression tests

Validation

The refreshed head is conflict-free and GitHub reports it as mergeable against current main.

Hosted CI, Docker, and Nix workflows were triggered for the refreshed head but currently show action_required, so they require maintainer approval before they can execute for this external-contributor PR.

Checklist

  • Based on current upstream main
  • Focused only on the Responses replay identity bug
  • Preserves independent msg_* IDs
  • Adds focused regression coverage
  • Does not change architecture, configuration, or user-facing behavior outside this fix
  • Hosted CI completed on the refreshed head (awaiting maintainer workflow approval)

@alt-glitch alt-glitch added type/bug Something isn't working P2 Medium — degraded but workaround exists comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint provider/openai OpenAI / Codex Responses API sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state labels Aug 28, 2026

@ehz0ah ehz0ah left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed exact head e3a6ba38287c. The change restores the Responses replay invariant without reintroducing store-false reasoning ID lookups: linked msg_* IDs are removed whenever their rs_* identity cannot be replayed, while message content, phase, status, and tool-call pairing remain intact. The same rule is applied to conversion, preflight, normalization, stale-replay pruning, and invalid-encrypted-content recovery.

I reproduced the invalid graph on current main and confirmed that this head produces the dependency-closed graph. I also ran the focused 10-file suite through scripts/run_tests.sh: 672 passed. Ruff, compileall, and git diff --check passed. I found no correctness blockers in the submitted implementation. The branch currently conflicts with main, so it will need a rebase before merge.

Enforce dependency closure at the final Responses provider-call boundary so stateless reasoning replay cannot send a native msg_* identity after its required rs_* identity has been stripped. Preserve independent message IDs and degrade all typed message IDs when encrypted reasoning replay is disabled. Add focused regression coverage.
@teknium1

Copy link
Copy Markdown
Collaborator

Thanks @salehelsayed — the work in this PR has landed on main via:

Your contribution is credited there (cherry-picked authorship / co-author trailer or credit in the PR body; see the linked PR for what was kept and what was trimmed). Closing this one as landed / superseded so the backlog reflects reality. If something in your original diff is still missing on current main, please comment and we'll reopen or follow up.

@teknium1 teknium1 closed this Sep 19, 2026
QuixThe2nd pushed a commit to QuixThe2nd/hermes-ide that referenced this pull request Sep 20, 2026
QuixThe2nd pushed a commit to QuixThe2nd/hermes-ide that referenced this pull request Sep 20, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint P2 Medium — degraded but workaround exists provider/openai OpenAI / Codex Responses API sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: GPT-5.6 Responses rejects replayed message after Hermes strips its reasoning ID

4 participants