Repository navigation
fix(agent): stop repaired sessions from repeatedly missing cache prefixes - #82925
Open
fangliquanflq wants to merge 3 commits into
Open
fangliquanflq wants to merge 3 commits into
fangliquanflq wants to merge 3 commits into
Conversation
13 tasks done
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Sessions with persisted alternation violations were repaired only in memory, so every resume could repair the same rows again and repeatedly lose the stable prompt-cache prefix. This change atomically reconciles repair results back to SessionDB while retaining removed rows as inactive audit records and preserving display-only verification candidates.
Symptom
After a malformed transcript is restored, consecutive requests repeatedly report the same message-sequence repairs. Repaired turns repeatedly fall back to an older prompt-cache prefix instead of reusing the stable repaired history.
Impact
Affected resumed sessions pay the repair and cache-miss cost on every subsequent turn. The reported trace showed the repaired requests reusing only about 30-38% of the prompt while adjacent clean requests reused 88-99%.
Bug Cause
Trigger:
hermes_state.py:8413/SessionDB.get_messages_as_conversation(..., repair_alternation=True)Causal chain:
repair_message_sequencefixes only the decoded in-memory message list and the restore path does not update the corresponding SessionDB rows.api_content, repeats the repair, and changes the replay prefix again.Why it is wrong: Live replay treats the repaired list as authoritative for the current request but leaves durable state inconsistent with that list, so the same defensive repair cannot converge across resumes.
Working sibling / contrast: Clean transcripts already reload byte-stably because no repair mutation separates the in-memory replay from the persisted active rows. Inspection paths with
repair_alternation=Falseintentionally remain read-only.Ruled out: Compression alone is not the cause. A disposable SessionDB seeded directly with malformed rows reproduced repair counts
[2, 2]across close and reopen before this fix.Fix
Repair-enabled active restores now retain internal row identity, atomically update changed survivor content, tool calls, reasoning content, and
api_content, and soft-archive rows removed from model replay. Existing FTS update triggers keep merged content searchable. Display-only verification candidates remain active, inactive audit loads do not reconcile, and internal row IDs are stripped unless requested by the caller.Related Issue
Closes #82863
Type of Change
Changes Made
hermes_state.py- reconcile identity-preserving replay repairs into active durable rows and archive repair-removed rows without deleting audit history.tests/hermes_state/test_restore_alternation_repair.py- cover load-repair-reload stability, stale sidecar clearing, FTS updates, archived rows, and read-only restore behavior.How to Test
api_contentsidecars.repair_alternation=True, close and reopen the database, then confirm the second load performs zero repairs and returns the same active replay bytes.All 227 targeted tests passed on the final commit, including three focused real-SessionDB display/resume regressions.
Checklist
Code
fix(scope):,feat(scope):, etc.)Documentation & Housekeeping
docs/, docstrings) - or N/Acli-config.yaml.exampleif I added/changed config keys - or N/ACONTRIBUTING.mdorAGENTS.mdif I changed architecture or workflows - or N/AScreenshots / Logs
REAL_ENV verification on the final commit repaired three violations on the first load and zero after close/reopen. The active sequence was byte-stable; survivor fields, inactive audit rows, stale sidecar clearing, FTS results, unrelated-session isolation, and read-only restore behavior all matched the expected contracts.