Skip to content

fix(agent): stop repaired sessions from repeatedly missing cache prefixes - #82925

Open
fangliquanflq wants to merge 3 commits into
NousResearch:mainfrom
fangliquanflq:fix/session-message-repair-persistence
Open

fangliquanflq wants to merge 3 commits into
NousResearch:mainfrom
fangliquanflq:fix/session-message-repair-persistence

Conversation

@fangliquanflq

Copy link
Copy Markdown

What does this PR do?

Sessions with persisted alternation violations were repaired only in memory, so every resume could repair the same rows again and repeatedly lose the stable prompt-cache prefix. This change atomically reconciles repair results back to SessionDB while retaining removed rows as inactive audit records and preserving display-only verification candidates.

Symptom

After a malformed transcript is restored, consecutive requests repeatedly report the same message-sequence repairs. Repaired turns repeatedly fall back to an older prompt-cache prefix instead of reusing the stable repaired history.

Impact

Affected resumed sessions pay the repair and cache-miss cost on every subsequent turn. The reported trace showed the repaired requests reusing only about 30-38% of the prompt while adjacent clean requests reused 88-99%.

Bug Cause

Trigger: hermes_state.py:8413 / SessionDB.get_messages_as_conversation(..., repair_alternation=True)

Causal chain:

  1. A persisted transcript contains consecutive user or assistant rows, duplicate tool responses, or another repairable alternation violation.
  2. repair_message_sequence fixes only the decoded in-memory message list and the restore path does not update the corresponding SessionDB rows.
  3. A fresh load reads the original malformed rows and stale api_content, repeats the repair, and changes the replay prefix again.

Why it is wrong: Live replay treats the repaired list as authoritative for the current request but leaves durable state inconsistent with that list, so the same defensive repair cannot converge across resumes.

Working sibling / contrast: Clean transcripts already reload byte-stably because no repair mutation separates the in-memory replay from the persisted active rows. Inspection paths with repair_alternation=False intentionally remain read-only.

Ruled out: Compression alone is not the cause. A disposable SessionDB seeded directly with malformed rows reproduced repair counts [2, 2] across close and reopen before this fix.

Fix

Repair-enabled active restores now retain internal row identity, atomically update changed survivor content, tool calls, reasoning content, and api_content, and soft-archive rows removed from model replay. Existing FTS update triggers keep merged content searchable. Display-only verification candidates remain active, inactive audit loads do not reconcile, and internal row IDs are stripped unless requested by the caller.

Related Issue

Closes #82863

Type of Change

  • Bug fix (non-breaking change that fixes an issue)

Changes Made

  • hermes_state.py - reconcile identity-preserving replay repairs into active durable rows and archive repair-removed rows without deleting audit history.
  • tests/hermes_state/test_restore_alternation_repair.py - cover load-repair-reload stability, stale sidecar clearing, FTS updates, archived rows, and read-only restore behavior.

How to Test

  1. Seed a disposable SessionDB with consecutive user and assistant rows, a duplicate tool response, and stale api_content sidecars.
  2. Load with repair_alternation=True, close and reopen the database, then confirm the second load performs zero repairs and returns the same active replay bytes.
  3. Run the relevant suites:
scripts/run_tests.sh tests/test_hermes_state.py
scripts/run_tests.sh tests/hermes_state/test_restore_alternation_repair.py
uvx ruff check hermes_state.py tests/hermes_state/test_restore_alternation_repair.py

All 227 targeted tests passed on the final commit, including three focused real-SessionDB display/resume regressions.

Checklist

Code

  • I've read the Contributing Guide
  • My commit messages follow Conventional Commits (fix(scope):, feat(scope):, etc.)
  • I searched for existing PRs to make sure this isn't a duplicate
  • My PR contains only changes related to this fix/feature (no unrelated commits)
  • I've run the repository test entry on the relevant suites and all targeted tests pass
  • I've added tests for my changes (required for bug fixes, strongly encouraged for features)
  • I've tested on my platform: Windows 10

Documentation & Housekeeping

  • I've updated relevant documentation (README, docs/, docstrings) - or N/A
  • I've updated cli-config.yaml.example if I added/changed config keys - or N/A
  • I've updated CONTRIBUTING.md or AGENTS.md if I changed architecture or workflows - or N/A
  • I've considered cross-platform impact (Windows, macOS) per the compatibility guide - or N/A
  • I've updated tool descriptions/schemas if I changed tool behavior - or N/A

Screenshots / Logs

REAL_ENV verification on the final commit repaired three violations on the first load and zero after close/reopen. The active sequence was byte-stable; survivor fields, inactive audit rows, stale sidecar clearing, FTS results, unrelated-session isolation, and read-only restore behavior all matched the expected contracts.

@alt-glitch alt-glitch added type/bug Something isn't working comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint P2 Medium — degraded but workaround exists sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state labels Aug 10, 2026
@fangliquanflq fangliquanflq reopened this Aug 10, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint P2 Medium — degraded but workaround exists sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: repair_message_sequence does not reconcile state.db → durable alternation violations can re-fire and miss prompt-cache prefixes

2 participants