feat: align LCM with Codex continuity and whitepaper control flow - #215
100yenadmin wants to merge 2 commits into
Conversation
Add provider-native Codex continuity, deterministic map operators, shared storage ownership, durable lock recovery, large-file lineage, bounded restart reconciliation, and explicit soft/hard compaction behavior while preserving existing Hermes plugin and storage compatibility. Constraint: Preserve plugin name, context engine name, and database compatibility Rejected: Raise compression timeouts | leaves archive scanning unbounded and masks the restart defect Confidence: high Scope-risk: broad Reversibility: clean Directive: Keep opaque provider capsules out of FTS, summaries, inspection, and public expansion Tested: 2,846 tests passed with one expected skip; Ruff; compileall; shell syntax; live 590K-token gateway auto-resume Not-tested: GitHub Actions Python 3.13/3.14 runners before PR publication
Restarted gateway sessions can contain thousands of replay rows, transient host placeholders, and rewritten Discord messages. Reconcile them with linear-time prefix/suffix matching, a conservative durable-tail anchor for very large histories, and an indexed tool-result archive lookup. Constraint: Preserve ambiguous short deltas and immutable stored payloads Rejected: Raise the 30-second timeout | masks quadratic work and continued duplicate ingestion Confidence: high Scope-risk: moderate Reversibility: clean Directive: Keep the 4096/1024 large-replay anchor conservative unless production evidence and replay-safety tests justify changes Tested: 1071 core and engine tests; exact 8964-message production replay in 2.136 seconds; 9000-message mismatch in 0.14 seconds Not-tested: Destructive cleanup of pre-existing duplicate LCM rows
|
Triage: this is one of three open PRs implementing the same architectural change — flagging so it is not reviewed in isolation. #197, #202 and #215 each introduce a shared, reference-counted SQLite storage bundle across LCM engine clones. All three necessarily relax the same documented guarantee — that a clone owns isolated storage — by removing the six To be clear, that is not test-hiding, and I checked before saying so: all three add replacement assertions for the new invariant (9, 8 and 11 respectively — clone-local session state, one idempotent lease per engine, owner-first shutdown). This is a legitimate contract change with coverage. The motivation is real and measured, from #197: ten retained clones cost 60 file descriptors and roughly 100 ms of duplicate SQLite helper construction, which risks descriptor exhaustion under parallel child-agent workloads. That is the defect tracked as #9. So the problem is not any one of these PRs — it is that there are three. Reviewing them independently would mean deciding the same architecture question three times, and merging any one of them makes the other two conflict against a moved contract. The maintainer action is to pick one design, land it, and close the other two with evidence. This needs an owner/architecture call, because it changes a documented guarantee for every embedder of the engine. It is not a call I will make unilaterally as maintainer. Recorded on #9 as the tracking issue; holding all three until it is decided. |
|
Closed superseded by the owner-ratified storage-trio decision: #197 is the pick (memo + ratification on #9, 2026-08-20). Reopenable — this records a product decision, not a merit refutation. The Codex-continuity/whitepaper control-flow ideas herein remain referenced from the memo for future consideration. |
Important
This executable LCM-X PR was recreated by the migration operator from the exact upstream commit head. GitHub did not transfer the original PR actor, dates, review objects, or approval state.
Source and attribution
02428dde78f70a9bc101b9d3e70b2d7c5e6c4e02codejeet/hermes-lcm:feat/codex-whitepaper-reliabilitymainupstream/pr-527Falsedraft(a maintainer can mark it ready when current-head review is wanted)@coderabbitai ignore
Original commit authorship and history remain in the commits. Historical discussion and review text are imported below as attributed ordinary comments; they are not new approvals or change requests.
Original upstream PR description
Summary
Why
Long-running concurrent gateway sessions exposed two production failures: restart reconciliation repeatedly scanned a 12,799-file, roughly 540 MB archive, and cloned engines could contend over independently owned SQLite helpers. The first failure pushed a 590K-token resume beyond the 30-second host timeout. The bounded lookup and per-pass identity cache reduce that reconciliation to roughly 2 seconds while preserving every source row and externalized payload.
Validation
Notes