Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
65 commits
Select commit Hold shift + click to select a range
0cb7f37
test: pin source observation times in as_of-pinned trajectory fixtures
100yenadmin Jul 23, 2026
c4052df
fix: include time-contract sidecars in FTS search() rows
100yenadmin Jul 24, 2026
fc8ab23
fix: release query-view build lease on post-claim publish failure
100yenadmin Jul 24, 2026
7da3a21
fix: whitelist lcm_query_*/lcm_trajectory_* for schema-stamp repair
100yenadmin Jul 24, 2026
5832cfb
fix: drop schema-required proposal.operation that the runtime rejects
100yenadmin Jul 24, 2026
5a2e053
fix: hash requirement description into requirements_digest()
100yenadmin Jul 24, 2026
046764e
trajectory: Policy A lexical-floor nucleus slots (#127)
100yenadmin Jul 23, 2026
1f00780
trajectory: Policy D per-arm quota union (#127)
100yenadmin Jul 23, 2026
4a3ce7b
bench: provider-free H3.1 composition replay harness (#127)
100yenadmin Jul 23, 2026
bbbeb45
test: candidate-composition policies A + D on a synthetic magnet (#127)
100yenadmin Jul 23, 2026
55351c1
feat(composition): A+D hybrid (lexical_floor+arm_quota) — QUARANTINED…
100yenadmin Jul 23, 2026
54379ec
trajectory: H5(b) lexical-seed adjacency pool-expansion (#135)
100yenadmin Jul 23, 2026
d3517d5
trajectory: fetch full rows only for quota-admitted expansion states …
100yenadmin Jul 23, 2026
5f92ee4
bench: provider-free H5(b) recall replay harness (#135)
100yenadmin Jul 23, 2026
da1f8f0
trajectory: per-state semantic index + backfill + default-off pool ar…
100yenadmin Jul 23, 2026
1df8821
bench: state-embedding backfill CLI + state-semantic replay sweep (#142)
100yenadmin Jul 24, 2026
c9fc558
bench: fix state-semantic sweep harness bootstrap + rate guard (#142)
100yenadmin Jul 24, 2026
505a4d8
feat: add compact retrieval delivery controls
100yenadmin Jul 24, 2026
e99f342
trajectory: default-off anti-boilerplate MMR (G) + title-boost (H) re…
100yenadmin Jul 24, 2026
b2f228c
recall: reduce a raw question to FTS5 terms before MATCH (#168)
100yenadmin Jul 28, 2026
f960d9f
recall: scan the whole vector corpus in batches, not a 25k recency wi…
100yenadmin Jul 28, 2026
c13f4a2
review finding 1: judge the LIKE fallback on the SANITIZED query, not…
100yenadmin Jul 28, 2026
902fc37
review finding 2: stream multi-batch scans past the matrix LRU
100yenadmin Jul 28, 2026
c6fb1da
review finding 3: route the scan limits through the binary-prescreen …
100yenadmin Jul 28, 2026
4de8727
review finding 4: keep emoji on the LIKE path it was routed to
100yenadmin Jul 28, 2026
530bd98
review finding 5: neutralize bare boolean operators in the sanitized …
100yenadmin Jul 28, 2026
5b58154
review finding 6: compose the query before splitting it into FTS terms
100yenadmin Jul 28, 2026
e4577f1
review finding 5 (follow-up): mark the deliberate-operator mode inste…
100yenadmin Jul 28, 2026
169c3a3
review finding 2 (residual): release warmed matrices before a streame…
100yenadmin Jul 28, 2026
61d5b14
Merge pull request #169 from 100yenadmin/fix/phase1b-scan-and-query
100yenadmin Jul 28, 2026
18ba0a5
embed: configurable generous query-path spend guard (#123)
100yenadmin Jul 23, 2026
1b45d61
test: query-path spend guard exemption + typed rate-limit reason (#123)
100yenadmin Jul 23, 2026
afbb000
fix: preserve store_id on summary recall hits (#164)
100yenadmin Jul 28, 2026
b16f39b
chore: prepare v0.20.0 R2 train
100yenadmin Jul 28, 2026
543e9ea
Merge pull request #170 from 100yenadmin/release/r2-consolidated
100yenadmin Jul 28, 2026
a3a47e9
recall: reference-strict answer_ready delivery — never deliver an unc…
100yenadmin Jul 28, 2026
511d93e
test: pin the reference-strict delivery invariant (#164/F35)
100yenadmin Jul 28, 2026
0aab863
recall: carry summary-KNN relevance onto source messages (review find…
100yenadmin Jul 28, 2026
6f501f5
recall: validate and publish every delivered span (review findings 2,…
100yenadmin Jul 28, 2026
8ca43fe
recall: verify before admitting, not after (delta-2 findings 1, 2, 4)
100yenadmin Jul 29, 2026
5674cc0
recall: keep every citable representation of a fused row (delta-2 fin…
100yenadmin Jul 29, 2026
5db4e87
recall: a seen delta result must not spend a session slot (delta-3 fi…
100yenadmin Jul 29, 2026
d9ab8f9
recall: a missing row must not cascade waves or pin the corpus (delta…
100yenadmin Jul 29, 2026
1a0b0e2
recall: make selection accounting a ledger, not counters (delta-4 fin…
100yenadmin Jul 29, 2026
92d6fcd
recall: make the defer reserve a position, not a buffer (delta-5)
100yenadmin Jul 29, 2026
a8ada78
recall: expire cached state with its preconditions (delta-6 findings …
100yenadmin Jul 29, 2026
6049157
recall: derive the resume point instead of storing it (delta-7)
100yenadmin Jul 29, 2026
2edb8fc
Merge pull request #174 from 100yenadmin/fix/citable-delivery
100yenadmin Jul 29, 2026
247f811
bench-tooling: round-1 review fixes (bootstrap cleanup, trace cache, …
100yenadmin Jul 29, 2026
f5d893b
tests+stress: reconcile CI with the #168/#174 retrieval contract
100yenadmin Jul 29, 2026
cfbfa90
product: round-1 review fixes — 10 confirmed defects + cheap hygiene …
100yenadmin Jul 29, 2026
fb732bc
merge r1/ci-tests: CI reconciliation with #168/#174 contract + stress…
100yenadmin Jul 29, 2026
a861c36
merge r1/product-fixes: round-1 product review fixes (10 confirmed de…
100yenadmin Jul 29, 2026
7d025e8
merge r1/tooling-fixes: round-1 bench-tooling review fixes
100yenadmin Jul 29, 2026
f1eb364
product+tooling: round-2 review fixes (PR #175)
100yenadmin Jul 29, 2026
4bd8401
merge r2/fixes: round-2 review fixes (10 findings closed, 1 deferred …
100yenadmin Jul 29, 2026
d9a1ab7
round-3 review fixes: close the final 6 (PR #175)
100yenadmin Jul 29, 2026
fece680
merge r2/fixes: round-3 fixes (6/6 closed; zero HIGH+ found in round 3)
100yenadmin Jul 29, 2026
1361123
round-4 review fixes: all 8 findings from both reviewers (PR #175)
100yenadmin Jul 29, 2026
2e5ebed
merge r2/fixes: complete round-4 response (8/8 findings closed)
100yenadmin Jul 29, 2026
4b684c9
round-5 review fixes: revert unsound continuity change; close the ind…
100yenadmin Jul 29, 2026
05a46f6
round-6 review fixes: confirmation-pass findings (PR #175)
100yenadmin Jul 29, 2026
060a0df
merge r2/fixes: round-6 confirmation-pass fixes (4/4)
100yenadmin Jul 29, 2026
68b1b55
round-7 review fixes: final in-train batch (PR #175)
100yenadmin Jul 29, 2026
93a3ade
merge r2/fixes: round-7 final batch (6/6) — in-train fix cycle closes
100yenadmin Jul 29, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/ISSUE_TEMPLATE/bug_report.yml
Original file line number Diff line number Diff line change
Expand Up @@ -44,7 +44,7 @@ body:
attributes:
label: Version and branch
description: hermes-lcm version, branch, commit, or release tag.
placeholder: v0.19.0, main, or commit SHA
placeholder: v0.20.0, main, or commit SHA
validations:
required: true

Expand Down
8 changes: 8 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,14 @@ This repo also publishes GitHub Releases. This file is the repo-root release sur

## Unreleased

## v0.20.0 - 2026-07-29

Release focus: benchmark-driven retrieval scaling, evidence provenance, and bounded query embedding spend.

- Removed the large-corpus recall ceilings by scanning the full summary and chunk corpora in bounded batches and sanitizing raw natural-language FTS queries before fallback. (#169)
- Made the query-path embedding spend guard configurable with a generous default while preserving the exempt backfill contract. (stephenschoettler/hermes-lcm#434)
- Preserved a direct source `store_id` on summary recall hits so strict evidence renderers can validate their source identity. (#164)

## v0.19.0 - 2026-07-07

Release focus: data-safety hardening, operator diagnostics, import tooling, benchmarking, and the WS5 engine decomposition.
Expand Down
24 changes: 24 additions & 0 deletions DELTA-REVIEW-PACKET.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,24 @@
# DELTA REVIEW — PR #169 fix commits only (review ONLY, no modifications)

Prior full review (sol·max) returned APPROVE-WITH-FIXES with 6 findings; the author fixed all six
(commits c13f4a2, 902fc37, c6fb1da, 4de8727, 530bd98+e4577f1, 5b58154 — one per finding, regression test each,
built from your repros). Scope: `git diff f960d9f..e4577f1` (the fixes only). The PR comment maps commits to
findings. Verify:

1. Each finding actually closed by its commit (re-run your original repro logic mentally or via in-memory
SQLite probes; e.g. requires_like_fallback("art-related") now False; NFD naïve matches; "Portland, OR
hotel" has no operator semantics under the default mode).
2. **The mode split (finding 5's real fix):** `search(..., allow_operators=False)` default with harness
opt-in — sweep EVERY caller of search/sanitizer entry points: does any raw-user-query path get
allow_operators=True? Does any deliberate-FTS caller silently lose operators? (The author found
benchmarking/longmemeval.build_fts_query relies on OR — verify its opt-in is correct and no other caller
was missed.)
3. **Finding 2's fix:** multi-batch sweeps now stream past the LRU — verify peak memory is truly one batch
(no reference retention), single-batch paths still use the cache byte-identically, and no double-fetch.
4. The two flagged semantic changes: compound queries now conjunctive-on-index (was disjunctive-on-LIKE) —
any real caller for whom that is a regression? And the four retargeted LIKE-trigger tests — do they still
test their original contracts?
5. Any NEW hole opened by the fixes themselves.

Output: VERDICT APPROVE / APPROVE-WITH-FIXES (mandatory list) / REJECT, findings with file:line + severity,
max 8, real defects only. No hardening suggestions.
19 changes: 19 additions & 0 deletions FINDINGS-VERDICTS-R2.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
# Round-2 findings verdicts

| Item / finding | Verdict | Source validation and disposition | Focused proof |
|---|---|---|---|
| 1 / `3671040912` | CONFIRMED-FIXED | Fully synced summary/chunk binary profiles bypassed all deadline checks. Binary loads now use deadline-interruptible read connections; Hamming batches, survivor loads, and rescoring check the same absolute deadline. Live two-stage output remains `full_approx`; actual expiry returns `bounded`. | `test_deadline_bounds_a_synced_binary_summary_prescreen`; `test_chunk_deadline_bounds_a_synced_binary_prescreen`; existing `test_two_stage_full_approx_coverage_surfaces_as_approximate`. |
| 2 / `3671040917` | CONFIRMED-FIXED | Oversize state chunks were sliced only by `batch_max_items`. Packing now stops on either item count or cumulative `batch_token_budget`, using the same estimator as normal documents. | `test_oversize_chunks_pack_by_item_and_token_budgets`; existing chunk/estimator tests. |
| 3 / `3671040939` | CONFIRMED-FIXED | `lcm_query*` and `lcm_trajectory*` tables passed the broad prefix allowlist but had no classifier verifier; remediation could lower the stamp while preserving an unknown shape. Current table/column/object shapes are now verified exactly, and malformed/future preserved families fail closed without drops or re-stamping. | Current full query/base+optional trajectory shapes classify interim; partial, extra-column, and unknown-table shapes classify genuinely newer and refuse apply. |
| 4 / `3671040926` | CONFIRMED-FIXED | Reference-strict summary lineage walked recursive source IDs and hydrated messages after the wrapped KNN/hydration stages, with no deadline. The expansion now uses an interruptible read-only snapshot and checks the deadline through lineage, hydration, and shaping. | `test_summary_source_expansion_refuses_an_expired_deadline`; recall suite. |
| 5 / `3671040922` | CONFIRMED-FIXED | The paid dimension probe ran before statistics initialization. Probe calls and `last_usage_tokens` now seed `provider_calls` and `billed_tokens`. | Extended `test_dimension_probe_is_bounded_by_document_token_budget`. |
| 6 / `3671050099` | CONFIRMED-FIXED | Stress containment/scope checks inconsistently read only `results` although other paths recognized `results`, `matches`, and `data`. One helper now supplies every such check. | `test_lcm_grep_result_rows_collects_every_supported_container`; stress smoke. |
| 7 / `3671050097` | CONFIRMED-FIXED | Unknown models silently used `$0.06/M`, weakening `--cost-cap`. Unknown pricing now errors before provider/store work unless `--assume-rate` explicitly supplies a positive rate. | `tests/test_state_embedding_backfill_cli.py`. |
| 8 / `3671050112` | CONFIRMED-FIXED | The state semantic exception arm incremented only the legacy fallback counter. It now records the typed exception/reason with the same fenced fallback telemetry as the source arm. | Extended `test_state_query_provider_failure_degrades_to_lexical`. |
| 9 / `3671050109` | CONFIRMED-FIXED | `REGRESSION-REPORT.md` embedded workstation paths and stale pre-fix status. The command now uses an explicit environment root plus repository-relative script/output paths and records final green status. | Absolute-path scan is empty; stress smoke passes. |
| 10 / `3671050103` | CONFIRMED-FIXED | The round-1 verdict document still said all `benchmarking/` work was excluded although the stress marker normalization landed. It now records that exception and passing smoke. | Documentation diff plus stress smoke. |
| 11 / `3671040934` | DEFERRED-NO-CHANGE | The symbol-loss claim reproduces in current semantics, but `SPEC.md` explicitly freezes this query-semantics change for this train and tracks it in fork issue `#172`. `search_query.py` and its tests were not changed. | Source inspection only; no current-lane test or code by triage. |

V1 delivery flag: none. Unexpired prescreen ranking remains byte-identical and keeps `full_approx`; only actual deadline expiry changes coverage/content. The lineage change likewise stops only after expiry. Other fixes affect accounting, validation, diagnostics, benchmarks, or explicit CLI error handling rather than default V1 delivered-hit selection.

Validation: touched area `226 passed`; full CI replica `2687 passed, 35 failed, 1 skipped, 12 xfailed`. Compared with `laneA-postfix-full-failures.txt`, there are zero new failure names and the prior oversize-chunk failure is resolved. Ruff and `git diff --check` pass.
21 changes: 21 additions & 0 deletions FINDINGS-VERDICTS-R3.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
# Round 3 findings verdicts

Base: `fork/bench/w3b-on-wave1` at `4bd8401d96fb00fbcf37cc0ad0852a4ec7e0ea12`.

1. `query_view_store.py` query-family discovery — **CONFIRMED-FIXED**. The verifier now discovers the full gated `lcm_query%` prefix. The `lcm_querycache` regression classifies the database as genuinely newer and refuses downgrade.
2. `tools.py` reference-strict delta after response-cap eviction — **CONFIRMED-FIXED**. Delta refs and progress fields are rebuilt from surviving delivered hits after whole-hit eviction. The regression asserts delivered refs equal delta refs and omitted refs remain unseen. **V1-DELIVERY-AFFECTING: cap-eviction path only.**
3. `trajectory_store.py` chunked-state spend progress — **CONFIRMED-FIXED**. Every successful chunk request emits cumulative progress before the next request. The low-cap regression stops after one chunk and preserves its provider-call and billed-token spend in the callback ledger.
4. `tools.py` conditional `rows` binding — **CONFIRMED-FIXED**. `rows` is always bound to the batch result or `{}`. The related shallow-copy nit is also aligned by replacing `read_store._write_lock`.
5. `vector_store.py` deadline clock seam — **CONFIRMED-FIXED**. `_monotonic` is module-local and used by deadline/budget paths; affected tests patch the seam. The summary mid-prescreen test is no longer coupled to process-wide time or an exact call sequence.
6. Chunk-side mid-prescreen deadline coverage — **CONFIRMED-FIXED**. The chunk test now expires after the first meaningful prescreen batch and asserts `coverage="bounded"` with `scanned == 1`; `bounded_scan_rows=1` now controls that exercised batch.

Validation:

- Exact regressions: 7 passed.
- Touched-area modules: 144 passed.
- Clock-seam delta: 2 passed.
- Full CI replica: 2689 passed, 35 failed, 1 skipped, 12 xfailed.
- Baseline comparison: 35 actual failure names exactly match 35 expected; zero new and zero missing.
- `git diff --check`: clean.

No commit or push performed.
17 changes: 17 additions & 0 deletions FINDINGS-VERDICTS-R4.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
# Round 4 findings verdicts

Merge: `d9a1ab79d9a65470ed3501ff8b7eabaeed1d80cb` fast-forwarded to `fece680d5f016b2a22cefb69e7b634a88e4f70d2`.

| Item | Verdict | Disposition |
|---|---|---|
| 1 strict response cap | CONFIRMED-FIXED | Evict hits, then summary leads; rebuild delta refs from surviving hits only. |
| 2 in-memory deadline load | CONFIRMED-FIXED | `:memory:` and memory-URI loads use the current interruptible connection. |
| 3 profile rebuild continuity | CONFIRMED-FIXED | Keep the old profile active until the completed replacement cuts over atomically. |

V1-DELIVERY-AFFECTING: item 1 only, at cap eviction.
Focused regressions: 3 passed (`laneR4-logs/focused-r4.xml`).
Touched area: 181 passed (`laneR4-logs/touched-area-r4.xml`).
Ruff on changed source/tests: clean (`laneR4-logs/ruff-r4.txt`).
Full replica: 2691 passed, 35 failed, 1 skipped, 12 xfailed (`laneR4-logs/full-suite-r4.xml`).
Baseline comparison: 35 actual = 35 expected; zero new/missing failure names (`laneR4-logs/full-suite-r4-comparison.txt`).
`git diff --check`: clean. No commit or push performed.
19 changes: 19 additions & 0 deletions FINDINGS-VERDICTS-R4B.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
# Round 4B findings verdicts

PR `100yenadmin/hermes-lcm#175` @ `fece680d5f016b2a22cefb69e7b634a88e4f70d2`.
The spec's single-comment URL returned 404; all listed bodies were fetched by the canonical `pulls/comments/<id>` route.
Comment `3671346839` is the already-fixed R4 profile-continuity item; the five R4B items follow.

| Item | Verdict | Disposition |
|---|---|---|
| 1 query objects | CONFIRMED-FIXED | Reject extra `lcm_query*` indexes/triggers; both probes classify genuinely newer. Trajectory already rejects both extra object types. |
| 2 probe ledger | CONFIRMED-FIXED | Emit the probe's calls/tokens before any document request; callback can stop spend immediately. |
| 3 scan budget | CONFIRMED-FIXED | Config promises a hard scan budget; unlimited enumeration had no deadline. Start at operation entry and interrupt summary/chunk candidate reads. |
| 4 H3 golden gate | CONFIRMED-FIXED | Print failure and return 1 before ground truth, sweep, latency work, or artifact writing. |
| 5 temporal tokens | CONFIRMED-FIXED | Match temporal terms as whole tokens; `update` no longer matches `date`. |

Item 5 affects only sharp-compilation V2 (`TrajectoryStore.query` with `sharp_token_budget > 0`); V1 message-store recall is untouched.
Focused modules: 125 passed (`laneR4B-logs/focused-modules-final.xml`); Ruff and `git diff --check`: clean.
Full replica: 2697 passed, 35 failed, 1 skipped, 12 xfailed (`laneR4B-logs/full-suite-r4b.xml`).
Baseline: exact same 35 R4 failure names; zero new/missing (`laneR4B-logs/full-suite-r4b-comparison.txt`).
No commit or push performed.
8 changes: 8 additions & 0 deletions FINDINGS-VERDICTS-R6.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
# Findings verdicts — R6

- FIXED — adaptive_retrieval.py P1 (Codex 3671928491; EvaOS 3671869848): digest capacity is derived from all declared requirement limits with one max-item headroom; 12 max-sized requirements validate and start.
- FIXED — trajectory_store.py P2 (Codex 3671928473): normal-batch provider spend and progress emit before vector validation/persistence; forced persist failure retains the full spend ledger.
- FIXED — trajectory_store.py P2 (Codex 3671928480): termless lexical queries become an empty lexical pool when state semantics are enabled; emoji-only semantic recall succeeds and quota-zero telemetry stays identical.
- FIXED — benchmarking/h5_state_semantic_replay.py P2 (Codex 3671928484): output parents are created immediately after argument parsing, before provider construction, the golden gate, warm-up, or sweep.
- VALIDATION — 4 focused regressions and 38 touched-area tests pass; Ruff passes; full suite has 35/35 baseline failure names with zero new names.
- PROOF BOUNDARY — source and CI-replica local test proof only; no commit, push, PR update, merge, release, or runtime change.
13 changes: 13 additions & 0 deletions FINDINGS-VERDICTS-R7.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,13 @@
# Round 7 findings verdicts
1. Fixed: forced same-profile rebuilds delete prior profile rows before refill; partial-rebuild regression passes.
2. Fixed: state documents use the smaller document/request token budget for chunk routing; request-cap regression passes.
3. Fixed: Policy D arm-quota cells measure recoverability from their fused candidate pool; scoped-row regression passes.
4. Fixed: evidence requirement descriptions reject surrogate code points before canonical UTF-8 digesting; regression passes.
5. Fixed: the H5 out-dir test now fails if provider construction occurs before the API-key guard.
6. Fixed: termless queries skip source-semantic embedding while state-semantic seeding and telemetry remain intact.
Validation: 6 focused R7 regressions passed.
Validation: full suite 2705 passed, 35 failed, 1 skipped, 12 xfailed.
Baseline comparison: 35/35 prior failure names; 0 new and 0 missing — PASS.
Issue-filing note: generation-scoped/profile-scoped embedding rows remain the larger option for availability during rebuild; not included here.
Scope: default-off/quota-gated subsystem, bench tooling, and tests only; no delivery path changes.
Proof boundary: local CI-replica source/test evidence only; no commit, push, PR update, release, or runtime proof.
22 changes: 22 additions & 0 deletions FINDINGS-VERDICTS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
# Round-1 product findings verdicts

| Item | Verdict | Source validation and disposition | Focused proof |
|---|---|---|---|
| 1 | CONFIRMED-FIXED | `state_semantic_quota > 0 and rows` excluded the empty-FTS underfill case. The state-semantic arm now runs for any positive quota and remains additive; quota zero is unchanged. | `test_expansion_can_fill_a_pure_lexical_miss`; existing default-byte and additive-tail tests. |
| 2 | CONFIRMED-FIXED | A profile was activated before any state vector write. Profiles now remain inactive throughout resumable writes, pass a complete-row-count check, and activate only at completion. An interrupted run leaves staged rows but no incomplete active profile; resume reuses them. | `test_profile_activates_only_after_resumed_backfill_completes`; existing idempotent/resume test. |
| 3 | CONFIRMED-FIXED | `_scan_ranked` checked only its relative budget. The absolute recall deadline now reaches summary and chunk KNN and is checked before and after each batch, producing bounded coverage on an early stop. | `test_full_scan_absolute_deadline_stops_between_batches`; existing budget/full-scan tests. |
| 4 | CONFIRMED-FIXED | `_semantic_state_ranks` provider exceptions escaped the query path. The state arm now fences them, increments the existing fallback counter, and leaves lexical results intact. | `test_state_query_provider_failure_degrades_to_lexical`. |
| 5 | CONFIRMED-FIXED | State query embeddings bypassed `_semantic_usage`. Query calls and successful usage tokens are now counted consistently with source-semantic queries. | `test_state_query_embedding_is_counted_in_semantic_usage`. |
| 6 | CONFIRMED-FIXED | The dimension probe sent the first full state document. It now probes only the first provider-bounded token chunk. | `test_dimension_probe_is_bounded_by_document_token_budget`. |
| 7 | CONFIRMED-FIXED | The fallback `token_budget * 4` window could exceed the shared estimator, especially for non-ASCII text; chunks now use that estimator directly. The reported `chunked_states == 1` test expectation was wrong independently: both fixture documents exceed five cl100k tokens, so the corrected contract is `2`. | `test_fallback_chunks_obey_the_shared_token_estimator`; corrected `test_chunked_path_pools_oversize_documents`. |
| 8 | CONFIRMED-FIXED | The hint consumer selects `lcm_expand(node_id=...)` only when `from_current_session` is present. Summary leads now preserve that context. | `test_summary_leads_preserve_current_context_and_obey_response_limit`. |
| 9 | CONFIRMED-FIXED | Summary leads followed the 50–100 row retrieval window while the response limit is at most 25 and may be lower. The actual clamped response limit now bounds leads. | `test_summary_leads_preserve_current_context_and_obey_response_limit`. |
| 10 | CONFIRMED-FIXED | `mark_failed` could raise inside the publish exception handler and replace the original failure contract. Cleanup is now best-effort and exception-safe. | `test_persist_compiled_view_cleanup_failure_does_not_escape`. |
| 11 | CONFIRMED-FIXED | Fixed the real/cheap items: state-cache freshness marker plus regression, removed the unused recall constant, corrected the summary-arm annotation, centralized the message column count, corrected the chunk-KNN docstring, narrowed the clock patch, added direct nested/source DAG tests, and extended query/trajectory remediation apply coverage. REFUTED: the MMR `assert` is not load-bearing because every nonempty `remaining` iteration assigns `best_item`; state-semantic IDs are unique by the embeddings PK, ranking indices, and merge dedupe. DECLINED: no SQLite row-value failure exists on the supported Python 3.11+ CI matrix, so no compatibility branch was added without a reproduced supported-path failure. | `test_state_matrix_cache_refreshes_after_same_profile_rewrite`; `tests/test_dag_source_message_ids.py`; extended schema/vector tests; touched-file suite. |

V1 delivery flag: none. Items 1–7 remain default-off or change only an already-expired operation; item 8–11 changes do not alter default delivered-hit selection.

Round-1 follow-through: the stress-CLI FTS marker normalization in
`benchmarking/stress.py` landed with the round, and the final stress smoke passed.

Explicit exclusions were not changed: `store.py:1123` query semantics, arbitrary summary source-row selection, and `tests/test_lcm_engine.py`; no benchmarking change beyond the documented stress-CLI fix belonged to round 1.
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -205,7 +205,7 @@ Typical output:

```text
Plugins (1):
✓ hermes-lcm v0.19.0 (8 tools)
✓ hermes-lcm v0.20.0 (8 tools)

Provider Plugins:
Context Engine: lcm
Expand Down
Loading