Skip to content

fix(hindsight): fence prefetch publication to the owning session - #64745

Closed
yingliang-zhang wants to merge 1 commit into
NousResearch:mainfrom
yingliang-zhang:fix/hindsight-prefetch-session-identity
Closed

yingliang-zhang wants to merge 1 commit into
NousResearch:mainfrom
yingliang-zhang:fix/hindsight-prefetch-session-identity

Conversation

@yingliang-zhang

@yingliang-zhang yingliang-zhang commented Jul 15, 2026 •

Copy link
Copy Markdown

Summary

Rewritten on top of current main after a design review against the prefetch machinery that landed upstream while this PR aged.

The leak that remains on main: the background prefetch worker publishes its recall result into the session slot unconditionally. on_session_switch joins the worker for 3.0s and then clears _prefetch_result, but the worker can legitimately spend up to 10s in _wait_for_retains_drained plus up to 120s in the recall itself. When it outlives the join it writes the old session's memories into the new session's slot, and the new session's first turn injects another conversation's context. queue_prefetch also spawned a fresh thread per turn with no liveness check, so a slow daemon plus rapid turns stacked N concurrent recalls against one embedded daemon and let the last finisher win the slot.

Fix (minimal, on upstream's shape)

  • _prefetch_generation, bumped under the existing _prefetch_lock on every spawn, on on_session_switch, and on shutdown().
  • The worker captures the generation at spawn and (1) skips its recall entirely when superseded or shutting down, (2) publishes only when its generation still owns the slot.
  • queue_prefetch skips while a prior worker is alive — the same idiom _ensure_writer already uses in this file (and honcho for its prefetch).
  • prefetch(), _join_prefetch, _wait_for_retains_drained, and the join-then-clear ordering on switch are untouched; test_in_flight_prefetch_thread_drained_on_switch stays green by construction.

What the original branch carried, and why it's gone

The previous head (92613bd5e1) carried a full _PrefetchRequest admission redesign: request identity dataclasses, a condition-variable admission path, a 2-worker pool + pending slot, retry-once scheduling, and client reservation/deferred-close ownership. Reviewing it against current main:

  • the pool/retry/client-ownership machinery addressed races that main's _run_hindsight_operation (reconnect + retry-once) and single-worker shape already cover;
  • it deleted two upstream tests (test_in_flight_prefetch_thread_drained_on_switch, test_prefetch_returns_empty_when_no_result) — a hard blocker;
  • its per-session bank re-resolution made prefetch read a different bank than its own retains write under a {session} bank template — strictly worse than main.

Dropped in favour of the ~30-line fence above. Original head preserved locally as archive/64745-original for reference.

Verification (head 44e607be61)

  • tests/plugins/memory/test_hindsight_provider.py: 92 passed, 1 skipped (87 upstream + 5 new)
  • tests/plugins/memory/ tests/agent/test_memory_provider.py: 524 passed / 4 skipped / 2 failed — the 2 test_mem0_v3.py backend-routing failures are pre-existing environment failures that also fail on clean main (missing mem0 import)
  • ruff + py_compile clean
  • Mutation check: removing only the publish-fence condition turns 3 of the new tests RED (test_stale_worker_cannot_publish_after_switch_join_timeout, test_superseded_worker_cannot_overwrite_newer_result, test_shutdown_fences_inflight_prefetch_publish); restoring it returns them to green.

New regression tests

  • stale worker cannot publish after the switch's join times out
  • a superseded worker cannot overwrite a newer result
  • shutdown fences an in-flight publish (no late publish, no client resurrection)
  • queue_prefetch performs exactly one recall under a burst while a worker is busy
  • positive control: an unswitched turn still delivers its warmed result

@teknium1

Copy link
Copy Markdown
Collaborator

Thanks for the focused lifecycle work. The premise is confirmed on current main: plugins/memory/hindsight/__init__.py:1484 accepts session_id but the worker at :1498-1522 reads mutable provider state and writes a shared result; on_session_switch() only performs a bounded join before clearing it at :1885-1888. agent/memory_manager.py:495-542 supplies session IDs to this provider path.

The PR addresses that class by binding request identity and fencing publication (plugins/memory/hindsight/__init__.py:1621-1750 on PR head 86e8ca56d69b39a8c4f7fa70962d4686b17a62bc), with session rotation under the same admission condition (:2690-2713). It remains within the existing plugin/test surface and adds no core tool, configuration, or prompt-cache mutation.

The PR base 6997dc81 is an ancestor of the inspected HEAD, with no later changes to either scoped file, so salvage appears mechanical.

Automated hermes-sweeper review.

@alt-glitch alt-glitch added type/bug Something isn't working P3 Low — cosmetic, nice to have comp/plugins Plugin system and bundled plugins tool/memory Memory tool and memory providers sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state labels Jul 16, 2026
@teknium1 teknium1 added sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades sweeper:blast-moderate Sweeper blast radius: moderate — a subsystem or single platform area/sessions Session lifecycle, resume, persistence, history area/memory Memory subsystem: store, providers, sync, background reviews labels Jul 16, 2026
@yingliang-zhang
yingliang-zhang force-pushed the fix/hindsight-prefetch-session-identity branch from 86e8ca5 to 8a7cbcb Compare July 30, 2026 05:55
@yingliang-zhang
yingliang-zhang force-pushed the fix/hindsight-prefetch-session-identity branch 3 times, most recently from 9f728d2 to 2b85276 Compare August 8, 2026 04:10
@yingliang-zhang
yingliang-zhang force-pushed the fix/hindsight-prefetch-session-identity branch 4 times, most recently from aef1bbc to 4abeb30 Compare August 18, 2026 01:24
@yingliang-zhang
yingliang-zhang requested a review from a team August 18, 2026 01:24
@yingliang-zhang
yingliang-zhang force-pushed the fix/hindsight-prefetch-session-identity branch 3 times, most recently from afd5827 to 0ea4817 Compare August 20, 2026 23:06
@yingliang-zhang

Copy link
Copy Markdown
Author

Status update after the #102117 refactor: I audited this PR against the new decomposed plugins/memory/hindsight structure. Verdict: every core mechanism the PR adds is still absent on new main — the audit finds no absorbed part beyond the generic embedded retry-once helper.

Missing on new main:

  • Request-time identity binding: queue_prefetch (new init.py around line 917) ignores its session_id kwarg entirely; the worker recalls with run-time self._bank_id/self._budget (via _do_recall to _recall, which builds kwargs at call time). A mid-prefetch session switch publishes the old query into the new bank — observable with templated {session} banks.
  • Cache publication fence: results publish into shared _prefetch_result with no epoch/session check. Main's only mitigation, on_session_switch's join(3s)+clear, is the heuristic this PR replaces: a worker can outlive 3s (retain-drain wait up to 10s + recall timeout), and it also blocks the switch itself.
  • Worker capacity vs future ownership: single _prefetch_thread pointer; each queue_prefetch overwrites it, orphaning any wedged worker whose late result then publishes straight into the new session.
  • Reconnect/client-replacement serialization: _run_hindsight_operation does a bare unlocked self._client = None + recreate (old client leaked, never closed, no identity check) — races shutdown's close.
  • Reserve-before-schedule, exact-once rollback, client-local closed markers, identity-checked weakrefs: none exist on new main.

Plan: port the PR onto the new structure as one coordinated change together with #64499's lifecycle half — both PRs rebuild the same client-lifecycle machinery, and the audit flagged that porting them independently would duplicate it.

yingliang-zhang added a commit to yingliang-zhang/hermes-agent that referenced this pull request Sep 10, 2026
…64745), stage 1

A background prefetch worker previously read mutable provider state
(self._bank_id, self._budget) at RUN time and wrote a shared result
slot; on_session_switch() only did a bounded join before clearing
state. A mid-prefetch session switch could therefore publish the OLD
session's query results into the NEW session's context (and attribute
recall to the wrong bank).

Stage 1 (request identity + fenced publication + switch hardening):

- _PrefetchRequest: frozen snapshot of request-time identity
  (session/bank/budget/query + epoch, cache_key, inflight_key) taken
  under the admission lock; _PrefetchOperation tracks its worker.
- queue_prefetch()/prefetch() admit through _admit_prefetch_locked:
  dedupe by cache_key, bounded worker pool (_MAX_PREFETCH_WORKERS=2)
  with one coalesced pending request, stale explicit-session rejection.
- The worker recalls with the SNAPSHOT (request.bank_id/budget/query),
  never live state; _do_recall/_recall/_reflect thread bank_id/budget
  kwargs so the recall_sync path binds to the caller's session too.
- Publication is fenced to the exact request (_request_is_current_locked:
  epoch + session + bank + latest-request identity); prefetch() consumes
  only the matching request's result (deadline-bounded condition wait).
- on_session_switch() rotates session+prefetch identity atomically under
  the admission lock (epoch bump) instead of joining in-flight workers:
  late old-epoch completions are fenced out and never delay the switch.
- shutdown() clears admission state under the lock and drains the worker
  set with a bounded budget; _get_client() creation is serialized.
- _bind_legacy_prefetch_result_locked fences externally-seeded results
  (raw _prefetch_result writes) so legacy behavior stays consumable.

Tests: TestPrefetchSessionIdentity (admission atomicity, no-cross-publish
across a same-query switch, superseded-waiter release, wait-timeout keeps
late result, failure retires inflight key, worker bounding/pending
coalescing, stale-session admission rejected, shutdown recheck under the
admission lock) + switch tests (no spurious flush, stale result cleared,
in-flight prefetch does not delay the switch; prefetch with no cached
result runs the current query).
yingliang-zhang added a commit to yingliang-zhang/hermes-agent that referenced this pull request Sep 10, 2026
…, stage 2

The prefetch worker schedules its async recall on the shared loop as a
future whose done-callback fires off-thread; a client retired while
that future is unresolved (timeout, reconnect, shutdown) used to be
closed underneath it, and a closed client could be handed back out by
a racing _get_client/_run_hindsight_operation.

Stage 2 (client ownership around inflight prefetch work):

- _execute_prefetch_attempt: reserve the client for the exact request,
  register the scheduled future on the request's _PrefetchOperation, and
  publish from the future's done-callback; a timeout retires the inflight
  key but keeps the operation until the future settles.
- _retire_hindsight_client: claim-close once; while prefetch owners hold
  reservations the close is deferred (and later scheduled on its own
  thread) until the last future releases it.
- Closed-client markers: attribute when possible, weakref registry with
  GC-safe eviction (id()-reuse safe), bounded strong-ref fallback for
  unweakrefable clients — a closed client can never be re-served.
- _run_hindsight_operation / prefetch retry path publish replacement
  clients through _publish_replacement_client so concurrent recreation
  keeps exactly one current client.
- _close_hindsight_client(client) replaces _close_client() (arg-taking,
  best-effort); shutdown retires the current client and schedules any
  deferred closes, so concurrent shutdowns claim the close exactly once.

Tests: TestPrefetchClientLifecycle — timed-out future defers an owned
close and later settles it; close between schedule and publication is
owned (accepted/rejected scheduler arms); 64x reconnect loop is bounded,
identity-safe and fully collectable; late-future ownership never
consumes worker capacity; embedded prefetch reconnect retries once on
a fresh client; _get_client serializes concurrent creation; retain-side
reconnect defers a prefetch-owned close; concurrent shutdowns close the
client exactly once.
@yingliang-zhang
yingliang-zhang force-pushed the fix/hindsight-prefetch-session-identity branch from 0ea4817 to 92613bd Compare September 10, 2026 14:08
@yingliang-zhang

Copy link
Copy Markdown
Author

Rebuilt onto current main (6e07eb4838) — the original branch lost its merge-base after the #102117 history rewrite (behind 9109, merge_base: none); the reviewed mechanism is re-applied fresh onto the decomposed structure as three commits (195db279c4 request identity + fenced publication, 180854ea4 client lifecycle, 92613bd5e1 review hardening; +1981/−111 across the provider and its tests).

Port notes (structure moved, mechanism unchanged — verified against the current seams):

  • The core guarantee set is intact and in places stricter than the original: _PrefetchRequest snapshots session/bank/budget/query at admission (the original left budget live); the worker recalls exclusively with the snapshot and publishes under a 5-factor currency fence (shutdown ∧ epoch ∧ session ∧ bank ∧ latest-request identity); on_session_switch bumps the epoch inside the admission lock and empties the slot; prefetch() consumes only the matching request via a deadline-bounded condition wait.
  • Client lifecycle: closed-client markers (attr → weakref → bounded id() fallback), single-claimer close, deferred close while reserved, bounded 2-worker pool + 1 pending slot (newest-wins, no starvation), retry-once on retriable embedded errors, and serialized _get_client — which also closes the [Bug]: Hindsight still leaks aiohttp ClientSession/connector after fix #4762 #11923-class unclosed-session leak on retry paths.
  • Adaptations required by the refactor: construction stays in main's _new_embedded_client/_new_cloud_client split (the original's _create_client port is unnecessary); workers spawn via _context_thread (preserves the multiplex-profile secret-scope snapshot); _format_recall → main's fused _finish_prefetch; _do_recall/_recall/_reflect take bank_id/budget kwargs (None → live state) so tool handlers and the recall-sync path keep their current signatures.
  • Retain side untouched (sync_turn/_make_turn_retain_job snapshot semantics preserved; the writer takes no new locks).

Verification (local): full provider file 107 passed (22 session-identity + 9 lifecycle tests added; the two mutation-proof wire-call pins from review); tests/plugins/memory/ 434 passed + 2 pre-existing (No module named mem0, fails on clean main); tests/agent/test_memory_provider.py 73 passed; the previously-flaky close-assertion now 20/20 stable; py_compile, ruff, git diff --check clean. Four independent review legs ACCEPT-WITH-NITS (P1 flake + P2 unpinned wire hop — both fixed in the third commit, test-only). CI will need workflow approval as usual.

@yingliang-zhang

Copy link
Copy Markdown
Author

Head 92613bd5e1 (the 09-10/09-11 rebuild onto post-#102117 main) still has all three workflow runs queued behind the fork-PR approval gate (action_required: CI run 34487226659, Nix 34487224966, Docker 34487224893). Could a maintainer approve the workflow run? Once approved, CI should validate the rebuilt head — local verification for the rebuild is documented in the rebuild comment below.

…ration (NousResearch#64745)

The background prefetch worker published its recall into the session slot
unconditionally; a worker outliving on_session_switch's 3s join wrote the
old session's memories into the new session's slot. queue_prefetch also
spawned unbounded threads with the last finisher winning the slot. Workers
now capture a slot generation at spawn, queue_prefetch bumps it and skips
while a prior worker runs, on_session_switch/shutdown bump it to fence late
publishers, and the publish + recall are gated on the current generation.
@yingliang-zhang yingliang-zhang changed the title fix(hindsight): bind prefetch work to session identity fix(hindsight): fence prefetch publication to the owning session Sep 20, 2026
@yingliang-zhang
yingliang-zhang force-pushed the fix/hindsight-prefetch-session-identity branch from 92613bd to 44e607b Compare September 20, 2026 12:55
@yingliang-zhang

Copy link
Copy Markdown
Author

Heads-up for reviewers: this PR is now a full rewrite — please ignore the old diff.

The previous head (92613bd5e1, ~1000 lines across the provider + tests) carried a full admission redesign: _PrefetchRequest identity dataclasses, a condition-variable admission path, a 2-worker pool + pending slot, retry-once scheduling, and client reservation/deferred-close ownership. Re-reviewing it against the prefetch machinery that landed on main while this PR aged, most of that machinery was re-covering ground main already covers, and it deleted two upstream tests (test_in_flight_prefetch_thread_drained_on_switch, test_prefetch_returns_empty_when_no_result) — a hard blocker. Its per-session bank re-resolution also made prefetch read a different bank than its own retains write under a {session} bank template.

Head is now 44e607be61 — one commit, +174/-1, two files, on top of current main (5a0fb0f). The diff is small enough to read directly:

  • _prefetch_generation, bumped under the existing _prefetch_lock on spawn, on on_session_switch, and on shutdown().
  • The worker captures the generation at spawn, skips its recall entirely when superseded or shutting down, and publishes only when its generation still owns the slot.
  • queue_prefetch skips while a prior worker is alive (the idiom _ensure_writer already uses in this file).
  • prefetch(), _join_prefetch and the join-then-clear ordering on switch are untouched.

The leak this closes on main: the prefetch worker publishes its recall unconditionally, but on_session_switch joins it for only 3.0s while the worker can legitimately spend up to 10s in _wait_for_retains_drained plus up to 120s in the recall. A worker that outlives the join writes the old session's memories into the new session's slot, and the new session's first turn injects them.

Verification: test_hindsight_provider.py 92 passed / 1 skipped (87 upstream + 5 new); tests/plugins/memory/ + tests/agent/test_memory_provider.py 524 passed, with the 2 test_mem0_v3.py backend-routing failures being pre-existing on clean main (missing mem0 import); ruff + py_compile clean. As a mutation check, removing only the publish-fence condition turns test_stale_worker_cannot_publish_after_switch_join_timeout, test_superseded_worker_cannot_overwrite_newer_result and test_shutdown_fences_inflight_prefetch_publish RED; they return to green once restored.

The old head is preserved locally as archive/64745-original on my fork if anyone wants to diff the two approaches. CI still needs the usual fork-workflow approval — sorry for the re-review, and thanks.

@teknium1

Copy link
Copy Markdown
Collaborator

Closing: the bundled Hindsight provider this PR patches has moved out of this repo.

Thanks @yingliang-zhang for this contribution. In #119888 (merge 9d799e0531c; removal commit 4cbf862abe4) the in-tree plugins/memory/hindsight/ provider was removed — Hindsight now installs from the plugin catalog and its code lives in vectorize-io/hindsight hindsight-integrations/hermes (maintained by @nicoloboschi). There is no longer any code in this repo for the Hindsight half of this PR to patch, so we are closing every open PR against the bundled provider rather than leaving them stranded.

Triage notes:

  • Current head (44e607b, rewritten 2026-09-20 to a ~30-line generation fence) touches only plugins/memory/hindsight/init.py and its test file. No generation fence ever landed on main (git log -S _prefetch_generation on the removed path is empty). At the pin, queue_prefetch (init.py:1247-1262) still publishes without a fence and on_session_switch (:1525-1527) still relies on join(3.0)+clear, so the fix is still relevant upstream.
  • Still relevant at the catalog pin (dc750388)? Yes — the same code is at hindsight-integrations/hermes/__init__.py:1260 in the upstream tree. It is listed with your credit in Fixes from Hermes-side PRs worth carrying into hindsight-integrations/hermes vectorize-io/hindsight#4662 so it is not lost; if you want to carry the fix yourself, please open it against vectorize-io/hindsight — it would be welcome there.

If you believe this was closed in error, comment and we will reopen.

(Bulk-closed in the hindsight-move close pass.)

@teknium1 teknium1 closed this Sep 23, 2026
yingliang-zhang added a commit to yingliang-zhang/hindsight that referenced this pull request Sep 30, 2026
Ported from NousResearch/hermes-agent#64745 (closed in the hindsight-move
close pass). Tracked in vectorize-io#4662.

The background prefetch worker published its recall into the session slot
unconditionally; a worker outliving on_session_switch's 3s join wrote the
old session's memories into the new session's slot. queue_prefetch also
spawned unbounded threads with the last finisher winning the slot. Workers
now capture a slot generation at spawn, queue_prefetch bumps it and skips
while a prior worker runs, on_session_switch/shutdown bump it to fence late
publishers, and the publish + recall are gated on the current generation.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/memory Memory subsystem: store, providers, sync, background reviews area/sessions Session lifecycle, resume, persistence, history comp/plugins Plugin system and bundled plugins P3 Low — cosmetic, nice to have sweeper:blast-moderate Sweeper blast radius: moderate — a subsystem or single platform sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state tool/memory Memory tool and memory providers type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants