Repository navigation
fix(hermes): pass retain_async=True to aretain_batch in tool handler - #4671
yingliang-zhang wants to merge 6 commits into
Conversation
_client is read/written from the retain writer, the prefetch worker and
the turn/tool thread with no lock. _get_client() was check-then-act and
the embedded constructor takes seconds: two threads hitting a cold or
just-nulled client both construct, and the loser's client is orphaned
with an aiohttp session nothing ever closes ("Unclosed client session"
noise). The stale-daemon retry's inline null-and-rebuild widened the
window to seconds and concurrent retries clobbered each other's
replacement.
- _client_lock, a LEAF lock: _get_client() takes it only on the
construction/retire path; the fast path stays lock-free; never held
across _run_sync/operation(client) or while taking _prefetch_lock /
_pending_retain_ops_lock.
- _get_client(*, retire=<client>): the retry retires the EXACT client
the operation ran with — identity passed as an argument, never shared
broken-client state a concurrent retry could clobber — and rebuilds
exactly once under the lock (identity CAS: a sibling's fresh rebuild
is returned as-is instead of being orphaned).
- shutdown() retires the client under the lock BEFORE closing it
(_close_client_of(client), parameterised); a concurrent _get_client()
rebuilds instead of racing the close.
Ported from NousResearch/hermes-agent#117236 (closed in the hindsight-move
close pass). Tracked in vectorize-io#4662.
|
CI note: |
b177530 to
4d84463
Compare
|
Generated-files sync commit added atop the substantive commit(s): the |
Conflict resolution (both intents kept): hindsight-integrations/hermes/__init__.py merges the leaf-lock client lifecycle (_client_lock, _get_client(retire=...), _close_client_of(client), shutdown retire-then-close) with main's daemon-based embedded client (_embedded_url, HindsightEmbedded removal). _close_client_of keeps the PR's parameterized identity with main's simplified single-aclose body. skills/hindsight-docs/references/developer/oracle.md auto-merged to main's regenerated blob (a8d62fc): the PR's vectorize-io#4678 mirror is subsumed by main's generate-docs-skill.sh output from the same source (byte-verified). Checks: 37 hermes tests passed incl. test_client_lifecycle.py; ruff check+format clean per scripts/hooks/lint.sh's hermes path.
The hindsight_retain tool handler called _retain_batch without forwarding the configured self._retain_async, so tool retains silently dropped the async/sync choice and aretain_batch fell back to its own server default. retain_async is a call-level arg (never an item key), so the tool handler must pass it explicitly. Forward the configured mode, matching the auto-retain path. Ported from NousResearch/hermes-agent#60648 (closed in the hindsight-move close pass). Tracked in vectorize-io#4662.
203c705 to
ac0d77b
Compare
…leaf-lock # Conflicts: # hindsight-integrations/hermes/__init__.py
…in-async-tool-handler
|
Thanks for this! We're tracking this problem in #4662 and will fix it from there, so I'm closing this PR. We've stopped accepting pull requests from outside the team (see CONTRIBUTING.md). Any extra detail or steps to reproduce on the issue are very welcome. |
Problem
The
hindsight_retaintool handler calledself._retain_batch(item, bank_id=self._bank_id)without forwarding the configuredself._retain_async, so tool-initiated retains silently dropped the provider's async/sync choice —aretain_batchfell back to its own server default.retain_asyncis a call-level arg (never an item key —_retain_batch's contract), so the tool handler must pass it explicitly; the auto-retain path (_make_turn_retain_job→_retain_batch) already does.Concretely: a user with
retain_async: falsegot synchronous semantics for auto-retains but server-default semantics for tool retains — and on banks with significant data the LLM fact-extraction call can take 60–600+ s, far exceeding the 120 s default timeout, surfacing as an opaqueTimeoutError("Failed to store memory: "). With the flag forwarded,retain_async: truetool retains return as soon as the server accepts the request, matchingsync_turn.Fix
Forward the configured mode as a call argument, matching the auto-retain path:
This preserves the review feedback applied on the original PR (pass the captured
self._retain_async, not a hard-codedTrue, so an explicitretain_async: falseis honored on both paths).Tests
tests/test_retain_async_forwarding.py—test_retain_tool_forwards_configured_retain_async, parametrized overTrue/False, asserting the call-levelretain_asyncmatches the provider config (port of the original parameterized test, adapted to this tree's recordingFakeClient).Verification
uv run pytest tests/test_retain_async_forwarding.py -v(only the touched test file):Mutation check — reverting the one-line fix makes both cases RED:
uv run pytest tests/test_provider.py(existing suite exercising the touched path, unchanged): 15 passed in 0.28s.Provenance
Ported from NousResearch/hermes-agent#60648 (closed in the hindsight-move close pass; the bundled provider moved to
hindsight-integrations/hermesin NousResearch/hermes-agent#119888). Original fix and analysis by @yingliang-zhang; review feedback by @teknium1 (configured value instead of hard-codedTrue, parameterized coverage) is incorporated. Tracked in #4662.Stacked PR (pre-chained 2026-09-30): the diff vs main is cumulative with #4674; this PR's own change is forwarding the configured
retain_asynctoaretain_batchin thehindsight_retaintool handler. Merge top-down: 4674 -> 4671 -> 4672 -> 4673 -> 4675 -> 4676. Identical shared commits merge cleanly in any order (git dedups same-patch-both-sides).