Skip to content

fix(hermes): queue asynchronous manual memory retains - #4670

Closed
yu-xin-c wants to merge 2 commits into
vectorize-io:mainfrom
yu-xin-c:codex/hermes-retain-async
Closed

yu-xin-c wants to merge 2 commits into
vectorize-io:mainfrom
yu-xin-c:codex/hermes-retain-async

Conversation

@yu-xin-c

@yu-xin-c yu-xin-c commented Sep 23, 2026 •

Copy link
Copy Markdown

Summary

With the default retain_async=true, an explicit Hermes hindsight_retain call still runs inline and omits the client's retain_async argument. A busy backend can therefore block the agent until the retain timeout, even after accepting the memory.

Route asynchronous manual retains through the existing single-writer queue, passing the configured flag at call level and tracking returned operation IDs for the next prefetch's visibility barrier. Capture the item and destination before enqueueing. The tool now reports "Memory queued for storage."; retain_async=false still waits and returns backend errors directly.

This carries forward my NousResearch/hermes-agent#61512 after the bundled provider moved here. Addresses the _tool_retain async entry in #4662 and NousResearch/hermes-agent#61442; #4662 contains other independent fixes and should stay open. The operation tracking and acknowledgement distinction also incorporate feedback from @russellbrenner on the original PR.

Validation

  • Integration unit suite after the rebase: 31 passed, Python 3.11.
  • Before the implementation, the new regressions produced 4 failed, 1 passed, including a blocked tool call while the fake backend withheld its acknowledgement.
  • Tests cover a real writer-thread round trip, operation tracking, queued payload/destination snapshots and FIFO ordering, synchronous flag forwarding, and error handling for both modes.
  • Ruff check and format check with the repository's ruff.toml pass for the Hermes integration; git diff --check passes.
  • Current head's CI run 36301585571 passed the enabled Hermes integration, Hermes compatibility, documentation, and generated-file/lint jobs. Unrelated jobs were skipped; this is not a claim that every monorepo test ran locally.

Actual backend verification (2026-09-27)

Ran a separate native hindsight-api process and real pg0 PostgreSQL, using the API and client built from this checkout. Loaded the external plugin through the actual Hermes memory-provider registry from Hermes babdfff62940150e55a93273f426f3d3ddddd651; no fake Hermes modules or FakeClient. This targeted integration used an isolated Python 3.11.15 environment, not a full Hermes CLI installation.

The provider talked over loopback HTTP through a forwarding proxy to the real API. The proxy held the retain request for one second; the upstream system-test model service then held fact extraction for one second. These controlled holds verify blocking behavior, not production throughput.

Plugin Setting Actual HTTP async Returns while HTTP/model held Measured tool latency
Base ccfe85b485 true false no 2065.96 ms
Patched ba22b0f9fc true true yes 1.54 ms
Base false false no 2041.11 ms
Patched false false no 2058.94 ms

For the patched async case, the real server receipt was tracked after the local queue drained; background prefetch stayed blocked and sent no recall request while extraction was held. After release, prefetch returned the stored fact, the operation tracking set drained, and an independent client both listed and recalled the persisted memory. All four cases stored and recalled their expected fact. API, PostgreSQL and temporary banks/homes were cleaned up.

Scope of the September 27 run: LLM, embeddings and reranking used the repository's deterministic system-test provider over HTTP. No real model API key was available at that time, so that run does not validate model quality, production latency, cloud service behavior or long-term durability across database restarts. The read-after-write probe explicitly configured recall_types=["world"]; it does not promise immediate visibility of separately consolidated observations under the default observation-only recall filter. No production implementation change was needed after this verification.

Live model verification (2026-09-30)

Repeated the before/after test with real Bailian qwen3.7-plus extraction/consolidation and qwen3.7-text-embedding (1024 dimensions). This used the actual Hermes registry/provider, Hindsight client, native API/worker and pg0 PostgreSQL, with the same pinned plugin revisions above. No synthetic model responses or artificial HTTP/model delays. Neural reranking was disabled via the production rrf passthrough provider.

Plugin retain_async Actual HTTP async Tool return time Returns before HTTP completes
Base true false 10859.70 ms no
Patched true true 1.55 ms yes
Base false false 10135.04 ms no
Patched false false 10209.96 ms no

All four final cases passed. The fixture says that CedarLamp's staging service uses TCP port 48173. An independent client listed the stored facts, recalled that port with types=["world"], and later recalled it with types=["observation"] after consolidation settled.

For the patched async case, the queue drained in 43.55 ms and the real server operation was still pending. The wire trace confirmed that prefetch issued recall only after observing a completed receipt; it returned the newly extracted port fact. The tool's 1.55 ms response means queued, not persisted: retain plus consolidation settled after 24728.31 ms in that run. The sync path still waited for actual retain completion.

Limits: single-fixture, single-run behavior checks, not a model-quality or latency benchmark. Immediate prefetch uses world facts; observation recall is tested after consolidation, not guaranteed by the retain receipt barrier. Neural reranking, a full interactive Hermes agent turn and restart durability remain untested. Only the standalone verification harness needed setup corrections (worker slot reservation and the client's fact_type field); production code did not change. Test processes and temporary banks/homes were cleaned up. No credentials or workspace endpoint are included here.

After merge, the Hermes catalog SHA needs updating to distribute this fix.

@tulioteixeira2020-debug

Copy link
Copy Markdown

Confirming this in production with hermes-plugin-hindsight 1.0.1 (catalog sha 176f8c2d) against a self-hosted Hindsight 0.10.1 (local_external), retain_async: true in config.json.

Observed impact before the fix: manual hindsight_retain tool calls took 13–54 s each in agent.log (tool hindsight_retain completed (36.53s, 41 chars)), i.e. fact extraction ran inline in the request. Auto-retain on the same instance returns in milliseconds.

Why it matters beyond latency: with async=false the server runs retain_batch_async inside the API process, not in the worker. That bypasses:

  • HINDSIGHT_API_WORKER_MAX_SLOTS=1 serialization (a manual retain can run concurrently with a worker retain/consolidation on the same local LLM);
  • any OperationValidatorExtension that defers work via DeferOperation, since deferral only applies on the worker path (RetainContext carries no async flag, so the validator can't tell the two apart).

Validation of this PR's approach locally: applied the _tool_retain change from this PR verbatim; plugin test suite 25/25 plus a parametrized check (retain_async True/False → forwarded as call arg; tool output reports "queued" only when async) all pass. Note the existing test_retain_tool_stores_content_with_per_call_tags needs the shutdown() moved before the assertion, as this PR already does.

Preferring this over #4671: forwarding the flag alone still blocks the tool call on the HTTP round-trip and doesn't register the op for the prefetch visibility barrier.

@yu-xin-c
yu-xin-c force-pushed the codex/hermes-retain-async branch from 1ffd59d to 95d84dd Compare September 27, 2026 06:49
@yu-xin-c

Copy link
Copy Markdown
Author

Rebased onto ccfe85b and regenerated both Hermes documentation outputs. Current head: ba22b0f.

The initial generated-files failure was unrelated Oracle documentation drift from the base, already fixed upstream in ef51970; rebasing incorporates that fix without adding Oracle changes to this PR. The newer README synchronization gate also required regenerating docs-integrations/hermes.md and the corresponding docs-skill reference. Both are now included.

Validation:

  • All 31 Hermes integration tests pass locally on Python 3.11; Ruff and format checks pass.
  • README/doc synchronization, full docs-skill generation and link validation pass. Repeating generation leaves the worktree clean.
  • CI on the current head is green: docs build, Hermes integration tests, live Hermes compatibility, and generated-file verification including the repository-wide lint step.

Thanks for the production validation above. The implementation still queues the HTTP acknowledgement off the tool-call path and tracks server-side operations for the existing prefetch barrier; retain_async=False remains synchronous and reports errors directly.

GitHub reports this head cleanly mergeable. Ready for maintainer review.

@yu-xin-c

Copy link
Copy Markdown
Author

Added a before/after backend integration run to the PR description. This used the actual Hermes provider registry, real Hindsight client over loopback HTTP, a separately launched API/worker, and real pg0 PostgreSQL, not FakeClient or stubbed Hermes modules.

With a controlled HTTP hold followed by an extraction hold, the base (ccfe85b485) configured with retain_async=true sent async=false and blocked for 2065.96 ms. This head sent async=true and returned in 1.54 ms while those boundaries were still held. retain_async=false remained blocking (2058.94 ms). The async receipt was tracked, prefetch sent no recall while extraction was held, and after release an independent client listed and recalled the persisted fact. Both before/after settings stored their expected memory. Processes and temporary databases were cleaned up.

Scope matters: the model/embedding/reranker services were the repository's deterministic HTTP test providers, and the latency numbers are controlled-delay observations, not a production benchmark. The visibility probe explicitly used recall_types=["world"], not the default observation-only filter. No live-model quality, cloud-backend or restart-durability claim. No implementation changes were needed; this adds runtime evidence for the current head.

@yu-xin-c

Copy link
Copy Markdown
Author

Coordination note after the September 30 stacking update on #4671: I checked its current diff. The manual-retain portion now forwards the configured retain_async, but still calls _retain_batch inline, discards the returned operation receipt, and reports "Memory stored successfully." The client-construction locking inherited from #4674 is separate from this behavior.

#4670 overlaps on flag forwarding, but additionally:

  • moves the HTTP acknowledgement off the tool-call path using the existing writer queue;
  • captures the item, bank and policy before enqueueing;
  • registers real operation IDs with the existing prefetch visibility barrier;
  • distinguishes queue acceptance from completed storage, while retaining synchronous errors for retain_async=false.

The before/after real API + PostgreSQL verification is in the PR description and runtime comment, including its deterministic-model and raw-fact-recall limits. This is not a claim that the default observation-only recall filter waits for subsequent consolidation.

These overlapping manual-retain hunks should be reconciled rather than merged twice unchanged. If the smaller flag-forwarding change is preferred first, the writer-queue/receipt-tracking portion remains a separate follow-up; flag forwarding alone does not cover those two boundaries. Please retain those distinctions when deciding whether this PR is superseded.

@yu-xin-c

Copy link
Copy Markdown
Author

Follow-up real-model evidence for ba22b0f9fc, compared with base ccfe85b485: all four final before/after cases passed using live Bailian qwen3.7-plus and qwen3.7-text-embedding (1024 dimensions), real Hermes provider loading, real Hindsight API/worker/client, and pg0 PostgreSQL. No synthetic model replies or artificial delay; neural reranking disabled via production RRF passthrough.

Plugin retain_async HTTP async Tool return time
Base true false 10859.70 ms
Patched true true 1.55 ms
Base false false 10135.04 ms
Patched false false 10209.96 ms

The async response is an enqueue acknowledgement, not a storage-completion claim. Its queue drained in 43.55 ms with the server receipt still pending; prefetch waited for a completed receipt before recalling the newly stored fact. Retain plus consolidation settled after 24728.31 ms in that case.

All four cases independently listed and recalled the fixture's TCP port 48173 as a world fact, and recalled it as an observation after consolidation. Immediate observation-only visibility is not promised by the retain barrier. These are one-shot behavioral checks, not a quality/performance benchmark; neural reranking, a full interactive agent turn and restart durability remain outside scope.

The PR description now separates this live-model evidence from the September 27 controlled-delay test. No production code changes were needed. API/PostgreSQL processes and temporary test data were cleaned up; credentials and workspace endpoint are omitted.

@nicoloboschi

Copy link
Copy Markdown
Collaborator

Thanks for this! We're handling this fix in another change, so I'm closing this one.

We've also stopped accepting pull requests from outside the team. If you run into anything else, please open an issue with steps to reproduce (see CONTRIBUTING.md).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants