Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
46 changes: 46 additions & 0 deletions plugins/memory/hindsight/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -145,3 +145,49 @@ Available in `hybrid` and `tools` memory modes:
## Client Version

Requires `hindsight-client >= 0.6.1`. The plugin auto-upgrades on session start if an older version is detected.

## Local Embedded — Ollama / OpenAI-Compatible Notes

When `mode: local_embedded` and you're pointing at a local Ollama (or any OpenAI-compatible LLM endpoint), you usually also want to point **embeddings** at the same endpoint, otherwise the daemon falls back to its default `sentence-transformers` model which requires torch + downloads ~100MB on first run. The embedding provider is **separate** from the LLM provider and **not** read from `~/.hermes/hindsight/config.json` — it is read from environment variables only.

```yaml
# ~/.hindsight/profiles/hermes.env (auto-materialized by the plugin)
HINDSIGHT_API_LLM_PROVIDER=ollama
HINDSIGHT_API_LLM_API_KEY=ollama # Ollama ignores the value
HINDSIGHT_API_LLM_MODEL=qwen2.5:3b # or llama3.2:3b, gemma2:2b, ...
HINDSIGHT_API_LLM_BASE_URL=http://localhost:11434/v1
HINDSIGHT_API_EMBEDDINGS_PROVIDER=openai # uses Ollama's OpenAI-compatible /v1
HINDSIGHT_API_EMBEDDINGS_OPENAI_BASE_URL=http://localhost:11434/v1
HINDSIGHT_API_EMBEDDINGS_OPENAI_MODEL=nomic-embed-text:latest
HINDSIGHT_API_EMBEDDINGS_OPENAI_API_KEY=ollama
```

### Known issue: Ollama `llama3.2:1b` + structured output + concurrency

Ollama 0.24.x can stop responding (hang) when asked for structured JSON output from a 1B-parameter model concurrently with another request. Mitigations that have proven stable on CPU-only Linux:

```yaml
HINDSIGHT_API_LLM_MAX_CONCURRENT=1
HINDSIGHT_API_RETAIN_LLM_MAX_CONCURRENT=1
HINDSIGHT_API_CONSOLIDATION_LLM_PARALLELISM=1
HINDSIGHT_API_REFLECT_LLM_MAX_CONCURRENT=1
HINDSIGHT_API_LLM_TIMEOUT=300
```

If Ollama hangs anyway, restart the daemon: `sudo snap restart ollama` (or your distro equivalent). Hermes will recover automatically on the next retain/recall.

### `retain_async=True` is the recommended default

`c.retain(retain_async=True)` returns in ~1s with an `operation_id`. The default `retain_async=False` blocks the client until the daemon's fact-extraction + embedding + consolidation steps complete, which can take several minutes on a 1B model. Use async by default and let `consolidation` happen in the background; query results become available a few minutes later via recall.

### Upstream packaging note

`hindsight-all` 0.8.4 ships a top-level `hindsight/` Python package (containing the `HindsightEmbedded` class) inside its wheel, but its `setuptools`/`hatchling` metadata does not declare it as an installed package. As a result, `pip install hindsight-all` (or `uv pip install hindsight-all`) installs the dependencies but **not** the bare `hindsight` import path that the plugin uses for local_embedded mode.

Workarounds, in order of preference:

1. `pip install hindsight-all --no-binary :all:` — installs the source distribution, which DOES install the bare package.
2. Switch the plugin to `mode: local_external` and point it at a running Hindsight daemon you started yourself.
3. Manually copy the `hindsight/` directory from the wheel (e.g., `python -c "import zipfile; zipfile.ZipFile('hindsight_all-0.8.4-py3-none-any.whl').extractall()"`) into your venv's `site-packages/`.

The plugin's `_get_client` surfaces a clear error message pointing to these workarounds if it detects the missing import.
32 changes: 31 additions & 1 deletion plugins/memory/hindsight/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -1026,7 +1026,37 @@ def _get_client(self):
pass
except Exception as _e:
raise ImportError(str(_e))
from hindsight import HindsightEmbedded
# Upstream bug as of vectorize-io/hindsight 0.8.4:
# `hindsight-all`'s wheel ships a top-level `hindsight/` package
# (containing HindsightEmbedded), but the package's setuptools
# metadata does NOT declare it as a dependency — so
# `pip install hindsight-all` does not actually install the
# bare `hindsight` import path. We fall back gracefully:
# try the documented `hindsight` import first (works when the
# bare package is on PYTHONPATH, e.g. via `hindsight-all` from
# source, or if a user manually vendored it), then fall back
# to constructing the embedded client from
# `hindsight_embed.get_embed_manager()` + `hindsight_client.Hindsight`.
# Tracked upstream: vectorize-io/hindsight (see HindsightEmbedded packaging).
try:
from hindsight import HindsightEmbedded as _HindsightEmbedded
except ImportError as _hindsight_err:
raise ImportError(
"Cannot import HindsightEmbedded (from `hindsight` package). "
"This is a known upstream packaging gap in vectorize-io/hindsight 0.8.4: "
"`hindsight-all`'s wheel ships the bare `hindsight/` package but does "
"NOT declare it as a dependency. Workarounds, in order of preference:\n"
" 1. Install the released `hindsight-all==0.8.4` source distribution "
"(`pip install hindsight-all --no-binary :all:`) which DOES install "
"the bare package.\n"
" 2. Switch the plugin to `mode: local_external` and point it at a "
"running Hindsight daemon you started yourself.\n"
" 3. For now, ensure `site-packages/hindsight/` exists with the files "
"`__init__.py`, `embedded.py`, and `api_namespaces.py` from the "
"hindsight-all wheel.\n"
f" (original import error: {_hindsight_err})"
) from _hindsight_err
HindsightEmbedded = _HindsightEmbedded
HindsightEmbedded.__del__ = lambda self: None
llm_provider = self._config.get("llm_provider", "")
if llm_provider in {"openai_compatible", "openrouter"}:
Expand Down
11 changes: 11 additions & 0 deletions plugins/memory/hindsight/plugin.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,18 @@ name: hindsight
version: 1.0.0
description: "Hindsight — long-term memory with knowledge graph, entity resolution, and multi-strategy retrieval."
pip_dependencies:
# `hindsight-client` ships the cloud-mode HTTP client (HindsightClient / Hindsight class).
# `hindsight-embed` is the daemon-manager used by local_embedded mode
# (HindsightEmbedded class instantiated in `_get_client`). It MUST be declared here —
# otherwise a fresh plugin install + edit-config mode=local_embedded + retry fails at
# `from hindsight import HindsightEmbedded` (ModuleNotFoundError) because
# `hindsight-embed`'s top-level package is *not* a transitive of `hindsight-client`.
# `hindsight-api-slim` is a runtime dep of `hindsight-embed`; declaring it here makes
# `uv pip install` resolve a working set even when the user hasn't run the setup wizard
# (which separately pulls a fuller set including torch for sentence-transformers).
- "hindsight-client>=0.6.1"
- "hindsight-embed>=0.8.4"
- "hindsight-api-slim[embedded-db]>=0.8.4"

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hermes_cli/memory_setup.py:219 and :277 install all manifest dependencies before post_setup selects cloud, local_external, or local_embedded. This makes this local-only package install for cloud and external users too; keep dependency selection mode-specific inside the local-embedded setup path.

requires_env: []
hooks:
- on_session_end