Repository navigation
Conversation
`_EXTERNAL_PREFETCH_TIMEOUT_S = 8.0` was hardcoded, but the right value depends entirely on the provider's recall latency, which the core cannot know. Expose it as `memory.external_prefetch_timeout` (default 8.0, so existing behaviour is unchanged) and pass it through from agent_init. This is not just a "recall arrives late" nicety. In MemoryManager._prefetch_external (memory_manager.py:611-632) a timed-out worker thread is NOT cancelled: join() gives up, the call returns "", and the thread keeps running. The next turn sees that thread still alive and skips prefetch entirely. So when a provider's recall reliably exceeds the timeout, prefetch is not merely delayed — it is never applied at all, and the only symptom is a repeating warning. Observed with the Hindsight provider on a bank whose local CPU cross-encoder reranked 300 candidates per recall: recall measured 15-22s against the 8s ceiling, and 135 warnings accumulated with prefetch effectively disabled the whole time. Raising the timeout restored it immediately. Settings belong in config.yaml rather than an env var, and a bad value must not break startup, so a non-numeric or non-positive value falls back to the existing default instead of raising. MemoryManager's own positive-value guard is left intact for direct callers. Tests: tests/agent/test_memory_prefetch_timeout_config.py covers the resolution chain (configured float/int/string, absent, None, zero, negative, garbage, empty config), asserts the documented default matches the runtime fallback so the two cannot drift, and guards the real agent_init call site — the parametrised cases mirror agent_init's logic and would stay green if it reverted to a bare MemoryManager(), so that case is asserted separately. Verified by mutation: breaking the wiring fails the guard, restoring it passes. 88 tests pass alongside the existing memory suite.
|
I've posted the field evidence (Bedrock-only Docker gateway with a local Ollama embedder: 8s timeouts coinciding with 18.4-47.1s embeds and a 400 after 60s, 101/197 agent starts with no injected recall) on #87028, since the triage bot already flagged this PR as a duplicate of it and #87028 is the earlier, more complete head. Two of the same gaps apply here: |
|
Closing as a duplicate of open #87028, which covers the same external_prefetch_timeout registration/handoff with stronger validation. This is not fixed on main. The keeper still needs rebase, actual _init_memory wiring tests, sample/docs, and coordination with provider-specific defaults. Preserve configured/invalid/missing-value tests in that work. |
Problem
agent/memory_manager.pyhardcodes_EXTERNAL_PREFETCH_TIMEOUT_S = 8.0, but the correct value depends on the external memory provider's recall latency — something the core cannot know.This is not only a "memory arrives late" issue. In
MemoryManager._prefetch_external(memory_manager.py:611-632) a timed-out worker thread is not cancelled:The thread keeps running, and lines 611-619 skip prefetch on any later turn while it is still alive. So when a provider's recall reliably exceeds the timeout, prefetch is never applied at all — not merely delayed. The only symptom is a repeating warning, which reads like noise rather than "your memory provider is switched off."
Observed with the Hindsight provider against a bank whose local CPU cross-encoder reranked 300 candidates per recall: recall measured 15-22s against the 8s ceiling, 135 warnings accumulated, and prefetch was effectively disabled for the entire period. Raising the timeout restored it immediately.
Change
Adds
memory.external_prefetch_timeout(default8.0, so existing behaviour is unchanged) and threads it fromagent_initintoMemoryManager.Deliberate choices:
config.yaml, not an env var — this is a behavioural setting, not a credential.MemoryManager's own positive-value guard is left intact for direct callers.Tests
tests/agent/test_memory_prefetch_timeout_config.py:None, zero, negative, garbage, empty config.config_defaults.pymatches_EXTERNAL_PREFETCH_TIMEOUT_S, so the two cannot silently drift.agent_initcall site. The parametrised cases mirroragent_init's resolution logic, which means they would stay green ifagent_initreverted to a bareMemoryManager()— so that case is asserted separately against the source.Mutation-verified: breaking the wiring fails the guard test (exit 1); restoring it passes (exit 0). 88 tests pass alongside the existing
tests/agent/test_memory_provider.pysuite.Verification
Exercised on a live install rather than only in tests: set to
15, restarted the gateway, confirmed the value resolves throughload_config()→MemoryManageras15.0, and confirmed realhindsight_recallcalls now complete instead of being discarded. Zero timeout warnings from sessions created after the restart.One behaviour worth noting for reviewers: the timeout is read when a session's agent is constructed, so already-running long-lived sessions keep their previous value until they end. That is consistent with how other
memory.*settings behave.