Skip to content

fix(agent): surface silent pre-tool-call hook stalls (#32460) - #32613

Open
xxxigm wants to merge 3 commits into
NousResearch:mainfrom
xxxigm:fix/32460-terminal-session-reuse-hang
Open

xxxigm wants to merge 3 commits into
NousResearch:mainfrom
xxxigm:fix/32460-terminal-session-reuse-hang

Conversation

@xxxigm

@xxxigm xxxigm commented May 26, 2026 •

Copy link
Copy Markdown

What does this PR do?

Makes the silent 50-65s wall-clock gap reported in #32460 visible in agent.log instead of looking like the agent froze.

The bug surfaces as a terminal-tool hang on every subsequent call after the first in a CLI/TUI session, while Feishu / gateway sessions run fast. Tool durations in the log show ≤0.5s for trivial commands like ping / echo, but the wall-clock gap between API call #N completed and tool terminal completed clusters tightly around 50-65s — a textbook fixed-timeout pattern.

Root cause: agent/shell_hooks.py runs every configured pre_tool_call hook synchronously on the path between the LLM response and tool execution with DEFAULT_TIMEOUT_SECONDS = 60. When a hook stalls for the full window, subprocess.run waits silently and the bridge only logs a single shell hook timed out WARNING after the timer expires. The matching tool execution then completes in milliseconds, so the user-visible tool ... completed (Xs) line under-reports the stall and there is nothing in the log that names the offending hook. CLI registers shell hooks by default; the gateway side does not unless they are allow-listed — which explains the CLI-only / Feishu-fast split.

Three small additions wired together turn the stall from silent into actionable:

  1. agent/shell_hooks.py — spawn each hook subprocess alongside a daemon watchdog thread. Once the hook has been running for 5 seconds the watchdog emits a WARNING naming the event and command, repeated every 10s until the hook returns. The threshold is well under DEFAULT_TIMEOUT_SECONDS, so a hung hook surfaces immediately instead of after the full 60s window. The watchdog stop signal is wired through finally so success, timeout, and spawn-error paths all clean it up.
  2. agent/tool_dispatch_helpers.py — new pre_tool_call_block_message_with_latency wraps get_pre_tool_call_block_message in a time.monotonic window and logs a WARNING when dispatch exceeds 2s. The warning names the tool so the operator can correlate the stall with a specific invocation regardless of what's blocking (hook, plugin, MCP probe, etc.). The wrapper is purely observational — return value passes through verbatim and exceptions re-raise.
  3. Call-site routing — sequential + concurrent paths in agent/tool_executor.py, invoke_tool in agent/agent_runtime_helpers.py, and the non-skip path in model_tools.handle_function_call all swap to the new wrapper. Every pre_tool_call invocation now contributes to the slow-dispatch warning.

Related Issue

Fixes #32460

Type of Change

  • 🐛 Bug fix (non-breaking change that surfaces a previously-silent stall)

Changes Made

  • agent/shell_hooks.py — add _SLOW_HOOK_THRESHOLD_SECONDS=5 / _SLOW_HOOK_REPEAT_SECONDS=10, new _watch_slow_hook daemon, wire watchdog around subprocess.run in _spawn with try/finally cleanup so the success / TimeoutExpired / spawn-error paths all stop it.
  • agent/tool_dispatch_helpers.py — new pre_tool_call_block_message_with_latency wrapper + _PRE_TOOL_DISPATCH_SLOW_THRESHOLD_SECONDS=2.
  • agent/tool_executor.py — sequential + concurrent paths route through the wrapper.
  • agent/agent_runtime_helpers.py::invoke_tool — routes through the wrapper.
  • model_tools.py::handle_function_call — non-skip path routes through the wrapper.
  • tests/agent/test_shell_hooks_slow_watchdog.py — 5 new tests pinning fast-hook silence, slow-hook warning content, watchdog stop on success, watchdog stop on TimeoutExpired, and spawn-failure silence.
  • tests/agent/test_pre_tool_dispatch_latency.py — 5 new tests pinning fast-dispatch silence, slow-dispatch warning with tool name + elapsed time, block-message pass-through, exception propagation while still logging, and a production-threshold behavioural pin.

Backwards compatible: no public API removed, no schema migration, no new config keys. Hooks that already returned promptly keep behaving identically; the only user-visible change is a WARNING that names the offending hook when it stalls.

How to Test

# New regression suites (10 tests total, ~7s)
.venv/bin/python -m pytest tests/agent/test_shell_hooks_slow_watchdog.py tests/agent/test_pre_tool_dispatch_latency.py -v

# Related areas exercised by the change (131 tests, ~7s)
.venv/bin/python -m pytest tests/agent/test_shell_hooks_consent.py \
    tests/hermes_cli/test_plugins.py \
    tests/test_model_tools.py \
    tests/run_agent/test_tool_call_guardrail_runtime.py \
    tests/run_agent/test_background_review_toolset_restriction.py -q

Behaviour after the fix — agent.log on a hook that stalls 60s:

2026-05-26 14:26:04,070 agent.conversation_loop: API call #3 completed
2026-05-26 14:26:09,073 agent.shell_hooks: shell hook still running after 5.0s
    (event=pre_tool_call command=/home/me/pre-tool-scan.sh timeout=60s); this
    stalls every tool call dispatch until the hook returns
2026-05-26 14:26:19,074 agent.shell_hooks: shell hook still running after 15.0s
    (event=pre_tool_call command=/home/me/pre-tool-scan.sh timeout=60s)
2026-05-26 14:26:29,074 agent.shell_hooks: shell hook still running after 25.0s
    (event=pre_tool_call command=/home/me/pre-tool-scan.sh timeout=60s)
…
2026-05-26 14:27:04,071 agent.shell_hooks: shell hook timed out after 60.00s
    (event=pre_tool_call command=/home/me/pre-tool-scan.sh)
2026-05-26 14:27:04,072 agent.tool_dispatch_helpers: pre_tool_call dispatch
    took 60.0s for tool 'terminal' — a slow shell hook or plugin is stalling
    the agent loop. Check ~/.hermes/config.yaml 'hooks.pre_tool_call' entries
    and recently installed plugins.
2026-05-26 14:27:04,219 agent.tool_executor: tool terminal completed (0.15s, 42 chars)

The 50-65s wall-clock gap is now named, timestamped, attributed to a specific tool, and traceable to a specific hook command — operator can act on it instead of guessing.

Checklist

  • Conventional Commits (feat(shell_hooks):, feat(agent):, refactor(dispatch):)
  • 3 focused commits, single author (xxxigm)
  • 10 new tests pass; 131 related existing tests pass; pre-existing macOS flakes (test_anthropic_adapter.py::TestRunOauthSetupToken, test_shell_hooks.py::TestCallbackSubprocess::test_timeout_returns_none, test_vision_routing_31179.py under contention) confirmed unchanged on upstream/main
  • Tested on macOS 15.6 (darwin 24.6.0), Python 3.12.5
  • No new config keys, no architecture change, no platform-specific calls
  • Backwards compatible — well-behaved hooks stay silent

xxxigm added 3 commits May 26, 2026 20:08
…Research#32460)

A configured ``pre_tool_call`` hook runs synchronously on the hot path
between an LLM response and tool execution.  When the hook stalls for
the full ``DEFAULT_TIMEOUT_SECONDS = 60`` window it appears to the
operator as a silent 50-65s wall-clock gap with no log line at all —
only a single ``shell hook timed out`` WARNING fires *after* the timer
expires.  This is exactly the symptom reported in NousResearch#32460.

Spawn the subprocess alongside a daemon watchdog thread that emits a
WARNING once the hook crosses ``_SLOW_HOOK_THRESHOLD_SECONDS`` (5s) and
repeats every ``_SLOW_HOOK_REPEAT_SECONDS`` (10s) until the hook
returns.  The warning names the event and command so the operator can
identify the offending script without waiting for the full timeout.

Cleanup is wired through ``finally`` so the watchdog stops on every
exit path — success, timeout, or spawn error.
…search#32460)

The agent calls ``get_pre_tool_call_block_message`` for every tool
invocation before the tool starts measuring its own duration.  A slow
plugin or shell hook on that path produces a silent wall-clock gap
between an LLM response and the resulting tool execution — the
``tool ... completed (Xs)`` line reports a tiny duration while the
real stall happened upstream.

Add ``pre_tool_call_block_message_with_latency`` to wrap the call in a
``time.monotonic`` window and emit a single WARNING when the dispatch
exceeds ``_PRE_TOOL_DISPATCH_SLOW_THRESHOLD_SECONDS`` (2s).  The
warning names the tool so the operator can correlate the stall with a
specific invocation, and the helper preserves the underlying return
value (including block directives) and re-raises exceptions verbatim
so it stays purely observational.

Threshold lives next to the helper as a module-level constant with a
behavioural-pin test so it cannot drift back toward the 60s window
that produced the original symptom.
…Research#32460)

Sequential + concurrent tool executors, ``invoke_tool`` in the runtime
helpers, and the non-skip path in ``model_tools.handle_function_call``
all called ``get_pre_tool_call_block_message`` directly.  Swap each
site to the latency-aware wrapper so every code path that fires a
``pre_tool_call`` hook contributes to the slow-dispatch WARNING.

No behaviour change for fast hooks; slow ones now emit a single
WARNING that names the offending tool instead of producing a silent
multi-second wall-clock gap.
@alt-glitch alt-glitch added type/bug Something isn't working P2 Medium — degraded but workaround exists comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint comp/tools Tool registry, model_tools, toolsets labels May 26, 2026

@teknium1 teknium1 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for targeting an otherwise silent hook-stall path. Current main still blocks synchronously in agent/shell_hooks.py:462-474, but this branch predates the current pre-tool approval contract and needs a deliberate salvage.

Problems

  • agent/tool_dispatch_helpers.py:78 wraps deprecated get_pre_tool_call_block_message(). Current main's dispatch chokepoint is resolve_pre_tool_block() (hermes_cli/plugins.py:2226-2274), which escalates approve directives and fails closed. The wrapper must preserve that behavior.
  • agent/shell_hooks.py:396 says a slow hook stalls every tool dispatch, although _spawn() is used for all hook events; tests/agent/test_shell_hooks.py:402-417 exercises pre_llm_call.
  • #32460 reports terminal-session reuse and only hypothesizes synchronization/readiness polling. It does not establish configured shell hooks as the cause.

Suggested changes

  • Time resolve_pre_tool_block() at the four current dispatch sites and forward the full context metadata.
  • Test approval denial/error through the instrumented path, and make watchdog messaging event-accurate.

Automated hermes-sweeper review.


t0 = time.monotonic()
try:
return get_pre_tool_call_block_message(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Current main makes get_pre_tool_call_block_message() a deprecated block-only shim (hermes_cli/plugins.py:2201-2223); live dispatch uses resolve_pre_tool_block() to enforce approve directives fail-closed. Salvage this wrapper around resolve_pre_tool_block() and forward the full dispatch metadata, otherwise approval directives can be treated as allow.

Comment thread agent/shell_hooks.py
logger.warning(
"shell hook still running after %.1fs "
"(event=%s command=%s timeout=%ss); "
"this stalls every tool call dispatch until the hook returns",

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

_spawn() is shared by every shell-hook event, including pre_llm_call, so this wording is inaccurate outside pre_tool_call. Either limit this watchdog to pre-tool hooks or describe the specific event that is currently blocked.

@teknium1 teknium1 added sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades sweeper:blast-broad Sweeper blast radius: broad — a core path most sessions hit labels Jul 13, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint comp/tools Tool registry, model_tools, toolsets P2 Medium — degraded but workaround exists sweeper:blast-broad Sweeper blast radius: broad — a core path most sessions hit sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: terminal tool hangs 50-65s when reusing existing session in CLI/TUI

3 participants