Skip to content

feat(hooks): pre_tool_call content transformation - #28953

Closed
NikolaRHristov wants to merge 24 commits into
NousResearch:mainfrom
NikolaRHristov:feat/pre-tool-call-content-transform
Closed

feat(hooks): pre_tool_call content transformation#28953
NikolaRHristov wants to merge 24 commits into
NousResearch:mainfrom
NikolaRHristov:feat/pre-tool-call-content-transform

Conversation

@NikolaRHristov

@NikolaRHristov NikolaRHristov commented May 19, 2026

Copy link
Copy Markdown
Contributor

Summary

Adds a modify response type to pre_tool_call hooks, letting a hook
transform tool arguments before the tool executes rather than repairing the
result afterwards via post_tool_call. This removes the mtime-warning cycle
that occurs when a post-call hook rewrites a file the agent has already written.

Note

This PR previously also registered shell hooks in the TUI and ACP entry
points. That half has been removed in 539974b
so it no longer overlaps #41464 and the _make_agent cluster — see
#41457. What remains is the
argument-rewriting feature alone.


Wire protocol

A hook signals a rewrite on stdout. Both forms are accepted, so existing Claude
Code-style hooks work unchanged:

hook stdout

{"action": "modify", "args": {...}}

hook stdout (Claude Code-compatible)

{"decision": "modify", "tool_input": {...}}

The returned mapping is merged into the existing arguments — a hook can
rewrite one field without restating the rest.


Design

The central constraint is that pre_tool_call must fire exactly once per
tool execution. Fetching a block decision and a modification separately would
double-fire every hook and double-count every observer, so both travel through a
single call.

hermes_cli.plugins._dispatch_pre_tool_call_hooks() is that call. It invokes
the hooks once and returns (block_message, modified_args), reusing the same
approval-gate logic as resolve_pre_tool_block — including its fail-closed
behaviour, where an approve directive whose gate errors, denies, or times out
becomes a block.

The three dispatch sites migrate onto it:

Call site Path
model_tools.py handle_function_call
agent/tool_executor.py sequential/concurrent tool dispatch
agent/agent_runtime_helpers.py runtime helper dispatch

Existing entry points are untouched. get_pre_tool_call_block_message() and
resolve_pre_tool_block() keep their current behaviour and signatures; callers
that only need block detection need no changes. modify is purely additive — a
hook that never emits it behaves exactly as before.


Files changed

+288/−23 across 9 files.

File Change
hermes_cli/plugins.py _dispatch_pre_tool_call_hooks() — single invocation point returning (block_message, modified_args)
agent/shell_hooks.py _parse_response() accepts both modify wire formats
model_tools.py dispatch site migrated; applies returned modifications
agent/tool_executor.py dispatch site migrated
agent/agent_runtime_helpers.py dispatch site migrated
tests/hermes_cli/test_plugins.py +117 — modify dispatch, merge semantics, precedence
tests/agent/test_shell_hooks.py +28 — response parsing for both formats
website/docs/user-guide/features/hooks.md documents the modify action
docs/observability/README.md observer metadata note

Testing

Terminal

pytest tests/hermes_cli/test_plugins.py tests/agent/test_shell_hooks.py

66 passing.

Beyond the unit tests, the feature was verified end-to-end against a live agent
rather than a stubbed dispatcher. A pre_tool_call plugin hook returns
{"action": "modify", ...} rewriting a write_file path from DECOY.txt to
REWRITTEN.txt and prepending a banner to the content. The agent is told only
to write DECOY.txt and is never informed of the rewrite, so the outcome is a
filesystem fact rather than a log assertion:

this branch main
hook fires
hook returns modify
REWRITTEN.txt created, banner present
DECOY.txt created

The control is the informative half: on main the hook fires and returns the
identical directive, and main discards it, having no code to honour the
action. Same plugin, same prompt, same sandbox — only the codebase differs.

Also verified: blocking still works (block and the approve escalation path),
and get_pre_tool_call_block_message() is unchanged for existing callers.


Related

@NikolaRHristov

Copy link
Copy Markdown
Contributor Author

Related: issue #18148 requests a full Python-level runtime extension hook system (intercept_tool_call with pass | block | rewrite semantics). This PR delivers the rewrite / modify half of that at the shell hook layer — the existing, already-deployed mechanism — without introducing a new extension architecture. If #18148 lands, the two approaches are complementary: shell hooks for lightweight script-based transforms, Python-level hooks for lower-latency, stateful extension logic.

@alt-glitch alt-glitch added type/feature New feature or request P2 Medium — degraded but workaround exists comp/tui Terminal UI (ui-tui/ + tui_gateway/) comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint comp/plugins Plugin system and bundled plugins labels May 19, 2026
@alt-glitch

Copy link
Copy Markdown
Collaborator

Related: part 2 of this PR (TUI gateway hook registration) subsumes #13854 and #24237 which address the same missing register_from_config call in tui_gateway/entry.py.

@NikolaRHristov

Copy link
Copy Markdown
Contributor Author

Summary of changes + follow-up recommendations

What's done in this PR

agent/shell_hooks.py — extends pre_tool_call with a "modify" decision alongside the existing "block"/"continue". A hook script can now return {"decision": "modify", "args": {...}} to rewrite tool arguments before dispatch. Two new internals power this:

  • _parse_response(event, stdout) — normalises hook stdout into a canonical {decision, reason, args} dict; accepts the legacy action key and the Claude-Code-style message alias; returns None on empty/non-JSON output.
  • _dispatch_pre_tool_call_hooks(specs, tool_name, args) — returns (block_message: str | None, modified_args: dict); last writer wins per key on "modify"; block takes precedence over modify.

get_pre_tool_call_block_message() is now a thin wrapper — its signature and return type are unchanged, so neither call site in tool_executor.py nor agent_runtime_helpers.py requires modification.

tui_gateway/entry.py — fixes the bootstrap sequence: hook registration was firing at module import time, before the transport was wired and before Python plugins were discovered. Moved into main() in the correct order (discover_plugins()register_from_config()), mirroring gateway/run.py. Added a module-level logger so failures produce logger.debug(exc_info=True) instead of silent pass.


Follow-up PRs recommended

A (highest value): Wire modified_args through the two call sites in tool_executor.py and agent_runtime_helpers.py. Hook-driven argument rewriting is fully implemented on the hook side but silently discarded by callers until this lands. The public API is already shaped for it — small, isolated change.

B: Analogous "modify" decision for post_tool_call shell hooks (rewrite tool output), bringing shell-hook parity with the Python plugin transform_tool_result surface.

C: Unit tests for _parse_response and _dispatch_pre_tool_call_hooks — all five wire-protocol shapes, last-writer-wins merge, block-over-modify precedence. tests/agent/test_shell_hooks.py is the correct home.

D: Update website/docs/user-guide/features/hooks.md, which currently documents pre_tool_call as block-only. The "modify" decision shape and merge semantics need a documentation entry.

pre_tool_call hooks can now return a modify directive to transform tool arguments before dispatch, not just block execution. This closes the gap where observability-only hooks had no way to correct or rewrite outgoing calls.

Added _dispatch_pre_tool_call_hooks() as the canonical single-fire entry point. It returns a (block_message, modified_args) tuple: block preserves the existing "first block wins" semantics, while modify performs a shallow merge of the returned args over the original function_args, letting the first modifier win. Existing get_pre_tool_call_block_message() now delegates to _dispatch_pre_tool_call_hooks() for backward compatibility.

Updated all dispatch sites to apply modified_args when present:
- agent/agent_runtime_helpers.py invoke_tool()
- agent/tool_executor.py execute_tool_calls_concurrent() and execute_tool_calls_sequential()
- model_tools.py handle_function_call()

agent/shell_hooks._parse_response() now recognizes both canonical format ({"action":"modify","args":{...}}) and Claude-Code-style format ({"decision":"modify","tool_input":{...}}) so shell-based hooks can participate in the same contract.

tui_gateway now registers discovered Python plugins and declarative shell hooks during startup and before each agent build, matching the CLI and messaging gateway initialization sequence so TUI sessions observe pre_tool_call/post_tool_call hooks by default when consent is configured.
@NikolaRHristov
NikolaRHristov force-pushed the feat/pre-tool-call-content-transform branch from 5cee156 to 7ce97af Compare May 30, 2026 15:42
…e TUI bootstrap registration

In agent/tool_executor.py, add a _ts_scope_block guard in execute_tool_calls_concurrent(): when the current task scopes out the target tool, reject the call immediately before pre_tool_call hooks and guardrails run. Previously hooks fired unconditionally, allowing modification logic to act on calls that should have been rejected at the boundary.

In tui_gateway/entry.py, drop the eager discover_plugins() and register_from_config() step from TUI bootstrap. Hook and plugin initialization now belongs to the agent build sequence (matching CLI and messaging gateway timing), eliminating the brittle config/TTY dependency created by duplicating this registration at process startup.
…sequential execution

Add a _ts_scope_block guard in execute_tool_calls_sequential() that resolves tool_search calls to their underlying tool and validates session scope before pre_tool_call hooks and guardrails run. Previously, hooks fired unconditionally on resolved calls, allowing modify logic to rewrite arguments for tools that should have been rejected at the scope boundary.

The new logic intercepts TOOL_CALL_NAME, resolves the underlying tool via tool_search.resolve_underlying_call(), and checks it against the agent's scoped tool names. If the underlying tool is out of scope, execution is blocked with an error directing the model to use tool_search to find available tools. A _effective_block is composed from either the scope block or the plugin hook block message, and is used consistently for guardrail bypass and error synthesis, ensuring scope enforcement happens before any transformation or guardrail evaluation.
…shot

Add a _mcp_discovery_thread module-level variable to track the background MCP tool-discovery thread. The startup sequence can now briefly join this thread during the first agent build, ensuring fast-spawning MCP servers complete registration before the agent snapshots its tool list.

@teknium1 teknium1 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for putting this together — both parts address real gaps on current main.

Problems

  • hermes_cli/plugins.py:1691 adds _dispatch_pre_tool_call_hooks() with only task_id/session_id/tool_call_id. Current main’s get_pre_tool_call_block_message() also passes turn_id, api_request_id, and middleware_trace into invoke_hook() (hermes_cli/plugins.py:1866-1895), so a salvage needs to preserve those observer fields.
  • hermes_cli/plugins.py:1743-1745 rebuilds each modify result from the original args, so multiple modify hooks do not actually merge per key; a later partial modify drops keys changed by earlier hooks.
  • agent/tool_executor.py:592 in the sequential path still dispatches pre_tool_call hooks even when _ts_scope_block is set, then blocks afterward via _effective_block. The stated security property is to reject out-of-scope tool_search calls before hooks run.
  • The sibling ACP path still lacks hook registration on current main (acp_adapter/session.py:571-623), while the PR summary claims support across all execution paths.

Suggested changes

  • Preserve current-main hook metadata in the new dispatcher.
  • Add tests for modify parsing, merge semantics, block precedence, and dispatch-time argument rewriting.
  • Update hook docs; current docs still describe pre_tool_call as block-only (docs/observability/README.md:54, website/docs/user-guide/features/hooks.md:1204-1215).

Automated hermes-sweeper review.

Comment thread hermes_cli/plugins.py
Comment thread hermes_cli/plugins.py Outdated
Comment thread agent/tool_executor.py Outdated
@teknium1 teknium1 added sweeper:risk-security-boundary Sweeper risk: may affect sandboxing, auth, credentials, or sensitive data sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades sweeper:blast-contained Sweeper blast radius: contained — one narrow path / opt-in / few users labels Jul 13, 2026
…ntent-transform

Signed-off-by: Nikola Hristov <Nikola@PlayForm.Cloud>
…from hooks

Introduce a new `modify` action for `pre_tool_call` hooks, allowing plugins and shell hooks to transform tool arguments before dispatch rather than only blocking or approving execution. This closes the gap where observability-only hooks had no mechanism to correct or rewrite outgoing calls.

Key changes:
- Added `_dispatch_pre_tool_call_hooks()` in `hermes_cli/plugins.py` as the canonical single-fire entry point for pre_tool_call hooks. It returns a `(block_message, modified_args)` tuple, replacing the previous exclusive focus on block detection.
- Extended `_PreToolCallDirective` with a `modified_args` field to carry transformed arguments through the directive resolution chain.
- Updated `_get_pre_tool_call_directive_details()` to recognise a `modify` action (both canonical `{"action":"modify","args":{...}}` and Claude-Code-style `{"decision":"modify","tool_input":{...}}` formats). The first modifier wins, performing a shallow merge of the returned args over the original `function_args`.
- Updated all dispatch sites to apply `modified_args` when present:
  - `agent/agent_runtime_helpers.py:invoke_tool()`
  - `agent/tool_executor.py:execute_tool_calls_concurrent()` and `execute_tool_calls_sequential()`
  - `model_tools.py:handle_function_call()`
- Removed duplicate plugin discovery and shell-hook registration from `tui_gateway/entry.py` startup. These registrations now belong solely to the agent build sequence, matching CLI and messaging gateway timing and eliminating the brittle config/TTY dependency.

Closes: #<feature-gap>
…dispatch

Extend _dispatch_pre_tool_call_hooks() invocations across all three dispatch sites with additional routing metadata — session_id, tool_call_id, turn_id, api_request_id, and middleware_trace — enabling hooks to make context-aware decisions about tool argument modification.

Previously, modify-oriented hooks had no visibility into which turn, session, or API request triggered the tool call, limiting their ability to apply targeted transformations or trace routing paths.

Changes:
- agent/agent_runtime_helpers.py: Pass routing context and middleware_trace to _dispatch_pre_tool_call_hooks() in invoke_tool().
- agent/tool_executor.py: Add the same context in execute_tool_calls_concurrent() and execute_tool_calls_sequential(), propagating middleware_trace where available.
- hermes_cli/plugins.py: Fix modify-accumulation logic in _get_pre_tool_call_directive_details() — instead of resetting merged args on each modifier, build from the original args on first hit and shallow-merge subsequent modifiers into the accumulator, preserving original keys across multiple modify directives.
…ormation

Update observability README and website hooks documentation to describe
the new `modify` action on `pre_tool_call` hooks, alongside the existing
`block` action. Cover both the Hermes-canonical and Claude-Code-compatible
return formats, explain shallow-merge accumulation semantics, and update
use-case descriptions to include argument sanitization and path rewriting.

Add comprehensive unit tests in both `test_shell_hooks.py` (shell hook
callback parsing for canonical and Claude-Code modify formats) and
`test_plugins.py` (Python plugin modify dispatch covering single modify,
multi-hook accumulation, last-writer-wins on same key, modify+block
interaction, block-first precedence, None args handling, and invalid
args filtering).

This completes the documentation and test coverage for the modify
feature introduced in the parent commit.
…l execution path

Fix a NameError in execute_tool_calls_sequential() where the middleware_trace variable was referenced in the post-call emission path but never declared in that scope. The concurrent execution path initialised it earlier via upstream routing context propagation, but the sequential path — which does not receive middleware_trace from its callers — left the variable undefined.

This commit adds an explicit empty-list initialisation for middleware_trace in the sequential path, ensuring the variable exists in scope for downstream consumers that expect it without altering the existing behaviour (the sequential path intentionally keeps middleware_trace empty).
…ution path

Distinguish between plugin-level blocks and tool-scope blocks in the sequential tool execution path by introducing a dedicated `_block_error_type` variable. Previously, both block sources used the same default error type, making it impossible to determine from the error context whether a tool call was blocked by a plugin hook or by a tool-scope restriction.

This commit adds `_block_error_type` initialisation alongside the existing `_block_msg` variable: it defaults to `"plugin_block"` and is overridden to `"tool_scope_block"` when the block originates from the tool-scope check. This enables downstream error handlers and observability hooks to produce accurate, source-specific diagnostics without requiring additional parsing of the block message text.
@NikolaRHristov

Copy link
Copy Markdown
Contributor Author

Hey @teknium1, thank you for the thorough review and sweep. All three points are
addressed below.

Missing hook metadata

90a83ea

_dispatch_pre_tool_call_hooks already had turn_id, api_request_id,
middleware_trace in its signature. The gap was at the call sites. All four
now pass the full metadata they have available, matching what
resolve_pre_tool_block sends to invoke_hook today.

Modify merge accumulation

90a83ea

Changed from rebuilding dict(args) on every modify hit to building once on
first hit then .update()ing subsequent partials. Hook A setting path and
hook B setting content now both survive. Same-key collisions still last-wins
(standard shallow merge).

Sequential scope gate

The hook dispatch was already wrapped in the else: branch, so out-of-scope
tool_search calls never reach _dispatch_pre_tool_call_hooks. This matches
the concurrent path's guard exactly. Applied during the re-merge in
2d55075.

Docs and tests

3922b7a

Added 10 tests and updated the docs. 8 modify-semantics tests in
test_plugins.py, 2 shell hook parsing tests in test_shell_hooks.py, and
doc updates to docs/observability/README.md and
website/docs/user-guide/features/hooks.md documenting modify alongside
block.

Merge artifact fixes

Two issues from the origin/main merge that broke the sequential executor
were discovered during live testing and fixed:

  • 27fe63e
    The merge accidentally dropped middleware_trace initialization from
    execute_tool_calls_sequential, causing NameError on every tool call.
    Restored with middleware_trace: list[dict[str, Any]] = [].

  • 3e60ab2
    Same merge dropped _block_error_type disambiguation ("plugin_block" vs
    "tool_scope_block") which was still referenced downstream. Restored.

Live test results

Tested end-to-end on the feat branch with a pre_tool_call shell hook that
prepends a diagnostic banner to every terminal command:

  • {"action": "modify", "args": {"command": "# [pre_tool_call hook modified] ..."}} correctly transforms the command before execution. Banner visible in live terminal output.
  • Non-terminal calls (read_file, write_file, patch, process, aphrodite_*) pass through the hook unmodified. The hook returns {} for them.
  • Observer payload includes full metadata: tool_name, tool_input, session_id, cwd, extra with task_id, tool_call_id, turn_id, api_request_id, middleware_trace, telemetry_schema_version.
  • 10/10 unit tests pass (TestPreToolCallModify plus shell hook parsing tests).
  • No tools fail under normal operation.

Let me know if anything else needs work.

ACP sessions that use shell hooks (e.g., pre/post_tool_call hooks) lost
them entirely because the ACP adapter's agent factory never called
register_from_config().  The CLI (cli.py) and gateway (gateway/run.py)
already wire hooks in before creating their agents, but the ACP adapter
path was missed, creating an inconsistency where hooks worked in some
interfaces but not others when the same config was used.

This commit adds a register_from_config call to the ACP adapter's
_create_agent_factory path, using accept_hooks=False so consent
resolution follows the standard path (HERMES_ACCEPT_HOOKS env var
and hooks_auto_accept config setting), matching every other entry
point in the codebase.  Failure is non-fatal — a failed hook load
is logged at DEBUG level with traceback, and agent creation proceeds
normally.
@NikolaRHristov

Copy link
Copy Markdown
Contributor Author

ACP adapter hook registration

111c625

The sweeper correctly flagged that acp_adapter/session.py also lacked
register_from_config. Current state on main before this fix:

Entry point Hook registration Status
hermes_cli/main.py register_from_config at startup Works
gateway/run.py register_from_config at startup Works
tui_gateway/server.py register_from_config in _make_agent Fixed in PR
acp_adapter/session.py None Fixed here

The ACP adapter loads config in _make_agent() (line 607) but never called
register_from_config, so hooks were silently absent in ACP sessions. Added the
same try/except-guarded call with accept_hooks=False (consent resolved from
HERMES_ACCEPT_HOOKS and hooks_auto_accept) matching the TUI fix.

Note: acp_adapter/entry.py main() also does not register hooks, following the
per-agent-build model established by commit 2 (which removed startup-time
registration from tui_gateway/entry.py as brittle). Hooks register lazily at
agent creation time in _make_agent instead.

@NikolaRHristov

NikolaRHristov commented Jul 27, 2026

Copy link
Copy Markdown
Contributor Author

Related PRs, Issues, and Commits

Statuses verified via GitHub API 2026-07-27. Comments read for all.

Issues

Directly requests our feature

# Author Title Comments
#18988 ta3pks Feature request: pre_tool_call rewrite action (rewrite tool args from a hook) Offered to PR: "~26 LOC + ~150 tests." Subset of #18148. Never implemented.
#18148 jeremypetz Runtime extension hooks (transform_user_input, before_llm_call, intercept_tool_call) ta3pks: "#18988 is a strict subset of intercept_tool_call from this proposal."
#56969 lff231118 Tool Call Pre-Execution Hook — Enable URL Routing Rules Before Tool Dispatch Reframed as probabilistic skill triggering issue.
#61656 AdamClaassens Add opt-in fail-closed mode for required pre_tool_call policy hooks rmacbot: verified use case. #61902 PR for this.

Hook registration gaps

# Author Title Comments
#61806 tinuxongit shell hooks never registered in hermes serve alt-glitch: distinct/broader than other registration issues.
#43823 - Shell hooks never registered in desktop TUI entry point alt-glitch: duplicate of #41457. liuhao1024: "Fixed by #43826" (closed).
#41457 EdderTalmor Shell hooks not registered in TUI gateway and ACP adapter Original bug report.

Related pre_tool_call enhancements

# Author Title Comments
#43201 Mrshengjie When pre_tool_call returns block, support terminating current agent loop No comments.
#47092 m24927605 add pre_agent_invocation hook subagent_start exists but is source-specific.
#64654 xlionjuan pass assistant_response to pre_tool_call hook Analyzed code paths, confirmed design.
#64329 lizhao888 Pre-tool skill-require gate alt-glitch: duplicate of #40662.
#40662 wzgrx PreToolUse enforcement hook Open.
#359 teknium1 Enhanced Extension System with Tool Interception & Lifecycle Events Original feature request.

PRs — Hook Registration (10 total, all open/closed for same gap)

# Title Author State Registration Notes
#13854 fix(tui_gateway): register shell hooks at startup witt3rd Closed entry.py (startup)
#34112 fix: register shell hooks in TUI gateway orcool Open entry.py (startup)
#41464 fix: register shell hooks in TUI gateway and ACP adapter EdderTalmor Open entry.py + session.py Approved by tonydwb. No tests.
#41555 fix(security): register shell hooks in TUI gateway and ACP adapter liuhao1024 Closed entry.py + entry.py 110 lines of tests. Closed for #41464.
#43826 fix(tui_gateway): register shell hooks at desktop/TUI startup liuhao1024 Closed entry.py (startup) Referenced by #43823.
#43828 fix(tui): register shell hooks at TUI gateway startup liuhao1024 Closed entry.py (startup)
#53894 fix(hooks): register shell hooks on dashboard/TUI startup + utf-8 hook I/O Ne0teric Open entry.py (startup) Adds Windows utf-8 fix.
#57020 fix(tui_gateway): register config shell hooks on the desktop surface grimmjoww Open server.py (per-agent) Earliest per-agent model. #63781 is duplicate.
#63781 fix(gateway): register declarative shell hooks in Desktop/TUI backend gcunharodrigues Open server.py + session.py Duplicate of #57020.
#70461 fix: add 'serve' to _AGENT_COMMANDS webtecnica Open - Complementary path fix.

PRs — Tool Call Modification

# Title Author State Notes
#54810 feat: add pre_tool_call mutate action support Celth-Wu Closed "Extend pre_tool_call to support mutate action." Closed without merge.
#61902 feat(plugins): add opt-in required pre-tool policies embwl0x Open Complements #61656. Policy layer, not args mutation.

Our PR

# Title State Tests Closes
#28953 feat: pre_tool_call modify + TUI/ACP hook registration Open 10 #18988 #41457 #61806 #43823

Capability Matrix

                             ┌──────────┬──────────┬──────────┬──────────┬──────────┬──────────┐
                             │  #41464  │  #41555  │  #57020  │  #63781  │  #54810  │  #28953  │
┌────────────────────────────┼──────────┼──────────┼──────────┼──────────┼──────────┼──────────┤
│ TUI hook registration      │    ✓     │    ✓     │    ✓     │    ✓     │    ·     │    ✓     │
│ ACP hook registration      │    ✓     │    ✓     │    ·     │    ✓     │    ·     │    ✓     │
│ pre_tool_call modify       │    ·     │    ·     │    ·     │    ·     │    ✓     │    ✓     │
│ Modify merge accumulation  │    ·     │    ·     │    ·     │    ·     │    ·     │    ✓     │
│ Claude Code wire compat    │    ·     │    ·     │    ·     │    ·     │    ·     │    ✓     │
│ Sequential scope gate      │    ·     │    ·     │    ·     │    ·     │    ·     │    ✓     │
│ middleware_trace fix       │    ·     │    ·     │    ·     │    ·     │    ·     │    ✓     │
│ _block_error_type fix      │    ·     │    ·     │    ·     │    ·     │    ·     │    ✓     │
│ Full metadata pass-through │    ·     │    ·     │    ·     │    ·     │    ·     │    ✓     │
│ Per-agent-build model      │    ·     │    ·     │    ✓     │    ✓     │    ·     │    ✓     │
│ Unit tests                 │    ·     │    ✓     │    ·     │    ·     │    ·     │    ✓     │
│ Docs updated               │    ·     │    ·     │    ·     │    ·     │    ·     │    ✓     │
│ Maintainer approved        │    ✓     │    ·     │    ·     │    ·     │    ·     │    ·     │
└────────────────────────────┴──────────┴──────────┴──────────┴──────────┴──────────┴──────────┘

Merged (on main)

Commit Title
f512d6f020 feat(plugins): pre_tool_call approve action escalates to human gate
3a30c605b3 feat(plugins): add thread-local tool whitelist to pre_tool_call gate
9be3ab1a5b fix(plugins): stop firing pre_tool_call hook twice (#17611)
7427b9d581 fix(tool-search): scope bridge catalog to session's toolsets
8fbe2e388f feat(tool_search): probe-validate blind tool_call args

Our Commits

Commit Title
2d5507540f feat(pre-tool-call): add modify action
90a83ea36f feat(pre-tool-call): propagate routing context
3922b7a247 docs(hooks): document modify action
27fe63ec1c fix(tool-executor): initialise middleware_trace
3e60ab2903 fix(tool-executor): disambiguate block error types
111c625c20 fix(acp): register shell hooks before agent creation

@NikolaRHristov
NikolaRHristov requested a review from teknium1 July 27, 2026 19:38
In _run_agent_tool_execution_middleware, the _resolve_pre_tool_block closure reassigns final_args when a pre-tool-call hook returns modified arguments. The nonlocal declaration for final_args was previously placed inside the `if modified_args is not None` branch rather than at the top of the nested function.

Hoisting `nonlocal final_args` to the start of _resolve_pre_tool_block guarantees the closure correctly binds the enclosing-scope variable for reassignment on every path, avoiding UnboundLocalError / latent scoping fragility if the assignment path is later reordered or extended. No behavioral change to the hook-dispatch flow; purely a correctness/clarity fix for the closure's variable binding.
@NikolaRHristov NikolaRHristov changed the title feat(hooks): pre_tool_call content transformation + TUI gateway hook registration feat(hooks): pre_tool_call content transformation Aug 5, 2026
…and TUI agent startup

Removed the explicit register_from_config(config, accept_hooks=False) call (and its guarding try/except) from both SessionManager._make_agent in acp_adapter/session.py and _make_agent in tui_gateway/server.py.

The pre/post_tool_call shell hooks are already wired into the plugin manager through the central discovery path (discover_plugins() runs as a side effect of importing model_tools.py, which both entry points eventually exercise). The manual registration at agent startup was therefore defensive duplicate work that could also double-register hooks if the central path later changed. Dropping it keeps ACP and TUI consistent with the gateway/run.py registration model and removes a dead try/except that only swallowed errors during a now-unnecessary call.

No behavioral change to hook dispatch: hooks are still resolved from HERMES_ACCEPT_HOOKS and hooks_auto_accept in config exactly as before, just via the shared discovery mechanism rather than a per-entry-point bootstrap.
@NikolaRHristov

Copy link
Copy Markdown
Contributor Author

As per:

Keep #62681 open with the requested real-import profile test, and require
author action on #28953 to split the rewrite feature from registration or
rebase the registration portion while resolving its contributor review
findings; the already-closed #13854, #41548, #41555, #43826, and #43828 need
no further state change.

From: #41457 (comment)

Done - I took the rebase option. 539974b
removes the shell-hook registration from both tui_gateway/server.py (−10) and
acp_adapter/session.py (−14); both files are now byte-identical to main, so
registration no longer appears in this PR's diff at all.

#28953 is now solely the pre_tool_call modify feature - +288/−23 across
hermes_cli/plugins.py, agent/shell_hooks.py, agent/tool_executor.py,
agent/agent_runtime_helpers.py, model_tools.py, plus tests and docs. No
gateway files touched. That should also clear the overlap with #41464 and the
_make_agent cluster (#48770 / #57020 / #63781 / #67084) - this PR no longer
competes with any of them.

Note this had to be a forward removal rather than a revert: the ACP registration
was a standalone commit (111c625), but the TUI
registration lived inside 7ce97af, the same
commit that introduces the modify feature. The net diff is clean either way.

Two things from testing this that may be useful to the wider complex:

1. The registration point leaks hooks across profiles. This is the same
defect flagged on #57020 and #63781, and it applies to any fix registering at
the TUI _make_agent chokepoint. shell_hooks._registered is a module-level
set and manager._hooks is a process-global dict, while one TUI process serves
multiple profiles. I reproduced it: profile A declares a hook, profile B
declares none, and after both agents are built profile B fires profile A's hook.
Since hooks execute arbitrary shell commands and the consent allowlist is also
process-global, that's a consent-boundary issue rather than just a correctness
wart.

2. I have the _make_agent factory-level regression test that several of
these PRs are noted as missing. It patches only _load_cfg and calls the real
tui_gateway.server._make_agent, with a sentinel halting execution just past
the registration point - so it reproduces the bug on main
(hook_registered: false) and proves the fix (true). Happy to contribute it
to #41464 or wherever it's most useful.

Separately, while rebasing I hit a SyntaxError in agent/tool_executor.py -
nonlocal final_args declared after use, which made the module unimportable. It
came from merge e502f20, not from the feature
commits (7ce97af,
2d55075,
90a83ea all compile). Fixed in
a1d8868.

Finally, one correction to the triage summary: it lists this PR's registration
as landing in tui_gateway/entry.py. That was true of an earlier approach,
removed in 1eecfb9; the registration actually
lived in tui_gateway/server.py and acp_adapter/session.py. My PR description
was stale and I've since corrected it.

teknium1 pushed a commit that referenced this pull request Aug 16, 2026
Adds a `modify` response type to pre_tool_call hooks so a hook can
transform tool arguments before the tool executes, instead of repairing
results afterwards via post_tool_call.

- hermes_cli/plugins.py: _dispatch_pre_tool_call_hooks() fires hooks once
  and returns (block_message, modified_args); modify directives
  shallow-merge into an accumulated dict built from the original args.
- agent/shell_hooks.py: _parse_response() accepts both the canonical
  {"action": "modify", "args": {...}} and Claude Code-compatible
  {"decision": "modify", "tool_input": {...}} wire formats.
- model_tools.py, agent/tool_executor.py, agent/agent_runtime_helpers.py:
  dispatch sites migrated; modified args applied before execution.
- Docs + 10 new tests (merge semantics, precedence, block interplay).

Salvaged from PR #28953. Best fix for #18988.
teknium1 added a commit that referenced this pull request Aug 16, 2026
Follow-up to the #28953 salvage:

- Extract _resolve_block_from_details() so resolve_pre_tool_block and
  _dispatch_pre_tool_call_hooks share ONE fail-closed approval-gate
  implementation. This also gives the new dispatcher the observability
  context wrapping around request_tool_approval that the original PR's
  inlined copy lacked.
- Update sibling tests that patched resolve_pre_tool_block at the three
  migrated dispatch sites to patch _dispatch_pre_tool_call_hooks with the
  (block_message, modified_args) tuple contract.

Verified: 448 targeted tests green; E2E with a real shell hook in an
isolated HERMES_HOME rewrote a live write_file call (path + content)
through handle_function_call, with block and negative paths intact.
@teknium1

Copy link
Copy Markdown
Contributor

Merged via PR #87482#87482 (rebase-merge, your authorship preserved on the feature commit; merged at 411903b).

Thanks for the clean split per the #41457 triage, the sustained main-merges, and the live E2E validation table — that made this an easy salvage. The only change on top: we extracted the fail-closed approve-gate logic into a shared _resolve_block_from_details() so the new dispatcher and resolve_pre_tool_block can't drift, and updated the sibling tests at the migrated dispatch sites.

The _make_agent factory-level regression test you offered would be genuinely useful on #41464 — please do post it there. The cross-profile hook-leak repro is noted as well; that's a real consent-boundary concern for the registration cluster.

@teknium1 teknium1 closed this Aug 16, 2026
@NikolaRHristov
NikolaRHristov deleted the feat/pre-tool-call-content-transform branch August 16, 2026 06:05
vashkartik added a commit to vashkartik/hermes-agent that referenced this pull request Aug 17, 2026
* feat(config): wire compression.tail_mode + docs (en/zh)

#87326 shipped the lean-compaction capability on the compressor; this adds
the config.yaml surface (compression.tail_mode: legacy|lean, default
legacy), DEFAULT_CONFIG entry, and docs on both the dev-guide compression
page and the user-guide configuration page, with zh-Hans parity.

* fix(computer-use): zero-rect AX bounds serialize as unknown, not a position

KDE/Qt apps report [0,0,0,0] bounds for elements that are perfectly
clickable by index (live QA: all 50 of kcalc's zero-rect elements,
including every radio button). Serializing that as a plausible rect
invites a model to derive coordinate=[0,0] and click the screen corner.

- _element_to_dict: zero rect -> bounds: null
- _format_elements: '@ bounds-unknown (click by element index)' instead
  of the fake rect in the summary line
- malformed bounds fail open (unchanged serialization)

Live-proven on real kcalc (cua-driver 0.20.0): 50 elements now null, 0
zero-rect leftovers, summary annotated, real rects preserved, and a
null-bounds radio button still clicks fine by index.

* feat(desktop): drag-resizable panes on the Capabilities Skills tab

The Skills tab's three panes are now all drag-resizable:

- List/detail column seam: MasterDetail grows an optional resizeId that
  turns the seam between the rail and the detail pane into a vertical
  drag sash (same visual language as DetailPane's top-edge sash). The
  rail width persists in the shared pane store under that id;
  double-click resets to the default 0.75fr track. Skills and Tools
  tabs share one id so the split stays consistent across tabs.
- Skills Hub section: the embedded hub picker's top edge is now a drag
  sash — pull the hub pane up to grow it (the skills list above
  absorbs the change). Height persists through the same pane store,
  double-click resets, and the cross-origin iframe gets
  pointer-events:none during the gesture so it can't swallow the drag.
  Replaces the old native CSS corner-resize handle.
- The skill editor bottom pane already resized via DetailPane's sash.

MASTER_DETAIL_WIDE_COLS keeps its exported shape (MCP tab reads it) —
the --md-split var falls back to the declared track when unset, so
grids without a sash render exactly as before.

* feat(hooks): pre_tool_call content transformation via `modify` directive

Adds a `modify` response type to pre_tool_call hooks so a hook can
transform tool arguments before the tool executes, instead of repairing
results afterwards via post_tool_call.

- hermes_cli/plugins.py: _dispatch_pre_tool_call_hooks() fires hooks once
  and returns (block_message, modified_args); modify directives
  shallow-merge into an accumulated dict built from the original args.
- agent/shell_hooks.py: _parse_response() accepts both the canonical
  {"action": "modify", "args": {...}} and Claude Code-compatible
  {"decision": "modify", "tool_input": {...}} wire formats.
- model_tools.py, agent/tool_executor.py, agent/agent_runtime_helpers.py:
  dispatch sites migrated; modified args applied before execution.
- Docs + 10 new tests (merge semantics, precedence, block interplay).

Salvaged from PR #28953. Best fix for #18988.

* refactor(hooks): share fail-closed approve logic; update sibling tests

Follow-up to the #28953 salvage:

- Extract _resolve_block_from_details() so resolve_pre_tool_block and
  _dispatch_pre_tool_call_hooks share ONE fail-closed approval-gate
  implementation. This also gives the new dispatcher the observability
  context wrapping around request_tool_approval that the original PR's
  inlined copy lacked.
- Update sibling tests that patched resolve_pre_tool_block at the three
  migrated dispatch sites to patch _dispatch_pre_tool_call_hooks with the
  (block_message, modified_args) tuple contract.

Verified: 448 targeted tests green; E2E with a real shell hook in an
isolated HERMES_HOME rewrote a live write_file call (path + content)
through handle_function_call, with block and negative paths intact.

* chore: contributor mapping for NikolaRHristov

* fix(mcp): handle DCR clients with secrets across all OAuth paths

Some MCP OAuth providers (notably Supabase) return a client_secret from
dynamic client registration but omit token_endpoint_auth_method. The MCP
SDK defaults the missing method to "none", so the token exchange omits
client_secret and the server rejects it (HTTP 422 "Required parameter:
client_secret"), looping the browser consent page.

This resolves the whole class, not just one provider:

- Storage layer (HermesTokenStorage): coerce secret-bearing client info
  with missing/none auth method to client_secret_post on both read and
  write, persisting the corrected shape.
- Both live provider paths (tools/mcp_oauth.py HermesOAuthClientProvider
  and tools/mcp_oauth_manager.py HermesMCPOAuthProvider): coerce
  in-memory client info immediately before token exchange and refresh.
- Accept the full 2xx range on token and refresh responses (Supabase
  returns 201 Created), instead of the SDK's exact-200 check.
- Redact token response bodies from error messages and logs on
  malformed responses.

The Figma-specific request-time default (apply_oauth_provider_defaults)
remains; this generalizes the same bug class for every DCR provider.

Fixes #29680. Supersedes #34274 and #35700 (201-only variants).

* chore: map contributor email for 5Hyeons

* fix(compression): watermark commit — appends flow freely, concurrent tail survives compaction

Redesign of the #75316 class (supersedes the approach in PR #87307).

Root cause family: the compression lock fenced ORDINARY transcript appends
for the whole slow provider-summary call. Turns died as
session_persistence_failed whenever a message overlapped a compression
(#74568, #77386, #75083), stale dead-PID locks blocked writes for the full
TTL, and the busy-wait mitigation (#75264) was an order of magnitude shorter
than real summaries. Separately, the commit archived from a pre-call
snapshot, so rows appended mid-compression were swept into the archive.

Design: the commit transaction is already exclusive — no lock phases needed.

1. Appends never check compression_locks. The lock's only job is stopping
   two compressions colliding; it keeps that job. The whole stale-lock /
   busy-wait symptom family dies as a class.
2. Watermark captured in the DB at compression start
   (get_active_message_watermark = MAX(id) of active rows) — not from
   in-memory message dicts, which carry no row ids in production.
3. archive_and_compact(watermark=, lock_holder=): one transaction verifies
   the holder still owns an unexpired lease (a reclaimed lease cannot
   publish a stale compaction), archives the snapshot, inserts the compacted
   set, and re-sequences the concurrent tail (id > watermark) via a
   pure-SQL column clone — every column except id survives byte-exact
   (api_content, platform_message_id, reasoning sidecars, token counts),
   FTS triggers index the clones naturally, originals stay archived and
   recoverable. watermark=None preserves the historical behavior.

Removed: the append-side compression fence in _check_transcript_write_guards
(with rationale note), making the _COMPRESSION_BUSY_WAIT_S retry lane
unreachable from append paths (kept for other callers).

Tests: 12 new (watermark contract, column-exact clone, commit fence incl.
lease-lost/expired/rollback failure injection, append-vs-commit race);
busy-retry suite flipped to pin the new contract; sabotage-verified (5 fail
with the watermark disabled, 12 pass restored); E2E through the real
compress_context seam with a mid-summary append landing and surviving.

* fix(compression): rotation path clones the concurrent tail into the child

CI caught the sibling site the in-place fix missed: legacy (non-in-place)
compression rotates via publish_compression_child, where a mid-summary
append previously stranded in the closed parent. Same watermark + pure-SQL
column clone as archive_and_compact, with session_id rewritten to the child.
Lineage-guard test flipped to pin the appends-flow-freely contract; rotation
watermark tests added (tail follows the child; None = historical behavior).

* fix(compression): bound the rotation tail clone below the rotator's own flush

CI caught two rotation-path regressions from the unbounded clone: the #47202
pre-publish flush writes the rotator's OWN input transcript to the parent
(above the start-watermark), and the clone was duplicating it into the child
alongside the handoff. publish_compression_child gains watermark_ceiling —
the MAX(id) captured immediately BEFORE that flush — so only rows in
(watermark, ceiling] (genuinely foreign concurrent appends) clone across.
Ceiling capture failure falls back to no tail preservation (historical
behavior) rather than risking duplication. Ceiling-exclusion test added.

* fix(terminal): warn when exit_code 0 masks a piped build/test failure

`cargo build 2>&1 | tail -20` exits with tail's 0 even when the build
failed — bash without pipefail reports the last pipeline command's
status, and `cmd || echo failed` swallows the status the same way. The
model reads exit_code: 0 as a strong success signal and can conclude a
build passed while the visible output says it failed (community report,
Windows Rust builds; not platform-specific).

Two-part fix, mirroring OpenCode's prompt-side approach plus a
result-side backstop they don't have:

- Tool description now forbids piping builds/tests through
  tail/head/cat (output is already auto-truncated + spilled to a file)
  and warns that pipes/|| fallbacks mask exit codes.
- New annotate_masked_success() in tools/terminal_hints.py: when
  exit_code == 0, the command shape can mask an upstream status
  (top-level pipe into a passthrough consumer, or || echo/true), AND
  the output carries strong tool-specific failure shapes (rustc,
  cargo, pytest, gcc, npm, make, ninja), attach an advisory 'hint'
  telling the model to treat the run as failed and re-run bare.
  exit_code itself is never modified. Search/content heads
  (grep/rg/echo/printf/...) are excluded to avoid false positives on
  pipelines whose output legitimately contains error text.

E2E-verified through the real terminal tool path: hint fires on masked
cargo-style failures, silent on bare commands, clean pipes, and
grep/printf pipelines. 42 targeted tests pass.

* fix(desktop): defer credential-warning onboarding to the first chat attempt

Switching to a profile with no provider configured popped the blocking
onboarding overlay (Nous Portal / provider picker, "gateway isn't
ready") the moment the profile's runtime info arrived — punishing the
user for merely looking at an unconfigured bot/profile.

The passive credential_warning (session create/activate/resume info,
stream heartbeats) is now stashed instead of opening the overlay.
The submit path consumes it when the user actually tries to chat and
opens onboarding then, before the doomed send; the draft stays in the
composer. A warning-free session event clears the stash, so healed or
switched-away profiles never fire stale onboarding. Turn-error paths
(a real failed send) still open onboarding immediately, unchanged.

* fmt(js): `npm run fix` on merge (#87505)

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>

* fix(agent): reject masked verification results

* fix(cli): map all Ctrl/Alt/Shift+key combos under modifyOtherKeys level 2

Commit 4c34eeb416 stopped pushing the Kitty keyboard protocol (CSI >1u)
because Ctrl+C arrived as ESC[99;5u instead of \x03, breaking SIGINT.
But modifyOtherKeys level 2 (CSI >4;2m) was kept so Shift+Enter stays
distinguishable from Enter.

Under modifyOtherKeys=2, terminals re-encode EVERY Ctrl+key combo as
ESC[27;5;<codepoint>~ instead of the raw control byte. prompt_toolkit
3.x only maps ESC[27;5;13~ (Ctrl+Enter = Ctrl+M); all other Ctrl+letter
combos are unmapped and leak as literal text or get swallowed — breaking
Ctrl+A, Ctrl+C, Ctrl+D, Ctrl+E, Ctrl+K, Ctrl+R, Ctrl+U, Ctrl+W, Ctrl+Z,
etc. Shift+letter combos (ESC[27;2;<codepoint>~) have the same problem,
causing the 'caps locked sessions' symptom where typed text appears
corrupted or stuck.

Fix: add install_modify_other_keys_aliases() to pt_input_extras.py that
populates prompt_toolkit's ANSI_SEQUENCES dict with 294 mappings covering:
- Ctrl+letter (a-z): ESC[27;5;<code>~ and ESC[<code>;5u -> Keys.ControlA..Z
- Ctrl+digit (0-9): same formats -> Keys.Control0..9
- Ctrl+symbol ([ \ ] ^ _ @ Space): same formats -> matching Keys.Control*
- Alt+letter (a-z, A-Z): both formats -> (Escape, <letter>) tuple
- Shift+letter (a-z, A-Z): both formats -> uppercase character

Uses setdefault semantics — never clobbers existing mappings from
install_shift_enter_alias or install_ctrl_enter_alias. The Ink TUI
(Node.js) already handles this via a regex parser; prompt_toolkit 3.x
uses dict lookup only, so we populate the dict.

Refs #56684, #87711.

* refactor: extract _install_paired helper, fix misleading comment, isolate test fixture

Apply findings from /simplify-code 3-agent review:

1. Extract _install_paired() inner helper — the Ctrl, Alt, and Shift
   sections all repeated the same mok+csiu sequence generation pattern
   (~30 lines of duplication). Now each section builds a dict and
   delegates to _install_paired(modifier, mapping).

2. Replace 10 hardcoded Ctrl+digit lines with a loop matching the
   Ctrl+letter pattern above it.

3. Fix misleading comment: claimed 'Ctrl+0 doesn't produce a control
   byte' but chr(ord('0') & 0x1F) = 0x10 = ControlP. The code was
   correct (maps directly to Keys.Control0..9); only the comment was
   wrong.

4. Add comment explaining why Shift+letter maps both lowercase and
   uppercase codepoints (some terminals send the already-shifted
   codepoint with modifier=2).

5. Test fixture: snapshot/restore ANSI_SEQUENCES in teardown so 294
   mappings don't leak into sibling test files (global mutable state).

* chore: map justin@bowes.org to @justinbowes

Attribution mapping for the PR #84982 salvage. The commit email is not
linked to a public GitHub account, so contributor_audit --strict fails
without it; login confirmed from the PR author field.

* feat(desktop): unify MCP Servers and Catalog into one coherent list

The MCP tab's left column previously split the configured fleet and the
Nous-approved catalog behind a Servers/Catalog tab toggle. Installed
entries appeared in both views and the install button lived a tab flip
away from the list it fed.

Now one scrolling column: configured servers (live status, toggles,
probes) on top, a Catalog section below offering only entries not yet
installed. Installing moves the entry up into the fleet list; the
zero-servers empty state keeps the catalog visible beneath it instead
of hiding it behind a full-page invitation.

Bot Mode's Advanced view embeds this same McpTab via the plugin SDK, so
the unification mirrors there automatically.

- removed leftView state + TextTab toggle; section headers reuse the
  existing tabServers/tabCatalog strings (no i18n changes)
- availableCatalog memo filters installed/name-clashing entries
- catalog memoized to satisfy react-hooks/exhaustive-deps

* perf(desktop): make session resume incremental

* test(tui): wait for deferred profile DB close

* test(desktop): preserve bounded deferred resume

* test(tui): isolate deferred hydration worker

* docs(tui): document defer_history vs omit_messages precedence

Follow-up to the #62799 salvage: Desktop sends both defer_history and
omit_messages on a cold resume. Make explicit in the deferred branch that
defer_history supersedes omit_messages — the single history read happens in
the background hydration worker and the synchronous omit_messages read on
the cold-resume default path is skipped entirely, so the transcript is
never loaded twice for one resume.

* feat(desktop-sdk): host.deleteProfile — teardown-routed profile delete for plugins

Plugins deleting profiles via `cli.exec ['profile','delete',…]` bypass the
Electron-side DELETE /api/profiles interception (prepareProfileDeleteRequest),
so a live pool backend — e.g. one the roster's hover pre-warm just woke —
holds the profile dir open and the renderer's reconnect respawns it
mid-delete, recreating the directory (#52279). Bot Mode's right-click Delete
hits this every time because right-click hovers the row first.

Add host.deleteProfile(name) to the plugin SDK: routes through the same
teardown-routed REST path core's DeleteProfileDialog uses (backend teardown
first, next request routed away), rejects on 'default', and re-homes the app
to the default profile when the deleted profile was the live gateway's —
mirroring the core dialog's ordering.

Reported by @BkashJosi (Bot Mode: deleting a bot errors while its session
is awake).

* fix(gateway): release loop ticks after empty responses

* fix(gateway): isolate post-turn loop failures

* fix(install): capture npm output on failure for diagnosable errors (#87340)

Both install-blocking npm install call sites (browser tools and TUI) ran
with --silent and no output capture, so failures printed only a generic
error message with no npm diagnostics.

Apply the same pattern used by the camofox install path: redirect npm
output to a temp file and replay it on failure, so users can see the
actual error (EBADENGINE, ETARGET, network timeout, registry 5xx, etc.).

Fixes #87340

* fix(install): validate --commit SHA and fail hard on fetch/checkout errors (#87268)

Three problems with install.sh --commit:

1. No validation: non-hex or too-short arguments passed through to git,
   producing misleading errors.

2. Fetch failure swallowed by || true: abbreviated SHAs are refused by
   GitHub's server ("couldn't find remote ref"), but the error was
   silently ignored.

3. Checkout failure not checked: git checkout --detach with a missing
   object produces a misleading "does not take a path argument" error
   and the install continues unpinned, exiting 0.

Fix:
- Validate --commit is a 7-40 hex string up front
- Remove || true from fetch; fail with actionable message directing
  users to full 40-char SHAs
- Check git checkout --detach result and fail hard on error

Fixes #87268

* fix(webhook): authenticate Linear deliveries via linear-signature HMAC (#87348)

* fix: persist computer_use provider selection so Desktop picker survives refresh

The `computer_use` toolset (cua-driver) had no persistence branch in
_write_provider_config or _is_provider_active. When the Desktop GUI
sent PUT /api/tools/toolsets/computer_use/provider, the config write
was a no-op and _is_provider_active always returned False -- so the
"Use this backend" CTA reappeared on every refresh.

Add a computer_use_backend marker to the cua-driver provider entry
and handle it in the same pattern as web_backend / browser_backend:
- _write_provider_config now sets computer_use.backend = "cua"
- _is_provider_active now checks computer_use.backend
- _reconfigure_provider (interactive CLI) also persists the key

Fixes #86962

* fix(runtime): exempt loopback custom-provider pool credentials from the usable-secret floor

Fixes #86864.

Legacy custom_providers configs commonly used short/placeholder
api_keys ('123', 'm') for local no-auth services like Ollama --
harmless for the endpoint itself, since Ollama accepts any key or no
key. A stricter has_usable_secret(value, min_length=4) gate added
later now rejects these, but only the credential-POOL resolution path
lacked the same "no-key-required" exemption every OTHER resolution
path in this file already has for exactly this scenario:

- The config-based custom_providers fallback (non-pool path) already
  ends with `api_key or "no-key-required"`.
- The "actual" provider's local-offline path already injects
  ACTUAL_LOCAL_NOAUTH_PLACEHOLDER before the usable-secret gate for a
  loopback base_url.
- _try_resolve_from_custom_pool() was the one gap: it returned the raw
  short pool credential unchanged, which then failed the downstream
  has_usable_secret() gate with a generic "No usable credentials found
  for custom" error that contradicts setup.status ("configured
  credentials" vs "runtime failed"), sending users hunting in the
  wrong direction.

Fixed by substituting the same "no-key-required" placeholder when the
pool's stored credential fails has_usable_secret() AND the base_url
resolves to a loopback hostname (using the existing _loopback_hostname
helper, matching the exemption scope the issue itself requested:
localhost/127.0.0.1/::1 only, not arbitrary remote endpoints with a
genuinely-too-short key).

Added 4 regression tests extending the existing
test_runtime_provider_resolution.py file, following its established
credential-pool mocking pattern: the exact reported 3-char repro
('123'), a 1-char case, a non-loopback sanity check confirming the
exemption stays scoped (a short key for a remote endpoint is NOT
silently exempted), and a sanity check that a genuinely usable
loopback key passes through unmodified. Verified as a genuine
regression by reverting the fix and confirming 2 tests fail with the
exact raw short key leaking through unchanged.

59/59 pass in the extended test file; 14/14 across two more related
custom-provider test files (no regression).

* fix(cron): do not treat directories as unsafe lifecycle scripts (#86753)

Docker Desktop writes fpath=(~/.docker/completions ...) into .zshrc.
The referenced-script walk then opened that directory, saw a non-regular
file, and fail-closed — blocking source ~/.zshrc on every terminal
command. Directories are not scripts; devices stay fail-closed.

* fix(gateway): read routed profile model config

* fix(batch_runner): propagate fatal and validation errors as non-zero exit codes

Python Fire serializes the return value of a wrapped function but does not
use that value as the process exit code. Error paths in main() that used
 or  therefore caused the process to exit 0, swallowing
fatal errors and argument-validation failures.

Raise SystemExit(1) on every error path so batch_runner returns a non-zero
exit code when it cannot run. Success paths (e.g. --list_distributions) are
left unchanged.

Closes NousResearch/hermes-agent#86524.

* fix(gateway): surface actionable message for local model server connection errors

* fix(gateway): keep broad connection phrases out of the provider-error gate

Fixes #86570

* fix(gateway): catch BaseException in _process_message_background to notify on SystemExit

The fire-and-forget handler only caught asyncio.CancelledError and
Exception, so a SystemExit/KeyboardInterrupt escaping a turn (e.g. a
plugin calling sys.exit() in a tool call or summary-LLM path) skipped
the user-facing failure notification and surfaced only as 'Task
exception was never retrieved' — radio silence for the user.

Catch BaseException instead; send the failure notification first, then
re-raise SystemExit/KeyboardInterrupt to preserve shutdown semantics
for the loop's own signal handling. Other BaseExceptions stay contained.

Closes #86651

* fix(cli): use custom provider default_model when --provider is set

hermes chat --provider <name> without -m sent the global model.default
to the custom endpoint. Named custom entries already expose
default_model via _get_named_custom_provider(); honor that when the
user selected the provider and did not pass an explicit model.

Fixes #86978

* fix(cli): warn when --provider default_model cannot be resolved

A named --provider without -m used to swallow lookup failures and
silently keep the global model.default. Log the resolution error so
the fallback is visible.

* fix(cli): allow persisted contributor tier consent

* fix(mcp): prefer server-native tool over generated utility on name collision (#87112)

An MCP server exposing a native tool named read_resource (or
list_resources/list_prompts/get_prompt) collided with the auto-generated
resource/prompt utility of the same name. The registration collision
handler flagged the pair as ambiguous and skipped BOTH entries, so the
server's own tool became silently unavailable on every gateway boot.

Resolve this specific native-vs-utility collision in favour of the native
tool: keep it and drop the shadowed utility, which is only convenience
sugar for servers that expose no such tool of their own. The conservative
skip-everything path still applies to genuinely ambiguous collisions (two
or more native tools normalizing to one name), which we cannot
disambiguate. Add a regression test covering the native-tool-wins path.

Fixes #87112

* fix(desktop): restore mod-chord keybinds while typing in inputs

#86586 replaced the combo-based input gate (any Cmd/Ctrl chord fires
while typing) with an action allowlist that dropped session.new and
every other mod-chord not explicitly listed. ⌘N/⌘T/⌘⇧N and friends
became dead keys whenever focus was in the composer.

Restore the pre-regression rule: primary-modifier chords stay global
even in text fields; the allowlist now gates only bare/Shift/Alt combos,
so rebound letter keys can never hijack typing. Text-navigation chords
(Ctrl+Arrow/PgUp/PgDn) still stay with the input.

* fix(desktop): reject bare modifier combos in the input gate

Review NIT: a malformed stored binding of just 'mod'/'ctrl' (never
produced by comboFromEvent) could pass the shape-only mod/ctrl check.
Reject bare-modifier bases in actionAllowedInInput and pin it in the
suite.

* fix(agent): preserve local reasoning timeout opt-out

* fix(approval): deterministic approvals.single_query_mode for -q sessions

hermes chat -q sets HERMES_INTERACTIVE=1 (for interactive sudo prompts) but
runs one turn with no user waiting to answer approval prompts. Previously a
dangerous command triggered the interactive gate, waited the full 300s
timeout, then failed closed — and the agent was effectively forced to work
around the block, often silently auto-approving via execute_code (which
auto-approves in non-gateway mode).

Add approvals.single_query_mode (default deny, mirror of cron_mode):
  deny    — block dangerous commands and execute_code deterministically with
            a clear 'no user present' message (no 300s wait)
  approve — auto-approve dangerous commands/execute_code in -q mode

cli.py marks the session with HERMES_SINGLE_QUERY_SESSION; the shared gate
(_run_approval_gate, check_all_command_guards, check_execute_code_guard)
treats -q as a deterministic non-interactive context when that marker is set.
execute_code, the -q escape hatch, now honors single_query_mode instead of
auto-approving headlessly. Includes tirith parity in the combined guard and
docs. Fixes #86878.

* fix(gateway): treat Ready scheduled tasks as Windows supervisors

After the VBS/cmd launcher exits, Task Scheduler marks
Hermes_Gateway_* Ready while the detached gateway keeps running.
The orphan reaper only bailed on Running, then fail-opened the
parent-chain check and killed the live bot on desktop serve start.

Fixes #87001

* fix(gateway): preserve launchd supervisor marker across stderr_timestamp wrapper

launchd only stamps XPC_SERVICE_NAME on its direct child. The timestamp
wrapper is that child, so the grandchild gateway sees XPC_SERVICE_NAME=0
and the supervised-conflict guard refuses the service's own spawn.

Forward HERMES_GATEWAY_EXTERNAL_SUPERVISOR=1 when the wrapper itself is
launchd-supervised. Interactive XPC_SERVICE_NAME=0 starts stay unmarked.

Fixes #86893

* fix(gateway): put --external-supervisor on launchd gateway argv

hermes update decides restart ownership from the live grandchild argv,
not from an env marker. Newly generated plists now include the flag.
The stderr_timestamp wrapper upgrades only historical Hermes gateway
run shapes for stale plists and leaves arbitrary launchd children unmarked.

* fix(agent): preserve stalled-provider escalation

* fix(agent): cover provider wait teardown paths

* fix(sessions): surface open sessions skipped by prune

* fix(sessions): address prune skip review notes

* fix(sessions): align prune filter derivation

* fix(agent+discord): guard truncated-response continuation loops and cap Discord split delivery (#86581)

* fix(agent): bound worker finalization when iteration budget exhausted (#87096)

Adds a bounded fallback path in turn_finalizer.py that always records
a terminal timed_out outcome via _record_task_failure (CAS receipt path)
when the iteration budget is exhausted, regardless of whether the normal
fallback paths (interrupted/failed/anomalous exit_reason) were eligible.

Previously, a kanban worker whose budget was exhausted but whose turn was
interrupted, failed, or exited with an anomalous reason would silently
leave its task in an ambiguous lifecycle state — the dispatcher would
eventually detect it as a crashed or protocol-violation worker, but the
failure was not bounded and could take a full tick cycle to reconcile.

The CAS invariant in _end_run (WHERE ended_at IS NULL) guarantees
idempotence: if another path already closed the run, the call is a no-op.

Extracted the inline kanban-budget-exhausted recording into a shared
helper function (_record_kanban_budget_exhausted) used by both the
existing iteration_limit_fallback path and the new bounded fallback path.

Closes #87096

* fix: guard exit watchdog against mid-cleanup overlap

* fix(desktop): hard-exit a lock-losing second instance before ready

app.quit() does not stop a lock-losing instance from reaching whenReady:
the before-quit teardown coordinator defers the quit (event.preventDefault
+ async backend shutdown), and ready fires in that window. The losing
instance then runs the full startup whose reapOrphans() SIGTERMs the
running instance's live backend (#87295).

The lock-loser holds no state and no backend — requestSingleInstanceLock()
has already delivered the argv to the primary by the time it returns false —
so there is nothing to clean up. app.exit(0) terminates immediately, before
ready, so a second launch routes into the running window and never touches
backend machinery.

* fix(desktop): keep startHermes inert without the single-instance lock

startHermes() is the only entry point that can reap, spawn, claim, and
therefore destroy a backend. Belt-and-suspenders on the exact killing line:
even if some future path reaches it in a lock-losing process (a refactor, a
dev harness, a race), the instance stays inert — no reap, no spawn, no
claim — instead of SIGTERMing the running instance's backend (#87295).

* fix(desktop): never reap a backend whose parent Electron is alive

reapOrphans() treats any recorded backend with a matching process identity
as orphan-reapable — including a backend owned by another live instance. The
ownership file is shared across instances, so a second launch that reaches
reap (even without the lock) SIGTERMs the running instance's backend.

Claims now record the spawning Electron (parentPid + parentStartMarker, the
same values already passed to the backend as HERMES_PARENT_PID /
HERMES_PARENT_START_MARKER), and reapOrphans() skips any entry whose parent
is still running. Even a second instance that wins a stale lock can never
kill a live instance's backend. Legacy entries without parent data keep the
old behaviour.

Known tradeoff: a parent-liveness probe failure preserves the record, so a
genuinely orphaned backend under a still-running parent is leaked until it
dies naturally. That is preferable to killing a live instance's backend.

Tests: parent-aware reap in backend-ownership.test.ts (live parent
preserved, dead parent still reaped, probe failure preserved, parent
identity round-trips through claim/parse).

* fix(agent): include timezone and UTC offset in system prompt timestamp

The "Conversation started:" line carried a bare date (%A, %B %d, %Y). Tools
that accept instants -- nutrition, calendar and similar MCP servers -- reject
naive datetimes and require an explicit UTC offset, so the model had to infer
EST vs EDT from the date alone. Near a DST boundary that is a coin flip, and a
wrong guess does not error: it silently writes the record onto the wrong day.

Append the IANA zone (when configured), the zone abbreviation and the UTC
offset, e.g.:

  Conversation started: Saturday, August 15, 2026 (America/New_York, EDT, UTC-04:00)

get_timezone() returns None when no timezone is configured; in that case the
line falls back to the abbreviation and offset of the server-local (still
tz-aware) time, so behaviour is unchanged for users who never set one:

  Conversation started: Saturday, August 15, 2026 (EDT, UTC-04:00)

Daily byte-stability is preserved -- the property the date-only format exists
to protect (PR #20451). Zone name, abbreviation and offset are all constant for
the whole day; they shift only at a DST transition, where a change is correct.
The static-prefix reconstruction guard in _restore_plugin_sections matches on
"\n\nConversation started:" and is unaffected by a suffix after the date.

test_datetime_is_date_only_not_minute_precision used `re.search(r":\d{2}")`
over the whole line as a proxy for "no time-of-day". A UTC offset also matches
that pattern, so the check now applies to the date portion (everything before
the zone parenthetical) and the invariant is tightened rather than relaxed:

- test_datetime_includes_utc_offset asserts the offset is present
- test_datetime_line_is_stable_across_rebuilds asserts two rebuilds in the
  same day produce a byte-identical line

Fixes #87403

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(gateway): make managed Node suppress PATH fallback

* fix(agent): canonicalize duplicate tool call arguments

* fix(agent): harden canonical tool call deduplication

* fix(gateway): let an explicit workspace move win for a running session

session.workspace.move refused a running live session with 4009
(session busy), but the desktop's Move-to-project flow calls exactly
this RPC — so the UI updated its local grouping while state.db kept the
old cwd and the agent's tools kept running in the old workspace. Two
sources of truth disagreed (#86626).

An explicit move now wins: the stored row and the live session re-anchor
together. In-flight tool calls keep the cwd they were launched with; the
next tool call uses the new workspace.

* fix(ui): ignore inherited esbuild binary overrides

* fix(ui): scrub esbuild override during builds

* fix(codex): clamp xAI max/ultra aliases to the model's reasoning ceiling (#87279)

* fix(backup): bound locked database snapshot waits

* fix(gemini): raise maxOutputTokens when thinking is enabled

Gemini bills thought tokens against maxOutputTokens/max_tokens, so a
global 4096 cap can be fully consumed by thinking on the first
request, leaving zero content tokens and aborting after 4
continuations. When thinking is enabled, raise the effective output
cap to the 65,535 ceiling on both the native and chat-completions
paths.

Refs #83915

* fix(gateway): don't crash on a foreign XDG_RUNTIME_DIR in user-systemd preflight (#86558)

runuser/su/sudo -u from a root shell leaks XDG_RUNTIME_DIR=/run/user/0 into
the child. _user_systemd_socket_ready() stat-ed sockets under it with a bare
Path.exists(), which only suppresses ENOENT/ENOTDIR/EBADF/ELOOP — EACCES on
the 0700 root-owned dir escaped as a raw PermissionError traceback instead of
the documented UserSystemdUnavailableError remediation path.

- _path_exists_safe(): Path.exists() that treats EACCES as absent; used at
  both the readiness and DBUS-detection call sites.
- _ensure_user_systemd_env(): drop an XDG_RUNTIME_DIR that is unset or owned
  by another user in favour of our own /run/user/{uid}, so the restart
  actually succeeds after su/sudo -u instead of only failing cleanly.

Regression tests cover the EACCES readiness probe, foreign-dir replacement,
and that preflight raises UserSystemdUnavailableError (not PermissionError).

* fix(agent): do not let hygiene idle timeouts block in-agent compression

Session hygiene persists compression_failure_cooldown_until after a
30s no-progress watchdog so the pre-agent pass can skip. The
in-conversation compressor read the same column and then refused to
run even though its own budget is sufficient.

Ignore hygiene idle-timeout errors on the in-agent path. Real
aux-model faults such as rate limits still block.

Fixes #86972

* fix(agent): clear in-memory cooldown when hygiene overwrites the shared row

A later hygiene idle-timeout write can replace an aux-model cooldown on
the shared column. Drop the in-memory timer on that refresh so the
in-agent compressor is not still blocked after the DB row is hygiene.

* fix(gateway): persist hygiene failure cooldown rung

* fix(gateway): offload hygiene streak persistence

* fix(skills): resolve skills dir before relative_to so junction installs work (#86971)

* fix(telegram): honor group_allowed_chats in early auth under multiplex profiles (#87132)

With gateway.multiplex_profiles enabled, the primary Telegram message
handler is the closure returned by _make_default_profile_message_handler(),
so its __self__ is absent. The early intake filter
(_is_user_authorized_from_message) recovered the GatewayRunner via
self._message_handler.__self__ and, finding none, fell back to env-only
authorization — never evaluating the configured chat allowlist through
GatewayRunner._is_user_authorized(). Every non-global sender was then
default-denied in an explicitly allowlisted group.

Prefer the platform-bound authorization callback registered via
set_authorization_check(): it routes through the runner's full auth chain
(platform + group allowlists, pairing store, allow-all) and survives the
closure wrapping, whereas the bound-handler lookup does not. The bound
handler remains the fallback for setups without a registered callback, and
the pairing-passthrough guard for unknown DMs is preserved.

Fixes #87132

* fix(desktop): close safe preview blockers before update

* fix(desktop): identify unsafe update blockers

* fix(gateway): keep persisted model routes consistent

* fix(desktop): don't cancel the running turn on Esc while an overlay is open

The composer's global Esc-to-cancel listener (useComposerEscCancel) fires
whenever the turn is busy and the active composer matches — but overlays
(Settings, Command Center, agents, cron, …) cover the chat while the
composer stays mounted and 'active' beneath them, so pressing Esc on any
of those pages interrupted the session the user wasn't even looking at.
OverlayView's own escape-layer Esc-to-close fired too, but the stream was
already dead.

Stand Esc down with composerFocusBlockedBySurface() — the same signal the
type-to-focus path uses (BLOCKING_OVERLAY includes OverlayView's
[data-overlay-surface] marker). Esc on an overlay now closes the overlay
via its escape layer instead of canceling the stream beneath it.

Fixes #82618

* fix(caching): engage prompt caching for LiteLLM Claude on the OpenAI wire

anthropic_prompt_cache_policy() only granted Anthropic cache_control
markers to LiteLLM over the native Anthropic wire
(api_mode == "anthropic_messages"). A LiteLLM deployment exposing the
OpenAI-compatible surface instead (/v1/chat/completions, /v1/messages
-> 404) matched no grant branch and fell through to (False, False): no
cache_control injected, the system prompt sent as a plain string, and
the provider serving zero cache hits -- the entire prompt re-billed at
full price on every turn. Silent: no error, no warning, usage simply
shows 100% uncached input forever.

Add one branch after the is_anthropic_wire/is_claude case that grants
caching to Claude-family models on a LiteLLM endpoint regardless of
wire, with the native inner-block layout. Same failure class already
documented in-function for Qwen/DashScope.

Design:
- Gated on the Claude family only (is_claude); a Gemini/GPT/Qwen route
  through the same proxy must not receive markers (they may reject the
  cache_control block format -- cf. the DeepSeek/OpenCode exclusion).
- Matches on provider string OR base_url host, since provider naming
  varies per install (litellm, custom:litellm, or a bare custom alias
  pointed at a LiteLLM host).
- prompt_caching.cache_ttl: false still wins (the _cache_disabled early
  return is untouched).
- Generic strict OpenAI-wire custom providers (e.g. Fireworks) remain
  excluded -- verified by the existing over-reach regression test.

Tests: adds TestLiteLLMOpenAIWire covering the grant (several model
spellings x provider/host signals), no-over-reach (non-Claude on the
same proxy get nothing; operator disable wins), and adjacent behavior
(LiteLLM in Anthropic proxy mode still native layout). Full module:
43 passed.

Closes #84506. Original diagnosis, patch design, and measurements by
@ottosulin.

* fix(caching): use the envelope layout for LiteLLM Claude on the OpenAI wire

Follow-up to the salvaged LiteLLM cache grant. The grant itself is right;
four things about how it was scoped were not.

1. Layout. The branch returned the native inner-block layout
   (use_native_layout=True) on api_mode == "chat_completions". That layout
   writes a TOP-LEVEL msg["cache_control"] on role:tool and empty-content
   messages and depends on the Anthropic adapter to relocate it into the
   block — but that adapter only runs for api_mode == "anthropic_messages"
   (agent/transports/anthropic.py registers there), and the
   chat_completions transport does no relocation. Measured on a 3-tool-turn
   transcript: 2 of the 4 available breakpoints landed on markers the
   provider never sees. Worse, when LiteLLM itself relocates a top-level
   marker for an OpenRouter-backed Claude route
   (OpenrouterConfig._move_cache_control_to_content), the marker lands on
   an empty assistant turn and produces a cache_control-marked empty text
   block — the HTTP 400 "text content blocks must contain" shape already
   guarded in agent/anthropic_adapter.py (#69512). Switched to the envelope
   layout, matching every other OpenAI-wire grant in this function:
   4 of 4 breakpoints honored, zero empty blocks.

2. Host matching. `"litellm" in base_url_hostname(...)` is the substring
   false-positive class base_url_hostname's own docstring warns against; it
   granted Anthropic markers to notlitellm.example.com,
   foolitellmbar.example and friends. Replaced with a label-token match in
   a named helper, so "litellm" must be a whole dot- or hyphen-delimited
   token. All three of the original test hosts still match; a "litellm"
   path segment on an unrelated host still does not.

3. Transport gate. `not is_anthropic_wire` also swept in codex_responses,
   bedrock_converse and codex_app_server. Gated on
   api_mode == "chat_completions" explicitly.

4. Operator override. The grant is inferred from a provider/host name, but
   the custom-provider capability lookup was gated on is_anthropic_wire, so
   an explicit `prompt_caching: false` for the route+model was honored on
   /v1/messages and silently ignored on /v1/chat/completions. The lookup
   now also runs for a LiteLLM route, and its layout follows the transport
   rather than the declaration (an explicit `true` must not promote a
   chat_completions request to the native layout).

Tests: 64 passed. Adds the wire-shape contract the original matrix was
missing (asserts no breakpoint sits on the message envelope, rather than
only checking the returned tuple), plus lookalike-host, other-transport,
and both operator-override directions. All five guards mutation-checked —
reverting each fix turns the corresponding test red.

* fix(caching): match the litellm provider id token-wise too

Self-review follow-up. The previous commit fixed substring matching on the
HOST but left the provider-id side as a bare substring, so a user-named
provider like `custom:notlitellm` or `mylitellmthing` still matched and was
handed Anthropic markers — the same bug class, half-fixed.

Both signals now match `litellm` as a whole delimited token via a shared
helper. Real spellings (`litellm`, `custom:litellm`, `litellm-router`, and
the already-lowercased `LiteLLM`) still match; lookalikes no longer do.

Tests: 71 passed. Adds lookalike-provider and real-spelling guards; both
new guards mutation-checked. Differential matrix over 2688 configs vs
origin/main: 60 changes, every one a Claude model on a genuine LiteLLM
route getting the envelope layout, zero pre-existing routes altered.

* perf(caching): narrow the widened capability lookup to the LiteLLM grant

Self-review follow-up, caught by benchmarking the previous commit.

Widening the custom-provider capability-lookup gate to `is_anthropic_wire or
_is_litellm_route(...)` made EVERY chat_completions route with a litellm-ish
provider/host enter the lookup, including non-Claude models that the grant
branch below can never match. Measured on a route with no config.yaml
(the uncached worst case) that was ~7.5us -> ~1528us per evaluation.

Narrowed the gate to the exact condition the LiteLLM branch grants on
(chat_completions + Claude + litellm route), computed once into a local and
reused by the branch itself so the predicate no longer runs twice.

Measured with a realistic config.yaml present (mtime cache warm), vs
origin/main:
  live-agent policy      20.6us -> 61.7us
  destination planning  219.3us -> 347.7us

Sub-millisecond and scoped to the routes that actually opted in. The
earlier 1.5ms figures were a tempdir artifact: load_config_readonly's
mtime cache cannot engage when no config.yaml exists, which is never true
of a real install. Non-LiteLLM and non-Claude routes are unaffected
(openrouter Claude measured flat at ~7.9us).

Tests: 82 passed across the policy and TTL-propagation modules.

* test(caching): pin signal precedence and the openrouter-host opt-out

Review follow-up. Three coverage gaps in the LiteLLM matrix:

- The operator opt-out on a litellm-named provider pointed at an OpenRouter
  host. That route previously took the OpenRouter branch and ignored an
  explicit per-model `prompt_caching: false`; it is the only cell in the
  differential matrix where the salvage REMOVES caching, so pin it as
  intended rather than leaving it to be read as a regression.
- Signal precedence: an explicitly litellm-named provider grants even on a
  lookalike host, because the provider id is an independent signal and only
  the host-derived signal is token-gated. Intentional, now documented.
- A hyphen-delimited host label (`my-litellm-gw.internal.example.com`),
  which the token matcher handles but nothing exercised.

Traded the redundant `claude-3-7-sonnet` parametrize cell for the new host
case, so the matrix covers more shapes with the same cell count.

Tests: 83 passed. All three production fixes re-mutation-checked against
the final stack.

* fix(web_server): discover root user plugins under profile-scoped processes

When the backend is spawned profile-scoped (`--profile <name>` sets
HERMES_HOME=<root>/profiles/<name>), _discover_dashboard_plugins()
scanned only get_process_hermes_home()/plugins — the profile directory,
which has no plugins/ content. Pooled per-profile backends therefore
discovered zero user plugins, mounted no plugin API routes, and every
plugin REST call fell through to the SPA catch-all 404.

Also scan get_default_hermes_root()/plugins (which unwraps
<root>/profiles/<name> to <root> and leaves a custom HERMES_HOME
untouched when it is itself the root), matching how hermes_cli.plugins
resolves install locations. The profile home is scanned first, so a
profile-local plugin of the same name stays authoritative via the
existing seen_names dedupe.

Adds regression tests for root-plugin discovery under a profile-scoped
process and for profile-over-root precedence.

Fixes #87197 (plugin discovery half — the misleading /api/* catch-all
half is addressed separately in #87270).

* fix(telegram): keep /loop and synthetic sends in the active DM topic

Fixes #87051

* fix(desktop): show failed status for timed-out subagents in fallback stream path

Fixes #87200

* fix(desktop): match custom provider aliases in model catalog menu

Fixes #87035

* chore(contributors): map emails for P2-sweep salvage wave

* fix(desktop): avoid PowerShell parent marker boot gate

* fix(desktop): give Windows start-marker PowerShell probe a 30s budget

PowerShell 5.1 cold starts take 2.4-8s on affected Windows hosts, so the
shared 3s execText timeout hard-failed the parent start-marker probe for
any PID that still needs the PowerShell path (e.g. backend children).
Make execText's timeout overridable and raise the marker probe to 30s.

Fixes #87169

* perf(desktop): hydrate transcripts with a small tail page + on-demand older-page backfill

Replace the fixed 500-message REST hydration (getLatestSessionMessages)
with a 120-row newest-first tail page. When the page comes back full, a
new per-session tail store records "possibly truncated + next offset";
"Show earlier" — once the DOM budget and the in-memory store window are
both exhausted — fetches the next older page via the new
getOlderSessionMessages helper (order latest + offset, matching the
backend's back-from-newest paging semantics) and prepends it to the
session store, deduped by durable row id and race-guarded against
session switches. Legacy backends without pagination metadata fall back
to the one-shot full transcript and retire the action.

Tail-page refreshes (background sync, post-turn rehydrate, re-activate,
cold-resume prefetch) graft the refreshed tail onto any backfilled
prefix instead of clobbering it, preserving reference identity on
no-ops. includeCompacted stays on every read — compaction-archived rows
remain part of the durable display history.

* feat(desktop): MCP fleet cost/usage overlay with schema token estimates and 30-day usage

Each configured server row on the MCP Capabilities page now shows what it
costs and whether it earns its keep:

- ~per-call token estimate of the server's tool schemas, summed over ENABLED
  tools only (ceil(schema_chars/4) via the existing include/exclude filter)
- 30-day usage count from getUsageAnalytics(30), cached per scope profile
  like the Toolsets tab's toolCallsCache, mapped to servers via the
  mcp__<server>__<tool> registry-name convention (tools/mcp_tool.py)
- a subtle muted "unused" pill on enabled, probed-ok servers with nonzero
  schema cost and zero 30-day uses — never a dialog

Backend: the /api/mcp/servers/{name}/test probe now fills an additive
per-tool `schema_chars` (length of the SAME converted registry schema the
agent registers). Older backends omit it → renderer shows counts only;
older renderers ignore the extra key. Display-only: nothing changes what
schemas are sent to models, no config knobs.

i18n keys (costTokens/usage30d/unusedPill) added to types/en/zh/zh-hant/ja
(ar inherits en via defineLocale overrides). Pure math lives in
lib/mcp-cost.ts with unit tests; Python wire shape pinned in
tests/hermes_cli/test_web_server_profile_unification.py.

* fix(tui): map lineage edit ordinals past compression prefix

Desktop/TUI count full displayed lineage after compression, but
prompt.submit validated truncate ordinals against tip-only history.
Translate via display_history_prefix and recover stale 4018s on Desktop.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(desktop): hoist GatewayMock type into #82462 edit recovery suite

CI typecheck failed because GatewayMock lived only in the previous describe.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(desktop): hermes:// deep link to install MCP servers with explicit confirmation

Adds hermes://mcp/install?name=NAME&config=B64 (base64url or standard
base64 JSON), mirroring Cursor's mcp/install deep link, so vendors and
docs can offer an "Add to Hermes" button.

- Electron: the existing generic hermes:// handler already forwards
  {kind, name, params}; only its comment is updated (no new handler).
- Renderer: use-desktop-integrations routes kind=mcp/name=install into
  a pending-install store; a new confirmation dialog shows the server
  name and the FULL pretty-printed config (attacker-controllable input),
  with a prominent caution for stdio command entries. Nothing is written
  until the user confirms; existing names require a rename or cancel.
  On confirm the server is merged over a fresh fetch of the current map
  via saveMcpServers, then navigation lands on /skills?tab=mcp&server=…
  so useDeepLinkHighlight focuses the new row.
- Validation: name ^[A-Za-z0-9._-]{1,64}$; config must decode to an
  object with a string http(s) `url` or a string `command` (never both);
  payloads over 32KB rejected; failures surface as a toast.
- Pure parser in src/lib/mcp-deeplink.ts with unit tests (url shape,
  command shape, bad base64, non-object, javascript: URL, oversized).
- i18n keys in types + en/zh/zh-hant/ja/ar.
- Docs: "Add to Hermes link" section in the MCP config reference.

* fmt(js): `npm run fix` on merge (#87599)

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>

* feat(desktop): background MCP health checks with re-auth nudges

MCP server problems (expired OAuth tokens especially) were only
discovered when the user visited the MCP page and a probe ran. Now a
renderer-side background checker (store/mcp-health.ts) sweeps the
active profile's enabled HTTP/SSE MCP servers on gateway connect and
every 30 minutes, and fires an in-app notification with a "Sign in"
action ("<name> MCP needs re-authentication") that navigates to the
MCP page with ?server=<name> so useDeepLinkHighlight focuses the
server and its Authenticate button. Navigation only — OAuth flows are
never auto-launched.

stdio servers are deliberately excluded: probing a stdio server SPAWNS
a local process, so a background timer must never touch them. Only
url-shaped servers (where OAuth expiry lives) are swept, sequentially.

The tab's probeCache/serverFingerprint/probeKey/NEEDS_AUTH_RE moved to
a shared lib/mcp-probe-cache.ts (behavior identical) so the page and
the checker share one probe cache and its 5-minute TTL — neither
surface re-probes what the other just learned.

Notifications fire only on a TRANSITION into needs-auth/error (pure
state machine, unit-tested), hard-capped at one per server per app
session, keyed per profile. Profile switches drop pending timers and
re-arm for the new profile; sweeps never run while the gateway is
disconnected. No new config knobs. i18n keys added across
en/zh/zh-hant/ja/ar + types.

* fmt(js): `npm run fix` on merge (#87607)

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>

* fix(desktop): scope messaging to active remote profile

Complete the sidebar profile-scope contract across remote Electron routing, older-backend fallbacks, standalone messaging refreshes, and pagination. Reject stale profile responses and keep the explicit all-profiles view unified.

Co-authored-by: 墨綠BG <s5460703@gmail.com>

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>

* fix(desktop): retain messaging totals per profile

Key resolved platform totals by Desktop profile and source so profile switches neither inherit another profile's count nor discard a count that was already resolved. Keep the full reset for connection configuration changes.

Co-authored-by: frendo <frendo.wu@gmail.com>

* fix(desktop): ignore stale messaging page responses

Sequence per-profile platform pagination so an older overlapping response cannot replace a newer, larger page.

* docs(desktop): document profile scope helpers

Add JSDoc to the exported helpers introduced by the profile-scoped sidebar change.

* fix(desktop): reject stale profile refreshes

* fix(desktop): harden profile-scoped refreshes

* fix(desktop): resolve sessions from sidebar caches

Consult messaging and cron caches before the by-id fallback so opening a sidebar row neither depends on a redundant network lookup nor duplicates it into regular recents.

Co-authored-by: protas-box <protas.box@icloud.com>

* fix(tools-config): stop reconfigure flow clobbering image_gen.use_gateway on managed FAL rows

The Nous Subscription image_gen row carries imagegen_backend="fal", so
_reconfigure_provider's post-model-picker step ran

    img_cfg["use_gateway"] = False

unconditionally at two sites, immediately after the managed branch had
written use_gateway=True. A user who picked Nous Subscription and then
re-entered `hermes tools` to change the model was silently flipped onto
their personal FAL_KEY.

Same bug class as fe63353cb, which fixed the plugin-provider selector
but missed these two legacy-backend sites in the reconfigure flow. Both
now write bool(managed_feature), matching the existing correct site in
_configure_provider.

Adds regression tests driven through the real TOOL_CATEGORIES managed
row; sabotage-verified (tests fail with the old behavior restored).

* feat(desktop): make the multi-gateway Connections registry discoverable

The multi-connection registry (Settings -> Connections) shipped with no
entry point outside the settings nav, and the product-owner report was
blunt: 'I didn't see any obvious way to hook up multiple gateways.'

- Profile rail: a plug pill pinned beside Manage ('Connect another
  Hermes gateway...') deep-links to /settings?tab=connections. Always
  visible, including for single-profile first-run users.
- Command palette: Settings -> Connections is now a searchable entry
  (keywords: add gateway, remote, ssh, cloud, instances, registry).
- i18n: profiles.connectGateway added to en/types/zh; other locales
  fall back through defineLocale.
- Tests: profile-rail-connect.test.tsx covers the deep link and the
  single-profile visibility guarantee.

* docs: full multi-gateway setup guide for Hermes Desktop

Expand user-guide/multi-connection-desktop.md into a complete setup
walkthrough: where to find the pane (settings nav, profile-rail plug,
command palette), the exact add-connection editor fields (Name,
Gateway URL, Authentication: Session token/OAuth, SSH host), Primary /
This device pills, Test semantics, agent roster + profile-rail
switching and per-profile session/cron/messaging scoping, token
storage via Electron safeStorage with the keyring-less Linux plain-
text opt-in, and troubleshooting. All quoted labels match the desktop
i18n strings. Cross-link the rail entry point from desktop.md.

* fix(desktop): never resolve a missing named gateway scope to the primary

activeGateway() fell back to the primary gateway when the active key named
a registry-agent scope (conn:<id>::<profile>) whose secondaries entry had
been evicted — e.g. closeSecondaryGateways() during a soft gateway switch —
so sends and session ops silently executed against the WRONG machine.

A named scope now resolves to its own socket or null, and every eviction
path (closeSecondaryGateways, pruneSecondaryGateways) explicitly restores
the primary as active when it evicts the active scope, keeping the
'activeKey always resolves' invariant with the atoms following.

* fix(desktop): sync connection atoms and share the switch mutex for agent activation

ensureGatewayForAgent (the SDK ensureAgent door) skipped the two invariants
the profile path provides:

- $connection / $activeGatewayProfile were only updated when a socket was
  freshly dialed (setConnection inside openSecondary), so activating an
  ALREADY-OPEN registry agent left both describing the previous backend —
  /api/fs, /api/media and image.attach routed to the wrong machine (same
  class as #46651) and newSessionInProfile targeted the stale profile.
- Activations bypassed the gatewaySwitch mutex, so a rapid agent/profile
  interleave could complete out of order with the earlier setActive()
  landing last.

Add profile.ts ensureGatewayAgent: the (connectionId, profile) analogue of
ensureGatewayProfile that shares the same gatewaySwitch mutex, moves
$activeGatewayProfile on every activation, and resyncs $connection from
getConnectionFor (best-effort, like the profile path). The SDK ensureAgent
now routes through it; local/null connectionId falls through to the
profile path unchanged.

* feat(desktop): expose busy turn flags on plugin SDK

Plugins can now read host.state.busy and host.state.awaitingResponse
for the focused chat. These follow the same session slice the chat pane
uses, so a draft falls back to the global flags and a background turn
does not leak.

* fix(desktop): make plugin SDK turn flags follow the focused chat

Follow-up to the salvaged #87558 commit: the PR's docs promised the flags
follow "the focused chat", but PRIMARY_SESSION_VIEW is the primary
workspace tab only — a focused session TILE would read the wrong chat.
Wire host.state.busy / host.state.awaitingResponse through the focused
slice ($focusedStoredSessionId / $focusedSessionState), same semantics
as the statusbar busy pulse, with the primary view (and its draft
fallback) while the workspace holds focus.

Adds a tile-focus vitest case and corrects the docs wording.

* feat(desktop): paste-anything MCP server import

Add a compact Import popover to the MCP Capabilities page that accepts
anything a user might copy from an MCP server README and infers the
server config:

- mcp.json snippets (mcpServers-wrapped, bare name->config maps, single
  unnamed server objects, Cursor/Claude `type` normalized to `transport`)
- bare npx/bunx/uvx/node/docker command lines (name inferred from the
  package basename, e.g. server-filesystem -> filesystem)
- `claude mcp add NAME [--transport http|sse] [-e K=V] [-H ...] [--] CMD
  ARGS...` and `claude mcp add NAME URL`
- bare http(s) URLs (name inferred from the hostname)
- Cursor deeplinks (cursor://anysphere.cursor-deeplink/mcp/install with
  a base64-encoded JSON config payload)

The parser is a pure module (src/lib/mcp-import.ts) with unit tests for
every format plus garbage input. The popover previews the inferred
name + config and, on confirm, merges the entries into the editor draft
exactly like addServer's starter entry: unique keys, dirty (unsaved)
draft, first new block focused. Placeholder env values (YOUR_KEY,
TOKEN_HERE, ...) are kept verbatim for the user to edit in the editor
before saving.

i18n keys added under settings.mcp for en, zh, zh-hant, ja (ar falls
back through defineLocale).

* feat(desktop): running is not busy

Gate composer submit and plugin host busy on the target session slice, not a leftover foreground busyRef. Staff can keep typing while a worker session is running.

Includes the follow-up test that submit uses the target session busy flag.

* fix(desktop): lint and map contributor email for running-is-not-busy

Drop the redundant Boolean() on selected in $primaryBusy and add the
professorpalmer9@gmail.com mapping so attribution CI can resolve the PR.

* feat(desktop-sdk): expose focused-session state atoms to plugins

Disk plugins read app state exclusively through host.state, which only
exposed the primary workspace tab ($activeSessionId). In the multi-tile
layout, clicking a tile never touches that atom — and tile focus is a
pure renderer concern, invisible to both gateway RPC and the event
stream — so a plugin cannot follow the session the user is actually
looking at.

The core statusbar solves this same problem by reading the focused-
session atoms (use-statusbar-items.tsx). Widen the generic plugin
surface with the same signals, per the contribution rubric:

- host.state.focusedSessionId — runtime id of the focused session
  (interacted tile, else the primary), the key for session.* RPC
- host.state.focusedStoredSessionId — durable id for navigation and
  session-list matching
- host.state.focusedUsage — live streamed UsageStats projection
  (context_used/max/percent, tokens, cost_usd), no RPC needed

Additive only; no existing behavior changes. tsc --noEmit clean.
Verified end-to-end with a disk plugin that now tracks the focused
session across tiles.

* test(desktop-sdk): contract-test the focused-session host.state atoms

Locks the plugin-facing contract: the focused atoms exist as readonly
nanostores, mirror the primary session while no tile is focused, project
the focused session's usage, and — the behavior this PR exists for —
follow the interacted tile while the primary-only $activeSessionId
stays put.

* fix(desktop-sdk): type focusedUsage as Partial<UsageStats>, fix expect arity

ClientSessionState.usage is Partial<UsageStats> (app/types.ts) — the
backend streams whichever fields changed — so the computed produces
ReadableAtom<Partial<UsageStats> | null>. Annotate the entry honestly
instead of claiming full UsageStats, and document the fallback rule for
plugin authors. Also collapse the three-argument expect() calls in the
contract test (vitest takes one message arg). Addresses triage review on
PR #80461.

* fix(desktop-sdk): address adversarial review — type honesty, real tile coverage, docs

Independent second-pass review found three gaps:

- focusedUsage is null | UsageStats, not Partial — ClientSessionState.usage
  is the full type (app/types.ts) and its only write site seeds the four
  required fields before merging (gateway-event.ts). The earlier Partial
  annotation traced the wrong type (SessionRuntimeInfo, an RPC payload).
  Comment now names the genuinely optional fields instead.
- The tile-focus contract test never seeded $sessionTiles/$sessionStates,
  so it proved focusedStoredSessionId…
cwliao pushed a commit to cwliao/hermes-agent that referenced this pull request Aug 17, 2026
Follow-up to the NousResearch#28953 salvage:

- Extract _resolve_block_from_details() so resolve_pre_tool_block and
  _dispatch_pre_tool_call_hooks share ONE fail-closed approval-gate
  implementation. This also gives the new dispatcher the observability
  context wrapping around request_tool_approval that the original PR's
  inlined copy lacked.
- Update sibling tests that patched resolve_pre_tool_block at the three
  migrated dispatch sites to patch _dispatch_pre_tool_call_hooks with the
  (block_message, modified_args) tuple contract.

Verified: 448 targeted tests green; E2E with a real shell hook in an
isolated HERMES_HOME rewrote a live write_file call (path + content)
through handle_function_call, with block and negative paths intact.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint comp/plugins Plugin system and bundled plugins comp/tui Terminal UI (ui-tui/ + tui_gateway/) P2 Medium — degraded but workaround exists sweeper:blast-contained Sweeper blast radius: contained — one narrow path / opt-in / few users sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades sweeper:risk-security-boundary Sweeper risk: may affect sandboxing, auth, credentials, or sensitive data type/feature New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants