Skip to content

fix(terminal): cap watch_patterns notifications over a process's lifetime - #93532

Closed
chelsealong wants to merge 1 commit into
NousResearch:mainfrom
chelsealong:fix/watch-patterns-lifetime-cap-93513
Closed

chelsealong wants to merge 1 commit into
NousResearch:mainfrom
chelsealong:fix/watch-patterns-lifetime-cap-93513

Conversation

@chelsealong

Copy link
Copy Markdown

What does this PR do?

Closes the gap in watch_patterns rate limiting that lets a recurring
pattern force an unbounded number of full-context agent turns.

ProcessRegistry._check_watch_patterns() already has a per-session rate
limit (1 notification / 15s) plus a 3-consecutive-strikes auto-disable, but
that strike counter only advances when a match is dropped inside an
active cooldown window. A match that arrives after the cooldown has already
expired is treated as "clean" and resets the strike counter to 0 — so a
pattern that recurs at a cadence just above the 15s floor (a service
restarted repeatedly over hours/days, for example) never trips the
strike-limit disable. Every one of those matches still synthesizes an
[IMPORTANT: Background process ... matched watch pattern ...] message and
forces a brand-new agent turn carrying the entire conversation.

This PR adds a lifetime cap independent of the strike counter:
WATCH_LIFETIME_MAX_HITS (8). Once a session has delivered that many
watch_match notifications over its whole life — regardless of how cleanly
spaced they were — watch_patterns is disabled and the session falls back
to notify_on_complete semantics (one notification when the process
actually exits), reusing the exact same disable/promotion path the
strike-limit already uses.

This targets the part of #93513 that the existing three-layer defense
(97d54f0e4d) doesn't cover: sparse-but-recurring matches. Deeper parts of
that issue — a token-budget guard before triggering a turn at all, and the
separate WebSocket re-attach defect — are larger, architectural changes to
the gateway's turn-triggering and connection-recovery paths and are left
for a maintainer/dedicated follow-up rather than bundled into this small,
mechanical fix.

Related Issue

Fixes #93513 (partial — see scope note above)

Type of Change

  • 🐛 Bug fix (non-breaking change that fixes an issue)

Changes Made

  • tools/process_registry.py — add WATCH_LIFETIME_MAX_HITS constant and enforce it in _check_watch_patterns(), queuing the same watch_disabled summary event used by the strike-limit path.
  • tools/terminal_tool.py — update the watch_patterns tool-schema description to mention the lifetime cap.
  • tests/tools/test_watch_patterns.py — add TestLifetimeCap covering sparsely-spaced matches (cooldown reset before every match, so the strike counter never advances) that must still stop after the cap.

How to Test

  1. python3 -m pytest tests/tools/test_watch_patterns.py -q
  2. To see the new test fail without the fix: git checkout HEAD~1 -- tools/process_registry.py tools/terminal_tool.py && python3 -m pytest tests/tools/test_watch_patterns.py -k TestLifetimeCap -q (fails with ImportError: cannot import name 'WATCH_LIFETIME_MAX_HITS'), then git checkout HEAD -- tools/process_registry.py tools/terminal_tool.py to restore.

Actual output

$ python3 -m pytest tests/tools/test_watch_patterns.py -q
.....................                                                    [100%]
21 passed in 0.63s

$ git checkout HEAD~1 -- tools/process_registry.py tools/terminal_tool.py
$ python3 -m pytest tests/tools/test_watch_patterns.py -k TestLifetimeCap -q
F                                                                        [100%]
FAILED tests/tools/test_watch_patterns.py::TestLifetimeCap::test_sparse_matches_disable_after_lifetime_cap - ImportError: cannot import name 'WATCH_LIFETIME_MAX_HITS' from 'tools.process_registry'
1 failed, 20 deselected in 0.43s
$ git checkout HEAD -- tools/process_registry.py tools/terminal_tool.py

$ python3 -m ruff check tools/process_registry.py tools/terminal_tool.py tests/tools/test_watch_patterns.py
All checks passed!

tests/tools/test_process_registry.py and other neighboring suites were
also run; the handful of pre-existing failures there (missing psutil /
ptyprocess in this sandbox, a live-system-guard block on real os.kill)
reproduce identically on unmodified main and are unrelated to this change.

Checklist

Code

Documentation & Housekeeping

  • I've updated relevant documentation (README, docs/, docstrings) — docstring in _check_watch_patterns updated
  • I've updated cli-config.yaml.example if I added/changed config keys — N/A, no new config key
  • I've updated CONTRIBUTING.md or AGENTS.md if I changed architecture or workflows — N/A
  • I've considered cross-platform impact (Windows, macOS) — N/A, pure in-memory rate-limit logic, no platform-specific behavior
  • I've updated tool descriptions/schemas if I changed tool behavior — updated watch_patterns description in TERMINAL_SCHEMA

AI assistance disclosure

This PR was prepared by an autonomous coding agent (Claude, via an
OSS-contribution pipeline) that read the issue, located the relevant rate
limiting code, wrote the fix and the failing-first test, and verified it.

…time

Per-session rate limiting only counts consecutive strike windows, so a
pattern that recurs at a cadence just above WATCH_MIN_INTERVAL_SECONDS
(e.g. a service restarted repeatedly over a day) never trips the
existing strike-limit disable — each match lands in its own clean
cooldown window. Every one of those matches still forces a full-context
agent turn, which stalls the event loop on large sessions (NousResearch#93513).

Add WATCH_LIFETIME_MAX_HITS: once a session has delivered this many
watch_match notifications over its whole life, disable watch_patterns
and fall back to notify_on_complete, reusing the existing disable path.
@alt-glitch alt-glitch added type/bug Something isn't working comp/tools Tool registry, model_tools, toolsets tool/terminal Terminal execution and process management P2 Medium — degraded but workaround exists labels Aug 24, 2026

@trevorgordon981 trevorgordon981 left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed from a Hermes audit pass (tested the new LifetimeCap test on this branch: passes). The cap is correctly independent of the consecutive-strike path and the hit that reaches the cap is still delivered. One accounting issue inline, plus a persistence note.

Separate note on persistence: _watch_hits and _watch_disabled are not written by _write_checkpoint, so a gateway restart while a long-lived watch process is still alive rebuilds the recovered session with the counter reset (recover_from_checkpoint) and re-arms the watch. The cap therefore does not survive a restart — the same limitation the existing strike counter already has, so not a regression, but worth flagging if the cap is meant to be a hard floor against recurring-restart spam.

Comment thread tools/process_registry.py
# Lifetime cap: this match is delivered (it already earned it),
# but disable further ones regardless of how cleanly spaced
# they are — see WATCH_LIFETIME_MAX_HITS above.
lifetime_exhausted = session._watch_hits >= WATCH_LIFETIME_MAX_HITS

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

_watch_hits increments at L261 before _global_watch_admit (L297) can reject the delivery, so a hit suppressed by the global circuit-breaker still counts toward the cap. When the global breaker trips, this block can disable the session — and the watch_disabled message claims "8 delivered matches" — after fewer than 8 notifications ever reached the user. Consider counting only globally-admitted deliveries, or drop "delivered" from the message.

teknium1 pushed a commit that referenced this pull request Aug 24, 2026
…ounting, Nth-delivery promotion, docstring

Follow-ups on top of the cherry-picked #93532 cap:
- Regression tests: suppressed (in-cooldown) matches must NOT consume the
  lifetime budget; the cap trips exactly at the Nth DELIVERED match and
  promotes to notify_on_complete with the watch_disabled summary queued
  right after the final match.
- Extract _emit_lifetime_watch_disabled() and emit the summary even when
  the global breaker drops the final match, so the user always learns why
  watching went quiet (parity with the strike-limit path).
- Mention the lifetime cap in the terminal tool docstring (the schema text
  was already updated by #93532).

Refs #93513
@teknium1

Copy link
Copy Markdown
Collaborator

Merged via #93670 with your commit cherry-picked (authorship preserved); we hardened delivered-only counting and documented the cap in the tool schema. Thanks @chelsealong!

@teknium1 teknium1 closed this Aug 24, 2026
RFingAdam added a commit to RFingAdam/hermes-agent that referenced this pull request Aug 27, 2026
…ic OAuth refresh races) (#3)

* fix(dashboard): follow scroll on implicit active-session resume (#93518)

pty_ws already fell back to the per-channel active-session file when a
/chat WS connects with no ?resume= param, replaying the whole session
into the PTY, but the frontend only pinned xterm's viewport to the
bottom when resumeParam came from the URL (#59591). The implicit path
had no way to learn a replay was happening, so the viewport stayed at
the top of the scrollback.

pty_ws now sends a one-off JSON control frame naming the session id it
resolved from the active-session file, before any PTY bytes; PTY
output itself always arrives as binary frames, so this is unambiguous
on the wire. ChatPage tracks an `effectiveResume` value seeded from
resumeParam and updated when this control frame arrives, and the
existing follow-scroll/sanitizer/hydration logic keys off it instead
of the URL param alone.

Fixes #93518.

* fix(desktop): bound reconnect awaits so a stuck IPC round-trip can't latch the UI frozen

After a liveness-probe-triggered reconnect on a remote gateway,
attemptReconnect() awaits desktop.getConnection() and resolveGatewayWsUrl()
with no timeout. If either stalls (e.g. main process wedged mid-revalidation
even though the backend itself is reachable), the `reconnecting` guard never
clears, so every later scheduleReconnect()/attemptReconnect() early-returns
forever and the UI stays stuck in "reconnecting" until the app is restarted.

Bound both awaits with a 20s timeout so a stall rejects instead of hanging;
the existing catch/finally already clears the guard and resumes backoff on
rejection. gateway.connect() keeps its own separate connect timeout.

Fixes #93454

* fix(desktop): bound the revalidateConnection() await too (#93454)

attemptReconnect() awaited desktop.revalidateConnection?.() unbounded,
immediately before the two IPC calls the previous commit wrapped in
withTimeout(). A wedged revalidation after a liveness-probe trip -
the exact trigger #93454 and this file's own comment describe - hung
that await forever, so the reconnecting guard never cleared and the
prior fix never got reached.

Wrap it in the same 20s withTimeout() (still swallowing the result via
.catch, matching its existing best-effort semantics) and extend the
regression test to hang revalidateConnection() specifically, proving
getConnection() and the socket still proceed once the stall times out.

* fix(desktop): bound the boot and gateway-switch resolveGatewayWsUrl awaits too (#93454)

Follow-up to the reconnect-loop fix: the same unbounded ticket-mint await
exists on the soft gateway-switch path and the initial boot() path. Bound
both with the same withTimeout/RECONNECT_ATTEMPT_TIMEOUT_MS so a wedged
IPC round-trip fails into the existing retry paths instead of hanging the
switch or the 'Starting Hermes…' screen forever.

* fix(desktop): group chats no longer crash when a member's connection is deleted

botRosterMeta() calls botConnectionRoute() for every sourceScoped/remoteSource
row to look up its metadata. That's a passive display lookup, but
botConnectionRoute() throws whenever connectionId can't be resolved -- which
is exactly what a stale group-chat roster row looks like once its connection
is deleted (its persisted descriptor keeps remoteSource: true but loses
connectionId). Since botRosterMeta() is called for every member on every
group-chat render, opening a group that still references a deleted
connection threw on render and crashed the pane's error boundary in a loop
that survived app restarts (the poisoned row is in Local Storage).

botConnectionRoute()'s fail-closed throw is correct and stays for its actual
callers -- routing a real request to a bot (requestForBot, session
creation, etc., covered by remote-routing-races.test.mjs). botRosterMeta()
now catches that throw and treats the row as having no resolvable route,
same as a bot with no meta at all, instead of letting it blow up rendering.

Fixes #93492

* fix(desktop): split strict connection routing from passive roster lookup

botConnectionRoute() stays the strict, throwing dispatch path for real
routing (requestForBot, session creation). botRosterMeta() is passive
display code and previously reached that throw through a bare catch,
which would have swallowed any unrelated failure the same way. It now
calls a new non-throwing resolveBotConnectionRoute() and branches on a
typed resolved | owner_removed | not_scoped status instead.

Adds witnesses for the split: the typed statuses themselves, that
strict dispatch still fails closed on an orphaned row, and that an
unrelated failure while resolving meta for a live route still
propagates instead of being swallowed.

* fix(desktop): sweep group-chat rosters when a connection is removed

Root cause of #93492: deleting a cloud/remote connection disposed its
gateways (store/gateway.ts) but never touched the persisted 'group-chats'
storage, so every member descriptor referencing the deleted connection
stayed behind as a poisoned row (remoteSource: true, connection gone) that
render-path route lookups tripped over forever.

Subscribe to the connection registry's 'removed' lifecycle push
(window.hermesDesktop.connections.onChanged, feature-detected — older
Electron mains don't emit it) and annotate every persisted group-chat
member owned by the deleted connection. Rows are marked
(sourceMissing/sourceReachable), never silently deleted: the member keeps
its identity and panes render the existing degraded 'Gateway removed'
botSourceStatus state. Writes ride updateGroupChat so the durable record
keeps its full shape, and the listener unbinds on plugin dispose.

* fix(desktop): annotate group-chat members already orphaned before hydrate

Rows poisoned before the removed-connection sweep existed (their
connection was deleted while an older Desktop ran, so no lifecycle push
ever swept them) are what made #93492 survive app restarts. After the
persisted 'group-chats' hydrate, run a pure annotate pass over the rooms:

- a descriptor that lost its connectionId (route unresolvable — the exact
  shape that threw on render) is always marked;
- a descriptor whose connectionId is absent from the live connection
  registry is marked only when the registry could actually be read —
  an unavailable registry must not read as 'everything is orphaned'.

Marked rows keep their identity and degrade to the existing 'Gateway
removed' state; nothing is deleted.

* fix(desktop): guard render-reachable route lookups against orphaned rows

Audit of the remaining unguarded botConnectionRoute() callers a pane
render can reach (#93492 follow-up to the botRosterMeta split). Each now
uses the non-throwing resolveBotConnectionRoute() and degrades on an
owner_removed row instead of throwing into the pane's error boundary:

- botWorkspaceOwnerKey / setBotsWorkspaceOwner: sidebar visibility
  listener, Bots home open, and roster context menus recompute these on
  passive UI edges; an orphaned selection now yields the name-keyed owner
  and the blocked workspace target.
- durableGroupChatMembers: rebuilt on every group send over the whole
  seated roster; one orphaned member no longer aborts the room update,
  and a swept member's degraded mark now survives the rebuild.
- useModelOptions: hook body runs during render; the query is disabled
  for an orphaned row and the picker paints its error/disabled state.
- AdvancedProfileConfig: dialog falls back to the bot's own name scope.

Strict dispatch callers (requestForBot, session creation, deleteBot,
duplicateBot, ensureBotMetadata, routines) intentionally keep the
fail-closed throw — remote-routing-races.test.mjs still asserts it.

Adds orphaned-connection-members.test.mjs covering the removed-connection
sweep, the hydrate annotate (with/without a readable registry), the
degraded 'Gateway removed' rendering of swept rows, and every guarded
caller.

* fix(desktop): isolate a failed lazy syntax-diff import from the workspace pane

React.lazy(() => import('./syntax-diff')) only has its pending state
covered by Suspense. When the dynamic import rejects (e.g. a packaged
app whose renderer window resolves to the app.asar copy of dist/ while
the chunk exists only in app.asar.unpacked, #93479), the rejection
throws past Suspense to the nearest error boundary, which is the whole
workspace ContribBoundary. One missing highlighter chunk then blanks
the entire chat transcript instead of just the diff falling back to
the plain colored DiffBody, the way markdown-text.tsx already isolates
this failure class for markdown.

Wraps LazySyntaxDiff in a local ErrorBoundary that renders DiffBody on
catch, so a failed highlight chunk degrades in place.

* fix(desktop): prefer the unpacked web dist over the asar-internal renderer index when packaged

The renderer index resolver tried APP_ROOT/dist/index.html — inside app.asar
when packaged — before the app.asar.unpacked copy that asarUnpack (dist/**)
ships and that resolveWebDist() already prefers for the embedded dashboard.
Loading the asar-internal index is how lazily imported chunks (syntax-diff-*,
shiki-*, mermaid-embed-*) end up fetched from a path that cannot serve them,
killing the workspace pane (#93479).

Reorder the candidate ladder to prefer the unpacked web dist when packaged,
following the unpackedPathFor/resolveWebDist precedent. All window loaders
(main, overlay, quick) share resolveRendererIndex, so one reorder covers
every surface. Dev behavior is unchanged: outside an asar both candidates
collapse to APP_ROOT/dist and keep the original order.

* fix(desktop): teach the torn-bundle guard to see missing lazy chunks

missingRendererAssets only checked the module refs index.html itself names
(<script type=module> + modulepreload), so a torn install whose boot-critical
files were intact but whose lazy chunks were gone passed the generation check
and died minutes later on the first React.lazy() route with 'Failed to fetch
dynamically imported module' (#93479: syntax-diff-*, shiki-*, mermaid-embed-*).

Walk the generation's module graph: for every present JS chunk, parse its
inline __vite__mapDeps filename table (the lazy-import manifest Vite bakes
into each chunk) and check those files too, transitively and cycle-safe.
resolveRendererIndex now skips a lazy-chunk-torn candidate in favor of the
intact copy instead of shipping a delayed crash.

Tests cover the mapDeps parser (definition table vs index-only call sites,
CDN refs), the exact #93479 tear shape, transitive/cyclic walks, and the
torn-vs-intact preference end to end.

* fix(terminal): cap watch_patterns notifications over a process's lifetime

Per-session rate limiting only counts consecutive strike windows, so a
pattern that recurs at a cadence just above WATCH_MIN_INTERVAL_SECONDS
(e.g. a service restarted repeatedly over a day) never trips the
existing strike-limit disable — each match lands in its own clean
cooldown window. Every one of those matches still forces a full-context
agent turn, which stalls the event loop on large sessions (#93513).

Add WATCH_LIFETIME_MAX_HITS: once a session has delivered this many
watch_match notifications over its whole life, disable watch_patterns
and fall back to notify_on_complete, reusing the existing disable path.

* test(terminal): harden watch_patterns lifetime cap — delivered-only counting, Nth-delivery promotion, docstring

Follow-ups on top of the cherry-picked #93532 cap:
- Regression tests: suppressed (in-cooldown) matches must NOT consume the
  lifetime budget; the cap trips exactly at the Nth DELIVERED match and
  promotes to notify_on_complete with the watch_disabled summary queued
  right after the final match.
- Extract _emit_lifetime_watch_disabled() and emit the summary even when
  the global breaker drops the final match, so the user always learns why
  watching went quiet (parity with the strike-limit path).
- Mention the lifetime cap in the terminal tool docstring (the schema text
  was already updated by #93532).

Refs #93513

* fix(auth): malformed OpenRouter env key no longer shadows valid credential-pool key

A malformed OPENROUTER_API_KEY in ~/.hermes/.env (truncated paste, wrong
provider's key) passed has_usable_secret's length/placeholder check and was
returned by _resolve_api_key_provider_secret before the credential-pool
fallback was ever reached, producing opaque '401 Missing Authentication
header' errors even when a valid pool entry existed (#93593).

- Add KNOWN_PROVIDER_KEY_PREFIXES (openrouter: sk-or-) and skip env values
  that mismatch a declared prefix, logging a WARNING naming the env var and
  expected prefix, then continuing to the next env var / pool fallback.
- Iterate credential-pool entries (peek first, then entries()) instead of
  only peek(), so one malformed pool entry doesn't block a valid one.
- Providers without a declared prefix are fail-open: unknown key formats
  are never rejected. Valid env keys still win over the pool (precedence
  unchanged).

Fixes #93593

* fix(pricing): support Gemini context-tiered rates in pricing snapshot (#93469)

The pricing snapshot could only express flat per-million rates, so
gemini-3.1-pro sessions with prompts over 200k tokens under-counted
input 2x ($2 vs $4/M) and output 1.5x ($12 vs $18/M).

- Add optional tier fields to PricingEntry: tier_threshold_tokens,
  input/output/cache_read_cost_per_million_above (None = flat, falls
  back to base rate per-field).
- estimate_usage_cost selects the above-threshold rates for the WHOLE
  request once usage.prompt_tokens (input + cache read + cache write)
  exceeds the threshold, matching Google's billing semantics.
- Populate gemini-3.1-pro (4.00/18.00/0.40 above 200k; alias
  gemini-3.1-pro-preview inherits) and gemini-2.5-pro (2.50/15.00
  above 200k).
- Flat entries are untouched: no threshold means no behavior change.

Reported and tier-field shape designed by @tornike14 (#93469).

Tests: below/at threshold unchanged, above-threshold tiered whole-request
pricing, cache-read tier rate and base-rate fallback, preview alias,
flat entries unaffected.

* fix(desktop): stop bot-relay drain loop from redialing a WebSocket per connection per tick

The bot relay's drain loop RPCs every registered connection through
requestGatewayForAgent's per-request lease. With no other consumer
holding the route, the refcount hit 0 after every tick and the pooled
secondary was disposed — a fresh WebSocket dial + teardown per
connection every 4s, flooding the gateway logs with connect/disconnect
pairs (#93594).

Two changes, both directions from the issue:

- Retained relay-route secondaries: retainGatewayForRelay pins a
  route's pooled socket with a counted retention (never clobbering the
  foreground 'retained' flag) for the relay's active lifetime, reusing
  the existing scheduleReconnect/full-jitter machinery on drops. The
  plugin pins each registered connection once via the new feature-
  detected host.retainProfileSocket door, reconciles pins with the
  current connection set on every drain, and releases everything in
  stopBotRelay/dispose. Local routes (null/'local') are exempt so the
  idle reaper can still reclaim spawned local backends. The live-work
  pruner also respects the pin.

- RELAY_DRAIN_INTERVAL_MS 4s -> 30s: the push path (#93091,
  bot_relay.outbox.pending) carries envelope latency, so the poll is
  purely a backstop — 30s matches LIVE_SESSION_STATUS_BACKSTOP_INTERVAL_MS.

Tests: relay-push-drain updated to the new backstop semantics; new
gateway-relay-retention.test.ts proves one socket construction across
5 drain ticks (vs 3 constructions for 3 unretained ticks) and that
release/prune/local-exemption behave; new relay-socket-retention
plugin test pins the pin-once / release-on-departure / stop-releases
contracts.

* fix(desktop): exclude cron sessions from the titlebar unread badge

Cron runs finish unwatched by design, so counting them in
$unreadSessionCount turned the titlebar badge into a permanently-lit
cron run counter (#93552). The badge now counts regular + messaging
sessions only; cron unread state stays visible on the sidebar cron
section rows, and 'Mark all as read' (markAllSessionsRead +
ackAllSessionsRead, which iterates cron rows) still clears them.

Fixes #93552

* fix(desktop): stack subsequent preview tiles as tabs instead of new right splits

Every opened file registered its preview pane with dock dir 'right', so
each open split a new zone off the right edge — three file opens made
three ever-narrower columns (#93610). The first preview still opens its
own zone docked beside main; every subsequent preview now anchors to an
existing preview-tile pane with dir 'center', so it stacks as a tab in
the same preview zone. Covers files, artifacts, and the Browser tab
alike (all flow through openPreview/$previewTabs); session tiles are
untouched.

Fixes #93610

* fix(hindsight): send configured event timestamps

Use Hermes timezone-aware timestamps for retained events and turn messages. Pass the public timestamp field supported by hindsight-client 0.6.1 and cover the final serialized request field.

* fix(hindsight): harden event timestamps

* fix(hindsight): let hindsight_retain convey event time via occurred_at

Adds an optional occurred_at (ISO-8601 date/datetime) parameter to the
hindsight_retain tool schema, threaded into the retain item's timestamp
field. When absent, the item timestamp defaults to the configured event
clock (base from PR #82928 by @ragingbulld, authorship preserved) so the
Hindsight server can resolve relative time phrases; previously no item
timestamp was ever sent and temporal memories landed with null
occurred_start/occurred_end.

Fixes #93568. Salvages #82928.

* fix(desktop): give keyboard focus a visible affordance where the global no-ring reset hides it

The unlayered *:focus-visible reset in styles.css intentionally zeroes
--tw-ring-shadow ('No focus rings, anywhere'), so any control that relied
solely on focus-visible:ring-* had no visible keyboard focus state at all.
Mirror each control's hover treatment as a focus-visible background/text
affordance instead, keeping the global reset intact:

- ui/sidebar.tsx: group label, group action, menu button, menu action,
  menu sub-button get focus-visible:bg-sidebar-accent + accent foreground
- ui/tabs.tsx: TabsTrigger gets focus-visible:bg-background + text-foreground
- ui/text-tab.tsx: focus-visible:text-foreground (matches its hover)
- chat/composer/micro-actions.tsx: pill gets focus-visible chrome-action-hover
- right-sidebar/index.tsx HEADER_ACTION_CLASS: focus-visible sidebar-accent
- right-sidebar/terminal/rail.tsx RAIL_ACTION: focus-visible chrome-action-hover
- chat/sidebar/cron-jobs-section.tsx (row body + run rows): focus-visible
  chrome-action-hover
- chat/sidebar/session-row.tsx <time>: focus-visible:text-foreground

Sweep verified: remaining focus-visible:ring-* usages under apps/desktop/src
already pair with a border/bg/text companion (button/checkbox/switch/input,
starmap share-controls) or are covered by PR #93460's row-hover work
(cron/index.tsx run rows).

Fixes #93462. Reported by @fred0m.

* fix(desktop): open a branched session in the main workspace, not just a tile

forkBranch ended by opening the branch as a session-tile and leaving the
primary selection on the parent (#69750). In the default layout there is
no visible tile pane, so branching only added a sidebar row with no
feedback in the main area — and openSessionTile no-ops when the target
is already the selected session, the common case of branching the chat
you're viewing.

Load the branch as the primary session via resumeSession instead, which
reuses the runtime already warm-cached by forkBranch's
ensureSessionState/updateSessionState calls, so it doesn't cost an extra
resume RPC.

Fixes #93444

* fix(desktop): scope branch-opens-primary to the currently selected session

forkBranch was unconditionally routing every branched session into the
main pane via resumeSession, including sidebar/background branches of
a session the user isn't currently viewing. That reintroduces the
#69750 focus-stealing bug for that path: branching a different session
from the sidebar yanked the active view away from whatever was open.
Only take over the main pane when the branch's parent is the session
already selected; otherwise keep opening it as its own tile.

* test(desktop): pin auto-speak silence across an Edge TTS fallback id rewrite

#93515 reports auto-speak reading each reply twice when the Edge TTS
streaming attempt falls back to the POST endpoint and the reply's
renderer id gets rewritten to its durable id mid-flight. That was true
before 63565fa26b, but resolveSpokenReply()'s ordinal-anchored dedupe
(landed 2026-08-19, five days before this issue was filed) already
follows the rewrite. No source change — this pins the behavior with a
regression test at the hook/store integration level, one layer above
the existing spoken-reply.ts unit tests.

* fix(desktop): hold a per-turn socket lease so group-chat member turns survive the runtime-session reaper (#93602)

A group member turn is a session-scoped RPC sequence (resume → attach →
prompt.submit → poll) issued with the runtime id its first RPC minted, but
requestForBot routes every RPC through its own request-scoped socket lease
(retained:false secondaries in store/gateway). Between two RPCs the refcount
hits 0, the leased socket closes, the gateway detaches the runtime session on
WS disconnect, the orphan reaper frees it after grace, and the next RPC —
prompt.submit, unwrapped — dies 4001 'not in memory'. The member turn aborts
and the sub-profile bot goes silent in the room.

- store/gateway: retainGatewayForAgent(connectionId, profile) — refcounted
  hold on the pooled socket with an idempotent release, mirroring the
  existing request-lease machinery.
- sdk: host.retainProfile(route) exposes the retain to plugins
  (feature-detected by consumers; older hosts keep working).
- hermes-bots plugin: runGroupChatMemberTurn acquires the lease before
  ensureGroupChatSession's first RPC and releases in finally, so the socket
  that minted the runtime id stays open across attach+submit+poll; and
  prompt.submit gets a one-shot catch-and-retry that re-resumes via the
  STORED session id on 4001-class failures (belt-and-braces for routes the
  lease can't cover). 4007 'never existed' keeps flowing to session.create.

Tests: simulated 4001 on first submit recovers via re-resume and delivers;
lease held across attach+submit (mock refcount never hits 0 mid-turn); lease
released after success AND failure; no-retainProfile host feature detection;
store-level retain/release + idempotent double-release + the unretained
disposal race.

* test: drop unused param in group-turn lease mock (lint)

* fix(desktop-update): use the system default browser for the update shim

The Windows update hand-off shim was hardcoded to Microsoft Edge
(Find-EdgeExe), so machines whose default browser is Chrome still got
an Edge --app progress window, and every run leaked a throwaway
browser profile (hermes-update-ui-<pid>) under %TEMP% that was never
removed.

- Get-DefaultBrowserExe replaces Find-EdgeExe: resolves the OS default
  browser from the UserChoice ProgId (https first, http fallback).
  ChromeHTML -> Chrome, MSEdgeHTM -> Edge; any other ProgId returns
  $null and degrades to the existing WinForms card.
- The dedicated --user-data-dir profile is now removed when the shim
  closes, and stale hermes-update-ui-* leftovers from interrupted
  runs are swept from %TEMP% in the same pass.
- --app + --user-data-dir is Chromium-only, so the whitelist is
  intentionally limited to chrome/msedge; Edge keeps
  --disable-features=msImplicitSignin to suppress the implicit MSA
  sign-in that leaks into shim windows (#88410).

* feat(gateway): add gateway.ping heartbeat wire contract

Additive WebSocket wire contract for a client-driven heartbeat.

The gateway.ready payload now advertises "heartbeat": True so clients can
discover the capability, and the WS read loop answers a gateway.ping request
with a {"ok": True} pong short-circuited BEFORE method dispatch (no method is
invoked). The WSTransport gains closed / last_inbound_at properties and a
mark_inbound() hook, updated on every inbound frame, for later liveness checks.

Backward-compatible in both directions: old clients never send gateway.ping,
and old servers simply never advertise the heartbeat flag. This is the first
slice of a WebSocket-recovery series; the follow-on slices consume this
contract (server-side transport rebind, TUI/desktop clients).

Receipts:
  bash scripts/run_tests.sh tests/test_tui_gateway_ws.py -q
  => 1 file, 7 tests passed, 0 failed (100%) in 0.8s; exit 0

* feat(shared): heartbeat and socket-generation invalidation in JsonRpcGatewayClient

The shared-client half of the gateway.ping heartbeat contract (#89958);
tracks lastInboundAt, sends pings, invalidates a silently-dead socket.
Part of #83166.

* fix(tui): heartbeat and bounded reconnect for silent WebSocket drops

the client half of the gateway.ping heartbeat contract (#89958); detects a silently-dropped socket via missed ping-acks and reconnects with bounded backoff; part of the #83166 recovery series.

* fix(cron): surface gateway liveness in cronjob tool results (#87033)

The builtin cron ticker only runs inside the gateway process. The CLI
surfaces this ('hermes cron list' / 'hermes cron status' both warn when
no gateway is running), but the model-facing cronjob tool returned a
clean success on create even with no gateway running - so the agent
confidently told the user a recurring task was scheduled while the job
could never fire.

Mirror the CLI's liveness heuristic in the tool's create path and attach
a tri-state gateway_running field to the result:

- true  -> gateway running (or a non-builtin scheduler provider owns
           firing, e.g. Chronos, which is exempt by design)
- false -> explicit warning telling the model the job is saved but will
           NOT fire until the gateway starts, so it can relay that to
           the user instead of reporting unqualified success
- null ->  probe failed; claim neither way

Fixes #87033

* fix(cron): share liveness helper with CLI and extend it to cronjob list

Follow-ups for the salvaged #93098:
- Move the tri-state liveness heuristic into hermes_cli.cron
  (_builtin_gateway_liveness) so the CLI warning and the cronjob tool
  share one implementation instead of two drifting copies.
- Surface gateway_running/warning on the list action too — an agent
  inspecting jobs in a gateway-less environment has the same silent-
  inert-job failure mode (#87033) as create. Empty lists stay quiet.

* refactor(cron): parameterize liveness warning plurality, drop fragile string replace

Simplify-pass follow-ups on the #87033 fix:
- _gateway_liveness_notice(plural=) authors both wording variants at one
  site; removes the exact-substring .replace() that would silently no-op
  if the create-path text is ever edited.
- Collapse the operator-precedence-trap conditional in list to a plain
  'if jobs' — an empty list has nothing inert and now skips the probe.
- Fix docstring/code mismatch (builder returns gateway_running: True on
  the happy path) and drop the dead try/except in
  _warn_if_gateway_not_running (the helper never raises).

* refactor(telegram): migrate _await_with_thread_deadline onto agent.deadline.run_bounded_async (#85125 2f)

The adapter's private thread-deadline helper was the ancestor of the
unified deadline layer's run_bounded_async (#85147 was extracted from
it, plus the caller-cancellation leak fix the original still lacked).
Consolidate: the helper body becomes a thin wrapper mapping
BoundedResult.timed_out back to the asyncio.TimeoutError its 9 call
sites (the PTB retry ladder) expect. ~90 duplicated lines die, along
with the adapter-local copies of the abandon-cleanup runner and the
blocked-loop faulthandler diagnostics (both live in agent/deadline.py).

Everything the call sites rely on is preserved by the unified layer:
- thread-timer deadline that survives a blocked event loop (#63309)
- abandonment of cancellation-shielded tasks (PTB/httpcore anyio init)
- detached best-effort on_abandon cleanup (no httpx pool leak per retry)
- off-loop stack dump when the loop never processes the expiry
Plus one behavior IMPROVEMENT inherited from the shared copy: a caller
cancelling the wrapper no longer leaks the inner task unobserved (the
telegram original had that leak; the extraction fixed it).

test_telegram_init_deadline.py: the #63309 diagnostics probe now pins
the shared layer's dump hook (label "telegram-init") — same contract,
new seam. Wedge + cleanup-crash tests pass unchanged.

* fix(mcp): resolve tool-call timeouts via the unified deadline layer (#85125 2g)

Both readers of the per-server MCP tool timeout (the connection's run()
and the cache-path registration) read config.get("timeout", 300) as
their own private resolution. Route them through _resolve_tool_timeout:
per-server mcp_servers.<name>.timeout still ALWAYS wins (most specific),
then timeouts.mcp.tool_call from the unified timeouts: section, then
the unchanged 300s default. Values pass through resolve_timeout's
platform clamp; resolution failure falls back to the historical default.

Default-behavior invariance pinned by contract tests (nothing
configured -> exactly 300, per-server beats section, section beats
default, invalid/failed resolution falls back).

* fix(desktop): move the Linux HUD with a native compositor drag

Wayland clients cannot place themselves, so the JS setBounds drag is a
no-op there. Make the composer bar a -webkit-app-region drag handle on
Linux (input carved out with no-drag) and let the compositor move it.

Co-authored-by: Tony Simons <214744153+asimons81@users.noreply.github.com>

* fix(desktop): debounce and re-verify zoom on Linux Wayland

Focus events fire for intra-app shifts on Wayland and Cosmic tiled
resizes can drop a just-applied zoom. Debounce focus with resize/move
and re-check a few times after the window settles.

Co-authored-by: joe0508 <75520452+joe050860@users.noreply.github.com>
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>

* fix(desktop): keep the Linux HUD clickable and recoverable

X11 cannot restore a window that has ignored the mouse, so stay solid
there. Native Wayland keeps click-through via the cursor poll. Add
desktop.ozone_platform_hint so COSMIC users can opt into XWayland for
always-on-top, and a layout reset that restores the default size.

Co-authored-by: Codex Metatron <47930664+BlakeB254@users.noreply.github.com>
Co-authored-by: DeseretSaint <202557515+DeseretSaint@users.noreply.github.com>
Co-authored-by: Shawn Wang <32839114+enwaiax@users.noreply.github.com>

* fix(desktop): add a HUD layout reset control

A persisted tall/narrow size has no way out on Linux. Put a reset next
to Exit HUD so the default size (and position, where the compositor
allows it) is one click away.

Co-authored-by: Shawn Wang <32839114+enwaiax@users.noreply.github.com>

* docs(desktop): document Linux and Wayland HUD behavior

Spell out native-compositor drag, the ozone_platform_hint escape hatch
for COSMIC always-on-top, and the snap-to-pointer no-op on Wayland.

* fix(desktop): polish HUD movement and resizing on X11

* fix(serve): emit BACKEND_PORT_IN_USE sentinel + exit 75 on port bind conflict (#93608)

A held port made 'hermes serve' print only uvicorn's bare
'ERROR: [Errno 98/10048] error while attempting to bind on address'
and exit 1 — indistinguishable from a broken backend for the desktop
spawn and wrapping scripts.

- Preflight bind probe (matching uvicorn's SO_REUSEADDR bind flags)
  before uvicorn.Server; on conflict print machine-readable
  'BACKEND_PORT_IN_USE port=<port>' + a human hint naming likely
  holders, exit 75 (EX_TEMPFAIL — existing repo convention, see
  gateway/restart.py, kanban_db.py).
- Probe-to-bind race covered: SystemExit(1) from uvicorn's own bind
  failure is re-checked and translated on both POSIX and Windows
  runner paths.
- --port 0 (ephemeral) short-circuits the probe: unchanged behavior.
- HERMES_BACKEND_READY contract untouched.
- Tests: real held-socket repro (sentinel + exit 75, sabotage-proven
  to fail as bare exit 1 without the fix), free-port boot regression,
  ephemeral-port regression, probe/classification units.
- Docs: port-conflict paragraph under 'hermes serve' in
  reference/cli-commands.md.

* fix(serve): Windows conflict probe uses SO_EXCLUSIVEADDRUSE (SO_REUSEADDR binds over live listeners on WinSock)

* fix(desktop): never kill a healthy backend on a claim probe failure; surface real stderr (#93608)

A start-marker probe failure (Get-Process timing out on a PowerShell 5.1
cold start, #87169) in claimBackendChild used to stop the freshly spawned
backend and rethrow — killing a healthy backend, triggering the renderer's
repair respawn, and looping. And because stderr piping only attached after
the claim, every before-ready failure surfaced as a bare exit code.

- extract probe + claim policy into electron/backend-claim.ts:
  processStartMarker/execText (moved verbatim from main.ts), probeStartMarker,
  and a pure claimDecision(childAlive, probe) a Windows CI lane can drive
  with real PowerShell
- probe failure + LIVE child now degrades to PID-only identity
  (pid-only:<pid> marker, WARNING logged), matching the existing
  createParentStartMarkerResolver degrade pattern; processIdentityMatches
  verifies degraded identities by PID liveness (command check still layers
  on top in backendIdentityMatches)
- probe failure + DEAD child keeps the fail-closed throw, now carrying the
  child's buffered stderr/stdout tail
- ring-buffered ~8KB output tail attached at spawn time in BOTH spawn paths
  (pool + primary); tail appended to claim errors, before-ready exit
  messages, and backend-ready's exited-before-port-announcement errors so
  the real exit reason reaches desktop.log and the boot UI
- tests: claimDecision matrix (degrade test fails against the old
  stop+throw behavior), real processStartMarker probe, ring-buffer caps,
  and output-tail suffixes on backend-ready exit errors

* fix(desktop): float and pin the HUD on Hyprland

Omarchy tiles the HUD like any other toplevel, so always-on-top and
xdg_toplevel.move never apply. Ask Hyprland to float+pin after map,
trying classic dispatch then Lua for 0.55+ configs.

* fmt(js): `npm run fix` on merge (#94046)

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>

* fix(codex): identify Hermes requests

* feat(desktop): make Cmd/Ctrl+L focus the composer from anywhere

The chord previously only acted when a terminal or preview selection
existed. A new bubble-phase window fallback now moves focus to the
composer on an unclaimed press, like the address-bar chord in a browser.

Existing owners keep priority: selection handlers claim the press on the
capture phase, a user-rebound action marks the event handled, and a
focused terminal with no selection keeps Ctrl+L as clear-screen via
composerFocusBlockedBySurface().

The chord matcher moves from the terminal feature to
src/lib/keybinds/chords.ts as isComposerChord: it now has three
consumers and the old name (isAddSelectionShortcut) was wrong at the
composer call site. The fixed panel row view.terminalSelection becomes
view.selectionToComposer because it covers preview selections too.

* refactor(desktop): derive HUD OS behavior from one windowing profile

Move, ignore-mouse, placement, resize edges, snap, cursor feed, and
overlay promote all read the same Ozone-normalized capabilities instead
of re-deriving linux/Wayland/X11 at each call site.

* fix(desktop): sort HUD windowing imports for eslint

* fmt(js): `npm run fix` on merge (#94140)

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>

* feat(desktop): let the in-app browser hold more than one tab

A URL tab used to be a singleton — every link navigated the one Browser,
so there was no way to keep a page open beside another and the strip's
"+" never appeared next to it.

A Browser tab is now a vessel with its own id: links still land in the
browser you are looking at (an agent opening five pages must not leave
five tabs behind), while the "+" mints another one on request. Tabs name
themselves after the page they are showing, since three tabs reading
"Browser" name nothing.

The "+" itself is now a pane capability rather than a session-only
button, so any pane kind that can make more of itself contributes one.

* fix(desktop): hold the typed address in the browser bar until the page moves

Committing an address dropped the field back to the url of the page you
were leaving, so typing baby.com over google.com flashed google.com back
before baby.com arrived — and nothing said a load was underway.

The address you asked for now stays in the field until the page actually
lands somewhere (a redirect supersedes it, as it should), and progress
spins inside the field beside it. The pane owns that loading state from
the moment it accepts the address, because the reach probe it runs first
delays did-start-loading.

* refactor(desktop): mint browser tab ids at random rather than by slot

Browser tabs took the lowest free slot, so an id was reused once its tab
closed. That is only safe while every store keyed by the id is wiped on
close — true today, but a discipline rather than a guarantee, and stale
state would resurface under an unrelated tab the day it lapses.

Mint like a terminal does instead: no id is ever handed out twice.

* fix(desktop): keep Home new sessions detached from the last project

Home's "+" passes path/cwd null on purpose, but null was falsy and fell
through into resolveNewSessionCwd(), so "New session in Home" (especially
the openTab path while main chat is occupied) still created under the
previous project folder and showed its branch.

* test(desktop): lock Home new-session detach against stale project cwd

Cover the null-path draft path, Home-scope createBackendSessionForSend,
and openNewSessionTile({ cwd: null }) so the last project folder cannot
leak back into a Home chat.

* fmt(js): `npm run fix` on merge (#94176)

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>

* fix(mcp): recover poisoned connections + fail fast on dead stdio transports (#85125 3b)

Fixes the four poisoned-connection classes (#81051, #77765, #84132,
#81995) with the SuspectableBackend cheap-mark/lazy-verify contract:

- mark_suspect/ensure_healthy protocol (agent/deadline.py): noticing a
  poisoned state never does I/O; the NEXT caller pays once for a health
  probe that clears the suspicion or forces a reconnect. A single
  teardown-vs-keepalive race or auth-lock corruption can no longer park
  a connection permanently — park stays reserved for genuinely
  exhausted reconnect budgets.
- keepalive failure marks the connection suspect before requesting
  reconnect; the next tool call probes and recycles if unhealthy.
- auth-classified permanent failures on a previously-proven session get
  a suspect+reconnect path instead of an immediate park.
- fast-fail (#81995): stdio child pids are tracked at spawn and an
  in-flight RPC races a child-watcher task, so a dead subprocess fails
  the call immediately with a retryable timeout instead of riding out
  the full 300s. Deliberate teardown/reconnect also fails in-flight
  calls now instead of leaving them attached to a dying transport.

Dispatch-boundary hardening for test doubles: stubbed sessions
(MagicMock/non-awaitable call_tool, absent child-watcher) fall back to
the exact pre-change inline-await semantics, so only real transports
gain the race guard.

Salvage credit: in-flight approach from #73377 (@luijoc, wedged
transport recovery) and #48069 (@arminanton, keepalive/in-flight
interaction); both PRs' bases predate main's current park/reconnect
architecture, so this is a fresh implementation of their contracts.

Tests: tests/tools/ -k mcp = 639 passed (was 22 new failures during
development; final tree zero).

* refactor(deadline): consolidate site-local tree-kills onto agent.deadline.kill_process_tree (#85125 4d)

Per-site decisions:

1. hermes_cli/_subprocess_compat.py kill_process_tree(proc) -> None:
   MIGRATED. Body now delegates to agent.deadline.kill_process_tree(proc.pid)
   via a function-local import; keeps the swallow-everything fail-open
   contract and the (proc) -> None signature (agent/shell_hooks.py imports
   it by name; _kill_git_process_tree alias preserved). The old body is kept
   verbatim as _legacy_kill_process_tree and used as fallback when the
   delegation import/call fails. A final proc.kill() is retained on the
   happy path so Popen bookkeeping sees the exit (matches old behavior).

2. tools/browser_tool.py _kill_process_tree(proc): MIGRATED, same pattern
   (delegate + _legacy_kill_process_tree fallback). Behavior delta: the old
   body sent SIGTERM then SIGKILL with zero grace between them; the shared
   primitive sends SIGKILL only. With no grace period the observable effect
   is identical, and the psutil descendant sweep now also reaches
   agent-browser's setsid'd daemon grandchild, which killpg alone missed.
   tests/tools/test_browser_npx_warmup.py's TestKillProcessTree repointed at
   the legacy fallback (its assertions describe the fallback's internals).

3. tools/code_execution_tool.py _kill_process_group(proc, escalate):
   MIGRATED. It was a plain parent+descendants terminate (then wait 5s +
   kill when escalate=True) — expressed as two delegated calls:
   kill_process_tree(pid, sig=SIGTERM), then on escalate-timeout
   kill_process_tree(pid, sig=SIGKILL). Delegation failure degrades to
   proc.kill(), mirroring the old psutil-failure fallback. Delta: the old
   body terminated children before the parent; the shared primitive
   signals the group atomically (child is a session leader via
   start_new_session=True) plus an identity-aware descendant sweep —
   strictly wider coverage, same signals.

4. gateway/status.py: KEPT BOTH SITES.
   - terminate_pid (~l305) taskkill wrapper: NOT migrated. Its contract is
     incompatible with the shared primitive — it must RAISE OSError with
     taskkill's stderr on non-zero exit (callers branch on that), falls back
     to os.kill on FileNotFoundError, and its POSIX branch is deliberately a
     single-PID SIGTERM/SIGKILL, not a tree kill. Wrapping the bool-returning
     fail-soft primitive would invert the error contract.
   - reap_gateway_children (~l2029): NOT migrated. It operates on a
     pre-snapshotted child list from a parent that is already dead
     (psutil.Process(pid) on the parent would fail), and every signal is
     wrapped in identity/ownership checks the primitive lacks: is_running()
     identity, zombie skip, and the skip-if-ppid-still-equals-parent guard,
     plus SIGTERM -> wait_procs -> SIGKILL staging and a reaped-count return.
     The coupling is the feature; migrating would delete the safety logic.

5. scripts/run_tests_parallel.py _kill_process_tree (~l253): NOT migrated.
   Dev tooling that intentionally kills by CAPTURED pgid because the direct
   child is usually already reaped (psutil/pid-based primitive cannot find
   it), and it avoids the psutil import on the test-runner hot path. Its
   docstring already documents why psutil is the wrong tool there.

New tests: tests/agent/test_treekill_consolidation.py — delegation +
raise-swallowing tests per migrated wrapper, consumer-identity checks, and
a live end-to-end probe (setsid grandchild dies through the compat wrapper,
zero survivors).

* fix(terminal): sweep setsid descendants after local timeout group-kill (#85125 4b)

LocalEnvironment._kill_process kills the process GROUP (SIGTERM ->
1s wait -> SIGKILL -> 2s wait), but a descendant that called setsid
escapes the group and survives — the #71148 orphan class, terminal
flavor (issue #84967's local sibling).

Fix: snapshot the descendant set via psutil BEFORE the first SIGTERM
(children reparent to init once the wrapper dies, so a later parent
walk finds nothing — same snapshot-before-signal design as
agent/deadline.py kill_process_tree), then after the existing group
escalation completes, SIGKILL any snapshotted survivor whose pgid is
no longer the (now-dead) group. The TERM->KILL grace window for
in-group members is preserved (interrupts use this path too), the
Windows branch is untouched, and the snapshot is fully guarded — a
broken psutil never breaks the kill path (unit-tested).

Tests: live_system_guard_bypass acceptance test spawning a setsid
grandchild and forcing the timeout path (RED on unmodified file,
GREEN after), plus a psutil-failure unit test.

Docker design note (#84967 open question 1, condensed; full note at
/tmp/4b-docker-design-note.md): the docker backend inherits
base.py:1378 _kill_process, which only proc.kill()s the HOST-side
`docker exec` client — the in-container tree (child of containerd-
shim, not the client) survives every timeout entirely. Option A,
`docker exec <cid> kill -- -<pgid>` with TERM->KILL escalation using
a PGID captured at command start, is surgical and preserves container
state but needs a live container + shell and still misses in-container
setsid escapees. Option B, container restart, is absolute (PID-
namespace teardown kills everything) but destroys all in-container
state mid-session and punishes every other consumer of the shared
persistent container. Recommendation: Option A as a best-effort
_kill_process override (degrade to today's behavior on failure);
reserve restart for the existing container-gone recovery path.

* fix(desktop): stop transcript jumps when a turn settles

Clarify remounted as a tool row once session.info flipped running=false, thinking previews collapsed their body, and the duration line grew the footer.

* fix(desktop): demote unanswered clarify cards on Stop

Latch the pending card on submit, not on seeing a request, so Stop still
collapses an unanswered question instead of leaving a disabled panel.

* fix(desktop): re-arm pending clarify cards in place

A hydrated Ask/clarify row stays complete after session or bot switch, so
the live card never mounts. Re-arm the existing transcript row and keep
the provider tool id instead of appending a duplicate at the tail.

Co-authored-by: frendo <frendo.wu@gmail.com>

* fix(desktop): restore pending_clarify snapshots on activate and resume

Replay single-question and batch snapshots from session.activate/resume,
including locked answers, and extract the helper so the session-actions
god-file is not the only owner of that wire shape.

Co-authored-by: ClintonEmok <54935030+ClintonEmok@users.noreply.github.com>
Co-authored-by: frendo <frendo.wu@gmail.com>

* style(desktop): space sibling restore import for eslint

* fix(computer-use): recreate CUA session suspect after MCP timeout (#74799)

An MCP call_tool deadline hit left the cua-driver session wedged for
all later computer-use calls. Mark the session suspect on a
concurrent.futures.TimeoutError and tear down + recreate it before the
next non-lifecycle call; healthy sessions are never restarted.

Fail-closed: the timed-out action may still have taken effect on the
remote screen, so it is never silently replayed — the error result
carries structuredContent.code=timeout_outcome_unknown with
next_step=fresh_state.

Informed by #74877 by BlackishGreen33.

Co-authored-by: BlackishGreen33 <s5460703@gmail.com>

* fix(lsp): retire clients when the protocol reader exits

* fix(lsp): abort diagnostics waits after transport death

* fix(models): OpenRouter :nitro/:floor routing variants no longer rejected by /model validation

OpenRouter's :nitro, :floor, :exacto, and :online suffixes are request-time
routing modifiers valid on any model id — /models lists only the base model.
validate_requested_model() compared the full suffixed id against the listing,
so a valid variant was either rejected outright or fuzzy-auto-corrected to
the base id, silently stripping the user's routing opt-in.

Now, for OpenRouter only, a recognized variant suffix validates the BASE id
against the live listing (and the curated-catalog soft-accept and static-
catalog fallback paths) while preserving the suffixed id for persistence and
API requests — checked BEFORE fuzzy correction. :free/:batch/:thinking
remain direct catalog SKUs and keep exact-match semantics; unknown suffixes
and unknown bases are still rejected.

Reported by JEB (Jakob's Hermes Agent) via Discord.

* feat(tui-gateway): seq-stamped event replay for lossless desktop reconnect

Server: per-session monotonic seq on every routed event frame, bounded
512-frame replay ring (64 sessions, FIFO eviction), plus two new RPCs —
session.events.since (replay newer-than-watermark, reports latest_seq +
truncated so clients detect gaps) and session.events.stats (telemetry).

Client: per-session seq watermarks recorded from live frames; after any
successful reconnect a fire-and-forget fetchReplay() drains missed events
through the normal dispatch path (recordSeq ignores non-increasing seqs,
so stale replay can never regress a watermark); focus-triggered reconnect
nudge in use-gateway-boot for the Electron unfocused case where macOS wake
skips visibilitychange.

Replay failures are swallowed by design: lossless resume is an upgrade
over the previous lossy reconnect, never a new failure mode.

* test: drop unused afterEach import (CI eslint)

* fmt(js): `npm run fix` on merge (#94230)

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>

* fix(web): keyword-align titles of six high-impression docs pages

GSC page-level data shows configuration (171K impr, pos 5.3), quickstart
(135K, 5.3), providers (123K, 5.3), web-dashboard (68K, 6.5), docker
(37K, 6.3) and the desktop app page (98K, 3.9) all losing ranking
headroom because their title/H1 are bare nouns instead of the terms
people search.

- Frontmatter title + H1 now carry the query terms on all six pages
- Desktop docs page links back to the new marketing /desktop product
  page, joining the two official properties Google sees for the query

Done by Hermes Agent (deepseek-v4-pro via nous), Nous Research.

* fix(mcp): psutil.pid_exists for stdio children liveness — Windows footgun (#85125 CI)

* fix(terminal): annotate sweep as POSIX-only for the killpg guard lint (#85125 CI)

* test: pin that install-root siblings still get parent-dir hardening

The install-tree exclusion added in #93757 has a positive test (paths
inside the tree are skipped) but no negative boundary test. The guard
compares path components, so a prefix-named sibling like
/opt/hermes-data must still be chmod'd 0700 — but a rewrite to a
string-prefix match would silently drop that hardening with the suite
staying green.

Add test_install_tree_siblings_still_hardened covering a prefix-named
sibling and an ordinary sibling of the install root. Verified by
mutation: replacing the guard with str(parent).startswith(...) turns
the new test red.

Follow-up to #93757.

* fix: log a warning when parent-dir hardening is skipped for the install tree

The install-tree exclusion in secure_parent_dir() (#93757) returned
silently. A credential file being written inside the install tree is
exactly the misconfiguration signal that produced the production
lockouts the exclusion guards against, and it also means a previously
hardened path (e.g. a hermes home nested inside a git clone) silently
loses its 0700 parent tightening.

Emit a single warning naming the skipped directory and the install
root so the condition is diagnosable from logs.

Follow-up to #93757.

* docs: update secure_parent_dir docstring and caller comments for the install-tree exclusion

The docstring and all four caller comments still said the helper
refuses only / and top-level directories. Since #93757 it also refuses
the entire hermes-agent install tree. Bring the docstring and the
comments at the four credential-write call sites in line with the
actual behavior so future changes are not misled by a stale safety
description.

Follow-up to #93757.

* docs: add operator remediation for install dirs locked to 0700 by older images

The Dockerfile fix in #93757 only helps newly built images, and an
image upgrade (container recreate) resets the permission because
/opt/hermes lives in the image layer. The one stranded case is an old
image whose container was stopped and restarted after the lockout: it
keeps the 0700 install dir and runs code without the guard.

Document the one-line in-place recovery (chmod 0755 /opt/hermes as
root) in the Docker troubleshooting section.

Follow-up to #93757.

* fix(desktop): stop the HUD frosting the window while a turn runs

The frost is the whole window rectangle and the `[data-hud-glass]` scrim is
what makes it readable, but the two ran on different gates: the scrim is
focus-only, while the caller widened the frost to "recent or held" — i.e. for
the whole of a turn. Thinking with the composer unfocused therefore raised a
bare native material with no scrim over it, which on a light theme is a white
slab under the band's unconditionally white ink.

Put the frost back on the scrim's gate, and re-run it on the window's own
focus changes: clicking away to another app fires no focusout, so the scrim
would go while the frost stayed behind.

* fix(desktop): say why window enumeration failed instead of swallowing it

`read_window_below` answers "could not enumerate windows on this system" on
macOS and Windows whatever went wrong, and the three failure paths behind it
discarded their errors — so a report where the HUD could see nothing had no
way to distinguish the module failing to load, the helper failing to spawn,
and the OS answering with nothing. Three different fixes, one sentence.

Enumeration now returns the reason, the tool's error carries it, and the HUD's
game-overlay watch logs it once before it gives up (it retries twice and then
goes quiet forever, which is the other half of why the log said nothing).
Linux keeps its environment-derived advice, which is more actionable than the
raw exception.

* feat(desktop): add Settings toggle for vibe hearts

Floating affection hearts were always on with no off switch. Message
Reactions in Appearance looks related but only gates message-row
tapbacks. Add a separate Vibe Hearts preference (default on) next to it.

* test(desktop): pin vibe-hearts toggle across the pet-overlay forward path

The overlay window's playVibeHearts() only fires on a reaction forwarded by
burstVibeHearts, so the single gate covers it — these tests pin that so a
future direct caller shows up as a red test.

* fmt(js): `npm run fix` on merge (#94346)

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>

* fix(desktop): keep modal context menus inside dialogs

* fix(desktop): stop gating edit-menu Paste on the clipboard probe (#91553)

The dom context menu disabled Paste unless a renderer-side
readClipboard() probe reported text when the menu opened. The items
action never consumes that probe: editableCommand("paste") dispatches
webContents.paste() in main - the same Chromium path Ctrl+V takes, which
resolves the system clipboard itself. On Windows the Win32
clipboard.readText() bridge can return empty while that path succeeds,
so Paste stayed grayed out even though pasting would have worked; probe
errors were swallowed the same way (.catch(() => undefined)).

Fail open instead: drop the gate and the now-unused clipboardHasText
fact from the dom menu shape, so opening an editable menu no longer
makes the IPC round-trip at all. Pasting with an empty clipboard is a
harmless no-op, matching Chromiums own menu, which keeps Paste enabled
for editables. The terminal paste item keeps its gate - its action
inserts the readClipboard() text into the PTY directly, so there the
probe and the action share one mechanism and the gate stays honest.

Fixes #91553

* fix(desktop): keep UI scale across in-page route navigation

Desktop is a HashRouter over one file:// document, so every route is a
distinct URL to Chromium's per-URL zoom store. A route the user never
zoomed on has no record at all and resolves to the host default (100%) —
that is every fresh session and every never-visited settings tab.

In-page navigation fires neither did-finish-load nor any window event,
so nothing re-asserted the persisted level. The window dropped to 100%
while the Appearance control kept reading the chosen scale, because the
renderer only learns of zoom changes through 'hermes:zoom:changed',
which never fired. Touching the setting sent a fresh apply, which is
why it appeared to fix itself.

Re-assert the persisted level on main-frame did-navigate-in-page.
Verified on real Electron 40.10.2 / Chromium 144 (win32): a recordless
hash route reports 100% at the event, so the existing drift-guard sees
the drop and re-applies, and still no-ops when the route's record
already matches.

Fixes #48658
Fixes #38854
Fixes #79863

Co-authored-by: Brooklyn Nicholson <brooklyn.bb.nicholson@gmail.com>

* test(desktop): cover UI scale across recordless hash routes

Drives the reported path rather than the helper: set a non-default
scale, then navigate to routes Chromium holds no zoom record for, which
is what opening a new session looks like to the per-URL store. Keeps the
Cmd/Ctrl+N case alongside it.

Co-authored-by: Clark Vines <38430798+clarkvines@users.noreply.github.com>

* fix(signal): chunk long cron deliveries instead of truncating

* fix(signal): chunk long standalone sends and cover both delivery paths (salvage #57929 + #67279)

Follow-up to lkz-de's adapter chunking commit: long Signal messages no
longer truncate on ANY delivery path.

- tools/send_message_tool.py: register Signal's 8000-char limit in
  _MAX_LENGTHS (imported from the adapter module so the two paths can't
  drift) so hermes send / cron standalone / MCP sends split via the
  shared truncate_message() pass instead of signal-cli rejecting them.
  Standalone-path idea credited to @5L-hermes01 (#67279).
- tests: regression test proving standalone Signal sends chunk at the
  adapter limit with no truncation footer (fails on pre-fix main).
- docs: Long Messages section on the Signal page (en + zh-Hans).

Both fixes verified by sabotage A/B (tests fail with the respective
half reverted to origin/main) and a real-import E2E: 27k-char message
with emoji + cross-boundary bold + code blocks -> 4 chunks, all styles
in-range UTF-16, lossless reassembly.

* feat(terminal): pluggable terminal environment backends via plugin registry

Third-party sandbox vendors can now ship a terminal backend as a standalone
plugin instead of landing in core. Adds the five-piece pluggable-subsystem
pattern for terminal environments:

- agent/terminal_env_provider.py — TerminalEnvironmentProvider ABC with
  declarative classification flags (is_remote, is_container,
  skip_container_guards, cache_path_base, strip_env_keys,
  session_isolated_when_nonpersistent) so every historical
  frozenset-of-names classification site consults the registry instead
- agent/terminal_env_registry.py — thread-safe scoped registry; built-in
  backend names are reserved and unregistrable
- PluginContext.register_terminal_environment_provider() mirroring
  register_browser_provider
- _create_environment falls through to registered providers; unknown-backend
  errors list plugin names
- Classification sites wired: approval guard skip, container path/cwd
  handling (terminal/file/code-exec), prompt-builder env hints + probe,
  host env probe suppression, skills remote-env note, cache path
  translation, subprocess secret stripping (both spawn paths),
  per-session isolation for name-resumed sandboxes
- Surfaces: hermes setup picker + doctor + status rows, dashboard
  terminal-backend picker rows/probe/validation, terminal.backend schema
  options recomputed per request
- Docs: developer-guide/terminal-environment-plugin.md + sidebar + plugins
  capability table

* feat: browser snapshots drop LLM summarization — truncate-and-store like web_extract; auxiliary.web_extract slot removed

web_extract stopped using an auxiliary LLM long ago (deterministic
truncate-and-store), but browser snapshots still routed oversized
accessibility trees through the auxiliary web_extract model, keeping a
dead-looking aux slot alive across every config/picker surface.

- tools/browser_tool.py: remove _extract_relevant_content and
  _get_extraction_model; oversized snapshots always truncate at line
  boundaries, store the full tree to cache/web, and append a read_file
  pointer (element refs beyond the cut live in the file)
- tools/browser_camofox.py: same — no LLM path
- Remove auxiliary.web_extract slot: config_defaults (removal note, same
  pattern as session_search/PR #27590), cli.py defaults + env bridge,
  gateway/run.py bridged keys, hermes config display, hermes model picker,
  dashboard REST slots, desktop + web AUX_TASKS, i18n labels (en/zh/
  zh-hant/ja/ar)
- Docs: env-vars, configuration, fallback-providers, browser + zh-Hans
  mirrors (web-search zh-Hans was stale on the old LLM pipeline — synced
  to truncate-and-store truth)
- Tests updated: aux bridge uses approval slot, browser tests assert the
  LLM path is gone and stored files are secret-redacted

* fix(tui-gateway): make WS reconnect replay actually deliver events (follow-up to #94219)

The #94219 replay was a production no-op: the server returned full
JSON-RPC envelopes from session.events.since while the client's replay
loop dispatches only elements with a top-level 'type' — every replayed
event was silently skipped. Each side's tests validated its own
assumption, so both suites stayed green.

- server: events_since() now returns bare event objects (the frame's
  params), the exact shape the live dispatch path consumes; ring stores
  params directly; cross-language contract test added on both sides.
- client: live frames racing an in-flight replay are parked and flushed
  seq-gated afterward — no double dispatch of deltas, no gap-skip from
  a watermark advanced past the replay window.
- restart poisoning: seq counters are in-process, so a backend restart
  reset them while clients kept high watermarks (replay forever empty,
  truncated=false). New replay_epoch advertised in gateway.ready and
  echoed by session.events.since; the client clears watermarks on epoch
  change.
- methods_session no longer reaches into event_replay privates
  (is_truncated() accessor).

Live repro: pre-fix, 3 stamped frames -> 0 dispatchable by the client
gate; post-fix 3/3. Tests: 16 py (replay+ws), 8 vitest, tsc clean, ruff
clean.

* fmt(js): `npm run fix` on merge (#94410)

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>

* fix: system prompt no longer references tools/skills the session can't use; hermes-agent skill is always kept

Audit finding (Blank Slate): the system prompt advertised web_search,
skill_view, todo, and the hermes-agent skill even when the toolset had
none of them — the model chases phantoms it can't call.

- hermes-agent skill is now essential: cannot be disabled (config reads
  strip it, hermes tools writes drop it), cannot be deleted by
  skill_manage, is re-seeded past curator suppression, and is seeded
  even on .no-bundled-skills profiles (Blank Slate / --no-skills).
- Blank Slate core toolsets grow from file+terminal to
  file+terminal+vision+skills: read_file cannot read images and points
  at vision_analyze; the essential skill needs skill_view to load.
- HERMES_AGENT_HELP_GUIDANCE degrades to a docs-URL-only variant when
  skill tools are absent.
- Execution-discipline guidance drops its web_search lines when web
  tools are off (execution_guidance_text renderer).
- Skills-index preamble says 'basic tools like terminal' instead of
  naming web_search when web tools are off.
- Coding operating brief drops the todo-tracking sentence when the todo
  tool isn't loaded.

All gating keys off agent.valid_tool_names, fixed at session
construction — prompt stays byte-stable per session (cache-safe).

* test: opted-out profile seeding now exercises the essential-only sync subprocess

* fix(tests): e2e group-restart test no longer flakes on cold SessionDB init

The /goal post-turn hook constructs a real SessionDB on an executor
thread at the turn boundary. On a cold or loaded CI runner that
state.db init can exceed send_and_capture's 2s poll window, so the
send lands after the assertion and the test reports the bare
'Expected mock to have been called once. Called 0 times.' (#92130).
Mock _run_post_turn_hooks in the e2e runner — these tests exercise
gateway command dispatch, not goal hooks.

Also scrub TELEGRAM_GROUP_ALLOWED_CHATS / *_GROUP_ALLOWED_USERS / QQ
allowlist env vars in the hermetic conftest: a developer shell with
those set flips _get_unauthorized_dm_behavior to 'ignore' and fails
the pairing e2e test locally.

* fix(agent): gate memory provider system_prompt_block on toolset config (#81014)

The external memory provider's `system_prompt_block()` was injected
unconditionally into the system prompt, while the provider's tools were
gated by `memory_provi…
and7777 pushed a commit to and7777/hermes-agent that referenced this pull request Aug 27, 2026
…ounting, Nth-delivery promotion, docstring

Follow-ups on top of the cherry-picked NousResearch#93532 cap:
- Regression tests: suppressed (in-cooldown) matches must NOT consume the
  lifetime budget; the cap trips exactly at the Nth DELIVERED match and
  promotes to notify_on_complete with the watch_disabled summary queued
  right after the final match.
- Extract _emit_lifetime_watch_disabled() and emit the summary even when
  the global breaker drops the final match, so the user always learns why
  watching went quiet (parity with the strike-limit path).
- Mention the lifetime cap in the terminal tool docstring (the schema text
  was already updated by NousResearch#93532).

Refs NousResearch#93513
melon-xf added a commit to melon-xf/hermes-agent that referenced this pull request Sep 3, 2026
…ounting, Nth-delivery promotion, docstring

Follow-ups on top of the cherry-picked NousResearch#93532 cap:
- Regression tests: suppressed (in-cooldown) matches must NOT consume the
  lifetime budget; the cap trips exactly at the Nth DELIVERED match and
  promotes to notify_on_complete with the watch_disabled summary queued
  right after the final match.
- Extract _emit_lifetime_watch_disabled() and emit the summary even when
  the global breaker drops the final match, so the user always learns why
  watching went quiet (parity with the strike-limit path).
- Mention the lifetime cap in the terminal tool docstring (the schema text
  was already updated by NousResearch#93532).

Refs NousResearch#93513
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/tools Tool registry, model_tools, toolsets P2 Medium — degraded but workaround exists tool/terminal Terminal execution and process management type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

watch_patterns on background processes repeatedly triggers full-context turns, stalling the event loop and detaching desktop sessions

4 participants