Repository navigation
fix(terminal): cap watch_patterns notifications over a process's lifetime - #93532
chelsealong wants to merge 1 commit into
Conversation
…time Per-session rate limiting only counts consecutive strike windows, so a pattern that recurs at a cadence just above WATCH_MIN_INTERVAL_SECONDS (e.g. a service restarted repeatedly over a day) never trips the existing strike-limit disable — each match lands in its own clean cooldown window. Every one of those matches still forces a full-context agent turn, which stalls the event loop on large sessions (NousResearch#93513). Add WATCH_LIFETIME_MAX_HITS: once a session has delivered this many watch_match notifications over its whole life, disable watch_patterns and fall back to notify_on_complete, reusing the existing disable path.
trevorgordon981
left a comment
There was a problem hiding this comment.
Reviewed from a Hermes audit pass (tested the new LifetimeCap test on this branch: passes). The cap is correctly independent of the consecutive-strike path and the hit that reaches the cap is still delivered. One accounting issue inline, plus a persistence note.
Separate note on persistence: _watch_hits and _watch_disabled are not written by _write_checkpoint, so a gateway restart while a long-lived watch process is still alive rebuilds the recovered session with the counter reset (recover_from_checkpoint) and re-arms the watch. The cap therefore does not survive a restart — the same limitation the existing strike counter already has, so not a regression, but worth flagging if the cap is meant to be a hard floor against recurring-restart spam.
| # Lifetime cap: this match is delivered (it already earned it), | ||
| # but disable further ones regardless of how cleanly spaced | ||
| # they are — see WATCH_LIFETIME_MAX_HITS above. | ||
| lifetime_exhausted = session._watch_hits >= WATCH_LIFETIME_MAX_HITS |
There was a problem hiding this comment.
_watch_hits increments at L261 before _global_watch_admit (L297) can reject the delivery, so a hit suppressed by the global circuit-breaker still counts toward the cap. When the global breaker trips, this block can disable the session — and the watch_disabled message claims "8 delivered matches" — after fewer than 8 notifications ever reached the user. Consider counting only globally-admitted deliveries, or drop "delivered" from the message.
…ounting, Nth-delivery promotion, docstring Follow-ups on top of the cherry-picked #93532 cap: - Regression tests: suppressed (in-cooldown) matches must NOT consume the lifetime budget; the cap trips exactly at the Nth DELIVERED match and promotes to notify_on_complete with the watch_disabled summary queued right after the final match. - Extract _emit_lifetime_watch_disabled() and emit the summary even when the global breaker drops the final match, so the user always learns why watching went quiet (parity with the strike-limit path). - Mention the lifetime cap in the terminal tool docstring (the schema text was already updated by #93532). Refs #93513
|
Merged via #93670 with your commit cherry-picked (authorship preserved); we hardened delivered-only counting and documented the cap in the tool schema. Thanks @chelsealong! |
…ic OAuth refresh races) (#3) * fix(dashboard): follow scroll on implicit active-session resume (#93518) pty_ws already fell back to the per-channel active-session file when a /chat WS connects with no ?resume= param, replaying the whole session into the PTY, but the frontend only pinned xterm's viewport to the bottom when resumeParam came from the URL (#59591). The implicit path had no way to learn a replay was happening, so the viewport stayed at the top of the scrollback. pty_ws now sends a one-off JSON control frame naming the session id it resolved from the active-session file, before any PTY bytes; PTY output itself always arrives as binary frames, so this is unambiguous on the wire. ChatPage tracks an `effectiveResume` value seeded from resumeParam and updated when this control frame arrives, and the existing follow-scroll/sanitizer/hydration logic keys off it instead of the URL param alone. Fixes #93518. * fix(desktop): bound reconnect awaits so a stuck IPC round-trip can't latch the UI frozen After a liveness-probe-triggered reconnect on a remote gateway, attemptReconnect() awaits desktop.getConnection() and resolveGatewayWsUrl() with no timeout. If either stalls (e.g. main process wedged mid-revalidation even though the backend itself is reachable), the `reconnecting` guard never clears, so every later scheduleReconnect()/attemptReconnect() early-returns forever and the UI stays stuck in "reconnecting" until the app is restarted. Bound both awaits with a 20s timeout so a stall rejects instead of hanging; the existing catch/finally already clears the guard and resumes backoff on rejection. gateway.connect() keeps its own separate connect timeout. Fixes #93454 * fix(desktop): bound the revalidateConnection() await too (#93454) attemptReconnect() awaited desktop.revalidateConnection?.() unbounded, immediately before the two IPC calls the previous commit wrapped in withTimeout(). A wedged revalidation after a liveness-probe trip - the exact trigger #93454 and this file's own comment describe - hung that await forever, so the reconnecting guard never cleared and the prior fix never got reached. Wrap it in the same 20s withTimeout() (still swallowing the result via .catch, matching its existing best-effort semantics) and extend the regression test to hang revalidateConnection() specifically, proving getConnection() and the socket still proceed once the stall times out. * fix(desktop): bound the boot and gateway-switch resolveGatewayWsUrl awaits too (#93454) Follow-up to the reconnect-loop fix: the same unbounded ticket-mint await exists on the soft gateway-switch path and the initial boot() path. Bound both with the same withTimeout/RECONNECT_ATTEMPT_TIMEOUT_MS so a wedged IPC round-trip fails into the existing retry paths instead of hanging the switch or the 'Starting Hermes…' screen forever. * fix(desktop): group chats no longer crash when a member's connection is deleted botRosterMeta() calls botConnectionRoute() for every sourceScoped/remoteSource row to look up its metadata. That's a passive display lookup, but botConnectionRoute() throws whenever connectionId can't be resolved -- which is exactly what a stale group-chat roster row looks like once its connection is deleted (its persisted descriptor keeps remoteSource: true but loses connectionId). Since botRosterMeta() is called for every member on every group-chat render, opening a group that still references a deleted connection threw on render and crashed the pane's error boundary in a loop that survived app restarts (the poisoned row is in Local Storage). botConnectionRoute()'s fail-closed throw is correct and stays for its actual callers -- routing a real request to a bot (requestForBot, session creation, etc., covered by remote-routing-races.test.mjs). botRosterMeta() now catches that throw and treats the row as having no resolvable route, same as a bot with no meta at all, instead of letting it blow up rendering. Fixes #93492 * fix(desktop): split strict connection routing from passive roster lookup botConnectionRoute() stays the strict, throwing dispatch path for real routing (requestForBot, session creation). botRosterMeta() is passive display code and previously reached that throw through a bare catch, which would have swallowed any unrelated failure the same way. It now calls a new non-throwing resolveBotConnectionRoute() and branches on a typed resolved | owner_removed | not_scoped status instead. Adds witnesses for the split: the typed statuses themselves, that strict dispatch still fails closed on an orphaned row, and that an unrelated failure while resolving meta for a live route still propagates instead of being swallowed. * fix(desktop): sweep group-chat rosters when a connection is removed Root cause of #93492: deleting a cloud/remote connection disposed its gateways (store/gateway.ts) but never touched the persisted 'group-chats' storage, so every member descriptor referencing the deleted connection stayed behind as a poisoned row (remoteSource: true, connection gone) that render-path route lookups tripped over forever. Subscribe to the connection registry's 'removed' lifecycle push (window.hermesDesktop.connections.onChanged, feature-detected — older Electron mains don't emit it) and annotate every persisted group-chat member owned by the deleted connection. Rows are marked (sourceMissing/sourceReachable), never silently deleted: the member keeps its identity and panes render the existing degraded 'Gateway removed' botSourceStatus state. Writes ride updateGroupChat so the durable record keeps its full shape, and the listener unbinds on plugin dispose. * fix(desktop): annotate group-chat members already orphaned before hydrate Rows poisoned before the removed-connection sweep existed (their connection was deleted while an older Desktop ran, so no lifecycle push ever swept them) are what made #93492 survive app restarts. After the persisted 'group-chats' hydrate, run a pure annotate pass over the rooms: - a descriptor that lost its connectionId (route unresolvable — the exact shape that threw on render) is always marked; - a descriptor whose connectionId is absent from the live connection registry is marked only when the registry could actually be read — an unavailable registry must not read as 'everything is orphaned'. Marked rows keep their identity and degrade to the existing 'Gateway removed' state; nothing is deleted. * fix(desktop): guard render-reachable route lookups against orphaned rows Audit of the remaining unguarded botConnectionRoute() callers a pane render can reach (#93492 follow-up to the botRosterMeta split). Each now uses the non-throwing resolveBotConnectionRoute() and degrades on an owner_removed row instead of throwing into the pane's error boundary: - botWorkspaceOwnerKey / setBotsWorkspaceOwner: sidebar visibility listener, Bots home open, and roster context menus recompute these on passive UI edges; an orphaned selection now yields the name-keyed owner and the blocked workspace target. - durableGroupChatMembers: rebuilt on every group send over the whole seated roster; one orphaned member no longer aborts the room update, and a swept member's degraded mark now survives the rebuild. - useModelOptions: hook body runs during render; the query is disabled for an orphaned row and the picker paints its error/disabled state. - AdvancedProfileConfig: dialog falls back to the bot's own name scope. Strict dispatch callers (requestForBot, session creation, deleteBot, duplicateBot, ensureBotMetadata, routines) intentionally keep the fail-closed throw — remote-routing-races.test.mjs still asserts it. Adds orphaned-connection-members.test.mjs covering the removed-connection sweep, the hydrate annotate (with/without a readable registry), the degraded 'Gateway removed' rendering of swept rows, and every guarded caller. * fix(desktop): isolate a failed lazy syntax-diff import from the workspace pane React.lazy(() => import('./syntax-diff')) only has its pending state covered by Suspense. When the dynamic import rejects (e.g. a packaged app whose renderer window resolves to the app.asar copy of dist/ while the chunk exists only in app.asar.unpacked, #93479), the rejection throws past Suspense to the nearest error boundary, which is the whole workspace ContribBoundary. One missing highlighter chunk then blanks the entire chat transcript instead of just the diff falling back to the plain colored DiffBody, the way markdown-text.tsx already isolates this failure class for markdown. Wraps LazySyntaxDiff in a local ErrorBoundary that renders DiffBody on catch, so a failed highlight chunk degrades in place. * fix(desktop): prefer the unpacked web dist over the asar-internal renderer index when packaged The renderer index resolver tried APP_ROOT/dist/index.html — inside app.asar when packaged — before the app.asar.unpacked copy that asarUnpack (dist/**) ships and that resolveWebDist() already prefers for the embedded dashboard. Loading the asar-internal index is how lazily imported chunks (syntax-diff-*, shiki-*, mermaid-embed-*) end up fetched from a path that cannot serve them, killing the workspace pane (#93479). Reorder the candidate ladder to prefer the unpacked web dist when packaged, following the unpackedPathFor/resolveWebDist precedent. All window loaders (main, overlay, quick) share resolveRendererIndex, so one reorder covers every surface. Dev behavior is unchanged: outside an asar both candidates collapse to APP_ROOT/dist and keep the original order. * fix(desktop): teach the torn-bundle guard to see missing lazy chunks missingRendererAssets only checked the module refs index.html itself names (<script type=module> + modulepreload), so a torn install whose boot-critical files were intact but whose lazy chunks were gone passed the generation check and died minutes later on the first React.lazy() route with 'Failed to fetch dynamically imported module' (#93479: syntax-diff-*, shiki-*, mermaid-embed-*). Walk the generation's module graph: for every present JS chunk, parse its inline __vite__mapDeps filename table (the lazy-import manifest Vite bakes into each chunk) and check those files too, transitively and cycle-safe. resolveRendererIndex now skips a lazy-chunk-torn candidate in favor of the intact copy instead of shipping a delayed crash. Tests cover the mapDeps parser (definition table vs index-only call sites, CDN refs), the exact #93479 tear shape, transitive/cyclic walks, and the torn-vs-intact preference end to end. * fix(terminal): cap watch_patterns notifications over a process's lifetime Per-session rate limiting only counts consecutive strike windows, so a pattern that recurs at a cadence just above WATCH_MIN_INTERVAL_SECONDS (e.g. a service restarted repeatedly over a day) never trips the existing strike-limit disable — each match lands in its own clean cooldown window. Every one of those matches still forces a full-context agent turn, which stalls the event loop on large sessions (#93513). Add WATCH_LIFETIME_MAX_HITS: once a session has delivered this many watch_match notifications over its whole life, disable watch_patterns and fall back to notify_on_complete, reusing the existing disable path. * test(terminal): harden watch_patterns lifetime cap — delivered-only counting, Nth-delivery promotion, docstring Follow-ups on top of the cherry-picked #93532 cap: - Regression tests: suppressed (in-cooldown) matches must NOT consume the lifetime budget; the cap trips exactly at the Nth DELIVERED match and promotes to notify_on_complete with the watch_disabled summary queued right after the final match. - Extract _emit_lifetime_watch_disabled() and emit the summary even when the global breaker drops the final match, so the user always learns why watching went quiet (parity with the strike-limit path). - Mention the lifetime cap in the terminal tool docstring (the schema text was already updated by #93532). Refs #93513 * fix(auth): malformed OpenRouter env key no longer shadows valid credential-pool key A malformed OPENROUTER_API_KEY in ~/.hermes/.env (truncated paste, wrong provider's key) passed has_usable_secret's length/placeholder check and was returned by _resolve_api_key_provider_secret before the credential-pool fallback was ever reached, producing opaque '401 Missing Authentication header' errors even when a valid pool entry existed (#93593). - Add KNOWN_PROVIDER_KEY_PREFIXES (openrouter: sk-or-) and skip env values that mismatch a declared prefix, logging a WARNING naming the env var and expected prefix, then continuing to the next env var / pool fallback. - Iterate credential-pool entries (peek first, then entries()) instead of only peek(), so one malformed pool entry doesn't block a valid one. - Providers without a declared prefix are fail-open: unknown key formats are never rejected. Valid env keys still win over the pool (precedence unchanged). Fixes #93593 * fix(pricing): support Gemini context-tiered rates in pricing snapshot (#93469) The pricing snapshot could only express flat per-million rates, so gemini-3.1-pro sessions with prompts over 200k tokens under-counted input 2x ($2 vs $4/M) and output 1.5x ($12 vs $18/M). - Add optional tier fields to PricingEntry: tier_threshold_tokens, input/output/cache_read_cost_per_million_above (None = flat, falls back to base rate per-field). - estimate_usage_cost selects the above-threshold rates for the WHOLE request once usage.prompt_tokens (input + cache read + cache write) exceeds the threshold, matching Google's billing semantics. - Populate gemini-3.1-pro (4.00/18.00/0.40 above 200k; alias gemini-3.1-pro-preview inherits) and gemini-2.5-pro (2.50/15.00 above 200k). - Flat entries are untouched: no threshold means no behavior change. Reported and tier-field shape designed by @tornike14 (#93469). Tests: below/at threshold unchanged, above-threshold tiered whole-request pricing, cache-read tier rate and base-rate fallback, preview alias, flat entries unaffected. * fix(desktop): stop bot-relay drain loop from redialing a WebSocket per connection per tick The bot relay's drain loop RPCs every registered connection through requestGatewayForAgent's per-request lease. With no other consumer holding the route, the refcount hit 0 after every tick and the pooled secondary was disposed — a fresh WebSocket dial + teardown per connection every 4s, flooding the gateway logs with connect/disconnect pairs (#93594). Two changes, both directions from the issue: - Retained relay-route secondaries: retainGatewayForRelay pins a route's pooled socket with a counted retention (never clobbering the foreground 'retained' flag) for the relay's active lifetime, reusing the existing scheduleReconnect/full-jitter machinery on drops. The plugin pins each registered connection once via the new feature- detected host.retainProfileSocket door, reconciles pins with the current connection set on every drain, and releases everything in stopBotRelay/dispose. Local routes (null/'local') are exempt so the idle reaper can still reclaim spawned local backends. The live-work pruner also respects the pin. - RELAY_DRAIN_INTERVAL_MS 4s -> 30s: the push path (#93091, bot_relay.outbox.pending) carries envelope latency, so the poll is purely a backstop — 30s matches LIVE_SESSION_STATUS_BACKSTOP_INTERVAL_MS. Tests: relay-push-drain updated to the new backstop semantics; new gateway-relay-retention.test.ts proves one socket construction across 5 drain ticks (vs 3 constructions for 3 unretained ticks) and that release/prune/local-exemption behave; new relay-socket-retention plugin test pins the pin-once / release-on-departure / stop-releases contracts. * fix(desktop): exclude cron sessions from the titlebar unread badge Cron runs finish unwatched by design, so counting them in $unreadSessionCount turned the titlebar badge into a permanently-lit cron run counter (#93552). The badge now counts regular + messaging sessions only; cron unread state stays visible on the sidebar cron section rows, and 'Mark all as read' (markAllSessionsRead + ackAllSessionsRead, which iterates cron rows) still clears them. Fixes #93552 * fix(desktop): stack subsequent preview tiles as tabs instead of new right splits Every opened file registered its preview pane with dock dir 'right', so each open split a new zone off the right edge — three file opens made three ever-narrower columns (#93610). The first preview still opens its own zone docked beside main; every subsequent preview now anchors to an existing preview-tile pane with dir 'center', so it stacks as a tab in the same preview zone. Covers files, artifacts, and the Browser tab alike (all flow through openPreview/$previewTabs); session tiles are untouched. Fixes #93610 * fix(hindsight): send configured event timestamps Use Hermes timezone-aware timestamps for retained events and turn messages. Pass the public timestamp field supported by hindsight-client 0.6.1 and cover the final serialized request field. * fix(hindsight): harden event timestamps * fix(hindsight): let hindsight_retain convey event time via occurred_at Adds an optional occurred_at (ISO-8601 date/datetime) parameter to the hindsight_retain tool schema, threaded into the retain item's timestamp field. When absent, the item timestamp defaults to the configured event clock (base from PR #82928 by @ragingbulld, authorship preserved) so the Hindsight server can resolve relative time phrases; previously no item timestamp was ever sent and temporal memories landed with null occurred_start/occurred_end. Fixes #93568. Salvages #82928. * fix(desktop): give keyboard focus a visible affordance where the global no-ring reset hides it The unlayered *:focus-visible reset in styles.css intentionally zeroes --tw-ring-shadow ('No focus rings, anywhere'), so any control that relied solely on focus-visible:ring-* had no visible keyboard focus state at all. Mirror each control's hover treatment as a focus-visible background/text affordance instead, keeping the global reset intact: - ui/sidebar.tsx: group label, group action, menu button, menu action, menu sub-button get focus-visible:bg-sidebar-accent + accent foreground - ui/tabs.tsx: TabsTrigger gets focus-visible:bg-background + text-foreground - ui/text-tab.tsx: focus-visible:text-foreground (matches its hover) - chat/composer/micro-actions.tsx: pill gets focus-visible chrome-action-hover - right-sidebar/index.tsx HEADER_ACTION_CLASS: focus-visible sidebar-accent - right-sidebar/terminal/rail.tsx RAIL_ACTION: focus-visible chrome-action-hover - chat/sidebar/cron-jobs-section.tsx (row body + run rows): focus-visible chrome-action-hover - chat/sidebar/session-row.tsx <time>: focus-visible:text-foreground Sweep verified: remaining focus-visible:ring-* usages under apps/desktop/src already pair with a border/bg/text companion (button/checkbox/switch/input, starmap share-controls) or are covered by PR #93460's row-hover work (cron/index.tsx run rows). Fixes #93462. Reported by @fred0m. * fix(desktop): open a branched session in the main workspace, not just a tile forkBranch ended by opening the branch as a session-tile and leaving the primary selection on the parent (#69750). In the default layout there is no visible tile pane, so branching only added a sidebar row with no feedback in the main area — and openSessionTile no-ops when the target is already the selected session, the common case of branching the chat you're viewing. Load the branch as the primary session via resumeSession instead, which reuses the runtime already warm-cached by forkBranch's ensureSessionState/updateSessionState calls, so it doesn't cost an extra resume RPC. Fixes #93444 * fix(desktop): scope branch-opens-primary to the currently selected session forkBranch was unconditionally routing every branched session into the main pane via resumeSession, including sidebar/background branches of a session the user isn't currently viewing. That reintroduces the #69750 focus-stealing bug for that path: branching a different session from the sidebar yanked the active view away from whatever was open. Only take over the main pane when the branch's parent is the session already selected; otherwise keep opening it as its own tile. * test(desktop): pin auto-speak silence across an Edge TTS fallback id rewrite #93515 reports auto-speak reading each reply twice when the Edge TTS streaming attempt falls back to the POST endpoint and the reply's renderer id gets rewritten to its durable id mid-flight. That was true before 63565fa26b, but resolveSpokenReply()'s ordinal-anchored dedupe (landed 2026-08-19, five days before this issue was filed) already follows the rewrite. No source change — this pins the behavior with a regression test at the hook/store integration level, one layer above the existing spoken-reply.ts unit tests. * fix(desktop): hold a per-turn socket lease so group-chat member turns survive the runtime-session reaper (#93602) A group member turn is a session-scoped RPC sequence (resume → attach → prompt.submit → poll) issued with the runtime id its first RPC minted, but requestForBot routes every RPC through its own request-scoped socket lease (retained:false secondaries in store/gateway). Between two RPCs the refcount hits 0, the leased socket closes, the gateway detaches the runtime session on WS disconnect, the orphan reaper frees it after grace, and the next RPC — prompt.submit, unwrapped — dies 4001 'not in memory'. The member turn aborts and the sub-profile bot goes silent in the room. - store/gateway: retainGatewayForAgent(connectionId, profile) — refcounted hold on the pooled socket with an idempotent release, mirroring the existing request-lease machinery. - sdk: host.retainProfile(route) exposes the retain to plugins (feature-detected by consumers; older hosts keep working). - hermes-bots plugin: runGroupChatMemberTurn acquires the lease before ensureGroupChatSession's first RPC and releases in finally, so the socket that minted the runtime id stays open across attach+submit+poll; and prompt.submit gets a one-shot catch-and-retry that re-resumes via the STORED session id on 4001-class failures (belt-and-braces for routes the lease can't cover). 4007 'never existed' keeps flowing to session.create. Tests: simulated 4001 on first submit recovers via re-resume and delivers; lease held across attach+submit (mock refcount never hits 0 mid-turn); lease released after success AND failure; no-retainProfile host feature detection; store-level retain/release + idempotent double-release + the unretained disposal race. * test: drop unused param in group-turn lease mock (lint) * fix(desktop-update): use the system default browser for the update shim The Windows update hand-off shim was hardcoded to Microsoft Edge (Find-EdgeExe), so machines whose default browser is Chrome still got an Edge --app progress window, and every run leaked a throwaway browser profile (hermes-update-ui-<pid>) under %TEMP% that was never removed. - Get-DefaultBrowserExe replaces Find-EdgeExe: resolves the OS default browser from the UserChoice ProgId (https first, http fallback). ChromeHTML -> Chrome, MSEdgeHTM -> Edge; any other ProgId returns $null and degrades to the existing WinForms card. - The dedicated --user-data-dir profile is now removed when the shim closes, and stale hermes-update-ui-* leftovers from interrupted runs are swept from %TEMP% in the same pass. - --app + --user-data-dir is Chromium-only, so the whitelist is intentionally limited to chrome/msedge; Edge keeps --disable-features=msImplicitSignin to suppress the implicit MSA sign-in that leaks into shim windows (#88410). * feat(gateway): add gateway.ping heartbeat wire contract Additive WebSocket wire contract for a client-driven heartbeat. The gateway.ready payload now advertises "heartbeat": True so clients can discover the capability, and the WS read loop answers a gateway.ping request with a {"ok": True} pong short-circuited BEFORE method dispatch (no method is invoked). The WSTransport gains closed / last_inbound_at properties and a mark_inbound() hook, updated on every inbound frame, for later liveness checks. Backward-compatible in both directions: old clients never send gateway.ping, and old servers simply never advertise the heartbeat flag. This is the first slice of a WebSocket-recovery series; the follow-on slices consume this contract (server-side transport rebind, TUI/desktop clients). Receipts: bash scripts/run_tests.sh tests/test_tui_gateway_ws.py -q => 1 file, 7 tests passed, 0 failed (100%) in 0.8s; exit 0 * feat(shared): heartbeat and socket-generation invalidation in JsonRpcGatewayClient The shared-client half of the gateway.ping heartbeat contract (#89958); tracks lastInboundAt, sends pings, invalidates a silently-dead socket. Part of #83166. * fix(tui): heartbeat and bounded reconnect for silent WebSocket drops the client half of the gateway.ping heartbeat contract (#89958); detects a silently-dropped socket via missed ping-acks and reconnects with bounded backoff; part of the #83166 recovery series. * fix(cron): surface gateway liveness in cronjob tool results (#87033) The builtin cron ticker only runs inside the gateway process. The CLI surfaces this ('hermes cron list' / 'hermes cron status' both warn when no gateway is running), but the model-facing cronjob tool returned a clean success on create even with no gateway running - so the agent confidently told the user a recurring task was scheduled while the job could never fire. Mirror the CLI's liveness heuristic in the tool's create path and attach a tri-state gateway_running field to the result: - true -> gateway running (or a non-builtin scheduler provider owns firing, e.g. Chronos, which is exempt by design) - false -> explicit warning telling the model the job is saved but will NOT fire until the gateway starts, so it can relay that to the user instead of reporting unqualified success - null -> probe failed; claim neither way Fixes #87033 * fix(cron): share liveness helper with CLI and extend it to cronjob list Follow-ups for the salvaged #93098: - Move the tri-state liveness heuristic into hermes_cli.cron (_builtin_gateway_liveness) so the CLI warning and the cronjob tool share one implementation instead of two drifting copies. - Surface gateway_running/warning on the list action too — an agent inspecting jobs in a gateway-less environment has the same silent- inert-job failure mode (#87033) as create. Empty lists stay quiet. * refactor(cron): parameterize liveness warning plurality, drop fragile string replace Simplify-pass follow-ups on the #87033 fix: - _gateway_liveness_notice(plural=) authors both wording variants at one site; removes the exact-substring .replace() that would silently no-op if the create-path text is ever edited. - Collapse the operator-precedence-trap conditional in list to a plain 'if jobs' — an empty list has nothing inert and now skips the probe. - Fix docstring/code mismatch (builder returns gateway_running: True on the happy path) and drop the dead try/except in _warn_if_gateway_not_running (the helper never raises). * refactor(telegram): migrate _await_with_thread_deadline onto agent.deadline.run_bounded_async (#85125 2f) The adapter's private thread-deadline helper was the ancestor of the unified deadline layer's run_bounded_async (#85147 was extracted from it, plus the caller-cancellation leak fix the original still lacked). Consolidate: the helper body becomes a thin wrapper mapping BoundedResult.timed_out back to the asyncio.TimeoutError its 9 call sites (the PTB retry ladder) expect. ~90 duplicated lines die, along with the adapter-local copies of the abandon-cleanup runner and the blocked-loop faulthandler diagnostics (both live in agent/deadline.py). Everything the call sites rely on is preserved by the unified layer: - thread-timer deadline that survives a blocked event loop (#63309) - abandonment of cancellation-shielded tasks (PTB/httpcore anyio init) - detached best-effort on_abandon cleanup (no httpx pool leak per retry) - off-loop stack dump when the loop never processes the expiry Plus one behavior IMPROVEMENT inherited from the shared copy: a caller cancelling the wrapper no longer leaks the inner task unobserved (the telegram original had that leak; the extraction fixed it). test_telegram_init_deadline.py: the #63309 diagnostics probe now pins the shared layer's dump hook (label "telegram-init") — same contract, new seam. Wedge + cleanup-crash tests pass unchanged. * fix(mcp): resolve tool-call timeouts via the unified deadline layer (#85125 2g) Both readers of the per-server MCP tool timeout (the connection's run() and the cache-path registration) read config.get("timeout", 300) as their own private resolution. Route them through _resolve_tool_timeout: per-server mcp_servers.<name>.timeout still ALWAYS wins (most specific), then timeouts.mcp.tool_call from the unified timeouts: section, then the unchanged 300s default. Values pass through resolve_timeout's platform clamp; resolution failure falls back to the historical default. Default-behavior invariance pinned by contract tests (nothing configured -> exactly 300, per-server beats section, section beats default, invalid/failed resolution falls back). * fix(desktop): move the Linux HUD with a native compositor drag Wayland clients cannot place themselves, so the JS setBounds drag is a no-op there. Make the composer bar a -webkit-app-region drag handle on Linux (input carved out with no-drag) and let the compositor move it. Co-authored-by: Tony Simons <214744153+asimons81@users.noreply.github.com> * fix(desktop): debounce and re-verify zoom on Linux Wayland Focus events fire for intra-app shifts on Wayland and Cosmic tiled resizes can drop a just-applied zoom. Debounce focus with resize/move and re-check a few times after the window settles. Co-authored-by: joe0508 <75520452+joe050860@users.noreply.github.com> Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com> * fix(desktop): keep the Linux HUD clickable and recoverable X11 cannot restore a window that has ignored the mouse, so stay solid there. Native Wayland keeps click-through via the cursor poll. Add desktop.ozone_platform_hint so COSMIC users can opt into XWayland for always-on-top, and a layout reset that restores the default size. Co-authored-by: Codex Metatron <47930664+BlakeB254@users.noreply.github.com> Co-authored-by: DeseretSaint <202557515+DeseretSaint@users.noreply.github.com> Co-authored-by: Shawn Wang <32839114+enwaiax@users.noreply.github.com> * fix(desktop): add a HUD layout reset control A persisted tall/narrow size has no way out on Linux. Put a reset next to Exit HUD so the default size (and position, where the compositor allows it) is one click away. Co-authored-by: Shawn Wang <32839114+enwaiax@users.noreply.github.com> * docs(desktop): document Linux and Wayland HUD behavior Spell out native-compositor drag, the ozone_platform_hint escape hatch for COSMIC always-on-top, and the snap-to-pointer no-op on Wayland. * fix(desktop): polish HUD movement and resizing on X11 * fix(serve): emit BACKEND_PORT_IN_USE sentinel + exit 75 on port bind conflict (#93608) A held port made 'hermes serve' print only uvicorn's bare 'ERROR: [Errno 98/10048] error while attempting to bind on address' and exit 1 — indistinguishable from a broken backend for the desktop spawn and wrapping scripts. - Preflight bind probe (matching uvicorn's SO_REUSEADDR bind flags) before uvicorn.Server; on conflict print machine-readable 'BACKEND_PORT_IN_USE port=<port>' + a human hint naming likely holders, exit 75 (EX_TEMPFAIL — existing repo convention, see gateway/restart.py, kanban_db.py). - Probe-to-bind race covered: SystemExit(1) from uvicorn's own bind failure is re-checked and translated on both POSIX and Windows runner paths. - --port 0 (ephemeral) short-circuits the probe: unchanged behavior. - HERMES_BACKEND_READY contract untouched. - Tests: real held-socket repro (sentinel + exit 75, sabotage-proven to fail as bare exit 1 without the fix), free-port boot regression, ephemeral-port regression, probe/classification units. - Docs: port-conflict paragraph under 'hermes serve' in reference/cli-commands.md. * fix(serve): Windows conflict probe uses SO_EXCLUSIVEADDRUSE (SO_REUSEADDR binds over live listeners on WinSock) * fix(desktop): never kill a healthy backend on a claim probe failure; surface real stderr (#93608) A start-marker probe failure (Get-Process timing out on a PowerShell 5.1 cold start, #87169) in claimBackendChild used to stop the freshly spawned backend and rethrow — killing a healthy backend, triggering the renderer's repair respawn, and looping. And because stderr piping only attached after the claim, every before-ready failure surfaced as a bare exit code. - extract probe + claim policy into electron/backend-claim.ts: processStartMarker/execText (moved verbatim from main.ts), probeStartMarker, and a pure claimDecision(childAlive, probe) a Windows CI lane can drive with real PowerShell - probe failure + LIVE child now degrades to PID-only identity (pid-only:<pid> marker, WARNING logged), matching the existing createParentStartMarkerResolver degrade pattern; processIdentityMatches verifies degraded identities by PID liveness (command check still layers on top in backendIdentityMatches) - probe failure + DEAD child keeps the fail-closed throw, now carrying the child's buffered stderr/stdout tail - ring-buffered ~8KB output tail attached at spawn time in BOTH spawn paths (pool + primary); tail appended to claim errors, before-ready exit messages, and backend-ready's exited-before-port-announcement errors so the real exit reason reaches desktop.log and the boot UI - tests: claimDecision matrix (degrade test fails against the old stop+throw behavior), real processStartMarker probe, ring-buffer caps, and output-tail suffixes on backend-ready exit errors * fix(desktop): float and pin the HUD on Hyprland Omarchy tiles the HUD like any other toplevel, so always-on-top and xdg_toplevel.move never apply. Ask Hyprland to float+pin after map, trying classic dispatch then Lua for 0.55+ configs. * fmt(js): `npm run fix` on merge (#94046) Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> * fix(codex): identify Hermes requests * feat(desktop): make Cmd/Ctrl+L focus the composer from anywhere The chord previously only acted when a terminal or preview selection existed. A new bubble-phase window fallback now moves focus to the composer on an unclaimed press, like the address-bar chord in a browser. Existing owners keep priority: selection handlers claim the press on the capture phase, a user-rebound action marks the event handled, and a focused terminal with no selection keeps Ctrl+L as clear-screen via composerFocusBlockedBySurface(). The chord matcher moves from the terminal feature to src/lib/keybinds/chords.ts as isComposerChord: it now has three consumers and the old name (isAddSelectionShortcut) was wrong at the composer call site. The fixed panel row view.terminalSelection becomes view.selectionToComposer because it covers preview selections too. * refactor(desktop): derive HUD OS behavior from one windowing profile Move, ignore-mouse, placement, resize edges, snap, cursor feed, and overlay promote all read the same Ozone-normalized capabilities instead of re-deriving linux/Wayland/X11 at each call site. * fix(desktop): sort HUD windowing imports for eslint * fmt(js): `npm run fix` on merge (#94140) Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> * feat(desktop): let the in-app browser hold more than one tab A URL tab used to be a singleton — every link navigated the one Browser, so there was no way to keep a page open beside another and the strip's "+" never appeared next to it. A Browser tab is now a vessel with its own id: links still land in the browser you are looking at (an agent opening five pages must not leave five tabs behind), while the "+" mints another one on request. Tabs name themselves after the page they are showing, since three tabs reading "Browser" name nothing. The "+" itself is now a pane capability rather than a session-only button, so any pane kind that can make more of itself contributes one. * fix(desktop): hold the typed address in the browser bar until the page moves Committing an address dropped the field back to the url of the page you were leaving, so typing baby.com over google.com flashed google.com back before baby.com arrived — and nothing said a load was underway. The address you asked for now stays in the field until the page actually lands somewhere (a redirect supersedes it, as it should), and progress spins inside the field beside it. The pane owns that loading state from the moment it accepts the address, because the reach probe it runs first delays did-start-loading. * refactor(desktop): mint browser tab ids at random rather than by slot Browser tabs took the lowest free slot, so an id was reused once its tab closed. That is only safe while every store keyed by the id is wiped on close — true today, but a discipline rather than a guarantee, and stale state would resurface under an unrelated tab the day it lapses. Mint like a terminal does instead: no id is ever handed out twice. * fix(desktop): keep Home new sessions detached from the last project Home's "+" passes path/cwd null on purpose, but null was falsy and fell through into resolveNewSessionCwd(), so "New session in Home" (especially the openTab path while main chat is occupied) still created under the previous project folder and showed its branch. * test(desktop): lock Home new-session detach against stale project cwd Cover the null-path draft path, Home-scope createBackendSessionForSend, and openNewSessionTile({ cwd: null }) so the last project folder cannot leak back into a Home chat. * fmt(js): `npm run fix` on merge (#94176) Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> * fix(mcp): recover poisoned connections + fail fast on dead stdio transports (#85125 3b) Fixes the four poisoned-connection classes (#81051, #77765, #84132, #81995) with the SuspectableBackend cheap-mark/lazy-verify contract: - mark_suspect/ensure_healthy protocol (agent/deadline.py): noticing a poisoned state never does I/O; the NEXT caller pays once for a health probe that clears the suspicion or forces a reconnect. A single teardown-vs-keepalive race or auth-lock corruption can no longer park a connection permanently — park stays reserved for genuinely exhausted reconnect budgets. - keepalive failure marks the connection suspect before requesting reconnect; the next tool call probes and recycles if unhealthy. - auth-classified permanent failures on a previously-proven session get a suspect+reconnect path instead of an immediate park. - fast-fail (#81995): stdio child pids are tracked at spawn and an in-flight RPC races a child-watcher task, so a dead subprocess fails the call immediately with a retryable timeout instead of riding out the full 300s. Deliberate teardown/reconnect also fails in-flight calls now instead of leaving them attached to a dying transport. Dispatch-boundary hardening for test doubles: stubbed sessions (MagicMock/non-awaitable call_tool, absent child-watcher) fall back to the exact pre-change inline-await semantics, so only real transports gain the race guard. Salvage credit: in-flight approach from #73377 (@luijoc, wedged transport recovery) and #48069 (@arminanton, keepalive/in-flight interaction); both PRs' bases predate main's current park/reconnect architecture, so this is a fresh implementation of their contracts. Tests: tests/tools/ -k mcp = 639 passed (was 22 new failures during development; final tree zero). * refactor(deadline): consolidate site-local tree-kills onto agent.deadline.kill_process_tree (#85125 4d) Per-site decisions: 1. hermes_cli/_subprocess_compat.py kill_process_tree(proc) -> None: MIGRATED. Body now delegates to agent.deadline.kill_process_tree(proc.pid) via a function-local import; keeps the swallow-everything fail-open contract and the (proc) -> None signature (agent/shell_hooks.py imports it by name; _kill_git_process_tree alias preserved). The old body is kept verbatim as _legacy_kill_process_tree and used as fallback when the delegation import/call fails. A final proc.kill() is retained on the happy path so Popen bookkeeping sees the exit (matches old behavior). 2. tools/browser_tool.py _kill_process_tree(proc): MIGRATED, same pattern (delegate + _legacy_kill_process_tree fallback). Behavior delta: the old body sent SIGTERM then SIGKILL with zero grace between them; the shared primitive sends SIGKILL only. With no grace period the observable effect is identical, and the psutil descendant sweep now also reaches agent-browser's setsid'd daemon grandchild, which killpg alone missed. tests/tools/test_browser_npx_warmup.py's TestKillProcessTree repointed at the legacy fallback (its assertions describe the fallback's internals). 3. tools/code_execution_tool.py _kill_process_group(proc, escalate): MIGRATED. It was a plain parent+descendants terminate (then wait 5s + kill when escalate=True) — expressed as two delegated calls: kill_process_tree(pid, sig=SIGTERM), then on escalate-timeout kill_process_tree(pid, sig=SIGKILL). Delegation failure degrades to proc.kill(), mirroring the old psutil-failure fallback. Delta: the old body terminated children before the parent; the shared primitive signals the group atomically (child is a session leader via start_new_session=True) plus an identity-aware descendant sweep — strictly wider coverage, same signals. 4. gateway/status.py: KEPT BOTH SITES. - terminate_pid (~l305) taskkill wrapper: NOT migrated. Its contract is incompatible with the shared primitive — it must RAISE OSError with taskkill's stderr on non-zero exit (callers branch on that), falls back to os.kill on FileNotFoundError, and its POSIX branch is deliberately a single-PID SIGTERM/SIGKILL, not a tree kill. Wrapping the bool-returning fail-soft primitive would invert the error contract. - reap_gateway_children (~l2029): NOT migrated. It operates on a pre-snapshotted child list from a parent that is already dead (psutil.Process(pid) on the parent would fail), and every signal is wrapped in identity/ownership checks the primitive lacks: is_running() identity, zombie skip, and the skip-if-ppid-still-equals-parent guard, plus SIGTERM -> wait_procs -> SIGKILL staging and a reaped-count return. The coupling is the feature; migrating would delete the safety logic. 5. scripts/run_tests_parallel.py _kill_process_tree (~l253): NOT migrated. Dev tooling that intentionally kills by CAPTURED pgid because the direct child is usually already reaped (psutil/pid-based primitive cannot find it), and it avoids the psutil import on the test-runner hot path. Its docstring already documents why psutil is the wrong tool there. New tests: tests/agent/test_treekill_consolidation.py — delegation + raise-swallowing tests per migrated wrapper, consumer-identity checks, and a live end-to-end probe (setsid grandchild dies through the compat wrapper, zero survivors). * fix(terminal): sweep setsid descendants after local timeout group-kill (#85125 4b) LocalEnvironment._kill_process kills the process GROUP (SIGTERM -> 1s wait -> SIGKILL -> 2s wait), but a descendant that called setsid escapes the group and survives — the #71148 orphan class, terminal flavor (issue #84967's local sibling). Fix: snapshot the descendant set via psutil BEFORE the first SIGTERM (children reparent to init once the wrapper dies, so a later parent walk finds nothing — same snapshot-before-signal design as agent/deadline.py kill_process_tree), then after the existing group escalation completes, SIGKILL any snapshotted survivor whose pgid is no longer the (now-dead) group. The TERM->KILL grace window for in-group members is preserved (interrupts use this path too), the Windows branch is untouched, and the snapshot is fully guarded — a broken psutil never breaks the kill path (unit-tested). Tests: live_system_guard_bypass acceptance test spawning a setsid grandchild and forcing the timeout path (RED on unmodified file, GREEN after), plus a psutil-failure unit test. Docker design note (#84967 open question 1, condensed; full note at /tmp/4b-docker-design-note.md): the docker backend inherits base.py:1378 _kill_process, which only proc.kill()s the HOST-side `docker exec` client — the in-container tree (child of containerd- shim, not the client) survives every timeout entirely. Option A, `docker exec <cid> kill -- -<pgid>` with TERM->KILL escalation using a PGID captured at command start, is surgical and preserves container state but needs a live container + shell and still misses in-container setsid escapees. Option B, container restart, is absolute (PID- namespace teardown kills everything) but destroys all in-container state mid-session and punishes every other consumer of the shared persistent container. Recommendation: Option A as a best-effort _kill_process override (degrade to today's behavior on failure); reserve restart for the existing container-gone recovery path. * fix(desktop): stop transcript jumps when a turn settles Clarify remounted as a tool row once session.info flipped running=false, thinking previews collapsed their body, and the duration line grew the footer. * fix(desktop): demote unanswered clarify cards on Stop Latch the pending card on submit, not on seeing a request, so Stop still collapses an unanswered question instead of leaving a disabled panel. * fix(desktop): re-arm pending clarify cards in place A hydrated Ask/clarify row stays complete after session or bot switch, so the live card never mounts. Re-arm the existing transcript row and keep the provider tool id instead of appending a duplicate at the tail. Co-authored-by: frendo <frendo.wu@gmail.com> * fix(desktop): restore pending_clarify snapshots on activate and resume Replay single-question and batch snapshots from session.activate/resume, including locked answers, and extract the helper so the session-actions god-file is not the only owner of that wire shape. Co-authored-by: ClintonEmok <54935030+ClintonEmok@users.noreply.github.com> Co-authored-by: frendo <frendo.wu@gmail.com> * style(desktop): space sibling restore import for eslint * fix(computer-use): recreate CUA session suspect after MCP timeout (#74799) An MCP call_tool deadline hit left the cua-driver session wedged for all later computer-use calls. Mark the session suspect on a concurrent.futures.TimeoutError and tear down + recreate it before the next non-lifecycle call; healthy sessions are never restarted. Fail-closed: the timed-out action may still have taken effect on the remote screen, so it is never silently replayed — the error result carries structuredContent.code=timeout_outcome_unknown with next_step=fresh_state. Informed by #74877 by BlackishGreen33. Co-authored-by: BlackishGreen33 <s5460703@gmail.com> * fix(lsp): retire clients when the protocol reader exits * fix(lsp): abort diagnostics waits after transport death * fix(models): OpenRouter :nitro/:floor routing variants no longer rejected by /model validation OpenRouter's :nitro, :floor, :exacto, and :online suffixes are request-time routing modifiers valid on any model id — /models lists only the base model. validate_requested_model() compared the full suffixed id against the listing, so a valid variant was either rejected outright or fuzzy-auto-corrected to the base id, silently stripping the user's routing opt-in. Now, for OpenRouter only, a recognized variant suffix validates the BASE id against the live listing (and the curated-catalog soft-accept and static- catalog fallback paths) while preserving the suffixed id for persistence and API requests — checked BEFORE fuzzy correction. :free/:batch/:thinking remain direct catalog SKUs and keep exact-match semantics; unknown suffixes and unknown bases are still rejected. Reported by JEB (Jakob's Hermes Agent) via Discord. * feat(tui-gateway): seq-stamped event replay for lossless desktop reconnect Server: per-session monotonic seq on every routed event frame, bounded 512-frame replay ring (64 sessions, FIFO eviction), plus two new RPCs — session.events.since (replay newer-than-watermark, reports latest_seq + truncated so clients detect gaps) and session.events.stats (telemetry). Client: per-session seq watermarks recorded from live frames; after any successful reconnect a fire-and-forget fetchReplay() drains missed events through the normal dispatch path (recordSeq ignores non-increasing seqs, so stale replay can never regress a watermark); focus-triggered reconnect nudge in use-gateway-boot for the Electron unfocused case where macOS wake skips visibilitychange. Replay failures are swallowed by design: lossless resume is an upgrade over the previous lossy reconnect, never a new failure mode. * test: drop unused afterEach import (CI eslint) * fmt(js): `npm run fix` on merge (#94230) Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> * fix(web): keyword-align titles of six high-impression docs pages GSC page-level data shows configuration (171K impr, pos 5.3), quickstart (135K, 5.3), providers (123K, 5.3), web-dashboard (68K, 6.5), docker (37K, 6.3) and the desktop app page (98K, 3.9) all losing ranking headroom because their title/H1 are bare nouns instead of the terms people search. - Frontmatter title + H1 now carry the query terms on all six pages - Desktop docs page links back to the new marketing /desktop product page, joining the two official properties Google sees for the query Done by Hermes Agent (deepseek-v4-pro via nous), Nous Research. * fix(mcp): psutil.pid_exists for stdio children liveness — Windows footgun (#85125 CI) * fix(terminal): annotate sweep as POSIX-only for the killpg guard lint (#85125 CI) * test: pin that install-root siblings still get parent-dir hardening The install-tree exclusion added in #93757 has a positive test (paths inside the tree are skipped) but no negative boundary test. The guard compares path components, so a prefix-named sibling like /opt/hermes-data must still be chmod'd 0700 — but a rewrite to a string-prefix match would silently drop that hardening with the suite staying green. Add test_install_tree_siblings_still_hardened covering a prefix-named sibling and an ordinary sibling of the install root. Verified by mutation: replacing the guard with str(parent).startswith(...) turns the new test red. Follow-up to #93757. * fix: log a warning when parent-dir hardening is skipped for the install tree The install-tree exclusion in secure_parent_dir() (#93757) returned silently. A credential file being written inside the install tree is exactly the misconfiguration signal that produced the production lockouts the exclusion guards against, and it also means a previously hardened path (e.g. a hermes home nested inside a git clone) silently loses its 0700 parent tightening. Emit a single warning naming the skipped directory and the install root so the condition is diagnosable from logs. Follow-up to #93757. * docs: update secure_parent_dir docstring and caller comments for the install-tree exclusion The docstring and all four caller comments still said the helper refuses only / and top-level directories. Since #93757 it also refuses the entire hermes-agent install tree. Bring the docstring and the comments at the four credential-write call sites in line with the actual behavior so future changes are not misled by a stale safety description. Follow-up to #93757. * docs: add operator remediation for install dirs locked to 0700 by older images The Dockerfile fix in #93757 only helps newly built images, and an image upgrade (container recreate) resets the permission because /opt/hermes lives in the image layer. The one stranded case is an old image whose container was stopped and restarted after the lockout: it keeps the 0700 install dir and runs code without the guard. Document the one-line in-place recovery (chmod 0755 /opt/hermes as root) in the Docker troubleshooting section. Follow-up to #93757. * fix(desktop): stop the HUD frosting the window while a turn runs The frost is the whole window rectangle and the `[data-hud-glass]` scrim is what makes it readable, but the two ran on different gates: the scrim is focus-only, while the caller widened the frost to "recent or held" — i.e. for the whole of a turn. Thinking with the composer unfocused therefore raised a bare native material with no scrim over it, which on a light theme is a white slab under the band's unconditionally white ink. Put the frost back on the scrim's gate, and re-run it on the window's own focus changes: clicking away to another app fires no focusout, so the scrim would go while the frost stayed behind. * fix(desktop): say why window enumeration failed instead of swallowing it `read_window_below` answers "could not enumerate windows on this system" on macOS and Windows whatever went wrong, and the three failure paths behind it discarded their errors — so a report where the HUD could see nothing had no way to distinguish the module failing to load, the helper failing to spawn, and the OS answering with nothing. Three different fixes, one sentence. Enumeration now returns the reason, the tool's error carries it, and the HUD's game-overlay watch logs it once before it gives up (it retries twice and then goes quiet forever, which is the other half of why the log said nothing). Linux keeps its environment-derived advice, which is more actionable than the raw exception. * feat(desktop): add Settings toggle for vibe hearts Floating affection hearts were always on with no off switch. Message Reactions in Appearance looks related but only gates message-row tapbacks. Add a separate Vibe Hearts preference (default on) next to it. * test(desktop): pin vibe-hearts toggle across the pet-overlay forward path The overlay window's playVibeHearts() only fires on a reaction forwarded by burstVibeHearts, so the single gate covers it — these tests pin that so a future direct caller shows up as a red test. * fmt(js): `npm run fix` on merge (#94346) Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> * fix(desktop): keep modal context menus inside dialogs * fix(desktop): stop gating edit-menu Paste on the clipboard probe (#91553) The dom context menu disabled Paste unless a renderer-side readClipboard() probe reported text when the menu opened. The items action never consumes that probe: editableCommand("paste") dispatches webContents.paste() in main - the same Chromium path Ctrl+V takes, which resolves the system clipboard itself. On Windows the Win32 clipboard.readText() bridge can return empty while that path succeeds, so Paste stayed grayed out even though pasting would have worked; probe errors were swallowed the same way (.catch(() => undefined)). Fail open instead: drop the gate and the now-unused clipboardHasText fact from the dom menu shape, so opening an editable menu no longer makes the IPC round-trip at all. Pasting with an empty clipboard is a harmless no-op, matching Chromiums own menu, which keeps Paste enabled for editables. The terminal paste item keeps its gate - its action inserts the readClipboard() text into the PTY directly, so there the probe and the action share one mechanism and the gate stays honest. Fixes #91553 * fix(desktop): keep UI scale across in-page route navigation Desktop is a HashRouter over one file:// document, so every route is a distinct URL to Chromium's per-URL zoom store. A route the user never zoomed on has no record at all and resolves to the host default (100%) — that is every fresh session and every never-visited settings tab. In-page navigation fires neither did-finish-load nor any window event, so nothing re-asserted the persisted level. The window dropped to 100% while the Appearance control kept reading the chosen scale, because the renderer only learns of zoom changes through 'hermes:zoom:changed', which never fired. Touching the setting sent a fresh apply, which is why it appeared to fix itself. Re-assert the persisted level on main-frame did-navigate-in-page. Verified on real Electron 40.10.2 / Chromium 144 (win32): a recordless hash route reports 100% at the event, so the existing drift-guard sees the drop and re-applies, and still no-ops when the route's record already matches. Fixes #48658 Fixes #38854 Fixes #79863 Co-authored-by: Brooklyn Nicholson <brooklyn.bb.nicholson@gmail.com> * test(desktop): cover UI scale across recordless hash routes Drives the reported path rather than the helper: set a non-default scale, then navigate to routes Chromium holds no zoom record for, which is what opening a new session looks like to the per-URL store. Keeps the Cmd/Ctrl+N case alongside it. Co-authored-by: Clark Vines <38430798+clarkvines@users.noreply.github.com> * fix(signal): chunk long cron deliveries instead of truncating * fix(signal): chunk long standalone sends and cover both delivery paths (salvage #57929 + #67279) Follow-up to lkz-de's adapter chunking commit: long Signal messages no longer truncate on ANY delivery path. - tools/send_message_tool.py: register Signal's 8000-char limit in _MAX_LENGTHS (imported from the adapter module so the two paths can't drift) so hermes send / cron standalone / MCP sends split via the shared truncate_message() pass instead of signal-cli rejecting them. Standalone-path idea credited to @5L-hermes01 (#67279). - tests: regression test proving standalone Signal sends chunk at the adapter limit with no truncation footer (fails on pre-fix main). - docs: Long Messages section on the Signal page (en + zh-Hans). Both fixes verified by sabotage A/B (tests fail with the respective half reverted to origin/main) and a real-import E2E: 27k-char message with emoji + cross-boundary bold + code blocks -> 4 chunks, all styles in-range UTF-16, lossless reassembly. * feat(terminal): pluggable terminal environment backends via plugin registry Third-party sandbox vendors can now ship a terminal backend as a standalone plugin instead of landing in core. Adds the five-piece pluggable-subsystem pattern for terminal environments: - agent/terminal_env_provider.py — TerminalEnvironmentProvider ABC with declarative classification flags (is_remote, is_container, skip_container_guards, cache_path_base, strip_env_keys, session_isolated_when_nonpersistent) so every historical frozenset-of-names classification site consults the registry instead - agent/terminal_env_registry.py — thread-safe scoped registry; built-in backend names are reserved and unregistrable - PluginContext.register_terminal_environment_provider() mirroring register_browser_provider - _create_environment falls through to registered providers; unknown-backend errors list plugin names - Classification sites wired: approval guard skip, container path/cwd handling (terminal/file/code-exec), prompt-builder env hints + probe, host env probe suppression, skills remote-env note, cache path translation, subprocess secret stripping (both spawn paths), per-session isolation for name-resumed sandboxes - Surfaces: hermes setup picker + doctor + status rows, dashboard terminal-backend picker rows/probe/validation, terminal.backend schema options recomputed per request - Docs: developer-guide/terminal-environment-plugin.md + sidebar + plugins capability table * feat: browser snapshots drop LLM summarization — truncate-and-store like web_extract; auxiliary.web_extract slot removed web_extract stopped using an auxiliary LLM long ago (deterministic truncate-and-store), but browser snapshots still routed oversized accessibility trees through the auxiliary web_extract model, keeping a dead-looking aux slot alive across every config/picker surface. - tools/browser_tool.py: remove _extract_relevant_content and _get_extraction_model; oversized snapshots always truncate at line boundaries, store the full tree to cache/web, and append a read_file pointer (element refs beyond the cut live in the file) - tools/browser_camofox.py: same — no LLM path - Remove auxiliary.web_extract slot: config_defaults (removal note, same pattern as session_search/PR #27590), cli.py defaults + env bridge, gateway/run.py bridged keys, hermes config display, hermes model picker, dashboard REST slots, desktop + web AUX_TASKS, i18n labels (en/zh/ zh-hant/ja/ar) - Docs: env-vars, configuration, fallback-providers, browser + zh-Hans mirrors (web-search zh-Hans was stale on the old LLM pipeline — synced to truncate-and-store truth) - Tests updated: aux bridge uses approval slot, browser tests assert the LLM path is gone and stored files are secret-redacted * fix(tui-gateway): make WS reconnect replay actually deliver events (follow-up to #94219) The #94219 replay was a production no-op: the server returned full JSON-RPC envelopes from session.events.since while the client's replay loop dispatches only elements with a top-level 'type' — every replayed event was silently skipped. Each side's tests validated its own assumption, so both suites stayed green. - server: events_since() now returns bare event objects (the frame's params), the exact shape the live dispatch path consumes; ring stores params directly; cross-language contract test added on both sides. - client: live frames racing an in-flight replay are parked and flushed seq-gated afterward — no double dispatch of deltas, no gap-skip from a watermark advanced past the replay window. - restart poisoning: seq counters are in-process, so a backend restart reset them while clients kept high watermarks (replay forever empty, truncated=false). New replay_epoch advertised in gateway.ready and echoed by session.events.since; the client clears watermarks on epoch change. - methods_session no longer reaches into event_replay privates (is_truncated() accessor). Live repro: pre-fix, 3 stamped frames -> 0 dispatchable by the client gate; post-fix 3/3. Tests: 16 py (replay+ws), 8 vitest, tsc clean, ruff clean. * fmt(js): `npm run fix` on merge (#94410) Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> * fix: system prompt no longer references tools/skills the session can't use; hermes-agent skill is always kept Audit finding (Blank Slate): the system prompt advertised web_search, skill_view, todo, and the hermes-agent skill even when the toolset had none of them — the model chases phantoms it can't call. - hermes-agent skill is now essential: cannot be disabled (config reads strip it, hermes tools writes drop it), cannot be deleted by skill_manage, is re-seeded past curator suppression, and is seeded even on .no-bundled-skills profiles (Blank Slate / --no-skills). - Blank Slate core toolsets grow from file+terminal to file+terminal+vision+skills: read_file cannot read images and points at vision_analyze; the essential skill needs skill_view to load. - HERMES_AGENT_HELP_GUIDANCE degrades to a docs-URL-only variant when skill tools are absent. - Execution-discipline guidance drops its web_search lines when web tools are off (execution_guidance_text renderer). - Skills-index preamble says 'basic tools like terminal' instead of naming web_search when web tools are off. - Coding operating brief drops the todo-tracking sentence when the todo tool isn't loaded. All gating keys off agent.valid_tool_names, fixed at session construction — prompt stays byte-stable per session (cache-safe). * test: opted-out profile seeding now exercises the essential-only sync subprocess * fix(tests): e2e group-restart test no longer flakes on cold SessionDB init The /goal post-turn hook constructs a real SessionDB on an executor thread at the turn boundary. On a cold or loaded CI runner that state.db init can exceed send_and_capture's 2s poll window, so the send lands after the assertion and the test reports the bare 'Expected mock to have been called once. Called 0 times.' (#92130). Mock _run_post_turn_hooks in the e2e runner — these tests exercise gateway command dispatch, not goal hooks. Also scrub TELEGRAM_GROUP_ALLOWED_CHATS / *_GROUP_ALLOWED_USERS / QQ allowlist env vars in the hermetic conftest: a developer shell with those set flips _get_unauthorized_dm_behavior to 'ignore' and fails the pairing e2e test locally. * fix(agent): gate memory provider system_prompt_block on toolset config (#81014) The external memory provider's `system_prompt_block()` was injected unconditionally into the system prompt, while the provider's tools were gated by `memory_provi…
…ounting, Nth-delivery promotion, docstring Follow-ups on top of the cherry-picked NousResearch#93532 cap: - Regression tests: suppressed (in-cooldown) matches must NOT consume the lifetime budget; the cap trips exactly at the Nth DELIVERED match and promotes to notify_on_complete with the watch_disabled summary queued right after the final match. - Extract _emit_lifetime_watch_disabled() and emit the summary even when the global breaker drops the final match, so the user always learns why watching went quiet (parity with the strike-limit path). - Mention the lifetime cap in the terminal tool docstring (the schema text was already updated by NousResearch#93532). Refs NousResearch#93513
…ounting, Nth-delivery promotion, docstring Follow-ups on top of the cherry-picked NousResearch#93532 cap: - Regression tests: suppressed (in-cooldown) matches must NOT consume the lifetime budget; the cap trips exactly at the Nth DELIVERED match and promotes to notify_on_complete with the watch_disabled summary queued right after the final match. - Extract _emit_lifetime_watch_disabled() and emit the summary even when the global breaker drops the final match, so the user always learns why watching went quiet (parity with the strike-limit path). - Mention the lifetime cap in the terminal tool docstring (the schema text was already updated by NousResearch#93532). Refs NousResearch#93513
What does this PR do?
Closes the gap in
watch_patternsrate limiting that lets a recurringpattern force an unbounded number of full-context agent turns.
ProcessRegistry._check_watch_patterns()already has a per-session ratelimit (1 notification / 15s) plus a 3-consecutive-strikes auto-disable, but
that strike counter only advances when a match is dropped inside an
active cooldown window. A match that arrives after the cooldown has already
expired is treated as "clean" and resets the strike counter to 0 — so a
pattern that recurs at a cadence just above the 15s floor (a service
restarted repeatedly over hours/days, for example) never trips the
strike-limit disable. Every one of those matches still synthesizes an
[IMPORTANT: Background process ... matched watch pattern ...]message andforces a brand-new agent turn carrying the entire conversation.
This PR adds a lifetime cap independent of the strike counter:
WATCH_LIFETIME_MAX_HITS(8). Once a session has delivered that manywatch_matchnotifications over its whole life — regardless of how cleanlyspaced they were —
watch_patternsis disabled and the session falls backto
notify_on_completesemantics (one notification when the processactually exits), reusing the exact same disable/promotion path the
strike-limit already uses.
This targets the part of #93513 that the existing three-layer defense
(
97d54f0e4d) doesn't cover: sparse-but-recurring matches. Deeper parts ofthat issue — a token-budget guard before triggering a turn at all, and the
separate WebSocket re-attach defect — are larger, architectural changes to
the gateway's turn-triggering and connection-recovery paths and are left
for a maintainer/dedicated follow-up rather than bundled into this small,
mechanical fix.
Related Issue
Fixes #93513 (partial — see scope note above)
Type of Change
Changes Made
tools/process_registry.py— addWATCH_LIFETIME_MAX_HITSconstant and enforce it in_check_watch_patterns(), queuing the samewatch_disabledsummary event used by the strike-limit path.tools/terminal_tool.py— update thewatch_patternstool-schema description to mention the lifetime cap.tests/tools/test_watch_patterns.py— addTestLifetimeCapcovering sparsely-spaced matches (cooldown reset before every match, so the strike counter never advances) that must still stop after the cap.How to Test
python3 -m pytest tests/tools/test_watch_patterns.py -qgit checkout HEAD~1 -- tools/process_registry.py tools/terminal_tool.py && python3 -m pytest tests/tools/test_watch_patterns.py -k TestLifetimeCap -q(fails withImportError: cannot import name 'WATCH_LIFETIME_MAX_HITS'), thengit checkout HEAD -- tools/process_registry.py tools/terminal_tool.pyto restore.Actual output
tests/tools/test_process_registry.pyand other neighboring suites werealso run; the handful of pre-existing failures there (missing
psutil/ptyprocessin this sandbox, a live-system-guard block on realos.kill)reproduce identically on unmodified
mainand are unrelated to this change.Checklist
Code
fix(scope):,feat(scope):, etc.)Documentation & Housekeeping
docs/, docstrings) — docstring in_check_watch_patternsupdatedcli-config.yaml.exampleif I added/changed config keys — N/A, no new config keyCONTRIBUTING.mdorAGENTS.mdif I changed architecture or workflows — N/Awatch_patternsdescription inTERMINAL_SCHEMAAI assistance disclosure
This PR was prepared by an autonomous coding agent (Claude, via an
OSS-contribution pipeline) that read the issue, located the relevant rate
limiting code, wrote the fix and the failing-first test, and verified it.