fix(dashboard): reload loopback tabs after stale session-token closes - #54022
fix(dashboard): reload loopback tabs after stale session-token closes#54022izumi0uu wants to merge 1 commit into
Conversation
Code Review Summary (Hermes Agent)Verdict: Approved — Dashboard auth reload extraction. Refactors session-storage reload logic into a shared Changes:
Assessment:
|
tonydwb
left a comment
There was a problem hiding this comment.
Code Review Summary
Verdict: Approved
Handles dashboard reload for stale session tokens on loopback WebSocket auth failures. The new dashboard-auth-reload module provides attemptDashboardTokenReloadOnce (latch + reload), clearDashboardTokenReloadAttempt (clear latch on success), and maybeReloadForLoopbackWsAuthFailure (reload once for 4401 close codes). Wired into both fetchJSON (HTTP 401) and GatewayClient/ChatSidebar (WS 4401 close). Comprehensive test coverage across api.ts, gatewayClient.ts, and ChatPage.tsx.
Reviewed by Hermes Agent
|
Thanks for isolating this to the existing one-shot client recovery path. The underlying defect is still present: Problems
Suggested changes
Automated hermes-sweeper review. |
8a4350e to
92278b1
Compare
Loopback dashboard tabs now share one one-shot stale-token recovery path across REST 401s, the PTY socket, the structured event socket, and the shared JSON-RPC gateway wrapper. The shared client exposes only an optional close-event interception hook; the dashboard remains responsible for deciding that loopback 4401 means reload. Constraint: Current main delegates the web gateway to apps/shared JsonRpcGatewayClient, and NousResearch#54022 review requires a shared-client-compatible close-code hook plus direct ChatSidebar event-socket coverage. Rejected: Restore the dashboard's old direct WebSocket implementation | stale against the shared JSON-RPC client and would duplicate transport behavior. Confidence: high Scope-risk: moderate Directive: Keep stale-token policy dashboard-specific; the shared JSON-RPC client should expose close events without learning dashboard auth semantics. Tested: npm --workspace web test (21 files, 106 tests); focused stale-token tests (5 files, 14 tests); npm --workspace web run typecheck; npm --workspace @hermes/shared run lint; npm --workspace @hermes/shared run typecheck; focused web eslint; git diff --check. Not-tested: Manual browser smoke test across a real dashboard restart.
92278b1 to
b212fd0
Compare
|
Addressed in b212fd0. The shared JsonRpcGatewayClient now exposes a narrow onSocketClose interception hook, allowing the dashboard to handle loopback 4401 stale-token failures without restoring the old direct WebSocket implementation or embedding dashboard auth policy in the shared transport. The one-shot reload path now covers REST 401s, shared JSON-RPC closes, PTY closes, and the independent ChatSidebar /api/events socket. I added focused regressions for the shared close hook and ChatSidebar event socket, and kept the component tests production-build-safe by using the existing React runtime instead of a workspace-hoisted test dependency. |
Sibling site missed by PR NousResearch#54022 — /api/console WebSocket in HermesConsoleModal.tsx has the same buildWsUrl → stale-token → 4401 close path as the PTY and events WebSockets. Without this guard, opening the console after a dashboard restart shows 'Console closed (4401). auth: token_mismatch' with no recovery.
|
Merged via #78313. Your commits cherry-picked with authorship preserved (rebase-merge). The salvage also wires the stale-token reload guard into a third WebSocket close handler (HermesConsoleModal /api/console) that was missed in the original PR. Thanks for the clean fix! |
Loopback dashboard tabs now share one one-shot stale-token recovery path across REST 401s, the PTY socket, the structured event socket, and the shared JSON-RPC gateway wrapper. The shared client exposes only an optional close-event interception hook; the dashboard remains responsible for deciding that loopback 4401 means reload. Constraint: Current main delegates the web gateway to apps/shared JsonRpcGatewayClient, and #54022 review requires a shared-client-compatible close-code hook plus direct ChatSidebar event-socket coverage. Rejected: Restore the dashboard's old direct WebSocket implementation | stale against the shared JSON-RPC client and would duplicate transport behavior. Confidence: high Scope-risk: moderate Directive: Keep stale-token policy dashboard-specific; the shared JSON-RPC client should expose close events without learning dashboard auth semantics. Tested: npm --workspace web test (21 files, 106 tests); focused stale-token tests (5 files, 14 tests); npm --workspace web run typecheck; npm --workspace @hermes/shared run lint; npm --workspace @hermes/shared run typecheck; focused web eslint; git diff --check. Not-tested: Manual browser smoke test across a real dashboard restart.
Sibling site missed by PR #54022 — /api/console WebSocket in HermesConsoleModal.tsx has the same buildWsUrl → stale-token → 4401 close path as the PTY and events WebSockets. Without this guard, opening the console after a dashboard restart shows 'Console closed (4401). auth: token_mismatch' with no recovery.
* chore: map bot@bkstock.dev to BKStock
* perf(session): route SQLite PRAGMAs through central apply_database_pragmas
Addresses review from @teknium1 on PR #71755:
- Extended apply_database_pragmas() to handle cache_size, mmap_size,
and temp_store from config.yaml (alongside existing wal_autocheckpoint
and journal_size_limit). No hardcoded defaults — all values are
opt-in via config.yaml, avoiding policy conflicts with other PRs.
- Applied to ALL connection types: writer (_connect_and_init),
read_only cross-profile attach, and WAL per-thread readers
(_get_read_conn). Previously PRAGMAs only ran on the writer path.
- Removed inline PRAGMAs from _connect_and_init — single source of
truth in apply_database_pragmas().
- Documented config keys with examples in function docstring.
* fix(pr): remove remnant local PRAGMAs from PR branch
* test(session): guard config-gated performance PRAGMAs across all connection types
E2E guard for the salvaged PR #71755: database.cache_size/mmap_size/
temp_store from config.yaml must reach the writer connection, the
read-only cross-profile attach, and the WAL per-thread reader — and a
default install (no database: keys) must keep byte-identical SQLite
defaults on every connection type. Also covers integer-coercion
rejection of garbage values for the three new keys.
cache_size uses -16000 (not the doc example -2000) because -2000 is
SQLite's compiled-in default and would not discriminate a regression.
* fix(desktop): flush queued deltas on window focus
* perf(desktop): stop scroll and status loops in busy sessions
* perf(desktop): pause hidden-pane timers in agents view, cron sidebar, and floating pet
Partial pick of the surviving renderer hunks from #75395 (perf commit
6502e441d plus fixup 3fbbc9c1d): gate the 500ms subagent now-ticker and
the cron sidebar 1s ticker/run-poll on usePaneVisible, and skip the
legacy floating-pet poll while the document is hidden. Dropped hunks
(electron/main.ts, vitest.setup.ts/config) intentionally excluded.
* style(desktop): restore alphabetical import order in agents/index.tsx
* refactor(desktop): shared pulse beat + fully-gated cron peek (simplify folds)
Two findings from the simplify pass on the final trio diff:
- status-pulse: one pause controller + one aligned period timer shared by all StatusPulse instances (ref-counted), instead of N x (document/window/bridge listeners + unsynchronized 5s wakes) — a sidebar can show dozens of pulsing dots. Pause still cancels in-flight animations so the compositor sleeps immediately.
- cron-jobs-section: the runs-peek effect created its interval even while the pane was hidden (callback no-oped but the timer still woke the renderer every 8s/60s per expanded job). Early-return when hidden — visibility is already in the dep array, so becoming visible restarts load + timer.
* fix(lint): import sort + eslint-disable for timer-handle ref clear in effect
CI-caught: cron-jobs-section had an extra blank line between sorted imports; use-message-stream's visibility-flush effect assigns flushHandleRef.current=null inside a useEffect (legitimate timer-clear, not an atom mirror) — eslint-disable-next-line per the rule's documented convention.
* fix(state): narrow FTS UPDATE triggers with AFTER UPDATE OF + migration
Retarget #73639 onto the SessionDB mixin split (hermes_state_common /
hermes_state_schema). Fresh installs create UPDATE OF content/tool_*
triggers; existing broad AFTER UPDATE triggers are inspected and
replaced under schema init without an FTS rebuild (WHEN clauses already
guarded content correctness; OF skips non-content status writes that
saturated disk I/O on large state.db).
Tests: tests/test_fts_update_of_narrowing.py (4)
* fix(state): fail closed on CJK trigger migration
* fix(state): quarantine CJK when ensure soft-fails after OF migration
_ensure_fts_cjk_schema never raises on OperationalError; post-condition
after dropping messages_fts_cjk_update now requires a narrowed UPDATE
trigger or durable fts_cjk_stale + unavailable. Covers the production
soft-fail path the raise-only handler missed.
* refactor(state): drop unreachable regex guard in trigger migration
Simplify-pass fold: to_drop names come from the literal update_names\nallowlist via IN binding, so the [A-Za-z0-9_]+ fullmatch could never\nfail — and if it somehow did, its `continue` would miscount (the\nskipped trigger stayed in len(to_drop)/the log while CREATE TRIGGER\nIF NOT EXISTS silently kept the broad variant). Delete the guard and\nits function-local re import; keep the invariant as a comment.
* fix(security): reject always-blocked OpenViking endpoints
## Summary
- Normalize OpenViking endpoints through `is_always_blocked_url` and fall back to the default local endpoint when poisoned.
- Keep intentional loopback / LAN self-host working.
- Add focused unit tests.
## Salvage / credit
Memory-provider endpoint floor sibling of RetainDB/Supermemory always-blocked hardening (avoids over-broad #4984-style private-IP bans).
(cherry picked from commit 8fa607d0aedb8c5fca398d7f112b1b25ade54fa2)
* fix(openviking): fail closed on blocked endpoints
(cherry picked from commit 389a90b81c9c2c89810f2fa7461f8faa9a5c9578)
* fix(openviking): don't spawn a second server onto a live port
`_start_local_openviking_server()` spawned `openviking-server`
unconditionally. Both callers — `initialize()` and the runtime
unreachable handler — reach it from a health probe, and that probe can
time out client-side while the server is up and serving. The spawned
process then loses the data-directory lock and exits immediately with
`DataDirectoryLocked`; because the probe keeps timing out, the cycle
repeats every cooldown window (~5 min observed).
The existing 30s `_failed_refresh` cooldown paces the loop but cannot
stop it, since it expires while the underlying condition persists.
Probe the target host:port before spawning and treat an occupied port as
already-started. This guards both call sites at their single convergence
point. The probe deliberately tests only that a listener owns the port —
enough to know a second server would lose the lock — and says nothing
about that listener's health.
The parse/probe now precedes the PATH lookup, so a reachable server is
reported as running even when `openviking-server` is not on PATH.
Fixes #74846
(cherry picked from commit b49427d85fd6628eb4a7fe099e5c390c5c4cc935)
* fix(openviking): drop stale "disabled for this Hermes run" warnings
The provider used to disable OpenViking permanently when the server was
unreachable. That was fixed: `_ensure_client()` now reconnects lazily,
with a 30s cooldown gate in `_ensure_client_locked`.
Only one of the seven user-facing warnings was updated to match. The
other six still told the user memory was "disabled for this Hermes run",
which is no longer true — every one of those paths is retried on the next
access. A user who reads the old message has no reason to retry, which is
very likely how #5721 ("never recovers") came to be filed against
behaviour that already recovers.
All six sites were traced to confirm none is terminal for the run: the
`initialize()`-time and waiter-thread failures never arm `_failed_refresh`
(only line 2439 does), so they retry on the very next access with no
cooldown at all.
The replacement wording deliberately omits the "(after cooldown)"
parenthetical used at the already-correct site — that detail is only
accurate where `_failed_refresh` was just armed. The neutral phrasing is
true at all six.
Also promotes two clause separators to periods to avoid "…; …disabled;"
collisions.
(cherry picked from commit 8346403a4b97af503d26b0f7905ff513828d821e)
* fix(openviking): re-arm the commit guard after in-place compression
`_committed_session_ids` is a permanent per-sid latch, and
`_session_needs_commit` checks it before the turn counter by design — a
racing sync_turn can re-increment `_turn_count` after commit+reset, so
the guard must win to stop a double-commit.
That is correct for a session being left behind. It is wrong for one
that keeps its id. `compress_context()` commits before rewriting the
transcript in both modes, and with `compression.in_place: true` (the
default) `on_session_switch` receives the same id and does not rotate.
The latch then rejects every later commit for a still-live session — the
next compression, /new, normal session end, startup recovery — so every
post-compression turn is silently never extracted.
Rotation mode is unaffected because a fresh child id is minted and
starts clean, which is what confirms the latch's intent was only ever to
dedupe the departing id.
Clear the latch when compression completes without rotation. Turns
arriving after that point are genuinely new, and this is a defined
moment rather than a race. The rotation path is untouched, so the old
id stays latched and its _finalize_session_async still dedupes against
the compression commit.
Fixes #74695
(cherry picked from commit d1e5c3dc33ef0d43d021662674e1a7cd5e43eecd)
* test(openviking): cover the compression lifecycle, not a hand-set latch
Review feedback: the previous test called _mark_session_committed
directly, so it verified the guard's behavior but not the wiring that
sets it — a future break in the commit_memory_session -> same-id
compression-boundary path would not be caught.
Add a lifecycle regression that drives the real sequence: on_session_end
commits through the actual path, on_session_switch(same id,
reason="compression") crosses the boundary, sync_turn records a genuinely
new turn, and a second on_session_end must produce a second commit POST.
Without the fix it fails showing exactly one commit call, which is the
reported data loss: every turn after the first compression is dropped.
The rotation and /undo tests stay as scope guards.
(cherry picked from commit 0ca5a330630a30b105cbbc32e8a23f2c5ffe0eab)
* fix(memory): read non-secret provider config from config.yaml for OpenViking and RetainDB
OpenViking is_available() only consulted env vars and use_ovcli_config, so an
endpoint saved to config.yaml (e.g. by the Dashboard) reported needs_config;
_resolve_connection_settings() likewise never folded config.yaml's non-secret
fields into its chain. RetainDB initialize() read base_url/project from the
environment only, ignoring the values the Dashboard writes to config.yaml.
Both now resolve non-secret fields as env -> (ovcli ->) config.yaml -> default;
secrets still come from the environment. Adds regression tests for both.
Fixes #68209
(cherry picked from commit dca57915b97b5705b30927a062e1d0f2f23d3841)
* fix(openviking): read recall settings from config.yaml first, env vars as fallback
_recall_config() previously read all settings (recall_limit, score_threshold,
recall_resources, etc.) exclusively from environment variables. This forced
users to store behavioural configuration in .env, violating the Hermes
convention that .env is for secrets only.
The infrastructure to load config.yaml -> memory.openviking was already in
place via _load_hermes_openviking_config(), but _recall_config() never
called it.
Fix: call _load_hermes_openviking_config() and pass its values as the
default parameter to _env_int/_env_float/_env_bool. Env vars still override
config.yaml values, preserving backward compatibility.
Closes #62540
(cherry picked from commit 6aadf1256835745e0302aa3d3b5ae0660b368637)
* test(openviking): cover config.yaml recall settings with temp-HERMES_HOME tests
Add three tests to TestOpenVikingConfigSchema:
1. test_recall_config_reads_from_config_yaml — writes memory.openviking
settings in config.yaml and verifies _recall_config() consumes them.
2. test_recall_config_env_overrides_config_yaml — writes both config.yaml
and OPENVIKING_RECALL_* env vars, verifies env takes precedence.
3. test_recall_config_partial_config_yaml — partially populated config.yaml
falls back to defaults for omitted keys.
All 46 openviking_plugin tests pass (43 existing + 3 new).
(cherry picked from commit b8d7834caf06c6912004333c270faa248eaed4cd)
* fix(openviking): integrate reliability and configuration hardening
* chore(contributors): map OpenViking source authors
* test(retaindb): guard scoped secret config resolution
* fix(openviking): verify servers before sending credentials
* fix(openviking): catch endpoint errors in setup validation functions
Review follow-up for salvaged PR #76782. Three setup-wizard
validation functions called _normalize_openviking_url outside their
try/except blocks. Since _normalize_openviking_url now raises
_OpenVikingEndpointError for blocked or malformed endpoints, an
invalid endpoint would crash the wizard instead of returning a
friendly (False, message) tuple.
- _validate_openviking_auth: move _normalize_openviking_url inside try
- _validate_openviking_root_access: same
- _validate_openviking_setup_values: catch _OpenVikingEndpointError explicitly
- Remove dead ternary in _normalize_openviking_url safety check (candidate
always has http/https scheme by that point)
- Replace redundant float('-inf') < x < float('inf') with math.isfinite()
in _setting_float; drop the redundant infinity check from _setting_int
(is_integer() already rejects inf/nan)
* fix(state): deduplicate session system prompts
* chore: map cicav legacy noreply email
* fix(tui): avoid writable Kanban opens on empty polls
* fix(context): dedupe subdirectory hints by content digest and skip backup/vendor dirs
SubdirectoryHintTracker re-injected identical context files whenever the same
AGENTS.md was reachable through more than one path. Symlinked shared
workspaces, hardlinks, and timestamped backup copies all alias a single file,
so a normal session could ship the same 8KB of instructions two or three
times. Nothing deduped it and nothing excluded directories that only ever
hold copies.
Two changes:
* Track a sha256 of every injected hint body. Repeat content is skipped, and
the working directory's own context file is seeded at construction so the
copy prompt_builder already loaded at startup is never sent again.
* Skip directories that hold copies rather than authoritative context
(backups, node_modules, venv, site-packages, .git, .Trash, vendor, caches).
Screening is relative to working_dir, so a project that legitimately lives
under vendor/ keeps discovering its own subdirectory hints.
Measured on a real session that touched a symlinked shared workspace:
3 injections / ~24,000 chars before, 1 injection / 8,112 chars after.
14 new tests cover symlink aliasing, byte-identical copies, working-dir
seeding, distinct content still being injected, each excluded directory name,
excluded ancestors, and the working-dir-inside-excluded-name case.
* perf(state): batch the turn flush into one SQLite transaction
Re-derivation of #23254 (@devsart95) on today's flush loop. The turn
flush in _flush_messages_to_session_db wrote one BEGIN IMMEDIATE
transaction per message row; a typical agent turn (user + assistant +
tool results) paid 3-8 transactions -- and, off WAL (the default on
macOS while the WAL-reset guard is active), 3-8 fsyncs -- per turn.
Adds SessionDB.append_messages_batch: same row shape as append_message
(shared _prepare_message_row serializer + _MESSAGE_INSERT_SQL column
list, so the two writers cannot drift), same compression-lock and
compression-closed guards, one aggregated session-counter UPDATE, one
transaction for the whole batch. Row serialization stays outside the
write lock.
The flush loop now collects the turn's new rows and writes them in one
call. All-or-nothing pairs exactly with the persisted-marker stamping:
on failure no rows landed and no markers were stamped, so the next
flush re-writes the whole tail (same recovery contract as before,
minus the partial-prefix case that could double-count).
Measured (same harness, 5-message turn, journal_mode=DELETE,
synchronous=FULL): 2.32ms -> 0.83ms median per turn flush (64% faster,
5 fsyncs -> 1). On WAL the win is smaller but the atomicity fix holds.
* perf(tui-gateway): batch branch-seed history copies (whole-bug-class)
Sibling sites of the per-message flush pattern: both branch-seed
paths (session.branch in methods_session.py and the lazy seed persist
in server.py) copied the parent history row-by-row -- one transaction
per row, and a branch seed can be hundreds of rows. Route both through
SessionDB.append_messages_batch. The server.py path also gains real
atomicity: _branch_seed_persisted assumed every row landed, which the
per-row loop could not guarantee.
* test(run-agent): update flush-path fakes and assertions for batched writes
The flush now goes through append_messages_batch; MagicMock-based
assertions and barrier fakes that hooked append_message observed
nothing (the flush's try/except swallowed the AttributeError). Assert
on the batch payload instead.
* refactor(state): fold simplify findings — reuse _insert_message_rows, share guards, chunk seeds
Simplify-pass folds on the #23254 salvage:
- REUSE (HIGH): append_messages_batch now delegates row serialization to
the pre-existing _insert_message_rows helper (already shared by
replace_messages / archive_and_compact / portability import) instead
of adding a third serialization path (_prepare_message_row +
_MESSAGE_INSERT_SQL are gone). One row-writer for every multi-row
path; the row-ID return was consumed by no production caller, so the
batch returns the inserted count.
- QUALITY (HIGH): the compression-lock + compression-closed admission
guards are extracted into _check_transcript_write_guards, shared by
append_message and append_messages_batch (previously duplicated 23
lines that had already needed targeted fixes, #74478). The role-gated
reasoning filtering is no longer duplicated in run_agent.py — it
lives at its one site inside _insert_message_rows.
- EFFICIENCY (MEDIUM, measured): unbounded seed copies hold one BEGIN
IMMEDIATE for seconds (10k rows ~= 2.4s; FTS triggers dominate) and
monopolize the in-process write lock. append_messages_batch grows a
chunk_rows param; all seed/copy call sites use chunk_rows=500. Same
recovery semantics as the old per-row loops, bounded lock holds.
- REUSE (MEDIUM): the two remaining per-row branch-copy loops found by
the pass (gateway/slash_commands.py /branch, hermes_cli
cli_commands_mixin.py branch) are converted to chunked batches too
(AsyncSessionDB's generic to_thread forwarder covers the async site).
Turn-flush benchmark unchanged after the refactor: 2.43 -> 0.87 ms
median per 5-message flush (64% faster).
* fix(tests): update two more append_message.call_args assertions to append_messages_batch
CI-caught: test_verification_stop_caching and test_tui_gateway_server::test_native_vision_turn_persists_a_renderable_image_ref both assert on append_message.call_args, but the flush loop now calls append_messages_batch. Same class of test-fake fallout fixed in 5 other files — these two were missed.
* perf(tui): memoize useSessionLifecycle return (idea from #38491)
Re-derivation of #38491 by @stremtec onto current main (the original is
10,119 commits behind; the hook moved into ui-tui/src/app/). The hook
returned a fresh object literal every render, defeating memoization in
useMainApp's consumers; useMemo over the (all-useCallback-stable)
handles makes the return referentially stable.
Dep array covers ALL nine returned handles incl. trimTail (the
re-derivation initially omitted it - stale-closure class).
* ci: retry uv python install
* fix(state): route session-resume reads through the WAL read-only connection
get_messages_as_conversation, get_resume_conversations, and
get_ancestor_display_prefix still took self._lock — the same global
choke point the read-path split (WAL per-thread read-only connections)
was meant to remove from every recall/browse read. These three are the
hottest reads in the file: every session resume across the gateway,
CLI, and ACP adapter goes through one of them, so a resume racing a
burst of concurrent-session writer flushes still convoys behind them
exactly like the fixed paths used to.
_session_lineage_root_to_tip (the lineage walk shared by all three,
plus get_conversation_root) had its own independent self._lock use and
needed the same conversion — without it the outer functions still
blocked on the very first line.
Verified empirically: a reader thread calling all three functions
while another thread holds self._lock blocked for the writer's full
hold duration before the fix, and returned immediately after (SQLite
3.50.4 in this dev venv falls back to journal_mode=DELETE per the
WAL-reset-bug guard, so the requires_wal-marked regression test is
exercised via a local WAL-forced script instead; it still runs and
passes on any runtime where WAL is actually active).
* chore: add contributor email mapping for ArcherQAQ
* fix(model_metadata): rewrite localhost->IPv4 for the remaining local probe sites
fetch_endpoint_model_metadata's generic (non-LM-Studio) /models fetch and
its llama.cpp /v1/props context-length follow-up built request URLs
straight from the unrewritten candidate, unlike every other local-probe
site. Both retained the multi-second dual-stack IPv6 connect penalty
that _localhost_to_ipv4() exists to skip (measured on macOS: localhost
32.9ms vs 127.0.0.1 0.1ms on a dead port; ~2s on Windows). normalized
stays the cache key so caching behavior is unchanged; only the outbound
request target is rewritten.
Re-derived from PR #61528 onto current main (original no longer applied
cleanly).
* fix(model_metadata): guard _localhost_to_ipv4 against non-string urls
CI slice 3/7 failures: run_conversation tests pass MagicMock base_urls
through the metadata probe path; re.sub raised TypeError where the old
code let non-strings flow through. Preserve that contract.
* perf(cold-start): mitigate ~14s GIL stall during backend init (#60800)
Three fixes for the Desktop/TUI cold-start stall where the event loop
is blocked for ~14s between HERMES_BACKEND_READY and the first
prompt (#60800):
1. copilot_auth: skip subprocess fallback when any
Copilot env var is explicitly set (even if invalid). The user
expressed token intent via env var; silently substituting a CLI
token is surprising and the subprocess adds up to 5s on Windows.
2. tui_gateway/ws: run resolve_skin() via asyncio.to_thread so config
loading + skin engine init do not block the WS read loop during
the cold-start RPC burst.
3. web_server: extend _warm_gateway_module to pre-import the heavy
module chains (auth, copilot_auth, runtime_provider, skin_engine,
inventory, model_switch) that the first WS connection + RPC burst
would otherwise import on the loop thread. These trigger .pyc
compilation and Defender scans on Windows (15-30s per the existing
comment) and were not covered by the original gateway-only warm.
Tests: 5 new tests in test_cold_start_gil_stall.py + 2 new tests in
test_copilot_auth.py. All 36 copilot_auth tests + 16 ws/web_server
tests pass.
* test: harden cold-start regression tests + debug-log the env-var skip
Review folds on the #60807 salvage:
- resolve_skin tests are behavioral (thread-ident probe + ready-frame
wiring check) instead of pure source inspection, per the #72720
pattern; a source assertion remains as belt-and-braces.
- The warm-list test does REAL imports and checks sys.modules —
_warm_gateway_module swallows ImportError by design, so the PR's
tracking-stub test would pass even with a typo'd module name.
- resolve_copilot_token logs a debug line when the env-var
short-circuit skips the gh-CLI fallback (behavioral change made
observable).
* perf(gateway): per-platform skip_context_files to cut agent build latency
Salvage of #26860 (hunk 2, ported \u2014 the PR's base predates the current
gateway layout by ~11.9K commits). Messaging platforms can set
gateway.platforms.<key>.skip_context_files: true to skip the
filesystem-heavy context-file discovery (SOUL.md, AGENTS.md,
.cursorrules walks) during AIAgent construction \u2014 10-100x slower
stat()/walk costs on Windows made this a real per-turn tax. Soul
identity is still loaded (single small file), so the persona survives.
The flag participates in _agent_config_signature so toggling it
rebuilds the cached agent instead of silently reusing a prompt built
under the other setting (prompt-cache correctness).
The PR's hunk 1 (mtime-caching the per-turn dotenv reload) was dropped:
df51ad797 mtime-cached load_config/read_raw_config and c2eda92fd
removed the per-turn deepcopies, capturing most of that win; the
function has since gained a multiplex early-return and managed-scope
overlay that the original whole-function skip would have bypassed.
* fix(relay): route Discord tool-progress into the auto-thread, not the parent channel (#77830)
When a Discord channel message initiates a relay auto-thread, the thread does
not exist at ingest (source.thread_id is None) — the connector creates it on
its FIRST send and auto-threads any outbound carrying the reply anchor. The
final reply carries that anchor, so it lands in the thread. But the
tool-progress / status bubbles (the "Searching the web for..." updates and the
streaming preamble) were sent with _progress_metadata=None and
_progress_reply_to=None: _resolve_progress_thread_id returns None for Discord
(only slack/mattermost get a synthetic thread), so the progress send had no
anchor and the connector posted it FLAT in the parent channel. Result: the
search-status updates leaked outside the thread while the answer threaded
(staging repro 2026-08-02).
The connector now stamps prospective_thread_id on the inbound (the anchor
message id == the id of the thread it will create). Reuse it: when a
relay-delivered Discord channel-initiate carries prospective_thread_id and has
no real thread yet, carry the reply anchor (event_message_id) on both the
progress metadata (reply_to_message_id) and the progress reply_to, so the
connector routes the progress bubble into the SAME auto-thread as the final
reply. Applied to both the tool-progress path (_progress_metadata /
_progress_reply_to) and the status/interim callback path
(_status_thread_metadata). Events already arriving in a real thread, DMs, and
non-relay sources are untouched (guarded on delivered_via_upstream_relay +
prospective_thread_id + not thread_id).
Tests: two new cases in test_run_progress_topics.py — a relay Discord
channel-initiate asserts every progress send carries the anchor (reply_to +
metadata.reply_to_message_id + non_conversational), and an event already in a
real thread asserts the synthetic-anchor path does NOT engage. Full gateway
progress + relay + session suites green (228 passed).
* fix(agent): stop re-probing endpoints that blackhole TCP connects
Salvage of #71282 (Fixes #71281): a routable-but-dead endpoint (corp
LAN address while off-VPN) blackholes TCP SYNs, so every probe in the
model-metadata waterfall waits out its full connect timeout — 20+
seconds of stall per startup across detect_local_server_type,
fetch_endpoint_model_metadata, and the per-model probes.
A module-level blackhole cache keyed on host:port is populated when
any probe observes a ConnectTimeout (httpx or requests; read timeouts
deliberately excluded — an accepted connection is not a blackhole) and
consulted at the top of each guarded function. 30s TTL: long enough to
collapse one startup burst, short enough that VPN recovery is picked
up without a restart. Guard ordering: blackhole check -> disk L2 ->
HTTP waterfall, and a blackholed leg aborts the remaining legs instead
of letting each stall in turn.
Squash of the PR's two real commits (the branch's merge commits made
it un-rebase-merge-able; content verified identical via merge-tree).
* chore: release v0.20.0 (2026.8.3)
The Herald Release — voice (streaming TTS, barge-in, wake words), A2A v1.0,
outbound webhooks, grounded citations, desktop platform wave. ~3,650 commits,
~1,400 PRs, ~1,200 issues closed, 650+ contributors since v0.19.0.
Also: contributor audit additions (18 email mappings, bot-filter widening).
* chore: add contributor email mapping for Ahmett101
* perf(moa): cache resolved preset + per-slot runtime to cut cold-start latency (#66793)
* fix(discord): leave voice channels before cancelling the bot task
`DiscordAdapter.disconnect()` cancelled the bot task before tearing down voice
clients. `leave_voice_channel()` ends in `await vc.disconnect()`, and discord.py
sends a voice state update over the main gateway websocket and then waits for the
voice socket to close. The bot task is the loop running that gateway connection,
so cancelling it first left the handshake with no transport: it could never
complete and blocked until the caller's shutdown timeout fired.
The effect was a fixed ~5s penalty on every shutdown with a voice connection
open, ending in "discord disconnect timed out after 5.0s - forcing continue",
with the voice disconnect abandoned rather than completed.
Measured on a live gateway with a voice connection open in both cases:
before: timed out after 5.0s, all adapters disconnected at +5.29s
after: discord disconnected (0.12s), all adapters disconnected at +0.46s
Moving the voice-cleanup loop above `_cancel_bot_task()` preserves the
zombie-client protection its comment describes: the bot task is still cancelled
before `client.close()`, just after voice teardown rather than before it. Voice
teardown is the one step that still requires a live gateway.
Adds a regression test asserting the ordering. It fails on the previous ordering
at index 1 with `cancel_bot_task != leave_voice_channel:111`.
Fixes #76044
* feat(image): parallelize image_generate batches
* fix(file-sync): serialize concurrent sync cycles
* fix(tool-executor): unpack 5-tuple runnable_calls in _max_workers_for_tool_batch
* fix: exponential backoff for rate-limit fallback cooldown
Replace the fixed 60-second cooldown with exponential backoff:
30min → 1h → 2h → 4h cap.
The counter is reset by restore_primary_runtime on successful
primary-provider recovery, so the backoff is strictly for
consecutive failures within a single degradation window.
Closes #29702
* fix(backoff): keep 60s first-hit cooldown, escalate only on consecutive rate-limits
Review follow-up on the #30223 salvage: the original changed the base
cooldown from 60s to 1800s, benching the primary for 30 minutes on the
FIRST 429 (30x regression in primary-restore latency) and breaking the
existing test_rate_limit_exhaustion_keeps_60s_cooldown contract.
Keep upstream's 60s base and escalate per consecutive rate-limit:
60s -> 2m -> 4m -> 8m -> ... capped at 4h. Counter still resets on
successful primary restore (cicae's mechanism, unchanged).
New tests: escalation doubling, 14400s cap, reset-on-restore.
Existing 60s contract test passes UNCHANGED. Mutation-checked:
escalation disabled -> 2 fail; reset disabled -> 1 fails.
* fix(catalog): wire api_key auth headers for http MCP servers
When an optional-mcps manifest declares transport.type=http with
auth.type=api_key, install_entry() prompts for the key and saves it to
.env, but _build_server_config() only handled the oauth case — the
api_key case produced a bare url entry with no headers, so every
request to the server was unauthenticated (-> 401).
Reuse _bearer_auth_headers(entry.name) from mcp_config.py so the
catalog path emits the same 'Authorization: Bearer ${MCP_..._API_KEY}'
template as the manual 'hermes mcp add --url' path.
Salvaged from #70782 (production hunk applied clean; tests re-anchored
onto current main). Credit: JonthanaHanh.
* perf(compressor): release allocator pages after successful compaction
A successful compaction frees the largest allocation a long session ever
drops (the compressed-away message dicts), but Python's arena allocator
keeps those pages in the heap — RSS retains the pre-compaction
high-water mark until exit. #76905's trim_memory lifecycle covers the
gateway/TUI housekeeping loops but not the CLI compression path.
Call trim_memory(reason='post-compression') at the compression-success
point in ContextCompressor.compress(), following the house pattern
(lazy import in try, debug-level log on failure). The helper is
glibc-gated, config-gated and rate-limited, so it is a safe no-op on
other platforms and cannot fail compression.
Re-expresses the intent of #70782 (JonthanaHanh), which reached for a
bare gc.collect(); trim_memory is the house mechanism and already
wraps a collect.
* fix(catalog): validate http+api_key manifests declare the header's env key
Simplify-pass follow-up on the #70782 salvage: _bearer_auth_headers
hard-emits ${MCP_<NAME>_API_KEY} but install_entry only persists
auth.env-declared vars — a manifest naming its key differently (the
shipped n8n style) would install cleanly yet send a literal-placeholder
header at connect time (silent 401, the #37792 bug class). Enforce the
naming contract at parse time. Also pins the secret-stays-in-.env
property in the install test (raw config.yaml carries the template,
never the secret). Mutation-checked: validation disabled -> guard test
fails.
* fix(agent): cap auxiliary LLM concurrency per task
* fix: thread extra_headers through the call_llm split
The PR's concurrency wrapper splits call_llm into a semaphore-guarded
entry + _call_llm_impl; main added extra_headers to call_llm's
signature after the PR's base, so the split has to forward it too
(dropped silently otherwise — Azure Foundry and custom-endpoint
callers set it).
* fix(nix): tie devShell's HERMES_PYTHON to the venv actually on PATH
`nix/devShell.nix` collected `devShellHook` by scanning every package:
nonNpmHooks = map (p: p.passthru.devShellHook or "") packages;
But `minimal` and `messaging` are `.override` variants of `default`, so
each carries its own `devShellHook` exporting its own HERMES_PYTHON. The
scan therefore concatenated three conflicting exports and forced Nix to
evaluate and realise three separate uv2nix editable venvs on every
`nix develop`.
`attrValues` is alphabetical, so the last hook won (`minimal`) while
`python`/VIRTUAL_ENV came from `default`'s devDeps:
HERMES_PYTHON = ...dimim2... (minimal — no optional deps)
python / VIRTUAL_ENV = ...85r28... (full)
Inside the shell `$HERMES_PYTHON -c "import anthropic"` failed while
`python -c "import anthropic"` succeeded. Worse, `scripts/run_tests.sh`
prefers HERMES_PYTHON, so the suite ran against the minimal venv. Its
guard did not catch this: it only checks that HERMES_PYTHON has pytest,
and minimal's venv does (pytest is in the `dev` group), so the wrong
interpreter was silently accepted.
Tying the hook to `packages.default` — the same package whose `devDeps`
are installed — keeps HERMES_PYTHON, `python`, and VIRTUAL_ENV pointing
at one venv by construction.
editable venvs referenced 3 -> 1
their combined closure 421 MB -> 140 MB
test failures 85 -> 32
The venv mismatch was masking 53 failures; e.g. test_web_tools_config.py
goes 2-failed -> 38-passed. Full suite is now 25369 passed / 32 failed,
and those 32 reproduce identically on a pristine HEAD worktree with no
nix/ changes under the same interpreter (mostly NixOS artifacts — tests
spawning bare `python3` in a scrubbed env exit 127).
* fix(system_prompt): move skills index to the volatile band
The skills index is runtime-mutable: the agent adds and patches skills mid-session, so it is not byte-stable. Keeping it in the stable band breaks that band prefix-cache contract, because every skill edit changes the stable band and invalidates the entire cached prefix in front of it. Move it to the front of the volatile band so the stable scaffold (identity, tool guidance, model guidance) stays cacheable across skill edits.
* docs(system_prompt): fix stale reconstruct_static_prefix docstring example
Simplify-pass finding: the safety note still cited 'skills edited' as a stable-tier input whose change mismatches the rebuilt prefix — after this PR a skill edit changes only the volatile tail (that's the point). Swap the example for genuinely stable-tier inputs.
* fix(prompt_size): search volatile tier for skills block after the stable->volatile move
CI-caught: compute_prompt_breakdown still looked for <available_skills> in the stable tier, but #37117 moved it to volatile. Search volatile first, fall back to stable for older sessions.
* chore: add contributor email mapping for zabih-sudo
* chore: add contributor email mapping for HAOWANG116
* fix(backup): serialize and atomically publish snapshots
* chore: add contributor email mapping for ElSnacko
* fix: prefer explicit anthropic api key
Cherry-picked from PR #58560 by @itsflownium, adapted to current main
(_getenv instead of os.getenv). Moves ANTHROPIC_API_KEY check ahead of
Claude Code credential file and credential_pool auto-discovery so an
explicitly configured key is never shadowed by auto-discovered OAuth.
Fixes #58546
* docs: document /personality none|default|neutral reset across personality docs
The reset keywords have existed in both CLI and gateway handlers since
June but were undocumented — users couldn't find how to cancel a
personality overlay. Adds a 'Resetting to the default' section to the
personality feature page and mentions the reset in the CLI guide,
slash-command reference (both tables), and messaging command table.
* fix(relay): avoid concurrent turn scope corruption
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
* fix(relay): preserve legacy turn shims
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
* fix(relay): gate skipped turn metrics
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
* test(relay): enforce LIFO in overlap regression
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
* fix(relay): preserve skipped turn context
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
* fix comment about relay workaround
* feat(models): add qwen3.8-max to Nous portal + OpenRouter catalogs, replacing qwen3.7-max
Qwen3.8 Max is live on both OpenRouter and the Nous portal
(qwen/qwen3.8-max, 1M context, 131K max output). Per the
newest-max-replaces-last-max convention, it takes qwen3.7-max's slot
in both curated lists.
- hermes_cli/models.py: OPENROUTER_MODELS + _PROVIDER_MODELS[nous]
swap qwen/qwen3.7-max -> qwen/qwen3.8-max
- agent/model_metadata.py: DEFAULT_CONTEXT_LENGTHS entry for
qwen3.8-max at 1,000,000 (verified against OpenRouter live
metadata and Nous /v1/models 2026-08-03)
- tests/test_empty_model_fallback.py: swap incidental catalog fixture
to the surviving slug
- website/static/api/model-catalog.json: regenerated
Pricing snapshot skipped: both routes bill via official_models_api
(live pricing), verified with resolve_billing_route. Reasoning
timeout floor already covered by the qwen3 prefix (180s).
* test: swap context-switch-guard fixture off qwen3.8-max-preview
test_custom_provider_context_avoids_false_shrink_warning used
qwen3.8-max-preview as a slug that deliberately falls through to the
generic 'qwen' 131K catalog match. The new qwen3.8-max
DEFAULT_CONTEXT_LENGTHS entry (1M) now substring-matches the preview
slug too, so the no-custom-providers branch stopped warning. Swap the
fixture to qwen3.9-max-preview, which still hits the generic fallback
— the test's intent (custom_providers threading) is unchanged.
* fix(agent): keep context_length pin for named custom providers
Empty model.base_url plus a runtime custom-provider URL was treated as a
route mismatch, so gateway session-reset banners dropped model.context_length
and fell back to the Qwen family default (131K) while /status still showed
the configured 262K pin.
* test(gateway): cover named-custom context pin on session-info banner
* fix(model_metadata): read llama.cpp context from meta.n_ctx + accept sole model
* chore: add contributor email mapping for johnrazmus
* fix(gateway): keep event loop alive during /compress and Relay drain
Offload manual /compress temporary-agent cleanup through the existing
bounded off-loop helper so a slow agent.close() cannot freeze the
gateway event loop, heartbeat, or platform polling.
Guarantee Relay transport teardown even when the runner cancels
adapter.disconnect() during go_idle: shielded finally, 2s drain-path
idle ACK budget under the 5s outer disconnect budget, and bounded
supervisor/reader/ws.close awaits.
Original commits:
- fix(gateway): offload manual /compress cleanup from the event loop
- fix(gateway): tear down Relay transport even if go_idle is cancelled
- fix(gateway): keep Relay disconnect budgets inside the runner window
By @Dannyzen (PR #78027), salvaged onto current main.
* fix(gateway): bound go_dormant ws.close with teardown timeout
Sibling site to the disconnect() fix: go_dormant() still did an
unbounded await self._ws.close(), the exact same pattern bounded in
disconnect(). go_dormant runs on the scale-to-zero suspend path (Fly
autostop), which also has timeout constraints. Apply the same 1s
wait_for treatment using _TEARDOWN_AWAIT_TIMEOUT_S.
Found during review of PR #78027.
* fix: close the Codex app-server session on agent teardown
Salvage of #65260's b7d7cfd0e (ported — the PR's close() predates ~4K
commits of teardown-step churn, so the hunk is re-anchored after step
6b rather than cherry-picked).
agent/codex_runtime.py already drops _codex_session on turn crash and
on retirement, but AIAgent.close() — the hard teardown for /new,
/reset, and session expiry — had no owner for it, so the app-server
child process survived until interpreter exit. Long-lived gateways
accumulate one leaked subprocess per ended Codex session.
The attribute is cleared BEFORE close() so a concurrent reader can't
observe a half-closed session and a raising close() can't strand a
stale reference (tested).
Tests extend the author's original lifecycle test with the
raising-close and no-codex-session cases.
* fix(agent): discard bare tool-call marker before fallback/persistence (#78148)
Local tool-call templates can emit a bare bracketed token (e.g. "[memory]")
as assistant content alongside a function call. The loop treated that
protocol scaffolding as visible content: it got cached as the post-tool
fallback, and when the next turn came back empty, the marker was replayed
as the final response and written into the persisted transcript. Later
context compaction preserved that history, letting the model repeat the
marker in subsequent turns.
Detect content that is only a bracketed marker (`[name]`) when the
response also carries tool_calls, and drop it before it can be cached
or persisted. Scoped narrowly: only fires alongside tool_calls, so a
genuine final response of "[memory]" without a tool call is unaffected.
* fix(agent): repair sessions already contaminated with stale tool-call markers (#78148)
The conversation_loop fix (previous commit) stops new "[memory]"-style
bare tool-call markers from being cached/persisted, but sessions written
before that fix can still carry rows where a bare marker was saved as
the assistant's "final response".
Add a load-on-read repair pass in hermes_state.py, mirroring the existing
_strip_background_review_harness defense-in-depth: on session restore,
any assistant row whose content is only a bracketed marker (e.g.
"[memory]", "[skill_manage]") AND that carries tool_calls has its content
blanked before the history re-enters the model's context. The tool call
and its result are left untouched so provider tool_call/tool_result
pairing stays intact. Sessions with no affected rows pass through the
normal path unchanged.
* feat(cli): add sessions clean-markers to permanently purge stale tool-call markers (#78148)
The load-on-read repair (_strip_stale_tool_call_markers) fixes affected
sessions in memory on every resume, but never touches the DB — long-lived
sessions re-scan and re-repair the same rows on every load, and the
contaminated bytes stay in state.db (and any backup/cache snapshot of it)
indefinitely.
Add SessionDB.purge_stale_tool_call_markers(dry_run=False): a one-time,
idempotent UPDATE that permanently blanks the content column on affected
rows. Only content is touched — tool_calls and every other column are
left untouched, so provider tool_call/tool_result pairing survives.
dry_run reads through the no-lock read path and never writes.
Wire it up as `hermes sessions clean-markers [--dry-run]`, mirroring the
existing optimize/repair subcommands. Verified end-to-end against a real
temp state.db: dry-run reports the row without writing, the real run
clears it and preserves tool_calls, and a second run is a no-op.
* fix(cli): back up state.db before clean-markers writes by default
purge_stale_tool_call_markers ran a permanent, irreversible UPDATE with
no backup — inconsistent with repair_state_db_schema's backup-by-default
convention for destructive state.db operations elsewhere in this file.
Take a full snapshot via VACUUM INTO (safe against a live connection,
unlike the raw-copy _backup_db_file used for malformed-schema repair)
before the write, timestamped beside state.db. Skipped when dry_run or
when there's nothing to change. Add --no-backup to `hermes sessions
clean-markers`, mirroring `sessions repair`.
Verified end-to-end: the CLI run against a real temp state.db produces
the backup file before printing the cleared-row count.
* refactor: dedup stale-marker regex — use compiled _STALE_MARKER_RE in conversation_loop
The bracketed-marker regex was inlined in conversation_loop.py as
re.fullmatch(r"\[...", ...) while hermes_state.py defines the same
pattern as _STALE_TOOL_CALL_MARKER_RE. Both must agree on what counts
as a stale marker — a drift here means the runtime guard silently
disagrees with the load-on-read repair and CLI purge in hermes_state.
Consolidate onto a single compiled constant (_STALE_MARKER_RE) at
module level in conversation_loop.py, with a comment noting it must
mirror _STALE_TOOL_CALL_MARKER_RE in hermes_state.py. A direct import
from hermes_state was tried first but caused a regression: hermes_state
initializes DEFAULT_DB_PATH = get_hermes_home() / 'state.db' at module
import time, which breaks tests that monkeypatch get_hermes_home() to
return a str (test_slash_worker_accepts_profile_home).
Follow-up to PR #78175 (@JoaoMarcos44).
* fix(conversation_loop): compress messages on output-cap retry path (#55546)
The output-cap retry loop reduced max_tokens by 64 tokens per attempt but
never called _compress_context(), so the compressor never fired. Input
growth (~65 tokens/attempt) canceled the savings, leaving the session
stuck at 200,001 tokens — 1 over the 200,000 ceiling.
The fix adds compression to the output-cap retry path. The compressor
drops the middle window, freeing ~50% of tokens. If compression makes
>=5% savings, the session continues; otherwise vision payloads are
stripped or the session ends with compression_exhausted=True.
Also adds CHANGELOG.md entry and bug fix report.
* fix(conversation_loop): prune dead vision-strip fallback; harden output-cap retry tests
* chore: drop CHANGELOG.md and docs/reports/ — not shipped with salvage PRs
* perf: reuse request_input_estimate instead of recomputing estimate_request_tokens_rough
The output-cap error handler already computes request_input_estimate at
line 4722 via estimate_request_tokens_rough(api_messages, tools=...).
The new compression block ~50 lines below was calling the same function
with the same inputs again. Reuse the existing local.
* chore: AUTHOR_MAP — add BobClawblaw for PR #77870 salvage
Bare noreply email (no NNN+ prefix) needs explicit mapping.
* chore(ci): rerun checks
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
* perf(gateway): prewarm /model picker cache on TUI startup
The classic CLI run() loop calls prewarm_picker_cache_async() during the
idle window after the banner is shown, so the first /model open hits a warm
provider-models disk cache and renders in ~100ms. The stdio TUI entry point
never did this, so the first /model open in a TUI session blocked on serial
/v1/models fetches for every authenticated provider.
Mirror the CLI behaviour: kick off the same off-thread prewarm right after
gateway.ready is emitted (banner shown, user about to type). Fire-and-forget,
guarded once-per-process, fully exception-isolated so a slow or offline
provider can never affect TUI startup.
* test(tui): pin picker-cache prewarm wiring in entry.main()
teknium's review gap on #72021: the helper's worker/once-guard was
covered, but nothing asserted the stdio TUI entry point actually
invokes prewarm_picker_cache_async() — or that it does so in the right
place. Add a focused entrypoint test that runs the real entry.main()
with stubbed collaborators (same monkeypatch-module-attrs harness as
test_tui_entry_mcp_owner.py), spies on the helper in
hermes_cli.model_switch (the lazy-import source), and asserts:
- prewarm fires exactly once, strictly AFTER the gateway.ready write
- startup stays non-blocking: main() reaches the stdin loop and
returns on EOF
- a prewarm failure is swallowed (fire-and-forget) without breaking
startup
Mutation-checked: deleting the prewarm hunk from entry.py fails both
tests.
* perf(cli): check local auth.json/config before slow provider registry sweep
_has_any_provider_configured() probed every api_key provider (gh subprocess
for copilot alone takes 5s; full sweep ~18s) before consulting auth.json and
config.yaml, which are instant local reads. Desktop setup.status calls
blocked past the UI's timeout, causing the connect/disconnect boot loop.
Reorder so cheap local checks run first. Same semantics, ~35x faster here.
* test(cli): regression tests pinning auth-first ordering skips registry sweep
Teknium's review on #63457: existing tests pin the final boolean but not
that the slow PROVIDER_REGISTRY sweep is skipped. Add three tests that
booby-trap hermes_cli.auth.get_auth_status and verify
_has_any_provider_configured() short-circuits on:
- config.yaml model.provider
- config.yaml base_url/api_key (custom endpoint shape)
- auth.json active_provider (sweep-only call-pattern guard)
Mutation-checked: reverting the reorder makes all three fail.
* perf(desktop): keep spinner frames out of React commits
Advance the existing animated status glyph through its DOM text node instead of React state, and pause its timer for hidden panes or inactive windows. Cover frame advancement, zero update-phase commits, and timer suspension with behavior tests.
* test(desktop): cover minimized/hidden window-state + visibilitychange pause for GlyphSpinner
Regression coverage requested in review of #74357: mock
window.hermesDesktop.onWindowStateChanged (pattern from
persistent.test.tsx) and assert minimized/hidden clears the spinner
interval while restore resumes it; also cover document.visibilityState
hidden/visible via visibilitychange.
* fmt(js): `npm run fix` on merge (#78271)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
* perf(gateway): replace SSE poll loop with call_soon_threadsafe-fed asyncio.Queue
_write_sse_chat_completion and _write_sse_responses bridged their
stream_delta_callback queue into the event loop via
`await loop.run_in_executor(None, lambda: stream_q.get(timeout=0.5))`
in a while-True poll — a thread-pool round trip on every 0.5s tick even
when idle, plus up to 500ms of tail latency between a delta landing in
the queue and it reaching the SSE response.
Add ThreadSafeAsyncQueue (asyncio.Queue + a put_threadsafe() that wraps
call_soon_threadsafe), used by both streaming producer closures
(_on_delta, tool start/complete callbacks — all invoked from the worker
thread running run_conversation via loop.run_in_executor). Consumers
now do a plain `await asyncio.wait_for(stream_q.get(), timeout=0.5)` —
woken immediately when a delta arrives, no executor hop, no poll
interval.
Updated tests/gateway/test_sse_agent_cancel.py's 7 call sites to
construct ThreadSafeAsyncQueue inside the running loop (required, since
it captures asyncio.get_running_loop() at construction) instead of a
bare queue.Queue() at test-method scope.
* refactor(gateway): extract _sse_frame() helper, dedup 5 inline SSE encode call sites
_write_sse_chat_completion had five near-identical
f"data: {json.dumps(...)}\n\n".encode() (and one event-tagged variant)
scattered across its role/content/finish/error chunk writes. Pure
extract-method, no behavior change: encoding is byte-identical for every
call site touched.
Left the pre-serialized-string writers elsewhere (_write_event's
json.dumps(..., ensure_ascii=False) path, the /v1/runs SSE writer) alone
— routing them through this helper's plain json.dumps(data) would
silently change their unicode-escaping behavior, which is out of scope
for a pure dedup.
* refactor(gateway): route all three SSE writers through _sse_frame()
_extend _sse_frame with an explicit ensure_ascii param (default True,
byte-identical to a bare json.dumps) and route the two sibling writers
through it: _write_sse_responses._write_event and the /v1/runs event
stream. This completes the dedup PR #65009 — previously only the five
_write_sse_chat_completion sites used the helper, leaving the other two
writers on inline json.dumps with no shared shape.
No behavior change: every writer's emitted bytes are unchanged (verified
byte-for-byte, including non-ASCII payloads where the default
ensure_ascii=True matches the original inline encoders). The ensure_ascii
option is exposed so a future writer can opt into raw non-ASCII bytes
without fractalizing the format again.
Adds tests/gateway/test_sse_frame.py asserting the byte-contract
invariant between _sse_frame and the historical inline encoders.
* refactor(gateway): route session event stream through _sse_frame (ensure_ascii=False)
The session event stream (api_server.py:~2236) was the one genuinely
unicode-distinct SSE writer — json.dumps(payload, ensure_ascii=False) +
.encode('utf-8'). Every other writer uses plain json.dumps. Route it
through _sse_frame(..., ensure_ascii=False) so _sse_frame is now the single
source of truth for ALL SSE frame serialization in the module (chat-
completion, responses._write_event, /v1/runs, and the session stream).
Byte-identical for non-ASCII payloads: verified against the historical
inline encoder (raw bytes preserved). The ensure_ascii=False path is now
exercised by test_sse_frame_ensure_ascii_false_reproduces_session_event_stream.
* test: add cross-thread put_threadsafe + long-reasoning tail tests
Addresses teknium1 sweeper review (2026-07-30) requiring coverage of:
1. ThreadSafeAsyncQueue.put_threadsafe() off-loop boundary: a real
daemon thread pushes into the queue from outside the owning event loop
while the consumer awaits get(), mirroring the run_conversation
worker-thread producer path. Includes a 20-concurrent-thread
no-drop regression.
2. Long-reasoning bound stability for thinkingPreview: 100k-char input
plus empty/collapsed cases must not crash and must retain the visible
tail marker inside the bounded 24k clean window.
* fix: reconstruct fused test after conflict resolution
The conflict-marker strip fused test_agent_task_raises with the body of
test_failed_result_dict — restore both as separate tests (content from
the PR head, verified verbatim).
* perf(tui): bound reasoning-clean input to the displayed tail
CI caught a split defect in this salvage: the PR's long-reasoning tail
test was kept but its production hunk was dropped as 'cosmetics'. It
isn't — cleanThinkingText runs several full-string regex passes and
reasoning grows on every streamed token, so re-cleaning the whole
accumulated string per chunk is O(n) per token / O(n^2) per stream.
Only the tail is displayed (boundedLiveRenderText caps it downstream),
so bound the input to 1.5x LIVE_RENDER_MAX_CHARS first.
Restores the one text.ts hunk from cd99e65fc (author preserved); the
italic-thinking display change and profile script from that commit
remain out of scope.
* test: exercise the production _loop_ref path in put_threadsafe tests
Gate finding (/simplify-code pass): both cross-thread tests passed
loop=loop explicitly, but no production caller does — all six
(_on_delta, _on_tool_*) rely on the queue resolving its own
_loop_ref in __init__. The kwarg made the tests vacuous: a broken
_loop_ref still passed them.
Dropping the kwarg exercises the real path. Verified by mutation:
with self._loop_ref = asyncio.new_event_loop() (wrong loop), both
tests now FAIL; they passed before this change.
* fix(test): feed the SSE writers an asyncio queue, not queue.Queue
CI caught a missed caller-shape update. Both PRODUCTION callers of
_write_sse_chat_completion / _write_sse_responses were converted to
ThreadSafeAsyncQueue, but two pre-existing tests in
tests/gateway/test_api_server.py construct the writer's queue
themselves and still passed a stdlib queue.Queue.
The consumer now does 'await asyncio.wait_for(stream_q.get(), ...)',
which on a queue.Queue blocks the thread forever:
test_stream_cancelled_persists_incomplete_snapshot hung until
pytest-timeout killed it (CI reported the whole file as 'no tests
ran (timeout before collection)'). The sibling disconnect test only
survived because it pre-fills before the first await.
tests/gateway/test_api_server.py: 99 passed (was 1 failed + a 60s
hang); with the SSE/api_server suites: 147 passed.
* fix(credential-pool): re-select in acquire_lease after a deferred refresh
select() re-selects once deferred single-use-token refreshes complete;
acquire_lease() performed the refresh but returned its pre-refresh
answer. Since _acquire_lease_under_lock returns early exactly when a
refresh is pending (if not available: return None, pending_refresh),
a pool whose entries all needed a refresh always returned None — the
caller failed an answerable request right after the refresh succeeded.
Retry once, only when the first pass was empty and a refresh ran.
Post-merge gate-sweep finding on the #71775 salvage (#77714).
* fix(credential-pool): lock the quarantine read-modify-write of _entries
#71775 moved deferred single-use-token refreshes outside the pool lock
(correct — they hold a cross-process flock plus network I/O). But
_refresh_entry_impl's three terminal-auth-failure quarantine paths do a
bare read-modify-write of self._entries. Those used to run with the
caller holding self._lock; on the deferred path they run unlocked, so a
concurrent mutation between the read and the write is silently lost.
Wrap all three in 'with self._lock' (an RLock, so locked callers
re-enter safely) and correct the _refresh_pending_entries docstring,
which claimed the mutations were already self-locking.
Post-merge gate-sweep finding on the #71775 salvage (#77714).
Sibling to the acquire_lease re-select fix.
* perf(file-ops): eliminate redundant subprocess calls in write_file and V4A patch path
write_file currently spawns up to 6 subprocesses per call:
1. mkdir -p (separate call before atomic write)
2. cat (to read pre-content for lint/BOM/line-ending detection)
3. _atomic_write (mktemp + write + mv — the essential one)
4. wc -c (to measure bytes written)
5. _check_lint_delta (post-write lint — also essential)
6. LSP snapshot (also essential)
This PR removes three of them without changing any observable behavior:
1. Fold mkdir -p into _atomic_write shell script (−1 subprocess/write)
The atomic write script already runs a single shell; adding mkdir -p
to it costs zero extra processes.
2. Add optional pre_content parameter to write_file (−1 subprocess/patch)
patch_replace and V4A _apply_update already read the file for fuzzy
matching. Passing that content as pre_content skips the redundant cat
inside write_file. Fully backward-compatible: callers that don't pass
pre_content still read from disk as before.
3. Replace wc -c with len(content.encode('utf-8')) (−1 subprocess/write)
We already have the content in memory; encoding it to get the byte count
is equivalent to wc -c for UTF-8 text.
4. Remove redundant _check_lint loop in apply_v4a_operations (−N subprocesses/V4A)
write_file already runs _check_lint_delta internally. The old code ran a
bare _check_lint(f) loop over all modified files — a re-read + re-lint
without post_content context. Now lint results propagate from write_file
via a four-tuple return, zeroing out the extra subprocesses.
Net effect:
- write_file: 6 → 3 subprocesses per call (new files)
- patch_replace: 6 → 5 subprocesses per call (pre_content skips cat)
- V4A multi-file patches: saves 1 subprocess per modified file
- A typical 4-file V4A patch drops from ~28 to ~16 subprocess calls
* fix(file-ops): decouple BOM detection from pre_content, add V4A backward compat
Bug 1 (UTF-8 BOM loss on V4A UPDATE):
_file_has_bom() trusted pre_content for BOM detection, but the most
common pre_content provider — read_file_raw() — deliberately strips
BOMs so the agent never sees U+FEFF glyphs. Passing BOM-stripped
content through pre_content caused a false-negative: the method
returned False and write_file() silently removed the marker on rewrite.
Fix: _file_has_bom() now always probes the first 3 bytes on disk
(head -c 3), ignoring pre_content for BOM purposes. pre_content is
still used by two other consumers — line-ending detection and lint/LSP
delta computation — neither of which is affected by BOM stripping.
Bug 2 (backward compatibility):
_apply_update() called write_file(path, content, pre_content=...) as a
keyword argument. Duck-typed file_ops implementations that only
implement the two-argument write_file(path, content) contract would
raise TypeError.
Fix: wrap the call in try/except TypeError, falling back to the
two-argument form when the keyword is not accepted.
Also declare tomli in pyproject.toml (pre-existing conditional import
for pre-3.11 Python, caught by the pre-commit dep scan after staging
file_operations.py).
Tests:
Add TestV4ABomRoundTrip with two cases:
- UPDATE on BOM-bearing file preserves the marker
- UPDATE on plain file does not inject a BOM
Addresses teknium1 review on PR #55661.
* fix(file-ops): surrogatepass in bytes_written encode (review finding)
Content that flowed through a surrogateescape decode (backend output via
patch_replace) can carry lone surrogates; a strict encode raises
UnicodeEncodeError where the old wc -c path could not. Mirrors the
existing sha256 verification encode.
* refactor(file-ops): fold simplify-pass findings
- write_file: encode content once, share bytes between bytes_written and
the sha256 verification (drops a second full-content encode per write)
- patch_parser: replace the except-TypeError retry around
write_file(pre_content=...) with signature-based feature detection so a
TypeError raised inside a capable implementation propagates instead of
triggering a duplicate write; tests for both duck-typing contracts
- tests: real-ops V4A BOM round-trip + _file_has_bom disk-probe guard
(the teknium1-review regression previously only covered by a fake)
- comment: document dirs_created's long-standing "parent ensured" meaning
* chore: update uv.lock for tomli dependency (rebase fix)
* chore: remove dead tomli dependency declaration
requires-python is >=3.11 so tomllib is always in stdlib; the
tomli fallback branch in _lint_toml_inproc was unreachable. Removes
the dependency from pyproject.toml + uv.lock and deletes the dead
try/except ImportError fallback in the code.
* fix(telegram+sqlite): resolve polling conflict loop + misleading WAL warning
#75017: Telegram polling conflict retry used drop_pending_updates=False,
starting a new getUpdates session that immediately got 409'd by the
previous still-expiring session — creating the very conflict it was
trying to recover from. Switch to drop_pending_updates=True so Telegram
terminates stale sessions. Also add a recovery-generation guard so the
first transient getUpdates success after a retry doesn't reset the
conflict counter back to 0 (defense-in-depth from PR #75096).
#75153: The WAL-reset warning always said 'hermes update can repair'
even for git/pip/system Python installs where it can't. Now uses
detect_install_method() + recommended_update_command_for_method() to
give a context-appropriate hint (hermes update for git, docker pull for
docker, nix message for nix, generic install hint as fallback).
* chore: map jun@junho.co to junhohong
* fix(dashboard): reload loopback tabs after stale session-token closes
Loopback dashboard tabs now share one one-shot stale-token recovery path across REST 401s, the PTY socket, the structured event socket, and the shared JSON-RPC gateway wrapper. The shared client exposes only an optional close-event interception hook; the dashboard remains responsible for deciding that loopback 4401 means reload.
Constraint: Current main delegates the web gateway to apps/shared JsonRpcGatewayClient, and #54022 review requires a shared-client-compatible close-code hook plus direct ChatSidebar event-socket coverage.
Rejected: Restore the dashboard's old direct WebSocket implementation | stale against the sh…
…85) (#148) * fix(lint): import sort + eslint-disable for timer-handle ref clear in effect CI-caught: cron-jobs-section had an extra blank line between sorted imports; use-message-stream's visibility-flush effect assigns flushHandleRef.current=null inside a useEffect (legitimate timer-clear, not an atom mirror) — eslint-disable-next-line per the rule's documented convention. * fix(state): narrow FTS UPDATE triggers with AFTER UPDATE OF + migration Retarget #73639 onto the SessionDB mixin split (hermes_state_common / hermes_state_schema). Fresh installs create UPDATE OF content/tool_* triggers; existing broad AFTER UPDATE triggers are inspected and replaced under schema init without an FTS rebuild (WHEN clauses already guarded content correctness; OF skips non-content status writes that saturated disk I/O on large state.db). Tests: tests/test_fts_update_of_narrowing.py (4) * fix(state): fail closed on CJK trigger migration * fix(state): quarantine CJK when ensure soft-fails after OF migration _ensure_fts_cjk_schema never raises on OperationalError; post-condition after dropping messages_fts_cjk_update now requires a narrowed UPDATE trigger or durable fts_cjk_stale + unavailable. Covers the production soft-fail path the raise-only handler missed. * refactor(state): drop unreachable regex guard in trigger migration Simplify-pass fold: to_drop names come from the literal update_names\nallowlist via IN binding, so the [A-Za-z0-9_]+ fullmatch could never\nfail — and if it somehow did, its `continue` would miscount (the\nskipped trigger stayed in len(to_drop)/the log while CREATE TRIGGER\nIF NOT EXISTS silently kept the broad variant). Delete the guard and\nits function-local re import; keep the invariant as a comment. * fix(security): reject always-blocked OpenViking endpoints ## Summary - Normalize OpenViking endpoints through `is_always_blocked_url` and fall back to the default local endpoint when poisoned. - Keep intentional loopback / LAN self-host working. - Add focused unit tests. ## Salvage / credit Memory-provider endpoint floor sibling of RetainDB/Supermemory always-blocked hardening (avoids over-broad #4984-style private-IP bans). (cherry picked from commit 8fa607d0aedb8c5fca398d7f112b1b25ade54fa2) * fix(openviking): fail closed on blocked endpoints (cherry picked from commit 389a90b81c9c2c89810f2fa7461f8faa9a5c9578) * fix(openviking): don't spawn a second server onto a live port `_start_local_openviking_server()` spawned `openviking-server` unconditionally. Both callers — `initialize()` and the runtime unreachable handler — reach it from a health probe, and that probe can time out client-side while the server is up and serving. The spawned process then loses the data-directory lock and exits immediately with `DataDirectoryLocked`; because the probe keeps timing out, the cycle repeats every cooldown window (~5 min observed). The existing 30s `_failed_refresh` cooldown paces the loop but cannot stop it, since it expires while the underlying condition persists. Probe the target host:port before spawning and treat an occupied port as already-started. This guards both call sites at their single convergence point. The probe deliberately tests only that a listener owns the port — enough to know a second server would lose the lock — and says nothing about that listener's health. The parse/probe now precedes the PATH lookup, so a reachable server is reported as running even when `openviking-server` is not on PATH. Fixes #74846 (cherry picked from commit b49427d85fd6628eb4a7fe099e5c390c5c4cc935) * fix(openviking): drop stale "disabled for this Hermes run" warnings The provider used to disable OpenViking permanently when the server was unreachable. That was fixed: `_ensure_client()` now reconnects lazily, with a 30s cooldown gate in `_ensure_client_locked`. Only one of the seven user-facing warnings was updated to match. The other six still told the user memory was "disabled for this Hermes run", which is no longer true — every one of those paths is retried on the next access. A user who reads the old message has no reason to retry, which is very likely how #5721 ("never recovers") came to be filed against behaviour that already recovers. All six sites were traced to confirm none is terminal for the run: the `initialize()`-time and waiter-thread failures never arm `_failed_refresh` (only line 2439 does), so they retry on the very next access with no cooldown at all. The replacement wording deliberately omits the "(after cooldown)" parenthetical used at the already-correct site — that detail is only accurate where `_failed_refresh` was just armed. The neutral phrasing is true at all six. Also promotes two clause separators to periods to avoid "…; …disabled;" collisions. (cherry picked from commit 8346403a4b97af503d26b0f7905ff513828d821e) * fix(openviking): re-arm the commit guard after in-place compression `_committed_session_ids` is a permanent per-sid latch, and `_session_needs_commit` checks it before the turn counter by design — a racing sync_turn can re-increment `_turn_count` after commit+reset, so the guard must win to stop a double-commit. That is correct for a session being left behind. It is wrong for one that keeps its id. `compress_context()` commits before rewriting the transcript in both modes, and with `compression.in_place: true` (the default) `on_session_switch` receives the same id and does not rotate. The latch then rejects every later commit for a still-live session — the next compression, /new, normal session end, startup recovery — so every post-compression turn is silently never extracted. Rotation mode is unaffected because a fresh child id is minted and starts clean, which is what confirms the latch's intent was only ever to dedupe the departing id. Clear the latch when compression completes without rotation. Turns arriving after that point are genuinely new, and this is a defined moment rather than a race. The rotation path is untouched, so the old id stays latched and its _finalize_session_async still dedupes against the compression commit. Fixes #74695 (cherry picked from commit d1e5c3dc33ef0d43d021662674e1a7cd5e43eecd) * test(openviking): cover the compression lifecycle, not a hand-set latch Review feedback: the previous test called _mark_session_committed directly, so it verified the guard's behavior but not the wiring that sets it — a future break in the commit_memory_session -> same-id compression-boundary path would not be caught. Add a lifecycle regression that drives the real sequence: on_session_end commits through the actual path, on_session_switch(same id, reason="compression") crosses the boundary, sync_turn records a genuinely new turn, and a second on_session_end must produce a second commit POST. Without the fix it fails showing exactly one commit call, which is the reported data loss: every turn after the first compression is dropped. The rotation and /undo tests stay as scope guards. (cherry picked from commit 0ca5a330630a30b105cbbc32e8a23f2c5ffe0eab) * fix(memory): read non-secret provider config from config.yaml for OpenViking and RetainDB OpenViking is_available() only consulted env vars and use_ovcli_config, so an endpoint saved to config.yaml (e.g. by the Dashboard) reported needs_config; _resolve_connection_settings() likewise never folded config.yaml's non-secret fields into its chain. RetainDB initialize() read base_url/project from the environment only, ignoring the values the Dashboard writes to config.yaml. Both now resolve non-secret fields as env -> (ovcli ->) config.yaml -> default; secrets still come from the environment. Adds regression tests for both. Fixes #68209 (cherry picked from commit dca57915b97b5705b30927a062e1d0f2f23d3841) * fix(openviking): read recall settings from config.yaml first, env vars as fallback _recall_config() previously read all settings (recall_limit, score_threshold, recall_resources, etc.) exclusively from environment variables. This forced users to store behavioural configuration in .env, violating the Hermes convention that .env is for secrets only. The infrastructure to load config.yaml -> memory.openviking was already in place via _load_hermes_openviking_config(), but _recall_config() never called it. Fix: call _load_hermes_openviking_config() and pass its values as the default parameter to _env_int/_env_float/_env_bool. Env vars still override config.yaml values, preserving backward compatibility. Closes #62540 (cherry picked from commit 6aadf1256835745e0302aa3d3b5ae0660b368637) * test(openviking): cover config.yaml recall settings with temp-HERMES_HOME tests Add three tests to TestOpenVikingConfigSchema: 1. test_recall_config_reads_from_config_yaml — writes memory.openviking settings in config.yaml and verifies _recall_config() consumes them. 2. test_recall_config_env_overrides_config_yaml — writes both config.yaml and OPENVIKING_RECALL_* env vars, verifies env takes precedence. 3. test_recall_config_partial_config_yaml — partially populated config.yaml falls back to defaults for omitted keys. All 46 openviking_plugin tests pass (43 existing + 3 new). (cherry picked from commit b8d7834caf06c6912004333c270faa248eaed4cd) * fix(openviking): integrate reliability and configuration hardening * chore(contributors): map OpenViking source authors * test(retaindb): guard scoped secret config resolution * fix(openviking): verify servers before sending credentials * fix(openviking): catch endpoint errors in setup validation functions Review follow-up for salvaged PR #76782. Three setup-wizard validation functions called _normalize_openviking_url outside their try/except blocks. Since _normalize_openviking_url now raises _OpenVikingEndpointError for blocked or malformed endpoints, an invalid endpoint would crash the wizard instead of returning a friendly (False, message) tuple. - _validate_openviking_auth: move _normalize_openviking_url inside try - _validate_openviking_root_access: same - _validate_openviking_setup_values: catch _OpenVikingEndpointError explicitly - Remove dead ternary in _normalize_openviking_url safety check (candidate always has http/https scheme by that point) - Replace redundant float('-inf') < x < float('inf') with math.isfinite() in _setting_float; drop the redundant infinity check from _setting_int (is_integer() already rejects inf/nan) * fix(state): deduplicate session system prompts * chore: map cicav legacy noreply email * fix(tui): avoid writable Kanban opens on empty polls * fix(context): dedupe subdirectory hints by content digest and skip backup/vendor dirs SubdirectoryHintTracker re-injected identical context files whenever the same AGENTS.md was reachable through more than one path. Symlinked shared workspaces, hardlinks, and timestamped backup copies all alias a single file, so a normal session could ship the same 8KB of instructions two or three times. Nothing deduped it and nothing excluded directories that only ever hold copies. Two changes: * Track a sha256 of every injected hint body. Repeat content is skipped, and the working directory's own context file is seeded at construction so the copy prompt_builder already loaded at startup is never sent again. * Skip directories that hold copies rather than authoritative context (backups, node_modules, venv, site-packages, .git, .Trash, vendor, caches). Screening is relative to working_dir, so a project that legitimately lives under vendor/ keeps discovering its own subdirectory hints. Measured on a real session that touched a symlinked shared workspace: 3 injections / ~24,000 chars before, 1 injection / 8,112 chars after. 14 new tests cover symlink aliasing, byte-identical copies, working-dir seeding, distinct content still being injected, each excluded directory name, excluded ancestors, and the working-dir-inside-excluded-name case. * perf(state): batch the turn flush into one SQLite transaction Re-derivation of #23254 (@devsart95) on today's flush loop. The turn flush in _flush_messages_to_session_db wrote one BEGIN IMMEDIATE transaction per message row; a typical agent turn (user + assistant + tool results) paid 3-8 transactions -- and, off WAL (the default on macOS while the WAL-reset guard is active), 3-8 fsyncs -- per turn. Adds SessionDB.append_messages_batch: same row shape as append_message (shared _prepare_message_row serializer + _MESSAGE_INSERT_SQL column list, so the two writers cannot drift), same compression-lock and compression-closed guards, one aggregated session-counter UPDATE, one transaction for the whole batch. Row serialization stays outside the write lock. The flush loop now collects the turn's new rows and writes them in one call. All-or-nothing pairs exactly with the persisted-marker stamping: on failure no rows landed and no markers were stamped, so the next flush re-writes the whole tail (same recovery contract as before, minus the partial-prefix case that could double-count). Measured (same harness, 5-message turn, journal_mode=DELETE, synchronous=FULL): 2.32ms -> 0.83ms median per turn flush (64% faster, 5 fsyncs -> 1). On WAL the win is smaller but the atomicity fix holds. * perf(tui-gateway): batch branch-seed history copies (whole-bug-class) Sibling sites of the per-message flush pattern: both branch-seed paths (session.branch in methods_session.py and the lazy seed persist in server.py) copied the parent history row-by-row -- one transaction per row, and a branch seed can be hundreds of rows. Route both through SessionDB.append_messages_batch. The server.py path also gains real atomicity: _branch_seed_persisted assumed every row landed, which the per-row loop could not guarantee. * test(run-agent): update flush-path fakes and assertions for batched writes The flush now goes through append_messages_batch; MagicMock-based assertions and barrier fakes that hooked append_message observed nothing (the flush's try/except swallowed the AttributeError). Assert on the batch payload instead. * refactor(state): fold simplify findings — reuse _insert_message_rows, share guards, chunk seeds Simplify-pass folds on the #23254 salvage: - REUSE (HIGH): append_messages_batch now delegates row serialization to the pre-existing _insert_message_rows helper (already shared by replace_messages / archive_and_compact / portability import) instead of adding a third serialization path (_prepare_message_row + _MESSAGE_INSERT_SQL are gone). One row-writer for every multi-row path; the row-ID return was consumed by no production caller, so the batch returns the inserted count. - QUALITY (HIGH): the compression-lock + compression-closed admission guards are extracted into _check_transcript_write_guards, shared by append_message and append_messages_batch (previously duplicated 23 lines that had already needed targeted fixes, #74478). The role-gated reasoning filtering is no longer duplicated in run_agent.py — it lives at its one site inside _insert_message_rows. - EFFICIENCY (MEDIUM, measured): unbounded seed copies hold one BEGIN IMMEDIATE for seconds (10k rows ~= 2.4s; FTS triggers dominate) and monopolize the in-process write lock. append_messages_batch grows a chunk_rows param; all seed/copy call sites use chunk_rows=500. Same recovery semantics as the old per-row loops, bounded lock holds. - REUSE (MEDIUM): the two remaining per-row branch-copy loops found by the pass (gateway/slash_commands.py /branch, hermes_cli cli_commands_mixin.py branch) are converted to chunked batches too (AsyncSessionDB's generic to_thread forwarder covers the async site). Turn-flush benchmark unchanged after the refactor: 2.43 -> 0.87 ms median per 5-message flush (64% faster). * fix(tests): update two more append_message.call_args assertions to append_messages_batch CI-caught: test_verification_stop_caching and test_tui_gateway_server::test_native_vision_turn_persists_a_renderable_image_ref both assert on append_message.call_args, but the flush loop now calls append_messages_batch. Same class of test-fake fallout fixed in 5 other files — these two were missed. * perf(tui): memoize useSessionLifecycle return (idea from #38491) Re-derivation of #38491 by @stremtec onto current main (the original is 10,119 commits behind; the hook moved into ui-tui/src/app/). The hook returned a fresh object literal every render, defeating memoization in useMainApp's consumers; useMemo over the (all-useCallback-stable) handles makes the return referentially stable. Dep array covers ALL nine returned handles incl. trimTail (the re-derivation initially omitted it - stale-closure class). * ci: retry uv python install * fix(state): route session-resume reads through the WAL read-only connection get_messages_as_conversation, get_resume_conversations, and get_ancestor_display_prefix still took self._lock — the same global choke point the read-path split (WAL per-thread read-only connections) was meant to remove from every recall/browse read. These three are the hottest reads in the file: every session resume across the gateway, CLI, and ACP adapter goes through one of them, so a resume racing a burst of concurrent-session writer flushes still convoys behind them exactly like the fixed paths used to. _session_lineage_root_to_tip (the lineage walk shared by all three, plus get_conversation_root) had its own independent self._lock use and needed the same conversion — without it the outer functions still blocked on the very first line. Verified empirically: a reader thread calling all three functions while another thread holds self._lock blocked for the writer's full hold duration before the fix, and returned immediately after (SQLite 3.50.4 in this dev venv falls back to journal_mode=DELETE per the WAL-reset-bug guard, so the requires_wal-marked regression test is exercised via a local WAL-forced script instead; it still runs and passes on any runtime where WAL is actually active). * chore: add contributor email mapping for ArcherQAQ * fix(model_metadata): rewrite localhost->IPv4 for the remaining local probe sites fetch_endpoint_model_metadata's generic (non-LM-Studio) /models fetch and its llama.cpp /v1/props context-length follow-up built request URLs straight from the unrewritten candidate, unlike every other local-probe site. Both retained the multi-second dual-stack IPv6 connect penalty that _localhost_to_ipv4() exists to skip (measured on macOS: localhost 32.9ms vs 127.0.0.1 0.1ms on a dead port; ~2s on Windows). normalized stays the cache key so caching behavior is unchanged; only the outbound request target is rewritten. Re-derived from PR #61528 onto current main (original no longer applied cleanly). * fix(model_metadata): guard _localhost_to_ipv4 against non-string urls CI slice 3/7 failures: run_conversation tests pass MagicMock base_urls through the metadata probe path; re.sub raised TypeError where the old code let non-strings flow through. Preserve that contract. * perf(cold-start): mitigate ~14s GIL stall during backend init (#60800) Three fixes for the Desktop/TUI cold-start stall where the event loop is blocked for ~14s between HERMES_BACKEND_READY and the first prompt (#60800): 1. copilot_auth: skip subprocess fallback when any Copilot env var is explicitly set (even if invalid). The user expressed token intent via env var; silently substituting a CLI token is surprising and the subprocess adds up to 5s on Windows. 2. tui_gateway/ws: run resolve_skin() via asyncio.to_thread so config loading + skin engine init do not block the WS read loop during the cold-start RPC burst. 3. web_server: extend _warm_gateway_module to pre-import the heavy module chains (auth, copilot_auth, runtime_provider, skin_engine, inventory, model_switch) that the first WS connection + RPC burst would otherwise import on the loop thread. These trigger .pyc compilation and Defender scans on Windows (15-30s per the existing comment) and were not covered by the original gateway-only warm. Tests: 5 new tests in test_cold_start_gil_stall.py + 2 new tests in test_copilot_auth.py. All 36 copilot_auth tests + 16 ws/web_server tests pass. * test: harden cold-start regression tests + debug-log the env-var skip Review folds on the #60807 salvage: - resolve_skin tests are behavioral (thread-ident probe + ready-frame wiring check) instead of pure source inspection, per the #72720 pattern; a source assertion remains as belt-and-braces. - The warm-list test does REAL imports and checks sys.modules — _warm_gateway_module swallows ImportError by design, so the PR's tracking-stub test would pass even with a typo'd module name. - resolve_copilot_token logs a debug line when the env-var short-circuit skips the gh-CLI fallback (behavioral change made observable). * perf(gateway): per-platform skip_context_files to cut agent build latency Salvage of #26860 (hunk 2, ported \u2014 the PR's base predates the current gateway layout by ~11.9K commits). Messaging platforms can set gateway.platforms.<key>.skip_context_files: true to skip the filesystem-heavy context-file discovery (SOUL.md, AGENTS.md, .cursorrules walks) during AIAgent construction \u2014 10-100x slower stat()/walk costs on Windows made this a real per-turn tax. Soul identity is still loaded (single small file), so the persona survives. The flag participates in _agent_config_signature so toggling it rebuilds the cached agent instead of silently reusing a prompt built under the other setting (prompt-cache correctness). The PR's hunk 1 (mtime-caching the per-turn dotenv reload) was dropped: df51ad797 mtime-cached load_config/read_raw_config and c2eda92fd removed the per-turn deepcopies, capturing most of that win; the function has since gained a multiplex early-return and managed-scope overlay that the original whole-function skip would have bypassed. * fix(relay): route Discord tool-progress into the auto-thread, not the parent channel (#77830) When a Discord channel message initiates a relay auto-thread, the thread does not exist at ingest (source.thread_id is None) — the connector creates it on its FIRST send and auto-threads any outbound carrying the reply anchor. The final reply carries that anchor, so it lands in the thread. But the tool-progress / status bubbles (the "Searching the web for..." updates and the streaming preamble) were sent with _progress_metadata=None and _progress_reply_to=None: _resolve_progress_thread_id returns None for Discord (only slack/mattermost get a synthetic thread), so the progress send had no anchor and the connector posted it FLAT in the parent channel. Result: the search-status updates leaked outside the thread while the answer threaded (staging repro 2026-08-02). The connector now stamps prospective_thread_id on the inbound (the anchor message id == the id of the thread it will create). Reuse it: when a relay-delivered Discord channel-initiate carries prospective_thread_id and has no real thread yet, carry the reply anchor (event_message_id) on both the progress metadata (reply_to_message_id) and the progress reply_to, so the connector routes the progress bubble into the SAME auto-thread as the final reply. Applied to both the tool-progress path (_progress_metadata / _progress_reply_to) and the status/interim callback path (_status_thread_metadata). Events already arriving in a real thread, DMs, and non-relay sources are untouched (guarded on delivered_via_upstream_relay + prospective_thread_id + not thread_id). Tests: two new cases in test_run_progress_topics.py — a relay Discord channel-initiate asserts every progress send carries the anchor (reply_to + metadata.reply_to_message_id + non_conversational), and an event already in a real thread asserts the synthetic-anchor path does NOT engage. Full gateway progress + relay + session suites green (228 passed). * fix(agent): stop re-probing endpoints that blackhole TCP connects Salvage of #71282 (Fixes #71281): a routable-but-dead endpoint (corp LAN address while off-VPN) blackholes TCP SYNs, so every probe in the model-metadata waterfall waits out its full connect timeout — 20+ seconds of stall per startup across detect_local_server_type, fetch_endpoint_model_metadata, and the per-model probes. A module-level blackhole cache keyed on host:port is populated when any probe observes a ConnectTimeout (httpx or requests; read timeouts deliberately excluded — an accepted connection is not a blackhole) and consulted at the top of each guarded function. 30s TTL: long enough to collapse one startup burst, short enough that VPN recovery is picked up without a restart. Guard ordering: blackhole check -> disk L2 -> HTTP waterfall, and a blackholed leg aborts the remaining legs instead of letting each stall in turn. Squash of the PR's two real commits (the branch's merge commits made it un-rebase-merge-able; content verified identical via merge-tree). * chore: release v0.20.0 (2026.8.3) The Herald Release — voice (streaming TTS, barge-in, wake words), A2A v1.0, outbound webhooks, grounded citations, desktop platform wave. ~3,650 commits, ~1,400 PRs, ~1,200 issues closed, 650+ contributors since v0.19.0. Also: contributor audit additions (18 email mappings, bot-filter widening). * chore: add contributor email mapping for Ahmett101 * perf(moa): cache resolved preset + per-slot runtime to cut cold-start latency (#66793) * fix(discord): leave voice channels before cancelling the bot task `DiscordAdapter.disconnect()` cancelled the bot task before tearing down voice clients. `leave_voice_channel()` ends in `await vc.disconnect()`, and discord.py sends a voice state update over the main gateway websocket and then waits for the voice socket to close. The bot task is the loop running that gateway connection, so cancelling it first left the handshake with no transport: it could never complete and blocked until the caller's shutdown timeout fired. The effect was a fixed ~5s penalty on every shutdown with a voice connection open, ending in "discord disconnect timed out after 5.0s - forcing continue", with the voice disconnect abandoned rather than completed. Measured on a live gateway with a voice connection open in both cases: before: timed out after 5.0s, all adapters disconnected at +5.29s after: discord disconnected (0.12s), all adapters disconnected at +0.46s Moving the voice-cleanup loop above `_cancel_bot_task()` preserves the zombie-client protection its comment describes: the bot task is still cancelled before `client.close()`, just after voice teardown rather than before it. Voice teardown is the one step that still requires a live gateway. Adds a regression test asserting the ordering. It fails on the previous ordering at index 1 with `cancel_bot_task != leave_voice_channel:111`. Fixes #76044 * feat(image): parallelize image_generate batches * fix(file-sync): serialize concurrent sync cycles * fix(tool-executor): unpack 5-tuple runnable_calls in _max_workers_for_tool_batch * fix: exponential backoff for rate-limit fallback cooldown Replace the fixed 60-second cooldown with exponential backoff: 30min → 1h → 2h → 4h cap. The counter is reset by restore_primary_runtime on successful primary-provider recovery, so the backoff is strictly for consecutive failures within a single degradation window. Closes #29702 * fix(backoff): keep 60s first-hit cooldown, escalate only on consecutive rate-limits Review follow-up on the #30223 salvage: the original changed the base cooldown from 60s to 1800s, benching the primary for 30 minutes on the FIRST 429 (30x regression in primary-restore latency) and breaking the existing test_rate_limit_exhaustion_keeps_60s_cooldown contract. Keep upstream's 60s base and escalate per consecutive rate-limit: 60s -> 2m -> 4m -> 8m -> ... capped at 4h. Counter still resets on successful primary restore (cicae's mechanism, unchanged). New tests: escalation doubling, 14400s cap, reset-on-restore. Existing 60s contract test passes UNCHANGED. Mutation-checked: escalation disabled -> 2 fail; reset disabled -> 1 fails. * fix(catalog): wire api_key auth headers for http MCP servers When an optional-mcps manifest declares transport.type=http with auth.type=api_key, install_entry() prompts for the key and saves it to .env, but _build_server_config() only handled the oauth case — the api_key case produced a bare url entry with no headers, so every request to the server was unauthenticated (-> 401). Reuse _bearer_auth_headers(entry.name) from mcp_config.py so the catalog path emits the same 'Authorization: Bearer ${MCP_..._API_KEY}' template as the manual 'hermes mcp add --url' path. Salvaged from #70782 (production hunk applied clean; tests re-anchored onto current main). Credit: JonthanaHanh. * perf(compressor): release allocator pages after successful compaction A successful compaction frees the largest allocation a long session ever drops (the compressed-away message dicts), but Python's arena allocator keeps those pages in the heap — RSS retains the pre-compaction high-water mark until exit. #76905's trim_memory lifecycle covers the gateway/TUI housekeeping loops but not the CLI compression path. Call trim_memory(reason='post-compression') at the compression-success point in ContextCompressor.compress(), following the house pattern (lazy import in try, debug-level log on failure). The helper is glibc-gated, config-gated and rate-limited, so it is a safe no-op on other platforms and cannot fail compression. Re-expresses the intent of #70782 (JonthanaHanh), which reached for a bare gc.collect(); trim_memory is the house mechanism and already wraps a collect. * fix(catalog): validate http+api_key manifests declare the header's env key Simplify-pass follow-up on the #70782 salvage: _bearer_auth_headers hard-emits ${MCP_<NAME>_API_KEY} but install_entry only persists auth.env-declared vars — a manifest naming its key differently (the shipped n8n style) would install cleanly yet send a literal-placeholder header at connect time (silent 401, the #37792 bug class). Enforce the naming contract at parse time. Also pins the secret-stays-in-.env property in the install test (raw config.yaml carries the template, never the secret). Mutation-checked: validation disabled -> guard test fails. * fix(agent): cap auxiliary LLM concurrency per task * fix: thread extra_headers through the call_llm split The PR's concurrency wrapper splits call_llm into a semaphore-guarded entry + _call_llm_impl; main added extra_headers to call_llm's signature after the PR's base, so the split has to forward it too (dropped silently otherwise — Azure Foundry and custom-endpoint callers set it). * fix(nix): tie devShell's HERMES_PYTHON to the venv actually on PATH `nix/devShell.nix` collected `devShellHook` by scanning every package: nonNpmHooks = map (p: p.passthru.devShellHook or "") packages; But `minimal` and `messaging` are `.override` variants of `default`, so each carries its own `devShellHook` exporting its own HERMES_PYTHON. The scan therefore concatenated three conflicting exports and forced Nix to evaluate and realise three separate uv2nix editable venvs on every `nix develop`. `attrValues` is alphabetical, so the last hook won (`minimal`) while `python`/VIRTUAL_ENV came from `default`'s devDeps: HERMES_PYTHON = ...dimim2... (minimal — no optional deps) python / VIRTUAL_ENV = ...85r28... (full) Inside the shell `$HERMES_PYTHON -c "import anthropic"` failed while `python -c "import anthropic"` succeeded. Worse, `scripts/run_tests.sh` prefers HERMES_PYTHON, so the suite ran against the minimal venv. Its guard did not catch this: it only checks that HERMES_PYTHON has pytest, and minimal's venv does (pytest is in the `dev` group), so the wrong interpreter was silently accepted. Tying the hook to `packages.default` — the same package whose `devDeps` are installed — keeps HERMES_PYTHON, `python`, and VIRTUAL_ENV pointing at one venv by construction. editable venvs referenced 3 -> 1 their combined closure 421 MB -> 140 MB test failures 85 -> 32 The venv mismatch was masking 53 failures; e.g. test_web_tools_config.py goes 2-failed -> 38-passed. Full suite is now 25369 passed / 32 failed, and those 32 reproduce identically on a pristine HEAD worktree with no nix/ changes under the same interpreter (mostly NixOS artifacts — tests spawning bare `python3` in a scrubbed env exit 127). * fix(system_prompt): move skills index to the volatile band The skills index is runtime-mutable: the agent adds and patches skills mid-session, so it is not byte-stable. Keeping it in the stable band breaks that band prefix-cache contract, because every skill edit changes the stable band and invalidates the entire cached prefix in front of it. Move it to the front of the volatile band so the stable scaffold (identity, tool guidance, model guidance) stays cacheable across skill edits. * docs(system_prompt): fix stale reconstruct_static_prefix docstring example Simplify-pass finding: the safety note still cited 'skills edited' as a stable-tier input whose change mismatches the rebuilt prefix — after this PR a skill edit changes only the volatile tail (that's the point). Swap the example for genuinely stable-tier inputs. * fix(prompt_size): search volatile tier for skills block after the stable->volatile move CI-caught: compute_prompt_breakdown still looked for <available_skills> in the stable tier, but #37117 moved it to volatile. Search volatile first, fall back to stable for older sessions. * chore: add contributor email mapping for zabih-sudo * chore: add contributor email mapping for HAOWANG116 * fix(backup): serialize and atomically publish snapshots * chore: add contributor email mapping for ElSnacko * fix: prefer explicit anthropic api key Cherry-picked from PR #58560 by @itsflownium, adapted to current main (_getenv instead of os.getenv). Moves ANTHROPIC_API_KEY check ahead of Claude Code credential file and credential_pool auto-discovery so an explicitly configured key is never shadowed by auto-discovered OAuth. Fixes #58546 * docs: document /personality none|default|neutral reset across personality docs The reset keywords have existed in both CLI and gateway handlers since June but were undocumented — users couldn't find how to cancel a personality overlay. Adds a 'Resetting to the default' section to the personality feature page and mentions the reset in the CLI guide, slash-command reference (both tables), and messaging command table. * fix(relay): avoid concurrent turn scope corruption Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com> * fix(relay): preserve legacy turn shims Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com> * fix(relay): gate skipped turn metrics Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com> * test(relay): enforce LIFO in overlap regression Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com> * fix(relay): preserve skipped turn context Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com> * fix comment about relay workaround * feat(models): add qwen3.8-max to Nous portal + OpenRouter catalogs, replacing qwen3.7-max Qwen3.8 Max is live on both OpenRouter and the Nous portal (qwen/qwen3.8-max, 1M context, 131K max output). Per the newest-max-replaces-last-max convention, it takes qwen3.7-max's slot in both curated lists. - hermes_cli/models.py: OPENROUTER_MODELS + _PROVIDER_MODELS[nous] swap qwen/qwen3.7-max -> qwen/qwen3.8-max - agent/model_metadata.py: DEFAULT_CONTEXT_LENGTHS entry for qwen3.8-max at 1,000,000 (verified against OpenRouter live metadata and Nous /v1/models 2026-08-03) - tests/test_empty_model_fallback.py: swap incidental catalog fixture to the surviving slug - website/static/api/model-catalog.json: regenerated Pricing snapshot skipped: both routes bill via official_models_api (live pricing), verified with resolve_billing_route. Reasoning timeout floor already covered by the qwen3 prefix (180s). * test: swap context-switch-guard fixture off qwen3.8-max-preview test_custom_provider_context_avoids_false_shrink_warning used qwen3.8-max-preview as a slug that deliberately falls through to the generic 'qwen' 131K catalog match. The new qwen3.8-max DEFAULT_CONTEXT_LENGTHS entry (1M) now substring-matches the preview slug too, so the no-custom-providers branch stopped warning. Swap the fixture to qwen3.9-max-preview, which still hits the generic fallback — the test's intent (custom_providers threading) is unchanged. * fix(agent): keep context_length pin for named custom providers Empty model.base_url plus a runtime custom-provider URL was treated as a route mismatch, so gateway session-reset banners dropped model.context_length and fell back to the Qwen family default (131K) while /status still showed the configured 262K pin. * test(gateway): cover named-custom context pin on session-info banner * fix(model_metadata): read llama.cpp context from meta.n_ctx + accept sole model * chore: add contributor email mapping for johnrazmus * fix(gateway): keep event loop alive during /compress and Relay drain Offload manual /compress temporary-agent cleanup through the existing bounded off-loop helper so a slow agent.close() cannot freeze the gateway event loop, heartbeat, or platform polling. Guarantee Relay transport teardown even when the runner cancels adapter.disconnect() during go_idle: shielded finally, 2s drain-path idle ACK budget under the 5s outer disconnect budget, and bounded supervisor/reader/ws.close awaits. Original commits: - fix(gateway): offload manual /compress cleanup from the event loop - fix(gateway): tear down Relay transport even if go_idle is cancelled - fix(gateway): keep Relay disconnect budgets inside the runner window By @Dannyzen (PR #78027), salvaged onto current main. * fix(gateway): bound go_dormant ws.close with teardown timeout Sibling site to the disconnect() fix: go_dormant() still did an unbounded await self._ws.close(), the exact same pattern bounded in disconnect(). go_dormant runs on the scale-to-zero suspend path (Fly autostop), which also has timeout constraints. Apply the same 1s wait_for treatment using _TEARDOWN_AWAIT_TIMEOUT_S. Found during review of PR #78027. * fix: close the Codex app-server session on agent teardown Salvage of #65260's b7d7cfd0e (ported — the PR's close() predates ~4K commits of teardown-step churn, so the hunk is re-anchored after step 6b rather than cherry-picked). agent/codex_runtime.py already drops _codex_session on turn crash and on retirement, but AIAgent.close() — the hard teardown for /new, /reset, and session expiry — had no owner for it, so the app-server child process survived until interpreter exit. Long-lived gateways accumulate one leaked subprocess per ended Codex session. The attribute is cleared BEFORE close() so a concurrent reader can't observe a half-closed session and a raising close() can't strand a stale reference (tested). Tests extend the author's original lifecycle test with the raising-close and no-codex-session cases. * fix(agent): discard bare tool-call marker before fallback/persistence (#78148) Local tool-call templates can emit a bare bracketed token (e.g. "[memory]") as assistant content alongside a function call. The loop treated that protocol scaffolding as visible content: it got cached as the post-tool fallback, and when the next turn came back empty, the marker was replayed as the final response and written into the persisted transcript. Later context compaction preserved that history, letting the model repeat the marker in subsequent turns. Detect content that is only a bracketed marker (`[name]`) when the response also carries tool_calls, and drop it before it can be cached or persisted. Scoped narrowly: only fires alongside tool_calls, so a genuine final response of "[memory]" without a tool call is unaffected. * fix(agent): repair sessions already contaminated with stale tool-call markers (#78148) The conversation_loop fix (previous commit) stops new "[memory]"-style bare tool-call markers from being cached/persisted, but sessions written before that fix can still carry rows where a bare marker was saved as the assistant's "final response". Add a load-on-read repair pass in hermes_state.py, mirroring the existing _strip_background_review_harness defense-in-depth: on session restore, any assistant row whose content is only a bracketed marker (e.g. "[memory]", "[skill_manage]") AND that carries tool_calls has its content blanked before the history re-enters the model's context. The tool call and its result are left untouched so provider tool_call/tool_result pairing stays intact. Sessions with no affected rows pass through the normal path unchanged. * feat(cli): add sessions clean-markers to permanently purge stale tool-call markers (#78148) The load-on-read repair (_strip_stale_tool_call_markers) fixes affected sessions in memory on every resume, but never touches the DB — long-lived sessions re-scan and re-repair the same rows on every load, and the contaminated bytes stay in state.db (and any backup/cache snapshot of it) indefinitely. Add SessionDB.purge_stale_tool_call_markers(dry_run=False): a one-time, idempotent UPDATE that permanently blanks the content column on affected rows. Only content is touched — tool_calls and every other column are left untouched, so provider tool_call/tool_result pairing survives. dry_run reads through the no-lock read path and never writes. Wire it up as `hermes sessions clean-markers [--dry-run]`, mirroring the existing optimize/repair subcommands. Verified end-to-end against a real temp state.db: dry-run reports the row without writing, the real run clears it and preserves tool_calls, and a second run is a no-op. * fix(cli): back up state.db before clean-markers writes by default purge_stale_tool_call_markers ran a permanent, irreversible UPDATE with no backup — inconsistent with repair_state_db_schema's backup-by-default convention for destructive state.db operations elsewhere in this file. Take a full snapshot via VACUUM INTO (safe against a live connection, unlike the raw-copy _backup_db_file used for malformed-schema repair) before the write, timestamped beside state.db. Skipped when dry_run or when there's nothing to change. Add --no-backup to `hermes sessions clean-markers`, mirroring `sessions repair`. Verified end-to-end: the CLI run against a real temp state.db produces the backup file before printing the cleared-row count. * refactor: dedup stale-marker regex — use compiled _STALE_MARKER_RE in conversation_loop The bracketed-marker regex was inlined in conversation_loop.py as re.fullmatch(r"\[...", ...) while hermes_state.py defines the same pattern as _STALE_TOOL_CALL_MARKER_RE. Both must agree on what counts as a stale marker — a drift here means the runtime guard silently disagrees with the load-on-read repair and CLI purge in hermes_state. Consolidate onto a single compiled constant (_STALE_MARKER_RE) at module level in conversation_loop.py, with a comment noting it must mirror _STALE_TOOL_CALL_MARKER_RE in hermes_state.py. A direct import from hermes_state was tried first but caused a regression: hermes_state initializes DEFAULT_DB_PATH = get_hermes_home() / 'state.db' at module import time, which breaks tests that monkeypatch get_hermes_home() to return a str (test_slash_worker_accepts_profile_home). Follow-up to PR #78175 (@JoaoMarcos44). * fix(conversation_loop): compress messages on output-cap retry path (#55546) The output-cap retry loop reduced max_tokens by 64 tokens per attempt but never called _compress_context(), so the compressor never fired. Input growth (~65 tokens/attempt) canceled the savings, leaving the session stuck at 200,001 tokens — 1 over the 200,000 ceiling. The fix adds compression to the output-cap retry path. The compressor drops the middle window, freeing ~50% of tokens. If compression makes >=5% savings, the session continues; otherwise vision payloads are stripped or the session ends with compression_exhausted=True. Also adds CHANGELOG.md entry and bug fix report. * fix(conversation_loop): prune dead vision-strip fallback; harden output-cap retry tests * chore: drop CHANGELOG.md and docs/reports/ — not shipped with salvage PRs * perf: reuse request_input_estimate instead of recomputing estimate_request_tokens_rough The output-cap error handler already computes request_input_estimate at line 4722 via estimate_request_tokens_rough(api_messages, tools=...). The new compression block ~50 lines below was calling the same function with the same inputs again. Reuse the existing local. * chore: AUTHOR_MAP — add BobClawblaw for PR #77870 salvage Bare noreply email (no NNN+ prefix) needs explicit mapping. * chore(ci): rerun checks Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com> * perf(gateway): prewarm /model picker cache on TUI startup The classic CLI run() loop calls prewarm_picker_cache_async() during the idle window after the banner is shown, so the first /model open hits a warm provider-models disk cache and renders in ~100ms. The stdio TUI entry point never did this, so the first /model open in a TUI session blocked on serial /v1/models fetches for every authenticated provider. Mirror the CLI behaviour: kick off the same off-thread prewarm right after gateway.ready is emitted (banner shown, user about to type). Fire-and-forget, guarded once-per-process, fully exception-isolated so a slow or offline provider can never affect TUI startup. * test(tui): pin picker-cache prewarm wiring in entry.main() teknium's review gap on #72021: the helper's worker/once-guard was covered, but nothing asserted the stdio TUI entry point actually invokes prewarm_picker_cache_async() — or that it does so in the right place. Add a focused entrypoint test that runs the real entry.main() with stubbed collaborators (same monkeypatch-module-attrs harness as test_tui_entry_mcp_owner.py), spies on the helper in hermes_cli.model_switch (the lazy-import source), and asserts: - prewarm fires exactly once, strictly AFTER the gateway.ready write - startup stays non-blocking: main() reaches the stdin loop and returns on EOF - a prewarm failure is swallowed (fire-and-forget) without breaking startup Mutation-checked: deleting the prewarm hunk from entry.py fails both tests. * perf(cli): check local auth.json/config before slow provider registry sweep _has_any_provider_configured() probed every api_key provider (gh subprocess for copilot alone takes 5s; full sweep ~18s) before consulting auth.json and config.yaml, which are instant local reads. Desktop setup.status calls blocked past the UI's timeout, causing the connect/disconnect boot loop. Reorder so cheap local checks run first. Same semantics, ~35x faster here. * test(cli): regression tests pinning auth-first ordering skips registry sweep Teknium's review on #63457: existing tests pin the final boolean but not that the slow PROVIDER_REGISTRY sweep is skipped. Add three tests that booby-trap hermes_cli.auth.get_auth_status and verify _has_any_provider_configured() short-circuits on: - config.yaml model.provider - config.yaml base_url/api_key (custom endpoint shape) - auth.json active_provider (sweep-only call-pattern guard) Mutation-checked: reverting the reorder makes all three fail. * perf(desktop): keep spinner frames out of React commits Advance the existing animated status glyph through its DOM text node instead of React state, and pause its timer for hidden panes or inactive windows. Cover frame advancement, zero update-phase commits, and timer suspension with behavior tests. * test(desktop): cover minimized/hidden window-state + visibilitychange pause for GlyphSpinner Regression coverage requested in review of #74357: mock window.hermesDesktop.onWindowStateChanged (pattern from persistent.test.tsx) and assert minimized/hidden clears the spinner interval while restore resumes it; also cover document.visibilityState hidden/visible via visibilitychange. * fmt(js): `npm run fix` on merge (#78271) Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> * perf(gateway): replace SSE poll loop with call_soon_threadsafe-fed asyncio.Queue _write_sse_chat_completion and _write_sse_responses bridged their stream_delta_callback queue into the event loop via `await loop.run_in_executor(None, lambda: stream_q.get(timeout=0.5))` in a while-True poll — a thread-pool round trip on every 0.5s tick even when idle, plus up to 500ms of tail latency between a delta landing in the queue and it reaching the SSE response. Add ThreadSafeAsyncQueue (asyncio.Queue + a put_threadsafe() that wraps call_soon_threadsafe), used by both streaming producer closures (_on_delta, tool start/complete callbacks — all invoked from the worker thread running run_conversation via loop.run_in_executor). Consumers now do a plain `await asyncio.wait_for(stream_q.get(), timeout=0.5)` — woken immediately when a delta arrives, no executor hop, no poll interval. Updated tests/gateway/test_sse_agent_cancel.py's 7 call sites to construct ThreadSafeAsyncQueue inside the running loop (required, since it captures asyncio.get_running_loop() at construction) instead of a bare queue.Queue() at test-method scope. * refactor(gateway): extract _sse_frame() helper, dedup 5 inline SSE encode call sites _write_sse_chat_completion had five near-identical f"data: {json.dumps(...)}\n\n".encode() (and one event-tagged variant) scattered across its role/content/finish/error chunk writes. Pure extract-method, no behavior change: encoding is byte-identical for every call site touched. Left the pre-serialized-string writers elsewhere (_write_event's json.dumps(..., ensure_ascii=False) path, the /v1/runs SSE writer) alone — routing them through this helper's plain json.dumps(data) would silently change their unicode-escaping behavior, which is out of scope for a pure dedup. * refactor(gateway): route all three SSE writers through _sse_frame() _extend _sse_frame with an explicit ensure_ascii param (default True, byte-identical to a bare json.dumps) and route the two sibling writers through it: _write_sse_responses._write_event and the /v1/runs event stream. This completes the dedup PR #65009 — previously only the five _write_sse_chat_completion sites used the helper, leaving the other two writers on inline json.dumps with no shared shape. No behavior change: every writer's emitted bytes are unchanged (verified byte-for-byte, including non-ASCII payloads where the default ensure_ascii=True matches the original inline encoders). The ensure_ascii option is exposed so a future writer can opt into raw non-ASCII bytes without fractalizing the format again. Adds tests/gateway/test_sse_frame.py asserting the byte-contract invariant between _sse_frame and the historical inline encoders. * refactor(gateway): route session event stream through _sse_frame (ensure_ascii=False) The session event stream (api_server.py:~2236) was the one genuinely unicode-distinct SSE writer — json.dumps(payload, ensure_ascii=False) + .encode('utf-8'). Every other writer uses plain json.dumps. Route it through _sse_frame(..., ensure_ascii=False) so _sse_frame is now the single source of truth for ALL SSE frame serialization in the module (chat- completion, responses._write_event, /v1/runs, and the session stream). Byte-identical for non-ASCII payloads: verified against the historical inline encoder (raw bytes preserved). The ensure_ascii=False path is now exercised by test_sse_frame_ensure_ascii_false_reproduces_session_event_stream. * test: add cross-thread put_threadsafe + long-reasoning tail tests Addresses teknium1 sweeper review (2026-07-30) requiring coverage of: 1. ThreadSafeAsyncQueue.put_threadsafe() off-loop boundary: a real daemon thread pushes into the queue from outside the owning event loop while the consumer awaits get(), mirroring the run_conversation worker-thread producer path. Includes a 20-concurrent-thread no-drop regression. 2. Long-reasoning bound stability for thinkingPreview: 100k-char input plus empty/collapsed cases must not crash and must retain the visible tail marker inside the bounded 24k clean window. * fix: reconstruct fused test after conflict resolution The conflict-marker strip fused test_agent_task_raises with the body of test_failed_result_dict — restore both as separate tests (content from the PR head, verified verbatim). * perf(tui): bound reasoning-clean input to the displayed tail CI caught a split defect in this salvage: the PR's long-reasoning tail test was kept but its production hunk was dropped as 'cosmetics'. It isn't — cleanThinkingText runs several full-string regex passes and reasoning grows on every streamed token, so re-cleaning the whole accumulated string per chunk is O(n) per token / O(n^2) per stream. Only the tail is displayed (boundedLiveRenderText caps it downstream), so bound the input to 1.5x LIVE_RENDER_MAX_CHARS first. Restores the one text.ts hunk from cd99e65fc (author preserved); the italic-thinking display change and profile script from that commit remain out of scope. * test: exercise the production _loop_ref path in put_threadsafe tests Gate finding (/simplify-code pass): both cross-thread tests passed loop=loop explicitly, but no production caller does — all six (_on_delta, _on_tool_*) rely on the queue resolving its own _loop_ref in __init__. The kwarg made the tests vacuous: a broken _loop_ref still passed them. Dropping the kwarg exercises the real path. Verified by mutation: with self._loop_ref = asyncio.new_event_loop() (wrong loop), both tests now FAIL; they passed before this change. * fix(test): feed the SSE writers an asyncio queue, not queue.Queue CI caught a missed caller-shape update. Both PRODUCTION callers of _write_sse_chat_completion / _write_sse_responses were converted to ThreadSafeAsyncQueue, but two pre-existing tests in tests/gateway/test_api_server.py construct the writer's queue themselves and still passed a stdlib queue.Queue. The consumer now does 'await asyncio.wait_for(stream_q.get(), ...)', which on a queue.Queue blocks the thread forever: test_stream_cancelled_persists_incomplete_snapshot hung until pytest-timeout killed it (CI reported the whole file as 'no tests ran (timeout before collection)'). The sibling disconnect test only survived because it pre-fills before the first await. tests/gateway/test_api_server.py: 99 passed (was 1 failed + a 60s hang); with the SSE/api_server suites: 147 passed. * fix(credential-pool): re-select in acquire_lease after a deferred refresh select() re-selects once deferred single-use-token refreshes complete; acquire_lease() performed the refresh but returned its pre-refresh answer. Since _acquire_lease_under_lock returns early exactly when a refresh is pending (if not available: return None, pending_refresh), a pool whose entries all needed a refresh always returned None — the caller failed an answerable request right after the refresh succeeded. Retry once, only when the first pass was empty and a refresh ran. Post-merge gate-sweep finding on the #71775 salvage (#77714). * fix(credential-pool): lock the quarantine read-modify-write of _entries #71775 moved deferred single-use-token refreshes outside the pool lock (correct — they hold a cross-process flock plus network I/O). But _refresh_entry_impl's three terminal-auth-failure quarantine paths do a bare read-modify-write of self._entries. Those used to run with the caller holding self._lock; on the deferred path they run unlocked, so a concurrent mutation between the read and the write is silently lost. Wrap all three in 'with self._lock' (an RLock, so locked callers re-enter safely) and correct the _refresh_pending_entries docstring, which claimed the mutations were already self-locking. Post-merge gate-sweep finding on the #71775 salvage (#77714). Sibling to the acquire_lease re-select fix. * perf(file-ops): eliminate redundant subprocess calls in write_file and V4A patch path write_file currently spawns up to 6 subprocesses per call: 1. mkdir -p (separate call before atomic write) 2. cat (to read pre-content for lint/BOM/line-ending detection) 3. _atomic_write (mktemp + write + mv — the essential one) 4. wc -c (to measure bytes written) 5. _check_lint_delta (post-write lint — also essential) 6. LSP snapshot (also essential) This PR removes three of them without changing any observable behavior: 1. Fold mkdir -p into _atomic_write shell script (−1 subprocess/write) The atomic write script already runs a single shell; adding mkdir -p to it costs zero extra processes. 2. Add optional pre_content parameter to write_file (−1 subprocess/patch) patch_replace and V4A _apply_update already read the file for fuzzy matching. Passing that content as pre_content skips the redundant cat inside write_file. Fully backward-compatible: callers that don't pass pre_content still read from disk as before. 3. Replace wc -c with len(content.encode('utf-8')) (−1 subprocess/write) We already have the content in memory; encoding it to get the byte count is equivalent to wc -c for UTF-8 text. 4. Remove redundant _check_lint loop in apply_v4a_operations (−N subprocesses/V4A) write_file already runs _check_lint_delta internally. The old code ran a bare _check_lint(f) loop over all modified files — a re-read + re-lint without post_content context. Now lint results propagate from write_file via a four-tuple return, zeroing out the extra subprocesses. Net effect: - write_file: 6 → 3 subprocesses per call (new files) - patch_replace: 6 → 5 subprocesses per call (pre_content skips cat) - V4A multi-file patches: saves 1 subprocess per modified file - A typical 4-file V4A patch drops from ~28 to ~16 subprocess calls * fix(file-ops): decouple BOM detection from pre_content, add V4A backward compat Bug 1 (UTF-8 BOM loss on V4A UPDATE): _file_has_bom() trusted pre_content for BOM detection, but the most common pre_content provider — read_file_raw() — deliberately strips BOMs so the agent never sees U+FEFF glyphs. Passing BOM-stripped content through pre_content caused a false-negative: the method returned False and write_file() silently removed the marker on rewrite. Fix: _file_has_bom() now always probes the first 3 bytes on disk (head -c 3), ignoring pre_content for BOM purposes. pre_content is still used by two other consumers — line-ending detection and lint/LSP delta computation — neither of which is affected by BOM stripping. Bug 2 (backward compatibility): _apply_update() called write_file(path, content, pre_content=...) as a keyword argument. Duck-typed file_ops implementations that only implement the two-argument write_file(path, content) contract would raise TypeError. Fix: wrap the call in try/except TypeError, falling back to the two-argument form when the keyword is not accepted. Also declare tomli in pyproject.toml (pre-existing conditional import for pre-3.11 Python, caught by the pre-commit dep scan after staging file_operations.py). Tests: Add TestV4ABomRoundTrip with two cases: - UPDATE on BOM-bearing file preserves the marker - UPDATE on plain file does not inject a BOM Addresses teknium1 review on PR #55661. * fix(file-ops): surrogatepass in bytes_written encode (review finding) Content that flowed through a surrogateescape decode (backend output via patch_replace) can carry lone surrogates; a strict encode raises UnicodeEncodeError where the old wc -c path could not. Mirrors the existing sha256 verification encode. * refactor(file-ops): fold simplify-pass findings - write_file: encode content once, share bytes between bytes_written and the sha256 verification (drops a second full-content encode per write) - patch_parser: replace the except-TypeError retry around write_file(pre_content=...) with signature-based feature detection so a TypeError raised inside a capable implementation propagates instead of triggering a duplicate write; tests for both duck-typing contracts - tests: real-ops V4A BOM round-trip + _file_has_bom disk-probe guard (the teknium1-review regression previously only covered by a fake) - comment: document dirs_created's long-standing "parent ensured" meaning * chore: update uv.lock for tomli dependency (rebase fix) * chore: remove dead tomli dependency declaration requires-python is >=3.11 so tomllib is always in stdlib; the tomli fallback branch in _lint_toml_inproc was unreachable. Removes the dependency from pyproject.toml + uv.lock and deletes the dead try/except ImportError fallback in the code. * fix(telegram+sqlite): resolve polling conflict loop + misleading WAL warning #75017: Telegram polling conflict retry used drop_pending_updates=False, starting a new getUpdates session that immediately got 409'd by the previous still-expiring session — creating the very conflict it was trying to recover from. Switch to drop_pending_updates=True so Telegram terminates stale sessions. Also add a recovery-generation guard so the first transient getUpdates success after a retry doesn't reset the conflict counter back to 0 (defense-in-depth from PR #75096). #75153: The WAL-reset warning always said 'hermes update can repair' even for git/pip/system Python installs where it can't. Now uses detect_install_method() + recommended_update_command_for_method() to give a context-appropriate hint (hermes update for git, docker pull for docker, nix message for nix, generic install hint as fallback). * chore: map jun@junho.co to junhohong * fix(dashboard): reload loopback tabs after stale session-token closes Loopback dashboard tabs now share one one-shot stale-token recovery path across REST 401s, the PTY socket, the structured event socket, and the shared JSON-RPC gateway wrapper. The shared client exposes only an optional close-event interception hook; the dashboard remains responsible for deciding that loopback 4401 means reload. Constraint: Current main delegates the web gateway to apps/shared JsonRpcGatewayClient, and #54022 review requires a shared-client-compatible close-code hook plus direct ChatSidebar event-socket coverage. Rejected: Restore the dashboard's old direct WebSocket implementation | stale against the shared JSON-RPC client and would duplicate transport behavior. Confidence: high Scope-risk: moderate Directive: Keep stale-token policy dashboard-specific; the shared JSON-RPC client should expose close events without learning dashboard auth semantics. Tested: npm --workspace web test (21 files, 106 tests); focused stale-token tests (5 files, 14 tests); npm --workspace web run typecheck; npm --workspace @hermes/shared run lint; npm --workspace @hermes/shared run typecheck; focused web eslint; git diff --check. Not-tested: Manual browser smoke test across a real dashboard restart. * fix: update ChatPage test import for react-router v7 react-router v7 exports MemoryRouter from 'react-router', not 'react-router-dom'. The test was written when the repo still imported from 'react-router-dom' (4000+ commits ago). * fix: wire HermesConsoleModal WS into stale-token reload guard Sibling site missed by PR #54022 — /api/console WebSocket in HermesConsoleModal.tsx has the same buildWsUrl → stale-token → 4401 close path as the PTY and events WebSockets. Without this guard, opening the console after a dashboard restart shows 'Console closed (4401). auth: token_mismatch' with no recovery. * fix(xai): honor configured web search backend on Responses path When Grok runs on xAI Responses, only swap to native server-side web_search when the active/configured backend is xai. For Firecrawl and other Hermes providers, keep client dispatch under a renamed wire tool so Grok cannot hijack web_search and ignore user config. * test(xai): cover Firecrawl vs native web_search on Responses Lock in backend preference, wire-name aliasing, and normalize mapping so configured non-xai search providers stay on the Hermes client path. Also init conflict-recovery generation on the telegram bare-adapter helper so CI polling progress tests do not AttributeError. * refactor(xai): simplify _xai_prefers_native_web_search to use registry Drop the manual web.search_backend / web.backend config-reading block that duplicated _read_config_key in web_search_registry.py. The function now delegates directly to get_active_search_provider() (which reads the same config keys via the registry's canonical resolver) and falls back to _get_search_backend() only when the registry has no providers loaded. Also updates the TestXaiWebSearchBackendPreference tests to monkeypatch the registry instead of load_config_readonly, and adds two new tests for the legacy fallback path (no provider registered -> _get_search_backend). * fix(relay): gate skipped task completion Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com> * fix(model-switch): treat models dict as metadata, not allowlist hermes model saves custom_providers models: {default: {context_length}} for local Ollama. That dict shape was treated as an explicit catalog, so no-key end…
Loopback dashboard tabs now share one one-shot stale-token recovery path across REST 401s, the PTY socket, the structured event socket, and the shared JSON-RPC gateway wrapper. The shared client exposes only an optional close-event interception hook; the dashboard remains responsible for deciding that loopback 4401 means reload. Constraint: Current main delegates the web gateway to apps/shared JsonRpcGatewayClient, and NousResearch#54022 review requires a shared-client-compatible close-code hook plus direct ChatSidebar event-socket coverage. Rejected: Restore the dashboard's old direct WebSocket implementation | stale against the shared JSON-RPC client and would duplicate transport behavior. Confidence: high Scope-risk: moderate Directive: Keep stale-token policy dashboard-specific; the shared JSON-RPC client should expose close events without learning dashboard auth semantics. Tested: npm --workspace web test (21 files, 106 tests); focused stale-token tests (5 files, 14 tests); npm --workspace web run typecheck; npm --workspace @hermes/shared run lint; npm --workspace @hermes/shared run typecheck; focused web eslint; git diff --check. Not-tested: Manual browser smoke test across a real dashboard restart.
Sibling site missed by PR NousResearch#54022 — /api/console WebSocket in HermesConsoleModal.tsx has the same buildWsUrl → stale-token → 4401 close path as the PTY and events WebSockets. Without this guard, opening the console after a dashboard restart shows 'Console closed (4401). auth: token_mismatch' with no recovery.
Loopback dashboard tabs now share one one-shot stale-token recovery path across REST 401s, the PTY socket, the structured event socket, and the shared JSON-RPC gateway wrapper. The shared client exposes only an optional close-event interception hook; the dashboard remains responsible for deciding that loopback 4401 means reload. Constraint: Current main delegates the web gateway to apps/shared JsonRpcGatewayClient, and NousResearch#54022 review requires a shared-client-compatible close-code hook plus direct ChatSidebar event-socket coverage. Rejected: Restore the dashboard's old direct WebSocket implementation | stale against the shared JSON-RPC client and would duplicate transport behavior. Confidence: high Scope-risk: moderate Directive: Keep stale-token policy dashboard-specific; the shared JSON-RPC client should expose close events without learning dashboard auth semantics. Tested: npm --workspace web test (21 files, 106 tests); focused stale-token tests (5 files, 14 tests); npm --workspace web run typecheck; npm --workspace @hermes/shared run lint; npm --workspace @hermes/shared run typecheck; focused web eslint; git diff --check. Not-tested: Manual browser smoke test across a real dashboard restart.
Sibling site missed by PR NousResearch#54022 — /api/console WebSocket in HermesConsoleModal.tsx has the same buildWsUrl → stale-token → 4401 close path as the PTY and events WebSockets. Without this guard, opening the console after a dashboard restart shows 'Console closed (4401). auth: token_mismatch' with no recovery.
What does this PR do?
Fixes the loopback dashboard case where a restart rotates the injected session token, so an already-open tab keeps reconnecting with a stale token and gets stuck on repeated 401/4401 failures until the user manually reloads.
This is intentionally narrower than #54003: it keeps stale-token policy in the dashboard and reuses one one-shot reload path instead of persisting session tokens across restarts. The shared JSON-RPC client only exposes a transport-level close interception hook and remains auth-policy-neutral.
Related Issue
Fixes #53972
Type of Change
Changes Made
web/src/lib/dashboard-auth-reload.tsonSocketCloseinterception hook to the sharedJsonRpcGatewayClientwithout embedding dashboard auth semanticsChatSidebar/api/eventssocket 4401ChatSidebarregression testHow to Test
npm --workspace web test(21 files, 106 tests).npm --workspace web run typecheck.npm --workspace web run build.npm --workspace @hermes/shared run typecheck.npm --workspace @hermes/shared run lint./chatopen, restart the dashboard backend, and confirm the tab reloads once instead of getting stuck on repeated 401/4401 auth failures.Checklist
Code
fix(scope):,feat(scope):, etc.)pytest tests/ -qand all tests passDocumentation & Housekeeping
Screenshots / Logs
N/A
Manual browser restart smoke testing was not performed locally.