Sync upstream hermes-agent main into NousAI-Assistant (375 commits, through af585de28e) — clean, no conflicts - #58
Merged
uaixo merged 377 commits intoAug 15, 2026
Conversation
…ils; test degradation Review follow-up (cc3f181): if AIAgent() raises inside _build_child_agent the freshly-opened dedicated handle has no owner and no child close() will ever run — release it on the exception path so the sqlite fds don't outlive the failed spawn. Also pin the degradation contract with a test: a parent without a SessionDB still yields session_db=None children.
…b_path A bare SessionDB() resolves the launch profile's default state.db, but parents can hold non-default per-profile handles (tui_gateway opens SessionDB(db_path=<profile_home>/state.db) for non-launch profiles and hands them to agents via _transfer_db_to_agent). A child of such a parent would write its transcript into the WRONG database — cross- profile leakage that breaks parent_session_id lineage and session_search. Open the dedicated handle at the parent handle's db_path instead (AsyncSessionDB forwards .db_path via __getattr__, so the gateway wrapper path works too). Regression test verified RED on the pre-fix code.
…op chains A turn writing against a session already closed by compression died with session_persistence_failed and a misleading "this is often a full disk" dialog, even though the store was healthy and a live continuation existed (NousResearch#82001). Depth-1 recovery (find_live_compression_child) could not resolve lineages with >=2 compression hops (root -> mid -> tip), reproduced independently on two- and three-hop chains. - run_agent.py flush chokepoint: on CompressionSessionClosedError, resolve tip = db.get_compression_tip(old_id) (canonical bounded transitive walk), adopt only when tip != old_id AND the tip row is live, retry the flush exactly once (adoption budget); otherwise fail closed. - gateway/session.py append_to_transcript: replace the depth-1 live-child lookup with the same tip + liveness contract, so gateway transcript reroutes follow full chains. - agent/conversation_compression.py _adopt_live_compression_child: turn-start recovery preflight now resolves via get_compression_tip with the same liveness check, closing the last depth-1 consumer in this family. - classify_persistence_error: new "compression_closed" bucket; the turn-end explanation names compression rotation and tells the client to refresh the session id instead of blaming a full disk. Tests: depth-1 adoption, multi-hop chain adoption (agent + gateway), fail closed with no continuation / stale-closed (ws_orphan_reap) tip, exactly-once adoption budget, and error-wording guards (compression-closed never mentions disk; real disk failures keep disk guidance). Closes NousResearch#82001 Co-authored-by: Al3xand3r1987 <125030427+Al3xand3r1987@users.noreply.github.com> Co-authored-by: yuzilongleif-collab <235949691+yuzilongleif-collab@users.noreply.github.com>
…air exception paths (NousResearch#83226) Two call sites create SessionDB instances without closing them on error: 1. gateway/slash_commands.py: /insights command - db.close() was on the success path but not in a finally block, so exceptions between SessionDB() and db.close() leak the connection. 2. hermes_cli/sessions_cmd.py: sessions repair - SessionDB() created inline with no .close() at all, leaking the FD on every call. Salvage note: the original PR (NousResearch#83237) also added a __del__ safety-net finalizer to SessionDB; review showed the atexit hook registered by queue_token_counts() strongly retains the instance, so the finalizer never fires for the leak class it claimed to cover. Dropped here in favor of the deterministic constructor-finally ownership repair salvaged from NousResearch#83620.
…usResearch#83226) SessionDB could leave native SQLite handles open when construction failed partway through schema/pragma/FTS/repair/lock/interrupt handling. Other short-lived callers (MCP reads/polling, session search, reactions, trace upload, insights, shutdown recovery) opened temporary SessionDB handles without a complete ownership boundary. API-server profile caches and RetainDB shutdown had similar late-close races. Under sustained load this exhausted file descriptors (EMFILE). - Close partially initialized SessionDB connections on every constructor exception path via a finally block guarded by an initialization-complete flag. - Close temporary/cross-profile SessionDB handles in finally blocks across CLI, MCP, search, trace, reactions, insights, and recovery paths. - Add API-server per-profile cache ownership and disconnect cleanup. - Make RetainDB writer-queue shutdown exception-safe: track connections per thread, close on worker exit, reject new enqueues after shutdown starts, and sweep any connections left by short-lived threads. - Add regression coverage for constructor failures, worker-thread readers, API disconnect failures, shutdown recovery, RetainDB late enqueue, and foreign-loop async clients. Salvage notes: the original PR's per-thread WAL-reader ownership changes were superseded by main's read-connection pool (permits + checkout/return); its cron timeout-abandon fix is credited separately to NousResearch#72822's earlier identical fix. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…imeout-abandoned worker (NousResearch#72782) run_job() submits SessionDB() to a one-worker executor and abandons the worker (shutdown(wait=False)) when init exceeds the cron timeout. If the constructor later completes inside that abandoned worker, the Future's result — an open SessionDB holding .db/WAL/SHM handles — was orphaned and never closed, leaking descriptors until EMFILE. Attach a done-callback on the timeout path that retrieves and closes any eventual late result. Salvage note: the lazy-recall ownership half of NousResearch#72822 (_owns_session_db tracked on AIAgent, owned handle closed in close()) already landed on main; this carries the remaining cron timeout-abandon half with its regression test.
…d by read pool - tests/cron/test_sessiondb_init_hang.py: add threading/time imports the salvaged late-close regression tests rely on. - tests/test_hermes_state.py: drop test_close_closes_wal_read_connection_created_on_worker_thread — main replaced per-thread WAL reader ownership with the pooled read-connection design (permits + checkout/return), so cross-thread reader draining no longer exists in the form the test asserted.
_reap_unsupervised_gateway_orphans() kills every gateway PID found by find_gateway_pids() on hosts without systemd (macOS launchd, Windows Scheduled Task). This includes service-managed gateways that are NOT orphans — they are supervised by launchd/systemd and should never be killed during a stale-process sweep. Add own |= _get_service_pids() to the exclusion set before scanning, so launchd/systemd-supervised gateways are preserved. True orphans (reparented leftovers not present in launchctl/systemctl) are still found and reaped, preserving the NousResearch#77276 protection. Fixes NousResearch#85344 (macOS launchd gateway killed by desktop serve startup) Fixes NousResearch#85044 (Windows Scheduled Task gateway killed by desktop serve) Fixes NousResearch#84855 (Permission denied to kill orphaned gateway PID) Fixes NousResearch#85368 (gateway process repeatedly killed, messaging offline)
…per on Windows The orphan reaper kills a healthy gateway (and its Scheduled-Task bootstrap parent chain) every time the Desktop backend starts on Windows, because _get_service_pids() only implements systemd/launchd and returns an empty set on Windows — a supervised gateway is therefore indistinguishable from an unsupervised orphan. Exempt the recorded healthy gateway PID and its parent chain from the orphan scan on Windows, mirroring the macOS launchd exemption (NousResearch#85913). The Scheduled-Task bootstrap's argv matches the gateway scan, so without exempting the parent chain killing the bootstrap takes the detached gateway down with it. Fixes NousResearch#86098
…r to all platforms Compose the service-PID exclusion (NousResearch#85743, RelaxJonh) and the recorded-PID + parent-chain exemption (NousResearch#86100, arccat-114) into one cross-platform rule: - _get_service_pids() exclusion now runs unconditionally, not only under is_macos() — it is the authoritative "supervised" signal for launchd and any systemd unit visible on a host that got past the systemd gate. - The recorded-healthy-gateway (get_running_pid()) + parent-chain exemption now runs on every platform, not only Windows. A recorded, liveness-verified gateway is by definition not an orphan "the pidfile/runtime record can't see", so the reaper must never target it — this covers Windows Scheduled Task / Startup VBS supervision, standalone launcher-started gateways (the case NousResearch#85743 alone would miss), and macOS/WSL equivalents. True orphans (no service registration, no valid runtime record) are still found and reaped, preserving the NousResearch#51325/NousResearch#75936 duplicate-port protection. Existing macOS regression tests updated to pin get_running_pid to None for their scenario; Windows regression tests from NousResearch#86100 carry over unchanged. Bug class: NousResearch#83683 (root), NousResearch#86287, NousResearch#86098, NousResearch#85738, NousResearch#85368, NousResearch#85344, NousResearch#85044, NousResearch#84855, NousResearch#84824, NousResearch#84200.
…ousResearch#78821) Filter manual dashboard/serve respawn candidates after update: skip ephemeral --port 0 backends (Desktop-owned), dedupe normalized cmdlines, and cap one restart per profile/HERMES_HOME so orphan counts no longer grow across successive updates.
`agent.restart_drain_timeout` defaults to 0 and governed every class of in-flight work at once. That default is deliberate for chat turns: the gateway announces the restart to the user and pre-marks the session resume_pending, so interrupting one is cheap and recoverable. A cron run has neither property. Nobody is waiting on it, it is written to jobs.json as a permanent failure, and a recurring job simply skips to its next schedule. Sharing the chat budget meant `_drain_active_agents()` short-circuited on `timeout <= 0` before entering the wait loop, so the drain reported `drain took 0.00s, timed_out=True, cron_at_start=1, cron_now=1` — it detected the job and killed it anyway. Cron work now drains on its own deadline, `agent.cron_drain_timeout` (default 30s, 0 opts out). The floor is clamped to the shutdown-watchdog leash minus a teardown reserve, so the longer wait can never consume the post-drain cleanup window: being SIGKILLed mid-cleanup would leave the job wedged at `last_status=running`, strictly worse than the bug. Being bounded also means a cron-triggered restart cannot deadlock on itself. The `timeout <= 0` special case is gone — an expired deadline expresses the legacy "interrupt immediately" behaviour, so `timed_out` is always computed from real state instead of asserted up front. The drain-timeout warning now reports the elapsed wait rather than the configured budget, which is what made "timed out after 0.0s" so confusing in the report. Chat-only shutdowns are unchanged: `restart_drain_timeout: 0` still interrupts chat turns immediately. Relates to NousResearch#82161 (complements NousResearch#82195, which removes the `hermes update` self-deadlock that triggered the reported instance).
… doubles CI slice 5/12 caught two ways the new cron budget broke `_stop_impl_body` for callers that are not real GatewayRunner instances: - `_FakeGateway` in test_shutdown_cache_cleanup.py borrows `_stop_impl` without subclassing, so it never picked up the class-level `_cron_drain_timeout` default and raised AttributeError. Read it through the getattr-guard convention the same function already uses for its liveness-guard machinery. - The same double overrides `_drain_active_agents(self, timeout)`, so passing the cron budget raised "takes 2 positional arguments but 3 were given". The double now mirrors the real optional parameter. It is the only override in the tree; test_startup_restart_race.py uses AsyncMock, which accepts any signature. Verified against a stashed clean tree: the 22 gateway test files that still fail locally fail identically with and without this branch (80 = 80, empty set difference both ways) — they are pre-existing Windows-only failures (setsid, POSIX modes) unrelated to this change.
…nect When the shutdown drain times out and kills an in-flight cron job, the job's owner is never told. The cron worker does try: `_is_interrupted()` forces the failure path with an honest "interrupted by gateway shutdown" error, and failed jobs always deliver. But that worker is a thread, it reaches `_deliver_result()` asynchronously, and by then `_bounded_adapter_teardown()` has closed the transport. The reporter of Worse, the loss is silent twice over: `_consume_interrupted_flag()` returns True — the gateway already wrote `last_status` — so `mark_job_run()` is skipped, and the `delivery_error` from the failed send is discarded with it. The run's only trace is a generic line in jobs.json. The gateway already owns the right window. `_notify_active_sessions_of_ shutdown()` runs while adapters are up, precisely so shutdown messages can be sent — but it iterates `_running_agents`, and cron work lives on the scheduler's own thread pool. Same structural blindness already fixed for counting (NousResearch#60432) and draining (NousResearch#63529), never fixed for notifying. So notify from the post-interrupt phase, which is the last point where the transport is still up: `_kill_tool_subprocesses()` now returns the job IDs it marked, and `_notify_interrupted_cron_jobs()` sends each one's owner a notice on the job's own resolved delivery targets. Adapter teardown order is untouched — it is load-bearing for NousResearch#53175 and NousResearch#8202. Jobs with `deliver: local`, and `deliver: origin` jobs with no resolvable origin (NousResearch#43014), resolve to zero targets and stay silent. Per-platform `gateway_restart_notification: false` is honoured, matching the chat path. Every failure is swallowed so a wedged adapter cannot extend shutdown. Second, when the interrupted flag short-circuits `mark_job_run()`, the delivery failure is now persisted on its own via `update_job()`, so a notice that still cannot be sent is at least recorded. `update_job()` rather than a second `mark_job_run()`: the latter also advances `next_run_at` and the repeat counter, and running that twice for one run would skip a fire or auto-delete the job early. Fixes NousResearch#82232. Related: NousResearch#82161, NousResearch#82224.
…send mark_running_jobs_interrupted skipped legacy fires without a registered durable owner entirely — correct for the persisted last_status write (no owner fence to protect a replacement run), but the gateway shutdown path also uses the returned ID list to deliver interrupted-cron notices while adapters are still connected (NousResearch#82232). Keep the persistence skip, but include the job in the returned list so the user is still told.
…'t erase config `hermes import` wrote every zip member with `open(target, "wb")` followed by `dst.write(src.read())`, at both restore sites in `run_import`. Opening for write truncates the user's existing file to zero *before* any replacement bytes exist, so a Ctrl-C, an ENOSPC, a corrupt zip member, or a crash leaves `config.yaml`, `.env`, or an external provider config (e.g. `~/.honcho/config.json`) empty with nothing behind it — during the disaster-recovery path the user is running precisely because they already lost something. The `_external/` branch writes outside HERMES_HOME, into third-party configs under the user's home, so the blast radius is not confined to Hermes state. Both sites now stage the member into the target's own directory, fsync it, and publish with `utils.atomic_replace`, so the target only ever moves from its old contents to the complete new contents. `atomic_replace` rather than a bare `os.replace`: it resolves a symlinked target first, so deployments that link `config.yaml` into a dotfiles repo keep the link instead of having it silently swapped for a regular file (NousResearch#16743), and it falls back to copy/fsync/unlink on EXDEV/EBUSY for cross-device and bind-mount installs. Members stream through `shutil.copyfileobj` instead of being read whole into memory. The temp file is removed on any failure so a partial import leaves no residue, and permission bits are carried across the replace so mkstemp's 0600 does not silently tighten restored files. This extends the module's own established idiom — `backup.py` already publishes atomically via `os.replace` in `_atomic_output_path` and in the snapshot writer — into the one path that still overwrote user files in place.
…0 transit window Follow-up on the atomic-import restore, delegating both metadata concerns to the shared helpers instead of half-handling them locally. Owner preservation was missing entirely. `tempfile.mkstemp` + `atomic_replace` publishes a temp file owned by the *writing* user, so `sudo hermes import` re-owned every restored file to root — on the disaster-recovery path, and on exactly the Docker/NAS volume installs `utils._restore_file_owner` was added for. `_extract_member_atomically` now captures `_preserve_file_owner(target)` before staging and calls `_restore_file_owner` after the replace, before the mode restore (chown clears setuid/setgid, so the mode has to go back last). Mode handling was also only half applied before the replace: the `os.fchmod` branch applied it to the temp fd, but the platforms without `fchmod` fell through to a best-effort post-replace chmod, leaving the published file at mkstemp's 0600 until that chmod landed — permanently if the process died in between — and making `atomic_replace`'s EXDEV/EBUSY `shutil.copystat` fallback copy 0600 onto the target. The mode is now applied to the temp file on both branches, with the post-replace `_restore_file_mode` kept as the belt-and- braces path. This is the same shape `atomic_write_text` and `atomic_yaml_write` already carry after 3556728 and 43fc865; capture and restore now reuse `utils._preserve_file_mode` / `_preserve_file_owner` / `_restore_file_mode` / `_restore_file_owner` rather than re-deriving them, which also drops the local `import stat`. Tests (tests/hermes_cli/test_backup.py, class TestImportAtomicWrites): - test_restore_preserves_existing_file_owner — forces a uid/gid so it does not need root; asserts chown fires once, with the captured owner, on the pre-existing file only (a newly created member has no prior owner). Mutation-checked: dropping only the `_restore_file_owner` call reds it. - test_mode_is_applied_before_the_replace_without_fchmod — `monkeypatch.delattr` on `os.fchmod`, spies the temp file's mode at replace time. Reads 0o600 without the fix, 0o644 with it. Mutation-checked the same way.
…truncates The atomicity claim in _extract_member_atomically's docstring holds on the os.replace path but not on atomic_replace's EXDEV/EBUSY fallback, which uses shutil.copyfile and so opens the destination 'wb'. That is pre-existing behaviour shared by every atomic writer in the repo, and it is reachable here for a symlinked target whose real file lives on another filesystem. Scope the docstring to what the helper actually guarantees instead of overstating it; the fallback itself is a utils.atomic_replace change.
…files ``_extract_member_atomically`` carries the replaced file's permissions across the publish so that routing through mkstemp does not change what the caller would otherwise have produced. But ``_preserve_file_mode`` returns ``stat.S_IMODE``, which is all twelve bits, and this restore is deliberate on both sides of the replace: the mode is fchmod'd onto the temp before ``atomic_replace`` and re-applied afterwards because chown clears the elevated bits. So a target sitting at 0o4755 comes out of ``hermes import`` still at 0o4755 — with contents supplied by the zip. That is a regression introduced by the atomic rewrite rather than a pre-existing one. The overwrite it replaced was an in-place ``open(target, "wb")``, and an in-place write by a process without CAP_FSETID has the elevated bits stripped by the kernel, so the old path left 0o4755 as 0o755. The blast radius is not limited to Hermes' own state: the ``_external/`` branch of ``run_import`` publishes members anywhere under ``$HOME``, and this is the path that documents ``sudo`` use so ownership survives a restore. An archive that happens to contain a member matching some existing privileged file would take over the identity that file runs as. Mask the two bits off the preserved mode. The masking happens once, before the temp file is chmod'd, so there is no transient elevation either. The sticky bit is kept — it is inert on a regular file. The ordinary permission bits are unaffected, so the Docker/NAS installs the preservation exists for still get their broader modes back. This is the one write path in the repo where the bytes are untrusted; the ``utils`` writers that preserve the full mode re-serialize content the process itself produced, and are correct as they stand.
… bits Pre-creates a 0o6755 target, imports a member over it, and asserts the published file is 0o755 with both elevated bits gone — plus that the staged temp file never carried them either, so there is no window where archive content sits behind an elevated mode. The existing coverage in this class cannot see the failure: every mode assertion masks with ``& 0o777``, which discards exactly the bits at issue, and the fixtures chmod their targets to ordinary modes that never had them set. Without the mask on the preserved mode this test reports the published file still holding S_ISUID. Skipped where the platform or filesystem refuses setuid on a user-owned file, so the assertion never depends on running as root.
Node 22 ships npm 11.16.0, which engines.npm rejects (11.10–11.16 ignore min-release-age-exclude). Fresh Hermes-managed installs then fail npm ci with EBADENGINE. Node 26 ships 11.17.0.
…wn (macOS) The Desktop-spawned hand-off consistently died during Electron's quit teardown on macOS: the orchestrator process group was terminated right after `running: hermes update ...`, so no exit code, result file, bundle swap, or relaunch ever happened, and the loopback shim window surfaced the death as ERR_CONNECTION_REFUSED or "Aw, Snap!" error code 15 (reproductions in NousResearch#66753). - Re-exec the orchestrator through a one-shot setsid child and let the direct Electron child exit immediately; the real orchestrator is owned by launchd (PPID 1), outside Electron's teardown, same marker/result protocol. - Hold TERM ignored across the `hermes update` invocation and log-and-ignore the single teardown TERM that can still arrive after the desktop PID dies (durable SIGNAL breadcrumb for diagnosis). - Delay start_ui until the desktop PID is gone plus 1s so the shim server/window are never born inside the teardown window. - Run both UI processes in their own sessions; keep SIGTERM/SIGHUP ignored in the shim server and stop it with SIGKILL, so a stray TERM can no longer leave the progress window on a dead loopback URL while the update continues. Verified on a production git install (macOS arm64, Darwin 27.0, v0.20.1): six consecutive Desktop-triggered/production-shape updates completed end-to-end including a full desktop rebuild + codesign; the shim survived a deliberately injected TERM+HUP mid-update and a full `hermes desktop --force-build --build-only` running alongside it. Fixes the macOS reproductions in NousResearch#66753. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…de-newer bricking The relative exclude-newer = "14 days" cutoff bricks installs whenever the resolver cannot see (or accept) a package's upload date: - defusedxml / python-olm / unpaddedbase64 (NousResearch#80387, NousResearch#79434): ancient frozen releases (2021-2023) whose upload dates are often absent from mirror indexes and stale uv HTTP caches. uv then filters them entirely ("there are no versions of defusedxml"), breaking [youtube]/[wecom]/ [matrix] resolution and daily `uv sync --locked` runs. - setuptools / pillow / mcp (NousResearch#78227, NousResearch#75992, NousResearch#76020): exact-pinned deps. When the pinned version's upload date is invisible, the resolver filters the ONLY acceptable candidate — setuptools==83.0.0 in [build-system].requires meant the project could not even be built from a git checkout on released v0.20.0. Exempting an exact pin costs nothing: the version cannot float without a reviewed pin bump. Changes: - pyproject.toml: add all six to the existing exclude-newer-package whitelist, with rationale comments per class. - uv.lock: regenerated; diff is the whitelist metadata only (verified zero version drift, still 249 packages). - tests/test_packaging_metadata.py: new standing guard test_build_system_requires_exempt_from_exclude_newer — every [build-system].requires package must be whitelisted while a relative exclude-newer cutoff is configured. Verified both directions (fails when setuptools is removed from the whitelist). - scripts/install.sh: fix the stale tier-name comparison ("all (with RL/matrix extras)" vs actual "all") that mislabeled every successful Tier-1 install as a fallback-tier install (NousResearch#79434 bonus finding). Verification: uv lock --check green on uv 0.11.19 and 0.12.5; uv sync --extra all --locked green; uv pip install -e '.[all]' resolves; whitelist mechanism A/B-proven on a minimal project (unsatisfiable -> resolves; build-requires variant: uv build fails -> succeeds). Reported-by: MichaelClawHub (NousResearch#80387), liujianqiu (NousResearch#79434), maxonliu (NousResearch#78227)
…script Session hygiene could persist a compressed transcript LARGER than the original (observed: 427K -> 598K), when the generated summary was bigger than the middle it replaced. Compare like-for-like (both rough estimates) before persisting a rotated transcript; on growth, keep the original unchanged so a failed compression is a strict no-op, never a net increase.
… (in-place path) The gateway rotation guard (NousResearch#83339) only protects the rotate path, but in-place compaction commits inside compress_context() via archive_and_compact — before the gateway can inspect the result. Add the anti-growth check at the commit site so both paths are covered: a compression whose rough output exceeds its input is a strict no-op (original transcript kept durable, session identity untouched). Covers the observed failure where session hygiene persisted 426 -> 426 messages and ~379K -> ~688K tokens.
…ment Simplify-code pass: gateway/run.py called estimate_messages_tokens_rough 6x on the same data in the anti-growth guard (condition + warning f-string). Bind to _hyg_in_toks/_hyg_out_toks locals like the conversation_compression.py guard already does. Also trim the comment from 10 lines to 4 (keep the WHY, drop the WHAT) and remove an extra blank line before TestCompactedTurnsStaySearchable.
(cherry picked from commit ef1d2cc)
…tion Residual hunk from PR NousResearch#84832's picker-rebind commit; the functional surfaces.tsx change landed on main via NousResearch#86250. (cherry picked from commit 990936d, docstring hunk only)
(cherry picked from commit 57c51cc)
Read the durable display transcript when creating a branch instead of copying the compacted model projection. Hydrate the Desktop branch boundary from persisted history, avoid stale whole-chat counts, and seed the new tile from the backend snapshot. Add regression coverage for compacted histories, visible-message counts, selected prefixes, and hydration races. (cherry picked from commit c3d2d75)
…ry CLI start The main.py decomposition re-exported the sessions/update/dashboard command surface with eager from-imports, so every hermes invocation (including hermes --version) paid for update_cmd's dependency chain (jwt, click, cryptography). Resolve the re-exports through the existing PEP 562 module __getattr__ (same pattern as _PROVIDER_MODELS) so each module loads on first actual use. Internal call sites go through a _self() helper because bare-name lookups do not trigger __getattr__; _self() imports sys locally since update tests patch hermes_cli.main.sys. The _warn_stale_dashboard_processes back-compat alias moves into the lazy surface, and the sessions argparse dispatch defers sessions_cmd to call time. Monkeypatching hermes_cli.main.<name> keeps working: a patch sets a real module attribute, which shadows __getattr__. Measured (Windows 11, Python 3.11, median of 7 warm runs): import hermes_cli.main 253ms -> 196ms (-22%). (cherry picked from commit cad1083)
The /tools and /personality completers run on every keystroke while the user types those commands (complete_while_typing), and both re-read + re-parse the full config on each keypress: - _tools_completions called load_config() — the defensive deepcopy (~340us/call on cache hit) even though it only reads toolset enable state + MCP server names. Switched to load_config_readonly() (the perf(agent) NousResearch#74322 pattern; this per-keystroke site was missed). - _personality_completions called load_cli_config() — a full YAML parse + deep merge of the built-in defaults (~110us) — on every keystroke. Memoised keyed on the config file path+mtime (same pattern as load_env / _nous_auth_status_cache), so the parse runs once per config state. Measured: /tools 357us -> 18us per keystroke; /personality parse drops from 1-per-keystroke to 1-per-config-change (500 keystrokes -> 1 parse). Regression tests: _tools_completions uses the readonly loader (deepcopy loader never called); personality memo parses once across repeated completions and re-parses once after a config mtime change. (cherry picked from commit 2b3f897)
…re read get_default_hermes_root() resolves HERMES_HOME against the platform native home (~80us of path resolution) on EVERY call and is called at 31+ sites — every _load_global_auth_store() (per provider row in the /model picker), kanban, backup, gateway, update. Its result depends only on (HERMES_HOME, native home), so memoise it keyed on those two inputs, compared for free on each call (freshness-correct even if a test or plugin mutates HERMES_HOME mid-process). _load_global_auth_store() re-read + re-parsed the global auth.json on every call; read_credential_pool() -> load_pool() runs it once per provider row in the /model picker even when the profile has entries and the global fallback never fires. Memoise keyed on the global auth file's path+mtime (same pattern as _nous_auth_status_cache); the store only changes when a global-scope auth write touches the file. Measured (profile mode, 30-provider global store): get_default_hermes_root 81us -> 10us; _load_global_auth_store 128us -> 66us; load_pool 165us -> 137us per call — ~2ms saved per /model picker render (20 provider rows). Regression tests: hermes_constants memo pin (no path resolution on repeat calls, HERMES_HOME change forces a fresh resolution); global-store memo pins (store read once across repeats, mtime bump re-reads once, absent store stays cheap). (cherry picked from commit be348f3)
(cherry picked from commit 4822dae)
…string (cherry picked from commit 64de1a1)
resolve_toolset() recursively walks toolset includes and, with include_registry=True, merges registry-registered tools on every call — each external call re-runs the includes walk and takes a fresh registry snapshot under the registry lock. It is called dozens of times per _get_platform_tools() (every /tools completion keystroke, per picker render) and at 19 call sites across the CLI/gateway. Memoise the external-entry result keyed on (name, include_registry, registry id, registry generation). tools.registry exposes a monotonic _generation counter bumped on every register/deregister/alias/MCP refresh (its docstring explicitly invites generation-keyed memoisation), so a cache entry is valid until the registry changes. External callers never pass visited, so the memo engages exactly at the public entry and the internal cycle-detection recursion is untouched. Measured: _get_platform_tools drops 165us -> 59us per call (3x) with the xAI credential fix simulated; /tools completion ~2ms -> ~77us/keystroke combined. Regression tests: repeat resolution is a memo hit (get_toolset called once), a generation bump forces a fresh resolve, and the memoised result is identical to a fresh resolution. (cherry picked from commit 3d36ecb)
_file_mutation_verifier_enabled and _turn_completion_explainer_enabled re-read config.yaml on every call via load_config() (~1ms deepcopy per call). finalize_turn runs these gates at the end of every turn, so each turn paid two redundant config deepcopies. The sibling _credits_notices_enabled already caches on self; mirror that pattern. The env-var override stays authoritative and uncached, so runtime flips still work. Config flips now apply on the next session, matching the documented sibling semantics. (cherry picked from commit f2d0e00)
(cherry picked from commit 57a7044)
(cherry picked from commit e2e0edd)
…olve_toolset memo - The readonly-loader completer test now stubs get_portable_mcp_server_names_nowait — real plugin discovery runs load_config() during one-time process init, which is not the per-keystroke read the test guards against. - Cap _resolve_toolset_memo at 256 entries: generation-keyed entries from stale generations are never hit again, so clear on overflow to keep long sessions bounded.
Addresses both review findings on the remote-gateway download PR: 1. Unbounded buffering (finding #1). fetchBuffer / fetchBufferViaOauthSession accumulated the entire response (then copied it again via Buffer.concat) before saveGatewayFile even opened the save dialog, so a large gateway file could exhaust the native process. Both auth paths now stream: once response headers arrive the connect timeout is cleared, the filename is derived, the save dialog is shown, and the body is piped to the chosen destination with backpressure. A read/write error tears down the stream and unlinks the partial file. The byte-moving, data-URL decoding, and filename/path helpers are extracted into gateway-file-download.ts so they're unit-testable without Electron. 2. No fallback for older gateways (finding #2). saveGatewayFile required the new /api/fs/download route. Desktop and the remote gateway update independently, so a gateway predating this PR 404s. Added a 404-only compatibility fallback to the existing capped /api/fs/read-data-url route (bounded, so it only serves smaller files — enough to keep older backends working). Tests: gateway-file-download.test.ts covers streaming, backpressure, error-cleanup (unlink on write/response error), data-URL decoding, filename derivation (incl. traversal reduction), and 404 detection; gateway-file-download-transport.test.ts asserts both transports stream (no whole-body Buffer.concat) and that the 404 fallback is wired. Both registered in the desktop platform test list. Server-side /api/fs/download tests (streaming + sensitive-file reject) already pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…oject The salvaged branch predates the electron test project's node:test -> vitest migration (test:desktop:platforms is now `vitest run --project electron`). Import `test` from vitest so the suites are collected; assertions stay on node:assert/strict per the existing electron test convention.
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
…hrough af585de) Clean merge - no conflicts. Replay (git merge-tree) produced zero conflicts and the committed tree is byte-identical to the replay tree. Semantic verification on the merged tree (deeper than usual given the batch size - 375 commits, the largest sync to date): - drift grep for 'Hermes Desktop' in apps/desktop/{src,electron}: empty - every brand carve-out value intact: productName/executableName 'NousAI', appId ai.nous.desktop, artifactName NousAI-..., APP_NAME default, WORDMARK, DEFAULT_SKIN_NAME 'nousai', index.html title, brand-mark asset, one nousai dashboard theme row, brand-identity.test.ts - hermes_cli/main.py's brand-agnostic macOS packaged-app lookup survived 12 upstream commits to that file byte-for-byte - apps/desktop/package.json (carve-out) differs from our branch ONLY by upstream's two new repro:short-session-hang scripts - nothing reverted - no new i18n locales (still ar/en/ja/zh/zh-hant), so no rebrand sweep needed; upstream's new e2e specs pin only toContain('Hermes'), which our '<title>NousAI - Hermes</title>' satisfies by design No CI-sensitive workflow files and no package-lock.json changes, so neither review gate applies. uv.lock changes are exclude-newer-package booleans only (no dependency version bumps). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KFEJ7TzKwG4tQujNy3CWjT
૮ >ﻌ< ა ci reviewran on 4df11ee — fix(tests): align lost-and-found sessions pin with the 56-co
|
Upstream's d16326b bumped this test's sessions-schema pin for the new git_metadata_generation column (54 -> 55), but landed on a tree that already carried the 'hidden' column from fbaea9b, so the real layout is 56 and upstream main is red on test_mapper_rebuilds_sessiondb_from_synthetic_lost_and_found (assert 56 == 55). Verified byte-identical to upstream and reproduced locally. Bump the width in the four coupled places the value is load-bearing - the pin, max_fields (the lost_and_found table width), the current-layout insert(), and its comment - mirroring exactly what d16326b itself did. Bumping the pin alone would only move the failure into the row builders. Verified locally: pins pass and the 56/52/14/23-field inserts are all accepted. The correct value is objectively fixed by the schema (this is a schema pin, not a product-intent choice), so when upstream re-aligns it the next sync takes upstream's version of these hunks. Recorded in CLAUDE.md (user-approved 2026-08-15). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KFEJ7TzKwG4tQujNy3CWjT
uaixo
marked this pull request as ready for review
August 15, 2026 08:33
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Squashing a sync PR flattens upstream history and breaks future syncs. Merge with the "Create a merge commit" method only.
What this is
Routine upstream sync:
nousresearch/hermes-agent:main→NousAI-Assistant, picking up where PR #57 left off.bab9a85b67(exactly PR Sync upstream hermes-agent main into NousAI-Assistant (154 commits, through bab9a85b67) — clean; includes PR #56's deferred commits #57's sync point) through tipaf585de28egit merge-treereplay clean (single merge base, deep fetch verified), and the merge commit's tree is byte-identical to the replay tree (85cc408d6c)Follow-up commit: upstream-red schema pin (user-approved 2026-08-15)
First CI run failed only
Python tests / slice 8/12:test_mapper_rebuilds_sessiondb_from_synthetic_lost_and_foundasserts thesessionstable has 55 columns, but it has 56 (assert 56 == 55). Verified byte-identical to upstream and reproduced locally — upstreammainis red, not a merge artifact.Bisected precisely:
bab9a85b67(this PR's base)fbaea9bddc(hidden-session flag)d16326bb25af585de28ed16326bb25("align lost-and-found schema pins with git_metadata_generation column") bumped the pin 54 → 55 for its own new column, but landed on a tree that already carriedhidden, leaving the pin one short.Commit
4df11ee258bumps the width in the four coupled places the value is load-bearing — the pin,max_fields(thelost_and_foundtable width), the current-layoutinsert(...), and its comment — mirroring exactly whatd16326bb25itself did. Bumping the pin alone would only move the failure into the row builders. Verified locally: pins pass and the 56/52/14/23-field inserts are all accepted.Temporary divergence: this is a schema pin, not a product-intent choice, so the correct value is objectively fixed by the schema. When upstream re-aligns it, the next sync takes upstream's version of these hunks (recorded in CLAUDE.md with that resolution rule).
Semantic verification (deeper than usual, given the batch size)
"Hermes Desktop"inapps/desktop/{src,electron}: emptyproductName/executableNameNousAI,appId ai.nous.desktop,artifactName NousAI-…,APP_NAMEdefault,WORDMARK,DEFAULT_SKIN_NAME 'nousai',<title>NousAI — Hermes</title>,nousai-mark.png, onenousaidashboard theme row,brand-identity.test.tshermes_cli/main.py's brand-agnostic macOS packaged-app lookup survived 12 upstream commits to that file byte-for-byteapps/desktop/package.json(carve-out) differs from our branch only by upstream's two newrepro:short-session-hangscripts — nothing revertedtoContain('Hermes'), which our title satisfies by designfalse &&guards still in ci.yml)Gates
Neither review gate applies: no
.githubchanges at all, andpackage-lock.jsonuntouched.uv.lock's only change isexclude-newer-packageboolean entries — configuration, no dependency version bumps.Upstream highlights
max_concurrent_childrendefault 3 → 10 (with migration); truncated-subagent-result marking🤖 Generated with Claude Code
https://claude.ai/code/session_01KFEJ7TzKwG4tQujNy3CWjT