Skip to content

Sync upstream hermes-agent main into NousAI-Assistant (375 commits, through af585de28e) — clean, no conflicts - #58

Merged
uaixo merged 377 commits into
NousAI-Assistantfrom
claude/nousai-sync-upstream-20260815-2
Aug 15, 2026
Merged

uaixo merged 377 commits into
NousAI-Assistantfrom
claude/nousai-sync-upstream-20260815-2

Conversation

@uaixo

@uaixo uaixo commented Aug 15, 2026 •

Copy link
Copy Markdown
Owner

⚠️ MERGE WITH MERGE COMMIT — DO NOT SQUASH

Squashing a sync PR flattens upstream history and breaks future syncs. Merge with the "Create a merge commit" method only.

What this is

Routine upstream sync: nousresearch/hermes-agent:main → NousAI-Assistant, picking up where PR #57 left off.

Follow-up commit: upstream-red schema pin (user-approved 2026-08-15)

First CI run failed only Python tests / slice 8/12: test_mapper_rebuilds_sessiondb_from_synthetic_lost_and_found asserts the sessions table has 55 columns, but it has 56 (assert 56 == 55). Verified byte-identical to upstream and reproduced locally — upstream main is red, not a merge artifact.

Bisected precisely:

upstream commit schema cols test pin
bab9a85b67 (this PR's base) 54 54 ✅
fbaea9bddc (hidden-session flag) 55 55 ✅ last green
d16326bb25 56 55 ❌ break starts
… → tip af585de28e 56 55 ❌

d16326bb25 ("align lost-and-found schema pins with git_metadata_generation column") bumped the pin 54 → 55 for its own new column, but landed on a tree that already carried hidden, leaving the pin one short.

Commit 4df11ee258 bumps the width in the four coupled places the value is load-bearing — the pin, max_fields (the lost_and_found table width), the current-layout insert(...), and its comment — mirroring exactly what d16326bb25 itself did. Bumping the pin alone would only move the failure into the row builders. Verified locally: pins pass and the 56/52/14/23-field inserts are all accepted.

Temporary divergence: this is a schema pin, not a product-intent choice, so the correct value is objectively fixed by the schema. When upstream re-aligns it, the next sync takes upstream's version of these hunks (recorded in CLAUDE.md with that resolution rule).

Semantic verification (deeper than usual, given the batch size)

  • Drift grep "Hermes Desktop" in apps/desktop/{src,electron}: empty
  • Every brand carve-out value intact: productName/executableName NousAI, appId ai.nous.desktop, artifactName NousAI-…, APP_NAME default, WORDMARK, DEFAULT_SKIN_NAME 'nousai', <title>NousAI — Hermes</title>, nousai-mark.png, one nousai dashboard theme row, brand-identity.test.ts
  • hermes_cli/main.py's brand-agnostic macOS packaged-app lookup survived 12 upstream commits to that file byte-for-byte
  • apps/desktop/package.json (carve-out) differs from our branch only by upstream's two new repro:short-session-hang scripts — nothing reverted
  • No new i18n locales (still ar/en/ja/zh/zh-hant), so no rebrand sweep needed
  • Upstream's new/changed e2e specs pin no product name; the only brand assertions repo-wide remain toContain('Hermes'), which our title satisfies by design
  • Desktop E2E remains disabled upstream (both false && guards still in ci.yml)

Gates

Neither review gate applies: no .github changes at all, and package-lock.json untouched. uv.lock's only change is exclude-newer-package boolean entries — configuration, no dependency version bumps.

Upstream highlights

  • Desktop multi-connection registry — named agent sources (schema v2 + IPC + Settings UI)
  • Delegation: max_concurrent_children default 3 → 10 (with migration); truncated-subagent-result marking
  • Generic hidden session flag (sidebar-hide, still resumable); session list density modes; delete-session confirmation
  • Plugin SDK gains Dialog/ConfirmDialog/Toast/useToast/useConfirmDelete
  • Prompt cache scoping; reasoning-block auto-collapse in TUI and desktop
  • 50 new contributor mappings (from upstream)

🤖 Generated with Claude Code

https://claude.ai/code/session_01KFEJ7TzKwG4tQujNy3CWjT

thatssoheil and others added 30 commits August 14, 2026 21:39
…ils; test degradation

Review follow-up (cc3f181): if AIAgent() raises inside _build_child_agent
the freshly-opened dedicated handle has no owner and no child close() will
ever run — release it on the exception path so the sqlite fds don't
outlive the failed spawn. Also pin the degradation contract with a test:
a parent without a SessionDB still yields session_db=None children.
…b_path

A bare SessionDB() resolves the launch profile's default state.db, but
parents can hold non-default per-profile handles (tui_gateway opens
SessionDB(db_path=<profile_home>/state.db) for non-launch profiles and
hands them to agents via _transfer_db_to_agent). A child of such a
parent would write its transcript into the WRONG database — cross-
profile leakage that breaks parent_session_id lineage and
session_search. Open the dedicated handle at the parent handle's
db_path instead (AsyncSessionDB forwards .db_path via __getattr__, so
the gateway wrapper path works too). Regression test verified RED on
the pre-fix code.
…op chains

A turn writing against a session already closed by compression died with
session_persistence_failed and a misleading "this is often a full disk"
dialog, even though the store was healthy and a live continuation existed
(NousResearch#82001). Depth-1 recovery (find_live_compression_child) could not resolve
lineages with >=2 compression hops (root -> mid -> tip), reproduced
independently on two- and three-hop chains.

- run_agent.py flush chokepoint: on CompressionSessionClosedError, resolve
  tip = db.get_compression_tip(old_id) (canonical bounded transitive walk),
  adopt only when tip != old_id AND the tip row is live, retry the flush
  exactly once (adoption budget); otherwise fail closed.
- gateway/session.py append_to_transcript: replace the depth-1 live-child
  lookup with the same tip + liveness contract, so gateway transcript
  reroutes follow full chains.
- agent/conversation_compression.py _adopt_live_compression_child: turn-start
  recovery preflight now resolves via get_compression_tip with the same
  liveness check, closing the last depth-1 consumer in this family.
- classify_persistence_error: new "compression_closed" bucket; the turn-end
  explanation names compression rotation and tells the client to refresh the
  session id instead of blaming a full disk.

Tests: depth-1 adoption, multi-hop chain adoption (agent + gateway), fail
closed with no continuation / stale-closed (ws_orphan_reap) tip, exactly-once
adoption budget, and error-wording guards (compression-closed never mentions
disk; real disk failures keep disk guidance).

Closes NousResearch#82001

Co-authored-by: Al3xand3r1987 <125030427+Al3xand3r1987@users.noreply.github.com>
Co-authored-by: yuzilongleif-collab <235949691+yuzilongleif-collab@users.noreply.github.com>
…air exception paths (NousResearch#83226)

Two call sites create SessionDB instances without closing them on error:

1. gateway/slash_commands.py: /insights command - db.close() was on the
   success path but not in a finally block, so exceptions between
   SessionDB() and db.close() leak the connection.

2. hermes_cli/sessions_cmd.py: sessions repair - SessionDB() created
   inline with no .close() at all, leaking the FD on every call.

Salvage note: the original PR (NousResearch#83237) also added a __del__ safety-net
finalizer to SessionDB; review showed the atexit hook registered by
queue_token_counts() strongly retains the instance, so the finalizer never
fires for the leak class it claimed to cover. Dropped here in favor of the
deterministic constructor-finally ownership repair salvaged from NousResearch#83620.
…usResearch#83226)

SessionDB could leave native SQLite handles open when construction failed
partway through schema/pragma/FTS/repair/lock/interrupt handling. Other
short-lived callers (MCP reads/polling, session search, reactions, trace
upload, insights, shutdown recovery) opened temporary SessionDB handles
without a complete ownership boundary. API-server profile caches and
RetainDB shutdown had similar late-close races. Under sustained load this
exhausted file descriptors (EMFILE).

- Close partially initialized SessionDB connections on every constructor
  exception path via a finally block guarded by an initialization-complete
  flag.
- Close temporary/cross-profile SessionDB handles in finally blocks across
  CLI, MCP, search, trace, reactions, insights, and recovery paths.
- Add API-server per-profile cache ownership and disconnect cleanup.
- Make RetainDB writer-queue shutdown exception-safe: track connections per
  thread, close on worker exit, reject new enqueues after shutdown starts,
  and sweep any connections left by short-lived threads.
- Add regression coverage for constructor failures, worker-thread readers,
  API disconnect failures, shutdown recovery, RetainDB late enqueue, and
  foreign-loop async clients.

Salvage notes: the original PR's per-thread WAL-reader ownership changes
were superseded by main's read-connection pool (permits + checkout/return);
its cron timeout-abandon fix is credited separately to NousResearch#72822's earlier
identical fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…imeout-abandoned worker (NousResearch#72782)

run_job() submits SessionDB() to a one-worker executor and abandons the
worker (shutdown(wait=False)) when init exceeds the cron timeout. If the
constructor later completes inside that abandoned worker, the Future's
result — an open SessionDB holding .db/WAL/SHM handles — was orphaned and
never closed, leaking descriptors until EMFILE. Attach a done-callback on
the timeout path that retrieves and closes any eventual late result.

Salvage note: the lazy-recall ownership half of NousResearch#72822 (_owns_session_db
tracked on AIAgent, owned handle closed in close()) already landed on main;
this carries the remaining cron timeout-abandon half with its regression
test.
…d by read pool

- tests/cron/test_sessiondb_init_hang.py: add threading/time imports the
  salvaged late-close regression tests rely on.
- tests/test_hermes_state.py: drop
  test_close_closes_wal_read_connection_created_on_worker_thread — main
  replaced per-thread WAL reader ownership with the pooled read-connection
  design (permits + checkout/return), so cross-thread reader draining no
  longer exists in the form the test asserted.
_reap_unsupervised_gateway_orphans() kills every gateway PID found by
find_gateway_pids() on hosts without systemd (macOS launchd, Windows
Scheduled Task). This includes service-managed gateways that are NOT
orphans — they are supervised by launchd/systemd and should never be
killed during a stale-process sweep.

Add own |= _get_service_pids() to the exclusion set before scanning,
so launchd/systemd-supervised gateways are preserved. True orphans
(reparented leftovers not present in launchctl/systemctl) are still
found and reaped, preserving the NousResearch#77276 protection.

Fixes NousResearch#85344 (macOS launchd gateway killed by desktop serve startup)
Fixes NousResearch#85044 (Windows Scheduled Task gateway killed by desktop serve)
Fixes NousResearch#84855 (Permission denied to kill orphaned gateway PID)
Fixes NousResearch#85368 (gateway process repeatedly killed, messaging offline)
…per on Windows

The orphan reaper kills a healthy gateway (and its Scheduled-Task bootstrap
parent chain) every time the Desktop backend starts on Windows, because
_get_service_pids() only implements systemd/launchd and returns an empty
set on Windows — a supervised gateway is therefore indistinguishable from
an unsupervised orphan.

Exempt the recorded healthy gateway PID and its parent chain from the
orphan scan on Windows, mirroring the macOS launchd exemption (NousResearch#85913).
The Scheduled-Task bootstrap's argv matches the gateway scan, so without
exempting the parent chain killing the bootstrap takes the detached
gateway down with it.

Fixes NousResearch#86098
…r to all platforms

Compose the service-PID exclusion (NousResearch#85743, RelaxJonh) and the recorded-PID +
parent-chain exemption (NousResearch#86100, arccat-114) into one cross-platform rule:

- _get_service_pids() exclusion now runs unconditionally, not only under
  is_macos() — it is the authoritative "supervised" signal for launchd and
  any systemd unit visible on a host that got past the systemd gate.
- The recorded-healthy-gateway (get_running_pid()) + parent-chain exemption
  now runs on every platform, not only Windows. A recorded, liveness-verified
  gateway is by definition not an orphan "the pidfile/runtime record can't
  see", so the reaper must never target it — this covers Windows Scheduled
  Task / Startup VBS supervision, standalone launcher-started gateways
  (the case NousResearch#85743 alone would miss), and macOS/WSL equivalents.

True orphans (no service registration, no valid runtime record) are still
found and reaped, preserving the NousResearch#51325/NousResearch#75936 duplicate-port protection.

Existing macOS regression tests updated to pin get_running_pid to None for
their scenario; Windows regression tests from NousResearch#86100 carry over unchanged.

Bug class: NousResearch#83683 (root), NousResearch#86287, NousResearch#86098, NousResearch#85738, NousResearch#85368, NousResearch#85344, NousResearch#85044,
NousResearch#84855, NousResearch#84824, NousResearch#84200.
…ousResearch#78821)

Filter manual dashboard/serve respawn candidates after update: skip
ephemeral --port 0 backends (Desktop-owned), dedupe normalized cmdlines,
and cap one restart per profile/HERMES_HOME so orphan counts no longer
grow across successive updates.
`agent.restart_drain_timeout` defaults to 0 and governed every class of
in-flight work at once. That default is deliberate for chat turns: the
gateway announces the restart to the user and pre-marks the session
resume_pending, so interrupting one is cheap and recoverable.

A cron run has neither property. Nobody is waiting on it, it is written
to jobs.json as a permanent failure, and a recurring job simply skips to
its next schedule. Sharing the chat budget meant `_drain_active_agents()`
short-circuited on `timeout <= 0` before entering the wait loop, so the
drain reported `drain took 0.00s, timed_out=True, cron_at_start=1,
cron_now=1` — it detected the job and killed it anyway.

Cron work now drains on its own deadline, `agent.cron_drain_timeout`
(default 30s, 0 opts out). The floor is clamped to the shutdown-watchdog
leash minus a teardown reserve, so the longer wait can never consume the
post-drain cleanup window: being SIGKILLed mid-cleanup would leave the
job wedged at `last_status=running`, strictly worse than the bug. Being
bounded also means a cron-triggered restart cannot deadlock on itself.

The `timeout <= 0` special case is gone — an expired deadline expresses
the legacy "interrupt immediately" behaviour, so `timed_out` is always
computed from real state instead of asserted up front. The drain-timeout
warning now reports the elapsed wait rather than the configured budget,
which is what made "timed out after 0.0s" so confusing in the report.

Chat-only shutdowns are unchanged: `restart_drain_timeout: 0` still
interrupts chat turns immediately.

Relates to NousResearch#82161 (complements NousResearch#82195, which removes the `hermes update`
self-deadlock that triggered the reported instance).
… doubles

CI slice 5/12 caught two ways the new cron budget broke `_stop_impl_body`
for callers that are not real GatewayRunner instances:

- `_FakeGateway` in test_shutdown_cache_cleanup.py borrows `_stop_impl`
  without subclassing, so it never picked up the class-level
  `_cron_drain_timeout` default and raised AttributeError. Read it through
  the getattr-guard convention the same function already uses for its
  liveness-guard machinery.
- The same double overrides `_drain_active_agents(self, timeout)`, so
  passing the cron budget raised "takes 2 positional arguments but 3 were
  given". The double now mirrors the real optional parameter. It is the
  only override in the tree; test_startup_restart_race.py uses AsyncMock,
  which accepts any signature.

Verified against a stashed clean tree: the 22 gateway test files that
still fail locally fail identically with and without this branch (80 = 80,
empty set difference both ways) — they are pre-existing Windows-only
failures (setsid, POSIX modes) unrelated to this change.
…nect

When the shutdown drain times out and kills an in-flight cron job, the
job's owner is never told. The cron worker does try: `_is_interrupted()`
forces the failure path with an honest "interrupted by gateway shutdown"
error, and failed jobs always deliver. But that worker is a thread, it
reaches `_deliver_result()` asynchronously, and by then
`_bounded_adapter_teardown()` has closed the transport. The reporter of

Worse, the loss is silent twice over: `_consume_interrupted_flag()`
returns True — the gateway already wrote `last_status` — so
`mark_job_run()` is skipped, and the `delivery_error` from the failed
send is discarded with it. The run's only trace is a generic line in
jobs.json.

The gateway already owns the right window. `_notify_active_sessions_of_
shutdown()` runs while adapters are up, precisely so shutdown messages
can be sent — but it iterates `_running_agents`, and cron work lives on
the scheduler's own thread pool. Same structural blindness already fixed
for counting (NousResearch#60432) and draining (NousResearch#63529), never fixed for notifying.

So notify from the post-interrupt phase, which is the last point where
the transport is still up: `_kill_tool_subprocesses()` now returns the
job IDs it marked, and `_notify_interrupted_cron_jobs()` sends each one's
owner a notice on the job's own resolved delivery targets. Adapter
teardown order is untouched — it is load-bearing for NousResearch#53175 and NousResearch#8202.

Jobs with `deliver: local`, and `deliver: origin` jobs with no resolvable
origin (NousResearch#43014), resolve to zero targets and stay silent. Per-platform
`gateway_restart_notification: false` is honoured, matching the chat
path. Every failure is swallowed so a wedged adapter cannot extend
shutdown.

Second, when the interrupted flag short-circuits `mark_job_run()`, the
delivery failure is now persisted on its own via `update_job()`, so a
notice that still cannot be sent is at least recorded. `update_job()`
rather than a second `mark_job_run()`: the latter also advances
`next_run_at` and the repeat counter, and running that twice for one run
would skip a fire or auto-delete the job early.

Fixes NousResearch#82232. Related: NousResearch#82161, NousResearch#82224.
…send

mark_running_jobs_interrupted skipped legacy fires without a registered
durable owner entirely — correct for the persisted last_status write
(no owner fence to protect a replacement run), but the gateway shutdown
path also uses the returned ID list to deliver interrupted-cron notices
while adapters are still connected (NousResearch#82232). Keep the persistence skip,
but include the job in the returned list so the user is still told.
…'t erase config

`hermes import` wrote every zip member with `open(target, "wb")` followed by
`dst.write(src.read())`, at both restore sites in `run_import`. Opening for
write truncates the user's existing file to zero *before* any replacement
bytes exist, so a Ctrl-C, an ENOSPC, a corrupt zip member, or a crash leaves
`config.yaml`, `.env`, or an external provider config (e.g.
`~/.honcho/config.json`) empty with nothing behind it — during the
disaster-recovery path the user is running precisely because they already
lost something. The `_external/` branch writes outside HERMES_HOME, into
third-party configs under the user's home, so the blast radius is not
confined to Hermes state.

Both sites now stage the member into the target's own directory, fsync it,
and publish with `utils.atomic_replace`, so the target only ever moves from
its old contents to the complete new contents.

`atomic_replace` rather than a bare `os.replace`: it resolves a symlinked
target first, so deployments that link `config.yaml` into a dotfiles repo
keep the link instead of having it silently swapped for a regular file
(NousResearch#16743), and it falls back to copy/fsync/unlink on EXDEV/EBUSY for
cross-device and bind-mount installs. Members stream through
`shutil.copyfileobj` instead of being read whole into memory. The temp file
is removed on any failure so a partial import leaves no residue, and
permission bits are carried across the replace so mkstemp's 0600 does not
silently tighten restored files.

This extends the module's own established idiom — `backup.py` already
publishes atomically via `os.replace` in `_atomic_output_path` and in the
snapshot writer — into the one path that still overwrote user files in place.
…0 transit window

Follow-up on the atomic-import restore, delegating both metadata concerns to
the shared helpers instead of half-handling them locally.

Owner preservation was missing entirely. `tempfile.mkstemp` + `atomic_replace`
publishes a temp file owned by the *writing* user, so `sudo hermes import`
re-owned every restored file to root — on the disaster-recovery path, and on
exactly the Docker/NAS volume installs `utils._restore_file_owner` was added
for. `_extract_member_atomically` now captures `_preserve_file_owner(target)`
before staging and calls `_restore_file_owner` after the replace, before the
mode restore (chown clears setuid/setgid, so the mode has to go back last).

Mode handling was also only half applied before the replace: the `os.fchmod`
branch applied it to the temp fd, but the platforms without `fchmod` fell
through to a best-effort post-replace chmod, leaving the published file at
mkstemp's 0600 until that chmod landed — permanently if the process died in
between — and making `atomic_replace`'s EXDEV/EBUSY `shutil.copystat` fallback
copy 0600 onto the target. The mode is now applied to the temp file on both
branches, with the post-replace `_restore_file_mode` kept as the belt-and-
braces path.

This is the same shape `atomic_write_text` and `atomic_yaml_write` already
carry after 3556728 and 43fc865; capture and restore now reuse
`utils._preserve_file_mode` / `_preserve_file_owner` / `_restore_file_mode` /
`_restore_file_owner` rather than re-deriving them, which also drops the local
`import stat`.

Tests (tests/hermes_cli/test_backup.py, class TestImportAtomicWrites):
- test_restore_preserves_existing_file_owner — forces a uid/gid so it does not
  need root; asserts chown fires once, with the captured owner, on the
  pre-existing file only (a newly created member has no prior owner).
  Mutation-checked: dropping only the `_restore_file_owner` call reds it.
- test_mode_is_applied_before_the_replace_without_fchmod — `monkeypatch.delattr`
  on `os.fchmod`, spies the temp file's mode at replace time. Reads 0o600
  without the fix, 0o644 with it. Mutation-checked the same way.
…truncates

The atomicity claim in _extract_member_atomically's docstring holds on the
os.replace path but not on atomic_replace's EXDEV/EBUSY fallback, which uses
shutil.copyfile and so opens the destination 'wb'. That is pre-existing
behaviour shared by every atomic writer in the repo, and it is reachable here
for a symlinked target whose real file lives on another filesystem. Scope the
docstring to what the helper actually guarantees instead of overstating it;
the fallback itself is a utils.atomic_replace change.
…files

``_extract_member_atomically`` carries the replaced file's permissions
across the publish so that routing through mkstemp does not change what
the caller would otherwise have produced. But ``_preserve_file_mode``
returns ``stat.S_IMODE``, which is all twelve bits, and this restore is
deliberate on both sides of the replace: the mode is fchmod'd onto the
temp before ``atomic_replace`` and re-applied afterwards because chown
clears the elevated bits. So a target sitting at 0o4755 comes out of
``hermes import`` still at 0o4755 — with contents supplied by the zip.

That is a regression introduced by the atomic rewrite rather than a
pre-existing one. The overwrite it replaced was an in-place
``open(target, "wb")``, and an in-place write by a process without
CAP_FSETID has the elevated bits stripped by the kernel, so the old path
left 0o4755 as 0o755.

The blast radius is not limited to Hermes' own state: the ``_external/``
branch of ``run_import`` publishes members anywhere under ``$HOME``, and
this is the path that documents ``sudo`` use so ownership survives a
restore. An archive that happens to contain a member matching some
existing privileged file would take over the identity that file runs as.

Mask the two bits off the preserved mode. The masking happens once,
before the temp file is chmod'd, so there is no transient elevation
either. The sticky bit is kept — it is inert on a regular file. The
ordinary permission bits are unaffected, so the Docker/NAS installs the
preservation exists for still get their broader modes back.

This is the one write path in the repo where the bytes are untrusted;
the ``utils`` writers that preserve the full mode re-serialize content
the process itself produced, and are correct as they stand.
… bits

Pre-creates a 0o6755 target, imports a member over it, and asserts the
published file is 0o755 with both elevated bits gone — plus that the
staged temp file never carried them either, so there is no window where
archive content sits behind an elevated mode.

The existing coverage in this class cannot see the failure: every mode
assertion masks with ``& 0o777``, which discards exactly the bits at
issue, and the fixtures chmod their targets to ordinary modes that never
had them set. Without the mask on the preserved mode this test reports
the published file still holding S_ISUID.

Skipped where the platform or filesystem refuses setuid on a user-owned
file, so the assertion never depends on running as root.
Node 22 ships npm 11.16.0, which engines.npm rejects (11.10–11.16
ignore min-release-age-exclude). Fresh Hermes-managed installs then
fail npm ci with EBADENGINE. Node 26 ships 11.17.0.
…wn (macOS)

The Desktop-spawned hand-off consistently died during Electron's quit
teardown on macOS: the orchestrator process group was terminated right
after `running: hermes update ...`, so no exit code, result file, bundle
swap, or relaunch ever happened, and the loopback shim window surfaced
the death as ERR_CONNECTION_REFUSED or "Aw, Snap!" error code 15
(reproductions in NousResearch#66753).

- Re-exec the orchestrator through a one-shot setsid child and let the
  direct Electron child exit immediately; the real orchestrator is owned
  by launchd (PPID 1), outside Electron's teardown, same marker/result
  protocol.
- Hold TERM ignored across the `hermes update` invocation and
  log-and-ignore the single teardown TERM that can still arrive after
  the desktop PID dies (durable SIGNAL breadcrumb for diagnosis).
- Delay start_ui until the desktop PID is gone plus 1s so the shim
  server/window are never born inside the teardown window.
- Run both UI processes in their own sessions; keep SIGTERM/SIGHUP
  ignored in the shim server and stop it with SIGKILL, so a stray TERM
  can no longer leave the progress window on a dead loopback URL while
  the update continues.

Verified on a production git install (macOS arm64, Darwin 27.0,
v0.20.1): six consecutive Desktop-triggered/production-shape updates
completed end-to-end including a full desktop rebuild + codesign; the
shim survived a deliberately injected TERM+HUP mid-update and a full
`hermes desktop --force-build --build-only` running alongside it.

Fixes the macOS reproductions in NousResearch#66753.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…de-newer bricking

The relative exclude-newer = "14 days" cutoff bricks installs whenever the
resolver cannot see (or accept) a package's upload date:

- defusedxml / python-olm / unpaddedbase64 (NousResearch#80387, NousResearch#79434): ancient frozen
  releases (2021-2023) whose upload dates are often absent from mirror
  indexes and stale uv HTTP caches. uv then filters them entirely
  ("there are no versions of defusedxml"), breaking [youtube]/[wecom]/
  [matrix] resolution and daily `uv sync --locked` runs.

- setuptools / pillow / mcp (NousResearch#78227, NousResearch#75992, NousResearch#76020): exact-pinned deps.
  When the pinned version's upload date is invisible, the resolver filters
  the ONLY acceptable candidate — setuptools==83.0.0 in
  [build-system].requires meant the project could not even be built from a
  git checkout on released v0.20.0. Exempting an exact pin costs nothing:
  the version cannot float without a reviewed pin bump.

Changes:
- pyproject.toml: add all six to the existing exclude-newer-package
  whitelist, with rationale comments per class.
- uv.lock: regenerated; diff is the whitelist metadata only (verified
  zero version drift, still 249 packages).
- tests/test_packaging_metadata.py: new standing guard
  test_build_system_requires_exempt_from_exclude_newer — every
  [build-system].requires package must be whitelisted while a relative
  exclude-newer cutoff is configured. Verified both directions (fails
  when setuptools is removed from the whitelist).
- scripts/install.sh: fix the stale tier-name comparison ("all (with
  RL/matrix extras)" vs actual "all") that mislabeled every successful
  Tier-1 install as a fallback-tier install (NousResearch#79434 bonus finding).

Verification: uv lock --check green on uv 0.11.19 and 0.12.5;
uv sync --extra all --locked green; uv pip install -e '.[all]' resolves;
whitelist mechanism A/B-proven on a minimal project (unsatisfiable ->
resolves; build-requires variant: uv build fails -> succeeds).

Reported-by: MichaelClawHub (NousResearch#80387), liujianqiu (NousResearch#79434), maxonliu (NousResearch#78227)
…script

Session hygiene could persist a compressed transcript LARGER than the
original (observed: 427K -> 598K), when the generated summary was bigger
than the middle it replaced. Compare like-for-like (both rough estimates)
before persisting a rotated transcript; on growth, keep the original
unchanged so a failed compression is a strict no-op, never a net increase.
… (in-place path)

The gateway rotation guard (NousResearch#83339) only protects the rotate path, but
in-place compaction commits inside compress_context() via
archive_and_compact — before the gateway can inspect the result. Add the
anti-growth check at the commit site so both paths are covered: a
compression whose rough output exceeds its input is a strict no-op
(original transcript kept durable, session identity untouched).

Covers the observed failure where session hygiene persisted 426 -> 426
messages and ~379K -> ~688K tokens.
…ment

Simplify-code pass: gateway/run.py called estimate_messages_tokens_rough
6x on the same data in the anti-growth guard (condition + warning f-string).
Bind to _hyg_in_toks/_hyg_out_toks locals like the conversation_compression.py
guard already does. Also trim the comment from 10 lines to 4 (keep the WHY,
drop the WHAT) and remove an extra blank line before TestCompactedTurnsStaySearchable.
mjolley9 and others added 25 commits August 15, 2026 00:35
…tion

Residual hunk from PR NousResearch#84832's picker-rebind commit; the functional
surfaces.tsx change landed on main via NousResearch#86250.

(cherry picked from commit 990936d, docstring hunk only)
Read the durable display transcript when creating a branch instead of copying the compacted model projection. Hydrate the Desktop branch boundary from persisted history, avoid stale whole-chat counts, and seed the new tile from the backend snapshot. Add regression coverage for compacted histories, visible-message counts, selected prefixes, and hydration races.

(cherry picked from commit c3d2d75)
…ry CLI start

The main.py decomposition re-exported the sessions/update/dashboard command
surface with eager from-imports, so every hermes invocation (including
hermes --version) paid for update_cmd's dependency chain (jwt, click,
cryptography). Resolve the re-exports through the existing PEP 562 module
__getattr__ (same pattern as _PROVIDER_MODELS) so each module loads on
first actual use. Internal call sites go through a _self() helper because
bare-name lookups do not trigger __getattr__; _self() imports sys locally
since update tests patch hermes_cli.main.sys. The
_warn_stale_dashboard_processes back-compat alias moves into the lazy
surface, and the sessions argparse dispatch defers sessions_cmd to call
time. Monkeypatching hermes_cli.main.<name> keeps working: a patch sets a
real module attribute, which shadows __getattr__.

Measured (Windows 11, Python 3.11, median of 7 warm runs):
import hermes_cli.main 253ms -> 196ms (-22%).

(cherry picked from commit cad1083)
The /tools and /personality completers run on every keystroke while the
user types those commands (complete_while_typing), and both re-read +
re-parse the full config on each keypress:

- _tools_completions called load_config() — the defensive deepcopy
  (~340us/call on cache hit) even though it only reads toolset enable
  state + MCP server names. Switched to load_config_readonly() (the
  perf(agent) NousResearch#74322 pattern; this per-keystroke site was missed).
- _personality_completions called load_cli_config() — a full YAML parse
  + deep merge of the built-in defaults (~110us) — on every keystroke.
  Memoised keyed on the config file path+mtime (same pattern as load_env
  / _nous_auth_status_cache), so the parse runs once per config state.

Measured: /tools 357us -> 18us per keystroke; /personality parse drops
from 1-per-keystroke to 1-per-config-change (500 keystrokes -> 1 parse).

Regression tests: _tools_completions uses the readonly loader (deepcopy
loader never called); personality memo parses once across repeated
completions and re-parses once after a config mtime change.

(cherry picked from commit 2b3f897)
…re read

get_default_hermes_root() resolves HERMES_HOME against the platform
native home (~80us of path resolution) on EVERY call and is called at
31+ sites — every _load_global_auth_store() (per provider row in the
/model picker), kanban, backup, gateway, update. Its result depends
only on (HERMES_HOME, native home), so memoise it keyed on those two
inputs, compared for free on each call (freshness-correct even if a
test or plugin mutates HERMES_HOME mid-process).

_load_global_auth_store() re-read + re-parsed the global auth.json on
every call; read_credential_pool() -> load_pool() runs it once per
provider row in the /model picker even when the profile has entries and
the global fallback never fires. Memoise keyed on the global auth
file's path+mtime (same pattern as _nous_auth_status_cache); the store
only changes when a global-scope auth write touches the file.

Measured (profile mode, 30-provider global store): get_default_hermes_root
81us -> 10us; _load_global_auth_store 128us -> 66us; load_pool 165us ->
137us per call — ~2ms saved per /model picker render (20 provider rows).

Regression tests: hermes_constants memo pin (no path resolution on
repeat calls, HERMES_HOME change forces a fresh resolution); global-store
memo pins (store read once across repeats, mtime bump re-reads once,
absent store stays cheap).

(cherry picked from commit be348f3)
resolve_toolset() recursively walks toolset includes and, with
include_registry=True, merges registry-registered tools on every call —
each external call re-runs the includes walk and takes a fresh registry
snapshot under the registry lock. It is called dozens of times per
_get_platform_tools() (every /tools completion keystroke, per picker
render) and at 19 call sites across the CLI/gateway.

Memoise the external-entry result keyed on (name, include_registry,
registry id, registry generation). tools.registry exposes a monotonic
_generation counter bumped on every register/deregister/alias/MCP
refresh (its docstring explicitly invites generation-keyed memoisation),
so a cache entry is valid until the registry changes. External callers
never pass visited, so the memo engages exactly at the public entry
and the internal cycle-detection recursion is untouched.

Measured: _get_platform_tools drops 165us -> 59us per call (3x) with the
xAI credential fix simulated; /tools completion ~2ms -> ~77us/keystroke
combined. Regression tests: repeat resolution is a memo hit (get_toolset
called once), a generation bump forces a fresh resolve, and the memoised
result is identical to a fresh resolution.

(cherry picked from commit 3d36ecb)
_file_mutation_verifier_enabled and _turn_completion_explainer_enabled
re-read config.yaml on every call via load_config() (~1ms deepcopy per
call). finalize_turn runs these gates at the end of every turn, so each
turn paid two redundant config deepcopies. The sibling
_credits_notices_enabled already caches on self; mirror that pattern.

The env-var override stays authoritative and uncached, so runtime flips
still work. Config flips now apply on the next session, matching the
documented sibling semantics.

(cherry picked from commit f2d0e00)
…olve_toolset memo

- The readonly-loader completer test now stubs
  get_portable_mcp_server_names_nowait — real plugin discovery runs
  load_config() during one-time process init, which is not the
  per-keystroke read the test guards against.
- Cap _resolve_toolset_memo at 256 entries: generation-keyed entries
  from stale generations are never hit again, so clear on overflow to
  keep long sessions bounded.
Addresses both review findings on the remote-gateway download PR:

1. Unbounded buffering (finding #1). fetchBuffer / fetchBufferViaOauthSession
   accumulated the entire response (then copied it again via Buffer.concat)
   before saveGatewayFile even opened the save dialog, so a large gateway file
   could exhaust the native process. Both auth paths now stream: once response
   headers arrive the connect timeout is cleared, the filename is derived, the
   save dialog is shown, and the body is piped to the chosen destination with
   backpressure. A read/write error tears down the stream and unlinks the
   partial file. The byte-moving, data-URL decoding, and filename/path helpers
   are extracted into gateway-file-download.ts so they're unit-testable without
   Electron.

2. No fallback for older gateways (finding #2). saveGatewayFile required the new
   /api/fs/download route. Desktop and the remote gateway update independently,
   so a gateway predating this PR 404s. Added a 404-only compatibility fallback
   to the existing capped /api/fs/read-data-url route (bounded, so it only
   serves smaller files — enough to keep older backends working).

Tests: gateway-file-download.test.ts covers streaming, backpressure,
error-cleanup (unlink on write/response error), data-URL decoding, filename
derivation (incl. traversal reduction), and 404 detection;
gateway-file-download-transport.test.ts asserts both transports stream (no
whole-body Buffer.concat) and that the 404 fallback is wired. Both registered
in the desktop platform test list. Server-side /api/fs/download tests
(streaming + sensitive-file reject) already pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…oject

The salvaged branch predates the electron test project's node:test -> vitest
migration (test:desktop:platforms is now `vitest run --project electron`).
Import `test` from vitest so the suites are collected; assertions stay on
node:assert/strict per the existing electron test convention.
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
…hrough af585de)

Clean merge - no conflicts. Replay (git merge-tree) produced zero
conflicts and the committed tree is byte-identical to the replay tree.

Semantic verification on the merged tree (deeper than usual given the
batch size - 375 commits, the largest sync to date):
- drift grep for 'Hermes Desktop' in apps/desktop/{src,electron}: empty
- every brand carve-out value intact: productName/executableName
  'NousAI', appId ai.nous.desktop, artifactName NousAI-..., APP_NAME
  default, WORDMARK, DEFAULT_SKIN_NAME 'nousai', index.html title,
  brand-mark asset, one nousai dashboard theme row, brand-identity.test.ts
- hermes_cli/main.py's brand-agnostic macOS packaged-app lookup survived
  12 upstream commits to that file byte-for-byte
- apps/desktop/package.json (carve-out) differs from our branch ONLY by
  upstream's two new repro:short-session-hang scripts - nothing reverted
- no new i18n locales (still ar/en/ja/zh/zh-hant), so no rebrand sweep
  needed; upstream's new e2e specs pin only toContain('Hermes'), which
  our '<title>NousAI - Hermes</title>' satisfies by design

No CI-sensitive workflow files and no package-lock.json changes, so
neither review gate applies. uv.lock changes are exclude-newer-package
booleans only (no dependency version bumps).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KFEJ7TzKwG4tQujNy3CWjT
@github-actions

github-actions Bot commented Aug 15, 2026 •

Copy link
Copy Markdown

૮ >ﻌ< ა ci review

ran on 4df11ee — fix(tests): align lost-and-found sessions pin with the 56-co

⚠️ Warnings

OSV vulnerability scan · View job

5 known vulnerabilities found in pinned dependencies.

How to fix:

Review the findings in the Security tab. Update the affected dependencies if a patched version is available.


debug info

CI timings

CI timings · View report · View job

Wall time 4m10s vs 3m50s (+8.7%). 15 job(s) slower, 18 faster, 4 unchanged.

  • JS & TS checks / apps/desktop / check:test:ui:shard-1of3: -26.0s
  • Python lints / ruff enforcement (blocking): +21.0s
  • Python tests / Run tests slice 3/12: -16.0s
  • Python tests / Run tests slice 8/12: +15.0s
  • Python tests / Run tests slice 7/12: -13.0s

Upstream's d16326b bumped this test's sessions-schema pin for the
new git_metadata_generation column (54 -> 55), but landed on a tree
that already carried the 'hidden' column from fbaea9b, so the real
layout is 56 and upstream main is red on
test_mapper_rebuilds_sessiondb_from_synthetic_lost_and_found
(assert 56 == 55). Verified byte-identical to upstream and reproduced
locally.

Bump the width in the four coupled places the value is load-bearing -
the pin, max_fields (the lost_and_found table width), the
current-layout insert(), and its comment - mirroring exactly what
d16326b itself did. Bumping the pin alone would only move the
failure into the row builders. Verified locally: pins pass and the
56/52/14/23-field inserts are all accepted.

The correct value is objectively fixed by the schema (this is a
schema pin, not a product-intent choice), so when upstream re-aligns
it the next sync takes upstream's version of these hunks. Recorded in
CLAUDE.md (user-approved 2026-08-15).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KFEJ7TzKwG4tQujNy3CWjT
@uaixo
uaixo marked this pull request as ready for review August 15, 2026 08:33
@uaixo
uaixo merged commit b7b0e05 into NousAI-Assistant Aug 15, 2026
56 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.