Skip to content

Merge latest NousResearch/hermes-agent main - #5

Merged
Nag112 merged 279 commits into
mainfrom
cursor/merge-upstream-hermes-latest-b024
Sep 18, 2026
Merged

Nag112 merged 279 commits into
mainfrom
cursor/merge-upstream-hermes-latest-b024

Conversation

@Nag112

@Nag112 Nag112 commented Sep 18, 2026

Copy link
Copy Markdown
Owner

What does this PR do?

Merges the latest NousResearch/hermes-agent main (e83b1d51f1) into this fork.

This is the next upstream sync after PR #4. Git merged cleanly (no conflicts). ~278 commits / 514 files.

Open PR #3 (NousResearch:main → Nag112:main) is the same direction; once this lands, close #3.

Type of Change

  • ♻️ Refactor (no behavior change) — upstream sync

Changes Made

  • Merged upstream/main through e83b1d51f1.
  • Fork-only skills kept: skills/productivity/calendar/, skills/productivity/vikunja/.
  • plugins/google_meet and plugins/spotify stay removed.

How to Test

  1. After merge, fork main should match current Hermes plus the two productivity skills.
  2. Close Merge from Nous #3 once this lands.

Checklist

  • Clean merge against current NousResearch main
  • Calendar and vikunja skills preserved
  • google_meet/spotify plugin removals preserved
Open in Web Open in Cursor 

ayushnangia and others added 30 commits September 17, 2026 20:33
…enu and the keybind

Follow-up to Ayush Nangia's feature: the filter menu kept its own literal
list of groupings while the new cycle keybind walked a second one, with only
a test comment saying they should agree. The menu now builds its options
from `SIDEBAR_GROUPING_ORDER`, so reordering or adding a grouping is one
edit and the keybind cannot drift from what the menu shows.
…variant cycle test

`SidebarGrouping` was a literal union kept in step with `SIDEBAR_GROUPING_ORDER`
by hand — a fifth grouping added to the type would compile while the cycle
silently skipped it. The array is now the single source (`as const`) and the
type is derived from it. The cycle test no longer pins the exact order (a
change-detector on the constant); it asserts one lap visits every grouping
once and returns to the start, and still fails when the cycle skips a step.
NousResearch#113535)

_command_detection_variants yielded one FULL-LENGTH variant per quoted or
escaped command word. A heredoc body of quoted lines ("key": "value", …) is
hundreds of quoted command words, so both detection passes scanned
O(words * len) characters: on a 16 KB / 460-line command the hardline pass
alone took 7.8 s and the dangerous pass 15 s even with the launchctl lookahead
anchored (33 KB: 36 s hardline), all while holding the GIL on the gateway loop.

Build a single variant with every command word deobfuscated instead. The
obfuscation catch ($(echo rm), r''m, ${0/x/r}m …) is unchanged — the same
deobfuscated words appear, in one string — and the variant count on the
issue's input drops from 923 to 4 (36 s -> 0.26 s hardline, verdict
unchanged). detect_hardline_command needs no bound and keeps its YOLO
semantics: a benign heredoc is neither blocked nor prompted.

Root cause identified in NousResearch#113943 by @Tranquil-Flow (variant explosion as the
second compounding factor); its cumulative-work budget is superseded by
removing the explosion.

Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
…er test

A bare `python -c` resolves `tools` through the venv's editable install (the
primary clone), not the worktree pytest runs in, so the regression guard could
silently test a different tree. Pass cwd + PYTHONPATH like
tests/tools/test_async_delegation.py does.
…ange

load_env_file() is now the canonical .env tokenizer after c849bc3 collapsed
six hand parsers onto it. Profile scopes, dashboard boundary discovery, skill
secret capture, managed .env and setup readers share the same decoding and
parsing rules, but direct calls still re-read and re-parse the whole file.

build_profile_secret_scope() puts that work on the gateway's hottest paths:
every turn, every cron job, MCP and browser lifecycle adoption, the TUI prompt
turn, and the 60s housekeeping drain of the durable cron delivery queue.
Memoising the shared tokenizer now removes redundant work for the consolidated
readers as well as the motivating profile-scope path.

config.load_env() keeps its existing outer memo over a call that is now itself
memoised; this leaves its menu-render shortcut intact without redesigning it.

Freshness is the hard part, because a stale credential map is a far worse
failure than a slow one. Every load_env_file() call still OPENS the file and
keys on os.fstat() of that descriptor rather than os.stat() of the path:

  * the open preserves close-to-open revalidation on NFS;
  * an open or read failure returns {} without caching it, and drops any warm
    entry so a transient EACCES cannot become a persistent empty profile;
  * the descriptor pins one inode, so a symlink repointed mid-read cannot file
    one file's contents under another file's identity.

The key is (mtime_ns, size, inode, device), re-checked on the same descriptor
after reading. A rewrite through that inode is not stored under the pre-read
fingerprint. Unavailable metadata means "do not cache", never "cannot read".
A generation counter checked under the lock prevents a slow reader from
repopulating an entry that a writer invalidated while the read was in flight.

Read bytes from the open descriptor and preserve upstream's decoding exactly:
strip a UTF-8 BOM before trying UTF-8, then fall back to latin-1 for invalid
UTF-8. Both the cached reader and the uncached reference use that decoder and
the shared text parser. Invalid UTF-8 is a cacheable result, not a read error.
The value and inline-comment helpers remain local to secret_scope.

invalidate_env_cache() clears this memo too, so save_env_value,
remove_env_value and sanitize_env_file invalidate both layers. The per-path
LRU holds at most 64 entries and every caller gets its own dict.

Accepted limitation: a write keeping length, inode and nanosecond mtime all
identical is not detected without explicit invalidation. This is a new
limitation for previously uncached readers. Parsed secrets also remain
reachable in memory between use and eviction for up to 64 paths, where the
outer config memo holds one.

The pre-rebase measurement on a 508-line, 25432-byte .env was 0.868 ms ->
0.011 ms per call. That measurement predates the shared binary decoder.

Add decode/cache invariants for BOM and plain files, latin-1 fallback,
same-size writes switching UTF-8 -> latin-1 -> UTF-8, and agreement with the
uncached reference. All eight decoding, cache and freshness mutations fail
the new tests, and all 40 cache tests pass after restoration.

The requested agent, CLI, cron and gateway suites report 30,168 passed,
479 failed and 303 skipped in the restricted local environment. All 140
failing files were checked against upstream/main at 5eb99eb: all 479
failed-test IDs and 210 collection/setup/teardown error IDs match exactly.
…sted

The 5s goals poller calls restore_heartbeat_watches(), which listed every
routed session and entered _profile_runtime_scope for each distinct origin
— a full config.yaml/.env parse plus secrets hydration per session per
scan. With zero heartbeats persisted (the common idle case) the entire
sweep is wasted work; on a 400-session multiplex gateway it dominated
idle CPU (~60% of on-CPU time in YAML parsing).

Gate the sweep on an indexed heartbeat:* prefix probe of each served
profile's SessionDB (HERMES_HOME contextvar only — no config/secret
scopes). Probe failures fail open to the historical full sweep.
…ofile's store

The handoff watcher (2s interval) and loop wakeup watcher (15s) entered
_async_profile_runtime_scope for EVERY served profile on EVERY tick — a
full config.yaml/.env parse plus secrets hydration per entry — even when
the profile's state.db held zero pending handoffs or active loops. On a
multiplex gateway this, not message traffic, was the idle CPU driver
(py-spy: ~75% of on-CPU time in YAML parsing from these two paths).

Both watchers now probe the profile's cached SessionDB first (only the
HERMES_HOME contextvar installed, off the loop thread via the runner's
executor hop) and enter the scope only when work exists:
- handoff watcher: new SessionDB.has_pending_handoffs() existence probe
- loop watcher: list_active_loops() under the probe context

All probes fail OPEN (error/unknown -> historical always-enter
behavior); runners without the executor hop (test stand-ins) skip the
gate, so existing watcher tests are unaffected. The startup
stale-handoff reclaim stays ungated: it runs once per boot and must
also see 'running' leftovers, which a pending-only probe would miss.
…nd ignore cleared rows

The loop watcher's probe used ``list_active_loops()``, which collapses an unavailable
SessionDB (and any read error) to ``[]`` — the gate then skipped that profile on every tick
(fail CLOSED). The heartbeat probe keyed on ``heartbeat:*`` key existence, but ``clear`` and
``pause`` keep their rows, so one past /heartbeat re-enabled the full sweep forever.

Both probes now read the rows through one ``_profile_meta_rows`` helper (None = store
unavailable → treat as work present) and parse the status themselves: only an ACTIVE row
opens the gate; a corrupt row is "unknown" and keeps the historical scan. The heartbeat
gate reuses ``_handoff_watch_scopes`` for the served homes instead of a second resolver.
Heartbeat: no active row → no sweep; a cleared row → still no sweep; an active row → sweep;
an unopenable store → sweep. Loop watcher: idle profile never scoped, active loop scoped,
unopenable store scoped. The handoff-watcher test from NousResearch#109497 is dropped: same shape, and
the salvage bar is two tests per fix.
The memo had its own (mtime_ns, size, inode, device) tuple and documented a
same-length/same-inode/pinned-mtime rewrite as an accepted gap. utils.file_signature is the
repo's change-detection key (config.load_env, skill_utils, prompt_builder already use it) and
includes ctime_ns, which user space cannot backdate — so the gap closes and one helper goes.
load_env() kept its own (path, file_signature) memo on top of load_env_file(), which now memoises
on the same key: two stat calls and two caches to keep coherent for one answer. Drop the outer
layer; invalidate_env_cache() stays as the writers' knob and forwards to the one memo.
…named-profile multiplexer

_watched_homes reused _handoff_watch_scopes, which deliberately omits "default" because the
handoff/loop watchers' root poll covers the launch home. For a `hermes -p work` multiplexer the
launch home is profiles/work, so ~/.hermes was probed by nobody and an active heartbeat there
answered "idle" every tick. Probe the served set the sweep can actually resolve to.
…ive with loops/heartbeat

The probe helpers had been appended to the gateway/run.py facade and the three watchers each
carried their own copy of the "probe → fail open → offload" shell, with gateway code parsing
loop/heartbeat rows itself. One sibling now owns the gate shape (_gate + off_loop_gate), and
hermes_cli.loops.store_has_active_loop / hermes_cli.heartbeat.store_has_active_heartbeat own
what an ACTIVE row is (heartbeat gains the _META_PREFIX loops already had). One fail-open layer
per read: has_pending_handoffs is a bare bounded query, the gate catches. The loop watcher calls
its executor hop directly — it already did so unconditionally for the scan, so the
"runner without a hop" fallback there was dead.
…ome case and the handoff gate

Loop watcher: a cleared row must not open the gate (a key-existence mutation now fails), then
the unavailable-store fail-open. Heartbeat restore: routing home = named profile with the only
active row in the default store must still sweep. The handoff-watcher test from NousResearch#109497 is
restored (its gate and the has_pending_handoffs query had no coverage).
… update

Port of NousResearch#59942 onto post-NousResearch#102117 main. The refactor moved the desktop
helpers from hermes_cli/main.py into hermes_cli/main_desktop.py, so the
new install machinery (_app_asar_hash, _macos_adhoc_sign_bundle,
_desktop_bundle_install_supported, _swap_in_new_macos_bundle,
_install_rebuilt_desktop_app) lives there now, re-exported through
main.py's frozen updater surface block so _m() resolves them.

_rebuild_desktop_after_update (hermes_cli/update_cmd_deps.py) now, on a
successful --build-only build, installs the rebuilt bundle to the
system location (macOS .app via ditto staged copy + atomic swap with
rollback; Windows via copytree), skipping when the installed app.asar
hash already matches. Linux stays owned by its package mechanism.

Without this, a CLI hermes update leaves the installed app stale —
the in-app updater covers its own path, but the CLI path never swapped
the installed bundle.

Tests: 17/17 ported suite (import/patch targets repointed to
main_desktop) + test_cmd_update.py 47/47 green.
…gnature, never swap under a live app

Rework of the salvaged NousResearch#59942 hook so it fixes the whole NousResearch#52339 class without
regressing what the Desktop updater already gets right:

- macOS only. The Windows arm rmtree'd a possibly-running NSIS install
  (partial deletion under a lock) and windows.ps1 already owns that swap;
  Linux packages stay with their package manager.
- No re-sign. The rebuilt release/ bundle already carries the stable local
  signing identity from _desktop_macos_relaunchable_fixup; a deep `codesign -s -`
  on the installed copy replaced it with a fresh ad-hoc cdhash and reset every
  TCC grant. ditto preserves the signature, so nothing is signed here.
- Running bundles are reported, not swapped: Electron loads app.asar and helper
  apps lazily, so renaming the bundle away and deleting the old tree crashes
  the live app. The detached updater waits for exit; a terminal `hermes update`
  with the app open now prints what to do instead.
- Failures are printed as warnings; the old `Path | None` return read every
  failure as "Desktop app up to date".
- The refresh also runs on the "build stamp current" path, so a stale
  /Applications copy left by an earlier update heals on the next `hermes update`
  even when there is nothing to rebuild.
- Core is host-independent (_install_rebuilt_macos_bundles takes paths as data);
  the two invariant tests run on every OS instead of `skipif(darwin)` tests that
  ran nowhere.

Covers the Desktop-button path too: posix.sh runs `hermes update`, so an app
running from apps/desktop/release/ now refreshes the /Applications copy Finder
launches (the Discord report: new shell right after the update, old shell on
the next Dock launch).
Cherry-pick of helix4u's NousResearch#88836 (75742db) resolved onto current main.

The POSIX runtime cutover parks the live venv with a directory rename. On
Windows that can never succeed from `hermes update`: the updater executes
from venv\Scripts\python.exe and Windows keeps that image mapped, so the
rename fails with WinError 5 on every attempt (NousResearch#93032). Instead, atomically
replace the live venv's pyvenv.cfg with the candidate's: Scripts\python.exe
is a launcher that reads `home` on every start, so every fresh process runs
the new generation while nothing touches the locked directory. A failed
post-cutover smoke restores the original config; a failed rollback keeps
the generation the live config now references.
The self-lock/holder preflight (NousResearch#99711) deferred the repair on the theory
that the updater's own mapped python.exe makes the venv rename impossible.
Live on windows-latest, a process executing from venv\Scripts\python.exe
(and the repo's real .venv with cp311 .pyd extensions loaded) does NOT block
the rename; what blocks it with WinError 5 is any ordinary handle under the
tree: a process cwd, an open file, a sync client. Since sys.executable is
always under the venv on Windows, the deferral fired on every run and no
Windows install could repair from `hermes update`. Retire it and its tests;
the pyvenv.cfg repoint needs no rename and has none of that exposure. The
wine2e lane now runs the cutover probes instead.

The repoint keeps the live venv's site-packages, so provisioning's
fall-forward to the next minor (3.11 -> 3.12, NousResearch#76106) must not be pointed at
a cp311 tree: refuse before touching pyvenv.cfg. The live Windows test holds
the venv the way the field does (a child with cwd inside it), asserts the
rename path fails (the symptom) and the repoint succeeds; the rollback and
minor-guard tests run on every host.
…ousResearch#85194)

`auxiliary.title_generation.enabled: false` turned off both title stages, so an
operator who only wanted to stop the background model call (unavailable or
metered endpoint) also lost the instant derived title. `model_upgrade_enabled:
false` keeps the derived title and starts no `auto-title` thread; missing keeps
the two-stage default and `enabled: false` still disables both.

Salvaged from NousResearch#85401 onto current main: the gate reuses `_title_config()` and
sits before `spawn_context_thread` (the thread seam moved off `threading.Thread`).
Non-noreply author email on the salvaged NousResearch#85401 commit.
…rade

The gate also sat inside `generate_title`, so the explicit operator repair
command `hermes sessions retitle-skills` returned None for every row when the
toggle was off, although the operator asked for a model call. The toggle's
contract is "never spend a model call upgrading the instant title": only
`maybe_auto_title` (the background path) is gated now; a direct
`generate_title` call still asks the model. Docs sentence adjusted.

Review follow-up on NousResearch#113955; the existing toggle test now also asserts the
explicit path still titles (red on the previous head).
…usResearch#72351)

Title generation hardcoded temperature=0.3 in its call_llm() call.
Models like GPT-5.6 only accept their server-side default temperature
and reject explicit values with "Unsupported value: 'temperature'".
While call_llm() has a retry that strips temperature on error, the
daemon thread races with session cleanup in short-lived CLI sessions,
causing the retry to fail with a connection error.

Fix: pass temperature=None so the provider uses its own default.  This
avoids the unsupported-temperature error entirely and eliminates the
need for the retry path.

Fixes NousResearch#72351
…ack requests

Reasoning models reject several request fields at once (gpt-5: temperature AND
max_tokens, NousResearch#78273), a reasoning-strip retry can then 400 on temperature
(NousResearch#72351), and strict-schema gateways such as Fireworks reject the generic
extra_body.reasoning fallback with "Extra inputs are not permitted, field:
'reasoning'" (NousResearch#109774) — a phrasing none of the unsupported-parameter predicates
recognised, so the reasoning rung never fired and every title call 400'd.

The parameter rungs were a fixed single-pass order (temperature → structured
output → reasoning → max_tokens): a field rejected AFTER an earlier rung had
already run was never stripped, and _param_rung_accepts did not admit a
temperature 400 raised by a later retry. They now form a table walked until no
rung matches, each field stripped at most once, so N rejected fields recover in
N retries in whatever order the provider raises them.

Fallback candidates (per-task chain, main chain, discovery) got a single shot
with the caller's temperature/max_tokens/reasoning fields and raised on the
first parameter 400; agent/auxiliary_fallback_recovery.py now runs the same
parameter rungs around a candidate's request (sync + async), leaving auth /
payment / connection errors to the caller's existing handling.

Live: real gpt-5-mini title_generation call (reasoning_effort 400 → temperature
400 → 200), real gpt-5-mini fallback candidate sync+async (temperature 400 →
200), and a stand-in replaying Fireworks' documented 400 (reasoning stripped →
200) — all failed on origin/main.
A positional _LadderRoute(...) in auxiliary_fallback_recovery breaks as soon as the
route tuple gains a field (NousResearch#113968 adds timeout); construct by name so fields the
parameter ladder never reads default to None regardless of tuple width.
…lds the route by name

base_info is the endpoint identity every rung reads for per-route bookkeeping; the
fallback ladder left it empty. The test mirrored the positional construction that
broke under a wider tuple.
teknium1 and others added 27 commits September 17, 2026 12:08
…ications

Squash of the 54 commits on victor-kyriazakos:feat/user-channel-warning-suppression
(PR NousResearch#112302, head f45c640) so the contributor's authorship survives a rebase-merge;
the commits interleave with a cron delivery-ledger rework that the salvage removes in
follow-up commits, so per-commit cherry-picks were not practical.

Adds display.suppress_warning_notifications (global + per-platform, default false):
one resolver (gateway/warning_notifications.py), BasePlatformAdapter.emit_warning /
emit_media_warning / warning_text, a notification_category classification carried
through wakes, queues and persistence, and render/present boundaries for CLI/TUI.
…nifest ledger rework

The suppression feature needs one cron cell: a failure notice whose target hides warning
notifications is recorded as `suppressed` (bot_chat_pending record, deliveries queue row,
execution delivery_outcome) instead of being sent. That is kept.

Everything else the PR added to cron/ is an independent ledger rework and is removed here:
execution delivery manifests + filesystem manifest journal, `schema_meta` fence and the
"legacy intent adoption" layer, incident occurrence generations + trigger, jobs projection
CAS (`bind_delivery_execution`/`update_delivery_projection`), the per-tick projection
reconciler, and the `_deliver_targets`/`_settle_manifest` split. origin/main has none of
the state that layer migrates from (zero occurrences of delivery_manifest / manifest_journal
/ schema_meta); the "pre-flag old writer" the r5-r9 tests simulate is a vendored snapshot of
this PR's own earlier revision (tests/cron/_r5_prior_executions.py). Those fixes may have
merit on their own and should land as separate PRs with a main-reproducing test each.

Also restores main's `_classify_delivery_outcome` precedence (`failed` before `queued`)
and drops the SHA-pinned `git show <PR commit>` test, which would go red the moment a
rebase-merge rewrote that commit.
Four places changed behaviour for users who never touched the setting:

- `_interim_send` was stamped on every `warn` status and media-failure notice, and the
  Slack/relay egress doors learned to skip stream sealing for it. Main's status sends carry
  no interim mark at all, so the gap is class-wide (every status kind), and fixing it for
  warnings alone is an undeclared streaming-contract change. Reverted here; the whole-class
  fix belongs in its own PR against gateway/AGENTS.md rule 3.
- The entire post-handler delivery (unwrap, TTS, final text, attachments, delivery-ledger
  writes) ran inside `_media_delivery_scope`. Under multiplex that binds the routed home, so
  delivery obligations landed in the routed profile's state.db while boot-time
  `_claim_pending_obligations` still reads the launch home. Only the policy reads
  (`diagnostic_wake_muted`, `warning_text`) bind the routed scope now; delivery stays where
  main ran it.
- The turn-crash notice is rebuilt the same way: scope around the policy read, send outside.
- The "delivery failed after multiple attempts" notice is unconditional again: the requested
  result itself was lost and this line is its only signal, so it is not a diagnostic.

Tests that asserted the reverted behaviours are removed; the reviewer-round test file is
renamed for what it covers.
…one diagnostic-metadata helper

- `is_diagnostic_notice()` replaces three drifting copies: the gateway muted every
  `credits.*` notice, the TUI and CLI only `warn`/`error`, so `credits.restored` was hidden
  on Telegram and shown in the TUI for the same config. Every credit-service notice is an
  automatic diagnostic (a "restored" line after a hidden depletion notice is orphan noise).
- `effective_user_config()` is the single fail-open effective-config read; the two extra
  `deepcopy`s per foreground turn go (the loader already returns a fresh copy and the
  snapshot is read-only).
- `diagnostic_metadata(event)` replaces the repeated
  `{"notification_category": "diagnostic"} if event.internal and ... else {}` literal in
  gateway/run_turn.py; the gateway-side imports of the resolver are module-level (no cycle:
  it imports only gateway.display_config).
- `display.suppress_warning_notifications` is listed with its sibling display keys in
  cli-config.yaml.example and the configuration reference; the messaging guide states that a
  muted diagnostic wake still runs (and bills) its agent turn.
…wake predicate once

`_notification_suppressed_targets` lived on the job dict but had no reader outside
`_deliver_result`, and `_deliver_to_bot_chat` snapshots `dict(job)` into the durable
deferred record mid-loop, so a suppressed native target earlier in the loop leaked the list
into bot_chat_pending. A local counter has the same meaning and no durable footprint.

`diagnostic_wake_muted` is imported at module level in base.py instead of inside the
per-message handler and the turn-error path (no cycle: the resolver imports only
gateway.display_config).
Bot Chat is the TUI/Desktop transcript, so its warning policy is `display.platforms.tui`;
the bare `"tui"` literal at two sites read as a typo beside `BOT_CHAT_PLATFORM = "bot-chat"`.

`_deliver_to_bot_chat` popped `_notification_all_targets_suppressed` on entry, but both
callers already guarantee the key is absent (`_deliver_result` pops it before and after the
call; `_drain`'s record JSON was snapshotted before the flag is ever set).
…e base

`_run_prompt_submit` built `run_thread = threading.Thread(target=run)` and then started the
turn through `_start_session_work(run, ...)` — main had already moved to the latter, and the
merge kept both. The orphan Thread was never started; the real-threading requeue test records
every Thread the module constructs and joined that one first, which raises
"cannot join thread before it is started" (CI red on both salvage heads).
A turn can retain its session lease indefinitely when the worker has not published a usable activity snapshot, because the independent watchdog skips every poll. Fall back to elapsed worker time until a valid activity clock is available.

Refs NousResearch#104303.
…sing snapshot

The reachable gap is agent_holder=[None] during turn setup (run_turn_runner
fills the holder later); a real AIAgent always has _last_activity_ts, so
seconds_since_activity=None is not a state production produces. Replace the
None-returning fake with a snapshot-raises fail-safe case and add the
empty-holder test. Both are red against the pre-fallback watchdog loop.
…ad of growing on every reload

load_hermes_dotenv() runs per gateway turn, per cron fire, and from plugins /
MCP config / `hermes send`, each time re-applying .env with override. dotenv
interpolates `${VAR}` against the live os.environ, which already holds the
previous load's output, so `PATH=/x:${PATH}` gains one `/x:` per reload until
child spawns fail with E2BIG (NousResearch#109902).

Resolve inside _load_dotenv_with_fallback — the one seam every caller goes
through — against a private copy of os.environ in which this process's own
earlier dotenv output is peeled back to the value the variable had before we
first published it. The record is one process-wide per-variable table, not a
per-home/per-project scope: os.environ is process-wide, so a gateway reload
(project .env) alternating with a cron reload (none), or home A then B, must
peel each other's output too. Only values that still hold exactly what we
published are peeled, so shell exports, the terminal config bridge, and
external secret sources changing a value between reloads are honoured, not
frozen. Layers within one load_hermes_dotenv share a pass so the project and
managed .env still build on the user .env as before.

Co-authored-by: Hukla <129692708+huklaa@users.noreply.github.com>
Co-authored-by: Sahilvishnaliya <141555468+salch-cred@users.noreply.github.com>
…ill import env_loader

Gateway tests replace sys.modules["dotenv"] with a bare module exposing only
load_dotenv; gateway.run imports env_loader at import time, so the module-level
dotenv.main / dotenv.variables imports broke every one of those files in CI.
…y gone (NousResearch#109824)

_refresh_tools called self.session.list_tools unguarded. run() resets
self.session to None on every transport teardown (clean reconnect,
error backoff, park, cancel) and the parked-probe path does the same,
none of it under _rpc_lock or _refresh_lock. A tools/list_changed
refresh scheduled on the dying transport therefore wakes up behind the
lock mid-restart and crashes the background task:

    ERROR tools.mcp_tool: MCP server 'treadmill': dynamic tool refresh failed
    AttributeError: 'NoneType' object has no attribute 'list_tools'

once per profile that includes the server, on every gateway restart with
a slow-to-connect MCP server.

Snapshot the session only once _rpc_lock is held and return at debug
level when it is None. The snapshot sits inside the RPC lock because
that is the point at which the refresh has actually won the right to
talk to the transport; a check before acquiring it can still observe a
session that teardown nulls while the refresh waits. Skipping is
correct, not a loss: the reconnect's own discovery re-lists and
re-registers tools, and the next tools/list_changed re-arms the refresh
against the live session. The previous registration stays intact rather
than being nuked mid-restart.

Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
Co-authored-by: ildunari <ildunari@users.noreply.github.com>
Co-authored-by: Andrex Ibiza <84248988+andrexibiza@users.noreply.github.com>
Co-authored-by: ly6751 <liuyu890412@gmail.com>
…sion restart (NousResearch#109824)

Three behaviour contracts on _refresh_tools / _schedule_tools_refresh:

- A refresh awaited with session None completes, leaves the registry,
  the toolset alias and the owned-names bookkeeping untouched.
- The real background task path (_schedule_tools_refresh) finishes with
  no ERROR record on tools.mcp_tool, and a subsequent refresh on the
  reconnected session republishes normally. Pre-fix this logged
  "dynamic tool refresh failed" with the NoneType AttributeError.
- Once the session is restored a refresh from an empty registration
  publishes the live tools, so the skip is scoped to a missing session.

Co-authored-by: ly6751 <liuyu890412@gmail.com>
Co-authored-by: salch-cred <salch-cred@users.noreply.github.com>
Since NousResearch#110934 the warning counts writable handles only; joining every
member's creation site put read-only attaches (dashboard routers, one-shot
lookups) in a list meant to explain that number. Filter the listing with
the same predicate as the count, and pin it in the test with a read-only
attach that must stay out of the warning.
…-mnemosyne-rekey

catalog: mnemosyne key goes to mnemosyne-oss; unaffiliated Devs-Foundation entry delisted
`_resolve_bedrock_context_length` means to persist only probe-derived context
windows -- its comment says "Only persist probe-derived values (region present);
a pure table fallback must not poison the cache" -- but the code under it cannot
tell the two apart. `get_bedrock_context_length(model, region=region,
probe=bool(region))` returns the probed window when the probe succeeds and the
static table (or the 128K default) when it does not, in the same int, and the
guard `if ctx and region:` tests the region, not the probe. The region is never
empty: `resolve_bedrock_region()` ends in `or "us-east-1"`.

Observed on a Bedrock deployment whose SSO session expires every 8 hours: the
first turn after expiry probes without credentials, `probe_bedrock_context_length`
returns None, and

    context_lengths:
      '<opus-5 global inference profile id>@Bedrock:': 128000

is written (the synthetic `bedrock://` key lands on disk as `@bedrock:`, since
`_context_cache_key` strips trailing slashes) -- `BEDROCK_DEFAULT_CONTEXT_LENGTH`, because that model has no
`BEDROCK_CONTEXT_LENGTHS` row (NousResearch#74263, addressed by NousResearch#75824). The resolver returns
any cached value before it considers probing, so the probe never runs again for
that model and the compressor works from a 128K window on a 1M model.

The resolver now asks the probe itself and persists only what the probe said:

  probe returned a window   persist under model@base_url (or model@bedrock://), return it
  probe returned None       return the static table / default, persist nothing,
                            memoise the failure in memory for 5 minutes
  cached value present      serve it, as before

The memo keeps this cheap. `probe_bedrock_context_length` pads prompts of 1.3M and
2.2M tokens and attempts up to two `converse` calls (`_BEDROCK_PROBE_TIERS`; the
second tier is sent only when the first yields no parseable limit, i.e. on the
failure path), and `_cached_client` is a bare `boto3.client` that never validates
credentials, so a model whose probe keeps returning None WITH working credentials
(un-enabled model, opaque InternalServerException, unparseable length error)
would otherwise re-send both on every resolution -- and resolution is not memoised
per instance (model switch, every turn containing '@', each fallback candidate,
each /models API call). `_BEDROCK_PROBE_FAILURE_CACHE` /
`_BEDROCK_PROBE_FAILURE_TTL_SECONDS` follow the file's existing convention for
this shape (`_ENDPOINT_PROBE_FAILURE_TTL_SECONDS`, `_LOCAL_CTX_PROBE_CACHE`):
negative results in memory only, bounded, so credentials that come back are
noticed. A successful probe drops the entry, and
`_invalidate_cached_context_length` clears it for the model alongside the
local-probe memos, since a dropped entry is the reason to probe again.

Prior report: NousResearch#68049 (open, 2026-07-20) fixes the same persistence defect with
the same core replacement (probe directly, persist only a probed window, return
`get_bedrock_context_length(model, probe=False)` otherwise). Its hunk targets the
inline code in `get_model_context_length` that has since moved into
`_resolve_bedrock_context_length`, so it no longer applies to main. This commit
adds, on top of that mechanism, the in-memory failure memo, its clearing on
invalidation, and a test of the memo's TTL.

Deliberately unchanged: the cached-value path of `_resolve_bedrock_context_length`
(a persisted value is served as on main; the step-1 table floor is not applied
to the `bedrock://` key, because a probed window may legitimately be below the
table and flooring it would re-probe every second call);
`get_bedrock_context_length` (signature, `probe=` parameter, table lookup and its
tests); `BEDROCK_CONTEXT_LENGTHS` (the missing Opus 5 row is NousResearch#75824);
`save_context_length` / `get_cached_context_length`; the step ordering in
`get_model_context_length`.

tests/agent/test_model_metadata.py::TestBedrockContextCachePersistence, 2 cases:
a failed probe for a model with no table row persists nothing, and a failed probe
is memoised so a second resolution inside the TTL costs no probe and writes
nothing, then once the memo lapses the probe runs again and its window is what
gets persisted. Both fail on main, which persists 128000 on the first call.

(cherry picked from commit 5bc156a)
… same-name pip entry point

The pyproject wrapper shape (NousResearch#113851) depends on a pip package that often ships its own
hermes_agent.plugins entry point under the same name. Discovery appended entry points last with
"later wins", so after a catalog install the plugin row became source=entrypoint with no install
dir or catalog provenance: Desktop lost "installed from catalog @ version / update to ..." for
exactly the shape the catalog now recommends. Entry points no longer displace a directory plugin
of the same key, in the loader and in the CLI/RPC listing.

Live: catalog install of mnemosyne, row before = entrypoint / mnemosyne_hermes:register / no
catalog fields; after = user / ~/.hermes/plugins/mnemosyne / catalog 0.7.0 @ 95ef3be2; provider
still discovered. Test red on base.
…-turn

Slack seals a native stream server-side after a few minutes (live-observed
at ~5m20s on three independent long turns, 2026-09-15/16; the lifetime is
not documented). The next chat.appendStream on the card fails with
message_not_in_streaming_state. The adapter returned a bare failure, the
TurnRunner latched native_failed, and the rest of the turn rendered as an
edited text bullet list. Long autonomous turns lost the card UX exactly
when it mattered.

On message_not_in_streaming_state from appendStream, drop the dead
stream_ts and chat.startStream a fresh plan-mode card in the same thread,
then append the current frame there. Every frame already carries the full
visible task projection, so no task state is lost. One reopen per update;
a second rejection surfaces as a real failure. The sealed card is a plain
message now, so no stopStream is sent to it; the turn-final stop targets
the reopened card. Error matching reads SlackApiError.response["error"],
never the message text.

Tests assert the wire sequence (start, append, rejected append, start,
append on the new ts), the reopened frame's task states, the cache pointing
at the new card, and the stop targeting it; plus the one-reopen bound.
Mutation: forcing the expiry branch off turns both tests red.
…rting slack_sdk

CI stubs slack_sdk as a bare module, so `from slack_sdk.errors import
SlackApiError` fails there. The adapter only reads exc.response["error"];
the tests now raise a local exception with that shape. Verified with
slack_sdk blocked from import.
Third sync after PR #4. Fast-forwards 278 upstream commits onto
current fork main with no conflicts. Keeps calendar/vikunja skills
and the google_meet/spotify plugin removals.

Co-authored-by: Nagc <Nag112@users.noreply.github.com>
@github-actions

github-actions Bot commented Sep 18, 2026 •

Copy link
Copy Markdown

૮ >ﻌ< ა ci review

running on 72836a7 — Merge NousResearch/hermes-agent main into this fork.


Still running 4 jobs: JS & TS checks / JS & TS checks, OS-specific tests / Windows-only tests, Python tests / Run tests, Rust tests / cargo test (bootstrap installer)

⚠️ Action required

CI-sensitive file review · View job

This PR changes CI-sensitive files (eslint config, workflow YAMLs, or composite actions). These influence what the js-autofix job executes and pushes to main.

Sensitive files changed:

How to fix:

Add the ci-reviewed label after verifying:

  • no new eslint rules with custom fix functions that write outside linted paths,
  • no workflow changes that widen permissions or remove guards,
  • no composite action changes that alter what gets executed.

@Nag112
Nag112 marked this pull request as ready for review September 18, 2026 03:02
@Nag112
Nag112 merged commit 5d158f3 into main Sep 18, 2026
28 of 34 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.