Merge latest NousResearch/hermes-agent main - #5
Merged
Merged
Conversation
…enu and the keybind Follow-up to Ayush Nangia's feature: the filter menu kept its own literal list of groupings while the new cycle keybind walked a second one, with only a test comment saying they should agree. The menu now builds its options from `SIDEBAR_GROUPING_ORDER`, so reordering or adding a grouping is one edit and the keybind cannot drift from what the menu shows.
…variant cycle test `SidebarGrouping` was a literal union kept in step with `SIDEBAR_GROUPING_ORDER` by hand — a fifth grouping added to the type would compile while the cycle silently skipped it. The array is now the single source (`as const`) and the type is derived from it. The cycle test no longer pins the exact order (a change-detector on the constant); it asserts one lap visits every grouping once and returns to the start, and still fails when the cycle skips a step.
NousResearch#113535) _command_detection_variants yielded one FULL-LENGTH variant per quoted or escaped command word. A heredoc body of quoted lines ("key": "value", …) is hundreds of quoted command words, so both detection passes scanned O(words * len) characters: on a 16 KB / 460-line command the hardline pass alone took 7.8 s and the dangerous pass 15 s even with the launchctl lookahead anchored (33 KB: 36 s hardline), all while holding the GIL on the gateway loop. Build a single variant with every command word deobfuscated instead. The obfuscation catch ($(echo rm), r''m, ${0/x/r}m …) is unchanged — the same deobfuscated words appear, in one string — and the variant count on the issue's input drops from 923 to 4 (36 s -> 0.26 s hardline, verdict unchanged). detect_hardline_command needs no bound and keeps its YOLO semantics: a benign heredoc is neither blocked nor prompted. Root cause identified in NousResearch#113943 by @Tranquil-Flow (variant explosion as the second compounding factor); its cumulative-work budget is superseded by removing the explosion. Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
…er test A bare `python -c` resolves `tools` through the venv's editable install (the primary clone), not the worktree pytest runs in, so the regression guard could silently test a different tree. Pass cwd + PYTHONPATH like tests/tools/test_async_delegation.py does.
…ange load_env_file() is now the canonical .env tokenizer after c849bc3 collapsed six hand parsers onto it. Profile scopes, dashboard boundary discovery, skill secret capture, managed .env and setup readers share the same decoding and parsing rules, but direct calls still re-read and re-parse the whole file. build_profile_secret_scope() puts that work on the gateway's hottest paths: every turn, every cron job, MCP and browser lifecycle adoption, the TUI prompt turn, and the 60s housekeeping drain of the durable cron delivery queue. Memoising the shared tokenizer now removes redundant work for the consolidated readers as well as the motivating profile-scope path. config.load_env() keeps its existing outer memo over a call that is now itself memoised; this leaves its menu-render shortcut intact without redesigning it. Freshness is the hard part, because a stale credential map is a far worse failure than a slow one. Every load_env_file() call still OPENS the file and keys on os.fstat() of that descriptor rather than os.stat() of the path: * the open preserves close-to-open revalidation on NFS; * an open or read failure returns {} without caching it, and drops any warm entry so a transient EACCES cannot become a persistent empty profile; * the descriptor pins one inode, so a symlink repointed mid-read cannot file one file's contents under another file's identity. The key is (mtime_ns, size, inode, device), re-checked on the same descriptor after reading. A rewrite through that inode is not stored under the pre-read fingerprint. Unavailable metadata means "do not cache", never "cannot read". A generation counter checked under the lock prevents a slow reader from repopulating an entry that a writer invalidated while the read was in flight. Read bytes from the open descriptor and preserve upstream's decoding exactly: strip a UTF-8 BOM before trying UTF-8, then fall back to latin-1 for invalid UTF-8. Both the cached reader and the uncached reference use that decoder and the shared text parser. Invalid UTF-8 is a cacheable result, not a read error. The value and inline-comment helpers remain local to secret_scope. invalidate_env_cache() clears this memo too, so save_env_value, remove_env_value and sanitize_env_file invalidate both layers. The per-path LRU holds at most 64 entries and every caller gets its own dict. Accepted limitation: a write keeping length, inode and nanosecond mtime all identical is not detected without explicit invalidation. This is a new limitation for previously uncached readers. Parsed secrets also remain reachable in memory between use and eviction for up to 64 paths, where the outer config memo holds one. The pre-rebase measurement on a 508-line, 25432-byte .env was 0.868 ms -> 0.011 ms per call. That measurement predates the shared binary decoder. Add decode/cache invariants for BOM and plain files, latin-1 fallback, same-size writes switching UTF-8 -> latin-1 -> UTF-8, and agreement with the uncached reference. All eight decoding, cache and freshness mutations fail the new tests, and all 40 cache tests pass after restoration. The requested agent, CLI, cron and gateway suites report 30,168 passed, 479 failed and 303 skipped in the restricted local environment. All 140 failing files were checked against upstream/main at 5eb99eb: all 479 failed-test IDs and 210 collection/setup/teardown error IDs match exactly.
…sted The 5s goals poller calls restore_heartbeat_watches(), which listed every routed session and entered _profile_runtime_scope for each distinct origin — a full config.yaml/.env parse plus secrets hydration per session per scan. With zero heartbeats persisted (the common idle case) the entire sweep is wasted work; on a 400-session multiplex gateway it dominated idle CPU (~60% of on-CPU time in YAML parsing). Gate the sweep on an indexed heartbeat:* prefix probe of each served profile's SessionDB (HERMES_HOME contextvar only — no config/secret scopes). Probe failures fail open to the historical full sweep.
…ofile's store The handoff watcher (2s interval) and loop wakeup watcher (15s) entered _async_profile_runtime_scope for EVERY served profile on EVERY tick — a full config.yaml/.env parse plus secrets hydration per entry — even when the profile's state.db held zero pending handoffs or active loops. On a multiplex gateway this, not message traffic, was the idle CPU driver (py-spy: ~75% of on-CPU time in YAML parsing from these two paths). Both watchers now probe the profile's cached SessionDB first (only the HERMES_HOME contextvar installed, off the loop thread via the runner's executor hop) and enter the scope only when work exists: - handoff watcher: new SessionDB.has_pending_handoffs() existence probe - loop watcher: list_active_loops() under the probe context All probes fail OPEN (error/unknown -> historical always-enter behavior); runners without the executor hop (test stand-ins) skip the gate, so existing watcher tests are unaffected. The startup stale-handoff reclaim stays ungated: it runs once per boot and must also see 'running' leftovers, which a pending-only probe would miss.
…nd ignore cleared rows The loop watcher's probe used ``list_active_loops()``, which collapses an unavailable SessionDB (and any read error) to ``[]`` — the gate then skipped that profile on every tick (fail CLOSED). The heartbeat probe keyed on ``heartbeat:*`` key existence, but ``clear`` and ``pause`` keep their rows, so one past /heartbeat re-enabled the full sweep forever. Both probes now read the rows through one ``_profile_meta_rows`` helper (None = store unavailable → treat as work present) and parse the status themselves: only an ACTIVE row opens the gate; a corrupt row is "unknown" and keeps the historical scan. The heartbeat gate reuses ``_handoff_watch_scopes`` for the served homes instead of a second resolver.
Heartbeat: no active row → no sweep; a cleared row → still no sweep; an active row → sweep; an unopenable store → sweep. Loop watcher: idle profile never scoped, active loop scoped, unopenable store scoped. The handoff-watcher test from NousResearch#109497 is dropped: same shape, and the salvage bar is two tests per fix.
The memo had its own (mtime_ns, size, inode, device) tuple and documented a same-length/same-inode/pinned-mtime rewrite as an accepted gap. utils.file_signature is the repo's change-detection key (config.load_env, skill_utils, prompt_builder already use it) and includes ctime_ns, which user space cannot backdate — so the gap closes and one helper goes.
load_env() kept its own (path, file_signature) memo on top of load_env_file(), which now memoises on the same key: two stat calls and two caches to keep coherent for one answer. Drop the outer layer; invalidate_env_cache() stays as the writers' knob and forwards to the one memo.
…named-profile multiplexer _watched_homes reused _handoff_watch_scopes, which deliberately omits "default" because the handoff/loop watchers' root poll covers the launch home. For a `hermes -p work` multiplexer the launch home is profiles/work, so ~/.hermes was probed by nobody and an active heartbeat there answered "idle" every tick. Probe the served set the sweep can actually resolve to.
…ive with loops/heartbeat The probe helpers had been appended to the gateway/run.py facade and the three watchers each carried their own copy of the "probe → fail open → offload" shell, with gateway code parsing loop/heartbeat rows itself. One sibling now owns the gate shape (_gate + off_loop_gate), and hermes_cli.loops.store_has_active_loop / hermes_cli.heartbeat.store_has_active_heartbeat own what an ACTIVE row is (heartbeat gains the _META_PREFIX loops already had). One fail-open layer per read: has_pending_handoffs is a bare bounded query, the gate catches. The loop watcher calls its executor hop directly — it already did so unconditionally for the scan, so the "runner without a hop" fallback there was dead.
…ome case and the handoff gate Loop watcher: a cleared row must not open the gate (a key-existence mutation now fails), then the unavailable-store fail-open. Heartbeat restore: routing home = named profile with the only active row in the default store must still sweep. The handoff-watcher test from NousResearch#109497 is restored (its gate and the has_pending_handoffs query had no coverage).
… update Port of NousResearch#59942 onto post-NousResearch#102117 main. The refactor moved the desktop helpers from hermes_cli/main.py into hermes_cli/main_desktop.py, so the new install machinery (_app_asar_hash, _macos_adhoc_sign_bundle, _desktop_bundle_install_supported, _swap_in_new_macos_bundle, _install_rebuilt_desktop_app) lives there now, re-exported through main.py's frozen updater surface block so _m() resolves them. _rebuild_desktop_after_update (hermes_cli/update_cmd_deps.py) now, on a successful --build-only build, installs the rebuilt bundle to the system location (macOS .app via ditto staged copy + atomic swap with rollback; Windows via copytree), skipping when the installed app.asar hash already matches. Linux stays owned by its package mechanism. Without this, a CLI hermes update leaves the installed app stale — the in-app updater covers its own path, but the CLI path never swapped the installed bundle. Tests: 17/17 ported suite (import/patch targets repointed to main_desktop) + test_cmd_update.py 47/47 green.
…gnature, never swap under a live app Rework of the salvaged NousResearch#59942 hook so it fixes the whole NousResearch#52339 class without regressing what the Desktop updater already gets right: - macOS only. The Windows arm rmtree'd a possibly-running NSIS install (partial deletion under a lock) and windows.ps1 already owns that swap; Linux packages stay with their package manager. - No re-sign. The rebuilt release/ bundle already carries the stable local signing identity from _desktop_macos_relaunchable_fixup; a deep `codesign -s -` on the installed copy replaced it with a fresh ad-hoc cdhash and reset every TCC grant. ditto preserves the signature, so nothing is signed here. - Running bundles are reported, not swapped: Electron loads app.asar and helper apps lazily, so renaming the bundle away and deleting the old tree crashes the live app. The detached updater waits for exit; a terminal `hermes update` with the app open now prints what to do instead. - Failures are printed as warnings; the old `Path | None` return read every failure as "Desktop app up to date". - The refresh also runs on the "build stamp current" path, so a stale /Applications copy left by an earlier update heals on the next `hermes update` even when there is nothing to rebuild. - Core is host-independent (_install_rebuilt_macos_bundles takes paths as data); the two invariant tests run on every OS instead of `skipif(darwin)` tests that ran nowhere. Covers the Desktop-button path too: posix.sh runs `hermes update`, so an app running from apps/desktop/release/ now refreshes the /Applications copy Finder launches (the Discord report: new shell right after the update, old shell on the next Dock launch).
Cherry-pick of helix4u's NousResearch#88836 (75742db) resolved onto current main. The POSIX runtime cutover parks the live venv with a directory rename. On Windows that can never succeed from `hermes update`: the updater executes from venv\Scripts\python.exe and Windows keeps that image mapped, so the rename fails with WinError 5 on every attempt (NousResearch#93032). Instead, atomically replace the live venv's pyvenv.cfg with the candidate's: Scripts\python.exe is a launcher that reads `home` on every start, so every fresh process runs the new generation while nothing touches the locked directory. A failed post-cutover smoke restores the original config; a failed rollback keeps the generation the live config now references.
The self-lock/holder preflight (NousResearch#99711) deferred the repair on the theory that the updater's own mapped python.exe makes the venv rename impossible. Live on windows-latest, a process executing from venv\Scripts\python.exe (and the repo's real .venv with cp311 .pyd extensions loaded) does NOT block the rename; what blocks it with WinError 5 is any ordinary handle under the tree: a process cwd, an open file, a sync client. Since sys.executable is always under the venv on Windows, the deferral fired on every run and no Windows install could repair from `hermes update`. Retire it and its tests; the pyvenv.cfg repoint needs no rename and has none of that exposure. The wine2e lane now runs the cutover probes instead. The repoint keeps the live venv's site-packages, so provisioning's fall-forward to the next minor (3.11 -> 3.12, NousResearch#76106) must not be pointed at a cp311 tree: refuse before touching pyvenv.cfg. The live Windows test holds the venv the way the field does (a child with cwd inside it), asserts the rename path fails (the symptom) and the repoint succeeds; the rollback and minor-guard tests run on every host.
…ousResearch#85194) `auxiliary.title_generation.enabled: false` turned off both title stages, so an operator who only wanted to stop the background model call (unavailable or metered endpoint) also lost the instant derived title. `model_upgrade_enabled: false` keeps the derived title and starts no `auto-title` thread; missing keeps the two-stage default and `enabled: false` still disables both. Salvaged from NousResearch#85401 onto current main: the gate reuses `_title_config()` and sits before `spawn_context_thread` (the thread seam moved off `threading.Thread`).
Non-noreply author email on the salvaged NousResearch#85401 commit.
…rade The gate also sat inside `generate_title`, so the explicit operator repair command `hermes sessions retitle-skills` returned None for every row when the toggle was off, although the operator asked for a model call. The toggle's contract is "never spend a model call upgrading the instant title": only `maybe_auto_title` (the background path) is gated now; a direct `generate_title` call still asks the model. Docs sentence adjusted. Review follow-up on NousResearch#113955; the existing toggle test now also asserts the explicit path still titles (red on the previous head).
…usResearch#72351) Title generation hardcoded temperature=0.3 in its call_llm() call. Models like GPT-5.6 only accept their server-side default temperature and reject explicit values with "Unsupported value: 'temperature'". While call_llm() has a retry that strips temperature on error, the daemon thread races with session cleanup in short-lived CLI sessions, causing the retry to fail with a connection error. Fix: pass temperature=None so the provider uses its own default. This avoids the unsupported-temperature error entirely and eliminates the need for the retry path. Fixes NousResearch#72351
…ack requests Reasoning models reject several request fields at once (gpt-5: temperature AND max_tokens, NousResearch#78273), a reasoning-strip retry can then 400 on temperature (NousResearch#72351), and strict-schema gateways such as Fireworks reject the generic extra_body.reasoning fallback with "Extra inputs are not permitted, field: 'reasoning'" (NousResearch#109774) — a phrasing none of the unsupported-parameter predicates recognised, so the reasoning rung never fired and every title call 400'd. The parameter rungs were a fixed single-pass order (temperature → structured output → reasoning → max_tokens): a field rejected AFTER an earlier rung had already run was never stripped, and _param_rung_accepts did not admit a temperature 400 raised by a later retry. They now form a table walked until no rung matches, each field stripped at most once, so N rejected fields recover in N retries in whatever order the provider raises them. Fallback candidates (per-task chain, main chain, discovery) got a single shot with the caller's temperature/max_tokens/reasoning fields and raised on the first parameter 400; agent/auxiliary_fallback_recovery.py now runs the same parameter rungs around a candidate's request (sync + async), leaving auth / payment / connection errors to the caller's existing handling. Live: real gpt-5-mini title_generation call (reasoning_effort 400 → temperature 400 → 200), real gpt-5-mini fallback candidate sync+async (temperature 400 → 200), and a stand-in replaying Fireworks' documented 400 (reasoning stripped → 200) — all failed on origin/main.
A positional _LadderRoute(...) in auxiliary_fallback_recovery breaks as soon as the route tuple gains a field (NousResearch#113968 adds timeout); construct by name so fields the parameter ladder never reads default to None regardless of tuple width.
…lds the route by name base_info is the endpoint identity every rung reads for per-route bookkeeping; the fallback ladder left it empty. The test mirrored the positional construction that broke under a wider tuple.
…114356 salvage resolution
…ications Squash of the 54 commits on victor-kyriazakos:feat/user-channel-warning-suppression (PR NousResearch#112302, head f45c640) so the contributor's authorship survives a rebase-merge; the commits interleave with a cron delivery-ledger rework that the salvage removes in follow-up commits, so per-commit cherry-picks were not practical. Adds display.suppress_warning_notifications (global + per-platform, default false): one resolver (gateway/warning_notifications.py), BasePlatformAdapter.emit_warning / emit_media_warning / warning_text, a notification_category classification carried through wakes, queues and persistence, and render/present boundaries for CLI/TUI.
…nifest ledger rework The suppression feature needs one cron cell: a failure notice whose target hides warning notifications is recorded as `suppressed` (bot_chat_pending record, deliveries queue row, execution delivery_outcome) instead of being sent. That is kept. Everything else the PR added to cron/ is an independent ledger rework and is removed here: execution delivery manifests + filesystem manifest journal, `schema_meta` fence and the "legacy intent adoption" layer, incident occurrence generations + trigger, jobs projection CAS (`bind_delivery_execution`/`update_delivery_projection`), the per-tick projection reconciler, and the `_deliver_targets`/`_settle_manifest` split. origin/main has none of the state that layer migrates from (zero occurrences of delivery_manifest / manifest_journal / schema_meta); the "pre-flag old writer" the r5-r9 tests simulate is a vendored snapshot of this PR's own earlier revision (tests/cron/_r5_prior_executions.py). Those fixes may have merit on their own and should land as separate PRs with a main-reproducing test each. Also restores main's `_classify_delivery_outcome` precedence (`failed` before `queued`) and drops the SHA-pinned `git show <PR commit>` test, which would go red the moment a rebase-merge rewrote that commit.
Four places changed behaviour for users who never touched the setting: - `_interim_send` was stamped on every `warn` status and media-failure notice, and the Slack/relay egress doors learned to skip stream sealing for it. Main's status sends carry no interim mark at all, so the gap is class-wide (every status kind), and fixing it for warnings alone is an undeclared streaming-contract change. Reverted here; the whole-class fix belongs in its own PR against gateway/AGENTS.md rule 3. - The entire post-handler delivery (unwrap, TTS, final text, attachments, delivery-ledger writes) ran inside `_media_delivery_scope`. Under multiplex that binds the routed home, so delivery obligations landed in the routed profile's state.db while boot-time `_claim_pending_obligations` still reads the launch home. Only the policy reads (`diagnostic_wake_muted`, `warning_text`) bind the routed scope now; delivery stays where main ran it. - The turn-crash notice is rebuilt the same way: scope around the policy read, send outside. - The "delivery failed after multiple attempts" notice is unconditional again: the requested result itself was lost and this line is its only signal, so it is not a diagnostic. Tests that asserted the reverted behaviours are removed; the reviewer-round test file is renamed for what it covers.
…one diagnostic-metadata helper
- `is_diagnostic_notice()` replaces three drifting copies: the gateway muted every
`credits.*` notice, the TUI and CLI only `warn`/`error`, so `credits.restored` was hidden
on Telegram and shown in the TUI for the same config. Every credit-service notice is an
automatic diagnostic (a "restored" line after a hidden depletion notice is orphan noise).
- `effective_user_config()` is the single fail-open effective-config read; the two extra
`deepcopy`s per foreground turn go (the loader already returns a fresh copy and the
snapshot is read-only).
- `diagnostic_metadata(event)` replaces the repeated
`{"notification_category": "diagnostic"} if event.internal and ... else {}` literal in
gateway/run_turn.py; the gateway-side imports of the resolver are module-level (no cycle:
it imports only gateway.display_config).
- `display.suppress_warning_notifications` is listed with its sibling display keys in
cli-config.yaml.example and the configuration reference; the messaging guide states that a
muted diagnostic wake still runs (and bills) its agent turn.
…wake predicate once `_notification_suppressed_targets` lived on the job dict but had no reader outside `_deliver_result`, and `_deliver_to_bot_chat` snapshots `dict(job)` into the durable deferred record mid-loop, so a suppressed native target earlier in the loop leaked the list into bot_chat_pending. A local counter has the same meaning and no durable footprint. `diagnostic_wake_muted` is imported at module level in base.py instead of inside the per-message handler and the turn-error path (no cycle: the resolver imports only gateway.display_config).
Bot Chat is the TUI/Desktop transcript, so its warning policy is `display.platforms.tui`; the bare `"tui"` literal at two sites read as a typo beside `BOT_CHAT_PLATFORM = "bot-chat"`. `_deliver_to_bot_chat` popped `_notification_all_targets_suppressed` on entry, but both callers already guarantee the key is absent (`_deliver_result` pops it before and after the call; `_drain`'s record JSON was snapshotted before the flag is ever set).
…e base `_run_prompt_submit` built `run_thread = threading.Thread(target=run)` and then started the turn through `_start_session_work(run, ...)` — main had already moved to the latter, and the merge kept both. The orphan Thread was never started; the real-threading requeue test records every Thread the module constructs and joined that one first, which raises "cannot join thread before it is started" (CI red on both salvage heads).
A turn can retain its session lease indefinitely when the worker has not published a usable activity snapshot, because the independent watchdog skips every poll. Fall back to elapsed worker time until a valid activity clock is available. Refs NousResearch#104303.
…sing snapshot The reachable gap is agent_holder=[None] during turn setup (run_turn_runner fills the holder later); a real AIAgent always has _last_activity_ts, so seconds_since_activity=None is not a state production produces. Replace the None-returning fake with a snapshot-raises fail-safe case and add the empty-holder test. Both are red against the pre-fallback watchdog loop.
…ad of growing on every reload
load_hermes_dotenv() runs per gateway turn, per cron fire, and from plugins /
MCP config / `hermes send`, each time re-applying .env with override. dotenv
interpolates `${VAR}` against the live os.environ, which already holds the
previous load's output, so `PATH=/x:${PATH}` gains one `/x:` per reload until
child spawns fail with E2BIG (NousResearch#109902).
Resolve inside _load_dotenv_with_fallback — the one seam every caller goes
through — against a private copy of os.environ in which this process's own
earlier dotenv output is peeled back to the value the variable had before we
first published it. The record is one process-wide per-variable table, not a
per-home/per-project scope: os.environ is process-wide, so a gateway reload
(project .env) alternating with a cron reload (none), or home A then B, must
peel each other's output too. Only values that still hold exactly what we
published are peeled, so shell exports, the terminal config bridge, and
external secret sources changing a value between reloads are honoured, not
frozen. Layers within one load_hermes_dotenv share a pass so the project and
managed .env still build on the user .env as before.
Co-authored-by: Hukla <129692708+huklaa@users.noreply.github.com>
Co-authored-by: Sahilvishnaliya <141555468+salch-cred@users.noreply.github.com>
…ill import env_loader Gateway tests replace sys.modules["dotenv"] with a bare module exposing only load_dotenv; gateway.run imports env_loader at import time, so the module-level dotenv.main / dotenv.variables imports broke every one of those files in CI.
…y gone (NousResearch#109824) _refresh_tools called self.session.list_tools unguarded. run() resets self.session to None on every transport teardown (clean reconnect, error backoff, park, cancel) and the parked-probe path does the same, none of it under _rpc_lock or _refresh_lock. A tools/list_changed refresh scheduled on the dying transport therefore wakes up behind the lock mid-restart and crashes the background task: ERROR tools.mcp_tool: MCP server 'treadmill': dynamic tool refresh failed AttributeError: 'NoneType' object has no attribute 'list_tools' once per profile that includes the server, on every gateway restart with a slow-to-connect MCP server. Snapshot the session only once _rpc_lock is held and return at debug level when it is None. The snapshot sits inside the RPC lock because that is the point at which the refresh has actually won the right to talk to the transport; a check before acquiring it can still observe a session that teardown nulls while the refresh waits. Skipping is correct, not a loss: the reconnect's own discovery re-lists and re-registers tools, and the next tools/list_changed re-arms the refresh against the live session. The previous registration stays intact rather than being nuked mid-restart. Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com> Co-authored-by: ildunari <ildunari@users.noreply.github.com> Co-authored-by: Andrex Ibiza <84248988+andrexibiza@users.noreply.github.com> Co-authored-by: ly6751 <liuyu890412@gmail.com>
…sion restart (NousResearch#109824) Three behaviour contracts on _refresh_tools / _schedule_tools_refresh: - A refresh awaited with session None completes, leaves the registry, the toolset alias and the owned-names bookkeeping untouched. - The real background task path (_schedule_tools_refresh) finishes with no ERROR record on tools.mcp_tool, and a subsequent refresh on the reconnected session republishes normally. Pre-fix this logged "dynamic tool refresh failed" with the NoneType AttributeError. - Once the session is restored a refresh from an empty registration publishes the live tools, so the skip is scoped to a missing session. Co-authored-by: ly6751 <liuyu890412@gmail.com> Co-authored-by: salch-cred <salch-cred@users.noreply.github.com>
Since NousResearch#110934 the warning counts writable handles only; joining every member's creation site put read-only attaches (dashboard routers, one-shot lookups) in a list meant to explain that number. Filter the listing with the same predicate as the count, and pin it in the test with a read-only attach that must stay out of the warning.
…-mnemosyne-rekey catalog: mnemosyne key goes to mnemosyne-oss; unaffiliated Devs-Foundation entry delisted
`_resolve_bedrock_context_length` means to persist only probe-derived context
windows -- its comment says "Only persist probe-derived values (region present);
a pure table fallback must not poison the cache" -- but the code under it cannot
tell the two apart. `get_bedrock_context_length(model, region=region,
probe=bool(region))` returns the probed window when the probe succeeds and the
static table (or the 128K default) when it does not, in the same int, and the
guard `if ctx and region:` tests the region, not the probe. The region is never
empty: `resolve_bedrock_region()` ends in `or "us-east-1"`.
Observed on a Bedrock deployment whose SSO session expires every 8 hours: the
first turn after expiry probes without credentials, `probe_bedrock_context_length`
returns None, and
context_lengths:
'<opus-5 global inference profile id>@Bedrock:': 128000
is written (the synthetic `bedrock://` key lands on disk as `@bedrock:`, since
`_context_cache_key` strips trailing slashes) -- `BEDROCK_DEFAULT_CONTEXT_LENGTH`, because that model has no
`BEDROCK_CONTEXT_LENGTHS` row (NousResearch#74263, addressed by NousResearch#75824). The resolver returns
any cached value before it considers probing, so the probe never runs again for
that model and the compressor works from a 128K window on a 1M model.
The resolver now asks the probe itself and persists only what the probe said:
probe returned a window persist under model@base_url (or model@bedrock://), return it
probe returned None return the static table / default, persist nothing,
memoise the failure in memory for 5 minutes
cached value present serve it, as before
The memo keeps this cheap. `probe_bedrock_context_length` pads prompts of 1.3M and
2.2M tokens and attempts up to two `converse` calls (`_BEDROCK_PROBE_TIERS`; the
second tier is sent only when the first yields no parseable limit, i.e. on the
failure path), and `_cached_client` is a bare `boto3.client` that never validates
credentials, so a model whose probe keeps returning None WITH working credentials
(un-enabled model, opaque InternalServerException, unparseable length error)
would otherwise re-send both on every resolution -- and resolution is not memoised
per instance (model switch, every turn containing '@', each fallback candidate,
each /models API call). `_BEDROCK_PROBE_FAILURE_CACHE` /
`_BEDROCK_PROBE_FAILURE_TTL_SECONDS` follow the file's existing convention for
this shape (`_ENDPOINT_PROBE_FAILURE_TTL_SECONDS`, `_LOCAL_CTX_PROBE_CACHE`):
negative results in memory only, bounded, so credentials that come back are
noticed. A successful probe drops the entry, and
`_invalidate_cached_context_length` clears it for the model alongside the
local-probe memos, since a dropped entry is the reason to probe again.
Prior report: NousResearch#68049 (open, 2026-07-20) fixes the same persistence defect with
the same core replacement (probe directly, persist only a probed window, return
`get_bedrock_context_length(model, probe=False)` otherwise). Its hunk targets the
inline code in `get_model_context_length` that has since moved into
`_resolve_bedrock_context_length`, so it no longer applies to main. This commit
adds, on top of that mechanism, the in-memory failure memo, its clearing on
invalidation, and a test of the memo's TTL.
Deliberately unchanged: the cached-value path of `_resolve_bedrock_context_length`
(a persisted value is served as on main; the step-1 table floor is not applied
to the `bedrock://` key, because a probed window may legitimately be below the
table and flooring it would re-probe every second call);
`get_bedrock_context_length` (signature, `probe=` parameter, table lookup and its
tests); `BEDROCK_CONTEXT_LENGTHS` (the missing Opus 5 row is NousResearch#75824);
`save_context_length` / `get_cached_context_length`; the step ordering in
`get_model_context_length`.
tests/agent/test_model_metadata.py::TestBedrockContextCachePersistence, 2 cases:
a failed probe for a model with no table row persists nothing, and a failed probe
is memoised so a second resolution inside the TTL costs no probe and writes
nothing, then once the memo lapses the probe runs again and its window is what
gets persisted. Both fail on main, which persists 128000 on the first call.
(cherry picked from commit 5bc156a)
… same-name pip entry point The pyproject wrapper shape (NousResearch#113851) depends on a pip package that often ships its own hermes_agent.plugins entry point under the same name. Discovery appended entry points last with "later wins", so after a catalog install the plugin row became source=entrypoint with no install dir or catalog provenance: Desktop lost "installed from catalog @ version / update to ..." for exactly the shape the catalog now recommends. Entry points no longer displace a directory plugin of the same key, in the loader and in the CLI/RPC listing. Live: catalog install of mnemosyne, row before = entrypoint / mnemosyne_hermes:register / no catalog fields; after = user / ~/.hermes/plugins/mnemosyne / catalog 0.7.0 @ 95ef3be2; provider still discovered. Test red on base.
…-turn Slack seals a native stream server-side after a few minutes (live-observed at ~5m20s on three independent long turns, 2026-09-15/16; the lifetime is not documented). The next chat.appendStream on the card fails with message_not_in_streaming_state. The adapter returned a bare failure, the TurnRunner latched native_failed, and the rest of the turn rendered as an edited text bullet list. Long autonomous turns lost the card UX exactly when it mattered. On message_not_in_streaming_state from appendStream, drop the dead stream_ts and chat.startStream a fresh plan-mode card in the same thread, then append the current frame there. Every frame already carries the full visible task projection, so no task state is lost. One reopen per update; a second rejection surfaces as a real failure. The sealed card is a plain message now, so no stopStream is sent to it; the turn-final stop targets the reopened card. Error matching reads SlackApiError.response["error"], never the message text. Tests assert the wire sequence (start, append, rejected append, start, append on the new ts), the reopened frame's task states, the cache pointing at the new card, and the stop targeting it; plus the one-reopen bound. Mutation: forcing the expiry branch off turns both tests red.
…rting slack_sdk CI stubs slack_sdk as a bare module, so `from slack_sdk.errors import SlackApiError` fails there. The adapter only reads exc.response["error"]; the tests now raise a local exception with that shape. Verified with slack_sdk blocked from import.
Third sync after PR #4. Fast-forwards 278 upstream commits onto current fork main with no conflicts. Keeps calendar/vikunja skills and the google_meet/spotify plugin removals. Co-authored-by: Nagc <Nag112@users.noreply.github.com>
૮ >ﻌ< ა ci reviewrunning on 72836a7 — Merge NousResearch/hermes-agent main into this fork. Still running 4 jobs:
|
Nag112
marked this pull request as ready for review
September 18, 2026 03:02
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Merges the latest NousResearch/hermes-agent
main(e83b1d51f1) into this fork.This is the next upstream sync after PR #4. Git merged cleanly (no conflicts). ~278 commits / 514 files.
Open PR #3 (
NousResearch:main→Nag112:main) is the same direction; once this lands, close #3.Type of Change
Changes Made
upstream/mainthroughe83b1d51f1.skills/productivity/calendar/,skills/productivity/vikunja/.plugins/google_meetandplugins/spotifystay removed.How to Test
mainshould match current Hermes plus the two productivity skills.Checklist
main