Skip to content

fix(delegation): keep credential pools endpoint-coherent - #2

Open
oferlaor wants to merge 10000 commits into
waritbb:fix/delegation-runtime-pool-coherence-prfrom
oferlaor:fix/pr39862-endpoint-coherence
Open

oferlaor wants to merge 10000 commits into
waritbb:fix/delegation-runtime-pool-coherence-prfrom
oferlaor:fix/pr39862-endpoint-coherence

Conversation

@oferlaor

Copy link
Copy Markdown

Summary

This current-main patch extends the provider-coherence fix from NousResearch#39862 to cover same-provider endpoint mismatches.

  • reuses credential_pool_matches_provider(..., base_url=...) so named custom pools remain supported;
  • requires production credential-pool entries to match the delegated child’s active endpoint before pool sharing or lease binding;
  • prevents an openai-api pool for https://api.openai.com/v1 from overwriting a child configured for an Azure OpenAI endpoint;
  • revalidates the pool immediately before lease binding as a defense-in-depth boundary;
  • preserves legacy unscoped pool adapters;
  • adds Azure/public-OpenAI mismatch and named-custom-pool regressions.

Once merged into fix/delegation-runtime-pool-coherence-pr, the existing upstream PR NousResearch#39862 will update automatically.

Validation

  • python -m pytest tests/tools/test_delegate.py -q -o 'addopts=' → 161 passed
  • Ruff on changed files → passed
  • Python compilation → passed
  • git diff --check → passed
  • Production CredentialPool probe:
    • Azure child + public OpenAI pool → rejected
    • Azure child + matching Azure pool → shared

@oferlaor

ghost commented Jul 20, 2026

Copy link
Copy Markdown
Author

Tracking issue: NousResearch#68237

This companion is the current-main endpoint-aware extension of NousResearch#39862.

@oferlaor

ghost commented Jul 20, 2026

Copy link
Copy Markdown
Author

Superseded by the direct upstream PR:
NousResearch#68240

Tracking issue: NousResearch#68237

@oferlaor
oferlaor force-pushed the fix/pr39862-endpoint-coherence branch 2 times, most recently from 4b83e2b to 24e71b8 Compare August 3, 2026 19:33
@oferlaor
oferlaor force-pushed the fix/pr39862-endpoint-coherence branch 3 times, most recently from d7d57d7 to 826fc47 Compare August 10, 2026 08:33
@oferlaor
oferlaor force-pushed the fix/pr39862-endpoint-coherence branch from 826fc47 to 9031c96 Compare August 18, 2026 08:28
@oferlaor
oferlaor force-pushed the fix/pr39862-endpoint-coherence branch from 9031c96 to 9681f36 Compare September 1, 2026 07:46
Teknium and others added 18 commits September 11, 2026 06:21
The dashboard's `_eager_reconcile_own_session_db` did an unconditional
writable `acquire()` at every startup. When the gateway shares that
state.db the dashboard became a second long-lived writable SessionDB
owner: a close-time WAL checkpoint plus a possible FTS rebuild in
`_init_fts`, the two-writer vector behind the corruption reports in
NousResearch#107688 and NousResearch#100896 ("5 live SessionDB handles" precursor, gateway +
dashboard both holding the WAL).

Route the startup reconcile through `_open_session_db_at_path(...,
read_only=True)`, which already bootstraps a missing store and heals a
stale/malformed schema through exactly ONE writable open before
reopening read-only. A healthy store now gets zero writable opens from
the dashboard while the NousResearch#79531/NousResearch#80037 "bring schema current before the
first poll" contract is kept (existing heal test unchanged).

Live repro (healthy store, count writable SessionDB.__init__ calls from
the startup worker): before=1 after=0.

Reported-by: NousResearch#107688, NousResearch#100896 (@kokhlo diagnosis)
Refs NousResearch#107688 NousResearch#100896
`hermes profile delete` leaves `profiles/.deleted/<name>` behind.
`HostedRoomService.local_profiles()` fed every subdirectory name to the
roster, so `.deleted` failed `validate_roster`'s identifier check and
`plan_next_task` raised on every cycle for every room until the
directory was removed by hand (NousResearch#106847, bug 2). Skip dot-dirs and
tombstoned profiles, using the same `named_profile_is_deleted` predicate
`hermes_cli.profiles` uses to list live profiles.

Refs NousResearch#106847
…b; label sidecars

Invariant test for the NousResearch#103489 salvage: a profile gateway and the root
gateway resolve the same hosted-room coordination file at the install
root, and that file is never the root SessionDB (`state.db`). Every
profile gateway starts the hosted-room worker unconditionally, so the
old mapping made each profile process a long-lived writer on the master
session store — the simultaneous-restart corruption in NousResearch#102120 and the
main offender in the NousResearch#103339 / NousResearch#103490 fleet reports.

Also: WAL-fallback labels say `shared-state.db`, and test fakes take the
store path from `default_db_path()` instead of hardcoding `state.db`.

Reported-by: NousResearch#102120 NousResearch#103339 (@RChina) NousResearch#103490
Refs NousResearch#102120 NousResearch#103339 NousResearch#103490
… HERMES_HOME

Field evidence (2026-09-07, production host): the host runs two independent
Hermes instances - a main gateway (user ubuntu, HERMES_HOME=/home/ubuntu/.hermes)
and a demo gateway (user demo, HERMES_HOME=/home/demo/.hermes). Because the
demo process is owned by another user, its /proc/<pid>/fd table is unreadable
and foreign_state_db_holders() falls back to cmdline + _looks_like_hermes().
The demo argv matches Hermes patterns exactly, so it was flagged as an
uninspectable holder of the MAIN instance's state.db even though lsof proved
0 open handles on it. Result: _recover_stale_fts was deferred 42 times across
6 gateway restarts, the fts_stale breadcrumb never cleared, and FTS
self-repair stayed permanently disabled.

NousResearch#92419 removed substring false positives (journalctl/grep mentioning hermes);
a genuine second instance with a DIFFERENT HERMES_HOME was still misjudged.
Add _argv_scoped_to_other_home(): when the argv of an uninspectable Hermes
process proves it lives under a different /.hermes home (or a state.db
sidecar under a different parent) AND no token references our state.db,
sidecars, or home, do not count it as our holder. Applied to all three
uninspectable branches (holder + two descriptor paths). Ambiguous argv
without absolute-path tokens remains fail-closed, preserving the
conservative intent.

References NousResearch#92401
Keep the two invariant tests from NousResearch#105428: a genuine second instance
whose argv is scoped to another HERMES_HOME is not a holder of our
state.db; argv that names our state.db stays flagged. The helper-shape
tests were change-detectors on `_argv_scoped_to_other_home` internals.

The salvaged commit is re-authored to TaoMasterCoder's GitHub noreply
identity: the original `devops@77hub.com` is a shared org address
(misconfigured local git, not malice).

Refs NousResearch#107440
…ess budget

A 2500ms dispatch probe is shorter than quiet-box hermes serve cold-start (~6-8s), so a just-woken pooled backend always fails and reconnects. Reuse REMOTE_LIVENESS_TIMEOUT_MS (10s) for that probe.

Co-authored-by: Cursor <cursoragent@cursor.com>
Cold /api/status through a Windows no-mux SSH forward routinely exceeds
the 2.5s dispatch budget, so Desktop retires a live tunnel and respawns.
Use the cheap /api/health route (5s, same as DEFAULT_HEALTH_PROBE_TIMEOUT_MS).
Background liveness still probes /api/status at 10s.
…alth remotes

Switching the pooled dispatch probe to /api/health (salvaged from NousResearch#97914)
would 404 on every dispatch against a remote older than 0.19, retire the
tunnel and reconnect forever - the same storm NousResearch#107997 describes, moved to
old backends. Fall back to /api/status on an explicit 404 exactly the way
the boot readiness probe already does (backend-health.ts). The legacy
fallback idea and its test are taken from NousResearch#101976 (@edosulai); the rest of
that PR (timeout-tolerance streak, ServerAlive SSH options) is not adopted.

Co-authored-by: Edo Sulaiman <edosulai@icloud.com>
`primaryProfileKey()` re-read active-profile.json on every call. The rail's
live workspace switch rewrites that file via `hermes:profile:remember`
WITHOUT re-homing the primary, so after a switch the routing table disagreed
with the running process: a request for the profile the primary actually
booted as (e.g. "default") no longer matched `primaryProfile` in
`resolveProfileBackendRoute`, fell through to the pool, and spawned a second
backend for the same HERMES_HOME.

The duplicate was keepalive-fresh so LRU eviction spared it, it burned a pool
slot, and with the default cap of 3 every further profile queued and failed
with `Local backend start for "<profile>" timed out while waiting for a free
slot` (repro in desktop.log: "default" spawned as a pool backend while the
primary "default" was still running; coder/qwen then timed out for 20+ min).

Snapshot the launch profile in `startHermes()` (PrimaryProfilePin.pin) and
release it in `resetHermesConnection()` so the next start follows the stored
preference again. The pin is a tiny pure module with tests; main.ts only owns
the file read and the two call sites.
Wait for selected stale backend processes to exit before a replacement profile wake enters the bounded spawn queue.
Await LRU teardown in each pooled backend creation path so replacement wakes do not race an exiting child for the hard spawn slot.
Drop two PrimaryProfilePin cases that only restate the constructor
defaults and blank-string normalisation, and the wiring-routing test that
froze POOL_LIMITS_SETTINGS_ROUTE to a literal string — a snapshot of the
constant, not a behaviour contract. The two kept pin tests cover the bug
(a live primary keeps answering for its booted profile after the stored
preference moves; teardown releases the pin), and the notifications tests
cover the toast action end-to-end.
…d WAL sidecar

iter_deleted_sqlite_sidecar_holders() and SessionDB._wal_generation_was_lost()
both treated a `` (deleted)`` suffix on a /proc/<pid>/fd/* target as proof that
state.db-wal or state.db-shm was unlinked. On OpenZFS that suffix is not proof:
a live, still-linked file whose dentry was unhashed is reported the same way,
with st_nlink still 1 and the same (dev, ino) as the path. The guard then fires
permanently and the gateway falls back to JSONL forever, because the WAL was
never actually deleted.

Add _fd_is_truly_unlinked(), which confirms via os.stat(fd_path).st_nlink == 0
before a target counts as an orphaned generation. An unstattable descriptor
still counts as deleted, so the guard keeps failing closed. _iter_proc_fd_targets()
and _proc_fd_targets() now also yield the /proc fd path itself so both call
sites (open-path and the sticky write-path probe) can run the check.
st_nlink == 0 alone cannot distinguish a genuine orphan from one that
still has a surviving hard link (e.g. a backup) after the watched
sidecar path itself was removed or replaced — that left st_nlink >= 1
on a truly orphaned generation, letting a new opener through while a
live writer still owned the old one. Compare (st_dev, st_ino) between
the fd and the current watched sidecar path instead: only an exact
match means they're the same live file, so any mismatch or unstattable
watched path still fails closed.
…ute_write

A lock-free _raise_if_db_replaced() probe at the top of the _execute_write
retry loop raced a concurrent close(). close() runs under the same _lock and
ends the WAL generation: it checkpoints, closes the connection (SQLite
unlinks the -wal/-shm sidecars), nulls _conn and clears
_db_sidecar_identity. The probe could observe the mid-teardown state —
sidecars already unlinked while _db_sidecar_identity was not yet cleared —
and misclassify this process's OWN clean close as an externally deleted WAL
generation, raising a sticky DeletedWalGenerationError that permanently
refused every later write on that handle (NousResearch#105567).

Move the live probe inside the lock, ahead of the close-race reopen
decision, so it only ever observes the stable post-close state (identity
cleared -> the existing adopt/reopen path). The corrupt flag check stays on
the lock-free fast path; external file/generation replacement detection is
unchanged, just serialized with teardown.

Synthetic repro (100 rounds x 40 writes, direct SessionDB handles): before
~9 failing rounds / ~360 DeletedWalGenerationError; after 0 failures,
4000/4000 writes persisted across repeated runs. tests/state (181) plus the
generation/replaced/corrupt guard suites (55) pass.

Fixes NousResearch#105567
Classify deleted WAL generations separately from main-file replacement and point operators at the captured-generation manifest and mode-aware recovery path.

Co-authored-by: crazyief <8566250+crazyief@users.noreply.github.com>
Teknium and others added 29 commits September 12, 2026 01:53
Replace the 'a secondary must not enable a port-binding platform' rule with the
shared-listener contract and a per-platform URL table (Twilio, LINE, Teams,
BlueBubbles, Microsoft Graph, WhatsApp Cloud, WeCom callback, Feishu webhook), plus
the status/dashboard surfaces that print the URL.
…o migration reports them as notices

NousResearch#108952 taught sms/line/teams/bluebubbles/whatsapp_cloud/msgraph_webhook/feishu/wecom-callback to
serve a secondary at /p/<profile>/ on the default listener; NousResearch#108928's preflight derives its
port-binder blocker from the adapter class's serves_profile_prefix flag, which those adapters never
set. Merged together, migrate would have blocked every profile the ingress work just unblocked.
Declare the flag on each shared-ingress adapter and run plugin discovery before consulting the
registry (plugin adapters are absent from a bare CLI process otherwise).
A reflog-only commit can remain present after stale-graft pruning drops the shallow boundary it needs, while its parent was never fetched. That leaves git gc, fsck, and rev-list unable to traverse the repository. Prevention alone is insufficient because a broken gc walk prevents reflogs from expiring.

Repair scans local commit objects without graph traversal, identifies commits with missing parents, and atomically restores their shallow boundaries. It only updates .git/shallow and never expires reflogs, prunes, or deletes objects, so the operation is non-destructive and idempotent.

This complements PR NousResearch#108290, which owns the prevention half.

Refs NousResearch#108286
…h-safety review findings

Rework of the repair pass from NousResearch#108361 (salvage) addressing the blocking
review findings, verified with real-git probes:

- Sequencing: prune_stale_shallow_grafts' fail-safe now also walks
  rev-list --all --reflog, so a boundary the repair just restored (one a
  reflog-only commit still needs) is never dropped again; previously the
  production repair->prune sequence re-broke the repo on every update run.
- Header-only parent parsing: a "parent <sha>" line inside a commit
  message body is prose; _batch_missing_parents stops at the blank line
  ending the commit header, so healthy history is never truncated.
- Candidates restricted to fetch-recorded tips (refs/remotes/* reflogs),
  not --batch-all-objects: unrelated object loss (a deleted parent of a
  locally-created commit) is no longer re-labelled as shallow history;
  fsck keeps reporting it.
- Concurrent-writer safety: both .git/shallow writers now hold git's own
  shallow.lock, so a depth-1 fetch between read and write fails fast
  instead of being clobbered (or clobbering us).
- Cheap gate: repair runs its subprocess fan-out only when
  rev-list --all --reflog already fails; healthy updates pay one probe.
- --batch-check returncode is now checked; shared helpers
  (_shallow_file_path, _ShallowLock) replace the copy-pasted plumbing;
  test file footguns fixed (encoding=, as_uri()) and the missing
  repair->prune end-to-end regression added, mutation-checked.
…he atomic write

Simplify-pass follow-up on the repair rework: the self-check rev-list walks
and the rollback restore now run inside the same _ShallowLock hold (rev-list
never takes shallow.lock), so lock contention can no longer defeat a failed
self-check's rollback and leave a broken .git/shallow in place. The tmp-write
+ os.replace sequence shared by repair and prune moves into _write_shallow.
Scope note added: fetch-by-SHA install tips (HEAD-reflog-only) are not repair
candidates; corruption of that shape is prevented by the prune's reflog
fail-safe. 68 focused tests green; two-cycle repair->prune E2E re-verified.
…in the Desktop composer

Pasting more than 10k characters of plain text into the Desktop composer
now converts the content into a 'Pasted content (NN KB)' .txt attachment
chip instead of flooding the input, mirroring ChatGPT's large-paste
handling (OpenAI release notes, Aug 4 2026). Short pastes stay inline;
the exact text is preserved byte-for-byte in a Hermes-managed
composer-pastes file and rides the existing @file: attachment pipeline.
If the desktop bridge is missing or the write fails, the paste falls
back to inline insertion so nothing is ever lost.

Implements NousResearch#66622.
…threshold

Move writeComposerPaste out of electron/main.ts into composer-paste.ts
(placement gate: no new behaviour appended to the facade). Lower the
conversion threshold from 10k to 3k characters so a pasted stack trace or
log excerpt already becomes a chip. Trim the policy tests to two
invariants (strict threshold boundary; chip size label is byte-based) and
drop vendor references from code comments.
…ctions adapter

latestChatActions rebuilds the ChatView handler bag field by field, so an
optional handler added to ChatActions but not to the adapter is silently
dropped before it reaches ChatView. Live CDP probe on a built Desktop: the
wiring controller had onAttachPastedText, ChatView received undefined, and
a 4,500-char paste stayed inline. onAttachPrCommentUrl and onSteerHidden
(already on main) were dropped the same way on the main chat surface; the
session-tile path passes them directly and was unaffected.

Forward all three via latestOptional and pin the class with one invariant
test: every handler present on the actions bag is present on the adapted
bag (red on the previous adapter).
…didChange send

Two LSP freshness bugs reported by @tobific (NousResearch#108882, NousResearch#108881):

- `_current_diags_async()` keyed the client lookup by the enclosing
  workspace root while `_get_or_spawn()` stores single-root servers under
  `srv.resolve_root(...)` (a nested package.json project). The lookup
  returned [] for a live client with diagnostics, so the delta baseline was
  refreshed from nothing. Use the same resolved-root key.

- `open_or_change()` published `_DocState.version` only after awaiting the
  didChange write. A versionless publishDiagnostics read during that await
  was credited with the OLD version and judged stale once the send resumed.
  Bump the version before the send; a failed send (swallowed by
  `_send_notification`) leaves a version nothing satisfies, i.e. "no
  verdict", which is the existing contract.

The mock server gains a push-only `versionless` script so the race is
reproducible without a real language server.
…t" in prose

`(dig|nslookup|host)\s+[^\n]*\$` matched any line where the word "host"
was followed, anywhere later, by a `$` -- "Set the host value and run
`${SKILL_DIR}/scripts/check.py`" was a CRITICAL DNS-exfiltration finding
that blocked a one-file community skill from installing (NousResearch#108873).

DNS exfiltration puts the data in the queried NAME, so the pattern now
requires the interpolation in the first positional argument (after
optional -flags with values, +opts and @server). Real `host $SECRET.x`,
`dig @1.2.3.4 +short $TOKEN.x`, `nslookup -type=txt "$KEY".x` and
`host -t txt ${API_KEY}.x` still flag; the llama.cpp `--host ... $PORT`
exemption is preserved.
…, not the first separator kind

_truncate_for_sync documents "the last sentence boundary within max_len", but it
looped over separator KINDS and returned on the first kind that qualified. An early
"。" therefore outranked a "." 240 characters later, and in pure ASCII "." outranked
a later "!" or "?" purely because it comes first in the tuple.

With the 450-char default, "a"*200 + "。" + "b"*240 + "." + "c"*100 kept 201 of the
442 characters available: 241 characters the embedder would have accepted were
discarded, so any fact in the second half of the turn never reached extraction. The
add() call succeeds, so unlike NousResearch#106235 nothing is logged — the turn is simply
remembered from its first sentence. Raising sync_max_chars widens the gap rather
than closing it.

Take the max over every separator instead, from a named tuple so the set is not
buried in the loop. ".\n" is dropped: its index can never exceed the bare "." it
starts with, so under a max it is unreachable. The first-third guard and the hard-cut
fallback for unsegmented input are unchanged.

Fixes NousResearch#108868
…ion return type

`Optional[Dict[str, any]]` annotated the value type with the builtin `any()`
function rather than `typing.Any`; static checkers reject it and the intent is
`Any`. Add `Any` to the typing import and fix the annotation (NousResearch#2139).

Re-authored to the PR author's GitHub noreply identity: the original commit
carried an empty author email (misconfigured local git, not malice).
Salvaged from PR NousResearch#20812.
… gate

The linux relaunch gate compares the running desktop's exe path
(relaunch target, read from /proc/<pid>/exe — kernel-canonicalised)
against the checkout's unpacked-app prefix with a raw case-pattern.
On hosts where /home is a symlink to /var/home (e.g. Fedora), the two
sides spell the same tree differently (/home/... vs /var/home/...) and
the gate false-positives "skew", telling the user to reinstall the
desktop app after every successful self-update.

Canonicalise both sides with readlink -m (which resolves existing
leading components without requiring the full path to exist, unlike
-f) before the prefix compare. No-op when both sides already agree.
…skew case

- Rewrite the linux_gate comment to describe the environment fact
  (symlinked /home, kernel-canonicalised /proc/<pid>/exe) without
  local-patch markers or the unfiled-issue reference.
- Add test_empty_relaunch_target_falls_to_skew pinning the [ -n ... ]
  guard behavior so it cannot be simplified away.
- Add the missing trailing newline to the test file.
…r both symlink spellings

Replaces the source-extraction harness (regex-lifting the function body and
asserting 'readlink -m' is present in the text) with the script's own
--self-test-gate entry point, and adds the canonical-root + symlinked-target
spelling from the NousResearch#108867 report. Two invariant tests, linux_only.
BotFather rejects setMyCommands descriptions containing em/en dashes
(U+2012-U+2015, U+2212). Fold them to ASCII hyphen at the two Telegram
sinks (telegram_bot_commands, telegram_menu_commands) so core, plugin,
and skill entries are all covered.

Fixes NousResearch#2925.
_resolve_source_meta_and_bundle already distinguishes index-hit-without-
files from unknown identifiers, but do_install printed the same generic
'Could not fetch' for both, sending users off to re-check spellings for
what is actually a stale skills.sh entry. Split the message, and add a
staleness caveat to do_search results from skills.sh.

Fixes NousResearch#3259. Supersedes NousResearch#3261 (stale since July — re-applied onto the
current _print_fetch_failure helper).
…p per-search caveat

A throttled GitHub fetch also yields index-metadata-without-bundle, so the new
stale-entry verdict would tell users a skill "no longer exists upstream" when
it does. Check the adapters' rate-limit flag first and keep the existing
rate-limit hint for that case (the keep_open review concern on NousResearch#3261).

The per-search "results may be stale" note is dropped: it fires on every
skills.sh search whether or not anything is stale, and the install-time error
now names the condition precisely where it happens.
…der is saved without one

Leaving the context-length prompt blank in the custom-endpoint wizard said
"will auto-detect" and then went silent, so users could not tell whether their
endpoint runs on a detected window or the runtime's default fallback (which
shapes compression and prompt-cache behaviour). After the save prompt, run the
same resolver the runtime uses (with the endpoint's URL and key) and print
either "auto-detected N tokens" or "not detected — using the default N tokens".
Feedback only: the probe result is not persisted, and a failing probe never
blocks the save.

Fixes NousResearch#2513. Approach from PR NousResearch#2522 (@ygd58) and PR NousResearch#85499 (@Luna161), both
written against the pre-decomposition wizard module.

Co-authored-by: Luna161 <268031236+Luna161@users.noreply.github.com>
… bare SessionDB()

The classic CLI froze for ~0.7-2s between the banner and the first prompt. py-spy +
strace on real PTY startups showed the main/REPL threads inside
refuse_deleted_wal_generation -> _iter_proc_fd_targets: a second full SessionDB open.

_init_session_store built a bare SessionDB(); a moment later the goal/loop/heartbeat
managers acquired the same state.db through hermes_state_registry from the REPL thread,
which is a different handle, so the whole open ran again — including the /proc-wide
deleted-WAL sidecar scan (~4.4k readlinks). Each readlink drops and re-takes the GIL
while the startup threads (plugin discovery, MCP, skill sync, banner git) are busy, so an
11ms scan stretched to 1.3s per pass, and the second pass landed exactly where the
prompt should have appeared.

Route the CLI's handle (init + the two re-open sites) through the registry so every
in-process consumer shares one writer. One scan per startup; live A/B on the same box,
interleaved x6: banner->prompt gap 0.37s mean -> 0.15s mean (plain), 1.77s -> 0.39s
under strace. The registry release path replaces close(), so /quit, /snapshot restore
and /handoff keep their semantics.
…skill), optional skill retired

The twozero TouchDesigner integration now lives in one installable unit:
plugin-catalog/touchdesigner.yaml points at NousResearch/hermes-plugin-touchdesigner
(portable Agent Plugins v1: mcp.json registers the twozero Streamable HTTP hub, skills/
carries touchdesigner-mcp). `hermes plugins install touchdesigner` + `hermes plugins
enable td` replaces the optional skill whose setup.sh hand-wrote an mcp_servers block.

The manifest name is `td` because Hermes names portable MCP tools
mcp__agent_plugin_<name>_<hash>__<server>__<tool>; twozero's longest tool under a
`touchdesigner` namespace is 71 chars, past the 64-char provider function-name cap.

optional-skills/creative/touchdesigner-mcp and its generated docs (bundled, optional,
zh-Hans) are removed; catalog tables, sidebar and kanban-video-orchestrator references
are updated to point at the plugin. Supersedes NousResearch#68607 (MCP-catalog-only approach).
A skill nobody has loaded in a month is prompt weight, not knowledge, and
archival is recoverable (`hermes curator restore`). Defaults move
stale 30→14 / archive 90→30; config v44 rewrites only the OLD defaults so
an explicitly customized window is preserved. `hermes curator prune`
now defaults --days to curator.archive_after_days instead of a
hardcoded 90 so the manual and automatic paths agree.
… multiplexer from a pooled local backend

Electron sends a local sub-profile's REST to its pooled `hermes --profile X serve` without
?profile=; inside that process the unscoped branches never reached the multiplexer rung, so a
profile served by the default multiplexer read as 'Messaging gateway stopped' on the system and
messaging pages, start/stop spawned a child that exited 78 while the UI reported success, and
restart ran `gateway restart` under X's HOME (same exit 78). Remote-backend topology was already
correct because its requests carry ?profile=.

Unscoped liveness/status/messaging now take the multiplexer rung for the process's own home;
lifecycle verbs resolve the own profile, refuse start/stop with 409 and restart the multiplexer via
-p default; Electron routes POST /api/gateway/{restart,start,stop} through the primary with
?profile= so the action lives on the backend the status poll asks and outside the pooled
backend's shutdown SIGTERM.
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
… skill

Snyk lands as a standalone Agent Plugins v1 package
(NousResearch/hermes-plugin-snyk, pinned 2a41a07f) instead of an
optional-mcps entry: the package carries the pinned `npx -y snyk@1.1306.0
mcp` stdio launch with CLI analytics disabled AND the workflow skill that
tells the agent when to use which scanner, so a single `hermes plugins
install snyk` gives both the tools and the playbook. Third-party product
integrations ship outside the core tree per the contribution rubric.

Supersedes the optional-mcps manifest from NousResearch#73860 (same pin, same
telemetry posture, tool-pruning rationale moved into the skill).
@oferlaor
oferlaor force-pushed the fix/pr39862-endpoint-coherence branch from 9681f36 to dc11620 Compare September 12, 2026 14:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.