Skip to content

Upstream catch-up: pristine NousResearch main + kanban worker toolset fix - #14

Merged
girnarholdings merged 6168 commits into
mainfrom
sync/upstream-catchup-20260819
Aug 20, 2026
Merged

girnarholdings merged 6168 commits into
mainfrom
sync/upstream-catchup-20260819

Conversation

@girnarholdings

Copy link
Copy Markdown
Owner

What this is

Upstream-first catch-up per operator directive: take pristine latest upstream/main (de32244) and add exactly ONE fork commit on top — the kanban worker toolset fix.

The one fork commit

  • fix(kanban): workers get mutation toolsets despite interactive least-privilege policy (5e74cc7)
    • nima-infra apply_profile_toolset_policy blocks terminal/file/code_execution/delegation/cronjob/kanban/session_search for domain profiles, crippling dispatcher-spawned kanban workers (no bash, no file tools, no board tools).
    • Re-enables mutation toolsets for worker spawns (HERMES_KANBAN_TASK-gated) while the Telegram gateway allowlist keeps interactive surface read-only.
    • Tests: tests/hermes_cli/test_kanban_worker_spawn_toolsets.py (4/4 pass).

Explicitly NOT carried (per directive: latest upstream wins; flagged for follow-up)

All 13 prior fork customizations were dropped rather than overriding upstream:

Verification done

  • Boot imports: cron/gateway/prompt_builder/kanban_db all OK
  • tests/hermes_cli/test_kanban_worker_spawn_toolsets.py: 4 passed
  • tests/cron: 805 passed, 0 failed (full re-run)
  • tests/agent/test_prompt_builder.py: 74 passed (upstream-native)
  • Live jobs.json: parses under upstream scheduler (64 jobs); upstream recovery re-armed a wedged job
  • deliver:/no_agent/script fields all natively supported; no profile: dependency

Merge gate

CI must be green before merge. After merge, the live tree fast-forward is a separate human-gated maintenance step.

jackulau and others added 30 commits August 18, 2026 15:04
The agent path already declined to publish when applyActive() rejected its
activation, but the profile path discarded the same boolean and published
unconditionally. applyActive() returns false when its captured epoch has
been superseded, which happens whenever a newer switch or a teardown lands
while this preparation is still awaiting its route or socket.

The result was not a torn publication. batch() makes those writes
observer-atomic either way. It was something subtler: ONE complete,
internally inconsistent tuple, the CURRENT gateway paired with the stale
target's profile pointer and descriptor. Atomicity cannot make a rejected
activation correct, so the caller has to decline to publish at all.

prepareGatewayForProfile now returns Promise<() => boolean> like its agent
counterpart. The primary and shared-primary thunks return applyActive()
directly; the secondary thunk reports whether the prepared entry was still
current AND the epoch was accepted, keeping the descriptor publish
conditional on having a cached connection so an accepted activation with no
descriptor still moves the companions.

prepareGatewayForAgent's genuinely-local fallthrough now returns the profile
thunk unchanged instead of wrapping it to return an unconditional true,
which had been reporting a rejected activation to the agent caller as a
successful one.

Two regressions on the profile door: a superseded activation leaves all
three stores on the existing complete route with no subscriber notified at
all, and an accepted one still publishes, so a thunk that always reported
false could not pass. The mock thunks in profile.test.ts now return true,
since a bare vi.fn() returns undefined and would read as "superseded".
The comment above ensureGatewayAgent carried both the old and the new
contract on consecutive lines: "a local/null connectionId falls through
to the profile path verbatim", immediately contradicted by "only a null
connectionId falls through, explicit local is a registry identity".
Dropped the stale line.

Same wording above prepareGatewayForAgent in gateway.ts, tightened to
match what the code actually does: registryBackendScopeKey only collapses
to the bare profile key for a null or empty id, so an explicit local id
scopes to conn:local::<profile> and stays on the registry route.

Comments only, no behavior change. tsc --noEmit, eslint and the three
affected suites (43 passed) re-verified.

Refs NousResearch#82140
…uns the loop-thread grace window

Third CI hit today for tests/gateway/test_goal_verdict_send.py (twice on
salvage PRs, once on main's own push run), always the same shape:
adapter.sends == [] after the full drain.

Mechanism (reproduced, not log-read): the tests call GoalManager.set() on
the event-loop thread. _get_session_db() refuses to construct SessionDB on
a loop thread (loop-liveness guard from the 2026-08-14 crash-loop fix) and
waits only _DB_BOOTSTRAP_LOOP_WAIT_S=0.25s for the background bootstrap.
On a loaded CI runner the init overruns that window, set() degrades to a
silent no-op by design, the goal never exists, and the continuation path
correctly does nothing — so no amount of drain-waiting helps (the NousResearch#88975
de-flake addressed a different, downstream race).

Fix: the hermes_home fixture pre-warms the SessionDB cache from its sync
context (direct construction path), so the loop-thread set() always finds
a cached DB. Production behavior untouched.

Proof: injecting a 0.4s-slow SessionDB.__init__ reproduces the exact CI
failure on the old fixture and passes 2/2 with the pre-warm.
… bot sees

Group rooms were text-only on the ingest side: the composer and
runGroupChatMemberTurn submitted prompt.submit {text} with no attach step,
so a user screenshot could never reach the members' models (community
report from Osiris). The gateway already ships the staging pipeline
(image.attach_bytes -> attached_images -> next prompt.submit) and the 1:1
canonical chat uses it — this wires the group surface to the same path.

- Composer + thread reply boxes: attach button, Ctrl/Cmd-V paste, pending
  chips with preview/remove, image-only sends allowed.
- Attachments are downscaled (long edge 1568px) and stored on the room-log
  entry, so reloads keep showing what members were shown.
- Turn drive stages the delta's images into EVERY responding member's own
  per-group session via image.attach_bytes before its prompt.submit —
  works cross-connection through requestForBot; a failed attach degrades
  that member to text-only rather than failing the turn. Watermarks
  guarantee an image is staged at most once per member.
- formatGroupChatLine names attachments ([attached image: name]) so the
  transcript delta and the staged pixels line up for every viewer.
- Room log renders attached images on user entries.

Tests: 6 new vm-harness tests (fan-out staging order, mention-scoped
attachment routing, image-only sends, no re-attach across turns,
transcript naming, invalid-attachment degradation); 276 total pass.
…s the

wake instead of the full runtime boot (NousResearch#89206 class)

zero trust's third bundle (on ae6578a, both prior fixes present) showed the
remaining failure: cold profile backends on slower Windows machines take
47-120s to fully boot, while the wake path's fixed budgets (20s hydration,
~15s resume retries) raced the whole boot and lost — "errors waking up BOTS"
while the backend came up healthy moments later.

Rather than raising timeouts, make the wake cheap:

- waitForFocusedSessionHydration: a history-bearing chat is hydrated when
  the persisted transcript is PAINTED on the right session. The REST
  prefetch delivers that seconds after the backend's HTTP is up; the full
  runtime resume (agent build, MCP discovery, 114-skill load) keeps warming
  in the background and binds the composer when it lands. Only an
  expected-empty chat still waits for the runtime (nothing to paint).
- On hydration timeout, log a [bot-wake] phase breakdown (activation ms,
  hydration ms, which conditions were unmet) to the renderer console so the
  next support bundle pinpoints the slow phase directly.
- web_server: flush the headless "listening" line — block-buffered on the
  Desktop's piped stdout, it surfaced minutes late and made boots look far
  slower than they were in support bundles (the 120s "gap" in this bundle
  was partly this artifact).

Sabotage-proven: restoring the runtime-gated wait fails the new paint-first
test by timing out — the exact field shape.
Second layer of the "make hydration feel instant" work (on top of the
paint-first wait): persist a bounded tail (40 msgs / 256KB / 50-session
LRU) of every reconciled transcript in localStorage, keyed by durable
stored-session id.

- Cold resume paints the cached tail immediately — the wake is visually
  complete before any network I/O; the REST prefetch / runtime resume
  reconcile the authoritative transcript over it when they land.
- The cached paint is DISPLAY-ONLY: reconciliation treats the view as
  empty (viewMessagesForReconcile), so authoritative content replaces the
  provisional paint wholesale — no grafting onto stale rows. Failure
  latches also treat it as empty, so a cached paint can never mask a
  genuinely stranded resume; a resume that proves the session empty rolls
  the paint back and drops the poisoned entry.
- Saves happen only post-reconcile (cold path + warm activate path);
  deletes drop the entry; a gateway/mode re-home wipes the cache (another
  backend can recycle stored ids).

Sabotage-proven: disabling the cache load fails 5/7 cache tests; the
display-only contract is covered by the existing resume reconciliation
suite (663 tests green).
… the agent's name

New agents now get a blobatar — a deterministic soft-body face generated
from the bot's name (same name, same face, forever) — as the default
shapes mode, with full manual control:

- Face follows the name live while typing in New Agent
- Randomize re-rolls the seed; Lock face pins the current one so a later
  rename can't change it (Unlock returns to name-following)
- Any of the six silhouettes (round/organic/boxy/nub/cloud/sun) can be
  pinned via frozen-per-major trait positions while the rest stays
  name-derived
- Classic geometric shapes remain one click away, and existing bots keep
  their stored looks untouched

Wiring: blobatar@0.2.0 (zero deps, ~3.7KB) exported through the plugin
SDK (blobatarSvg / Blobatar), feature-detected in plugin.js with a
legacy-shape fallback for older desktops. Blob shape strings are
'blobatar[:seed[:kind]]' inside the existing meta.shape field, so
persistence, cross-machine ui_meta sync, and the roster's PNG backfill
(data-bot-face tag preserved) all work unchanged.
A fresh state.db init (schema DDL, FTS tables, first config import)
measures ~300ms warm on a fast machine. The gateway constructs
GoalManager on the event-loop thread, and a cold cache ran that init
behind a 0.25s bootstrap grace window: on a slow CI box the /goal set
path's waits expired and save_goal silently no-oped — the reply said
"Goal set (7-turn budget)..." but nothing persisted, and a fresh
GoalManager read back no state (first assertion passes, second fails).

Two changes, one per caller shape:

- Async callers (_get_goal_manager_for_event,
  _get_heartbeat_manager_for_event, _post_turn_goal_continuation, and
  the heartbeat poller) warm the SessionDB cache off-loop through the
  context-preserving executor before constructing the manager (shared
  _warm_goals_session_db helper). The loop never blocks and the first
  write lands at any init duration. A bare to_thread would lose the
  per-turn profile home override under multiplex; the executor hop
  keeps it (same pattern as the goal judge path).
- Sync callers (heartbeat persistence, _goal_still_active_for_session)
  cannot await, so the bootstrap windows stay: the call that starts the
  bootstrap waits a one-time init window (1.5s) instead of the short
  per-call window (0.25s), giving healthy cold inits room to land while
  a contended migration still degrades to None with only a bounded
  one-time stall. The bootstrap thread binds the caller's home as a
  contextvar override so a multiplexed worker cannot cache the default
  profile's DB under another profile's key.

save_goal and heartbeat save_state now log at WARNING when they drop a
write, because the reply has already told the user the state was set.
Regression test pins the contract: init past the window, write
persists, loop gap under 2s (the flake-policy floor for wall-clock
bounds; the slow-init margin grew to match, so the test still tells
on-loop from off-loop).

Independent diagnosis + measurement by jackulau (NousResearch#88965 review); the
off-loop warm-up shape follows their harness table. Simplify-code
review (4-agent) contributed the helper extraction and the poller
warm-up.
The loops delegation (previous commit) moved /loop onto the shared
bootstrap windows, but the loop-class gateway callers still constructed
LoopManager on the loop thread with no warm-up — the same false-ack
class as /goal, one sibling over:

- The /loop command handler: a cold init past the window made
  save_loop discard the write while the reply claimed the loop was
  set. Reproduced with a 2s init: "↻ Loop set" with nothing persisted;
  with the warm-up the loop persists.
- _post_turn_loop_completion: a cold cache at the turn boundary
  stalled the loop for the init duration and could drop the
  tick-completion write.
- _loop_wakeup_watcher: the scan reads every persisted loop, so a cold
  cache ran the state.db init on the loop thread before the first
  read.

All three now warm the cache off-loop first (same helper, same reason
as the goal paths). save_loop also logs at WARNING when it drops a
write, matching save_goal: the reply has already told the user the
loop was set.

Review findings (PR NousResearch#88965): the goal and heartbeat command paths
warmed off-loop, the loop-class paths did not.
Review fixups for NousResearch#88965. The goal, loop, and heartbeat managers each
had a copy of the same WARNING text. The shared _warn_dropped_write
helper in goals.py keeps the three logs identical and greppable as one
bug class. The _warm_goals_session_db parameter is now label. The old
name ctx said context, but the value is a log label.
Review follow-up on the salvaged NousResearch#88965 work:

- test_goal_command_slow_db_init_still_persists: drop the 4s slow-init
  loop-gap harness (wall-clock gap assertions on shared CI runners are
  their own flake class; loop-freeze bounds are already covered in
  test_goals_db_bootstrap_off_loop.py). The persistence contract keeps
  its discriminating power by shrinking the monkeypatched init window
  (0.2s) under a 0.8s slow init — the window-only path still fails it.
- test_slow_construction_does_not_block_the_loop: monkeypatch both
  bootstrap windows down (0.3s/0.05s) and shrink the blocking init from
  6s to 1.5s; the two-window contract is what's under test, not the
  production constants. Adds an elapsed ordering assertion so the
  kick-vs-in-flight window distinction stays pinned.

Combined wall time for the pair: ~10.5s -> ~1.4s.
Completes group/1:1 attachment parity (NousResearch#88983). PR NousResearch#89486 covered images;
this adds the remaining half:

- The composer picker accepts any file type; kind (image/pdf/file) decides
  the staging RPC. Paste handlers accept non-image files too.
- Drag & drop anywhere on the room drops into the active composer (open
  reply box, else main), with a drop overlay naming the target.
- Member turns stage PDFs via pdf.attach (rendered per-page into vision
  tiles by the gateway) and other files via file.attach; each returned
  @file: ref is appended to that member's turn prompt so file tools can
  read the artifact. Failed attaches still degrade to text-only.
- Transcript markers distinguish [attached PDF: x] / [attached file: x] /
  [attached image: x]; room log renders non-image attachments as named
  chips, pending chips show type icons.

Tests: 3 new vm-harness tests (per-kind RPC routing across members,
@file: ref injection into turn prompts, transcript labels); 279 total pass.
…er-replay sidecar

Bot-mode interrupted member turns with no visible assistant text persisted an
empty assistant row (content="" + display_kind="hidden"). The pre-call
sanitizer repair_empty_non_final_messages() re-healed that row on every later
call (wire copy only), so the loop never converged (NousResearch#88955).

Stamp api_content="[response interrupted]" (the canonical
_INTERRUPTED_PLACEHOLDER) on the hidden placeholder instead. display_kind is
stripped before sanitization, but api_content is projected back into content
for historical assistant rows, so the provider sees a non-empty neutral turn
and the sanitizer stops touching the row — while the durable transcript stays
hidden and empty. Uses the neutral interruption text, never the
_INTERRUPTED_SCAFFOLD_MARKER, which replaying as assistant text caused NousResearch#81841.

Adds regression coverage proving (A) the placeholder carries the replay
sidecar, (B) two consecutive projections converge without sanitizer healing,
(C) the sanitizer still repairs genuinely-empty unmarked assistants.

Refs NousResearch#88955
…entries

The tracked contributors/emails/agent@Agents-Mac-mini.local and
agent@agents-Mac-mini.local differ only by case, which cannot materialize on
case-insensitive filesystems and surfaces one of them as perpetually modified
in git status, breaking clean checkouts. These entries are stale agent
identifiers, not real contributors; remove both.
…payload at projection time (NousResearch#88955)

The salvaged writer-side fix stamps api_content on NEW hidden redirect
placeholders, but rows persisted before it (content="" + display_kind=hidden,
no sidecar) would keep re-triggering repair_empty_non_final_messages on every
call forever. Substitute [response interrupted] on the wire copy at the
api_content/display_kind projection stage so legacy sessions converge too.
Never the interrupt scaffold (NousResearch#81841). Durable transcript untouched.

Regression tests drive run_conversation end-to-end with a spied sanitizer:
the projection must leave the sanitizer nothing to heal (its per-turn warning
spam is the bug), verified failing via sabotage run against the writer-only
fix.

Projection-side approach credit: @JoaoMarcos44 (PR NousResearch#88996).
The deeplink-driven plugin install flow shipped in NousResearch#89464 (salvage of
NousResearch#82735 by @serefyarar) had no docs. Adds:

- user-guide/features/plugins.md: "One-click install links (Desktop)"
  section under Managing plugins — link forms (repo/enable/force), the
  confirm-first dialog contract (never auto-installs, same install-time
  security scanning as the CLI), hybrid-repo behavior, legacy
  plugin-agent/plugin-desktop routing, hermes-dev:// in dev builds, and
  the no-SDK anchor example. Cross-links the MCP "Add to Hermes link"
  equivalent.
- developer-guide/desktop-plugin-sdk.md: "Distributing with an install
  link" section so plugin authors find the link form next to the
  packaging docs.
…pending turns (NousResearch#86580)

The salvaged fix persisted the recovery note unconditionally — correct for
the synthesized empty auto-resume turn, but a user who typed real text while
resume was pending would get the [System note: ...] scaffold persisted as
their own words, leaking scaffolding into the durable transcript (the same
class as NousResearch#81841 on the assistant side).

_prepare_resume_pending_message now persists the note only when the original
message is blank; real text persists verbatim while the model still receives
the wrapped note. Whitespace-only counts as blank. Tests cover all three
shapes.
Add a pane-level opt-out for the hover close button and apply it to the persistent Sessions and Bots navigation panes. Keep their existing close handlers and other tab behavior intact.\n\nFixes NousResearch#89546
…tures, right-click + Cmd-K toggles

Builds on NousResearch#89551 (@calvinnwq, cherry-picked): his showCloseButton flag
hid the hover X; this completes the model so standing chrome can never
be closed at all, only shown/hidden (NousResearch#89546).

- hideOnly pane chrome (sessions + Bots): no hover X, no middle/meta
  click close, no Close verbs in the tab menu, excluded from
  close-others/right/all sweeps
- zone right-click menu gains Show/Hide rows for the strip's chrome
  tabs (Hide bots / Show sessions, localized in 6 locales)
- Cmd-K palette: auto-registered "Toggle <tab> tab" rows for every
  hideOnly pane, on-screen truth semantics, plugin panes included via
  registry subscription
- hides persist across launches (survive the enforced dock re-adopt);
  reveal intent and Layout reset clear them
- last-visible-tab guard: hiding the zone's last shown tab is refused
  with a toast, so the strip can never become an empty dead zone
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
…opic transport

Plugin structured completions (plugin_llm.complete_structured) build an
OpenAI Chat Completions response_format payload in extra_body. The
anthropic_messages transport forwarded it verbatim, and strict
Anthropic-compatible gateways reject it with HTTP 400:

  response_format: OpenAI Chat Completions structured-output shape is
  not supported. Use output_config.format = {"type": "json_schema", ...}

Observed live: every discord-thread-autotitle structured call failed
for 2+ days (1,600+ logged errors) once the main provider became an
anthropic_messages gateway.

Fix: _translate_anthropic_response_format converts
- json_schema  -> output_config.format = {type: json_schema, schema: S}
- json_object  -> permissive object schema (SDK 0.87.0 has no
  schema-less JSON mode)

merging into any existing output_config (adaptive-thinking effort
coexists) and excluding response_format from the raw extra_body
passthrough alongside the existing reasoning exclusion. The async
adapter delegates to the sync adapter via asyncio.to_thread and is
covered by a test. Non-Anthropic transports are unchanged.
…adapter

The adapter builds the Messages body from a fixed allow-list of kwargs.
A caller that passes response_format as a top-level kwarg (the OpenAI
SDK call shape) got it dropped on the floor. The request succeeded, but
the schema contract silently became prompt compliance. No in-tree
caller uses this shape today. The pin-test makes sure that a future
refactor cannot open this leak again.

The top-level kwarg gets the same output_config.format translation as
the extra_body shape. When a caller sends both shapes, the extra_body
value wins because every in-tree caller uses that shape.

Pin-test pattern from PR NousResearch#85626 review follow-up.

Co-authored-by: Matt McClean <mmcclean@amazon.com>
Some providers reject the structured-output request field with a hard
400. The error classifier marks a 400 as non-retryable, so one rejected
field failed the whole auxiliary call. Session titles stayed derived
forever (NousResearch#82816), and no fallback fired.

Three rejection shapes are covered, from live reports:
- vLLM gateways translate response_format into guided_grammar and fail
  when the grammar backend is absent (compile_grammar_error: No module
  named 'xgrammar').
- Some OpenAI-compatible endpoints answer "This response_format type
  is unavailable now".
- Anthropic-compatible gateways that predate structured outputs reject
  the translated field: "output_config: Extra inputs are not
  permitted". The documented case is the bedrock-mantle Messages
  endpoint.

The fix is reactive, the same pattern as the temperature and
max_tokens rungs: when the provider rejects the field, retry once
without it. Callers tolerate an unconstrained reply — the title prompt
demands bare JSON and _extract_title_text has a loose-JSON fallback —
so the call succeeds with prompt compliance instead of failing. The
retry only fires when the request carried the field, and both the sync
and async paths get the same rung.

Closes NousResearch#82816
Hermes is an agent for one person. The credentials, the memory, the
sessions and the cron jobs all belong to that person. But the only
declarative path was a NixOS system service. Issue NousResearch#9056 asks for the
user-level equivalent. 25 public Nix configurations already write one by
hand, and several of them copy nix/nixosModules.nix and edit the systemd
part.

This module is not a second copy of that file. The code that both modules
share moves into nix/moduleCommon.nix:

  - the options
  - the renderers for config.yaml, .env and the documents
  - the activation body
  - the command lines of the processes

nixosModules.nix keeps only the parts that need root. Those parts are the
service user, stateDir, addToSystemPackages, container mode and tmpfiles.
The file goes from 1008 lines to 666.

`services.hermes-agent` is now the same option set on both modules. A
NixOS example works on Home Manager without a change, and an option added
one time appears on both.

The Home Manager module is different only where it must be. It uses
systemd.user.services on Linux and launchd.agents on Darwin. It uses
home.activation and not system.activationScripts. It sets HERMES_HOME
directly, with the default ~/.hermes, so an existing directory continues
to work. It uses the modes 0600 and 0700, because the state has one user
and does not need the group-shared umask of the NixOS module. It does not
support container mode, which needs root and the Docker socket.

The change also makes four corrections that apply to both modules:

- backend.mode runs `hermes serve` or `hermes dashboard`. Both modules
  had only the gateway. But Hermes Desktop and the web dashboard connect
  to a different process, so six of the configurations in public repos
  add a second unit by hand. serve and dashboard are one entry point with
  one flag of difference, and you can run only one of them. Thus the
  option is an enum. The NixOS module asserts against container mode with
  a backend, and does not make a unit that cannot start.

- hermesHomeFiles installs files into HERMES_HOME. The `documents` option
  installs into the working directory, which is correct for AGENTS.md but
  wrong for SOUL.md and memories/. Hermes reads those files from
  HERMES_HOME, in agent/prompt_builder.py:2095. A SOUL.md in `documents`
  made a workspace file that Hermes never loaded as the identity. The
  documentation said this in prose, but two directory diagrams showed the
  opposite. This change corrects both. A key in either option can now
  contain subdirectories.

- `documents` needs an explicit `workingDirectory`. The default of that
  option is bad on both modules. It is the home directory of the user on
  Home Manager, and ${stateDir}/workspace on NixOS. A user who declares
  workspace files without a directory therefore gets a place that the
  user did not select. The place is also different on each module. The
  modules now refuse that combination.

  The test is on the priority of the option and not on its value. An
  option that nothing sets keeps the priority of its own default, and
  each definition from a user is stronger. Thus a directory with the same
  text as the default still counts as a selection, and so does a
  mkDefault. A comparison of values detects neither case.

- Each activation writes .env again from a base in the Nix store, and
  does not add to the file that exists. Thus a second activation cannot
  put the same secret in the file two times, and a removed
  environmentFile goes away. environmentFiles keeps the type `listOf
  str` and not `path`, so Nix cannot copy a sops-nix or agenix path into
  the Nix store, which all users can read.

- HERMES_MANAGED and the .managed marker now hold the name of the system
  that manages the install. Thus a refusal says "managed by home-manager"
  and not "managed by NixOS", and `hermes update` gives the Nix guidance
  for both shapes. The CLI does not print a rebuild command for each
  system. It names the owner, and the user knows their own tool. A bare
  `true` and an empty marker still mean NixOS, so this does not change an
  existing install.

Verification. Six new checks, all built:

  nixos-module           evaluates the module with evalModules and the
                         NixOS module list. It asserts both units, one
                         HERMES_HOME, and that the module refuses
                         container mode with a backend.
  home-manager-module    evaluates the module with the
                         homeManagerConfiguration function of
                         home-manager. The process assertions run against
                         systemd units on Linux and launchd agents on
                         Darwin.
  module-option-parity   asserts that each shared option is on both
                         modules, and that the two exclusion lists name
                         only options that exist.
  env-file-assembly      runs the real .env script and checks the
                         contents, the mode, that a second run gives the
                         same bytes, and that a removed file goes away.
  workspace-files-need-a-directory
                         checks that the module refuses `documents`
                         without a directory, and accepts a directory
                         that has the same text as the default.
  service-argv           runs each command line that the modules build
                         through the real parser of the CLI, with one
                         sentinel flag added, and requires that argparse
                         refuses only the sentinel.

`nix flake check` passes, with 21 checks in total.

The CLI branches that treat an install as a Nix install move to one
helper, is_nix_install_method. Four call sites in main.py, web_server.py,
update_cmd.py and doctor.py tested the literal set {"nix", "nixos"}, and
each one missed home-manager. recommended_update_command asks the managed
state before the code-scoped stamp again, because a managed install can
carry a stale stamp that names an update path the managed guard refuses.
The metrics contract gets a home-manager bucket, so a Home Manager
install does not report as unknown.

Each check was mutation-probed. 22 faults were injected, and the checks
caught all 22:

  - a lost --no-open
  - a backend that runs the gateway
  - an overwritten config.yaml
  - documents in the wrong directory
  - a different HERMES_HOME on the two processes
  - a lost HERMES_HOME export
  - a missing backend unit
  - a removed assertion
  - an .env file that grows at each activation
  - an install that reports NixOS
  - an empty .managed marker
  - an option on the NixOS module only
  - a stale entry in an exclusion list
  - a renamed subcommand
  - an unknown flag
  - the workspace-files assertion always passes
  - the assertion compares values instead of priorities
  - an off-by-one that lets an untouched default through
  - the assertion also fires for hermesHomeFiles
  - a mkDefault no longer counts as a selection
  - the Home Manager module stops wiring the assertion
  - the NixOS module stops wiring the assertion

The 16 Python tests in tests/hermes_cli/test_managed_install_shapes.py
were probed the same way. 8 faults were injected and 8 were caught.

These tests fail on this tree. They fail in the same way on the stashed
HEAD, and they have no relation to Nix:

  - test_git_probe_tree_kill.py (2 tests)
  - test_update_import_guard.py (1 test)
  - test_telegram_media_read_timeout.py (2 tests)
  - test_teams.py (a collection error)

Closes NousResearch#9056

# Conflicts:
#	hermes_cli/main.py
#	hermes_cli/update_cmd.py
#	hermes_cli/web_server.py
The docker.yml gate held its own copy of the build formula, in shell.
classify_changes.py now owns a derived docker lane, and the nix lane in
the next commit derives from the same file. Two formulas in two
languages drift apart, and one Python function with tests does not.
The workflow owns its triggers and ci.yml does not call it. A
reusable-workflow call holds the caller run in progress for the full
build, and GitHub refuses `gh run rerun` on a run that is still in
progress. A separate run reruns and cancels on its own.

The job restores /nix/store from the GitHub Actions cache and saves from
main only. A cache that a PR writes is visible to that PR alone, so a
save there spends the quota of the repository and helps no later run.
detect_install_method reads the stamp against an allowlist. The
allowlist held "nixos" but not "home-manager", and a stamp that names
home-manager gave "unknown". The managed path (step 3) returned the
correct name, so the gap was invisible: it appeared only for an install
that carries a stamp.

An install with the value "unknown" gets "hermes update" as its update
guidance. That command is the one command a managed install refuses, so
the user gets a dead end.

The test for this was also environment-dependent. It called the real
get_project_root(), and it passed here only because this worktree
carries no stamp. A checkout from the curl installer carries a "git"
stamp, and the assertion then failed for the contributor and not for
us. The test now detects against a temporary install tree.

The new test stamps each managed system and asserts the value that
comes back. With the allowlist reverted, the home-manager case fails
with "assert 'unknown' == 'home-manager'". The nixos case passes,
because that name was already in the allowlist.
The clarify tool gets an optional questions parameter (2-5 independent
questions, issue NousResearch#18450). Batch-capable platform callbacks receive the
normalized list in one call and reply with per-question answers. Legacy
callbacks are looped one question at a time. The loop stops on timeout
so the user is not asked the remaining questions after they walk away.
Locked answers survive a timeout: the result carries them plus a
timed_out flag, and unanswered entries have an empty user_response.

The single-question path is byte-identical to the previous behavior.
One clarify.request carries the question list (qid, question, choices,
multi_select per entry). clarify.respond gains an optional question_id:
each respond locks one answer, a repeat respond overwrites it, and the
batch resolves when every question is locked. A respond without
question_id keeps its existing meaning (cancel the whole prompt).

Locked answers survive the deadline: a timed-out batch returns the
partial answer map with a timed_out flag instead of an empty string.
The reconnect replay snapshot also carries the locked answers, so a
reattached client restores its per-question state.

Both agent-side clarify dispatch sites forward the questions arg.
teknium1 and others added 25 commits August 19, 2026 16:10
New tests/tools/test_strict_provider_selection.py covers read_selection
semantics (legacy use_gateway interpretation, seeded stt local, empty
strings, browser.backend vs cloud_provider) and the three strict
behaviors per category: managed 'nous' selection wins over present
direct keys, a vendor selection with missing credentials raises the
selection-naming error with NO managed call, and never-configured
installs keep today's autodetect. Updated the tests that pinned the old
credential-first precedence (TTS resolver gateway override, STT silent
managed fallback, web invalid-backend reroute, video_gen picker writes).
Sabotage-verified: reverting the image FAL strict switch makes the new
managed-selection tests fail.
…er provider-string migration

Two real gaps the CI-red sibling tests exposed:

- read_selection() treated EVERY raw stt.provider: local as the legacy
  DEFAULT_CONFIG seed and reported no-selection — but the seed never
  reached config.yaml (save_config strips schema defaults), so a
  picker- or hand-written local pick was silently discarded and the
  autodetect ladder could route an explicit local user to cloud STT.
  A raw 'local' is now a genuine selection; the merged-view ambiguity
  note replaces the over-broad shim (mirror comment updated in
  nous_subscription._selected_provider and _get_provider).

- _reconfigure_provider was half-migrated: the tts/stt/browser/web
  branches and the managed-category fallthrough still wrote
  use_gateway flags and vendor names for managed rows. They now write
  the single provider string ('nous' for managed rows) and pop the
  legacy key, matching _write_provider_config.
Update the sibling tests that pinned the old use_gateway-writing
contract: image/video selector and reconfigure rows now assert the
single provider string ('nous' managed / 'fal' BYOK) plus legacy-key
popping, the stt/video picker writes drop the use_gateway expectation,
the web_server managed-browser select asserts the persisted 'nous'
cloud_provider, and explicit-local STT pins no-cloud-fallback against
a stored raw-config selection.
… picker normalize

The SDK now returns the registry rows per its documented contract
(salvaged NousResearch#89893), while desktops predating the SDK unwrap resolve the
raw registry envelope. The plugin normalize accepts both, so the picker
works across the transition; regression test updated to pin the
dual-shape normalize.
…ls (config model.execution_guidance)

Un-fences OPENAI_MODEL_EXECUTION_GUIDANCE from the gpt/codex/grok substring
check and gives it its own injection gate, independent of
tool_use_enforcement, controlled by config.yaml `agent.execution_guidance`
(auto/true/false/list — same semantics as tool_use_enforcement). The "auto"
list (EXECUTION_GUIDANCE_MODELS) now also covers deepseek, kimi, qwen, glm,
minimax, mimo, and mistral.

Composio agentic-eval traces showed Hermes+DeepSeek/Kimi failing where
competitors passed: financial math done in prose, no read-back after
external writes, malformed identifiers "repaired", completeness claimed
despite count mismatches. The discipline block existed but those models
never received it.

The block is extended with compact clauses distilled from that analysis:
- external-write read-back (tool-call success is not task success; internal
  file edits already confirmed by the tool are not re-verified)
- count reconciliation (declared totals/has_more are hard assertions)
- literal preservation (never normalize identifiers that fail a stated
  format; lookup success does not validate a malformed token)
- retry-differently (empty/partial/suspiciously narrow results get a
  broader retry before concluding)
- completion gated on verification (done = every named acceptance
  criterion verified, never a plausible subset)

The todo tool description now encourages enumeration-as-checklist for
"all N items" tasks and gates completed status on verified work, never
intent.

Guidance is chosen once at session start keyed on model name, so the
system prompt stays byte-stable for the life of a conversation.

Supersedes/absorbs prior contributor proposals: NousResearch#20588, NousResearch#35087, NousResearch#41874
(MiMo), NousResearch#53847 (GLM tool-calls-as-text stall).

Co-authored-by: Mat-London <56627804+Mat-London@users.noreply.github.com>
Co-authored-by: intelac <8803887+intelac@users.noreply.github.com>
Co-authored-by: 6ylqq <51219463+6ylqq@users.noreply.github.com>
Co-authored-by: tauros1983 <267660491+tauros1983@users.noreply.github.com>
Composio-style MCP servers return un-paginated 22-47K-char payloads that
sail under the generic 100K per-result spillover threshold, bloating
context and ballooning per-turn reasoning time on long conversations.
Competitors cap harder (OpenCode/pi 50KB, Claude Code 30K, Codex ~10K
tokens). Three changes:

- mcp_* tools spill at a tighter 50K default (BudgetConfig.mcp_result_size,
  config-overridable via tool_budget.mcp_result_size_chars; pinned and
  per-tool overrides still win; capped by the context-scaled default).
- The persisted-output preview now teaches recovery: page the saved file
  with read_file or process with execute_code instead of re-requesting the
  same data from the remote API.
- Untrusted/MCP string results are scanned (bounded, first 64KB) for
  provider-side elision markers ('...N more items', "has_more": true,
  'saved to sandbox', data_preview) and get ONE cache-safe incompleteness
  notice appended at result-construction time, before untrusted wrapping —
  so the model stops treating provider-elided enumerations as complete.
- Hard 2M-char allocation cap in mcp_tool.py (text, error, and
  structuredContent paths) so a pathological multi-MB server payload is
  bounded before it propagates, while ordinary large results reach
  spillover intact. Distilled from NousResearch#56060/NousResearch#56072/NousResearch#56511 (issue NousResearch#56059);
  supersedes their 50K lossy truncation with spillover-friendly semantics.

Docs: configuration.md spillover-budget section + cli-config.yaml.example.

Co-authored-by: Stoltemberg <215755014+Stoltemberg@users.noreply.github.com>
Co-authored-by: AlexFucuson9 <295703459+AlexFucuson9@users.noreply.github.com>
Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
The persistent terminal is a position:fixed overlay that chases its slot's
rect, and the whole tracker — visibility included — was gated behind the
renderer pause. Switching tabs while the window is unfocused therefore left
the overlay parked over the zone at full opacity with pointerEvents:auto, so
the chat underneath was unreachable until something refocused the window.

Visibility is correctness rather than perf, so sample it on every wake even
while paused; the rect chase, which is the part that forces layout, stays
gated.
…caled stale timeouts (agent.run_budget_seconds / --run-budget)
…-intent recovery (agent.stall_guards)

Composio eval traces showed Hermes wasting turns re-issuing identical tool
calls (same tool, same args, same result — 3x/4x in one run) and ending
turns by announcing an action it never took. Two conservative, config-gated
guards (agent.stall_guards, default true):

- Identical-call loop breaker: ToolCallGuardrailController.observe_identical_call
  tracks the consecutive streak of (tool, canonical args, result-hash); on
  the 3rd identical call a compact one-line notice is appended to that tool
  RESULT at construction time (cache-safe — tool results are append-only).
  Never blocks the call. Pollers (process, *_get_result, *_poll) are exempt
  via STALL_GUARD_REPEATABLE_TOOLS. Streak resets on any different call,
  changed result, or new turn. Observed on the raw result before the
  tool-loop warning suffix so its changing count can't defeat matching.

- Said-continue-but-stopped recovery: trailing_continue_intent() detects a
  short reply ENDING on an announced next action ('Let me now…', 'I will
  now…', 'Next, I…'); the conversation loop feeds it into the EXISTING
  intent-ack continuation path (same interim-assistant + user-nudge
  mechanism, same codex_ack_continuations cap of 2), preserving message
  alternation — no parallel recovery machinery.

Config: agent.stall_guards in DEFAULT_CONFIG; docs in configuration.md;
unit tests for streak/allowlist/reset/gate and detector pos/neg cases.
…ab-trap

fix(desktop): terminal pane no longer traps the window when you switch tabs
…p uses

The HUD asked for vibrancy directly and always with the 'hud' material —
one of the two rungs the macOS census rejected, because it collapses into
under-window on blur and so changed the frost the moment another app took
focus. It also ignored the translucency setting entirely: Glass off still
frosted, and Windows got nothing at all.

hudFrostFor is the mapping for a transparent window, beside vibrancyFor in
the shared module both processes read. Two gates give it its answer: the
renderer's report that the band actually covers the window, and the user's
Glass setting. Off resolves to no material rather than a resting one, since
a transparent window has no opaque page to hide an unwanted frost behind.

Windows 11 rides setBackgroundMaterial through the same call, so the HUD
follows the frost ladder on both platforms. Main self-diffs and keys the
latch to the window, so a Settings change re-frosts a live HUD, a tint drag
touches nothing native, and a HUD respawned on another profile is not
mistaken for the window that already carried the material.
…lass

The band wore its own card tint at a hardcoded 80/92%, so a HUD beside the
docked window read as a lookalike rather than the same surface, and the Tint
slider moved one and not the other. It now paints --ui-bg-chrome at
--translucency-glass-keep: one painter, one token, one lever.

That needed the setting and the surface rewrite to stop being one flag.
data-hermes-glass means "this window's field surfaces may be rewritten" and
is deliberately false in the HUD, which owns its own backgrounds; the new
data-hermes-glass-on means "the user's Glass setting is live" and is
published everywhere, along with the tint number the band reads.

The 0.5rem side inset drops to zero while glass is on. It exists to keep an
opaque sheet clear of the bar's corner controls, but the frost is the whole
window — an inset sheet left a hairline of bare untinted material down both
sides.

An open completion drawer now drops the frost along with the band it belongs
to. The drawer takes the band to 25% and blurs it while the native material
stayed at full strength, which is the same bare slab in a different
disguise. It mounts without a focus change, so it is observed rather than
passed in, coalesced to a frame because the shell mutates with every
streamed token.
Dictation, spoken replies, the wake word and start-conversation were four
separate icon buttons in a Spotlight bar a few hundred pixels wide — most of
the row spent on toggles that are set once and rarely touched. In the HUD
they collapse into a single menu; the docked composer has the width and
keeps them inline, same controls and same state.

The trigger is not a static glyph. It reports the loudest live voice state —
recording, transcribing, listening for the wake word, speaking replies — and
lights while any is on, because a folded menu that looked idle with the mic
open would be a worse trade than the space it saves. The three toggles are
checkbox rows that hold the menu open on select, so the state you just
changed is the state you can see.

The shared control class names move to a module of their own so the row and
the menus it renders can wear them without importing each other, and the
pressed-toggle tint stops being written out at each of its four sites.
… above it

The exit chip floated over the composer in a 26px transparent strip reserved
for it (--hud-chip-strip), hidden until you hovered the bar. Under glass that
strip is bare untinted material across the top of the HUD — a band of chrome
above the surface, present in every state, holding a control you cannot see.

It rides the composer's controls row now, next to send. That costs no
reserved space and takes about 120 lines of CSS with it: the chip needed its
own placement, hover reveal, leave-hold, and an opaque card to stay legible
over an unknown desktop. None of that applies to a button on the bar, which
is already our surface — the problem was the placement, not the control.

Trade-off worth naming: the way out is now always visible in the HUD rather
than revealed on hover. It is one more permanent glyph on a Spotlight bar, in
exchange for an escape hatch that no longer depends on discovering it.
…privilege policy

The nima-infra apply_profile_toolset_policy blocks terminal/file/code_execution/
delegation/cronjob/kanban/session_search globally for domain profiles. That
cripples dispatcher-spawned kanban workers assigned to those profiles: the
worker session had no bash, no file tools, and no kanban board tools, so tasks
could not be read or worked ('essentially useless').

Kanban workers are sanctioned mutation contexts: the dispatcher scopes each
worker to its task workspace / branch / run claim. Re-enable the mutation
toolsets for the worker spawn (kanban_db._resolve_worker_cli_toolsets) and
exempt them from the session's disabled_toolsets when HERMES_KANBAN_TASK is set
(cli.py). The Telegram gateway allowlist still keeps the interactive surface
read-only.
…-fix cherry-pick

The kanban worker toolset fix cherry-picked onto pristine upstream
accidentally deleted the local 'from agent.skill_utils import
parse_config_string_list' (the fork's original had it imported at a
different location). Upstream's loader requires it at the call site;
without it cli.py raises NameError when parsing disabled_toolsets.

CI caught it: tests/cli/test_cli_secret_capture.py NameError at cli.py:5139.
Verified locally: failing test passes, kanban worker toolset tests 4/4.
@girnarholdings
girnarholdings merged commit fa95c40 into main Aug 20, 2026
57 of 60 checks passed
girnarholdings added a commit that referenced this pull request Aug 20, 2026
* test(cron): pin profile-adapter routing seams (regression guard for 2026-08-19)

The upstream catchup rebase (merge PR #14, 2026-08-19) silently dropped the
fork's profile-adapter routing fix (PR #13) — profile cron jobs delivered
through the root bot to the root DM for a day and every per-profile channel
went silent. No test pinned the three seams, so the drop was invisible to CI.

Add test_multiplex_ticker_passes_profile_specific_adapters:
- asserts start() and _start_multiplex() carry profile_adapters
- asserts _start_multiplex selects per-profile adapter dicts
- asserts the root fallback for adapter-less profiles
- behaviorally verifies each profile tick receives ITS OWN adapters

Verified: fails on fa95c40 (pre-fix), passes on 57bfb25 (fixed).

* ci: re-trigger checks

---------

Co-authored-by: nima <nima@girnarholdings.com>
@girnarholdings
girnarholdings deleted the sync/upstream-catchup-20260819 branch August 23, 2026 18:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.