Skip to content

Sync fork with upstream/main + consolidate fork work (and fix Anthropic OAuth refresh races) - #3

Merged
RFingAdam merged 714 commits into
mainfrom
consolidate/upstream-sync
Aug 27, 2026
Merged

RFingAdam merged 714 commits into
mainfrom
consolidate/upstream-sync

Conversation

@RFingAdam

@RFingAdam RFingAdam commented Aug 26, 2026 •

Copy link
Copy Markdown
Owner

Brings the fork up to date with upstream/main (692 commits) and consolidates every fork-unique change onto that base, so main stops drifting.

What's on this branch

Rebuilt from upstream/main with 10 commits replayed.

Fork-original work:

Carried forward because it is not in upstream/main: the auteur and discernment-nudge optional skills, and the Python repr redaction fixes.

New here: fix(auth) for the Anthropic OAuth refresh race, described below.

The auth fix

Anthropic OAuth refresh tokens are single-use and rotate on use. The anthropic provider was excluded from the cross-process auth-store lock and in-lock re-sync that openai-codex and xai-oauth already use. With several processes sharing one auth.json (gateway, dashboard, per-profile gateways), the first refresher rotates the token and every other process then POSTs a stale one, gets HTTP 404, and benches the credential. In practice Claude went offline about an hour after each login and traffic fell through the fallback chain.

This applied cleanly to upstream's base, so upstream still carries the bug.

Conflict resolutions worth a look

runGroupChatRounds: upstream rewrote it to be thread-aware, with stranded-reply harvesting and cancel recording. Kept upstream's rewrite and re-applied the fork's ceilings (maxRounds, maxMessages, deadline) on top rather than taking one side.

Room prompt rules: the fork's rules[] still held upstream's older wording. Folded upstream's current text in, so the toolCapable chat-only branch survives without reverting upstream's improvement.

Disband button: kept the fork's mode dropdown, took upstream's newer Tip-wrapped button over the fork's older title-based duplicate.

Test export list: unioned to 53 names, dropping GROUP_CHAT_MAX_ROUNDS and GROUP_CHAT_MAX_MESSAGES, which no longer exist.

Dropped as redundant

03a448e6a2 (transcript jumps) is already upstream as ead9d8e3d4, same author and timestamp, and upstream's version is a superset.

8397a186ed (Bot Mode invariants doc) came out empty. Upstream's AGENTS.md now documents a registry-based canonical-chat contract that supersedes the pin-based description.

Tests

  • group-chat.test.mjs: 103/103, including the fork-specific room-mode tests
  • full hermes-bots/tests/*.test.mjs: 580/580
  • auth suites (provider gate, xai and codex oauth, precedence): 44 passed

Not yet deployed.

chelsealong and others added 30 commits August 24, 2026 03:21
…Research#93518)

pty_ws already fell back to the per-channel active-session file when a
/chat WS connects with no ?resume= param, replaying the whole session
into the PTY, but the frontend only pinned xterm's viewport to the
bottom when resumeParam came from the URL (NousResearch#59591). The implicit path
had no way to learn a replay was happening, so the viewport stayed at
the top of the scrollback.

pty_ws now sends a one-off JSON control frame naming the session id it
resolved from the active-session file, before any PTY bytes; PTY
output itself always arrives as binary frames, so this is unambiguous
on the wire. ChatPage tracks an `effectiveResume` value seeded from
resumeParam and updated when this control frame arrives, and the
existing follow-scroll/sanitizer/hydration logic keys off it instead
of the URL param alone.

Fixes NousResearch#93518.
…latch the UI frozen

After a liveness-probe-triggered reconnect on a remote gateway,
attemptReconnect() awaits desktop.getConnection() and resolveGatewayWsUrl()
with no timeout. If either stalls (e.g. main process wedged mid-revalidation
even though the backend itself is reachable), the `reconnecting` guard never
clears, so every later scheduleReconnect()/attemptReconnect() early-returns
forever and the UI stays stuck in "reconnecting" until the app is restarted.

Bound both awaits with a 20s timeout so a stall rejects instead of hanging;
the existing catch/finally already clears the guard and resumes backoff on
rejection. gateway.connect() keeps its own separate connect timeout.

Fixes NousResearch#93454
…h#93454)

attemptReconnect() awaited desktop.revalidateConnection?.() unbounded,
immediately before the two IPC calls the previous commit wrapped in
withTimeout(). A wedged revalidation after a liveness-probe trip -
the exact trigger NousResearch#93454 and this file's own comment describe - hung
that await forever, so the reconnecting guard never cleared and the
prior fix never got reached.

Wrap it in the same 20s withTimeout() (still swallowing the result via
.catch, matching its existing best-effort semantics) and extend the
regression test to hang revalidateConnection() specifically, proving
getConnection() and the socket still proceed once the stall times out.
…waits too (NousResearch#93454)

Follow-up to the reconnect-loop fix: the same unbounded ticket-mint await
exists on the soft gateway-switch path and the initial boot() path. Bound
both with the same withTimeout/RECONNECT_ATTEMPT_TIMEOUT_MS so a wedged
IPC round-trip fails into the existing retry paths instead of hanging the
switch or the 'Starting Hermes…' screen forever.
…is deleted

botRosterMeta() calls botConnectionRoute() for every sourceScoped/remoteSource
row to look up its metadata. That's a passive display lookup, but
botConnectionRoute() throws whenever connectionId can't be resolved -- which
is exactly what a stale group-chat roster row looks like once its connection
is deleted (its persisted descriptor keeps remoteSource: true but loses
connectionId). Since botRosterMeta() is called for every member on every
group-chat render, opening a group that still references a deleted
connection threw on render and crashed the pane's error boundary in a loop
that survived app restarts (the poisoned row is in Local Storage).

botConnectionRoute()'s fail-closed throw is correct and stays for its actual
callers -- routing a real request to a bot (requestForBot, session
creation, etc., covered by remote-routing-races.test.mjs). botRosterMeta()
now catches that throw and treats the row as having no resolvable route,
same as a bot with no meta at all, instead of letting it blow up rendering.

Fixes NousResearch#93492
botConnectionRoute() stays the strict, throwing dispatch path for real
routing (requestForBot, session creation). botRosterMeta() is passive
display code and previously reached that throw through a bare catch,
which would have swallowed any unrelated failure the same way. It now
calls a new non-throwing resolveBotConnectionRoute() and branches on a
typed resolved | owner_removed | not_scoped status instead.

Adds witnesses for the split: the typed statuses themselves, that
strict dispatch still fails closed on an orphaned row, and that an
unrelated failure while resolving meta for a live route still
propagates instead of being swallowed.
Root cause of NousResearch#93492: deleting a cloud/remote connection disposed its
gateways (store/gateway.ts) but never touched the persisted 'group-chats'
storage, so every member descriptor referencing the deleted connection
stayed behind as a poisoned row (remoteSource: true, connection gone) that
render-path route lookups tripped over forever.

Subscribe to the connection registry's 'removed' lifecycle push
(window.hermesDesktop.connections.onChanged, feature-detected — older
Electron mains don't emit it) and annotate every persisted group-chat
member owned by the deleted connection. Rows are marked
(sourceMissing/sourceReachable), never silently deleted: the member keeps
its identity and panes render the existing degraded 'Gateway removed'
botSourceStatus state. Writes ride updateGroupChat so the durable record
keeps its full shape, and the listener unbinds on plugin dispose.
…rate

Rows poisoned before the removed-connection sweep existed (their
connection was deleted while an older Desktop ran, so no lifecycle push
ever swept them) are what made NousResearch#93492 survive app restarts. After the
persisted 'group-chats' hydrate, run a pure annotate pass over the rooms:

- a descriptor that lost its connectionId (route unresolvable — the exact
  shape that threw on render) is always marked;
- a descriptor whose connectionId is absent from the live connection
  registry is marked only when the registry could actually be read —
  an unavailable registry must not read as 'everything is orphaned'.

Marked rows keep their identity and degrade to the existing 'Gateway
removed' state; nothing is deleted.
Audit of the remaining unguarded botConnectionRoute() callers a pane
render can reach (NousResearch#93492 follow-up to the botRosterMeta split). Each now
uses the non-throwing resolveBotConnectionRoute() and degrades on an
owner_removed row instead of throwing into the pane's error boundary:

- botWorkspaceOwnerKey / setBotsWorkspaceOwner: sidebar visibility
  listener, Bots home open, and roster context menus recompute these on
  passive UI edges; an orphaned selection now yields the name-keyed owner
  and the blocked workspace target.
- durableGroupChatMembers: rebuilt on every group send over the whole
  seated roster; one orphaned member no longer aborts the room update,
  and a swept member's degraded mark now survives the rebuild.
- useModelOptions: hook body runs during render; the query is disabled
  for an orphaned row and the picker paints its error/disabled state.
- AdvancedProfileConfig: dialog falls back to the bot's own name scope.

Strict dispatch callers (requestForBot, session creation, deleteBot,
duplicateBot, ensureBotMetadata, routines) intentionally keep the
fail-closed throw — remote-routing-races.test.mjs still asserts it.

Adds orphaned-connection-members.test.mjs covering the removed-connection
sweep, the hydrate annotate (with/without a readable registry), the
degraded 'Gateway removed' rendering of swept rows, and every guarded
caller.
…pace pane

React.lazy(() => import('./syntax-diff')) only has its pending state
covered by Suspense. When the dynamic import rejects (e.g. a packaged
app whose renderer window resolves to the app.asar copy of dist/ while
the chunk exists only in app.asar.unpacked, NousResearch#93479), the rejection
throws past Suspense to the nearest error boundary, which is the whole
workspace ContribBoundary. One missing highlighter chunk then blanks
the entire chat transcript instead of just the diff falling back to
the plain colored DiffBody, the way markdown-text.tsx already isolates
this failure class for markdown.

Wraps LazySyntaxDiff in a local ErrorBoundary that renders DiffBody on
catch, so a failed highlight chunk degrades in place.
…derer index when packaged

The renderer index resolver tried APP_ROOT/dist/index.html — inside app.asar
when packaged — before the app.asar.unpacked copy that asarUnpack (dist/**)
ships and that resolveWebDist() already prefers for the embedded dashboard.
Loading the asar-internal index is how lazily imported chunks (syntax-diff-*,
shiki-*, mermaid-embed-*) end up fetched from a path that cannot serve them,
killing the workspace pane (NousResearch#93479).

Reorder the candidate ladder to prefer the unpacked web dist when packaged,
following the unpackedPathFor/resolveWebDist precedent. All window loaders
(main, overlay, quick) share resolveRendererIndex, so one reorder covers
every surface. Dev behavior is unchanged: outside an asar both candidates
collapse to APP_ROOT/dist and keep the original order.
missingRendererAssets only checked the module refs index.html itself names
(<script type=module> + modulepreload), so a torn install whose boot-critical
files were intact but whose lazy chunks were gone passed the generation check
and died minutes later on the first React.lazy() route with 'Failed to fetch
dynamically imported module' (NousResearch#93479: syntax-diff-*, shiki-*, mermaid-embed-*).

Walk the generation's module graph: for every present JS chunk, parse its
inline __vite__mapDeps filename table (the lazy-import manifest Vite bakes
into each chunk) and check those files too, transitively and cycle-safe.
resolveRendererIndex now skips a lazy-chunk-torn candidate in favor of the
intact copy instead of shipping a delayed crash.

Tests cover the mapDeps parser (definition table vs index-only call sites,
CDN refs), the exact NousResearch#93479 tear shape, transitive/cyclic walks, and the
torn-vs-intact preference end to end.
…time

Per-session rate limiting only counts consecutive strike windows, so a
pattern that recurs at a cadence just above WATCH_MIN_INTERVAL_SECONDS
(e.g. a service restarted repeatedly over a day) never trips the
existing strike-limit disable — each match lands in its own clean
cooldown window. Every one of those matches still forces a full-context
agent turn, which stalls the event loop on large sessions (NousResearch#93513).

Add WATCH_LIFETIME_MAX_HITS: once a session has delivered this many
watch_match notifications over its whole life, disable watch_patterns
and fall back to notify_on_complete, reusing the existing disable path.
…ounting, Nth-delivery promotion, docstring

Follow-ups on top of the cherry-picked NousResearch#93532 cap:
- Regression tests: suppressed (in-cooldown) matches must NOT consume the
  lifetime budget; the cap trips exactly at the Nth DELIVERED match and
  promotes to notify_on_complete with the watch_disabled summary queued
  right after the final match.
- Extract _emit_lifetime_watch_disabled() and emit the summary even when
  the global breaker drops the final match, so the user always learns why
  watching went quiet (parity with the strike-limit path).
- Mention the lifetime cap in the terminal tool docstring (the schema text
  was already updated by NousResearch#93532).

Refs NousResearch#93513
…ntial-pool key

A malformed OPENROUTER_API_KEY in ~/.hermes/.env (truncated paste, wrong
provider's key) passed has_usable_secret's length/placeholder check and was
returned by _resolve_api_key_provider_secret before the credential-pool
fallback was ever reached, producing opaque '401 Missing Authentication
header' errors even when a valid pool entry existed (NousResearch#93593).

- Add KNOWN_PROVIDER_KEY_PREFIXES (openrouter: sk-or-) and skip env values
  that mismatch a declared prefix, logging a WARNING naming the env var and
  expected prefix, then continuing to the next env var / pool fallback.
- Iterate credential-pool entries (peek first, then entries()) instead of
  only peek(), so one malformed pool entry doesn't block a valid one.
- Providers without a declared prefix are fail-open: unknown key formats
  are never rejected. Valid env keys still win over the pool (precedence
  unchanged).

Fixes NousResearch#93593
…NousResearch#93469)

The pricing snapshot could only express flat per-million rates, so
gemini-3.1-pro sessions with prompts over 200k tokens under-counted
input 2x ($2 vs $4/M) and output 1.5x ($12 vs $18/M).

- Add optional tier fields to PricingEntry: tier_threshold_tokens,
  input/output/cache_read_cost_per_million_above (None = flat, falls
  back to base rate per-field).
- estimate_usage_cost selects the above-threshold rates for the WHOLE
  request once usage.prompt_tokens (input + cache read + cache write)
  exceeds the threshold, matching Google's billing semantics.
- Populate gemini-3.1-pro (4.00/18.00/0.40 above 200k; alias
  gemini-3.1-pro-preview inherits) and gemini-2.5-pro (2.50/15.00
  above 200k).
- Flat entries are untouched: no threshold means no behavior change.

Reported and tier-field shape designed by @tornike14 (NousResearch#93469).

Tests: below/at threshold unchanged, above-threshold tiered whole-request
pricing, cache-read tier rate and base-rate fallback, preview alias,
flat entries unaffected.
…r connection per tick

The bot relay's drain loop RPCs every registered connection through
requestGatewayForAgent's per-request lease. With no other consumer
holding the route, the refcount hit 0 after every tick and the pooled
secondary was disposed — a fresh WebSocket dial + teardown per
connection every 4s, flooding the gateway logs with connect/disconnect
pairs (NousResearch#93594).

Two changes, both directions from the issue:

- Retained relay-route secondaries: retainGatewayForRelay pins a
  route's pooled socket with a counted retention (never clobbering the
  foreground 'retained' flag) for the relay's active lifetime, reusing
  the existing scheduleReconnect/full-jitter machinery on drops. The
  plugin pins each registered connection once via the new feature-
  detected host.retainProfileSocket door, reconciles pins with the
  current connection set on every drain, and releases everything in
  stopBotRelay/dispose. Local routes (null/'local') are exempt so the
  idle reaper can still reclaim spawned local backends. The live-work
  pruner also respects the pin.

- RELAY_DRAIN_INTERVAL_MS 4s -> 30s: the push path (NousResearch#93091,
  bot_relay.outbox.pending) carries envelope latency, so the poll is
  purely a backstop — 30s matches LIVE_SESSION_STATUS_BACKSTOP_INTERVAL_MS.

Tests: relay-push-drain updated to the new backstop semantics; new
gateway-relay-retention.test.ts proves one socket construction across
5 drain ticks (vs 3 constructions for 3 unretained ticks) and that
release/prune/local-exemption behave; new relay-socket-retention
plugin test pins the pin-once / release-on-departure / stop-releases
contracts.
Cron runs finish unwatched by design, so counting them in
$unreadSessionCount turned the titlebar badge into a permanently-lit
cron run counter (NousResearch#93552). The badge now counts regular + messaging
sessions only; cron unread state stays visible on the sidebar cron
section rows, and 'Mark all as read' (markAllSessionsRead +
ackAllSessionsRead, which iterates cron rows) still clears them.

Fixes NousResearch#93552
…ight splits

Every opened file registered its preview pane with dock dir 'right', so
each open split a new zone off the right edge — three file opens made
three ever-narrower columns (NousResearch#93610). The first preview still opens its
own zone docked beside main; every subsequent preview now anchors to an
existing preview-tile pane with dir 'center', so it stacks as a tab in
the same preview zone. Covers files, artifacts, and the Browser tab
alike (all flow through openPreview/$previewTabs); session tiles are
untouched.

Fixes NousResearch#93610
Use Hermes timezone-aware timestamps for retained events and turn messages. Pass the public timestamp field supported by hindsight-client 0.6.1 and cover the final serialized request field.
Adds an optional occurred_at (ISO-8601 date/datetime) parameter to the
hindsight_retain tool schema, threaded into the retain item's timestamp
field. When absent, the item timestamp defaults to the configured event
clock (base from PR NousResearch#82928 by @ragingbulld, authorship preserved) so the
Hindsight server can resolve relative time phrases; previously no item
timestamp was ever sent and temporal memories landed with null
occurred_start/occurred_end.

Fixes NousResearch#93568. Salvages NousResearch#82928.
…al no-ring reset hides it

The unlayered *:focus-visible reset in styles.css intentionally zeroes
--tw-ring-shadow ('No focus rings, anywhere'), so any control that relied
solely on focus-visible:ring-* had no visible keyboard focus state at all.
Mirror each control's hover treatment as a focus-visible background/text
affordance instead, keeping the global reset intact:

- ui/sidebar.tsx: group label, group action, menu button, menu action,
  menu sub-button get focus-visible:bg-sidebar-accent + accent foreground
- ui/tabs.tsx: TabsTrigger gets focus-visible:bg-background + text-foreground
- ui/text-tab.tsx: focus-visible:text-foreground (matches its hover)
- chat/composer/micro-actions.tsx: pill gets focus-visible chrome-action-hover
- right-sidebar/index.tsx HEADER_ACTION_CLASS: focus-visible sidebar-accent
- right-sidebar/terminal/rail.tsx RAIL_ACTION: focus-visible chrome-action-hover
- chat/sidebar/cron-jobs-section.tsx (row body + run rows): focus-visible
  chrome-action-hover
- chat/sidebar/session-row.tsx <time>: focus-visible:text-foreground

Sweep verified: remaining focus-visible:ring-* usages under apps/desktop/src
already pair with a border/bg/text companion (button/checkbox/switch/input,
starmap share-controls) or are covered by PR NousResearch#93460's row-hover work
(cron/index.tsx run rows).

Fixes NousResearch#93462. Reported by @fred0m.
… a tile

forkBranch ended by opening the branch as a session-tile and leaving the
primary selection on the parent (NousResearch#69750). In the default layout there is
no visible tile pane, so branching only added a sidebar row with no
feedback in the main area — and openSessionTile no-ops when the target
is already the selected session, the common case of branching the chat
you're viewing.

Load the branch as the primary session via resumeSession instead, which
reuses the runtime already warm-cached by forkBranch's
ensureSessionState/updateSessionState calls, so it doesn't cost an extra
resume RPC.

Fixes NousResearch#93444
…ssion

forkBranch was unconditionally routing every branched session into the
main pane via resumeSession, including sidebar/background branches of
a session the user isn't currently viewing. That reintroduces the
NousResearch#69750 focus-stealing bug for that path: branching a different session
from the sidebar yanked the active view away from whatever was open.
Only take over the main pane when the branch's parent is the session
already selected; otherwise keep opening it as its own tile.
…rewrite

NousResearch#93515 reports auto-speak reading each reply twice when the Edge TTS
streaming attempt falls back to the POST endpoint and the reply's
renderer id gets rewritten to its durable id mid-flight. That was true
before 63565fa, but resolveSpokenReply()'s ordinal-anchored dedupe
(landed 2026-08-19, five days before this issue was filed) already
follows the rewrite. No source change — this pins the behavior with a
regression test at the hook/store integration level, one layer above
the existing spoken-reply.ts unit tests.
…try-carrier

fix: /retry and /undo no longer replay an older message after compaction (NousResearch#81233, salvage NousResearch#81234)
…k-transform-bypass-93650-v2

fix: route codex payloads around the SDK's GIL-holding request transform (NousResearch#93650)
… survive the runtime-session reaper (NousResearch#93602)

A group member turn is a session-scoped RPC sequence (resume → attach →
prompt.submit → poll) issued with the runtime id its first RPC minted, but
requestForBot routes every RPC through its own request-scoped socket lease
(retained:false secondaries in store/gateway). Between two RPCs the refcount
hits 0, the leased socket closes, the gateway detaches the runtime session on
WS disconnect, the orphan reaper frees it after grace, and the next RPC —
prompt.submit, unwrapped — dies 4001 'not in memory'. The member turn aborts
and the sub-profile bot goes silent in the room.

- store/gateway: retainGatewayForAgent(connectionId, profile) — refcounted
  hold on the pooled socket with an idempotent release, mirroring the
  existing request-lease machinery.
- sdk: host.retainProfile(route) exposes the retain to plugins
  (feature-detected by consumers; older hosts keep working).
- hermes-bots plugin: runGroupChatMemberTurn acquires the lease before
  ensureGroupChatSession's first RPC and releases in finally, so the socket
  that minted the runtime id stays open across attach+submit+poll; and
  prompt.submit gets a one-shot catch-and-retry that re-resumes via the
  STORED session id on 4001-class failures (belt-and-braces for routes the
  lease can't cover). 4007 'never existed' keeps flowing to session.create.

Tests: simulated 4001 on first submit recovers via re-resume and delivers;
lease held across attach+submit (mock refcount never hits 0 mid-turn); lease
released after success AND failure; no-retainProfile host feature detection;
store-level retain/release + idempotent double-release + the unretained
disposal race.
teknium1 and others added 13 commits August 25, 2026 16:39
Two interaction seams between the NousResearch#92693 salvage (merged as NousResearch#95050) and
this branch: the source-label indexing test now compares in token space
(the stemmer shortens 'catalogsource' to 'catalogsourc'), and the
unregistered-core-name describe test forces the unregistered condition
via monkeypatch instead of depending on which sibling test file imported
model_tools first.
Follow-up to the salvaged NousResearch#94296: the two guards covered the repair and
confirmed-update branches, but when cua-driver is enabled yet not
installed at all, control still reached _run_cua_driver_installer() and
an automatic 'hermes update' would launch the interactive install.ps1
anyway. Add the same defer before the installer run, keep POSIX
behavior unchanged, and give the confirmed-update message a natural
fallback when latest_version is unknown.
…ixes

Port from OpenHands/software-agent-sdk#4508: their dict-entry secret
redaction was uppercase-only and leaked mixed-case keys (UserPassword,
sessionToken). Apply the same case-insensitive treatment to the Python
mapping-repr pass: a casefolded credential suffix (apikey/token/secret/
password/passwd/credential) now qualifies a key, while metadata names
(TOKEN_COUNT, password_policy, tokenizer) stay untouched.

(cherry picked from commit bbcaad4)
… system text

Rewrite bare Hermes tool-name mentions (skill_manage, session_search, etc.)
in system-prompt prose to the mcp__ wire-name form, and swap product-name
strings (Hermes Agent -> Claude Code, Nous Research -> Anthropic) so an
Anthropic OAuth subscription request is billed against the Claude Max plan
instead of being misclassified as third-party-agent traffic onto the empty
"extra usage" pool (observed live: bare snake_case tool fingerprints in
system text alone triggered HTTP 400 "You're out of extra usage" even
though the actual tool schemas were already using mcp__ names).

Local addition, not upstream -- billing-classifier behavior is specific to
this deployment's OAuth identity.

(cherry picked from commit 38147e5)
…iagnostic

Two local hardening changes for the kanban dispatcher, both motivated by
real incidents on this board:

1. Require a GitHub PR URL in task_comments before hermes_cli kanban
   complete / the mcp__kanban_complete tool will mark a project-linked
   (project_id set) task done. Guards the fabricated/undelivered-"done"
   pattern this board hit repeatedly (t_67182643: 2 fabricated
   rust-worker completions; t_0f867309: self-closed the same turn it
   admitted 2/3 planned steps were undone; 7+ other caught instances).
   Reviews, specs, and ops-triage cards (no project_id) are unaffected.

2. New kanban_diagnostics rule (_rule_stuck_in_ready_unassigned): surface
   a task sitting in 'ready' with an empty assignee past the stranded
   threshold as an actionable diagnostic. Previously the dispatcher only
   logged "Skipped (unassigned)" to console on every tick -- invisible
   and easy to leave rotting forever.

3. tools/kanban_tools.py's _handle_complete now distinguishes "unmet
   parent dependency" from a generic "could not complete" failure when
   complete_task() returns False. Root-caused via a real incident
   (t_f7fa01b0, see kanban board ops t_3bd77dee): a parent link added
   to a task AFTER it was already claimed/running silently vetoes
   completion even though kanban_show still reports status=running with
   a matching run_id the whole time -- previously surfaced as the opaque
   "unknown id or already terminal" message, which reads exactly like a
   dispatcher/DB bug and caused a worker to abandon real, verified work.
   The new message names the unmet parent(s), states plainly this is
   expected dependency gating (not corruption), and confirms run state
   is untouched so the worker knows it's safe to retry once the parent
   finishes.

Local additions, not upstream -- specific to this board's dispatch
patterns and incident history. Both changed files' relevant tests pass
(33/33 in test_kanban_tools.py, including the new regression test for
the late-parent-link case).

(cherry picked from commit 5ae0807)
Bot Mode's group-chat room loop (`apps/desktop/src/plugins/hermes-bots/plugin.js`) gets an opt-in **"Extended rounds"** setting:

- **Default (OFF): exact upstream stock behavior** — 3 rounds / 10 messages / no wall-clock cap. A fresh install of this fork is byte-identical to `NousResearch/hermes-agent` until a user explicitly opts in.
- **ON: raised ceilings + a new protection** — 24 rounds / 120 messages, plus a 30-minute wall-clock cap that stock never had.
- New UI toggle in the Bots pane header (stopwatch icon, next to the activity-toast bell).
- Both the toggle and an optional fine-tune override (`{ maxRounds, maxMessages, wallClockMinutes }`, hard-clamped, malformed/out-of-range input rejected per-field) persist via plugin storage.

The first commit on this branch just raised the shipped defaults globally. Adam's direction was better: keep Nous's own behavior as the untouched default, and make "run longer" an explicit, protected opt-in — so this fork never silently diverges from upstream for anyone who doesn't ask for it, and the extended path gets its own independent safety net rather than just bigger numbers.

- Flipping extended mode off always restores exact stock constants — no override can leak through.
- The wall-clock cap is new and independent of the existing "everyone passed this round" settle exit, which is conversation-shaped (only fires when literally every member passes/times out). Extended mode's much higher round ceiling widens the exposure window if a member ever got stuck producing real-looking, non-settling text every turn — the wall-clock deadline bounds worst-case cost/duration regardless of whether that heuristic ever fires.
- Override values are clamped to hard ranges (rounds 1-100, messages 1-500, wall-clock 1s-2h) and rejected (not silently clamped to a boundary) when out of range or malformed.

- Rewrote the fork's regression tests for the new design: default-is-stock, toggle raises/restores correctly, override clamps per-field in both directions, an end-to-end test that a tuned-down extended cap actually stops the loop, an end-to-end test that the wall-clock deadline ends a non-settling drive early via real `setTimeout`-based turn latency, and a test proving stock mode ignores a staged extended-mode override entirely.
- Fixed the test harness's `prompt.submit` mock to `await turnScript` (previously sync-only, silently no-op'd on an async turnScript — needed for the real-latency wall-clock test).
- Full `hermes-bots` plugin suite: **244/244 passing** (`node --test src/plugins/hermes-bots/tests/*.test.mjs`).
- Secret scan clean on both changed files.

Targets our fork's own working branch (`feat/kanban-completion-guards-and-oauth-sanitize`, which carries pre-existing local kanban/OAuth patches), not `NousResearch/hermes-agent` upstream `main` — this is a fork-local feature we're managing ourselves, not (yet) proposed upstream.

The gateway-side per-platform `require_mention` mention-gate (Discord/Telegram/etc adapters) that production group chats actually run on is a different mechanism (per-message gating, not a round-count ceiling) and isn't touched by this PR.

(cherry picked from commit 791fc66)
…esets

Replace the install-wide stock/extended ceiling toggle as the only governor
with optional per-room modes that override it when set. Unset rooms keep
exact prior global-toggle behavior (zero regression).

- rooms[name].mode: 'build' | 'decide' | 'standing' | unset
- Message budget scales as messagesPerMemberPerRound * members * rounds
  (fixes the PR #1 5-bot stock bug where a flat 10-msg cap died at round 2)
- decide: chat-only prompt rule + deterministic system summary on hard cap
- build: long ceilings, tool-capable (prompt does not restrict tools)
- standing: moderate multi-trigger coordination ceilings
- Room-header dropdown to set/clear mode (persists on the room record)

Kanban: t_ee2e6fe4
(cherry picked from commit e212071)
…m port)

Port of anthropics/skills discernment-nudge (Apache-2.0, added upstream
Aug 17 2026), snapshot 3b3fad96. After a substantive, actionable answer,
append 2-3 short targeted follow-up questions that help the user check
key facts, probe the reasoning, and notice missing context — at most
once per conversation, with explicit skip rules (trivial lookups,
formatting, code, creative writing). Reframed as an opt-in output habit,
not an identity change; upstream LICENSE.txt carried verbatim.

- optional-skills/productivity/discernment-nudge: SKILL.md + LICENSE.txt
- tests/skills/test_discernment_nudge_skill.py
- Docs: catalog row, sidebar entry, generated skill page (scoped regen).

(cherry picked from commit ca713cc)
…xecutable anti-slop gates

Port of agiwhitelist/auteur (MIT, ~1k stars), snapshot 9bca227d. Three
registers (build / direct / system) on one taste core: commit-sheet-first
art direction, asset generation via image_generate + local CLIs, and
node-based quality gates (slopscan anti-slop linter, motionqa frame-drop
check, systemscan cross-route drift) run through playwright.

- optional-skills/creative/auteur: SKILL.md (de-Clauded, Hermes tool
  framing), LICENSE (upstream MIT), 11 references, 8 verbatim upstream
  .mjs scripts (all pass node --check; slopscan smoke-run verified),
  6 templates. README gallery assets not vendored (size cap).
- tests/skills/test_auteur_skill.py: frontmatter, path-annotation
  invariant, de-Claude residue, related_skills resolution.
- Docs: catalog row, sidebar entry, generated skill page (scoped regen).

(cherry picked from commit b00b60f)
Anthropic refresh tokens are single-use and rotate on use, but the
anthropic provider was excluded from the cross-process auth-store lock
and in-lock re-sync already applied to openai-codex and xai-oauth. With
several Hermes processes sharing one auth.json (gateway + dashboard +
per-profile gateways), the first refresher rotated the token and every
other process POSTed a stale one, got HTTP 404, and benched the
credential - taking Claude offline about an hour after each login.

Route anthropic through the same locked sync-POST-write-back path and
widen the pool-store re-sync gate to cover it.
@RFingAdam
RFingAdam force-pushed the consolidate/upstream-sync branch from 6e7aed6 to 41a3df2 Compare August 26, 2026 00:44
@github-actions

github-actions Bot commented Aug 26, 2026 •

Copy link
Copy Markdown

૮ >ﻌ< ა ci review

ran on 31c2337 — ci: bump Python test suite timeout 30->60min (standard runne

❌ Job failures

JS & TS checks / JS & TS checks · View job

Job JS & TS checks / JS & TS checks failed.


Python tests / Run tests · View job

Job Python tests / Run tests failed.


⚠️ Action required

package-lock.json · View job

Locked npm dependency versions changed.

package-lock.json

Package Before After
@playwright/test 1.58.2 1.62.1
playwright 1.58.2 1.62.1
playwright-core 1.58.2 1.62.1

How to fix:

Add the ci-reviewed label after verifying the version changes are expected.


⚠️ Warnings

OSV vulnerability scan · View job

7 known vulnerabilities found in pinned dependencies.

How to fix:

Review the findings in the Security tab. Update the affected dependencies if a patched version is available.


ℹ️ Info

CI-sensitive file review · View job

PR touches sensitive files, but the ci-reviewed label has been added, approving them.

Sensitive files changed:


debug info

CI timings

CI timings · View report · View job

Wall time 46m (no baseline yet).

A member that takes on a task now keeps the floor across rounds until it
decides it is finished, instead of stopping after one message.

Protocol mirrors the existing (pass) contract: end a turn with (working)
to keep the floor, or (done)/(blocked) with a report to release it. A
member that says nothing special behaves exactly as before.

There is no turn ceiling on an open claim. The room's round, message and
wall-clock ceilings are bypassed for the member holding it, so work is
never cut off mid-task. Two exits remain:

  - repeating itself for GROUP_WORK_NO_PROGRESS_LIMIT turns releases the
    claim and posts a system note saying why
  - a hard backstop at GROUP_WORK_HARD_TURN_BACKSTOP rounds, matching
    delegation.max_iterations

Only an explicit (working) keeps the floor, so a member that forgets the
sentinel stops rather than holding the room to the backstop.

A member mid-task is continuing its own work rather than reacting to new
messages, so an empty delta no longer skips its turn; it gets a
continue-your-task prompt instead.

Claims live in the room record beside stranded/holds and are carried by
all three serializers, so a window restart resumes the loop. A held
member's claim is ignored until it is released.

On by default in every room, with a room-header toggle to opt out;
workLoop is persisted only when off, so existing rooms need no migration.
The loop does not grant tools: a chat-first room still loops for
multi-turn reasoning, and running tools stays an explicit build-mode
choice.
…merge

The bots pane failed to render with "$botSessionsWorkspace is not
defined". Three references survived the upstream consolidation while the
code that defined them did not:

  - $botSessionsWorkspace: upstream removed the sessions-workspace pane
    outright. The surviving read had no remaining consumer, so it goes
    with the feature.
  - showHidden: upstream renamed it hiddenExpanded. Six reads in the
    roster eye-toggle still used the old name. showHiddenSection and
    showHiddenRows are separate values and are untouched.
  - hiddenUnread: only ever defined in the fork. Restored next to
    hiddenBots, keyed by botSelectionKey to match how unread is stored
    now rather than the fork's bot.name.

The test suite passed throughout: it exercises the group-chat engine but
never renders the roster or pane components, so a dangling identifier in
JSX is invisible to it. eslint no-undef does catch it, and now reports no
undefined project identifiers in this file.
The work loop deliberately has no turn ceiling, which leaves spend
unbounded. This adds the one guard that was missing.

The desktop has no usage or billing RPC, so the figure is an ESTIMATE
from prompt and reply characters (~4 chars per token), not a billed
count. It is a smoke alarm, not an invoice: it stops a drive that is
burning far more than expected and says plainly that the number is an
estimate.

Unlike the round and wall-clock ceilings, this one DOES apply to a member
holding a claim. A claim means "let me finish", not "spend without
limit". On trip, open claims are released so nothing silently resumes and
a system note records the spend and the ceiling.

Budget resolves as: explicit per-room tokenBudget, then the mode preset
(build 400k, standing 120k, decide 60k), then 200k. Zero or negative
disables it.
An unanswered approval returned "The user has NOT consented to this
action", which is false: nobody refused it, the request went unanswered.
A well-behaved agent reads that as a decision, abandons the work
permanently, and reports it to the user as rejected. Observed in a live
run where an authorised `gh pr create` timed out under CPU starvation:
the agent stopped, declined to retry via any route, and reported a
consent denial that never happened.

The fail-closed behaviour is correct and unchanged - the command does not
run, and the agent still must not retry or route around it without fresh
approval. Only the claim about what the human did is corrected, plus an
explicit instruction to report it as blocked-awaiting-approval rather
than as rejected work.

Tests assert the new wording and now guard against the regression: a
timeout message must never contain "has NOT consented".
…r lane

Coordination in a room was purely conversational: a bot saying "assigning
t_123 to @backend" was a statement of intent that nothing enforced, so two
members could each decide the same task was theirs and open parallel work
on it.

A thread with an assignee is now that member's lane. Selection filters to
the assignee regardless of who was @mentioned, so a non-assignee is never
dispatched into someone else's work. Clearing the assignee reopens the
thread.

Assignments live beside working/holds in the room record and are carried
by all three serializers, so a lane survives a reload.
`sessions.model` is the CONFIGURED model. After a provider fallback it is
no longer the model answering, so a client showing it reports a
reassuring lie - a room can run an entire session on a fallback provider
with nothing on screen saying so. Observed live: a bot ran a whole
session on gemini while its config, and the UI, said claude-sonnet-5.

Adds SessionDB.get_session_usage_summary(), reading session_model_usage
(the per-call record) for the most recent route, every distinct route
seen, and token/cost totals. Deliberately most-recent rather than
dominant: a late fallback is exactly the case worth surfacing, even when
it served the fewest calls.

The shared live-session payload carries it as `usage`, so session.resume
exposes it without a new RPC. The room records it per member and renders
the serving model beside each turn, amber when the session has used more
than one route.

The room's spend ceiling now prefers these real totals, measured as the
delta from each member's total at its first turn of the drive, and falls
back to the character estimate only when the backend reports nothing -
the system note says which figure it used.
…s you

Six bots produce a wall of text with no answer to the only question that
matters: is any of it waiting on me. Status is now derived from the same
turn intent the drive already parses, rather than a second protocol the
model has to remember - the (working) sentinel was used once in an entire
six-bot run, so anything needing the model to emit extra ceremony does
not survive contact.

(done) and (blocked) are now distinct intents. Merging them reported work
a bot could not finish as finished, which is the opposite of useful. Each
member carries working / review / blocked / idle, persisted like the rest
of the room record, and the header summarises "N to review, M blocked, K
working" - amber for blocked, accent for review, and nothing at all when
the room is quiet.

The turn prompt now states the distinction explicitly so (done) is not
used for work that stalled.
…GH account)

ubuntu-latest-32-core / -96-core and windows-latest-32-core are paid
GitHub Team/Enterprise runner tiers not provisioned on this personal
account (confirmed: gh api users/RFingAdam/actions/hosted-runners -> 404).
Every job pinned to one of these sat queued indefinitely with zero
runner ever assigned -- reproduced twice across separate CI runs, not
a transient scheduling delay. Downgraded to standard ubuntu-latest/
windows-latest across all 4 stalled workflows (tests.yml, tests-os.yml,
js-tests.yml, nix.yml) plus 3 more not yet exercised by this PR but
carrying the same landmine (e2e-desktop.yml, docker.yml, rust-tests.yml).
… less parallelism than the 96-core it was tuned for)
@RFingAdam
RFingAdam merged commit a9a0c61 into main Aug 27, 2026
57 of 63 checks passed
@RFingAdam
RFingAdam deleted the consolidate/upstream-sync branch August 27, 2026 03:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.