Repository navigation
Sync fork with upstream/main + consolidate fork work (and fix Anthropic OAuth refresh races) - #3
Merged
Merged
Conversation
…Research#93518) pty_ws already fell back to the per-channel active-session file when a /chat WS connects with no ?resume= param, replaying the whole session into the PTY, but the frontend only pinned xterm's viewport to the bottom when resumeParam came from the URL (NousResearch#59591). The implicit path had no way to learn a replay was happening, so the viewport stayed at the top of the scrollback. pty_ws now sends a one-off JSON control frame naming the session id it resolved from the active-session file, before any PTY bytes; PTY output itself always arrives as binary frames, so this is unambiguous on the wire. ChatPage tracks an `effectiveResume` value seeded from resumeParam and updated when this control frame arrives, and the existing follow-scroll/sanitizer/hydration logic keys off it instead of the URL param alone. Fixes NousResearch#93518.
…latch the UI frozen After a liveness-probe-triggered reconnect on a remote gateway, attemptReconnect() awaits desktop.getConnection() and resolveGatewayWsUrl() with no timeout. If either stalls (e.g. main process wedged mid-revalidation even though the backend itself is reachable), the `reconnecting` guard never clears, so every later scheduleReconnect()/attemptReconnect() early-returns forever and the UI stays stuck in "reconnecting" until the app is restarted. Bound both awaits with a 20s timeout so a stall rejects instead of hanging; the existing catch/finally already clears the guard and resumes backoff on rejection. gateway.connect() keeps its own separate connect timeout. Fixes NousResearch#93454
…h#93454) attemptReconnect() awaited desktop.revalidateConnection?.() unbounded, immediately before the two IPC calls the previous commit wrapped in withTimeout(). A wedged revalidation after a liveness-probe trip - the exact trigger NousResearch#93454 and this file's own comment describe - hung that await forever, so the reconnecting guard never cleared and the prior fix never got reached. Wrap it in the same 20s withTimeout() (still swallowing the result via .catch, matching its existing best-effort semantics) and extend the regression test to hang revalidateConnection() specifically, proving getConnection() and the socket still proceed once the stall times out.
…waits too (NousResearch#93454) Follow-up to the reconnect-loop fix: the same unbounded ticket-mint await exists on the soft gateway-switch path and the initial boot() path. Bound both with the same withTimeout/RECONNECT_ATTEMPT_TIMEOUT_MS so a wedged IPC round-trip fails into the existing retry paths instead of hanging the switch or the 'Starting Hermes…' screen forever.
…is deleted botRosterMeta() calls botConnectionRoute() for every sourceScoped/remoteSource row to look up its metadata. That's a passive display lookup, but botConnectionRoute() throws whenever connectionId can't be resolved -- which is exactly what a stale group-chat roster row looks like once its connection is deleted (its persisted descriptor keeps remoteSource: true but loses connectionId). Since botRosterMeta() is called for every member on every group-chat render, opening a group that still references a deleted connection threw on render and crashed the pane's error boundary in a loop that survived app restarts (the poisoned row is in Local Storage). botConnectionRoute()'s fail-closed throw is correct and stays for its actual callers -- routing a real request to a bot (requestForBot, session creation, etc., covered by remote-routing-races.test.mjs). botRosterMeta() now catches that throw and treats the row as having no resolvable route, same as a bot with no meta at all, instead of letting it blow up rendering. Fixes NousResearch#93492
botConnectionRoute() stays the strict, throwing dispatch path for real routing (requestForBot, session creation). botRosterMeta() is passive display code and previously reached that throw through a bare catch, which would have swallowed any unrelated failure the same way. It now calls a new non-throwing resolveBotConnectionRoute() and branches on a typed resolved | owner_removed | not_scoped status instead. Adds witnesses for the split: the typed statuses themselves, that strict dispatch still fails closed on an orphaned row, and that an unrelated failure while resolving meta for a live route still propagates instead of being swallowed.
Root cause of NousResearch#93492: deleting a cloud/remote connection disposed its gateways (store/gateway.ts) but never touched the persisted 'group-chats' storage, so every member descriptor referencing the deleted connection stayed behind as a poisoned row (remoteSource: true, connection gone) that render-path route lookups tripped over forever. Subscribe to the connection registry's 'removed' lifecycle push (window.hermesDesktop.connections.onChanged, feature-detected — older Electron mains don't emit it) and annotate every persisted group-chat member owned by the deleted connection. Rows are marked (sourceMissing/sourceReachable), never silently deleted: the member keeps its identity and panes render the existing degraded 'Gateway removed' botSourceStatus state. Writes ride updateGroupChat so the durable record keeps its full shape, and the listener unbinds on plugin dispose.
…rate Rows poisoned before the removed-connection sweep existed (their connection was deleted while an older Desktop ran, so no lifecycle push ever swept them) are what made NousResearch#93492 survive app restarts. After the persisted 'group-chats' hydrate, run a pure annotate pass over the rooms: - a descriptor that lost its connectionId (route unresolvable — the exact shape that threw on render) is always marked; - a descriptor whose connectionId is absent from the live connection registry is marked only when the registry could actually be read — an unavailable registry must not read as 'everything is orphaned'. Marked rows keep their identity and degrade to the existing 'Gateway removed' state; nothing is deleted.
Audit of the remaining unguarded botConnectionRoute() callers a pane render can reach (NousResearch#93492 follow-up to the botRosterMeta split). Each now uses the non-throwing resolveBotConnectionRoute() and degrades on an owner_removed row instead of throwing into the pane's error boundary: - botWorkspaceOwnerKey / setBotsWorkspaceOwner: sidebar visibility listener, Bots home open, and roster context menus recompute these on passive UI edges; an orphaned selection now yields the name-keyed owner and the blocked workspace target. - durableGroupChatMembers: rebuilt on every group send over the whole seated roster; one orphaned member no longer aborts the room update, and a swept member's degraded mark now survives the rebuild. - useModelOptions: hook body runs during render; the query is disabled for an orphaned row and the picker paints its error/disabled state. - AdvancedProfileConfig: dialog falls back to the bot's own name scope. Strict dispatch callers (requestForBot, session creation, deleteBot, duplicateBot, ensureBotMetadata, routines) intentionally keep the fail-closed throw — remote-routing-races.test.mjs still asserts it. Adds orphaned-connection-members.test.mjs covering the removed-connection sweep, the hydrate annotate (with/without a readable registry), the degraded 'Gateway removed' rendering of swept rows, and every guarded caller.
…pace pane
React.lazy(() => import('./syntax-diff')) only has its pending state
covered by Suspense. When the dynamic import rejects (e.g. a packaged
app whose renderer window resolves to the app.asar copy of dist/ while
the chunk exists only in app.asar.unpacked, NousResearch#93479), the rejection
throws past Suspense to the nearest error boundary, which is the whole
workspace ContribBoundary. One missing highlighter chunk then blanks
the entire chat transcript instead of just the diff falling back to
the plain colored DiffBody, the way markdown-text.tsx already isolates
this failure class for markdown.
Wraps LazySyntaxDiff in a local ErrorBoundary that renders DiffBody on
catch, so a failed highlight chunk degrades in place.
…derer index when packaged The renderer index resolver tried APP_ROOT/dist/index.html — inside app.asar when packaged — before the app.asar.unpacked copy that asarUnpack (dist/**) ships and that resolveWebDist() already prefers for the embedded dashboard. Loading the asar-internal index is how lazily imported chunks (syntax-diff-*, shiki-*, mermaid-embed-*) end up fetched from a path that cannot serve them, killing the workspace pane (NousResearch#93479). Reorder the candidate ladder to prefer the unpacked web dist when packaged, following the unpackedPathFor/resolveWebDist precedent. All window loaders (main, overlay, quick) share resolveRendererIndex, so one reorder covers every surface. Dev behavior is unchanged: outside an asar both candidates collapse to APP_ROOT/dist and keep the original order.
missingRendererAssets only checked the module refs index.html itself names (<script type=module> + modulepreload), so a torn install whose boot-critical files were intact but whose lazy chunks were gone passed the generation check and died minutes later on the first React.lazy() route with 'Failed to fetch dynamically imported module' (NousResearch#93479: syntax-diff-*, shiki-*, mermaid-embed-*). Walk the generation's module graph: for every present JS chunk, parse its inline __vite__mapDeps filename table (the lazy-import manifest Vite bakes into each chunk) and check those files too, transitively and cycle-safe. resolveRendererIndex now skips a lazy-chunk-torn candidate in favor of the intact copy instead of shipping a delayed crash. Tests cover the mapDeps parser (definition table vs index-only call sites, CDN refs), the exact NousResearch#93479 tear shape, transitive/cyclic walks, and the torn-vs-intact preference end to end.
…time Per-session rate limiting only counts consecutive strike windows, so a pattern that recurs at a cadence just above WATCH_MIN_INTERVAL_SECONDS (e.g. a service restarted repeatedly over a day) never trips the existing strike-limit disable — each match lands in its own clean cooldown window. Every one of those matches still forces a full-context agent turn, which stalls the event loop on large sessions (NousResearch#93513). Add WATCH_LIFETIME_MAX_HITS: once a session has delivered this many watch_match notifications over its whole life, disable watch_patterns and fall back to notify_on_complete, reusing the existing disable path.
…ounting, Nth-delivery promotion, docstring Follow-ups on top of the cherry-picked NousResearch#93532 cap: - Regression tests: suppressed (in-cooldown) matches must NOT consume the lifetime budget; the cap trips exactly at the Nth DELIVERED match and promotes to notify_on_complete with the watch_disabled summary queued right after the final match. - Extract _emit_lifetime_watch_disabled() and emit the summary even when the global breaker drops the final match, so the user always learns why watching went quiet (parity with the strike-limit path). - Mention the lifetime cap in the terminal tool docstring (the schema text was already updated by NousResearch#93532). Refs NousResearch#93513
…ntial-pool key A malformed OPENROUTER_API_KEY in ~/.hermes/.env (truncated paste, wrong provider's key) passed has_usable_secret's length/placeholder check and was returned by _resolve_api_key_provider_secret before the credential-pool fallback was ever reached, producing opaque '401 Missing Authentication header' errors even when a valid pool entry existed (NousResearch#93593). - Add KNOWN_PROVIDER_KEY_PREFIXES (openrouter: sk-or-) and skip env values that mismatch a declared prefix, logging a WARNING naming the env var and expected prefix, then continuing to the next env var / pool fallback. - Iterate credential-pool entries (peek first, then entries()) instead of only peek(), so one malformed pool entry doesn't block a valid one. - Providers without a declared prefix are fail-open: unknown key formats are never rejected. Valid env keys still win over the pool (precedence unchanged). Fixes NousResearch#93593
…NousResearch#93469) The pricing snapshot could only express flat per-million rates, so gemini-3.1-pro sessions with prompts over 200k tokens under-counted input 2x ($2 vs $4/M) and output 1.5x ($12 vs $18/M). - Add optional tier fields to PricingEntry: tier_threshold_tokens, input/output/cache_read_cost_per_million_above (None = flat, falls back to base rate per-field). - estimate_usage_cost selects the above-threshold rates for the WHOLE request once usage.prompt_tokens (input + cache read + cache write) exceeds the threshold, matching Google's billing semantics. - Populate gemini-3.1-pro (4.00/18.00/0.40 above 200k; alias gemini-3.1-pro-preview inherits) and gemini-2.5-pro (2.50/15.00 above 200k). - Flat entries are untouched: no threshold means no behavior change. Reported and tier-field shape designed by @tornike14 (NousResearch#93469). Tests: below/at threshold unchanged, above-threshold tiered whole-request pricing, cache-read tier rate and base-rate fallback, preview alias, flat entries unaffected.
…r connection per tick The bot relay's drain loop RPCs every registered connection through requestGatewayForAgent's per-request lease. With no other consumer holding the route, the refcount hit 0 after every tick and the pooled secondary was disposed — a fresh WebSocket dial + teardown per connection every 4s, flooding the gateway logs with connect/disconnect pairs (NousResearch#93594). Two changes, both directions from the issue: - Retained relay-route secondaries: retainGatewayForRelay pins a route's pooled socket with a counted retention (never clobbering the foreground 'retained' flag) for the relay's active lifetime, reusing the existing scheduleReconnect/full-jitter machinery on drops. The plugin pins each registered connection once via the new feature- detected host.retainProfileSocket door, reconciles pins with the current connection set on every drain, and releases everything in stopBotRelay/dispose. Local routes (null/'local') are exempt so the idle reaper can still reclaim spawned local backends. The live-work pruner also respects the pin. - RELAY_DRAIN_INTERVAL_MS 4s -> 30s: the push path (NousResearch#93091, bot_relay.outbox.pending) carries envelope latency, so the poll is purely a backstop — 30s matches LIVE_SESSION_STATUS_BACKSTOP_INTERVAL_MS. Tests: relay-push-drain updated to the new backstop semantics; new gateway-relay-retention.test.ts proves one socket construction across 5 drain ticks (vs 3 constructions for 3 unretained ticks) and that release/prune/local-exemption behave; new relay-socket-retention plugin test pins the pin-once / release-on-departure / stop-releases contracts.
Cron runs finish unwatched by design, so counting them in $unreadSessionCount turned the titlebar badge into a permanently-lit cron run counter (NousResearch#93552). The badge now counts regular + messaging sessions only; cron unread state stays visible on the sidebar cron section rows, and 'Mark all as read' (markAllSessionsRead + ackAllSessionsRead, which iterates cron rows) still clears them. Fixes NousResearch#93552
…ight splits Every opened file registered its preview pane with dock dir 'right', so each open split a new zone off the right edge — three file opens made three ever-narrower columns (NousResearch#93610). The first preview still opens its own zone docked beside main; every subsequent preview now anchors to an existing preview-tile pane with dir 'center', so it stacks as a tab in the same preview zone. Covers files, artifacts, and the Browser tab alike (all flow through openPreview/$previewTabs); session tiles are untouched. Fixes NousResearch#93610
Use Hermes timezone-aware timestamps for retained events and turn messages. Pass the public timestamp field supported by hindsight-client 0.6.1 and cover the final serialized request field.
Adds an optional occurred_at (ISO-8601 date/datetime) parameter to the hindsight_retain tool schema, threaded into the retain item's timestamp field. When absent, the item timestamp defaults to the configured event clock (base from PR NousResearch#82928 by @ragingbulld, authorship preserved) so the Hindsight server can resolve relative time phrases; previously no item timestamp was ever sent and temporal memories landed with null occurred_start/occurred_end. Fixes NousResearch#93568. Salvages NousResearch#82928.
…al no-ring reset hides it
The unlayered *:focus-visible reset in styles.css intentionally zeroes
--tw-ring-shadow ('No focus rings, anywhere'), so any control that relied
solely on focus-visible:ring-* had no visible keyboard focus state at all.
Mirror each control's hover treatment as a focus-visible background/text
affordance instead, keeping the global reset intact:
- ui/sidebar.tsx: group label, group action, menu button, menu action,
menu sub-button get focus-visible:bg-sidebar-accent + accent foreground
- ui/tabs.tsx: TabsTrigger gets focus-visible:bg-background + text-foreground
- ui/text-tab.tsx: focus-visible:text-foreground (matches its hover)
- chat/composer/micro-actions.tsx: pill gets focus-visible chrome-action-hover
- right-sidebar/index.tsx HEADER_ACTION_CLASS: focus-visible sidebar-accent
- right-sidebar/terminal/rail.tsx RAIL_ACTION: focus-visible chrome-action-hover
- chat/sidebar/cron-jobs-section.tsx (row body + run rows): focus-visible
chrome-action-hover
- chat/sidebar/session-row.tsx <time>: focus-visible:text-foreground
Sweep verified: remaining focus-visible:ring-* usages under apps/desktop/src
already pair with a border/bg/text companion (button/checkbox/switch/input,
starmap share-controls) or are covered by PR NousResearch#93460's row-hover work
(cron/index.tsx run rows).
Fixes NousResearch#93462. Reported by @fred0m.
… a tile forkBranch ended by opening the branch as a session-tile and leaving the primary selection on the parent (NousResearch#69750). In the default layout there is no visible tile pane, so branching only added a sidebar row with no feedback in the main area — and openSessionTile no-ops when the target is already the selected session, the common case of branching the chat you're viewing. Load the branch as the primary session via resumeSession instead, which reuses the runtime already warm-cached by forkBranch's ensureSessionState/updateSessionState calls, so it doesn't cost an extra resume RPC. Fixes NousResearch#93444
…ssion forkBranch was unconditionally routing every branched session into the main pane via resumeSession, including sidebar/background branches of a session the user isn't currently viewing. That reintroduces the NousResearch#69750 focus-stealing bug for that path: branching a different session from the sidebar yanked the active view away from whatever was open. Only take over the main pane when the branch's parent is the session already selected; otherwise keep opening it as its own tile.
…rewrite NousResearch#93515 reports auto-speak reading each reply twice when the Edge TTS streaming attempt falls back to the POST endpoint and the reply's renderer id gets rewritten to its durable id mid-flight. That was true before 63565fa, but resolveSpokenReply()'s ordinal-anchored dedupe (landed 2026-08-19, five days before this issue was filed) already follows the rewrite. No source change — this pins the behavior with a regression test at the hook/store integration level, one layer above the existing spoken-reply.ts unit tests.
…try-carrier fix: /retry and /undo no longer replay an older message after compaction (NousResearch#81233, salvage NousResearch#81234)
…k-transform-bypass-93650-v2 fix: route codex payloads around the SDK's GIL-holding request transform (NousResearch#93650)
… survive the runtime-session reaper (NousResearch#93602) A group member turn is a session-scoped RPC sequence (resume → attach → prompt.submit → poll) issued with the runtime id its first RPC minted, but requestForBot routes every RPC through its own request-scoped socket lease (retained:false secondaries in store/gateway). Between two RPCs the refcount hits 0, the leased socket closes, the gateway detaches the runtime session on WS disconnect, the orphan reaper frees it after grace, and the next RPC — prompt.submit, unwrapped — dies 4001 'not in memory'. The member turn aborts and the sub-profile bot goes silent in the room. - store/gateway: retainGatewayForAgent(connectionId, profile) — refcounted hold on the pooled socket with an idempotent release, mirroring the existing request-lease machinery. - sdk: host.retainProfile(route) exposes the retain to plugins (feature-detected by consumers; older hosts keep working). - hermes-bots plugin: runGroupChatMemberTurn acquires the lease before ensureGroupChatSession's first RPC and releases in finally, so the socket that minted the runtime id stays open across attach+submit+poll; and prompt.submit gets a one-shot catch-and-retry that re-resumes via the STORED session id on 4001-class failures (belt-and-braces for routes the lease can't cover). 4007 'never existed' keeps flowing to session.create. Tests: simulated 4001 on first submit recovers via re-resume and delivers; lease held across attach+submit (mock refcount never hits 0 mid-turn); lease released after success AND failure; no-retainProfile host feature detection; store-level retain/release + idempotent double-release + the unretained disposal race.
Two interaction seams between the NousResearch#92693 salvage (merged as NousResearch#95050) and this branch: the source-label indexing test now compares in token space (the stemmer shortens 'catalogsource' to 'catalogsourc'), and the unregistered-core-name describe test forces the unregistered condition via monkeypatch instead of depending on which sibling test file imported model_tools first.
Follow-up to the salvaged NousResearch#94296: the two guards covered the repair and confirmed-update branches, but when cua-driver is enabled yet not installed at all, control still reached _run_cua_driver_installer() and an automatic 'hermes update' would launch the interactive install.ps1 anyway. Add the same defer before the installer run, keep POSIX behavior unchanged, and give the confirmed-update message a natural fallback when latest_version is unknown.
(cherry picked from commit b1c03c8)
…ixes Port from OpenHands/software-agent-sdk#4508: their dict-entry secret redaction was uppercase-only and leaked mixed-case keys (UserPassword, sessionToken). Apply the same case-insensitive treatment to the Python mapping-repr pass: a casefolded credential suffix (apikey/token/secret/ password/passwd/credential) now qualifies a key, while metadata names (TOKEN_COUNT, password_policy, tokenizer) stay untouched. (cherry picked from commit bbcaad4)
(cherry picked from commit 2af31a0)
… system text Rewrite bare Hermes tool-name mentions (skill_manage, session_search, etc.) in system-prompt prose to the mcp__ wire-name form, and swap product-name strings (Hermes Agent -> Claude Code, Nous Research -> Anthropic) so an Anthropic OAuth subscription request is billed against the Claude Max plan instead of being misclassified as third-party-agent traffic onto the empty "extra usage" pool (observed live: bare snake_case tool fingerprints in system text alone triggered HTTP 400 "You're out of extra usage" even though the actual tool schemas were already using mcp__ names). Local addition, not upstream -- billing-classifier behavior is specific to this deployment's OAuth identity. (cherry picked from commit 38147e5)
…iagnostic Two local hardening changes for the kanban dispatcher, both motivated by real incidents on this board: 1. Require a GitHub PR URL in task_comments before hermes_cli kanban complete / the mcp__kanban_complete tool will mark a project-linked (project_id set) task done. Guards the fabricated/undelivered-"done" pattern this board hit repeatedly (t_67182643: 2 fabricated rust-worker completions; t_0f867309: self-closed the same turn it admitted 2/3 planned steps were undone; 7+ other caught instances). Reviews, specs, and ops-triage cards (no project_id) are unaffected. 2. New kanban_diagnostics rule (_rule_stuck_in_ready_unassigned): surface a task sitting in 'ready' with an empty assignee past the stranded threshold as an actionable diagnostic. Previously the dispatcher only logged "Skipped (unassigned)" to console on every tick -- invisible and easy to leave rotting forever. 3. tools/kanban_tools.py's _handle_complete now distinguishes "unmet parent dependency" from a generic "could not complete" failure when complete_task() returns False. Root-caused via a real incident (t_f7fa01b0, see kanban board ops t_3bd77dee): a parent link added to a task AFTER it was already claimed/running silently vetoes completion even though kanban_show still reports status=running with a matching run_id the whole time -- previously surfaced as the opaque "unknown id or already terminal" message, which reads exactly like a dispatcher/DB bug and caused a worker to abandon real, verified work. The new message names the unmet parent(s), states plainly this is expected dependency gating (not corruption), and confirms run state is untouched so the worker knows it's safe to retry once the parent finishes. Local additions, not upstream -- specific to this board's dispatch patterns and incident history. Both changed files' relevant tests pass (33/33 in test_kanban_tools.py, including the new regression test for the late-parent-link case). (cherry picked from commit 5ae0807)
Bot Mode's group-chat room loop (`apps/desktop/src/plugins/hermes-bots/plugin.js`) gets an opt-in **"Extended rounds"** setting:
- **Default (OFF): exact upstream stock behavior** — 3 rounds / 10 messages / no wall-clock cap. A fresh install of this fork is byte-identical to `NousResearch/hermes-agent` until a user explicitly opts in.
- **ON: raised ceilings + a new protection** — 24 rounds / 120 messages, plus a 30-minute wall-clock cap that stock never had.
- New UI toggle in the Bots pane header (stopwatch icon, next to the activity-toast bell).
- Both the toggle and an optional fine-tune override (`{ maxRounds, maxMessages, wallClockMinutes }`, hard-clamped, malformed/out-of-range input rejected per-field) persist via plugin storage.
The first commit on this branch just raised the shipped defaults globally. Adam's direction was better: keep Nous's own behavior as the untouched default, and make "run longer" an explicit, protected opt-in — so this fork never silently diverges from upstream for anyone who doesn't ask for it, and the extended path gets its own independent safety net rather than just bigger numbers.
- Flipping extended mode off always restores exact stock constants — no override can leak through.
- The wall-clock cap is new and independent of the existing "everyone passed this round" settle exit, which is conversation-shaped (only fires when literally every member passes/times out). Extended mode's much higher round ceiling widens the exposure window if a member ever got stuck producing real-looking, non-settling text every turn — the wall-clock deadline bounds worst-case cost/duration regardless of whether that heuristic ever fires.
- Override values are clamped to hard ranges (rounds 1-100, messages 1-500, wall-clock 1s-2h) and rejected (not silently clamped to a boundary) when out of range or malformed.
- Rewrote the fork's regression tests for the new design: default-is-stock, toggle raises/restores correctly, override clamps per-field in both directions, an end-to-end test that a tuned-down extended cap actually stops the loop, an end-to-end test that the wall-clock deadline ends a non-settling drive early via real `setTimeout`-based turn latency, and a test proving stock mode ignores a staged extended-mode override entirely.
- Fixed the test harness's `prompt.submit` mock to `await turnScript` (previously sync-only, silently no-op'd on an async turnScript — needed for the real-latency wall-clock test).
- Full `hermes-bots` plugin suite: **244/244 passing** (`node --test src/plugins/hermes-bots/tests/*.test.mjs`).
- Secret scan clean on both changed files.
Targets our fork's own working branch (`feat/kanban-completion-guards-and-oauth-sanitize`, which carries pre-existing local kanban/OAuth patches), not `NousResearch/hermes-agent` upstream `main` — this is a fork-local feature we're managing ourselves, not (yet) proposed upstream.
The gateway-side per-platform `require_mention` mention-gate (Discord/Telegram/etc adapters) that production group chats actually run on is a different mechanism (per-message gating, not a round-count ceiling) and isn't touched by this PR.
(cherry picked from commit 791fc66)
…esets Replace the install-wide stock/extended ceiling toggle as the only governor with optional per-room modes that override it when set. Unset rooms keep exact prior global-toggle behavior (zero regression). - rooms[name].mode: 'build' | 'decide' | 'standing' | unset - Message budget scales as messagesPerMemberPerRound * members * rounds (fixes the PR #1 5-bot stock bug where a flat 10-msg cap died at round 2) - decide: chat-only prompt rule + deterministic system summary on hard cap - build: long ceilings, tool-capable (prompt does not restrict tools) - standing: moderate multi-trigger coordination ceilings - Room-header dropdown to set/clear mode (persists on the room record) Kanban: t_ee2e6fe4 (cherry picked from commit e212071)
…m port) Port of anthropics/skills discernment-nudge (Apache-2.0, added upstream Aug 17 2026), snapshot 3b3fad96. After a substantive, actionable answer, append 2-3 short targeted follow-up questions that help the user check key facts, probe the reasoning, and notice missing context — at most once per conversation, with explicit skip rules (trivial lookups, formatting, code, creative writing). Reframed as an opt-in output habit, not an identity change; upstream LICENSE.txt carried verbatim. - optional-skills/productivity/discernment-nudge: SKILL.md + LICENSE.txt - tests/skills/test_discernment_nudge_skill.py - Docs: catalog row, sidebar entry, generated skill page (scoped regen). (cherry picked from commit ca713cc)
…xecutable anti-slop gates Port of agiwhitelist/auteur (MIT, ~1k stars), snapshot 9bca227d. Three registers (build / direct / system) on one taste core: commit-sheet-first art direction, asset generation via image_generate + local CLIs, and node-based quality gates (slopscan anti-slop linter, motionqa frame-drop check, systemscan cross-route drift) run through playwright. - optional-skills/creative/auteur: SKILL.md (de-Clauded, Hermes tool framing), LICENSE (upstream MIT), 11 references, 8 verbatim upstream .mjs scripts (all pass node --check; slopscan smoke-run verified), 6 templates. README gallery assets not vendored (size cap). - tests/skills/test_auteur_skill.py: frontmatter, path-annotation invariant, de-Claude residue, related_skills resolution. - Docs: catalog row, sidebar entry, generated skill page (scoped regen). (cherry picked from commit b00b60f)
Anthropic refresh tokens are single-use and rotate on use, but the anthropic provider was excluded from the cross-process auth-store lock and in-lock re-sync already applied to openai-codex and xai-oauth. With several Hermes processes sharing one auth.json (gateway + dashboard + per-profile gateways), the first refresher rotated the token and every other process POSTed a stale one, got HTTP 404, and benched the credential - taking Claude offline about an hour after each login. Route anthropic through the same locked sync-POST-write-back path and widen the pool-store re-sync gate to cover it.
RFingAdam
force-pushed
the
consolidate/upstream-sync
branch
from
August 26, 2026 00:44
6e7aed6 to
41a3df2
Compare
૮ >ﻌ< ა ci reviewran on 31c2337 — ci: bump Python test suite timeout 30->60min (standard runne ❌ Job failuresJS & TS checks / JS & TS checks · View jobJob JS & TS checks / JS & TS checks failed. Python tests / Run tests · View jobJob Python tests / Run tests failed.
|
| Package | Before | After |
|---|---|---|
| @playwright/test | 1.58.2 |
1.62.1 |
| playwright | 1.58.2 |
1.62.1 |
| playwright-core | 1.58.2 |
1.62.1 |
How to fix:
Add the ci-reviewed label after verifying the version changes are expected.
⚠️ Warnings
OSV vulnerability scan · View job
7 known vulnerabilities found in pinned dependencies.
- CVE-2026-67213 in package-lock.json
- CVE-2026-67213 in website/package-lock.json
- CVE-2026-71554 in uv.lock
- CVE-2026-70608 in package-lock.json
- CVE-2026-56876 in package-lock.json
- CVE-2026-70606 in package-lock.json
- CVE-2026-71554 in uv.lock
How to fix:
Review the findings in the Security tab. Update the affected dependencies if a patched version is available.
ℹ️ Info
CI-sensitive file review · View job
PR touches sensitive files, but the ci-reviewed label has been added, approving them.
Sensitive files changed:
.github/workflows/docker.yml.github/workflows/e2e-desktop.yml.github/workflows/js-tests.yml.github/workflows/nix.yml.github/workflows/rust-tests.yml.github/workflows/tests-os.yml.github/workflows/tests.yml
debug info
A member that takes on a task now keeps the floor across rounds until it
decides it is finished, instead of stopping after one message.
Protocol mirrors the existing (pass) contract: end a turn with (working)
to keep the floor, or (done)/(blocked) with a report to release it. A
member that says nothing special behaves exactly as before.
There is no turn ceiling on an open claim. The room's round, message and
wall-clock ceilings are bypassed for the member holding it, so work is
never cut off mid-task. Two exits remain:
- repeating itself for GROUP_WORK_NO_PROGRESS_LIMIT turns releases the
claim and posts a system note saying why
- a hard backstop at GROUP_WORK_HARD_TURN_BACKSTOP rounds, matching
delegation.max_iterations
Only an explicit (working) keeps the floor, so a member that forgets the
sentinel stops rather than holding the room to the backstop.
A member mid-task is continuing its own work rather than reacting to new
messages, so an empty delta no longer skips its turn; it gets a
continue-your-task prompt instead.
Claims live in the room record beside stranded/holds and are carried by
all three serializers, so a window restart resumes the loop. A held
member's claim is ignored until it is released.
On by default in every room, with a room-header toggle to opt out;
workLoop is persisted only when off, so existing rooms need no migration.
The loop does not grant tools: a chat-first room still loops for
multi-turn reasoning, and running tools stays an explicit build-mode
choice.
…merge
The bots pane failed to render with "$botSessionsWorkspace is not
defined". Three references survived the upstream consolidation while the
code that defined them did not:
- $botSessionsWorkspace: upstream removed the sessions-workspace pane
outright. The surviving read had no remaining consumer, so it goes
with the feature.
- showHidden: upstream renamed it hiddenExpanded. Six reads in the
roster eye-toggle still used the old name. showHiddenSection and
showHiddenRows are separate values and are untouched.
- hiddenUnread: only ever defined in the fork. Restored next to
hiddenBots, keyed by botSelectionKey to match how unread is stored
now rather than the fork's bot.name.
The test suite passed throughout: it exercises the group-chat engine but
never renders the roster or pane components, so a dangling identifier in
JSX is invisible to it. eslint no-undef does catch it, and now reports no
undefined project identifiers in this file.
The work loop deliberately has no turn ceiling, which leaves spend unbounded. This adds the one guard that was missing. The desktop has no usage or billing RPC, so the figure is an ESTIMATE from prompt and reply characters (~4 chars per token), not a billed count. It is a smoke alarm, not an invoice: it stops a drive that is burning far more than expected and says plainly that the number is an estimate. Unlike the round and wall-clock ceilings, this one DOES apply to a member holding a claim. A claim means "let me finish", not "spend without limit". On trip, open claims are released so nothing silently resumes and a system note records the spend and the ceiling. Budget resolves as: explicit per-room tokenBudget, then the mode preset (build 400k, standing 120k, decide 60k), then 200k. Zero or negative disables it.
An unanswered approval returned "The user has NOT consented to this action", which is false: nobody refused it, the request went unanswered. A well-behaved agent reads that as a decision, abandons the work permanently, and reports it to the user as rejected. Observed in a live run where an authorised `gh pr create` timed out under CPU starvation: the agent stopped, declined to retry via any route, and reported a consent denial that never happened. The fail-closed behaviour is correct and unchanged - the command does not run, and the agent still must not retry or route around it without fresh approval. Only the claim about what the human did is corrected, plus an explicit instruction to report it as blocked-awaiting-approval rather than as rejected work. Tests assert the new wording and now guard against the regression: a timeout message must never contain "has NOT consented".
…r lane Coordination in a room was purely conversational: a bot saying "assigning t_123 to @backend" was a statement of intent that nothing enforced, so two members could each decide the same task was theirs and open parallel work on it. A thread with an assignee is now that member's lane. Selection filters to the assignee regardless of who was @mentioned, so a non-assignee is never dispatched into someone else's work. Clearing the assignee reopens the thread. Assignments live beside working/holds in the room record and are carried by all three serializers, so a lane survives a reload.
`sessions.model` is the CONFIGURED model. After a provider fallback it is no longer the model answering, so a client showing it reports a reassuring lie - a room can run an entire session on a fallback provider with nothing on screen saying so. Observed live: a bot ran a whole session on gemini while its config, and the UI, said claude-sonnet-5. Adds SessionDB.get_session_usage_summary(), reading session_model_usage (the per-call record) for the most recent route, every distinct route seen, and token/cost totals. Deliberately most-recent rather than dominant: a late fallback is exactly the case worth surfacing, even when it served the fewest calls. The shared live-session payload carries it as `usage`, so session.resume exposes it without a new RPC. The room records it per member and renders the serving model beside each turn, amber when the session has used more than one route. The room's spend ceiling now prefers these real totals, measured as the delta from each member's total at its first turn of the drive, and falls back to the character estimate only when the backend reports nothing - the system note says which figure it used.
…s you Six bots produce a wall of text with no answer to the only question that matters: is any of it waiting on me. Status is now derived from the same turn intent the drive already parses, rather than a second protocol the model has to remember - the (working) sentinel was used once in an entire six-bot run, so anything needing the model to emit extra ceremony does not survive contact. (done) and (blocked) are now distinct intents. Merging them reported work a bot could not finish as finished, which is the opposite of useful. Each member carries working / review / blocked / idle, persisted like the rest of the room record, and the header summarises "N to review, M blocked, K working" - amber for blocked, accent for review, and nothing at all when the room is quiet. The turn prompt now states the distinction explicitly so (done) is not used for work that stalled.
…GH account) ubuntu-latest-32-core / -96-core and windows-latest-32-core are paid GitHub Team/Enterprise runner tiers not provisioned on this personal account (confirmed: gh api users/RFingAdam/actions/hosted-runners -> 404). Every job pinned to one of these sat queued indefinitely with zero runner ever assigned -- reproduced twice across separate CI runs, not a transient scheduling delay. Downgraded to standard ubuntu-latest/ windows-latest across all 4 stalled workflows (tests.yml, tests-os.yml, js-tests.yml, nix.yml) plus 3 more not yet exercised by this PR but carrying the same landmine (e2e-desktop.yml, docker.yml, rust-tests.yml).
… less parallelism than the 96-core it was tuned for)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Brings the fork up to date with
upstream/main(692 commits) and consolidates every fork-unique change onto that base, somainstops drifting.What's on this branch
Rebuilt from
upstream/mainwith 10 commits replayed.Fork-original work:
Carried forward because it is not in
upstream/main: the auteur and discernment-nudge optional skills, and the Python repr redaction fixes.New here:
fix(auth)for the Anthropic OAuth refresh race, described below.The auth fix
Anthropic OAuth refresh tokens are single-use and rotate on use. The
anthropicprovider was excluded from the cross-process auth-store lock and in-lock re-sync thatopenai-codexandxai-oauthalready use. With several processes sharing oneauth.json(gateway, dashboard, per-profile gateways), the first refresher rotates the token and every other process then POSTs a stale one, gets HTTP 404, and benches the credential. In practice Claude went offline about an hour after each login and traffic fell through the fallback chain.This applied cleanly to upstream's base, so upstream still carries the bug.
Conflict resolutions worth a look
runGroupChatRounds: upstream rewrote it to be thread-aware, with stranded-reply harvesting and cancel recording. Kept upstream's rewrite and re-applied the fork's ceilings (maxRounds,maxMessages,deadline) on top rather than taking one side.Room prompt rules: the fork's
rules[]still held upstream's older wording. Folded upstream's current text in, so thetoolCapablechat-only branch survives without reverting upstream's improvement.Disband button: kept the fork's mode dropdown, took upstream's newer
Tip-wrapped button over the fork's oldertitle-based duplicate.Test export list: unioned to 53 names, dropping
GROUP_CHAT_MAX_ROUNDSandGROUP_CHAT_MAX_MESSAGES, which no longer exist.Dropped as redundant
03a448e6a2(transcript jumps) is already upstream asead9d8e3d4, same author and timestamp, and upstream's version is a superset.8397a186ed(Bot Mode invariants doc) came out empty. Upstream's AGENTS.md now documents a registry-based canonical-chat contract that supersedes the pin-based description.Tests
group-chat.test.mjs: 103/103, including the fork-specific room-mode testshermes-bots/tests/*.test.mjs: 580/580Not yet deployed.