Skip to content

Re-pin to upstream v2026.8.27 (fix Codex OAuth refresh race) - #2

Closed
tars-withvariable wants to merge 8757 commits into
mainfrom
repin/v2026.8.27
Closed

tars-withvariable wants to merge 8757 commits into
mainfrom
repin/v2026.8.27

Conversation

@tars-withvariable

Copy link
Copy Markdown

Summary

Re-pins the fork to upstream tag v2026.8.27 (commit 5fc308a707, Aug 27 2026). The fork was previously pinned at an old upstream base with 2 fork-specific commits that cancel each other out (approval timeout add + revert, zero net diff).

Problem

Intermittent Codex OAuth 401 errors cause tars-hermes to fall back from gpt-5.6-luna (openai-codex) to z-ai/glm-5.2 (openrouter). Root cause: single-use OAuth refresh tokens get rotated by racing gateway threads/processes, leaving the loser with a consumed refresh_token and no recovery path.

Fixes picked up

  1. 0ab4cdc (Jul 23) - adopt refresh_token from auth.json even without access_token ([Bug]: openai-codex pool replays a consumed (rotated) refresh_token and goes terminally DEAD despite adoptable fresh tokens; 'auth refreshed after 401' logged when refresh failed NousResearch/hermes-agent#70097). Fixes two defects: credential pool sync skipped adoption when auth store had no access_token, and _try_refresh_codex_client_credentials returned True even when refresh failed.

  2. 7380b48 (Aug 3) - widen source guard to include manual:device_code entries. The above fix was unreachable for source=manual:device_code. Caused a 12-of-16 fleet outage on Aug 1.

  3. 9cd605f (Jul 26) - defer single-use-token refresh outside threading lock. Pool held self._lock during OAuth refresh (20+ seconds), blocking all gateway threads.

Verification

  • Confirmed all 3 fix commits are ancestors of v2026.8.27
  • Confirmed the 2 fork-specific commits produce zero net diff
  • Merge commit contains the full upstream history with no conflicts

beplee and others added 30 commits August 26, 2026 07:28
Enough1122 review points on NousResearch#94417:
1. Precedence hazard fixed: the busy-guard assertion now locates the
   rebind helper body precisely and asserts the guard INSIDE it, instead
   of a 2000-char window with an (m and X) or Y precedence trap.
2. stored_session_id guarantee: documented + pinned — the gateway always
   stamps it ('stored_session_id': session_key or "" in server.py), and
   the rebind's typeof check refuses non-string/empty values, so an
   unnamed rebuilt runtime is never adopted as lineage proof.
3. New third assertion pins that refusal contract.

Structural smoke tests remain structural by design; the behavior
contract for the rebind is exercised end-to-end by the model-switch
manual repro path — a vitest harness driving handleSessionInfoEvent is
the follow-up candidate noted in the reply.
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
…talogs

Slots below glm-5.3, above glm-5.2 in both curated lists; regenerates
model-catalog.json. No new metadata entries needed: context resolves via
the existing glm-5.3 fuzzy key (1,048,576 — matches OpenRouter live), and
both routes bill via official_models_api (live pricing).
…heir recorded endpoints (NousResearch#63206)

A manually-launched `hermes serve --host <ip>` powering a remote Desktop
was invisible to the entire update pipeline: not in the runtime
inventory, a permanent exit-2 dead-end at the Windows venv-holder guard,
and — when anything killed it — never relaunched, stranding the remote
client on a dead endpoint (NousResearch#63206). Serve backends were also visible to
`hermes dashboard --stop` but hidden from `--status` (NousResearch#81564's
asymmetry), so operators could kill what they couldn't see.

Built on the spawn ledger (positive identity, never argv guessing):

- process_identity.py: LedgerEntry gains structured host/port/profile
  (backward-compatible — readers .get()); register_self accepts detail=;
  argv capture widened 6→10 tokens so profiled launches survive.
- web_server.py: serve/dashboard registration moved AFTER the bind and
  now records the ACTUAL bound host/port/profile.
- update_inventory.py: serve/dashboard collector reading the ledger —
  manual backends inventory as supervisor=manual-serve with
  restart_via=respawn-argv; Desktop-owned ones (live recorded spawner)
  as desktop. Plan/receipts/fleet matrix see them for free.
- update_cmd.py: new venv-guard rung — manual serve/dashboard holders
  are stopped for the update and relaunched via an idempotent atexit
  token built from structured identity (same contract as the gateway
  pause/resume); receipts record serve_pause/serve_relaunch.
  Desktop-owned backends keep the refusal (the app respawns what we
  kill).
- dashboard_procs.py: the process scan is augmented with live ledger
  rows, so profiled launches (`hermes --profile p serve ...`) that match
  no substring pattern are finally visible to kill/respawn.
- main.py: `--status` now lists serve-mode backends too, tagged [serve]
  — closing the NousResearch#81564 status/stop asymmetry.

Salvage note: detection deliberately does NOT reuse NousResearch#70742's psutil
cmdline-pattern scan (the argv-guessing class this campaign retires);
its resume-token lifecycle (atexit + idempotent flag) and don't-replay
guard shaped the relaunch contract here — credit @Tranquil-Flow.

Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
…earch#81564)

The old assertions pinned the phrasing that HID serve backends — the
exact asymmetry NousResearch#81564 reports. Re-pinned to the new message and
strengthened: a serve-mode row must now appear, tagged [serve].
…ence results

Live tool-use A/B for session_search schema changes: arms are git refs
(tools/session_search_tool.py extracted per ref), tasks run a minimal
agent loop over OpenRouter against a freshly seeded temp session DB with
programmatic oracles — discovery, forced forward-scroll, AND-miss
broadening, verbatim link emission, profile-link resolution, browse.

Checked-in results/pr95570/ holds the 108-run battery (3 models, 3 reps,
2 arms) that validated the PR NousResearch#95570 schema diet before merge:
base 49/54 vs diet 52/54, avg tokens/task -25%.
…esearch#94030)

The stale-dashboard sweep at the end of hermes update snapshots each killed
backend's HERMES_HOME (_hermes_home_for_pid) but only used it as the per-profile
dedupe key. _respawn_dashboard_processes replays the argv with no env=, so a
backend belonging to a second install (e.g. a launchd KeepAlive sidecar) came
back running on the updating install's default home and stole the sidecar's
fixed port: the supervisor crash-looped on EADDRINUSE and clients on that port
silently talked to the wrong backend.

Drop such candidates in _filter_dashboard_respawn_candidates: a backend whose
captured HERMES_HOME differs from the updater's own get_hermes_home() is not
replayed at all — its own supervisor/user owns its lifecycle. Homes are
normalized the same way _profile_key_for_respawn normalizes home: keys, so
symlinked roots compare equal. An unreadable home (None) stays eligible,
keeping the pre-fix fail-open behaviour.
Desktop composer ignored gateway's confirm_required response for
contributor / expensive models. It painted the target optimistically
then invalidated model-options and refetched the still-active session,
so muse-spark-1.2-contributor appeared to instantly snap back to
gpt-5.6-sol with no explanation.

Now handle the confirmation protocol: on confirm_required rollback the
optimistic state, surface the backend's confirm_message as a warning
notification with a Confirm action, and on Confirm retry config.set with
confirm_expensive_model:true. Preserves data-training consent, keeps
the fix session-scoped and test-covered.
…ack, clarify pending return

- staleness guard in applyConfirmedSwitch: bail and dismiss if
  current model/session no longer matches snapshot this warning was
  created for, preventing stale Confirm from clobbering newer pick
  (Enough1122 review NousResearch#92492)
- neutral fallback for missing confirm_message: 'Confirm this model
  switch?' instead of modelSwitchFailed
- document selectModel boolean: false means not-applied (pending
  confirmation or failed) — pending already shows warning, not an error
- test mock: add dismissNotification mock for guarded flow

Addresses NousResearch#92492 (comment)
- braces for single-statement if(confirmNotificationId) guards
- blank line before return in staleness branch (padding-line-between-statements)
…randing the wake on an unsatisfiable profile gate

Fixes NousResearch#89843. On a shared-remote connection every profile is served
through the primary socket, so waitForFocusedSessionHydration's
profileMatches gate could never become true — a bot chat whose stored
transcript painted within seconds still burned the whole 20s hydration
budget and then stranded the pane with 'Timed out loading <bot>'s
session history'.

The wake now resolves paint-first: once the stored transcript is painted
on exactly the target session, the content is its own proof — the pane
opens immediately and a subtle 'Syncing…' badge (new $hydrationSyncProfile
atom + ChatSyncBadge) shows until the profile gate catches up in the
background. Fail-closed everywhere content is not its own proof: a
superseded/conflicting concurrent wake still rejects, and an
expected-empty chat still waits for the full runtime gate.
… on one strip

With several gateways registered, the Sessions profile rail only ever showed
the active gateway's profiles; reaching a bot on another machine meant a
gateway switch first, then a click on the rail that appeared afterwards. Bot
Mode (NousResearch#91134) and Capabilities already read the union agent roster; the rail
is now its third consumer.

- Every registered gateway's profiles sit on the one strip, in registry order
  (This device first, then by label), each group headed by that gateway's
  kind glyph. The active gateway's squares are unchanged; the others are
  "at rest" (dimmed) with tooltips/accessible names qualified by machine
  (`inbox · Homelab`), so same-named profiles never read alike.
- Clicking an at-rest square performs the same dial → commit → re-home as
  the statusbar switcher, landing on that exact (gateway, profile):
  `selectConnection(id, { profile })`. The spinner sits on the clicked
  square; the previous source stays painted until the target answers.
  Groups keep their slots whichever gateway is active, so a square never
  moves under the pointer that clicked it.
- Right-click on an at-rest square: Switch to / Color / Rename / Edit
  SOUL.md / Delete, executed on the owning gateway (renameProfile,
  getProfileSoul and updateProfileSoul accept the same scope deleteProfile
  already had); the delete confirmation names the machine. The legacy
  per-profile "Connect to a remote host…" item is hidden on multi-gateway
  setups, where the rail shows machines directly.
- Unreachable gateways keep their squares with an amber dot on the glyph;
  two registrations of one backend collapse to one group; past thirteen
  squares across the fleet the strip condenses into a menu sectioned by
  gateway. Roster is fetched on mount / focus / registry change only — no
  periodic fleet polling.
- Single-gateway Desktops render exactly as before: no roster fetch, same DOM.

Also fixes a boot race the e2e surfaced: initializeConnectionsRegistry()
"restored" the launch-mode source over a switch the user had already made
while boot was settling (same class as NousResearch#91047). The restore now yields when
a switch is pending or already landed.

Tests: pure grouping (fleet-rail.test.ts), rail component fleet mode
(profile-rail-fleet.test.tsx), store (explicit profile pick; restore yields),
and a Playwright e2e (fleet-profile-rail.spec.ts) that boots Desktop with two
REAL backends — the local one plus a second `hermes serve` registered as a
remote URL connection — and verifies layout, a real re-home, gateway-scoped
actions, and order stability.

Docs: multi-connection-desktop.md describes the fleet rail.

Refs NousResearch#89304, NousResearch#92384, NousResearch#91047, NousResearch#94724

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…itial-connect failure

MCPServerTask.run() used `_ready.is_set()` to tell a genuine first
connection attempt from a later reconnect. `_ready` is cleared on every
reconnect cycle, so once a server has already registered its tools and
then drops (keepalive failure, transient TaskGroup exit, etc.), the next
failed reconnect attempt is misclassified as "never connected" and burns
the 3-attempt initial-connect ladder instead of the 5-attempt reconnect
budget, parking the server much sooner and logging "failed initial
connection after 3 attempts" even though tools were already registered.

Add a sticky `_ever_connected` flag, set once alongside `_ready.set()`
right after a successful `_discover_tools()` call and never cleared, and
gate the initial-vs-reconnect branch on it instead.

Fixes NousResearch#94654
These pre-existing tests fake a successful first connect by calling
only _ready.set(), which is what the real code did before this PR.
Now that run() gates the initial-vs-reconnect branch on the new sticky
_ever_connected flag instead, their later simulated reconnect failures
were misclassified as never-connected and hit the 3-attempt ladder,
failing test_reconnect_counter_resets_after_successful_session,
test_parked_server_self_probes_and_revives, and
test_retry_attempts_log_debug_transitions_warn in CI. Set the flag
alongside _ready.set() to mirror the real success sites, same as the
new test added in tools/mcp_tool.py's own PR.
Drop the try/except AttributeError guard in the new regression test
now that the slot is always defined, and note in the run() comment
that _ever_connected is set once and never cleared.
…ty, not auth mode (NousResearch#92183)

Two registered basic-auth gateways shared the single
persist:hermes-remote-oauth cookie jar, so signing in to gateway B
evicted gateway A's session cookies (Chromium jars ignore the port) and
A's cookie was silently presented to B on every request. Non-primary v2
registry remotes with cookie auth now ride a per-connection partition
(persist:hermes-remote-oauth:conn:<id>) resolved at the jar boundary;
the registry primary, v1 remote, cloud cascade, and portal flows keep
the legacy shared jar so upgrades do not sign anyone out. Fail closed:
a connection's requests can never see another connection's cookies.
… instead of dropping or nuking the registry (NousResearch#94246)

- normalizeRegistry now preserves every malformed entry (unknown kind,
  url-less remote/cloud, host-less ssh, mangled non-object items, and
  any entry whose normalization throws) under a capped 'quarantined'
  key that survives write cycles — healthy entries keep loading and
  user data is never silently deleted.
- A whole-file parse failure preserves the original bytes in a
  connections.json.corrupt-<ts> sidecar BEFORE the drift reconciler or
  a save can overwrite the file with the degraded local-only registry.
- Loads log a quarantine notice and sanitizeConnectionsRegistry
  surfaces reason+label summaries (never raw entries/token envelopes).
… primary is remote (NousResearch#91564, NousResearch#90316)

'Make primary' on a registered remote/cloud/ssh gateway only rewrites
connections.json — the v1 config.mode stays 'local', so startHermes()
resolved no remote route and spawned a loopback 'hermes serve' the
desktop never uses (full MCP set duplicated, port squat, respawn on
poll). resolveDesktopRemoteRoute gains a lowest-precedence registry-
primary rung (source: 'registry', existing v1/env/profile precedence
untouched), and globalRemoteActive() now recognizes a remote registry
primary so local-entry routes force pooled local children instead of
delegating into a primary that dials remote. A 'local' registry
primary still resolves null — genuinely-local desktops unchanged, and
local-profile secondaries keep their forced-local pooled backends.
…rofile switches (NousResearch#92434 close-candidate)

Reproduces the reported Bot↔Default switch shape at the gateway.ts
activation seam: a switch-back that lands while the outgoing switch's
WS handshake is still pending keeps the route, the late-completing dial
neither steals the foreground nor breaks its socket, and re-activating
the bot works without an app restart. The guard (activation epochs +
open-socket-publish, landed via NousResearch#89622/NousResearch#92265/NousResearch#81094) already prevents
the reported permanent break; this pins it so it cannot regress.
… the graceful quit wait (NousResearch#91668 remainder)

The NousResearch#95085 quit teardown kills the owned serve --isolated before the
SSH tunnel closes, but a backend mid-turn (in-flight LLM call, live MCP
children) can ride out SIGTERM past cleanupStale's 5s graceful wait.
The old code then gave up (threw, kept the lockfile) and before-quit's
6s race closed SSH anyway — reparenting the still-running serve to
pid 1: the reported leak, now specific to quit-during-active-turn.
Escalate to kill -9 with a confirmed-exit wait; only an unkillable pid
(D-state, permissions) still throws and preserves the lock record so
the next connect's reap pass retries.
The watcher now skips the quiesce when self_suspend_available() is
False, but its docstring and the one on self_suspend_available() still
said the suspend step is skipped and dormancy still happens. Say what
happens instead: the gateway stays connected until the platform freezes
it.
The three existing assertions are absence checks, so the test also
passed when the loop got no iteration inside the sleep window (with
interval=5.0 it passes without the gate ever executing). Checking
_scale_to_zero_no_suspend_logged proves the branch was taken.
teknium1 and others added 21 commits August 27, 2026 03:56
…#91868, NousResearch#94569)

stopGroupThread(group, thread, members?) is the room's first true
cancellation primitive: it bumps the room epoch (the driving loop bails
at its next member boundary), sets NousResearch#93129 holds for every member (no
future turns until an explicit release), records a 'stopped' activity
event on the new epoch, and sends session.interrupt to the member
currently on turn via its own route — previously the plugin issued zero
interrupt RPCs, so 'stop' meant waiting out the in-flight model call.

The runGroupChatMemberTurnLeased poll loop now abandons a turn whose
dispatch epoch went stale WHILE its member is held — the stop signature.
An ordinary newer-send epoch bump without a hold still polls to
completion so late work keeps landing (NousResearch#93127 commit check unchanged).
…ch#94570)

Salvaged from NousResearch#94570 (@ShonnQ): the Activity bar gains a Stop button
while a round is running (room.running), plus an inline Stop on the
expanded 'working' activity row. Rewired from the original per-member
session.interrupt spray onto the stopGroupThread primitive so the round
loop actually stops (epoch bump + holds + on-turn interrupt) instead of
marching to the next member; labels are plain English like the rest of
the plugin's UI strings.

Co-authored-by: Hermes Agent <agent@nousresearch.com>
…itive

Source-contract tests (the group-room-ux pattern): the workspace renders
the Stop button only while room.running, wires it to stopGroupThread
(not the NousResearch#94570 per-member interrupt spray), and carries no hardcoded
CJK label.
…pass)

runGroupChatMemberTurn (and harvestStrandedGroupReply) selected only the
last assistant message in a finished turn. A Codex intent-ack continuation
nudge can land a complete, substantive room answer and then get a
synthetic "(pass)" reply to the nudge itself — the terminal message picked
by the old scan, which silently discarded the real answer (NousResearch#94376).

Both call sites now scan the messages appended this turn for the last
substantive (non-pass) assistant reply, falling back to a pass only when
no substantive answer exists in that window.
…path

Address review feedback on NousResearch#94386: the new tests only exercised
runGroupChatMemberTurn's use of pickGroupTurnReply. Add the analogous
case for harvestStrandedGroupReply (substantive answer -> synthetic
continuation nudge -> (pass) tail) and document the pass-only tie-break
(newest wins) in pickGroupTurnReply's docstring.
… group chat

The agent loop writes an internal "(empty)" sentinel when the
nudge/prefill/empties/fallback ladder all fail. The gateway converts it
into a user-friendly notice at delivery, but the desktop group-chat
bridge appended the raw sentinel into the room log (seen posting
"(empty)" in a Bot Mode group room), and it synced to the shared
ui_meta for mobile.

Normalize at the single choke point, appendGroupChatEntry, mirroring
gateway/run.py substitution so group chat and gateway surfaces show the
same text. (pass)/empty silence semantics unchanged. Includes a
regression test proven to fail on the pre-fix code.

Fixes NousResearch#94308
OpenRouter and Nous already list z-ai/glm-5.3-flash (NousResearch#95621). The
native z.ai picker, OpenCode Go/Zen fallbacks, setup wizard, and
Coding Plan probes did not. Context still resolves through the
existing glm-5.3 1M key.
…enRouter)

Live /api/v1/models probe (2026-08-27) confirms the id is gone from the
catalog, so the curated picker entry was a dead pick. Manifest
regenerated. No provider-agnostic metadata existed for the slug.

Delist credit: @orouge97 flagged this in PR NousResearch#80036.
…tdout

Since 6d4e851 the serve startup path imports tui_gateway.server (for the
flush-on-SIGTERM handlers) before the READY sentinel is printed. That module
redirects sys.stdout to sys.stderr at import time, so the
HERMES_(BACKEND|DASHBOARD)_READY port=<n> sentinel landed on stderr while the
Electron desktop spawn watches child.stdout only — the desktop timed out
after 90s and killed a perfectly healthy backend (issue NousResearch#96282).

Write the sentinel to the real stdout file descriptor (fd 1 is untouched by
the Python-level redirect), with a print() fallback.

Adds a regression test that captures stdout/stderr separately — the existing
E2E suite merges them, which is exactly how this slipped past CI.
…site

The same stdout redirect that rerouted the READY sentinel (NousResearch#96282) also
reroutes the machine-parsed BACKEND_PORT_IN_USE sentinel printed by
_report_port_in_use() — both preflight and probe-to-bind-race callers run
after tui_gateway.server's sys.stdout=sys.stderr swap. Extract the fd-1
write into _write_machine_sentinel_line() and use it at both sentinel
sites; human-facing hint lines stay on print().
…L stderr in split-stream test

- _write_machine_sentinel_line: wrap the print() fallback so a closed
  redirected stream (ValueError, not OSError) can't propagate out of the
  ready path and kill a healthy serve; document that pythonw port
  discovery relies on the HERMES_DESKTOP_READY_FILE channel, not stdout
- regression test: stderr=DEVNULL instead of PIPE — with the stdout
  redirect active all server logging lands on stderr, and an unread
  stderr pipe can fill and block the child before the sentinel, flaking
  the test at the 120s timeout
…sResearch#78183)

httpx timeout exceptions (ReadTimeout, WriteTimeout) stringify to "",
which defeats _is_timeout_error's first-line guard (if not error: return
False).  The base-layer plain-text fallback then re-sends an already-
delivered message — the user receives it twice.

Replace error=str(exc) with error=str(exc) or type(exc).__name__ at every
httpx-based adapter boundary so the existing matcher ("readtimeout",
"writetimeout") still fires.  ConnectTimeout intentionally stays
unmatched: if the connection never opened the message was not delivered,
so retry/fallback remains correct.

Applies to BlueBubbles (send + _create_chat_for_handle), WhatsApp Cloud
(text + interactive + media), QQ Bot (send chunk + keyboard + media), and
Yuanbao media handler — the same latent bug exists in every adapter that
stores error=str(exc) from an httpx call.
…ing the Phase 3a Protocol

agent/deadline.py defined SuspectableBackend twice: the Phase 3a Protocol
(sync ensure_healthy(self) -> bool) and, further down the same module, an
unrelated concrete class with the same name (async
ensure_healthy(self, timeout=5.0)) added later by the MCP Phase 3b adopter.
Since Python executes class statements top-to-bottom, the second definition
silently shadowed the first at module scope.

Nothing in the tree imports or subclasses either by name today — the MCP
adopter duck-types the same-shaped contract directly on its own connection
class rather than referencing agent.deadline.SuspectableBackend — so this
caused no live behavior change. But it left the wrong (and differently
shaped) class resolvable under that name for the next Phase 3b adopter that
does import it for a type hint.
Merges upstream tag v2026.8.27 (commit 5fc308a) into the fork.

The fork was previously pinned at an old upstream base with 2 fork-specific
commits (approval timeout add + revert) that cancel each other out (zero net diff).

This merge picks up 3 critical Codex OAuth refresh race fixes:
- 0ab4cdc: adopt refresh_token from auth.json even without access_token (NousResearch#70097)
- 7380b48: widen source guard to include manual:device_code entries
- 9cd605f: defer single-use-token refresh outside threading lock
…able

Allow HERMES_CODEX_REFRESH_SKEW_SECONDS (default 120) and
HERMES_AUTH_LOCK_TIMEOUT_SECONDS (default 15) to be overridden via
environment variables. This lets production deployments increase the
proactive refresh window and cross-process lock timeout to prevent
401 fallbacks when multiple gateway threads share Codex OAuth tokens.

No behavior change with defaults.
jakeoliver-withvariable pushed a commit that referenced this pull request Sep 22, 2026
Kanban cards have no length limit, but the session title store rejects
titles past SessionDB.MAX_TITLE_LENGTH with ValueError. _persist_session_title
reads that as a unique-title collision, retries with a "#N" suffix (longer
still), and the caller suppresses the second failure - so a worker spawned on
a >100-char card ended up with no title at all, where main at least gave it a
derived one. Trim the card title (with room for the "#N" retry suffix) before
persisting; a retried card now gets "<trimmed> #2" within the cap.

Review finding: >100-char card title left the kanban worker session untitled.
jakeoliver-withvariable pushed a commit that referenced this pull request Sep 22, 2026
…, with or without the multiplex flag

Two authority gaps in served_profile_child_env (NousResearch#111617 review, andrexibiza P1 #1/#2,
kvnloo finding 1):

- The base was hermes_subprocess_env(inherit_credentials=True) = the launch environ's
  provider credentials; strip_launch_profile_env only knows names with .env/source
  provenance, so a key systemd/Compose/the shell injected into the launch process
  survived into profile B's child whenever B did not define the same name. Now a ROUTED
  target scrubs every Tier-1/Tier-2 credential from the base regardless of provenance
  before B's own scope is overlaid (the child boundary gets get_secret's contract: a
  scoped miss is no credential, never ambient fallback). The launch profile's own child
  keeps its env. bot_relay's base=os.environ goes through the same scrub.
- strip_launch_profile_env / the scrub keyed on is_multiplex_active(); the Desktop and
  dashboard backends serve ?profile=B by installing the HERMES_HOME override without
  that flag, so B's slash worker / helper children kept A's .env and settings. The
  authority test is now "is the target a routed home" (target != process home).
- _build_browser_env resolved the passthrough keys via get_secret, which falls through
  to os.environ on a scoped miss while multiplexing is inactive: a routed B with no
  Firecrawl key got A's. Under serves_routed_profile() the bound scope is the only source.
- served_profile_child_env(inherit_credentials=True) with no target and no scope bound
  under multiplex minted with the launch credentials (key_cmd TTL refresh on a worker
  thread); it now raises UnscopedSecretError like get_secret.

tests/tui_gateway/test_served_profile_child_env_authority.py: ambient-only A key + B
missing it (mux on), flag-off routed B (helper child + browser), real child observation.
3/3 red on base.
jakeoliver-withvariable pushed a commit that referenced this pull request Sep 22, 2026
`_TERMINAL_KANBAN_TOOLS` listed only `kanban_complete` / `kanban_block`,
so the turn-end guard fired at workers that had already handed the card
off correctly:

- A build worker that calls `kanban_request_review` moves the card from
  `running` to `review` (tools/kanban_tools.py `_handle_request_review`),
  yet `session_called_kanban_terminal()` returned False and the nudge
  told it "Task is still `running`" — false by then — and to call
  `kanban_complete`, which would close a card that must go through
  review. Both goal-mode prompts name that tool explicitly
  (hermes_cli/goals.py `KANBAN_GOAL_CONTINUATION_TEMPLATE` and
  `KANBAN_GOAL_FINALIZE_TEMPLATE`), so the prompt and the guard
  contradicted each other.

- Review agents hit the same wall. The review lane spawns through the
  same `_default_spawn`, which sets `HERMES_KANBAN_TASK`, so the guard is
  active for them, and the force-loaded sdlc-review skill's decision
  table ends the request-changes path on `kanban_request_changes`.

Adds both handoff tools to the set. Workers that obey their prompt now
exit cleanly; the guard still fires for a non-terminal board tool, which
is the case it exists for.

Covers lifecycle mismatch #2 of NousResearch#94916 only — the `dispatch --dry-run`
half is a separate change and is discussed on the issue.
jakeoliver-withvariable pushed a commit that referenced this pull request Sep 22, 2026
… in the flag table

- Dialog 2 now shows the read-then-answer step for one benign prompt and
  warns against blind timed Enter, and names --permission-mode acceptEdits
  as the narrower opt-in (idea from PR NousResearch#113456).
- Quick Reference row and pitfall #2 carry the opt-in wording instead of
  teaching Down+Enter as the expected move.
- Regenerated the claude-code docs page.
jakeoliver-withvariable pushed a commit that referenced this pull request Sep 22, 2026
Adopts the DNS-rebinding-pinned transport hardening (repo issues #2/#3,
PR #3) and the starter-feed/settings failure-surfacing + SSRF-gated icon
proxy fix (issue #6, PR #8). Full range in
tony-simons-aiowa/hermes-newswire deccdc4..e6b438e (13 commits):

Security-relevant highlights:
- All outbound fetches (feeds, redirects, icons) now go through a pinned
  transport: the SSRF gate's validated address set is bound to the actual
  connection — no second DNS lookup, so DNS rebinding/TOCTOU has no
  window; the plugin fails closed if the pin seam changes.
- New GET /icon.json proxies favicons through the same gate and returns
  base64 data URLs — the renderer's <img> no longer performs unpinned
  DNS resolutions of feed-controlled hostnames. 64 KB cap enforced
  mid-transfer; image content-type allowlist; bounded, normalized TTL
  cache.
- Renderer surfaces backend failures (settings/sources banners,
  starter-feed inline errors) instead of silent no-ops.

Capabilities unchanged (all empty — dashboard plugin, no tools/hooks/
env). Verification at the new pin: 129 pytest, 33 renderer interaction
checks, 26 ESM render smoke, hermes plugins validate clean.
jakeoliver-withvariable pushed a commit that referenced this pull request Sep 22, 2026
…xes)

Picks up the plugin's merged fixes: empty-input tool calls no longer fail the
request (#2), Haiku 4.5 requests no longer send adaptive thinking, top-level
schema combinators are stripped before the native validator, the claude CLI is
resolved on the env PATH, and relay/native failures now name their real cause
instead of only the admission denial. Plugin version unchanged (0.3.0); the
card art URL follows the pin.
jakeoliver-withvariable pushed a commit that referenced this pull request Sep 22, 2026
gc_blobs carried two byte-identical `warning("malformed ledger line; blob
GC skipped"); return 0, 0` blocks — one under `except json.JSONDecodeError`,
one under `if not isinstance(row, dict)` (tools/skill_ledger.py:350-357;
simplify reuse #2, efficiency note, re-gate G2 S2). Funnel the decode
failure into the type guard (`row = None`) so the abort exists once.
Behaviour is unchanged for both the [malformed-json] and [non-dict-row]
parametrizations: any line that is not a JSON object still aborts the
sweep with the same warning.
jakeoliver-withvariable pushed a commit that referenced this pull request Sep 22, 2026
argparse's _get_value always calls a `type=` callable with the raw str
token, and int(<str>) can only raise ValueError, so the TypeError arm in
`except (TypeError, ValueError)` is unreachable. The try itself stays:
without it argparse would print "invalid _nonnegative_int value: 'thirty'",
leaking the private helper name into the usage error.

Invariant: `_nonnegative_int` is only referenced as an argparse `type=`
(hermes_cli/kanban_parser.py) and by the unit test, which passes str.

Finding: simplify/D.quality.md #2 (hermes_cli/kanban_parser.py:50).
Dead-branch deletion; existing test_gc_parser_rejects_negative_retention_days
still covers the "thirty" -> "must be an integer" path (green).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.