chore: sync upstream (49 commits behind) - #3
Merged
POWERFULMOVES merged 49 commits intoAug 10, 2026
Merged
POWERFULMOVES merged 49 commits into
POWERFULMOVES merged 49 commits into
Conversation
…stop matching lifecycle words inside SQL/text Two live failures on the same guard (cron/lifecycle_guard.py), both of which blocked legitimate diagnostics from inside the gateway: 1. Crash class: the referenced-script walk read compiled binaries as if they were shell scripts. Reading/inspecting a referenced file is now best-effort by construction: executable magic numbers (ELF, PE, Mach-O fat/thin) short-circuit before any full read via a 4KB sniff, NUL-bearing heads are skipped as non-scripts, and unreadable paths of every kind (NUL bytes in the token, ENAMETOOLONG, missing files) degrade to "nothing to scan" instead of raising. A second fail-safe layer wraps the pure-string fallback so the boundary function stays total even if the tokenizer itself fails. 2. False-positive class: the lifecycle regex matched its command shapes inside DATA arguments — SQL string literals passed to sqlite3/psql and grep/rg/journalctl patterns hunting for the lifecycle string in logs. Added a fail-closed second-pass exemption: on a raw regex hit, re-scan with data-sink executables' arguments masked; only a match that survives (i.e. sits in command position) blocks. Masking is skipped for pipes into shells/xargs, command/process substitution, sqlite3 dot-commands and psql backslash escapes, so it can only ever allow, never miss. Behavioral tests: exact live false-positive shapes as negatives, the smuggling shapes as positives, the kill-primitive positive catalog unchanged, and an adversarial never-raises suite (NUL bytes, non-UTF-8, /dev/*, directories, missing files, magic-prefix binaries).
…peners The WAL-reset-vulnerability gate (NousResearch#70055 lineage) could flip a LIVE WAL database to journal_mode=DELETE while another process was writing to it. Observed on state.db (Aug 5): a pytest process on the repo .venv (SQLite 3.50.4, vulnerable) opened the live ~/.hermes/state.db while the gateway (SQLite 3.53.1, WAL) held it, downgraded the journal mode, and destroyed the gateway's committed-but-uncheckpointed WAL transactions (disk rows went 10 -> 0 while memory held 185). cron/executions.db already had the "leave WAL in place, no live downgrade under concurrent openers" rule via the on-disk WAL probe; state.db and every other store shared the hole whenever the mode PROBE itself was blocked by a concurrent opener's locks ("could not read the mode" was treated as "not WAL" -> flip anyway). Generalized in the single journal-mode owner (apply_wal_with_fallback), covering ALL call sites (state.db, kanban.db, projects.db, cron/executions.db, delivery_ledger, async_delegation, verification_evidence, discord recovery, response_store.db, memory_store.db): - _set_journal_mode_no_wait(): the only journal-mode switch primitive for non-WAL targets. Forces busy_timeout=0 around the pragma so SQLite's own exclusivity requirement for leaving WAL becomes the concurrent-opener detector — any other opener (this process or another) makes the flip fail immediately instead of waiting out a busy timeout and sneaking the flip in under a live writer. - Vulnerable-SQLite gate: an unreadable journal mode (probe blocked) now means "ownership not provably exclusive" — leave the mode untouched and warn, never flip. A lock conflict on the flip itself likewise leaves the mode alone. - Configured journal_mode=delete: refuses (raises) rather than downgrading blind when the mode cannot be verified under a concurrent opener. - Filesystem-incompat fallback: re-raises instead of downgrading when the on-disk mode cannot be verified. - New/exclusively-owned DBs on vulnerable builds behave exactly as before (DELETE gate retained per NousResearch#70055). Behavioral tests use a REAL second process (and a real second connection holding an exclusive lock) with the blocked-state assertions running WHILE the holder owns the DB, plus exclusive-ownership downgrade-still-happens coverage.
…ive-DB isolation guard)
Forensics on a live developer machine found pytest fixture rows inside the
REAL ~/.hermes/state.db — sessions with chat_id 'chat-1', '123', 'wx-chat',
and gateway_routing rows whose scope was literally under /tmp/pytest-of-*/.
A pytest-spawned process also opened the live DB and flipped its journal
mode (journal_mode=DELETE fallback on SQLite 3.50.4) under the WAL-mode
gateway writer, destroying committed transcripts ("Persisted transcript
lagged live cached history ... possible FTS write corruption", 15+
occurrences). The existing live-system guard covers kill primitives but not
the SessionDB/SessionStore write paths.
Root cause (leak vector): the session-level HERMES_HOME sandbox in
tests/conftest.py only created a tempdir when HERMES_HOME was UNSET. On a
machine where the shell (e.g. gateway-launched, or an exported
HERMES_HOME=~/.hermes) hands pytest the production home, the sandbox was
skipped entirely — every argless SessionDB()/SessionStore() and every
collection-time DEFAULT_DB_PATH froze onto the real state.db.
Fixes (fail the class, one owner):
* hermes_state._ensure_test_isolation(): single choke point wired into
SessionDB.__init__ (every construction, incl. read_only). Under pytest
(PYTEST_CURRENT_TEST / PYTEST_VERSION — inherited by subprocess
children), a db path resolving to <real-root>/state.db or
<real-root>/profiles/<name>/state.db raises RuntimeError('live-system
guard: ...') before any connection, mkdir, or journal-mode pragma.
* tests/conftest.py: session sandbox now also tempdir-redirects a pre-set
HERMES_HOME that points at the production root (the actual escape
vector); kanban deny-list capture updated to match. New autouse
_state_db_write_guard fixture honors the existing
@pytest.mark.live_system_guard_bypass marker as the escape hatch and
feeds custom (non-~/.hermes) production roots into the guard deny-list.
* gateway/session.py: SessionStore.__init__ no longer swallows the guard's
RuntimeError into the JSONL fallback — guard trips are loud.
* tests/hermes_state/test_live_db_isolation_guard.py: behavioral
regression tests — production paths (direct, profile, read-only,
unnormalized, default-resolution) raise; tmp HERMES_HOME works; bypass
marker works; SessionStore re-raises guard errors but still degrades on
ordinary failures; subprocess child without HERMES_HOME is refused while
a hermetic child succeeds.
No new HERMES_* env vars; no hardcoded ~/.hermes (platform root comes from
hermes_constants._get_platform_default_hermes_home()).
… Path.home monkeypatches Tests like tests/gateway/test_goal_verdict_send.py monkeypatch Path.home() to a tmpdir; resolving the guard's 'real root' through Path.home() made the test's own hermetic home look like production (false positive). Resolve via os.path.expanduser/LOCALAPPDATA instead — the hermetic conftest never rewrites HOME, so this always names the actual production root.
…ing it The turn finalizer already hands back steer text that queued after the final tool batch — result["pending_steer"], with the comment "hand it back to the caller so it can be delivered as the next user turn instead of being silently lost." Every interactive surface honors that contract (cli.py, gateway/run.py, tui_gateway/server.py all requeue it). The delegation layer doesn't: _run_single_child never reads it, so a steer queued into a delegated child that finishes first vanishes with no trace in the completion entry. There is also no sanctioned sender: the registry has interrupt_subagent() but no redirection-side mirror, and session.steer cannot reach children (lazy watch sessions have agent=None, so it 4010s). Complete the contract for delegated children — both halves: - steer_subagent(subagent_id, text): redirection-side mirror of interrupt_subagent(). Resolves the live child in _active_subagents and queues text via AIAgent.steer(). True means queued, not delivered. - missed_steer retention: when the child's result carries pending_steer, _run_single_child names it on the completion entry (missed_steer field plus a summary note) so the parent can re-issue the guidance instead of trusting it landed. This is what makes adding a sender safe: without it the finish-before-drain race silently loses the text — the exact loss the finalizer contract exists to prevent. - subagent.steer gateway RPC beside subagent.interrupt so programmatic hosts (dashboard, voice layers, ACP bridges) get an in-tree caller; catalogued in programmatic-integration.md. - docs: "Steering a Running Subagent" section in delegation.md covering the queued-vs-delivered semantics. Tests: registry-level steer coverage (delivery, unknown id, empty text, dead record, raising agent), the finish-before-drain race retaining missed_steer, and the RPC contract (4000/4002 validation, queued and rejected envelopes).
Per the 'when in doubt, optional' rule — niche prediction-market data skill that sees no regular use; belongs alongside stocks in the finance optional category rather than the default bundle. Install via: hermes skills install official/finance/polymarket
Reload chose the in-process bootout/bootstrap path based on POSIX ancestry, but bootout tears down the job's process coalition, and coalition membership is inherited at spawn and survives reparenting. A gateway-spawned process reparented to PID 1 is no longer an ancestor yet still dies with the coalition, so the retry loop was killed mid-bootstrap and nothing re-registered the label (KeepAlive can't revive a job launchd no longer knows about). - always prefer the detached transient-job helper; it's also correct when genuinely outside the coalition, just asynchronous - wait for the old gateway PID to exit before bootstrapping; bootout only sends SIGTERM and every bootstrap during the drain fails EIO - fall through to the in-process path when the helper can't spawn instead of leaving the plist rewritten but never reloaded
The reload retry loop treated `launchctl list <label>` exit 0 as success, but exit 0 also covers a registered-but-not-running definition (macOS 26+ `state = not running`) — the same trap _probe_launchd_service_running already guards against. Require a PID so success means launchd is supervising a live process, in both the Python loop and the shell helper. Verified against live launchd: a RunAtLoad=false job reports exit 0 with no PID, which the old check accepted and the new one rejects. Note this is NOT what distinguishes a draining instance — measured, the label deregisters within ~1s of bootout while the old process drains on. Waiting for the old PID to exit is what covers that.
Review follow-ups on the salvage: - The helper script's success check accepted any "PID" line, including the "PID" = -1 a recently-crashed job reports — while the in-process path's _parse_launchd_pid_from_list_output rejects non-positive PIDs. Both bash sites now require a positive PID (grep -qE '"PID" = [0-9]+;') so the two paths enforce the same supervised-PID standard. - _graceful_restart_via_sigusr1's drain-wait tail was a duplicate of the new _wait_for_pid_exit — now delegates to it. - Stale comments: the ancestry-detection framing at the top of the reload block, and the exhaustion log's '(refresh ran outside gateway process tree)' which is false on the new helper-spawn-failure fallback path (now '(in-process fallback path)').
… too The pane, in-app browser, and reaction tools were gated on HERMES_DESKTOP=1 — an env var set only on backends Electron spawns itself (local and SSH). A desktop client connected to a plain URL gateway or Hermes Cloud lost all six: they were stripped from the schema before the model saw them, on the same backend whose platform hint was telling it "you are chatting inside the Hermes desktop app". open_preview, read_preview, read_terminal, close_terminal, focus_pane, and react_to_message were all silently absent. The client is not the host. Capability now resolves from the session's own source, which session.create already carries: - The six tools move into a `desktop_ui` toolset, off _HERMES_CORE_TOOLS so no other platform pays their schema. - _gui_surface_toolsets(platform) folds `desktop_ui` (and the existing `project` tools) into the GUI gateway's resolution when the session's platform is the desktop app — the same answer on every topology. - check_fn drops the env probe. It kept the one thing that is genuinely a per-process/user fact: react_to_message's display.message_reactions opt-in, which the desktop mirrors onto whichever gateway it is connected to. react_to_message was doubly broken: it read that toggle behind the env gate, so even a local-backend user's Settings toggle could not reach a remote session. The embedded terminal pane keeps working correctly the other way round: it runs `hermes --tui` against a desktop-spawned backend, and a tui-sourced session gets no GUI tools even though HERMES_DESKTOP=1 is set on that process.
…ess env The rule the preview-tool bug broke, written down so the next GUI-adjacent tool does not rediscover it: the client and the backend are separate machines, so "was this process spawned by Electron?" cannot answer "is a GUI watching?". Names the working pattern (toolset gates the surface, check_fn answers only reachability or user opt-in), the process-wide check_fn TTL cache that makes it the wrong home for a per-session answer, and the test that would have caught it — assert the GUI session gets the tool with the env var absent.
The six GUI tools moved out of `terminal` into their own toolset; the tables still described them as check_fn-gated members of it, and as available to every hermes-* platform bundle.
PATCH /api/sessions/{id} only accepted title and end_reason, so the
`pinned` flag the desktop sends was rejected as an unsupported field —
and the client swallows that error. Pins lived in one app's localStorage
and never reached state.db, which also meant the server-side auto-archive
sweep was free to hide the chats a pin exists to keep.
Accept pinned and archived as booleans, route them to the SessionDB
setters that already existed, and include both in the serialized session
so clients can reconcile against server truth.
The list endpoints deliberately back-fill pinned conversations past their LIMIT, then the client sliced the response back down to that same limit and threw them away — so only pins that happened to land inside the most recent page ever rendered, which reads as a cap on how many sessions you can pin. Keep the back-filled rows when trimming, and discount them from the "window came back full" test that drives Load more. Counting a back-fill as a loaded row invented a page that could never be fetched, leaving a Load more button that refetched the same rows forever.
The guard that stops a stale list page from reverting a fresh pin was released on the PATCH's own ack. A list request issued just before the write is slower than the write, so it lands after the ack still carrying the old value, with no guard left to fence it: the pin flips back and the next reconcile pushes that wrong value to the server, making it durable. Keep the guard until a page actually confirms the value written, with a cooldown so a row that never returns can't fence itself forever, and drop it outright when the write fails — the server never changed, so it stays authoritative. Also reset the mirror bookkeeping on a gateway switch. mirrored/pending are per-backend facts; carrying them across a re-home told us the pins were already pushed to a backend that has never seen them.
Two ways a pin got misfiled. The duplicate: a pin is stored on the durable lineage root, but recents, the messaging slice and the backend project tree are three independent fetches and each can surface the same conversation under either its live tip or its root — so the filter compared one identity against the other, missed, and the session rendered in both Pinned and its project group. Match on every id the pin is reachable under. The lost reorder: a drag only reports the pins whose row is loaded, and setPinnedSessionOrder required that list to match the stored one in length, so a single unresolved pin discarded the whole reorder. Treat it as a permutation of a subset — re-slot the named ids, leave the rest.
Dragging one row switched the entire sidebar into a frozen manual mode with no date dividers at all — permanently, for every session, because the manual order replaced the recency sort outright instead of layering on it. Chronology and ranking are separate concerns: keep the calendar buckets where recency put them and apply the hand-picked order only within a bucket, so a drag ranks a chat among its own day's chats and the dividers survive. Rows move as clusters, so a reorder can't strand a branch child from its parent, and a session the saved order doesn't name keeps the slot recency gave it. Two supporting fixes fall out: dnd-kit now receives the ids it actually renders (it was handed the unrendered session order, so a drop computed its target against a list the user wasn't looking at), and an older page that loads no longer jumps above the hand-picked rows — new ids fold in by position rather than all hoisting to the top.
The transcript decides what a tool call draws and the render budget has to price it. Both sides need the same answer, so the classification moves out of the tool renderer into its own module rather than the budget importing the formatting and i18n weight of fallback-model to ask one question. Adds isSilentTool for the rows that render nothing at all: todo is hoisted to its own panel, and a reaction's UI is the emoji on the bubble.
One weight function served two budgets that protect different things. The store window protects the heap: every message it admits is normalized into the runtime repository whether or not the transcript collapses it, so it has to price the payload it holds. The DOM budget protects the paint, and what a turn mounts is decided by the grouping, not by the bytes behind it. Charging the DOM budget for payload made it count work that never happens. A settled run of twelve reads is one grey summary line, a thought is one collapsed disclosure, a todo is hoisted out of the transcript, and an image is one img however long its data URL — all of it priced as if fully expanded. messageStoreWeight keeps the payload price for the window. messagePaintWeight prices what mounts: collapsed rows flat, silent rows free, cards fixed, and markdown and diffs by size, since those really do build DOM. Both share one character ceiling per message rather than one per part.
On real sessions the button showed up two or three turns from the bottom, over a screen and a half of transcript that had barely painted anything. The budget now spends paint weight, which is what the DOM actually mounts, and 600 units of it — 10-20 agentic turns measured, where a tool-heavy turn prices at 30-90 and a plain exchange at 5-10. A floor of 8 turns covers the session of enormous turns that a weight-only cut still truncates hard; it applies to a real page only, so the small first-paint commit stays small and the backfill a frame later fills the rest. Measured on four stored sessions at the same budget: one went from 3 turns visible to 12, another from 3 to 4, two unchanged. The store window still caps what the DOM can reach at all.
Pinned was capped at half the viewport by its own nested scroller, so past roughly a dozen pins the rest were reachable only by scrolling inside a scroller — a pin you have to go hunting for isn't doing its job. Drop the cap and let the section grow into the sidebar's existing scroll, and stop virtualizing Pinned: virtualization needs a bounded viewport to measure against, which is exactly what's being removed. No count badge, no "show more" — pin as many as you want and they all render. Also back-fill pins on the API-server list route, which was the one list path still windowing purely on recency.
…er-threshold Show earlier messages no longer hides most of a session
…n-sort Pin as many sessions as you want, and they stay where you put them
…ocess-elapsed Keep elapsed status text from overlapping
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
…chat wire Reasoning-summary models emit one reasoning_content delta per completed summary part, each a self-contained bold heading. The Responses API delimits those parts with summary_index; the OpenAI chat wire carries no such field — verified live against Nous Portal, whose reasoning chunks contain nothing but delta.reasoning_content — so concatenating them glued every part into one unspaced, half-bold paragraph. Re-derive the boundary from the signal the wire does carry: a delta opening a closed bold heading against a mid-line tail. This matches Hermes own Responses adapter, which already joins its summary parts with a blank line.
The native Responses stream does carry summary_index, so the part boundary is structured data here rather than something to infer. Break on a change of index, and leave streams that send no index (plain reasoning_text) untouched.
Repairs what is already in the transcript: reasoning persisted before the backend fix, and any provider still gluing its parts. Handles both shapes — heading-onto-heading (the **** run) and prose-onto-heading (vercel/ai#6742). Verified against 46 real glued messages from a gpt-5.6-sol session; all repair cleanly and idempotently.
It was a 1:1 rebuild of the core statusbar gateway item and shipped enabled by default, so the pill showed up twice. Core chrome stays in shell; demos that clone it belong in hermes-example-plugins.
…eway-pill fix(desktop): stop double-painting the gateway pill
…summary-blocks Reasoning steps read as separate blocks again instead of one glued paragraph
…ssion-integrity fix: preserve session history when a turn crashes
After `hermes update`, the desktop sidebar showed "No sessions yet" until the user's first message. NousResearch#72424 added sessions.last_activity_at, which list_sessions_rich now selects — but column adds only land through _reconcile_columns() in the writable _init_schema, and read-only opens skip that by design. Every sidebar read path opens state.db read-only, so each poll raised "no such column: s.last_activity_at" until the first prompt's lazy session-row persist forced a writable open and reconciled. A heal for exactly this class already existed (_open_session_db_for_profile probes the read-only handle and does a one-time writable reopen on staleness), but its probe was a hand-written four-column list that never learned last_activity_at — it went stale three days after shipping. And the batched sidebar route (/api/profiles/sessions/sidebar) bypassed the helper entirely, swallowing per-profile failures into an errors array the desktop never surfaces, so the incident produced an empty sidebar with clean logs. The fix removes the maintenance burden instead of paying it once more: - hermes_state_schema.schema_read_probe_statements() derives one `SELECT <every declared column> FROM <table> LIMIT 0` per table from SCHEMA_SQL via the existing _parse_schema_columns() — the same source of truth the writable reconciler diffs against, so any future ADD COLUMN is probed with no list to update. Column references are table-qualified: an unqualified double-quoted identifier that fails to resolve silently degrades to a string literal (SQLite's double-quoted-string misfeature) and would make the probe pass on exactly the store it exists to catch. - web_server splits the heal into a path-level _open_session_db_at_path (semantics unchanged) so the cross-profile session routes can share it; both profiles.py loops and _count_status_active_sessions (the remaining raw read-only sibling) now open through it. The heal stays a helper rather than a SessionDB classmethod on purpose: escalation-to-writable must remain an explicit caller decision — update_cmd.py opens read-only mid-update and must never write. - Exhaustion guard: if the writable heal SUCCEEDS and the re-probe still fails (a schema problem ADD COLUMN cannot express), the store is marked exhausted — warn once, skip the probe, serve reads probe-less — instead of re-running the full writable init on every poll against a possibly live DB. A FAILED writable open (transient lock) is deliberately not recorded, so the next poll retries the heal. - The per-profile swallow sites in profiles.py now also log a deduplicated warning, so a persistent read failure is loud in errors.log even though the response errors array stays invisible to the sidebar. Tests: probe/SCHEMA_SQL coverage invariants (tests/test_schema_read_probe.py), last_activity_at added to the /api/sessions heal parametrize, a sidebar-route heal test reproducing the shipped symptom (errors == [] and the session returned against a store missing the column), and an exhaustion test pinning exactly one writable open. The sidebar and last_activity_at tests fail on main.
…tale-schema-probe fix(desktop): derive the stale-schema read probe from SCHEMA_SQL
…ng standards The in-repo skill-authoring skill taught the validator's ceilings (1024-char descriptions, 'Use when ...' phrasing) instead of the repo's review standards, so agents following it produced skills that fail review: 240+ char descriptions, author 'Hermes Agent' with no human credit, no bundled-vs- optional decision, dangling related_skills, no platforms audit, no tests, no docs regen, and machine-local /home/bb/... paths baked into prose. Rewritten to teach the hardline standards from AGENTS.md: - description <= 60 chars, one sentence, ends with period - author credits the human contributor first - bundled vs optional tier decision (5+ sessions/month bar; default optional) - no router/index/hub skills - platforms: audited against actual scripts, POSIX-signal table - related_skills must resolve in-repo - Hermes-tool framing instead of raw shell prose - tests at tests/skills/ + docs regen with scope discipline - removed machine-local paths; validator limits marked as NOT the standard
…pora Inspired by virgiliojr94/book-to-skill (MIT): /learn now picks the skill shape by the source. Workflows and small sources still get one tight SKILL.md; books, paper stacks, specs, and large doc corpora get a knowledge-base layout — a lean always-loaded SKILL.md index plus one distilled file per chapter/topic under references/, loaded on demand via skill_view so query cost stays proportional to the answer. - agent/learn_prompt.py: new _KNOWLEDGE_SKILL_STANDARDS block (index + per-chapter references/, structure-not-summary distillation, never reproduce source passages, fold-in instead of duplicating) and a _SOURCE_HYGIENE block pinning extracted source text as data and dropping invisible/bidi Unicode (Trojan Source class). Clarified that the ~200-line cap and hub-skill ban apply to SKILL.md itself, not a knowledge skill's own references/ files. - tests: contracts for the knowledge-base layout, the three embedded standards blocks, and the source-hygiene coverage. - docs: skills.md documents the knowledge-base shape.
A manual cronjob run executed the job synchronously on the calling agent's tool thread. A cron job is a full agent run that routinely takes minutes to hours, so the parent turn sat inside ONE tool call the whole time: uninterruptible (the interrupt flag is only checked between loop iterations) and serial (a batch of manual runs executed one by one). A Telegram session that kicked off dozens of new jobs 'right now' was wedged for hours ignoring every interrupt. action='run' now rides the async-delegation rail delegate_task background mode uses: the at-most-once claim is taken synchronously (so paused/missing/already-firing jobs still report immediately), the run executes on the shared daemon executor, the tool returns at once with a delegation handle, and the job's outcome re-enters the conversation as a type='async_delegation' completion event through the existing completion-queue drains (CLI + gateway) — preserving message-role alternation and the prompt cache. Sync fallbacks preserved: - no routable session (direct Python callers, hermes cron run) - async delivery unsupported (hermes -z, cron child sessions, Kanban workers, stateless HTTP) - dispatch pool at capacity (claim already taken — runs inline rather than stranding it) The completion block reports ok/failure, delivery target, next scheduled run, and an excerpt of the job's saved output.
Salvaged from PR NousResearch#53395 by @izumi0uu: the fire claim's 300s TTL is routinely outlived by real cron jobs, so claim_job_for_fire alone cannot stop a manual cronjob(action='run') from double-firing a job the ticker (or another manual run) is still executing. Extract the ticker's _submit_with_guard running-set check into shared module-level helpers (try_register_running_job / release_running_job) and register manual runs through the same set — one dedupe owner, no drift. Manual runs also become visible to get_running_job_ids (the gateway shutdown drain, NousResearch#60432) and mark_running_jobs_interrupted, which previously could not see them. The background dispatch path pre-checks the running set so a mid-run job reports 'already running' in the tool response immediately instead of as a delayed error completion event; the authoritative atomic check remains in _run_claimed_job on the worker. Co-authored-by: izumi0uu <izumi0uu@gmail.com>
Pins the scheduler-boundary contract: extra_prompt is appended under '## Run Context', does not mutate job['prompt'], and the header is absent when extra_prompt is omitted. Addresses review feedback from harjothkhara on PR NousResearch#57342.
…esearch#57331) Salvaged from PR NousResearch#57342 by @liuhao1024 (with the injection-scan half from PR NousResearch#57360 by @ghedeselmabot): cronjob(action='run', prompt=...) silently discarded the prompt argument — per-run context never reached the spawned cron session. The prompt is now threaded as extra_prompt through the whole chain (cronjob run action → _try_dispatch_background_run/_execute_job_now → _run_claimed_job → run_one_job → run_job → _build_job_prompt) and appended to the stored prompt under a '## Run Context' header for that single fire only — never persisted to the job definition. It passes the same strict _scan_cron_prompt injection scan as stored prompts before firing, and works identically on the background and sync fallback paths. Test fakes across tests/cron/ updated to accept the new kwargs (sibling-test blast radius from the signature change). Co-authored-by: liuhao1024 <liuhao1024@users.noreply.github.com>
… standalone send Fixes NousResearch#61495 When manually triggering cron jobs from a live Matrix session, delivery would fail with "Timeout context manager should be used inside a task" because the aiohttp.ClientTimeout context manager requires a proper asyncio task context. Use asyncio.wait_for() instead of aiohttp.ClientTimeout to avoid this error, following the same pattern as the Weixin platform (gateway/platforms/weixin.py). Changes: - Remove aiohttp.ClientTimeout(total=30) from ClientSession constructor - Wrap the send operation in a nested async function (_do_send) - Use asyncio.wait_for(_do_send(), timeout=30) for timeout handling - Catch asyncio.TimeoutError explicitly and return clear error message
Covers the behavior shipped in NousResearch#80807 (background dispatch for cronjob action='run') and NousResearch#80838 (per-run '## Run Context' prompt, gateway-loop delivery): immediate return with handle, completion re-entering the conversation, in-flight dedupe, transient context injection with prompt scanning, and the sync fallbacks.
POWERFULMOVES
merged commit Aug 10, 2026
40765f7
into
PMOVES.AI-Edition-Hardened
50 of 52 checks passed
2 tasks
POWERFULMOVES
added a commit
that referenced
this pull request
Sep 23, 2026
#15) * fix(agent): keep _serialize_for_summary byte-identical to main The record split added .rstrip("\n") to every serialized record. That changes the bytes _serialize_for_summary returns for legacy tail_mode and for agent/micro_compaction.py's _micro_summarize_one, which never sample and never needed it: a message with trailing newlines was no longer reproduced verbatim in the summarizer prompt. Sampling receives records structurally, so an empty pseudo-record can no longer become the tail anchor and the strip has no remaining purpose. * refactor(agent): sampler takes serialized records only _sample_summary_input accepted either a flat string or a record sequence and, for strings, rebuilt records with content.split("\n\n"). The PR's own third commit established that "\n\n" is not a record delimiter (message bodies contain blank lines), so that fallback would re-introduce split records for any caller that reached it. The only production caller already passes records to _sample_summary_records; drop the dual-signature wrapper and move the test callers to records. * test: keep two behavioural lean-sampling tests Replace the eight sampler tests with two invariants. The boundary test now drives _generate_summary so it fails on main for the actual defect (a retained section starting mid-record) instead of on a helper name, and folds in the multi-paragraph case since blank lines inside a message were what broke the string-split approach. Dropped: covers_multiple_ regions and generate_summary_lean_samples_structural_records (green on main, no teeth), omission_markers_account_separators_exactly (asserted only that a marker exists), newest_message_ending_in_blank_lines (moot once records are structural), and preserves_oversized_last_record (a record above the whole cap is unreachable through _serialize_for_summary because _CONTENT_MAX bounds message bodies). The remaining oversized test asserts behaviour — tail kept, oversized record bounded — not the marker wording. * feat(agent): record summary-input coverage telemetry (from #118390) Lean sampling is the documented budget mechanism, but nothing reported how much of a region actually reached the summarizer, which is the gap the reporter narrowed #118362 to. _sample_summary_records now returns record-level coverage counters alongside the bounded transcript and _generate_summary writes them into the active compression telemetry (summary_input_chars / sampled_chars / omitted_chars / record_count / sampled_record_count / elided_record_count), so the existing compression_attempt JSON log line carries them with no transcript content. Char counters cover record content only, not separators or markers. Elision markers also name the one-based record range they stand for, so a reader can tell how many turns a gap hides, not only how many characters. Co-authored-by: KoNit-K <konit.block@protonmail.com> * refactor(agent): drop unreachable re-trim in _bound_oversized_record Delete the `if len(res) > limit` re-trim at the end of `_bound_oversized_record` (agent/context_compressor.py:3488-3489). Why it can never fire: `marker_reserve` is the marker width formatted with `elided=len(record)`, the largest value `elided` can take, so the real marker is <= marker_reserve; `head` is `record[:head_len].rstrip` (<= head_len) and `tail` is `record[-tail_len:].lstrip` (<= tail_len); head_len + tail_len == limit - marker_reserve. Hence len(head + marker + tail) <= limit always. The branch was untested and unreachable (gate 2c W2). * refactor(agent): remove unreachable overflow-trim loop in _sample_summary_records Delete the post-render overflow-trim loop (agent/context_compressor.py :3591-3607) and state the invariant in its place. Why it is unreachable: every display record is pre-bounded to `target` via _bound_oversized_record; each slice's greedy fill stops before `size + sep + next > target`, so a slice renders to <= target chars and the n slices to <= n*target (a merge of overlapping/adjacent slices is their union: overlap only removes chars, adjacency adds one 2-char separator but drops one marker of >= 2 chars). There are at most n-1 markers, each <= marker_len because marker_len is formatted with first=last=len(records) and elided=total_len, the maxima. Since target = (cap - (n-1)*marker_len) // n, the rendered length is <= n*target + (n-1)*marker_len <= cap. Gate mutation M5 (loop disabled) left all 7 sampler tests green; 30 random shapes never entered it. * fix(agent): spend leftover lean-sampling budget on whole neighbouring records The whole-record greedy fill in `_sample_summary_records` (agent/context_compressor.py:3541,3555-3563) stops a slice as soon as the next record would cross `target` (~20K), so each slice can be short by up to one record. With mid-sized records that wastes a large share of the 160K cap: probe worst case (10K records) filled 90,660/160,000 (57%), 8K records 129,021 (81%); main filled 159,988 by splitting records. Add a bounded extension pass after `selected` is built: newest slice grows backward, older slices grow forward, one whole record at a time, only while the exact rendered length stays <= cap and the slice does not run into its neighbour (adjacent slices are merged so no separator is lost). Records are never split and the cap is never exceeded. After: worst random shape 150,581 (94%, remaining headroom < one 10K record), 8K records 153,117 (96%), 1551-record region 159,996 with 968 whole records, 0 partial, 8 buckets within +-10%, telemetry input_chars == sampled_chars + omitted_chars (probe: $D/fold/probe_C_fold.py). The existing keeps_record_boundaries fixture (1.2K records) fills 99.1% without this pass, so no fill assert was added there — it would be toothless. * refactor(agent): pre-declare summary_input_* telemetry keys `_begin_compression_telemetry` (agent/context_compressor.py:1947-1960) did not seed the six `summary_input_*` counters that `_record_summary_input_coverage` adds on the lean path, so the emitted `compression_attempt` line (conversation_compression.py:1451-1470) had a different key set for lean and legacy attempts. Seed them as None so the schema is stable; probe: lean and legacy attempts now emit identical key sets (gate 2c suggestion, L1947-1960). * docs(agent): say sampled_chars counts display chars in _sample_summary_records The coverage docstring (agent/context_compressor.py:3510-3511) said the counters count "record content only", but `sampled_chars` sums the *display* records (post `_bound_oversized_record` truncation) while `input_chars` sums the raw records. Spell that out so telemetry consumers do not compute `omitted` two different ways (gate 2c suggestion, L3512-3513/3524-3525). * test: assert the newest record token directly in lean sampling test tests/agent/test_lean_single_aux_call.py:163 accepted either `<seg09>` or a `content[-500:]` substring. The last slice is anchored to the newest record, so the fallback can never be what passes; assert `"<seg09>" in out` outright and drop the now-unused `content`. Red when the tail anchor is defeated (end = len(records) - 1): 1 failed; green at head (gate 2c suggestion). * fix(agent): hand out lean-sampling headroom round-robin across slices The budget-extension pass in _sample_summary_records grew the newest slice until it hit the cap and only then moved to older slices, so on mid-size records all headroom went to one region: 10K x 100 records -> records/slice [1,1,1,1,1,1,1,8] (newest 4.3x the mean). That contradicts the design comment above (regions must not consume each other's budget) and the docs' "evenly sampled". Wrap the per-slice loop in a `while grew` round so each slice adds at most ONE whole record per round (newest grows backward, others forward); the cap pre-check and exact re-render check are unchanged. Before -> after (records per slice, fill unchanged): 10000x100 [1,1,1,1,1,1,1,8] -> [1,2,2,2,2,2,2,2] fill 0.944 8000x100 [2,2,2,2,2,2,2,5] -> [2,2,2,2,2,3,3,3] fill 0.957 4000x300 [4,4,4,4,4,4,4,11] -> [4,5,5,5,5,5,5,5] fill 0.987 2000x600 [9,9,9,9,9,9,9,15] -> [9,9,10,10,10,10,10,10] fill 0.995 500x2000 [37,...,37,39] -> [37,...,37,38,38] fill 0.998 Re-gate probe (40 random shapes): 0 violations, worst fill 0.892 -> 0.930, worst char drift 4.27 -> 1.83; under-cap output byte-identical to base. Test: test_lean_sampling_oversized_middle_record_does_not_evict_tail now asserts no slice holds more than mean+1 records (RED with the previous newest-first order: [16,16,16,16,16,16,18]; GREEN: [16,16,16,16,16,17,17]). Finding: $D/regate/C.md W1 (agent/context_compressor.py:3599-3623). * docs(agent): say why the extension pass re-renders after the cap pre-check The comment claimed "marker widths shrink as records leave a gap" — but shrinkage can only make the render smaller, which the pre-check already tolerates. The exact re-render exists for the opposite case: a gap's first index can gain a digit or thousands separator (999 -> 1,000, +2 chars) when the added record moves a marker, and that growth can push a render sitting at cap over it. Comment only; no behaviour change. Finding: $D/regate/C.md suggestion 1 (agent/context_compressor.py:3619). * test(agent): cover slices merged when the extension pass closes a gap No fixture exercised the final `selected = _merged(selected)` in _sample_summary_records: the re-gate mutant that dropped it survived all 8 bounding tests and 40 random probe shapes. Add 20 x 8.5K records (170K > 160K cap, headroom > one 8.5K gap) so the round-robin extension closes the gaps between the newer slices. Assert consecutive shown records are joined by exactly "\n\n" (an unmerged adjacent pair renders as "...xxx[USER]: record-..." with no separator), remaining gaps carry a correctly numbered marker, fewer than 7 markers survive, and sampled_record_count equals the whole records shown. RED with the final `_merged` call removed: AssertionError (14, 15); GREEN at head. Finding: $D/regate/C.md suggestion 2 (M-c, agent/context_compressor.py:3623). * refactor(agent): merge lean-sampling slices once via _merged instead of inline The slice-build loop in _sample_summary_records hand-rolled the same "overlapping/touching slice -> extend previous" fold that the `_merged` closure 25 lines below implements (agent/context_compressor.py:3563-3568 vs :3590-3597). Hoist `_merged` above the loop, append raw (start, end) pairs, and merge once before the extension pass; the first `_render` no longer wraps `selected` in a redundant `_merged` since it is already merged. One interval-merge implementation instead of two. Output is byte-identical: the merge is applied once to the same ordered slice list, so the extension pass starts from the same `selected`. Verified with a 47-shape capture (40 probe_C2 shapes + 7 edge shapes) diffed before/after: identical sampled text and coverage; probe_C2 40 shapes -> 0 violations; under-cap base identity x4 True. * refactor(agent): drop unreachable trailing-gap branch from lean-sampling _render The `if cursor < len(records)` block at the end of `_render` (agent/context_compressor.py:3580-3583) duplicated the gap-marker arithmetic of the in-loop branch but had no path to it. Invariant: `_render` is only reached with non-empty `records` (empty returns early), and n >= 1, so the slice-build loop always runs its last iteration, which anchors `end = len(records)` and `start = end - 1 >= 0`, hence `end > start` and the final slice ends at `len(records)`. `_merged` keeps `max(end)` for the last interval, and the extension pass only grows the newest slice backward (`(s - 1, e)`) or older slices forward but never past the next slice's start, so the last slice's end stays `len(records)`. Therefore `cursor == len(records)` after the loop for every input. Byte-identical across the 47-shape capture (40 probe_C2 shapes + 7 edge shapes incl. single-record and empty); probe_C2 -> 0 violations; under-cap base identity x4 True. * refactor(agent): integer head/tail split in _bound_oversized_record `_bound_oversized_record` (agent/context_compressor.py:3485-3488) split `remaining` with float arithmetic (`int(remaining * 0.5)`) and guarded the tail slice with `if tail_len else ""`. `remaining = limit - marker_reserve` is >= 1 because the line above already returned when `limit <= marker_reserve`, so `head_len = remaining // 2 <= remaining - 1` and `tail_len = remaining - head_len >= 1` always: the guard can never fire. Use `remaining // 2` and drop the dead guard. Byte-identical: `int(r * 0.5) == r // 2` holds for every non-negative int (checked exhaustively for r in 0..10**6), and the guard was dead. Capture of 800 bounded-record cases (4 record shapes x remaining 1..200) plus the 47-shape sampling capture is identical before/after. * test(agent): derive the lean-sampling marker bound from _SAMPLED_INPUT_SLICES tests/agent/test_context_compressor.py:3397 asserted `sampled.count("chars elided") < 7`, hardcoding the n-1 gap count that the 8-slice constant implies and that the comment explained in prose. Use `ContextCompressor._SAMPLED_INPUT_SLICES - 1` so the invariant ("the extension pass closed at least one initial gap") survives a change to the slice count instead of silently becoming a change detector. `_SAMPLED_INPUT_SLICES - 1 == 7` today, so the assertion value is unchanged. * refactor(agent): inline the single-use _finish closure in _sample_summary_records `_finish` (agent/context_compressor.py:3586-3588) only wrapped `_coverage` and re-derived `shown` from the final `selected`; it was called once, at the return. With `_merged` hoisted above the slice loop the classmethod now has two closures (`_merged`, `_render`) plus `_coverage` instead of four. Same expressions, same evaluation order (`_render` first, then the coverage counters), so the output and coverage dict are byte-identical — verified on the 47-shape capture and probe_C2 (0 violations, base identity x4 True). * fix(stt): isolate CAF conversions from source recordings (cherry picked from commit 5a87336db193740ebf4911262370a5ce0dbb2b4e) * test: fold CAF isolation test into TestCafConversion, trim to 2 cases The invariant (source .caf untouched, sibling .wav not overwritten, owned output removed) belongs beside the existing CAF tests rather than in a new file. The conversion-error case exercised the same ExitStack return-inside-with path as provider-error, so it is dropped to stay at two cases. Also reverts the unrelated encoding="utf-8" drive-by on the .env fixture in TestTranscribeCredentialReadGuard. * fix(stt): tolerate CAF work-dir cleanup errors Pass ignore_cleanup_errors=True to the TemporaryDirectory that holds the converted CAF->WAV file (tools/transcription_tools.py:420). WHY: the sibling trim-dir cleanup a few lines below already swallows errors (shutil.rmtree(..., ignore_errors=True)), but the CAF work dir did not. On Windows an AV scanner or the search indexer briefly holding the freshly written WAV makes rmtree raise out of TemporaryDirectory.__exit__, turning an otherwise successful transcription into an exception. The kwarg is available since 3.10; pyproject requires-python >= 3.11, and plugins/teams_pipeline/meetings.py:199 already uses it. No new test: the behaviour is stdlib. * test(stt): drop dead shutil.which patch in CAF neighbor test Remove `monkeypatch.setattr(audio.shutil, "which", lambda _name: None)` from test_caf_conversion_preserves_neighbors_and_removes_owned_output (tests/tools/test_transcription_tools.py:1360). Why it is dead: `_convert_caf_to_wav` (tools/transcription_audio.py:168-177) tries candidates in order ffmpeg-then-afconvert and returns on the first success. The test already patches `_find_ffmpeg_binary -> "ffmpeg"` and `_run_quiet -> encode`, which never raises, so the ffmpeg candidate always succeeds and the afconvert candidate is never executed. Whether `shutil.which("afconvert")` returns a path or None cannot change the outcome. The patch also mutated the process-global `shutil.which` for the test's duration for no benefit. Proof: `-k Caf` selection is 4 passed before and after. * fix(skills): preserve rollback blobs when ledger reads fail (cherry picked from commit e88e7cbbacbc9f00b0260c1c42b05086a7414d45) * fix(skills): warn when blob GC is skipped because the ledger is unreadable The new read-failure branch in gc_blobs() returned (0, 0) silently, so `hermes curator ledger --compact` printed "0 unreferenced blob(s) removed" as if the sweep had run. Log a warning with the underlying error so an operator can tell "nothing to reap" from "could not look". The kept invariant test now also asserts the warning via caplog. * fix(skills): guard compact_ledger() against undecodable ledger bytes compact_ledger() caught OSError on the read but decoded the bytes outside the guard, so a ledger with invalid UTF-8 raised UnicodeDecodeError out of `hermes curator ledger --compact`. Same escape class the gc_blobs() fix closes; treat it identically (no-op, ledger left untouched). * test: drop incidental encoding= edits from an unrelated ledger test Three read_text() calls in test_entry_stores_only_changed_paths_and_rollback_still_restores gained encoding="utf-8" alongside the gc_blobs fix. They are unrelated Windows hygiene (the footgun checker does not flag them); keep the PR scoped to the blob-GC read-failure bug. * fix(skills): warn when blob GC is skipped over a malformed ledger line gc_blobs() aborts the sweep on a json.JSONDecodeError (tools/skill_ledger.py:351-352) but did so silently, while the sibling read-failure abort a few lines above logs "ledger unreadable (...); blob GC skipped". `hermes curator ledger --compact` therefore printed "0 blobs removed" as if the sweep had run and found nothing. Emit the same warning family so the operator learns the ledger needs repair before blobs can be GC'd. Test: `malformed-json` param on the kept invariant test — (0, 0), blobs intact, and "blob GC skipped" in caplog. RED with the warning removed (assert 'blob GC skipped' in ''), GREEN at head. * fix(skills): abort blob GC on a non-dict ledger row instead of raising After json.loads() succeeds, gc_blobs() called row.get(...) unconditionally (tools/skill_ledger.py:353). A syntactically valid but non-object line such as `[]` or `"x"` raised AttributeError out of gc_blobs() and up through `hermes curator ledger --compact`. list_entries() already tolerates such rows with an isinstance(row, dict) check (~L414); mirror it here and abort the sweep with the same "malformed ledger line; blob GC skipped" warning — a row we cannot interpret may still hold blob references, so nothing may be deleted. Test: `non-dict-row` param on the kept invariant test (ledger + b"[]\n"). RED with the guard removed (AttributeError: 'list' object has no attribute 'get'), GREEN at head. * fix(skills): warn when compaction is skipped because the ledger is unreadable compact_ledger() returns (0, 0, 0) when the ledger cannot be read or decoded (tools/skill_ledger.py:308-309) but did so silently, so `hermes curator ledger --compact` reported a no-op compaction as if the ledger were simply empty. gc_blobs() already logs "ledger unreadable (...); blob GC skipped" for the identical condition; emit the symmetric warning here so the operator sees the ledger needs attention. Manual probe (HERMES_HOME=<tmp>, ledger bytes b"\xff"): WARNING skill_ledger: ledger unreadable ('utf-8' codec can't decode byte 0xff in position 0: invalid start byte); compaction skipped compact_ledger() -> (0, 0, 0); ledger bytes still b"\xff". The regression test that pins this arrives in the following commit. * test(skills): pin compact_ledger() no-op on an undecodable ledger Gate mutation M2 (moving `raw.decode("utf-8")` outside the try in compact_ledger(), tools/skill_ledger.py:306-307) survived the existing suite: nothing exercised compaction over ledger bytes that are not UTF-8. Add the 4-line regression: with the ledger set to b"\xff", compact_ledger() returns (0, 0, 0), the bytes are left exactly as they were, and the "compaction skipped" warning from the previous commit is emitted. RED with the decode moved outside the try (UnicodeDecodeError escapes), GREEN at head. * fix(skills): treat an undecodable ledger as empty in list_entries() list_entries() caught only OSError (tools/skill_ledger.py:409), so a ledger whose bytes are not valid UTF-8 raised UnicodeDecodeError out of `hermes curator ledger`, get_entry() and rollback_entry() — the very corrupt state gc_blobs() and compact_ledger() now tolerate on this branch. Widen to (OSError, UnicodeError) and return [], matching the existing "unreadable ledger == empty" contract; the read path is read-only, so nothing is lost by treating the file as having no entries. Manual probe (HERMES_HOME=<tmp>, ledger bytes b"\xff"): before: list_entries() -> UnicodeDecodeError: 'utf-8' codec can't decode byte 0xff after: list_entries() -> []; get_entry('abc') -> None; rollback_entry('abc') -> (False, "no ledger entry with id 'abc'") * test(skills): make the unreadable-ledger fixture transport-agnostic The `unreadable` param of test_gc_keeps_rollback_blobs_when_the_ledger_cannot_be_read (tests/tools/test_skill_ledger_delta.py:122-130) monkeypatched Path.read_text class-wide to raise PermissionError for the ledger path. That couples the test to one read primitive: switching gc_blobs() to read_bytes()/open() would make the param pass vacuously without testing anything. Replace it with a real filesystem condition — a directory at the ledger path — which fails read_text() with an OSError on every platform (IsADirectoryError on POSIX, PermissionError on Windows) regardless of how the file is read. The monkeypatch fixture and the with-block indentation go away; the directory is removed before the ledger bytes are restored for the rollback check. RED with the guard defeated (`return 0, 0` -> `lines = []`): [unreadable] assert (2, 14) == (0, 0) (blobs were deleted); GREEN at head. * docs(skills): document that an unreadable ledger also aborts blob GC The gc_blobs() docstring (tools/skill_ledger.py:334) only mentioned malformed lines as a reason to abort the sweep; this branch added the unreadable/undecodable-ledger abort without updating it. State both conditions and the invariant (blobs are kept). * fix(skills): warn when list_entries() hits a corrupt ledger, and pin the UnicodeError guard list_entries() (tools/skill_ledger.py:414) caught (OSError, UnicodeError) but returned [] silently, so a ledger that EXISTS but cannot be decoded made `hermes curator ledger` print "ledger is empty" and `hermes curator rollback` report "no ledger entry with id" with no hint that the file is damaged — while gc_blobs()/compact_ledger() already warn on the same condition (:309, :345). Now warn for any read failure except FileNotFoundError (a missing ledger is the normal empty state). `%s`-lazy logger call, same "ledger unreadable" prefix. Also adds the regression test the UnicodeError widening (2213062284) lacked: re-gate mutation M8 reverted the except clause to `except OSError:` and the whole suite stayed green. New tests write b"\xff" to ledger_path() and assert list_entries() == [], get_entry() is None and "listing empty" in caplog; a second test pins that a merely missing ledger stays silent. Red/green: `except OSError:` -> UnicodeDecodeError (1 red); warning dropped -> caplog assert red; warning made unconditional -> missing-ledger test red. * refactor(skills): share one _read_ledger() helper across compact/gc/list compact_ledger, gc_blobs and list_entries each stamped the same prelude: read the ledger, catch (OSError, UnicodeError), warn "skill_ledger: ledger unreadable (%s); <what>", bail (tools/skill_ledger.py:305-310, :341-346, :412-417 — simplify reuse #1). Fold them into a private sibling _read_ledger(what, *, quiet_missing=False) -> Optional[bytes] that reads bytes, validates the UTF-8 decode and warns once unless quiet_missing and the ledger is merely absent (list_entries: a fresh install has no ledger). Returns bytes rather than str so compact_ledger keeps reporting the on-disk size in bytes_before instead of re-encoding. Warning texts ("compaction skipped", "blob GC skipped", "listing empty") are unchanged, so the existing invariant tests pass verbatim. * refactor(skills): abort blob GC from a single malformed-row guard gc_blobs carried two byte-identical `warning("malformed ledger line; blob GC skipped"); return 0, 0` blocks — one under `except json.JSONDecodeError`, one under `if not isinstance(row, dict)` (tools/skill_ledger.py:350-357; simplify reuse #2, efficiency note, re-gate G2 S2). Funnel the decode failure into the type guard (`row = None`) so the abort exists once. Behaviour is unchanged for both the [malformed-json] and [non-dict-row] parametrizations: any line that is not a JSON object still aborts the sweep with the same warning. * test(skills): assert the missing-ledger case emits no warning records test_list_entries_is_silent_when_the_ledger_is_merely_missing checked `"listing empty" not in caplog.text` (tests/tools/test_skill_ledger_delta.py:40-44; simplify quality #1) — a negative substring match that goes vacuously green if the message is ever reworded, and stays green if a *different* warning fires. Assert on caplog.records at WARNING or above instead, which checks the actual invariant (no warning for a merely absent ledger) independent of wording. * docs(skills): explain why an undecodable ledger warns, without review jargon The docstring of test_list_entries_treats_an_undecodable_ledger_as_empty_and_warns cited "(re-gate M8/S1)" (tests/tools/test_skill_ledger_delta.py:88; re-gate G2 S1) — a review-process artifact meaningless to future readers. State the behavioural reason instead: silence would make a corrupt ledger look like a fresh install with no history. * fix(kanban): bound gc retention days so -N/0 cannot mass-delete hermes kanban gc (and /kanban gc from a chat session) passed the raw --event-retention-days/--log-retention-days values straight into gc_events/gc_worker_logs as seconds. A negative value builds a future cutoff that matches every terminal-task event row and every worker log; 0 deletes everything older than now. Neither is a sane retention. - _cmd_gc errors (rc 2) on days < 0 and treats 0 as disabling that sweep, matching the "0 disables" convention done_sub_retention_days documents and purge_stale_done_notify_subs implements. - gc_events/gc_worker_logs raise ValueError on older_than_seconds < 0 at the library boundary, the same guard shape _prune_where uses. (cherry picked from commit a9c92da1005c4dc77f8fd4057bcf6e7d99e3241f) * kanban gc: validate retention flags before the workspace sweep (cherry picked from commit f7e6505fff08955ee0b46c9972002ab1c015e287) * fix(kanban): reject negative retention at the argparse boundary (from #118615) `hermes kanban gc --event-retention-days -1` is now a usage error (rc 2, argparse's own message) before `_cmd_gc` ever runs, instead of relying on the command body to notice. The `_cmd_gc` guard from #118506 stays: slash and programmatic callers build a Namespace by hand and bypass argparse. gc_events/gc_worker_logs now int()-normalise `older_than_seconds` before the `< 0` compare; gc_worker_logs previously fed the raw value into `time.time() - older_than_seconds`, so a str would TypeError deep in the sweep rather than at the guard. Deliberately not taken from #118615: the helper-level `seconds == 0 -> return 0`. "0 disables" is CLI policy and lives in `_cmd_gc`; at the library layer 0 means "everything older than now", which is a valid request. Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com> * test: collapse duplicated kanban gc retention tests into parametrize The -1 / 0 / 30 _cmd_gc cases and the -1 / 0 slash cases each ran the same fixture and assertions with one value swapped; a parametrize keeps every case and every red-on-base assertion while the file reads as six invariants instead of nine near-copies. Also drops the `capsys` fixture that was requested but never read. * refactor(kanban): share the negative-retention ValueError message gc_events and gc_worker_logs raised the same two-line f-string with one noun swapped; one module constant keeps the two sweeps from drifting. * docs(kanban): state gc retention semantics for 0 and negative values website/docs/user-guide/features/kanban.md:987-988 listed the two gc retention flags without saying what the edge values do; add "(negative N is rejected; 0 disables that sweep)" so users don't have to read the code. hermes_cli/kanban_db.py:4279,4293 — gc_events/gc_worker_logs accept older_than_seconds=0 as "everything older than now" while _cmd_gc maps days=0 to "disabled" before calling them. The docstrings did not state that split, so a library caller could assume 0 is a no-op. One line each. Gate findings: D.2ab.md:27, D.2c.md:32. * test(kanban): fix slash gc docstring — -1 stops at the parser type tests/hermes_cli/test_kanban_gc_retention.py (test_slash_kanban_gc_retention_bounds) said "-1" goes through "the same _cmd_gc guard"; via the slash path only the argparse _nonnegative_int type fires for -1, so the guard is never reached. Only 0 reaches _cmd_gc, which treats it as "disabled". Reword; assertions unchanged (docstring accuracy only). Gate finding: D.2c.md:19. * refactor(kanban): drop redundant int() in gc_events cutoff hermes_cli/kanban_db.py:4288 — older_than_seconds is already normalised with int(...) on the line above the negative check, so the second int() in the cutoff expression is a no-op. Sibling gc_worker_logs already subtracts the bare value (float cutoff for st_mtime), so nothing to change there. No behaviour change. Gate finding: spec-D fold item 4. * refactor(kanban): share one _retention_seconds helper across both gc sweeps gc_events and gc_worker_logs each carried the same three lines (int() coerce, `< 0` check, raise ValueError(_NEGATIVE_RETENTION_MSG.format(...))); only the message string was hoisted. Replace the constant with a sibling-local helper `_retention_seconds(older_than_seconds) -> int` that owns coerce + check + raise, and call it from both sweeps. WHY: the comment explaining the rule ("a negative window puts the cutoff in the future, so 'older than cutoff' matches everything") is the reason for the check, so it belongs on the check rather than on a format string; and a third sweep now has one place to reuse instead of a block to paste. Message text is unchanged, so the existing `pytest.raises(ValueError, match="older_than_seconds")` tests pass as-is. Finding: simplify/D.reuse.md #1 + simplify/D.quality.md #1 (hermes_cli/kanban_db.py:4285-4287, :4303-4305). Proof: mutating the helper's `< 0` to `< -10**12` fails test_gc_events_rejects_negative_window and test_gc_worker_logs_rejects_negative_window (DID NOT RAISE); head green. * refactor(kanban): narrow _nonnegative_int except to ValueError argparse's _get_value always calls a `type=` callable with the raw str token, and int(<str>) can only raise ValueError, so the TypeError arm in `except (TypeError, ValueError)` is unreachable. The try itself stays: without it argparse would print "invalid _nonnegative_int value: 'thirty'", leaking the private helper name into the usage error. Invariant: `_nonnegative_int` is only referenced as an argparse `type=` (hermes_cli/kanban_parser.py) and by the unit test, which passes str. Finding: simplify/D.quality.md #2 (hermes_cli/kanban_parser.py:50). Dead-branch deletion; existing test_gc_parser_rejects_negative_retention_days still covers the "thirty" -> "must be an integer" path (green). * test(kanban): fold the workspace-survival check into the gc bounds parametrize test_cmd_gc_negative_days_leaves_workspaces_untouched and test_cmd_gc_retention_bounds[negative-refuses] killed the same mutant (the _cmd_gc guard turned into `if False`) and differed only in which fixture they asserted survived. Create the archived scratch workspace in the parametrized body (small helper) and assert on it alongside the event/log checks; drop the standalone test. WHY: one invariant, one test. The ordering claim ("refuse before ANY sweep" — the unconditional workspace sweep runs first in _cmd_gc) is still proven by the negative case; the 0/positive cases now also confirm that valid values let the workspace sweep run. The file goes from 6 to 5 test functions with no lost mutation coverage. Finding: simplify/D.quality.md #3 (tests/hermes_cli/test_kanban_gc_retention.py:82-97). Proof: with the _cmd_gc guard mutated to `if False`, test_cmd_gc_retention_bounds[negative-refuses] fails (1 failed, 7 passed); head is green (8 passed). * fix(plugins): per-plugin load deadline so a hung register() no longer hangs startup A plugin whose import or register() never returns (an infinite loop, a blocking network call) held PluginManager.discover_and_load() forever, and with it every synchronous caller: `hermes chat`, gateway startup, ACP session/new (#108139). Each plugin's import + register() now runs under `plugins.load_timeout_seconds` (default 10, 0 disables, max 600) on a daemon worker. On overrun the plugin is recorded as failed with "load timed out after Ns" (same channel as every other load failure: startup WARNING, `/plugins`, `list_plugins()`), its pre-hang registrations are disposed, and discovery continues with the next plugin. The abandoned worker's later `ctx.register_*`/`subscribe`/`on_unload` calls are refused with a WARNING (the context is marked abandoned), so a late registration can never land in a registry the failure path already swept. Abandoned loaders are capped per process (8); past the cap further loads are refused with a named reason rather than run inline, which would recreate the hang (#98382 shape). Because the worker cannot own the caller's RLocks: the deferred-platform eager fallback now runs outside the replacement transaction, discovery re-entered from a loader worker returns on the already-set discovered flag instead of blocking on the sweep's lock, and such a worker never joins the background discovery thread that is waiting on it. * feat(platform): hermes_platform.host with machine facts, runtime predicates, SoC recognizer Host facts and runtime predicates need one stdlib-only source that stays correct under architecture emulation. Hardware recognition reads host facts, not environment variables or GPU proxies. The package boundary gives later resolvers and catalog checks a shared foundation. Linear NS-920. * refactor(cua): WSL detection through hermes_platform.host.runtime.is_wsl One runtime predicate keeps computer-use overlay policy from holding a separate WSL result. The shared result is cached per process. * test: repoint WSL/container probe patches to hermes_platform.host.runtime (group B) The probe cache now belongs to hermes_platform.host.runtime. Call-site name patches stay where production resolves them. * test(platform): host facts, products, and import hygiene Contracts for native-architecture detection, hardware recognition, and the package's stdlib-only boundary. * test: repoint WSL/container probe patches to hermes_platform.host.runtime (group A) The probe cache now belongs to hermes_platform.host.runtime. Call-site name patches stay where production resolves them. * refactor(voice): WSL detection through hermes_platform.host.runtime.is_wsl One runtime predicate keeps WSL detection from diverging between call sites. The shared result is cached per process. * test: resolution ratchet for bare which and known-path tables, fully allowlisted A checked-in baseline blocks new resource lookups outside hermes_platform while existing owners migrate. * refactor(platform): one native_arch for local-runtime assets, desktop PE gate, managed Node The three architecture readers could disagree under emulation. One shared native-architecture result prevents wrong asset selection. * test: known-path ratchet walks nested join() arguments; allowlist moves to tests/fixtures Known-path tables build rows with nested os.path.join() calls, which the first scanner missed. Walking nested arguments covers those tables. The allowlist lives in tests/fixtures/ because tests/data/ is gitignored. * test: managed-Node heal tests pin native_arch where production reads it The production reader ignores PROCESSOR_ARCHITECTURE by design. Patching native_arch() makes these tests independent of the host running them. * test: resolution allowlist follows main (webhook_filters, gateway_service_unit, git_credentials, mcp PATHEXT) Main moved or removed four allowlisted sites and added three. Updating the rows preserves the ratchet's stale-entry check. * docs(agents): machine facts and resource lookup go through hermes_platform Document the shared source for control-host facts and the resolution ratchet so new callers do not create competing implementations. * test: allowlist follows main's gateway wizard split (_setup_signal moved to gateway_setup_wizard.py) main moved _setup_signal out of hermes_cli/gateway.py after this branch was cut; the ratchet's stale-row test fails on the merge result without this rename. * feat(plugins): fire pre/post_auxiliary_call events on every auxiliary LLM call (#79733) Auxiliary LLM calls (titling, compression, MoA advisors/aggregator, vision, approval, ...) never reached any plugin hook: hook-based observability and cost plugins were structurally blind to them. Teknium's ruling on #79733: NEW events rather than reusing the turn-scoped pre/post_api_request pair, so existing subscribers keep their per-turn semantics. - agent/auxiliary_hooks.py (new sibling): builds the pre_api_request / post_api_request payload shape plus `aux_task`, `api_request_id` (`aux-...`, shared by every attempt of one logical call), `retry_count`, `streaming`, parent-turn `session_id`/`task_id`/`turn_id` when a main turn is in flight; fail-open (a raising/hung subscriber is logged and the aux task proceeds); post carries `error`/`error_type` on failure. - agent/auxiliary_client.py: the three relay funnels every physical attempt shares (_relay_sync_completion / _relay_async_completion / _relay_sync_stream) run under the hook pair — retries and fallbacks included. Main-loop *_api_request events do not fire for aux calls. - Catalogue: VALID_HOOKS, bounded-timeout hook set, `hermes hooks test` sample payloads, hooks.md / plugins index / observer-hooks / plugins.md tables, agent + plugins AGENTS.md. - tests/agent/test_auxiliary_hooks.py: 2 invariants (pair fires with aux_task and no api_request events; raising subscriber never breaks the call). First is red on origin/main. Supersedes #32416 (@zrmnelson), #68060 (@JonZal), #77518 (@hsy5571615), #79826 (@webtecnica) — their relay-boundary placement, usage normalisation and fail-open policy shaped this implementation. Co-authored-by: zrmnelson <zacharynelson1@gmail.com> Co-authored-by: Jonas Zalys <jonas@tryholo.ai> Co-authored-by: saitsuki <nukuom976228@gmail.com> Co-authored-by: webtecnica <webtecnica@gmail.com> * fix(aux-hooks): tolerate test seams that stub relay metadata without task keys test_chat_sdk_transform_bypass stubs _relay_auxiliary_metadata with an empty metadata dict; read aux_task/api_mode with .get so the seam keeps working. * feat(connectors): the backend API for the desktop Connectors page; connect an app without a chat session (#115191) * feat(connectors): the backend serves a connector's tool list, cached for 24 hours The Connectors page opens one app and shows every tool it has. The backend had no way to read that list. - `tools/connectors/portal/`: a client for the portal's tool-list route and a JSON cache under the Hermes home, one file per portal origin and connector. An entry is fresh for 24 hours. After that the read revalidates with the stored ETag: 304 keeps the list, 404 deletes the entry, an upstream failure serves the stored list marked stale, and a 401 never serves the cache. - `connectors.tools {slug, refresh}`: account-level, routed by `profile`, no chat session. Errors carry a fixed `reason` from one closed set on the rail. - Every connector model that is not operation state moves into `tui_gateway/contracts/connectors.py`. Handlers that no chat session owns live in `tui_gateway/methods_connectors_account.py`. The wire model is tolerant: an unknown facet reads as unclassified and one odd tool never blanks a connector. * feat(connectors): catalog, accounts and member tool rules by RPC The Connectors page needs the app catalog, the connected account of one app, a way to disconnect it, and the member's own on/off rules. None had an RPC. - `connectors.catalog`: name, description, category and logo of each app. - `connectors.accounts`, `connectors.accounts.remove`: read the accounts at the tool gateway and remove one by id. - `connectors.policy.get`: the rule layers that apply to the member, widest first. The body is a union on `mode`, so a reader can name who turned a tool off. - `connectors.policy.set`: one change, a union on `type` (the tools of one connector, or one connector on or off), with the revision the user saw. A stale revision answers `POLICY_CONFLICT`. The backend composes the upstream write in one pure function, so no renderer learns the upstream rules. - Bundled MCP manifests can name their hosted twin with `connector:`, so the page can show one card per app. * feat(connectors): connect an app without a chat session Every connector RPC took a `session_id`, and a connect that did not come from the model's tool call minted a link with no watcher. The Connectors page has no chat session, and its card must flip to connected by itself. - `connectors.list`, `connectors.connect`, `connectors.operation.status`, `connectors.operation.wake` and `connection.respond` take `owner`, a union on `type`: `session` (today's behaviour and authorization) or `account` (routed by `profile`, authorized by the live transport like `mcp.*`). `session_id` is gone from these params; every desktop caller sends `owner`. - An account connect runs the same operation lifecycle on a background thread, under the profile's scope, so the watcher reads the account and settles the operation. A second connect for an app that is already connecting returns the open operation and mints nothing. - `connection.update` carries `owner`. An account operation has no session to address, so its updates go out on the session-less broadcast path. * feat(mcp-catalog): eighteen more bundled entries name their hosted connector A bundled MCP entry and a hosted connector for the same app are one card on the Connectors page only when the manifest names its hosted twin. Linear and Notion had the field. These entries get it too: airtable, asana, attio, calendly, dropbox, figma, railway, supabase, todoist, betterstack, canva, cloudflare, datadog, intercom, neon, sentry, stripe and vercel. Atlassian maps to two hosted connectors and Prisma Postgres is not clearly the same app, so both stay without one. * refactor(connectors): the account handlers share one gate, one params model and one write table The six account-level handlers each repeated the availability gate, the auth catch and the catch-all reply. One decorator now owns that, and each handler validates its params with its contract model instead of a ladder of isinstance checks. The five connection RPCs share one guard for the unexpected-failure reply. The four write composers for the member rules were the same function with a different list key and polarity. They are one table now. The owner union lives in contracts/common.py, so the params side and the event side stop declaring it twice and the import cycle is gone. An account operation start carries one event and a flag, so the wait for the sign-in link blocks instead of polling every 50 ms. run_operation loses its two account-only parameters; drive_operation is the second entry point. Tests: four deleted (they exercised pydantic or the mock), three merged into tables, two added (a client that still sends the old top-level session_id is refused; all six account RPCs run off the server loop). The shared reply helper and the HTTP and managed-client fakes move to one place each. Comments are one line or gone. * fix(connectors): a missing tool-list route reads as "unavailable", not "connector gone" The tool-list read treated every 404 as the portal's "this connector is not in the catalog" answer. It deleted the cache entry and answered CONNECTOR_NOT_FOUND, so a page would offer to remove an app that is connected and works. A portal that does not serve the route yet answers a bare 404 for every app. Only the portal's own {"error": "connector_not_found"} means the connector is gone. Any other 404 is now a tool-list outage: the cached list is served as stale, or the RPC answers TOOLS_UNAVAILABLE. * fix(connectors): a connect from the page returns to the app after sign-in The sign-in link carries a return target only when the session's surface is the desktop. A chat session binds that surface. An account-owned call has no chat session, so nothing bound it: the link was minted without a return target and the browser ended on the portal's done page instead of coming back to Hermes. Every account-owned call now runs with the process's own surface bound, next to its profile scope. The operation thread copies that context, so the first link and every reissued link carry the return target and the operation id. * test(connectors): defer the new connector RPC coverage The tests for the new account RPCs, the portal client, the tool-list cache and the rule composer leave this PR and come back in one later change, after the API is settled. The same was done for #111008. Kept: the edits that existing tests need because the five connection RPCs now take `owner` instead of `session_id`, and the rename of the managed client seam. Removed: six new test files, their two fakes and the gateway conftest, and the new cases in test_mcp_catalog.py, test_connectors_gateway_client.py, gateway-rpc.test.ts and notifications.test.ts. Reverting this commit restores all of them. * fix(cli): the connection panel hands the tool thread back at once The classic CLI's connection callback waited on a queue for the user's first decision. The operation's watcher starts only after the callback returns, and the watcher is what polls a hosted account, runs the 300-second deadline and sees Ctrl+C. For a hosted connector the panel opens on the sign-in link, where the only key that filled the queue was Cancel. The account was never polled: the user signed in, the panel never changed, and Esc reported the app as skipped. Ctrl+C set the interrupt flag but left the thread parked on the queue, so the turn never ended. The callback now opens the panel and returns, as the gateway's callback does for the desktop and the Ink TUI. The panel's actions already reach the operation through apply_answer on the UI thread, so the queue is removed. An install with a form still waits for Connect, because the backend starts no work for a pending row. Ctrl+C now settles the operation as `interrupt`, and open rows become `not_connected`. Checked on the e2e rig with the fake tool gateway: hosted connect completes on the third status read; Ctrl+C ends the turn and the polling stops; an MCP install with a plain and a secret field still saves config and both values. * fix(connectors): "run it again" lives in the library, so the classic CLI can use it Making a new sign-in link for a failed or expired hosted connector was implemented only in the JSON-RPC layer (`_reissue`). The classic CLI does not go through JSON-RPC: its Connect button on a failed row called apply_answer, which does nothing for a hosted operation because it has no MCP runner. The panel showed "Waiting…" until the deadline. `tools.connectors.run.reissue(operation, names)` now holds the checks and the per-kind action, and returns a refusal reason or None. The gateway maps each reason to the same JSON-RPC error as before. The CLI calls it for a hosted row; a refusal is shown on the row. MCP rows keep their path, because Connect on a failed MCP row re-sends the form values. Checked on the e2e rig: a scripted failed sign-in, then Connect: a second mint with `reinitiate: true`, a new link with a new connection id, then connected. * feat(connectors): the account list and disconnect go through the portal `connectors.accounts` and `connectors.accounts.remove` called the tool gateway. They now call the portal's account-management routes (`GET /api/v1/connectors/accounts`, `DELETE /api/v1/connectors/accounts/{id}`), which apply the organisation membership checks and write the disconnect audit row. There is no fallback to the gateway when the portal is unavailable, and a removal is never retried. The read of ONE account stays on the gateway (`GET v1/connectors/accounts/{id}`): the portal has no such route, and the operation watcher polls it once per second. `ConnectorClient.list_accounts` and `delete_account` are removed. The removed account's reply model carries `connector`, which both services send. * fix(connectors): the account RPCs answer what the portal really sends Checked against the portal source and against the staging and production services. - Errors are read from the upstream error code, not the HTTP status. A rule write answered 409 for a stale revision and for a user with no organisation; both read as "the policy changed". `org_required` is now `ORG_REQUIRED` and 403 `no_access` is `ORG_ACCESS_DENIED` on every account RPC; only a rejected sign-in is `NEEDS_NOUS_AUTH`. `connectors.list` and `connectors.connect` with the account owner map these too. - `connectors.policy.get` and `connectors.policy.set` carry `effective`: the portal's own result for this user, with its stamp and without provider or subject ids. Nothing is recomputed locally. - A rule write needs the revision the user saw: `expected_revision` is required and must be a revision string; a bad one is refused before any HTTP call. - A tool row carries `no_auth`; a list without the upstream flag is an invalid answer, not `false`. - `connectors.accounts.remove` returns the app of the removed account. An invalid id is `INVALID_PARAMS`. - The tool-list cache is per signed-in member (a hash of the token's `sub`), so two Nous accounts on one profile do not share entries. - A malformed slug is a local error, not a 404 from a server nobody called. Live, staging: no revision and a malformed revision refused locally; a good revision wrote one disabled Gmail tool and returned it in `effective`; the same revision again answered `POLICY_CONFLICT`; the list row showed the tool; the restore brought the member rules back to the start. Live, staging and production, read-only: all 60 tool lists (5483 tools) parse. * fix(connectors): the operation RPCs match their contract; a settled card cannot start a new link Found by two adversarial reviews of the RPC layer and its types. - `connectors.connect` from a chat session with no open operation is refused (`UNKNOWN_OPERATION`). It used to call `manage_connections` through the tool registry with no card: it made a link nobody watched, returned a reply without the required `settled` field, and named an operation that was never registered. There is one way into an operation: the agent's call, or the account owner's `connectors.connect`. "Run it again" inside an open operation is unchanged. - `connection.update` for a session is routed by session key AND profile; two profiles with the same key no longer cross-deliver a sign-in link. The event payload gets the same redaction as the RPC replies. - `connection.respond` runs on the long-handler pool: an approval can start MCP OAuth discovery, which blocked every RPC of the gateway while it ran. - `connectors.list` rows are a closed snake_case model: `connector`, `enabled`, `connected`, `connection_status`, `status_reason`, `gateway_disabled_tools`. The last one is display data: the gateway enforces the rules, the backend only passes the list on. The phantom `name` and `description` are gone, and the desktop uses the generated types instead of hand-written copies. - `tools_listing` (model-only data) no longer rides on `connectors.operation.status`. - `unavailable` is removed from the target states and settle reasons: nothing produces it. The contract generator now fails when a contract enum and its domain enum differ. - `ConnectorErrorReason` is part of the generated TypeScript and OpenRPC. - The desktop sends `connection.respond` on the socket that holds the session, as wake and reissue already did. - Contract violations are logged every time, at error level. - An account connect whose prepare step is slow returns the live operation instead of an error while the operation keeps running. - The MCP-manifest `connector` field leaves this PR (it moves to a later one on top of the catalog-reader change). `hermes_cli/mcp_catalog.py` and `optional-mcps/` are untouched by this PR again. anti-slop: no net-new findings (15 touched files). * fix(connectors): the model gets no sign-in link wherever a card exists; side agents cannot connect The flag that tells the model "a connection card exists" was the session platform (`== "desktop"`). The Ink TUI and the classic CLI also draw a card, so there a connector call on an unconnected app handed the model the raw `connect_url` and told it to pass the link to the user. - The agent turn now declares how a link can reach the user (`tools/connectors/turn.py`): CARD when the agent was built with a connection callback, SIDE for a subagent or a background turn, LINK for a headless run (`-q`, cron, ACP, api_server, messaging). It is set once per tool batch in the agent loop and read by the connector dispatch path, which never sees the agent. The session platform decides return-to-app only. - CARD: the result carries `connect_card_available` and our hint, never the link and never the gateway's own hint. - SIDE: subagents (`delegate_tool`), gateway background turns and the classic CLI `/bg` are built with `side_agent=True`. They hold no `manage_connections` tool on any path that derives the tool list, and a connector call on an unconnected app gets no link, only "report this to the main agent". - LINK is unchanged. - The hosted path with no card builds a detached operation, as the MCP path does, so no `connection.update` is emitted for an operation no client asked for. Names and docstrings that said "off desktop" now say "no card". - A settled card is dead on the desktop: `reissueConnectionTarget` and `respondToConnectionRequest` share one guard and send nothing for a settled or unknown operation. - The model-facing settled result no longer carries `connection_id`; the model repeated it to the user. Shown on the real clients with a real model (rig, fake tool gateway): Ink TUI and classic CLI get `connect_card_available` and no link, the model opens the card, the account connects, the retried call succeeds; `-q` still gets the link; a subagent and a background turn have no `manage_connections` and get the no-link hint; on the desktop a card settled with Continue has no enabled control and sends no RPC. * feat(tools): every call made through tool_search + tool_call shows a real label on all three clients A bridged call showed as a generic `tool_call` row in the Ink TUI and as `⚡ tool_call` in the classic CLI, because the display looked the name up in the tool registry and bridged names are made at run time. The desktop labelled only batches that were all hosted connector calls, by parsing names itself. - `tools/tool_labels.py` is the one place that turns a bridged call into a label: kind, app, action, emoji and text. Hosted: `connectors__gmail__GMAIL_SEND_EMAIL` → "Gmail · send email". MCP: "Linear · list issues". A local deferred tool keeps its own emoji, verb and primary-argument preview. A batch gets exactly one label per entry, always; an entry with no name gets a generic label. - Classic CLI: one row per inner call; the duration on the last row; the failure text on the row of the call that failed. With friendly labels off it prints what it printed before. - Gateway: tool start, progress and complete events and stored transcript rows carry a typed `labels` field. It does not depend on the classic CLI's display setting. Clients no longer parse tool names. - Ink TUI: rows from the labels; the verbose trail keeps Args and Result. - Desktop: `ConnectorExecution` renders hosted, MCP and mixed turns from the labels, one row per call. The labels reach the row under a key no tool argument can use. The connect card it drew under a failed tool result is gone: after `CONNECTION_REQUIRED` the one way in is the agent's own `manage_connections` call. - `tool_search` and `tool_describe` rows read "Searching tools · <query>" and "Reading tool details · N tools". Shown on the real desktop (video and screenshots), the Ink TUI and the classic CLI with the rig: hosted rows, MCP rows, a two-entry batch, a failed entry, a `CONNECTION_REQUIRED` row with no card under it, labels after a reload, and the desktop rows with the classic CLI setting off. * fix(connectors): the model can tell "hosted tools unavailable" from "no such tool"; manage_connections routes MCP names correctly - A failed hosted search or describe used to return nothing, by design, so the model saw only local tools and told the user that a connected app was missing. The local results are unchanged; when the hosted leg failed, the `tool_search` and `tool_describe` results carry `connectors: {status: "unavailable", reason: "unreachable" | "sign_in_expired"}` and one hint line. A rejected token is `sign_in_expired`; an entitlement refusal or a shut gate adds nothing. `tool_describe` no longer lists those names under `not_found` next to "search again". - NS-932. The description now says which side a name belongs to: a bare name is a hosted connector account; `mcp: true` only when the user asks for an MCP server, a local server or an install, or when the name exists only in the catalog; connect and reconnect are hosted verbs, install, enable and authorize are MCP verbs. It names the three clients that draw a card. - A misrouted target is refused with the call that works. Only when the gateway does not know the connector (confirmed on that failure path) and the name is a catalog entry does the target fail with "X is a local MCP server. Call manage_connections with action install ...". It is a per-target outcome: other targets of the same call keep their links and their card. A vendor failure on a name both sides know stays an ordinary failed row. The MCP side mirrors it, and never for an entry that is only not installed. - "Do not re-ask after a skip or a timeout" no longer stops the model when the USER asks for that app again; the description and the settled-result notes say so. A builder saw the model refuse a direct user request. Shown on the Ink TUI and the classic CLI with a real model: a dead gateway and a 401; "connect fxmail" goes hosted; "install the fx-noauth MCP server" goes MCP; "connect fx-noauth" reaches the MCP install card in one corrective round with no hosted mint; a two-target call where one is misrouted still connects the other with exactly one mint. * fix(tui): the connection card answers every key, shows what is happening, and is dead once settled Reproduced on the real Ink TUI with the rig, then fixed: - The keyboard was dead during the sign-in wait: the card kept a `submitting` flag that the normal OAuth path never cleared, and Esc went through the same guard. The in-flight state now belongs to the answered row and clears when that row moves, when any later frame of the operation arrives, or after five seconds. Esc skips the row in every phase; Ctrl+C interrupts the turn (the input handler had no branch for this overlay); Shift+arrows scroll the transcript and the card ignores them; arrow keys no longer move the text cursor and the field focus at once. - The card was lost at turn idle: the overlay flag was cleared while the operation stayed in the store, and a resume dropped the pending card. The flag survives idle, a resume shows the pending card again, a session switch clears it. - States with no branch: `not_connected` and a row with no link fell into the credential form; `expired` vanished with no note. The title and the row text now name the action (connect, reconnect, install, enable, authorize); a failed or expired row with no fields offers Try again / Skip; a failed row WITH fields reopens the form over the typed draft, with the failure above it. - A settled card is dead: at settle the overlay closes and one transcript line per app states the outcome. A settled or dismissed operation id is remembered, so no replay or resume can reopen its card. Esc in the last "Finishing…" moment hides the card and still writes the outcome lines. - A failed `connection.respond` and a browser that did not open are shown on the card in one sentence. Also: `tui_gateway/connector_payload.py` redacted the BOOLEAN `secret` flag of a credential field to the string "[REDACTED]". On the desktop every credential field therefore rendered as a password and lost its prefilled default. A boolean is no longer redacted. * chore(connectors): remove the comments and docstrings this branch added Deletions only. Kept: tool directives (`# noqa`, `// eslint-disable`, ...), `// SAFETY:` lines, and the docstrings of the contract models under `tui_gateway/contracts/`, which become the descriptions in the generated OpenRPC and TypeScript. Checked that no code changed: every Python file has the same AST as before once docstrings and `pass` are ignored (62 files), and every TypeScript file prints the same with comments stripped by the TypeScript printer (32 files). The generated contract files are unchanged. * fix(connectors): a card restored after a reload answers again; every account RPC names auth and org failures Found by the end-to-end runs on the pushed head. - Desktop: after a window reload, Continue on the restored card sent nothing. The answer looked up the backend that holds the session with the runtime session id, the lookup wants the stored id, and a failed lookup returned silently. When the lookup gives no owner the answer now goes out on the window's active socket, which is what main does. - `connectors.policy.get` answered `POLICY_UNAVAILABLE` for a rejected sign-in, a refused scope, a non-member and a missing organisation alike: the handler runs with the gateway's globals and did not import the reason enum, so its own error mapping raised. `connectors.accounts.remove` caught auth failures in its generic branch. `org_required` was mapped on `policy.set` only. All six account RPCs now answer `NEEDS_NOUS_AUTH`, `FORBIDDEN_SCOPE`, `ORG_ACCESS_DENIED` and `ORG_REQUIRED` for those four upstream answers. * fmt(js): `npm run fix` on merge (#118939) Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> * feat(desktop): render plugin-declared settings in the Plugins tab (#46600, #87934) A plugin manifest's `config_schema` now reaches the Desktop: `plugins.manage list` returns each plugin's schema with the current `plugins.entries.<id>.settings` values (`settings_schema`), and a new `settings` action writes edits through `hermes_cli.plugins_state.save_plugin_setting` — the writer extracted from `PluginContext.set_config`, so the plugin, the CLI and the Desktop share one config path, one lock and the same managed-install / managed-key refusals. The Plugins tab grows a gear per plugin with a schema; the inline form is table-driven (`FIELD_CONTROLS` / `INITIAL_TEXT` / `COERCE` keyed on the wire field type) for string / number / boolean / enum / json / secret. Secrets are declared with `type: secret`: the row carries only the `.env` name and a presence flag, the client writes the value through the existing `PUT /api/env` credential route, and the RPC refuses secret keys so nothing lands in config.yaml. Contracts regenerated; docs gain a "Settings form in the Desktop" section. * fix(config): user-installed platform plugins feed env-var metadata (#46600, redo of #46964) `_platform_plugin_manifests()` scanned only the repo's `plugins/platforms/*`, so a third-party platform plugin under `<HERMES_HOME>/plugins/` never reached `OPTIONAL_ENV_VARS`: the Desktop Gateway form and `hermes config` showed bare variable names with no prompt, description or password masking. It now also walks `<HERMES_HOME>/plugins/platforms/*` and flat `<HERMES_HOME>/plugins/*` manifests that declare `kind: platform`. Slim redo of #46964 by @LeonSGP43 onto the refactored helper (the original predates `_platform_plugin_manifests` and replaced `fast_safe_load`). Co-authored-by: LeonSGP43 <cine.dreamer.one@gmail.com> * fix(relay): use a reaction Telegram actuall…
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Upstream Sync
Merged locally in the runner (blobless clone). Review the diff to confirm PMOVES customizations are preserved.
Automated by fork-sync workflow