sync(main): promote the PMOVES overlay line — upstream through 2026-09-20 + 48-commit overlay, AUDIT INSIDE - #11
Conversation
…stop matching lifecycle words inside SQL/text Two live failures on the same guard (cron/lifecycle_guard.py), both of which blocked legitimate diagnostics from inside the gateway: 1. Crash class: the referenced-script walk read compiled binaries as if they were shell scripts. Reading/inspecting a referenced file is now best-effort by construction: executable magic numbers (ELF, PE, Mach-O fat/thin) short-circuit before any full read via a 4KB sniff, NUL-bearing heads are skipped as non-scripts, and unreadable paths of every kind (NUL bytes in the token, ENAMETOOLONG, missing files) degrade to "nothing to scan" instead of raising. A second fail-safe layer wraps the pure-string fallback so the boundary function stays total even if the tokenizer itself fails. 2. False-positive class: the lifecycle regex matched its command shapes inside DATA arguments — SQL string literals passed to sqlite3/psql and grep/rg/journalctl patterns hunting for the lifecycle string in logs. Added a fail-closed second-pass exemption: on a raw regex hit, re-scan with data-sink executables' arguments masked; only a match that survives (i.e. sits in command position) blocks. Masking is skipped for pipes into shells/xargs, command/process substitution, sqlite3 dot-commands and psql backslash escapes, so it can only ever allow, never miss. Behavioral tests: exact live false-positive shapes as negatives, the smuggling shapes as positives, the kill-primitive positive catalog unchanged, and an adversarial never-raises suite (NUL bytes, non-UTF-8, /dev/*, directories, missing files, magic-prefix binaries).
…peners The WAL-reset-vulnerability gate (NousResearch#70055 lineage) could flip a LIVE WAL database to journal_mode=DELETE while another process was writing to it. Observed on state.db (Aug 5): a pytest process on the repo .venv (SQLite 3.50.4, vulnerable) opened the live ~/.hermes/state.db while the gateway (SQLite 3.53.1, WAL) held it, downgraded the journal mode, and destroyed the gateway's committed-but-uncheckpointed WAL transactions (disk rows went 10 -> 0 while memory held 185). cron/executions.db already had the "leave WAL in place, no live downgrade under concurrent openers" rule via the on-disk WAL probe; state.db and every other store shared the hole whenever the mode PROBE itself was blocked by a concurrent opener's locks ("could not read the mode" was treated as "not WAL" -> flip anyway). Generalized in the single journal-mode owner (apply_wal_with_fallback), covering ALL call sites (state.db, kanban.db, projects.db, cron/executions.db, delivery_ledger, async_delegation, verification_evidence, discord recovery, response_store.db, memory_store.db): - _set_journal_mode_no_wait(): the only journal-mode switch primitive for non-WAL targets. Forces busy_timeout=0 around the pragma so SQLite's own exclusivity requirement for leaving WAL becomes the concurrent-opener detector — any other opener (this process or another) makes the flip fail immediately instead of waiting out a busy timeout and sneaking the flip in under a live writer. - Vulnerable-SQLite gate: an unreadable journal mode (probe blocked) now means "ownership not provably exclusive" — leave the mode untouched and warn, never flip. A lock conflict on the flip itself likewise leaves the mode alone. - Configured journal_mode=delete: refuses (raises) rather than downgrading blind when the mode cannot be verified under a concurrent opener. - Filesystem-incompat fallback: re-raises instead of downgrading when the on-disk mode cannot be verified. - New/exclusively-owned DBs on vulnerable builds behave exactly as before (DELETE gate retained per NousResearch#70055). Behavioral tests use a REAL second process (and a real second connection holding an exclusive lock) with the blocked-state assertions running WHILE the holder owns the DB, plus exclusive-ownership downgrade-still-happens coverage.
…ive-DB isolation guard)
Forensics on a live developer machine found pytest fixture rows inside the
REAL ~/.hermes/state.db — sessions with chat_id 'chat-1', '123', 'wx-chat',
and gateway_routing rows whose scope was literally under /tmp/pytest-of-*/.
A pytest-spawned process also opened the live DB and flipped its journal
mode (journal_mode=DELETE fallback on SQLite 3.50.4) under the WAL-mode
gateway writer, destroying committed transcripts ("Persisted transcript
lagged live cached history ... possible FTS write corruption", 15+
occurrences). The existing live-system guard covers kill primitives but not
the SessionDB/SessionStore write paths.
Root cause (leak vector): the session-level HERMES_HOME sandbox in
tests/conftest.py only created a tempdir when HERMES_HOME was UNSET. On a
machine where the shell (e.g. gateway-launched, or an exported
HERMES_HOME=~/.hermes) hands pytest the production home, the sandbox was
skipped entirely — every argless SessionDB()/SessionStore() and every
collection-time DEFAULT_DB_PATH froze onto the real state.db.
Fixes (fail the class, one owner):
* hermes_state._ensure_test_isolation(): single choke point wired into
SessionDB.__init__ (every construction, incl. read_only). Under pytest
(PYTEST_CURRENT_TEST / PYTEST_VERSION — inherited by subprocess
children), a db path resolving to <real-root>/state.db or
<real-root>/profiles/<name>/state.db raises RuntimeError('live-system
guard: ...') before any connection, mkdir, or journal-mode pragma.
* tests/conftest.py: session sandbox now also tempdir-redirects a pre-set
HERMES_HOME that points at the production root (the actual escape
vector); kanban deny-list capture updated to match. New autouse
_state_db_write_guard fixture honors the existing
@pytest.mark.live_system_guard_bypass marker as the escape hatch and
feeds custom (non-~/.hermes) production roots into the guard deny-list.
* gateway/session.py: SessionStore.__init__ no longer swallows the guard's
RuntimeError into the JSONL fallback — guard trips are loud.
* tests/hermes_state/test_live_db_isolation_guard.py: behavioral
regression tests — production paths (direct, profile, read-only,
unnormalized, default-resolution) raise; tmp HERMES_HOME works; bypass
marker works; SessionStore re-raises guard errors but still degrades on
ordinary failures; subprocess child without HERMES_HOME is refused while
a hermetic child succeeds.
No new HERMES_* env vars; no hardcoded ~/.hermes (platform root comes from
hermes_constants._get_platform_default_hermes_home()).
… Path.home monkeypatches Tests like tests/gateway/test_goal_verdict_send.py monkeypatch Path.home() to a tmpdir; resolving the guard's 'real root' through Path.home() made the test's own hermetic home look like production (false positive). Resolve via os.path.expanduser/LOCALAPPDATA instead — the hermetic conftest never rewrites HOME, so this always names the actual production root.
…ing it The turn finalizer already hands back steer text that queued after the final tool batch — result["pending_steer"], with the comment "hand it back to the caller so it can be delivered as the next user turn instead of being silently lost." Every interactive surface honors that contract (cli.py, gateway/run.py, tui_gateway/server.py all requeue it). The delegation layer doesn't: _run_single_child never reads it, so a steer queued into a delegated child that finishes first vanishes with no trace in the completion entry. There is also no sanctioned sender: the registry has interrupt_subagent() but no redirection-side mirror, and session.steer cannot reach children (lazy watch sessions have agent=None, so it 4010s). Complete the contract for delegated children — both halves: - steer_subagent(subagent_id, text): redirection-side mirror of interrupt_subagent(). Resolves the live child in _active_subagents and queues text via AIAgent.steer(). True means queued, not delivered. - missed_steer retention: when the child's result carries pending_steer, _run_single_child names it on the completion entry (missed_steer field plus a summary note) so the parent can re-issue the guidance instead of trusting it landed. This is what makes adding a sender safe: without it the finish-before-drain race silently loses the text — the exact loss the finalizer contract exists to prevent. - subagent.steer gateway RPC beside subagent.interrupt so programmatic hosts (dashboard, voice layers, ACP bridges) get an in-tree caller; catalogued in programmatic-integration.md. - docs: "Steering a Running Subagent" section in delegation.md covering the queued-vs-delivered semantics. Tests: registry-level steer coverage (delivery, unknown id, empty text, dead record, raising agent), the finish-before-drain race retaining missed_steer, and the RPC contract (4000/4002 validation, queued and rejected envelopes).
Per the 'when in doubt, optional' rule — niche prediction-market data skill that sees no regular use; belongs alongside stocks in the finance optional category rather than the default bundle. Install via: hermes skills install official/finance/polymarket
Reload chose the in-process bootout/bootstrap path based on POSIX ancestry, but bootout tears down the job's process coalition, and coalition membership is inherited at spawn and survives reparenting. A gateway-spawned process reparented to PID 1 is no longer an ancestor yet still dies with the coalition, so the retry loop was killed mid-bootstrap and nothing re-registered the label (KeepAlive can't revive a job launchd no longer knows about). - always prefer the detached transient-job helper; it's also correct when genuinely outside the coalition, just asynchronous - wait for the old gateway PID to exit before bootstrapping; bootout only sends SIGTERM and every bootstrap during the drain fails EIO - fall through to the in-process path when the helper can't spawn instead of leaving the plist rewritten but never reloaded
The reload retry loop treated `launchctl list <label>` exit 0 as success, but exit 0 also covers a registered-but-not-running definition (macOS 26+ `state = not running`) — the same trap _probe_launchd_service_running already guards against. Require a PID so success means launchd is supervising a live process, in both the Python loop and the shell helper. Verified against live launchd: a RunAtLoad=false job reports exit 0 with no PID, which the old check accepted and the new one rejects. Note this is NOT what distinguishes a draining instance — measured, the label deregisters within ~1s of bootout while the old process drains on. Waiting for the old PID to exit is what covers that.
Review follow-ups on the salvage: - The helper script's success check accepted any "PID" line, including the "PID" = -1 a recently-crashed job reports — while the in-process path's _parse_launchd_pid_from_list_output rejects non-positive PIDs. Both bash sites now require a positive PID (grep -qE '"PID" = [0-9]+;') so the two paths enforce the same supervised-PID standard. - _graceful_restart_via_sigusr1's drain-wait tail was a duplicate of the new _wait_for_pid_exit — now delegates to it. - Stale comments: the ancestry-detection framing at the top of the reload block, and the exhaustion log's '(refresh ran outside gateway process tree)' which is false on the new helper-spawn-failure fallback path (now '(in-process fallback path)').
… too The pane, in-app browser, and reaction tools were gated on HERMES_DESKTOP=1 — an env var set only on backends Electron spawns itself (local and SSH). A desktop client connected to a plain URL gateway or Hermes Cloud lost all six: they were stripped from the schema before the model saw them, on the same backend whose platform hint was telling it "you are chatting inside the Hermes desktop app". open_preview, read_preview, read_terminal, close_terminal, focus_pane, and react_to_message were all silently absent. The client is not the host. Capability now resolves from the session's own source, which session.create already carries: - The six tools move into a `desktop_ui` toolset, off _HERMES_CORE_TOOLS so no other platform pays their schema. - _gui_surface_toolsets(platform) folds `desktop_ui` (and the existing `project` tools) into the GUI gateway's resolution when the session's platform is the desktop app — the same answer on every topology. - check_fn drops the env probe. It kept the one thing that is genuinely a per-process/user fact: react_to_message's display.message_reactions opt-in, which the desktop mirrors onto whichever gateway it is connected to. react_to_message was doubly broken: it read that toggle behind the env gate, so even a local-backend user's Settings toggle could not reach a remote session. The embedded terminal pane keeps working correctly the other way round: it runs `hermes --tui` against a desktop-spawned backend, and a tui-sourced session gets no GUI tools even though HERMES_DESKTOP=1 is set on that process.
…ess env The rule the preview-tool bug broke, written down so the next GUI-adjacent tool does not rediscover it: the client and the backend are separate machines, so "was this process spawned by Electron?" cannot answer "is a GUI watching?". Names the working pattern (toolset gates the surface, check_fn answers only reachability or user opt-in), the process-wide check_fn TTL cache that makes it the wrong home for a per-session answer, and the test that would have caught it — assert the GUI session gets the tool with the env var absent.
The six GUI tools moved out of `terminal` into their own toolset; the tables still described them as check_fn-gated members of it, and as available to every hermes-* platform bundle.
The transcript decides what a tool call draws and the render budget has to price it. Both sides need the same answer, so the classification moves out of the tool renderer into its own module rather than the budget importing the formatting and i18n weight of fallback-model to ask one question. Adds isSilentTool for the rows that render nothing at all: todo is hoisted to its own panel, and a reaction's UI is the emoji on the bubble.
One weight function served two budgets that protect different things. The store window protects the heap: every message it admits is normalized into the runtime repository whether or not the transcript collapses it, so it has to price the payload it holds. The DOM budget protects the paint, and what a turn mounts is decided by the grouping, not by the bytes behind it. Charging the DOM budget for payload made it count work that never happens. A settled run of twelve reads is one grey summary line, a thought is one collapsed disclosure, a todo is hoisted out of the transcript, and an image is one img however long its data URL — all of it priced as if fully expanded. messageStoreWeight keeps the payload price for the window. messagePaintWeight prices what mounts: collapsed rows flat, silent rows free, cards fixed, and markdown and diffs by size, since those really do build DOM. Both share one character ceiling per message rather than one per part.
On real sessions the button showed up two or three turns from the bottom, over a screen and a half of transcript that had barely painted anything. The budget now spends paint weight, which is what the DOM actually mounts, and 600 units of it — 10-20 agentic turns measured, where a tool-heavy turn prices at 30-90 and a plain exchange at 5-10. A floor of 8 turns covers the session of enormous turns that a weight-only cut still truncates hard; it applies to a real page only, so the small first-paint commit stays small and the backfill a frame later fills the rest. Measured on four stored sessions at the same budget: one went from 3 turns visible to 12, another from 3 to 4, two unchanged. The store window still caps what the DOM can reach at all.
PATCH /api/sessions/{id} only accepted title and end_reason, so the
`pinned` flag the desktop sends was rejected as an unsupported field —
and the client swallows that error. Pins lived in one app's localStorage
and never reached state.db, which also meant the server-side auto-archive
sweep was free to hide the chats a pin exists to keep.
Accept pinned and archived as booleans, route them to the SessionDB
setters that already existed, and include both in the serialized session
so clients can reconcile against server truth.
The list endpoints deliberately back-fill pinned conversations past their LIMIT, then the client sliced the response back down to that same limit and threw them away — so only pins that happened to land inside the most recent page ever rendered, which reads as a cap on how many sessions you can pin. Keep the back-filled rows when trimming, and discount them from the "window came back full" test that drives Load more. Counting a back-fill as a loaded row invented a page that could never be fetched, leaving a Load more button that refetched the same rows forever.
The guard that stops a stale list page from reverting a fresh pin was released on the PATCH's own ack. A list request issued just before the write is slower than the write, so it lands after the ack still carrying the old value, with no guard left to fence it: the pin flips back and the next reconcile pushes that wrong value to the server, making it durable. Keep the guard until a page actually confirms the value written, with a cooldown so a row that never returns can't fence itself forever, and drop it outright when the write fails — the server never changed, so it stays authoritative. Also reset the mirror bookkeeping on a gateway switch. mirrored/pending are per-backend facts; carrying them across a re-home told us the pins were already pushed to a backend that has never seen them.
Two ways a pin got misfiled. The duplicate: a pin is stored on the durable lineage root, but recents, the messaging slice and the backend project tree are three independent fetches and each can surface the same conversation under either its live tip or its root — so the filter compared one identity against the other, missed, and the session rendered in both Pinned and its project group. Match on every id the pin is reachable under. The lost reorder: a drag only reports the pins whose row is loaded, and setPinnedSessionOrder required that list to match the stored one in length, so a single unresolved pin discarded the whole reorder. Treat it as a permutation of a subset — re-slot the named ids, leave the rest.
Dragging one row switched the entire sidebar into a frozen manual mode with no date dividers at all — permanently, for every session, because the manual order replaced the recency sort outright instead of layering on it. Chronology and ranking are separate concerns: keep the calendar buckets where recency put them and apply the hand-picked order only within a bucket, so a drag ranks a chat among its own day's chats and the dividers survive. Rows move as clusters, so a reorder can't strand a branch child from its parent, and a session the saved order doesn't name keeps the slot recency gave it. Two supporting fixes fall out: dnd-kit now receives the ids it actually renders (it was handed the unrendered session order, so a drop computed its target against a list the user wasn't looking at), and an older page that loads no longer jumps above the hand-picked rows — new ids fold in by position rather than all hoisting to the top.
Pinned was capped at half the viewport by its own nested scroller, so past roughly a dozen pins the rest were reachable only by scrolling inside a scroller — a pin you have to go hunting for isn't doing its job. Drop the cap and let the section grow into the sidebar's existing scroll, and stop virtualizing Pinned: virtualization needs a bounded viewport to measure against, which is exactly what's being removed. No count badge, no "show more" — pin as many as you want and they all render. Also back-fill pins on the API-server list route, which was the one list path still windowing purely on recency.
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
It was a 1:1 rebuild of the core statusbar gateway item and shipped enabled by default, so the pill showed up twice. Core chrome stays in shell; demos that clone it belong in hermes-example-plugins.
…chat wire Reasoning-summary models emit one reasoning_content delta per completed summary part, each a self-contained bold heading. The Responses API delimits those parts with summary_index; the OpenAI chat wire carries no such field — verified live against Nous Portal, whose reasoning chunks contain nothing but delta.reasoning_content — so concatenating them glued every part into one unspaced, half-bold paragraph. Re-derive the boundary from the signal the wire does carry: a delta opening a closed bold heading against a mid-line tail. This matches Hermes own Responses adapter, which already joins its summary parts with a blank line.
The native Responses stream does carry summary_index, so the part boundary is structured data here rather than something to infer. Break on a change of index, and leave streams that send no index (plain reasoning_text) untouched.
Salvaged from PR NousResearch#53395 by @izumi0uu: the fire claim's 300s TTL is routinely outlived by real cron jobs, so claim_job_for_fire alone cannot stop a manual cronjob(action='run') from double-firing a job the ticker (or another manual run) is still executing. Extract the ticker's _submit_with_guard running-set check into shared module-level helpers (try_register_running_job / release_running_job) and register manual runs through the same set — one dedupe owner, no drift. Manual runs also become visible to get_running_job_ids (the gateway shutdown drain, NousResearch#60432) and mark_running_jobs_interrupted, which previously could not see them. The background dispatch path pre-checks the running set so a mid-run job reports 'already running' in the tool response immediately instead of as a delayed error completion event; the authoritative atomic check remains in _run_claimed_job on the worker. Co-authored-by: izumi0uu <izumi0uu@gmail.com>
Pins the scheduler-boundary contract: extra_prompt is appended under '## Run Context', does not mutate job['prompt'], and the header is absent when extra_prompt is omitted. Addresses review feedback from harjothkhara on PR NousResearch#57342.
…esearch#57331) Salvaged from PR NousResearch#57342 by @liuhao1024 (with the injection-scan half from PR NousResearch#57360 by @ghedeselmabot): cronjob(action='run', prompt=...) silently discarded the prompt argument — per-run context never reached the spawned cron session. The prompt is now threaded as extra_prompt through the whole chain (cronjob run action → _try_dispatch_background_run/_execute_job_now → _run_claimed_job → run_one_job → run_job → _build_job_prompt) and appended to the stored prompt under a '## Run Context' header for that single fire only — never persisted to the job definition. It passes the same strict _scan_cron_prompt injection scan as stored prompts before firing, and works identically on the background and sync fallback paths. Test fakes across tests/cron/ updated to accept the new kwargs (sibling-test blast radius from the signature change). Co-authored-by: liuhao1024 <liuhao1024@users.noreply.github.com>
… standalone send Fixes NousResearch#61495 When manually triggering cron jobs from a live Matrix session, delivery would fail with "Timeout context manager should be used inside a task" because the aiohttp.ClientTimeout context manager requires a proper asyncio task context. Use asyncio.wait_for() instead of aiohttp.ClientTimeout to avoid this error, following the same pattern as the Weixin platform (gateway/platforms/weixin.py). Changes: - Remove aiohttp.ClientTimeout(total=30) from ClientSession constructor - Wrap the send operation in a nested async function (_do_send) - Use asyncio.wait_for(_do_send(), timeout=30) for timeout handling - Catch asyncio.TimeoutError explicitly and return clear error message
Covers the behavior shipped in NousResearch#80807 (background dispatch for cronjob action='run') and NousResearch#80838 (per-run '## Run Context' prompt, gateway-loop delivery): immediate return with handle, completion re-entering the conversation, in-flight dedupe, transient context injection with prompt scanning, and the sync fallbacks.
… for Hermes (#4) * feat(pmoves-bootstrap): loader + tools_bridge + subscriber stub CLAIM the Hermes-agent fork's slice of the Mavis harness v0 (3-repo coordinated). Companion to POWERFULMOVES/PMOVES.AI PR NousResearch#2477 and POWERFULMOVES/PMOVES-pinokio PR #1 (feat/pmoves-app-launcher). The CGP (pmoves.bootstrap/v1) is the contract that ties the 3 forks together: PMOVES.AI writes it, PMOVES-hermes-agent reads it at session init + registers PMOVES tools alongside the native toolset, PMOVES-pinokio reads it when launching a PMOVES-tagged app. What this commit ships: - pmoves_bootstrap/loader.py - the CGP reader. Accepts both YAML and JSON (Hermes has pyyaml==6.0.3 in core deps, so YAML is the natural format; JSON is supported for PMOVES_BOOTSTRAP_CGP raw- string env vars). 4 input sources in priority order: path arg, source arg, PMOVES_BOOTSTRAP_CGP[_PATH] env var, vendored example. Validates structurally (the vendored v1.schema.json is the source of truth, but no jsonschema dep is added - the thin structural check covers the 80% case). Returns a typed Bootstrap object with has_tool/has_mcp/has_constraint/service/route_for accessors. Stub Bootstrap (safe defaults, all 6 constraints) when no CGP is present - the non-breaking fallback. - pmoves_bootstrap/tools_bridge.py - the PMOVES tools bridge. Reads bootstrap.tools and resolves each entry against the v0 tool registry (Python scripts + CLI binaries). Returns a BridgeResult with registered/skipped/disabled lists. The session init code (a future slice) merges registered tools into Hermes's active toolset. PMOVES_TOOLS_DISABLE env var is a per-tool deny list. Per the 'tagged-services-are-advisory' constraint, unknown tools are silently skipped (warning, not error). - pmoves_bootstrap/subscriber.py - the optional NATS subscriber. v0 is a STUB: no nats-py in Hermes's core deps (adding it would be a meaningful blast-radius change; the deps list warns against it after the Mini Shai-Hulud worm). subscribe() always returns a SubscriberStatus with enabled=False and a clear reason. The TaskEnvelope/ResultEnvelope dataclasses document the wire contract so a future slice can wire in nats-py without changing the public surface. Subjects: pmoves.agent.task.v1 (input), pmoves.agent.result.v1 (output), pmoves.bpm.phase.v1 + pmoves.bpm.pomodoro.v1 (observability, not consumed by Hermes). - pmoves_bootstrap/__init__.py - the public surface. Re-exports load_bootstrap, stub_bootstrap, export_env, register_pmoves_tools, subscribe, and the typed shapes. Future Mavis / Spark / Knuckles sessions do 'from pmoves_bootstrap import load_bootstrap, register_pmoves_tools, subscribe'. - pmoves_bootstrap/cgp_schema/v1.schema.json - vendored copy of the PMOVES.AI schema. The hermes-agent fork doesn't depend on the PMOVES.AI repo at install time. - pmoves_bootstrap/cgp_schema/example.cgp.yaml - vendored YAML example (the same data as the PMOVES.AI example.cgp.yaml). Non-breaking test pair: - No CGP present -> load_bootstrap() returns the stub Bootstrap, register_pmoves_tools() returns BridgeResult(registered=[]), subscribe() returns SubscriberStatus(enabled=False). Existing Hermes behavior unchanged. - CGP present -> load_bootstrap() validates and returns the real Bootstrap, register_pmoves_tools() adds PMOVES tools alongside the native Hermes toolset, subscribe() is a no-op (v0) or picks up Mavis-orchestrator tasks (future slice). Cross-fork plan: - PMOVES.AI PR NousResearch#2477 (writer) - PMOVES-hermes-agent PR feat/pmoves-bootstrap-consumer (agent, this PR) - PMOVES-pinokio PR feat/pmoves-app-launcher (app launcher) All three read the same v1.schema.json - the schema is the contract. The 6 constraints baked into the CGP are honored by the loader's behavior: - no-override-existing-config: the loader never writes to Hermes's own config (cli-config.yaml, hermes_state, etc.) - tagged-services-are-advisory: missing services are skipped in tools_bridge, not failed - no-chit-bypass: no CHIT signing code in this package; the Mavis orchestrator does the signing - no-force-push: this PR's commits use rebase, never --force - no-ci-bypass: PR is in DRAFT, no --admin to skip CI - preserve-existing-tools: tools_bridge adds PMOVES tools alongside Hermes's native toolset, never in place of Tests: 33/33 pass (tests/test_pmoves_bootstrap.py, run with 'python -m pytest tests/test_pmoves_bootstrap.py -o addopts='). No new core dependencies. pyyaml is already a Hermes core dep (see pyproject.toml); no new packages are added. * test(pmoves-bootstrap): 33 pytest tests across 9 groups The test suite for the hermes-side CGP consumer. Mirrors the PMOVES.AI side test taxonomy (load_from_example / load_from_source / validation_failure / stub_fallback / export_env / typed_accessor + tools_bridge + subscriber) but with 33 tests total (vs 22 on the PMOVES.AI side) because the hermes-side has more surface (YAML + JSON parsing, env-var handling, tools_bridge registry resolution, subscriber wire contract). Test groups: - A. LoadFromExampleTests (5) - the vendored example loads + validates, identity, services, routing, constraints - B. LoadFromSourceTests (4) - raw YAML, raw JSON, PMOVES_BOOTSTRAP_CGP env var, PMOVES_BOOTSTRAP_CGP_PATH env var - C. ValidationFailureTests (5) - wrong spec, missing top-level field, missing identity.agent, bad role, non-empty super_nodes - D. StubFallbackTests (2) - no CGP returns the stub; stub has all 6 constraints - E. ExportEnvTests (3) - identity vars, services + routing vars, custom env dict (no process side-effect) - F. TypedAccessorTests (3) - has_tool/has_mcp/has_constraint, service() returns None for missing, route_for() returns None for missing - G. ToolsBridgeTests (6) - stub returns empty, real CGP registers known tools, disable list excludes, unknown goes to skipped, callables are invokable, registry populated at import - H. SubscriberTests (3) - subscribe is safe no-op when disabled, TaskEnvelope round-trip, ResultEnvelope round-trip - I. Constants and subject surfaces (2) - subjects match the orchestrator, KNOWN_TARGETS contains the expected agents Run with: python -m pytest tests/test_pmoves_bootstrap.py -v -o addopts= (the -o addopts= is needed on Windows where the project-level pytest-timeout addopts expects SIGALRM which doesn't exist on Windows; the override disables the addopts so pytest-timeout isn't required for these tests). The autouse _isolate_env fixture strips PMOVES_BOOTSTRAP_*, PMOVES_SUBSCRIBER_*, and PMOVES_TOOLS_* env vars before every test, so the tests are order-independent and don't leak state across test files. * docs(pmoves-bootstrap): README + integration notes The high-level map of the pmoves_bootstrap package + the 3 files (loader, tools_bridge, subscriber) + the non-breaking contract + the design choices (YAML+JSON, no jsonschema, nats-py as follow-up) + the cross-fork plan. Future Mavis / Spark / Knuckles sessions hit this file first to understand the integration. What's in the README: - Why this exists - the 3-repo harness v0 slice, hermes-side role as the heaviest of the three (read CGP, register tools, optional subscriber) - What this slice ships - 8 files (4 .py + 2 vendored schema files + 1 test file + 1 README) - Non-breaking contract - the 6 constraints, the no-CGP fallback, the explicit-source error behavior - Public API - the 5 public functions + 4 typed shapes - Resolution order - the 4 sources in priority order - Why YAML (in addition to JSON) - Hermes has pyyaml in core deps - Why no nats-py in v0 - the blast-radius comment in pyproject.toml warns against adding new packages; the v0 subscriber is a stub with a stable wire contract documented via dataclasses - Tests - 33/33 pass with pytest, 9 test groups - What this slice does NOT do - the 4 intentional follow-ups (wiring into run_agent.py, real nats-py, CHIT trail signing, per-session tool allow-list) - Cross-fork plan - the 3 PRs and the schema as the contract * fix(bootstrap): rename Bootstrap.source to load_source, guard non-string tool_ids, drop dead imports The verifier's review of PR #4 surfaced 6 pre-merge findings; this commit applies the cleanup for all 6: 1. Semantic-naming drift (Bootstrap.source to load_source): the field name 'source' collided with meta.source (the producer). Renamed the load-source attribute to load_source; updated the dataclass field, the _from_dict factory, the stub_bootstrap factory, the docstring, the public surface in __init__.py, and the 5 call sites in test_pmoves_bootstrap.py. 2. Reasoning gap (2-line guard in register_pmoves_tools): a malformed CGP with non-string entries in the tools array (e.g. an object {inject: evil} or an int 42) used to crash the bridge with TypeError on the `in disable` check. Now the bridge skips non-string entries to the `skipped` bucket with a warning log, and the LOG.info(skipped) call uses key=str to sort mixed-type lists. 3. Defense-in-depth (sort key=str): the LOG.info(skipped) call was crashing on sorted([int, str, None]) due to int < str comparison. Added key=str to handle mixed types. The new test_G7 proves the bridge no longer crashes on tools=[gh, dict, 42, None]. 4. Cleanup (dead imports in loader.py): dropped unused re, sys, Iterable from the typing import. 5. Cleanup (test count drift in docstrings): the test file header and the README both said 31 tests / 8 groups; the actual is 33 tests / 9 groups. Updated to 33 / 9. 6. Nit (line endings on vendored JSON): added pmoves_bootstrap/cgp_schema/*.json text eol=lf to .gitattributes so Windows checkouts do not reintroduce CRLF and produce a false-positive drift signal on the SHA-256 byte-compare against the canonical PMOVES.AI copy. Also re-vendored v1.schema.json with the PMOVES.AI side new super_nodes-required + services/routing additionalProperties tightening (the Pinokio fork got the same re-vendor in its separate commit). Test count: 33 to 34 (added test_G7_non_string_tool_id_does_not_crash_bridge). All 34 pass. * chore(mailmap): add Mavis@pmoves.local -> Mavis@users.noreply.github.com The Contributor Attribution Check (CI) was failing because my local commit author (Mavis@pmoves.local) wasn't in the .mailmap. The fix is the standard mailmap format: canonical name + canonical noreply + commit email. Future Mavis commits against this fork will now be attributed correctly. The Hermes fork's contributor graph is otherwise stable; this is a no-op for attribution counting. * chore(release): add Mavis@pmoves.local to AUTHOR_MAP The Contributor Attribution Check (CI) was failing because my commit author Mavis@pmoves.local isn't in scripts/release.py AUTHOR_MAP. The check uses AUTHOR_MAP (not the .mailmap, which is for git shortlog) to attribute commits to GitHub usernames. Added the mapping Mavis@pmoves.local -> Mavis-PMOVES. The .github/PULL_REQUEST_TEMPLATE.md / CONTRIBUTING.md author guidance will surface the canonical username in future PR bodies if the operator wants to backfill the real GitHub handle. Also added a .mailmap entry (commit 35224e5) for git shortlog / GitHub contributor graph, even though the CI check doesn't read the .mailmap. --------- Co-authored-by: Mavis <Mavis@pmoves.local> (cherry picked from commit e68b4ce)
…lution, item validation Review follow-up on the un-strand PR (chatgpt-codex-connector findings): 1. Packaging (P1): pmoves_bootstrap was absent from [tool.setuptools.packages.find] include and its vendored schema/example from package-data. Verified with the same HERMES_NIX_BUILD=1 wheel path the reviewer used: 6/6 entries now present (4 modules + schema + example). Before: zero entries — sealed installs silently dropped the consumer. 2. Lazy candidate resolution (P1): the four CGP sources were read eagerly before the priority loop, so a malformed lower-priority source (bad PMOVES_BOOTSTRAP_CGP env var) could break a call that passed a valid explicit path. Candidates are now resolved inside the loop; read/parse failures fall through like validation failures. 3. String-item validation (P2): _validate_cgp only checked that tools/mcps/constraints were lists, so non-string items passed and TypeError'd later in export_env()'s comma-join. Now rejected at validation with a clear BootstrapError. Tests: 38 passed (34 original + G7 rewritten to the new contract + J1-J4 regression tests covering all three fixes). Wheel-content check done.
…t gate fix Same fix as the contributor-check.yml change on the sync branch: scan the PR's actual base (not hardcoded origin/main) and only the first-parent commit line, so upstream history merged in via sync PRs is never flagged for mapping in this fork.
The check-attribution gate hardcoded 'git merge-base origin/main HEAD' as the scan root. That is wrong twice over on this fork, which has two long-lived bases (main for upstream syncs, PMOVES.AI-Edition-Hardened for the submodule pin): 1. Hardened-base PRs scanned the entire upstream delta embedded in hardened (~24k commits at the Aug 2026 sync) and demanded mappings for upstream authors who have never contributed to this fork. 2. Sync PRs to main scanned all upstream commits the sync brings in — again upstream authors, not fork contributors. Fix (both in the workflow and in scripts/audit_pr_attribution.py, which the gate says to keep in sync): - Scan root is now the PR's actual base branch (github.event.pull_request.base.ref; falls back to main on push events). - --first-parent: only the branch's own commit line counts. Upstream history arrives via a sync merge's SECOND parent and is excluded; fork-authored commits (including conflict-resolution commits after a sync merge) remain on the first-parent line and are still counted. Verified locally by simulating both PR shapes: - feature PR off hardened -> exactly the fork's 2 patch commits - sync PR to main -> only merge + carry-forward commits Upstream contributor emails are attributed in upstream's tree and do not belong in this fork's contributors/ directory (the 17 mappings added in 1e8a69f were reverted in c2cd76c; only the fork owner's own commit email mapping remains, and Mavis@pmoves.local from PR #4).
…rection + explicit refspec
Self-review (delivery agent) with the diff treated as hostile:
P1: BASE_REF was interpolated inline via ${{ github.event.pull_request.base.ref }}
into the run: shell — the classic Actions expression-injection pattern. base.ref
is event-controlled; the workflow header itself warns this workflow runs
PR-controlled code. Moved to env: indirection so the value never touches the
shell parser as code.
P2: 'git fetch origin $BASE_REF' left refs/remotes/origin/$BASE_REF existence
to opportunistic ref updates. Explicit refspec
+refs/heads/$BASE_REF:refs/remotes/origin/$BASE_REF guarantees resolution.
Semantics unchanged: same scan root (PR base branch), same --first-parent
range, verified earlier against both real PR shapes.
feat(pmoves-bootstrap): un-strand the CGP consumer onto the hardened branch
…ed on runners (#9) Provenance: 2026-09-20 session, PR #8 pilot, DARKXSIDE registry-first doctrine. Local transports failed 5x (early-EOF, blobless negotiation, zombie packs); the sync runs where connectivity is native. Resolution: -X theirs (main), PMOVES allowlist keeps Hardened (.gitattributes .mailmap AGENTS.md contributor-check.yml .github/scripts/*). Safety: MERGE_HEAD commit conclusion + parent-count verification close the phantom-push hole; unresolved conflicts fail loudly (exit 1, never push old tip). Danger-room rooms applied: stale-checkout-heal, phantom-deliverable.
…dened # Conflicts: # tests/test_sqlite_wal_reset_gate.py
…ermes_state/ The migrated copy landed via the merge: 419 lines, targets hermes_state_wal (0 stale setattr(hermes_state,) sites), covers all 11 scenarios of the old copy. Carrying the stale duplicate trips upstream compat scanning (14 deprecated setattr sites) and is the stale-copy-carrying anti-pattern. Provenance: hardened-sync run 35605169124 keep-ours resolution, superseded by upstream migration receipt (main tests/hermes_state/ 131-file layout).
The 5 heavy jobs starved 2.5h+ queued on labels this account cannot serve (empty runner assignments; receipts in PR #10). Map: *-32-core and 96-core -> ubuntu-latest, windows-32-core -> windows-latest. Atomic commit via Git Data API: one CI trigger, no half-downgraded intermediates. Provenance: hardened-sync run 35605169124, PR #10.
The Hardened<-main merge (run 35605169124) silently duplicated content in files both sides touched - no textual conflict, so the duplicated export surfaced only as a rolldown parse failure killing 106 desktop test files (CI run 35654743553). Sweep found and fixes: - session-pin-sync.ts: resetSessionPinMirror x2 -> main verbatim (1) - conftest.py: _state_db_write_guard x2 -> 1 (ast-gated) - cronjob_tools.py: _latest_job_output_excerpt + _try_dispatch_background_run x2 -> 1 each (kept main-matching blocks) - AGENTS.md: session-capability doctrine section x2 -> 1 methods_session.py triaged separately (57 top-level 'def _' on both sides compared; exclusion recorded in the follow-up lane). Provenance: hardened-sync run 35605169124, PR #10 pilot, DARKXSIDE registry-first doctrine. Danger-room rooms: phantom-deliverable, template-brace-double-escape (ast gates not eyeballs).
sync(hardened): absorb upstream main through 2026-09-20 — 48-commit overlay reconciliation, AUDIT TABLE INSIDE
|
Important Review skippedAuto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Repository: POWERFULMOVES/PMOVES-hermes-agent/.coderabbit.yaml Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
૮ >ﻌ< ა ci reviewran on ac85cd2 — Merge pull request #10 from POWERFULMOVES/sync/hardened-reso ℹ️ InfoCI-sensitive file review · View jobPR touches sensitive files, but the Sensitive files changed:
debug infoCI timingsCI timings · View report · View jobWall time 63m17s vs 1440m20s (-95.6%). 8 job(s) slower, 9 faster, 2 unchanged.
|
Label audit:
|
| File | Change | Verification |
|---|---|---|
docker.yml |
+2/−2 | ubuntu-latest-32-core ×2 → ubuntu-latest |
e2e-desktop.yml |
+1/−1 | ubuntu-latest-32-core → ubuntu-latest |
js-tests.yml |
+1/−1 | ubuntu-latest-32-core → ubuntu-latest |
nix.yml |
+1/−1 | ubuntu-latest-32-core → ubuntu-latest |
rust-tests.yml |
+1/−1 | ubuntu-latest-32-core → ubuntu-latest |
tests-os.yml |
+1/−1 | windows-latest-32-core → windows-latest |
tests.yml |
+1/−1 | ubuntu-latest-96-core → ubuntu-latest |
Checks applied per file: every removed line carried an unservable label · every added line is a hosted default · zero riders (no permissions:, no secrets., no self-hosted, no curl-pipe-to-shell) · removed/added counts balanced.
Why: these labels starved 2.5h+ queued with empty runner assignments (PR #10 receipts) — they never allocated on this account. Posture fix, not regression; restore path = re-add labels when larger runners exist.
Non-CI-sensitive files (7) — survival-table content, listed for completeness
pyproject.toml (wheel include += pmoves_bootstrap — fixes a main-side packaging gap) · tests/conftest.py (+16, wal-gate guard additions) · tests/agent/test_run_agent_codex_responses.py (+1) · tests/plugins/.../test_fireworks_profile.py (+4) · tests/test_schema_read_probe.py (new, +100) · tests/test_tui_gateway_server_crash_history.py (new, +37) · tui_gateway/methods_session.py (+46, session-pin)
Precedent + authority
PR #9 sanctioned pattern (same authoring agent, label accepted, merged). PR #10 carried the same label with a 33-file audit table (comment 5768548326). Operator authority: DARKXSIDE ("approved for option 2" + standing directives). Merge remains operator-gated per this PR's body — the label certifies the CI-sensitive-file review only.
✅ Danger Room bake receipt — acceptance gate CLOSED (25/25 PASSED)The migrated wal-gate suite's first live execution just happened — on this exact promoted tip, and it passed clean.
What executed live (first time — this closes the acceptance gap)The entire Full acceptance state
Residuals (receipted, queued)
The acceptance gate this chain existed to close is closed. Chain: PR #9 (workflow) → #10 (merge pilot, ruleset pause+restore) → #11 (promotion, squash per |
Fork main absorbed upstream through v2026.9.14 on 2026-09-20 (PR #11 lineage). Hourly ticks stay quiet until upstream ships the next release.
Sibling of hardened-sync (PR #9), opposite direction: hourly release-watch against upstream releases/latest + workflow_dispatch catch-up. Overlay footprint computed at runtime keeps ours on conflict; upstream canonical elsewhere; modify/delete both directions resolved; unresolved conflicts fail loudly (never push stale tip); MERGE_HEAD + parent-count gates close the phantom-push hole; output = new sync/fork-resolved-<n> ref + PR, merge OPERATOR-GATED. Companion state seed: .github/fork-sync-state.json (v2026.9.14 = fork main already absorbed through that release on 2026-09-20). Provenance: DARKXSIDE directive (fork receives upstream updates at dev release cadence); local transports failed 5x on the operator seat, CI is the transport (PRs #9/#10/#11 receipts).
…t) (#12) * ci(fork-sync): registry-first upstream->fork sync workflow Sibling of hardened-sync (PR #9), opposite direction: hourly release-watch against upstream releases/latest + workflow_dispatch catch-up. Overlay footprint computed at runtime keeps ours on conflict; upstream canonical elsewhere; modify/delete both directions resolved; unresolved conflicts fail loudly (never push stale tip); MERGE_HEAD + parent-count gates close the phantom-push hole; output = new sync/fork-resolved-<n> ref + PR, merge OPERATOR-GATED. Companion state seed: .github/fork-sync-state.json (v2026.9.14 = fork main already absorbed through that release on 2026-09-20). Provenance: DARKXSIDE directive (fork receives upstream updates at dev release cadence); local transports failed 5x on the operator seat, CI is the transport (PRs #9/#10/#11 receipts). * ci(fork-sync): seed release-watch state at v2026.9.14 Fork main absorbed upstream through v2026.9.14 on 2026-09-20 (PR #11 lineage). Hourly ticks stay quiet until upstream ships the next release.
Summary
Promotes the PMOVES overlay line into fork
main: Hardened ahead 54 / behind 0 — fully additive, nothing on main is lost (behind-0 receipt via compare API, this session).This lands upstream NousResearch through 2026-09-20 (release v2026.9.14) + the 48-commit PMOVES overlay (pmoves_bootstrap CGP consumer, session-pin feature, contributor-check first-parent gate, wal-gate test migration) + the sync/dedup fixes, as one merge commit.
Provenance chain
hardened-syncworkflow (PR #9)sync/hardened-resolved-2faa3077f5— fixed duplicateresetSessionPinMirror(JS parse blocker, 106 test files), duplicate_state_db_write_guard, 2 cronjob functions, duplicate AGENTS.md doctrine sectionac85cd284, 2 parents verified, ruleset 20589548 paused+restored (serialized snapshot)tests/agent/,tests/tools/,tests/tui_gateway/,tests/hermes_state/)Survival table (PMOVES overlay delta landing on main)
pmoves_bootstrap/— CGP consumer + tools_bridge + subscriber (loader blobee53c7235byte-verified)pyproject.toml— wheel include +=pmoves_bootstrap(fixes a main-side packaging gap: main ships the package but omits it from include)AGENTS.md— session-capability doctrine section (upstream-canonical wording; PMOVES variant preserved verbatim in PR sync(hardened): absorb upstream main through 2026-09-20 — 48-commit overlay reconciliation, AUDIT TABLE INSIDE #10 comment 5767129740).github/workflows/contributor-check.yml— first-parent/base-ref gate (expression-injection hardening).github/workflows/hardened-sync.yml— the sync automation itself.gitattributescgp_schema LF rule,.mailmapMavis entry,tests/hermes_state/wal-gate suite, session-pin feature files, test deltasWhat this PR's CI buys (the acceptance run)
tests/hermes_state/test_sqlite_wal_reset_gate.py— the migrated copy exists on main; whether main's own CI has exercised it is unverified, so this run is the overlay-lineage acceptance receipt)resetSessionPinMirrordedup resolved the 106-file parse failureKnown residuals (receipted, not blockers)
local-models-settingstext-matcher, 1/937) exists on main already — if it re-appears here: retry/disposition, don't chase it into the sync-X theirslossiness bound to the 11-file survival tableMerge: OPERATOR-GATED — two policy choices land with it
Fork main carries
required_linear_history: true(receipted this session), so a merge commit is forbidden on the normal road. Two landing options:ac85cd284, 2 parents receipt-verified), PR sync(hardened): absorb upstream main through 2026-09-20 — 48-commit overlay reconciliation, AUDIT TABLE INSIDE #10's record, and this PR's diff.--merge --admin: preserves the merge-commit ancestry in-line; requires admin bypass of linear history (enforce_admins: false— the bypass exists).Named plainly:
required_approving_review_count: 1and the PR author (POWERFULMOVES) cannot self-approve — GitHub blocks it — so admin merge is the sanctioned mechanism under session doctrine either way.required_conversation_resolution: truemeans bot threads resolve pre-merge. Required checks (9 contexts) exclude JS & TS, so the upstreamlocal-models-settingssnapshot flake cannot block this PR.Post-merge: Danger Room bake before this tip touches any live install (this seat runs v0.21.3 = upstream current).