Skip to content

fix: normalize foreground terminal heartbeat defaults - #119201

Closed
KoNit-K wants to merge 1 commit into
NousResearch:mainfrom
KoNit-K:fix/119196-terminal-heartbeat-foreground
Closed

KoNit-K wants to merge 1 commit into
NousResearch:mainfrom
KoNit-K:fix/119196-terminal-heartbeat-foreground

Conversation

@KoNit-K

@KoNit-K KoNit-K commented Sep 22, 2026

Copy link
Copy Markdown

What does this PR do?

Foreground terminal calls can arrive with schema-materialized background defaults such as notify=false and heartbeat=60. Before this change, the heartbeat caused a validation error before the command executed, allowing repeated retry loops. This PR ignores those non-operative fields for foreground commands while retaining validation for requested foreground notifications and heartbeat delivery for background commands.

Related Issue

Fixes #119196

Type of Change

  • Bug fix

Changes Made

  • tools/terminal_tool.py — normalize heartbeat to zero for foreground terminal calls after rejecting active notification requests.
  • tests/tools/test_process_heartbeat.py — add the complete foreground argument regression and controls for foreground notify=true rejection and background heartbeat forwarding.

How to Test

Run scripts/run_tests.sh tests/tools/test_process_heartbeat.py -q.

Result: 1 passed, 1 skipped (Linux-only test skipped on macOS).

Evidence

  • BEFORE RED: python3 -m pytest tests/tools/test_process_heartbeat.py -q failed with the foreground background=false, notify=false, heartbeat=60 regression case (1 failed, 1 skipped).
  • AFTER GREEN: scripts/run_tests.sh tests/tools/test_process_heartbeat.py -q passed the focused suite (1 passed, 1 skipped).
  • CONTROL: the same test confirms notify=true remains rejected in foreground mode and background=true, heartbeat=120 still forwards the heartbeat and enables completion notification.
  • Lanes: the changed Python files enable the Python lane; ruff check tools/terminal_tool.py tests/tools/test_process_heartbeat.py passed. Docker, Nix, and Linux-only coverage remain for CI because this macOS run does not execute those environments.

Checklist

  • Code normalizes only non-operative foreground background defaults.
  • Code retains rejection for active foreground notification requests.
  • Background heartbeat behavior remains covered by a control assertion.
  • Regression test covers the reported complete argument shape.
  • Focused Python test passes locally.
  • Ruff passes for both changed Python files.
  • Housekeeping: no unrelated files or lockfiles were changed.

@alt-glitch alt-glitch added type/bug Something isn't working P2 Medium — degraded but workaround exists tool/terminal Terminal execution and process management labels Sep 22, 2026

KoNit-K commented Oct 1, 2026

Copy link
Copy Markdown
Author

Closing in favor of 4317ed0e7139, the cherry-pick of #119202 that makes heartbeat=0 schema-valid and the default for ordinary foreground calls. Upstream explicitly retained rejection of positive foreground heartbeat values, so silently clearing them as this PR does would change the chosen behavior.

@KoNit-K KoNit-K closed this Oct 1, 2026
teknium1 added a commit that referenced this pull request Oct 7, 2026
…ntent still is

Follow-up to the cherry-picked #119201 (@KoNit-K): keep main's relocated
test file, flip the one invariant that pinned the refusal (a foreground
call with heartbeat>0 now runs with heartbeat=0 while notify=true stays
refused), reword the refusal so it no longer names heartbeat, and say in
the schema that heartbeat is ignored on foreground commands.

Why: the schema fix (4317ed0) only helps providers that materialize the
schema default. Models that copy a background call shape still send
heartbeat=60 on ordinary commands; this week 46 subagent sessions hit the
refusal 4 times each (173 refusals), 38 of them never ran a single shell
command, and the agents burned $316 / 1,498 model calls on retries.
teknium1 added a commit that referenced this pull request Oct 7, 2026
…ntent still is

Follow-up to the cherry-picked #119201 (@KoNit-K): keep main's relocated
test file, flip the one invariant that pinned the refusal (a foreground
call with heartbeat>0 now runs with heartbeat=0 while notify=true stays
refused), reword the refusal so it no longer names heartbeat, and say in
the schema that heartbeat is ignored on foreground commands.

Why: the schema fix (4317ed0) only helps providers that materialize the
schema default. Models that copy a background call shape still send
heartbeat=60 on ordinary commands; this week 46 subagent sessions hit the
refusal 4 times each (173 refusals), 38 of them never ran a single shell
command, and the agents burned $316 / 1,498 model calls on retries.
jigo-hyunjin added a commit to jigo-ai-team/hermes-agent that referenced this pull request Oct 8, 2026
* feat(plugin-catalog): bump stt-vocab to 1.1.0

New languages setting (de, es, fr, it, nl, pt, tr) so memory notes written
in those languages give clean names: their sentence openers, note words and
month names are no longer glued onto the names. Pin 6145d06 -> 54e83f9.

* feat(plugin-catalog): bump aux-ledger to 1.1.0

Time windows (/aux 6h, /aux 7d, any number of hours or days) and grouping by provider (/aux providers). What the plugin records is unchanged.

* feat(plugin-catalog): bump stream-speed to 1.1.0

C:/Program Files/Git/speed session, time windows (/speed 6h, /speed 7d, any number of hours or days) and grouping by provider (/speed providers). What the plugin records is unchanged.

* chore(plugin-catalog): crew 0.7.9, re-pin to 317c08c

* feat(plugin-catalog): stalkchain banner and update

* chore(catalog): update Local System One to 0.6.0

* plugin-catalog: add banner and bump pin for honcho

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* feat(plugin-catalog): add klipper-print-watch

* chore(plugin-catalog): bump klipper-print-watch to a970970

* chore(catalog): review disclosure for klipper-print-watch

* feat(plugin-catalog): add pastdotdev memory provider

* chore(catalog): pastdotdev disclosure form + contributor map for Kiloris

* catalog: add kimchi-acp-provider (Kimchi harness over ACP stdio)

Pins getkimchi/kimchi-hermes model-providers/kimchi-acp @
f3da18f16dff7220c909227f687b5885ab729f66 (community tier, models
category, requires_env none — the harness subprocess owns auth).
Owner-submitted per the admission rules. Companion to the
kimchi-provider entry.

Co-Authored-By: Kimchi <noreply@kimchi.dev>

* catalog: pin kimchi-hermes@d6ffe628 — make --yolo opt-in per review

* chore(catalog): review disclosure for kimchi-acp-provider

* chore(plugin-catalog): bump pinned-folders entry sha to v0.1.4

* feat(plugin-catalog): add alice-voice plugin

* chore(catalog): review disclosure for alice-voice

* Re-pin agora to v2.0.10

* feat(plugin-catalog): add gbrain-pointer (community GBrain memory provider)

* chore(catalog): review disclosure for gbrain-pointer

* chore: map contributor email for praggybuilds

* feat(catalog): add Jot notes for Hermes Desktop

* feat(catalog): pin Jot 0.2.0 with current screenshots

Use the reviewed main commit with refreshed bilingual Desktop captures
and precise documentation of undoing the latest round of agent edits.
Keep the banner as the card image and English screenshots in the gallery.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(catalog): add Kinprove genealogy research recipes

* chore: map contributor email for lazyants

* feat(plugins): a portable package can ask Hermes to gate its MCP server

Agent Plugins v1 has no trust field, and Hermes's portable mcp.json reader
accepts only type/url/headers for a remote server, so a package server always
ran at the default trust: full. A package whose tools trade, spend or send
(senpi resolving trade approvals on a user's Hyperliquid account) could only
ask the model to check with the user; Hermes itself never prompted.

plugin.json's Hermes extension now accepts
  extensions.com.nousresearch.hermes.servers.<name>.trust: untrusted
which is applied to that mcp.json server exactly like a config.yaml
'trust: untrusted': every write-capable call asks first and fails closed
when unattended. 'full' (the default) is accepted and changes nothing, any
other value refuses the package, so a package can narrow access but never
widen it, and a config.yaml server with the same name still replaces the
package's entry. 'hermes plugins validate' lists the gated servers.

* feat(plugin-catalog): add switchbot-control

* chore(plugin-catalog): bump switchbot-control to bb68cb4

* chore(catalog): review disclosure for switchbot-control

* fix(update): a stash with no Python is never rejected by the import probe (#130101)

The stash-restore health check compared two independent critical-module
import probes and rejected the restore on any difference, even when the
stash restored no Python at all. The probe's outcome also moves with
install state between the two runs (launch preparation relaunching under
the live update surfaces as SystemExit(0); a dependency sync triggered by
the first probe), so a docs-only stash was rolled back, the stash parked,
the gateway restart skipped and the update exited 1.

Only compare imports when the restore touched Python (the same
restored-Python set the syntax check already uses). A restored .py that
breaks an import is still rejected.

* fix(update): skip the restore import check only for non-runtime files

The #130101 exemption skipped the post-restore import comparison whenever
the stash restored no .py file. A restored native extension (toolsets.so
shadowing toolsets.py), bytecode, a .pth or a data file read at import is
not .py, so such a restore was accepted, the stash dropped, and run_agent,
model_tools and toolsets were left failing to import.

Skip the comparison only when every path the stash restores (read from the
stash commit: worktree, index and untracked trees, renames split) has a
closed non-runtime suffix (docs, images), is not a symlink, and lives in a
directory HEAD already tracks. Anything else is compared as before. A
docs-only restore still succeeds.

* docs: stop presenting Desktop Light as a shipped download

The installation page said "Light is a remote-only build variant" next to
the download instructions, which read as if a remote-only Desktop client
can be downloaded. Light exists only as a build target with no release
leg, and there are no plans to publish it. Replace the sentence with how
remote use actually works today (install a package, connect from Settings
-> Gateways), drop "Light packages" from the uninstall note, and mark the
variant as unpublished in the desktop README.

* plugin-catalog: add Agent Hold Em

* chore(catalog): review disclosure for agent-hold-em

* feat(plugin-catalog): add hermes-field-notes — pitfalls and local core patches

* plugin-catalog: hermes-field-notes 1.2.0 (patch recipes)

* plugin-catalog: hermes-field-notes 1.2.1 (review fixes)

* plugin-catalog: hermes-field-notes — README author section

* chore(catalog): review disclosure for hermes-field-notes

* chore: map alrcatraz for salvage of #132751

* fix(agent): let an explicit tool_reason label a soft interrupt

interrupt() derived the published tool-interrupt reason from whether a message
was attached, so on the soft path a system producer had no way to label itself:
passing a message booked the stop as "user sent a new message", passing none
booked it as "user interrupt". Both live in USER_INTERRUPT_REASONS, so
interrupt_issuer() returned None and the turn exit reason became
interrupted_by_user — a system abort recorded as a human stop.

An explicit tool_reason now wins on the soft path as it already did for
hard_cancel, keeping the message heuristic only as the fallback.

(cherry picked from commit 495f1a84a15a4bf88295af1e06f225e25d57715c)

* fix(agent): attribute the terminal batch-timeout abort to the guard

The sequential-batch timeout guard stopped the turn with a bare
interrupt(message), which the soft-path heuristic read as "user sent a new
message". The abort was therefore recorded as interrupted_by_user and the
skipped calls told the user they had sent a message they never wrote.

Pass the guard's own tool_reason so the exit reason names the real issuer and
the skipped-call notice can describe it.

(cherry picked from commit a756f781b2630ea3059c8e56df132c5d8ba35844)

* fix(agent): render skipped-call notices from the recorded reason

Both sequential skip sites hardcoded a user-initiated cause ("User sent a new
message" / "skipped due to user interrupt"), so any system abort that reached
them asserted an action the user never took.

Render the notice from the published reason instead: a system issuer describes
itself, and only the genuinely user-initiated reasons keep the user-facing
wording.

(cherry picked from commit 8efe429e2a485a183c7ad46522d9a5c99cef6a9c)

* fix(agent): hide the synthetic interrupt close from the transcript

close_interrupted_tool_sequence() appends a synthetic assistant turn when an
interrupted tail is a raw tool result, so the next user message does not land as
tool -> user. The row carried a user-visible "Operation interrupted." string and
was persisted as an ordinary assistant message, so background preemption —
session reaping, watchdogs, batch guards — surfaced in the transcript as an
interruption the user was expected to act on.

Store it in the same hidden shape turn_api_call.py already uses for interrupt
placeholders: content="" plus display_kind="hidden" keeps it out of rendered
transcripts, while the api_content sidecar carries the LLM-visible text and is
substituted at API-build time. The alternation guarantee is unchanged.

(cherry picked from commit 22502b1c40c26a08eefc3d35cee067bd1d889bfc)

* test(agent): cover soft-interrupt attribution and skipped-call wording

The exit reason and the skipped-call notice must both name the real issuer: a
system producer that stops the turn softly labels itself via tool_reason, while
a message-carrying soft interrupt with no reason stays a human stop. The wording
derives from the recorded reason, so a system abort never tells the user they
did something they did not.

(cherry picked from commit ec65b95cfb312ebb7a276a26b163df0eaad65b55)

* fix(agent): render concurrent skip notices from the recorded reason

The sequential path was converted to interrupt_skip_wording() but the
concurrent batch kept three hardcoded user-blaming strings: the pre-batch
skip (execute_tool_calls_concurrent), the unfilled-slot notice
(_unfinished_tool_result) and emit_cancelled's hook message. A system
producer that stops a parallel batch — watchdog, lease loss, SSE
disconnect, or any segment routed through execute_tool_calls_segmented —
still told the model "skipped due to user interrupt" while
interrupt_issuer() booked the same turn as interrupted_by_system. Render
all three from the recorded reason like the sequential siblings; the
observability enums (error_type/hook_error_type) stay untouched.

Also harden the interpolation the wording rides on: _append_skipped_tool_results
substitutes {name} with str.replace instead of str.format, and
interrupt_skip_wording escapes braces in caller-supplied reasons, so a
free-form tool_reason can never raise KeyError/IndexError mid-turn.

(cherry picked from commit b785bccbca6fd997921411198d277c39d3d93261)

* fix(agent): hide only the placeholder close, keep caller banners visible

close_interrupted_tool_sequence hid every synthetic closing turn behind
display_kind="hidden", but four of its five callers pass a real user-facing
notice — the truncation terminal ("Response truncated — …"), the context-
overflow partial text, tool-validation partials and abort_turn_on_interrupt's
"Operation interrupted: …" lines. Hiding those removed the turn's only
explanation from every transcript surface.

Narrow the hidden shape to what it was for: an empty final_response or the
bare interrupt placeholder (which previously rendered as the "Operation
interrupted." bubble users never asked for). Any other text is rendered
verbatim as before. Alternation safety is unchanged in both branches.

(cherry picked from commit a637af2ef44c1f0213c3138c458d6f6c48650988)

* fix(agent): stop the interrupt placeholder echoing back as a reply

A hidden assistant row is hidden from the transcript, not from the request:
the projection substitutes its api_content sidecar into content on every call.
The value chosen for that slot in #88955 was "[response interrupted]" - a short
natural-language phrase sitting in the model's own prior-turn position, which
models reproduce verbatim: an ordinary instruction answered with clean
finish_reason=stop and nothing but the placeholder (#132949). The same hazard
was already recognised twice in this codebase - the redirect scaffold is
barred from assistant text because "the model echoes it" (#81841), and a
runaway repetition partial is replaced by a self-describing label rather than
replayed (#112764) - but the placeholder was never given the same treatment.

Change the placeholder to a structural label the model has no reason to say,
and retire rows already persisted under the pre-change spellings: the prep
filter that dropped scaffold ghosts now also drops hidden rows carrying the
legacy placeholder or the scaffold marker.

Dropping is not always safe. A hidden close appended after a tool result is
the only thing separating that tool message from the next user turn; removing
it recreates the tool -> user alternation violation that makes strict providers
ignore prior context (#48879) - the exact failure the close exists to prevent,
and one repair_message_sequence deliberately leaves alone. When the neighbours
would form that pair, rewrite the row's text instead of dropping it. The
rewrite targets the current placeholder, so a neutralised row never re-triggers
the filter.

(cherry picked from commit 227e9a1ed64e0d75594d031423ecd2f5c74cd0b9)

* test(agent): cover replay-echo retirement, and pin behaviour not wording

The retirement filter gets five cases: a legacy placeholder row between two
user turns is dropped; the same row after a tool result is kept with its text
rewritten (dropping it would hand the provider the tool -> user pair that
close_interrupted_tool_sequence exists to prevent); a row carrying the current
placeholder is left untouched, which is what keeps #88955's re-heal guard
intact and proves the rewrite cannot re-trigger itself; a scaffold ghost still
drops as before; and a visible caller banner is never treated as a ghost.

Verified discriminating: reverting the filter to the scaffold-only version
fails all five, restoring it passes.

The existing #88955 assertions compared against the literal
"[response interrupted]" in four places. That couples the tests to one chosen
spelling - any wording change, including this one, breaks them for the wrong
reason. They now compare against the constant, asserting what the contract
actually says: the durable row stays empty and hidden, the wire copy carries
the neutral payload the writer stamped, and neither is ever the interrupt
scaffold (#81841).

(cherry picked from commit e4ebf0138b43547dbf1db36eaa84beb5c80bbfb7)

* fix(agent): drop the diagnostic message from the batch-guard interrupt

The terminal batch-timeout guard still passed "terminal batch tool did not
complete" as the interrupt message. That text lands in _interrupt_message,
which the gateway (_run_agent_drain_pending) and the CLI
(_chat_resolve_interrupt) re-queue as the user's next turn, so a guard abort
injected a fake user message. Call interrupt(tool_reason=...) with no
message, and use a slug-safe reason so the issuer reads
interrupted_by_system(terminal_batch_timeout) rather than
terminal_batch_aborted:_call_timed_out.

Message-less guard shape from #130209.

Co-authored-by: Yuan Li <dskwelmcy@163.com>

* test(agent): align skip-notice fixtures with the recorded-reason wording

Skipped-call notices now render from _tool_interrupt_reason instead of a
hardcoded "skipped due to user interrupt". The concurrent pre-flight stub
never recorded a reason, so it fell to the "Turn interrupted" fallback and
the stale assertion failed; record the new-message reason the real soft
path publishes and assert its wording. The ACP fixture is cosmetic but
should show the shape the executor actually emits.

* fix(agent): retire replay-echo ghosts without mutating the transcript

_neutralise_replay_echo_ghosts rewrote the caller's list in place
(seq[:] = out) and overwrote content/api_content on durable row dicts.
Those dicts are the live session rows the SessionDB flush cursor and
_DB_PERSISTED_MARKER track, so an in-place edit silently diverges the
in-memory row from what state.db holds. BASE built a new list here; keep
that contract: return a filtered list and neutralise a shallow copy of the
row, which only shapes this request's replay.

* test(agent): trim the soft-interrupt salvage tests to two invariants

Keep one attribution test that goes red on BASE (a tool_reason soft
interrupt is booked to the system issuer with no requeued message) and one
replay-echo test (a legacy placeholder row next to a tool tail is
neutralised on a copy, never dropped into tool -> user, and the durable row
is left untouched). The remaining wording/brace/close-row cases pin
implementation detail already covered by the updated finalizer, steer and
concurrent-interrupt tests.

* fix(agent): render skip wording verbatim instead of brace-escaping it

Fold finding 3 (B02 gate): interrupt_skip_wording doubled braces on the
assumption the wording reaches str.format, but every consumer substitutes
{name} with str.replace, so a reason like "a{b}" rendered as "a{{b}}".
Drop the escape and its docstring claim, and look user-stop wording up
directly in _USER_STOP_WORDING.

* fix(agent): attribute Ctrl-C tool cancellation to a user interrupt

Fold finding 2 (B02 gate, sibling of #130207): both KeyboardInterrupt
handlers in tool_executor call agent.interrupt("keyboard interrupt").
With a message and no tool_reason, interrupt() records "user sent a new
message", so the skipped-call notices for a Ctrl-C claimed the user had
sent a new message. Pass tool_reason=_REASON_USER_INTERRUPT so the
recorded reason (and the rendered wording) is "User interrupt"; the
issuer stays None, so attribution to the human is unchanged.

* fix(agent): neutralise legacy interrupt placeholders instead of dropping them

Fold finding 1 (B02 gate, regression): the replay-echo filter now matches
the legacy "[response interrupted]" placeholder and DROPPED the hidden row
unless its neighbours were tool -> user. The common persisted redirect
shape user -> hidden(placeholder) -> user(correction with checkpoint
api_content) then became user -> user, and _merge_consecutive_users merged
the pair: the correction's checkpoint sidecar was discarded and the first
stored user row was rewritten in place.

Always neutralise a copy ({content: "", api_content: placeholder}) and
never drop: role-safe in every neighbour shape, and the echoable text is
still gone. The kept replay test now also covers the redirect shape
(3 rows survive, checkpoint intact, stored dicts untouched).

* refactor(agent): tidy the interrupt-placeholder fold-ups

Fold finding 4 (B02 gate, low): behaviour-neutral cleanups.
- close_interrupted_tool_sequence: the hidden branch only runs when the
  stripped text is empty or already the placeholder, so `stripped or
  _INTERRUPTED_PLACEHOLDER` is always the placeholder; say so.
- tool_executor: compute interrupt_skip_wording once for the cancelled
  outcome instead of twice.
- move _neutralise_replay_echo_ghosts out of prepare_iteration to module
  level (lazy imports kept inside to avoid the conversation_loop cycle).
- the finalizer test asserts the exact placeholder, not just non-empty.

* test(agent): follow the renamed interrupt placeholder in the send-time pad tests

Gate r2 finding 1 (High, regression): a4c687126c renamed the non-final
empty-turn placeholder from "[response interrupted]" to
"[interrupt: no assistant output for this turn]", but two send-time pad
tests in test_partial_stream_finish_reason.py still asserted the old
literal and went red on the stack head. Compare against
agent.agent_runtime_helpers._INTERRUPTED_PLACEHOLDER so the tests track
the constant, not a spelling.

A full grep of tests/ and prod for the old wordings also found a healed
prefill fixture in test_thinking_only_sanitizer.py (still green, but it
models a stub the repair can no longer produce) and a stale docstring in
test_steer.py; both now use the current placeholder.

* test(agent): drive the real batch-timeout guard in the soft tool_reason test

Gate r2 finding 2 (toothless): 482a748431 stopped the terminal batch-timeout
guard from passing a diagnostic message to interrupt(), because
_interrupt_message is what the gateway drain and the CLI re-queue as the
user's next turn. The existing test only called agent.interrupt() directly,
so reverting the guard hunk left it green.

Route the existing test through _run_sequential_tool_execution_middleware
with a prepared terminal call whose worker never settles, so the guard
itself publishes the interrupt; the issuer and `_interrupt_message is None`
asserts now go red when the guard hunk is reverted.

* fix(agent): publish the Ctrl-C reason before the sequential cancel hook

Gate r2 finding 3: in _run_sequential_call's KeyboardInterrupt handler,
ref.emit_cancelled() ran before agent.interrupt(..., tool_reason=user
interrupt), so the post_tool_call hook rendered its wording from an unset
reason and reported "Tool execution cancelled. Turn interrupted" while the
concurrent handler (and 47db3231e2's intent) report "User interrupt".
Publish the interrupt first, matching the concurrent handler.

The existing cancelled-hook test now pins the hook's error_message.

* test(agent): pin the visible-banner branch of close_interrupted_tool_sequence

Gate r2 finding 4 (toothless): f8aea0e81e keeps a caller-supplied banner
("Response truncated — …", partial-delivery text) visible and only hides
the bare placeholder close, but nothing exercised the banner branch, so
reverting it to the always-hidden shape stayed green.

Extend the existing alternation test: a banner close must land as plain
visible content with no display_kind. Red with f8aea0e81e's hunk reverted.

* refactor(agent): match the replay-echo test docstring to the neutralise-only filter

Gate r2 finding 5: d482ce366a made the prep filter neutralise legacy hidden
rows on a copy instead of dropping them, but the module docstring still
described retiring rows except where removal would form tool -> user.
State the actual contract (never drop: removal could form tool -> user or a
user -> user pair that repair merges) and drop _hidden_row's content /
api_content kwargs that no caller passes.

* refactor(agent): move interrupt placeholders into agent_runtime_helpers_placeholders to keep files under the size cap

Code-health ratchet: agent/agent_runtime_helpers.py was 3787 > 3776 and
tests/agent/test_run_agent.py 7014 > 7012 (both over the 2,000 target, so
they may only shrink vs main).

- _INTERRUPTED_PLACEHOLDER / _LEGACY_INTERRUPTED_PLACEHOLDER (and their
  rationale comment) move byte-identically into the leaf sibling
  agent/agent_runtime_helpers_placeholders.py. Every importer (runtime
  helpers via late import, conversation_loop, message_sanitization,
  turn_api_call, turn_iteration_prep and five test files) now reads the
  sibling; no shim is left in the facade. Nothing patches the old path.
- test_keyboard_interrupt_emits_cancelled_post_tool_hook moves unchanged
  into tests/agent/test_run_agent_interrupt_hook.py, reusing the facade's
  agent fixture.

* fix(agent): only relabel a prepared-batch abort as a batch timeout on timeout

The batch-guard tail after _poll_sequential_future is shared by the
"timeout" and "interrupted" outcomes. On a user stop (Ctrl-C, /stop, a
new message) whose prepared terminal worker had not settled after the 3s
grace, the unconditional agent.interrupt(tool_reason="terminal batch
timeout") republished the interrupt: the turn was booked
interrupted_by_system(terminal_batch_timeout), _interrupt_message (the
user's queued next message) and _pending_redirect were nulled, and the
wrong reason was re-fanned to workers.

Keep batch.close() unconditional, but only call interrupt() on the
timeout branch; on the interrupted branch the stop is already published.

Public review blocker on #133271 (@ahrazzle, executed repro; @Enough1122).

* refactor(agent): correct the stale placeholder-consistency comment

The moved comment on _INTERRUPTED_PLACEHOLDER claimed it was "kept
identical to the stub placeholder in chat_completion_helpers"; that
module has no such literal (the claim was already stale on main and was
moved verbatim). State what the constant actually is instead.

Trivial note from the public review on #133271 (@ahrazzle).

* fix(agent): stop terminal batch preparation from overwriting a user stop

terminal_approval_batch caught both _CancelledPreparation and TimeoutError
and called agent.interrupt(str(exc)). _CancelledPreparation is raised
because a user interrupt is already pending, so the user's queued message
and redirect were replaced by "Terminal approval preparation cancelled…",
which the gateway then re-queued as a fake user turn. The TimeoutError
branch booked a system timeout as a user stop with a fake requeue — the
#130207 bug class itself, at the sibling site the batch guard fix missed.

Now: close the batch always; only on timeout publish a message-less
interrupt with tool_reason="terminal batch preparation timeout"; do
nothing on cancellation (the stop is already published).

Gate rP finding 1 (2c).

* test(agent): pin system attribution in sequential skip notices

The two _execute_tool_calls_sequential skip notices (stop before a call,
and "remaining tool call(s)" after one) were rewritten to render the
recorded interrupt reason (6eee1fb6b0), but reverting either back to the
hardcoded "due to user interrupt" / "User sent a new message" left every
kept test green. Extend the system-attribution test to drive both sites
after a tool_reason stop; each reverted site now fails the test.

Gate rP finding 2 (2ab).

* test(agent): pin system attribution in unfinished concurrent slot results

_unfinished_tool_result renders the recorded interrupt reason for a slot
no worker filled (ec51b7d021), but hardcoding the old user wording back
left every kept test green. Extend the system-attribution test to call it
after a tool_reason stop and assert both the row and the post_tool_call
error_message say "Turn aborted — …".

Gate rP finding 3 (2ab).

* test(agent): pin the Ctrl-C attribution on the concurrent tool path

fed76c1df2 made both KeyboardInterrupt sites publish
tool_reason="user interrupt", but only the sequential one was covered:
reverting the concurrent worker's tool_reason left the suite green
(the hook read "User sent a new message"). Parametrize the existing
cancelled-hook test over the concurrent path with two calls.

Gate rP finding 4 (2ab).

* refactor(agent): derive user-stop skip wording from USER_INTERRUPT_REASONS

_USER_STOP_WORDING restated the three USER_INTERRUPT_REASONS keys with
their values capitalised, so a new human-stop reason had to be added in
two places or it would render as "Turn aborted — …". Drop the mirror
and capitalise the recorded reason; output is identical for all three
keys (probe-B02-5.py).

Gate rP finding 5 (reuse).

* refactor(agent): share the hidden interrupt placeholder row

message_sanitization.close_interrupted_tool_sequence and the
turn_api_call repetition-loop exit each spelled the same hidden
assistant row literal (content "", display_kind hidden, api_content
_INTERRUPTED_PLACEHOLDER), and the conversation loop's redirect built it
field by field. Add hidden_interrupt_placeholder_row() to the leaf
placeholders module and use it at all three sites so the shape cannot
drift between them.

Gate rP finding 6 (reuse).

* refactor(agent): import the interrupt placeholder at module level

agent_runtime_helpers (two sites) and message_sanitization imported
_INTERRUPTED_PLACEHOLDER inside the function body. The placeholders
module is a leaf with no imports, so there is no cycle to dodge; import
it once at module level (each module still imports cleanly in a fresh
interpreter). conversation_loop keeps its local import.

Gate rP finding 7 (quality).

* docs(agent): say skipped-result content uses str.replace for {name}

The _append_skipped_tool_results docstring still said content "is
formatted with {name}", but the body substitutes with str.replace so a
recorded interrupt reason containing braces renders verbatim. State
that, so nobody "simplifies" it back to str.format.

Gate rP finding 8 (quality).

* test(agent): move the brace-rendering check after the attribution scenarios

The verbatim "{x}" wording check sat between the batch-guard timeout and
the user-stop scenario, mutating _tool_interrupt_reason mid-story. Move
it to the end, alongside the other wording assertions.

Gate rP finding 9 (quality).

* refactor(agent): hoist the hidden placeholder row import and say why a timeout carries no message

Gate rQ quality Lows: the placeholders module imports nothing, so conversation_loop can import it at module level like turn_api_call; the terminal-approval comment now says callers re-queue _interrupt_message, so a system stop must not set one.

* fix(cron): fire an outage-skipped one-shot whose fire claim never landed

Follow-up to #133283 (greptile review). The due scan stamps a one-shot's
run_claim; when the later fire-claim save fails the run never starts. The
next saved scan skips it on that live claim and ends the recovery window,
so once the claim expires _retire_expired_oneshot saw a stale claim past
grace and skipped it forever: neither fired nor retired.

The scan now marks a run claim it stamps past grace (only an outage lets a
one-shot through there) with outage=true; such a claim without a fire claim
keeps the job covered, so it fires once the claim expires. A fire claim
still means the run may have started, so at-most-once is unchanged.

* fix(cron): report an unwritable tick lock in cron status and doctor

Follow-up to #133283 (greptile review). When only cron/.tick.lock is
unwritable (e.g. root-owned after a sudo run, in a writable cron dir) the
tick degrades and returns 0, so the ticker records a success; probe_store
checked only jobs.json, so `hermes cron status` said jobs would fire and
`hermes doctor` reported a writable store while nothing ran.

probe_store now treats an existing tick lock it cannot write the same way
as a read-only jobs.json target: both status and doctor surface it with the
usual fix hint.

* fix(gateway): send cron store notices for a symlinked cron/ directory

Follow-up to #133283 (greptile review). Store records are keyed by the
resolved store path, but the notice sender rebuilt the owning profile from
that path's parent. A profile whose cron/ is a symlink to another disk has
a foreign parent, so neither its outage nor recovery notice reached any
home channel; forget_homes had the same mismatch and never dropped such a
departed profile's record.

Both now match by resolving each served home's own cron/ path.

* fix(cron): keep the outage marker when a skipped one-shot's dispatch fails

clear_run_claim (the dispatch-failure path: interpreter shutdown, execution
creation failure, pool.submit failure) set run_claim to None. The claiming scan
had already ended the recovery window, and a restart drops the in-memory outage
records, so the next scan saw a past-grace one-shot with no claim and retired it
as missed: the job the outage skipped never fired.

Keep {"outage": True} instead. With no "at" it is a stale claim, so the job is
re-dispatched on the next tick, and _retire_expired_oneshot keeps it. At-most-once
is unchanged: the fire_claim CAS still fences every re-dispatch.

Gate r1 (PM12, 2c) Low finding.

* fix(cron): drop a departed profile's store state by its resolved store path

forget_homes rebuilt each store path from a departed home KEY
(Path(key) / "cron"). cron/AGENTS.md forbids rebuilding a home from a key
(hermes_home_key normcases), and the realpath only follows a symlinked cron/
while the home still exists: a profile whose home was deleted kept its degraded
record, holding the host-wide hermes.cron.store.writable gauge at 0.

register_ticked_homes now records each home's resolved store (new public
store_health.store_key) while the home exists and hands the departed stores to
forget_homes. gateway/cron_store_notices.py uses the same store_key against
record.store (already resolved), so one normaliser decides "is this the served
home's store" everywhere.

Gate r1 (PM12, 2c + quality + 2ab toothless) Low finding.

* test(cron): pin that the tick-lock probe names the lock file

The tick-lock case only asserted probe_store() returned an error, which any other
probe failure would also satisfy. Assert the error names .tick.lock so the case
proves that branch fired.

Gate r1 (PM12, quality) Low finding.

* test(cron): derive the run-claim expiry step from the TTL helper

timedelta(minutes=31) silently encoded ONESHOT_RUN_CLAIM_TTL_SECONDS = 1800; use
_oneshot_run_claim_ttl_seconds() + 60 as tests/cron/test_jobs.py already does, so
a TTL change cannot leave the step inside the live-claim window.

Gate r1 (PM12, quality) Low finding.

* fix(cron): keep store state another still-ticked profile shares through a symlink

Two profiles whose cron/ resolves to one store share a store key; when one left, its departure dropped the degraded record the remaining profile still needed (gauge flipped to writable, a second outage notice, and the earlier outage start lost for one-shot cover). Departed stores now exclude the ones still ticked. The notices test pins it (red without the set difference).

* refactor(cron): name forget_stores for what it takes; fix the clear_run_claim early-return comment

* test(cron): pass encoding= to every text read/write in the store-outage tests

* feat(telemetry): updates that stop before applying say why

Every pre-apply exit of `hermes update` records one closed token on the
receipt at the exit itself (update_receipt.record_stop_reason ->
stop_class), and shared_metrics_update.update_failure_class maps it to
the run's failure_class. aborted_before_apply stays only as the fallback
for an exit with no recorded reason.

Exits that fire before the receipt opens (update-lock refusal, Git
operation in progress, managed install) write no receipt, as before, but
now record a metrics-only row with the same closed class.

The parked receipt copy keeps stop_class, exit_code, the stop reason's
leading label and the restart/user-action flags, so a parked run
classifies the same as the in-process one. A run that committed and is
owed only an interrupt or the user's parked changes reads outcome
partial (failure_class interrupted / local_changes_parked), never failed.

* fixup: keep main.py and update_failure_class under their health caps

* test(telemetry): pre-apply exit classes, parked parity, receipt-less refusals; docs + smoke

- test_shared_metrics_update_stop_reasons: every stop token reads as its own class in-process
  and through the parked copy (no free text kept), unknown tokens and post-apply tokens fall
  back, a committed run that only owes parked changes is partial; the lock and Git-operation
  refusals each record one row and leave the holder's latest.json byte-identical.
- test_shared_metrics_install_failures: a PermissionError stop reason now reads its own class
  permission_denied (was os_error); ENOSPC reads disk_full; other OSErrors keep os_error.
- relay-shared-metrics.md: the new classes, partial outcome and parked fields.
- relay smoke: the failed update receipt names its fetch failure and the row asserts fetch_failed.

* fix(telemetry): review findings on pre-apply stop reasons

- The parked copy keeps a stop reason's label only when it is an exception type name (plus its
  errno token) or one of Hermes' two fixed phrases; any other text before a colon reads '-'
  (a free phrase like 'my project name:' used to survive).
- _exit_after_failed_branch_switch's branch_missing/checkout_move_failed probe cannot raise, so
  the exit it names is unchanged even when git refuses.
- hermes update --plan/--check/--list-venv-holders on a managed install records no update run.
- Docs/contract say 'before the apply stage mark' and name the tokens that fire after git moved
  and restored the tree; tests drive git_error_stop_class and _zip_stop_class directly.

* fix(telemetry): pre-receipt exits keep their initiator; index.lock permission errors; bare-interpreter updates park

- record_stop_without_receipt reads the update marker raw and tags initiator=desktop only when the
  marker names this run's hand-off (delegate line = us, or HERMES_UPDATE_HANDOFF_PID / our parent),
  so a Desktop-started early refusal reads kind=desktop and a CLI run refused by someone else's
  update stays cli.
- _move_checkout_to classifies its git error with the shared classifier (git_output_stop_class):
  index.lock 'File exists' -> git_index_locked, 'Permission denied'/EACCES -> permission_denied
  (now a stop class), else checkout_move_failed.
- _publish_shared_metrics: an interpreter that cannot import the config reader (the -I -S bootstrap
  that finalizes a dependency-preparation failure) never emits; it parks the bounded receipt in
  pending_updates unless it can tell collection is off (loaded config says off, or no config.yaml),
  and the next normal start applies the existing gate (report_pending_updates / begin_process purge).

* docs(telemetry): how update dashboards count a partial run

* test(desktop-update): a hand-off log read racing Add-Content retries, not fails

test_script_killed_before_publishing_the_delegate_runs_no_update failed on
the Windows arm64 lane with PermissionError reading
logs/desktop-update-handoff.log: windows.ps1 appends each line with
Add-Content, and a poll that lands while that write holds the file is a
Windows sharing violation, not a missing line. The poll helper now
retries the read (5 s budget) instead of failing the test on it.
Same test is green on main's latest arm64 run; nothing in this PR
touches the desktop-update scripts.

* test(e2e/desktop): give the build-fail updater its manual-outcome grace before asserting it exited

The build-fail spec polls 30 s for posix.sh to exit after the relaunched
app reports the outcome. That outcome is "manual" (the Desktop build is
owed), and finish() then runs launch_app (1.5 s acceptance) and
stop_ui leave-window, which sleeps the 15 s shim grace before tearing the
UI down: ~13 s of margin, and a slow Linux runner overran it (posix.sh
still alive at the 30 s mark, run 37619760207). The same code passed this
spec on bb177627828 and c4c2d9d8b54; the head that failed differs only by
a Windows test helper. 90 s keeps "no updater is left running" a real
check without racing the grace window.

* fix(cron): live-owner stale-claim reclaim measures silence, not run length

The #115692 reclaim released any running attempt whose claim was older than
max(3 x HERMES_CRON_TIMEOUT, script timeout, 2 h), even when the owner was
healthy and busy. Every cron job that legitimately ran longer than that was
marked `unknown` mid-run, its fire claim yanked from under it and the run
aborted with "lost its durable fire claim ownership" (two weekly jobs on one
install in a single day).

The run monitor now stamps `progress_at` on the execution row every 60 s
while the agent reports recent activity, and the sweep measures the stale
bound from that stamp (claim age for legacy rows). A wedged worker stops
stamping and is still reclaimed once the bound passes.

Shape: the stamper and the inactivity watchdog loop (which reads the same
idle signal) live in a new topical sibling cron/scheduler_liveness.py, so
cron/scheduler.py shrinks (4457 -> 4436 lines) instead of growing.

* fix(tools): normalize foreground terminal heartbeat

* fix(terminal): foreground heartbeat is dropped, not refused; notify intent still is

Follow-up to the cherry-picked #119201 (@KoNit-K): keep main's relocated
test file, flip the one invariant that pinned the refusal (a foreground
call with heartbeat>0 now runs with heartbeat=0 while notify=true stays
refused), reword the refusal so it no longer names heartbeat, and say in
the schema that heartbeat is ignored on foreground commands.

Why: the schema fix (4317ed0e71) only helps providers that materialize the
schema default. Models that copy a background call shape still send
heartbeat=60 on ordinary commands; this week 46 subagent sessions hit the
refusal 4 times each (173 refusals), 38 of them never ran a single shell
command, and the agents burned $316 / 1,498 model calls on retries.

* fix(desktop): setup cards stop redrawing in a loop

useSetupRows built a new row array on every render for the app-owned
lists (accent, layout, theme). That recreated the card's stage callback,
whose effect wrote $setupChooseStages, which re-rendered the card. The
rows are memoized on their inputs now, and stageSetupChoose skips the
atom write when a patch changes nothing.

* fix(onboarding): the first-task handoff follows a renamed owner profile

The setup marker saves the profile setup was made from. After that
profile was renamed, the saved name pointed at a missing directory, so
start_chat and reset targeted nothing. The owner now resolves through the
rename history (profile.yaml previous_names) that rename already records.

* fix(free-tier): a short free-tier 429 no longer blocks for a whole quota window

welcome_refusal_from_headers took the hourly bucket's reset as the wait
even when that bucket had requests left, so Retry-After: 4 with a
healthy bucket became a 2,000 s rate_limited refusal and tripped the
shared breaker. The wait now comes from an empty bucket's reset, else
Retry-After, else the body's wait.

* fix(desktop): a sign-in offer from the previous backend never opens

requestGateway keeps one identity across connection and profile
switches, so a free_tier.claim_nudge reply or status read that landed
after a switch could open the offer, or overwrite $freeTierStatus, for
the new backend. The offer timer, the claim and the finished-turn
re-read are tied to gatewayActivationEpoch() and drop a reply once the
route changes.

* fix(auth): persist profile single-use OAuth credentials when root store is empty

Fixes #103694: When a named profile runs auth add for a single-use
refresh provider (e.g. Anthropic, Codex, xAI) and the root store has
no existing rows for that provider, load_pool() sets _borrowed_root_ids = set().

Previously, add_entry() checked if borrowed_ids:, which evaluated to
False on an empty set. It fell back to _persist() -> persist_pool_entries(),
which routed to the update-only _update_root_pool_rows() on root. Because
root had no rows to update, the credential was dropped with no error and
never written to the profile's auth.json.

Checking if borrowed_ids is not None: ensures that profile-scoped pools
persist their new credential directly to the profile store via write_credential_pool(),
transitioning the profile to owning rows for that provider.

(cherry picked from commit b7467a2999832749f540d86530cf463ca0945b04)
(cherry picked from commit 9c604999e233cf6d74a83e4d7c141785b341b0d6)

* fix(auth): persist first profile OAuth row with empty root

(cherry picked from commit 5293489825af98e3e3f0e6cb485ffffa72d14833)
(cherry picked from commit 4144e14bfa7c021a92db5b0f66a75e2e0ac767d1)

* fix(auth): save a profile's first single-use login, and never print Added for an unsaved row

- Decide the profile claim when the row is written (the check from #101827)
  and drop the None sentinel from #103712. On current main load_pool() leaves
  a profile that borrows from an empty root at the default, so the sentinel
  alone never reached the claim path (its test stays red).
- add_entry() re-reads the store after writing. A missing row raises
  CredentialNotSavedError; `hermes auth add` exits with that message instead
  of printing "Added", and the dashboard pool endpoint returns it as a 400.
- Tests: the two contributor tests become one invariant over every
  SINGLE_USE_REFRESH_POOL_PROVIDERS provider (the row is readable, root
  auth.json is byte-identical), plus the CLI success-line contract.
- Docs: profiles page, EN and zh-Hans.

Fixes #103694

(cherry picked from commit 0e250a005a3353564642ce70f100d89174a583f2)

* fix(auth): the forked-grant heal keeps a profile's own Anthropic login

`hermes -p <profile> auth add anthropic --type oauth` saves a
`manual:hermes_pkce` row in the profile. On the next load_pool(), the
forked-grant heal took every row whose source ends in `hermes_pkce` for a
copy of root's .anthropic_oauth.json whenever root's auth.json held no
Anthropic OAuth row. It moved the profile's token pair into root's file
and deleted the row from the profile: the CLI printed "Added", the login
was gone on the next turn, and root's file now held the profile's
account.

Only a `hermes_pkce` row comes from that file. `manual:hermes_pkce` is
owned by the pool and never written to the singleton
(_commit_anthropic_rotation matches the exact source for the same
reason), so the heal now matches the exact source too. The heal's own
docstring already promises to leave an independent `auth add` grant
alone.

This covers the empty-root case the previous commit enables, and main's
existing variant where root holds only an Anthropic API-key row next to
its .anthropic_oauth.json.

Co-authored-by: Enough <10966420+Enough1122@users.noreply.github.com>

* fix(auth): a profile's auth add never copies root's login into the profile

Root's login for nous, openai-codex and xai-oauth can live only in its
auth.json `providers.<id>` block, with no pool rows yet. In a profile
that borrows from root, load_pool() seeds it as a `device_code` entry
through the root fallback. The write-time claim in add_entry() kept
every entry except root's pool rows, so the profile's first `auth add`
wrote root's single-use refresh token into the profile's auth.json: a
forked grant (#100339) that the heal cannot match, because root has no
pool row for it.

- add_entry() writes only the added row when a profile claims a
  single-use provider. Rows the profile seeds from its own sources are
  seeded again on the next load.
- _seed_tokens_singleton() (openai-codex, xai-oauth) no longer seeds
  root's providers block into a profile that owns its own rows, the
  rule _seed_nous_singleton already applies. On current main such a
  profile copied root's refresh token into its own pool on every load.

* fix(telemetry): update stage rows count the stage a failed run died in

hermes.update.run blamed 971 failed CLI runs in Sep 28 - Oct 5 on the apply
stage while hermes.update.stage held zero failed apply rows. Stage marks are
END marks: every pre-apply exit (fetch, channel, branch, merge, HEAD checks)
leaves no mark for the stage it died in, so the run row inferred the stage and
no stage row said it failed. A failed run whose failed_stage names a stage it
never marked now gets one failed row for it.

Rebuilt on the updater overhaul's followups model (#132386): a committed run
that owes follow-ups is a run-level success and its failed stage already has
its own mark, so only failed runs get the row; refusals never do.

* test: detached-writer custody tests read a pid from a beat that can never be empty

Mechanism: the detached build writer in the two "keeps custody until a
detached writer is gone" tests rewrote its heartbeat file in place with
pathlib.write_text, which opens with O_TRUNC and then writes. Custody
SIGKILLs that writer at an arbitrary instant; when the kill lands between
the truncating open and the write, `beat` is left empty and the teardown's
`beat.read_text().split()[0]` raises IndexError (suppress only covered
ProcessLookupError/ValueError). CI hit it on
test_a_ctrl_c_keeps_custody_until_a_detached_writer_is_gone[verbose-slow-cleanup]
(job 112764854735).

Repro: scratch copy of the writer with a 0.1 s sleep between open() and
write() -> both ctrl-c cases and the group-kill sibling fail with
IndexError at the teardown read on every run (5/5).

Fix: one shared _heartbeat_writer helper writes each beat to `<beat>.tmp`
and os.replace()s it into place, so `beat` always holds a complete
"<pid> <n>" once it exists. With the same injection moved onto the temp
write, all three cases pass 5/5. The ValueError suppress that papered over
a partial read is dropped.

* fix(tests): runner applies passthrough --ignore-glob; windows marker test stops sitting out 150s

Windows-only tests job 112764655375 (run 37612501107, a YAML-only catalog PR)
went red with "667 passed, 0 failed": test_desktop_update_windows_marker.py
hit the runner's 300s per-file cap mid-file ("(300s exceeded; process tree
SIGKILL'd)"), so its 19 tests were never counted and it was listed under
"no tests ran". Run 37612584210 died the same way. Clean runs of this file
take 220-300s (arm64 hit 296.5s and 299.8s), so it is at the cap on every run.

Two defects:

1. The desktop_updater lane gate never worked. tests-os.yml passes
   --ignore-glob='*test_desktop_update_windows_*.py' after `--`, and
   run_tests_parallel.py forwards it to `pytest <file>`. pytest applies
   --ignore/--ignore-glob only while recursing directories, never to an
   explicit file argument, so the flag did nothing. These hand-off tests ran
   on every PR, including ones that never touched the updater, even though
   the job printed "desktop_updater lane off: skipping ...". The runner now
   applies --ignore/--ignore-glob to its own file list (pytest's fnmatch on
   the absolute path, relative patterns anchored at the repo root, which is
   the per-file pytest's cwd) and says how many files it excluded.

2. test_desktop_that_never_exits_is_not_relaunched_over waited out the
   production 150s Desktop-exit ceiling. That one wait is half the file's
   budget. The script already honours HERMES_UPDATE_DESKTOP_EXIT_SECONDS
   ("so the self-tests need not sit it out"), and the posix twin sets it to
   2. Set it to 5 and assert that the script's refusal names 5s, so the
   override is proven to reach the wait. The verdict under test (refuse
   with 4, never relaunch) is the same at any ceiling.

Repro (Linux, CI-shape args):
  python scripts/run_tests_parallel.py --files "<windows_marker>:<windows_cwd>:<other>" \
    -- --ignore-glob='*test_desktop_update_windows_*.py' -m "platforms and not integration"
  base:  "Running 3 test files", both gated files run
  fixed: "note: --ignore/--ignore-glob excluded 2 test files", "Running 1 test files"
New test_passthrough_ignore_drops_files_the_runner_hands_pytest_explicitly
(--files and discovery x glob/glob-spaced/path) fails 6/6 on base and
passes 6/6 fixed.

* fix(desktop): show a queued cron run row immediately after trigger

Run History only rendered sessions the backend had already materialized,
so an accepted Trigger now was invisible for the tens of seconds the
scheduler took to create the run session (#70826). The queued row is
painted from the click, settles when Run History observes a run that
started after it (any load — a fast 1s poll while queued, the cron.changed
broadcast on event-capable backends, or the visibility poll), and drops
at a bounded 90s timeout so a never-materializing run cannot wedge the
button. A failed request drops the row immediately. Pause deliberately
does not settle it: the pause response only persists enabled=false and
never cancels an already-claimed execution.

Queued design (queued-run row + polling + settle semantics) originates in
#70840 by @nv-cho; this ports it onto current main's
createCronTriggerController wiring.

Fixes #70826

* refactor(desktop-i18n): move billingBlock copy into en_billing sibling

Offset the cron queuedRun string growth in the over-cap en.ts/types.ts
facades per the code health ratchet: moved code keeps its cap.

* chore(desktop): refresh locales/_keys.desktop.json for cron.queuedRun

* fix(update): sweep stale git locks at the start of the run, not only before the fetch

The stale-lock / aborted-pack sweep ran only in the apply path, after the
snapshot. A run that died in that window left `.git/index.lock` behind with
nothing in the next run to remove it, so the following update failed with
"File exists" until an operator cleaned it up by hand. The sweep moves to the
top of `_cmd_update_impl`, ahead of the plan and the backup;
`clear_git_debris` already wraps locks and aborted-transfer pack temps (and
the partial-clone maintenance keys), and `_sweep_stale` skips everything
while a git process holds it, so the earlier position is safe.

(Salvaged from #132089 work; the receipt stop-reason half of the original
commit is now owned by the overhaul's receipt API and is not carried here.)

* fix(update): a killed git's index.lock no longer blocks the next update

`.git/index.lock` left by a killed git refused every later merge with
"File exists": the stale-lock sweep keeps any lock younger than 10 minutes,
so a `hermes update` killed mid-run (Windows crash cell `mid_fetch`, run
37573293632) or a version probe killed by its own 3 s timeout (#132089)
failed the next update at "Pulling updates...".

- `gitlock.release_dead_index_lock`: drop the lock as soon as the launch
  repair's ownership proof (`_early_recovery._release_dead_index_lock`:
  /proc fds + lock-keeping gits on Linux, lsof on macOS, the unlink itself on
  Windows) shows no live git holds it. An interrupted tree move keeps its own
  lock judgement. `_cmd_update_impl` calls it first, under the update lock.
- `version_info._git_version_info`: the read-only status probe runs as
  `git --no-optional-locks status`, so a timed-out probe strands no lock.

Part of #132089 (the stale index.lock half; naming why a pre-apply exit fired is separate work)

* fix(update): on Windows a killed git's index.lock waits only for gits working in this checkout

Windows has no /proc to name a lock's holder, and its unlink-as-probe cannot see a
lock-keeping git with its fd closed (a `git commit` in the editor), so the reclaim keeps the
lock while a git.exe has its cwd or a path argument inside the checkout, and whenever the
scan cannot tell (no psutil, an uninspectable git). Any other git on the machine no longer
blocks it, unlike the age-floor sweep's machine-wide check.

* fix(update): only a killed update's index.lock is reclaimed at once

A lock nobody holds open is not proof its git is dead: a git that just created it may not have
opened it yet. The start-of-run reclaim now takes .git/index.lock only when the latest update run
died unfinished (receipt running/interrupted, owner gone) and the lock is newer than that run's
start. Every other lock is left to the age-floor sweep, and the update refuses on it truthfully
(e2e-upgrade test_live_index_lock_is_refused_truthfully went red on fd2214b0ad6).

* fix(update): the start-of-update index.lock reclaim never takes a live git's lock

Review round 2 on #133077. A killed update's receipt plus a younger lock let
release_dead_index_lock delete the lock of a live `/usr/lib/git-core/git-commit -a`
waiting in its editor: the lock fd is closed there, and the Linux holder scan only
knew `comm == git` with an argv subcommand, so the dashed form read as "no holder".
A competing `git add` then got in and the user's commit died "unable to write new
index file".

- The receipt gate stays necessary, never sufficient. The reclaim now also needs
  proof that no git of ANY form works in this checkout (`_held_open(any_git=True)`):
  cwd, path arguments (relative ones against the git's cwd) or GIT_DIR/GIT_WORK_TREE/
  GIT_INDEX_FILE inside the checkout or its git dir, compared by path components.
  A git we cannot place (unreadable cwd/environ, hidepid) keeps the lock.
- Git forms: `git`, a dashed `git-<sub>` from git-core, `git.exe`/`git-<sub>.exe`,
  `-C`/`-c`/`--git-dir` globals, aliases (an unknown subcommand counts). The
  launch-time repair's own scan inverts its list: only known readers are exempt.
- macOS: an empty lsof is no proof any more; `ps` lists every git (comm, so dashed
  forms too) and any lock-keeping one keeps the lock.
- Windows: the psutil scan matches dashed and .exe gits, reads environ, and compares
  by path components (`hermes-backup` was read as inside `hermes`).
- A failed retry no longer erases the evidence: `_killed_update_owns` reads the durable
  per-run records (read-only) and asks about the last run that STARTED before the lock
  was written, so a later failed run cannot discharge the killed one.
- The FIFO test helper kills its blocked git when its own setup assertion fails.

* fix(update): on macOS only a git working in this checkout keeps a dead index.lock

The lsof branch's ps scan counted every lock-keeping git on the machine,
because ps cannot say where a process works. A git commit open in any
other repository (or a sibling test worker's git on CI) therefore kept a
dead index.lock and blocked the launch-time repair after a killed update:
the macOS-only lane failed
test_a_merge_killed_writing_hermes_constants_is_repaired_by_the_next_launch
with "a running git holds the index".

Each candidate git's cwd now comes from `lsof -d cwd`, compared to the
checkout by path components (as on Linux); a candidate whose cwd cannot
be read while it still runs makes the answer unknowable, so the lock is
kept. Probe with real lsof/ps and a dashed git-commit waiting in its
editor: in another repo, 4b3c28c816c named it a holder and the fix does
not; in the checkout, both keep the lock.

* fix(update): the macOS lock scan finds ps when the launcher's PATH has none

The launch-time repair runs with whatever PATH the launcher got; the
macOS-only lane runs it with a PATH holding no git (as on the Windows
install that motivated the repair), so the bare "ps" call failed, the
scan answered "unknowable" and the dead index.lock of a killed update
was kept: test_a_merge_killed_writing_hermes_constants_is_repaired_by_the_next_launch
stayed red. ps is now resolved like lsof already is (which, then /bin/ps,
/usr/bin/ps). Probe: dead lock, PATH with neither git nor ps, at
ba2965261a0 the scan returned None (lock kept); now False (reclaimed).

* fix(update): review the macOS lock scan's lsof/ps lookups; keep it stdlib-only

The resolution-allowlist ratchet flagged the lsof lookup (moved from
_held_open to _held_open_lsof) and the new ps lookup in _ps_git_holder:
both are the stdlib-only launch-time repair asking the OS's own tools,
with an absent tool meaning "unknown" and the lock kept. Rows renamed
and added with that justification.

_still_running also imported psutil inside the launch-time repair, which
runs from hermes_bootstrap before app dependencies can import; it is now
a POSIX os.kill(pid, 0) probe (the lsof/ps branch never runs on Windows).

* fix(desktop): a setup card's typed-pick callback keeps a stable identity

The name card crashed the setup chat on a real Windows bundle (React #185,
maximum update depth): the card wrote its stage callback into
$setupChooseStages, which it also reads, and the callback depended on
rowLabels, a new object every render. 40512011b6 stopped the rows from
changing identity; this removes the dependency itself, so the callback
changes only with kind, mode, multi-select or the request. The labels
effect already stages the row names the callback used to rewrite.

The new test renders the real card: red on 17646902e1 (the crashing
build), green here and with either guard alone.

* test(ci): the CI replay names a silent step's exit code

workflow_steps.py is a helper module, not a test module, so pytest does not
rewrite its asserts: a replayed bash step that exits non-zero with empty
stdout/stderr fails as a bare 'AssertionError' with no exit code. Two
Windows arm64 os-tests runs (jobs 112635777217, 112764629268) failed this way
on unmodified baseline replays, and the NTSTATUS needed to tell a crashed
child (0xC0000142 / 0xC0000005) from a bash failure was lost. Put rc (decimal
and hex), the bash used, and both streams in the message.

* test(ci): the CI replay does not run Git Bash under x64 emulation on Windows arm64

Mechanism: Git for Windows on arm64 ships native git.exe but an x86-64 MSYS
usr/bin/bash.exe, which runs under the WoA x64 emulator. With the os-tests
lane's 16 parallel workers, a replayed step's bash dies before running a line
with 0xC000026F (STATUS_WX86_INTERNAL_ERROR) or 0xC0000005 and empty
stdout/stderr, so an unmodified baseline replay fails as a bare
AssertionError (jobs 112635777217 and 112764629268, two different nodeids).

Repro (throwaway workflows on the real runners, previous commit's diagnostics):
- windows-latest-32-arm-core, 16 copies of the file's replay tests at once,
  4 rounds: 21/64 copies red, every failure 0xC000026F.
- Same load, windows-latest-32-core (x64): 0/64.
- Passing SYSTEMROOT/WINDIR or the full Windows env to the step: 12/48 and
  17/48 vs 15/48 base, so the stripped env is not the cause.
- One process, 32 threads, 600 replays: 0 crashes; it takes concurrent
  processes, i.e. host-wide emulator load.

Seam: _NATIVE_WINDOWS_TOO keeps platforms("linux", "windows") and adds a
skipif for win32 + native_arch() == arm64. The x64 Windows row still runs the
replay on the Windows layout (python.exe-only venv, Git Bash), Linux arm64
keeps it (native bash), and the routing guard that pins the marker passes.

* fix(free-tier): a setup chat turn does not start the sign-in offer clock

The setup chat's last turn (the fork answer and the handoff) finished clean
on the free tier, so it was recorded as the first task. The offer then came
due three minutes after setup ended, which lands in the middle of the first
real task when that task runs longer than three minutes. Seen live on a
Windows dev build: first_task_at was the setup chat's end, not the task's.
Turns in the setup profile's sessions are no longer tasks, as the offer
module already intended ("normally the task setup hands off to").

* i18n(desktop): translate cron.queuedRun in every overlay

Main added the Run History 'Queued run' label (44c1ad7f84) in English only;
this branch's overlay-completeness ratchet fails on a new untranslated key.

* Fix reasoning markup rendering in dashboard

* fix(web): scope structured reasoning UI to assistant messages

Only assistant content with known reasoning tags uses StructuredReasoning.
User and tool messages with literal <action>/<result> tags stay on Markdown.

* test(web): behaviour coverage for structured reasoning transcript rendering

RED on main: the expanded session transcript renders assistant
reasoning with literal <thinking>/<action>/<result> wrapper tags.
Keeps the user-role invariant (literal tags stay on the Markdown path).

* refactor(web): move session source config to SessionsPage_sources sibling; satisfy code-health ratchet

* fix(desktop): hide unknown cloud agent status

Omit the gateway status description when cloud discovery cannot provide a meaningful state, while preserving valid status labels. Add renderer coverage for both unknown and known states.

* test(desktop): cover empty cloud agent status

* fix(pm): reclaim dependency state of deleted checkouts (pm gc + startup worktree prune)

Every checkout that boots Hermes commits its own dependency state under
<home>/installs/<sha256(path)[:16]>/ (venv, test venv, PM runtime; ~200 MB
each) and nothing removed it when the checkout went away. On a host that
runs `hermes -w` campaigns the directory held 499 entries / 92 GB, of which
402 (80 GB) belonged to worktrees and scratch clones that no longer exist.

The state dir already records its checkout in inputs/.project-root. A new
pm/install_states.py lists entries whose recorded checkout is gone and
removes them unless something could still be reading them (install lock
taken, or a generation lease held). `pm gc` runs it and reports the count;
the startup worktree pruner runs it after its own reaping so a removed
worktree's state follows the tree instead of outliving it.

* fix(pm): orphan reclaim runs on every prune pass and keeps checkouts it cannot see

Follow-up to the review of this PR:

- The startup pruner returned early (no .worktrees/, or no tree past 24h) before the
  reclaim ran, so a tree removed by hand kept its state until some other tree aged out.
  _prune_stale_worktrees now runs the reclaim in a finally around the prune phases.
- A missing recorded path is not proof of deletion: a container sharing the data root
  (-v ~/.hermes:/opt/data) cannot see the host's checkout, and gc there deleted the
  host's live dependency environment. A state is an orphan only when its checkout sat
  under the data root that owns installs/, or directly in a .worktrees/ dir that still
  exists.
- PM runtime generation leases (pm-runtime/generations/*/.leases) now count as held,
  so a worker still running from a deleted checkout's runtime keeps it.

* fix(pm): an interrupted orphan reclaim leaves a dir the next pass still finds

The startup reclaim runs on a daemon thread, so a hermes exit can cut a removal
short, and Windows refuses to unlink open files. shutil.rmtree takes children in
scandir order, so inputs/.project-root (the only record naming the checkout)
often went first and the remainder was invisible to every later pass.
Everything else is removed first; the record goes only once nothing else is left.

* fix(kanban): only sweep a parent workspace once the parent itself is terminal

The deferred parent sweep added for #33774 reaps a parent's scratch or
worktree workspace as soon as it has no active children left, without
checking the parent's own status (#133501). Archiving or finishing the
last active child of a parent that is still running/blocked/review
deletes the parent's live workspace from under its worker; the shared
wor…
mrkillbob added a commit to mrkillbob/hermes-agent that referenced this pull request Oct 9, 2026
* chore(plugin-catalog): bump switchbot-control to bb68cb4

* chore(catalog): review disclosure for switchbot-control

* fix(update): a stash with no Python is never rejected by the import probe (#130101)

The stash-restore health check compared two independent critical-module
import probes and rejected the restore on any difference, even when the
stash restored no Python at all. The probe's outcome also moves with
install state between the two runs (launch preparation relaunching under
the live update surfaces as SystemExit(0); a dependency sync triggered by
the first probe), so a docs-only stash was rolled back, the stash parked,
the gateway restart skipped and the update exited 1.

Only compare imports when the restore touched Python (the same
restored-Python set the syntax check already uses). A restored .py that
breaks an import is still rejected.

* fix(update): skip the restore import check only for non-runtime files

The #130101 exemption skipped the post-restore import comparison whenever
the stash restored no .py file. A restored native extension (toolsets.so
shadowing toolsets.py), bytecode, a .pth or a data file read at import is
not .py, so such a restore was accepted, the stash dropped, and run_agent,
model_tools and toolsets were left failing to import.

Skip the comparison only when every path the stash restores (read from the
stash commit: worktree, index and untracked trees, renames split) has a
closed non-runtime suffix (docs, images), is not a symlink, and lives in a
directory HEAD already tracks. Anything else is compared as before. A
docs-only restore still succeeds.

* docs: stop presenting Desktop Light as a shipped download

The installation page said "Light is a remote-only build variant" next to
the download instructions, which read as if a remote-only Desktop client
can be downloaded. Light exists only as a build target with no release
leg, and there are no plans to publish it. Replace the sentence with how
remote use actually works today (install a package, connect from Settings
-> Gateways), drop "Light packages" from the uninstall note, and mark the
variant as unpublished in the desktop README.

* plugin-catalog: add Agent Hold Em

* chore(catalog): review disclosure for agent-hold-em

* feat(plugin-catalog): add hermes-field-notes — pitfalls and local core patches

* plugin-catalog: hermes-field-notes 1.2.0 (patch recipes)

* plugin-catalog: hermes-field-notes 1.2.1 (review fixes)

* plugin-catalog: hermes-field-notes — README author section

* chore(catalog): review disclosure for hermes-field-notes

* chore: map alrcatraz for salvage of #132751

* fix(agent): let an explicit tool_reason label a soft interrupt

interrupt() derived the published tool-interrupt reason from whether a message
was attached, so on the soft path a system producer had no way to label itself:
passing a message booked the stop as "user sent a new message", passing none
booked it as "user interrupt". Both live in USER_INTERRUPT_REASONS, so
interrupt_issuer() returned None and the turn exit reason became
interrupted_by_user — a system abort recorded as a human stop.

An explicit tool_reason now wins on the soft path as it already did for
hard_cancel, keeping the message heuristic only as the fallback.

(cherry picked from commit 495f1a84a15a4bf88295af1e06f225e25d57715c)

* fix(agent): attribute the terminal batch-timeout abort to the guard

The sequential-batch timeout guard stopped the turn with a bare
interrupt(message), which the soft-path heuristic read as "user sent a new
message". The abort was therefore recorded as interrupted_by_user and the
skipped calls told the user they had sent a message they never wrote.

Pass the guard's own tool_reason so the exit reason names the real issuer and
the skipped-call notice can describe it.

(cherry picked from commit a756f781b2630ea3059c8e56df132c5d8ba35844)

* fix(agent): render skipped-call notices from the recorded reason

Both sequential skip sites hardcoded a user-initiated cause ("User sent a new
message" / "skipped due to user interrupt"), so any system abort that reached
them asserted an action the user never took.

Render the notice from the published reason instead: a system issuer describes
itself, and only the genuinely user-initiated reasons keep the user-facing
wording.

(cherry picked from commit 8efe429e2a485a183c7ad46522d9a5c99cef6a9c)

* fix(agent): hide the synthetic interrupt close from the transcript

close_interrupted_tool_sequence() appends a synthetic assistant turn when an
interrupted tail is a raw tool result, so the next user message does not land as
tool -> user. The row carried a user-visible "Operation interrupted." string and
was persisted as an ordinary assistant message, so background preemption —
session reaping, watchdogs, batch guards — surfaced in the transcript as an
interruption the user was expected to act on.

Store it in the same hidden shape turn_api_call.py already uses for interrupt
placeholders: content="" plus display_kind="hidden" keeps it out of rendered
transcripts, while the api_content sidecar carries the LLM-visible text and is
substituted at API-build time. The alternation guarantee is unchanged.

(cherry picked from commit 22502b1c40c26a08eefc3d35cee067bd1d889bfc)

* test(agent): cover soft-interrupt attribution and skipped-call wording

The exit reason and the skipped-call notice must both name the real issuer: a
system producer that stops the turn softly labels itself via tool_reason, while
a message-carrying soft interrupt with no reason stays a human stop. The wording
derives from the recorded reason, so a system abort never tells the user they
did something they did not.

(cherry picked from commit ec65b95cfb312ebb7a276a26b163df0eaad65b55)

* fix(agent): render concurrent skip notices from the recorded reason

The sequential path was converted to interrupt_skip_wording() but the
concurrent batch kept three hardcoded user-blaming strings: the pre-batch
skip (execute_tool_calls_concurrent), the unfilled-slot notice
(_unfinished_tool_result) and emit_cancelled's hook message. A system
producer that stops a parallel batch — watchdog, lease loss, SSE
disconnect, or any segment routed through execute_tool_calls_segmented —
still told the model "skipped due to user interrupt" while
interrupt_issuer() booked the same turn as interrupted_by_system. Render
all three from the recorded reason like the sequential siblings; the
observability enums (error_type/hook_error_type) stay untouched.

Also harden the interpolation the wording rides on: _append_skipped_tool_results
substitutes {name} with str.replace instead of str.format, and
interrupt_skip_wording escapes braces in caller-supplied reasons, so a
free-form tool_reason can never raise KeyError/IndexError mid-turn.

(cherry picked from commit b785bccbca6fd997921411198d277c39d3d93261)

* fix(agent): hide only the placeholder close, keep caller banners visible

close_interrupted_tool_sequence hid every synthetic closing turn behind
display_kind="hidden", but four of its five callers pass a real user-facing
notice — the truncation terminal ("Response truncated — …"), the context-
overflow partial text, tool-validation partials and abort_turn_on_interrupt's
"Operation interrupted: …" lines. Hiding those removed the turn's only
explanation from every transcript surface.

Narrow the hidden shape to what it was for: an empty final_response or the
bare interrupt placeholder (which previously rendered as the "Operation
interrupted." bubble users never asked for). Any other text is rendered
verbatim as before. Alternation safety is unchanged in both branches.

(cherry picked from commit a637af2ef44c1f0213c3138c458d6f6c48650988)

* fix(agent): stop the interrupt placeholder echoing back as a reply

A hidden assistant row is hidden from the transcript, not from the request:
the projection substitutes its api_content sidecar into content on every call.
The value chosen for that slot in #88955 was "[response interrupted]" - a short
natural-language phrase sitting in the model's own prior-turn position, which
models reproduce verbatim: an ordinary instruction answered with clean
finish_reason=stop and nothing but the placeholder (#132949). The same hazard
was already recognised twice in this codebase - the redirect scaffold is
barred from assistant text because "the model echoes it" (#81841), and a
runaway repetition partial is replaced by a self-describing label rather than
replayed (#112764) - but the placeholder was never given the same treatment.

Change the placeholder to a structural label the model has no reason to say,
and retire rows already persisted under the pre-change spellings: the prep
filter that dropped scaffold ghosts now also drops hidden rows carrying the
legacy placeholder or the scaffold marker.

Dropping is not always safe. A hidden close appended after a tool result is
the only thing separating that tool message from the next user turn; removing
it recreates the tool -> user alternation violation that makes strict providers
ignore prior context (#48879) - the exact failure the close exists to prevent,
and one repair_message_sequence deliberately leaves alone. When the neighbours
would form that pair, rewrite the row's text instead of dropping it. The
rewrite targets the current placeholder, so a neutralised row never re-triggers
the filter.

(cherry picked from commit 227e9a1ed64e0d75594d031423ecd2f5c74cd0b9)

* test(agent): cover replay-echo retirement, and pin behaviour not wording

The retirement filter gets five cases: a legacy placeholder row between two
user turns is dropped; the same row after a tool result is kept with its text
rewritten (dropping it would hand the provider the tool -> user pair that
close_interrupted_tool_sequence exists to prevent); a row carrying the current
placeholder is left untouched, which is what keeps #88955's re-heal guard
intact and proves the rewrite cannot re-trigger itself; a scaffold ghost still
drops as before; and a visible caller banner is never treated as a ghost.

Verified discriminating: reverting the filter to the scaffold-only version
fails all five, restoring it passes.

The existing #88955 assertions compared against the literal
"[response interrupted]" in four places. That couples the tests to one chosen
spelling - any wording change, including this one, breaks them for the wrong
reason. They now compare against the constant, asserting what the contract
actually says: the durable row stays empty and hidden, the wire copy carries
the neutral payload the writer stamped, and neither is ever the interrupt
scaffold (#81841).

(cherry picked from commit e4ebf0138b43547dbf1db36eaa84beb5c80bbfb7)

* fix(agent): drop the diagnostic message from the batch-guard interrupt

The terminal batch-timeout guard still passed "terminal batch tool did not
complete" as the interrupt message. That text lands in _interrupt_message,
which the gateway (_run_agent_drain_pending) and the CLI
(_chat_resolve_interrupt) re-queue as the user's next turn, so a guard abort
injected a fake user message. Call interrupt(tool_reason=...) with no
message, and use a slug-safe reason so the issuer reads
interrupted_by_system(terminal_batch_timeout) rather than
terminal_batch_aborted:_call_timed_out.

Message-less guard shape from #130209.

Co-authored-by: Yuan Li <dskwelmcy@163.com>

* test(agent): align skip-notice fixtures with the recorded-reason wording

Skipped-call notices now render from _tool_interrupt_reason instead of a
hardcoded "skipped due to user interrupt". The concurrent pre-flight stub
never recorded a reason, so it fell to the "Turn interrupted" fallback and
the stale assertion failed; record the new-message reason the real soft
path publishes and assert its wording. The ACP fixture is cosmetic but
should show the shape the executor actually emits.

* fix(agent): retire replay-echo ghosts without mutating the transcript

_neutralise_replay_echo_ghosts rewrote the caller's list in place
(seq[:] = out) and overwrote content/api_content on durable row dicts.
Those dicts are the live session rows the SessionDB flush cursor and
_DB_PERSISTED_MARKER track, so an in-place edit silently diverges the
in-memory row from what state.db holds. BASE built a new list here; keep
that contract: return a filtered list and neutralise a shallow copy of the
row, which only shapes this request's replay.

* test(agent): trim the soft-interrupt salvage tests to two invariants

Keep one attribution test that goes red on BASE (a tool_reason soft
interrupt is booked to the system issuer with no requeued message) and one
replay-echo test (a legacy placeholder row next to a tool tail is
neutralised on a copy, never dropped into tool -> user, and the durable row
is left untouched). The remaining wording/brace/close-row cases pin
implementation detail already covered by the updated finalizer, steer and
concurrent-interrupt tests.

* fix(agent): render skip wording verbatim instead of brace-escaping it

Fold finding 3 (B02 gate): interrupt_skip_wording doubled braces on the
assumption the wording reaches str.format, but every consumer substitutes
{name} with str.replace, so a reason like "a{b}" rendered as "a{{b}}".
Drop the escape and its docstring claim, and look user-stop wording up
directly in _USER_STOP_WORDING.

* fix(agent): attribute Ctrl-C tool cancellation to a user interrupt

Fold finding 2 (B02 gate, sibling of #130207): both KeyboardInterrupt
handlers in tool_executor call agent.interrupt("keyboard interrupt").
With a message and no tool_reason, interrupt() records "user sent a new
message", so the skipped-call notices for a Ctrl-C claimed the user had
sent a new message. Pass tool_reason=_REASON_USER_INTERRUPT so the
recorded reason (and the rendered wording) is "User interrupt"; the
issuer stays None, so attribution to the human is unchanged.

* fix(agent): neutralise legacy interrupt placeholders instead of dropping them

Fold finding 1 (B02 gate, regression): the replay-echo filter now matches
the legacy "[response interrupted]" placeholder and DROPPED the hidden row
unless its neighbours were tool -> user. The common persisted redirect
shape user -> hidden(placeholder) -> user(correction with checkpoint
api_content) then became user -> user, and _merge_consecutive_users merged
the pair: the correction's checkpoint sidecar was discarded and the first
stored user row was rewritten in place.

Always neutralise a copy ({content: "", api_content: placeholder}) and
never drop: role-safe in every neighbour shape, and the echoable text is
still gone. The kept replay test now also covers the redirect shape
(3 rows survive, checkpoint intact, stored dicts untouched).

* refactor(agent): tidy the interrupt-placeholder fold-ups

Fold finding 4 (B02 gate, low): behaviour-neutral cleanups.
- close_interrupted_tool_sequence: the hidden branch only runs when the
  stripped text is empty or already the placeholder, so `stripped or
  _INTERRUPTED_PLACEHOLDER` is always the placeholder; say so.
- tool_executor: compute interrupt_skip_wording once for the cancelled
  outcome instead of twice.
- move _neutralise_replay_echo_ghosts out of prepare_iteration to module
  level (lazy imports kept inside to avoid the conversation_loop cycle).
- the finalizer test asserts the exact placeholder, not just non-empty.

* test(agent): follow the renamed interrupt placeholder in the send-time pad tests

Gate r2 finding 1 (High, regression): a4c687126c renamed the non-final
empty-turn placeholder from "[response interrupted]" to
"[interrupt: no assistant output for this turn]", but two send-time pad
tests in test_partial_stream_finish_reason.py still asserted the old
literal and went red on the stack head. Compare against
agent.agent_runtime_helpers._INTERRUPTED_PLACEHOLDER so the tests track
the constant, not a spelling.

A full grep of tests/ and prod for the old wordings also found a healed
prefill fixture in test_thinking_only_sanitizer.py (still green, but it
models a stub the repair can no longer produce) and a stale docstring in
test_steer.py; both now use the current placeholder.

* test(agent): drive the real batch-timeout guard in the soft tool_reason test

Gate r2 finding 2 (toothless): 482a748431 stopped the terminal batch-timeout
guard from passing a diagnostic message to interrupt(), because
_interrupt_message is what the gateway drain and the CLI re-queue as the
user's next turn. The existing test only called agent.interrupt() directly,
so reverting the guard hunk left it green.

Route the existing test through _run_sequential_tool_execution_middleware
with a prepared terminal call whose worker never settles, so the guard
itself publishes the interrupt; the issuer and `_interrupt_message is None`
asserts now go red when the guard hunk is reverted.

* fix(agent): publish the Ctrl-C reason before the sequential cancel hook

Gate r2 finding 3: in _run_sequential_call's KeyboardInterrupt handler,
ref.emit_cancelled() ran before agent.interrupt(..., tool_reason=user
interrupt), so the post_tool_call hook rendered its wording from an unset
reason and reported "Tool execution cancelled. Turn interrupted" while the
concurrent handler (and 47db3231e2's intent) report "User interrupt".
Publish the interrupt first, matching the concurrent handler.

The existing cancelled-hook test now pins the hook's error_message.

* test(agent): pin the visible-banner branch of close_interrupted_tool_sequence

Gate r2 finding 4 (toothless): f8aea0e81e keeps a caller-supplied banner
("Response truncated — …", partial-delivery text) visible and only hides
the bare placeholder close, but nothing exercised the banner branch, so
reverting it to the always-hidden shape stayed green.

Extend the existing alternation test: a banner close must land as plain
visible content with no display_kind. Red with f8aea0e81e's hunk reverted.

* refactor(agent): match the replay-echo test docstring to the neutralise-only filter

Gate r2 finding 5: d482ce366a made the prep filter neutralise legacy hidden
rows on a copy instead of dropping them, but the module docstring still
described retiring rows except where removal would form tool -> user.
State the actual contract (never drop: removal could form tool -> user or a
user -> user pair that repair merges) and drop _hidden_row's content /
api_content kwargs that no caller passes.

* refactor(agent): move interrupt placeholders into agent_runtime_helpers_placeholders to keep files under the size cap

Code-health ratchet: agent/agent_runtime_helpers.py was 3787 > 3776 and
tests/agent/test_run_agent.py 7014 > 7012 (both over the 2,000 target, so
they may only shrink vs main).

- _INTERRUPTED_PLACEHOLDER / _LEGACY_INTERRUPTED_PLACEHOLDER (and their
  rationale comment) move byte-identically into the leaf sibling
  agent/agent_runtime_helpers_placeholders.py. Every importer (runtime
  helpers via late import, conversation_loop, message_sanitization,
  turn_api_call, turn_iteration_prep and five test files) now reads the
  sibling; no shim is left in the facade. Nothing patches the old path.
- test_keyboard_interrupt_emits_cancelled_post_tool_hook moves unchanged
  into tests/agent/test_run_agent_interrupt_hook.py, reusing the facade's
  agent fixture.

* fix(agent): only relabel a prepared-batch abort as a batch timeout on timeout

The batch-guard tail after _poll_sequential_future is shared by the
"timeout" and "interrupted" outcomes. On a user stop (Ctrl-C, /stop, a
new message) whose prepared terminal worker had not settled after the 3s
grace, the unconditional agent.interrupt(tool_reason="terminal batch
timeout") republished the interrupt: the turn was booked
interrupted_by_system(terminal_batch_timeout), _interrupt_message (the
user's queued next message) and _pending_redirect were nulled, and the
wrong reason was re-fanned to workers.

Keep batch.close() unconditional, but only call interrupt() on the
timeout branch; on the interrupted branch the stop is already published.

Public review blocker on #133271 (@ahrazzle, executed repro; @Enough1122).

* refactor(agent): correct the stale placeholder-consistency comment

The moved comment on _INTERRUPTED_PLACEHOLDER claimed it was "kept
identical to the stub placeholder in chat_completion_helpers"; that
module has no such literal (the claim was already stale on main and was
moved verbatim). State what the constant actually is instead.

Trivial note from the public review on #133271 (@ahrazzle).

* fix(agent): stop terminal batch preparation from overwriting a user stop

terminal_approval_batch caught both _CancelledPreparation and TimeoutError
and called agent.interrupt(str(exc)). _CancelledPreparation is raised
because a user interrupt is already pending, so the user's queued message
and redirect were replaced by "Terminal approval preparation cancelled…",
which the gateway then re-queued as a fake user turn. The TimeoutError
branch booked a system timeout as a user stop with a fake requeue — the
#130207 bug class itself, at the sibling site the batch guard fix missed.

Now: close the batch always; only on timeout publish a message-less
interrupt with tool_reason="terminal batch preparation timeout"; do
nothing on cancellation (the stop is already published).

Gate rP finding 1 (2c).

* test(agent): pin system attribution in sequential skip notices

The two _execute_tool_calls_sequential skip notices (stop before a call,
and "remaining tool call(s)" after one) were rewritten to render the
recorded interrupt reason (6eee1fb6b0), but reverting either back to the
hardcoded "due to user interrupt" / "User sent a new message" left every
kept test green. Extend the system-attribution test to drive both sites
after a tool_reason stop; each reverted site now fails the test.

Gate rP finding 2 (2ab).

* test(agent): pin system attribution in unfinished concurrent slot results

_unfinished_tool_result renders the recorded interrupt reason for a slot
no worker filled (ec51b7d021), but hardcoding the old user wording back
left every kept test green. Extend the system-attribution test to call it
after a tool_reason stop and assert both the row and the post_tool_call
error_message say "Turn aborted — …".

Gate rP finding 3 (2ab).

* test(agent): pin the Ctrl-C attribution on the concurrent tool path

fed76c1df2 made both KeyboardInterrupt sites publish
tool_reason="user interrupt", but only the sequential one was covered:
reverting the concurrent worker's tool_reason left the suite green
(the hook read "User sent a new message"). Parametrize the existing
cancelled-hook test over the concurrent path with two calls.

Gate rP finding 4 (2ab).

* refactor(agent): derive user-stop skip wording from USER_INTERRUPT_REASONS

_USER_STOP_WORDING restated the three USER_INTERRUPT_REASONS keys with
their values capitalised, so a new human-stop reason had to be added in
two places or it would render as "Turn aborted — …". Drop the mirror
and capitalise the recorded reason; output is identical for all three
keys (probe-B02-5.py).

Gate rP finding 5 (reuse).

* refactor(agent): share the hidden interrupt placeholder row

message_sanitization.close_interrupted_tool_sequence and the
turn_api_call repetition-loop exit each spelled the same hidden
assistant row literal (content "", display_kind hidden, api_content
_INTERRUPTED_PLACEHOLDER), and the conversation loop's redirect built it
field by field. Add hidden_interrupt_placeholder_row() to the leaf
placeholders module and use it at all three sites so the shape cannot
drift between them.

Gate rP finding 6 (reuse).

* refactor(agent): import the interrupt placeholder at module level

agent_runtime_helpers (two sites) and message_sanitization imported
_INTERRUPTED_PLACEHOLDER inside the function body. The placeholders
module is a leaf with no imports, so there is no cycle to dodge; import
it once at module level (each module still imports cleanly in a fresh
interpreter). conversation_loop keeps its local import.

Gate rP finding 7 (quality).

* docs(agent): say skipped-result content uses str.replace for {name}

The _append_skipped_tool_results docstring still said content "is
formatted with {name}", but the body substitutes with str.replace so a
recorded interrupt reason containing braces renders verbatim. State
that, so nobody "simplifies" it back to str.format.

Gate rP finding 8 (quality).

* test(agent): move the brace-rendering check after the attribution scenarios

The verbatim "{x}" wording check sat between the batch-guard timeout and
the user-stop scenario, mutating _tool_interrupt_reason mid-story. Move
it to the end, alongside the other wording assertions.

Gate rP finding 9 (quality).

* refactor(agent): hoist the hidden placeholder row import and say why a timeout carries no message

Gate rQ quality Lows: the placeholders module imports nothing, so conversation_loop can import it at module level like turn_api_call; the terminal-approval comment now says callers re-queue _interrupt_message, so a system stop must not set one.

* fix(cron): fire an outage-skipped one-shot whose fire claim never landed

Follow-up to #133283 (greptile review). The due scan stamps a one-shot's
run_claim; when the later fire-claim save fails the run never starts. The
next saved scan skips it on that live claim and ends the recovery window,
so once the claim expires _retire_expired_oneshot saw a stale claim past
grace and skipped it forever: neither fired nor retired.

The scan now marks a run claim it stamps past grace (only an outage lets a
one-shot through there) with outage=true; such a claim without a fire claim
keeps the job covered, so it fires once the claim expires. A fire claim
still means the run may have started, so at-most-once is unchanged.

* fix(cron): report an unwritable tick lock in cron status and doctor

Follow-up to #133283 (greptile review). When only cron/.tick.lock is
unwritable (e.g. root-owned after a sudo run, in a writable cron dir) the
tick degrades and returns 0, so the ticker records a success; probe_store
checked only jobs.json, so `hermes cron status` said jobs would fire and
`hermes doctor` reported a writable store while nothing ran.

probe_store now treats an existing tick lock it cannot write the same way
as a read-only jobs.json target: both status and doctor surface it with the
usual fix hint.

* fix(gateway): send cron store notices for a symlinked cron/ directory

Follow-up to #133283 (greptile review). Store records are keyed by the
resolved store path, but the notice sender rebuilt the owning profile from
that path's parent. A profile whose cron/ is a symlink to another disk has
a foreign parent, so neither its outage nor recovery notice reached any
home channel; forget_homes had the same mismatch and never dropped such a
departed profile's record.

Both now match by resolving each served home's own cron/ path.

* fix(cron): keep the outage marker when a skipped one-shot's dispatch fails

clear_run_claim (the dispatch-failure path: interpreter shutdown, execution
creation failure, pool.submit failure) set run_claim to None. The claiming scan
had already ended the recovery window, and a restart drops the in-memory outage
records, so the next scan saw a past-grace one-shot with no claim and retired it
as missed: the job the outage skipped never fired.

Keep {"outage": True} instead. With no "at" it is a stale claim, so the job is
re-dispatched on the next tick, and _retire_expired_oneshot keeps it. At-most-once
is unchanged: the fire_claim CAS still fences every re-dispatch.

Gate r1 (PM12, 2c) Low finding.

* fix(cron): drop a departed profile's store state by its resolved store path

forget_homes rebuilt each store path from a departed home KEY
(Path(key) / "cron"). cron/AGENTS.md forbids rebuilding a home from a key
(hermes_home_key normcases), and the realpath only follows a symlinked cron/
while the home still exists: a profile whose home was deleted kept its degraded
record, holding the host-wide hermes.cron.store.writable gauge at 0.

register_ticked_homes now records each home's resolved store (new public
store_health.store_key) while the home exists and hands the departed stores to
forget_homes. gateway/cron_store_notices.py uses the same store_key against
record.store (already resolved), so one normaliser decides "is this the served
home's store" everywhere.

Gate r1 (PM12, 2c + quality + 2ab toothless) Low finding.

* test(cron): pin that the tick-lock probe names the lock file

The tick-lock case only asserted probe_store() returned an error, which any other
probe failure would also satisfy. Assert the error names .tick.lock so the case
proves that branch fired.

Gate r1 (PM12, quality) Low finding.

* test(cron): derive the run-claim expiry step from the TTL helper

timedelta(minutes=31) silently encoded ONESHOT_RUN_CLAIM_TTL_SECONDS = 1800; use
_oneshot_run_claim_ttl_seconds() + 60 as tests/cron/test_jobs.py already does, so
a TTL change cannot leave the step inside the live-claim window.

Gate r1 (PM12, quality) Low finding.

* fix(cron): keep store state another still-ticked profile shares through a symlink

Two profiles whose cron/ resolves to one store share a store key; when one left, its departure dropped the degraded record the remaining profile still needed (gauge flipped to writable, a second outage notice, and the earlier outage start lost for one-shot cover). Departed stores now exclude the ones still ticked. The notices test pins it (red without the set difference).

* refactor(cron): name forget_stores for what it takes; fix the clear_run_claim early-return comment

* test(cron): pass encoding= to every text read/write in the store-outage tests

* feat(telemetry): updates that stop before applying say why

Every pre-apply exit of `hermes update` records one closed token on the
receipt at the exit itself (update_receipt.record_stop_reason ->
stop_class), and shared_metrics_update.update_failure_class maps it to
the run's failure_class. aborted_before_apply stays only as the fallback
for an exit with no recorded reason.

Exits that fire before the receipt opens (update-lock refusal, Git
operation in progress, managed install) write no receipt, as before, but
now record a metrics-only row with the same closed class.

The parked receipt copy keeps stop_class, exit_code, the stop reason's
leading label and the restart/user-action flags, so a parked run
classifies the same as the in-process one. A run that committed and is
owed only an interrupt or the user's parked changes reads outcome
partial (failure_class interrupted / local_changes_parked), never failed.

* fixup: keep main.py and update_failure_class under their health caps

* test(telemetry): pre-apply exit classes, parked parity, receipt-less refusals; docs + smoke

- test_shared_metrics_update_stop_reasons: every stop token reads as its own class in-process
  and through the parked copy (no free text kept), unknown tokens and post-apply tokens fall
  back, a committed run that only owes parked changes is partial; the lock and Git-operation
  refusals each record one row and leave the holder's latest.json byte-identical.
- test_shared_metrics_install_failures: a PermissionError stop reason now reads its own class
  permission_denied (was os_error); ENOSPC reads disk_full; other OSErrors keep os_error.
- relay-shared-metrics.md: the new classes, partial outcome and parked fields.
- relay smoke: the failed update receipt names its fetch failure and the row asserts fetch_failed.

* fix(telemetry): review findings on pre-apply stop reasons

- The parked copy keeps a stop reason's label only when it is an exception type name (plus its
  errno token) or one of Hermes' two fixed phrases; any other text before a colon reads '-'
  (a free phrase like 'my project name:' used to survive).
- _exit_after_failed_branch_switch's branch_missing/checkout_move_failed probe cannot raise, so
  the exit it names is unchanged even when git refuses.
- hermes update --plan/--check/--list-venv-holders on a managed install records no update run.
- Docs/contract say 'before the apply stage mark' and name the tokens that fire after git moved
  and restored the tree; tests drive git_error_stop_class and _zip_stop_class directly.

* fix(telemetry): pre-receipt exits keep their initiator; index.lock permission errors; bare-interpreter updates park

- record_stop_without_receipt reads the update marker raw and tags initiator=desktop only when the
  marker names this run's hand-off (delegate line = us, or HERMES_UPDATE_HANDOFF_PID / our parent),
  so a Desktop-started early refusal reads kind=desktop and a CLI run refused by someone else's
  update stays cli.
- _move_checkout_to classifies its git error with the shared classifier (git_output_stop_class):
  index.lock 'File exists' -> git_index_locked, 'Permission denied'/EACCES -> permission_denied
  (now a stop class), else checkout_move_failed.
- _publish_shared_metrics: an interpreter that cannot import the config reader (the -I -S bootstrap
  that finalizes a dependency-preparation failure) never emits; it parks the bounded receipt in
  pending_updates unless it can tell collection is off (loaded config says off, or no config.yaml),
  and the next normal start applies the existing gate (report_pending_updates / begin_process purge).

* docs(telemetry): how update dashboards count a partial run

* test(desktop-update): a hand-off log read racing Add-Content retries, not fails

test_script_killed_before_publishing_the_delegate_runs_no_update failed on
the Windows arm64 lane with PermissionError reading
logs/desktop-update-handoff.log: windows.ps1 appends each line with
Add-Content, and a poll that lands while that write holds the file is a
Windows sharing violation, not a missing line. The poll helper now
retries the read (5 s budget) instead of failing the test on it.
Same test is green on main's latest arm64 run; nothing in this PR
touches the desktop-update scripts.

* test(e2e/desktop): give the build-fail updater its manual-outcome grace before asserting it exited

The build-fail spec polls 30 s for posix.sh to exit after the relaunched
app reports the outcome. That outcome is "manual" (the Desktop build is
owed), and finish() then runs launch_app (1.5 s acceptance) and
stop_ui leave-window, which sleeps the 15 s shim grace before tearing the
UI down: ~13 s of margin, and a slow Linux runner overran it (posix.sh
still alive at the 30 s mark, run 37619760207). The same code passed this
spec on bb177627828 and c4c2d9d8b54; the head that failed differs only by
a Windows test helper. 90 s keeps "no updater is left running" a real
check without racing the grace window.

* fix(cron): live-owner stale-claim reclaim measures silence, not run length

The #115692 reclaim released any running attempt whose claim was older than
max(3 x HERMES_CRON_TIMEOUT, script timeout, 2 h), even when the owner was
healthy and busy. Every cron job that legitimately ran longer than that was
marked `unknown` mid-run, its fire claim yanked from under it and the run
aborted with "lost its durable fire claim ownership" (two weekly jobs on one
install in a single day).

The run monitor now stamps `progress_at` on the execution row every 60 s
while the agent reports recent activity, and the sweep measures the stale
bound from that stamp (claim age for legacy rows). A wedged worker stops
stamping and is still reclaimed once the bound passes.

Shape: the stamper and the inactivity watchdog loop (which reads the same
idle signal) live in a new topical sibling cron/scheduler_liveness.py, so
cron/scheduler.py shrinks (4457 -> 4436 lines) instead of growing.

* fix(tools): normalize foreground terminal heartbeat

* fix(terminal): foreground heartbeat is dropped, not refused; notify intent still is

Follow-up to the cherry-picked #119201 (@KoNit-K): keep main's relocated
test file, flip the one invariant that pinned the refusal (a foreground
call with heartbeat>0 now runs with heartbeat=0 while notify=true stays
refused), reword the refusal so it no longer names heartbeat, and say in
the schema that heartbeat is ignored on foreground commands.

Why: the schema fix (4317ed0e71) only helps providers that materialize the
schema default. Models that copy a background call shape still send
heartbeat=60 on ordinary commands; this week 46 subagent sessions hit the
refusal 4 times each (173 refusals), 38 of them never ran a single shell
command, and the agents burned $316 / 1,498 model calls on retries.

* fix(desktop): setup cards stop redrawing in a loop

useSetupRows built a new row array on every render for the app-owned
lists (accent, layout, theme). That recreated the card's stage callback,
whose effect wrote $setupChooseStages, which re-rendered the card. The
rows are memoized on their inputs now, and stageSetupChoose skips the
atom write when a patch changes nothing.

* fix(onboarding): the first-task handoff follows a renamed owner profile

The setup marker saves the profile setup was made from. After that
profile was renamed, the saved name pointed at a missing directory, so
start_chat and reset targeted nothing. The owner now resolves through the
rename history (profile.yaml previous_names) that rename already records.

* fix(free-tier): a short free-tier 429 no longer blocks for a whole quota window

welcome_refusal_from_headers took the hourly bucket's reset as the wait
even when that bucket had requests left, so Retry-After: 4 with a
healthy bucket became a 2,000 s rate_limited refusal and tripped the
shared breaker. The wait now comes from an empty bucket's reset, else
Retry-After, else the body's wait.

* fix(desktop): a sign-in offer from the previous backend never opens

requestGateway keeps one identity across connection and profile
switches, so a free_tier.claim_nudge reply or status read that landed
after a switch could open the offer, or overwrite $freeTierStatus, for
the new backend. The offer timer, the claim and the finished-turn
re-read are tied to gatewayActivationEpoch() and drop a reply once the
route changes.

* fix(auth): persist profile single-use OAuth credentials when root store is empty

Fixes #103694: When a named profile runs auth add for a single-use
refresh provider (e.g. Anthropic, Codex, xAI) and the root store has
no existing rows for that provider, load_pool() sets _borrowed_root_ids = set().

Previously, add_entry() checked if borrowed_ids:, which evaluated to
False on an empty set. It fell back to _persist() -> persist_pool_entries(),
which routed to the update-only _update_root_pool_rows() on root. Because
root had no rows to update, the credential was dropped with no error and
never written to the profile's auth.json.

Checking if borrowed_ids is not None: ensures that profile-scoped pools
persist their new credential directly to the profile store via write_credential_pool(),
transitioning the profile to owning rows for that provider.

(cherry picked from commit b7467a2999832749f540d86530cf463ca0945b04)
(cherry picked from commit 9c604999e233cf6d74a83e4d7c141785b341b0d6)

* fix(auth): persist first profile OAuth row with empty root

(cherry picked from commit 5293489825af98e3e3f0e6cb485ffffa72d14833)
(cherry picked from commit 4144e14bfa7c021a92db5b0f66a75e2e0ac767d1)

* fix(auth): save a profile's first single-use login, and never print Added for an unsaved row

- Decide the profile claim when the row is written (the check from #101827)
  and drop the None sentinel from #103712. On current main load_pool() leaves
  a profile that borrows from an empty root at the default, so the sentinel
  alone never reached the claim path (its test stays red).
- add_entry() re-reads the store after writing. A missing row raises
  CredentialNotSavedError; `hermes auth add` exits with that message instead
  of printing "Added", and the dashboard pool endpoint returns it as a 400.
- Tests: the two contributor tests become one invariant over every
  SINGLE_USE_REFRESH_POOL_PROVIDERS provider (the row is readable, root
  auth.json is byte-identical), plus the CLI success-line contract.
- Docs: profiles page, EN and zh-Hans.

Fixes #103694

(cherry picked from commit 0e250a005a3353564642ce70f100d89174a583f2)

* fix(auth): the forked-grant heal keeps a profile's own Anthropic login

`hermes -p <profile> auth add anthropic --type oauth` saves a
`manual:hermes_pkce` row in the profile. On the next load_pool(), the
forked-grant heal took every row whose source ends in `hermes_pkce` for a
copy of root's .anthropic_oauth.json whenever root's auth.json held no
Anthropic OAuth row. It moved the profile's token pair into root's file
and deleted the row from the profile: the CLI printed "Added", the login
was gone on the next turn, and root's file now held the profile's
account.

Only a `hermes_pkce` row comes from that file. `manual:hermes_pkce` is
owned by the pool and never written to the singleton
(_commit_anthropic_rotation matches the exact source for the same
reason), so the heal now matches the exact source too. The heal's own
docstring already promises to leave an independent `auth add` grant
alone.

This covers the empty-root case the previous commit enables, and main's
existing variant where root holds only an Anthropic API-key row next to
its .anthropic_oauth.json.

Co-authored-by: Enough <10966420+Enough1122@users.noreply.github.com>

* fix(auth): a profile's auth add never copies root's login into the profile

Root's login for nous, openai-codex and xai-oauth can live only in its
auth.json `providers.<id>` block, with no pool rows yet. In a profile
that borrows from root, load_pool() seeds it as a `device_code` entry
through the root fallback. The write-time claim in add_entry() kept
every entry except root's pool rows, so the profile's first `auth add`
wrote root's single-use refresh token into the profile's auth.json: a
forked grant (#100339) that the heal cannot match, because root has no
pool row for it.

- add_entry() writes only the added row when a profile claims a
  single-use provider. Rows the profile seeds from its own sources are
  seeded again on the next load.
- _seed_tokens_singleton() (openai-codex, xai-oauth) no longer seeds
  root's providers block into a profile that owns its own rows, the
  rule _seed_nous_singleton already applies. On current main such a
  profile copied root's refresh token into its own pool on every load.

* fix(telemetry): update stage rows count the stage a failed run died in

hermes.update.run blamed 971 failed CLI runs in Sep 28 - Oct 5 on the apply
stage while hermes.update.stage held zero failed apply rows. Stage marks are
END marks: every pre-apply exit (fetch, channel, branch, merge, HEAD checks)
leaves no mark for the stage it died in, so the run row inferred the stage and
no stage row said it failed. A failed run whose failed_stage names a stage it
never marked now gets one failed row for it.

Rebuilt on the updater overhaul's followups model (#132386): a committed run
that owes follow-ups is a run-level success and its failed stage already has
its own mark, so only failed runs get the row; refusals never do.

* test: detached-writer custody tests read a pid from a beat that can never be empty

Mechanism: the detached build writer in the two "keeps custody until a
detached writer is gone" tests rewrote its heartbeat file in place with
pathlib.write_text, which opens with O_TRUNC and then writes. Custody
SIGKILLs that writer at an arbitrary instant; when the kill lands between
the truncating open and the write, `beat` is left empty and the teardown's
`beat.read_text().split()[0]` raises IndexError (suppress only covered
ProcessLookupError/ValueError). CI hit it on
test_a_ctrl_c_keeps_custody_until_a_detached_writer_is_gone[verbose-slow-cleanup]
(job 112764854735).

Repro: scratch copy of the writer with a 0.1 s sleep between open() and
write() -> both ctrl-c cases and the group-kill sibling fail with
IndexError at the teardown read on every run (5/5).

Fix: one shared _heartbeat_writer helper writes each beat to `<beat>.tmp`
and os.replace()s it into place, so `beat` always holds a complete
"<pid> <n>" once it exists. With the same injection moved onto the temp
write, all three cases pass 5/5. The ValueError suppress that papered over
a partial read is dropped.

* fix(tests): runner applies passthrough --ignore-glob; windows marker test stops sitting out 150s

Windows-only tests job 112764655375 (run 37612501107, a YAML-only catalog PR)
went red with "667 passed, 0 failed": test_desktop_update_windows_marker.py
hit the runner's 300s per-file cap mid-file ("(300s exceeded; process tree
SIGKILL'd)"), so its 19 tests were never counted and it was listed under
"no tests ran". Run 37612584210 died the same way. Clean runs of this file
take 220-300s (arm64 hit 296.5s and 299.8s), so it is at the cap on every run.

Two defects:

1. The desktop_updater lane gate never worked. tests-os.yml passes
   --ignore-glob='*test_desktop_update_windows_*.py' after `--`, and
   run_tests_parallel.py forwards it to `pytest <file>`. pytest applies
   --ignore/--ignore-glob only while recursing directories, never to an
   explicit file argument, so the flag did nothing. These hand-off tests ran
   on every PR, including ones that never touched the updater, even though
   the job printed "desktop_updater lane off: skipping ...". The runner now
   applies --ignore/--ignore-glob to its own file list (pytest's fnmatch on
   the absolute path, relative patterns anchored at the repo root, which is
   the per-file pytest's cwd) and says how many files it excluded.

2. test_desktop_that_never_exits_is_not_relaunched_over waited out the
   production 150s Desktop-exit ceiling. That one wait is half the file's
   budget. The script already honours HERMES_UPDATE_DESKTOP_EXIT_SECONDS
   ("so the self-tests need not sit it out"), and the posix twin sets it to
   2. Set it to 5 and assert that the script's refusal names 5s, so the
   override is proven to reach the wait. The verdict under test (refuse
   with 4, never relaunch) is the same at any ceiling.

Repro (Linux, CI-shape args):
  python scripts/run_tests_parallel.py --files "<windows_marker>:<windows_cwd>:<other>" \
    -- --ignore-glob='*test_desktop_update_windows_*.py' -m "platforms and not integration"
  base:  "Running 3 test files", both gated files run
  fixed: "note: --ignore/--ignore-glob excluded 2 test files", "Running 1 test files"
New test_passthrough_ignore_drops_files_the_runner_hands_pytest_explicitly
(--files and discovery x glob/glob-spaced/path) fails 6/6 on base and
passes 6/6 fixed.

* fix(desktop): show a queued cron run row immediately after trigger

Run History only rendered sessions the backend had already materialized,
so an accepted Trigger now was invisible for the tens of seconds the
scheduler took to create the run session (#70826). The queued row is
painted from the click, settles when Run History observes a run that
started after it (any load — a fast 1s poll while queued, the cron.changed
broadcast on event-capable backends, or the visibility poll), and drops
at a bounded 90s timeout so a never-materializing run cannot wedge the
button. A failed request drops the row immediately. Pause deliberately
does not settle it: the pause response only persists enabled=false and
never cancels an already-claimed execution.

Queued design (queued-run row + polling + settle semantics) originates in
#70840 by @nv-cho; this ports it onto current main's
createCronTriggerController wiring.

Fixes #70826

* refactor(desktop-i18n): move billingBlock copy into en_billing sibling

Offset the cron queuedRun string growth in the over-cap en.ts/types.ts
facades per the code health ratchet: moved code keeps its cap.

* chore(desktop): refresh locales/_keys.desktop.json for cron.queuedRun

* fix(update): sweep stale git locks at the start of the run, not only before the fetch

The stale-lock / aborted-pack sweep ran only in the apply path, after the
snapshot. A run that died in that window left `.git/index.lock` behind with
nothing in the next run to remove it, so the following update failed with
"File exists" until an operator cleaned it up by hand. The sweep moves to the
top of `_cmd_update_impl`, ahead of the plan and the backup;
`clear_git_debris` already wraps locks and aborted-transfer pack temps (and
the partial-clone maintenance keys), and `_sweep_stale` skips everything
while a git process holds it, so the earlier position is safe.

(Salvaged from #132089 work; the receipt stop-reason half of the original
commit is now owned by the overhaul's receipt API and is not carried here.)

* fix(update): a killed git's index.lock no longer blocks the next update

`.git/index.lock` left by a killed git refused every later merge with
"File exists": the stale-lock sweep keeps any lock younger than 10 minutes,
so a `hermes update` killed mid-run (Windows crash cell `mid_fetch`, run
37573293632) or a version probe killed by its own 3 s timeout (#132089)
failed the next update at "Pulling updates...".

- `gitlock.release_dead_index_lock`: drop the lock as soon as the launch
  repair's ownership proof (`_early_recovery._release_dead_index_lock`:
  /proc fds + lock-keeping gits on Linux, lsof on macOS, the unlink itself on
  Windows) shows no live git holds it. An interrupted tree move keeps its own
  lock judgement. `_cmd_update_impl` calls it first, under the update lock.
- `version_info._git_version_info`: the read-only status probe runs as
  `git --no-optional-locks status`, so a timed-out probe strands no lock.

Part of #132089 (the stale index.lock half; naming why a pre-apply exit fired is separate work)

* fix(update): on Windows a killed git's index.lock waits only for gits working in this checkout

Windows has no /proc to name a lock's holder, and its unlink-as-probe cannot see a
lock-keeping git with its fd closed (a `git commit` in the editor), so the reclaim keeps the
lock while a git.exe has its cwd or a path argument inside the checkout, and whenever the
scan cannot tell (no psutil, an uninspectable git). Any other git on the machine no longer
blocks it, unlike the age-floor sweep's machine-wide check.

* fix(update): only a killed update's index.lock is reclaimed at once

A lock nobody holds open is not proof its git is dead: a git that just created it may not have
opened it yet. The start-of-run reclaim now takes .git/index.lock only when the latest update run
died unfinished (receipt running/interrupted, owner gone) and the lock is newer than that run's
start. Every other lock is left to the age-floor sweep, and the update refuses on it truthfully
(e2e-upgrade test_live_index_lock_is_refused_truthfully went red on fd2214b0ad6).

* fix(update): the start-of-update index.lock reclaim never takes a live git's lock

Review round 2 on #133077. A killed update's receipt plus a younger lock let
release_dead_index_lock delete the lock of a live `/usr/lib/git-core/git-commit -a`
waiting in its editor: the lock fd is closed there, and the Linux holder scan only
knew `comm == git` with an argv subcommand, so the dashed form read as "no holder".
A competing `git add` then got in and the user's commit died "unable to write new
index file".

- The receipt gate stays necessary, never sufficient. The reclaim now also needs
  proof that no git of ANY form works in this checkout (`_held_open(any_git=True)`):
  cwd, path arguments (relative ones against the git's cwd) or GIT_DIR/GIT_WORK_TREE/
  GIT_INDEX_FILE inside the checkout or its git dir, compared by path components.
  A git we cannot place (unreadable cwd/environ, hidepid) keeps the lock.
- Git forms: `git`, a dashed `git-<sub>` from git-core, `git.exe`/`git-<sub>.exe`,
  `-C`/`-c`/`--git-dir` globals, aliases (an unknown subcommand counts). The
  launch-time repair's own scan inverts its list: only known readers are exempt.
- macOS: an empty lsof is no proof any more; `ps` lists every git (comm, so dashed
  forms too) and any lock-keeping one keeps the lock.
- Windows: the psutil scan matches dashed and .exe gits, reads environ, and compares
  by path components (`hermes-backup` was read as inside `hermes`).
- A failed retry no longer erases the evidence: `_killed_update_owns` reads the durable
  per-run records (read-only) and asks about the last run that STARTED before the lock
  was written, so a later failed run cannot discharge the killed one.
- The FIFO test helper kills its blocked git when its own setup assertion fails.

* fix(update): on macOS only a git working in this checkout keeps a dead index.lock

The lsof branch's ps scan counted every lock-keeping git on the machine,
because ps cannot say where a process works. A git commit open in any
other repository (or a sibling test worker's git on CI) therefore kept a
dead index.lock and blocked the launch-time repair after a killed update:
the macOS-only lane failed
test_a_merge_killed_writing_hermes_constants_is_repaired_by_the_next_launch
with "a running git holds the index".

Each candidate git's cwd now comes from `lsof -d cwd`, compared to the
checkout by path components (as on Linux); a candidate whose cwd cannot
be read while it still runs makes the answer unknowable, so the lock is
kept. Probe with real lsof/ps and a dashed git-commit waiting in its
editor: in another repo, 4b3c28c816c named it a holder and the fix does
not; in the checkout, both keep the lock.

* fix(update): the macOS lock scan finds ps when the launcher's PATH has none

The launch-time repair runs with whatever PATH the launcher got; the
macOS-only lane runs it with a PATH holding no git (as on the Windows
install that motivated the repair), so the bare "ps" call failed, the
scan answered "unknowable" and the dead index.lock of a killed update
was kept: test_a_merge_killed_writing_hermes_constants_is_repaired_by_the_next_launch
stayed red. ps is now resolved like lsof already is (which, then /bin/ps,
/usr/bin/ps). Probe: dead lock, PATH with neither git nor ps, at
ba2965261a0 the scan returned None (lock kept); now False (reclaimed).

* fix(update): review the macOS lock scan's lsof/ps lookups; keep it stdlib-only

The resolution-allowlist ratchet flagged the lsof lookup (moved from
_held_open to _held_open_lsof) and the new ps lookup in _ps_git_holder:
both are the stdlib-only launch-time repair asking the OS's own tools,
with an absent tool meaning "unknown" and the lock kept. Rows renamed
and added with that justification.

_still_running also imported psutil inside the launch-time repair, which
runs from hermes_bootstrap before app dependencies can import; it is now
a POSIX os.kill(pid, 0) probe (the lsof/ps branch never runs on Windows).

* fix(desktop): a setup card's typed-pick callback keeps a stable identity

The name card crashed the setup chat on a real Windows bundle (React #185,
maximum update depth): the card wrote its stage callback into
$setupChooseStages, which it also reads, and the callback depended on
rowLabels, a new object every render. 40512011b6 stopped the rows from
changing identity; this removes the dependency itself, so the callback
changes only with kind, mode, multi-select or the request. The labels
effect already stages the row names the callback used to rewrite.

The new test renders the real card: red on 17646902e1 (the crashing
build), green here and with either guard alone.

* test(ci): the CI replay names a silent step's exit code

workflow_steps.py is a helper module, not a test module, so pytest does not
rewrite its asserts: a replayed bash step that exits non-zero with empty
stdout/stderr fails as a bare 'AssertionError' with no exit code. Two
Windows arm64 os-tests runs (jobs 112635777217, 112764629268) failed this way
on unmodified baseline replays, and the NTSTATUS needed to tell a crashed
child (0xC0000142 / 0xC0000005) from a bash failure was lost. Put rc (decimal
and hex), the bash used, and both streams in the message.

* test(ci): the CI replay does not run Git Bash under x64 emulation on Windows arm64

Mechanism: Git for Windows on arm64 ships native git.exe but an x86-64 MSYS
usr/bin/bash.exe, which runs under the WoA x64 emulator. With the os-tests
lane's 16 parallel workers, a replayed step's bash dies before running a line
with 0xC000026F (STATUS_WX86_INTERNAL_ERROR) or 0xC0000005 and empty
stdout/stderr, so an unmodified baseline replay fails as a bare
AssertionError (jobs 112635777217 and 112764629268, two different nodeids).

Repro (throwaway workflows on the real runners, previous commit's diagnostics):
- windows-latest-32-arm-core, 16 copies of the file's replay tests at once,
  4 rounds: 21/64 copies red, every failure 0xC000026F.
- Same load, windows-latest-32-core (x64): 0/64.
- Passing SYSTEMROOT/WINDIR or the full Windows env to the step: 12/48 and
  17/48 vs 15/48 base, so the stripped env is not the cause.
- One process, 32 threads, 600 replays: 0 crashes; it takes concurrent
  processes, i.e. host-wide emulator load.

Seam: _NATIVE_WINDOWS_TOO keeps platforms("linux", "windows") and adds a
skipif for win32 + native_arch() == arm64. The x64 Windows row still runs the
replay on the Windows layout (python.exe-only venv, Git Bash), Linux arm64
keeps it (native bash), and the routing guard that pins the marker passes.

* fix(free-tier): a setup chat turn does not start the sign-in offer clock

The setup chat's last turn (the fork answer and the handoff) finished clean
on the free tier, so it was recorded as the first task. The offer then came
due three minutes after setup ended, which lands in the middle of the first
real task when that task runs longer than three minutes. Seen live on a
Windows dev build: first_task_at was the setup chat's end, not the task's.
Turns in the setup profile's sessions are no longer tasks, as the offer
module already intended ("normally the task setup hands off to").

* i18n(desktop): translate cron.queuedRun in every overlay

Main added the Run History 'Queued run' label (44c1ad7f84) in English only;
this branch's overlay-completeness ratchet fails on a new untranslated key.

* Fix reasoning markup rendering in dashboard

* fix(web): scope structured reasoning UI to assistant messages

Only assistant content with known reasoning tags uses StructuredReasoning.
User and tool messages with literal <action>/<result> tags stay on Markdown.

* test(web): behaviour coverage for structured reasoning transcript rendering

RED on main: the expanded session transcript renders assistant
reasoning with literal <thinking>/<action>/<result> wrapper tags.
Keeps the user-role invariant (literal tags stay on the Markdown path).

* refactor(web): move session source config to SessionsPage_sources sibling; satisfy code-health ratchet

* fix(desktop): hide unknown cloud agent status

Omit the gateway status description when cloud discovery cannot provide a meaningful state, while preserving valid status labels. Add renderer coverage for both unknown and known states.

* test(desktop): cover empty cloud agent status

* fix(pm): reclaim dependency state of deleted checkouts (pm gc + startup worktree prune)

Every checkout that boots Hermes commits its own dependency state under
<home>/installs/<sha256(path)[:16]>/ (venv, test venv, PM runtime; ~200 MB
each) and nothing removed it when the checkout went away. On a host that
runs `hermes -w` campaigns the directory held 499 entries / 92 GB, of which
402 (80 GB) belonged to worktrees and scratch clones that no longer exist.

The state dir already records its checkout in inputs/.project-root. A new
pm/install_states.py lists entries whose recorded checkout is gone and
removes them unless something could still be reading them (install lock
taken, or a generation lease held). `pm gc` runs it and reports the count;
the startup worktree pruner runs it after its own reaping so a removed
worktree's state follows the tree instead of outliving it.

* fix(pm): orphan reclaim runs on every prune pass and keeps checkouts it cannot see

Follow-up to the review of this PR:

- The startup pruner returned early (no .worktrees/, or no tree past 24h) before the
  reclaim ran, so a tree removed by hand kept its state until some other tree aged out.
  _prune_stale_worktrees now runs the reclaim in a finally around the prune phases.
- A missing recorded path is not proof of deletion: a container sharing the data root
  (-v ~/.hermes:/opt/data) cannot see the host's checkout, and gc there deleted the
  host's live dependency environment. A state is an orphan only when its checkout sat
  under the data root that owns installs/, or directly in a .worktrees/ dir that still
  exists.
- PM runtime generation leases (pm-runtime/generations/*/.leases) now count as held,
  so a worker still running from a deleted checkout's runtime keeps it.

* fix(pm): an interrupted orphan reclaim leaves a dir the next pass still finds

The startup reclaim runs on a daemon thread, so a hermes exit can cut a removal
short, and Windows refuses to unlink open files. shutil.rmtree takes children in
scandir order, so inputs/.project-root (the only record naming the checkout)
often went first and the remainder was invisible to every later pass.
Everything else is removed first; the record goes only once nothing else is left.

* fix(kanban): only sweep a parent workspace once the parent itself is terminal

The deferred parent sweep added for #33774 reaps a parent's scratch or
worktree workspace as soon as it has no active children left, without
checking the parent's own status (#133501). Archiving or finishing the
last active child of a parent that is still running/blocked/review
deletes the parent's live workspace from under its worker; the shared
workspace guard does not catch it because _OTHER_LIVE_PATHS_SQL excludes
the task being cleaned, i.e. the parent itself.

Gate the sweep on the parent being terminal (the same status set as
_ACTIVE_CHILDREN_SQL), so a live parent keeps its dir and reaps it at
its own terminal transition. #33774 semantics are unchanged: a parent
that finished while children still needed its handoff files is still
swept when the last child drains.

(cherry picked from commit 81f6588bbc5346097ab17e3660636c8a3be7319a)

* fix(mcp): reject fabricated DCR for providers without RFC 7591 registration (#78190)

Google's hosted Gmail/Drive MCP servers advertise no registration_endpoint
in their ASM, so the mcp SDK's create_client_registration_request falls
back to guessing POST {server_origin}/register. Those servers answer with
an opaque 404 (OAuthRegistrationError: Registration failed: 404) and the
gateway burns the reconnect ladder on an unrecoverable failure while the
CLI looks healthy because tools/list is unauthenticated on Google servers.

- Guard: HermesMCPOAuthProvider intercepts the fabricated registration
  POST in the bidirectional auth_flow bridge and raises an actionable
  OAuthRegistrationError (create an OAuth client, add oauth.client_id/
  client_secret under mcp_servers.<name>.oauth in config.yaml) instead of
  sending the request. Fires only when the URL is the exact SDK fallback
  AND the ASM lacks a registration_endpoint, so servers advertising a
  real RFC 7591 endpoint pass through untouched.
- Surfacing: needs_reauth tool errors now include the registration
  guidance, and the generic 'MCP call failed' path unwraps TaskGroup
  exception groups so the root cause reaches the model instead of
  'unhandled errors in a TaskGroup (1 sub-exception)'.
- humanize_oauth_registration_error: 404 registration failures get the
  same pre-registered-client guidance as 403 (CLI login path).

Tests: full discovery dance (401 -> PRM -> ASM without registration_endpoint
-> fabricated POST) drives the bridge and asserts the guard raises before
any network send; humanize 404 + passthrough cases; _auth_error_detail and
group unwrap surfacing.

Rebuilt on current main, where mcp_tool.py was decomposed and mcp 2.0
rides on httpx2: the surfacing helpers live in tools/mcp_tool_errors.py
and feed the shared _dispatch / _NEEDS_REAUTH_MSG paths in
tools/mcp_tool_handlers.py; the guard compares with the request's own URL
type and only fires once ASM was actually discovered (an undiscovered
ASM keeps main's discovery-context error, #113771), and closes the inner
SDK generator before raising so context.lock is released immediately.

* Report bounded npm probe and terminal GUI installer failures

* fix(ts): npm run fix

* fix(free-tier): HERMES_PREVIEW_FULL_CONNECTORS=1 turns the preview cohort on

The preview flag only read exactly true/false, so a bundle built with
=1 (HERMES_GUEST_ONBOARDING's spelling) silently sent {} and every guest
landed in the external cohort. 1/true now send true and 0/false send
false; anything else still omits the field.

* fix(desktop): explain missing remote profiles

* fix(desktop): lint missing-profile error and move its tests to a sibling file

Keeps remote-lifecycle.test.ts under its code-health line cap and clears
the curly/no-useless-escape errors from check:lint.

* fix(models): TokenHub's saved endpoint is not a relay, so /model keeps the curated Hy list

#134320 gave tencent-tokenhub a profile whose base_url is empty (its auth registry row owns
the endpoint). _configured_relay_base_url decides "relay or canonical" from the profile's
base_url only, so a model.base_url that `hermes model` saved as TokenHub's normal endpoint
read as a relay: the picke…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

P2 Medium — degraded but workaround exists tool/terminal Terminal execution and process management type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: terminal heartbeat schema causes foreground validation loops and notification flooding

2 participants