Skip to content

fix(auxiliary): map Fireworks reasoning disable - #109807

Closed
huklaa wants to merge 1 commit into
NousResearch:mainfrom
huklaa:fix-109774-aux-reasoning
Closed

huklaa wants to merge 1 commit into
NousResearch:mainfrom
huklaa:fix-109774-aux-reasoning

Conversation

@huklaa

@huklaa huklaa commented Sep 13, 2026

Copy link
Copy Markdown

Fixes #109774.

Summary

  • translate Hermes' disabled reasoning config into Fireworks' supported top-level reasoning_effort: "none"
  • prevent the auxiliary client from falling back to the nested extra_body.reasoning object that Fireworks rejects
  • add regression coverage through the real provider-profile projection used by title_generation

Root cause

The Fireworks profile inherited the base no-op reasoning hook. The auxiliary kwargs builder therefore treated it as reasoning-unaware and injected its generic nested reasoning body, which Fireworks rejects with HTTP 400.

Validation

  • uv run --with pytest pytest tests/plugins/model_providers/test_fireworks_profile.py -q — 10 passed
  • targeted Ruff checks passed
  • Python compilation and git diff --check passed

Scope

The change is isolated to the Fireworks provider profile. Other providers and explicit caller extra_body settings are unchanged.

@alt-glitch alt-glitch added type/bug Something isn't working P3 Low — cosmetic, nice to have comp/plugins Plugin system and bundled plugins comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint labels Sep 13, 2026
@huklaa
huklaa force-pushed the fix-109774-aux-reasoning branch 2 times, most recently from 45b7b88 to 4ff6efa Compare September 13, 2026 20:59
@huklaa
huklaa force-pushed the fix-109774-aux-reasoning branch from 4ff6efa to 8cb6483 Compare September 14, 2026 15:19
Signed-off-by: huklaa <129692708+huklaa@users.noreply.github.com>
@huklaa
huklaa force-pushed the fix-109774-aux-reasoning branch from 8cb6483 to 076578e Compare September 14, 2026 23:34
teknium1 added a commit that referenced this pull request Sep 17, 2026
… reasoning_effort

The pre-ladder max_tokens rung accepted payment|connection|rate_limit errors on
its retry, so a 429 after a stripped retry reached the credential and
provider-fallback rungs. The ladder's shared `_param_rung_accepts` dropped
rate_limit: a 429 on the retry raised out of the primary call and skipped the
whole fallback chain. Accept it there too, with a rung-level test that the
stripped kwargs and the 429 are handed on rather than raised.

The body claimed `Fixes #109774` while the aux client still sent the generic
`extra_body.reasoning` to Fireworks on every call and only recovered
reactively (an extra 400 round-trip per request). Port @huklaa's profile
override from #109807: Fireworks documents top-level `reasoning_effort`
(`none` disables thinking), and overriding `build_api_kwargs_extras` marks the
profile reasoning-aware so the transport omits the generic fallback on this
route. Extends the salvage with the effort mapping so enabled-with-effort takes
the same wire, with a control that a profile-less route keeps the fallback.

Salvages #109807 (@huklaa).

Co-authored-by: Hukla <129692708+huklaa@users.noreply.github.com>
teknium1 added a commit that referenced this pull request Sep 17, 2026
… reasoning_effort

The pre-ladder max_tokens rung accepted payment|connection|rate_limit errors on
its retry, so a 429 after a stripped retry reached the credential and
provider-fallback rungs. The ladder's shared `_param_rung_accepts` dropped
rate_limit: a 429 on the retry raised out of the primary call and skipped the
whole fallback chain. Accept it there too, with a rung-level test that the
stripped kwargs and the 429 are handed on rather than raised.

The body claimed `Fixes #109774` while the aux client still sent the generic
`extra_body.reasoning` to Fireworks on every call and only recovered
reactively (an extra 400 round-trip per request). Port @huklaa's profile
override from #109807: Fireworks documents top-level `reasoning_effort`
(`none` disables thinking), and overriding `build_api_kwargs_extras` marks the
profile reasoning-aware so the transport omits the generic fallback on this
route. Extends the salvage with the effort mapping so enabled-with-effort takes
the same wire, with a control that a profile-less route keeps the fallback.

Salvages #109807 (@huklaa).

Co-authored-by: Hukla <129692708+huklaa@users.noreply.github.com>
@teknium1

Copy link
Copy Markdown
Collaborator

Thanks @huklaa — this landed on main via #113958 (cbe2413); your Fireworks preflight was ported with a Co-authored-by credit. Closing this PR as merged-through-salvage; the fix is on main.

@teknium1 teknium1 closed this Sep 17, 2026
Nag112 added a commit to Nag112/hermes-agent that referenced this pull request Sep 18, 2026
* fix(desktop): hide no-op pop-out for file previews

* feat(desktop): add rebindable sidebar grouping cycle

* refactor(desktop): one ordering for sidebar groupings shared by the menu and the keybind

Follow-up to Ayush Nangia's feature: the filter menu kept its own literal
list of groupings while the new cycle keybind walked a second one, with only
a test comment saying they should agree. The menu now builds its options
from `SIDEBAR_GROUPING_ORDER`, so reordering or adding a grouping is one
edit and the keybind cannot drift from what the menu shows.

* refactor(desktop): derive SidebarGrouping from the grouping order; invariant cycle test

`SidebarGrouping` was a literal union kept in step with `SIDEBAR_GROUPING_ORDER`
by hand — a fifth grouping added to the type would compile while the cycle
silently skipped it. The array is now the single source (`as const`) and the
type is derived from it. The cycle test no longer pins the exact order (a
change-detector on the constant); it asserts one lap visits every grouping
once and returns to the start, and still fails when the cycle skips a step.

* chore: map contributor email for #104781 salvage

* fix(approval): anchor launchctl lookaheads to prevent GIL starvation

* fix(approval): deobfuscate every command word in one detection variant (#113535)

_command_detection_variants yielded one FULL-LENGTH variant per quoted or
escaped command word. A heredoc body of quoted lines ("key": "value", …) is
hundreds of quoted command words, so both detection passes scanned
O(words * len) characters: on a 16 KB / 460-line command the hardline pass
alone took 7.8 s and the dangerous pass 15 s even with the launchctl lookahead
anchored (33 KB: 36 s hardline), all while holding the GIL on the gateway loop.

Build a single variant with every command word deobfuscated instead. The
obfuscation catch ($(echo rm), r''m, ${0/x/r}m …) is unchanged — the same
deobfuscated words appear, in one string — and the variant count on the
issue's input drops from 923 to 4 (36 s -> 0.26 s hardline, verdict
unchanged). detect_hardline_command needs no bound and keeps its YOLO
semantics: a benign heredoc is neither blocked nor prompted.

Root cause identified in #113943 by @Tranquil-Flow (variant explosion as the
second compounding factor); its cumulative-work budget is superseded by
removing the explosion.

Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>

* test(approval): pin the launchctl perf subprocess to the checkout under test

A bare `python -c` resolves `tools` through the venv's editable install (the
primary clone), not the worktree pytest runs in, so the regression guard could
silently test a different tree. Pass cwd + PYTHONPATH like
tests/tools/test_async_delegation.py does.

* chore: map schwinguin's commit email for #109497 salvage

* fix(secret_scope): memoise the shared .env tokenizer once per file change

load_env_file() is now the canonical .env tokenizer after c849bc383a collapsed
six hand parsers onto it. Profile scopes, dashboard boundary discovery, skill
secret capture, managed .env and setup readers share the same decoding and
parsing rules, but direct calls still re-read and re-parse the whole file.

build_profile_secret_scope() puts that work on the gateway's hottest paths:
every turn, every cron job, MCP and browser lifecycle adoption, the TUI prompt
turn, and the 60s housekeeping drain of the durable cron delivery queue.
Memoising the shared tokenizer now removes redundant work for the consolidated
readers as well as the motivating profile-scope path.

config.load_env() keeps its existing outer memo over a call that is now itself
memoised; this leaves its menu-render shortcut intact without redesigning it.

Freshness is the hard part, because a stale credential map is a far worse
failure than a slow one. Every load_env_file() call still OPENS the file and
keys on os.fstat() of that descriptor rather than os.stat() of the path:

  * the open preserves close-to-open revalidation on NFS;
  * an open or read failure returns {} without caching it, and drops any warm
    entry so a transient EACCES cannot become a persistent empty profile;
  * the descriptor pins one inode, so a symlink repointed mid-read cannot file
    one file's contents under another file's identity.

The key is (mtime_ns, size, inode, device), re-checked on the same descriptor
after reading. A rewrite through that inode is not stored under the pre-read
fingerprint. Unavailable metadata means "do not cache", never "cannot read".
A generation counter checked under the lock prevents a slow reader from
repopulating an entry that a writer invalidated while the read was in flight.

Read bytes from the open descriptor and preserve upstream's decoding exactly:
strip a UTF-8 BOM before trying UTF-8, then fall back to latin-1 for invalid
UTF-8. Both the cached reader and the uncached reference use that decoder and
the shared text parser. Invalid UTF-8 is a cacheable result, not a read error.
The value and inline-comment helpers remain local to secret_scope.

invalidate_env_cache() clears this memo too, so save_env_value,
remove_env_value and sanitize_env_file invalidate both layers. The per-path
LRU holds at most 64 entries and every caller gets its own dict.

Accepted limitation: a write keeping length, inode and nanosecond mtime all
identical is not detected without explicit invalidation. This is a new
limitation for previously uncached readers. Parsed secrets also remain
reachable in memory between use and eviction for up to 64 paths, where the
outer config memo holds one.

The pre-rebase measurement on a 508-line, 25432-byte .env was 0.868 ms ->
0.011 ms per call. That measurement predates the shared binary decoder.

Add decode/cache invariants for BOM and plain files, latin-1 fallback,
same-size writes switching UTF-8 -> latin-1 -> UTF-8, and agreement with the
uncached reference. All eight decoding, cache and freshness mutations fail
the new tests, and all 40 cache tests pass after restoration.

The requested agent, CLI, cron and gateway suites report 30,168 passed,
479 failed and 303 skipped in the restricted local environment. All 140
failing files were checked against upstream/main at 5eb99eb284: all 479
failed-test IDs and 210 collection/setup/teardown error IDs match exactly.

* fix(gateway): skip heartbeat restore sweep when no heartbeat is persisted

The 5s goals poller calls restore_heartbeat_watches(), which listed every
routed session and entered _profile_runtime_scope for each distinct origin
— a full config.yaml/.env parse plus secrets hydration per session per
scan. With zero heartbeats persisted (the common idle case) the entire
sweep is wasted work; on a 400-session multiplex gateway it dominated
idle CPU (~60% of on-CPU time in YAML parsing).

Gate the sweep on an indexed heartbeat:* prefix probe of each served
profile's SessionDB (HERMES_HOME contextvar only — no config/secret
scopes). Probe failures fail open to the historical full sweep.

* fix(gateway): gate per-profile watcher ticks on actual work in the profile's store

The handoff watcher (2s interval) and loop wakeup watcher (15s) entered
_async_profile_runtime_scope for EVERY served profile on EVERY tick — a
full config.yaml/.env parse plus secrets hydration per entry — even when
the profile's state.db held zero pending handoffs or active loops. On a
multiplex gateway this, not message traffic, was the idle CPU driver
(py-spy: ~75% of on-CPU time in YAML parsing from these two paths).

Both watchers now probe the profile's cached SessionDB first (only the
HERMES_HOME contextvar installed, off the loop thread via the runner's
executor hop) and enter the scope only when work exists:
- handoff watcher: new SessionDB.has_pending_handoffs() existence probe
- loop watcher: list_active_loops() under the probe context

All probes fail OPEN (error/unknown -> historical always-enter
behavior); runners without the executor hop (test stand-ins) skip the
gate, so existing watcher tests are unaffected. The startup
stale-handoff reclaim stays ungated: it runs once per boot and must
also see 'running' leftovers, which a pending-only probe would miss.

* fix(gateway): watcher idle probes fail open on an unavailable store and ignore cleared rows

The loop watcher's probe used ``list_active_loops()``, which collapses an unavailable
SessionDB (and any read error) to ``[]`` — the gate then skipped that profile on every tick
(fail CLOSED). The heartbeat probe keyed on ``heartbeat:*`` key existence, but ``clear`` and
``pause`` keep their rows, so one past /heartbeat re-enabled the full sweep forever.

Both probes now read the rows through one ``_profile_meta_rows`` helper (None = store
unavailable → treat as work present) and parse the status themselves: only an ACTIVE row
opens the gate; a corrupt row is "unknown" and keeps the historical scan. The heartbeat
gate reuses ``_handoff_watch_scopes`` for the served homes instead of a second resolver.

* test(gateway): two invariant tests for the watcher idle gates

Heartbeat: no active row → no sweep; a cleared row → still no sweep; an active row → sweep;
an unopenable store → sweep. Loop watcher: idle profile never scoped, active loop scoped,
unopenable store scoped. The handoff-watcher test from #109497 is dropped: same shape, and
the salvage bar is two tests per fix.

* refactor(secret_scope): key the .env memo on utils.file_signature

The memo had its own (mtime_ns, size, inode, device) tuple and documented a
same-length/same-inode/pinned-mtime rewrite as an accepted gap. utils.file_signature is the
repo's change-detection key (config.load_env, skill_utils, prompt_builder already use it) and
includes ctime_ns, which user space cannot backdate — so the gap closes and one helper goes.

* refactor(config): load_env is a thin wrapper over the memoised tokenizer

load_env() kept its own (path, file_signature) memo on top of load_env_file(), which now memoises
on the same key: two stat calls and two caches to keep coherent for one answer. Drop the outer
layer; invalidate_env_cache() stays as the writers' knob and forwards to the one memo.

* fix(gateway): heartbeat idle gate probes the default profile under a named-profile multiplexer

_watched_homes reused _handoff_watch_scopes, which deliberately omits "default" because the
handoff/loop watchers' root poll covers the launch home. For a `hermes -p work` multiplexer the
launch home is profiles/work, so ~/.hermes was probed by nobody and an active heartbeat there
answered "idle" every tick. Probe the served set the sweep can actually resolve to.

* refactor(gateway): idle gates live in run_idle_gates; row semantics live with loops/heartbeat

The probe helpers had been appended to the gateway/run.py facade and the three watchers each
carried their own copy of the "probe → fail open → offload" shell, with gateway code parsing
loop/heartbeat rows itself. One sibling now owns the gate shape (_gate + off_loop_gate), and
hermes_cli.loops.store_has_active_loop / hermes_cli.heartbeat.store_has_active_heartbeat own
what an ACTIVE row is (heartbeat gains the _META_PREFIX loops already had). One fail-open layer
per read: has_pending_handoffs is a bare bounded query, the gate catches. The loop watcher calls
its executor hop directly — it already did so unconditionally for the scan, so the
"runner without a hop" fallback there was dead.

* test(gateway): idle-gate tests read status, cover the named-routing-home case and the handoff gate

Loop watcher: a cleared row must not open the gate (a key-existence mutation now fails), then
the unavailable-store fail-open. Heartbeat restore: routing home = named profile with the only
active row in the default store must still sweep. The handoff-watcher test from #109497 is
restored (its gate and the has_pending_handoffs query had no coverage).

* fix(desktop): install rebuilt app to system location after CLI hermes update

Port of #59942 onto post-#102117 main. The refactor moved the desktop
helpers from hermes_cli/main.py into hermes_cli/main_desktop.py, so the
new install machinery (_app_asar_hash, _macos_adhoc_sign_bundle,
_desktop_bundle_install_supported, _swap_in_new_macos_bundle,
_install_rebuilt_desktop_app) lives there now, re-exported through
main.py's frozen updater surface block so _m() resolves them.

_rebuild_desktop_after_update (hermes_cli/update_cmd_deps.py) now, on a
successful --build-only build, installs the rebuilt bundle to the
system location (macOS .app via ditto staged copy + atomic swap with
rollback; Windows via copytree), skipping when the installed app.asar
hash already matches. Linux stays owned by its package mechanism.

Without this, a CLI hermes update leaves the installed app stale —
the in-app updater covers its own path, but the CLI path never swapped
the installed bundle.

Tests: 17/17 ported suite (import/patch targets repointed to
main_desktop) + test_cmd_update.py 47/47 green.

* fix(update): scope the installed-bundle refresh to macOS, keep the signature, never swap under a live app

Rework of the salvaged #59942 hook so it fixes the whole #52339 class without
regressing what the Desktop updater already gets right:

- macOS only. The Windows arm rmtree'd a possibly-running NSIS install
  (partial deletion under a lock) and windows.ps1 already owns that swap;
  Linux packages stay with their package manager.
- No re-sign. The rebuilt release/ bundle already carries the stable local
  signing identity from _desktop_macos_relaunchable_fixup; a deep `codesign -s -`
  on the installed copy replaced it with a fresh ad-hoc cdhash and reset every
  TCC grant. ditto preserves the signature, so nothing is signed here.
- Running bundles are reported, not swapped: Electron loads app.asar and helper
  apps lazily, so renaming the bundle away and deleting the old tree crashes
  the live app. The detached updater waits for exit; a terminal `hermes update`
  with the app open now prints what to do instead.
- Failures are printed as warnings; the old `Path | None` return read every
  failure as "Desktop app up to date".
- The refresh also runs on the "build stamp current" path, so a stale
  /Applications copy left by an earlier update heals on the next `hermes update`
  even when there is nothing to rebuild.
- Core is host-independent (_install_rebuilt_macos_bundles takes paths as data);
  the two invariant tests run on every OS instead of `skipif(darwin)` tests that
  ran nowhere.

Covers the Desktop-button path too: posix.sh runs `hermes update`, so an app
running from apps/desktop/release/ now refreshes the /Applications copy Finder
launches (the Discord report: new shell right after the update, old shell on
the next Dock launch).

* fix(update): repoint Windows managed runtime without renaming live venv

Cherry-pick of helix4u's #88836 (75742db5a42) resolved onto current main.

The POSIX runtime cutover parks the live venv with a directory rename. On
Windows that can never succeed from `hermes update`: the updater executes
from venv\Scripts\python.exe and Windows keeps that image mapped, so the
rename fails with WinError 5 on every attempt (#93032). Instead, atomically
replace the live venv's pyvenv.cfg with the candidate's: Scripts\python.exe
is a launcher that reads `home` on every start, so every fresh process runs
the new generation while nothing touches the locked directory. A failed
post-cutover smoke restores the original config; a failed rollback keeps
the generation the live config now references.

* fix(update): let the Windows repoint run; refuse a minor-line jump

The self-lock/holder preflight (#99711) deferred the repair on the theory
that the updater's own mapped python.exe makes the venv rename impossible.
Live on windows-latest, a process executing from venv\Scripts\python.exe
(and the repo's real .venv with cp311 .pyd extensions loaded) does NOT block
the rename; what blocks it with WinError 5 is any ordinary handle under the
tree: a process cwd, an open file, a sync client. Since sys.executable is
always under the venv on Windows, the deferral fired on every run and no
Windows install could repair from `hermes update`. Retire it and its tests;
the pyvenv.cfg repoint needs no rename and has none of that exposure. The
wine2e lane now runs the cutover probes instead.

The repoint keeps the live venv's site-packages, so provisioning's
fall-forward to the next minor (3.11 -> 3.12, #76106) must not be pointed at
a cp311 tree: refuse before touching pyvenv.cfg. The live Windows test holds
the venv the way the field does (a child with cwd inside it), asserts the
rename path fails (the symptom) and the repoint succeeds; the rollback and
minor-guard tests run on every host.

* fix(title): keep the derived title while skipping the model upgrade (#85194)

`auxiliary.title_generation.enabled: false` turned off both title stages, so an
operator who only wanted to stop the background model call (unavailable or
metered endpoint) also lost the instant derived title. `model_upgrade_enabled:
false` keeps the derived title and starts no `auto-title` thread; missing keeps
the two-stage default and `enabled: false` still disables both.

Salvaged from #85401 onto current main: the gate reuses `_title_config()` and
sits before `spawn_context_thread` (the thread seam moved off `threading.Thread`).

* chore: map contributor email for lightcloud00

Non-noreply author email on the salvaged #85401 commit.

* fix(aux): model_upgrade_enabled only silences the automatic title upgrade

The gate also sat inside `generate_title`, so the explicit operator repair
command `hermes sessions retitle-skills` returned None for every row when the
toggle was off, although the operator asked for a model call. The toggle's
contract is "never spend a model call upgrading the instant title": only
`maybe_auto_title` (the background path) is gated now; a direct
`generate_title` call still asks the model. Docs sentence adjusted.

Review follow-up on #113955; the existing toggle test now also asserts the
explicit path still titles (red on the previous head).

* fix(agent): use provider-default temperature for title generation (#72351)

Title generation hardcoded temperature=0.3 in its call_llm() call.
Models like GPT-5.6 only accept their server-side default temperature
and reject explicit values with "Unsupported value: 'temperature'".
While call_llm() has a retry that strips temperature on error, the
daemon thread races with session cleanup in short-lived CLI sessions,
causing the retry to fail with a connection error.

Fix: pass temperature=None so the provider uses its own default.  This
avoids the unsupported-temperature error entirely and eliminates the
need for the retry path.

Fixes #72351

* fix: chain auxiliary parameter-rejection retries on primary and fallback requests

Reasoning models reject several request fields at once (gpt-5: temperature AND
max_tokens, #78273), a reasoning-strip retry can then 400 on temperature
(#72351), and strict-schema gateways such as Fireworks reject the generic
extra_body.reasoning fallback with "Extra inputs are not permitted, field:
'reasoning'" (#109774) — a phrasing none of the unsupported-parameter predicates
recognised, so the reasoning rung never fired and every title call 400'd.

The parameter rungs were a fixed single-pass order (temperature → structured
output → reasoning → max_tokens): a field rejected AFTER an earlier rung had
already run was never stripped, and _param_rung_accepts did not admit a
temperature 400 raised by a later retry. They now form a table walked until no
rung matches, each field stripped at most once, so N rejected fields recover in
N retries in whatever order the provider raises them.

Fallback candidates (per-task chain, main chain, discovery) got a single shot
with the caller's temperature/max_tokens/reasoning fields and raised on the
first parameter 400; agent/auxiliary_fallback_recovery.py now runs the same
parameter rungs around a candidate's request (sync + async), leaving auth /
payment / connection errors to the caller's existing handling.

Live: real gpt-5-mini title_generation call (reasoning_effort 400 → temperature
400 → 200), real gpt-5-mini fallback candidate sync+async (temperature 400 →
200), and a stand-in replaying Fireworks' documented 400 (reasoning stripped →
200) — all failed on origin/main.

* fix(aux): build the fallback ladder route by field name

A positional _LadderRoute(...) in auxiliary_fallback_recovery breaks as soon as the
route tuple gains a field (#113968 adds timeout); construct by name so fields the
parameter ladder never reads default to None regardless of tuple width.

* fix(aux): fallback ladder route carries the client base_url; test builds the route by name

base_info is the endpoint identity every rung reads for per-route bookkeeping; the
fallback ladder left it empty. The test mirrored the positional construction that
broke under a wider tuple.

* fix(aux): parameter rungs hand a 429 on, and Fireworks gets top-level reasoning_effort

The pre-ladder max_tokens rung accepted payment|connection|rate_limit errors on
its retry, so a 429 after a stripped retry reached the credential and
provider-fallback rungs. The ladder's shared `_param_rung_accepts` dropped
rate_limit: a 429 on the retry raised out of the primary call and skipped the
whole fallback chain. Accept it there too, with a rung-level test that the
stripped kwargs and the 429 are handed on rather than raised.

The body claimed `Fixes #109774` while the aux client still sent the generic
`extra_body.reasoning` to Fireworks on every call and only recovered
reactively (an extra 400 round-trip per request). Port @huklaa's profile
override from #109807: Fireworks documents top-level `reasoning_effort`
(`none` disables thinking), and overriding `build_api_kwargs_extras` marks the
profile reasoning-aware so the transport omits the generic fallback on this
route. Extends the salvage with the effort mapping so enabled-with-effort takes
the same wire, with a control that a profile-less route keeps the fallback.

Salvages #109807 (@huklaa).

Co-authored-by: Hukla <129692708+huklaa@users.noreply.github.com>

* fix(aux): max_tokens rung retries even when the wire kwargs no longer carry the cap

The rung table refused to re-send an unchanged request, but the Codex Responses
route translates the caller cap away and gateways inject their own: the 400 names
max_tokens while the kwargs show none, and the identical retry is what completes
(tests/agent/test_injected_param_strip_retry_registry.py, red in CI on 99ff3b9).

* fix(anthropic): merged user turns keep each turn as its own text block

_concat_content joined two string user contents into one "a\nb" string when
_merge_consecutive_roles collapsed adjacent user turns for the Anthropic
Messages wire. For a MoA aggregator on that wire, iteration 1 of a turn ends
[user(task), user(guidance)] and was sent as user("task\n<guidance>"), while
iteration 2 replays user("task") alone, so the prompt-cache prefix diverged
at the first user block and the #112358 collapse persisted there (the
first-pass fix in #113175 only covered the OpenAI-compatible wire).

Merged turns are now always a block list with each side's blocks intact (a
string becomes one text block), matching what the list+list and list+str
shapes already did. The task block is byte-identical to the standalone turn
later iterations replay, a cache_control marker on it stays put, and the
guidance follows as its own text block. Assistant merges are unaffected
(assistant content is already a block list).

Bedrock Converse and native Gemini already merge at block/part granularity,
so the docs' Anthropic/Converse/Gemini fold caveat is replaced with the
accurate statement and the byte-identical-extension claim no longer needs
the OpenAI-compatible scope.

Part of #112358

* test(mcp): drive a poisoned client.json through _configure_callback_port end-to-end

The #112568 coverage exercised _cached_redirect and get_client_info in
isolation; the issue promised a flow-level case. This drives the login
path's actual sequence -- _configure_callback_port(cfg, storage) then the
SDK's storage.get_client_info() -- with the issue's exact bad-port payload
and a non-dict client.json, asserting a freshly parked ephemeral port and
"no registration", so a regression on either step fails the caller path,
not just the private helper. Red before f8f89b20d27 (both cases) and
before e166791f9ff (non-dict case); green on main.

Closes the remaining atom of #112568.

* fix(skills): skill_manage advertises one shape per action; a misfiled text slot fails before any op applies

The operations[] item schema was one flat object with four coexisting text
slots (content / new_string / file_content / file_path). A 27B local model
that had just used write_file's file_content kept emitting it on create and
patch ops; the call validated against the advertised schema, the handler
failed on "content is required", and the whole batch rolled back — eight
identical retries until the tool-loop guardrail tripped (#112677).

- items is now an anyOf of self-contained per-action op objects (create,
  patch targeted, patch full-rewrite, write_file, remove_file, delete), each
  with additionalProperties: false. The wire shape of a correct call is
  unchanged (still a flat op with name/action/...), so transcripts, staging
  and replay are untouched; grammar-constrained backends can no longer emit
  another action's slot, and schema-validating providers reject it up front.
  Nested (non-top-level) anyOf survives every sanitizer (schema_sanitizer,
  Gemini legacy translator). Cost: parameters JSON 1352 -> 2414 bytes.
- _validate_batch_ops runs the per-op argument-shape check (_op_shape_error)
  before any op is applied, so a misfiled op[1] no longer applies op[0] and
  then rolls the batch back; the error carries the same misplaced-key hint.
- The hint is attached to argument-shape misses only: a patch whose real
  problem is an unmatched old_string is no longer told to "move that text to
  'content' (full rewrite)", the escape the patch error itself warns against.
- Shape tables (_REQUIRED_ARGS, text-slot maps, _misplaced_text_hint,
  _op_shape_error) move out of the facade into skill_manager_batch.py, the
  op-validation sibling; _patch_skill shares the old_string guidance text.
- Docs: skills.md Actions table states the one-slot-per-action contract.

* fix(skills): keep absorbed_into in the delete shape; decide every patch shape miss pre-effect

The per-action schema made the delete branch `additionalProperties: false`
with no `absorbed_into`, so a schema-validating or grammar-constrained
backend could no longer emit the curator's consolidation delete and
`_curator_consolidation_delete_guard` fail-closed every consolidation.
Advertise `absorbed_into` on the delete branch and cover the batch path
that forwards it to the guard.

Also fold the two remaining patch shape checks (missing new_string,
content mixed with old_string/new_string) into `_op_shape_error`, so a
batch rejects them before applying any sibling instead of creating op[0]
and rolling it back. Drop the stale `edit` vocabulary from skills.md:440
and the curator prompt, and move the schema-diet test helpers above the
`__main__` guard.

* test(curator): stop pinning the unadvertised edit alias in the read-before-write prompt

The prompt now names patch/write_file/remove_file only; edit is the legacy
alias of a full-rewrite patch and no longer appears in the advertised
vocabulary, so the prompt assertion must not require it.

* chore: map sy0u1ti for salvage of #114024

* fix(sessions): include live session activity in prune recency

* refactor(sessions): pinned back-fill uses the canonical last-active expression

The list_sessions pinned back-fill query still hand-rolled
COALESCE(MAX(m2.timestamp), s.started_at) AS last_active while the main
query a few lines above already used _sql_session_last_active("s").
Use the shared helper so pinned rows report the same freshest-of
(last_activity_at / latest message / started_at) recency as every other
listing and prune path, instead of ignoring live activity.

* docs(sessions): recency wording matches the freshest-of rule

Prune recency is now the freshest of last_activity_at / latest message /
started_at (via _sql_session_last_active), but four doc sites still said
"latest message, else started_at": the list_prune_candidates docstring,
the --older-than/--newer-than help in session_filters, and two comments
in config_defaults. Reword them to the phrasing archive_stale_sessions
already uses so operators are not told live activity is ignored.

Also make archive_stale_sessions use the module constant _LAST_ACTIVE_SQL
instead of an inline _sql_session_last_active("s") call — same alias,
identical SQL, one fewer place to drift.

* docs(sessions): user guide describes the freshest-of recency rule

The --older-than/--newer-than and prune-retention paragraphs still said recency is the latest message (falling back to session start). After the prune-recency fix, both paths use the freshest of live activity / latest message / session start; reword to match the phrasing in hermes_cli/session_filters.py.

* fix(tools): refuse text writes into database files and binary overwrites

The text tools' read->modify->write round-trip re-encodes binary bytes
lossily: the terminal transport decodes stdout with errors=replace, so
a patch on a SQLite database reads mojibake and writes it back, silently
destroying the file (observed in the wild: a kanban board DB lost 550k
bytes / 26 days of history this way on 2026-09-15).

read_file already refuses to display binary files (extension + byte
sniffing), but write_file/patch had no write-side equivalent, and
patch_replace reads via plain cat so the read-side guard never fires
for it.

Guard rules (extends _check_binary_document_write, the existing
binary-document write guard):

- .db/.sqlite/.sqlite3 and their -wal/-shm/-journal sidecars: ALWAYS
  refused (even new-file creation) — plain text can never be a valid
  database, mirroring the opaque-document rule. The sidecars need a
  basename regex because os.path.splitext('x.db-wal') yields '.db-wal',
  which is not in BINARY_EXTENSIONS.
- Other BINARY_EXTENSIONS (.png, .zip, ...): refused when OVERWRITING an
  existing file, mirroring the existing .pdf-overwrite rule — the model
  can only have read mojibake, so writing back destroys it. Creating a
  NEW file with such an extension stays allowed.

Tests: new test_binary_db_write_guard.py covers the regex (main names,
sidecars, non-db names), the guard function, and the write_file/patch
tool paths against real SQLite files (integrity_check still ok, bytes
untouched). 29 new+existing guard tests pass; file-tools/patch suites
green (196 tests).

* refactor(tools): recognise SQLite sidecars in binary_extensions, drop write-guard regex

Move the .db-wal/-shm/-journal detection from a write-guard-local regex
into tools/binary_extensions._has_extension_in so has_binary_extension
(read guard AND write guard) agree: a sidecar counts as its database's
extension. read_file now refuses sidecars instead of returning lossy
text, which also gives write_file the no-baseline overwrite refusal.

Behaviour change vs the contributor's commit: creating a NEW .db /
.sqlite file via write_file stays allowed (main allows it and text
fixtures named *.db exist) — only overwriting an existing binary is
refused. The generic has_binary_extension overwrite branch is kept and
merged with the PDF branch because patch has no full-read baseline
check: on main, patch on an existing .db with a matching old_string
rewrites the header in place.

Tests folded into the existing guard test classes (one write_file, one
patch refusal on a real WAL sidecar); the PR's 12-test file and its
private-regex assertions are dropped.

* test(tools): pin the binary refusal for write_file into a SQLite -wal sidecar

Any error would also be produced by the no-baseline overwrite guard on
main, so the test could not detect a regression of the sidecar
detection. Asserting the binary-file refusal makes it fail when
binary_extensions stops recognising .db-wal.

* fix(tools): a SQLite sidecar path is refused even when no sidecar file exists yet

The refactor that moved sidecar detection into binary_extensions dropped the
contributor's unconditional refusal for the db family and made every binary
extension a plain overwrite guard. That is right for the MAIN .db/.sqlite
(text fixtures named *.db exist), but a -wal/-shm/-journal path is never a
legitimate text target: a checkpointed database has no sidecar on disk, so
write_file("kanban.db-wal", text) silently dropped a garbage WAL next to a
live database.

Add is_sqlite_sidecar(path) and refuse such paths in
_check_binary_document_write regardless of is_file(). Scope the marker
stripping in _has_extension_in to SQLite suffixes so report.docx-wal no
longer counts as an opaque document. Hoist the duplicated is_pdf_path call.
Parametrize the write_file WAL test over sidecar existing/absent.

* fix(tools): sandbox read paths recognise SQLite sidecars as binary

file_operations.py checked `os.path.splitext(path)[1].lower() in
BINARY_EXTENSIONS` at five sites, so a `.db-wal` / `.sqlite3-shm` sidecar
(final suffix never in the set) was read and edited as text in the sandbox
paths. Route all five through has_binary_extension(), which strips the
sidecar marker. Behaviour is otherwise identical: same lower-cased final
suffix membership (rfind vs splitext differ only on leading-dot basenames
like `.bashrc`, which are in neither set). binary_extensions imports nothing
from tools, so no cycle.

* test(tools): honest sidecar fixture

_make_wal_db claimed sqlite keeps the -wal until checkpoint, but SQLite
deletes -wal/-shm on last-connection close, so the fake-bytes fallback was
what actually ran. Make the fixture a context manager that holds the
connection open across the tool call so the sidecar is a real WAL, and drop
the PRAGMA integrity_check assertion (SQLite ignores a garbage WAL, so it
passed either way); keep the byte-identity check.

In the patch test assert "binary" in the error so the no-baseline guard
cannot mask a regression in sidecar detection.

* refactor(tools): one suffix helper for the binary-extension checks; drop the dead .db3 member

_has_extension_in and is_sqlite_sidecar each re-implemented the same
rfind('.') / -1 / .lower() suffix extraction, and the write guard used a
third spelling (filepath[filepath.rfind("."):]) for its refusal messages
next to os.path.splitext in the overwrite branch. Add _lower_suffix()
("" when no dot) and route both predicates through it; the write guard
now uses os.path.splitext for every message's displayed extension.

_SQLITE_EXTENSIONS was `{.db,.sqlite,.sqlite3,.db3} & BINARY_EXTENSIONS`,
which silently dropped .db3 (not a binary extension anywhere in the
codebase, so `x.db3-wal` was never a sidecar). Spell the set as the three
members that actually take effect and assert the subset invariant instead
of hiding it behind an intersection. Behaviour is unchanged.

* test(tools): overwriting the database file itself is still refused as binary

The test trim dropped the headline case: text written over an EXISTING .db. Sidecar paths return at is_sqlite_sidecar before the overwrite branch, so a mutant that drops has_binary_extension from the overwrite condition survived every remaining test. Add a third parametrize case that targets the held-open WAL database file itself and asserts the binary refusal with bytes unchanged.

Same commit: the _check_binary_document_write docstring now names SQLite sidecars (always) and every BINARY_EXTENSIONS suffix (on overwrite), and the suffix is computed once at the top instead of three times.

* chore: map holny for salvage credit of #113643

* fix(telegram): rebuild adapter on confirmed polling stall instead of reusing unquiesced updater (#113618)

* refactor(telegram): type the polling-stall signal instead of matching log text

Replace `_looks_like_polling_stall` (a three-substring classifier over
`str(error)`) with `_PollingStallError(RuntimeError)`. The watchdog and the
post-reconnect verifier raise it at their two stall sites; the reconnect
ladder checks `isinstance` through `_iter_exception_graph` before handing
the adapter to the supervisor.

WHY: a substring match couples recovery routing to log wording (rename a
message, silently lose the fatal handoff) and can misfire on unrelated
errors that quote the same words. A typed exception keeps #113657's
behaviour with no behavioural change beyond the classifier's source.

* test(telegram): fold stall handoff cases into one parametrized invariant

Merge the watchdog and verifier stall tests into a single parametrized
`test_confirmed_stall_hands_off_before_reusing_updater` and assert the
adapter is also marked degraded, keeping two invariant tests for #113618
(stall -> fatal handoff; generic network error -> in-place reconnect).

* refactor(telegram): watchdog stall reuses _schedule_polling_recovery

`_check_polling_stall` hand-rolled the body of `_schedule_polling_recovery`
(set `_send_path_degraded`, `_mark_degraded()` when running, spawn
`_handle_polling_network_error`). Call the helper instead, matching the
post-reconnect verifier stall site.

WHY: one recovery entry point keeps degraded-state marking and the
"polling degraded (reason)" log line consistent across every scheduler.
The helper's early-return guards (`_teardown_started or has_fatal_error`,
`_recovery_in_flight()`) are already asserted at the top of
`_check_polling_stall` with no await in between, so they are no-ops here
and behaviour is unchanged.

* fix(telegram): a confirmed stall hands off before the reconnect backoff, not after it

Hoist the `_PollingStallError` check in `_handle_polling_network_error` to
right after the teardown/fatal guard, before the retry counter increment,
the exponential sleep, `_stop_updater_or_go_fatal` and both connection
drains. Use a plain `isinstance` (both raise sites construct the error
directly; nothing wraps it). The stall test now also asserts no sleep, no
retry-counter bump, no in-place `updater.stop()` and no drain.

WHY: the check sat after `await asyncio.sleep(delay)` (5-60 s), the
counter bump and a bounded `updater.stop()` (up to 15 s), so a gateway
already confirmed deaf stayed deaf 5-75 s longer and consumed a retry
slot for something that is not a retry.

The in-place `updater.stop()` before the handoff is dropped deliberately:
`_go_fatal_network` -> `_handoff_polling_fatal_error` -> supervisor
rebuild runs `disconnect()`, which performs the same bounded
`updater.stop()` (`_UPDATER_STOP_TIMEOUT`, falls through on timeout) plus
`app.stop()/shutdown()`, so stopping here only duplicated that work on the
slow path. Docstrings for `_handle_polling_network_error` and
`_check_polling_stall` no longer describe the stall as a reconnect-ladder
escalation.

* refactor(telegram): a confirmed stall logs one hand-off line, not a retry promise

_schedule_polling_recovery promised the gateway 'stays alive and will retry' for every error, but a _PollingStallError goes straight to _go_fatal_network (supervisor rebuild). Branch the wording on the error type, drop the watchdog's own pre-log so a stall yields exactly one error-level line (from _go_fatal_network), and carry stalled_for/generation in the stall error text instead. Test module docstring and test name updated to match the hand-off semantics.

* fix(mcp): avoid default keepalive on stdio servers

* fix(mcp): wake recycle deadlines after active RPCs

* docs(mcp): clarify transport keepalive defaults

* fix(mcp): an idle stdio server still proves its session without a keepalive ping

After the stdio keepalive default became None, the lifecycle loop's
`if keepalive_interval is None: continue` skipped the only idle path to
`_mark_session_proven()`. `_session_proven` is otherwise set only on a
tool-call success or a suspect health-check pass, and each new session
resets it, so an idle default-config stdio server never became proven:
every later child death/reconnect charged the rapid-drop budget that a
180 s idle survival used to clear (#62212 regression).

While the session is unproven, keep the `asyncio.wait` timeout at
`_DEFAULT_KEEPALIVE_INTERVAL` for stdio too; when it fires with the child
still alive, mark the session proven without pinging. Once proven the
stdio timeout returns to None, so a healthy idle pipe is never probed.
Shutdown/reconnect precedence and the rpc_idle waiter are untouched.

Test: `_captured_interval` now runs two lifecycle cycles (first times
out, second shuts down) through a shared helper so the stdio test asserts
timeouts [180, None], zero pings and `_session_proven` set, instead of
copy-pasting the fake `asyncio.wait`. Docstring/comment updated to the
new rule.

* refactor(mcp): lifecycle loop reads its inputs once and stops re-deriving the recycle predicate

WHAT: `_wait_for_lifecycle_event` hoists `is_http = self._is_http()` (it only
reads the immutable `"url" in self._config`) and reads `keepalive_interval`
from the config once instead of `in` + `.get()`. The one-shot RPC-idle waiter
is now spawned on `not is_http and self._rpc_lock.locked()` alone, dropping
the inline `_idle_timeout_seconds is not None or _max_lifetime_seconds is not
None` clause. The waiter list is built by one comprehension per iteration
(rpc_idle_task changes between iterations) and the `finally` reuses the last
value; `_cancel_waiters` already skips done tasks. The docstring folds the
#17003 keepalive paragraph into the first paragraph instead of describing the
keepalive twice.

WHY: the dropped clause was the inverse of
`MCPServerHealthMixin._stdio_recycle_deadlines` re-derived by reaching into
health-mixin attributes. Without it, a stdio server with no limits and an
active RPC spawns one trivially-completing waiter per RPC; on lock release
the loop takes one extra iteration where `_recycle_if_due()` is False and
`_next_stdio_recycle_deadline()` is None, so it simply recomputes the timeout
and waits again — harmless.

* test(mcp): one recycle-wake case

WHAT: drop the idle-vs-lifetime parametrize on
`test_stdio_recycle_wakes_after_active_rpc`, keeping the idle case, and drop
its `_recycled_reason == reason` assertion.

WHY: both parameters exercised the same new invariant (releasing the RPC lock
wakes the lifecycle loop so the now-visible deadline recycles the server); the
reason assertion pinned pre-existing `_stdio_recycle_reason` behaviour that is
not this stack's concern. The stack now adds exactly two test functions.

* fix(mcp): an idle stdio server is proven only after a full default interval

The no-keepalive proof branch ran on ANY timeout wake. With idle_timeout_seconds=60 the wait is min(180, 60); an overlapping RPC hides the recycle deadline so _recycle_if_due() is False and the session was marked proven after 60 s, clearing the #62212 rapid-drop budget early. Fix a proof_at deadline before the loop and gate the branch on it; the unproven timeout is the remaining distance to proof_at. Tests: the stdio test now shows the first (idle-limit) wake does not prove; the interval test covers an explicit keepalive_interval on stdio.

* fix(desktop): reasoning blocks stop eating streamed prose

preprocessMarkdown runs on the accumulated text on every streaming flush, so
the plain REASONING_BLOCK_RE replace damaged the visible answer in two ways a
settled message never shows:

- a model that inlines its chain of thought in the answer channel streams the
  open tag long before the close one, and the regex needs both, so the reasoning
  rendered as chat prose until the close tag arrived — and when it did the whole
  span vanished in one frame, taking text the reader had already read with it
  (#62774: the mid-reply "truncation" on long answers).
- the regex ate the whitespace on its seam, so a block sitting between two words
  fused them: `no` + `Hermes` rendered as `noHermes`.

An unterminated block now strips at a block boundary — the same rule
agent/think_scrubber.py draws, so a real reasoning preamble (always its own
block) disappears while prose that merely mentions `<thinking>` mid-sentence
survives — and a removal between two prose fragments leaves one space instead of
deleting the seam.

Tests: markdown-reasoning-stream.test.ts (module level: seam, unterminated
blocks, per-delta leak + visibility monotonicity) and
streaming-text-fidelity.test.tsx (real surface: accents/emoji and plain prose
streamed in 1- and 3-character deltas, plus the chain of thought never painting
a frame).

* refactor(desktop): align reasoning tags with think_scrubber, trim #113302 tests

WHAT: one REASONING_TAGS alternation feeds both regexes and now also covers
`thought` and `reasoning_scratchpad` (agent/think_scrubber.py THINK_TAG_NAMES);
the desktop-only `scratchpad`/`analysis` stay. Behaviour change beyond the
contributor's: closed or unterminated `<thought>`/`<reasoning_scratchpad>`
blocks are now stripped on the desktop too.

Doc comments trimmed to the WHY (block-boundary rule, why an unterminated block
is held back). Tests cut to the ≤2-invariant shape: two unit cases (seam,
unterminated-vs-mention) and one DOM frame-by-frame case; the accents/plain-prose
DOM cases were protection for unrelated code and are dropped.

* perf(desktop): reasoning-block seam check reads two chars, not two string copies

stripReasoningBlocks runs on the accumulated message text every streaming
flush (markdown-text.tsx). Its replacer sliced the whole text twice per
closed block to test whether a space is needed at the seam — O(n) copies plus
a forward scan per block per flush. Read the single char on each side of the
match instead.

The seam test also had a second defect: with two adjacent closed blocks the
second block's `before` ended in the first block's `>`, so a second space was
emitted (`no  Hermes`). Matching a run of adjacent blocks as one match makes
the two-edge check see the real prose on both sides. Covered by an extra
expect in the existing seam test.

* fix(desktop): a half-arrived reasoning tag at a block boundary is held back too

OPEN_REASONING_BLOCK_RE needs the full `<tag>`, so a frame ending in `<thin`
at a block boundary painted `<thin` as prose for one frame and then erased it
— the paint/un-paint class of #62774, one frame long. The fidelity test carved
an escape hatch around exactly those frames.

Add a third pass that drops a trailing partial open tag at a block boundary,
restricted to prefixes of the known reasoning tag names (built from
REASONING_TAGS) so `<div` at a line start still renders. This is the desktop
counterpart of agent/think_scrubber.py `_hold_partial`/`_max_partial_suffix`,
cited rather than ported. The fidelity test's escape hatch is deleted so the
monotonic-prefix invariant holds on every frame; the unit test gains the
`<thin` positive and `<div` negative cases.

* test(desktop): reasoning tests follow the sibling naming

The lib test sits beside markdown-preprocess.directives.test.ts as
markdown-preprocess.reasoning.test.ts, and the frame-fidelity test joins
its markdown-text.*.test.tsx siblings as markdown-text.reasoning.test.tsx,
so a glance at the directory shows what each file exercises. Also trim the
REASONING_TAGS comment: how the two tag lists converged is git history, not
something a reader of the constant needs.

* chore: map cswrld-net for salvage of #112791

* fix(desktop): answer server→client requests when a handler crashes instead of stalling the backend

A clarify request renders as an eternal spinner when anything in the
renderer's handler chain throws: deliverRequest had no error handling,
the WS message listener let the exception escape as an uncaught error,
and both dispatch sites silently no-oped via optional chaining when the
registry was absent. In every case the backend (clarify_tool blocks up
to 3600s) never receives any frame — no result, no error — and waits
out its whole deadline.

- json-rpc-channel: wrap each handler invocation; a crash now answers
  -32603 ("server request handler crashed: <method>"), fires the
  onUnhandledRequest hook, logs the stack, and stops. The unhandled
  path still answers -32601.
- json-rpc-gateway: guard the socket message listener so no frame can
  escape as an uncaught error.
- store/gateway (primary + secondary wiring): missing registry now
  fails the request immediately with -32601 instead of dropping it.

Backend already handles {"error"} response frames
(tui_gateway/server_requests.resolve_response), so fail-fast answers
settle the tool at once; no backend change needed.

Regression test drives boom→-32603, unknown→-32601, and a working
request after the crash.

* refactor(desktop): drop the socket-listener try/catch and share the registry-missing guard

Follow-up to the salvage of #112791.

- json-rpc-gateway.ts: revert the try/catch around `channel.handleFrame`
  to main. deliverRequest is the single chokepoint that dispatches into
  feature handlers and now answers -32603 itself; a second catch in the
  socket listener is defense-in-depth that would also hide bugs in event
  and response handling that should surface as uncaught errors.
- store/gateway.ts: the primary and secondary dispatch sites carried the
  same copy-pasted "no registry -> fail -32601" block. Fold both into one
  file-local `dispatchServerRequest(request, profile, connectionId)`.
  The secondary keeps main's tagging (its own connectionId, no fallback
  to the active connection), which the contributor's version changed.

* refactor(shared): a crashed server-request handler reports through an owner hook, not console.error

The deliverRequest catch branch wrote to a module-level console.error sink
(the only direct console.* in apps/shared/src) and fired onUnhandledRequest,
whose contract is "nobody handled it, already answered -32601" — so the TUI
logged a -32603 crash as "unhandled server request".

Add onRequestHandlerError(error, request) to JsonRpcRequestChannelOptions
beside onHeartbeatFailure, call it from the catch after answering -32603,
and drop the console sink. Wire both owners: HermesGateway (desktop, via a
GatewayClientOptions passthrough) logs to console.error like its dial-failure
sink; ui-tui gatewayClient pushes a [protocol] log line. Collapse the two
normalisation arms into the existing `error instanceof Error ? … : new
Error(String(error))` idiom and restore the early `return true` instead of
the handled flag + break — nothing runs after the loop but the -32601
fallthrough.

Test: the crash case now asserts onRequestHandlerError fires once for the
-32603 request and onUnhandledRequest only for the -32601 one.

* test(desktop): registry-missing server request answers -32601

Mutation check: reverting store/gateway.ts to the merge-base left the desktop
suite green — nothing exercised dispatchServerRequest. Cover both arms through
dispatchPrimaryServerRequest: a registry configured without onServerRequest
answers request.fail(-32601, /registry/) exactly once; with onServerRequest
the request is forwarded carrying the profile and fail is never called.

* refactor(desktop): a missing server-request registry lets the channel answer -32601

dispatchServerRequest hand-rolled request.fail(JSON_RPC_METHOD_NOT_FOUND,
'Hermes Desktop has no server-request registry yet'), but the channel already
owns that reply: JsonRpcRequestChannel.deliverRequest answers -32601 and fires
onUnhandledRequest when a ServerRequestHandler returns false, and
GatewayBootOptions.handleServerRequest already documents "false = no handler
(the channel answers -32601)".

Make dispatchServerRequest return false when the registry has no
onServerRequest and true after forwarding; dispatchPrimaryServerRequest and
both gateway.onRequest registrations (store/gateway.ts secondary sockets,
use-gateway-boot.ts primary) now propagate that value to the channel. Keep
the desktop-specific wording by wiring onUnhandledRequest on HermesGateway
beside the onRequestHandlerError sink (console.warn), which needs the same
GatewayClientOptions passthrough onRequestHandlerError got. Drop the now
unused JSON_RPC_METHOD_NOT_FOUND import from the store and export
JSON_RPC_INTERNAL_ERROR from the shared barrel beside it.

Test: the store test asserts the false return with no fail() call when the
registry is missing, and true + forwarded profile when it is present.

Follow-ups (same class, outside this stack): use-gateway-boot.ts ~L889-891
still hand-rolls -32601 when the registry is present but has no handler;
ui-tui has its own copy.

* fix(gateway): resolve session_key in shutdown-flush recovery

recover_pending_to_db skipped every real flush file: _serialise_value
captures only the text field from adapter MessageEvent objects (they
have no session_id attribute), so recovery always hit the 'no
session_id' skip branch — messages queued during a gateway drain were
written to disk and then silently never re-ingested, despite the
user-facing 'queued for the next turn' promise. The existing tests
masked it by hand-writing session_id into payloads real events never
produce. Add an optional session_resolver parameter and wire it to
SessionStore.peek_session_id at the startup call site; files remain
preserved (with the warning) when a key genuinely can't resolve.

* fix(gateway): resolve flush session_key to session_id at recovery

* fix(gateway): keep one failed flush file from aborting pending-message recovery

`recover_pending_to_db` walks the shutdown flush spool and appends each
pending message back into state.db. Its per-file `try` listed
`except BaseException:` above `except Exception as exc:`. Exception is a
BaseException subclass, so the ordinary-error handler was unreachable and
every ordinary failure took the interrupt path: close the owned DB and
re-raise.

The effect is that a single unrecoverable spool file aborts the entire
recovery pass. Because the file is only unlinked after a successful
append, it survives and re-poisons the next boot, and the pass is walked
in `sorted(glob("*.json"))` order over uuid4-named files, so a different
subset of the user's pending messages is stranded each time. The caller in
`gateway/run.py` wraps the call in `except Exception: pass`, so nothing is
logged — the messages simply never come back.

Reordering the two handlers restores the documented per-file behaviour
("Leave the file for next startup retry") while keeping the interrupt
contract exactly as written: a KeyboardInterrupt, SystemExit or
CancelledError still closes an owned SessionDB and propagates.

* test(gateway): cover recovery skipping a failing payload and continuing

Pins that an ordinary append failure on one spool file only skips that
file: the remaining files are still recovered, the good file is unlinked,
and the failing one is preserved for the next startup retry.

Fails before the handler reorder — the RuntimeError propagates out of
recover_pending_to_db and the second message is never recovered.

* refactor(gateway): shutdown-flush resolver — routing map first, exact-key row fallback fenced by flush time

Follow-up to the #75536 + #106112 picks; changes vs #106112:

- Routing map first: peek_session_id (sessions.json) is authoritative for
  the session a message was routed to at shutdown. The durable-row lookup
  (find_latest_gateway_session_for_peer under the exact key) is only the
  fallback for a pruned map.
- Fence added per gysyl's #106112 review: a row whose started_at is later
  than the flush payload's ts (passed as not_after by _recover_one_payload)
  cannot be the message's origin and is never adopted; None is returned so
  the flush file is preserved instead of appending into a newer session.
- WhatsApp phone<->LID alias expansion dropped: main's session_recovery
  does no alias expansion anywhere else, and re-implementing it here made
  the resolver diverge from the recovery path it sits beside. Exact-key
  match only.
- The private _db_for_key reach stays inside SessionRecoveryMixin so the
  (session_id, db) pair lands the append in the owning profile partition;
  shutdown_flush never touches store internals.

* test(gateway): trim flush-resolver tests to positive + fence invariants

Keep test_recover_payload_without_session_id_uses_resolver_and_deletes_file
(positive path) and add one negative fence test on the resolver: a row
started after the flush ts is rejected, an older one is adopted with its
owning db. Drop the WhatsApp alias tests (feature removed), the duplicate
positive test and the redundant no-resolver/None-resolver file-preserved
tests, which were exercising the same branch three ways.

* refactor(gateway): one peer-finder helper for the two recovery lookups

_find_gateway_session_row and resolve_session_id_for_key each carried the
same guarded call: getattr the finder off a possibly-None db, check
callable, try/except -> debug log -> None. Extract it as the mixin-private
static `_peer_row(db, *, source, session_key, raise_on_lookup_error=False,
**peer)` and call it from both sites.

Behaviour is unchanged at both call sites: the tuple site passes
user_id/chat_id/chat_type/thread_id through **peer with the same
allow_peer_fallback gating and keeps raise_on_lookup_error; the exact-key
site passes only source + session_key, so the finder's tuple fallback
still never runs there. The only observable difference is the resolver's
failure debug line, which now reads "Gateway session DB recovery failed"
(shared) instead of "Session key->id resolution failed"; no test asserts
on either.

The platform-slot parse (parts[2]) in resolve_session_id_for_key is left
in place: the only key-parsing helper on the mixin,
_profile_from_session_key, returns the profile namespace (parts[1]) and
nothing returns the platform slot, so there is nothing to fold onto.

* docs(gateway): say why the exact-key resolver skips the scope fences

resolve_session_id_for_key calls the peer finder directly, while the
sibling _query_recoverable_row also applies _recovered_row_matches_source_scope
and _recovered_row_allowed_for_active_profile. Read both fences against the
resolver's call shape; both are redundant there, so this adds a WHY comment
at the call site rather than a fence call.

- Profile fence: it returns True immediately when recovered.session_key ==
  requested key. The resolver passes no chat_id/chat_type, so
  find_latest_gateway_session_for_peer returns after the _PEER_BY_KEY_SQL
  branch (`s.session_key = ?`) or None; the tuple fallback is unreachable.
  Every hit therefore carries the requested key and the fence is a no-op.
- Slack workspace fence: it compares origin_json.scope_id with
  source.scope_id. The resolver has no SessionSource (only the key), and a
  scoped Slack key embeds the scope_id as its own slot
  (`<ns>:slack:<chat_type>:<scope_id>:...`, build_session_key). An exact
  key match therefore already pins the workspace, regardless of whether
  the row lives in a profile store or a non-multiplexed root store — the
  store choice (_db_for_key) partitions by profile, not by workspace, and
  the key does the workspace work in either store. The sibling needs the
  fence only because it also looks up the UNscoped legacy key, whose row
  can belong to any workspace; the resolver never does that lookup.

* fix(gateway): an unresolvable profile store preserves the flush file instead of writing to the root store

`_db_for_key` deliberately returns None for a named-profile key whose home cannot be
resolved (fail closed, #66887/#102157), but `resolve_session_id_for_key` returned
`(session_id, None)` on a routing-map hit, and `_recover_one_payload` treated that None
as "fall back to session_db" — the owned ROOT state.db at the production call site. A
profile message with an unopenable home was appended to the root store: the very
split-identity write the fail-closed return exists to prevent.

- resolver: an unresolvable store answers None before the peek and the row fallback
  (None already means "preserve the flush file").
- recovery: the resolver's db is authoritative for resolver-resolved payloads; the owned
  default store only serves payloads that already carry a session_id (and cap-drop spool).
- `started_at` is `REAL NOT NULL` in the schema, so the `float()` fence stays as is.

* fix(gateway): the flush-time fence compares whole seconds

The flush payload's ``ts`` is ``int(time.time())`` while ``sessions.started_at`` is a REAL, so
``started_at > not_after`` rejected a row minted later in the same second as the flush — the
common "message arrives, session minted, SIGTERM" shape lost recoverability (file preserved, not
misrouted). Compare ``floor(started_at)`` against the whole-second ``ts`` so same-second rows
are adopted and only rows from a later second are refused.

Spotted by the round-2 gate on the salvage stack.

* fix(config): allow unseeded runtime config keys

* test(config): pin the same-section-typo trade-off of the unseeded-key fix

Fold the trade-off into the salvaged parametrized test instead of adding a
new test function: `agent.max_turnz` (suggestion `agent.max_turns`) is now
written with the post-write notice rather than refused, because the schema
walk cannot tell a same-section typo from a deliberately unseeded
runtime-read key such as `stt.provider`. Assert the notice and the
suggestion so reviewers see the behaviour change; refresh the class
docstring that still described the pre-#114107 refusal rule.

* docs(config): config set docs describe the wrong-prefix-only refusal

The CLI reference, the configuration guide and the set_config_value
docstring still said that any unknown path under a known section is
refused with a did-you-mean. After b47f40e699 only a path whose suffix is
itself a known key (gateway.discord.foo -> discord.foo) is refused; every
other unknown path under a known section — a same-section typo or an
unseeded runtime-read key — is written with a did-you-mean notice, because
DEFAULT_CONFIG is an incomplete schema and cannot tell the two apart.

Reword all three sites to state the refusal that actually exists and the
write-with-notice fallback, so users are not told a typo will be blocked.

* test(config): the unseeded-key rows state their real did-you-mean

The display.tool_progress row passed suggestion=None, but
_validate_config_key("display.tool_progress") returns
(False, "display.tool_progress_command") because only the _command sibling
is seeded in DEFAULT_CONFIG. The `if suggestion:` guard let the row pass
without ever checking the notice, so the misleading "Did you mean" this
real runtime key (cli.py:2568) receives was neither asserted nor visible.

Pin the actual suggestion for that row and assert the notice text exactly
for every row: present with the expected sibling, or absent. Fix the
comment too — only stt.provider comes from da942e4483's list; the rows are
unseeded runtime-read keys at agent/agent_init.py:1324, cli.py:2568 and
tools/transcription_tools.py:241.

Seeding display.tool_progress in config_defaults.py (hermes_cli/AGENTS.md:
"every reader a registry entry") is the proper follow-up that removes the
misleading suggestion; it is out of scope for this salvage.

* test(config): unseeded-key rows pin the invariant, not a known-defective suggestion

Drop the display.tool_progress row (it pinned a did-you-mean the comment itself called misleading; seeding the key would fail it for an unrelated reason) and the skills.creation_nudge_interval row (same invariant as stt.provider, the key named in the issue).

* fix(config): a structural wrong-prefix match beats a fuzzy sibling suggestion

In _validate_config_key the fuzzy same-level sibling (cutoff 0.6) was tried before the
structural "path minus its wrong prefix is a known key" check. While every unknown sub-key was
refused that order did not matter; now that only the wrong-prefix case is refused, a path whose
stray middle segment fuzzy-matched a sibling was WRITTEN with a misleading did-you-mean:
`agent.gateway.strict` -> (False, 'agent.gateway_timeout') instead of being refused as
`gateway.strict`. A structural match is proof, a fuzzy match is a guess; check it first.

* fix(config): the custom-top-level footer prints only for top-level keys

_print_unknown_key_notice always appended "Custom top-level keys are supported and bridged to
the environment". Before the unseeded-key change no nested path reached this notice; now
`stt.provider` or `agent.max_turnz` did and were told they are env-bridged, which they are not.
Nested paths get only the --force hint.

* fix: goal loop names the Nous auth failure instead of an opaque judge error

With auxiliary.goal_judge.provider pinned to nous and a dead refresh token
(invalid_grant), the status line read "judge error: RuntimeError": the ladder's
_resolve_nous_runtime_api swallowed the resolver's AuthError at DEBUG and the
generic "no API key" RuntimeError was reduced to its type name by judge_goal.
Users were sent to context-length / model debugging instead of `hermes model`.

Now the failure is remembered in agent/auxiliary_unavailable.py (one WARNING per
distinct message, preferring the persisted quarantine marker because the poo…
emirsaffar-collab added a commit to emirsaffar-collab/hermes-agent that referenced this pull request Sep 18, 2026
…cks from titler input (#4)

* refactor(desktop): one ordering for sidebar groupings shared by the menu and the keybind

Follow-up to Ayush Nangia's feature: the filter menu kept its own literal
list of groupings while the new cycle keybind walked a second one, with only
a test comment saying they should agree. The menu now builds its options
from `SIDEBAR_GROUPING_ORDER`, so reordering or adding a grouping is one
edit and the keybind cannot drift from what the menu shows.

* refactor(desktop): derive SidebarGrouping from the grouping order; invariant cycle test

`SidebarGrouping` was a literal union kept in step with `SIDEBAR_GROUPING_ORDER`
by hand — a fifth grouping added to the type would compile while the cycle
silently skipped it. The array is now the single source (`as const`) and the
type is derived from it. The cycle test no longer pins the exact order (a
change-detector on the constant); it asserts one lap visits every grouping
once and returns to the start, and still fails when the cycle skips a step.

* chore: map contributor email for #104781 salvage

* fix(approval): anchor launchctl lookaheads to prevent GIL starvation

* fix(approval): deobfuscate every command word in one detection variant (#113535)

_command_detection_variants yielded one FULL-LENGTH variant per quoted or
escaped command word. A heredoc body of quoted lines ("key": "value", …) is
hundreds of quoted command words, so both detection passes scanned
O(words * len) characters: on a 16 KB / 460-line command the hardline pass
alone took 7.8 s and the dangerous pass 15 s even with the launchctl lookahead
anchored (33 KB: 36 s hardline), all while holding the GIL on the gateway loop.

Build a single variant with every command word deobfuscated instead. The
obfuscation catch ($(echo rm), r''m, ${0/x/r}m …) is unchanged — the same
deobfuscated words appear, in one string — and the variant count on the
issue's input drops from 923 to 4 (36 s -> 0.26 s hardline, verdict
unchanged). detect_hardline_command needs no bound and keeps its YOLO
semantics: a benign heredoc is neither blocked nor prompted.

Root cause identified in #113943 by @Tranquil-Flow (variant explosion as the
second compounding factor); its cumulative-work budget is superseded by
removing the explosion.

Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>

* test(approval): pin the launchctl perf subprocess to the checkout under test

A bare `python -c` resolves `tools` through the venv's editable install (the
primary clone), not the worktree pytest runs in, so the regression guard could
silently test a different tree. Pass cwd + PYTHONPATH like
tests/tools/test_async_delegation.py does.

* chore: map schwinguin's commit email for #109497 salvage

* fix(secret_scope): memoise the shared .env tokenizer once per file change

load_env_file() is now the canonical .env tokenizer after c849bc383a collapsed
six hand parsers onto it. Profile scopes, dashboard boundary discovery, skill
secret capture, managed .env and setup readers share the same decoding and
parsing rules, but direct calls still re-read and re-parse the whole file.

build_profile_secret_scope() puts that work on the gateway's hottest paths:
every turn, every cron job, MCP and browser lifecycle adoption, the TUI prompt
turn, and the 60s housekeeping drain of the durable cron delivery queue.
Memoising the shared tokenizer now removes redundant work for the consolidated
readers as well as the motivating profile-scope path.

config.load_env() keeps its existing outer memo over a call that is now itself
memoised; this leaves its menu-render shortcut intact without redesigning it.

Freshness is the hard part, because a stale credential map is a far worse
failure than a slow one. Every load_env_file() call still OPENS the file and
keys on os.fstat() of that descriptor rather than os.stat() of the path:

  * the open preserves close-to-open revalidation on NFS;
  * an open or read failure returns {} without caching it, and drops any warm
    entry so a transient EACCES cannot become a persistent empty profile;
  * the descriptor pins one inode, so a symlink repointed mid-read cannot file
    one file's contents under another file's identity.

The key is (mtime_ns, size, inode, device), re-checked on the same descriptor
after reading. A rewrite through that inode is not stored under the pre-read
fingerprint. Unavailable metadata means "do not cache", never "cannot read".
A generation counter checked under the lock prevents a slow reader from
repopulating an entry that a writer invalidated while the read was in flight.

Read bytes from the open descriptor and preserve upstream's decoding exactly:
strip a UTF-8 BOM before trying UTF-8, then fall back to latin-1 for invalid
UTF-8. Both the cached reader and the uncached reference use that decoder and
the shared text parser. Invalid UTF-8 is a cacheable result, not a read error.
The value and inline-comment helpers remain local to secret_scope.

invalidate_env_cache() clears this memo too, so save_env_value,
remove_env_value and sanitize_env_file invalidate both layers. The per-path
LRU holds at most 64 entries and every caller gets its own dict.

Accepted limitation: a write keeping length, inode and nanosecond mtime all
identical is not detected without explicit invalidation. This is a new
limitation for previously uncached readers. Parsed secrets also remain
reachable in memory between use and eviction for up to 64 paths, where the
outer config memo holds one.

The pre-rebase measurement on a 508-line, 25432-byte .env was 0.868 ms ->
0.011 ms per call. That measurement predates the shared binary decoder.

Add decode/cache invariants for BOM and plain files, latin-1 fallback,
same-size writes switching UTF-8 -> latin-1 -> UTF-8, and agreement with the
uncached reference. All eight decoding, cache and freshness mutations fail
the new tests, and all 40 cache tests pass after restoration.

The requested agent, CLI, cron and gateway suites report 30,168 passed,
479 failed and 303 skipped in the restricted local environment. All 140
failing files were checked against upstream/main at 5eb99eb284: all 479
failed-test IDs and 210 collection/setup/teardown error IDs match exactly.

* fix(gateway): skip heartbeat restore sweep when no heartbeat is persisted

The 5s goals poller calls restore_heartbeat_watches(), which listed every
routed session and entered _profile_runtime_scope for each distinct origin
— a full config.yaml/.env parse plus secrets hydration per session per
scan. With zero heartbeats persisted (the common idle case) the entire
sweep is wasted work; on a 400-session multiplex gateway it dominated
idle CPU (~60% of on-CPU time in YAML parsing).

Gate the sweep on an indexed heartbeat:* prefix probe of each served
profile's SessionDB (HERMES_HOME contextvar only — no config/secret
scopes). Probe failures fail open to the historical full sweep.

* fix(gateway): gate per-profile watcher ticks on actual work in the profile's store

The handoff watcher (2s interval) and loop wakeup watcher (15s) entered
_async_profile_runtime_scope for EVERY served profile on EVERY tick — a
full config.yaml/.env parse plus secrets hydration per entry — even when
the profile's state.db held zero pending handoffs or active loops. On a
multiplex gateway this, not message traffic, was the idle CPU driver
(py-spy: ~75% of on-CPU time in YAML parsing from these two paths).

Both watchers now probe the profile's cached SessionDB first (only the
HERMES_HOME contextvar installed, off the loop thread via the runner's
executor hop) and enter the scope only when work exists:
- handoff watcher: new SessionDB.has_pending_handoffs() existence probe
- loop watcher: list_active_loops() under the probe context

All probes fail OPEN (error/unknown -> historical always-enter
behavior); runners without the executor hop (test stand-ins) skip the
gate, so existing watcher tests are unaffected. The startup
stale-handoff reclaim stays ungated: it runs once per boot and must
also see 'running' leftovers, which a pending-only probe would miss.

* fix(gateway): watcher idle probes fail open on an unavailable store and ignore cleared rows

The loop watcher's probe used ``list_active_loops()``, which collapses an unavailable
SessionDB (and any read error) to ``[]`` — the gate then skipped that profile on every tick
(fail CLOSED). The heartbeat probe keyed on ``heartbeat:*`` key existence, but ``clear`` and
``pause`` keep their rows, so one past /heartbeat re-enabled the full sweep forever.

Both probes now read the rows through one ``_profile_meta_rows`` helper (None = store
unavailable → treat as work present) and parse the status themselves: only an ACTIVE row
opens the gate; a corrupt row is "unknown" and keeps the historical scan. The heartbeat
gate reuses ``_handoff_watch_scopes`` for the served homes instead of a second resolver.

* test(gateway): two invariant tests for the watcher idle gates

Heartbeat: no active row → no sweep; a cleared row → still no sweep; an active row → sweep;
an unopenable store → sweep. Loop watcher: idle profile never scoped, active loop scoped,
unopenable store scoped. The handoff-watcher test from #109497 is dropped: same shape, and
the salvage bar is two tests per fix.

* refactor(secret_scope): key the .env memo on utils.file_signature

The memo had its own (mtime_ns, size, inode, device) tuple and documented a
same-length/same-inode/pinned-mtime rewrite as an accepted gap. utils.file_signature is the
repo's change-detection key (config.load_env, skill_utils, prompt_builder already use it) and
includes ctime_ns, which user space cannot backdate — so the gap closes and one helper goes.

* refactor(config): load_env is a thin wrapper over the memoised tokenizer

load_env() kept its own (path, file_signature) memo on top of load_env_file(), which now memoises
on the same key: two stat calls and two caches to keep coherent for one answer. Drop the outer
layer; invalidate_env_cache() stays as the writers' knob and forwards to the one memo.

* fix(gateway): heartbeat idle gate probes the default profile under a named-profile multiplexer

_watched_homes reused _handoff_watch_scopes, which deliberately omits "default" because the
handoff/loop watchers' root poll covers the launch home. For a `hermes -p work` multiplexer the
launch home is profiles/work, so ~/.hermes was probed by nobody and an active heartbeat there
answered "idle" every tick. Probe the served set the sweep can actually resolve to.

* refactor(gateway): idle gates live in run_idle_gates; row semantics live with loops/heartbeat

The probe helpers had been appended to the gateway/run.py facade and the three watchers each
carried their own copy of the "probe → fail open → offload" shell, with gateway code parsing
loop/heartbeat rows itself. One sibling now owns the gate shape (_gate + off_loop_gate), and
hermes_cli.loops.store_has_active_loop / hermes_cli.heartbeat.store_has_active_heartbeat own
what an ACTIVE row is (heartbeat gains the _META_PREFIX loops already had). One fail-open layer
per read: has_pending_handoffs is a bare bounded query, the gate catches. The loop watcher calls
its executor hop directly — it already did so unconditionally for the scan, so the
"runner without a hop" fallback there was dead.

* test(gateway): idle-gate tests read status, cover the named-routing-home case and the handoff gate

Loop watcher: a cleared row must not open the gate (a key-existence mutation now fails), then
the unavailable-store fail-open. Heartbeat restore: routing home = named profile with the only
active row in the default store must still sweep. The handoff-watcher test from #109497 is
restored (its gate and the has_pending_handoffs query had no coverage).

* fix(desktop): install rebuilt app to system location after CLI hermes update

Port of #59942 onto post-#102117 main. The refactor moved the desktop
helpers from hermes_cli/main.py into hermes_cli/main_desktop.py, so the
new install machinery (_app_asar_hash, _macos_adhoc_sign_bundle,
_desktop_bundle_install_supported, _swap_in_new_macos_bundle,
_install_rebuilt_desktop_app) lives there now, re-exported through
main.py's frozen updater surface block so _m() resolves them.

_rebuild_desktop_after_update (hermes_cli/update_cmd_deps.py) now, on a
successful --build-only build, installs the rebuilt bundle to the
system location (macOS .app via ditto staged copy + atomic swap with
rollback; Windows via copytree), skipping when the installed app.asar
hash already matches. Linux stays owned by its package mechanism.

Without this, a CLI hermes update leaves the installed app stale —
the in-app updater covers its own path, but the CLI path never swapped
the installed bundle.

Tests: 17/17 ported suite (import/patch targets repointed to
main_desktop) + test_cmd_update.py 47/47 green.

* fix(update): scope the installed-bundle refresh to macOS, keep the signature, never swap under a live app

Rework of the salvaged #59942 hook so it fixes the whole #52339 class without
regressing what the Desktop updater already gets right:

- macOS only. The Windows arm rmtree'd a possibly-running NSIS install
  (partial deletion under a lock) and windows.ps1 already owns that swap;
  Linux packages stay with their package manager.
- No re-sign. The rebuilt release/ bundle already carries the stable local
  signing identity from _desktop_macos_relaunchable_fixup; a deep `codesign -s -`
  on the installed copy replaced it with a fresh ad-hoc cdhash and reset every
  TCC grant. ditto preserves the signature, so nothing is signed here.
- Running bundles are reported, not swapped: Electron loads app.asar and helper
  apps lazily, so renaming the bundle away and deleting the old tree crashes
  the live app. The detached updater waits for exit; a terminal `hermes update`
  with the app open now prints what to do instead.
- Failures are printed as warnings; the old `Path | None` return read every
  failure as "Desktop app up to date".
- The refresh also runs on the "build stamp current" path, so a stale
  /Applications copy left by an earlier update heals on the next `hermes update`
  even when there is nothing to rebuild.
- Core is host-independent (_install_rebuilt_macos_bundles takes paths as data);
  the two invariant tests run on every OS instead of `skipif(darwin)` tests that
  ran nowhere.

Covers the Desktop-button path too: posix.sh runs `hermes update`, so an app
running from apps/desktop/release/ now refreshes the /Applications copy Finder
launches (the Discord report: new shell right after the update, old shell on
the next Dock launch).

* fix(update): repoint Windows managed runtime without renaming live venv

Cherry-pick of helix4u's #88836 (75742db5a42) resolved onto current main.

The POSIX runtime cutover parks the live venv with a directory rename. On
Windows that can never succeed from `hermes update`: the updater executes
from venv\Scripts\python.exe and Windows keeps that image mapped, so the
rename fails with WinError 5 on every attempt (#93032). Instead, atomically
replace the live venv's pyvenv.cfg with the candidate's: Scripts\python.exe
is a launcher that reads `home` on every start, so every fresh process runs
the new generation while nothing touches the locked directory. A failed
post-cutover smoke restores the original config; a failed rollback keeps
the generation the live config now references.

* fix(update): let the Windows repoint run; refuse a minor-line jump

The self-lock/holder preflight (#99711) deferred the repair on the theory
that the updater's own mapped python.exe makes the venv rename impossible.
Live on windows-latest, a process executing from venv\Scripts\python.exe
(and the repo's real .venv with cp311 .pyd extensions loaded) does NOT block
the rename; what blocks it with WinError 5 is any ordinary handle under the
tree: a process cwd, an open file, a sync client. Since sys.executable is
always under the venv on Windows, the deferral fired on every run and no
Windows install could repair from `hermes update`. Retire it and its tests;
the pyvenv.cfg repoint needs no rename and has none of that exposure. The
wine2e lane now runs the cutover probes instead.

The repoint keeps the live venv's site-packages, so provisioning's
fall-forward to the next minor (3.11 -> 3.12, #76106) must not be pointed at
a cp311 tree: refuse before touching pyvenv.cfg. The live Windows test holds
the venv the way the field does (a child with cwd inside it), asserts the
rename path fails (the symptom) and the repoint succeeds; the rollback and
minor-guard tests run on every host.

* fix(title): keep the derived title while skipping the model upgrade (#85194)

`auxiliary.title_generation.enabled: false` turned off both title stages, so an
operator who only wanted to stop the background model call (unavailable or
metered endpoint) also lost the instant derived title. `model_upgrade_enabled:
false` keeps the derived title and starts no `auto-title` thread; missing keeps
the two-stage default and `enabled: false` still disables both.

Salvaged from #85401 onto current main: the gate reuses `_title_config()` and
sits before `spawn_context_thread` (the thread seam moved off `threading.Thread`).

* chore: map contributor email for lightcloud00

Non-noreply author email on the salvaged #85401 commit.

* fix(aux): model_upgrade_enabled only silences the automatic title upgrade

The gate also sat inside `generate_title`, so the explicit operator repair
command `hermes sessions retitle-skills` returned None for every row when the
toggle was off, although the operator asked for a model call. The toggle's
contract is "never spend a model call upgrading the instant title": only
`maybe_auto_title` (the background path) is gated now; a direct
`generate_title` call still asks the model. Docs sentence adjusted.

Review follow-up on #113955; the existing toggle test now also asserts the
explicit path still titles (red on the previous head).

* fix(agent): use provider-default temperature for title generation (#72351)

Title generation hardcoded temperature=0.3 in its call_llm() call.
Models like GPT-5.6 only accept their server-side default temperature
and reject explicit values with "Unsupported value: 'temperature'".
While call_llm() has a retry that strips temperature on error, the
daemon thread races with session cleanup in short-lived CLI sessions,
causing the retry to fail with a connection error.

Fix: pass temperature=None so the provider uses its own default.  This
avoids the unsupported-temperature error entirely and eliminates the
need for the retry path.

Fixes #72351

* fix: chain auxiliary parameter-rejection retries on primary and fallback requests

Reasoning models reject several request fields at once (gpt-5: temperature AND
max_tokens, #78273), a reasoning-strip retry can then 400 on temperature
(#72351), and strict-schema gateways such as Fireworks reject the generic
extra_body.reasoning fallback with "Extra inputs are not permitted, field:
'reasoning'" (#109774) — a phrasing none of the unsupported-parameter predicates
recognised, so the reasoning rung never fired and every title call 400'd.

The parameter rungs were a fixed single-pass order (temperature → structured
output → reasoning → max_tokens): a field rejected AFTER an earlier rung had
already run was never stripped, and _param_rung_accepts did not admit a
temperature 400 raised by a later retry. They now form a table walked until no
rung matches, each field stripped at most once, so N rejected fields recover in
N retries in whatever order the provider raises them.

Fallback candidates (per-task chain, main chain, discovery) got a single shot
with the caller's temperature/max_tokens/reasoning fields and raised on the
first parameter 400; agent/auxiliary_fallback_recovery.py now runs the same
parameter rungs around a candidate's request (sync + async), leaving auth /
payment / connection errors to the caller's existing handling.

Live: real gpt-5-mini title_generation call (reasoning_effort 400 → temperature
400 → 200), real gpt-5-mini fallback candidate sync+async (temperature 400 →
200), and a stand-in replaying Fireworks' documented 400 (reasoning stripped →
200) — all failed on origin/main.

* fix(aux): build the fallback ladder route by field name

A positional _LadderRoute(...) in auxiliary_fallback_recovery breaks as soon as the
route tuple gains a field (#113968 adds timeout); construct by name so fields the
parameter ladder never reads default to None regardless of tuple width.

* fix(aux): fallback ladder route carries the client base_url; test builds the route by name

base_info is the endpoint identity every rung reads for per-route bookkeeping; the
fallback ladder left it empty. The test mirrored the positional construction that
broke under a wider tuple.

* fix(aux): parameter rungs hand a 429 on, and Fireworks gets top-level reasoning_effort

The pre-ladder max_tokens rung accepted payment|connection|rate_limit errors on
its retry, so a 429 after a stripped retry reached the credential and
provider-fallback rungs. The ladder's shared `_param_rung_accepts` dropped
rate_limit: a 429 on the retry raised out of the primary call and skipped the
whole fallback chain. Accept it there too, with a rung-level test that the
stripped kwargs and the 429 are handed on rather than raised.

The body claimed `Fixes #109774` while the aux client still sent the generic
`extra_body.reasoning` to Fireworks on every call and only recovered
reactively (an extra 400 round-trip per request). Port @huklaa's profile
override from #109807: Fireworks documents top-level `reasoning_effort`
(`none` disables thinking), and overriding `build_api_kwargs_extras` marks the
profile reasoning-aware so the transport omits the generic fallback on this
route. Extends the salvage with the effort mapping so enabled-with-effort takes
the same wire, with a control that a profile-less route keeps the fallback.

Salvages #109807 (@huklaa).

Co-authored-by: Hukla <129692708+huklaa@users.noreply.github.com>

* fix(aux): max_tokens rung retries even when the wire kwargs no longer carry the cap

The rung table refused to re-send an unchanged request, but the Codex Responses
route translates the caller cap away and gateways inject their own: the 400 names
max_tokens while the kwargs show none, and the identical retry is what completes
(tests/agent/test_injected_param_strip_retry_registry.py, red in CI on 99ff3b9).

* fix(anthropic): merged user turns keep each turn as its own text block

_concat_content joined two string user contents into one "a\nb" string when
_merge_consecutive_roles collapsed adjacent user turns for the Anthropic
Messages wire. For a MoA aggregator on that wire, iteration 1 of a turn ends
[user(task), user(guidance)] and was sent as user("task\n<guidance>"), while
iteration 2 replays user("task") alone, so the prompt-cache prefix diverged
at the first user block and the #112358 collapse persisted there (the
first-pass fix in #113175 only covered the OpenAI-compatible wire).

Merged turns are now always a block list with each side's blocks intact (a
string becomes one text block), matching what the list+list and list+str
shapes already did. The task block is byte-identical to the standalone turn
later iterations replay, a cache_control marker on it stays put, and the
guidance follows as its own text block. Assistant merges are unaffected
(assistant content is already a block list).

Bedrock Converse and native Gemini already merge at block/part granularity,
so the docs' Anthropic/Converse/Gemini fold caveat is replaced with the
accurate statement and the byte-identical-extension claim no longer needs
the OpenAI-compatible scope.

Part of #112358

* test(mcp): drive a poisoned client.json through _configure_callback_port end-to-end

The #112568 coverage exercised _cached_redirect and get_client_info in
isolation; the issue promised a flow-level case. This drives the login
path's actual sequence -- _configure_callback_port(cfg, storage) then the
SDK's storage.get_client_info() -- with the issue's exact bad-port payload
and a non-dict client.json, asserting a freshly parked ephemeral port and
"no registration", so a regression on either step fails the caller path,
not just the private helper. Red before f8f89b20d27 (both cases) and
before e166791f9ff (non-dict case); green on main.

Closes the remaining atom of #112568.

* fix(skills): skill_manage advertises one shape per action; a misfiled text slot fails before any op applies

The operations[] item schema was one flat object with four coexisting text
slots (content / new_string / file_content / file_path). A 27B local model
that had just used write_file's file_content kept emitting it on create and
patch ops; the call validated against the advertised schema, the handler
failed on "content is required", and the whole batch rolled back — eight
identical retries until the tool-loop guardrail tripped (#112677).

- items is now an anyOf of self-contained per-action op objects (create,
  patch targeted, patch full-rewrite, write_file, remove_file, delete), each
  with additionalProperties: false. The wire shape of a correct call is
  unchanged (still a flat op with name/action/...), so transcripts, staging
  and replay are untouched; grammar-constrained backends can no longer emit
  another action's slot, and schema-validating providers reject it up front.
  Nested (non-top-level) anyOf survives every sanitizer (schema_sanitizer,
  Gemini legacy translator). Cost: parameters JSON 1352 -> 2414 bytes.
- _validate_batch_ops runs the per-op argument-shape check (_op_shape_error)
  before any op is applied, so a misfiled op[1] no longer applies op[0] and
  then rolls the batch back; the error carries the same misplaced-key hint.
- The hint is attached to argument-shape misses only: a patch whose real
  problem is an unmatched old_string is no longer told to "move that text to
  'content' (full rewrite)", the escape the patch error itself warns against.
- Shape tables (_REQUIRED_ARGS, text-slot maps, _misplaced_text_hint,
  _op_shape_error) move out of the facade into skill_manager_batch.py, the
  op-validation sibling; _patch_skill shares the old_string guidance text.
- Docs: skills.md Actions table states the one-slot-per-action contract.

* fix(skills): keep absorbed_into in the delete shape; decide every patch shape miss pre-effect

The per-action schema made the delete branch `additionalProperties: false`
with no `absorbed_into`, so a schema-validating or grammar-constrained
backend could no longer emit the curator's consolidation delete and
`_curator_consolidation_delete_guard` fail-closed every consolidation.
Advertise `absorbed_into` on the delete branch and cover the batch path
that forwards it to the guard.

Also fold the two remaining patch shape checks (missing new_string,
content mixed with old_string/new_string) into `_op_shape_error`, so a
batch rejects them before applying any sibling instead of creating op[0]
and rolling it back. Drop the stale `edit` vocabulary from skills.md:440
and the curator prompt, and move the schema-diet test helpers above the
`__main__` guard.

* test(curator): stop pinning the unadvertised edit alias in the read-before-write prompt

The prompt now names patch/write_file/remove_file only; edit is the legacy
alias of a full-rewrite patch and no longer appears in the advertised
vocabulary, so the prompt assertion must not require it.

* chore: map sy0u1ti for salvage of #114024

* fix(sessions): include live session activity in prune recency

* refactor(sessions): pinned back-fill uses the canonical last-active expression

The list_sessions pinned back-fill query still hand-rolled
COALESCE(MAX(m2.timestamp), s.started_at) AS last_active while the main
query a few lines above already used _sql_session_last_active("s").
Use the shared helper so pinned rows report the same freshest-of
(last_activity_at / latest message / started_at) recency as every other
listing and prune path, instead of ignoring live activity.

* docs(sessions): recency wording matches the freshest-of rule

Prune recency is now the freshest of last_activity_at / latest message /
started_at (via _sql_session_last_active), but four doc sites still said
"latest message, else started_at": the list_prune_candidates docstring,
the --older-than/--newer-than help in session_filters, and two comments
in config_defaults. Reword them to the phrasing archive_stale_sessions
already uses so operators are not told live activity is ignored.

Also make archive_stale_sessions use the module constant _LAST_ACTIVE_SQL
instead of an inline _sql_session_last_active("s") call — same alias,
identical SQL, one fewer place to drift.

* docs(sessions): user guide describes the freshest-of recency rule

The --older-than/--newer-than and prune-retention paragraphs still said recency is the latest message (falling back to session start). After the prune-recency fix, both paths use the freshest of live activity / latest message / session start; reword to match the phrasing in hermes_cli/session_filters.py.

* fix(tools): refuse text writes into database files and binary overwrites

The text tools' read->modify->write round-trip re-encodes binary bytes
lossily: the terminal transport decodes stdout with errors=replace, so
a patch on a SQLite database reads mojibake and writes it back, silently
destroying the file (observed in the wild: a kanban board DB lost 550k
bytes / 26 days of history this way on 2026-09-15).

read_file already refuses to display binary files (extension + byte
sniffing), but write_file/patch had no write-side equivalent, and
patch_replace reads via plain cat so the read-side guard never fires
for it.

Guard rules (extends _check_binary_document_write, the existing
binary-document write guard):

- .db/.sqlite/.sqlite3 and their -wal/-shm/-journal sidecars: ALWAYS
  refused (even new-file creation) — plain text can never be a valid
  database, mirroring the opaque-document rule. The sidecars need a
  basename regex because os.path.splitext('x.db-wal') yields '.db-wal',
  which is not in BINARY_EXTENSIONS.
- Other BINARY_EXTENSIONS (.png, .zip, ...): refused when OVERWRITING an
  existing file, mirroring the existing .pdf-overwrite rule — the model
  can only have read mojibake, so writing back destroys it. Creating a
  NEW file with such an extension stays allowed.

Tests: new test_binary_db_write_guard.py covers the regex (main names,
sidecars, non-db names), the guard function, and the write_file/patch
tool paths against real SQLite files (integrity_check still ok, bytes
untouched). 29 new+existing guard tests pass; file-tools/patch suites
green (196 tests).

* refactor(tools): recognise SQLite sidecars in binary_extensions, drop write-guard regex

Move the .db-wal/-shm/-journal detection from a write-guard-local regex
into tools/binary_extensions._has_extension_in so has_binary_extension
(read guard AND write guard) agree: a sidecar counts as its database's
extension. read_file now refuses sidecars instead of returning lossy
text, which also gives write_file the no-baseline overwrite refusal.

Behaviour change vs the contributor's commit: creating a NEW .db /
.sqlite file via write_file stays allowed (main allows it and text
fixtures named *.db exist) — only overwriting an existing binary is
refused. The generic has_binary_extension overwrite branch is kept and
merged with the PDF branch because patch has no full-read baseline
check: on main, patch on an existing .db with a matching old_string
rewrites the header in place.

Tests folded into the existing guard test classes (one write_file, one
patch refusal on a real WAL sidecar); the PR's 12-test file and its
private-regex assertions are dropped.

* test(tools): pin the binary refusal for write_file into a SQLite -wal sidecar

Any error would also be produced by the no-baseline overwrite guard on
main, so the test could not detect a regression of the sidecar
detection. Asserting the binary-file refusal makes it fail when
binary_extensions stops recognising .db-wal.

* fix(tools): a SQLite sidecar path is refused even when no sidecar file exists yet

The refactor that moved sidecar detection into binary_extensions dropped the
contributor's unconditional refusal for the db family and made every binary
extension a plain overwrite guard. That is right for the MAIN .db/.sqlite
(text fixtures named *.db exist), but a -wal/-shm/-journal path is never a
legitimate text target: a checkpointed database has no sidecar on disk, so
write_file("kanban.db-wal", text) silently dropped a garbage WAL next to a
live database.

Add is_sqlite_sidecar(path) and refuse such paths in
_check_binary_document_write regardless of is_file(). Scope the marker
stripping in _has_extension_in to SQLite suffixes so report.docx-wal no
longer counts as an opaque document. Hoist the duplicated is_pdf_path call.
Parametrize the write_file WAL test over sidecar existing/absent.

* fix(tools): sandbox read paths recognise SQLite sidecars as binary

file_operations.py checked `os.path.splitext(path)[1].lower() in
BINARY_EXTENSIONS` at five sites, so a `.db-wal` / `.sqlite3-shm` sidecar
(final suffix never in the set) was read and edited as text in the sandbox
paths. Route all five through has_binary_extension(), which strips the
sidecar marker. Behaviour is otherwise identical: same lower-cased final
suffix membership (rfind vs splitext differ only on leading-dot basenames
like `.bashrc`, which are in neither set). binary_extensions imports nothing
from tools, so no cycle.

* test(tools): honest sidecar fixture

_make_wal_db claimed sqlite keeps the -wal until checkpoint, but SQLite
deletes -wal/-shm on last-connection close, so the fake-bytes fallback was
what actually ran. Make the fixture a context manager that holds the
connection open across the tool call so the sidecar is a real WAL, and drop
the PRAGMA integrity_check assertion (SQLite ignores a garbage WAL, so it
passed either way); keep the byte-identity check.

In the patch test assert "binary" in the error so the no-baseline guard
cannot mask a regression in sidecar detection.

* refactor(tools): one suffix helper for the binary-extension checks; drop the dead .db3 member

_has_extension_in and is_sqlite_sidecar each re-implemented the same
rfind('.') / -1 / .lower() suffix extraction, and the write guard used a
third spelling (filepath[filepath.rfind("."):]) for its refusal messages
next to os.path.splitext in the overwrite branch. Add _lower_suffix()
("" when no dot) and route both predicates through it; the write guard
now uses os.path.splitext for every message's displayed extension.

_SQLITE_EXTENSIONS was `{.db,.sqlite,.sqlite3,.db3} & BINARY_EXTENSIONS`,
which silently dropped .db3 (not a binary extension anywhere in the
codebase, so `x.db3-wal` was never a sidecar). Spell the set as the three
members that actually take effect and assert the subset invariant instead
of hiding it behind an intersection. Behaviour is unchanged.

* test(tools): overwriting the database file itself is still refused as binary

The test trim dropped the headline case: text written over an EXISTING .db. Sidecar paths return at is_sqlite_sidecar before the overwrite branch, so a mutant that drops has_binary_extension from the overwrite condition survived every remaining test. Add a third parametrize case that targets the held-open WAL database file itself and asserts the binary refusal with bytes unchanged.

Same commit: the _check_binary_document_write docstring now names SQLite sidecars (always) and every BINARY_EXTENSIONS suffix (on overwrite), and the suffix is computed once at the top instead of three times.

* chore: map holny for salvage credit of #113643

* fix(telegram): rebuild adapter on confirmed polling stall instead of reusing unquiesced updater (#113618)

* refactor(telegram): type the polling-stall signal instead of matching log text

Replace `_looks_like_polling_stall` (a three-substring classifier over
`str(error)`) with `_PollingStallError(RuntimeError)`. The watchdog and the
post-reconnect verifier raise it at their two stall sites; the reconnect
ladder checks `isinstance` through `_iter_exception_graph` before handing
the adapter to the supervisor.

WHY: a substring match couples recovery routing to log wording (rename a
message, silently lose the fatal handoff) and can misfire on unrelated
errors that quote the same words. A typed exception keeps #113657's
behaviour with no behavioural change beyond the classifier's source.

* test(telegram): fold stall handoff cases into one parametrized invariant

Merge the watchdog and verifier stall tests into a single parametrized
`test_confirmed_stall_hands_off_before_reusing_updater` and assert the
adapter is also marked degraded, keeping two invariant tests for #113618
(stall -> fatal handoff; generic network error -> in-place reconnect).

* refactor(telegram): watchdog stall reuses _schedule_polling_recovery

`_check_polling_stall` hand-rolled the body of `_schedule_polling_recovery`
(set `_send_path_degraded`, `_mark_degraded()` when running, spawn
`_handle_polling_network_error`). Call the helper instead, matching the
post-reconnect verifier stall site.

WHY: one recovery entry point keeps degraded-state marking and the
"polling degraded (reason)" log line consistent across every scheduler.
The helper's early-return guards (`_teardown_started or has_fatal_error`,
`_recovery_in_flight()`) are already asserted at the top of
`_check_polling_stall` with no await in between, so they are no-ops here
and behaviour is unchanged.

* fix(telegram): a confirmed stall hands off before the reconnect backoff, not after it

Hoist the `_PollingStallError` check in `_handle_polling_network_error` to
right after the teardown/fatal guard, before the retry counter increment,
the exponential sleep, `_stop_updater_or_go_fatal` and both connection
drains. Use a plain `isinstance` (both raise sites construct the error
directly; nothing wraps it). The stall test now also asserts no sleep, no
retry-counter bump, no in-place `updater.stop()` and no drain.

WHY: the check sat after `await asyncio.sleep(delay)` (5-60 s), the
counter bump and a bounded `updater.stop()` (up to 15 s), so a gateway
already confirmed deaf stayed deaf 5-75 s longer and consumed a retry
slot for something that is not a retry.

The in-place `updater.stop()` before the handoff is dropped deliberately:
`_go_fatal_network` -> `_handoff_polling_fatal_error` -> supervisor
rebuild runs `disconnect()`, which performs the same bounded
`updater.stop()` (`_UPDATER_STOP_TIMEOUT`, falls through on timeout) plus
`app.stop()/shutdown()`, so stopping here only duplicated that work on the
slow path. Docstrings for `_handle_polling_network_error` and
`_check_polling_stall` no longer describe the stall as a reconnect-ladder
escalation.

* refactor(telegram): a confirmed stall logs one hand-off line, not a retry promise

_schedule_polling_recovery promised the gateway 'stays alive and will retry' for every error, but a _PollingStallError goes straight to _go_fatal_network (supervisor rebuild). Branch the wording on the error type, drop the watchdog's own pre-log so a stall yields exactly one error-level line (from _go_fatal_network), and carry stalled_for/generation in the stall error text instead. Test module docstring and test name updated to match the hand-off semantics.

* fix(mcp): avoid default keepalive on stdio servers

* fix(mcp): wake recycle deadlines after active RPCs

* docs(mcp): clarify transport keepalive defaults

* fix(mcp): an idle stdio server still proves its session without a keepalive ping

After the stdio keepalive default became None, the lifecycle loop's
`if keepalive_interval is None: continue` skipped the only idle path to
`_mark_session_proven()`. `_session_proven` is otherwise set only on a
tool-call success or a suspect health-check pass, and each new session
resets it, so an idle default-config stdio server never became proven:
every later child death/reconnect charged the rapid-drop budget that a
180 s idle survival used to clear (#62212 regression).

While the session is unproven, keep the `asyncio.wait` timeout at
`_DEFAULT_KEEPALIVE_INTERVAL` for stdio too; when it fires with the child
still alive, mark the session proven without pinging. Once proven the
stdio timeout returns to None, so a healthy idle pipe is never probed.
Shutdown/reconnect precedence and the rpc_idle waiter are untouched.

Test: `_captured_interval` now runs two lifecycle cycles (first times
out, second shuts down) through a shared helper so the stdio test asserts
timeouts [180, None], zero pings and `_session_proven` set, instead of
copy-pasting the fake `asyncio.wait`. Docstring/comment updated to the
new rule.

* refactor(mcp): lifecycle loop reads its inputs once and stops re-deriving the recycle predicate

WHAT: `_wait_for_lifecycle_event` hoists `is_http = self._is_http()` (it only
reads the immutable `"url" in self._config`) and reads `keepalive_interval`
from the config once instead of `in` + `.get()`. The one-shot RPC-idle waiter
is now spawned on `not is_http and self._rpc_lock.locked()` alone, dropping
the inline `_idle_timeout_seconds is not None or _max_lifetime_seconds is not
None` clause. The waiter list is built by one comprehension per iteration
(rpc_idle_task changes between iterations) and the `finally` reuses the last
value; `_cancel_waiters` already skips done tasks. The docstring folds the
#17003 keepalive paragraph into the first paragraph instead of describing the
keepalive twice.

WHY: the dropped clause was the inverse of
`MCPServerHealthMixin._stdio_recycle_deadlines` re-derived by reaching into
health-mixin attributes. Without it, a stdio server with no limits and an
active RPC spawns one trivially-completing waiter per RPC; on lock release
the loop takes one extra iteration where `_recycle_if_due()` is False and
`_next_stdio_recycle_deadline()` is None, so it simply recomputes the timeout
and waits again — harmless.

* test(mcp): one recycle-wake case

WHAT: drop the idle-vs-lifetime parametrize on
`test_stdio_recycle_wakes_after_active_rpc`, keeping the idle case, and drop
its `_recycled_reason == reason` assertion.

WHY: both parameters exercised the same new invariant (releasing the RPC lock
wakes the lifecycle loop so the now-visible deadline recycles the server); the
reason assertion pinned pre-existing `_stdio_recycle_reason` behaviour that is
not this stack's concern. The stack now adds exactly two test functions.

* fix(mcp): an idle stdio server is proven only after a full default interval

The no-keepalive proof branch ran on ANY timeout wake. With idle_timeout_seconds=60 the wait is min(180, 60); an overlapping RPC hides the recycle deadline so _recycle_if_due() is False and the session was marked proven after 60 s, clearing the #62212 rapid-drop budget early. Fix a proof_at deadline before the loop and gate the branch on it; the unproven timeout is the remaining distance to proof_at. Tests: the stdio test now shows the first (idle-limit) wake does not prove; the interval test covers an explicit keepalive_interval on stdio.

* fix(desktop): reasoning blocks stop eating streamed prose

preprocessMarkdown runs on the accumulated text on every streaming flush, so
the plain REASONING_BLOCK_RE replace damaged the visible answer in two ways a
settled message never shows:

- a model that inlines its chain of thought in the answer channel streams the
  open tag long before the close one, and the regex needs both, so the reasoning
  rendered as chat prose until the close tag arrived — and when it did the whole
  span vanished in one frame, taking text the reader had already read with it
  (#62774: the mid-reply "truncation" on long answers).
- the regex ate the whitespace on its seam, so a block sitting between two words
  fused them: `no` + `Hermes` rendered as `noHermes`.

An unterminated block now strips at a block boundary — the same rule
agent/think_scrubber.py draws, so a real reasoning preamble (always its own
block) disappears while prose that merely mentions `<thinking>` mid-sentence
survives — and a removal between two prose fragments leaves one space instead of
deleting the seam.

Tests: markdown-reasoning-stream.test.ts (module level: seam, unterminated
blocks, per-delta leak + visibility monotonicity) and
streaming-text-fidelity.test.tsx (real surface: accents/emoji and plain prose
streamed in 1- and 3-character deltas, plus the chain of thought never painting
a frame).

* refactor(desktop): align reasoning tags with think_scrubber, trim #113302 tests

WHAT: one REASONING_TAGS alternation feeds both regexes and now also covers
`thought` and `reasoning_scratchpad` (agent/think_scrubber.py THINK_TAG_NAMES);
the desktop-only `scratchpad`/`analysis` stay. Behaviour change beyond the
contributor's: closed or unterminated `<thought>`/`<reasoning_scratchpad>`
blocks are now stripped on the desktop too.

Doc comments trimmed to the WHY (block-boundary rule, why an unterminated block
is held back). Tests cut to the ≤2-invariant shape: two unit cases (seam,
unterminated-vs-mention) and one DOM frame-by-frame case; the accents/plain-prose
DOM cases were protection for unrelated code and are dropped.

* perf(desktop): reasoning-block seam check reads two chars, not two string copies

stripReasoningBlocks runs on the accumulated message text every streaming
flush (markdown-text.tsx). Its replacer sliced the whole text twice per
closed block to test whether a space is needed at the seam — O(n) copies plus
a forward scan per block per flush. Read the single char on each side of the
match instead.

The seam test also had a second defect: with two adjacent closed blocks the
second block's `before` ended in the first block's `>`, so a second space was
emitted (`no  Hermes`). Matching a run of adjacent blocks as one match makes
the two-edge check see the real prose on both sides. Covered by an extra
expect in the existing seam test.

* fix(desktop): a half-arrived reasoning tag at a block boundary is held back too

OPEN_REASONING_BLOCK_RE needs the full `<tag>`, so a frame ending in `<thin`
at a block boundary painted `<thin` as prose for one frame and then erased it
— the paint/un-paint class of #62774, one frame long. The fidelity test carved
an escape hatch around exactly those frames.

Add a third pass that drops a trailing partial open tag at a block boundary,
restricted to prefixes of the known reasoning tag names (built from
REASONING_TAGS) so `<div` at a line start still renders. This is the desktop
counterpart of agent/think_scrubber.py `_hold_partial`/`_max_partial_suffix`,
cited rather than ported. The fidelity test's escape hatch is deleted so the
monotonic-prefix invariant holds on every frame; the unit test gains the
`<thin` positive and `<div` negative cases.

* test(desktop): reasoning tests follow the sibling naming

The lib test sits beside markdown-preprocess.directives.test.ts as
markdown-preprocess.reasoning.test.ts, and the frame-fidelity test joins
its markdown-text.*.test.tsx siblings as markdown-text.reasoning.test.tsx,
so a glance at the directory shows what each file exercises. Also trim the
REASONING_TAGS comment: how the two tag lists converged is git history, not
something a reader of the constant needs.

* chore: map cswrld-net for salvage of #112791

* fix(desktop): answer server→client requests when a handler crashes instead of stalling the backend

A clarify request renders as an eternal spinner when anything in the
renderer's handler chain throws: deliverRequest had no error handling,
the WS message listener let the exception escape as an uncaught error,
and both dispatch sites silently no-oped via optional chaining when the
registry was absent. In every case the backend (clarify_tool blocks up
to 3600s) never receives any frame — no result, no error — and waits
out its whole deadline.

- json-rpc-channel: wrap each handler invocation; a crash now answers
  -32603 ("server request handler crashed: <method>"), fires the
  onUnhandledRequest hook, logs the stack, and stops. The unhandled
  path still answers -32601.
- json-rpc-gateway: guard the socket message listener so no frame can
  escape as an uncaught error.
- store/gateway (primary + secondary wiring): missing registry now
  fails the request immediately with -32601 instead of dropping it.

Backend already handles {"error"} response frames
(tui_gateway/server_requests.resolve_response), so fail-fast answers
settle the tool at once; no backend change needed.

Regression test drives boom→-32603, unknown→-32601, and a working
request after the crash.

* refactor(desktop): drop the socket-listener try/catch and share the registry-missing guard

Follow-up to the salvage of #112791.

- json-rpc-gateway.ts: revert the try/catch around `channel.handleFrame`
  to main. deliverRequest is the single chokepoint that dispatches into
  feature handlers and now answers -32603 itself; a second catch in the
  socket listener is defense-in-depth that would also hide bugs in event
  and response handling that should surface as uncaught errors.
- store/gateway.ts: the primary and secondary dispatch sites carried the
  same copy-pasted "no registry -> fail -32601" block. Fold both into one
  file-local `dispatchServerRequest(request, profile, connectionId)`.
  The secondary keeps main's tagging (its own connectionId, no fallback
  to the active connection), which the contributor's version changed.

* refactor(shared): a crashed server-request handler reports through an owner hook, not console.error

The deliverRequest catch branch wrote to a module-level console.error sink
(the only direct console.* in apps/shared/src) and fired onUnhandledRequest,
whose contract is "nobody handled it, already answered -32601" — so the TUI
logged a -32603 crash as "unhandled server request".

Add onRequestHandlerError(error, request) to JsonRpcRequestChannelOptions
beside onHeartbeatFailure, call it from the catch after answering -32603,
and drop the console sink. Wire both owners: HermesGateway (desktop, via a
GatewayClientOptions passthrough) logs to console.error like its dial-failure
sink; ui-tui gatewayClient pushes a [protocol] log line. Collapse the two
normalisation arms into the existing `error instanceof Error ? … : new
Error(String(error))` idiom and restore the early `return true` instead of
the handled flag + break — nothing runs after the loop but the -32601
fallthrough.

Test: the crash case now asserts onRequestHandlerError fires once for the
-32603 request and onUnhandledRequest only for the -32601 one.

* test(desktop): registry-missing server request answers -32601

Mutation check: reverting store/gateway.ts to the merge-base left the desktop
suite green — nothing exercised dispatchServerRequest. Cover both arms through
dispatchPrimaryServerRequest: a registry configured without onServerRequest
answers request.fail(-32601, /registry/) exactly once; with onServerRequest
the request is forwarded carrying the profile and fail is never called.

* refactor(desktop): a missing server-request registry lets the channel answer -32601

dispatchServerRequest hand-rolled request.fail(JSON_RPC_METHOD_NOT_FOUND,
'Hermes Desktop has no server-request registry yet'), but the channel already
owns that reply: JsonRpcRequestChannel.deliverRequest answers -32601 and fires
onUnhandledRequest when a ServerRequestHandler returns false, and
GatewayBootOptions.handleServerRequest already documents "false = no handler
(the channel answers -32601)".

Make dispatchServerRequest return false when the registry has no
onServerRequest and true after forwarding; dispatchPrimaryServerRequest and
both gateway.onRequest registrations (store/gateway.ts secondary sockets,
use-gateway-boot.ts primary) now propagate that value to the channel. Keep
the desktop-specific wording by wiring onUnhandledRequest on HermesGateway
beside the onRequestHandlerError sink (console.warn), which needs the same
GatewayClientOptions passthrough onRequestHandlerError got. Drop the now
unused JSON_RPC_METHOD_NOT_FOUND import from the store and export
JSON_RPC_INTERNAL_ERROR from the shared barrel beside it.

Test: the store test asserts the false return with no fail() call when the
registry is missing, and true + forwarded profile when it is present.

Follow-ups (same class, outside this stack): use-gateway-boot.ts ~L889-891
still hand-rolls -32601 when the registry is present but has no handler;
ui-tui has its own copy.

* fix(gateway): resolve session_key in shutdown-flush recovery

recover_pending_to_db skipped every real flush file: _serialise_value
captures only the text field from adapter MessageEvent objects (they
have no session_id attribute), so recovery always hit the 'no
session_id' skip branch — messages queued during a gateway drain were
written to disk and then silently never re-ingested, despite the
user-facing 'queued for the next turn' promise. The existing tests
masked it by hand-writing session_id into payloads real events never
produce. Add an optional session_resolver parameter and wire it to
SessionStore.peek_session_id at the startup call site; files remain
preserved (with the warning) when a key genuinely can't resolve.

* fix(gateway): resolve flush session_key to session_id at recovery

* fix(gateway): keep one failed flush file from aborting pending-message recovery

`recover_pending_to_db` walks the shutdown flush spool and appends each
pending message back into state.db. Its per-file `try` listed
`except BaseException:` above `except Exception as exc:`. Exception is a
BaseException subclass, so the ordinary-error handler was unreachable and
every ordinary failure took the interrupt path: close the owned DB and
re-raise.

The effect is that a single unrecoverable spool file aborts the entire
recovery pass. Because the file is only unlinked after a successful
append, it survives and re-poisons the next boot, and the pass is walked
in `sorted(glob("*.json"))` order over uuid4-named files, so a different
subset of the user's pending messages is stranded each time. The caller in
`gateway/run.py` wraps the call in `except Exception: pass`, so nothing is
logged — the messages simply never come back.

Reordering the two handlers restores the documented per-file behaviour
("Leave the file for next startup retry") while keeping the interrupt
contract exactly as written: a KeyboardInterrupt, SystemExit or
CancelledError still closes an owned SessionDB and propagates.

* test(gateway): cover recovery skipping a failing payload and continuing

Pins that an ordinary append failure on one spool file only skips that
file: the remaining files are still recovered, the good file is unlinked,
and the failing one is preserved for the next startup retry.

Fails before the handler reorder — the RuntimeError propagates out of
recover_pending_to_db and the second message is never recovered.

* refactor(gateway): shutdown-flush resolver — routing map first, exact-key row fallback fenced by flush time

Follow-up to the #75536 + #106112 picks; changes vs #106112:

- Routing map first: peek_session_id (sessions.json) is authoritative for
  the session a message was routed to at shutdown. The durable-row lookup
  (find_latest_gateway_session_for_peer under the exact key) is only the
  fallback for a pruned map.
- Fence added per gysyl's #106112 review: a row whose started_at is later
  than the flush payload's ts (passed as not_after by _recover_one_payload)
  cannot be the message's origin and is never adopted; None is returned so
  the flush file is preserved instead of appending into a newer session.
- WhatsApp phone<->LID alias expansion dropped: main's session_recovery
  does no alias expansion anywhere else, and re-implementing it here made
  the resolver diverge from the recovery path it sits beside. Exact-key
  match only.
- The private _db_for_key reach stays inside SessionRecoveryMixin so the
  (session_id, db) pair lands the append in the owning profile partition;
  shutdown_flush never touches store internals.

* test(gateway): trim flush-resolver tests to positive + fence invariants

Keep test_recover_payload_without_session_id_uses_resolver_and_deletes_file
(positive path) and add one negative fence test on the resolver: a row
started after the flush ts is rejected, an older one is adopted with its
owning db. Drop the WhatsApp alias tests (feature removed), the duplicate
positive test and the redundant no-resolver/None-resolver file-preserved
tests, which were exercising the same branch three ways.

* refactor(gateway): one peer-finder helper for the two recovery lookups

_find_gateway_session_row and resolve_session_id_for_key each carried the
same guarded call: getattr the finder off a possibly-None db, check
callable, try/except -> debug log -> None. Extract it as the mixin-private
static `_peer_row(db, *, source, session_key, raise_on_lookup_error=False,
**peer)` and call it from both sites.

Behaviour is unchanged at both call sites: the tuple site passes
user_id/chat_id/chat_type/thread_id through **peer with the same
allow_peer_fallback gating and keeps raise_on_lookup_error; the exact-key
site passes only source + session_key, so the finder's tuple fallback
still never runs there. The only observable difference is the resolver's
failure debug line, which now reads "Gateway session DB recovery failed"
(shared) instead of "Session key->id resolution failed"; no test asserts
on either.

The platform-slot parse (parts[2]) in resolve_session_id_for_key is left
in place: the only key-parsing helper on the mixin,
_profile_from_session_key, returns the profile namespace (parts[1]) and
nothing returns the platform slot, so there is nothing to fold onto.

* docs(gateway): say why the exact-key resolver skips the scope fences

resolve_session_id_for_key calls the peer finder directly, while the
sibling _query_recoverable_row also applies _recovered_row_matches_source_scope
and _recovered_row_allowed_for_active_profile. Read both fences against the
resolver's call shape; both are redundant there, so this adds a WHY comment
at the call site rather than a fence call.

- Profile fence: it returns True immediately when recovered.session_key ==
  requested key. The resolver passes no chat_id/chat_type, so
  find_latest_gateway_session_for_peer returns after the _PEER_BY_KEY_SQL
  branch (`s.session_key = ?`) or None; the tuple fallback is unreachable.
  Every hit therefore carries the requested key and the fence is a no-op.
- Slack workspace fence: it compares origin_json.scope_id with
  source.scope_id. The resolver has no SessionSource (only the key), and a
  scoped Slack key embeds the scope_id as its own slot
  (`<ns>:slack:<chat_type>:<scope_id>:...`, build_session_key). An exact
  key match therefore already pins the workspace, regardless of whether
  the row lives in a profile store or a non-multiplexed root store — the
  store choice (_db_for_key) partitions by profile, not by workspace, and
  the key does the workspace work in either store. The sibling needs the
  fence only because it also looks up the UNscoped legacy key, whose row
  can belong to any workspace; the resolver never does that lookup.

* fix(gateway): an unresolvable profile store preserves the flush file instead of writing to the root store

`_db_for_key` deliberately returns None for a named-profile key whose home cannot be
resolved (fail closed, #66887/#102157), but `resolve_session_id_for_key` returned
`(session_id, None)` on a routing-map hit, and `_recover_one_payload` treated that None
as "fall back to session_db" — the owned ROOT state.db at the production call site. A
profile message with an unopenable home was appended to the root store: the very
split-identity write the fail-closed return exists to prevent.

- resolver: an unresolvable store answers None before the peek and the row fallback
  (None already means "preserve the flush file").
- recovery: the resolver's db is authoritative for resolver-resolved payloads; the owned
  default store only serves payloads that already carry a session_id (and cap-drop spool).
- `started_at` is `REAL NOT NULL` in the schema, so the `float()` fence stays as is.

* fix(gateway): the flush-time fence compares whole seconds

The flush payload's ``ts`` is ``int(time.time())`` while ``sessions.started_at`` is a REAL, so
``started_at > not_after`` rejected a row minted later in the same second as the flush — the
common "message arrives, session minted, SIGTERM" shape lost recoverability (file preserved, not
misrouted). Compare ``floor(started_at)`` against the whole-second ``ts`` so same-second rows
are adopted and only rows from a later second are refused.

Spotted by the round-2 gate on the salvage stack.

* fix(config): allow unseeded runtime config keys

* test(config): pin the same-section-typo trade-off of the unseeded-key fix

Fold the trade-off into the salvaged parametrized test instead of adding a
new test function: `agent.max_turnz` (suggestion `agent.max_turns`) is now
written with the post-write notice rather than refused, because the schema
walk cannot tell a same-section typo from a deliberately unseeded
runtime-read key such as `stt.provider`. Assert the notice and the
suggestion so reviewers see the behaviour change; refresh the class
docstring that still described the pre-#114107 refusal rule.

* docs(config): config set docs describe the wrong-prefix-only refusal

The CLI reference, the configuration guide and the set_config_value
docstring still said that any unknown path under a known section is
refused with a did-you-mean. After b47f40e699 only a path whose suffix is
itself a known key (gateway.discord.foo -> discord.foo) is refused; every
other unknown path under a known section — a same-section typo or an
unseeded runtime-read key — is written with a did-you-mean notice, because
DEFAULT_CONFIG is an incomplete schema and cannot tell the two apart.

Reword all three sites to state the refusal that actually exists and the
write-with-notice fallback, so users are not told a typo will be blocked.

* test(config): the unseeded-key rows state their real did-you-mean

The display.tool_progress row passed suggestion=None, but
_validate_config_key("display.tool_progress") returns
(False, "display.tool_progress_command") because only the _command sibling
is seeded in DEFAULT_CONFIG. The `if suggestion:` guard let the row pass
without ever checking the notice, so the misleading "Did you mean" this
real runtime key (cli.py:2568) receives was neither asserted nor visible.

Pin the actual suggestion for that row and assert the notice text exactly
for every row: present with the expected sibling, or absent. Fix the
comment too — only stt.provider comes from da942e4483's list; the rows are
unseeded runtime-read keys at agent/agent_init.py:1324, cli.py:2568 and
tools/transcription_tools.py:241.

Seeding display.tool_progress in config_defaults.py (hermes_cli/AGENTS.md:
"every reader a registry entry") is the proper follow-up that removes the
misleading suggestion; it is out of scope for this salvage.

* test(config): unseeded-key rows pin the invariant, not a known-defective suggestion

Drop the display.tool_progress row (it pinned a did-you-mean the comment itself called misleading; seeding the key would fail it for an unrelated reason) and the skills.creation_nudge_interval row (same invariant as stt.provider, the key named in the issue).

* fix(config): a structural wrong-prefix match beats a fuzzy sibling suggestion

In _validate_config_key the fuzzy same-level sibling (cutoff 0.6) was tried before the
structural "path minus its wrong prefix is a known key" check. While every unknown sub-key was
refused that order did not matter; now that only the wrong-prefix case is refused, a path whose
stray middle segment fuzzy-matched a sibling was WRITTEN with a misleading did-you-mean:
`agent.gateway.strict` -> (False, 'agent.gateway_timeout') instead of being refused as
`gateway.strict`. A structural match is proof, a fuzzy match is a guess; check it first.

* fix(config): the custom-top-level footer prints only for top-level keys

_print_unknown_key_notice always appended "Custom top-level keys are supported and bridged to
the environment". Before the unseeded-key change no nested path reached this notice; now
`stt.provider` or `agent.max_turnz` did and were told they are env-bridged, which they are not.
Nested paths get only the --force hint.

* fix: goal loop names the Nous auth failure instead of an opaque judge error

With auxiliary.goal_judge.provider pinned to nous and a dead refresh token
(invalid_grant), the status line read "judge error: RuntimeError": the ladder's
_resolve_nous_runtime_api swallowed the resolver's AuthError at DEBUG and the
generic "no API key" RuntimeError was reduced to its type name by judge_goal.
Users were sent to context-length / model debugging instead of `hermes model`.

Now the failure is remembered in agent/auxiliary_unavailable.py (one WARNING per
distinct message, preferring the persisted quarantine marker because the pool
rung wipes the dead tokens before the auth-store resolver run…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint comp/plugins Plugin system and bundled plugins P3 Low — cosmetic, nice to have type/bug Something isn't working

Projects

None yet

3 participants