fix: keep watcher arms and reply listeners alive through slow cycles - #6103
Merged
Merged
Conversation
…aking supervision - fm_pending_reply_tick selects the records it has work for in one awk pass, so settled records cost no lock or fork and the walk no longer grows with the never-pruned store. - An attached arm keeps following a live, identity-matched holder whose beacon went stale until the lock changes or the shared stall bound (fm_watcher_stall_bound), then fails with a typed stalled-holder line so the retry replaces the holder. - The remote-reply adapter reports the job worker's preemption (exit 76) as a closed window, so the listener keeps its claim and polls again instead of being relaunched every watcher cycle.
|
This was referenced Sep 30, 2026
andrewesweet
pushed a commit
to andrewesweet/firstmate
that referenced
this pull request
Sep 30, 2026
…unchenguid#6103) * fix(bin): keep slow watcher cycles and preempted reply polls from breaking supervision - fm_pending_reply_tick selects the records it has work for in one awk pass, so settled records cost no lock or fork and the walk no longer grows with the never-pruned store. - An attached arm keeps following a live, identity-matched holder whose beacon went stale until the lock changes or the shared stall bound (fm_watcher_stall_bound), then fails with a typed stalled-holder line so the retry replaces the holder. - The remote-reply adapter reports the job worker's preemption (exit 76) as a closed window, so the listener keeps its claim and polls again instead of being relaunched every watcher cycle. * no-mistakes(document): Clarify watcher grace and attached-arm documentation
ironerumi
added a commit
to ironerumi/firstmate
that referenced
this pull request
Oct 2, 2026
* fix: grant Claude workers access to Firstmate task channels (#5884)
* fix(bin): grant Claude workers their task-channel dirs via --add-dir
Since Claude Code 2.1.257, a file-tool read (Read/Glob/Grep, and an
Edit's mandatory prior read) of a path outside the working directories
parks --permission-mode auto panes on a one-time interactive question,
and a "Block" answer lands permissions.blockReadsOutsideWorkingDirectories
in user settings, refusing the same reads even under bypass. Firstmate
launches Claude with no --add-dir, so a secondmate's parent-home steering
inbox and a ship or scout worker's launch record, steering inbox, brief
dir, and code-root .agents/skills were all outside: workers wedged on
the question the first time they read a steer.
Every Claude launch, spawn and relaunch, in both permission modes, now
grants exactly the task's channel directories: state/<id>.inbox for a
secondmate (in the parent home), or state/operational-inbox,
state/<id>.inbox, data/<id>, and the code root's .agents/skills for a
ship or scout. Paths resolve to real paths and lazily created channel
dirs are made before launch so the grant never names a not-yet-existing
directory; the whole state/ is deliberately never granted.
The grant keeps the bypass-mode launch argv changed on purpose: it also
protects bypass workers against a machine-recorded Block answer.
* no-mistakes(document): Consolidate Claude launch guidance in configuration reference
* no-mistakes(document): Clarify Claude permission documentation reference
* fix(bin): prevent idle recovery loops without stranding wakes (#4819)
* fix(supervision): prevent idle recovery loops without stranding wakes
* no-mistakes(review): Remove unused wake-append rollback helper
* fix: reduce remote worker and polling helper process churn (#5889)
* fix(bin): stop the remote-job worker busy-polling an idle queue
The serving loop slept 50ms between passes and re-ran state preparation
(chmod on every queue directory), the heartbeat publish, and the stale sweep
on every pass. It now blocks on a worker.wake FIFO that staging,
cancellation, and lane exit nudge, keeps a short fast-poll window after
activity, refreshes the heartbeat at most once a second, and runs the sweep
(which re-applies the queue directories' 0700 modes) at startup and then on
a bounded interval. Lane-owned records are no longer re-read every pass.
Measured with a fork/execve-interposing counter on a --serve worker in a
disposable HOME and queue, bash 3.2, 20-second windows (the counter slows
the old loop to about 5 passes a second, so real-host rates were higher):
idle worker 146 forks/s, 61 execs/s -> 4.8 forks/s, 3.1 execs/s
one running long job 232 forks/s, 100 execs/s -> 15 forks/s, 11 execs/s
Stage-to-result latency for a no-op job, idle and back to back, stayed at
about 0.8-1.2s in both versions (dominated by job execution, not pickup).
* perf(bin): drop per-cycle forks from watcher, drain, and lock helpers
The watcher, drain, inactive-reconcile scan, and branch-outcome reads forked
small external commands on every cycle where bash can do the same work.
- fm-wake-lib.sh gains fm_dirname_to, fm_basename_to, and
fm_epoch_seconds_to, exact stand-ins for $(dirname --), $(basename --),
and $(date +%s); the clock uses printf %(%s)T on bash 4.2+ and still forks
date exactly once on stock macOS bash 3.2.
- fm_lock_abs_path, fm_wake_signal_seen_path, fm_path_age, the watcher's
age_of and wedge timer, and the recovery-marker line count use them or
plain reads instead of dirname/basename/tr/date/wc.
- window_to_task reads a meta file once instead of two
grep | tail -1 | cut -d= -f2- pipelines per file per call.
- fm-classify-lib.sh reads uname -s once at source time instead of in every
status stat helper.
- Libraries sourced every cycle derive their own directory without forking
dirname, including the backend adapter siblings a subshell re-sources on
each probe.
tests/fm-fork-free-helpers.test.sh pins each replacement against the command
it replaces on edge-case inputs, under every available bash and both the C
and a UTF-8 locale; CI's stock macOS bash lane runs it under /bin/bash 3.2.
Measured with a fork/execve-interposing counter in a disposable home, one
tmux crew task, FM_POLL=1 (forks and execs per watcher cycle, per run
otherwise):
watcher cycle bash 5.3 299/138 -> 199/66 bash 3.2 341/146 -> 224/80
drain bash 5.3 492/238 -> 430/200 bash 3.2 567/250 -> 491/212
inactive scan bash 5.3 27/14 -> 17/4 bash 3.2 37/14 -> 17/4
branch-outcome bash 5.3 40/21 -> 35/16 bash 3.2 48/24 -> 38/19
* test: note the interpreter-expanded version probe for shellcheck
* no-mistakes(review): Fix worker idle bounds and fork-free contributions snapshot
* no-mistakes(review): Coalesce buffered worker wake nudges into one wake
* no-mistakes(review): Coalesce wake nudges via pending marker so publishers never block
* no-mistakes(review): Claim wake nudges atomically via noclobber pending marker
* no-mistakes(review): Release abandoned wake claims only after a 30-second bound
* no-mistakes(review): Drop worker wake FIFO; load path helpers side-effect free
* no-mistakes(document): Document remote worker polling and preemption cadence
* no-mistakes(ci): Fixed both failing CI shards: isolated remote and teardown test fixtures now include fm-path-lib.sh, which fm-wake-lib.sh requires. The three affected tests, fm-lint.sh, and git diff --check pass locally
* fix: keep remote reply listeners and watcher cycles running (#5941)
* fix: keep a successor watcher and remote-reply listeners across the gaps that dropped them
A main-only supervision pass-through exited without leaving a watcher, and each remote-reply poll released its claim until the next cycle, so short-lived listeners stayed down.
* no-mistakes(document): Clarify listener and supervision continuity documentation
* no-mistakes(ci): Fixed the three Greptile findings: failed ingestion leaves one durable capture, failed reads exit instead of relistening, and the disposable-checkout guard rejects state paths outside the marked lab. Added behavioral tests; the remote-reply and watcher-lock suites, shell syntax checks, and git diff checks passed
* no-mistakes(ci): Fixed a race in the wake-queue interruption test: it now waits for the drain to own the lock and enter handling before signaling it. The wake-queue suite, shell syntax check, and diff check pass
* fix: make attended cutover outcome re-presentation check-first (#5925)
* fix: date replayed branch outcomes and ask main to check current state first
A captain outcome main never acknowledged is presented again, which after a
harness or posture switch, or the first drain after the upgrade whose earlier
presenter never advanced the read cursor, can be days after its situation
settled. The replay read as fresh news, so a PR since merged looked ready.
bin/fm-branch-outcome.sh now adds a "recordedAgo" age (minutes, hours, then
days) to present and unprocessed rows, one owner of that wording for both
presenters. The drain's BRANCH OUTCOMES captain lines and the Pi branch's
processing request name that age and ask main to check the task's current
state first; an outcome already settled needs only the acknowledgement, with
nothing relayed to the captain. Nothing is adopted as processed, so a fresh
home's first outcome is still presented until acknowledged.
* no-mistakes(review): Absent processed marker reads 0; never adopt read cursor
* no-mistakes(review): Require recordedAgo in Pi requests; report undated rows to main
* no-mistakes(review): Keep recordedAgo on captain rows only in present output
* no-mistakes(document): Correct cutover documentation and retire stale migration guidance
* fix: keep settled branch outcomes out of main's reply to the captain
A live Pi primary that took over a host-drain home received the carried-over
outcomes dated and check-first, but its processing reply still told the
captain about an outcome whose decision had since been answered. The request
also claimed every outcome was already shown as an anchor entry in this
transcript, which is false for an outcome carried over from before a restart
or a switch of primary.
The Pi processing request now says each outcome was recorded earlier and may
already have been seen or handled, and that a settled outcome gets no
captain-facing mention at all in the reply or any recap, not even that it is
settled. The drain's BRANCH OUTCOMES header and the supervision docs state the
same rule, and the tests check both delivered texts.
* fix: scope main's outcome reply to what is still open
Telling main what not to say about a settled outcome was not enough: in two
live Pi trials the processing reply still told the captain that an answered
decision was settled. Main now sorts the outcomes by current state first, and
its reply to the captain covers only the still-open ones, written as if the
settled ones had never been listed. With that framing three live Pi trials
kept the settled outcome out of the reply and relayed the open one each time.
The drain's BRANCH OUTCOMES header and the supervision docs use the same
framing, and the tests check both delivered texts.
* no-mistakes(review): Clarify that main acknowledges every presented captain outcome
* no-mistakes(document): Clarify outcome cursor ownership across Pi and host
* no-mistakes(ci): Fixed Pi replay by batching unprocessed captain outcomes oldest-first and acknowledging only through each batch. Verified a marker-less backlog over 1 MiB replays through all batches. A real-drain regression confirms an older keyed decision remains under OPEN DECISIONS after a newer branch row is acknowledged; the check-first instruction now names those decisions. Relevant targeted tests and branch-supervision tests passed; the full host suite timed out
* no-mistakes(ci): Fixed the host drain’s check-first wording in bin/fm-wake-drain.sh; the CI fixture now passes. The full host suite passed the affected fixtures but timed out later. Syntax and diff checks passed
* no-mistakes(ci): Fixed ci-1: abbreviated Pi outcome summaries now stay within 1,024 characters and include a row-specific full-outcome lookup command. The delivered instruction requires reading the full outcome before acting, relaying, or acknowledging it. The new extension-driver regression failed before the fix and passes now; the Pi and supervision-host suites pass
* no-mistakes(ci): Corrected the batching sentence in docs/pi-supervision-branch.md. The cancelled CI check needs no code fix; its clean rerun passed. The Pi branch extension suite and git diff check passed
* fix: restore primary rewakes after attended main-only closes (#5961)
* fix: wake an idle Claude primary for attended main-only hand-backs
An attended main-only pass-through confirmed a handling handoff for the
successor it leaves running, which flipped the recovery marker to handling.
The Claude Stop hook only rewakes main while that marker reads downtime, so
the close reached no one and an idle primary slept with wakes queued.
The pass-through now leaves the marker at downtime, and a close that turns
main-only at its turn hands the consumed handoff back to downtime before it
reaches main. Regression tests drive the real Stop hook around the real host
on both paths and for the successor's own later close, and a new opt-in live
guard proves it against an idle interactive Claude primary with a pre-fix
negative control.
* no-mistakes(review): Assert live lab Stop-hook registration via parsed settings JSON
* no-mistakes(document): Correct supervision hand-back documentation
* no-mistakes(ci): Fixed the failed downtime-write path so the Stop hook notifies main instead of silently dropping the close. Corrected the live guard’s tracked-hook check and added the requested at-turn main-only scenario. The new regression failed before the fix and passed after it; the host suite, syntax checks, and diff check pass. The credentialed live guard was not run in this CI phase because it writes outside the worktree
* no-mistakes(ci): Fixed the Stop hook’s retry ordering: a crashed host gets its bounded retry before a non-crash hand-back failure is reported. The Stop-hook suite passes, including the crash regression. The host suite passed the failed-marker-write regression but timed out before completing; syntax and diff checks pass
* Clarify live Claude login for opted-in tests (#5975)
* fix(bin): ensure resumed worker launches enter their recorded worktree (#5916)
* Fix worker launches to enter recorded worktrees
* no-mistakes(review): placeholder
* no-mistakes(document): Update agent-control.md worktree-refusal note to match new universal cd+assert
* no-mistakes(review): Add regression tests for Orca spawn/relaunch worktree carve-outs
* Fix PR relaunch and prelaunch cwd verification
* no-mistakes(document): Fix docs/agent-control.md: worktree cwd check is pre-launch, not post-launch
* fix: slim worktree launch change onto upstream main
* fix(bin): retire task-keyed watcher markers and orphan journals at teardown (#5997)
* WIP: retire task-keyed watcher markers and orphan journals at teardown
Re-applies old PR #5584 on current main: teardown retires the
turn-ended .seen-* signature and an orphaned Herdr presentation
journal whose workspace is already gone, and the wake-drain rotates
its own dead scratch files. Not yet validated through no-mistakes.
* no-mistakes(ci): Fixed the Greptile P1 finding in bin/backends/herdr.sh. fm_backend_herdr_projection_token_workspace_gone used `! ... jq -e ... 2>&1`, which swallowed a jq runtime error (thrown when a non-object workspace entry, e.g. a number before a live token-bearing workspace, hits `.label`) and flipped it to a "gone" verdict, causing teardown to delete a still-live v1 presentation journal. Invariant: a workspace-query error/ambiguity must never be read as token absence; only a cleanly-parsed list with no token-bearing label is "gone". Replaced the body with a single jq verdict (unknown/present/gone): a non-array list or any non-object/non-string-label entry yields "unknown", jq errors/empty output fall through `|| return 1` to unknown, and only "gone" returns 0. Sibling fm_backend_herdr_projection_endpoint_matches_journal already fails safe on jq error (empty match -> journal kept), so it needed no change, matching the author's scoping. Added test_teardown_retains_v1_journal_when_workspace_query_ambiguous driving real teardown with a malformed workspace-list entry, proving the journal is kept and no workspace close occurs. Verified the old logic returns GONE on that input (test fails before, passes after); full tests/fm-teardown.test.sh suite passes (exit 0) and shellcheck is clean. Marker-naming finding left untouched per explicit out-of-scope instruction
* test: guard the live harness gate against Claude Code's auto-updater (#6002)
* fix(tests): disable Claude Code's auto-updater during live harness runs
fm_live_gate let a live run proceed without ever setting
DISABLE_AUTOUPDATER, so a live Claude test could let the real updater
repoint ~/.local/bin/claude into a temporary directory and stop every
Claude process on the machine from starting. Export
DISABLE_AUTOUPDATER=1 on every path where the gate lets a live run
proceed, and assert the export in tests/fm-live-gate.test.sh, including
that it reaches a child process the same way a real harness pane would
inherit it.
* no-mistakes(ci): Greptile flagged that the PR's DISABLE_AUTOUPDATER inheritance test only checked a `bash -c` direct child, not the fm-spawn.sh launch path. The user chose to fix it with a regression on that path. In tests/fm-live-gate.test.sh I replaced the generic child test with test_disable_autoupdater_reaches_the_claude_pane_on_the_fm_spawn_launch_path: it drives the real fm-spawn claude launch through the spawn fixtures, captures the exact staged launch command, and runs it as a synthetic pane whose only `claude` is a stub recording the inherited DISABLE_AUTOUPDATER, asserting it saw 1. Switched the file to source fixtures.sh (pulls in lib.sh, guarded) for the spawn helpers. Verified it is a real guard: the stub records `1` when the ambient var is set and `unset` when absent, so it fails if fm-spawn ever scrubbed the variable (e.g. env -i or -u). This confirms fm-spawn's launch construction never references the name and passes it through via ordinary ambient inheritance with no allowlist. Full suite passes (12 tests ok), shellcheck clean. Note for the outer executor: I embedded the daemon caveat as a code comment in the test, but the finding also asks the PR body to state that a backend daemon already running before the gate exported the variable does not inherit it and fully covering that would need launcher support - that forge-side PR-body sentence is outside this CI phase's scope
* no-mistakes(ci): Fixed Greptile finding ci-1. Root cause: fm-spawn.sh handed its launch command to an already-running backend daemon that never inherited the test process's exported DISABLE_AUTOUPDATER, so ambient inheritance dropped it and Claude's auto-updater could still run. Fix (bin/fm-spawn.sh): when DISABLE_AUTOUPDATER is set in the spawn's own environment, embed `export DISABLE_AUTOUPDATER=<value>;` into the LAUNCH command text (same idiom as the adjacent COMPACT_ADVISER_DISABLE export), so it survives a daemon-built pane, the env -i allowlist path, and relaunch alike; gated on presence so ordinary spawns are unchanged. Added regression test test_disable_autoupdater_survives_a_daemon_pane_that_never_inherited_it in tests/fm-live-gate.test.sh: stages a real claude launch with DISABLE_AUTOUPDATER set in the spawn env, then runs that exact command in a synthetic pane with `env -u DISABLE_AUTOUPDATER` and asserts the claude stub still recorded autoupdater=1. Verified the test fails (autoupdater=unset) without the fix and passes with it; the round-1 ambient test stays green either way. Full suite passes (13 ok); test file and isolated snippet shellcheck-clean (full fm-spawn.sh shellcheck kept getting terminated by the memory-constrained host, not by findings). Forge-side note for the outer executor: the PR-body caveat that fully covering the daemon case would need launcher support no longer applies to the Claude launch path and should be corrected
* fix(bin): ring the worker inbox doorbell only for a newly written procevent record (#6010)
* fix(bin): ring the inbox doorbell only for a newly published procevent result
publish_result rewrote a worker's captured Lavish round idempotently on
every reconcile, unconditionally moved an already-acknowledged inbox
record back out of handled/, and rang the doorbell every time - so an
already-processed round rang the owning worker on every cycle. Snapshot
the existing active and handled records before the idempotent write and
ring, or move anything, only when the write actually created a fresh
record; re-delivery of a still-open round is left to the inbox's own
re-ring ladder.
* no-mistakes(document): docs: reflect worker-board doorbell rings only on fresh inbox record
* no-mistakes(ci): Fixed Greptile finding ci-1 in tests/fm-procevent.test.sh (test-only change). The redelivery regression previously moved the delivered note into handled/ before any repeated reconciles, so it only proved an acknowledged note stays quiet and would still pass if an unchanged active note rang every cycle. Per the user's instruction, I inserted (before the mv into handled/) five repeated `pe reconcile` runs with the note still in the active inbox and asserted the ring log holds exactly one line and 001.msg remains active; the existing acknowledged-note assertion after the move is kept unchanged. No product code changed. bash -n confirms syntax is valid; the block mirrors the already-passing post-move reconcile/ring-count assertion directly below it
* fix: prevent manual Claude Stop hook calls from arming supervision (#6032)
* fix(bin): make the Claude Stop auto-arm refuse arguments before arming
A model running bin/fm-claude-stop-autoarm.sh --help mid-turn armed a real
supervision-host park owned by its short-lived tool process, leaving
supervision down once that process exited. The Stop hook passes no
arguments, so -h/--help now prints usage and any other argument is refused
before anything is sourced, read, or armed.
* no-mistakes(document): Clarify Claude Stop hook documentation for manual invocations
* no-mistakes(ci): Updated the argument-run regression test to compare checksums of state files as well as entry names. The Stop auto-arm test suite passes, and git diff --check is clean
* docs: restore the bin/ toolbelt intro's manual-use clause
The document step dropped "interactive entrypoints work by hand too" from
docs/scripts.md, which still holds for most bin/ scripts.
* no-mistakes(document): Clarify Claude auto-arm manual-use guidance
* feat(firstmate-calm): show supervision notes in Claude Code (#6039)
* feat(calm): show supervision sailboat and anchor notes on Claude Code
The Calm mod follows a bounded display tail copy of the outcome store,
which bin/fm-branch-outcome.sh append now refreshes, and the supervision
host's latch, and appends one dim transcript line per visible routine
outcome, captain outcome, and latch change, replaying unread and
unprocessed outcomes at session start. It shows them whenever the mod is
active, regardless of config/calm, and never marks anything read.
* fix(calm): show each supervision note once per session on Claude Code
Claude Code 2.1.283 stores ui.log lines in the session and restores them
on --continue, so the mod records how far each session has followed the
outcome store and a resume replays only newer outcomes. It also checks
file existence before reads so absent files do not log debug errors.
The live guard gains the supervision-notes scenario and the dated
2.1.283 record documents the observed behavior.
* docs: name the Claude supervision note row as the engine draws it
* no-mistakes(review): Seed outcome tail on present and anchor first tail on markers
* no-mistakes(review): Seed outcome tail at session start; replay against start markers
* no-mistakes(review): Bound outcome tail by bytes; reread recently changed files
* no-mistakes(review): Skip store validation when outcome tail already exists
* no-mistakes(document): Clarify bounded Claude supervision note replay
* no-mistakes(ci): Fixed seed-tail to validate only a bounded suffix of complete store rows and write it through the existing byte- and row-limited tail writer. Added a regression test with malformed history outside that window and updated the script header. Outcome tests and shellcheck passed; the full session-start suite timed out after 240 seconds
* fix: stop quiet mode from holding requested actions for return (#6033)
* fix(bin): read a quiet-mode record as a present captain, never hold-for-return
Daemon-backed quiet mode writes the away-posture record marked mode: quiet,
but the entry announcement, read-back, and session-start digest rendered it
as "hold-for-return only", and the spend cap and PR merge gate treated it as
away. A present captain's requested actions could then be held for a return
that was not coming.
bin/fm-afk-contract.sh now owns which posture a record is (the mode
subcommand, fm_afk_contract_mode, fm_afk_contract_away_present). A quiet
record announces, reads back, and appears in the digest as a present captain
holding nothing; merges under it stay attended and it binds no spend cap. An
away record is unchanged, an /afk entry over quiet mode rewrites the record
as away, and a quiet entry never turns a standing away record quiet.
* no-mistakes(document): Clarify quiet-mode authority and remove stale away guidance
* no-mistakes(ci): The CI failure came from a race in the supervision-host test: its restart fixture could observe a watcher left by the preceding cycle. The test now retires that watcher and waits for the fixture arm to report its own started cycle. The focused test passed three times; the full suite was attempted but stopped at a separate intermittent test failure
* no-mistakes(ci): Fixed daemon refresh mode selection so an unset-mode refresh follows the posture record: /afk over a running quiet daemon now changes state/.afk to away, while a plain quiet refresh stays quiet. Added script-level regression coverage for start and start-native and corrected a quiet-refresh fixture. The launch test suite, syntax checks, and diff check passed
* no-mistakes(ci): Herdr was blocked before tests ran by a GitHub HTTP 500 downloading pinned Treehouse; no code change was warranted for that check. Fixed the Lint 1 ShellCheck warning in tests/fm-afk-launch.test.sh by annotating the intentional background PID capture. The focused test suite, ShellCheck, syntax check, and diff check passed
* Say ahoy impact order is the first mate's pick. (#6065)
* fix(bin): accept a task's next PR after its bound PR merges in fm-pr-merge (#6053)
* fix(bin): accept a task's next PR once fm-pr-merge confirms the bound one merged
require_recorded_pr_identity now checks fm_pr_poll_merge_already_notified for
the recorded pr= before refusing a different URL, so a task's later PR is
accepted once its earlier PR's merge is confirmed, while it keeps refusing
while the bound PR is still unmerged.
* no-mistakes(document): docs(fm-pr-merge): note next-PR accepted after bound PR merges
* fix: treat quiet records as attended across supervision (#6064)
* fix(bin): read a live quiet record as a present captain at the host and watcher
A quiet record left without its daemon (a quiet start that never ran or was
interrupted) was read as away by the supervision host, so it parked a present
captain's main and held captain outcomes for a return that never comes, and
the watcher and daemon silenced captain-held rechecks on record presence.
The host's posture checks, the watcher's and daemon's captain-held silencing,
and the host's outcome path (branch report, drain BRANCH OUTCOMES, relocated
branch authority, the owners' away wake note, and the Codex checkpoint bound)
now ask the record owner's away-or-quiet reading, so only an away record is
away. A live away record keeps today's behavior.
* no-mistakes(document): Correct quiet-record documentation and supervision guidance
* no-mistakes(document): Clarify quiet-record posture and captain-held rechecks
* no-mistakes(document): Clarify quiet-record posture in documentation
* fix: report supervision host latches accurately in away return briefs (#6043)
* fix(bin): name an in-window engine latch in the return brief and drop the false handling GAP line
The away return brief said nothing had failed after the supervision host
latched on engine errors during the window, and printed a GAP: watcher
downtime line whenever a wake was merely being handled or queued at return.
The failures section now reads the host ledger and latch record and names
the latch time, the window's engine-error count, and whether the session
is still paused or recovered. An open recovery episode is reported as
information, and as a gap only when a queued episode outlived the return
grace or the marker cannot be read.
* no-mistakes(review): Fix latch trip time, drop marker-age grace, bound error count
* no-mistakes(review): Report paused latch without ledger trip row; bound errors
* no-mistakes(review): Never report a failed probe's latch row as trip time
* no-mistakes(review): Only a retained trip row marks a pre-window latch
* no-mistakes(document): Clarify return-brief latch and watcher-gap documentation
* no-mistakes(ci): Fixed Lint 1 by marking the shared cooldown constant as used by sourcing scripts. The repository lint command and diff check pass; the return test run was stopped by a 180-second timeout after its completed cases passed
* no-mistakes(ci): Fixed the return brief so the trip time and error count come from the same initial latch row, and ledger rows before the current session’s lock boundary cannot affect its latch report. Added real-script regressions for both findings. The return test suite, repository lint, and diff check pass
* no-mistakes(ci): Fixed the return brief’s restart cutoff so it retains in-window failures, prints one line per initial-trip row, and omits zero-error count wording. Added real-script restart regressions. The return test suite, ShellCheck, and diff check pass
* no-mistakes(ci): Fixed the return brief so a recorded trip followed by recovery stays recovered, while a later pause with a lost trip append gets a separate “trip time unavailable” line. Added a real-script regression that failed before the fix. The return test suite, ShellCheck, syntax checks, and diff check pass
* no-mistakes(ci): Fixed the false second latch during recovery. A real-script regression failed before the fix and passes now; the lost-second-trip test still passes. The return test suite, ShellCheck, syntax checks, and diff check pass
* fix: shorten Claude Code Calm supervision note label (#6086)
* fix(calm): name the Claude Code Calm plugin fm so supervision notes read "fm: "
Claude Code labels every mod transcript line with the plugin name, so the
notes rendered as "firstmate-calm: ⚓ ...". Rename the plugin to fm, update
the live guard to assert the fm: label, and document the one-time replay for
sessions resumed across the rename.
* no-mistakes(document): Clarify Calm plugin rename in documentation
* feat(bin): add a disposable live supervision lab builder (#6037)
* feat(bin): add fm-live-lab.sh, a one-command live supervision lab builder
* fix(bin): exact lab windows, per-lab task ids, self-safe teardown
* fix(bin): target lab windows by id, stop lab descendants, add readiness tests
* fix(bin): keep Claude's auto-updater off in live labs; list fm-live-lab.sh
* fix(bin): start the lab tmux server without user config
* no-mistakes(review): Scope lab teardown to its store, root, and task ids
* no-mistakes(review): Record selected user stores at up for check and down
* no-mistakes(document): Clarify live lab documentation and remove stale narratives
* no-mistakes(ci): Fixed the CI failure by checking for an existing lab root before looking up the harness executable. The affected behavioral test and shell syntax check pass; the refusal also works with Claude absent from PATH
* no-mistakes(ci): Fixed all four Greptile findings: teardown signals only recorded lab processes and their descendants; the worker gate is in its granted task directory and its path is exposed; readiness uses current crew state; and mate and worker IDs use 12 nonce hex digits. The CLI behavior tests pass, as do shell syntax, ShellCheck, and diff checks. The Claude no-host path is unchanged
* no-mistakes(ci): Fixed the CI test’s dependence on an installed Claude binary by supplying a test-local stub. The full fm-live-lab test, shell syntax check, and diff check pass
* no-mistakes(ci): Fixed all three selected findings in bin/fm-live-lab.sh: down waits for recorded processes and escalates before cleanup, PID roots are checked against recorded start times, and Claude primary trust is rechecked after mate/worker readiness. Added behavioral tests in tests/fm-live-lab.test.sh. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass
* no-mistakes(ci): Fixed the pre-primary settle wait, worker gate instructions, unused retry variable, and teardown PID revalidation in bin/fm-live-lab.sh. Added behavioral tests in tests/fm-live-lab.test.sh. Both requested commands pass: tests/fm-live-lab.test.sh and bin/fm-lint.sh
* no-mistakes(ci): Fixed teardown to track pre-kill lab processes by PID and start time, including children orphaned when a root exits. Up now rejects an empty pane PID before calling ps. Added regression tests and a Linux-safe worker fixture. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass
* no-mistakes(ci): Fixed teardown tracking for children spawned during shutdown and made the worker fixture verify its exact window with a Linux-available shell. Both requested checks pass. The lab test takes about 66 seconds locally, so the under-one-minute target remains unmet
* no-mistakes(ci): Fixed ci-2 and ci-4 in bin/fm-live-lab.sh and tests/fm-live-lab.test.sh. Teardown now tracks identity-checked members of captured lab process groups, including children orphaned during shutdown, without signaling the caller’s group or unrelated processes. Lint passed, and the lab test passed four times
* no-mistakes(ci): Fixed teardown so an observed-empty process group is permanently dropped, preventing a reused group ID from signalling unrelated work. Added a ps-shim regression test. The lab test, lint, and diff checks pass
* no-mistakes(ci): Fixed ci-1 in bin/fm-live-lab.sh and tests/fm-live-lab.test.sh. The TERM-born-child fixture now waits until its handler is installed before calling down. Down sends SIGKILL to identity-valid survivors on every pass from pass 20 onward and includes survivor process details if it must refuse cleanup. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass locally; Linux CI remains to be verified
* no-mistakes(ci): Fixed down’s teardown wait to require two empty identity-checked scans separated by 0.5 seconds, and removed the unused test loop variable without changing the TERM-born-child test. The lab test, lint, and diff check pass locally
* fix: keep watcher arms and reply listeners alive through slow cycles (#6103)
* fix(bin): keep slow watcher cycles and preempted reply polls from breaking supervision
- fm_pending_reply_tick selects the records it has work for in one awk pass,
so settled records cost no lock or fork and the walk no longer grows with
the never-pruned store.
- An attached arm keeps following a live, identity-matched holder whose beacon
went stale until the lock changes or the shared stall bound
(fm_watcher_stall_bound), then fails with a typed stalled-holder line so the
retry replaces the holder.
- The remote-reply adapter reports the job worker's preemption (exit 76) as a
closed window, so the listener keeps its claim and polls again instead of
being relaunched every watcher cycle.
* no-mistakes(document): Clarify watcher grace and attached-arm documentation
* fix(bin): retry fm-pr-merge a bounded number of times when GitHub mergeable is UNKNOWN (#6110)
* fix(bin): retry a bounded number of times when GitHub mergeable is UNKNOWN
Fixes #6020
bin/fm-pr-merge.sh refused a GitHub merge whenever the pull request's
mergeable field was not literally MERGEABLE. GitHub reports UNKNOWN for
a short while after a push or a base-branch change while it recomputes
mergeability, so a green, conflict-free pull request was refused as if
it could not be merged.
github_verify_mergeable now returns a distinct status when mergeable is
the only failing condition and reads UNKNOWN. The caller retries up to
5 times, 3 seconds apart (overridable in tests), re-reading and
re-checking every live condition on each attempt. Once the bound is
spent it reports mergeability as still being computed rather than
unmergeable, with the same nonzero exit as before. Every other refusal
(closed, draft, conflicting, red or missing checks, away authority,
queue protection) is unchanged and never retried.
* no-mistakes(ci): I fixed both review findings the way you asked. The full suite (`bash tests/fm-pr-merge.test.sh`) ran to completion. Its last lines showed all `ok`, and any failure would have stopped the run early. I watched the output through `tail`, so I didn't see the new test's own `ok` line directly. **ci-2 (`bin/fm-pr-merge.sh`), retry delay not validated.** What must hold: the retry wait is always a short, valid `sleep` argument, so a bad `FM_PR_GITHUB_MERGEABLE_RETRY_DELAY` can never trip `set -e` or hold the task lock for a long time. The retry loop is the only place that reads this variable. The script now reads the value once before the loop and accepts only whole numbers from 0 to 10. Anything else (empty, `abc`, `-1`, `1.5`, `11`, a huge number, leading spaces) falls back to 3. I ran those values through the check by hand and each came out as expected. The loop now sleeps on that checked value. **ci-1 (`tests/fm-pr-merge.test.sh`), no test for a check changing between UNKNOWN reads.** What must hold: every retry re-checks all live conditions, not just mergeable. The fake `gh pr view` in the test can now take an optional second word on each line of the mergeable sequence, which sets the first check's result. The new test `test_github_mergeable_unknown_retry_rechecks_checks` feeds `UNKNOWN`, then `UNKNOWN FAILURE`. It asserts: - exit code 1 after exactly 2 reads, - the refusal names `check 'ci' is not green`, - the message does not say mergeability is still being computed, - `pr merge` was never called. If a later change made the retry look only at mergeable, the loop would read UNKNOWN 5 times, end with the "still being computed" message, and this test would fail. I didn't run it against a deliberately broken script to confirm that. `bash -n` passes. `shellcheck` reports only the existing info-level notes about files it can't follow. Only `bin/fm-pr-merge.sh` and `tests/fm-pr-merge.test.sh` changed
* fix(bin): converge every open owner onto a known terminal contribution (#6112)
* fix(bin): converge every open owner onto a known terminal contribution
settle_final only cleared a stale error on retry, so an owner whose saved
row still said open kept projecting a merged or closed pull request as
open after another owner's row had already recorded the terminal
observation. Copy the known terminal observation to every owner whose
saved row is not itself terminal, keeping that owner's own pending and
notified state, and clear its error.
* no-mistakes(review): Carry terminal checked_at when converging existing owner rows
* no-mistakes(ci): I fixed Greptile finding ci-2 as you asked, with a change to tests/fm-contributions.test.sh only. The rule it enforces: when a retry converges an owner onto a URL that is already merged or closed, that owner gets the terminal owner's whole observation, not just its state. The same weak check appeared twice in test_interrupted_multi_owner_poll_settles_every_owner, so I fixed both: - **Open owner (line 784):** the check now also requires `.observation == $terminal[0].records[0].observation`. The existing checks for error, checked_at, pending and notified are unchanged. - **Errored owner (just below):** it only checked state and error before. It now reads the terminal owner's file and makes the same full-observation comparison. Adding the comparison alone would not have caught anything. The test fixtures gave both owners identical observations apart from `state`, so copying only the state would still have passed. In both cases I also set the terminal owner's observation head to HEAD_B, so the two observations now really differ. Verification: - The focused test passes against the current bin/fm-contributions.sh. - I temporarily changed `settle_final` so it copied only the state. The test then failed, reporting the owner still on the old head (HEAD_A). I restored the file afterwards, and `git status` shows only the test file modified. - The full tests/fm-contributions.test.sh suite exits 0. No product code changed. The other CI finding (ci-1, "Behavior portable serial 9") was left alone because you chose to ignore it
* feat: enable supervision host by default for Claude primaries (#6124)
* feat: run the supervision host by default on a Claude primary
An absent config/supervision-host on a Claude primary now reads as on with
the default engine, and a file holding `off` opts any home out. Cursor,
OpenCode, omp, Grok, and Codex stay file-gated, with `off` read as disabled
there too. Every reader asks fm_supervision_host_enabled instead of testing
the file, and non-bash readers query it through the lib's `enabled` entry.
A primary's `off` is not inherited by secondmates: each home keeps its own
supervision posture.
* test: pin the watcher-path posture in fixtures that assume no supervision host
Fixtures that drive the watcher arm or assert a non-host drain now write
an explicit off file, and fixtures that copy the Stop auto-arm or the
supervision instructions carry the engine lib they now source. The two
drain suites also stop reading the code root's config.
* fix: name the opt-out when an off home passes an attended wake to main
A host parked when the home writes off now logs that the home does not run
the supervision host, rather than claiming it has no engine.
* no-mistakes(document): Clarify Claude supervision defaults and historical evidence
* no-mistakes(ci): Fixed process leaks in the two added host tests. Each case now stops its recorded watcher and host/arm processes; fake hook sessions exit through session.stop. The full host suite passed before the final cleanup refinement, and both affected cases, bash syntax, ShellCheck, and diff checks passed afterward. CI runtime still needs confirmation
* fix(bin): create the state dir on a fresh primary before the session-start scope check (#6125)
* fix(bin): create the state dir on a fresh primary before the session-start scope check
fm_primary_scope_matches required an already-existing state directory, so
bin/fm-sessionstart-run.sh stood down on a fresh clone before anything could
create it. Split out fm_primary_root_matches so the run wrapper can confirm
primary-home identity first, create the gitignored state dir when it is
missing, and only then run the unchanged scope check.
* no-mistakes(document): Document session-start state dir creation on fresh clones
* no-mistakes(ci): I fixed the Greptile P1 the way you asked. When a fresh primary can't create `state/`, the run wrapper no longer stands down silently. **Invariant:** when an otherwise eligible fresh primary cannot create `state/`, startup must never fail silently. This path has only one site: the mkdir in `bin/fm-sessionstart-run.sh`. Other hooks and the nudge wrapper never create `state/`, so they have no equivalent failure. **What changed:** - **Run wrapper** (`bin/fm-sessionstart-run.sh`): it captures mkdir's error and prints one line to stderr before standing down as before (exit 0, or 3 for the Pi prerequisite). The line looks like `fm-sessionstart-run: startup could not create the state directory <path>: <reason>`. - **Test** (`tests/fm-sessionstart-nudge.test.sh`): the new case `test_run_reports_a_state_dir_it_cannot_create` uses a fresh primary with no `state/` and a read-only (0500) root. It checks four things: exit 0, no digest on stdout, no state dir created, and exactly one stderr line ending in "Permission denied". It fails without the fix and passes with it. - **Docs** (`docs/sessionstart-nudge.md`): I added one sentence describing the stderr line and one describing what the new test proves. **Verification:** I ran `tests/fm-sessionstart-nudge.test.sh`, and every test passes. `bin/fm-lint.sh` on the changed scripts (pinned ShellCheck 0.11.0) and `tests/fm-documentation-audiences.test.sh` also pass. As you asked, the wrapper still stands down with the ineligible-checkout status afterwards. It does not report this as a failed eligible startup, which is what the bot suggested
* fix(bin): measure pending-reply grace from turn completion, not delivery (#6126)
* fix(bin): measure pending-reply grace from turn completion, not delivery
Fixes #6057
The pending-reply guard demanded a repost ("REPOST REQUIRED: previous
marked request had no correlated parent report") while the second
mate's correlated reply was already on its way.
fm_pending_reply_send_recovery measured its grace window from delivery
instead of from the request turn's completion, so any turn longer than
the grace fired the demand the moment the turn ended, before the reply
could have landed. The missed-report escalation had the same gap: it
fired the instant the recovery turn's completion was observed, with no
grace at all.
Both now measure grace from the relevant turn's completion (request
turn for the recovery repost, recovery turn for the escalation), and
both take one fresh, uncached read of the parent status file
immediately before firing, accepting a correlated line regardless of
its verb. Transport-failure escalations stay immediate, and the
one-repost limit is unchanged.
* no-mistakes(review): Document grace window as measured from turn completion
* no-mistakes(ci): Both Greptile findings were real and caused by this PR, so I fixed them. The full `tests/fm-pending-reply.test.sh` suite passes. **ci-1 (a reply could be overwritten by a repost).** The rule that must hold: a recovery send is recorded only if the record is still unresolved, checked under the same per-correlation lock that resolution uses. The escalation path already did this (`_fm_pending_reply_maybe_escalate_locked` reads fresh and publishes under one lock). The recovery path did not: `fm_pending_reply_send_recovery` did its fresh read through `fm_pending_reply_try_resolve`, which let go of the lock before the send was recorded. A reply landing in that gap could be overwritten, and the repost would go out anyway. Now `send_recovery` takes the lock once and, while holding it, re-checks that the phase is still `awaiting_report`, runs the fresh uncached read, and records the send (sender pid and identity, attempt time, phase `recovery_sending`). It releases the lock before actually sending, so the lock is not held during the send. It uses the same lock helpers the other lock wrappers use. Grace timing, the one-repost limit and the escalation path are unchanged. **ci-2 (the test would pass even without the fix).** In `test_recovery_fresh_status_read_resolves_before_firing`, the reply is still appended to the status file, but the stored file signature is then set to the file's new signature. That stands in for a same-size rewrite that the signature cache cannot see. The test first checks that a normal cached read misses the reply, then that the fresh read before sending catches it. I also added the same check for the fresh read before escalation, which the review said was uncovered. The test now sets its own send hook, so it no longer depends on one left over from an earlier test (that leftover had made failures exit silently). **Checks:** - I removed the fresh-read bypass at each site in turn and reran the suite. With it gone from recovery, the test fails with "recovery must not fire once a correlated reply has landed". With it gone from escalation, it fails with "the fresh pre-escalation read should have resolved the record, got escalated". With both in place, all tests pass. - Shellcheck with `-x` timed out locally. Without `-x` and ignoring SC1091, the only warnings are SC2034 on the existing `maybe_escalate` lock wrapper, which is not part of this change. The new code adds no warnings. Changes are in `bin/fm-pending-reply-lib.sh` and `tests/fm-pending-reply.test.sh`. Nothing is committed yet; a plain commit message such as "fix(bin): record the pending-reply recovery send under the fresh-read lock" fits the instruction
* no-mistakes(ci): ci-1 was real and caused by this PR. The same bug was also in the escalation path, so both are fixed. The full tests/fm-pending-reply.test.sh suite passes. The rule that must hold: a recovery repost or an escalation goes out only if the record's phase, read after the fresh-read resolve, is still what it was before. The resolver writes phase=resolved first and only then writes the other resolution fields. If one of those later writes fails, it returns an error even though the record is already resolved. Places this rule applies, both fixed: - Recovery (fm_pending_reply_send_recovery): the fresh-read resolve now runs first, and the phase is re-read right after it, whatever it returned. The send is recorded and made only if the phase is still exactly awaiting_report. This replaces the earlier phase check rather than adding a second one. - Escalation (_fm_pending_reply_maybe_escalate_locked): same bug. After a failed resolve it went on to publish the blocked line and set phase=escalated. One added line after the resolve call returns 1 without publishing if the phase has changed. Test: added test_partial_resolve_write_blocks_firing. It forces a failure on the resolved_epoch write after a correlated reply has landed. It checks that the recovery send hook is never called, that no escalation line is published, and that the phase stays resolved. The forced failure runs in a subshell so it can't affect later tests. Checks: - With the recovery fix reverted, the new test fails with "recovery must not fire after a partial resolve". - With the escalation fix reverted, it fails with "partial resolve should block escalation, got escalated". - With both fixes in, every test passes. - Shellcheck was run with SC1091 excluded and without -x, not through the repo's lint script. The only new message is one SC2329 info on the test's override function; other test overrides in the same file already get that same info, unsuppressed. Changed files: bin/fm-pending-reply-lib.sh and tests/fm-pending-reply.test.sh. Nothing is committed. Suggested plain commit message: "fix(bin): recheck pending-reply phase after the fresh read before sending
* fix(bin): stop provider-table lookup from writing broken-pipe errors to stderr (#6001)
* fix: provider-table lookup never writes a broken-pipe error to stderr
Fixes #5956
fm_quota_single_provider_for_harness returned from its while read loop
as soon as it found a match, closing the pipe while
fm_quota_single_provider_table's printf could still be writing.
Where SIGPIPE is ignored, as on GitHub Actions runners, bash then
prints "printf: write error: Broken pipe" on the resolver's stderr,
which intermittently broke the one-diagnostic-line assertions in
tests/fm-dispatch-resolve.test.sh.
Read the whole table before answering, the way
fm_control_harness_supported already does, so the writer always
finishes. Return values and output are unchanged.
Reproduced by running tests/fm-dispatch-resolve.test.sh with SIGPIPE
ignored on a single pinned core under CPU contention: 30 of 30 runs
failed before the fix, 0 of 30 after. Note: reproducing requires
setting the trap inside the tested shell because nice(1) resets an
inherited SIGPIPE ignore to SIG_DFL. tests/fm-quota-choose.test.sh
passes and bin/fm-lint.sh is clean.
* no-mistakes(ci): Fixed both Greptile findings the user chose to address. ci-1 (bin/fm-quota-axi-lib.sh:154). Invariant: looking up a harness must always end with status 0 and print the provider, even when the caller runs under `set -e`. The loop body `[ -z "$found" ] && [ "$harness" = "$1" ] && found=$provider` now ends in `|| :`. Every iteration succeeds and the whole table is still read. Only `fm_quota_single_provider_for_harness` loops over the table this way, so this is the one place the fix was needed. One caveat: on bash 5.3 the old code did not actually exit under `set -e`, because the `while` loop is not the function's last command, so the new `set -e` test would have passed before this fix too. The change makes the loop's success explicit, as the user asked. ci-2 (regression coverage). I added three cases to the existing `tests/fm-quota-choose.test.sh`, all calling the public lookup function after sourcing the library: 1. With SIGPIPE ignored (`trap "" PIPE`), it looks up every harness 200 times and checks that nothing reaches stderr. 2. A deterministic version of the race: the table function is wrapped so it writes the first row, pauses 0.2 s, then writes the rest. With SIGPIPE ignored, it checks that looking up `claude` prints `claude` and writes nothing to stderr. The stress loop alone reproduced the bug in only about 1 of 5 local runs, which is why this case exists. 3. A direct call under `set -e` prints `claude`. Verification: - `bash tests/fm-quota-choose.test.sh`: all pass. - Same test against the pre-PR library (fa48367, via `FM_ROOT_OVERRIDE`): fails with `printf: write error: Broken pipe`. The deterministic case failed in one run and the stress loop caught it in another. - `shellcheck` on both files: clean. - `tests/fm-dispatch-resolve.test.sh`: passes
* fix: restore portable CI behavior across Pi rendering and remote provisioning (#6162)
* fix: survive Pi 0.99 rendering and Git 2.55 local-clone races
Pi 0.99 puts arguments on the stock tool header and leaves hidden custom messages in the export conversation column. Match that header, and keep Calm's boundary on the visible column. Clone a remote home with --no-local so a prune during Git's loose-object copy cannot fail the seed.
* no-mistakes(review): Stop SIGPIPE write errors; cover older Pi export and project clones
* no-mistakes(document): Clarify Calm export visibility and tool rendering
* no-mistakes(ci): Fixed the dispatch diagnostic to list every provider-less use/default profile in one line and added a multi-profile behavior test. Shortened supervision fixtures using the existing engine-grace and park-clock knobs; removed stray scratch files. Dispatch tests, syntax checks, and three targeted supervision cases passed. CI’s prior supervision duration was 751s; the single permitted local full-suite run timed out at 1200s, so an after-duration is not established. The cancelled serial check had no failure verdict. The outer executor should record the measured before/after duration in the PR body when available
* no-mistakes(review): Gate Pi 0.99 call headers by version; drop hidden-row assertion
* no-mistakes(review): Test stock call headers under Pi 0.87 and 0.99 stubs
* no-mistakes(test): Fix older-Pi queued-row test and verify park-boundary behavior
* no-mistakes(document): Clarify Pi Calm export and queued-turn documentation
* no-mistakes(ci): Fixed the stock macOS Bash 3.2 parse failure in tests/fm-calm-pi-extension.test.sh; its parse check passes. The watcher CI failure is in unchanged code: the isolated five-minute/66-minute case passes locally, but the CI log omits the drain error needed to establish its cause. No speculative watcher fix was made. The full local watcher suite timed out after 500 seconds
* fix: confirm Lavish board replies before worker handoff (#6169)
* Prevent premature Lavish board handoffs
* Prove Lavish arm lacks reply acknowledgement
* Confirm Lavish replies before arming worker boards
* no-mistakes(review): Post Lavish reply only after locked arm eligibility checks
* no-mistakes(review): Fail Lavish reply closed on unknown version
* no-mistakes(document): Correct Lavish reply documentation and remove stale guidance
* no-mistakes(document): Clarify Lavish reply routing and remove duplicate version guidance
* fix: inherit supervision host opt-out across secondmates (#6154)
* feat: inherit the supervision-host opt-out from the primary
Move the supervision host's off opt-out out of config/supervision-host into
its own presence flag, config/supervision-host-off, and add that flag to the
primary-authoritative inherited config set. A primary that opts out now opts
every secondmate home out at spawn and convergence, and clearing it converges
them back. config/supervision-host stays the home-local engine choice.
Shape: config/supervision-host mixed two things, a fleet posture (off) and a
per-home engine and model. Only the posture should follow the primary, so it
becomes a separate presence flag that rides the existing inherited-config
mechanism (FM_INHERITABLE_CONFIG in bin/fm-config-inherit-lib.sh) with no new
machinery, while the engine line stays local. The parse stays in its one
owner, fm_supervision_host_enabled. There is no migration or compatibility
handling for a home that still holds off in config/supervision-host.
Primary off, mate on: inherited material is primary-authoritative by design,
so a mate cannot keep the host while the primary is opted out, and a mate's
own opt-out is removed at the next convergence while the primary has none.
Running the host on a mate is the primary's choice for the fleet; no override
mechanism is added.
Live validation (disposable bin/fm-live-lab.sh lab, Claude primary with a
real seeded secondmate, --supervision-host off):
- up: every readiness check ok, including "host: none running, as expected"
and a live mate session; the spawned mate home held the inherited
config/supervision-host-off and the gate read primary OFF, mate OFF.
- primary removed its opt-out, then bin/fm-config-push.sh reported
"supervision-host-off: pushed - mirrored primary absence" and a config
reread sent; the gate read primary ON, mate ON, and the live mate handled
the reread.
- primary opted out again and pushed: "supervision-host-off: pushed", mate
gate OFF.
- down stopped every lab process and left no lab process running.
Out of scope, follow-up: default-on for the other harnesses, away-daemon
retirement, rollout.
* no-mistakes(document): Document inherited supervision-host opt-out ownership
* no-mistakes(ci): Fixed ci-4: with `--supervision-host off --mate`, lab readiness now requires the inherited flag in the mate home and a disabled mate supervision-host gate. The focused behavior test, shellcheck, and diff checks pass. Left ci-1–ci-3 untouched as directed
* no-mistakes(test): Fix mate readiness HOST_OFF initialization in lab up
* no-mistakes(ci): Fixed Lint 2 by making the new test’s fixtures source resolvable to ShellCheck; its off/on readiness test and ShellCheck now pass locally. Behavior portable serial 5 failed in the unchanged remote-reply test at generation 7. That test passes locally, and no PR-caused defect was identified, so no remote-reply code was changed
* fix: reduce supervision exit latency and stabilize host tests (#6179)
* fix(tests): cut the fixed sleeps in supervision-host cycles
The serial CI lane keeps brushing its 30-minute cap because
fm-supervision-host.test.sh spends ~903s of the job, and per the
run-36635306527 case profile the top nine cases are all multi-cycle
ones (3-10 park/close/turn cycles each): every close waits out the
host's sleep $POLL in await_close plus a watcher sleep $FM_POLL scan
cycle, and every engine turn waits out the fixed sleep 1 descendant
snapshot. That is ~3s of pure sleep per cycle before any real work.
The host poll now accepts positive decimal seconds through a new
seconds_or validator (FM_SUPERVISION_HOST_POLL), and the engine turn's
snapshot loop takes FM_SUPERVISION_ENGINE_SNAPSHOT_SECONDS, also a
positive decimal defaulting to one second - the smallest seam at each
wait's single owner. The suite drives them at 0.2 alongside the
existing FM_POLL=0.5 and FM_ARM_ATTACH_POLL=0.2 knobs, so the real
poll loops still run. The park-boundary case moves onto the injected
test clock instead of a real 3s wait, per-case cleanup polls the host
pid rather than sleeping a full second, and the proof-by-absence
windows (flood re-escalation, successor re-announce, watcher
persistence, recovery staying off main) shrink from 2-3s to 1s, which
still spans two watcher polls at the test cadence.
Every assertion, process lifecycle, and reaping path is unchanged;
production defaults stay at one second. Isolated case timings on a
contended host, base vs branch: attended-latch 54.3->34.6s,
undelivered-dialog 67.7->59.1s, away-latch 46.5->30.5s, held-cadence
47.9->21.6s, unreadable-mirror 39.2->38.5s, park-limit 18.2->12.3s,
registration-fallback 14.1->10.0s, first-cycle-status 12.6->8.4s,
latch-scope 16.7->16.3s. Full suite: 65/65 pass. fm-lint and
shellcheck clean.
* no-mistakes(review): Wait for scan lock release before duplicate check
* no-mistakes(document): Correct supervision snapshot cadence documentation
* fix(tests): keep production poll cadence, probe exits at 0.1s
The fractional poll cadences multiplied the cost of each loop body:
full process-table scans in the engine turn and process refreshes in
await_close ran five times more often, which swamped the thin CI runner
and nearly doubled every multi-cycle case (serial 5 was cancelled at its
30-minute limit on run 36635306527's successor). Restore the production
cadence and notice arm/engine exits with a cheap kill -0 probe at a
tenth of a second between the one-second bodies instead: strictly less
dead time than baseline with no added CPU.
Also hold each injected-clock park bound well past its case's
wall-clock checks so a host that ignored the test clock fails instead
of silently passing at a real-time boundary, and restore the shortened
proof windows (watcher liveness, recovery-off-main absence, first-cycle
stream) to their baseline depth.
* no-mistakes(document): Clarify supervision engine snapshot documentation
* ci: rebalance portable test groups and enforce a packing budget (#6192)
* fix: rebalance portable CI from current duration measurements
* no-mistakes(test): Test serial packing boundary and verify endpoint timeout cleanup
* no-mistakes(document): Clarify timeout guidance and remove duplicated packing estimates
* fix(bin): run no repository hook when core.hooksPath is empty (#6216)
* fix(bin): run no repository hook when core.hooksPath is empty
The per-task hook wrapper refused every commit in a repository whose own
config sets core.hooksPath to the empty string, because git rev-parse
--git-path hooks fails on it. Plain git reads that setting as no hooks, so
the wrapper now runs none; every other lookup failure still refuses and
shows git's error.
Fixes #6171
* no-mistakes(review): Refuse commits when core.hooksPath is a valueless key
* no-mistakes(document): Document empty core.hooksPath handling in commit attribution docs
* no-mistakes(ci): When the wrapper refuses a commit, Git's hook-lookup error now shows up once instead of twice. That required changing one line in the wrapper, and the tests were extended so both bad-config cases would catch the duplicate. Invariant: when the wrapper refuses, Git's lookup error must appear exactly once. In the failure path, the only Git call besides the deliberate second lookup is the `git config --get --type=path core.hooksPath` check in `runtime_chain_body` (`bin/fm-git-strip-ai-trailers.sh:168`). That check prints the same error, so it was the one place to fix. I added `2>/dev/null` to it. Its exit status still decides the outcome: an empty value still runs no hook, and anything else goes on to the second lookup, which prints Git's error once, and the commit is refused. Tests (`tests/fm-git-strip-ai-trailers.test.sh`): - The unresolvable-path test (`~fm-no-such-user-6171/hooks`) now requires `failed to expand user dir` to appear exactly once in the refused commit's output. - The valueless-key test now requires `missing value for 'core.hookspath'` to appear exactly once. - Pre-existing bug in the unresolvable-path test: its `git add` ran after the bad config was set, so it failed silently (exit 128) and the "refused commit" had nothing staged. The test now stages the file before writing the config, the same way the valueless test does, so a real commit gets refused. - The empty-string test is unchanged and still passes, so an empty `core.hooksPath` still runs no hook. Verification: - With the wrapper change reverted, both new checks fail with `expected '1', got '2'`. With the change in place, the whole suite passes. - `bash -n` passes. shellcheck shows only an info-level SC1091 note about sourcing `lib.sh`, which was already there before this change. - `git status` lists only the two intended files
* fix(bin): let a stale record on a reassigned slot retire records-only (#6213)
* fix(bin): let a stale record on a reassigned slot retire records-only
When a pool slot's owner claim names another task, the stale record's
teardown touches nothing under the slot, so the exclusive-slot record scan
no longer refuses it. Full teardowns of a slot this task still claims, or
one with no claim, keep the refusal.
Fixes #6184
* no-mistakes(document): Note claim-over-record precedence for reassigned teardown slots
* fix(bin): keep the steering doorbell short under deep homes (#6240)
* fix(bin): keep the steering doorbell short under deep homes
The doorbell printed the task inbox's absolute path twice, so under a deep
home it grew to about 290 characters and a Herdr submit reported it never
reached the pane on every re-ring. It now names the inbox once by its short
<task>.inbox name and points at the full path the worker's brief already
gives, so its length no longer depends on the home's depth.
Fixes #6120
* no-mistakes(review): Export FM_TASK_INBOX at launch and name it in doorbell
* no-mistakes(ci): ci-1 (Behavior portable serial 9) was caused by this PR, and I fixed it in the test. tests/fm-claude-trust.test.sh failed with "the launch command did not carry a brief doorbell". Its claude_launch_doorbell helper stripped exactly two leading `export ...;` statements before reading the final prompt argument. This PR adds a third one (`export FM_TASK_INBOX=...`) to every launch, so the helper was reading the wrong command. The invariant: a test that parses the launch command must skip every leading export statement, however many there are. I checked every test that parses the launch this way. The only other ones are the two helper…
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
A
Context: that is the choice to ship parts 1+2+3 in one PR, as recommended by the two watcher-stall scout reports (data/fm-watcher-stall-s1/report.md and data/fm-watcher-cycle-cost-s1/report.md in the fmdev-f1 home). In the main home, watcher poll cycles take 130-270 s (one about 328 s) against a 300 s heartbeat grace, because every cycle walks all ~2,030 never-pruned pending-reply records with per-record locks, forks, and re-sourcing; under load a cycle passed the grace twice on 2026-09-28 and the Stop-hook auto-arm failed each time. The three parts:
Reproduce first in a real setup with a large pending-reply store, and prove each part with observable behavior (cycle time, attach behavior, listener stability). Not approved, follow-ups only: archiving long-resolved records, and decoupling the contributions poll from the check cadence.
What Changed
Risk Assessment
High (review): the review found no code defect. Its one finding, that the real-setup reproduction and cycle-time proof were absent from the diff, was resolved by the lab evidence in "Reproduction and measurements" below; that evidence stays outside the diff by decision.
Testing
No baseline commands were supplied for this phase. The three targeted end-to-end suites passed: settled records were skipped despite a held lock, an attached arm followed a slow watcher and handled a stalled holder, and a preempted remote-reply listener kept its claim without a false caught-up watermark or relaunch. A separate disposable 2,030-record tick took 0.132–0.141 seconds across three runs. The test fixtures were cleaned up; this measures the pending-reply tick, not a full watcher cycle or a before/after comparison.
bash tests/fm-pending-reply.test.sh;2030-record-cycle.logbash tests/fm-watch-arm.test.sh— slow-holder wake scenariobash tests/fm-watch-arm.test.sh— stalled-holder replacement scenariobash tests/fm-remote-reply.test.sh— preemption and listener-stability scenariosEvidence: 2,030-record pending-reply tick timings
Source: 2,030-record pending-reply tick timings
Reproduction and measurements
Every part was reproduced on the base code (
b5fdf74d, a pristine copy) before the fix and re-run on the fix, in disposable homes on the Mac mini (load 5-8); no live home was touched.Part 1 - one-pass record selection
Lab: 2,030 resolved records (every 7th with a closed escalation) plus 3 needing work.
fm_pending_reply_tick, 2,033 recordsfm-watch-arm.sh+fm-watch.sh,FM_POLL=15, 2,030 recordsOld and new ticks on identical 2,033-record labs with a fixed clock left identical record and status files; the only differences were the wall-clock
[at=]stamp on the appended escalation close and a file identity.Part 2 - attached arm follows a live holder
Real Stop hook -> supervision host -> arm -> watcher under a fake
claudesession (the scouts' Appendix A harness), grace scaled to 20 s: old attach give-up at 31 s, stall bound 60 s (main: 300 / 311 / 900 s). The second park's first arm attaches to the main-only pass-through's detached successor, as on main.attached-cycle-ended beacon_age=31, retrynonzero-exit beacon_age=32against the live holder, hook exit 2 withfirstmate watcher auto-arm FAILEDsignal:lineattached-holder-stalled beacon_age=60; retry evicted the holder (stalled-holder-replaced) and main got an ordinarycheck: rearm-resurfacewakewatcher_live=no hook_rerun=no queued_rows=2 arms_following=0)Part 3 - preempted reply long-poll
Real adapter, runner,
fm-on.sh, remote entrypoint, and job worker (only ssh faked); a same-home non-preemptible job, the liveness probe's shape, preempts the listener's poll.no-result: remote-reply-ios (exit 76), claim releasedfm-procevent.sh reconcilestarted=1started=0Regression tests
Each new test fails on the base code and passes on the fix:
test_tick_leaves_settled_records_alone("the tick blocked on a settled record's foreign-held lock"),test_attached_arm_follows_a_slow_live_holderandtest_attached_arm_hands_a_stalled_holder_to_its_replacement(watcher: FAILED - cycle ended without an actionable reason), and the remote-reply preemption cases ("a preempted reply poll did not report a closed window: 76").A 2,000-record timing assertion is deliberately kept out of the unit suite; the cycle-time proof is the lab measurement above.
Follow-ups and known limitations
queued_rows=2 arms_following=0.Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
tests/fm-pending-reply.test.sh:1118- The intent requires: “Reproduce first in a real setup with a large pending-reply store, and prove each part with observable behavior (cycle time, attach behavior, listener stability).” This test builds 600 settled records and checks that the tick finishes within 60 seconds, but the change provides no before/after watcher-cycle timing for the reported ~2,030-record setup. The attach and listener tests provide behavioral evidence; the required real-setup reproduction and cycle-time proof remain absent from this change. Please provide that evidence or ask the author whether this criterion can be deferred.✅ **Test** - passed
✅ No issues found.
bash tests/fm-pending-reply.test.sh;2030-record-cycle.logbash tests/fm-watch-arm.test.sh— slow-holder wake scenariobash tests/fm-watch-arm.test.sh— stalled-holder replacement scenariobash tests/fm-remote-reply.test.sh— preemption and listener-stability scenariosbash tests/fm-pending-reply.test.shbash tests/fm-watch-arm.test.shbash tests/fm-remote-reply.test.shRan threefm_pending_reply_tickcycles against a disposable 2,030-record store and measured elapsed time.✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.