Repository navigation
Merge upstream kunchenguid/firstmate main into the fork, keeping #5895, #5911 and #5929 - #1
Merged
Merged
Conversation
…configurable bound (kunchenguid#5917) Each per-task endpoint read in the session-start fleet digest runs in its own crash-isolated child under the existing timeout helper with a configurable FM_SESSION_START_ENDPOINT_TIMEOUT bound (default 10s). A hung or killed read becomes that task's endpoint error while the digest continues, and the wrapper banners any abnormal child exit naming its status.
…nchenguid#5886) * Keep lab tmux sockets on short private paths * no-mistakes(review): Fix Linux stat checks and fail-closed lab tmux teardown * no-mistakes(document): Update lab helper documentation for isolated tmux sockets
* Allow silent task-level no-change outcomes * no-mistakes(review): Exclude silent outcomes from captain-return handoffs * fix: look up supervision receipts by exact sequence * no-mistakes(review): Suppress silent notes in away-return brief * no-mistakes(review): Clarify visible notes; remove unused mode * no-mistakes(review): Clarify silent outcomes and avoid false drain promises * no-mistakes(document): Clarify silent supervision outcome documentation
…chenguid#5928) * feat: make /quiet a statement where the attended supervision host runs On a home that opted into the supervision host, quiet mode is what the attended host already does, so /quiet now enters nothing there instead of launching the quiet daemon and writing a record that would park a present captain's main. - bin/fm-afk-launch.sh quiet-check says quiet mode needs nothing where the attended host runs, or that the session is paused while its broken-session latch holds; a quiet enter refuses there before writing anything. - Where the home opted in but the attended host lacks a part (engine, tools, verified mirror writer, identifiable main session, valid mirror), quiet-check names it and quiet mode falls back to the daemon. - Under a live away record on that home, quiet-check and a quiet enter refuse and name the record, so the return runs first, whatever state/.afk says. - A quiet enter records mode: quiet in the posture record, so start and start-native launch the quiet daemon without FM_AFK_MODE, and the away refusal wording fires only for away. - bin/fm-host-mirror.sh check validates the dialog mirror read-only and exits 1 on a missing, unreadable, or invalid mirror. - The quiet and afk skills and the supervision-host docs describe the new behavior; homes without the opt-in and Pi homes keep the daemon path. * no-mistakes(review): Archive the quiet record when a quiet daemon start fails * no-mistakes(document): Clarify quiet-mode documentation and remove stale duplicates
…uid#5884) * fix(bin): grant Claude workers their task-channel dirs via --add-dir Since Claude Code 2.1.257, a file-tool read (Read/Glob/Grep, and an Edit's mandatory prior read) of a path outside the working directories parks --permission-mode auto panes on a one-time interactive question, and a "Block" answer lands permissions.blockReadsOutsideWorkingDirectories in user settings, refusing the same reads even under bypass. Firstmate launches Claude with no --add-dir, so a secondmate's parent-home steering inbox and a ship or scout worker's launch record, steering inbox, brief dir, and code-root .agents/skills were all outside: workers wedged on the question the first time they read a steer. Every Claude launch, spawn and relaunch, in both permission modes, now grants exactly the task's channel directories: state/<id>.inbox for a secondmate (in the parent home), or state/operational-inbox, state/<id>.inbox, data/<id>, and the code root's .agents/skills for a ship or scout. Paths resolve to real paths and lazily created channel dirs are made before launch so the grant never names a not-yet-existing directory; the whole state/ is deliberately never granted. The grant keeps the bypass-mode launch argv changed on purpose: it also protects bypass workers against a machine-recorded Block answer. * no-mistakes(document): Consolidate Claude launch guidance in configuration reference * no-mistakes(document): Clarify Claude permission documentation reference
…nguid#4819) * fix(supervision): prevent idle recovery loops without stranding wakes * no-mistakes(review): Remove unused wake-append rollback helper
…id#5889) * fix(bin): stop the remote-job worker busy-polling an idle queue The serving loop slept 50ms between passes and re-ran state preparation (chmod on every queue directory), the heartbeat publish, and the stale sweep on every pass. It now blocks on a worker.wake FIFO that staging, cancellation, and lane exit nudge, keeps a short fast-poll window after activity, refreshes the heartbeat at most once a second, and runs the sweep (which re-applies the queue directories' 0700 modes) at startup and then on a bounded interval. Lane-owned records are no longer re-read every pass. Measured with a fork/execve-interposing counter on a --serve worker in a disposable HOME and queue, bash 3.2, 20-second windows (the counter slows the old loop to about 5 passes a second, so real-host rates were higher): idle worker 146 forks/s, 61 execs/s -> 4.8 forks/s, 3.1 execs/s one running long job 232 forks/s, 100 execs/s -> 15 forks/s, 11 execs/s Stage-to-result latency for a no-op job, idle and back to back, stayed at about 0.8-1.2s in both versions (dominated by job execution, not pickup). * perf(bin): drop per-cycle forks from watcher, drain, and lock helpers The watcher, drain, inactive-reconcile scan, and branch-outcome reads forked small external commands on every cycle where bash can do the same work. - fm-wake-lib.sh gains fm_dirname_to, fm_basename_to, and fm_epoch_seconds_to, exact stand-ins for $(dirname --), $(basename --), and $(date +%s); the clock uses printf %(%s)T on bash 4.2+ and still forks date exactly once on stock macOS bash 3.2. - fm_lock_abs_path, fm_wake_signal_seen_path, fm_path_age, the watcher's age_of and wedge timer, and the recovery-marker line count use them or plain reads instead of dirname/basename/tr/date/wc. - window_to_task reads a meta file once instead of two grep | tail -1 | cut -d= -f2- pipelines per file per call. - fm-classify-lib.sh reads uname -s once at source time instead of in every status stat helper. - Libraries sourced every cycle derive their own directory without forking dirname, including the backend adapter siblings a subshell re-sources on each probe. tests/fm-fork-free-helpers.test.sh pins each replacement against the command it replaces on edge-case inputs, under every available bash and both the C and a UTF-8 locale; CI's stock macOS bash lane runs it under /bin/bash 3.2. Measured with a fork/execve-interposing counter in a disposable home, one tmux crew task, FM_POLL=1 (forks and execs per watcher cycle, per run otherwise): watcher cycle bash 5.3 299/138 -> 199/66 bash 3.2 341/146 -> 224/80 drain bash 5.3 492/238 -> 430/200 bash 3.2 567/250 -> 491/212 inactive scan bash 5.3 27/14 -> 17/4 bash 3.2 37/14 -> 17/4 branch-outcome bash 5.3 40/21 -> 35/16 bash 3.2 48/24 -> 38/19 * test: note the interpreter-expanded version probe for shellcheck * no-mistakes(review): Fix worker idle bounds and fork-free contributions snapshot * no-mistakes(review): Coalesce buffered worker wake nudges into one wake * no-mistakes(review): Coalesce wake nudges via pending marker so publishers never block * no-mistakes(review): Claim wake nudges atomically via noclobber pending marker * no-mistakes(review): Release abandoned wake claims only after a 30-second bound * no-mistakes(review): Drop worker wake FIFO; load path helpers side-effect free * no-mistakes(document): Document remote worker polling and preemption cadence * no-mistakes(ci): Fixed both failing CI shards: isolated remote and teardown test fixtures now include fm-path-lib.sh, which fm-wake-lib.sh requires. The three affected tests, fm-lint.sh, and git diff --check pass locally
…uid#5941) * fix: keep a successor watcher and remote-reply listeners across the gaps that dropped them A main-only supervision pass-through exited without leaving a watcher, and each remote-reply poll released its claim until the next cycle, so short-lived listeners stayed down. * no-mistakes(document): Clarify listener and supervision continuity documentation * no-mistakes(ci): Fixed the three Greptile findings: failed ingestion leaves one durable capture, failed reads exit instead of relistening, and the disposable-checkout guard rejects state paths outside the marked lab. Added behavioral tests; the remote-reply and watcher-lock suites, shell syntax checks, and git diff checks passed * no-mistakes(ci): Fixed a race in the wake-queue interruption test: it now waits for the drain to own the lock and enter handling before signaling it. The wake-queue suite, shell syntax check, and diff check pass
…enguid#5925) * fix: date replayed branch outcomes and ask main to check current state first A captain outcome main never acknowledged is presented again, which after a harness or posture switch, or the first drain after the upgrade whose earlier presenter never advanced the read cursor, can be days after its situation settled. The replay read as fresh news, so a PR since merged looked ready. bin/fm-branch-outcome.sh now adds a "recordedAgo" age (minutes, hours, then days) to present and unprocessed rows, one owner of that wording for both presenters. The drain's BRANCH OUTCOMES captain lines and the Pi branch's processing request name that age and ask main to check the task's current state first; an outcome already settled needs only the acknowledgement, with nothing relayed to the captain. Nothing is adopted as processed, so a fresh home's first outcome is still presented until acknowledged. * no-mistakes(review): Absent processed marker reads 0; never adopt read cursor * no-mistakes(review): Require recordedAgo in Pi requests; report undated rows to main * no-mistakes(review): Keep recordedAgo on captain rows only in present output * no-mistakes(document): Correct cutover documentation and retire stale migration guidance * fix: keep settled branch outcomes out of main's reply to the captain A live Pi primary that took over a host-drain home received the carried-over outcomes dated and check-first, but its processing reply still told the captain about an outcome whose decision had since been answered. The request also claimed every outcome was already shown as an anchor entry in this transcript, which is false for an outcome carried over from before a restart or a switch of primary. The Pi processing request now says each outcome was recorded earlier and may already have been seen or handled, and that a settled outcome gets no captain-facing mention at all in the reply or any recap, not even that it is settled. The drain's BRANCH OUTCOMES header and the supervision docs state the same rule, and the tests check both delivered texts. * fix: scope main's outcome reply to what is still open Telling main what not to say about a settled outcome was not enough: in two live Pi trials the processing reply still told the captain that an answered decision was settled. Main now sorts the outcomes by current state first, and its reply to the captain covers only the still-open ones, written as if the settled ones had never been listed. With that framing three live Pi trials kept the settled outcome out of the reply and relayed the open one each time. The drain's BRANCH OUTCOMES header and the supervision docs use the same framing, and the tests check both delivered texts. * no-mistakes(review): Clarify that main acknowledges every presented captain outcome * no-mistakes(document): Clarify outcome cursor ownership across Pi and host * no-mistakes(ci): Fixed Pi replay by batching unprocessed captain outcomes oldest-first and acknowledging only through each batch. Verified a marker-less backlog over 1 MiB replays through all batches. A real-drain regression confirms an older keyed decision remains under OPEN DECISIONS after a newer branch row is acknowledged; the check-first instruction now names those decisions. Relevant targeted tests and branch-supervision tests passed; the full host suite timed out * no-mistakes(ci): Fixed the host drain’s check-first wording in bin/fm-wake-drain.sh; the CI fixture now passes. The full host suite passed the affected fixtures but timed out later. Syntax and diff checks passed * no-mistakes(ci): Fixed ci-1: abbreviated Pi outcome summaries now stay within 1,024 characters and include a row-specific full-outcome lookup command. The delivered instruction requires reading the full outcome before acting, relaying, or acknowledging it. The new extension-driver regression failed before the fix and passes now; the Pi and supervision-host suites pass * no-mistakes(ci): Corrected the batching sentence in docs/pi-supervision-branch.md. The cancelled CI check needs no code fix; its clean rerun passed. The Pi branch extension suite and git diff check passed
…guid#5961) * fix: wake an idle Claude primary for attended main-only hand-backs An attended main-only pass-through confirmed a handling handoff for the successor it leaves running, which flipped the recovery marker to handling. The Claude Stop hook only rewakes main while that marker reads downtime, so the close reached no one and an idle primary slept with wakes queued. The pass-through now leaves the marker at downtime, and a close that turns main-only at its turn hands the consumed handoff back to downtime before it reaches main. Regression tests drive the real Stop hook around the real host on both paths and for the successor's own later close, and a new opt-in live guard proves it against an idle interactive Claude primary with a pre-fix negative control. * no-mistakes(review): Assert live lab Stop-hook registration via parsed settings JSON * no-mistakes(document): Correct supervision hand-back documentation * no-mistakes(ci): Fixed the failed downtime-write path so the Stop hook notifies main instead of silently dropping the close. Corrected the live guard’s tracked-hook check and added the requested at-turn main-only scenario. The new regression failed before the fix and passed after it; the host suite, syntax checks, and diff check pass. The credentialed live guard was not run in this CI phase because it writes outside the worktree * no-mistakes(ci): Fixed the Stop hook’s retry ordering: a crashed host gets its bounded retry before a non-crash hand-back failure is reported. The Stop-hook suite passes, including the crash regression. The host suite passed the failed-marker-write regression but timed out before completing; syntax and diff checks pass
kunchenguid#5916) * Fix worker launches to enter recorded worktrees * no-mistakes(review): placeholder * no-mistakes(document): Update agent-control.md worktree-refusal note to match new universal cd+assert * no-mistakes(review): Add regression tests for Orca spawn/relaunch worktree carve-outs * Fix PR relaunch and prelaunch cwd verification * no-mistakes(document): Fix docs/agent-control.md: worktree cwd check is pre-launch, not post-launch * fix: slim worktree launch change onto upstream main
…ardown (kunchenguid#5997) * WIP: retire task-keyed watcher markers and orphan journals at teardown Re-applies old PR kunchenguid#5584 on current main: teardown retires the turn-ended .seen-* signature and an orphaned Herdr presentation journal whose workspace is already gone, and the wake-drain rotates its own dead scratch files. Not yet validated through no-mistakes. * no-mistakes(ci): Fixed the Greptile P1 finding in bin/backends/herdr.sh. fm_backend_herdr_projection_token_workspace_gone used `! ... jq -e ... 2>&1`, which swallowed a jq runtime error (thrown when a non-object workspace entry, e.g. a number before a live token-bearing workspace, hits `.label`) and flipped it to a "gone" verdict, causing teardown to delete a still-live v1 presentation journal. Invariant: a workspace-query error/ambiguity must never be read as token absence; only a cleanly-parsed list with no token-bearing label is "gone". Replaced the body with a single jq verdict (unknown/present/gone): a non-array list or any non-object/non-string-label entry yields "unknown", jq errors/empty output fall through `|| return 1` to unknown, and only "gone" returns 0. Sibling fm_backend_herdr_projection_endpoint_matches_journal already fails safe on jq error (empty match -> journal kept), so it needed no change, matching the author's scoping. Added test_teardown_retains_v1_journal_when_workspace_query_ambiguous driving real teardown with a malformed workspace-list entry, proving the journal is kept and no workspace close occurs. Verified the old logic returns GONE on that input (test fails before, passes after); full tests/fm-teardown.test.sh suite passes (exit 0) and shellcheck is clean. Marker-naming finding left untouched per explicit out-of-scope instruction
…unchenguid#6002) * fix(tests): disable Claude Code's auto-updater during live harness runs fm_live_gate let a live run proceed without ever setting DISABLE_AUTOUPDATER, so a live Claude test could let the real updater repoint ~/.local/bin/claude into a temporary directory and stop every Claude process on the machine from starting. Export DISABLE_AUTOUPDATER=1 on every path where the gate lets a live run proceed, and assert the export in tests/fm-live-gate.test.sh, including that it reaches a child process the same way a real harness pane would inherit it. * no-mistakes(ci): Greptile flagged that the PR's DISABLE_AUTOUPDATER inheritance test only checked a `bash -c` direct child, not the fm-spawn.sh launch path. The user chose to fix it with a regression on that path. In tests/fm-live-gate.test.sh I replaced the generic child test with test_disable_autoupdater_reaches_the_claude_pane_on_the_fm_spawn_launch_path: it drives the real fm-spawn claude launch through the spawn fixtures, captures the exact staged launch command, and runs it as a synthetic pane whose only `claude` is a stub recording the inherited DISABLE_AUTOUPDATER, asserting it saw 1. Switched the file to source fixtures.sh (pulls in lib.sh, guarded) for the spawn helpers. Verified it is a real guard: the stub records `1` when the ambient var is set and `unset` when absent, so it fails if fm-spawn ever scrubbed the variable (e.g. env -i or -u). This confirms fm-spawn's launch construction never references the name and passes it through via ordinary ambient inheritance with no allowlist. Full suite passes (12 tests ok), shellcheck clean. Note for the outer executor: I embedded the daemon caveat as a code comment in the test, but the finding also asks the PR body to state that a backend daemon already running before the gate exported the variable does not inherit it and fully covering that would need launcher support - that forge-side PR-body sentence is outside this CI phase's scope * no-mistakes(ci): Fixed Greptile finding ci-1. Root cause: fm-spawn.sh handed its launch command to an already-running backend daemon that never inherited the test process's exported DISABLE_AUTOUPDATER, so ambient inheritance dropped it and Claude's auto-updater could still run. Fix (bin/fm-spawn.sh): when DISABLE_AUTOUPDATER is set in the spawn's own environment, embed `export DISABLE_AUTOUPDATER=<value>;` into the LAUNCH command text (same idiom as the adjacent COMPACT_ADVISER_DISABLE export), so it survives a daemon-built pane, the env -i allowlist path, and relaunch alike; gated on presence so ordinary spawns are unchanged. Added regression test test_disable_autoupdater_survives_a_daemon_pane_that_never_inherited_it in tests/fm-live-gate.test.sh: stages a real claude launch with DISABLE_AUTOUPDATER set in the spawn env, then runs that exact command in a synthetic pane with `env -u DISABLE_AUTOUPDATER` and asserts the claude stub still recorded autoupdater=1. Verified the test fails (autoupdater=unset) without the fix and passes with it; the round-1 ambient test stays green either way. Full suite passes (13 ok); test file and isolated snippet shellcheck-clean (full fm-spawn.sh shellcheck kept getting terminated by the memory-constrained host, not by findings). Forge-side note for the outer executor: the PR-body caveat that fully covering the daemon case would need launcher support no longer applies to the Claude launch path and should be corrected
…cevent record (kunchenguid#6010) * fix(bin): ring the inbox doorbell only for a newly published procevent result publish_result rewrote a worker's captured Lavish round idempotently on every reconcile, unconditionally moved an already-acknowledged inbox record back out of handled/, and rang the doorbell every time - so an already-processed round rang the owning worker on every cycle. Snapshot the existing active and handled records before the idempotent write and ring, or move anything, only when the write actually created a fresh record; re-delivery of a still-open round is left to the inbox's own re-ring ladder. * no-mistakes(document): docs: reflect worker-board doorbell rings only on fresh inbox record * no-mistakes(ci): Fixed Greptile finding ci-1 in tests/fm-procevent.test.sh (test-only change). The redelivery regression previously moved the delivered note into handled/ before any repeated reconciles, so it only proved an acknowledged note stays quiet and would still pass if an unchanged active note rang every cycle. Per the user's instruction, I inserted (before the mv into handled/) five repeated `pe reconcile` runs with the note still in the active inbox and asserted the ring log holds exactly one line and 001.msg remains active; the existing acknowledged-note assertion after the move is kept unchanged. No product code changed. bash -n confirms syntax is valid; the block mirrors the already-passing post-move reconcile/ring-count assertion directly below it
…unchenguid#6032) * fix(bin): make the Claude Stop auto-arm refuse arguments before arming A model running bin/fm-claude-stop-autoarm.sh --help mid-turn armed a real supervision-host park owned by its short-lived tool process, leaving supervision down once that process exited. The Stop hook passes no arguments, so -h/--help now prints usage and any other argument is refused before anything is sourced, read, or armed. * no-mistakes(document): Clarify Claude Stop hook documentation for manual invocations * no-mistakes(ci): Updated the argument-run regression test to compare checksums of state files as well as entry names. The Stop auto-arm test suite passes, and git diff --check is clean * docs: restore the bin/ toolbelt intro's manual-use clause The document step dropped "interactive entrypoints work by hand too" from docs/scripts.md, which still holds for most bin/ scripts. * no-mistakes(document): Clarify Claude auto-arm manual-use guidance
…uid#6039) * feat(calm): show supervision sailboat and anchor notes on Claude Code The Calm mod follows a bounded display tail copy of the outcome store, which bin/fm-branch-outcome.sh append now refreshes, and the supervision host's latch, and appends one dim transcript line per visible routine outcome, captain outcome, and latch change, replaying unread and unprocessed outcomes at session start. It shows them whenever the mod is active, regardless of config/calm, and never marks anything read. * fix(calm): show each supervision note once per session on Claude Code Claude Code 2.1.283 stores ui.log lines in the session and restores them on --continue, so the mod records how far each session has followed the outcome store and a resume replays only newer outcomes. It also checks file existence before reads so absent files do not log debug errors. The live guard gains the supervision-notes scenario and the dated 2.1.283 record documents the observed behavior. * docs: name the Claude supervision note row as the engine draws it * no-mistakes(review): Seed outcome tail on present and anchor first tail on markers * no-mistakes(review): Seed outcome tail at session start; replay against start markers * no-mistakes(review): Bound outcome tail by bytes; reread recently changed files * no-mistakes(review): Skip store validation when outcome tail already exists * no-mistakes(document): Clarify bounded Claude supervision note replay * no-mistakes(ci): Fixed seed-tail to validate only a bounded suffix of complete store rows and write it through the existing byte- and row-limited tail writer. Added a regression test with malformed history outside that window and updated the script header. Outcome tests and shellcheck passed; the full session-start suite timed out after 240 seconds
…enguid#6033) * fix(bin): read a quiet-mode record as a present captain, never hold-for-return Daemon-backed quiet mode writes the away-posture record marked mode: quiet, but the entry announcement, read-back, and session-start digest rendered it as "hold-for-return only", and the spend cap and PR merge gate treated it as away. A present captain's requested actions could then be held for a return that was not coming. bin/fm-afk-contract.sh now owns which posture a record is (the mode subcommand, fm_afk_contract_mode, fm_afk_contract_away_present). A quiet record announces, reads back, and appears in the digest as a present captain holding nothing; merges under it stay attended and it binds no spend cap. An away record is unchanged, an /afk entry over quiet mode rewrites the record as away, and a quiet entry never turns a standing away record quiet. * no-mistakes(document): Clarify quiet-mode authority and remove stale away guidance * no-mistakes(ci): The CI failure came from a race in the supervision-host test: its restart fixture could observe a watcher left by the preceding cycle. The test now retires that watcher and waits for the fixture arm to report its own started cycle. The focused test passed three times; the full suite was attempted but stopped at a separate intermittent test failure * no-mistakes(ci): Fixed daemon refresh mode selection so an unset-mode refresh follows the posture record: /afk over a running quiet daemon now changes state/.afk to away, while a plain quiet refresh stays quiet. Added script-level regression coverage for start and start-native and corrected a quiet-refresh fixture. The launch test suite, syntax checks, and diff check passed * no-mistakes(ci): Herdr was blocked before tests ran by a GitHub HTTP 500 downloading pinned Treehouse; no code change was warranted for that check. Fixed the Lint 1 ShellCheck warning in tests/fm-afk-launch.test.sh by annotating the intentional background PID capture. The focused test suite, ShellCheck, syntax check, and diff check passed
…merge (kunchenguid#6053) * fix(bin): accept a task's next PR once fm-pr-merge confirms the bound one merged require_recorded_pr_identity now checks fm_pr_poll_merge_already_notified for the recorded pr= before refusing a different URL, so a task's later PR is accepted once its earlier PR's merge is confirmed, while it keeps refusing while the bound PR is still unmerged. * no-mistakes(document): docs(fm-pr-merge): note next-PR accepted after bound PR merges
…6064) * fix(bin): read a live quiet record as a present captain at the host and watcher A quiet record left without its daemon (a quiet start that never ran or was interrupted) was read as away by the supervision host, so it parked a present captain's main and held captain outcomes for a return that never comes, and the watcher and daemon silenced captain-held rechecks on record presence. The host's posture checks, the watcher's and daemon's captain-held silencing, and the host's outcome path (branch report, drain BRANCH OUTCOMES, relocated branch authority, the owners' away wake note, and the Codex checkpoint bound) now ask the record owner's away-or-quiet reading, so only an away record is away. A live away record keeps today's behavior. * no-mistakes(document): Correct quiet-record documentation and supervision guidance * no-mistakes(document): Clarify quiet-record posture and captain-held rechecks * no-mistakes(document): Clarify quiet-record posture in documentation
…kunchenguid#6043) * fix(bin): name an in-window engine latch in the return brief and drop the false handling GAP line The away return brief said nothing had failed after the supervision host latched on engine errors during the window, and printed a GAP: watcher downtime line whenever a wake was merely being handled or queued at return. The failures section now reads the host ledger and latch record and names the latch time, the window's engine-error count, and whether the session is still paused or recovered. An open recovery episode is reported as information, and as a gap only when a queued episode outlived the return grace or the marker cannot be read. * no-mistakes(review): Fix latch trip time, drop marker-age grace, bound error count * no-mistakes(review): Report paused latch without ledger trip row; bound errors * no-mistakes(review): Never report a failed probe's latch row as trip time * no-mistakes(review): Only a retained trip row marks a pre-window latch * no-mistakes(document): Clarify return-brief latch and watcher-gap documentation * no-mistakes(ci): Fixed Lint 1 by marking the shared cooldown constant as used by sourcing scripts. The repository lint command and diff check pass; the return test run was stopped by a 180-second timeout after its completed cases passed * no-mistakes(ci): Fixed the return brief so the trip time and error count come from the same initial latch row, and ledger rows before the current session’s lock boundary cannot affect its latch report. Added real-script regressions for both findings. The return test suite, repository lint, and diff check pass * no-mistakes(ci): Fixed the return brief’s restart cutoff so it retains in-window failures, prints one line per initial-trip row, and omits zero-error count wording. Added real-script restart regressions. The return test suite, ShellCheck, and diff check pass * no-mistakes(ci): Fixed the return brief so a recorded trip followed by recovery stays recovered, while a later pause with a lost trip append gets a separate “trip time unavailable” line. Added a real-script regression that failed before the fix. The return test suite, ShellCheck, syntax checks, and diff check pass * no-mistakes(ci): Fixed the false second latch during recovery. A real-script regression failed before the fix and passes now; the lost-second-trip test still passes. The return test suite, ShellCheck, syntax checks, and diff check pass
* fix(calm): name the Claude Code Calm plugin fm so supervision notes read "fm: " Claude Code labels every mod transcript line with the plugin name, so the notes rendered as "firstmate-calm: ⚓ ...". Rename the plugin to fm, update the live guard to assert the fm: label, and document the one-time replay for sessions resumed across the rename. * no-mistakes(document): Clarify Calm plugin rename in documentation
…#6037) * feat(bin): add fm-live-lab.sh, a one-command live supervision lab builder * fix(bin): exact lab windows, per-lab task ids, self-safe teardown * fix(bin): target lab windows by id, stop lab descendants, add readiness tests * fix(bin): keep Claude's auto-updater off in live labs; list fm-live-lab.sh * fix(bin): start the lab tmux server without user config * no-mistakes(review): Scope lab teardown to its store, root, and task ids * no-mistakes(review): Record selected user stores at up for check and down * no-mistakes(document): Clarify live lab documentation and remove stale narratives * no-mistakes(ci): Fixed the CI failure by checking for an existing lab root before looking up the harness executable. The affected behavioral test and shell syntax check pass; the refusal also works with Claude absent from PATH * no-mistakes(ci): Fixed all four Greptile findings: teardown signals only recorded lab processes and their descendants; the worker gate is in its granted task directory and its path is exposed; readiness uses current crew state; and mate and worker IDs use 12 nonce hex digits. The CLI behavior tests pass, as do shell syntax, ShellCheck, and diff checks. The Claude no-host path is unchanged * no-mistakes(ci): Fixed the CI test’s dependence on an installed Claude binary by supplying a test-local stub. The full fm-live-lab test, shell syntax check, and diff check pass * no-mistakes(ci): Fixed all three selected findings in bin/fm-live-lab.sh: down waits for recorded processes and escalates before cleanup, PID roots are checked against recorded start times, and Claude primary trust is rechecked after mate/worker readiness. Added behavioral tests in tests/fm-live-lab.test.sh. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass * no-mistakes(ci): Fixed the pre-primary settle wait, worker gate instructions, unused retry variable, and teardown PID revalidation in bin/fm-live-lab.sh. Added behavioral tests in tests/fm-live-lab.test.sh. Both requested commands pass: tests/fm-live-lab.test.sh and bin/fm-lint.sh * no-mistakes(ci): Fixed teardown to track pre-kill lab processes by PID and start time, including children orphaned when a root exits. Up now rejects an empty pane PID before calling ps. Added regression tests and a Linux-safe worker fixture. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass * no-mistakes(ci): Fixed teardown tracking for children spawned during shutdown and made the worker fixture verify its exact window with a Linux-available shell. Both requested checks pass. The lab test takes about 66 seconds locally, so the under-one-minute target remains unmet * no-mistakes(ci): Fixed ci-2 and ci-4 in bin/fm-live-lab.sh and tests/fm-live-lab.test.sh. Teardown now tracks identity-checked members of captured lab process groups, including children orphaned during shutdown, without signaling the caller’s group or unrelated processes. Lint passed, and the lab test passed four times * no-mistakes(ci): Fixed teardown so an observed-empty process group is permanently dropped, preventing a reused group ID from signalling unrelated work. Added a ps-shim regression test. The lab test, lint, and diff checks pass * no-mistakes(ci): Fixed ci-1 in bin/fm-live-lab.sh and tests/fm-live-lab.test.sh. The TERM-born-child fixture now waits until its handler is installed before calling down. Down sends SIGKILL to identity-valid survivors on every pass from pass 20 onward and includes survivor process details if it must refuse cleanup. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass locally; Linux CI remains to be verified * no-mistakes(ci): Fixed down’s teardown wait to require two empty identity-checked scans separated by 0.5 seconds, and removed the unused test loop variable without changing the TERM-born-child test. The lab test, lint, and diff check pass locally
…unchenguid#6103) * fix(bin): keep slow watcher cycles and preempted reply polls from breaking supervision - fm_pending_reply_tick selects the records it has work for in one awk pass, so settled records cost no lock or fork and the walk no longer grows with the never-pruned store. - An attached arm keeps following a live, identity-matched holder whose beacon went stale until the lock changes or the shared stall bound (fm_watcher_stall_bound), then fails with a typed stalled-holder line so the retry replaces the holder. - The remote-reply adapter reports the job worker's preemption (exit 76) as a closed window, so the listener keeps its claim and polls again instead of being relaunched every watcher cycle. * no-mistakes(document): Clarify watcher grace and attached-arm documentation
…geable is UNKNOWN (kunchenguid#6110) * fix(bin): retry a bounded number of times when GitHub mergeable is UNKNOWN Fixes kunchenguid#6020 bin/fm-pr-merge.sh refused a GitHub merge whenever the pull request's mergeable field was not literally MERGEABLE. GitHub reports UNKNOWN for a short while after a push or a base-branch change while it recomputes mergeability, so a green, conflict-free pull request was refused as if it could not be merged. github_verify_mergeable now returns a distinct status when mergeable is the only failing condition and reads UNKNOWN. The caller retries up to 5 times, 3 seconds apart (overridable in tests), re-reading and re-checking every live condition on each attempt. Once the bound is spent it reports mergeability as still being computed rather than unmergeable, with the same nonzero exit as before. Every other refusal (closed, draft, conflicting, red or missing checks, away authority, queue protection) is unchanged and never retried. * no-mistakes(ci): I fixed both review findings the way you asked. The full suite (`bash tests/fm-pr-merge.test.sh`) ran to completion. Its last lines showed all `ok`, and any failure would have stopped the run early. I watched the output through `tail`, so I didn't see the new test's own `ok` line directly. **ci-2 (`bin/fm-pr-merge.sh`), retry delay not validated.** What must hold: the retry wait is always a short, valid `sleep` argument, so a bad `FM_PR_GITHUB_MERGEABLE_RETRY_DELAY` can never trip `set -e` or hold the task lock for a long time. The retry loop is the only place that reads this variable. The script now reads the value once before the loop and accepts only whole numbers from 0 to 10. Anything else (empty, `abc`, `-1`, `1.5`, `11`, a huge number, leading spaces) falls back to 3. I ran those values through the check by hand and each came out as expected. The loop now sleeps on that checked value. **ci-1 (`tests/fm-pr-merge.test.sh`), no test for a check changing between UNKNOWN reads.** What must hold: every retry re-checks all live conditions, not just mergeable. The fake `gh pr view` in the test can now take an optional second word on each line of the mergeable sequence, which sets the first check's result. The new test `test_github_mergeable_unknown_retry_rechecks_checks` feeds `UNKNOWN`, then `UNKNOWN FAILURE`. It asserts: - exit code 1 after exactly 2 reads, - the refusal names `check 'ci' is not green`, - the message does not say mergeability is still being computed, - `pr merge` was never called. If a later change made the retry look only at mergeable, the loop would read UNKNOWN 5 times, end with the "still being computed" message, and this test would fail. I didn't run it against a deliberately broken script to confirm that. `bash -n` passes. `shellcheck` reports only the existing info-level notes about files it can't follow. Only `bin/fm-pr-merge.sh` and `tests/fm-pr-merge.test.sh` changed
kunchenguid#6112) * fix(bin): converge every open owner onto a known terminal contribution settle_final only cleared a stale error on retry, so an owner whose saved row still said open kept projecting a merged or closed pull request as open after another owner's row had already recorded the terminal observation. Copy the known terminal observation to every owner whose saved row is not itself terminal, keeping that owner's own pending and notified state, and clear its error. * no-mistakes(review): Carry terminal checked_at when converging existing owner rows * no-mistakes(ci): I fixed Greptile finding ci-2 as you asked, with a change to tests/fm-contributions.test.sh only. The rule it enforces: when a retry converges an owner onto a URL that is already merged or closed, that owner gets the terminal owner's whole observation, not just its state. The same weak check appeared twice in test_interrupted_multi_owner_poll_settles_every_owner, so I fixed both: - **Open owner (line 784):** the check now also requires `.observation == $terminal[0].records[0].observation`. The existing checks for error, checked_at, pending and notified are unchanged. - **Errored owner (just below):** it only checked state and error before. It now reads the terminal owner's file and makes the same full-observation comparison. Adding the comparison alone would not have caught anything. The test fixtures gave both owners identical observations apart from `state`, so copying only the state would still have passed. In both cases I also set the terminal owner's observation head to HEAD_B, so the two observations now really differ. Verification: - The focused test passes against the current bin/fm-contributions.sh. - I temporarily changed `settle_final` so it copied only the state. The test then failed, reporting the owner still on the old head (HEAD_A). I restored the file afterwards, and `git status` shows only the test file modified. - The full tests/fm-contributions.test.sh suite exits 0. No product code changed. The other CI finding (ci-1, "Behavior portable serial 9") was left alone because you chose to ignore it
…nguid#6124) * feat: run the supervision host by default on a Claude primary An absent config/supervision-host on a Claude primary now reads as on with the default engine, and a file holding `off` opts any home out. Cursor, OpenCode, omp, Grok, and Codex stay file-gated, with `off` read as disabled there too. Every reader asks fm_supervision_host_enabled instead of testing the file, and non-bash readers query it through the lib's `enabled` entry. A primary's `off` is not inherited by secondmates: each home keeps its own supervision posture. * test: pin the watcher-path posture in fixtures that assume no supervision host Fixtures that drive the watcher arm or assert a non-host drain now write an explicit off file, and fixtures that copy the Stop auto-arm or the supervision instructions carry the engine lib they now source. The two drain suites also stop reading the code root's config. * fix: name the opt-out when an off home passes an attended wake to main A host parked when the home writes off now logs that the home does not run the supervision host, rather than claiming it has no engine. * no-mistakes(document): Clarify Claude supervision defaults and historical evidence * no-mistakes(ci): Fixed process leaks in the two added host tests. Each case now stops its recorded watcher and host/arm processes; fake hook sessions exit through session.stop. The full host suite passed before the final cleanup refinement, and both affected cases, bash syntax, ShellCheck, and diff checks passed afterward. CI runtime still needs confirmation
…start scope check (kunchenguid#6125) * fix(bin): create the state dir on a fresh primary before the session-start scope check fm_primary_scope_matches required an already-existing state directory, so bin/fm-sessionstart-run.sh stood down on a fresh clone before anything could create it. Split out fm_primary_root_matches so the run wrapper can confirm primary-home identity first, create the gitignored state dir when it is missing, and only then run the unchanged scope check. * no-mistakes(document): Document session-start state dir creation on fresh clones * no-mistakes(ci): I fixed the Greptile P1 the way you asked. When a fresh primary can't create `state/`, the run wrapper no longer stands down silently. **Invariant:** when an otherwise eligible fresh primary cannot create `state/`, startup must never fail silently. This path has only one site: the mkdir in `bin/fm-sessionstart-run.sh`. Other hooks and the nudge wrapper never create `state/`, so they have no equivalent failure. **What changed:** - **Run wrapper** (`bin/fm-sessionstart-run.sh`): it captures mkdir's error and prints one line to stderr before standing down as before (exit 0, or 3 for the Pi prerequisite). The line looks like `fm-sessionstart-run: startup could not create the state directory <path>: <reason>`. - **Test** (`tests/fm-sessionstart-nudge.test.sh`): the new case `test_run_reports_a_state_dir_it_cannot_create` uses a fresh primary with no `state/` and a read-only (0500) root. It checks four things: exit 0, no digest on stdout, no state dir created, and exactly one stderr line ending in "Permission denied". It fails without the fix and passes with it. - **Docs** (`docs/sessionstart-nudge.md`): I added one sentence describing the stderr line and one describing what the new test proves. **Verification:** I ran `tests/fm-sessionstart-nudge.test.sh`, and every test passes. `bin/fm-lint.sh` on the changed scripts (pinned ShellCheck 0.11.0) and `tests/fm-documentation-audiences.test.sh` also pass. As you asked, the wrapper still stands down with the ineligible-checkout status afterwards. It does not report this as a failed eligible startup, which is what the bot suggested
…ery (kunchenguid#6126) * fix(bin): measure pending-reply grace from turn completion, not delivery Fixes kunchenguid#6057 The pending-reply guard demanded a repost ("REPOST REQUIRED: previous marked request had no correlated parent report") while the second mate's correlated reply was already on its way. fm_pending_reply_send_recovery measured its grace window from delivery instead of from the request turn's completion, so any turn longer than the grace fired the demand the moment the turn ended, before the reply could have landed. The missed-report escalation had the same gap: it fired the instant the recovery turn's completion was observed, with no grace at all. Both now measure grace from the relevant turn's completion (request turn for the recovery repost, recovery turn for the escalation), and both take one fresh, uncached read of the parent status file immediately before firing, accepting a correlated line regardless of its verb. Transport-failure escalations stay immediate, and the one-repost limit is unchanged. * no-mistakes(review): Document grace window as measured from turn completion * no-mistakes(ci): Both Greptile findings were real and caused by this PR, so I fixed them. The full `tests/fm-pending-reply.test.sh` suite passes. **ci-1 (a reply could be overwritten by a repost).** The rule that must hold: a recovery send is recorded only if the record is still unresolved, checked under the same per-correlation lock that resolution uses. The escalation path already did this (`_fm_pending_reply_maybe_escalate_locked` reads fresh and publishes under one lock). The recovery path did not: `fm_pending_reply_send_recovery` did its fresh read through `fm_pending_reply_try_resolve`, which let go of the lock before the send was recorded. A reply landing in that gap could be overwritten, and the repost would go out anyway. Now `send_recovery` takes the lock once and, while holding it, re-checks that the phase is still `awaiting_report`, runs the fresh uncached read, and records the send (sender pid and identity, attempt time, phase `recovery_sending`). It releases the lock before actually sending, so the lock is not held during the send. It uses the same lock helpers the other lock wrappers use. Grace timing, the one-repost limit and the escalation path are unchanged. **ci-2 (the test would pass even without the fix).** In `test_recovery_fresh_status_read_resolves_before_firing`, the reply is still appended to the status file, but the stored file signature is then set to the file's new signature. That stands in for a same-size rewrite that the signature cache cannot see. The test first checks that a normal cached read misses the reply, then that the fresh read before sending catches it. I also added the same check for the fresh read before escalation, which the review said was uncovered. The test now sets its own send hook, so it no longer depends on one left over from an earlier test (that leftover had made failures exit silently). **Checks:** - I removed the fresh-read bypass at each site in turn and reran the suite. With it gone from recovery, the test fails with "recovery must not fire once a correlated reply has landed". With it gone from escalation, it fails with "the fresh pre-escalation read should have resolved the record, got escalated". With both in place, all tests pass. - Shellcheck with `-x` timed out locally. Without `-x` and ignoring SC1091, the only warnings are SC2034 on the existing `maybe_escalate` lock wrapper, which is not part of this change. The new code adds no warnings. Changes are in `bin/fm-pending-reply-lib.sh` and `tests/fm-pending-reply.test.sh`. Nothing is committed yet; a plain commit message such as "fix(bin): record the pending-reply recovery send under the fresh-read lock" fits the instruction * no-mistakes(ci): ci-1 was real and caused by this PR. The same bug was also in the escalation path, so both are fixed. The full tests/fm-pending-reply.test.sh suite passes. The rule that must hold: a recovery repost or an escalation goes out only if the record's phase, read after the fresh-read resolve, is still what it was before. The resolver writes phase=resolved first and only then writes the other resolution fields. If one of those later writes fails, it returns an error even though the record is already resolved. Places this rule applies, both fixed: - Recovery (fm_pending_reply_send_recovery): the fresh-read resolve now runs first, and the phase is re-read right after it, whatever it returned. The send is recorded and made only if the phase is still exactly awaiting_report. This replaces the earlier phase check rather than adding a second one. - Escalation (_fm_pending_reply_maybe_escalate_locked): same bug. After a failed resolve it went on to publish the blocked line and set phase=escalated. One added line after the resolve call returns 1 without publishing if the phase has changed. Test: added test_partial_resolve_write_blocks_firing. It forces a failure on the resolved_epoch write after a correlated reply has landed. It checks that the recovery send hook is never called, that no escalation line is published, and that the phase stays resolved. The forced failure runs in a subshell so it can't affect later tests. Checks: - With the recovery fix reverted, the new test fails with "recovery must not fire after a partial resolve". - With the escalation fix reverted, it fails with "partial resolve should block escalation, got escalated". - With both fixes in, every test passes. - Shellcheck was run with SC1091 excluded and without -x, not through the repo's lint script. The only new message is one SC2329 info on the test's override function; other test overrides in the same file already get that same info, unsuppressed. Changed files: bin/fm-pending-reply-lib.sh and tests/fm-pending-reply.test.sh. Nothing is committed. Suggested plain commit message: "fix(bin): recheck pending-reply phase after the fresh read before sending
…nguid#6179) * fix(tests): cut the fixed sleeps in supervision-host cycles The serial CI lane keeps brushing its 30-minute cap because fm-supervision-host.test.sh spends ~903s of the job, and per the run-36635306527 case profile the top nine cases are all multi-cycle ones (3-10 park/close/turn cycles each): every close waits out the host's sleep $POLL in await_close plus a watcher sleep $FM_POLL scan cycle, and every engine turn waits out the fixed sleep 1 descendant snapshot. That is ~3s of pure sleep per cycle before any real work. The host poll now accepts positive decimal seconds through a new seconds_or validator (FM_SUPERVISION_HOST_POLL), and the engine turn's snapshot loop takes FM_SUPERVISION_ENGINE_SNAPSHOT_SECONDS, also a positive decimal defaulting to one second - the smallest seam at each wait's single owner. The suite drives them at 0.2 alongside the existing FM_POLL=0.5 and FM_ARM_ATTACH_POLL=0.2 knobs, so the real poll loops still run. The park-boundary case moves onto the injected test clock instead of a real 3s wait, per-case cleanup polls the host pid rather than sleeping a full second, and the proof-by-absence windows (flood re-escalation, successor re-announce, watcher persistence, recovery staying off main) shrink from 2-3s to 1s, which still spans two watcher polls at the test cadence. Every assertion, process lifecycle, and reaping path is unchanged; production defaults stay at one second. Isolated case timings on a contended host, base vs branch: attended-latch 54.3->34.6s, undelivered-dialog 67.7->59.1s, away-latch 46.5->30.5s, held-cadence 47.9->21.6s, unreadable-mirror 39.2->38.5s, park-limit 18.2->12.3s, registration-fallback 14.1->10.0s, first-cycle-status 12.6->8.4s, latch-scope 16.7->16.3s. Full suite: 65/65 pass. fm-lint and shellcheck clean. * no-mistakes(review): Wait for scan lock release before duplicate check * no-mistakes(document): Correct supervision snapshot cadence documentation * fix(tests): keep production poll cadence, probe exits at 0.1s The fractional poll cadences multiplied the cost of each loop body: full process-table scans in the engine turn and process refreshes in await_close ran five times more often, which swamped the thin CI runner and nearly doubled every multi-cycle case (serial 5 was cancelled at its 30-minute limit on run 36635306527's successor). Restore the production cadence and notice arm/engine exits with a cheap kill -0 probe at a tenth of a second between the one-second bodies instead: strictly less dead time than baseline with no added CPU. Also hold each injected-clock park bound well past its case's wall-clock checks so a host that ignored the test clock fails instead of silently passing at a real-time boundary, and restore the shortened proof windows (watcher liveness, recovery-off-main absence, first-cycle stream) to their baseline depth. * no-mistakes(document): Clarify supervision engine snapshot documentation
…henguid#6192) * fix: rebalance portable CI from current duration measurements * no-mistakes(test): Test serial packing boundary and verify endpoint timeout cleanup * no-mistakes(document): Clarify timeout guidance and remove duplicated packing estimates
…nguid#6216) * fix(bin): run no repository hook when core.hooksPath is empty The per-task hook wrapper refused every commit in a repository whose own config sets core.hooksPath to the empty string, because git rev-parse --git-path hooks fails on it. Plain git reads that setting as no hooks, so the wrapper now runs none; every other lookup failure still refuses and shows git's error. Fixes kunchenguid#6171 * no-mistakes(review): Refuse commits when core.hooksPath is a valueless key * no-mistakes(document): Document empty core.hooksPath handling in commit attribution docs * no-mistakes(ci): When the wrapper refuses a commit, Git's hook-lookup error now shows up once instead of twice. That required changing one line in the wrapper, and the tests were extended so both bad-config cases would catch the duplicate. Invariant: when the wrapper refuses, Git's lookup error must appear exactly once. In the failure path, the only Git call besides the deliberate second lookup is the `git config --get --type=path core.hooksPath` check in `runtime_chain_body` (`bin/fm-git-strip-ai-trailers.sh:168`). That check prints the same error, so it was the one place to fix. I added `2>/dev/null` to it. Its exit status still decides the outcome: an empty value still runs no hook, and anything else goes on to the second lookup, which prints Git's error once, and the commit is refused. Tests (`tests/fm-git-strip-ai-trailers.test.sh`): - The unresolvable-path test (`~fm-no-such-user-6171/hooks`) now requires `failed to expand user dir` to appear exactly once in the refused commit's output. - The valueless-key test now requires `missing value for 'core.hookspath'` to appear exactly once. - Pre-existing bug in the unresolvable-path test: its `git add` ran after the bad config was set, so it failed silently (exit 128) and the "refused commit" had nothing staged. The test now stages the file before writing the config, the same way the valueless test does, so a real commit gets refused. - The empty-string test is unchanged and still passes, so an empty `core.hooksPath` still runs no hook. Verification: - With the wrapper change reverted, both new checks fail with `expected '1', got '2'`. With the change in place, the whole suite passes. - `bash -n` passes. shellcheck shows only an info-level SC1091 note about sourcing `lib.sh`, which was already there before this change. - `git status` lists only the two intended files
…kunchenguid#6213) * fix(bin): let a stale record on a reassigned slot retire records-only When a pool slot's owner claim names another task, the stale record's teardown touches nothing under the slot, so the exclusive-slot record scan no longer refuses it. Full teardowns of a slot this task still claims, or one with no claim, keep the refusal. Fixes kunchenguid#6184 * no-mistakes(document): Note claim-over-record precedence for reassigned teardown slots
…uid#6240) * fix(bin): keep the steering doorbell short under deep homes The doorbell printed the task inbox's absolute path twice, so under a deep home it grew to about 290 characters and a Herdr submit reported it never reached the pane on every re-ring. It now names the inbox once by its short <task>.inbox name and points at the full path the worker's brief already gives, so its length no longer depends on the home's depth. Fixes kunchenguid#6120 * no-mistakes(review): Export FM_TASK_INBOX at launch and name it in doorbell * no-mistakes(ci): ci-1 (Behavior portable serial 9) was caused by this PR, and I fixed it in the test. tests/fm-claude-trust.test.sh failed with "the launch command did not carry a brief doorbell". Its claude_launch_doorbell helper stripped exactly two leading `export ...;` statements before reading the final prompt argument. This PR adds a third one (`export FM_TASK_INBOX=...`) to every launch, so the helper was reading the wrong command. The invariant: a test that parses the launch command must skip every leading export statement, however many there are. I checked every test that parses the launch this way. The only other ones are the two helpers in tests/fm-spawn-dispatch-profile.test.sh, and they already loop over all exports. The kimi and dispatch-profile exact-string checks were updated earlier in this PR. The fix makes claude_launch_doorbell use the same loop (`while [[ "$command" == export\ *\;* ]]; do command=${command#*; }; done`) and then take the last argument. The ordinary path still works: the claude spawn test and the secondmate-clone spawn test both resolve the brief record through the same helper. Verified locally: `bash tests/fm-claude-trust.test.sh` exits 0 with no failing cases. ci-2 (Behavior tests (Herdr)) was not caused by this change, and I made no code change for it. In tests/fm-backend-herdr-presentation-e2e.test.sh, the concurrent secondmate recovery failed with "herdr presentation recovery could not acquire its session lock; refusing a concurrent resume". Two reasons it is not this PR: - The same failure, in the same test and case, happened on run 36655209015 for the unrelated branch fm/fm-contributions-old-gh-compat about 14 hours earlier. - This PR's change cannot lengthen how long the lock is held. The launch is written to a file and sent to the pane as `. launch.N.sh`, so the extra export changes neither the pane submit nor the lock hold time. The cause is a race that was already there: spawn_herdr_presentation_order_lock_acquire gives up after 5 seconds, and a concurrent real-Herdr recovery can hold the lock longer. Fixing that means changing the product's lock timeout, which is outside this PR. It should be tracked separately, and a rerun of the Herdr job is expected to pass. The only file changed is tests/fm-claude-trust.test.sh
… no turns (kunchenguid#4859) * fix(dod): drive no-mistakes with one foreground call, not a background poll The brief told workers to background the drive call and poll `axi status` because one call "routinely outlives what your harness lets a single command run". That advice contradicts the tool it drives: `no-mistakes axi run --help` documents `--wait` with an 8m default, existing precisely "so an agent harness with a 10-minute tool cap gets a structured return instead of an unbounded hang". Following the old text, a worker could never idle - a backgrounded call returns in milliseconds, so it does not wait at all - and each attempt leaked a live timer that later fired as a paid wake. Tell workers to make one foreground call, let it block, and repeat it when it returns on elapsed wait rather than on a gate or outcome. Also drops the generalisation that told workers on any unestablished harness to assume a command cap and use the same shape, which exported the defect to harnesses with no such cap. * fix(bin): let a waiting worker spend no turns until it is answered A worker waiting on a decision, a pipeline gate, CI, or a heavy-test slot kept taking model turns: the brief told it to list its inbox at any natural checkpoint, and six automatic senders nudged secondmates whatever their open decisions. - The ship and scout briefs gain one Waiting section: end the turn after needs-decision or blocked, and hold an external wait inside ONE blocking command bounded by the harness's own command ceiling. The checkpoint clause is deleted. Forbidding the wrong shapes is not enough on its own, so the section also names the blocking foreground `until` loop as the wait a Claude Code worker may use, because that harness can refuse a sleep-then-check command while pointing at backgrounding, which is the one shape a waiting worker must not take. - fm-send --automatic defers (exit 4, nothing written or rung) while the target has an open decision or blocker of its own; every automatic sender passes it and keeps its retry state, and the pending-reply recovery waits the same way. - The two senders that report the result classified it by matching the text of the send's captured output against `deferred:*`. fm-send runs bin/fm-guard.sh as a supervision warning, and that guard prints its worktree-tangle banner whenever the primary checkout is on a feature branch, which is exactly what a CI pull-request checkout is. The banner lands ahead of the `deferred:` line, so the match fell through and a waiting mate was reported as a failed send, with the banner as the reason. Both senders now classify on fm-send's exit status, which is the contract the deferral is actually stated in, and select the `deferred:` line out of the output rather than assuming it came first. The third root cause, a no-mistakes definition of done that backgrounded the drive call and polled axi status, is fixed by this branch's parent commit "drive no-mistakes with one foreground call, not a background poll"; this commit takes that text as is and adds the regression test. Upstream's spawn abort path no longer calls the lease-return helper at all, so the fork's missing-helper guard and its pin-feature test line are moot here and are not ported. The command ceilings each harness enforces, and the probes behind the named Claude Code wait, are recorded in docs/verification/runtime-backends.md. * no-mistakes(review): Exempt captain holds, quiet deferred reconcile, clarify worker pauses * no-mistakes(document): Document deferred automatic nudges, rereads, and reply recovery * no-mistakes(document): Ring unlanded fire-and-forget steers exactly once more * no-mistakes(ci): The failing check, "PR must be raised via no-mistakes", reads the pipeline's attestation record, which says document=skipped. No file in the repository can change that record, so I did not touch the check or the PR body. As you said, the no-mistakes rerun after this run finishes will re-execute the document step and record document=completed. The one change is the documentation sentence you ordered. It adds a line to docs/remote-secondmates.md, right after the line saying the remote host runs no re-ring ladder of its own: "A fire-and-forget record, such as a reconcile ask, gets its single retry ring only on the local plane: the remote steer leg owes no re-ring, so a swallowed remote doorbell for one waits for the next ring into that inbox, and a remote-side retry is known follow-up scope." No behavior changed. Checks: tests/fm-documentation-audiences.test.sh passes (4/4) and bin/fm-lint.sh is clean. The change is left uncommitted in the working tree for the pipeline to pick up * no-mistakes(review): Hold automatic wakes until a mate's own decision closes * no-mistakes(document): Document watcher delivery of deferred remote re-read nudges * no-mistakes(review): Merge duplicate elapsed-wait reattach instructions in DOD * no-mistakes(test): Resolve merged default decision in remote-reply recovery fixture * no-mistakes(test): Source classify lib so config-push retry-deferred honors open decisions * no-mistakes(ci): Fixed a flaky test that also fails on main. Neither this PR's bin/fm-brief.sh nor its bin/fm-dod-lib.sh change is involved: bin/fm-dispatch-resolve.sh sources neither file. Another branch (fm-attended-cutover-smoothing-s1, run 36343879084) failed the same shard 8 check the same way, on a different case ("a rule-criterion match prints one diagnostic line, got 2"). Root cause: `fm_quota_single_provider_for_harness` in bin/fm-quota-axi-lib.sh returned from its `while read` loop as soon as it found a match. That closed the pipe while `fm_quota_single_provider_table`'s `printf` was sometimes still writing. GitHub Actions runners ignore SIGPIPE, so bash printed `fm-quota-axi-lib.sh: line 138: printf: write error: Broken pipe` to the resolver's stderr. That is the extra line. I reproduced it locally by running the test with SIGPIPE ignored: 2 of 20 runs failed, one with the resolver's diagnostic line plus two broken-pipe lines. Invariant: looking up a harness in the provider table must never make the table writer fail. The only reader of that table is this function, and all of the resolver's lookups (line 208 without stderr redirected, line 222 with it) go through it. So the fix is in that one place: read the whole table, then print the match. The same file now shows it reads the full table first, like `fm_control_harness_supported` does. Return values and output are unchanged. Verification: with SIGPIPE ignored, tests/fm-dispatch-resolve.test.sh failed 0 of 30 runs after the fix (2 of 20 before). tests/fm-dispatch-resolve.test.sh, tests/fm-brief.test.sh, tests/fm-send-inbox.test.sh, tests/fm-quota-choose.test.sh and tests/fm-quota-array-dispatch-live-e2e.test.sh all pass, and shellcheck is clean. tests/fm-procevent-quota.test.sh fails locally with or without the change ("process-event state root is not a private directory"), so that failure comes from the local environment, not from this fix. No new test was added: the existing one-diagnostic-line assertions already catch this whenever SIGPIPE is ignored, as it is in CI * Revert "no-mistakes(ci): Fixed a flaky test that also fails on main. Neither this PR's bin/fm-brief.sh nor its bin/fm-dod-lib.sh change is involved: bin/fm-dispatch-resolve.sh sources neither file. Another branch (fm-attended-cutover-smoothing-s1, run 36343879084) failed the same shard 8 check the same way, on a different case ("a rule-criterion match prints one diagnostic line, got 2"). Root cause: `fm_quota_single_provider_for_harness` in bin/fm-quota-axi-lib.sh returned from its `while read` loop as soon as it found a match. That closed the pipe while `fm_quota_single_provider_table`'s `printf` was sometimes still writing. GitHub Actions runners ignore SIGPIPE, so bash printed `fm-quota-axi-lib.sh: line 138: printf: write error: Broken pipe` to the resolver's stderr. That is the extra line. I reproduced it locally by running the test with SIGPIPE ignored: 2 of 20 runs failed, one with the resolver's diagnostic line plus two broken-pipe lines. Invariant: looking up a harness in the provider table must never make the table writer fail. The only reader of that table is this function, and all of the resolver's lookups (line 208 without stderr redirected, line 222 with it) go through it. So the fix is in that one place: read the whole table, then print the match. The same file now shows it reads the full table first, like `fm_control_harness_supported` does. Return values and output are unchanged. Verification: with SIGPIPE ignored, tests/fm-dispatch-resolve.test.sh failed 0 of 30 runs after the fix (2 of 20 before). tests/fm-dispatch-resolve.test.sh, tests/fm-brief.test.sh, tests/fm-send-inbox.test.sh, tests/fm-quota-choose.test.sh and tests/fm-quota-array-dispatch-live-e2e.test.sh all pass, and shellcheck is clean. tests/fm-procevent-quota.test.sh fails locally with or without the change ("process-event state root is not a private directory"), so that failure comes from the local environment, not from this fix. No new test was added: the existing one-diagnostic-line assertions already catch this whenever SIGPIPE is ignored, as it is in CI" This reverts commit c719928. * no-mistakes(review): Retry deferred local instruction nudges via the watcher * no-mistakes(review): Document watcher retry for deferred local instruction nudges * no-mistakes(ci): I fixed both review findings you selected (ci-1 and ci-3). I did not touch the deferral check in bin/fm-send.sh. ci-1 (bin/fm-config-push.sh, retry_deferred_rereads) - Rule that must hold: a deferred reread stays flagged until it is actually delivered. - Before the fix, the flag was removed before any of the steps that can skip a mate: the remote lock-path lookup, validate_secondmate_home, the local lock-path lookup, and the lock acquire. A skip at any of those dropped the flag, so the watcher lost track of the reread. - Now the flag is removed in one place only, when the send succeeds (rc 0). A skipped home, a busy lock, a deferred send (rc 4) or a failed send all leave it in place. The re-mark calls on a busy lock and on rc 4 were no longer needed, so I removed them. I updated the comment above the function to match. - Side effect: a send that keeps failing now stays flagged, so the watcher retries it on every poll and logs each failure. That follows your "don't clear until delivered" rule, but it replaces the old behaviour of leaving a failed send to the next config push or session start. - New test in tests/fm-secondmate-sync.test.sh: T8j "a deferred flag survives a skipped invalid home and is retried once it validates". It takes the home's marker away to make validation fail, checks that nothing is sent and the flag stays, then puts the marker back and checks that the nudge is delivered and both the flag and the retry marker are cleared. It fails on the old code and passes now. ci-3 (bin/fm-secondmate-restart.sh) - Rule that must hold: no automatic send wakes a mate that is waiting on its own open decision. - The two automatic sends in this script are the fallback reread nudge (fall_back_to_nudge) and the persist request. Both now pass --automatic. If a persist request is deferred, its correlation is discarded and the mate goes to the fallback nudge, which is also deferred, so the mate is reported as unreached. - New test in tests/fm-secondmate-restart.test.sh: T3b. It gives a mate an open needs-decision and runs a restart. It checks that both sends report as deferred, the mate's doorbell is never rung, its inbox gets no message, nothing is stopped, and the mate is reported as unreached with exit status 3. It fails on the old code and passes now. - The test marks the watcher as alive first. Without that, the watcher-down warning is printed first and becomes the reported reason instead of the deferral message. Verification - tests/fm-secondmate-sync.test.sh passes. - tests/fm-secondmate-restart.test.sh passes. - tests/fm-secondmate-harness.test.sh (the other test that exercises --retry-deferred) passes. - The fm-send-inbox test that covers automatic deferral passes. I only looked at the last lines of that run, not the whole file. - `shellcheck -x` on the four changed files is clean * Pin autoarm supervision model in secondmate restart T3b The fresh watcher beat the test writes proves a live watcher only under the autoarm model; on CI hosts with no detected harness the persistent model demands a lock-holding watcher, so the watcher-down banner became the reported reason and the deferral assertion failed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Keep deferred secondmate nudges retryable under the inheritance lock. A bootstrap instruction nudge could write its deferral flag outside the lock the watcher retry holds, so a concurrent retry could delete a flag that had just been set. A restart fallback that is deferred now records the same marker and flag, so the watcher delivers it once the decision closes. * no-mistakes(document): Document watcher retry of deferred restart re-read nudges * Send secondmate reread and restart nudges immediately again. Deferring those nudges let a later config push drop an incomplete transfer once the decision closed. They now send as they do on main. * Make the no-turn wait opt-in behind config/wait-no-turns. Homes that do not create the file keep the previous briefs, drive text, and sends. * no-mistakes(document): Document wait-no-turns inbox wording change in configuration * no-mistakes(review): Keep checkpoint inbox check; forbid only polling while waiting * no-mistakes(ci): Fixed ci-2 (Greptile: a concurrent retry marker gets lost). The rule that was broken: the watcher may remove only the `.retry-ring` mark for the record it just processed. A newer mark written in the meantime is owed its own retry. `fm_task_inbox_clear_retry` is the one shared function that removes the mark, and I fixed it there. In `bin/fm-task-inbox-lib.sh` it now takes the record path. It compares the mark's content with that record's name and removes the mark only when they match. When the mark names a different record it returns success and leaves the mark alone. It still fails only when the processed record's own mark can't be removed. Both callers in `bin/fm-watch.sh` now pass `"$rec"`: the dead or missing pane path and the path after a retry ring. So the fix holds at both removal sites. Tests, in `tests/fm-task-inbox.test.sh`: - I added an optional `FM_RING_MARKS_RETRY` hook to the fake tmux. It writes a newer record's mark while the doorbell is being typed, which reproduces the race deterministically. - I added `test_watcher_retry_keeps_a_newer_mark`. The owed retry rings once, the newer mark survives, and a later check rings the newer record once and then clears its mark. The test fails without the fix ("the spent retry removed a newer record's mark written during its ring") and passes with it. - I updated the direct `clear_retry` call in the existing unit test to pass the record. Results: `tests/fm-task-inbox.test.sh` passes in full and `tests/fm-send-inbox.test.sh` passes 15/15. Shellcheck reports only SC1091 "not following sourced file" notices. As instructed, I didn't change the brief inbox wording * no-mistakes(document): Fix stale wait-no-turns inbox wording in inbox lib comment --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
…note (kunchenguid#6140) * fix(bin): record Gerrit change URLs as close notes Teardown's backlog_done_args hands every ship's recorded pr= URL to fm_backlog_done as --pr, and tasks-axi refuses any --pr that is not a canonical GitHub or Forgejo pull request. A Gerrit change URL therefore left the item In flight after cleanup, and the pending backlog-close record replayed into the same refusal at every session start. fm_backlog_done now rewrites a --pr whose value fm_pr_url_parse reads as a Gerrit change into --note "Gerrit change <url>". The mapping sits at the tasks-axi call rather than in the pending-close record, so records already written with --pr replay to a close unchanged. The captain-held retain path records the URL in its deliverable line and skips the update --pr it cannot make. * no-mistakes(review): Note retained Gerrit change URL when captain answers early * no-mistakes(document): Document Gerrit change URL handling in captain-hold retention
* perf(remote): separate active job sampling from dispatcher cadence * no-mistakes(document): Link remote wait timing to its authoritative contract * no-mistakes(ci): Fixed ci-1 with two narrowly scoped SC2030 annotations documenting intentional subshell-local legacy and active cadence overrides in tests/fm-remote-job.test.sh. Runtime behavior is unchanged. Reproduced the lint failure before the fix; afterward ShellCheck 0.11.0 with source following, Bash syntax validation, the complete remote-job behavior suite, and git diff --check all passed * perf(supervision): reduce park, delta and dispatcher polling * no-mistakes(document): Clarify poll latency contracts and authoritative documentation pointers
…henguid#6221) * fix(bin): load backend sibling libraries under zsh fm_backend_source kept each backend's sibling list in one space-separated string and iterated it unquoted. zsh does not word-split an unquoted expansion, so the readability check saw the whole list as one path and refused every backend with more than one sibling. Hold the list in the function's positional parameters instead, which needs no word splitting in Bash 3.2, Bash 5, or zsh. The existing zsh case in tests/fm-backend.test.sh covers it wherever zsh is installed. * test: run the Calm mod suite on stock Bash 3.2 The suite injected shell values into its generated Node scripts with the ${value@Q} transformation, which needs Bash 4.4. Stock macOS Bash 3.2 reports a bad substitution, so every case failed before it asserted anything. Build each JavaScript string literal with JSON.stringify through a small helper instead, which works on any Bash and is a valid literal for any value. * no-mistakes(review): fix(bin): rename zsh-special path local in fm_backend_source * test: narrow the zsh backend claim to name matching Under zsh the adapters locate their siblings through BASH_SOURCE, so a successful fm_backend_source is not a full load. Assert only what the contract states, and pass js_string values after -- so node never reads a leading-dash value as its own option. --------- Co-authored-by: Nova Agent B <novaagentb@gmail.com>
…tatus scans (kunchenguid#5263) * fix(bin): exclude a remote mate's own parent channel from self-home scans A remote secondmate home's outbound parent channel lives at state/parent-replies.status inside its own state dir, so the watcher's signal scan enumerated it as a task status file and the open-decisions fold classified it as a phantom task named parent-replies: every parent-channel append spun a spurious signal wake and a phantom open decision in the mate's own home. fm-parent-channel-lib.sh gains fm_parent_channel_outbound_status, which resolves the channel into the mate's own state dir for the remote route only, and fm-classify-lib.sh's status_scan_parent_channel_exclude wraps it for the fleet-wide scans. The watcher's scan_signals and heartbeat fail-safe backstop, the whole-file and incremental open-decisions folds, the presentation snapshot, and the unread-surface scan now skip exactly that resolved path. The exclusion is home-shape-aware: a parent-replies.status in a main home or a local mate is an ordinary task log and keeps waking and folding, and every other status file is untouched. * no-mistakes(review): exclude a remote mate's parent channel from the daemon heartbeat scan * no-mistakes(document): Document remote mate parent-channel scan exclusion * ci: retrigger portable serial 4 * no-mistakes(ci): CI check 'Behavior portable serial 7' failed in tests/fm-contributions.test.sh ('reservation poll failed'). CI stderr showed bin/fm-contributions.sh:345 arithmetic 'DEADLINE - 6\n90077104: syntax error in expression': the fixture's fake date returned a torn two-line clock value. Root cause: the fake forge wrapper in wrap_forge advances the shared controllable clock via a non-atomic read-modify-write ('$(cat $FORGE/clock) + 6' with truncate-in-place '> $FORGE/clock') while concurrent background gh calls run and the fake date reads the same file; an interleaved truncate+write publishes a half-written value (CI's torn '6\n90077104', tail of 1790077104) or an emptied-read value ('6'), which either breaks the poll's arithmetic (nonzero exit -> 'reservation poll failed') or defeats the 15-second reservation defer. This is a pre-existing test-fixture race, not caused by the PR's diff (base..target touches no contributions code; the same commit passed this shard in run 35711207830 earlier the same day). Fixed the flaky fixture at its root: clock_bump() now writes each new value to a per-process mktemp file in the same directory and publishes it with mv (atomic rename), so concurrent forge callers and the fake date always read one complete old-or-new clock; fault patterns and deltas are unchanged. Verified: minimal 3-way concurrency repro shows the old wrapper corrupting (12/32/38 outcomes incl. empty-read) while the rename-based wrapper never corrupts (20/20 clean); the full tests/fm-contributions.test.sh passes twice (all 38 assertions ok, incl. the reservation, budget-exhaustion, genuine-failure, shared-once, and latency tests); 10 isolated reservation runs pass; shellcheck rc=0; worktree contains only this one-file change * no-mistakes(document): drop stale file-set copy in daemon catch-all comment
…6307) * fix(bin): name the accepted verdict actors in fm-contributions help and refusal * fix(ci): Updated tests/fm-contributions.test.sh to assert exactly captain, fleet, maintainer, and nobody in command-emitted help and refusal output. Three focused regressions passed; all three extra-actor mutations were rejected. ShellCheck, syntax, and diff checks passed. Production code remains unchanged
…chenguid#6306) * fix(bin): recognise a clone root git names with different path spelling fm-fleet-sync compared git's --show-toplevel with pwd -P as strings, so a clone root that git recorded with different casing (case-insensitive volume) was skipped as not a clone root and never refreshed. Compare filesystem identity instead, which also covers symlink spelling. * fix(document): Remove stale clone-root comparison comment
) * test(calm): pin Pi's regular TUI mode where pane assertions read scrollback Pi 1.0.0 defaults its TUI to a fullscreen alternate-screen mode whose scrollable transcript is application-owned, so rows that leave the viewport never enter terminal scrollback and tmux capture-pane -S can no longer see them. The Pi Calm e2e launches now pass --tui-mode regular wherever the flag exists so the transcript assertions keep reading real scrollback on both the Pi 1.0.0 line and earlier Pi lines, which have no such flag and render regular-only anyway. * no-mistakes(document): Correct Pi TUI documentation and scrollback rationale
…ries (kunchenguid#6331) * fix(bin): encode captain-hold reasons and reject self-inventory in complete hold now stores a reason with parentheses, line breaks, or percent signs through a reversible percent encoding that every reader decodes, instead of refusing it. hold --origin records the origin on the held task, and complete refuses the origin as its own inventory entry and an entry held for a different origin; holds with no recorded origin are accepted and flagged. * fix(review): Decode marked hold reasons consistently across readers * fix(review): Remove unnecessary lifecycle test dispatch * fix(review): Correct hold origin identity and inventory recovery * fix(review): Record origins before placing backend holds * fix(document): Clarify captain-hold validation and reason reader documentation * fix(ci): Fixed both findings: failed backend holds restore the previous origin, and invalid base64/UTF-8 reasons remain verbatim. Added regressions and documented valid-literal ambiguity. Both failures were reproduced before fixes. Verification: 54 lifecycle tests and 9 wrapper tests passed; 7 Beads-specific cases skipped because tasks-axi is markdown-only. Focused lint and diff checks passed. No pipeline or publication actions performed
* fix(bin): take over the watcher cycle a main-only pass-through leaves An attended main-only pass-through leaves a successor watcher cycle running through main's handling turn. The session's next park attached to that cycle instead of owning it, so the successor's arm, orphaned by its host's exit, kept owning the watcher while the new park's arm polled it twice a second until the next close or the park boundary, hours later in a quiet second mate. A second-mate restart hit this every time, since its persist request is a main-only close. The host now records the successor it leaves for main, and the next host's first cycle runs bin/fm-watch-arm.sh --take-over on it: when that arm still owns the healthy watcher, the new arm stops it, reports a reason the cycle delivered first, and otherwise owns a fresh cycle. The stop's own downtime publication is undone over an acknowledged episode when no wake was appended in between, so the handover wakes nobody. * no-mistakes(review): Keep left-arm record until the orphaned arm is gone * no-mistakes(review): Relinquish successor arm only after durably recording it * no-mistakes(review): Relinquish successor only after its record reads back * no-mistakes(document): Clarify successor takeover guarantees and authoritative documentation * no-mistakes(ci): Fixed ci-1: acknowledgement restore now requires the exact taken-over arm/watcher ledger row with signal=TERM, awaited within a short bound. Otherwise takeover proceeds without erasing downtime. Added a self-exit regression confirmed failing before the fix and passing afterward; ordinary takeover tests and the full watcher-arm suite pass. Updated Generation reuse documentation. Syntax, diff checks, and ShellCheck pass with existing SC1091/SC2034 warnings excluded. ci-2 remains unchanged * no-mistakes(document): Clarify watcher take-over recovery and restart limits
…cessor already closed (kunchenguid#6355) * fix(bin): restore supervision host hand-back continuity * no-mistakes(review): Scope host hand-back failure fallback to lost pending:handling * no-mistakes(review): Remove stray scratch test copy tests/.tmp-rest.test.sh * no-mistakes(test): Initialise successor globals so early hand-back survives set -u * no-mistakes(document): Document host hand-back downtime failure and Claude lost-handback notice * no-mistakes(ci): I fixed the Greptile finding. The rule that must hold: when the supervision host hands back an actionable wake, its rewake is refused, and no watcher is healthy, the hand-back still has to reach main as a delivered notice. That must be true whether the recovery marker is `pending:handling` or `announced:handling`. Only one place applies this check: the lost hand-back fallback in `bin/fm-claude-stop-autoarm.sh`. **Fix:** that check now accepts both `pending:handling:*` and `announced:handling:*` tokens (a one-line change). Nothing else in the fallback changed: - Refusals on any other marker, such as an already acknowledged one, still exit 0 silently and open no failure episode. - The notice is still sent once per episode, and repeats are recorded as `failed-suppressed`. **Tests:** - `tests/fm-claude-stop-autoarm.test.sh`: the lost hand-back test now runs as a shared helper with two variants, one writing a `pending:handling` marker and a new one writing `announced:handling` (`test_host_lost_announced_handback_notifies_once_per_episode`). - `tests/fm-supervision-host.test.sh`: the end-to-end test where downtime restoration fails is now a shared helper too, with a new `announced` variant (`test_claude_stop_hook_notifies_when_closed_announced_successor_downtime_restore_fails`). It moves the handling episode to `announced` before the host hands back. Without the fix the hook would exit 0 here; the test requires exit 2, `outcome=failed` and a delivered failure notice. **Verification (all under nice -n 10):** - The full `tests/fm-claude-stop-autoarm.test.sh` suite passed (rc=0), including both lost hand-back variants and the benign-refusal test. - In `tests/fm-supervision-host.test.sh` I ran only the four hand-back test functions, all passing (rc=0). The suite can't run single functions, so I used a temporary copy with a trimmed test list and deleted it afterwards; `git status` shows only the 3 intended files changed. - shellcheck is clean on all three changed files. I did not run `bin/fm-lint.sh`. - I did not run the new tests against the unfixed code; the claim that they fail without the fix comes from reading the old check
* perf: cut remote-job idle process creation in the three hot loops Post-update host measurement still attributes most idle churn to three per-sample loops: result-consumer state reads and date calls, the delta reader's capture/hash pass on every poll, and the lane preemption scan's per-field pipelines. This drops each to its minimum without touching the contracts around them. * fm_remote_job_read_state gains an optional result-variable form backed by fm_remote_job_read_line, a builtin-only bounded record read (regular non-symlink file, byte bound, one newline-terminated line, tolerated unterminated tail, no carriage returns). fm_remote_job_wait samples state and the SECONDS clock with no per-sample children; one date call converts the epoch deadline once. * fm-remote-delta-read stats the log each poll and re-runs the bounded capture and hashing only when size, mtime, ctime, inode, or device change. The snapshot's own stat writes the comparison key, so a log that moves between the gate and the capture is never read as stable. * worker_preempting_waiter_exists reads state, home, and the staged argv head with builtins only. The now-unused worker_job_command goes away. The bounded reads use -d '' -n, which behaves identically on the macOS stock bash 3.2 and current bash; -N does not exist on 3.2. Tests cover the malformed-record corpus, delta identity gating, fork-free lane scanning through counting PATH shims, and same-home versus cross-home preemption. No signal traps or sleep contracts change. * no-mistakes(review): Restore subsecond delta keys and byte-bounded builtin record reads * no-mistakes(document): Clarify delta snapshot caching and coarse-timestamp fallback * no-mistakes(lint): Scope UTF-8 regression locales to individual function calls * no-mistakes(ci): Fixed both lint failures by applying the documented production-library analysis boundary at the two affected test imports. Runtime behavior is unchanged; the library remains independently linted. Canonical full-analysis lint passed for the library and both suites, as did bash syntax checks and git diff --check
…ommands (kunchenguid#5963) * fix(composer): read a titled Claude top rule as the composer's edge A named Claude Code session draws its title into the composer's top rule. The strict separator predicate rejected that row, so the closing rule read as a lower unmatched separator and an idle, empty composer classified unknown on every cursorless backend, refusing fm-send, exit, and relaunch. Spare a bare agent-glyph row sandwiched between a width-proven titled rule and the screen's only unmatched separator directly below it. The strict separator predicate, dead-shell rule, and blank-row posture are unchanged. Fixes kunchenguid#5601 Fixes kunchenguid#5558 * no-mistakes(test): Keep Claude's grey slash command in Herdr payload proof * no-mistakes(test): Make missing-herdr version check hermetic to installed herdr
…#6387) * fix: pre-approve Pi trust for seeded secondmate homes Unattended first launches of Firstmate-seeded Pi secondmate homes stalled on "Trust project folder?" until Enter. Probe --approve like --tui-mode and pass it only for --secondmate when help advertises it (.fm-secondmate-home signal), leaving ordinary workers and older Pi unchanged. * no-mistakes(document): Consolidate Pi seeded-home trust documentation ownership * no-mistakes(ci): Fixed Lint 1’s unused polling counter. Diagnosed Behavior portable serial 4 as a pre-existing delta-reader test clock race; replaced timing-dependent rewrite and deletion with deterministic executable-boundary synchronization. ShellCheck, Bash syntax checks, and git diff --check passed. Delta-reader tests passed three consecutive runs; all three live Pi trust cases passed. Production behavior unchanged
…--external-sources (kunchenguid#6443) * fix(lint): retry memory-bound roots without external sources * no-mistakes(review): Make fallback tests portable and correct source-following telemetry * no-mistakes(review): Remove committed parity fixtures and use disposable test roots * no-mistakes(document): Document ShellCheck memory fallback and telemetry * no-mistakes(review): Cover bounded and unbounded fallback RSS behavior * no-mistakes(document): Correct stale lint fallback documentation * no-mistakes(document): Correct stale lint test documentation * no-mistakes(ci): The memory fallback (the retry without --external-sources) now gets only the time left in its root's original deadline, so it can no longer outlast the CI job. Invariant: one root's first attempt plus its fallback must fit inside a single FM_LINT_ROOT_SECONDS deadline, plus the cleanup grace. Only one site started a new deadline: the fallback call in fm_lint_run_root. The deadline is the only budget involved, because the memory limit already applies to each process separately. Changes in bin/fm-lint.sh: - fm_lint_exec_root now takes a <seconds> argument instead of always reading FM_LINT_INTERNAL_ROOT_SECS. - The first attempt passes the full deadline. - The fallback passes floor((start + deadline - now) / 1000) seconds. - When bounds are enforced and less than 1 second is left, no retry starts. fm_exec_timed rejects 0 seconds, so the retry cannot run with no time. The root keeps reason=memory, and the shard output says "no time left in its Ns deadline to retry without it". - Unbounded local runs have no deadline and behave as before. - The header comment now describes the shared deadline. Changes in tests/fm-lint.test.sh: a new test, test_memory_fallback_spends_only_the_remaining_root_deadline, runs only on hosts that can enforce bounds. It uses a 6 s deadline and 1 s grace. - Case 1: the first attempt runs 3 s and then fails with memory status 251. The test asserts one fallback ran, reported reason=timeout, and the root's recorded duration is under 7000 ms. - Case 2: the first attempt runs 5.2 s. The test asserts no fallback starts, the skip is explained, and the sidecar records memory with source-following 1. Verification: - Full `nice -n 10 bash tests/fm-lint.test.sh` passed, including the new test, in about 5 minutes. - Case 1 run against the HEAD script: the root took 9168 ms, so the under-7000 ms check fails before the fix. - `bin/fm-lint.sh bin/fm-lint.sh tests/fm-lint.test.sh` reported no findings. - The CI workflow is unchanged, so the Test step still runs only tests/fm-lint.test.sh with nice -n 10 and the 12 GiB ShellCheck limit
…tension log opt-in (kunchenguid#5489) * Fix Pi watcher successor-gap confirmations and add extension log Accept an already-acknowledged handling confirmation as a no-op when the generation matches, confirm the restoration's own recovery token with a superseded (not rejected) outcome on generation mismatch, retire an arm on confirm failure only when the failed token names that exact pid, and record restore attempts, readiness timeouts, and confirm results in the bounded state/.watch-extension.log. Regression tests: already-acked no-op plus mismatch/dead-pid/lock-mismatch rejections and the manual-restart churn contract in fm-watch-arm.test.sh, and a mid-restore marker advance with no rejection appendix in fm-pi-watch-extension.test.sh. * Treat a dead arm child as an empty slot so repair and retry recover startArm and scheduleRetry answered unchanged while holding a ChildProcess whose OS process was already gone but whose close had not fired, so neither the repair tool nor the retry timer started anything until that close fired. Gate slot occupancy on a liveness check (exit/signal codes plus pid probe) and start a fresh arm instead, with a regression test driving the repair tool against a dead-but-unclosed child. * no-mistakes(document): Document new Pi extension log knob * no-mistakes(review): Fix confirm-failure retire token match, add distinct-pid test * no-mistakes(document): Clarify retire guard needs pid and generation * Make the Pi extension diagnostic log opt-in and default-off Only a positive FM_WATCH_EXTENSION_LOG_KEEP_LINES enables state/.watch-extension.log. Unset, empty, non-numeric, zero, and negative values disable logging entirely, so the default run writes nothing and never creates the file. The shared positiveInteger fallback semantics stay untouched for the retry and timeout knobs. docs/configuration.md owns the knob contract and docs/watcher-continuity.md points at it. Tests: the superseded-delivery case runs opted in, and a new case proves unset, zero, and non-numeric values create no log file while delivery still succeeds. * no-mistakes(document): Qualify extension-log coverage bullet as opt-in * no-mistakes(ci): The two reported checks (CI run 36372002913, Require no-mistakes run 36372069208) show conclusion action_required with 0 jobs and no logs because this is a fork PR (RibatTRW/firstmate) and GitHub is holding the workflow runs awaiting maintainer approval; that gate is external to the code and needs a maintainer to approve the runs. While verifying the change locally I found a real defect the approved CI run would hit: the PR's new test test_handling_delivered_rejects_a_superseded_generation failed deterministically. Invariant violated: reopen-announced (a non-successor/manual arm start) mints a fresh recovery generation only when the durable wake queue holds unrecovered work; an announced episode with an empty queue must be left untouched so idle arm starts never churn generations (the [ -s queue ] guard from kunchenguid#4819, relied on by bin/fm-watch.sh:2413 and covered by the append-reopens and announcement-bound sibling tests). The test called reopen with an empty queue and expected churn, so the fix establishes the queued-work precondition (append one wake, re-announce, re-read the generation) before asserting the reopen mints and the old confirmation mismatches. Test-only change, 14 lines in tests/fm-watch-arm.test.sh. Verified: fm-watch-arm.test.sh 25/25 ok on two consecutive runs, fm-pi-watch-extension.test.sh 55/55 ok, and shellcheck reports only one pre-existing warning outside the edited region * no-mistakes(document): Restore blank line in watcher-continuity docs * Route superseded Pi deliveries like confirmed ones and cover the retire guard A superseded handling confirmation now falls through to the normal delivery path, so an accepting supervision branch owns the wake instead of main. Scope the watcher-continuity token-pinned confirmation and narrowed retire rule to Pi, since omp and OpenCode still confirm against the current successor. Add tests that fail when the retire guard, the scheduled-retry gate, or the deferred-close gate is reverted, relabel the churned-generation characterization test, and use a reaped pid for the dead-pid rejection.
…es (kunchenguid#5863) * fix(pi): silence unacknowledged processing retry replies Suppress autonomous processing prose before persistence and during streaming while retaining tool calls, signed reasoning, and retryable outcomes. Restore ordinary output after acknowledgement or a user message. Fixes kunchenguid#4954 * no-mistakes(review): Silence only processing retries, keep first presentation visible * fix(pi): preserve differing processing retry replies
…henguid#6530) Pi 1.0.1's createToolHtmlRenderer reads getToolRenderers and ignores getToolDefinition. Calm /export still includes stock grep HTML; the fixture has to pass the lookup key the installed Pi actually reads.
* fix(procevent-quota): tolerate consecutive slow quota-axi reads The quota allowance poll treated any quota_json failure as terminal, so one slow quota-axi --json (measured max ~29s under a 48s derived bound) shut the watch down until someone re-armed it, and the detail always said "missing/incompatible". Tolerate three consecutive failed or timed-out reads before going terminal, reset the streak on any good read, and report a timeout distinctly from a missing or incompatible tool. Each timed poll runs exactly one bounded --version and one bounded --json: validate the captured version text through fm_quota_axi_version_compatible rather than launching a second probe, and describe a mixed failure streak by count plus last cause. * no-mistakes(document): Document quota polling failure tolerance * no-mistakes(ci): Fixed ci-2 and ci-3. Permanent quota read failures (rc 2 missing, rc 3 incompatible) now report on the first poll, while transient rc 1/4 failures retain the existing three-failure retry behavior. The missing-binary test now uses an isolated PATH without quota-axi and asserts both permanent failures stop at condition_polls: 1. Verification passed: tests/fm-procevent-quota.test.sh, canonical fast lint for both changed files, bash syntax checks, and git diff --check * no-mistakes(review): Classify untimed quota version failures as transient * no-mistakes(document): Clarify quota polling failure budget
…enguid#6505) * fix(teardown): clarify scratch guidance and dirty worktree refusals Keep ship proof material outside the task worktree and distinguish untracked-only leftovers from tracked edits without changing cleanup guards. Fixes kunchenguid#6319 * fix(ci): Fixed ci-3 only. Both promotion outputs now replace the scout restriction and require external scratch storage and a clean worktree before done. Verification: 27 delivery tests passed, five mutations caught, restored test passed, pinned ShellCheck and syntax/whitespace checks passed. ci-1, ci-2, and ci-4 remain untouched
…nguid#6518) * Escalate inbox instructions stuck behind a busy worker Count consecutive busy-deferred due doorbells durably and escalate at the configured bound without typing into the worker pane. Fixes kunchenguid#6445 * fix(review): Fix inbox escalation deduplication and busy streak resets * fix(review): Preserve busy inbox escalations through daemon supervision * fix(document): Correct busy-inbox escalation documentation * fix(ci): Fixed SC2034 in tests/fm-task-inbox.test.sh by including the loop counter in the failure diagnostic. Source-aware lint, all 34 inbox tests, and git diff --check pass. Behavior portable serial 6 reproduces identically on base 1f3e769 and target 78156b8 with Pi 1.0.1: an unrelated renderer API change breaks the unchanged Calm test. No Calm changes made; that failure is addressed separately by kunchenguid#6516. Logs retained in scratchpad-ci/ * fix(ci): Fixed ci-2, ci-3, and ci-4: successor failures surface, reset alerts deduplicate, and oversized busy limits fall back to two. Passed 47 inbox tests, 7 focused daemon checks, all 13 mutation checks, lint, documentation checks, and diff checks. Evidence: scratchpad-ci-selected/summary.json. ci-1 remains unchanged and unwaived. Fresh live Herdr proof remains with the outer driver
* fix: prevent healthy remote job worker turnover * no-mistakes(review): Serialize full LaunchAgent repair and verify launchd-tracked owners * no-mistakes(review): Let launchd-tracked unpublished spawns start before reloading * no-mistakes(document): Document remote worker heartbeat and serialized LaunchAgent recovery * no-mistakes(ci): Full CI log showed the idle-worker regression exceeded its outdated command budget (82 versus 80) after independent heartbeat ownership checks were added. Raised the budget to 120 while retaining the separate busy-poll sleep limit. Remote-job and LaunchAgent executable tests passed, as did bash syntax validation and git diff --check. No production behavior changed * no-mistakes(ci): Fixed missing readiness recovery under verified live ownership, preserving the serving PID and private file mode. Added executable regressions for deletion during a blocked sweep and stale readiness diagnostics without LaunchAgent reload. Deletion regression failed before the fix. Both remote-job test suites, bash syntax validation, and git diff --check passed * no-mistakes(ci): Full CI log identified a flaky ownership-loss test racing an already-authorized heartbeat refresh. Replaced backdating and a fixed sleep with bounded observation of readiness expiry through the public probe. Production behavior unchanged. Remote-job and LaunchAgent executable suites passed; bash syntax validation and git diff --check passed * fix: wait for launchd bootout cleanup * no-mistakes(document): Document remote worker recovery and read-only turnover verification * no-mistakes(review): Publish worker identity before lock owner records * no-mistakes(document): Document worker identity publication safety invariant * no-mistakes(ci): Fixed ci-1: replacement workers discard predecessor readiness before publishing identity and roll back identity if lock-owner recording fails. Added executable regressions reproducing both defects. LaunchAgent and orphan-reap suites, shell syntax checks, and git diff --check passed. No live service state was modified * no-mistakes(ci): Restored lock-owner-before-identity publication and removed the identity rollback and reordering-only tests. Retained an executable regression proving predecessor readiness is rejected until replacement startup completes. LaunchAgent and orphan-reap suites, shell syntax checks, and git diff --check passed. Other changes remain intact; the retained-identity interrupted-repair edge remains out of scope. No live service state was modified
* fix(remote-job): reap expired seq claims with one directory walk The hourly claim sweep forked uname+stat per .seq-claims entry and blocked serving for ~85s at ~17k dirs. Delete expired empty claim dirs with a single find -exec rmdir batch and cache the host uname for remaining mtime reads. * no-mistakes(review): Restore original path mtime helper and drop uname cache * no-mistakes(test): Restore claim retention eligibility; focused retention and serving tests pass * no-mistakes(document): Document single-walk claim cleanup and regression entrypoints * no-mistakes(ci): Fixed Lint 2’s reproduced SC1091 by adding the tests/lib.sh ShellCheck source directive to the retention test. Runtime behavior is unchanged. ShellCheck passed for both new claim tests; Bash syntax, the retention behavior test, and git diff --check passed * fix(tests): keep fixture registries out of git worktree roots A TMPDIR pointed at a repository root placed live .fm-test-* registries beside tracked files, and a concurrent git add during the claim-walk CI fix round committed three of them. Route registries and fixture roots through a TMPDIR that refuses git worktree roots, remove the stray files, and pin the escape with a behavioral cleanup test. * no-mistakes(review): Preserve whole-second claim expiry in single-walk sweep * no-mistakes(document): Correct temporary-directory resolution documentation * no-mistakes(ci): Fixed ci-3 by changing only the stale, fresh, and read-only orphan fixture paths in tests/fm-test-fixture-cleanup.test.sh to use $FM_TEST_TMPDIR. All seven tests passed both normally and with TMPDIR set to the worktree root. Shell syntax and git diff --check passed
…henguid#5895, kunchenguid#5911 and kunchenguid#5929 # Conflicts: # docs/watcher-continuity.md
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Merges upstream
kunchenguid/firstmatemain into this fork (all PR numbers below are upstream pull requests) with a real merge commit (no rebase, no squash), keeping the three fixes the fork carries ahead of upstream.050a4464(merge base, 2026-09-27) tof470a01c(2026-10-04), 62 commits, all non-merge.fm_exec_timed.Conflict and resolution
One content conflict, in
docs/watcher-continuity.md("Recovery episodes" paragraph).Upstream rewrote the non-successor arm rule (an empty-queue arm now leaves an announced generation announced, and a later durable row opens a fresh pending generation).
The resolution keeps upstream's new text verbatim and re-appends only the two sentences PR 5929 added about a handling successor standing down only for its inherited generation.
The other two PR 5929 hunks in that file merged cleanly.
No other file conflicted.
Upstream's reworked
_fm_recovery_marker_arm_checkstill turns apending:downtime:<gen>marker intorecover, which is what the PR 5929 successor path relies on when it announces a newer generation.Evidence
Carried fixes are intact
git diff origin/main HEAD --statlists exactly the 18 files thatgit diff origin/main...fork/main --statlisted before the merge, with identical line counts (18 files changed, 294 insertions(+), 50 deletions(-)):For every one of the 18 files, the added and removed lines of
git diff origin/main HEAD -- <file>are byte-identical to those ofgit diff origin/main...fork/main -- <file>(only hunk offsets moved).So the merge changes nothing beyond what upstream ships plus the three carried fixes.
Tests (local macOS, stock
/bin/bash3.2.57)All runs used a scrubbed environment (
env -iwith only HOME, PATH, USER, locale, TMPDIR) andFM_LIVE=0, so no live Herdr, Claude, or Pi lifecycle was launched; the live-harness guards reportskip: live: disabled by FM_LIVE=0.This host has no tmux, so tmux-backed tests capability-skip or fail as noted below.
Lint
bin/fm-lint.sh(changed-file mode): pass.CI=true bin/fm-lint.sh(full canonical set with extended analysis, plus actionlint): pass.Carried-fix tests, each run directly under
/bin/bash:tests/fm-dispatch-resolve.test.shtests/fm-timeout-lib.test.shtests/fm-bootstrap.test.shtests/fm-watch-recovery-loop.test.shtests/fm-test-run.test.shtests/fm-watch-triage.test.shPortable behaviour suite via
bin/fm-test-run.sh --lane portable-parallel-1|portable-parallel-2|portable-serial(231 scripts):f470a01cin this same checkout (detached), so they are upstream behaviour on this macOS host, not merge regressions:fm-pi-primary-typesfm-afk-returnfm-cursor-harnessfm-gate-refusefm-git-strip-ai-trailerscore.hooksPathinstallfm-gotmpfm-harness-precedencefm-muse-harnessfm-proceventfm-remote-doctor--fixleft a repairable host unreadyfm-remote-herdr-guardfm-watch-triagefm-supervision-hostexceeded the suite's 1500s per-script bound on this Mac (59 of 73 cases passed before the bound; its CI hint is 789s on Linux).Rerun alone with a 3600s bound, cases 1-55 passed and case 56 (
a captain return during a failed engine turn still hands that turn's outcomes to main) failed its 25swait_untilonce.That case then passed 5 of 5 isolated runs on this branch and 5 of 5 on an export of upstream plus PR 5911, and a run of cases 56-73 on this branch passed all 18.
Every one of the 73 cases has passed on this branch; case 56 failed in 1 of 9 runs, a timing-window failure on this host.
Notable difference from plain upstream: on plain upstream
f470a01c,fm-supervision-hostfails its first case deterministically under bash 3.2 (3 of 3 runs:not ok - the host must hand its captain outcome to main, withthe engine turn failed (exit 1)).An export of upstream with only the PR 5911
bin/fm-timeout-lib.shchange applied passes that case and the next ones.So upstream's supervision-host engine turn depends on the carried PR 5911 fix to work under stock macOS bash 3.2.
Upstream changes that alter behaviour on a Claude-primary home
config/supervision-hoston a Claude primary now means the host runs with the Claude engine at its default model.Opt out with the presence file
config/supervision-host-off, which secondmates inherit from the primary.An
offline insideconfig/supervision-hostis no longer an opt-out and has no migration./quietis a no-op statement where the attended host runs (PR 5928), quiet records treated as attended (PR 6064), routine no-change outcomes silenced (PR 5808).--add-dirfor the task's channel directories (operational inbox, steering inbox, task data, code-root skills), so file-tool reads of those paths no longer stop on a permission question (PR 5884).--helpno longer arms supervision (PR 6032); titled Claude top rules and grey slash commands are recognised in composer reads (PR 5963).config/wait-no-turnslets a waiting worker spend no turns; worker briefs now drive no-mistakes with one foreground call (PR 4859).fm-pr-mergeretries when GitHub mergeability is UNKNOWN (PR 6110), captain-hold reasons percent-encoded (PR 6331), configurable per-task endpoint read bound (PR 5917), emptycore.hooksPathruns no hook (PR 6216).Pi 1.0 support
--tui-mode regularwhere the flag exists (PR 6338), and the Calm export fixture follows Pi 1.0.1's renamed renderer lookup (PR 6530).--approvewhen Pi's help advertises it, so first launches do not stop on "Trust project folder?" (PR 6387).Merge Danger
Door: two-way.
A merge commit on the fork's main; reverting it (
git revert -m 1) restores the previous fork state.Running homes pick it up only when they fast-forward and restart.
Blast Radius: every home running from this fork after it updates.
The largest behaviour change is the supervision host defaulting on for Claude primaries; a home already holding
config/supervision-host-offkeeps it off.The three carried fixes are textually unchanged, and the supervision host's engine turn depends on the PR 5911 fix under stock macOS bash 3.2.