feat(bin): sync upstream Firstmate for the Herdr 0.9 upgrade - #16
Merged
doitdigital0495 merged 183 commits intoSep 11, 2026
Merged
Conversation
…enguid#2707) * fix(bearings): always show decision options and a close/drop control Freeform-only Captain's Call cards hid the option buttons the board was designed around, and there was no way to drop a stale hold without inventing an answer. Require selectable options, keep freeform as a supplement, and route the reserved __drop__ answer through decline so the hold leaves Captain's Call. * no-mistakes(review): Fix drop closure and decision-only option validation * no-mistakes(review): Preserve answerability for non-decision cards * no-mistakes(document): Clarify decision drop documentation
Signature-only PRs can hide skipped review, test, or document steps. Fail unless no-mistakes >= 1.46.0 attests those three steps completed.
…#2728) * feat(captain-hold): collapse the decisions concept into tasks held for the captain A decision is no longer a separate type: it is an ordinary backlog task held for the captain, identified by its task id. bin/fm-captain-hold.sh owns the surviving behaviors - guarded hold creation, the recorded-answer close (answer/answers with a release mode for captain-gated work), the source bindings, and the investigation completion gate - and bin/fm-decision-hold.sh becomes a one-release compatibility shim over it. The fleet snapshot now parses hold-until and computes captain_actionable as queued + captain-held + unblocked + due, independent of row kind, plus a presentation-only deferred_marker for prose-deferred rows. Bearings renders every due captain-held task in Captain's Call, date-deferred holds as dated Charted Next gates, suppresses prose-deferred rows from default views with an omitted disclosure, and excludes from Recently Landed anything that closed while still held for the captain. Legacy compatibility: pre-collapse <origin>-decision-<key> rows are already plain task ids and keep working; short keys in recorded metadata, concrete origin bindings, chat --resolve-key fallbacks, and old resolution records all resolve in place. * no-mistakes(review): Fix captain answer replay and body preservation * no-mistakes(review): Fix captain hold idempotency and legacy replay * no-mistakes(review): Validate card close modes and compatibility routing * no-mistakes(review): Enforce release replay mode matching * no-mistakes(review): Prevent duplicate decision cards and released replay mismatches * no-mistakes(review): Preserve answer columns and legacy resolve replays * no-mistakes(document): Document strict replay and legacy compatibility * no-mistakes(lint): Quote done literals to satisfy ShellCheck * no-mistakes: apply CI fixes * fix(rebase): keep collapsed captain hold board semantics
…id#2733) * fix(watch): announce recovery once per generation and keep successors supervising A lost Pi/OpenCode handling handshake re-announced the same recovery generation on every cycle and spent the successor's first ~55s blind, so a real crew event could be ignored and then dropped. Record the announcement in the durable marker, confirm the handshake before the follow-up without swallowing failure, and enter the poll loop immediately. * no-mistakes(review): Tighten recovery event timing regression * no-mistakes(document): Document recovery-loop supervision guarantees
* fix(bin): signal a captain call resolved in the log but still held A captain call has two records and closing one has never closed the other: a `resolved [key=...]` line closes the status-log fold, while the backlog task held for the captain closes only through `fm-captain-hold.sh answer`. Answering on the status side alone left no trace of the disagreement - the fold went quiet, the durable record kept saying the captain owed an answer, and nothing warned. The defect was never the separation; it was the silence. Add `fm-captain-hold.sh diverged`, a read-only report of that contradiction, and print it from `fm-wake-drain.sh` as a bounded RECORD DIVERGENCE section beside OPEN DECISIONS on every drain. It flags one condition: a task still open and still carrying the captain-hold annotations whose key was closed on the status side by the resolve verb, under the collapsed identity or the legacy derived one. It closes nothing, ever. A captain call closed wrongly leaves review entirely, which is worse than the noise, so both reconciliation directions stay human-owned and the printed hint names both - a resolution is not proof the captain ruled, since a call can dissolve on a false premise or turn out to have been a question of fact. Three states are deliberately not divergence: a `captain-held` close is the verified transfer `complete` writes, a still-open keyed decision belongs to the OPEN DECISIONS fold, and a captain call with no routed work item is legitimate rather than incomplete, so routed work is no part of the test. `fm-classify-lib.sh` gains `status_key_closing_verb`, which reports how the status side currently reads one key by replaying the existing `_fm_decision_fold_line` rule rather than re-deriving it, so the two closing verbs stay distinguishable in one place. The per-wake cost is one `tasks-axi list`, one key scan per status log, and the precise per-key fold only for a key that already names a still-open task; the call is hard-bounded so a slow backlog tool can never delay wake presentation. * fix(document): Correct divergence lifecycle documentation * fix(document): Neutralize divergence lifecycle prose
…escalation while a worktree is written (kunchenguid#2524) * fix(watch): re-arm supervision after an abandoned auto-arm claim A Claude auto-arm cycle that armed, delivered one rewake, and exited left its single-flight lock behind. Both Stop-event participants then deferred to that lock forever, because its recorded pid was still live: the turn-end guard read it as recovery under way and allowed the stop, and the next Stop firing treated it as another owner and declined to arm. On 2026-08-14 a home with two tasks in flight lost supervision for about 40 minutes with no watcher process and no watcher lock, its beacon frozen at the one delivery, and both crewmates' finished reports sat in the durable queue until an operator drained it by hand. Abandonment is now proven from the epoch ledger instead of inferred from pid liveness. A lock whose holder pid matches the ledger's own owner_pid while the recorded outcome is anything other than arming has already finished its decision, so that claim is reclaimed under the lock's steal mutex, stops counting as recovery ownership in the guard, and is cleared by the guard's terminal check rather than deferred to. A failed clear re-blocks instead of allowing a blind stop, and an arming entry stays in flight however old it is, because its owner foregrounds the arm for the whole watcher cycle. Issue kunchenguid#2251's PR kunchenguid#2263 does not cover this failure. It is closed and unmerged, lives entirely in bin/fm-watch-arm.sh, and retires the stalled watcher and matching stale watcher lock of an arm that is currently running. Here no arm and no watcher were running and no watcher lock existed, so it has nothing to retire and the home stays blind. tests/fm-claude-stop-autoarm.test.sh covers the reclaim, the still-arming and unnamed-owner cases that must keep the gate closed, and the failed clear. tests/fm-turnend-guard.test.sh covers the guard side of the same boundary. Both fail without this change. * fix(watch): defer a wedge escalation while the task worktree is written The wedge detector had two inputs, rendered pane quietness and the run step, and neither can see a crew that is writing source, then tests, then documentation behind a static pane. On 2026-08-14 one crewmate produced eight consecutive possible-wedge escalations in a single afternoon, three of them demanding deep inspection, while it was demonstrably working and then committed. Every one of them cost a supervision turn to disprove by hand. Add write activity inside the crew's own recorded worktree as a third liveness input. crew_worktree_written_since compares the worktree against the caller's existing idle-window timer file, so -newer needs no clock arithmetic, no temp file, and no portable mtime write. The probe runs only inside the branch that was about to escalate, which bounds it to one pruned, depth-bounded walk per window per FM_STALE_ESCALATE_SECS and leaves the per-poll stale sweep exactly as cheap as before. Positive evidence defers rather than cancels. The idle timer restarts so the next window probes again, the escalation counter is neither advanced nor reset so a later genuine wedge keeps the demand-deep-inspection history it earned, and a .writing-since marker ages the whole deferral chain so the pane still re-surfaces once per FM_PAUSE_RESURFACE_SECS, through the same throttle shape a declared pause already uses, labeled as a recheck rather than a wedge. This can only reduce false positives: every absence of evidence, including no recorded worktree, a torn-down worktree, a missing anchor, and a failed walk, falls through to the unchanged escalation schedule, so a crew that writes nothing still escalates on the existing timetable. What the signal cannot see, by design or by construction: - CPU burn with no writes, such as a long compaction, is invisible. That case keeps the old behavior exactly. - A commit-only phase writes only .git, which is pruned first so that firstmate's own read-only git commands against the worktree can never make the probe self-fulfilling. - Writes under the pruned generated trees, or deeper than FM_WORKTREE_WRITE_MAXDEPTH, do not count. - The probe cannot attribute a write to the crew, so a background build or another process touching the tree looks the same. The hourly re-surface is what bounds that, and a churny file cannot buy silence. - The away-mode daemon's own escalation path is deliberately untouched. tests/fm-watch-triage.test.sh covers the classifier including the .git prune, both halves of the live case on one fixture (quiet plus writing defers, quiet plus silent still escalates and counts), and the bounded re-surface. All three fail without this change. * no-mistakes(review): prove autoarm claims by identity; skip mate-home write probe * no-mistakes(document): document away-mode wedge boundary and probe filesystem limit * no-mistakes(document): qualify turn-end recovery condition for abandoned auto-arm claims * fix(watch): keep a write deferral scoped to its own idle window Two consistency gaps in the worktree write probe, both found while reviewing the wedge-deferral change on this branch. A write deferral is a bounded chain: its .writing-since marker ages the whole chain so a churning worktree still re-surfaces once per resurface window. That is only sound while the chain belongs to the current quiet stretch, so every path that restarts the idle-window timer has to drop it too. Two did not: the corrupt-timer repair in wedge_timer_check, and both first-sight branches for a captain-relevant status. A chain left over from an earlier quiet stretch made the first deferral of the new window re-surface immediately instead of after a full fresh window. FM_WORKTREE_WRITE_PRUNE is a skip list, so clearing it reads as "skip nothing" and is the obvious way to widen the probe to the whole depth-bounded tree. Instead an empty list reported no evidence at all, quietly costing the wedge detector its third liveness input on a home that meant to widen the walk. An empty list now widens the walk, and the header says so. Neither change alters when a stall that writes nothing escalates. Regressions in tests/fm-watch-triage.test.sh cover all three paths and each one fails on the pre-fix code. * no-mistakes(review): honor an empty write-prune, bound the probe, share window_key * no-mistakes(document): align probe knob count and guard regression-coverage ownership * no-mistakes(lint): silence deliberate single-quote SC2016 in write-prune env test
…lared pause (kunchenguid#2748) * fix(bin): give a captain hold the same bounded pause cadence as a declared pause Two supervisors read a finished task's last status line and disagreed about which declarations mean an idle endpoint is expected. bin/fm-inactive-reconcile.sh suppresses its inactive-outcome scan only on `captain-held`, while the away-mode daemon's wedge path gated deferral on `paused` alone. Both read the LAST line, so the two verbs are mutually exclusive and no finished task waiting on a person could satisfy both at once. Marking 11 such tasks `captain-held:` silenced the 900s outcome scan and immediately produced five possible-wedge escalations in one batch, because the 240s wedge detector no longer saw a pause verb. fm-classify-lib.sh's status_is_paused_or_captain_held already owns the combined question, and bin/fm-watch.sh's ordinary-crew wedge path already asked it. This extends that same answer to the paths still asking the narrower one: - bin/fm-supervise-daemon.sh, all six sites, which form one subsystem and have to move together. classify_stale returns the pause action, reconcile_pause_tracking and migrate_watcher_pause_markers record and migrate the marker, and housekeeping defers the wedge and then re-surfaces the recheck. Changing only the stale-persistence gate would defer the escalation while reconcile_pause_tracking recorded nothing, so the wedge marker would persist and the sweep would `continue` past it forever: quiet, but never re-surfacing. - bin/fm-watch.sh's secondmate stale gate, whose downstream owner pause_state_class already treats both declarations identically. - bin/fm-push-transition-lib.sh's absorb, where either declaration already names the human the transition would report and the wait is already durably recorded. Quieting alone would be half a fix, so the bounded re-surface had to reach a hold too. A hold has no current-state mapping, unlike `paused`, so authoritative crew state reports it as unknown and pause_state_class received `none`. An ordinary crew recovers pause classification from that state through confirmed agent death, which proves no live decision gate is being silenced. A secondmate's endpoint liveness is deliberately never read there, because an idle mate is healthy by design, so that confirmation is unavailable by construction and cannot be required: without recovering the classification for a mate, every caller silenced a held mate outright and its hold would rot invisibly. That promotion is bounded by the declared-wait guard at the top of the function, so it can only reclassify a task that already declared a wait and shows no positive working evidence. Two narrow `status_is_paused` calls are deliberately left alone. bin/fm-crew-state.sh's map_log_state is a current-state reporting contract, not a wedge path; reporting a hold as `paused` would erase the distinction status_key_closing_verb and fm-captain-hold.sh depend on, where a `captain-held` close is a verified durable transfer and a `resolved` close claims outright settlement. fm-classify-lib.sh's call inside status_is_captain_relevant needs no change because that function's own case list already returns non-relevant for `captain-held`. bin/fm-inactive-reconcile.sh keeps its `captain-held` suppression as it is. Its guard exists because a finished task's crew state still reports done from a higher-priority source than the log, and a declared pause needs no such guard: the scan only reports done or failed, and nothing else reaches its record path. Widening it would change a separate subsystem's reporting contract, which this defect does not require. Coverage extends the existing colocated patterns for these predicates and asserts both halves. tests/fm-daemon.test.sh covers the classification, the wedge marker converting to pause tracking with no escalation, the bounded re-surface with its window reset, and the boundary case where an answered hold stops claiming the cadence. tests/fm-watch-triage.test.sh covers a held secondmate re-surfacing on the same bounded cadence without being labeled a wedge. tests/fm-supervision-events.test.sh covers the absorbed push transition. Every one of these fails on the pre-fix code except the answered-hold boundary case, which is there to pin that the quieting was not widened too far. The `paused:` workaround appended to those 11 tasks is live supervision state and is untouched here. It can be retired once this lands. * no-mistakes(review): name the captain in a held task's bounded recheck * no-mistakes(document): extend declared-wait supervision docs to captain-held holds
…guid#2758) * fix(lint): name the installer when ShellCheck or actionlint is missing A missing actionlint exited 127 like a bare command-not-found. Fail with exit 1 and point at the pinned installer, matching the missing-ShellCheck path, without weakening the version pin. * test: isolate kimi and muse detection from inherited Cursor markers Harness detection checks CURSOR_AGENT before ancestry, so these markerless-adapter cases failed when the suite itself ran under Cursor. Clear the verified markers the same way the secondmate harness tests already do. * no-mistakes(document): Document Muse Cursor marker cleanup
…lled but inert (kunchenguid#2684) * feat(checks): report tool updates that are available or installed but inert Firstmate had no way to notice that tooling this home depends on needs an update, and no way at all to notice the worse case: an update that installed correctly and then did nothing. That second case is why this exists. A tool that self-installs into ~/.local/bin while a version manager keeps its own older copy earlier on PATH looks completely up to date to anything that asks only "is a newer version published". On 2026-08-20 a Herdr update landed at 0.8.2 while an older 0.8.0 copy stayed earlier on PATH, so every Herdr command failed on a protocol mismatch and firstmate could not read its own fleet. bin/fm-tool-update-check.sh reports the two conditions separately: <tool> update available a newer version exists at the update source. <tool> update not in effect a newer copy is installed on this host, but PATH still resolves an older one. PATH skew is measured, never inferred. Every executable copy of a watched command on PATH is asked for its own version and those answers are compared, so one lookup cannot hide the skew, and a directory name is never read as a version because a version manager's "latest" directory can hold an older build. A copy that will not report a version is a check failure, not a pass. The watched tools live in local, gitignored config/watched-tools.json, so adding a tool is a config edit rather than a code change, and the file is never propagated to another home. Update sources cover both shapes: a local clone's commit distance from its remote branch, and a command's own version and update announcement, including a tool like no-mistakes that prints its version on one command and announces a new release on another. The check prints one line when something needs attention and prints nothing otherwise, so it rides the existing watcher state-check contract with its trust binding instead of introducing a schedule of its own, and state/.tool-updates keeps the same pending update from being reported on every poll. The check only reports. It never installs, updates, reorders PATH, touches a version manager, or fetches into a watched repository; every git probe is read-only. Tests cover the skew case as a regression, and it was verified by mutation: removing the skew report, or stopping after the first PATH hit as a single lookup would, each make that test fail. * no-mistakes(review): fix tool update check probe reporting, budget, and shim write * no-mistakes(review): keep sweeps alive on broken patterns and oversized budgets * no-mistakes(review): roll back failed arm, widen budget clamp, bound repo probe * no-mistakes(review): guard git probes at the budget, record uncut findings * no-mistakes(document): fix stale watched-tool report-record wording in docs and header * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes The behavior shard's watch-triage suite failed on the new worktree-write wedge tests. Those five tests are the only ones in the file that do not use its standard waits. They give a fixed 3 second liveness budget to the one poll that now spawns the bounded worktree walk, and 4 seconds to an escalating watcher where every other test in the file gives 10. On a loaded runner that poll outlives the fixed budget, so the round is reaped before the deferral it asserts on is recorded, and the test reports a lost deferral instead of the deferral under test. Wait for a completed poll cycle through the file's own wait_poll_cycle, which is what its header documents this hazard for, and use the file's standard 100 tick exit budget. Verified against a load that reproduces the failure: 11 of 12 runs failed before, 8 of 8 pass after. Verified by mutation too, so the waits still prove the behavior: removing the write deferral, and keeping a finished deferral chain across an idle-timer repair, each still fail their test.
* fix: treat yolo as merge authority only, not ask-user finding authority Yolo on/off was documented as also deciding no-mistakes ask-user findings, which hid firstmate's duty to judge unambiguous-toward-design findings itself. Keep every safety boundary; this is a contract clarification, not a relaxation. * no-mistakes(document): Clarify yolo documentation ownership and merge posture
… work over (kunchenguid#2767) * feat(voice): spoken round trip on Nova Sonic 2 with a measured relay cost Step one of the spoken interface: the laptop captures and plays audio, this desktop holds the model session, and no AWS credential leaves the desktop. Measured, amazon.nova-2-sonic-v1:0 in eu-north-1, end of speech to first byte of reply audio, 6 runs each, all answered, on a question that forces a records read: relay path 1.229 1.379 1.428 1.447 1.481 1.516 median 1.438 direct 1.147 1.179 1.203 1.237 1.244 1.317 median 1.220 The relay costs about 0.22s of the median. The direct figure reproduces the earlier survey, which is what makes it a usable control. Excluded: the captain's own ssh round trip, microphone capture, and speaker output. This desktop has no microphone and no speaker, so every run used audio files. Three pieces: bin/fm-voice-relay.py holds the conversation on this host bin/fm_voice_records.py what a spoken answer may read, and the handover bin/fm-voice-client.py the laptop end; audio devices UNVERIFIED bin/fm_voice_frame.py the wire format both machines share Real work is handed to the existing bin/fm-inbox.sh rather than a second queueing surface, and the agent says it is handing over rather than answering as firstmate. Read scope: Done history and free-form note bodies are never assembled at any scope, so the wide default cannot reach the places commercial detail accumulates. config/voice-read-scope narrows it to counts only, and config/voice-read-deny excludes a named item in one line. The boundary is an executable test that widening the reader fails. Push to talk is the default because it is cheaper and the choice is still open; --listen open-mic is the single flip. Two traps worth knowing: a clip with no trailing silence is never answered, and the end of a reply is contentEnd with stopReason END_TURN, not completionEnd. A second user turn in one session is treated as barge-in unconditionally, and an interrupted turn that calls a tool is lost, so the session reconnects per turn and gives up conversational memory. That is the concrete thing step three has to solve. * no-mistakes(review): fix voice relay credential reuse, frame validation and record parsing * no-mistakes(review): test uplink header guard, bound unknown expiry, align state dir * no-mistakes(review): decide deny per item, guard turn failures, bound ambient credentials * no-mistakes(review): read account config from home, harden deny and turn failures * no-mistakes(review): close status verb set, fix inbox help, pair data override * no-mistakes(review): keep profile-free relay alive, unblock loop, fix dead assertion * no-mistakes(review): hide finished pull requests, refuse open mic, keep suite offline * no-mistakes(review): survive reader failures, release devices, fix claims A failure while handling a model event, or while sending a tool result, left the reader task dead with ended and turn_done clear, and close() re-raised the stored failure on every await. One dropped stream became a relay that could never build another session. The reader now reports the session over in a finally whatever killed it, and close() absorbs the task the same way it already absorbed its sends. The laptop client releases what it already started when a later startup step refuses, SystemExit from the handshake wait included, and names a device refusal instead of leaking a raw PortAudio error. Whether it releases correctly against a real device is still unverified here. The records docstring claimed every reading was filtered to open ids. Only the pull request count and list are; the worker count and the state histogram cover every live runtime record, finished ids included, because a meta file still on disk still needs tearing down. The finished-work deny half of the suite asserted things that held with the deny list absent. It is replaced by a deny on an open title, which removes the row and says so while the count stays honest. * no-mistakes(review): name reader failures, split file and device refusals A failure inside the model reader released the waiting turn and told nobody. The session was not marked spent, no notice reached the client, and the client waits for a reply end or a notice, so the captain got their whole timeout of silence and then a record saying the turn went unanswered with nothing about why. Both ends of the relay now name a failed turn through one function, once per turn, and --self-test carries the cause in relay_error the way the client's own record does. Two things that are not failures stay that way. A stream that simply ends is the end of a session, which serve still reads on its own terms. A stream that goes away because close() asked it to is an ordinary renew, and announcing it would have put a failure notice in front of the captain on every turn. On the laptop end, the refusal that became a device error covered the file-backed playback and capture too, so a mistyped --in-file was reported as an audio device failure and the advice named the flag that had just failed. The file ends now report the path and the flag that chose it and stay an OSError; the device ends keep the device advice and name the flag for that end. The device paths remain unrun here, so only the file halves are covered by a test. * no-mistakes(test): survive model session end, order client turn frames * no-mistakes(document): sync voice relay docs with reviewed relay behavior * no-mistakes(document): re-measure relay latency and correct its cause * no-mistakes(document): correct measurement date and name the unmeasured SSH hop * no-mistakes(document): describe the unpublished control measurement, fix list formatting * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes
…unchenguid#2763) * fix: keep Relay public loops open until retire Delivering a promised-final reply was deleting the only record that tied a public thread to later work, so a follow-on ship silently owed no closing reply. Retain the registration after delivery, rechain follow-on work onto the same thread, and make retire --reason the only close. * no-mistakes(review): Propagate public follow-up registration removal failures * no-mistakes(review): Persist retire receipts and align parent resolution * no-mistakes(review): Make rechain resumable after partial obligation creation * no-mistakes(review): Repair follow-up state, briefs, and expiry escalation * no-mistakes(review): Serialize follow-up delivery stamps with retirement * no-mistakes(review): Serialize rechain claims and protect registration terminal states * no-mistakes(review): Avoid reporting retired delivery loops as open * no-mistakes(document): Refresh public-loop documentation and verification evidence * no-mistakes: apply CI fixes * no-mistakes(review): Preserve delivered follow-up bindings during registration replay * no-mistakes(review): Harden public follow-up retirement and rechain races * no-mistakes(review): Fail closed on unresolved secondmate retirement * no-mistakes(review): Bind secondmate cleanup to its recorded canonical home * no-mistakes(review): Fix rechain command output and expiry validation * no-mistakes(review): Validate brief keys and warn on remote promotion * no-mistakes(document): Document retained public follow-up loops * no-mistakes(lint): Remove unused bounded-wait loop variable
…ath (kunchenguid#2779) * feat(bin): merge GitLab merge requests through the guarded PR merge path bin/fm-pr-lib.sh already parses a GitLab merge request URL for the watcher, but bin/fm-pr-merge.sh refused every non-github provider, so a merge request had to be merged by hand and got none of the recording, guards, or audit trail a pull request gets. The merge path now dispatches on the parsed provider. A GitHub URL keeps its exact previous behavior. A GitLab URL is addressed through glab by the project URL rebuilt from the parsed host and path, so a merge request on any instance resolves and no host is hardcoded, and no merge-method flag is added because the project's own merge method is what should apply. A GitLab merge happens only after one live read of the merge request confirms it is open, detailed_merge_status is mergeable, has_conflicts is false, blocking_discussions_resolved is true, and the head pipeline succeeded at the exact current head. Every failing condition is reported, not just the first. The verified head is bound to the merge with glab's --sha, so a push landing between the read and the merge fails the merge instead of landing commits nothing verified. Recorded metadata is never the authority for any of this: a rebase moves the head and leaves a recorded value stale, so a recorded head that disagrees with the live one is reported rather than trusted, and the recorded value is read before the recording step because that step drops a GitLab head it cannot resolve. * no-mistakes(review): reject bundled -R clusters and make tool-absence cases host-independent * no-mistakes(test): state authorised GitHub narrowing of bundled -R guard This branch NARROWS GitHub behaviour. The narrowing was authorised deliberately rather than slipping in by accident, and it applies to both providers, GitHub and GitLab alike, because a script that guards one provider and not the other is a trap for the next reader. What bin/fm-pr-merge.sh now refuses is extra merge arguments containing a bundled short-option cluster that includes R, for example "-dR other/repo". The forge CLIs expand such a cluster one character at a time, so it carries "--repo other/repo", and that later value wins over the repository the URL named. Before this change, "fm-pr-merge.sh <task> <github-url> -- -dR other/repo" reached "gh-axi pr merge 12 --repo example/repo --squash -dR other/repo" and exited 0 with pr= recorded and the merge poll armed. It now exits 1 with "extra merge arguments must not override the repository", records nothing, and invokes no forge merge command. Every other GitHub invocation is byte-identical to the base commit. Closing that hole honours the existing rule rather than departing from it. The file header already forbids --repo and -R because the repository must come only from the URL, so a bundled cluster carrying a repository override was never legitimate behaviour to preserve: it was that guard being evaded. Redirecting a merge to a repository the URL does not name is exactly what the guard exists to prevent. The refusal is already pinned on both paths by the existing case test_bundled_repo_override_args_refuse_before_recording in tests/fm-pr-merge.test.sh. On GitHub ("-dR wrong/repo") and on GitLab ("-yR https://other.example/g/p") it asserts exit 1, the refusal wording, no pr= in the task meta, no armed merge poll, and no forge merge command invoked, with a control case proving a cluster that carries no repository override still reaches the forge. No duplicate assertion was added. Both assertions were confirmed to have teeth by narrowing the guard back to a bare -R and watching each path fail. This commit carries no file change: the guard and its coverage landed in 614853d, and this message exists so the pull request description states the narrowing. * no-mistakes(document): fix README pointer for GitLab watch and merge doc * no-mistakes: apply CI fixes
kunchenguid#2788) * no-mistakes: apply CI fixes * fix(bin): drop a private record citation and narrow the review rule Three corrections to the spoken interface that landed in kunchenguid#2767, plus one fix carried over from that branch after its pull request had already been merged. The confidentiality fix. The module docstring of bin/fm-voice-relay.py cited a private, gitignored fleet record by exact path and section number. That widens what this public repository points at, and it cannot resolve for any reader here, because the path has never been in the repository. Both traps it pointed at are already described in full in the list immediately below it, and docs/voice-relay.md carries the same two for operators with no citation at all, so the pointer is removed and no claim is weakened by losing it. Two comments that referred to "the survey" as though it were something a reader could open are reworded the same way. Neither exposed a path, so that half is comprehensibility rather than confidentiality. The review rule. .greptile/rules.md is kept, because its conditions are right and deleting it would leave the next reviewer to re-litigate a decision already argued out. What was wrong with it is narrower than its existence: it read as settled repository policy, when whether VISION.md itself should be reconciled is an open question belonging to the captain. One sentence now says so, and says that the conditions listed below it are what the interpretation depends on. That narrows the claim rather than widening it. The carried-over fix. The first commit on this branch is 7f98e79 from fm/voice-relay-build-v4, taken verbatim rather than rewritten. It closes the window where a transport failure was recorded and then erased, so a run could be emitted as answered false with relay_error null. That matters more than it looks: relay_error is the field that keeps an infrastructure failure from being averaged into a latency figure, so the failure mode is a dead connection wearing the costume of a slow reply. It landed fifteen minutes after kunchenguid#2767 merged and so never reached the default branch. * no-mistakes(review): name a reason on every unanswered-turn close path * no-mistakes(review): guard the downlink body and pin frames to their turn * no-mistakes(review): attribute reply audio to its own turn and tell endings apart * no-mistakes(review): tell a cut-short reply from an unanswered turn * no-mistakes(review): discard reply audio arriving after the output closes * no-mistakes(review): count discarded reply audio on the speaker path too * no-mistakes(review): keep a reason off a turn already answered in full * no-mistakes(review): say a reset cut a reply short, not that none arrived * no-mistakes(review): read one turn's audio count once, and hush a tidy exit * no-mistakes(document): fix stale session-end relay_error claim in voice-relay guide
…unchenguid#2811) A pi worker parked on an interactive prompt - a permission dialog, a question menu, a trust dialog - reports agent_status=blocked, because it is waiting on a human keystroke. Pi draws that menu above its separator pair, so the composer region between the rules is blank and structure alone looks like a free composer. _fm_composer_pi_verdict admitted blocked alongside idle and done, so the shared classifier reported an affirmatively empty composer for exactly the pane where typing is unsafe. Every "is it safe to type here?" consumer reads that verdict and proceeds only on an affirmative empty, so both are told yes on a parked prompt: the away-mode injection guard in bin/fm-supervise-daemon.sh, and fm-send's pre-type refusal. The keys then answer the menu instead of composing a message - the highlighted default is selected, the text is discarded, and the record attributes a decision to a human who never made it. blocked now defers to unknown, which every consumer already treats as fail-closed. idle and done still prove an empty composer, so ordinary steering is unchanged, and Cursor is unaffected because its always-blocked panes never reach this pi-only branch. Regression coverage lands first at both levels: the verdict owner (a blocked pi defers) and the herdr adapter (a parked pi prompt is not an empty composer).
…2849) * fix(bin): require a clone root before fleet-sync touches a project Git repository discovery walks upward, so `git -C projects/<dir>` on a plain directory nested under projects/ resolves to the enclosing repository - in a firstmate home, the firstmate checkout itself. fm-fleet-sync.sh guarded its candidates with `rev-parse --is-inside-work-tree`, which such a directory passes, so every later git call read, pruned and fast-forwarded firstmate's own default branch and reported it under the project directory's label. A running session's AGENTS.md changed underneath it, and the report named a project that had nothing to do with the change. Require each candidate to be the root of its own work tree before any other git command: compare `rev-parse --show-toplevel` against the directory's own physical path. Both sides are physical, so a symlinked clone still compares equal. Anything else is skipped by name, naming the repository that would have been touched, and bootstrap relays that as a FLEET_SYNC line. Regression coverage reproduces the wrong-repo fast-forward against a home nested inside another repository, in both the whole-fleet and single-project forms, and pins that a symlinked clone dir still syncs. * no-mistakes(review): Keep enclosing fixture clean during clone-root regression
* fix(procevent): retry a transient Lavish poll interruption quietly
A live Lavish listener can be cut short by the server with exactly
error: Lavish Editor poll response was interrupted
code: SERVER_ERROR
while the session's marks remain available. Firstmate registered raw
`lavish-axi poll` output, so the generic process-event runner captured
that transient response as a result and woke the whole fleet over what is
really an internal retry.
The Lavish adapter now registers its own listener command, which reruns
the published blocking poll up to 12 times at 5 second intervals for that
one exact two-line response. The match is deliberately narrow: real
feedback, ended and missing sessions, any other SERVER_ERROR, and the same
interruption still standing once the bound is spent all pass straight
through and are captured and announced as before. The retry is a Lavish
fact, so the generic runner stays adapter-agnostic.
`FM_LAVISH_POLL_RETRY_DELAY` is a bounded 0 to 60 second override for the
interval only, refused rather than rounded when malformed, so a test can
exercise the real bound without waiting it out.
* no-mistakes(review): Harden Lavish retry matching, validation, and cleanup
* no-mistakes(review): Bound Lavish retry staging and stabilize regression
* no-mistakes(document): docs: explain Lavish retry adoption
* no-mistakes(lint): Restore Lavish trap ShellCheck suppression
… gate (kunchenguid#2838) The unguarded Herdr declaration quoted `{TASK}` in its own prose while the scaffold instructs firstmate to replace every `{TASK}` placeholder. The documented global replace therefore spliced the whole task body into the middle of the safety gate's sentence, silently destroying the one contract that exists precisely because the scaffold cannot inspect the task text. Reword the gate to refer to the task text filled in above, leaving the placeholder only at its genuine fill site. Rewording rather than renaming the token keeps the unfilled-charter guards in fm-home-seed.sh and fm-remote-home-seed.sh working unchanged. Add a regression test that performs the documented global fill on ship and scout scaffolds and asserts the body lands once and the gate survives.
…tat form (kunchenguid#2837) The writer lock's stale-lock branch read the lock's mtime with `stat -f %m ... || stat -c %Y ...`. On GNU coreutils `-f` is filesystem stat, so it consumed the format string as a path, complained on stderr, printed a partial filesystem dump (" File: ...") on stdout, and still exited 0. The GNU form in the fallback therefore never ran, and the following arithmetic evaluated the word `File`, aborting the writer under `set -u` with "File: unbound variable". fm-teardown.sh died there after returning the worktree, leaving state/<id>.meta, .status, .busy-gen, .busy-state, .busy-state.lock/ and .turn-ended behind. The surviving metadata kept the watcher monitoring an endpoint whose agent was gone, so a finished task produced stale wakes forever, and every re-run died identically because the abandoned lock was never broken. Detect the platform once and pick the right stat form, the pattern bin/fm-watch.sh already documents, and treat any non-numeric result as "just created" so a future portability surprise degrades to a lock-timeout refusal rather than killing teardown mid-way.
* fix(stow): give memory decay a per-pass horizon so the clock fires The tiered decay clocks were wall-clock only, while admission is per-pass: each /stow admits the findings that pass produced. In a home that stows daily those two rates diverge by the stow cadence, an entry the fleet keeps exercising never reaches 30 days unreinforced, and memory only grows while the pass reports decay evaluated. Give each dated marker an optional unreinforced-pass counter and make both tiers stale at whichever horizon comes first: 10 passes or 30 days for aging, 3 passes or 7 days for perishable. Reinforcement clears the counter and nothing else does, so the existing evidence-based restamp rule stays the only way an entry renews its lease. An absent /N means zero, so entries that stay exercised carry no extra marker bytes, and a rarely stowed home keeps its current behaviour through the unchanged date horizon. * no-mistakes(document): Align stow workflow with dual decay clocks * fix(stow): make the per-pass decay horizon opt-in The unreinforced-pass horizon shipped as a new default archival cadence, which is a product default rather than a restoration of the existing wall-clock contract. Keep the 30-day and 7-day horizons as the only default clock, and put the 10-pass and 3-pass horizons behind an explicit opt-in: config/stow-pass-horizon for the firstmate home, and the file's own header pointer for the public skill. With the opt-in absent no counter is written and no counter is read, so a home that does not ask for it decays exactly as it does today. * no-mistakes(review): Preserve frozen counters and correct archive provenance
…artup (kunchenguid#2876) tests/fm-watcher-lock.test.sh passed in isolation but failed intermittently under full-suite and ambient concurrent load. bin/fm-watch-arm.sh computes its confirmation deadline immediately after forking the real child watcher, so the child's entire fork, exec, lock acquisition and beacon publication has to land inside that wall clock. Two cases shrank that budget to one second, leaving a two-second window for work measured at 3.1-4.9s under CPU oversubscription, so the arm honestly reported "FAILED - no live watcher with a fresh beacon" and their premises collapsed. A third case ran on the production budget, but its child must also execute a registered check before exiting: measured at 1.9-2.3s idle and 9.1-13.1s under load, against an 11s budget. The two cases that must confirm a real child now hold the arm to production's own budget instead of a shrunken fixture one, the immediate-wake case gets an explicit budget with headroom over its measured loaded cost, and the two waits for the arm's typed failure are sized off the largest production default rather than a fixed eight seconds. No bin/ change and no default behavior change: the lock's fail-closed semantics, SIGSTOP handling, stale-heartbeat detection and the arm's typed failures are untouched. Verified 4/4 green at 3x CPU oversubscription (loadavg 75-80) after 3/3 red before the change, and CONTRIBUTING.md records the convention.
* fix(bin): order discovered tool installs by the shell's own expansion fm_remote_job_compose_operator_path built the asdf and mise install directories with `compgen -G`, which does not sort. Bash sorts glob matches in pathexp.c, on the shell's own pathname-expansion path only; `compgen -G` reaches the same glob_filename through pcomplete.c, which sorts nothing. On bash 3.2 (macOS /bin/bash) and every bash before 5.3 that handed the composition raw readdir order, so which install of a multi-version tool a remote job resolved was decided by directory order on disk rather than by this composition. Expand the globs at the call sites and let the function take the matches, so the composition and the documented portable-PATH contract are the same operation. Quoting the account home at the call site also stops a home whose name contains glob metacharacters from being reinterpreted. The colocated regression pins both the order and the mechanism: bash 5.3 moved sorting into the glob library, so an order-only assertion cannot see the defect there. * no-mistakes(review): Remove source-reading PATH regression guard
…2848) * fix: surface stalled secondmate queues and wake handoffs * no-mistakes(review): Make handoff wakes retryable and stall alerts crash-safe * no-mistakes(review): Prevent duplicate handoff wakes and cover remote delivery * no-mistakes(review): Serialize local handoffs and preserve pre-move wake intent * no-mistakes(review): Serialize teardown with handoffs and retain remote wake confirmation * no-mistakes(review): Reconcile correlated handoff wake delivery after crashes * no-mistakes(review): Keep failed wakes retryable and isolate stall receipts * no-mistakes(review): Reset known-undelivered wake attempts for durable retries * no-mistakes(review): Refuse duplicate sends for unresolved delivery attempts * no-mistakes(review): Atomically restore retryability after reconciled send failures * no-mistakes(review): Serialize delivery confirmation with reconciliation * no-mistakes(document): Document routed wake and stall supervision * no-mistakes(lint): Fix ShellCheck expansion and subshell warnings * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes(review): Retire stale wake state and defer pre-move wakes * no-mistakes(review): Secure markers, bind batches, and preserve teardown routes * no-mistakes(review): Preserve unresolved prepared wakes across unrelated handoffs * no-mistakes(review): Preserve prepared wakes before unrelated moving handoffs * no-mistakes(document): Document prepared wake batch ownership * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes(review): Make local wake retirement recoverable * no-mistakes(document): Clarify handoff recovery and teardown documentation
…guid#2856) * feat(bin): steer local tasks by durable inbox record plus constant doorbell Stage 1 (local steers) of the captain-adopted reframe in data/fm-send-reliability-reframe-s1/report.md: an ordinary fm-send text steer to a task recorded in this home is appended as a sequenced durable record under state/<id>.inbox/ and the terminal receives only one constant self-describing doorbell line, best-effort. The worker acknowledges by moving the record into handled/; the watcher re-rings an unacknowledged message on an idle pane and escalates once as an ordinary stale wake. --resolve-key closes decisions at enqueue time, because the durable enqueue IS delivery to the task's record. bin/fm-task-inbox-lib.sh owns the record format, doorbell line, and re-ring ladder. The typed plane remains for what must reach the terminal itself: lifecycle keys, harness-native slash and codex $-skill invocations, explicit backend targets, and the remote secondmate leg (unchanged until the remote inbox leg ships separately). The composer classifier is demoted from delivery proof to an advisory ring guard that skips only on a proven pending verdict. Verified live against claude, codex, opencode, pi, grok, and muse: each real worker read its record, acted, and acked with the mv (docs/verification/runtime-backends.md "Steering-inbox doorbell"). * docs(verification): flag the grok 1.0.5 composer-matrix staleness observed by the doorbell run * test(captain-hold): read the chat-channel answer from the durable inbox record * test: migrate fm-control's marker contrast to the inbox record and fix macOS wc padding in the tool-update suite * no-mistakes(review): Harden inbox locking, teardown races, and acknowledgements * no-mistakes(review): Serialize watcher actions with inbox acknowledgements * no-mistakes(review): Bound metadata locking and tighten acknowledgement rechecks * no-mistakes(review): Preserve exact inbox bytes and harden delivery recovery * no-mistakes(review): Harden watcher bookkeeping against concurrent inbox teardown * no-mistakes(document): Update inbox and typed-plane documentation * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * revert(pipeline): keep parser-native secondmate marking and the both-failed exit out of stage 1 The CI monitor's fix changed the secondmate marking contract for parser-native invocations (appending the marker after the text) and softened the both-commit-and-marker-failed branch to exit 0. The merge authority ruled the marking question out of scope for this stage-1 transport PR (follow-up: fm-send-secondmate-harness-invocation-r1) and ruled the both-failed case a loud nonzero local failure. Restore both, keeping the monitor's legitimate migrations and hardening. * no-mistakes(document): Document inbox and typed-plane boundaries * no-mistakes(document): Scope backend transport docs to typed plane * no-mistakes(document): Clarify inbox attempt-budget documentation * no-mistakes: apply CI fixes * fix(send): the durable record alone governs the inbox exit status Captain-refined ruling on the F2/Greptile finding: the durable inbox record is what delivers the steer, so pending-reply bookkeeping trouble after a successful enqueue never exits nonzero - a resend-inviting status would make automated callers enqueue the delivered instruction again under a new sequence. With the recovery marker stored the watcher reconciles silently; with the commit and marker both lost the send surfaces a distinct reply-tracking-degraded do-not-resend warning and still exits 0. Nonzero remains only where nothing was delivered (or a decision close needs its manual command). Regression: record durable + both bookkeeping writes lost -> exit 0, one record, no duplicate. * no-mistakes(review): Preserve inbox ordering with drain-all doorbells * no-mistakes(review): Surface unwritable inbox ladder bookkeeping * no-mistakes(review): Silence ladder failures after inbox acknowledgement * no-mistakes(document): Update steering inbox documentation * no-mistakes: apply CI fixes
* feat: add fast local lint mode * fix: preserve complete fm-lint help * fix: isolate fast lint mode * no-mistakes(document): Clarify lint mode documentation ownership * no-mistakes: apply CI fixes
…#2901) * feat(bin): deliver remote secondmate steers through durable task inboxes Stage 2 of the inbox+doorbell steer channel (stage 1: kunchenguid#2856). A remote secondmate steer now crosses fm-on.sh as a durable record written idempotently into the remote home's steering inbox plus a best-effort remote doorbell, and the last typed-payload steer transport is deleted: - fm-remote-secondmate-control.sh cmd_send writes the record via the new fm_task_inbox_write_idempotent and rings the doorbell; it no longer types the payload through an inner fm-send at an explicit pane target. - fm-send.sh routes every remote text steer (harness-native included, which marking already reduced to chat) onto the remote inbox leg, retries the identical leg once on ssh 255, closes --resolve-key decisions at enqueue for remote too, and preserves a marked request's reply expectation when completion stays unknown. The exit-3-as- delivered remap, the 255 do-not-resend trap, and the remote typed submit block are removed. - fm-task-inbox-lib.sh owns the idempotent enqueue: an exact-body re-run lands on the existing record, handled or not, so an ambiguous transport can always be safely re-run. - Tests pin the new contract end to end (record + doorbell + no typed payload across ssh, one-record idempotence under an ambiguous transport, enqueue-time decision close, loud real failures, and the deleted typed-payload behaviors gone), and AGENTS.md plus docs/remote-secondmates.md describe the remote leg's new semantics. * no-mistakes(review): Harden remote inbox delivery against lifecycle races * no-mistakes(review): Enable correlation-preserving remote steer resends * no-mistakes(review): Fail closed on stale correlation resends * no-mistakes(review): Include home context in remote resend commands * no-mistakes(review): Lock and revalidate remote parent routes * no-mistakes(document): Clarify remote steer retry documentation * no-mistakes: apply CI fixes
* wip: forked supervision on Pi (checkpoint before docs) * fix(pi-branch): harden mirror delivery, fallback encoding, and session replacement Peek-then-shift mirror flush so a failed append retries instead of dropping; durable mirror cursor commits only after delivery into the branch; the main fallback wake is operational-encoded like every watcher injection; session_shutdown quiesces the generation and session_start re-arms, so /new and /resume no longer kill the branch permanently. Registers the extension in the strict typecheck, adds the dispatch handshake test, the branch extension suite, the bash-level regression suite, the session-start replay test, and the opt-in real-SDK live guard. * test(fixtures): carry the branch-dispatch lib and lease lib into isolated fixtures The watcher extension now imports lib/fm-branch-dispatch.ts and fm-teardown sources fm-lease-lib.sh, so every fixture that copies or symlinks those files in isolation gains the new sibling. * no-mistakes(review): Prevent shutdown wake loss and serialize lease claims * no-mistakes(review): Durably hand off wakes and retain portable leases * no-mistakes(review): Require durable reports and clear disposed branch leases * no-mistakes(review): Enforce per-wake outcomes and quiescent lease cleanup * no-mistakes(review): Require wake acknowledgements and tighten branch lifecycle boundaries * no-mistakes(review): Require complete acknowledgements and replay cleanup failures * no-mistakes(review): Bind supervision to lock ownership and durable delivery * no-mistakes(review): Activate branch lazily after session lock acquisition * no-mistakes(review): Preserve undelivered mirror context across extension rebinds * no-mistakes(review): Acknowledge startup replay only after main delivery * no-mistakes(review): Isolate replay metadata from untrusted digest content * no-mistakes(review): Reject duplicate reports for active wake sequences * no-mistakes(review): Retain failed fallbacks and deduplicate outcome replay * no-mistakes(review): Deduplicate durable outcomes and cache delivery receipts * no-mistakes(review): Anchor wake sequence matching to outcome fields * no-mistakes(document): Clarify Pi supervision durability contracts * no-mistakes(lint): Fix ShellCheck issues in branch supervision scripts * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * refactor(pi-branch): collapse to confused-agent-grade guards per captain decision Captain decision A: the lease/actor guards target the CONFUSED-AGENT threat model bin/fm-gate-refuse-lib.sh already documents; adversarial-grade separation is impossible in the shared-process design and is filed as separate follow-up work. Rip out the machinery that chased it: the generation fence and shell-provenance markers, the wrapper-tagged ancestry walks, guard auto-claim with per-script release traps, the pending-wake files and ack-receipt correlation (the durable wake queue already re-presents anything unacknowledged), the delivery-receipt store with contiguous cursor advancement, the session-start replay-metadata channel, and the branch tool quiescence counters. Keep the behaviors the board requires, each on its simplest implementation: lazy per-action session-lock ownership (cold start activates after the lock lands; a secondary session stays inert), mirror durability across extension rebinds via the durable cursor, replay-exactly-once from the one read cursor, the awaited operational-encoded fallback, per-generation stray-lease cleanup, session-lock-bound lease liveness (a recycled pid or a non-Pi home never honors a leftover lease), the loud accidental-override guards (readonly actor prelude, cross-actor claim refusal), and the role-partition refinements (no forced teardown, no direct relaunch for the branch). Default-on-for-Pi is unchanged. * no-mistakes(review): Enforce lock ownership and serialize lease mutations * no-mistakes(review): Synchronize guard cleanup and bind leases to lock owner * no-mistakes(review): Report outcomes before acknowledging durable wakes * no-mistakes(review): Restrict leases to Pi and instruct main claims * no-mistakes(review): Reject malformed lease locks and torn outcome tails * no-mistakes(review): Validate complete outcome tails before appending * no-mistakes(review): Guard branch side effects across session replacements * no-mistakes(document): Update Pi supervision durability and lease documentation * no-mistakes(lint): Suppress intentional nested-shell expansion warning * no-mistakes: apply CI fixes * fix(pi-branch): authorize lease releases by caller * fix(lint): break redundant source-analysis path in fm-lease-lib.sh fm-lease-lib.sh's lazy fallback source of fm-wake-lib.sh gave ShellCheck's --external-sources traversal a second path into an already 1540-line file that fm-send.sh and fm-teardown.sh also source directly, blowing up the recursive analysis past CI's lint timeout. Mark it a source=/dev/null analysis boundary, matching the existing fm-task-inbox-lib.sh convention. Also restores bin/fm-lint.sh and tests/fm-lint.test.sh to the shared serial-lint definition (dropping an unrelated parallel-sharding change that was itself hanging and masked this root cause). * no-mistakes(document): Correct lease caller-authorization documentation * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes
* feat(bin): parallelize session-start remote secondmate network sweeps Run per-secondmate liveness and convergence probes concurrently and overlap clone refresh, while replaying each mate's fail-closed diagnostic in original order. Ignore scratchpad* so untracked scratch no longer blocks remote sync. Co-authored-by: Cursor <cursoragent@cursor.com> * no-mistakes(document): Document parallel startup network sweeps * no-mistakes(lint): Fix empty environment assignment lint warning * no-mistakes: apply CI fixes --------- Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(tests): count declared-pause wakes without crashing on an absent queue The exited-declared-pause case counts queued stale wakes by handing state/.wake-queue straight to awk. A watcher that queues nothing never creates that file, and awk aborts on a missing path before its END rule runs, so the count collapses to the empty string. The next comparison then fails as an integer-expression error and surfaces as a wake flood with no number, hiding the real contract breach the following grep names. Read the queue the way the drain-count assertion at the end of this file already does: silence awk's open error and default an absent queue to zero. Applied to all four counts in this case, including the live external-decision gate pair whose queue an acknowledged drain can also leave behind. An absent queue now reports "did not use the bounded paused recheck", while a genuine flood still fails with its real count. Fixes kunchenguid#2628 * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes
…unchenguid#3950) * fix(bin): make every counted wake queue row presentable or retired A wake row could be counted as queued while no drain would ever present it, leaving the operator told to "drain them before anything else" by a command that printed nothing and offered no acknowledgement. Two independent paths produced that state. A row reserved by a live supervision-branch grant is excluded from a main drain by design, but fm-guard.sh counted the whole queue, so main was warned about rows only the branch could present, on every guarded command for as long as the grant was held. A row that lost its five appended fields or its numeric sequence can never be claimed, presented, or named by an --ack-through cutoff, yet it still counted as queued, wedging the queue permanently. The guard now counts only the rows the calling actor can itself present or retire, and a main drain retires unusable rows under the queue lock, reporting them in bounded escaped form before removal so the evidence survives for the separate row-generating defects. A retirement failure is reported loudly and never suppresses unrelated consumable work. A main drain whose remaining rows are all branch-held says so in one bounded line instead of exiting silently. Grant row-list and owner-record reads move into fm-wake-lib.sh so the drain, the grant publisher, and the guard share one implementation. Ownership is unchanged: a branch drain still touches nothing outside its grant and never retires a row, and main still cannot present or acknowledge an active grant. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GghGvsa4JDB1E5FuznX2i1 * no-mistakes(test): keep SIGTERM-safe arithmetic in wake queue retirement pass * no-mistakes(review): add guard advisory for branch-held wake rows * no-mistakes(document): document per-actor wake counting and unusable-row retirement * no-mistakes(ci): Addressed the Greptile P1 on bin/fm-wake-lib.sh:1840 ("Unreadable queue suppresses alarms"). Root cause: fm_wake_actor_pending_count inferred "the queue could not be counted" only from awk's printed output (`case "$count" in ''|*[!0-9]*) count=1`). That relies on awk aborting before its END rule when the input cannot be opened. An awk that reaches END after a failed open prints `0`, which the fallback accepts as a genuine count; both actor counts then read zero and bin/fm-guard.sh emits neither the queued-wake warning nor the branch-held advisory for a queue nobody proved empty. Fix (bin/fm-wake-lib.sh:1828,1835): both counting awk invocations now set `count=''` on a non-zero awk exit status, so the existing "cannot be counted => report a pending row" fallback is driven by awk's exit status instead of an implementation-defined detail of what it printed. No new code path or behavior; the pre-existing fallback just becomes unconditional. Comment updated to state why. Regression test (tests/fm-wake-queue.test.sh: test_uncountable_queue_still_raises_the_pending_alarm, registered in the run list): runs the real bin/fm-guard.sh against a non-empty, unreadable queue with a PATH-injected awk emulating an END-running implementation (prints 0, exits 2; execs the real awk otherwise) and asserts "queued wakes pending" is still emitted; disconfirming half asserts the same fake awk over a readable, provably empty queue stays silent. Fails on the pre-fix library ("not ok - a queue that could not be counted silenced the queued-wake alarm"), passes after. Verified locally: bin/fm-test-run.sh tests/fm-wake-queue.test.sh -> 0 failed; tests/fm-guard-stale-banner.test.sh + tests/fm-watcher-lock.test.sh -> 0 failed; bin/fm-lint.sh (ShellCheck 0.11.0 + actionlint 1.7.12) clean. Caveat reported honestly: on the awks available/known here (mawk locally, plus gawk and BWK/macOS awk, all of which treat an unopenable input as fatal and skip END) the alarm was not actually suppressed - I reproduced the unreadable-queue case and the warning fired. The change removes the code's dependence on that awk detail rather than repairing an outage observed on this platform. Both intent constraints still hold: every counted row remains presentable or retirable, and no alarm-suppressing path was added --------- Co-authored-by: Alex William <awilliam@v2202608403614505120.powersrv.de> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(bin): use /usr/bin/stat on Darwin to survive GNU stat shadowing * fix(bin): extend /usr/bin/stat prefix to Darwin stat -f sites added on main * test(bin): make fm-stat-shadowing skip visible on non-Darwin and isolate fm-watch state * ci: re-trigger after Chrome headless timeout in calm HTML export test * ci: re-trigger serial-5 after second Chrome headless timeout in calm HTML export test * test(bin): skip PATH-based stat fault injection on Darwin where stat is /usr/bin/stat * no-mistakes(document): Refresh stat and shard docs
…henguid#3945) The captain's attribution policy (no Co-Authored-By trailer, no Claude-Session link, no generated-with line) lives in Claude Code's `user` settings scope. A spawned worker's settings sources are not guaranteed to load that scope, so a launched worker could write attribution trailers into its commits and PR bodies regardless of the captain's own configuration. launch_template()'s claude case now carries the same policy ("attribution": {"commit": "", "pr": "", "sessionUrl": false}) directly in its inline --settings JSON, so every claude launch keeps attribution off independent of which settings scopes end up loaded. Tests assert the policy on the rendered launch command for both a crewmate and a secondmate spawn. Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
…nchenguid#3946) * fix(bin): derive the away-mode beacon grace from the poll cadence fm-turnend-guard.sh's away-mode branch required the watcher beacon to be fresh within the flat FM_GUARD_GRACE default (300s), but the daemon starts a fresh one-shot watcher only after it finishes handling the previous wake, and that handling can legitimately outrun a fixed 300s window under load (a slow registered check, a busy supervisor pane) with the daemon perfectly healthy throughout. That misread a live, correctly-cycling daemon as down and blocked the turn. Add fm_poll_derived_grace, the single owner of the max(300, FM_POLL + 60) formula, and have the away-mode branch, fm-claude-stop-autoarm.sh, and fm-watch.sh's own runtime beacon-staleness check all derive their default grace from it instead of the flat default. A dead daemon pid or a beacon older than that grace still blocks, so a genuinely lapsed away mode still alarms; every other check is unchanged. fm-claude-stop-autoarm.sh computed the derived grace into GRACE but its two fm-watch-arm.sh invocations called the wrapper bare, so the wrapper fell back to its own flat 300s default and could reject a healthy long-poll watcher. Both invocations now pass FM_GUARD_GRACE="$GRACE" through explicitly, and a new test proves a long FM_POLL with FM_GUARD_GRACE unset reaches fm-watch-arm.sh with the derived value. Also drops fm_last_activity_age, added alongside the derivation but never called anywhere in the tree; fm-inactive-reconcile.sh already owns that computation. * no-mistakes(review): Remove dead WATCHER_STALE_GRACE assignment in fm-watch.sh * no-mistakes(document): Update FM_WATCHER_STALE_GRACE default note for poll-derived grace --------- Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
…henguid#3904) * fix(procevent): bind a source runner to the session that owns it A process-event source runner is detached into its own process group so a persistent source survives the turn that armed it. Nothing bounded that detachment, so a runner could reparent to init and keep its blocking child - and every process that child spawned - running with nothing left to reap it. One such runner outlived its home for about a day; the cost was not the runner but the exec churn of the poll stubs under it, which stalled every fresh process launch on the host. Each runner now starts a small guard beside it, in a separate process group, that re-reads its home's process-event lease and stops the runner's whole process group once that lease can no longer be proved fresh. Every ordinary entry point an owning session runs refreshes the lease, and the watcher's reconcile cycle keeps it fresh in a live home; nothing a runner spawns can refresh it, so a source cannot certify its own owner. Scope is the owning state root and one runner generation, never a script or process name, so a live source in another home is untouched and a live home simply starts a replacement runner on its next cycle. The test scaffolding that starts real runners could not reap them either: the bearings-board and board-render suites tracked their homes in a shell array appended to inside a command substitution, so the array was always empty and every listener they started survived the run. Home registration moves to a `$$`-keyed registry in tests/lib.sh, which sweeps it from every cleanup path, now including HUP and QUIT, and the blocking fixture stubs stop themselves at a bound so an escaped one cannot keep spawning processes indefinitely. Adds a regression test that reproduces the orphan shape - a reparented listener with a live descendant tree under it - and proves the whole group and its process churn stop once its session is gone, that an identical listener in a home whose session is still there is untouched, and that retirement still reaches a reparented listener and everything under it. * no-mistakes(review): Bound source launches and fail closed on guard startup * no-mistakes(document): Document runner lease and storm containment * no-mistakes(ci): Fixed the Greptile watchdog finding: failed runner cleanup now retries on each watchdog tick instead of abandoning the orphaned process group. The shared state-root lease behavior remains unchanged because it is an explicitly accepted ownership policy. Verified with bash syntax checks, git diff checks, and the complete fm-procevent test suite * test(procevent): pin that an unprovable stop is retried, not abandoned The owner guard used to call stop_runner_pid and exit unconditionally, so a stop it could not prove - a descendant still finishing uninterruptible work outlives even the group signal, and an unreadable process identity proves nothing - left a still-running expired runner with nothing watching it. That is the best-effort reaping this mechanism exists to remove, and the fix that made the guard retry landed without a test holding it in place. The unprovable attempt is injected through the signal the real path actually reads: `ps` answers exactly one process-group query for the runner with a group it does not lead, which is how a stop that cannot be proved is reported, and every other call is the real command. The test also asserts that the injected attempt happened, so it cannot pass vacuously if the fixture stops arming. Fails against the exit-after-one-attempt guard, where the runner survives its expired lease, and passes once the guard retries on its check cadence. * docs(procevent): scope the no-self-refresh rule to confused-agent grade The runner-lease documentation asserted as an absolute that nothing a runner spawns can refresh the lease, so a source cannot certify its own owner. That overclaims what the inherited FM_PROCEVENT_IN_RUNNER marker actually enforces. The marker holds at confused-agent grade: a runner and its ordinary children inherit it and skip every refresh, which is exactly the accidental case this boundary exists for. A source that deliberately strips the marker from its environment can still refresh, so adversarial-grade unforgeability is explicitly out of scope and tracked as separate follow-up design work. This states the real scope in docs/configuration.md, which owns the operating contract, and corrects the two matching comments in bin/fm-procevent.sh. The process-event-sources skill keeps its cross-reference and gains one line in its never-to-be-claimed list so the overclaim is not reintroduced from the agent-facing side. The lease mechanism itself is unchanged. * no-mistakes(review): Fix process-event lease and launch pacing edge cases * no-mistakes(review): Scope launch pacing and clarify lease boundaries * no-mistakes(review): Reap leftover groups and use monotonic launch pacing * no-mistakes(review): Keep guards alive across runner PID reuse * no-mistakes(review): Use monotonic leases and simplify launch generation identity * no-mistakes(review): Prevent pacing identity reuse and bound reused-group guards * no-mistakes(review): Preserve active pacing state on failed registration * no-mistakes(review): Reap reused runner groups with registration evidence * no-mistakes(review): Avoid ambiguous group kills and encode pacing identities * no-mistakes(review): Abort kill escalation after runner identity reuse * no-mistakes(review): Gate group signals and prune stale pacing state * no-mistakes(review): Document bounded PID reuse signaling safety * no-mistakes(review): Align leaderless group ambiguity guidance * no-mistakes(review): Expire reboot stamps and preserve publication success * no-mistakes(review): Bind owner leases to physical state roots * no-mistakes(document): Clarify process-event lease and pacing contracts * fix(procevent): drop a platform-dependent post-TERM test assertion CI ran red on two lanes that the local gate could not see. Lint failed with SC2034 on two reads in cmd_owner_watchdog that destructure the state-root identity into five fields while using only the device and inode. Local changed-file mode suppresses the cross-file codes that need --external-sources, so the warning cleared the pre-push lint step and failed CI's full analysis, exactly as bin/fm-lint.sh's header describes. The unused fields now read into `_`. The behavior shard failed on this suite's own post-TERM assertion, which required the stubbed identity source to be consulted more than once. Whether that happens is platform-dependent: where the runner leader keeps waiting on its TERM-ignoring source child, the post-TERM check sees a live leader whose identity no longer matches, and where the leader dies promptly it sees a leaderless group carrying the same numeric id. fm_procevent_pid_state reaches that second verdict without consulting process identity at all, so the identity source is never read twice and the count assertion fails through no fault of the behavior. The case now asserts the invariant both forms share: retirement refuses, and the ambiguous group is not signalled. Scoping a mutation to this fixture and making the refusal signal instead confirms the case still fails, so dropping the count does not leave it passing vacuously. * no-mistakes(review): Prevent superseded runners recreating stale pacing stamps * no-mistakes(document): Document pacing and ambiguity boundaries * fix(procevent): retire under the recorded identity source and state the home-scoped lease The reused-group case started its runner with the proc-root override in place, so the runner recorded a ps-derived identity, then retired it without that override. Where /proc exists the retirement read identity from a different source than the one recorded, the guard correctly refused an identity it could not confirm, and the case failed on Linux while passing on macOS. It now retires under the same source, and clearing the stub marker first turns that cleanup into the complementary assertion: once the ambiguity is gone, retirement reaps the whole group instead of leaving it behind. The lease prose claimed a runner is bound to the session that owns it, while the mechanism binds it to the home. That gap is what makes a replacement session or an inspection command look like a defect: any activity in the same home refreshes the lease. The granularity is deliberate, because a persistent source is meant to outlive the session that armed it, and binding a runner to that session would stop the sources this mechanism exists to keep running. A runner whose source is no longer wanted in a live home is stopped by reconcile when that source is retired, independently of the lease, so the lease is the backstop for a home that is gone - the torn-down sandbox this change bounds - and the residual is recorded as a known limit. * no-mistakes(review): Rate-limit polls and skip superseded runner launches * fix(procevent): build the claim-only sweep case as a runnerless owned claim A superseded generation now observes the registration-identity mismatch, self-retires, and releases its claim, which is the behavior we want: it clears its own residue rather than leaving a claim with no runner for the home sweep to find. The claim-only sweep case was built by deleting a registration out from under a live runner, which used to leave that runner in place. It now makes the runner retire itself, so the sweep raced that exit and retired one source or two depending on which won. The case failed three runs in four, alternating between a preflight-count failure and `attempted=1`. It now builds the state it means to test: kill the runner's group so it cannot run its own cleanup, assert the owned claim survived that kill, and only then drop the registration. Coverage is unchanged - a runnerless owned claim must still be swept - and the result no longer depends on whether the runner had exited yet. Three consecutive runs pass. The superseded exit also skipped the runner-marker cleanup the normal path performs. The marker is written before the launch floor is waited on, and a home sweep counts a marker with no owned claim as a preflight failure, so exiting without clearing it would make that home refuse to sweep. `FM_LAVISH_POLL_RETRY_DELAY= ` trips SC1007 under the full analysis CI runs, though not under the changed-file mode the pre-push gate uses. * test(procevent): retire a quiet reparented listener instead of racing a storm Explicit retirement was exercised against the spawn-churning stub, which made it nondeterministic. Retirement refuses rather than signalling when it cannot confirm the runner's identity, that identity is read through `ps`, and the stub's 0.1s spawn loop starves that read often enough that a single attempt is a race - the suite failed on this case roughly one run in four, reporting `cannot confirm runner identity; source remains registered`. The refusal is correct: it is the documented preserve-for-retry contract, and a separate case already asserts it. So this is a fixture problem, not a behavior problem. The storm is still covered where the evidence for it lives. The owner-loss home keeps the churning stub and still asserts its tick log stops, which is what proves the churn ended rather than one pid going away. The retirement home never asserted ticks; it only ever read the descendant pid, so the spawn loop bought this case nothing while costing it determinism. It now uses a quiet stub that still reparents and still holds a real descendant in its process group, so the assertions are unchanged: retiring the source must reap the reparented listener's whole group and the descendant under it. Four consecutive runs pass. * no-mistakes(review): Serialize registration replacement through source child launch * chore(no-mistakes): require honest test-step scenario marking The test step recorded scenarios as passing that were only reached through a stubbed dependency or the executable suite, and its validator refused them, because `pass` asserts a scenario was verified against the real live product. That refusal is correct, so the fix is to mark honestly rather than to weaken the gate: a scenario driven live stays a pass and cites its live transcript, while one reached only through a stub or the suite is recorded as untested with the reason and a pointer to its executable coverage. Untested scenarios are reported rather than treated as failures, so real coverage stays visible without claiming verification that did not happen. The instruction also forbids dropping a scenario to avoid marking it untested, since that would hide the gap instead of stating it. * no-mistakes(review): Remove unrelated test scenario policy * no-mistakes(document): Clarify process-event home lease documentation * no-mistakes(document): Correct owner guard failure wording
…nguid#4037) * test(lib): set fixture mtimes through one portable epoch helper On macOS the visible symptom was ONE red case in the turn-end guard suite. The actual damage was TWO cases that had quietly stopped testing their subject. The red one was the harmless half - people read one red case as one broken thing, and here that intuition is wrong. `touch -d @<epoch>` is a GNU extension; BSD touch rejects it outright and leaves the file at its current mtime. So on macOS the three away-mode beacon cases never aged their beacon at all. The 400s case exists to pin that 400s is stale under the flat 300s default but fresh under the poll-derived grace (660s at FM_POLL=600). Deleting the grace it guards (FM_POLL=60, so max(300,120)=300) and re-running proves what it was worth on this platform: pre-fix input (beacon left at now): ok - passes with the feature DELETED post-fix input (beacon 400s old): not ok - expected exit 0, got 2 It was green while measuring nothing, and could not have caught a regression in the grace it names. Only the 700s case broke loudly. `touch -t [[CC]YY]MMDDhhmm[.SS]` is POSIX and both platforms accept it, so the only host-specific step left is formatting the epoch into that stamp, which date(1) spells two incompatible ways. fm_touch_epoch in tests/lib.sh owns that probe once and fails loudly rather than leaving an unset timestamp behind - the failure mode that caused this. Verified on BSD touch/date here and on GNU coreutils 9.7 in a container. Three real sites, and one consistency change - not four fixes. The stale destination lock in tests/fm-remote-backlog-handoff.test.sh was never a defect: its `uname = Darwin` branch made the `touch -d` line unreachable on macOS, and BSD touch accepts that space-separated form anyway. `touch -t` takes the date directly on both platforms, so the branch goes rather than standing as a second copy of the same platform assumption. Known limit: the third case (away mode off) is only HALF recovered here. It now receives the input its name claims, but it is still insensitive after this fix - its verdict is identical with a 0s and a 400s beacon, because the fixture records a daemon lock and no watcher lock, and with away mode off the daemon lock proves nothing. Not fixed here; tracked separately, with the requirement that any fix be shown to FAIL when the protection is removed. FULL SUITE ON macOS: 189 scripts, four red, none of them this change. The turn-end guard and remote-handoff suites are clean. Attribution was established by running the four failures at the base commit and at this head on an idle machine, because base-idle against head-under-load moves two variables at once: script head/loaded base/idle head/idle verdict fm-calm-pi-extension red red red pre-existing fm-backlog-atomicity red red red pre-existing fm-procevent red red red pre-existing fm-startup-network red green green cause unestablished, load-sensitive under a full run Reported, not fixed. fm-calm-pi-extension deserves its own note: it FAILS because Chrome is absent instead of declaring the capability it needs and standing aside, so its verdict is about the machine rather than its subject - the same family as the defect above, with the red at least announcing itself. Neighbouring class, reported not changed: `file_mode()` - a verbatim `uname = Darwin ? stat -f %Lp : stat -c %a` - is copy-pasted across at least five test scripts plus a `reread_mode` variant, and epoch-mtime reads are open-coded as `date -r … || stat -c %Y` in three more; same one-owner shape as the defect above. `git init` without `-b main` depends on the host's init.defaultBranch in several scripts (branch-name case, tracked elsewhere). timeout, sha256sum and sed -i uses are all correctly guarded where checked. Observed while building the check rather than the fix: the first watcher I wrote to wait for the suite matched its own command line, so it was waiting on its own existence and could never fire. Same shape as the cases above - machinery answering confidently about something other than its subject, by including itself in the evidence it was meant to judge. The file sentinel it was replaced with cannot be produced by the observer that reads it. * fix(review): Pin fixture timestamps to UTC across DST transitions * fix(document): Clarify shared fixture suite coverage
…ision (kunchenguid#4038) * fix(pi): preserve native Codex effort and guarded supervision * test(pi): identify native compatibility guard versions * no-mistakes(review): share native-main follow rule between build and picker * no-mistakes(document): document native progress marker and ultra effort owners * no-mistakes(ci): Failing check "Behavior portable serial 1" was caused by this PR. The shard ran the default-on live guard tests/fm-pi-branch-responsiveness-live-e2e.test.sh, whose idle arm loads .pi/extensions/fm-branch-supervision.ts into a scratch project with a fixed list of copied libs. This PR added `import { registerFirstmateTool } from "./lib/fm-native-contract.ts"` to that extension and updated every other loading fixture's copy list, but missed this guard. Pi 0.85.1 therefore refused to load the extension ("Cannot find module './lib/fm-native-contract.ts'"), never drew its TUI, the test failed with "Pi 0.85.1 never drew its TUI in the idle arm", and the job hit its 20-minute cap. Fix (one line): added fm-native-contract to the lib copy loop in tests/fm-pi-branch-responsiveness-live-e2e.test.sh. Swept all other suites referencing fm-branch-supervision.ts / fm-primary-pi-watch.ts; the remaining ones without the new lib only hash, path-reference, or string-match the files and do not load them into Pi, so no further fixture changes are needed. Verification: reproduced the mechanism against the installed Pi 0.85.1 by building the lab copy with the old lib list (load error as above) and the fixed list (loads cleanly). tmux is not installed on this machine, so the live guard itself gate-skips locally ("skip: live: tmux absent") and could not be run end to end here; CI (which has tmux and Pi) will exercise it. shellcheck is clean on the edited file --------- Co-authored-by: Talon Stark <talonstark@gmail.com>
…guid#4041) * fix(herdr): step around a stale client the running server refuses A remote host can carry a self-updated herdr in ~/.local/bin beside a package-managed one, and the fixed remote-job PATH resolves ~/.local/bin first. After the server upgraded to 0.9.0 (protocol 22) the stale 0.8.2 client (protocol 20) was answered with protocol_mismatch on every command, which the read classifiers folded into `unreadable`: the live remote secondmate read unknown, every doorbell into it failed, and both the spawn and relaunch recovery paths refused, so the defect trapped itself. The adapter's session-scoped CLI wrapper now recognizes that refusal, reads status per session from each distinct herdr on PATH, adopts the first one the running server reports compatible, retries once, and keeps it for the process. The happy path makes no extra call and no other failure reselects. An endpoint that still reads unreadable names the refused client, both protocols, and the fix on stderr; the remote state read, fm-crew-state, and the launch refusal carry that reason, and fm-remote-doctor reports the selected client and rebinds the launch agent to it. Regression coverage: fake two-client hosts in the herdr unit suite, the doctor suite, the crew-state remote arm, and the real host-local control script in the remote lifecycle e2e; the real-herdr smoke refreshes the status shape the selection reads. * no-mistakes(review): Reselect Herdr client after every protocol mismatch * no-mistakes(review): Remove unrequired Herdr diagnostics and launch-agent rebinding * no-mistakes(review): Scope cached Herdr clients to their selected session * no-mistakes(review): Restrict herdr client selection to reactive CLI calls * no-mistakes(document): Document session-scoped Herdr client reselection * no-mistakes(document): Clarify Herdr client selection documentation * no-mistakes(ci): Updated the trusted fm-remote-doctor.sh SHA-256 in bin/fm-remote-entrypoint.sh after the PR changed the doctor, restoring git-unavailable bootstrap authentication. Verified tests/fm-on.test.sh, tests/fm-backend-herdr.test.sh, bin/fm-lint.sh, and git diff --check all pass
* feat(afk): record the away posture and its lifecycle (phase 1) Away mode becomes a posture of the one supervision session, recorded in state/.afk-contract by the new bin/fm-afk-contract.sh: the one owner of the record schema, the mandate-clause grammar and compiler, refusal naming the missing part, the read-back rendering, the entry announcement (hold-for-return only, no phone channel), and the archive at return. This release records clauses and does not execute them; the announcement and return brief say so. bin/fm-afk-launch.sh gains propose and confirm, confirms the record before any daemon launch, refuses to launch the daemon on Pi and pi-signed, and archives the record last on stop. bin/fm-afk-return.sh snapshots supervisor health before shutdown, renders the return brief (health, mandate, waiting on the captain, could not fix, handled, cost) from the archived record, the outcome store, the held set, and the status logs, and shrinks the blocker gate to what the away session could not fix. While the record exists the watcher and the daemon never recheck an item held for the captain. Declared external waits get a four-hour default cadence and honor `until <UTC ISO 8601>` on the paused line, in both postures, bounded by FM_PAUSE_UNTIL_MAX_SECS. The /afk skill, AGENTS.md's layout and away-mode stub, the session-start digest, and the architecture, Pi branch, configuration, and scripts docs describe the record. The Pi/Herdr e2e now proves the no-daemon posture on a real Pi primary; its verification record carries the 2026-09-08 run. * no-mistakes(review): Fix AFK confirmation, grammar, waits, and return gating * no-mistakes(review): Harden AFK authority and posture lifecycle * no-mistakes(review): Preserve AFK history and tighten authority grammar * refactor(afk): record clause fields with no natural-language parser By the captain's mandate the away-posture record keeps no static parser that tries to understand natural language. A mandate clause is now given as explicit fields (--action, --object, --when, optional --stop) that bin/fm-afk-contract.sh records verbatim. The structural check asserts only that the action, object, and precondition fields are present and that the action is a listed verb; whether a precondition holds is the supervision session's judgment at execution time in a later phase. The never-set stays as a forbidden-concept safety scan: fields mentioning credentials, passwords, logins, legal or financial acceptance, payments, invoices, one-time codes, or an attended prompt are refused, matched at token prefixes after punctuation normalization so compound and plural spellings are caught. The red-check grammar, class-word rejection, unconditional-word detection, clause-reference resolution, and condition aliases are removed. --words-file keeps the captain's words verbatim, trailing newline included. The skill, docs, launcher help, and tests describe the field form. * no-mistakes(review): Preserve AFK words and tighten safety refusals * no-mistakes(review): Preserve clause bytes and honor declared waits * no-mistakes(review): Harden deny-list and gate unreadable outcomes * no-mistakes(review): Demote never-set scan and clarify authority * no-mistakes(review): Gate return on unreadable held and status data * no-mistakes(review): Validate posture archives and enforce Pi detection * fix(afk): make the never-set a non-refusing flag and keep return fail-safe Per the captain's decision the never-set scan is a coarse best-effort flag, never a refusal and never the gate: a clause naming a listed concept is still recorded with a flag the read-back, announcement, and return brief show, and the scan matches listed terms exactly or with a plain inflection at punctuation-delimited token boundaries, so unrelated names such as ping-service or tokenize-worker are never flagged and joined compounds remain a documented miss. Authoritative never-set and forbidden-action enforcement is the supervision session's judgment at execution time in phase 4. A replacement copies the superseded record through a temporary name and renames it atomically so a failed copy leaves no partial archive, the record owner gains validate and flags subcommands, and the return keeps catch-up gated when a superseded archive cannot be read. * no-mistakes(review): Harden AFK record validation and return reconciliation * no-mistakes(review): Harden AFK record validation and simplify commands * no-mistakes(review): Harden mandate validation and retain missing records * no-mistakes(review): Refuse blank explicit mandate stops * no-mistakes(review): Recover restored posture epoch before return * no-mistakes(review): Prevent return brief status symlink reads * no-mistakes(document): Refresh AFK posture documentation * no-mistakes(ci): Fixed both CI failures: updated lint telemetry for the new fourth source directive, quoted the hyphenated fixture value, and removed unreachable test cleanup. Verified with tests/fm-lint.test.sh, targeted CI-mode ShellCheck, bin/fm-lint.sh, bash syntax checks, and git diff checks * no-mistakes(ci): Bound structured pause deadlines by FM_PAUSE_RESURFACE_SECS in watcher and daemon housekeeping, added distinct bounded-horizon reasons, regression coverage for near, passed, and wrong-year deadlines, and updated documentation. Verified targeted behavior tests, full daemon tests, ShellCheck source-following lint, syntax, and diff checks
) Opt-in IMAP/SMTP mail plane (fm-mail.sh / fm-mail-check.sh). Absent FM_MAIL_* stays off. Speaking as Kun's firstmate: this is merged. Thank you @feilipu — really appreciate you taking the time on this.
* fix(remote): start the fm-remote Herdr agent through a login shell Launchd was exec-ing herdr directly, so the Aqua agent inherited a background session without login-keychain access. Start it via /bin/zsh -lc exec so panes keep login env and can refresh OAuth tokens after reboot. * fix(remote): start fm-remote Herdr via the account login shell Resolve UserShell from Directory Services and invoke it with separate -l and -c so bash, fish, and zsh all get login-keychain access. Fall back to SHELL, then /bin/zsh, then /bin/sh without failing the render. * no-mistakes(review): Fix launch-agent shell fallback resolution * no-mistakes(review): Preserve and escape Directory Services shell paths * no-mistakes(document): Document login-shell LaunchAgent behavior * no-mistakes(ci): Updated the trusted fm-remote-doctor SHA-256 identity in bin/fm-remote-entrypoint.sh. Verified with tests/fm-on.test.sh, tests/fm-remote-doctor.test.sh, bash syntax checks, and git diff --check * no-mistakes(ci): Resolved the login shell exactly once per doctor invocation and threaded it through plist rendering, installed/loaded contract validation, repair reporting, and post-repair checks. Added a regression test proving repeated repair remains healthy and performs no reload when a hypothetical second Directory Services lookup would differ. Updated the trusted doctor hash. Verified doctor, fm-on, remote-entrypoint, lint tests, ShellCheck, syntax, and diff checks * no-mistakes(ci): Made Darwin shell resolution hermetic with executable injection and a 2-second Directory Services timeout. Updated tests to inject shells by default, isolate dscl-specific cases, parse plists semantically, and verify stalled dscl fallback. Updated the trusted doctor hash. Doctor, fm-on, entrypoint, syntax, hash, and diff checks pass * no-mistakes(ci): Raised portable serial CI timeout from 20 to 30 minutes, refreshed the specified timing hints, added missing hints, and recomputed shard documentation. Verified coverage, runner behavior tests, workflow lint tests, shell syntax, requested timing maxima, and diff checks
…enguid#3710) * fix(bearings): keep captain-approved deliveries in Recently Landed A closed task is never held: tasks-axi clears the held flag when a task closes and keeps hold-kind and the hold reason as the record of the call that was made. Recently Landed excluded every Done row whose hold-kind was captain, so the marker it treated as "closed while still waiting on the captain" was in fact the proof that the captain had approved the work. Every merge routed through a captain decision disappeared from the list of what shipped, including under --all-landed. The selector now asks whether the closed row delivered something. Recently Landed is merged PRs, completed scouts, and finished local-only merges, so a row carrying one of those artifacts belongs there whoever approved it. A captain question closes with an answer and no artifact of its own, and that is what still stays out, so an answered question is never rendered as shipped work. The same rule was written twice - the bearings projection selects this home's Done rows and the fleet snapshot selects each secondmate home's Done rows into the roll-up the same section merges in - which is why one defect hid deliveries in every home. Both now share bin/fm-landed-lib.sh. * fix(review): Normalize landed evidence and exclude answered captain questions * fix(review): Normalize captain delivery evidence across relocated data * fix(review): Record authoritative delivery provenance with legacy fallback * fix(review): Harden delivery provenance across forced and pruned completions * fix(review): Replace premature merge closure with existing release contract * fix(review): Document provenance-based Recently Landed selection * fix(document): Align documentation with completion provenance * fix(lint): Fix targeted ShellCheck warnings * fix(ci): order the pinned tasks-axi install before its stock-Bash consumers In `.github/workflows/ci.yml` the pinned tasks-axi install now precedes both stock-Bash consumers, and the Bearings expectation is updated from 49 to 50 tests. Verified with macOS Bash 3.2: snapshot 16/16, Bearings 50/50, public-followup 1/1. Full repository lint and all three workflow validations pass, and `git diff --check` is clean. * fix(review): Make completion provenance unambiguous * fix(review): Make completion verdict authoritative over quoted provenance * fix(review): Preserve retained artifacts through resumed captain closes * fix(review): Unify completion provenance ordering across writer and reader * fix(review): Preserve artifacts across failed captain closes * fix(review): Refresh v1 assertions; provenance authority remains unresolved * fix(review): Remove unreliable provenance while preserving landed deliveries * fix(review): Reject stale home summaries visibly * fix(review): Restore retained deliverable recording * fix(review): Match landed artifacts and restore retention documentation * fix(review): Disambiguate captain calls and restore landed artifact matching * fix(review): Persist retained report and PR artifacts * fix(review): Preserve staged artifacts before captain answers * fix(review): Avoid wedging answers on unsupported report paths * fix(review): Exclude unreleased captain-held pull requests * fix(review): Exclude held local-only answers from landed * fix(review): Preserve retained scout reports across snapshot rendering * fix(document): Align landed lifecycle documentation with release semantics * fix(review): Enforce landed artifact-kind ownership * fix(review): Infer canonical task kinds in snapshots * fix(review): Require captain-hold release before merges * fix(review): Qualify merge lifecycle regression evidence * fix(review): Serialize captain holds with merge operations * fix(review): Document merge cleanup residuals honestly * fix(test): Replace vacuous Bearings regression with behavioral cases * fix(document): Align Bearings verification and merge lifecycle documentation * fix(review): Serialize merges and exclude captain calls from landed * fix(review): Harden merge identity and landed selection * fix(document): Clarify landed selector compatibility filtering * fix(bin): keep merge entrypoints usable on records without an incarnation The merge identity guard refused any task record with no spawn_gen field. That field identifies one exact incarnation, so comparing it across the wait for the merge lock is what catches a task relaunched while the merge was queued. Requiring it to be present is a different rule, and it refused every record written before the field existed: a legacy task could no longer be merged at all, and five behaviour suites refused before reaching the check they were written to exercise. The comparison only needs to notice a change. An absent field is now read as an empty incarnation and compared like any other value, so a record that gains, loses, or alters one is still refused, while a record that simply predates the field merges. An ambiguous or unreadable field stays an error, because a record that cannot name one incarnation cannot be compared. The missing-record message each entrypoint had before the guard is restored, so a genuinely absent record still says so in its own words. The role partition now precedes reading the record. Refusing the supervision branch is a statement about the actor, not about the task, so it cannot depend on a record the wrong actor may not have. A backlog file that does not exist meant "no longer an open captain call". For a caller that asked to tell absence apart it now means absent, so a board card whose home carries no backlog stays visible instead of being dropped as resolved. Fixture repositories pin their initial branch instead of inheriting init.defaultBranch, which resolved to main on a developer machine and master on a runner, so a fixture naming main failed only in CI. * fix(review): read local-only note from body; surface pending-close failures * fix(review): keep kindless local-only landings in Recently Landed * fix(review): bind local-only note scan to the tasks-axi note line * fix(review): Guard unavailable captain-hold authority records * fix(document): Document unreadable authority predicate outcome * fix(bin): read an absent backlog as absence, not an unreadable record The merge gate refused every task whose home carries no backlog file. A backlog that does not exist holds no captain call, so nothing can be held and the merge is safe; only a backlog that exists and cannot be read may hide a live hold. Those two states were collapsed into one refusal, which stopped merges in any home that keeps no backlog. The predicate now reports a missing backlog file as absence, alongside a row the backlog does not carry. A record that exists but cannot be read still leaves by the existing cannot-tell path, which both merge entrypoints already refuse, so the restrictive direction is unchanged. That leaves no way to reach the separate unavailable-record result, so the result and the two branches that handled it are removed rather than left describing an outcome that can no longer occur. The lifecycle documentation loses the same claim. Regressions cover both directions in each entrypoint: a home with a task record and no backlog merges, and a backlog present but unreadable refuses without reaching the forge. * fix(review): Fail closed unreadable backend configuration * fix(tests): pin the bare origin's initial branch in the remote seed fixture The fixture created its bare origin with no initial branch, so that repository's HEAD followed init.defaultBranch while the source repository pushed the branch fm_git_init_commit pins. On a host that still defaults to master the two disagreed: the bare origin's HEAD named a branch the push never created, cloning it warned that the remote HEAD referred to a nonexistent ref and checked out nothing, and the seed assertion for the cloned README failed. A machine whose default is already main paired the two by accident and hid it, which is why the fixture passed locally and failed on the runner. Pinning the bare origin to the same branch removes the dependency on the ambient default from both sides. Verified under both conditions: with init.defaultBranch set to master, and set to main, the suite passes 26 of 26. * fix(review): Fail closed unreadable user backend configuration * fix(bin): republish the home summary as v1 and record two load-bearing rules The published home-summary schema had moved to v3, which routed every secondmate home still emitting the earlier version to the stale branch: their landed rows, open decisions and holds all came back empty and their state read as unknown until each home was updated. The payload never justified that. Its field set, field order, truncations and the landed array construction are byte-identical to v1, so only which rows the selector places in landed differs, and a v1 consumer reads that the same way. Republishing as v1 removes the rollout regression and, with it, the tolerance machinery that existed only to soften the bump: the stale-schema predicate, its two collection branches, the flag and its provenance branch, the omitted surface that can no longer be reached, and the fixtures and assertions that covered them. Two rules that a scope review proposed removing are kept, each now carrying the reason it exists, because both were measured to be load-bearing: The artifact-kind ownership clause is what keeps an explicit scout that recorded no report out of Recently Landed. Without it such a row has none of the three artifacts, satisfies the compatibility fallback and renders as shipped work with an empty artifact. The kind fallback is needed because tasks-axi omits the kind metadata entirely when a title begins with a canonical keyword. Without it a scout titled "SCOUT ..." reports no kind, its recorded report stops counting as a delivery, and it drops out of the section this selector exists to repair. * fix(bin): move the scout guard note onto the rule and drop two dead pieces The LOAD-BEARING note sat on an unreachable branch. Measured in both directions: removing that branch together with the kind-is-not-scout guards lets an explicit reportless scout into Recently Landed and fails tests/fm-captain-hold-lifecycle.test.sh, while removing the branch alone leaves that suite passing at 49 assertions. The guards carry the rule, so the note now sits on them and the unreachable branch is gone. A note pointing a later reader at the wrong line is the hazard this change corrects elsewhere. summary_file_has_schema lost its only caller when the stale-schema machinery was removed, so it goes with it. * fix(review): Fix legacy report artifacts and canonical keyword boundaries * fix(review): Update pinned Bearings test count to 56 * fix(document): Clarify landed summary compatibility documentation * fix(review): Preserve unreadable backend configuration errors * fix(review): Honor backend resolution errors at existing call sites * fix(test): Stabilize remote collector tests under host load * fix(document): Document backend resolution failure contracts * fix(lint): Suppress intentional deferred probe expansion warnings * fix(ci): Captain, quoted the two literal test IDs in tests/fm-backlog-atomicity.test.sh to fix SC2100 without changing behavior. Both warnings reproduced before the fix; the targeted fm-lint.sh run now passes with ShellCheck 0.11.0. Bash syntax and git diff --check also pass * fix(ci): Fixed the resolver’s two configuration-parent checks to return 2 for inaccessible directories while preserving genuine absence. Added two behavioral tests; RED/GREEN and both requested mutation proofs confirmed. All 10 focused checks, targeted lint, syntax, and whitespace checks passed. Broader merge suite stopped after 10 passing cases under host load. Declined portable checks and merge-authority code remain unchanged
…nguid#4090) * fix(remote): let the Aqua launch agent own the fm-remote Herdr session A herdr server keeps the macOS audit session of whatever started it, and only the Aqua login session (gui/<uid>) can read the login keychain without a prompt. Herdr's SSH remote attach starts the fm-remote server as its own child when it finds none, wins the socket at boot because sshd accepts connections before the login session exists, and every claude pane under that server then gets `security` exit 36, falls back to a stale plaintext credentials file, and reports "Login expired". launchd's own job lost the socket on every KeepAlive retry and the doctor still reported the session ready because it only asked whether any server answered. - Add bin/fm-remote-herdr-guard.sh, the launch agent's exec target: start the server in the foreground when nothing owns the socket, exit 0 when an Aqua-born server does, and otherwise stop the foreign server, wait for the socket, and exec the server at once. - Add bin/fm-remote-herdr-owner-lib.sh, the single owner of socket-owner discovery (lsof; pgrep cannot see herdr's argv on macOS) and the birth markers (SSH_*, XPC_SERVICE_NAME, FM_REMOTE_JOB_ACTIVE, sshd or remote-client-bridge ancestry matched on argv[0] and whole arguments). - Render the agent as the login shell exec'ing the guard with KeepAlive={SuccessfulExit=false} and ThrottleInterval=10, check the loaded job's successful-exit semaphore, and report a session served outside the Aqua login session as fixable so --fix retakes it through launchd; the reload waits for an Aqua-born owner rather than any running server. - Correct the doctor and docs: the launch shell provides environment parity, the launchd domain provides keychain access. - Pin the guard's decision table and the doctor's verdicts against real marker-carrying processes, and record the dated audit-session evidence. * no-mistakes(review): Verify Aqua ownership through launchd domains * no-mistakes(document): Document macOS lsof ownership requirement
…uid#2881) * fix(bin): prefer a live no-mistakes run over a terminal one A worktree can bind to more than one recorded no-mistakes run at once. The branch-and-code-identity rule in bin/fm-nm-run-lib.sh accepts both an exact-equal commit and a worktree-is-an-ancestor match, but never stated which wins when both bind, so the tie fell to whichever candidate the caller reached first. Observed on a live fleet: a crashed validation daemon left a FAILED run at the worktree's own commit while the live run that replaced it validated a descendant commit on the same branch. Bare `axi status` answers with the most-recently-touched run - the corpse - and it bound by the equal-commit rule, so every recomputation reported `failed` for a task whose real run was healthy. The same label had also read `failed` earlier while the work was genuinely stalled, so the signal was wrong in both directions. State the live-over-terminal policy in the matching rule's own contract, where the equal-commit and ancestor rules already live, and add fm_nm_run_status_class as the one classifier that decides liveness from a recorded status word. fm-crew-state.sh applies it on both selection paths: the runs listing now scans past a terminal row for a live one, and a terminal `axi status` answer is provisional until the listing has been asked whether this worktree also has a live run. Same-liveness-class candidates keep the listing's newest-first precedence, and a status word the classifier cannot place keeps the caller's own ordering rather than displacing a known result, so a single-run task and a task whose runs are all terminal are unchanged. Regression coverage reproduces the proven case (terminal run at the worktree's exact commit plus a live run descending from it) and its runs-list twin; both fail under the old tie-break. Two companion cases pin the no-widening half - two terminal rows still resolve newest-first, and a terminal run with no live sibling keeps its full run-step detail - and both pass before and after the change. * no-mistakes(review): accept unfetched live sibling anchored at exact worktree head * docs(bin): name both ledger reads behind the runs-limit setting The FM_CREW_STATE_RUNS_LIMIT comment in bin/fm-crew-state.sh still described the runs ledger as scanned only by the cross-branch fallback, but the live-over-terminal fix also consults it as the live-sibling probe behind a terminal axi status answer. Point the comment at docs/configuration.md as the setting's owner instead of restating a second copy. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…chenguid#4009) * fix(procevent): make the ordinary stop signal actually stop a runner The owner guard that shipped in kunchenguid#3904 half-reaps. Against a poll child that handles the ordinary stop signal and keeps waiting, the guard signals the group, loses the runner leader to its own signal, then reads that success as a leaderless group and exits without escalating. It destroys the only proof of ownership that would have authorised the forced signal, so the survivor becomes unreachable by retire, reconcile, sweep-home and the guard alike. A guard that turns a leaking-but-identifiable generation into a permanently unreachable one is worse than no guard at all. Two defects, and they hid each other: - The escalation re-derived ownership from the leader. `runner_group_signal` now takes a `proved` mode, passed only by the escalation inside the stop that already proved and signalled that exact generation moments earlier. A leader dying to our own signal is the ordinary outcome, not fresh ambiguity. - Every stop held the per-source lock across its wait while the runner's own exit cleanup waited unboundedly for that same lock. That circular wait was broken only by the forced signal, so the forced signal silently became the normal path - and, by keeping the leader alive through the whole window, it masked the escalation defect above. The runner's exit cleanup now refuses that lock instead of waiting for it, which is what its existing `return 0` already said it did. Fixing the lock alone would have turned every stop of a signal-proof child into a refusal that leaves it running, so both land together and the tests pin that. Measured on macOS with a stand-in poll child that traps TERM, INT and HUP: the guard left it running past 70s and now clears the group within the lease plus one check; retiring a healthy runner fell from ~2.8s with a forced group signal every time to ~0.6s on the ordinary signal alone. Unchanged and stated deliberately: a leader lost to anything other than the stop's own signal still leaves a group that retire, reconcile, sweep-home and the guard all refuse, permanently - and that source stops listening without saying so. Whether such a group may ever be signalled is an open decision and is not answered here. * fix(review): Fix proved escalation race and stop regression assertions * fix(review): Preserve proved escalation through transient identity failures * fix(review): Simplify proved escalation and correct guard timing documentation * fix(document): Clarify process-event stop ownership and cleanup limits * fix(document): Clarify process-event stop ownership and fixture comments * revert(skills): restore the leaderless-ambiguity limit to the loaded skill An automatic documentation step in this branch's validation edited .agents/skills/process-event-sources/SKILL.md, which no instruction in this change asked it to touch. That file is not documentation about the code: it is the agent-loaded instruction surface, what an agent reads to know what it is permitted to do. The step deleted this line: - leaderless PID/PGID-reuse ambiguity preserves the claim without signalling or replacement, as owned by the operating contract in [`docs/configuration.md`](../../../docs/configuration.md#process-to-event-sources-stateprocevent); and folded it, with its neighbour, into a generic "registration and ownership transitions, stop authority, and claim reclamation follow the operating contract". That deleted line states a PROHIBITION - that such a group is preserved WITHOUT SIGNALLING - and it is the exact limit an open captain decision currently rests on. Folded into a pointer, an agent reading the skill to learn what it may do would have to chase a second document to discover it may not signal. A prohibition that requires a second lookup is not a prohibition. The effect was to weaken, in the instructions themselves, the boundary that keeps one home from signalling another's process group - while the question of whether that boundary should move at all is still open. This is a deliberate revert, not an oversight, and it restores the file exactly to its pre-branch state. The full statement also survives in docs/configuration.md; that does not rescue it, because the agent handling a process-event wake loads the skill and not the documentation. * revert(procevent): restore the open-question marking beside the escalation The same automatic documentation step that edited the loaded skill also removed this from the comment above runner_group_signal: A leaderless group nobody in this call ever proved remains refused too, for every caller. That untouched refusal is what makes a crashed leader's group permanent, and relaxing it is a separate open question, not something this path assumes. and replaced it with a pointer to docs/configuration.md. This one fails differently from the skill deletion, which is why it is restored separately. There, a prohibition was moved out of the reader's path, and a missing prohibition gets violated. Here the prohibition survives in code - the unproved path still refuses - and what was removed is the fact that the limit is UNDECIDED. A prohibition that has quietly lost its "this is still open" reads as settled design, and settled design gets relied on, extended, and eventually relaxed by someone confident they understand why it is there. That question is open right now. The rule this branch's four instances produce, stated once here because this is the point of decision: an unresolved question must be marked unresolved AT THE POINT OF DECISION, not only where the contract is documented. A reader who does not know something is open will treat it as closed, and that default is stronger than any pointer overcomes. The pointer added by that step is kept alongside; this restores what it replaced rather than reverting it. * docs(verification): restore the measured guard bound and its reason The document step's rewrite of this record dropped the concrete figure while keeping the surrounding measurements. What went missing was the bound itself - lease plus two consecutive failed checks plus the stop's grace, roughly 630 seconds at the shipped 600-second lease and 15-second check - together with the reason there are two checks rather than one: a single unreadable read must not be enough to kill a live runner. The mechanism survived elsewhere and the reason survived in docs/configuration.md, so nothing was lost from the repository. The concreteness was, and that is what this restores. A number recorded without why it is that number is the one a later reader shortens; the reason is the whole safety argument for the debounce, and the debounce is what stops the reaper killing a live runner on one bad read. * fix(document): Replace stale stop-authority summaries with owner pointers * test(procevent): make the guard-bound case able to fail for its own reason An automated reviewer observed that this case allowed sixty seconds for a bound of roughly eight, so it could not go red for the reason it names: it would have passed a guard that took fifty-five seconds. That is correct, and it is the same family as the defect the case exists to defend against - a check that is green because it cannot fail, rather than because the thing it guards is working. The deadline is now derived from the bound itself - the lease, plus the two consecutive failed checks the guard debounces on, plus the stop's own ordinary and forced signal windows - rather than from a flat wall-clock number, and the shortened lease and check the fixtures run under have a single definition so a derived deadline cannot silently diverge from the settings the guard is given. The doubling that remains is a load allowance and is documented as one; widening it to make a slow guard pass would convert the assertion back into decoration. Proven by mutation rather than by argument. Against the repaired case: correct code ok guard debounces on 20 misses instead of 2 not ok - "still holding the group after 16s, against a documented bound of 8s" proved escalation removed (the original defect) not ok - same code restored ok The previous sixty-second version passes every one of those mutations. The reviewer's other claim, that the guard can survive past the announced bound when an owner disappears immediately after a check, was measured and does not hold against what this branch announces. Sweeping the phase deliberately at 0.0, 0.2, 0.4, 0.6 and 0.8 of a check interval gave 7.21s, 7.31s, 6.75s, 6.49s and 6.31s, worst 7.31s, against the announced lease plus two consecutive failed checks plus stop grace, which is up to 8s at those settings. The mechanism the reviewer describes is real and is the announced mechanism; the bound it was measured against is a phrasing this branch no longer carries. * revert(scope): return the instruction surfaces to their base state This delivery is being split. It carries the two proven process fixes alone; the instruction text travels separately, through a run that removes the documentation step rather than refusing it at its gate. Two surfaces are therefore returned to exactly what the base branch has, so this delivery neither adds to them nor removes from them: .agents/skills/process-event-sources/SKILL.md - identical to base again. Three bullets an automatic documentation step had folded into a pointer, including that leaderless PID/PGID-reuse ambiguity preserves the claim WITHOUT SIGNALLING and that there is one identity-matched owner per canonical source across homes sharing one store. The header comment block of bin/fm-procevent.sh, which is what the script prints as its own help. Seven lines were removed from it: that a live owner is never displaced, that only a claim whose stale owner and independently absent process group prove its whole generation gone is reclaimed, that a crashed leader or reused pid whose process group still has members cannot relax ownership cleanup, and that reconcile signals only a live identity-matched runner group and otherwise keeps the claim without starting a replacement. The help output is now byte-identical to base. Neither removal was requested by any instruction in this change, and both were made to text that predates it. Returning them is scoping, not a third restoration: nothing is being added to those files here. * fix(ci): Captain, live CI revealed a fixture deadlock: it suspended the runner before startup released its lock. Added a public-list synchronization barrier in tests/fm-procevent.test.sh. Forced-delay reproduction detected the deadlock before the fix; all four cases passed afterward. Targeted lint, Bash syntax, and whitespace checks passed. Greptile’s watchdog requirement conflicts with the recorded R2 decision; runtime behavior and documentation remain unchanged. Full CI rerun belongs to the outer executor * test(procevent): make the post-TERM cases report what they saw when they fail On the failure path only, these cases now print what they actually saw: the identity recorded at claim time, the identity readable at that moment, the size of the signals file, the leader's state and wchan, every live member of the runner's process group with its own state and wchan, the elapsed time since the stop began, and what retire said. None of it runs when a case passes. WHY THIS IS KEPT, stated accurately rather than by its original reason. It was written to make an unexplained CI failure verifiable. That failure is now explained - it was a fixture deadlock, diagnosed and repaired in the preceding commit - so that justification has expired and is not the reason given here. The reason it stays is smaller and independent of that failure: it is already written, it is small, it sits in the file whose assertion this change reworked, and an assertion that could not say why it failed cost most of a morning to diagnose from the outside. The next failure will not be this one. WHAT A PASSING RUN WOULD NOT MEAN: a pass is a sample of behaviour already observed many times, not proof that anything is fixed. Only a failure carrying the evidence above establishes a cause. * fix(document): Clarify process-event fixture diagnostic rationale * fix(ci): Captain, fixed two cleanup races in tests/fm-procevent.test.sh: removed premature child completion and waited for runner exit before retiring the restart fixture. Controlled Linux reproductions demonstrated failure before and success after. The full Linux process-event suite, six focused macOS checks, targeted ShellCheck, Bash syntax, and whitespace checks passed. Runtime behavior, guard debounce, and documentation remain unchanged. CI rerun belongs to the outer executor * fix(procevent): bound owner-guard cleanup at one check interval, not two A THIRD WAY, not a capitulation to the reviewer and not a refusal of it. The automated reviewer's grievance was the LOOSENESS OF THE BOUND, not the number of observations the guard makes before it acts. It asked for a single read because that was the only route it could see to an acceptable bound. There was another route, and this change takes it: the bound is reached and both reads are kept. TIGHTENED - the SPACING of the guard's two reads, not their number. The owner watchdog now sleeps half the configured check interval and still requires two consecutive failing reads, so the pair completes inside one check interval instead of costing two. Worst-case detection falls from the lease term plus TWO check intervals to the lease term plus ONE. At the shipped 600s lease and 15s interval the stated bound falls from ~635s to ~620s. PRESERVED - the second read. bin/fm-procevent.sh's two-consecutive-miss rule is untouched. WHY IT PROTECTS: the guard's inputs are a lease read and a state-root identity read, and either can fail transiently on a live, healthy home. Acting on the first failure would let one isolated unreadable read kill a live service. Requiring a second, independent read is what makes that impossible, and it is a protection rather than padding. Nothing was traded away to reach the bound. Both properties are now guarded by their own case, and each was proven by MUTATION rather than asserted: - putting a full interval back between the two reads fails the bound case: "still running 17.0s after the last owner activity, against a documented bound of 15s"; - acting on one failed read fails the new debounce case: "one unreadable lease read ended a runner whose home was still alive" - while the bound case then passes FASTER, 9.9s against 13.1s. The unsafe variant being the quicker one is exactly why these are two cases: one elapsed-time case would have registered the removal of the protection as an improvement. MEASURED, sampling the phase between the guard's check clock and the lease clock across eight runs per variant, on macOS (Darwin 25.5.0). Reaping an orphaned listener whose home stopped refreshing its lease: lease 2s / interval 1s: 4.41-5.29s before, 3.48-4.65s after lease 2s / interval 4s: 7.69-8.12s before, 5.94-6.13s after The 4s configuration is the informative one: the gap is about one check interval, which is precisely the term that was removed. A previously unstated term of the bound surfaced while measuring: the lease age is compared in whole seconds, so a configured lease of N is honoured until that age reads N+1. It is now part of the documented bound and of the regression's derivation instead of being absorbed into a fudge factor. The bound regression derives its deadline from the documented bound instead of a flat number, and PINS the phase between the guard's check clock and the lease clock rather than sampling it, because with a sampled phase a guard spending two intervals passes about half the time on a lucky alignment. Its load slack is additive and stays under half a check interval, so an extra whole interval cannot hide inside it. The two flat deadlines that were there before (40s and 20s) and the doubling allowance on the derived one are gone; that looseness was the reviewer's third complaint. The stop's own grace is untouched: 2s for the ordinary signal, then 2s for the forced one. It is a ceiling paid only by a group that outlives the signal it was sent, not a delay every stop pays - a healthy runner's whole retire measures 0.40-0.66s on this host. The reviewer's literal "lease plus one tick" is unreachable by any implementation, since signalling a process and giving it any chance to exit takes non-zero time; detection now meets it and the stop runs inside its own ceiling, and the contract says so rather than glossing it. NECESSARY BUT NOT SUFFICIENT, and written BEFORE this head's integration runs start rather than after they report. On the previous head, "Behavior portable serial 1" and "Behavior portable serial 4" were both CANCELLED at the job ceiling, independently of this finding. A new head triggers fresh runs, so those two lanes MAY complete this time. IF THEY DO, THAT IS NOT EVIDENCE THE CEILING DEFECT IS FIXED. It is one more sample of a lane that has been cut repeatedly and sometimes is not; the shard-packing repair for it is open separately. Do not reread a lucky pass here as a resolution. Relatedly, and deliberately: the per-script duration hint in bin/fm-test-run.sh was NOT updated even though the two new cases add ~19s of wall clock. docs/fm-test-portable-shards.md says those hints are replaced wholesale from CI timing artifacts of green runs, and that repair is the open request doing it; a hand-edited estimate here would collide with it and silently repack the shards. This suite runs in portable serial shard 3, which was green in the last run. Verification: tests/fm-procevent.test.sh green, plus tests/fm-captain-hold-lifecycle.test.sh, the test-coverage guard, and bin/fm-lint.sh. The unrelated "reconcile stops a runner whose registration was removed" case flaked in 4 of 7 local full runs; an isolated 20-trial reproduction measured it at 13/20 unclean before this change and 11/20 after, so it is issue 4080 and is not aggravated here. * fix(procevent): repair our decimal-interval regression and enforce the timing phase REPAIRED BEFORE PUBLICATION, AND IT WAS OURS. The half-interval arithmetic added by the previous commit read a zero-prefixed interval as octal: 010 halved to 4 instead of 5, and 08 was not a number at all, so the owner guard died before reporting ready and the runner failed closed and never listened. The validator accepts those values and `[` compares them as decimal, so this broke a configuration that worked before. Introduced by this delivery, found in review, repaired here. Forcing base ten before the arithmetic is the whole runtime fix. Proven by driving it rather than by reading the source: a new case starts a real listener at 08 and at 010 and observes the guard's actual sleep argument - 4s and 5s. Removing the normalisation turns that case red with "a zero-prefixed decimal interval (08) prevented the listener from starting". THE TIMING PHASE IS NOW OBSERVED AND ENFORCED, NOT ASSUMED. The bound case pinned its phase by CONSTRUCTION, from an assumed startup time, and enforced nothing. Review was right that this is not enough: once startup reaches about two seconds the expiry lands in a different part of the interval and the case silently stops rejecting a two-interval guard while still reporting success. A bound that cannot fail for the reason it names is the defect this whole delivery exists to correct, so it must not ship inside the fix for it. Now the lease is synchronised to the guard's own FIRST observed lease read, every later real read is recorded, and the case REFUSES unless one recorded read proves the required phase: it read the synchronised reference, it was still fresh, and it began late enough that two further full intervals could not finish before the deadline. An unestablished precondition refuses; it does not proceed on trust. The derived deadline, the two-read debounce and the additive slack are unchanged, and the slack invariant is now asserted rather than left to a comment. Review also found the deadline was only ever checked while the group was still alive, so a sampler descheduled past it would see the group gone and certify success. The observed completion time is now checked too. PROVEN BY MUTATION, each one run against this code: - remove the decimal normalisation -> the interval case fails on 08; - a full interval between the two reads -> "the guard exceeded its bound: group still running 17.1s ... against a documented bound of 15s"; - a full interval WITH startup forced to ~2.5s, which is exactly the condition the old construction pin could not survive -> still red, same message; - the same ~2.5s startup with the correct guard -> still passes, 13.0s against the 15s bound, so the delay alone does not break the case; - phase evidence made unavailable -> "could not establish the required pre-expiry guard-read phase", a refusal rather than a pass, even though the group stopped quickly; - act on one failed read -> the debounce case fails and the bound case passes FASTER, 9.5s against 12.7s, which is why these remain separate cases. Verification: full tests/fm-procevent.test.sh green, and bin/fm-lint.sh clean. * fix(document): Correct process-event timing and debounce comments
…kunchenguid#3417) * fix(backlog): honor configured task adapters * no-mistakes(review): Harden backend purity lint against prefixed Beads calls * no-mistakes(document): Document configured backend lifecycle transitions * fix(backlog): preserve markdown exemptions * no-mistakes(review): Enforce backend purity for explicit lint paths * no-mistakes(document): Update lifecycle backend documentation * no-mistakes(lint): Remove redundant backend lint pattern * fix(backlog): close adapter routing gaps * no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint * no-mistakes(document): Document environment-selected backlog adapters * no-mistakes(lint): Fix empty local variable assignment * fix(backlog): close quoted path gaps * no-mistakes(review): Reject partially quoted direct Beads commands * no-mistakes(document): Align lifecycle documentation with configured adapters * test(backlog): keep structural cases markdown-only * fix(backlog): honor configured task adapters * no-mistakes(review): Harden backend purity lint against prefixed Beads calls * no-mistakes(document): Document configured backend lifecycle transitions * fix(backlog): preserve markdown exemptions * no-mistakes(review): Enforce backend purity for explicit lint paths * no-mistakes(document): Update lifecycle backend documentation * no-mistakes(lint): Remove redundant backend lint pattern * fix(backlog): close adapter routing gaps * no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint * no-mistakes(document): Document environment-selected backlog adapters * no-mistakes(lint): Fix empty local variable assignment * fix(backlog): close quoted path gaps * no-mistakes(review): Reject partially quoted direct Beads commands * no-mistakes(document): Align lifecycle documentation with configured adapters * test(backlog): keep structural cases markdown-only * no-mistakes(review): Harden markdown lifecycle routing and close recovery * fix(lint): catch dollar-quoted beads commands * no-mistakes(review): Drop unused authorized-data-dir param; canonicalize fm-lint ROOT with pwd -P * no-mistakes(document): Document fm-lint backend-purity check in header and CONTRIBUTING * no-mistakes(ci): Two failing checks, one code-caused and one attestation-only. 1. Greptile Review (P1: ANSI-C quoting bypasses lint) — REAL DEFECT, FIXED. The backend-purity normalizer in bin/fm-lint.sh (invokes_bd awk function) stripped $'...' quote delimiters without decoding ANSI-C escapes, so a core script containing $'\x62\x64' close fm-example executes as `bd close` while the lint accepted it. Fix: the normalizer now tracks whether a single-quote context came from $' (ansi flag) and decodes ANSI-C escapes while building the command word: \xHH (1-2 hex), \uHHHH / \UHHHHHHHH (4/8 hex), \NNN (1-3 octal), \cX control characters, simple escapes (a/b/e/E/f/n/r/t/v decode to a placeholder that can never spell bd), NUL (value 0) truncates the word per bash C-string semantics, and unknown escapes drop the backslash per bash. Values outside printable ASCII decode to a placeholder so they can neither falsely match nor collide into `bd`. Regression coverage added to the existing lint-interface test test_rejects_direct_beads_cli_invocations in tests/fm-lint.test.sh: $'\x62\x64', $'\142\144' (octal), and b$'\x64' (split word), each written into a fixture and asserted rejected through the real fm-lint.sh executable. Verified regression property: with the fix stashed, the new hex case passes lint (reproduces the reported bypass); with the fix, it is rejected. Verification: all 29 tests in tests/fm-lint.test.sh pass, and the full CI-parity lint (CI=true bin/fm-lint.sh: ShellCheck 0.11.0 full analysis over the canonical set, backend-purity check, and actionlint workflow validation) exits 0. One iteration was needed because an awk comment containing a literal $'\x62\x64' sample broke shell-level quoting (bash -n / SC1001/SC2026); the comment was reworded without quotes. 2. PR must be raised via no-mistakes — NOT code-caused. The check failed with 'Pipeline attestation head_sha does not match the current PR head': the PR body attestation binds to 82b41c7 while the PR head is 1995cdb because a later pipeline push moved the head. This is exactly the stale-attestation condition the user intent describes; it clears when the outer pipeline re-runs 'git push no-mistakes' and re-binds the attestation to the new head (which now includes this Greptile fix). No code change can or should address it. Files changed: bin/fm-lint.sh (ANSI-C escape decoding in the backend-purity normalizer), tests/fm-lint.test.sh (three encoded-bd rejection cases) * no-mistakes(review): fix tasks.toml hang, root authorization, lint quoting * no-mistakes(review): validate tasks config before exemption; fix lint quote gap * fix(backlog): address the markdown backlog as <data>/backlog.md Resolving the markdown backlog through a configured `[markdown] path` was scope this task never asked for. It is absent from main, which addresses `<data>/backlog.md` everywhere, and it came from an earlier review round rather than the task brief. Making it effective on the transition path alone put that path at odds with every other consumer of the same backlog - fm-captain-hold.sh, fm-session-start.sh, fm-fleet-snapshot.sh, fm-inbox.sh, fm-backlog-handoff.sh - which all still address `<data>/backlog.md`. In fm-captain-hold.sh the split was live: its reads had already moved to the shared gate while its writes had not, so the two could address different files. Address `<data>/backlog.md` from the shared gate, delete the unused resolver, and drop the two tests that pinned the withdrawn behaviour. What this task actually changes is unaffected: a configured non-markdown adapter is still addressed by its own root, without `--file`. * fix(backlog): honor configured task adapters * no-mistakes(review): Harden backend purity lint against prefixed Beads calls * no-mistakes(document): Document configured backend lifecycle transitions * fix(backlog): preserve markdown exemptions * no-mistakes(review): Enforce backend purity for explicit lint paths * no-mistakes(document): Update lifecycle backend documentation * no-mistakes(lint): Remove redundant backend lint pattern * fix(backlog): close adapter routing gaps * no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint * no-mistakes(document): Document environment-selected backlog adapters * no-mistakes(lint): Fix empty local variable assignment * fix(backlog): close quoted path gaps * no-mistakes(review): Reject partially quoted direct Beads commands * no-mistakes(document): Align lifecycle documentation with configured adapters * test(backlog): keep structural cases markdown-only * fix(lint): catch dollar-quoted beads commands * no-mistakes(review): Drop unused authorized-data-dir param; canonicalize fm-lint ROOT with pwd -P * no-mistakes(document): Document fm-lint backend-purity check in header and CONTRIBUTING * no-mistakes(ci): Two failing checks, one code-caused and one attestation-only. 1. Greptile Review (P1: ANSI-C quoting bypasses lint) — REAL DEFECT, FIXED. The backend-purity normalizer in bin/fm-lint.sh (invokes_bd awk function) stripped $'...' quote delimiters without decoding ANSI-C escapes, so a core script containing $'\x62\x64' close fm-example executes as `bd close` while the lint accepted it. Fix: the normalizer now tracks whether a single-quote context came from $' (ansi flag) and decodes ANSI-C escapes while building the command word: \xHH (1-2 hex), \uHHHH / \UHHHHHHHH (4/8 hex), \NNN (1-3 octal), \cX control characters, simple escapes (a/b/e/E/f/n/r/t/v decode to a placeholder that can never spell bd), NUL (value 0) truncates the word per bash C-string semantics, and unknown escapes drop the backslash per bash. Values outside printable ASCII decode to a placeholder so they can neither falsely match nor collide into `bd`. Regression coverage added to the existing lint-interface test test_rejects_direct_beads_cli_invocations in tests/fm-lint.test.sh: $'\x62\x64', $'\142\144' (octal), and b$'\x64' (split word), each written into a fixture and asserted rejected through the real fm-lint.sh executable. Verified regression property: with the fix stashed, the new hex case passes lint (reproduces the reported bypass); with the fix, it is rejected. Verification: all 29 tests in tests/fm-lint.test.sh pass, and the full CI-parity lint (CI=true bin/fm-lint.sh: ShellCheck 0.11.0 full analysis over the canonical set, backend-purity check, and actionlint workflow validation) exits 0. One iteration was needed because an awk comment containing a literal $'\x62\x64' sample broke shell-level quoting (bash -n / SC1001/SC2026); the comment was reworded without quotes. 2. PR must be raised via no-mistakes — NOT code-caused. The check failed with 'Pipeline attestation head_sha does not match the current PR head': the PR body attestation binds to 82b41c7 while the PR head is 1995cdb because a later pipeline push moved the head. This is exactly the stale-attestation condition the user intent describes; it clears when the outer pipeline re-runs 'git push no-mistakes' and re-binds the attestation to the new head (which now includes this Greptile fix). No code change can or should address it. Files changed: bin/fm-lint.sh (ANSI-C escape decoding in the backend-purity normalizer), tests/fm-lint.test.sh (three encoded-bd rejection cases) * fix(bin): preserve captain calls during teardown (kunchenguid#3595) * fix(bin): never close a captain call during cleanup A scout that held its own work item for the captain, which is what captain-hold-lifecycle prefers ("hold the work item the question gates"), was closed by bin/fm-teardown.sh's automatic backlog transition. The completion gate passed, cleanup ran, and the captain's question moved to Done with no recorded answer: the one thing the policy says must never happen. `tasks-axi done` closes a held row silently, and nothing in teardown asked whether the row was the captain's own call. bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when the task is still an open captain call, 1 when it is not, 2 when that cannot be established. It reads the row through the transition library's backend-aware probe, so it addresses the same backlog teardown does; the script's other commands now address the configured data directory the same way instead of FM_HOME, which also fixes captain holds in a home with a relocated data directory. Teardown asks `open` before any destructive step and refuses on 2. On 0 only the close changes: after cleanup and still under the task's own lock, the row gets one "Deliverable of the finished work" line at the end of its body and returns to Queued through `tasks-axi reopen`, keeping its hold, so it lands in Captain's Call instead of reading as work under way. --force does not lift this: it authorizes discarding unlanded work, never the captain's question. The deliverable goes into the body because `tasks-axi update --report` rewrites the title of a row that is not Done. The crash window reuses the pending-close record teardown already stages: a `mode=retain` line makes the existing replay record the deliverable and reopen instead of closing, with the same validator, stale-generation check, cleanup-incomplete marking, and non-blocking bootstrap lock as an ordinary close. A retained row the captain answered first simply retires the record. No parallel record type, recovery command, or second bootstrap loop is introduced. Regressions run the real executables: the captain-held scout survives cleanup queued, held, with its deliverable and on the board, only `answer` closes it, --force keeps it open, and an ordinary scout still closes with its report; an interrupted cleanup leaves the row untouched and the next session start retains it; a relocated backlog keeps the retention in its one configured file; and a ship row whose hold cannot be read refuses cleanup before anything destructive. Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np * no-mistakes(review): Serialize captain holds and fix backend-aware listing * no-mistakes(document): Update captain-call retention documentation * no-mistakes(document): Fix relocated captain-hold backlog diagnostics * fix(backlog): honor configured task adapters * no-mistakes(document): Update lifecycle backend documentation * fix(backlog): close adapter routing gaps * no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint * no-mistakes(lint): Fix empty local variable assignment * fix(backlog): close quoted path gaps * test(backlog): keep structural cases markdown-only * no-mistakes(review): Harden markdown lifecycle routing and close recovery * no-mistakes(review): fix tasks.toml hang, root authorization, lint quoting * no-mistakes(review): validate tasks config before exemption; fix lint quote gap * fix(backlog): address the markdown backlog as <data>/backlog.md Resolving the markdown backlog through a configured `[markdown] path` was scope this task never asked for. It is absent from main, which addresses `<data>/backlog.md` everywhere, and it came from an earlier review round rather than the task brief. Making it effective on the transition path alone put that path at odds with every other consumer of the same backlog - fm-captain-hold.sh, fm-session-start.sh, fm-fleet-snapshot.sh, fm-inbox.sh, fm-backlog-handoff.sh - which all still address `<data>/backlog.md`. In fm-captain-hold.sh the split was live: its reads had already moved to the shared gate while its writes had not, so the two could address different files. Address `<data>/backlog.md` from the shared gate, delete the unused resolver, and drop the two tests that pinned the withdrawn behaviour. What this task actually changes is unaffected: a configured non-markdown adapter is still addressed by its own root, without `--file`. * no-mistakes(review): restore home boundary guard and tighten purity lint * no-mistakes(review): authorize home boundary for every backlog adapter * no-mistakes(test): complete tasks-axi stubs in fm-gotmp teardown fixtures * no-mistakes(document): align backlog transition docs with adapter-neutral addressing * no-mistakes(review): label adapter data-dir authorization, drop dead row_probe local * no-mistakes(review): pin markdown backend at relocated-data addressing roots * no-mistakes(document): point lint-definition mention at fm-lint.sh header * no-mistakes(document): point mutate comment at adapter addressing owner * no-mistakes(review): Fix leftover-symlink refusal on non-markdown homes; hoist config check and lint/dedup cleanups * no-mistakes(document): Align fm-lint purity scope header with bin/backends --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
…guid#4119) * fix(bin): stop false watcher-down alarms on long Claude turns A healthy Stop auto-arm rewake or open claim already explains a mid-turn beacon that has aged past grace, because turn-end will re-arm. Keep the supervision-off banner for a missing, failed, or exhausted generation. * no-mistakes(review): Bind Claude rewakes to active recovery generation * no-mistakes(document): Document Claude long-turn supervision exception * no-mistakes(lint): Fix empty ShellCheck assignment * no-mistakes(ci): Fixed all reported CI failures: quoted the hyphenated recovery-delivery value to satisfy ShellCheck SC2100, and updated the session-lock auto-arm fixture to emit the recovery marker and watcher beacon now required for a valid rewake. Verified fm-session-lock-ancestry, fm-test-run, stale-banner, Claude auto-arm, targeted lint, ShellCheck, and workflow lint checks pass
* fix(herdr): classify a gone session's endpoint as recoverable A task whose Herdr endpoint could not be read was classified `unreadable`, which blocks recovery by design. The commonest reason that read fails is that the recorded session's server is not running at all - a host reboot, a server exit, a session never restored - and that is authoritative absence for every pane in that session, not an ambiguous answer about one of them. Tasks in that state had no sanctioned way back. The recovery-grade read now settles an uninterpretable pane read with the session server's own `.server.running` state: positively stopped reads `missing`, while a running server, or a server state that cannot itself be read, still reads `unreadable`. Resting the verdict on that field rather than on the `server_not_running` error code is what keeps it working across Herdr 0.8.x and 0.9.0, since the field is present on both and the code is not. Only that one boundary is widened. The husk classifier under it stays strict, so duplicate prevention, rollback, and teardown - the paths that can destroy something - keep refusing on exactly the reads they refused on before. Separately, a relaunch refused outright when the endpoint's shell had drifted out of the recorded worktree. An agent's own exit routinely leaves its shell somewhere else, so that refusal stranded tasks whose work was sitting untouched on disk. The shell is now told once to return, and only a shell that will not go refuses; the replacement still never starts outside the copy holding the work. Herdr 0.8.x is not installed on this host, so protocol-20 coverage is structural plus the adapter fixture exercising both response shapes, and is recorded as such rather than as a live result. Fixes kunchenguid#4091. * no-mistakes(review): Restrict drift recovery to Herdr endpoints * no-mistakes(review): Correct Herdr recovery verification coverage * no-mistakes(document): Document Herdr endpoint recovery boundaries
kunchenguid#4033) * fix(bin): keep an escalated undelivered handoff wake retryable A remote backlog handoff holds its outbox until the backlog receipt and the receiver wake are both confirmed, and retries the wake under the same pending-reply correlation on every resume. When that wake's remote transport was lost, the correlation stayed undelivered in delivery_unknown and the watcher's next pending-reply tick escalated it. Both the reuse predicate and the known-undelivered reset refused an escalated record, so the resume refused to resend the wake forever and every later handoff to that mate jammed behind the outbox. Treat an escalated record with no confirmed delivery as the undelivered correlation it is: fm_pending_reply_corr_reusable accepts it for its own task and fm_pending_reply_reset_known_undelivered returns it to awaiting_report for the idempotent remote resend, while a delivered record is still never reset and a missed-report escalation keeps its meaning. The published delivery-unknown decision stays open until the record resolves, so a repeat loss neither re-notifies nor strands it. Reproduce the deadlock end to end in the remote handoff test (lost wake transport, watcher escalation, resume) and pin the predicate contract in the pending-reply suite; the fm-send fixture that pinned the refusal now uses a genuinely stale delivered escalation. * no-mistakes(review): Decouple durable outboxes from best-effort wake retries * no-mistakes(review): Align handoff documentation with durable receipt release policy * no-mistakes(review): Handle unrecordable wake state as dropped * no-mistakes(review): Prevent stale wake markers blocking handoffs * no-mistakes(review): Prevent stale delivered markers suppressing new wakes * no-mistakes(document): Clarify retry escalation decision lifecycle * no-mistakes(document): Document pending receiver wake retries
* docs: bound the mandatory captain address to the chat channel AGENTS.md's opening address rule said "address the user as captain at least once in every response" and never said what a response is. The artefact exclusion two lines below governed only the optional nautical seasoning, not the mandatory address. An agent that reads this file without being the first mate - a pipeline corrector agent running inside a copy of this repo - therefore read the obligation as applying everywhere and the exclusion as applying only to flavour, and opened its delivery message with "Captain,". That reading was correct. Patch the existing owner rather than adding a rule elsewhere: - bound the obligation to chat messages sent to the captain; - state the artefact exclusion once, explicitly binding every agent that reads this file whether or not it is the first mate, and naming commit messages, PR and issue descriptions, briefs, code and comments; - fold the seasoning under the same bound instead of carrying a second, narrower copy of the exclusion. The obligation itself is unchanged: the captain is still addressed in every chat message. AGENTS.md goes from 603 to 602 lines: the redundant "never send a response with zero direct address" clause and the duplicated seasoning exclusion pay for the new bound. The two cross-references that paraphrased the unbounded wording (bin/fm-parent-channel-lib.sh's header and docs/secondmate-parent-channel.md's problem statement) now match the owner; neither restates the rule. * fix(review): Limit address exclusions to artifacts while preserving public replies * fix(document): Consolidate captain address guidance
…nguid#4131) * fix(herdr): close persisted-focused tabs when no live client is attached The teardown active-tab guard treated Herdr's last-focused pointer as a live viewer, so detached sessions could not close panes on that tab. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(herdr): allow detached seeded-tab prune after live-client gate Projection create still restored the persisted focused tab after a successful prune, so a detached last-focused seeded tab still quarantined the spawn. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(herdr): probe live client after seeded prune only when that tab was focused The extra title-clear read after every prune shifted canned CLI fixtures and failed projection create. Co-authored-by: Cursor <cursoragent@cursor.com> * no-mistakes(review): Tighten Herdr active-tab close guard * no-mistakes(review): Guard Herdr mutations with fresh target focus * no-mistakes(document): Document Herdr live-viewer teardown guard --------- Co-authored-by: Cursor <cursoragent@cursor.com>
…guid#4169) * fix(bin): escalate decision-owned wakes once as the decision The away-mode daemon treated a needs-decision: queued payload as an unknown wake, so suppression markers never committed and the same open decision re-escalated on every poll. Classify that payload through the existing signal path so it escalates once, labelled as the decision, and an unchanged repeat is suppressed on the same terms as any other signal. Fixes kunchenguid#4096 * no-mistakes(review): Escalate captain-held decision-owned rows once as the decision * no-mistakes(review): Self-handle captain-held decision-owned rows instead of escalating them * no-mistakes(document): Name away daemon as needs-decision payload reader
* test(herdr): pin leftover-shell vs live-idle via agent get Herdr 0.9.0 already distinguishes a Pi that exits to a surviving pane shell from a sibling live idle occupant. Pin that pair through agent get and the recovery classifier so a lagged pane-get status cannot silently reclaim the leftover shell as alive. Co-authored-by: Cursor <cursoragent@cursor.com> * no-mistakes(document): Document Herdr leftover-shell liveness regression * no-mistakes(ci): Fixed Lint failure SC2034 by replacing the unused wait-loop variable with `_`. Verified with the pinned project lint command, Bash syntax check, and git diff check --------- Co-authored-by: Cursor <cursoragent@cursor.com>
…nchenguid#4151) * ci: rebalance the portable parallel lanes on measured runner durations Both portable parallel lanes are capped at 10 minutes. Lane 1 was cancelled at that cap on every request raised on 2026-09-10 while lane 2 finished in about 3.5 minutes, so no request could go green. CONTRACT CLASS: RESTORE. The workflow already promises two duration-balanced lanes and the shard documentation already claims a measured wall; this re-establishes both against what the lanes now cost, and changes no lane count, no cap, and no scope of what runs. The counter-argument, so nobody has to take that on trust: two pieces here are genuinely new rather than restored, and either could be argued to make this a NEW-behavior change. `--list-scheduled` now ranks a parallel lane on measured durations where it previously handed every parallel script the serial default weight and returned an alphabetical order; and `--check-coverage` gains three reported fields. I classify the change RESTORE because both exist only to make the already-promised property checkable, but they are named here rather than folded into the restoration. === PART 1: THE TOTAL, AND HOW IT WAS OBTAINED === This section stands on its own. It establishes what the parallel set costs. It derives no packing; Part 2 does that, from this number. THE TOTAL: 828568 ms, about 13 min 49 s of serial work across the 24 scripts. Lane 1 held 624299 ms of it and lane 2 held 204269 ms, a 3.06:1 split. HOW IT WAS OBTAINED. The difficulty was that lane 1 had never finished, so its duration did not exist as a recorded figure anywhere and no timing artifact was expected for it. It turned out to be recoverable from the real lane without estimating, by two routes, across six CI runs on 2026-09-10 (34459949083, 34460760299, 34462530836, 34462758357, 34466966385, 34470382458): - Run 34462758357's lane-1 job finished its suite 18 s BEFORE the wall and uploaded a complete fm-test-timing-portable-parallel-1 artifact carrying all 11 scripts, FM_TEST_SUMMARY total=11 failed=0 duration_ms=598225. The upload step is if: always(), so the cancellation did not suppress it. This is one full, untruncated lane-1 measurement. - The five other lane-1 jobs were cancelled mid-suite, but each logs every script that had already finished as an FM_TEST_END duration_ms= marker. Those per-script records are complete measurements of completed scripts; only the script in flight at cancellation is lost, and it differs by run. Lane 2 completed in all six runs, so its scripts come from the six uploaded fm-test-timing-portable-parallel-2 artifacts. Every one of the 24 scripts therefore carries at least one untruncated measurement: 20 of them measured in all six runs, two in three or four runs, and two (fm-brief, fm-transition-lib, the tail of lane 1) in the single complete run. Each hint is the SLOWEST value that script reached, so the total is an upper envelope rather than an average. NO FIGURE IN IT IS DERIVED FROM A TRUNCATED LANE, and no lower bound was ever extrapolated into a total. THE ENVIRONMENT, AND WHETHER IT TRANSFERS. Every hint is a serial run of the real portable parallel lane on a GitHub ubuntu-latest runner, produced by the lane's own CI job. It transfers because it is not a proxy for the lane; it is the lane. Nothing in the total came from this machine or from any harness of mine. That mattered, and here is what it would have cost. A same-day macOS cross-check of the same scripts ran 1.7x to 5.0x slower with the ratio varying per script (fm-test-run 157420 ms against 92944 ms, fm-x-mode 67217 ms against 31870 ms, fm-composer-ghost 10521 ms against 2120 ms). Local timings therefore do not scale the lane, they REORDER it, so a packing derived from them would have balanced the wrong thing while looking clean. WHAT IT REPLACES, which is the root cause. The lanes were packed from the 2026-08-20 concurrent isolation proof: 24 candidates across four LOCAL workers. That record answers whether the candidates are isolation-safe, not how long a SERIAL CI lane runs, so it was structurally incapable of representing lane wall clock even when it was fresh. It was also never refreshed while the set grew about 3.2x. Both the wrong instrument and the staleness are fixed here: the hints now come from the lane itself and carry their run ids and date. === PART 2: THE SPLIT DERIVED FROM THAT TOTAL === Longest-processing-time assignment over those hints gives 414269 ms and 414299 ms, 30 ms apart, against 624299/204269 before. tests/fm-pi-primary-types.test.sh stays in lane 1 because that is the job which installs the Pi package, so ci.yml needs no step changes. === PART 3: DOES THE MARGIN SURVIVE MACHINE VARIANCE === Stated explicitly, because 6.90 min against a 10 min cap is 69% of cap before any variance is applied, and the cap covers the whole job rather than the suite. worst lane, script time 414299 ms 6.90 min job overhead, measured on the real lane ~18 s (see below) expected healthy job ~432300 ms 7.21 min x1.29 on the script time, plus overhead ~552400 ms 9.21 min cap 600000 ms 10.00 min room left after the multiplication ~47.6 s 7.9% of cap The 1.29x is the runner variance measured today on the SIBLING SERIAL lane, as supplied; it is not this lane's own figure. This lane family does have its own, and it is tighter: the six full lane-2 sums today span 192939 ms to 203451 ms, a spread of 1.054x. At that figure the worst lane lands near 7.58 min with about 2.4 min of room. I have used the LARGER, borrowed 1.29x for the verdict rather than the tighter one this lane actually shows, and note that the hints are already per-script maxima, so 1.29x on top is conservative twice over. THE MARGIN SURVIVES THE MULTIPLICATION, so this proceeds rather than stopping. The 18 s overhead is measured, not assumed: in run 34462758357 the lane-1 job ran 10 min 16 s against a 598.2 s suite, and lane 2 ran 3 min 21 s against a 192.9 s suite, a ~10 s difference that matches lane 1's extra Pi package install. The cap is unchanged, the lane count is unchanged, and nothing in the serial lane, its shard count, its guard or its hint table is touched. === PART 4: THE RECORDED FACT === The workflow comment no longer restates the shard wall as a literal, which is how "~1 min of serial sum" survived a 10x change without announcing it. It now points at bin/fm-test-run.sh --check-coverage, which prints parallel_max_ms, parallel_imbalance_ms and parallel_unhinted derived from the hint table, so the current number is computed on demand. The shard documentation carries the dated run ids, which route it was taken by, and the local cross-check that shows why local numbers are not admissible as hints. Two regressions pin what rotted: lane membership must be stored longest-measured-first, and the lanes must be fully hinted and packed within 5% of each other. Both were run against the old composition and both fail on it (420030 ms imbalance against a 624299 ms worst lane). The ordering assertion they replace named a specific script by hand and had itself gone stale. === PART 5: NAMED AND LEFT, OUTSIDE THIS REBALANCE === tests/fm-captain-hold-lifecycle.test.sh alone is 296481 ms, 36% of the whole set, so it is the floor of any two-lane split: no repacking can put a lane below it. After this rebalance the cap is about 1.45x the healthy lane where the sibling serial lane keeps roughly 2x. Nothing refuses a stale parallel hint the way PORTABLE_SERIAL_MAX_UNHINTED_PERCENT bounds the serial lane. parallel_unhinted is reported, not enforced, which is what let this drift for three weeks unnoticed. * fix(review): Restrict parallel scheduling hints to portable parallel lanes * fix(document): Clarify parallel lane scheduling and timing evidence
Upstream's backlog adapter routing added a spawn case that refuses a data directory resolving outside the home. The fork's home session and account pin runs earlier and cannot write its pin through that directory, so spawn blamed the session instead of the data directory. When the pin fails and the read-only backlog preflight also refuses the data directory, report that precise reason; every other pin failure is reported as before.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Captain said:
yes do full firstmate update and update other tooling too.Refresh the captain's Firstmate fork with current upstream changes needed for the authorized Herdr 0.9 upgrade while preserving every deliberate captain-specific policy and fix.
Keep work in remote custody and do not change the running Firstmate or Herdr server before the PR is merged and validated.
What Changed
869ae905and keeps the fork's own policies and fixes. The Herdr backend (bin/backends/herdr.sh) now:status --jsonand recovers worker endpoints that have gone away or drifted.cdback, and only refuses if the shell stays outside.fm-mail.sh,fm-mail.py) andfm-mail-check.sh, which sets up a standing mail poll.fm-remote-herdr-guard.sh,fm-remote-herdr-owner-lib.sh) that keeps thefm-remoteHerdr server in the macOS Aqua login session.fm-remote-doctor.shrenders and checks it.fm-landed-lib.shthat tells landed deliveries apart from resolved captain calls.bin/, with matching docs and tests:fm-spawn.shreports an inaccessible backlog data directory instead of the session-mismatch error when the home pin is absent, but still prefers the session-mismatch error when that is the real cause.🤖 Generated with Claude Code
Risk Assessment
show/absent contract, and correctly keeps the session-mismatch error with a regression test that reproduces the prior bug. The overall branch is still a large upstream merge (91 files, about 12k lines) that touches spawn, backlog and home-identity paths, so it merits normal follow-up attention.Testing
I ran the suite with the new session-mismatch regression (fm-home-identity) and the backlog suite with the two data-directory-outside-the-home cases (fm-backlog-atomicity). Both passed on the target commit. Against the pre-review fm-spawn.sh, the new regression fails with the misattributed backlog error. The Herdr backend and remote-guard suites for the upstream refresh also passed. This is a CLI and shell change with no UI, so the evidence is the CLI transcripts from the test runs. I did not change the running Firstmate or Herdr, and the worktree is left clean.
Evidence: Test run: new foreign-session-vs-backlog regression and the existing session tests (target commit)
Source: Test run: new foreign-session-vs-backlog regression and the existing session tests (target commit)
Evidence: Same test against the pre-review fm-spawn.sh: the regression fails, which shows the bug it catches
Source: Same test against the pre-review fm-spawn.sh: the regression fails, which shows the bug it catches
not ok - a refusing backlog preflight must not hide the session mismatch (missing: 'does not belong to the current session') error: task task-h cannot be dispatched because its backlog data directory is inaccessible: .../spawn-foreign-backlog/home/data (tasks-axi config resolves outside its authorized directory at .../home/.tasks.toml)Evidence: Backlog atomicity suite, including the data-directory-outside-the-home refusals
Source: Backlog atomicity suite, including the data-directory-outside-the-home refusals
Evidence: Remote Herdr guard suite
Source: Remote Herdr guard suite
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
bin/fm-spawn.sh:1190- The new fix blames the backlog data directory whenever the home pin fails AND the full backlog check (fm_backlog_transition_applies) returns 2. That check also returns 2 for reasons that have nothing to do with the data directory, such as an old tasks-axi ('automatic backlog transitions require tasks-axi ... or newer') or a tasks-axi backend/config error.Failing case: a home pinned to another Herdr session, spawned from the wrong session, on a host with an incompatible tasks-axi.
fm-home-identity.sh ensurerefuses (exit 3, session mismatch), then the check returns 2 for the tasks-axi reason. The last error line becomes 'task X cannot be dispatched because its backlog data directory is inaccessible: <data> (automatic backlog transitions require tasks-axi...)' instead of the session-ownership refusal. That contradicts the commit's claim that 'every other pin failure is reported as before'.On
--relaunchthere is a second problem: KIND is still the defaultshipat this point, because it is only read from the task meta at line 1352. A secondmate relaunch is therefore checked as a backlog item.Fix: only use the backlog reason when the data directory is actually why no pin could be written. Either check just the data-directory resolution and authorization (e.g.
[ -L "$DATA" ] || [ ! -d "$DATA" ]plus the authorized-root check), or skip the override when the pin output is a session/store mismatch.🔧 Fix: Keep session-mismatch spawn error ahead of backlog refusal
1 info still open:
bin/fm-spawn.sh:1193- The session-mismatch case is now fixed.fm-home-identity.sh showprintsabsentonly when no pin exists, and a mismatch needs a pin to exist, so a genuine session or account mismatch always keeps the 'does not belong to the current session' error. The new regression test fails on the previous commit (8c4f80e) and passes now.One narrow mislabel remains when there is no pin yet.
ensurecan fail with an absent pin for reasons unrelated to the data directory, such as an unusable HERDR_SESSION name or a relative CLAUDE_CONFIG_DIR. If the backlog preflight also returns 2 for its own unrelated reason (for example an incompatible tasks-axi), the last line reads 'backlog data directory is inaccessible: <data> (automatic backlog transitions require tasks-axi ...)'.This is low impact: fm-home-identity's own refusal is printed first, the spawn would be refused by the later backlog preflight anyway, and the previous fallback text was equally inaccurate for these causes. No action needed unless the label should list only data-directory causes.
✅ **Test** - passed
✅ No issues found.
bash tests/fm-home-identity.test.shon the target commit: passed, including the newtest_a_foreign_session_refusal_is_not_blamed_on_a_refusing_backlogBefore-the-fix check: took a /tmp copy of the target commit, swapped inbin/fm-spawn.shfrom 8c4f80e (before the review fix) and ranbash tests/fm-home-identity.test.sh. The new regression fails with the wrong 'backlog data directory is inaccessible' error, which shows the test catches the bugbash tests/fm-backlog-atomicity.test.sh: passed, including 'spawn refuses a data directory symlinked outside the home' and 'a configured adapter refuses a data directory outside the home'bash tests/fm-backend-herdr.test.sh: 194 checks passedbash tests/fm-remote-herdr-guard.test.sh: 9 checks passed✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.