feat(bin): add supervision host engine, per-home pool roots, and forge integration - #6
Closed
mehulbhagwani wants to merge 54 commits into
Closed
mehulbhagwani wants to merge 54 commits into
mehulbhagwani wants to merge 54 commits into
Conversation
…ns (kunchenguid#5389) The sibling secondmate stall cases in tests/fm-wake-queue.test.sh now wait for the watcher's recorded observation instead of a one-second wall-clock checkpoint, so they can neither fail nor pass vacuously under load. Deterministic proof with a 5s watcher launch delay: before the fix 4 cases passed vacuously and 6 failed; after it all 10 pass on the recorded observation. Also includes a CI flake fix from validation: fm_control_harness_supported in bin/fm-control-lib.sh finishes reading the harness allowlist before returning, removing intermittent broken-pipe diagnostics. Behavior is unchanged.
… tail (kunchenguid#5336) * fix(bin): refuse a Herdr submit that would send only a message tail A long typed payload can sit in the composer as a suffix, or as a paste placeholder plus a remainder, and the following Enter was still reported as delivered. Prove the selected composer holds the payload before Enter, and report failure when it does not. * no-mistakes(review): Scope Herdr payload proof to Claude, clear composer on refusal * no-mistakes(test): Clear refused Herdr composer drafts one wrapped row per press * no-mistakes(test): Accept Claude's multi-line paste placeholder in Herdr submit proof * no-mistakes(review): Accept Claude read-back that drops U+2063 in Herdr proof * no-mistakes(document): Document Herdr proof ignoring U+2063 operational mark * no-mistakes(ci): I made a one-line test change. The failing check comes from a timing race in an existing test that this PR doesn't touch. **What failed:** `tests/fm-procevent.test.sh` failed at "the superseded paced runner invoked its stale command" (line ~3313). The PR only changes the Herdr files and their tests, and the same shard passed on main at the base commit. **Why it can fail:** the fixture starts a second runner with a 3-second launch floor (the minimum wait since the source's last launch). That runner sleeps for the rest of the floor and only then checks whether its registration was replaced (`fm_procevent_launch_floor_wait` in `bin/fm-procevent-lib.sh`). The test then waits for the claim and re-registers the source. If that takes longer than about 3 seconds after the first launch, the old runner wakes up, finds its registration still current, and runs the stale command. That produces the second log line the test reports. The CI shard was slow (this one test took 160 s). **Fix:** in `tests/fm-procevent.test.sh` I raised the superseded runner's floor from 3 to 15 seconds and added a comment explaining why. The floor now outlasts the fixture setup even on a loaded runner. Nothing else changed: the first launch and the later fresh-registration start still use a 3-second floor, and no product code changed. **Verification:** - The full test file can't give a reliable result on this machine (load average about 64 on 8 cores). It failed earlier, at the reconcile assertion around line 1680, before it reached this section. - I ran the changed section by itself (file setup plus the pacing-race block) five times with the fix: all passed, in about 9-13 s each. - The original code also passed five out of five, so the race didn't reproduce locally. The diagnosis rests on the code path and the CI log. - I haven't seen the full file or the CI shard pass with the fix yet * no-mistakes(test): Accept Claude folder-trust prompt via down+enter in live e2e * no-mistakes(document): Note unreadable Claude composer refusal in Herdr docs
…unchenguid#5427) Speaking as Kun's firstmate: squash-merging — opt-in (forge=gerrit registry-gated; default project-mode stdout restored to two words), attestation MATCH, CI+NM green, safe review, MERGEABLE.
…uid#5358) * feat(bin): add an opt-in per-home worker account pin A home that mixes work and personal accounts for one runner had no way to say which account its workers launch on: Claude workers inherited whatever CLAUDE_CONFIG_DIR the supervising process had, Pi workers the pane's ambient root, and an ambient API key outranked both, with no signal at launch. config/claude-account and config/pi-account now pin that choice per home. With neither file every launch is unchanged. With one, every launch of that runner from the home (ship, scout, local secondmate, raw Claude command, and relaunch) runs under the declared root, and the spawn refuses before any endpoint exists when the file is malformed or the runner's own check (claude auth status, pi auth check with a model-listing fallback) says the pinned account is not signed in. The check runs in a cleared environment so an ambient credential cannot answer for an empty root. A pinned Claude launch sheds the environment credentials Claude ranks above a stored login; a pinned Pi launch needs an explicit <provider>/<id> model for a declared provider and also carries --provider. The chosen account is printed on the spawned line and recorded in the task record, and relaunch checks the pin before stopping the running agent. * test(secondmate): give the concurrent config-push wait room for a slow host test_config_reread_serializes_concurrent_pushes waited about two seconds for the first fm-config-push.sh to reach its first send-keys. On a slower host that push takes four to five seconds, so the test failed on main before the push ever got there. The loop still leaves as soon as the marker appears, so the larger bound costs nothing where the push is fast. * no-mistakes(review): Refuse raw Claude account overrides under a pin
…ts (kunchenguid#5470) * fix(bin): keep the Herdr lab session option before a -- delimiter fm-herdr-lab.sh run appended --session <lab> after every argument, so a command with a passthrough delimiter such as agent start ... -- <agent args> handed the session flag to the agent and Herdr routed the call by the caller's ambient socket instead of the lab. The helper now inserts --session <lab> immediately before the first -- delimiter and keeps the trailing form otherwise. * no-mistakes(document): Clarify Herdr lab session option placement
* feat(bin): guard the partition, harness pin, and bounded exec for a non-Pi supervision host Lease liveness is now the pure record test in every calling context, so an unmarked main honors a live branch lease held by a separate process, and a lease file engages the guard's claim serialization for any caller; a home with no lease files still takes no lock. bin/fm-harness.sh honors FM_SUPERVISION_PRIMARY_HARNESS while FM_SUPERVISION_ACTOR=branch, so a supervision branch running under another harness resolves own, crew, and secondmate to the primary's harness. fm_tasks_axi's watchdog moves into bin/fm-timeout-lib.sh as fm_exec_timed with a separate grace: the perl watchdog is preferred, runs the command in its own process group against wall-clock deadlines, forwards TERM/INT/HUP, and reaps the group, so a descendant holding captured output can no longer keep the caller waiting past the bound on a host without timeout. The Claude Stop auto-arm header records that Claude drops the exit 2 of a hook it terminated at the configured timeout, re-measured on Claude Code 2.1.281. * fix(bin): state that fm_exec_timed cannot reach a descendant in its own process group Live runs of real Claude and Pi engine turns under the bound showed both CLIs start every tool command in a process group of its own, so those processes end through the engine's own TERM handling rather than the group signal or reap. Also clears the new timeout test's ShellCheck findings. * no-mistakes(document): Clarify cross-harness lease documentation
…henguid#2648) * feat(bin): make the ship-branch prefix configurable per project fm-brief.sh hardcoded every generated ship branch to fm/<task-id>, which leaks that firstmate produced the branch/PR - unwanted for a third-party public repo that does not use this tooling. Add an optional --branch-prefix flag to fm-brief.sh (default "fm/", so existing installs are unaffected) and teach fm-project-mode.sh - the registry's single-owner parser - to resolve a project's optional "branch=<prefix>" data/projects.md annotation via a new --branch-prefix query, order-independent with the existing mode/+yolo tokens. Firstmate resolves the override at task intake and passes it explicitly, mirroring how --mode already works; fm-brief.sh itself never reads the registry. An empty override resolves to a bare "<task-id>" branch rather than a leading slash. All five previously hardcoded fm/$ID sites (branch creation, never-push rule text, definition-of-done text, and the status message) now render the resolved prefix consistently. * no-mistakes(review): Wire branch-prefix intake in AGENTS.md; fix fm-merge-local.sh hardcoded fm/ prefix * no-mistakes(document): docs: document configurable ship-branch prefix in architecture.md * no-mistakes(review): Persist immutable branch contracts * no-mistakes(document): Document configurable ship branch prefixes * no-mistakes(lint): Captain: fix ShellCheck test warnings * fix(bin): map bearings PR rows to their recorded ship branch (kunchenguid#1887) fm-bearings-snapshot.sh keyed a PR back to its task by string-matching the headRefName against the fm/ prefix, so any project whose branch prefix was overridden (e.g. via kunchenguid#2648's branch=<prefix> registry annotation) had its PRs silently drop to task "-" in the bearings view, exactly the third fm/-assumption issue kunchenguid#1887 named alongside fm-merge-local.sh and fm-bearings-snapshot.sh itself. fm-fleet-snapshot.sh now surfaces each task's recorded branch= metadata field in its JSON task rows, and fm-bearings-snapshot.sh cross-references a PR's headRefName against those recorded branches before falling back to the legacy fm/ prefix heuristic, so a custom branch prefix maps a PR back to its real task. Adds a regression test proving a PR opened against a fix/<task-id> branch resolves to that task instead of "-"; confirmed it fails on the prior startswith("fm/") logic and passes with this change. ShellCheck clean; full fm-bearings-snapshot.test.sh and fm-fleet-snapshot-view.test.sh suites pass. * fix(ci): align lint arithmetic-looking assignment and stale Bearings snapshot count - Quote the --branch-prefix want_value assignment in fm-brief.sh, fm-promote.sh, and fm-spawn.sh so ShellCheck SC2100 no longer misreads the plain string 'branch-prefix' as arithmetic shorthand. - Bump the Stock macOS Bash snapshot job's hardcoded Bearings test-count assertion from 59 to 60: this PR added a Bearings test, so the count was stale, not the feature. * no-mistakes(review): fix(bin): honor recorded ship branch in relaunch and review-diff * no-mistakes(document): docs: complete branch-prefix flag in brief and promote headers * fix(lint): quote branch-prefix parser token; drop unused BRANCH_Q after rebase * no-mistakes(review): Restore %q branch escaping in promotion instructions with regression test * no-mistakes(document): document recorded ship branch and prefix flag fm-review-diff.sh's header is the owner of its branch-resolution contract; it still described only the legacy local-branch behavior after the change made review-diff honor state/<id>.meta's recorded ship branch. README's feature bullet enumerates the registry's optional flags and was missing the new branch=<prefix> override. * no-mistakes(lint): Silence SC2016 on intentional single-quoted sed expression * no-mistakes(review): address branch-prefix review findings in DoD and project-mode * no-mistakes(test): branch-prefix suites pass under tasks-axi 0.2.6; environment-only failure * no-mistakes(document): purge stale fm/ branch naming from docs and headers * fix(test): assert the merged epoch status wording in the branch-prefix override test The rebase resolution of tests/fm-brief.test.sh kept the branch's pre-merge \`done: ready in branch ...\` assertion while the merged fm-dod-lib.sh (carrying main's epoch-stamped status line) renders \`done [at=<epoch>]: ready in branch ...\`. Align the assertion so the override-consistency test matches the behavior it verifies. * no-mistakes(review): Address remaining branch-prefix findings in four bin scripts * no-mistakes(test): skip real-tasks-axi tests below the repo's 0.2.6 floor * no-mistakes(document): document spawn's branch-prefix registry deviation notice
…d per-rule confidence floors (kunchenguid#5478) * feat(bin): send dispatch resolver only the brief's task sections * Sent Jev only the scaffolded Captain's intent and Firstmate spec sections, falling back to the whole brief when neither heading is present, so the identical setup, rules, and definition-of-done boilerplate no longer reads as a signal about the task * Added an optional per-rule min_confidence that replaces the global 0.6 floor for that rule; a picked rule below its own floor falls to the most probable other option that clears its floor, or returns ambiguous * Kept files with no declared floor on the exact previous behavior and kept the model blind to the new field * Recorded the live old-versus-new comparison over scaffolded fixtures * no-mistakes(review): share brief heading parser, add kind line, fix floors * no-mistakes(test): stop sending ship delivery mode to jev, keep scout tag * no-mistakes(document): docs: list shared brief heading lib in scripts inventory
* feat(bin): supervision host core behind config/supervision-host Add the supervision host (bin/fm-supervision-host.sh): beside a Claude primary it owns the watcher cycle for the Stop auto-arm and, while the away-posture record exists, hands each wake to a bounded headless Claude engine session that runs the supervision branch's contract - the same generated prompt, row eligibility, wake grant, per-actor drain, outcome store, leases, and away relocation the Pi branch uses. Attended wakes pass straight to main. Every path that cannot finish a wake hands it to main with a supervision-host line; the park ends itself before the Stop hook timeout with a cycle-boundary wake. - bin/fm-supervision-engine-lib.sh: opt-in parse, verified engines (claude, default sonnet), one bounded engine turn, and a reap of engine tool processes that sit in their own process groups. - bin/fm-branch-report.sh: the command twin of fm_branch_report, scoped to the tasks the current host turn claimed. - bin/fm-branch-dispatch.mjs: command entry to the Pi dispatch module, so eligibility and the wake prompt have one owner. - bin/fm-claude-stop-autoarm.sh runs the host in the arm's place when config/supervision-host exists; nothing changes without the file. - bin/fm-watch-arm.sh --stop: home-scoped stop without a re-arm. - bin/fm-lease-lib.sh: an opted-in home takes the lease-command lock for unmarked main too, closing the first-claim race; the refusal tells the caller to leave the lease alone and retry. - /afk launches no away daemon on an opted-in Claude home; /quiet still does. Session start renders the host's main-side protocol there. * fix(bin): relay a host turn's outcomes when the captain returns mid-turn, and log per-turn engine cost Live validation found two supervision host gaps. A captain who returns while an engine turn is running gets a return brief rendered before that turn's outcomes exist, so the host now hands the close to main with those outcomes. Claude reports a resumed conversation's running cost, so the engine lib now derives each turn's cost from the total the host records, and the host log records every close's destination. * docs(verification): record the supervision host's live evidence The dated live results behind docs/supervision-host.md: the Claude engine's live guard, the away-wake cases against real workers, the engine's cost reporting, and the flag-off before-and-after regression. * docs: describe the supervision host ledger as covering every close * no-mistakes(review): Harden supervision host ownership, boundary, ack, and late outcomes * no-mistakes(review): Recheck park boundary just before starting an engine turn * no-mistakes(review): Cap park boundary, deliver all host lines, reject incomplete results * no-mistakes(document): Correct supervision host documentation and stale pointers
…t-in, and delivery (kunchenguid#5506) Attestation MATCH; contract-class restore; CI/NM green. Squash-merged by Kun's firstmate.
…kunchenguid#5528) * fix(bin): bound the startup-network worker's lock waits by its budget Fixes kunchenguid#5377 The deferred startup network worker bounded its sweeps with a stage budget but took the publish lock and the fleet-lock lease with an unbounded wait, so a live holder of that lock kept the detached worker alive for hours past its timeout with its output discarded at the end. Every wait now goes through the bounded acquire and shares the remaining stage or delivery budget; a lock a live process still holds at the deadline ends the worker with a failed record naming the holder and the rerun command, and a wake so the result surfaces. * no-mistakes(review): propagate publish exit code from cmd_run terminal paths
…dth (kunchenguid#5517) * fix(bin): keep the ps fallback identity independent of terminal width fm_pid_identity's portable fallback read the command column at the ambient COLUMNS width, so an identity recorded from a wide shell never matched the one recomputed inside a narrow hook and the continuity guard denied every fleet command. Pass -ww so the column is never cut. Fixes kunchenguid#799 * no-mistakes(ci): Fixed CI failure in Behavior portable serial 4. Root cause: the -ww flag added in commit ac7ab5d to fm_pid_identity (bin/fm-wake-lib.sh) shifted the ps argv so $1 became -ww instead of -p, breaking the positional fake-ps fixtures in tests/fm-procevent.test.sh (lines 2736, 3571) which then fell through to real ps and failed the fm-procevent test. Fix (already applied in the worktree, matching the authoritative user instruction exactly): replaced -ww with a COLUMNS=10000 environment pin so the call is COLUMNS=10000 LC_ALL=C ps -p "$pid" -o lstart= -o command=, mirroring fm_pending_reply_pid_identity in bin/fm-pending-reply-lib.sh:982. argv is back to -p PID -o lstart= -o command=, so the fixtures match again with no fixture edits. Comments above the call in bin/fm-wake-lib.sh and in test_pid_identity_is_terminal_width_invariant (tests/fm-watcher-lock.test.sh) now describe the COLUMNS pin instead of -ww; the regression test still asserts narrow-vs-wide byte equality and the full command. Verified: the terminal-width-invariant regression test passes. The only local not-ok results were flaky, run-varying timing tests (procevent launch/claim confirmation, listener reparenting) that differ each run and are unrelated to the ps argv change
…thorized intent (kunchenguid#5526) Fixes kunchenguid#3608 When a scout is promoted to a ship, the captain's authorized intent is extracted from a legacy `# Task` body by matching `Captain:` and `[captain]` lines anywhere in the body, including inside fenced code blocks and indented examples, while the heading reader already tracks fences. A fenced `Captain:` example therefore passed the provenance gate and became the ship contract's intent while the real ask was dropped. Make the captain-words extractor fence-aware like the heading reader: a line inside a ``` or ~~~ fenced block, or indented four spaces or a tab, is never a marked line. The promotion and spawn callers need no change. The regression test covers both the extractor and the promotion provenance gate refusing a brief whose only Captain lines are fenced or indented examples.
…riefs (kunchenguid#2868) * fix(bin): forbid administering the shared worktree pool in crewmate briefs A crewmate ran a `git worktree remove` loop over the treehouse pool its own worktree came from, destroying five worktrees - four belonging to tasks that were running mid-pipeline. The generated brief's rule 2, "stay inside this worktree; modify nothing outside it", is a rule about files: removing a worktree is administration of shared state, not an edit outside a directory, so the sentence never reached the act. The worker satisfied its brief completely. Rule 7 already named one piece of shared infrastructure - the no-mistakes daemon, one instance serving every lane - with the reason stated plainly. The worktree pool is the same class of thing and was unnamed. Fold the pool into that existing rule rather than adding a second warning: state the constraint around the act (create, remove, return, prune, move, reassign a worktree or pool slot; write into a sibling slot), keep concrete commands as examples rather than as the definition so no single provider is pinned, and give the prohibition a real exit through `blocked:`. The rule is emitted from one shared string interpolated into both crewmate scaffolds, so the ship and scout copies cannot drift apart. The secondmate charter deliberately omits it: that home runs its own fleet and legitimately allocates and returns slots for its own crewmates. Contract text only; no runtime enforcement layer. * no-mistakes(document): Distill pool-safety comment rationale * no-mistakes(review): align pool-rule test grep patterns with emitted [at=<epoch>] text
… a dispatch record (kunchenguid#5524) * fix(bin): refuse tasks-axi add --start so In flight always has a dispatch record Fixes kunchenguid#4753 Dispatch (bin/fm-spawn.sh) is the only path that moves a backlog row to In flight, because it creates the task record, status file, and inbox that go with the row. A row hand-placed there through the wrapper's `add --start` had none of those, and nothing later noticed, so the live-task count included work nobody was doing. The wrapper now refuses `add --start` (exit 2) and names the dispatch path; plain `add` and `start <id>` pass through unchanged, and the lifecycle transitions address tasks-axi directly so dispatch is unaffected. The issue's other half, a reconcile sweep in bin/fm-inactive-reconcile.sh that notices an In flight row with no task record, is left as is; this change closes the only path that creates such a row. * no-mistakes(review): refuse create --start alias, not just add --start * no-mistakes(review): reword add --start guard docs to drop only-path overclaim * no-mistakes(review): scope add/create --start guard docs, drop universal claim
…#5503) * feat(bin): run the supervision host beside the other non-Pi primaries while away Cursor's stop-hook park, the OpenCode plugin, the omp watch extension, Grok's model-owned background arm, and Codex's foreground checkpoint now run bin/fm-supervision-host.sh in the watcher arm's place when the home opted in with config/supervision-host, so the host's Claude engine takes away-posture wakes beside those primaries exactly as it does beside Claude. Without the file nothing changes. - The host streams its first cycle's status line, accepts --restart and the owner's predecessor arm for its first cycle, and prints each exit in one write, so owners that wait for arm readiness and restart their own successor (OpenCode, omp) keep their handling handoff. - Codex's checkpoint passes its bound to the host as the park boundary, raises it to FM_CODEX_WATCH_CHECKPOINT_AWAY (3600 s) while the away record exists, and lets an engine turn that starts before the bound finish after it (FM_SUPERVISION_HOST_PARK_LIMIT). - /afk launches no away daemon on an opted-in home of those harnesses and says so at entry when the file selects no engine for that primary. - Session start renders the host protocol for each arm owner, and Grok's arm command becomes the host. * fix(bin): keep the watcher-down banner away from the supervision branch actor A supervision host's engine turn runs guarded commands after its successor watcher cycle may already have closed on a newer wake, so the guard showed it the watcher-down banner with the primary's repair line. Under a Codex primary pin that line is the checkpoint, and a live Codex lab run showed the away session running it mid-turn (the nested host stood down on its ownership check). The branch actor never owns watcher continuity, so the banner, its reminder, and the episode state now leave that actor out, as the queued-wake warning already does. The lint telemetry fixture counts bin/fm-afk-launch.sh's source directives, which the host engine note raised from four to five. * fix(bin): queue away-session outcomes recorded after the return for main A Cursor park superseded by the captain's return stops its host as the engine turn ends, so the host's own handoff of that turn's outcomes was never printed and the outcomes never reached main. The report surface now queues every outcome it records after the away record is gone as a durable check wake; the return owner archives the record before it reads the store, so each outcome is in the return brief, queued, or both. A host stopped mid-turn also removes its turn's result and error files. The stream test now acknowledges its first close and accepts a restarted cycle that closes on its resurface before the arm confirms it. * fix(bin): clear a hard-killed host's turn at the next activation A Cursor park superseded mid-turn can kill its host outright, which runs no cleanup, so the turn's result, error, and descendant files stayed behind and any tool process the engine started was left running. The next host's activation now reaps the descendants that turn recorded and removes its files. The host suite also registers its homes in a file, because make_home runs in a command substitution, so its cleanup now stops every host a case leaves running. * fix(bin): leave rows that arrive after main's drain unclaimed at its acknowledgement Main's acknowledgement re-claimed every unreserved queued row, including one that arrived after the drain above the acknowledged cutoff. That row stayed main's without ever being shown to it, so while away the supervision host refused every later wake that included it and handed each back to main until main drained again. The acknowledgement now claims only unreserved rows at or below its cutoff. * docs: name the killed turn's engine and files in the host's failure direction * docs: record live supervision host runs on the non-Pi primaries * no-mistakes(review): Replay host-only supervision boundaries across omp session replacement * no-mistakes(review): Deliver omp supervision-host wakes only at the host's close * no-mistakes(document): Correct supervision host documentation for non-Pi primaries * no-mistakes(ci): Fixed the CI failure by naming FM_CODEX_WATCH_CHECKPOINT_AWAY in the rendered Codex host instructions. The focused instruction and checkpoint suites pass
) * feat(bin): auto-relaunch dead persistent secondmates during ordinary supervision A persistent secondmate whose primary agent exits mid-session previously stayed down until the next session-start liveness sweep. Extract the sweep's probe/classify/relaunch mechanics into a shared library and drive the same contract from a cadence-gated watcher tick, so a positively dead or missing endpoint is relaunched through the guarded spawn path within a poll cycle instead of an hour later. Only the recovery-grade `dead` and `missing` verdicts authorize relaunch; ambiguous, unreadable, unverified, and unreachable-remote reads stay fail-closed and a remote route is never replaced by a local endpoint. Each relaunch emits exactly one `check` wake and appends to a durable per-mate ledger; a mate exceeding the bounded attempt budget is parked behind a marker until a live probe rearms it. A per-mate liveness lock serializes the tick against a concurrent session-start sweep. * no-mistakes(review): Fail closed on relaunch ledger errors; clear state on remote teardown * no-mistakes(review): Share ledger read guard; retire relaunch state under liveness lock * no-mistakes(review): Lazy-load wake lib; live rearm restores full relaunch budget * no-mistakes(review): Finish liveness tick for every mate before waking once * no-mistakes(review): Keep liveness tick scanning past per-mate errors, then wake * no-mistakes(review): Wake only on queued rows; teardown holds liveness lock * no-mistakes(review): Queue liveness outcome wake before releasing mate lock * no-mistakes(document): Update secondmate liveness documentation for mid-session recovery * no-mistakes(lint): Fix empty assignments flagged by ShellCheck * no-mistakes(ci): Added ShellCheck analysis boundaries for the shared liveness library in both callers and marked its result globals as intentional library outputs. Changed-file lint passed; full CI partitions were not run locally * no-mistakes(ci): Fixed Lint 2 by removing an unused test variable in tests/fm-wake-queue.test.sh. ShellCheck, bash syntax, and the full wake-queue test script pass * no-mistakes(ci): Fixed the CI wake-queue fixture: stall-only watcher legs now seed the liveness cadence marker, preventing the new endpoint probe from interfering with their assertions. The full wake-queue test, ShellCheck, and diff checks pass locally
…cycle ends (kunchenguid#5550) * fix(bin): start a successor when the Claude Stop-hook arm's attached cycle ends Fixes kunchenguid#2381 When the Claude Stop hook's foreground arm attached to a peer watcher cycle and that cycle ended, the arm reported the delivered wake and the hook exited 2 without starting a successor, so the handling turn ran with no watcher. Pi, omp, and OpenCode start the next arm before delivering the wake and pass the closed arm's pid as FM_WATCH_PREDECESSOR_ARM_PID; the Claude hook never passed that predecessor identity. The hook now runs its arm as a tracked child it waits on, so it holds that arm's pid, and after any actionable close starts one handling-successor bin/fm-watch-arm.sh with the closed arm's pid as FM_WATCH_PREDECESSOR_ARM_PID. The successor is launched the one way a process outlives a Claude hook's exit-2 rewake (nohup, detached stdio, own process group, the shape bin/fm-startup-network.sh already uses); the hook waits for its status line and adds one banner line when no live watcher was confirmed, never withholding the wake. The supervision-host path is unchanged, as is the arm wrapper. The regression test drives the real hook against an arm fixture whose attached peer cycle ends: it fails on the previous tip because no successor starts, and now asserts the successor names the closed arm as its predecessor and outlives the rewake. A second case pins the unconfirmed-successor banner line. docs/watcher-continuity.md no longer records the Claude asymmetry. * no-mistakes(ci): Serial-4 failure was a real regression: tests/fm-session-lock-ancestry.test.sh asserts exact cumulative arm-invocation counts while driving the real fm-claude-stop-autoarm.sh hook against a stubbed fm-watch-arm.sh. This PR makes the hook start a handling successor after an actionable close, so every owned actionable phase now records TWO arm invocations (foreground arm + successor) instead of one, breaking "healthy chain: expected 1 arm(s), got 2". Fixed by updating the cumulative expectations to match the new behavior: owned phases 1/2/6 -> 2/4/6, foreign carry phases 3/4/5 -> 4, plus a comment explaining the +2-per-owned-phase model. Verified: phase-1 (the CI failure point) now passes on every run, syntax checks clean, and sibling arm-count tests (fm-claude-stop-autoarm.test.sh, fm-cursor-primary, fm-turnend-guard) pass unchanged. The only remaining local not-ok is a WSL-only environmental artifact (orphan reparents to a subreaper, not PID 1) that passes on the CI runner. Parallel-1 failure is an unrelated flake: its 11 tests (fm-lint, fm-pr-merge, fm-test-run, fm-cd-pretool-check, fm-pi-primary-types, fm-grok-harness, fm-composer-lib, fm-review-diff, fm-tmux-submit-busy, fm-composer-ghost, fm-brief) do not include fm-session-lock-ancestry and none reads any file this PR touches; all pass locally. It should clear on CI re-run. Made the smallest root-cause fix (one test file, 6 count updates + a clarifying comment). Validation of the branch continues through the no-mistakes pipeline, which owns re-running CI
…ning (kunchenguid#5566) * fix(bin): report a Lavish source armed only after its listener is running Registration alone was treated as ready, so arm could succeed before anything was collecting from the board. * no-mistakes(review): Guard Lavish arm launches, keep retire refusals, report live prior listener * no-mistakes(review): Keep polling through window before reporting a still-live prior listener * test: wait for a capture's claim to drop before the next arm The result is stored before the runner exits, so a re-arm in that gap was meeting a live claim. * no-mistakes(document): Record Lavish arm readiness evidence in verification doc * no-mistakes(ci): Both failures were caused by this PR, and both are fixed with test-only edits. Lint 2 (ShellCheck SC2034): this branch removed the only use of `reply_id` (a `start "$reply_id"` call) from tests/fm-procevent.test.sh, which left the assignment at line 1450 unused. I deleted that assignment. It was the only `reply_id` in the file. ShellCheck is now clean on both test files. Behavior portable serial 4: the failing test was tests/fm-bearings-board.test.sh, in the check "registration consumed its answer before the any-origin binding existed". I reproduced it locally: the hold was still `state: queued` when the test checked it. - What must hold: the test's check that the hold is closed must run after the listener has captured the answer. - Why it broke: the test used a stand-in adapter that ran `fm-procevent.sh start` in the foreground after `arm`, so capture finished before build returned. On this branch, `arm` starts the listener itself in the background, so the real listener captures the answer and closes the hold a moment after build returns. - Fix: removed the now-redundant stand-in adapter, the copied runtime directory, and its extra environment variables. The test now runs the real build through the existing `run_board` helper and waits up to about 10s for the hold to reach `state: done`. The checks that follow are unchanged: `Resolution mode: answered` and the any-origin binding. - Other tests: this was the only test in the file that stood in for the adapter this way. The shard's other pure-contract-unit test (tests/fm-trace-context-lib.test.sh) passed unchanged. Verification: - tests/fm-bearings-board.test.sh passed 3 times in a row via bin/fm-test-run.sh, all 18 checks, about 53s per run. - tests/fm-procevent.test.sh was not rerun, because the lint fix only removed an unused assignment
…elivered (kunchenguid#5599) * fix(bin): acknowledge a delivered unknown-wake escalation The same unrecognized wake was escalated again after it had already been handled, because delivery never recorded that identity. * no-mistakes(review): Scope unknown-wake acknowledgements to one away session * no-mistakes(review): Clear delivered digest when unknown-wake ack write fails * no-mistakes(review): Limit unknown-wake suppression to acknowledged lines * no-mistakes(document): List unknown-wake ack file among away-session artifacts
…ess wait (kunchenguid#5587) * fix(bin): keep a stated default retraction from cancelling a keyless wait A resolved line that names the shared default decision bucket was closing the keyless live wait that only prints as that same key. Keyless self-retraction still closes the keyless wait. * no-mistakes(review): Keep declared waits standing past foreign-key resolved lines * no-mistakes(review): Bound declared-wait read and share one decision-key parser * no-mistakes(document): Document supervisors' key-aware declared-wait read
…unchenguid#5544) * fix(bin): terminate a remote job worker that lost ownership when it receives TERM A serving worker whose lock directory is gone can no longer quarantine shutdown, and resuming service publishes a false ready heartbeat. Exit after stopping only that worker's own command tree, without removing a replacement owner's lock. * no-mistakes(review): Check worker lock ownership before publishing shutdown quarantine * no-mistakes(document): Correct worker shutdown comment on replacement-owned lock * fix(bin): keep an ousted remote job worker off the replacement quarantine Shutdown can lose the lock after the first ownership check and before it writes or clears quarantine. Bind both operations to the directory object this process still owns so a replacement's quarantine stays untouched. * no-mistakes(review): Make ousted-worker shutdown test reliably reach quarantine clear * no-mistakes(document): Reattach worker_shutdown doc comment to its function * no-mistakes(ci): Fixed the failing check (Behavior portable serial 7) with a test-only change to the stall test in tests/fm-remote-job.test.sh. Product code is unchanged; no other test changed. Cause: after the decoy dies, both workers run the same check-exists, read, delete sequence on the job records. On the CI runner the replacement deleted a record between the ousted worker's check and its read. The ousted worker exited 125, and because the file runs under set -e the unguarded `wait` ended the test with 125. The exit trap then killed the replacement, which produced the "Killed" line. Reproduction: a temporary 0.3 s delay between the check and the read, applied to the ousted worker only, made the committed test fail exactly as in CI (exit 125 and the "Killed" line). The new test passed with the same delay. The delay is reverted, along with a similar debug hook that the timed-out attempt had left in bin/fm-remote-job-worker.sh. Test changes: - The replacement is frozen (and confirmed stopped) before the decoy is killed and resumed only after the ousted worker exits, so only one worker touches the job records at a time. - The ousted worker is stopped only once its quarantine exists and its lane is reaped, which places it inside its stop loop. - Every fixed poll loop is now a wait on a named condition with a 30 s deadline and an explicit failure message. Exit detection also handles zombies. - The exit trap kills and waits for the decoy and both workers on every path. - A non-zero exit from the ousted worker now fails with its exit code and stderr instead of silently ending the file. The test still proves that the resumed ousted worker exits 0 and leaves the replacement's lock, quarantine contents and quarantine inode unchanged. Verification: the full test file passed four times on its own and three times under nice -n 10 with four busy-loop CPU hogs; bin/fm-lint.sh passes. Changes are not committed * no-mistakes(ci): I fixed the failing check (Behavior portable serial 7) by changing only the stall test in tests/fm-remote-job.test.sh. Product code is unchanged. **What failed:** "an ousted worker in shutdown leaves the replacement quarantine untouched" failed on CI with the ousted worker exiting 125 ("could not stop the active command tree"). **Why:** during shutdown, the worker retries the still-running decoy command group a fixed 100 times, 0.01 s apart, then gives up and exits 125. The test tried to freeze the worker partway through those retries by sending SIGSTOP from outside. On a slow runner the retries ran out before the stop arrived, so the worker had already given up. The invariant is that the test must hold the ousted worker inside that retry loop until the replacement owns the lock. That was the only place the test depended on timing. The other waits already watch for a named state change with a 30 s deadline. **Fix:** - The ousted worker now starts with a small `sleep` wrapper at the front of its PATH, and the SIGSTOP race is gone. - The wrapper only holds a `sleep` called directly by that worker's own process (it checks its parent pid against a hold file) while its quarantine file exists. - The only such `sleep` is the first retry in the shutdown stop loop, so the worker waits there as long as needed. - The wrapper writes a marker when it starts holding. The test waits for that marker, then hands the lock to the replacement, freezes the replacement, and kills the decoy. - The test releases the worker by deleting the hold file. Deleting the whole temp directory also releases it, so a failed run cannot leave the wrapper looping. - A process leak: the test overwrites the job's command-group record with the decoy, so no worker ever stopped the job's real command. `fm-hold-job.sh` and its `sleep 30` stayed running for up to 30 s after the test. The test now records that group before overwriting it and kills it at the end of the test and in the exit cleanup. - The test still asserts the same things: the ousted worker exits 0, and the replacement's lock, quarantine contents and quarantine inode are unchanged. **Verification:** - The full file passed twice on its own, twice under `nice -n 10` with six busy-loop CPU hogs, and twice more after the leak fix. - `pgrep` found no leftover processes afterwards. - With the worker from just before the fix commit (cf45cb6^), the test still fails with "the ousted worker wrote or cleared the replacement quarantine during shutdown", so it still proves the fix. - `bin/fm-lint.sh` passes. - I did not reproduce the CI failure locally. The cause comes from the fixed retry limit and the CI error message. The changes are not committed * no-mistakes(ci): I changed only the stall test ("an ousted worker in shutdown leaves the replacement quarantine untouched") in tests/fm-remote-job.test.sh. Product code is unchanged, and so is every other test. **Invariant:** the pid written to the job's group record must be a process-group leader whose group dies when that one process is killed. Otherwise the worker's bounded stop loop never sees the group die, gives up, and exits 125 ("could not stop the active command tree") before it reaches the lost-ownership exit. The decoy is the only place in this test that depends on this. **Fix:** - The decoy used to be `set -m; sleep 30 &`. It now starts as `perl -MPOSIX=setsid -e 'setsid() >= 0 or exit 1; exec @argv' sleep 30 &`, which gets its own session and group without shell job control. tests/fm-procevent.test.sh already uses the same idiom. - The test now waits, with the file's usual 30 s deadline and a named failure, until `ps -o pgid=` of the decoy equals its pid before writing it into the group record. This way the worker can never read the record before `setsid` has run. - The existing steps are unchanged: the test kills the decoy, reaps it with `wait` before releasing the hold file, and the exit trap still kills and reaps the decoy and both workers. - The assertions are unchanged: the ousted worker exits 0, and the replacement's lock pid, quarantine text and quarantine inode stay the same. **Cleanup:** I reverted a debug `printf` hook that the timed-out previous attempt had left in bin/fm-remote-job-worker.sh, and deleted its untracked `.tmp-repro/` directory. Neither was committed. **Verification:** - The full tests/fm-remote-job.test.sh passed twice normally and once under `setsid -w` with stdin from /dev/null (no controlling terminal). - `bin/fm-lint.sh` passes. - No leftover `sleep 30` processes afterwards. **Not reproduced:** I could not reproduce the CI failure locally. On this host `set -m` made the decoy its own group leader even without a controlling terminal, so the cause on the runner is not confirmed. The change removes the test's reliance on shell job control, as the user asked. Changes are not committed * no-mistakes(ci): I changed only the stall test ("an ousted worker in shutdown leaves the replacement quarantine untouched") in tests/fm-remote-job.test.sh. Product code is unchanged, and so is every other test. **Invariant:** the group record the ousted worker checks in its stop loop must stay the job's own command group, and the test must stop that group before it releases the hold. Otherwise the bounded retry keeps seeing a live group, gives up, and exits 125 ("could not stop the active command tree") before it reaches the lost-ownership exit. The test overwrote this record in one place (the decoy) and stopped the group in one place (killing the decoy); both are changed. **Fix:** - I removed the setsid decoy and the overwrite of `.claim/group`. The record keeps the job's real command group, which the test still saves as `STALL_JOB_GROUP`. - The two-line `group_start` stays. It is still needed: without it the worker kills the real group on its first pass, before the replacement takes over, so the hold would never matter. - The `sleep` wrapper that holds the worker at its first stop-loop retry is unchanged. - After the replacement owns the lock, its quarantine is planted and it is frozen, the test runs `kill -KILL -- -$STALL_JOB_GROUP`. It then waits, with the file's usual 30 s deadline and a named failure, until `kill -0` on the group fails. Only then does it remove the hold file. The worker therefore always sees its own command already stopped and never races its retry budget. - The exit trap still kills the saved command group if the test fails. It can't `wait` on that group because the group is not a child of the test shell. The decoy variable and its cleanup entry are gone. - The assertions are unchanged: the ousted worker exits 0, and the replacement's lock pid, quarantine text and quarantine inode stay the same. **Verification:** - The full tests/fm-remote-job.test.sh passed twice normally. - It passed once under `setsid -w` with stdin from /dev/null (no controlling terminal). - It passed once under `nice -n 10` with six busy-loop CPU hogs. - With the worker from before the fix (cf45cb6^), the test still fails with "the ousted worker wrote or cleared the replacement quarantine during shutdown", so it still proves the fix. - No `fm-hold-job` or `sleep 30` processes were left afterwards. - `bin/fm-lint.sh` passes. **Not reproduced:** I couldn't reproduce the CI failure locally; the decoy version also passed on this host. So I can't confirm why the decoy group stayed alive on the runner. The new wait turns any leftover live group into a clear named failure instead of an exit 125. The changes are not committed * fix(bin): keep a dead command group dead on bash 5.2 A bare return inside the liveness check drops the failing kill status when the check runs in a conditional, so shutdown keeps treating a stopped group as alive and exits 125. * no-mistakes(review): Use bash 3.2 fd syntax and fix trap return comments
…henguid#5589) * docs: make configuration settings easier to find and understand * no-mistakes(review): Restore dropped qualifiers and fix misplaced config doc labels * no-mistakes(review): Restore three dropped qualifiers in configuration reference
) * fix: bound worker edits of project AGENTS.md/CLAUDE.md to factual corrections These files are loaded into every agent session of a project, so additions should be a deliberate human choice rather than automated task output. The ship brief's project-memory section and AGENTS.md section 6 previously invited workers to record durable knowledge, which let project AGENTS.md files accrete detail the codebase or README already carries. Workers now edit only to fix factually wrong content - including content their own change made wrong - and fm-ensure-agents-md.sh runs only alongside such a correction. Stow no longer routes project-memory additions through ship tasks, and the generated skeleton no longer invites discovery-driven additions. * no-mistakes(review): Stop running fm-ensure-agents-md.sh on memory-file corrections * no-mistakes(document): Clarify manual project-memory initialization and remove duplicate guidance
…enguid#5635) * fix(bin): let gate agents drive lifecycle against marked lab homes Part 2 of the kunchenguid#5615 split. A no-mistakes gate agent runs inside a checkout carrying the fleet-captain identity, so fm-gate-refuse-lib refuses fleet mutation on the gate signal. That refusal was absolute, which kept gate validation from ever exercising the real lifecycle. Stamp a disposable lab FM_HOME with a .fm-lab-home marker file that only bin/fm-lab-home.sh writes, and only onto a fresh empty dir, so no call path can mark a populated real home. fm_refuse_if_gate_agent then permits lifecycle only when FM_HOME carries the marker and is driven through its stock layout - any FM_*_OVERRIDE relocation stays refused so part of the "lab" cannot be split back onto the real fleet. The threat model is a confused agent touching the real fleet, not deliberate forgery, so the marker is a plain token file rather than a bound record. FM_GATE_REFUSE_BYPASS is unchanged: it still serves the test harness, which cannot mark hundreds of temp homes. Teardown's slot-ownership scan compared state-dir paths textually while fm_firstmate_root_home canonicalizes, so a lab home under a symlinked TMPDIR scanned its own record twice and self-collided; compare file identity (-ef) instead. * no-mistakes(review): Refuse unlistable lab homes and hardlinked slot records * no-mistakes(review): Mint lab markers only on verified-empty fresh dirs * no-mistakes(document): Clarify lab-home gate documentation and comment contracts * no-mistakes(document): Clarify lab-home gate documentation and remove stale claims * no-mistakes(document): Clarify gate lab-home documentation and boundary wording
…ad of refusing every re-arm (kunchenguid#5594) * fix(bin): replace a watcher whose beacon stalls past a hard bound instead of refusing every re-arm A fleet watcher that is alive but whose liveness beacon has gone stale could never be replaced: every re-arm was refused because the lock holder was a live pid, and the holder was never evicted because it was not dead. Add FM_WATCHER_STALL_BOUND (default 3x the stale grace): below it the refusal is unchanged; at or past it the arm re-verifies the holder against the lock's recorded identity, sends TERM, waits boundedly, and takes the lock the normal way, ledgering a stalled-holder-replaced row. A holder that survives TERM keeps the old refusal. Fixes kunchenguid#4400 * no-mistakes(test): poll for replacement message to fix watcher-lock test flake * no-mistakes(document): document FM_WATCHER_STALL_BOUND in config inventory
…n can keep them (kunchenguid#5563) * fix(pi): hide queued Firstmate notifications under Calm only when the session can keep them Calm now keeps authenticated Firstmate operational inputs out of Pi's queued-message listing, but only after proving the live session exposes every member needed to keep them across Escape. A session missing any of them keeps stock rows and Escape and shows one generic warning. Escape and the dequeue key return only captain-authored messages to the editor and re-queue hidden notifications in order; after an abort that kept any in Pi's agent queue, the adapter starts the delivery turn itself because Pi 0.87.1 does not continue an aborted run. Compaction-held notifications stay with Pi's compaction flush and never start or announce a turn. Fixes kunchenguid#1588 * docs(calm): record Pi 0.87.1 queued-row retention verification * no-mistakes(review): Deliver kept Calm notifications after tree-navigation aborts too * no-mistakes(review): Defer Calm notification turn until tree navigation finishes * no-mistakes(lint): Silence SC2016 for literal JavaScript in queue-retention e2e test
…uid#5548) * fix(bin): refuse teardown when a required source disappears A missing sibling was sourced after cleanup had started, so Bash 3.2 exited 0 from the EXIT trap and Bash 5 continued and reported success. * no-mistakes(review): Remove unused FM_TEST_ONLY hook from teardown tests * no-mistakes(review): Check task backend sources before any teardown cleanup * test(gotmp): give teardown fixtures every tmux adapter sibling Teardown now refuses when a sibling the recorded backend's adapter sources is missing, so the fake bin must carry fm-session-lock-lib.sh, fm-agent-process-lib.sh and fm-gemini-lib.sh.
Restructure the supervision host doc's prose into shorter sections, lists, and tables without changing documented behavior. Every original heading, anchor, identifier, number, quoted string, and link target is preserved.
* docs: make herdr-backend easier to read Restructure the Herdr backend doc's prose into shorter sections, lists, numbered procedures, and tables without changing documented behavior. Every original heading, anchor, fenced code block, link target, and documented fact is kept. * no-mistakes(document): Restore composer-proof reason and complete Herdr topic table
Restructure the prose into sections, lists, and tables without changing documented behavior. Every original heading and anchor, inline-code span, link target, number, and quoted string is kept, and each sentence sits on its own line. Adds a topic navigation table and short subsections under the existing headings.
* docs: make watcher-continuity easier to read Restructure the prose into sections, lists, and tables without changing documented behavior. Every original heading, anchor, identifier, link target, and number is kept. * no-mistakes(review): Fix actor and supervision-host scope in watcher-continuity doc * no-mistakes(review): Make readiness TERM and retry conditional on unready successor
* docs: make sessionstart-nudge easier to read Restructure the prose into sections, lists, and tables without changing documented behavior. Every original heading, inline-code span, link target, number, and fact is preserved, and a harness-to-tier table now sits near the top. * no-mistakes(review): Drop helm glossary line and dedupe exit-code lead-in
* docs: make captain-hold-lifecycle easier to read Restructure the captain-hold lifecycle prose into sections, lists, and tables without changing documented behavior. Every original heading, anchor, identifier, number, quoted string, and link target is kept. * no-mistakes(review): Fix verification record subjects and grouping headings * no-mistakes(review): Clarify task-body read-back cases belong to the suite
* docs: make remote-secondmates easier to read Restructure the remote second mates prose into sections, lists, numbered procedures, and tables without changing documented behavior. Every original heading, anchor, fenced code block, identifier, link target, and qualifier is preserved. * no-mistakes(review): Merge remote-home table cell into one sentence * no-mistakes(review): Tighten readiness lead-in, restore causal link, fix dangling reference
…nguid#5554) * fix(bin): bound the away digest and log why a delivery failed The away daemon joined every buffered escalation into one unbounded digest. A start-up catch-all span can exceed what one transport argument carries (tmux rejects the send-keys command; Linux refuses to exec any argument above 131,071 bytes, which is how herdr receives it), so the initial send failed on every housekeeping pass and was logged as an unconfirmed Enter with text possibly in the composer. escalate_flush now builds the injected digest under a fixed byte budget: each event is cut at a UTF-8 boundary with an omitted-bytes marker, the joined events stop with a "+K more event(s)" tail, and a bounded digest names a state/.subsuper-digests/ file that keeps every buffered event verbatim. The buffer itself is untouched, so the return catch-up stays complete. The tmux submit core and the herdr literal send now replay the transport's stderr on failure, and inject_msg logs the failing stage (initial send versus Enter confirmation) with the byte count and that stderr. The wedge alarm line and marker carry the last failure reason. Fixes kunchenguid#4382 * no-mistakes(review): Drop digest pruning; label send-failed as send-or-Enter stage * no-mistakes(review): Keep digest full text once submit ran; reuse on retry * no-mistakes(lint): Count digest files with find instead of ls --------- Co-authored-by: firstmate-oss <firstmate@kunchenguid.local>
…henguid#5638) * feat(tests): add FM_TEST_SEAM launch seam and gate lab-primary recipe Part 1 of the kunchenguid#5615 split: the pieces that let the no-mistakes pipeline live-validate firstmate changes, without the gate-refusal rescoping. - bin/fm-afk-launch.sh: FM_TEST_HARNESS pins the detected harness only alongside the FM_TEST_SEAM=1 marker test suites set, so a leaked variable in a real primary's environment stays inert and unknown tokens fall through to real detection. - tests/lib.sh: export FM_TEST_SEAM=1 for every suite. - .no-mistakes.yaml: per-harness recipe for running a real fixture primary from a gate run - a plain mktemp lab FM_HOME on a private tmux socket, with FM_GATE_REFUSE_BYPASS=1 scoped to it and NO_MISTAKES_GATE scrubbed. - tests/fm-wake-queue.test.sh: stop the owned watcher fixture with KILL and clear its lifecycle state so the next leg starts clean; TERM could leave bash waiting in a child on some runners. - tests/fm-remote-secondmate-lifecycle-e2e.test.sh: wait for the liveness lock holder's post-acquire marker instead of the lock dir, which is published before the claim finishes. * no-mistakes(review): Scrub lab home overrides and require FM_TEST_SEAM separately * no-mistakes(document): Clarify test seam and disposable lab bypass documentation * no-mistakes(document): Clarify lab isolation and test-seam documentation * no-mistakes(ci): Fixed the CI failure: test cleanup killed the remote worker child but left its supervisor able to restart it during fixture removal. Cleanup now stops the worker tree. The lifecycle test passed locally; ShellCheck and diff checks passed
… home is gone (kunchenguid#5552) * fix(bin): refuse watchers from disposable checkouts and exit when the home is gone Fixes kunchenguid#321 Fixes kunchenguid#4760 A watcher armed from a disposable no-mistakes validation checkout under .no-mistakes/worktrees/ outlived the validation step and kept writing the real home's state, and a running watcher never noticed when its home, state directory, or code root disappeared. The arm now refuses from such a checkout with the typed failure line, the watcher checks once per poll that its home, state directory (or its own lock holder record), and bin directory still exist and exits with a logged reason scoped to itself, and the shared test helpers reap every watcher a suite armed for a temporary home through the home-scoped stop. * no-mistakes(lint): fix SC1007 by assigning empty string in watch-arm test * no-mistakes(ci): Found and fixed a genuine, reproducible hang introduced by this branch's test-watcher reaper, which is what killed both CI checks (serial-2 cancelled at the 30-min cap; Lint 2 exit 143 = the suite's own TERM-trap code). Root cause: test_drain_asserts_watcher_liveness (tests/fm-wake-queue.test.sh) fabricates a .watch.lock whose pid is the test runner's own $$ with the runner's real identity, to make the drain believe a live watcher exists. The new make_case tracking registers that state dir for reaping, so at fm_test_cleanup the new fm_test_reap_watchers drives fm-watch-arm.sh --stop; its identity check matches (the fixture recorded the runner's identity) and it kill -TERMs the test runner. tests/lib.sh:231 is `trap 'fm_test_cleanup; exit 143' TERM`, so the TERM re-enters cleanup -> reap -> kills $$ again -> infinite loop until the runner cap. I reproduced this locally: the suite ran all tests then looped forever in cleanup spawning fm-watch-arm.sh --stop against a lock naming its own PID. Fix (tests/lib.sh, +5 lines): in fm_test_reap_watchers, skip any tracked lock whose pid equals our own $$ before driving --stop. This is the single shared reap boundary; seven $$-self-lock fixtures across four test files are all covered by the one guard, and real armed watchers (pid != $$) are still reaped. Invariant: the test reaper must only signal real armed watcher processes, never the test runner itself. Verified locally: tests/fm-wake-queue.test.sh -> EXIT 0 (63 ok, no hang); tests/fm-watch-arm.test.sh -> EXIT 0 (21 ok, including test_reaper_stops_a_tracked_watcher, confirming the guard does not over-skip). Lint 2's exit 143 was the same shard/cap signature; a fresh CI run on this new commit will re-evaluate it --------- Co-authored-by: firstmate-oss <firstmate@kunchenguid.local>
…m as silence (kunchenguid#5588) * fix(bin): surface an unrecognized status prefix instead of dropping it A parked or holding declaration, and a verb whose correlation token did not parse, never became an event, so the supervisor still saw the earlier line. * no-mistakes(review): Require verb-shaped unrecognized status prefixes, add continuation tests * no-mistakes(document): Document unrecognized status prefix escalation in afk skill * no-mistakes(ci): I reproduced the "Behavior portable serial 6" failure locally and fixed it by changing the test data in one test. No product code changed. **What failed:** `tests/fm-session-start.test.sh`, in `test_orphan_status_logs_are_printed`, with "matched status log was printed 2 times". **Why:** the test writes status lines with made-up prefixes, `matched: surfaced once` and `orphan: step N`. The test only uses them as placeholder text. It checks that the session-start digest prints each task's status tail exactly once. This PR (kunchenguid#4763) deliberately makes an unrecognized one-word lowercase prefix a status event. So those lines now surface as captain-relevant events, and the wake queue's STATUS OUTCOME BACKSTOP section prints them a second time. The code under review is behaving as the issue asks. Only the test's placeholder data had become meaningful. **Rule the test depends on:** its status lines must not be captain-relevant, so the digest is the only place they are printed. Both lines in this test broke that rule. The orphan line would have failed the same count check right after the matched line did. **Fix:** in that test only, I switched both lines to the recognized, non-captain verb `working:`: `working: surfaced once` and `working: orphan step 1..6`. I updated the matching assertions and counts to use the new text. What the test checks is unchanged: orphan logs are labelled, the tail is bounded, the log path is printed, and each tail appears once. **Verification:** before the fix, the test failed locally the same way as in CI. After it, `bash tests/fm-session-start.test.sh` reports "all assertions passed * no-mistakes(review): Detect unrecognized prefixes on unstamped lines; share verb list --------- Co-authored-by: Kun's firstmate <kunchenguid+firstmate@users.noreply.github.com>
…henguid#5658) Fixes kunchenguid#5295 Session start now reports a remote inheritance failure using the push's own error line instead of the first unchanged item that happened to print before it, and the shared captain preferences header check now names the first required phrase it did not find, on both the local and remote inheritance paths.
…yloads (kunchenguid#5657) * fix(bin): stand down the Claude Stop auto-arm on pi-code-delivered payloads pi-code loads the tracked Claude settings but has no asyncRewake, so it awaits every Stop hook; without a stand-down the auto-arm runs synchronously inside Pi's turn end and holds it open for the declared multi-hour timeout. Stand down when the payload's transcript_path contains a /.pi/ path component, the same discriminator the closed-but- unmerged fix in kunchenguid#3352 used, with an explicit string-type check on the jq filter. Fixes kunchenguid#3343 * no-mistakes(document): document pi-code stand-down in harness integrations reference
…kunchenguid#5659) * fix(bin): match whole multi-word project names in the registry lookup bin/fm-project-mode.sh matched a registered project name against only the first whitespace-delimited token of a registry row, so a name containing a space never matched, silently defaulting the project to no-mistakes off instead of its declared posture. The lookup now matches the whole registered name against the raw line text, so a name is compared literally (never as a regex) and a name that is a leading prefix of another registered name still resolves to its own row. * no-mistakes(document): docs already accurate for multiword registry name match * chore: drop accidental empty err file Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com> --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
…d#5546) * fix(bin): classify the stdin program of `bash -s` with operands in the arm policy With -s, sh/bash/zsh read the program from stdin even when operands follow; the operands are only positional parameters. The arm policy treated the first operand as a script path, so heredoc and here-string payloads were never classified and a hidden bin/fm-watch.sh execution was allowed. A protected path in the operand position still fails closed as before. Fixes kunchenguid#1489 * no-mistakes(document): Clarify stdin shell operand documentation * no-mistakes(ci): Captain, fixed `shellInvocation` so `bash -- -s` treats `-s` as a script name, and updated R21 to test the exact command. The targeted policy suite, lint, documentation check, and diff check pass. Both hosted workflows show `action_required` before any jobs ran; that external approval state remains unresolved * fix(bin): keep main's handling of words after a leading `--` Revert the pipeline CI-step change that made the first word after a leading `--` always a script. It turned forms that main denies today into allow (for example `bash -- -c 'bin/fm-watch.sh'`), which is outside kunchenguid#1489 and loosens a fail-closed policy. `--` after `-s` still ends option parsing.
…nchenguid#5695) * fix(bin): strip AI co-author trailers from fleet-launched commits Cursor and other non-Claude runtimes append the trailer after the typed message. A per-task commit-msg hook removes it and leaves human co-authors and the author identity untouched. * no-mistakes(review): Export pane hooksPath override and drop generated-with stripping * no-mistakes(ci): This PR caused all three CI failures, and the fix is test-only: 4 test files change, no product code. **Cause.** `fm-spawn.sh` now installs the AI-trailer strip hooks for every spawn, secondmates included. The installer refuses a worktree that is not a git repository, and the PR deliberately keeps that fail-closed rule because real secondmate homes are firstmate clones. Four test fixtures still gave secondmates a plain directory as their home, so each spawn failed with "not a git worktree ... could not install the AI-trailer strip hooks": - serial 5: `tests/fm-backlog-atomicity.test.sh` ("secondmate spawn failed"). - serial 8: `tests/fm-secondmate-harness.test.sh` ("split: no meta written"). - Herdr: `tests/fm-backend-herdr-launcher-workspace-e2e.test.sh` and `tests/fm-backend-herdr-workspace-per-home-e2e.test.sh`. **Rule that must hold.** Every home a test spawns as a secondmate must be a git worktree. I checked the other places in the changed area: the only secondmate spawns in these tests are the ones listed. The earlier rounds already fixed the other fixtures (`fm-secondmate-liveness`, `fm-secondmate-safety`) the same way. **Fix.** - Each of those four secondmate homes now gets the same `.gitignore` plus `git init -q -b main` that the liveness and safety tests already use. - The two Herdr tests clean up with their own plain `rm -rf "$TMP_ROOT"`, not the shared `tests/lib.sh` helper. Because the installer leaves each `state/<id>.git-hooks` directory read-only, that cleanup printed "Permission denied" and left the directories behind. Both cleanups now restore the owner's write bit on every directory before removing (`find ... -exec chmod u+rwx`), which is what `fm_test_remove_tree` in `tests/lib.sh` does. **Verification.** - `tests/fm-secondmate-harness.test.sh` passes. - `tests/fm-backlog-atomicity.test.sh` passes (99 ok, exit 0). - shellcheck is clean on all four files. - I could not run the two real-Herdr tests locally: the Herdr lab on this host refuses to start because it needs exactly one running default session, and I did not change the host's Herdr state to get around that. Instead I checked their two changed steps directly: the installer succeeds on a home set up the new way, and the new cleanup removes the read-only hooks directory completely. Those two tests will only be proven on CI
…id#5683) * fix(bin): treat Pi's dollar-first cost footer as furniture An idle Pi status row opening with $0.000 was read as a dead-shell prompt, so exit and relaunch refused on an empty composer. * test: wait for the draining holder to exec sleep before reading its identity The procevent drain fixture read fm_pid_identity immediately after backgrounding setsid sleep, racing the child's exec chain. Mid-exec the cmdline can read empty, failing the fixture on a loaded CI runner. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…5534) * fix(bin): refuse a merge when a required check never reported fm-pr-merge.sh built its GitHub refusals only from checks present in statusCheckRollup, so a required check that never ran was simply absent and the merge proceeded on the subset that reported, contradicting its own "every required check green" claim. The GitHub verify now reads the base branch's required contexts from the forge itself - the classic branch protection summary on GET repos/{o}/{r}/branches/{b} and the active ruleset rules on GET repos/{o}/{r}/rules/branches/{b} - and refuses when a required context has no entry in the same rollup, at the same head, that the merge is bound to. Absence reads as unknown, never green. The required-set read joins the existing refusal list, so a draft, a red check, and an unreported required check are all reported together. Could not read vs nothing required: both endpoints need only repository read access. The admin-only GET .../branches/{b}/protection endpoint is deliberately not used: it answers a non-admin token with the same 404 an unprotected branch gets (observed live on kunchenguid/firstmate main with this token), which would read a missing permission as "nothing required". Any failed or malformed read of either source (auth, missing fine-grained permission, rate limit, network, 404, unexpected shape) refuses the merge with a line naming the unreadable source. The one exception is GitHub's plan-gated 403 on the rules endpoint ("Upgrade to GitHub Pro or make this repository public"), which already means "this repository has no branch rules" for the merge-queue reader; that check moves into one shared helper and the classic summary still decides for such a repository. Attended waiver: --allow-missing <check-name> is the twin of --allow-red and follows the same design and recording path: once, separate name argument, waives only that exact unreported required check, still requires every other required check reported and every check green, never waives an unreadable required set, refused while the away-posture record exists, and refused on GitLab. Merge-state BLOCKED policy is unchanged. How this differs from the withdrawn kunchenguid#5353 (read from its diff): - kunchenguid#5353 read the admin-only branches/{b}/protection endpoint and treated its 404 as "no required checks", so for any non-admin token the required set silently read as empty; this change reads the read-access branch summary and treats every failure as unreadable. - kunchenguid#5353 ignored rulesets; this change also reads required_status_checks rules from the effective branch rules. - kunchenguid#5353 made separate per-head REST reads of statuses and check-runs capped at per_page=100 with no pagination; this change checks presence in the same statusCheckRollup view the red-check gate already reads at the verified head. - kunchenguid#5353 stopped at the first unreadable read; this change reports it as one refusal among all the others. - kunchenguid#5353 also claimed kunchenguid#5345 (lock stealing) and changed 39 files, most unrelated; this change is kunchenguid#5344 only. Live proof, read-only (a gh wrapper refused every merge and mutating call): - cli/cli#14474 (trunk requires 3 classic build contexts, none ran): refused, naming build (macos-latest), build (ubuntu-latest), build (windows-latest); with --allow-missing "build (macos-latest)" it still refused, naming the other two. - cli/cli#13665 (ran build (ubuntu-24.04-firewall) instead): refused, naming build (ubuntu-latest). - hashicorp/terraform#39262 (ruleset-required checks absent): refused, naming Code Consistency Checks, End-to-end Tests, Race Tests, Unit Tests. - cli/cli#14485 (all required reported and green): verified; the wrapper blocked the merge call and the pull request read back open. Fixes kunchenguid#5344 * fix(review): Preserve required-check producers and aggregate independent read failures * fix(document): Clarify required-check verification and waiver documentation * fix(bin): match an app-bound required commit status by name The producer-identity check resolved an app-bound required context only against check runs, so a required context that the required app reports as a commit status could never match and always read as "has not reported". A commit status carries no app id to compare, so an app-bound requirement that arrives as a status now matches by name, as before producer binding; check runs keep requiring the configured producer app. Live, read-only: hashicorp/terraform#39262 requires license/cla from integration 865473, reported green as a commit status by the CLA app. The previous head refused it as unreported; this head no longer does, while still naming the four required check runs that never ran there. Refs kunchenguid#5344 * fix(document): Clarify accepted commit-status producer verification limitation --------- Co-authored-by: firstmate-oss <firstmate-oss@kunchenguid.local>
…uid#5696) * fix(bin): never offer a persistent secondmate for teardown The return brief's "Landed, cleanup due" scan listed every state/*.meta record carrying a pr= and a merge-notified marker without regard to kind, so a secondmate record holding a relayed child's merged PR put the mate itself up for "bin/fm-teardown.sh <mate>" cleanup. A secondmate is a persistent worker, never landed work. - bin/fm-afk-return.sh: skip kind=secondmate in the landed-cleanup scan. - bin/fm-pr-check.sh: refuse to record pr= or arm a merge watch on a kind=secondmate record before any side effect; a PR reported on its routed status channel belongs to a task in the mate's own home, which arms its own watch. - bin/fm-watch.sh: a merged result from a poll already armed on a secondmate retires the poll silently - no merge outcome, marker, or wake. * no-mistakes(document): Document secondmate merge-watch and return-brief exclusions * no-mistakes(ci): CI failed in an unchanged watcher-shutdown test whose three-second wait was sensitive to runner load. Increased the wait for both state- and home-deletion cases without changing watcher behavior. The full fm-watch-arm suite passed locally; syntax and diff checks passed
…unchenguid#5702) * fix(bin): refuse unknown dash-leading args in public-posting fm-x scripts fm-x-reply.sh collected any unrecognized argument into the positional pool and took the first one as the reply text, so an invocation like "fm-x-reply.sh <id> --followup --final <text>" posted the literal string "--final" to X and silently dropped the real text. Make argument parsing strict in every script that can post publicly: an unknown dash-leading argument, a dash-leading request_id/task id, a dash-leading option value, or a surplus positional now exits 2 with a usage error before any config load, outbox write, or network call. Reply text starting with '-' is still accepted via --text-file or stdin, and --help is honored wherever it appears instead of becoming text (a --help forwarded through fm-x-followup.sh would have counted as a posted follow-up and mutated the link). fm-x-link.sh and the fm-public-followup scripts already refuse unknown arguments; fm-x-poll.sh takes none. * no-mistakes(review): Refuse surplus follow-up text sources; drop post-ID help branches * no-mistakes(document): Clarify reply and follow-up argument usage * no-mistakes(review): Refuse dash-leading --text-file operands in fm-x-reply * no-mistakes(document): Correct follow-up argument parsing comment * no-mistakes(document): Document dismiss argument rejection in script header
…nchenguid#5701) * feat(bin): latch the supervision host after repeated engine errors Rung 3c-1 of the PR 5631 re-cut: the host copies the Pi branch's broken-session policy. Two consecutive engine errors latch the session; every away wake then reaches main with one supervision-host line for a five-minute cooldown, after which one wake probes the engine, and each failed probe doubles the cooldown up to one hour. A reported turn without an engine error clears it. The latch is kept per main session, engine, and model in state/.supervision-host-health, and the engine conversation now uses the same main-session key, which includes the lock holder's process identity so a recycled pid never shares either. Lifted from the validated 5631 tree and adapted to main's away-only host: the attended recovery line and attended cooldown pass-through are left for the attended core, so a recovery is only logged. * no-mistakes(document): Consolidate supervision-host latch documentation
…kunchenguid#5535) * fix(bin): absorb routine second-mate progress while surfacing routed replies Fixes kunchenguid#2959 A kind=secondmate task's status signal was never absorbable, so a healthy mate's routine working: and paused: appends woke the primary every time. signal_crew_provably_working now reads the mate's lines new since the watcher's classified position: a decision, blocker, terminal outcome, note:, correlation-marked line, or unknown verb still surfaces regardless of busy evidence, while unmarked working:, paused:, and resolved: fall through to the same provably-working absorb an ordinary crewmate gets. * no-mistakes(review): narrow secondmate routine absorb to working and paused
* fix(bin): use gh-axi for the ship DoD draft check * no-mistakes(review): use PR number not URL in gh-axi draft check
… only (kunchenguid#5520) * fix(bin): drop status prose from the inactive-outcome dedupe identity The inactive-outcome receipt fingerprint included the child's sanitized last status line, so a persistent child appending routine prose after one terminal outcome minted a fresh parent event per sentence. Bind the identity to incarnation, task id, terminal state, and PR only, keeping the last line in the record as status_head evidence. Fixes kunchenguid#2960 * no-mistakes(document): note structured-only inactive receipt identity in regression coverage
Treehouse keys a worktree pool by the repository's resolved origin rather than by the clone, so every home holding its own clone of one repository allocated from a single shared pool. Homes competed for the same numbered slots, and a returned slot read free while its checkout was still a linked worktree of another home's clone; fm-spawn.sh's isolation assertion then refused it and the task stopped until a human released the slot from the home that owned it. Derive a per-home worktree root from the home's own resolved path and pass it to both acquisition sites, so two homes never see each other's slots. Returns stay flagless on purpose: treehouse resolves a return's pool from the worktree path it is handed, so every worktree already leased under the previously shared root stays returnable and nothing is migrated. Bootstrap now probes the global --root flag alongside the durable lease, and the CI treehouse pin moves to v2.3.0 because v2.0.1 has no such flag.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closed — duplicate of upstream kunchenguid#4932. Attested head and CI are tracked on the upstream PR.