feat: integrate Claude supervision and runtime updates - #96
Merged
Merged
Conversation
… untouched (kunchenguid#4243) * fix(teardown): refuse to return a Treehouse pool slot reassigned to another task A pool slot is reused across tasks, so a finished task's worktree= line can name a slot a different, live task now holds. Teardown already refused when a second task record named the same live path, but that scan cannot prove the record it is tearing down is the current owner: the task that took the slot next may leave no record the scan can reach - its own worker may have exited and its record been cleaned up, or it may live in a home this machine does not register. Teardown then killed every process under the path, hard-reset it and returned it, and its unlanded-work refusal never fired because it was inspecting a directory that no longer belonged to the task being torn down (observed 2026-09-07). Treehouse's own state file cannot answer the ownership question. It records a slot's owner as a live process lease (owner_pid plus owner_started_at, with `treehouse status` reporting in-use from the processes actually running under the path), which names no task and is released by the very event that makes a record stale - the worker exiting. An unleased slot therefore reads identical whether it is still this task's or has since been handed on, and a slot whose new holder has also exited but left uncommitted work reads as free. So the identity source is Firstmate's own claim, not Treehouse's lease. fm-spawn writes that claim - the task id - into the slot at the moment it takes it, under the same project lock that allocates the slot, and fm-teardown drops it only after the slot is genuinely returned. It lives at <pool>/<slot>/.fm-slot-owner, a sibling of the repo checkout rather than a file inside it, so claiming a slot can never dirty the copy the landed-work checks inspect. A claim naming another task, or one that cannot be read, refuses; --force does not lift either refusal, because --force authorizes discarding this task's unlanded work, never another task's live work. A slot that cannot be claimed refuses the spawn instead. An absent claim proceeds on exactly the record-scan protection it had before: slots taken before claims existed, and slots already returned, carry none, and refusing those would strand every task in flight across this change on no evidence at all. The refusal is deliberately all-or-nothing rather than partially completing the task's own cleanup. state/<id>.meta is the only durable record naming the worktree and endpoint, so removing it would destroy the evidence needed to reconcile which record is wrong, and its removal is one step with the backlog transition. Nothing is stranded: clearing the stale worktree= line leaves a record with no slot to release, which then tears down normally, and the refusal names that remedy. Repairing the previous claimant's stale worktree= line at spawn time is left for separate work. It would have the new owner write another task's record - the same class of cross-task mutation this bug is - and would need that record's own meta lock; with the claim in place teardown refuses on evidence rather than depending on the stale pointer having been scrubbed. For the same reason the relaunch path writes no claim: it holds no allocation lock, and a record whose worktree= is already stale would stamp the wrong task's claim onto a live sibling's slot. The regression reproduces the reuse sequence with only one discoverable record, including a clean, fully landed ship copy torn down without --force - the shape of the real incident, which the previous code returned to the pool - and fails against the previous code; the existing two-record, cross-home, own-slot and no-claim cases still pass unchanged. This builds ON upstream b028e8b (kunchenguid#3837), which is already in this branch's base (origin/main 40c50ea) and owns the record-exclusivity scan. Nothing here replaces that scan; the claim is the positive proof it cannot supply. Claude-Session: https://claude.ai/code/session_01JTBmuqKugaPUj7k9TXQwFS * no-mistakes(review): teardown leaves reassigned slot; spawn abort drops claim * no-mistakes(review): narrow Treehouse lease evidence; gate abort claim release on lock * no-mistakes(review): pin spawn-side slot claim; narrow abort-release header * no-mistakes(document): docs: point slot-claim rationale at fm-wake-lib owner
…nguid#4247) The reviewer treats Captain's intent as acceptance criteria, so a widened ask there drives over-built work; the spec should carry only what the ask requires.
* feat(bearings): name the Underway rows and order Charted Next newest filed first The fleet board's Underway rows led with the run status alone, so a scan told the captain where a pipeline stood but never which task the row was, and Charted Next rendered in backlog order rather than by when work was filed. The snapshot now projects the durable task name onto every in_flight row - from this home's backlog title, and from a secondmate home's own ledger for an active child - and the durable filed date onto every gate. The board's Underway row leads with that name and keeps the run status on its second line, and Charted Next renders newest filed first, with rows carrying no comparable date keeping their payload order after every dated row. The payload validator requires an explicit name marker on every Underway row and refuses a filed value that is not an ISO date, so the board can never sort on garbage or invent a label. * no-mistakes(review): Fix Bearings labels, bounds, and filed validation * no-mistakes(review): Fix Bearings identifiers and eligible queue bounds * no-mistakes(document): Document Bearings labels and newest-first bounds * no-mistakes(ci): Updated the stock macOS Bash CI expectation from 56 to 59 Bearings tests. Verified the suite under /bin/bash 3.2: all 59 tests pass. git diff --check also passes
…nchenguid#4248) * fix(bearings): report the away-return catch-up instead of refusing A captain returning from away and asking for bearings got zero bytes and an error: fm-bearings-snapshot.sh ran the away-return guard with `|| exit $?` before reading any fleet state, so the mere existence of the catch-up gate killed every bearings mode (and /ahoy with them). Bearings now consults that guard rather than obeying it. fm-afk-return.sh separates its two refusal branches by exit status, so an ACTIVE away window still refuses exactly as before - the right answer there is to run the return first - while return catch-up (exit 4) lets collection and projection proceed and is disclosed as one action-free `(return-catchup)` gate row, following the existing `(main-inventory)` precedent. It stays out of decisions_open: these blockers are firstmate-actionable, not the captain's own call, and the per-task blockers already project as their own Underway rows. The guard's refusal text also stops promising a blocker list it cannot produce: a gate retained for a lifecycle reason alone now names that retention reason, and bearings carries the same reason in the gate row's title. Reporting is not ordinary work. AGENTS.md already scopes the return hold to work rather than reporting, so only the /afk and bearings skills needed the correction. * no-mistakes(document): Refresh away-return Bearings verification * no-mistakes(review): Reserve catch-up gate outside Bearings truncation * no-mistakes(review): Preserve filed dates in catch-up gate output * no-mistakes(document): Document reserved catch-up gate projection
…forked code-root copy (kunchenguid#4223) * fix(backlog): address the home's backlog from any directory and detect a forked code-root copy A home outside the code root forks its queue: the tracked .tasks.toml names data/backlog.md relative to tasks-axi's working directory, so a bare tasks-axi call from the code root writes the code root's data/ while session start, spawn, and teardown use $FM_HOME/data. Linking the code-root copy into the home does not hold, because tasks-axi 0.2.4 writes by renaming a temp file over its target and rename(2) replaces a symlink: add, start, hold, and done from the code root each turn the link back into a regular file. The archive path is resolved against the working directory too, even with --file. bin/fm-tasks-axi.sh runs tasks-axi against this home's backlog from any directory, using the lifecycle transitions' existing addressing (run from the data directory's parent, pin <data>/backlog.md through TASKS_AXI_FILE). It keeps relative --to/--*-file arguments meaning the caller's paths, and refuses a caller --file, an unresolvable home, and a symlinked home backlog. The fm-send hold lookup, fm-public-followup, and the fm-decision-hold shim, which relied on cwd discovery, now go through it with an explicit FM_HOME and a cleared data override, so they keep addressing exactly $FM_HOME/data and an ambient TASKS_AXI_FILE cannot divert them; every agent-facing backlog command names it instead of bare tasks-axi. Bootstrap gains a detect-only BACKLOG_RECONCILE check, also run read-only: when the home's data directory is not the code root's, a code-root data/backlog.md or data/done-archive.md that is not the home's own file is reported as a fork, with the merge procedure in bootstrap-diagnostics. * test(teardown): assert the completion hint names bin/fm-tasks-axi.sh ready The completion hint now points at the home-addressed command instead of a bare tasks-axi call, so the dependency-cleared follow-up assertion checks for that command. * no-mistakes(test): clear ambient tasks-axi env in tests/lib.sh * no-mistakes(document): drop bare tasks-axi example from cd-guard doc * no-mistakes(lint): replace ls -A decoy listing with find for SC2012 * no-mistakes: apply CI fixes * revert: keep the compliance gate unchanged; the synchronize race is filed separately
* fix(spawn): pre-register Claude workspace trust for secondmate homes A claude --secondmate launch skipped workspace-trust registration entirely, so a standalone-clone secondmate home (an explicit ~/fm-homes/<id> path) had no store entry and its pane wedged on the "Is this a project you trust?" dialog before it read its charter. The step was gated on the task kind rather than on the harness, so the spawn's fail-closed guard had nothing to run against and reported a launch that could never start work. fm-claude-trust.sh gains a secondmate-home mode. A secondmate home is a whole firstmate instance, produced either as a leased worktree or as a standalone clone, so the linked-worktree test cannot decide it and the seed is the evidence instead: the .fm-secondmate-home marker must be a regular file this user owns naming exactly the id being spawned, the home must hold AGENTS.md and bin/, and each operational directory must resolve inside the home. That is the set fm-home-seed.sh writes and fm-spawn.sh's own home validation re-checks, so nothing wider than a home a secondmate spawn would launch into can earn home-level trust. The worktree path is unchanged, and still refuses a home. fm-spawn.sh now runs the registration for every claude launch and keeps refusing the spawn when it fails, rather than launching an agent that would wedge. * no-mistakes(document): Correct Claude secondmate trust guidance
* fix(pr-merge): judge each required check by its current run When the base branch advances, GitHub cancels a pull request's in-flight run and re-triggers it. The cancelled run stays in statusCheckRollup beside the passing re-run, so the rollup can hold several runs of one check name at the same head while GitHub itself reports the pull request CLEAN. github_checks_not_green judged every run independently, so that superseded failure refused a genuinely mergeable pull request and pushed the operator toward a needless --allow-red. Group the rollup by the reported name and judge each check by its current run. Supersession is proven, never assumed: a name leaves the red set only when every one of its non-green runs is strictly older than one of its green runs, dated by the forge's own settled timestamp - a check run's completedAt once its status is COMPLETED, or a status context's createdAt - and only in the whole-second UTC form GitHub emits, which is the one spelling that orders correctly as plain text. A run with no such timestamp is never superseded, so a still-running, queued or undated run keeps its check red, and a name with no green run at all stays red. An unnamed entry is grouped alone so two unrelated unnamed checks are never treated as one. Every comparison is one-directional: it can only clear a failure a later success provably replaced, and never clears a check whose current run failed, is pending, or is missing. No other guard moves - the pull request must still be open, undrafted, mergeable, conflict-free and head-bound, and --allow-red still waives exactly its named check with every other check green. Live reproduction: PR kunchenguid#4224 read CLEAN with an old FAILURE and a newer SUCCESS for one check name and was refused; it now verifies, while kunchenguid#4208 and kunchenguid#4210, whose latest runs failed, still refuse. * no-mistakes(review): Use check-run start times for safe supersession * no-mistakes(document): Clarify GitHub check-rollup documentation
…guid#4266) * fix(merge): persist the merge authority on poll-detected merge outcomes The merge ledger tags a merge with the authority that permitted it while the away-posture record existed, but only the direct attended merge in bin/fm-pr-merge.sh recorded it. A merge the forge queued, or one the merge poll detected after the fact, published an untagged row, so exactly the merges no agent watched were the least auditable. bin/fm-merge-authority-lib.sh now owns that answer, read from the same structured sources the merge gate already used: the task's recorded yolo posture and the away-posture record's mechanical grant list, never prose. bin/fm-pr-merge.sh keeps its own refusal wording and gates on that answer; bin/fm-watch.sh only records it on the row its poll publishes, so reading the authority never becomes a second path to a merge. An unresolved answer records an untagged row rather than dropping the outcome or inventing an authority. * no-mistakes(review): Persist canonical merge authority for queued poll outcomes * no-mistakes(review): Harden merge authority persistence against lifecycle races * no-mistakes(review): Serialize poll authority publication with teardown * no-mistakes(document): Clarify persisted merge authority lifecycle * no-mistakes(ci): Added targeted SC2034 suppressions for the two public result assignments in bin/fm-merge-authority-lib.sh. Verified successfully with `CI=true bin/fm-lint.sh`
…4281) The 2026-09-12 Actions starvation incident found firstmate CI with no concurrency deduplication, so every superseded PR head kept its full 13-job fan-out, and four jobs with no timeout at all. Add per-PR supersession keyed on the PR number for pull_request events and on the unique run id for push events, cancelling only pull_request runs, so a new PR head replaces its own in-flight CI while every main push keeps its own group and is never cancelled. Add hang tripwires to the four previously unbounded jobs: 25 minutes for lint (measured at 14-16 minutes) and 5 minutes each for the coverage guard, the timing aggregate, and the repo invariants. Measured lane bounds are unchanged. tests/fm-ci-workflow.test.sh resolves the workflow's concurrency expressions against simulated pull_request and push contexts and holds every job's finite timeout.
…henguid#4288) Every other make_hold_home caller in this file skips when tasks-axi is absent; this test was the one unguarded call, so hosts without tasks-axi hard-fail the fixture build instead of skipping.
…nnot blind a session start (kunchenguid#4027) * fix(bin): bound each backlog row read so one wedged backend cannot blind a session start bin/fm-bootstrap.sh's reconcile and close-replay sweeps read the backlog backend once per item through fm_backlog_row_show, and that read was unbounded. A single wedged `tasks-axi show` therefore consumed the whole FM_SESSION_START_TIMEOUT and truncated the digest before the wake queue, supervision instructions, fleet state, and context sections ever printed, leaving the fleet unsupervised with no live watcher. The harm was a blind startup, not a slow one. Bound the read with the existing shared timeout primitive (bin/fm-timeout-lib.sh), so a wedged backend degrades to a loud partial reconcile: the sweep's existing BACKLOG_RECONCILE diagnostic names the item it could not read and the loop continues to the next one. The first bound hit also latches FM_BACKLOG_ROW_SHOW_WEDGED, so a sweep over many items pays one bound rather than one per item and still names every item it skipped, which is what keeps the digest whole on a home carrying a large fleet. The bound holds regardless of any particular tasks-axi install, so it does not depend on the 0.2.5 `show` hang being resolved separately. * fix(bin): set the wedged-backend latch where it survives, and prove it The latch added with the read bound was inert. fm_backlog_row_show runs inside a command substitution in both of its status-capturing callers, so the subshell read the inherited value correctly but its write died with the subshell. Every item still paid a full bound and reported `exceeded`, never `skipped`, which left the large-fleet case the latch existed to cover completely uncovered. Move the write to the two callers that capture the read's status and own the surviving shell, and leave fm_backlog_row_show reading the latch only. Correct the comments that claimed an ownership the function never had. The test that was supposed to cover this asserted only that the second read finished under a generous ceiling, which is true whether or not the latch works. Assert instead that a latched read is strictly faster than one bound and that it reports its own item as skipped, so an inert latch fails the test. * test: cover every item the wedged-backend latch skips The latch assertion exercised a single skipped item, so "every skipped item is still named" was inferred rather than tested. Probe three items instead and assert each skipped one names itself and costs less than a bound. Verified as a real guard by removing both latch writes: the suite then fails on the first skipped item instead of passing. * no-mistakes(review): distinguish backlog read-bound hits from absent rows * no-mistakes(review): preserve read-bound status through the captain verify gates * no-mistakes(review): Preserve backlog read-bound hits through resolve_entry and reconcile instead of spending them as absent rows * no-mistakes(review): Preserve backlog read-bound 124 through migrated-prefix scan and remaining task_show call sites * no-mistakes(document): Document bounded backlog row reads and FM_BACKLOG_ROW_TIMEOUT_SECS * no-mistakes(ci): Fixed all four failing CI checks with one root-cause fix plus one test-heredity fix. (1) bin/fm-captain-hold.sh: task_show carries the row in TASK_SHOW_OUTPUT and emits no stdout, but four call sites still used the stale command-substitution convention show=$(task_show ...), leaving show empty: task_show_or_fail (every captain hold failed with 'did not retain its hold-set stamp' - broke fm-captain-hold-lifecycle in parallel 1 and fm-bearings-board in serial 3), resolve_migrated_entry (migrated-prefix resolution could never match), reconcile-requests (existing rows were refused as absent), and command_open --identity (printed a constant '#0' identity, so fm-watch-triage's re-held captain call inherited the previous call's silence in serial 1). This is also the Greptile P1. Fixed by invoking task_show in the current shell and reading show=$TASK_SHOW_OUTPUT, the convention the other eight call sites already use; read-bound hits still stop loudly by name. (2) tests/fm-backlog-read-bound.test.sh (serial 4, unclassified family): the new e2e half implicitly relied on the author's process tree containing a harness process so fm-lock.sh would grant the fleet lock; on CI runners the lock is refused, the reconcile sweep is skipped, and the final BACKLOG_RECONCILE assertion fails. Reproduced by simulating a CI ancestry via a ps shim, fixed by pinning the lock evidence with the established fake-ps harness fixture pattern from tests/fm-session-start.test.sh. Verified: shellcheck clean; parallel-1, serial-3, and serial-4 lanes fully green locally (failed=0); serial-1 lane green except fm-gemini-harness, which fails only under local Node v26 (comm=node-MainThread); CI's default Node 22 reports comm=node, the branch that test passes on, so it is not a CI failure * no-mistakes(document): Verified bounded backlog read docs accurate across branch
…guid#4285) * fix(merge): serialize the away-authority check with a synchronous merge bin/fm-pr-merge.sh read the away-posture record for merge authority (the per-task merge grant and the yolo/away-grant decision) and handed the merge to the forge afterwards. An archive at the captain's return or a grant revoked by a replacement record could land in between, so a merge could proceed on away authority that no longer held. The away record now carries a cross-subsystem lock, built on the existing bounded lock primitive rather than a new lock format: the record-mutating subcommands hold it across their mutation, and the merge holds it across both its authority read and the forge command. Because a queued or auto merge returns before the pull request lands, and would therefore outlive the lock, an away merge is now refused whenever it could land asynchronously: a requested --auto, a base branch whose merge-queue state does not prove an immediate merge, and GitLab's asynchronous flags and configuration. What remains permitted while away is the synchronous merge that lands inside the lock. This closes the common away-record/merge race against a live lock owner. It does not make the merge atomic in every case, and two narrow races are accepted and documented at their sites rather than hidden, both confused-agent-grade in the sense bin/fm-lease-lib.sh already uses: - A merge-queue rule change or a PR base change in the window between the queue-free preflight and the forge call can still enqueue the merge, which can then land after its grant lapses. - Killing the lock-owning shell while its gh or glab child is still running lets stale-owner recovery reclaim the lock and the record be archived or replaced, after which the orphaned child can complete the merge on lapsed authority. Closing either one needs landing verification or an ownership handoff, which is deliberately out of scope here. No existing gate is relaxed. The lock is taken after the live green-at-head verify and the captain-hold check, the in-lock authority read is unchanged, and a lock that cannot be taken refuses the merge rather than proceeding unlocked. The away grant stays a structured field; no prose is parsed. * no-mistakes(review): Fix GitHub rollup fixture base branch * no-mistakes(document): Document atomic away-authority merge locking * no-mistakes(ci): Updated two executable GitHub API fixtures to include the required baseRefName. Both previously failing test suites now pass: fm-captain-hold-lifecycle.test.sh and fm-pr-check-security.test.sh. git diff --check also passes
…unchenguid#4200) * feat(agy): verify Antigravity CLI as third worker/scout adapter Detection by anchored ancestry in fm-harness.sh (no marker of its own); bootstrap harness and effort validation; launch template with model and effort mapping plus reachable-catalog model validation; rendered-tail busy fallback in fm-busy-lib.sh with delivery footer in fm-composer-lib.sh; control mechanics with crewmate/scout-only refusal; tmux liveness naming; router entry with concise adapter reference; dated verification record; portable regression plus opt-in live drift guard. Verified live on agy 1.2.0: supervised spawn, durable steering, same-copy relaunch, and exit, with Herdr-native busy agreement. * no-mistakes(review): bound agy model probe, gate trust dialog, narrow busy signature * no-mistakes(review): pre-register agy workspace trust, make readiness gate strict * no-mistakes(review): Close Orca terminal on gate failure; isolate live-guard HOME; tighten agy matching * no-mistakes(document): Document agy adapter in stale harness enumerations * no-mistakes(review): Clamp non-positive FM_AGY_MODELS_TIMEOUT to the default bound * no-mistakes(document): Fix stale test-shard snapshots after agy lane additions * no-mistakes(ci): Fixed ci-3 (tests/fm-agy-harness.test.sh:519). Root cause: the agy spawn fixture's default base PATH (/usr/bin:/bin:/usr/sbin:/sbin) omits node's directory, but the spawn drives the real bin/fm-agy-trust.sh (which hard-requires node to record trust) and the fixture's fake tmux trust lookup (node -e) under that PATH. On the ubuntu-latest CI runner node lives in the toolcache (/usr/local/bin), so trust pre-registration failed on portable serial 2; on typical Arch hosts node is in /usr/bin, masking the defect. Fix (smallest, following the existing tests/fm-kimi-harness.test.sh precedent of carrying the interpreter's resolved directory): resolve node from the invoking environment (failing the test with 'test needs node' if absent, as kimi does for python3) and prepend its directory to the fixture's default base PATH; the FM_TEST_BASE_PATH override contract is untouched. Verified locally: (1) pre-fix reproduction with a CI-shaped base PATH (system bins minus node) produced exactly the reported failure — 'node is required to record workspace trust and was not found on PATH' plus the fake tmux 'node: command not found'; (2) post-fix, all 29 tests in the file pass both with node available only via a leading non-standard dir in the base PATH (CI's shape) and with the default base PATH on this host. bash -n clean; ShellCheck is not installed in this worktree (previously recorded as environmental) * no-mistakes(test): Give agy typed sends a longer submit-confirm budget * no-mistakes(document): Document agy send budget, trust gate, and control coverage * no-mistakes(document): Document agy busy fallback inventory and send-timing evidence
…uid#4337) * feat(afk): add quiet supervision mode for a present captain Adds a first-class quiet supervision mode alongside /afk for kunchenguid#2356: the same away-mode daemon, injection, busy/composer guards, classification policy, and reliability properties, but the captain staying present and chatting no longer exits it - only an explicit /quiet off does. state/.afk's first line now declares its mode (away, the default, or quiet); fm_afk_mode() in bin/fm-wake-lib.sh is the single reader, falling back to away for missing/empty/unreadable/unrecognized content (including the legacy bare-epoch-timestamp format written before mode existed) so nothing regresses. fm_afk_flag_write() preserves the on-disk mode on a bare refresh (no explicit mode given) rather than defaulting to away, which is what keeps the daemon's own redundant terminal-side re-write from silently resetting a captain's quiet mode back to away underneath them. New .agents/skills/quiet/SKILL.md is a thin wrapper cross-referencing /afk for every shared mechanism, per the one-owner rule. AGENTS.md gains the state/.afk table entry and section 8's exit-trigger line. bin/fm-supervision-instructions.sh, bin/fm-session-start.sh, and bin/fm-guard.sh's stale-watcher banner all become mode-aware so a quiet-mode captain is never misdirected to /afk in captain-facing text. Closes kunchenguid#2356 * no-mistakes(review): Fix AFK epoch parsing and quiet-mode digest wording for two-line flag * no-mistakes(document): Fix turnend-guard.md daemon-ownership contract for quiet mode --------- Co-authored-by: NewAiCoder <claude@theinbtw.com> Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
…kunchenguid#3578) * fix(bin): let verified harness ancestry outrank retained markers (#3) * fix(bin): let a structural harness ancestor outrank a retained marker bin/fm-harness.sh treated a verified environment marker as unconditionally authoritative, so a Codex session started from an environment that had retained CLAUDECODE=1 detected as claude. Session start then emitted Claude's Stop-owned supervision protocol to a Codex primary, and every turn end was blocked for missing Claude recovery. The defect is the precedence boundary, not any one harness. codex, opencode, kimi, and muse publish no identity marker at all, so with markers winning outright any retained CLAUDECODE renamed them; the Cursor-before-Claude ordering was a point patch on the same class of problem, and the launch-time marker clearing only ever covered sessions fm-spawn started. Markers and ancestry are now separate evidence layers that detect_own arbitrates: - no ancestry match, or no marker: the single available layer answers, unchanged; - same harness family: the marker's finer verdict stands, so a launch-selected pi-signed is not flattened to pi by an ancestry walk that can only see the shared launcher name; - different harness with a structural (command-name) ancestor: ancestry wins, because only ancestry proves who owns the process tree; - different harness with only a bare-interpreter script-path match: the marker wins, since a harness-shaped path in some node process's arguments is weaker evidence than a harness publishing its own identity. The correction is symmetric: a retained CURSOR_AGENT no longer renames a claude worker nested under cursor either. Adds fm-harness.sh ancestry [<pid>], ancestry evidence with no marker layer, so a real harness process can be asked what the walk makes of it. tests/fm-harness-precedence.test.sh is the portable regression, built from real renamed processes with no harness installed. Every case drives the two layers apart and asserts each alone as well as the combination, so no case can pass vacuously; it also pins Codex's real two-process install topology, since the fix depends on the native binary being what a tool subprocess meets first. The opt-in drift guard gains the matching live half: each installed harness's real running process must still be identified by the ancestry walk, and it fails naming the harness and version when a release changes that name. Documentation follows the corrected contract in the script header, the harness-adapters detection section, the codex, opencode, kimi, and cursor references, and a dated verification record. * fix(tests): drop the unused argument pass-through in the shim-topology helper bin/fm-lint.sh refused the branch: run_shim declared a `[ancestry]` argument and forwarded "$@", but every call site that varies the environment or passes the ancestry subcommand invokes the shim entry point directly, so the helper is only ever called with no arguments (ShellCheck SC2120/SC2119). Behavior is unchanged: with no arguments "$@" expanded to nothing. * fix(bin): examine the top of the process chain instead of assuming init harness_ancestry stopped as soon as the next pid was 1, on the assumption that pid 1 is always init and can never be a harness. Inside a PID namespace that assumption inverts: the harness itself is pid 1, so the walk never examined the one process that proves who owns the tree, reported no ancestry at all, and handed the verdict straight back to a retained marker. A real Codex session under `codex sandbox`, holding CLAUDECODE=1 and CLAUDE_CODE_ENTRYPOINT=cli, is exactly that shape: it resolved claude and rendered Claude's Stop-owned supervision protocol even with the marker-vs-ancestry precedence boundary in place. The same probe now resolves codex and renders the Codex foreground checkpoint. A host's real pid 1 (init, systemd, launchd) matches no harness name, so examining it costs one ps call and can introduce no false positive; the walk still stops once that top process has been read, and a non-numeric or zero ppid still ends it. tests/fm-harness-precedence.test.sh pins the namespace shape with a fake ps that reports every process as bash with ppid 1 and pid 1 as the harness. The case asserts the marker still answers alone when pid 1 is host-shaped, so it cannot pass vacuously, and it fails against the previous stop condition. * docs(verification): record the real-Codex retained-marker evidence The existing record proved the precedence boundary with the portable regression and recorded each installed harness's process name behind the ancestry walk, but it had no evidence from a real Codex process actually holding a retained Claude marker, which is the failure the boundary exists for. Adds the dated before/after result from codex-cli 0.152.0 under `codex sandbox`, with the exact command and the decisive verdict and rendered protocol on each side, and records the second boundary that shape exposed: the walk must examine the top of the process chain, because inside a PID namespace the harness is pid 1. Refreshes the portable regression's observed output for the case it gained. * no-mistakes(review): blind ancestry in marker-pinned harness tests * no-mistakes(review): blind ancestry in the Pi guard-routing test * no-mistakes(review): classify precedence suite, dedupe ps stub, soften claims * no-mistakes(review): model the spawn-and-wait Codex shim topology * no-mistakes(document): correct stale muse marker-clearing detection claims * no-mistakes: apply CI fixes * fix(bin): examine the top of the chain in the lock and nudge walks too The pid-1 defect corrected in bin/fm-harness.sh survived unchanged in the two other harness-ancestry walks, on the exact topology the branch verified against a real Codex process. bin/fm-session-lock-lib.sh's fm_harness_ancestry_pids stopped as soon as the next pid was 1, so a firstmate whose harness is pid 1 of its own PID namespace could not find that harness at all and did not recognize its own session lock. bin/fm-sessionstart-nudge.sh carried the same stop plus a blanket rejection of a lock pid of 1, so the same session was told to run session start again on every turn. Both walks now compare the top process before stopping, matching the shape used in bin/fm-harness.sh. For the lock walk this is safe because fm_harness_process_matches rejects a host's real pid 1. For the nudge, `kill -0` still gates the lock pid, and on a host an unprivileged `kill -0 1` fails, so a lock file that wrongly names pid 1 leaves the hook silent rather than acting on init. Each walk gains one regression case. The lock case drives a deterministic process table whose pid 1 is the harness and asserts a host-shaped pid 1 still finds nothing, so it cannot pass vacuously. The nudge case needs a real PID namespace, because the builtin `kill -0` gate cannot be reached through a fake ps, and it first proves the same fixture nudges with no lock present; it skips explicitly where unprivileged namespaces are unavailable. * no-mistakes(review): assert comm-strength detection from subprocess vantage in drift guard * fix(bin): verify the live harness guard at the strength the guarantee needs The marker-versus-ancestry boundary this branch ships is a strength claim: detect_own hands an args-strength verdict straight back to a retained foreign marker, so a harness is only protected where the ancestry walk reaches it at comm strength. The installed-harness drift guard probed the pane process alone. Under an interpreter shim the pane process IS the shim, whose own script path is args strength, while the native binary that carries comm strength is its child. The guard therefore observed args for Codex, passed, and would have kept passing if a release stopped spawning that native child at all, while real sessions silently regressed to the original bug. fm-harness.sh gains `ancestry-subtree`, which asks the walk from the pane process and every descendant of it, the vantage a tool subprocess actually occupies. The guard now requires comm strength somewhere in that set and requires every vantage to name the same harness. This supersedes the preceding commit's in-guard leaf walk, which reached the same vantage but left the logic inside the test file, where CI could not pin it and nothing else could reuse it. A harness-dependent check needs both halves: `tests/fm-harness-precedence.test.sh` now carries a portable case proving the subtree probe reaches a strength the top-of-session probe cannot, mutation checked twice, once against the pre-change script and once by disabling descendant enumeration. The subtree walk also avoids depending on tty and process-group semantics that differ between Linux and macOS. Verified live: codex-cli 0.152.0 reports [args codex;comm codex] and Claude Code 2.1.257 reports [comm claude]. * no-mistakes(review): narrow drift guard to the upward vantage path * no-mistakes(review): judge only comm-strength vantages in drift guard * no-mistakes(document): drop duplicated rationale in detection precedence evidence * no-mistakes(review): fix pid-1 nudge case vacuity and descent no-arg expansion * no-mistakes(document): drop branch-relative phrasing in detection precedence evidence * no-mistakes(review): guard remaining empty positional expansions in fm-harness * no-mistakes(document): scope cursor marker-ordering claim to the marker layer * no-mistakes(review): Prefer comm-strength leaves in equal-depth descent ties * no-mistakes(document): Document comm-strength descent tie-break --------- * no-mistakes(review): Blind ancestry in stale gemini/rovo marker-precedence tests * no-mistakes(document): Add missing equal-depth-tie test line to precedence evidence transcript * no-mistakes(review): Fix stale/vacuous agy precedence test, add agy to precedence suite and docs * no-mistakes(document): Fix stale kimi.md marker doc missed by ancestry-precedence fix --------- Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
…rkers (kunchenguid#3944) Claude Code's external-imports check (hasClaudeMdExternalIncludesApproved) reads only the canonical git-root project entry in ~/.claude.json, which its own worktree-to-primary-checkout canonicalization means is never the task worktree fm-claude-trust.sh registered. The trust dialog kept working previously only because its check has an ancestor-walk fallback that happens to reach the worktree entry; the external-imports check has no such fallback. Verified by disassembling the installed claude binary and reproducing in an isolated three-way tmux launch: identical flags registered only at the worktree key still showed the external-imports dialog, and registering them at the primary checkout key suppressed both dialogs. fm-claude-trust.sh now registers all three flags on both the worktree entry and the primary-checkout entry in one atomic write, and refuses when the <project> argument is not itself a primary checkout (its own write target would then be wrong). Extends the harness-adapters Claude reference and the trust test suite. Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
…nguid#4355) The marker lifecycle (fm-wake-lib.sh _fm_recovery_marker_ack) leaves state/.watcher-down behind in an acked:* state after a downtime episode is handled. health_snapshot's presence check reported that as an open gap on every later return, so a handled episode kept surfacing as a false GAP forever.
…kunchenguid#4361) * fix(update): rebind fm-procevent-when watches after a self-update A self-update fast-forwards bin/ in place, changing an armed watch's action executable bytes with no tampering involved. The watch's trust binding was hashed at arm time, so the very next fire was refused as not matching the registered binding and the watch died silently. Add fm-procevent-when.sh rebind-all: it re-hashes and republishes the trust binding for every watch whose action executable lives under FM_ROOT, using the same spec/trust validation as an ordinary fire, and leaves any watch whose action lives outside FM_ROOT untouched. Wire it into fm-update.sh right after a successful fast-forward, for both the primary home and any local secondmate home that advances. * no-mistakes(review): Canonicalize FM_ROOT for rebind-all's containment check * no-mistakes(document): Document fm-update.sh's automatic watch rebind and its verification evidence * no-mistakes(lint): fix(tests): double-quote printf scripts to satisfy shellcheck SC2016 * no-mistakes(review): Reload trust binding from disk before firing to reach live pollers * no-mistakes(review): Lock the fire-time trust reload against rebind_one's publish race * no-mistakes(document): Document rebind-all's self-update guarantee and its two review-round test rows --------- Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
…kunchenguid#4424) * fix(pr-merge): treat plan-gated 403 on branch rules as no merge queue (#42) * fix(pr-merge): read a plan-gated 403 on branch rules as no merge queue github_read_queue_method left status=unreadable for every failed rules read, including a 403 whose body is GitHub's own "Upgrade to GitHub Pro or make this repository public" message. A repository whose plan cannot expose branch rules cannot have a merge_queue rule either, so that specific 403 now resolves to status=none instead of unreadable - unblocking the away-merge grant on private repos without GitHub Pro. Any other failure (auth, rate limit, network, 404, unrelated 403) still reads as unreadable. * no-mistakes(document): Update stale away-merge queue-grant comment for plan-gated 403 --------- Co-authored-by: NewAiCoder <claude@theinbtw.com> * no-mistakes(review): Fix misleading away-queue-grant comment in fm-pr-merge and its test * no-mistakes(document): Update architecture.md for plan-gated-403 merge queue exception --------- Co-authored-by: NewAiCoder <claude@theinbtw.com>
…unchenguid#4246) * fix(tests): select readers of a changed top-level test fixture bin/fm-test-run.sh --changed recognised shared test helpers by an explicit list, tests/lib.sh|tests/*-helpers.sh|tests/fixtures.sh. A top-level tests/*-fixture.sh matched none of those, fell through to the tests/* catch-all, and was marked unmapped, so selection aborted with "no changed-test mapping for source path" and the run selected nothing at all. tests/herdr-client-pair-fixture.sh and tests/remote-herdr-fixture.sh are real shared fixtures with real consumers, so any branch touching one of them left a validation pipeline driving --changed with a hard abort rather than a narrowed selection. Extend the helper arm to tests/*-fixture.sh rather than routing it through the tests/fixtures/*/* arm. Both arms resolve consumers with the same reference scan, and that scan is what selects the right suites here: it finds exactly the tests that read the fixture. The fixtures/ arm adds only a directory-keying step, which has nothing to key on for a top-level file, so the helper arm is the same behaviour with no extra machinery. A tests/ path nothing reads still reaches the catch-all and still refuses loudly. Refs kunchenguid#4100 * no-mistakes(test): order nested fixtures arm before top-level fixture glob * no-mistakes(document): document tests/ shared-file mapping contract and arm order * no-mistakes(review): drop vacuous test phase, correct header claim, restore comment
… asked, not declined (kunchenguid#4387) * fix(bin): read Claude Code's default external-imports flags as never asked, not declined (kunchenguid#4378) fm-claude-trust.sh refused the whole trust registration whenever the project-root entry carried hasClaudeMdExternalIncludesApproved === false, on the premise that Claude Code writes that value only on an explicit "No, disable". Claude Code's default project entry carries Approved and WarningShown both false before the dialog is ever shown, so every such project refused every spawn. Only Approved === false with WarningShown === true — the pair the dialog writes on a decline — now counts as a decline. false/false behaves like an absent flag: trust is registered and no import consent is manufactured. New case test_project_root_entry_default_import_flags_are_not_a_decline fails on b182d0f with the refusal and passes with the fix; tests/fm-claude-trust.test.sh 31/31, bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 and actionlint 1.7.12. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * no-mistakes(review): Correct harness doc's external-imports decline predicate --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…chenguid#4445) * fix(brief): keep operator address out of composed intent Teach raw-word authoring for intent sections and mid-task relays, with a neutral [captain] provenance marker for legacy mixed tasks. Keep headings and contract prose outside the serialized intent body. The legacy selector already excluded the old speaker labels from its output; preserve that read compatibility. The reproduced leak comes from adding labels inside a modern intent body, not from the legacy selector. Do not scrub actual request content. Add exact serialized-input and generated-contract regressions, retaining refusal of unmarked legacy tasks and coverage of scout promotion. Fixes kunchenguid#3882 * no-mistakes(review): Refuse operator-address lines in Captain's intent body * no-mistakes(document): Document operator-address refusal in intent contract comments
…as a proven empty composer (kunchenguid#4455) * fix(composer): accept Grok title overhang * no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage
…ailure (kunchenguid#4474) * fix(bin): recover Claude auto-arm after timeout * no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list
* fix(spawn): establish Claude task channel authority * no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference
…or pending text (kunchenguid#4458) * fix: guard relaunch exit against pending input * no-mistakes(review): Verifying test run in progress * no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard * no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now
…unchenguid#4460) * fix: reconcile diverged secondmate updates * no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md * no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md
…d#4497) * fix(dispatch): support Codex Luna max effort * no-mistakes(review): use portable CODEX_HOME path in codex effort reference
…ardown (kunchenguid#5997) * WIP: retire task-keyed watcher markers and orphan journals at teardown Re-applies old PR kunchenguid#5584 on current main: teardown retires the turn-ended .seen-* signature and an orphaned Herdr presentation journal whose workspace is already gone, and the wake-drain rotates its own dead scratch files. Not yet validated through no-mistakes. * no-mistakes(ci): Fixed the Greptile P1 finding in bin/backends/herdr.sh. fm_backend_herdr_projection_token_workspace_gone used `! ... jq -e ... 2>&1`, which swallowed a jq runtime error (thrown when a non-object workspace entry, e.g. a number before a live token-bearing workspace, hits `.label`) and flipped it to a "gone" verdict, causing teardown to delete a still-live v1 presentation journal. Invariant: a workspace-query error/ambiguity must never be read as token absence; only a cleanly-parsed list with no token-bearing label is "gone". Replaced the body with a single jq verdict (unknown/present/gone): a non-array list or any non-object/non-string-label entry yields "unknown", jq errors/empty output fall through `|| return 1` to unknown, and only "gone" returns 0. Sibling fm_backend_herdr_projection_endpoint_matches_journal already fails safe on jq error (empty match -> journal kept), so it needed no change, matching the author's scoping. Added test_teardown_retains_v1_journal_when_workspace_query_ambiguous driving real teardown with a malformed workspace-list entry, proving the journal is kept and no workspace close occurs. Verified the old logic returns GONE on that input (test fails before, passes after); full tests/fm-teardown.test.sh suite passes (exit 0) and shellcheck is clean. Marker-naming finding left untouched per explicit out-of-scope instruction
…unchenguid#6002) * fix(tests): disable Claude Code's auto-updater during live harness runs fm_live_gate let a live run proceed without ever setting DISABLE_AUTOUPDATER, so a live Claude test could let the real updater repoint ~/.local/bin/claude into a temporary directory and stop every Claude process on the machine from starting. Export DISABLE_AUTOUPDATER=1 on every path where the gate lets a live run proceed, and assert the export in tests/fm-live-gate.test.sh, including that it reaches a child process the same way a real harness pane would inherit it. * no-mistakes(ci): Greptile flagged that the PR's DISABLE_AUTOUPDATER inheritance test only checked a `bash -c` direct child, not the fm-spawn.sh launch path. The user chose to fix it with a regression on that path. In tests/fm-live-gate.test.sh I replaced the generic child test with test_disable_autoupdater_reaches_the_claude_pane_on_the_fm_spawn_launch_path: it drives the real fm-spawn claude launch through the spawn fixtures, captures the exact staged launch command, and runs it as a synthetic pane whose only `claude` is a stub recording the inherited DISABLE_AUTOUPDATER, asserting it saw 1. Switched the file to source fixtures.sh (pulls in lib.sh, guarded) for the spawn helpers. Verified it is a real guard: the stub records `1` when the ambient var is set and `unset` when absent, so it fails if fm-spawn ever scrubbed the variable (e.g. env -i or -u). This confirms fm-spawn's launch construction never references the name and passes it through via ordinary ambient inheritance with no allowlist. Full suite passes (12 tests ok), shellcheck clean. Note for the outer executor: I embedded the daemon caveat as a code comment in the test, but the finding also asks the PR body to state that a backend daemon already running before the gate exported the variable does not inherit it and fully covering that would need launcher support - that forge-side PR-body sentence is outside this CI phase's scope * no-mistakes(ci): Fixed Greptile finding ci-1. Root cause: fm-spawn.sh handed its launch command to an already-running backend daemon that never inherited the test process's exported DISABLE_AUTOUPDATER, so ambient inheritance dropped it and Claude's auto-updater could still run. Fix (bin/fm-spawn.sh): when DISABLE_AUTOUPDATER is set in the spawn's own environment, embed `export DISABLE_AUTOUPDATER=<value>;` into the LAUNCH command text (same idiom as the adjacent COMPACT_ADVISER_DISABLE export), so it survives a daemon-built pane, the env -i allowlist path, and relaunch alike; gated on presence so ordinary spawns are unchanged. Added regression test test_disable_autoupdater_survives_a_daemon_pane_that_never_inherited_it in tests/fm-live-gate.test.sh: stages a real claude launch with DISABLE_AUTOUPDATER set in the spawn env, then runs that exact command in a synthetic pane with `env -u DISABLE_AUTOUPDATER` and asserts the claude stub still recorded autoupdater=1. Verified the test fails (autoupdater=unset) without the fix and passes with it; the round-1 ambient test stays green either way. Full suite passes (13 ok); test file and isolated snippet shellcheck-clean (full fm-spawn.sh shellcheck kept getting terminated by the memory-constrained host, not by findings). Forge-side note for the outer executor: the PR-body caveat that fully covering the daemon case would need launcher support no longer applies to the Claude launch path and should be corrected
…cevent record (kunchenguid#6010) * fix(bin): ring the inbox doorbell only for a newly published procevent result publish_result rewrote a worker's captured Lavish round idempotently on every reconcile, unconditionally moved an already-acknowledged inbox record back out of handled/, and rang the doorbell every time - so an already-processed round rang the owning worker on every cycle. Snapshot the existing active and handled records before the idempotent write and ring, or move anything, only when the write actually created a fresh record; re-delivery of a still-open round is left to the inbox's own re-ring ladder. * no-mistakes(document): docs: reflect worker-board doorbell rings only on fresh inbox record * no-mistakes(ci): Fixed Greptile finding ci-1 in tests/fm-procevent.test.sh (test-only change). The redelivery regression previously moved the delivered note into handled/ before any repeated reconciles, so it only proved an acknowledged note stays quiet and would still pass if an unchanged active note rang every cycle. Per the user's instruction, I inserted (before the mv into handled/) five repeated `pe reconcile` runs with the note still in the active inbox and asserted the ring log holds exactly one line and 001.msg remains active; the existing acknowledged-note assertion after the move is kept unchanged. No product code changed. bash -n confirms syntax is valid; the block mirrors the already-passing post-move reconcile/ring-count assertion directly below it
…unchenguid#6032) * fix(bin): make the Claude Stop auto-arm refuse arguments before arming A model running bin/fm-claude-stop-autoarm.sh --help mid-turn armed a real supervision-host park owned by its short-lived tool process, leaving supervision down once that process exited. The Stop hook passes no arguments, so -h/--help now prints usage and any other argument is refused before anything is sourced, read, or armed. * no-mistakes(document): Clarify Claude Stop hook documentation for manual invocations * no-mistakes(ci): Updated the argument-run regression test to compare checksums of state files as well as entry names. The Stop auto-arm test suite passes, and git diff --check is clean * docs: restore the bin/ toolbelt intro's manual-use clause The document step dropped "interactive entrypoints work by hand too" from docs/scripts.md, which still holds for most bin/ scripts. * no-mistakes(document): Clarify Claude auto-arm manual-use guidance
…uid#6039) * feat(calm): show supervision sailboat and anchor notes on Claude Code The Calm mod follows a bounded display tail copy of the outcome store, which bin/fm-branch-outcome.sh append now refreshes, and the supervision host's latch, and appends one dim transcript line per visible routine outcome, captain outcome, and latch change, replaying unread and unprocessed outcomes at session start. It shows them whenever the mod is active, regardless of config/calm, and never marks anything read. * fix(calm): show each supervision note once per session on Claude Code Claude Code 2.1.283 stores ui.log lines in the session and restores them on --continue, so the mod records how far each session has followed the outcome store and a resume replays only newer outcomes. It also checks file existence before reads so absent files do not log debug errors. The live guard gains the supervision-notes scenario and the dated 2.1.283 record documents the observed behavior. * docs: name the Claude supervision note row as the engine draws it * no-mistakes(review): Seed outcome tail on present and anchor first tail on markers * no-mistakes(review): Seed outcome tail at session start; replay against start markers * no-mistakes(review): Bound outcome tail by bytes; reread recently changed files * no-mistakes(review): Skip store validation when outcome tail already exists * no-mistakes(document): Clarify bounded Claude supervision note replay * no-mistakes(ci): Fixed seed-tail to validate only a bounded suffix of complete store rows and write it through the existing byte- and row-limited tail writer. Added a regression test with malformed history outside that window and updated the script header. Outcome tests and shellcheck passed; the full session-start suite timed out after 240 seconds
…enguid#6033) * fix(bin): read a quiet-mode record as a present captain, never hold-for-return Daemon-backed quiet mode writes the away-posture record marked mode: quiet, but the entry announcement, read-back, and session-start digest rendered it as "hold-for-return only", and the spend cap and PR merge gate treated it as away. A present captain's requested actions could then be held for a return that was not coming. bin/fm-afk-contract.sh now owns which posture a record is (the mode subcommand, fm_afk_contract_mode, fm_afk_contract_away_present). A quiet record announces, reads back, and appears in the digest as a present captain holding nothing; merges under it stay attended and it binds no spend cap. An away record is unchanged, an /afk entry over quiet mode rewrites the record as away, and a quiet entry never turns a standing away record quiet. * no-mistakes(document): Clarify quiet-mode authority and remove stale away guidance * no-mistakes(ci): The CI failure came from a race in the supervision-host test: its restart fixture could observe a watcher left by the preceding cycle. The test now retires that watcher and waits for the fixture arm to report its own started cycle. The focused test passed three times; the full suite was attempted but stopped at a separate intermittent test failure * no-mistakes(ci): Fixed daemon refresh mode selection so an unset-mode refresh follows the posture record: /afk over a running quiet daemon now changes state/.afk to away, while a plain quiet refresh stays quiet. Added script-level regression coverage for start and start-native and corrected a quiet-refresh fixture. The launch test suite, syntax checks, and diff check passed * no-mistakes(ci): Herdr was blocked before tests ran by a GitHub HTTP 500 downloading pinned Treehouse; no code change was warranted for that check. Fixed the Lint 1 ShellCheck warning in tests/fm-afk-launch.test.sh by annotating the intentional background PID capture. The focused test suite, ShellCheck, syntax check, and diff check passed
…merge (kunchenguid#6053) * fix(bin): accept a task's next PR once fm-pr-merge confirms the bound one merged require_recorded_pr_identity now checks fm_pr_poll_merge_already_notified for the recorded pr= before refusing a different URL, so a task's later PR is accepted once its earlier PR's merge is confirmed, while it keeps refusing while the bound PR is still unmerged. * no-mistakes(document): docs(fm-pr-merge): note next-PR accepted after bound PR merges
…6064) * fix(bin): read a live quiet record as a present captain at the host and watcher A quiet record left without its daemon (a quiet start that never ran or was interrupted) was read as away by the supervision host, so it parked a present captain's main and held captain outcomes for a return that never comes, and the watcher and daemon silenced captain-held rechecks on record presence. The host's posture checks, the watcher's and daemon's captain-held silencing, and the host's outcome path (branch report, drain BRANCH OUTCOMES, relocated branch authority, the owners' away wake note, and the Codex checkpoint bound) now ask the record owner's away-or-quiet reading, so only an away record is away. A live away record keeps today's behavior. * no-mistakes(document): Correct quiet-record documentation and supervision guidance * no-mistakes(document): Clarify quiet-record posture and captain-held rechecks * no-mistakes(document): Clarify quiet-record posture in documentation
…kunchenguid#6043) * fix(bin): name an in-window engine latch in the return brief and drop the false handling GAP line The away return brief said nothing had failed after the supervision host latched on engine errors during the window, and printed a GAP: watcher downtime line whenever a wake was merely being handled or queued at return. The failures section now reads the host ledger and latch record and names the latch time, the window's engine-error count, and whether the session is still paused or recovered. An open recovery episode is reported as information, and as a gap only when a queued episode outlived the return grace or the marker cannot be read. * no-mistakes(review): Fix latch trip time, drop marker-age grace, bound error count * no-mistakes(review): Report paused latch without ledger trip row; bound errors * no-mistakes(review): Never report a failed probe's latch row as trip time * no-mistakes(review): Only a retained trip row marks a pre-window latch * no-mistakes(document): Clarify return-brief latch and watcher-gap documentation * no-mistakes(ci): Fixed Lint 1 by marking the shared cooldown constant as used by sourcing scripts. The repository lint command and diff check pass; the return test run was stopped by a 180-second timeout after its completed cases passed * no-mistakes(ci): Fixed the return brief so the trip time and error count come from the same initial latch row, and ledger rows before the current session’s lock boundary cannot affect its latch report. Added real-script regressions for both findings. The return test suite, repository lint, and diff check pass * no-mistakes(ci): Fixed the return brief’s restart cutoff so it retains in-window failures, prints one line per initial-trip row, and omits zero-error count wording. Added real-script restart regressions. The return test suite, ShellCheck, and diff check pass * no-mistakes(ci): Fixed the return brief so a recorded trip followed by recovery stays recovered, while a later pause with a lost trip append gets a separate “trip time unavailable” line. Added a real-script regression that failed before the fix. The return test suite, ShellCheck, syntax checks, and diff check pass * no-mistakes(ci): Fixed the false second latch during recovery. A real-script regression failed before the fix and passes now; the lost-second-trip test still passes. The return test suite, ShellCheck, syntax checks, and diff check pass
* fix(calm): name the Claude Code Calm plugin fm so supervision notes read "fm: " Claude Code labels every mod transcript line with the plugin name, so the notes rendered as "firstmate-calm: ⚓ ...". Rename the plugin to fm, update the live guard to assert the fm: label, and document the one-time replay for sessions resumed across the rename. * no-mistakes(document): Clarify Calm plugin rename in documentation
…#6037) * feat(bin): add fm-live-lab.sh, a one-command live supervision lab builder * fix(bin): exact lab windows, per-lab task ids, self-safe teardown * fix(bin): target lab windows by id, stop lab descendants, add readiness tests * fix(bin): keep Claude's auto-updater off in live labs; list fm-live-lab.sh * fix(bin): start the lab tmux server without user config * no-mistakes(review): Scope lab teardown to its store, root, and task ids * no-mistakes(review): Record selected user stores at up for check and down * no-mistakes(document): Clarify live lab documentation and remove stale narratives * no-mistakes(ci): Fixed the CI failure by checking for an existing lab root before looking up the harness executable. The affected behavioral test and shell syntax check pass; the refusal also works with Claude absent from PATH * no-mistakes(ci): Fixed all four Greptile findings: teardown signals only recorded lab processes and their descendants; the worker gate is in its granted task directory and its path is exposed; readiness uses current crew state; and mate and worker IDs use 12 nonce hex digits. The CLI behavior tests pass, as do shell syntax, ShellCheck, and diff checks. The Claude no-host path is unchanged * no-mistakes(ci): Fixed the CI test’s dependence on an installed Claude binary by supplying a test-local stub. The full fm-live-lab test, shell syntax check, and diff check pass * no-mistakes(ci): Fixed all three selected findings in bin/fm-live-lab.sh: down waits for recorded processes and escalates before cleanup, PID roots are checked against recorded start times, and Claude primary trust is rechecked after mate/worker readiness. Added behavioral tests in tests/fm-live-lab.test.sh. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass * no-mistakes(ci): Fixed the pre-primary settle wait, worker gate instructions, unused retry variable, and teardown PID revalidation in bin/fm-live-lab.sh. Added behavioral tests in tests/fm-live-lab.test.sh. Both requested commands pass: tests/fm-live-lab.test.sh and bin/fm-lint.sh * no-mistakes(ci): Fixed teardown to track pre-kill lab processes by PID and start time, including children orphaned when a root exits. Up now rejects an empty pane PID before calling ps. Added regression tests and a Linux-safe worker fixture. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass * no-mistakes(ci): Fixed teardown tracking for children spawned during shutdown and made the worker fixture verify its exact window with a Linux-available shell. Both requested checks pass. The lab test takes about 66 seconds locally, so the under-one-minute target remains unmet * no-mistakes(ci): Fixed ci-2 and ci-4 in bin/fm-live-lab.sh and tests/fm-live-lab.test.sh. Teardown now tracks identity-checked members of captured lab process groups, including children orphaned during shutdown, without signaling the caller’s group or unrelated processes. Lint passed, and the lab test passed four times * no-mistakes(ci): Fixed teardown so an observed-empty process group is permanently dropped, preventing a reused group ID from signalling unrelated work. Added a ps-shim regression test. The lab test, lint, and diff checks pass * no-mistakes(ci): Fixed ci-1 in bin/fm-live-lab.sh and tests/fm-live-lab.test.sh. The TERM-born-child fixture now waits until its handler is installed before calling down. Down sends SIGKILL to identity-valid survivors on every pass from pass 20 onward and includes survivor process details if it must refuse cleanup. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass locally; Linux CI remains to be verified * no-mistakes(ci): Fixed down’s teardown wait to require two empty identity-checked scans separated by 0.5 seconds, and removed the unused test loop variable without changing the TERM-born-child test. The lab test, lint, and diff check pass locally
…unchenguid#6103) * fix(bin): keep slow watcher cycles and preempted reply polls from breaking supervision - fm_pending_reply_tick selects the records it has work for in one awk pass, so settled records cost no lock or fork and the walk no longer grows with the never-pruned store. - An attached arm keeps following a live, identity-matched holder whose beacon went stale until the lock changes or the shared stall bound (fm_watcher_stall_bound), then fails with a typed stalled-holder line so the retry replaces the holder. - The remote-reply adapter reports the job worker's preemption (exit 76) as a closed window, so the listener keeps its claim and polls again instead of being relaunched every watcher cycle. * no-mistakes(document): Clarify watcher grace and attached-arm documentation
…geable is UNKNOWN (kunchenguid#6110) * fix(bin): retry a bounded number of times when GitHub mergeable is UNKNOWN Fixes kunchenguid#6020 bin/fm-pr-merge.sh refused a GitHub merge whenever the pull request's mergeable field was not literally MERGEABLE. GitHub reports UNKNOWN for a short while after a push or a base-branch change while it recomputes mergeability, so a green, conflict-free pull request was refused as if it could not be merged. github_verify_mergeable now returns a distinct status when mergeable is the only failing condition and reads UNKNOWN. The caller retries up to 5 times, 3 seconds apart (overridable in tests), re-reading and re-checking every live condition on each attempt. Once the bound is spent it reports mergeability as still being computed rather than unmergeable, with the same nonzero exit as before. Every other refusal (closed, draft, conflicting, red or missing checks, away authority, queue protection) is unchanged and never retried. * no-mistakes(ci): I fixed both review findings the way you asked. The full suite (`bash tests/fm-pr-merge.test.sh`) ran to completion. Its last lines showed all `ok`, and any failure would have stopped the run early. I watched the output through `tail`, so I didn't see the new test's own `ok` line directly. **ci-2 (`bin/fm-pr-merge.sh`), retry delay not validated.** What must hold: the retry wait is always a short, valid `sleep` argument, so a bad `FM_PR_GITHUB_MERGEABLE_RETRY_DELAY` can never trip `set -e` or hold the task lock for a long time. The retry loop is the only place that reads this variable. The script now reads the value once before the loop and accepts only whole numbers from 0 to 10. Anything else (empty, `abc`, `-1`, `1.5`, `11`, a huge number, leading spaces) falls back to 3. I ran those values through the check by hand and each came out as expected. The loop now sleeps on that checked value. **ci-1 (`tests/fm-pr-merge.test.sh`), no test for a check changing between UNKNOWN reads.** What must hold: every retry re-checks all live conditions, not just mergeable. The fake `gh pr view` in the test can now take an optional second word on each line of the mergeable sequence, which sets the first check's result. The new test `test_github_mergeable_unknown_retry_rechecks_checks` feeds `UNKNOWN`, then `UNKNOWN FAILURE`. It asserts: - exit code 1 after exactly 2 reads, - the refusal names `check 'ci' is not green`, - the message does not say mergeability is still being computed, - `pr merge` was never called. If a later change made the retry look only at mergeable, the loop would read UNKNOWN 5 times, end with the "still being computed" message, and this test would fail. I didn't run it against a deliberately broken script to confirm that. `bash -n` passes. `shellcheck` reports only the existing info-level notes about files it can't follow. Only `bin/fm-pr-merge.sh` and `tests/fm-pr-merge.test.sh` changed
kunchenguid#6112) * fix(bin): converge every open owner onto a known terminal contribution settle_final only cleared a stale error on retry, so an owner whose saved row still said open kept projecting a merged or closed pull request as open after another owner's row had already recorded the terminal observation. Copy the known terminal observation to every owner whose saved row is not itself terminal, keeping that owner's own pending and notified state, and clear its error. * no-mistakes(review): Carry terminal checked_at when converging existing owner rows * no-mistakes(ci): I fixed Greptile finding ci-2 as you asked, with a change to tests/fm-contributions.test.sh only. The rule it enforces: when a retry converges an owner onto a URL that is already merged or closed, that owner gets the terminal owner's whole observation, not just its state. The same weak check appeared twice in test_interrupted_multi_owner_poll_settles_every_owner, so I fixed both: - **Open owner (line 784):** the check now also requires `.observation == $terminal[0].records[0].observation`. The existing checks for error, checked_at, pending and notified are unchanged. - **Errored owner (just below):** it only checked state and error before. It now reads the terminal owner's file and makes the same full-observation comparison. Adding the comparison alone would not have caught anything. The test fixtures gave both owners identical observations apart from `state`, so copying only the state would still have passed. In both cases I also set the terminal owner's observation head to HEAD_B, so the two observations now really differ. Verification: - The focused test passes against the current bin/fm-contributions.sh. - I temporarily changed `settle_final` so it copied only the state. The test then failed, reporting the owner still on the old head (HEAD_A). I restored the file afterwards, and `git status` shows only the test file modified. - The full tests/fm-contributions.test.sh suite exits 0. No product code changed. The other CI finding (ci-1, "Behavior portable serial 9") was left alone because you chose to ignore it
…nguid#6124) * feat: run the supervision host by default on a Claude primary An absent config/supervision-host on a Claude primary now reads as on with the default engine, and a file holding `off` opts any home out. Cursor, OpenCode, omp, Grok, and Codex stay file-gated, with `off` read as disabled there too. Every reader asks fm_supervision_host_enabled instead of testing the file, and non-bash readers query it through the lib's `enabled` entry. A primary's `off` is not inherited by secondmates: each home keeps its own supervision posture. * test: pin the watcher-path posture in fixtures that assume no supervision host Fixtures that drive the watcher arm or assert a non-host drain now write an explicit off file, and fixtures that copy the Stop auto-arm or the supervision instructions carry the engine lib they now source. The two drain suites also stop reading the code root's config. * fix: name the opt-out when an off home passes an attended wake to main A host parked when the home writes off now logs that the home does not run the supervision host, rather than claiming it has no engine. * no-mistakes(document): Clarify Claude supervision defaults and historical evidence * no-mistakes(ci): Fixed process leaks in the two added host tests. Each case now stops its recorded watcher and host/arm processes; fake hook sessions exit through session.stop. The full host suite passed before the final cleanup refinement, and both affected cases, bash syntax, ShellCheck, and diff checks passed afterward. CI runtime still needs confirmation
…start scope check (kunchenguid#6125) * fix(bin): create the state dir on a fresh primary before the session-start scope check fm_primary_scope_matches required an already-existing state directory, so bin/fm-sessionstart-run.sh stood down on a fresh clone before anything could create it. Split out fm_primary_root_matches so the run wrapper can confirm primary-home identity first, create the gitignored state dir when it is missing, and only then run the unchanged scope check. * no-mistakes(document): Document session-start state dir creation on fresh clones * no-mistakes(ci): I fixed the Greptile P1 the way you asked. When a fresh primary can't create `state/`, the run wrapper no longer stands down silently. **Invariant:** when an otherwise eligible fresh primary cannot create `state/`, startup must never fail silently. This path has only one site: the mkdir in `bin/fm-sessionstart-run.sh`. Other hooks and the nudge wrapper never create `state/`, so they have no equivalent failure. **What changed:** - **Run wrapper** (`bin/fm-sessionstart-run.sh`): it captures mkdir's error and prints one line to stderr before standing down as before (exit 0, or 3 for the Pi prerequisite). The line looks like `fm-sessionstart-run: startup could not create the state directory <path>: <reason>`. - **Test** (`tests/fm-sessionstart-nudge.test.sh`): the new case `test_run_reports_a_state_dir_it_cannot_create` uses a fresh primary with no `state/` and a read-only (0500) root. It checks four things: exit 0, no digest on stdout, no state dir created, and exactly one stderr line ending in "Permission denied". It fails without the fix and passes with it. - **Docs** (`docs/sessionstart-nudge.md`): I added one sentence describing the stderr line and one describing what the new test proves. **Verification:** I ran `tests/fm-sessionstart-nudge.test.sh`, and every test passes. `bin/fm-lint.sh` on the changed scripts (pinned ShellCheck 0.11.0) and `tests/fm-documentation-audiences.test.sh` also pass. As you asked, the wrapper still stands down with the ineligible-checkout status afterwards. It does not report this as a failed eligible startup, which is what the bot suggested
…ery (kunchenguid#6126) * fix(bin): measure pending-reply grace from turn completion, not delivery Fixes kunchenguid#6057 The pending-reply guard demanded a repost ("REPOST REQUIRED: previous marked request had no correlated parent report") while the second mate's correlated reply was already on its way. fm_pending_reply_send_recovery measured its grace window from delivery instead of from the request turn's completion, so any turn longer than the grace fired the demand the moment the turn ended, before the reply could have landed. The missed-report escalation had the same gap: it fired the instant the recovery turn's completion was observed, with no grace at all. Both now measure grace from the relevant turn's completion (request turn for the recovery repost, recovery turn for the escalation), and both take one fresh, uncached read of the parent status file immediately before firing, accepting a correlated line regardless of its verb. Transport-failure escalations stay immediate, and the one-repost limit is unchanged. * no-mistakes(review): Document grace window as measured from turn completion * no-mistakes(ci): Both Greptile findings were real and caused by this PR, so I fixed them. The full `tests/fm-pending-reply.test.sh` suite passes. **ci-1 (a reply could be overwritten by a repost).** The rule that must hold: a recovery send is recorded only if the record is still unresolved, checked under the same per-correlation lock that resolution uses. The escalation path already did this (`_fm_pending_reply_maybe_escalate_locked` reads fresh and publishes under one lock). The recovery path did not: `fm_pending_reply_send_recovery` did its fresh read through `fm_pending_reply_try_resolve`, which let go of the lock before the send was recorded. A reply landing in that gap could be overwritten, and the repost would go out anyway. Now `send_recovery` takes the lock once and, while holding it, re-checks that the phase is still `awaiting_report`, runs the fresh uncached read, and records the send (sender pid and identity, attempt time, phase `recovery_sending`). It releases the lock before actually sending, so the lock is not held during the send. It uses the same lock helpers the other lock wrappers use. Grace timing, the one-repost limit and the escalation path are unchanged. **ci-2 (the test would pass even without the fix).** In `test_recovery_fresh_status_read_resolves_before_firing`, the reply is still appended to the status file, but the stored file signature is then set to the file's new signature. That stands in for a same-size rewrite that the signature cache cannot see. The test first checks that a normal cached read misses the reply, then that the fresh read before sending catches it. I also added the same check for the fresh read before escalation, which the review said was uncovered. The test now sets its own send hook, so it no longer depends on one left over from an earlier test (that leftover had made failures exit silently). **Checks:** - I removed the fresh-read bypass at each site in turn and reran the suite. With it gone from recovery, the test fails with "recovery must not fire once a correlated reply has landed". With it gone from escalation, it fails with "the fresh pre-escalation read should have resolved the record, got escalated". With both in place, all tests pass. - Shellcheck with `-x` timed out locally. Without `-x` and ignoring SC1091, the only warnings are SC2034 on the existing `maybe_escalate` lock wrapper, which is not part of this change. The new code adds no warnings. Changes are in `bin/fm-pending-reply-lib.sh` and `tests/fm-pending-reply.test.sh`. Nothing is committed yet; a plain commit message such as "fix(bin): record the pending-reply recovery send under the fresh-read lock" fits the instruction * no-mistakes(ci): ci-1 was real and caused by this PR. The same bug was also in the escalation path, so both are fixed. The full tests/fm-pending-reply.test.sh suite passes. The rule that must hold: a recovery repost or an escalation goes out only if the record's phase, read after the fresh-read resolve, is still what it was before. The resolver writes phase=resolved first and only then writes the other resolution fields. If one of those later writes fails, it returns an error even though the record is already resolved. Places this rule applies, both fixed: - Recovery (fm_pending_reply_send_recovery): the fresh-read resolve now runs first, and the phase is re-read right after it, whatever it returned. The send is recorded and made only if the phase is still exactly awaiting_report. This replaces the earlier phase check rather than adding a second one. - Escalation (_fm_pending_reply_maybe_escalate_locked): same bug. After a failed resolve it went on to publish the blocked line and set phase=escalated. One added line after the resolve call returns 1 without publishing if the phase has changed. Test: added test_partial_resolve_write_blocks_firing. It forces a failure on the resolved_epoch write after a correlated reply has landed. It checks that the recovery send hook is never called, that no escalation line is published, and that the phase stays resolved. The forced failure runs in a subshell so it can't affect later tests. Checks: - With the recovery fix reverted, the new test fails with "recovery must not fire after a partial resolve". - With the escalation fix reverted, it fails with "partial resolve should block escalation, got escalated". - With both fixes in, every test passes. - Shellcheck was run with SC1091 excluded and without -x, not through the repo's lint script. The only new message is one SC2329 info on the test's override function; other test overrides in the same file already get that same info, unsuppressed. Changed files: bin/fm-pending-reply-lib.sh and tests/fm-pending-reply.test.sh. Nothing is committed. Suggested plain commit message: "fix(bin): recheck pending-reply phase after the fresh read before sending
…to stderr (kunchenguid#6001) * fix: provider-table lookup never writes a broken-pipe error to stderr Fixes kunchenguid#5956 fm_quota_single_provider_for_harness returned from its while read loop as soon as it found a match, closing the pipe while fm_quota_single_provider_table's printf could still be writing. Where SIGPIPE is ignored, as on GitHub Actions runners, bash then prints "printf: write error: Broken pipe" on the resolver's stderr, which intermittently broke the one-diagnostic-line assertions in tests/fm-dispatch-resolve.test.sh. Read the whole table before answering, the way fm_control_harness_supported already does, so the writer always finishes. Return values and output are unchanged. Reproduced by running tests/fm-dispatch-resolve.test.sh with SIGPIPE ignored on a single pinned core under CPU contention: 30 of 30 runs failed before the fix, 0 of 30 after. Note: reproducing requires setting the trap inside the tested shell because nice(1) resets an inherited SIGPIPE ignore to SIG_DFL. tests/fm-quota-choose.test.sh passes and bin/fm-lint.sh is clean. * no-mistakes(ci): Fixed both Greptile findings the user chose to address. ci-1 (bin/fm-quota-axi-lib.sh:154). Invariant: looking up a harness must always end with status 0 and print the provider, even when the caller runs under `set -e`. The loop body `[ -z "$found" ] && [ "$harness" = "$1" ] && found=$provider` now ends in `|| :`. Every iteration succeeds and the whole table is still read. Only `fm_quota_single_provider_for_harness` loops over the table this way, so this is the one place the fix was needed. One caveat: on bash 5.3 the old code did not actually exit under `set -e`, because the `while` loop is not the function's last command, so the new `set -e` test would have passed before this fix too. The change makes the loop's success explicit, as the user asked. ci-2 (regression coverage). I added three cases to the existing `tests/fm-quota-choose.test.sh`, all calling the public lookup function after sourcing the library: 1. With SIGPIPE ignored (`trap "" PIPE`), it looks up every harness 200 times and checks that nothing reaches stderr. 2. A deterministic version of the race: the table function is wrapped so it writes the first row, pauses 0.2 s, then writes the rest. With SIGPIPE ignored, it checks that looking up `claude` prints `claude` and writes nothing to stderr. The stress loop alone reproduced the bug in only about 1 of 5 local runs, which is why this case exists. 3. A direct call under `set -e` prints `claude`. Verification: - `bash tests/fm-quota-choose.test.sh`: all pass. - Same test against the pre-PR library (fa48367, via `FM_ROOT_OVERRIDE`): fails with `printf: write error: Broken pipe`. The deterministic case failed in one run and the stress loop caught it in another. - `shellcheck` on both files: clean. - `tests/fm-dispatch-resolve.test.sh`: passes
…isioning (kunchenguid#6162) * fix: survive Pi 0.99 rendering and Git 2.55 local-clone races Pi 0.99 puts arguments on the stock tool header and leaves hidden custom messages in the export conversation column. Match that header, and keep Calm's boundary on the visible column. Clone a remote home with --no-local so a prune during Git's loose-object copy cannot fail the seed. * no-mistakes(review): Stop SIGPIPE write errors; cover older Pi export and project clones * no-mistakes(document): Clarify Calm export visibility and tool rendering * no-mistakes(ci): Fixed the dispatch diagnostic to list every provider-less use/default profile in one line and added a multi-profile behavior test. Shortened supervision fixtures using the existing engine-grace and park-clock knobs; removed stray scratch files. Dispatch tests, syntax checks, and three targeted supervision cases passed. CI’s prior supervision duration was 751s; the single permitted local full-suite run timed out at 1200s, so an after-duration is not established. The cancelled serial check had no failure verdict. The outer executor should record the measured before/after duration in the PR body when available * no-mistakes(review): Gate Pi 0.99 call headers by version; drop hidden-row assertion * no-mistakes(review): Test stock call headers under Pi 0.87 and 0.99 stubs * no-mistakes(test): Fix older-Pi queued-row test and verify park-boundary behavior * no-mistakes(document): Clarify Pi Calm export and queued-turn documentation * no-mistakes(ci): Fixed the stock macOS Bash 3.2 parse failure in tests/fm-calm-pi-extension.test.sh; its parse check passes. The watcher CI failure is in unchanged code: the isolated five-minute/66-minute case passes locally, but the CI log omits the drain error needed to establish its cause. No speculative watcher fix was made. The full local watcher suite timed out after 500 seconds
…6169) * Prevent premature Lavish board handoffs * Prove Lavish arm lacks reply acknowledgement * Confirm Lavish replies before arming worker boards * no-mistakes(review): Post Lavish reply only after locked arm eligibility checks * no-mistakes(review): Fail Lavish reply closed on unknown version * no-mistakes(document): Correct Lavish reply documentation and remove stale guidance * no-mistakes(document): Clarify Lavish reply routing and remove duplicate version guidance
…#6154) * feat: inherit the supervision-host opt-out from the primary Move the supervision host's off opt-out out of config/supervision-host into its own presence flag, config/supervision-host-off, and add that flag to the primary-authoritative inherited config set. A primary that opts out now opts every secondmate home out at spawn and convergence, and clearing it converges them back. config/supervision-host stays the home-local engine choice. Shape: config/supervision-host mixed two things, a fleet posture (off) and a per-home engine and model. Only the posture should follow the primary, so it becomes a separate presence flag that rides the existing inherited-config mechanism (FM_INHERITABLE_CONFIG in bin/fm-config-inherit-lib.sh) with no new machinery, while the engine line stays local. The parse stays in its one owner, fm_supervision_host_enabled. There is no migration or compatibility handling for a home that still holds off in config/supervision-host. Primary off, mate on: inherited material is primary-authoritative by design, so a mate cannot keep the host while the primary is opted out, and a mate's own opt-out is removed at the next convergence while the primary has none. Running the host on a mate is the primary's choice for the fleet; no override mechanism is added. Live validation (disposable bin/fm-live-lab.sh lab, Claude primary with a real seeded secondmate, --supervision-host off): - up: every readiness check ok, including "host: none running, as expected" and a live mate session; the spawned mate home held the inherited config/supervision-host-off and the gate read primary OFF, mate OFF. - primary removed its opt-out, then bin/fm-config-push.sh reported "supervision-host-off: pushed - mirrored primary absence" and a config reread sent; the gate read primary ON, mate ON, and the live mate handled the reread. - primary opted out again and pushed: "supervision-host-off: pushed", mate gate OFF. - down stopped every lab process and left no lab process running. Out of scope, follow-up: default-on for the other harnesses, away-daemon retirement, rollout. * no-mistakes(document): Document inherited supervision-host opt-out ownership * no-mistakes(ci): Fixed ci-4: with `--supervision-host off --mate`, lab readiness now requires the inherited flag in the mate home and a disabled mate supervision-host gate. The focused behavior test, shellcheck, and diff checks pass. Left ci-1–ci-3 untouched as directed * no-mistakes(test): Fix mate readiness HOST_OFF initialization in lab up * no-mistakes(ci): Fixed Lint 2 by making the new test’s fixtures source resolvable to ShellCheck; its off/on readiness test and ShellCheck now pass locally. Behavior portable serial 5 failed in the unchanged remote-reply test at generation 7. That test passes locally, and no PR-caused defect was identified, so no remote-reply code was changed
…nguid#6179) * fix(tests): cut the fixed sleeps in supervision-host cycles The serial CI lane keeps brushing its 30-minute cap because fm-supervision-host.test.sh spends ~903s of the job, and per the run-36635306527 case profile the top nine cases are all multi-cycle ones (3-10 park/close/turn cycles each): every close waits out the host's sleep $POLL in await_close plus a watcher sleep $FM_POLL scan cycle, and every engine turn waits out the fixed sleep 1 descendant snapshot. That is ~3s of pure sleep per cycle before any real work. The host poll now accepts positive decimal seconds through a new seconds_or validator (FM_SUPERVISION_HOST_POLL), and the engine turn's snapshot loop takes FM_SUPERVISION_ENGINE_SNAPSHOT_SECONDS, also a positive decimal defaulting to one second - the smallest seam at each wait's single owner. The suite drives them at 0.2 alongside the existing FM_POLL=0.5 and FM_ARM_ATTACH_POLL=0.2 knobs, so the real poll loops still run. The park-boundary case moves onto the injected test clock instead of a real 3s wait, per-case cleanup polls the host pid rather than sleeping a full second, and the proof-by-absence windows (flood re-escalation, successor re-announce, watcher persistence, recovery staying off main) shrink from 2-3s to 1s, which still spans two watcher polls at the test cadence. Every assertion, process lifecycle, and reaping path is unchanged; production defaults stay at one second. Isolated case timings on a contended host, base vs branch: attended-latch 54.3->34.6s, undelivered-dialog 67.7->59.1s, away-latch 46.5->30.5s, held-cadence 47.9->21.6s, unreadable-mirror 39.2->38.5s, park-limit 18.2->12.3s, registration-fallback 14.1->10.0s, first-cycle-status 12.6->8.4s, latch-scope 16.7->16.3s. Full suite: 65/65 pass. fm-lint and shellcheck clean. * no-mistakes(review): Wait for scan lock release before duplicate check * no-mistakes(document): Correct supervision snapshot cadence documentation * fix(tests): keep production poll cadence, probe exits at 0.1s The fractional poll cadences multiplied the cost of each loop body: full process-table scans in the engine turn and process refreshes in await_close ran five times more often, which swamped the thin CI runner and nearly doubled every multi-cycle case (serial 5 was cancelled at its 30-minute limit on run 36635306527's successor). Restore the production cadence and notice arm/engine exits with a cheap kill -0 probe at a tenth of a second between the one-second bodies instead: strictly less dead time than baseline with no added CPU. Also hold each injected-clock park bound well past its case's wall-clock checks so a host that ignored the test clock fails instead of silently passing at a real-time boundary, and restore the shortened proof windows (watcher liveness, recovery-off-main absence, first-cycle stream) to their baseline depth. * no-mistakes(document): Clarify supervision engine snapshot documentation
…henguid#6192) * fix: rebalance portable CI from current duration measurements * no-mistakes(test): Test serial packing boundary and verify endpoint timeout cleanup * no-mistakes(document): Clarify timeout guidance and remove duplicated packing estimates
Upstream's refreshed duration hints put the honest three-lane estimate at about 445 s, above the fork's seven-minute target.
…t cleanup now wait up to 30 seconds for the shared Herdr session lock; fresh creates retain the 5-second flat-fallback bound. ShellCheck and `bash -n` pass. The real-Herdr test could not reach the reported scenario locally: installed Herdr 0.9.0 fails an earlier workspace-order expectation; CI used 0.7.4. Lint 2’s ShellCheck OOM occurred after all roots completed successfully and appears environmental
…y limit from 14 GiB to 20 GiB, updated its sizing rationale and the telemetry test expectation. `bash tests/fm-lint.test.sh`, shell syntax checks, and `git diff --check` passed. The local host could not enforce the bounded memory envelope, so the pinned memory-envelope test was skipped; Linux CI verification remains needed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Update firstmate from upstream kunchenguid/firstmate, because upstream made valuable Claude updates ("lets update firstmate since he made some great claude updates").
Context: this repository is a fork (origin yelenplays/firstmate) of upstream kunchenguid/firstmate. Upstream main has 34 commits the fork does not yet contain, including Claude Code supervision work (supervision host on by default for Claude primaries kunchenguid#6124, Calm supervision notes in Claude Code kunchenguid#6039/kunchenguid#6086, the manual Stop hook arming guard kunchenguid#6032, Claude worker access to task channels kunchenguid#5884) plus supervision latency, pending-reply, fm-pr-merge and CI fixes. The fork's own behavior, including its 79 fork-only commits, must survive the update.
Upstream sync
Upstream port point: eb77f02
This is a real two-parent merge of upstream kunchenguid/firstmate main (merge base f93f0d3); every conflict hunk was decided individually.
Upstream's refreshed duration hints pushed the honest three-lane parallel estimate to about 445 s, so the fork's modeled parallel packing target moves from seven to eight minutes (480000 ms) instead of adding a CI lane.
What Changed
Risk Assessment
🚨 High: The shared operational-inbox grant creates a reachable cross-task disclosure path, and whether that access is permitted needs human authorization.
Testing
Non-live targeted tests exercised Stop auto-arming, Claude Calm behavior, branch supervision, and spawn behavior; the supervision-host and spawn test command chains timed out before completion. No live product scenario or visual artifact was produced:
tmuxwas unavailable, and the live Claude setup would require out-of-worktree configuration writes.tmuxis not on PATH, and live Claude setup would require out-of-worktree trust/config writes. Providetmuxand an authorized disposable Claude config/login isolated to the lab.tmuxis not on PATH, and live Claude setup would require out-of-worktree trust/config writes. Providetmuxand an authorized disposable Claude config/login isolated to the lab.tmuxis not on PATH; the permitted live setup also lacks an authorized isolated Claude config/login. Provide both for a live rerun.tmuxis not on PATH, and no authorized isolated Claude config/login was available for the live task. Provide these for a live rerun.tmuxis unavailable and isolated Claude configuration was not authorized.Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
⏭️ **Rebase** - skipped
.agents/skills/afk/SKILL.md- merge conflict rebasing onto origin/mainbin/fm-classify-lib.sh- merge conflict rebasing onto origin/mainbin/fm-push-transition-lib.sh- merge conflict rebasing onto origin/mainbin/fm-supervise-daemon.sh- merge conflict rebasing onto origin/mainbin/fm-watch.sh- merge conflict rebasing onto origin/maindocs/architecture.md- merge conflict rebasing onto origin/maindocs/configuration.md- merge conflict rebasing onto origin/maindocs/herdr-backend.md- merge conflict rebasing onto origin/maintests/fm-daemon.test.sh- merge conflict rebasing onto origin/mainbin/fm-spawn.sh:2862- Every ship/scout Claude worker receives--add-dirfor the sharedstate/operational-inbox, not just its own launch record. A worker can therefore use Claude’s file tools to read other tasks’ operational envelopes; launch-brief records contain the encoded brief and remain in that directory for up to seven days (bin/fm-operational-input.sh:33-36, 349-363). The other grants here are task-scoped, but this one is not. Repository instructions and source do not establish whether cross-task disclosure is allowed. Please decide whether workers may read other tasks’ records; narrowing the grant may require an authorized task-specific storage or access design.tmuxis not available on PATH, and the available live-lab route would require Claude trust/config writes outside the worktree, which this run is not authorized to make. Providetmuxon PATH and an authorized disposable Claude config/login isolated to the lab, then rerun the live scenarios. Two targeted test command chains also exceeded their 240s and 600s timeouts, so their full runs are unverified.tmuxis not on PATH, and live Claude setup would require out-of-worktree trust/config writes. Providetmuxand an authorized disposable Claude config/login isolated to the lab.tmuxis not on PATH, and live Claude setup would require out-of-worktree trust/config writes. Providetmuxand an authorized disposable Claude config/login isolated to the lab.tmuxis not on PATH; the permitted live setup also lacks an authorized isolated Claude config/login. Provide both for a live rerun.tmuxis not on PATH, and no authorized isolated Claude config/login was available for the live task. Provide these for a live rerun.tmuxis unavailable and isolated Claude configuration was not authorized.bash tests/fm-claude-stop-autoarm.test.shbash tests/fm-calm-claude-mod.test.shbash tests/fm-branch-supervision.test.shbash tests/fm-supervision-host.test.sh(exceeded 600s; incomplete)bash tests/fm-spawn-dispatch-profile.test.sh && bash tests/fm-spawn-orca-worktree.test.sh && bash tests/fm-pending-reply.test.sh && bash tests/fm-pr-merge.test.sh(exceeded 240s; incomplete)git diff --check c9098ab99e6692ec01184e084e1c769504a5055b 0cf78b5100c7bcc510c76d1f38b5499d50a6e9b2Checked availability oftmux,treehouse, andclaude;tmuxwas unavailable.✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.