feat: sync upstream Firstmate runtime and guidance - #24
Merged
Merged
Conversation
…3821) * docs: correct stale tmux/herdr backend maturity claims Herdr now has 21 test files, its own required CI job (tests-herdr) that installs a pinned build and hard-fails on "skip: herdr not found", while tmux has 3 test files and is only required as a dependency of the portable-serial e2e lane. zellij, orca, and cmux still have no CI lane at all. AGENTS.md and docs/herdr-backend.md still called Herdr merely "experimental" alongside those three, misleading every session and reader about actual coverage. Update AGENTS.md's config/backend entry, the opening lines of docs/herdr-backend.md and docs/tmux-backend.md, the runtime-backend section of docs/configuration.md, and the matching claims in docs/architecture.md, CONTRIBUTING.md, and README.md so they agree and distinguish tmux (default), herdr (own required CI lane, largest suite, Windows still spike-only), and zellij/orca/cmux (still experimental, no CI lane). No behavior, selection order, or dispatch logic changes. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GDYsWPEfwTuPNjQ2nGBCcj * no-mistakes(review): docs: fix stale herdr label and CI-lane wording * no-mistakes(document): docs: align tmux adapter label in scripts.md * no-mistakes(review): docs: drop duplicated herdr CI claim from tmux page * no-mistakes(review): docs: drop windows claim, align contributing backend wording * no-mistakes(review): docs: trim duplicated CI claim from herdr opening line * no-mistakes(review): docs: drop unguarded largest-test-suite superlative * no-mistakes(review): docs: restore tmux verified label and README experimental scope --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* fix: delete deterministic no-mistakes test baseline, restore intent-targeted Test PR kunchenguid#3644 pinned commands.test to a fm-test-run.sh --changed walk of the repository's 75-162 tests/*.test.sh scripts. no-mistakes runs commands.test verbatim and unconditionally after every fix round, so that walk multiplied by round count: measured at 32.7 minutes per validation versus 3.6 minutes intent-targeted. Delete the pin and restore the 3.6-minute posture. Add tests/fm-nm-test-contract.test.sh as a regression guard, parsing .no-mistakes.yaml as YAML (ruby's bundled Psych, matching the parser tests/fm-test-run.test.sh already uses for ci.yml) rather than grepping its text, restoring in legal form what PR kunchenguid#823 added and PR kunchenguid#1282 removed. Record the rule in docs/configuration.md's "Gate defaults" section (the authoritative owner CONTRIBUTING.md already points at) and strengthen CONTRIBUTING.md's existing local-Test guidance to state it plainly: never configure commands.test to a deterministic test command, complete or partial. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CSt9JvrQMVc4u3jPyUFCFC * no-mistakes(review): Centralize no-mistakes test policy and narrow guard --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* fix(bin): verify pool-slot ownership before returning a worktree slot Workers were killed when cleanup returned a Treehouse pool slot that a different, live task had already taken. Teardown now proves the slot is genuinely this task's before releasing it: it refuses when another task record claims the same live worktree path, or when the endpoint's working directory contradicts the recorded slot, and that refusal holds under --force. Slot allocation, metadata publication, ownership verification, and slot return are serialized across linked firstmate homes, and forced secondmate cleanup verifies descendant slot ownership before returning any child worktree. Regression coverage drives the scripts with two task records naming one slot path and asserts the live worker survives and its slot is not reset. * no-mistakes(review): Protect slots across cloned Firstmate homes * no-mistakes(test): Gate teardown locking on genuine Treehouse slots * no-mistakes(test): Clarify pooled descendant slot gating * no-mistakes(test): Synchronize watcher re-arm test on process exit * no-mistakes(test): Wait for watcher cleanup before timeout escalation * no-mistakes(document): Document pool-slot ownership safeguards * no-mistakes(ci): Fixed all reported CI issues: normalized bare local Git origins to the same Treehouse project-lock identity as absolute clone origins; resolved ShellCheck SC1091 with explicit conditional sourcing; and taught concurrent Herdr teardown coverage to retry expected Treehouse lock contention. Added behavioral regression coverage for bare/absolute origin lock identity. Verified endpoint-safety tests, watcher tests, full CI lint, and the previously failing Herdr teardown assertion * fix(bin): resolve relative origins from repository root * no-mistakes(ci): Fixed teardown so an exact recorded endpoint may change cwd without falsely vetoing cleanup. Removed cwd-based ownership refusal while preserving cross-home record exclusivity and project locking. Updated behavioral coverage for both foreign slot ownership refusal and moved-cwd teardown success. Endpoint-safety, backend, watcher, checkpoint, and targeted lint checks pass. Real Herdr presentation E2E progressed successfully but exceeded the 600s local timeout
…ndmate, and primary (kunchenguid#3867) * feat: add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary Add omp as a verified harness: anchored process-name detection with a Firstmate-owned FM_OMP_HARNESS launch marker that needs real omp ancestry, the fm-spawn launch template with foreign-marker clearing, the tracked .omp/fm-worker-overlay.yml posture overlay, --auto-approve, --cwd, and pre-launch model validation scoped to providers 'omp models --json' lists. Workers get a state-resident busy-state extension keyed on agent_end without willContinue (omp has no agent_settled). The primary gets two tracked .omp/extensions: a turn-end guard that answers omp's blocking session_stop hook by compelling one continuation per turn, with the pre-tool seatbelts and Run-tier session-start delivery, and a watcher extension ported from the Pi one with fm_watch_arm_omp. Control tables, composer busy footers, omp's status row as a bare-composer boundary, the extension supervision model with an omp-keyed ownership proof, the session-start diagnostic, and the supervision protocol snippet follow. Verified live on omp 18.1.11 with openai-codex/gpt-6-astra: a Herdr scout through spawn, busy state, steer, interrupt, exit, and teardown, and the isolated rpc primary lab through extension auto-discovery, digest delivery, lock identity, watcher arm, successor and wake delivery, and the compelled guard continuation. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh * test: prove the omp guard continuation through a guard spy Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh * fix(spawn): clear the gemini marker at the omp launch boundary Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh * test(omp): force the guard stage by freezing the watcher and clear lint findings Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh * test(omp): reap the live lab by path and record omp's rpc shutdown as a note Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh * test(omp): spawn a real secondmate for the discovery rule and classify the omp surfaces Replace the template-extraction check with a genuine --secondmate launch pinned to the fake tmux backend, assert the worker extension's handler set through the executable rather than its bytes, classify the two new omp surfaces in the documentation inventory, and record the Herdr worker evidence in the runtime-backends verification doc. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh * no-mistakes(review): omp: unverify remote routes, narrow busy regex, drop overlay approval pin * no-mistakes(review): omp: validate config-pinned model, correct remote and marker docs * no-mistakes(review): omp: pin config-model validation with a test, trim overlay * no-mistakes(review): omp: sync guard evidence, drop dead param, map quota family * no-mistakes(review): omp quota: refuse unmapped prefixes, match bare model scopes * no-mistakes(document): docs: cover omp in cd-guard, quota, continuity, tmux * no-mistakes(document): docs: add omp subagent-guard row, fix live test header * no-mistakes(ci): Fixed both failing behavior shards and the Greptile P1 in bin/fm-composer-lib.sh. Root cause of "Behavior portable serial 1" and "Behavior portable parallel 2": the omp busy regex (FM_DELIVERY_OMP_BUSY_REGEX_DEFAULT) and omp status-row furniture regex (FM_COMPOSER_OMP_STATUS_RE_DEFAULT) used the bracket range [⠁-⣿]; BSD grep on macOS accepts it but GNU grep on Linux CI aborts with "Invalid collation character", failing every omp busy/furniture read (3 assertions across fm-omp-harness, fm-tmux-submit-busy, fm-composer-lib). Replaced the range with one shared explicit alternation FM_OMP_SPINNER_FRAMES_RE of omp 18.1.11's unicode-preset spinner frames (status set ⣾⣽⣻⢿⡿⣟⣯⣷ + activity set ⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏, read from the installed binary), the same pattern the Kimi busy regex already uses in CI. For Greptile's finding (the harness-agnostic furniture rule's first alternative matched any 1–4-byte token + ' · ', so wrapped typed input like 'fix · tests' with the cursor on it regressed from pending to unknown; reproduced locally vs base), pinned that alternative to omp's identity cell (π||pi, the icon.omp of each preset in the 18.1.11 binary). Tests: fm-composer-lib.test.sh asserts 'fix · tests' is not furniture, a status-set spinner row is furniture, and the wrapped composer screen reads pending under both locales (CAPS_TMUX cursor 3); fm-omp-harness.test.sh asserts a status-set frame reads busy. New negative cases fail against the pre-fix lib and pass after. Verified: fm-omp-harness, fm-tmux-submit-busy pass via bin/fm-test-run.sh; fm-composer-lib passes all cases except one pre-existing, unrelated local failure (Herdr half-block test uses printf '▀', unsupported by macOS bash 3.2; fails identically on a pristine HEAD export, passes on CI bash 5); shellcheck and bin/fm-lint.sh clean. Caveat: GNU grep is unavailable locally, so the Linux compile was not run directly; the fix uses only constructs already proven on CI's GNU grep (multibyte literal alternations, incl. under LC_ALL=C). Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh, tests/fm-omp-harness.test.sh. No docs needed changes (they describe the rule generically) * no-mistakes(ci): Greptile Review: fixed. The omp status-row furniture regex FM_COMPOSER_OMP_STATUS_RE_DEFAULT in bin/fm-composer-lib.sh still accepted a literal `pi ·` opening, so wrapped composer input beginning with `pi ·` was truncated and misclassified. Read the installed omp 18.1.11 binary: the ascii preset's `icon.omp` is `pi` but its `sep.dot` separator is ` - ` (unicode/nerd use ` · `), so a real ascii status row never contains `pi ·` and that alternative could only ever match typed text. Removal-first fix: dropped `pi` from the identity alternation (now `(π|)`) and updated the comment to record why the ascii preset is excluded. Tests (tests/fm-composer-lib.test.sh): added a negative furniture case for 'pi · e · phi as the three constants' and a wrapped-screen assertion (CAPS_TMUX, cursor 3) that a continuation row opening `pi ·` reads pending in both locales; the new case fails against the unfixed lib and passes after. Verified: composer test with the half-block case skipped passes all 33 cases including the omp matrix; bin/fm-test-run.sh tests/fm-omp-harness.test.sh passes; shellcheck -x clean on both files; bin/fm-lint.sh clean. The full composer test via the runner fails locally only on the pre-existing half-block case (bash 3.2 printf cannot emit ▀; passes on CI bash 5), identical to before this change. Docs unchanged (they describe the rule generically and never mention the ascii identity cell). PR must be raised via no-mistakes: not caused by code. attestation.head_sha is cdddc60 while the PR head is cd51cf4 because the pipeline's ci-phase push moved the head; the outer executor's re-push will re-bind the attestation. No file change for that check. Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…#3843) * Fix Pi shell invocation on native Windows * no-mistakes(document): Document Pi Windows Bash transport * no-mistakes(ci): Captain, staged a narrow fix: register the Pi Windows regression for both extension paths, make Windows mode emulation non-failing, and enforce LF shell checkouts. Mapping and coverage checks pass; CI/Require no-mistakes were approval-gated externally * no-mistakes(review): Cover async Windows branch-outcome Bash invocation * no-mistakes(review): Preserve Cygwin checks and refresh Windows timing * no-mistakes(test): Invoke OpenCode operational-input owner through Bash on Windows * validation-fixture * no-mistakes(document): Document Windows Bash helper invocation * no-mistakes(ci): Fixed PR-caused changed-selection failure by removing the malformed tracked evidence artifact and allowing deleted, unconsumed source paths to retire cleanly while preserving fail-closed behavior for live unmapped paths. Added regression coverage. Verified native-Windows Pi shell-seam test passes and --changed selects the Windows regression --------- Co-authored-by: test <test@example.invalid>
…unchenguid#3870) * fix(bin): recognise squash-merged rebased work as landed at teardown A pipeline rebase can leave the local worktree on pre-rebase commits while GitHub squash-merges the rebased head. The landed-work test then compared those stale commits against a squashed main and refused cleanup of work that had already landed. When the forge reports the recorded PR merged and its merge commit is on the default branch, treat a local branch that only repeats paths from the pipeline push as stale rather than unlanded. If the forge is unreachable, the same coverage check runs against a PR head whose content is already on default. Extra local paths still refuse. * fix(bin): drop unprovable squash-rebase landed-work coverage Path-set coverage treated a diverged local branch as landed whenever it touched the same files as the squash merge. That accepts the reviewer's failing sequence: same path, different content, work discarded. git cherry and merge-tree containment were already too strict on the real rebase-fold case. No remaining check is both safe and permissive enough to recognise a stale pre-rebase copy without also accepting unlanded edits, so that case still refuses. Keep the proofs that hold: a merged PR head that contains local work, or a clean content-in-default tree match. Tests now refuse same-path different content and extra unlanded commits, and still allow a local branch that followed the pipeline rebase. * no-mistakes(review): drop recorded-pr-head fallback and reverted-design leftovers * no-mistakes(review): silence squash-merge stdout corrupting test PR head * no-mistakes(review): make unlanded follow-up commit sole cause of refusal * no-mistakes(document): correct stale squash-rebase fixture comments in teardown tests * no-mistakes(ci): Split the three reported checks: - CI (run 34061098467) and Require no-mistakes (run 34061098460) both concluded `action_required` — approval-gated workflow runs that never executed a step. Not caused by this PR's code; no change can clear them. - Greptile Review was a genuine defect in the new tests: the three new refusal cases (tests/fm-teardown.test.sh) asserted only exit status 1 and a REFUSED line, so a teardown regression that destroyed the worktree, branch, and task record before reporting refusal would still pass. Fix (tests only): added one `assert_refusal_retained_task_state` helper and called it from `test_squash_merged_same_file_different_content_refuses`, `test_squash_merged_rebased_local_with_unlanded_commit_refuses`, and `test_squash_merged_stale_local_refuses_when_forge_unreachable`, each capturing the worktree HEAD before `run_teardown`. It pins that the refusal left the isolated copy on disk, the task branch still checked out at the same unlanded commit, and state/task-x1.meta intact. Verification: the four squash tests pass; a sensitivity probe ran the ALLOW fixture (teardown completes) and pointed the same helper at the outcome — it fires, because a completed teardown detaches/deletes the branch and removes the task record, proving the assertions discriminate. Full tests/fm-teardown.test.sh: 83 passing. bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 + actionlint 1.7.12 (plus an explicit --external-sources pass on the changed file). bin/fm-test-run.sh --check-coverage ok. Caveat: test_herdr_flat_teardown_preflight_refuses_before_changes (mode missing-adapter) fails on this machine. Verified it fails identically on base commit f91a950 via `git archive`, so it is a pre-existing local environment difference untouched by this diff; skipped to run the rest of the suite, not modified --------- Co-authored-by: Morten Gad <mogad@itm8.com>
…nguid#3860) * fix(bin): keep supervision armed for registered custom checks A custom check bound by bin/fm-check-register.sh only ever runs inside the watcher's check sweep, but fm_supervision_status counted in-flight tasks, the relay poll shim, and process-event sources as supervision need, and not registered checks. Tearing down the last task therefore stopped every home-level check silently until the next spawn. Count a state/<id>.check.sh that carries its state/<id>.check-trust binding as supervision need. The relay shim keeps its own trust path and task PR polls carry no such binding and are torn down with their task, so neither arms a home by accident. Presence of the binding is the whole test: the sweep validates the bytes at execution time and wakes firstmate when it rejects one, which is the outcome an idle home needs. Closes kunchenguid#3856 * no-mistakes(review): name registered checks in turn-end block banner and doc invariant * no-mistakes(review): narrow PR poll predicate test to what it proves * no-mistakes(document): point Grok re-arm step at supervision-need owner
…nguid#3883) * fix(bin): resolve the shared Treehouse project lock inside remote secondmate homes Every spawn and teardown inside a remote-seeded secondmate home refused, because the project lock's anchor could not be resolved there. fm_firstmate_root_home walks a home's parent bindings upward to find the anchor the lock lives in, and treated a remote parent binding as an error. A remote-seeded home's parent is on another machine, so that walk can never succeed from there - and neither can the home's own local descendants, whose chain terminates at the same record. Both fail closed on every Treehouse-backed spawn and every pool-slot teardown. A remote parent now terminates the walk at the home holding it, which is the correct anchor: a lock taken on this filesystem is neither held nor observable across that boundary, and that home is already the top of the local tree teardown's collect_local_firstmate_states enumerates, since that walk skips remote registry entries for the same reason. Mutual exclusion is unchanged - every home reachable through local parent links still derives one identical lock file per project, and an unreadable binding, an unsupported route, an unreachable local parent, a cycle, and an over-deep chain all still refuse. Origin-less local-only projects keep resolving through their worktree top. Regression coverage pins the anchor for the main-home layout, a local secondmate, a remote-seeded home, and its local child; drives teardown end-to-end in a remote-seeded home; keeps the cross-home slot-ownership refusal across that boundary; and proves two homes still serialize on the one shared lock file. * no-mistakes(document): Clarify machine-local Treehouse lock ownership
…nchenguid#3872) * fix(bearings): repair the board's listening, card hygiene, and reconcile path Three defects made the fleet board go quiet and then lie about what still needs the captain. Never arm a poll on a session that is not live. `lavish-axi <file>` exits 0 even when it refuses to reopen a session the captain ended from the browser, reporting `status: user-ended` with the same session id, so the build's exit-status check accepted a dead session, printed `already-armed`, and left the board reading "not listening". The build now proves the session is live from a fresh authoritative listing immediately before arming - not from the establish call's status alone, which is already stale by then - reopens once when it finds the session ended, and refuses rather than arming when it stays ended. A reopen also replaces the pre-reopen source generation before reporting success, so a runner on its way out cannot be mistaken for a listener, and a board whose source is registered but unowned gets a replacement started before the build returns. Let a dead generation's ownership actually move. Reclaiming a claim ran its capture-reservation cleanup first, and that cleanup re-verifies the recorded state-root identity, so a claim naming a pid and a process group that were both provably gone could not be cleared: reconcile reported a start while nothing attached, and retire refused with "cannot release source ownership". Reservation records are keyed by claim token and every replacement claims a fresh one, so they are hygiene, not an ownership invariant. Reclamation now additionally requires the owning process group to be absent independently, which keeps a reused pid whose poll child still runs from ever reading as a gone generation. A live owner and a crashed leader whose owned group survives are still never reclaimed. Stop carding decisions whose subject already landed. The build drops a decision card whose work item or PR appears in the payload's own landed rows, and one whose task is no longer an open captain call, naming each drop on stderr. A task whose state cannot be established is kept, because a call wrongly hidden is worse than a card wrongly shown. Add the reconcile choice, and make it structurally incapable of closing a call. Every decision card carries a standard `reconcile` option, injected by the build rather than left to the composer. The board now emits the picked option and any freeform note as separate structured fields instead of fusing them, so a reconcile selection is not expressible as an answer value at all - the defect that let `reconcile - <note>` reach the intake as an ordinary answer. The adapter routes selections from that structured field, creation of a reconcile request is bound to a verified board source rather than the shared keyed-answer intake, and the intake still refuses the reserved value on every channel. Each authorization is bound to the captain-hold generation that produced the card, so an obsolete card cannot close a later call, and both terminal outcomes require a pending request: `reconcile close` records the evidence under its own `reconciled` mode so it never reads as the captain's words, and `reconcile note` leaves the call open. Anything unprovable - an unversioned row, a missing generation, an unreadable state - refuses rather than acting. Regression coverage fails without each fix, and pins every leak path: a bare reconcile, a standalone close or note with no pending request, an any-channel reconcile, an annotated selection from a freeform card, and a generation-skewed authorization. An opt-in guard re-proves the lavish-axi shapes and the reopen against the installed tool. * fix(bin): quote the done comparison in the reconcile intake shellcheck SC1010 reads the bare word as the loop keyword. The failed run never reached its lint step, so this shipped in the recovered content. * no-mistakes(review): Publish reconciled parent resolution before request retirement * no-mistakes(review): Clarify committed cleanup and reconcile reservation scope * no-mistakes(review): Preserve remote cards and legacy answer compatibility * no-mistakes(test): Separate live claim release from stale reclamation * no-mistakes(test): Allow terminal self-retirement during active capture * no-mistakes(document): Document Bearings repair contracts * no-mistakes(ci): Stabilized the failing Herdr presentation E2E by serializing test-harness Treehouse allocator calls, preventing concurrent recovery spawns from claiming the same pool slot while preserving Herdr concurrency coverage. Verified with the full E2E suite on Herdr 0.8.2, bash syntax checks, ShellCheck, and git diff checks
* fix(bin): skip pooled-worktree freshness fetch when no origin is configured An origin-less local-only project has nothing remote to be stale against, so fm-spawn's freshen_spawn_worktree_base refused to launch crews for it. Detect a missing origin remote and skip the fetch freshness gate entirely; an existing-but-unreachable origin keeps refusing as before. * no-mistakes(review): Preserve pool safety for absent and unusable origins * no-mistakes(review): Refuse empty origin configurations during pooled spawn * no-mistakes(review): Detect empty origin sections across config includes * no-mistakes(review): Honor globbed includes when detecting origin configuration * no-mistakes(review): Document conservative conditional include handling * no-mistakes(review): Use Git-resolved config files for origin detection * no-mistakes(review): Document included empty-origin detection boundary * no-mistakes(document): Document originless pooled spawn behavior
…enguid#3889) * feat(tests): run live harness guards by default where the harness is installed The 24 live-harness guards each opened with their own env check, so on the machine that has every harness - the one the product and its validation actually run on - all of them skipped and passed. Fourteen had never been run by the pipeline at all. tests/lib.sh gains fm_live_gate as the single owner of that decision: a guard that spends no model tokens runs wherever its tools are installed, a guard that submits prompts stays opt-in, an absent tool is a named capability skip, and a guard's own variable or FM_LIVE forces it on (turning an absent tool into a failure) or off. Every live guard now opens with it, which also carries the test-suite gate-refusal bypass into the guards that never sourced the shared helpers and were therefore refused whenever a gate agent ran them. bin/fm-test-run.sh records what a skip means: the family's expected class is live-capability rather than a bare env opt-in, and each gate skip's reason is logged and written to the timing artifact, so a lane can say which tool this host could not exercise. Only the token-free guards flip to default-on: composer-matrix, the harness liveness drift guard, and the Herdr version floor. cursor-primary submits three prompts, so it stays opt-in. Running the drift guard unasked immediately found a real defect it existed to catch: it resolved the harness through a generic `command -v cursor`, which on a machine that also has the Cursor editor finds the editor launcher rather than cursor-agent. That binary exits at once, leaving a bare shell in the pane and a liveness-drift failure no classifier change could fix. It now asks fm_cursor_resolve_binary first, the same verified owner fm-spawn uses. CI installs the public Pi package in the portable serial lane and fails on its skip token, so the Pi extension tests stop passing silently against a package that is not there. No secret is added. Verified on macOS 26.5.2 arm64: the drift guard runs with no variable set and classifies 8 installed harnesses alive; the Herdr version-floor guard runs by default and checks 4 real releases; every live guard refuses together under FM_LIVE=0. * fix(tests): keep the composer-matrix guard opt-in Running it unasked is red on a healthy machine for reasons no code change here removes: a harness that has not trusted this checkout sits on its own trust dialog, which the guard treats as an unreadable composer and correctly fails. The opencode 1.18.29 and grok 1.0.13 composer drift it also surfaced reproduces identically on main and is filed as separate work. So this token-free guard stays opt-in with the reason stated in its header, and the coding guidelines record the narrow exception: a guard whose verdict depends on host state that installing its tools does not establish may stay opt-in, because one that is permanently red is one the fleet learns to ignore. The other two token-free guards keep running by default. * no-mistakes(review): Wire bearings guard and remove composer exception policy * no-mistakes(review): Run Pi responsiveness guard by default * no-mistakes(review): Gate AFK Pi Herdr through authoritative family sweep * no-mistakes(review): Sanitize live gate test environments * no-mistakes(document): Document default-on live guard behavior * no-mistakes(ci): Fixed CI by installing the Pi package in portable-parallel-1, where fm-pi-primary-types.test.sh runs, and enforcing its package-missing gate skip there. Verified with fm-lint.sh, workflow actionlint, coverage partition checks, lane membership, and git diff checks * no-mistakes(ci): Fixed CI’s Pi typecheck skip enforcement by giving npm, tsc, and Pi-package capability skips a shared prefix and configuring both relevant CI lanes to fail on that prefix. Verified missing tsc emits the expected skip, missing Pi package becomes a runner failure, and actionlint, ShellCheck, and git diff checks pass
… is set (kunchenguid#3891) * fix(bin): refuse the behavior suite in the repository primary checkout A task worker's isolated worktree placement is verified exactly once, when its task starts, and nothing re-checks it afterwards. A worker that later changes directory into the repository's primary checkout runs its Git commands, and this branch-switching suite, against the one checkout every linked worktree resolves against and every landing merges into. A run that dies mid-suite can leave that checkout on a stray branch. bin/fm-test-run.sh now refuses that case. When FM_TASK_ID marks a task worker and the runner resolves to the primary checkout, every executing mode exits non-zero before selecting a suite, with one line naming the primary path and pointing at the assigned task worktree. The predicate is the one bin/fm-spawn.sh already uses for launch placement: the working tree's own git dir is the repository's common git dir, which separates the primary from every linked worktree even when their top levels differ. A run with no FM_TASK_ID set is unchanged, and so are the inspection modes, which execute nothing. When git resolves neither directory - a non-repository fixture, a detached copy - nothing proves this is the primary, so the run proceeds. bin/fm-spawn.sh sets the marker: ship and scout launches export FM_TASK_ID into the pane shell on the same pre-launch channel as GOTMPDIR, and the name joins the sanitized launch environment allowlist so an isolated launch keeps it. * no-mistakes(review): clear inherited task marker in test lib; name resolved ROOT * no-mistakes(document): docs: record FM_TASK_ID marker and runner placement refusal --------- Co-authored-by: Talon Stark <talonstark@gmail.com>
) * fix(bin): bound a stale alarm with the backlog hold, not only the status line A legitimate wait has two records and the stale alarm reads only one. `status_is_paused_or_captain_held` takes a status line, so it sees a wait the worker declared. It cannot see the wait firstmate records when it hands work to the captain: `bin/fm-captain-hold.sh hold` writes that into the backlog and leaves the status log alone, so a delivered task keeps `done: PR ...` as its last line for the whole time the captain is deciding. Both stale branches were blind to it, and each churned a new pane hash back into its own alarm: a `done:` line is captain-relevant and reaches the terminal-stale branch, while a held task whose last line is `working:` reaches `surface_nonterminal_stale` and fails its declared-wait test. Consult that second record where the watcher is about to alarm, through `bin/fm-captain-hold.sh open`, which already owns the predicate's semantics, and bound the alarm on the shared `.paused-resurfaced-<key>` marker and `PAUSE_RESURFACE_SECS` window the declared-wait absorb already uses. The first sight still alarms, the window's end alarms once more, and a held crew that goes genuinely silent still escalates through the wedge timer. Only an established open captain call bounds anything: an unreadable backlog, an absent or incompatible tasks-axi, a row this home does not carry, and every task with no hold keep alarming exactly as before. The backlog hold is deliberately not recorded as a declared pause, because the loop-top reconciliation and `pause_state_class` both read the status line and would clear a flag that line does not support. Extends the fix in kunchenguid#3443, which closed the forms of this loop that the status line itself can express. * fix(bin): identify the captain call a stale alarm is bounded by Three gaps in the bound added by the previous commit, all in how the throttle is scoped and where the backlog is consulted. The scope carried only the status-log signature. A task can be held, answered with `--release`, and re-held as a genuinely different captain call without any status append, so the second call inherited the first one's marker and its first sight was absorbed - the one thing this bound must never do. The task id is not the call: `bin/fm-captain-hold.sh open` gains `--identity`, which reports the call's own lifecycle - its hold-set stamp and the number of recorded answers - on an exit 0 and only then, leaving the silent predicate every existing caller reads unchanged. The throttle scope now carries that identity. The terminal path recorded the throttle before publishing the durable wake. A failed append exits the watcher with nothing queued, and the next sighting then read that fresh marker and absorbed the retry, turning a delayed alarm into a lost one. Recording moves behind the append, as the non-terminal path already had it, and the comment claiming the marker could not outlive its wake is gone because it was false. The backlog was consulted only on a new terminal pane hash. A captain call can open after a hash was absorbed as provably working, changing neither the pane nor the status log, so nothing re-read the backlog and the wedge timer kept firing possible-wedge alarms through a legitimate wait. That timer now consults the call at its own alarm boundary and takes the same bounded cadence - and only at that boundary, so an ordinary repeat poll under the bound stays the local-only read it was. Regression coverage for each, all driving churn through one watcher process rather than relaunching per pane change: relaunch cost dominated the earlier shape, and an absorbing watcher stays in its poll loop across churn in production anyway. An unheld task still alarms on every new hash, and an elapsed wedge timer with no open captain call still escalates as a possible wedge. * fix(review): Compose stale throttles with captain-call lifecycle identity * fix(review): Preserve bounded same-hash captain-call resurfacing * revert(bin): narrow the captain-hold stale bound to its observed defect Lifts the lifecycle-identity and cadence-ownership work back out, leaving the change at the shape that matches the defect actually observed: the stale alarm did not consult the backlog captain hold, on either stale branch. Reviewing the wider version surfaced a series of adjacent gaps in the watcher's alarm state machine - a call opening after the first alarm, marker invalidation at the hold lifecycle boundary, and which deadline a terminal timer represents. They are real, but fixing them turns a small extension into a state-machine change to the alarm path, which is a different review on a subsystem that is being actively reworked. They are named as known limitations rather than carried here, and none of them is load-bearing for what remains: the bound does strictly less than the reverted version, leaves the wedge path escalating on STALE_ESCALATE_SECS exactly as before, and introduces no silence that the existing terminal-alarm path did not already have. Kept from the reverted work is the record-after-append ordering, because that is a defect in the code being shipped rather than an adjacent one: recording the cadence marker before publishing the durable wake let a failed append lose an alarm outright instead of delaying it. History is preserved: the earlier commits stay on the branch and this removal sits on top of them. * fix(review): Document secondmate captain-hold scope boundary * fix(document): Document captain-hold stale alarm scope * fix(bin): bind the stale throttle to the captain call, not the status log The throttle this change introduces was scoped to the task's status-log signature. Answering a call with `--release` and holding the task again creates a genuinely different captain call without necessarily appending to that log, so the second call inherited the first one's marker and its first sight was absorbed. That is the one alarm this bound must never swallow. A delivery announced twice is noise; a decision waiting on the captain that is never surfaced is invisible, because nobody asks for what they do not know to ask for. Measured rather than assumed, on the same fixture - a delivered task held for the captain, released, and re-held with no status append, driven through bin/fm-watch.sh: base c499f84 call-1 first=ALARM call-1 churn=ALARM new call first sight=ALARM before this fix call-1 first=ALARM call-1 churn=absorbed new call first sight=absorbed after call-1 first=ALARM call-1 churn=absorbed new call first sight=ALARM Base never suppresses the new call, so the suppression came from this change and closing it completes the fix rather than widening it. `bin/fm-captain-hold.sh open` gains `--identity`, printing the call's lifecycle - its hold-set stamp and count of recorded answers - on an exit 0 and only then, so the silent predicate bin/fm-teardown.sh reads is untouched. The throttle scope carries that identity beside the status signature. The sibling case was measured too and is NOT included: on the status-declared path, where the last line is `captain-held:`, base already absorbs a re-held call's first sight. That behaviour predates this change and stays documented as a known limitation rather than repaired here. * fix(document): Document captain-call throttle lifecycle scope * fix(ci): isolate the Herdr restart fixtures from a claimed worktree The Herdr behaviour test intermittently reused a local worktree still claimed by an earlier fixture after a restart. The restart scenarios now use an isolated Treehouse project. The full Herdr test passes on Herdr 0.8.2; bash -n and git diff --check pass as well.
…ranch (kunchenguid#3871) * Let the supervision branch resolve extension-registered providers The isolated branch ModelRuntime cannot see providers an extension registered into main's runtime at run time, so a pin on pi-devin-auth's devin/swe-1-7 (or an unpinned branch following a main session on devin) failed with "unavailable to the isolated branch runtime". Capture main's ModelRegistry alongside mainModel and copy each extension-registered provider config into the branch runtime at model-resolution time. The config carries the provider's own streamSimple and oauth wiring by reference, so the custom gRPC transport reaches the branch unchanged instead of being reimplemented. The /supervision-model picker uses the same copy so those models are offered. Update configuration.md and pi-supervision-branch.md, which previously stated extension-registered providers were not offered. * no-mistakes(document): docs: own devin provider carve-out in branch architecture doc * no-mistakes(ci): Fixed the Greptile P1 finding: the /supervision-model picker copied extension-registered providers into the branch ModelRuntime but checked hasConfiguredAuth without refreshing them, so providers with provisional post-registration auth were omitted from the picker while the pin-resolution path (which did refresh) accepted them. Root-cause fix in .pi/extensions/fm-branch-supervision.ts: moved the `refresh({ providers, allowNetwork: false })` call into `copyExtensionProviders` (now async, refreshing every provider it copied) and removed the duplicate per-provider refresh from `resolveBranchModel`. Both the picker and the resolution path now share one copy-and-refresh step, so hasConfiguredAuth is real in both. Regression coverage in tests/fm-pi-branch-extension.test.sh: the stubbed ModelRuntime now mirrors the real runtime by leaving a registered provider's auth pending until `refresh()` runs for it. With that stub, the existing extension-registered-provider case fails against the pre-fix extension (picker offers only anthropic/main-model) and passes with the fix. Verification: tests/fm-pi-branch-extension.test.sh passes (42 ok, no failures); tests/fm-branch-supervision.test.sh passes; tests/fm-pi-primary-types.test.sh skips locally because tsc is not installed (the refresh signature reused is the one the existing code already called). Intent constraints preserved: isolation flags untouched, carve-out still scoped to provider registration, graceful fallthrough when no providers are registered * ci: retrigger flaky Herdr/serial-1 lanes * no-mistakes(document): docs already cover branch extension-provider copy
…gress (kunchenguid#3943) * fix(watch): detect stalled secondmate queue progress * no-mistakes(review): gate secondmate stall on active turns and progress episodes * no-mistakes(document): align secondmate wake-stall docs with progress-episode detector * no-mistakes(document): clarify active-turn gate in wake-stall config docs * no-mistakes(ci): Fixed the one real defect behind the failing checks. ROOT CAUSE (Greptile P1, real code defect in this PR): `secondmate_wake_stall_tick` in bin/fm-watch.sh reset the no-progress timer only when the oldest actionable queue sequence INCREASED (`[ "$seq" -gt "$observed_seq" ]`). When a secondmate is retired and reprovisioned under the same task ID, its fresh home's queue sequence restarts BELOW the recorded position, so the comparison is false, no reset happens, and the new queue inherits the retired generation's already-expired idle interval — emitting a false `secondmate wake-loop stalled` on its very first observation. That is precisely the false-alarm class the user intent requires this PR to remove. FIX (smallest, removal-first): bin/fm-watch.sh:749 now resets when the drain position MOVES AT ALL (`-ne` instead of `-gt`). Draining moves it up, reprovisioning moves it down; neither is a continued no-progress episode. The asymmetric `-gt` branch is removed rather than special-cased or hardened. Updated the function header comment plus the two doc sentences in docs/architecture.md and docs/configuration.md that stated the old advance-only semantics. REGRESSION TEST: added `test_secondmate_reprovisioned_queue_starts_a_fresh_interval` to tests/fm-wake-queue.test.sh (registered in the invocation list). It drives the real watcher through retired generation (seq 9) -> reprovision (seq 3, later clock) -> freeze, asserting observable wake-queue output, no source-text inspection. VERIFICATION: - Fails before / passes after: with the fix reverted the suite aborts on `not ok - a reprovisioned queue generation inherited the retired generation's idle interval and alerted`; with the fix it passes, and its third leg confirms the restarted generation still escalates on a genuine freeze (row=3 idle=2s), so the fix does not merely mute the alarm. - `bash tests/fm-wake-queue.test.sh`: exit 0, 38/38 pass, all five secondmate cases green. - `bin/fm-lint.sh`: clean (ShellCheck 0.11.0, actionlint 1.7.12). CHECKS NOT CAUSED BY THE CODE: the CI and "Require no-mistakes" runs (34152740610, 34152740604) both ended with conclusion `action_required` — workflow approval pending, not a test/build failure. Separately, `bin/fm-test-run.sh --check-coverage` exits 1 in this environment, but I confirmed by stashing my changes that it fails identically on the unmodified base tree (locale-related `comm: input is not in sorted order`); it is pre-existing and this change adds no new test file for the partition to account for. Changes are left uncommitted in the worktree * no-mistakes(ci): Fixed the one real code defect behind the failing checks. ROOT CAUSE (Greptile P1, second round, on the head commit e1304e6): `secondmate_wake_stall_tick` in bin/fm-watch.sh identified the queue's drain position by the sequence number ALONE. The previous round changed the comparison from `-gt` to `-ne`, which handles a reprovisioned queue that restarts BELOW the recorded position, but not one that restarts ON it. A mate retired and reprovisioned under the same task id gets a fresh home whose wake-queue sequence counter restarts at 1 — and the retained parent progress marker very plausibly holds a low sequence too (a queue frozen on its first row records seq 1). Equal sequence ⇒ no reset ⇒ the brand-new queue inherits the retired generation's long-expired idle interval and emits a false `secondmate wake-loop stalled` on its very first observation. That is exactly the false-alarm class this PR exists to remove. FIX (smallest, removal-first): the file already defines the identity of a queue row once, as `row_key="$epoch-$seq"` (used for stall receipts, the stall marker, and the notify key). The progress marker's separate, weaker seq-only identity is removed: `row_key` is now computed once right after the row is parsed, stored in the progress marker, and compared with `!=`. Across generations the epoch differs (the new generation's rows are appended later), so no sequence collision can carry a stale interval; within a generation the key is stable exactly while the position does not move. bin/fm-wake-lib.sh's `fm_wake_secondmate_progress_marker_write` now takes `<oldest-row-key>` and validates it the same way the two neighbouring row-key writers do. Updated the function header comment and the two doc sentences (docs/architecture.md, docs/configuration.md) that described the old sequence-only semantics. REGRESSION TEST: `test_secondmate_reprovisioned_queue_starts_a_fresh_interval` in tests/fm-wake-queue.test.sh now drives the reported case — the reprovisioned generation restarts on the SAME sequence 9 (epoch 200) that the retired generation recorded (epoch 100), at a later clock — and asserts observable watcher output only. Its third leg still confirms the restarted generation escalates on a genuine freeze (row=9 idle=2s), so the fix does not merely mute the alarm. Three seeded progress markers in the symlink, crash-window and prefix-receipt tests were updated to the epoch-sequence form. VERIFICATION: - Fails before / passes after: with bin/fm-watch.sh and bin/fm-wake-lib.sh reverted to HEAD and the new test in place, the suite aborts on `not ok - a reprovisioned queue generation inherited the retired generation's idle interval and alerted` (exit 1); with the fix, `bash tests/fm-wake-queue.test.sh` exits 0 with 38/38 pass, all six secondmate cases green. - `bin/fm-lint.sh`: clean (ShellCheck 0.11.0, actionlint 1.7.12). - I also started `bin/fm-test-run.sh tests/fm-watch-checkpoint.test.sh tests/fm-watch-triage.test.sh tests/fm-watch-recovery-loop.test.sh` as a blast-radius check; it was still running when this phase had to return, so its result is not included. No other suite references the stall detector or the progress marker (grep over tests/ for `wake-loop stall|SECONDMATE_WAKE_STALL|secondmate-wake-progress` matches only fm-wake-queue.test.sh), and the changed lib function has exactly one caller. CHECKS NOT CAUSED BY THE CODE: the CI and "Require no-mistakes" runs on the head commit (34154091191, 34154091236, 34154091945) all ended with conclusion `action_required` — pending workflow approval, not a test/build failure. `bin/fm-test-run.sh --check-coverage` still exits 1 in this environment for the pre-existing locale reason recorded in the previous phase (`comm: input is not in sorted order` on the unmodified base tree); this change adds no new test file. Changes are left uncommitted in the worktree: bin/fm-watch.sh, bin/fm-wake-lib.sh, docs/architecture.md, docs/configuration.md, tests/fm-wake-queue.test.sh --------- Co-authored-by: Alex William <awilliam@v2202608403614505120.powersrv.de>
…idence (kunchenguid#3952) * fix(tests): make the Calm export-DOM render step retry and report The Calm suite's rendered-export-DOM assertion started breaking CI with a bare "could not render calm-mode HTML export DOM", which read like a Pi 0.85 rendering change. It is not one. Calm's rendered rows are identical across Pi 0.84.4, 0.85.0, and 0.85.1, and the CI break appeared in exactly one of the thirteen most recent runs, all on the same Pi 0.85.1, with the main runs immediately before and after it passing. What actually failed is headless Chrome's start-up. The render step made a single unattended attempt and discarded both Chrome's stderr and its exit status, so the log held nothing to tell a Chrome crash apart from a real change in Pi's export shape. Rendering is a vendor-tool step; the DOM assertions that follow it are what protect the Calm conversation boundary. So the step now retries a bounded number of Chrome start-ups on a fresh profile, drops Chrome's background network and /dev/shm dependencies without changing what a local file renders to, and, when every attempt fails, reports the Chrome binary, its version, the installed Pi version, each attempt's exit status, and Chrome's own stderr. test_export_dom_render_guard pins that with real processes and no browser: one clean render, one that only succeeds after a start-up failure, and one that never renders and must report enough to diagnose itself. The verification record adds the 0.85.1 evidence this contract is now pinned to, the cross-version comparison run through isolated installs, and the Pi 0.85.0 packaging gap - its dist/experimental/server.js statically imports @earendil-works/pi-server, which 0.85.0 does not declare - that made the contract look version-sensitive in the first place. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Y7rr2DHf51MjMvy7sMRauo * no-mistakes(review): docs: attribute Pi 0.85 calm contract adaptation to renderer change * no-mistakes(review): tests: drop inert chrome flags, report render timeouts * no-mistakes(document): docs: fix stale Pi version facts and doc-lint link --------- Co-authored-by: Alex William <awilliam@v2202608403614505120.powersrv.de> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…unchenguid#3950) * fix(bin): make every counted wake queue row presentable or retired A wake row could be counted as queued while no drain would ever present it, leaving the operator told to "drain them before anything else" by a command that printed nothing and offered no acknowledgement. Two independent paths produced that state. A row reserved by a live supervision-branch grant is excluded from a main drain by design, but fm-guard.sh counted the whole queue, so main was warned about rows only the branch could present, on every guarded command for as long as the grant was held. A row that lost its five appended fields or its numeric sequence can never be claimed, presented, or named by an --ack-through cutoff, yet it still counted as queued, wedging the queue permanently. The guard now counts only the rows the calling actor can itself present or retire, and a main drain retires unusable rows under the queue lock, reporting them in bounded escaped form before removal so the evidence survives for the separate row-generating defects. A retirement failure is reported loudly and never suppresses unrelated consumable work. A main drain whose remaining rows are all branch-held says so in one bounded line instead of exiting silently. Grant row-list and owner-record reads move into fm-wake-lib.sh so the drain, the grant publisher, and the guard share one implementation. Ownership is unchanged: a branch drain still touches nothing outside its grant and never retires a row, and main still cannot present or acknowledge an active grant. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GghGvsa4JDB1E5FuznX2i1 * no-mistakes(test): keep SIGTERM-safe arithmetic in wake queue retirement pass * no-mistakes(review): add guard advisory for branch-held wake rows * no-mistakes(document): document per-actor wake counting and unusable-row retirement * no-mistakes(ci): Addressed the Greptile P1 on bin/fm-wake-lib.sh:1840 ("Unreadable queue suppresses alarms"). Root cause: fm_wake_actor_pending_count inferred "the queue could not be counted" only from awk's printed output (`case "$count" in ''|*[!0-9]*) count=1`). That relies on awk aborting before its END rule when the input cannot be opened. An awk that reaches END after a failed open prints `0`, which the fallback accepts as a genuine count; both actor counts then read zero and bin/fm-guard.sh emits neither the queued-wake warning nor the branch-held advisory for a queue nobody proved empty. Fix (bin/fm-wake-lib.sh:1828,1835): both counting awk invocations now set `count=''` on a non-zero awk exit status, so the existing "cannot be counted => report a pending row" fallback is driven by awk's exit status instead of an implementation-defined detail of what it printed. No new code path or behavior; the pre-existing fallback just becomes unconditional. Comment updated to state why. Regression test (tests/fm-wake-queue.test.sh: test_uncountable_queue_still_raises_the_pending_alarm, registered in the run list): runs the real bin/fm-guard.sh against a non-empty, unreadable queue with a PATH-injected awk emulating an END-running implementation (prints 0, exits 2; execs the real awk otherwise) and asserts "queued wakes pending" is still emitted; disconfirming half asserts the same fake awk over a readable, provably empty queue stays silent. Fails on the pre-fix library ("not ok - a queue that could not be counted silenced the queued-wake alarm"), passes after. Verified locally: bin/fm-test-run.sh tests/fm-wake-queue.test.sh -> 0 failed; tests/fm-guard-stale-banner.test.sh + tests/fm-watcher-lock.test.sh -> 0 failed; bin/fm-lint.sh (ShellCheck 0.11.0 + actionlint 1.7.12) clean. Caveat reported honestly: on the awks available/known here (mawk locally, plus gawk and BWK/macOS awk, all of which treat an unopenable input as fatal and skip END) the alarm was not actually suppressed - I reproduced the unreadable-queue case and the warning fired. The change removes the code's dependence on that awk detail rather than repairing an outage observed on this platform. Both intent constraints still hold: every counted row remains presentable or retirable, and no alarm-suppressing path was added --------- Co-authored-by: Alex William <awilliam@v2202608403614505120.powersrv.de> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(bin): use /usr/bin/stat on Darwin to survive GNU stat shadowing * fix(bin): extend /usr/bin/stat prefix to Darwin stat -f sites added on main * test(bin): make fm-stat-shadowing skip visible on non-Darwin and isolate fm-watch state * ci: re-trigger after Chrome headless timeout in calm HTML export test * ci: re-trigger serial-5 after second Chrome headless timeout in calm HTML export test * test(bin): skip PATH-based stat fault injection on Darwin where stat is /usr/bin/stat * no-mistakes(document): Refresh stat and shard docs
…henguid#3945) The captain's attribution policy (no Co-Authored-By trailer, no Claude-Session link, no generated-with line) lives in Claude Code's `user` settings scope. A spawned worker's settings sources are not guaranteed to load that scope, so a launched worker could write attribution trailers into its commits and PR bodies regardless of the captain's own configuration. launch_template()'s claude case now carries the same policy ("attribution": {"commit": "", "pr": "", "sessionUrl": false}) directly in its inline --settings JSON, so every claude launch keeps attribution off independent of which settings scopes end up loaded. Tests assert the policy on the rendered launch command for both a crewmate and a secondmate spawn. Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
…nchenguid#3946) * fix(bin): derive the away-mode beacon grace from the poll cadence fm-turnend-guard.sh's away-mode branch required the watcher beacon to be fresh within the flat FM_GUARD_GRACE default (300s), but the daemon starts a fresh one-shot watcher only after it finishes handling the previous wake, and that handling can legitimately outrun a fixed 300s window under load (a slow registered check, a busy supervisor pane) with the daemon perfectly healthy throughout. That misread a live, correctly-cycling daemon as down and blocked the turn. Add fm_poll_derived_grace, the single owner of the max(300, FM_POLL + 60) formula, and have the away-mode branch, fm-claude-stop-autoarm.sh, and fm-watch.sh's own runtime beacon-staleness check all derive their default grace from it instead of the flat default. A dead daemon pid or a beacon older than that grace still blocks, so a genuinely lapsed away mode still alarms; every other check is unchanged. fm-claude-stop-autoarm.sh computed the derived grace into GRACE but its two fm-watch-arm.sh invocations called the wrapper bare, so the wrapper fell back to its own flat 300s default and could reject a healthy long-poll watcher. Both invocations now pass FM_GUARD_GRACE="$GRACE" through explicitly, and a new test proves a long FM_POLL with FM_GUARD_GRACE unset reaches fm-watch-arm.sh with the derived value. Also drops fm_last_activity_age, added alongside the derivation but never called anywhere in the tree; fm-inactive-reconcile.sh already owns that computation. * no-mistakes(review): Remove dead WATCHER_STALE_GRACE assignment in fm-watch.sh * no-mistakes(document): Update FM_WATCHER_STALE_GRACE default note for poll-derived grace --------- Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
…henguid#3904) * fix(procevent): bind a source runner to the session that owns it A process-event source runner is detached into its own process group so a persistent source survives the turn that armed it. Nothing bounded that detachment, so a runner could reparent to init and keep its blocking child - and every process that child spawned - running with nothing left to reap it. One such runner outlived its home for about a day; the cost was not the runner but the exec churn of the poll stubs under it, which stalled every fresh process launch on the host. Each runner now starts a small guard beside it, in a separate process group, that re-reads its home's process-event lease and stops the runner's whole process group once that lease can no longer be proved fresh. Every ordinary entry point an owning session runs refreshes the lease, and the watcher's reconcile cycle keeps it fresh in a live home; nothing a runner spawns can refresh it, so a source cannot certify its own owner. Scope is the owning state root and one runner generation, never a script or process name, so a live source in another home is untouched and a live home simply starts a replacement runner on its next cycle. The test scaffolding that starts real runners could not reap them either: the bearings-board and board-render suites tracked their homes in a shell array appended to inside a command substitution, so the array was always empty and every listener they started survived the run. Home registration moves to a `$$`-keyed registry in tests/lib.sh, which sweeps it from every cleanup path, now including HUP and QUIT, and the blocking fixture stubs stop themselves at a bound so an escaped one cannot keep spawning processes indefinitely. Adds a regression test that reproduces the orphan shape - a reparented listener with a live descendant tree under it - and proves the whole group and its process churn stop once its session is gone, that an identical listener in a home whose session is still there is untouched, and that retirement still reaches a reparented listener and everything under it. * no-mistakes(review): Bound source launches and fail closed on guard startup * no-mistakes(document): Document runner lease and storm containment * no-mistakes(ci): Fixed the Greptile watchdog finding: failed runner cleanup now retries on each watchdog tick instead of abandoning the orphaned process group. The shared state-root lease behavior remains unchanged because it is an explicitly accepted ownership policy. Verified with bash syntax checks, git diff checks, and the complete fm-procevent test suite * test(procevent): pin that an unprovable stop is retried, not abandoned The owner guard used to call stop_runner_pid and exit unconditionally, so a stop it could not prove - a descendant still finishing uninterruptible work outlives even the group signal, and an unreadable process identity proves nothing - left a still-running expired runner with nothing watching it. That is the best-effort reaping this mechanism exists to remove, and the fix that made the guard retry landed without a test holding it in place. The unprovable attempt is injected through the signal the real path actually reads: `ps` answers exactly one process-group query for the runner with a group it does not lead, which is how a stop that cannot be proved is reported, and every other call is the real command. The test also asserts that the injected attempt happened, so it cannot pass vacuously if the fixture stops arming. Fails against the exit-after-one-attempt guard, where the runner survives its expired lease, and passes once the guard retries on its check cadence. * docs(procevent): scope the no-self-refresh rule to confused-agent grade The runner-lease documentation asserted as an absolute that nothing a runner spawns can refresh the lease, so a source cannot certify its own owner. That overclaims what the inherited FM_PROCEVENT_IN_RUNNER marker actually enforces. The marker holds at confused-agent grade: a runner and its ordinary children inherit it and skip every refresh, which is exactly the accidental case this boundary exists for. A source that deliberately strips the marker from its environment can still refresh, so adversarial-grade unforgeability is explicitly out of scope and tracked as separate follow-up design work. This states the real scope in docs/configuration.md, which owns the operating contract, and corrects the two matching comments in bin/fm-procevent.sh. The process-event-sources skill keeps its cross-reference and gains one line in its never-to-be-claimed list so the overclaim is not reintroduced from the agent-facing side. The lease mechanism itself is unchanged. * no-mistakes(review): Fix process-event lease and launch pacing edge cases * no-mistakes(review): Scope launch pacing and clarify lease boundaries * no-mistakes(review): Reap leftover groups and use monotonic launch pacing * no-mistakes(review): Keep guards alive across runner PID reuse * no-mistakes(review): Use monotonic leases and simplify launch generation identity * no-mistakes(review): Prevent pacing identity reuse and bound reused-group guards * no-mistakes(review): Preserve active pacing state on failed registration * no-mistakes(review): Reap reused runner groups with registration evidence * no-mistakes(review): Avoid ambiguous group kills and encode pacing identities * no-mistakes(review): Abort kill escalation after runner identity reuse * no-mistakes(review): Gate group signals and prune stale pacing state * no-mistakes(review): Document bounded PID reuse signaling safety * no-mistakes(review): Align leaderless group ambiguity guidance * no-mistakes(review): Expire reboot stamps and preserve publication success * no-mistakes(review): Bind owner leases to physical state roots * no-mistakes(document): Clarify process-event lease and pacing contracts * fix(procevent): drop a platform-dependent post-TERM test assertion CI ran red on two lanes that the local gate could not see. Lint failed with SC2034 on two reads in cmd_owner_watchdog that destructure the state-root identity into five fields while using only the device and inode. Local changed-file mode suppresses the cross-file codes that need --external-sources, so the warning cleared the pre-push lint step and failed CI's full analysis, exactly as bin/fm-lint.sh's header describes. The unused fields now read into `_`. The behavior shard failed on this suite's own post-TERM assertion, which required the stubbed identity source to be consulted more than once. Whether that happens is platform-dependent: where the runner leader keeps waiting on its TERM-ignoring source child, the post-TERM check sees a live leader whose identity no longer matches, and where the leader dies promptly it sees a leaderless group carrying the same numeric id. fm_procevent_pid_state reaches that second verdict without consulting process identity at all, so the identity source is never read twice and the count assertion fails through no fault of the behavior. The case now asserts the invariant both forms share: retirement refuses, and the ambiguous group is not signalled. Scoping a mutation to this fixture and making the refusal signal instead confirms the case still fails, so dropping the count does not leave it passing vacuously. * no-mistakes(review): Prevent superseded runners recreating stale pacing stamps * no-mistakes(document): Document pacing and ambiguity boundaries * fix(procevent): retire under the recorded identity source and state the home-scoped lease The reused-group case started its runner with the proc-root override in place, so the runner recorded a ps-derived identity, then retired it without that override. Where /proc exists the retirement read identity from a different source than the one recorded, the guard correctly refused an identity it could not confirm, and the case failed on Linux while passing on macOS. It now retires under the same source, and clearing the stub marker first turns that cleanup into the complementary assertion: once the ambiguity is gone, retirement reaps the whole group instead of leaving it behind. The lease prose claimed a runner is bound to the session that owns it, while the mechanism binds it to the home. That gap is what makes a replacement session or an inspection command look like a defect: any activity in the same home refreshes the lease. The granularity is deliberate, because a persistent source is meant to outlive the session that armed it, and binding a runner to that session would stop the sources this mechanism exists to keep running. A runner whose source is no longer wanted in a live home is stopped by reconcile when that source is retired, independently of the lease, so the lease is the backstop for a home that is gone - the torn-down sandbox this change bounds - and the residual is recorded as a known limit. * no-mistakes(review): Rate-limit polls and skip superseded runner launches * fix(procevent): build the claim-only sweep case as a runnerless owned claim A superseded generation now observes the registration-identity mismatch, self-retires, and releases its claim, which is the behavior we want: it clears its own residue rather than leaving a claim with no runner for the home sweep to find. The claim-only sweep case was built by deleting a registration out from under a live runner, which used to leave that runner in place. It now makes the runner retire itself, so the sweep raced that exit and retired one source or two depending on which won. The case failed three runs in four, alternating between a preflight-count failure and `attempted=1`. It now builds the state it means to test: kill the runner's group so it cannot run its own cleanup, assert the owned claim survived that kill, and only then drop the registration. Coverage is unchanged - a runnerless owned claim must still be swept - and the result no longer depends on whether the runner had exited yet. Three consecutive runs pass. The superseded exit also skipped the runner-marker cleanup the normal path performs. The marker is written before the launch floor is waited on, and a home sweep counts a marker with no owned claim as a preflight failure, so exiting without clearing it would make that home refuse to sweep. `FM_LAVISH_POLL_RETRY_DELAY= ` trips SC1007 under the full analysis CI runs, though not under the changed-file mode the pre-push gate uses. * test(procevent): retire a quiet reparented listener instead of racing a storm Explicit retirement was exercised against the spawn-churning stub, which made it nondeterministic. Retirement refuses rather than signalling when it cannot confirm the runner's identity, that identity is read through `ps`, and the stub's 0.1s spawn loop starves that read often enough that a single attempt is a race - the suite failed on this case roughly one run in four, reporting `cannot confirm runner identity; source remains registered`. The refusal is correct: it is the documented preserve-for-retry contract, and a separate case already asserts it. So this is a fixture problem, not a behavior problem. The storm is still covered where the evidence for it lives. The owner-loss home keeps the churning stub and still asserts its tick log stops, which is what proves the churn ended rather than one pid going away. The retirement home never asserted ticks; it only ever read the descendant pid, so the spawn loop bought this case nothing while costing it determinism. It now uses a quiet stub that still reparents and still holds a real descendant in its process group, so the assertions are unchanged: retiring the source must reap the reparented listener's whole group and the descendant under it. Four consecutive runs pass. * no-mistakes(review): Serialize registration replacement through source child launch * chore(no-mistakes): require honest test-step scenario marking The test step recorded scenarios as passing that were only reached through a stubbed dependency or the executable suite, and its validator refused them, because `pass` asserts a scenario was verified against the real live product. That refusal is correct, so the fix is to mark honestly rather than to weaken the gate: a scenario driven live stays a pass and cites its live transcript, while one reached only through a stub or the suite is recorded as untested with the reason and a pointer to its executable coverage. Untested scenarios are reported rather than treated as failures, so real coverage stays visible without claiming verification that did not happen. The instruction also forbids dropping a scenario to avoid marking it untested, since that would hide the gap instead of stating it. * no-mistakes(review): Remove unrelated test scenario policy * no-mistakes(document): Clarify process-event home lease documentation * no-mistakes(document): Correct owner guard failure wording
…nguid#4037) * test(lib): set fixture mtimes through one portable epoch helper On macOS the visible symptom was ONE red case in the turn-end guard suite. The actual damage was TWO cases that had quietly stopped testing their subject. The red one was the harmless half - people read one red case as one broken thing, and here that intuition is wrong. `touch -d @<epoch>` is a GNU extension; BSD touch rejects it outright and leaves the file at its current mtime. So on macOS the three away-mode beacon cases never aged their beacon at all. The 400s case exists to pin that 400s is stale under the flat 300s default but fresh under the poll-derived grace (660s at FM_POLL=600). Deleting the grace it guards (FM_POLL=60, so max(300,120)=300) and re-running proves what it was worth on this platform: pre-fix input (beacon left at now): ok - passes with the feature DELETED post-fix input (beacon 400s old): not ok - expected exit 0, got 2 It was green while measuring nothing, and could not have caught a regression in the grace it names. Only the 700s case broke loudly. `touch -t [[CC]YY]MMDDhhmm[.SS]` is POSIX and both platforms accept it, so the only host-specific step left is formatting the epoch into that stamp, which date(1) spells two incompatible ways. fm_touch_epoch in tests/lib.sh owns that probe once and fails loudly rather than leaving an unset timestamp behind - the failure mode that caused this. Verified on BSD touch/date here and on GNU coreutils 9.7 in a container. Three real sites, and one consistency change - not four fixes. The stale destination lock in tests/fm-remote-backlog-handoff.test.sh was never a defect: its `uname = Darwin` branch made the `touch -d` line unreachable on macOS, and BSD touch accepts that space-separated form anyway. `touch -t` takes the date directly on both platforms, so the branch goes rather than standing as a second copy of the same platform assumption. Known limit: the third case (away mode off) is only HALF recovered here. It now receives the input its name claims, but it is still insensitive after this fix - its verdict is identical with a 0s and a 400s beacon, because the fixture records a daemon lock and no watcher lock, and with away mode off the daemon lock proves nothing. Not fixed here; tracked separately, with the requirement that any fix be shown to FAIL when the protection is removed. FULL SUITE ON macOS: 189 scripts, four red, none of them this change. The turn-end guard and remote-handoff suites are clean. Attribution was established by running the four failures at the base commit and at this head on an idle machine, because base-idle against head-under-load moves two variables at once: script head/loaded base/idle head/idle verdict fm-calm-pi-extension red red red pre-existing fm-backlog-atomicity red red red pre-existing fm-procevent red red red pre-existing fm-startup-network red green green cause unestablished, load-sensitive under a full run Reported, not fixed. fm-calm-pi-extension deserves its own note: it FAILS because Chrome is absent instead of declaring the capability it needs and standing aside, so its verdict is about the machine rather than its subject - the same family as the defect above, with the red at least announcing itself. Neighbouring class, reported not changed: `file_mode()` - a verbatim `uname = Darwin ? stat -f %Lp : stat -c %a` - is copy-pasted across at least five test scripts plus a `reread_mode` variant, and epoch-mtime reads are open-coded as `date -r … || stat -c %Y` in three more; same one-owner shape as the defect above. `git init` without `-b main` depends on the host's init.defaultBranch in several scripts (branch-name case, tracked elsewhere). timeout, sha256sum and sed -i uses are all correctly guarded where checked. Observed while building the check rather than the fix: the first watcher I wrote to wait for the suite matched its own command line, so it was waiting on its own existence and could never fire. Same shape as the cases above - machinery answering confidently about something other than its subject, by including itself in the evidence it was meant to judge. The file sentinel it was replaced with cannot be produced by the observer that reads it. * fix(review): Pin fixture timestamps to UTC across DST transitions * fix(document): Clarify shared fixture suite coverage
…ision (kunchenguid#4038) * fix(pi): preserve native Codex effort and guarded supervision * test(pi): identify native compatibility guard versions * no-mistakes(review): share native-main follow rule between build and picker * no-mistakes(document): document native progress marker and ultra effort owners * no-mistakes(ci): Failing check "Behavior portable serial 1" was caused by this PR. The shard ran the default-on live guard tests/fm-pi-branch-responsiveness-live-e2e.test.sh, whose idle arm loads .pi/extensions/fm-branch-supervision.ts into a scratch project with a fixed list of copied libs. This PR added `import { registerFirstmateTool } from "./lib/fm-native-contract.ts"` to that extension and updated every other loading fixture's copy list, but missed this guard. Pi 0.85.1 therefore refused to load the extension ("Cannot find module './lib/fm-native-contract.ts'"), never drew its TUI, the test failed with "Pi 0.85.1 never drew its TUI in the idle arm", and the job hit its 20-minute cap. Fix (one line): added fm-native-contract to the lib copy loop in tests/fm-pi-branch-responsiveness-live-e2e.test.sh. Swept all other suites referencing fm-branch-supervision.ts / fm-primary-pi-watch.ts; the remaining ones without the new lib only hash, path-reference, or string-match the files and do not load them into Pi, so no further fixture changes are needed. Verification: reproduced the mechanism against the installed Pi 0.85.1 by building the lab copy with the old lib list (load error as above) and the fixed list (loads cleanly). tmux is not installed on this machine, so the live guard itself gate-skips locally ("skip: live: tmux absent") and could not be run end to end here; CI (which has tmux and Pi) will exercise it. shellcheck is clean on the edited file --------- Co-authored-by: Talon Stark <talonstark@gmail.com>
…guid#4041) * fix(herdr): step around a stale client the running server refuses A remote host can carry a self-updated herdr in ~/.local/bin beside a package-managed one, and the fixed remote-job PATH resolves ~/.local/bin first. After the server upgraded to 0.9.0 (protocol 22) the stale 0.8.2 client (protocol 20) was answered with protocol_mismatch on every command, which the read classifiers folded into `unreadable`: the live remote secondmate read unknown, every doorbell into it failed, and both the spawn and relaunch recovery paths refused, so the defect trapped itself. The adapter's session-scoped CLI wrapper now recognizes that refusal, reads status per session from each distinct herdr on PATH, adopts the first one the running server reports compatible, retries once, and keeps it for the process. The happy path makes no extra call and no other failure reselects. An endpoint that still reads unreadable names the refused client, both protocols, and the fix on stderr; the remote state read, fm-crew-state, and the launch refusal carry that reason, and fm-remote-doctor reports the selected client and rebinds the launch agent to it. Regression coverage: fake two-client hosts in the herdr unit suite, the doctor suite, the crew-state remote arm, and the real host-local control script in the remote lifecycle e2e; the real-herdr smoke refreshes the status shape the selection reads. * no-mistakes(review): Reselect Herdr client after every protocol mismatch * no-mistakes(review): Remove unrequired Herdr diagnostics and launch-agent rebinding * no-mistakes(review): Scope cached Herdr clients to their selected session * no-mistakes(review): Restrict herdr client selection to reactive CLI calls * no-mistakes(document): Document session-scoped Herdr client reselection * no-mistakes(document): Clarify Herdr client selection documentation * no-mistakes(ci): Updated the trusted fm-remote-doctor.sh SHA-256 in bin/fm-remote-entrypoint.sh after the PR changed the doctor, restoring git-unavailable bootstrap authentication. Verified tests/fm-on.test.sh, tests/fm-backend-herdr.test.sh, bin/fm-lint.sh, and git diff --check all pass
* feat(afk): record the away posture and its lifecycle (phase 1) Away mode becomes a posture of the one supervision session, recorded in state/.afk-contract by the new bin/fm-afk-contract.sh: the one owner of the record schema, the mandate-clause grammar and compiler, refusal naming the missing part, the read-back rendering, the entry announcement (hold-for-return only, no phone channel), and the archive at return. This release records clauses and does not execute them; the announcement and return brief say so. bin/fm-afk-launch.sh gains propose and confirm, confirms the record before any daemon launch, refuses to launch the daemon on Pi and pi-signed, and archives the record last on stop. bin/fm-afk-return.sh snapshots supervisor health before shutdown, renders the return brief (health, mandate, waiting on the captain, could not fix, handled, cost) from the archived record, the outcome store, the held set, and the status logs, and shrinks the blocker gate to what the away session could not fix. While the record exists the watcher and the daemon never recheck an item held for the captain. Declared external waits get a four-hour default cadence and honor `until <UTC ISO 8601>` on the paused line, in both postures, bounded by FM_PAUSE_UNTIL_MAX_SECS. The /afk skill, AGENTS.md's layout and away-mode stub, the session-start digest, and the architecture, Pi branch, configuration, and scripts docs describe the record. The Pi/Herdr e2e now proves the no-daemon posture on a real Pi primary; its verification record carries the 2026-09-08 run. * no-mistakes(review): Fix AFK confirmation, grammar, waits, and return gating * no-mistakes(review): Harden AFK authority and posture lifecycle * no-mistakes(review): Preserve AFK history and tighten authority grammar * refactor(afk): record clause fields with no natural-language parser By the captain's mandate the away-posture record keeps no static parser that tries to understand natural language. A mandate clause is now given as explicit fields (--action, --object, --when, optional --stop) that bin/fm-afk-contract.sh records verbatim. The structural check asserts only that the action, object, and precondition fields are present and that the action is a listed verb; whether a precondition holds is the supervision session's judgment at execution time in a later phase. The never-set stays as a forbidden-concept safety scan: fields mentioning credentials, passwords, logins, legal or financial acceptance, payments, invoices, one-time codes, or an attended prompt are refused, matched at token prefixes after punctuation normalization so compound and plural spellings are caught. The red-check grammar, class-word rejection, unconditional-word detection, clause-reference resolution, and condition aliases are removed. --words-file keeps the captain's words verbatim, trailing newline included. The skill, docs, launcher help, and tests describe the field form. * no-mistakes(review): Preserve AFK words and tighten safety refusals * no-mistakes(review): Preserve clause bytes and honor declared waits * no-mistakes(review): Harden deny-list and gate unreadable outcomes * no-mistakes(review): Demote never-set scan and clarify authority * no-mistakes(review): Gate return on unreadable held and status data * no-mistakes(review): Validate posture archives and enforce Pi detection * fix(afk): make the never-set a non-refusing flag and keep return fail-safe Per the captain's decision the never-set scan is a coarse best-effort flag, never a refusal and never the gate: a clause naming a listed concept is still recorded with a flag the read-back, announcement, and return brief show, and the scan matches listed terms exactly or with a plain inflection at punctuation-delimited token boundaries, so unrelated names such as ping-service or tokenize-worker are never flagged and joined compounds remain a documented miss. Authoritative never-set and forbidden-action enforcement is the supervision session's judgment at execution time in phase 4. A replacement copies the superseded record through a temporary name and renames it atomically so a failed copy leaves no partial archive, the record owner gains validate and flags subcommands, and the return keeps catch-up gated when a superseded archive cannot be read. * no-mistakes(review): Harden AFK record validation and return reconciliation * no-mistakes(review): Harden AFK record validation and simplify commands * no-mistakes(review): Harden mandate validation and retain missing records * no-mistakes(review): Refuse blank explicit mandate stops * no-mistakes(review): Recover restored posture epoch before return * no-mistakes(review): Prevent return brief status symlink reads * no-mistakes(document): Refresh AFK posture documentation * no-mistakes(ci): Fixed both CI failures: updated lint telemetry for the new fourth source directive, quoted the hyphenated fixture value, and removed unreachable test cleanup. Verified with tests/fm-lint.test.sh, targeted CI-mode ShellCheck, bin/fm-lint.sh, bash syntax checks, and git diff checks * no-mistakes(ci): Bound structured pause deadlines by FM_PAUSE_RESURFACE_SECS in watcher and daemon housekeeping, added distinct bounded-horizon reasons, regression coverage for near, passed, and wrong-year deadlines, and updated documentation. Verified targeted behavior tests, full daemon tests, ShellCheck source-following lint, syntax, and diff checks
) Opt-in IMAP/SMTP mail plane (fm-mail.sh / fm-mail-check.sh). Absent FM_MAIL_* stays off. Speaking as Kun's firstmate: this is merged. Thank you @feilipu — really appreciate you taking the time on this.
* fix(remote): start the fm-remote Herdr agent through a login shell Launchd was exec-ing herdr directly, so the Aqua agent inherited a background session without login-keychain access. Start it via /bin/zsh -lc exec so panes keep login env and can refresh OAuth tokens after reboot. * fix(remote): start fm-remote Herdr via the account login shell Resolve UserShell from Directory Services and invoke it with separate -l and -c so bash, fish, and zsh all get login-keychain access. Fall back to SHELL, then /bin/zsh, then /bin/sh without failing the render. * no-mistakes(review): Fix launch-agent shell fallback resolution * no-mistakes(review): Preserve and escape Directory Services shell paths * no-mistakes(document): Document login-shell LaunchAgent behavior * no-mistakes(ci): Updated the trusted fm-remote-doctor SHA-256 identity in bin/fm-remote-entrypoint.sh. Verified with tests/fm-on.test.sh, tests/fm-remote-doctor.test.sh, bash syntax checks, and git diff --check * no-mistakes(ci): Resolved the login shell exactly once per doctor invocation and threaded it through plist rendering, installed/loaded contract validation, repair reporting, and post-repair checks. Added a regression test proving repeated repair remains healthy and performs no reload when a hypothetical second Directory Services lookup would differ. Updated the trusted doctor hash. Verified doctor, fm-on, remote-entrypoint, lint tests, ShellCheck, syntax, and diff checks * no-mistakes(ci): Made Darwin shell resolution hermetic with executable injection and a 2-second Directory Services timeout. Updated tests to inject shells by default, isolate dscl-specific cases, parse plists semantically, and verify stalled dscl fallback. Updated the trusted doctor hash. Doctor, fm-on, entrypoint, syntax, hash, and diff checks pass * no-mistakes(ci): Raised portable serial CI timeout from 20 to 30 minutes, refreshed the specified timing hints, added missing hints, and recomputed shard documentation. Verified coverage, runner behavior tests, workflow lint tests, shell syntax, requested timing maxima, and diff checks
…enguid#3710) * fix(bearings): keep captain-approved deliveries in Recently Landed A closed task is never held: tasks-axi clears the held flag when a task closes and keeps hold-kind and the hold reason as the record of the call that was made. Recently Landed excluded every Done row whose hold-kind was captain, so the marker it treated as "closed while still waiting on the captain" was in fact the proof that the captain had approved the work. Every merge routed through a captain decision disappeared from the list of what shipped, including under --all-landed. The selector now asks whether the closed row delivered something. Recently Landed is merged PRs, completed scouts, and finished local-only merges, so a row carrying one of those artifacts belongs there whoever approved it. A captain question closes with an answer and no artifact of its own, and that is what still stays out, so an answered question is never rendered as shipped work. The same rule was written twice - the bearings projection selects this home's Done rows and the fleet snapshot selects each secondmate home's Done rows into the roll-up the same section merges in - which is why one defect hid deliveries in every home. Both now share bin/fm-landed-lib.sh. * fix(review): Normalize landed evidence and exclude answered captain questions * fix(review): Normalize captain delivery evidence across relocated data * fix(review): Record authoritative delivery provenance with legacy fallback * fix(review): Harden delivery provenance across forced and pruned completions * fix(review): Replace premature merge closure with existing release contract * fix(review): Document provenance-based Recently Landed selection * fix(document): Align documentation with completion provenance * fix(lint): Fix targeted ShellCheck warnings * fix(ci): order the pinned tasks-axi install before its stock-Bash consumers In `.github/workflows/ci.yml` the pinned tasks-axi install now precedes both stock-Bash consumers, and the Bearings expectation is updated from 49 to 50 tests. Verified with macOS Bash 3.2: snapshot 16/16, Bearings 50/50, public-followup 1/1. Full repository lint and all three workflow validations pass, and `git diff --check` is clean. * fix(review): Make completion provenance unambiguous * fix(review): Make completion verdict authoritative over quoted provenance * fix(review): Preserve retained artifacts through resumed captain closes * fix(review): Unify completion provenance ordering across writer and reader * fix(review): Preserve artifacts across failed captain closes * fix(review): Refresh v1 assertions; provenance authority remains unresolved * fix(review): Remove unreliable provenance while preserving landed deliveries * fix(review): Reject stale home summaries visibly * fix(review): Restore retained deliverable recording * fix(review): Match landed artifacts and restore retention documentation * fix(review): Disambiguate captain calls and restore landed artifact matching * fix(review): Persist retained report and PR artifacts * fix(review): Preserve staged artifacts before captain answers * fix(review): Avoid wedging answers on unsupported report paths * fix(review): Exclude unreleased captain-held pull requests * fix(review): Exclude held local-only answers from landed * fix(review): Preserve retained scout reports across snapshot rendering * fix(document): Align landed lifecycle documentation with release semantics * fix(review): Enforce landed artifact-kind ownership * fix(review): Infer canonical task kinds in snapshots * fix(review): Require captain-hold release before merges * fix(review): Qualify merge lifecycle regression evidence * fix(review): Serialize captain holds with merge operations * fix(review): Document merge cleanup residuals honestly * fix(test): Replace vacuous Bearings regression with behavioral cases * fix(document): Align Bearings verification and merge lifecycle documentation * fix(review): Serialize merges and exclude captain calls from landed * fix(review): Harden merge identity and landed selection * fix(document): Clarify landed selector compatibility filtering * fix(bin): keep merge entrypoints usable on records without an incarnation The merge identity guard refused any task record with no spawn_gen field. That field identifies one exact incarnation, so comparing it across the wait for the merge lock is what catches a task relaunched while the merge was queued. Requiring it to be present is a different rule, and it refused every record written before the field existed: a legacy task could no longer be merged at all, and five behaviour suites refused before reaching the check they were written to exercise. The comparison only needs to notice a change. An absent field is now read as an empty incarnation and compared like any other value, so a record that gains, loses, or alters one is still refused, while a record that simply predates the field merges. An ambiguous or unreadable field stays an error, because a record that cannot name one incarnation cannot be compared. The missing-record message each entrypoint had before the guard is restored, so a genuinely absent record still says so in its own words. The role partition now precedes reading the record. Refusing the supervision branch is a statement about the actor, not about the task, so it cannot depend on a record the wrong actor may not have. A backlog file that does not exist meant "no longer an open captain call". For a caller that asked to tell absence apart it now means absent, so a board card whose home carries no backlog stays visible instead of being dropped as resolved. Fixture repositories pin their initial branch instead of inheriting init.defaultBranch, which resolved to main on a developer machine and master on a runner, so a fixture naming main failed only in CI. * fix(review): read local-only note from body; surface pending-close failures * fix(review): keep kindless local-only landings in Recently Landed * fix(review): bind local-only note scan to the tasks-axi note line * fix(review): Guard unavailable captain-hold authority records * fix(document): Document unreadable authority predicate outcome * fix(bin): read an absent backlog as absence, not an unreadable record The merge gate refused every task whose home carries no backlog file. A backlog that does not exist holds no captain call, so nothing can be held and the merge is safe; only a backlog that exists and cannot be read may hide a live hold. Those two states were collapsed into one refusal, which stopped merges in any home that keeps no backlog. The predicate now reports a missing backlog file as absence, alongside a row the backlog does not carry. A record that exists but cannot be read still leaves by the existing cannot-tell path, which both merge entrypoints already refuse, so the restrictive direction is unchanged. That leaves no way to reach the separate unavailable-record result, so the result and the two branches that handled it are removed rather than left describing an outcome that can no longer occur. The lifecycle documentation loses the same claim. Regressions cover both directions in each entrypoint: a home with a task record and no backlog merges, and a backlog present but unreadable refuses without reaching the forge. * fix(review): Fail closed unreadable backend configuration * fix(tests): pin the bare origin's initial branch in the remote seed fixture The fixture created its bare origin with no initial branch, so that repository's HEAD followed init.defaultBranch while the source repository pushed the branch fm_git_init_commit pins. On a host that still defaults to master the two disagreed: the bare origin's HEAD named a branch the push never created, cloning it warned that the remote HEAD referred to a nonexistent ref and checked out nothing, and the seed assertion for the cloned README failed. A machine whose default is already main paired the two by accident and hid it, which is why the fixture passed locally and failed on the runner. Pinning the bare origin to the same branch removes the dependency on the ambient default from both sides. Verified under both conditions: with init.defaultBranch set to master, and set to main, the suite passes 26 of 26. * fix(review): Fail closed unreadable user backend configuration * fix(bin): republish the home summary as v1 and record two load-bearing rules The published home-summary schema had moved to v3, which routed every secondmate home still emitting the earlier version to the stale branch: their landed rows, open decisions and holds all came back empty and their state read as unknown until each home was updated. The payload never justified that. Its field set, field order, truncations and the landed array construction are byte-identical to v1, so only which rows the selector places in landed differs, and a v1 consumer reads that the same way. Republishing as v1 removes the rollout regression and, with it, the tolerance machinery that existed only to soften the bump: the stale-schema predicate, its two collection branches, the flag and its provenance branch, the omitted surface that can no longer be reached, and the fixtures and assertions that covered them. Two rules that a scope review proposed removing are kept, each now carrying the reason it exists, because both were measured to be load-bearing: The artifact-kind ownership clause is what keeps an explicit scout that recorded no report out of Recently Landed. Without it such a row has none of the three artifacts, satisfies the compatibility fallback and renders as shipped work with an empty artifact. The kind fallback is needed because tasks-axi omits the kind metadata entirely when a title begins with a canonical keyword. Without it a scout titled "SCOUT ..." reports no kind, its recorded report stops counting as a delivery, and it drops out of the section this selector exists to repair. * fix(bin): move the scout guard note onto the rule and drop two dead pieces The LOAD-BEARING note sat on an unreachable branch. Measured in both directions: removing that branch together with the kind-is-not-scout guards lets an explicit reportless scout into Recently Landed and fails tests/fm-captain-hold-lifecycle.test.sh, while removing the branch alone leaves that suite passing at 49 assertions. The guards carry the rule, so the note now sits on them and the unreachable branch is gone. A note pointing a later reader at the wrong line is the hazard this change corrects elsewhere. summary_file_has_schema lost its only caller when the stale-schema machinery was removed, so it goes with it. * fix(review): Fix legacy report artifacts and canonical keyword boundaries * fix(review): Update pinned Bearings test count to 56 * fix(document): Clarify landed summary compatibility documentation * fix(review): Preserve unreadable backend configuration errors * fix(review): Honor backend resolution errors at existing call sites * fix(test): Stabilize remote collector tests under host load * fix(document): Document backend resolution failure contracts * fix(lint): Suppress intentional deferred probe expansion warnings * fix(ci): Captain, quoted the two literal test IDs in tests/fm-backlog-atomicity.test.sh to fix SC2100 without changing behavior. Both warnings reproduced before the fix; the targeted fm-lint.sh run now passes with ShellCheck 0.11.0. Bash syntax and git diff --check also pass * fix(ci): Fixed the resolver’s two configuration-parent checks to return 2 for inaccessible directories while preserving genuine absence. Added two behavioral tests; RED/GREEN and both requested mutation proofs confirmed. All 10 focused checks, targeted lint, syntax, and whitespace checks passed. Broader merge suite stopped after 10 passing cases under host load. Declined portable checks and merge-authority code remain unchanged
…nguid#4090) * fix(remote): let the Aqua launch agent own the fm-remote Herdr session A herdr server keeps the macOS audit session of whatever started it, and only the Aqua login session (gui/<uid>) can read the login keychain without a prompt. Herdr's SSH remote attach starts the fm-remote server as its own child when it finds none, wins the socket at boot because sshd accepts connections before the login session exists, and every claude pane under that server then gets `security` exit 36, falls back to a stale plaintext credentials file, and reports "Login expired". launchd's own job lost the socket on every KeepAlive retry and the doctor still reported the session ready because it only asked whether any server answered. - Add bin/fm-remote-herdr-guard.sh, the launch agent's exec target: start the server in the foreground when nothing owns the socket, exit 0 when an Aqua-born server does, and otherwise stop the foreign server, wait for the socket, and exec the server at once. - Add bin/fm-remote-herdr-owner-lib.sh, the single owner of socket-owner discovery (lsof; pgrep cannot see herdr's argv on macOS) and the birth markers (SSH_*, XPC_SERVICE_NAME, FM_REMOTE_JOB_ACTIVE, sshd or remote-client-bridge ancestry matched on argv[0] and whole arguments). - Render the agent as the login shell exec'ing the guard with KeepAlive={SuccessfulExit=false} and ThrottleInterval=10, check the loaded job's successful-exit semaphore, and report a session served outside the Aqua login session as fixable so --fix retakes it through launchd; the reload waits for an Aqua-born owner rather than any running server. - Correct the doctor and docs: the launch shell provides environment parity, the launchd domain provides keychain access. - Pin the guard's decision table and the doctor's verdicts against real marker-carrying processes, and record the dated audit-session evidence. * no-mistakes(review): Verify Aqua ownership through launchd domains * no-mistakes(document): Document macOS lsof ownership requirement
…uid#2881) * fix(bin): prefer a live no-mistakes run over a terminal one A worktree can bind to more than one recorded no-mistakes run at once. The branch-and-code-identity rule in bin/fm-nm-run-lib.sh accepts both an exact-equal commit and a worktree-is-an-ancestor match, but never stated which wins when both bind, so the tie fell to whichever candidate the caller reached first. Observed on a live fleet: a crashed validation daemon left a FAILED run at the worktree's own commit while the live run that replaced it validated a descendant commit on the same branch. Bare `axi status` answers with the most-recently-touched run - the corpse - and it bound by the equal-commit rule, so every recomputation reported `failed` for a task whose real run was healthy. The same label had also read `failed` earlier while the work was genuinely stalled, so the signal was wrong in both directions. State the live-over-terminal policy in the matching rule's own contract, where the equal-commit and ancestor rules already live, and add fm_nm_run_status_class as the one classifier that decides liveness from a recorded status word. fm-crew-state.sh applies it on both selection paths: the runs listing now scans past a terminal row for a live one, and a terminal `axi status` answer is provisional until the listing has been asked whether this worktree also has a live run. Same-liveness-class candidates keep the listing's newest-first precedence, and a status word the classifier cannot place keeps the caller's own ordering rather than displacing a known result, so a single-run task and a task whose runs are all terminal are unchanged. Regression coverage reproduces the proven case (terminal run at the worktree's exact commit plus a live run descending from it) and its runs-list twin; both fail under the old tie-break. Two companion cases pin the no-widening half - two terminal rows still resolve newest-first, and a terminal run with no live sibling keeps its full run-step detail - and both pass before and after the change. * no-mistakes(review): accept unfetched live sibling anchored at exact worktree head * docs(bin): name both ledger reads behind the runs-limit setting The FM_CREW_STATE_RUNS_LIMIT comment in bin/fm-crew-state.sh still described the runs ledger as scanned only by the cross-branch fallback, but the live-over-terminal fix also consults it as the live-sibling probe behind a terminal axi status answer. Point the comment at docs/configuration.md as the setting's owner instead of restating a second copy. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
… only (kunchenguid#5520) * fix(bin): drop status prose from the inactive-outcome dedupe identity The inactive-outcome receipt fingerprint included the child's sanitized last status line, so a persistent child appending routine prose after one terminal outcome minted a fresh parent event per sentence. Bind the identity to incarnation, task id, terminal state, and PR only, keeping the last line in the record as status_head evidence. Fixes kunchenguid#2960 * no-mistakes(document): note structured-only inactive receipt identity in regression coverage
…unchenguid#5707) * feat(bin): record the supervision host's dialog mirror on Claude and Cursor Add bin/fm-host-mirror.sh, the one owner of the supervision host's dialog mirror file, cursor, lock, and feed, plus the main-session key it keys entries to. The tracked Claude UserPromptSubmit and Stop hooks and the Cursor beforeSubmitPrompt and afterAgentResponse hooks record the captain's prompt and main's reply, only on a home with config/supervision-host, from a genuine primary checkout, for the lock-owning session. The mirror lands inert: writers record and nothing reads it yet; attended supervision on the host is the later step that consumes the feed. Codex, Grok, OpenCode, and omp have no writer here. * no-mistakes(review): Scope mirror dedup to session, atomic appends, marker-inclusive caps * no-mistakes(document): Clarify dialog mirror scope and remove duplicate contract details * no-mistakes(document): Correct Cursor hook documentation for dialog mirror registration * no-mistakes(review): Pass mirrored dialog text to jq via stdin * no-mistakes(document): Clarify dialog mirror documentation and remove duplicate claims * no-mistakes(review): Preserve internal dialog whitespace; drop mirror check and verified modes * no-mistakes(review): Drop only identical mirror repeats; remove redundant chmod guard
… escalations are not repeated (kunchenguid#5731) * fix(bin): retire check-row receipts on branch acks and report an unchanged situation once * fix(bin): scope a branch acknowledgement's check-row receipt retirement to its granted sequences The away posture lifts the attended partition's check/decision exclusions, so a branch grant can name check-kind rows - but the branch-actor ack still assumed check rows were main-only and skipped every receipt scan. The queue row was consumed while its terminal-outcome .pending receipt stayed behind, and each inactive-reconcile cadence scan re-queued the same fingerprint. In the first real away window on the supervision host that re-escalated one unchanged held-PR situation on every cycle (~1,734 of 4,149 outcomes). A branch ack now scans inactive-outcome and inactive-reconcile receipts and commits secondmate stall receipts against exactly the sequences in its eligible-row snapshot - the same rows it consumes - instead of none. Attended grants still name no check row, so the scans find nothing. * fix(bin): store a repeated captain verdict as routine while the task's durable situation is provably unchanged fm-branch-outcome.sh append computes a mechanical situation key per captain row - metadata bytes, captured status-log endpoint and identity, live crew-state verb, worktree head - and anchors it in state/.<task>.branch-captain-key. A later captain verdict whose recomputed key matches is stored as routine with "unchanged since seq <N>:" prefixed to its summary, so one situation escalates once until something provably changes. A task with no readable status ledger is never demoted, an unreadable record fails toward reporting, and teardown removes the sidecar with the task's other branch records. The append-only store schema is unchanged. This covers both hosts: the Pi supervision branch and the supervision host both funnel reports through append. * docs: check rows are main-owned only while attended; the away posture grants them to the branch, whose ack retires their receipts exactly * test: the away-flood reproduction as a regression test (branch ack retires the receipt and later scans stay quiet), store-level dedupe coverage, and a branch-ack secondmate stall receipt case * fix(bin): restore the secondmate child devin-config cleanup path The branch-captain-key sidecar addition mistyped the sibling entry as .$child_id.devin-config.json, so a forced secondmate teardown would have stopped removing each child's real <id>.devin-config.json. Restore the original path and add a behavioral test that stops the child sweep mid-loop on a refused close, proving the cleaned child's devin config and captain anchor are both removed while the unconsumed child's records are retained. * no-mistakes(review): Key captain dedupe on the covered wake rows' fingerprint * no-mistakes(review): Drop captain-key demotion; prove one escalation on both surfaces * no-mistakes(review): Drop unrelated teardown test; cite both receipt test files * no-mistakes(document): Docs already match branch-ack check-receipt retirement
…oorbell (kunchenguid#5664) * fix(calm): deliver Claude-bound operational input as a record-backed doorbell Claude Code 2.1.280 removes U+2063 from every submitted prompt, so a typed operational envelope reaches a Claude Code primary as plain text. The away daemon now writes the envelope to a record under state/operational-inbox and types only a plain doorbell naming it; the /afk return check and the Calm mod recognize the doorbell only when that record holds a current envelope. Marker- preserving harnesses keep the typed envelope. The live Calm guard accepts the 2.1.280 module-load log line, drives the doorbell, and asserts thinking stays hidden. * no-mistakes(review): Fix operational record retention at 7 days and document prune limit * no-mistakes(document): Point Calm bounds at 2.1.280 evidence; fix afk-exit comment * no-mistakes(lint): Pick newest Calm e2e transcript without parsing ls * docs(calm): add a minimal turning-Calm-on step for Claude Code * fix(spawn): deliver the Claude launch brief as a record-backed doorbell Claude Code strips U+2063 from the launch-prompt argument too, so a worker's launch brief arrived with its operational marker removed. Publish the brief as a record in the receiving home's operational inbox - a secondmate's own state, not the primary's - and pass only the printable doorbell naming it, falling back to the typed envelope when the record cannot be published so the brief body still delivers. Unwrap doorbell-carried digests in the daemon digest tests that still read the raw send log under the claude pin, and update the documented bounds now that launch briefs hide like the other operational rows. * test(spawn): cover a secondmate's launch-brief record landing in its own home The record-backed doorbell resolves its state through the receiving pane's home, so prove a claude secondmate launch publishes into the seeded secondmate's operational inbox and never leaks a record into the primary's. * no-mistakes(review): Pass primary harness to daemon, tighten retention, refresh verdicts * no-mistakes(review): Prune operational records by exact seven-day elapsed age * no-mistakes(review): Batch record pruning so large inboxes still expire * no-mistakes(review): Refuse Claude spawn when brief record cannot publish * no-mistakes(review): Drop thinking probe from Claude Calm live test and docs * no-mistakes(review): Record dated Claude Code 2.1.282 reproduction evidence * no-mistakes(document): Clarify operational doorbell documentation and record expiry * no-mistakes(document): Correct AFK escalation carrier guidance * no-mistakes(review): Describe operational record retention as about seven days * no-mistakes(document): Clarify Calm delivery and operational record retention * no-mistakes(review): Remove out-of-scope Calm launch guide from Claude docs * no-mistakes(document): Document Claude launch-brief delivery and refusal * no-mistakes(document): Correct stale operational-input documentation * no-mistakes(ci): Fixed the stale Claude trust test to verify that worker and secondmate launches deliver readable, record-backed briefs instead of expecting brief paths in their commands. Annotated the daemon’s output variable for ShellCheck without changing behavior. The affected tests, daemon tests, ShellCheck, and diff check pass locally * no-mistakes(ci): parse rebased Claude launch after trailer hook prefix * no-mistakes(review): Trust launch-brief record and restore thinking bound doc * no-mistakes(review): Parse final Claude launch statement; drop Stop-hook docs --------- Co-authored-by: Mike Sewell <maikunari@protonmail.com> Co-authored-by: no-mistakes <no-mistakes@localhost>
kunchenguid#5583) A host-local relaunch rewrote only the far endpoint, so this home kept the old harness, model, and effort, and appending those keys after pr= broke pull-request poll authentication.
…ound (kunchenguid#5516) tests/fm-watch-triage.test.sh finishes in about 434s alone and about 698s under CI load, so the 900s bound the changed-suite runner applies produced a false timeout under ordinary concurrent validation. Raise the automatic bound to 1500s, which keeps every measured script under it while staying below the 30-minute normal CI tier so a genuinely hung script still fails here with its output before the job cap cancels the lane. Fixes kunchenguid#3869 Refs kunchenguid#3565
…kunchenguid#5728) * Fix nested watcher lock reclaim * no-mistakes(review): Elect a single steal-mutex reaper and bound arm TERM wait * no-mistakes(review): Reclaim self-held steal mutex and unify autoarm steal reaping * no-mistakes(review): Resume own interrupted steal reap from its tombstone
…nguid#5710) * test: hold the back-to-back boundary close on the host's own clock test_park_boundary_holds_under_back_to_back_closes assumed two engine turns fit in the ~16s pre-refusal window and that the stub finished a turn in 3s. Under load the stub's real drain, report, and acknowledgement take ~13s, so the turn either died at its bound (which hands the wake to main, no boundary line) or the second close landed past the window and the fixture failed while the boundary held. 3 failures in 5 runs at a load average near 11. Hold the first turn on a release file instead: once the engine is in flight, a second close is appended mid-turn and the turn is released as the refusal window opens (park bound minus turn bound and grace, read off the host's own start record). The queued close can then only wait for the boundary on any machine speed, which is what the test asserts: the boundary line ends the output, the demo.status row stays queued for main, and no second engine turn ever starts. A host too loaded to start the turn at all hands the first close to the same boundary exit. After: 12/12 at load ~15-42. * no-mistakes(review): Print boundary test deadline as a decimal integer * no-mistakes(review): Hold boundary test turn on a FIFO, require full sequence * no-mistakes(review): Remove stray before/after supervision-host test copies * test: hold the late close's render until the refusal window opens The boundary recheck test's node shim slept a fixed 10s, which assumed the first close was read before the host's refusal window opened. Under load the close arrived after the refusal check, so the host correctly refused it before the successor started and the render snapshot never appeared. Block the wake-prompt render on a FIFO released at the refusal-open instant read from the host's own start record, so the pre-turn recheck must refuse on any machine speed. * no-mistakes(review): Derive minimal park bounds and refresh supervision-host shard hint * no-mistakes(review): Drive park-boundary tests from a seam-gated host test clock
* fix(bin): stage remote home clones before publishing them A remote home provision cloned the code root directly into the public FM_HOME path while rollback() claimed rm -rf of that same path on any failure. Bash defers trapped signals past a foreground child, but any other cleanup or lifecycle path that removes the home directory races the live clone's object copy, producing the CI flake "fatal: failed to copy file to .../.git/objects/...: No such file or directory". Clone into a private staging directory beside the home and publish with an atomic rename once complete, so no cleanup can remove a directory a live clone is still writing; a home that appears mid-provision now dies cleanly instead of inheriting torn state. The regression coverage holds a real clone mid-copy, removes the public path, and requires the provision to finish and publish intact. * no-mistakes(review): Prove home ownership by sentinel and hold only a live clone * no-mistakes(review): Assert raced provision publishes a complete, intact clone * no-mistakes(document): Document remote home staging and publication safety * no-mistakes(lint): Fix ShellCheck warning in clone integrity assertion * no-mistakes(document): Clarify remote home publication and rollback guarantees
* fix(control): keep a relaunched Pi worker's herdr pane status authority alive Defect: after `bin/fm-control.sh <id> relaunch` (observed live on a herdr Pi crewmate whose pane read idle while it ran its validation pipeline), the pane froze at whatever its previous agent had last reported. Cause, measured on herdr 0.9.1 against a real Pi: a pane has one status authority, and for Pi with its integration installed that authority is the lifecycle hooks, so herdr also skips screen detection for the pane. In the crew shape the registration outlives its agent process (upstream issue kunchenguid#4115; docs/herdr-backend.md "Restart and liveness behavior"), and herdr applies only reports carrying the session identity it bound. A replacement started fresh in that pane reports a NEW session, so its state reports are ignored and the pane stays frozen. Nothing from outside repairs it: `pane report-agent-session` and `pane report-agent` for `herdr:pi` are accepted (rc=0) without being applied unless the reporter is the registered pane agent, and `pane release-agent` on the stale record changes nothing. Fix: a relaunch preserves the binding instead of fighting it. The launch owner reads the session reference the endpoint's own runtime recorded (`fm_backend_herdr_pane_agent_session_ref`) and passes it back as Pi's own `--session <path-or-id>` (`relaunch_resume_args`; `fm_control_relaunch_resume_flag` owns which adapters and which registered-agent labels qualify). That is the same reference herdr itself resumes Pi panes with after a server restart, and the resumed session's reports land again, which the live check confirmed: the pane returned to working while the replacement worked and idle when it settled, on the same session identity. Safety: relaunch-only (a fresh spawn binds nothing), herdr-only (the one adapter that records a per-pane session), Pi-family only, and only when the registration's own agent label matches - so no other adapter's conversation can be handed to a Pi launch. An unreadable, missing, or malformed reference degrades to exactly the fresh-session launch that existed before. No lifecycle, liveness, isolation, or merge guard is touched, and an empty result leaves every non-Pi launch byte-identical. `resume` remains a refused verb; docs/agent-control.md and the harness-adapters references are corrected where they claimed Pi had no verified resume form at all. * no-mistakes(document): docs: correct relaunch session-authority ownership and skill paths * no-mistakes(document): docs: correct stale control-plane ownership claim * no-mistakes(document): docs: drop unverified Herdr restart resume claim * no-mistakes(test): Added offline Herdr Pi session-authority relaunch coverage * no-mistakes(document): Document Herdr Pi relaunch session continuity * no-mistakes(ci): The failing remote relaunch test tried to arm a PR poll for a secondmate, which `fm-pr-check.sh` correctly refuses. Removed that invalid test scenario; the remaining remote relaunch tests pass, and `git diff --check` is clean
…nchenguid#5758) Main has been red since fm-pr-check.sh began refusing to arm a merge poll on a kind=secondmate record (kunchenguid#5696): the relaunch-ordering case in tests/fm-remote-secondmate-relaunch.test.sh armed its fixture through that entry point and could no longer be set up. The ordering guarantee still matters: a secondmate record armed before the refusal can legitimately carry a trailing pr=/pr_head= identity block until the watcher retires it, and fm-remote-secondmate-relaunch.sh must still keep that block last when republishing harness/model/effort. Seed the fixture the way such a record was really written - pr= appended last to the meta, then the poll artifacts published through the same fm_pr_poll_prepare/fm_pr_poll_publish_prepared pair fm-pr-check.sh uses, a pattern tests/fm-pr-check-security.test.sh already follows - and drop the now-unused fake gh fixture. The kunchenguid#5696 refusal itself stays pinned by the security suite's secondmate-record case.
…id#5748) * feat: run attended supervision on the host for Claude and Cursor On a home opted into config/supervision-host with a Claude or Cursor primary, the supervision host now takes the attended wakes the Pi branch would take: routine outcomes stay off main, and a captain outcome wakes main once with a branch-outcome line and waits in the drain's new BRANCH OUTCOMES section until main acknowledges it with mark-processed. - The offer rule moves into branchOfferForWake, shared by the Pi watcher and the host through bin/fm-branch-dispatch.mjs offer. - The host feeds the dialog mirror at the head of each attended wake and passes a close through unchanged when it is main-only, the engine or a tool is missing, the primary has no verified mirror, the main session cannot be identified, or the session is cooling down. - The drain presents captain outcomes first, one line per task, never behind older routine outcomes, and collapses routine overflow into a count that is marked read. - The return advances the store's read cursor through the away window once the brief has rendered, so the first drain does not replay it. - The branch prompt's mirror wording is host-neutral, and the rule to report what main must act on as captain, once per unchanged situation, applies only to the attended posture on the host. * docs: record the attended supervision host live check * no-mistakes(review): Present pre-window unread outcomes and contiguous captain prefix * no-mistakes(review): Return brief presents every row it marks read * no-mistakes(review): Return brief lists every unread outcome in one list * no-mistakes(review): Keep return list in store order and gate cursor failures * no-mistakes(review): Make the drain the only branch-outcome presenter after return * no-mistakes(review): Gate return on drain outcome failures; byte-count outcome budgets * no-mistakes(review): Gate drain on projection failures; UTF-8-safe byte cuts * no-mistakes(review): Fail drain without jq; hand unreadable prompt mirror to main * no-mistakes(document): Correct supervision-host return and drain documentation * no-mistakes(review): Recheck attended offer at turn start; honest failed-drain brief * no-mistakes(document): Correct supervision-host posture and drain documentation * no-mistakes(document): Documentation remains accurate for attended supervision
…unchenguid#5753) Each '# shellcheck source=' directive makes ShellCheck's external-source traversal expand that library's whole transitive graph again at the site. fm-pending-reply-lib carried three directed lazy sources of fm-wake-lib and two of fm-parent-channel-lib on identical per-call re-source sites, so one file analysis peaked above 4 GiB and every caller (fm-watch, fm-teardown) inherited the multiplier - the root cause of the PR kunchenguid#5732 Lint 1 OOM kill. Keep the runtime '.' commands byte-identical: the lazy re-source under 'local STATE FM_WAKE_QUEUE FM_WAKE_QUEUE_LOCK' is real behavior. Drop the duplicate directives so each library expands once per unit, and drop the tmux/classify directives since classify already arrives through the kept fm-wake-lib expansion and no tmux symbol is referenced here. The directive above the lib-dir assignment is kept - it binds the bin/ prefix so the undirected sites still resolve without SC1091. Measured peak RSS, ShellCheck 0.11.0 -x on Linux arm64: bin/fm-pending-reply-lib.sh 4.06 GiB -> 1.96 GiB, zero findings
kunchenguid#5773) * fix: split bash 5.2 sibling $() in recovery mint and delivery log Sibling command substitutions on one line can empty a recovery generation under bash 5.2 when a CHLD trap is set. Mint pid/epoch sequentially, refuse empty tokens before write, and clean delivery fields before printf. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: split bash 5.2 sibling $() in recovery mint and delivery log Sibling command substitutions on one line can empty a recovery generation under bash 5.2 when a CHLD trap is set. Mint pid/epoch sequentially, refuse empty tokens before write, and clean delivery fields before printf. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: keep recovery mint failure semantics after sibling $() split Remove the new pid/date refusal and grammar guard so a mint miss still yields a grammar-valid token and a durable wake row, matching accepted review intent. Drop the fake-failing-date case that locked in the refuse. Co-authored-by: Cursor <cursoragent@cursor.com> * no-mistakes(document): Point recovery-mint hazard comment at its regression test --------- Co-authored-by: Cursor <cursoragent@cursor.com>
…stop calling Escape safe there (kunchenguid#5791) Fixes kunchenguid#4520
…unchenguid#5790) Fixes kunchenguid#4756 The voice status reader in bin/fm_voice_records.py reports each worker's state from the last non-blank line of its status log. When a worker appends a status line and then a line of plain prose, the reader reported "note" with the prose line instead of the declared state, diverging from bin/fm-classify-lib.sh's shell scan. Scan back through the tail for the newest line whose prefix is a single lowercase verb-shaped word (letters and hyphens), and report that event's verb instead of always taking the last line. An unrecognised verb-shaped prefix still reports "note" rather than letting an earlier recognised line answer for it, and free text with no colon is skipped as prose. When the tail holds no such event, the last line is reported exactly as before.
…newer version (kunchenguid#5786) * fix(bin): stop reporting an already-installed version as an available update An update announcement named its version first ("current -> new"), so reading the first dotted number as the announced version compared the current version against itself and always looked newer. Read the last dotted number instead, and only report an available update when that announced version is newer than the newest installed copy found; when that version is already installed, report only PATH skew. Fixes kunchenguid#5151 * no-mistakes(document): docs: gate announce update-available report on newer-than-installed
…h their launch config (kunchenguid#5799) * fix(bin): pass the profile effort to OpenCode workers through their launch config The dispatch profile's effort axis was recorded in task metadata but never reached an OpenCode worker: the launch wrote only a permission grant into the config it constructs. OpenCode 1.18.32's config schema carries per-model reasoning effort as agent.<name>.variant, so the chosen effort is now merged into the same OPENCODE_CONFIG_CONTENT JSON as the default build agent's variant, keyed to the resolved model. With no effort chosen the launch stays byte-identical. Fixes kunchenguid#1373 * no-mistakes(review): gate OpenCode effort variant by model provider family * no-mistakes(document): docs(opencode): note provider-family gating for effort variant
…nchenguid#5815) * fix: preserve cancellation as no verdict in crew state Reuse the green-delivery safeguard for cancelled CI monitors and permit a skipped rebase. Other cancelled outcomes and coarse ledger records use the existing unknown state. Four delivered-PR regressions failed before the fix and pass afterward. The isolated public resolver and fleet-summary tests prove that undelivered cancellation no longer creates a failure contradiction, while preserving historical records and the terminal_in_flight invariant. Evidence uses fixture no-mistakes responses, not a live daemon cancellation. Update the existing coarse cancellation assertion from failed to unknown because it encoded this defect; retain its newest-run precedence check. Full fm-crew-state suite and pinned lint pass. * fix(review): Verify PR disposition before reclassifying terminal validation runs * fix(test): Add captured cancellation replay coverage for resolver and fleet * fix(document): Clarify cancellation and terminal delivery documentation
…henguid#5812) * fix: declare worker background and pipeline waits Require ship and scout workers to declare owned-work waits with the existing paused verb before ending a turn or waiting on a pipeline or long command. Keep the first-sight alert and existing liveness classification unchanged; subsequent inspection follows the existing long pause cadence. Validation: emitted brief regression failed before the instruction change and passes afterward. Public watcher/drain regressions cover the first alert, repeated wedge suppression, bounded rechecks, and undeclared idle alarms using isolated backend fixtures. Brief suite, pinned lint, Bash syntax, documentation inventory, and whitespace checks pass. No real worker harness was exercised for wait behavior. * fix(document): Clarify declared worker waits and documentation ownership * fix(ci): Captain, fixed the cadence test to age both the declaration and first-alert throttle while preserving declaration identity. Reproduced the CI failure using stable identity; corrected tests pass with both stable identity and native macOS behavior. Focused ShellCheck and diff checks pass. Production behavior is unchanged; Linux CI was not rerun locally
…5770) * fix(bin): bound each lint root in its own ShellCheck process CI job "Lint 1" died twice at about ten minutes because the two shard workers each packed about 110 canonical roots into one unbounded ShellCheck process, and a byte-weight rebalance moved the analysis-heavy fm-watch.sh into a partition with other heavy roots, so the pair outgrew the 16 GiB runner before anything could name a culprit. Run one canonical root per ShellCheck process under an enforced envelope: a wall deadline plus terminate-then-kill grace via the shared fm-timeout-lib.sh watchdog, and a per-root rlimit spec applied inside the child before exec (default a 4 GiB address-space cap, so two workers stay inside a 16 GiB job with headroom). A root that exceeds the envelope fails by name with a recorded reason - timeout, memory, signal, or limit-unavailable - instead of taking the runner down. The per-root watchdog runs in its own process group so the owner's group sweep cannot orphan the bounded subtree, and fm_exec_timed now starts the same escalation when its parent dies before it can be signalled. FM_LINT_REQUIRE_BOUNDS=1, set in CI, refuses the run outright when a configured bound cannot be enforced on the host rather than lint uncapped. Each root's begin/end, reason, duration, and peak RSS stream to stderr in partition mode and append to a retained <telemetry>.roots.tsv sidecar uploaded beside the partition telemetry. Coverage is unchanged: pinned ShellCheck 0.11.0, --norc, --external-sources full analysis, complete and disjoint partition inventory, workflow lint, and the backend-purity check, with byte-identical diagnostics across jobs=1/2 proven by tests/fm-lint.test.sh. * fix(bin): fail closed on unenforceable lint bounds and size the cap Required-bounds mode (FM_LINT_REQUIRE_BOUNDS=1, set by CI) now refuses the run with named errors before any root starts: a missing fm-timeout-lib.sh, a watchdog that cannot actually bound a probe command, or a host that rejects the address-space limit all stop the run rather than lint uncapped. The generalized FM_LINT_ROOT_RLIMITS flag:value interface is replaced by a single FM_LINT_ROOT_MEMORY_KIB, and the roots sidecar and telemetry record the run's final exit status after backend-purity and workflow checks instead of the pre-check lint status. The default cap is 6 GiB of address space per root, not 4 GiB: ulimit -v bounds virtual address space rather than resident memory, and ShellCheck's GHC runtime keeps roughly a third of that space as reservation, so 6 GiB yields about a 4 GiB working heap budget. A Linux measurement during this change showed eleven real canonical roots running out of memory under the earlier 4 GiB cap while the largest passing root peaked near 2.8 GiB resident; two 6 GiB roots plus runner overhead still fit the 16 GiB job. Roots that still exceed the cap keep failing by name, and the sidecar's per-root peak RSS keeps roots approaching the budget visible. tests/fm-lint.test.sh now proves the memory primitive where it can be proven: on hosts that accept ulimit -v a perl allocator is refused under a 256 MiB limit and reported by name as a memory death, the pinned ShellCheck lints a small file under the configured cap and is named when a far smaller cap binds it, and a watchdog-less copy refuses under REQUIRE_BOUNDS; the bounded cases skip on macOS, which cannot enforce the address-space limit. * no-mistakes(review): Prove memory cap binds, pass watchdog owner, drop unused modes * no-mistakes(review): Capture watchdog owner before startup for every fm_exec_timed caller * no-mistakes(document): Clarify bounded lint documentation and telemetry * no-mistakes(document): Correct bounded lint documentation and sidecar path * docs(bin): restore the per-root memory cap sizing rationale The pipeline's document step rewrote the ROOT_MEMORY_KIB comment and dropped the sizing reasoning the change is required to record: address space vs resident memory, the GHC reservation share, the measured 4 GiB failures and ~2.8 GiB peak, and the two-roots-plus-runner capacity arithmetic. Restore it beside the default while keeping the corrected "not a resident-memory ceiling" framing. * no-mistakes(review): Document memory cap RSS reduction threshold and first candidate * no-mistakes(review): Scope owner-death escalation docs to the perl watchdog * no-mistakes(document): Clarify bounded lint and timeout documentation * no-mistakes(review): Install perl watchdog signal handlers before forking the command * no-mistakes(document): Correct bounded lint documentation and stale watcher comments * no-mistakes(ci): Fixed the supervision-host test’s obsolete expectation: the watchdog now reaps an engine when its host dies. The timeout and supervision-host tests pass locally; the watcher test also passes locally. Lint 1 and 2 remain unresolved: seven canonical roots exceeded the required 6 GiB address-space cap in CI. I did not raise the cap, exempt roots, or reduce source-following coverage to make those failures disappear * no-mistakes(ci): The two lint checks failed when eight canonical roots hit the enforced memory cap. I reduced repeated ShellCheck source-graph expansion while keeping runtime imports and the canonical root inventory intact. Pinned ShellCheck passes for all changed roots; the relevant local tests pass. The 6 GiB Linux CI run remains unverified * no-mistakes(review): Restore source directives, raise cap to 8 GiB, classify OOM * no-mistakes(review): Classify memory deaths from root stderr, not source excerpts * no-mistakes(review): Match only whole runtime memory-error lines for memory reason * no-mistakes(document): Clarify lint memory classification in script documentation * no-mistakes(ci): Fixed both lint checks’ memory-limit failure: each CI lint job now runs one root at a time with a 12 GiB address-space cap. Kept the local two-worker default and updated the sizing comment and test expectation. The lint tests and workflow validation pass locally; Linux CI remains to confirm the heavy roots
…uid#5732) * fix(bin): bound the watcher cleanup marker-lock wait tests/fm-watch-triage.test.sh intermittently failed serial CI shard 1 with "watcher pid <pid> did not exit within 10s of TERM". The watcher had processed the TERM and was inside watcher_cleanup, where the recovery-marker publish waits on state/.watcher-down.lock through an unbounded fm_lock_acquire_wait. A live foreign holder of that lock leaves the TERM'd watcher spinning in its own EXIT trap until the lock frees or a second signal short-circuits the trap. fm_recovery_transition now takes an optional bound and both release-lock paths plus publish honour it through a new in-process fm_lock_acquire_wait_max. watcher_cleanup passes FM_WATCHER_CLEANUP_LOCK_BOUND (default 2s); on timeout the publish is skipped, the singleton stays behind as ordinary dead-pid evidence, and the next arm's clear-stale-lock still republishes it. Regression test drives a real watcher with .watcher-down.lock held by a live foreign process and asserts a single TERM still stops it. * no-mistakes(review): Parse watcher cleanup lock bound as decimal, zero defaults * no-mistakes(review): Pin cleanup bound tests to observed marker-lock contention * no-mistakes(document): Document bounded watcher cleanup and recovery * no-mistakes(document): Clarify bounded watcher cleanup and recovery documentation * no-mistakes: apply agent fixes * no-mistakes(review): Arm marker-lock FIFO before TERM; drop FM_TEST_ONLY_LATE * no-mistakes(review): Hold marker lock through a failed cleanup acquire
…kunchenguid#5845) * test: stop the leaked unreachable watcher before remote e2e cleanup The remote secondmate lifecycle e2e backgrounded fm-watch.sh through the remote_env shell function, so $! named the function's subshell rather than the watcher. Killing that subshell left the unreachable-leg watcher running, and its one-second liveness probe kept invoking the fake ssh, which rewrites ssh.count in the temp root. When a probe landed while the EXIT trap was removing the root, rm failed with "Directory not empty" after every assertion had passed. Exec the watcher from the backgrounded function so the recorded pid is the watcher itself, and assert the stopped watcher stops probing and writing its state. Cleanup also stops a watcher left running by a failed assertion and removes the root through fm_test_remove_tree, so a run that fails before retirement does not strand the read-only spawn hooks directory. Closes kunchenguid#5836 * no-mistakes(review): Clear reaped watcher PIDs and restore plain temp-root removal * no-mistakes(review): Let in-flight probe settle before stopped-watcher baseline
…nchenguid#4806) * fix: stop quarantining ordinary shared-captain source updates * no-mistakes(document): Rewrap remote inherit header so usage prints fully
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
can you check our agents.md and claude.md files. aren't they outdated?
Context: this fork (peterOC26/firstmate) last took the original project's updates (kunchenguid/firstmate) on 2026-09-20 in #20, which brought in upstream changes and kept this fork's fleet-board overlays. Since then upstream has changed AGENTS.md about 30 times, together with the scripts, docs and skills it describes (for example: project AGENTS.md edits limited to factual corrections, merges refused while a required check has not reported, a ship done report refused when its named head exists only in the worker copy, AI co-author trailers stripped from worker commits, a per-project ship-branch prefix, a Devin worker adapter, an optional supervision host). CLAUDE.md is only an @AGENTS.md pointer. Updating AGENTS.md alone would describe scripts this fork does not have, so the update has to bring in the code with it.
The approved answer (Q80 = Y): start one job that brings in the latest updates from the original project, keeps this fork's own changes (the fleet board and the rest), and opens a PR; it is not merged without explicit approval.
What Changed
Risk Assessment
Testing
Focused CLI and host checks passed; manual prompt, project setup, Git commit, branch-prefix, and named Herdr-lab artifacts show real output. Fixture-backed board and ship-head tests passed, but no live cancelled run or worker was available. The merge test was blocked by
tasks-axi0.2.5, and the older Herdr smoke test assumed an empty workspace that Herdr 0.9 does not provide. No screenshot was captured because the board is a terminal-text surface and this phase could not create a real cancelled run. The worktree is clean and the lab was torn down.bash tests/fm-ensure-agents-md.test.shbash tests/fm-branch-supervision.test.shbash tests/fm-supervision-host.test.shbash tests/fm-git-strip-ai-trailers.test.shfix/and an unconfigured project returnsfm/.bash tests/fm-task-delivery.test.shtasks-axiis 0.2.5 and the merge guard requires a compatible version. Provide version 0.2.6 or newer on PATH within the permitted test environment.devinCLI is not on PATH. Provide a repository-local executable path and isolated credentials to run a real worker.firstmateworkspace, while the unchanged smoke test expects the adapter to create it. A Herdr 0.9-compatible test setup and an isolated worker account are needed. The named…Evidence: Generated AGENTS.md and CLAUDE.md pointer
Source: Generated AGENTS.md and CLAUDE.md pointer
Evidence: Pi and supervision-host prompt output
Source: Pi and supervision-host prompt output
Evidence: Committed Git object after trailer stripping
Source: Committed Git object after trailer stripping
Evidence: Registered ship-branch prefixes
Source: Registered ship-branch prefixes
Evidence: Named Herdr lab output
Source: Named Herdr lab output
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
🔧 **Review** - 3 issues found → auto-fixed ✅
bin/fm-branch-prompt.sh:77- The fork's continuation handoff (fix: route pending supervision continuations to MAIN #21) does not reach the supervision host that this merge brings in from upstream. The prompt says (lines 13-16) that one text serves both the Pi branch and the supervision host, and fm-supervision-host.sh:619 feeds it to the headless engine. Lines 76-79 now tell that engine to 'set fm_branch_report.continuation', and line 78 says an attended branch 'must use it when continuation needs a new worker'. The host's report surface does not support this. bin/fm-branch-report.sh:61-72 rejects--continuationthrough its*) usagebranch with exit 2, so nothing is recorded. bin/fm-branch-report.sh:117 never forwards a continuation to fm-branch-outcome.sh append. bin/fm-wake-drain.sh:611-614 builds each captain BRANCH OUTCOMES line from seq, task and summary only, so a stored.continuationis never shown to MAIN off Pi. Concrete case: a home with config/supervision-host and a Claude primary gets a wake for a completed stage that needs a new worker. The engine follows the prompt and runsfm-branch-report.sh --continuation ..., which fails with a usage error. Either the wake goes back to MAIN with no outcome recorded, or the engine drops the handoff and reports a plain captain summary. Either way MAIN never receives the structured next-action/authority handoff that the fork's fix exists to guarantee. Choosing the remedy needs a product decision, which is why this is ask-user: (a) extend the continuation to the host path by adding--continuationto fm-branch-report.sh and showing it in the drain's captain lines, or (b) make the prompt paragraph Pi-specific and give the host its own routing wording.bin/fm-bearings-snapshot.sh:117- Upstream 9d56cf6, brought in by this merge, changed bin/fm-crew-state.sh:1076/1099/1132 so a cancelled run now reports stateunknownwith the detail 'run cancelled: no verdict'. The fleet board overlay still assumes the oldcancelledstate. The header comment at line 117 says a failed/cancelled run-step state 'reads as a validation park'.validation_parkat line 458 andunder_way_detailat line 465 still compare.state == "cancelled", which can no longer happen. The merge rewrote the overlay test (tests/fm-bearings-snapshot.test.sh, cancelled case) to expect 'current state unclear' instead of 'parked after validation stop'. So the board's wording for a cancelled validation run has quietly changed while the comment still describes the old behavior. Remedy: drop the deadcancelledcomparisons at 458/465 and update the comment at 117-120 to say a cancelled run has no verdict and shows as unclear.docs/captain-hold-lifecycle.md:477- When the fork's overlay was applied, it replaced "Bearings' Captain's Call" and 'Recently Landed' in the 'It proves' paragraph with 'decision projection' and 'landed projection' to match the six-column board. Upstream then split that paragraph into bullets, and the merge kept upstream's wording. Line 477 ('visible in Bearings' Captain's Call'), line 483 ('without losing the delivery from Recently Landed') and line 490 ('Recently Landed publishes it') now name chat sections that the fork's six-column Bearings digest no longer has. Remedy: reapply the overlay's wording to these three bullets.🔧 Fix applied.
✅ Re-checked - no issues remain.
✅ **Test** - passed
✅ No issues found.
bash tests/fm-ensure-agents-md.test.shbash tests/fm-branch-supervision.test.shbash tests/fm-supervision-host.test.shbash tests/fm-git-strip-ai-trailers.test.shfix/and an unconfigured project returnsfm/.bash tests/fm-task-delivery.test.shtasks-axiis 0.2.5 and the merge guard requires a compatible version. Provide version 0.2.6 or newer on PATH within the permitted test environment.devinCLI is not on PATH. Provide a repository-local executable path and isolated credentials to run a real worker.firstmateworkspace, while the unchanged smoke test expects the adapter to create it. A Herdr 0.9-compatible test setup and an isolated worker account are needed. The named…bash tests/fm-ensure-agents-md.test.shbash tests/fm-branch-supervision.test.shbash tests/fm-supervision-host.test.shbash tests/fm-bearings-snapshot.test.shbash tests/fm-git-strip-ai-trailers.test.shbash tests/fm-task-delivery.test.shbash tests/fm-dod-lib.test.shbash tests/fm-crew-state.test.shbash tests/fm-devin-harness.test.shbash tests/fm-pr-merge.test.sh(blocked bytasks-axi0.2.5)bash tests/fm-backend-herdr-smoke.test.sh(failed its Herdr 0.9 workspace assumption)Manualfm-ensure-agents-md.shcreation and repeat run in a disposable projectManualfm-branch-prompt.shPi/host output andfm-branch-report.sh --continuationrefusalManual Git commit throughfm-git-strip-ai-trailers.shand inspection of the committed objectManualfm-project-mode.sh --branch-prefixqueriesbin/fm-herdr-lab.shnamed provision, workspace creation, teardown, and session-list verification✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.