Skip to content

Reconcile upstream Firstmate at 29213a09 - #77

Merged
timbarreto merged 62 commits into
mainfrom
reconcile/upstream-2026-09-28-29213a09-9bf222f0
Sep 28, 2026
Merged

timbarreto merged 62 commits into
mainfrom
reconcile/upstream-2026-09-28-29213a09-9bf222f0

Conversation

@timbarreto

@timbarreto timbarreto commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Reconciliation status

All 25 automatic checks passed for head 221f6c60d40c47414e0e57ff0acaea4bae4a0b10.
The PR was landed with a merge commit, preserving canonical upstream ancestry.
GitHub records the merge at September 28, 2026, 20:25:59 UTC, after both CI workflows completed successfully.
This reconciliation session did not merge the PR, enable auto-merge, or change repository settings.

This is a normal, two-parent merge performed entirely in a separate worktree.
The original local main remains clean at 22517d81ede215c5946a054a08a4ff1063ba0c6c, with its index byte-unchanged.
No live Firstmate supervisor, worker fleet, or credentialed vendor session was started.

Frozen identities and prior synchronization

Identity Commit
Frozen fork and actual PR base, verified against GitHub 22517d81ede215c5946a054a08a4ff1063ba0c6c
Proven prior canonical upstream b805823a6beef7f59b2a10fe50ae6054c39cf571
Frozen canonical upstream 29213a09322a35d87b2f8acc29c415f3f860cb64
Two-parent reconciliation merge b82db25db5bfb0d4aaef0731c4a57a895161ef65
Current PR head 221f6c60d40c47414e0e57ff0acaea4bae4a0b10
Externally observed PR landing merge 011c291f47942f7d054d059da9d8c78f5a7e6f6f

PR #76 records the previous synchronization.
Its head, 85c1d779065e7447727aecf5af2c15a1870444e5, carries the prior upstream trailer and is reachable from the frozen fork.
The prior upstream object is an ancestor of both frozen inputs and is their exact merge base.
PR #76's merge commit is the frozen fork SHA above.
Canonical upstream was fetched once and remained frozen throughout this run.

The reconciliation commit's first parent is the frozen fork and its second parent is the frozen upstream.
The committed tree was checked against the reviewed index, and Git's trailer formatter returned the exact upstream SHA.
No reconstruction or ancestry-only anchor was needed.
The current head adds three separate CI-evidenced follow-up commits; the two-parent merge remains unchanged and reachable.
The PR landing merge has the frozen fork and the verified PR head as its two parents, so it also retains the canonical upstream ancestry.

Conflict and integration decisions

All 35 conflicted files were resolved by composing intent rather than selecting whole-file ours/theirs.
The reviewed scope is the original 208-path merge manifest plus eight explicit current-owner/test/CI additions, for 216 reviewed paths and 215 changed paths.
The seventh addition is the existing Fork CI workflow, which runs native Git-hook and imported Claude-hook regressions in its existing Windows management lane without adding a check producer.
The eighth addition is tests/fm-copilot-harness.test.sh, extending the existing imported-hook regression to upstream's mirror hooks.
The reviewed incoming change to tests/fm-gotmp.test.sh was withdrawn because its new symlinks duplicate the fork's existing fixture installer; that file intentionally equals the frozen fork and is absent from the final base diff.

Concern Resolution and retained owner
Fork-free helpers Adopt upstream's component and epoch helpers and cheaper source-directory discovery. Keep fm-path-lib.sh's Windows spelling/normalization boundary, the absolute-lock-path fast path, and explicit module-load failures.
Stale locks Keep upstream's elected tombstone reaper so a stale link reaper cannot delete a successor. Compose it with bounded, leaf-first recovery of legacy directory chains and retain the Windows fm-lock-fast.pl path. No new recursive .steal.steal acquisition.
Startup Adopt per-endpoint bounds and truthful abnormal-child-exit banners. Retain streamed diagnostics, per-stage timings, and the distinction between a current-stage breadcrumb and a measured bottleneck. Do not prescribe widening a deadline as the default repair.
Launch and relaunch Preserve guarded missing-endpoint recovery, native command transport, queue-lock cleanup, and lifecycle-owned publication. Import Claude's record-backed brief/add-directory handling, attribution opt-out and hook lifecycle, OpenCode's model-scoped variant, and fresh-launch cwd proof. Put Pi's runtime-reference placeholder and resume policy in bin/harnesses/pi.sh; keep Copilot's no-resume policy explicit and signed Pi on its legacy path.
Teardown Preserve native close confirmation and the fork's close-before-worktree-return transaction. Add upstream's seen-marker cleanup and retirement of an exact-pane journal or a version-1 journal whose token-bearing workspace is proven gone. Ambiguous or differently bound journals remain.
PR and contribution publication Refuse persistent-secondmate PR watches before remote observation. Preserve Azure identity/source-head handling and secure-before-write native privacy. Compose contribution-budget propagation with private atomic shim publication.
Classification and cost Keep the extracted, batched fresh-file/status readers and existing native management cost boundaries rather than restoring older inline implementations. The hash-pinned remote-doctor compatibility path was not migrated.
Test metadata and fixtures Transplant six upstream family registrations into tests/catalog/core.tsv, not the obsolete inline runner table. Keep the 1500-second generic changed-suite policy and measured Windows exceptions. Add explicit Python guard/voice-record routes and catalog regressions. Preserve named-case registries, the cancellation group's isolated aggregate-failure behavior, and copied catalog/path dependency closure.
Instructions and docs Adopt the upstream instruction/skill reorganization and operational carrier clarification without dropping Copilot, ownership safeguards, native teardown guarantees, or the fork's diagnostic guidance.

Relevant upstream intent includes kunchenguid#5889, kunchenguid#5728, kunchenguid#5917, kunchenguid#5859, and kunchenguid#5696.

Local evidence and explicit limitations

Repository gates and focused checks used the existing skills/reconcile-firstmate-upstream/scripts/validate-local.mjs controller and the recorded Git-for-Windows Bash executable, with process-local FM_LIVE=0.
These local results predate all CI follow-ups; no local lint, coverage, case inventory, or test execution resumed after the circuit breaker.
Caching was disabled.
Both invocations shared the 2400-second budget and two-timeout circuit breaker; the second received only the remaining budget and one remaining timeout.
Measured controller time was 909.155 seconds across the two invocations, plus Git/probe overhead for which 60 seconds was reserved.
Execution stopped on the second timeout, not because the budget was reset or a production deadline was widened.

Executed command/check Result Seconds
Shell syntax sweep over bin/fm-lint.sh --list-files First pass found an auto-merged executable loop inside the crew-state case array 20.595
bash bin/fm-lint.sh Failed on that same syntax defect; workflow lint reported all three YAML files valid 342.523
bash bin/fm-doc-audience-check.sh Passed: 124 surfaces, 758 local links 8.958
bash bin/fm-test-run.sh --check-coverage Timed out at its local 120-second allowance; no passing baseline claimed 121.830
bash bin/fm-test-run.sh --list --changed --base 22517d81ede215c5946a054a08a4ff1063ba0c6c First attempt exposed the Python guard mapping; follow-up exposed the voice-record mapping 28.962 / 21.218
First case-listing loop Listed AFK return, Herdr, control, and relaunch; stopped at the Cursor fixture's missing POSIX C compiler 41.782
Repaired shell syntax sweep Passed before the final catalog/fixture follow-ups; final-head verification belongs to CI 53.188
Follow-up case-listing loop Listed crew-state and inactive reconciliation, then reached Kimi's top-level fixture and its Windows private-hook permission failure 29.134
bash bin/fm-test-run.sh --jobs 1 tests/fm-harness-contract.test.sh Ten success messages preceded the 180-second suite timeout; not a suite pass 184.633

The syntax sweep used:

set -e
scripts=$(bin/fm-lint.sh --list-files)
while IFS= read -r script; do /bin/bash -n "$script" || exit; done <<< "$scripts"

The case-listing commands set FM_TEST_LIST_CASES=1 for each script.
The first loop named afk-return backend-herdr control control-relaunch cursor-primary inactive-reconcile kimi-harness lint pi-branch-extension pr-check-security spawn-dispatch-profile supervision-host teardown test-run watcher-lock session-start.
The follow-up named crew-state inactive-reconcile kimi-harness lint pi-branch-extension pr-check-security spawn-dispatch-profile supervision-host teardown test-run watcher-lock session-start.
Their exact Bash command bodies were:

set -e; for suite in afk-return backend-herdr control control-relaunch cursor-primary inactive-reconcile kimi-harness lint pi-branch-extension pr-check-security spawn-dispatch-profile supervision-host teardown test-run watcher-lock session-start; do FM_TEST_LIST_CASES=1 /bin/bash "tests/fm-$suite.test.sh" >/dev/null; printf 'registered: %s\n' "$suite"; done

set -e; for suite in crew-state inactive-reconcile kimi-harness lint pi-branch-extension pr-check-security spawn-dispatch-profile supervision-host teardown test-run watcher-lock session-start; do printf 'CASE_REGISTRY %s\n' "$suite"; FM_TEST_LIST_CASES=1 /bin/bash "tests/fm-$suite.test.sh"; done

These are partial observations, not completed inventories.
Kimi is not a registered case-listing suite, so that attempt executed its existing top-level fixture; no listing-only or passing-suite claim is made.
Its failing producer is unchanged from the frozen fork, but no matching baseline execution proved Windows compatibility, and the observation remains unresolved.

Successful prerequisite probes identified Bash 5.3.15, Git 2.55.0.windows.5, ShellCheck 0.11.0, actionlint 1.7.12, GNU Awk 5.4.1, Node 24.19.0, and jq 1.8.2.
No cc, gcc, or clang was found on the host PATH.
Signing, hooks, and safe.bareRepository safeguards were not disabled.
No frozen-upstream differential was run, and no interrupted/deferred result was reused as a pass.

Source follow-ups fixed both observed Python mappings, the cancellation registry composition, the new timeout-test fixture's catalog installation, and a copied wake fixture's missing path module.
The fixture repair is not asserted to prove the cause of the pilot timeout.
The final changed selection and final lint were not rerun after the circuit breaker.
The changed-selection count is therefore unconfirmed locally rather than fabricated from a different inventory.

CI findings and isolated follow-up

The initial PR head, b82db25db5bfb0d4aaef0731c4a57a895161ef65, started both automatic workflows.
Its completed job logs and timing artifacts exposed the following integration defects.
Commit f23f9fb90fc975ccda1e0326f910fa86b49e68a4 made the source changes below.
Its completed matrix recorded 22 successful checks and 3 failed checks; successes from that head are not reused as proof of the current head.

Initial-head evidence Follow-up
Windows reconciliation (copilot-launch), also portable serial 7: the shared registry rejected undefined test_opencode_threads_model_and_ignores_effort_axis before the selected case ran Register upstream's renamed OpenCode effort-variant case once at the original position; retain all of its behavior assertions.
Portable parallel 1: the lint suite's incoming exclusion-completeness registration had no function body Restore the complete helper and regression from the frozen upstream, rather than dropping the registered coverage.
Portable parallel 1: the runner still asserted a 900-second generic timeout although the observed policy was 1500 Update the remaining generic/non-Windows expectations to 1500; retain the 1800/4500/7200 Windows exceptions unchanged.
Stock macOS Bash: the imported differential reference canonicalized an already-absolute doubled-separator lock path Preserve the frozen fork's absolute-spelling/no-canonicalization contract in that reference, including nonexistent-parent and dot-component cases. Relative components still compare against the command implementation.
Windows Copilot management: the new hook installer passed a raw MSYS worktree path containing glob characters to native Git Use the existing fm_path_native_argument owner for the installer argument and launch's Git hooksPath value. Add a native commit-object/previous-hook regression to the existing management lane; no stripping, repository-hook, or privacy check is bypassed.
Portable parallel 2: the capped-inventory regression still expected cancellation alone to mean failed Expect upstream's explicit unknown / run cancelled: no verdict result, retaining the read-only and single-read guards.
Portable serial 7: an incoming AFK return fixture could not load the fork's extracted harness detector and therefore did not distinguish Pi Install the existing session-lock and harness/process dependency closure in the fixture, preserving the Pi/non-Pi assertions.

The Windows management log also showed setup diagnostics from the unchanged control-inspection fixture before that case reported success.
That observation is not treated as a proven baseline incompatibility or as complete path-fixture verification.
The local Kimi observation and interrupted local inventories remain disclosed above.
Full logs and downloaded timing artifacts remain in session evidence, not in the repository.

Remaining failures and second follow-up

Shared CI run 36467573440 failed on f23f9fb90fc975ccda1e0326f910fa86b49e68a4; Fork CI run 36467573425 succeeded.
The three failed Linux shards contained four failing assertions.
Commit 70751eca3618d536618a097bded6ae451dbf6f07 addresses those specific failures:

Failed-head evidence Second follow-up
Portable serial 4: imported Claude hooks attempted to run the missing fm-host-mirror.sh and exited 127 Give both mirror-hook registrations the existing Bash true and PowerShell exit 0 Copilot overrides, while keeping their Claude commands active. Extend the fixture and all import counts to eight hooks, and run the case in the existing Windows management job.
Portable serial 6: the bounded endpoint-read assertion counted one hanging fixture Herdr process Exercise the same bounded digest read through --reemit, without separately owned startup sweeps querying that fake backend. Keep the 2-second configured bound, padded-zero fallback, and zero-leftover assertion unchanged.
Portable serial 6: the replacement-launch assertion still expected Claude's former inline brief encoding Update the two Claude-specific assertions to the record-backed operational-input carrier; retain the Copilot inline-encoding assertions and exactly-one-launch requirement.
Portable serial 8: fixture setup copied fm-path-lib.sh onto itself through a redundant symlink Remove the two incoming symlinks because the existing private-path fixture helper already installs that module. This returns the file to frozen-fork content.

The endpoint isolation rationale is supported by the startup call graph, but the original leftover process's parent/group was not captured.
The unchanged endpoint-hang and padded-zero cleanup assertions both passed at 70751eca; this does not conclusively identify the originally counted process or rule out every production timeout leak.
No production timeout, process cleanup, privacy safeguard, or hook-chaining policy was weakened, and the exhausted local validation allowance was not reset.

Final serial-shard follow-up

The completed 70751eca matrix had 24 successful checks and 1 failed check: Behavior portable serial 6.
Its log confirms both endpoint cleanup cases and both previously stale Claude carrier assertions passed.
The Windows management log separately confirms the eight imported hooks are silent no-ops in Bash and PowerShell while Claude's commands remain active.
Shared CI run 36473181147 then reached two later assertions previously hidden by those failures; Fork CI run 36473181171 succeeded.

Commit 221f6c60d40c47414e0e57ff0acaea4bae4a0b10 addresses the remaining composition:

Evidence Final follow-up
The abnormal-death output correctly reported exit 143 and the lock stage, but its test expected upstream's earlier wording Match the composed diagnostic and explicitly reject a false cumulative-deadline message, retaining the exit-status, missing-stage, completion-marker, and no-deadline-widening assertions.
Reviewing that same banner against the frozen fork and the existing runtime-bound regression exposed lost cumulative-deadline wording and its timing heading Restore those fork diagnostics only for exit 124; keep upstream's distinct abnormal-death banner for other nonzero exits and the shared distinction between a breadcrumb and a measured bottleneck.
Herdr recovery's copied code root contained bin/ but lacked the .agents/skills directory required by Claude's task-scoped grant Create that directory in the isolated transport fixture, as other copied-code fixtures do. Keep the real grant validation, private lock namespace, work preservation, and endpoint-recovery assertions unchanged.

The actual timeout implementation, configured deadlines, and process cleanup remain unchanged.
The current head's complete CI matrix passed; the prior head's successes remain historical evidence, not substituted results.

Deferred verification and owning CI

GitHub Actions owns the complete cross-platform matrix.
The resolved workflows still define 25 automatic checks: 18 shared CI checks and 7 Fork CI checks.
Workflow-qualified concurrency remains intact.
Catalog registration has not been used to grant concurrency: the existing runner and isolation-proof owner retain admission and shard composition.

Locally deferred or unresolved subject Existing producer
Final lint, including the crew-state/test-run/wake follow-ups CI: lint matrix, Lint 1 and Lint 2
Complete inventory partition and coverage timeout CI: test-coverage, Test coverage guard
Changed-timeout and cancellation regressions CI: tests-portable-parallel-1 / tests-portable-parallel-2
Python routes, imported family gates, and copied fixture closure CI: portable lanes selected by the existing runner/proof owner; catalog and harness-contract suites
Fork-free helper suite CI: portable serial matrix and Stock macOS Bash snapshot compatibility
test_lock_steal_reap_cannot_remove_successor and test_lock_stale_steal_hierarchy_converges_without_growing CI: tests-portable-serial
Native fast-lock suite Fork CI: Windows Copilot management
test_windows_paths_strip_and_chain_hooks and native/mixed/POSIX relaunch path cases Fork CI: Windows Copilot management
test_claude_settings_are_inert_for_copilot, including all eight Bash/PowerShell imports CI: portable serial 4; Fork CI: Windows Copilot management
test_bootstrap_diagnostics_stream_before_probe_completion Fork CI: Windows reconciliation (core)
Endpoint death/hang and abnormal-digest regressions CI: portable serial matrix, with POSIX process facilities
test_claude_launch_brief_publishes_record_doorbell and test_herdr_relaunch_resumes_only_the_registered_pi_session CI: portable serial matrix
test_teardown_retires_task_watcher_markers_and_orphan_journal and test_secondmate_record_refuses_a_pr_watch CI: portable serial matrix
Cursor's compiled process fixture and Kimi's unresolved Windows observation Linux portable behavior jobs; Windows Kimi parity remains an explicitly unproven surface
Real Herdr/Treehouse behavior CI: Behavior tests (Herdr), with its pinned tools and session tripwire
Native updater, Copilot launch/management, rollback, Azure PR completion, and path handling Fork CI: Windows self-update entry point, four Windows reconciliation subjects, Windows Copilot management
Pi/native declarations and package compatibility Fork CI: Harness package compatibility; shared portable package consumers
Stock Bash and repository structure CI: Stock macOS Bash snapshot compatibility and Repo invariants
Credentialed live harnesses and the Windows Herdr experiment Manual opt-in only; neither was run or counted as automatic coverage

Exact-head final CI result

At September 28, 2026, 20:26:44 UTC, head 221f6c60d40c47414e0e57ff0acaea4bae4a0b10 had 25 successful checks, 0 pending, 0 failed, and 0 absent.
CI run 36476513087 completed successfully at 20:25:39 UTC; Fork CI run 36476512989 completed successfully at 20:07:51 UTC.
Every expected check had exactly one GitHub Actions producer from its owning workflow, with no unexpected or duplicate producers, and every check matched the literal head SHA.
The serial-6 log directly records both endpoint cleanup cases, the abnormal-death banner, the cumulative-bound recovery, and the Herdr whole-task reclaim passing.
The complete fm-session-start and fm-control-relaunch suites exited 0 in 168683 ms and 172703 ms respectively, both with gate_skip=false.
This is the complete automatic-matrix result, not a claim that the separately gated live experiments or the disclosed local Windows Kimi observation were verified.

Expected automatic check Result Visible producers
Lint 1 Successful 1
Lint 2 Successful 1
Test coverage guard Successful 1
Behavior portable parallel 1 Successful 1
Behavior portable parallel 2 Successful 1
Behavior portable serial 1 Successful 1
Behavior portable serial 2 Successful 1
Behavior portable serial 3 Successful 1
Behavior portable serial 4 Successful 1
Behavior portable serial 5 Successful 1
Behavior portable serial 6 Successful 1
Behavior portable serial 7 Successful 1
Behavior portable serial 8 Successful 1
Behavior portable serial 9 Successful 1
Behavior tests (Herdr) Successful 1
Behavior timing aggregate Successful 1
Stock macOS Bash snapshot compatibility Successful 1
Repo invariants Successful 1
Windows self-update entry point Successful 1
Windows reconciliation (core) Successful 1
Windows reconciliation (copilot-launch) Successful 1
Windows reconciliation (legacy-rollback) Successful 1
Windows reconciliation (pr-completion) Successful 1
Windows Copilot management Successful 1
Harness package compatibility Successful 1

The following twelve planned local commands were all deferred by the circuit breaker.
They have no elapsed result because they did not start; the table above records their CI dispositions.

bash bin/fm-test-run.sh --jobs 1 tests/fm-fork-free-helpers.test.sh
FM_TEST_ONLY=test_lock_steal_reap_cannot_remove_successor bash bin/fm-test-run.sh --jobs 1 tests/fm-watcher-lock.test.sh
FM_TEST_ONLY=test_lock_stale_steal_hierarchy_converges_without_growing bash bin/fm-test-run.sh --jobs 1 tests/fm-watcher-lock.test.sh
bash bin/fm-test-run.sh --jobs 1 tests/fm-lock-fast.test.sh
FM_TEST_ONLY=test_bootstrap_diagnostics_stream_before_probe_completion bash bin/fm-test-run.sh --jobs 1 tests/fm-session-start.test.sh
FM_TEST_ONLY=test_claude_launch_brief_publishes_record_doorbell bash bin/fm-test-run.sh --jobs 1 tests/fm-spawn-dispatch-profile.test.sh
FM_TEST_ONLY=test_herdr_relaunch_resumes_only_the_registered_pi_session bash bin/fm-test-run.sh --jobs 1 tests/fm-control-relaunch.test.sh
FM_TEST_ONLY=test_teardown_retires_task_watcher_markers_and_orphan_journal bash bin/fm-test-run.sh --jobs 1 tests/fm-teardown.test.sh
FM_TEST_ONLY=test_secondmate_record_refuses_a_pr_watch bash bin/fm-test-run.sh --jobs 1 tests/fm-pr-check-security.test.sh
FM_TEST_ONLY=test_changed_bound_gives_slow_watcher_suites_headroom bash bin/fm-test-run.sh --jobs 1 tests/fm-test-run.test.sh
FM_TEST_ONLY=test_cancelled_run_groups bash bin/fm-test-run.sh --jobs 1 tests/fm-crew-state.test.sh
bash bin/fm-lint.sh tests/fm-crew-state.test.sh tests/fm-test-run.test.sh bin/fm-wake-lib.sh

No-renames divergence and locality audit

All comparisons use literal frozen identities and the committed head.
No broad line-ending normalization or unrelated worktree staging was performed.

Comparison Paths Added lines Deleted lines Binary paths
Prior upstream to frozen fork: retained fork before reconciliation 276 27470 5802 1
Frozen target upstream to frozen fork, before reconciliation 390 29163 23308 1
Frozen target upstream to reconciled head 279 27429 5618 1
Actual PR base/frozen fork to reconciled head 215 18040 2084 0

The smaller apparent divergence is not treated as behavioral proof.
Pi/Copilot facts remain in the closed adapter modules; signed Pi and other nonpilots retain their compatibility paths.
Native process/path/private-publication machinery remains until upstream provides equivalent Windows semantics and cost guarantees.
The catalog remains a metadata owner, not an execution/proof owner.
Legacy directory-chain recovery and the native fast-lock helper remain until equivalent recovery and process-cost behavior is established upstream.
Fresh classification/status observations and Azure/private publication remain at their existing owners rather than duplicated into callers.
docs/fork/architecture.md and docs/fork/verification.md retain the owner/removal conditions.

The frozen patches, NUL-delimited intended/candidate manifests, registry-composition records, validation plans/logs/ledger, and committed no-renames audit patches are retained outside the repository in the CLI session artifacts.
Final scope, clean branch, both merge parents, exact tree, upstream trailer, and original-main fingerprint were checked before publication.

Firstmate-Upstream-SHA: 29213a09322a35d87b2f8acc29c415f3f860cb64

tiago-peixoto and others added 30 commits September 25, 2026 11:21
…nchenguid#5695)

* fix(bin): strip AI co-author trailers from fleet-launched commits

Cursor and other non-Claude runtimes append the trailer after the typed
message. A per-task commit-msg hook removes it and leaves human co-authors
and the author identity untouched.

* no-mistakes(review): Export pane hooksPath override and drop generated-with stripping

* no-mistakes(ci): This PR caused all three CI failures, and the fix is test-only: 4 test files change, no product code. **Cause.** `fm-spawn.sh` now installs the AI-trailer strip hooks for every spawn, secondmates included. The installer refuses a worktree that is not a git repository, and the PR deliberately keeps that fail-closed rule because real secondmate homes are firstmate clones. Four test fixtures still gave secondmates a plain directory as their home, so each spawn failed with "not a git worktree ... could not install the AI-trailer strip hooks": - serial 5: `tests/fm-backlog-atomicity.test.sh` ("secondmate spawn failed"). - serial 8: `tests/fm-secondmate-harness.test.sh` ("split: no meta written"). - Herdr: `tests/fm-backend-herdr-launcher-workspace-e2e.test.sh` and `tests/fm-backend-herdr-workspace-per-home-e2e.test.sh`. **Rule that must hold.** Every home a test spawns as a secondmate must be a git worktree. I checked the other places in the changed area: the only secondmate spawns in these tests are the ones listed. The earlier rounds already fixed the other fixtures (`fm-secondmate-liveness`, `fm-secondmate-safety`) the same way. **Fix.** - Each of those four secondmate homes now gets the same `.gitignore` plus `git init -q -b main` that the liveness and safety tests already use. - The two Herdr tests clean up with their own plain `rm -rf "$TMP_ROOT"`, not the shared `tests/lib.sh` helper. Because the installer leaves each `state/<id>.git-hooks` directory read-only, that cleanup printed "Permission denied" and left the directories behind. Both cleanups now restore the owner's write bit on every directory before removing (`find ... -exec chmod u+rwx`), which is what `fm_test_remove_tree` in `tests/lib.sh` does. **Verification.** - `tests/fm-secondmate-harness.test.sh` passes. - `tests/fm-backlog-atomicity.test.sh` passes (99 ok, exit 0). - shellcheck is clean on all four files. - I could not run the two real-Herdr tests locally: the Herdr lab on this host refuses to start because it needs exactly one running default session, and I did not change the host's Herdr state to get around that. Instead I checked their two changed steps directly: the installer succeeds on a home set up the new way, and the new cleanup removes the read-only hooks directory completely. Those two tests will only be proven on CI
…id#5683)

* fix(bin): treat Pi's dollar-first cost footer as furniture

An idle Pi status row opening with $0.000 was read as a dead-shell prompt, so exit and relaunch refused on an empty composer.

* test: wait for the draining holder to exec sleep before reading its identity

The procevent drain fixture read fm_pid_identity immediately after
backgrounding setsid sleep, racing the child's exec chain. Mid-exec the
cmdline can read empty, failing the fixture on a loaded CI runner.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…5534)

* fix(bin): refuse a merge when a required check never reported

fm-pr-merge.sh built its GitHub refusals only from checks present in
statusCheckRollup, so a required check that never ran was simply absent and
the merge proceeded on the subset that reported, contradicting its own
"every required check green" claim.

The GitHub verify now reads the base branch's required contexts from the
forge itself - the classic branch protection summary on
GET repos/{o}/{r}/branches/{b} and the active ruleset rules on
GET repos/{o}/{r}/rules/branches/{b} - and refuses when a required context
has no entry in the same rollup, at the same head, that the merge is bound
to. Absence reads as unknown, never green. The required-set read joins the
existing refusal list, so a draft, a red check, and an unreported required
check are all reported together.

Could not read vs nothing required: both endpoints need only repository
read access. The admin-only GET .../branches/{b}/protection endpoint is
deliberately not used: it answers a non-admin token with the same 404 an
unprotected branch gets (observed live on kunchenguid/firstmate main with
this token), which would read a missing permission as "nothing required".
Any failed or malformed read of either source (auth, missing fine-grained
permission, rate limit, network, 404, unexpected shape) refuses the merge
with a line naming the unreadable source. The one exception is GitHub's
plan-gated 403 on the rules endpoint ("Upgrade to GitHub Pro or make this
repository public"), which already means "this repository has no branch
rules" for the merge-queue reader; that check moves into one shared helper
and the classic summary still decides for such a repository.

Attended waiver: --allow-missing <check-name> is the twin of --allow-red and
follows the same design and recording path: once, separate name argument,
waives only that exact unreported required check, still requires every
other required check reported and every check green, never waives an
unreadable required set, refused while the away-posture record exists, and
refused on GitLab. Merge-state BLOCKED policy is unchanged.

How this differs from the withdrawn kunchenguid#5353 (read from its diff):
- kunchenguid#5353 read the admin-only branches/{b}/protection endpoint and treated
  its 404 as "no required checks", so for any non-admin token the required
  set silently read as empty; this change reads the read-access branch
  summary and treats every failure as unreadable.
- kunchenguid#5353 ignored rulesets; this change also reads required_status_checks
  rules from the effective branch rules.
- kunchenguid#5353 made separate per-head REST reads of statuses and check-runs capped
  at per_page=100 with no pagination; this change checks presence in the
  same statusCheckRollup view the red-check gate already reads at the
  verified head.
- kunchenguid#5353 stopped at the first unreadable read; this change reports it as one
  refusal among all the others.
- kunchenguid#5353 also claimed kunchenguid#5345 (lock stealing) and changed 39 files, most
  unrelated; this change is kunchenguid#5344 only.

Live proof, read-only (a gh wrapper refused every merge and mutating call):
- cli/cli#14474 (trunk requires 3 classic build contexts, none ran):
  refused, naming build (macos-latest), build (ubuntu-latest),
  build (windows-latest); with --allow-missing "build (macos-latest)" it
  still refused, naming the other two.
- cli/cli#13665 (ran build (ubuntu-24.04-firewall) instead): refused,
  naming build (ubuntu-latest).
- hashicorp/terraform#39262 (ruleset-required checks absent): refused,
  naming Code Consistency Checks, End-to-end Tests, Race Tests, Unit Tests.
- cli/cli#14485 (all required reported and green): verified; the wrapper
  blocked the merge call and the pull request read back open.

Fixes kunchenguid#5344

* fix(review): Preserve required-check producers and aggregate independent read failures

* fix(document): Clarify required-check verification and waiver documentation

* fix(bin): match an app-bound required commit status by name

The producer-identity check resolved an app-bound required context only
against check runs, so a required context that the required app reports as
a commit status could never match and always read as "has not reported".
A commit status carries no app id to compare, so an app-bound requirement
that arrives as a status now matches by name, as before producer binding;
check runs keep requiring the configured producer app.

Live, read-only: hashicorp/terraform#39262 requires license/cla from
integration 865473, reported green as a commit status by the CLA app. The
previous head refused it as unreported; this head no longer does, while
still naming the four required check runs that never ran there.

Refs kunchenguid#5344

* fix(document): Clarify accepted commit-status producer verification limitation

---------

Co-authored-by: firstmate-oss <firstmate-oss@kunchenguid.local>
…uid#5696)

* fix(bin): never offer a persistent secondmate for teardown

The return brief's "Landed, cleanup due" scan listed every state/*.meta
record carrying a pr= and a merge-notified marker without regard to kind, so
a secondmate record holding a relayed child's merged PR put the mate itself
up for "bin/fm-teardown.sh <mate>" cleanup. A secondmate is a persistent
worker, never landed work.

- bin/fm-afk-return.sh: skip kind=secondmate in the landed-cleanup scan.
- bin/fm-pr-check.sh: refuse to record pr= or arm a merge watch on a
  kind=secondmate record before any side effect; a PR reported on its routed
  status channel belongs to a task in the mate's own home, which arms its
  own watch.
- bin/fm-watch.sh: a merged result from a poll already armed on a secondmate
  retires the poll silently - no merge outcome, marker, or wake.

* no-mistakes(document): Document secondmate merge-watch and return-brief exclusions

* no-mistakes(ci): CI failed in an unchanged watcher-shutdown test whose three-second wait was sensitive to runner load. Increased the wait for both state- and home-deletion cases without changing watcher behavior. The full fm-watch-arm suite passed locally; syntax and diff checks passed
…unchenguid#5702)

* fix(bin): refuse unknown dash-leading args in public-posting fm-x scripts

fm-x-reply.sh collected any unrecognized argument into the positional
pool and took the first one as the reply text, so an invocation like
"fm-x-reply.sh <id> --followup --final <text>" posted the literal
string "--final" to X and silently dropped the real text.

Make argument parsing strict in every script that can post publicly:
an unknown dash-leading argument, a dash-leading request_id/task id, a
dash-leading option value, or a surplus positional now exits 2 with a
usage error before any config load, outbox write, or network call. Reply
text starting with '-' is still accepted via --text-file or stdin, and
--help is honored wherever it appears instead of becoming text (a --help
forwarded through fm-x-followup.sh would have counted as a posted
follow-up and mutated the link).

fm-x-link.sh and the fm-public-followup scripts already refuse unknown
arguments; fm-x-poll.sh takes none.

* no-mistakes(review): Refuse surplus follow-up text sources; drop post-ID help branches

* no-mistakes(document): Clarify reply and follow-up argument usage

* no-mistakes(review): Refuse dash-leading --text-file operands in fm-x-reply

* no-mistakes(document): Correct follow-up argument parsing comment

* no-mistakes(document): Document dismiss argument rejection in script header
…nchenguid#5701)

* feat(bin): latch the supervision host after repeated engine errors

Rung 3c-1 of the PR 5631 re-cut: the host copies the Pi branch's
broken-session policy. Two consecutive engine errors latch the session;
every away wake then reaches main with one supervision-host line for a
five-minute cooldown, after which one wake probes the engine, and each
failed probe doubles the cooldown up to one hour. A reported turn without
an engine error clears it. The latch is kept per main session, engine, and
model in state/.supervision-host-health, and the engine conversation now
uses the same main-session key, which includes the lock holder's process
identity so a recycled pid never shares either.

Lifted from the validated 5631 tree and adapted to main's away-only host:
the attended recovery line and attended cooldown pass-through are left for
the attended core, so a recovery is only logged.

* no-mistakes(document): Consolidate supervision-host latch documentation
…kunchenguid#5535)

* fix(bin): absorb routine second-mate progress while surfacing routed replies

Fixes kunchenguid#2959

A kind=secondmate task's status signal was never absorbable, so a healthy
mate's routine working: and paused: appends woke the primary every time.
signal_crew_provably_working now reads the mate's lines new since the
watcher's classified position: a decision, blocker, terminal outcome, note:,
correlation-marked line, or unknown verb still surfaces regardless of busy
evidence, while unmarked working:, paused:, and resolved: fall through to the
same provably-working absorb an ordinary crewmate gets.

* no-mistakes(review): narrow secondmate routine absorb to working and paused
* fix(bin): use gh-axi for the ship DoD draft check

* no-mistakes(review): use PR number not URL in gh-axi draft check
… only (kunchenguid#5520)

* fix(bin): drop status prose from the inactive-outcome dedupe identity

The inactive-outcome receipt fingerprint included the child's sanitized
last status line, so a persistent child appending routine prose after one
terminal outcome minted a fresh parent event per sentence. Bind the
identity to incarnation, task id, terminal state, and PR only, keeping
the last line in the record as status_head evidence.

Fixes kunchenguid#2960

* no-mistakes(document): note structured-only inactive receipt identity in regression coverage
…unchenguid#5707)

* feat(bin): record the supervision host's dialog mirror on Claude and Cursor

Add bin/fm-host-mirror.sh, the one owner of the supervision host's dialog
mirror file, cursor, lock, and feed, plus the main-session key it keys
entries to. The tracked Claude UserPromptSubmit and Stop hooks and the
Cursor beforeSubmitPrompt and afterAgentResponse hooks record the captain's
prompt and main's reply, only on a home with config/supervision-host, from
a genuine primary checkout, for the lock-owning session. The mirror lands
inert: writers record and nothing reads it yet; attended supervision on the
host is the later step that consumes the feed.

Codex, Grok, OpenCode, and omp have no writer here.

* no-mistakes(review): Scope mirror dedup to session, atomic appends, marker-inclusive caps

* no-mistakes(document): Clarify dialog mirror scope and remove duplicate contract details

* no-mistakes(document): Correct Cursor hook documentation for dialog mirror registration

* no-mistakes(review): Pass mirrored dialog text to jq via stdin

* no-mistakes(document): Clarify dialog mirror documentation and remove duplicate claims

* no-mistakes(review): Preserve internal dialog whitespace; drop mirror check and verified modes

* no-mistakes(review): Drop only identical mirror repeats; remove redundant chmod guard
… escalations are not repeated (kunchenguid#5731)

* fix(bin): retire check-row receipts on branch acks and report an unchanged situation once

* fix(bin): scope a branch acknowledgement's check-row receipt retirement to
  its granted sequences

The away posture lifts the attended partition's check/decision exclusions, so
a branch grant can name check-kind rows - but the branch-actor ack still
assumed check rows were main-only and skipped every receipt scan. The queue
row was consumed while its terminal-outcome .pending receipt stayed behind,
and each inactive-reconcile cadence scan re-queued the same fingerprint. In
the first real away window on the supervision host that re-escalated one
unchanged held-PR situation on every cycle (~1,734 of 4,149 outcomes).

A branch ack now scans inactive-outcome and inactive-reconcile receipts and
commits secondmate stall receipts against exactly the sequences in its
eligible-row snapshot - the same rows it consumes - instead of none. Attended
grants still name no check row, so the scans find nothing.

* fix(bin): store a repeated captain verdict as routine while the task's
  durable situation is provably unchanged

fm-branch-outcome.sh append computes a mechanical situation key per captain
row - metadata bytes, captured status-log endpoint and identity, live
crew-state verb, worktree head - and anchors it in
state/.<task>.branch-captain-key. A later captain verdict whose recomputed key
matches is stored as routine with "unchanged since seq <N>:" prefixed to its
summary, so one situation escalates once until something provably changes. A
task with no readable status ledger is never demoted, an unreadable record
fails toward reporting, and teardown removes the sidecar with the task's
other branch records. The append-only store schema is unchanged.

This covers both hosts: the Pi supervision branch and the supervision host
both funnel reports through append.

* docs: check rows are main-owned only while attended; the away posture grants
  them to the branch, whose ack retires their receipts exactly

* test: the away-flood reproduction as a regression test (branch ack retires
  the receipt and later scans stay quiet), store-level dedupe coverage, and a
  branch-ack secondmate stall receipt case

* fix(bin): restore the secondmate child devin-config cleanup path

The branch-captain-key sidecar addition mistyped the sibling entry as
.$child_id.devin-config.json, so a forced secondmate teardown would have
stopped removing each child's real <id>.devin-config.json. Restore the
original path and add a behavioral test that stops the child sweep mid-loop
on a refused close, proving the cleaned child's devin config and captain
anchor are both removed while the unconsumed child's records are retained.

* no-mistakes(review): Key captain dedupe on the covered wake rows' fingerprint

* no-mistakes(review): Drop captain-key demotion; prove one escalation on both surfaces

* no-mistakes(review): Drop unrelated teardown test; cite both receipt test files

* no-mistakes(document): Docs already match branch-ack check-receipt retirement
…oorbell (kunchenguid#5664)

* fix(calm): deliver Claude-bound operational input as a record-backed doorbell

Claude Code 2.1.280 removes U+2063 from every submitted prompt, so a typed
operational envelope reaches a Claude Code primary as plain text. The away
daemon now writes the envelope to a record under state/operational-inbox and
types only a plain doorbell naming it; the /afk return check and the Calm mod
recognize the doorbell only when that record holds a current envelope. Marker-
preserving harnesses keep the typed envelope. The live Calm guard accepts the
2.1.280 module-load log line, drives the doorbell, and asserts thinking stays
hidden.

* no-mistakes(review): Fix operational record retention at 7 days and document prune limit

* no-mistakes(document): Point Calm bounds at 2.1.280 evidence; fix afk-exit comment

* no-mistakes(lint): Pick newest Calm e2e transcript without parsing ls

* docs(calm): add a minimal turning-Calm-on step for Claude Code

* fix(spawn): deliver the Claude launch brief as a record-backed doorbell

Claude Code strips U+2063 from the launch-prompt argument too, so a
worker's launch brief arrived with its operational marker removed.
Publish the brief as a record in the receiving home's operational
inbox - a secondmate's own state, not the primary's - and pass only
the printable doorbell naming it, falling back to the typed envelope
when the record cannot be published so the brief body still delivers.

Unwrap doorbell-carried digests in the daemon digest tests that still
read the raw send log under the claude pin, and update the documented
bounds now that launch briefs hide like the other operational rows.

* test(spawn): cover a secondmate's launch-brief record landing in its own home

The record-backed doorbell resolves its state through the receiving
pane's home, so prove a claude secondmate launch publishes into the
seeded secondmate's operational inbox and never leaks a record into
the primary's.

* no-mistakes(review): Pass primary harness to daemon, tighten retention, refresh verdicts

* no-mistakes(review): Prune operational records by exact seven-day elapsed age

* no-mistakes(review): Batch record pruning so large inboxes still expire

* no-mistakes(review): Refuse Claude spawn when brief record cannot publish

* no-mistakes(review): Drop thinking probe from Claude Calm live test and docs

* no-mistakes(review): Record dated Claude Code 2.1.282 reproduction evidence

* no-mistakes(document): Clarify operational doorbell documentation and record expiry

* no-mistakes(document): Correct AFK escalation carrier guidance

* no-mistakes(review): Describe operational record retention as about seven days

* no-mistakes(document): Clarify Calm delivery and operational record retention

* no-mistakes(review): Remove out-of-scope Calm launch guide from Claude docs

* no-mistakes(document): Document Claude launch-brief delivery and refusal

* no-mistakes(document): Correct stale operational-input documentation

* no-mistakes(ci): Fixed the stale Claude trust test to verify that worker and secondmate launches deliver readable, record-backed briefs instead of expecting brief paths in their commands. Annotated the daemon’s output variable for ShellCheck without changing behavior. The affected tests, daemon tests, ShellCheck, and diff check pass locally

* no-mistakes(ci): parse rebased Claude launch after trailer hook prefix

* no-mistakes(review): Trust launch-brief record and restore thinking bound doc

* no-mistakes(review): Parse final Claude launch statement; drop Stop-hook docs

---------

Co-authored-by: Mike Sewell <maikunari@protonmail.com>
Co-authored-by: no-mistakes <no-mistakes@localhost>
kunchenguid#5583)

A host-local relaunch rewrote only the far endpoint, so this home kept the old harness, model, and effort, and appending those keys after pr= broke pull-request poll authentication.
…ound (kunchenguid#5516)

tests/fm-watch-triage.test.sh finishes in about 434s alone and about 698s
under CI load, so the 900s bound the changed-suite runner applies produced
a false timeout under ordinary concurrent validation. Raise the automatic
bound to 1500s, which keeps every measured script under it while staying
below the 30-minute normal CI tier so a genuinely hung script still fails
here with its output before the job cap cancels the lane.

Fixes kunchenguid#3869
Refs kunchenguid#3565
…kunchenguid#5728)

* Fix nested watcher lock reclaim

* no-mistakes(review): Elect a single steal-mutex reaper and bound arm TERM wait

* no-mistakes(review): Reclaim self-held steal mutex and unify autoarm steal reaping

* no-mistakes(review): Resume own interrupted steal reap from its tombstone
…nguid#5710)

* test: hold the back-to-back boundary close on the host's own clock

test_park_boundary_holds_under_back_to_back_closes assumed two engine
turns fit in the ~16s pre-refusal window and that the stub finished a
turn in 3s. Under load the stub's real drain, report, and
acknowledgement take ~13s, so the turn either died at its bound (which
hands the wake to main, no boundary line) or the second close landed
past the window and the fixture failed while the boundary held. 3
failures in 5 runs at a load average near 11.

Hold the first turn on a release file instead: once the engine is in
flight, a second close is appended mid-turn and the turn is released as
the refusal window opens (park bound minus turn bound and grace, read
off the host's own start record). The queued close can then only wait
for the boundary on any machine speed, which is what the test asserts:
the boundary line ends the output, the demo.status row stays queued for
main, and no second engine turn ever starts. A host too loaded to start
the turn at all hands the first close to the same boundary exit.

After: 12/12 at load ~15-42.

* no-mistakes(review): Print boundary test deadline as a decimal integer

* no-mistakes(review): Hold boundary test turn on a FIFO, require full sequence

* no-mistakes(review): Remove stray before/after supervision-host test copies

* test: hold the late close's render until the refusal window opens

The boundary recheck test's node shim slept a fixed 10s, which assumed
the first close was read before the host's refusal window opened. Under
load the close arrived after the refusal check, so the host correctly
refused it before the successor started and the render snapshot never
appeared. Block the wake-prompt render on a FIFO released at the
refusal-open instant read from the host's own start record, so the
pre-turn recheck must refuse on any machine speed.

* no-mistakes(review): Derive minimal park bounds and refresh supervision-host shard hint

* no-mistakes(review): Drive park-boundary tests from a seam-gated host test clock
* fix(bin): stage remote home clones before publishing them

A remote home provision cloned the code root directly into the public
FM_HOME path while rollback() claimed rm -rf of that same path on any
failure. Bash defers trapped signals past a foreground child, but any
other cleanup or lifecycle path that removes the home directory races
the live clone's object copy, producing the CI flake "fatal: failed to
copy file to .../.git/objects/...: No such file or directory".

Clone into a private staging directory beside the home and publish with
an atomic rename once complete, so no cleanup can remove a directory a
live clone is still writing; a home that appears mid-provision now dies
cleanly instead of inheriting torn state. The regression coverage holds
a real clone mid-copy, removes the public path, and requires the
provision to finish and publish intact.

* no-mistakes(review): Prove home ownership by sentinel and hold only a live clone

* no-mistakes(review): Assert raced provision publishes a complete, intact clone

* no-mistakes(document): Document remote home staging and publication safety

* no-mistakes(lint): Fix ShellCheck warning in clone integrity assertion

* no-mistakes(document): Clarify remote home publication and rollback guarantees
* fix(control): keep a relaunched Pi worker's herdr pane status authority alive

Defect: after `bin/fm-control.sh <id> relaunch` (observed live on a herdr
Pi crewmate whose pane read idle while it ran its validation pipeline),
the pane froze at whatever its previous agent had last reported.

Cause, measured on herdr 0.9.1 against a real Pi: a pane has one status
authority, and for Pi with its integration installed that authority is
the lifecycle hooks, so herdr also skips screen detection for the pane.
In the crew shape the registration outlives its agent process (upstream
issue kunchenguid#4115; docs/herdr-backend.md "Restart and liveness behavior"), and
herdr applies only reports carrying the session identity it bound. A
replacement started fresh in that pane reports a NEW session, so its
state reports are ignored and the pane stays frozen. Nothing from
outside repairs it: `pane report-agent-session` and `pane report-agent`
for `herdr:pi` are accepted (rc=0) without being applied unless the
reporter is the registered pane agent, and `pane release-agent` on the
stale record changes nothing.

Fix: a relaunch preserves the binding instead of fighting it. The launch
owner reads the session reference the endpoint's own runtime recorded
(`fm_backend_herdr_pane_agent_session_ref`) and passes it back as Pi's
own `--session <path-or-id>` (`relaunch_resume_args`;
`fm_control_relaunch_resume_flag` owns which adapters and which
registered-agent labels qualify). That is the same reference herdr
itself resumes Pi panes with after a server restart, and the resumed
session's reports land again, which the live check confirmed: the pane
returned to working while the replacement worked and idle when it
settled, on the same session identity.

Safety: relaunch-only (a fresh spawn binds nothing), herdr-only (the one
adapter that records a per-pane session), Pi-family only, and only when
the registration's own agent label matches - so no other adapter's
conversation can be handed to a Pi launch. An unreadable, missing, or
malformed reference degrades to exactly the fresh-session launch that
existed before. No lifecycle, liveness, isolation, or merge guard is
touched, and an empty result leaves every non-Pi launch byte-identical.
`resume` remains a refused verb; docs/agent-control.md and the
harness-adapters references are corrected where they claimed Pi had no
verified resume form at all.

* no-mistakes(document): docs: correct relaunch session-authority ownership and skill paths

* no-mistakes(document): docs: correct stale control-plane ownership claim

* no-mistakes(document): docs: drop unverified Herdr restart resume claim

* no-mistakes(test): Added offline Herdr Pi session-authority relaunch coverage

* no-mistakes(document): Document Herdr Pi relaunch session continuity

* no-mistakes(ci): The failing remote relaunch test tried to arm a PR poll for a secondmate, which `fm-pr-check.sh` correctly refuses. Removed that invalid test scenario; the remaining remote relaunch tests pass, and `git diff --check` is clean
…nchenguid#5758)

Main has been red since fm-pr-check.sh began refusing to arm a merge poll
on a kind=secondmate record (kunchenguid#5696): the relaunch-ordering case in
tests/fm-remote-secondmate-relaunch.test.sh armed its fixture through that
entry point and could no longer be set up.

The ordering guarantee still matters: a secondmate record armed before the
refusal can legitimately carry a trailing pr=/pr_head= identity block until
the watcher retires it, and fm-remote-secondmate-relaunch.sh must still keep
that block last when republishing harness/model/effort. Seed the fixture the
way such a record was really written - pr= appended last to the meta, then
the poll artifacts published through the same
fm_pr_poll_prepare/fm_pr_poll_publish_prepared pair fm-pr-check.sh uses, a
pattern tests/fm-pr-check-security.test.sh already follows - and drop the
now-unused fake gh fixture. The kunchenguid#5696 refusal itself stays pinned by the
security suite's secondmate-record case.
…id#5748)

* feat: run attended supervision on the host for Claude and Cursor

On a home opted into config/supervision-host with a Claude or Cursor
primary, the supervision host now takes the attended wakes the Pi branch
would take: routine outcomes stay off main, and a captain outcome wakes
main once with a branch-outcome line and waits in the drain's new
BRANCH OUTCOMES section until main acknowledges it with mark-processed.

- The offer rule moves into branchOfferForWake, shared by the Pi watcher
  and the host through bin/fm-branch-dispatch.mjs offer.
- The host feeds the dialog mirror at the head of each attended wake and
  passes a close through unchanged when it is main-only, the engine or a
  tool is missing, the primary has no verified mirror, the main session
  cannot be identified, or the session is cooling down.
- The drain presents captain outcomes first, one line per task, never
  behind older routine outcomes, and collapses routine overflow into a
  count that is marked read.
- The return advances the store's read cursor through the away window
  once the brief has rendered, so the first drain does not replay it.
- The branch prompt's mirror wording is host-neutral, and the rule to
  report what main must act on as captain, once per unchanged situation,
  applies only to the attended posture on the host.

* docs: record the attended supervision host live check

* no-mistakes(review): Present pre-window unread outcomes and contiguous captain prefix

* no-mistakes(review): Return brief presents every row it marks read

* no-mistakes(review): Return brief lists every unread outcome in one list

* no-mistakes(review): Keep return list in store order and gate cursor failures

* no-mistakes(review): Make the drain the only branch-outcome presenter after return

* no-mistakes(review): Gate return on drain outcome failures; byte-count outcome budgets

* no-mistakes(review): Gate drain on projection failures; UTF-8-safe byte cuts

* no-mistakes(review): Fail drain without jq; hand unreadable prompt mirror to main

* no-mistakes(document): Correct supervision-host return and drain documentation

* no-mistakes(review): Recheck attended offer at turn start; honest failed-drain brief

* no-mistakes(document): Correct supervision-host posture and drain documentation

* no-mistakes(document): Documentation remains accurate for attended supervision
…unchenguid#5753)

Each '# shellcheck source=' directive makes ShellCheck's external-source
traversal expand that library's whole transitive graph again at the site.
fm-pending-reply-lib carried three directed lazy sources of fm-wake-lib and
two of fm-parent-channel-lib on identical per-call re-source sites, so one
file analysis peaked above 4 GiB and every caller (fm-watch, fm-teardown)
inherited the multiplier - the root cause of the PR kunchenguid#5732 Lint 1 OOM kill.

Keep the runtime '.' commands byte-identical: the lazy re-source under
'local STATE FM_WAKE_QUEUE FM_WAKE_QUEUE_LOCK' is real behavior. Drop the
duplicate directives so each library expands once per unit, and drop the
tmux/classify directives since classify already arrives through the kept
fm-wake-lib expansion and no tmux symbol is referenced here. The directive
above the lib-dir assignment is kept - it binds the bin/ prefix so the
undirected sites still resolve without SC1091.

Measured peak RSS, ShellCheck 0.11.0 -x on Linux arm64:
  bin/fm-pending-reply-lib.sh  4.06 GiB -> 1.96 GiB, zero findings
kunchenguid#5773)

* fix: split bash 5.2 sibling $() in recovery mint and delivery log

Sibling command substitutions on one line can empty a recovery generation
under bash 5.2 when a CHLD trap is set. Mint pid/epoch sequentially, refuse
empty tokens before write, and clean delivery fields before printf.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: split bash 5.2 sibling $() in recovery mint and delivery log

Sibling command substitutions on one line can empty a recovery generation
under bash 5.2 when a CHLD trap is set. Mint pid/epoch sequentially, refuse
empty tokens before write, and clean delivery fields before printf.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: keep recovery mint failure semantics after sibling $() split

Remove the new pid/date refusal and grammar guard so a mint miss still
yields a grammar-valid token and a durable wake row, matching accepted
review intent. Drop the fake-failing-date case that locked in the refuse.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(document): Point recovery-mint hazard comment at its regression test

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
…unchenguid#5790)

Fixes kunchenguid#4756

The voice status reader in bin/fm_voice_records.py reports each
worker's state from the last non-blank line of its status log. When a
worker appends a status line and then a line of plain prose, the
reader reported "note" with the prose line instead of the declared
state, diverging from bin/fm-classify-lib.sh's shell scan.

Scan back through the tail for the newest line whose prefix is a
single lowercase verb-shaped word (letters and hyphens), and report
that event's verb instead of always taking the last line. An
unrecognised verb-shaped prefix still reports "note" rather than
letting an earlier recognised line answer for it, and free text with
no colon is skipped as prose. When the tail holds no such event, the
last line is reported exactly as before.
…newer version (kunchenguid#5786)

* fix(bin): stop reporting an already-installed version as an available update

An update announcement named its version first ("current -> new"), so
reading the first dotted number as the announced version compared the
current version against itself and always looked newer. Read the last
dotted number instead, and only report an available update when that
announced version is newer than the newest installed copy found; when
that version is already installed, report only PATH skew.

Fixes kunchenguid#5151

* no-mistakes(document): docs: gate announce update-available report on newer-than-installed
…h their launch config (kunchenguid#5799)

* fix(bin): pass the profile effort to OpenCode workers through their launch config

The dispatch profile's effort axis was recorded in task metadata but never
reached an OpenCode worker: the launch wrote only a permission grant into
the config it constructs.

OpenCode 1.18.32's config schema carries per-model reasoning effort as
agent.<name>.variant, so the chosen effort is now merged into the same
OPENCODE_CONFIG_CONTENT JSON as the default build agent's variant, keyed to
the resolved model. With no effort chosen the launch stays byte-identical.

Fixes kunchenguid#1373

* no-mistakes(review): gate OpenCode effort variant by model provider family

* no-mistakes(document): docs(opencode): note provider-family gating for effort variant
…nchenguid#5815)

* fix: preserve cancellation as no verdict in crew state

Reuse the green-delivery safeguard for cancelled CI monitors and permit a skipped rebase. Other cancelled outcomes and coarse ledger records use the existing unknown state.

Four delivered-PR regressions failed before the fix and pass afterward. The isolated public resolver and fleet-summary tests prove that undelivered cancellation no longer creates a failure contradiction, while preserving historical records and the terminal_in_flight invariant. Evidence uses fixture no-mistakes responses, not a live daemon cancellation.

Update the existing coarse cancellation assertion from failed to unknown because it encoded this defect; retain its newest-run precedence check. Full fm-crew-state suite and pinned lint pass.

* fix(review): Verify PR disposition before reclassifying terminal validation runs

* fix(test): Add captured cancellation replay coverage for resolver and fleet

* fix(document): Clarify cancellation and terminal delivery documentation
…henguid#5812)

* fix: declare worker background and pipeline waits

Require ship and scout workers to declare owned-work waits with the existing
paused verb before ending a turn or waiting on a pipeline or long command.
Keep the first-sight alert and existing liveness classification unchanged;
subsequent inspection follows the existing long pause cadence.

Validation: emitted brief regression failed before the instruction change
and passes afterward. Public watcher/drain regressions cover the first
alert, repeated wedge suppression, bounded rechecks, and undeclared idle
alarms using isolated backend fixtures. Brief suite, pinned lint, Bash
syntax, documentation inventory, and whitespace checks pass.
No real worker harness was exercised for wait behavior.

* fix(document): Clarify declared worker waits and documentation ownership

* fix(ci): Captain, fixed the cadence test to age both the declaration and first-alert throttle while preserving declaration identity. Reproduced the CI failure using stable identity; corrected tests pass with both stable identity and native macOS behavior. Focused ShellCheck and diff checks pass. Production behavior is unchanged; Linux CI was not rerun locally
…5770)

* fix(bin): bound each lint root in its own ShellCheck process

CI job "Lint 1" died twice at about ten minutes because the two shard
workers each packed about 110 canonical roots into one unbounded ShellCheck
process, and a byte-weight rebalance moved the analysis-heavy fm-watch.sh
into a partition with other heavy roots, so the pair outgrew the 16 GiB
runner before anything could name a culprit.

Run one canonical root per ShellCheck process under an enforced envelope:
a wall deadline plus terminate-then-kill grace via the shared
fm-timeout-lib.sh watchdog, and a per-root rlimit spec applied inside the
child before exec (default a 4 GiB address-space cap, so two workers stay
inside a 16 GiB job with headroom). A root that exceeds the envelope fails
by name with a recorded reason - timeout, memory, signal, or
limit-unavailable - instead of taking the runner down. The per-root
watchdog runs in its own process group so the owner's group sweep cannot
orphan the bounded subtree, and fm_exec_timed now starts the same
escalation when its parent dies before it can be signalled.
FM_LINT_REQUIRE_BOUNDS=1, set in CI, refuses the run outright when a
configured bound cannot be enforced on the host rather than lint uncapped.
Each root's begin/end, reason, duration, and peak RSS stream to stderr in
partition mode and append to a retained <telemetry>.roots.tsv sidecar
uploaded beside the partition telemetry.

Coverage is unchanged: pinned ShellCheck 0.11.0, --norc, --external-sources
full analysis, complete and disjoint partition inventory, workflow lint,
and the backend-purity check, with byte-identical diagnostics across
jobs=1/2 proven by tests/fm-lint.test.sh.

* fix(bin): fail closed on unenforceable lint bounds and size the cap

Required-bounds mode (FM_LINT_REQUIRE_BOUNDS=1, set by CI) now refuses the
run with named errors before any root starts: a missing fm-timeout-lib.sh,
a watchdog that cannot actually bound a probe command, or a host that
rejects the address-space limit all stop the run rather than lint uncapped.
The generalized FM_LINT_ROOT_RLIMITS flag:value interface is replaced by a
single FM_LINT_ROOT_MEMORY_KIB, and the roots sidecar and telemetry record
the run's final exit status after backend-purity and workflow checks
instead of the pre-check lint status.

The default cap is 6 GiB of address space per root, not 4 GiB: ulimit -v
bounds virtual address space rather than resident memory, and ShellCheck's
GHC runtime keeps roughly a third of that space as reservation, so 6 GiB
yields about a 4 GiB working heap budget. A Linux measurement during this
change showed eleven real canonical roots running out of memory under the
earlier 4 GiB cap while the largest passing root peaked near 2.8 GiB
resident; two 6 GiB roots plus runner overhead still fit the 16 GiB job.
Roots that still exceed the cap keep failing by name, and the sidecar's
per-root peak RSS keeps roots approaching the budget visible.

tests/fm-lint.test.sh now proves the memory primitive where it can be
proven: on hosts that accept ulimit -v a perl allocator is refused under a
256 MiB limit and reported by name as a memory death, the pinned ShellCheck
lints a small file under the configured cap and is named when a far smaller
cap binds it, and a watchdog-less copy refuses under REQUIRE_BOUNDS; the
bounded cases skip on macOS, which cannot enforce the address-space limit.

* no-mistakes(review): Prove memory cap binds, pass watchdog owner, drop unused modes

* no-mistakes(review): Capture watchdog owner before startup for every fm_exec_timed caller

* no-mistakes(document): Clarify bounded lint documentation and telemetry

* no-mistakes(document): Correct bounded lint documentation and sidecar path

* docs(bin): restore the per-root memory cap sizing rationale

The pipeline's document step rewrote the ROOT_MEMORY_KIB comment and
dropped the sizing reasoning the change is required to record: address
space vs resident memory, the GHC reservation share, the measured 4 GiB
failures and ~2.8 GiB peak, and the two-roots-plus-runner capacity
arithmetic. Restore it beside the default while keeping the corrected
"not a resident-memory ceiling" framing.

* no-mistakes(review): Document memory cap RSS reduction threshold and first candidate

* no-mistakes(review): Scope owner-death escalation docs to the perl watchdog

* no-mistakes(document): Clarify bounded lint and timeout documentation

* no-mistakes(review): Install perl watchdog signal handlers before forking the command

* no-mistakes(document): Correct bounded lint documentation and stale watcher comments

* no-mistakes(ci): Fixed the supervision-host test’s obsolete expectation: the watchdog now reaps an engine when its host dies. The timeout and supervision-host tests pass locally; the watcher test also passes locally. Lint 1 and 2 remain unresolved: seven canonical roots exceeded the required 6 GiB address-space cap in CI. I did not raise the cap, exempt roots, or reduce source-following coverage to make those failures disappear

* no-mistakes(ci): The two lint checks failed when eight canonical roots hit the enforced memory cap. I reduced repeated ShellCheck source-graph expansion while keeping runtime imports and the canonical root inventory intact. Pinned ShellCheck passes for all changed roots; the relevant local tests pass. The 6 GiB Linux CI run remains unverified

* no-mistakes(review): Restore source directives, raise cap to 8 GiB, classify OOM

* no-mistakes(review): Classify memory deaths from root stderr, not source excerpts

* no-mistakes(review): Match only whole runtime memory-error lines for memory reason

* no-mistakes(document): Clarify lint memory classification in script documentation

* no-mistakes(ci): Fixed both lint checks’ memory-limit failure: each CI lint job now runs one root at a time with a 12 GiB address-space cap. Kept the local two-worker default and updated the sizing comment and test expectation. The lint tests and workflow validation pass locally; Linux CI remains to confirm the heavy roots
tmchow and others added 28 commits September 26, 2026 22:13
* docs: make turnend-guard easier to read

Restructure the turn-end guard doc's prose into shorter sections, lists, and tables without changing documented behavior. Every original heading, anchor, inline identifier, link target, and number is kept.

* no-mistakes(review): Restore legacy-only scope on TERM retirement sentence

* no-mistakes(review): Name Cursor park behavior in live e2e test line
…henguid#5872)

* docs: move situational AGENTS.md sections into on-demand skills

Backpass memory optimization: shrink the always-loaded AGENTS.md by moving
situational contracts (home layout, session-start recovery, validation and
landing supervision, scout completion, away/quiet supervision, Relay
ownership) into agent-only skills loaded at their triggers, with a trigger
index skill.

* docs: classify the new on-demand skills' documentation audience

Register the seven new agent-only skills as agent-runtime docs and fix a
link in validation-supervision that kept its AGENTS.md-relative path.

* docs: close load-timing gaps found by the live regression check

- load validation-supervision whenever an ask-user finding is decided or
  answered, so forbid --yes and process-every-return reach the worker
- keep the mid-task captain-ask rule, the unconfirmed network-checks rule,
  and the worker account pin rule inline in AGENTS.md
- fix cross-references that still pointed at moved AGENTS.md sections
…guid#5879)

* fix: route second-mate signal wakes by their new status span

A second mate's status log is a shared channel carrying many independently
keyed decisions, so judging its signal rows by every decision still open in
the whole log pinned each routine update to main behind any unrelated
parked hold. scopeForUnreadWake (the one owner for Pi and the attended
supervision host) now judges a second-mate signal row by the lines presented
since the last drain, bounded by the existing status-presentation cursor:
a decision, blocked, resolution, or captain-held line, or a line declaring
the key of a still-open decision, keeps the whole row on main, and any
cursor problem falls back to the whole log. Keys are read only at the
status parser's declared positions, with readable time stamps stripped as
bin/fm-classify-lib.sh does. Single-task crewmate and stale routing are
unchanged, and stale and signal rows for one mate keep independent verdicts.

The supervision branch now treats a second mate's done and merged lines as
relayed child outcomes, and fm-teardown refuses the branch actor second-mate
retirement through the existing role-partition helper in both postures.

* no-mistakes(review): Route second-mate resolutions to main only when closing open decision

* no-mistakes(review): Guard bare-verb fold lines and fold resolution spans incrementally

* no-mistakes(document): Clarify second-mate wake routing and retirement documentation

* no-mistakes(ci): The CI failure was a timing-sensitive watcher teardown test, not the PR’s signal-scope code. Extended the bounded wait for both state-directory and home removal. The watcher suite passes locally, and git diff --check is clean
…extension (kunchenguid#5882)

register-extension took the extension lifecycle lock and then the source
lock, while reconcile republishing an unhandled extension result holds the
source lock and reaches the lifecycle lock through the extension host's
process-event path. Both waits are unbounded and both owners stay alive, so
the two could wait on each other forever and freeze the home's monitoring
cycle.

register-extension now takes the source lock first, matching every other
path that holds both. The lifecycle lock still spans binding resolution
through registration publication, so binding retirement stays serialized.

A new lifecycle-order section in the extension-binding suite, run in the
default aggregate, holds a re-registration inside binding resolution while
reconcile republishes that source's unhandled result and requires both to
finish within a bound.

Fixes kunchenguid#5866
* fix(bin): run the repository's own hooks when git -c carries the per-task hooksPath

The per-task hooks wrappers cleared only the GIT_CONFIG_COUNT override before
looking up the repository's own hooks directory. When core.hooksPath reached git
through git -c (GIT_CONFIG_PARAMETERS), directly or inherited by a child
process, the lookup found the wrapper directory again and exited 0, so the
repository's real hook - such as a pre-push publish guard - never ran and the
push succeeded.

The lookup now ignores GIT_CONFIG_PARAMETERS too, so only the repository's
config files decide its hooks directory, and a failed lookup exits nonzero
instead of skipping the hook. AI-trailer stripping is unchanged.

Fixes kunchenguid#5871

* no-mistakes(document): Document git hook chaining and lookup failure behavior
…unchenguid#5744)

* feat(bin): add an optional never-send list to typed dispatch resolution

* Added config/dispatch-never-send, an optional local list of literal
  values and re: regular expressions checked against every string of
  the resolver request before it is sent to typesafe.ai
* A match, an unreadable list, or an empty or invalid pattern now stops
  the request and falls back to the off path, so firstmate dispatches
  through its existing intake; the one stderr diagnostic names at most
  the list line number and never the value
* No list, or a list with no match, leaves resolution unchanged

* no-mistakes(review): Match never-send literals across whitespace, drop regex mode

* no-mistakes(review): Inherit the never-send list into secondmate homes

* no-mistakes(ci): ci-2 (Lint 2), caused by this PR and now fixed. The rule that broke: the test script must not share shell variables with a library it sources. The new secondmate-inheritance test in tests/fm-dispatch-resolve.test.sh sourced bin/fm-config-inherit-lib.sh inside a `( ... )` subshell. That lib assigns `out`, so ShellCheck flagged every later `$out` in the test with SC2031. The subshell was the only place this PR sources that lib. The fix runs the propagation in a child shell instead (`bash -c '. "$1" && propagate_inheritable_config "$2" "$3"' _ lib from to`), so the test shell never sources the lib. `bin/fm-lint.sh tests/fm-dispatch-resolve.test.sh` now exits 0, and `bash tests/fm-dispatch-resolve.test.sh` passes. That includes the check that an inherited list is enforced in a secondmate home. ci-1 (Behavior portable serial 8), not caused by this PR, so no code change for it. The only failing test is tests/fm-remote-secondmate-relaunch.test.sh, which fails with "not ok - could not arm the PR poll fixture for the relaunch-ordering test". This PR doesn't touch that test or the code it runs. I reproduced the same failure locally on the base commit ea7c7f7. Main's own CI run on ea7c7f7 (run 36212602588) fails only this job, with the same message. The recent main runs before it also concluded failure
…hrashing guard (kunchenguid#5903)

* feat(jev): add the guard framework and the memory RSS/swap thrashing guard

A Jev guard is a bounded read-only host diagnostic that turns one class of
resource pressure into a machine-readable audit record and a one-line verdict.
This lands the framework contract (docs/jev-guards.md) with one representative
family (Pattern 46, memory RSS and swap thrashing) as the pilot: thin bash
wrapper, stdlib-only python engine, and a behavioral test through the CLI.

* no-mistakes(review): fix jev mem guard fail-open unknown and contract

* no-mistakes(review): fix mem guard rss pairing, memfree, tests, docs

* no-mistakes(review): fix inverted UNKNOWN branch in fail-forcing test

* no-mistakes(review): register docs/jev-guards.md in audience inventory

* no-mistakes(review): simplify wrapper, drop dead top_n, single-source thresholds

* no-mistakes(review): assert exact exit code in fail-forcing test leg

* no-mistakes(review): tolerate any stdout encoding in text output

* no-mistakes(document): Fix guard contract dash style and output wording
…enguid#5900)

* fix(bin): stop slow GitHub reads from starving and waking the contributions poll

The poll's fixed budget cut the tail URL's reads short, and any read killed at the bound was recorded as a forge failure, so routine GitHub slowness produced an observation-unavailable wake. A configured FM_CONTRIBUTIONS_BUDGET now rides the generated check shim and is cut down to the watcher's per-check bound with a margin, a read killed at the bound or the deadline is budget refusal that leaves records untouched, and the poll moves to the next URL that still has a full observation reserve instead of ending.

* fix(bin): report the bound when a signal death leaks through fm_run_timed

fm_run_external_timeout trusted the wrapper's recorded status over the runner's own bound verdict, so a read killed by the bound's TERM could surface as 143 and callers such as the contributions poll classified a budget-bound read as a forge failure. When the runner reports the bound, a signal-death status is the bound's own TERM racing the wrapper's bookkeeping and is reported as 124; a natural exit still passes through. The replace-shell probe also moves to the repo's portable BASHPID fallback so the suite runs on the macOS system bash.

* no-mistakes(review): Rotate contribution polling and verify generated budget behavior

* no-mistakes(review): Stabilize contribution rotation across successful observation refreshes

* no-mistakes(review): Exclude settled contributions from live observation rotation

* no-mistakes(document): Document contribution poll rotation and observation reserves

* no-mistakes(ci): Captain, fixed timeout handling with a one-line change preserving natural exit 137. Full session-start and contributions suites, targeted regressions, and lint passed. The timeout suite still fails on the known Bash 3.2 BASHPID issue, left unchanged as instructed. Logs retained in .no-mistakes/ci-evidence/. Remote CI was not rerun

* no-mistakes(ci): Fixed the fixture’s lock-acquisition race with a one-line bounded wait. Forced contention reproduced the CI error before the fix and passed afterward; the ordinary held-lock case, lint, syntax, and diff checks passed. Local Bash 3.2 failures remain: “cleanup lock bound 08 gave up before the marker lock freed” (also reproduced without the fix) and “TERM did not stop a watcher blocked inside a poll”. Remote CI was not rerun
…railers (kunchenguid#5859)

* feat: add keep AI trailers setting

* no-mistakes(document): Drop duplicate keep-ai-trailers line, update Cursor attribution note

* no-mistakes(document): Honor keep-ai-trailers for Devin worker attribution

* no-mistakes(document): Qualify Devin attribution note with keep-ai-trailers flag

* no-mistakes(review): Inherit keep-ai-trailers into secondmate homes
…sh-command popup cannot hide the composer (kunchenguid#5876)

* Fix Herdr composer reads blinded by the slash-command popup

Capture the full visible viewport for every Herdr composer state and content read instead of a bounded tail.
Claude Code renders its slash-command popup between the composer and the pane bottom, which pushes the composer outside a tail window.
The pre-Enter payload proof then read an empty composer, judged a typed command unsent, and cleared it without pressing Enter.
The proof-lines value now bounds only the clear cost, not the capture size.
Growing the window adds rows above the composer only, so bottom-most shape selection and prior verdicts are unchanged.
The suffix refusal is kept, and the shared inbox pending-line read stays a bounded tail.
Portable regressions cover the popup-below-composer layout, and the live submit-confirmation guard gains a third exit scenario.
The runtime-backends verification record documents the Herdr 0.9.0 and Claude Code 2.1.283 run.

Closes kunchenguid#5533

* no-mistakes(document): Clarify composer capture bound ownership

* no-mistakes(review): Drop unrequired bracketed-paste Enter fallback from live guard

* no-mistakes(review): Latch trust prompt, check idle composer arm first
…configurable bound (kunchenguid#5917)

Each per-task endpoint read in the session-start fleet digest runs in its own crash-isolated child under the existing timeout helper with a configurable FM_SESSION_START_ENDPOINT_TIMEOUT bound (default 10s).

A hung or killed read becomes that task's endpoint error while the digest continues, and the wrapper banners any abnormal child exit naming its status.
…nchenguid#5886)

* Keep lab tmux sockets on short private paths

* no-mistakes(review): Fix Linux stat checks and fail-closed lab tmux teardown

* no-mistakes(document): Update lab helper documentation for isolated tmux sockets
* Allow silent task-level no-change outcomes

* no-mistakes(review): Exclude silent outcomes from captain-return handoffs

* fix: look up supervision receipts by exact sequence

* no-mistakes(review): Suppress silent notes in away-return brief

* no-mistakes(review): Clarify visible notes; remove unused mode

* no-mistakes(review): Clarify silent outcomes and avoid false drain promises

* no-mistakes(document): Clarify silent supervision outcome documentation
…chenguid#5928)

* feat: make /quiet a statement where the attended supervision host runs

On a home that opted into the supervision host, quiet mode is what the
attended host already does, so /quiet now enters nothing there instead of
launching the quiet daemon and writing a record that would park a present
captain's main.

- bin/fm-afk-launch.sh quiet-check says quiet mode needs nothing where the
  attended host runs, or that the session is paused while its
  broken-session latch holds; a quiet enter refuses there before writing
  anything.
- Where the home opted in but the attended host lacks a part (engine,
  tools, verified mirror writer, identifiable main session, valid mirror),
  quiet-check names it and quiet mode falls back to the daemon.
- Under a live away record on that home, quiet-check and a quiet enter
  refuse and name the record, so the return runs first, whatever
  state/.afk says.
- A quiet enter records mode: quiet in the posture record, so start and
  start-native launch the quiet daemon without FM_AFK_MODE, and the away
  refusal wording fires only for away.
- bin/fm-host-mirror.sh check validates the dialog mirror read-only and
  exits 1 on a missing, unreadable, or invalid mirror.
- The quiet and afk skills and the supervision-host docs describe the new
  behavior; homes without the opt-in and Pi homes keep the daemon path.

* no-mistakes(review): Archive the quiet record when a quiet daemon start fails

* no-mistakes(document): Clarify quiet-mode documentation and remove stale duplicates
…uid#5884)

* fix(bin): grant Claude workers their task-channel dirs via --add-dir

Since Claude Code 2.1.257, a file-tool read (Read/Glob/Grep, and an
Edit's mandatory prior read) of a path outside the working directories
parks --permission-mode auto panes on a one-time interactive question,
and a "Block" answer lands permissions.blockReadsOutsideWorkingDirectories
in user settings, refusing the same reads even under bypass. Firstmate
launches Claude with no --add-dir, so a secondmate's parent-home steering
inbox and a ship or scout worker's launch record, steering inbox, brief
dir, and code-root .agents/skills were all outside: workers wedged on
the question the first time they read a steer.

Every Claude launch, spawn and relaunch, in both permission modes, now
grants exactly the task's channel directories: state/<id>.inbox for a
secondmate (in the parent home), or state/operational-inbox,
state/<id>.inbox, data/<id>, and the code root's .agents/skills for a
ship or scout. Paths resolve to real paths and lazily created channel
dirs are made before launch so the grant never names a not-yet-existing
directory; the whole state/ is deliberately never granted.

The grant keeps the bypass-mode launch argv changed on purpose: it also
protects bypass workers against a machine-recorded Block answer.

* no-mistakes(document): Consolidate Claude launch guidance in configuration reference

* no-mistakes(document): Clarify Claude permission documentation reference
…nguid#4819)

* fix(supervision): prevent idle recovery loops without stranding wakes

* no-mistakes(review): Remove unused wake-append rollback helper
…id#5889)

* fix(bin): stop the remote-job worker busy-polling an idle queue

The serving loop slept 50ms between passes and re-ran state preparation
(chmod on every queue directory), the heartbeat publish, and the stale sweep
on every pass. It now blocks on a worker.wake FIFO that staging,
cancellation, and lane exit nudge, keeps a short fast-poll window after
activity, refreshes the heartbeat at most once a second, and runs the sweep
(which re-applies the queue directories' 0700 modes) at startup and then on
a bounded interval. Lane-owned records are no longer re-read every pass.

Measured with a fork/execve-interposing counter on a --serve worker in a
disposable HOME and queue, bash 3.2, 20-second windows (the counter slows
the old loop to about 5 passes a second, so real-host rates were higher):
  idle worker             146 forks/s, 61 execs/s -> 4.8 forks/s, 3.1 execs/s
  one running long job    232 forks/s, 100 execs/s -> 15 forks/s, 11 execs/s
Stage-to-result latency for a no-op job, idle and back to back, stayed at
about 0.8-1.2s in both versions (dominated by job execution, not pickup).

* perf(bin): drop per-cycle forks from watcher, drain, and lock helpers

The watcher, drain, inactive-reconcile scan, and branch-outcome reads forked
small external commands on every cycle where bash can do the same work.

- fm-wake-lib.sh gains fm_dirname_to, fm_basename_to, and
  fm_epoch_seconds_to, exact stand-ins for $(dirname --), $(basename --),
  and $(date +%s); the clock uses printf %(%s)T on bash 4.2+ and still forks
  date exactly once on stock macOS bash 3.2.
- fm_lock_abs_path, fm_wake_signal_seen_path, fm_path_age, the watcher's
  age_of and wedge timer, and the recovery-marker line count use them or
  plain reads instead of dirname/basename/tr/date/wc.
- window_to_task reads a meta file once instead of two
  grep | tail -1 | cut -d= -f2- pipelines per file per call.
- fm-classify-lib.sh reads uname -s once at source time instead of in every
  status stat helper.
- Libraries sourced every cycle derive their own directory without forking
  dirname, including the backend adapter siblings a subshell re-sources on
  each probe.

tests/fm-fork-free-helpers.test.sh pins each replacement against the command
it replaces on edge-case inputs, under every available bash and both the C
and a UTF-8 locale; CI's stock macOS bash lane runs it under /bin/bash 3.2.

Measured with a fork/execve-interposing counter in a disposable home, one
tmux crew task, FM_POLL=1 (forks and execs per watcher cycle, per run
otherwise):
  watcher cycle      bash 5.3  299/138 -> 199/66   bash 3.2  341/146 -> 224/80
  drain              bash 5.3  492/238 -> 430/200  bash 3.2  567/250 -> 491/212
  inactive scan      bash 5.3   27/14  ->  17/4    bash 3.2   37/14  ->  17/4
  branch-outcome     bash 5.3   40/21  ->  35/16   bash 3.2   48/24  ->  38/19

* test: note the interpreter-expanded version probe for shellcheck

* no-mistakes(review): Fix worker idle bounds and fork-free contributions snapshot

* no-mistakes(review): Coalesce buffered worker wake nudges into one wake

* no-mistakes(review): Coalesce wake nudges via pending marker so publishers never block

* no-mistakes(review): Claim wake nudges atomically via noclobber pending marker

* no-mistakes(review): Release abandoned wake claims only after a 30-second bound

* no-mistakes(review): Drop worker wake FIFO; load path helpers side-effect free

* no-mistakes(document): Document remote worker polling and preemption cadence

* no-mistakes(ci): Fixed both failing CI shards: isolated remote and teardown test fixtures now include fm-path-lib.sh, which fm-wake-lib.sh requires. The three affected tests, fm-lint.sh, and git diff --check pass locally
…uid#5941)

* fix: keep a successor watcher and remote-reply listeners across the gaps that dropped them

A main-only supervision pass-through exited without leaving a watcher, and each remote-reply poll released its claim until the next cycle, so short-lived listeners stayed down.

* no-mistakes(document): Clarify listener and supervision continuity documentation

* no-mistakes(ci): Fixed the three Greptile findings: failed ingestion leaves one durable capture, failed reads exit instead of relistening, and the disposable-checkout guard rejects state paths outside the marked lab. Added behavioral tests; the remote-reply and watcher-lock suites, shell syntax checks, and git diff checks passed

* no-mistakes(ci): Fixed a race in the wake-queue interruption test: it now waits for the drain to own the lock and enter handling before signaling it. The wake-queue suite, shell syntax check, and diff check pass
…enguid#5925)

* fix: date replayed branch outcomes and ask main to check current state first

A captain outcome main never acknowledged is presented again, which after a
harness or posture switch, or the first drain after the upgrade whose earlier
presenter never advanced the read cursor, can be days after its situation
settled. The replay read as fresh news, so a PR since merged looked ready.

bin/fm-branch-outcome.sh now adds a "recordedAgo" age (minutes, hours, then
days) to present and unprocessed rows, one owner of that wording for both
presenters. The drain's BRANCH OUTCOMES captain lines and the Pi branch's
processing request name that age and ask main to check the task's current
state first; an outcome already settled needs only the acknowledgement, with
nothing relayed to the captain. Nothing is adopted as processed, so a fresh
home's first outcome is still presented until acknowledged.

* no-mistakes(review): Absent processed marker reads 0; never adopt read cursor

* no-mistakes(review): Require recordedAgo in Pi requests; report undated rows to main

* no-mistakes(review): Keep recordedAgo on captain rows only in present output

* no-mistakes(document): Correct cutover documentation and retire stale migration guidance

* fix: keep settled branch outcomes out of main's reply to the captain

A live Pi primary that took over a host-drain home received the carried-over
outcomes dated and check-first, but its processing reply still told the
captain about an outcome whose decision had since been answered. The request
also claimed every outcome was already shown as an anchor entry in this
transcript, which is false for an outcome carried over from before a restart
or a switch of primary.

The Pi processing request now says each outcome was recorded earlier and may
already have been seen or handled, and that a settled outcome gets no
captain-facing mention at all in the reply or any recap, not even that it is
settled. The drain's BRANCH OUTCOMES header and the supervision docs state the
same rule, and the tests check both delivered texts.

* fix: scope main's outcome reply to what is still open

Telling main what not to say about a settled outcome was not enough: in two
live Pi trials the processing reply still told the captain that an answered
decision was settled. Main now sorts the outcomes by current state first, and
its reply to the captain covers only the still-open ones, written as if the
settled ones had never been listed. With that framing three live Pi trials
kept the settled outcome out of the reply and relayed the open one each time.

The drain's BRANCH OUTCOMES header and the supervision docs use the same
framing, and the tests check both delivered texts.

* no-mistakes(review): Clarify that main acknowledges every presented captain outcome

* no-mistakes(document): Clarify outcome cursor ownership across Pi and host

* no-mistakes(ci): Fixed Pi replay by batching unprocessed captain outcomes oldest-first and acknowledging only through each batch. Verified a marker-less backlog over 1 MiB replays through all batches. A real-drain regression confirms an older keyed decision remains under OPEN DECISIONS after a newer branch row is acknowledged; the check-first instruction now names those decisions. Relevant targeted tests and branch-supervision tests passed; the full host suite timed out

* no-mistakes(ci): Fixed the host drain’s check-first wording in bin/fm-wake-drain.sh; the CI fixture now passes. The full host suite passed the affected fixtures but timed out later. Syntax and diff checks passed

* no-mistakes(ci): Fixed ci-1: abbreviated Pi outcome summaries now stay within 1,024 characters and include a row-specific full-outcome lookup command. The delivered instruction requires reading the full outcome before acting, relaying, or acknowledging it. The new extension-driver regression failed before the fix and passes now; the Pi and supervision-host suites pass

* no-mistakes(ci): Corrected the batching sentence in docs/pi-supervision-branch.md. The cancelled CI check needs no code fix; its clean rerun passed. The Pi branch extension suite and git diff check passed
…guid#5961)

* fix: wake an idle Claude primary for attended main-only hand-backs

An attended main-only pass-through confirmed a handling handoff for the
successor it leaves running, which flipped the recovery marker to handling.
The Claude Stop hook only rewakes main while that marker reads downtime, so
the close reached no one and an idle primary slept with wakes queued.

The pass-through now leaves the marker at downtime, and a close that turns
main-only at its turn hands the consumed handoff back to downtime before it
reaches main. Regression tests drive the real Stop hook around the real host
on both paths and for the successor's own later close, and a new opt-in live
guard proves it against an idle interactive Claude primary with a pre-fix
negative control.

* no-mistakes(review): Assert live lab Stop-hook registration via parsed settings JSON

* no-mistakes(document): Correct supervision hand-back documentation

* no-mistakes(ci): Fixed the failed downtime-write path so the Stop hook notifies main instead of silently dropping the close. Corrected the live guard’s tracked-hook check and added the requested at-turn main-only scenario. The new regression failed before the fix and passed after it; the host suite, syntax checks, and diff check pass. The credentialed live guard was not run in this CI phase because it writes outside the worktree

* no-mistakes(ci): Fixed the Stop hook’s retry ordering: a crashed host gets its bounded retry before a non-crash hand-back failure is reported. The Stop-hook suite passes, including the crash regression. The host suite passed the failed-marker-write regression but timed out before completing; syntax and diff checks pass
kunchenguid#5916)

* Fix worker launches to enter recorded worktrees

* no-mistakes(review): placeholder

* no-mistakes(document): Update agent-control.md worktree-refusal note to match new universal cd+assert

* no-mistakes(review): Add regression tests for Orca spawn/relaunch worktree carve-outs

* Fix PR relaunch and prelaunch cwd verification

* no-mistakes(document): Fix docs/agent-control.md: worktree cwd check is pre-launch, not post-launch

* fix: slim worktree launch change onto upstream main
…ardown (kunchenguid#5997)

* WIP: retire task-keyed watcher markers and orphan journals at teardown

Re-applies old PR kunchenguid#5584 on current main: teardown retires the
turn-ended .seen-* signature and an orphaned Herdr presentation
journal whose workspace is already gone, and the wake-drain rotates
its own dead scratch files. Not yet validated through no-mistakes.

* no-mistakes(ci): Fixed the Greptile P1 finding in bin/backends/herdr.sh. fm_backend_herdr_projection_token_workspace_gone used `! ... jq -e ... 2>&1`, which swallowed a jq runtime error (thrown when a non-object workspace entry, e.g. a number before a live token-bearing workspace, hits `.label`) and flipped it to a "gone" verdict, causing teardown to delete a still-live v1 presentation journal. Invariant: a workspace-query error/ambiguity must never be read as token absence; only a cleanly-parsed list with no token-bearing label is "gone". Replaced the body with a single jq verdict (unknown/present/gone): a non-array list or any non-object/non-string-label entry yields "unknown", jq errors/empty output fall through `|| return 1` to unknown, and only "gone" returns 0. Sibling fm_backend_herdr_projection_endpoint_matches_journal already fails safe on jq error (empty match -> journal kept), so it needed no change, matching the author's scoping. Added test_teardown_retains_v1_journal_when_workspace_query_ambiguous driving real teardown with a malformed workspace-list entry, proving the journal is kept and no workspace close occurs. Verified the old logic returns GONE on that input (test fails before, passes after); full tests/fm-teardown.test.sh suite passes (exit 0) and shellcheck is clean. Marker-naming finding left untouched per explicit out-of-scope instruction
…unchenguid#6002)

* fix(tests): disable Claude Code's auto-updater during live harness runs

fm_live_gate let a live run proceed without ever setting
DISABLE_AUTOUPDATER, so a live Claude test could let the real updater
repoint ~/.local/bin/claude into a temporary directory and stop every
Claude process on the machine from starting. Export
DISABLE_AUTOUPDATER=1 on every path where the gate lets a live run
proceed, and assert the export in tests/fm-live-gate.test.sh, including
that it reaches a child process the same way a real harness pane would
inherit it.

* no-mistakes(ci): Greptile flagged that the PR's DISABLE_AUTOUPDATER inheritance test only checked a `bash -c` direct child, not the fm-spawn.sh launch path. The user chose to fix it with a regression on that path. In tests/fm-live-gate.test.sh I replaced the generic child test with test_disable_autoupdater_reaches_the_claude_pane_on_the_fm_spawn_launch_path: it drives the real fm-spawn claude launch through the spawn fixtures, captures the exact staged launch command, and runs it as a synthetic pane whose only `claude` is a stub recording the inherited DISABLE_AUTOUPDATER, asserting it saw 1. Switched the file to source fixtures.sh (pulls in lib.sh, guarded) for the spawn helpers. Verified it is a real guard: the stub records `1` when the ambient var is set and `unset` when absent, so it fails if fm-spawn ever scrubbed the variable (e.g. env -i or -u). This confirms fm-spawn's launch construction never references the name and passes it through via ordinary ambient inheritance with no allowlist. Full suite passes (12 tests ok), shellcheck clean. Note for the outer executor: I embedded the daemon caveat as a code comment in the test, but the finding also asks the PR body to state that a backend daemon already running before the gate exported the variable does not inherit it and fully covering that would need launcher support - that forge-side PR-body sentence is outside this CI phase's scope

* no-mistakes(ci): Fixed Greptile finding ci-1. Root cause: fm-spawn.sh handed its launch command to an already-running backend daemon that never inherited the test process's exported DISABLE_AUTOUPDATER, so ambient inheritance dropped it and Claude's auto-updater could still run. Fix (bin/fm-spawn.sh): when DISABLE_AUTOUPDATER is set in the spawn's own environment, embed `export DISABLE_AUTOUPDATER=<value>;` into the LAUNCH command text (same idiom as the adjacent COMPACT_ADVISER_DISABLE export), so it survives a daemon-built pane, the env -i allowlist path, and relaunch alike; gated on presence so ordinary spawns are unchanged. Added regression test test_disable_autoupdater_survives_a_daemon_pane_that_never_inherited_it in tests/fm-live-gate.test.sh: stages a real claude launch with DISABLE_AUTOUPDATER set in the spawn env, then runs that exact command in a synthetic pane with `env -u DISABLE_AUTOUPDATER` and asserts the claude stub still recorded autoupdater=1. Verified the test fails (autoupdater=unset) without the fix and passes with it; the round-1 ambient test stays green either way. Full suite passes (13 ok); test file and isolated snippet shellcheck-clean (full fm-spawn.sh shellcheck kept getting terminated by the memory-constrained host, not by findings). Forge-side note for the outer executor: the PR-body caveat that fully covering the daemon case would need launcher support no longer applies to the Claude launch path and should be corrected
Merge the frozen canonical upstream while preserving the fork's native
Windows transports, private publication, extracted harness and catalog
owners, startup cost contracts, and guarded lifecycle recovery.

Compose successor-safe stale-link reaping with legacy directory-chain
recovery, carry Pi runtime references through its adapter, and retain all
applicable upstream and fork regression cases.

Local evidence is bounded and incomplete after the two-timeout circuit
breaker. The ordinary pull request records unresolved observations and
routes the complete cross-platform matrix to GitHub Actions.

Firstmate-Upstream-SHA: 29213a0
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Upstream renamed the effort-variant regression, but the registry union
retained its obsolete name as well as the replacement. The shared case
runner validates every registered function before selecting a case, so
even the focused Windows Copilot launch case failed before execution.

Register the replacement once in the original position. Restore the
upstream lint exclusion regression whose registration survived without its
body, and update the remaining generic timeout assertions to 1500 seconds
without changing the measured Windows exceptions.
Keep the capped-inventory race regression aligned with upstream's explicit
no-verdict cancellation state. Supply the extracted harness dependency
closure to the incoming AFK return fixtures so Pi detection is exercised.

Keep the fork's absolute lock-path spelling and no-canonicalization fast
path in the imported differential test, including nonexistent parents.

The new hook installer passed raw MSYS paths containing glob characters to
native Git. Route that worktree argument and the launch's hooksPath value
through the existing native path owner. Add a native commit-object and
hook-chaining regression to the existing Windows management lane.

CI owns runtime verification because the local two-timeout circuit breaker
has already stopped further local test execution.

Firstmate-Upstream-SHA: 29213a0
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The new Claude host-mirror hooks omitted the platform-specific no-op fields
that prevent Copilot from importing duplicate operational hooks. Retain
Claude's commands and cover all eight imports, including native PowerShell
in the existing Windows management lane.

Update the two remaining Claude relaunch assertions to the record-backed
launch brief, preserving the inline-encoding assertions for Copilot.
Remove duplicate path-library symlinks from gotmp fixtures; the existing
private-path fixture installer already owns those copies.

Scope endpoint-hang cleanup assertions to a context re-emit, which performs
the same bounded digest reads without separately owned startup probes
querying the fake backend. Keep both timeout values and leftover-process
assertions unchanged; do not widen production deadlines.

The completed CI run at f23f9fb had 22 passing checks and three failed
Linux shards. CI owns verification of this follow-up; the exhausted local
validation circuit breaker remains in force.

Firstmate-Upstream-SHA: 29213a0
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The completed 70751ec CI matrix passed 24 checks. Its remaining serial
shard reached two assertions hidden behind the earlier failures: the
abnormal-startup test expected upstream's old wording, and the copied
Herdr recovery root omitted the skills directory Claude must grant.

Match the composed abnormal-exit diagnostic while keeping explicit
negative deadline assertions. Restore the frozen fork's cumulative-bound
wording and timing heading only for a real timeout, retaining upstream's
distinct abnormal-death banner and the shared breadcrumb guidance.
The existing runtime-bound regression already requires this distinction.

Create the expected skills directory in the copied recovery fixture,
without weakening Claude's task-channel grant or copying live state.
The earlier endpoint cleanup and Claude carrier assertions passed in CI.
No local validation was restarted after the two-timeout circuit breaker;
the exact new head's automatic matrix owns verification.

Firstmate-Upstream-SHA: 29213a0
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@timbarreto
timbarreto merged commit 011c291 into main Sep 28, 2026
25 checks passed
@timbarreto
timbarreto deleted the reconcile/upstream-2026-09-28-29213a09-9bf222f0 branch September 28, 2026 20:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.