Skip to content

Merge upstream firstmate into the fork - #47

Merged
cloud-practitioner merged 19 commits into
mainfrom
fm/fm-merge-upstream-5
Oct 8, 2026
Merged

cloud-practitioner merged 19 commits into
mainfrom
fm/fm-merge-upstream-5

Conversation

@cloud-practitioner

@cloud-practitioner cloud-practitioner commented Oct 8, 2026 •

Copy link
Copy Markdown
Owner

Intent

Bring kunchenguid/firstmate main (tip b062eb94) into fork main as a true merge that carries upstream changes only, so fork history and fork-only commits stay intact.
This PR must land as a merge commit (bin/fm-pr-merge.sh ... -- --merge), never squashed, rebased, or through GitHub "Sync fork", the same rules as fork PRs #35, #39 and #46.
The merge commit cd92819c has exactly two parents: fork main 5bff92af and upstream b062eb94.

What changed

The 14 upstream PRs merged since fork PR #46:

Conflict resolutions (9 files)

  • bin/fm-exclude-tools-lib.sh, the exclude-tools comments and JSON encoding in bin/fm-spawn.sh, docs/configuration.md validation, and all 9 hunks of tests/fm-spawn-dispatch-profile.test.sh: upstream side. Upstream feat: add per-home worker tool exclusions kunchenguid/firstmate#6750 is the refined version of fork PR feat: add per-home worker tool exclusions #44.
  • Herdr recovery lock (bin/fm-spawn.sh, docs/herdr-backend.md): combine option. The fork's 120 s recovery wait stays the default, and upstream's --herdr-resume-lock-wait is added for an unbounded wait. spawn_herdr_presentation_order_lock_acquire takes [wait-seconds|wait]; the doc paragraphs keep both sides' facts.
  • Coherence edits so the combined behavior reads consistently: the fm-spawn.sh header paragraph, docs/verification/runtime-backends.md, and the default-refusal case in tests/fm-backend-herdr-presentation-e2e.test.sh now uses FM_TEST_HERDR_RECOVERY_LOCK_WAIT=1.
  • tests/fm-supervision-host.test.sh: fork run_case form plus upstream's 2 new cases. tests/fm-test-fixtures.test.sh and tests/fm-watch-arm.test.sh: keep both sides' tests. tests/fm-control-herdr-smoke.test.sh: upstream's cleanup structure with the fork's two removal helpers.

Changes beyond the merge resolutions, all made by the no-mistakes pipeline and approved:

  • 8d0b33be (review fix): the two new recovery setups in tests/fm-backend-herdr-presentation-e2e.test.sh apply use_legacy_restart_record, and the opt-in invocation uses the same shortened recovery wait as the default-refusal one.
  • f7be2507 (test stabilization): a longer bounded readiness wait in tests/fm-control-relaunch.test.sh.
  • de4a4197 (document step): documentation wording on Herdr recovery and Calm coexistence.
  • ec5b6c8e (CI fix): a race in the downtime-write failure fixture of tests/fm-supervision-host.test.sh.

Fleet risks

  • The Herdr presentation-lock path changes to a per-account directory (fix(bin): make the Herdr presentation lock namespace per OS account kunchenguid/firstmate#6780). Until every home sharing a Herdr session runs the new code, an old-code and a new-code home do not serialize presentation changes. Run /updatefirstmate when no Herdr spawn, recovery, or cleanup is under way so all homes switch in one pass.
  • Recovery behavior is unchanged for running homes: the 120 s default stays.
  • No config, data, backlog, status, or task-record schema changes; config/project-capacity is new and optional.
  • No skill is renamed or removed. Running watchers, supervision hosts, and Pi extensions keep old code until restarted.

Follow-ups outside this PR

Testing

  • bash -n on every touched shell file; bin/fm-test-run.sh --check-coverage ok; bin/fm-lint.sh on the conflicted and coherence-edited files, rc=0.
  • Targeted suites with --jobs 1, Herdr pane identity, FM_HOME, FM_BACKEND and forge credentials unset, all passing: fixtures, watch-arm, control-relaunch, project-capacity, spawn-dispatch-profile, backend-herdr (fake Herdr), teardown, supervision-host, pending-reply.
  • Real-Herdr suites were not run locally; CI's Herdr lane is the acceptance evidence for the recovery-lock behavior.
  • CI on this PR: all checks green.

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 2 warnings
  • 🚨 tests/fm-backend-herdr-presentation-e2e.test.sh:1550 - The new recovery fixture retains herdr_process_identity across stop/provision. The restored shell has a different identity, so the fork's ownership check in bin/backends/herdr.sh:1045 refuses the old-pane close; reclaim rolls back and falls back flat, failing the same-workspace assertion at tests/fm-backend-herdr-presentation-e2e.test.sh:1567. Both new restart setups (:1470 and :1526) omit use_legacy_restart_record, unlike the existing recovery cases at :1259, :1286, :1330 and :1369–1370. Apply that existing fixture helper before recovery rather than weakening the fork's ownership protection.
  • ⚠️ tests/fm-backend-herdr-presentation-e2e.test.sh:1532 - After correcting the fixture, the opt-in regression still passes if --herdr-resume-lock-wait is ignored: its 30-second holder releases within the retained 120-second default, satisfying both success and elapsed-time assertions. The default-refusal invocation at :1491 shortens its bound, but the opt-in invocation at :1549–1550 does not. Give both invocations the same shortened FM_TEST_HERDR_RECOVERY_LOCK_WAIT so the existing test distinguishes bounded refusal from flag-enabled waiting.
  • ⚠️ bin/fm-watch.sh:2198 - The claimed pane-capture TERM fix leaves reachable captures outside the owned process group. A delivered local pending reply reaches bin/fm-watch.sh:2737 → bin/fm-pending-reply-lib.sh:1659 → :819; a blocked capture there still delays TERM on Bash 3.2 and is not tracked for cleanup. Other routes remain through crew-state reads at bin/fm-watch.sh:744, :2983 and :3126 (bin/fm-crew-state.sh:322/:339), and empty-tail busy fallbacks following the changed captures at bin/fm-watch.sh:552, :874, :893 and :3071 (bin/fm-busy-lib.sh:1109/:1128/:1150). Extend the existing cancellable, owned-background execution boundary to these helper-driven pane reads; replacing only direct capture calls does not establish the stated invariant.
  • ⚠️ bin/fm-wake-lib.sh:1526 - Capacity admission can silently undercount when a registry cannot be opened. With capacity 1, one worker in a registered local secondmate home, and a regular but unreadable root data/secondmates.md, the registry guard accepts the file, the loop's redirection fails, and the walk nevertheless returns success without that home. The new spawn then admits a second worker. Related unchecked reads are bin/fm-project-capacity-lib.sh:159 (declaration open failure becomes an uncapped result) and :201/:203/:204 (kind, pr and project reads use a helper that suppresses I/O errors). Refuse failed reads in the shared home walk and capacity reader before treating their results as a complete holder count.

🔧 Fix applied.
2 warnings still open:

  • ⚠️ bin/fm-watch.sh:2198 - The claimed pane-capture TERM fix leaves reachable captures outside the owned process group. A delivered local pending reply reaches bin/fm-watch.sh:2737 → bin/fm-pending-reply-lib.sh:1659 → :819; a blocked capture there still delays TERM on Bash 3.2 and is not tracked for cleanup. Other routes remain through crew-state reads at bin/fm-watch.sh:744, :2983 and :3126 (bin/fm-crew-state.sh:322/:339), and empty-tail busy fallbacks following the changed captures at bin/fm-watch.sh:552, :874, :893 and :3071 (bin/fm-busy-lib.sh:1109/:1128/:1150). Extend the existing cancellable, owned-background execution boundary to these helper-driven pane reads; replacing only direct capture calls does not establish the stated invariant.
  • ⚠️ bin/fm-wake-lib.sh:1526 - Capacity admission can silently undercount when a registry cannot be opened. With capacity 1, one worker in a registered local secondmate home, and a regular but unreadable root data/secondmates.md, the registry guard accepts the file, the loop's redirection fails, and the walk nevertheless returns success without that home. The new spawn then admits a second worker. Related unchecked reads are bin/fm-project-capacity-lib.sh:159 (declaration open failure becomes an uncapped result) and :201/:203/:204 (kind, pr and project reads use a helper that suppresses I/O errors). Refuse failed reads in the shared home walk and capacity reader before treating their results as a complete holder count.
⚠️ **Test** - 2 warnings
  • ⚠️ The Test agent did not finish within its invocation budget. Reported: agent run tests timed out after 30m0s: agent last produced output 22s ago (262 observed); agent reported: pi exited: exit status 143. This is a budget or provider-slowness cut, not a code failure. Re-running the same request costs another full budget, so no further attempt is made automatically. If this repository's targeted tests or evidence gathering routinely approach the default 30m0s, raise test_agent_timeout in global config. Respond with fix to spend another budget: a repair turn runs only for selected findings other than this budget cut, then validation re-runs. Or abort and retry after raising the budget.
  • 🚨 Approval is refused: the run worktree at ~/.no-mistakes/worktrees/450411b3e67c/01M4DMESQPKH4HM8802D7T5VFZ holds work no Test turn validated, and the steps after Test would commit and publish it. It holds uncommitted changes to .test-local/, l.Vx4OHn/ (inspect with git -C ~/.no-mistakes/worktrees/450411b3e67c/01M4DMESQPKH4HM8802D7T5VFZ status and git -C ~/.no-mistakes/worktrees/450411b3e67c/01M4DMESQPKH4HM8802D7T5VFZ diff). Respond with fix to validate it, or abort.

🔧 No changes applied.
3 warnings still open:

  • ⚠️ tests/fm-control-relaunch.test.sh:1403 - The relaunch suite intermittently failed with 'relaunch did not reach its pre-publication endpoint check', then passed unchanged on retry. Its approximately two-second synchronization window did not reliably reach the staged check on this host. Stabilize that bounded handshake before relying on this test as a deterministic gate.
  • ⚠️ The spawn/dispatch-profile and teardown suites exceeded the imposed 240-second per-script bound while reporting successful cases. Their remaining checks were not validated; these interruptions are not confirmed product failures. Complete both authorized suites in CI or a separately authorized run with sufficient time before approval.
  • ⚠️ live validation verdict: inconclusive (0 of 9 scenarios were driven live against the product); untested: Resume a Herdr task using the fork's default recovery wait and refuse when its bound expires, Resume with --herdr-resume-lock-wait beyond the default bound and reclaim the same workspace, Dispatch concurrent workers at project capacity without allocating an extra task, Launch Pi workers with home-specific exclusions and reject invalid or leaking exclusions, Reject stale or foreign Herdr endpoint identities without steering or closing another task, Re-arm a watcher after downtime, surface durable wakes, and exclude snapshot-scoped environment, Reject another mate's echoed correlation token while accepting the owning mate's reply, Relaunch a worker transactionally while retaining its work and concurrent durable metadata, Complete safe teardown while preserving unlanded work and unconfirmed endpoint evidence
  • Live validation: ⚠️ inconclusive - 0 of 9 scenarios driven live against the product
Scenario Result Live Evidence
Resume a Herdr task using the fork's default recovery wait and refuse when its bound expires ⏸️ untested no The recorded decision prohibits running the real-Herdr suite or lifecycle commands locally. Static checks cannot prove recovery timing. The authorized real-Herdr CI lane must execute the recovery scen…
Resume with --herdr-resume-lock-wait beyond the default bound and reclaim the same workspace ⏸️ untested no Real lock contention and same-workspace recovery require the real-Herdr presentation suite, which the recorded decision reserves for CI. No live lifecycle attempt was made against an operator session…
Dispatch concurrent workers at project capacity without allocating an extra task ⏸️ untested no The prior payload did not establish a live result: tests/fm-project-capacity.test.sh completed successfully, and targeted-suites.log records last-place concurrency, mutation-free deferral, shared-home…
Launch Pi workers with home-specific exclusions and reject invalid or leaking exclusions ⏸️ untested no The authorized fixture suite did not complete registry-reporting and subsequent exclusion-boundary checks, and it launches no real Pi worker. Completing that suite and live agent validation requires a…
Reject stale or foreign Herdr endpoint identities without steering or closing another task ⏸️ untested no The prior payload did not establish a live result: tests/fm-backend-herdr.test.sh completed successfully, including ownership, restored-endpoint, process-identity, and closure safeguards, but used Her…
Re-arm a watcher after downtime, surface durable wakes, and exclude snapshot-scoped environment ⏸️ untested no The prior payload did not establish a live result: tests/fm-watch-arm.test.sh completed successfully, and watcher-cli-output.txt preserves actual recovery acknowledgement commands, but the watcher scr…
Reject another mate's echoed correlation token while accepting the owning mate's reply ⏸️ untested no The prior payload did not establish a live result: tests/fm-pending-reply.test.sh completed successfully, including echoed-token and longer-token adversarial cases, but only deterministic helper and c…
Relaunch a worker transactionally while retaining its work and concurrent durable metadata ⏸️ untested no The prior payload did not establish a live result: tests/fm-control-relaunch.test.sh passed unchanged on retry, and relaunch-recheck.log records successful completion, but the session provider and age…
Complete safe teardown while preserving unlanded work and unconfirmed endpoint evidence ⏸️ untested no The authorized fixture suite did not complete its remaining pipeline/process-cleanup checks. No live Herdr teardown was permitted; complete coverage requires the remaining targeted execution and the a…
  • Removed the user-identified .test-local/ and l.Vx4OHn/ scratch directories.
  • bash -n tests/fm-backend-herdr-presentation-e2e.test.sh
  • bin/fm-lint.sh tests/fm-backend-herdr-presentation-e2e.test.sh — the specifically authorized R1/R2 lint check.
  • bin/fm-test-run.sh --jobs 1 --per-script-timeout-secs 240 --json ~/.no-mistakes/evidence/01M4DMESQPKH4HM8802D7T5VFZ/targeted-timing.json tests/fm-test-fixtures.test.sh tests/fm-watch-arm.test.sh tests/fm-control-relaunch.test.sh tests/fm-project-capacity.test.sh tests/fm-spawn-dispatch-profile.test.sh tests/fm-backend-herdr.test.sh tests/fm-teardown.test.sh tests/fm-pending-reply.test.sh — with inherited FM, Herdr identity, task-backlog overrides, and forge credential variables unset.
  • bin/fm-test-run.sh --jobs 1 --per-script-timeout-secs 150 tests/fm-control-relaunch.test.sh — unchanged retry under the same sanitized environment; passed.
  • Captured actual watcher acknowledgement output and interruption diagnostics in the evidence directory.
  • Cleaned test-owned relaunch fixture remnants and verified git status --porcelain=v1 and git diff --stat were empty.

🔧 Fix applied.
2 warnings still open:

  • ⚠️ The required real-Herdr recovery behavior—retaining the 120-second default while allowing opt-in waiting beyond its bound—has no live evidence from this step. The recorded execution restriction permits fixture suites and presentation-test syntax/lint checks, but excludes real Herdr lifecycle validation. Resolve this acceptance gap through the required Herdr CI lane or a separately authorized isolated lab run. Rendered Pi Calm behavior also remains unverified.
  • ⚠️ live validation verdict: inconclusive (0 of 12 scenarios were driven live against the product); untested: Resume a Herdr task within the default recovery wait, then refuse a holder that exceeds its bound, Resume with --herdr-resume-lock-wait beyond the default bound and reclaim the same workspace, Dispatch concurrent workers at project capacity and defer excess work without allocating resources, Launch Pi workers with home-specific tool exclusions and reject invalid or unsupported exclusion lists, Reject stale or foreign Herdr endpoint identities without steering or closing another task, Re-arm supervision after downtime, surface durable wakes, and exclude snapshot-scoped environment, Reject another mate's echoed correlation token while accepting the owning mate's reply, Relaunch a worker transactionally while preserving unfinished work and concurrent durable metadata, Complete safe teardown while preserving unlanded work and unconfirmed endpoint evidence, Keep a successor watcher alive after Claude tears down its stop-hook process group, Load Firstmate Calm alongside standalone Pi Calm without duplicate boats or removing the replacement widget, Read Lavish feedback with captain-message counts distinct from session-ending-message counts
  • Live validation: ⚠️ inconclusive - 0 of 12 scenarios driven live against the product
Scenario Result Live Evidence
Resume a Herdr task within the default recovery wait, then refuse a holder that exceeds its bound ⏸️ untested no Checked the corrected fixture without executing its recovery lifecycle. The recorded decision explicitly excludes real Herdr suites and lifecycle commands, so an isolated lab cannot be driven within t…
Resume with --herdr-resume-lock-wait beyond the default bound and reclaim the same workspace ⏸️ untested no The corrected opt-in fixture was syntax/lint checked only. Real lock holding, restart, and workspace reclamation are expressly excluded from this turn. Provide the required Herdr CI result or authoriz…
Dispatch concurrent workers at project capacity and defer excess work without allocating resources ⏸️ untested no The authorized suite exercised admission and refusal against disposable repositories with fake tmux/treehouse endpoints. That is not live worker dispatch. A real isolated backend run requires addition…
Launch Pi workers with home-specific tool exclusions and reject invalid or unsupported exclusion lists ⏸️ untested no The suites captured generated launches and exercised registry fixtures without launching a real Pi worker. Although Pi is available, execution was restricted to the named suites. Authorize a disposabl…
Reject stale or foreign Herdr endpoint identities without steering or closing another task ⏸️ untested no The authorized ownership suite uses fake Herdr responses and process fixtures, not a live Herdr session. Real endpoint lifecycle checks remain forbidden in this turn. The pinned-Herdr CI lane or an au…
Re-arm supervision after downtime, surface durable wakes, and exclude snapshot-scoped environment ⏸️ untested no Watcher protocols were exercised through the authorized fixtures, but no real primary harness was started. The execution restriction prevents that workaround, and tmux is absent from PATH. Authorize a…
Reject another mate's echoed correlation token while accepting the owning mate's reply ⏸️ untested no The authorized pending-reply suite exercised persisted protocol fixtures and fake delivery, not two running mates. Additional live lifecycle execution is outside the recorded scope. Authorize disposab…
Relaunch a worker transactionally while preserving unfinished work and concurrent durable metadata ⏸️ untested no The suite ran real control code and disposable Git worktrees with fake terminal/harness tools. No actual worker was replaced, so this does not qualify as live. Authorize an isolated real-harness lifec…
Complete safe teardown while preserving unlanded work and unconfirmed endpoint evidence ⏸️ untested no The authorized suite exercised real disposable Git state with mocked forge and endpoint responses. It did not close a real worker endpoint. Real lifecycle execution is excluded; use the dedicated live…
Keep a successor watcher alive after Claude tears down its stop-hook process group ⏸️ untested no The recorded decision excludes the supervision-host suite and additional real-harness runs. Fixture watcher-arm coverage cannot demonstrate Claude's hook process-group teardown. Authorize the focused…
Load Firstmate Calm alongside standalone Pi Calm without duplicate boats or removing the replacement widget ⏸️ untested no The available Pi executable was identified, but the permitted execution list does not include a real Pi UI run or Calm suite. Authorize a disposable non-zero-sized PTY session with both extensions and…
Read Lavish feedback with captain-message counts distinct from session-ending-message counts ⏸️ untested no The changed public reader was inspected, but the recorded restriction permits only the eight named suites and presentation checks. None drives this reader scenario. Authorize a focused reader invocati…
  • git diff 5bff92af338d03c7b6dd0744503011e366cd94c9 --stat and targeted runtime/test diffs to derive scenarios.
  • bash -n tests/fm-backend-herdr-presentation-e2e.test.sh — passed.
  • bin/fm-lint.sh tests/fm-backend-herdr-presentation-e2e.test.sh — passed; limited to the explicitly authorized file.
  • Stopped the initial fixture invocation after approximately 15 seconds to additionally unset inherited Bitbucket email, then restarted with Herdr identity, FM_HOME, FM_BACKEND, fleet-path overrides, and forge credentials removed.
  • bash bin/fm-test-run.sh --jobs 1 --per-script-timeout-secs 600 --json ~/.no-mistakes/evidence/01M4DMESQPKH4HM8802D7T5VFZ/targeted-round3-verified-timing.json tests/fm-test-fixtures.test.sh tests/fm-watch-arm.test.sh tests/fm-control-relaunch.test.sh tests/fm-project-capacity.test.sh tests/fm-spawn-dispatch-profile.test.sh tests/fm-backend-herdr.test.sh tests/fm-teardown.test.sh tests/fm-pending-reply.test.sh — all selected suites completed successfully.
  • Confirmed the relaunch suite exercised the pre-publication rollback handshake and preserved concurrent durable metadata without reproducing the previous synchronization failure.
  • Verified disposable test scratch was removed; final git status --porcelain=v1 was empty and git diff --exit-code succeeded. Neither .test-local/ nor l.Vx4OHn/ remained.
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

ironerumi and others added 19 commits October 6, 2026 23:42
* fix(calm): share the standalone Pi Calm working-ship widget slot

Firstmate Calm and the user-global standalone Pi Calm both install an
animated working-ship widget during agent runs. Each claimed its own Pi
widget key, so a session loading both (the main Firstmate home) rendered
two boats. Pi replaces widgets under one key, so claiming the shared
"calm-working-ship" slot keeps dual-install sessions to a single boat
while a Firstmate-only session is unchanged.

Pins the shared slot contract in the working-ship module test so the key
cannot silently diverge again.

* test(calm): pin the shared working-ship widget key in CI, document dual-install

The key-parity assertion inside the Pi fixture only runs where the
@earendil-works/pi-coding-agent package is installed, so CI never
exercised it. Add a source-level twin that needs nothing but the
tracked file, and note in docs/calm.md that the boat shares the
standalone Pi Calm working-row widget slot.

* no-mistakes(review): Add executable dual-install widget replacement coverage

* no-mistakes(review): Guard shared widget cleanup with disposal ownership

* no-mistakes(document): Document shared Calm working-ship slot behavior

* test(calm): read the standalone Calm slot from its own module

The dual-install check registered both boats itself under the shared
slot, so it could only prove that Pi replaces a widget under one key: it
would still pass if the standalone Pi Calm extension installed its boat
under a different key, which is the two-boat regression the check exists
to prevent.

Read the standalone extension's own working-ship module when it is
installed - FM_STANDALONE_CALM_SHIP, else ~/.pi/agent/extensions/calm -
and drive the check with the key that module exports, so a rename on
either side registers two widgets and fails naming both keys. A pinned
shared-slot contract still covers a machine without the extension, and
the run reports which side it used instead of passing silently over an
absent extension.

Verified: the touched Pi Calm suite passes and reads the installed
standalone extension; with a copy of it whose key is renamed to
calm-working-ship-v2 the suite fails naming the drift.

* no-mistakes(review): Gate stock-row restoration by shared-widget ownership

* no-mistakes(review): Removed redundant widget-key source assertions

* no-mistakes(document): Document shared Calm working-ship widget ownership
…id#6649)

* fix(herdr): make exact-resume presentation-lock wait instead of a bounded timeout

The exact-resume path in bin/fm-spawn.sh used the same 50-attempt-then-
give-up lock acquire as the new-task-create path, but the two paths are
not equivalent on contention: a create has no prior state to strand and
can safely fall back to a flat layout, while a resume is recovering a
specific existing identity that a concurrent recovery may legitimately
be holding the lock for. Giving up there does not degrade gracefully,
it hard-fails the resume outright. The suite's own concurrent
cross-home recoveries test already asserts both concurrent recoveries
succeed with a genuine reclaim, and the file's header comment already
(inaccurately) claimed lock contention falls back to the ordinary flat
layout for both paths alike, so the intended contract was always that
recoveries serialize and both succeed, not that either one refuses
under a short bound.

Give spawn_herdr_presentation_order_lock_acquire a wait mode that uses
this file's own established fm_lock_acquire_wait idiom (already used
for its other fleet-shared locks) instead of the bounded loop, and use
it only at the exact-resume call site. The new-task-create call site
is unchanged and keeps its bounded-then-flat-fallback behavior, which
is already covered by its own passing test. Dead-owner PID-liveness
reclaim inside fm_lock_try_acquire still bounds the wait against a
holder that crashed mid-hold.

Adds a deterministic regression test that holds the shared session
lock from an unrelated process for well past the old bound, then
asserts the resume succeeds with a genuine reclaim and took close to
the full hold duration, so a fix that merely widens the bound rather
than genuinely waiting is still caught. The existing concurrent
cross-home recovery test exercises this under real timing but does not
reliably outlast a fixed bound on its own.

Corrects the header comment's claim that create and resume share one
bounded-then-flat-fallback behavior on lock contention; they no longer
do.

* no-mistakes(document): Document Herdr recovery waiting for presentation lock

* no-mistakes(document): Update stale hard-refusal claim in verification log

* no-mistakes(ci): Fixed the Greptile finding on tests/fm-backend-herdr-presentation-e2e.test.sh:1389 by bounding the resume lock-wait regression's spawn_task call. Added an optional 4th `deadline_seconds` arg to the `spawn_task` helper (defaults to empty, so all ~20 other existing call sites are unaffected and unwrapped by `timeout`). The lock-wait test now passes `LOCK_WAIT_HOLD_SECONDS + 60` (90s) as the deadline, and a dedicated check for exit code 124 emits a clear "hung for over Xs instead of waiting out a Ys lock hold" diagnostic before falling through to the existing pass/fail assertions, which are unchanged. No product code was touched. Verified with `bash -n`, `shellcheck -x` (no warnings), a standalone reproduction of the timeout/no-timeout/success paths, the project's `bin/fm-lint.sh --fast` on the file (clean), and the full `tests/fm-lint.test.sh` suite (all 46 assertions pass)

* no-mistakes(ci): Replaced the direct `timeout "$deadline_seconds"` call in `spawn_task()` (tests/fm-backend-herdr-presentation-e2e.test.sh) with the repo's portable bounded-execution helper: sourced `bin/fm-timeout-lib.sh` at the top of the file and changed `deadline_cmd=(timeout "$deadline_seconds")` to `deadline_cmd=(fm_run_timed "$deadline_seconds")`. This removes the GNU/BSD `timeout` dependency that would fail with exit 127 on a stock macOS host without coreutils, while preserving identical semantics (exit 124 on bound-hit, command's own exit otherwise), which the existing `[ "$LOCK_WAIT_STATUS" -eq 124 ]` diagnostic check already relies on. Verified: `bash -n` syntax check, `bin/fm-lint.sh --fast` clean, full `tests/fm-lint.test.sh` suite (46/46 pass), and a standalone repro confirming `fm_run_timed` returns 124 on timeout and 0 on success identically to the prior `timeout` call. No other direct `timeout` calls exist in this file or elsewhere in the PR's diff, so no sibling sites remain

* fix(herdr): gate exact-resume lock wait behind --herdr-resume-lock-wait

Keep refuse-by-default on presentation-order lock contention for Herdr
exact resume. Callers that need concurrent recoveries to serialize must
pass --herdr-resume-lock-wait; unbounded blocking on a third-party session
lock is never the default.

Update docs and the real-Herdr e2e suite so the default path asserts the
refusal and the opt-in path asserts the wait.

* no-mistakes(test): Fix e2e test's lost exit status after if/fi with no else branch

* docs(herdr): stop advertising --herdr-resume-lock-wait on --relaunch

The relaunch path reuses the recorded endpoint and never takes the
presentation-order lock, so the flag is inert there. Drop it from the
--relaunch usage line and state where the flag applies.

* no-mistakes(review): Clarify lock-wait docs; simplify bash-3.2-safe spawn_task helper

* no-mistakes(ci): Fixed ci-1 (Greptile P2). In tests/fm-backend-herdr-presentation-e2e.test.sh, the failure cleanup `cleanup_all` stopped only `LOCK_CONTENTION_OWNER_PID`. It now also stops `LOCK_REFUSE_HOLDER_PID` and `LOCK_WAIT_HOLDER_PID`, the holders of the two new contention cases, so a `fail` before their explicit `wait` no longer leaves them running. Both new PIDs are initialised empty next to the existing one, and each is cleared right after its successful `wait` so cleanup never touches a finished PID. I changed nothing else. `bash -n` passes. The real Herdr e2e run passed both new cases ("default resumed identity refuses session lock contention" and "--herdr-resume-lock-wait waits out session lock contention instead of refusing"). The full run hit my 550s timeout in a later, unrelated case, after the new cases passed
….2 (kunchenguid#6762)

* fix(bin): let TERM stop a watcher blocked in a pane capture on bash 3.2

Stock macOS bash 3.2 holds a HUP or TERM until a running command
substitution's child exits, and the watcher read every pane through
$(fm_backend_capture ...). A blocked backend read therefore held the
watcher's stop for as long as the read lasted, and a stopped watcher left
the hung read orphaned. tests/fm-watch-triage.test.sh
test_term_stops_a_watcher_blocked_inside_a_poll failed on /bin/bash 3.2
for this reason while passing on bash 5.

Pane captures now go through watcher_capture, which runs the read as a
waited background process group recorded like a check's, so the stop is
honored at once and watcher_cleanup stops a read still in flight along
with its per-call output file.

* no-mistakes(review): Run drain-ring idle capture in watcher shell, add regression test

* no-mistakes(document): Document watcher TERM handling for blocked checks and captures

* no-mistakes(test): Silence bash 3.2 setpgid race noise from watcher captures

* fix(bin): verify the capture group and scope the stop claim to pane reads

watcher_capture now confirms its background read leads its own process
group, as run_check_capture already does, so watcher_cleanup never relies
on a group that set -m failed to create. The comment and continuity doc
now say only fm_backend_capture pane reads go through watcher_capture;
agent-state and composer-state reads still run inside command
substitutions.
* feat(spawn): add per-home worker tool exclusions

Add an optional per-home config/crew-exclude-tools file listing tool names to hide from workers, one per line, with blank lines and # comments allowed.
It applies to every ship and scout launch and relaunch in that home, is never inherited by another home, and does not affect secondmate agents.
Pi and pi-signed apply it through --exclude-tools, which also covers MCP tool names.
Any other runtime, and a raw launch command, refuses the launch when the list is non-empty rather than ignoring it.
Malformed entries are refused before provisioning, and before a relaunch stops a running worker.
Exclusions that match no tool in the worker's loaded registry are reported as unverified warnings in its status record instead of refusing the worker.

Closes kunchenguid#6744

* no-mistakes(review): Preserve UTF-8 exclusion paths and verify Pi lifecycle behavior

* no-mistakes(document): Clarify worker tool exclusion documentation

* no-mistakes(ci): Fixed ci-1 in bin/fm-exclude-tools-lib.sh: a failed read now returns an error before printing names, so all shared launch and relaunch callers refuse rather than silently dropping exclusions. Added deterministic regression coverage for a file disappearing after readability checks across Pi/pi-signed ship and scout launches. Reproduced the original failure; verified 83 spawn checks, 77 relaunch checks, direct parser/runtime failure cases, full targeted lint, Bash syntax, and git diff --check. Relaunch tests passed with existing fixture-cleanup permission warnings. ci-2 remains unchanged per the user's decision; the outer executor owns the fresh CI run
kunchenguid#5343)

* refactor(bin): share the local Firstmate home walk from the wake library

Teardown's walk over the root home and its registered local secondmate homes
moves into bin/fm-wake-lib.sh as fm_local_firstmate_state_dirs, next to
fm_firstmate_root_home, so a second consumer can count task records across
this machine's homes without a copy. Teardown keeps its exact refusal wording
through a thin wrapper.

* feat(bin): defer spawns beyond a project's declared machine capacity

A project whose machine-local resource only serves a few workers at once had
no way to tell Firstmate so: every queued item was launched, and the surplus
workers spent full-context turns retrying the resource.

config/project-capacity in the root home now declares how many workers each
named project admits at once on this machine. bin/fm-spawn.sh counts the ship
and scout records on the same project origin across the root and its local
secondmate homes, skipping ones whose ready PR is recorded, while holding the
shared project lock through publication. A spawn with every place held exits 75
before any brief render, endpoint, worktree, record, or backlog move, so the
item stays queued; batches report it as deferred. Undeclared projects keep
today's uncapped dispatch, and an unreadable declaration refuses rather than
guessing the limit.

Refs kunchenguid#4237

* no-mistakes(review): Document that capacity matches the clone directory name

* no-mistakes(document): Rewrap stale fm-wake-lib root-home doc comment

* no-mistakes(review): Dedupe local state dirs by identity to avoid double-counting

* no-mistakes(document): Rewrap fm_local_firstmate_state_dirs error doc comment

* no-mistakes(ci): I fixed all four Greptile findings. All 14 tests in tests/fm-project-capacity.test.sh pass, and shellcheck at warning level is clean on the changed files. Each new test failed against the old code and passes now. - **ci-1 (spaced names):** a declaration line must give a name its capacity whenever the name is a valid clone directory name. `fm_project_capacity_lookup` now trims each line, skips blank lines and lines whose first non-blank character is `#`, and takes the last field as the capacity. Everything before that field is the name, so it may contain spaces. The old error cases still refuse: a single field is rejected, and trailing text leaves a last field that is not an integer. The library header and docs/configuration.md now say a name starting with `#` cannot be declared. New test `test_spaced_project_name_is_declared` declares `my heavy project 1` next to an indented comment line and gets a deferral. - **ci-2 (unreadable records):** the holder count must never silently leave out a holder. `fm_project_capacity_occupants` now refuses when a local home's state directory exists but cannot be read or listed, or when a `.meta` file cannot be read. The error names the path, and `fm-spawn.sh` shows it in its existing refusal message. New test `test_unreadable_holders_refuse_admission` covers an unreadable record in the root home and an unreadable state directory in a registered local secondmate home, then checks that the spawn is admitted once both are readable. The test is skipped when run as root. - **ci-3 (Orca lock):** any spawn that can become a holder for a capped project must take that project's lock. The lookup now also reports whether the declaration caps any project at all, and an Orca spawn takes the per-origin lock whenever it does. This covers every capped same-origin clone. It also covers some cases where no same-origin clone is capped, because a spawn cannot find clones under other directory names without searching for them. With no declaration file, Orca still skips the lock. The comments in the library and in the `fm-spawn.sh` header are updated. The Orca test now clones the origin as `project-2`, which has no declaration, and checks that its Orca spawn refuses while the lock is held and publishes no record. - **ci-4 (worktrees):** `assert_nothing_created` now also compares the project's `git worktree list` from before and after a deferred spawn. Both tests that call it take that snapshot first. Files changed: bin/fm-project-capacity-lib.sh, bin/fm-spawn.sh, docs/configuration.md, tests/fm-project-capacity.test.sh

* fix(bin): declare capacity for a project name that begins with #

A clone directory whose name begins with # was skipped as a comment, so that project stayed uncapped. A line is a declaration when the # is written against the rest of the name and the line ends with a capacity; a # followed by whitespace stays a comment.

* no-mistakes(document): Rewrap project-capacity library header comment

* no-mistakes(ci): Lint 2 fails because this PR's code pushes ShellCheck past its memory cap. ShellCheck ran out of memory analyzing bin/fm-teardown.sh in CI (reason=memory, rc=251, peak about 8.39 GB). On current main the same file passes at about 7.29 GB. **Cause:** the new `fm_local_firstmate_state_dirs` function in bin/fm-wake-lib.sh had a conditional `. fm-secondmate-registry-lib.sh` with a `# shellcheck source=` directive inside the function. ShellCheck followed that source again, inside a function scope, wherever fm-wake-lib.sh is sourced, and bin/fm-teardown.sh is the heaviest root that sources it. Measured locally with `shellcheck --norc --external-sources bin/fm-teardown.sh`: - current main (fd325b1): 7.29 GB - main merged with this PR: 7.86 GB - the same merge without the in-function source: 7.27 GB **Rule this restores:** this change must not make any lint root heavier than it is on main. That function holds the only new nested source in the change. **Fix:** I removed the in-function source, which no caller needs. Both callers already load the registry library at top level before calling the function: - bin/fm-teardown.sh sources it directly. - bin/fm-spawn.sh, the only user of bin/fm-project-capacity-lib.sh, gets it through bin/fm-ff-lib.sh. I also documented the requirement in the function's comment and in the "Requires" note in bin/fm-project-capacity-lib.sh. No behaviour changes. **Verification:** - ShellCheck on head: bin/fm-teardown.sh peaks at 7.12 GB and bin/fm-spawn.sh at 6.68 GB, both with rc=0. bin/fm-wake-lib.sh and bin/fm-project-capacity-lib.sh lint clean. - tests/fm-project-capacity.test.sh, tests/fm-teardown.test.sh (102 ok) and tests/fm-teardown-endpoint-safety.test.sh all pass. Files changed: bin/fm-wake-lib.sh, bin/fm-project-capacity-lib.sh

* fix(bin): release the Herdr session lock when reclaim finishes

A concurrent resume in another home waits five seconds for that lock.
Reclaim is the last presentation change on the recovery path, so holding
the lock through the launch tail made the waiter time out. The contributions
arm check also freezes its one-second clock, the same way the budget tests
do, because an unfrozen clock can tick past before the first forge read.

* no-mistakes(review): Keep Herdr session lock through launch handoff after reclaim

* no-mistakes(review): Skip the spawning task's own record in capacity count

* no-mistakes(review): Restore release test comment above its test

* docs: scope PR-ready re-evaluation to a declared project capacity

A ready pull request frees a place only when that project declares capacity, so the always-loaded backlog contract should re-evaluate on that handoff only in that case.
… section (kunchenguid#6785)

fm-procevent-lavish.sh read labels tag=message rows SESSION-ENDING MESSAGE
only when session_ended is true and CAPTAIN MESSAGE otherwise, but the
count line always said session_ending_message_count. Several composer
messages on a still-open board were therefore counted as session-ending.

The count line now follows the same session_ended switch:
session_ending_message_count once the session ended, captain_message_count
otherwise. Message rows stay out of the annotation count, per triage.

Fixes kunchenguid#6743
…henguid#6780)

The Herdr presentation lock namespace was the fixed machine-global
/tmp/firstmate-herdr-presentation, so on a host where two OS users run
Firstmate on Herdr the first account to create it owned it and every
teardown from the other account was refused with no way to clear it.

Suffix the namespace with the account uid. The owner-uid and mode-700
checks are unchanged, so a foreign-owned or wrong-mode name at this
account's path is still refused and never adopted, chowned, or removed.

Fixes kunchenguid#4716.
…te (kunchenguid#6809)

The OpenCode session plugin's shouldArm kept its own copy of the need
test that only looked for in-flight task records, while the turn-end
guard decides with fm_supervision_needed in bin/fm-supervision-lib.sh,
which also counts registered process-event sources and trusted custom
checks. With an empty fleet but any registered source or check, the
guard blocked every turn end while the plugin declined to arm - a loop
the guard's own repair line could not resolve because it names the
plugin as the fix.

The plugin now delegates the decision to the shared predicate through
bash, keeping the local away-record decline and the x-mode.env arm
override. OpenCode plugin test fixtures now carry the real predicate
their arming path sources, and the arm suite gains six cases asserting
the plugin's decision against the shared verdict over the same
synthetic state directories.

Co-authored-by: Mia Sun <mia@Bigs-Mac-mini.localdomain>
kunchenguid#6792)

* fix(bin): resolve a pending reply only from its own task's status line

Remote reply ingestion handed every corr= token in a mate's payload to
fm_pending_reply_try_resolve together with that mate's own status log,
so one mate echoing another mate's token resolved the other request.
Honor a status-file override only when it is the record's own
parent_status, and match the corr= token as a whole word.

Fixes kunchenguid#6538

* no-mistakes(document): docs: scope remote reply settlement to the asked mate
…nguid#6784)

agent-skill-trigger-index claims to be the complete agent-only trigger
index but omitted operational-home-layout, session-start-recovery,
validation-supervision, ship-landing, scout-completion, and
away-quiet-supervision. Add each with its own description's trigger,
placed beside the related entries.

The decision-hold-lifecycle redirect stub stays out, per triage.

Fixes kunchenguid#6503
…nchenguid#6814)

* test: share a rename-safe agent stand-in across liveness suites

On Ubuntu 26.04, `sleep` is the uutils multicall binary, which refuses to
run when invoked through a symlink named after another utility. The Herdr
descendant process-walk tests built their agent-named process as a `pi`
symlink to the host `sleep`, so the process exited at once, its parent shell
was gone before the walk ran, and both cases read `unknown unreadable` and
failed on that host. The suite stops at its first failure, so every later
case went unrun. The Herdr control smoke test's `claude` symlink has the
same construction.

The tmux liveness suite already solved this with a host-compiled spinner and
a survival-checked `sleep` fallback. That builder moves into tests/lib.sh as
fm_agent_standin, and the tmux suite, both Herdr descendant cases, and the
Herdr control smoke test now use it. When no stand-in can survive a foreign
name, a case skips with the reason instead of failing.

tests/fm-test-fixtures.test.sh gains a portable regression with a fake
multicall `sleep`, so it bites on hosts whose own `sleep` is single-purpose.

* no-mistakes(document): Correct Herdr verification fixture reference

* ci: retrigger cancelled shard
…he Stop hook's group is torn down (kunchenguid#6787)

* fix(bin): keep the supervision host's pass-through successor out of the hook's process group

The successor a main-only pass-through leaves for main shared the Stop hook's
process group, so the harness tearing that group down after the exit-2 rewake
stopped it. The stop published downtime and the next park's first cycle
announced an empty check: rearm-resurface, which woke main again in a loop.
Start that successor in a process group of its own, as the hook's own
handling successor already is.

* no-mistakes(review): Give the at-turn successor left for main its own group

* no-mistakes(document): Document own-group successor for turn-start hand-back too

* no-mistakes(ci): I made the change you asked for: both new teardown tests in tests/fm-supervision-host.test.sh now call the existing `stop_home_processes "$home"` just before `pass`. The tests are `test_successor_left_at_the_turn_survives_the_hook_process_group_teardown` and `test_pass_through_successor_survives_the_hook_process_group_teardown`. No production code and no other tests changed. The rule broken was that a test must not leave a home's watcher or arm processes running after it passes. These two were the only cases in the changed area that broke it. The other host+hook tests already stop their home, and `test_successor_close_during_main_turn_is_delivered_at_the_next_turn_end` leaves its watcher behind too, but it is an older test you said not to touch. The only reason anything was left over is that the successor's arm now sits in its own process group, outside the hook's teardown. `stop_home_processes` kills the watcher by the pid in its lock file, which stops it no matter which group it is in. **Checks run:** - I ran just these two tests from a scratch copy of the suite (since deleted). Both pass in about 13 seconds. - After each test, a process listing filtered to that test's home directory came back empty once the processes had about a second to exit after TERM. - `bash -n` on the test file passes. - `shellcheck` is not installed here, so I did not lint the file. - I did not run the full serial-2 suite locally. Whether it now finishes under its 30-minute limit will only show on the next CI run
…ters (kunchenguid#6823)

* test: use idle composer readiness for Claude tmux guards

* no-mistakes(test): Fix attended supervision test expectations and isolate worker state

* no-mistakes(document): Correct live guard coverage and readiness documentation

* no-mistakes(ci): Captain, fixed SC2100 by quoting the cursor-agent assignment in tests/fm-host-mirror-live-e2e.test.sh. Reproduced the failure before editing; pinned ShellCheck lint on both PR test files, bash syntax checks, and git diff --check now pass

* test: preserve attended successor close assertions

* no-mistakes(test): Fix attended live test watcher takeover expectations

* no-mistakes(document): Correct stale attended guard documentation

* Revert "no-mistakes(document): Correct stale attended guard documentation"

This reverts commit 8e59d89.

* Revert "no-mistakes(test): Fix attended live test watcher takeover expectations"

This reverts commit c0b8510.
…he downtime-write failure fixture now uses the existing quiet eligibility transition instead of emitting an extra status event that could close its successor. Added a real watcher-health assertion. The fixture invariant failed before the fix under delayed scheduling and passed three times afterward. Eight targeted cases, syntax, pinned ShellCheck lint, and git diff --check passed. The exact CI timeout was not reproduced locally. No production changes or untracked files remain
@cloud-practitioner cloud-practitioner changed the title feat: merge upstream capacity controls and runtime fixes Merge upstream firstmate into the fork Oct 8, 2026
@cloud-practitioner
cloud-practitioner merged commit afd6bd5 into main Oct 8, 2026
20 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

10 participants