Skip to content

chore: update pull request - #14

Merged
jbalke merged 189 commits into
mainfrom
fm/fm-reconcile-upstream-2026-09-30
Oct 1, 2026
Merged

jbalke merged 189 commits into
mainfrom
fm/fm-reconcile-upstream-2026-09-30

Conversation

@jbalke

@jbalke jbalke commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

Intent

The captain asked to check for updates from firstmate upstream. He was told
that the upstream remote kunchenguid/firstmate has 182 commits this fork
lacks, that the fork is 71 commits ahead, so reaching upstream is a merge and
not a fast-forward, and that a worker would merge it on a branch, resolve
conflicts, run the checks and open a PR on the fork that he merges himself.
He chose to sync now.

Measured at intake, 2026-09-30:

The ask is to bring those commits in while preserving this fork's deliberate
local behaviour. The fork's local commits are divergence on purpose, not
drift. Notable upstream themes: attended supervision for Claude and Cursor
hosts, enabled by default for Claude primaries; supervision exit latency and
watcher-cycle robustness; fm-pr-merge retry on UNKNOWN mergeable and accepting
a task's next PR; quiet-mode fixes; teardown marker cleanup; AGENTS.md
situational sections moved into on-demand skills.

What Changed

Final changed paths and statuses:

M	.agents/skills/afk/SKILL.md
A	.agents/skills/agent-skill-trigger-index/SKILL.md
M	.agents/skills/ahoy/SKILL.md
A	.agents/skills/away-quiet-supervision/SKILL.md
M	.agents/skills/bootstrap-diagnostics/SKILL.md
M	.agents/skills/captain-hold-lifecycle/SKILL.md
M	.agents/skills/firstmate-codexapp/SKILL.md
M	.agents/skills/firstmate-coding-guidelines/SKILL.md
M	.agents/skills/fmx-respond/SKILL.md
M	.agents/skills/harness-adapters/SKILL.md
M	.agents/skills/harness-adapters/references/common/control-and-recovery.md
M	.agents/skills/harness-adapters/references/common/primary-hooks.md
M	.agents/skills/harness-adapters/references/harness/claude.md
M	.agents/skills/harness-adapters/references/harness/codex.md
M	.agents/skills/harness-adapters/references/harness/cursor.md
A	.agents/skills/harness-adapters/references/harness/devin.md
M	.agents/skills/harness-adapters/references/harness/grok.md
M	.agents/skills/harness-adapters/references/harness/omp.md
M	.agents/skills/harness-adapters/references/harness/opencode.md
M	.agents/skills/harness-adapters/references/harness/pi.md
A	.agents/skills/operational-home-layout/SKILL.md
M	.agents/skills/process-event-sources/SKILL.md
M	.agents/skills/project-management/SKILL.md
M	.agents/skills/quiet/SKILL.md
M	.agents/skills/quota-array-dispatch/SKILL.md
A	.agents/skills/scout-completion/SKILL.md
M	.agents/skills/secondmate-provisioning/SKILL.md
A	.agents/skills/session-start-recovery/SKILL.md
A	.agents/skills/ship-landing/SKILL.md
M	.agents/skills/stow/SKILL.md
M	.agents/skills/stuck-crewmate-recovery/SKILL.md
A	.agents/skills/validation-supervision/SKILL.md
M	.claude/mods/firstmate-calm/.claude-plugin/plugin.json
M	.claude/mods/firstmate-calm/hooks/register.ts
A	.claude/mods/firstmate-calm/lib/fm-branch-notes.ts
M	.claude/mods/firstmate-calm/lib/fm-calm-presentation.ts
M	.claude/mods/firstmate-calm/lib/fm-operational-input.ts
A	.claude/mods/firstmate-calm/tests/branch-notes.test.ts
M	.claude/mods/firstmate-calm/tests/calm.test.ts
M	.claude/mods/firstmate-calm/tests/support.ts
M	.claude/settings.json
M	.cursor/hooks.json
M	.github/workflows/ci.yml
M	.github/workflows/no-mistakes-required.yml
M	.no-mistakes.yaml
M	.omp/extensions/fm-primary-omp-watch.ts
M	.opencode/plugins/fm-primary-watch-arm.js
M	.pi/extensions/fm-branch-supervision.ts
M	.pi/extensions/fm-calm.ts
M	.pi/extensions/fm-primary-pi-watch.ts
M	.pi/extensions/lib/fm-branch-dispatch.ts
M	.pi/extensions/lib/fm-calm-operational-user-layout.ts
A	.pi/extensions/lib/fm-calm-pending-operational-layout.ts
M	.pi/extensions/lib/fm-operational-input.ts
M	AGENTS.md
M	README.md
M	VISION.md
M	bin/backends/herdr.sh
M	bin/fm-afk-contract.sh
M	bin/fm-afk-launch.sh
M	bin/fm-afk-return.sh
M	bin/fm-afk-start.sh
M	bin/fm-agent-process-lib.sh
M	bin/fm-arm-command-policy.mjs
M	bin/fm-backend.sh
M	bin/fm-backlog-transition-lib.sh
M	bin/fm-bearings-snapshot.sh
M	bin/fm-bootstrap.sh
A	bin/fm-branch-dispatch.mjs
M	bin/fm-branch-outcome.sh
M	bin/fm-branch-prompt.sh
A	bin/fm-branch-report.sh
A	bin/fm-brief-heading-lib.sh
M	bin/fm-brief.sh
M	bin/fm-busy-lib.sh
M	bin/fm-captain-hold.sh
M	bin/fm-classify-lib.sh
M	bin/fm-claude-stop-autoarm.sh
M	bin/fm-claude-trust.sh
M	bin/fm-composer-lib.sh
M	bin/fm-config-inherit-lib.sh
M	bin/fm-contributions.jq
M	bin/fm-contributions.sh
M	bin/fm-control-lib.sh
M	bin/fm-control.sh
M	bin/fm-crew-state.sh
A	bin/fm-devin-config.sh
M	bin/fm-dispatch-resolve.sh
M	bin/fm-dod-lib.sh
M	bin/fm-ensure-agents-md.sh
M	bin/fm-ff-lib.sh
A	bin/fm-fleet-ledger.sh
M	bin/fm-fleet-snapshot.sh
M	bin/fm-fleet-sync.sh
A	bin/fm-forge-detect.sh
M	bin/fm-gate-refuse-lib.sh
A	bin/fm-git-strip-ai-trailers.sh
M	bin/fm-guard.sh
M	bin/fm-harness.sh
M	bin/fm-herdr-lab.sh
M	bin/fm-home-seed.sh
A	bin/fm-host-mirror.sh
M	bin/fm-inactive-reconcile.sh
M	bin/fm-inbox.sh
A	bin/fm-jev-mem-guard.py
A	bin/fm-jev-mem-guard.sh
A	bin/fm-lab-home.sh
M	bin/fm-lease-lib.sh
M	bin/fm-lint.sh
A	bin/fm-live-lab.sh
M	bin/fm-lock.sh
M	bin/fm-merge-authority-lib.sh
M	bin/fm-merge-local.sh
M	bin/fm-merge-outcome-lib.sh
M	bin/fm-nm-run-lib.sh
M	bin/fm-operational-input.sh
M	bin/fm-parent-channel-lib.sh
A	bin/fm-path-lib.sh
M	bin/fm-pending-reply-lib.sh
M	bin/fm-pr-check.sh
M	bin/fm-pr-lib.sh
M	bin/fm-pr-merge.sh
M	bin/fm-pr-poll.sh
M	bin/fm-primary-scope-lib.sh
M	bin/fm-procevent-lavish.sh
M	bin/fm-procevent-lib.sh
M	bin/fm-procevent-quota.sh
M	bin/fm-procevent-remote-reply.sh
M	bin/fm-procevent.sh
M	bin/fm-project-mode.sh
M	bin/fm-promote.sh
M	bin/fm-public-followup-emit.sh
M	bin/fm-public-followup-lib.sh
M	bin/fm-public-followup.sh
M	bin/fm-push-transition-lib.sh
M	bin/fm-quota-axi-lib.sh
M	bin/fm-quota-choose.sh
M	bin/fm-remote-home-provision.sh
M	bin/fm-remote-home-seed.sh
M	bin/fm-remote-inherit-push.sh
M	bin/fm-remote-inherit.sh
M	bin/fm-remote-job-lib.sh
M	bin/fm-remote-job-worker.sh
M	bin/fm-remote-secondmate-control.sh
A	bin/fm-remote-secondmate-relaunch.sh
M	bin/fm-review-diff.sh
A	bin/fm-secondmate-liveness-lib.sh
M	bin/fm-secondmate-report.sh
M	bin/fm-secondmate-restart.sh
M	bin/fm-send.sh
M	bin/fm-session-lock-lib.sh
M	bin/fm-session-start.sh
M	bin/fm-sessionstart-run.sh
M	bin/fm-spawn.sh
M	bin/fm-startup-network.sh
M	bin/fm-supervise-daemon.sh
A	bin/fm-supervision-engine-lib.sh
A	bin/fm-supervision-host.sh
M	bin/fm-supervision-instructions.sh
M	bin/fm-task-inbox-lib.sh
M	bin/fm-tasks-axi-lib.sh
M	bin/fm-tasks-axi.sh
M	bin/fm-teardown.sh
M	bin/fm-test-run.sh
M	bin/fm-timeout-lib.sh
M	bin/fm-tmux-lib.sh
M	bin/fm-tool-update-check.sh
M	bin/fm-turnend-guard-cursor.sh
M	bin/fm-turnend-guard.sh
M	bin/fm-wake-drain.sh
M	bin/fm-wake-lib.sh
M	bin/fm-watch-arm.sh
M	bin/fm-watch-checkpoint.sh
M	bin/fm-watch.sh
A	bin/fm-worker-account-lib.sh
M	bin/fm-x-dismiss.sh
M	bin/fm-x-followup.sh
M	bin/fm-x-poll.sh
M	bin/fm-x-reply.sh
M	bin/fm_voice_records.py
M	docs/agent-control.md
M	docs/architecture.md
M	docs/arm-pretool-check.md
M	docs/calm-mode-feasibility.md
M	docs/calm.md
M	docs/captain-hold-lifecycle.md
M	docs/configuration.md
M	docs/documentation-audiences.json
A	docs/fleet-ledger.md
M	docs/fm-test-portable-shards.md
A	docs/gerrit-change-watch.md
A	docs/gerrit-forge-integration.md
M	docs/gitlab-merge-watch.md
M	docs/herdr-backend.md
A	docs/jev-guards.md
M	docs/pi-supervision-branch.md
M	docs/remote-secondmates.md
M	docs/scripts.md
M	docs/secondmate-parent-channel.md
M	docs/sessionstart-nudge.md
A	docs/supervision-host.md
M	docs/supervision-protocols/claude.md
M	docs/supervision-protocols/cursor.md
M	docs/supervision-protocols/grok.md
M	docs/supervision-protocols/omp.md
M	docs/supervision-protocols/pi.md
A	docs/supervision-protocols/supervision-host.md
M	docs/tmux-backend.md
M	docs/trace-context.md
M	docs/turnend-guard.md
A	docs/verification/devin.md
M	docs/verification/dispatch-auth.md
M	docs/verification/dispatch-resolve.md
M	docs/verification/process-event-sources.md
M	docs/verification/public-followup.md
M	docs/verification/runtime-backends.md
M	docs/verification/secondmate-parent-channel.md
M	docs/verification/supervision.md
M	docs/verification/trace-context.md
M	docs/voice-relay.md
M	docs/watcher-continuity.md
M	tests/captures/no-mistakes-v1.70.1/README.md
M	tests/fixtures.sh
M	tests/fm-afk-contract.test.sh
M	tests/fm-afk-inject-e2e.test.sh
M	tests/fm-afk-inject-herdr-e2e.test.sh
M	tests/fm-afk-launch.test.sh
M	tests/fm-afk-pi-herdr-return-e2e.test.sh
M	tests/fm-afk-return.test.sh
M	tests/fm-agy-harness.test.sh
M	tests/fm-arm-pretool-check.test.sh
M	tests/fm-backend-autodetect-smoke.test.sh
M	tests/fm-backend-herdr-launcher-workspace-e2e.test.sh
M	tests/fm-backend-herdr-workspace-per-home-e2e.test.sh
M	tests/fm-backend-herdr.test.sh
M	tests/fm-backend-orca.test.sh
M	tests/fm-backend.test.sh
M	tests/fm-backlog-atomicity.test.sh
M	tests/fm-backlog-read-bound.test.sh
M	tests/fm-bearings-board-lavish-live-e2e.test.sh
M	tests/fm-bearings-board-render.test.sh
M	tests/fm-bearings-board.test.sh
M	tests/fm-bearings-snapshot.test.sh
M	tests/fm-bootstrap-network-parallel.test.sh
M	tests/fm-bootstrap.test.sh
M	tests/fm-branch-supervision.test.sh
M	tests/fm-brief.test.sh
M	tests/fm-busy-state.test.sh
M	tests/fm-calm-claude-mod-live-e2e.test.sh
M	tests/fm-calm-claude-mod-plugin.test.sh
M	tests/fm-calm-claude-mod.test.sh
M	tests/fm-calm-pi-extension.test.sh
A	tests/fm-calm-pi-queue-retention-live-e2e.test.sh
M	tests/fm-captain-hold-lifecycle.test.sh
M	tests/fm-ci-workflow.test.sh
M	tests/fm-classify-corr-token.test.sh
M	tests/fm-classify-decision-key.test.sh
M	tests/fm-claude-stop-autoarm-live-e2e.test.sh
M	tests/fm-claude-stop-autoarm.test.sh
M	tests/fm-claude-trust.test.sh
M	tests/fm-cmux-claude-composer-live-e2e.test.sh
M	tests/fm-composer-lib.test.sh
M	tests/fm-composer-matrix-live-e2e.test.sh
M	tests/fm-contributions.test.sh
M	tests/fm-control-relaunch.test.sh
M	tests/fm-control.test.sh
M	tests/fm-crew-state.test.sh
M	tests/fm-cursor-primary.test.sh
M	tests/fm-daemon.test.sh
A	tests/fm-devin-harness.test.sh
A	tests/fm-devin-signals-live-e2e.test.sh
M	tests/fm-dispatch-resolve.test.sh
A	tests/fm-dod-lib.test.sh
M	tests/fm-extension-binding.test.sh
A	tests/fm-fleet-ledger.test.sh
M	tests/fm-fleet-snapshot-view.test.sh
M	tests/fm-fleet-sync.test.sh
A	tests/fm-forge-detect.test.sh
A	tests/fm-fork-free-helpers.test.sh
M	tests/fm-gate-refuse.test.sh
M	tests/fm-gemini-harness.test.sh
A	tests/fm-git-strip-ai-trailers.test.sh
M	tests/fm-gotmp.test.sh
M	tests/fm-guard-stale-banner.test.sh
M	tests/fm-harness-precedence.test.sh
M	tests/fm-herdr-lab.test.sh
M	tests/fm-herdr-submit-confirm-live-e2e.test.sh
A	tests/fm-host-mirror-live-e2e.test.sh
A	tests/fm-host-mirror.test.sh
M	tests/fm-inactive-reconcile.test.sh
A	tests/fm-inbox.test.sh
A	tests/fm-jev-mem-guard.test.sh
M	tests/fm-kimi-harness.test.sh
A	tests/fm-launch-prompt-signals-live-e2e.test.sh
M	tests/fm-lint.test.sh
M	tests/fm-live-gate.test.sh
A	tests/fm-live-lab-up-mate.test.sh
A	tests/fm-live-lab.test.sh
M	tests/fm-mail-check.test.sh
M	tests/fm-muse-harness.test.sh
M	tests/fm-omp-harness.test.sh
M	tests/fm-on.test.sh
M	tests/fm-operational-input.test.sh
M	tests/fm-pending-reply.test.sh
M	tests/fm-pi-branch-extension.test.sh
M	tests/fm-pi-codex-native.test.sh
M	tests/fm-pi-primary-live-e2e.test.sh
M	tests/fm-pi-primary-types.test.sh
M	tests/fm-pi-watch-extension.test.sh
M	tests/fm-pr-check-security.test.sh
M	tests/fm-pr-merge.test.sh
M	tests/fm-procevent-quota.test.sh
M	tests/fm-procevent.test.sh
M	tests/fm-public-followup.test.sh
M	tests/fm-quota-choose.test.sh
M	tests/fm-remote-backlog-handoff.test.sh
M	tests/fm-remote-doctor.test.sh
M	tests/fm-remote-job.test.sh
M	tests/fm-remote-reply.test.sh
M	tests/fm-remote-secondmate-lifecycle-e2e.test.sh
M	tests/fm-remote-secondmate-parent-binding.test.sh
A	tests/fm-remote-secondmate-relaunch.test.sh
M	tests/fm-remote-secondmate-trace-context.test.sh
M	tests/fm-remote-transport-lanes.test.sh
M	tests/fm-review-diff.test.sh
M	tests/fm-rovo-harness.test.sh
M	tests/fm-secondmate-harness.test.sh
M	tests/fm-secondmate-lifecycle-e2e.test.sh
M	tests/fm-secondmate-liveness.test.sh
M	tests/fm-secondmate-reconcile.test.sh
M	tests/fm-secondmate-restart.test.sh
M	tests/fm-secondmate-safety.test.sh
M	tests/fm-secondmate-sync.test.sh
M	tests/fm-send-remote-delivery.test.sh
M	tests/fm-send-resolve-key.test.sh
M	tests/fm-session-lock-ancestry.test.sh
M	tests/fm-session-start.test.sh
M	tests/fm-sessionstart-nudge.test.sh
M	tests/fm-shared-captain-inheritance.test.sh
A	tests/fm-spawn-compact-adviser-disable-remote.test.sh
A	tests/fm-spawn-compact-adviser-disable.test.sh
M	tests/fm-spawn-dispatch-profile.test.sh
A	tests/fm-spawn-orca-worktree.test.sh
M	tests/fm-spawn-worktree-settle.test.sh
M	tests/fm-startup-memory-budget.test.sh
M	tests/fm-startup-network.test.sh
A	tests/fm-supervision-host-attended-live-e2e.test.sh
A	tests/fm-supervision-host-live-e2e.test.sh
A	tests/fm-supervision-host.test.sh
M	tests/fm-supervision-instructions.test.sh
M	tests/fm-tangle-guard.test.sh
M	tests/fm-task-delivery.test.sh
M	tests/fm-task-inbox.test.sh
M	tests/fm-tasks-axi.test.sh
M	tests/fm-teardown-endpoint-safety.test.sh
M	tests/fm-teardown.test.sh
M	tests/fm-test-run.test.sh
A	tests/fm-timeout-lib.test.sh
M	tests/fm-tmux-agent-liveness.test.sh
M	tests/fm-tool-update-check.test.sh
M	tests/fm-trace-context-spawn.test.sh
M	tests/fm-turnend-foreign-owner-repro.py
M	tests/fm-turnend-guard.test.sh
M	tests/fm-voice-relay.test.sh
M	tests/fm-wake-drain-outcome-backstop.test.sh
M	tests/fm-wake-drain-unread-status.test.sh
M	tests/fm-wake-queue.test.sh
M	tests/fm-watch-arm.test.sh
M	tests/fm-watch-checkpoint.test.sh
M	tests/fm-watch-triage.test.sh
M	tests/fm-watcher-lock.test.sh
A	tests/fm-worker-account-live-e2e.test.sh
A	tests/fm-worker-account.test.sh
M	tests/fm-x-mode.test.sh
M	tests/lib.sh
M	tests/secondmate-helpers.sh
M	tests/wake-helpers.sh

Risk Assessment

✅ Low: The fix round only changes tests and correctly addresses both accepted findings: fm_dod_block now gets its third task-dir argument, and the brief paths go through fm_test_task_dir/fm_test_task_brief, which call fm_task_data_find across every project slug. Because of that, the slightly mismatched project arguments in a few calls cannot point a presence or absence check at the wrong path.

Testing

I ran the six targeted suites for the merged areas (dod-lib, brief, task-delivery, fleet-ledger, pr-merge, teardown) and ran fm-brief.sh live in an isolated firstmate home. The scaffolds land in the fork's grouped data/tasks/<project>/<id>/ layout, with task.meta and nothing at the legacy flat path. All six suites pass with tasks-axi 0.2.6. With the machine's installed 0.2.5, one case each in pr-merge and teardown fails because the merge raised the tasks-axi minimum to 0.2.6; that is an environment upgrade, not a code defect. Only the brief-scaffold scenario was driven live; the others were covered by script tests and are recorded as untested for live validation. The temporary tasks-axi install and the FM home were removed afterwards, and no worktree files changed. There is no UI surface, so the evidence is CLI transcripts and test logs.

  • Live validation: ✅ go - 1 of 7 scenarios driven live against the product
Scenario Result Live Evidence
Captain scaffolds a ship brief and a scout brief: both land in the fork's grouped data/tasks/<project>/<id>/ with task.meta, nothing at the legacy flat data/<id>/ ✅ pass live fm-brief-grouped-layout-transcript.txt
PR-based DoD block renders when given the fork's task-data-dir signature (fixed upstream test) ⏸️ untested no The prior payload covered this only through the fm-dod-lib library test harness and did not establish a live result against a running fleet
Merged upstream brief, delivery and ledger tests read briefs from the grouped layout (branch-prefix, forge, DoD) ⏸️ untested no The prior payload covered this only through the repo's script tests in sandboxed homes and did not establish a live result against a running fleet
fm-pr-merge retries on UNKNOWN mergeable, merges once resolved, and reports pending when retries run out ⏸️ untested no The prior payload drove this only against mocked gh; a live result needs a real GitHub PR and gh credentials
Teardown retires the task's own watcher markers and leaves other tasks' markers alone ⏸️ untested no The prior payload covered this only through the teardown script test in a sandboxed home and did not establish a live result on a real fleet task
Adversarial: merge and teardown with the machine's installed tasks-axi 0.2.5 (below the new 0.2.6 minimum) ⏸️ untested no The prior payload saw this refusal only in script tests, not in a live merge or teardown. The refusal is the intended upstream version floor, and upgrading the global tasks-axi to 0.2.6 fixes it
Attended supervision for Claude and Cursor hosts on a real primary ⏸️ untested no Needs a live Claude or Cursor primary session supervising crew in tmux or herdr; this gate agent has project settings disabled and must not drive the fleet
Evidence: fm-brief grouped-layout live transcript

Source: fm-brief grouped-layout live transcript

$ FM_HOME=$H bin/fm-brief.sh sync-demo-a1 some-proj --mode no-mistakes --title "Sync demo"
scaffolded: /tmp/fmhome.OEge/data/tasks/some-proj/sync-demo-a1/brief.md (ship, mode=no-mistakes; replace {TASK} and {FIRSTMATE_SPEC})
exit=0
$ FM_HOME=$H bin/fm-brief.sh scout-demo-b2 some-proj --scout
scaffolded: /tmp/fmhome.OEge/data/tasks/some-proj/scout-demo-b2/brief.md (scout; replace {TASK} and {FIRSTMATE_SPEC})
exit=0
$ find $H/data -type f
data/tasks/some-proj/scout-demo-b2/brief.md
data/tasks/some-proj/scout-demo-b2/task.meta
data/tasks/some-proj/sync-demo-a1/brief.md
data/tasks/some-proj/sync-demo-a1/task.meta
$ cat data/tasks/*/sync-demo-a1/task.meta
project=some-proj
title=Sync demo
kind=ship
created=2026-10-01
$ grep -n "Definition of done" -A12 brief.md
41:   turn after it; continue the same stage until a defined `done:` gate under Definition of done.
42-   Use `paused: {why}` - distinct from `blocked:` - when deliberately waiting for work or an external condition expected to clear on its own, including your own validation round.
43-   Before ending your turn with your own background shell or monitor still running, or before waiting on your own pipeline run or a long foreground command, append `paused [at=<epoch>]: {job and completion condition}` to the status file.
44-   Name what you are waiting for and what will let you resume; do not repeat the declaration on every poll.
45-   Do not declare active implementation or reasoning as a wait.
46-   Firstmate may still raise one first-sight alert; the declared wait then uses the existing long recheck cadence instead of repeated possible-wedge alarms.
47-   When you know when the wait clears, include `until <YYYY-MM-DDTHH:MMZ>` (UTC) for a recheck at that time.
48-   Follow the resolution rule below when the wait clears, then resume the task.
49-   Use `blocked:` when you are stuck and need help.
50-
51-5. If you hit the same obstacle twice, append `blocked [at=<epoch>]: {why}` and stop; firstmate will help.
52-6. If a decision belongs above the implementation worker (product choices, destructive actions),
53-   append `needs-decision [at=<epoch>]: {summary of options}` and stop. Firstmate will reply with the decision.
--
96:# Definition of done
97-Delivery contract: mode=no-mistakes
98-Ship branch: fm/sync-demo-a1
99-The task is complete only when committed on your branch.
100-When you believe it is complete, append `done [at=<epoch>]: {summary}` to the status file and stop.
101-Firstmate will then instruct you to run /no-mistakes to validate and ship a PR.
$ legacy flat path exists?
"/tmp/fmhome.OEge/data/sync-demo-a1": No such file or directory (os error 2)
Evidence: fm-pr-merge run with tasks-axi 0.2.5 (version-floor failure)

Source: fm-pr-merge run with tasks-axi 0.2.5 (version-floor failure)

not ok - github-zero-exit-queue-required: refusal did not name the concrete observed state
Evidence: fm-teardown run with tasks-axi 0.2.5 (version-floor failure)

Source: fm-teardown run with tasks-axi 0.2.5 (version-floor failure)

ok - a missing teardown startup source refuses before cleanup
ok - an unreadable teardown startup source refuses before cleanup
ok - a missing adapter sibling refuses before cleanup
ok - a forced descendant with a missing adapter sibling refuses before cleanup
ok - a forced secondmate with a missing own adapter sibling refuses before child cleanup
ok - present required sources still reach the ordinary teardown refusal
ok - local-only worktree with HEAD on a fork remote is torn down and the home summary is refreshed
error: task task-x1 cannot be torn down because its backlog data directory is inaccessible: /private/var/folders/kc/90bft43x3sj0s76yp56ylh7c0000gn/T/fm-teardown-tests.5wm22u/tasks-axi-close/data (automatic backlog transitions require tasks-axi 0.2.6 or newer with the required update and mv features)
not ok - teardown failed with a real backlog
Evidence: fm-pr-merge run with tasks-axi 0.2.6 (UNKNOWN retry, all pass)

Source: fm-pr-merge run with tasks-axi 0.2.6 (UNKNOWN retry, all pass)

ok - fm-pr-merge reports exact queue retry flags after a zero-exit false success
ok - fm-pr-merge omits merge-queue retry guidance for a closed GitHub PR
ok - fm-pr-merge aggregates agreeing merge-queue rules
ok - fm-pr-merge reports ambiguity for conflicting merge-queue rules
ok - fm-pr-merge records pr= and pr_head= for a verified GitHub merge
ok - fm-pr-merge records pr= before the forge call can land the merge
ok - fm-pr-merge propagates a real merge failure without silently succeeding
ok - fm-pr-merge refuses a GitHub merge call that leaves the PR open and unqueued
ok - fm-pr-merge retries a bounded number of times when mergeable is UNKNOWN and merges once it resolves
ok - fm-pr-merge reports mergeability still pending after its bounded UNKNOWN retry is spent
ok - fm-pr-merge refuses on the UNKNOWN re-check when a check turned red between reads
ok - fm-pr-merge refuses a genuine mergeable conflict immediately, without retrying
ok - fm-pr-merge keeps PR bookkeeping when it cannot read a successful merge call's outcome
ok - fm-pr-merge refuses with the forge's own output quoted apart from its verdict
ok - fm-pr-merge quotes the forge output when it cannot read the outcome either
ok - fm-pr-merge does not echo back queue flags the caller already used
ok - fm-pr-merge still names retry flags when the caller used a different method
ok - fm-pr-merge names the queue requirement even when its method is unrecognised
ok - fm-pr-merge distinguishes unreadable branch rules from a base with no merge queue
ok - fm-pr-merge reads a plan-gated 403 on branch rules as no merge queue, not unreadable
ok - fm-pr-merge says nothing about a merge queue when the base branch has no queue rule
ok - fm-pr-merge accepts only a proved merge from the gh-axi fallback
ok - fm-pr-merge explains an armed auto-merge that landed nothing on a queue-less base
ok - fm-pr-merge never reports auto-merge as armed when the merge command failed
ok - fm-pr-merge claims no acceptance for a failed merge command carrying queue flags
ok - fm-pr-merge falls back to the gh-axi view when gh's read fails
ok - fm-pr-merge names a landed state hiding behind a failed GitHub merge command
ok - fm-pr-merge refuses a GitHub merge when gh is missing, before recording
ok - fm-pr-merge refuses a GitHub merge when gh is missing rather than merging blind
ok - fm-pr-merge verifies a genuinely merged GitHub pull request
ok - fm-pr-merge refuses to claim a merge when poll recording fails
ok - fm-pr-merge accepts and accurately reports a GitHub merge-queue entry
ok - fm-pr-merge explains how to retry with the required GitHub merge queue method
ok - fm-pr-merge refuses branch deletion unless --attended-override is passed
ok - fm-pr-merge refuses before merging when task meta is missing
ok - fm-pr-merge refuses malformed PR URLs before calling gh-axi
ok - fm-pr-merge refuses unsafe PR URL segments before recording state
ok - fm-pr-merge refuses repo override args before recording state
ok - fm-pr-merge refuses a bundled short-option repo override and refuses -d unless attended
ok - fm-pr-merge does not add default --squash when the caller passes an explicit merge method
ok - fm-pr-merge respects --method=<value> as an explicit merge method
ok - fm-pr-merge parses a GitHub PR URL into gh-axi number and --repo arguments
ok - fm-pr-merge refuses a caller --sha on GitHub because the head comes from the live read
ok - fm-pr-merge merges a GitLab merge request through glab instead of refusing it
ok - fm-pr-merge takes the GitLab instance from the URL rather than assuming one
ok - fm-pr-merge imposes no merge method on GitLab, leaving the project's own one
ok - fm-pr-merge refuses GitLab source-branch deletion unless --attended-override is passed
ok - fm-pr-merge propagates a real glab merge failure without silently succeeding
ok - fm-pr-merge refuses on each GitLab pre-merge condition independently
ok - fm-pr-merge reports every failing GitLab condition, not only the first
ok - fm-pr-merge reports a stale recorded head and verifies the live one
ok - fm-pr-merge refuses an unreadable GitLab merge request state rather than merging blind
ok - fm-pr-merge refuses a GitLab head commit it cannot validate
ok - fm-pr-merge refuses before recording anything when glab or jq is absent
ok - fm-pr-merge refuses a GitLab head override before recording state
ok - a merge a secondmate home performs itself is reported upward exactly once
ok - a locally routed secondmate home reports the landed PR into its parent's own channel
ok - a landed GitLab merge request is reported upward on the same channel
ok - a queued GitLab merge stays silent and leaves confirmation to the armed poll
ok - a refused or failed merge reports no outcome
ok - a GitLab merge refused before the forge call reports no outcome
ok - a merge a main home performs itself leaves one durable wake naming the PR
ok - a queued GitHub merge stays silent and leaves confirmation to the armed poll
ok - distinct merged PRs for one task retain distinct captain-facing wakes
ok - an uncommitted marker retry preserves at least one durable outcome
ok - a secondmate home that cannot report upward says so instead of merging in silence
ok - fm-pr-merge proceeds when the home carries no backlog at all
ok - fm-pr-merge refuses when the backlog exists but cannot be read
ok - fm-pr-merge refuses when its configured backend cannot be read
ok - fm-pr-merge refuses when its user backend configuration cannot be read
ok - fm-pr-merge refuses when its user backend configuration directory cannot be traversed
ok - fm-pr-merge proceeds once when its user configuration directory and backlog are genuinely absent
ok - fm-pr-merge honors a backend override over an unreadable user configuration
ok - fm-pr-merge refuses red GitHub checks and waives only a named --allow-red check
ok - fm-pr-merge refuses a draft pull request and one with no boolean draft state
ok - fm-pr-merge merges when a failed check run was replaced by a passing re-run
ok - fm-pr-merge never lets a check run supersede a legacy status context
ok - fm-pr-merge still refuses when a check's current run failed after an earlier pass
ok - fm-pr-merge uses start order when the old success finishes last
ok - fm-pr-merge supersedes an old cancellation that finishes last
ok - fm-pr-merge keeps a check red while its re-run is still in flight
ok - fm-pr-merge never lets one check's pass clear another check's failure
ok - fm-pr-merge clears a failure only on a proven later pass of the same check
ok - fm-pr-merge keeps --allow-red scoped to its named check beside a superseded failure
ok - fm-pr-merge rechecks away presence before an attended red merge
ok - fm-pr-merge keeps a quiet-mode home's merges attended, the named red-check waiver included
ok - fm-pr-merge accepts exactly one separately named red-check waiver
ok - while the away-posture record exists any green merge lands under away authority, yolo or not, and attended merges stay untagged
ok - under the away-posture record the branch merges a green task, is refused on a red check with or without --allow-red, and is refused at the partition while attended
ok - a branch merge refuses under the lock when the away record is archived during preflight
ok - away posture permits immediate merges but refuses every asynchronous path
ok - away merge proceeds on a plan-gated 403 because that repository cannot have a merge queue
ok - the away record does not bypass red checks, and a recorded pr= must match the URL
ok - an unreadable away-posture record refuses the merge instead of skipping the record
ok - no away-record archive or replacement lands between the authority read and the merge
ok - a record made unreadable before the merge's own authority read refuses the merge
ok - a merge that cannot lock the away record refuses instead of merging unlocked
ok - fm-pr-merge refuses --allow-red on GitLab
ok - fm-pr-merge refuses when a required check from branch protection or a ruleset never reported
ok - fm-pr-merge merges when every required check reported and is green
ok - fm-pr-merge reports a red check and an unreported required check together with every other failure
ok - fm-pr-merge refuses when the required checks cannot be read, and tells a plan without rules apart
ok - fm-pr-merge --allow-missing waives only its named unreported check
ok - fm-pr-merge --allow-missing is single use, attended-only, and GitHub-only like --allow-red
ok - fm-pr-merge enforces required producer identity and named waivers
ok - fm-pr-merge matches an app-bound required commit status by name
ok - fm-pr-merge reports known missing checks and all independent read errors
Evidence: fm-teardown run with tasks-axi 0.2.6 (marker cleanup, all pass)

Source: fm-teardown run with tasks-axi 0.2.6 (marker cleanup, all pass)

ok - a missing teardown startup source refuses before cleanup
ok - an unreadable teardown startup source refuses before cleanup
ok - a missing adapter sibling refuses before cleanup
ok - a forced descendant with a missing adapter sibling refuses before cleanup
ok - a forced secondmate with a missing own adapter sibling refuses before child cleanup
ok - present required sources still reach the ordinary teardown refusal
ok - local-only worktree with HEAD on a fork remote is torn down and the home summary is refreshed
ok - teardown closes its own backlog item before reporting success
ok - teardown honors config/backlog-backend=manual and still finishes cleanly
ok - local-only worktree with truly unpushed work is refused (safety preserved)
ok - local-only worktree with work merged into local main is torn down (no regression)
ok - no-mistakes worktree with HEAD on origin is torn down (no regression)
ok - no-mistakes worktree with genuinely unlanded work is refused (safety preserved)
ok - local-only worktree with unpushed work is torn down under --force (escape hatch)
ok - fm-pr-check publishes the PR-ready line on a secondmate's parent channel once
ok - a secondmate home's teardown delivers the child's final line or refuses until it can
ok - teardown completes when an exact busy-state sidecar is already absent
ok - herdr teardown removes pane-owned escalation dedupe state
ok - herdr flat teardown refuses before returning the isolated copy under lock contention and the retry completes cleanly
ok - herdr flat teardown never erases records when pane presence is unparseable
ok - herdr flat teardown preflight refuses before every destructive change
ok - forced secondmate teardown preflights every Herdr child before cleanup mutation
ok - forced secondmate teardown holds every descendant lifecycle and metadata lock
ok - forced secondmate teardown retains Herdr child identity until exact pane disappearance
ok - forced teardown retains a nested secondmate home and its grandchild's Herdr identity when the grandchild close is unconfirmed
ok - herdr projection teardown retires its journal only after confirming the exact recorded pane is gone
ok - herdr projection teardown retains every record when post-close presence is unknown
ok - herdr projection teardown surfaces failed focus restoration without turning confirmed cleanup into a hard failure
ok - teardown retires the task's own watcher markers and orphaned presentation journal, leaving other tasks' markers alone
ok - teardown retains a presentation journal bound to a pane other than the closed endpoint
ok - teardown retires a v1 presentation journal once its token workspace is confirmed gone
ok - teardown retains a v1 presentation journal while its token workspace is still present
ok - teardown retains a v1 presentation journal when the workspace query is ambiguous
ok - squash-merged + deleted-branch worktree (PR merged) is torn down (the fix)
ok - squash-merged PR accepts a local HEAD that is an ancestor of the final PR head
ok - teardown discovers a merged PR by branch name and tears down when no pr= was ever recorded
ok - squash-merged PR accepts replayed unpushed local patches contained in the PR head
ok - merged PR does not allow teardown after a later local commit
ok - squash-merged task whose local branch followed the pipeline rebase is torn down
ok - squash-merged same-path different content still refuses
ok - squash-merged rebased local still refuses a genuinely unlanded follow-up commit
ok - squash-merged stale local still refuses when the forge is unreachable
fm-contributions: data directory unavailable
contributions: observation not armed; coverage is unconfirmed
fm-contributions: data directory unavailable
contributions: observation not armed; coverage is unconfirmed
ok - fm-pr-check does not refresh PR head after HEAD moves
fm-contributions: data directory unavailable
contributions: observation not armed; coverage is unconfirmed
ok - fm-pr-check records the remote PR head when the local worktree lags
ok - worktree whose content already landed in the default branch is torn down (content fallback)
ok - content fallback refreshes origin default before comparing trees
ok - dirty worktree is refused even when its committed work has landed (dirty always wins)
ok - gh lookup error with content not in default refuses (fail-safe)
ok - a record predating spawn_gen refuses teardown until --legacy-record is passed
ok - a windowless leftover with no spawn_gen and no worktree tears down without --legacy-record
ok - a windowless leftover with no spawn_gen also tears down when --legacy-record is passed
ok - a windowless leftover still refuses while its worktree holds unlanded work
ok - a windowless record with a spawn_gen, a non-tmux backend or endpoint identity, no backlog validation, or ambiguous, foreign, or malformed identity still refuses
ok - a windowless leftover retries its retained legacy stamp without --legacy-record
ok - a landed legacy record with a dead endpoint tears down and logs its accepted incarnation
ok - --legacy-record never relaxes the unlanded-work refusal
ok - an endpoint that cannot be confidently read as dead refuses --legacy-record teardown
ok - --legacy-record teardown rolls its stamp back when the close marker write fails
ok - a legacy stamp a failed rollback left behind still faces the endpoint gate
ok - a corrupt spawn_gen is never accepted as a legacy record
ok - provably-stale worktree index.lock (old, no live holder) is cleared and teardown succeeds
ok - live-held worktree index.lock is never removed and teardown refuses
ok - lsof errors leave worktree index.lock in place and refuse teardown
ok - stale lock cleanup rechecks and refuses dirty worktree before return
ok - normal repo index.lock is resolved from the worktree and cleared when stale
ok - index-lock mtime fault injection is PATH-based; skipped on Darwin where stat is /usr/bin/stat
ok - transient index.lock cleared after first failed return is retried successfully without force-remove
ok - persistent index.lock exhausts retries and refuses without force-removing the lock
ok - empty retry wait overrides use the default without aborting teardown
ok - fractional legacy retry wait remains supported without arithmetic
ok - a task's own parked no-mistakes run is aborted, not orphaned, before the worker is removed
ok - a run that lands on passed-with-override after abort is still recognized as terminal
ok - a run that lands on passed-with-skips after abort is still recognized as terminal
ok - a parked run the pipeline advanced past the task copy is still concluded from the runs ledger, not orphaned
ok - a ledger row for a different head never authorizes a parked-run abort
ok - a malformed ledger row never authorizes a parked-run abort
ok - an impossible ledger date never authorizes a parked-run abort
ok - a terminal status with a stale gate never reaches ledger cleanup
ok - an advanced head present locally aborts through the strict rule alone - the ledger fallback stays dormant
ok - an unresolvable active row with no same-branch anchor is never concluded (conservative refusal)
ok - an ancestor-only anchor never binds an advanced parked run to this task
ok - a terminal unfetched-head row is stale history and never concludes a run
ok - a terminal newest row anchored at this worktree's head never authorizes an abort
ok - a resolvable diverged newer same-branch row makes every older row stale history; no run is concluded
ok - consecutive unresolvable rows are ambiguous and never conclude a run
ok - a ledger-proven continuation is still left alone while the run is autonomously active
ok - teardown refuses before reap or removal when a task-owned run remains parked
ok - a different run cannot confirm the targeted abort
ok - empty post-abort status is not accepted as confirmation
ok - the CLI's exact run-not-found signal confirms completion
ok - a parked run on another branch is never aborted by this task's teardown (ownership is precise)
ok - a task-owned autonomous running step is left alone rather than aborted
ok - a leaked descendant process rooted under the task's worktree is reaped by teardown, not left surviving
ok - a leaked descendant process rooted under the task's per-task tasktmp is reaped by teardown too
ok - missing lsof falls back to reaping the tmux pane process group
ok - an erroring lsof scan refuses teardown and preserves the task
ok - a reused pid with a different start time is never force-killed
ok - an exec change preserves birth identity and the process is reaped
ok - a process spawned during grace is reaped on a later pass
ok - persistent leaked processes refuse teardown after bounded retries
ok - a process exiting during identity lookup does not block teardown
ok - the run abort and the leaked-process reap both complete before the destructive worktree return
Evidence: fm-dod-lib run

Source: fm-dod-lib run

ok - scout done: is not gated
ok - unpushed ship done: is refused
ok - no-mistakes pre-validation done: is not gated
ok - named head on a remote-tracking ref is accepted
ok - a moved remote branch that lacks the named head is refused
ok - a 40-hex token in the note is not the named head
ok - a recorded merged PR satisfies the gate after prune
ok - the merged-PR short-circuit applies only to the recorded PR the done line names
ok - a forge-recorded head for the named PR is accepted without a local object
ok - a direct-PR recorded head does not cover a later unpushed commit
ok - no-mistakes CI-ready done: with extra text is gated
ok - keyed and spaced ship done: lines are gated
ok - local-only linked named branch is reachable from the project clone
ok - local-only detached HEAD only in the disposable copy is refused
ok - standalone local-only done: requires the named head in the project clone
ok - non-done lines are not gated
ok - fenced and indented Captain lines are not authorized intent
ok - PR-based DoD draft check uses gh-axi
all fm-dod-lib tests passed
Evidence: fm-brief run

Source: fm-brief run

ok - fm-brief: scaffolds leave the worker role scope to the launch boundary and keep the secondmate contract
ok - fm-brief.sh: bash -n succeeds
/private/var/folders/kc/90bft43x3sj0s76yp56ylh7c0000gn/T/fm-brief.jHpLUh/heredoc-in-substitution.sh:2
ok - fm-brief.sh: no heredoc is nested inside a command substitution (Bash 3.2 parse-safe)
ok - fm-brief.sh: --help renders the complete header
ok - fm-brief.sh: every ship mode generates cleanly and requires a durable report plus saved evidence
ok - fm-brief.sh: ship --mode is required and closed-set validated
ok - fm-brief.sh: the explicit ship mode wins over the registered posture
ok - fm-brief.sh: --yolo and scout/secondmate --mode are refused, never silently dropped
ok - fm-brief.sh: faster paths use configured authority without stacked review
ok - fm-brief.sh: no-mistakes DOD keeps its apostrophe prose and bans --yes outright
ok - fm-brief.sh: no-mistakes DOD detects a green PR from the drive call, not a status poll
ok - fm-brief.sh: PR-based done requires a non-draft PR; a deliberate draft declares a wait
ok - fm-brief.sh: no-mistakes ask-user findings use one event plus a verbatim snapshot
ok - fm-brief.sh: ship project-memory wording bounds edits to corrections of wrong information
ok - fm-brief.sh: context rules reach every crewmate variant and repo-artifact rules only where a repo artifact exists
ok - fm-brief.sh: --herdr-lab emits the complete hard safety contract
ok - fm-brief.sh: --herdr-lab uses its quoted Firstmate-owned helper path
ok - fm-brief.sh: ship and scout scaffolds make omitted Herdr intent fail-visible
ok - fm-brief.sh: the documented {TASK} and {FIRSTMATE_SPEC} fills cannot corrupt the Herdr safety gate
ok - fm-brief.sh: Herdr lab contract covers scouts and rejects secondmate misuse
ok - fm-brief.sh: --no-projects scaffolds a project-less charter and guards misuse
ok - fm-brief.sh: marked requests avoid generic acknowledgements and preserve material reporting
ok - fm-brief.sh: relative directory inputs ignore CDPATH, render stable absolute charter paths, or fail loudly
ok - fm-brief.sh: custom pause verb renders in every scaffold
ok - fm-brief.sh: ship and scout scaffolds teach validation-round pauses
ok - fm-brief.sh: investigation and visual-review completions load the shared decision policy
ok - fm-brief: scout and secondmate code paths still scaffold well-formed briefs
ok - fm-brief.sh: scout Lavish hosting follows the bootstrap lavish-axi floor
ok - fm-brief.sh: the home brief include lands last on ship and scout, verbatim, and fails closed
ok - fm-brief.sh: --branch-prefix omitted defaults every ship mode to fm/<task-id>
ok - fm-brief.sh: a --branch-prefix override renders identically across every generated section
ok - fm-brief.sh: an empty --branch-prefix override resolves to a bare <task-id> branch
ok - fm-brief.sh: --branch-prefix is refused on scout and secondmate scaffolds
ok - fm-brief.sh: --branch-prefix value is validated against embedded spaces and a leading dash
Switched to a new branch '$(touch${IFS}/private/var/folders/kc/90bft43x3sj0s76yp56ylh7c0000gn/T/fm-brief.jHpLUh/branch-prefix-shell-safe-marker)brief-branch-safe-g3'
ok - fm-brief.sh: ref-format-valid shell metacharacters stay literal in generated branch commands
ok - fm-brief.sh: every crewmate scaffold forbids administering the shared worktree pool
Evidence: fm-task-delivery run

Source: fm-task-delivery run

ok - fm-spawn/fm-promote: authorized intent preserves exact words and refuses operator-address lines
ok - fm-spawn: every legacy worker receives scoped role instructions without changing project or primary instructions
ok - fm-spawn: a ship spawn requires a valid explicit mode and yolo before anything is created
ok - fm-spawn: scout and secondmate spawns refuse ship delivery flags
ok - fm-spawn: the brief's recorded mode and the spawn's explicit mode must agree
ok - fm-spawn: a rigor downgrade against the registered posture is announced, never blocked
ok - fm-spawn: a scout spawn resolves no delivery posture from the registry
ok - fm-promote: promotion requires the delivery contract and records it exactly once
ok - fm-promote: a symlinked task record is refused and its target is left untouched
ok - fm-promote: a promoted worker receives the same mode-specific delivery contract a briefed one does
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
●  WATCHER DOWN - SUPERVISION IS OFF
●  1 task(s) in flight, but no watcher has a fresh beacon (last beat: never, grace 300s).
●  Trust the emitted supervision protocol for this harness; do not use shell & for watcher repair.
●  This is a supervision warning only; the guarded operation WILL still run.
●  watcher supervision needs Stop-owned automatic recovery; inspect the hook registration and startup status before ending the turn.
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
ok - fm-promote: a selected branch prefix reaches both worker instructions and durable task state
Switched to a new branch '$(touch${IFS}/private/var/folders/kc/90bft43x3sj0s76yp56ylh7c0000gn/T/fm-task-delivery.JVI3JX/promote-branch-shell-safe-marker)/promote-branch-safe-e3'
ok - fm-promote: ref-format-valid shell metacharacters stay literal in promotion branch commands
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
●  WATCHER DOWN - SUPERVISION IS OFF
●  1 task(s) in flight, but no watcher has a fresh beacon (last beat: never, grace 300s).
●  Trust the emitted supervision protocol for this harness; do not use shell & for watcher repair.
●  This is a supervision warning only; the guarded operation WILL still run.
●  watcher supervision needs Stop-owned automatic recovery; inspect the hook registration and startup status before ending the turn.
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
ok - fm-merge-local: a registry change cannot redirect an in-flight local-only task
ok - fm-project-mode: the registry lookup matches a whole multi-word name, not just its first token
ok - fm-project-mode: the conditional policy is accepted, mapped for mechanical callers, and readable raw
ok - fm-project-mode: the forge binds from its own token and is reported only through --forge
ok - fm-project-mode: only a malformed forge binding refuses; every other token keeps its old tolerance
ok - forge=gerrit: yolo is refused with its reason, never silently dropped
ok - forge=gerrit: no-mistakes runs with its forge steps skipped, recovers its fixes, then publishes one change
ok - forge=gerrit: direct-PR publishes one squashed change and a stack is refused with its reason
ok - fm-spawn: a registered forge must reach the worker's brief
ok - fm-spawn: the brief must carry the spawn's selected ship branch, and the selection is validated before anything is created
ok - fm-spawn: a ship branch that deviates from the registered prefix is announced, never blocked
ok - fm-spawn: a registry forge token the parser refuses stops the launch, reason included
ok - fm-promote: a promoted worker receives the project's registered forge contract with no flag to remember
ok - fm-spawn/fm-promote: leftover Task placeholders are refused until both subsections are filled
ok - fm-project-mode: --branch-prefix resolves order-independently and defaults to the legacy fm/ prefix
# all fm-task-delivery tests passed
Evidence: fm-fleet-ledger run

Source: fm-fleet-ledger run

ok - flag on: dispatch, polled status lines, the local merge after its task's pending lines, and cleanup are recorded in order
ok - flag on: a PR merge is recorded once, after the task's pending status lines
ok - flag on: registering a PR records task.pr_ready with its full URL after the task's pending status lines, and the merge-time re-record adds nothing
ok - flag on: a worker's status command records its line at once, and the watcher backstop does not record it again
ok - flag on, state override outside the home: the worker's status command records its line at once
ok - flag on, relative config override: a worker running elsewhere still records its line at once
ok - append failing: the worker's status command exits nonzero and records nothing
ok - ledger failing: the worker's status line still lands exactly, quietly, and the backstop records it later
ok - flag off: the worker's status command is a plain append and leaves no ledger file, offset, or lock
ok - flag off: the whole lifecycle leaves no ledger file, offset, or lock
- Outcome: ⚠️ 1 warning across 1 run (18m39s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 3 issues found → auto-fixed ✅
  • 🚨 tests/fm-dod-lib.test.sh:376 - The upstream test test_pr_based_dod_draft_check_uses_gh_axi calls fm_dod_block &#34;$mode&#34; dod-draft-task with two arguments. The merge changed the signature to &lt;mode&gt; &lt;task-id&gt; &lt;task-data-dir&gt; [branch] [forge], and fm_dod_block reads local ... task_dir=$3. The test runs under set -u, so that read fails with 'unbound variable' and aborts the whole suite before 'all fm-dod-lib tests passed' prints. A bash repro confirms the abort (exit 127). Fix: pass a task data directory, e.g. fm_dod_block &#34;$mode&#34; dod-draft-task &#34;$TMP_ROOT/dod-task&#34;.
  • 🚨 tests/fm-brief.test.sh:361 - Several newly merged upstream tests scaffold a brief with the real bin/fm-brief.sh and then read it from the legacy flat path $home/data/$id/brief.md. The fork scaffolds into data/tasks/&lt;project&gt;/&lt;id&gt;/ (fm_task_data_dir), so these reads find nothing: assert_present fails, or fill_brief_subsections writes to a missing file. Affected: tests/fm-brief.test.sh:361, 443, 1270, 1288, 1297, 1309, 1335, 1397 and 1425 (DoD non-draft, green detection, branch-prefix and forge tests), and tests/fm-task-delivery.test.sh:431, 461 and 1502-1528 (branch-prefix and forge promotion). Commit 92d197d fixed only the same problem in fm-fleet-ledger.test.sh. Fix: resolve these paths with the existing fm_test_task_brief helper, as the fork's older tests in the same files already do (e.g. fm-task-delivery.test.sh:48, 329).
  • ⚠️ bin/fm-public-followup.sh:464 - The new upstream public-followup subsystem pre-fills the report_path deliverable as data/$work_id/report.md. Its validator (bin/fm-public-followup-lib.sh:380, ^data/[A-Za-z0-9][A-Za-z0-9._-]*/report\.md$) and format hint (fm-public-followup-lib.sh:250) accept only a single path segment. In this fork, a scout scaffolded by fm-brief.sh writes its report to data/tasks/&lt;project&gt;/&lt;id&gt;/report.md. So a report-ready obligation bound to a fork task gets a pre-filled path to a file that does not exist, and the correct grouped path is refused by fm_pf_deliverable_problem. The merge adapted every other report-path consumer (fm-inactive-reconcile, fm-teardown, fm-fleet-snapshot, fm-captain-hold, backlog-transition) to the grouped layout, but not this one. The pre-fill should come from fm_task_data_dir and the validator should accept the grouped form. That requires confirming that tasks-axi's public-followup work-event accepts a multi-segment data/.../report.md path, as the fork's backlog-transition comments say its row validator does. That external contract is why this is ask-user.

🔧 Fix applied.
✅ Re-checked - no issues remain.

⚠️ **Test** - 1 warning
  • ⚠️ bin/fm-tasks-axi-lib.sh:45 - The merge raises FM_TASKS_AXI_MIN from 0.2.4 (base) to 0.2.6 (upstream). The tasks-axi installed globally on this machine is 0.2.5, so with it tests/fm-pr-merge.test.sh (github-zero-exit-queue-required) and tests/fm-teardown.test.sh ('teardown failed with a real backlog') fail. fm-captain-hold refuses with 'compatible tasks-axi is required', and a live merge or teardown would refuse the same way. Both suites pass fully with tasks-axi 0.2.6, installed into a temporary prefix for this run. Before using the merged fleet, the captain should upgrade the global tasks-axi to at least 0.2.6, e.g. npm i -g tasks-axi@0.2.6.
  • Live validation: ✅ go - 1 of 7 scenarios driven live against the product
Scenario Result Live Evidence
Captain scaffolds a ship brief and a scout brief: both land in the fork's grouped data/tasks/<project>/<id>/ with task.meta, nothing at the legacy flat data/<id>/ ✅ pass live fm-brief-grouped-layout-transcript.txt
PR-based DoD block renders when given the fork's task-data-dir signature (fixed upstream test) ⏸️ untested no The prior payload covered this only through the fm-dod-lib library test harness and did not establish a live result against a running fleet
Merged upstream brief, delivery and ledger tests read briefs from the grouped layout (branch-prefix, forge, DoD) ⏸️ untested no The prior payload covered this only through the repo's script tests in sandboxed homes and did not establish a live result against a running fleet
fm-pr-merge retries on UNKNOWN mergeable, merges once resolved, and reports pending when retries run out ⏸️ untested no The prior payload drove this only against mocked gh; a live result needs a real GitHub PR and gh credentials
Teardown retires the task's own watcher markers and leaves other tasks' markers alone ⏸️ untested no The prior payload covered this only through the teardown script test in a sandboxed home and did not establish a live result on a real fleet task
Adversarial: merge and teardown with the machine's installed tasks-axi 0.2.5 (below the new 0.2.6 minimum) ⏸️ untested no The prior payload saw this refusal only in script tests, not in a live merge or teardown. The refusal is the intended upstream version floor, and upgrading the global tasks-axi to 0.2.6 fixes it
Attended supervision for Claude and Cursor hosts on a real primary ⏸️ untested no Needs a live Claude or Cursor primary session supervising crew in tmux or herdr; this gate agent has project settings disabled and must not drive the fleet
  • bash tests/fm-dod-lib.test.sh (passes the fixed PR-based DoD draft test with the 3-argument fm_dod_block)
  • bash tests/fm-brief.test.sh (brief-path fixes: DoD, branch-prefix and forge scaffolds)
  • bash tests/fm-task-delivery.test.sh (brief-path fixes: promotion, branch-prefix and forge)
  • bash tests/fm-fleet-ledger.test.sh
  • bash tests/fm-pr-merge.test.sh with tasks-axi 0.2.5 (1 failure) and 0.2.6 (all pass)
  • bash tests/fm-teardown.test.sh with tasks-axi 0.2.5 (1 failure) and 0.2.6 (all pass)
  • Debug re-run of the failing pr-merge case to capture its stderr (root cause: tasks-axi version floor)
  • Live run in an isolated FM_HOME: bin/fm-brief.sh sync-demo-a1 some-proj --mode no-mistakes --title &#39;Sync demo&#39; and bin/fm-brief.sh scout-demo-b2 some-proj --scout, then inspected the files created
⚠️ **Document** - 1 info
  • ⚠️ bin/fm-spawn.sh:2789 - The newly merged upstream claude_add_dirs_flag gives Claude ship and scout workers --add-dir for the flat $data_real/$id path, and it also mkdirs that path. In this fork, fm-brief.sh scaffolds the brief and report into data/tasks/&lt;project&gt;/&lt;id&gt;/ (fm_task_data_dir), so that grant misses the worker's real brief and report directory. Under --permission-mode auto, the first read of the brief can park the pane on an outside-directory question. The rovo grant was already adapted to use BRIEF_DIR_REAL. The fix is a code change, outside this docs phase: pass BRIEF_DIR_REAL in the same way. docs/configuration.md:821 describes the current code (data/&lt;id&gt;), so it was left as is and should be updated together with that fix.

🔧 Fix applied.
1 info still open:

  • ℹ️ tests/fm-secondmate-harness.test.sh:721 - Applied the approved add-dir fix. In bin/fm-spawn.sh, claude_add_dirs_flag now grants BRIEF_DIR_REAL instead of the flat data/<id> and no longer creates the flat folder. docs/configuration.md:821 and the tests/fm-spawn-dispatch-profile.test.sh helper now name the grouped task folder. fm-spawn-dispatch-profile passed in full. fm-backend-orca's add-dir spawn test passed; the suite still exits 1 later, at a teardown test, on the known tasks-axi 0.2.5 problem, which was declined earlier as tasks-axi-floor-raised. fm-secondmate-harness was still running when this phase returned: every test up to B12b passed, but there is no final result, so the test phase should rerun it.
✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

Merge inventory

Merge base 9bc051f (2026-09-17), upstream head 65c75b0, fork head 66cbdd9.
182 upstream commits, 71 local commits, three-dot diff 376 files, +66,253 / -6,965.
49 files moved on both sides; 25 had textual conflicts.
Rule applied: a local change is never reverted; upstream wins only where the local hunk is provably the older version of the same idea, and each such case is named below.

Conflict resolutions

File Resolution
AGENTS.md Upstream moved the layout block into operational-home-layout; took the pointer and moved the fork's data/tasks/<project>/<id>/ lines into that skill. Fork's scout report line kept.
bin/fm-backlog-transition-lib.sh Took upstream fm_exec_timed. It already group-kills the bounded child, which is what the fork's perl setpgrp change added. Fork's report-artifact predicate kept.
bin/fm-brief.sh Both: --title plus --branch-prefix/--forge/--shape, task-id validation then branch check, grouped TASK_DIR plus brief include, upstream rule 7 (SHARED_INFRA_RULE) then fork rules 8-12, fm_dod_block gets task dir plus branch and forge.
bin/fm-dod-lib.sh fm_dod_block <mode> <id> <task-data-dir> [branch] [forge]; ask-user block takes task dir with upstream [at=<epoch>]; fm_ci_rule kept beside upstream fm_nm_driving_block.
bin/fm-promote.sh Upstream forge lookup plus fork's grouped scout brief path and CI rule comment; dod call carries task dir, branch, forge.
bin/fm-spawn.sh Upstream export prefixes collected in LAUNCH_EXPORTS and joined after the fork's env -u TRACEPARENT clear, so the clear lands on the command, never on an export. Upstream staged launch file plus fork's verified delivery.
bin/fm-composer-lib.sh Fork's Codex single-dot particle rule kept; upstream's wider braille furniture removed again; upstream Pi cost-footer rule kept; upstream footer-zone furniture routed through the particle rule under a › envelope.
bin/fm-contributions.sh Upstream budget model and helpers; fork's grouped records, safe_record_path, nameable owner kept. Fork's retry-once (cdbefaa) dropped: upstream d051e6f is the newer version of the same idea.
bin/fm-inactive-reconcile.sh Both libraries sourced; fork's grouped report pointer kept with upstream wording.
docs/captain-hold-lifecycle.md Upstream structure; fork's report-depth facts kept. tasks-axi 0.2.6 verified to accept grouped report paths on backlog rows.
docs/herdr-backend.md Upstream lists; fork's orphaned-journal re-projection kept.
docs/remote-secondmates.md Upstream structure; grouped charter brief wording kept.
docs/scripts.md, secondmate-parent-channel.md, trace-context.md, verification/runtime-backends.md Union of both rows or entries; footer furniture wording matched to the particle rule.
tests/* (9) Fork's grouped paths with upstream stamps and fixtures; procevent keeps both wait helpers; liveness keeps fork's wider server-state rows; contributions drops the retry test with the retry.

Adapted outside conflicts

  • bin/fm-live-lab.sh (new upstream): worker brief resolved through fm-task-data-lib.sh.
  • bin/fm-teardown.sh comment: task data folder wording.
  • tests trace clear assertions: the clear must lead the command after export statements; fish proof extended.
  • Gate fixes in this run: merged upstream tests adapted to the fork's fm_dod_block signature and grouped brief paths; Claude --add-dir grant now covers the grouped task data folder (BRIEF_DIR_REAL) instead of the flat data/<id>, with docs/configuration.md updated.

Environment requirement

The merge raises the tasks-axi minimum from 0.2.4 to 0.2.6 (upstream). A home running tasks-axi 0.2.5 refuses merges and teardowns until tasks-axi is upgraded to 0.2.6 or newer.

Local failures attributed

  • tests/fm-procevent.test.sh: fails on macOS because setsid is absent; fails identically on unmodified upstream.
  • tests/fm-remote-reply.test.sh: timing-sensitive recapture wait; failed once on this branch and passed on rerun, and fails on unmodified upstream on the same machine.

Follow-up work

  • public-followup report_path does not support the grouped task-data layout; needs a tasks-axi change first. tasks-axi 0.2.5 and 0.2.6 both reject a data/tasks/<project>/<id>/report.md report_path on a public-followup work event, so the upstream single-segment data/<id>/report.md behaviour is kept unchanged here.

RooseveltAdvisors and others added 30 commits September 18, 2026 11:48
…id#4854)

Captain holds have no due semantics and are a hold kind, not a Beads issue
type. The create path now waives due.required and maps to native type task.

Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(bin): launch every spawned agent with the compact adviser disabled

Every crewmate, scout, and secondmate Firstmate launches now starts with
COMPACT_ADVISER_DISABLE=1, on a fresh spawn and on a relaunch alike, so an
unattended session never activates the compact adviser.
The value is unconditional: no configuration file gates it and there is no
override, unlike the trace carrier beside it.

Three carriers deliver it, because no single one covers every launch shape.
The pane shell receives an export beside GOTMPDIR, so the agent's own children
inherit it too.
The launch command carries an explicit assignment, prepended outermost so it
wins over any ambient value the pane already held.
The cleared launch environment sets it again at the `env -i` boundary and keeps
COMPACT_ADVISER_DISABLE in the fixed operational floor, which is what preserves
the switch when config/launch-env-allowlist empties the environment, and what
delivers it on a remote host that never had the value.

bin/fm-control.sh relaunch, the bootstrap secondmate relaunch, and the remote
secondmate transport all rebuild their launch through bin/fm-spawn.sh, so they
inherit the same floor.
The captain's own primary session is untouched.

The two new suites drive the real spawn and then execute the launch command the
pane actually received, with the harness replaced by a probe that prints its own
environment, rather than matching script text.
They cover ship and secondmate launches with the allowlist absent and enabled,
the pane export and its ordering, fm-control.sh relaunch, and the full parent to
remote-host chain.

* no-mistakes(review): Export compact-adviser disable across compound launches

* no-mistakes(document): Document spawned-agent compact-adviser environment guarantee
…henguid#4894)

* fix(bin): let a background Claude session keep owning its session lock

Session-lock ownership was decided by process ancestry alone. Under an
unattended Claude session the model loop runs in a transient bg-spare
bridged to the front-end by a shared daemon; when that bridge is
recycled the contiguous claude-named ancestry from a hook to the
recorded owner breaks while the owner pid stays alive, so the Stop
auto-arm stood down as a foreign live owner, the turn-end guard ended
every turn with its read-only diagnostic, and fm-lock.sh refused - a
self-sustaining outage until restart.

Ownership is now ancestry membership OR a trusted same-session id,
never id-first:

- fm-session-lock-lib.sh accepts CLAUDE_CODE_SESSION_ID only when
  CLAUDE_PID is a Claude-shaped member of the current contiguous run,
  compares it against the id recorded in state/.lock-session, and
  requires the recorded pid to still be a live harness. No id, no
  sidecar, an untrusted id, a different id, or a dead recorded pid
  leaves the ancestry verdict unchanged. Ids are never read from ps
  argv.
- fm-lock.sh accepts a same-session holder at both refusal sites,
  writes, refreshes, and clears the sidecar only under its claim lock
  (including the early already-mine exit, skipped only while the
  deferred startup sweep leases that lock), keeps it byte-identical
  across a same-session confirmation, records CLAUDE_PID on lock line 1
  for a session with a trusted id so a shared daemon or front-end that
  outlives the session never keeps a dead session's lock alive, never
  rewrites a live line 1 on a same-session confirmation, and names the
  recorded id in the live-owner refusal.
- The .lock line-1 format is unchanged, so every reader that takes the
  whole first line as the pid keeps working; the guard's foreign-owner
  exit is unchanged and inherits the fix through the shared predicate.

Tests: the ancestry suite drives the ancestry and id signals apart in a
deterministic process table (asserting the divergence) and runs a real
orphaned front-end/daemon/pty-host/spare tree through six phases with
the real lock, auto-arm, and guard scripts; the foreign-owner repro
keeps its negative control and adds a same-id positive control.

Disclosure: no live unattended Claude background session ran on the
verifying machine. The topology is documented by the real process
listings in kunchenguid#3902, kunchenguid#2314, kunchenguid#3398, and kunchenguid#4066; coverage is the structural
predicate plus the executable fixtures, not a live pass.

Residual: bin/fm-sessionstart-nudge.sh keeps its own private ancestry
walk (it only decides whether to print a nudge) and may nudge on a
resume in the recycled case.

Out of scope, deliberately: no structured lock format, no guard budget
changes, no daemon-identity rejection, no fork lineage.

* no-mistakes(review): Wait for claim lock; revert failed sidecars

* no-mistakes(review): Revalidate ownership after wait; restore sidecars

* no-mistakes(review): Roll back sidecar by publication phase

* no-mistakes(review): Restore sidecar only if lock line is unchanged

* no-mistakes(review): Trust session ids without a spelling allowlist

* no-mistakes(review): Disarm sidecar rollback before backup cleanup

* no-mistakes(document): Updated session-lock ownership documentation
* feat: park main under the away posture on Pi

While the away-posture record exists on a Pi primary, the supervision branch
takes every actionable wake, no processing turn opens on main, captain rows
accumulate for the return brief, and main's standing authority relocates to
the branch through the existing guarded scripts.

- lib/fm-branch-dispatch.ts: read the record at every routing decision; while
  it exists claim check, decision-owned, and heartbeat rows too, keeping the
  two broken-queue vetoes; expose checkSeqs so a claimed check row lifts task
  scoping.
- fm-primary-pi-watch.ts: offer every actionable row under the record; a
  declined wake and every watcher-failure alarm still reach main.
- fm-branch-supervision.ts: drop the legacy .afk decline; append a fixed
  POSTURE: AWAY tail carrying the record's read-back verbatim per wake; open no
  processing request while the record exists, re-checked immediately before a
  request would open and at every run boundary; present the accumulated rows
  at the first run boundary after archive.
- fm-lease-lib.sh: fm_lease_forbid_branch passes the branch for opted-in
  actions only while fm-afk-contract.sh validate succeeds on a confirmed live
  record; PR merge, fresh spawn, and decision answer opt in, local landing
  never does.
- fm-send.sh: a --resolve-key naming an open needs-decision or captain-held
  task is a decision answer and meets the partition; blocked: keys stay
  steering.
- fm-spawn.sh: enforce the record's spend cap for a fresh ordinary spawn by
  either actor; relaunches and secondmates exempt.
- fm-branch-prompt.sh: fixed Postures section and the verbatim
  ask-user-authority policy; the prefix stays byte-stable.
- fm-afk-return.sh: count what the away session handled from the store.
- docs, afk skill, AGENTS.md stub: main parked on Pi, green merge gate
  absolute while away.
- tests: watcher and branch extension suites, fleet-record, merge, and
  decision-answer suites cover the relocation, the vetoes, the tail, the
  parked processing turn, the cancellation, the re-presentation, and the
  spend cap; dated live-guard evidence recorded.

* no-mistakes(review): Refuse branch merge after preflight archive race

* no-mistakes(review): Fix away wake, spawn, and processing races

* no-mistakes(review): Suppress parked processing; narrow away-only rejection

* no-mistakes(review): Abort dedicated processing; gate branch spawn once

* no-mistakes(review): Stamp away-only on the dispatch offer

* no-mistakes(review): Treat invalid away records as spend-cap absence

* no-mistakes(review): Drop spawn test hook; abort processing-opened runs

* no-mistakes(review): Bind abort to opening prompt; cap-read absence

* no-mistakes(review): Limit away branch spawn to queued work only

* no-mistakes(document): Correct AFK posture documentation
* ci: simplify CI job timeouts to a three-tier policy

Replace the scattered per-job timeout values (10m parallel, 25m lint, 30m
serial, 10m macOS) with three readable tiers, each a hang tripwire with
headroom rather than a packing estimate:

- fast (5m): coverage guard, repo invariants, timing aggregate
- normal (30m, one shared budget): lint partitions, portable parallel
  shards, portable serial shards, macOS stock Bash
- heavy (Herdr only): 20m step tripwire on the family run so always()
  cleanup still runs, under a 75m job-level last-resort backstop

The workflow's header comment states the policy and points at
docs/fm-test-portable-shards.md "Timeouts", which now owns it, and each
job names its tier beside timeout-minutes. tests/fm-ci-workflow.test.sh
asserts the policy against the parsed workflow instead of the old
per-job minute values: every job joins exactly one tier, exactly three
distinct job-level values exist, the fast tier stays within 5-10
minutes, the normal budget stays at least double the modeled parallel
lane sum reported by fm-test-run.sh --check-coverage, and the Herdr step
tripwire stays below its job backstop with an always() cleanup after it.

Concurrency supersession, shard counts, lane membership, and fail-fast
settings are unchanged.

* no-mistakes(review): Decouple the normal timeout from packing estimates

* no-mistakes(review): Assert Herdr teardown follows the family run

* no-mistakes(review): Pin Herdr family-run timeout to 20 minutes

* no-mistakes(review): Ignore comments when identifying Herdr steps

* no-mistakes(review): Identify Herdr steps by declarative ids

* no-mistakes(document): Clarify authoritative three-tier timeout policy
…nchenguid#4895)

* fix(bin): keep supervisor status closes from waking the same home

A drain that already folded OPEN DECISIONS has presented those bytes even
when the watcher has no matching seen marker. Treat that fold, and the
presentation cursor, as known so the bookkeeping close stays quiet while
later worker lines still signal.

* no-mistakes(review): Keep folded worker failures waking past supervisor closes

* no-mistakes(review): Wake on unlisted folded worker lines; batch multi-key closes

* no-mistakes(review): Stop folded worker resolved lines from counting as already read

* no-mistakes(document): Correct self-announced close marker contract in docs
* Stop steering operators away from Herdr

* no-mistakes(review): Neutralize remaining Herdr opt-out documentation wording
…enguid#4973)

* fix(bin): treat a live no-mistakes run as current after rebase

A running run on the task's branch is authoritative regardless of head.
Matching only the local head made a rebased in-flight run look failed.

* no-mistakes(review): restrict coarse live-any-head to foreign-branch answers

* no-mistakes(review): reject gate-parked runs from the executing predicate

* no-mistakes(review): hoist gate-marker patterns into single run-lib owner

* no-mistakes(review): require live daemon for head-free run binding

* no-mistakes(review): require answered daemon-down before unbinding live runs

* no-mistakes(review): extend daemon guard to anchored continuation routes

* no-mistakes(review): delete live-any-head; restore dead-daemon verdict

* no-mistakes(review): keep parked gates parked; name dead daemon everywhere

* no-mistakes(review): set dead-daemon verdict instead of emitting early

* no-mistakes(review): align selected route with legacy dead-daemon handling

* no-mistakes(review): drop unproven-record binds; narrow coarse gate reading

* no-mistakes(review): narrow header, drop vestigial guard, retarget tests

* no-mistakes(review): revert coarse gate override; require answered-down probe

* no-mistakes(review): cache one daemon probe; stop duplicating run id

* no-mistakes(review): restrict coarse dead-daemon verdict to moved-off rows

* no-mistakes(review): delete coarse dead-daemon extension and gate note

* no-mistakes(review): delete remaining coarse dead-daemon block and stale docs

* no-mistakes(document): document rebase-safe live-run bind and unverified-record verdict
…4994)

* fix(bin): stage the launch command in a private file and type a short source line

A long launch line typed while the fresh pane shell is still busy waits in the
terminal's canonical line buffer, which drops input past about 1,024 bytes on
macOS, so the pane was left at an unfinished command with no agent running.
fm-spawn now writes the assembled command to the task's own temp root under
umask 077 and types only a short line that sources it.

Refs kunchenguid#4559

* fix(bin): keep the per-task temp root private before staging the launch command

The root lives at a predictable path under /tmp and now holds the whole launch
command. Create it with mode 0700, refuse one that already exists as anything but
a directory owned by this user that nobody else can write, and tighten an owned
one, so no other local user can plant or swap the staged file.

Refs kunchenguid#4559

* fix(bin): enforce private staged launch file mode

* test(spawn): cover long staged Claude launches

* no-mistakes(review): Namespace launch files and prove truncation staging

* no-mistakes(review): Use immutable per-spawn launch filenames

* no-mistakes(document): Document staged launch delivery safeguards

* no-mistakes(ci): Updated eight behavior tests/fakes to execute or inspect immutable staged launch files instead of expecting inline launch commands. This restores Muse, secondmate lifecycle/restart, remote trace/parent binding, compact-adviser, and Orca coverage. All affected tests, dispatch-profile regression, fixture tests, syntax checks, ShellCheck, and git diff checks pass

---------

Co-authored-by: Vytautas Stankus <svycka@gmail.com>
* Add isolated Herdr runbook to test instructions

* no-mistakes(review): Drop substring matching from test.instructions contract

* no-mistakes(review): Assert commands.test key absence in YAML

* Drop unit-first sentence and instructions contract test

Captain-scoped follow-up on the Herdr-lab test.instructions ship:
keep the lab safety runbook only, and leave the no-mistakes contract
test focused on commands.test absence.
…uid#4873) (kunchenguid#5001)

* docs(vision): accept vendor-semantics and 9k contract-ceiling amendments (kunchenguid#4873)

Replace the pixels-of-today's-UI rule with a quarantined, version-pinned
surface-adapter exception recorded as standing debt. Cap the always-loaded
contract at 9,000 words and require prune-or-trigger before a crossing change
lands.

Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

* docs(vision): restore accepted three-sentence vendor-semantics form (kunchenguid#4873)

Replace the compressed paraphrase with the issue's accepted wording:
a named quarantined version-pinned adapter, expected to break, recorded
as standing debt that never hardens into a shared contract.

Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
…or-owed gate (kunchenguid#4974)

* fix(watch): recheck a gate awaiting a human instead of wedge-escalating it

A lane whose validation run is parked at a gate waiting on a human
decision is correctly quiet, but nothing in its status line says so: the
evidence is the pipeline's own gate state rather than anything the worker
wrote. The wedge timer read that silence as a suspected wedge and climbed
the escalation ladder for as long as the wait lasted, and each escalation
cost a supervising turn. The landed declared-wait consult does not reach
it, because a live ordinary crewmate never reports a declared pause, and
raising FM_STALE_ESCALATE_SECS would delay genuine wedge detection for
every lane by the same amount.

The threshold now reads a second, independent record when the status line
accounts for nothing: whether the crew's current state is a gate whose
answer is owed by a human. That is minted only from the gate's own
findings table, by a row whose `action` column is exactly `ask-user`,
located by position out of the table header the way nm_gate_step_row
already reads its row - never searched for over the run payload, where a
finding's free-text description or a branch name satisfies a search just
as well. A gate awaiting the CREWMATE's own answer keeps the unchanged
escalation schedule, reason and demand-deep-inspection wording, because a
crewmate that goes quiet before answering its own gate is exactly the
wedge the ladder exists to catch.

Each kind of wait now carries the human it is on, the action that clears
it, and whether that human is the captain as data alongside the verdict,
rather than as wording chosen per branch where the recheck is written, so
the deferral cannot word one kind of wait as another and a new kind
cannot ship without deciding all of them. A parked gate has no written
record of when its wait began, so its recheck publishes no wait age at
all rather than one read from the quiet window this deferral resets on
every pass, which would report the same small number for a gate of any
age. Like every other captain-facing recheck here it is absorbed in
silence while the away-posture record exists, arming no throttle, so the
recheck is owed in full the moment the record is archived.

The consult runs only in the at-threshold branch that was about to
escalate, beside the worktree walk already there, and only for lanes
whose status line explained nothing.

Closes kunchenguid#3055

* no-mistakes(review): require an unanswered decision before deferring a parked gate

* no-mistakes(review): reset the away-silenced timer, fail-safe findings parse, US-joined wait records

* test(watch): pass the pane hash wedge_timer_check now takes

Upstream gave wedge_timer_check a sixth <pane-hash> argument for its
dead-record probe. The malformed-wait-record rounds drive the real function
directly, so they pass one, and stub fm_backend_agent_state to a live agent so
the probe that runs after a refused deferral keeps the unchanged ladder rather
than reading a backend the child shell has none of.

* no-mistakes(review): Bind parked-gate wait to its run, owe it firstmate

* no-mistakes(document): correct wait-kind count, crew-state reader scope, gate-key coupling

* feat(watch): make the parked-gate wait deferral opt-in

The wedge timer deferring a lane parked at a validation gate is new
supervision behaviour rather than a restored one, and it decides which
lanes give up the escalation ladder, so it now ships as a default-off
per-home option instead of changing every home on upgrade.

config/wedge-defer-parked-gate arms it. The flag is read before the
decision fold, so an unconfigured home spends no fold or current-state
read, writes no record, and keeps the unchanged escalation schedule,
reasons and demand-deep-inspection wording; a test counts the reader
calls in both directions to pin that.

It is not inherited by secondmate homes: each home supervises its own
crew and owns that trade separately, the same reason
config/turnend-churn-absorb is home-local.

The away-posture absorb returns to leaving the idle timer alone, which
it had restarted only because the costly consult could reach it. A
parked-gate wait is owed to the supervisor rather than the captain, so
it never enters that branch, and the recheck owed on return is again
owed in full the moment the record is archived.

* test(watch): pin that the away-silenced hold leaves the idle timer alone

The absorb no longer restarts the timer, so the recheck owed on return is
owed in full rather than a cadence into the return. Nothing asserted
that, so a restart could be reintroduced silently.

* no-mistakes(review): document away-silence rationale, pin captured gate component

* no-mistakes(test): anchor gate row scan to the braced findings header

* no-mistakes(document): pin same-block gate row invariant in crew-state comment
…uid#5007)

* fix(control): let the owning seat reclaim a task whose endpoint is gone

A destroyed pane or workspace made `missing` a terminal state. Relaunch
accepted only `dead` and said to stop the agent first; exit refused
`missing` and said to reconcile the task first; there is no reconcile
verb. Each command named the other as its prerequisite, so a task whose
terminal went away could not be reclaimed by anything, and a no-mistakes
approval it was parked on had no seat left to answer it.

`missing` is agent-free a fortiori: there is no endpoint, so there is no
agent in it. Widen the existing guards rather than add a verb.

- fm-spawn --relaunch accepts a positively proven `missing` and creates
  one fresh endpoint in the recorded worktree; the record it already
  republishes rebinds the task to it. A `dead` endpoint is still adopted
  in place.
- fm-control exit reports `endpoint-gone` instead of dying, so the
  relaunch transaction's stop step no longer dead-ends, and re-resolves
  the endpoint from the record before verifying the replacement.

The duplicate-agent refusal is untouched: both verdicts come from the
same recovery-grade classifier, which claims `missing` only from positive
absence, so `alive`, `ambiguous`, and `unreadable` all still refuse. The
backends' own create paths refuse a live same-labeled endpoint as a
second independent guard. The worktree, its branch, commits, uncommitted
changes, armed poll and registration, record rows, and status log are all
untouched - a reclaim is a recovery, never a teardown.

A secondmate is excluded: its gone-endpoint recovery already has one
owner in the session-start liveness sweep, so relaunch refuses and names
it rather than becoming a second path to the same outcome.

Tests reproduce both halves of the deadlock, the reclaim succeeding,
unlanded work surviving it, and the refusals that still hold.

* no-mistakes(review): prove endpoint absence per backend before reclaim rebinds

* no-mistakes(review): give exit and relaunch one absence proof; pin herdr rebind session

* no-mistakes(review): narrow endpoint reclaim to herdr; tmux refuses honestly

* no-mistakes(review): stop refusals and docs asserting unestablished causes

* no-mistakes(review): stop herdr fixture helper losing tmp-root registration

* no-mistakes(review): document workspace drift and absence-probe server residue

* no-mistakes(review): correct rebind limitation to its one reachable case

* no-mistakes(review): stop claiming reclaim leaves instructions untouched

* no-mistakes(document): scope fm-control-lib purity claim, note reclaim coverage

* no-mistakes(rebase): read the staged launch file in the herdr fixture

Rebasing onto main picked up kunchenguid#4994, which stages a long worker launch
command into a script and delivers the short `. '<path>'` line instead of
the literal command. The tmux fake and tests/fixtures.sh were updated for
that; the herdr fake this branch adds was written before it and still
keyed "an agent now exists on this pane" off the literal
`encode launch-brief` text, so after the rebase it never marked the
rebound pane live and the reclaim's alive-wait read `dead`.

Dereference the staged file first, exactly as the tmux fake above does.
Test-fixture only; no production path changes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(document): note reclaim placement in herdr and scripts inventories

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…3764)

* test(status): reproduce missing event emission time

* wip(status): preserve optional event emission time

* test(status): document indirect clock stub invocation

* no-mistakes(review): Preserve historical status bytes during reply recovery

* no-mistakes(test): Fix timestamped status assertions and remote fixture dependencies

* no-mistakes(review): Preserve captain regex overrides for timestamped status events

* no-mistakes(document): Clarify status event timing and publication contracts

* no-mistakes(lint): Quote literal done to satisfy ShellCheck

* no-mistakes(ci): Captain, updated .github/workflows/ci.yml to expect 19 snapshot tests instead of 18, matching the PR’s added regression. Reproduced the failure before the fix. Stock Bash 3.2.57 verification passed: parse sweep, 19 snapshot tests, 53 Bearings tests, and the public-followup regression. Workflow lint and diff checks passed

* no-mistakes(test): Preserve terminal notifications with malformed timestamp tags

* no-mistakes(test): Stamp Rovo spawn failures with emission time

* no-mistakes(document): Verify status event documentation

* no-mistakes(lint): Fix ShellCheck quoting in status emission-time tests

* no-mistakes(ci): Captain, fixed four lifecycle assertions to accept emission timestamps while preserving publication and retry checks. Reproduced the CI failure before the fix. The lifecycle suite now passes with six Beads capability skips; syntax, targeted ShellCheck, and diff checks passed

* no-mistakes(ci): Captain, fixed malformed timestamp colons hiding actionable events using shared normalization. Original bytes and unknown ages are preserved. Regression reproduced before the fix; classifier and remote-reply suites, targeted lint, syntax, and diff checks passed

* no-mistakes(review): Stamp remote escalations at call sites, drop new flag

* no-mistakes(review): Accept stamped escalation and close lines in test assertions

* no-mistakes(review): Restore reserved-key answered-note guard for stamped closes

* test(status): accept optional emission time in PR-provenance assertions

The kunchenguid#4148 provenance test landed on main with exact unstamped greps.
Parent-channel lines from this branch carry [at=<epoch>], so strip only
that tag before the same exact match. No production change.

* no-mistakes(review): Accept stamped ready signal in PR fallback scrape

* no-mistakes(review): Drop relay flag, stamp parent events at call sites

* no-mistakes(review): Stamp worker terminal-signal instructions, revert fm-on fixture

* no-mistakes(review): Accept optional stamp in live cmux drift guard

* no-mistakes(review): Restore original test invocation order in two suites

* no-mistakes(review): Strip only well-formed numeric status time tags

* no-mistakes(document): Drop stale unstamped PR-ready line spelling from channel doc

* no-mistakes(review): Stamp agy spawn-failure status lines with event time

* fix(bin): normalize status event times in-shell and freeze the budget test clock

Two paths made a status event's emission time cost more than it should.

The captain-relevance fallback piped every line through awk to drop a
well-formed `[at=<epoch>]` tag before matching, so a supervisor sweep paid a
fork per line just to prepare a regex match. Shell parameter expansion does the
same strip with no fork, and the retry-dedup scan now reuses that one helper
instead of carrying a second copy of the rule in awk. The copies had already
drifted: the shell side stripped tags from lines with no colon, which the awk
rule left whole, so a colonless line could be mistaken for one already
recorded. One definition, checked against the awk rule it replaces over the
edge cases and a 4000-line fuzz.

tests/fm-contributions.test.sh froze its fixture clock only in exhaust mode. In
hang mode the poll set DEADLINE to the real now plus a one-second budget, and
when the second ticked before the first forge call the loop broke without ever
calling gh: forge/calls was never written and the assertion failed reading a
missing file. Freezing the clock in both modes removes the dependence on wall
time; the bounded call is still cut by the real timeout, so the observation the
test asserts still starts.

Emission time stays optional on new status records, and legacy or malformed
lines keep an unknown age.

* no-mistakes(review): Stamp ask-user escalation line and fix Kimi status assertion

* no-mistakes(document): Drop stale unstamped done-line spelling from watcher docs

* test: fold emission-time snapshot coverage into the fixture case

Drop the incidental ci.yml 18-to-19 count hunk so the PR no longer
touches workflows. Keep every emission-time assertion by folding it
into test_fixture_snapshot_json.

* no-mistakes(review): replace brief date substitution with epoch placeholder; drop emitted_at_epoch

* no-mistakes(review): align untimed normalizer with epoch parser; tolerate placeholder stamp in PR scrape

* no-mistakes(review): strip undelimited at-tags; correct brief stamp header

* no-mistakes(review): normalize stamps at both captain-regex sites; restore mtime freshness

* no-mistakes(review): strip colon-bearing stamps for relevance; fix headers and test oracles

* no-mistakes(review): narrow escalation match to stamp tolerance; pin note verb

* no-mistakes(review): read note and key past colon-bearing stamps

* test(status): keep inactive reconcile assertions stamp-tolerant

These two oracles were made stamp-tolerant while resolving one of the
branch's merges from main. The rebase drops merge commits, so that
adaptation was lost and both assertions went back to matching an exact
substring that a stamped line no longer contains: the tag lands before
the colon, so "failed [key=k]: ..." is now "failed [key=k] [at=N]: ...".
Strip a well-formed tag before matching, as the branch's other oracles do.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): unstamp fold colon tests; reserve stamp width in cap

* no-mistakes(document): correct stale unstamped status-line spellings in docs

* no-mistakes(document): quote brief-test literals for lint; correct stamp-helper contract comments

* no-mistakes(ci): rename subshell-local epoch in delivery-race stub

The serialization test overrides fm_pending_reply_mark_delivered inside a
(..) subshell. Its `epoch` local collided with the same name in
status_line_at_epoch/status_stamp_line, which this branch added and this
suite now calls at top level, so ShellCheck 0.11.0 reported SC2030 and
failed Lint 2. The stub already prefixes its other locals with `pending_`
for the same reason; `epoch` was the leftover.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: ship clean Lavish host fixes

* no-mistakes(review): Fix Lavish classifications and fail-closed host loading

* no-mistakes(review): Restore Lavish host state across retries and launches

* no-mistakes(review): Preserve destination Lavish host when configuration is absent

* no-mistakes(document): Document Lavish status and host guarantees
…#5076)

* feat(afk): make the captain's away words the whole mandate

Retire the clause fields, verb list, never-set scan, refused records, and
the per-task merge-grant list from the away-posture record. The record is
now version 2: the captain's words verbatim plus expected return, spend
cap, and reach line; a version 1 record still validates, reads, and
archives so a live away window is never broken by the upgrade.

The supervision branch reads the words at the tail of every wake and acts
on them by its own judgment through the guarded scripts under standing
authority, never by analogy, holding for the return on doubt, and opens
each such outcome summary with "per your away instructions:" so the
return brief can render the words beside the session's account. While the
record exists any green merge runs under away authority (ledger tag
"away"); red merges, --allow-red, asynchronous and queued merges, and
local-only landing stay refused. The branch may file a backlog item the
words explicitly call for before dispatching it under the spend cap.

Tests drive fm-afk-contract.sh, fm-afk-launch.sh, fm-afk-return.sh, and
fm-pr-merge.sh as commands: version 2 written, version 1 read, retired
flags and subcommands refused by name, green merges landing under the
record, red and waived-red refused, the record lock still closing the
authority-read window, and the Pi away tail carrying the words.

* no-mistakes(review): carry the away read-back to the session verbatim

* no-mistakes(review): match the exact away-action marker in the return brief

* no-mistakes(review): refuse a words block truncated by a damaged line

* no-mistakes(document): Refresh away-role contract documentation
…unchenguid#5049)

* fix(bin): render the remote charter's steering-inbox path host-local

A freshly provisioned remote secondmate read a parent-home absolute
steering-inbox path in its charter - a location that exists on no route -
and spent its first turn discovering the gap and filing a blocked
decision for what was a render defect. The seed's remote-copy rewrite now
maps the inbox to the route's host-local parent-route inbox, exactly as
it already maps the reply-log path, so every mention - bare path, listing,
and handled/ acknowledgement - lands host-local.

Both rewrites also become plain assignments, because a quoted substitution
nested inside a double-quoted printf argument leaks literal quotes into
the replacement text on stock macOS bash. The lifecycle suite pins the
corrected render both directions against the real seed, provisioning,
and delivery route, sharing one fixture value between the render truth
and the delivery truth.

Closes kunchenguid#5012

* no-mistakes(document): document remote charter's host-local steering inbox
)

* feat(procevent): route worker-owned Lavish rounds

* no-mistakes(review): drop duplicate artifact field from task-owned registration

* no-mistakes(review): post worker reply once, fix ring label, keep re-arm atomic

* no-mistakes(review): keep worker board owned until terminal round acknowledged

* no-mistakes(review): refuse every retirement of an open worker-owned round

* no-mistakes(review): use real lavish reply flag, isolate reply generations

* no-mistakes(review): drop .posted marker for best-effort reply posting

* no-mistakes(review): consume staged reply after listener setup, refuse orphaned captures

* no-mistakes(review): require a reachable owner, redeliver open rounds, roll back failed re-arms

* no-mistakes(review): re-arm only to acknowledge an open round

* no-mistakes(review): conclude only a still-open terminal round

* no-mistakes(review): record the acknowledgement before retiring the board

* no-mistakes(review): retain the registration across a conclude, qualify terminal docs

* no-mistakes(document): Document worker-owned Lavish round lifecycle
…unchenguid#5107)

* fix(bin): reserve contribution observation budget

* no-mistakes(review): Strengthen slow-read regression test to exceed the poll budget
…ness JSON (kunchenguid#5103)

* feat(bin): add idempotent inbox orders, receipts, replies, and readiness

Let a caller supply a request id when publishing a captain inbox note so a
retry returns the original note instead of creating a second one, including
across the crash window between save and wake announcement. Separate saved
from announced so a failed wake is repairable without enqueueing again.
Add bounded receipts JSON with omission disclosure, a durable primary reply
against a note id, and a read-only readiness projection that can say
unknown instead of inferring liveness from a lock file.

* no-mistakes(review): fix(bin): honest inbox announce, reply cursor, and readiness verdict

* fix(bin): resolve ready from lock-holder ancestry; drop lock status --json

Remove the extra JSON surface from fm-lock.sh so its human status still
always exits zero. Have the readiness projection classify the inspected
home from the lock-holder pid via fm-harness.sh ancestry, with an explicit
FM_SUPERVISION_MODEL still winning and an unknown model when there is no
holder. Prove the yes path when that ancestry names a known harness.

* no-mistakes(review): Harden inbox announce, receipts reads, and reply sequence cursor

* no-mistakes(document): Note read-only lock inspection in scripts inventory

* no-mistakes(lint): Pass missing id argument to malformed-reply test printf

---------

Co-authored-by: cliflacata-svg <304148223+cliflacata-svg@users.noreply.github.com>
…ending text (kunchenguid#5118)

* fix(composer): stop a harness footer row from reading as a composer holding text

A harness draws its own furniture below the composer - a user statusLine, a
permission-mode hint - and the cursorless "bottom-most shape wins" rule looks
exactly there. `→` (U+2192) is Cursor's prompt glyph but ordinary text
everywhere else, so a statusLine opening with `→` was selected as a bare
composer, swallowed the hint row beneath it as wrapped input, and answered
`pending` on a visibly empty pane. `fm_task_inbox_ring` defers on exactly that
verdict, and `bin/fm-watch.sh`'s re-ring calls the same function, so the first
doorbell and every retry were skipped and the worker never saw the steer.

Measured live on 2026-09-20: three of five Claude Code 2.1.236 worker panes on
Herdr 0.8.0 had genuinely empty composers and every one of them was refused.

A separator pair that closed over a bare agent-glyph row is a proven composer
container, so the contiguous non-blank rows below its closing rule are that
composer's footer and are no longer composer candidates. The demotion is bounded
by all three of its own preconditions: a blank row ends the zone, a pair that
closed over no glyph row demotes nothing, and a shape with no separator pair at
all (Cursor's half-block rules) is untouched. Real unsubmitted text in that same
composer, including a stray SGR mouse report left by a click in the pane, still
reads `pending`.

Pinned by two portable regressions and by a new cursorless arm on the live
composer-matrix guard, which re-reads each harness's already-proven-idle pane
the way every non-tmux backend reads it and fails naming the harness and
version when that read is `pending`.

* no-mistakes(review): make composer footer-zone demotion shape-independent

* no-mistakes(review): make footer-zone demotion refuse-only and drop rescan

* no-mistakes(lint): quote probe-absent sentinel to clear ShellCheck SC2100

---------

Co-authored-by: Koen Muller <koen@catapult.nl>
…5115)

Co-authored-by: guanchengh-lgtm <271917158+guanchengh-lgtm@users.noreply.github.com>
… an unreadable runs table (kunchenguid#5114)

* fix(bin): stop misreading a no-run branch as an unreadable runs table

Defect: when `no-mistakes axi status`'s overview is truncated (a task's
own branch has zero rows among the shown ones), fm_nm_select_run's
Python fallback derived the repo identity for its direct SQLite query
from a `repo: <path>` line it expected in the overview text. The real
CLI never emits that line, truncated or not (see the genuine capture at
tests/captures/no-mistakes-v1.70.1/overview.toon, which has only
`count:`/`runs[...]:`), so the lookup always failed and reported
"unreadable runs table" for a task that simply has no run on its
branch. On a fleet with many concurrent runs, every idle-branch task
hits the truncated-overview path routinely, so this fired every few
minutes and drowned genuine unreadable/blocked verdicts in noise.

Fix: derive the repo identity from the task worktree path instead,
which is exactly the value `no-mistakes` records as a repo's
`working_path` (confirmed against the existing capped-overview test
fixtures, which already register repos by worktree path). A worktree
path that is not absolute cannot be matched and still reads as
unreadable rather than being guessed at. Also raise the reader's
SQLite busy timeout from 1s to 30s so ordinary lock contention on a
busy fleet cannot masquerade as an unreadable database.

Safety: every other verdict byte-for-byte unchanged - the repo lookup
still requires exactly one matching row (a genuinely corrupt or
mismatched repos table still reports unreadable, per the existing
`repo` failure-mode test), the branch query and row validation are
untouched, and a zero-row result for the branch still flows through
the same recursive re-parse that already turns an empty `runs[0]{...}`
table into `absent`. Added a regression test
(test_capped_overview_without_repo_line_and_no_runs_reports_absent)
that reproduces the real overview shape - capped, zero rows for the
task's branch, no `repo: ` line - and asserts the crew state falls
through to the pane/busy verdict instead of reporting unknown or
"unreadable". Full fm-crew-state.test.sh suite passes unchanged
otherwise.

* fix: recovered same-branch inventory awk misreads empty result as unreadable

fm_nm_select_run's deep SQLite reader rebuilds a `count:`/`runs[...]:`
overview and re-runs it through the same awk selection pass. When that
rebuilt inventory has zero rows for the branch, the row-matching loop never
executes, so its counters (`seen`) stay at awk's uninitialized empty string
while `expected` and `shown` are plain strings parsed from the header text.
Comparing an uninitialized value against a non-numeric string uses string
comparison, so "" != "0" is true, and the END block takes the "unreadable
runs table" branch instead of falling through to the correct "absent"
verdict for a branch with genuinely zero runs.

Coerce the affected END comparisons with `+0` so they are always numeric,
matching seen/expected/shown/total regardless of whether awk classified
them as strings or numeric strings. A truncated or genuinely malformed
inventory still differs numerically and still reports unreadable.

* no-mistakes(review): bound capped-overview inventory reader and canonicalize worktree lookup

* no-mistakes(review): match recorded repo path first, tolerate duplicate spellings

* no-mistakes(review): revert repo lookup to exact working_path match

* no-mistakes(document): note state-db inventory read under crew-state nm timeout
…ort (kunchenguid#5141)

* fix(bin): require a non-draft pull request before a PR-based done report

A PR-based ship could report done, and merge monitoring could be armed, while the pull request was still a draft. A draft cannot be merged, so the poll waited for an event that could not occur and nobody was asked to merge.

The PR-based definitions of done now require reading the pull request back from the forge and confirming it is not a draft, and a lane that deliberately holds a draft declares a wait instead of done.
bin/fm-pr-check.sh refuses to arm merge monitoring on a draft, naming the draft state, and treats an unreadable draft state as before.
The draft reading now lives in bin/fm-pr-lib.sh and bin/fm-pr-merge.sh uses it, with its refusal to merge a draft unchanged.

Closes kunchenguid#4757

* fix(review): Skip arm-time draft refusal when fm-pr-merge records metadata
* fix(bin): accept quota-axi schema 6 snapshots keyed by provider + accountKey

quota-axi 0.1.47 emits schemaVersion 6 once a provider expands to more
than one account: every provider row carries an accountKey and one
provider id may appear on several rows. fm_quota_json_valid accepted
only schema 5 with unique provider ids, so fm-dispatch-resolve.sh,
fm-quota-choose.sh, and fm-procevent-quota.sh all rejected the live
snapshot and quota-informed dispatch was dead against the current tool.

- bin/fm-quota-axi-lib.sh: the validator accepts schema 6 with
  accountKey required on every row and uniqueness on
  provider + accountKey; schema 5 keeps its exact rules. FM_QUOTA_ROW_JQ
  is the one join every consumer uses: schema 5 binds by provider alone,
  schema 6 binds to the row keyed by the candidate's Pi lane, else the
  provider's default row, else no row (unmeasured, never blocked, never
  by position or summed across accounts).
- bin/fm-quota-choose.sh: accepts schema 6 JSON and the TOON accountKey
  column, and joins through the shared function.
- bin/fm-dispatch-resolve.sh and bin/fm-procevent-quota.sh: join through
  the shared function; an expanded provider with no row for the
  candidate's account is reported as such.
- tests: schema 6 fixtures shaped like the real snapshot, each paired
  with a schema 5 case on the same path; every new case fails on the
  previous scripts and passes now.
- docs: the two sentences naming the row join describe the schema 6 key.

* no-mistakes(review): Fix native Codex quota and expanded provider watches

* no-mistakes(review): Align native Codex account matching across dispatch paths

* no-mistakes(document): Align quota documentation with account-aware snapshots

* no-mistakes(document): Align quota dispatch documentation with account matching

* fix(bin): keep CI lint and the quota watch test portable

- bin/fm-quota-axi-lib.sh: FM_QUOTA_ROW_JQ is read only by the scripts
  that source this library, so full-mode ShellCheck reported SC2034 on
  the assignment; mark it alongside the existing SC2016 disable.
- tests/fm-procevent-quota.test.sh: the schema 6 provider-watch
  assertions used rg, which CI runners do not install, so the case
  failed with 'rg: command not found' rather than on behavior; use grep
  like the rest of the file.

* no-mistakes(document): Documented schema-version account-row compatibility
* test: repair Claude live auto-arm regression

* no-mistakes(review): Assert SessionStart digest completeness within its hook_response event

* no-mistakes(document): Consolidate Claude live verification references
Roll the shared require-no-mistakes action to the tagged v1.80.1 SHA and grant pull-requests: read so the check can read PR bodies.
…nchenguid#5174)

* fix: preserve Pi watcher ownership across session replacement

* no-mistakes(document): Scope Pi predecessor retention away from omp

* no-mistakes(ci): Diagnosed all three failing checks; only one was code-caused. (ci-3, genuine) Stock macOS Bash snapshot compatibility: `tests/fm-pi-watch-extension.test.sh` failed the macOS Bash 3.2 `bash -n` parse sweep with `line 4265: unexpected EOF while looking for matching '`. I built GNU Bash 3.2.0 from source locally and reproduced it. Root cause: the PR added a comment containing an apostrophe (`// Replacement shutdown deliberately retains module 2's established arm until`) inside a quoted here-document (`<<'EOF'`) nested inside a `$(...)` command substitution. Bash 3.2 has a parser bug (fixed in later bash) where an unmatched single quote inside such a here-doc body is treated as opening a shell quote and never closed, aborting the whole file parse. The base commit parses cleanly under Bash 3.2, confirming this PR introduced the break. Minimal fix: reworded the comment to remove the apostrophe (`... retains the established module-2 arm until`), preserving meaning. Verified `bin/fm-lint.sh --list-files` (the 6 changed shell files) now all pass `/tmp/bash-3.2/bash -n`; Bash 5 also parses. (ci-1, infrastructure) Behavior portable serial 8: GitHub API shows the `Run portable serial shard 8` step conclusion=success; only `Upload portable serial shard 8 timing artifact` failed with `Failed to FinalizeArtifact ... (403) Forbidden`. This is a transient artifact-service/cancellation failure, not a test or code failure. No change. (ci-2, infrastructure) Lint 1: fetched the job log via the GitHub API; it ends with `##[error]The runner has received a shutdown signal...` then exit 143. The step was cancelled mid-run, not a ShellCheck finding. Independently ran `bin/fm-lint.sh --partition 1of2 --telemetry ...` locally with pinned ShellCheck 0.11.0 and actionlint 1.7.12: exited rc=0 (no findings). No change. The only code change is the apostrophe removal in tests/fm-pi-watch-extension.test.sh; no other files modified
…d#5236)

* fix(bin): retire windowless leftovers and stop claiming a Pi daemon teardown

Catch-up correctly refuses while a leftover task record has no status file.
Cleanup used to deadlock on those same records when they also had no spawn_gen and no window, so they lingered and wedged every later away-mode return. Teardown now treats a windowless leftover as a missing-endpoint legacy record, and stop reports that no daemon terminal was running when none was launched.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Narrow windowless teardown exception to tmux legacy leftovers

* no-mistakes(review): Validate windowless leftover identity via shared endpoint validator

* no-mistakes(review): Refuse windowless leftovers carrying other backends' endpoint identity

* no-mistakes(document): Clarify windowless teardown retry documentation

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
kunchenguid and others added 27 commits September 28, 2026 15:15
…merge (kunchenguid#6053)

* fix(bin): accept a task's next PR once fm-pr-merge confirms the bound one merged

require_recorded_pr_identity now checks fm_pr_poll_merge_already_notified for
the recorded pr= before refusing a different URL, so a task's later PR is
accepted once its earlier PR's merge is confirmed, while it keeps refusing
while the bound PR is still unmerged.

* no-mistakes(document): docs(fm-pr-merge): note next-PR accepted after bound PR merges
…6064)

* fix(bin): read a live quiet record as a present captain at the host and watcher

A quiet record left without its daemon (a quiet start that never ran or was
interrupted) was read as away by the supervision host, so it parked a present
captain's main and held captain outcomes for a return that never comes, and
the watcher and daemon silenced captain-held rechecks on record presence.

The host's posture checks, the watcher's and daemon's captain-held silencing,
and the host's outcome path (branch report, drain BRANCH OUTCOMES, relocated
branch authority, the owners' away wake note, and the Codex checkpoint bound)
now ask the record owner's away-or-quiet reading, so only an away record is
away. A live away record keeps today's behavior.

* no-mistakes(document): Correct quiet-record documentation and supervision guidance

* no-mistakes(document): Clarify quiet-record posture and captain-held rechecks

* no-mistakes(document): Clarify quiet-record posture in documentation
…kunchenguid#6043)

* fix(bin): name an in-window engine latch in the return brief and drop the false handling GAP line

The away return brief said nothing had failed after the supervision host
latched on engine errors during the window, and printed a GAP: watcher
downtime line whenever a wake was merely being handled or queued at return.

The failures section now reads the host ledger and latch record and names
the latch time, the window's engine-error count, and whether the session
is still paused or recovered. An open recovery episode is reported as
information, and as a gap only when a queued episode outlived the return
grace or the marker cannot be read.

* no-mistakes(review): Fix latch trip time, drop marker-age grace, bound error count

* no-mistakes(review): Report paused latch without ledger trip row; bound errors

* no-mistakes(review): Never report a failed probe's latch row as trip time

* no-mistakes(review): Only a retained trip row marks a pre-window latch

* no-mistakes(document): Clarify return-brief latch and watcher-gap documentation

* no-mistakes(ci): Fixed Lint 1 by marking the shared cooldown constant as used by sourcing scripts. The repository lint command and diff check pass; the return test run was stopped by a 180-second timeout after its completed cases passed

* no-mistakes(ci): Fixed the return brief so the trip time and error count come from the same initial latch row, and ledger rows before the current session’s lock boundary cannot affect its latch report. Added real-script regressions for both findings. The return test suite, repository lint, and diff check pass

* no-mistakes(ci): Fixed the return brief’s restart cutoff so it retains in-window failures, prints one line per initial-trip row, and omits zero-error count wording. Added real-script restart regressions. The return test suite, ShellCheck, and diff check pass

* no-mistakes(ci): Fixed the return brief so a recorded trip followed by recovery stays recovered, while a later pause with a lost trip append gets a separate “trip time unavailable” line. Added a real-script regression that failed before the fix. The return test suite, ShellCheck, syntax checks, and diff check pass

* no-mistakes(ci): Fixed the false second latch during recovery. A real-script regression failed before the fix and passes now; the lost-second-trip test still passes. The return test suite, ShellCheck, syntax checks, and diff check pass
* fix(calm): name the Claude Code Calm plugin fm so supervision notes read "fm: "

Claude Code labels every mod transcript line with the plugin name, so the
notes rendered as "firstmate-calm: ⚓ ...". Rename the plugin to fm, update
the live guard to assert the fm: label, and document the one-time replay for
sessions resumed across the rename.

* no-mistakes(document): Clarify Calm plugin rename in documentation
…#6037)

* feat(bin): add fm-live-lab.sh, a one-command live supervision lab builder

* fix(bin): exact lab windows, per-lab task ids, self-safe teardown

* fix(bin): target lab windows by id, stop lab descendants, add readiness tests

* fix(bin): keep Claude's auto-updater off in live labs; list fm-live-lab.sh

* fix(bin): start the lab tmux server without user config

* no-mistakes(review): Scope lab teardown to its store, root, and task ids

* no-mistakes(review): Record selected user stores at up for check and down

* no-mistakes(document): Clarify live lab documentation and remove stale narratives

* no-mistakes(ci): Fixed the CI failure by checking for an existing lab root before looking up the harness executable. The affected behavioral test and shell syntax check pass; the refusal also works with Claude absent from PATH

* no-mistakes(ci): Fixed all four Greptile findings: teardown signals only recorded lab processes and their descendants; the worker gate is in its granted task directory and its path is exposed; readiness uses current crew state; and mate and worker IDs use 12 nonce hex digits. The CLI behavior tests pass, as do shell syntax, ShellCheck, and diff checks. The Claude no-host path is unchanged

* no-mistakes(ci): Fixed the CI test’s dependence on an installed Claude binary by supplying a test-local stub. The full fm-live-lab test, shell syntax check, and diff check pass

* no-mistakes(ci): Fixed all three selected findings in bin/fm-live-lab.sh: down waits for recorded processes and escalates before cleanup, PID roots are checked against recorded start times, and Claude primary trust is rechecked after mate/worker readiness. Added behavioral tests in tests/fm-live-lab.test.sh. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass

* no-mistakes(ci): Fixed the pre-primary settle wait, worker gate instructions, unused retry variable, and teardown PID revalidation in bin/fm-live-lab.sh. Added behavioral tests in tests/fm-live-lab.test.sh. Both requested commands pass: tests/fm-live-lab.test.sh and bin/fm-lint.sh

* no-mistakes(ci): Fixed teardown to track pre-kill lab processes by PID and start time, including children orphaned when a root exits. Up now rejects an empty pane PID before calling ps. Added regression tests and a Linux-safe worker fixture. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass

* no-mistakes(ci): Fixed teardown tracking for children spawned during shutdown and made the worker fixture verify its exact window with a Linux-available shell. Both requested checks pass. The lab test takes about 66 seconds locally, so the under-one-minute target remains unmet

* no-mistakes(ci): Fixed ci-2 and ci-4 in bin/fm-live-lab.sh and tests/fm-live-lab.test.sh. Teardown now tracks identity-checked members of captured lab process groups, including children orphaned during shutdown, without signaling the caller’s group or unrelated processes. Lint passed, and the lab test passed four times

* no-mistakes(ci): Fixed teardown so an observed-empty process group is permanently dropped, preventing a reused group ID from signalling unrelated work. Added a ps-shim regression test. The lab test, lint, and diff checks pass

* no-mistakes(ci): Fixed ci-1 in bin/fm-live-lab.sh and tests/fm-live-lab.test.sh. The TERM-born-child fixture now waits until its handler is installed before calling down. Down sends SIGKILL to identity-valid survivors on every pass from pass 20 onward and includes survivor process details if it must refuse cleanup. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass locally; Linux CI remains to be verified

* no-mistakes(ci): Fixed down’s teardown wait to require two empty identity-checked scans separated by 0.5 seconds, and removed the unused test loop variable without changing the TERM-born-child test. The lab test, lint, and diff check pass locally
…unchenguid#6103)

* fix(bin): keep slow watcher cycles and preempted reply polls from breaking supervision

- fm_pending_reply_tick selects the records it has work for in one awk pass,
  so settled records cost no lock or fork and the walk no longer grows with
  the never-pruned store.
- An attached arm keeps following a live, identity-matched holder whose beacon
  went stale until the lock changes or the shared stall bound
  (fm_watcher_stall_bound), then fails with a typed stalled-holder line so the
  retry replaces the holder.
- The remote-reply adapter reports the job worker's preemption (exit 76) as a
  closed window, so the listener keeps its claim and polls again instead of
  being relaunched every watcher cycle.

* no-mistakes(document): Clarify watcher grace and attached-arm documentation
…geable is UNKNOWN (kunchenguid#6110)

* fix(bin): retry a bounded number of times when GitHub mergeable is UNKNOWN

Fixes kunchenguid#6020

bin/fm-pr-merge.sh refused a GitHub merge whenever the pull request's
mergeable field was not literally MERGEABLE. GitHub reports UNKNOWN for
a short while after a push or a base-branch change while it recomputes
mergeability, so a green, conflict-free pull request was refused as if
it could not be merged.

github_verify_mergeable now returns a distinct status when mergeable is
the only failing condition and reads UNKNOWN. The caller retries up to
5 times, 3 seconds apart (overridable in tests), re-reading and
re-checking every live condition on each attempt. Once the bound is
spent it reports mergeability as still being computed rather than
unmergeable, with the same nonzero exit as before. Every other refusal
(closed, draft, conflicting, red or missing checks, away authority,
queue protection) is unchanged and never retried.

* no-mistakes(ci): I fixed both review findings the way you asked. The full suite (`bash tests/fm-pr-merge.test.sh`) ran to completion. Its last lines showed all `ok`, and any failure would have stopped the run early. I watched the output through `tail`, so I didn't see the new test's own `ok` line directly. **ci-2 (`bin/fm-pr-merge.sh`), retry delay not validated.** What must hold: the retry wait is always a short, valid `sleep` argument, so a bad `FM_PR_GITHUB_MERGEABLE_RETRY_DELAY` can never trip `set -e` or hold the task lock for a long time. The retry loop is the only place that reads this variable. The script now reads the value once before the loop and accepts only whole numbers from 0 to 10. Anything else (empty, `abc`, `-1`, `1.5`, `11`, a huge number, leading spaces) falls back to 3. I ran those values through the check by hand and each came out as expected. The loop now sleeps on that checked value. **ci-1 (`tests/fm-pr-merge.test.sh`), no test for a check changing between UNKNOWN reads.** What must hold: every retry re-checks all live conditions, not just mergeable. The fake `gh pr view` in the test can now take an optional second word on each line of the mergeable sequence, which sets the first check's result. The new test `test_github_mergeable_unknown_retry_rechecks_checks` feeds `UNKNOWN`, then `UNKNOWN FAILURE`. It asserts: - exit code 1 after exactly 2 reads, - the refusal names `check 'ci' is not green`, - the message does not say mergeability is still being computed, - `pr merge` was never called. If a later change made the retry look only at mergeable, the loop would read UNKNOWN 5 times, end with the "still being computed" message, and this test would fail. I didn't run it against a deliberately broken script to confirm that. `bash -n` passes. `shellcheck` reports only the existing info-level notes about files it can't follow. Only `bin/fm-pr-merge.sh` and `tests/fm-pr-merge.test.sh` changed
kunchenguid#6112)

* fix(bin): converge every open owner onto a known terminal contribution

settle_final only cleared a stale error on retry, so an owner whose saved
row still said open kept projecting a merged or closed pull request as
open after another owner's row had already recorded the terminal
observation. Copy the known terminal observation to every owner whose
saved row is not itself terminal, keeping that owner's own pending and
notified state, and clear its error.

* no-mistakes(review): Carry terminal checked_at when converging existing owner rows

* no-mistakes(ci): I fixed Greptile finding ci-2 as you asked, with a change to tests/fm-contributions.test.sh only. The rule it enforces: when a retry converges an owner onto a URL that is already merged or closed, that owner gets the terminal owner's whole observation, not just its state. The same weak check appeared twice in test_interrupted_multi_owner_poll_settles_every_owner, so I fixed both: - **Open owner (line 784):** the check now also requires `.observation == $terminal[0].records[0].observation`. The existing checks for error, checked_at, pending and notified are unchanged. - **Errored owner (just below):** it only checked state and error before. It now reads the terminal owner's file and makes the same full-observation comparison. Adding the comparison alone would not have caught anything. The test fixtures gave both owners identical observations apart from `state`, so copying only the state would still have passed. In both cases I also set the terminal owner's observation head to HEAD_B, so the two observations now really differ. Verification: - The focused test passes against the current bin/fm-contributions.sh. - I temporarily changed `settle_final` so it copied only the state. The test then failed, reporting the owner still on the old head (HEAD_A). I restored the file afterwards, and `git status` shows only the test file modified. - The full tests/fm-contributions.test.sh suite exits 0. No product code changed. The other CI finding (ci-1, "Behavior portable serial 9") was left alone because you chose to ignore it
…nguid#6124)

* feat: run the supervision host by default on a Claude primary

An absent config/supervision-host on a Claude primary now reads as on with
the default engine, and a file holding `off` opts any home out. Cursor,
OpenCode, omp, Grok, and Codex stay file-gated, with `off` read as disabled
there too. Every reader asks fm_supervision_host_enabled instead of testing
the file, and non-bash readers query it through the lib's `enabled` entry.
A primary's `off` is not inherited by secondmates: each home keeps its own
supervision posture.

* test: pin the watcher-path posture in fixtures that assume no supervision host

Fixtures that drive the watcher arm or assert a non-host drain now write
an explicit off file, and fixtures that copy the Stop auto-arm or the
supervision instructions carry the engine lib they now source. The two
drain suites also stop reading the code root's config.

* fix: name the opt-out when an off home passes an attended wake to main

A host parked when the home writes off now logs that the home does not run
the supervision host, rather than claiming it has no engine.

* no-mistakes(document): Clarify Claude supervision defaults and historical evidence

* no-mistakes(ci): Fixed process leaks in the two added host tests. Each case now stops its recorded watcher and host/arm processes; fake hook sessions exit through session.stop. The full host suite passed before the final cleanup refinement, and both affected cases, bash syntax, ShellCheck, and diff checks passed afterward. CI runtime still needs confirmation
…start scope check (kunchenguid#6125)

* fix(bin): create the state dir on a fresh primary before the session-start scope check

fm_primary_scope_matches required an already-existing state directory, so
bin/fm-sessionstart-run.sh stood down on a fresh clone before anything could
create it. Split out fm_primary_root_matches so the run wrapper can confirm
primary-home identity first, create the gitignored state dir when it is
missing, and only then run the unchanged scope check.

* no-mistakes(document): Document session-start state dir creation on fresh clones

* no-mistakes(ci): I fixed the Greptile P1 the way you asked. When a fresh primary can't create `state/`, the run wrapper no longer stands down silently. **Invariant:** when an otherwise eligible fresh primary cannot create `state/`, startup must never fail silently. This path has only one site: the mkdir in `bin/fm-sessionstart-run.sh`. Other hooks and the nudge wrapper never create `state/`, so they have no equivalent failure. **What changed:** - **Run wrapper** (`bin/fm-sessionstart-run.sh`): it captures mkdir's error and prints one line to stderr before standing down as before (exit 0, or 3 for the Pi prerequisite). The line looks like `fm-sessionstart-run: startup could not create the state directory <path>: <reason>`. - **Test** (`tests/fm-sessionstart-nudge.test.sh`): the new case `test_run_reports_a_state_dir_it_cannot_create` uses a fresh primary with no `state/` and a read-only (0500) root. It checks four things: exit 0, no digest on stdout, no state dir created, and exactly one stderr line ending in "Permission denied". It fails without the fix and passes with it. - **Docs** (`docs/sessionstart-nudge.md`): I added one sentence describing the stderr line and one describing what the new test proves. **Verification:** I ran `tests/fm-sessionstart-nudge.test.sh`, and every test passes. `bin/fm-lint.sh` on the changed scripts (pinned ShellCheck 0.11.0) and `tests/fm-documentation-audiences.test.sh` also pass. As you asked, the wrapper still stands down with the ineligible-checkout status afterwards. It does not report this as a failed eligible startup, which is what the bot suggested
…ery (kunchenguid#6126)

* fix(bin): measure pending-reply grace from turn completion, not delivery

Fixes kunchenguid#6057

The pending-reply guard demanded a repost ("REPOST REQUIRED: previous
marked request had no correlated parent report") while the second
mate's correlated reply was already on its way.
fm_pending_reply_send_recovery measured its grace window from delivery
instead of from the request turn's completion, so any turn longer than
the grace fired the demand the moment the turn ended, before the reply
could have landed. The missed-report escalation had the same gap: it
fired the instant the recovery turn's completion was observed, with no
grace at all.

Both now measure grace from the relevant turn's completion (request
turn for the recovery repost, recovery turn for the escalation), and
both take one fresh, uncached read of the parent status file
immediately before firing, accepting a correlated line regardless of
its verb. Transport-failure escalations stay immediate, and the
one-repost limit is unchanged.

* no-mistakes(review): Document grace window as measured from turn completion

* no-mistakes(ci): Both Greptile findings were real and caused by this PR, so I fixed them. The full `tests/fm-pending-reply.test.sh` suite passes. **ci-1 (a reply could be overwritten by a repost).** The rule that must hold: a recovery send is recorded only if the record is still unresolved, checked under the same per-correlation lock that resolution uses. The escalation path already did this (`_fm_pending_reply_maybe_escalate_locked` reads fresh and publishes under one lock). The recovery path did not: `fm_pending_reply_send_recovery` did its fresh read through `fm_pending_reply_try_resolve`, which let go of the lock before the send was recorded. A reply landing in that gap could be overwritten, and the repost would go out anyway. Now `send_recovery` takes the lock once and, while holding it, re-checks that the phase is still `awaiting_report`, runs the fresh uncached read, and records the send (sender pid and identity, attempt time, phase `recovery_sending`). It releases the lock before actually sending, so the lock is not held during the send. It uses the same lock helpers the other lock wrappers use. Grace timing, the one-repost limit and the escalation path are unchanged. **ci-2 (the test would pass even without the fix).** In `test_recovery_fresh_status_read_resolves_before_firing`, the reply is still appended to the status file, but the stored file signature is then set to the file's new signature. That stands in for a same-size rewrite that the signature cache cannot see. The test first checks that a normal cached read misses the reply, then that the fresh read before sending catches it. I also added the same check for the fresh read before escalation, which the review said was uncovered. The test now sets its own send hook, so it no longer depends on one left over from an earlier test (that leftover had made failures exit silently). **Checks:** - I removed the fresh-read bypass at each site in turn and reran the suite. With it gone from recovery, the test fails with "recovery must not fire once a correlated reply has landed". With it gone from escalation, it fails with "the fresh pre-escalation read should have resolved the record, got escalated". With both in place, all tests pass. - Shellcheck with `-x` timed out locally. Without `-x` and ignoring SC1091, the only warnings are SC2034 on the existing `maybe_escalate` lock wrapper, which is not part of this change. The new code adds no warnings. Changes are in `bin/fm-pending-reply-lib.sh` and `tests/fm-pending-reply.test.sh`. Nothing is committed yet; a plain commit message such as "fix(bin): record the pending-reply recovery send under the fresh-read lock" fits the instruction

* no-mistakes(ci): ci-1 was real and caused by this PR. The same bug was also in the escalation path, so both are fixed. The full tests/fm-pending-reply.test.sh suite passes. The rule that must hold: a recovery repost or an escalation goes out only if the record's phase, read after the fresh-read resolve, is still what it was before. The resolver writes phase=resolved first and only then writes the other resolution fields. If one of those later writes fails, it returns an error even though the record is already resolved. Places this rule applies, both fixed: - Recovery (fm_pending_reply_send_recovery): the fresh-read resolve now runs first, and the phase is re-read right after it, whatever it returned. The send is recorded and made only if the phase is still exactly awaiting_report. This replaces the earlier phase check rather than adding a second one. - Escalation (_fm_pending_reply_maybe_escalate_locked): same bug. After a failed resolve it went on to publish the blocked line and set phase=escalated. One added line after the resolve call returns 1 without publishing if the phase has changed. Test: added test_partial_resolve_write_blocks_firing. It forces a failure on the resolved_epoch write after a correlated reply has landed. It checks that the recovery send hook is never called, that no escalation line is published, and that the phase stays resolved. The forced failure runs in a subshell so it can't affect later tests. Checks: - With the recovery fix reverted, the new test fails with "recovery must not fire after a partial resolve". - With the escalation fix reverted, it fails with "partial resolve should block escalation, got escalated". - With both fixes in, every test passes. - Shellcheck was run with SC1091 excluded and without -x, not through the repo's lint script. The only new message is one SC2329 info on the test's override function; other test overrides in the same file already get that same info, unsuppressed. Changed files: bin/fm-pending-reply-lib.sh and tests/fm-pending-reply.test.sh. Nothing is committed. Suggested plain commit message: "fix(bin): recheck pending-reply phase after the fresh read before sending
…to stderr (kunchenguid#6001)

* fix: provider-table lookup never writes a broken-pipe error to stderr

Fixes kunchenguid#5956

fm_quota_single_provider_for_harness returned from its while read loop
as soon as it found a match, closing the pipe while
fm_quota_single_provider_table's printf could still be writing.
Where SIGPIPE is ignored, as on GitHub Actions runners, bash then
prints "printf: write error: Broken pipe" on the resolver's stderr,
which intermittently broke the one-diagnostic-line assertions in
tests/fm-dispatch-resolve.test.sh.

Read the whole table before answering, the way
fm_control_harness_supported already does, so the writer always
finishes. Return values and output are unchanged.

Reproduced by running tests/fm-dispatch-resolve.test.sh with SIGPIPE
ignored on a single pinned core under CPU contention: 30 of 30 runs
failed before the fix, 0 of 30 after. Note: reproducing requires
setting the trap inside the tested shell because nice(1) resets an
inherited SIGPIPE ignore to SIG_DFL. tests/fm-quota-choose.test.sh
passes and bin/fm-lint.sh is clean.

* no-mistakes(ci): Fixed both Greptile findings the user chose to address. ci-1 (bin/fm-quota-axi-lib.sh:154). Invariant: looking up a harness must always end with status 0 and print the provider, even when the caller runs under `set -e`. The loop body `[ -z "$found" ] && [ "$harness" = "$1" ] && found=$provider` now ends in `|| :`. Every iteration succeeds and the whole table is still read. Only `fm_quota_single_provider_for_harness` loops over the table this way, so this is the one place the fix was needed. One caveat: on bash 5.3 the old code did not actually exit under `set -e`, because the `while` loop is not the function's last command, so the new `set -e` test would have passed before this fix too. The change makes the loop's success explicit, as the user asked. ci-2 (regression coverage). I added three cases to the existing `tests/fm-quota-choose.test.sh`, all calling the public lookup function after sourcing the library: 1. With SIGPIPE ignored (`trap "" PIPE`), it looks up every harness 200 times and checks that nothing reaches stderr. 2. A deterministic version of the race: the table function is wrapped so it writes the first row, pauses 0.2 s, then writes the rest. With SIGPIPE ignored, it checks that looking up `claude` prints `claude` and writes nothing to stderr. The stress loop alone reproduced the bug in only about 1 of 5 local runs, which is why this case exists. 3. A direct call under `set -e` prints `claude`. Verification: - `bash tests/fm-quota-choose.test.sh`: all pass. - Same test against the pre-PR library (fa48367, via `FM_ROOT_OVERRIDE`): fails with `printf: write error: Broken pipe`. The deterministic case failed in one run and the stress loop caught it in another. - `shellcheck` on both files: clean. - `tests/fm-dispatch-resolve.test.sh`: passes
…isioning (kunchenguid#6162)

* fix: survive Pi 0.99 rendering and Git 2.55 local-clone races

Pi 0.99 puts arguments on the stock tool header and leaves hidden custom messages in the export conversation column. Match that header, and keep Calm's boundary on the visible column. Clone a remote home with --no-local so a prune during Git's loose-object copy cannot fail the seed.

* no-mistakes(review): Stop SIGPIPE write errors; cover older Pi export and project clones

* no-mistakes(document): Clarify Calm export visibility and tool rendering

* no-mistakes(ci): Fixed the dispatch diagnostic to list every provider-less use/default profile in one line and added a multi-profile behavior test. Shortened supervision fixtures using the existing engine-grace and park-clock knobs; removed stray scratch files. Dispatch tests, syntax checks, and three targeted supervision cases passed. CI’s prior supervision duration was 751s; the single permitted local full-suite run timed out at 1200s, so an after-duration is not established. The cancelled serial check had no failure verdict. The outer executor should record the measured before/after duration in the PR body when available

* no-mistakes(review): Gate Pi 0.99 call headers by version; drop hidden-row assertion

* no-mistakes(review): Test stock call headers under Pi 0.87 and 0.99 stubs

* no-mistakes(test): Fix older-Pi queued-row test and verify park-boundary behavior

* no-mistakes(document): Clarify Pi Calm export and queued-turn documentation

* no-mistakes(ci): Fixed the stock macOS Bash 3.2 parse failure in tests/fm-calm-pi-extension.test.sh; its parse check passes. The watcher CI failure is in unchanged code: the isolated five-minute/66-minute case passes locally, but the CI log omits the drain error needed to establish its cause. No speculative watcher fix was made. The full local watcher suite timed out after 500 seconds
…6169)

* Prevent premature Lavish board handoffs

* Prove Lavish arm lacks reply acknowledgement

* Confirm Lavish replies before arming worker boards

* no-mistakes(review): Post Lavish reply only after locked arm eligibility checks

* no-mistakes(review): Fail Lavish reply closed on unknown version

* no-mistakes(document): Correct Lavish reply documentation and remove stale guidance

* no-mistakes(document): Clarify Lavish reply routing and remove duplicate version guidance
…#6154)

* feat: inherit the supervision-host opt-out from the primary

Move the supervision host's off opt-out out of config/supervision-host into
its own presence flag, config/supervision-host-off, and add that flag to the
primary-authoritative inherited config set. A primary that opts out now opts
every secondmate home out at spawn and convergence, and clearing it converges
them back. config/supervision-host stays the home-local engine choice.

Shape: config/supervision-host mixed two things, a fleet posture (off) and a
per-home engine and model. Only the posture should follow the primary, so it
becomes a separate presence flag that rides the existing inherited-config
mechanism (FM_INHERITABLE_CONFIG in bin/fm-config-inherit-lib.sh) with no new
machinery, while the engine line stays local. The parse stays in its one
owner, fm_supervision_host_enabled. There is no migration or compatibility
handling for a home that still holds off in config/supervision-host.

Primary off, mate on: inherited material is primary-authoritative by design,
so a mate cannot keep the host while the primary is opted out, and a mate's
own opt-out is removed at the next convergence while the primary has none.
Running the host on a mate is the primary's choice for the fleet; no override
mechanism is added.

Live validation (disposable bin/fm-live-lab.sh lab, Claude primary with a
real seeded secondmate, --supervision-host off):
- up: every readiness check ok, including "host: none running, as expected"
  and a live mate session; the spawned mate home held the inherited
  config/supervision-host-off and the gate read primary OFF, mate OFF.
- primary removed its opt-out, then bin/fm-config-push.sh reported
  "supervision-host-off: pushed - mirrored primary absence" and a config
  reread sent; the gate read primary ON, mate ON, and the live mate handled
  the reread.
- primary opted out again and pushed: "supervision-host-off: pushed", mate
  gate OFF.
- down stopped every lab process and left no lab process running.

Out of scope, follow-up: default-on for the other harnesses, away-daemon
retirement, rollout.

* no-mistakes(document): Document inherited supervision-host opt-out ownership

* no-mistakes(ci): Fixed ci-4: with `--supervision-host off --mate`, lab readiness now requires the inherited flag in the mate home and a disabled mate supervision-host gate. The focused behavior test, shellcheck, and diff checks pass. Left ci-1–ci-3 untouched as directed

* no-mistakes(test): Fix mate readiness HOST_OFF initialization in lab up

* no-mistakes(ci): Fixed Lint 2 by making the new test’s fixtures source resolvable to ShellCheck; its off/on readiness test and ShellCheck now pass locally. Behavior portable serial 5 failed in the unchanged remote-reply test at generation 7. That test passes locally, and no PR-caused defect was identified, so no remote-reply code was changed
…nguid#6179)

* fix(tests): cut the fixed sleeps in supervision-host cycles

The serial CI lane keeps brushing its 30-minute cap because
fm-supervision-host.test.sh spends ~903s of the job, and per the
run-36635306527 case profile the top nine cases are all multi-cycle
ones (3-10 park/close/turn cycles each): every close waits out the
host's sleep $POLL in await_close plus a watcher sleep $FM_POLL scan
cycle, and every engine turn waits out the fixed sleep 1 descendant
snapshot. That is ~3s of pure sleep per cycle before any real work.

The host poll now accepts positive decimal seconds through a new
seconds_or validator (FM_SUPERVISION_HOST_POLL), and the engine turn's
snapshot loop takes FM_SUPERVISION_ENGINE_SNAPSHOT_SECONDS, also a
positive decimal defaulting to one second - the smallest seam at each
wait's single owner. The suite drives them at 0.2 alongside the
existing FM_POLL=0.5 and FM_ARM_ATTACH_POLL=0.2 knobs, so the real
poll loops still run. The park-boundary case moves onto the injected
test clock instead of a real 3s wait, per-case cleanup polls the host
pid rather than sleeping a full second, and the proof-by-absence
windows (flood re-escalation, successor re-announce, watcher
persistence, recovery staying off main) shrink from 2-3s to 1s, which
still spans two watcher polls at the test cadence.

Every assertion, process lifecycle, and reaping path is unchanged;
production defaults stay at one second. Isolated case timings on a
contended host, base vs branch: attended-latch 54.3->34.6s,
undelivered-dialog 67.7->59.1s, away-latch 46.5->30.5s, held-cadence
47.9->21.6s, unreadable-mirror 39.2->38.5s, park-limit 18.2->12.3s,
registration-fallback 14.1->10.0s, first-cycle-status 12.6->8.4s,
latch-scope 16.7->16.3s. Full suite: 65/65 pass. fm-lint and
shellcheck clean.

* no-mistakes(review): Wait for scan lock release before duplicate check

* no-mistakes(document): Correct supervision snapshot cadence documentation

* fix(tests): keep production poll cadence, probe exits at 0.1s

The fractional poll cadences multiplied the cost of each loop body:
full process-table scans in the engine turn and process refreshes in
await_close ran five times more often, which swamped the thin CI runner
and nearly doubled every multi-cycle case (serial 5 was cancelled at its
30-minute limit on run 36635306527's successor). Restore the production
cadence and notice arm/engine exits with a cheap kill -0 probe at a
tenth of a second between the one-second bodies instead: strictly less
dead time than baseline with no added CPU.

Also hold each injected-clock park bound well past its case's
wall-clock checks so a host that ignored the test clock fails instead
of silently passing at a real-time boundary, and restore the shortened
proof windows (watcher liveness, recovery-off-main absence, first-cycle
stream) to their baseline depth.

* no-mistakes(document): Clarify supervision engine snapshot documentation
…henguid#6192)

* fix: rebalance portable CI from current duration measurements

* no-mistakes(test): Test serial packing boundary and verify endpoint timeout cleanup

* no-mistakes(document): Clarify timeout guidance and remove duplicated packing estimates
…nguid#6216)

* fix(bin): run no repository hook when core.hooksPath is empty

The per-task hook wrapper refused every commit in a repository whose own
config sets core.hooksPath to the empty string, because git rev-parse
--git-path hooks fails on it. Plain git reads that setting as no hooks, so
the wrapper now runs none; every other lookup failure still refuses and
shows git's error.

Fixes kunchenguid#6171

* no-mistakes(review): Refuse commits when core.hooksPath is a valueless key

* no-mistakes(document): Document empty core.hooksPath handling in commit attribution docs

* no-mistakes(ci): When the wrapper refuses a commit, Git's hook-lookup error now shows up once instead of twice. That required changing one line in the wrapper, and the tests were extended so both bad-config cases would catch the duplicate. Invariant: when the wrapper refuses, Git's lookup error must appear exactly once. In the failure path, the only Git call besides the deliberate second lookup is the `git config --get --type=path core.hooksPath` check in `runtime_chain_body` (`bin/fm-git-strip-ai-trailers.sh:168`). That check prints the same error, so it was the one place to fix. I added `2>/dev/null` to it. Its exit status still decides the outcome: an empty value still runs no hook, and anything else goes on to the second lookup, which prints Git's error once, and the commit is refused. Tests (`tests/fm-git-strip-ai-trailers.test.sh`): - The unresolvable-path test (`~fm-no-such-user-6171/hooks`) now requires `failed to expand user dir` to appear exactly once in the refused commit's output. - The valueless-key test now requires `missing value for 'core.hookspath'` to appear exactly once. - Pre-existing bug in the unresolvable-path test: its `git add` ran after the bad config was set, so it failed silently (exit 128) and the "refused commit" had nothing staged. The test now stages the file before writing the config, the same way the valueless test does, so a real commit gets refused. - The empty-string test is unchanged and still passes, so an empty `core.hooksPath` still runs no hook. Verification: - With the wrapper change reverted, both new checks fail with `expected '1', got '2'`. With the change in place, the whole suite passes. - `bash -n` passes. shellcheck shows only an info-level SC1091 note about sourcing `lib.sh`, which was already there before this change. - `git status` lists only the two intended files
…kunchenguid#6213)

* fix(bin): let a stale record on a reassigned slot retire records-only

When a pool slot's owner claim names another task, the stale record's
teardown touches nothing under the slot, so the exclusive-slot record scan
no longer refuses it. Full teardowns of a slot this task still claims, or
one with no claim, keep the refusal.

Fixes kunchenguid#6184

* no-mistakes(document): Note claim-over-record precedence for reassigned teardown slots
Reconcile 182 upstream commits with 71 fork-only commits, keeping both
histories as merge parents so provenance survives.

Kept from the fork: the project-grouped task data layout
(data/tasks/<project>/<id>/) now recorded in the operational-home-layout
skill that replaced the AGENTS.md layout block, the durable report and
saved evidence trailer and repo-artifact rules in every ship definition
of done, the per-contract CI rule and context-spend rules in briefs, the
narrow Codex single-dot particle rule for composer furniture, the
shell-agnostic `env -u TRACEPARENT` pane clear, the verified launch
delivery, Herdr re-projection of an orphaned journal, the tasks-axi
report artifact predicate, and the contribution nameable-owner rule.

Kept from upstream: the attended supervision host, branch-prefix and
Gerrit forge support in briefs and promotion, the staged launch file,
status event stamps, the home brief include, the Pi cost footer
furniture rule, the shared fm_exec_timed bound, and the contribution
poll budget model.

Adapted where both sides met:
- fm_dod_block takes <mode> <task-id> <task-data-dir> [branch] [forge].
- Launch export statements are collected apart from the agent command so
  the env -u trace clear still lands on the command, not on an export.
- The composer footer zone counts Codex particle rows only under a
  `›`-proven envelope, in place of the wider braille furniture rule.
- fm-live-lab.sh resolves its worker brief through fm-task-data-lib.sh.

Taken from upstream over a local change: the contribution read retry
(cdbefaa) is the older version of the same idea. Upstream d051e6f treats
a read killed at its five-second cap as budget refusal that leaves records
untouched, so a slow read no longer raises an alert, and a retry would
overrun its fifteen-second per-URL reserve.
Upstream now opens every launch with export statements, so the clear sits on the command after them. The assertions check that position, and the fish proof runs the export-prefixed shape.
…ayout

The new upstream suite read data/<id>/brief.md, but the fork scaffolds briefs into data/tasks/<project>/<id>/, so the status command was never found.
…ts/fm-contributions.test.sh assumed the flat data/<id>/ layout. The fork stores task data under data/tasks/<project>/<id>/ and refuses task ids it cannot name. This is the same adaptation gap already fixed in other suites. Fixes, all in tests/fm-contributions.test.sh: - Reservation test: read the `filed` record through `fm_task_data_find` instead of `data/filed/`. A new task is written to the grouped path. - Sustained-slow-refresh test: read the `late-owner` record through `fm_task_data_find`, for the same reason. - Task-identity test: dropped the `-dash` and `nl\n` directory names. The fork's `fm_task_data_valid_id` refuses those names on purpose, and the existing "a task id the data layer cannot name" test already covers that. Each refused name counts as an unreadable record, which makes `pending` refuse as a whole, so the test got empty output. A comment now says why those names are excluded. No product code changed. Verification: the whole `bash tests/fm-contributions.test.sh` suite passed locally, with no `not ok`. In one earlier local run, while the machine was heavily loaded (load average about 8), two unrelated timing tests failed: the watcher surfacing a contribution signal, and generated-check budget arming. Both passed on rerun and both pass on CI, so I left them alone
@jbalke
jbalke merged commit 10fcb6f into main Oct 1, 2026
20 checks passed
@jbalke
jbalke deleted the fm/fm-reconcile-upstream-2026-09-30 branch October 1, 2026 10:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.