Skip to content

feat: sync upstream firstmate main into the fork - #11

Merged
zakna merged 33 commits into
mainfrom
fm/fm-upstream-sync-o1
Sep 27, 2026
Merged

zakna merged 33 commits into
mainfrom
fm/fm-upstream-sync-o1

Conversation

@zakna

@zakna zakna commented Sep 27, 2026

Copy link
Copy Markdown
Owner

Intent

Words from 2026-09-27: "lets update firstmate too get the upstream + our forks changes".

Context: this firstmate home runs from the fork zakna/firstmate (remote origin). The original project is kunchenguid/firstmate (remote upstream). The home already runs the fork's latest main. On 2026-09-27 the fork's main had 35 commits that upstream lacks, and upstream's main had 30 commits that the fork lacks; their merge base is 4299683 from 2026-09-25. A read-only test merge of upstream/main into origin/main conflicted only in AGENTS.md; every other file merged automatically.

What Changed

  • Merges the 30 upstream commits from kunchenguid/firstmate main into the fork. They add attended supervision for Claude and Cursor hosts in bin/fm-supervision-host.sh, plus a new bin/fm-remote-secondmate-relaunch.sh that republishes parent metadata after a remote secondmate relaunch. They also bring many bin/ fixes, including:

    • record-backed doorbell delivery for Claude-bound operational input
    • withholding never-send values from dispatch resolver requests
    • required declared waits for workers that are waiting on their own work
    • watcher lock and cleanup bounds, and stopping cancelled validation runs from reporting false failures
    • OpenCode dispatch effort passthrough, and preserving Herdr status on Pi relaunch
    • shared-captain drift reporting and wake routing by status span

    Each fix comes with its own test coverage.

  • Resolves the only merge conflict, in AGENTS.md, by taking upstream's restructure. That restructure moves situational sections into new on-demand skills under .agents/skills/ (operational-home-layout, session-start-recovery, ship-landing, validation-supervision, scout-completion, away-quiet-supervision, agent-skill-trigger-index) and applies it on top of the fork's own changes. Also brings in upstream's reworked docs/calm.md and docs/turnend-guard.md.

  • The ship-brief scope allowance in bin/fm-dod-lib.sh and bin/fm-brief.sh, and its test, now points to .agents/skills/validation-supervision/SKILL.md instead of the removed AGENTS.md section 7.

Risk Assessment

✅ Low: The round-1 fix only repoints the scope-allowance owner reference to .agents/skills/validation-supervision/SKILL.md, which does hold that allowance. The other "AGENTS.md section 7" references left in place are about delivery mode, which is still in AGENTS.md section 7, and the updated test checks the generated brief output.

Testing

Checked the merge topology and conflict resolution. In a disposable lab home, scaffolded ship briefs in all three delivery modes and confirmed the scope-allowance pointer resolves to a real skill file that holds the allowance. Launched a real Claude primary in the lab on the private tmux socket; it named the validation-supervision skill and fm-dod-lib.sh as the owners and quoted the fork's no-CI ready-signal pattern word for word. Ran the six focused suites covering files both sides changed; all passed. The merge-topology and fork-feature scenarios are recorded as untested because they were not driven live. The lab was torn down and the worktree is clean. Overall: go.

  • Live validation: ✅ go - 2 of 4 scenarios driven live against the product
Scenario Result Live Evidence
Merge brings in all upstream commits and keeps fork main, with no leftover conflict markers ⏸️ untested no The prior payload established this only through a repository-state check (git show/merge-base/grep), not a live product run, so there is no live result.
Scaffolding a ship brief in any delivery mode gives a scope allowance that points to an existing skill file containing the allowance ✅ pass live fm-brief-scaffold.txt, scope-allowance-pointer.txt
A lab firstmate primary running the merged instructions finds the scope-allowance owner and the fork's no-CI ready-signal wording ✅ pass live lab-primary-transcript.txt
Fork features still work after the merge: intent check, pool-slot anchor refusal, review-comment inventory, ready-signal evidence, external-merge/crew-state, teardown ⏸️ untested no The prior payload covered this only through existing executable test suites (all passing), not a live fleet run, so there is no live result.
Evidence: Ship brief scaffold transcript (lab home, 3 modes)

Source: Ship brief scaffold transcript (lab home, 3 modes)

scaffolded: /var/folders/hk/lxr6ppq904ng8dxkxskjs5180000gn/T//fm-lab.WF7ivg/data/scope-no-mistakes/brief.md (ship, mode=no-mistakes; replace {TASK} and {FIRSTMATE_SPEC})
rc=0
scaffolded: /var/folders/hk/lxr6ppq904ng8dxkxskjs5180000gn/T//fm-lab.WF7ivg/data/scope-direct-PR/brief.md (ship, mode=direct-PR; replace {TASK} and {FIRSTMATE_SPEC})
rc=0
scaffolded: /var/folders/hk/lxr6ppq904ng8dxkxskjs5180000gn/T//fm-lab.WF7ivg/data/scope-local-only/brief.md (ship, mode=local-only; replace {TASK} and {FIRSTMATE_SPEC})
rc=0
Evidence: Scope allowance pointer resolves to validation-supervision skill

Source: Scope allowance pointer resolves to validation-supervision skill

== no-mistakes
# Scope allowance
However narrowly the task above states its scope, the smallest downstream changes needed to keep already accepted behavior correct, add behavioral tests where an executable contract exists, or keep documentation accurate stay within this task even in files it does not name; Firstmate's `~/.no-mistakes/worktrees/0b1ab7dd6182/01M3HTH1B393AQ2FDGCPHSTJ6Z/.agents/skills/validation-supervision/SKILL.md` owns this allowance.
referenced path: ~/.no-mistakes/worktrees/0b1ab7dd6182/01M3HTH1B393AQ2FDGCPHSTJ6Z/.agents/skills/validation-supervision/SKILL.md
EXISTS; skill carries allowance:
1
== direct-PR
# Scope allowance
However narrowly the task above states its scope, the smallest downstream changes needed to keep already accepted behavior correct, add behavioral tests where an executable contract exists, or keep documentation accurate stay within this task even in files it does not name; Firstmate's `~/.no-mistakes/worktrees/0b1ab7dd6182/01M3HTH1B393AQ2FDGCPHSTJ6Z/.agents/skills/validation-supervision/SKILL.md` owns this allowance.
referenced path: ~/.no-mistakes/worktrees/0b1ab7dd6182/01M3HTH1B393AQ2FDGCPHSTJ6Z/.agents/skills/validation-supervision/SKILL.md
EXISTS; skill carries allowance:
1
== local-only
# Scope allowance
However narrowly the task above states its scope, the smallest downstream changes needed to keep already accepted behavior correct, add behavioral tests where an executable contract exists, or keep documentation accurate stay within this task even in files it does not name; Firstmate's `~/.no-mistakes/worktrees/0b1ab7dd6182/01M3HTH1B393AQ2FDGCPHSTJ6Z/.agents/skills/validation-supervision/SKILL.md` owns this allowance.
referenced path: ~/.no-mistakes/worktrees/0b1ab7dd6182/01M3HTH1B393AQ2FDGCPHSTJ6Z/.agents/skills/validation-supervision/SKILL.md
EXISTS; skill carries allowance:
1
Evidence: Live lab Claude primary transcript

Source: Live lab Claude primary transcript

 ▐▛███▛█   Claude Code v2.1.283
▝▜██████▀  Opus 5.5 with xhigh effort · Claude Max
 ▝▝   ▝▝   ~/.no-mistakes/worktrees/0b1ab7dd6182/01M3HTH1B393AQ2FDGCPHSTJ6Z
❯ Quick check, do not start any fleet work or spawn anything: per your instructions, which file owns (a) the ship-task scope allowance for smallest downstream changes and (b) the no-mistakes ready
  signal wording including the 'no CI' form? Quote the no-CI ready-signal pattern verbatim. Answer in 3 lines.
  Searched for 1 pattern
⏺ Captain, here's the answer:
  1. Smallest-downstream-changes allowance: .agents/skills/validation-supervision/SKILL.md:15 owns it. bin/fm-dod-lib.sh:577 says so when it adds the allowance to each worker's instructions.
  2. No-mistakes ready-signal wording, including the no-CI form: bin/fm-dod-lib.sh owns it (lines 764–774, checked at line 806). ship-landing points to it as the owner.
  3. No-CI pattern, verbatim: done [at=<epoch>]: PR {url} no CI: test step {n}/{m} scenarios live; untested: {each scenario the test step could not drive, or none}
✻ Sautéed for 15s · done 6:27 PM
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
❯ 
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
  ➜  worktrees/0b1ab7dd6182/01M3HTH1B393AQ2FDGCPHSTJ6Z ctx 48809/1000000  Opus xhigh | 5h [██░░░░] 27% →20:30 | 7d [█░░░░░] 18% →Sat 19:00
  ⏵⏵ auto mode on (shift+tab to cycle) · ← for agents
Evidence: fm-brief test log

Source: fm-brief test log

ok - fm-brief: scaffolds leave the worker role scope to the launch boundary and keep the secondmate contract
ok - fm-brief.sh: bash -n succeeds
/private/var/folders/hk/lxr6ppq904ng8dxkxskjs5180000gn/T/fm-brief.YST3Vf/heredoc-in-substitution.sh:2
ok - fm-brief.sh: no heredoc is nested inside a command substitution (Bash 3.2 parse-safe)
ok - fm-brief.sh: --help renders the complete header
ok - fm-brief.sh: no-mistakes/direct-PR/local-only briefs generate cleanly
ok - fm-brief.sh: ship --mode is required and closed-set validated
ok - fm-brief.sh: the explicit ship mode wins over the registered posture
ok - fm-brief.sh: --yolo and scout/secondmate --mode are refused, never silently dropped
ok - fm-brief.sh: faster paths use configured authority without stacked review
ok - fm-brief.sh: no-mistakes DOD keeps its apostrophe prose and bans --yes outright
ok - fm-brief.sh: no-mistakes DOD detects a green PR from the drive call, not a status poll
ok - fm-brief.sh: the ready line names its CI, no-CI, and held evidence
ok - fm-brief.sh: PR-based done requires a non-draft PR; a deliberate draft declares a wait
ok - fm-brief.sh: no-mistakes ask-user findings use one event plus a verbatim snapshot
ok - fm-brief.sh: ship project-memory wording bounds edits to corrections of wrong information
ok - fm-brief.sh: ship briefs keep the documentation-accuracy scope allowance visible
ok - fm-brief.sh: --herdr-lab emits the complete hard safety contract
ok - fm-brief.sh: --herdr-lab uses its quoted Firstmate-owned helper path
ok - fm-brief.sh: ship and scout scaffolds make omitted Herdr intent fail-visible
ok - fm-brief.sh: the documented {TASK} and {FIRSTMATE_SPEC} fills cannot corrupt the Herdr safety gate
ok - fm-brief.sh: Herdr lab contract covers scouts and rejects secondmate misuse
ok - fm-brief.sh: --no-projects scaffolds a project-less charter and guards misuse
ok - fm-brief.sh: marked requests avoid generic acknowledgements and preserve material reporting
ok - fm-brief.sh: relative directory inputs ignore CDPATH, render stable absolute charter paths, or fail loudly
ok - fm-brief.sh: custom pause verb renders in every scaffold
ok - fm-brief.sh: ship and scout scaffolds teach validation-round pauses
ok - fm-brief.sh: investigation and visual-review completions load the shared decision policy
ok - fm-brief: scout and secondmate code paths still scaffold well-formed briefs
ok - fm-brief.sh: scout Lavish hosting follows the bootstrap lavish-axi floor
ok - fm-brief.sh: the home brief include lands last on ship and scout, verbatim, and fails closed
ok - fm-brief.sh: --branch-prefix omitted defaults every ship mode to fm/<task-id>
ok - fm-brief.sh: a --branch-prefix override renders identically across every generated section
ok - fm-brief.sh: an empty --branch-prefix override resolves to a bare <task-id> branch
ok - fm-brief.sh: --branch-prefix is refused on scout and secondmate scaffolds
ok - fm-brief.sh: --branch-prefix value is validated against embedded spaces and a leading dash
Switched to a new branch '$(touch${IFS}/private/var/folders/hk/lxr6ppq904ng8dxkxskjs5180000gn/T/fm-brief.YST3Vf/branch-prefix-shell-safe-marker)brief-branch-safe-g3'
ok - fm-brief.sh: ref-format-valid shell metacharacters stay literal in generated branch commands
ok - fm-brief.sh: every crewmate scaffold forbids administering the shared worktree pool
ok - fm-brief.sh: ship briefs require an existing automated-review comment inventory
Evidence: fm-crew-state test log

Source: fm-crew-state test log

ok - captured AXI replacement status replays through crew-state
ok - captured AXI parked status replays through crew-state
ok - captured AXI failed status replays through crew-state
ok - captured capped inventory replays selection, ambiguity, and unavailable lookup
ok - captured status formats reject a synthetic authority transition
ok - captured completed status yields to synthetic subsequent development
ok - captured cancelled run leaves fleet inventory unverified without a failure contradiction
ok - failed-outcome/github/open: terminal delivery uses current disposition
ok - failed-outcome/github/merged: terminal delivery uses current disposition
ok - failed-outcome/github/closed: terminal delivery uses current disposition
ok - failed-outcome/github/unreadable: terminal delivery uses current disposition
ok - failed-outcome/github/skipped: terminal delivery uses current disposition
ok - failed-outcome/github/no-identity: terminal delivery uses current disposition
ok - failed-outcome/gitlab/open: terminal delivery uses current disposition
ok - failed-outcome/gitlab/merged: terminal delivery uses current disposition
ok - failed-outcome/gitlab/closed: terminal delivery uses current disposition
ok - failed-outcome/gitlab/unreadable: terminal delivery uses current disposition
ok - failed-outcome/gitlab/skipped: terminal delivery uses current disposition
ok - failed-outcome/gitlab/no-identity: terminal delivery uses current disposition
ok - failed-outcome/gerrit/open: terminal delivery uses current disposition
ok - failed-outcome/gerrit/merged: terminal delivery uses current disposition
ok - failed-outcome/gerrit/closed: terminal delivery uses current disposition
ok - failed-outcome/gerrit/unreadable: terminal delivery uses current disposition
ok - failed-outcome/gerrit/skipped: terminal delivery uses current disposition
ok - failed-outcome/gerrit/no-identity: terminal delivery uses current disposition
ok - failed-status/github/open: terminal delivery uses current disposition
ok - failed-status/github/merged: terminal delivery uses current disposition
ok - failed-status/github/closed: terminal delivery uses current disposition
ok - failed-status/github/unreadable: terminal delivery uses current disposition
ok - failed-status/github/skipped: terminal delivery uses current disposition
ok - failed-status/github/no-identity: terminal delivery uses current disposition
ok - failed-status/gitlab/open: terminal delivery uses current disposition
ok - failed-status/gitlab/merged: terminal delivery uses current disposition
ok - failed-status/gitlab/closed: terminal delivery uses current disposition
ok - failed-status/gitlab/unreadable: terminal delivery uses current disposition
ok - failed-status/gitlab/skipped: terminal delivery uses current disposition
ok - failed-status/gitlab/no-identity: terminal delivery uses current disposition
ok - failed-status/gerrit/open: terminal delivery uses current disposition
ok - failed-status/gerrit/merged: terminal delivery uses current disposition
ok - failed-status/gerrit/closed: terminal delivery uses current disposition
ok - failed-status/gerrit/unreadable: terminal delivery uses current disposition
ok - failed-status/gerrit/skipped: terminal delivery uses current disposition
ok - failed-status/gerrit/no-identity: terminal delivery uses current disposition
ok - cancelled-outcome/github/open: terminal delivery uses current disposition
ok - cancelled-outcome/github/merged: terminal delivery uses current disposition
ok - cancelled-outcome/github/closed: terminal delivery uses current disposition
ok - cancelled-outcome/github/unreadable: terminal delivery uses current disposition
ok - cancelled-outcome/github/skipped: terminal delivery uses current disposition
ok - cancelled-outcome/github/no-identity: terminal delivery uses current disposition
ok - cancelled-outcome/gitlab/open: terminal delivery uses current disposition
ok - cancelled-outcome/gitlab/merged: terminal delivery uses current disposition
ok - cancelled-outcome/gitlab/closed: terminal delivery uses current disposition
ok - cancelled-outcome/gitlab/unreadable: terminal delivery uses current disposition
ok - cancelled-outcome/gitlab/skipped: terminal delivery uses current disposition
ok - cancelled-outcome/gitlab/no-identity: terminal delivery uses current disposition
ok - cancelled-outcome/gerrit/open: terminal delivery uses current disposition
ok - cancelled-outcome/gerrit/merged: terminal delivery uses current disposition
ok - cancelled-outcome/gerrit/closed: terminal delivery uses current disposition
ok - cancelled-outcome/gerrit/unreadable: terminal delivery uses current disposition
ok - cancelled-outcome/gerrit/skipped: terminal delivery uses current disposition
ok - cancelled-outcome/gerrit/no-identity: terminal delivery uses current disposition
ok - cancelled-status/github/open: terminal delivery uses current disposition
ok - cancelled-status/github/merged: terminal delivery uses current disposition
ok - cancelled-status/github/closed: terminal delivery uses current disposition
ok - cancelled-status/github/unreadable: terminal delivery uses current disposition
ok - cancelled-status/github/skipped: terminal delivery uses current disposition
ok - cancelled-status/github/no-identity: terminal delivery uses current disposition
ok - cancelled-status/gitlab/open: terminal delivery uses current disposition
ok - cancelled-status/gitlab/merged: terminal delivery uses current disposition
ok - cancelled-status/gitlab/closed: terminal delivery uses current disposition
ok - cancelled-status/gitlab/unreadable: terminal delivery uses current disposition
ok - cancelled-status/gitlab/skipped: terminal delivery uses current disposition
ok - cancelled-status/gitlab/no-identity: terminal delivery uses current disposition
ok - cancelled-status/gerrit/open: terminal delivery uses current disposition
ok - cancelled-status/gerrit/merged: terminal delivery uses current disposition
ok - cancelled-status/gerrit/closed: terminal delivery uses current disposition
ok - cancelled-status/gerrit/unreadable: terminal delivery uses current disposition
ok - cancelled-status/gerrit/skipped: terminal delivery uses current disposition
ok - cancelled-status/gerrit/no-identity: terminal delivery uses current disposition
ok - outcome: cancellation without delivery carries no verdict
ok - status: cancellation without delivery carries no verdict
ok - selected: cancellation without delivery carries no verdict
ok - coarse: cancellation without delivery carries no verdict
ok - no-ci-log: cancellation without delivery carries no verdict
ok - red-ci: cancellation without delivery carries no verdict
ok - cancelled-test: cancellation without delivery carries no verdict
ok - skipped-test: cancellation without delivery carries no verdict
ok - synthetic cancelled run leaves fleet inventory unverified without a failure contradiction
ok - cancelled-outcome: terminal delivery reports only observed evidence
ok - cancelled-status: terminal delivery reports only observed evidence
ok - skipped-rebase: terminal delivery reports only observed evidence
ok - cancelled-skipped-rebase: terminal delivery reports only observed evidence
ok - passed: terminal delivery reports only observed evidence
ok - active run-step is authoritative
ok - stale needs-decision over active run is superseded
ok - stale blocked over active run is superseded
ok - daemon/timeout blocked claim over a live fixing run reads as run alive
ok - socket refusal or missing socket over a stale fixing run reports blocked
ok - socket refusal over a terminal attributed run reports blocked
ok - socket-down evidence outranks a live run only while it is the log's latest event
ok - broken-pipe blocker over a live run keeps the plain superseded reading
ok - genuine daemon-down blocked line still reports blocked
ok - a busy secondmate keeps its open blocker until that exact key closes
ok - the most recently opened decision supplies the reported state and detail
ok - ship and scout terminal declarations supersede stale decisions
ok - latest status retains legacy completion events and shared captain matching
ok - latest status subprocess work stays bounded an

... [3512 bytes truncated] ...

erted adapter never reads working from rendered footer text
ok - grok still reads working through its isolated rendered-tail fallback
ok - herdr's native busy verdict reads working with no record present
ok - a herdr CLI that fails to answer reads unknown/unreachable, never gone
ok - an alive endpoint whose scrollback read failed stays working
ok - a husk pane (agent gone) still reads gone for reclaim
ok - a mid-tool-call crew stays working because its record outranks herdr's generation state
ok - an idle record with idle agent_status stays not-busy (no regression for a human-blocked agent)
ok - no run + idle pane uses the status-log verb
ok - no run + idle pane parses keyed status syntax
ok - no run + idle pane on a paused: status reports state: paused with its reason
ok - no run + idle pane honors the configured paused verb
ok - a trailing resolved: event does not corrupt state render (idle stays idle)
ok - dead window ignores stale status log
ok - a tmux that fails to answer reads unknown/unreachable, never gone
ok - closed pane still reports a terminal run-step
ok - closed pane still reports an active run-step
ok - no timeout command uses perl bound
ok - scout skips the run lookup
ok - torn-down worktree is handled gracefully
ok - fm-crew-state remote: alive endpoint falls through to the routed status log
ok - fm-crew-state remote: an idle alive endpoint reads alive, never gone or dead
ok - fm-crew-state remote: an unreachable host reads unknown-remote, never gone or dead
ok - fm-crew-state remote: the remote host's own dead verdict is reported truthfully
ok - missing meta is handled gracefully
ok - crew_is_provably_working absorbs a validating crew found only via the runs-list fallback
ok - crew_is_provably_working still surfaces a genuinely stopped crew (safety property preserved)
ok - usage error exits 2
ok - historical same-branch rewritten head is not attributed as current
ok - active run with valid descendant fix head remains current
ok - local work advanced past run head invalidates attribution
ok - pipeline-owned active run binds without head equality and beats the failed row
ok - a genuinely failed run with no later run is not hidden
ok - coarse scan anchors the unresolvable active row instead of falling to an older one
ok - coarse scan with a mismatched anchor stays unknown and lets the pane answer
ok - coarse terminal row at a foreign head is not attributed
ok - an executing run binds regardless of branch_sync state
ok - a parked run keeps the strict head rule without pipeline_owned
ok - run_parked_scalar_gate_running keeps the strict head rule despite its live status word
ok - run_parked_in_gate_block keeps the strict head rule despite its live status word
ok - the exemption never applies to a terminal run
ok - missing run head falls back instead of matching by branch
ok - active fix round with an unfetched pipeline head reads working
ok - unanchored unverifiable active row is attributed because it is live
ok - unresolvable terminal row never reads as current
ok - runs-list continuation attribution works when axi answers another branch
ok - herdr stale registration over a shell-only pane reads agent gone, not alive
ok - herdr stale working record never reports a shell-only pane busy
ok - capped overview retains both competing same-branch run ids
ok - same-branch identity survives both runs falling outside the overview
ok - a capped overview with zero same-branch rows reports absent, not unreadable
ok - no run for this branch beside a live run elsewhere reads absent, not unreadable
ok - the capped inventory reader is bounded by the crew read budget
ok - a repo spelling the inventory does not record reads unreadable
ok - a linked worktree green PR in merge monitoring reads held for merge
ok - complete inventory preserves the replacement gate without writes
ok - missing complete-inventory failure reports unknown
ok - corrupt complete-inventory failure reports unknown
ok - schema complete-inventory failure reports unknown
ok - repo complete-inventory failure reports unknown
ok - count complete-inventory failure reports unknown
ok - norepo complete-inventory failure reports unknown
ok - R6 complete selection ignores unrelated branch semantics
ok - R6 requested branches use exact identity without a whitelist
ok - R6 capped inventory ignores unrelated semantics and names both ids
ok - R6 complete identity lookup precedes partial-row semantic rejection
ok - R6 capped inventory preserves quoted requested-branch identity
ok - R6 structural completeness and requested-run validation remain enforced
ok - R5 complete inventory without Python keeps the replacement gate
ok - R5 complete ambiguity without Python names both ids
ok - R5 capped lookup without Python preserves available ids
ok - R5 capped lookup without SQLite support preserves available ids
ok - R1 both directions of inventory liveness disagreement read unknown
ok - R2 uninitialized busy workers retain pane reporting
ok - R2 uninitialized idle workers retain status reporting
ok - R3 historical inventory yields to the current busy pane
ok - R3 historical inventory yields to current worker status
ok - superseded cancelled run preserves the replacement review gate
ok - a live rebased run beats an older failed run at the local head
ok - pending run with a rebased head reads working
ok - running run with a rebased head reads working
ok - legacy live rebased run is authoritative over an older failed row
ok - legacy surface binds a fixing run at a rebased head
ok - legacy surface binds a ci run at a rebased head
ok - an unproven record at a diverged head does not answer for the crew
ok - an unproven record with a dead daemon never overrides a busy pane
ok - an ordinary blocked tip over a coarse live row keeps the superseded reading
ok - a socket-refused blocker survives the dead-daemon verdict
ok - an ordinary blocked tip survives the dead-daemon verdict
ok - a parked gate survives a dead daemon with its findings intact
ok - an unproven record at a diverged head does not answer on the selected route
ok - a live record at a diverged head binds while the daemon answers
ok - the anchored continuation binds while the daemon answers
ok - a head-tied coarse record keeps its working reading and its original note
ok - a coarse failed record with a dead daemon reads unknown
ok - the selected-run anchored continuation reports the dead daemon, not an identity failure
ok - the selected-run anchored continuation binds while the daemon answers
ok - the selected route keeps an anchored parked run's gate with a dead daemon
ok - an open decision survives the dead-daemon verdict on the selected route
ok - a coarse pending ledger word reads unknown
ok - an unanswered probe never turns a failed coarse record into a gate
ok - the selected-route dead-daemon verdict names the run once
ok - a head-tied row reads working when axi names the self run
ok - a head-tied row reads working when axi names the other run
ok - a head-tied coarse row is exempt even when the record head diverged
ok - an unrecognised ledger word keeps the ordinary supersede note
ok - an unanswered daemon probe leaves a live rebased run bound
ok - a head-tied coarse live row is exempt from the dead-daemon verdict
ok - a coarse live row over an open decision keeps the original supersede note
ok - a coarse live row at a rebased head is not attributed
ok - a terminal run at a diverged head keeps the strict head rule
ok - competing live runs report unknown with both run ids
ok - newer failed run remains failed beside an older live run
ok - missing run selection reports unknown with candidate ids
ok - wrong-id run selection reports unknown with candidate ids
ok - wrong-branch run selection reports unknown with candidate ids
ok - wrong-head run selection reports unknown with candidate ids
ok - missing-status run selection reports unknown with candidate ids
ok - malformed-table run selection reports unknown with candidate ids
ok - inventory-error run selection reports unknown with candidate ids
ok - selected-error run selection reports unknown with candidate ids
ok - legacy conflicting run records report unknown
all fm-crew-state tests passed
Evidence: fm-intent-check test log (fork feature)

Source: fm-intent-check test log (fork feature)

ok - a clean intent and a sentence-level subset pass
ok - second-person product wording, links, and task items pass; worker address does not
ok - each forbidden class is refused and names its line
ok - only a quote a speech word attributes is refused
ok - only fleet terms the repository does not use are refused
ok - fenced examples and inline code are exempt from the refusal classes
ok - only this repository's named issues count as references
ok - scrub refuses and names each refused line, and a clean scrub passes the check
ok - a refusal names its source file, and rewording that file clears it
ok - scrub keeps a Refs line the brief names without repeating it
ok - later captain words and resolved substance are authorized but still checked
ok - a legacy provenance-marked Task authorizes only its marked words
ok - a task directory resolves to the launch overlay, then promotion instructions
ok - the no-mistakes brief tells the worker to run the check on the task directory
Evidence: fm-spawn-pool-slot-anchor test log (fork feature)

Source: fm-spawn-pool-slot-anchor test log (fork feature)

ok - a spawn for one clone refuses a Treehouse slot anchored to another clone of the same origin
ok - the same slot still launches and is claimed for the clone that owns it
# all fm-spawn-pool-slot-anchor tests passed

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

⏭️ **Rebase** - skipped
  • ⚠️ AGENTS.md - merge conflict rebasing onto origin/main
🔧 **Review** - 1 issue found → auto-fixed ✅
  • ⚠️ bin/fm-dod-lib.sh:577 - This is drift from the merge itself. The fork's ship-brief # Scope allowance block tells workers that "Firstmate's &lt;home&gt;/AGENTS.md section 7 owns this allowance". Upstream docs: move situational AGENTS.md sections into on-demand skills kunchenguid/firstmate#5872 moved the Validate text, including the "smallest downstream changes ... stay within the current task" allowance, out of AGENTS.md and into .agents/skills/validation-supervision/SKILL.md:15 (also .agents/skills/ask-user-authority/SKILL.md:29). After the merge, AGENTS.md section 7 ### Validate (AGENTS.md:241-243) only says to load validation-supervision. A worker or reviewer who follows the pointer to AGENTS.md section 7 will not find the allowance there. The allowance text is still inlined in the brief, so this does not break anything: the pointer is misleading, not broken. Fix: name .agents/skills/validation-supervision/SKILL.md as the owner. The same stale pointer also appears in bin/fm-brief.sh:14-15 (header comment), bin/fm-brief.sh:44, 91 and 252, bin/fm-spawn.sh:9 and 832 (comments), and tests/fm-brief.test.sh:556 and 572 (which assert the emitted pointer string).

🔧 Fix applied.
✅ Re-checked - no issues remain.

✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 2 of 4 scenarios driven live against the product
Scenario Result Live Evidence
Merge brings in all upstream commits and keeps fork main, with no leftover conflict markers ⏸️ untested no The prior payload established this only through a repository-state check (git show/merge-base/grep), not a live product run, so there is no live result.
Scaffolding a ship brief in any delivery mode gives a scope allowance that points to an existing skill file containing the allowance ✅ pass live fm-brief-scaffold.txt, scope-allowance-pointer.txt
A lab firstmate primary running the merged instructions finds the scope-allowance owner and the fork's no-CI ready-signal wording ✅ pass live lab-primary-transcript.txt
Fork features still work after the merge: intent check, pool-slot anchor refusal, review-comment inventory, ready-signal evidence, external-merge/crew-state, teardown ⏸️ untested no The prior payload covered this only through existing executable test suites (all passing), not a live fleet run, so there is no live result.
  • git show --format=%P -s 1d85c04e, git merge-base, git grep for conflict markers (none)
  • bin/fm-lab-home.sh create $LAB then FM_HOME=$LAB bin/fm-brief.sh scope-&lt;mode&gt; demo --mode &lt;no-mistakes|direct-PR|local-only&gt;, then resolve the # Scope allowance path and check that it exists and contains the allowance text
  • Live lab Claude primary started on the private fm-lab tmux socket (FM_HOME=$LAB), asked which files own the scope allowance and the no-CI ready signal; answer captured with capture-pane, then kill-server and rm -rf $LAB
  • bash tests/fm-brief.test.sh
  • bash tests/fm-crew-state.test.sh
  • bash tests/fm-inactive-reconcile.test.sh
  • bash tests/fm-intent-check.test.sh
  • bash tests/fm-spawn-pool-slot-anchor.test.sh
  • bash tests/fm-teardown.test.sh
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

kunchenguid and others added 30 commits September 25, 2026 17:51
… escalations are not repeated (kunchenguid#5731)

* fix(bin): retire check-row receipts on branch acks and report an unchanged situation once

* fix(bin): scope a branch acknowledgement's check-row receipt retirement to
  its granted sequences

The away posture lifts the attended partition's check/decision exclusions, so
a branch grant can name check-kind rows - but the branch-actor ack still
assumed check rows were main-only and skipped every receipt scan. The queue
row was consumed while its terminal-outcome .pending receipt stayed behind,
and each inactive-reconcile cadence scan re-queued the same fingerprint. In
the first real away window on the supervision host that re-escalated one
unchanged held-PR situation on every cycle (~1,734 of 4,149 outcomes).

A branch ack now scans inactive-outcome and inactive-reconcile receipts and
commits secondmate stall receipts against exactly the sequences in its
eligible-row snapshot - the same rows it consumes - instead of none. Attended
grants still name no check row, so the scans find nothing.

* fix(bin): store a repeated captain verdict as routine while the task's
  durable situation is provably unchanged

fm-branch-outcome.sh append computes a mechanical situation key per captain
row - metadata bytes, captured status-log endpoint and identity, live
crew-state verb, worktree head - and anchors it in
state/.<task>.branch-captain-key. A later captain verdict whose recomputed key
matches is stored as routine with "unchanged since seq <N>:" prefixed to its
summary, so one situation escalates once until something provably changes. A
task with no readable status ledger is never demoted, an unreadable record
fails toward reporting, and teardown removes the sidecar with the task's
other branch records. The append-only store schema is unchanged.

This covers both hosts: the Pi supervision branch and the supervision host
both funnel reports through append.

* docs: check rows are main-owned only while attended; the away posture grants
  them to the branch, whose ack retires their receipts exactly

* test: the away-flood reproduction as a regression test (branch ack retires
  the receipt and later scans stay quiet), store-level dedupe coverage, and a
  branch-ack secondmate stall receipt case

* fix(bin): restore the secondmate child devin-config cleanup path

The branch-captain-key sidecar addition mistyped the sibling entry as
.$child_id.devin-config.json, so a forced secondmate teardown would have
stopped removing each child's real <id>.devin-config.json. Restore the
original path and add a behavioral test that stops the child sweep mid-loop
on a refused close, proving the cleaned child's devin config and captain
anchor are both removed while the unconsumed child's records are retained.

* no-mistakes(review): Key captain dedupe on the covered wake rows' fingerprint

* no-mistakes(review): Drop captain-key demotion; prove one escalation on both surfaces

* no-mistakes(review): Drop unrelated teardown test; cite both receipt test files

* no-mistakes(document): Docs already match branch-ack check-receipt retirement
…oorbell (kunchenguid#5664)

* fix(calm): deliver Claude-bound operational input as a record-backed doorbell

Claude Code 2.1.280 removes U+2063 from every submitted prompt, so a typed
operational envelope reaches a Claude Code primary as plain text. The away
daemon now writes the envelope to a record under state/operational-inbox and
types only a plain doorbell naming it; the /afk return check and the Calm mod
recognize the doorbell only when that record holds a current envelope. Marker-
preserving harnesses keep the typed envelope. The live Calm guard accepts the
2.1.280 module-load log line, drives the doorbell, and asserts thinking stays
hidden.

* no-mistakes(review): Fix operational record retention at 7 days and document prune limit

* no-mistakes(document): Point Calm bounds at 2.1.280 evidence; fix afk-exit comment

* no-mistakes(lint): Pick newest Calm e2e transcript without parsing ls

* docs(calm): add a minimal turning-Calm-on step for Claude Code

* fix(spawn): deliver the Claude launch brief as a record-backed doorbell

Claude Code strips U+2063 from the launch-prompt argument too, so a
worker's launch brief arrived with its operational marker removed.
Publish the brief as a record in the receiving home's operational
inbox - a secondmate's own state, not the primary's - and pass only
the printable doorbell naming it, falling back to the typed envelope
when the record cannot be published so the brief body still delivers.

Unwrap doorbell-carried digests in the daemon digest tests that still
read the raw send log under the claude pin, and update the documented
bounds now that launch briefs hide like the other operational rows.

* test(spawn): cover a secondmate's launch-brief record landing in its own home

The record-backed doorbell resolves its state through the receiving
pane's home, so prove a claude secondmate launch publishes into the
seeded secondmate's operational inbox and never leaks a record into
the primary's.

* no-mistakes(review): Pass primary harness to daemon, tighten retention, refresh verdicts

* no-mistakes(review): Prune operational records by exact seven-day elapsed age

* no-mistakes(review): Batch record pruning so large inboxes still expire

* no-mistakes(review): Refuse Claude spawn when brief record cannot publish

* no-mistakes(review): Drop thinking probe from Claude Calm live test and docs

* no-mistakes(review): Record dated Claude Code 2.1.282 reproduction evidence

* no-mistakes(document): Clarify operational doorbell documentation and record expiry

* no-mistakes(document): Correct AFK escalation carrier guidance

* no-mistakes(review): Describe operational record retention as about seven days

* no-mistakes(document): Clarify Calm delivery and operational record retention

* no-mistakes(review): Remove out-of-scope Calm launch guide from Claude docs

* no-mistakes(document): Document Claude launch-brief delivery and refusal

* no-mistakes(document): Correct stale operational-input documentation

* no-mistakes(ci): Fixed the stale Claude trust test to verify that worker and secondmate launches deliver readable, record-backed briefs instead of expecting brief paths in their commands. Annotated the daemon’s output variable for ShellCheck without changing behavior. The affected tests, daemon tests, ShellCheck, and diff check pass locally

* no-mistakes(ci): parse rebased Claude launch after trailer hook prefix

* no-mistakes(review): Trust launch-brief record and restore thinking bound doc

* no-mistakes(review): Parse final Claude launch statement; drop Stop-hook docs

---------

Co-authored-by: Mike Sewell <maikunari@protonmail.com>
Co-authored-by: no-mistakes <no-mistakes@localhost>
kunchenguid#5583)

A host-local relaunch rewrote only the far endpoint, so this home kept the old harness, model, and effort, and appending those keys after pr= broke pull-request poll authentication.
…ound (kunchenguid#5516)

tests/fm-watch-triage.test.sh finishes in about 434s alone and about 698s
under CI load, so the 900s bound the changed-suite runner applies produced
a false timeout under ordinary concurrent validation. Raise the automatic
bound to 1500s, which keeps every measured script under it while staying
below the 30-minute normal CI tier so a genuinely hung script still fails
here with its output before the job cap cancels the lane.

Fixes kunchenguid#3869
Refs kunchenguid#3565
…kunchenguid#5728)

* Fix nested watcher lock reclaim

* no-mistakes(review): Elect a single steal-mutex reaper and bound arm TERM wait

* no-mistakes(review): Reclaim self-held steal mutex and unify autoarm steal reaping

* no-mistakes(review): Resume own interrupted steal reap from its tombstone
…nguid#5710)

* test: hold the back-to-back boundary close on the host's own clock

test_park_boundary_holds_under_back_to_back_closes assumed two engine
turns fit in the ~16s pre-refusal window and that the stub finished a
turn in 3s. Under load the stub's real drain, report, and
acknowledgement take ~13s, so the turn either died at its bound (which
hands the wake to main, no boundary line) or the second close landed
past the window and the fixture failed while the boundary held. 3
failures in 5 runs at a load average near 11.

Hold the first turn on a release file instead: once the engine is in
flight, a second close is appended mid-turn and the turn is released as
the refusal window opens (park bound minus turn bound and grace, read
off the host's own start record). The queued close can then only wait
for the boundary on any machine speed, which is what the test asserts:
the boundary line ends the output, the demo.status row stays queued for
main, and no second engine turn ever starts. A host too loaded to start
the turn at all hands the first close to the same boundary exit.

After: 12/12 at load ~15-42.

* no-mistakes(review): Print boundary test deadline as a decimal integer

* no-mistakes(review): Hold boundary test turn on a FIFO, require full sequence

* no-mistakes(review): Remove stray before/after supervision-host test copies

* test: hold the late close's render until the refusal window opens

The boundary recheck test's node shim slept a fixed 10s, which assumed
the first close was read before the host's refusal window opened. Under
load the close arrived after the refusal check, so the host correctly
refused it before the successor started and the render snapshot never
appeared. Block the wake-prompt render on a FIFO released at the
refusal-open instant read from the host's own start record, so the
pre-turn recheck must refuse on any machine speed.

* no-mistakes(review): Derive minimal park bounds and refresh supervision-host shard hint

* no-mistakes(review): Drive park-boundary tests from a seam-gated host test clock
* fix(bin): stage remote home clones before publishing them

A remote home provision cloned the code root directly into the public
FM_HOME path while rollback() claimed rm -rf of that same path on any
failure. Bash defers trapped signals past a foreground child, but any
other cleanup or lifecycle path that removes the home directory races
the live clone's object copy, producing the CI flake "fatal: failed to
copy file to .../.git/objects/...: No such file or directory".

Clone into a private staging directory beside the home and publish with
an atomic rename once complete, so no cleanup can remove a directory a
live clone is still writing; a home that appears mid-provision now dies
cleanly instead of inheriting torn state. The regression coverage holds
a real clone mid-copy, removes the public path, and requires the
provision to finish and publish intact.

* no-mistakes(review): Prove home ownership by sentinel and hold only a live clone

* no-mistakes(review): Assert raced provision publishes a complete, intact clone

* no-mistakes(document): Document remote home staging and publication safety

* no-mistakes(lint): Fix ShellCheck warning in clone integrity assertion

* no-mistakes(document): Clarify remote home publication and rollback guarantees
* fix(control): keep a relaunched Pi worker's herdr pane status authority alive

Defect: after `bin/fm-control.sh <id> relaunch` (observed live on a herdr
Pi crewmate whose pane read idle while it ran its validation pipeline),
the pane froze at whatever its previous agent had last reported.

Cause, measured on herdr 0.9.1 against a real Pi: a pane has one status
authority, and for Pi with its integration installed that authority is
the lifecycle hooks, so herdr also skips screen detection for the pane.
In the crew shape the registration outlives its agent process (upstream
issue kunchenguid#4115; docs/herdr-backend.md "Restart and liveness behavior"), and
herdr applies only reports carrying the session identity it bound. A
replacement started fresh in that pane reports a NEW session, so its
state reports are ignored and the pane stays frozen. Nothing from
outside repairs it: `pane report-agent-session` and `pane report-agent`
for `herdr:pi` are accepted (rc=0) without being applied unless the
reporter is the registered pane agent, and `pane release-agent` on the
stale record changes nothing.

Fix: a relaunch preserves the binding instead of fighting it. The launch
owner reads the session reference the endpoint's own runtime recorded
(`fm_backend_herdr_pane_agent_session_ref`) and passes it back as Pi's
own `--session <path-or-id>` (`relaunch_resume_args`;
`fm_control_relaunch_resume_flag` owns which adapters and which
registered-agent labels qualify). That is the same reference herdr
itself resumes Pi panes with after a server restart, and the resumed
session's reports land again, which the live check confirmed: the pane
returned to working while the replacement worked and idle when it
settled, on the same session identity.

Safety: relaunch-only (a fresh spawn binds nothing), herdr-only (the one
adapter that records a per-pane session), Pi-family only, and only when
the registration's own agent label matches - so no other adapter's
conversation can be handed to a Pi launch. An unreadable, missing, or
malformed reference degrades to exactly the fresh-session launch that
existed before. No lifecycle, liveness, isolation, or merge guard is
touched, and an empty result leaves every non-Pi launch byte-identical.
`resume` remains a refused verb; docs/agent-control.md and the
harness-adapters references are corrected where they claimed Pi had no
verified resume form at all.

* no-mistakes(document): docs: correct relaunch session-authority ownership and skill paths

* no-mistakes(document): docs: correct stale control-plane ownership claim

* no-mistakes(document): docs: drop unverified Herdr restart resume claim

* no-mistakes(test): Added offline Herdr Pi session-authority relaunch coverage

* no-mistakes(document): Document Herdr Pi relaunch session continuity

* no-mistakes(ci): The failing remote relaunch test tried to arm a PR poll for a secondmate, which `fm-pr-check.sh` correctly refuses. Removed that invalid test scenario; the remaining remote relaunch tests pass, and `git diff --check` is clean
…nchenguid#5758)

Main has been red since fm-pr-check.sh began refusing to arm a merge poll
on a kind=secondmate record (kunchenguid#5696): the relaunch-ordering case in
tests/fm-remote-secondmate-relaunch.test.sh armed its fixture through that
entry point and could no longer be set up.

The ordering guarantee still matters: a secondmate record armed before the
refusal can legitimately carry a trailing pr=/pr_head= identity block until
the watcher retires it, and fm-remote-secondmate-relaunch.sh must still keep
that block last when republishing harness/model/effort. Seed the fixture the
way such a record was really written - pr= appended last to the meta, then
the poll artifacts published through the same
fm_pr_poll_prepare/fm_pr_poll_publish_prepared pair fm-pr-check.sh uses, a
pattern tests/fm-pr-check-security.test.sh already follows - and drop the
now-unused fake gh fixture. The kunchenguid#5696 refusal itself stays pinned by the
security suite's secondmate-record case.
…id#5748)

* feat: run attended supervision on the host for Claude and Cursor

On a home opted into config/supervision-host with a Claude or Cursor
primary, the supervision host now takes the attended wakes the Pi branch
would take: routine outcomes stay off main, and a captain outcome wakes
main once with a branch-outcome line and waits in the drain's new
BRANCH OUTCOMES section until main acknowledges it with mark-processed.

- The offer rule moves into branchOfferForWake, shared by the Pi watcher
  and the host through bin/fm-branch-dispatch.mjs offer.
- The host feeds the dialog mirror at the head of each attended wake and
  passes a close through unchanged when it is main-only, the engine or a
  tool is missing, the primary has no verified mirror, the main session
  cannot be identified, or the session is cooling down.
- The drain presents captain outcomes first, one line per task, never
  behind older routine outcomes, and collapses routine overflow into a
  count that is marked read.
- The return advances the store's read cursor through the away window
  once the brief has rendered, so the first drain does not replay it.
- The branch prompt's mirror wording is host-neutral, and the rule to
  report what main must act on as captain, once per unchanged situation,
  applies only to the attended posture on the host.

* docs: record the attended supervision host live check

* no-mistakes(review): Present pre-window unread outcomes and contiguous captain prefix

* no-mistakes(review): Return brief presents every row it marks read

* no-mistakes(review): Return brief lists every unread outcome in one list

* no-mistakes(review): Keep return list in store order and gate cursor failures

* no-mistakes(review): Make the drain the only branch-outcome presenter after return

* no-mistakes(review): Gate return on drain outcome failures; byte-count outcome budgets

* no-mistakes(review): Gate drain on projection failures; UTF-8-safe byte cuts

* no-mistakes(review): Fail drain without jq; hand unreadable prompt mirror to main

* no-mistakes(document): Correct supervision-host return and drain documentation

* no-mistakes(review): Recheck attended offer at turn start; honest failed-drain brief

* no-mistakes(document): Correct supervision-host posture and drain documentation

* no-mistakes(document): Documentation remains accurate for attended supervision
…unchenguid#5753)

Each '# shellcheck source=' directive makes ShellCheck's external-source
traversal expand that library's whole transitive graph again at the site.
fm-pending-reply-lib carried three directed lazy sources of fm-wake-lib and
two of fm-parent-channel-lib on identical per-call re-source sites, so one
file analysis peaked above 4 GiB and every caller (fm-watch, fm-teardown)
inherited the multiplier - the root cause of the PR kunchenguid#5732 Lint 1 OOM kill.

Keep the runtime '.' commands byte-identical: the lazy re-source under
'local STATE FM_WAKE_QUEUE FM_WAKE_QUEUE_LOCK' is real behavior. Drop the
duplicate directives so each library expands once per unit, and drop the
tmux/classify directives since classify already arrives through the kept
fm-wake-lib expansion and no tmux symbol is referenced here. The directive
above the lib-dir assignment is kept - it binds the bin/ prefix so the
undirected sites still resolve without SC1091.

Measured peak RSS, ShellCheck 0.11.0 -x on Linux arm64:
  bin/fm-pending-reply-lib.sh  4.06 GiB -> 1.96 GiB, zero findings
kunchenguid#5773)

* fix: split bash 5.2 sibling $() in recovery mint and delivery log

Sibling command substitutions on one line can empty a recovery generation
under bash 5.2 when a CHLD trap is set. Mint pid/epoch sequentially, refuse
empty tokens before write, and clean delivery fields before printf.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: split bash 5.2 sibling $() in recovery mint and delivery log

Sibling command substitutions on one line can empty a recovery generation
under bash 5.2 when a CHLD trap is set. Mint pid/epoch sequentially, refuse
empty tokens before write, and clean delivery fields before printf.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: keep recovery mint failure semantics after sibling $() split

Remove the new pid/date refusal and grammar guard so a mint miss still
yields a grammar-valid token and a durable wake row, matching accepted
review intent. Drop the fake-failing-date case that locked in the refuse.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(document): Point recovery-mint hazard comment at its regression test

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
…unchenguid#5790)

Fixes kunchenguid#4756

The voice status reader in bin/fm_voice_records.py reports each
worker's state from the last non-blank line of its status log. When a
worker appends a status line and then a line of plain prose, the
reader reported "note" with the prose line instead of the declared
state, diverging from bin/fm-classify-lib.sh's shell scan.

Scan back through the tail for the newest line whose prefix is a
single lowercase verb-shaped word (letters and hyphens), and report
that event's verb instead of always taking the last line. An
unrecognised verb-shaped prefix still reports "note" rather than
letting an earlier recognised line answer for it, and free text with
no colon is skipped as prose. When the tail holds no such event, the
last line is reported exactly as before.
…newer version (kunchenguid#5786)

* fix(bin): stop reporting an already-installed version as an available update

An update announcement named its version first ("current -> new"), so
reading the first dotted number as the announced version compared the
current version against itself and always looked newer. Read the last
dotted number instead, and only report an available update when that
announced version is newer than the newest installed copy found; when
that version is already installed, report only PATH skew.

Fixes kunchenguid#5151

* no-mistakes(document): docs: gate announce update-available report on newer-than-installed
…h their launch config (kunchenguid#5799)

* fix(bin): pass the profile effort to OpenCode workers through their launch config

The dispatch profile's effort axis was recorded in task metadata but never
reached an OpenCode worker: the launch wrote only a permission grant into
the config it constructs.

OpenCode 1.18.32's config schema carries per-model reasoning effort as
agent.<name>.variant, so the chosen effort is now merged into the same
OPENCODE_CONFIG_CONTENT JSON as the default build agent's variant, keyed to
the resolved model. With no effort chosen the launch stays byte-identical.

Fixes kunchenguid#1373

* no-mistakes(review): gate OpenCode effort variant by model provider family

* no-mistakes(document): docs(opencode): note provider-family gating for effort variant
…nchenguid#5815)

* fix: preserve cancellation as no verdict in crew state

Reuse the green-delivery safeguard for cancelled CI monitors and permit a skipped rebase. Other cancelled outcomes and coarse ledger records use the existing unknown state.

Four delivered-PR regressions failed before the fix and pass afterward. The isolated public resolver and fleet-summary tests prove that undelivered cancellation no longer creates a failure contradiction, while preserving historical records and the terminal_in_flight invariant. Evidence uses fixture no-mistakes responses, not a live daemon cancellation.

Update the existing coarse cancellation assertion from failed to unknown because it encoded this defect; retain its newest-run precedence check. Full fm-crew-state suite and pinned lint pass.

* fix(review): Verify PR disposition before reclassifying terminal validation runs

* fix(test): Add captured cancellation replay coverage for resolver and fleet

* fix(document): Clarify cancellation and terminal delivery documentation
…henguid#5812)

* fix: declare worker background and pipeline waits

Require ship and scout workers to declare owned-work waits with the existing
paused verb before ending a turn or waiting on a pipeline or long command.
Keep the first-sight alert and existing liveness classification unchanged;
subsequent inspection follows the existing long pause cadence.

Validation: emitted brief regression failed before the instruction change
and passes afterward. Public watcher/drain regressions cover the first
alert, repeated wedge suppression, bounded rechecks, and undeclared idle
alarms using isolated backend fixtures. Brief suite, pinned lint, Bash
syntax, documentation inventory, and whitespace checks pass.
No real worker harness was exercised for wait behavior.

* fix(document): Clarify declared worker waits and documentation ownership

* fix(ci): Captain, fixed the cadence test to age both the declaration and first-alert throttle while preserving declaration identity. Reproduced the CI failure using stable identity; corrected tests pass with both stable identity and native macOS behavior. Focused ShellCheck and diff checks pass. Production behavior is unchanged; Linux CI was not rerun locally
…5770)

* fix(bin): bound each lint root in its own ShellCheck process

CI job "Lint 1" died twice at about ten minutes because the two shard
workers each packed about 110 canonical roots into one unbounded ShellCheck
process, and a byte-weight rebalance moved the analysis-heavy fm-watch.sh
into a partition with other heavy roots, so the pair outgrew the 16 GiB
runner before anything could name a culprit.

Run one canonical root per ShellCheck process under an enforced envelope:
a wall deadline plus terminate-then-kill grace via the shared
fm-timeout-lib.sh watchdog, and a per-root rlimit spec applied inside the
child before exec (default a 4 GiB address-space cap, so two workers stay
inside a 16 GiB job with headroom). A root that exceeds the envelope fails
by name with a recorded reason - timeout, memory, signal, or
limit-unavailable - instead of taking the runner down. The per-root
watchdog runs in its own process group so the owner's group sweep cannot
orphan the bounded subtree, and fm_exec_timed now starts the same
escalation when its parent dies before it can be signalled.
FM_LINT_REQUIRE_BOUNDS=1, set in CI, refuses the run outright when a
configured bound cannot be enforced on the host rather than lint uncapped.
Each root's begin/end, reason, duration, and peak RSS stream to stderr in
partition mode and append to a retained <telemetry>.roots.tsv sidecar
uploaded beside the partition telemetry.

Coverage is unchanged: pinned ShellCheck 0.11.0, --norc, --external-sources
full analysis, complete and disjoint partition inventory, workflow lint,
and the backend-purity check, with byte-identical diagnostics across
jobs=1/2 proven by tests/fm-lint.test.sh.

* fix(bin): fail closed on unenforceable lint bounds and size the cap

Required-bounds mode (FM_LINT_REQUIRE_BOUNDS=1, set by CI) now refuses the
run with named errors before any root starts: a missing fm-timeout-lib.sh,
a watchdog that cannot actually bound a probe command, or a host that
rejects the address-space limit all stop the run rather than lint uncapped.
The generalized FM_LINT_ROOT_RLIMITS flag:value interface is replaced by a
single FM_LINT_ROOT_MEMORY_KIB, and the roots sidecar and telemetry record
the run's final exit status after backend-purity and workflow checks
instead of the pre-check lint status.

The default cap is 6 GiB of address space per root, not 4 GiB: ulimit -v
bounds virtual address space rather than resident memory, and ShellCheck's
GHC runtime keeps roughly a third of that space as reservation, so 6 GiB
yields about a 4 GiB working heap budget. A Linux measurement during this
change showed eleven real canonical roots running out of memory under the
earlier 4 GiB cap while the largest passing root peaked near 2.8 GiB
resident; two 6 GiB roots plus runner overhead still fit the 16 GiB job.
Roots that still exceed the cap keep failing by name, and the sidecar's
per-root peak RSS keeps roots approaching the budget visible.

tests/fm-lint.test.sh now proves the memory primitive where it can be
proven: on hosts that accept ulimit -v a perl allocator is refused under a
256 MiB limit and reported by name as a memory death, the pinned ShellCheck
lints a small file under the configured cap and is named when a far smaller
cap binds it, and a watchdog-less copy refuses under REQUIRE_BOUNDS; the
bounded cases skip on macOS, which cannot enforce the address-space limit.

* no-mistakes(review): Prove memory cap binds, pass watchdog owner, drop unused modes

* no-mistakes(review): Capture watchdog owner before startup for every fm_exec_timed caller

* no-mistakes(document): Clarify bounded lint documentation and telemetry

* no-mistakes(document): Correct bounded lint documentation and sidecar path

* docs(bin): restore the per-root memory cap sizing rationale

The pipeline's document step rewrote the ROOT_MEMORY_KIB comment and
dropped the sizing reasoning the change is required to record: address
space vs resident memory, the GHC reservation share, the measured 4 GiB
failures and ~2.8 GiB peak, and the two-roots-plus-runner capacity
arithmetic. Restore it beside the default while keeping the corrected
"not a resident-memory ceiling" framing.

* no-mistakes(review): Document memory cap RSS reduction threshold and first candidate

* no-mistakes(review): Scope owner-death escalation docs to the perl watchdog

* no-mistakes(document): Clarify bounded lint and timeout documentation

* no-mistakes(review): Install perl watchdog signal handlers before forking the command

* no-mistakes(document): Correct bounded lint documentation and stale watcher comments

* no-mistakes(ci): Fixed the supervision-host test’s obsolete expectation: the watchdog now reaps an engine when its host dies. The timeout and supervision-host tests pass locally; the watcher test also passes locally. Lint 1 and 2 remain unresolved: seven canonical roots exceeded the required 6 GiB address-space cap in CI. I did not raise the cap, exempt roots, or reduce source-following coverage to make those failures disappear

* no-mistakes(ci): The two lint checks failed when eight canonical roots hit the enforced memory cap. I reduced repeated ShellCheck source-graph expansion while keeping runtime imports and the canonical root inventory intact. Pinned ShellCheck passes for all changed roots; the relevant local tests pass. The 6 GiB Linux CI run remains unverified

* no-mistakes(review): Restore source directives, raise cap to 8 GiB, classify OOM

* no-mistakes(review): Classify memory deaths from root stderr, not source excerpts

* no-mistakes(review): Match only whole runtime memory-error lines for memory reason

* no-mistakes(document): Clarify lint memory classification in script documentation

* no-mistakes(ci): Fixed both lint checks’ memory-limit failure: each CI lint job now runs one root at a time with a 12 GiB address-space cap. Kept the local two-worker default and updated the sizing comment and test expectation. The lint tests and workflow validation pass locally; Linux CI remains to confirm the heavy roots
…uid#5732)

* fix(bin): bound the watcher cleanup marker-lock wait

tests/fm-watch-triage.test.sh intermittently failed serial CI shard 1
with "watcher pid <pid> did not exit within 10s of TERM". The watcher
had processed the TERM and was inside watcher_cleanup, where the
recovery-marker publish waits on state/.watcher-down.lock through an
unbounded fm_lock_acquire_wait. A live foreign holder of that lock
leaves the TERM'd watcher spinning in its own EXIT trap until the lock
frees or a second signal short-circuits the trap.

fm_recovery_transition now takes an optional bound and both
release-lock paths plus publish honour it through a new in-process
fm_lock_acquire_wait_max. watcher_cleanup passes
FM_WATCHER_CLEANUP_LOCK_BOUND (default 2s); on timeout the publish is
skipped, the singleton stays behind as ordinary dead-pid evidence, and
the next arm's clear-stale-lock still republishes it.

Regression test drives a real watcher with .watcher-down.lock held by
a live foreign process and asserts a single TERM still stops it.

* no-mistakes(review): Parse watcher cleanup lock bound as decimal, zero defaults

* no-mistakes(review): Pin cleanup bound tests to observed marker-lock contention

* no-mistakes(document): Document bounded watcher cleanup and recovery

* no-mistakes(document): Clarify bounded watcher cleanup and recovery documentation

* no-mistakes: apply agent fixes

* no-mistakes(review): Arm marker-lock FIFO before TERM; drop FM_TEST_ONLY_LATE

* no-mistakes(review): Hold marker lock through a failed cleanup acquire
…kunchenguid#5845)

* test: stop the leaked unreachable watcher before remote e2e cleanup

The remote secondmate lifecycle e2e backgrounded fm-watch.sh through the
remote_env shell function, so $! named the function's subshell rather than
the watcher. Killing that subshell left the unreachable-leg watcher running,
and its one-second liveness probe kept invoking the fake ssh, which rewrites
ssh.count in the temp root. When a probe landed while the EXIT trap was
removing the root, rm failed with "Directory not empty" after every
assertion had passed.

Exec the watcher from the backgrounded function so the recorded pid is the
watcher itself, and assert the stopped watcher stops probing and writing its
state. Cleanup also stops a watcher left running by a failed assertion and
removes the root through fm_test_remove_tree, so a run that fails before
retirement does not strand the read-only spawn hooks directory.

Closes kunchenguid#5836

* no-mistakes(review): Clear reaped watcher PIDs and restore plain temp-root removal

* no-mistakes(review): Let in-flight probe settle before stopped-watcher baseline
…nchenguid#4806)

* fix: stop quarantining ordinary shared-captain source updates

* no-mistakes(document): Rewrap remote inherit header so usage prints fully
* docs: make calm easier to read

Restructure the Calm mode prose into sections, lists, and tables without changing documented behavior. Every original heading, anchor, fenced code block, inline-code span, link target, number, and quoted string is preserved.

* docs: restore reload case in calm override lead-in

The Built-in tool override collisions lead-in covers a session that reloads with Calm already on, as the original text did.
* docs: make turnend-guard easier to read

Restructure the turn-end guard doc's prose into shorter sections, lists, and tables without changing documented behavior. Every original heading, anchor, inline identifier, link target, and number is kept.

* no-mistakes(review): Restore legacy-only scope on TERM retirement sentence

* no-mistakes(review): Name Cursor park behavior in live e2e test line
…henguid#5872)

* docs: move situational AGENTS.md sections into on-demand skills

Backpass memory optimization: shrink the always-loaded AGENTS.md by moving
situational contracts (home layout, session-start recovery, validation and
landing supervision, scout completion, away/quiet supervision, Relay
ownership) into agent-only skills loaded at their triggers, with a trigger
index skill.

* docs: classify the new on-demand skills' documentation audience

Register the seven new agent-only skills as agent-runtime docs and fix a
link in validation-supervision that kept its AGENTS.md-relative path.

* docs: close load-timing gaps found by the live regression check

- load validation-supervision whenever an ask-user finding is decided or
  answered, so forbid --yes and process-every-return reach the worker
- keep the mid-task captain-ask rule, the unconfirmed network-checks rule,
  and the worker account pin rule inline in AGENTS.md
- fix cross-references that still pointed at moved AGENTS.md sections
…guid#5879)

* fix: route second-mate signal wakes by their new status span

A second mate's status log is a shared channel carrying many independently
keyed decisions, so judging its signal rows by every decision still open in
the whole log pinned each routine update to main behind any unrelated
parked hold. scopeForUnreadWake (the one owner for Pi and the attended
supervision host) now judges a second-mate signal row by the lines presented
since the last drain, bounded by the existing status-presentation cursor:
a decision, blocked, resolution, or captain-held line, or a line declaring
the key of a still-open decision, keeps the whole row on main, and any
cursor problem falls back to the whole log. Keys are read only at the
status parser's declared positions, with readable time stamps stripped as
bin/fm-classify-lib.sh does. Single-task crewmate and stale routing are
unchanged, and stale and signal rows for one mate keep independent verdicts.

The supervision branch now treats a second mate's done and merged lines as
relayed child outcomes, and fm-teardown refuses the branch actor second-mate
retirement through the existing role-partition helper in both postures.

* no-mistakes(review): Route second-mate resolutions to main only when closing open decision

* no-mistakes(review): Guard bare-verb fold lines and fold resolution spans incrementally

* no-mistakes(document): Clarify second-mate wake routing and retirement documentation

* no-mistakes(ci): The CI failure was a timing-sensitive watcher teardown test, not the PR’s signal-scope code. Extended the bounded wait for both state-directory and home removal. The watcher suite passes locally, and git diff --check is clean
…extension (kunchenguid#5882)

register-extension took the extension lifecycle lock and then the source
lock, while reconcile republishing an unhandled extension result holds the
source lock and reaches the lifecycle lock through the extension host's
process-event path. Both waits are unbounded and both owners stay alive, so
the two could wait on each other forever and freeze the home's monitoring
cycle.

register-extension now takes the source lock first, matching every other
path that holds both. The lifecycle lock still spans binding resolution
through registration publication, so binding retirement stays serialized.

A new lifecycle-order section in the extension-binding suite, run in the
default aggregate, holds a re-registration inside binding resolution while
reconcile republishes that source's unhandled result and requires both to
finish within a bound.

Fixes kunchenguid#5866
* fix(bin): run the repository's own hooks when git -c carries the per-task hooksPath

The per-task hooks wrappers cleared only the GIT_CONFIG_COUNT override before
looking up the repository's own hooks directory. When core.hooksPath reached git
through git -c (GIT_CONFIG_PARAMETERS), directly or inherited by a child
process, the lookup found the wrapper directory again and exited 0, so the
repository's real hook - such as a pre-push publish guard - never ran and the
push succeeded.

The lookup now ignores GIT_CONFIG_PARAMETERS too, so only the repository's
config files decide its hooks directory, and a failed lookup exits nonzero
instead of skipping the hook. AI-trailer stripping is unchanged.

Fixes kunchenguid#5871

* no-mistakes(document): Document git hook chaining and lookup failure behavior
…unchenguid#5744)

* feat(bin): add an optional never-send list to typed dispatch resolution

* Added config/dispatch-never-send, an optional local list of literal
  values and re: regular expressions checked against every string of
  the resolver request before it is sent to typesafe.ai
* A match, an unreadable list, or an empty or invalid pattern now stops
  the request and falls back to the off path, so firstmate dispatches
  through its existing intake; the one stderr diagnostic names at most
  the list line number and never the value
* No list, or a list with no match, leaves resolution unchanged

* no-mistakes(review): Match never-send literals across whitespace, drop regex mode

* no-mistakes(review): Inherit the never-send list into secondmate homes

* no-mistakes(ci): ci-2 (Lint 2), caused by this PR and now fixed. The rule that broke: the test script must not share shell variables with a library it sources. The new secondmate-inheritance test in tests/fm-dispatch-resolve.test.sh sourced bin/fm-config-inherit-lib.sh inside a `( ... )` subshell. That lib assigns `out`, so ShellCheck flagged every later `$out` in the test with SC2031. The subshell was the only place this PR sources that lib. The fix runs the propagation in a child shell instead (`bash -c '. "$1" && propagate_inheritable_config "$2" "$3"' _ lib from to`), so the test shell never sources the lib. `bin/fm-lint.sh tests/fm-dispatch-resolve.test.sh` now exits 0, and `bash tests/fm-dispatch-resolve.test.sh` passes. That includes the check that an inherited list is enforced in a secondmate home. ci-1 (Behavior portable serial 8), not caused by this PR, so no code change for it. The only failing test is tests/fm-remote-secondmate-relaunch.test.sh, which fails with "not ok - could not arm the PR poll fixture for the relaunch-ordering test". This PR doesn't touch that test or the code it runs. I reproduced the same failure locally on the base commit ea7c7f7. Main's own CI run on ea7c7f7 (run 36212602588) fails only this job, with the same message. The recent main runs before it also concluded failure
Upstream moved AGENTS.md's PR-ready and landing block into the ship-landing skill; the fork's ready-signal evidence wording now lives there.
Copilot AI lite review requested due to automatic review settings September 27, 2026 16:29

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Resolve the bin/fm-brief.sh header conflict by keeping main's gate-delegation warning note and the scope-allowance pointer to the validation-supervision skill.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 5b9bae9082

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread bin/fm-afk-return.sh
# so the brief counts them and points there instead of listing them, or says
# they await a successful drain when this return's drain failed.
# shellcheck source=bin/fm-supervision-engine-lib.sh
if . "$SCRIPT_DIR/fm-supervision-engine-lib.sh" \

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Stage the new library in the shared return-test fixture

When any pre-existing fm-afk-return fixture reaches render_return_brief, this unconditional source resolves under the fixture's isolated bin/, but the shared install_runner does not copy fm-supervision-engine-lib.sh. Only the two newly added cases stage it ad hoc, so bash tests/fm-afk-return.test.sh aborts in the earlier cases with No such file or directory and leaves the test suite red. Add the new dependency and its required libraries to the shared fixture, or avoid sourcing it unconditionally.

Useful? React with 👍 / 👎.

Comment thread bin/fm-spawn.sh
low | medium | high | xhigh | max) printf -- '--thinking %s ' "$(shell_quote "$effort")" ;;
esac
;;
opencode)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Accept OpenCode efforts in dispatch-profile validation

When OpenCode is selected through config/crew-dispatch.json with the newly documented model and supported effort, both bin/fm-bootstrap.sh's and bin/fm-dispatch-resolve.sh's effort_ok functions still categorically reject every OpenCode effort; the bootstrap test even continues asserting that opencode plus high is invalid. Consequently, configured routing cannot reach this new agent.build.variant launch path even though direct spawn flags can, so update both profile validators to accept the model-family effort combinations implemented here.

AGENTS.md reference: AGENTS.md:L94-L95

Useful? React with 👍 / 👎.

[ -n "$REMOTE_HOST" ] \
|| die "task $ID is not a remotely placed secondmate; use bin/fm-control.sh $ID relaunch instead"

RELAUNCH_OUT=$("$SCRIPT_DIR/fm-on.sh" "$ID" fm-remote-secondmate-control.sh \

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Serialize the parent-side remote relaunch transaction

When two remote relaunches for the same mate overlap, this starts the host mutation before acquiring any parent-side lifecycle lock. Although the remote host serializes its own operations, parent publication can reorder afterward—for example, relaunch A returns, relaunch B completes and publishes, and then A publishes—leaving state/<id>.meta naming A while the endpoint actually runs B. Acquire the parent's .control-$ID.lock before reading metadata and hold it through the remote operation and metadata publication.

Useful? React with 👍 / 👎.


# Agent-only reference skills

These skills are not captain-invocable; load them only at their precise triggers.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Include every agent-only skill in the trigger index

When this skill is loaded for the advertised complete trigger audit, its list omits the newly added user-invocable: false skills operational-home-layout, session-start-recovery, away-quiet-supervision, validation-supervision, ship-landing, and scout-completion. Those skills contain required triggers for startup recovery, away supervision, validation decisions, landing, and cleanup, so an audit driven by this new index can incorrectly conclude that those runtime contracts are covered and allow their triggers to drift. Include all agent-only skills or stop calling this the complete index.

AGENTS.md reference: AGENTS.md:L405-L406

Useful? React with 👍 / 👎.

ours=$(quote_for_hook "$ours")
name=\${0##*/}
orig=\$(git rev-parse --path-format=absolute --git-path hooks) || exit 0
orig=\$(unset GIT_CONFIG_PARAMETERS; git rev-parse --path-format=absolute --git-path hooks) || {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Preserve command-scoped Git safety configuration in hook chaining

When a worker commits in a differently owned or container-mounted repository using git -c safe.directory=<repo>, this unsets the entire GIT_CONFIG_PARAMETERS value before the generated hook's nested lookup, rather than removing only the recursive core.hooksPath entry. Git 2.43's git --help lists [-c <name>=<value>] as a global option, and the failure reproduces here: the outer commit accepts safe.directory and enters the hook, then git rev-parse loses it, reports dubious ownership, and aborts the commit. Preserve unrelated command-scoped settings or resolve the repository hook path without discarding them.

Useful? React with 👍 / 👎.

Comment thread bin/fm-wake-drain.sh
printf 'BRANCH OUTCOMES SKIPPED: jq is not installed, so the outcome store cannot be presented; nothing was marked read, and these outcomes are presented once jq is back.\n' >&2
return 1
fi
if ! rows=$("$SCRIPT_DIR/fm-branch-outcome.sh" present 2>/dev/null); then

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Run the processed-marker migration before presenting outcomes

When a supervision-host home already has a read cursor but no .branch-outcomes-processed marker, this new call treats every historical captain row as unprocessed unless processed-init has run. The only drain call to processed-init is inside the status-outcome-backstop path, which print_status_presentation skips when there are no live status snapshots, as normally occurs after task cleanup. Reusing a Pi-originated home under Claude or Cursor in that state therefore reissues already-delivered captain outcomes as new actions; call processed-init before present rather than relying on an unrelated status section.

Useful? React with 👍 / 👎.

@zakna
zakna merged commit acde313 into main Sep 27, 2026
18 of 19 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.