Skip to content

ci: bound Herdr behavior shard with a 20-minute step timeout - #2413

Merged
kunchenguid merged 1 commit into
mainfrom
fm/fm-herdr-shard-step-timeout-r1
Aug 15, 2026
Merged

kunchenguid merged 1 commit into
mainfrom
fm/fm-herdr-shard-step-timeout-r1

Conversation

@kunchenguid

@kunchenguid kunchenguid commented Aug 15, 2026 •

Copy link
Copy Markdown
Owner

Intent

Bound firstmate's Herdr behavior CI shard with a step-level timeout so a hung suite fails fast with evidence instead of occupying a runner until the 75-minute job cap.

From the 2026-08-14 fleet CI audit (captain-approved to dispatch): firstmate's Behavior tests (Herdr) CI job relies on a 75-minute JOB cap (added in firstmate#2191) as its only stop - but healthy runs finish in ~7 min, so 75 min is a HANG TRIPWIRE, not a target: a wedged shard holds a runner for the full 75 min before cancel (e.g. run 31524483885, ~92-min wall).

Fix: add a STEP-LEVEL timeout INSIDE the Herdr behavior shard/script so a hung shard FAILS FAST with evidence instead of occupying a runner to the 75-min cap. Choose a bound comfortably above the ~7-min healthy time but far below 75 min. Keep the job cap as a backstop. Do NOT touch the 479 action_required runs - those are fork-PR approvals, not flakes, and out of scope.

This edits firstmate's shared CI / Herdr material. Ship no-mistakes +yolo; merge on green. Do not auto-fix unrelated flaky tests such as tests/fm-procevent.test.sh.

What Changed

  • Add a 20-minute timeout-minutes on the required Herdr family-run CI step so a hung suite fails that step (cleanup and timing artifacts still upload) instead of occupying a runner until the 75-minute job cap.
  • Keep the 75-minute job timeout as a last-resort backstop and document the two-level Herdr bound in the portable-shards timeout table.
  • Add a YAML-parsed regression test that asserts the family-run step timeout is 20 minutes and the job backstop stays 75.

Risk Assessment

✅ Low: The Herdr family-run step is now bounded at 20 minutes under the kept 75-minute job backstop, which is the requested fail-fast contract and leaves always() cleanup and artifact upload intact.

Testing

Parsed the real tests-herdr workflow as YAML, reproduced the missing step timeout on the base commit, confirmed the official contract test passes after the fix, and simulated a wedged family-run so a hang now fails at 20 minutes with cleanup and timing artifacts still uploaded while the 75-minute job cap stays the backstop.

Evidence: Reviewer-facing hang-tripwire comparison (base vs target)

Source: Reviewer-facing hang-tripwire comparison (base vs target)

# Herdr behavior shard: step-level hang tripwire

End-user of this change: a CI operator watching a wedged `Behavior tests (Herdr)` job.

## What the runner will do

Parsed `.github/workflows/ci.yml` as YAML (Psych / the same document GitHub Actions consumes). Not a text grep.

| | Base `6789876` | Target `221aebd` |
|---|---|---|
| Job `tests-herdr` `timeout-minutes` | 75 | 75 (kept as backstop) |
| Family-run step `timeout-minutes` | **missing** | **20** |
| Cleanup `if:` | `always()` | `always()` |
| Upload timing/diagnostics `if:` | `always()` | `always()` |

GitHub Actions semantics (docs.github.com workflow syntax):

- **Step** `timeout-minutes`: "The maximum number of minutes to run the step before killing the process."
- **Job** `timeout-minutes`: "The maximum number of minutes to let a job run before GitHub automatically cancels it."

A step kill leaves the job running, so the two `if: always()` steps after the family-run still execute. A job-cap cancel is the whole job stopping — that is the 75-minute hang the audit called out (run 31524483885).

20 minutes is above the ~7-minute healthy wall and far below the 75-minute backstop.

## Contract test (official suite function)

`tests/fm-test-run.test.sh::test_herdr_ci_family_run_has_a_step_timeout`

- Against **base** workflow: `not ok - could not parse tests-herdr timeouts from ci.yml` (`family-run step has no timeout-minutes`). Exit 1.
- Against **target** workflow: `ok - Herdr CI family-run step times out at 20 min under a 75 min job backstop`. Exit 0.

The reported failure exists before the fix and is gone after it.

## Hung-suite experience (scaled clock)

1 CI minute = 80 ms wall so a 75-minute occupation is observable without burning a runner.

`` `
BASE (job cap only)
  t=    0 min  family-run starts and wedges
  t=   75 min  job timeout-minutes=75 cancels the whole job (only stop)
  t=   75 min  cleanup/upload not a reliable evidence path
  => runner occupied 75 CI minutes; evidence_uploaded=false

TARGET (this change)
  t=    0 min  family-run starts and wedges
  t=   20 min  step timeout-minutes=20 fires; hung family-run is killed
  t= 20.5 min  Cleanup job-owned Herdr lab sessions ran (if: always())
  t= 21.0 min  Upload Herdr timing and diagnostics ran (if: always())
  => runner occupied 20 CI minutes; evidence_uploaded=true
`` `

A hang now fails ~55 runner-minutes sooner, with cleanup and the timing/diagnostics artifact still uploaded. The 75-minute job cap remains the last-resort backstop.

## Scope (forbidden work)

Changed files vs base: `.github/workflows/ci.yml`, `docs/fm-test-portable-shards.md`, `tests/fm-test-run.test.sh`.

- `tests/fm-procevent.test.sh` not touched.
- Diff does not mention `action_required` (the 479 fork-PR approvals stay out of scope).
Evidence: Scaled hung-suite timeline: 75-minute job cancel vs 20-minute step fail-fast with evidence

Source: Scaled hung-suite timeline: 75-minute job cancel vs 20-minute step fail-fast with evidence

BASE (job cap only) t= 0 min family-run starts and wedges t= 75 min job timeout-minutes=75 cancels the whole job (only stop) => runner occupied 75 CI minutes; evidence_uploaded=False TARGET (this change) t= 0 min family-run starts and wedges t= 20 min step timeout-minutes=20 fires; hung family-run is killed t= 20.5 min Cleanup job-owned Herdr lab sessions ran (if: always()) t= 21.0 min Upload Herdr timing and diagnostics ran (if: always()) => runner occupied 20 CI minutes; evidence_uploaded=True

Herdr hang tripwire — compressed end-user simulation
Time scale: 1 CI minute = 0.08s wall (so a 75-min occupation is observable)

BASE (before the change): tests-herdr has only timeout-minutes: 75 on the job.
A wedged family-run holds the runner until that job cap. Cleanup/upload are
not a reliable evidence path once the job itself is cancelled.

  t=    0 min  BASE (job cap only): family-run starts and wedges (suite never returns)
  t=   75 min  job timeout-minutes=75 cancels the whole job (only stop)
  t=   75 min  job cancelled at cap; cleanup/upload not a reliable evidence path
  => runner occupied 75 CI minutes; evidence_uploaded=False

TARGET (this change): family-run step timeout-minutes: 20; job stays 75.
A wedged family-run is killed at 20 minutes. The job is still running, so
if: always() cleanup and artifact upload still execute.

  t=    0 min  TARGET (step tripwire + job backstop): family-run starts and wedges (suite never returns)
  t=   20 min  step timeout-minutes=20 fires; hung family-run is killed
  t= 20.5 min  Cleanup job-owned Herdr lab sessions ran (if: always())
  t= 21.0 min  Upload Herdr timing and diagnostics ran (if: always())
  => runner occupied 20 CI minutes; evidence_uploaded=True

Saved ~55 runner-minutes on a hang. Job cap remains the last-resort backstop.
Evidence: Official contract test on the target workflow

Source: Official contract test on the target workflow

ok - Herdr CI family-run step times out at 20 min under a 75 min job backstop exit=0

repo=/Users/kunchen/.no-mistakes/worktrees/016d88035d58/01M01T1AM9F508XFJAHTK1GN6S
2026-08-15T04:28:24Z
ok - Herdr CI family-run step times out at 20 min under a 75 min job backstop
exit=0
Evidence: Official contract test on the base workflow (must fail)

Source: Official contract test on the base workflow (must fail)

family-run step has no timeout-minutes (RuntimeError) not ok - could not parse tests-herdr timeouts from ci.yml exit=1

repo=/var/folders/0k/bf8mwt2n5qddzk24r20gfk0c0000gn/T/no-mistakes-evidence/01M01T1AM9F508XFJAHTK1GN6S/base-tree
2026-08-15T04:28:24Z
-e:8:in `<main>': family-run step has no timeout-minutes (RuntimeError)
not ok - could not parse tests-herdr timeouts from ci.yml
exit=1
Evidence: Parsed tests-herdr timeout model (target)

Source: Parsed tests-herdr timeout model (target)

{
  "job": "tests-herdr",
  "job_timeout_minutes": 75,
  "family_run_step": {
    "name": "Run real-Herdr family (serial, required)",
    "timeout_minutes": 20,
    "has_timeout": true
  },
  "cleanup_step": {
    "name": "Cleanup job-owned Herdr lab sessions",
    "if": "always()",
    "runs_after_step_failure_or_cancel": true
  },
  "upload_step": {
    "name": "Upload Herdr timing and diagnostics",
    "if": "always()",
    "runs_after_step_failure_or_cancel": true
  },
  "contract": {
    "job_backstop_minutes": 75,
    "step_tripwire_minutes": 20,
    "step_is_tighter_than_job": true,
    "healthy_runs_about_minutes": 7,
    "step_above_healthy_runtime": true,
    "step_far_below_job_cap": true,
    "cleanup_and_upload_survive_step_timeout": true
  }
}

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

✅ **Test** - passed

✅ No issues found.

  • ruby parse-herdr-timeouts.rb semantic YAML model of .github/workflows/ci.yml jobs.tests-herdr (target 221aebd vs base 6789876)
  • bash test_herdr_ci_family_run_has_a_step_timeout.sh against the target workflow — ok - Herdr CI family-run step times out at 20 min under a 75 min job backstop
  • bash test_herdr_ci_family_run_has_a_step_timeout.sh against the base workflow — not ok / family-run step has no timeout-minutes (regression exists before the fix)
  • python3 simulate-herdr-hang.py scaled hung-suite timeline: base holds the runner 75 CI minutes with no reliable evidence path; target kills the step at 20 minutes and still runs if: always() cleanup + artifact upload
  • scope check that the diff does not touch tests/fm-procevent.test.sh or action_required
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

A wedged family-run step was occupying the runner until the 75-minute
job cap; bound that step so cleanup and timing artifacts still upload.
@kunchenguid
kunchenguid merged commit f1a4af4 into main Aug 15, 2026
24 of 25 checks passed
@kunchenguid
kunchenguid deleted the fm/fm-herdr-shard-step-timeout-r1 branch August 15, 2026 04:32
prajwal-395 added a commit to prajwal-395/firstmate that referenced this pull request Aug 15, 2026
* fix(ci): fail hung Herdr behavior runs in 20 minutes (kunchenguid#2413)

A wedged family-run step was occupying the runner until the 75-minute
job cap; bound that step so cleanup and timing artifacts still upload.

* test: add spawn-guard placeholder regression test for fm-brief.sh

Add a regression test to tests/fm-brief.test.sh that prevents
bin/fm-brief.sh from ever again emitting a scaffold token that its
sibling guard in bin/fm-spawn.sh then rejects.

The spawn guard (PR 5) refuses to launch when any {[A-Z_]+} token
remains in the brief file.  PR 8 fixed a case where the Herdr
NOT-ENABLED section contained a literal {TASK} in prose, creating a
second occurrence that survived firstmate filling the real placeholder
and blocked dispatch for every non---herdr-lab brief.

Coverage spans every scaffold variant:
- ship (no-mistakes, direct-PR, local-only)
- scout
- secondmate charter (project-list and --no-projects)
- each with and without --herdr-lab where applicable

For each variant the test asserts exact per-token occurrence counts
against the spawn guard regex (read from bin/fm-spawn.sh, not
restated), catching both novel tokens and extra occurrences of
expected ones - the exact original bug class.

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: Firstmate Tests <tests@example.invalid>
undeemed pushed a commit to undeemed/firstmate that referenced this pull request Aug 15, 2026
A wedged family-run step was occupying the runner until the 75-minute
job cap; bound that step so cleanup and timing artifacts still upload.
prajwal-395 added a commit to prajwal-395/firstmate that referenced this pull request Aug 15, 2026
* fix(ci): fail hung Herdr behavior runs in 20 minutes (kunchenguid#2413)

A wedged family-run step was occupying the runner until the 75-minute
job cap; bound that step so cleanup and timing artifacts still upload.

* Add Herdr orphan pane reaper

Leaked Herdr panes accumulate when a task's agent exits, is stopped, or
its record is removed without teardown - from that point nothing owns
the pane and nothing ever sweeps for it.

bin/fm-herdr-orphan-reaper.sh identifies orphaned panes by comparing
Herdr's live tab list (scoped to this home's own workspace) against the
herdr_pane_id= lines in state/*.meta.  Safety refusals:
- Never closes a pane claimed by any state/<id>.meta in this home.
- Never closes a pane with a live agent (working or blocked).
- Never touches panes outside this home's workspace.
- Never considers tabs without the fm-<id> label (captain terminals).
- Gates itself on backend=herdr; non-herdr homes get an early exit.

Supports --report (default, list-only) and --close (actually close).

Wired into bin/fm-bootstrap.sh as a local-phase mutating sweep, gated
on the session lock (FM_BOOTSTRAP_DETECT_ONLY), so orphans are swept
automatically at every locked session start.

tests/fm-herdr-orphan-reaper.test.sh covers all four safety refusals,
orphan detection, close mode, mixed panes, non-herdr skip, and the
home-scoping boundary with a stateful fake herdr CLI.

* Silence routine reaper output on non-herdr and zero-orphan paths

Bootstrap deliberately keeps routine confirmations silent. The reaper's
non-herdr skip, no-workspace-found, and zero-orphan messages are all
routine no-ops that violated this contract, breaking three existing
tests that assert silence on those paths.

Remove all three diagnostic lines. The reaper now produces output only
when it has something actionable to report: orphan candidates found,
panes closed, or refusals with reasons.

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
prajwal-395 added a commit to prajwal-395/firstmate that referenced this pull request Aug 15, 2026
* fix(ci): fail hung Herdr behavior runs in 20 minutes (kunchenguid#2413)

A wedged family-run step was occupying the runner until the 75-minute
job cap; bound that step so cleanup and timing artifacts still upload.

* feat: Add quota memory extraction for agy harness

Extracts model, quota, and reset time from agy pane footer and stores
it persistently in state/ so dispatch rules can read it.
Fails open on ambiguous reads and resets on staleness.

* fix: Support set -u with optional now_override arg in fm_agy_quota_read

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
joliverMI added a commit to joliverMI/firstmate that referenced this pull request Aug 16, 2026
…ng state (#2)

* fix(ci): fail hung Herdr behavior runs in 20 minutes (kunchenguid#2413)

A wedged family-run step was occupying the runner until the 75-minute
job cap; bound that step so cleanup and timing artifacts still upload.

* fix: keep the public promise reachable when work is routed to a second mate (kunchenguid#2457)

The lightweight Relay follow-up link lives in the answering home's own
state/<task-id>.meta, so it can only bind work that home owns. When a
Relay-linked request is routed to a second mate, the task record lives in the
second mate's home, fm-x-link.sh failed with a bare "no such task ...meta", and
nothing else picked the promise up: only the soft acknowledgement was ever
posted. The typed promised-final path already supports --work-home
secondmate:<id>; the playbook simply never chose it.

- fmx-respond now states the routing rule crisply: a task in this home takes the
  lightweight link, and second-mate-routed work takes a promised-final
  commitment bound to that home, registered up front with the brief command
  carried into the routed worker's instructions.
- fm-x-link.sh refuses a task with no local record by naming the registered
  second mate whose home actually holds it and printing the promised-final
  registration command, with the exact --work-home when the match is
  unambiguous. A home with no registered second mates keeps the plain error.
- fm-backlog-handoff.sh reports, after a successful move, any moved key that
  still owes a public reply bound to main/<key>, since that binding no longer
  names the home owning the work. The move itself is never blocked.

Docs and the secondmate handoff prose follow the same rule. Tests cover the
refusal, its scoping, the unchanged local-link path, and both handoff outcomes
at the script boundary.

* docs(skills): add remote-secondmate recovery hint for false-negative verdicts (kunchenguid#2456)

* fix(skills): hint that remote secondmate liveness verdicts false-negative

fm-crew-state and fm-send routinely misreport a live remote secondmate
as dead; confirm against the pane before relaunching, and relaunch only
through fm-spawn.sh, never raw herdr pane surgery.

* no-mistakes: apply CI fixes

* fix(calm): keep Pi's export confirmation visible (kunchenguid#2461)

Pi 0.83.0 added a status line to every tool-expansion change, and Pi
updates the previous status line in place when two status messages
arrive back to back. Calm's post-export redraw cycled tool expansion on
the macrotask right after Pi printed "Session exported to: <path>", so
both expansion status lines coalesced over that confirmation and the
captain was left with no record of where their export landed.

Calm now repaints only the tool rows it presents, by invalidating each
row through the render context Pi hands its render slots, and requests
the surrounding redraw through setStatus. Neither appends to the
transcript. The repaint is still needed because Pi can re-render a row
asynchronously - the built-in edit row invalidates itself once its diff
is ready - and that re-render can land inside the window where /export
forces stock rendering.

The real-terminal /export case now asserts the confirmation is still on
screen after the redraw has settled, and that the redraw restored every
Calm-hidden row, instead of only racing the moment the confirmation
first appeared.

* feat(bin): render the captain's four-section status board from existing state

fm-status-board.sh prints one HTML page (Needs you to continue, In
progress, Waiting, Recently completed) rendered entirely from
fm-bearings-snapshot.sh's already-correct cross-home classification and
fm-fleet-snapshot.sh's untruncated per-item detail, with no second store
for agents to keep in sync.

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: joliverMI <joliver@sensibletech.biz>
joliverMI added a commit to joliverMI/firstmate that referenced this pull request Aug 16, 2026
* fix(ci): fail hung Herdr behavior runs in 20 minutes (kunchenguid#2413)

A wedged family-run step was occupying the runner until the 75-minute
job cap; bound that step so cleanup and timing artifacts still upload.

* fix: keep the public promise reachable when work is routed to a second mate (kunchenguid#2457)

The lightweight Relay follow-up link lives in the answering home's own
state/<task-id>.meta, so it can only bind work that home owns. When a
Relay-linked request is routed to a second mate, the task record lives in the
second mate's home, fm-x-link.sh failed with a bare "no such task ...meta", and
nothing else picked the promise up: only the soft acknowledgement was ever
posted. The typed promised-final path already supports --work-home
secondmate:<id>; the playbook simply never chose it.

- fmx-respond now states the routing rule crisply: a task in this home takes the
  lightweight link, and second-mate-routed work takes a promised-final
  commitment bound to that home, registered up front with the brief command
  carried into the routed worker's instructions.
- fm-x-link.sh refuses a task with no local record by naming the registered
  second mate whose home actually holds it and printing the promised-final
  registration command, with the exact --work-home when the match is
  unambiguous. A home with no registered second mates keeps the plain error.
- fm-backlog-handoff.sh reports, after a successful move, any moved key that
  still owes a public reply bound to main/<key>, since that binding no longer
  names the home owning the work. The move itself is never blocked.

Docs and the secondmate handoff prose follow the same rule. Tests cover the
refusal, its scoping, the unchanged local-link path, and both handoff outcomes
at the script boundary.

* docs(skills): add remote-secondmate recovery hint for false-negative verdicts (kunchenguid#2456)

* fix(skills): hint that remote secondmate liveness verdicts false-negative

fm-crew-state and fm-send routinely misreport a live remote secondmate
as dead; confirm against the pane before relaunching, and relaunch only
through fm-spawn.sh, never raw herdr pane surgery.

* no-mistakes: apply CI fixes

* fix(calm): keep Pi's export confirmation visible (kunchenguid#2461)

Pi 0.83.0 added a status line to every tool-expansion change, and Pi
updates the previous status line in place when two status messages
arrive back to back. Calm's post-export redraw cycled tool expansion on
the macrotask right after Pi printed "Session exported to: <path>", so
both expansion status lines coalesced over that confirmation and the
captain was left with no record of where their export landed.

Calm now repaints only the tool rows it presents, by invalidating each
row through the render context Pi hands its render slots, and requests
the surrounding redraw through setStatus. Neither appends to the
transcript. The repaint is still needed because Pi can re-render a row
asynchronously - the built-in edit row invalidates itself once its diff
is ready - and that re-render can land inside the window where /export
forces stock rendering.

The real-terminal /export case now asserts the confirmation is still on
screen after the redraw has settled, and that the redraw restored every
Calm-hidden row, instead of only racing the moment the confirmation
first appeared.

* fix(procevent): deliver captured results at most once

A Lavish-captured answer reached a secondmate five times because
forwarding it ran fm-send.sh and only afterward, manually, remembered to
call `handled` - so every re-announcement of the still-unacknowledged
capture (a restart, a compaction, a re-read wake) repeated the forward.
The retry loop was ours, not lavish-axi's: fm-procevent.sh already proves
capture is exactly-once, but nothing paired a downstream effect with its
acknowledgement atomically.

Add `fm-procevent.sh deliver <source-id> <sequence> -- <command>...`,
which checks handled status, runs the command, and marks the generation
handled as one call under the source lock, so a repeated delivery attempt
against an already-delivered generation runs the command zero times and a
failed command stays eligible for retry instead of being dropped. Point
the process-event-sources skill and the runner's operating contract at it
for exactly this class of forward.

Five delivery attempts against one captured generation now run the
downstream command once instead of five times - the read-and-discard
cost of the other four is eliminated structurally. The historical
incident's own token cost was not measured at the time and is not
reconstructable after the fact.

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: joliverMI <joliver@sensibletech.biz>
zeeshaanahmad added a commit to zeeshaanahmad/firstmate that referenced this pull request Aug 17, 2026
…12)

* fix(bin): require quota-axi 0.1.25 (kunchenguid#2300)

* fix: raise quota-axi floor to 0.1.25 for Cursor CLI quota awareness

Homes on latest main need quota-axi kunchenguid#87 so Desktop-absent CLI machines report a fresh Cursor quota instead of a false sign-in-required.

* no-mistakes(document): Update quota floor documentation pointer

* fix(bin): prevent false Pi watcher alarms during hand-offs (kunchenguid#2304)

* fix(guard): stop the false send-time watcher-down alarm on Pi primaries

On a Pi primary the watcher process is not the liveness signal. The Pi
extension tears the watcher down on every actionable wake and spawns the
replacement itself, so the singleton lock is legitimately unheld between
cycles: every one of the 799 cycles in a live primary's ledger ends with
lock_after=pid:none, and a live capture caught the guard verdict flipping to
no-watcher during one hand-off with the beacon 63s old.

bin/fm-guard.sh classified Pi as a persistent-watcher harness, which demands a
live identity-matched lock holder at all times, so any guarded command landing
in a hand-off painted the full WATCHER DOWN - SUPERVISION IS OFF banner and
told firstmate to repair a cycle the extension already owns and is restoring.

Add an extension supervision model for pi and pi-signed. A live
identity-matched watcher stays the ordinary healthy state; an unheld lock is
healthy only while the beacon is fresh within grace AND a live Pi session
provably owns continuity - both primary extensions recorded in their state
markers at their current on-disk builds by the process named in state/.lock,
with that process still alive. Without that proof the banner fires exactly as
before, so an unloaded, version-drifted, or exited Pi session is loud
immediately and a cycle the extension never restores is loud once the beacon
passes grace. The queued-wake warning, the PID-strict turn-end guard, and
every other primary's detection are untouched.

Fold session-start's duplicate Pi marker predicate into the shared library so
the ownership contract has one owner.

* no-mistakes(review): Restrict Pi hand-off tolerance to unheld watcher locks

* no-mistakes(document): Document Pi watcher hand-off supervision

* feat: support Cursor Agent CLI as a primary harness (kunchenguid#2305)

* feat(cursor): add Cursor Agent CLI primary hooks, park supervision, and session start

Register a tracked project-scope .cursor/hooks.json for Cursor's stop,
sessionStart, preCompact, and preToolUse steps.

bin/fm-turnend-guard-cursor.sh owns Cursor's turn boundary as a park: it
foregrounds the watcher arm, holds the boundary open until an actionable close,
and returns that wake as one follow-up. Exit 2 is a silent no-op on Cursor's
stop step, so the adapter never uses it. The follow-up loop is bounded twice,
by Cursor's own loop_limit and by the payload's loop_count.

bin/fm-sessionstart-cursor.sh delivers the digest as additional_context at
sessionStart, and stages it for the next turn boundary at preCompact, which
cannot inject context.

Cursor also loads the tracked Claude settings, so bin/fm-hook-host-lib.sh lets
each tracked Claude-shaped entrypoint stand down on a Cursor-delivered payload
rather than running every covered event twice.

bin/fm-tmux-lib.sh reclassifies a Cursor pane's composer cursorlessly, because
Cursor parks its terminal cursor outside the composer, which restores a genuine
composer-empty proof and unblocks away-mode escalation delivery.

* feat(cursor): make Cursor Agent CLI a verified primary harness

Resolve Cursor in the session-lock ancestry through bin/fm-cursor-lib.sh, which
a Cursor primary needs before it can hold its own home lock, and classify its
stop-hook park under the autoarm supervision model so the mid-turn pull guard
stops reporting a healthy between-turns watcher as down.

Read a Cursor pane's composer cursorlessly on tmux, gated on Cursor's own
structural process identity, which restores a genuine composer-empty proof and
lets away-mode escalations reach a Cursor primary with no daemon change.

Lift the secondmate refusals in bin/fm-spawn.sh and bin/fm-control-lib.sh now
that the supervision protocol exists and is recorded.

Cover the whole surface with a portable regression over real processes, an
opt-in live guard against the installed cursor-agent, and dated per-harness
evidence.

* docs(cursor): record Cursor as a verified primary across the owning surfaces

Update the turn-end guard, session-start, arm-seatbelt, cd-guard, watcher
continuity, architecture, configuration, README, and harness-adapters owners,
and add dated live evidence to the supervision and runtime-backend verification
records. Correct the recorded Cursor tmux composer verdict: the cursor-anchored
read is still blind, but the composite reader is no longer unknown.

Lift the remaining remote-secondmate refusal missed in the previous commit, and
add the new libs to the existing fixtures that copy a fixed dependency list.

* refactor(cursor): name the park's stand-down condition for both its causes

Also record that Cursor's preCompact firing itself is not yet live-verified,
while the static evidence that it cannot inject context, and the staging path
that follows from it, both are.

* test: give the pretool fixtures their new dependency and one lint owner

The cd-guard fixture copies a fixed dependency list and now needs the shared
hook-host predicate. Both pretool suites also asserted cleanliness with a bare
shellcheck call, a second and weaker copy of the lint definition that
bin/fm-lint.sh owns: it omits --external-sources, so it failed the moment these
checkers sourced a shared library. They now delegate to that owner.

* test: assert the cursor secondmate contract instead of its removed refusal

A cursor secondmate now launches, so the suite asserts what its park actually
needs: --trust so the home's project hooks load at all, its own home pinned as
the workspace, and the autoarm supervision model inherited across the launch.

* no-mistakes(review): Serialize Cursor wakes and bind staged context

* no-mistakes(review): Serialize Cursor context and nag state commits

* no-mistakes(review): Enforce Cursor ceiling before staged context delivery

* no-mistakes(review): Serialize Cursor claims and staged context

* no-mistakes(review): Serialize Cursor ownership and state commits

* no-mistakes(review): Protect Cursor context across session takeover

* no-mistakes(review): Preserve Cursor context across session takeover

* no-mistakes(review): Enforce owner-keyed Cursor staged context

* no-mistakes(review): Atomically claim Cursor follow-ups and staged context

* no-mistakes(review): Defer Cursor preCompact staging and simplify supersession

* no-mistakes(review): Serialize Cursor park commits and defer preCompact

* no-mistakes(review): Stop Cursor parks after session takeover

* no-mistakes(test): Route Cursor preCompact context through stop follow-up

* no-mistakes(document): Update Cursor primary documentation

* revert(cursor): cut preCompact staging from this change

Carrying a compaction digest across two concurrently running stop hooks kept
producing races that could deliver it twice or strand it indefinitely, and
closing them kept enlarging a critical section inside a hook Cursor awaits at
the turn boundary. Native preCompact firing was never observed either, so the
surface has no empirical basis yet.

Remove the adapter, its registration, its staged path in the park, and its
tests, and record the surface as deferred and uncovered alongside the Codex
interactive TUI. A regression now asserts preCompact stays unregistered so it
cannot return without its own design and evidence.

This change ships the proven core only: the turn-end follow-up park, the
run-tier session start, and away-mode delivery.

* no-mistakes(review): Correct Cursor park supersession documentation

* no-mistakes(document): Clarify Cursor run-tier verification ownership

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

---------

Co-authored-by: kunchenguid <kun-1@kunchenguid.com>

* feat(bin): add decline and repair paths for decision holds (kunchenguid#2330)

* feat(bin): add unrouted close paths to the captain decision gate

A captain who declines a held decision leaves no follow-up work to route,
so `resolve` could not express that answer: it requires at least one
`--routed-to` task. The only way to close such a hold was a direct
`tasks-axi done`, which never writes the durable resolution record the
completion gate reads, so the originating investigation could no longer
pass `verify` and its cleanup stayed blocked.

Add two close paths that route no work:

- `decline` closes an actively held hold with a recorded captain decision
  and no routed task. It refuses while any task is still blocked by the
  hold, because releasing routed work without recording it is `resolve`'s
  job.
- `repair` records the missing resolution block on a hold that was already
  closed outside this script. It never reopens a hold and never clears a
  dependency edge, and it refuses a hold that is still actively held.

Both require a non-empty captain decision file and share `resolve`'s
digest-based retry identity, so an exact retry is idempotent while a
changed decision is rejected. The recorded body now also names which path
closed the hold, and each routed entry regains its own line.

The gate itself is unchanged: an unanswered decision still fails
completion and blocks teardown, and neither new path can close a hold
without the captain's recorded word.

* fix(bin): require captain-hold provenance before repairing a decision

`repair` checked only that the backlog item was kind captain and Done, so
an ordinary captain-kind task that was never held for the captain could be
closed, repaired, and then pass the completion gate.

tasks-axi keeps `hold_kind` through a close, so it is the surviving proof
that an identity really was a captain hold. Require it before writing the
resolution record, and cover the case in the gate regression.

* no-mistakes(document): Correct decision-hold lifecycle documentation

* fix(bin): surface buried wake status lines once (kunchenguid#2331)

* fix(bin): surface buried status notes on wake drain

A note: answer immediately followed by a routine note was dropped because
annotations kept only the newest line and note: never enters OPEN DECISIONS.
Present every unread note and pending-reply resolution since the last drain
cursor, and annotate every unread line on a queued signal.

* no-mistakes(review): Fix unread status cursor races and overflow

* no-mistakes(review): Preserve cursors when status span reads fail

* no-mistakes(review): Make status presentation transactional under I/O failures

* no-mistakes(review): Simplify unread status cursor and presentation locking

* no-mistakes(review): Align cursor failure regressions with transactional presentation

* no-mistakes(review): Retire stale presentation cursors during task teardown

* no-mistakes(review): Preserve routine status until signal annotation

* no-mistakes(review): Correct unread status cap documentation

* no-mistakes(document): Document unread wake status presentation

* no-mistakes(lint): Fix wake surfacing ShellCheck warnings

* no-mistakes: apply CI fixes

* feat: add max Calm presentation level (kunchenguid#2334)

* feat(calm): add a max presentation level that hides mid-turn working notes

Calm's home-local preference becomes a three-state level instead of a
boolean: "off" is stock Pi, "on" is today's Calm, and "max" is Calm plus
hiding the assistant text of messages the model did not end its response
with. `/calm max` selects it from any state, a plain `/calm` steps max
back to ordinary Calm and otherwise keeps the existing on/off cycle, and
any other argument keeps that cycle too.

`config/calm` now persists "max" as its own literal value, so a session
start, resume, fork, or reload restores the stored level rather than
treating it as unrecognized and dropping to off.

The hide rule keys on Pi's intrinsic per-message stopReason: "toolUse",
or "length" with tool calls present. Streaming ("pending") text is never
filtered, because suppressing it would also stop a genuine reply from
streaming. The existing assistant layout adapter filters the blocks out
of the same shallow presentation copy it already uses for collapsed
thinking, so the message, model context, session storage, /export, and
delivery are untouched and a hidden mid-turn row collapses to zero
height. The new "assistant-working-note" class keeps that choice in the
visibility policy owner, where ordinary Calm keeps it visible.

* no-mistakes(document): Clarify Calm max persistence and taxonomy

* feat(calm): hide mid-turn working notes by default (kunchenguid#2339)

* feat(calm): make hiding mid-turn working notes the ordinary Calm state

Calm collapses back to the two-state on/off toggle it was before the max
presentation level, with max's hide rule promoted into ordinary Calm.
Calm on now hides mid-turn assistant working notes in addition to what it
already hid, and the /calm command parses no argument again.

The hide rule itself is unchanged: assistant text is removed from the
shallow presentation copy when the message's own stopReason is "toolUse",
or "length" with tool calls present. Streaming ("pending") text is never
filtered, so a genuine reply still streams. The message, model context,
session storage, /export, and delivery remain untouched.

config/calm persists only "on" and "off" again, but the reader still maps
a persisted "max" to on so a home upgraded from the removed level keeps
Calm on instead of dropping to off.

The mid-turn hide is now default behavior rather than an opt-in level, so
docs/calm.md documents it for users, docs/configuration.md records the
two written values plus the legacy max mapping, and the feasibility
taxonomy drops its level-scoped wording.

* no-mistakes(document): Document ordinary Calm working-note hiding

* chore: store no-mistakes test evidence in the repo (kunchenguid#2355)

* chore: ignore scratchpad/ at the repo root (kunchenguid#2359)

* fix(ci): fail hung Herdr behavior runs in 20 minutes (kunchenguid#2413)

A wedged family-run step was occupying the runner until the 75-minute
job cap; bound that step so cleanup and timing artifacts still upload.

* fix: keep the public promise reachable when work is routed to a second mate (kunchenguid#2457)

The lightweight Relay follow-up link lives in the answering home's own
state/<task-id>.meta, so it can only bind work that home owns. When a
Relay-linked request is routed to a second mate, the task record lives in the
second mate's home, fm-x-link.sh failed with a bare "no such task ...meta", and
nothing else picked the promise up: only the soft acknowledgement was ever
posted. The typed promised-final path already supports --work-home
secondmate:<id>; the playbook simply never chose it.

- fmx-respond now states the routing rule crisply: a task in this home takes the
  lightweight link, and second-mate-routed work takes a promised-final
  commitment bound to that home, registered up front with the brief command
  carried into the routed worker's instructions.
- fm-x-link.sh refuses a task with no local record by naming the registered
  second mate whose home actually holds it and printing the promised-final
  registration command, with the exact --work-home when the match is
  unambiguous. A home with no registered second mates keeps the plain error.
- fm-backlog-handoff.sh reports, after a successful move, any moved key that
  still owes a public reply bound to main/<key>, since that binding no longer
  names the home owning the work. The move itself is never blocked.

Docs and the secondmate handoff prose follow the same rule. Tests cover the
refusal, its scoping, the unchanged local-link path, and both handoff outcomes
at the script boundary.

* docs(skills): add remote-secondmate recovery hint for false-negative verdicts (kunchenguid#2456)

* fix(skills): hint that remote secondmate liveness verdicts false-negative

fm-crew-state and fm-send routinely misreport a live remote secondmate
as dead; confirm against the pane before relaunching, and relaunch only
through fm-spawn.sh, never raw herdr pane surgery.

* no-mistakes: apply CI fixes

* fix(calm): keep Pi's export confirmation visible (kunchenguid#2461)

Pi 0.83.0 added a status line to every tool-expansion change, and Pi
updates the previous status line in place when two status messages
arrive back to back. Calm's post-export redraw cycled tool expansion on
the macrotask right after Pi printed "Session exported to: <path>", so
both expansion status lines coalesced over that confirmation and the
captain was left with no record of where their export landed.

Calm now repaints only the tool rows it presents, by invalidating each
row through the render context Pi hands its render slots, and requests
the surrounding redraw through setStatus. Neither appends to the
transcript. The repaint is still needed because Pi can re-render a row
asynchronously - the built-in edit row invalidates itself once its diff
is ready - and that re-render can land inside the window where /export
forces stock rendering.

The real-terminal /export case now asserts the confirmation is still on
screen after the redraw has settled, and that the redraw restored every
Calm-hidden row, instead of only racing the moment the confirmation
first appeared.

* Apply upstream-preference policy to the decision-hold resolution

Captain policy (2026-08-17): where upstream fixed something we also fixed, take
upstream's version; keep only what is exclusively ours, and only after proving
upstream's version does not already cover it.

Re-resolved the three decision-hold surfaces so our divergence from upstream is
exactly our exclusive precondition-recheck work, with no upstream text overridden.

bin/fm-decision-hold.sh - our whole change (base..ba6f281) is purely additive: two
header lines, print_precondition_reminder(), and two call sites. It modifies no line
upstream touched, so there is nothing to hand over. Restored our two header lines
verbatim instead of the reworded version, dropping the decline/repair clause I had
added, so the delta is exactly the original exclusive lines.

.agents/skills/decision-hold-lifecycle/SKILL.md - base had steps 6-8. We never
changed base steps 6 and 7; upstream did, so upstream's wording wins outright.
The one genuinely contested sentence is base step 8, which both sides edited.
Upstream's version of it is now kept verbatim as step 12. Our precondition
verification, which upstream's sentence does not cover, is retained as its own
narrow step 11 rather than as a hybrid sentence.

docs/decision-hold-lifecycle.md - restored our single advisory-reminder line
verbatim, dropping the routed-path clause I had added.

Equivalence evidence for what was dropped:

- open-decisions-cursor retirement in bin/fm-teardown.sh is a same-problem overlap.
  Upstream's status_retire_presentation_task runs
  `rm -f -- "$state/$task.status" "$state/.$task.open-decisions-cursor"` under the
  presentation lock and also retires the manifest row, so it strictly covers the
  inline rm ours used. Ours dropped, upstream's kept.

- liveness.sh / liveness-trust retirement is exclusively ours: upstream has zero
  liveness references in both bin/fm-teardown.sh and bin/fm-classify-lib.sh, so
  nothing covers it. Retained.

- precondition recheck is exclusively ours: upstream/main has zero
  precondition/recheck/reminder references in bin/fm-decision-hold.sh, the skill,
  and the regression, against 5/3/4 on our side. Retained.

- bin/fm-test-run.sh is not a competing fix: each side added a different test file
  to the same shard list. Union kept.

* test(watch-arm): gate re-arm exit on the condition, not a fixed sleep

tests/fm-watch-arm.test.sh asserted that a re-arm carrying durable wakes exits on
its own by sleeping a fixed 0.25s and then failing if the arm was still live. That
window elapses under load before a healthy arm can start, surface the wakes and
close, so a correct run fails.

Measured on this host before the change, same failure message both sides:
merged 9/25 runs failed, pre-merge baseline 8/25 failed. The test file is
byte-identical at the merge base, our tip, upstream/main and the merge commit, and
the fixed sleep is present in all four, so this is inherited test debt rather than
anything either side introduced.

Gate on the arm actually exiting instead, using the file's existing condition-based
wait_for_exit helper and its established 80-iteration bound. wait_for_exit returns
the arm's own exit status, so the separate wait/status read is no longer needed and
the 124 timeout result drives the failure path. This is the same correction
83eb27a made to the watch-triage wait gates for the same reason.

After the change: 33 consecutive runs with no occurrence of this failure, at system
load averages between 9 and 13, the range that reproduced it reliably before.

A second, much rarer failure in this file, "marker publication failure discarded
stale-lock recovery evidence", is NOT addressed here. It appeared roughly twice in
about seventy runs across every configuration, including zero times in twenty
pre-merge baseline runs and zero times in twenty-six unfixed merged runs, so its
mechanism is not established. It is reported rather than gated, because adding a
readiness gate there could hide a real watcher-cleanup defect.

* no-mistakes(document): docs(decision-hold,scripts): fix stale audience-check counts, document 3 new Cursor scripts

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: kunchenguid <kun-1@kunchenguid.com>
huynhtandat223 added a commit to huynhtandat223/firstmate that referenced this pull request Aug 18, 2026
* feat(bin): add deterministic agent lifecycle control (#1568)

* feat(bin): add deterministic agent lifecycle control

Separate firstmate's data plane from its control plane.

bin/fm-send.sh is the data plane: conversational text, always
routing-marked for a kind=secondmate target. That marking is right for a
message and wrong for a lifecycle command - a marked "/quit" arrives as
ordinary chat the agent reasons about instead of executing.

bin/fm-control.sh is the control plane: allowlisted interrupt, exit, and
transactional relaunch verbs addressed to an exact task id, with
per-harness mechanics owned by the executable bin/fm-control-lib.sh
rather than improvised in agent prose, and a verified postcondition for
every action. There is no arbitrary-text and no raw-key entry point.

relaunch runs as a transaction with a durable journal: it resolves the
profile, proves the work it must preserve is recoverable, records the
required progress note, stops the old agent, then delegates the launch
to its single owner, bin/fm-spawn.sh --relaunch, which adopts the
recorded endpoint and worktree instead of creating either. A refusal
before the stop leaves the record and instructions byte-identical; a
failure after it reports the concrete state rather than claiming an
agent that is not running. Teardown and discard stay separate and
explicit.

exit and relaunch require a backend with a recovery-grade agent-state
classifier, so zellij, orca, and cmux are refused rather than reported
as successful blind. A remotely placed secondmate is refused by name,
because its agent runs on a host where none of these postconditions can
be read.

* fix(control): resolve a recorded harness to its adapter before retiring wiring

fm-spawn arms per-task harness wiring on prefixes, because a task
launched from a raw command records that command's basename rather than
the exact adapter name. The control plane's retirement tables are keyed
by the exact adapter, so a task recorded as `grok-2` had its turn-end
token, private registry entry, and worktree hook pointer armed and never
retired - leaving a registry entry that outlived the agent that owned
it.

State the prefix rule once, in the capability owner, and resolve the
recorded value through it before every table lookup. bin/fm-send.sh's
composer-clear lookup reads the same owner instead of keeping its own
copy of which adapters need one.

* test(control): pin muse session-binding retirement across a harness switch

* no-mistakes(review): Resolve prefixed harnesses across lifecycle control verbs

* no-mistakes(review): Report interrupt delivery without fabricating cancellation state

* no-mistakes(review): Clear disabled relaunch trace context atomically

* no-mistakes(review): Clarify control interrupts and restore legacy send state

* no-mistakes(review): Refuse ambiguous relaunches and report exit delivery

* no-mistakes(review): Revalidate interrupts and accept interrupt-stopped exits

* no-mistakes(review): Lock descendant tasks before forced recursive teardown

* no-mistakes(document): Align lifecycle adapter documentation with control plane

* no-mistakes: apply CI fixes

* fix(bin): serialize fresh task publication with forced teardown

Forced secondmate teardown enumerated a home's task set, locked what it
found, then re-enumerated while removing. A fresh spawn takes only its
own per-task lock, so a record published inside that window was
invisible to the preflight and visible to the cleanup: it was
destructively processed while never lifecycle-locked.

Reproduced with real agents. A record published 0.249s after teardown
began was removed, its window closed, and its worktree returned to the
pool - while both commands reported success. A per-task lock cannot
protect a task that does not exist yet.

Add a per-home task-set lock guarding WHICH tasks a home has, as opposed
to the metadata lock guarding one task's record. Teardown takes it per
home, parent before child, before enumerating and holds it through
cleanup. A fresh spawn takes it before its own per-task locks and holds
it through publication; a relaunch is exempt, because it republishes an
existing task already covered by that task's control lock.

Either the spawn publishes first and the teardown's preflight covers it,
or the teardown owns the set and the spawn refuses. Both directions fail
closed, and both are pinned by tests that hold the lock rather than
racing on timing.

* no-mistakes(review): Serialize remote secondmate publication with forced teardown

* no-mistakes(review): Preserve remote spawn routing and state initialization

* no-mistakes(review): Serialize teardown when descendant state is absent

* no-mistakes(review): Cover symlinked descendant state refusal

* no-mistakes(document): Document task-set serialization safeguards

* no-mistakes(lint): Isolate task-set lock path resolution

* no-mistakes: apply CI fixes

* feat(stow): add tiered decaying memory management (#1984)

* feat(stow): tiered decaying memory with captain-gated offload to local excluded skills

Implement the captain-adopted /stow redesign from the v2 tiering report as
amended by the adoption decision:

- Per-entry trailing HTML-comment markers with three tiers named for their
  handling: pinned (no clock, no eviction), aging (stale after 30 days),
  perishable (stale after 7 days, mandatory checkable expiry condition).
- File-scoped defaults (captain.md and captain-shared.md pinned,
  learnings.md aging) with a self-describing legend line per file header.
- Reinforcement requires session evidence; re-reading memory never counts.
- Archive-not-delete: stale and budget-evicted entries move with provenance
  to the never-injected data/memory-archive.md; prune always means the cold
  tier, and a stale unique fact is never deleted.
- Captain-gated over-budget offload: staleness evaluated before scope, the
  sweep runs only when still over budget after decay and consolidation,
  proposals go through the receipt plus one durable captain-held backlog
  item, migration runs through the destination's normal path, and the
  memory entry leaves only once the destination is live.
- Offload destination per the adoption decision: a user-owned skill under
  .agents/skills/<freeform-name>/ excluded via the local .git/info/exclude,
  with the hard rule that stow never creates or writes a tracked skill.
- Five graduation moves, receipt verbs archived and proposed-offload, and
  the one-time non-destructive migration of unmarked legacy entries.

The public skills/stow/SKILL.md mirrors the generic parts (markers, decay,
archive exit, user-approved on-demand offload exit, migration) with no
firstmate-specific paths.

The load-bearing assumption that a git-excluded skill is still discovered
was verified empirically against Claude Code 2.1.226 (direct
.git/info/exclude scratch-repo test plus an in-repo ignored-probe test);
the dated evidence is recorded in docs/verification/stow-memory.md.

The graduation list's deletion move is deliberately narrowed to duplicates
already preserved by a stronger owner, reconciling the v2 report's retained
'deletion of a stale entry' wording with its own prune-always-archives
rule.

* no-mistakes(review): Persist legacy migration grace across stow passes

* no-mistakes(review): Enforce archival invariants and exempt default-pinned legacy entries

* no-mistakes(review): Clarify offload scope, archive placement, and marker boundaries

* no-mistakes(review): Enforce aging fallback and verify excluded skill loading

* no-mistakes(review): Fix stow decay, pinned offload, and archival safeguards

* no-mistakes(review): Preserve pinned entries, approvals, and archive provenance

* no-mistakes(review): Restrict stow mutations to editable memory files

* no-mistakes(review): Clarify skill destinations, collision checks, and migration legends

* no-mistakes(review): Resolve exclude paths for linked worktrees

* no-mistakes(review): Secure per-home excluded skill migration

* no-mistakes(test): Require explicit tier markers on new stow entries

* no-mistakes(test): Route missing shared legends to primary owner

* no-mistakes(document): Align stow documentation with tiered memory

* fix(stow): converge the pass on an over-budget home (dogfood D1-D3)

The dogfood run against a copy of the real over-budget home showed the
pass increasing the deficit from 624 to 1,107 estimated tokens and the
relief ladder provably unable to reach budget. Three skill-text fixes:

- D1: markers become single-token spellings (<!--a:DATE-->, <!--p:DATE-->,
  <!--P-->, <!--g-->), entries matching a pinned file default carry no
  marker, the per-file policy legend collapses to a one-line pointer
  naming the stow skill as the scheme owner, and marker/pointer bytes are
  explicitly counted content - roughly 76% less metadata cost on the
  dogfooded home's first installment.
- D2: the eviction rung gains a convergence precondition - total the
  eligible pool first, and when archiving all of it cannot reach budget,
  skip eviction entirely, archive nothing for budget reasons, and report
  the exempt pinned floor as the concrete inability in the final step.
- D3: budget eviction considers only dated aging entries; <!--g-->
  legacy-grace entries are ineligible until their grace cycle resolves,
  so eviction cannot cancel promised grace or invert against validation.

Public skill mirrors the D1 marker/pointer changes; D2/D3 are internal
because the public skill has no budget ladder.

* no-mistakes(test): Enforce evidence-only reinforcement during stow migration

* no-mistakes(document): Clarify stow receipt marker actions

* docs: add project vision (#1997)

* docs: add firstmate vision

* no-mistakes(test): Classify VISION.md as public product documentation

* no-mistakes(document): Restore approved one-file vision diff

* no-mistakes: apply CI fixes

* fix(spawn): force regular Pi TUI for crews (#2005)

* fix(spawn): force regular Pi TUI for crews

* no-mistakes(document): Documented Pi regular TUI launch mode

* fix(cmux): classify borderless Claude composers (#2029)

* fix(cmux): classify borderless Claude composer

* no-mistakes(review): Normalize cmux NBSP prompts across locales

* no-mistakes(document): Document cmux borderless Claude composer classification

* docs(stow): generalize read-before-write in the public stow skill (#2091)

The public installer-facing stow skill scoped its classify-then-replace
discipline to TODO/BACKLOG items only, so findings routed to a memory file
had no stated rule against a blind append or a wholesale overwrite.

Step 6 now classifies every finding against the destination's current
contents as new, duplicate, superseding, or obsolete, and states the
considered replacement each classification implies. The outcomes follow the
tiered-memory contract already in the file: an obsolete entry is refreshed,
archived, or replaced in a way that preserves its fact, a duplicate folds
into the entry that already carries it, and a superseded body worth keeping
leaves through step 7's existing exits rather than a second recovery
mechanism.

* fix: resurface durable supervision work after re-arm (#2065)

* fix(watcher): resurface durable work after downtime

* no-mistakes(review): Make watcher rearm recovery durable and cursor-safe

* no-mistakes(review): Persist safe recovery markers across migration lock recovery

* no-mistakes(review): Retain stale lock when recovery marker publication fails

* no-mistakes(review): Preserve delivery-gap recovery and quarantine malformed markers

* no-mistakes(review): Serialize recovery consumption and report acknowledgment failures

* no-mistakes(review): Centralize recovery publication before clearing watcher evidence

* no-mistakes(review): Guarantee recovery evidence across queue and lock handoffs

* no-mistakes(review): Publish recovery evidence before durable wake commits

* no-mistakes(review): Replace recovery marker Perl dependency with Node

* no-mistakes(review): Keep interrupted wakes durable until handling acknowledgment

* no-mistakes(review): Add post-handling durable wake acknowledgements

* no-mistakes(review): Enforce post-handling acknowledgement across recovery and AFK return

* no-mistakes(review): Bind wake acknowledgements to recovery generations

* no-mistakes(review): Align wake regressions with generation-bound acknowledgements

* no-mistakes(document): Document durable re-arm recovery semantics

* no-mistakes(lint): Resolve ShellCheck warnings in recovery and watcher tests

* no-mistakes: apply CI fixes

* test(watcher): assert post-handling wake replay

* no-mistakes(review): Prevent successor loops and adopt legacy wake generations

* no-mistakes(review): Rearm durable wakes without recursive successor recovery

* no-mistakes(review): Align recovery tests with handling marker state

* no-mistakes(review): Delay handling transition until successor launch is established

* no-mistakes(review): Confirm wake handling only after successful prompt delivery

* no-mistakes(review): Acknowledge AFK wakes only after evidence publication

* no-mistakes(review): Prevent AFK wake loss before post-handling acknowledgement

* no-mistakes(document): Document durable wake acknowledgement semantics

* no-mistakes(lint): Suppress false positive for recovery action output

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* ci: measure Herdr automation on Windows runners (#2100)

* ci: add Windows Herdr automation spike

* ci: run Windows spike on its pull request

* fix: wait for Windows Herdr command output

* fix: run ANSI probe in pane shell

* ci: keep Windows Herdr spike manually triggered

* docs: clarify Windows Herdr spike verdict

* feat(ahoy): guide captains through open decisions (#2099)

* Add guided ahoy decision flow

* no-mistakes(document): Document guided Ahoy decision flow

* fix(stow): enforce startup-memory budget decisions (#2110)

* Harden stow memory budget policy

* Refine internal stow offload policy

* no-mistakes(review): Enforce shared-budget decisions and autonomous offload

* fix(spawn): refresh pooled worktrees from origin before launch (#2116)

* fix(spawn): refresh pooled worktree base

* no-mistakes(document): Document spawn base-freshness invariant

* no-mistakes: apply CI fixes

* fix(composer): unify safe classification across backends (#2102)

* refactor(composer): one shape owner behind thin capture adapters, whole matrix fixed

Consolidate every composer shape - bordered boxes (all families, geometry,
titled bottom borders), bare agent-glyph rows and their wrap regions,
opencode's left bar, and pi's identity-gated separator pair - into
fm_composer_classify_screen in bin/fm-composer-lib.sh. Adapters now
contribute only a capture and a declarative capability descriptor
(styled/cursor/identity/rows); capability differences change how confidently
a shape is judged, never what the shapes are, so a new harness shape is
teachable in exactly one place.

Correctness fixes landed as part of the consolidation (audit
data/fm-composer-consolidation-audit-s1):
- locale-safe Unicode-space normalization in the shared owner (closes the
  fleet-wide half of #1988; cmux's local byte-exact NBSP case deleted;
  naming converges with PR #1995's normalization primitive)
- muse's bare glyph joins the shared set, unbreaking muse on herdr/cmux/orca
- orca learns the borderless bare shape, drops its backward-paged composer
  window, and can no longer classify a stale startup banner as the composer
- tmux tolerates a titled bottom border, unbreaking grok steering
- the left-bar shape makes opencode readable on every backend
- zellij gets a real classifier through dump-screen --ansi, replacing the
  content-diff submit heuristic that could confirm an undelivered message
  and close a --resolve-key decision (the fleet's only false positive)
- fm-spawn's kimi launch-readiness regex (the fourth shape copy) now routes
  through the shared classifier

The strict blank-row posture applies fleet-wide (captain decision
blank-row-injection-posture): no positive container proof = unknown = defer,
replacing tmux's permissive blank-cursor-row rule. Away-mode injection was
re-validated end to end on real tmux (defer on partial input and unproven
rows, clean delivery with swallowed-Enter retry into proven-empty
composers). The tmux submit core gains a baseline-idle turn-started
conversion so pi steering stays confirmed while its working screen hides
the composer; busy conversion without that baseline remains forbidden.

Plain-capture backends now degrade a glyph row carrying trailing text to
unknown instead of a false pending, per the approved capability rule.

Portable regressions pin the full byte-capture matrix from the audit under
a UTF-8 locale and LC_ALL=C, the strict-vs-permissive divergence, and
deliberate signal separation; the opt-in live guard
(tests/fm-composer-matrix-live-e2e.test.sh) verified every installed
harness against the real classifier, recorded in
docs/verification/runtime-backends.md.

* no-mistakes(review): Fix Pi glyph ambiguity and complete profile matrix

* no-mistakes(review): Preserve bare verdict when Pi identity probe is absent

* no-mistakes(review): Harden composer structure and titled-border geometry

* no-mistakes(review): Require proven idle baseline and strict Zellij guard

* no-mistakes(review): Reject box bottom borders as composer input rows

* no-mistakes(review): Prove Zellij probe typing before classifier retries

* no-mistakes(review): Preserve Pi identity uncertainty and scan full left-bar drafts

* no-mistakes(review): Verify Zellij text lands before submitting

* no-mistakes(review): Scope Zellij typing verification to selected composer content

* no-mistakes(review): Verify Zellij pastes through composer-scoped content deltas

* no-mistakes(review): Prove wrapped bare Zellij pastes through composer extraction

* no-mistakes(review): Invalidate stale cursorless composers below dead shell prompts

* no-mistakes(review): Handle shell prompt placeholders in composer extraction

* no-mistakes(review): Classify cursorless bare continuation regions safely

* no-mistakes(review): Reject stale cursorless containers below live activity

* no-mistakes(review): Preserve prompt glyphs in wrapped Zellij pastes

* no-mistakes(review): Reject live shell rows during composer extraction

* no-mistakes(review): Preserve wrapped glyph continuations through submit retries

* no-mistakes(review): Scope idle placeholders to proven positions

* no-mistakes(review): Restore boxed placeholders and live prompt reanchoring

* no-mistakes(review): Fix Zellij placeholder and wrapped glyph paste proof

* no-mistakes(document): Align composer architecture documentation

* no-mistakes(lint): Fix ShellCheck warnings in composer refactor

* no-mistakes: apply CI fixes

* docs(verification): record the trusted-checkout live matrix rerun

The pipeline's isolated gate worktree is untrusted, so claude, grok, and
muse stopped at first-launch trust dialogs there (the guard refuses to
confirm them by design). This rerun from the trusted checkout at the final
validated head verified all six installed harnesses, the strict blank-row
deferral, and the hardened zellij false-positive probe live.

* no-mistakes(document): Align composer verification evidence

* no-mistakes: apply CI fixes

* no-mistakes(review): Restore proven box bottom-cursor classification

* no-mistakes(review): Preserve styled placeholder-like drafts as pending

* no-mistakes(document): Align composer safety and Zellij delivery documentation

* no-mistakes: apply CI fixes

* docs(verification): refresh the live matrix with the final-head trusted rerun

The post-validation rerun from the trusted checkout verified all six
installed harnesses at the branch's final head, including Claude 2.1.227
(auto-updated since the audit's captures) and Grok, which the untrusted
gate worktree could not verify past their first-launch trust dialogs.

* fix(spawn): gate Pi TUI mode by CLI capability (#2117)

* fix(spawn): gate Pi regular TUI flag by capability

* no-mistakes(review): Document conditional Pi TUI capability detection

* no-mistakes(review): Pin Pi probing and launch to one executable

* no-mistakes(review): Preserve literal pinned Pi paths and update documentation

* no-mistakes(review): Defer pinned Pi path insertion until final substitution

* no-mistakes(document): Document version-safe Pi launch probing

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* docs(vision): elevate experience, pain narrative, and distro virtues (#2147)

* docs(vision): elevate experience, pain narrative, and distro virtues

Fold the captain's public vision framing into VISION.md: peace of mind as a
primary goal, multi-session context-switch pain as the problem one interface
solves, clone-and-run setup ease, self-evolution including community, and
explicit harness/backend orthogonality. Reconcile experience-as-garnish into
experience-as-purpose and update aligns/resists accordingly.

* docs(vision): state the experience goal positively

Drop the negative "not a smart workflow / useful tool / impressive technology"
pretext. Lead straight into the positive experience north star.

* feat(bin): reconcile inactive terminal crew outcomes (#2167)

* fix: reconcile inactive terminal outcomes

* fix: stream secondmate summary inputs

* no-mistakes(review): Fix reconciliation locking and request delivery retries

* no-mistakes(review): Prevent retries after unknown request delivery

* no-mistakes(document): Clarify inactive reconciliation cadence and receipts

* no-mistakes(lint): Quote terminal status arguments in reconciliation tests

* refactor: simplify inactive outcome reconciliation

* no-mistakes(review): Bound inactive reconciliation scans with durable progress

* no-mistakes(review): Bound reconciliation and deduplicate recovery notices

* no-mistakes(document): Document inactive outcome reconciliation contracts

* no-mistakes(review): Reject relative local secondmate parent routes

* no-mistakes(review): Key terminal receipts by spawn incarnation

* no-mistakes(review): Stabilize legacy receipts and lock reconciliation snapshots

* no-mistakes(review): Fail closed on invalid secondmate identity markers

* no-mistakes(document): Document durable inactive-outcome reconciliation

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* ci: raise Herdr test timeout (#2191)

* fix: refresh stale Pi instructions after compaction (#2163)

* fix(session-start): refresh drifted instructions on stale rebuilds

* test(session-start): prove Pi instruction refresh end to end

* no-mistakes(review): Fix stale instruction refresh and baseline integrity

* no-mistakes(review): Preserve true-start baselines across Pi continuations

* no-mistakes(review): Correct Pi continuation classification and live expectation

* no-mistakes(review): Correct Pi continuation coverage documentation

* no-mistakes(review): Fix read-only refresh and exact Pi session restores

* no-mistakes(review): Classify Pi create-if-missing sessions correctly

* no-mistakes(review): Classify named Pi sessions using immutable headers

* no-mistakes(review): Correct Codex interactive coverage diagnostic

* no-mistakes(document): Document immutable Pi compaction instruction refresh

* no-mistakes(document): Correct Pi refresh documentation and validation claims

* feat: add deterministic condition-to-action watcher (#2200)

* feat(bin): add deterministic condition->action watch adapter on the process-event channel

Register a (condition, action) pair once with bin/fm-procevent-when.sh and the
existing process-to-event runner polls the condition tokenlessly, fires the
action at most once on a stable true, and wakes firstmate exactly once with the
captured outcome - instead of burning an agent turn per re-check.

The pair is stored privately under state/when/ and hash-bound by a trust record
the same way fm-check-register.sh binds a custom check, so a mutated spec is
refused without executing anything. A durable exclusive fired marker claimed
before the action makes restarts and re-polls unable to double-fire; every
failure path (mutated spec, condition error past budget, expired deadline,
failed action, uncaptured earlier fire) ends in a terminal captured outcome
that wakes firstmate rather than a silent retry. Eligibility stays a firstmate
judgment: only exact, safe, reversible actions may be bound, and judgment-
needing or destructive actions keep the wake-and-decide flow.

* no-mistakes(review): Harden when watcher concurrency, deadlines, timeouts, and output

* no-mistakes(test): Bind watcher actions to registered executable bytes

* no-mistakes(document): Correct condition-action watcher documentation

* no-mistakes(document): Clarify outcome wake re-announcement

* no-mistakes: apply CI fixes

* fix(bin): honor a decision key stated after the verb colon (#2202)

The open-decisions fold only recognized a [key=<slug>] token between the
verb and the colon (needs-decision [key=x]: note). The common worker
shape with the colon first (needs-decision: [key=x] note) silently
folded its stated key into the shared "default" bucket, so two open
decisions could collapse into one record and fm-send --resolve-key <x>
refused to close the decision it plainly named.

A complete token at the head of the note is now an equivalent stated-key
position for every keyed verb, shared by the whole-file and incremental
folds through the one _fm_decision_key owner. The documented
before-colon position wins when both are present, a token deeper in the
note stays prose, a bare keyless line still folds to "default", and a
stated-but-malformed slug is rejected rather than rewritten to
"default". A consumed note-head token is stripped from the note so both
positions yield identical records, and the incremental fold version is
bumped so persisted cursors folded under the old interpretation are
rebuilt from the authoritative log.

Fixes #2109

* fix(bin): prevent watcher recovery acknowledgement livelock (#2212)

* fix(bin): keep a recovery acknowledgement valid across republication

A watcher cycle that opened and closed while the model handled its drained
wakes minted a fresh recovery generation, which invalidated the exact
acknowledgement the drain had just printed. That acknowledgement then consumed
nothing, so the marker stayed pending and every later arm spent its whole cycle
re-announcing the same recovery instead of supervising - a livelock the home
could not leave on its own.

A downtime publication now reuses the generation of an outstanding handling
episode, so a close during the handling window cannot orphan the printed
acknowledgement. The acknowledgement itself separates its two facts: queue-row
consumption is bound to the monotonic --ack-through sequence and always
happens, while only retiring the episode is bound to --recovery-generation. A
generation that moved on is a non-fatal result that names its own remedy
instead of a refusal that consumes nothing.

* no-mistakes(review): Preserve recovery generations and consume stale acknowledgements safely

* no-mistakes(document): Document sequence-bound recovery acknowledgements

* feat(fmx-respond): consume Relay conversation chains (#2206)

* feat(fmx-respond): consume in_reply_to_chain conversation context

The relay's poll payload can carry in_reply_to_chain, an oldest-first
transcript of the surrounding conversation, but the mention-handling
procedure only ever read the immediate in_reply_to parent, so referents
like "this" in a standalone mention stayed unresolvable even when
context was delivered.

Teach fmx-respond to read the chain when present (optional and
backward-compatible: often absent today, kind label not required),
resolve referents against the whole transcript, and extend the
untrusted-content framing to every chain entry including the upcoming
kind=history entries. Document the field's wire shape in
docs/configuration.md as the firstmate-side owner.

* no-mistakes(document): Document Relay chain context ownership

* fix: parse decision verbs before status metadata tags (#2280)

* fix(bin): strip every bracket tag, not just [key=...], from a status verb

status_line_verb only stripped a leading "[key=...]" token before the
colon, so a remote secondmate reply's leading "[corr=...]" correlation
tag stayed glued onto the returned verb word ("needs-decision
[corr=...]" instead of "needs-decision"). The open-decisions fold's
verb match then silently failed to recognize the line at all, so
fm-send --resolve-key refused to close a decision that was plainly
open on the status line.

Generalize the parser to strip every "[name=value]" tag before the
colon, in any order and count, so local and remote replies fold
identically.

* no-mistakes(review): Invalidate stale decision cursors after parser fix

* no-mistakes(document): Clarify status metadata verb parsing

* fix(bin): collapse duplicate supervision wakes (#2287)

* fix: collapse duplicate supervision wakes without losing legitimate updates

One remote-secondmate note produced two handling turns (a procevent check
wake published before autohandle, then a signal wake for the same mirrored
bytes), already-ingested replays such as a cursor-loss whole-log recapture
still woke with nothing to do, this home's own bookkeeping closes (fm-send
--resolve-key, the pending-reply escalation close, the captain-held
transfer) re-woke the session that wrote them, and turn-ended-only wakes
were annotated with already-announced status lines that looked like fresh
progress.

Dedup rules, each at its layer's one owner:
- fm-procevent.sh: an adapter may declare 'self-announcing'; the runner
  then applies first and publishes a check wake only for what remains
  unhandled. fm-procevent-remote-reply.sh declares it: the mirrored status
  append is the single announcement, so a fully applied capture publishes
  nothing and a byte-identical replay stays completely quiet. All other
  adapters keep strict publish-before-apply.
- fm-wake-lib.sh: fm_wake_signal_sig/seen_path/seen_current now own the
  watcher's signal signature and .seen-* marker format, plus
  fm_wake_status_append_self_announced, the guarded bookkeeping append
  that advances the marker only over exactly its own bytes and fails
  toward waking on any pending or interleaved foreign write.
- fm-send.sh, fm-pending-reply-lib.sh, fm-decision-hold.sh: bookkeeping
  closes go through that guarded append; escalation opens stay plain
  appends because a new blocker must wake.
- fm-wake-lib.sh annotations: a historical (turn-ended-only) row skips its
  status annotation only when the file's signature provably matches the
  seen marker; anything unannounced keeps annotating.
- fm-classify-lib.sh: a kind=secondmate task's status signal is never
  absorbed as provably-working, because that stream is the routed-reply
  channel the parent must read.

Also fixes a pre-existing exit-path deadlock the regression run reproduced:
a TERM inside a recovery-marker critical section left fm_lock_try_acquire
spinning against this same process's abandoned hold; a self-held lock is
now reclaimed (a subshell still waits on its parent's live hold).

Regression tests drive the real wake functions and executables in both
directions: each duplicate case collapses, while a new remote reply, new
decision, new blocker, merge result, failure, first status change, and a
later different note on the same task all still wake.

* no-mistakes(document): Document wake deduplication contracts

* feat: add Cursor CLI crew harness (#2238)

* feat(harness): add Cursor Agent CLI adapter

# Conflicts:
#	bin/fm-spawn.sh

* fix(composer): read cursor-agent's reverse-video placeholder as idle

cursor-agent renders its idle composer placeholder dim (SGR 2) but paints the
cell under the terminal cursor in reverse video (SGR 0;7). Reverse video is
neither dim nor a dark truecolor foreground, so the shared ghost stripper keeps
that one character and an idle composer reduces to a lone `P`. Judged on its
own, that remnant reads `pending` on a genuinely idle pane, which defers
away-mode escalation indefinitely on the styled cursorless backends.

Teach the ONE fleet-wide classifier the shape instead of adding an adapter-local
copy: register `→` as an agent prompt glyph so the composer row is structurally
findable at all (without it the bottom-most shape is a stale shell prompt echo
in the scrollback), add both verified placeholders to the idle set, and consult
the styling-independent plain row when the styled row is only a remnant.

The plain-row branch demands the remnant be a proper, strictly shorter substring
of a plain row matching a fully anchored placeholder. Real typed text is
uniformly bright, so stripping leaves it equal to the plain row and it stays
`pending` - verified live against a pane where the typed text was exactly the
placeholder string.

Verified live on cursor-agent 2026.08.11-e8db854; the regression pins the real
captured bytes and asserts the remnant survives stripping, so the case cannot go
vacuous if the stripper later learns SGR 7.

Co-authored-by: Amplify Logic AI <lars@sockinator.co>

* feat(cursor): narrow cursor identity and order its marker before CLAUDECODE

Cursor ships two executable names - `cursor-agent` and the legacy alias `agent`
- and runs as a bundled node script, so tmux reports the pane command as a bare
`node`. Neither `agent` nor `node` can be trusted by name, so identity gets one
owner in bin/fm-cursor-lib.sh that demands cursor's own name or install tree in
the path or argv[0], from the structural signal only. Probing an arbitrary pid's
executable during a liveness poll would execute a stranger's binary, which is
the hazard that rule exists to close.

Two consequences wired up:

Detection. cursor-agent does NOT clear an inherited CLAUDECODE, so a cursor
worker launched under a claude primary carries both markers and whichever is
tested first wins. The cursor markers are ordered ahead of the CLAUDECODE check;
fm-spawn additionally clears foreign markers at the launch boundary. Both are
kept deliberately - launch sanitization only covers sessions fm-spawn started,
while the ordering also covers a cursor session started by hand. Verified live
that CURSOR_INVOKED_AS is set on the agent process and CURSOR_AGENT=1 on the
child/tool processes fm-harness.sh actually runs as.

Pane liveness. A cursor pane now classifies `agent`. An unrelated node or agent
stays `other`, which the liveness callers already fold into `ambiguous` rather
than `dead`, so a stranger's node pane is never reported agent-free.

Resolution prints the STABLE launcher rather than the canonical target: identity
is proven through canonicalization, but cursor's canonical path carries a
version its own auto-update replaces, and pinning that would strand a task on a
version that can vanish.

The regression drives the two identity signals apart - a cursor-named executable
outside any cursor tree, and a non-cursor-named alias inside one - and asserts
each carries a verdict alone, so no single vendor string is load-bearing. Its
negative controls are real spawned processes, not fixtures.

Verified live on cursor-agent 2026.08.11-e8db854.

Co-authored-by: Ville Penttinen <villem.penttinen@gmail.com>

* feat(cursor): classify cursor busy state from its own turn transcript

Cursor shipped as "unknown cursor-unverified" on the premise that it exposes no
semantic turn lifecycle, only a rendered "Working" footer. That premise is
wrong: cursor-agent persists an append-only JSONL transcript per conversation
and brackets every submitted turn with a role:user open and a typed turn_ended
close. Verified live on 2026.08.11-e8db854, including the interrupt path, where
Escape closes the turn with status "aborted" - so this source covers manual
interruption, which Claude's Stop hook does not.

That makes it a genuine pull source in the muse mould rather than the rendered
text the redesign forbids: no writer, no arm, no gen, nothing seeded that could
never be cleared. Cursor's `ctrl+c to stop` footer stays out of the verdict, and
herdr's narrower native streaming state cannot stand in for it either.

Binding deliberately does not reconstruct cursor's workspace-slug directory
name. That slug collapses path separators, so rebuilding it would be a guess
that could bind the wrong pane; cursor records the exact absolute workspace path
in each project's .workspace-trusted, and the binding matches on that. A
conversation recorded as prior at spawn is excluded, so a relaunch in a reused
worktree folds its own turn rather than its predecessor's. Requiring a unique
remaining conversation keeps zero and several both unknown, because neither
proves anything about the current turn.

The regression pins the fold with real transcript files and asserts the
dangerous direction stays closed: an unresolvable binding, a record-free file,
an unclaimed workspace, and a workspace-path PREFIX all read unknown, never
idle. The prefix case uses an opaque fixture slug so a slug-rebuilding
implementation cannot pass it.

Co-authored-by: Ville Penttinen <villem.penttinen@gmail.com>

* feat(cursor): make the cursor launch runnable and give it lifecycle control

Five gaps that together kept a cursor crewmate from being drivable end to end.

Launch. The template invoked `cursor agent`, but `cursor` is not the CLI - the
installed names are `cursor-agent` and the legacy alias `agent` - so the command
could not run at all on a machine with a normal cursor install. It now resolves
through the verified owner, which also refuses a spawn loudly instead of leaving
a pane that dies with command-not-found and reads as a wedged worker.

Session binding. fm-spawn writes state/<id>.cursor-session so the busy fold can
find this pane's transcript, and teardown removes it.

Lifecycle control. No cursor PR touched fm-control-lib.sh, so
`fm-control <id> interrupt|exit|relaunch` could not drive a cursor worker at
all. Verified live: interrupt is a single Escape, exit is /exit, and cursor does
NOT repollute its composer with the cancelled prompt, so unlike muse it needs no
clear key. Secondmate is refused, matching the spawn refusal.

Submit acknowledgement. cursor parks its terminal cursor outside its composer,
so the composer verdict on tmux is always `unknown` and a submit could never be
acknowledged from the composer alone. The submit core's existing idle-to-busy
transition covers that case, but only if the pane's busy footer is recognised,
so cursor's `ctrl+c to stop` joins the harness-less default union the submit
cores read. The TOKEN is matched rather than the spinner verb: the same version
rendered both `Working` and `Running` in consecutive turns.

Bootstrap. A configured cursor crew harness with no cursor executable is now a
loud MISSING diagnostic rather than a first-spawn failure, and it accepts either
installed name.

Interrupt cancellation is deliberately left unconfirmed. The transcript does
type an aborted close, but its post-interrupt write latency measured as
variable - sometimes seconds, sometimes not within twenty - so a claim built on
it would be unreliable. Normal turn completion is prompt, which is what the busy
fold actually depends on.

Two inherited tests are corrected rather than deleted: the busy test asserted
cursor could have no semantic source, and the launch test pinned the literal
`cursor agent` string. Both now pin the verified behaviour, including that the
launch never allocates a second worktree.

Co-authored-by: ABHISHAKE KUMAR BOJJA <abojja@uvic.ca>
Co-authored-by: Ville Penttinen <villem.penttinen@gmail.com>

* docs(cursor): record the verified crewmate facts and extend the drift guard

The inherited cursor entry was written against 2026.08.04-aaa8809 and several of
its claims no longer hold: it named `cursor agent` as the binary (not the CLI
name), listed six Grok model ids of which the live catalog now returns two, and
recorded busy state, exit, interrupt, and skill invocation as unverified.

Replaced with what was measured against 2026.08.11-e8db854, including the two
facts most likely to be rediscovered painfully: cursor runs as a bundled node
script so its pane title is a bare `node`, and it parks its terminal cursor
outside its composer, which makes the tmux composer verdict permanently
`unknown` by design rather than a defect to chase.

Model ids now route to `--list-models` for the account instead of a fixed list,
since that list is exactly what drifted.

The live drift guard covers cursor, resolving it through the same verified owner
fm-spawn uses and passing --trust so the probe cannot hang on the workspace
prompt. Run against every installed harness: 8 checked, all alive, with cursor
reporting title='node' foreground=[.../cursor-agent] - the drift shape this
guard exists to catch.

Co-authored-by: Ville Penttinen <villem.penttinen@gmail.com>

* docs(agents): record the cursor session-binding state file

The state/ layout section is the inventory every session reads; a busy-source
binding that fm-spawn writes and teardown removes belongs in it alongside muse's.

* no-mistakes(review): Sanitize ambient Cursor marker in harness tests

* no-mistakes(review): Validate Cursor models against live catalog

* no-mistakes(review): Reject unsupported secondmates before binary preflight

* no-mistakes(review): Narrow Cursor ancestry detection to structured process identity

* no-mistakes(review): Parse Cursor transcripts and sanitize inherited markers

* no-mistakes(review): Handle malformed Cursor transcript records safely

* no-mistakes(review): Validate malformed Cursor closes in fallback parser

* no-mistakes(review): Retire stale Cursor bindings during relaunch

* no-mistakes(review): Fix Cursor drift guard command variable

* no-mistakes(review): Narrow Cursor identity to versioned install trees

* no-mistakes(document): Document Cursor harness boundaries

* refactor(composer): move the delivery busy footers to the shared owner

The per-harness rendered busy footers lived in bin/fm-tmux-lib.sh under
FM_TMUX_* names, so cursor's `ctrl+c to stop` signature - and every other
harness's - was reachable only from tmux. That placement was wrong on its own
terms: herdr, zellij, cmux, and orca run the same harnesses and face the same
question these footers answer, which is whether a submitted Enter actually
landed. Nothing about the signature is tmux-specific.

Moved verbatim into bin/fm-composer-lib.sh, the shared composer/delivery owner
every backend already sources, and renamed to FM_DELIVERY_* so the names stop
claiming a scope they never had. All five adapters now reach cursor's signature;
verified per adapter rather than assumed.

The boundary the move must not blur is stated where it now lives: this is a
DELIVERY guard, never a worker-state source. Confirming a keystroke landed is a
different question from asking what a worker is doing, and bin/fm-busy-lib.sh
remains the semantic owner that forbids classifying a harness from rendered
text. Cursor still classifies only from its transcript fold, which is already
backend-agnostic because it folds a file rather than reading a pane - the same
verdict on all six backends.

The old FM_TMUX_* aliases are dropped rather than kept as dead shims: nothing
outside the moved block referenced them except fm-busy-lib.sh's grok fallback,
which now reads the new name. The documented operator override, FM_BUSY_REGEX,
is untouched.

Also removes a dead duplicate CURSOR_INVOKED_AS check in bin/fm-harness.sh,
unreachable behind the marker check above it.

* no-mistakes(review): Correct shared delivery guard ownership references

* no-mistakes(document): Document shared delivery guards and Cursor backend limits

* no-mistakes: apply CI fixes

* fix(composer): bound a bare composer's wrap region at a half-block rule

A live cursor crewmate on herdr classified its IDLE composer as `pending`, and
fm-send consequently exited 1 with "delivery unconfirmed" on a message that had
actually landed. The cause is not cursor-specific.

Herdr draws a composer's top and bottom rules with the half-block glyphs U+2584
and U+2580 rather than the box-drawing family. fm_composer_row_has_edge knew
only the box-drawing set, so no box was detected; the composer was found as a
BARE row, and its wrap region - which extends while rows are non-blank and carry
no structural edge - walked straight through the composer's own closing rule and
swallowed the model and path footer below it. That footer is real text, so the
region classified pending on a genuinely idle pane.

Teaching the shared edge detector the half-block glyphs bounds the region at the
closing rule. Measured on the captured bytes of a real herdr cursor pane: the
same capture that read `pending` now reads `empty`.

This is a shared shape-path change, so it is deliberately narrow - it adds
glyphs to the edge vocabulary and changes no verdict logic - and the whole
composer and backend suite is green, including the other harnesses' herdr
fixtures.

The regression pins the real captured shape and asserts the footer content is
genuinely present, so the case cannot pass vacuously if the region were ever
bounded for some unrelated reason.

* fix(herdr): confirm a cursor submit from the rendered-footer transition

Herdr's composer-shape fix made an idle cursor pane classify `empty`, but
`fm-send` still exited 1 with "delivery unconfirmed" on messages that had
actually landed. Live measurement found the second, independent cause.

Herdr reports a cursor pane `agent_status=blocked` in EVERY state - idle,
mid-turn, and after - so the submit path's idle-baseline native confirmation is
structurally unreachable for cursor and every send falls into the composer
branch. That branch reads cursor's mid-turn composer row, which renders its own
`Add a follow-up` placeholder beside a right-aligned `ctrl+c to stop`. That
token is composer content, so the verdict is `pending` on a composer holding no
user text at all, and the Enter-retry budget then reports pending.

The escape is the same semantic signal the native path uses, read from the
pane's verified busy footer instead of native agent-state, and it is the
rendered-footer twin of the tmux submit core's turn-started confirmation: an
idle-to-busy transition ACROSS our Enter proves the harness accepted the
submission. The baseline is taken before the first Enter and only when the
native baseline was not legibly idle, so the idle-baseline path still never
reads pane content and a pane already mid-turn before we typed keeps reporting
`pending` rather than borrowing another turn as proof of this delivery.

The composer verdict is deliberately NOT relaxed. A right-aligned status token
on the composer row stays content for every other caller, including the
away-mode pre-injection guard, and the shared cursorless submit core is left
untouched so zellij, cmux, and Orca keep the behavior their own follow-up owns.

Verified live on herdr 0.8.0 and cursor-agent 2026.08.11-e8db854 in an isolated
lab session: `fm-send` now exits 0 and the steer executes, interrupt cancels a
running turn, `/exit` stops the agent, and teardown clears the record. All seven
panes of the running default session classify identically before and after the
shape fix, so no other harness regressed.

* no-mistakes(review): Prevent working Herdr baselines from falsely confirming delivery

* no-mistakes(document): Correct Cursor harness and backend documentation

---------

Co-authored-by: ABHISHAKE KUMAR BOJJA <abojja@uvic.ca>
Co-authored-by: Amplify Logic AI <lars@sockinator.co>
Co-authored-by: Ville Penttinen <villem.penttinen@gmail.com>

* fix(bin): require quota-axi 0.1.25 (#2300)

* fix: raise quota-axi floor to 0.1.25 for Cursor CLI quota awareness

Homes on latest main need quota-axi #87 so Desktop-absent CLI machines report a fresh Cursor quota instead of a false sign-in-required.

* no-mistakes(document): Update quota floor documentation pointer

* fix(bin): prevent false Pi watcher alarms during hand-offs (#2304)

* fix(guard): stop the false send-time watcher-down alarm on Pi primaries

On a Pi primary the watcher process is not the liveness signal. The Pi
extension tears the watcher down on every actionable wake and spawns the
replacement itself, so the singleton lock is legitimately unheld between
cycles: every one of the 799 cycles in a live primary's ledger ends with
lock_after=pid:none, and a live capture caught the guard verdict flipping to
no-watcher during one hand-off with the beacon 63s old.

bin/fm-guard.sh classified Pi as a persistent-watcher harness, which demands a
live identity-matched lock holder at all times, so any guarded command landing
in a hand-off painted the full WATCHER DOWN - SUPERVISION IS OFF banner and
told firstmate to repair a cycle the extension already owns and is restoring.

Add an extension supervision model for pi and pi-signed. A live
identity-matched watcher stays the ordinary healthy state; an unheld lock is
healthy only while the beacon is fresh within grace AND a live Pi session
provably owns continuity - both primary extensions recorded in their state
markers at their current on-disk builds by the process named in state/.lock,
with that process still alive. Without that proof the banner fires exactly as
before, so an unloaded, version-drifted, or exited Pi session is loud
immediately and a cycle the extension never restores is loud once the beacon
passes grace. The queued-wake warning, the PID-strict turn-end guard, and
every other primary's detection are untouched.

Fold session-start's duplicate Pi marker predicate into the shared library so
the ownership contract has one owner.

* no-mistakes(review): Restrict Pi hand-off tolerance to unheld watcher locks

* no-mistakes(document): Document Pi watcher hand-off supervision

* feat: support Cursor Agent CLI as a primary harness (#2305)

* feat(cursor): add Cursor Agent CLI primary hooks, park supervision, and session start

Register a tracked project-scope .cursor/hooks.json for Cursor's stop,
sessionStart, preCompact, and preToolUse steps.

bin/fm-turnend-guard-cursor.sh owns Cursor's turn boundary as a park: it
foregrounds the watcher arm, holds the boundary open until an actionable close,
and returns that wake as one follow-up. Exit 2 is a silent no-op on Cursor's
stop step, so the adapter never uses it. The follow-up loop is bounded twice,
by Cursor's own loop_limit and by the payload's loop_count.

bin/fm-sessionstart-cursor.sh delivers the digest as additional_context at
sessionStart, and stages it for the next turn boundary at preCompact, which
cannot inject context.

Cursor also loads the tracked Claude settings, so bin/fm-hook-host-lib.sh lets
each tracked Claude-shaped entrypoint stand down on a Cursor-delivered payload
rather than running every covered event twice.

bin/fm-tmux-lib.sh reclassifies a Cursor pane's composer cursorlessly, because
Cursor parks its terminal cursor outside the composer, which restores a genuine
composer-empty proof and unblocks away-mode escalation delivery.

* feat(cursor): make Cursor Agent CLI a verified primary harness

Resolve Cursor in the session-lock ancestry through bin/fm-cursor-lib.sh, which
a Cursor primary needs before it can hold its own home lock, and classify its
stop-hook park under the autoarm supervision model so the mid-turn pull guard
stops reporting a healthy between-turns watcher as down.

Read a Cursor pane's composer cursorlessly on tmux, gated on Cursor's own
structural process identity, which restores a genuine composer-empty proof and
lets away-mode escalations reach a Cursor primary with no daemon change.

Lift the secondmate refusals in bin/fm-spawn.sh and bin/fm-control-lib.sh now
that the supervision protocol exists and is recorded.

Cover the whole surface with a portable regression over real processes, an
opt-in live guard against the installed cursor-agent, and dated per-harness
evidence.

* docs(cursor): record Cursor as a verified primary across the owning surfaces

Update the turn-end guard, session-start, arm-seatbelt, cd-guard, watcher
continuity, architecture, configuration, README, and harness-adapters owners,
and add dated live evidence to the supervision and runtime-backend verification
records. Correct the recorded Cursor tmux composer verdict: the cursor-anchored
read is still blind, but the composite reader is no longer unknown.

Lift the remaining remote-secondmate refusal missed in the previous commit, and
add the new libs to the existing fixtures that copy a fixed dependency list.

* refactor(cursor): name the park's stand-down condition for both its causes

Also record that Cursor's preCompact firing itself is not yet live-verified,
while the static evidence that it cannot inject context, and the staging path
that follows from it, both are.

* test: give the pretool fixtures their new dependency and one lint owner

The cd-guard fixture copies a fixed dependency list and now needs the shared
hook-host predicate. Both pretool suites also asserted cleanliness with a bare
shellcheck call, a second and weaker copy of the lint definition that
bin/fm-lint.sh owns: it omits --external-sources, so it failed the moment these
checkers sourced a shared library. They now delegate to that owner.

* test: assert the cursor secondmate contract instead of its removed refusal

A cursor secondmate now launches, so the suite asserts what its park actually
needs: --trust so the home's project hooks load at all, its own home pinned as
the workspace, and the autoarm supervision model inherited across the launch.

* no-mistakes(review): Serialize Cursor wakes and bind staged context

* no-mistakes(review): Serialize Cursor context and nag state commits

* no-mistakes(review): Enforce Cursor ceiling before staged context delivery

* no-mistakes(review): Serialize Cursor claims and staged context

* no-mistakes(review): Serialize Cursor ownership and state commits

* no-mistakes(review): Protect Cursor context across session takeover

* no-mistakes(review): Preserve Cursor context across session takeover

* no-mistakes(review): Enforce owner-keyed Cursor staged context

* no-mistakes(review): Atomically claim Cursor follow-ups and staged context

* no-mistakes(review): Defer Cursor preCompact staging and simplify supersession

* no-mistakes(review): Serialize Cursor park commits and defer preCompact

* no-mistakes(review): Stop Cursor parks after session takeover

* no-mistakes(test): Route Cursor preCompact context through stop follow-up

* no-mistakes(document): Update Cursor primary documentation

* revert(cursor): cut preCompact staging from this change

Carrying a compaction digest across two concurrently running stop hooks kept
producing races that could deliver it twice or strand it indefinitely, and
closing them kept enlarging a critical section inside a hook Cursor awaits at
the turn boundary. Native preCompact firing was never observed either, so the
surface has no empirical basis yet.

Remove the adapter, its registration, its staged path in the park, and its
tests, and record the surface as deferred and uncovered alongside the Codex
interactive TUI. A regression now asserts preCompact stays unregistered so it
cannot return without its own design and evidence.

This change ships the proven core only: the turn-end follow-up park, the
run-tier session start, and away-mode delivery.

* no-mistakes(review): Correct Cursor park supersession documentation

* no-mistakes(document): Clarify Cursor run-tier verification ownership

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

---------

Co-authored-by: kunchenguid <kun-1@kunchenguid.com>

* feat(bin): add decline and repair paths for decision holds (#2330)

* feat(bin): add unrouted close paths to the captain decision gate

A captain who declines a held decision leaves no follow-up work to route,
so `resolve` could not express that answer: it requires at least one
`--routed-to` task. The only way to close such a hold was a direct
`tasks-axi done`, which never writes the durable resolution record the
completion gate reads, so the originating investigation could no longer
pass `verify` and its cleanup stayed blocked.

Add two close paths that route no work:

- `decline` closes an actively held hold with a recorded captain decision
  and no routed task. It refuses while any task is still blocked by the
  hold, because releasing routed work without recording it is `resolve`'s
  job.
- `repair` records the missing resolution block on a hold that was already
  closed outside this script. It never reopens a hold and never clears a
  dependency edge, and it refuses a hold that is still actively held.

Both require a non-empty captain decision file and share `resolve`'s
digest-based retry identity, so an exact retry is idempotent while a
changed decision is rejected. The recorded body now also names which path
closed the hold, and each routed entry regains its own line.

The gate itself is unchanged: an unanswered decision still fails
completion and blocks teardown, and neither new path can close a hold
without the captain's recorded word.

* fix(bin): require captain-hold provenance before repairing a decision

`repair` checked only that the backlog item was kind captain and Done, so
an ordinary captain-kind task that was never held for the captain could be
closed, repaired, and then pass the completion gate.

tasks-axi keeps `hold_kind` through a close, so it is the surviving proof
that an identity really was a captain hold. Require it before writing the
resolution record, and cover the case in the gate regression.

* no-mistakes(document): Correct decision-hold lifecycle documentation

* fix(bin): surface buried wake status lines once (#2331)

* fix(bin): surface buried status notes on wake drain

A note: answer immediately followed by a routine note was dropped because
annotations kept only the newest line and note: never enters OPEN DECISIONS.
Present every unread note and pending-reply resolution since the last drain
cursor, and annotate every unread line on a queued signal.

* no-mistakes(review): Fix unread status cursor races and overflow

* no-mistakes(review): Preserve cursors when status span reads fail

* no-mistakes(review): Make status presentation transactional under I/O failures

* no-mistakes(review): Simplify unread status cursor and presentation locking

* no-mistakes(review): Align cursor failure regressions with transactional presentation

* no-mistakes(review): Retire stale presentation cursors during task teardown

* no-mistakes(review): Preserve routine status until signal annotation

* no-mistakes(review): Correct unread status cap documentation

* no-mistakes(document): Document unread wake status presentation

* no-mistakes(lint): Fix wake surfacing ShellCheck warnings

* no-mistakes: apply CI fixes

* feat: add max Calm presentation level (#2334)

* feat(calm): add a max presentation level that hides mid-turn working notes

Calm's home-local preference becomes a three-state level instead of a
boolean: "off" is stock Pi, "on" is today's Calm, and "max" is Calm plus
hiding the assistant text of messages the model did not end its response
with. `/calm max` selects it from any state, a plain `/calm` steps max
back to ordinary Calm and otherwise keeps the existing on/off cycle, and
any other argument keeps that cycle too.

`config/calm` now persists "max" as its own literal value, so a session
start, resume, fork, or reload restores the stored level rather than
treating it as unrecognized and dropping to off.

The hide rule keys on Pi's intrinsic per-message stopReason: "toolUse",
or "length" with tool calls present. Streaming ("pending") text is never
filtered, because suppressing it would also stop a genuine reply from
streaming. The existing assistant layout adapter filters the blocks out
of the same shallow presentation copy it already uses for collapsed
thinking, so the message, model context, session storage, /export, and
delivery are untouched and a hidden mid-turn row collapses to zero
height. The new "assistant-working-note" class keeps that choice in the
visibility policy owner, where ordinary Calm keeps it visible.

* no-mistakes(document): Clarify Calm max persistence and taxonomy

* feat(calm): hide mid-turn working notes by default (#2339)

* feat(calm): make hiding mid-turn working notes the ordinary Calm state

Calm collapses back to the two-state on/off toggle it was before the max
presentation level, with max's hide rule promoted into ordinary Calm.
Calm on now hides mid-turn assistant working notes in addition to what it
already hid, and the /calm command parses no argument again.

The hide rule itself is unchanged: assistant text is removed from the
shallow presentation copy when the message's own stopReason is "toolUse",
or "length" with tool calls present. Streaming ("pending") text is never
filtered, so a genuine reply still streams. The message, model context,
session storage, /export, and delivery remain untouched.

config/calm persists only "on" and "off" again, but the reader still maps
a persisted "max" to on so a home upgraded from the removed level keeps
Calm on instead of dropping to off.

The mid-turn hide is now default behavior rather than an opt-in level, so
docs/calm.md documents it for users, docs/configuration.md records the
two written values plus the legacy max mapping, and the feasibility
taxonomy drops its level-scoped wording.

* no-mistakes(document): Document ordinary Calm working-note hiding

* chore: store no-mistakes test evidence in the repo (#2355)

* chore: ignore scratchpad/ at the repo root (#2359)

* fix(ci): fail hung Herdr behavior runs in 20 minutes (#2413)

A wedged family-run step was occupying the runner until the 75-minute
job cap; bound that step so cleanup and timing artifacts still upload.

* fix: keep the public promise reachable when work is routed to a second mate (#2457)

The lightweight Relay follow-up link lives in the answering home's own
state/<task-id>.meta, so it can only bind work that home owns. When a
Relay-linked request is routed to a second mate, the task record lives in the
second mate's home, fm-x-link.sh failed with a bare "no such task ...meta", and
nothing else picked the promise up: only the soft acknowledgement was ever
posted. The typed promised-final path already supports --work-home
secondmate:<id>; the playbook simply never chose it.

- fmx-respond now states the routing rule crisply: a task in this home takes the
  lightweight link, and second-mate-routed work takes a promised-final
  commitment bound to that home, registered up front with the brief command
  carried into the routed worker's instructions.
- fm-x-link.sh refuses a task with no local record by naming the registered
  second mate whose home actually holds it and printing the promised-final
  registration command, with the exact --work-home when the match is
  unambiguous. A home with no registered second mates keeps the plain error.
- fm-backlog-handoff.sh reports, after a successful move, any moved key that
  still owes a public reply bound to main/<key>, since that binding no longer
  names the home owning the work. The move itself is never blocked.

Docs and the secondmate handoff prose follow the same rule. Tests cover the
refusal, its scoping, the unchanged local-link path, and both handoff outcomes
at the script boundary.

* docs(skills): add remote-secondmate recovery hint for false-negative verdicts (#2456)

* fix(skills): hint that remote secondmate liveness verdicts false-negative

fm-crew-state and fm-send routinely misreport a live remote secondmate
as dead; confirm against the pane before relaunching, and relaunch only
through fm-spawn.sh, never raw herdr pane surgery.

* no-mistakes: apply CI fixes

* fix(calm): keep Pi's export confirmation visible (#2461)

Pi 0.83.0 added a status line to every tool-expansion change, and Pi
updates the previous status line in place when two status messages
arrive back to back. Calm's post-export redraw cycled tool expansion on
the macrotask right after Pi printed "Session exported to: <path>", so
both expansion status lines coalesced over that confirmation and the
captain was left with no record of where their export landed.

Calm now repaints only the tool rows it presents, by invalidating each
row through the render context Pi hands its render slots, and requests
the surrounding redraw through setStatus. Neither appends to the
transcript. The repaint is still needed because Pi can re-render a row
asynchronously - the built-in edit row invalidates itself once its diff
is ready - and that re-render can land inside the window where /export
forces stock rendering.

The real-terminal /export case now asserts the confirmation is still on
screen after the redraw has settled, and that the redraw restored every
Calm-hidden row, instead of only racing the moment the confirmation
first appeared.

* feat(stow): add open-record persistence to /stow before reset (#2488)

* feat(stow): persist the open records a session is holding

/stow curated memory and captured session knowledge, but never touched
record state, while AGENTS.md called it an "unfinished-work sweep" and the
receipt declared the session "safe to re…
BohnBawerick added a commit to BohnBawerick/firstmate that referenced this pull request Aug 18, 2026
* fix(bin): prevent false Pi watcher alarms during hand-offs (kunchenguid#2304)

* fix(guard): stop the false send-time watcher-down alarm on Pi primaries

On a Pi primary the watcher process is not the liveness signal. The Pi
extension tears the watcher down on every actionable wake and spawns the
replacement itself, so the singleton lock is legitimately unheld between
cycles: every one of the 799 cycles in a live primary's ledger ends with
lock_after=pid:none, and a live capture caught the guard verdict flipping to
no-watcher during one hand-off with the beacon 63s old.

bin/fm-guard.sh classified Pi as a persistent-watcher harness, which demands a
live identity-matched lock holder at all times, so any guarded command landing
in a hand-off painted the full WATCHER DOWN - SUPERVISION IS OFF banner and
told firstmate to repair a cycle the extension already owns and is restoring.

Add an extension supervision model for pi and pi-signed. A live
identity-matched watcher stays the ordinary healthy state; an unheld lock is
healthy only while the beacon is fresh within grace AND a live Pi session
provably owns continuity - both primary extensions recorded in their state
markers at their current on-disk builds by the process named in state/.lock,
with that process still alive. Without that proof the banner fires exactly as
before, so an unloaded, version-drifted, or exited Pi session is loud
immediately and a cycle the extension never restores is loud once the beacon
passes grace. The queued-wake warning, the PID-strict turn-end guard, and
every other primary's detection are untouched.

Fold session-start's duplicate Pi marker predicate into the shared library so
the ownership contract has one owner.

* no-mistakes(review): Restrict Pi hand-off tolerance to unheld watcher locks

* no-mistakes(document): Document Pi watcher hand-off supervision

* feat: support Cursor Agent CLI as a primary harness (kunchenguid#2305)

* feat(cursor): add Cursor Agent CLI primary hooks, park supervision, and session start

Register a tracked project-scope .cursor/hooks.json for Cursor's stop,
sessionStart, preCompact, and preToolUse steps.

bin/fm-turnend-guard-cursor.sh owns Cursor's turn boundary as a park: it
foregrounds the watcher arm, holds the boundary open until an actionable close,
and returns that wake as one follow-up. Exit 2 is a silent no-op on Cursor's
stop step, so the adapter never uses it. The follow-up loop is bounded twice,
by Cursor's own loop_limit and by the payload's loop_count.

bin/fm-sessionstart-cursor.sh delivers the digest as additional_context at
sessionStart, and stages it for the next turn boundary at preCompact, which
cannot inject context.

Cursor also loads the tracked Claude settings, so bin/fm-hook-host-lib.sh lets
each tracked Claude-shaped entrypoint stand down on a Cursor-delivered payload
rather than running every covered event twice.

bin/fm-tmux-lib.sh reclassifies a Cursor pane's composer cursorlessly, because
Cursor parks its terminal cursor outside the composer, which restores a genuine
composer-empty proof and unblocks away-mode escalation delivery.

* feat(cursor): make Cursor Agent CLI a verified primary harness

Resolve Cursor in the session-lock ancestry through bin/fm-cursor-lib.sh, which
a Cursor primary needs before it can hold its own home lock, and classify its
stop-hook park under the autoarm supervision model so the mid-turn pull guard
stops reporting a healthy between-turns watcher as down.

Read a Cursor pane's composer cursorlessly on tmux, gated on Cursor's own
structural process identity, which restores a genuine composer-empty proof and
lets away-mode escalations reach a Cursor primary with no daemon change.

Lift the secondmate refusals in bin/fm-spawn.sh and bin/fm-control-lib.sh now
that the supervision protocol exists and is recorded.

Cover the whole surface with a portable regression over real processes, an
opt-in live guard against the installed cursor-agent, and dated per-harness
evidence.

* docs(cursor): record Cursor as a verified primary across the owning surfaces

Update the turn-end guard, session-start, arm-seatbelt, cd-guard, watcher
continuity, architecture, configuration, README, and harness-adapters owners,
and add dated live evidence to the supervision and runtime-backend verification
records. Correct the recorded Cursor tmux composer verdict: the cursor-anchored
read is still blind, but the composite reader is no longer unknown.

Lift the remaining remote-secondmate refusal missed in the previous commit, and
add the new libs to the existing fixtures that copy a fixed dependency list.

* refactor(cursor): name the park's stand-down condition for both its causes

Also record that Cursor's preCompact firing itself is not yet live-verified,
while the static evidence that it cannot inject context, and the staging path
that follows from it, both are.

* test: give the pretool fixtures their new dependency and one lint owner

The cd-guard fixture copies a fixed dependency list and now needs the shared
hook-host predicate. Both pretool suites also asserted cleanliness with a bare
shellcheck call, a second and weaker copy of the lint definition that
bin/fm-lint.sh owns: it omits --external-sources, so it failed the moment these
checkers sourced a shared library. They now delegate to that owner.

* test: assert the cursor secondmate contract instead of its removed refusal

A cursor secondmate now launches, so the suite asserts what its park actually
needs: --trust so the home's project hooks load at all, its own home pinned as
the workspace, and the autoarm supervision model inherited across the launch.

* no-mistakes(review): Serialize Cursor wakes and bind staged context

* no-mistakes(review): Serialize Cursor context and nag state commits

* no-mistakes(review): Enforce Cursor ceiling before staged context delivery

* no-mistakes(review): Serialize Cursor claims and staged context

* no-mistakes(review): Serialize Cursor ownership and state commits

* no-mistakes(review): Protect Cursor context across session takeover

* no-mistakes(review): Preserve Cursor context across session takeover

* no-mistakes(review): Enforce owner-keyed Cursor staged context

* no-mistakes(review): Atomically claim Cursor follow-ups and staged context

* no-mistakes(review): Defer Cursor preCompact staging and simplify supersession

* no-mistakes(review): Serialize Cursor park commits and defer preCompact

* no-mistakes(review): Stop Cursor parks after session takeover

* no-mistakes(test): Route Cursor preCompact context through stop follow-up

* no-mistakes(document): Update Cursor primary documentation

* revert(cursor): cut preCompact staging from this change

Carrying a compaction digest across two concurrently running stop hooks kept
producing races that could deliver it twice or strand it indefinitely, and
closing them kept enlarging a critical section inside a hook Cursor awaits at
the turn boundary. Native preCompact firing was never observed either, so the
surface has no empirical basis yet.

Remove the adapter, its registration, its staged path in the park, and its
tests, and record the surface as deferred and uncovered alongside the Codex
interactive TUI. A regression now asserts preCompact stays unregistered so it
cannot return without its own design and evidence.

This change ships the proven core only: the turn-end follow-up park, the
run-tier session start, and away-mode delivery.

* no-mistakes(review): Correct Cursor park supersession documentation

* no-mistakes(document): Clarify Cursor run-tier verification ownership

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

---------

Co-authored-by: kunchenguid <kun-1@kunchenguid.com>

* feat(bin): add decline and repair paths for decision holds (kunchenguid#2330)

* feat(bin): add unrouted close paths to the captain decision gate

A captain who declines a held decision leaves no follow-up work to route,
so `resolve` could not express that answer: it requires at least one
`--routed-to` task. The only way to close such a hold was a direct
`tasks-axi done`, which never writes the durable resolution record the
completion gate reads, so the originating investigation could no longer
pass `verify` and its cleanup stayed blocked.

Add two close paths that route no work:

- `decline` closes an actively held hold with a recorded captain decision
  and no routed task. It refuses while any task is still blocked by the
  hold, because releasing routed work without recording it is `resolve`'s
  job.
- `repair` records the missing resolution block on a hold that was already
  closed outside this script. It never reopens a hold and never clears a
  dependency edge, and it refuses a hold that is still actively held.

Both require a non-empty captain decision file and share `resolve`'s
digest-based retry identity, so an exact retry is idempotent while a
changed decision is rejected. The recorded body now also names which path
closed the hold, and each routed entry regains its own line.

The gate itself is unchanged: an unanswered decision still fails
completion and blocks teardown, and neither new path can close a hold
without the captain's recorded word.

* fix(bin): require captain-hold provenance before repairing a decision

`repair` checked only that the backlog item was kind captain and Done, so
an ordinary captain-kind task that was never held for the captain could be
closed, repaired, and then pass the completion gate.

tasks-axi keeps `hold_kind` through a close, so it is the surviving proof
that an identity really was a captain hold. Require it before writing the
resolution record, and cover the case in the gate regression.

* no-mistakes(document): Correct decision-hold lifecycle documentation

* fix(bin): surface buried wake status lines once (kunchenguid#2331)

* fix(bin): surface buried status notes on wake drain

A note: answer immediately followed by a routine note was dropped because
annotations kept only the newest line and note: never enters OPEN DECISIONS.
Present every unread note and pending-reply resolution since the last drain
cursor, and annotate every unread line on a queued signal.

* no-mistakes(review): Fix unread status cursor races and overflow

* no-mistakes(review): Preserve cursors when status span reads fail

* no-mistakes(review): Make status presentation transactional under I/O failures

* no-mistakes(review): Simplify unread status cursor and presentation locking

* no-mistakes(review): Align cursor failure regressions with transactional presentation

* no-mistakes(review): Retire stale presentation cursors during task teardown

* no-mistakes(review): Preserve routine status until signal annotation

* no-mistakes(review): Correct unread status cap documentation

* no-mistakes(document): Document unread wake status presentation

* no-mistakes(lint): Fix wake surfacing ShellCheck warnings

* no-mistakes: apply CI fixes

* feat: add max Calm presentation level (kunchenguid#2334)

* feat(calm): add a max presentation level that hides mid-turn working notes

Calm's home-local preference becomes a three-state level instead of a
boolean: "off" is stock Pi, "on" is today's Calm, and "max" is Calm plus
hiding the assistant text of messages the model did not end its response
with. `/calm max` selects it from any state, a plain `/calm` steps max
back to ordinary Calm and otherwise keeps the existing on/off cycle, and
any other argument keeps that cycle too.

`config/calm` now persists "max" as its own literal value, so a session
start, resume, fork, or reload restores the stored level rather than
treating it as unrecognized and dropping to off.

The hide rule keys on Pi's intrinsic per-message stopReason: "toolUse",
or "length" with tool calls present. Streaming ("pending") text is never
filtered, because suppressing it would also stop a genuine reply from
streaming. The existing assistant layout adapter filters the blocks out
of the same shallow presentation copy it already uses for collapsed
thinking, so the message, model context, session storage, /export, and
delivery are untouched and a hidden mid-turn row collapses to zero
height. The new "assistant-working-note" class keeps that choice in the
visibility policy owner, where ordinary Calm keeps it visible.

* no-mistakes(document): Clarify Calm max persistence and taxonomy

* feat(calm): hide mid-turn working notes by default (kunchenguid#2339)

* feat(calm): make hiding mid-turn working notes the ordinary Calm state

Calm collapses back to the two-state on/off toggle it was before the max
presentation level, with max's hide rule promoted into ordinary Calm.
Calm on now hides mid-turn assistant working notes in addition to what it
already hid, and the /calm command parses no argument again.

The hide rule itself is unchanged: assistant text is removed from the
shallow presentation copy when the message's own stopReason is "toolUse",
or "length" with tool calls present. Streaming ("pending") text is never
filtered, so a genuine reply still streams. The message, model context,
session storage, /export, and delivery remain untouched.

config/calm persists only "on" and "off" again, but the reader still maps
a persisted "max" to on so a home upgraded from the removed level keeps
Calm on instead of dropping to off.

The mid-turn hide is now default behavior rather than an opt-in level, so
docs/calm.md documents it for users, docs/configuration.md records the
two written values plus the legacy max mapping, and the feasibility
taxonomy drops its level-scoped wording.

* no-mistakes(document): Document ordinary Calm working-note hiding

* chore: store no-mistakes test evidence in the repo (kunchenguid#2355)

* chore: ignore scratchpad/ at the repo root (kunchenguid#2359)

* fix(ci): fail hung Herdr behavior runs in 20 minutes (kunchenguid#2413)

A wedged family-run step was occupying the runner until the 75-minute
job cap; bound that step so cleanup and timing artifacts still upload.

* fix: keep the public promise reachable when work is routed to a second mate (kunchenguid#2457)

The lightweight Relay follow-up link lives in the answering home's own
state/<task-id>.meta, so it can only bind work that home owns. When a
Relay-linked request is routed to a second mate, the task record lives in the
second mate's home, fm-x-link.sh failed with a bare "no such task ...meta", and
nothing else picked the promise up: only the soft acknowledgement was ever
posted. The typed promised-final path already supports --work-home
secondmate:<id>; the playbook simply never chose it.

- fmx-respond now states the routing rule crisply: a task in this home takes the
  lightweight link, and second-mate-routed work takes a promised-final
  commitment bound to that home, registered up front with the brief command
  carried into the routed worker's instructions.
- fm-x-link.sh refuses a task with no local record by naming the registered
  second mate whose home actually holds it and printing the promised-final
  registration command, with the exact --work-home when the match is
  unambiguous. A home with no registered second mates keeps the plain error.
- fm-backlog-handoff.sh reports, after a successful move, any moved key that
  still owes a public reply bound to main/<key>, since that binding no longer
  names the home owning the work. The move itself is never blocked.

Docs and the secondmate handoff prose follow the same rule. Tests cover the
refusal, its scoping, the unchanged local-link path, and both handoff outcomes
at the script boundary.

* docs(skills): add remote-secondmate recovery hint for false-negative verdicts (kunchenguid#2456)

* fix(skills): hint that remote secondmate liveness verdicts false-negative

fm-crew-state and fm-send routinely misreport a live remote secondmate
as dead; confirm against the pane before relaunching, and relaunch only
through fm-spawn.sh, never raw herdr pane surgery.

* no-mistakes: apply CI fixes

* fix(calm): keep Pi's export confirmation visible (kunchenguid#2461)

Pi 0.83.0 added a status line to every tool-expansion change, and Pi
updates the previous status line in place when two status messages
arrive back to back. Calm's post-export redraw cycled tool expansion on
the macrotask right after Pi printed "Session exported to: <path>", so
both expansion status lines coalesced over that confirmation and the
captain was left with no record of where their export landed.

Calm now repaints only the tool rows it presents, by invalidating each
row through the render context Pi hands its render slots, and requests
the surrounding redraw through setStatus. Neither appends to the
transcript. The repaint is still needed because Pi can re-render a row
asynchronously - the built-in edit row invalidates itself once its diff
is ready - and that re-render can land inside the window where /export
forces stock rendering.

The real-terminal /export case now asserts the confirmation is still on
screen after the redraw has settled, and that the redraw restored every
Calm-hidden row, instead of only racing the moment the confirmation
first appeared.

* feat(stow): add open-record persistence to /stow before reset (kunchenguid#2488)

* feat(stow): persist the open records a session is holding

/stow curated memory and captured session knowledge, but never touched
record state, while AGENTS.md called it an "unfinished-work sweep" and the
receipt declared the session "safe to reset" - wording that implied a
record-correctness guarantee stow does not make. A shipped PR with no
backlog item, a queued umbrella whose phases had merged, and four decision
holds left open after their answers shipped all survived repeated stows.

Add a bounded pass that files record state from the same volatile input the
rest of stow already uses: the open threads in context, minutes before the
reset destroys them. It creates a record for an unfiled thread and corrects
one the session knows is wrong, through the owning path, and states its
boundary as part of the contract - it never enumerates the backlog, lists
holds, or queries a forge, because it cannot be a reconciliation and must
not be read as one.

Correct the wording in AGENTS.md and the completion receipt so reset-safe
means what it actually guarantees: nothing this session knew was lost.

* no-mistakes(review): correct stow decision-hold inspection to read hold via tasks-axi

* no-mistakes(document): note /stow open-record persistence in README command catalog

* refactor(stow): state open-record persistence as principle, not procedure

The first version enumerated triggers, named commands, and prescribed an
ordered procedure. That is too rigid for an agent skill: it invites literal
execution of a checklist instead of judgment, and every enumerated example
is a way for the guidance to go stale.

Reduce it to the intent - before a reset, the important open work you are
holding in context must end up durably recorded rather than dying with the
session, filing what is unfiled and correcting what is stale - and let the
agent judge importance, the record, and the owning write path.

Keep the scope bound, since it is a decided contract and not a mechanic:
this covers the open work the session is holding, never a reconciliation of
durable records against repository or forge reality. The wording
corrections in AGENTS.md and the completion receipt are unchanged.

* fix(decisions): close decision holds at answer time via one general keyed-answer path (kunchenguid#2490)

* fix(decisions): close captain holds at answer time

Firstmate had two "a decision is open" ledgers with asymmetric closing
mechanics. The live status-log ledger closes atomically at answer time,
because bin/fm-send.sh --resolve-key makes answering a decision be the
act that closes it. The durable backlog hold ledger had no such coupling:
answering and recording were two separate acts, and only the first was
forced by the workflow.

That asymmetry lost four real captain decisions. Their answers were
captured durably to disk, keyed character for character by the hold
decision keys, acknowledged, and even implemented and shipped, yet the
holds stayed open for two days and the captain was asked to re-answer
decisions already on his own disk.

Give the hold ledger the same answer-time-closure property:

- bin/fm-decision-hold.sh gains an `answer` subcommand, the hold ledger's
  counterpart to --resolve-key. It shares one unrouted close
  implementation with `decline`, so it carries every existing guard - the
  captain decision file, the active-hold requirement, retry identity, and
  the refusal to release still-routed work - and differs only in the
  resolution mode it records. `decline` keeps its stronger meaning that
  the answer routes no follow-up work at all.
- bin/fm-procevent-lavish.sh wires the channel that actually carried the
  lost answers. `arm --decisions-origin` binds a deck to the origin whose
  holds it carries, `answers` reads the structured choices out of a
  captured poll result, `close-decisions` maps each key to its hold and
  closes it through the command above, and `autohandle` lets the runner
  apply that at capture time.

Safety is preserved rather than traded away. Only rows tagged `choice`
are read, so freeform captain prose cannot forge a decision key. Closure
is confined to the one bound origin. The decision text is a pure function
of the captured result, so a replayed capture is idempotent. A hold that
is absent, already closed, or still blocking routed work is skipped and
left for `resolve`, never forced. A deck armed without the binding
touches no hold at all. And autohandle deliberately never reports full
handling, because recording an answer is transcription while acting on it
is firstmate's judgement - so the check wake still reaches the handler.

fm-send --resolve-key is untouched.

* no-mistakes(document): document state/lavish-decisions binding dir in AGENTS.md state inventory

* refactor(decisions): make keyed-answer closure one general capability

The previous pass gave holds answer-time closure but built it as bespoke
Lavish wiring: the review adapter carried the source-to-origin binding,
mapped keys to hold identities, wrote decision records, decided what to
skip, and closed holds itself. That treated a review deck as a special
decision source. It is not - it is an ephemeral discussion format that
happens to carry answers.

Collapse it into ONE general capability with one owner.

bin/fm-decision-hold.sh now owns the whole of "a keyed answer closes its
matching hold":
- `answers <origin> --source <provenance>` is the channel-agnostic
  intake. It reads key/answer/label lines on stdin, maps each key to its
  hold, and closes it through the same `answer` path, so every guard
  applies identically whatever channel the answer came from. --source is
  provenance recorded in the decision, never a behavior switch; there is
  no per-channel branch and no knowledge of chat, decks, or transports.
- `bind`/`unbind`/`binding` own the source-to-origin binding for any
  channel whose answers arrive detached from their origin.

Every channel is now an ordinary caller that only turns what it received
into keyed lines:
- bin/fm-send.sh (chat) feeds the intake for a key that names an active
  hold. This also fixes a real gap: once `complete` transfers a decision
  to its hold it closes the live status copy, so --resolve-key alone
  could never answer a transferred decision.
- bin/fm-procevent.sh feeds it generically. A bound source's captured
  result goes to `<adapter> answers <result-file>` and whatever that
  prints is piped into the intake. The runner names no adapter, parses
  no result, and carries no decision rule, so any future adapter with an
  `answers` command works with no change here.
- bin/fm-procevent-lavish.sh keeps only `answers`, which reports the
  structured choices a review captured and stops. It maps nothing to a
  hold and closes nothing; it lost ~160 lines of decision logic.

Feeding is independent of handling, so it never acknowledges a result
and never suppresses a wake - recording an answer is transcription,
acting on it stays firstmate's judgement.

The regression that proves closure now drives a FIXTURE adapter that is
not the review adapter, so what is proven is that any bound channel
reaches the intake rather than that one channel is wired specially. A
new regression drives the real fm-send over a stubbed transport for the
chat side. Every prior guarantee still holds, and fm-send's status-log
behavior is unchanged.

* no-mistakes(review): test(decisions): drop source-content grep from hold-closure regression

* fix(memory): emit a real @AGENTS.md pointer instead of a CLAUDE.md symlink (kunchenguid#2512)

A Write aimed at CLAUDE.md followed the symlink and destroyed AGENTS.md.
The installer now creates and migrates to a recoverable two-line pointer file.

* fix(ci): keep CLAUDE.md pointer check valid (kunchenguid#2515)

* ci: gate GitHub workflows with pinned actionlint (kunchenguid#2517)

* fix(lint): catch malformed GitHub workflows before merge

A self-broken ci.yml cannot report its own breakage, so parse every
workflow in the local lint path that no-mistakes already runs.

* fix(lint): pin actionlint instead of Ruby for workflow lint

A self-broken ci.yml still has to fail in the local lint path, and the
named tool for that gate is actionlint, not a new Ruby runtime.

* no-mistakes(document): Clarify pinned workflow lint documentation

* fix: install pinned lint tools across supported platforms (kunchenguid#2546)

* fix: install pinned shellcheck and actionlint on macOS and linux arm64

The installers were hardcoded to linux amd64 and sha256sum, so a Mac
dev could not satisfy the refuse-on-mismatch lint gate. Select the
official per-platform archive and checksum, and fall back to shasum -a 256.

* no-mistakes(document): Document cross-platform pinned lint installers

* docs: reconcile test-evidence docs with store_in_repo: true (kunchenguid#2548)

.no-mistakes.yaml has set test.evidence.store_in_repo: true since kunchenguid#2355, but
CONTRIBUTING.md, docs/configuration.md, and docs/architecture.md still described
the old policy of keeping evidence out of the repo in a temp directory.

The current no-mistakes behavior for store_in_repo: true is to publish each run's
test evidence to the orphan no-mistakes/evidence branch and link it from the PR
body. That branch shares no history with code branches, so evidence never enters
a pushed feature branch or the default branch, and CI's tracked personal fleet
paths rule stays accurate.

Docs only. No change to .no-mistakes.yaml or any workflow.

* docs: clarify test evidence branch storage (kunchenguid#2549)

* docs: correct test evidence storage comment in .no-mistakes.yaml

* no-mistakes: apply CI fixes

* docs: hint that live scouts may host their own Lavish review loop (kunchenguid#2563)

Make that a first-class option in always-loaded instructions so firstmate does not default to mediating and tearing the scout down between iteration rounds.

* fix(bin): report remote secondmate delivery and state truthfully (kunchenguid#2570)

* fix(bin): report remote secondmate delivery and state truthfully

A steer to a remote secondmate crosses fm-on.sh to a host-local fm-send
leg whose unconfirmed submit read-back (verdict=pending, typically a busy
mate whose harness queues the steer) was flattened into exit 1, so the
parent printed "error: text not submitted" / "error: text not sent" and
discarded the pending-reply expectation for a steer that had actually
landed. fm-send now carries the verdict across the ssh boundary as a
documented delivered-unconfirmed exit 3: the parent reports the steer as
delivered with confirmation pending, exits 0, keeps the expectation armed
(awaiting_report), and closes --resolve-key decisions, while transport
loss (ssh 255) and real remote failures keep failing loudly with the
remote leg's stderr attached. A local unconfirmed submit now also exits 3
with an honest non-error message and still never closes a decision key.

fm-crew-state.sh and fm-peek.sh no longer read a remote mate's endpoint
through local probes (which misreported a healthy mate as "worktree gone"
/ "can't find session: remote"): both now use the true remote source over
fm-on.sh, and an unreachable or unreadable remote reads as unknown-remote,
never as gone or dead.

* no-mistakes(document): Document remote delivery and state truth

* no-mistakes: apply CI fixes

* fix(bin): distinguish unreadable away-mode panes from gone ones

Away-mode housekeeping treated every failed capture as a gone pane and
dropped the marker with no escalation. A redraw, timeout, or backend
hiccup then silently stopped watching a worker that was still there,
which is the failure this path exists to prevent.

Both the stale-wedge and pause-resurface sites now share
stale_window_recheck: retry the capture twice (0.4s apart) before
verdict, then ask fm_backend_agent_state. Only an authoritatively
missing endpoint is gone. A present dead shell is ordinary idle. Every
other state, including an unreadable or unverified probe, escalates and
keeps the marker on the same cadence because the watcher cannot
recapture an unreadable pane. target_exists is not used as a gone proof:
tmux can fall back to the active window, and Orca's check is itself a
capture.

Tests cover gone, unreadable-present (alive/unreadable/unverified),
retry-then-ordinary, and dead-is-not-gone at both call sites.

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: kunchenguid <kun-1@kunchenguid.com>
kaan-sirin pushed a commit to kaan-sirin/firstmate that referenced this pull request Aug 18, 2026
A wedged family-run step was occupying the runner until the 75-minute
job cap; bound that step so cleanup and timing artifacts still upload.
BohnBawerick pushed a commit to BohnBawerick/firstmate that referenced this pull request Aug 19, 2026
A wedged family-run step was occupying the runner until the 75-minute
job cap; bound that step so cleanup and timing artifacts still upload.
joliverMI added a commit to joliverMI/firstmate that referenced this pull request Aug 19, 2026
…ng state (#2)

* fix(ci): fail hung Herdr behavior runs in 20 minutes (kunchenguid#2413)

A wedged family-run step was occupying the runner until the 75-minute
job cap; bound that step so cleanup and timing artifacts still upload.

* fix: keep the public promise reachable when work is routed to a second mate (kunchenguid#2457)

The lightweight Relay follow-up link lives in the answering home's own
state/<task-id>.meta, so it can only bind work that home owns. When a
Relay-linked request is routed to a second mate, the task record lives in the
second mate's home, fm-x-link.sh failed with a bare "no such task ...meta", and
nothing else picked the promise up: only the soft acknowledgement was ever
posted. The typed promised-final path already supports --work-home
secondmate:<id>; the playbook simply never chose it.

- fmx-respond now states the routing rule crisply: a task in this home takes the
  lightweight link, and second-mate-routed work takes a promised-final
  commitment bound to that home, registered up front with the brief command
  carried into the routed worker's instructions.
- fm-x-link.sh refuses a task with no local record by naming the registered
  second mate whose home actually holds it and printing the promised-final
  registration command, with the exact --work-home when the match is
  unambiguous. A home with no registered second mates keeps the plain error.
- fm-backlog-handoff.sh reports, after a successful move, any moved key that
  still owes a public reply bound to main/<key>, since that binding no longer
  names the home owning the work. The move itself is never blocked.

Docs and the secondmate handoff prose follow the same rule. Tests cover the
refusal, its scoping, the unchanged local-link path, and both handoff outcomes
at the script boundary.

* docs(skills): add remote-secondmate recovery hint for false-negative verdicts (kunchenguid#2456)

* fix(skills): hint that remote secondmate liveness verdicts false-negative

fm-crew-state and fm-send routinely misreport a live remote secondmate
as dead; confirm against the pane before relaunching, and relaunch only
through fm-spawn.sh, never raw herdr pane surgery.

* no-mistakes: apply CI fixes

* fix(calm): keep Pi's export confirmation visible (kunchenguid#2461)

Pi 0.83.0 added a status line to every tool-expansion change, and Pi
updates the previous status line in place when two status messages
arrive back to back. Calm's post-export redraw cycled tool expansion on
the macrotask right after Pi printed "Session exported to: <path>", so
both expansion status lines coalesced over that confirmation and the
captain was left with no record of where their export landed.

Calm now repaints only the tool rows it presents, by invalidating each
row through the render context Pi hands its render slots, and requests
the surrounding redraw through setStatus. Neither appends to the
transcript. The repaint is still needed because Pi can re-render a row
asynchronously - the built-in edit row invalidates itself once its diff
is ready - and that re-render can land inside the window where /export
forces stock rendering.

The real-terminal /export case now asserts the confirmation is still on
screen after the redraw has settled, and that the redraw restored every
Calm-hidden row, instead of only racing the moment the confirmation
first appeared.

* feat(bin): render the captain's four-section status board from existing state

fm-status-board.sh prints one HTML page (Needs you to continue, In
progress, Waiting, Recently completed) rendered entirely from
fm-bearings-snapshot.sh's already-correct cross-home classification and
fm-fleet-snapshot.sh's untruncated per-item detail, with no second store
for agents to keep in sync.

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: joliverMI <joliver@sensibletech.biz>
joliverMI added a commit to joliverMI/firstmate that referenced this pull request Aug 19, 2026
* fix(ci): fail hung Herdr behavior runs in 20 minutes (kunchenguid#2413)

A wedged family-run step was occupying the runner until the 75-minute
job cap; bound that step so cleanup and timing artifacts still upload.

* fix: keep the public promise reachable when work is routed to a second mate (kunchenguid#2457)

The lightweight Relay follow-up link lives in the answering home's own
state/<task-id>.meta, so it can only bind work that home owns. When a
Relay-linked request is routed to a second mate, the task record lives in the
second mate's home, fm-x-link.sh failed with a bare "no such task ...meta", and
nothing else picked the promise up: only the soft acknowledgement was ever
posted. The typed promised-final path already supports --work-home
secondmate:<id>; the playbook simply never chose it.

- fmx-respond now states the routing rule crisply: a task in this home takes the
  lightweight link, and second-mate-routed work takes a promised-final
  commitment bound to that home, registered up front with the brief command
  carried into the routed worker's instructions.
- fm-x-link.sh refuses a task with no local record by naming the registered
  second mate whose home actually holds it and printing the promised-final
  registration command, with the exact --work-home when the match is
  unambiguous. A home with no registered second mates keeps the plain error.
- fm-backlog-handoff.sh reports, after a successful move, any moved key that
  still owes a public reply bound to main/<key>, since that binding no longer
  names the home owning the work. The move itself is never blocked.

Docs and the secondmate handoff prose follow the same rule. Tests cover the
refusal, its scoping, the unchanged local-link path, and both handoff outcomes
at the script boundary.

* docs(skills): add remote-secondmate recovery hint for false-negative verdicts (kunchenguid#2456)

* fix(skills): hint that remote secondmate liveness verdicts false-negative

fm-crew-state and fm-send routinely misreport a live remote secondmate
as dead; confirm against the pane before relaunching, and relaunch only
through fm-spawn.sh, never raw herdr pane surgery.

* no-mistakes: apply CI fixes

* fix(calm): keep Pi's export confirmation visible (kunchenguid#2461)

Pi 0.83.0 added a status line to every tool-expansion change, and Pi
updates the previous status line in place when two status messages
arrive back to back. Calm's post-export redraw cycled tool expansion on
the macrotask right after Pi printed "Session exported to: <path>", so
both expansion status lines coalesced over that confirmation and the
captain was left with no record of where their export landed.

Calm now repaints only the tool rows it presents, by invalidating each
row through the render context Pi hands its render slots, and requests
the surrounding redraw through setStatus. Neither appends to the
transcript. The repaint is still needed because Pi can re-render a row
asynchronously - the built-in edit row invalidates itself once its diff
is ready - and that re-render can land inside the window where /export
forces stock rendering.

The real-terminal /export case now asserts the confirmation is still on
screen after the redraw has settled, and that the redraw restored every
Calm-hidden row, instead of only racing the moment the confirmation
first appeared.

* fix(procevent): deliver captured results at most once

A Lavish-captured answer reached a secondmate five times because
forwarding it ran fm-send.sh and only afterward, manually, remembered to
call `handled` - so every re-announcement of the still-unacknowledged
capture (a restart, a compaction, a re-read wake) repeated the forward.
The retry loop was ours, not lavish-axi's: fm-procevent.sh already proves
capture is exactly-once, but nothing paired a downstream effect with its
acknowledgement atomically.

Add `fm-procevent.sh deliver <source-id> <sequence> -- <command>...`,
which checks handled status, runs the command, and marks the generation
handled as one call under the source lock, so a repeated delivery attempt
against an already-delivered generation runs the command zero times and a
failed command stays eligible for retry instead of being dropped. Point
the process-event-sources skill and the runner's operating contract at it
for exactly this class of forward.

Five delivery attempts against one captured generation now run the
downstream command once instead of five times - the read-and-discard
cost of the other four is eliminated structurally. The historical
incident's own token cost was not measured at the time and is not
reconstructable after the fact.

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: joliverMI <joliver@sensibletech.biz>
prajwal-395 added a commit to prajwal-395/firstmate that referenced this pull request Aug 19, 2026
…-ahead behaviour (#21)

* fix(ci): fail hung Herdr behavior runs in 20 minutes (kunchenguid#2413)

A wedged family-run step was occupying the runner until the 75-minute
job cap; bound that step so cleanup and timing artifacts still upload.

* fix: keep the public promise reachable when work is routed to a second mate (kunchenguid#2457)

The lightweight Relay follow-up link lives in the answering home's own
state/<task-id>.meta, so it can only bind work that home owns. When a
Relay-linked request is routed to a second mate, the task record lives in the
second mate's home, fm-x-link.sh failed with a bare "no such task ...meta", and
nothing else picked the promise up: only the soft acknowledgement was ever
posted. The typed promised-final path already supports --work-home
secondmate:<id>; the playbook simply never chose it.

- fmx-respond now states the routing rule crisply: a task in this home takes the
  lightweight link, and second-mate-routed work takes a promised-final
  commitment bound to that home, registered up front with the brief command
  carried into the routed worker's instructions.
- fm-x-link.sh refuses a task with no local record by naming the registered
  second mate whose home actually holds it and printing the promised-final
  registration command, with the exact --work-home when the match is
  unambiguous. A home with no registered second mates keeps the plain error.
- fm-backlog-handoff.sh reports, after a successful move, any moved key that
  still owes a public reply bound to main/<key>, since that binding no longer
  names the home owning the work. The move itself is never blocked.

Docs and the secondmate handoff prose follow the same rule. Tests cover the
refusal, its scoping, the unchanged local-link path, and both handoff outcomes
at the script boundary.

* docs(skills): add remote-secondmate recovery hint for false-negative verdicts (kunchenguid#2456)

* fix(skills): hint that remote secondmate liveness verdicts false-negative

fm-crew-state and fm-send routinely misreport a live remote secondmate
as dead; confirm against the pane before relaunching, and relaunch only
through fm-spawn.sh, never raw herdr pane surgery.

* no-mistakes: apply CI fixes

* fix(calm): keep Pi's export confirmation visible (kunchenguid#2461)

Pi 0.83.0 added a status line to every tool-expansion change, and Pi
updates the previous status line in place when two status messages
arrive back to back. Calm's post-export redraw cycled tool expansion on
the macrotask right after Pi printed "Session exported to: <path>", so
both expansion status lines coalesced over that confirmation and the
captain was left with no record of where their export landed.

Calm now repaints only the tool rows it presents, by invalidating each
row through the render context Pi hands its render slots, and requests
the surrounding redraw through setStatus. Neither appends to the
transcript. The repaint is still needed because Pi can re-render a row
asynchronously - the built-in edit row invalidates itself once its diff
is ready - and that re-render can land inside the window where /export
forces stock rendering.

The real-terminal /export case now asserts the confirmation is still on
screen after the redraw has settled, and that the redraw restored every
Calm-hidden row, instead of only racing the moment the confirmation
first appeared.

* feat(stow): add open-record persistence to /stow before reset (kunchenguid#2488)

* feat(stow): persist the open records a session is holding

/stow curated memory and captured session knowledge, but never touched
record state, while AGENTS.md called it an "unfinished-work sweep" and the
receipt declared the session "safe to reset" - wording that implied a
record-correctness guarantee stow does not make. A shipped PR with no
backlog item, a queued umbrella whose phases had merged, and four decision
holds left open after their answers shipped all survived repeated stows.

Add a bounded pass that files record state from the same volatile input the
rest of stow already uses: the open threads in context, minutes before the
reset destroys them. It creates a record for an unfiled thread and corrects
one the session knows is wrong, through the owning path, and states its
boundary as part of the contract - it never enumerates the backlog, lists
holds, or queries a forge, because it cannot be a reconciliation and must
not be read as one.

Correct the wording in AGENTS.md and the completion receipt so reset-safe
means what it actually guarantees: nothing this session knew was lost.

* no-mistakes(review): correct stow decision-hold inspection to read hold via tasks-axi

* no-mistakes(document): note /stow open-record persistence in README command catalog

* refactor(stow): state open-record persistence as principle, not procedure

The first version enumerated triggers, named commands, and prescribed an
ordered procedure. That is too rigid for an agent skill: it invites literal
execution of a checklist instead of judgment, and every enumerated example
is a way for the guidance to go stale.

Reduce it to the intent - before a reset, the important open work you are
holding in context must end up durably recorded rather than dying with the
session, filing what is unfiled and correcting what is stale - and let the
agent judge importance, the record, and the owning write path.

Keep the scope bound, since it is a decided contract and not a mechanic:
this covers the open work the session is holding, never a reconciliation of
durable records against repository or forge reality. The wording
corrections in AGENTS.md and the completion receipt are unchanged.

* fix(decisions): close decision holds at answer time via one general keyed-answer path (kunchenguid#2490)

* fix(decisions): close captain holds at answer time

Firstmate had two "a decision is open" ledgers with asymmetric closing
mechanics. The live status-log ledger closes atomically at answer time,
because bin/fm-send.sh --resolve-key makes answering a decision be the
act that closes it. The durable backlog hold ledger had no such coupling:
answering and recording were two separate acts, and only the first was
forced by the workflow.

That asymmetry lost four real captain decisions. Their answers were
captured durably to disk, keyed character for character by the hold
decision keys, acknowledged, and even implemented and shipped, yet the
holds stayed open for two days and the captain was asked to re-answer
decisions already on his own disk.

Give the hold ledger the same answer-time-closure property:

- bin/fm-decision-hold.sh gains an `answer` subcommand, the hold ledger's
  counterpart to --resolve-key. It shares one unrouted close
  implementation with `decline`, so it carries every existing guard - the
  captain decision file, the active-hold requirement, retry identity, and
  the refusal to release still-routed work - and differs only in the
  resolution mode it records. `decline` keeps its stronger meaning that
  the answer routes no follow-up work at all.
- bin/fm-procevent-lavish.sh wires the channel that actually carried the
  lost answers. `arm --decisions-origin` binds a deck to the origin whose
  holds it carries, `answers` reads the structured choices out of a
  captured poll result, `close-decisions` maps each key to its hold and
  closes it through the command above, and `autohandle` lets the runner
  apply that at capture time.

Safety is preserved rather than traded away. Only rows tagged `choice`
are read, so freeform captain prose cannot forge a decision key. Closure
is confined to the one bound origin. The decision text is a pure function
of the captured result, so a replayed capture is idempotent. A hold that
is absent, already closed, or still blocking routed work is skipped and
left for `resolve`, never forced. A deck armed without the binding
touches no hold at all. And autohandle deliberately never reports full
handling, because recording an answer is transcription while acting on it
is firstmate's judgement - so the check wake still reaches the handler.

fm-send --resolve-key is untouched.

* no-mistakes(document): document state/lavish-decisions binding dir in AGENTS.md state inventory

* refactor(decisions): make keyed-answer closure one general capability

The previous pass gave holds answer-time closure but built it as bespoke
Lavish wiring: the review adapter carried the source-to-origin binding,
mapped keys to hold identities, wrote decision records, decided what to
skip, and closed holds itself. That treated a review deck as a special
decision source. It is not - it is an ephemeral discussion format that
happens to carry answers.

Collapse it into ONE general capability with one owner.

bin/fm-decision-hold.sh now owns the whole of "a keyed answer closes its
matching hold":
- `answers <origin> --source <provenance>` is the channel-agnostic
  intake. It reads key/answer/label lines on stdin, maps each key to its
  hold, and closes it through the same `answer` path, so every guard
  applies identically whatever channel the answer came from. --source is
  provenance recorded in the decision, never a behavior switch; there is
  no per-channel branch and no knowledge of chat, decks, or transports.
- `bind`/`unbind`/`binding` own the source-to-origin binding for any
  channel whose answers arrive detached from their origin.

Every channel is now an ordinary caller that only turns what it received
into keyed lines:
- bin/fm-send.sh (chat) feeds the intake for a key that names an active
  hold. This also fixes a real gap: once `complete` transfers a decision
  to its hold it closes the live status copy, so --resolve-key alone
  could never answer a transferred decision.
- bin/fm-procevent.sh feeds it generically. A bound source's captured
  result goes to `<adapter> answers <result-file>` and whatever that
  prints is piped into the intake. The runner names no adapter, parses
  no result, and carries no decision rule, so any future adapter with an
  `answers` command works with no change here.
- bin/fm-procevent-lavish.sh keeps only `answers`, which reports the
  structured choices a review captured and stops. It maps nothing to a
  hold and closes nothing; it lost ~160 lines of decision logic.

Feeding is independent of handling, so it never acknowledges a result
and never suppresses a wake - recording an answer is transcription,
acting on it stays firstmate's judgement.

The regression that proves closure now drives a FIXTURE adapter that is
not the review adapter, so what is proven is that any bound channel
reaches the intake rather than that one channel is wired specially. A
new regression drives the real fm-send over a stubbed transport for the
chat side. Every prior guarantee still holds, and fm-send's status-log
behavior is unchanged.

* no-mistakes(review): test(decisions): drop source-content grep from hold-closure regression

* fix(memory): emit a real @AGENTS.md pointer instead of a CLAUDE.md symlink (kunchenguid#2512)

A Write aimed at CLAUDE.md followed the symlink and destroyed AGENTS.md.
The installer now creates and migrates to a recoverable two-line pointer file.

* fix(ci): keep CLAUDE.md pointer check valid (kunchenguid#2515)

* ci: gate GitHub workflows with pinned actionlint (kunchenguid#2517)

* fix(lint): catch malformed GitHub workflows before merge

A self-broken ci.yml cannot report its own breakage, so parse every
workflow in the local lint path that no-mistakes already runs.

* fix(lint): pin actionlint instead of Ruby for workflow lint

A self-broken ci.yml still has to fail in the local lint path, and the
named tool for that gate is actionlint, not a new Ruby runtime.

* no-mistakes(document): Clarify pinned workflow lint documentation

* fix: install pinned lint tools across supported platforms (kunchenguid#2546)

* fix: install pinned shellcheck and actionlint on macOS and linux arm64

The installers were hardcoded to linux amd64 and sha256sum, so a Mac
dev could not satisfy the refuse-on-mismatch lint gate. Select the
official per-platform archive and checksum, and fall back to shasum -a 256.

* no-mistakes(document): Document cross-platform pinned lint installers

* docs: reconcile test-evidence docs with store_in_repo: true (kunchenguid#2548)

.no-mistakes.yaml has set test.evidence.store_in_repo: true since kunchenguid#2355, but
CONTRIBUTING.md, docs/configuration.md, and docs/architecture.md still described
the old policy of keeping evidence out of the repo in a temp directory.

The current no-mistakes behavior for store_in_repo: true is to publish each run's
test evidence to the orphan no-mistakes/evidence branch and link it from the PR
body. That branch shares no history with code branches, so evidence never enters
a pushed feature branch or the default branch, and CI's tracked personal fleet
paths rule stays accurate.

Docs only. No change to .no-mistakes.yaml or any workflow.

* docs: clarify test evidence branch storage (kunchenguid#2549)

* docs: correct test evidence storage comment in .no-mistakes.yaml

* no-mistakes: apply CI fixes

* docs: hint that live scouts may host their own Lavish review loop (kunchenguid#2563)

Make that a first-class option in always-loaded instructions so firstmate does not default to mediating and tearing the scout down between iteration rounds.

* fix(bin): report remote secondmate delivery and state truthfully (kunchenguid#2570)

* fix(bin): report remote secondmate delivery and state truthfully

A steer to a remote secondmate crosses fm-on.sh to a host-local fm-send
leg whose unconfirmed submit read-back (verdict=pending, typically a busy
mate whose harness queues the steer) was flattened into exit 1, so the
parent printed "error: text not submitted" / "error: text not sent" and
discarded the pending-reply expectation for a steer that had actually
landed. fm-send now carries the verdict across the ssh boundary as a
documented delivered-unconfirmed exit 3: the parent reports the steer as
delivered with confirmation pending, exits 0, keeps the expectation armed
(awaiting_report), and closes --resolve-key decisions, while transport
loss (ssh 255) and real remote failures keep failing loudly with the
remote leg's stderr attached. A local unconfirmed submit now also exits 3
with an honest non-error message and still never closes a decision key.

fm-crew-state.sh and fm-peek.sh no longer read a remote mate's endpoint
through local probes (which misreported a healthy mate as "worktree gone"
/ "can't find session: remote"): both now use the true remote source over
fm-on.sh, and an unreachable or unreadable remote reads as unknown-remote,
never as gone or dead.

* no-mistakes(document): Document remote delivery and state truth

* no-mistakes: apply CI fixes

* feat: adopt spendPriority for quota dispatch (kunchenguid#2574)

* Adopt quota-axi 0.1.29 spendPriority-primary array dispatch.

quota-axi 0.1.29 publishes schema 5 with selection.spendPriority as the primary comparative signal and demotes derivation fields out of default --json. Rank comparable-fit candidates on that scalar, keep runway versus the completion horizon as a hard gate, and raise the compatibility floor so a pre-consolidation build cannot reach dispatch intake.

* no-mistakes(review): Correct schema fixtures and remove prescriptive selection prompts

* no-mistakes(document): Correct quota verification evidence chronology

* Collapse quota-array-dispatch onto TOON-first spendPriority ranking.

Decide from quota-axi's default TOON; keep --json as a rare defensive fallback.
Rank by spendPriority after eligibility, reasoning-class, and runway-feasibility gates, and drop the hand-computed Pareto, pace, reserve, and window-id layers.

* no-mistakes(review): Permit ambiguous JSON fallback and correct reset fixtures

* no-mistakes(review): Correct runway semantics and escalate unresolved uncertainty

* no-mistakes(document): Document TOON-first quota dispatch evidence

* docs: add GROK_BOT.md Grok Bot system prompt (kunchenguid#2590)

* docs: add GROK_BOT.md Grok Bot system prompt

* docs: amend GROK_BOT.md with charter report-back and delegation marker

* docs: classify GROK_BOT.md as public-product

* docs: make GROK_BOT.md the plain Grok Bot system prompt

* docs: update GROK_BOT.md nautical terms and self-improvement (kunchenguid#2592)

* doc: Update language in GROK_BOT.md for clarity

Refine language for clarity and consistency in instructions.

* fix(bin): preserve inactive reconciliation scan progress (kunchenguid#2595)

* fix(bin): guarantee inactive-reconcile scan progress under second quantization

The inactive-outcome scan computed its aggregate deadline in whole seconds,
so a 1-second budget's effective value lands anywhere in (0,1]; a scan
starting just before a wall-clock second boundary rounded its whole budget
away mid-scan and exited having visited no child, while the durable cursor
had already advanced past the never-examined child. This is the CI flake
behind tests/fm-inactive-reconcile.test.sh's 'next bounded scan did not
resume with the following child' (watcher-wake-lock family, portable
serial 2, seen on the PR kunchenguid#2590 run).

Every scan now visits at least its first due child with the per-child
state-read bound floored at one second, so no invocation can be a zero-work
no-op. The outer process-group kill moves to budget+1s: the scan's own
deadline enforces the budget, and the kill is a backstop for a scan wedged
in an unbounded wait instead of a racer that routinely preempts the clean
bounded exit. The wake-lock-wait test bound tracks the backstop (3s -> 4s);
the previously flaky assertion is unchanged.

* no-mistakes(document): Document inactive-reconcile deadline backstop

* doc: Revise Firstmate delegation and communication guidelines

Refactor the guidelines for Firstmate's role and delegation process, emphasizing the importance of crewmates and asynchronous work.

* doc: Update work delegation and secret management instructions

Clarified guidelines for handing off work to crewmates and managing secrets.

* test(procevent): make the process-event suite's detached-runner assertions deterministic (kunchenguid#2617)

Three assertions in tests/fm-procevent.test.sh depended on a detached runner
having finished work that the command starting it does not wait for.

reconcile's replacement runner is started through detach_runner, which only
forks: reconcile returns and counts the start before that runner has claimed
its source or exec'd its child. Any assertion taken straight after reconcile
therefore samples a race.

- The publish-before-apply recovery section left its always-ready /bin/echo
  source registered across the recovery reconcile, so that reconcile launched
  a competing detached poll (observed: started=1) that then raced every later
  assertion for the source claim, the next capture sequence, and this home's
  applied record, and outlived the section holding a live claim. It is now
  retired before that reconcile - re-announcement is proven from the durable
  inbox alone and needs no registration - and started=0 is asserted so a
  competing poll cannot be reintroduced unnoticed. This is the same
  retire-before-reconcile discipline the self-announcing section already
  carries; that section acquired it after the identical race made its
  "not-autohandled: self-src" assertion read "already owned: self-src".

- The crashed-leader replacement section snapshotted the replacement's claim
  file and execution log behind a fixed 0.5s settle window. On a loaded
  machine that window expires first, which is the CI flake behind "a
  replacement runner started without recording its own claim" and "reconcile
  did not start exactly one replacement source". Both effects are now waited
  for with the suite's bounded wait helpers; the exact one-replacement count
  is still asserted afterwards, unchanged.

- The duplicate-start section slept 0.5s for reconcile's runner to record
  ownership before asserting that a second start loses to it. It now waits
  for that claim.

Also tighten one assertion that could not fail as written: "autohandled:
self-src" is a substring of "not-autohandled: self-src", so the applied path
was accepted even when the runner reported the capture left for the handler.

Evidence: on the unmodified suite, 128 full runs at 6-8x concurrency produced
6 failing runs, all in the crashed-leader section. On the fixed suite, 216
full runs under the same load produced none. Reverting the self-announcing
section's retire-before-reconcile line reproduces "already owned: self-src"
on the first iteration, confirming the shared mechanism.

* fix: preserve pending replies and defer remote reposts (kunchenguid#2618)

* fix(bin): keep pending-reply expectations honest on both send legs

Two related asymmetries let the parent-owned secondmate reply guard drop or
nag requests it should not have.

Local delivered-unconfirmed dropped the expectation. A marked request whose
submit read-back stayed unconfirmed (verdict=pending) is the same
not-a-failure outcome the remote leg reports as delivered, but fm-send
discarded the parent's pending-reply record for it, so a request that very
likely landed stopped being tracked entirely. The record now stays armed on
its unconfirmed-delivery marker: a correlated report still resolves it, and
an unanswered one still surfaces through the library's own reconciliation.
Exit 3 and the local rule that an unconfirmed answer never closes a decision
key are unchanged.

Remote replies were nagged for a repost they did not need. A remote mate's
report reaches the parent's status log only through the asynchronous mirror
in fm-procevent-remote-reply.sh, yet the guard read an absent correlated
line as proof the mate never reported - even while the answer was still in
flight, which is the common case because the mirror's poll window is
comparable to the recovery grace. The mirror now publishes one caught-up
watermark from a quiet window, and the guard admits a missing report as
evidence only once that watermark passes the turn that should have produced
it. A genuinely missed report still gets exactly one repost, and a channel
that is behind, unarmed, or broken leaves the request durably open and
un-nagged rather than nagging blind; the mirror escalates its own continuity
failures as before.

Tests: a local unconfirmed secondmate send keeps its expectation armed and
resolvable; a mirrored correlated remote reply resolves with no repost; a
stale or absent watermark withholds the repost while a fresh one still
releases it; a quiet remote window publishes the watermark and retirement
clears it.

* no-mistakes(review): Distinguish preempted polls from quiet windows

* no-mistakes(document): Clarify remote reply channel freshness

* no-mistakes(lint): Annotate shared remote preemption exit constant

* fix(bin): honor declared pauses in busy-pane wedge checks (kunchenguid#2619)

* fix(watch): honor a declared pause on a busy pane's completed-turn bound

A worker that declares an external wait (`paused:`) and then blocks in one
long foreground call - a review-hosting scout parked in a single blocking
`lavish-axi poll`, a bounded watch loop, a rate-limit sleep - keeps its pane
BUSY, so the stale path that already honors declared pauses never ran for it.
The busy-pane completed-turn bound instead routed it straight into
wedge_timer_check, which re-escalated "possible wedge, escalation N" (and, past
the threshold, demand-deep-inspection) every FM_STALE_ESCALATE_SECS for as long
as the review stayed open.

busy_turn_bound_check now owns which absorber takes a crossed bound: a crew
whose own last status line declares an external wait or a verified captain-held
transfer takes the bounded FM_PAUSE_RESURFACE_SECS recheck, and everything else
keeps the unchanged wedge timer. The discriminator is the declaration together
with liveness (the caller has already confirmed the pane is busy), never a
blanket silencing - a crew that declared nothing, or whose pane is not live,
escalates exactly as before, and a declared pause still re-surfaces once per
long cadence so a forgotten wait cannot rot invisibly. Away mode is untouched:
the daemon owns pause triage there and already reads the same vocabulary.

The two call sites also no longer clear pause bookkeeping in the same poll the
pause cadence recorded it, which would have erased the re-surface throttle and
turned the long cadence back into a per-poll re-surface.

Tests: a three-phase regression fixture pins the absorbed pause, its long-cadence
recheck, and the restored wedge escalation once the declaration is lifted on the
same busy over-age pane.

Also de-flakes tests/fm-watch-triage.test.sh, which failed spuriously on a loaded
machine: fixed liveness budgets were reaping watchers mid-startup, so assertions
on post-poll state passed vacuously or failed spuriously. Waits that describe a
poll's outcome now wait for a completed poll cycle via the liveness beacon, the
heartbeat test waits for the heartbeat it asserts on, and every wait_for_exit
budget is the uniform 10s already used elsewhere in the file.

* no-mistakes(review): Fail poll-cycle waits on timeout

* no-mistakes(review): Prevent poll timeout test hangs

* no-mistakes(document): Clarify paused busy-pane supervision

* fix(ci): rebalance portable serial shard weight hints from measured CI timings

The old weight hints (from 2026-08-02, 69 scripts) left 52 tests at the
20000ms default. Many live-e2e tests actually run in <100ms while several
new tests run over 100s, so the LPT algorithm's apparently balanced
assignment was wildly skewed in practice: shard 1 took ~14.1 min vs
7.9-9.8 for the others, hitting the 15-min job timeout.

Refresh all hints from CI run 32302063973 (2026-08-19, 119 scripts,
~39 min total). Three unmeasured new tests use 5000ms conservative
estimates. The rebalanced shards each estimate ~589s (~9.8 min) with
a 2ms imbalance.

Also bump the CI timeout from 15 to 20 minutes (~2x margin on the
~10 min balanced target), and update docs/fm-test-portable-shards.md
to match.

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Bre77 added a commit to Bre77/firstmate that referenced this pull request Aug 22, 2026
…s) (#70)

* fix(ci): fail hung Herdr behavior runs in 20 minutes (#2413)

A wedged family-run step was occupying the runner until the 75-minute
job cap; bound that step so cleanup and timing artifacts still upload.

* fix: keep the public promise reachable when work is routed to a second mate (#2457)

The lightweight Relay follow-up link lives in the answering home's own
state/<task-id>.meta, so it can only bind work that home owns. When a
Relay-linked request is routed to a second mate, the task record lives in the
second mate's home, fm-x-link.sh failed with a bare "no such task ...meta", and
nothing else picked the promise up: only the soft acknowledgement was ever
posted. The typed promised-final path already supports --work-home
secondmate:<id>; the playbook simply never chose it.

- fmx-respond now states the routing rule crisply: a task in this home takes the
  lightweight link, and second-mate-routed work takes a promised-final
  commitment bound to that home, registered up front with the brief command
  carried into the routed worker's instructions.
- fm-x-link.sh refuses a task with no local record by naming the registered
  second mate whose home actually holds it and printing the promised-final
  registration command, with the exact --work-home when the match is
  unambiguous. A home with no registered second mates keeps the plain error.
- fm-backlog-handoff.sh reports, after a successful move, any moved key that
  still owes a public reply bound to main/<key>, since that binding no longer
  names the home owning the work. The move itself is never blocked.

Docs and the secondmate handoff prose follow the same rule. Tests cover the
refusal, its scoping, the unchanged local-link path, and both handoff outcomes
at the script boundary.

* docs(skills): add remote-secondmate recovery hint for false-negative verdicts (#2456)

* fix(skills): hint that remote secondmate liveness verdicts false-negative

fm-crew-state and fm-send routinely misreport a live remote secondmate
as dead; confirm against the pane before relaunching, and relaunch only
through fm-spawn.sh, never raw herdr pane surgery.

* no-mistakes: apply CI fixes

* fix(calm): keep Pi's export confirmation visible (#2461)

Pi 0.83.0 added a status line to every tool-expansion change, and Pi
updates the previous status line in place when two status messages
arrive back to back. Calm's post-export redraw cycled tool expansion on
the macrotask right after Pi printed "Session exported to: <path>", so
both expansion status lines coalesced over that confirmation and the
captain was left with no record of where their export landed.

Calm now repaints only the tool rows it presents, by invalidating each
row through the render context Pi hands its render slots, and requests
the surrounding redraw through setStatus. Neither appends to the
transcript. The repaint is still needed because Pi can re-render a row
asynchronously - the built-in edit row invalidates itself once its diff
is ready - and that re-render can land inside the window where /export
forces stock rendering.

The real-terminal /export case now asserts the confirmation is still on
screen after the redraw has settled, and that the redraw restored every
Calm-hidden row, instead of only racing the moment the confirmation
first appeared.

* feat(stow): add open-record persistence to /stow before reset (#2488)

* feat(stow): persist the open records a session is holding

/stow curated memory and captured session knowledge, but never touched
record state, while AGENTS.md called it an "unfinished-work sweep" and the
receipt declared the session "safe to reset" - wording that implied a
record-correctness guarantee stow does not make. A shipped PR with no
backlog item, a queued umbrella whose phases had merged, and four decision
holds left open after their answers shipped all survived repeated stows.

Add a bounded pass that files record state from the same volatile input the
rest of stow already uses: the open threads in context, minutes before the
reset destroys them. It creates a record for an unfiled thread and corrects
one the session knows is wrong, through the owning path, and states its
boundary as part of the contract - it never enumerates the backlog, lists
holds, or queries a forge, because it cannot be a reconciliation and must
not be read as one.

Correct the wording in AGENTS.md and the completion receipt so reset-safe
means what it actually guarantees: nothing this session knew was lost.

* no-mistakes(review): correct stow decision-hold inspection to read hold via tasks-axi

* no-mistakes(document): note /stow open-record persistence in README command catalog

* refactor(stow): state open-record persistence as principle, not procedure

The first version enumerated triggers, named commands, and prescribed an
ordered procedure. That is too rigid for an agent skill: it invites literal
execution of a checklist instead of judgment, and every enumerated example
is a way for the guidance to go stale.

Reduce it to the intent - before a reset, the important open work you are
holding in context must end up durably recorded rather than dying with the
session, filing what is unfiled and correcting what is stale - and let the
agent judge importance, the record, and the owning write path.

Keep the scope bound, since it is a decided contract and not a mechanic:
this covers the open work the session is holding, never a reconciliation of
durable records against repository or forge reality. The wording
corrections in AGENTS.md and the completion receipt are unchanged.

* fix(decisions): close decision holds at answer time via one general keyed-answer path (#2490)

* fix(decisions): close captain holds at answer time

Firstmate had two "a decision is open" ledgers with asymmetric closing
mechanics. The live status-log ledger closes atomically at answer time,
because bin/fm-send.sh --resolve-key makes answering a decision be the
act that closes it. The durable backlog hold ledger had no such coupling:
answering and recording were two separate acts, and only the first was
forced by the workflow.

That asymmetry lost four real captain decisions. Their answers were
captured durably to disk, keyed character for character by the hold
decision keys, acknowledged, and even implemented and shipped, yet the
holds stayed open for two days and the captain was asked to re-answer
decisions already on his own disk.

Give the hold ledger the same answer-time-closure property:

- bin/fm-decision-hold.sh gains an `answer` subcommand, the hold ledger's
  counterpart to --resolve-key. It shares one unrouted close
  implementation with `decline`, so it carries every existing guard - the
  captain decision file, the active-hold requirement, retry identity, and
  the refusal to release still-routed work - and differs only in the
  resolution mode it records. `decline` keeps its stronger meaning that
  the answer routes no follow-up work at all.
- bin/fm-procevent-lavish.sh wires the channel that actually carried the
  lost answers. `arm --decisions-origin` binds a deck to the origin whose
  holds it carries, `answers` reads the structured choices out of a
  captured poll result, `close-decisions` maps each key to its hold and
  closes it through the command above, and `autohandle` lets the runner
  apply that at capture time.

Safety is preserved rather than traded away. Only rows tagged `choice`
are read, so freeform captain prose cannot forge a decision key. Closure
is confined to the one bound origin. The decision text is a pure function
of the captured result, so a replayed capture is idempotent. A hold that
is absent, already closed, or still blocking routed work is skipped and
left for `resolve`, never forced. A deck armed without the binding
touches no hold at all. And autohandle deliberately never reports full
handling, because recording an answer is transcription while acting on it
is firstmate's judgement - so the check wake still reaches the handler.

fm-send --resolve-key is untouched.

* no-mistakes(document): document state/lavish-decisions binding dir in AGENTS.md state inventory

* refactor(decisions): make keyed-answer closure one general capability

The previous pass gave holds answer-time closure but built it as bespoke
Lavish wiring: the review adapter carried the source-to-origin binding,
mapped keys to hold identities, wrote decision records, decided what to
skip, and closed holds itself. That treated a review deck as a special
decision source. It is not - it is an ephemeral discussion format that
happens to carry answers.

Collapse it into ONE general capability with one owner.

bin/fm-decision-hold.sh now owns the whole of "a keyed answer closes its
matching hold":
- `answers <origin> --source <provenance>` is the channel-agnostic
  intake. It reads key/answer/label lines on stdin, maps each key to its
  hold, and closes it through the same `answer` path, so every guard
  applies identically whatever channel the answer came from. --source is
  provenance recorded in the decision, never a behavior switch; there is
  no per-channel branch and no knowledge of chat, decks, or transports.
- `bind`/`unbind`/`binding` own the source-to-origin binding for any
  channel whose answers arrive detached from their origin.

Every channel is now an ordinary caller that only turns what it received
into keyed lines:
- bin/fm-send.sh (chat) feeds the intake for a key that names an active
  hold. This also fixes a real gap: once `complete` transfers a decision
  to its hold it closes the live status copy, so --resolve-key alone
  could never answer a transferred decision.
- bin/fm-procevent.sh feeds it generically. A bound source's captured
  result goes to `<adapter> answers <result-file>` and whatever that
  prints is piped into the intake. The runner names no adapter, parses
  no result, and carries no decision rule, so any future adapter with an
  `answers` command works with no change here.
- bin/fm-procevent-lavish.sh keeps only `answers`, which reports the
  structured choices a review captured and stops. It maps nothing to a
  hold and closes nothing; it lost ~160 lines of decision logic.

Feeding is independent of handling, so it never acknowledges a result
and never suppresses a wake - recording an answer is transcription,
acting on it stays firstmate's judgement.

The regression that proves closure now drives a FIXTURE adapter that is
not the review adapter, so what is proven is that any bound channel
reaches the intake rather than that one channel is wired specially. A
new regression drives the real fm-send over a stubbed transport for the
chat side. Every prior guarantee still holds, and fm-send's status-log
behavior is unchanged.

* no-mistakes(review): test(decisions): drop source-content grep from hold-closure regression

* fix(memory): emit a real @AGENTS.md pointer instead of a CLAUDE.md symlink (#2512)

A Write aimed at CLAUDE.md followed the symlink and destroyed AGENTS.md.
The installer now creates and migrates to a recoverable two-line pointer file.

* fix(ci): keep CLAUDE.md pointer check valid (#2515)

* ci: gate GitHub workflows with pinned actionlint (#2517)

* fix(lint): catch malformed GitHub workflows before merge

A self-broken ci.yml cannot report its own breakage, so parse every
workflow in the local lint path that no-mistakes already runs.

* fix(lint): pin actionlint instead of Ruby for workflow lint

A self-broken ci.yml still has to fail in the local lint path, and the
named tool for that gate is actionlint, not a new Ruby runtime.

* no-mistakes(document): Clarify pinned workflow lint documentation

* fix: install pinned lint tools across supported platforms (#2546)

* fix: install pinned shellcheck and actionlint on macOS and linux arm64

The installers were hardcoded to linux amd64 and sha256sum, so a Mac
dev could not satisfy the refuse-on-mismatch lint gate. Select the
official per-platform archive and checksum, and fall back to shasum -a 256.

* no-mistakes(document): Document cross-platform pinned lint installers

* docs: reconcile test-evidence docs with store_in_repo: true (#2548)

.no-mistakes.yaml has set test.evidence.store_in_repo: true since #2355, but
CONTRIBUTING.md, docs/configuration.md, and docs/architecture.md still described
the old policy of keeping evidence out of the repo in a temp directory.

The current no-mistakes behavior for store_in_repo: true is to publish each run's
test evidence to the orphan no-mistakes/evidence branch and link it from the PR
body. That branch shares no history with code branches, so evidence never enters
a pushed feature branch or the default branch, and CI's tracked personal fleet
paths rule stays accurate.

Docs only. No change to .no-mistakes.yaml or any workflow.

* docs: clarify test evidence branch storage (#2549)

* docs: correct test evidence storage comment in .no-mistakes.yaml

* no-mistakes: apply CI fixes

* docs: hint that live scouts may host their own Lavish review loop (#2563)

Make that a first-class option in always-loaded instructions so firstmate does not default to mediating and tearing the scout down between iteration rounds.

* fix(bin): report remote secondmate delivery and state truthfully (#2570)

* fix(bin): report remote secondmate delivery and state truthfully

A steer to a remote secondmate crosses fm-on.sh to a host-local fm-send
leg whose unconfirmed submit read-back (verdict=pending, typically a busy
mate whose harness queues the steer) was flattened into exit 1, so the
parent printed "error: text not submitted" / "error: text not sent" and
discarded the pending-reply expectation for a steer that had actually
landed. fm-send now carries the verdict across the ssh boundary as a
documented delivered-unconfirmed exit 3: the parent reports the steer as
delivered with confirmation pending, exits 0, keeps the expectation armed
(awaiting_report), and closes --resolve-key decisions, while transport
loss (ssh 255) and real remote failures keep failing loudly with the
remote leg's stderr attached. A local unconfirmed submit now also exits 3
with an honest non-error message and still never closes a decision key.

fm-crew-state.sh and fm-peek.sh no longer read a remote mate's endpoint
through local probes (which misreported a healthy mate as "worktree gone"
/ "can't find session: remote"): both now use the true remote source over
fm-on.sh, and an unreachable or unreadable remote reads as unknown-remote,
never as gone or dead.

* no-mistakes(document): Document remote delivery and state truth

* no-mistakes: apply CI fixes

* feat: adopt spendPriority for quota dispatch (#2574)

* Adopt quota-axi 0.1.29 spendPriority-primary array dispatch.

quota-axi 0.1.29 publishes schema 5 with selection.spendPriority as the primary comparative signal and demotes derivation fields out of default --json. Rank comparable-fit candidates on that scalar, keep runway versus the completion horizon as a hard gate, and raise the compatibility floor so a pre-consolidation build cannot reach dispatch intake.

* no-mistakes(review): Correct schema fixtures and remove prescriptive selection prompts

* no-mistakes(document): Correct quota verification evidence chronology

* Collapse quota-array-dispatch onto TOON-first spendPriority ranking.

Decide from quota-axi's default TOON; keep --json as a rare defensive fallback.
Rank by spendPriority after eligibility, reasoning-class, and runway-feasibility gates, and drop the hand-computed Pareto, pace, reserve, and window-id layers.

* no-mistakes(review): Permit ambiguous JSON fallback and correct reset fixtures

* no-mistakes(review): Correct runway semantics and escalate unresolved uncertainty

* no-mistakes(document): Document TOON-first quota dispatch evidence

* docs: add GROK_BOT.md Grok Bot system prompt (#2590)

* docs: add GROK_BOT.md Grok Bot system prompt

* docs: amend GROK_BOT.md with charter report-back and delegation marker

* docs: classify GROK_BOT.md as public-product

* docs: make GROK_BOT.md the plain Grok Bot system prompt

* docs: update GROK_BOT.md nautical terms and self-improvement (#2592)

* doc: Update language in GROK_BOT.md for clarity

Refine language for clarity and consistency in instructions.

* fix(bin): preserve inactive reconciliation scan progress (#2595)

* fix(bin): guarantee inactive-reconcile scan progress under second quantization

The inactive-outcome scan computed its aggregate deadline in whole seconds,
so a 1-second budget's effective value lands anywhere in (0,1]; a scan
starting just before a wall-clock second boundary rounded its whole budget
away mid-scan and exited having visited no child, while the durable cursor
had already advanced past the never-examined child. This is the CI flake
behind tests/fm-inactive-reconcile.test.sh's 'next bounded scan did not
resume with the following child' (watcher-wake-lock family, portable
serial 2, seen on the PR #2590 run).

Every scan now visits at least its first due child with the per-child
state-read bound floored at one second, so no invocation can be a zero-work
no-op. The outer process-group kill moves to budget+1s: the scan's own
deadline enforces the budget, and the kill is a backstop for a scan wedged
in an unbounded wait instead of a racer that routinely preempts the clean
bounded exit. The wake-lock-wait test bound tracks the backstop (3s -> 4s);
the previously flaky assertion is unchanged.

* no-mistakes(document): Document inactive-reconcile deadline backstop

* doc: Revise Firstmate delegation and communication guidelines

Refactor the guidelines for Firstmate's role and delegation process, emphasizing the importance of crewmates and asynchronous work.

* doc: Update work delegation and secret management instructions

Clarified guidelines for handing off work to crewmates and managing secrets.

* test(procevent): make the process-event suite's detached-runner assertions deterministic (#2617)

Three assertions in tests/fm-procevent.test.sh depended on a detached runner
having finished work that the command starting it does not wait for.

reconcile's replacement runner is started through detach_runner, which only
forks: reconcile returns and counts the start before that runner has claimed
its source or exec'd its child. Any assertion taken straight after reconcile
therefore samples a race.

- The publish-before-apply recovery section left its always-ready /bin/echo
  source registered across the recovery reconcile, so that reconcile launched
  a competing detached poll (observed: started=1) that then raced every later
  assertion for the source claim, the next capture sequence, and this home's
  applied record, and outlived the section holding a live claim. It is now
  retired before that reconcile - re-announcement is proven from the durable
  inbox alone and needs no registration - and started=0 is asserted so a
  competing poll cannot be reintroduced unnoticed. This is the same
  retire-before-reconcile discipline the self-announcing section already
  carries; that section acquired it after the identical race made its
  "not-autohandled: self-src" assertion read "already owned: self-src".

- The crashed-leader replacement section snapshotted the replacement's claim
  file and execution log behind a fixed 0.5s settle window. On a loaded
  machine that window expires first, which is the CI flake behind "a
  replacement runner started without recording its own claim" and "reconcile
  did not start exactly one replacement source". Both effects are now waited
  for with the suite's bounded wait helpers; the exact one-replacement count
  is still asserted afterwards, unchanged.

- The duplicate-start section slept 0.5s for reconcile's runner to record
  ownership before asserting that a second start loses to it. It now waits
  for that claim.

Also tighten one assertion that could not fail as written: "autohandled:
self-src" is a substring of "not-autohandled: self-src", so the applied path
was accepted even when the runner reported the capture left for the handler.

Evidence: on the unmodified suite, 128 full runs at 6-8x concurrency produced
6 failing runs, all in the crashed-leader section. On the fixed suite, 216
full runs under the same load produced none. Reverting the self-announcing
section's retire-before-reconcile line reproduces "already owned: self-src"
on the first iteration, confirming the shared mechanism.

* fix: preserve pending replies and defer remote reposts (#2618)

* fix(bin): keep pending-reply expectations honest on both send legs

Two related asymmetries let the parent-owned secondmate reply guard drop or
nag requests it should not have.

Local delivered-unconfirmed dropped the expectation. A marked request whose
submit read-back stayed unconfirmed (verdict=pending) is the same
not-a-failure outcome the remote leg reports as delivered, but fm-send
discarded the parent's pending-reply record for it, so a request that very
likely landed stopped being tracked entirely. The record now stays armed on
its unconfirmed-delivery marker: a correlated report still resolves it, and
an unanswered one still surfaces through the library's own reconciliation.
Exit 3 and the local rule that an unconfirmed answer never closes a decision
key are unchanged.

Remote replies were nagged for a repost they did not need. A remote mate's
report reaches the parent's status log only through the asynchronous mirror
in fm-procevent-remote-reply.sh, yet the guard read an absent correlated
line as proof the mate never reported - even while the answer was still in
flight, which is the common case because the mirror's poll window is
comparable to the recovery grace. The mirror now publishes one caught-up
watermark from a quiet window, and the guard admits a missing report as
evidence only once that watermark passes the turn that should have produced
it. A genuinely missed report still gets exactly one repost, and a channel
that is behind, unarmed, or broken leaves the request durably open and
un-nagged rather than nagging blind; the mirror escalates its own continuity
failures as before.

Tests: a local unconfirmed secondmate send keeps its expectation armed and
resolvable; a mirrored correlated remote reply resolves with no repost; a
stale or absent watermark withholds the repost while a fresh one still
releases it; a quiet remote window publishes the watermark and retirement
clears it.

* no-mistakes(review): Distinguish preempted polls from quiet windows

* no-mistakes(document): Clarify remote reply channel freshness

* no-mistakes(lint): Annotate shared remote preemption exit constant

* fix(bin): honor declared pauses in busy-pane wedge checks (#2619)

* fix(watch): honor a declared pause on a busy pane's completed-turn bound

A worker that declares an external wait (`paused:`) and then blocks in one
long foreground call - a review-hosting scout parked in a single blocking
`lavish-axi poll`, a bounded watch loop, a rate-limit sleep - keeps its pane
BUSY, so the stale path that already honors declared pauses never ran for it.
The busy-pane completed-turn bound instead routed it straight into
wedge_timer_check, which re-escalated "possible wedge, escalation N" (and, past
the threshold, demand-deep-inspection) every FM_STALE_ESCALATE_SECS for as long
as the review stayed open.

busy_turn_bound_check now owns which absorber takes a crossed bound: a crew
whose own last status line declares an external wait or a verified captain-held
transfer takes the bounded FM_PAUSE_RESURFACE_SECS recheck, and everything else
keeps the unchanged wedge timer. The discriminator is the declaration together
with liveness (the caller has already confirmed the pane is busy), never a
blanket silencing - a crew that declared nothing, or whose pane is not live,
escalates exactly as before, and a declared pause still re-surfaces once per
long cadence so a forgotten wait cannot rot invisibly. Away mode is untouched:
the daemon owns pause triage there and already reads the same vocabulary.

The two call sites also no longer clear pause bookkeeping in the same poll the
pause cadence recorded it, which would have erased the re-surface throttle and
turned the long cadence back into a per-poll re-surface.

Tests: a three-phase regression fixture pins the absorbed pause, its long-cadence
recheck, and the restored wedge escalation once the declaration is lifted on the
same busy over-age pane.

Also de-flakes tests/fm-watch-triage.test.sh, which failed spuriously on a loaded
machine: fixed liveness budgets were reaping watchers mid-startup, so assertions
on post-poll state passed vacuously or failed spuriously. Waits that describe a
poll's outcome now wait for a completed poll cycle via the liveness beacon, the
heartbeat test waits for the heartbeat it asserts on, and every wait_for_exit
budget is the uniform 10s already used elsewhere in the file.

* no-mistakes(review): Fail poll-cycle waits on timeout

* no-mistakes(review): Prevent poll timeout test hangs

* no-mistakes(document): Clarify paused busy-pane supervision

* doc: Enhance communication guidelines for decision-making

Added guidelines for decision communication to the captain.

* doc: Update task delegation and communication guidelines

Clarify communication protocols with crewmates regarding task delegation and reporting.

* fix(bin): reliably confirm herdr steer submission (#2647)

* fix(herdr): confirm local steers that native agent-state misses

Herdr can leave agent_status idle for a landed Claude turn and can keep
queued Enter text visible while busy, so fm-send was reporting false
swallows. Confirm those cases through the shared queued-Enter verdict
and a cleared composer, and keep a genuine idle pending composer as
unconfirmed.

* no-mistakes(review): Stop Herdr Enter retries on unreadable composers

* no-mistakes(review): Reject queued delivery when all Herdr Enter sends fail

* no-mistakes(review): Prevent confirmation after failed Herdr Enter

* no-mistakes(review): Pace Herdr retries and clarify submit fallback

* no-mistakes(review): Align Herdr submit docs with idle fallback

* no-mistakes(document): Correct Herdr submit-confirmation documentation

* feat(bearings): add interactive Lavish fleet board (#2659)

* feat(bin): accept any-origin decision bindings with full-identity keys

An aggregation surface (the bearings board) carries captain answers for holds
across origins, but a binding was one-origin-per-source and the Lavish adapter
capped question keys at 64 chars while real full hold identities measure 69-81.

- fm-decision-hold.sh: bind <source-id> --any-origin records the (any) marker;
  binding prints it verbatim and answers accepts it, so the runner's feed seam
  carries an any-origin source with no runner change. In any-origin mode each
  key is a full hold identity <origin>-decision-<key>, split at its first
  -decision-; a key with no separator (merge/dispatch instructions) is skipped
  and feeds nothing, keeping non-decision answers out of the hold ledger by
  construction. Every existing close guard applies unchanged.
- fm-procevent-lavish.sh: raise the question-key cap 64 -> 128 so a full hold
  identity fits; the slug-shape security property is unchanged.
- tests: cross-origin closure through the real runner seam, an 81-char
  identity through the adapter, cap and shape refusals, routed-work skips,
  nonexistent-identity skips, and idempotent replay.

* feat(bearings): add the /bearings lavish interactive fleet board

/bearings lavish renders the bearings snapshot onto a shipped, reusable board
template and arms it as a Lavish process-event source, so the captain answers
Captain's Call items on the board and firstmate is woken by an ordinary check
wake - no conversational turn ever blocks on a poll.

- .agents/skills/bearings/assets/board-template.html: the shipped template
  (myfirstmate design system inlined, one fm-bearings-board.v1 JSON slot,
  fail-closed schema guard that renders an error card instead of an empty
  fleet). Per-invocation agent work is composing the payload only.
- bin/fm-bearings-board.sh: build/refresh owner - fail-closed payload
  validation, slot injection with a round-trip check and \u003c escaping,
  stable board path, any-origin bind ALWAYS before arm, arm-if-absent.
- bearings SKILL.md: the lavish invocation option, board composition rules,
  board-wake handling, and the captain-ruled merge-click authorization with
  its mandatory safeguards (PR resolved from the task's own meta record,
  wake-time green re-verification, never a red or changed PR, merges only
  through bin/fm-pr-merge.sh, chat echo with the full PR URL).
- process-event-sources SKILL.md: one-line board-wake routing trigger.
- tests: payload refusals, injection round-trip, bind-before-arm, idempotent
  re-arm, and template slot integrity.

Fleet pickup: homes receive this after merge plus a firstmate self-update;
landing timing is coordinated with the main firstmate.

* no-mistakes(review): Harden bearings board validation and wake handling

* no-mistakes(review): Require HTTPS for bearings board PR links

* no-mistakes(review): Fail closed and bound bearings board answers

* no-mistakes(review): Enforce UTF-8 byte limits for board answers

* no-mistakes(review): Serve bearings board before arming and reject empty actions

* no-mistakes(review): Prove bind-before-arm ordering through live answer consumption

* no-mistakes(document): Document bearings board and cross-origin answers

* fix(bearings): restore decision options and add close controls (#2707)

* fix(bearings): always show decision options and a close/drop control

Freeform-only Captain's Call cards hid the option buttons the board was designed around, and there was no way to drop a stale hold without inventing an answer. Require selectable options, keep freeform as a supplement, and route the reserved __drop__ answer through decline so the hold leaves Captain's Call.

* no-mistakes(review): Fix drop closure and decision-only option validation

* no-mistakes(review): Preserve answerability for non-decision cards

* no-mistakes(document): Clarify decision drop documentation

* ci: require no-mistakes pipeline step attestation (#2710)

Signature-only PRs can hide skipped review, test, or document steps. Fail unless no-mistakes >= 1.46.0 attests those three steps completed.

* feat: collapse decisions into tasks held for the captain (#2728)

* feat(captain-hold): collapse the decisions concept into tasks held for the captain

A decision is no longer a separate type: it is an ordinary backlog task held
for the captain, identified by its task id. bin/fm-captain-hold.sh owns the
surviving behaviors - guarded hold creation, the recorded-answer close
(answer/answers with a release mode for captain-gated work), the source
bindings, and the investigation completion gate - and bin/fm-decision-hold.sh
becomes a one-release compatibility shim over it.

The fleet snapshot now parses hold-until and computes captain_actionable as
queued + captain-held + unblocked + due, independent of row kind, plus a
presentation-only deferred_marker for prose-deferred rows. Bearings renders
every due captain-held task in Captain's Call, date-deferred holds as dated
Charted Next gates, suppresses prose-deferred rows from default views with an
omitted disclosure, and excludes from Recently Landed anything that closed
while still held for the captain.

Legacy compatibility: pre-collapse <origin>-decision-<key> rows are already
plain task ids and keep working; short keys in recorded metadata, concrete
origin bindings, chat --resolve-key fallbacks, and old resolution records all
resolve in place.

* no-mistakes(review): Fix captain answer replay and body preservation

* no-mistakes(review): Fix captain hold idempotency and legacy replay

* no-mistakes(review): Validate card close modes and compatibility routing

* no-mistakes(review): Enforce release replay mode matching

* no-mistakes(review): Prevent duplicate decision cards and released replay mismatches

* no-mistakes(review): Preserve answer columns and legacy resolve replays

* no-mistakes(document): Document strict replay and legacy compatibility

* no-mistakes(lint): Quote done literals to satisfy ShellCheck

* no-mistakes: apply CI fixes

* fix(rebase): keep collapsed captain hold board semantics

* fix: bound recovery announcements and preserve supervision (#2733)

* fix(watch): announce recovery once per generation and keep successors supervising

A lost Pi/OpenCode handling handshake re-announced the same recovery
generation on every cycle and spent the successor's first ~55s blind, so
a real crew event could be ignored and then dropped. Record the
announcement in the durable marker, confirm the handshake before the
follow-up without swallowing failure, and enter the poll loop immediately.

* no-mistakes(review): Tighten recovery event timing regression

* no-mistakes(document): Document recovery-loop supervision guarantees

* fix(bin): surface captain-call record divergence (#2744)

* fix(bin): signal a captain call resolved in the log but still held

A captain call has two records and closing one has never closed the
other: a `resolved [key=...]` line closes the status-log fold, while the
backlog task held for the captain closes only through
`fm-captain-hold.sh answer`. Answering on the status side alone left no
trace of the disagreement - the fold went quiet, the durable record kept
saying the captain owed an answer, and nothing warned. The defect was
never the separation; it was the silence.

Add `fm-captain-hold.sh diverged`, a read-only report of that
contradiction, and print it from `fm-wake-drain.sh` as a bounded RECORD
DIVERGENCE section beside OPEN DECISIONS on every drain. It flags one
condition: a task still open and still carrying the captain-hold
annotations whose key was closed on the status side by the resolve verb,
under the collapsed identity or the legacy derived one.

It closes nothing, ever. A captain call closed wrongly leaves review
entirely, which is worse than the noise, so both reconciliation
directions stay human-owned and the printed hint names both - a
resolution is not proof the captain ruled, since a call can dissolve on a
false premise or turn out to have been a question of fact.

Three states are deliberately not divergence: a `captain-held` close is
the verified transfer `complete` writes, a still-open keyed decision
belongs to the OPEN DECISIONS fold, and a captain call with no routed
work item is legitimate rather than incomplete, so routed work is no part
of the test.

`fm-classify-lib.sh` gains `status_key_closing_verb`, which reports how
the status side currently reads one key by replaying the existing
`_fm_decision_fold_line` rule rather than re-deriving it, so the two
closing verbs stay distinguishable in one place. The per-wake cost is one
`tasks-axi list`, one key scan per status log, and the precise per-key
fold only for a key that already names a still-open task; the call is
hard-bounded so a slow backlog tool can never delay wake presentation.

* fix(document): Correct divergence lifecycle documentation

* fix(document): Neutralize divergence lifecycle prose

* fix(bin): re-arm after an abandoned auto-arm claim and defer a wedge escalation while a worktree is written (#2524)

* fix(watch): re-arm supervision after an abandoned auto-arm claim

A Claude auto-arm cycle that armed, delivered one rewake, and exited left
its single-flight lock behind. Both Stop-event participants then deferred
to that lock forever, because its recorded pid was still live: the
turn-end guard read it as recovery under way and allowed the stop, and the
next Stop firing treated it as another owner and declined to arm. On
2026-08-14 a home with two tasks in flight lost supervision for about 40
minutes with no watcher process and no watcher lock, its beacon frozen at
the one delivery, and both crewmates' finished reports sat in the durable
queue until an operator drained it by hand.

Abandonment is now proven from the epoch ledger instead of inferred from
pid liveness. A lock whose holder pid matches the ledger's own owner_pid
while the recorded outcome is anything other than arming has already
finished its decision, so that claim is reclaimed under the lock's steal
mutex, stops counting as recovery ownership in the guard, and is cleared
by the guard's terminal check rather than deferred to. A failed clear
re-blocks instead of allowing a blind stop, and an arming entry stays in
flight however old it is, because its owner foregrounds the arm for the
whole watcher cycle.

Issue #2251's PR #2263 does not cover this failure. It is closed and
unmerged, lives entirely in bin/fm-watch-arm.sh, and retires the stalled
watcher and matching stale watcher lock of an arm that is currently
running. Here no arm and no watcher were running and no watcher lock
existed, so it has nothing to retire and the home stays blind.

tests/fm-claude-stop-autoarm.test.sh covers the reclaim, the still-arming
and unnamed-owner cases that must keep the gate closed, and the failed
clear. tests/fm-turnend-guard.test.sh covers the guard side of the same
boundary. Both fail without this change.

* fix(watch): defer a wedge escalation while the task worktree is written

The wedge detector had two inputs, rendered pane quietness and the run
step, and neither can see a crew that is writing source, then tests, then
documentation behind a static pane. On 2026-08-14 one crewmate produced
eight consecutive possible-wedge escalations in a single afternoon, three
of them demanding deep inspection, while it was demonstrably working and
then committed. Every one of them cost a supervision turn to disprove by
hand.

Add write activity inside the crew's own recorded worktree as a third
liveness input. crew_worktree_written_since compares the worktree against
the caller's existing idle-window timer file, so -newer needs no clock
arithmetic, no temp file, and no portable mtime write. The probe runs only
inside the branch that was about to escalate, which bounds it to one
pruned, depth-bounded walk per window per FM_STALE_ESCALATE_SECS and
leaves the per-poll stale sweep exactly as cheap as before.

Positive evidence defers rather than cancels. The idle timer restarts so
the next window probes again, the escalation counter is neither advanced
nor reset so a later genuine wedge keeps the demand-deep-inspection
history it earned, and a .writing-since marker ages the whole deferral
chain so the pane still re-surfaces once per FM_PAUSE_RESURFACE_SECS,
through the same throttle shape a declared pause already uses, labeled as
a recheck rather than a wedge. This can only reduce false positives: every
absence of evidence, including no recorded worktree, a torn-down worktree,
a missing anchor, and a failed walk, falls through to the unchanged
escalation schedule, so a crew that writes nothing still escalates on the
existing timetable.

What the signal cannot see, by design or by construction:

- CPU burn with no writes, such as a long compaction, is invisible. That
  case keeps the old behavior exactly.
- A commit-only phase writes only .git, which is pruned first so that
  firstmate's own read-only git commands against the worktree can never
  make the probe self-fulfilling.
- Writes under the pruned generated trees, or deeper than
  FM_WORKTREE_WRITE_MAXDEPTH, do not count.
- The probe cannot attribute a write to the crew, so a background build or
  another process touching the tree looks the same. The hourly re-surface
  is what bounds that, and a churny file cannot buy silence.
- The away-mode daemon's own escalation path is deliberately untouched.

tests/fm-watch-triage.test.sh covers the classifier including the .git
prune, both halves of the live case on one fixture (quiet plus writing
defers, quiet plus silent still escalates and counts), and the bounded
re-surface. All three fail without this change.

* no-mistakes(review): prove autoarm claims by identity; skip mate-home write probe

* no-mistakes(document): document away-mode wedge boundary and probe filesystem limit

* no-mistakes(document): qualify turn-end recovery condition for abandoned auto-arm claims

* fix(watch): keep a write deferral scoped to its own idle window

Two consistency gaps in the worktree write probe, both found while reviewing
the wedge-deferral change on this branch.

A write deferral is a bounded chain: its .writing-since marker ages the whole
chain so a churning worktree still re-surfaces once per resurface window. That
is only sound while the chain belongs to the current quiet stretch, so every
path that restarts the idle-window timer has to drop it too. Two did not: the
corrupt-timer repair in wedge_timer_check, and both first-sight branches for a
captain-relevant status. A chain left over from an earlier quiet stretch made
the first deferral of the new window re-surface immediately instead of after a
full fresh window.

FM_WORKTREE_WRITE_PRUNE is a skip list, so clearing it reads as "skip nothing"
and is the obvious way to widen the probe to the whole depth-bounded tree.
Instead an empty list reported no evidence at all, quietly costing the wedge
detector its third liveness input on a home that meant to widen the walk. An
empty list now widens the walk, and the header says so.

Neither change alters when a stall that writes nothing escalates.

Regressions in tests/fm-watch-triage.test.sh cover all three paths and each
one fails on the pre-fix code.

* no-mistakes(review): honor an empty write-prune, bound the probe, share window_key

* no-mistakes(document): align probe knob count and guard regression-coverage ownership

* no-mistakes(lint): silence deliberate single-quote SC2016 in write-prune env test

* fix(bin): give a captain hold the same bounded pause cadence as a declared pause (#2748)

* fix(bin): give a captain hold the same bounded pause cadence as a declared pause

Two supervisors read a finished task's last status line and disagreed about which
declarations mean an idle endpoint is expected. bin/fm-inactive-reconcile.sh
suppresses its inactive-outcome scan only on `captain-held`, while the away-mode
daemon's wedge path gated deferral on `paused` alone. Both read the LAST line, so
the two verbs are mutually exclusive and no finished task waiting on a person
could satisfy both at once. Marking 11 such tasks `captain-held:` silenced the
900s outcome scan and immediately produced five possible-wedge escalations in one
batch, because the 240s wedge detector no longer saw a pause verb.

fm-classify-lib.sh's status_is_paused_or_captain_held already owns the combined
question, and bin/fm-watch.sh's ordinary-crew wedge path already asked it. This
extends that same answer to the paths still asking the narrower one:

- bin/fm-supervise-daemon.sh, all six sites, which form one subsystem and have to
  move together. classify_stale returns the pause action, reconcile_pause_tracking
  and migrate_watcher_pause_markers record and migrate the marker, and
  housekeeping defers the wedge and then re-surfaces the recheck. Changing only
  the stale-persistence gate would defer the escalation while
  reconcile_pause_tracking recorded nothing, so the wedge marker would persist and
  the sweep would `continue` past it forever: quiet, but never re-surfacing.
- bin/fm-watch.sh's secondmate stale gate, whose downstream owner
  pause_state_class already treats both declarations identically.
- bin/fm-push-transition-lib.sh's absorb, where either declaration already names
  the human the transition would report and the wait is already durably recorded.

Quieting alone would be half a fix, so the bounded re-surface had to reach a hold
too. A hold has no current-state mapping, unlike `paused`, so authoritative crew
state reports it as unknown and pause_state_class received `none`. An ordinary
crew recovers pause classification from that state through confirmed agent death,
which proves no live decision gate is being silenced. A secondmate's endpoint
liveness is deliberately never read there, because an idle mate is healthy by
design, so that confirmation is unavailable by construction and cannot be
required: without recovering the classification for a mate, every caller silenced
a held mate outright and its hold would rot invisibly. That promotion is bounded
by the declared-wait guard at the top of the function, so it can only reclassify a
task that already declared a wait and shows no positive working evidence.

Two narrow `status_is_paused` calls are deliberately left alone.
bin/fm-crew-state.sh's map_log_state is a current-state reporting contract, not a
wedge path; reporting a hold as `paused` would erase the distinction
status_key_closing_verb and fm-captain-hold.sh depend on, where a `captain-held`
close is a verified durable transfer and a `resolved` close claims outright
settlement. fm-classify-lib.sh's call inside status_is_captain_relevant needs no
change because that function's own case list already returns non-relevant for
`captain-held`.

bin/fm-inactive-reconcile.sh keeps its `captain-held` suppression as it is. Its
guard exists because a finished task's crew state still reports done from a
higher-priority source than the log, and a declared pause needs no such guard: the
scan only reports done or failed, and nothing else reaches its record path.
Widening it would change a separate subsystem's reporting contract, which this
defect does not require.

Coverage extends the existing colocated patterns for these predicates and asserts
both halves. tests/fm-daemon.test.sh covers the classification, the wedge marker
converting to pause tracking with no escalation, the bounded re-surface with its
window reset, and the boundary case where an answered hold stops claiming the
cadence. tests/fm-watch-triage.test.sh covers a held secondmate re-surfacing on
the same bounded cadence without being labeled a wedge.
tests/fm-supervision-events.test.sh covers the absorbed push transition. Every one
of these fails on the pre-fix code except the answered-hold boundary case, which
is there to pin that the quieting was not widened too far.

The `paused:` workaround appended to those 11 tasks is live supervision state and
is untouched here. It can be retired once this lands.

* no-mistakes(review): name the captain in a held task's bounded recheck

* no-mistakes(document): extend declared-wait supervision docs to captain-held holds

* fix(bin): make lint prerequisites and harness tests reliable (#2758)

* fix(lint): name the installer when ShellCheck or actionlint is missing

A missing actionlint exited 127 like a bare command-not-found. Fail with
exit 1 and point at the pinned installer, matching the missing-ShellCheck
path, without weakening the version pin.

* test: isolate kimi and muse detection from inherited Cursor markers

Harness detection checks CURSOR_AGENT before ancestry, so these
markerless-adapter cases failed when the suite itself ran under Cursor.
Clear the verified markers the same way the secondmate harness tests already do.

* no-mistakes(document): Document Muse Cursor marker cleanup

* feat(bin): report watched tooling updates that are available or installed but inert (#2684)

* feat(checks): report tool updates that are available or installed but inert

Firstmate had no way to notice that tooling this home depends on needs an
update, and no way at all to notice the worse case: an update that installed
correctly and then did nothing.

That second case is why this exists. A tool that self-installs into
~/.local/bin while a version manager keeps its own older copy earlier on PATH
looks completely up to date to anything that asks only "is a newer version
published". On 2026-08-20 a Herdr update landed at 0.8.2 while an older 0.8.0
copy stayed earlier on PATH, so every Herdr command failed on a protocol
mismatch and firstmate could not read its own fleet.

bin/fm-tool-update-check.sh reports the two conditions separately:

  <tool> update available      a newer version exists at the update source.
  <tool> update not in effect  a newer copy is installed on this host, but
                               PATH still resolves an older one.

PATH skew is measured, never inferred. Every executable copy of a watched
command on PATH is asked for its own version and those answers are compared,
so one lookup cannot hide the skew, and a directory name is never read as a
version because a version manager's "latest" directory can hold an older
build. A copy that will not report a version is a check failure, not a pass.

The watched tools live in local, gitignored config/watched-tools.json, so
adding a tool is a config edit rather than a code change, and the file is
never propagated to another home. Update sources cover both shapes: a local
clone's commit distance from its remote branch, and a command's own version
and update announcement, including a tool like no-mistakes that prints its
version on one command and announces a new release on another.

The check prints one line when something needs attention and prints nothing
otherwise, so it rides the existing watcher state-check contract with its
trust binding instead of introducing a schedule of its own, and
state/.tool-updates keeps the same pending update from being reported on
every poll.

The check only reports. It never installs, updates, reorders PATH, touches a
version manager, or fetches into a watched repository; every git probe is
read-only.

Tests cover the skew case as a regression, and it was verified by mutation:
removing the skew report, or stopping after the first PATH hit as a single
lookup would, each make that test fail.

* no-mistakes(review): fix tool update check probe reporting, budget, and shim write

* no-mistakes(review): keep sweeps alive on broken patterns and oversized budgets

* no-mistakes(review): roll back failed arm, widen budget clamp, bound repo probe

* no-mistakes(review): guard git probes at the budget, record uncut findings

* no-mistakes(document): fix stale watched-tool report-record wording in docs and header

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

The behavior shard's watch-triage suite failed on the new worktree-write wedge
tests. Those five tests are the only ones in the file that do not use its
standard waits. They give a fixed 3 second liveness budget to the one poll that
now spawns the bounded worktree walk, and 4 seconds to an escalating watcher
where every other test in the file gives 10. On a loaded runner that poll
outlives the fixed budget, so the round is reaped before the deferral it asserts
on is recorded, and the test reports a lost deferral instead of the deferral
under test. Wait for a completed poll cycle through the file's own
wait_poll_cycle, which is what its header documents this hazard for, and use the
file's standard 100 tick exit budget.

Verified against a load that reproduces the failure: 11 of 12 runs failed
before, 8 of 8 pass after. Verified by mutation too, so the waits still prove
the behavior: removing the write deferral, and keeping a finished deferral chain
across an idle-timer repair, each still fail their test.

* fix: decouple ask-user decisions from yolo (#2764)

* fix: treat yolo as merge authority only, not ask-user finding authority

Yolo on/off was documented as also deciding no-mistakes ask-user findings, which hid firstmate's duty to judge unambiguous-toward-design findings itself. Keep every safety boundary; this is a contract clarification, not a relaxation.

* no-mistakes(document): Clarify yolo documentation ownership and merge posture

* feat(bin): add a spoken interface that answers from records and hands work over (#2767)

* feat(voice): spoken round trip on Nova Sonic 2 with a measured relay cost

Step one of the spoken interface: the laptop captures and plays audio, this
desktop holds the model session, and no AWS credential leaves the desktop.

Measured, amazon.nova-2-sonic-v1:0 in eu-north-1, end of speech to first byte
of reply audio, 6 runs each, all answered, on a question that forces a records
read:

  relay path   1.229 1.379 1.428 1.447 1.481 1.516  median 1.438
  direct       1.147 1.179 1.203 1.237 1.244 1.317  median 1.220

The relay costs about 0.22s of the median. The direct figure reproduces the
earlier survey, which is what makes it a usable control. Excluded: the
captain's own ssh round trip, microphone capture, and speaker output. This
desktop has no microphone and no speaker, so every run used audio files.

Three pieces:

  bin/fm-voice-relay.py    holds the conversation on this host
  bin/fm_voice_records.py  what a spoken answer may read, and the handover
  bin/fm-voice-client.py   the laptop end; audio devices UNVERIFIED
  bin/fm_voice_frame.py    the wire format both machines share

Real work is handed to the existing bin/fm-inbox.sh rather than a second
queueing surface, and the agent says it is handing over rather than answering
as firstmate.

Read scope: Done history and free-form note bodies are never assembled at any
scope, so the wide default cannot reach the places commercial detail
accumulates. config/voice-read-scope narrows it to counts only, and
config/voice-read-deny excludes a named item in one line. The boundary is an
executable test that widening the reader fails.

Push to talk is the default because it is cheaper and the choice is still open;
--listen open-mic is the single flip.

Two traps worth knowing: a clip with no trailing silence is never answered, and
the end of a reply is contentEnd with stopReason END_TURN, not completionEnd.
A second user turn in one session is treated as barge-in unconditionally, and
an interrupted turn that calls a tool is lost, so the session reconnects per
turn and gives up conversational memory. That is the concrete thing step three
has to solve.

* no-mistakes(review): fix voice relay credential reuse, frame validation and record parsing

* no-mistakes(review): test uplink header guard, bound unknown expiry, align state dir

* no-mistakes(review): decide deny per item, guard turn failures, bound ambient credentials

* no-mistakes(review): read account config from home, harden deny and turn failures

* no-mistakes(review): close status verb set, fix inbox help, pair data override

* no-mistakes(review): keep profile-free relay alive, unblock loop, fix dead assertion

* no-mistakes(review): hide finished pull requests, refuse open mic, keep suite offline

* no-mistakes(review): survive reader failures, release devices, fix claims

A failure while handling a model event, or while sending a tool result,
left the reader task dead with ended and turn_done clear, and close()
re-raised the stored failure on every await. One dropped stream became a
relay that could never build another session. The reader now reports the
session over in a finally whatever killed it, and close() absorbs the
task the same way it already absorbed its sends.

The laptop client releases what it already started when a later startup
step refuses, SystemExit from the handshake wait included, and names a
device refusal instead of leaking a raw PortAudio error. Whether it
releases correctly against a real device is still unverified here.

The records docstring claimed every reading was filtered to open ids.
Only the pull request count and list are; the worker count and the state
histogram cover every live runtime record, finished ids included,
because a meta file still on disk still needs tearing down.

The finished-work deny half of the suite asserted things that held with
the deny list absent. It is replaced by a deny on an open title, which
removes the row and says so while the count stays honest.

* no-mistakes(review): name reader failures, split file and device refusals

A failure inside the model reader released the waiting turn and told
nobody. The session was not marked spent, no notice reached the client,
and the client waits for a reply end or a notice, so the captain got
their whole timeout of silence and then a record saying the turn went
unanswered with nothing about why. Both ends of the relay now name a
failed turn through one function, once per turn, and --self-test carries
the cause in relay_error the way the client's own record does.

Two things that are not failures stay that way. A stream that simply
ends is the end of a session, which serve still reads on its own terms.
A stream that goes away because close() asked it to is an ordinary
renew, and announcing it would have put a failure notice in front of the
captain on every turn.

On the laptop end, the refusal that became a device error covered the
file-backed playback and capture too, so a mistyped --in-file was
reported as an audio device failure and the advice named the flag that
had just failed. The file ends now report the path and the flag that
chose it and stay an OSError; the device ends keep the device advice and
name the flag for that end. The device paths remain unrun here, so only
the file halves are covered by a test.

* no-mistakes(test): survive model session end, order client turn frames

* no-mistakes(document): sync voice relay docs with reviewed relay behavior

* no-mistakes(document): re-measure relay latency and correct its cause

* no-mistakes(document): correct measurement date and name the unmeasured SSH hop

* no-mistakes(document): describe the unpublished control measurement, fix list formatting

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* fix(bin): preserve Relay follow-up loops until explicit disposition (#2763)

* fix: keep Relay public loops open until retire

Delivering a promised-final reply was deleting the only record that tied a public thread to later work, so a follow-on ship silently owed no closing reply. Retain the registration after delivery, rechain follow-on work onto the same thread, and make retire --reason the only close.

* no-mistakes(review): Propagate public follow-up registration removal failures

* no-mistakes(review): Persist retire receipts and align parent resolution

* no-mistakes(review): Make rechain resumable after partial obligation creation

* no-mistakes(review): Repair follow-up state, briefs, and expiry escalation

* no-mistakes(review): Serialize follow-up delivery stamps with retirement

* no-mistakes(review): Serialize rechain claims and protect registration terminal states

* no-mistakes(review): Avoid reporting retired delivery loops as open

* no-mistakes(document): Refresh public-loop documentation and verification evidence

* no-mistakes: apply CI fixes

* no-mistakes(review): Preserve delivered follow-up bindings during registration replay

* no-mistakes(review): Harden public follow-up retirement and rechain races

* no-mistakes(review): Fail closed on unresolved secondmate retirement

* no-mistakes(review): Bind secondmate cleanup to its recorded canonical home

* no-mistakes(review): Fix rechain command output and expiry validation

* no-mistakes(review): Validate brief keys and warn on remote promotion

* no-mistakes(document): Document retained public follow-up loops

* no-mistakes(lint): Remove unused bounded-wait loop variable

* feat(bin): merge GitLab merge requests through the guarded PR merge path (#2779)

* feat(bin): merge GitLab merge requests through the guarded PR merge path

bin/fm-pr-lib.sh already parses a GitLab merge request URL for the watcher,
but bin/fm-pr-merge.sh refused every non-github provider, so a merge request
had to be merged by hand and got none of the recording, guards, or audit
trail a pull request gets.

The merge path now dispatches on the parsed provider. A GitHub URL keeps its
exact previous behavior. A GitLab URL is addressed through glab by the project
URL rebuilt from the parsed host and path, so a merge request on any instance
resolves and no host is hardcoded, and no merge-method flag is added because
the project's own merge method is what should apply.

A GitLab merge happens only after one live read of the merge request confirms
it is open, detailed_merge_status is mergeable, has_conflicts is false,
blocking_discussions_resolved is true, and the head pipeline succeeded at the
exact current head. Every failing condition is reported, not just the first.
The verified head is bound to the merge with glab's --sha, so a push landing
between the read and the merge fails the merge instead of landing commits
nothing verified. Recorded metadata is never the authority for any of this: a
rebase moves the head and leaves a recorded value stale, so a recorded head
that disagrees with the live one is reported rather than trusted, and the
recorded value is read before the recording step because that step drops a
GitLab head it cannot resolve.

* no-mistakes(review): reject bundled -R clusters and make tool-absence cases host-independent

* no-mistakes(test): state authorised GitHub narrowing of bundled -R guard

This branch NARROWS GitHub behaviour. The narrowing was authorised
deliberately rather than slipping in by accident, and it applies to both
providers, GitHub and GitLab alike, because a script that guards one provider
and not the other is a trap for the next reader.

What bin/fm-pr-merge.sh now refuses is extra merge arguments containing a
bundled short-option cluster that includes R, for example "-dR other/repo".
The forge CLIs expand such a cluster one character at a time, so it carries
"--repo other/repo", and that later value wins over the repository the URL
named. Before this change, "fm-pr-merge.sh <task> <github-url> -- -dR
other/repo" reached "gh-axi pr merge 12 --repo example/repo --squash -dR
other/repo" and exited 0 with pr= recorded and the merge poll armed. It now
exits 1 with "extra merge arguments must not override the repository", records
nothing, and invokes no forge merge command. Every other GitHub invocation is
byte-identical to the base commit.

Closing that hole honours the existing rule rather than departing from it. The
file header already forbids --repo and -R because the repository must come
only from the URL, so a bundled cluster carrying a repository override was
never legitimate behaviour to preserve: it was that guard being evaded.
Redirecting a merge to a repository the URL does not name is exactly what the
guard exists to prevent.

The refusal is already pinned on both paths by the existing case
test_bundled_repo_override_args_refuse_before_recording in
tests/fm-pr-merge.test.sh. On GitHub ("-dR wrong/repo") and on GitLab ("-yR
https://other.example/g/p") it asserts exit 1, the refusal wording, no pr= in
the task meta, no armed merge poll, and no forge merge command invoked, with a
control case proving a cluster that carries no repository override still
reaches the forge. No duplicate assertion was added. Both assertions were
confirmed to have teeth by narrowing the guard back to a bare -R and watching
each path fail.

This commit carries no file change: the guard and its coverage landed in
614853d, and this message exists so the pull request description states the
narrowing.

* no-mistakes(document): fix README pointer for GitLab watch and merge doc

* no-mistakes: apply CI fixes

* fix(bin): record a lost relay connection instead of an unanswered turn (#2788)

* no-mistakes: apply CI fixes

* fix(bin): drop a private record citation and narrow the review rule

Three corrections to the spoken interface that landed in #2767, plus one
fix carried over from that branch after its pull request had already been
merged.

The confidentiality fix. The module docstring of bin/fm-voice-relay.py
cited a private, gitignored fleet record by exact path and section number.
That widens what this public repository points at, and it cannot resolve
for any reader here, because the path has never been in the repository.
Both traps it pointed at are already described in full in the list
immediately below it, and docs/voice-relay.md carries the same two for
operators with no citation at all, so the pointer is removed and no claim
is weakened by losing it. Two comments that referred to "the survey" as
though it were something a reader could open are reworded the same way.
Neither exposed a path, so that half is comprehensibility rather than
confidentiality.

The review rule. .greptile/rules.md is kept, because its conditions are
right and deleting it would leave the next reviewer to re-litigate a
decision already argued out. What was wrong with it is narrower than its
existence: it read as settled repository policy, when whether VISION.md
itself should be reconciled is an open question belonging to the captain.
One sentence now says so, and says that the conditions listed below it are
what the interpretation depends on. That narrows the claim rather than
widening it.

The carried-over fix. The first commit on this branch is 7f98e797 from
fm/voice-relay-build-v4, taken verbatim rather than rewritten. It closes
the window where a transport failure was recorded and then erased, so a
run could be emitted as answered false with relay_error null. That matters
more than it looks: relay_error is the field that keeps an infrastructure
failure from being averaged into a latency figure, so the failure mode is
a dead connection wearing the costume of a slow reply. It landed fifteen
minutes after…
ironerumi added a commit to ironerumi/firstmate that referenced this pull request Aug 24, 2026
* fix(bin): prevent false Pi watcher alarms during hand-offs (#2304)

* fix(guard): stop the false send-time watcher-down alarm on Pi primaries

On a Pi primary the watcher process is not the liveness signal. The Pi
extension tears the watcher down on every actionable wake and spawns the
replacement itself, so the singleton lock is legitimately unheld between
cycles: every one of the 799 cycles in a live primary's ledger ends with
lock_after=pid:none, and a live capture caught the guard verdict flipping to
no-watcher during one hand-off with the beacon 63s old.

bin/fm-guard.sh classified Pi as a persistent-watcher harness, which demands a
live identity-matched lock holder at all times, so any guarded command landing
in a hand-off painted the full WATCHER DOWN - SUPERVISION IS OFF banner and
told firstmate to repair a cycle the extension already owns and is restoring.

Add an extension supervision model for pi and pi-signed. A live
identity-matched watcher stays the ordinary healthy state; an unheld lock is
healthy only while the beacon is fresh within grace AND a live Pi session
provably owns continuity - both primary extensions recorded in their state
markers at their current on-disk builds by the process named in state/.lock,
with that process still alive. Without that proof the banner fires exactly as
before, so an unloaded, version-drifted, or exited Pi session is loud
immediately and a cycle the extension never restores is loud once the beacon
passes grace. The queued-wake warning, the PID-strict turn-end guard, and
every other primary's detection are untouched.

Fold session-start's duplicate Pi marker predicate into the shared library so
the ownership contract has one owner.

* no-mistakes(review): Restrict Pi hand-off tolerance to unheld watcher locks

* no-mistakes(document): Document Pi watcher hand-off supervision

* feat: support Cursor Agent CLI as a primary harness (#2305)

* feat(cursor): add Cursor Agent CLI primary hooks, park supervision, and session start

Register a tracked project-scope .cursor/hooks.json for Cursor's stop,
sessionStart, preCompact, and preToolUse steps.

bin/fm-turnend-guard-cursor.sh owns Cursor's turn boundary as a park: it
foregrounds the watcher arm, holds the boundary open until an actionable close,
and returns that wake as one follow-up. Exit 2 is a silent no-op on Cursor's
stop step, so the adapter never uses it. The follow-up loop is bounded twice,
by Cursor's own loop_limit and by the payload's loop_count.

bin/fm-sessionstart-cursor.sh delivers the digest as additional_context at
sessionStart, and stages it for the next turn boundary at preCompact, which
cannot inject context.

Cursor also loads the tracked Claude settings, so bin/fm-hook-host-lib.sh lets
each tracked Claude-shaped entrypoint stand down on a Cursor-delivered payload
rather than running every covered event twice.

bin/fm-tmux-lib.sh reclassifies a Cursor pane's composer cursorlessly, because
Cursor parks its terminal cursor outside the composer, which restores a genuine
composer-empty proof and unblocks away-mode escalation delivery.

* feat(cursor): make Cursor Agent CLI a verified primary harness

Resolve Cursor in the session-lock ancestry through bin/fm-cursor-lib.sh, which
a Cursor primary needs before it can hold its own home lock, and classify its
stop-hook park under the autoarm supervision model so the mid-turn pull guard
stops reporting a healthy between-turns watcher as down.

Read a Cursor pane's composer cursorlessly on tmux, gated on Cursor's own
structural process identity, which restores a genuine composer-empty proof and
lets away-mode escalations reach a Cursor primary with no daemon change.

Lift the secondmate refusals in bin/fm-spawn.sh and bin/fm-control-lib.sh now
that the supervision protocol exists and is recorded.

Cover the whole surface with a portable regression over real processes, an
opt-in live guard against the installed cursor-agent, and dated per-harness
evidence.

* docs(cursor): record Cursor as a verified primary across the owning surfaces

Update the turn-end guard, session-start, arm-seatbelt, cd-guard, watcher
continuity, architecture, configuration, README, and harness-adapters owners,
and add dated live evidence to the supervision and runtime-backend verification
records. Correct the recorded Cursor tmux composer verdict: the cursor-anchored
read is still blind, but the composite reader is no longer unknown.

Lift the remaining remote-secondmate refusal missed in the previous commit, and
add the new libs to the existing fixtures that copy a fixed dependency list.

* refactor(cursor): name the park's stand-down condition for both its causes

Also record that Cursor's preCompact firing itself is not yet live-verified,
while the static evidence that it cannot inject context, and the staging path
that follows from it, both are.

* test: give the pretool fixtures their new dependency and one lint owner

The cd-guard fixture copies a fixed dependency list and now needs the shared
hook-host predicate. Both pretool suites also asserted cleanliness with a bare
shellcheck call, a second and weaker copy of the lint definition that
bin/fm-lint.sh owns: it omits --external-sources, so it failed the moment these
checkers sourced a shared library. They now delegate to that owner.

* test: assert the cursor secondmate contract instead of its removed refusal

A cursor secondmate now launches, so the suite asserts what its park actually
needs: --trust so the home's project hooks load at all, its own home pinned as
the workspace, and the autoarm supervision model inherited across the launch.

* no-mistakes(review): Serialize Cursor wakes and bind staged context

* no-mistakes(review): Serialize Cursor context and nag state commits

* no-mistakes(review): Enforce Cursor ceiling before staged context delivery

* no-mistakes(review): Serialize Cursor claims and staged context

* no-mistakes(review): Serialize Cursor ownership and state commits

* no-mistakes(review): Protect Cursor context across session takeover

* no-mistakes(review): Preserve Cursor context across session takeover

* no-mistakes(review): Enforce owner-keyed Cursor staged context

* no-mistakes(review): Atomically claim Cursor follow-ups and staged context

* no-mistakes(review): Defer Cursor preCompact staging and simplify supersession

* no-mistakes(review): Serialize Cursor park commits and defer preCompact

* no-mistakes(review): Stop Cursor parks after session takeover

* no-mistakes(test): Route Cursor preCompact context through stop follow-up

* no-mistakes(document): Update Cursor primary documentation

* revert(cursor): cut preCompact staging from this change

Carrying a compaction digest across two concurrently running stop hooks kept
producing races that could deliver it twice or strand it indefinitely, and
closing them kept enlarging a critical section inside a hook Cursor awaits at
the turn boundary. Native preCompact firing was never observed either, so the
surface has no empirical basis yet.

Remove the adapter, its registration, its staged path in the park, and its
tests, and record the surface as deferred and uncovered alongside the Codex
interactive TUI. A regression now asserts preCompact stays unregistered so it
cannot return without its own design and evidence.

This change ships the proven core only: the turn-end follow-up park, the
run-tier session start, and away-mode delivery.

* no-mistakes(review): Correct Cursor park supersession documentation

* no-mistakes(document): Clarify Cursor run-tier verification ownership

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

---------

Co-authored-by: kunchenguid <kun-1@kunchenguid.com>

* feat(bin): add decline and repair paths for decision holds (#2330)

* feat(bin): add unrouted close paths to the captain decision gate

A captain who declines a held decision leaves no follow-up work to route,
so `resolve` could not express that answer: it requires at least one
`--routed-to` task. The only way to close such a hold was a direct
`tasks-axi done`, which never writes the durable resolution record the
completion gate reads, so the originating investigation could no longer
pass `verify` and its cleanup stayed blocked.

Add two close paths that route no work:

- `decline` closes an actively held hold with a recorded captain decision
  and no routed task. It refuses while any task is still blocked by the
  hold, because releasing routed work without recording it is `resolve`'s
  job.
- `repair` records the missing resolution block on a hold that was already
  closed outside this script. It never reopens a hold and never clears a
  dependency edge, and it refuses a hold that is still actively held.

Both require a non-empty captain decision file and share `resolve`'s
digest-based retry identity, so an exact retry is idempotent while a
changed decision is rejected. The recorded body now also names which path
closed the hold, and each routed entry regains its own line.

The gate itself is unchanged: an unanswered decision still fails
completion and blocks teardown, and neither new path can close a hold
without the captain's recorded word.

* fix(bin): require captain-hold provenance before repairing a decision

`repair` checked only that the backlog item was kind captain and Done, so
an ordinary captain-kind task that was never held for the captain could be
closed, repaired, and then pass the completion gate.

tasks-axi keeps `hold_kind` through a close, so it is the surviving proof
that an identity really was a captain hold. Require it before writing the
resolution record, and cover the case in the gate regression.

* no-mistakes(document): Correct decision-hold lifecycle documentation

* fix(bin): surface buried wake status lines once (#2331)

* fix(bin): surface buried status notes on wake drain

A note: answer immediately followed by a routine note was dropped because
annotations kept only the newest line and note: never enters OPEN DECISIONS.
Present every unread note and pending-reply resolution since the last drain
cursor, and annotate every unread line on a queued signal.

* no-mistakes(review): Fix unread status cursor races and overflow

* no-mistakes(review): Preserve cursors when status span reads fail

* no-mistakes(review): Make status presentation transactional under I/O failures

* no-mistakes(review): Simplify unread status cursor and presentation locking

* no-mistakes(review): Align cursor failure regressions with transactional presentation

* no-mistakes(review): Retire stale presentation cursors during task teardown

* no-mistakes(review): Preserve routine status until signal annotation

* no-mistakes(review): Correct unread status cap documentation

* no-mistakes(document): Document unread wake status presentation

* no-mistakes(lint): Fix wake surfacing ShellCheck warnings

* no-mistakes: apply CI fixes

* feat: add max Calm presentation level (#2334)

* feat(calm): add a max presentation level that hides mid-turn working notes

Calm's home-local preference becomes a three-state level instead of a
boolean: "off" is stock Pi, "on" is today's Calm, and "max" is Calm plus
hiding the assistant text of messages the model did not end its response
with. `/calm max` selects it from any state, a plain `/calm` steps max
back to ordinary Calm and otherwise keeps the existing on/off cycle, and
any other argument keeps that cycle too.

`config/calm` now persists "max" as its own literal value, so a session
start, resume, fork, or reload restores the stored level rather than
treating it as unrecognized and dropping to off.

The hide rule keys on Pi's intrinsic per-message stopReason: "toolUse",
or "length" with tool calls present. Streaming ("pending") text is never
filtered, because suppressing it would also stop a genuine reply from
streaming. The existing assistant layout adapter filters the blocks out
of the same shallow presentation copy it already uses for collapsed
thinking, so the message, model context, session storage, /export, and
delivery are untouched and a hidden mid-turn row collapses to zero
height. The new "assistant-working-note" class keeps that choice in the
visibility policy owner, where ordinary Calm keeps it visible.

* no-mistakes(document): Clarify Calm max persistence and taxonomy

* feat(calm): hide mid-turn working notes by default (#2339)

* feat(calm): make hiding mid-turn working notes the ordinary Calm state

Calm collapses back to the two-state on/off toggle it was before the max
presentation level, with max's hide rule promoted into ordinary Calm.
Calm on now hides mid-turn assistant working notes in addition to what it
already hid, and the /calm command parses no argument again.

The hide rule itself is unchanged: assistant text is removed from the
shallow presentation copy when the message's own stopReason is "toolUse",
or "length" with tool calls present. Streaming ("pending") text is never
filtered, so a genuine reply still streams. The message, model context,
session storage, /export, and delivery remain untouched.

config/calm persists only "on" and "off" again, but the reader still maps
a persisted "max" to on so a home upgraded from the removed level keeps
Calm on instead of dropping to off.

The mid-turn hide is now default behavior rather than an opt-in level, so
docs/calm.md documents it for users, docs/configuration.md records the
two written values plus the legacy max mapping, and the feasibility
taxonomy drops its level-scoped wording.

* no-mistakes(document): Document ordinary Calm working-note hiding

* chore: store no-mistakes test evidence in the repo (#2355)

* chore: ignore scratchpad/ at the repo root (#2359)

* fix(ci): fail hung Herdr behavior runs in 20 minutes (#2413)

A wedged family-run step was occupying the runner until the 75-minute
job cap; bound that step so cleanup and timing artifacts still upload.

* fix: keep the public promise reachable when work is routed to a second mate (#2457)

The lightweight Relay follow-up link lives in the answering home's own
state/<task-id>.meta, so it can only bind work that home owns. When a
Relay-linked request is routed to a second mate, the task record lives in the
second mate's home, fm-x-link.sh failed with a bare "no such task ...meta", and
nothing else picked the promise up: only the soft acknowledgement was ever
posted. The typed promised-final path already supports --work-home
secondmate:<id>; the playbook simply never chose it.

- fmx-respond now states the routing rule crisply: a task in this home takes the
  lightweight link, and second-mate-routed work takes a promised-final
  commitment bound to that home, registered up front with the brief command
  carried into the routed worker's instructions.
- fm-x-link.sh refuses a task with no local record by naming the registered
  second mate whose home actually holds it and printing the promised-final
  registration command, with the exact --work-home when the match is
  unambiguous. A home with no registered second mates keeps the plain error.
- fm-backlog-handoff.sh reports, after a successful move, any moved key that
  still owes a public reply bound to main/<key>, since that binding no longer
  names the home owning the work. The move itself is never blocked.

Docs and the secondmate handoff prose follow the same rule. Tests cover the
refusal, its scoping, the unchanged local-link path, and both handoff outcomes
at the script boundary.

* docs(skills): add remote-secondmate recovery hint for false-negative verdicts (#2456)

* fix(skills): hint that remote secondmate liveness verdicts false-negative

fm-crew-state and fm-send routinely misreport a live remote secondmate
as dead; confirm against the pane before relaunching, and relaunch only
through fm-spawn.sh, never raw herdr pane surgery.

* no-mistakes: apply CI fixes

* fix(calm): keep Pi's export confirmation visible (#2461)

Pi 0.83.0 added a status line to every tool-expansion change, and Pi
updates the previous status line in place when two status messages
arrive back to back. Calm's post-export redraw cycled tool expansion on
the macrotask right after Pi printed "Session exported to: <path>", so
both expansion status lines coalesced over that confirmation and the
captain was left with no record of where their export landed.

Calm now repaints only the tool rows it presents, by invalidating each
row through the render context Pi hands its render slots, and requests
the surrounding redraw through setStatus. Neither appends to the
transcript. The repaint is still needed because Pi can re-render a row
asynchronously - the built-in edit row invalidates itself once its diff
is ready - and that re-render can land inside the window where /export
forces stock rendering.

The real-terminal /export case now asserts the confirmation is still on
screen after the redraw has settled, and that the redraw restored every
Calm-hidden row, instead of only racing the moment the confirmation
first appeared.

* feat(stow): add open-record persistence to /stow before reset (#2488)

* feat(stow): persist the open records a session is holding

/stow curated memory and captured session knowledge, but never touched
record state, while AGENTS.md called it an "unfinished-work sweep" and the
receipt declared the session "safe to reset" - wording that implied a
record-correctness guarantee stow does not make. A shipped PR with no
backlog item, a queued umbrella whose phases had merged, and four decision
holds left open after their answers shipped all survived repeated stows.

Add a bounded pass that files record state from the same volatile input the
rest of stow already uses: the open threads in context, minutes before the
reset destroys them. It creates a record for an unfiled thread and corrects
one the session knows is wrong, through the owning path, and states its
boundary as part of the contract - it never enumerates the backlog, lists
holds, or queries a forge, because it cannot be a reconciliation and must
not be read as one.

Correct the wording in AGENTS.md and the completion receipt so reset-safe
means what it actually guarantees: nothing this session knew was lost.

* no-mistakes(review): correct stow decision-hold inspection to read hold via tasks-axi

* no-mistakes(document): note /stow open-record persistence in README command catalog

* refactor(stow): state open-record persistence as principle, not procedure

The first version enumerated triggers, named commands, and prescribed an
ordered procedure. That is too rigid for an agent skill: it invites literal
execution of a checklist instead of judgment, and every enumerated example
is a way for the guidance to go stale.

Reduce it to the intent - before a reset, the important open work you are
holding in context must end up durably recorded rather than dying with the
session, filing what is unfiled and correcting what is stale - and let the
agent judge importance, the record, and the owning write path.

Keep the scope bound, since it is a decided contract and not a mechanic:
this covers the open work the session is holding, never a reconciliation of
durable records against repository or forge reality. The wording
corrections in AGENTS.md and the completion receipt are unchanged.

* fix(decisions): close decision holds at answer time via one general keyed-answer path (#2490)

* fix(decisions): close captain holds at answer time

Firstmate had two "a decision is open" ledgers with asymmetric closing
mechanics. The live status-log ledger closes atomically at answer time,
because bin/fm-send.sh --resolve-key makes answering a decision be the
act that closes it. The durable backlog hold ledger had no such coupling:
answering and recording were two separate acts, and only the first was
forced by the workflow.

That asymmetry lost four real captain decisions. Their answers were
captured durably to disk, keyed character for character by the hold
decision keys, acknowledged, and even implemented and shipped, yet the
holds stayed open for two days and the captain was asked to re-answer
decisions already on his own disk.

Give the hold ledger the same answer-time-closure property:

- bin/fm-decision-hold.sh gains an `answer` subcommand, the hold ledger's
  counterpart to --resolve-key. It shares one unrouted close
  implementation with `decline`, so it carries every existing guard - the
  captain decision file, the active-hold requirement, retry identity, and
  the refusal to release still-routed work - and differs only in the
  resolution mode it records. `decline` keeps its stronger meaning that
  the answer routes no follow-up work at all.
- bin/fm-procevent-lavish.sh wires the channel that actually carried the
  lost answers. `arm --decisions-origin` binds a deck to the origin whose
  holds it carries, `answers` reads the structured choices out of a
  captured poll result, `close-decisions` maps each key to its hold and
  closes it through the command above, and `autohandle` lets the runner
  apply that at capture time.

Safety is preserved rather than traded away. Only rows tagged `choice`
are read, so freeform captain prose cannot forge a decision key. Closure
is confined to the one bound origin. The decision text is a pure function
of the captured result, so a replayed capture is idempotent. A hold that
is absent, already closed, or still blocking routed work is skipped and
left for `resolve`, never forced. A deck armed without the binding
touches no hold at all. And autohandle deliberately never reports full
handling, because recording an answer is transcription while acting on it
is firstmate's judgement - so the check wake still reaches the handler.

fm-send --resolve-key is untouched.

* no-mistakes(document): document state/lavish-decisions binding dir in AGENTS.md state inventory

* refactor(decisions): make keyed-answer closure one general capability

The previous pass gave holds answer-time closure but built it as bespoke
Lavish wiring: the review adapter carried the source-to-origin binding,
mapped keys to hold identities, wrote decision records, decided what to
skip, and closed holds itself. That treated a review deck as a special
decision source. It is not - it is an ephemeral discussion format that
happens to carry answers.

Collapse it into ONE general capability with one owner.

bin/fm-decision-hold.sh now owns the whole of "a keyed answer closes its
matching hold":
- `answers <origin> --source <provenance>` is the channel-agnostic
  intake. It reads key/answer/label lines on stdin, maps each key to its
  hold, and closes it through the same `answer` path, so every guard
  applies identically whatever channel the answer came from. --source is
  provenance recorded in the decision, never a behavior switch; there is
  no per-channel branch and no knowledge of chat, decks, or transports.
- `bind`/`unbind`/`binding` own the source-to-origin binding for any
  channel whose answers arrive detached from their origin.

Every channel is now an ordinary caller that only turns what it received
into keyed lines:
- bin/fm-send.sh (chat) feeds the intake for a key that names an active
  hold. This also fixes a real gap: once `complete` transfers a decision
  to its hold it closes the live status copy, so --resolve-key alone
  could never answer a transferred decision.
- bin/fm-procevent.sh feeds it generically. A bound source's captured
  result goes to `<adapter> answers <result-file>` and whatever that
  prints is piped into the intake. The runner names no adapter, parses
  no result, and carries no decision rule, so any future adapter with an
  `answers` command works with no change here.
- bin/fm-procevent-lavish.sh keeps only `answers`, which reports the
  structured choices a review captured and stops. It maps nothing to a
  hold and closes nothing; it lost ~160 lines of decision logic.

Feeding is independent of handling, so it never acknowledges a result
and never suppresses a wake - recording an answer is transcription,
acting on it stays firstmate's judgement.

The regression that proves closure now drives a FIXTURE adapter that is
not the review adapter, so what is proven is that any bound channel
reaches the intake rather than that one channel is wired specially. A
new regression drives the real fm-send over a stubbed transport for the
chat side. Every prior guarantee still holds, and fm-send's status-log
behavior is unchanged.

* no-mistakes(review): test(decisions): drop source-content grep from hold-closure regression

* fix(memory): emit a real @AGENTS.md pointer instead of a CLAUDE.md symlink (#2512)

A Write aimed at CLAUDE.md followed the symlink and destroyed AGENTS.md.
The installer now creates and migrates to a recoverable two-line pointer file.

* fix(ci): keep CLAUDE.md pointer check valid (#2515)

* ci: gate GitHub workflows with pinned actionlint (#2517)

* fix(lint): catch malformed GitHub workflows before merge

A self-broken ci.yml cannot report its own breakage, so parse every
workflow in the local lint path that no-mistakes already runs.

* fix(lint): pin actionlint instead of Ruby for workflow lint

A self-broken ci.yml still has to fail in the local lint path, and the
named tool for that gate is actionlint, not a new Ruby runtime.

* no-mistakes(document): Clarify pinned workflow lint documentation

* fix: install pinned lint tools across supported platforms (#2546)

* fix: install pinned shellcheck and actionlint on macOS and linux arm64

The installers were hardcoded to linux amd64 and sha256sum, so a Mac
dev could not satisfy the refuse-on-mismatch lint gate. Select the
official per-platform archive and checksum, and fall back to shasum -a 256.

* no-mistakes(document): Document cross-platform pinned lint installers

* docs: reconcile test-evidence docs with store_in_repo: true (#2548)

.no-mistakes.yaml has set test.evidence.store_in_repo: true since #2355, but
CONTRIBUTING.md, docs/configuration.md, and docs/architecture.md still described
the old policy of keeping evidence out of the repo in a temp directory.

The current no-mistakes behavior for store_in_repo: true is to publish each run's
test evidence to the orphan no-mistakes/evidence branch and link it from the PR
body. That branch shares no history with code branches, so evidence never enters
a pushed feature branch or the default branch, and CI's tracked personal fleet
paths rule stays accurate.

Docs only. No change to .no-mistakes.yaml or any workflow.

* docs: clarify test evidence branch storage (#2549)

* docs: correct test evidence storage comment in .no-mistakes.yaml

* no-mistakes: apply CI fixes

* docs: hint that live scouts may host their own Lavish review loop (#2563)

Make that a first-class option in always-loaded instructions so firstmate does not default to mediating and tearing the scout down between iteration rounds.

* fix(bin): report remote secondmate delivery and state truthfully (#2570)

* fix(bin): report remote secondmate delivery and state truthfully

A steer to a remote secondmate crosses fm-on.sh to a host-local fm-send
leg whose unconfirmed submit read-back (verdict=pending, typically a busy
mate whose harness queues the steer) was flattened into exit 1, so the
parent printed "error: text not submitted" / "error: text not sent" and
discarded the pending-reply expectation for a steer that had actually
landed. fm-send now carries the verdict across the ssh boundary as a
documented delivered-unconfirmed exit 3: the parent reports the steer as
delivered with confirmation pending, exits 0, keeps the expectation armed
(awaiting_report), and closes --resolve-key decisions, while transport
loss (ssh 255) and real remote failures keep failing loudly with the
remote leg's stderr attached. A local unconfirmed submit now also exits 3
with an honest non-error message and still never closes a decision key.

fm-crew-state.sh and fm-peek.sh no longer read a remote mate's endpoint
through local probes (which misreported a healthy mate as "worktree gone"
/ "can't find session: remote"): both now use the true remote source over
fm-on.sh, and an unreachable or unreadable remote reads as unknown-remote,
never as gone or dead.

* no-mistakes(document): Document remote delivery and state truth

* no-mistakes: apply CI fixes

* feat: adopt spendPriority for quota dispatch (#2574)

* Adopt quota-axi 0.1.29 spendPriority-primary array dispatch.

quota-axi 0.1.29 publishes schema 5 with selection.spendPriority as the primary comparative signal and demotes derivation fields out of default --json. Rank comparable-fit candidates on that scalar, keep runway versus the completion horizon as a hard gate, and raise the compatibility floor so a pre-consolidation build cannot reach dispatch intake.

* no-mistakes(review): Correct schema fixtures and remove prescriptive selection prompts

* no-mistakes(document): Correct quota verification evidence chronology

* Collapse quota-array-dispatch onto TOON-first spendPriority ranking.

Decide from quota-axi's default TOON; keep --json as a rare defensive fallback.
Rank by spendPriority after eligibility, reasoning-class, and runway-feasibility gates, and drop the hand-computed Pareto, pace, reserve, and window-id layers.

* no-mistakes(review): Permit ambiguous JSON fallback and correct reset fixtures

* no-mistakes(review): Correct runway semantics and escalate unresolved uncertainty

* no-mistakes(document): Document TOON-first quota dispatch evidence

* docs: add GROK_BOT.md Grok Bot system prompt (#2590)

* docs: add GROK_BOT.md Grok Bot system prompt

* docs: amend GROK_BOT.md with charter report-back and delegation marker

* docs: classify GROK_BOT.md as public-product

* docs: make GROK_BOT.md the plain Grok Bot system prompt

* docs: update GROK_BOT.md nautical terms and self-improvement (#2592)

* doc: Update language in GROK_BOT.md for clarity

Refine language for clarity and consistency in instructions.

* fix(bin): preserve inactive reconciliation scan progress (#2595)

* fix(bin): guarantee inactive-reconcile scan progress under second quantization

The inactive-outcome scan computed its aggregate deadline in whole seconds,
so a 1-second budget's effective value lands anywhere in (0,1]; a scan
starting just before a wall-clock second boundary rounded its whole budget
away mid-scan and exited having visited no child, while the durable cursor
had already advanced past the never-examined child. This is the CI flake
behind tests/fm-inactive-reconcile.test.sh's 'next bounded scan did not
resume with the following child' (watcher-wake-lock family, portable
serial 2, seen on the PR #2590 run).

Every scan now visits at least its first due child with the per-child
state-read bound floored at one second, so no invocation can be a zero-work
no-op. The outer process-group kill moves to budget+1s: the scan's own
deadline enforces the budget, and the kill is a backstop for a scan wedged
in an unbounded wait instead of a racer that routinely preempts the clean
bounded exit. The wake-lock-wait test bound tracks the backstop (3s -> 4s);
the previously flaky assertion is unchanged.

* no-mistakes(document): Document inactive-reconcile deadline backstop

* doc: Revise Firstmate delegation and communication guidelines

Refactor the guidelines for Firstmate's role and delegation process, emphasizing the importance of crewmates and asynchronous work.

* doc: Update work delegation and secret management instructions

Clarified guidelines for handing off work to crewmates and managing secrets.

* test(procevent): make the process-event suite's detached-runner assertions deterministic (#2617)

Three assertions in tests/fm-procevent.test.sh depended on a detached runner
having finished work that the command starting it does not wait for.

reconcile's replacement runner is started through detach_runner, which only
forks: reconcile returns and counts the start before that runner has claimed
its source or exec'd its child. Any assertion taken straight after reconcile
therefore samples a race.

- The publish-before-apply recovery section left its always-ready /bin/echo
  source registered across the recovery reconcile, so that reconcile launched
  a competing detached poll (observed: started=1) that then raced every later
  assertion for the source claim, the next capture sequence, and this home's
  applied record, and outlived the section holding a live claim. It is now
  retired before that reconcile - re-announcement is proven from the durable
  inbox alone and needs no registration - and started=0 is asserted so a
  competing poll cannot be reintroduced unnoticed. This is the same
  retire-before-reconcile discipline the self-announcing section already
  carries; that section acquired it after the identical race made its
  "not-autohandled: self-src" assertion read "already owned: self-src".

- The crashed-leader replacement section snapshotted the replacement's claim
  file and execution log behind a fixed 0.5s settle window. On a loaded
  machine that window expires first, which is the CI flake behind "a
  replacement runner started without recording its own claim" and "reconcile
  did not start exactly one replacement source". Both effects are now waited
  for with the suite's bounded wait helpers; the exact one-replacement count
  is still asserted afterwards, unchanged.

- The duplicate-start section slept 0.5s for reconcile's runner to record
  ownership before asserting that a second start loses to it. It now waits
  for that claim.

Also tighten one assertion that could not fail as written: "autohandled:
self-src" is a substring of "not-autohandled: self-src", so the applied path
was accepted even when the runner reported the capture left for the handler.

Evidence: on the unmodified suite, 128 full runs at 6-8x concurrency produced
6 failing runs, all in the crashed-leader section. On the fixed suite, 216
full runs under the same load produced none. Reverting the self-announcing
section's retire-before-reconcile line reproduces "already owned: self-src"
on the first iteration, confirming the shared mechanism.

* fix: preserve pending replies and defer remote reposts (#2618)

* fix(bin): keep pending-reply expectations honest on both send legs

Two related asymmetries let the parent-owned secondmate reply guard drop or
nag requests it should not have.

Local delivered-unconfirmed dropped the expectation. A marked request whose
submit read-back stayed unconfirmed (verdict=pending) is the same
not-a-failure outcome the remote leg reports as delivered, but fm-send
discarded the parent's pending-reply record for it, so a request that very
likely landed stopped being tracked entirely. The record now stays armed on
its unconfirmed-delivery marker: a correlated report still resolves it, and
an unanswered one still surfaces through the library's own reconciliation.
Exit 3 and the local rule that an unconfirmed answer never closes a decision
key are unchanged.

Remote replies were nagged for a repost they did not need. A remote mate's
report reaches the parent's status log only through the asynchronous mirror
in fm-procevent-remote-reply.sh, yet the guard read an absent correlated
line as proof the mate never reported - even while the answer was still in
flight, which is the common case because the mirror's poll window is
comparable to the recovery grace. The mirror now publishes one caught-up
watermark from a quiet window, and the guard admits a missing report as
evidence only once that watermark passes the turn that should have produced
it. A genuinely missed report still gets exactly one repost, and a channel
that is behind, unarmed, or broken leaves the request durably open and
un-nagged rather than nagging blind; the mirror escalates its own continuity
failures as before.

Tests: a local unconfirmed secondmate send keeps its expectation armed and
resolvable; a mirrored correlated remote reply resolves with no repost; a
stale or absent watermark withholds the repost while a fresh one still
releases it; a quiet remote window publishes the watermark and retirement
clears it.

* no-mistakes(review): Distinguish preempted polls from quiet windows

* no-mistakes(document): Clarify remote reply channel freshness

* no-mistakes(lint): Annotate shared remote preemption exit constant

* fix(bin): honor declared pauses in busy-pane wedge checks (#2619)

* fix(watch): honor a declared pause on a busy pane's completed-turn bound

A worker that declares an external wait (`paused:`) and then blocks in one
long foreground call - a review-hosting scout parked in a single blocking
`lavish-axi poll`, a bounded watch loop, a rate-limit sleep - keeps its pane
BUSY, so the stale path that already honors declared pauses never ran for it.
The busy-pane completed-turn bound instead routed it straight into
wedge_timer_check, which re-escalated "possible wedge, escalation N" (and, past
the threshold, demand-deep-inspection) every FM_STALE_ESCALATE_SECS for as long
as the review stayed open.

busy_turn_bound_check now owns which absorber takes a crossed bound: a crew
whose own last status line declares an external wait or a verified captain-held
transfer takes the bounded FM_PAUSE_RESURFACE_SECS recheck, and everything else
keeps the unchanged wedge timer. The discriminator is the declaration together
with liveness (the caller has already confirmed the pane is busy), never a
blanket silencing - a crew that declared nothing, or whose pane is not live,
escalates exactly as before, and a declared pause still re-surfaces once per
long cadence so a forgotten wait cannot rot invisibly. Away mode is untouched:
the daemon owns pause triage there and already reads the same vocabulary.

The two call sites also no longer clear pause bookkeeping in the same poll the
pause cadence recorded it, which would have erased the re-surface throttle and
turned the long cadence back into a per-poll re-surface.

Tests: a three-phase regression fixture pins the absorbed pause, its long-cadence
recheck, and the restored wedge escalation once the declaration is lifted on the
same busy over-age pane.

Also de-flakes tests/fm-watch-triage.test.sh, which failed spuriously on a loaded
machine: fixed liveness budgets were reaping watchers mid-startup, so assertions
on post-poll state passed vacuously or failed spuriously. Waits that describe a
poll's outcome now wait for a completed poll cycle via the liveness beacon, the
heartbeat test waits for the heartbeat it asserts on, and every wait_for_exit
budget is the uniform 10s already used elsewhere in the file.

* no-mistakes(review): Fail poll-cycle waits on timeout

* no-mistakes(review): Prevent poll timeout test hangs

* no-mistakes(document): Clarify paused busy-pane supervision

* doc: Enhance communication guidelines for decision-making

Added guidelines for decision communication to the captain.

* doc: Update task delegation and communication guidelines

Clarify communication protocols with crewmates regarding task delegation and reporting.

* fix(bin): reliably confirm herdr steer submission (#2647)

* fix(herdr): confirm local steers that native agent-state misses

Herdr can leave agent_status idle for a landed Claude turn and can keep
queued Enter text visible while busy, so fm-send was reporting false
swallows. Confirm those cases through the shared queued-Enter verdict
and a cleared composer, and keep a genuine idle pending composer as
unconfirmed.

* no-mistakes(review): Stop Herdr Enter retries on unreadable composers

* no-mistakes(review): Reject queued delivery when all Herdr Enter sends fail

* no-mistakes(review): Prevent confirmation after failed Herdr Enter

* no-mistakes(review): Pace Herdr retries and clarify submit fallback

* no-mistakes(review): Align Herdr submit docs with idle fallback

* no-mistakes(document): Correct Herdr submit-confirmation documentation

* feat(bearings): add interactive Lavish fleet board (#2659)

* feat(bin): accept any-origin decision bindings with full-identity keys

An aggregation surface (the bearings board) carries captain answers for holds
across origins, but a binding was one-origin-per-source and the Lavish adapter
capped question keys at 64 chars while real full hold identities measure 69-81.

- fm-decision-hold.sh: bind <source-id> --any-origin records the (any) marker;
  binding prints it verbatim and answers accepts it, so the runner's feed seam
  carries an any-origin source with no runner change. In any-origin mode each
  key is a full hold identity <origin>-decision-<key>, split at its first
  -decision-; a key with no separator (merge/dispatch instructions) is skipped
  and feeds nothing, keeping non-decision answers out of the hold ledger by
  construction. Every existing close guard applies unchanged.
- fm-procevent-lavish.sh: raise the question-key cap 64 -> 128 so a full hold
  identity fits; the slug-shape security property is unchanged.
- tests: cross-origin closure through the real runner seam, an 81-char
  identity through the adapter, cap and shape refusals, routed-work skips,
  nonexistent-identity skips, and idempotent replay.

* feat(bearings): add the /bearings lavish interactive fleet board

/bearings lavish renders the bearings snapshot onto a shipped, reusable board
template and arms it as a Lavish process-event source, so the captain answers
Captain's Call items on the board and firstmate is woken by an ordinary check
wake - no conversational turn ever blocks on a poll.

- .agents/skills/bearings/assets/board-template.html: the shipped template
  (myfirstmate design system inlined, one fm-bearings-board.v1 JSON slot,
  fail-closed schema guard that renders an error card instead of an empty
  fleet). Per-invocation agent work is composing the payload only.
- bin/fm-bearings-board.sh: build/refresh owner - fail-closed payload
  validation, slot injection with a round-trip check and \u003c escaping,
  stable board path, any-origin bind ALWAYS before arm, arm-if-absent.
- bearings SKILL.md: the lavish invocation option, board composition rules,
  board-wake handling, and the captain-ruled merge-click authorization with
  its mandatory safeguards (PR resolved from the task's own meta record,
  wake-time green re-verification, never a red or changed PR, merges only
  through bin/fm-pr-merge.sh, chat echo with the full PR URL).
- process-event-sources SKILL.md: one-line board-wake routing trigger.
- tests: payload refusals, injection round-trip, bind-before-arm, idempotent
  re-arm, and template slot integrity.

Fleet pickup: homes receive this after merge plus a firstmate self-update;
landing timing is coordinated with the main firstmate.

* no-mistakes(review): Harden bearings board validation and wake handling

* no-mistakes(review): Require HTTPS for bearings board PR links

* no-mistakes(review): Fail closed and bound bearings board answers

* no-mistakes(review): Enforce UTF-8 byte limits for board answers

* no-mistakes(review): Serve bearings board before arming and reject empty actions

* no-mistakes(review): Prove bind-before-arm ordering through live answer consumption

* no-mistakes(document): Document bearings board and cross-origin answers

* fix(bearings): restore decision options and add close controls (#2707)

* fix(bearings): always show decision options and a close/drop control

Freeform-only Captain's Call cards hid the option buttons the board was designed around, and there was no way to drop a stale hold without inventing an answer. Require selectable options, keep freeform as a supplement, and route the reserved __drop__ answer through decline so the hold leaves Captain's Call.

* no-mistakes(review): Fix drop closure and decision-only option validation

* no-mistakes(review): Preserve answerability for non-decision cards

* no-mistakes(document): Clarify decision drop documentation

* ci: require no-mistakes pipeline step attestation (#2710)

Signature-only PRs can hide skipped review, test, or document steps. Fail unless no-mistakes >= 1.46.0 attests those three steps completed.

* feat: collapse decisions into tasks held for the captain (#2728)

* feat(captain-hold): collapse the decisions concept into tasks held for the captain

A decision is no longer a separate type: it is an ordinary backlog task held
for the captain, identified by its task id. bin/fm-captain-hold.sh owns the
surviving behaviors - guarded hold creation, the recorded-answer close
(answer/answers with a release mode for captain-gated work), the source
bindings, and the investigation completion gate - and bin/fm-decision-hold.sh
becomes a one-release compatibility shim over it.

The fleet snapshot now parses hold-until and computes captain_actionable as
queued + captain-held + unblocked + due, independent of row kind, plus a
presentation-only deferred_marker for prose-deferred rows. Bearings renders
every due captain-held task in Captain's Call, date-deferred holds as dated
Charted Next gates, suppresses prose-deferred rows from default views with an
omitted disclosure, and excludes from Recently Landed anything that closed
while still held for the captain.

Legacy compatibility: pre-collapse <origin>-decision-<key> rows are already
plain task ids and keep working; short keys in recorded metadata, concrete
origin bindings, chat --resolve-key fallbacks, and old resolution records all
resolve in place.

* no-mistakes(review): Fix captain answer replay and body preservation

* no-mistakes(review): Fix captain hold idempotency and legacy replay

* no-mistakes(review): Validate card close modes and compatibility routing

* no-mistakes(review): Enforce release replay mode matching

* no-mistakes(review): Prevent duplicate decision cards and released replay mismatches

* no-mistakes(review): Preserve answer columns and legacy resolve replays

* no-mistakes(document): Document strict replay and legacy compatibility

* no-mistakes(lint): Quote done literals to satisfy ShellCheck

* no-mistakes: apply CI fixes

* fix(rebase): keep collapsed captain hold board semantics

* fix: bound recovery announcements and preserve supervision (#2733)

* fix(watch): announce recovery once per generation and keep successors supervising

A lost Pi/OpenCode handling handshake re-announced the same recovery
generation on every cycle and spent the successor's first ~55s blind, so
a real crew event could be ignored and then dropped. Record the
announcement in the durable marker, confirm the handshake before the
follow-up without swallowing failure, and enter the poll loop immediately.

* no-mistakes(review): Tighten recovery event timing regression

* no-mistakes(document): Document recovery-loop supervision guarantees

* fix(bin): surface captain-call record divergence (#2744)

* fix(bin): signal a captain call resolved in the log but still held

A captain call has two records and closing one has never closed the
other: a `resolved [key=...]` line closes the status-log fold, while the
backlog task held for the captain closes only through
`fm-captain-hold.sh answer`. Answering on the status side alone left no
trace of the disagreement - the fold went quiet, the durable record kept
saying the captain owed an answer, and nothing warned. The defect was
never the separation; it was the silence.

Add `fm-captain-hold.sh diverged`, a read-only report of that
contradiction, and print it from `fm-wake-drain.sh` as a bounded RECORD
DIVERGENCE section beside OPEN DECISIONS on every drain. It flags one
condition: a task still open and still carrying the captain-hold
annotations whose key was closed on the status side by the resolve verb,
under the collapsed identity or the legacy derived one.

It closes nothing, ever. A captain call closed wrongly leaves review
entirely, which is worse than the noise, so both reconciliation
directions stay human-owned and the printed hint names both - a
resolution is not proof the captain ruled, since a call can dissolve on a
false premise or turn out to have been a question of fact.

Three states are deliberately not divergence: a `captain-held` close is
the verified transfer `complete` writes, a still-open keyed decision
belongs to the OPEN DECISIONS fold, and a captain call with no routed
work item is legitimate rather than incomplete, so routed work is no part
of the test.

`fm-classify-lib.sh` gains `status_key_closing_verb`, which reports how
the status side currently reads one key by replaying the existing
`_fm_decision_fold_line` rule rather than re-deriving it, so the two
closing verbs stay distinguishable in one place. The per-wake cost is one
`tasks-axi list`, one key scan per status log, and the precise per-key
fold only for a key that already names a still-open task; the call is
hard-bounded so a slow backlog tool can never delay wake presentation.

* fix(document): Correct divergence lifecycle documentation

* fix(document): Neutralize divergence lifecycle prose

* fix(bin): re-arm after an abandoned auto-arm claim and defer a wedge escalation while a worktree is written (#2524)

* fix(watch): re-arm supervision after an abandoned auto-arm claim

A Claude auto-arm cycle that armed, delivered one rewake, and exited left
its single-flight lock behind. Both Stop-event participants then deferred
to that lock forever, because its recorded pid was still live: the
turn-end guard read it as recovery under way and allowed the stop, and the
next Stop firing treated it as another owner and declined to arm. On
2026-08-14 a home with two tasks in flight lost supervision for about 40
minutes with no watcher process and no watcher lock, its beacon frozen at
the one delivery, and both crewmates' finished reports sat in the durable
queue until an operator drained it by hand.

Abandonment is now proven from the epoch ledger instead of inferred from
pid liveness. A lock whose holder pid matches the ledger's own owner_pid
while the recorded outcome is anything other than arming has already
finished its decision, so that claim is reclaimed under the lock's steal
mutex, stops counting as recovery ownership in the guard, and is cleared
by the guard's terminal check rather than deferred to. A failed clear
re-blocks instead of allowing a blind stop, and an arming entry stays in
flight however old it is, because its owner foregrounds the arm for the
whole watcher cycle.

Issue #2251's PR #2263 does not cover this failure. It is closed and
unmerged, lives entirely in bin/fm-watch-arm.sh, and retires the stalled
watcher and matching stale watcher lock of an arm that is currently
running. Here no arm and no watcher were running and no watcher lock
existed, so it has nothing to retire and the home stays blind.

tests/fm-claude-stop-autoarm.test.sh covers the reclaim, the still-arming
and unnamed-owner cases that must keep the gate closed, and the failed
clear. tests/fm-turnend-guard.test.sh covers the guard side of the same
boundary. Both fail without this change.

* fix(watch): defer a wedge escalation while the task worktree is written

The wedge detector had two inputs, rendered pane quietness and the run
step, and neither can see a crew that is writing source, then tests, then
documentation behind a static pane. On 2026-08-14 one crewmate produced
eight consecutive possible-wedge escalations in a single afternoon, three
of them demanding deep inspection, while it was demonstrably working and
then committed. Every one of them cost a supervision turn to disprove by
hand.

Add write activity inside the crew's own recorded worktree as a third
liveness input. crew_worktree_written_since compares the worktree against
the caller's existing idle-window timer file, so -newer needs no clock
arithmetic, no temp file, and no portable mtime write. The probe runs only
inside the branch that was about to escalate, which bounds it to one
pruned, depth-bounded walk per window per FM_STALE_ESCALATE_SECS and
leaves the per-poll stale sweep exactly as cheap as before.

Positive evidence defers rather than cancels. The idle timer restarts so
the next window probes again, the escalation counter is neither advanced
nor reset so a later genuine wedge keeps the demand-deep-inspection
history it earned, and a .writing-since marker ages the whole deferral
chain so the pane still re-surfaces once per FM_PAUSE_RESURFACE_SECS,
through the same throttle shape a declared pause already uses, labeled as
a recheck rather than a wedge. This can only reduce false positives: every
absence of evidence, including no recorded worktree, a torn-down worktree,
a missing anchor, and a failed walk, falls through to the unchanged
escalation schedule, so a crew that writes nothing still escalates on the
existing timetable.

What the signal cannot see, by design or by construction:

- CPU burn with no writes, such as a long compaction, is invisible. That
  case keeps the old behavior exactly.
- A commit-only phase writes only .git, which is pruned first so that
  firstmate's own read-only git commands against the worktree can never
  make the probe self-fulfilling.
- Writes under the pruned generated trees, or deeper than
  FM_WORKTREE_WRITE_MAXDEPTH, do not count.
- The probe cannot attribute a write to the crew, so a background build or
  another process touching the tree looks the same. The hourly re-surface
  is what bounds that, and a churny file cannot buy silence.
- The away-mode daemon's own escalation path is deliberately untouched.

tests/fm-watch-triage.test.sh covers the classifier including the .git
prune, both halves of the live case on one fixture (quiet plus writing
defers, quiet plus silent still escalates and counts), and the bounded
re-surface. All three fail without this change.

* no-mistakes(review): prove autoarm claims by identity; skip mate-home write probe

* no-mistakes(document): document away-mode wedge boundary and probe filesystem limit

* no-mistakes(document): qualify turn-end recovery condition for abandoned auto-arm claims

* fix(watch): keep a write deferral scoped to its own idle window

Two consistency gaps in the worktree write probe, both found while reviewing
the wedge-deferral change on this branch.

A write deferral is a bounded chain: its .writing-since marker ages the whole
chain so a churning worktree still re-surfaces once per resurface window. That
is only sound while the chain belongs to the current quiet stretch, so every
path that restarts the idle-window timer has to drop it too. Two did not: the
corrupt-timer repair in wedge_timer_check, and both first-sight branches for a
captain-relevant status. A chain left over from an earlier quiet stretch made
the first deferral of the new window re-surface immediately instead of after a
full fresh window.

FM_WORKTREE_WRITE_PRUNE is a skip list, so clearing it reads as "skip nothing"
and is the obvious way to widen the probe to the whole depth-bounded tree.
Instead an empty list reported no evidence at all, quietly costing the wedge
detector its third liveness input on a home that meant to widen the walk. An
empty list now widens the walk, and the header says so.

Neither change alters when a stall that writes nothing escalates.

Regressions in tests/fm-watch-triage.test.sh cover all three paths and each
one fails on the pre-fix code.

* no-mistakes(review): honor an empty write-prune, bound the probe, share window_key

* no-mistakes(document): align probe knob count and guard regression-coverage ownership

* no-mistakes(lint): silence deliberate single-quote SC2016 in write-prune env test

* fix(bin): give a captain hold the same bounded pause cadence as a declared pause (#2748)

* fix(bin): give a captain hold the same bounded pause cadence as a declared pause

Two supervisors read a finished task's last status line and disagreed about which
declarations mean an idle endpoint is expected. bin/fm-inactive-reconcile.sh
suppresses its inactive-outcome scan only on `captain-held`, while the away-mode
daemon's wedge path gated deferral on `paused` alone. Both read the LAST line, so
the two verbs are mutually exclusive and no finished task waiting on a person
could satisfy both at once. Marking 11 such tasks `captain-held:` silenced the
900s outcome scan and immediately produced five possible-wedge escalations in one
batch, because the 240s wedge detector no longer saw a pause verb.

fm-classify-lib.sh's status_is_paused_or_captain_held already owns the combined
question, and bin/fm-watch.sh's ordinary-crew wedge path already asked it. This
extends that same answer to the paths still asking the narrower one:

- bin/fm-supervise-daemon.sh, all six sites, which form one subsystem and have to
  move together. classify_stale returns the pause action, reconcile_pause_tracking
  and migrate_watcher_pause_markers record and migrate the marker, and
  housekeeping defers the wedge and then re-surfaces the recheck. Changing only
  the stale-persistence gate would defer the escalation while
  reconcile_pause_tracking recorded nothing, so the wedge marker would persist and
  the sweep would `continue` past it forever: quiet, but never re-surfacing.
- bin/fm-watch.sh's secondmate stale gate, whose downstream owner
  pause_state_class already treats both declarations identically.
- bin/fm-push-transition-lib.sh's absorb, where either declaration already names
  the human the transition would report and the wait is already durably recorded.

Quieting alone would be half a fix, so the bounded re-surface had to reach a hold
too. A hold has no current-state mapping, unlike `paused`, so authoritative crew
state reports it as unknown and pause_state_class received `none`. An ordinary
crew recovers pause classification from that state through confirmed agent death,
which proves no live decision gate is being silenced. A secondmate's endpoint
liveness is deliberately never read there, because an idle mate is healthy by
design, so that confirmation is unavailable by construction and cannot be
required: without recovering the classification for a mate, every caller silenced
a held mate outright and its hold would rot invisibly. That promotion is bounded
by the declared-wait guard at the top of the function, so it can only reclassify a
task that already declared a wait and shows no positive working evidence.

Two narrow `status_is_paused` calls are deliberately left alone.
bin/fm-crew-state.sh's map_log_state is a current-state reporting contract, not a
wedge path; reporting a hold as `paused` would erase the distinction
status_key_closing_verb and fm-captain-hold.sh depend on, where a `captain-held`
close is a verified durable transfer and a `resolved` close claims outright
settlement. fm-classify-lib.sh's call inside status_is_captain_relevant needs no
change because that function's own case list already returns non-relevant for
`captain-held`.

bin/fm-inactive-reconcile.sh keeps its `captain-held` suppression as it is. Its
guard exists because a finished task's crew state still reports done from a
higher-priority source than the log, and a declared pause needs no such guard: the
scan only reports done or failed, and nothing else reaches its record path.
Widening it would change a separate subsystem's reporting contract, which this
defect does not require.

Coverage extends the existing colocated patterns for these predicates and asserts
both halves. tests/fm-daemon.test.sh covers the classification, the wedge marker
converting to pause tracking with no escalation, the bounded re-surface with its
window reset, and the boundary case where an answered hold stops claiming the
cadence. tests/fm-watch-triage.test.sh covers a held secondmate re-surfacing on
the same bounded cadence without being labeled a wedge.
tests/fm-supervision-events.test.sh covers the absorbed push transition. Every one
of these fails on the pre-fix code except the answered-hold boundary case, which
is there to pin that the quieting was not widened too far.

The `paused:` workaround appended to those 11 tasks is live supervision state and
is untouched here. It can be retired once this lands.

* no-mistakes(review): name the captain in a held task's bounded recheck

* no-mistakes(document): extend declared-wait supervision docs to captain-held holds

* fix(bin): make lint prerequisites and harness tests reliable (#2758)

* fix(lint): name the installer when ShellCheck or actionlint is missing

A missing actionlint exited 127 like a bare command-not-found. Fail with
exit 1 and point at the pinned installer, matching the missing-ShellCheck
path, without weakening the version pin.

* test: isolate kimi and muse detection from inherited Cursor markers

Harness detection checks CURSOR_AGENT before ancestry, so these
markerless-adapter cases failed when the suite itself ran under Cursor.
Clear the verified markers the same way the secondmate harness tests already do.

* no-mistakes(document): Document Muse Cursor marker cleanup

* feat(bin): report watched tooling updates that are available or installed but inert (#2684)

* feat(checks): report tool updates that are available or installed but inert

Firstmate had no way to notice that tooling this home depends on needs an
update, and no way at all to notice the worse case: an update that installed
correctly and then did nothing.

That second case is why this exists. A tool that self-installs into
~/.local/bin while a version manager keeps its own older copy earlier on PATH
looks completely up to date to anything that asks only "is a newer version
published". On 2026-08-20 a Herdr update landed at 0.8.2 while an older 0.8.0
copy stayed earlier on PATH, so every Herdr command failed on a protocol
mismatch and firstmate could not read its own fleet.

bin/fm-tool-update-check.sh reports the two conditions separately:

  <tool> update available      a newer version exists at the update source.
  <tool> update not in effect  a newer copy is installed on this host, but
                               PATH still resolves an older one.

PATH skew is measured, never inferred. Every executable copy of a watched
command on PATH is asked for its own version and those answers are compared,
so one lookup cannot hide the skew, and a directory name is never read as a
version because a version manager's "latest" directory can hold an older
build. A copy that will not report a version is a check failure, not a pass.

The watched tools live in local, gitignored config/watched-tools.json, so
adding a tool is a config edit rather than a code change, and the file is
never propagated to another home. Update sources cover both shapes: a local
clone's commit distance from its remote branch, and a command's own version
and update announcement, including a tool like no-mistakes that prints its
version on one command and announces a new release on another.

The check prints one line when something needs attention and prints nothing
otherwise, so it rides the existing watcher state-check contract with its
trust binding instead of introducing a schedule of its own, and
state/.tool-updates keeps the same pending update from being reported on
every poll.

The check only reports. It never installs, updates, reorders PATH, touches a
version manager, or fetches into a watched repository; every git probe is
read-only.

Tests cover the skew case as a regression, and it was verified by mutation:
removing the skew report, or stopping after the first PATH hit as a single
lookup would, each make that test fail.

* no-mistakes(review): fix tool update check probe reporting, budget, and shim write

* no-mistakes(review): keep sweeps alive on broken patterns and oversized budgets

* no-mistakes(review): roll back failed arm, widen budget clamp, bound repo probe

* no-mistakes(review): guard git probes at the budget, record uncut findings

* no-mistakes(document): fix stale watched-tool report-record wording in docs and header

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

The behavior shard's watch-triage suite failed on the new worktree-write wedge
tests. Those five tests are the only ones in the file that do not use its
standard waits. They give a fixed 3 second liveness budget to the one poll that
now spawns the bounded worktree walk, and 4 seconds to an escalating watcher
where every other test in the file gives 10. On a loaded runner that poll
outlives the fixed budget, so the round is reaped before the deferral it asserts
on is recorded, and the test reports a lost deferral instead of the deferral
under test. Wait for a completed poll cycle through the file's own
wait_poll_cycle, which is what its header documents this hazard for, and use the
file's standard 100 tick exit budget.

Verified against a load that reproduces the failure: 11 of 12 runs failed
before, 8 of 8 pass after. Verified by mutation too, so the waits still prove
the behavior: removing the write deferral, and keeping a finished deferral chain
across an idle-timer repair, each still fail their test.

* fix: decouple ask-user decisions from yolo (#2764)

* fix: treat yolo as merge authority only, not ask-user finding authority

Yolo on/off was documented as also deciding no-mistakes ask-user findings, which hid firstmate's duty to judge unambiguous-toward-design findings itself. Keep every safety boundary; this is a contract clarification, not a relaxation.

* no-mistakes(document): Clarify yolo documentation ownership and merge posture

* feat(b…
guanchengh-lgtm added a commit to guanchengh-lgtm/firstmate that referenced this pull request Aug 25, 2026
* feat(bin): add deterministic agent lifecycle control (#1568)

* feat(bin): add deterministic agent lifecycle control

Separate firstmate's data plane from its control plane.

bin/fm-send.sh is the data plane: conversational text, always
routing-marked for a kind=secondmate target. That marking is right for a
message and wrong for a lifecycle command - a marked "/quit" arrives as
ordinary chat the agent reasons about instead of executing.

bin/fm-control.sh is the control plane: allowlisted interrupt, exit, and
transactional relaunch verbs addressed to an exact task id, with
per-harness mechanics owned by the executable bin/fm-control-lib.sh
rather than improvised in agent prose, and a verified postcondition for
every action. There is no arbitrary-text and no raw-key entry point.

relaunch runs as a transaction with a durable journal: it resolves the
profile, proves the work it must preserve is recoverable, records the
required progress note, stops the old agent, then delegates the launch
to its single owner, bin/fm-spawn.sh --relaunch, which adopts the
recorded endpoint and worktree instead of creating either. A refusal
before the stop leaves the record and instructions byte-identical; a
failure after it reports the concrete state rather than claiming an
agent that is not running. Teardown and discard stay separate and
explicit.

exit and relaunch require a backend with a recovery-grade agent-state
classifier, so zellij, orca, and cmux are refused rather than reported
as successful blind. A remotely placed secondmate is refused by name,
because its agent runs on a host where none of these postconditions can
be read.

* fix(control): resolve a recorded harness to its adapter before retiring wiring

fm-spawn arms per-task harness wiring on prefixes, because a task
launched from a raw command records that command's basename rather than
the exact adapter name. The control plane's retirement tables are keyed
by the exact adapter, so a task recorded as `grok-2` had its turn-end
token, private registry entry, and worktree hook pointer armed and never
retired - leaving a registry entry that outlived the agent that owned
it.

State the prefix rule once, in the capability owner, and resolve the
recorded value through it before every table lookup. bin/fm-send.sh's
composer-clear lookup reads the same owner instead of keeping its own
copy of which adapters need one.

* test(control): pin muse session-binding retirement across a harness switch

* no-mistakes(review): Resolve prefixed harnesses across lifecycle control verbs

* no-mistakes(review): Report interrupt delivery without fabricating cancellation state

* no-mistakes(review): Clear disabled relaunch trace context atomically

* no-mistakes(review): Clarify control interrupts and restore legacy send state

* no-mistakes(review): Refuse ambiguous relaunches and report exit delivery

* no-mistakes(review): Revalidate interrupts and accept interrupt-stopped exits

* no-mistakes(review): Lock descendant tasks before forced recursive teardown

* no-mistakes(document): Align lifecycle adapter documentation with control plane

* no-mistakes: apply CI fixes

* fix(bin): serialize fresh task publication with forced teardown

Forced secondmate teardown enumerated a home's task set, locked what it
found, then re-enumerated while removing. A fresh spawn takes only its
own per-task lock, so a record published inside that window was
invisible to the preflight and visible to the cleanup: it was
destructively processed while never lifecycle-locked.

Reproduced with real agents. A record published 0.249s after teardown
began was removed, its window closed, and its worktree returned to the
pool - while both commands reported success. A per-task lock cannot
protect a task that does not exist yet.

Add a per-home task-set lock guarding WHICH tasks a home has, as opposed
to the metadata lock guarding one task's record. Teardown takes it per
home, parent before child, before enumerating and holds it through
cleanup. A fresh spawn takes it before its own per-task locks and holds
it through publication; a relaunch is exempt, because it republishes an
existing task already covered by that task's control lock.

Either the spawn publishes first and the teardown's preflight covers it,
or the teardown owns the set and the spawn refuses. Both directions fail
closed, and both are pinned by tests that hold the lock rather than
racing on timing.

* no-mistakes(review): Serialize remote secondmate publication with forced teardown

* no-mistakes(review): Preserve remote spawn routing and state initialization

* no-mistakes(review): Serialize teardown when descendant state is absent

* no-mistakes(review): Cover symlinked descendant state refusal

* no-mistakes(document): Document task-set serialization safeguards

* no-mistakes(lint): Isolate task-set lock path resolution

* no-mistakes: apply CI fixes

* feat(stow): add tiered decaying memory management (#1984)

* feat(stow): tiered decaying memory with captain-gated offload to local excluded skills

Implement the captain-adopted /stow redesign from the v2 tiering report as
amended by the adoption decision:

- Per-entry trailing HTML-comment markers with three tiers named for their
  handling: pinned (no clock, no eviction), aging (stale after 30 days),
  perishable (stale after 7 days, mandatory checkable expiry condition).
- File-scoped defaults (captain.md and captain-shared.md pinned,
  learnings.md aging) with a self-describing legend line per file header.
- Reinforcement requires session evidence; re-reading memory never counts.
- Archive-not-delete: stale and budget-evicted entries move with provenance
  to the never-injected data/memory-archive.md; prune always means the cold
  tier, and a stale unique fact is never deleted.
- Captain-gated over-budget offload: staleness evaluated before scope, the
  sweep runs only when still over budget after decay and consolidation,
  proposals go through the receipt plus one durable captain-held backlog
  item, migration runs through the destination's normal path, and the
  memory entry leaves only once the destination is live.
- Offload destination per the adoption decision: a user-owned skill under
  .agents/skills/<freeform-name>/ excluded via the local .git/info/exclude,
  with the hard rule that stow never creates or writes a tracked skill.
- Five graduation moves, receipt verbs archived and proposed-offload, and
  the one-time non-destructive migration of unmarked legacy entries.

The public skills/stow/SKILL.md mirrors the generic parts (markers, decay,
archive exit, user-approved on-demand offload exit, migration) with no
firstmate-specific paths.

The load-bearing assumption that a git-excluded skill is still discovered
was verified empirically against Claude Code 2.1.226 (direct
.git/info/exclude scratch-repo test plus an in-repo ignored-probe test);
the dated evidence is recorded in docs/verification/stow-memory.md.

The graduation list's deletion move is deliberately narrowed to duplicates
already preserved by a stronger owner, reconciling the v2 report's retained
'deletion of a stale entry' wording with its own prune-always-archives
rule.

* no-mistakes(review): Persist legacy migration grace across stow passes

* no-mistakes(review): Enforce archival invariants and exempt default-pinned legacy entries

* no-mistakes(review): Clarify offload scope, archive placement, and marker boundaries

* no-mistakes(review): Enforce aging fallback and verify excluded skill loading

* no-mistakes(review): Fix stow decay, pinned offload, and archival safeguards

* no-mistakes(review): Preserve pinned entries, approvals, and archive provenance

* no-mistakes(review): Restrict stow mutations to editable memory files

* no-mistakes(review): Clarify skill destinations, collision checks, and migration legends

* no-mistakes(review): Resolve exclude paths for linked worktrees

* no-mistakes(review): Secure per-home excluded skill migration

* no-mistakes(test): Require explicit tier markers on new stow entries

* no-mistakes(test): Route missing shared legends to primary owner

* no-mistakes(document): Align stow documentation with tiered memory

* fix(stow): converge the pass on an over-budget home (dogfood D1-D3)

The dogfood run against a copy of the real over-budget home showed the
pass increasing the deficit from 624 to 1,107 estimated tokens and the
relief ladder provably unable to reach budget. Three skill-text fixes:

- D1: markers become single-token spellings (<!--a:DATE-->, <!--p:DATE-->,
  <!--P-->, <!--g-->), entries matching a pinned file default carry no
  marker, the per-file policy legend collapses to a one-line pointer
  naming the stow skill as the scheme owner, and marker/pointer bytes are
  explicitly counted content - roughly 76% less metadata cost on the
  dogfooded home's first installment.
- D2: the eviction rung gains a convergence precondition - total the
  eligible pool first, and when archiving all of it cannot reach budget,
  skip eviction entirely, archive nothing for budget reasons, and report
  the exempt pinned floor as the concrete inability in the final step.
- D3: budget eviction considers only dated aging entries; <!--g-->
  legacy-grace entries are ineligible until their grace cycle resolves,
  so eviction cannot cancel promised grace or invert against validation.

Public skill mirrors the D1 marker/pointer changes; D2/D3 are internal
because the public skill has no budget ladder.

* no-mistakes(test): Enforce evidence-only reinforcement during stow migration

* no-mistakes(document): Clarify stow receipt marker actions

* docs: add project vision (#1997)

* docs: add firstmate vision

* no-mistakes(test): Classify VISION.md as public product documentation

* no-mistakes(document): Restore approved one-file vision diff

* no-mistakes: apply CI fixes

* fix(spawn): force regular Pi TUI for crews (#2005)

* fix(spawn): force regular Pi TUI for crews

* no-mistakes(document): Documented Pi regular TUI launch mode

* fix(cmux): classify borderless Claude composers (#2029)

* fix(cmux): classify borderless Claude composer

* no-mistakes(review): Normalize cmux NBSP prompts across locales

* no-mistakes(document): Document cmux borderless Claude composer classification

* docs(stow): generalize read-before-write in the public stow skill (#2091)

The public installer-facing stow skill scoped its classify-then-replace
discipline to TODO/BACKLOG items only, so findings routed to a memory file
had no stated rule against a blind append or a wholesale overwrite.

Step 6 now classifies every finding against the destination's current
contents as new, duplicate, superseding, or obsolete, and states the
considered replacement each classification implies. The outcomes follow the
tiered-memory contract already in the file: an obsolete entry is refreshed,
archived, or replaced in a way that preserves its fact, a duplicate folds
into the entry that already carries it, and a superseded body worth keeping
leaves through step 7's existing exits rather than a second recovery
mechanism.

* fix: resurface durable supervision work after re-arm (#2065)

* fix(watcher): resurface durable work after downtime

* no-mistakes(review): Make watcher rearm recovery durable and cursor-safe

* no-mistakes(review): Persist safe recovery markers across migration lock recovery

* no-mistakes(review): Retain stale lock when recovery marker publication fails

* no-mistakes(review): Preserve delivery-gap recovery and quarantine malformed markers

* no-mistakes(review): Serialize recovery consumption and report acknowledgment failures

* no-mistakes(review): Centralize recovery publication before clearing watcher evidence

* no-mistakes(review): Guarantee recovery evidence across queue and lock handoffs

* no-mistakes(review): Publish recovery evidence before durable wake commits

* no-mistakes(review): Replace recovery marker Perl dependency with Node

* no-mistakes(review): Keep interrupted wakes durable until handling acknowledgment

* no-mistakes(review): Add post-handling durable wake acknowledgements

* no-mistakes(review): Enforce post-handling acknowledgement across recovery and AFK return

* no-mistakes(review): Bind wake acknowledgements to recovery generations

* no-mistakes(review): Align wake regressions with generation-bound acknowledgements

* no-mistakes(document): Document durable re-arm recovery semantics

* no-mistakes(lint): Resolve ShellCheck warnings in recovery and watcher tests

* no-mistakes: apply CI fixes

* test(watcher): assert post-handling wake replay

* no-mistakes(review): Prevent successor loops and adopt legacy wake generations

* no-mistakes(review): Rearm durable wakes without recursive successor recovery

* no-mistakes(review): Align recovery tests with handling marker state

* no-mistakes(review): Delay handling transition until successor launch is established

* no-mistakes(review): Confirm wake handling only after successful prompt delivery

* no-mistakes(review): Acknowledge AFK wakes only after evidence publication

* no-mistakes(review): Prevent AFK wake loss before post-handling acknowledgement

* no-mistakes(document): Document durable wake acknowledgement semantics

* no-mistakes(lint): Suppress false positive for recovery action output

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* ci: measure Herdr automation on Windows runners (#2100)

* ci: add Windows Herdr automation spike

* ci: run Windows spike on its pull request

* fix: wait for Windows Herdr command output

* fix: run ANSI probe in pane shell

* ci: keep Windows Herdr spike manually triggered

* docs: clarify Windows Herdr spike verdict

* feat(ahoy): guide captains through open decisions (#2099)

* Add guided ahoy decision flow

* no-mistakes(document): Document guided Ahoy decision flow

* fix(stow): enforce startup-memory budget decisions (#2110)

* Harden stow memory budget policy

* Refine internal stow offload policy

* no-mistakes(review): Enforce shared-budget decisions and autonomous offload

* fix(spawn): refresh pooled worktrees from origin before launch (#2116)

* fix(spawn): refresh pooled worktree base

* no-mistakes(document): Document spawn base-freshness invariant

* no-mistakes: apply CI fixes

* fix(composer): unify safe classification across backends (#2102)

* refactor(composer): one shape owner behind thin capture adapters, whole matrix fixed

Consolidate every composer shape - bordered boxes (all families, geometry,
titled bottom borders), bare agent-glyph rows and their wrap regions,
opencode's left bar, and pi's identity-gated separator pair - into
fm_composer_classify_screen in bin/fm-composer-lib.sh. Adapters now
contribute only a capture and a declarative capability descriptor
(styled/cursor/identity/rows); capability differences change how confidently
a shape is judged, never what the shapes are, so a new harness shape is
teachable in exactly one place.

Correctness fixes landed as part of the consolidation (audit
data/fm-composer-consolidation-audit-s1):
- locale-safe Unicode-space normalization in the shared owner (closes the
  fleet-wide half of #1988; cmux's local byte-exact NBSP case deleted;
  naming converges with PR #1995's normalization primitive)
- muse's bare glyph joins the shared set, unbreaking muse on herdr/cmux/orca
- orca learns the borderless bare shape, drops its backward-paged composer
  window, and can no longer classify a stale startup banner as the composer
- tmux tolerates a titled bottom border, unbreaking grok steering
- the left-bar shape makes opencode readable on every backend
- zellij gets a real classifier through dump-screen --ansi, replacing the
  content-diff submit heuristic that could confirm an undelivered message
  and close a --resolve-key decision (the fleet's only false positive)
- fm-spawn's kimi launch-readiness regex (the fourth shape copy) now routes
  through the shared classifier

The strict blank-row posture applies fleet-wide (captain decision
blank-row-injection-posture): no positive container proof = unknown = defer,
replacing tmux's permissive blank-cursor-row rule. Away-mode injection was
re-validated end to end on real tmux (defer on partial input and unproven
rows, clean delivery with swallowed-Enter retry into proven-empty
composers). The tmux submit core gains a baseline-idle turn-started
conversion so pi steering stays confirmed while its working screen hides
the composer; busy conversion without that baseline remains forbidden.

Plain-capture backends now degrade a glyph row carrying trailing text to
unknown instead of a false pending, per the approved capability rule.

Portable regressions pin the full byte-capture matrix from the audit under
a UTF-8 locale and LC_ALL=C, the strict-vs-permissive divergence, and
deliberate signal separation; the opt-in live guard
(tests/fm-composer-matrix-live-e2e.test.sh) verified every installed
harness against the real classifier, recorded in
docs/verification/runtime-backends.md.

* no-mistakes(review): Fix Pi glyph ambiguity and complete profile matrix

* no-mistakes(review): Preserve bare verdict when Pi identity probe is absent

* no-mistakes(review): Harden composer structure and titled-border geometry

* no-mistakes(review): Require proven idle baseline and strict Zellij guard

* no-mistakes(review): Reject box bottom borders as composer input rows

* no-mistakes(review): Prove Zellij probe typing before classifier retries

* no-mistakes(review): Preserve Pi identity uncertainty and scan full left-bar drafts

* no-mistakes(review): Verify Zellij text lands before submitting

* no-mistakes(review): Scope Zellij typing verification to selected composer content

* no-mistakes(review): Verify Zellij pastes through composer-scoped content deltas

* no-mistakes(review): Prove wrapped bare Zellij pastes through composer extraction

* no-mistakes(review): Invalidate stale cursorless composers below dead shell prompts

* no-mistakes(review): Handle shell prompt placeholders in composer extraction

* no-mistakes(review): Classify cursorless bare continuation regions safely

* no-mistakes(review): Reject stale cursorless containers below live activity

* no-mistakes(review): Preserve prompt glyphs in wrapped Zellij pastes

* no-mistakes(review): Reject live shell rows during composer extraction

* no-mistakes(review): Preserve wrapped glyph continuations through submit retries

* no-mistakes(review): Scope idle placeholders to proven positions

* no-mistakes(review): Restore boxed placeholders and live prompt reanchoring

* no-mistakes(review): Fix Zellij placeholder and wrapped glyph paste proof

* no-mistakes(document): Align composer architecture documentation

* no-mistakes(lint): Fix ShellCheck warnings in composer refactor

* no-mistakes: apply CI fixes

* docs(verification): record the trusted-checkout live matrix rerun

The pipeline's isolated gate worktree is untrusted, so claude, grok, and
muse stopped at first-launch trust dialogs there (the guard refuses to
confirm them by design). This rerun from the trusted checkout at the final
validated head verified all six installed harnesses, the strict blank-row
deferral, and the hardened zellij false-positive probe live.

* no-mistakes(document): Align composer verification evidence

* no-mistakes: apply CI fixes

* no-mistakes(review): Restore proven box bottom-cursor classification

* no-mistakes(review): Preserve styled placeholder-like drafts as pending

* no-mistakes(document): Align composer safety and Zellij delivery documentation

* no-mistakes: apply CI fixes

* docs(verification): refresh the live matrix with the final-head trusted rerun

The post-validation rerun from the trusted checkout verified all six
installed harnesses at the branch's final head, including Claude 2.1.227
(auto-updated since the audit's captures) and Grok, which the untrusted
gate worktree could not verify past their first-launch trust dialogs.

* fix(spawn): gate Pi TUI mode by CLI capability (#2117)

* fix(spawn): gate Pi regular TUI flag by capability

* no-mistakes(review): Document conditional Pi TUI capability detection

* no-mistakes(review): Pin Pi probing and launch to one executable

* no-mistakes(review): Preserve literal pinned Pi paths and update documentation

* no-mistakes(review): Defer pinned Pi path insertion until final substitution

* no-mistakes(document): Document version-safe Pi launch probing

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* docs(vision): elevate experience, pain narrative, and distro virtues (#2147)

* docs(vision): elevate experience, pain narrative, and distro virtues

Fold the captain's public vision framing into VISION.md: peace of mind as a
primary goal, multi-session context-switch pain as the problem one interface
solves, clone-and-run setup ease, self-evolution including community, and
explicit harness/backend orthogonality. Reconcile experience-as-garnish into
experience-as-purpose and update aligns/resists accordingly.

* docs(vision): state the experience goal positively

Drop the negative "not a smart workflow / useful tool / impressive technology"
pretext. Lead straight into the positive experience north star.

* feat(bin): reconcile inactive terminal crew outcomes (#2167)

* fix: reconcile inactive terminal outcomes

* fix: stream secondmate summary inputs

* no-mistakes(review): Fix reconciliation locking and request delivery retries

* no-mistakes(review): Prevent retries after unknown request delivery

* no-mistakes(document): Clarify inactive reconciliation cadence and receipts

* no-mistakes(lint): Quote terminal status arguments in reconciliation tests

* refactor: simplify inactive outcome reconciliation

* no-mistakes(review): Bound inactive reconciliation scans with durable progress

* no-mistakes(review): Bound reconciliation and deduplicate recovery notices

* no-mistakes(document): Document inactive outcome reconciliation contracts

* no-mistakes(review): Reject relative local secondmate parent routes

* no-mistakes(review): Key terminal receipts by spawn incarnation

* no-mistakes(review): Stabilize legacy receipts and lock reconciliation snapshots

* no-mistakes(review): Fail closed on invalid secondmate identity markers

* no-mistakes(document): Document durable inactive-outcome reconciliation

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* ci: raise Herdr test timeout (#2191)

* fix: refresh stale Pi instructions after compaction (#2163)

* fix(session-start): refresh drifted instructions on stale rebuilds

* test(session-start): prove Pi instruction refresh end to end

* no-mistakes(review): Fix stale instruction refresh and baseline integrity

* no-mistakes(review): Preserve true-start baselines across Pi continuations

* no-mistakes(review): Correct Pi continuation classification and live expectation

* no-mistakes(review): Correct Pi continuation coverage documentation

* no-mistakes(review): Fix read-only refresh and exact Pi session restores

* no-mistakes(review): Classify Pi create-if-missing sessions correctly

* no-mistakes(review): Classify named Pi sessions using immutable headers

* no-mistakes(review): Correct Codex interactive coverage diagnostic

* no-mistakes(document): Document immutable Pi compaction instruction refresh

* no-mistakes(document): Correct Pi refresh documentation and validation claims

* feat: add deterministic condition-to-action watcher (#2200)

* feat(bin): add deterministic condition->action watch adapter on the process-event channel

Register a (condition, action) pair once with bin/fm-procevent-when.sh and the
existing process-to-event runner polls the condition tokenlessly, fires the
action at most once on a stable true, and wakes firstmate exactly once with the
captured outcome - instead of burning an agent turn per re-check.

The pair is stored privately under state/when/ and hash-bound by a trust record
the same way fm-check-register.sh binds a custom check, so a mutated spec is
refused without executing anything. A durable exclusive fired marker claimed
before the action makes restarts and re-polls unable to double-fire; every
failure path (mutated spec, condition error past budget, expired deadline,
failed action, uncaptured earlier fire) ends in a terminal captured outcome
that wakes firstmate rather than a silent retry. Eligibility stays a firstmate
judgment: only exact, safe, reversible actions may be bound, and judgment-
needing or destructive actions keep the wake-and-decide flow.

* no-mistakes(review): Harden when watcher concurrency, deadlines, timeouts, and output

* no-mistakes(test): Bind watcher actions to registered executable bytes

* no-mistakes(document): Correct condition-action watcher documentation

* no-mistakes(document): Clarify outcome wake re-announcement

* no-mistakes: apply CI fixes

* fix(bin): honor a decision key stated after the verb colon (#2202)

The open-decisions fold only recognized a [key=<slug>] token between the
verb and the colon (needs-decision [key=x]: note). The common worker
shape with the colon first (needs-decision: [key=x] note) silently
folded its stated key into the shared "default" bucket, so two open
decisions could collapse into one record and fm-send --resolve-key <x>
refused to close the decision it plainly named.

A complete token at the head of the note is now an equivalent stated-key
position for every keyed verb, shared by the whole-file and incremental
folds through the one _fm_decision_key owner. The documented
before-colon position wins when both are present, a token deeper in the
note stays prose, a bare keyless line still folds to "default", and a
stated-but-malformed slug is rejected rather than rewritten to
"default". A consumed note-head token is stripped from the note so both
positions yield identical records, and the incremental fold version is
bumped so persisted cursors folded under the old interpretation are
rebuilt from the authoritative log.

Fixes #2109

* fix(bin): prevent watcher recovery acknowledgement livelock (#2212)

* fix(bin): keep a recovery acknowledgement valid across republication

A watcher cycle that opened and closed while the model handled its drained
wakes minted a fresh recovery generation, which invalidated the exact
acknowledgement the drain had just printed. That acknowledgement then consumed
nothing, so the marker stayed pending and every later arm spent its whole cycle
re-announcing the same recovery instead of supervising - a livelock the home
could not leave on its own.

A downtime publication now reuses the generation of an outstanding handling
episode, so a close during the handling window cannot orphan the printed
acknowledgement. The acknowledgement itself separates its two facts: queue-row
consumption is bound to the monotonic --ack-through sequence and always
happens, while only retiring the episode is bound to --recovery-generation. A
generation that moved on is a non-fatal result that names its own remedy
instead of a refusal that consumes nothing.

* no-mistakes(review): Preserve recovery generations and consume stale acknowledgements safely

* no-mistakes(document): Document sequence-bound recovery acknowledgements

* feat(fmx-respond): consume Relay conversation chains (#2206)

* feat(fmx-respond): consume in_reply_to_chain conversation context

The relay's poll payload can carry in_reply_to_chain, an oldest-first
transcript of the surrounding conversation, but the mention-handling
procedure only ever read the immediate in_reply_to parent, so referents
like "this" in a standalone mention stayed unresolvable even when
context was delivered.

Teach fmx-respond to read the chain when present (optional and
backward-compatible: often absent today, kind label not required),
resolve referents against the whole transcript, and extend the
untrusted-content framing to every chain entry including the upcoming
kind=history entries. Document the field's wire shape in
docs/configuration.md as the firstmate-side owner.

* no-mistakes(document): Document Relay chain context ownership

* fix: parse decision verbs before status metadata tags (#2280)

* fix(bin): strip every bracket tag, not just [key=...], from a status verb

status_line_verb only stripped a leading "[key=...]" token before the
colon, so a remote secondmate reply's leading "[corr=...]" correlation
tag stayed glued onto the returned verb word ("needs-decision
[corr=...]" instead of "needs-decision"). The open-decisions fold's
verb match then silently failed to recognize the line at all, so
fm-send --resolve-key refused to close a decision that was plainly
open on the status line.

Generalize the parser to strip every "[name=value]" tag before the
colon, in any order and count, so local and remote replies fold
identically.

* no-mistakes(review): Invalidate stale decision cursors after parser fix

* no-mistakes(document): Clarify status metadata verb parsing

* fix(bin): collapse duplicate supervision wakes (#2287)

* fix: collapse duplicate supervision wakes without losing legitimate updates

One remote-secondmate note produced two handling turns (a procevent check
wake published before autohandle, then a signal wake for the same mirrored
bytes), already-ingested replays such as a cursor-loss whole-log recapture
still woke with nothing to do, this home's own bookkeeping closes (fm-send
--resolve-key, the pending-reply escalation close, the captain-held
transfer) re-woke the session that wrote them, and turn-ended-only wakes
were annotated with already-announced status lines that looked like fresh
progress.

Dedup rules, each at its layer's one owner:
- fm-procevent.sh: an adapter may declare 'self-announcing'; the runner
  then applies first and publishes a check wake only for what remains
  unhandled. fm-procevent-remote-reply.sh declares it: the mirrored status
  append is the single announcement, so a fully applied capture publishes
  nothing and a byte-identical replay stays completely quiet. All other
  adapters keep strict publish-before-apply.
- fm-wake-lib.sh: fm_wake_signal_sig/seen_path/seen_current now own the
  watcher's signal signature and .seen-* marker format, plus
  fm_wake_status_append_self_announced, the guarded bookkeeping append
  that advances the marker only over exactly its own bytes and fails
  toward waking on any pending or interleaved foreign write.
- fm-send.sh, fm-pending-reply-lib.sh, fm-decision-hold.sh: bookkeeping
  closes go through that guarded append; escalation opens stay plain
  appends because a new blocker must wake.
- fm-wake-lib.sh annotations: a historical (turn-ended-only) row skips its
  status annotation only when the file's signature provably matches the
  seen marker; anything unannounced keeps annotating.
- fm-classify-lib.sh: a kind=secondmate task's status signal is never
  absorbed as provably-working, because that stream is the routed-reply
  channel the parent must read.

Also fixes a pre-existing exit-path deadlock the regression run reproduced:
a TERM inside a recovery-marker critical section left fm_lock_try_acquire
spinning against this same process's abandoned hold; a self-held lock is
now reclaimed (a subshell still waits on its parent's live hold).

Regression tests drive the real wake functions and executables in both
directions: each duplicate case collapses, while a new remote reply, new
decision, new blocker, merge result, failure, first status change, and a
later different note on the same task all still wake.

* no-mistakes(document): Document wake deduplication contracts

* feat: add Cursor CLI crew harness (#2238)

* feat(harness): add Cursor Agent CLI adapter

# Conflicts:
#	bin/fm-spawn.sh

* fix(composer): read cursor-agent's reverse-video placeholder as idle

cursor-agent renders its idle composer placeholder dim (SGR 2) but paints the
cell under the terminal cursor in reverse video (SGR 0;7). Reverse video is
neither dim nor a dark truecolor foreground, so the shared ghost stripper keeps
that one character and an idle composer reduces to a lone `P`. Judged on its
own, that remnant reads `pending` on a genuinely idle pane, which defers
away-mode escalation indefinitely on the styled cursorless backends.

Teach the ONE fleet-wide classifier the shape instead of adding an adapter-local
copy: register `→` as an agent prompt glyph so the composer row is structurally
findable at all (without it the bottom-most shape is a stale shell prompt echo
in the scrollback), add both verified placeholders to the idle set, and consult
the styling-independent plain row when the styled row is only a remnant.

The plain-row branch demands the remnant be a proper, strictly shorter substring
of a plain row matching a fully anchored placeholder. Real typed text is
uniformly bright, so stripping leaves it equal to the plain row and it stays
`pending` - verified live against a pane where the typed text was exactly the
placeholder string.

Verified live on cursor-agent 2026.08.11-e8db854; the regression pins the real
captured bytes and asserts the remnant survives stripping, so the case cannot go
vacuous if the stripper later learns SGR 7.

Co-authored-by: Amplify Logic AI <lars@sockinator.co>

* feat(cursor): narrow cursor identity and order its marker before CLAUDECODE

Cursor ships two executable names - `cursor-agent` and the legacy alias `agent`
- and runs as a bundled node script, so tmux reports the pane command as a bare
`node`. Neither `agent` nor `node` can be trusted by name, so identity gets one
owner in bin/fm-cursor-lib.sh that demands cursor's own name or install tree in
the path or argv[0], from the structural signal only. Probing an arbitrary pid's
executable during a liveness poll would execute a stranger's binary, which is
the hazard that rule exists to close.

Two consequences wired up:

Detection. cursor-agent does NOT clear an inherited CLAUDECODE, so a cursor
worker launched under a claude primary carries both markers and whichever is
tested first wins. The cursor markers are ordered ahead of the CLAUDECODE check;
fm-spawn additionally clears foreign markers at the launch boundary. Both are
kept deliberately - launch sanitization only covers sessions fm-spawn started,
while the ordering also covers a cursor session started by hand. Verified live
that CURSOR_INVOKED_AS is set on the agent process and CURSOR_AGENT=1 on the
child/tool processes fm-harness.sh actually runs as.

Pane liveness. A cursor pane now classifies `agent`. An unrelated node or agent
stays `other`, which the liveness callers already fold into `ambiguous` rather
than `dead`, so a stranger's node pane is never reported agent-free.

Resolution prints the STABLE launcher rather than the canonical target: identity
is proven through canonicalization, but cursor's canonical path carries a
version its own auto-update replaces, and pinning that would strand a task on a
version that can vanish.

The regression drives the two identity signals apart - a cursor-named executable
outside any cursor tree, and a non-cursor-named alias inside one - and asserts
each carries a verdict alone, so no single vendor string is load-bearing. Its
negative controls are real spawned processes, not fixtures.

Verified live on cursor-agent 2026.08.11-e8db854.

Co-authored-by: Ville Penttinen <villem.penttinen@gmail.com>

* feat(cursor): classify cursor busy state from its own turn transcript

Cursor shipped as "unknown cursor-unverified" on the premise that it exposes no
semantic turn lifecycle, only a rendered "Working" footer. That premise is
wrong: cursor-agent persists an append-only JSONL transcript per conversation
and brackets every submitted turn with a role:user open and a typed turn_ended
close. Verified live on 2026.08.11-e8db854, including the interrupt path, where
Escape closes the turn with status "aborted" - so this source covers manual
interruption, which Claude's Stop hook does not.

That makes it a genuine pull source in the muse mould rather than the rendered
text the redesign forbids: no writer, no arm, no gen, nothing seeded that could
never be cleared. Cursor's `ctrl+c to stop` footer stays out of the verdict, and
herdr's narrower native streaming state cannot stand in for it either.

Binding deliberately does not reconstruct cursor's workspace-slug directory
name. That slug collapses path separators, so rebuilding it would be a guess
that could bind the wrong pane; cursor records the exact absolute workspace path
in each project's .workspace-trusted, and the binding matches on that. A
conversation recorded as prior at spawn is excluded, so a relaunch in a reused
worktree folds its own turn rather than its predecessor's. Requiring a unique
remaining conversation keeps zero and several both unknown, because neither
proves anything about the current turn.

The regression pins the fold with real transcript files and asserts the
dangerous direction stays closed: an unresolvable binding, a record-free file,
an unclaimed workspace, and a workspace-path PREFIX all read unknown, never
idle. The prefix case uses an opaque fixture slug so a slug-rebuilding
implementation cannot pass it.

Co-authored-by: Ville Penttinen <villem.penttinen@gmail.com>

* feat(cursor): make the cursor launch runnable and give it lifecycle control

Five gaps that together kept a cursor crewmate from being drivable end to end.

Launch. The template invoked `cursor agent`, but `cursor` is not the CLI - the
installed names are `cursor-agent` and the legacy alias `agent` - so the command
could not run at all on a machine with a normal cursor install. It now resolves
through the verified owner, which also refuses a spawn loudly instead of leaving
a pane that dies with command-not-found and reads as a wedged worker.

Session binding. fm-spawn writes state/<id>.cursor-session so the busy fold can
find this pane's transcript, and teardown removes it.

Lifecycle control. No cursor PR touched fm-control-lib.sh, so
`fm-control <id> interrupt|exit|relaunch` could not drive a cursor worker at
all. Verified live: interrupt is a single Escape, exit is /exit, and cursor does
NOT repollute its composer with the cancelled prompt, so unlike muse it needs no
clear key. Secondmate is refused, matching the spawn refusal.

Submit acknowledgement. cursor parks its terminal cursor outside its composer,
so the composer verdict on tmux is always `unknown` and a submit could never be
acknowledged from the composer alone. The submit core's existing idle-to-busy
transition covers that case, but only if the pane's busy footer is recognised,
so cursor's `ctrl+c to stop` joins the harness-less default union the submit
cores read. The TOKEN is matched rather than the spinner verb: the same version
rendered both `Working` and `Running` in consecutive turns.

Bootstrap. A configured cursor crew harness with no cursor executable is now a
loud MISSING diagnostic rather than a first-spawn failure, and it accepts either
installed name.

Interrupt cancellation is deliberately left unconfirmed. The transcript does
type an aborted close, but its post-interrupt write latency measured as
variable - sometimes seconds, sometimes not within twenty - so a claim built on
it would be unreliable. Normal turn completion is prompt, which is what the busy
fold actually depends on.

Two inherited tests are corrected rather than deleted: the busy test asserted
cursor could have no semantic source, and the launch test pinned the literal
`cursor agent` string. Both now pin the verified behaviour, including that the
launch never allocates a second worktree.

Co-authored-by: ABHISHAKE KUMAR BOJJA <abojja@uvic.ca>
Co-authored-by: Ville Penttinen <villem.penttinen@gmail.com>

* docs(cursor): record the verified crewmate facts and extend the drift guard

The inherited cursor entry was written against 2026.08.04-aaa8809 and several of
its claims no longer hold: it named `cursor agent` as the binary (not the CLI
name), listed six Grok model ids of which the live catalog now returns two, and
recorded busy state, exit, interrupt, and skill invocation as unverified.

Replaced with what was measured against 2026.08.11-e8db854, including the two
facts most likely to be rediscovered painfully: cursor runs as a bundled node
script so its pane title is a bare `node`, and it parks its terminal cursor
outside its composer, which makes the tmux composer verdict permanently
`unknown` by design rather than a defect to chase.

Model ids now route to `--list-models` for the account instead of a fixed list,
since that list is exactly what drifted.

The live drift guard covers cursor, resolving it through the same verified owner
fm-spawn uses and passing --trust so the probe cannot hang on the workspace
prompt. Run against every installed harness: 8 checked, all alive, with cursor
reporting title='node' foreground=[.../cursor-agent] - the drift shape this
guard exists to catch.

Co-authored-by: Ville Penttinen <villem.penttinen@gmail.com>

* docs(agents): record the cursor session-binding state file

The state/ layout section is the inventory every session reads; a busy-source
binding that fm-spawn writes and teardown removes belongs in it alongside muse's.

* no-mistakes(review): Sanitize ambient Cursor marker in harness tests

* no-mistakes(review): Validate Cursor models against live catalog

* no-mistakes(review): Reject unsupported secondmates before binary preflight

* no-mistakes(review): Narrow Cursor ancestry detection to structured process identity

* no-mistakes(review): Parse Cursor transcripts and sanitize inherited markers

* no-mistakes(review): Handle malformed Cursor transcript records safely

* no-mistakes(review): Validate malformed Cursor closes in fallback parser

* no-mistakes(review): Retire stale Cursor bindings during relaunch

* no-mistakes(review): Fix Cursor drift guard command variable

* no-mistakes(review): Narrow Cursor identity to versioned install trees

* no-mistakes(document): Document Cursor harness boundaries

* refactor(composer): move the delivery busy footers to the shared owner

The per-harness rendered busy footers lived in bin/fm-tmux-lib.sh under
FM_TMUX_* names, so cursor's `ctrl+c to stop` signature - and every other
harness's - was reachable only from tmux. That placement was wrong on its own
terms: herdr, zellij, cmux, and orca run the same harnesses and face the same
question these footers answer, which is whether a submitted Enter actually
landed. Nothing about the signature is tmux-specific.

Moved verbatim into bin/fm-composer-lib.sh, the shared composer/delivery owner
every backend already sources, and renamed to FM_DELIVERY_* so the names stop
claiming a scope they never had. All five adapters now reach cursor's signature;
verified per adapter rather than assumed.

The boundary the move must not blur is stated where it now lives: this is a
DELIVERY guard, never a worker-state source. Confirming a keystroke landed is a
different question from asking what a worker is doing, and bin/fm-busy-lib.sh
remains the semantic owner that forbids classifying a harness from rendered
text. Cursor still classifies only from its transcript fold, which is already
backend-agnostic because it folds a file rather than reading a pane - the same
verdict on all six backends.

The old FM_TMUX_* aliases are dropped rather than kept as dead shims: nothing
outside the moved block referenced them except fm-busy-lib.sh's grok fallback,
which now reads the new name. The documented operator override, FM_BUSY_REGEX,
is untouched.

Also removes a dead duplicate CURSOR_INVOKED_AS check in bin/fm-harness.sh,
unreachable behind the marker check above it.

* no-mistakes(review): Correct shared delivery guard ownership references

* no-mistakes(document): Document shared delivery guards and Cursor backend limits

* no-mistakes: apply CI fixes

* fix(composer): bound a bare composer's wrap region at a half-block rule

A live cursor crewmate on herdr classified its IDLE composer as `pending`, and
fm-send consequently exited 1 with "delivery unconfirmed" on a message that had
actually landed. The cause is not cursor-specific.

Herdr draws a composer's top and bottom rules with the half-block glyphs U+2584
and U+2580 rather than the box-drawing family. fm_composer_row_has_edge knew
only the box-drawing set, so no box was detected; the composer was found as a
BARE row, and its wrap region - which extends while rows are non-blank and carry
no structural edge - walked straight through the composer's own closing rule and
swallowed the model and path footer below it. That footer is real text, so the
region classified pending on a genuinely idle pane.

Teaching the shared edge detector the half-block glyphs bounds the region at the
closing rule. Measured on the captured bytes of a real herdr cursor pane: the
same capture that read `pending` now reads `empty`.

This is a shared shape-path change, so it is deliberately narrow - it adds
glyphs to the edge vocabulary and changes no verdict logic - and the whole
composer and backend suite is green, including the other harnesses' herdr
fixtures.

The regression pins the real captured shape and asserts the footer content is
genuinely present, so the case cannot pass vacuously if the region were ever
bounded for some unrelated reason.

* fix(herdr): confirm a cursor submit from the rendered-footer transition

Herdr's composer-shape fix made an idle cursor pane classify `empty`, but
`fm-send` still exited 1 with "delivery unconfirmed" on messages that had
actually landed. Live measurement found the second, independent cause.

Herdr reports a cursor pane `agent_status=blocked` in EVERY state - idle,
mid-turn, and after - so the submit path's idle-baseline native confirmation is
structurally unreachable for cursor and every send falls into the composer
branch. That branch reads cursor's mid-turn composer row, which renders its own
`Add a follow-up` placeholder beside a right-aligned `ctrl+c to stop`. That
token is composer content, so the verdict is `pending` on a composer holding no
user text at all, and the Enter-retry budget then reports pending.

The escape is the same semantic signal the native path uses, read from the
pane's verified busy footer instead of native agent-state, and it is the
rendered-footer twin of the tmux submit core's turn-started confirmation: an
idle-to-busy transition ACROSS our Enter proves the harness accepted the
submission. The baseline is taken before the first Enter and only when the
native baseline was not legibly idle, so the idle-baseline path still never
reads pane content and a pane already mid-turn before we typed keeps reporting
`pending` rather than borrowing another turn as proof of this delivery.

The composer verdict is deliberately NOT relaxed. A right-aligned status token
on the composer row stays content for every other caller, including the
away-mode pre-injection guard, and the shared cursorless submit core is left
untouched so zellij, cmux, and Orca keep the behavior their own follow-up owns.

Verified live on herdr 0.8.0 and cursor-agent 2026.08.11-e8db854 in an isolated
lab session: `fm-send` now exits 0 and the steer executes, interrupt cancels a
running turn, `/exit` stops the agent, and teardown clears the record. All seven
panes of the running default session classify identically before and after the
shape fix, so no other harness regressed.

* no-mistakes(review): Prevent working Herdr baselines from falsely confirming delivery

* no-mistakes(document): Correct Cursor harness and backend documentation

---------

Co-authored-by: ABHISHAKE KUMAR BOJJA <abojja@uvic.ca>
Co-authored-by: Amplify Logic AI <lars@sockinator.co>
Co-authored-by: Ville Penttinen <villem.penttinen@gmail.com>

* fix(bin): require quota-axi 0.1.25 (#2300)

* fix: raise quota-axi floor to 0.1.25 for Cursor CLI quota awareness

Homes on latest main need quota-axi #87 so Desktop-absent CLI machines report a fresh Cursor quota instead of a false sign-in-required.

* no-mistakes(document): Update quota floor documentation pointer

* fix(bin): prevent false Pi watcher alarms during hand-offs (#2304)

* fix(guard): stop the false send-time watcher-down alarm on Pi primaries

On a Pi primary the watcher process is not the liveness signal. The Pi
extension tears the watcher down on every actionable wake and spawns the
replacement itself, so the singleton lock is legitimately unheld between
cycles: every one of the 799 cycles in a live primary's ledger ends with
lock_after=pid:none, and a live capture caught the guard verdict flipping to
no-watcher during one hand-off with the beacon 63s old.

bin/fm-guard.sh classified Pi as a persistent-watcher harness, which demands a
live identity-matched lock holder at all times, so any guarded command landing
in a hand-off painted the full WATCHER DOWN - SUPERVISION IS OFF banner and
told firstmate to repair a cycle the extension already owns and is restoring.

Add an extension supervision model for pi and pi-signed. A live
identity-matched watcher stays the ordinary healthy state; an unheld lock is
healthy only while the beacon is fresh within grace AND a live Pi session
provably owns continuity - both primary extensions recorded in their state
markers at their current on-disk builds by the process named in state/.lock,
with that process still alive. Without that proof the banner fires exactly as
before, so an unloaded, version-drifted, or exited Pi session is loud
immediately and a cycle the extension never restores is loud once the beacon
passes grace. The queued-wake warning, the PID-strict turn-end guard, and
every other primary's detection are untouched.

Fold session-start's duplicate Pi marker predicate into the shared library so
the ownership contract has one owner.

* no-mistakes(review): Restrict Pi hand-off tolerance to unheld watcher locks

* no-mistakes(document): Document Pi watcher hand-off supervision

* feat: support Cursor Agent CLI as a primary harness (#2305)

* feat(cursor): add Cursor Agent CLI primary hooks, park supervision, and session start

Register a tracked project-scope .cursor/hooks.json for Cursor's stop,
sessionStart, preCompact, and preToolUse steps.

bin/fm-turnend-guard-cursor.sh owns Cursor's turn boundary as a park: it
foregrounds the watcher arm, holds the boundary open until an actionable close,
and returns that wake as one follow-up. Exit 2 is a silent no-op on Cursor's
stop step, so the adapter never uses it. The follow-up loop is bounded twice,
by Cursor's own loop_limit and by the payload's loop_count.

bin/fm-sessionstart-cursor.sh delivers the digest as additional_context at
sessionStart, and stages it for the next turn boundary at preCompact, which
cannot inject context.

Cursor also loads the tracked Claude settings, so bin/fm-hook-host-lib.sh lets
each tracked Claude-shaped entrypoint stand down on a Cursor-delivered payload
rather than running every covered event twice.

bin/fm-tmux-lib.sh reclassifies a Cursor pane's composer cursorlessly, because
Cursor parks its terminal cursor outside the composer, which restores a genuine
composer-empty proof and unblocks away-mode escalation delivery.

* feat(cursor): make Cursor Agent CLI a verified primary harness

Resolve Cursor in the session-lock ancestry through bin/fm-cursor-lib.sh, which
a Cursor primary needs before it can hold its own home lock, and classify its
stop-hook park under the autoarm supervision model so the mid-turn pull guard
stops reporting a healthy between-turns watcher as down.

Read a Cursor pane's composer cursorlessly on tmux, gated on Cursor's own
structural process identity, which restores a genuine composer-empty proof and
lets away-mode escalations reach a Cursor primary with no daemon change.

Lift the secondmate refusals in bin/fm-spawn.sh and bin/fm-control-lib.sh now
that the supervision protocol exists and is recorded.

Cover the whole surface with a portable regression over real processes, an
opt-in live guard against the installed cursor-agent, and dated per-harness
evidence.

* docs(cursor): record Cursor as a verified primary across the owning surfaces

Update the turn-end guard, session-start, arm-seatbelt, cd-guard, watcher
continuity, architecture, configuration, README, and harness-adapters owners,
and add dated live evidence to the supervision and runtime-backend verification
records. Correct the recorded Cursor tmux composer verdict: the cursor-anchored
read is still blind, but the composite reader is no longer unknown.

Lift the remaining remote-secondmate refusal missed in the previous commit, and
add the new libs to the existing fixtures that copy a fixed dependency list.

* refactor(cursor): name the park's stand-down condition for both its causes

Also record that Cursor's preCompact firing itself is not yet live-verified,
while the static evidence that it cannot inject context, and the staging path
that follows from it, both are.

* test: give the pretool fixtures their new dependency and one lint owner

The cd-guard fixture copies a fixed dependency list and now needs the shared
hook-host predicate. Both pretool suites also asserted cleanliness with a bare
shellcheck call, a second and weaker copy of the lint definition that
bin/fm-lint.sh owns: it omits --external-sources, so it failed the moment these
checkers sourced a shared library. They now delegate to that owner.

* test: assert the cursor secondmate contract instead of its removed refusal

A cursor secondmate now launches, so the suite asserts what its park actually
needs: --trust so the home's project hooks load at all, its own home pinned as
the workspace, and the autoarm supervision model inherited across the launch.

* no-mistakes(review): Serialize Cursor wakes and bind staged context

* no-mistakes(review): Serialize Cursor context and nag state commits

* no-mistakes(review): Enforce Cursor ceiling before staged context delivery

* no-mistakes(review): Serialize Cursor claims and staged context

* no-mistakes(review): Serialize Cursor ownership and state commits

* no-mistakes(review): Protect Cursor context across session takeover

* no-mistakes(review): Preserve Cursor context across session takeover

* no-mistakes(review): Enforce owner-keyed Cursor staged context

* no-mistakes(review): Atomically claim Cursor follow-ups and staged context

* no-mistakes(review): Defer Cursor preCompact staging and simplify supersession

* no-mistakes(review): Serialize Cursor park commits and defer preCompact

* no-mistakes(review): Stop Cursor parks after session takeover

* no-mistakes(test): Route Cursor preCompact context through stop follow-up

* no-mistakes(document): Update Cursor primary documentation

* revert(cursor): cut preCompact staging from this change

Carrying a compaction digest across two concurrently running stop hooks kept
producing races that could deliver it twice or strand it indefinitely, and
closing them kept enlarging a critical section inside a hook Cursor awaits at
the turn boundary. Native preCompact firing was never observed either, so the
surface has no empirical basis yet.

Remove the adapter, its registration, its staged path in the park, and its
tests, and record the surface as deferred and uncovered alongside the Codex
interactive TUI. A regression now asserts preCompact stays unregistered so it
cannot return without its own design and evidence.

This change ships the proven core only: the turn-end follow-up park, the
run-tier session start, and away-mode delivery.

* no-mistakes(review): Correct Cursor park supersession documentation

* no-mistakes(document): Clarify Cursor run-tier verification ownership

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

---------

Co-authored-by: kunchenguid <kun-1@kunchenguid.com>

* feat(bin): add decline and repair paths for decision holds (#2330)

* feat(bin): add unrouted close paths to the captain decision gate

A captain who declines a held decision leaves no follow-up work to route,
so `resolve` could not express that answer: it requires at least one
`--routed-to` task. The only way to close such a hold was a direct
`tasks-axi done`, which never writes the durable resolution record the
completion gate reads, so the originating investigation could no longer
pass `verify` and its cleanup stayed blocked.

Add two close paths that route no work:

- `decline` closes an actively held hold with a recorded captain decision
  and no routed task. It refuses while any task is still blocked by the
  hold, because releasing routed work without recording it is `resolve`'s
  job.
- `repair` records the missing resolution block on a hold that was already
  closed outside this script. It never reopens a hold and never clears a
  dependency edge, and it refuses a hold that is still actively held.

Both require a non-empty captain decision file and share `resolve`'s
digest-based retry identity, so an exact retry is idempotent while a
changed decision is rejected. The recorded body now also names which path
closed the hold, and each routed entry regains its own line.

The gate itself is unchanged: an unanswered decision still fails
completion and blocks teardown, and neither new path can close a hold
without the captain's recorded word.

* fix(bin): require captain-hold provenance before repairing a decision

`repair` checked only that the backlog item was kind captain and Done, so
an ordinary captain-kind task that was never held for the captain could be
closed, repaired, and then pass the completion gate.

tasks-axi keeps `hold_kind` through a close, so it is the surviving proof
that an identity really was a captain hold. Require it before writing the
resolution record, and cover the case in the gate regression.

* no-mistakes(document): Correct decision-hold lifecycle documentation

* fix(bin): surface buried wake status lines once (#2331)

* fix(bin): surface buried status notes on wake drain

A note: answer immediately followed by a routine note was dropped because
annotations kept only the newest line and note: never enters OPEN DECISIONS.
Present every unread note and pending-reply resolution since the last drain
cursor, and annotate every unread line on a queued signal.

* no-mistakes(review): Fix unread status cursor races and overflow

* no-mistakes(review): Preserve cursors when status span reads fail

* no-mistakes(review): Make status presentation transactional under I/O failures

* no-mistakes(review): Simplify unread status cursor and presentation locking

* no-mistakes(review): Align cursor failure regressions with transactional presentation

* no-mistakes(review): Retire stale presentation cursors during task teardown

* no-mistakes(review): Preserve routine status until signal annotation

* no-mistakes(review): Correct unread status cap documentation

* no-mistakes(document): Document unread wake status presentation

* no-mistakes(lint): Fix wake surfacing ShellCheck warnings

* no-mistakes: apply CI fixes

* feat: add max Calm presentation level (#2334)

* feat(calm): add a max presentation level that hides mid-turn working notes

Calm's home-local preference becomes a three-state level instead of a
boolean: "off" is stock Pi, "on" is today's Calm, and "max" is Calm plus
hiding the assistant text of messages the model did not end its response
with. `/calm max` selects it from any state, a plain `/calm` steps max
back to ordinary Calm and otherwise keeps the existing on/off cycle, and
any other argument keeps that cycle too.

`config/calm` now persists "max" as its own literal value, so a session
start, resume, fork, or reload restores the stored level rather than
treating it as unrecognized and dropping to off.

The hide rule keys on Pi's intrinsic per-message stopReason: "toolUse",
or "length" with tool calls present. Streaming ("pending") text is never
filtered, because suppressing it would also stop a genuine reply from
streaming. The existing assistant layout adapter filters the blocks out
of the same shallow presentation copy it already uses for collapsed
thinking, so the message, model context, session storage, /export, and
delivery are untouched and a hidden mid-turn row collapses to zero
height. The new "assistant-working-note" class keeps that choice in the
visibility policy owner, where ordinary Calm keeps it visible.

* no-mistakes(document): Clarify Calm max persistence and taxonomy

* feat(calm): hide mid-turn working notes by default (#2339)

* feat(calm): make hiding mid-turn working notes the ordinary Calm state

Calm collapses back to the two-state on/off toggle it was before the max
presentation level, with max's hide rule promoted into ordinary Calm.
Calm on now hides mid-turn assistant working notes in addition to what it
already hid, and the /calm command parses no argument again.

The hide rule itself is unchanged: assistant text is removed from the
shallow presentation copy when the message's own stopReason is "toolUse",
or "length" with tool calls present. Streaming ("pending") text is never
filtered, so a genuine reply still streams. The message, model context,
session storage, /export, and delivery remain untouched.

config/calm persists only "on" and "off" again, but the reader still maps
a persisted "max" to on so a home upgraded from the removed level keeps
Calm on instead of dropping to off.

The mid-turn hide is now default behavior rather than an opt-in level, so
docs/calm.md documents it for users, docs/configuration.md records the
two written values plus the legacy max mapping, and the feasibility
taxonomy drops its level-scoped wording.

* no-mistakes(document): Document ordinary Calm working-note hiding

* chore: store no-mistakes test evidence in the repo (#2355)

* chore: ignore scratchpad/ at the repo root (#2359)

* fix(ci): fail hung Herdr behavior runs in 20 minutes (#2413)

A wedged family-run step was occupying the runner until the 75-minute
job cap; bound that step so cleanup and timing artifacts still upload.

* fix: keep the public promise reachable when work is routed to a second mate (#2457)

The lightweight Relay follow-up link lives in the answering home's own
state/<task-id>.meta, so it can only bind work that home owns. When a
Relay-linked request is routed to a second mate, the task record lives in the
second mate's home, fm-x-link.sh failed with a bare "no such task ...meta", and
nothing else picked the promise up: only the soft acknowledgement was ever
posted. The typed promised-final path already supports --work-home
secondmate:<id>; the playbook simply never chose it.

- fmx-respond now states the routing rule crisply: a task in this home takes the
  lightweight link, and second-mate-routed work takes a promised-final
  commitment bound to that home, registered up front with the brief command
  carried into the routed worker's instructions.
- fm-x-link.sh refuses a task with no local record by naming the registered
  second mate whose home actually holds it and printing the promised-final
  registration command, with the exact --work-home when the match is
  unambiguous. A home with no registered second mates keeps the plain error.
- fm-backlog-handoff.sh reports, after a successful move, any moved key that
  still owes a public reply bound to main/<key>, since that binding no longer
  names the home owning the work. The move itself is never blocked.

Docs and the secondmate handoff prose follow the same rule. Tests cover the
refusal, its scoping, the unchanged local-link path, and both handoff outcomes
at the script boundary.

* docs(skills): add remote-secondmate recovery hint for false-negative verdicts (#2456)

* fix(skills): hint that remote secondmate liveness verdicts false-negative

fm-crew-state and fm-send routinely misreport a live remote secondmate
as dead; confirm against the pane before relaunching, and relaunch only
through fm-spawn.sh, never raw herdr pane surgery.

* no-mistakes: apply CI fixes

* fix(calm): keep Pi's export confirmation visible (#2461)

Pi 0.83.0 added a status line to every tool-expansion change, and Pi
updates the previous status line in place when two status messages
arrive back to back. Calm's post-export redraw cycled tool expansion on
the macrotask right after Pi printed "Session exported to: <path>", so
both expansion status lines coalesced over that confirmation and the
captain was left with no record of where their export landed.

Calm now repaints only the tool rows it presents, by invalidating each
row through the render context Pi hands its render slots, and requests
the surrounding redraw through setStatus. Neither appends to the
transcript. The repaint is still needed because Pi can re-render a row
asynchronously - the built-in edit row invalidates itself once its diff
is ready - and that re-render can land inside the window where /export
forces stock rendering.

The real-terminal /export case now asserts the confirmation is still on
screen after the redraw has settled, and that the redraw restored every
Calm-hidden row, instead of only racing the moment the confirmation
first appeared.

* feat(stow): add open-record persistence to /stow before reset (#2488)

* feat(stow): persist the open records a session is holding

/stow curated memory and captured session knowledge, but never touched
record state, while AGENTS.md called it an "unfinished-work sweep" and the
receipt declare…
Bre77 added a commit to Bre77/firstmate that referenced this pull request Sep 2, 2026
…ts) (#72)

* fix(ci): fail hung Herdr behavior runs in 20 minutes (#2413)

A wedged family-run step was occupying the runner until the 75-minute
job cap; bound that step so cleanup and timing artifacts still upload.

* fix: keep the public promise reachable when work is routed to a second mate (#2457)

The lightweight Relay follow-up link lives in the answering home's own
state/<task-id>.meta, so it can only bind work that home owns. When a
Relay-linked request is routed to a second mate, the task record lives in the
second mate's home, fm-x-link.sh failed with a bare "no such task ...meta", and
nothing else picked the promise up: only the soft acknowledgement was ever
posted. The typed promised-final path already supports --work-home
secondmate:<id>; the playbook simply never chose it.

- fmx-respond now states the routing rule crisply: a task in this home takes the
  lightweight link, and second-mate-routed work takes a promised-final
  commitment bound to that home, registered up front with the brief command
  carried into the routed worker's instructions.
- fm-x-link.sh refuses a task with no local record by naming the registered
  second mate whose home actually holds it and printing the promised-final
  registration command, with the exact --work-home when the match is
  unambiguous. A home with no registered second mates keeps the plain error.
- fm-backlog-handoff.sh reports, after a successful move, any moved key that
  still owes a public reply bound to main/<key>, since that binding no longer
  names the home owning the work. The move itself is never blocked.

Docs and the secondmate handoff prose follow the same rule. Tests cover the
refusal, its scoping, the unchanged local-link path, and both handoff outcomes
at the script boundary.

* docs(skills): add remote-secondmate recovery hint for false-negative verdicts (#2456)

* fix(skills): hint that remote secondmate liveness verdicts false-negative

fm-crew-state and fm-send routinely misreport a live remote secondmate
as dead; confirm against the pane before relaunching, and relaunch only
through fm-spawn.sh, never raw herdr pane surgery.

* no-mistakes: apply CI fixes

* fix(calm): keep Pi's export confirmation visible (#2461)

Pi 0.83.0 added a status line to every tool-expansion change, and Pi
updates the previous status line in place when two status messages
arrive back to back. Calm's post-export redraw cycled tool expansion on
the macrotask right after Pi printed "Session exported to: <path>", so
both expansion status lines coalesced over that confirmation and the
captain was left with no record of where their export landed.

Calm now repaints only the tool rows it presents, by invalidating each
row through the render context Pi hands its render slots, and requests
the surrounding redraw through setStatus. Neither appends to the
transcript. The repaint is still needed because Pi can re-render a row
asynchronously - the built-in edit row invalidates itself once its diff
is ready - and that re-render can land inside the window where /export
forces stock rendering.

The real-terminal /export case now asserts the confirmation is still on
screen after the redraw has settled, and that the redraw restored every
Calm-hidden row, instead of only racing the moment the confirmation
first appeared.

* feat(stow): add open-record persistence to /stow before reset (#2488)

* feat(stow): persist the open records a session is holding

/stow curated memory and captured session knowledge, but never touched
record state, while AGENTS.md called it an "unfinished-work sweep" and the
receipt declared the session "safe to reset" - wording that implied a
record-correctness guarantee stow does not make. A shipped PR with no
backlog item, a queued umbrella whose phases had merged, and four decision
holds left open after their answers shipped all survived repeated stows.

Add a bounded pass that files record state from the same volatile input the
rest of stow already uses: the open threads in context, minutes before the
reset destroys them. It creates a record for an unfiled thread and corrects
one the session knows is wrong, through the owning path, and states its
boundary as part of the contract - it never enumerates the backlog, lists
holds, or queries a forge, because it cannot be a reconciliation and must
not be read as one.

Correct the wording in AGENTS.md and the completion receipt so reset-safe
means what it actually guarantees: nothing this session knew was lost.

* no-mistakes(review): correct stow decision-hold inspection to read hold via tasks-axi

* no-mistakes(document): note /stow open-record persistence in README command catalog

* refactor(stow): state open-record persistence as principle, not procedure

The first version enumerated triggers, named commands, and prescribed an
ordered procedure. That is too rigid for an agent skill: it invites literal
execution of a checklist instead of judgment, and every enumerated example
is a way for the guidance to go stale.

Reduce it to the intent - before a reset, the important open work you are
holding in context must end up durably recorded rather than dying with the
session, filing what is unfiled and correcting what is stale - and let the
agent judge importance, the record, and the owning write path.

Keep the scope bound, since it is a decided contract and not a mechanic:
this covers the open work the session is holding, never a reconciliation of
durable records against repository or forge reality. The wording
corrections in AGENTS.md and the completion receipt are unchanged.

* fix(decisions): close decision holds at answer time via one general keyed-answer path (#2490)

* fix(decisions): close captain holds at answer time

Firstmate had two "a decision is open" ledgers with asymmetric closing
mechanics. The live status-log ledger closes atomically at answer time,
because bin/fm-send.sh --resolve-key makes answering a decision be the
act that closes it. The durable backlog hold ledger had no such coupling:
answering and recording were two separate acts, and only the first was
forced by the workflow.

That asymmetry lost four real captain decisions. Their answers were
captured durably to disk, keyed character for character by the hold
decision keys, acknowledged, and even implemented and shipped, yet the
holds stayed open for two days and the captain was asked to re-answer
decisions already on his own disk.

Give the hold ledger the same answer-time-closure property:

- bin/fm-decision-hold.sh gains an `answer` subcommand, the hold ledger's
  counterpart to --resolve-key. It shares one unrouted close
  implementation with `decline`, so it carries every existing guard - the
  captain decision file, the active-hold requirement, retry identity, and
  the refusal to release still-routed work - and differs only in the
  resolution mode it records. `decline` keeps its stronger meaning that
  the answer routes no follow-up work at all.
- bin/fm-procevent-lavish.sh wires the channel that actually carried the
  lost answers. `arm --decisions-origin` binds a deck to the origin whose
  holds it carries, `answers` reads the structured choices out of a
  captured poll result, `close-decisions` maps each key to its hold and
  closes it through the command above, and `autohandle` lets the runner
  apply that at capture time.

Safety is preserved rather than traded away. Only rows tagged `choice`
are read, so freeform captain prose cannot forge a decision key. Closure
is confined to the one bound origin. The decision text is a pure function
of the captured result, so a replayed capture is idempotent. A hold that
is absent, already closed, or still blocking routed work is skipped and
left for `resolve`, never forced. A deck armed without the binding
touches no hold at all. And autohandle deliberately never reports full
handling, because recording an answer is transcription while acting on it
is firstmate's judgement - so the check wake still reaches the handler.

fm-send --resolve-key is untouched.

* no-mistakes(document): document state/lavish-decisions binding dir in AGENTS.md state inventory

* refactor(decisions): make keyed-answer closure one general capability

The previous pass gave holds answer-time closure but built it as bespoke
Lavish wiring: the review adapter carried the source-to-origin binding,
mapped keys to hold identities, wrote decision records, decided what to
skip, and closed holds itself. That treated a review deck as a special
decision source. It is not - it is an ephemeral discussion format that
happens to carry answers.

Collapse it into ONE general capability with one owner.

bin/fm-decision-hold.sh now owns the whole of "a keyed answer closes its
matching hold":
- `answers <origin> --source <provenance>` is the channel-agnostic
  intake. It reads key/answer/label lines on stdin, maps each key to its
  hold, and closes it through the same `answer` path, so every guard
  applies identically whatever channel the answer came from. --source is
  provenance recorded in the decision, never a behavior switch; there is
  no per-channel branch and no knowledge of chat, decks, or transports.
- `bind`/`unbind`/`binding` own the source-to-origin binding for any
  channel whose answers arrive detached from their origin.

Every channel is now an ordinary caller that only turns what it received
into keyed lines:
- bin/fm-send.sh (chat) feeds the intake for a key that names an active
  hold. This also fixes a real gap: once `complete` transfers a decision
  to its hold it closes the live status copy, so --resolve-key alone
  could never answer a transferred decision.
- bin/fm-procevent.sh feeds it generically. A bound source's captured
  result goes to `<adapter> answers <result-file>` and whatever that
  prints is piped into the intake. The runner names no adapter, parses
  no result, and carries no decision rule, so any future adapter with an
  `answers` command works with no change here.
- bin/fm-procevent-lavish.sh keeps only `answers`, which reports the
  structured choices a review captured and stops. It maps nothing to a
  hold and closes nothing; it lost ~160 lines of decision logic.

Feeding is independent of handling, so it never acknowledges a result
and never suppresses a wake - recording an answer is transcription,
acting on it stays firstmate's judgement.

The regression that proves closure now drives a FIXTURE adapter that is
not the review adapter, so what is proven is that any bound channel
reaches the intake rather than that one channel is wired specially. A
new regression drives the real fm-send over a stubbed transport for the
chat side. Every prior guarantee still holds, and fm-send's status-log
behavior is unchanged.

* no-mistakes(review): test(decisions): drop source-content grep from hold-closure regression

* fix(memory): emit a real @AGENTS.md pointer instead of a CLAUDE.md symlink (#2512)

A Write aimed at CLAUDE.md followed the symlink and destroyed AGENTS.md.
The installer now creates and migrates to a recoverable two-line pointer file.

* fix(ci): keep CLAUDE.md pointer check valid (#2515)

* ci: gate GitHub workflows with pinned actionlint (#2517)

* fix(lint): catch malformed GitHub workflows before merge

A self-broken ci.yml cannot report its own breakage, so parse every
workflow in the local lint path that no-mistakes already runs.

* fix(lint): pin actionlint instead of Ruby for workflow lint

A self-broken ci.yml still has to fail in the local lint path, and the
named tool for that gate is actionlint, not a new Ruby runtime.

* no-mistakes(document): Clarify pinned workflow lint documentation

* fix: install pinned lint tools across supported platforms (#2546)

* fix: install pinned shellcheck and actionlint on macOS and linux arm64

The installers were hardcoded to linux amd64 and sha256sum, so a Mac
dev could not satisfy the refuse-on-mismatch lint gate. Select the
official per-platform archive and checksum, and fall back to shasum -a 256.

* no-mistakes(document): Document cross-platform pinned lint installers

* docs: reconcile test-evidence docs with store_in_repo: true (#2548)

.no-mistakes.yaml has set test.evidence.store_in_repo: true since #2355, but
CONTRIBUTING.md, docs/configuration.md, and docs/architecture.md still described
the old policy of keeping evidence out of the repo in a temp directory.

The current no-mistakes behavior for store_in_repo: true is to publish each run's
test evidence to the orphan no-mistakes/evidence branch and link it from the PR
body. That branch shares no history with code branches, so evidence never enters
a pushed feature branch or the default branch, and CI's tracked personal fleet
paths rule stays accurate.

Docs only. No change to .no-mistakes.yaml or any workflow.

* docs: clarify test evidence branch storage (#2549)

* docs: correct test evidence storage comment in .no-mistakes.yaml

* no-mistakes: apply CI fixes

* docs: hint that live scouts may host their own Lavish review loop (#2563)

Make that a first-class option in always-loaded instructions so firstmate does not default to mediating and tearing the scout down between iteration rounds.

* fix(bin): report remote secondmate delivery and state truthfully (#2570)

* fix(bin): report remote secondmate delivery and state truthfully

A steer to a remote secondmate crosses fm-on.sh to a host-local fm-send
leg whose unconfirmed submit read-back (verdict=pending, typically a busy
mate whose harness queues the steer) was flattened into exit 1, so the
parent printed "error: text not submitted" / "error: text not sent" and
discarded the pending-reply expectation for a steer that had actually
landed. fm-send now carries the verdict across the ssh boundary as a
documented delivered-unconfirmed exit 3: the parent reports the steer as
delivered with confirmation pending, exits 0, keeps the expectation armed
(awaiting_report), and closes --resolve-key decisions, while transport
loss (ssh 255) and real remote failures keep failing loudly with the
remote leg's stderr attached. A local unconfirmed submit now also exits 3
with an honest non-error message and still never closes a decision key.

fm-crew-state.sh and fm-peek.sh no longer read a remote mate's endpoint
through local probes (which misreported a healthy mate as "worktree gone"
/ "can't find session: remote"): both now use the true remote source over
fm-on.sh, and an unreachable or unreadable remote reads as unknown-remote,
never as gone or dead.

* no-mistakes(document): Document remote delivery and state truth

* no-mistakes: apply CI fixes

* feat: adopt spendPriority for quota dispatch (#2574)

* Adopt quota-axi 0.1.29 spendPriority-primary array dispatch.

quota-axi 0.1.29 publishes schema 5 with selection.spendPriority as the primary comparative signal and demotes derivation fields out of default --json. Rank comparable-fit candidates on that scalar, keep runway versus the completion horizon as a hard gate, and raise the compatibility floor so a pre-consolidation build cannot reach dispatch intake.

* no-mistakes(review): Correct schema fixtures and remove prescriptive selection prompts

* no-mistakes(document): Correct quota verification evidence chronology

* Collapse quota-array-dispatch onto TOON-first spendPriority ranking.

Decide from quota-axi's default TOON; keep --json as a rare defensive fallback.
Rank by spendPriority after eligibility, reasoning-class, and runway-feasibility gates, and drop the hand-computed Pareto, pace, reserve, and window-id layers.

* no-mistakes(review): Permit ambiguous JSON fallback and correct reset fixtures

* no-mistakes(review): Correct runway semantics and escalate unresolved uncertainty

* no-mistakes(document): Document TOON-first quota dispatch evidence

* docs: add GROK_BOT.md Grok Bot system prompt (#2590)

* docs: add GROK_BOT.md Grok Bot system prompt

* docs: amend GROK_BOT.md with charter report-back and delegation marker

* docs: classify GROK_BOT.md as public-product

* docs: make GROK_BOT.md the plain Grok Bot system prompt

* docs: update GROK_BOT.md nautical terms and self-improvement (#2592)

* doc: Update language in GROK_BOT.md for clarity

Refine language for clarity and consistency in instructions.

* fix(bin): preserve inactive reconciliation scan progress (#2595)

* fix(bin): guarantee inactive-reconcile scan progress under second quantization

The inactive-outcome scan computed its aggregate deadline in whole seconds,
so a 1-second budget's effective value lands anywhere in (0,1]; a scan
starting just before a wall-clock second boundary rounded its whole budget
away mid-scan and exited having visited no child, while the durable cursor
had already advanced past the never-examined child. This is the CI flake
behind tests/fm-inactive-reconcile.test.sh's 'next bounded scan did not
resume with the following child' (watcher-wake-lock family, portable
serial 2, seen on the PR #2590 run).

Every scan now visits at least its first due child with the per-child
state-read bound floored at one second, so no invocation can be a zero-work
no-op. The outer process-group kill moves to budget+1s: the scan's own
deadline enforces the budget, and the kill is a backstop for a scan wedged
in an unbounded wait instead of a racer that routinely preempts the clean
bounded exit. The wake-lock-wait test bound tracks the backstop (3s -> 4s);
the previously flaky assertion is unchanged.

* no-mistakes(document): Document inactive-reconcile deadline backstop

* doc: Revise Firstmate delegation and communication guidelines

Refactor the guidelines for Firstmate's role and delegation process, emphasizing the importance of crewmates and asynchronous work.

* doc: Update work delegation and secret management instructions

Clarified guidelines for handing off work to crewmates and managing secrets.

* test(procevent): make the process-event suite's detached-runner assertions deterministic (#2617)

Three assertions in tests/fm-procevent.test.sh depended on a detached runner
having finished work that the command starting it does not wait for.

reconcile's replacement runner is started through detach_runner, which only
forks: reconcile returns and counts the start before that runner has claimed
its source or exec'd its child. Any assertion taken straight after reconcile
therefore samples a race.

- The publish-before-apply recovery section left its always-ready /bin/echo
  source registered across the recovery reconcile, so that reconcile launched
  a competing detached poll (observed: started=1) that then raced every later
  assertion for the source claim, the next capture sequence, and this home's
  applied record, and outlived the section holding a live claim. It is now
  retired before that reconcile - re-announcement is proven from the durable
  inbox alone and needs no registration - and started=0 is asserted so a
  competing poll cannot be reintroduced unnoticed. This is the same
  retire-before-reconcile discipline the self-announcing section already
  carries; that section acquired it after the identical race made its
  "not-autohandled: self-src" assertion read "already owned: self-src".

- The crashed-leader replacement section snapshotted the replacement's claim
  file and execution log behind a fixed 0.5s settle window. On a loaded
  machine that window expires first, which is the CI flake behind "a
  replacement runner started without recording its own claim" and "reconcile
  did not start exactly one replacement source". Both effects are now waited
  for with the suite's bounded wait helpers; the exact one-replacement count
  is still asserted afterwards, unchanged.

- The duplicate-start section slept 0.5s for reconcile's runner to record
  ownership before asserting that a second start loses to it. It now waits
  for that claim.

Also tighten one assertion that could not fail as written: "autohandled:
self-src" is a substring of "not-autohandled: self-src", so the applied path
was accepted even when the runner reported the capture left for the handler.

Evidence: on the unmodified suite, 128 full runs at 6-8x concurrency produced
6 failing runs, all in the crashed-leader section. On the fixed suite, 216
full runs under the same load produced none. Reverting the self-announcing
section's retire-before-reconcile line reproduces "already owned: self-src"
on the first iteration, confirming the shared mechanism.

* fix: preserve pending replies and defer remote reposts (#2618)

* fix(bin): keep pending-reply expectations honest on both send legs

Two related asymmetries let the parent-owned secondmate reply guard drop or
nag requests it should not have.

Local delivered-unconfirmed dropped the expectation. A marked request whose
submit read-back stayed unconfirmed (verdict=pending) is the same
not-a-failure outcome the remote leg reports as delivered, but fm-send
discarded the parent's pending-reply record for it, so a request that very
likely landed stopped being tracked entirely. The record now stays armed on
its unconfirmed-delivery marker: a correlated report still resolves it, and
an unanswered one still surfaces through the library's own reconciliation.
Exit 3 and the local rule that an unconfirmed answer never closes a decision
key are unchanged.

Remote replies were nagged for a repost they did not need. A remote mate's
report reaches the parent's status log only through the asynchronous mirror
in fm-procevent-remote-reply.sh, yet the guard read an absent correlated
line as proof the mate never reported - even while the answer was still in
flight, which is the common case because the mirror's poll window is
comparable to the recovery grace. The mirror now publishes one caught-up
watermark from a quiet window, and the guard admits a missing report as
evidence only once that watermark passes the turn that should have produced
it. A genuinely missed report still gets exactly one repost, and a channel
that is behind, unarmed, or broken leaves the request durably open and
un-nagged rather than nagging blind; the mirror escalates its own continuity
failures as before.

Tests: a local unconfirmed secondmate send keeps its expectation armed and
resolvable; a mirrored correlated remote reply resolves with no repost; a
stale or absent watermark withholds the repost while a fresh one still
releases it; a quiet remote window publishes the watermark and retirement
clears it.

* no-mistakes(review): Distinguish preempted polls from quiet windows

* no-mistakes(document): Clarify remote reply channel freshness

* no-mistakes(lint): Annotate shared remote preemption exit constant

* fix(bin): honor declared pauses in busy-pane wedge checks (#2619)

* fix(watch): honor a declared pause on a busy pane's completed-turn bound

A worker that declares an external wait (`paused:`) and then blocks in one
long foreground call - a review-hosting scout parked in a single blocking
`lavish-axi poll`, a bounded watch loop, a rate-limit sleep - keeps its pane
BUSY, so the stale path that already honors declared pauses never ran for it.
The busy-pane completed-turn bound instead routed it straight into
wedge_timer_check, which re-escalated "possible wedge, escalation N" (and, past
the threshold, demand-deep-inspection) every FM_STALE_ESCALATE_SECS for as long
as the review stayed open.

busy_turn_bound_check now owns which absorber takes a crossed bound: a crew
whose own last status line declares an external wait or a verified captain-held
transfer takes the bounded FM_PAUSE_RESURFACE_SECS recheck, and everything else
keeps the unchanged wedge timer. The discriminator is the declaration together
with liveness (the caller has already confirmed the pane is busy), never a
blanket silencing - a crew that declared nothing, or whose pane is not live,
escalates exactly as before, and a declared pause still re-surfaces once per
long cadence so a forgotten wait cannot rot invisibly. Away mode is untouched:
the daemon owns pause triage there and already reads the same vocabulary.

The two call sites also no longer clear pause bookkeeping in the same poll the
pause cadence recorded it, which would have erased the re-surface throttle and
turned the long cadence back into a per-poll re-surface.

Tests: a three-phase regression fixture pins the absorbed pause, its long-cadence
recheck, and the restored wedge escalation once the declaration is lifted on the
same busy over-age pane.

Also de-flakes tests/fm-watch-triage.test.sh, which failed spuriously on a loaded
machine: fixed liveness budgets were reaping watchers mid-startup, so assertions
on post-poll state passed vacuously or failed spuriously. Waits that describe a
poll's outcome now wait for a completed poll cycle via the liveness beacon, the
heartbeat test waits for the heartbeat it asserts on, and every wait_for_exit
budget is the uniform 10s already used elsewhere in the file.

* no-mistakes(review): Fail poll-cycle waits on timeout

* no-mistakes(review): Prevent poll timeout test hangs

* no-mistakes(document): Clarify paused busy-pane supervision

* doc: Enhance communication guidelines for decision-making

Added guidelines for decision communication to the captain.

* doc: Update task delegation and communication guidelines

Clarify communication protocols with crewmates regarding task delegation and reporting.

* fix(bin): reliably confirm herdr steer submission (#2647)

* fix(herdr): confirm local steers that native agent-state misses

Herdr can leave agent_status idle for a landed Claude turn and can keep
queued Enter text visible while busy, so fm-send was reporting false
swallows. Confirm those cases through the shared queued-Enter verdict
and a cleared composer, and keep a genuine idle pending composer as
unconfirmed.

* no-mistakes(review): Stop Herdr Enter retries on unreadable composers

* no-mistakes(review): Reject queued delivery when all Herdr Enter sends fail

* no-mistakes(review): Prevent confirmation after failed Herdr Enter

* no-mistakes(review): Pace Herdr retries and clarify submit fallback

* no-mistakes(review): Align Herdr submit docs with idle fallback

* no-mistakes(document): Correct Herdr submit-confirmation documentation

* feat(bearings): add interactive Lavish fleet board (#2659)

* feat(bin): accept any-origin decision bindings with full-identity keys

An aggregation surface (the bearings board) carries captain answers for holds
across origins, but a binding was one-origin-per-source and the Lavish adapter
capped question keys at 64 chars while real full hold identities measure 69-81.

- fm-decision-hold.sh: bind <source-id> --any-origin records the (any) marker;
  binding prints it verbatim and answers accepts it, so the runner's feed seam
  carries an any-origin source with no runner change. In any-origin mode each
  key is a full hold identity <origin>-decision-<key>, split at its first
  -decision-; a key with no separator (merge/dispatch instructions) is skipped
  and feeds nothing, keeping non-decision answers out of the hold ledger by
  construction. Every existing close guard applies unchanged.
- fm-procevent-lavish.sh: raise the question-key cap 64 -> 128 so a full hold
  identity fits; the slug-shape security property is unchanged.
- tests: cross-origin closure through the real runner seam, an 81-char
  identity through the adapter, cap and shape refusals, routed-work skips,
  nonexistent-identity skips, and idempotent replay.

* feat(bearings): add the /bearings lavish interactive fleet board

/bearings lavish renders the bearings snapshot onto a shipped, reusable board
template and arms it as a Lavish process-event source, so the captain answers
Captain's Call items on the board and firstmate is woken by an ordinary check
wake - no conversational turn ever blocks on a poll.

- .agents/skills/bearings/assets/board-template.html: the shipped template
  (myfirstmate design system inlined, one fm-bearings-board.v1 JSON slot,
  fail-closed schema guard that renders an error card instead of an empty
  fleet). Per-invocation agent work is composing the payload only.
- bin/fm-bearings-board.sh: build/refresh owner - fail-closed payload
  validation, slot injection with a round-trip check and \u003c escaping,
  stable board path, any-origin bind ALWAYS before arm, arm-if-absent.
- bearings SKILL.md: the lavish invocation option, board composition rules,
  board-wake handling, and the captain-ruled merge-click authorization with
  its mandatory safeguards (PR resolved from the task's own meta record,
  wake-time green re-verification, never a red or changed PR, merges only
  through bin/fm-pr-merge.sh, chat echo with the full PR URL).
- process-event-sources SKILL.md: one-line board-wake routing trigger.
- tests: payload refusals, injection round-trip, bind-before-arm, idempotent
  re-arm, and template slot integrity.

Fleet pickup: homes receive this after merge plus a firstmate self-update;
landing timing is coordinated with the main firstmate.

* no-mistakes(review): Harden bearings board validation and wake handling

* no-mistakes(review): Require HTTPS for bearings board PR links

* no-mistakes(review): Fail closed and bound bearings board answers

* no-mistakes(review): Enforce UTF-8 byte limits for board answers

* no-mistakes(review): Serve bearings board before arming and reject empty actions

* no-mistakes(review): Prove bind-before-arm ordering through live answer consumption

* no-mistakes(document): Document bearings board and cross-origin answers

* fix(bearings): restore decision options and add close controls (#2707)

* fix(bearings): always show decision options and a close/drop control

Freeform-only Captain's Call cards hid the option buttons the board was designed around, and there was no way to drop a stale hold without inventing an answer. Require selectable options, keep freeform as a supplement, and route the reserved __drop__ answer through decline so the hold leaves Captain's Call.

* no-mistakes(review): Fix drop closure and decision-only option validation

* no-mistakes(review): Preserve answerability for non-decision cards

* no-mistakes(document): Clarify decision drop documentation

* ci: require no-mistakes pipeline step attestation (#2710)

Signature-only PRs can hide skipped review, test, or document steps. Fail unless no-mistakes >= 1.46.0 attests those three steps completed.

* feat: collapse decisions into tasks held for the captain (#2728)

* feat(captain-hold): collapse the decisions concept into tasks held for the captain

A decision is no longer a separate type: it is an ordinary backlog task held
for the captain, identified by its task id. bin/fm-captain-hold.sh owns the
surviving behaviors - guarded hold creation, the recorded-answer close
(answer/answers with a release mode for captain-gated work), the source
bindings, and the investigation completion gate - and bin/fm-decision-hold.sh
becomes a one-release compatibility shim over it.

The fleet snapshot now parses hold-until and computes captain_actionable as
queued + captain-held + unblocked + due, independent of row kind, plus a
presentation-only deferred_marker for prose-deferred rows. Bearings renders
every due captain-held task in Captain's Call, date-deferred holds as dated
Charted Next gates, suppresses prose-deferred rows from default views with an
omitted disclosure, and excludes from Recently Landed anything that closed
while still held for the captain.

Legacy compatibility: pre-collapse <origin>-decision-<key> rows are already
plain task ids and keep working; short keys in recorded metadata, concrete
origin bindings, chat --resolve-key fallbacks, and old resolution records all
resolve in place.

* no-mistakes(review): Fix captain answer replay and body preservation

* no-mistakes(review): Fix captain hold idempotency and legacy replay

* no-mistakes(review): Validate card close modes and compatibility routing

* no-mistakes(review): Enforce release replay mode matching

* no-mistakes(review): Prevent duplicate decision cards and released replay mismatches

* no-mistakes(review): Preserve answer columns and legacy resolve replays

* no-mistakes(document): Document strict replay and legacy compatibility

* no-mistakes(lint): Quote done literals to satisfy ShellCheck

* no-mistakes: apply CI fixes

* fix(rebase): keep collapsed captain hold board semantics

* fix: bound recovery announcements and preserve supervision (#2733)

* fix(watch): announce recovery once per generation and keep successors supervising

A lost Pi/OpenCode handling handshake re-announced the same recovery
generation on every cycle and spent the successor's first ~55s blind, so
a real crew event could be ignored and then dropped. Record the
announcement in the durable marker, confirm the handshake before the
follow-up without swallowing failure, and enter the poll loop immediately.

* no-mistakes(review): Tighten recovery event timing regression

* no-mistakes(document): Document recovery-loop supervision guarantees

* fix(bin): surface captain-call record divergence (#2744)

* fix(bin): signal a captain call resolved in the log but still held

A captain call has two records and closing one has never closed the
other: a `resolved [key=...]` line closes the status-log fold, while the
backlog task held for the captain closes only through
`fm-captain-hold.sh answer`. Answering on the status side alone left no
trace of the disagreement - the fold went quiet, the durable record kept
saying the captain owed an answer, and nothing warned. The defect was
never the separation; it was the silence.

Add `fm-captain-hold.sh diverged`, a read-only report of that
contradiction, and print it from `fm-wake-drain.sh` as a bounded RECORD
DIVERGENCE section beside OPEN DECISIONS on every drain. It flags one
condition: a task still open and still carrying the captain-hold
annotations whose key was closed on the status side by the resolve verb,
under the collapsed identity or the legacy derived one.

It closes nothing, ever. A captain call closed wrongly leaves review
entirely, which is worse than the noise, so both reconciliation
directions stay human-owned and the printed hint names both - a
resolution is not proof the captain ruled, since a call can dissolve on a
false premise or turn out to have been a question of fact.

Three states are deliberately not divergence: a `captain-held` close is
the verified transfer `complete` writes, a still-open keyed decision
belongs to the OPEN DECISIONS fold, and a captain call with no routed
work item is legitimate rather than incomplete, so routed work is no part
of the test.

`fm-classify-lib.sh` gains `status_key_closing_verb`, which reports how
the status side currently reads one key by replaying the existing
`_fm_decision_fold_line` rule rather than re-deriving it, so the two
closing verbs stay distinguishable in one place. The per-wake cost is one
`tasks-axi list`, one key scan per status log, and the precise per-key
fold only for a key that already names a still-open task; the call is
hard-bounded so a slow backlog tool can never delay wake presentation.

* fix(document): Correct divergence lifecycle documentation

* fix(document): Neutralize divergence lifecycle prose

* fix(bin): re-arm after an abandoned auto-arm claim and defer a wedge escalation while a worktree is written (#2524)

* fix(watch): re-arm supervision after an abandoned auto-arm claim

A Claude auto-arm cycle that armed, delivered one rewake, and exited left
its single-flight lock behind. Both Stop-event participants then deferred
to that lock forever, because its recorded pid was still live: the
turn-end guard read it as recovery under way and allowed the stop, and the
next Stop firing treated it as another owner and declined to arm. On
2026-08-14 a home with two tasks in flight lost supervision for about 40
minutes with no watcher process and no watcher lock, its beacon frozen at
the one delivery, and both crewmates' finished reports sat in the durable
queue until an operator drained it by hand.

Abandonment is now proven from the epoch ledger instead of inferred from
pid liveness. A lock whose holder pid matches the ledger's own owner_pid
while the recorded outcome is anything other than arming has already
finished its decision, so that claim is reclaimed under the lock's steal
mutex, stops counting as recovery ownership in the guard, and is cleared
by the guard's terminal check rather than deferred to. A failed clear
re-blocks instead of allowing a blind stop, and an arming entry stays in
flight however old it is, because its owner foregrounds the arm for the
whole watcher cycle.

Issue #2251's PR #2263 does not cover this failure. It is closed and
unmerged, lives entirely in bin/fm-watch-arm.sh, and retires the stalled
watcher and matching stale watcher lock of an arm that is currently
running. Here no arm and no watcher were running and no watcher lock
existed, so it has nothing to retire and the home stays blind.

tests/fm-claude-stop-autoarm.test.sh covers the reclaim, the still-arming
and unnamed-owner cases that must keep the gate closed, and the failed
clear. tests/fm-turnend-guard.test.sh covers the guard side of the same
boundary. Both fail without this change.

* fix(watch): defer a wedge escalation while the task worktree is written

The wedge detector had two inputs, rendered pane quietness and the run
step, and neither can see a crew that is writing source, then tests, then
documentation behind a static pane. On 2026-08-14 one crewmate produced
eight consecutive possible-wedge escalations in a single afternoon, three
of them demanding deep inspection, while it was demonstrably working and
then committed. Every one of them cost a supervision turn to disprove by
hand.

Add write activity inside the crew's own recorded worktree as a third
liveness input. crew_worktree_written_since compares the worktree against
the caller's existing idle-window timer file, so -newer needs no clock
arithmetic, no temp file, and no portable mtime write. The probe runs only
inside the branch that was about to escalate, which bounds it to one
pruned, depth-bounded walk per window per FM_STALE_ESCALATE_SECS and
leaves the per-poll stale sweep exactly as cheap as before.

Positive evidence defers rather than cancels. The idle timer restarts so
the next window probes again, the escalation counter is neither advanced
nor reset so a later genuine wedge keeps the demand-deep-inspection
history it earned, and a .writing-since marker ages the whole deferral
chain so the pane still re-surfaces once per FM_PAUSE_RESURFACE_SECS,
through the same throttle shape a declared pause already uses, labeled as
a recheck rather than a wedge. This can only reduce false positives: every
absence of evidence, including no recorded worktree, a torn-down worktree,
a missing anchor, and a failed walk, falls through to the unchanged
escalation schedule, so a crew that writes nothing still escalates on the
existing timetable.

What the signal cannot see, by design or by construction:

- CPU burn with no writes, such as a long compaction, is invisible. That
  case keeps the old behavior exactly.
- A commit-only phase writes only .git, which is pruned first so that
  firstmate's own read-only git commands against the worktree can never
  make the probe self-fulfilling.
- Writes under the pruned generated trees, or deeper than
  FM_WORKTREE_WRITE_MAXDEPTH, do not count.
- The probe cannot attribute a write to the crew, so a background build or
  another process touching the tree looks the same. The hourly re-surface
  is what bounds that, and a churny file cannot buy silence.
- The away-mode daemon's own escalation path is deliberately untouched.

tests/fm-watch-triage.test.sh covers the classifier including the .git
prune, both halves of the live case on one fixture (quiet plus writing
defers, quiet plus silent still escalates and counts), and the bounded
re-surface. All three fail without this change.

* no-mistakes(review): prove autoarm claims by identity; skip mate-home write probe

* no-mistakes(document): document away-mode wedge boundary and probe filesystem limit

* no-mistakes(document): qualify turn-end recovery condition for abandoned auto-arm claims

* fix(watch): keep a write deferral scoped to its own idle window

Two consistency gaps in the worktree write probe, both found while reviewing
the wedge-deferral change on this branch.

A write deferral is a bounded chain: its .writing-since marker ages the whole
chain so a churning worktree still re-surfaces once per resurface window. That
is only sound while the chain belongs to the current quiet stretch, so every
path that restarts the idle-window timer has to drop it too. Two did not: the
corrupt-timer repair in wedge_timer_check, and both first-sight branches for a
captain-relevant status. A chain left over from an earlier quiet stretch made
the first deferral of the new window re-surface immediately instead of after a
full fresh window.

FM_WORKTREE_WRITE_PRUNE is a skip list, so clearing it reads as "skip nothing"
and is the obvious way to widen the probe to the whole depth-bounded tree.
Instead an empty list reported no evidence at all, quietly costing the wedge
detector its third liveness input on a home that meant to widen the walk. An
empty list now widens the walk, and the header says so.

Neither change alters when a stall that writes nothing escalates.

Regressions in tests/fm-watch-triage.test.sh cover all three paths and each
one fails on the pre-fix code.

* no-mistakes(review): honor an empty write-prune, bound the probe, share window_key

* no-mistakes(document): align probe knob count and guard regression-coverage ownership

* no-mistakes(lint): silence deliberate single-quote SC2016 in write-prune env test

* fix(bin): give a captain hold the same bounded pause cadence as a declared pause (#2748)

* fix(bin): give a captain hold the same bounded pause cadence as a declared pause

Two supervisors read a finished task's last status line and disagreed about which
declarations mean an idle endpoint is expected. bin/fm-inactive-reconcile.sh
suppresses its inactive-outcome scan only on `captain-held`, while the away-mode
daemon's wedge path gated deferral on `paused` alone. Both read the LAST line, so
the two verbs are mutually exclusive and no finished task waiting on a person
could satisfy both at once. Marking 11 such tasks `captain-held:` silenced the
900s outcome scan and immediately produced five possible-wedge escalations in one
batch, because the 240s wedge detector no longer saw a pause verb.

fm-classify-lib.sh's status_is_paused_or_captain_held already owns the combined
question, and bin/fm-watch.sh's ordinary-crew wedge path already asked it. This
extends that same answer to the paths still asking the narrower one:

- bin/fm-supervise-daemon.sh, all six sites, which form one subsystem and have to
  move together. classify_stale returns the pause action, reconcile_pause_tracking
  and migrate_watcher_pause_markers record and migrate the marker, and
  housekeeping defers the wedge and then re-surfaces the recheck. Changing only
  the stale-persistence gate would defer the escalation while
  reconcile_pause_tracking recorded nothing, so the wedge marker would persist and
  the sweep would `continue` past it forever: quiet, but never re-surfacing.
- bin/fm-watch.sh's secondmate stale gate, whose downstream owner
  pause_state_class already treats both declarations identically.
- bin/fm-push-transition-lib.sh's absorb, where either declaration already names
  the human the transition would report and the wait is already durably recorded.

Quieting alone would be half a fix, so the bounded re-surface had to reach a hold
too. A hold has no current-state mapping, unlike `paused`, so authoritative crew
state reports it as unknown and pause_state_class received `none`. An ordinary
crew recovers pause classification from that state through confirmed agent death,
which proves no live decision gate is being silenced. A secondmate's endpoint
liveness is deliberately never read there, because an idle mate is healthy by
design, so that confirmation is unavailable by construction and cannot be
required: without recovering the classification for a mate, every caller silenced
a held mate outright and its hold would rot invisibly. That promotion is bounded
by the declared-wait guard at the top of the function, so it can only reclassify a
task that already declared a wait and shows no positive working evidence.

Two narrow `status_is_paused` calls are deliberately left alone.
bin/fm-crew-state.sh's map_log_state is a current-state reporting contract, not a
wedge path; reporting a hold as `paused` would erase the distinction
status_key_closing_verb and fm-captain-hold.sh depend on, where a `captain-held`
close is a verified durable transfer and a `resolved` close claims outright
settlement. fm-classify-lib.sh's call inside status_is_captain_relevant needs no
change because that function's own case list already returns non-relevant for
`captain-held`.

bin/fm-inactive-reconcile.sh keeps its `captain-held` suppression as it is. Its
guard exists because a finished task's crew state still reports done from a
higher-priority source than the log, and a declared pause needs no such guard: the
scan only reports done or failed, and nothing else reaches its record path.
Widening it would change a separate subsystem's reporting contract, which this
defect does not require.

Coverage extends the existing colocated patterns for these predicates and asserts
both halves. tests/fm-daemon.test.sh covers the classification, the wedge marker
converting to pause tracking with no escalation, the bounded re-surface with its
window reset, and the boundary case where an answered hold stops claiming the
cadence. tests/fm-watch-triage.test.sh covers a held secondmate re-surfacing on
the same bounded cadence without being labeled a wedge.
tests/fm-supervision-events.test.sh covers the absorbed push transition. Every one
of these fails on the pre-fix code except the answered-hold boundary case, which
is there to pin that the quieting was not widened too far.

The `paused:` workaround appended to those 11 tasks is live supervision state and
is untouched here. It can be retired once this lands.

* no-mistakes(review): name the captain in a held task's bounded recheck

* no-mistakes(document): extend declared-wait supervision docs to captain-held holds

* fix(bin): make lint prerequisites and harness tests reliable (#2758)

* fix(lint): name the installer when ShellCheck or actionlint is missing

A missing actionlint exited 127 like a bare command-not-found. Fail with
exit 1 and point at the pinned installer, matching the missing-ShellCheck
path, without weakening the version pin.

* test: isolate kimi and muse detection from inherited Cursor markers

Harness detection checks CURSOR_AGENT before ancestry, so these
markerless-adapter cases failed when the suite itself ran under Cursor.
Clear the verified markers the same way the secondmate harness tests already do.

* no-mistakes(document): Document Muse Cursor marker cleanup

* feat(bin): report watched tooling updates that are available or installed but inert (#2684)

* feat(checks): report tool updates that are available or installed but inert

Firstmate had no way to notice that tooling this home depends on needs an
update, and no way at all to notice the worse case: an update that installed
correctly and then did nothing.

That second case is why this exists. A tool that self-installs into
~/.local/bin while a version manager keeps its own older copy earlier on PATH
looks completely up to date to anything that asks only "is a newer version
published". On 2026-08-20 a Herdr update landed at 0.8.2 while an older 0.8.0
copy stayed earlier on PATH, so every Herdr command failed on a protocol
mismatch and firstmate could not read its own fleet.

bin/fm-tool-update-check.sh reports the two conditions separately:

  <tool> update available      a newer version exists at the update source.
  <tool> update not in effect  a newer copy is installed on this host, but
                               PATH still resolves an older one.

PATH skew is measured, never inferred. Every executable copy of a watched
command on PATH is asked for its own version and those answers are compared,
so one lookup cannot hide the skew, and a directory name is never read as a
version because a version manager's "latest" directory can hold an older
build. A copy that will not report a version is a check failure, not a pass.

The watched tools live in local, gitignored config/watched-tools.json, so
adding a tool is a config edit rather than a code change, and the file is
never propagated to another home. Update sources cover both shapes: a local
clone's commit distance from its remote branch, and a command's own version
and update announcement, including a tool like no-mistakes that prints its
version on one command and announces a new release on another.

The check prints one line when something needs attention and prints nothing
otherwise, so it rides the existing watcher state-check contract with its
trust binding instead of introducing a schedule of its own, and
state/.tool-updates keeps the same pending update from being reported on
every poll.

The check only reports. It never installs, updates, reorders PATH, touches a
version manager, or fetches into a watched repository; every git probe is
read-only.

Tests cover the skew case as a regression, and it was verified by mutation:
removing the skew report, or stopping after the first PATH hit as a single
lookup would, each make that test fail.

* no-mistakes(review): fix tool update check probe reporting, budget, and shim write

* no-mistakes(review): keep sweeps alive on broken patterns and oversized budgets

* no-mistakes(review): roll back failed arm, widen budget clamp, bound repo probe

* no-mistakes(review): guard git probes at the budget, record uncut findings

* no-mistakes(document): fix stale watched-tool report-record wording in docs and header

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

The behavior shard's watch-triage suite failed on the new worktree-write wedge
tests. Those five tests are the only ones in the file that do not use its
standard waits. They give a fixed 3 second liveness budget to the one poll that
now spawns the bounded worktree walk, and 4 seconds to an escalating watcher
where every other test in the file gives 10. On a loaded runner that poll
outlives the fixed budget, so the round is reaped before the deferral it asserts
on is recorded, and the test reports a lost deferral instead of the deferral
under test. Wait for a completed poll cycle through the file's own
wait_poll_cycle, which is what its header documents this hazard for, and use the
file's standard 100 tick exit budget.

Verified against a load that reproduces the failure: 11 of 12 runs failed
before, 8 of 8 pass after. Verified by mutation too, so the waits still prove
the behavior: removing the write deferral, and keeping a finished deferral chain
across an idle-timer repair, each still fail their test.

* fix: decouple ask-user decisions from yolo (#2764)

* fix: treat yolo as merge authority only, not ask-user finding authority

Yolo on/off was documented as also deciding no-mistakes ask-user findings, which hid firstmate's duty to judge unambiguous-toward-design findings itself. Keep every safety boundary; this is a contract clarification, not a relaxation.

* no-mistakes(document): Clarify yolo documentation ownership and merge posture

* feat(bin): add a spoken interface that answers from records and hands work over (#2767)

* feat(voice): spoken round trip on Nova Sonic 2 with a measured relay cost

Step one of the spoken interface: the laptop captures and plays audio, this
desktop holds the model session, and no AWS credential leaves the desktop.

Measured, amazon.nova-2-sonic-v1:0 in eu-north-1, end of speech to first byte
of reply audio, 6 runs each, all answered, on a question that forces a records
read:

  relay path   1.229 1.379 1.428 1.447 1.481 1.516  median 1.438
  direct       1.147 1.179 1.203 1.237 1.244 1.317  median 1.220

The relay costs about 0.22s of the median. The direct figure reproduces the
earlier survey, which is what makes it a usable control. Excluded: the
captain's own ssh round trip, microphone capture, and speaker output. This
desktop has no microphone and no speaker, so every run used audio files.

Three pieces:

  bin/fm-voice-relay.py    holds the conversation on this host
  bin/fm_voice_records.py  what a spoken answer may read, and the handover
  bin/fm-voice-client.py   the laptop end; audio devices UNVERIFIED
  bin/fm_voice_frame.py    the wire format both machines share

Real work is handed to the existing bin/fm-inbox.sh rather than a second
queueing surface, and the agent says it is handing over rather than answering
as firstmate.

Read scope: Done history and free-form note bodies are never assembled at any
scope, so the wide default cannot reach the places commercial detail
accumulates. config/voice-read-scope narrows it to counts only, and
config/voice-read-deny excludes a named item in one line. The boundary is an
executable test that widening the reader fails.

Push to talk is the default because it is cheaper and the choice is still open;
--listen open-mic is the single flip.

Two traps worth knowing: a clip with no trailing silence is never answered, and
the end of a reply is contentEnd with stopReason END_TURN, not completionEnd.
A second user turn in one session is treated as barge-in unconditionally, and
an interrupted turn that calls a tool is lost, so the session reconnects per
turn and gives up conversational memory. That is the concrete thing step three
has to solve.

* no-mistakes(review): fix voice relay credential reuse, frame validation and record parsing

* no-mistakes(review): test uplink header guard, bound unknown expiry, align state dir

* no-mistakes(review): decide deny per item, guard turn failures, bound ambient credentials

* no-mistakes(review): read account config from home, harden deny and turn failures

* no-mistakes(review): close status verb set, fix inbox help, pair data override

* no-mistakes(review): keep profile-free relay alive, unblock loop, fix dead assertion

* no-mistakes(review): hide finished pull requests, refuse open mic, keep suite offline

* no-mistakes(review): survive reader failures, release devices, fix claims

A failure while handling a model event, or while sending a tool result,
left the reader task dead with ended and turn_done clear, and close()
re-raised the stored failure on every await. One dropped stream became a
relay that could never build another session. The reader now reports the
session over in a finally whatever killed it, and close() absorbs the
task the same way it already absorbed its sends.

The laptop client releases what it already started when a later startup
step refuses, SystemExit from the handshake wait included, and names a
device refusal instead of leaking a raw PortAudio error. Whether it
releases correctly against a real device is still unverified here.

The records docstring claimed every reading was filtered to open ids.
Only the pull request count and list are; the worker count and the state
histogram cover every live runtime record, finished ids included,
because a meta file still on disk still needs tearing down.

The finished-work deny half of the suite asserted things that held with
the deny list absent. It is replaced by a deny on an open title, which
removes the row and says so while the count stays honest.

* no-mistakes(review): name reader failures, split file and device refusals

A failure inside the model reader released the waiting turn and told
nobody. The session was not marked spent, no notice reached the client,
and the client waits for a reply end or a notice, so the captain got
their whole timeout of silence and then a record saying the turn went
unanswered with nothing about why. Both ends of the relay now name a
failed turn through one function, once per turn, and --self-test carries
the cause in relay_error the way the client's own record does.

Two things that are not failures stay that way. A stream that simply
ends is the end of a session, which serve still reads on its own terms.
A stream that goes away because close() asked it to is an ordinary
renew, and announcing it would have put a failure notice in front of the
captain on every turn.

On the laptop end, the refusal that became a device error covered the
file-backed playback and capture too, so a mistyped --in-file was
reported as an audio device failure and the advice named the flag that
had just failed. The file ends now report the path and the flag that
chose it and stay an OSError; the device ends keep the device advice and
name the flag for that end. The device paths remain unrun here, so only
the file halves are covered by a test.

* no-mistakes(test): survive model session end, order client turn frames

* no-mistakes(document): sync voice relay docs with reviewed relay behavior

* no-mistakes(document): re-measure relay latency and correct its cause

* no-mistakes(document): correct measurement date and name the unmeasured SSH hop

* no-mistakes(document): describe the unpublished control measurement, fix list formatting

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* fix(bin): preserve Relay follow-up loops until explicit disposition (#2763)

* fix: keep Relay public loops open until retire

Delivering a promised-final reply was deleting the only record that tied a public thread to later work, so a follow-on ship silently owed no closing reply. Retain the registration after delivery, rechain follow-on work onto the same thread, and make retire --reason the only close.

* no-mistakes(review): Propagate public follow-up registration removal failures

* no-mistakes(review): Persist retire receipts and align parent resolution

* no-mistakes(review): Make rechain resumable after partial obligation creation

* no-mistakes(review): Repair follow-up state, briefs, and expiry escalation

* no-mistakes(review): Serialize follow-up delivery stamps with retirement

* no-mistakes(review): Serialize rechain claims and protect registration terminal states

* no-mistakes(review): Avoid reporting retired delivery loops as open

* no-mistakes(document): Refresh public-loop documentation and verification evidence

* no-mistakes: apply CI fixes

* no-mistakes(review): Preserve delivered follow-up bindings during registration replay

* no-mistakes(review): Harden public follow-up retirement and rechain races

* no-mistakes(review): Fail closed on unresolved secondmate retirement

* no-mistakes(review): Bind secondmate cleanup to its recorded canonical home

* no-mistakes(review): Fix rechain command output and expiry validation

* no-mistakes(review): Validate brief keys and warn on remote promotion

* no-mistakes(document): Document retained public follow-up loops

* no-mistakes(lint): Remove unused bounded-wait loop variable

* feat(bin): merge GitLab merge requests through the guarded PR merge path (#2779)

* feat(bin): merge GitLab merge requests through the guarded PR merge path

bin/fm-pr-lib.sh already parses a GitLab merge request URL for the watcher,
but bin/fm-pr-merge.sh refused every non-github provider, so a merge request
had to be merged by hand and got none of the recording, guards, or audit
trail a pull request gets.

The merge path now dispatches on the parsed provider. A GitHub URL keeps its
exact previous behavior. A GitLab URL is addressed through glab by the project
URL rebuilt from the parsed host and path, so a merge request on any instance
resolves and no host is hardcoded, and no merge-method flag is added because
the project's own merge method is what should apply.

A GitLab merge happens only after one live read of the merge request confirms
it is open, detailed_merge_status is mergeable, has_conflicts is false,
blocking_discussions_resolved is true, and the head pipeline succeeded at the
exact current head. Every failing condition is reported, not just the first.
The verified head is bound to the merge with glab's --sha, so a push landing
between the read and the merge fails the merge instead of landing commits
nothing verified. Recorded metadata is never the authority for any of this: a
rebase moves the head and leaves a recorded value stale, so a recorded head
that disagrees with the live one is reported rather than trusted, and the
recorded value is read before the recording step because that step drops a
GitLab head it cannot resolve.

* no-mistakes(review): reject bundled -R clusters and make tool-absence cases host-independent

* no-mistakes(test): state authorised GitHub narrowing of bundled -R guard

This branch NARROWS GitHub behaviour. The narrowing was authorised
deliberately rather than slipping in by accident, and it applies to both
providers, GitHub and GitLab alike, because a script that guards one provider
and not the other is a trap for the next reader.

What bin/fm-pr-merge.sh now refuses is extra merge arguments containing a
bundled short-option cluster that includes R, for example "-dR other/repo".
The forge CLIs expand such a cluster one character at a time, so it carries
"--repo other/repo", and that later value wins over the repository the URL
named. Before this change, "fm-pr-merge.sh <task> <github-url> -- -dR
other/repo" reached "gh-axi pr merge 12 --repo example/repo --squash -dR
other/repo" and exited 0 with pr= recorded and the merge poll armed. It now
exits 1 with "extra merge arguments must not override the repository", records
nothing, and invokes no forge merge command. Every other GitHub invocation is
byte-identical to the base commit.

Closing that hole honours the existing rule rather than departing from it. The
file header already forbids --repo and -R because the repository must come
only from the URL, so a bundled cluster carrying a repository override was
never legitimate behaviour to preserve: it was that guard being evaded.
Redirecting a merge to a repository the URL does not name is exactly what the
guard exists to prevent.

The refusal is already pinned on both paths by the existing case
test_bundled_repo_override_args_refuse_before_recording in
tests/fm-pr-merge.test.sh. On GitHub ("-dR wrong/repo") and on GitLab ("-yR
https://other.example/g/p") it asserts exit 1, the refusal wording, no pr= in
the task meta, no armed merge poll, and no forge merge command invoked, with a
control case proving a cluster that carries no repository override still
reaches the forge. No duplicate assertion was added. Both assertions were
confirmed to have teeth by narrowing the guard back to a bare -R and watching
each path fail.

This commit carries no file change: the guard and its coverage landed in
614853d, and this message exists so the pull request description states the
narrowing.

* no-mistakes(document): fix README pointer for GitLab watch and merge doc

* no-mistakes: apply CI fixes

* fix(bin): record a lost relay connection instead of an unanswered turn (#2788)

* no-mistakes: apply CI fixes

* fix(bin): drop a private record citation and narrow the review rule

Three corrections to the spoken interface that landed in #2767, plus one
fix carried over from that branch after its pull request had already been
merged.

The confidentiality fix. The module docstring of bin/fm-voice-relay.py
cited a private, gitignored fleet record by exact path and section number.
That widens what this public repository points at, and it cannot resolve
for any reader here, because the path has never been in the repository.
Both traps it pointed at are already described in full in the list
immediately below it, and docs/voice-relay.md carries the same two for
operators with no citation at all, so the pointer is removed and no claim
is weakened by losing it. Two comments that referred to "the survey" as
though it were something a reader could open are reworded the same way.
Neither exposed a path, so that half is comprehensibility rather than
confidentiality.

The review rule. .greptile/rules.md is kept, because its conditions are
right and deleting it would leave the next reviewer to re-litigate a
decision already argued out. What was wrong with it is narrower than its
existence: it read as settled repository policy, when whether VISION.md
itself should be reconciled is an open question belonging to the captain.
One sentence now says so, and says that the conditions listed below it are
what the interpretation depends on. That narrows the claim rather than
widening it.

The carried-over fix. The first commit on this branch is 7f98e797 from
fm/voice-relay-build-v4, taken verbatim rather than rewritten. It closes
the window where a transport failure was recorded and then erased, so a
run could be emitted as answered false with relay_error null. That matters
more than it looks: relay_error is the field that keeps an infrastructure
failure from being averaged into a latency figure, so the failure mode is
a dead connection wearing the costume of a slow reply. It landed fifteen
minutes afte…
lytv pushed a commit to lytv/mymate that referenced this pull request Sep 8, 2026
A wedged family-run step was occupying the runner until the 75-minute
job cap; bound that step so cleanup and timing artifacts still upload.
cipherholdingsllc added a commit to cipherlab-ai/firstmate that referenced this pull request Sep 11, 2026
…w scoring (#7)

* fix(cmux): classify borderless Claude composers (#2029)

* fix(cmux): classify borderless Claude composer

* no-mistakes(review): Normalize cmux NBSP prompts across locales

* no-mistakes(document): Document cmux borderless Claude composer classification

* docs(stow): generalize read-before-write in the public stow skill (#2091)

The public installer-facing stow skill scoped its classify-then-replace
discipline to TODO/BACKLOG items only, so findings routed to a memory file
had no stated rule against a blind append or a wholesale overwrite.

Step 6 now classifies every finding against the destination's current
contents as new, duplicate, superseding, or obsolete, and states the
considered replacement each classification implies. The outcomes follow the
tiered-memory contract already in the file: an obsolete entry is refreshed,
archived, or replaced in a way that preserves its fact, a duplicate folds
into the entry that already carries it, and a superseded body worth keeping
leaves through step 7's existing exits rather than a second recovery
mechanism.

* fix: resurface durable supervision work after re-arm (#2065)

* fix(watcher): resurface durable work after downtime

* no-mistakes(review): Make watcher rearm recovery durable and cursor-safe

* no-mistakes(review): Persist safe recovery markers across migration lock recovery

* no-mistakes(review): Retain stale lock when recovery marker publication fails

* no-mistakes(review): Preserve delivery-gap recovery and quarantine malformed markers

* no-mistakes(review): Serialize recovery consumption and report acknowledgment failures

* no-mistakes(review): Centralize recovery publication before clearing watcher evidence

* no-mistakes(review): Guarantee recovery evidence across queue and lock handoffs

* no-mistakes(review): Publish recovery evidence before durable wake commits

* no-mistakes(review): Replace recovery marker Perl dependency with Node

* no-mistakes(review): Keep interrupted wakes durable until handling acknowledgment

* no-mistakes(review): Add post-handling durable wake acknowledgements

* no-mistakes(review): Enforce post-handling acknowledgement across recovery and AFK return

* no-mistakes(review): Bind wake acknowledgements to recovery generations

* no-mistakes(review): Align wake regressions with generation-bound acknowledgements

* no-mistakes(document): Document durable re-arm recovery semantics

* no-mistakes(lint): Resolve ShellCheck warnings in recovery and watcher tests

* no-mistakes: apply CI fixes

* test(watcher): assert post-handling wake replay

* no-mistakes(review): Prevent successor loops and adopt legacy wake generations

* no-mistakes(review): Rearm durable wakes without recursive successor recovery

* no-mistakes(review): Align recovery tests with handling marker state

* no-mistakes(review): Delay handling transition until successor launch is established

* no-mistakes(review): Confirm wake handling only after successful prompt delivery

* no-mistakes(review): Acknowledge AFK wakes only after evidence publication

* no-mistakes(review): Prevent AFK wake loss before post-handling acknowledgement

* no-mistakes(document): Document durable wake acknowledgement semantics

* no-mistakes(lint): Suppress false positive for recovery action output

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* ci: measure Herdr automation on Windows runners (#2100)

* ci: add Windows Herdr automation spike

* ci: run Windows spike on its pull request

* fix: wait for Windows Herdr command output

* fix: run ANSI probe in pane shell

* ci: keep Windows Herdr spike manually triggered

* docs: clarify Windows Herdr spike verdict

* feat(ahoy): guide captains through open decisions (#2099)

* Add guided ahoy decision flow

* no-mistakes(document): Document guided Ahoy decision flow

* fix(stow): enforce startup-memory budget decisions (#2110)

* Harden stow memory budget policy

* Refine internal stow offload policy

* no-mistakes(review): Enforce shared-budget decisions and autonomous offload

* fix(spawn): refresh pooled worktrees from origin before launch (#2116)

* fix(spawn): refresh pooled worktree base

* no-mistakes(document): Document spawn base-freshness invariant

* no-mistakes: apply CI fixes

* fix(composer): unify safe classification across backends (#2102)

* refactor(composer): one shape owner behind thin capture adapters, whole matrix fixed

Consolidate every composer shape - bordered boxes (all families, geometry,
titled bottom borders), bare agent-glyph rows and their wrap regions,
opencode's left bar, and pi's identity-gated separator pair - into
fm_composer_classify_screen in bin/fm-composer-lib.sh. Adapters now
contribute only a capture and a declarative capability descriptor
(styled/cursor/identity/rows); capability differences change how confidently
a shape is judged, never what the shapes are, so a new harness shape is
teachable in exactly one place.

Correctness fixes landed as part of the consolidation (audit
data/fm-composer-consolidation-audit-s1):
- locale-safe Unicode-space normalization in the shared owner (closes the
  fleet-wide half of #1988; cmux's local byte-exact NBSP case deleted;
  naming converges with PR #1995's normalization primitive)
- muse's bare glyph joins the shared set, unbreaking muse on herdr/cmux/orca
- orca learns the borderless bare shape, drops its backward-paged composer
  window, and can no longer classify a stale startup banner as the composer
- tmux tolerates a titled bottom border, unbreaking grok steering
- the left-bar shape makes opencode readable on every backend
- zellij gets a real classifier through dump-screen --ansi, replacing the
  content-diff submit heuristic that could confirm an undelivered message
  and close a --resolve-key decision (the fleet's only false positive)
- fm-spawn's kimi launch-readiness regex (the fourth shape copy) now routes
  through the shared classifier

The strict blank-row posture applies fleet-wide (captain decision
blank-row-injection-posture): no positive container proof = unknown = defer,
replacing tmux's permissive blank-cursor-row rule. Away-mode injection was
re-validated end to end on real tmux (defer on partial input and unproven
rows, clean delivery with swallowed-Enter retry into proven-empty
composers). The tmux submit core gains a baseline-idle turn-started
conversion so pi steering stays confirmed while its working screen hides
the composer; busy conversion without that baseline remains forbidden.

Plain-capture backends now degrade a glyph row carrying trailing text to
unknown instead of a false pending, per the approved capability rule.

Portable regressions pin the full byte-capture matrix from the audit under
a UTF-8 locale and LC_ALL=C, the strict-vs-permissive divergence, and
deliberate signal separation; the opt-in live guard
(tests/fm-composer-matrix-live-e2e.test.sh) verified every installed
harness against the real classifier, recorded in
docs/verification/runtime-backends.md.

* no-mistakes(review): Fix Pi glyph ambiguity and complete profile matrix

* no-mistakes(review): Preserve bare verdict when Pi identity probe is absent

* no-mistakes(review): Harden composer structure and titled-border geometry

* no-mistakes(review): Require proven idle baseline and strict Zellij guard

* no-mistakes(review): Reject box bottom borders as composer input rows

* no-mistakes(review): Prove Zellij probe typing before classifier retries

* no-mistakes(review): Preserve Pi identity uncertainty and scan full left-bar drafts

* no-mistakes(review): Verify Zellij text lands before submitting

* no-mistakes(review): Scope Zellij typing verification to selected composer content

* no-mistakes(review): Verify Zellij pastes through composer-scoped content deltas

* no-mistakes(review): Prove wrapped bare Zellij pastes through composer extraction

* no-mistakes(review): Invalidate stale cursorless composers below dead shell prompts

* no-mistakes(review): Handle shell prompt placeholders in composer extraction

* no-mistakes(review): Classify cursorless bare continuation regions safely

* no-mistakes(review): Reject stale cursorless containers below live activity

* no-mistakes(review): Preserve prompt glyphs in wrapped Zellij pastes

* no-mistakes(review): Reject live shell rows during composer extraction

* no-mistakes(review): Preserve wrapped glyph continuations through submit retries

* no-mistakes(review): Scope idle placeholders to proven positions

* no-mistakes(review): Restore boxed placeholders and live prompt reanchoring

* no-mistakes(review): Fix Zellij placeholder and wrapped glyph paste proof

* no-mistakes(document): Align composer architecture documentation

* no-mistakes(lint): Fix ShellCheck warnings in composer refactor

* no-mistakes: apply CI fixes

* docs(verification): record the trusted-checkout live matrix rerun

The pipeline's isolated gate worktree is untrusted, so claude, grok, and
muse stopped at first-launch trust dialogs there (the guard refuses to
confirm them by design). This rerun from the trusted checkout at the final
validated head verified all six installed harnesses, the strict blank-row
deferral, and the hardened zellij false-positive probe live.

* no-mistakes(document): Align composer verification evidence

* no-mistakes: apply CI fixes

* no-mistakes(review): Restore proven box bottom-cursor classification

* no-mistakes(review): Preserve styled placeholder-like drafts as pending

* no-mistakes(document): Align composer safety and Zellij delivery documentation

* no-mistakes: apply CI fixes

* docs(verification): refresh the live matrix with the final-head trusted rerun

The post-validation rerun from the trusted checkout verified all six
installed harnesses at the branch's final head, including Claude 2.1.227
(auto-updated since the audit's captures) and Grok, which the untrusted
gate worktree could not verify past their first-launch trust dialogs.

* fix(spawn): gate Pi TUI mode by CLI capability (#2117)

* fix(spawn): gate Pi regular TUI flag by capability

* no-mistakes(review): Document conditional Pi TUI capability detection

* no-mistakes(review): Pin Pi probing and launch to one executable

* no-mistakes(review): Preserve literal pinned Pi paths and update documentation

* no-mistakes(review): Defer pinned Pi path insertion until final substitution

* no-mistakes(document): Document version-safe Pi launch probing

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* docs(vision): elevate experience, pain narrative, and distro virtues (#2147)

* docs(vision): elevate experience, pain narrative, and distro virtues

Fold the captain's public vision framing into VISION.md: peace of mind as a
primary goal, multi-session context-switch pain as the problem one interface
solves, clone-and-run setup ease, self-evolution including community, and
explicit harness/backend orthogonality. Reconcile experience-as-garnish into
experience-as-purpose and update aligns/resists accordingly.

* docs(vision): state the experience goal positively

Drop the negative "not a smart workflow / useful tool / impressive technology"
pretext. Lead straight into the positive experience north star.

* feat(bin): reconcile inactive terminal crew outcomes (#2167)

* fix: reconcile inactive terminal outcomes

* fix: stream secondmate summary inputs

* no-mistakes(review): Fix reconciliation locking and request delivery retries

* no-mistakes(review): Prevent retries after unknown request delivery

* no-mistakes(document): Clarify inactive reconciliation cadence and receipts

* no-mistakes(lint): Quote terminal status arguments in reconciliation tests

* refactor: simplify inactive outcome reconciliation

* no-mistakes(review): Bound inactive reconciliation scans with durable progress

* no-mistakes(review): Bound reconciliation and deduplicate recovery notices

* no-mistakes(document): Document inactive outcome reconciliation contracts

* no-mistakes(review): Reject relative local secondmate parent routes

* no-mistakes(review): Key terminal receipts by spawn incarnation

* no-mistakes(review): Stabilize legacy receipts and lock reconciliation snapshots

* no-mistakes(review): Fail closed on invalid secondmate identity markers

* no-mistakes(document): Document durable inactive-outcome reconciliation

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* ci: raise Herdr test timeout (#2191)

* fix: refresh stale Pi instructions after compaction (#2163)

* fix(session-start): refresh drifted instructions on stale rebuilds

* test(session-start): prove Pi instruction refresh end to end

* no-mistakes(review): Fix stale instruction refresh and baseline integrity

* no-mistakes(review): Preserve true-start baselines across Pi continuations

* no-mistakes(review): Correct Pi continuation classification and live expectation

* no-mistakes(review): Correct Pi continuation coverage documentation

* no-mistakes(review): Fix read-only refresh and exact Pi session restores

* no-mistakes(review): Classify Pi create-if-missing sessions correctly

* no-mistakes(review): Classify named Pi sessions using immutable headers

* no-mistakes(review): Correct Codex interactive coverage diagnostic

* no-mistakes(document): Document immutable Pi compaction instruction refresh

* no-mistakes(document): Correct Pi refresh documentation and validation claims

* feat: add deterministic condition-to-action watcher (#2200)

* feat(bin): add deterministic condition->action watch adapter on the process-event channel

Register a (condition, action) pair once with bin/fm-procevent-when.sh and the
existing process-to-event runner polls the condition tokenlessly, fires the
action at most once on a stable true, and wakes firstmate exactly once with the
captured outcome - instead of burning an agent turn per re-check.

The pair is stored privately under state/when/ and hash-bound by a trust record
the same way fm-check-register.sh binds a custom check, so a mutated spec is
refused without executing anything. A durable exclusive fired marker claimed
before the action makes restarts and re-polls unable to double-fire; every
failure path (mutated spec, condition error past budget, expired deadline,
failed action, uncaptured earlier fire) ends in a terminal captured outcome
that wakes firstmate rather than a silent retry. Eligibility stays a firstmate
judgment: only exact, safe, reversible actions may be bound, and judgment-
needing or destructive actions keep the wake-and-decide flow.

* no-mistakes(review): Harden when watcher concurrency, deadlines, timeouts, and output

* no-mistakes(test): Bind watcher actions to registered executable bytes

* no-mistakes(document): Correct condition-action watcher documentation

* no-mistakes(document): Clarify outcome wake re-announcement

* no-mistakes: apply CI fixes

* fix(bin): honor a decision key stated after the verb colon (#2202)

The open-decisions fold only recognized a [key=<slug>] token between the
verb and the colon (needs-decision [key=x]: note). The common worker
shape with the colon first (needs-decision: [key=x] note) silently
folded its stated key into the shared "default" bucket, so two open
decisions could collapse into one record and fm-send --resolve-key <x>
refused to close the decision it plainly named.

A complete token at the head of the note is now an equivalent stated-key
position for every keyed verb, shared by the whole-file and incremental
folds through the one _fm_decision_key owner. The documented
before-colon position wins when both are present, a token deeper in the
note stays prose, a bare keyless line still folds to "default", and a
stated-but-malformed slug is rejected rather than rewritten to
"default". A consumed note-head token is stripped from the note so both
positions yield identical records, and the incremental fold version is
bumped so persisted cursors folded under the old interpretation are
rebuilt from the authoritative log.

Fixes #2109

* fix(bin): prevent watcher recovery acknowledgement livelock (#2212)

* fix(bin): keep a recovery acknowledgement valid across republication

A watcher cycle that opened and closed while the model handled its drained
wakes minted a fresh recovery generation, which invalidated the exact
acknowledgement the drain had just printed. That acknowledgement then consumed
nothing, so the marker stayed pending and every later arm spent its whole cycle
re-announcing the same recovery instead of supervising - a livelock the home
could not leave on its own.

A downtime publication now reuses the generation of an outstanding handling
episode, so a close during the handling window cannot orphan the printed
acknowledgement. The acknowledgement itself separates its two facts: queue-row
consumption is bound to the monotonic --ack-through sequence and always
happens, while only retiring the episode is bound to --recovery-generation. A
generation that moved on is a non-fatal result that names its own remedy
instead of a refusal that consumes nothing.

* no-mistakes(review): Preserve recovery generations and consume stale acknowledgements safely

* no-mistakes(document): Document sequence-bound recovery acknowledgements

* feat(fmx-respond): consume Relay conversation chains (#2206)

* feat(fmx-respond): consume in_reply_to_chain conversation context

The relay's poll payload can carry in_reply_to_chain, an oldest-first
transcript of the surrounding conversation, but the mention-handling
procedure only ever read the immediate in_reply_to parent, so referents
like "this" in a standalone mention stayed unresolvable even when
context was delivered.

Teach fmx-respond to read the chain when present (optional and
backward-compatible: often absent today, kind label not required),
resolve referents against the whole transcript, and extend the
untrusted-content framing to every chain entry including the upcoming
kind=history entries. Document the field's wire shape in
docs/configuration.md as the firstmate-side owner.

* no-mistakes(document): Document Relay chain context ownership

* fix: parse decision verbs before status metadata tags (#2280)

* fix(bin): strip every bracket tag, not just [key=...], from a status verb

status_line_verb only stripped a leading "[key=...]" token before the
colon, so a remote secondmate reply's leading "[corr=...]" correlation
tag stayed glued onto the returned verb word ("needs-decision
[corr=...]" instead of "needs-decision"). The open-decisions fold's
verb match then silently failed to recognize the line at all, so
fm-send --resolve-key refused to close a decision that was plainly
open on the status line.

Generalize the parser to strip every "[name=value]" tag before the
colon, in any order and count, so local and remote replies fold
identically.

* no-mistakes(review): Invalidate stale decision cursors after parser fix

* no-mistakes(document): Clarify status metadata verb parsing

* fix(bin): collapse duplicate supervision wakes (#2287)

* fix: collapse duplicate supervision wakes without losing legitimate updates

One remote-secondmate note produced two handling turns (a procevent check
wake published before autohandle, then a signal wake for the same mirrored
bytes), already-ingested replays such as a cursor-loss whole-log recapture
still woke with nothing to do, this home's own bookkeeping closes (fm-send
--resolve-key, the pending-reply escalation close, the captain-held
transfer) re-woke the session that wrote them, and turn-ended-only wakes
were annotated with already-announced status lines that looked like fresh
progress.

Dedup rules, each at its layer's one owner:
- fm-procevent.sh: an adapter may declare 'self-announcing'; the runner
  then applies first and publishes a check wake only for what remains
  unhandled. fm-procevent-remote-reply.sh declares it: the mirrored status
  append is the single announcement, so a fully applied capture publishes
  nothing and a byte-identical replay stays completely quiet. All other
  adapters keep strict publish-before-apply.
- fm-wake-lib.sh: fm_wake_signal_sig/seen_path/seen_current now own the
  watcher's signal signature and .seen-* marker format, plus
  fm_wake_status_append_self_announced, the guarded bookkeeping append
  that advances the marker only over exactly its own bytes and fails
  toward waking on any pending or interleaved foreign write.
- fm-send.sh, fm-pending-reply-lib.sh, fm-decision-hold.sh: bookkeeping
  closes go through that guarded append; escalation opens stay plain
  appends because a new blocker must wake.
- fm-wake-lib.sh annotations: a historical (turn-ended-only) row skips its
  status annotation only when the file's signature provably matches the
  seen marker; anything unannounced keeps annotating.
- fm-classify-lib.sh: a kind=secondmate task's status signal is never
  absorbed as provably-working, because that stream is the routed-reply
  channel the parent must read.

Also fixes a pre-existing exit-path deadlock the regression run reproduced:
a TERM inside a recovery-marker critical section left fm_lock_try_acquire
spinning against this same process's abandoned hold; a self-held lock is
now reclaimed (a subshell still waits on its parent's live hold).

Regression tests drive the real wake functions and executables in both
directions: each duplicate case collapses, while a new remote reply, new
decision, new blocker, merge result, failure, first status change, and a
later different note on the same task all still wake.

* no-mistakes(document): Document wake deduplication contracts

* feat: add Cursor CLI crew harness (#2238)

* feat(harness): add Cursor Agent CLI adapter

# Conflicts:
#	bin/fm-spawn.sh

* fix(composer): read cursor-agent's reverse-video placeholder as idle

cursor-agent renders its idle composer placeholder dim (SGR 2) but paints the
cell under the terminal cursor in reverse video (SGR 0;7). Reverse video is
neither dim nor a dark truecolor foreground, so the shared ghost stripper keeps
that one character and an idle composer reduces to a lone `P`. Judged on its
own, that remnant reads `pending` on a genuinely idle pane, which defers
away-mode escalation indefinitely on the styled cursorless backends.

Teach the ONE fleet-wide classifier the shape instead of adding an adapter-local
copy: register `→` as an agent prompt glyph so the composer row is structurally
findable at all (without it the bottom-most shape is a stale shell prompt echo
in the scrollback), add both verified placeholders to the idle set, and consult
the styling-independent plain row when the styled row is only a remnant.

The plain-row branch demands the remnant be a proper, strictly shorter substring
of a plain row matching a fully anchored placeholder. Real typed text is
uniformly bright, so stripping leaves it equal to the plain row and it stays
`pending` - verified live against a pane where the typed text was exactly the
placeholder string.

Verified live on cursor-agent 2026.08.11-e8db854; the regression pins the real
captured bytes and asserts the remnant survives stripping, so the case cannot go
vacuous if the stripper later learns SGR 7.

Co-authored-by: Amplify Logic AI <lars@sockinator.co>

* feat(cursor): narrow cursor identity and order its marker before CLAUDECODE

Cursor ships two executable names - `cursor-agent` and the legacy alias `agent`
- and runs as a bundled node script, so tmux reports the pane command as a bare
`node`. Neither `agent` nor `node` can be trusted by name, so identity gets one
owner in bin/fm-cursor-lib.sh that demands cursor's own name or install tree in
the path or argv[0], from the structural signal only. Probing an arbitrary pid's
executable during a liveness poll would execute a stranger's binary, which is
the hazard that rule exists to close.

Two consequences wired up:

Detection. cursor-agent does NOT clear an inherited CLAUDECODE, so a cursor
worker launched under a claude primary carries both markers and whichever is
tested first wins. The cursor markers are ordered ahead of the CLAUDECODE check;
fm-spawn additionally clears foreign markers at the launch boundary. Both are
kept deliberately - launch sanitization only covers sessions fm-spawn started,
while the ordering also covers a cursor session started by hand. Verified live
that CURSOR_INVOKED_AS is set on the agent process and CURSOR_AGENT=1 on the
child/tool processes fm-harness.sh actually runs as.

Pane liveness. A cursor pane now classifies `agent`. An unrelated node or agent
stays `other`, which the liveness callers already fold into `ambiguous` rather
than `dead`, so a stranger's node pane is never reported agent-free.

Resolution prints the STABLE launcher rather than the canonical target: identity
is proven through canonicalization, but cursor's canonical path carries a
version its own auto-update replaces, and pinning that would strand a task on a
version that can vanish.

The regression drives the two identity signals apart - a cursor-named executable
outside any cursor tree, and a non-cursor-named alias inside one - and asserts
each carries a verdict alone, so no single vendor string is load-bearing. Its
negative controls are real spawned processes, not fixtures.

Verified live on cursor-agent 2026.08.11-e8db854.

Co-authored-by: Ville Penttinen <villem.penttinen@gmail.com>

* feat(cursor): classify cursor busy state from its own turn transcript

Cursor shipped as "unknown cursor-unverified" on the premise that it exposes no
semantic turn lifecycle, only a rendered "Working" footer. That premise is
wrong: cursor-agent persists an append-only JSONL transcript per conversation
and brackets every submitted turn with a role:user open and a typed turn_ended
close. Verified live on 2026.08.11-e8db854, including the interrupt path, where
Escape closes the turn with status "aborted" - so this source covers manual
interruption, which Claude's Stop hook does not.

That makes it a genuine pull source in the muse mould rather than the rendered
text the redesign forbids: no writer, no arm, no gen, nothing seeded that could
never be cleared. Cursor's `ctrl+c to stop` footer stays out of the verdict, and
herdr's narrower native streaming state cannot stand in for it either.

Binding deliberately does not reconstruct cursor's workspace-slug directory
name. That slug collapses path separators, so rebuilding it would be a guess
that could bind the wrong pane; cursor records the exact absolute workspace path
in each project's .workspace-trusted, and the binding matches on that. A
conversation recorded as prior at spawn is excluded, so a relaunch in a reused
worktree folds its own turn rather than its predecessor's. Requiring a unique
remaining conversation keeps zero and several both unknown, because neither
proves anything about the current turn.

The regression pins the fold with real transcript files and asserts the
dangerous direction stays closed: an unresolvable binding, a record-free file,
an unclaimed workspace, and a workspace-path PREFIX all read unknown, never
idle. The prefix case uses an opaque fixture slug so a slug-rebuilding
implementation cannot pass it.

Co-authored-by: Ville Penttinen <villem.penttinen@gmail.com>

* feat(cursor): make the cursor launch runnable and give it lifecycle control

Five gaps that together kept a cursor crewmate from being drivable end to end.

Launch. The template invoked `cursor agent`, but `cursor` is not the CLI - the
installed names are `cursor-agent` and the legacy alias `agent` - so the command
could not run at all on a machine with a normal cursor install. It now resolves
through the verified owner, which also refuses a spawn loudly instead of leaving
a pane that dies with command-not-found and reads as a wedged worker.

Session binding. fm-spawn writes state/<id>.cursor-session so the busy fold can
find this pane's transcript, and teardown removes it.

Lifecycle control. No cursor PR touched fm-control-lib.sh, so
`fm-control <id> interrupt|exit|relaunch` could not drive a cursor worker at
all. Verified live: interrupt is a single Escape, exit is /exit, and cursor does
NOT repollute its composer with the cancelled prompt, so unlike muse it needs no
clear key. Secondmate is refused, matching the spawn refusal.

Submit acknowledgement. cursor parks its terminal cursor outside its composer,
so the composer verdict on tmux is always `unknown` and a submit could never be
acknowledged from the composer alone. The submit core's existing idle-to-busy
transition covers that case, but only if the pane's busy footer is recognised,
so cursor's `ctrl+c to stop` joins the harness-less default union the submit
cores read. The TOKEN is matched rather than the spinner verb: the same version
rendered both `Working` and `Running` in consecutive turns.

Bootstrap. A configured cursor crew harness with no cursor executable is now a
loud MISSING diagnostic rather than a first-spawn failure, and it accepts either
installed name.

Interrupt cancellation is deliberately left unconfirmed. The transcript does
type an aborted close, but its post-interrupt write latency measured as
variable - sometimes seconds, sometimes not within twenty - so a claim built on
it would be unreliable. Normal turn completion is prompt, which is what the busy
fold actually depends on.

Two inherited tests are corrected rather than deleted: the busy test asserted
cursor could have no semantic source, and the launch test pinned the literal
`cursor agent` string. Both now pin the verified behaviour, including that the
launch never allocates a second worktree.

Co-authored-by: ABHISHAKE KUMAR BOJJA <abojja@uvic.ca>
Co-authored-by: Ville Penttinen <villem.penttinen@gmail.com>

* docs(cursor): record the verified crewmate facts and extend the drift guard

The inherited cursor entry was written against 2026.08.04-aaa8809 and several of
its claims no longer hold: it named `cursor agent` as the binary (not the CLI
name), listed six Grok model ids of which the live catalog now returns two, and
recorded busy state, exit, interrupt, and skill invocation as unverified.

Replaced with what was measured against 2026.08.11-e8db854, including the two
facts most likely to be rediscovered painfully: cursor runs as a bundled node
script so its pane title is a bare `node`, and it parks its terminal cursor
outside its composer, which makes the tmux composer verdict permanently
`unknown` by design rather than a defect to chase.

Model ids now route to `--list-models` for the account instead of a fixed list,
since that list is exactly what drifted.

The live drift guard covers cursor, resolving it through the same verified owner
fm-spawn uses and passing --trust so the probe cannot hang on the workspace
prompt. Run against every installed harness: 8 checked, all alive, with cursor
reporting title='node' foreground=[.../cursor-agent] - the drift shape this
guard exists to catch.

Co-authored-by: Ville Penttinen <villem.penttinen@gmail.com>

* docs(agents): record the cursor session-binding state file

The state/ layout section is the inventory every session reads; a busy-source
binding that fm-spawn writes and teardown removes belongs in it alongside muse's.

* no-mistakes(review): Sanitize ambient Cursor marker in harness tests

* no-mistakes(review): Validate Cursor models against live catalog

* no-mistakes(review): Reject unsupported secondmates before binary preflight

* no-mistakes(review): Narrow Cursor ancestry detection to structured process identity

* no-mistakes(review): Parse Cursor transcripts and sanitize inherited markers

* no-mistakes(review): Handle malformed Cursor transcript records safely

* no-mistakes(review): Validate malformed Cursor closes in fallback parser

* no-mistakes(review): Retire stale Cursor bindings during relaunch

* no-mistakes(review): Fix Cursor drift guard command variable

* no-mistakes(review): Narrow Cursor identity to versioned install trees

* no-mistakes(document): Document Cursor harness boundaries

* refactor(composer): move the delivery busy footers to the shared owner

The per-harness rendered busy footers lived in bin/fm-tmux-lib.sh under
FM_TMUX_* names, so cursor's `ctrl+c to stop` signature - and every other
harness's - was reachable only from tmux. That placement was wrong on its own
terms: herdr, zellij, cmux, and orca run the same harnesses and face the same
question these footers answer, which is whether a submitted Enter actually
landed. Nothing about the signature is tmux-specific.

Moved verbatim into bin/fm-composer-lib.sh, the shared composer/delivery owner
every backend already sources, and renamed to FM_DELIVERY_* so the names stop
claiming a scope they never had. All five adapters now reach cursor's signature;
verified per adapter rather than assumed.

The boundary the move must not blur is stated where it now lives: this is a
DELIVERY guard, never a worker-state source. Confirming a keystroke landed is a
different question from asking what a worker is doing, and bin/fm-busy-lib.sh
remains the semantic owner that forbids classifying a harness from rendered
text. Cursor still classifies only from its transcript fold, which is already
backend-agnostic because it folds a file rather than reading a pane - the same
verdict on all six backends.

The old FM_TMUX_* aliases are dropped rather than kept as dead shims: nothing
outside the moved block referenced them except fm-busy-lib.sh's grok fallback,
which now reads the new name. The documented operator override, FM_BUSY_REGEX,
is untouched.

Also removes a dead duplicate CURSOR_INVOKED_AS check in bin/fm-harness.sh,
unreachable behind the marker check above it.

* no-mistakes(review): Correct shared delivery guard ownership references

* no-mistakes(document): Document shared delivery guards and Cursor backend limits

* no-mistakes: apply CI fixes

* fix(composer): bound a bare composer's wrap region at a half-block rule

A live cursor crewmate on herdr classified its IDLE composer as `pending`, and
fm-send consequently exited 1 with "delivery unconfirmed" on a message that had
actually landed. The cause is not cursor-specific.

Herdr draws a composer's top and bottom rules with the half-block glyphs U+2584
and U+2580 rather than the box-drawing family. fm_composer_row_has_edge knew
only the box-drawing set, so no box was detected; the composer was found as a
BARE row, and its wrap region - which extends while rows are non-blank and carry
no structural edge - walked straight through the composer's own closing rule and
swallowed the model and path footer below it. That footer is real text, so the
region classified pending on a genuinely idle pane.

Teaching the shared edge detector the half-block glyphs bounds the region at the
closing rule. Measured on the captured bytes of a real herdr cursor pane: the
same capture that read `pending` now reads `empty`.

This is a shared shape-path change, so it is deliberately narrow - it adds
glyphs to the edge vocabulary and changes no verdict logic - and the whole
composer and backend suite is green, including the other harnesses' herdr
fixtures.

The regression pins the real captured shape and asserts the footer content is
genuinely present, so the case cannot pass vacuously if the region were ever
bounded for some unrelated reason.

* fix(herdr): confirm a cursor submit from the rendered-footer transition

Herdr's composer-shape fix made an idle cursor pane classify `empty`, but
`fm-send` still exited 1 with "delivery unconfirmed" on messages that had
actually landed. Live measurement found the second, independent cause.

Herdr reports a cursor pane `agent_status=blocked` in EVERY state - idle,
mid-turn, and after - so the submit path's idle-baseline native confirmation is
structurally unreachable for cursor and every send falls into the composer
branch. That branch reads cursor's mid-turn composer row, which renders its own
`Add a follow-up` placeholder beside a right-aligned `ctrl+c to stop`. That
token is composer content, so the verdict is `pending` on a composer holding no
user text at all, and the Enter-retry budget then reports pending.

The escape is the same semantic signal the native path uses, read from the
pane's verified busy footer instead of native agent-state, and it is the
rendered-footer twin of the tmux submit core's turn-started confirmation: an
idle-to-busy transition ACROSS our Enter proves the harness accepted the
submission. The baseline is taken before the first Enter and only when the
native baseline was not legibly idle, so the idle-baseline path still never
reads pane content and a pane already mid-turn before we typed keeps reporting
`pending` rather than borrowing another turn as proof of this delivery.

The composer verdict is deliberately NOT relaxed. A right-aligned status token
on the composer row stays content for every other caller, including the
away-mode pre-injection guard, and the shared cursorless submit core is left
untouched so zellij, cmux, and Orca keep the behavior their own follow-up owns.

Verified live on herdr 0.8.0 and cursor-agent 2026.08.11-e8db854 in an isolated
lab session: `fm-send` now exits 0 and the steer executes, interrupt cancels a
running turn, `/exit` stops the agent, and teardown clears the record. All seven
panes of the running default session classify identically before and after the
shape fix, so no other harness regressed.

* no-mistakes(review): Prevent working Herdr baselines from falsely confirming delivery

* no-mistakes(document): Correct Cursor harness and backend documentation

---------

Co-authored-by: ABHISHAKE KUMAR BOJJA <abojja@uvic.ca>
Co-authored-by: Amplify Logic AI <lars@sockinator.co>
Co-authored-by: Ville Penttinen <villem.penttinen@gmail.com>

* fix(bin): require quota-axi 0.1.25 (#2300)

* fix: raise quota-axi floor to 0.1.25 for Cursor CLI quota awareness

Homes on latest main need quota-axi #87 so Desktop-absent CLI machines report a fresh Cursor quota instead of a false sign-in-required.

* no-mistakes(document): Update quota floor documentation pointer

* fix(bin): prevent false Pi watcher alarms during hand-offs (#2304)

* fix(guard): stop the false send-time watcher-down alarm on Pi primaries

On a Pi primary the watcher process is not the liveness signal. The Pi
extension tears the watcher down on every actionable wake and spawns the
replacement itself, so the singleton lock is legitimately unheld between
cycles: every one of the 799 cycles in a live primary's ledger ends with
lock_after=pid:none, and a live capture caught the guard verdict flipping to
no-watcher during one hand-off with the beacon 63s old.

bin/fm-guard.sh classified Pi as a persistent-watcher harness, which demands a
live identity-matched lock holder at all times, so any guarded command landing
in a hand-off painted the full WATCHER DOWN - SUPERVISION IS OFF banner and
told firstmate to repair a cycle the extension already owns and is restoring.

Add an extension supervision model for pi and pi-signed. A live
identity-matched watcher stays the ordinary healthy state; an unheld lock is
healthy only while the beacon is fresh within grace AND a live Pi session
provably owns continuity - both primary extensions recorded in their state
markers at their current on-disk builds by the process named in state/.lock,
with that process still alive. Without that proof the banner fires exactly as
before, so an unloaded, version-drifted, or exited Pi session is loud
immediately and a cycle the extension never restores is loud once the beacon
passes grace. The queued-wake warning, the PID-strict turn-end guard, and
every other primary's detection are untouched.

Fold session-start's duplicate Pi marker predicate into the shared library so
the ownership contract has one owner.

* no-mistakes(review): Restrict Pi hand-off tolerance to unheld watcher locks

* no-mistakes(document): Document Pi watcher hand-off supervision

* feat: support Cursor Agent CLI as a primary harness (#2305)

* feat(cursor): add Cursor Agent CLI primary hooks, park supervision, and session start

Register a tracked project-scope .cursor/hooks.json for Cursor's stop,
sessionStart, preCompact, and preToolUse steps.

bin/fm-turnend-guard-cursor.sh owns Cursor's turn boundary as a park: it
foregrounds the watcher arm, holds the boundary open until an actionable close,
and returns that wake as one follow-up. Exit 2 is a silent no-op on Cursor's
stop step, so the adapter never uses it. The follow-up loop is bounded twice,
by Cursor's own loop_limit and by the payload's loop_count.

bin/fm-sessionstart-cursor.sh delivers the digest as additional_context at
sessionStart, and stages it for the next turn boundary at preCompact, which
cannot inject context.

Cursor also loads the tracked Claude settings, so bin/fm-hook-host-lib.sh lets
each tracked Claude-shaped entrypoint stand down on a Cursor-delivered payload
rather than running every covered event twice.

bin/fm-tmux-lib.sh reclassifies a Cursor pane's composer cursorlessly, because
Cursor parks its terminal cursor outside the composer, which restores a genuine
composer-empty proof and unblocks away-mode escalation delivery.

* feat(cursor): make Cursor Agent CLI a verified primary harness

Resolve Cursor in the session-lock ancestry through bin/fm-cursor-lib.sh, which
a Cursor primary needs before it can hold its own home lock, and classify its
stop-hook park under the autoarm supervision model so the mid-turn pull guard
stops reporting a healthy between-turns watcher as down.

Read a Cursor pane's composer cursorlessly on tmux, gated on Cursor's own
structural process identity, which restores a genuine composer-empty proof and
lets away-mode escalations reach a Cursor primary with no daemon change.

Lift the secondmate refusals in bin/fm-spawn.sh and bin/fm-control-lib.sh now
that the supervision protocol exists and is recorded.

Cover the whole surface with a portable regression over real processes, an
opt-in live guard against the installed cursor-agent, and dated per-harness
evidence.

* docs(cursor): record Cursor as a verified primary across the owning surfaces

Update the turn-end guard, session-start, arm-seatbelt, cd-guard, watcher
continuity, architecture, configuration, README, and harness-adapters owners,
and add dated live evidence to the supervision and runtime-backend verification
records. Correct the recorded Cursor tmux composer verdict: the cursor-anchored
read is still blind, but the composite reader is no longer unknown.

Lift the remaining remote-secondmate refusal missed in the previous commit, and
add the new libs to the existing fixtures that copy a fixed dependency list.

* refactor(cursor): name the park's stand-down condition for both its causes

Also record that Cursor's preCompact firing itself is not yet live-verified,
while the static evidence that it cannot inject context, and the staging path
that follows from it, both are.

* test: give the pretool fixtures their new dependency and one lint owner

The cd-guard fixture copies a fixed dependency list and now needs the shared
hook-host predicate. Both pretool suites also asserted cleanliness with a bare
shellcheck call, a second and weaker copy of the lint definition that
bin/fm-lint.sh owns: it omits --external-sources, so it failed the moment these
checkers sourced a shared library. They now delegate to that owner.

* test: assert the cursor secondmate contract instead of its removed refusal

A cursor secondmate now launches, so the suite asserts what its park actually
needs: --trust so the home's project hooks load at all, its own home pinned as
the workspace, and the autoarm supervision model inherited across the launch.

* no-mistakes(review): Serialize Cursor wakes and bind staged context

* no-mistakes(review): Serialize Cursor context and nag state commits

* no-mistakes(review): Enforce Cursor ceiling before staged context delivery

* no-mistakes(review): Serialize Cursor claims and staged context

* no-mistakes(review): Serialize Cursor ownership and state commits

* no-mistakes(review): Protect Cursor context across session takeover

* no-mistakes(review): Preserve Cursor context across session takeover

* no-mistakes(review): Enforce owner-keyed Cursor staged context

* no-mistakes(review): Atomically claim Cursor follow-ups and staged context

* no-mistakes(review): Defer Cursor preCompact staging and simplify supersession

* no-mistakes(review): Serialize Cursor park commits and defer preCompact

* no-mistakes(review): Stop Cursor parks after session takeover

* no-mistakes(test): Route Cursor preCompact context through stop follow-up

* no-mistakes(document): Update Cursor primary documentation

* revert(cursor): cut preCompact staging from this change

Carrying a compaction digest across two concurrently running stop hooks kept
producing races that could deliver it twice or strand it indefinitely, and
closing them kept enlarging a critical section inside a hook Cursor awaits at
the turn boundary. Native preCompact firing was never observed either, so the
surface has no empirical basis yet.

Remove the adapter, its registration, its staged path in the park, and its
tests, and record the surface as deferred and uncovered alongside the Codex
interactive TUI. A regression now asserts preCompact stays unregistered so it
cannot return without its own design and evidence.

This change ships the proven core only: the turn-end follow-up park, the
run-tier session start, and away-mode delivery.

* no-mistakes(review): Correct Cursor park supersession documentation

* no-mistakes(document): Clarify Cursor run-tier verification ownership

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

---------

Co-authored-by: kunchenguid <kun-1@kunchenguid.com>

* feat(bin): add decline and repair paths for decision holds (#2330)

* feat(bin): add unrouted close paths to the captain decision gate

A captain who declines a held decision leaves no follow-up work to route,
so `resolve` could not express that answer: it requires at least one
`--routed-to` task. The only way to close such a hold was a direct
`tasks-axi done`, which never writes the durable resolution record the
completion gate reads, so the originating investigation could no longer
pass `verify` and its cleanup stayed blocked.

Add two close paths that route no work:

- `decline` closes an actively held hold with a recorded captain decision
  and no routed task. It refuses while any task is still blocked by the
  hold, because releasing routed work without recording it is `resolve`'s
  job.
- `repair` records the missing resolution block on a hold that was already
  closed outside this script. It never reopens a hold and never clears a
  dependency edge, and it refuses a hold that is still actively held.

Both require a non-empty captain decision file and share `resolve`'s
digest-based retry identity, so an exact retry is idempotent while a
changed decision is rejected. The recorded body now also names which path
closed the hold, and each routed entry regains its own line.

The gate itself is unchanged: an unanswered decision still fails
completion and blocks teardown, and neither new path can close a hold
without the captain's recorded word.

* fix(bin): require captain-hold provenance before repairing a decision

`repair` checked only that the backlog item was kind captain and Done, so
an ordinary captain-kind task that was never held for the captain could be
closed, repaired, and then pass the completion gate.

tasks-axi keeps `hold_kind` through a close, so it is the surviving proof
that an identity really was a captain hold. Require it before writing the
resolution record, and cover the case in the gate regression.

* no-mistakes(document): Correct decision-hold lifecycle documentation

* fix(bin): surface buried wake status lines once (#2331)

* fix(bin): surface buried status notes on wake drain

A note: answer immediately followed by a routine note was dropped because
annotations kept only the newest line and note: never enters OPEN DECISIONS.
Present every unread note and pending-reply resolution since the last drain
cursor, and annotate every unread line on a queued signal.

* no-mistakes(review): Fix unread status cursor races and overflow

* no-mistakes(review): Preserve cursors when status span reads fail

* no-mistakes(review): Make status presentation transactional under I/O failures

* no-mistakes(review): Simplify unread status cursor and presentation locking

* no-mistakes(review): Align cursor failure regressions with transactional presentation

* no-mistakes(review): Retire stale presentation cursors during task teardown

* no-mistakes(review): Preserve routine status until signal annotation

* no-mistakes(review): Correct unread status cap documentation

* no-mistakes(document): Document unread wake status presentation

* no-mistakes(lint): Fix wake surfacing ShellCheck warnings

* no-mistakes: apply CI fixes

* feat: add max Calm presentation level (#2334)

* feat(calm): add a max presentation level that hides mid-turn working notes

Calm's home-local preference becomes a three-state level instead of a
boolean: "off" is stock Pi, "on" is today's Calm, and "max" is Calm plus
hiding the assistant text of messages the model did not end its response
with. `/calm max` selects it from any state, a plain `/calm` steps max
back to ordinary Calm and otherwise keeps the existing on/off cycle, and
any other argument keeps that cycle too.

`config/calm` now persists "max" as its own literal value, so a session
start, resume, fork, or reload restores the stored level rather than
treating it as unrecognized and dropping to off.

The hide rule keys on Pi's intrinsic per-message stopReason: "toolUse",
or "length" with tool calls present. Streaming ("pending") text is never
filtered, because suppressing it would also stop a genuine reply from
streaming. The existing assistant layout adapter filters the blocks out
of the same shallow presentation copy it already uses for collapsed
thinking, so the message, model context, session storage, /export, and
delivery are untouched and a hidden mid-turn row collapses to zero
height. The new "assistant-working-note" class keeps that choice in the
visibility policy owner, where ordinary Calm keeps it visible.

* no-mistakes(document): Clarify Calm max persistence and taxonomy

* feat(calm): hide mid-turn working notes by default (#2339)

* feat(calm): make hiding mid-turn working notes the ordinary Calm state

Calm collapses back to the two-state on/off toggle it was before the max
presentation level, with max's hide rule promoted into ordinary Calm.
Calm on now hides mid-turn assistant working notes in addition to what it
already hid, and the /calm command parses no argument again.

The hide rule itself is unchanged: assistant text is removed from the
shallow presentation copy when the message's own stopReason is "toolUse",
or "length" with tool calls present. Streaming ("pending") text is never
filtered, so a genuine reply still streams. The message, model context,
session storage, /export, and delivery remain untouched.

config/calm persists only "on" and "off" again, but the reader still maps
a persisted "max" to on so a home upgraded from the removed level keeps
Calm on instead of dropping to off.

The mid-turn hide is now default behavior rather than an opt-in level, so
docs/calm.md documents it for users, docs/configuration.md records the
two written values plus the legacy max mapping, and the feasibility
taxonomy drops its level-scoped wording.

* no-mistakes(document): Document ordinary Calm working-note hiding

* chore: store no-mistakes test evidence in the repo (#2355)

* chore: ignore scratchpad/ at the repo root (#2359)

* fix(ci): fail hung Herdr behavior runs in 20 minutes (#2413)

A wedged family-run step was occupying the runner until the 75-minute
job cap; bound that step so cleanup and timing artifacts still upload.

* fix: keep the public promise reachable when work is routed to a second mate (#2457)

The lightweight Relay follow-up link lives in the answering home's own
state/<task-id>.meta, so it can only bind work that home owns. When a
Relay-linked request is routed to a second mate, the task record lives in the
second mate's home, fm-x-link.sh failed with a bare "no such task ...meta", and
nothing else picked the promise up: only the soft acknowledgement was ever
posted. The typed promised-final path already supports --work-home
secondmate:<id>; the playbook simply never chose it.

- fmx-respond now states the routing rule crisply: a task in this home takes the
  lightweight link, and second-mate-routed work takes a promised-final
  commitment bound to that home, registered up front with the brief command
  carried into the routed worker's instructions.
- fm-x-link.sh refuses a task with no local record by naming the registered
  second mate whose home actually holds it and printing the promised-final
  registration command, with the exact --work-home when the match is
  unambiguous. A home with no registered second mates keeps the plain error.
- fm-backlog-handoff.sh reports, after a successful move, any moved key that
  still owes a public reply bound to main/<key>, since that binding no longer
  names the home owning the work. The move itself is never blocked.

Docs and the secondmate handoff prose follow the same rule. Tests cover the
refusal, its scoping, the unchanged local-link path, and both handoff outcomes
at the script boundary.

* docs(skills): add remote-secondmate recovery hint for false-negative verdicts (#2456)

* fix(skills): hint that remote secondmate liveness verdicts false-negative

fm-crew-state and fm-send routinely misreport a live remote secondmate
as dead; confirm against the pane before relaunching, and relaunch only
through fm-spawn.sh, never raw herdr pane surgery.

* no-mistakes: apply CI fixes

* fix(calm): keep Pi's export confirmation visible (#2461)

Pi 0.83.0 added a status line to every tool-expansion change, and Pi
updates the previous status line in place when two status messages
arrive back to back. Calm's post-export redraw cycled tool expansion on
the macrotask right after Pi printed "Session exported to: <path>", so
both expansion status lines coalesced over that confirmation and the
captain was left with no record of where their export landed.

Calm now repaints only the tool rows it presents, by invalidating each
row through the render context Pi hands its render slots, and requests
the surrounding redraw through setStatus. Neither appends to the
transcript. The repaint is still needed because Pi can re-render a row
asynchronously - the built-in edit row invalidates itself once its diff
is ready - and that re-render can land inside the window where /export
forces stock rendering.

The real-terminal /export case now asserts the confirmation is still on
screen after the redraw has settled, and that the redraw restored every
Calm-hidden row, instead of only racing the moment the confirmation
first appeared.

* feat(stow): add open-record persistence to /stow before reset (#2488)

* feat(stow): persist the open records a session is holding

/stow curated memory and captured session knowledge, but never touched
record state, while AGENTS.md called it an "unfinished-work sweep" and the
receipt declared the session "safe to reset" - wording that implied a
record-correctness guarantee stow does not make. A shipped PR with no
backlog item, a queued umbrella whose phases had merged, and four decision
holds left open after their answers shipped all survived repeated stows.

Add a bounded pass that files record state from the same volatile input the
rest of stow already uses: the open threads in context, minutes before the
reset destroys them. It creates a record for an unfiled thread and corrects
one the session knows is wrong, through the owning path, and states its
boundary as part of the contract - it never enumerates the backlog, lists
holds, or queries a forge, because it cannot be a reconciliation and must
not be read as one.

Correct the wording in AGENTS.md and the completion receipt so reset-safe
means what it actually guarantees: nothing this session knew was lost.

* no-mistakes(review): correct stow decision-hold inspection to read hold via tasks-axi

* no-mistakes(document): note /stow open-record persistence in README command catalog

* refactor(stow): state open-record persistence as principle, not procedure

The first version enumerated triggers, named commands, and prescribed an
ordered procedure. That is too rigid for an agent skill: it invites literal
execution of a checklist instead of judgment, and every enumerated example
is a way for the guidance to go stale.

Reduce it to the intent - before a reset, the important open work you are
holding in context must end up durably recorded rather than dying with the
session, filing what is unfiled and correcting what is stale - and let the
agent judge importance, the record, and the owning write path.

Keep the scope bound, since it is a decided contract and not a mechanic:
this covers the open work the session is holding, never a reconciliation of
durable records against repository or forge reality. The wording
corrections in AGENTS.md and the completion receipt are unchanged.

* fix(decisions): close decision holds at answer time via one general keyed-answer path (#2490)

* fix(decisions): close captain holds at answer time

Firstmate had two "a decision is open" ledgers with asymmetric closing
mechanics. The live status-log ledger closes atomically at answer time,
because bin/fm-send.sh --resolve-key makes answering a decision be the
act that closes it. The durable backlog hold ledger had no such coupling:
answering and recording were two separate acts, and only the first was
forced by the workflow.

That asymmetry lost four real captain decisions. Their answers were
captured durably to disk, keyed character for character by the hold
decision keys, acknowledged, and even implemented and shipped, yet the
holds stayed open for two days and the captain was asked to re-answer
decisions already on his own disk.

Give the hold ledger the same answer-time-closure property:

- bin/fm-decision-hold.sh gains an `answer` subcommand, the hold ledger's
  counterpart to --resolve-key. It shares one unrouted close
  implementation with `decline`, so it carries every existing guard - the
  captain decision file, the active-hold requirement, retry identity, and
  the refusal to release still-routed work - and differs only in the
  resolution mode it records. `decline` keeps its stronger meaning that
  the answer routes no follow-up work at all.
- bin/fm-procevent-lavish.sh wires the channel that actually carried the
  lost answers. `arm --decisions-origin` binds a deck to the origin whose
  holds it carries, `answers` reads the structured choices out of a
  captured poll result, `close-decisions` maps each key to its hold and
  closes it through the command above, and `autohandle` lets the runner
  apply that at capture time.

Safety is preserved rather than traded away. Only rows tagged `choice`
are read, so freeform captain prose cannot forge a decision key. Closure
is confined to the one bound origin. The decision text is a pure function
of the captured result, so a replayed capture is idempotent. A hold that
is absent, already closed, or still blocking routed work is skipped and
left for `resolve`, never forced. A deck armed without the binding
touches no hold at all. And autohandle deliberately never reports full
handling, because recording an answer is transcription while acting on it
is firstmate's judgement - so the check wake still reaches the handler.

fm-send --resolve-key is untouched.

* no-mistakes(document): document state/lavish-decisions binding dir in AGENTS.md state inventory

* refactor(decisions): make keyed-answer closure one general capability

The previous pass gave holds answer-time closure but built it as bespoke
Lavish wiring: the review adapter carried the source-to-origin binding,
mapped keys to hold identities, wrote decision records, decided what to
skip, and closed holds itself. That treated a review deck as a special
decision source. It is not - it is an ephemeral discussion format that
happens to carry answers.

Collapse it into ONE general capability with one owner.

bin/fm-decision-hold.sh now owns the whole of "a keyed answer closes its
matching hold":
- `answers <origin> --source <provenance>` is the channel-agnostic
  intake. It reads key/answer/label lines on stdin, maps each key to its
  hold, and closes it through the same `answer` path, so every guard
  applies identically whatever channel the answer came from. --source is
  provenance recorded in the decision, never a behavior switch; there is
  no per-channel branch and no knowledge of chat, decks, or transports.
- `bind`/`unbind`/`binding` own the source-to-origin binding for any
  channel whose answers arrive detached from their origin.

Every channel is now an ordinary caller that only turns what it received
into keyed lines:
- bin/fm-send.sh (chat) feeds the intake for a key that names an active
  hold. This also fixes a real gap: once `complete` transfers a decision
  to its hold it closes the live status copy, so --resolve-key alone
  could never answer a transferred decision.
- bin/fm-procevent.sh feeds it generically. A bound source's captured
  result goes to `<adapter> answers <result-file>` and whatever that
  prints is piped into the intake. The runner names no adapter, parses
  no result, and carries no decision rule, so any future adapter with an
  `answers` command works with no change here.
- bin/fm-procevent-lavish.sh keeps only `answers`, which reports the
  structured choices a review captured and stops. It maps nothing to a
  hold and closes nothing; it lost ~160 lines of decision logic.

Feeding is independent of handling, so it never acknowledges a result
and never suppresses a wake - recording an answer is transcription,
acting on it stays firstmate's judgement.

The regression that proves closure now drives a FIXTURE adapter that is
not the review adapter, so what is proven is that any bound channel
reaches the intake rather than that one channel is wired specially. A
new regression drives the real fm-send over a stubbed transport for the
chat side. Every prior guarantee still holds, and fm-send's status-log
behavior is unchanged.

* no-mistakes(review): test(decisions): drop source-content grep from hold-closure regression

* fix(memory): emit a real @AGENTS.md pointer instead of a CLAUDE.md symlink (#2512)

A Write aimed at CLAUDE.md followed the symlink and destroyed AGENTS.md.
The installer now creates and migrates to a recoverable two-line pointer file.

* fix(ci): keep CLAUDE.md pointer check valid (#2515)

* ci: gate GitHub workflows with pinned actionlint (#2517)

* fix(lint): catch malformed GitHub workflows before merge

A self-broken ci.yml cannot report its own breakage, so parse every
workflow in the local lint path that no-mistakes already runs.

* fix(lint): pin actionlint instead of Ruby for workflow lint

A self-broken ci.yml still has to fail in the local lint path, and the
named tool for that gate is actionlint, not a new Ruby runtime.

* no-mistakes(document): Clarify pinned workflow lint documentation

* fix: install pinned lint tools across supported platforms (#2546)

* fix: install pinned shellcheck and actionlint on macOS and linux arm64

The installers were hardcoded to linux amd64 and sha256sum, so a Mac
dev could not satisfy the refuse-on-mismatch lint gate. Select the
official per-platform archive and checksum, and fall back to shasum -a 256.

* no-mistakes(document): Document cross-platform pinned lint installers

* docs: reconcile test-evidence docs with store_in_repo: true (#2548)

.no-mistakes.yaml has set test.evidence.store_in_repo: true since #2355, but
CONTRIBUTING.md, docs/configuration.md, and docs/architecture.md still described
the old policy of keeping evidence out of the repo in a temp directory.

The current no-mistakes behavior for store_in_repo: true is to publish each run's
test evidence to the orphan no-mistakes/evidence branch and link it from the PR
body. That branch shares no history with code branches, so evidence never enters
a pushed feature branch or the default branch, and CI's tracked personal fleet
paths rule stays accurate.

Docs only. No change to .no-mistakes.yaml or any workflow.

* docs: clarify test evidence branch storage (#2549)

* docs: correct test evidence storage comment in .no-mistakes.yaml

* no-mistakes: apply CI fixes

* docs: hint that live scouts may host their own Lavish review loop (#2563)

Make that a first-class option in always-loaded instructions so firstmate does not default to mediating and tearing the scout down between iteration rounds.

* fix(bin): report remote secondmate delivery and state truthfully (#2570)

* fix(bin): report remote secondmate delivery and state truthfully

A steer to a remote secondmate crosses fm-on.sh to a host-local fm-send
leg whose unconfirmed submit read-back (verdict=pending, typically a busy
mate whose harness queues the steer) was flattened into exit 1, so the
parent printed "error: text not submitted" / "error: text not sent" and
discarded the …
Alchemy86 added a commit to Alchemy86/firstmate that referenced this pull request Sep 14, 2026
* fix(remote-job): stop workers abandoned by a pruned code root (#1927)

29 fm-remote-job-worker.sh processes were found running at ppid 1, 1-2 days
old, each still polling and appending to a log inside a no-mistakes gate
worktree that had already been returned.

Three things combined to make that possible:

- The recorded worker.pid is the serving child, not the restart supervisor
  above it, so a teardown that stops that one pid only makes the supervisor
  respawn. The Linux start path also left the worker tree in the launching
  command's process group, so there was no group to signal instead.
- Neither the serving loop nor the supervisor ever rechecked whether its
  configured FM_ROOT still existed, so a worker launched from a worktree
  outlived that worktree indefinitely.
- The supervisor restarted a failing child with a fixed 0.1s delay and no
  bound, which is what grew the logs (~66MB/day measured).

The Linux start path now puts the worker tree in its own process group, and
fm_remote_job_stop_worker_tree signals that whole group - refusing any group
whose leader is not itself a worker, so a worker from an older build or from
launchd's own session is still stopped safely as a single process. The worker
stops itself once its code root stops being a Firstmate checkout, confirmed
across a grace window so an ordinary transient cannot stop a healthy worker.
The supervisor backs off and gives up rather than restarting forever.

bin/fm-remote-job-reap-orphans.sh is the belt-and-suspenders sweep for workers
already orphaned that way, wired into fm-teardown.sh. Its reap condition is
exactly "the code root named in the worker's own command line is gone", which
is why the account's healthy LaunchAgent worker and every live remote
secondmate worker are never candidates.

The two suites that leaked these in the first place now stop the worker tree
rather than the recorded pid alone.

* feat(bin): lint only the changed shard locally, full lint in CI (#1925)

* fix(bin): lint only the changed shard locally, full lint in CI

Two ships hitting fm-lint.sh at once could spike CPU to 190% and load
to 8.58 on a captain's Mac, even though each run finishes quickly.
fm-lint.sh now defaults to linting only the canonical-set files
changed since the merge-base with origin/main (including uncommitted
edits) on an ordinary local branch, using plain local git with no
network calls. It still lints the full canonical set in CI
(GITHUB_ACTIONS=true or CI=true), on the main branch, or whenever no
merge-base can be found, so CI coverage never depends on a local diff.
Explicit paths keep bypassing this selection entirely.

* no-mistakes: apply CI fixes

* feat(bin): add deterministic agent lifecycle control (#1568)

* feat(bin): add deterministic agent lifecycle control

Separate firstmate's data plane from its control plane.

bin/fm-send.sh is the data plane: conversational text, always
routing-marked for a kind=secondmate target. That marking is right for a
message and wrong for a lifecycle command - a marked "/quit" arrives as
ordinary chat the agent reasons about instead of executing.

bin/fm-control.sh is the control plane: allowlisted interrupt, exit, and
transactional relaunch verbs addressed to an exact task id, with
per-harness mechanics owned by the executable bin/fm-control-lib.sh
rather than improvised in agent prose, and a verified postcondition for
every action. There is no arbitrary-text and no raw-key entry point.

relaunch runs as a transaction with a durable journal: it resolves the
profile, proves the work it must preserve is recoverable, records the
required progress note, stops the old agent, then delegates the launch
to its single owner, bin/fm-spawn.sh --relaunch, which adopts the
recorded endpoint and worktree instead of creating either. A refusal
before the stop leaves the record and instructions byte-identical; a
failure after it reports the concrete state rather than claiming an
agent that is not running. Teardown and discard stay separate and
explicit.

exit and relaunch require a backend with a recovery-grade agent-state
classifier, so zellij, orca, and cmux are refused rather than reported
as successful blind. A remotely placed secondmate is refused by name,
because its agent runs on a host where none of these postconditions can
be read.

* fix(control): resolve a recorded harness to its adapter before retiring wiring

fm-spawn arms per-task harness wiring on prefixes, because a task
launched from a raw command records that command's basename rather than
the exact adapter name. The control plane's retirement tables are keyed
by the exact adapter, so a task recorded as `grok-2` had its turn-end
token, private registry entry, and worktree hook pointer armed and never
retired - leaving a registry entry that outlived the agent that owned
it.

State the prefix rule once, in the capability owner, and resolve the
recorded value through it before every table lookup. bin/fm-send.sh's
composer-clear lookup reads the same owner instead of keeping its own
copy of which adapters need one.

* test(control): pin muse session-binding retirement across a harness switch

* no-mistakes(review): Resolve prefixed harnesses across lifecycle control verbs

* no-mistakes(review): Report interrupt delivery without fabricating cancellation state

* no-mistakes(review): Clear disabled relaunch trace context atomically

* no-mistakes(review): Clarify control interrupts and restore legacy send state

* no-mistakes(review): Refuse ambiguous relaunches and report exit delivery

* no-mistakes(review): Revalidate interrupts and accept interrupt-stopped exits

* no-mistakes(review): Lock descendant tasks before forced recursive teardown

* no-mistakes(document): Align lifecycle adapter documentation with control plane

* no-mistakes: apply CI fixes

* fix(bin): serialize fresh task publication with forced teardown

Forced secondmate teardown enumerated a home's task set, locked what it
found, then re-enumerated while removing. A fresh spawn takes only its
own per-task lock, so a record published inside that window was
invisible to the preflight and visible to the cleanup: it was
destructively processed while never lifecycle-locked.

Reproduced with real agents. A record published 0.249s after teardown
began was removed, its window closed, and its worktree returned to the
pool - while both commands reported success. A per-task lock cannot
protect a task that does not exist yet.

Add a per-home task-set lock guarding WHICH tasks a home has, as opposed
to the metadata lock guarding one task's record. Teardown takes it per
home, parent before child, before enumerating and holds it through
cleanup. A fresh spawn takes it before its own per-task locks and holds
it through publication; a relaunch is exempt, because it republishes an
existing task already covered by that task's control lock.

Either the spawn publishes first and the teardown's preflight covers it,
or the teardown owns the set and the spawn refuses. Both directions fail
closed, and both are pinned by tests that hold the lock rather than
racing on timing.

* no-mistakes(review): Serialize remote secondmate publication with forced teardown

* no-mistakes(review): Preserve remote spawn routing and state initialization

* no-mistakes(review): Serialize teardown when descendant state is absent

* no-mistakes(review): Cover symlinked descendant state refusal

* no-mistakes(document): Document task-set serialization safeguards

* no-mistakes(lint): Isolate task-set lock path resolution

* no-mistakes: apply CI fixes

* feat(stow): add tiered decaying memory management (#1984)

* feat(stow): tiered decaying memory with captain-gated offload to local excluded skills

Implement the captain-adopted /stow redesign from the v2 tiering report as
amended by the adoption decision:

- Per-entry trailing HTML-comment markers with three tiers named for their
  handling: pinned (no clock, no eviction), aging (stale after 30 days),
  perishable (stale after 7 days, mandatory checkable expiry condition).
- File-scoped defaults (captain.md and captain-shared.md pinned,
  learnings.md aging) with a self-describing legend line per file header.
- Reinforcement requires session evidence; re-reading memory never counts.
- Archive-not-delete: stale and budget-evicted entries move with provenance
  to the never-injected data/memory-archive.md; prune always means the cold
  tier, and a stale unique fact is never deleted.
- Captain-gated over-budget offload: staleness evaluated before scope, the
  sweep runs only when still over budget after decay and consolidation,
  proposals go through the receipt plus one durable captain-held backlog
  item, migration runs through the destination's normal path, and the
  memory entry leaves only once the destination is live.
- Offload destination per the adoption decision: a user-owned skill under
  .agents/skills/<freeform-name>/ excluded via the local .git/info/exclude,
  with the hard rule that stow never creates or writes a tracked skill.
- Five graduation moves, receipt verbs archived and proposed-offload, and
  the one-time non-destructive migration of unmarked legacy entries.

The public skills/stow/SKILL.md mirrors the generic parts (markers, decay,
archive exit, user-approved on-demand offload exit, migration) with no
firstmate-specific paths.

The load-bearing assumption that a git-excluded skill is still discovered
was verified empirically against Claude Code 2.1.226 (direct
.git/info/exclude scratch-repo test plus an in-repo ignored-probe test);
the dated evidence is recorded in docs/verification/stow-memory.md.

The graduation list's deletion move is deliberately narrowed to duplicates
already preserved by a stronger owner, reconciling the v2 report's retained
'deletion of a stale entry' wording with its own prune-always-archives
rule.

* no-mistakes(review): Persist legacy migration grace across stow passes

* no-mistakes(review): Enforce archival invariants and exempt default-pinned legacy entries

* no-mistakes(review): Clarify offload scope, archive placement, and marker boundaries

* no-mistakes(review): Enforce aging fallback and verify excluded skill loading

* no-mistakes(review): Fix stow decay, pinned offload, and archival safeguards

* no-mistakes(review): Preserve pinned entries, approvals, and archive provenance

* no-mistakes(review): Restrict stow mutations to editable memory files

* no-mistakes(review): Clarify skill destinations, collision checks, and migration legends

* no-mistakes(review): Resolve exclude paths for linked worktrees

* no-mistakes(review): Secure per-home excluded skill migration

* no-mistakes(test): Require explicit tier markers on new stow entries

* no-mistakes(test): Route missing shared legends to primary owner

* no-mistakes(document): Align stow documentation with tiered memory

* fix(stow): converge the pass on an over-budget home (dogfood D1-D3)

The dogfood run against a copy of the real over-budget home showed the
pass increasing the deficit from 624 to 1,107 estimated tokens and the
relief ladder provably unable to reach budget. Three skill-text fixes:

- D1: markers become single-token spellings (<!--a:DATE-->, <!--p:DATE-->,
  <!--P-->, <!--g-->), entries matching a pinned file default carry no
  marker, the per-file policy legend collapses to a one-line pointer
  naming the stow skill as the scheme owner, and marker/pointer bytes are
  explicitly counted content - roughly 76% less metadata cost on the
  dogfooded home's first installment.
- D2: the eviction rung gains a convergence precondition - total the
  eligible pool first, and when archiving all of it cannot reach budget,
  skip eviction entirely, archive nothing for budget reasons, and report
  the exempt pinned floor as the concrete inability in the final step.
- D3: budget eviction considers only dated aging entries; <!--g-->
  legacy-grace entries are ineligible until their grace cycle resolves,
  so eviction cannot cancel promised grace or invert against validation.

Public skill mirrors the D1 marker/pointer changes; D2/D3 are internal
because the public skill has no budget ladder.

* no-mistakes(test): Enforce evidence-only reinforcement during stow migration

* no-mistakes(document): Clarify stow receipt marker actions

* docs: add project vision (#1997)

* docs: add firstmate vision

* no-mistakes(test): Classify VISION.md as public product documentation

* no-mistakes(document): Restore approved one-file vision diff

* no-mistakes: apply CI fixes

* fix(spawn): force regular Pi TUI for crews (#2005)

* fix(spawn): force regular Pi TUI for crews

* no-mistakes(document): Documented Pi regular TUI launch mode

* fix(cmux): classify borderless Claude composers (#2029)

* fix(cmux): classify borderless Claude composer

* no-mistakes(review): Normalize cmux NBSP prompts across locales

* no-mistakes(document): Document cmux borderless Claude composer classification

* docs(stow): generalize read-before-write in the public stow skill (#2091)

The public installer-facing stow skill scoped its classify-then-replace
discipline to TODO/BACKLOG items only, so findings routed to a memory file
had no stated rule against a blind append or a wholesale overwrite.

Step 6 now classifies every finding against the destination's current
contents as new, duplicate, superseding, or obsolete, and states the
considered replacement each classification implies. The outcomes follow the
tiered-memory contract already in the file: an obsolete entry is refreshed,
archived, or replaced in a way that preserves its fact, a duplicate folds
into the entry that already carries it, and a superseded body worth keeping
leaves through step 7's existing exits rather than a second recovery
mechanism.

* fix: resurface durable supervision work after re-arm (#2065)

* fix(watcher): resurface durable work after downtime

* no-mistakes(review): Make watcher rearm recovery durable and cursor-safe

* no-mistakes(review): Persist safe recovery markers across migration lock recovery

* no-mistakes(review): Retain stale lock when recovery marker publication fails

* no-mistakes(review): Preserve delivery-gap recovery and quarantine malformed markers

* no-mistakes(review): Serialize recovery consumption and report acknowledgment failures

* no-mistakes(review): Centralize recovery publication before clearing watcher evidence

* no-mistakes(review): Guarantee recovery evidence across queue and lock handoffs

* no-mistakes(review): Publish recovery evidence before durable wake commits

* no-mistakes(review): Replace recovery marker Perl dependency with Node

* no-mistakes(review): Keep interrupted wakes durable until handling acknowledgment

* no-mistakes(review): Add post-handling durable wake acknowledgements

* no-mistakes(review): Enforce post-handling acknowledgement across recovery and AFK return

* no-mistakes(review): Bind wake acknowledgements to recovery generations

* no-mistakes(review): Align wake regressions with generation-bound acknowledgements

* no-mistakes(document): Document durable re-arm recovery semantics

* no-mistakes(lint): Resolve ShellCheck warnings in recovery and watcher tests

* no-mistakes: apply CI fixes

* test(watcher): assert post-handling wake replay

* no-mistakes(review): Prevent successor loops and adopt legacy wake generations

* no-mistakes(review): Rearm durable wakes without recursive successor recovery

* no-mistakes(review): Align recovery tests with handling marker state

* no-mistakes(review): Delay handling transition until successor launch is established

* no-mistakes(review): Confirm wake handling only after successful prompt delivery

* no-mistakes(review): Acknowledge AFK wakes only after evidence publication

* no-mistakes(review): Prevent AFK wake loss before post-handling acknowledgement

* no-mistakes(document): Document durable wake acknowledgement semantics

* no-mistakes(lint): Suppress false positive for recovery action output

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* ci: measure Herdr automation on Windows runners (#2100)

* ci: add Windows Herdr automation spike

* ci: run Windows spike on its pull request

* fix: wait for Windows Herdr command output

* fix: run ANSI probe in pane shell

* ci: keep Windows Herdr spike manually triggered

* docs: clarify Windows Herdr spike verdict

* feat(ahoy): guide captains through open decisions (#2099)

* Add guided ahoy decision flow

* no-mistakes(document): Document guided Ahoy decision flow

* fix(stow): enforce startup-memory budget decisions (#2110)

* Harden stow memory budget policy

* Refine internal stow offload policy

* no-mistakes(review): Enforce shared-budget decisions and autonomous offload

* fix(spawn): refresh pooled worktrees from origin before launch (#2116)

* fix(spawn): refresh pooled worktree base

* no-mistakes(document): Document spawn base-freshness invariant

* no-mistakes: apply CI fixes

* fix(composer): unify safe classification across backends (#2102)

* refactor(composer): one shape owner behind thin capture adapters, whole matrix fixed

Consolidate every composer shape - bordered boxes (all families, geometry,
titled bottom borders), bare agent-glyph rows and their wrap regions,
opencode's left bar, and pi's identity-gated separator pair - into
fm_composer_classify_screen in bin/fm-composer-lib.sh. Adapters now
contribute only a capture and a declarative capability descriptor
(styled/cursor/identity/rows); capability differences change how confidently
a shape is judged, never what the shapes are, so a new harness shape is
teachable in exactly one place.

Correctness fixes landed as part of the consolidation (audit
data/fm-composer-consolidation-audit-s1):
- locale-safe Unicode-space normalization in the shared owner (closes the
  fleet-wide half of #1988; cmux's local byte-exact NBSP case deleted;
  naming converges with PR #1995's normalization primitive)
- muse's bare glyph joins the shared set, unbreaking muse on herdr/cmux/orca
- orca learns the borderless bare shape, drops its backward-paged composer
  window, and can no longer classify a stale startup banner as the composer
- tmux tolerates a titled bottom border, unbreaking grok steering
- the left-bar shape makes opencode readable on every backend
- zellij gets a real classifier through dump-screen --ansi, replacing the
  content-diff submit heuristic that could confirm an undelivered message
  and close a --resolve-key decision (the fleet's only false positive)
- fm-spawn's kimi launch-readiness regex (the fourth shape copy) now routes
  through the shared classifier

The strict blank-row posture applies fleet-wide (captain decision
blank-row-injection-posture): no positive container proof = unknown = defer,
replacing tmux's permissive blank-cursor-row rule. Away-mode injection was
re-validated end to end on real tmux (defer on partial input and unproven
rows, clean delivery with swallowed-Enter retry into proven-empty
composers). The tmux submit core gains a baseline-idle turn-started
conversion so pi steering stays confirmed while its working screen hides
the composer; busy conversion without that baseline remains forbidden.

Plain-capture backends now degrade a glyph row carrying trailing text to
unknown instead of a false pending, per the approved capability rule.

Portable regressions pin the full byte-capture matrix from the audit under
a UTF-8 locale and LC_ALL=C, the strict-vs-permissive divergence, and
deliberate signal separation; the opt-in live guard
(tests/fm-composer-matrix-live-e2e.test.sh) verified every installed
harness against the real classifier, recorded in
docs/verification/runtime-backends.md.

* no-mistakes(review): Fix Pi glyph ambiguity and complete profile matrix

* no-mistakes(review): Preserve bare verdict when Pi identity probe is absent

* no-mistakes(review): Harden composer structure and titled-border geometry

* no-mistakes(review): Require proven idle baseline and strict Zellij guard

* no-mistakes(review): Reject box bottom borders as composer input rows

* no-mistakes(review): Prove Zellij probe typing before classifier retries

* no-mistakes(review): Preserve Pi identity uncertainty and scan full left-bar drafts

* no-mistakes(review): Verify Zellij text lands before submitting

* no-mistakes(review): Scope Zellij typing verification to selected composer content

* no-mistakes(review): Verify Zellij pastes through composer-scoped content deltas

* no-mistakes(review): Prove wrapped bare Zellij pastes through composer extraction

* no-mistakes(review): Invalidate stale cursorless composers below dead shell prompts

* no-mistakes(review): Handle shell prompt placeholders in composer extraction

* no-mistakes(review): Classify cursorless bare continuation regions safely

* no-mistakes(review): Reject stale cursorless containers below live activity

* no-mistakes(review): Preserve prompt glyphs in wrapped Zellij pastes

* no-mistakes(review): Reject live shell rows during composer extraction

* no-mistakes(review): Preserve wrapped glyph continuations through submit retries

* no-mistakes(review): Scope idle placeholders to proven positions

* no-mistakes(review): Restore boxed placeholders and live prompt reanchoring

* no-mistakes(review): Fix Zellij placeholder and wrapped glyph paste proof

* no-mistakes(document): Align composer architecture documentation

* no-mistakes(lint): Fix ShellCheck warnings in composer refactor

* no-mistakes: apply CI fixes

* docs(verification): record the trusted-checkout live matrix rerun

The pipeline's isolated gate worktree is untrusted, so claude, grok, and
muse stopped at first-launch trust dialogs there (the guard refuses to
confirm them by design). This rerun from the trusted checkout at the final
validated head verified all six installed harnesses, the strict blank-row
deferral, and the hardened zellij false-positive probe live.

* no-mistakes(document): Align composer verification evidence

* no-mistakes: apply CI fixes

* no-mistakes(review): Restore proven box bottom-cursor classification

* no-mistakes(review): Preserve styled placeholder-like drafts as pending

* no-mistakes(document): Align composer safety and Zellij delivery documentation

* no-mistakes: apply CI fixes

* docs(verification): refresh the live matrix with the final-head trusted rerun

The post-validation rerun from the trusted checkout verified all six
installed harnesses at the branch's final head, including Claude 2.1.227
(auto-updated since the audit's captures) and Grok, which the untrusted
gate worktree could not verify past their first-launch trust dialogs.

* fix(spawn): gate Pi TUI mode by CLI capability (#2117)

* fix(spawn): gate Pi regular TUI flag by capability

* no-mistakes(review): Document conditional Pi TUI capability detection

* no-mistakes(review): Pin Pi probing and launch to one executable

* no-mistakes(review): Preserve literal pinned Pi paths and update documentation

* no-mistakes(review): Defer pinned Pi path insertion until final substitution

* no-mistakes(document): Document version-safe Pi launch probing

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* docs(vision): elevate experience, pain narrative, and distro virtues (#2147)

* docs(vision): elevate experience, pain narrative, and distro virtues

Fold the captain's public vision framing into VISION.md: peace of mind as a
primary goal, multi-session context-switch pain as the problem one interface
solves, clone-and-run setup ease, self-evolution including community, and
explicit harness/backend orthogonality. Reconcile experience-as-garnish into
experience-as-purpose and update aligns/resists accordingly.

* docs(vision): state the experience goal positively

Drop the negative "not a smart workflow / useful tool / impressive technology"
pretext. Lead straight into the positive experience north star.

* feat(bin): reconcile inactive terminal crew outcomes (#2167)

* fix: reconcile inactive terminal outcomes

* fix: stream secondmate summary inputs

* no-mistakes(review): Fix reconciliation locking and request delivery retries

* no-mistakes(review): Prevent retries after unknown request delivery

* no-mistakes(document): Clarify inactive reconciliation cadence and receipts

* no-mistakes(lint): Quote terminal status arguments in reconciliation tests

* refactor: simplify inactive outcome reconciliation

* no-mistakes(review): Bound inactive reconciliation scans with durable progress

* no-mistakes(review): Bound reconciliation and deduplicate recovery notices

* no-mistakes(document): Document inactive outcome reconciliation contracts

* no-mistakes(review): Reject relative local secondmate parent routes

* no-mistakes(review): Key terminal receipts by spawn incarnation

* no-mistakes(review): Stabilize legacy receipts and lock reconciliation snapshots

* no-mistakes(review): Fail closed on invalid secondmate identity markers

* no-mistakes(document): Document durable inactive-outcome reconciliation

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* ci: raise Herdr test timeout (#2191)

* fix: refresh stale Pi instructions after compaction (#2163)

* fix(session-start): refresh drifted instructions on stale rebuilds

* test(session-start): prove Pi instruction refresh end to end

* no-mistakes(review): Fix stale instruction refresh and baseline integrity

* no-mistakes(review): Preserve true-start baselines across Pi continuations

* no-mistakes(review): Correct Pi continuation classification and live expectation

* no-mistakes(review): Correct Pi continuation coverage documentation

* no-mistakes(review): Fix read-only refresh and exact Pi session restores

* no-mistakes(review): Classify Pi create-if-missing sessions correctly

* no-mistakes(review): Classify named Pi sessions using immutable headers

* no-mistakes(review): Correct Codex interactive coverage diagnostic

* no-mistakes(document): Document immutable Pi compaction instruction refresh

* no-mistakes(document): Correct Pi refresh documentation and validation claims

* feat: add deterministic condition-to-action watcher (#2200)

* feat(bin): add deterministic condition->action watch adapter on the process-event channel

Register a (condition, action) pair once with bin/fm-procevent-when.sh and the
existing process-to-event runner polls the condition tokenlessly, fires the
action at most once on a stable true, and wakes firstmate exactly once with the
captured outcome - instead of burning an agent turn per re-check.

The pair is stored privately under state/when/ and hash-bound by a trust record
the same way fm-check-register.sh binds a custom check, so a mutated spec is
refused without executing anything. A durable exclusive fired marker claimed
before the action makes restarts and re-polls unable to double-fire; every
failure path (mutated spec, condition error past budget, expired deadline,
failed action, uncaptured earlier fire) ends in a terminal captured outcome
that wakes firstmate rather than a silent retry. Eligibility stays a firstmate
judgment: only exact, safe, reversible actions may be bound, and judgment-
needing or destructive actions keep the wake-and-decide flow.

* no-mistakes(review): Harden when watcher concurrency, deadlines, timeouts, and output

* no-mistakes(test): Bind watcher actions to registered executable bytes

* no-mistakes(document): Correct condition-action watcher documentation

* no-mistakes(document): Clarify outcome wake re-announcement

* no-mistakes: apply CI fixes

* fix(bin): honor a decision key stated after the verb colon (#2202)

The open-decisions fold only recognized a [key=<slug>] token between the
verb and the colon (needs-decision [key=x]: note). The common worker
shape with the colon first (needs-decision: [key=x] note) silently
folded its stated key into the shared "default" bucket, so two open
decisions could collapse into one record and fm-send --resolve-key <x>
refused to close the decision it plainly named.

A complete token at the head of the note is now an equivalent stated-key
position for every keyed verb, shared by the whole-file and incremental
folds through the one _fm_decision_key owner. The documented
before-colon position wins when both are present, a token deeper in the
note stays prose, a bare keyless line still folds to "default", and a
stated-but-malformed slug is rejected rather than rewritten to
"default". A consumed note-head token is stripped from the note so both
positions yield identical records, and the incremental fold version is
bumped so persisted cursors folded under the old interpretation are
rebuilt from the authoritative log.

Fixes #2109

* fix(bin): prevent watcher recovery acknowledgement livelock (#2212)

* fix(bin): keep a recovery acknowledgement valid across republication

A watcher cycle that opened and closed while the model handled its drained
wakes minted a fresh recovery generation, which invalidated the exact
acknowledgement the drain had just printed. That acknowledgement then consumed
nothing, so the marker stayed pending and every later arm spent its whole cycle
re-announcing the same recovery instead of supervising - a livelock the home
could not leave on its own.

A downtime publication now reuses the generation of an outstanding handling
episode, so a close during the handling window cannot orphan the printed
acknowledgement. The acknowledgement itself separates its two facts: queue-row
consumption is bound to the monotonic --ack-through sequence and always
happens, while only retiring the episode is bound to --recovery-generation. A
generation that moved on is a non-fatal result that names its own remedy
instead of a refusal that consumes nothing.

* no-mistakes(review): Preserve recovery generations and consume stale acknowledgements safely

* no-mistakes(document): Document sequence-bound recovery acknowledgements

* feat(fmx-respond): consume Relay conversation chains (#2206)

* feat(fmx-respond): consume in_reply_to_chain conversation context

The relay's poll payload can carry in_reply_to_chain, an oldest-first
transcript of the surrounding conversation, but the mention-handling
procedure only ever read the immediate in_reply_to parent, so referents
like "this" in a standalone mention stayed unresolvable even when
context was delivered.

Teach fmx-respond to read the chain when present (optional and
backward-compatible: often absent today, kind label not required),
resolve referents against the whole transcript, and extend the
untrusted-content framing to every chain entry including the upcoming
kind=history entries. Document the field's wire shape in
docs/configuration.md as the firstmate-side owner.

* no-mistakes(document): Document Relay chain context ownership

* fix: parse decision verbs before status metadata tags (#2280)

* fix(bin): strip every bracket tag, not just [key=...], from a status verb

status_line_verb only stripped a leading "[key=...]" token before the
colon, so a remote secondmate reply's leading "[corr=...]" correlation
tag stayed glued onto the returned verb word ("needs-decision
[corr=...]" instead of "needs-decision"). The open-decisions fold's
verb match then silently failed to recognize the line at all, so
fm-send --resolve-key refused to close a decision that was plainly
open on the status line.

Generalize the parser to strip every "[name=value]" tag before the
colon, in any order and count, so local and remote replies fold
identically.

* no-mistakes(review): Invalidate stale decision cursors after parser fix

* no-mistakes(document): Clarify status metadata verb parsing

* fix(bin): collapse duplicate supervision wakes (#2287)

* fix: collapse duplicate supervision wakes without losing legitimate updates

One remote-secondmate note produced two handling turns (a procevent check
wake published before autohandle, then a signal wake for the same mirrored
bytes), already-ingested replays such as a cursor-loss whole-log recapture
still woke with nothing to do, this home's own bookkeeping closes (fm-send
--resolve-key, the pending-reply escalation close, the captain-held
transfer) re-woke the session that wrote them, and turn-ended-only wakes
were annotated with already-announced status lines that looked like fresh
progress.

Dedup rules, each at its layer's one owner:
- fm-procevent.sh: an adapter may declare 'self-announcing'; the runner
  then applies first and publishes a check wake only for what remains
  unhandled. fm-procevent-remote-reply.sh declares it: the mirrored status
  append is the single announcement, so a fully applied capture publishes
  nothing and a byte-identical replay stays completely quiet. All other
  adapters keep strict publish-before-apply.
- fm-wake-lib.sh: fm_wake_signal_sig/seen_path/seen_current now own the
  watcher's signal signature and .seen-* marker format, plus
  fm_wake_status_append_self_announced, the guarded bookkeeping append
  that advances the marker only over exactly its own bytes and fails
  toward waking on any pending or interleaved foreign write.
- fm-send.sh, fm-pending-reply-lib.sh, fm-decision-hold.sh: bookkeeping
  closes go through that guarded append; escalation opens stay plain
  appends because a new blocker must wake.
- fm-wake-lib.sh annotations: a historical (turn-ended-only) row skips its
  status annotation only when the file's signature provably matches the
  seen marker; anything unannounced keeps annotating.
- fm-classify-lib.sh: a kind=secondmate task's status signal is never
  absorbed as provably-working, because that stream is the routed-reply
  channel the parent must read.

Also fixes a pre-existing exit-path deadlock the regression run reproduced:
a TERM inside a recovery-marker critical section left fm_lock_try_acquire
spinning against this same process's abandoned hold; a self-held lock is
now reclaimed (a subshell still waits on its parent's live hold).

Regression tests drive the real wake functions and executables in both
directions: each duplicate case collapses, while a new remote reply, new
decision, new blocker, merge result, failure, first status change, and a
later different note on the same task all still wake.

* no-mistakes(document): Document wake deduplication contracts

* feat: add Cursor CLI crew harness (#2238)

* feat(harness): add Cursor Agent CLI adapter

# Conflicts:
#	bin/fm-spawn.sh

* fix(composer): read cursor-agent's reverse-video placeholder as idle

cursor-agent renders its idle composer placeholder dim (SGR 2) but paints the
cell under the terminal cursor in reverse video (SGR 0;7). Reverse video is
neither dim nor a dark truecolor foreground, so the shared ghost stripper keeps
that one character and an idle composer reduces to a lone `P`. Judged on its
own, that remnant reads `pending` on a genuinely idle pane, which defers
away-mode escalation indefinitely on the styled cursorless backends.

Teach the ONE fleet-wide classifier the shape instead of adding an adapter-local
copy: register `→` as an agent prompt glyph so the composer row is structurally
findable at all (without it the bottom-most shape is a stale shell prompt echo
in the scrollback), add both verified placeholders to the idle set, and consult
the styling-independent plain row when the styled row is only a remnant.

The plain-row branch demands the remnant be a proper, strictly shorter substring
of a plain row matching a fully anchored placeholder. Real typed text is
uniformly bright, so stripping leaves it equal to the plain row and it stays
`pending` - verified live against a pane where the typed text was exactly the
placeholder string.

Verified live on cursor-agent 2026.08.11-e8db854; the regression pins the real
captured bytes and asserts the remnant survives stripping, so the case cannot go
vacuous if the stripper later learns SGR 7.

Co-authored-by: Amplify Logic AI <lars@sockinator.co>

* feat(cursor): narrow cursor identity and order its marker before CLAUDECODE

Cursor ships two executable names - `cursor-agent` and the legacy alias `agent`
- and runs as a bundled node script, so tmux reports the pane command as a bare
`node`. Neither `agent` nor `node` can be trusted by name, so identity gets one
owner in bin/fm-cursor-lib.sh that demands cursor's own name or install tree in
the path or argv[0], from the structural signal only. Probing an arbitrary pid's
executable during a liveness poll would execute a stranger's binary, which is
the hazard that rule exists to close.

Two consequences wired up:

Detection. cursor-agent does NOT clear an inherited CLAUDECODE, so a cursor
worker launched under a claude primary carries both markers and whichever is
tested first wins. The cursor markers are ordered ahead of the CLAUDECODE check;
fm-spawn additionally clears foreign markers at the launch boundary. Both are
kept deliberately - launch sanitization only covers sessions fm-spawn started,
while the ordering also covers a cursor session started by hand. Verified live
that CURSOR_INVOKED_AS is set on the agent process and CURSOR_AGENT=1 on the
child/tool processes fm-harness.sh actually runs as.

Pane liveness. A cursor pane now classifies `agent`. An unrelated node or agent
stays `other`, which the liveness callers already fold into `ambiguous` rather
than `dead`, so a stranger's node pane is never reported agent-free.

Resolution prints the STABLE launcher rather than the canonical target: identity
is proven through canonicalization, but cursor's canonical path carries a
version its own auto-update replaces, and pinning that would strand a task on a
version that can vanish.

The regression drives the two identity signals apart - a cursor-named executable
outside any cursor tree, and a non-cursor-named alias inside one - and asserts
each carries a verdict alone, so no single vendor string is load-bearing. Its
negative controls are real spawned processes, not fixtures.

Verified live on cursor-agent 2026.08.11-e8db854.

Co-authored-by: Ville Penttinen <villem.penttinen@gmail.com>

* feat(cursor): classify cursor busy state from its own turn transcript

Cursor shipped as "unknown cursor-unverified" on the premise that it exposes no
semantic turn lifecycle, only a rendered "Working" footer. That premise is
wrong: cursor-agent persists an append-only JSONL transcript per conversation
and brackets every submitted turn with a role:user open and a typed turn_ended
close. Verified live on 2026.08.11-e8db854, including the interrupt path, where
Escape closes the turn with status "aborted" - so this source covers manual
interruption, which Claude's Stop hook does not.

That makes it a genuine pull source in the muse mould rather than the rendered
text the redesign forbids: no writer, no arm, no gen, nothing seeded that could
never be cleared. Cursor's `ctrl+c to stop` footer stays out of the verdict, and
herdr's narrower native streaming state cannot stand in for it either.

Binding deliberately does not reconstruct cursor's workspace-slug directory
name. That slug collapses path separators, so rebuilding it would be a guess
that could bind the wrong pane; cursor records the exact absolute workspace path
in each project's .workspace-trusted, and the binding matches on that. A
conversation recorded as prior at spawn is excluded, so a relaunch in a reused
worktree folds its own turn rather than its predecessor's. Requiring a unique
remaining conversation keeps zero and several both unknown, because neither
proves anything about the current turn.

The regression pins the fold with real transcript files and asserts the
dangerous direction stays closed: an unresolvable binding, a record-free file,
an unclaimed workspace, and a workspace-path PREFIX all read unknown, never
idle. The prefix case uses an opaque fixture slug so a slug-rebuilding
implementation cannot pass it.

Co-authored-by: Ville Penttinen <villem.penttinen@gmail.com>

* feat(cursor): make the cursor launch runnable and give it lifecycle control

Five gaps that together kept a cursor crewmate from being drivable end to end.

Launch. The template invoked `cursor agent`, but `cursor` is not the CLI - the
installed names are `cursor-agent` and the legacy alias `agent` - so the command
could not run at all on a machine with a normal cursor install. It now resolves
through the verified owner, which also refuses a spawn loudly instead of leaving
a pane that dies with command-not-found and reads as a wedged worker.

Session binding. fm-spawn writes state/<id>.cursor-session so the busy fold can
find this pane's transcript, and teardown removes it.

Lifecycle control. No cursor PR touched fm-control-lib.sh, so
`fm-control <id> interrupt|exit|relaunch` could not drive a cursor worker at
all. Verified live: interrupt is a single Escape, exit is /exit, and cursor does
NOT repollute its composer with the cancelled prompt, so unlike muse it needs no
clear key. Secondmate is refused, matching the spawn refusal.

Submit acknowledgement. cursor parks its terminal cursor outside its composer,
so the composer verdict on tmux is always `unknown` and a submit could never be
acknowledged from the composer alone. The submit core's existing idle-to-busy
transition covers that case, but only if the pane's busy footer is recognised,
so cursor's `ctrl+c to stop` joins the harness-less default union the submit
cores read. The TOKEN is matched rather than the spinner verb: the same version
rendered both `Working` and `Running` in consecutive turns.

Bootstrap. A configured cursor crew harness with no cursor executable is now a
loud MISSING diagnostic rather than a first-spawn failure, and it accepts either
installed name.

Interrupt cancellation is deliberately left unconfirmed. The transcript does
type an aborted close, but its post-interrupt write latency measured as
variable - sometimes seconds, sometimes not within twenty - so a claim built on
it would be unreliable. Normal turn completion is prompt, which is what the busy
fold actually depends on.

Two inherited tests are corrected rather than deleted: the busy test asserted
cursor could have no semantic source, and the launch test pinned the literal
`cursor agent` string. Both now pin the verified behaviour, including that the
launch never allocates a second worktree.

Co-authored-by: ABHISHAKE KUMAR BOJJA <abojja@uvic.ca>
Co-authored-by: Ville Penttinen <villem.penttinen@gmail.com>

* docs(cursor): record the verified crewmate facts and extend the drift guard

The inherited cursor entry was written against 2026.08.04-aaa8809 and several of
its claims no longer hold: it named `cursor agent` as the binary (not the CLI
name), listed six Grok model ids of which the live catalog now returns two, and
recorded busy state, exit, interrupt, and skill invocation as unverified.

Replaced with what was measured against 2026.08.11-e8db854, including the two
facts most likely to be rediscovered painfully: cursor runs as a bundled node
script so its pane title is a bare `node`, and it parks its terminal cursor
outside its composer, which makes the tmux composer verdict permanently
`unknown` by design rather than a defect to chase.

Model ids now route to `--list-models` for the account instead of a fixed list,
since that list is exactly what drifted.

The live drift guard covers cursor, resolving it through the same verified owner
fm-spawn uses and passing --trust so the probe cannot hang on the workspace
prompt. Run against every installed harness: 8 checked, all alive, with cursor
reporting title='node' foreground=[.../cursor-agent] - the drift shape this
guard exists to catch.

Co-authored-by: Ville Penttinen <villem.penttinen@gmail.com>

* docs(agents): record the cursor session-binding state file

The state/ layout section is the inventory every session reads; a busy-source
binding that fm-spawn writes and teardown removes belongs in it alongside muse's.

* no-mistakes(review): Sanitize ambient Cursor marker in harness tests

* no-mistakes(review): Validate Cursor models against live catalog

* no-mistakes(review): Reject unsupported secondmates before binary preflight

* no-mistakes(review): Narrow Cursor ancestry detection to structured process identity

* no-mistakes(review): Parse Cursor transcripts and sanitize inherited markers

* no-mistakes(review): Handle malformed Cursor transcript records safely

* no-mistakes(review): Validate malformed Cursor closes in fallback parser

* no-mistakes(review): Retire stale Cursor bindings during relaunch

* no-mistakes(review): Fix Cursor drift guard command variable

* no-mistakes(review): Narrow Cursor identity to versioned install trees

* no-mistakes(document): Document Cursor harness boundaries

* refactor(composer): move the delivery busy footers to the shared owner

The per-harness rendered busy footers lived in bin/fm-tmux-lib.sh under
FM_TMUX_* names, so cursor's `ctrl+c to stop` signature - and every other
harness's - was reachable only from tmux. That placement was wrong on its own
terms: herdr, zellij, cmux, and orca run the same harnesses and face the same
question these footers answer, which is whether a submitted Enter actually
landed. Nothing about the signature is tmux-specific.

Moved verbatim into bin/fm-composer-lib.sh, the shared composer/delivery owner
every backend already sources, and renamed to FM_DELIVERY_* so the names stop
claiming a scope they never had. All five adapters now reach cursor's signature;
verified per adapter rather than assumed.

The boundary the move must not blur is stated where it now lives: this is a
DELIVERY guard, never a worker-state source. Confirming a keystroke landed is a
different question from asking what a worker is doing, and bin/fm-busy-lib.sh
remains the semantic owner that forbids classifying a harness from rendered
text. Cursor still classifies only from its transcript fold, which is already
backend-agnostic because it folds a file rather than reading a pane - the same
verdict on all six backends.

The old FM_TMUX_* aliases are dropped rather than kept as dead shims: nothing
outside the moved block referenced them except fm-busy-lib.sh's grok fallback,
which now reads the new name. The documented operator override, FM_BUSY_REGEX,
is untouched.

Also removes a dead duplicate CURSOR_INVOKED_AS check in bin/fm-harness.sh,
unreachable behind the marker check above it.

* no-mistakes(review): Correct shared delivery guard ownership references

* no-mistakes(document): Document shared delivery guards and Cursor backend limits

* no-mistakes: apply CI fixes

* fix(composer): bound a bare composer's wrap region at a half-block rule

A live cursor crewmate on herdr classified its IDLE composer as `pending`, and
fm-send consequently exited 1 with "delivery unconfirmed" on a message that had
actually landed. The cause is not cursor-specific.

Herdr draws a composer's top and bottom rules with the half-block glyphs U+2584
and U+2580 rather than the box-drawing family. fm_composer_row_has_edge knew
only the box-drawing set, so no box was detected; the composer was found as a
BARE row, and its wrap region - which extends while rows are non-blank and carry
no structural edge - walked straight through the composer's own closing rule and
swallowed the model and path footer below it. That footer is real text, so the
region classified pending on a genuinely idle pane.

Teaching the shared edge detector the half-block glyphs bounds the region at the
closing rule. Measured on the captured bytes of a real herdr cursor pane: the
same capture that read `pending` now reads `empty`.

This is a shared shape-path change, so it is deliberately narrow - it adds
glyphs to the edge vocabulary and changes no verdict logic - and the whole
composer and backend suite is green, including the other harnesses' herdr
fixtures.

The regression pins the real captured shape and asserts the footer content is
genuinely present, so the case cannot pass vacuously if the region were ever
bounded for some unrelated reason.

* fix(herdr): confirm a cursor submit from the rendered-footer transition

Herdr's composer-shape fix made an idle cursor pane classify `empty`, but
`fm-send` still exited 1 with "delivery unconfirmed" on messages that had
actually landed. Live measurement found the second, independent cause.

Herdr reports a cursor pane `agent_status=blocked` in EVERY state - idle,
mid-turn, and after - so the submit path's idle-baseline native confirmation is
structurally unreachable for cursor and every send falls into the composer
branch. That branch reads cursor's mid-turn composer row, which renders its own
`Add a follow-up` placeholder beside a right-aligned `ctrl+c to stop`. That
token is composer content, so the verdict is `pending` on a composer holding no
user text at all, and the Enter-retry budget then reports pending.

The escape is the same semantic signal the native path uses, read from the
pane's verified busy footer instead of native agent-state, and it is the
rendered-footer twin of the tmux submit core's turn-started confirmation: an
idle-to-busy transition ACROSS our Enter proves the harness accepted the
submission. The baseline is taken before the first Enter and only when the
native baseline was not legibly idle, so the idle-baseline path still never
reads pane content and a pane already mid-turn before we typed keeps reporting
`pending` rather than borrowing another turn as proof of this delivery.

The composer verdict is deliberately NOT relaxed. A right-aligned status token
on the composer row stays content for every other caller, including the
away-mode pre-injection guard, and the shared cursorless submit core is left
untouched so zellij, cmux, and Orca keep the behavior their own follow-up owns.

Verified live on herdr 0.8.0 and cursor-agent 2026.08.11-e8db854 in an isolated
lab session: `fm-send` now exits 0 and the steer executes, interrupt cancels a
running turn, `/exit` stops the agent, and teardown clears the record. All seven
panes of the running default session classify identically before and after the
shape fix, so no other harness regressed.

* no-mistakes(review): Prevent working Herdr baselines from falsely confirming delivery

* no-mistakes(document): Correct Cursor harness and backend documentation

---------

Co-authored-by: ABHISHAKE KUMAR BOJJA <abojja@uvic.ca>
Co-authored-by: Amplify Logic AI <lars@sockinator.co>
Co-authored-by: Ville Penttinen <villem.penttinen@gmail.com>

* fix(bin): require quota-axi 0.1.25 (#2300)

* fix: raise quota-axi floor to 0.1.25 for Cursor CLI quota awareness

Homes on latest main need quota-axi #87 so Desktop-absent CLI machines report a fresh Cursor quota instead of a false sign-in-required.

* no-mistakes(document): Update quota floor documentation pointer

* fix(bin): prevent false Pi watcher alarms during hand-offs (#2304)

* fix(guard): stop the false send-time watcher-down alarm on Pi primaries

On a Pi primary the watcher process is not the liveness signal. The Pi
extension tears the watcher down on every actionable wake and spawns the
replacement itself, so the singleton lock is legitimately unheld between
cycles: every one of the 799 cycles in a live primary's ledger ends with
lock_after=pid:none, and a live capture caught the guard verdict flipping to
no-watcher during one hand-off with the beacon 63s old.

bin/fm-guard.sh classified Pi as a persistent-watcher harness, which demands a
live identity-matched lock holder at all times, so any guarded command landing
in a hand-off painted the full WATCHER DOWN - SUPERVISION IS OFF banner and
told firstmate to repair a cycle the extension already owns and is restoring.

Add an extension supervision model for pi and pi-signed. A live
identity-matched watcher stays the ordinary healthy state; an unheld lock is
healthy only while the beacon is fresh within grace AND a live Pi session
provably owns continuity - both primary extensions recorded in their state
markers at their current on-disk builds by the process named in state/.lock,
with that process still alive. Without that proof the banner fires exactly as
before, so an unloaded, version-drifted, or exited Pi session is loud
immediately and a cycle the extension never restores is loud once the beacon
passes grace. The queued-wake warning, the PID-strict turn-end guard, and
every other primary's detection are untouched.

Fold session-start's duplicate Pi marker predicate into the shared library so
the ownership contract has one owner.

* no-mistakes(review): Restrict Pi hand-off tolerance to unheld watcher locks

* no-mistakes(document): Document Pi watcher hand-off supervision

* feat: support Cursor Agent CLI as a primary harness (#2305)

* feat(cursor): add Cursor Agent CLI primary hooks, park supervision, and session start

Register a tracked project-scope .cursor/hooks.json for Cursor's stop,
sessionStart, preCompact, and preToolUse steps.

bin/fm-turnend-guard-cursor.sh owns Cursor's turn boundary as a park: it
foregrounds the watcher arm, holds the boundary open until an actionable close,
and returns that wake as one follow-up. Exit 2 is a silent no-op on Cursor's
stop step, so the adapter never uses it. The follow-up loop is bounded twice,
by Cursor's own loop_limit and by the payload's loop_count.

bin/fm-sessionstart-cursor.sh delivers the digest as additional_context at
sessionStart, and stages it for the next turn boundary at preCompact, which
cannot inject context.

Cursor also loads the tracked Claude settings, so bin/fm-hook-host-lib.sh lets
each tracked Claude-shaped entrypoint stand down on a Cursor-delivered payload
rather than running every covered event twice.

bin/fm-tmux-lib.sh reclassifies a Cursor pane's composer cursorlessly, because
Cursor parks its terminal cursor outside the composer, which restores a genuine
composer-empty proof and unblocks away-mode escalation delivery.

* feat(cursor): make Cursor Agent CLI a verified primary harness

Resolve Cursor in the session-lock ancestry through bin/fm-cursor-lib.sh, which
a Cursor primary needs before it can hold its own home lock, and classify its
stop-hook park under the autoarm supervision model so the mid-turn pull guard
stops reporting a healthy between-turns watcher as down.

Read a Cursor pane's composer cursorlessly on tmux, gated on Cursor's own
structural process identity, which restores a genuine composer-empty proof and
lets away-mode escalations reach a Cursor primary with no daemon change.

Lift the secondmate refusals in bin/fm-spawn.sh and bin/fm-control-lib.sh now
that the supervision protocol exists and is recorded.

Cover the whole surface with a portable regression over real processes, an
opt-in live guard against the installed cursor-agent, and dated per-harness
evidence.

* docs(cursor): record Cursor as a verified primary across the owning surfaces

Update the turn-end guard, session-start, arm-seatbelt, cd-guard, watcher
continuity, architecture, configuration, README, and harness-adapters owners,
and add dated live evidence to the supervision and runtime-backend verification
records. Correct the recorded Cursor tmux composer verdict: the cursor-anchored
read is still blind, but the composite reader is no longer unknown.

Lift the remaining remote-secondmate refusal missed in the previous commit, and
add the new libs to the existing fixtures that copy a fixed dependency list.

* refactor(cursor): name the park's stand-down condition for both its causes

Also record that Cursor's preCompact firing itself is not yet live-verified,
while the static evidence that it cannot inject context, and the staging path
that follows from it, both are.

* test: give the pretool fixtures their new dependency and one lint owner

The cd-guard fixture copies a fixed dependency list and now needs the shared
hook-host predicate. Both pretool suites also asserted cleanliness with a bare
shellcheck call, a second and weaker copy of the lint definition that
bin/fm-lint.sh owns: it omits --external-sources, so it failed the moment these
checkers sourced a shared library. They now delegate to that owner.

* test: assert the cursor secondmate contract instead of its removed refusal

A cursor secondmate now launches, so the suite asserts what its park actually
needs: --trust so the home's project hooks load at all, its own home pinned as
the workspace, and the autoarm supervision model inherited across the launch.

* no-mistakes(review): Serialize Cursor wakes and bind staged context

* no-mistakes(review): Serialize Cursor context and nag state commits

* no-mistakes(review): Enforce Cursor ceiling before staged context delivery

* no-mistakes(review): Serialize Cursor claims and staged context

* no-mistakes(review): Serialize Cursor ownership and state commits

* no-mistakes(review): Protect Cursor context across session takeover

* no-mistakes(review): Preserve Cursor context across session takeover

* no-mistakes(review): Enforce owner-keyed Cursor staged context

* no-mistakes(review): Atomically claim Cursor follow-ups and staged context

* no-mistakes(review): Defer Cursor preCompact staging and simplify supersession

* no-mistakes(review): Serialize Cursor park commits and defer preCompact

* no-mistakes(review): Stop Cursor parks after session takeover

* no-mistakes(test): Route Cursor preCompact context through stop follow-up

* no-mistakes(document): Update Cursor primary documentation

* revert(cursor): cut preCompact staging from this change

Carrying a compaction digest across two concurrently running stop hooks kept
producing races that could deliver it twice or strand it indefinitely, and
closing them kept enlarging a critical section inside a hook Cursor awaits at
the turn boundary. Native preCompact firing was never observed either, so the
surface has no empirical basis yet.

Remove the adapter, its registration, its staged path in the park, and its
tests, and record the surface as deferred and uncovered alongside the Codex
interactive TUI. A regression now asserts preCompact stays unregistered so it
cannot return without its own design and evidence.

This change ships the proven core only: the turn-end follow-up park, the
run-tier session start, and away-mode delivery.

* no-mistakes(review): Correct Cursor park supersession documentation

* no-mistakes(document): Clarify Cursor run-tier verification ownership

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

---------

Co-authored-by: kunchenguid <kun-1@kunchenguid.com>

* feat(bin): add decline and repair paths for decision holds (#2330)

* feat(bin): add unrouted close paths to the captain decision gate

A captain who declines a held decision leaves no follow-up work to route,
so `resolve` could not express that answer: it requires at least one
`--routed-to` task. The only way to close such a hold was a direct
`tasks-axi done`, which never writes the durable resolution record the
completion gate reads, so the originating investigation could no longer
pass `verify` and its cleanup stayed blocked.

Add two close paths that route no work:

- `decline` closes an actively held hold with a recorded captain decision
  and no routed task. It refuses while any task is still blocked by the
  hold, because releasing routed work without recording it is `resolve`'s
  job.
- `repair` records the missing resolution block on a hold that was already
  closed outside this script. It never reopens a hold and never clears a
  dependency edge, and it refuses a hold that is still actively held.

Both require a non-empty captain decision file and share `resolve`'s
digest-based retry identity, so an exact retry is idempotent while a
changed decision is rejected. The recorded body now also names which path
closed the hold, and each routed entry regains its own line.

The gate itself is unchanged: an unanswered decision still fails
completion and blocks teardown, and neither new path can close a hold
without the captain's recorded word.

* fix(bin): require captain-hold provenance before repairing a decision

`repair` checked only that the backlog item was kind captain and Done, so
an ordinary captain-kind task that was never held for the captain could be
closed, repaired, and then pass the completion gate.

tasks-axi keeps `hold_kind` through a close, so it is the surviving proof
that an identity really was a captain hold. Require it before writing the
resolution record, and cover the case in the gate regression.

* no-mistakes(document): Correct decision-hold lifecycle documentation

* fix(bin): surface buried wake status lines once (#2331)

* fix(bin): surface buried status notes on wake drain

A note: answer immediately followed by a routine note was dropped because
annotations kept only the newest line and note: never enters OPEN DECISIONS.
Present every unread note and pending-reply resolution since the last drain
cursor, and annotate every unread line on a queued signal.

* no-mistakes(review): Fix unread status cursor races and overflow

* no-mistakes(review): Preserve cursors when status span reads fail

* no-mistakes(review): Make status presentation transactional under I/O failures

* no-mistakes(review): Simplify unread status cursor and presentation locking

* no-mistakes(review): Align cursor failure regressions with transactional presentation

* no-mistakes(review): Retire stale presentation cursors during task teardown

* no-mistakes(review): Preserve routine status until signal annotation

* no-mistakes(review): Correct unread status cap documentation

* no-mistakes(document): Document unread wake status presentation

* no-mistakes(lint): Fix wake surfacing ShellCheck warnings

* no-mistakes: apply CI fixes

* feat: add max Calm presentation level (#2334)

* feat(calm): add a max presentation level that hides mid-turn working notes

Calm's home-local preference becomes a three-state level instead of a
boolean: "off" is stock Pi, "on" is today's Calm, and "max" is Calm plus
hiding the assistant text of messages the model did not end its response
with. `/calm max` selects it from any state, a plain `/calm` steps max
back to ordinary Calm and otherwise keeps the existing on/off cycle, and
any other argument keeps that cycle too.

`config/calm` now persists "max" as its own literal value, so a session
start, resume, fork, or reload restores the stored level rather than
treating it as unrecognized and dropping to off.

The hide rule keys on Pi's intrinsic per-message stopReason: "toolUse",
or "length" with tool calls present. Streaming ("pending") text is never
filtered, because suppressing it would also stop a genuine reply from
streaming. The existing assistant layout adapter filters the blocks out
of the same shallow presentation copy it already uses for collapsed
thinking, so the message, model context, session storage, /export, and
delivery are untouched and a hidden mid-turn row collapses to zero
height. The new "assistant-working-note" class keeps that choice in the
visibility policy owner, where ordinary Calm keeps it visible.

* no-mistakes(document): Clarify Calm max persistence and taxonomy

* feat(calm): hide mid-turn working notes by default (#2339)

* feat(calm): make hiding mid-turn working notes the ordinary Calm state

Calm collapses back to the two-state on/off toggle it was before the max
presentation level, with max's hide rule promoted into ordinary Calm.
Calm on now hides mid-turn assistant working notes in addition to what it
already hid, and the /calm command parses no argument again.

The hide rule itself is unchanged: assistant text is removed from the
shallow presentation copy when the message's own stopReason is "toolUse",
or "length" with tool calls present. Streaming ("pending") text is never
filtered, so a genuine reply still streams. The message, model context,
session storage, /export, and delivery remain untouched.

config/calm persists only "on" and "off" again, but the reader still maps
a persisted "max" to on so a home upgraded from the removed level keeps
Calm on instead of dropping to off.

The mid-turn hide is now default behavior rather than an opt-in level, so
docs/calm.md documents it for users, docs/configuration.md records the
two written values plus the legacy max mapping, and the feasibility
taxonomy drops its level-scoped wording.

* no-mistakes(document): Document ordinary Calm working-note hiding

* chore: store no-mistakes test evidence in the repo (#2355)

* chore: ignore scratchpad/ at the repo root (#2359)

* fix(ci): fail hung Herdr behavior runs in 20 minutes (#2413)

A wedged family-run step was occupying the runner until the 75-minute
job cap; bound that step so cleanup and timing artifacts still upload.

* fix: keep the public promise reachable when work is routed to a second mate (#2457)

The lightweight Relay follow-up link lives in the answering home's own
state/<task-id>.meta, so it can only bind work that home owns. When a
Relay-linked request is routed to a second mate, the task record lives in the
second mate's home, fm-x-link.sh failed with a bare "no such task ...meta", and
nothing else picked the promise up: only the soft acknowledgement was ever
posted. The typed promised-final path already supports --work-home
secondmate:<id>; the playbook simply never chose it.

- fmx-respond now states the routing rule crisply: a task in this home takes the
  lightweight link, and second-mate-route…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant