Skip to content

sync: finalize reconciled upstream merge - #62

Merged
DereKk8 merged 295 commits into
mainfrom
fm/fm-sync-finalize-merge-p3
Aug 15, 2026
Merged

DereKk8 merged 295 commits into
mainfrom
fm/fm-sync-finalize-merge-p3

Conversation

@DereKk8

@DereKk8 DereKk8 commented Aug 15, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • Integrated upstream/main into the fork in fresh two-parent merge commit 544d4d7.
  • Preserved the reconciled tree from 60f0c3f byte-for-byte.
  • Dispositions: 8 restored, 28 superseded, 3 reconciled.
  • Merge-content check: PASS with the documented intentional allowlist.
  • Reconciliation report: report.

Fix pass

  • jq_argjson_tempfile keeps tempfile-backed --slurpfile transport and falls back to Perl when mktemp is unavailable. fm-bearings-snapshot passes, and the stock macOS snapshot job passes.
  • Restored the /refactor-review README entry and the fork-only instruction contract needed by fm-instruction-owners; that test passes.
  • Updated the PR security test fixtures to model the merged guard's ordinary project and valid no-mistakes PR metadata. The real bin/fm-pr-check.sh guard was not weakened, and fm-pr-check-security passes.
  • Updated the restored decision-key test to the upstream stripped-token semantics; fm-wake-drain-open-decisions passes.
  • Updated the restored stow contract test to assert the current upstream tiered-memory lifecycle rather than superseded fork wording; it passes.
  • Updated secondmate safety fixtures with the valid project metadata required by the merged PR guard; that test passes.

Validation and attribution

The initial integrated full suite timed out at 1,800 seconds. Individual baseline verdicts against untouched upstream/main are recorded in the report:

  • fm-bearings-snapshot: MERGE REGRESSION, fixed by the guarded tempfile fallback.
  • fm-instruction-owners: fork-only test, fixed by restoring the retained documentation-inventory entry.
  • fm-on: NO REPRODUCTION in the individual rerun.
  • fm-pr-check-security: MERGE REGRESSION, fixed in the test fixture without weakening the guard.
  • fm-public-followup: NO REPRODUCTION in the individual rerun.
  • fm-remote-job-orphan-reap: PRE-EXISTING.
  • fm-remote-secondmate-lifecycle-e2e: PRE-EXISTING.
  • fm-calm-pi-extension: PRE-EXISTING.
  • Stock macOS Bash snapshot compatibility: MERGE REGRESSION in merge CI run 31908527065; untouched upstream CI run 31864677716 passed. The guarded tempfile fallback fixes it.

Run 31910945842 completed with the macOS job passing, but still contained the now-fixed stow contract mismatch and the now-fixed secondmate fixture failure, plus a non-reproducing fm-pi-watch-extension failure. CI run 31911959850 for commit af17e2b completed with every job passing except serial 4, where tests/fm-pi-watch-extension.test.sh failed once at hung-successor recovery.

Attribution of the remaining serial-4 failure: ten sequential local runs passed 10/10. The failing case has a 5-second outer polling window, 250 ms readiness timeout, 1 second retirement timeout, 5 ms and 10 ms retry delays, and a 100 ms stability check, leaving a wall-clock margin that a loaded CI runner can exceed. Both tests/fm-pi-watch-extension.test.sh and .pi/extensions/fm-primary-pi-watch.ts compare byte-for-byte identical with upstream/main. VERDICT: inherited upstream CI flakiness, not a merge regression. No assertion was loosened and the test was not disabled.

The separate body-compliance run reports the expected process failure because no-mistakes validation is intentionally held.

No-mistakes validation was NOT run because the captain requires explicit approval and a named model for every pipeline run. This PR must not be merged without that validation and the captain's explicit approval.

kunchenguid and others added 30 commits July 9, 2026 22:40
* fix(agents-md): inject self-governance section into existing AGENTS.md

fm-ensure-agents-md.sh only appended the canonical "## Maintaining this
file" section on skeleton create or CLAUDE.md promotion, so an existing
AGENTS.md that lacked it exited unchanged and forced hand-copying the
wording during a rollout across existing projects. Call the already-
idempotent ensure_maintenance_section on the existing-AGENTS.md paths and
report whether the file changed; a re-run and an already-complete file stay
byte-identical.

Also fixes kunchenguid#389: refuse a case-variant real memory file (e.g. a lowercase
agents.md) instead of silently emitting a CLAUDE.md symlink whose uppercase
literal target dangles once the tree lands on a case-sensitive filesystem.

Tests extend tests/fm-ensure-agents-md.test.sh; skeleton-create and
CLAUDE.md-promotion regressions still pass. Docs updated to match.

* no-mistakes(review): Captain: preserve CRLF maintenance-section idempotency

* no-mistakes(review): Preserve CRLF during maintenance-section injection

* no-mistakes(review): Captain: harden dangling-symlink regression coverage

* no-mistakes(document): Document agent-memory injection outcomes
* feat(secondmate): support project-less homes via --no-projects

fm-brief.sh --secondmate and fm-home-seed.sh now accept an explicit
--no-projects signal to scaffold, seed, and register a secondmate home
whose subject is the firstmate repo itself (no clones). The signal is
mutually exclusive with a project list; omitting both still fails loudly
so an accidental omission is never a silent project-less seed. The
registry line renders an empty projects: field, which spawn and the
snapshot already tolerate. Docs updated in the secondmate-provisioning
skill and both script headers.

* no-mistakes(review): Captain: document project-less secondmate flow

* no-mistakes(review): Captain: refuse project-less reseeding of populated homes

* fix(seed): fail closed on unreadable project data

* no-mistakes(review): Captain: reject stale projectful charters

* no-mistakes(review): Captain: fail closed on unsafe project paths

* no-mistakes(review): Captain: validate project-less charter clone sections

* no-mistakes(document): Document project-less secondmate seeding
* wip(handoff): record verified delegation design + tasks-axi mv blocker

No production code changed yet. tasks-axi mv (v0.2.1) cannot atomically
move a blocked-by-linked item set across backlogs (deadlocks both orders,
no batch/--force), which fm-secondmate-lifecycle-e2e requires. Parked
pending a tasks-axi connected-set mv enhancement; note captures the
verified design, semantics, test/CI/doc changes, and resume checklist.

* refactor(handoff): delegate the item move to tasks-axi mv

fm-backlog-handoff.sh's two-pass awk was a second parser of the backlog
format and the source of the PR kunchenguid#401 body-orphaning drift. Delete it and
delegate the move to `tasks-axi mv <id>... --to <dest>` (v0.2.2 atomic
multi-id), the single owner of the format: a connected set (blocker plus
dependents) moves together with blocked-by preserved, item blocks stay
byte-exact, and destination section placement holds. The helper keeps only
the fleet-level validation tasks-axi cannot know - secondmate-home
resolution, the seeded-home safety checks, the In-flight refusal, and
idempotent per-key reporting - and is atomic: on any move failure nothing
moves.

Tests: fm-backlog-handoff.test.sh keeps PR kunchenguid#401's regression matrix but now
exercises the delegated path and skips cleanly when tasks-axi is absent; the
two whole-file fixtures move to tasks-axi's canonical whitespace. The
lifecycle-e2e and safety move-cases gain the same skip guard. CI installs
tasks-axi so the delegated path is exercised. Docs state that
config/backlog-backend=manual governs firstmate's own hand-editing, not this
validated helper, which delegates fleet-wide because bootstrap requires
tasks-axi on PATH.

Remove the now-redundant WIP design note.

* no-mistakes(review): Captain: harden atomic backlog handoffs

* no-mistakes(review): Captain: enforce queued-only backlog handoffs

* no-mistakes(review): Captain: harden handoff section parsing

* no-mistakes(document): Document delegated backlog handoffs

* no-mistakes(lint): Silence ShellCheck source diagnostics
* fix: gitignore the secondmate home marker

bin/fm-home-seed.sh writes an untracked .fm-secondmate-home marker into
every seeded secondmate home. A secondmate home is a worktree of the
firstmate repo, so any plain `git status --porcelain` dirtiness check
counted the untracked marker and the home read as dirty forever:
fleet-sync reported it STUCK and the local fast-forward convergence
sweeps risked leaving it stale on firstmate updates.

Add .fm-secondmate-home to the tracked .gitignore so the marker is
invisible to every dirtiness check uniformly, without weakening
fleet-sync's deliberate untracked-counting for project clones.

Convergence chicken-and-egg: existing homes predate the fix and it only
arrives by fast-forward. The already-present marker-tolerant ff-skip
(ignore_seed_marker=yes, used by the bootstrap sweep, /updatefirstmate,
and spawn pre-launch) advances such a home past the fix commit, after
which .gitignore takes over - no hand intervention.

Tests in tests/fm-secondmate-sync.test.sh cover a freshly seeded home
reading clean, an existing marker-only home converging then reading
clean, and a genuinely dirty home still skipping.

* no-mistakes(review): Captain: document standalone-clone update path

* no-mistakes(document): Document secondmate marker migration
* fix(composer): stop reading dead-shell prompts as empty agent composers

Consolidate composer empty/pending/unknown classification into one shared
owner, bin/fm-composer-lib.sh's fm_composer_classify_content, delegated to by
all four backend adapters (tmux via fm-tmux-lib.sh, herdr, orca, cmux). This
replaces four drifting copies of the glyph decision.

Safety fix: a bare shell prompt glyph (> $ % #) on an unstructured row is now
classified unknown (a dead shell, unsafe for injection), not empty. It is only
empty inside a bordered composer box (the harness's own prompt). Agent glyphs
❯ (claude) and › (codex) read empty either way. The away-mode injector
(inject_msg) now requires an affirmatively-empty composer, deferring on pending
or unknown, so an escalation can never be typed into (or executed by) a pane
whose agent exited to its login shell.

Regression coverage: new tests/fm-composer-lib.test.sh pins the shared owner;
per-backend dead-shell tests in fm-daemon (tmux + injector), orca, and the
existing herdr/cmux suites. shellcheck clean; herdr incident regressions stay
green.

* no-mistakes(review): Captain: harden composer safety checks

* no-mistakes(test): Stabilize Herdr prune safety setup

* no-mistakes(document): Document composer injection safety

* no-mistakes(lint): Clean composer safety lint

* no-mistakes: apply CI fixes
* feat(watcher): add paused/awaiting-external crew state

A crew (or firstmate steering it) can declare a deliberate wait on a known
external dependency with a paused: <reason> status. Both the always-on watcher
and the away-mode daemon absorb such an idle pane through shared fm-classify-lib.sh
vocabulary instead of tripping the possible-wedge stale escalation, and re-surface
it for a recheck only on a long bounded cadence (FM_PAUSE_RESURFACE_SECS) so a
forgotten pause cannot rot invisibly. fm-crew-state.sh reports state: paused
distinctly. A crew that goes idle without declaring a pause classifies exactly as
before. Docs and brief scaffold state lists updated; tests colocated.

* no-mistakes(review): Captain: fix paused-state transitions

* init

* no-mistakes(review): Captain: fix paused-state supervision transitions

* no-mistakes(review): Captain: fix paused supervision handoffs

* no-mistakes(review): Reconcile paused supervision markers

* no-mistakes(review): Captain: prioritize paused states over captain relevance

* no-mistakes(review): Captain: preserve paused-working wedge timer

* no-mistakes(review): Captain: honor configured pause verb in briefs

* no-mistakes(test): Captain: fix AFK paused watcher handoff

* no-mistakes(document): Document declared external waits

* no-mistakes(lint): Clean paused-state lint

---------

Co-authored-by: fmtest <fmtest@example.invalid>
* fix(x-mode): make follow-up platform splitting immune to link ordering

A ~470-char Discord follow-up posted as a (1/2)(2/2) thread split at ~280
chars because fm-x-link only learned the platform from the inbox payload,
and the fmx-respond ack path can drain that inbox file before the task is
linked. A link recorded after cleanup silently lost the platform and the
splitter defaulted to the X 280-char budget.

Make platform resolution ordering-proof:

- fm-x-link now resolves the platform AUTHORITATIVELY by request_id via a
  new fmx_request_relay_context helper (POST /connector/request-context)
  when neither the inbox payload nor carry flags carry it. The request_id
  survives the inbox drain, so a post-cleanup link still learns the right
  split budget. Best-effort: no token/curl or a non-2xx relay degrades to
  the loud warning below rather than a silent X default.
- fm-x-link warns loudly when no platform source resolves, so the loss is
  never silent.
- The fmx-respond procedure now orders link-before-inbox-cleanup so the
  fast local path stays correct without a relay round-trip.

Colocated regression tests: a Discord follow-up >280 <2000 posts as ONE
message even when linked after inbox cleanup, and an unresolvable platform
warns loudly instead of splitting silently. docs/configuration.md documents
the request-context lookup.

The relay endpoint is the companion durable change (see done status); until
it ships, the link-before-cleanup reorder keeps the normal path correct.

* no-mistakes(document): Document X-mode platform recovery
* fix(composer): one ANSI-aware ghost owner covers claude dim + grok truecolor

Away-mode injection wedged all night on the primary claude-on-herdr pane:
the herdr composer classifier never stripped generic dim ghost text (only a
narrow codex bold-wrapped byte-pattern check), so claude's rotating
prompt-suggestion ghost - a bare "❯" then SGR-2 dim text, which herdr's ANSI
pane read preserves - read as real pending input and every escalation deferred
(6524 lifetime "pending input (non-empty composer)" defers; wedge 30623s).

Consolidate ghost extraction into one fleet-wide ANSI-aware owner,
fm_composer_strip_ghost (bin/fm-composer-lib.sh), that drops every
de-emphasised run - dim/faint (SGR 2: claude, codex) AND a dark/muted truecolor
foreground (grok's placeholder, luminance below FM_COMPOSER_GHOST_LUMA_MAX,
default 128, dark-theme assumption). Both ANSI-capable backends route through
it: fm_tmux_composer_state (fm_tmux_strip_ghost is now a thin adapter) and
fm_backend_herdr_composer_state. The herdr-only faint byte-pattern check is
removed and fm_backend_herdr_strip_ansi reduced to a thin adapter over the
shared fm_composer_strip_ansi. Bordered detection now reads the plain row so a
dark box border dropped with the ghost does not lose the composer shape.

This also closes the documented grok TRUECOLOR placeholder gap by the same
mechanism (harness-adapters skill note updated).

Empirical evidence (read-only live capture + isolated tmux, no herdr lifecycle)
and the incident write-up are in docs/herdr-backend.md; deterministic
regressions feed the exact captured bytes through the real classifiers
(tests/fm-backend-herdr.test.sh, tests/fm-composer-ghost.test.sh). Two prior
ghost-test fixtures that used a near-black 38;2;1;2;3 as "real" colored text
(never a realistic real-input color) are corrected to a bright 38;2;224;222;244,
preserving the truecolor payload-skip parser intent.

* no-mistakes(review): Preserve dark shell prompt safety

* no-mistakes(review): Harden erased shell prompt classification

* no-mistakes(document): Document shared composer ghost extraction

* no-mistakes(lint): Normalize tmux comment punctuation
…enguid#432)

* fix(session-start): isolate harness env markers in suite runner

Neutralize CLAUDECODE, PI_CODING_AGENT, and GROK_AGENT in
run_session_start so ambient interactive shells cannot override the
suite's fake ps harness (local-vs-CI split on the pi supervision case).

* no-mistakes(document): Correct Pi marker documentation
…nchenguid#435)

* fix(teardown): retry treehouse return on transient index.lock

Killed crew git ops can leave a short-lived worktree index.lock that
makes treehouse return fail. Retry on that error signature with a
bounded wait (env-overridable), never force-delete a live lock, and
only then fall back to the existing provably-stale cleanup path.

* no-mistakes(review): Harden teardown retry configuration

* no-mistakes(document): Document teardown index-lock retry behavior

* no-mistakes(lint): Fix empty shell variable assignments
* docs: de-feature the scripts.md and CONTRIBUTING test inventories

Slice 1 of the documentation redundancy cleanup wave (firstmate scope).

docs/scripts.md: every row is now one purpose clause; script headers
are the declared owner of behavior, flags, and contracts. Coverage
stays 61/61 scripts; bytes drop 19,922 -> 7,958.

CONTRIBUTING.md: the 54-row per-test inventory is gone; contributors
discover tests by listing tests/*.test.sh and reading each script's
own header, and gated tests print their own skip gates. The run
commands, symlink assertions, and watcher smoke line are unchanged.
Lines drop 135 -> 84 (18,797 -> 7,831 bytes).

Two facts that existed only as inventory rows moved into their
owners' headers first: fm-brief.sh's paused-vs-blocked scaffold
distinction and fm-session-start.sh's Pi extension-loaded check.

No instruction-surface or behavior change; AGENTS.md untouched.

* no-mistakes(review): Captain, fix brief help and Grok test discovery

* no-mistakes(review): Captain: document Grok lock-holder test coverage
* docs: consolidate universal backend contracts into configuration.md

Slice 2 of the documentation redundancy cleanup wave (firstmate scope).

docs/configuration.md is now the declared single owner of three
universal contracts, each with an explicit ownership sentence:
- the universal toolchain list (Toolchain), now also carrying the
  per-tool purpose clauses that previously lived only in the tmux guide;
- the task-selector vocabulary (Runtime backend);
- the tasks-axi compatibility definition (Backlog backend).

The five backend guides' prerequisites replace their verbatim
universal-requirements parentheticals (5 full copies) with a pointer
plus only backend-specific items; zellij/cmux selector restatements
and architecture.md's partial copy become pointers or are dropped;
CONTRIBUTING's compatibility sentence becomes a pointer; two
near-verbatim orca-bootstrap restatements (configuration.md Runtime
backend, orca guide) collapse into the Toolchain owner copy.

Backend-specific setup, behavior, target-string shapes, and every
empirical verification record are untouched. AGENTS.md untouched
(slice 3).

* docs: include git and GitHub auth in the toolchain owner list

The review flagged that the new universal-toolchain owner omitted git
and GitHub authentication while every backend guide now defers its
prerequisites here; bootstrap's NEEDS_GH_AUTH check makes them real
universal requirements.

* no-mistakes(review): Detect Git in bootstrap toolchain

* no-mistakes(document): Clarify GitHub CLI and centralize selector documentation
* feat(daemon): backend-independent active alert for the wedge alarm

When away-mode injection wedges past max-defer, inject_wedge_alarm only
actively signalled via the tmux status-line, which is skipped on non-tmux
backends. A wedged claude-on-herdr primary left only the passive
state/.subsuper-inject-wedged marker (2026-07-10 overnight incident).

Add a config-gated active alert (config/wedge-alarm, local/gitignored;
FM_WEDGE_ALARM_CHANNEL) that reaches the captain even when every pane and
its status-line is unreadable: an OS-level macOS notification (osascript),
a herdr notification, or a captain-supplied command. Default-on (auto) so
the alarm is never silent; each channel best-effort, degrading to the next
and never crashing the daemon loop. The tmux flash and durable marker stay.

The OS notifiers route through a single FM_WEDGE_ALARM_EXEC seam. When the
daemon is sourced (only tests do this; production execs it) the seam
defaults to "discard", and tests/wake-helpers.sh points it at a recorder,
so it is structurally impossible for any test to post a real notification.

Channels verified once manually on macOS 26.5.2 / herdr 0.7.3; see
docs/wedge-alarm.md.

* no-mistakes(review): Bound wedge alarm notifier execution

* no-mistakes(review): Captain: harden wedge alarm notifier safety

* no-mistakes(review): Captain: harden wedge alarm test notifier isolation

* no-mistakes(review): Captain: harden wedge alarm throttling

* no-mistakes(review): Redact wedge alarm directive logs

* no-mistakes(review): Harden wedge alarm notifier safety

* no-mistakes(review): Track notifier process groups through cleanup

* no-mistakes(document): Document wedge-alarm active alert behavior
* docs(agents): extract conditional AGENTS.md material to owned homes

Slice 3 of the documentation redundancy cleanup wave (firstmate scope):
the always-loaded instruction surface drops from 941 lines / 116,733
bytes (~29k tokens per session per fleet member) to 785 / 91,353
(~22.8k tokens), moving only audit-identified conditional and
situational material while preserving every load-bearing invariant at
its trigger point via the inline-stub pattern.

Moves, each to one declared owner plus an inline stub:
- section 3's bootstrap output-line handbook (~44 lines) -> new
  agent-only bootstrap-diagnostics skill, added to the section 13
  trigger index; the detect-consent-install rule and the
  do-not-dispatch gate stay inline as safety-critical.
- section 4's crew-dispatch JSON schema and field semantics ->
  docs/configuration.md 'Crew dispatch profiles' (pointer direction
  flipped); the intake procedure, precedence, backstop, and
  never-select-unverified rules stay inline.
- section 4's quota-balanced algorithm -> bin/fm-dispatch-select.sh
  header (now the declared owner; usage() converted to the dynamic
  header extraction pattern PR kunchenguid#438 established for fm-brief.sh).
- section 7's spawn resolution narrative and example sprawl ->
  bin/fm-spawn.sh header; the isolated-worktree assertion, refusal-is-
  a-blocker rule, and post-spawn duties stay inline.
- section 7's teardown landed-work mechanics -> bin/fm-teardown.sh
  header (section 1's containment pointer retargeted); the fork benign
  case and never-force rule stay inline.
- section 8's watcher classification narrative -> docs/architecture.md
  'Event-driven supervision' (already the owner); every operative rule
  (one live cycle, no turn ends blind, drain first, wake ladder,
  never-pkill, guard responses) stays inline.
- sections 3/4/6/7 secondmate sync, propagation, schema, and handoff
  restatements -> secondmate-provisioning skill, now the declared
  owner including the literal-file inheritance nuance.
- section 14's X-mode cadence mechanism -> docs/configuration.md
  'X mode (.env)', closing issue kunchenguid#363; activation semantics, the
  fmx-respond trigger, and the terminal-wake final-follow-up duty
  stay inline.

CLAUDE.md stays a symlink; no behavior or test change.

* no-mistakes(document): Centralize contract-owner documentation
* fix(cmux): close the last/selected workspace in a window at teardown

cmux keeps every window at >=1 workspace, so close-workspace on the only
workspace in a window silently no-ops (returns OK, workspace stays), and a
window holding a live session cannot be closed over the control socket.
That left a selected task workspace open at teardown (the last workspace
in a window is always the selected one).

Add fm_backend_cmux_window_of_workspace and have fm_backend_cmux_kill
create a throwaway default sibling in the target's window before closing
when the target is the last workspace there, so the close lands; the
window keeps a fresh default workspace (cmux's own "closed the last tab"
outcome). Non-last teardown closes directly, as before.

Cover both kill branches plus the helper with fake-CLI unit tests, add a
real-cmux window/count detection smoke assertion, and record the
empirical evidence in docs/cmux-backend.md.

* no-mistakes(review): Derive cmux count from membership snapshot

* no-mistakes(document): Document cmux last-workspace teardown behavior
…d#453)

* fix(fleet-sync): recover from an orphaned packed-refs.lock

A git ref rewrite (fetch --prune, pack-refs, branch -D) killed after
creating .git/packed-refs.lock but before renaming it - e.g. bootstrap's
timed-out fleet-sync kill or teardown's process kills - leaves a lock that
makes the next sync's fetch fail with "Unable to create
'...packed-refs.lock': File exists", leaving the clone unsynced.

On that signature only, fm-fleet-sync.sh now retries the fetch with a
bounded wait (transient locks self-clear), then removes the lock and
retries once more ONLY when it is provably stale: still present, mtime
age past a threshold, and no lsof holder of the lock file or of the clone
worktree itself (a live git keeps that as its cwd even in the window after
it closes the lock and before it exits). A live lock, a missing lsof, any
failed check, or any other fetch failure keeps today's behavior. Every
wait/retry/removal prints to stderr, and a successful recovery also prints
one "recovered:" summary to stdout so a session-start refresh - which
discards fleet-sync stderr and relays only stdout - still surfaces it.

The shared "is this git lock provably abandoned?" proof is extracted into
bin/fm-lock-lib.sh so it has one owner, used by both fm-teardown.sh and
fm-fleet-sync.sh. Constants are env-overridable knobs. tests/fm-gotmp.test.sh
gains the fm-lock-lib.sh symlink teardown now needs in its fake bin/.

* no-mistakes(review): Captain, remove obsolete teardown wake dependency

* no-mistakes(document): Document packed-refs lock recovery architecture
* feat(herdr): immediate blocked-state escalation via native events.subscribe push

Fold herdr's native pane.agent_status_changed stream into the single watcher so
a crew entering blocked wakes its supervisor sub-second (measured 0.129s)
instead of after the ~240s stale-pane wedge timer.

- bin/fm-transition-lib.sh: backend-neutral normalized-transition record shape
  plus the single-owner status->action policy table (blocked=actionable,
  working=absorb+clear-dedupe, idle/done=defer, else=fall back to polling).
- bin/backends/herdr.sh + herdr-eventwait.py: a raw AF_UNIX events.subscribe
  subscriber over one connection for all this home's herdr panes, subscribing to
  ALL statuses, returning the first fresh blocked edge, with a per-pane dedupe
  marker and a reconnect level-reconcile. Version/schema capability gate.
- bin/fm-backend.sh: has-push / events-capable / wait-transition dispatchers so
  the watcher stays backend-agnostic and the shape+policy are reusable.
- bin/fm-watch.sh: splice the bounded event wait in as the watcher's terminal
  wait primitive (replacing the blind sleep POLL for push-capable homes),
  behind a source guard so the splice is unit-testable; secondmate/paused
  exemptions; map pane->window->task and enqueue a stale wake. No second
  watcher process; the single-cycle invariant and every guard/beacon/turn-end
  mechanism are unchanged.
- Polling stays the permanent fail-closed backstop: below-capability, subscribe
  failure, and repeated runtime failures all degrade to sleep.
- Tests: fake-CLI units (fm-transition-lib, wait/apply/dedupe/reconcile/
  fallbacks in fm-backend-herdr, watcher exemptions in fm-supervision-events)
  plus an isolated real-herdr idle->blocked smoke. docs/herdr-backend.md carries
  the dated evidence and retires the old gap note.

* no-mistakes(review): Captain, fix Herdr disconnect handling and dedupe docs

* no-mistakes(review): Captain, commit markers after wake and reuse capability cache

* no-mistakes(review): Captain, clear stale markers and secure Herdr FIFOs

* no-mistakes(review): Captain, subscribe before Herdr reconciliation

* no-mistakes(review): Captain, make Herdr FIFO handling Bash 3.2-safe

* no-mistakes(test): Captain: include lock library in teardown fixture

* no-mistakes(document): Captain: document Herdr immediate blocked escalation

* fix: clarify shellcheck conditionals
* feat(bearings): deterministic bearings snapshot + durable decision model

Add bin/fm-bearings-snapshot.sh: a bounded TOON-by-default projection over the
canonical fm-fleet-snapshot. Default is local-only (zero network); live open-PR
discovery and checks happen only under --include-prs, which fails soft. Every
dropped surface is marked in omitted[] with the flag that reveals it, and the
prs: line states when checks were not requested, so absence is never silent.

Fix the unresolved-decision masking bug in the canonical layer. fm-classify-lib
gains status_open_decisions, the one authoritative keyed open/resolved fold over
the whole status stream: needs-decision/blocked opens a keyed entry, only an
explicit keyed resolution (or, for run-backed tasks, run-step advancement)
closes it, so a later unrelated done/paused can no longer mask a still-open
captain decision. fm-fleet-snapshot surfaces hints.open_decisions and derives
pending_decision/blocked_event from it; the canonical schema stays complete.

Point the /bearings skill at the one command; add the resolved: writer line to
ship, scout, and secondmate briefs. Register the script and add regression
tests for the output bound, TOON/JSON parity, local-only default, opt-in PR
fetch, partial-failure degradation, decision durability, and report pointers.

* fix(bearings): completed scout report is a pointer, not a pending decision

A completed scout that raised a needs-decision and then finished (done) without
a keyed resolution falsely surfaced as an open/pending decision (the Lavish-103
case). Root cause: the open-decision reconciliation in bin/fm-fleet-snapshot.sh
cleared a stale decision only for a live run-step/pane activity read, so a
terminal task whose current state is read from the status log (a scout or ship
that reached done/failed) never cleared its stale, never-keyed-resolved
needs-decision, and it lingered as pending.

The open-decision set is still derived purely from the keyed fold - never from a
report body or decision-like prose - and reconciled against the crew lifecycle.
Extend that reconciliation so a terminal done/failed state on a single-owner
task (scout or ship), whose deliverable is its report or PR, also clears the set;
a completed scout now surfaces only as a report pointer. Secondmates are excluded
from the terminal clear (persistent, multiplexed stream), which keeps the
unrelated-event masking fix intact. Add regression tests: a completed scout with
decision-like report prose is a pointer not pending (canonical + end-to-end), and
a scout still parked at a decision stays pending so the terminal clear never
over-fires.

* no-mistakes(review): Captain, preserve keyed decisions across shared status parsing

* no-mistakes(review): Captain, close blockers and harden keyed decision parsing

* no-mistakes(review): Captain, bound GitHub enrichment without coreutils timeout

* no-mistakes(review): Captain, bound bearings sections and fail closed

* no-mistakes(review): Captain, disclose capped per-repository PR results

* no-mistakes(document): Refresh bearings documentation and status contracts

* fix(bearings): avoid ambiguous worktree guard
* fix(lint): one shellcheck owner pinned to 0.11.0 for CI/local parity

Firstmate PRs passed local no-mistakes validation but failed CI's
"Lint shell scripts" job on shellcheck findings (SC2015, SC1007, SC2034).
Two divergences caused it:

1. The no-mistakes gate had no commands.lint, so its lint step never ran
   the deterministic shellcheck bin/*.sh bin/backends/*.sh tests/*.sh that
   CI runs. Confirmed from state.sqlite: the lint step_result recorded
   findings:null with no lint agent invocation.
2. CI's shellcheck floated with the runner image while local ran a newer
   build; shellcheck retired SC2015 in 0.11.0, so an older CI shellcheck
   rejected an SC2015 that the newer local one no longer emits.

Establish bin/fm-lint.sh as the single owner of the lint definition: the
file set, the config, and the pinned shellcheck version (0.11.0, printed
via --required-version). Both CI (.github/workflows/ci.yml) and the
no-mistakes gate (.no-mistakes.yaml commands.lint) invoke it; CI installs
the exact version it names and logs the resolved version, and fm-lint.sh
refuses to lint under any other version. This is not a CI relaxation: it
adopts shellcheck 0.11.0's rule set consistently, dropping only the
upstream-retired, false-positive-prone SC2015; default severity and every
still-supported finding stay enforced (no severity downgrade, no excludes).

tests/fm-lint.test.sh asserts both gates invoke the owner, that CI installs
and logs the pinned version, that the owner refuses a non-pinned shellcheck,
and that it rejects a real lint defect the old no-op gate passed.

* no-mistakes(review): Captain, harden deterministic ShellCheck parity

* no-mistakes(review): Captain, neutralize ambient ShellCheck overrides
* feat: add cd-guard PreToolUse seatbelt for the primary shell

A stray persistent top-level `cd projects/<clone>` in the primary firstmate
shell relocates the shell, so a later firstmate-owned command (a backlog write,
an fm-* lifecycle call, tasks-axi) runs inside a project clone instead of the
home. The cd-guard denies exactly that command shape before it runs, across all
five verified primary harnesses, mirroring the watcher-arm PreToolUse seatbelt.

- bin/fm-cd-command-policy.mjs: sole block/allow decision owner. Reuses the
  shell classifier exported from bin/fm-arm-command-policy.mjs (no duplicate
  lexer; that file's CLI now runs only when invoked directly).
- bin/fm-cd-pretool-check.sh: transport, strict-superset prefilter, harness
  output rendering, and primary-checkout scoping - fires in a secondmate's own
  primary session, inert in crew/scout child worktrees and non-firstmate repos.
- Wired into claude, codex, grok, opencode, and pi PreToolUse-equivalents;
  per-harness hooks only call the owner.
- Blocks top-level cd/pushd/popd (including cd to an absolute path, X=1 cd,
  and command cd). Allows git -C, subshell / bash -c / env -C / make -C /
  find -execdir, pipeline and background forms, and cd-as-data. Fails open on
  malformed input; agent-mistake threat model.
- tests/fm-cd-pretool-check.test.sh: 43-case x 5-harness-entry-form matrix,
  end-to-end cwd-leak regression, scoping, fail-open, prefilter, and wiring.
- docs/cd-guard.md: full contract plus live validation (claude, codex,
  opencode, pi blocked end-to-end; grok live run blocked by an API balance
  limit, with mechanism parity and deterministic coverage recorded).

* no-mistakes(review): Captain, fix cd-guard classification and prefilter coverage

* no-mistakes(review): Captain, allow path-qualified command wrappers

* no-mistakes(review): Captain, allow non-executing command queries

* no-mistakes(test): Captain, clarify cd-guard safe-path remediation

* docs: clarify cd guard guidance

* no-mistakes(document): Clarify cd-guard safe target guidance
…kunchenguid#267)

Crews must never stop, restart, or update the shared no-mistakes
daemon since one instance serves every firstmate lane/home; a restart
kills other lanes' in-flight pipeline runs and forces expensive
re-runs. Encodes this as a new numbered rule in both the ship-task and
scout-task brief scaffolds.

Co-authored-by: mielyemitchell <249051873+mielyemitchell@users.noreply.github.com>
* feat(bearings): four-section chat contract, accurate secondmate landed, resolved-event state render

/bearings skill (one owner of the chat-response format): mandate the four
always-present chat sections - Captain's Call, Recently Landed, Underway,
Charted Next - each with an explicit empty-state sentence, no At Anchor,
materially shorter than and linking to the report file. Resolves the ambiguous
Check first / Decisions pending split into one strict captain-action section.

fm-crew-state: the log fallback derives current state only from a real
run-state verb, so a trailing decision-closing resolved: event no longer
renders a healthy idle crew (typically a secondmate) as unknown with the
resolution prose as its detail. The keyed-decision contract in
fm-classify-lib.sh is untouched; map_log_state stays the one verb->state owner.

fm-fleet-snapshot: add a bounded, read-only secondmate_landed roll-up of Done
records from registered secondmate homes, reusing the single backlog parser and
the one secondmate-home enumerator (meta home= with data/secondmates.md
fallback); no network, per-home capped.

fm-bearings-snapshot: landed now merges main-home Done with the secondmate
roll-up, bounded by a per-home cap and an overall cap with omitted[] disclosure
(also fixing the previously-silent landed truncation); --all-landed reveals the
full set.

tests: resolved-event state render, secondmate landed aggregation with caps and
omitted[] disclosure, Captain's Call anti-leak, and the four-section contract.

* no-mistakes(review): Captain, ensure bearings reveals all landed work

* no-mistakes(document): Document bearings accuracy contracts
* fix: script-owned non-visible away-daemon launch + stale-artifact lifecycle

Away-mode entry left "make the daemon a tracked background terminal" to the
operator; on a pi/herdr primary that meant splitting the captain's active pane,
which visibly shrank it. Add bin/fm-afk-launch.sh, a single owner that launches
the daemon in a non-visible tracked terminal per backend (herdr dedicated
--no-focus workspace, detached tmux session), never a split, pins the captain
pane as FM_SUPERVISOR_TARGET/FM_SUPERVISOR_BACKEND, records the exact terminal
id, and tears it down or reconciles a leaked one by that id. No shell &.
Extract supervisor-pane discovery into bin/fm-supervisor-target-lib.sh, shared
with the daemon (one owner).

Fix the stale subsuper-artifact leak: clear the prior away session's delivery
cache on a fresh entry (fm_afk_clear_stale_artifacts), and stop the daemon
before clearing state/.afk so its shutdown flush runs instead of being a no-op.

Tests: tests/fm-afk-launch.test.sh (per-backend topology invariant in a lab
session, stale clear-on-entry vs refresh, exit ordering). Docs: /afk SKILL.md,
docs/herdr-backend.md (dated herdr evidence), AGENTS.md exit stub, docs/scripts.md.

* no-mistakes(review): Captain, serialize AFK launcher lifecycle safely

* no-mistakes(review): Captain, harden AFK launcher lifecycle races

* no-mistakes(review): Captain, ensure AFK daemon launch readiness

* no-mistakes(review): Captain, unify AFK lifecycle ownership and teardown

* no-mistakes(review): Captain, preserve AFK reconciliation records uniformly

* no-mistakes(review): Captain, harden AFK recovery state durability

* no-mistakes(review): Captain, harden AFK tmux ownership checks

* no-mistakes(review): Captain, simplify AFK lifecycle failure handling

* no-mistakes(review): Captain, require confirmed AFK daemon shutdown

* no-mistakes(review): Captain, confirm AFK exit by process identity

* no-mistakes(document): Align AFK launcher lifecycle documentation

* no-mistakes: apply CI fixes
…uid#518)

* feat: contain no-mistakes gate agents from driving the fleet

Add bin/fm-gate-refuse-lib.sh, sourced at the top of fm-spawn/fm-send/
fm-teardown before any fleet mutation. It fails closed when NO_MISTAKES_GATE
is set, and via an unspoofable git-common-dir backstop when invoked from a
no-mistakes gate worktree (.no-mistakes/repos/*.git) even with the marker
unset. A normal firstmate session has neither signal and is unaffected.

Set disable_project_settings: true in the tracked .no-mistakes.yaml so the
installed pipeline neutralizes gate agents' project instructions for this repo
(trusted-only, honored from the default branch).

firstmate's own suite runs from a gate worktree during validation, so the
shared test helpers set FM_GATE_REFUSE_BYPASS=1 to exempt it; the dedicated
tests/fm-gate-refuse.test.sh strips it to verify real refusal.

* no-mistakes(review): Captain, refuse empty no-mistakes gate markers

* no-mistakes(document): Document no-mistakes gate authority boundary
…uid#505)

* fix: guard secondmate own-home turn ends

Remove the .fm-secondmate-home early-exit in fm-turnend-guard.sh so the
'no turn ends blind' backstop fires in a secondmate's own primary session,
matching the cd-guard's scope: the own home is guarded, child crew/scout
worktrees stay exempt via the retained git-dir/git-common-dir test. This
was pure scoping from the guard's primary-only origin and guarded against
no secondmate-specific hazard.

Add secondmate regression tests (blind-turn block, idle-by-default,
stop_hook_active loop guard, deferred-death recovery loop, child-worktree
exemption) and record the autonomous background-notify re-invoke
measurement (Claude Code 2.1.207, 11s) in docs/turnend-guard.md.

* no-mistakes(document): Correct secondmate guard documentation, captain

* fix: force-include marked secondmate homes in turn-end guard

The prior remove-only form (just deleting the .fm-secondmate-home check)
left the DEFAULT secondmate topology unguarded: a treehouse-leased home is
a linked git worktree (git-dir != git-common-dir), which the retained
git-dir exemption still skipped, so its own primary session could still end
a turn blind. Invert the marker: a genuinely-marked home is force-included
as a guarded primary (treehouse-leased linked OR git-cloned plain), and the
git-dir exemption applies only to UNMARKED child worktrees. Marker
validation (regular non-symlink file, non-empty id-token content) blocks a
stray or empty marker from spoofing inclusion.

Add real linked-worktree regression tests: a treehouse-leased LINKED
secondmate home is guarded, a stray/empty marker stays exempt, and the
unmarked child worktree stays exempt - the topology the plain git-init
fixtures masked. Predicates, in-flight gate, and loop guard untouched.

* fix: force ASCII collation in secondmate marker validation

Add a function-scoped local LC_ALL=C in fm_root_is_secondmate_home so the
[A-Za-z0-9._-] id allowlist matches under C collation, not the ambient
locale - a locale-crafted non-ASCII marker id can no longer slip through
the range match and spoof force-inclusion of a linked child worktree.
Add a regression test proving a non-ASCII marker id is rejected and the
linked worktree stays exempt.

* no-mistakes(test): fix backend baseline gate-refusal dependency

* no-mistakes(document): Correct secondmate turn-end guard documentation
* fix: make bootstrap required-tool detection backend-aware

Bootstrap demanded tmux and treehouse for every backend except orca, so a
herdr/zellij/cmux home with tmux absent was wrongly told MISSING: tmux.

Required tools now follow the resolved backend via the single-owner
fm_backend_required_tools helper (bin/fm-backend.sh): each backend's own
session-provider CLI, jq for the JSON-emitting adapters (herdr/zellij/cmux),
and treehouse for session-provider-only backends (orca owns its worktree).
The treehouse lease-support check is gated to backends that use treehouse.

Adds install hints for herdr/zellij/cmux, regression tests for the full
backend dependency matrix (herdr-without-tmux repro plus each boundary),
and updates the authoritative Toolchain docs.

* no-mistakes(review): Captain, prevent executing Herdr install guidance

* no-mistakes(review): Captain, harden backend-aware bootstrap diagnostics

* no-mistakes(review): Captain, separate manual dependency remediation

* no-mistakes(review): Captain, align bootstrap diagnostic consumers

* no-mistakes(document): Align backend adapter dependency comments
…guid#520)

* fix: recover X/Discord follow-up platform after inbox cleanup

A milestone follow-up posted directly by request_id after the inbox was
drained - and with no task link, because one persistent secondmate's single
x_request slot collides across concurrent requests - resolved platform only
from the local inbox, so a >280 Discord reply silently defaulted to the X
280-char budget and threaded as (1/2).

- fm-x-poll records a durable per-request reply context
  (state/x-context/<rid>.json) at stash time, keyed by request_id so
  concurrent requests never overwrite each other; it survives inbox cleanup
  and restart.
- fm-x-reply resolves platform/budget through registry -> inbox -> relay
  (the relay lookup confined to a live follow-up), recovering the original
  platform independent of task-link availability.
- Fail-safe: a follow-up whose platform/budget cannot be authoritatively
  resolved and that would split is refused (exit 8) and held for retry,
  never wrongly split; fm-x-followup keeps the link on that exit.
- fm-x-dismiss clears the durable context for a dismissed mention.

Refactors reply-context extraction into a single owner and adds regression
coverage for all four cases.

* no-mistakes(review): Captain, fail closed on incomplete follow-up context

* no-mistakes(review): Captain, bound X context registry retention

* no-mistakes(review): Captain, align context retention with answer binding

* no-mistakes(document): Align X follow-up context documentation

* no-mistakes(document): Align durable X follow-up documentation
…id#533)

* fix: preserve secondmate routing markers

* no-mistakes(review): Captain, preserve trailing newlines in marked secondmate sends

* no-mistakes(test): Captain, tolerate bootstrap timeout elapsed drift

* no-mistakes(document): Refresh Herdr marker documentation
kunchenguid and others added 29 commits August 10, 2026 18:52
…enguid#2116)

* fix(spawn): refresh pooled worktree base

* no-mistakes(document): Document spawn base-freshness invariant

* no-mistakes: apply CI fixes
…#2102)

* refactor(composer): one shape owner behind thin capture adapters, whole matrix fixed

Consolidate every composer shape - bordered boxes (all families, geometry,
titled bottom borders), bare agent-glyph rows and their wrap regions,
opencode's left bar, and pi's identity-gated separator pair - into
fm_composer_classify_screen in bin/fm-composer-lib.sh. Adapters now
contribute only a capture and a declarative capability descriptor
(styled/cursor/identity/rows); capability differences change how confidently
a shape is judged, never what the shapes are, so a new harness shape is
teachable in exactly one place.

Correctness fixes landed as part of the consolidation (audit
data/fm-composer-consolidation-audit-s1):
- locale-safe Unicode-space normalization in the shared owner (closes the
  fleet-wide half of kunchenguid#1988; cmux's local byte-exact NBSP case deleted;
  naming converges with PR kunchenguid#1995's normalization primitive)
- muse's bare glyph joins the shared set, unbreaking muse on herdr/cmux/orca
- orca learns the borderless bare shape, drops its backward-paged composer
  window, and can no longer classify a stale startup banner as the composer
- tmux tolerates a titled bottom border, unbreaking grok steering
- the left-bar shape makes opencode readable on every backend
- zellij gets a real classifier through dump-screen --ansi, replacing the
  content-diff submit heuristic that could confirm an undelivered message
  and close a --resolve-key decision (the fleet's only false positive)
- fm-spawn's kimi launch-readiness regex (the fourth shape copy) now routes
  through the shared classifier

The strict blank-row posture applies fleet-wide (captain decision
blank-row-injection-posture): no positive container proof = unknown = defer,
replacing tmux's permissive blank-cursor-row rule. Away-mode injection was
re-validated end to end on real tmux (defer on partial input and unproven
rows, clean delivery with swallowed-Enter retry into proven-empty
composers). The tmux submit core gains a baseline-idle turn-started
conversion so pi steering stays confirmed while its working screen hides
the composer; busy conversion without that baseline remains forbidden.

Plain-capture backends now degrade a glyph row carrying trailing text to
unknown instead of a false pending, per the approved capability rule.

Portable regressions pin the full byte-capture matrix from the audit under
a UTF-8 locale and LC_ALL=C, the strict-vs-permissive divergence, and
deliberate signal separation; the opt-in live guard
(tests/fm-composer-matrix-live-e2e.test.sh) verified every installed
harness against the real classifier, recorded in
docs/verification/runtime-backends.md.

* no-mistakes(review): Fix Pi glyph ambiguity and complete profile matrix

* no-mistakes(review): Preserve bare verdict when Pi identity probe is absent

* no-mistakes(review): Harden composer structure and titled-border geometry

* no-mistakes(review): Require proven idle baseline and strict Zellij guard

* no-mistakes(review): Reject box bottom borders as composer input rows

* no-mistakes(review): Prove Zellij probe typing before classifier retries

* no-mistakes(review): Preserve Pi identity uncertainty and scan full left-bar drafts

* no-mistakes(review): Verify Zellij text lands before submitting

* no-mistakes(review): Scope Zellij typing verification to selected composer content

* no-mistakes(review): Verify Zellij pastes through composer-scoped content deltas

* no-mistakes(review): Prove wrapped bare Zellij pastes through composer extraction

* no-mistakes(review): Invalidate stale cursorless composers below dead shell prompts

* no-mistakes(review): Handle shell prompt placeholders in composer extraction

* no-mistakes(review): Classify cursorless bare continuation regions safely

* no-mistakes(review): Reject stale cursorless containers below live activity

* no-mistakes(review): Preserve prompt glyphs in wrapped Zellij pastes

* no-mistakes(review): Reject live shell rows during composer extraction

* no-mistakes(review): Preserve wrapped glyph continuations through submit retries

* no-mistakes(review): Scope idle placeholders to proven positions

* no-mistakes(review): Restore boxed placeholders and live prompt reanchoring

* no-mistakes(review): Fix Zellij placeholder and wrapped glyph paste proof

* no-mistakes(document): Align composer architecture documentation

* no-mistakes(lint): Fix ShellCheck warnings in composer refactor

* no-mistakes: apply CI fixes

* docs(verification): record the trusted-checkout live matrix rerun

The pipeline's isolated gate worktree is untrusted, so claude, grok, and
muse stopped at first-launch trust dialogs there (the guard refuses to
confirm them by design). This rerun from the trusted checkout at the final
validated head verified all six installed harnesses, the strict blank-row
deferral, and the hardened zellij false-positive probe live.

* no-mistakes(document): Align composer verification evidence

* no-mistakes: apply CI fixes

* no-mistakes(review): Restore proven box bottom-cursor classification

* no-mistakes(review): Preserve styled placeholder-like drafts as pending

* no-mistakes(document): Align composer safety and Zellij delivery documentation

* no-mistakes: apply CI fixes

* docs(verification): refresh the live matrix with the final-head trusted rerun

The post-validation rerun from the trusted checkout verified all six
installed harnesses at the branch's final head, including Claude 2.1.227
(auto-updated since the audit's captures) and Grok, which the untrusted
gate worktree could not verify past their first-launch trust dialogs.
* fix(spawn): gate Pi regular TUI flag by capability

* no-mistakes(review): Document conditional Pi TUI capability detection

* no-mistakes(review): Pin Pi probing and launch to one executable

* no-mistakes(review): Preserve literal pinned Pi paths and update documentation

* no-mistakes(review): Defer pinned Pi path insertion until final substitution

* no-mistakes(document): Document version-safe Pi launch probing

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes
…unchenguid#2147)

* docs(vision): elevate experience, pain narrative, and distro virtues

Fold the captain's public vision framing into VISION.md: peace of mind as a
primary goal, multi-session context-switch pain as the problem one interface
solves, clone-and-run setup ease, self-evolution including community, and
explicit harness/backend orthogonality. Reconcile experience-as-garnish into
experience-as-purpose and update aligns/resists accordingly.

* docs(vision): state the experience goal positively

Drop the negative "not a smart workflow / useful tool / impressive technology"
pretext. Lead straight into the positive experience north star.
* fix: reconcile inactive terminal outcomes

* fix: stream secondmate summary inputs

* no-mistakes(review): Fix reconciliation locking and request delivery retries

* no-mistakes(review): Prevent retries after unknown request delivery

* no-mistakes(document): Clarify inactive reconciliation cadence and receipts

* no-mistakes(lint): Quote terminal status arguments in reconciliation tests

* refactor: simplify inactive outcome reconciliation

* no-mistakes(review): Bound inactive reconciliation scans with durable progress

* no-mistakes(review): Bound reconciliation and deduplicate recovery notices

* no-mistakes(document): Document inactive outcome reconciliation contracts

* no-mistakes(review): Reject relative local secondmate parent routes

* no-mistakes(review): Key terminal receipts by spawn incarnation

* no-mistakes(review): Stabilize legacy receipts and lock reconciliation snapshots

* no-mistakes(review): Fail closed on invalid secondmate identity markers

* no-mistakes(document): Document durable inactive-outcome reconciliation

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes
* fix(session-start): refresh drifted instructions on stale rebuilds

* test(session-start): prove Pi instruction refresh end to end

* no-mistakes(review): Fix stale instruction refresh and baseline integrity

* no-mistakes(review): Preserve true-start baselines across Pi continuations

* no-mistakes(review): Correct Pi continuation classification and live expectation

* no-mistakes(review): Correct Pi continuation coverage documentation

* no-mistakes(review): Fix read-only refresh and exact Pi session restores

* no-mistakes(review): Classify Pi create-if-missing sessions correctly

* no-mistakes(review): Classify named Pi sessions using immutable headers

* no-mistakes(review): Correct Codex interactive coverage diagnostic

* no-mistakes(document): Document immutable Pi compaction instruction refresh

* no-mistakes(document): Correct Pi refresh documentation and validation claims
* feat(bin): add deterministic condition->action watch adapter on the process-event channel

Register a (condition, action) pair once with bin/fm-procevent-when.sh and the
existing process-to-event runner polls the condition tokenlessly, fires the
action at most once on a stable true, and wakes firstmate exactly once with the
captured outcome - instead of burning an agent turn per re-check.

The pair is stored privately under state/when/ and hash-bound by a trust record
the same way fm-check-register.sh binds a custom check, so a mutated spec is
refused without executing anything. A durable exclusive fired marker claimed
before the action makes restarts and re-polls unable to double-fire; every
failure path (mutated spec, condition error past budget, expired deadline,
failed action, uncaptured earlier fire) ends in a terminal captured outcome
that wakes firstmate rather than a silent retry. Eligibility stays a firstmate
judgment: only exact, safe, reversible actions may be bound, and judgment-
needing or destructive actions keep the wake-and-decide flow.

* no-mistakes(review): Harden when watcher concurrency, deadlines, timeouts, and output

* no-mistakes(test): Bind watcher actions to registered executable bytes

* no-mistakes(document): Correct condition-action watcher documentation

* no-mistakes(document): Clarify outcome wake re-announcement

* no-mistakes: apply CI fixes
…id#2202)

The open-decisions fold only recognized a [key=<slug>] token between the
verb and the colon (needs-decision [key=x]: note). The common worker
shape with the colon first (needs-decision: [key=x] note) silently
folded its stated key into the shared "default" bucket, so two open
decisions could collapse into one record and fm-send --resolve-key <x>
refused to close the decision it plainly named.

A complete token at the head of the note is now an equivalent stated-key
position for every keyed verb, shared by the whole-file and incremental
folds through the one _fm_decision_key owner. The documented
before-colon position wins when both are present, a token deeper in the
note stays prose, a bare keyless line still folds to "default", and a
stated-but-malformed slug is rejected rather than rewritten to
"default". A consumed note-head token is stripped from the note so both
positions yield identical records, and the incremental fold version is
bumped so persisted cursors folded under the old interpretation are
rebuilt from the authoritative log.

Fixes kunchenguid#2109
…uid#2212)

* fix(bin): keep a recovery acknowledgement valid across republication

A watcher cycle that opened and closed while the model handled its drained
wakes minted a fresh recovery generation, which invalidated the exact
acknowledgement the drain had just printed. That acknowledgement then consumed
nothing, so the marker stayed pending and every later arm spent its whole cycle
re-announcing the same recovery instead of supervising - a livelock the home
could not leave on its own.

A downtime publication now reuses the generation of an outstanding handling
episode, so a close during the handling window cannot orphan the printed
acknowledgement. The acknowledgement itself separates its two facts: queue-row
consumption is bound to the monotonic --ack-through sequence and always
happens, while only retiring the episode is bound to --recovery-generation. A
generation that moved on is a non-fatal result that names its own remedy
instead of a refusal that consumes nothing.

* no-mistakes(review): Preserve recovery generations and consume stale acknowledgements safely

* no-mistakes(document): Document sequence-bound recovery acknowledgements
* feat(fmx-respond): consume in_reply_to_chain conversation context

The relay's poll payload can carry in_reply_to_chain, an oldest-first
transcript of the surrounding conversation, but the mention-handling
procedure only ever read the immediate in_reply_to parent, so referents
like "this" in a standalone mention stayed unresolvable even when
context was delivered.

Teach fmx-respond to read the chain when present (optional and
backward-compatible: often absent today, kind label not required),
resolve referents against the whole transcript, and extend the
untrusted-content framing to every chain entry including the upcoming
kind=history entries. Document the field's wire shape in
docs/configuration.md as the firstmate-side owner.

* no-mistakes(document): Document Relay chain context ownership
* fix(bin): strip every bracket tag, not just [key=...], from a status verb

status_line_verb only stripped a leading "[key=...]" token before the
colon, so a remote secondmate reply's leading "[corr=...]" correlation
tag stayed glued onto the returned verb word ("needs-decision
[corr=...]" instead of "needs-decision"). The open-decisions fold's
verb match then silently failed to recognize the line at all, so
fm-send --resolve-key refused to close a decision that was plainly
open on the status line.

Generalize the parser to strip every "[name=value]" tag before the
colon, in any order and count, so local and remote replies fold
identically.

* no-mistakes(review): Invalidate stale decision cursors after parser fix

* no-mistakes(document): Clarify status metadata verb parsing
* fix: collapse duplicate supervision wakes without losing legitimate updates

One remote-secondmate note produced two handling turns (a procevent check
wake published before autohandle, then a signal wake for the same mirrored
bytes), already-ingested replays such as a cursor-loss whole-log recapture
still woke with nothing to do, this home's own bookkeeping closes (fm-send
--resolve-key, the pending-reply escalation close, the captain-held
transfer) re-woke the session that wrote them, and turn-ended-only wakes
were annotated with already-announced status lines that looked like fresh
progress.

Dedup rules, each at its layer's one owner:
- fm-procevent.sh: an adapter may declare 'self-announcing'; the runner
  then applies first and publishes a check wake only for what remains
  unhandled. fm-procevent-remote-reply.sh declares it: the mirrored status
  append is the single announcement, so a fully applied capture publishes
  nothing and a byte-identical replay stays completely quiet. All other
  adapters keep strict publish-before-apply.
- fm-wake-lib.sh: fm_wake_signal_sig/seen_path/seen_current now own the
  watcher's signal signature and .seen-* marker format, plus
  fm_wake_status_append_self_announced, the guarded bookkeeping append
  that advances the marker only over exactly its own bytes and fails
  toward waking on any pending or interleaved foreign write.
- fm-send.sh, fm-pending-reply-lib.sh, fm-decision-hold.sh: bookkeeping
  closes go through that guarded append; escalation opens stay plain
  appends because a new blocker must wake.
- fm-wake-lib.sh annotations: a historical (turn-ended-only) row skips its
  status annotation only when the file's signature provably matches the
  seen marker; anything unannounced keeps annotating.
- fm-classify-lib.sh: a kind=secondmate task's status signal is never
  absorbed as provably-working, because that stream is the routed-reply
  channel the parent must read.

Also fixes a pre-existing exit-path deadlock the regression run reproduced:
a TERM inside a recovery-marker critical section left fm_lock_try_acquire
spinning against this same process's abandoned hold; a self-held lock is
now reclaimed (a subshell still waits on its parent's live hold).

Regression tests drive the real wake functions and executables in both
directions: each duplicate case collapses, while a new remote reply, new
decision, new blocker, merge result, failure, first status change, and a
later different note on the same task all still wake.

* no-mistakes(document): Document wake deduplication contracts
* feat(harness): add Cursor Agent CLI adapter

# Conflicts:
#	bin/fm-spawn.sh

* fix(composer): read cursor-agent's reverse-video placeholder as idle

cursor-agent renders its idle composer placeholder dim (SGR 2) but paints the
cell under the terminal cursor in reverse video (SGR 0;7). Reverse video is
neither dim nor a dark truecolor foreground, so the shared ghost stripper keeps
that one character and an idle composer reduces to a lone `P`. Judged on its
own, that remnant reads `pending` on a genuinely idle pane, which defers
away-mode escalation indefinitely on the styled cursorless backends.

Teach the ONE fleet-wide classifier the shape instead of adding an adapter-local
copy: register `→` as an agent prompt glyph so the composer row is structurally
findable at all (without it the bottom-most shape is a stale shell prompt echo
in the scrollback), add both verified placeholders to the idle set, and consult
the styling-independent plain row when the styled row is only a remnant.

The plain-row branch demands the remnant be a proper, strictly shorter substring
of a plain row matching a fully anchored placeholder. Real typed text is
uniformly bright, so stripping leaves it equal to the plain row and it stays
`pending` - verified live against a pane where the typed text was exactly the
placeholder string.

Verified live on cursor-agent 2026.08.11-e8db854; the regression pins the real
captured bytes and asserts the remnant survives stripping, so the case cannot go
vacuous if the stripper later learns SGR 7.

Co-authored-by: Amplify Logic AI <lars@sockinator.co>

* feat(cursor): narrow cursor identity and order its marker before CLAUDECODE

Cursor ships two executable names - `cursor-agent` and the legacy alias `agent`
- and runs as a bundled node script, so tmux reports the pane command as a bare
`node`. Neither `agent` nor `node` can be trusted by name, so identity gets one
owner in bin/fm-cursor-lib.sh that demands cursor's own name or install tree in
the path or argv[0], from the structural signal only. Probing an arbitrary pid's
executable during a liveness poll would execute a stranger's binary, which is
the hazard that rule exists to close.

Two consequences wired up:

Detection. cursor-agent does NOT clear an inherited CLAUDECODE, so a cursor
worker launched under a claude primary carries both markers and whichever is
tested first wins. The cursor markers are ordered ahead of the CLAUDECODE check;
fm-spawn additionally clears foreign markers at the launch boundary. Both are
kept deliberately - launch sanitization only covers sessions fm-spawn started,
while the ordering also covers a cursor session started by hand. Verified live
that CURSOR_INVOKED_AS is set on the agent process and CURSOR_AGENT=1 on the
child/tool processes fm-harness.sh actually runs as.

Pane liveness. A cursor pane now classifies `agent`. An unrelated node or agent
stays `other`, which the liveness callers already fold into `ambiguous` rather
than `dead`, so a stranger's node pane is never reported agent-free.

Resolution prints the STABLE launcher rather than the canonical target: identity
is proven through canonicalization, but cursor's canonical path carries a
version its own auto-update replaces, and pinning that would strand a task on a
version that can vanish.

The regression drives the two identity signals apart - a cursor-named executable
outside any cursor tree, and a non-cursor-named alias inside one - and asserts
each carries a verdict alone, so no single vendor string is load-bearing. Its
negative controls are real spawned processes, not fixtures.

Verified live on cursor-agent 2026.08.11-e8db854.

Co-authored-by: Ville Penttinen <villem.penttinen@gmail.com>

* feat(cursor): classify cursor busy state from its own turn transcript

Cursor shipped as "unknown cursor-unverified" on the premise that it exposes no
semantic turn lifecycle, only a rendered "Working" footer. That premise is
wrong: cursor-agent persists an append-only JSONL transcript per conversation
and brackets every submitted turn with a role:user open and a typed turn_ended
close. Verified live on 2026.08.11-e8db854, including the interrupt path, where
Escape closes the turn with status "aborted" - so this source covers manual
interruption, which Claude's Stop hook does not.

That makes it a genuine pull source in the muse mould rather than the rendered
text the redesign forbids: no writer, no arm, no gen, nothing seeded that could
never be cleared. Cursor's `ctrl+c to stop` footer stays out of the verdict, and
herdr's narrower native streaming state cannot stand in for it either.

Binding deliberately does not reconstruct cursor's workspace-slug directory
name. That slug collapses path separators, so rebuilding it would be a guess
that could bind the wrong pane; cursor records the exact absolute workspace path
in each project's .workspace-trusted, and the binding matches on that. A
conversation recorded as prior at spawn is excluded, so a relaunch in a reused
worktree folds its own turn rather than its predecessor's. Requiring a unique
remaining conversation keeps zero and several both unknown, because neither
proves anything about the current turn.

The regression pins the fold with real transcript files and asserts the
dangerous direction stays closed: an unresolvable binding, a record-free file,
an unclaimed workspace, and a workspace-path PREFIX all read unknown, never
idle. The prefix case uses an opaque fixture slug so a slug-rebuilding
implementation cannot pass it.

Co-authored-by: Ville Penttinen <villem.penttinen@gmail.com>

* feat(cursor): make the cursor launch runnable and give it lifecycle control

Five gaps that together kept a cursor crewmate from being drivable end to end.

Launch. The template invoked `cursor agent`, but `cursor` is not the CLI - the
installed names are `cursor-agent` and the legacy alias `agent` - so the command
could not run at all on a machine with a normal cursor install. It now resolves
through the verified owner, which also refuses a spawn loudly instead of leaving
a pane that dies with command-not-found and reads as a wedged worker.

Session binding. fm-spawn writes state/<id>.cursor-session so the busy fold can
find this pane's transcript, and teardown removes it.

Lifecycle control. No cursor PR touched fm-control-lib.sh, so
`fm-control <id> interrupt|exit|relaunch` could not drive a cursor worker at
all. Verified live: interrupt is a single Escape, exit is /exit, and cursor does
NOT repollute its composer with the cancelled prompt, so unlike muse it needs no
clear key. Secondmate is refused, matching the spawn refusal.

Submit acknowledgement. cursor parks its terminal cursor outside its composer,
so the composer verdict on tmux is always `unknown` and a submit could never be
acknowledged from the composer alone. The submit core's existing idle-to-busy
transition covers that case, but only if the pane's busy footer is recognised,
so cursor's `ctrl+c to stop` joins the harness-less default union the submit
cores read. The TOKEN is matched rather than the spinner verb: the same version
rendered both `Working` and `Running` in consecutive turns.

Bootstrap. A configured cursor crew harness with no cursor executable is now a
loud MISSING diagnostic rather than a first-spawn failure, and it accepts either
installed name.

Interrupt cancellation is deliberately left unconfirmed. The transcript does
type an aborted close, but its post-interrupt write latency measured as
variable - sometimes seconds, sometimes not within twenty - so a claim built on
it would be unreliable. Normal turn completion is prompt, which is what the busy
fold actually depends on.

Two inherited tests are corrected rather than deleted: the busy test asserted
cursor could have no semantic source, and the launch test pinned the literal
`cursor agent` string. Both now pin the verified behaviour, including that the
launch never allocates a second worktree.

Co-authored-by: ABHISHAKE KUMAR BOJJA <abojja@uvic.ca>
Co-authored-by: Ville Penttinen <villem.penttinen@gmail.com>

* docs(cursor): record the verified crewmate facts and extend the drift guard

The inherited cursor entry was written against 2026.08.04-aaa8809 and several of
its claims no longer hold: it named `cursor agent` as the binary (not the CLI
name), listed six Grok model ids of which the live catalog now returns two, and
recorded busy state, exit, interrupt, and skill invocation as unverified.

Replaced with what was measured against 2026.08.11-e8db854, including the two
facts most likely to be rediscovered painfully: cursor runs as a bundled node
script so its pane title is a bare `node`, and it parks its terminal cursor
outside its composer, which makes the tmux composer verdict permanently
`unknown` by design rather than a defect to chase.

Model ids now route to `--list-models` for the account instead of a fixed list,
since that list is exactly what drifted.

The live drift guard covers cursor, resolving it through the same verified owner
fm-spawn uses and passing --trust so the probe cannot hang on the workspace
prompt. Run against every installed harness: 8 checked, all alive, with cursor
reporting title='node' foreground=[.../cursor-agent] - the drift shape this
guard exists to catch.

Co-authored-by: Ville Penttinen <villem.penttinen@gmail.com>

* docs(agents): record the cursor session-binding state file

The state/ layout section is the inventory every session reads; a busy-source
binding that fm-spawn writes and teardown removes belongs in it alongside muse's.

* no-mistakes(review): Sanitize ambient Cursor marker in harness tests

* no-mistakes(review): Validate Cursor models against live catalog

* no-mistakes(review): Reject unsupported secondmates before binary preflight

* no-mistakes(review): Narrow Cursor ancestry detection to structured process identity

* no-mistakes(review): Parse Cursor transcripts and sanitize inherited markers

* no-mistakes(review): Handle malformed Cursor transcript records safely

* no-mistakes(review): Validate malformed Cursor closes in fallback parser

* no-mistakes(review): Retire stale Cursor bindings during relaunch

* no-mistakes(review): Fix Cursor drift guard command variable

* no-mistakes(review): Narrow Cursor identity to versioned install trees

* no-mistakes(document): Document Cursor harness boundaries

* refactor(composer): move the delivery busy footers to the shared owner

The per-harness rendered busy footers lived in bin/fm-tmux-lib.sh under
FM_TMUX_* names, so cursor's `ctrl+c to stop` signature - and every other
harness's - was reachable only from tmux. That placement was wrong on its own
terms: herdr, zellij, cmux, and orca run the same harnesses and face the same
question these footers answer, which is whether a submitted Enter actually
landed. Nothing about the signature is tmux-specific.

Moved verbatim into bin/fm-composer-lib.sh, the shared composer/delivery owner
every backend already sources, and renamed to FM_DELIVERY_* so the names stop
claiming a scope they never had. All five adapters now reach cursor's signature;
verified per adapter rather than assumed.

The boundary the move must not blur is stated where it now lives: this is a
DELIVERY guard, never a worker-state source. Confirming a keystroke landed is a
different question from asking what a worker is doing, and bin/fm-busy-lib.sh
remains the semantic owner that forbids classifying a harness from rendered
text. Cursor still classifies only from its transcript fold, which is already
backend-agnostic because it folds a file rather than reading a pane - the same
verdict on all six backends.

The old FM_TMUX_* aliases are dropped rather than kept as dead shims: nothing
outside the moved block referenced them except fm-busy-lib.sh's grok fallback,
which now reads the new name. The documented operator override, FM_BUSY_REGEX,
is untouched.

Also removes a dead duplicate CURSOR_INVOKED_AS check in bin/fm-harness.sh,
unreachable behind the marker check above it.

* no-mistakes(review): Correct shared delivery guard ownership references

* no-mistakes(document): Document shared delivery guards and Cursor backend limits

* no-mistakes: apply CI fixes

* fix(composer): bound a bare composer's wrap region at a half-block rule

A live cursor crewmate on herdr classified its IDLE composer as `pending`, and
fm-send consequently exited 1 with "delivery unconfirmed" on a message that had
actually landed. The cause is not cursor-specific.

Herdr draws a composer's top and bottom rules with the half-block glyphs U+2584
and U+2580 rather than the box-drawing family. fm_composer_row_has_edge knew
only the box-drawing set, so no box was detected; the composer was found as a
BARE row, and its wrap region - which extends while rows are non-blank and carry
no structural edge - walked straight through the composer's own closing rule and
swallowed the model and path footer below it. That footer is real text, so the
region classified pending on a genuinely idle pane.

Teaching the shared edge detector the half-block glyphs bounds the region at the
closing rule. Measured on the captured bytes of a real herdr cursor pane: the
same capture that read `pending` now reads `empty`.

This is a shared shape-path change, so it is deliberately narrow - it adds
glyphs to the edge vocabulary and changes no verdict logic - and the whole
composer and backend suite is green, including the other harnesses' herdr
fixtures.

The regression pins the real captured shape and asserts the footer content is
genuinely present, so the case cannot pass vacuously if the region were ever
bounded for some unrelated reason.

* fix(herdr): confirm a cursor submit from the rendered-footer transition

Herdr's composer-shape fix made an idle cursor pane classify `empty`, but
`fm-send` still exited 1 with "delivery unconfirmed" on messages that had
actually landed. Live measurement found the second, independent cause.

Herdr reports a cursor pane `agent_status=blocked` in EVERY state - idle,
mid-turn, and after - so the submit path's idle-baseline native confirmation is
structurally unreachable for cursor and every send falls into the composer
branch. That branch reads cursor's mid-turn composer row, which renders its own
`Add a follow-up` placeholder beside a right-aligned `ctrl+c to stop`. That
token is composer content, so the verdict is `pending` on a composer holding no
user text at all, and the Enter-retry budget then reports pending.

The escape is the same semantic signal the native path uses, read from the
pane's verified busy footer instead of native agent-state, and it is the
rendered-footer twin of the tmux submit core's turn-started confirmation: an
idle-to-busy transition ACROSS our Enter proves the harness accepted the
submission. The baseline is taken before the first Enter and only when the
native baseline was not legibly idle, so the idle-baseline path still never
reads pane content and a pane already mid-turn before we typed keeps reporting
`pending` rather than borrowing another turn as proof of this delivery.

The composer verdict is deliberately NOT relaxed. A right-aligned status token
on the composer row stays content for every other caller, including the
away-mode pre-injection guard, and the shared cursorless submit core is left
untouched so zellij, cmux, and Orca keep the behavior their own follow-up owns.

Verified live on herdr 0.8.0 and cursor-agent 2026.08.11-e8db854 in an isolated
lab session: `fm-send` now exits 0 and the steer executes, interrupt cancels a
running turn, `/exit` stops the agent, and teardown clears the record. All seven
panes of the running default session classify identically before and after the
shape fix, so no other harness regressed.

* no-mistakes(review): Prevent working Herdr baselines from falsely confirming delivery

* no-mistakes(document): Correct Cursor harness and backend documentation

---------

Co-authored-by: ABHISHAKE KUMAR BOJJA <abojja@uvic.ca>
Co-authored-by: Amplify Logic AI <lars@sockinator.co>
Co-authored-by: Ville Penttinen <villem.penttinen@gmail.com>
* fix: raise quota-axi floor to 0.1.25 for Cursor CLI quota awareness

Homes on latest main need quota-axi #87 so Desktop-absent CLI machines report a fresh Cursor quota instead of a false sign-in-required.

* no-mistakes(document): Update quota floor documentation pointer
…id#2304)

* fix(guard): stop the false send-time watcher-down alarm on Pi primaries

On a Pi primary the watcher process is not the liveness signal. The Pi
extension tears the watcher down on every actionable wake and spawns the
replacement itself, so the singleton lock is legitimately unheld between
cycles: every one of the 799 cycles in a live primary's ledger ends with
lock_after=pid:none, and a live capture caught the guard verdict flipping to
no-watcher during one hand-off with the beacon 63s old.

bin/fm-guard.sh classified Pi as a persistent-watcher harness, which demands a
live identity-matched lock holder at all times, so any guarded command landing
in a hand-off painted the full WATCHER DOWN - SUPERVISION IS OFF banner and
told firstmate to repair a cycle the extension already owns and is restoring.

Add an extension supervision model for pi and pi-signed. A live
identity-matched watcher stays the ordinary healthy state; an unheld lock is
healthy only while the beacon is fresh within grace AND a live Pi session
provably owns continuity - both primary extensions recorded in their state
markers at their current on-disk builds by the process named in state/.lock,
with that process still alive. Without that proof the banner fires exactly as
before, so an unloaded, version-drifted, or exited Pi session is loud
immediately and a cycle the extension never restores is loud once the beacon
passes grace. The queued-wake warning, the PID-strict turn-end guard, and
every other primary's detection are untouched.

Fold session-start's duplicate Pi marker predicate into the shared library so
the ownership contract has one owner.

* no-mistakes(review): Restrict Pi hand-off tolerance to unheld watcher locks

* no-mistakes(document): Document Pi watcher hand-off supervision
* feat(cursor): add Cursor Agent CLI primary hooks, park supervision, and session start

Register a tracked project-scope .cursor/hooks.json for Cursor's stop,
sessionStart, preCompact, and preToolUse steps.

bin/fm-turnend-guard-cursor.sh owns Cursor's turn boundary as a park: it
foregrounds the watcher arm, holds the boundary open until an actionable close,
and returns that wake as one follow-up. Exit 2 is a silent no-op on Cursor's
stop step, so the adapter never uses it. The follow-up loop is bounded twice,
by Cursor's own loop_limit and by the payload's loop_count.

bin/fm-sessionstart-cursor.sh delivers the digest as additional_context at
sessionStart, and stages it for the next turn boundary at preCompact, which
cannot inject context.

Cursor also loads the tracked Claude settings, so bin/fm-hook-host-lib.sh lets
each tracked Claude-shaped entrypoint stand down on a Cursor-delivered payload
rather than running every covered event twice.

bin/fm-tmux-lib.sh reclassifies a Cursor pane's composer cursorlessly, because
Cursor parks its terminal cursor outside the composer, which restores a genuine
composer-empty proof and unblocks away-mode escalation delivery.

* feat(cursor): make Cursor Agent CLI a verified primary harness

Resolve Cursor in the session-lock ancestry through bin/fm-cursor-lib.sh, which
a Cursor primary needs before it can hold its own home lock, and classify its
stop-hook park under the autoarm supervision model so the mid-turn pull guard
stops reporting a healthy between-turns watcher as down.

Read a Cursor pane's composer cursorlessly on tmux, gated on Cursor's own
structural process identity, which restores a genuine composer-empty proof and
lets away-mode escalations reach a Cursor primary with no daemon change.

Lift the secondmate refusals in bin/fm-spawn.sh and bin/fm-control-lib.sh now
that the supervision protocol exists and is recorded.

Cover the whole surface with a portable regression over real processes, an
opt-in live guard against the installed cursor-agent, and dated per-harness
evidence.

* docs(cursor): record Cursor as a verified primary across the owning surfaces

Update the turn-end guard, session-start, arm-seatbelt, cd-guard, watcher
continuity, architecture, configuration, README, and harness-adapters owners,
and add dated live evidence to the supervision and runtime-backend verification
records. Correct the recorded Cursor tmux composer verdict: the cursor-anchored
read is still blind, but the composite reader is no longer unknown.

Lift the remaining remote-secondmate refusal missed in the previous commit, and
add the new libs to the existing fixtures that copy a fixed dependency list.

* refactor(cursor): name the park's stand-down condition for both its causes

Also record that Cursor's preCompact firing itself is not yet live-verified,
while the static evidence that it cannot inject context, and the staging path
that follows from it, both are.

* test: give the pretool fixtures their new dependency and one lint owner

The cd-guard fixture copies a fixed dependency list and now needs the shared
hook-host predicate. Both pretool suites also asserted cleanliness with a bare
shellcheck call, a second and weaker copy of the lint definition that
bin/fm-lint.sh owns: it omits --external-sources, so it failed the moment these
checkers sourced a shared library. They now delegate to that owner.

* test: assert the cursor secondmate contract instead of its removed refusal

A cursor secondmate now launches, so the suite asserts what its park actually
needs: --trust so the home's project hooks load at all, its own home pinned as
the workspace, and the autoarm supervision model inherited across the launch.

* no-mistakes(review): Serialize Cursor wakes and bind staged context

* no-mistakes(review): Serialize Cursor context and nag state commits

* no-mistakes(review): Enforce Cursor ceiling before staged context delivery

* no-mistakes(review): Serialize Cursor claims and staged context

* no-mistakes(review): Serialize Cursor ownership and state commits

* no-mistakes(review): Protect Cursor context across session takeover

* no-mistakes(review): Preserve Cursor context across session takeover

* no-mistakes(review): Enforce owner-keyed Cursor staged context

* no-mistakes(review): Atomically claim Cursor follow-ups and staged context

* no-mistakes(review): Defer Cursor preCompact staging and simplify supersession

* no-mistakes(review): Serialize Cursor park commits and defer preCompact

* no-mistakes(review): Stop Cursor parks after session takeover

* no-mistakes(test): Route Cursor preCompact context through stop follow-up

* no-mistakes(document): Update Cursor primary documentation

* revert(cursor): cut preCompact staging from this change

Carrying a compaction digest across two concurrently running stop hooks kept
producing races that could deliver it twice or strand it indefinitely, and
closing them kept enlarging a critical section inside a hook Cursor awaits at
the turn boundary. Native preCompact firing was never observed either, so the
surface has no empirical basis yet.

Remove the adapter, its registration, its staged path in the park, and its
tests, and record the surface as deferred and uncovered alongside the Codex
interactive TUI. A regression now asserts preCompact stays unregistered so it
cannot return without its own design and evidence.

This change ships the proven core only: the turn-end follow-up park, the
run-tier session start, and away-mode delivery.

* no-mistakes(review): Correct Cursor park supersession documentation

* no-mistakes(document): Clarify Cursor run-tier verification ownership

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

---------

Co-authored-by: kunchenguid <kun-1@kunchenguid.com>
…id#2330)

* feat(bin): add unrouted close paths to the captain decision gate

A captain who declines a held decision leaves no follow-up work to route,
so `resolve` could not express that answer: it requires at least one
`--routed-to` task. The only way to close such a hold was a direct
`tasks-axi done`, which never writes the durable resolution record the
completion gate reads, so the originating investigation could no longer
pass `verify` and its cleanup stayed blocked.

Add two close paths that route no work:

- `decline` closes an actively held hold with a recorded captain decision
  and no routed task. It refuses while any task is still blocked by the
  hold, because releasing routed work without recording it is `resolve`'s
  job.
- `repair` records the missing resolution block on a hold that was already
  closed outside this script. It never reopens a hold and never clears a
  dependency edge, and it refuses a hold that is still actively held.

Both require a non-empty captain decision file and share `resolve`'s
digest-based retry identity, so an exact retry is idempotent while a
changed decision is rejected. The recorded body now also names which path
closed the hold, and each routed entry regains its own line.

The gate itself is unchanged: an unanswered decision still fails
completion and blocks teardown, and neither new path can close a hold
without the captain's recorded word.

* fix(bin): require captain-hold provenance before repairing a decision

`repair` checked only that the backlog item was kind captain and Done, so
an ordinary captain-kind task that was never held for the captain could be
closed, repaired, and then pass the completion gate.

tasks-axi keeps `hold_kind` through a close, so it is the surviving proof
that an identity really was a captain hold. Require it before writing the
resolution record, and cover the case in the gate regression.

* no-mistakes(document): Correct decision-hold lifecycle documentation
* fix(bin): surface buried status notes on wake drain

A note: answer immediately followed by a routine note was dropped because
annotations kept only the newest line and note: never enters OPEN DECISIONS.
Present every unread note and pending-reply resolution since the last drain
cursor, and annotate every unread line on a queued signal.

* no-mistakes(review): Fix unread status cursor races and overflow

* no-mistakes(review): Preserve cursors when status span reads fail

* no-mistakes(review): Make status presentation transactional under I/O failures

* no-mistakes(review): Simplify unread status cursor and presentation locking

* no-mistakes(review): Align cursor failure regressions with transactional presentation

* no-mistakes(review): Retire stale presentation cursors during task teardown

* no-mistakes(review): Preserve routine status until signal annotation

* no-mistakes(review): Correct unread status cap documentation

* no-mistakes(document): Document unread wake status presentation

* no-mistakes(lint): Fix wake surfacing ShellCheck warnings

* no-mistakes: apply CI fixes
* feat(calm): add a max presentation level that hides mid-turn working notes

Calm's home-local preference becomes a three-state level instead of a
boolean: "off" is stock Pi, "on" is today's Calm, and "max" is Calm plus
hiding the assistant text of messages the model did not end its response
with. `/calm max` selects it from any state, a plain `/calm` steps max
back to ordinary Calm and otherwise keeps the existing on/off cycle, and
any other argument keeps that cycle too.

`config/calm` now persists "max" as its own literal value, so a session
start, resume, fork, or reload restores the stored level rather than
treating it as unrecognized and dropping to off.

The hide rule keys on Pi's intrinsic per-message stopReason: "toolUse",
or "length" with tool calls present. Streaming ("pending") text is never
filtered, because suppressing it would also stop a genuine reply from
streaming. The existing assistant layout adapter filters the blocks out
of the same shallow presentation copy it already uses for collapsed
thinking, so the message, model context, session storage, /export, and
delivery are untouched and a hidden mid-turn row collapses to zero
height. The new "assistant-working-note" class keeps that choice in the
visibility policy owner, where ordinary Calm keeps it visible.

* no-mistakes(document): Clarify Calm max persistence and taxonomy
* feat(calm): make hiding mid-turn working notes the ordinary Calm state

Calm collapses back to the two-state on/off toggle it was before the max
presentation level, with max's hide rule promoted into ordinary Calm.
Calm on now hides mid-turn assistant working notes in addition to what it
already hid, and the /calm command parses no argument again.

The hide rule itself is unchanged: assistant text is removed from the
shallow presentation copy when the message's own stopReason is "toolUse",
or "length" with tool calls present. Streaming ("pending") text is never
filtered, so a genuine reply still streams. The message, model context,
session storage, /export, and delivery remain untouched.

config/calm persists only "on" and "off" again, but the reader still maps
a persisted "max" to on so a home upgraded from the removed level keeps
Calm on instead of dropping to off.

The mid-turn hide is now default behavior rather than an opt-in level, so
docs/calm.md documents it for users, docs/configuration.md records the
two written values plus the legacy max mapping, and the feasibility
taxonomy drops its level-scoped wording.

* no-mistakes(document): Document ordinary Calm working-note hiding
A wedged family-run step was occupying the runner until the 75-minute
job cap; bound that step so cleanup and timing artifacts still upload.
@DereKk8
DereKk8 merged commit 82b517b into main Aug 15, 2026
11 of 16 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.