Skip to content

feat: add new crew adapters, Pi branch supervision, and durable captain operations - #33

Merged
yelenplays merged 252 commits into
mainfrom
fm/upstream-reconcile-v1
Sep 18, 2026
Merged

yelenplays merged 252 commits into
mainfrom
fm/upstream-reconcile-v1

Conversation

@yelenplays

@yelenplays yelenplays commented Sep 18, 2026 •

Copy link
Copy Markdown
Owner

What Changed

  • Add Oh My Pi as a verified primary plus Gemini, Rovo, and Antigravity crew adapters, flag-gated Claude Calm, and a default-on Pi supervision branch with /supervision-model model and effort pins.
  • Add a durable AFK posture, /quiet attended supervision, IMAP/SMTP mail, a spoken voice relay, and durable local/remote steering inboxes.
  • Route GitLab merges through the guarded PR path, add live-head merge gates, restart second mates after updates, and introduce trusted process-event extensions plus opt-in typed dispatch resolution.

Risk Assessment

⚠️ Medium: The reconcile itself is a coherent fork-on-upstream restoration, but the restored spawn {TASK} guard refuses already-supported filled briefs that mention the placeholder token.

Testing

Targeted fm-test-run of the fork-local restore suites all passed: task-delivery, live-gate (all 35 guards refuse together on FM_LIVE=0), fixtures, crew-state idle-working, and Pi watch recovery/extension. Devin and Pi-role live e2e correctly skipped as opt-in. A manual spawn against an isolated fake-tmux home showed an unedited fm-brief scaffold refused for leftover placeholders, a legacy # Task body of only {TASK} refused by the restored spawn-side guard with no task metadata, and a filled brief proceeding past those checks until the fake tmux backstop. FM_LIVE=0 on the re-sanitized default-on live e2e scripts emitted the shared skip. No product failures.

Evidence: Operator-facing spawn {TASK} guard transcript

Source: Operator-facing spawn {TASK} guard transcript

unedited scaffold: error: ... still contains {TASK} or {FIRSTMATE_SPEC}; fill ## Captain's intent and ## Firstmate spec before spawn (exit=1, no metadata) legacy # Task body only {TASK}: error: brief at .../launch-brief.md still carries the literal {TASK} placeholder; write the task description before spawning (exit=1, no metadata) filled subsections: standing-posture notice, then fake tmux refusal (placeholder checks passed)

=== firstmate spawn {TASK} guard (operator-facing) ===
worktree: ~/.no-mistakes/worktrees/548b8aa73bec/01M2TMGAY38PBK5W100EM4S15Y
head: 56862de chore: rebind no-mistakes attestation to a fresh head
fixture home: /tmp/fm-spawn-guard-demo.ylASxl/home

--- 1. unedited ship scaffold (fm-brief.sh default {TASK}/{FIRSTMATE_SPEC}) ---
scaffolded: /tmp/fm-spawn-guard-demo.ylASxl/home/data/unfilled-ship/brief.md (ship, mode=no-mistakes; replace {TASK} and {FIRSTMATE_SPEC})
Captain's intent / Firstmate spec bodies:
## Captain's intent
{TASK}

## Firstmate spec
## Firstmate spec
{FIRSTMATE_SPEC}

# Herdr lifecycle declaration - NOT ENABLED
spawn:
error: /tmp/fm-spawn-guard-demo.ylASxl/home/data/unfilled-ship/brief.md still contains {TASK} or {FIRSTMATE_SPEC}; fill ## Captain's intent and ## Firstmate spec before spawn
exit=1
task metadata present? no

--- 2. legacy # Task body that is only {TASK} (no subsections) ---
Task section:
# Task
{TASK}

# Definition of done
spawn:
error: brief at /tmp/fm-spawn-guard-demo.ylASxl/home/data/legacy-task-only/launch-brief.md still carries the literal {TASK} placeholder; write the task description before spawning
exit=1
task metadata present? no

--- 3. filled subsections, no leftover placeholder ---
Captain's intent / Firstmate spec bodies:
## Captain's intent
Ship the idle-working status disclosure.

## Firstmate spec
## Firstmate spec
Keep historical working events unknown for idle ships.

# Herdr lifecycle declaration - NOT ENABLED
spawn:
notice: filled-ship ships mode=direct-PR while the standing posture for proj is no-mistakes - less rigor than the captain's standing posture; proceed only on a current explicit captain instruction or an intake judgment you can state
tmux should not have been invoked
tmux should not have been invoked
tmux should not have been invoked
exit=1
Evidence: Live-gate skip transcript (opt-in unset and FM_LIVE=0)

Source: Live-gate skip transcript (opt-in unset and FM_LIVE=0)

tests/fm-devin-live-e2e.test.sh: skip: live: opt-in; set FM_DEVIN_LIVE=1 to run (exit=0) tests/fm-pi-role-agents-live-e2e.test.sh: skip: live: opt-in; set FM_PI_ROLE_AGENTS_LIVE=1 to run (exit=0) tests/fm-harness-liveness-drift-live-e2e.test.sh: skip: live: disabled by FM_LIVE=0 (exit=0) tests/fm-herdr-version-floor-live-e2e.test.sh: skip: live: disabled by FM_LIVE=0 (exit=0)

=== live-gate skip when FM_LIVE=0 / opt-in unset (operator-facing) ===
head: 56862de

--- tests/fm-devin-live-e2e.test.sh ---
skip: live: opt-in; set FM_DEVIN_LIVE=1 to run
exit=0

--- tests/fm-pi-role-agents-live-e2e.test.sh ---
skip: live: opt-in; set FM_PI_ROLE_AGENTS_LIVE=1 to run
exit=0

--- FM_LIVE=0 on default-on live e2e (sanitizer + shared gate) ---
tests/fm-harness-liveness-drift-live-e2e.test.sh:
skip: live: disabled by FM_LIVE=0
exit=0
tests/fm-herdr-version-floor-live-e2e.test.sh:
skip: live: disabled by FM_LIVE=0
exit=0
Evidence: Targeted fm-test-run timing artifact

Source: Targeted fm-test-run timing artifact

{
  "families": [
    {
      "count": 2,
      "duration_ms": 517,
      "failed": 0,
      "name": "live-harness-optin"
    },
    {
      "count": 2,
      "duration_ms": 152004,
      "failed": 0,
      "name": "pure-contract-unit"
    },
    {
      "count": 2,
      "duration_ms": 18044,
      "failed": 0,
      "name": "standalone"
    },
    {
      "count": 2,
      "duration_ms": 165788,
      "failed": 0,
      "name": "watcher-wake-lock"
    }
  ],
  "finished_at": "2026-09-18T16:25:30Z",
  "run_id": "fm-test-run-1789748468931-74530",
  "scripts": [
    {
      "duration_ms": 113948,
      "exit": 0,
      "expected_gate_skip": "none",
      "family": "pure-contract-unit",
      "gate_skip": false,
      "gate_skip_reason": "",
      "path": "tests/fm-crew-state.test.sh"
    },
    {
      "duration_ms": 68794,
      "exit": 0,
      "expected_gate_skip": "none",
      "family": "watcher-wake-lock",
      "gate_skip": false,
      "gate_skip_reason": "",
      "path": "tests/fm-pi-watch-extension.test.sh"
    },
    {
      "duration_ms": 96994,
      "exit": 0,
      "expected_gate_skip": "none",
      "family": "watcher-wake-lock",
      "gate_skip": false,
      "gate_skip_reason": "",
      "path": "tests/fm-watch-recovery-loop.test.sh"
    },
    {
      "duration_ms": 38056,
      "exit": 0,
      "expected_gate_skip": "none",
      "family": "pure-contract-unit",
      "gate_skip": false,
      "gate_skip_reason": "",
      "path": "tests/fm-task-delivery.test.sh"
    },
    {
      "duration_ms": 7169,
      "exit": 0,
      "expected_gate_skip": "none",
      "family": "standalone",
      "gate_skip": false,
      "gate_skip_reason": "",
      "path": "tests/fm-live-gate.test.sh"
    },
    {
      "duration_ms": 10875,
      "exit": 0,
      "expected_gate_skip": "none",
      "family": "standalone",
      "gate_skip": false,
      "gate_skip_reason": "",
      "path": "tests/fm-test-fixtures.test.sh"
    },
    {
      "duration_ms": 270,
      "exit": 0,
      "expected_gate_skip": "live-capability",
      "family": "live-harness-optin",
      "gate_skip": true,
      "gate_skip_reason": "live: opt-in; set FM_DEVIN_LIVE=1 to run",
      "path": "tests/fm-devin-live-e2e.test.sh"
    },
    {
      "duration_ms": 247,
      "exit": 0,
      "expected_gate_skip": "live-capability",
      "family": "live-harness-optin",
      "gate_skip": true,
      "gate_skip_reason": "live: opt-in; set FM_PI_ROLE_AGENTS_LIVE=1 to run",
      "path": "tests/fm-pi-role-agents-live-e2e.test.sh"
    }
  ],
  "selection": "scripts;jobs=4",
  "started_at": "2026-09-18T16:21:08Z",
  "summary": {
    "duration_ms": 261230,
    "failed": 0,
    "skipped_gate": 2,
    "total": 8
  }
}
Evidence: Targeted fm-test-run summary

Source: Targeted fm-test-run summary

total=8 failed=0 skipped_gate=2 duration_ms=261230

=== targeted fm-test-run summary ===
total=8 failed=0 skipped_gate=2 duration_ms=261230

tests/fm-crew-state.test.sh exit=0 duration_ms=113948 gate_skip=False reason=
tests/fm-pi-watch-extension.test.sh exit=0 duration_ms=68794 gate_skip=False reason=
tests/fm-watch-recovery-loop.test.sh exit=0 duration_ms=96994 gate_skip=False reason=
tests/fm-task-delivery.test.sh exit=0 duration_ms=38056 gate_skip=False reason=
tests/fm-live-gate.test.sh exit=0 duration_ms=7169 gate_skip=False reason=
tests/fm-test-fixtures.test.sh exit=0 duration_ms=10875 gate_skip=False reason=
tests/fm-devin-live-e2e.test.sh exit=0 duration_ms=270 gate_skip=True reason=live: opt-in; set FM_DEVIN_LIVE=1 to run
tests/fm-pi-role-agents-live-e2e.test.sh exit=0 duration_ms=247 gate_skip=True reason=live: opt-in; set FM_PI_ROLE_AGENTS_LIVE=1 to run

Pipeline

Updates from git push no-mistakes

⏭️ **intent** - skipped

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 1 error
  • 🚨 bin/fm-spawn.sh:2658 - The restored unedited-scaffold abort is a substring grep of the # Task range after launch-brief overlay, not an exact leftover-placeholder check. Concrete path: a ship brief whose ## Captain's intent is filled with Fix replacement of {TASK} in Herdr briefs. and whose ## Firstmate spec is filled (the case tests/fm-task-delivery.test.sh already treats as valid) passes fm_brief_task_placeholders_present (exact-body only) and fm_brief_task_content_valid, then this grep still matches {TASK} and spawn exits 1 with still carries the literal {TASK} placeholder. The same happens for any filled # Task that mentions the token. The claimed invariant is unedited scaffold must never launch; a leftover legacy body that is only {TASK} is still a real gap because fm_brief_task_placeholders_present currently returns false for briefs without the two subsections (bin/fm-dod-lib.sh:79) while content_valid accepts nonempty {TASK}. Close that at fm_brief_task_placeholders_present with the same exact-body rule already used for ## Captain's intent / ## Firstmate spec, and drop this spawn-side substring grep rather than layering another symptom patch.
✅ **Test** - passed

✅ No issues found.

  • bin/fm-test-run.sh --json ~/.no-mistakes/evidence/01M2TMGAY38PBK5W100EM4S15Y/fm-test-run.json tests/fm-task-delivery.test.sh tests/fm-live-gate.test.sh tests/fm-devin-live-e2e.test.sh tests/fm-pi-role-agents-live-e2e.test.sh tests/fm-test-fixtures.test.sh tests/fm-watch-recovery-loop.test.sh tests/fm-pi-watch-extension.test.sh tests/fm-crew-state.test.sh
  • manual fm-brief.sh + fm-spawn.sh on an unedited ship scaffold with leftover {TASK} / {FIRSTMATE_SPEC}
  • manual fm-spawn.sh on a legacy # Task body that is only {TASK}
  • manual fm-spawn.sh on filled subsections with no leftover placeholder (fake tmux backstop)
  • env -u FM_DEVIN_LIVE -u FM_LIVE bash tests/fm-devin-live-e2e.test.sh
  • env -u FM_PI_ROLE_AGENTS_LIVE -u FM_LIVE bash tests/fm-pi-role-agents-live-e2e.test.sh
  • FM_LIVE=0 bash tests/fm-harness-liveness-drift-live-e2e.test.sh
  • FM_LIVE=0 bash tests/fm-herdr-version-floor-live-e2e.test.sh
✅ **Document** - passed

✅ No issues found.

⚠️ **Lint** - 1 warning
  • ⚠️ linter found issues (exit code 1)
✅ **Push** - passed

✅ No issues found.

Inthuson and others added 30 commits August 21, 2026 15:55
…lled but inert (kunchenguid#2684)

* feat(checks): report tool updates that are available or installed but inert

Firstmate had no way to notice that tooling this home depends on needs an
update, and no way at all to notice the worse case: an update that installed
correctly and then did nothing.

That second case is why this exists. A tool that self-installs into
~/.local/bin while a version manager keeps its own older copy earlier on PATH
looks completely up to date to anything that asks only "is a newer version
published". On 2026-08-20 a Herdr update landed at 0.8.2 while an older 0.8.0
copy stayed earlier on PATH, so every Herdr command failed on a protocol
mismatch and firstmate could not read its own fleet.

bin/fm-tool-update-check.sh reports the two conditions separately:

  <tool> update available      a newer version exists at the update source.
  <tool> update not in effect  a newer copy is installed on this host, but
                               PATH still resolves an older one.

PATH skew is measured, never inferred. Every executable copy of a watched
command on PATH is asked for its own version and those answers are compared,
so one lookup cannot hide the skew, and a directory name is never read as a
version because a version manager's "latest" directory can hold an older
build. A copy that will not report a version is a check failure, not a pass.

The watched tools live in local, gitignored config/watched-tools.json, so
adding a tool is a config edit rather than a code change, and the file is
never propagated to another home. Update sources cover both shapes: a local
clone's commit distance from its remote branch, and a command's own version
and update announcement, including a tool like no-mistakes that prints its
version on one command and announces a new release on another.

The check prints one line when something needs attention and prints nothing
otherwise, so it rides the existing watcher state-check contract with its
trust binding instead of introducing a schedule of its own, and
state/.tool-updates keeps the same pending update from being reported on
every poll.

The check only reports. It never installs, updates, reorders PATH, touches a
version manager, or fetches into a watched repository; every git probe is
read-only.

Tests cover the skew case as a regression, and it was verified by mutation:
removing the skew report, or stopping after the first PATH hit as a single
lookup would, each make that test fail.

* no-mistakes(review): fix tool update check probe reporting, budget, and shim write

* no-mistakes(review): keep sweeps alive on broken patterns and oversized budgets

* no-mistakes(review): roll back failed arm, widen budget clamp, bound repo probe

* no-mistakes(review): guard git probes at the budget, record uncut findings

* no-mistakes(document): fix stale watched-tool report-record wording in docs and header

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

The behavior shard's watch-triage suite failed on the new worktree-write wedge
tests. Those five tests are the only ones in the file that do not use its
standard waits. They give a fixed 3 second liveness budget to the one poll that
now spawns the bounded worktree walk, and 4 seconds to an escalating watcher
where every other test in the file gives 10. On a loaded runner that poll
outlives the fixed budget, so the round is reaped before the deferral it asserts
on is recorded, and the test reports a lost deferral instead of the deferral
under test. Wait for a completed poll cycle through the file's own
wait_poll_cycle, which is what its header documents this hazard for, and use the
file's standard 100 tick exit budget.

Verified against a load that reproduces the failure: 11 of 12 runs failed
before, 8 of 8 pass after. Verified by mutation too, so the waits still prove
the behavior: removing the write deferral, and keeping a finished deferral chain
across an idle-timer repair, each still fail their test.
* fix: treat yolo as merge authority only, not ask-user finding authority

Yolo on/off was documented as also deciding no-mistakes ask-user findings, which hid firstmate's duty to judge unambiguous-toward-design findings itself. Keep every safety boundary; this is a contract clarification, not a relaxation.

* no-mistakes(document): Clarify yolo documentation ownership and merge posture
… work over (kunchenguid#2767)

* feat(voice): spoken round trip on Nova Sonic 2 with a measured relay cost

Step one of the spoken interface: the laptop captures and plays audio, this
desktop holds the model session, and no AWS credential leaves the desktop.

Measured, amazon.nova-2-sonic-v1:0 in eu-north-1, end of speech to first byte
of reply audio, 6 runs each, all answered, on a question that forces a records
read:

  relay path   1.229 1.379 1.428 1.447 1.481 1.516  median 1.438
  direct       1.147 1.179 1.203 1.237 1.244 1.317  median 1.220

The relay costs about 0.22s of the median. The direct figure reproduces the
earlier survey, which is what makes it a usable control. Excluded: the
captain's own ssh round trip, microphone capture, and speaker output. This
desktop has no microphone and no speaker, so every run used audio files.

Three pieces:

  bin/fm-voice-relay.py    holds the conversation on this host
  bin/fm_voice_records.py  what a spoken answer may read, and the handover
  bin/fm-voice-client.py   the laptop end; audio devices UNVERIFIED
  bin/fm_voice_frame.py    the wire format both machines share

Real work is handed to the existing bin/fm-inbox.sh rather than a second
queueing surface, and the agent says it is handing over rather than answering
as firstmate.

Read scope: Done history and free-form note bodies are never assembled at any
scope, so the wide default cannot reach the places commercial detail
accumulates. config/voice-read-scope narrows it to counts only, and
config/voice-read-deny excludes a named item in one line. The boundary is an
executable test that widening the reader fails.

Push to talk is the default because it is cheaper and the choice is still open;
--listen open-mic is the single flip.

Two traps worth knowing: a clip with no trailing silence is never answered, and
the end of a reply is contentEnd with stopReason END_TURN, not completionEnd.
A second user turn in one session is treated as barge-in unconditionally, and
an interrupted turn that calls a tool is lost, so the session reconnects per
turn and gives up conversational memory. That is the concrete thing step three
has to solve.

* no-mistakes(review): fix voice relay credential reuse, frame validation and record parsing

* no-mistakes(review): test uplink header guard, bound unknown expiry, align state dir

* no-mistakes(review): decide deny per item, guard turn failures, bound ambient credentials

* no-mistakes(review): read account config from home, harden deny and turn failures

* no-mistakes(review): close status verb set, fix inbox help, pair data override

* no-mistakes(review): keep profile-free relay alive, unblock loop, fix dead assertion

* no-mistakes(review): hide finished pull requests, refuse open mic, keep suite offline

* no-mistakes(review): survive reader failures, release devices, fix claims

A failure while handling a model event, or while sending a tool result,
left the reader task dead with ended and turn_done clear, and close()
re-raised the stored failure on every await. One dropped stream became a
relay that could never build another session. The reader now reports the
session over in a finally whatever killed it, and close() absorbs the
task the same way it already absorbed its sends.

The laptop client releases what it already started when a later startup
step refuses, SystemExit from the handshake wait included, and names a
device refusal instead of leaking a raw PortAudio error. Whether it
releases correctly against a real device is still unverified here.

The records docstring claimed every reading was filtered to open ids.
Only the pull request count and list are; the worker count and the state
histogram cover every live runtime record, finished ids included,
because a meta file still on disk still needs tearing down.

The finished-work deny half of the suite asserted things that held with
the deny list absent. It is replaced by a deny on an open title, which
removes the row and says so while the count stays honest.

* no-mistakes(review): name reader failures, split file and device refusals

A failure inside the model reader released the waiting turn and told
nobody. The session was not marked spent, no notice reached the client,
and the client waits for a reply end or a notice, so the captain got
their whole timeout of silence and then a record saying the turn went
unanswered with nothing about why. Both ends of the relay now name a
failed turn through one function, once per turn, and --self-test carries
the cause in relay_error the way the client's own record does.

Two things that are not failures stay that way. A stream that simply
ends is the end of a session, which serve still reads on its own terms.
A stream that goes away because close() asked it to is an ordinary
renew, and announcing it would have put a failure notice in front of the
captain on every turn.

On the laptop end, the refusal that became a device error covered the
file-backed playback and capture too, so a mistyped --in-file was
reported as an audio device failure and the advice named the flag that
had just failed. The file ends now report the path and the flag that
chose it and stay an OSError; the device ends keep the device advice and
name the flag for that end. The device paths remain unrun here, so only
the file halves are covered by a test.

* no-mistakes(test): survive model session end, order client turn frames

* no-mistakes(document): sync voice relay docs with reviewed relay behavior

* no-mistakes(document): re-measure relay latency and correct its cause

* no-mistakes(document): correct measurement date and name the unmeasured SSH hop

* no-mistakes(document): describe the unpublished control measurement, fix list formatting

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes
…unchenguid#2763)

* fix: keep Relay public loops open until retire

Delivering a promised-final reply was deleting the only record that tied a public thread to later work, so a follow-on ship silently owed no closing reply. Retain the registration after delivery, rechain follow-on work onto the same thread, and make retire --reason the only close.

* no-mistakes(review): Propagate public follow-up registration removal failures

* no-mistakes(review): Persist retire receipts and align parent resolution

* no-mistakes(review): Make rechain resumable after partial obligation creation

* no-mistakes(review): Repair follow-up state, briefs, and expiry escalation

* no-mistakes(review): Serialize follow-up delivery stamps with retirement

* no-mistakes(review): Serialize rechain claims and protect registration terminal states

* no-mistakes(review): Avoid reporting retired delivery loops as open

* no-mistakes(document): Refresh public-loop documentation and verification evidence

* no-mistakes: apply CI fixes

* no-mistakes(review): Preserve delivered follow-up bindings during registration replay

* no-mistakes(review): Harden public follow-up retirement and rechain races

* no-mistakes(review): Fail closed on unresolved secondmate retirement

* no-mistakes(review): Bind secondmate cleanup to its recorded canonical home

* no-mistakes(review): Fix rechain command output and expiry validation

* no-mistakes(review): Validate brief keys and warn on remote promotion

* no-mistakes(document): Document retained public follow-up loops

* no-mistakes(lint): Remove unused bounded-wait loop variable
…ath (kunchenguid#2779)

* feat(bin): merge GitLab merge requests through the guarded PR merge path

bin/fm-pr-lib.sh already parses a GitLab merge request URL for the watcher,
but bin/fm-pr-merge.sh refused every non-github provider, so a merge request
had to be merged by hand and got none of the recording, guards, or audit
trail a pull request gets.

The merge path now dispatches on the parsed provider. A GitHub URL keeps its
exact previous behavior. A GitLab URL is addressed through glab by the project
URL rebuilt from the parsed host and path, so a merge request on any instance
resolves and no host is hardcoded, and no merge-method flag is added because
the project's own merge method is what should apply.

A GitLab merge happens only after one live read of the merge request confirms
it is open, detailed_merge_status is mergeable, has_conflicts is false,
blocking_discussions_resolved is true, and the head pipeline succeeded at the
exact current head. Every failing condition is reported, not just the first.
The verified head is bound to the merge with glab's --sha, so a push landing
between the read and the merge fails the merge instead of landing commits
nothing verified. Recorded metadata is never the authority for any of this: a
rebase moves the head and leaves a recorded value stale, so a recorded head
that disagrees with the live one is reported rather than trusted, and the
recorded value is read before the recording step because that step drops a
GitLab head it cannot resolve.

* no-mistakes(review): reject bundled -R clusters and make tool-absence cases host-independent

* no-mistakes(test): state authorised GitHub narrowing of bundled -R guard

This branch NARROWS GitHub behaviour. The narrowing was authorised
deliberately rather than slipping in by accident, and it applies to both
providers, GitHub and GitLab alike, because a script that guards one provider
and not the other is a trap for the next reader.

What bin/fm-pr-merge.sh now refuses is extra merge arguments containing a
bundled short-option cluster that includes R, for example "-dR other/repo".
The forge CLIs expand such a cluster one character at a time, so it carries
"--repo other/repo", and that later value wins over the repository the URL
named. Before this change, "fm-pr-merge.sh <task> <github-url> -- -dR
other/repo" reached "gh-axi pr merge 12 --repo example/repo --squash -dR
other/repo" and exited 0 with pr= recorded and the merge poll armed. It now
exits 1 with "extra merge arguments must not override the repository", records
nothing, and invokes no forge merge command. Every other GitHub invocation is
byte-identical to the base commit.

Closing that hole honours the existing rule rather than departing from it. The
file header already forbids --repo and -R because the repository must come
only from the URL, so a bundled cluster carrying a repository override was
never legitimate behaviour to preserve: it was that guard being evaded.
Redirecting a merge to a repository the URL does not name is exactly what the
guard exists to prevent.

The refusal is already pinned on both paths by the existing case
test_bundled_repo_override_args_refuse_before_recording in
tests/fm-pr-merge.test.sh. On GitHub ("-dR wrong/repo") and on GitLab ("-yR
https://other.example/g/p") it asserts exit 1, the refusal wording, no pr= in
the task meta, no armed merge poll, and no forge merge command invoked, with a
control case proving a cluster that carries no repository override still
reaches the forge. No duplicate assertion was added. Both assertions were
confirmed to have teeth by narrowing the guard back to a bare -R and watching
each path fail.

This commit carries no file change: the guard and its coverage landed in
614853d, and this message exists so the pull request description states the
narrowing.

* no-mistakes(document): fix README pointer for GitLab watch and merge doc

* no-mistakes: apply CI fixes
kunchenguid#2788)

* no-mistakes: apply CI fixes

* fix(bin): drop a private record citation and narrow the review rule

Three corrections to the spoken interface that landed in kunchenguid#2767, plus one
fix carried over from that branch after its pull request had already been
merged.

The confidentiality fix. The module docstring of bin/fm-voice-relay.py
cited a private, gitignored fleet record by exact path and section number.
That widens what this public repository points at, and it cannot resolve
for any reader here, because the path has never been in the repository.
Both traps it pointed at are already described in full in the list
immediately below it, and docs/voice-relay.md carries the same two for
operators with no citation at all, so the pointer is removed and no claim
is weakened by losing it. Two comments that referred to "the survey" as
though it were something a reader could open are reworded the same way.
Neither exposed a path, so that half is comprehensibility rather than
confidentiality.

The review rule. .greptile/rules.md is kept, because its conditions are
right and deleting it would leave the next reviewer to re-litigate a
decision already argued out. What was wrong with it is narrower than its
existence: it read as settled repository policy, when whether VISION.md
itself should be reconciled is an open question belonging to the captain.
One sentence now says so, and says that the conditions listed below it are
what the interpretation depends on. That narrows the claim rather than
widening it.

The carried-over fix. The first commit on this branch is 7f98e79 from
fm/voice-relay-build-v4, taken verbatim rather than rewritten. It closes
the window where a transport failure was recorded and then erased, so a
run could be emitted as answered false with relay_error null. That matters
more than it looks: relay_error is the field that keeps an infrastructure
failure from being averaged into a latency figure, so the failure mode is
a dead connection wearing the costume of a slow reply. It landed fifteen
minutes after kunchenguid#2767 merged and so never reached the default branch.

* no-mistakes(review): name a reason on every unanswered-turn close path

* no-mistakes(review): guard the downlink body and pin frames to their turn

* no-mistakes(review): attribute reply audio to its own turn and tell endings apart

* no-mistakes(review): tell a cut-short reply from an unanswered turn

* no-mistakes(review): discard reply audio arriving after the output closes

* no-mistakes(review): count discarded reply audio on the speaker path too

* no-mistakes(review): keep a reason off a turn already answered in full

* no-mistakes(review): say a reset cut a reply short, not that none arrived

* no-mistakes(review): read one turn's audio count once, and hush a tidy exit

* no-mistakes(document): fix stale session-end relay_error claim in voice-relay guide
…unchenguid#2811)

A pi worker parked on an interactive prompt - a permission dialog, a
question menu, a trust dialog - reports agent_status=blocked, because it
is waiting on a human keystroke. Pi draws that menu above its separator
pair, so the composer region between the rules is blank and structure
alone looks like a free composer. _fm_composer_pi_verdict admitted
blocked alongside idle and done, so the shared classifier reported an
affirmatively empty composer for exactly the pane where typing is unsafe.

Every "is it safe to type here?" consumer reads that verdict and proceeds
only on an affirmative empty, so both are told yes on a parked prompt:
the away-mode injection guard in bin/fm-supervise-daemon.sh, and fm-send's
pre-type refusal. The keys then answer the menu instead of composing a
message - the highlighted default is selected, the text is discarded, and
the record attributes a decision to a human who never made it.

blocked now defers to unknown, which every consumer already treats as
fail-closed. idle and done still prove an empty composer, so ordinary
steering is unchanged, and Cursor is unaffected because its always-blocked
panes never reach this pi-only branch.

Regression coverage lands first at both levels: the verdict owner
(a blocked pi defers) and the herdr adapter (a parked pi prompt is not an
empty composer).
…2849)

* fix(bin): require a clone root before fleet-sync touches a project

Git repository discovery walks upward, so `git -C projects/<dir>` on a plain
directory nested under projects/ resolves to the enclosing repository - in a
firstmate home, the firstmate checkout itself. fm-fleet-sync.sh guarded its
candidates with `rev-parse --is-inside-work-tree`, which such a directory
passes, so every later git call read, pruned and fast-forwarded firstmate's own
default branch and reported it under the project directory's label. A running
session's AGENTS.md changed underneath it, and the report named a project that
had nothing to do with the change.

Require each candidate to be the root of its own work tree before any other git
command: compare `rev-parse --show-toplevel` against the directory's own
physical path. Both sides are physical, so a symlinked clone still compares
equal. Anything else is skipped by name, naming the repository that would have
been touched, and bootstrap relays that as a FLEET_SYNC line.

Regression coverage reproduces the wrong-repo fast-forward against a home nested
inside another repository, in both the whole-fleet and single-project forms, and
pins that a symlinked clone dir still syncs.

* no-mistakes(review): Keep enclosing fixture clean during clone-root regression
* fix(procevent): retry a transient Lavish poll interruption quietly

A live Lavish listener can be cut short by the server with exactly

    error: Lavish Editor poll response was interrupted
    code: SERVER_ERROR

while the session's marks remain available. Firstmate registered raw
`lavish-axi poll` output, so the generic process-event runner captured
that transient response as a result and woke the whole fleet over what is
really an internal retry.

The Lavish adapter now registers its own listener command, which reruns
the published blocking poll up to 12 times at 5 second intervals for that
one exact two-line response. The match is deliberately narrow: real
feedback, ended and missing sessions, any other SERVER_ERROR, and the same
interruption still standing once the bound is spent all pass straight
through and are captured and announced as before. The retry is a Lavish
fact, so the generic runner stays adapter-agnostic.

`FM_LAVISH_POLL_RETRY_DELAY` is a bounded 0 to 60 second override for the
interval only, refused rather than rounded when malformed, so a test can
exercise the real bound without waiting it out.

* no-mistakes(review): Harden Lavish retry matching, validation, and cleanup

* no-mistakes(review): Bound Lavish retry staging and stabilize regression

* no-mistakes(document): docs: explain Lavish retry adoption

* no-mistakes(lint): Restore Lavish trap ShellCheck suppression
… gate (kunchenguid#2838)

The unguarded Herdr declaration quoted `{TASK}` in its own prose while the
scaffold instructs firstmate to replace every `{TASK}` placeholder. The
documented global replace therefore spliced the whole task body into the
middle of the safety gate's sentence, silently destroying the one contract
that exists precisely because the scaffold cannot inspect the task text.

Reword the gate to refer to the task text filled in above, leaving the
placeholder only at its genuine fill site. Rewording rather than renaming the
token keeps the unfilled-charter guards in fm-home-seed.sh and
fm-remote-home-seed.sh working unchanged.

Add a regression test that performs the documented global fill on ship and
scout scaffolds and asserts the body lands once and the gate survives.
…tat form (kunchenguid#2837)

The writer lock's stale-lock branch read the lock's mtime with
`stat -f %m ... || stat -c %Y ...`. On GNU coreutils `-f` is filesystem
stat, so it consumed the format string as a path, complained on stderr,
printed a partial filesystem dump ("  File: ...") on stdout, and still
exited 0. The GNU form in the fallback therefore never ran, and the
following arithmetic evaluated the word `File`, aborting the writer under
`set -u` with "File: unbound variable".

fm-teardown.sh died there after returning the worktree, leaving
state/<id>.meta, .status, .busy-gen, .busy-state, .busy-state.lock/ and
.turn-ended behind. The surviving metadata kept the watcher monitoring an
endpoint whose agent was gone, so a finished task produced stale wakes
forever, and every re-run died identically because the abandoned lock was
never broken.

Detect the platform once and pick the right stat form, the pattern
bin/fm-watch.sh already documents, and treat any non-numeric result as
"just created" so a future portability surprise degrades to a lock-timeout
refusal rather than killing teardown mid-way.
* fix(stow): give memory decay a per-pass horizon so the clock fires

The tiered decay clocks were wall-clock only, while admission is per-pass:
each /stow admits the findings that pass produced. In a home that stows
daily those two rates diverge by the stow cadence, an entry the fleet keeps
exercising never reaches 30 days unreinforced, and memory only grows while
the pass reports decay evaluated.

Give each dated marker an optional unreinforced-pass counter and make both
tiers stale at whichever horizon comes first: 10 passes or 30 days for
aging, 3 passes or 7 days for perishable. Reinforcement clears the counter
and nothing else does, so the existing evidence-based restamp rule stays
the only way an entry renews its lease. An absent /N means zero, so entries
that stay exercised carry no extra marker bytes, and a rarely stowed home
keeps its current behaviour through the unchanged date horizon.

* no-mistakes(document): Align stow workflow with dual decay clocks

* fix(stow): make the per-pass decay horizon opt-in

The unreinforced-pass horizon shipped as a new default archival cadence,
which is a product default rather than a restoration of the existing
wall-clock contract. Keep the 30-day and 7-day horizons as the only
default clock, and put the 10-pass and 3-pass horizons behind an explicit
opt-in: config/stow-pass-horizon for the firstmate home, and the file's
own header pointer for the public skill.

With the opt-in absent no counter is written and no counter is read, so a
home that does not ask for it decays exactly as it does today.

* no-mistakes(review): Preserve frozen counters and correct archive provenance
…artup (kunchenguid#2876)

tests/fm-watcher-lock.test.sh passed in isolation but failed intermittently
under full-suite and ambient concurrent load. bin/fm-watch-arm.sh computes its
confirmation deadline immediately after forking the real child watcher, so the
child's entire fork, exec, lock acquisition and beacon publication has to land
inside that wall clock. Two cases shrank that budget to one second, leaving a
two-second window for work measured at 3.1-4.9s under CPU oversubscription, so
the arm honestly reported "FAILED - no live watcher with a fresh beacon" and
their premises collapsed. A third case ran on the production budget, but its
child must also execute a registered check before exiting: measured at 1.9-2.3s
idle and 9.1-13.1s under load, against an 11s budget.

The two cases that must confirm a real child now hold the arm to production's
own budget instead of a shrunken fixture one, the immediate-wake case gets an
explicit budget with headroom over its measured loaded cost, and the two waits
for the arm's typed failure are sized off the largest production default rather
than a fixed eight seconds.

No bin/ change and no default behavior change: the lock's fail-closed semantics,
SIGSTOP handling, stale-heartbeat detection and the arm's typed failures are
untouched. Verified 4/4 green at 3x CPU oversubscription (loadavg 75-80) after
3/3 red before the change, and CONTRIBUTING.md records the convention.
* fix(bin): order discovered tool installs by the shell's own expansion

fm_remote_job_compose_operator_path built the asdf and mise install
directories with `compgen -G`, which does not sort. Bash sorts glob
matches in pathexp.c, on the shell's own pathname-expansion path only;
`compgen -G` reaches the same glob_filename through pcomplete.c, which
sorts nothing. On bash 3.2 (macOS /bin/bash) and every bash before 5.3
that handed the composition raw readdir order, so which install of a
multi-version tool a remote job resolved was decided by directory order
on disk rather than by this composition.

Expand the globs at the call sites and let the function take the matches,
so the composition and the documented portable-PATH contract are the same
operation. Quoting the account home at the call site also stops a home
whose name contains glob metacharacters from being reinterpreted.

The colocated regression pins both the order and the mechanism: bash 5.3
moved sorting into the glob library, so an order-only assertion cannot
see the defect there.

* no-mistakes(review): Remove source-reading PATH regression guard
…2848)

* fix: surface stalled secondmate queues and wake handoffs

* no-mistakes(review): Make handoff wakes retryable and stall alerts crash-safe

* no-mistakes(review): Prevent duplicate handoff wakes and cover remote delivery

* no-mistakes(review): Serialize local handoffs and preserve pre-move wake intent

* no-mistakes(review): Serialize teardown with handoffs and retain remote wake confirmation

* no-mistakes(review): Reconcile correlated handoff wake delivery after crashes

* no-mistakes(review): Keep failed wakes retryable and isolate stall receipts

* no-mistakes(review): Reset known-undelivered wake attempts for durable retries

* no-mistakes(review): Refuse duplicate sends for unresolved delivery attempts

* no-mistakes(review): Atomically restore retryability after reconciled send failures

* no-mistakes(review): Serialize delivery confirmation with reconciliation

* no-mistakes(document): Document routed wake and stall supervision

* no-mistakes(lint): Fix ShellCheck expansion and subshell warnings

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes(review): Retire stale wake state and defer pre-move wakes

* no-mistakes(review): Secure markers, bind batches, and preserve teardown routes

* no-mistakes(review): Preserve unresolved prepared wakes across unrelated handoffs

* no-mistakes(review): Preserve prepared wakes before unrelated moving handoffs

* no-mistakes(document): Document prepared wake batch ownership

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes(review): Make local wake retirement recoverable

* no-mistakes(document): Clarify handoff recovery and teardown documentation
…guid#2856)

* feat(bin): steer local tasks by durable inbox record plus constant doorbell

Stage 1 (local steers) of the captain-adopted reframe in
data/fm-send-reliability-reframe-s1/report.md: an ordinary fm-send text
steer to a task recorded in this home is appended as a sequenced durable
record under state/<id>.inbox/ and the terminal receives only one constant
self-describing doorbell line, best-effort. The worker acknowledges by
moving the record into handled/; the watcher re-rings an unacknowledged
message on an idle pane and escalates once as an ordinary stale wake.
--resolve-key closes decisions at enqueue time, because the durable
enqueue IS delivery to the task's record. bin/fm-task-inbox-lib.sh owns
the record format, doorbell line, and re-ring ladder.

The typed plane remains for what must reach the terminal itself:
lifecycle keys, harness-native slash and codex $-skill invocations,
explicit backend targets, and the remote secondmate leg (unchanged until
the remote inbox leg ships separately). The composer classifier is
demoted from delivery proof to an advisory ring guard that skips only on
a proven pending verdict.

Verified live against claude, codex, opencode, pi, grok, and muse: each
real worker read its record, acted, and acked with the mv
(docs/verification/runtime-backends.md "Steering-inbox doorbell").

* docs(verification): flag the grok 1.0.5 composer-matrix staleness observed by the doorbell run

* test(captain-hold): read the chat-channel answer from the durable inbox record

* test: migrate fm-control's marker contrast to the inbox record and fix macOS wc padding in the tool-update suite

* no-mistakes(review): Harden inbox locking, teardown races, and acknowledgements

* no-mistakes(review): Serialize watcher actions with inbox acknowledgements

* no-mistakes(review): Bound metadata locking and tighten acknowledgement rechecks

* no-mistakes(review): Preserve exact inbox bytes and harden delivery recovery

* no-mistakes(review): Harden watcher bookkeeping against concurrent inbox teardown

* no-mistakes(document): Update inbox and typed-plane documentation

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* revert(pipeline): keep parser-native secondmate marking and the both-failed exit out of stage 1

The CI monitor's fix changed the secondmate marking contract for
parser-native invocations (appending the marker after the text) and
softened the both-commit-and-marker-failed branch to exit 0. The merge
authority ruled the marking question out of scope for this stage-1
transport PR (follow-up: fm-send-secondmate-harness-invocation-r1) and
ruled the both-failed case a loud nonzero local failure. Restore both,
keeping the monitor's legitimate migrations and hardening.

* no-mistakes(document): Document inbox and typed-plane boundaries

* no-mistakes(document): Scope backend transport docs to typed plane

* no-mistakes(document): Clarify inbox attempt-budget documentation

* no-mistakes: apply CI fixes

* fix(send): the durable record alone governs the inbox exit status

Captain-refined ruling on the F2/Greptile finding: the durable inbox
record is what delivers the steer, so pending-reply bookkeeping trouble
after a successful enqueue never exits nonzero - a resend-inviting status
would make automated callers enqueue the delivered instruction again
under a new sequence. With the recovery marker stored the watcher
reconciles silently; with the commit and marker both lost the send
surfaces a distinct reply-tracking-degraded do-not-resend warning and
still exits 0. Nonzero remains only where nothing was delivered (or a
decision close needs its manual command). Regression: record durable +
both bookkeeping writes lost -> exit 0, one record, no duplicate.

* no-mistakes(review): Preserve inbox ordering with drain-all doorbells

* no-mistakes(review): Surface unwritable inbox ladder bookkeeping

* no-mistakes(review): Silence ladder failures after inbox acknowledgement

* no-mistakes(document): Update steering inbox documentation

* no-mistakes: apply CI fixes
* feat: add fast local lint mode

* fix: preserve complete fm-lint help

* fix: isolate fast lint mode

* no-mistakes(document): Clarify lint mode documentation ownership

* no-mistakes: apply CI fixes
…#2901)

* feat(bin): deliver remote secondmate steers through durable task inboxes

Stage 2 of the inbox+doorbell steer channel (stage 1: kunchenguid#2856). A remote
secondmate steer now crosses fm-on.sh as a durable record written
idempotently into the remote home's steering inbox plus a best-effort
remote doorbell, and the last typed-payload steer transport is deleted:

- fm-remote-secondmate-control.sh cmd_send writes the record via the new
  fm_task_inbox_write_idempotent and rings the doorbell; it no longer
  types the payload through an inner fm-send at an explicit pane target.
- fm-send.sh routes every remote text steer (harness-native included,
  which marking already reduced to chat) onto the remote inbox leg,
  retries the identical leg once on ssh 255, closes --resolve-key
  decisions at enqueue for remote too, and preserves a marked request's
  reply expectation when completion stays unknown. The exit-3-as-
  delivered remap, the 255 do-not-resend trap, and the remote typed
  submit block are removed.
- fm-task-inbox-lib.sh owns the idempotent enqueue: an exact-body re-run
  lands on the existing record, handled or not, so an ambiguous
  transport can always be safely re-run.
- Tests pin the new contract end to end (record + doorbell + no typed
  payload across ssh, one-record idempotence under an ambiguous
  transport, enqueue-time decision close, loud real failures, and the
  deleted typed-payload behaviors gone), and AGENTS.md plus
  docs/remote-secondmates.md describe the remote leg's new semantics.

* no-mistakes(review): Harden remote inbox delivery against lifecycle races

* no-mistakes(review): Enable correlation-preserving remote steer resends

* no-mistakes(review): Fail closed on stale correlation resends

* no-mistakes(review): Include home context in remote resend commands

* no-mistakes(review): Lock and revalidate remote parent routes

* no-mistakes(document): Clarify remote steer retry documentation

* no-mistakes: apply CI fixes
* wip: forked supervision on Pi (checkpoint before docs)

* fix(pi-branch): harden mirror delivery, fallback encoding, and session replacement

Peek-then-shift mirror flush so a failed append retries instead of dropping;
durable mirror cursor commits only after delivery into the branch;
the main fallback wake is operational-encoded like every watcher injection;
session_shutdown quiesces the generation and session_start re-arms, so /new
and /resume no longer kill the branch permanently. Registers the extension in
the strict typecheck, adds the dispatch handshake test, the branch extension
suite, the bash-level regression suite, the session-start replay test, and
the opt-in real-SDK live guard.

* test(fixtures): carry the branch-dispatch lib and lease lib into isolated fixtures

The watcher extension now imports lib/fm-branch-dispatch.ts and fm-teardown
sources fm-lease-lib.sh, so every fixture that copies or symlinks those
files in isolation gains the new sibling.

* no-mistakes(review): Prevent shutdown wake loss and serialize lease claims

* no-mistakes(review): Durably hand off wakes and retain portable leases

* no-mistakes(review): Require durable reports and clear disposed branch leases

* no-mistakes(review): Enforce per-wake outcomes and quiescent lease cleanup

* no-mistakes(review): Require wake acknowledgements and tighten branch lifecycle boundaries

* no-mistakes(review): Require complete acknowledgements and replay cleanup failures

* no-mistakes(review): Bind supervision to lock ownership and durable delivery

* no-mistakes(review): Activate branch lazily after session lock acquisition

* no-mistakes(review): Preserve undelivered mirror context across extension rebinds

* no-mistakes(review): Acknowledge startup replay only after main delivery

* no-mistakes(review): Isolate replay metadata from untrusted digest content

* no-mistakes(review): Reject duplicate reports for active wake sequences

* no-mistakes(review): Retain failed fallbacks and deduplicate outcome replay

* no-mistakes(review): Deduplicate durable outcomes and cache delivery receipts

* no-mistakes(review): Anchor wake sequence matching to outcome fields

* no-mistakes(document): Clarify Pi supervision durability contracts

* no-mistakes(lint): Fix ShellCheck issues in branch supervision scripts

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* refactor(pi-branch): collapse to confused-agent-grade guards per captain decision

Captain decision A: the lease/actor guards target the CONFUSED-AGENT threat
model bin/fm-gate-refuse-lib.sh already documents; adversarial-grade
separation is impossible in the shared-process design and is filed as
separate follow-up work. Rip out the machinery that chased it: the
generation fence and shell-provenance markers, the wrapper-tagged ancestry
walks, guard auto-claim with per-script release traps, the pending-wake
files and ack-receipt correlation (the durable wake queue already
re-presents anything unacknowledged), the delivery-receipt store with
contiguous cursor advancement, the session-start replay-metadata channel,
and the branch tool quiescence counters.

Keep the behaviors the board requires, each on its simplest implementation:
lazy per-action session-lock ownership (cold start activates after the lock
lands; a secondary session stays inert), mirror durability across extension
rebinds via the durable cursor, replay-exactly-once from the one read
cursor, the awaited operational-encoded fallback, per-generation stray-lease
cleanup, session-lock-bound lease liveness (a recycled pid or a non-Pi home
never honors a leftover lease), the loud accidental-override guards
(readonly actor prelude, cross-actor claim refusal), and the role-partition
refinements (no forced teardown, no direct relaunch for the branch).
Default-on-for-Pi is unchanged.

* no-mistakes(review): Enforce lock ownership and serialize lease mutations

* no-mistakes(review): Synchronize guard cleanup and bind leases to lock owner

* no-mistakes(review): Report outcomes before acknowledging durable wakes

* no-mistakes(review): Restrict leases to Pi and instruct main claims

* no-mistakes(review): Reject malformed lease locks and torn outcome tails

* no-mistakes(review): Validate complete outcome tails before appending

* no-mistakes(review): Guard branch side effects across session replacements

* no-mistakes(document): Update Pi supervision durability and lease documentation

* no-mistakes(lint): Suppress intentional nested-shell expansion warning

* no-mistakes: apply CI fixes

* fix(pi-branch): authorize lease releases by caller

* fix(lint): break redundant source-analysis path in fm-lease-lib.sh

fm-lease-lib.sh's lazy fallback source of fm-wake-lib.sh gave ShellCheck's
--external-sources traversal a second path into an already 1540-line file
that fm-send.sh and fm-teardown.sh also source directly, blowing up the
recursive analysis past CI's lint timeout. Mark it a source=/dev/null
analysis boundary, matching the existing fm-task-inbox-lib.sh convention.

Also restores bin/fm-lint.sh and tests/fm-lint.test.sh to the shared
serial-lint definition (dropping an unrelated parallel-sharding change
that was itself hanging and masked this root cause).

* no-mistakes(document): Correct lease caller-authorization documentation

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes
* feat(bin): parallelize session-start remote secondmate network sweeps

Run per-secondmate liveness and convergence probes concurrently and overlap clone refresh, while replaying each mate's fail-closed diagnostic in original order. Ignore scratchpad* so untracked scratch no longer blocks remote sync.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(document): Document parallel startup network sweeps

* no-mistakes(lint): Fix empty environment assignment lint warning

* no-mistakes: apply CI fixes

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(tests): count declared-pause wakes without crashing on an absent queue

The exited-declared-pause case counts queued stale wakes by handing
state/.wake-queue straight to awk. A watcher that queues nothing never
creates that file, and awk aborts on a missing path before its END rule
runs, so the count collapses to the empty string. The next comparison
then fails as an integer-expression error and surfaces as a wake flood
with no number, hiding the real contract breach the following grep names.

Read the queue the way the drain-count assertion at the end of this file
already does: silence awk's open error and default an absent queue to
zero. Applied to all four counts in this case, including the live
external-decision gate pair whose queue an acknowledged drain can also
leave behind. An absent queue now reports "did not use the bounded
paused recheck", while a genuine flood still fails with its real count.

Fixes kunchenguid#2628

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes
… icon (kunchenguid#2934)

* style(pi): restyle supervision merge notes with a sailboat and matching pad

Secondary-session notes were flush against the TUI edge and fully tinted.
Use the sailboat prefix, Pi's default outputPad, boat-only color, and dim remainder so they sit like real messages.

* style(pi): distinguish routine and captain merge notes by icon only

Visible notes now lead with a sailboat or anchor, then only the dim outcome.
Drop the branch-merged wording and verdict brackets so the icon is the only kind signal.
…id#2938)

The markdown contract stays the owner; the still is only the visual of the idea.
* fix(bin): bound remote worker supervisors

* no-mistakes(review): release incumbent supervisor before starting its replacement

* no-mistakes(review): wait out a healthy same-root supervisor instead of replacing it

* no-mistakes(review): narrow remote worker change to restart accounting only

* no-mistakes(document): clarify supervisor restart guard is a lifetime total
* feat(bin,pi): per-actor wake consume, silent success gating, merge-poll dedup

Three related fixes to the shared wake-drain and Pi supervision-branch
dispatch machinery so a routine success is never main-blocking and a
mixed queue can safely split between actors.

1. Successful routine results no longer create main-blocking wake rows.
   fm-startup-network.sh only enqueues a check: startup-network wake when
   the deferred result is actionable (state is not "done", or the report
   carries a bootstrap-diagnostics actionable prefix); a clean success
   stays durable in the report file without ever waking the agent.

2. Per-actor wake-drain consume contract. bin/fm-wake-drain.sh now scopes
   presentation and --ack-through to the current actor
   (bin/fm-lease-lib.sh's fm_lease_actor): main keeps the original
   whole-queue cutoff behavior, unaffected. A branch actor
   (FM_SUPERVISION_ACTOR=branch, set only inside the Pi supervision
   branch's own bash tool calls) is scoped to an explicit eligible-row
   snapshot instead of a cutoff comparison, so it can never remove a row
   it was not granted - the fix for the swallow risk that used to force
   an all-or-nothing whole-queue fallback to main.
   .pi/extensions/lib/fm-branch-dispatch.ts's scopeForUnreadWake is the
   single owner of eligibility: a check-kind row (merge-confirmation
   polls, Relay mentions, credential/auth failures) is now excluded
   rather than vetoing the whole scan for a non-heartbeat wake, while a
   heartbeat review keeps its original all-or-nothing rule unchanged.
   writeEligibleRowsSnapshot publishes the exact eligible sequence
   numbers before every branch prompt; fm-primary-pi-watch.ts's offer
   still refuses a check-kind trigger outright so a main-only close is
   never itself routed to the branch.

3. A repeat identical merged-PR-poll result for an already-notified task
   is absorbed instead of enqueued again. A poll's own retirement state
   is scoped to one registration and cannot see a prior registration's
   outcome, so a task re-registered after its merge was already surfaced
   would otherwise wake main a second time for the same event.
   bin/fm-pr-lib.sh's new per-task pr-poll-merge-notified marker survives
   across re-registrations to catch that case; the first notification for
   a task still reaches main unchanged.

Regression tests colocated in tests/fm-startup-network.test.sh,
tests/fm-wake-queue.test.sh (including the mixed-queue no-swallow
property), tests/fm-pi-branch-extension.test.sh, and
tests/fm-pr-check-security.test.sh. docs/watcher-continuity.md and
docs/pi-supervision-branch.md updated for the new contracts.

* no-mistakes(review): Bind merge deduplication to canonical PR identity

* no-mistakes(review): Serialize wake row ownership across main and branch

* no-mistakes(review): Bind branch grants and deduplicate within actor claims

* no-mistakes(review): Fallback main-owned wake claims to main delivery

* no-mistakes(review): Clarify silent startup success guidance

* no-mistakes(review): Release residual branch grants after settled prompts

* no-mistakes(review): Reject truncated wake rows as corrupted

* no-mistakes(document): Document per-actor routing and silent startup success

* no-mistakes(lint): Fix ShellCheck findings in wake grant and startup test

* no-mistakes: apply CI fixes
* Hide branch outcome tool in Pi Calm

* no-mistakes(review): Preserve stock outcomes rendering and document tool audit

* no-mistakes(review): Document branch read tool audit disposition

* no-mistakes(review): Match stock outcomes output sanitization

* no-mistakes(document): Document Calm custom-tool visibility
* fix: delegate no-mistakes PR gate to pinned action

* no-mistakes(document): Document commit-bound no-mistakes attestations
…uid#3028)

* feat(pi): let operators pin a cheaper supervision-branch model

Supervision is an easier job than the captain's own conversation, so the
Pi supervision branch does not need main's model. A new /supervision-model
command opens Pi's own selector over Pi's own catalog of credentialed
models, plus a "Follow main" entry, and persists the pick as one
<provider>/<model-id> line in this home's gitignored
config/supervision-branch-model. Firstmate keeps no model catalog of its
own.

The branch resolves the pin at every branch build - the first wake of a
cold start and the reopen after /new, /resume, /fork, or reload - so the
choice survives all of them, and picking also releases the live branch so
the next wake reopens the same persistent branch conversation under the
new model. An absent, unreadable, or unparseable file means no pin and
keeps today's behavior byte for byte: no model option is passed and Pi
picks the branch's model exactly as before.

A pin naming a model Pi cannot hand back is never silently downgraded
onto main's model: the branch refuses to build and the wake falls back to
the captain-facing main path naming the unusable pin, which is the
extension's existing failure direction.

The choice is home-local and not part of secondmate inherited
configuration, matching the Pi Calm preference precedent.

docs/configuration.md owns the operator-facing schema. Portable
regressions cover pin-present on create and reopen, pin-absent default,
the command's persistence, cancellation, and live rebind, and both
unusable and unparseable pins. The opt-in real-SDK guard proves the
vendor surface the pin reads and that an explicit model wins over the
model a reopened session recorded.

* no-mistakes(review): Fix supervision model runtime and rebind races

* no-mistakes(review): Restrict supervision picker to isolated runtime models

* no-mistakes(document): Document supervision branch model selection

* fix(pi): make the supervision model pin authoritative on every reopen

Clearing the pin with "Follow main" removed the file but the next branch
build reopened the persistent branch session with no explicit model
override, so Pi restored the model that session had recorded - the old
pinned model - while the command reported that the branch now follows
main. The same gap meant an absent pin did not reliably mean
same-model-as-main once a home had pinned once.

The pin file's current state now decides the model on every branch build,
create and reopen alike, overriding Pi's session-state restore. With a
pin, that model. With no pin, main's own current model is applied
explicitly, tracked from the contexts Pi already hands the extension plus
its model_select event, since the branch is built at wake time with no
context of its own. Only when main's model is unknown, or this home's
stored credentials cannot run it in the isolated branch runtime, does a
build fall back to passing no override at all, which is the behavior from
before the pin existed; the branch is never refused over model choice.

The command's notification now reports the model actually applied, and
says plainly when clearing the pin could not apply main's model instead
of claiming a change that did not take effect.

No credential handling changes: the branch still relies entirely on the
stored credentials its own runtime already holds, and the picker stays
restricted to models that runtime can resolve.

Colocated regressions cover pin present on create and reopen, clearing
the pin returning a reopened branch to main's model and specifically not
the old pinned one, an unparseable pin behaving as no pin, and the
unknown-main-model fallback to no override.

* no-mistakes(review): Make unpinned supervision follow main model changes

* no-mistakes(document): Correct supervision model documentation
mremond and others added 27 commits September 16, 2026 07:43
kunchenguid#4627)

* fix: restore published contribution follow-up (Fixes kunchenguid#4469)

* fix(review): Fix contribution freshness and merge actor routing

* fix(review): Restore issue triage and scope contribution follow-up

* fix(test): test: assert one wake per contribution signal

* fix(document): Document contribution follow-up

* fix: restore truthful terminal delivery evidence

* fix(review): Disclose unsupported contributions and deduplicate watcher wakes

* fix(review): Preserve unmeasured unsupported contributions across Bearings

* fix(review): Deduplicate shared contribution wakes and isolate diagnostics

* fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower
…nguid#4658)

* fix(bin): make a remote-reply document gap self-clearing and re-attemptable

A remote mate's undelivered document raised a keyed `blocked` decision that
nothing could ever resolve, and any `data/*.md` substring in any mirrored line
was an unconditional fetch instruction. A mate announcing a report it had not
written yet therefore manufactured a permanent, factually false blocker, and
its own explanation of the false alarm manufactured more.

The reader has no permanence vocabulary: a report still being written refuses
exactly like a path that will never exist. So an undelivered document is now a
durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`,
re-attempted on the next delta and on the channel's own quiet poll, and retired
with a matching `resolved` line naming the local copy once it arrives. The
cursor still advances and no delta stalls on one bad pointer.

Only a structured `report=data/....md` pointer now offers a document, so a path
merely mentioned in prose - including one under another home's mirror tree,
which is provably not that mate's to serve - is never fetched. Offers are
deduplicated across the whole delta, the escalation names each missing document
once and carries the reader's own reason instead of discarding it, and a
strictly increasing notice ordinal keeps a later escalation from being
swallowed as duplicate bytes. A mirrored line still lands once whichever
pointer form it was first written under.

* no-mistakes(review): Require structured pointer token boundaries

* no-mistakes(review): Unify boundary-safe pointer extraction and rewriting

* fix(bin): identify a mirrored line independently of its delivery state

Two defects in the boundary-safe pointer work.

The at-most-once check compared only the all-remote and all-local renderings
of a line, so it could not recognize a mixed one. A line offering two documents
where only the first was deliverable mirrored as local-plus-remote; once the
second arrived, a cursor-loss whole-log recapture rendered the same line
all-local, matched neither alternate, and mirrored a second time. A line's
identity is now the canonical form every boundary-valid pointer would take once
delivered, derived by the same parser that does extraction and rewriting, so it
no longer depends on which documents happened to be deliverable at the time.

The pointer map was passed to awk through the process environment. A delta may
carry up to the configured 1 MiB bound, and an expanded map of delivered
pointers can exceed the platform's exec argument limit, so awk would fail to
start; because no caller checked, the empty result would have been appended as
blank lines while the cursor advanced past dropped status content. The map now
travels in a file, and every call site checks the exit status and stops the
ingest rather than committing a delta it could not render.

Both passes now run once per stream instead of twice per line.

* no-mistakes(review): Abort ingest when document pointer extraction fails

* no-mistakes(review): Exclude structured cross-home pointers from document transfer

* fix(bin): fail open on an undeliverable remote document instead of tracking it

Narrow the remote-reply document fix to the scope the diagnosis actually
requires, as decided after measuring a simpler alternative.

A document the reader cannot deliver now fails open. The mate's line is
mirrored with its own pointer, the cursor advances, and one unkeyed note
carries the reader's reason. A note never enters the open-decision fold, so it
cannot stand open the way the original keyed block did - which removes the
never-clearing false blocker by construction rather than by resolving it.

That makes the durable self-clearing obligation unnecessary, so it goes: the
per-mate pending-documents record, its notice ordinal and resolved
announcements, and the poll-side retry. Canonical line identity goes too, and
with it a way to silently drop a genuine status line; mirroring is back to
at-most-once on exact bytes. The cross-home exclusion goes as well: under
fail-open a cross-home report= either fails harmlessly or is a nested remote
report this mate genuinely holds, which is now relayed again.

Kept: fetching only on a structured report= pointer, the boundary-correct
parser, the file-based rewrite map, and checked extraction and rewrite exit
status. The parser now scans behind a sentinel byte so a rejected candidate can
no longer give the text right after it a false leading boundary.

The reported incident is covered end to end: a report path announced in prose
before it exists raises no decision, and the report still arrives through the
ledger publisher's structured offer once written.

* no-mistakes(review): Preserve source-line identity across remote reply replays

* no-mistakes(document): Document remote reply transfer and replay semantics

* no-mistakes(lint): Fix staging truncation lint checks
* Preserve substantive Calm mid-turn text

* no-mistakes(review): Distinguish newline-preserved replies from short narration

* no-mistakes(document): Document Calm mid-turn preservation boundaries

* no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`
…#4656)

* fix(bin): re-record PR poll identity after a volume device renumber (Fixes kunchenguid#4260)

A volume remount can renumber the state filesystem's st_dev while every
inode and byte stays the same; APFS does this across a reboot. A poll
registration records its sidecar and check as device:inode, so every poll
armed before the remount failed strict validation and the watcher refused
all of them as unauthenticated state checks until each was re-armed by hand.

There are two device comparisons. fm_pr_private_file_valid compares a live
file's device with the state directory's device read in the same invocation:
it refuses a file that is not on the state directory's own filesystem and
already survives a renumber, so it is unchanged. The registration's recorded
identity versus the live identity (from kunchenguid#556, reused by the kunchenguid#932 retirement
receipt) binds the registration to the exact files published in its own
transaction; its device part is what breaks.

When strict capture fails, the watcher now proves the device is the only
difference: every other artifact check passes (template bytes, both hashes,
private mode, single link, live device, metadata), both recorded identities
name one device, and each recorded inode equals its live inode. Only then,
under the task's control lock, does it rewrite the two identity lines,
repeating the whole proof and comparing the registration's file identity and
bytes just before the rename, and then capture strictly again. A swapped,
altered, re-moded, relinked, split-device, or foreign-device artifact still
fails a proof and is still refused, and a pending retirement receipt blocks
the rewrite.

Reproduction: on macOS a poll armed on an APFS disk image that was detached
and re-attached behind another image moved st_dev 16777239 -> 16777243 with
inodes, bytes, mode, and link count unchanged; the real watcher refused it on
main and reports its merge with this change. The portable regression test
rewrites a real registration's recorded device and drives the watcher.

Not changed here: the status presentation cursor keys rows by its own
device:inode identity in bin/fm-classify-lib.sh, a different helper that
needs its own fix; a retirement receipt left by a reboot between its
publication and removal still names the old device and stays refused; custom
check trust binds only a content hash and is unaffected.

* fix(review): Serialize PR poll publication writers

* fix(review): Bound PR poll publication lock scope
…llow-up to kunchenguid#4627) (kunchenguid#4661)

A budget that expires partway through an observation no longer records an
error or prints the unavailable wake; the URL keeps its prior record and is
observed first next poll. forge() flags budget exhaustion at the point it
refuses, or when a read is killed at the budget's own deadline, so a genuine
forge failure still records the error and wakes. Each distinct URL is now
observed once per poll and applied to every owning task.
…kunchenguid#4680)

* fix(bin): clear parent pending-replies on local secondmate retirement

Local secondmate teardown left resolved parent pending-reply records behind
after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced
retirement while any reply for that id is still unresolved, and delete every
matching record plus its delivery confirmation after a successful local or
remote retirement, matching the remote cleanup path.

* no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup

* no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung

* no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern

* no-mistakes(review): Pending-replies Basename und corr_id abgleichen

* no-mistakes(document): Clarify forced retirement pending-reply cleanup

---------

Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>
kunchenguid#4677)

* fix(bin): accept Orca's composite worktree id at teardown

Teardown refused every Orca-backed task because the endpoint validator
checked orca_worktree_id with the simple-atom rule meant for tmux-style
window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca
returns that id as `<orca id>::<absolute worktree path>`, so the colon and
slashes in every real value made validation fail and finished Orca tasks
could never be cleaned up.

Validate the field as the composite it is: both halves of the first `::`
split present, the path half absolute, and no embedded newline, carriage
return, or tab. The terminal field keeps the atom check, which is correct
for it, and no other backend's validation changes.

The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca
never returns, which is why the suite passed a check the real value fails.
They now carry the composite form, so the tests exercise the real value.

* no-mistakes(document): name Orca's repo id in the composite worktree id

* no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points
* feat(bin): add opt-in typed dispatch resolution through typesafe.ai

Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or
scout profile from a written brief with typesafe.ai's System One model:
one Choice question over the rules' `when` texts, then the confidence
floor, the rule's `approval` and `floor`, each profile's `provider` and
`floor`, one quota-axi snapshot, and the spendPriority argmax all in code.
It is off unless TYPESAFE_API_KEY is in the environment or the home's
gitignored .env; off means one stderr line, exit 0, and no network call,
so firstmate dispatches exactly as before. The key reaches curl on a file
descriptor, never argv.

Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and
the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new
tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates
the four new optional dispatch fields. Document the schema, the operator
contract, the AGENTS.md intake step, and the live and benchmark evidence.

* no-mistakes(review): Harden typed dispatch resolution and quota bounds

* no-mistakes(review): Validate dispatch floors and ranking evidence

* no-mistakes(review): Tighten dispatch response and floor evidence

* no-mistakes(review): Neutralize none matching and resolve defaults locally

* no-mistakes(review): Preserve providerless profiles outside typed resolution

* no-mistakes(review): Validate response usage and reject duplicate profiles

* no-mistakes(review): Escalate unverifiable floors and validate probabilities

* no-mistakes(review): Validate probability mass and unknown profile floors

* no-mistakes(review): Simplify resolver interface and preserve fallback routing

* no-mistakes(review): Fix constants and rank partial quota evidence

* no-mistakes(review): Add authoritative provider mapping and enforce explicit providers

* no-mistakes(review): Declare provider for documented Pi profile

* no-mistakes(review): Validate provider identifiers and support Gemini dispatch

* no-mistakes(review): Strictly anchor provider identifiers

* no-mistakes(review): Validate selectors and preserve fallback candidate evidence

* no-mistakes(review): Gate typed validation and harden resolver evidence

* no-mistakes(review): Preserve opt-in routing and harden candidate evidence

* no-mistakes(review): Prioritize known exhaustion over quota uncertainty

* no-mistakes(review): Isolate API secrets and preserve no-key diagnostics

* no-mistakes(review): Fallback safely when dispatch rules are absent

* no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets

* no-mistakes(document): Document typed dispatch safety and fallback behavior
…n decisions aren't lost (kunchenguid#3753)

* test: reproduce buried status declarations in shared readers

* fix: share status event reads and preserve open blockers

* fix: retain terminal scout and ship status declarations

* no-mistakes(review): Fix status chronology, legacy completions, and reader performance

* no-mistakes(review): Share terminal decision reconciliation across fleet snapshots

* no-mistakes(review): Unify terminal supersession across cached folds and consumers

* no-mistakes(review): Filter per-key status history while preserving terminal chronology

* no-mistakes(test): Preserve parent lock ownership in Bash 3.2 subshells

* no-mistakes(review): Anchor legacy status tokens so prose cannot hide pauses

* no-mistakes(document): Document latest-event status read and kind-scoped fold cursor

* no-mistakes(lint): Quote literal done in test for-lists for SC1010

* ci: expect 19 snapshot/fleet-view tests

This branch adds a fleet-snapshot regression, so the stock macOS Bash
lane's hardcoded guard of 18 'ok - ' lines fails on the new count.
Bump the guard and its message to 19.

* no-mistakes(review): Restore multiline child outcome reporting

* no-mistakes(review): Select ledger terminal events through bounded shared reader

* no-mistakes(review): Report newest open decision instead of preferring blocked

* no-mistakes(review): Require colon before ship/scout terminal supersession in fold

* no-mistakes(review): Gate socket-down override on latest event; drop lock matrix

* no-mistakes(review): Fold only colon-bearing or keyed lines as decision transitions

* no-mistakes(review): Pre-select candidate lines before per-key closing-verb fold

* no-mistakes(test): Update fleet-view expectations to newest-open-decision rule

* no-mistakes(document): Align status-read docs with fold-resolved crew state

* no-mistakes(document): Correct status-reader contracts in classify-lib and crew-state headers

* no-mistakes(ci): Greptile P1 (bin/fm-crew-state.sh:729, "Stale socket blocker survives") was a real defect introduced by commit b7c2183 on this branch, and is fixed. Root cause: the daemon-socket-down override took its verb check from `last_status_line "$LOG"` but its evidence and emitted detail from `$LOG_LINE` (status_current_line = the fold's newest still-open decision). Those are different lines whenever a later recognized `blocked:` event is one the decision fold declines. Reproduced by sourcing bin/fm-classify-lib.sh on `blocked: no-mistakes daemon socket is missing` followed by `blocked [key=pending-reply-t3]: still waiting on the answer` (reserved-namespace key whose note does not speak that vocabulary, so _fm_decision_key_transition_allowed rejects it): open set still holds the socket blocker, last_status_line returns the newer line, its verb is blocked, so the gate passed and the stale daemon-down evidence overrode a healthy attributed run. Fix (bin/fm-crew-state.sh): capture LOG_LATEST=$(last_status_line "$LOG") once and read verb, socket-down evidence, and the emitted note all off that same line, so the override fires only while the socket-down declaration is itself the log's latest recognized event — preserving the narrow override the prior round's user instruction asked for. Comment updated to state that contract. No new machinery; the two-line conflation was removed rather than papered over. Regression: extended tests/fm-crew-state.test.sh:test_socket_refusal_override_expires_when_the_crew_moves_on with the reproduced sequence, asserting the run-step reading (state: working, source: run-step) and absence of the override detail. It fails before the fix ("not ok - a later unfolded blocked event also hands the reading back to the run (missing: 'state: working')") and passes after. Verified locally: tests/fm-crew-state.test.sh, tests/fm-fleet-snapshot-view.test.sh, tests/fm-classify-decision-key.test.sh, tests/fm-watch-triage.test.sh, tests/fm-captain-hold-lifecycle.test.sh all pass; bin/fm-lint.sh (shellcheck 0.11.0 + actionlint) exits 0. Changes left uncommitted in the worktree

* test: fold terminal-cleanup snapshot coverage into the completed-scout case

Keep the ship/scout/secondmate supersession assertions without adding a
nineteenth top-level fleet-view test, so CI can stay at the upstream suite count.

* no-mistakes(document): Clarify socket-down override expiry in architecture doc

* ci: retrigger flaky contribution check
…nchenguid#4689)

* fix(spawn): launch codex crewmates with codex's hook layer disabled

A freshly launched Codex worker never reached its instructions. Codex
stopped it on an interactive "Hooks need review" modal whose selection
sits on "Review hooks", which is neither trusting nor declining.
Firstmate's key plane carries only Enter, Escape and Ctrl-C with no arrow
navigation, so the selection cannot be moved, and pre-accepting the
prompt by writing Codex's own trust store would record an operator
consent that was never given.

The hooks are the machine's own ~/.codex/hooks.json plus any project's
.codex/hooks.json. A crewmate needs neither: its turn-end signal is the
-c notify= program on the same launch, and Firstmate's project hooks are
primary-session infrastructure that stands down in a child worktree.

Crewmate and scout launches now pass --disable hooks. That is the
opposite of --dangerously-bypass-hook-trust, which RUNS the untrusted
hooks; disabling the feature runs none of them and leaves the operator's
~/.codex untouched. An unknown feature name is a hard Codex error, so a
release that drops the flag fails the launch loudly instead of silently
restoring the modal. A secondmate is a primary in its own home and keeps
the project hooks its turn-end guard and session-start digest ride on.

Verified on codex-cli 0.151.0: the modal is gone and the turn-end
notification still lands.

This unblocks the second review that every finished pull request is supposed to get.

Fixes kunchenguid#4673

* no-mistakes(review): Fix contradictory hook count in Codex verification record
…d#4669, Fixes kunchenguid#4670) (kunchenguid#4710)

* fix(bin): settle terminal contributions and wake once per read-failure episode

A contribution whose last good observation is merged or closed is final:
poll no longer re-reads it, projection keeps it fresh, and a stale error
recorded beside it is cleared once. A genuine forge-read failure on an open
contribution still records its error on every cycle but prints the
unavailable wake only when it starts a failure episode; a successful read
ends the episode. Open PRs linked from done tasks keep being observed.

The false unavailable beside a complete observation was budget exhaustion
mid-observation, already fixed by kunchenguid#4661.

* fix(review): Settle terminal contribution owners

* fix(review): Deduplicate shared contribution failure episodes

* fix(test): Preserve settled terminal contribution records
* fix(crew-state): select authoritative validation runs by identity

Use the AXI run overview and id-addressed status reads to preserve replacement review gates, report competing live runs as unknown, and retain newer failures. Keep the coarse ledger in creation order rather than preferring an older live row.

Refs: kunchenguid#3215

* fix(review): Resolve same-branch run identities beyond capped history

* fix(review): Fix run-selection compatibility, races, and worker-state fallbacks

* fix(review): Limit run validation to the requested branch

* fix(test): Anchor AXI fixtures and document remaining live evidence gaps

* fix(document): Clarify run selection documentation and capture ownership

* fix(lint): Fix ShellCheck diagnostics while preserving fixture isolation
* fix(AGENTS): send a captain-facing outcome instead of shipshape for finished requested work

MAIN answered a supervision-branch outcome for completed captain-requested
work (implementation done, PR ready for review and merge approval) with
"Captain, shipshape.", reading section 9's no-action reply as covering it
and reading the Pi protocol's "do not re-emit the anchor verbatim" as "no
captain-facing response is owed".

Section 9 now limits the shipshape reply to true no-ops (idle re-read,
empty heartbeat, consequence-free acknowledgement) and requires a short
outcome response naming what finished and what word is needed whenever
requested work finishes or a result needs the captain's word, even when a
transcript entry already shows the substance. The Pi protocol's re-emit
rule now says it bounds repetition only, and carries a worked example of
the ready-for-review outcome whose correct processing turn a shipshape
reply fails.

No executable contract evaluates the content of MAIN's captain-facing
reply, so the regression is the protocol example in the owner doc rather
than a text-match test.

* no-mistakes(document): Clarify captain-facing outcomes versus no-ops

* docs(pi): restore the ready-for-review regression example as a preserved-verbatim contract line

The document step condensed the Pi protocol's re-emit rule and dropped the
worked example of a finished, ready-for-review outcome whose correct
processing turn a "Captain, shipshape." reply fails. That example is the
contract's regression: no executable contract evaluates the content of
MAIN's captain-facing reply, so the owner doc's example is the test case.

Restore it directly under the re-emit rule, prefixed as a regression
example that is kept verbatim and never condensed or summarized away.

* no-mistakes(review): Clarify captain outcome and decision-word requirements

* no-mistakes(document): Clarify captain-facing completion outcomes

* docs(pi): require the PR URL in the visible captain-facing outcome reply

Captain review on the regression example: drop the sample reply string
and say only that the ready-for-review outcome requires relaying a
captain-facing outcome response, not just "Captain, shipshape.".

Fold in the visible-PR-handoff failure seen this session: after the
branch outcome reporting this fix green, MAIN's visible reply was only
"Awaiting your merge call." with no PR URL, leaning on the dim anchor.
Section 9's URL rule now also covers a review or merge ask and names the
visible reply as where the URL goes, sourced from the ready status, pr=
metadata, or the supervision branch's summary and never left to a
transcript entry. The Pi protocol adds the same-way failure and places
the captain-facing text in the final visible assistant reply after the
fm_branch_processed call, because Calm hides assistant text emitted in
the same step as a tool call as a working note.

Investigation verdict, evidence in the PR comment: no recent PR caused
the handoff failure; Pi has hidden same-step pre-tool assistant text
since kunchenguid#2339 (2026-08-13), kunchenguid#4655 changed only the Claude Code mod, and
kunchenguid#4658 touched only remote report transfer.

* no-mistakes(review): Restore safe outcome ordering and consolidate PR URLs

* no-mistakes(document): Clarify captain-facing supervision outcomes

* docs(AGENTS): keep the whenever-a-PR-is-mentioned trigger on the consolidated URL rule

The consolidated section 9 URL rule narrowed its trigger to a review or
merge ask, dropping the "whenever a PR is mentioned" catch-all from
kunchenguid#3648 that keeps every PR URL copied from a durable record and never
assembled from memory. Restore that trigger as a union with the review
or merge ask so the one consolidated rule covers both.
* Fix foreign-owner turn-end supervision loop

* no-mistakes(review): Scope foreign-owner safe exit to Claude guard

* no-mistakes(document): Document Claude foreign-owner safe exit
…orb (kunchenguid#4778)

Under set -u, stock macOS bash 3.2.57 treats "${arr[@]}" on an empty
indexed array as an unbound variable and aborts the shell. In
signal_turnend_panes_churned() the missing_keys loop was reachable with
an empty array whenever every churned key already held a fresh
.churn-since-* marker (a second churning turn-end inside an open
deferral window), so each watcher cycle died about half a minute in and
supervision restarted endlessly. The created_keys rollback loops had the
same latent crash on their error paths.

Audit of bin/ for the same pattern found one more confirmed-reachable
case: remote_handoff's noncanonical-body scan iterates to_move, which is
empty when a retried remote handoff finds every key already staged in
the outbox. All other "${arr[@]}" sites are either count-guarded,
guaranteed non-empty by construction, or unreachable while empty.

Guard the three reachable expansions with the repo's existing
"${arr[@]+...}" idiom. New regression test drives a real watcher
through the all-marked churn path; the macos-stock-bash CI lane runs it
under real /bin/bash 3.2 via FM_TEST_ONLY.
… lock. (kunchenguid#4783)

The synthetic harness was named synthetic-claude, which Linux procps truncates to synthetic-claud so fm-lock.sh never matched a harness or wrote state/.lock before the test read it.

Co-authored-by: Cursor <cursoragent@cursor.com>
* docs: require complete final responses across harnesses

* no-mistakes(document): Document complete final replies for Grok Bot

* docs: point Grok replies to the shared contract owner

* no-mistakes(review): Clarify final recap without batching decision asks
* fix(calm): preserve substantive Pi mid-turn text

* no-mistakes(review): Preserve substantive Pi Calm text per block

* no-mistakes(test): Cover shared Calm preservation boundaries behaviorally

* no-mistakes(document): Consolidate Calm preservation documentation
)

* Improve CI reliability and rebalance full-coverage validation

* no-mistakes(document): Clarify lint partition documentation
…guid#4799)

* Handle Kimi workspace trust dialog

* no-mistakes(review): Retry Kimi trust Enter and gate ready on dialog markers

* no-mistakes(review): Gate Kimi ready on any trust marker and clean captures

* no-mistakes(review): Read visible pane for Kimi trust and ready gates

* no-mistakes(review): Add per-backend visible-pane capture for Kimi trust gate

* no-mistakes(review): Harden Kimi viewport capture and trust dialog detection

* no-mistakes(document): Document Kimi spawn refusal on cmux and Orca
…er (kunchenguid#4775)

* fix(bin): report a record whose agent is gone once instead of escalating forever

The wedge escalation path never asked whether there was still an agent to be
wedged. A wedge is something stuck that might recover, so re-alarming it earns
its cost; an agent that is gone never moves again, its pane never churns, the
idle timer never resets, and the escalate path clears its own timer and re-arms
with nothing bounding the count.

Observed on a live fleet: two finished lanes reached 226 and 203 consecutive
escalations, roughly one every FM_STALE_ESCALATE_SECS, indefinitely - about 400
notifications a day from two lanes with no agent running at all. On one,
fm-control.sh exit answered already-stopped and fm-crew-state.sh read
"failed - run failed". Closing the Herdr pane did not stop it either: with the
pane genuinely gone and herdr pane read returning pane_not_found, the count kept
climbing, because the poll is driven by the record's window= line rather than by
the pane. The cost is not the repetition but that it drowns the alarms that
matter.

fm_backend_agent_state already separates a thinking agent from a gone one at
process level. In the branch that was about to escalate, read it once and treat
only its two recovery-grade verdicts - dead (endpoint present, no agent in it)
and missing (endpoint authoritatively absent) - as proof, reporting that record
once and not re-escalating it while it stays that way. Every other verdict,
including alive, ambiguous, unreadable, unverified, and a read that failed
outright, keeps the identical schedule, reason, and escalation count, so a
genuinely wedged live agent is unaffected. The probe costs at most one backend
read per window per threshold, the same budget the declared-wait consult and the
worktree write probe already take.

The report decides nothing about the record's fate: both lanes still held
unlanded work and teardown refusing them was correct, so retiring, relaunching,
or cleaning up stays with the supervisor. The once-only marker is owned entirely
by that function and is dropped by the same read the moment the endpoint stops
reading gone, so a replacement launched into the same window escalates normally
and its own later death is reported again.

Related, and not closed by this: kunchenguid#4412, kunchenguid#4482, kunchenguid#4316.

Tests drive the real watcher against a record whose endpoint does not exist and
pin both directions: dead and missing report once and never advance the count
across later thresholds, while alive, ambiguous, and unreadable endpoints keep
escalating with the identical reason and a climbing count.

* fix(bin): bind the once-only dead report to the pane it reported

Review of the parent commit found a reachable sequence where a later death in
the same window lost its promised report. The marker was keyed on the verdict
string alone and dropped only when a threshold probe read a non-gone verdict,
but probes run only at thresholds: a replacement launched into the same window
that dies without ever being probed alive - it crashes at startup, or works and
then crashes - was absorbed by the previous death's marker. The pane's first
sight yielded only the generic stale wake and every later threshold matched the
stale marker, so the second death never got the detailed once-report that both
the function's own comment and docs/architecture.md promise.

Record the verdict together with the pane hash it was reported for, and absorb a
repeat only while both still match. A replacement churns the pane, which resets
the stale suppressor, wedge timer, and escalation count while no reset site
touches this marker, so the pane half is what tells the second death apart from
the first. The live-probe drop stays as it was.

Clearing the marker at those reset sites instead would re-open unbounded
re-alarming for a dead pane whose display ever ticks, which is the exact defect
the parent commit exists to close.

The noise bound is unchanged: an unchanged dead pane still absorbs on every
later threshold and never advances the escalation count, and every verdict short
of proof still escalates exactly as before.

* no-mistakes(review): Key the dead-record once-marker on the busy incarnation token

* no-mistakes(document): Document dead-record escalation cap in stale-pane config entry

* no-mistakes(document): Add busy-state inventory line to AGENTS.md

* no-mistakes(document): Document dead-record probe on busy-turn-bound wedge path
Bring this fork current with kunchenguid/firstmate so upstream/main is an
ancestor again, land the opt-in typed dispatch resolver, and keep fork-local
capabilities on top of upstream's files.

Fork-local work that landed on top of upstream includes Devin-on-Herdr
crewmates, wiki-layer routing, unedited-brief refusal (now via upstream's
placeholder check, OpenSEO and orchestrated-delivery skills, approved-work
handoff, Matt plugin restore, Hermes Agent tools, decision records, ingress
provenance sanitization, and distinct same-key check results.
EOF
)
Keep the unedited-scaffold abort from c1a8cea in front of delivery_rigor_rank
so a leftover {TASK} in # Task still refuses spawn. Re-source the shared test
environment sanitizer in the two live e2e suites that lost it in the merge.
Honor an explicit FM_HOME in the teardown helper so secondmate parent-channel
fixtures still bind. Keep historical status-log working events as unknown for
idle ships, and give foreign pending-reply fixtures 16-hex names so upstream
recovery validation can accept them.
… Restored fm-lint.sh --source-path=tests so ShellCheck can follow tests/environment.sh (Lint 1/2 SC1091). Copied environment.sh into the Git-config fixture so the nested runner can source it (serial 5). Wired fork-local Devin and Pi-role live e2e tests to fm_live_gate so FM_LIVE=0 emits the shared skip (serial 4). In fm-primary-pi-watch.ts, return true for an already-handled wake (TS2322 / parallel 1) and read session_shutdown event?.reason so a missing event is a terminal quit instead of a crash (serial 9). Verified: live-gate (35 guards), fixtures, watch-recovery-loop, Pi typecheck, and ShellCheck on the previously failing files
Empty commit so git push no-mistakes can attest the current PR head
and fire a new pull_request synchronize event. No product changes.
@yelenplays yelenplays changed the title feat(bin): land upstream firstmate with typed dispatch feat: add new crew adapters, Pi branch supervision, and durable captain operations Sep 18, 2026
@yelenplays
yelenplays merged commit e73d3b6 into main Sep 18, 2026
19 of 20 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.