Skip to content

feat(bin): add cline/openhands adapters, cross-home claims, provider lane caps, and faster wake/session start - #6174

Open
keenvc wants to merge 53 commits into
kunchenguid:mainfrom
keenvc:fm/fm-slowness-scout
Open

keenvc wants to merge 53 commits into
kunchenguid:mainfrom
keenvc:fm/fm-slowness-scout

Conversation

@keenvc

@keenvc keenvc commented Sep 30, 2026 •

Copy link
Copy Markdown

Intent

Re-bind no-mistakes pipeline attestation on keenvc/firstmate PR12 to current head 37bad4e after prior run updated only upstream PR 6174; no further code changes.

What Changed

  • Adds new crew adapters and supervision features in bin/. There are verified cline and openhands crewmate/scout harness adapters, each with verification docs and skill references. Other additions: cross-home work claims (fm-claim.sh/fm-claim-lib.sh), recurring re-verification of aged captain holds (fm-hold-reverify.sh), per-billing-provider live-lane caps at spawn (fm-provider-lib.sh/fm-provider-load.sh), and a per-spawn Claude config-dir seat in fm-spawn.sh. This PR also adds a human-text-discipline skill and makes a before/after evidence pair part of the ship definition of done.
  • Fixes several reliability problems:
    • Herdr CLI calls now have a time limit (FM_BACKEND_HERDR_CLI_TIMEOUT), and projected workspaces are removed without stealing focus.
    • Provider quota walls are classified as their own quota state.
    • A captain hold now revokes standing merge authority.
    • agy's empty composer is confirmed by identity, so control exit works.
    • The secondmate liveness sweep covers every registered secondmate.
    • Cleared lane records are retired and their wakes stop.
    • Aged holds are read through backlog-json.
    • Stale-task teardown is allowed when a pool slot has a new owner.
    • opencode v2 workers launch with --standalone.
  • Speeds up the hot paths:
    • Wake drain, watch, and supervision scripts assign IFS=$'\t' directly instead of forking subshells.
    • fm-session-start.sh starts the home summary refresh early, in the background, so interactive startup doesn't wait on it.
    • The bootstrap reconcile sweep keeps its per-record backlog probes.
    • Tests are added or extended for each change.

🤖 Generated with Claude Code

Risk Assessment

✅ Low: The scout-specific delta applies the previously chosen fixes. The herdr cap is now per-read only via ENDPOINT_HERDR_TIMEOUT, with no export. The summary refresh is launched detached at the early position. The test polls are longer. The rest is a mechanical switch from IFS=$(printf '\t') to IFS=$'\t', and every touched file is bash. The intent asks only to re-bind the attestation, with no new code changes, and the change conforms.

Testing

I drove the real bin/fm-session-start.sh against a disposable lab home. The full digest ran cleanly and published home-summary.json. With the digest forced to truncate at its 5s bound in the wake-queue stage, the detached summary refresh still published. Non-integer and larger herdr cap values caused no shell errors and no truncation banner. The 1s and 3s bounds cut the digest in the lock stage, before the refresh launch point, so the summary was correctly absent there. All seven targeted test files passed. They cover the hung herdr read bound, tab-separated parsing and the backlog read bound, but they are scripted tests, not live runs, so those scenarios are marked untested. The lab home was torn down with no stray processes, and the worktree is clean.

  • Live validation: ✅ go - 3 of 6 scenarios driven live against the product
Scenario Result Live Evidence
Operator starts a session on a clean home: the digest completes and the detached refresh publishes home-summary.json ✅ pass live live-session-start-digest.txt; home-summary.json appeared with schema fm-secondmate-home-summary.v1 and home=$LAB
Adversarial: the digest hits FM_SESSION_START_TIMEOUT during the wake-queue stage and the home summary still gets republished ✅ pass live live-session-start-truncated-5s.txt shows 'STARTUP TRUNCATED ... wake-queue' stage; home-summary.json recreated after being deleted
Adversarial: a non-integer or oversized FM_BACKEND_HERDR_CLI_TIMEOUT causes no shell integer errors and does not break the digest ✅ pass live live-session-start-herdr-cap-{2.5,abc,30}.txt: rc=0, 0 'integer expression' lines, no truncation banner, NEXT STEP present
A hung herdr endpoint read is bounded per read and leaves no herdr processes behind ⏸️ untested no The prior payload did not establish a live result. Only the scripted tests/fm-session-start.test.sh with a fake hanging herdr fixture was run (rc=0). A live check needs a herdr lab session via bin/fm-…
Away-mode return brief, wake-drain backstop and decision cursors, host-mirror cursor and task-inbox ring-state still parse tab-separated rows after the IFS=$'\t' change ⏸️ untested no The prior payload did not establish a live result. Only the scripted end-to-end tests were run (fm-afk-return, fm-wake-drain-outcome-backstop, fm-wake-drain-open-decisions-cursor, fm-host-mirror, fm-t…
The session-start digest still bounds a hanging backlog read, with the git identity pinned in the fixture ⏸️ untested no The prior payload did not establish a live result. Only the scripted tests/fm-backlog-read-bound.test.sh with a fake hanging tasks-axi was run (rc=0). A live check needs a real hanging backlog backend…
Evidence: Live session-start digest (lab home, full run)

Source: Live session-start digest (lab home, full run)


================================================================================
SESSION START - /tmp/fm-lab.TWFKOi
================================================================================

LOCK
--------------------------------------------------------------------------------
lock acquired: harness pid 3748038

BOOTSTRAP
--------------------------------------------------------------------------------
BOOTSTRAP_INFO: lavish-axi >=0.1.80 enables confirmed board replies; this older compatible version retains the legacy reply path, but upgrade to prevent handing back a board before its reply is accepted

WAKE QUEUE
--------------------------------------------------------------------------------
(no queued wakes)
================================================================================
SUPERVISION OPERATING INSTRUCTIONS - primary harness: claude
================================================================================
Current state:
- Lock: held by this session; this session owns normal supervision unless away mode says otherwise.
- Away/quiet mode: inactive.
- X mode: inactive; use the default watcher cadence.
- Supervision host: on; it takes away-posture wakes and, where the dialog mirror is verified, eligible attended wakes itself, and hands the rest to you (protocol at the end of this block).
- Ordinary wake: the Stop-owned auto-arm (bin/fm-claude-stop-autoarm.sh) already owns watcher continuity; drain and handle the wake, and do not arm another cycle yourself.

Mode: Claude Stop-hook-owned supervision.

When this session owns supervision and away mode is not active:
1. Drain first with `bin/fm-wake-drain.sh`.
   After handling all emitted wakes and reconciling open decisions and unread status lines, run the exact `--ack-through` command printed as `WAKE_ACK_REQUIRED`; until then the work remains durable for idempotent re-handling after interruption.
2. Routine watcher arm and re-arm are owned by the Stop `asyncRewake` hook (`bin/fm-claude-stop-autoarm.sh`), never by you.
   Every turn end while supervision is needed launches or attaches one home-scoped watcher cycle with no model command and no model tokens.
   An actionable close wakes you through the hook's exit-2 rewake, delivered as a `Stop hook feedback` message.
3. On a `Stop hook feedback` wake (`signal:`, `stale:`, `check:`, or `heartbeat`), run `bin/fm-wake-drain.sh` first and handle the wake.
   Do not run `bin/fm-watch-arm.sh` after an ordinary wake; the next turn end re-arms automatically when supervision is still needed.
   Do not invent a wake from an attach-status line alone; drain and act only on real wake records, the drain's `OPEN DECISIONS` and `UNREAD STATUS` entries, or a real watcher reason line.
4. On the one `Stop hook feedback` automatic-mechanism failure notice (`firstmate watcher auto-arm FAILED ...`), drain, inspect the automatic mechanism failure, and do not turn the notice into a repeating manual-arm loop.
5. If the Stop hook does not claim the home or reports an exhausted failure, inspect its registration and watcher startup path before ending blind.
   Keep the Stop-owned automatic mechanism as the only Claude arm owner.
6. Treat `watcher: started ...` and `watcher: attached ...` inside automatic arm output as proof that one live cycle exists.
   On attach, the arm follows verified identity-matched successors instead of exiting when the first cycle ends.
7. The durable wake queue preserves actionable events between a rewake and the next Stop-launched arm, while the bounded turn-end guard prevents a blind Stop when recovery did not start.
   No PreToolUse hook denies fleet commands based on watcher status.
   [`watcher-continuity.md`](../watcher-continuity.md) owns the exact session-lock recovery boundary.
8. The turn-end guard (`bin/fm-turnend-guard.sh --claude`) remains the final backstop.
   It requires the PID-strict live-watcher and fresh-beacon predicate at the Stop boundary, except for the Claude-specific foreign-live-owner safe exit owned by [`turnend-guard.md`](../turnend-guard.md#guard-predicates); that document also owns the distinct model-aware mid-turn pull-guard rules.
   Otherwise, it allows the stop when a watcher is healthy or an open auto-arm generation claim owns recovery, while fresh failure epochs advance the bounded one-time attended fail-open progression described there.
9. Waiting on the hook-owned cycle is silent: do not send idle progress while the watcher is parked.

The watcher itself remains `bin/fm-watch.sh`, and `bin/fm-watch-arm.sh` remains the verified arm wrapper that the Stop hook foregrounds on a home that opted out of the [supervision host](../supervision-host.md) (`config/supervision-host` holding `off`).
Re-arm attaches to an existing healthy cycle when one is already present and follows its verified successor chain.
See [`watcher-continuity.md`](../watcher-continuity.md) for the arm-layer successor and clean-close failure contract and the Claude ownership model.

Supervision host: on for this home (`config/supervision-host` holding `off` turns it off; [`supervision-host.md`](../supervision-host.md) owns the design).
The Stop hook runs the supervision host in the arm's place, and everything above still holds with these additions:
1. Attended (no away record, including a quiet-mode record; see [Postures](../supervision-host.md#postures)): a headless supervision session takes the wakes the supervision branch may take and never wakes you for a routine outcome, so fewer wakes reach you; check wakes, decision wakes, and whatever it cannot take still reach you exactly as above.
   `supervision-host: branch-outcome: ...` means it handled a wake and recorded captain outcomes for you: run `bin/fm-wake-drain.sh`, process each entry of its `BRANCH OUTCOMES` section as firstmate from the task's current state, because each entry says how long ago it was recorded (tell the captain, land or merge what is ready, answer or escalate a decision, or act on a blocker; your reply covers only entries still open, as if a settled one, such as a PR since merged, had never been listed), then run the `mark-processed` acknowledgement it prints; every drain presents them again until you do.
   `supervision-host: the supervision session could not take this wake ...` means the wake is yours: handle it as above.
   A failing turn may include a `supervision-host:` health note about repeated engine errors: tell the captain when it matters and handle the handed-back wake as usual; during cooldown later attended closes reach you unchanged.
   Routine outcomes never wake you; your next drain lists only visible routine outcomes under `BRANCH OUTCOMES, ROUTINE` for awareness, with nothing to acknowledge. Silent rows do not appear there, but remain available through `bin/fm-branch-outcome.sh list`.
2. Away (an away record exists and no daemon runs): the host hands each wake to a headless away session that runs the supervision branch's contract under the record, and you are parked.
   Only a wake the host hands back reaches you, as `Stop hook feedback` carrying the close plus one `supervision-host: <why>` line.
   That wake is automatic supervision, not the captain's return: drain and handle it under the away posture, and never run the return from it.
   After the return, a `supervision-host:` line naming the captain's return during a turn means that turn has visible outcomes missing from the return brief, whether the wake was handled or handed back: relay every following `supervision-host: outcome ...` line to the captain (the rows also remain in `bin/fm-branch-outcome.sh list`), then drain and handle any queued wake before acknowledging.
   Each such visible outcome is also a queued `check: supervision-host outcome <n> ... was recorded after the captain returned` wake, which the drain presents until acknowledged: relay each outcome once, whichever arrives first, and acknowledge its `BRANCH OUTCOMES` entry too when it has one. Silent outcomes remain in the store but do not generate a handoff line or check wake.
3. `supervision-host: cycle boundary ...` means the host ended its park at its bound: run `bin/fm-wake-drain.sh`, handle whatever it presents, run its printed acknowledgement (an empty queue prints `--ack-through 0`), and end the turn; the next park starts at that turn end.
4. A guarded command that exits 6 naming the branch actor's lease means the supervision session is handling that task right now: leave the lease alone and retry after it releases, which it does when its turn ends.
5. Captain outcomes the away session records stay in the outcome store until the return drain presents them; the return brief (`bin/fm-afk-return.sh`) counts them and points you to the `BRANCH OUTCOMES` section for processing and acknowledgement ([Captain outcomes](../supervision-host.md#captain-outcomes)).
6. `/afk` writes only the record here (`bin/fm-afk-launch.sh start-native` refuses the away daemon on this home); for `/quiet`, follow the [quiet skill](../../.agents/skills/quiet/SKILL.md).


================================================================================
READ-ONCE CONTRACT
================================================================================
Everything below is printed in full for this session start: every state/*.meta,
a compact data/backlog.md listing, a bounded tail of every state/*.status,
data/projects.md, data/secondmates.md, data/captain.md, data/captain-shared.md,
and data/learnings.md.
Do NOT re-read any of them after reading this digest, and do NOT bulk-read
data/backlog.md or state/*.status: re-reading everything defeats the entire
point of this command.

Go to a source directly only when:
  - this digest flagged it ABSENT (then rebuild or create it per AGENTS.md),
  - its contents looked unparseable or corrupt,
  - an individual full status log is needed for older wake-event history, or a
    status line was capped and its tail matters (each task's full log path is
    printed with its tail),
  - a full task body is needed (bin/fm-tasks-axi.sh show <id> --full, or data/backlog.md),
  - the backlog listing disclosed omitted queued items and this turn needs them,
  - the NETWORK CHECKS section reported its checks still IN PROGRESS and this
    turn needs their verdict (bin/fm-startup-network.sh report),
  - or a STARTUP TRUNCATED banner named the stage that would have printed it, in
    which case that stage's sources were never emitted and must be reconciled.

================================================================================
FLEET STATE
================================================================================

data/backlog.md
--------------------------------------------------------------------------------
ABSENT

Work under way (state/*.meta)
--------------------------------------------------------------------------------
(none)

Orphan status logs (state/*.status without matching .meta)
--------------------------------------------------------------------------------
(none)

AFK
--------------------------------------------------------------------------------
absent

================================================================================
NETWORK CHECKS
================================================================================
completed off the startup path in 2s: GitHub authentication, dead-secondmate relaunch, secondmate convergence, pending handoff delivery, project clone refresh with its drift reporting, and inactive terminal-outcome reconciliation.
(silent - no problems found)

================================================================================
CONTEXT
================================================================================

data/projects.md
--------------------------------------------------------------------------------
ABSENT

data/secondmates.md
--------------------------------------------------------------------------------
ABSENT

data/captain.md
--------------------------------------------------------------------------------
ABSENT

data/captain-shared.md (shared, main-authoritative, read-only in secondmate homes)
--------------------------------------------------------------------------------
ABSENT

data/learnings.md
--------------------------------------------------------------------------------
ABSENT

================================================================================
NEXT STEP
================================================================================
Follow the supervision operating instructions block above for harness 'claude'.
This script never starts supervision itself.

The digest above is complete for this session start. The READ-ONCE CONTRACT
section near the top of it governs what may still be read from disk.
Evidence: Digest truncated at 5s bound in wake-queue stage (summary still published at 05:47:58)

Source: Digest truncated at 5s bound in wake-queue stage (summary still published at 05:47:58)


================================================================================
SESSION START - /tmp/fm-lab.TWFKOi
================================================================================

LOCK
--------------------------------------------------------------------------------
lock acquired: harness pid 3748038

BOOTSTRAP
--------------------------------------------------------------------------------
BOOTSTRAP_INFO: lavish-axi >=0.1.80 enables confirmed board replies; this older compatible version retains the legacy reply path, but upgrade to prevent handing back a board before its reply is accepted

WAKE QUEUE
--------------------------------------------------------------------------------

●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
●  STARTUP TRUNCATED - SESSION START HIT ITS 5s RUNTIME BOUND
●  It stopped during the "wake-queue" stage, so everything above is COMPLETE
●  only up to that point.
●  RECONCILE these stages before acting on anything they would have shown:
●    wake-queue supervision-instructions read-once fleet-state network-checks context next-step
●  Rerun bin/fm-session-start.sh now to finish taking the helm. If it truncates
●  again, raise FM_SESSION_START_TIMEOUT and report the slow stage - a stage that
●  cannot finish inside the bound is a fleet problem, not a reporting detail.
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Evidence: Digest truncated at 3s bound in lock stage (before refresh launch point)

Source: Digest truncated at 3s bound in lock stage (before refresh launch point)


================================================================================
SESSION START - /tmp/fm-lab.TWFKOi
================================================================================

LOCK
--------------------------------------------------------------------------------
lock acquired: harness pid 3748038

●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
●  STARTUP TRUNCATED - SESSION START HIT ITS 3s RUNTIME BOUND
●  It stopped during the "lock" stage, so everything above is COMPLETE
●  only up to that point.
●  RECONCILE these stages before acting on anything they would have shown:
●    lock bootstrap wake-queue supervision-instructions read-once fleet-state network-checks context next-step
●  Rerun bin/fm-session-start.sh now to finish taking the helm. If it truncates
●  again, raise FM_SESSION_START_TIMEOUT and report the slow stage - a stage that
●  cannot finish inside the bound is a fleet problem, not a reporting detail.
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Evidence: Digest with FM_BACKEND_HERDR_CLI_TIMEOUT=2.5 (no integer errors)

Source: Digest with FM_BACKEND_HERDR_CLI_TIMEOUT=2.5 (no integer errors)


================================================================================
SESSION START - /tmp/fm-lab.TWFKOi
================================================================================

LOCK
--------------------------------------------------------------------------------
lock acquired: harness pid 3748038

BOOTSTRAP
--------------------------------------------------------------------------------
BOOTSTRAP_INFO: lavish-axi >=0.1.80 enables confirmed board replies; this older compatible version retains the legacy reply path, but upgrade to prevent handing back a board before its reply is accepted

WAKE QUEUE
--------------------------------------------------------------------------------
(no queued wakes)
================================================================================
SUPERVISION OPERATING INSTRUCTIONS - primary harness: claude
================================================================================
Current state:
- Lock: held by this session; this session owns normal supervision unless away mode says otherwise.
- Away/quiet mode: inactive.
- X mode: inactive; use the default watcher cadence.
- Supervision host: on; it takes away-posture wakes and, where the dialog mirror is verified, eligible attended wakes itself, and hands the rest to you (protocol at the end of this block).
- Ordinary wake: the Stop-owned auto-arm (bin/fm-claude-stop-autoarm.sh) already owns watcher continuity; drain and handle the wake, and do not arm another cycle yourself.

Mode: Claude Stop-hook-owned supervision.

When this session owns supervision and away mode is not active:
1. Drain first with `bin/fm-wake-drain.sh`.
   After handling all emitted wakes and reconciling open decisions and unread status lines, run the exact `--ack-through` command printed as `WAKE_ACK_REQUIRED`; until then the work remains durable for idempotent re-handling after interruption.
2. Routine watcher arm and re-arm are owned by the Stop `asyncRewake` hook (`bin/fm-claude-stop-autoarm.sh`), never by you.
   Every turn end while supervision is needed launches or attaches one home-scoped watcher cycle with no model command and no model tokens.
   An actionable close wakes you through the hook's exit-2 rewake, delivered as a `Stop hook feedback` message.
3. On a `Stop hook feedback` wake (`signal:`, `stale:`, `check:`, or `heartbeat`), run `bin/fm-wake-drain.sh` first and handle the wake.
   Do not run `bin/fm-watch-arm.sh` after an ordinary wake; the next turn end re-arms automatically when supervision is still needed.
   Do not invent a wake from an attach-status line alone; drain and act only on real wake records, the drain's `OPEN DECISIONS` and `UNREAD STATUS` entries, or a real watcher reason line.
4. On the one `Stop hook feedback` automatic-mechanism failure notice (`firstmate watcher auto-arm FAILED ...`), drain, inspect the automatic mechanism failure, and do not turn the notice into a repeating manual-arm loop.
5. If the Stop hook does not claim the home or reports an exhausted failure, inspect its registration and watcher startup path before ending blind.
   Keep the Stop-owned automatic mechanism as the only Claude arm owner.
6. Treat `watcher: started ...` and `watcher: attached ...` inside automatic arm output as proof that one live cycle exists.
   On attach, the arm follows verified identity-matched successors instead of exiting when the first cycle ends.
7. The durable wake queue preserves actionable events between a rewake and the next Stop-launched arm, while the bounded turn-end guard prevents a blind Stop when recovery did not start.
   No PreToolUse hook denies fleet commands based on watcher status.
   [`watcher-continuity.md`](../watcher-continuity.md) owns the exact session-lock recovery boundary.
8. The turn-end guard (`bin/fm-turnend-guard.sh --claude`) remains the final backstop.
   It requires the PID-strict live-watcher and fresh-beacon predicate at the Stop boundary, except for the Claude-specific foreign-live-owner safe exit owned by [`turnend-guard.md`](../turnend-guard.md#guard-predicates); that document also owns the distinct model-aware mid-turn pull-guard rules.
   Otherwise, it allows the stop when a watcher is healthy or an open auto-arm generation claim owns recovery, while fresh failure epochs advance the bounded one-time attended fail-open progression described there.
9. Waiting on the hook-owned cycle is silent: do not send idle progress while the watcher is parked.

The watcher itself remains `bin/fm-watch.sh`, and `bin/fm-watch-arm.sh` remains the verified arm wrapper that the Stop hook foregrounds on a home that opted out of the [supervision host](../supervision-host.md) (`config/supervision-host` holding `off`).
Re-arm attaches to an existing healthy cycle when one is already present and follows its verified successor chain.
See [`watcher-continuity.md`](../watcher-continuity.md) for the arm-layer successor and clean-close failure contract and the Claude ownership model.

Supervision host: on for this home (`config/supervision-host` holding `off` turns it off; [`supervision-host.md`](../supervision-host.md) owns the design).
The Stop hook runs the supervision host in the arm's place, and everything above still holds with these additions:
1. Attended (no away record, including a quiet-mode record; see [Postures](../supervision-host.md#postures)): a headless supervision session takes the wakes the supervision branch may take and never wakes you for a routine outcome, so fewer wakes reach you; check wakes, decision wakes, and whatever it cannot take still reach you exactly as above.
   `supervision-host: branch-outcome: ...` means it handled a wake and recorded captain outcomes for you: run `bin/fm-wake-drain.sh`, process each entry of its `BRANCH OUTCOMES` section as firstmate from the task's current state, because each entry says how long ago it was recorded (tell the captain, land or merge what is ready, answer or escalate a decision, or act on a blocker; your reply covers only entries still open, as if a settled one, such as a PR since merged, had never been listed), then run the `mark-processed` acknowledgement it prints; every drain presents them again until you do.
   `supervision-host: the supervision session could not take this wake ...` means the wake is yours: handle it as above.
   A failing turn may include a `supervision-host:` health note about repeated engine errors: tell the captain when it matters and handle the handed-back wake as usual; during cooldown later attended closes reach you unchanged.
   Routine outcomes never wake you; your next drain lists only visible routine outcomes under `BRANCH OUTCOMES, ROUTINE` for awareness, with nothing to acknowledge. Silent rows do not appear there, but remain available through `bin/fm-branch-outcome.sh list`.
2. Away (an away record exists and no daemon runs): the host hands each wake to a headless away session that runs the supervision branch's contract under the record, and you are parked.
   Only a wake the host hands back reaches you, as `Stop hook feedback` carrying the close plus one `supervision-host: <why>` line.
   That wake is automatic supervision, not the captain's return: drain and handle it under the away posture, and never run the return from it.
   After the return, a `supervision-host:` line naming the captain's return during a turn means that turn has visible outcomes missing from the return brief, whether the wake was handled or handed back: relay every following `supervision-host: outcome ...` line to the captain (the rows also remain in `bin/fm-branch-outcome.sh list`), then drain and handle any queued wake before acknowledging.
   Each such visible outcome is also a queued `check: supervision-host outcome <n> ... was recorded after the captain returned` wake, which the drain presents until acknowledged: relay each outcome once, whichever arrives first, and acknowledge its `BRANCH OUTCOMES` entry too when it has one. Silent outcomes remain in the store but do not generate a handoff line or check wake.
3. `supervision-host: cycle boundary ...` means the host ended its park at its bound: run `bin/fm-wake-drain.sh`, handle whatever it presents, run its printed acknowledgement (an empty queue prints `--ack-through 0`), and end the turn; the next park starts at that turn end.
4. A guarded command that exits 6 naming the branch actor's lease means the supervision session is handling that task right now: leave the lease alone and retry after it releases, which it does when its turn ends.
5. Captain outcomes the away session records stay in the outcome store until the return drain presents them; the return brief (`bin/fm-afk-return.sh`) counts them and points you to the `BRANCH OUTCOMES` section for processing and acknowledgement ([Captain outcomes](../supervision-host.md#captain-outcomes)).
6. `/afk` writes only the record here (`bin/fm-afk-launch.sh start-native` refuses the away daemon on this home); for `/quiet`, follow the [quiet skill](../../.agents/skills/quiet/SKILL.md).


================================================================================
READ-ONCE CONTRACT
================================================================================
Everything below is printed in full for this session start: every state/*.meta,
a compact data/backlog.md listing, a bounded tail of every state/*.status,
data/projects.md, data/secondmates.md, data/captain.md, data/captain-shared.md,
and data/learnings.md.
Do NOT re-read any of them after reading this digest, and do NOT bulk-read
data/backlog.md or state/*.status: re-reading everything defeats the entire
point of this command.

Go to a source directly only when:
  - this digest flagged it ABSENT (then rebuild or create it per AGENTS.md),
  - its contents looked unparseable or corrupt,
  - an individual full status log is needed for older wake-event history, or a
    status line was capped and its tail matters (each task's full log path is
    printed with its tail),
  - a full task body is needed (bin/fm-tasks-axi.sh show <id> --full, or data/backlog.md),
  - the backlog listing disclosed omitted queued items and this turn needs them,
  - the NETWORK CHECKS section reported its checks still IN PROGRESS and this
    turn needs their verdict (bin/fm-startup-network.sh report),
  - or a STARTUP TRUNCATED banner named the stage that would have printed it, in
    which case that stage's sources were never emitted and must be reconciled.

================================================================================
FLEET STATE
================================================================================

data/backlog.md
--------------------------------------------------------------------------------
ABSENT

Work under way (state/*.meta)
--------------------------------------------------------------------------------
(none)

Orphan status logs (state/*.status without matching .meta)
--------------------------------------------------------------------------------
(none)

AFK
--------------------------------------------------------------------------------
absent

================================================================================
NETWORK CHECKS
================================================================================
completed off the startup path in 2s: GitHub authentication, dead-secondmate relaunch, secondmate convergence, pending handoff delivery, project clone refresh with its drift reporting, and inactive terminal-outcome reconciliation.
(silent - no problems found)

================================================================================
CONTEXT
================================================================================

data/projects.md
--------------------------------------------------------------------------------
ABSENT

data/secondmates.md
--------------------------------------------------------------------------------
ABSENT

data/captain.md
--------------------------------------------------------------------------------
ABSENT

data/captain-shared.md (shared, main-authoritative, read-only in secondmate homes)
--------------------------------------------------------------------------------
ABSENT

data/learnings.md
--------------------------------------------------------------------------------
ABSENT

================================================================================
NEXT STEP
================================================================================
Follow the supervision operating instructions block above for harness 'claude'.
This script never starts supervision itself.

The digest above is complete for this session start. The READ-ONCE CONTRACT
section near the top of it governs what may still be read from disk.
Evidence: Targeted test summary (rc per file)

Source: Targeted test summary (rc per file)

fm-session-start rc=0 225s
fm-afk-return rc=0 62s
fm-wake-drain-outcome-backstop rc=0 41s
fm-wake-drain-open-decisions-cursor rc=0 84s
fm-host-mirror rc=0 24s
fm-task-inbox rc=0 31s
fm-backlog-read-bound rc=0 24s
Evidence: Live summary results
full run: rc=0 elapsed=8.1s, home-summary.json schema=fm-secondmate-home-summary.v1
T=5: TRUNCATED during wake-queue stage -> home-summary.json published 05:47:58
herdr cap 2.5/abc/30: rc=0, integer-errors=0, no truncation banner
7/7 targeted test files rc=0

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

⚠️ **Rebase** - 1 warning

Confirm these commits belong in this PR before approving, or manually separate the intended work onto origin/main before gating.

🔧 No changes applied.
1 warning still open:

Confirm these commits belong in this PR before approving, or manually separate the intended work onto origin/main before gating.

no changes applied: bundled local-default commits require manual separation or explicit approval; the rebase conflict resolver cannot safely select commits to discard.

⚠️ **Review** - 1 info
  • ℹ️ bin/fm-session-start.sh:727 - The home-summary refresh is now detached, so it runs at the same time as fm-bootstrap.sh's detect_home_summary_publication (bin/fm-bootstrap.sh:1596) instead of finishing before it. The case: a home has at least FM_HOME_SUMMARY_FAILURE_REPORT (2) recorded failures since its last publication. Bootstrap can now print the 'HOME_SUMMARY: ... has not been republished' diagnostic on the same start whose in-flight refresh is about to succeed. Before this change, the synchronous refresh would already have republished, so the diagnostic stayed quiet. The effect is one stale-by-seconds operator diagnostic, which clears on the next start. This is an accepted consequence of the user-approved early detached launch; no action needed.
✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 3 of 6 scenarios driven live against the product
Scenario Result Live Evidence
Operator starts a session on a clean home: the digest completes and the detached refresh publishes home-summary.json ✅ pass live live-session-start-digest.txt; home-summary.json appeared with schema fm-secondmate-home-summary.v1 and home=$LAB
Adversarial: the digest hits FM_SESSION_START_TIMEOUT during the wake-queue stage and the home summary still gets republished ✅ pass live live-session-start-truncated-5s.txt shows 'STARTUP TRUNCATED ... wake-queue' stage; home-summary.json recreated after being deleted
Adversarial: a non-integer or oversized FM_BACKEND_HERDR_CLI_TIMEOUT causes no shell integer errors and does not break the digest ✅ pass live live-session-start-herdr-cap-{2.5,abc,30}.txt: rc=0, 0 'integer expression' lines, no truncation banner, NEXT STEP present
A hung herdr endpoint read is bounded per read and leaves no herdr processes behind ⏸️ untested no The prior payload did not establish a live result. Only the scripted tests/fm-session-start.test.sh with a fake hanging herdr fixture was run (rc=0). A live check needs a herdr lab session via bin/fm-…
Away-mode return brief, wake-drain backstop and decision cursors, host-mirror cursor and task-inbox ring-state still parse tab-separated rows after the IFS=$'\t' change ⏸️ untested no The prior payload did not establish a live result. Only the scripted end-to-end tests were run (fm-afk-return, fm-wake-drain-outcome-backstop, fm-wake-drain-open-decisions-cursor, fm-host-mirror, fm-t…
The session-start digest still bounds a hanging backlog read, with the git identity pinned in the fixture ⏸️ untested no The prior payload did not establish a live result. Only the scripted tests/fm-backlog-read-bound.test.sh with a fake hanging tasks-axi was run (rc=0). A live check needs a real hanging backlog backend…
  • bin/fm-lab-home.sh create $LAB (disposable marked lab home, removed afterwards with rm -rf $LAB)
  • FM_HOME=$LAB bin/fm-session-start.sh (full live digest, rc=0, ~8s), then waited for $LAB/state/home-summary.json and checked its schema/home with jq
  • FM_HOME=$LAB FM_SESSION_START_TIMEOUT=1|3|5 bin/fm-session-start.sh after deleting home-summary.json (forced truncation)
  • FM_HOME=$LAB FM_BACKEND_HERDR_CLI_TIMEOUT=2.5|abc|30 bin/fm-session-start.sh (grep for 'integer expression' errors and a truncation banner)
  • bash tests/fm-session-start.test.sh (rc=0, 225s)
  • bash tests/fm-afk-return.test.sh (rc=0, 62s)
  • bash tests/fm-wake-drain-outcome-backstop.test.sh (rc=0, 41s)
  • bash tests/fm-wake-drain-open-decisions-cursor.test.sh (rc=0, 84s)
  • bash tests/fm-host-mirror.test.sh (rc=0, 24s)
  • bash tests/fm-task-inbox.test.sh (rc=0, 31s)
  • bash tests/fm-backlog-read-bound.test.sh (rc=0, 24s)
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

keenvc and others added 30 commits September 16, 2026 16:15
…spatch (#1)

* feat(harness): add the cline crewmate/scout adapter with ClinePass dispatch

Add Cline CLI 3.0.62 as a verified crewmate/scout harness following the agy
pattern: ancestry detection on the native .cline process, a launch-then-send
TUI launch, per-task .cline/hooks busy/turn-end wiring under a new cline-hook
busy source, Escape interrupt, /exit, and control-plane tables. Wire the
ClinePass open-weights pool into the crew-dispatch example and document the
known composer-empty gap (placeholder luminance above the shared ghost
ceiling) with a tmux live guard as the refresh command.

* no-mistakes(review): fix(docs,quota): correct cline resume grouping and cline-pass family id

* no-mistakes(document): docs: cover cline in tmux liveness/anchoring list and configuration.md secondmate-refusal note

* no-mistakes(lint): {"summary": "lint: no code changes needed, fm-lint.sh passes with shellcheck on PATH"}

* chore(gitignore): drop the stray .omc handoff artifact and ignore .omc/

The no-mistakes gate agent's Claude Code oh-my-claudecode plugin writes
.omc/handoffs/last-session-end.md into the run worktree at session end, and a
later pipeline step committed it into this branch. Remove the committed file and
ignore .omc/ so a home-environment handoff artifact can never ride into a PR.

---------

Co-authored-by: firstmate-worker <worker@local>
Two Claude subscriptions need to run concurrently across lanes without
moving every claude spawn onto one account. --claude-config-dir picks the
CLAUDE_CONFIG_DIR one claude spawn's pane resolves into: validated before
any worktree or endpoint exists, recorded in the task's own meta, reused
unchanged on --relaunch, and threaded through both the pre-launch trust
registration and the launch's own environment so the two halves can never
land in different stores. A spawn naming no seat is byte-identical to
before.
Count live crewmate, scout, and local secondmate lanes grouped by the
billing provider a candidate actually draws on, expose the load for
dispatch intake, and refuse a spawn that would push a provider past its
configured cap.

Provider identity comes from the resolved model string, not the harness
name: a provider-qualified prefix or a model-id pattern decides the pool,
and the harness table is only the fallback. That keeps two models on one
pool counting together while a different pool stays separate. The mapping
lives once in bin/fm-provider-lib.sh, whose fallback reuses the existing
quota tables rather than restating them.

A lane occupies a seat unless its recorded endpoint is provably dead or
missing, so a cleared seat is visible before the next dispatch. The cap
comes from providerCaps in config/crew-dispatch.json (per provider, else
default), falling back to 4; bootstrap now rejects a malformed
providerCaps instead of silently ignoring it.

bin/fm-provider-load.sh prints the current per-provider used/cap for
intake. Stranded-record detection stays with fm-lane-account-dead-records;
this counter reads the current endpoint classifier.
Deliver bin/fm-hold-reverify.sh, an armed watcher check that re-checks each
captain hold past an age threshold against shipped reality and reports it as
dead, still_live, not_a_decision, or unestablishable - the reconciliation
vocabulary captain-hold-lifecycle already owns.

It reports only: it never calls answer and never closes or annotates a call,
so only the captain's own words or an explicit evidence-backed reconciliation
can resolve one. Dead is never inferred from absence or an unreadable source.

Aged holds come from the canonical local backlog projection
(fm-fleet-snapshot.sh --contribution-input); recorded pull requests are read
through fm-pr-lib.sh. Each sweep writes a docket and prints one line only when
the finding set changes, with a report record keyed on that set.
…rget

A firstmate home had no way to see that another home was already working a
shared external target, so the main home and a secondmate could both arm to
land the same PR with nothing to stop a double merge.

bin/fm-claim.sh records, releases, and inspects a work claim on a shared
external target - a PR, an issue id, or a declared file area. The store is a
machine-wide directory (FM_CLAIM_ROOT, default
${XDG_STATE_HOME:-$HOME/.local/state}/firstmate/claims), the sibling of the
existing process-event source claim root, because one owner per canonical
target cannot live inside a single home. Local homes share one filesystem; a
remote secondmate is a separate host and stays outside the mechanism.

Acquire is atomic and fails closed: a second home's live claim refuses rather
than racing. A claim is released on cleanup or reclaimed only when its holder
is provably gone (its home directory is absent, or its task record is absent
past FM_CLAIM_PENDING_GRACE). Any uncertainty keeps the claim.

bin/fm-spawn.sh --claim records the claim before any endpoint or task record
exists and refuses the spawn on conflict, recording the canonical keys on the
task as claims=; fm-teardown.sh releases them on cleanup. The flag is refused
on --secondmate, --relaunch, and a batch dispatch.

Tests: tests/fm-claim.test.sh drives the real CLI across two simulated homes
sharing one claim root, plus a real spawn that records its claim and a second
dispatch that is refused.
Merge authority was decided once at intake and never revisited, so a task
dispatched yolo=on kept that authority even after firstmate held it for the
captain, and the recorded authority and the merge path could disagree.

Fold the captain-hold predicate into fm_merge_authority_resolve, the single
owner of a task's standing merge authority, so a held task resolves to
captain-hold (or hold-unreadable) whatever its yolo posture or away grants say.
Remove bin/fm-pr-merge.sh's duplicate require_released_captain_hold and fold its
refusal into the shared gate, and have bin/fm-merge-local.sh share the same
predicate instead of repeating it.
A worker parked on a provider quota wall kept a live, painting harness
while its turn could not advance, so every one of them read as working
from its semantic busy record. The measured fleet incident had all of
one provider's workers stalled at the same weekly limit while
supervision saw a healthy fleet.

Recognize the wall from the pane text the busy reader already inspects.
The signal is built from two independent rendered families - a
wall-shaped limit phrase and a scheduled retry/reset phrase - within the
last few non-empty lines, so no single vendor string is load-bearing and
ordinary worker prose does not match. A busy verdict over that wall
reports `quota` instead of busy.

fm-crew-state.sh surfaces it as its own `state: quota` rather than
collapsing it into working or a declared pause, because a quota-killed
worker cannot be relaunched in place; the recovery skill now states that
preserve-and-replace under a new id is the path.

The portable regression pins the logic and its divergence cases over
synthetic transcripts. The live guard drives the real installed OpenCode
TUI against a local 429 stub so its own retry modal renders with no
model tokens spent, and proves the same task reads working before the
wall and quota after it.
Teardown refused any record whose endpoint was already cleared, so a lane
could never be retired once its window was gone and it kept occupying an
in-flight row. Accept an explicit endpoint_cleared stamp as stronger
agent-less evidence than a dead window, with no flag and no --force, while
keeping the unlanded-work refusal unchanged.

A projected Herdr teardown confirmed only the task pane was gone, so a
workspace whose recorded pane vanished before its close survived for a
restart to restore as a live agent in the primary clone. Remove the
workspace's remaining panes through the same focus-preserving pane close and
require the workspace gone before retiring the journal.

A dead pane whose display redrew re-alarmed on every new hash, a supervision
tax that grew with each dead lane. Absorb a redrawn dead display against the
existing once-record, while a relaunched agent re-arms the incarnation and
its own death still reports in full.
keenvc and others added 6 commits September 20, 2026 05:38
# Conflicts:
#	.agents/skills/harness-adapters/SKILL.md
#	AGENTS.md
#	bin/fm-agent-process-lib.sh
#	bin/fm-bootstrap.sh
#	bin/fm-busy-lib.sh
#	bin/fm-composer-lib.sh
#	bin/fm-control-lib.sh
#	bin/fm-harness.sh
#	bin/fm-spawn.sh
#	docs/agent-control.md
#	docs/architecture.md
#	docs/configuration.md
#	docs/trace-context.md
#	docs/verification/runtime-backends.md
#	tests/fm-control.test.sh
# Conflicts:
#	.gitignore
#	bin/fm-spawn.sh
* fix(bin): read aged holds via backlog-json, not contribution-input

The hold re-verify sweep only needs canonical backlog rows; contribution-input
also walks every task meta for merge-authority resolution, which took ~96s at
this fleet size and always exceeded the five-second projection bound. Add
fm-fleet-snapshot.sh --backlog-json for that narrower read and point the sweep
at it so failures stay loud without raising the timeout.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(bin): apply review decisions for aged hold re-verify sweep

Sort aged captain holds oldest-first before the per-sweep cap, let
FM_HOLD_REVERIFY_BUDGET_SECS govern the backlog projection bound, clamp
forge probes to remaining budget, skip probes for predetermined
not-a-decision rows, drop the classify subcommand, and document
--backlog-json.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(bin): use backlog_json output mode label for shellcheck

Co-authored-by: Cursor <cursoragent@cursor.com>

* docs: fold pipeline document-step hold-reverify pointers into fork tip

Adds the toolbelt row and prose alignment from the failed run's document
step without rebasing onto upstream main.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(document): correct hold-reverify docket contents description

* chore: gitignore .omc session state and document in AGENTS.md

Apply captain inbox 006 on the fork publication branch without rebasing onto upstream main.

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
* ci: expect 19 snapshot/fleet-view tests under stock macOS Bash

The fork's tests/fm-fleet-snapshot-view.test.sh carries
test_large_payloads_compose_through_files, the regression for
bin/fm-fleet-snapshot.sh composing payloads above 128KB through files
instead of argv. It landed in d9356ca together with that snapshot
change and passes under /bin/bash 3.2 on the macOS runner, which
already counted 19 ok lines, so the hard-coded 18 in the macOS job was
the only thing left behind.

* docs: declare the cline-pass provider on the documented cline profiles

The resolver refuses docs/examples/crew-dispatch.json with "profiles
whose harness lacks one authoritative provider family require
provider: cline", so the documented-example check in
tests/fm-dispatch-resolve.test.sh has failed since the cline profiles
were added to the example.

The example is the wrong side. The resolver's provider is the quota-axi
provider family it ranks candidates by, not the launch prefix cline
reads from its model id, and the single-provider table in
docs/configuration.md leaves cline out on purpose: the model prefix,
not the harness, decides who bills the run, exactly as for pi and omp.
The Pi default in the same example already declares provider: claude
for that reason. Add provider: cline-pass to the four cline profiles,
correct the one sentence in docs/configuration.md that claimed no field
was needed, and give the test's canned Choice answer the fourth rule
the example now has, since the resolver checks the answer against the
full option set.

* test: corrupt the claim record under test, not the first one find returns

The corrupt-record check picked its victim with find | head -n 1 while
three claims exist, so on a filesystem whose directory order differs
from the author's it corrupted o/r#7 or repos/example#9 and then asked
about owner/repo#11, whose intact record answered "held" with exit 0.
CI's stdout showed exactly that line. Select the record by its
documented key= line instead. The guard itself is intact: corrupting
the right record returns exit 5 on the same inputs.

* test: give the restart watchers a refresh bound a slow runner can meet

In the watcher-restart section of tests/fm-home-summary-refresh.test.sh
the watcher runs with FM_HOME_SUMMARY_INTERVAL=999999, but age_of
reports 999999 for a missing ledger, so its detached refresh fires on
every one-second poll. After the lock holder is killed that refresh
steals the dead lock ahead of the test's idle-only refresh, which then
returns without publishing, and with the section's
FM_HOME_SUMMARY_TIMEOUT=2 a slow runner kills the watcher's attempt
before it publishes. The next poll dies the same way and the ledger
never appears, which is the "a dead publication lock wedged
publication" failure in CI run 35683307194.

Raise the three restart watchers' bound to 30 seconds, the deadline
this file already uses for its accumulated-home publication. Under a
30% CPU quota the unchanged test fails on exactly that line and the
changed one passes all 21 checks. The idle-only refresh, the dead-lock
reclamation and the 10-second wait are untouched.
…wner (adopt kunchenguid#5993) (#7)

* Fix reassigned teardown slot collisions

(cherry picked from commit 115333b)

* no-mistakes(document): Document claim-first pool-slot teardown behavior

(cherry picked from commit dcbe51a)

* no-mistakes(document): Confirm teardown documentation reflects slot ownership behavior

(cherry picked from commit bdcedcf)

* fix(teardown): avoid presentation-lock races after claim-first slot cleanup

Reassigned-slot teardown with a dead Herdr husk no longer takes the shared
presentation session lock, restoring the pre-5993 contention profile for stale
records while keeping live-slot protection intact. Herdr presentation recovery
spawns now wait up to 120s for that lock and release it on abort so concurrent
cross-home recovery cannot fail the 5s try loop after a legitimate holder.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: QIanGua <15757826110@163.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
…in config (#9)

* Fix OpenCode v2 worker launch to use --standalone and config model.

OpenCode 2.x removed the interactive --model flag; carry the resolved model in OPENCODE_CONFIG_CONTENT and launch with --standalone so the model and permission block are honored off the shared service.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(bin): honor opencode v2 top-level model and gate --standalone

Always write the resolved model as OPENCODE_CONFIG_CONTENT top-level model on 2.x, drop unverified agent.build variant JSON, and keep the 1.x --model launch shape when opencode --version reports major 1.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test: stub opencode --version in shared spawn fakebin

Spawn tests prepend a fakebin to PATH; fm-spawn now probes opencode
--version for the v1/v2 launch gate, so every spawn fakebin must answer it.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
@greptile-apps

greptile-apps Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 0/5

[Medium risk] Adds support for new agent harnesses and updates session startup logic.

The PR is not safe to merge while the earlier claim, hold, provider-cap, and merge-authority defects remain.

Reviews (9) · Last reviewed commit: "no-mistakes(review): Scope herdr cap to ..."

Comment thread bin/fm-claim.sh
Comment on lines +302 to +304
fm_claim_read "$path" || continue
if [ "$FM_CLAIM_HOME" = "$HOME" ] && [ "$FM_CLAIM_TASK" = "$TASK" ]; then
if rm -f -- "$path"; then

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Cleanup can erase another claim
After teardown removes a task record, another home can reclaim its now-stale claim between this ownership check and rm. Cleanup then deletes the new home's live claim, allowing a third home to dispatch against the same target. The single-target release command has the same read-before-lock gap. Both paths need to check ownership while holding the target mutex.

Comment thread bin/fm-claim-lib.sh
Comment on lines +322 to +331
[ -n "$owner" ] && [ -n "$repo" ] && [ -n "$num" ] || return 1
case "$num" in
*[!0-9]*) return 1 ;;
esac
case "$owner$repo" in
*[!A-Za-z0-9._-]*) return 1 ;;
esac
owner=${owner,,}
repo=${repo,,}
printf '%s:%s/%s/%s#%s\n' "$kind" "$host" "$owner" "$repo" "$num"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Equivalent PR numbers get separate claims
owner/repo#01 passes the digit check but keeps its leading zero, producing a different claim file from owner/repo#1. Two homes can therefore both acquire and dispatch against the same PR or issue, breaking the cross-home exclusivity guarantee.

Comment thread bin/fm-hold-reverify.sh
Comment on lines +416 to +423
while IFS= read -r hold; do
[ -n "$hold" ] || continue
if [ "$EXAMINED" -ge "$MAX_HOLDS" ] || budget_exhausted; then
DEFERRED=$((DEFERRED + 1))
continue
fi
EXAMINED=$((EXAMINED + 1))
facts=$(gather_facts "$hold")

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Deferred holds never get checked
When more than FM_HOLD_REVERIFY_MAX_HOLDS aged holds exist, each sweep starts with the same oldest holds and skips the rest. Repeated scheduled checks never reach the deferred holds, so their changed or resolved decisions can remain unverified indefinitely. The sweep needs to advance through the deferred set.

Comment thread bin/fm-spawn.sh
Comment on lines +2614 to +2625
if [ -n "$HARNESS" ]; then
LANE_CAP_MODEL=$MODEL
LANE_CAP_EXCLUDE=
if [ "$RELAUNCH" -eq 1 ]; then
LANE_CAP_EXCLUDE=$ID
[ -n "$LANE_CAP_MODEL" ] || LANE_CAP_MODEL=$(fm_meta_get "$RELAUNCH_META" model)
fi
fm_provider_cap_refuse "$STATE" "$CONFIG" "$HARNESS" "$LANE_CAP_MODEL" "$LANE_CAP_EXCLUDE" || exit 1
fi
if [ "$HARNESS" = openhands ]; then
if [ -z "$MODEL" ] || [ "$MODEL" = default ]; then
if [ -n "${LLM_MODEL:-}" ]; then

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 OpenHands checks the wrong cap
If --model is omitted and LLM_MODEL selects a full provider such as fireworks_ai/..., this check runs before OpenHands copies LLM_MODEL into MODEL. It checks the openhands fallback bucket instead, then launches on the full provider, exceeding that provider's configured lane cap. Resolve the effective model before checking capacity.

Comment thread bin/fm-bootstrap.sh Outdated
@keenvc
keenvc force-pushed the fm/fm-slowness-scout branch from 4eb95cb to 450056a Compare September 30, 2026 01:23
Comment thread bin/fm-session-start.sh Outdated
@greptile-apps

This comment has been minimized.

@keenvc keenvc changed the title fix(bin): speed up session start and merge upstream firstmate main feat(bin): merge upstream reliability batch and speed up session start Sep 30, 2026
* feat(bin): defer the wedge escalation for a lane parked at a supervisor-owed gate (#4974)

* fix(watch): recheck a gate awaiting a human instead of wedge-escalating it

A lane whose validation run is parked at a gate waiting on a human
decision is correctly quiet, but nothing in its status line says so: the
evidence is the pipeline's own gate state rather than anything the worker
wrote. The wedge timer read that silence as a suspected wedge and climbed
the escalation ladder for as long as the wait lasted, and each escalation
cost a supervising turn. The landed declared-wait consult does not reach
it, because a live ordinary crewmate never reports a declared pause, and
raising FM_STALE_ESCALATE_SECS would delay genuine wedge detection for
every lane by the same amount.

The threshold now reads a second, independent record when the status line
accounts for nothing: whether the crew's current state is a gate whose
answer is owed by a human. That is minted only from the gate's own
findings table, by a row whose `action` column is exactly `ask-user`,
located by position out of the table header the way nm_gate_step_row
already reads its row - never searched for over the run payload, where a
finding's free-text description or a branch name satisfies a search just
as well. A gate awaiting the CREWMATE's own answer keeps the unchanged
escalation schedule, reason and demand-deep-inspection wording, because a
crewmate that goes quiet before answering its own gate is exactly the
wedge the ladder exists to catch.

Each kind of wait now carries the human it is on, the action that clears
it, and whether that human is the captain as data alongside the verdict,
rather than as wording chosen per branch where the recheck is written, so
the deferral cannot word one kind of wait as another and a new kind
cannot ship without deciding all of them. A parked gate has no written
record of when its wait began, so its recheck publishes no wait age at
all rather than one read from the quiet window this deferral resets on
every pass, which would report the same small number for a gate of any
age. Like every other captain-facing recheck here it is absorbed in
silence while the away-posture record exists, arming no throttle, so the
recheck is owed in full the moment the record is archived.

The consult runs only in the at-threshold branch that was about to
escalate, beside the worktree walk already there, and only for lanes
whose status line explained nothing.

Closes #3055

* no-mistakes(review): require an unanswered decision before deferring a parked gate

* no-mistakes(review): reset the away-silenced timer, fail-safe findings parse, US-joined wait records

* test(watch): pass the pane hash wedge_timer_check now takes

Upstream gave wedge_timer_check a sixth <pane-hash> argument for its
dead-record probe. The malformed-wait-record rounds drive the real function
directly, so they pass one, and stub fm_backend_agent_state to a live agent so
the probe that runs after a refused deferral keeps the unchanged ladder rather
than reading a backend the child shell has none of.

* no-mistakes(review): Bind parked-gate wait to its run, owe it firstmate

* no-mistakes(document): correct wait-kind count, crew-state reader scope, gate-key coupling

* feat(watch): make the parked-gate wait deferral opt-in

The wedge timer deferring a lane parked at a validation gate is new
supervision behaviour rather than a restored one, and it decides which
lanes give up the escalation ladder, so it now ships as a default-off
per-home option instead of changing every home on upgrade.

config/wedge-defer-parked-gate arms it. The flag is read before the
decision fold, so an unconfigured home spends no fold or current-state
read, writes no record, and keeps the unchanged escalation schedule,
reasons and demand-deep-inspection wording; a test counts the reader
calls in both directions to pin that.

It is not inherited by secondmate homes: each home supervises its own
crew and owns that trade separately, the same reason
config/turnend-churn-absorb is home-local.

The away-posture absorb returns to leaving the idle timer alone, which
it had restarted only because the costly consult could reach it. A
parked-gate wait is owed to the supervisor rather than the captain, so
it never enters that branch, and the recheck owed on return is again
owed in full the moment the record is archived.

* test(watch): pin that the away-silenced hold leaves the idle timer alone

The absorb no longer restarts the timer, so the recheck owed on return is
owed in full rather than a cadence into the return. Nothing asserted
that, so a restart could be reintroduced silently.

* no-mistakes(review): document away-silence rationale, pin captured gate component

* no-mistakes(test): anchor gate row scan to the braced findings header

* no-mistakes(document): pin same-block gate row invariant in crew-state comment

* fix(bin): reclaim a task whose herdr endpoint was destroyed (#5007)

* fix(control): let the owning seat reclaim a task whose endpoint is gone

A destroyed pane or workspace made `missing` a terminal state. Relaunch
accepted only `dead` and said to stop the agent first; exit refused
`missing` and said to reconcile the task first; there is no reconcile
verb. Each command named the other as its prerequisite, so a task whose
terminal went away could not be reclaimed by anything, and a no-mistakes
approval it was parked on had no seat left to answer it.

`missing` is agent-free a fortiori: there is no endpoint, so there is no
agent in it. Widen the existing guards rather than add a verb.

- fm-spawn --relaunch accepts a positively proven `missing` and creates
  one fresh endpoint in the recorded worktree; the record it already
  republishes rebinds the task to it. A `dead` endpoint is still adopted
  in place.
- fm-control exit reports `endpoint-gone` instead of dying, so the
  relaunch transaction's stop step no longer dead-ends, and re-resolves
  the endpoint from the record before verifying the replacement.

The duplicate-agent refusal is untouched: both verdicts come from the
same recovery-grade classifier, which claims `missing` only from positive
absence, so `alive`, `ambiguous`, and `unreadable` all still refuse. The
backends' own create paths refuse a live same-labeled endpoint as a
second independent guard. The worktree, its branch, commits, uncommitted
changes, armed poll and registration, record rows, and status log are all
untouched - a reclaim is a recovery, never a teardown.

A secondmate is excluded: its gone-endpoint recovery already has one
owner in the session-start liveness sweep, so relaunch refuses and names
it rather than becoming a second path to the same outcome.

Tests reproduce both halves of the deadlock, the reclaim succeeding,
unlanded work surviving it, and the refusals that still hold.

* no-mistakes(review): prove endpoint absence per backend before reclaim rebinds

* no-mistakes(review): give exit and relaunch one absence proof; pin herdr rebind session

* no-mistakes(review): narrow endpoint reclaim to herdr; tmux refuses honestly

* no-mistakes(review): stop refusals and docs asserting unestablished causes

* no-mistakes(review): stop herdr fixture helper losing tmp-root registration

* no-mistakes(review): document workspace drift and absence-probe server residue

* no-mistakes(review): correct rebind limitation to its one reachable case

* no-mistakes(review): stop claiming reclaim leaves instructions untouched

* no-mistakes(document): scope fm-control-lib purity claim, note reclaim coverage

* no-mistakes(rebase): read the staged launch file in the herdr fixture

Rebasing onto main picked up #4994, which stages a long worker launch
command into a script and delivers the short `. '<path>'` line instead of
the literal command. The tmux fake and tests/fixtures.sh were updated for
that; the herdr fake this branch adds was written before it and still
keyed "an agent now exists on this pane" off the literal
`encode launch-brief` text, so after the rebase it never marked the
rebound pane live and the reclaim's alive-wait read `dead`.

Dereference the staged file first, exactly as the tmux fake above does.
Test-fixture only; no production path changes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(document): note reclaim placement in herdr and scripts inventories

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(bin): stamp status events with their emission time (#3764)

* test(status): reproduce missing event emission time

* wip(status): preserve optional event emission time

* test(status): document indirect clock stub invocation

* no-mistakes(review): Preserve historical status bytes during reply recovery

* no-mistakes(test): Fix timestamped status assertions and remote fixture dependencies

* no-mistakes(review): Preserve captain regex overrides for timestamped status events

* no-mistakes(document): Clarify status event timing and publication contracts

* no-mistakes(lint): Quote literal done to satisfy ShellCheck

* no-mistakes(ci): Captain, updated .github/workflows/ci.yml to expect 19 snapshot tests instead of 18, matching the PR’s added regression. Reproduced the failure before the fix. Stock Bash 3.2.57 verification passed: parse sweep, 19 snapshot tests, 53 Bearings tests, and the public-followup regression. Workflow lint and diff checks passed

* no-mistakes(test): Preserve terminal notifications with malformed timestamp tags

* no-mistakes(test): Stamp Rovo spawn failures with emission time

* no-mistakes(document): Verify status event documentation

* no-mistakes(lint): Fix ShellCheck quoting in status emission-time tests

* no-mistakes(ci): Captain, fixed four lifecycle assertions to accept emission timestamps while preserving publication and retry checks. Reproduced the CI failure before the fix. The lifecycle suite now passes with six Beads capability skips; syntax, targeted ShellCheck, and diff checks passed

* no-mistakes(ci): Captain, fixed malformed timestamp colons hiding actionable events using shared normalization. Original bytes and unknown ages are preserved. Regression reproduced before the fix; classifier and remote-reply suites, targeted lint, syntax, and diff checks passed

* no-mistakes(review): Stamp remote escalations at call sites, drop new flag

* no-mistakes(review): Accept stamped escalation and close lines in test assertions

* no-mistakes(review): Restore reserved-key answered-note guard for stamped closes

* test(status): accept optional emission time in PR-provenance assertions

The #4148 provenance test landed on main with exact unstamped greps.
Parent-channel lines from this branch carry [at=<epoch>], so strip only
that tag before the same exact match. No production change.

* no-mistakes(review): Accept stamped ready signal in PR fallback scrape

* no-mistakes(review): Drop relay flag, stamp parent events at call sites

* no-mistakes(review): Stamp worker terminal-signal instructions, revert fm-on fixture

* no-mistakes(review): Accept optional stamp in live cmux drift guard

* no-mistakes(review): Restore original test invocation order in two suites

* no-mistakes(review): Strip only well-formed numeric status time tags

* no-mistakes(document): Drop stale unstamped PR-ready line spelling from channel doc

* no-mistakes(review): Stamp agy spawn-failure status lines with event time

* fix(bin): normalize status event times in-shell and freeze the budget test clock

Two paths made a status event's emission time cost more than it should.

The captain-relevance fallback piped every line through awk to drop a
well-formed `[at=<epoch>]` tag before matching, so a supervisor sweep paid a
fork per line just to prepare a regex match. Shell parameter expansion does the
same strip with no fork, and the retry-dedup scan now reuses that one helper
instead of carrying a second copy of the rule in awk. The copies had already
drifted: the shell side stripped tags from lines with no colon, which the awk
rule left whole, so a colonless line could be mistaken for one already
recorded. One definition, checked against the awk rule it replaces over the
edge cases and a 4000-line fuzz.

tests/fm-contributions.test.sh froze its fixture clock only in exhaust mode. In
hang mode the poll set DEADLINE to the real now plus a one-second budget, and
when the second ticked before the first forge call the loop broke without ever
calling gh: forge/calls was never written and the assertion failed reading a
missing file. Freezing the clock in both modes removes the dependence on wall
time; the bounded call is still cut by the real timeout, so the observation the
test asserts still starts.

Emission time stays optional on new status records, and legacy or malformed
lines keep an unknown age.

* no-mistakes(review): Stamp ask-user escalation line and fix Kimi status assertion

* no-mistakes(document): Drop stale unstamped done-line spelling from watcher docs

* test: fold emission-time snapshot coverage into the fixture case

Drop the incidental ci.yml 18-to-19 count hunk so the PR no longer
touches workflows. Keep every emission-time assertion by folding it
into test_fixture_snapshot_json.

* no-mistakes(review): replace brief date substitution with epoch placeholder; drop emitted_at_epoch

* no-mistakes(review): align untimed normalizer with epoch parser; tolerate placeholder stamp in PR scrape

* no-mistakes(review): strip undelimited at-tags; correct brief stamp header

* no-mistakes(review): normalize stamps at both captain-regex sites; restore mtime freshness

* no-mistakes(review): strip colon-bearing stamps for relevance; fix headers and test oracles

* no-mistakes(review): narrow escalation match to stamp tolerance; pin note verb

* no-mistakes(review): read note and key past colon-bearing stamps

* test(status): keep inactive reconcile assertions stamp-tolerant

These two oracles were made stamp-tolerant while resolving one of the
branch's merges from main. The rebase drops merge commits, so that
adaptation was lost and both assertions went back to matching an exact
substring that a stamped line no longer contains: the tag lands before
the colon, so "failed [key=k]: ..." is now "failed [key=k] [at=N]: ...".
Strip a well-formed tag before matching, as the branch's other oracles do.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* no-mistakes(review): unstamp fold colon tests; reserve stamp width in cap

* no-mistakes(document): correct stale unstamped status-line spellings in docs

* no-mistakes(document): quote brief-test literals for lint; correct stamp-helper contract comments

* no-mistakes(ci): rename subshell-local epoch in delivery-race stub

The serialization test overrides fm_pending_reply_mark_delivered inside a
(..) subshell. Its `epoch` local collided with the same name in
status_line_at_epoch/status_stamp_line, which this branch added and this
suite now calls at top level, so ShellCheck 0.11.0 reported SC2030 and
failed Lint 2. The stub already prefixes its other locals with `pending_`
for the same reason; `epoch` was the leftover.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(bin): unify Lavish host and disconnect handling (#5060)

* fix: ship clean Lavish host fixes

* no-mistakes(review): Fix Lavish classifications and fail-closed host loading

* no-mistakes(review): Restore Lavish host state across retries and launches

* no-mistakes(review): Preserve destination Lavish host when configuration is absent

* no-mistakes(document): Document Lavish status and host guarantees

* feat: act on captain's away words during AFK supervision (#5076)

* feat(afk): make the captain's away words the whole mandate

Retire the clause fields, verb list, never-set scan, refused records, and
the per-task merge-grant list from the away-posture record. The record is
now version 2: the captain's words verbatim plus expected return, spend
cap, and reach line; a version 1 record still validates, reads, and
archives so a live away window is never broken by the upgrade.

The supervision branch reads the words at the tail of every wake and acts
on them by its own judgment through the guarded scripts under standing
authority, never by analogy, holding for the return on doubt, and opens
each such outcome summary with "per your away instructions:" so the
return brief can render the words beside the session's account. While the
record exists any green merge runs under away authority (ledger tag
"away"); red merges, --allow-red, asynchronous and queued merges, and
local-only landing stay refused. The branch may file a backlog item the
words explicitly call for before dispatching it under the spend cap.

Tests drive fm-afk-contract.sh, fm-afk-launch.sh, fm-afk-return.sh, and
fm-pr-merge.sh as commands: version 2 written, version 1 read, retired
flags and subcommands refused by name, green merges landing under the
record, red and waived-red refused, the record lock still closing the
authority-read window, and the Pi away tail carrying the words.

* no-mistakes(review): carry the away read-back to the session verbatim

* no-mistakes(review): match the exact away-action marker in the return brief

* no-mistakes(review): refuse a words block truncated by a damaged line

* no-mistakes(document): Refresh away-role contract documentation

* fix(bin): render the remote charter's steering-inbox path host-local (#5049)

* fix(bin): render the remote charter's steering-inbox path host-local

A freshly provisioned remote secondmate read a parent-home absolute
steering-inbox path in its charter - a location that exists on no route -
and spent its first turn discovering the gap and filing a blocked
decision for what was a render defect. The seed's remote-copy rewrite now
maps the inbox to the route's host-local parent-route inbox, exactly as
it already maps the reply-log path, so every mention - bare path, listing,
and handled/ acknowledgement - lands host-local.

Both rewrites also become plain assignments, because a quoted substitution
nested inside a double-quoted printf argument leaks literal quotes into
the replacement text on stock macOS bash. The lifecycle suite pins the
corrected render both directions against the real seed, provisioning,
and delivery route, sharing one fixture value between the render truth
and the delivery truth.

Closes #5012

* no-mistakes(document): document remote charter's host-local steering inbox

* feat: route Lavish feedback directly to owning workers (#5099)

* feat(procevent): route worker-owned Lavish rounds

* no-mistakes(review): drop duplicate artifact field from task-owned registration

* no-mistakes(review): post worker reply once, fix ring label, keep re-arm atomic

* no-mistakes(review): keep worker board owned until terminal round acknowledged

* no-mistakes(review): refuse every retirement of an open worker-owned round

* no-mistakes(review): use real lavish reply flag, isolate reply generations

* no-mistakes(review): drop .posted marker for best-effort reply posting

* no-mistakes(review): consume staged reply after listener setup, refuse orphaned captures

* no-mistakes(review): require a reachable owner, redeliver open rounds, roll back failed re-arms

* no-mistakes(review): re-arm only to acknowledge an open round

* no-mistakes(review): conclude only a still-open terminal round

* no-mistakes(review): record the acknowledgement before retiring the board

* no-mistakes(review): retain the registration across a conclude, qualify terminal docs

* no-mistakes(document): Document worker-owned Lavish round lifecycle

* fix(bin): fit pull observation within the contribution poll budget (#5107)

* fix(bin): reserve contribution observation budget

* no-mistakes(review): Strengthen slow-read regression test to exceed the poll budget

* feat(bin): add idempotent inbox capture, replies, receipts, and readiness JSON (#5103)

* feat(bin): add idempotent inbox orders, receipts, replies, and readiness

Let a caller supply a request id when publishing a captain inbox note so a
retry returns the original note instead of creating a second one, including
across the crash window between save and wake announcement. Separate saved
from announced so a failed wake is repairable without enqueueing again.
Add bounded receipts JSON with omission disclosure, a durable primary reply
against a note id, and a read-only readiness projection that can say
unknown instead of inferring liveness from a lock file.

* no-mistakes(review): fix(bin): honest inbox announce, reply cursor, and readiness verdict

* fix(bin): resolve ready from lock-holder ancestry; drop lock status --json

Remove the extra JSON surface from fm-lock.sh so its human status still
always exits zero. Have the readiness projection classify the inspected
home from the lock-holder pid via fm-harness.sh ancestry, with an explicit
FM_SUPERVISION_MODEL still winning and an unknown model when there is no
holder. Prove the yes path when that ancestry names a known harness.

* no-mistakes(review): Harden inbox announce, receipts reads, and reply sequence cursor

* no-mistakes(document): Note read-only lock inspection in scripts inventory

* no-mistakes(lint): Pass missing id argument to malformed-reply test printf

---------

Co-authored-by: cliflacata-svg <304148223+cliflacata-svg@users.noreply.github.com>

* fix(bin): stop harness footer rows below a composer from reading as pending text (#5118)

* fix(composer): stop a harness footer row from reading as a composer holding text

A harness draws its own furniture below the composer - a user statusLine, a
permission-mode hint - and the cursorless "bottom-most shape wins" rule looks
exactly there. `→` (U+2192) is Cursor's prompt glyph but ordinary text
everywhere else, so a statusLine opening with `→` was selected as a bare
composer, swallowed the hint row beneath it as wrapped input, and answered
`pending` on a visibly empty pane. `fm_task_inbox_ring` defers on exactly that
verdict, and `bin/fm-watch.sh`'s re-ring calls the same function, so the first
doorbell and every retry were skipped and the worker never saw the steer.

Measured live on 2026-09-20: three of five Claude Code 2.1.236 worker panes on
Herdr 0.8.0 had genuinely empty composers and every one of them was refused.

A separator pair that closed over a bare agent-glyph row is a proven composer
container, so the contiguous non-blank rows below its closing rule are that
composer's footer and are no longer composer candidates. The demotion is bounded
by all three of its own preconditions: a blank row ends the zone, a pair that
closed over no glyph row demotes nothing, and a shape with no separator pair at
all (Cursor's half-block rules) is untouched. Real unsubmitted text in that same
composer, including a stray SGR mouse report left by a click in the pane, still
reads `pending`.

Pinned by two portable regressions and by a new cursorless arm on the live
composer-matrix guard, which re-reads each harness's already-proven-idle pane
the way every non-tmux backend reads it and fails naming the harness and
version when that read is `pending`.

* no-mistakes(review): make composer footer-zone demotion shape-independent

* no-mistakes(review): make footer-zone demotion refuse-only and drop rescan

* no-mistakes(lint): quote probe-absent sentinel to clear ShellCheck SC2100

---------

Co-authored-by: Koen Muller <koen@catapult.nl>

* feat(bin): append optional home-local include to briefs (#5115)

Co-authored-by: guanchengh-lgtm <271917158+guanchengh-lgtm@users.noreply.github.com>

* fix(bin): report a branch with no validation run as absent instead of an unreadable runs table (#5114)

* fix(bin): stop misreading a no-run branch as an unreadable runs table

Defect: when `no-mistakes axi status`'s overview is truncated (a task's
own branch has zero rows among the shown ones), fm_nm_select_run's
Python fallback derived the repo identity for its direct SQLite query
from a `repo: <path>` line it expected in the overview text. The real
CLI never emits that line, truncated or not (see the genuine capture at
tests/captures/no-mistakes-v1.70.1/overview.toon, which has only
`count:`/`runs[...]:`), so the lookup always failed and reported
"unreadable runs table" for a task that simply has no run on its
branch. On a fleet with many concurrent runs, every idle-branch task
hits the truncated-overview path routinely, so this fired every few
minutes and drowned genuine unreadable/blocked verdicts in noise.

Fix: derive the repo identity from the task worktree path instead,
which is exactly the value `no-mistakes` records as a repo's
`working_path` (confirmed against the existing capped-overview test
fixtures, which already register repos by worktree path). A worktree
path that is not absolute cannot be matched and still reads as
unreadable rather than being guessed at. Also raise the reader's
SQLite busy timeout from 1s to 30s so ordinary lock contention on a
busy fleet cannot masquerade as an unreadable database.

Safety: every other verdict byte-for-byte unchanged - the repo lookup
still requires exactly one matching row (a genuinely corrupt or
mismatched repos table still reports unreadable, per the existing
`repo` failure-mode test), the branch query and row validation are
untouched, and a zero-row result for the branch still flows through
the same recursive re-parse that already turns an empty `runs[0]{...}`
table into `absent`. Added a regression test
(test_capped_overview_without_repo_line_and_no_runs_reports_absent)
that reproduces the real overview shape - capped, zero rows for the
task's branch, no `repo: ` line - and asserts the crew state falls
through to the pane/busy verdict instead of reporting unknown or
"unreadable". Full fm-crew-state.test.sh suite passes unchanged
otherwise.

* fix: recovered same-branch inventory awk misreads empty result as unreadable

fm_nm_select_run's deep SQLite reader rebuilds a `count:`/`runs[...]:`
overview and re-runs it through the same awk selection pass. When that
rebuilt inventory has zero rows for the branch, the row-matching loop never
executes, so its counters (`seen`) stay at awk's uninitialized empty string
while `expected` and `shown` are plain strings parsed from the header text.
Comparing an uninitialized value against a non-numeric string uses string
comparison, so "" != "0" is true, and the END block takes the "unreadable
runs table" branch instead of falling through to the correct "absent"
verdict for a branch with genuinely zero runs.

Coerce the affected END comparisons with `+0` so they are always numeric,
matching seen/expected/shown/total regardless of whether awk classified
them as strings or numeric strings. A truncated or genuinely malformed
inventory still differs numerically and still reports unreadable.

* no-mistakes(review): bound capped-overview inventory reader and canonicalize worktree lookup

* no-mistakes(review): match recorded repo path first, tolerate duplicate spellings

* no-mistakes(review): revert repo lookup to exact working_path match

* no-mistakes(document): note state-db inventory read under crew-state nm timeout

* fix(bin): require a non-draft pull request before a PR-based done report (#5141)

* fix(bin): require a non-draft pull request before a PR-based done report

A PR-based ship could report done, and merge monitoring could be armed, while the pull request was still a draft. A draft cannot be merged, so the poll waited for an event that could not occur and nobody was asked to merge.

The PR-based definitions of done now require reading the pull request back from the forge and confirming it is not a draft, and a lane that deliberately holds a draft declares a wait instead of done.
bin/fm-pr-check.sh refuses to arm merge monitoring on a draft, naming the draft state, and treats an unreadable draft state as before.
The draft reading now lives in bin/fm-pr-lib.sh and bin/fm-pr-merge.sh uses it, with its refusal to merge a draft unchanged.

Closes #4757

* fix(review): Skip arm-time draft refusal when fm-pr-merge records metadata

* fix: support quota-axi schema 6 snapshots (#4904)

* fix(bin): accept quota-axi schema 6 snapshots keyed by provider + accountKey

quota-axi 0.1.47 emits schemaVersion 6 once a provider expands to more
than one account: every provider row carries an accountKey and one
provider id may appear on several rows. fm_quota_json_valid accepted
only schema 5 with unique provider ids, so fm-dispatch-resolve.sh,
fm-quota-choose.sh, and fm-procevent-quota.sh all rejected the live
snapshot and quota-informed dispatch was dead against the current tool.

- bin/fm-quota-axi-lib.sh: the validator accepts schema 6 with
  accountKey required on every row and uniqueness on
  provider + accountKey; schema 5 keeps its exact rules. FM_QUOTA_ROW_JQ
  is the one join every consumer uses: schema 5 binds by provider alone,
  schema 6 binds to the row keyed by the candidate's Pi lane, else the
  provider's default row, else no row (unmeasured, never blocked, never
  by position or summed across accounts).
- bin/fm-quota-choose.sh: accepts schema 6 JSON and the TOON accountKey
  column, and joins through the shared function.
- bin/fm-dispatch-resolve.sh and bin/fm-procevent-quota.sh: join through
  the shared function; an expanded provider with no row for the
  candidate's account is reported as such.
- tests: schema 6 fixtures shaped like the real snapshot, each paired
  with a schema 5 case on the same path; every new case fails on the
  previous scripts and passes now.
- docs: the two sentences naming the row join describe the schema 6 key.

* no-mistakes(review): Fix native Codex quota and expanded provider watches

* no-mistakes(review): Align native Codex account matching across dispatch paths

* no-mistakes(document): Align quota documentation with account-aware snapshots

* no-mistakes(document): Align quota dispatch documentation with account matching

* fix(bin): keep CI lint and the quota watch test portable

- bin/fm-quota-axi-lib.sh: FM_QUOTA_ROW_JQ is read only by the scripts
  that source this library, so full-mode ShellCheck reported SC2034 on
  the assignment; mark it alongside the existing SC2016 disable.
- tests/fm-procevent-quota.test.sh: the schema 6 provider-watch
  assertions used rg, which CI runners do not install, so the case
  failed with 'rg: command not found' rather than on behavior; use grep
  like the rest of the file.

* no-mistakes(document): Documented schema-version account-row compatibility

* test: fix Claude session-start drain live E2E (#5165)

* test: repair Claude live auto-arm regression

* no-mistakes(review): Assert SessionStart digest completeness within its hook_response event

* no-mistakes(document): Consolidate Claude live verification references

* ci: pin the no-mistakes required check to v1.80.1 (#5195)

Roll the shared require-no-mistakes action to the tagged v1.80.1 SHA and grant pull-requests: read so the check can read PR bodies.

* fix(bin): retain Pi watcher predecessor to stop false down alarms (#5174)

* fix: preserve Pi watcher ownership across session replacement

* no-mistakes(document): Scope Pi predecessor retention away from omp

* no-mistakes(ci): Diagnosed all three failing checks; only one was code-caused. (ci-3, genuine) Stock macOS Bash snapshot compatibility: `tests/fm-pi-watch-extension.test.sh` failed the macOS Bash 3.2 `bash -n` parse sweep with `line 4265: unexpected EOF while looking for matching '`. I built GNU Bash 3.2.0 from source locally and reproduced it. Root cause: the PR added a comment containing an apostrophe (`// Replacement shutdown deliberately retains module 2's established arm until`) inside a quoted here-document (`<<'EOF'`) nested inside a `$(...)` command substitution. Bash 3.2 has a parser bug (fixed in later bash) where an unmatched single quote inside such a here-doc body is treated as opening a shell quote and never closed, aborting the whole file parse. The base commit parses cleanly under Bash 3.2, confirming this PR introduced the break. Minimal fix: reworded the comment to remove the apostrophe (`... retains the established module-2 arm until`), preserving meaning. Verified `bin/fm-lint.sh --list-files` (the 6 changed shell files) now all pass `/tmp/bash-3.2/bash -n`; Bash 5 also parses. (ci-1, infrastructure) Behavior portable serial 8: GitHub API shows the `Run portable serial shard 8` step conclusion=success; only `Upload portable serial shard 8 timing artifact` failed with `Failed to FinalizeArtifact ... (403) Forbidden`. This is a transient artifact-service/cancellation failure, not a test or code failure. No change. (ci-2, infrastructure) Lint 1: fetched the job log via the GitHub API; it ends with `##[error]The runner has received a shutdown signal...` then exit 143. The step was cancelled mid-run, not a ShellCheck finding. Independently ran `bin/fm-lint.sh --partition 1of2 --telemetry ...` locally with pinned ShellCheck 0.11.0 and actionlint 1.7.12: exited rc=0 (no findings). No change. The only code change is the apostrophe removal in tests/fm-pi-watch-extension.test.sh; no other files modified

* fix(bin): allow cleanup of windowless legacy task records (#5236)

* fix(bin): retire windowless leftovers and stop claiming a Pi daemon teardown

Catch-up correctly refuses while a leftover task record has no status file.
Cleanup used to deadlock on those same records when they also had no spawn_gen and no window, so they lingered and wedged every later away-mode return. Teardown now treats a windowless leftover as a missing-endpoint legacy record, and stop reports that no daemon terminal was running when none was launched.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Narrow windowless teardown exception to tmux legacy leftovers

* no-mistakes(review): Validate windowless leftover identity via shared endpoint validator

* no-mistakes(review): Refuse windowless leftovers carrying other backends' endpoint identity

* no-mistakes(document): Clarify windowless teardown retry documentation

---------

Co-authored-by: Cursor <cursoragent@cursor.com>

* ci: exempt kunchenguid from the no-mistakes required check (#5256)

* fix(bin): surface launches parked on an interactive prompt as not-started (#5250)

* fix: surface parked launch prompts as not started

* no-mistakes(document): docs: record launch-prompt busy backstop classification

* no-mistakes(document): docs: align tail40 and rendered-text comments with launch-prompt backstop

* fix: record away posture immediately on /afk (#5260)

* feat(afk): make /afk itself the go with a same-turn record write

Collapse the propose-then-confirm away entry into one 'enter' step that
writes state/.afk-contract immediately and prints the announcement and
read-back after the record exists, never asking for a go. The retired
propose, confirm, and --proposal inputs are refused by name, and a stale
proposal left by an older version is removed rather than promoted.
Refresh and replace semantics, verbatim words, the single writer, the
never-set, and per-harness launch behavior are unchanged.

* no-mistakes(document): Refresh away-entry documentation evidence

* fix(bin): recognize passed-with-override as a passing outcome (#5294)

* fix(bin): map passed-with-override to done instead of unknown

no-mistakes' axi status emits outcome: passed-with-override for a run
that finished with an explicitly approved Test or CI exception. Both
bin/fm-crew-state.sh's outcome resolver and bin/fm-teardown.sh's
pre-teardown terminal-run check only matched the literal passed and
checks-passed tokens, so this outcome fell through to unknown/parked
and a finished worker awaiting merge kept getting re-alerted as stale,
while an abort race during teardown could also leave a finished run
misreported as still parked.

Map passed-with-override to the same done/terminal handling as a
clean passed in both places.

* fix(document): Replace stale outcome mapping with authoritative pointer

* fix(ci): Fixed a pre-existing mock-clock race in tests/fm-contributions.test.sh by advancing time only during the serial issue read. Reproduced the exact CI failure before fixing it. Forced-race replay, all 38 contribution scenarios, scoped ShellCheck, Bash syntax, and diff checks pass. Only the test fixture changed; CI rerun remains with the outer executor

* fix: clean up workers after their pull requests land (#5317)

* fix: close landed workers from supervision in both postures and at return

During the 2026-09-22 away window every exemption worker whose pull request
had merged was left sitting for nine hours. The supervision branch received
the stale wake, the merge-landed check, and the hourly inactive-outcome row
for each of them, ran the recovery playbook, found nothing to recover, and
reported "no further action". The branch prompt granted ordinary teardown of
a confirmed-landed task without ever naming the moment or the command, and
the playbook has no landed exit, so the stale path ended at "nothing to
recover". The return brief then listed only blockers, decisions, and the
latest five routine outcomes, so the landed workers stayed invisible after
the captain came back.

- bin/fm-branch-prompt.sh: name the merge-landed wake, and any later stale,
  inactive-outcome, or heartbeat row on a done task with a merged PR, as the
  moment to claim the lease and run bin/fm-teardown.sh with no flags; a
  refusal is reported, never forced or worked around. Add teardown to the
  handling tool list.
- stuck-crewmate-recovery: a landed worker is not a recovery case; point at
  the ordinary teardown owner for each actor.
- bin/fm-afk-return.sh: render a "Landed, cleanup due" section from durable
  records only (a live task record whose recorded PR carries the
  merge-notification marker), between could-not-fix and handled, without
  holding the gate; the afk skill's return step closes each listed task
  through ordinary teardown once the check clears.
- tests: pin the prompt rule in fm-branch-supervision and the brief section
  in fm-afk-return through the real marker writer.

* no-mistakes(document): Document landed-task cleanup ownership

* fix: surface green no-mistakes PRs awaiting merge (#5327)

* fix(bin): surface a green no-mistakes PR still in ci merge monitoring

A green PR could sit unreported because neither the worker nor the
supervisor could observe checks-green while the ci step kept monitoring
for the merge.

Supervisor read: fm_nm_select_run's capped-overview inventory reader looked
the repository up by the task worktree path, but no-mistakes registers a
repository once by its main clone path and resolves every linked worktree
to it, so on every task copy of a busy repo the lookup matched no row and
each read reported "complete same-branch run inventory unreadable". Key the
lookup on the overview's own top-level `repo:` line, which every axi
release emits as the resolved working_path.

Even with a readable run, the ci-log classifier treated "base branch
advanced ..., re-arming CI monitor timeout" as not-ready. The monitor logs
a checks state only when it changes and a base advance does not clear
readiness, so a green PR read as still validating for as long as main kept
advancing. Stop treating that line as a marker, matching no-mistakes' own
ci-log parser, and name the run's PR URL in the held-for-merge reading so
the existing inactive-outcome path can act on it without a worker report.

Worker contract: `axi status` never reports checks-passed while the ci
step monitors for merge, so the definition of done no longer makes a
status poll the wait for the next gate or outcome; the drive call's own
return is the green signal, reattached with `no-mistakes axi run` after a
bounded return.

* no-mistakes(review): read the full ci log when checking checks-green

* no-mistakes(review): correct stale ci log tail wording in docs

* no-mistakes(document): Document checks-green supervisor fallback

* fix: derive Lavish polling route from board session (#5334)

* fix: derive Lavish polling server from its board session

* no-mistakes(document): Document session-derived Lavish polling

* no-mistakes(document): Correct Lavish routing verification claims

* fix(bin): stop secondmate relaunch failing when watcher scratch files vanish (#4900)

* fix(bin): ignore vanished state scratch files on secondmate relaunch

Relaunch refused when find(1) exited non-zero while listing a secondmate
home's state directory. A live watcher can delete scratch files between
readdir and processing, which is not evidence that child *.meta records
are unreadable.

Prove the directory is listable from its mode and keep the existing
readable-meta loop as the child-record guarantee. Fixes #4765.

* no-mistakes(review): Skip chmod-000 unlistable-state relaunch test when running as root

* fix(bin): stop each keyed answer from re-waking this home (#4907)

* fix(bin): treat home-owned status closes as already read

Self-announced bookkeeping appends now record their exact byte ranges.
Later drains and signal scans skip those ranges, so two distinct
--resolve-key answers after an OPEN DECISIONS fold do not each wake the
supervisor. Worker-authored lines outside that ledger still signal.

* no-mistakes(review): Keep owned closes in unread status; lock ledger writes

* no-mistakes(review): Drop fold-lag wake suppression so folded worker decisions still wake

* no-mistakes(review): Require real owned growth before ledger marks status seen

* no-mistakes(document): Clarify home-appends ledger scope versus UNREAD STATUS

* no-mistakes(review): Restore fold-lag path, drop owned-range filters, fix test

* no-mistakes(review): Align ledger docs and scope ledger to wake path only

* no-mistakes(review): Restore stranded historical-annotation test comment to its function

* no-mistakes(review): Retire the home-appends lock alongside its ledger

* no-mistakes(document): Note ledger's lock-helper dependency in classify library

* no-mistakes(review): Append-and-coalesce home-appends ledger; fix stamped-line assertions

* no-mistakes(review): Drop redundant empty-span branch; make owned test pin ledger

* no-mistakes(document): Document covers' ascending-order dependency on home-appends ledger

* no-mistakes(document): Note owned-append skip in watcher signal-scan comment

* fix: deliver failed public follow-ups with updated AXI floors (#5350)

* chore(bin): raise tasks-axi, quota-axi, and lavish-axi floors to latest

Raise the minimum versions to tasks-axi 0.2.6, quota-axi 0.1.50, and
lavish-axi 0.1.77, pin CI's tasks-axi install to 0.2.6, and move the
floor-boundary test fixtures to the new versions.

tasks-axi 0.2.6 makes a failed relation deliverable for a promised-final
expecting pr-merged, so add the regression test: a bound work that ends
failed reports its honest outcome text through fm-public-followup-emit.sh,
consume marks the commitment ready, and deliver posts that text exactly
once.

Also make two hang-guard tests in fm-backlog-atomicity portable to hosts
without coreutils timeout, and stop an installed herdr from leaking into
the secondmate-liveness husk classifier test.

* no-mistakes(review): drop out-of-scope bounded_run hang-guard helper from atomicity test

* no-mistakes(review): pin quota-axi floor at 0.1.49 across fixtures

* no-mistakes(document): Document failed public-followup delivery behavior

* no-mistakes(ci): Updated quota-axi floor and all 0.1.49 fixtures to 0.1.51, corrected bootstrap boundaries to 0.1.51/0.1.52/0.1.50, and bumped the bearings lavish-axi stub to 0.1.77. Bearings, quota procevent, quota chooser, startup budget, and bootstrap floor coverage passed; the full bootstrap suite exceeded the 240-second local command limit after relevant checks passed. git diff --check passed

* fix(bin): refuse ship done: when the named head exists only in the worker copy (#4878)

* fix(bin): refuse ship done: when the named head lives only in the worker copy

A ship done: is not current-state done until that exact commit is reachable
outside the disposable copy. The check tests the named head, not whether
some branch moved.

* fix(bin): gate CI-ready ship done: on named-head reachability, not handoff

Keep no-mistakes' first done: as the pipeline handoff, apply the same shared
check when registering a PR and when a secondmate publishes ledger-first,
treat a recorded merged PR as landed after prune, and name the PR head
instead of scanning free-text SHAs.

* no-mistakes(review): Bind named-head gate to recorded PR and forge heads

* no-mistakes(review): Gate direct-PR forge heads and keep pending ledger deliveries

* no-mistakes(review): Align worker done wording, test mapping, pending-retry test

* no-mistakes(test): Raise watcher test time limit to stop load flake

* no-mistakes(document): Restore ledger-path fact and name named-head gate coverage

* ci: re-attest named-head ship-done gate for a fresh serial-3 verdict

* no-mistakes(review): Simplify local-only gate, gate keyed done lines, document recovery

* no-mistakes(document): Name fm-crew-state among named-head gate callers

* fix(bin): ring a proven-idle secondmate before raising a wake-loop stall alarm (#5204)

* fix(bin): ring a proven-idle secondmate before a wake-loop stall alarm

A leftover foreign-queue row on an idle, alive, ring-safe mate is still drainable in that home. Ring once, reset the observation interval, and keep the parent alarm for unknown, busy, or still-frozen rows.

* no-mistakes(review): Mark drain steer with from-firstmate fire-and-forget carrier

* test(watch-arm): size re-arm waits off the real loaded recovery cost (#5335)

The re-arm recovery cases judged "the watcher stayed live instead of
surfacing recovery" with fixed budgets below what a real stale-lock
recovery costs on a contended host: the arm's default 10s confirmation
deadline, a start helper that returned after about 4s whether or not the
arm had confirmed its watcher, and an 80-poll exit wait.
A changed-suite run beside other suites starves the recovery's many
short-lived processes while this suite's sleeping poll loops keep their
pace, so a watcher still surfacing its recovery read as one that stayed
live (issue #3793).
The original 0.25s window after confirmation was widened to 80 polls in
#3837, which left the same race at a larger size.

Following the CONTRIBUTING.md fixture-budget rule, the re-arm helper now
gives the arm an explicit 30s confirmation budget and waits for its
confirmation or exit within a ceiling that outlasts it, and every wait on
a re-armed watcher uses one named iteration-counted ceiling that outlasts
the same budget.
A passing case returns as soon as the arm reports or exits, and a watcher
that never surfaces its recovery still fails.

A new case delays every mktemp and readlink the re-armed watcher runs
after it publishes its beacon, so its first poll and exit take about 13s
on any host.
It fails with the reported symptom on the previous budgets and passes now.
No bin/ change.

* fix: stop watchers reliably during blocked polls (#5362)

* fix(bin): let one TERM always stop the watcher on bash 5.2

Bash 5.2 runs a pending trap from the parser entry of the next command
substitution it expands, where the trap body is parsed as the inside of
that substitution and fails ("trap: line 2: unexpected EOF while looking
for matching `)'") or is dropped silently, consuming the signal. The
watcher's `trap 'exit 1' HUP INT TERM` could therefore ignore a TERM and
keep polling while its stopper waited: the triage suite's reap waited
forever (CI jobs cancelled at 30 minutes), and the arm's signal path and
the away-mode daemon's shutdown wait for the watcher the same way.
Bash 5.3 fixed the parser; 5.2 is the stock bash on Ubuntu 24.04.

HUP and TERM now keep bash's native fatal-signal handling, which runs the
EXIT trap (watcher_cleanup) and exits on bash 3.2, 5.2, and 5.3. INT keeps
its trap because bash ignores a direct SIGINT while a child runs. The
check-spawn deferral window no longer contains a command substitution.

The triage suite's reap is now bounded and fails the case within 10s with
process evidence instead of hanging the job, and a new regression test
proves TERM stops a watcher blocked inside a poll's pane capture and still
releases its lock and records an acknowledgeable stop.

* no-mistakes(document): Clarify watcher stop-signal documentation

* fix: submit stuck inbox doorbells instead of skipping them (#5374)

* fix(bin): submit our own stuck doorbell instead of skipping every later ring

* no-mistakes(review): Confirm and retry Enter once on stuck-doorbell submit

* no-mistakes(document): Clarify doorbell retry and pending-composer documentation

* feat: add opt-in fleet activity ledger (#5375)

* feat(bin): add the opt-in fleet activity ledger

Homes that create config/fleet-ledger get an append-only JSONL file,
state/fleet-ledger.jsonl, recording task.dispatched, task.status,
task.merged, and task.cleaned_up so outside tools can follow a fleet.
With the flag absent each producer does one file test and nothing else.
docs/fleet-ledger.md owns the record contract and its documented limits.

* no-mistakes(review): Record task.status text verbatim after the first colon

* no-mistakes(document): Clarify fleet ledger status and setup documentation

* no-mistakes(ci): Fixed a timing race in tests/fm-pi-branch-extension.test.sh: the replacement-wake test now waits for the prompt to start before releasing it. The focused test passed twice, and git diff --check passed

* fix: validate public follow-up deliverables and wake on rejection (#5352)

* fix(bin): format, validate, and surface public-followup deliverables

brief pre-fills report_path=data/<work-id>/report.md and states the accepted
format of every value it cannot know instead of a bare <value> placeholder.
fm-public-followup-emit.sh refuses a deliverable tasks-axi would refuse, in
both the direct and staged destinations, naming the key, value, and format.
consume records the specific deliverable, outcome, or missing key behind a
tasks-axi refusal, and each refusal wakes the owning home once through the
existing relay poll.

* no-mistakes(review): refuse emits missing a required deliverable in both destinations

* no-mistakes(review): require promised deliverables and keep rejections recoverable

* no-mistakes(review): mirror tasks-axi's canonical pull request URL rule

* no-mistakes(review): keep a rejection wake whose line cannot be read

* no-mistakes(review): key emit-time rules on the promise, not the outcome

* no-mistakes(review): bound deliverable keys and values as tasks-axi does

* no-mistakes(review): state rejection wakes as at-least-once and pin it

* no-mistakes(review): enforce the promised contract tasks-axi holds at emit

* no-mistakes(review): stop inferring a staged promise from its outcome

* no-mistakes(document): Refresh public follow-up documentation

* no-mistakes(ci): Fixed both CI flakes. Watcher cleanup is now installed before singleton acquisition, preventing timeout races from leaving stale locks while preserving recovery-failure evidence. Bearings render fixtures now publish a valid isolated Lavish session store and retire each listener after rendering, eliminating false unowned-source races. Verified with checkpoint stress, fm-watch-checkpoint, fm-watcher-lock, repeated fm-bearings-board-render runs, project lint, syntax checks, and git diff checks

* Revert unrelated CI auto-fix edits to the watcher and bearings board test

The CI step's automatic repair changed bin/fm-watch.sh and
tests/fm-bearings-board-render.test.sh to chase two intermittent CI
failures that also occur on main and are not part of this change. Restore
both files so this branch carries only the public-followup deliverable fix.

* no-mistakes(review): Refuse a repeated --deliverable key at emit argument parsing

* no-mistakes(document): Clarify public-followup validation and rejection-wake documentation

* feat: add Devin CLI crewmate and scout adapter (#5380)

* Add verified Devin CLI worker adapter

* no-mistakes(review): Drop Devin resolver refusal and launch marker

* no-mistakes(review): Verify devin in bootstrap, fold kind rule, update docs

* no-mistakes(document): Document Devin sidecar, resume, and worker-only facts

* no-mistakes(document): Document Devin interrupt, liveness anchor, composer signals

* fix(control): never pair Devin interrupt presses on an idle agent

A fast double Escape on an idle Devin opens its /revert picker, where Enter
reverts file changes. fm-control now sends the second press only after the
first renders Devin's 'esc again to interrupt' armed hint, never sooner than
0.5 s, closes a revert picker a mistimed press opened with one Escape, and
refuses to type the exit command while that picker is open. An unarmed
interrupt reports cancel=not-running and leaves the busy record untouched.

* fix(devin): disable Claude hook import and commit attribution for workers

The per-task Devin config now forces read_config_from.claude=false, so a
worker no longer runs the user's or project's Claude Code hooks (including
Herdr's Claude agent-state hook), and attribution=false, so Devin adds no
Co-Authored-By trailer or Generated-with line to commits and PRs.

* test(devin): extend live guard and record Herdr and revert-picker evidence

The credentialed live guard now fails if an imported Claude Code hook runs,
if the worker's commit carries Devin attribution, if an idle interrupt sends
more than one press or opens the revert picker, or if an open picker lets
exit through or is closed with a revert. The Devin reference, agent-control
doc, and verification records carry the 2026-09-22 tmux and Herdr lab results,
including the Herdr exit refusal.

* no-mistakes(document): Correct Devin documentation links and lifecycle guidance

---------

Co-authored-by: Denis Beliaev <battler73@yandex.ru>

* fix(bin): recognize passed-with-skips as a passing outcome (#5322)

fm-crew-state classifies the no-mistakes outcome 'passed-with-skips' as
unknown, so a finished worker awaiting merge is re-alerted as stale. The
same blind spot lets fm-teardown's pre-teardown terminal-run check refuse
a legitimate abort race that lands on this outcome.

Map passed-with-skips to done in crew-state resolution, keeping the
skipped publication/CI verification visible in the detail rather than
reporting a clean pass, and recognize it as terminal during teardown.

* fix(bin): refuse unavailable backend adapters before sourcing (#5382)

* fix: refuse missing backend adapter before source

* no-mistakes(review): Gate backend precheck under stock Bash

* no-mistakes(document): Clarify adapter precheck docs

* no-mistakes(lint): Suppress intentional child Bash ShellCheck warning

* test: repair base-red liveness, export-DOM, and wake-queue self-tests (#5338)

* fix(test): repair tmux liveness and calm follow-up loaded_off regressions

Both self-tests fail on untouched main on a host whose coreutils are a
multicall binary and whose Chrome has no pre-warmed profile, and each failure
masks the other's file.

tests/fm-tmux-agent-liveness.test.sh - the stand-in harness processes were
symlinks to the host's `sleep`. A single-purpose `sleep` runs happily under
another name, but a multicall coreutils binary (uutils or busybox) resolves its
applet from argv[0]: `claude-link -> sleep` invoked under the harness name runs
the wrong applet and exits immediately, so no foreground process exists and
every positive case reads not-alive ("last verdict for liveness:agent was
missing (expected alive); title=sh comms=[sh ]"). Build a dedicated spinner as
the stand-in target, exactly the way the version-string case already builds its
executable, and require the fallback target to demonstrably survive the rename
before using it. Every assertion is untouched; the stand-in identity signal is
unchanged (the kernel still records the symlink name as the executable
identity).

tests/fm-calm-pi-extension.test.sh - render_export_dom pinned a brand-new
`--user-data-dir` per attempt. On Google Chrome for Testing 151.0.7922.34 that
pristine profile makes Chrome's first-run initialization never complete: the
browser and its renderers start, but --dump-dom never returns, so all three
bounded attempts end exit=0 timed_out=yes bytes=0 and the DOM assertions never
run ("could not render calm-mode HTML export DOM"). Chrome's own profile
creation under a fresh HOME renders the same document in about a second, so the
helper now gives Chrome a private per-attempt HOME instead of the explicit
profile flag. Each attempt still gets an isolated profile, and every DOM
assertion is unchanged.

Root-cause evidence: a pristine --user-data-dir with `--headless=new
--dump-dom` had not returned after 150s, while the same command with an empty
HOME and no --user-data-dir returned the full DOM in ~1s, and reusing an
already-populated profile also returned it in ~1s. The render failure masked
the rest of the file: with it repaired, the Pi follow-up loaded_off case passes
unmodified against an installed @earendil-works/pi-coding-agent package.

These two failures block downstream validation of every lane on hosts with
multicall coreutils or a fresh Chrome profile.

Verification:
- timeout 300 bash tests/fm-tmux-agent-liveness.test.sh -> exit 0, 16 assertions ok
- timeout 700 bash tests/fm-calm-pi-extension.test.sh -> exit 0, 13 assertions ok,
  including the Pi operational follow-up loaded_off case
- bash -n and shellcheck clean on both touched files
- rest of tests/: bin/fm-test-run.sh --all bounded by timeout 900 completed 17 files with 0 failures (fm-afk-contract.test.sh through fm-backend-herdr-launcher-workspace-e2e.test.sh), then the bound cut off the 18th (fm-backend-herdr-presentation-e2e.test.sh, a real-herdr-gated lab test) with no failure recorded

* fix(test): give wake-queue observation checkpoints the alerting ceiling

tests/fm-wake-queue.test.sh's secondmate stall case runs bounded foreground
watcher checkpoints whose job is to record an observation, with the alerting
checkpoint that follows asserting the stall. A checkpoint's exit publishes a
downtime marker, and the next checkpoint consumes it only by reaching the end of
the watcher's poll loop, where the recovery surfacing runs after the stall tick;
the observation itself is recorded by that same stall tick. On a loaded host a
1s ceiling sits under the cost of that iteration (which includes a pane capture
in the active-turn gate), so the observation was never recorded, the downtime
marker stayed pending, and the alerting checkpoint surfaced
`check: rearm-resurface` instead of the stall it asserts:

  not ok - a foreign queue with no progress did not alert: check: rearm-resurface
  not ok - a frozen reprovisioned queue generation was hidden: check: rearm-resurface

Give the observation checkpoints that feed a later alert the same 4s ceiling the
file already documents for alerting checkpoints. The ceiling is only a bound - a
checkpoint still returns on its first actionable wake - so no assertion is
weakened, and the quiet windows get longer, not shorter.

* no-mistakes(document): docs: correct export-DOM Chrome render root cause

* no-mistakes(review): Isolate Chrome profile on macOS, dedupe tmux CC_BIN lookup

* chore: re-trigger fork workflow approval for triage

---------

Co-authored-by: Captain <blackxwhite88@users.noreply.github.com>
Co-authored-by: kunchenguid <kunchenguid@users.noreply.github.com>

* fix: keep watcher status classification bounded to new log spans (#5383)

* fix(bin): classify a status span without re-folding the whole log

A watcher poll could take minutes, so its liveness beacon aged past the
guard's 300s grace and the Stop auto-arm reported the watcher down. On the
main home, cycles ended with beacon_age 91-235s while healthy and 534-706s
while the laptop was CPU-starved.

Cause: whenever a newly appended status span held a keyed needs-decision
or blocked line, status_span_first_actionable_record re-read and re-folded
the ENTIRE log to decide whether that opening was still live, forking
several subshells per line. On a remote second mate's mirrored parent
channel (1.2MB, ~2300 lines) that is 13-20k subshells, about 17s per log
per classification when idle, paid by every signal and heartbeat scan.

Nothing regressed recently: subshell counts per classification were
20,272 from #3268 (2026-08-29, which introduced the whole-log fold) and
13,188 from #3753 onward through HEAD. The cost grew with log size, since
parent-channel logs only grow.

Fix: fold only the captured span. An accepted opening does not depend on
earlier lines and only later lines close or supersede it, and every later
line lies inside the span, so the span fold names the same live openings
at a cost bounded by the span. Old and new classification outputs are
byte-identical across 51 span offsets of real-shaped secondmate and ship
logs.

A real-watcher regression test records every read the classification
makes through the span-reader seam and asserts none reaches before the
classified offset; it fails on the old code (5,157 bytes read from
offset 0 to classify an 84-byte span).

* no-mistakes(document): Clarify span classification and watcher regression coverage

* test: close pr-check watcher test gaps (original flake already fixed by #5362 and #4878) (#5381)

* test: fix watcher timing flakes in fm-pr-check-security

The bounded watcher's hang guard now counts only the watcher's own time: a
case marks the intervals where it holds the watcher on injected work or makes
it wait on concurrent work, and those no longer count against its budget. The
budget itself stays at main's sixty seconds. The helper also stops forcing a
one-second per-check timeout, which killed a correct merged poll whenever that
poll took longer than a second, so the watcher only retried it or exited on a
later check's wake without the merge.

The concurrent-publication case pauses the guard while its arming is in
flight, and its task now sorts before the contributions observer the arming
also registers, so the watcher stops on the poll under test before running
that unrelated fleet snapshot. The case also prints the watcher's stderr when
it fails.

The replacement case pauses the guard while the re-arm runs inside the
watcher, runs that injected arming with the fixture root every other arming
here uses, and waits on the replacement merge's process instead of a
two-second cap. Merged-poll runs retire the contributions observer before the
watcher starts, since no case here exercises it.

The returned-descendant case no longer races a four-second sleep or a TERM
landing at an arbitrary point in the watcher's idle loop: its descendant holds
until killed, and a second check in the same cycle witnesses that it was
drained and stops the watcher.

* no-mistakes(ci): Reproduced the intermittent board-render failure. Its Lavish stub listed an open session but omitted the session-state record required by the listener, so the build could race the listener’s exit. Added matching fixture state; the affected suite passed three consecutive runs, and shell syntax and diff checks passed

* Revert "no-mistakes(ci): Reproduced the intermittent board-render failure. Its Lavish stub listed an open session but omitted the session-state record required by the listener, so the build could race the listener’s exit. Added matching fixture state; the affected suite passed three consecutive runs, and shell syntax and diff checks passed"

This reverts commit 6a59859b2e2a3778f9b46faeea42d6de37468cd6.

* feat: record fleet status immediately and emit PR-ready events (#5385)

* feat: record task.pr_ready in the fleet ledger when a task PR is registered

* feat: record worker status lines in the fleet ledger as they are written

* no-mistakes(review): Keep worker status append failures and pass the resolved config to the ledger

* no-mistakes(review): Resolve relative config override before embedding in worker command

* no-mistakes(document): Clarify fleet ledger status capture timing

* test: synchronize foreign queue stall checks with watcher progress (#5386)

* test: synchronize foreign secondmate stall legs on the watcher's recorded observation

Each leg of test_secondmate_foreign_queue_stall_tracks_progress_and_alerts_once
ran the watcher under a 1s or 4s wall-clock checkpoint, but every later leg
depends on the progress observation the previous leg's watcher recorded. Under
load the watcher was killed before its first stall tick, the observation was
never written, and the next leg treated its own sighting as the first one, so
the stall alert never fired.

Run the watcher directly and end each leg on its observable outcome: the
progress marker recording the expected observation, or the watcher's own first
wake. Also move a comment orphaned above this test back to the drain liveness
test it describes.

* no-mistakes(review): Wait for full stall reset before stopping watcher leg

* test: isolate the bearings render fixture from the shared Lavish store (#5391)

The listener resolves its server from that store before it polls. Without a session for this bo…
@keenvc
keenvc force-pushed the fm/fm-slowness-scout branch from 8b76c35 to 2269539 Compare September 30, 2026 04:31
Comment thread bin/fm-session-start.sh Outdated
Comment thread bin/fm-session-start.sh Outdated
@keenvc keenvc changed the title feat(bin): merge upstream reliability batch and speed up session start fix(bin): speed up session start and wake drain, and sync the fork with upstream Sep 30, 2026
Comment on lines +69 to +77
if ! fm_afk_contract_away_present "$state"; then
FM_MERGE_AUTHORITY='attended'
FM_MERGE_AUTHORITY_REASON='attended'
return 0
fi
grants=$(FM_HOME="$home" FM_STATE_OVERRIDE="$state" \
"$_FM_MERGE_AUTHORITY_LIB_DIR/fm-afk-contract.sh" grants 2>/dev/null) || {
FM_MERGE_AUTHORITY_REASON='grants-unreadable'
return 1
}
while IFS= read -r grant; do
[ "$grant" = "$id" ] || continue
FM_MERGE_AUTHORITY='away-grant'
FM_MERGE_AUTHORITY_REASON='granted'
return 0
done <<EOF
$grants
EOF
FM_MERGE_AUTHORITY='away'
# shellcheck disable=SC2034 # Public results consumed by sourcing callers.
FM_MERGE_AUTHORITY_REASON='not-granted'
return 1
FM_MERGE_AUTHORITY_REASON='away'
return 0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 security Away record permits ungranted merges

When the captain is away, this resolver accepts a valid away record without checking whether the task has a merge grant or yolo=on. The merge gate can therefore merge a green PR without the task-specific consent previously required. That violates the repository requirement not to drift from VISION.md, which says merging requires the captain’s explicit word and autonomy must never be inferred.

How this was verified: The merge gate uses this resolver, which returns away authority based on record presence without checking a task grant.

Context Used: If there is a VISION.md file at the root of the repo, the PR must not conflict / diverge / drift from it. If the PR description has an "Intent" section, respect that as the accepted user intent. - Do make comments if anything in the implementation ... (source)

Knowledge Base Used: Session transitions and teardown

keenvc and others added 6 commits September 30, 2026 05:00
The fork sync (27a63ab) already carried upstream through a774c44;
this brings in the two later upstream commits. The Pi outcomes header
and export-boundary test take upstream's versions, which supersede the
fork's equivalent fix.
…ession start

Eliminate IFS subshell evaluations across wake drain, watch, and session scripts.
Bash ANSI C quoting IFS=$'\t' assigns the tab delimiter directly without forking printf subshells, preventing tens of thousands of process forks during large wake queue drains.

Detach home summary refresh during session start.
Running fm-home-summary-refresh.sh in the background after printing the session-start digest avoids holding the session lock or stalling interactive startup for 15-25s.

Batch queued backlog status lookups during bootstrap reconciliation.
Querying queued items once via tasks-axi list --state queued allows bootstrap to skip per-task show probes for non-queued records, reducing reconciliation latency by 20-25s.
The end-to-end half of this test seeds a git repo with an --allow-empty
commit, but never sets an author identity. A developer box with a global
git config passes; a CI runner has none, so the commit dies with
"fatal: empty ident name" and the digest the final assertion inspects is
never produced, failing "the wedged item must be reported by name as a
partial reconcile". Pass -c user.name/-c user.email on the commit, the
same pattern tests/fm-backlog-atomicity.test.sh already uses.
The queued-listing shortcut skipped fm_backlog_row_probe when listing
succeeded, so wedged show backends never named stuck items. Match upstream
behavior so partial reconcile lines still report each owned record.

Co-authored-by: Cursor <cursoragent@cursor.com>
…vert

Per-record backlog probes are required so wedged show backends still
surface partial reconcile lines; the single-list shortcut is not kept.

Co-authored-by: Cursor <cursoragent@cursor.com>
@keenvc
keenvc force-pushed the fm/fm-slowness-scout branch from f1d4fb6 to 37bad4e Compare September 30, 2026 05:35
@keenvc keenvc changed the title fix(bin): speed up session start and wake drain, and sync the fork with upstream feat(bin): land fork reliability batch, cline/openhands adapters, and session-start perf fixes Sep 30, 2026
@keenvc keenvc changed the title feat(bin): land fork reliability batch, cline/openhands adapters, and session-start perf fixes feat(bin): add cline/openhands adapters, cross-home claims, provider lane caps, and faster wake/session start Sep 30, 2026
@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate: triage on HEAD 37bad4e5f2a29d137434c8f645a81f6f72f70777.

contract-class: new-default. Unconfigured path evidence inspected tip-vs-main:

  • FM_PROVIDER_LANE_CAP_DEFAULT=4 in bin/fm-provider-lib.sh — spawn calls fm_provider_cap_refuse and refuses when live lanes ≥ 4 with no providerCaps config. New default-on refuse.
  • bin/fm-session-start.sh detaches home-summary refresh (behavior change on every session start).
  • New cline/openhands crew adapters expand the verified harness set (available when chosen).
  • Cross-home claims / fm-hold-reverify.sh arm are explicit opt-in surfaces, but they do not save the PR: the provider-cap default alone is new-default.
    No auto-merge.

VISION (brief): One captain/interface mixed (adapters deepen fleet; provider-cap refuse is new silent gate). Authority does not align for the default cap of 4 (consent not asked; unconfigured refuse). Scripts/agents aligns (mechanics in scripts). Restart aligns (durable claims/holds). Delegation mixed (new claim primitive). Fleet-outlives-vendor aligns (more verified adapters). Scope mixed (large surface growth toward workshop/adapters). Resist: default-on lane cap and session-start behavior change without captain grant.

Attestation: MATCH. workflow-zero: NO — .github/workflows/ci.yml only bumps expected snapshot test count 18→19 (benign; no RCE/sudo/pull_request_target/unpinned actions). CI: fork runs approved this pass (36676105868, 36676068202, 36674297149, 36674199888, 36674200004); NM edited runs SUCCESS; full CI queued.

Waiting on: CI. Even if green: new-default → Firstmate flag only once otherwise-ready (not yet).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants