Skip to content

fix: restore primary rewakes after attended main-only closes - #5961

Merged
kunchenguid merged 5 commits into
mainfrom
fm/fm-watcher-down-recurrence-s1
Sep 28, 2026
Merged

kunchenguid merged 5 commits into
mainfrom
fm/fm-watcher-down-recurrence-s1

Conversation

@kunchenguid

@kunchenguid kunchenguid commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Intent

The captain asked "are you missing any crewmate updates again?" - and yes: on 2026-09-27 the main home's supervision went down again and crewmate updates were missed for over an hour, shortly after #5941 (successor watcher on main-only pass-through + continuous remote-reply listeners) merged and the fleet updated (~16:31 PDT).

Fix the listener problem as soon as possible: ship the recommended smallest fix so an attended main-only pass-through rewakes the primary again while keeping #5941's successor, as a PR for the captain's merge.

This time the fix must be more carefully validated live.

The recommended smallest fix, from the investigation of that outage: since #5941, when the supervision host (bin/fm-supervision-host.sh) passes an attended main-only close through to main, it starts a successor watcher and also confirms a handling handoff for it, which moves the watcher recovery marker (state/.watcher-down) from downtime to handling. The Claude Stop hook re-arm owner (bin/fm-claude-stop-autoarm.sh) rewakes an idle primary only while that marker reads downtime, so the close reached no one and the idle primary slept while wakes queued. The fix drops that handoff confirmation in leave_successor_for_main while keeping the successor running, and on the path where a close turns main-only at its own turn (HANDLE_RC=2) republishes downtime after detach_successor.

What Changed

  • Keep the successor watcher running without confirming a handling handoff when an attended close passes through to main. Restore the downtime marker when a close becomes main-only at its turn, allowing the Stop hook to rewake the primary.
  • Add host and Stop-hook regression coverage plus an opt-in live test for repeated hand-backs, successor closes, and remote-listener delivery; update the supervision documentation.

Risk Assessment

✅ Low: The host fix is narrowly scoped to the two intended hand-back paths; the remaining finding concerns one test assertion, not the delivery logic.

Testing

The host regression run reached the changed hand-back cases but timed out later; the Stop-hook regression completed. No real idle Claude primary or negative control was driven, because the isolated configuration had no login and the supplied live guard would write outside the permitted boundary.

  • Live validation: ⚠️ inconclusive - 0 of 5 scenarios driven live against the product
Scenario Result Live Evidence
An idle Claude primary receives and acknowledges an attended main-only close while its successor watcher stays alive ⏸️ untested no The fixture-backed host-and-hook regression passed, but it is not an interactive Claude primary. A disposable worktree-local Claude configuration reported loggedIn=false. The supplied live guard requi…
After re-arming, a second main-only close and the successor’s own close reach the idle primary ⏸️ untested no The fixture-backed successor-close regression passed, but repeated interactive turns require an authenticated Claude primary. The isolated configuration had no login, and the supplied guard cannot run…
A close that becomes main-only at its turn rewakes main rather than remaining marked as handled ⏸️ untested no The host-and-hook regression passed with a fake harness, not a live Claude turn. The worktree-local Claude configuration was unauthenticated; a disposable login is needed for an interactive proof.
A remote-reply listener remains owned while repeated hand-backs are delivered ⏸️ untested no The listener’s repeated-delivery scenario depends on the supplied interactive guard. Its isolated Claude configuration had no login, while the guard as written would write outside the worktree.
The pre-fix host leaves the same idle Claude primary unwoken as a negative control ⏸️ untested no The control requires a real authenticated idle Claude primary. The disposable configuration reported loggedIn=false, and using the guard’s operator-login transcript store is outside the permitted boun…
Evidence: Host regression transcript, including changed cases and timeout boundary

Source: Host regression transcript, including changed cases and timeout boundary

ok - report surface: only the branch actor's current turn may report, and only on the tasks its wake names
ok - report surface: visible late outcomes queue a relay, while silent outcomes remain stored without a wake or note
ok - dispatch entry: the host reads branch eligibility, the offer rule, and the wake prompt from the Pi branch's own owner
ok - drain: BRANCH OUTCOMES runs only on an opted-in home whose primary is not Pi
ok - drain: captain outcomes come first, and routine overflow collapses into a count one drain clears
ok - drain: repeated captain outcomes collapse per task, and the byte cap presents only the run its acknowledgement covers
ok - drain: a long away window costs one short drain, captain outcomes collapsed per task and routine overflow counted, and nothing from it is shown again
ok - drain: the BRANCH OUTCOMES budgets count bytes, cutting multibyte summaries by whole characters in any locale
ok - drain: branch outcomes stay unread when a projection of the store fails
ok - drain: branch outcomes stay unread and the drain fails when jq is missing
ok - drain: branch outcomes stay unread when the drain cannot print them
ok - drain: a legacy backlog is presented with each outcome's age and a check-first instruction, never adopted
ok - drain: a keyed decision survives acknowledgement through a newer outcome for its task
ok - drain: an outcome carried across a switch off Pi comes back with its age, never adopted
ok - drain: an outcome nothing has shown is presented until acknowledged, and a repeated acknowledgement changes nothing
ok - drain: an outcome the host drain presented but main never acknowledged stays unprocessed across a switch to Pi
ok - drain: an outcome the host drain presented but main never acknowledged survives an outcome index repair
ok - host: an attended wake the branch may take is handled on the engine, and its routine outcome never wakes main
ok - host: an attended captain outcome wakes main once and stays in its drain until main acknowledges it
ok - host: a captain outcome recorded after the captain left waits for the return, then reaches main's drain
ok - host: an attended decision close stays main's exactly as the plain arm delivers it
ok - host: a main-only pass-through leaves the successor watcher running and the close undelivered for main
ok - host: an attended close whose main session cannot be identified reaches main and runs no engine turn
ok - host: a decision close accepted away whose turn starts attended still reaches main unchanged
ok - host: an attended close whose task turns main-only before its turn still reaches main unchanged
ok - host+hook: an attended main-only pass-through rewakes main and keeps its successor watcher
ok - host+hook: a close that turns main-only at its turn rewakes main and keeps its successor watcher
ok - host+hook: a successor close that lands during main's turn is delivered at the next turn end
ok - host: a primary with no verified dialog mirror keeps every attended close on main, and its away posture still runs
ok - host: each wake carries the captain's dialog since the last wake, without operational input
ok - host: the dialog mirror, its feed, and the wake file are owner-only
ok - host: dialog a turn never completed with its report (a park boundary, a stopped turn, no report) reaches the next turn
Evidence: Stop-hook regression transcript

Source: Stop-hook regression transcript

ok - auto-arm: inert in a linked child worktree even when in-flight
ok - auto-arm: inert with no session lock
ok - auto-arm: a demonstrably dead recorded session owner is reclaimed through fm-lock.sh before arming
ok - auto-arm: inert without arm, rewake, or lock replacement when another live harness owns the home
ok - auto-arm: inert while AFK owns supervision
ok - auto-arm: stale-owner recovery leaves the AFK and supervision-need gates unchanged
ok - auto-arm: resolves the outermost pid of a nested contiguous claude ancestry (bg-spare chain)
ok - auto-arm: inert with nothing in flight and no X-mode need
ok - auto-arm: actionable close translates to exactly one exit-2 rewake with reason
ok - auto-arm: actionable close survives a healthy successor without duplicate delivery
ok - auto-arm: an attached cycle's end starts a handling successor named after the closed arm before the rewake
ok - auto-arm: an unconfirmed handling successor is reported in the rewake instead of blocking it
ok - auto-arm: bounded failure verification emits one automatic-mechanism alarm
ok - auto-arm: consecutive failures keep Stop-owned retry without repeating notice
ok - auto-arm: marker-write refusal defers delivery until one successor commits the notice
ok - auto-arm: unverified clean close exhausts retries and fails closed
ok - auto-arm: post-alarm actionable outcomes cannot continue or reset failure state
ok - auto-arm: benign cycle end with a live watcher and fresh beacon stays silent across the next cycle
ok - auto-arm: budget contention preserves the episode and forces a reset retry
ok - auto-arm: owner-mutex contention preserves successor episode state
ok - auto-arm: X-mode poll need arms the cycle even with no tasks in flight
ok - auto-arm: a registered custom check arms the cycle even with no tasks in flight
ok - auto-arm: concurrent firings admit one owner and one rewake translation
ok - auto-arm: TERM mid-arm commits a durable failure and exits 2 for rewake
ok - auto-arm: an abandoned owner claim is reclaimed so a lapsed cycle re-arms
ok - auto-arm: an abandoned-claim reclaim reaps a dead steal mutex without nesting
ok - auto-arm: a legacy owner still arming is never reclaimed while its watcher keeps beating
ok - auto-arm: a fresh legacy arming claim is never reclaimed while its startup window is still open
ok - auto-arm: a live claim the ledger does not name is never reclaimed
ok - auto-arm: a claim whose pid was reused is reclaimed even while its ledger entry still reads arming
ok - auto-arm: a reused-pid claim is reclaimed even with no ledger entry to prove it
ok - auto-arm: an identity-matched owner still arming is never reclaimed
ok - auto-arm: the guard's terminal-check claim is never reclaimed
ok - auto-arm: a stuck live legacy owner is retired via TERM and its lock reclaimed
ok - auto-arm: a SIGSTOPped legacy owner is reclaimed with TERM pending and dies on continue
ok - auto-arm: a live open generation claim defers concurrent firings with no lock held
ok - auto-arm: a hung generation owner with no watcher beat is superseded so re-arming self-heals
ok - auto-arm: an identityless arming ledger never defers the gate (reused-pid loophole closed)
ok - auto-arm: a superseded owner never re-invokes the arm and leaves its successor's claim untouched
ok - auto-arm: a superseded owner goes silent - one supersession episode, one translation, no held mutex
ok - auto-arm: need vanishing mid-cycle closes without a rewake
ok - auto-arm: mid-cycle AFK hands triage to the daemon with no rewake
ok - auto-arm: active in a marked secondmate home
ok - auto-arm: a long FM_POLL with FM_GUARD_GRACE unset reaches fm-watch-arm.sh with the derived grace
ok - auto-arm: without config/supervision-host the hook runs the arm exactly as before
ok - auto-arm: an opted-in home runs the host bound to its generation, and a host line rewakes like a wake
ok - auto-arm: a wake the host hands back under the away record says it is automatic supervision, not a return
ok - auto-arm: without the host the rewake banner is unchanged, eight wake lines at most
ok - auto-arm: a host handback delivers every host line, while its wake lines keep their cap
ok - auto-arm: a host that stood down closes silently without a retry
ok - auto-arm: a host that died without a close is retried, then reported as a failure
ok - fm-lock: shared session-lock lib preserves the status path
ok - auto-arm: stands down only on a pi-code-delivered transcript_path (/.pi/)
Evidence: Disposable Claude configuration authentication status

Source: Disposable Claude configuration authentication status

{
  "loggedIn": false,
  "authMethod": "none",
  "apiProvider": "firstParty",
  "analyticsDisabled": false,
  "projectsDirectory": "~/.no-mistakes/worktrees/016d88035d58/01M3JWRBFJ1PFP2JB65SB8PRBH/.attended-auth-probe-44519/projects",
  "configDirectory": "~/.no-mistakes/worktrees/016d88035d58/01M3JWRBFJ1PFP2JB65SB8PRBH/.attended-auth-probe-44519"
}
- Outcome: ⏭️ skipped across 2 runs (27m34s)

Live validation on the final head

The pipeline's test step could not drive a real Claude primary, so the live proof was run outside the pipeline with the operator's Claude login, in disposable lab homes, on the final pushed head 2c9128820fffbe7799685a1c1970dd6218c0e2da (Claude Code 2.1.283, model haiku, macOS 26.6.2 arm64, 2026-09-27 22:25-22:29 PDT).

$ FM_SUPERVISION_HOST_ATTENDED_LIVE_E2E=1 FM_SUPERVISION_HOST_ATTENDED_LIVE_CONTROL_REF=origin/main tests/fm-supervision-host-attended-live-e2e.test.sh
Scenario Result Evidence
1. An idle real Claude primary is woken by an attended main-only close, with the successor watcher alive pass 22:27:49 host log pass-through attended main-only; successor watcher pid 63411 alive; marker announced:downtime:57583.1790573268.XT7j6U. 22:27:55 Stop hook rewake at 1790573270 with ledger outcome=rewake ... recovery_generation=57583.1790573268.XT7j6U; the idle primary ran bin/fm-wake-drain.sh then bin/fm-wake-drain.sh --ack-through 4 --recovery-generation 57583.1790573268.XT7j6U.
2. A second main-only close, and the successor's own close, are both delivered pass 22:27:57 the turn end re-armed onto successor 63411. 22:28:18 event 2 closed it (watcher_pid=63411 origin=attached reason=attached-delivered-wake), the host passed it through (pass-through attended main-only at 1790573293), and the primary was rewoken at 1790573294 and acknowledged it.
3. A close that becomes main-only at its turn rewakes main pass 22:28:50 a routine close on a second task was accepted for the supervision session and its successor handoff confirmed (marker announced:handling:52920.1790573328.fihGRJ); a decision landed before the turn started; the host logged pass-through attended main-only with no engine turn. 22:28:51 rewake at 1790573330 with ledger outcome=rewake ... recovery_generation=52920.1790573328.fihGRJ (the same generation, so downtime was restored); the primary drained and ran --ack-through 13.
4. Remote-reply listeners stay owned across repeated hand-backs pass The stand-in remote listener (a local append-only log) kept runner pid 58694 owned at every check from setup through scenario 3; 22:28:42 it mirrored the remote event (1 ingested), which was delivered to the idle primary at 1790573318.
5. Negative control: the pre-fix host on origin/main (3c2a91d7) leaves the idle primary unwoken pass 22:26:00 host log pass-through attended main-only; 22:27:30 after 90s no rewake and no primary command; ledger outcome=arming; marker announced:handling:23180.1790573158.IdFki2; 4 queued rows undelivered.

Both labs passed the Stop-hook registration check and reached their hand-backs. No prompt was typed into either primary after its setup prompt.

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 2 warnings
  • ⚠️ tests/fm-supervision-host-attended-live-e2e.test.sh:271 - The new registration check greps raw settings.json text. It can match an inactive or unrelated occurrence of the hook name without establishing that Claude has a Stop hook registered. Parse the JSON and assert the effective Stop-hook entry instead.

🔧 Fix applied.
2 warnings still open:

  • ⚠️ tests/fm-supervision-host-attended-live-e2e.test.sh:271 - The new registration check greps raw settings.json text. It can match an inactive or unrelated occurrence of the hook name without establishing that Claude has a Stop hook registered. Parse the JSON and assert the effective Stop-hook entry instead.
  • ⚠️ tests/fm-supervision-host-attended-live-e2e.test.sh:271 - Round 1's fix introduced a check that rejects the actual tracked Stop hook. Its command ends in /bin/fm-claude-stop-autoarm.sh" because the shell path is quoted, so endswith("/bin/fm-claude-stop-autoarm.sh") is false. Both control and positive labs fail at step 1 before exercising the hand-back. Match the parsed command's actual shell invocation.
⏭️ **Test** - skipped
  • ⚠️ tests/fm-supervision-host-attended-live-e2e.test.sh:62 - The required interactive Claude hand-back was not demonstrated live. The supplied guard creates a lab and deletes a Claude transcript directory outside the worktree, contrary to this run's workspace boundary. An isolated Claude login was not available without using the operator's credentials. The fixture-backed checks cannot establish that an idle Claude primary actually rewakes.
  • ⚠️ live validation verdict: inconclusive (0 of 5 scenarios were driven live against the product); untested: An idle Claude primary receives and acknowledges an attended main-only close while its successor watcher stays alive, After re-arming, a second main-only close and the successor's own close reach the idle primary, A close that becomes main-only at its turn rewakes main rather than remaining marked as handled, A remote-reply listener remains owned while repeated hand-backs are delivered, The pre-fix host leaves the same idle primary unwoken as a negative control
  • Live validation: ⚠️ inconclusive - 0 of 5 scenarios driven live against the product
Scenario Result Live Evidence
An idle Claude primary receives and acknowledges an attended main-only close while its successor watcher stays alive ⏸️ untested no The existing live guard requires out-of-worktree lab and Claude-store writes and cleanup. I checked that route and ran the fixture-backed host checks, but could not use the guard within the workspace…
After re-arming, a second main-only close and the successor's own close reach the idle primary ⏸️ untested no The interactive guard has the same out-of-worktree writes and isolated-login blocker. Fixture checks cannot prove the repeated Claude turn; provide a workspace-contained guard and disposable authentic…
A close that becomes main-only at its turn rewakes main rather than remaining marked as handled ⏸️ untested no I ran the executable fixture check, but it uses a fake harness. A live Claude turn could not be driven under the guard's workspace and login requirements; provide a workspace-contained authenticated l…
A remote-reply listener remains owned while repeated hand-backs are delivered ⏸️ untested no The listener scenario is part of the interactive guard, which could not be run without its out-of-worktree writes and an isolated Claude login. Provide a workspace-contained authenticated lab.
The pre-fix host leaves the same idle primary unwoken as a negative control ⏸️ untested no The control ref was available, but its live guard uses the same prohibited out-of-worktree writes and requires an isolated authenticated Claude login. Provide both through a workspace-contained guard.
  • bash tests/fm-supervision-host.test.sh (twice; timed out after 240s and 600s)
  • bash tests/fm-claude-stop-autoarm.test.sh (completed successfully)
  • Inspected the live guard's lab, credential, and cleanup requirements against the workspace boundary

🔧 No changes applied.
2 warnings still open:

  • ⚠️ The required live hand-back and pre-fix negative control remain unverified. An isolated Claude configuration reports loggedIn=false. The opt-in guard uses the operator’s existing login and writes a transcript outside the worktree, so running it as written would violate this run’s workspace boundary. Provide a disposable Claude login confined to the worktree, or authorize a separate validation run with the required credential and filesystem access.
  • ⚠️ live validation verdict: inconclusive (0 of 5 scenarios were driven live against the product); untested: An idle Claude primary receives and acknowledges an attended main-only close while its successor watcher stays alive, After re-arming, a second main-only close and the successor’s own close reach the idle primary, A close that becomes main-only at its turn rewakes main rather than remaining marked as handled, A remote-reply listener remains owned while repeated hand-backs are delivered, The pre-fix host leaves the same idle Claude primary unwoken as a negative control
  • Live validation: ⚠️ inconclusive - 0 of 5 scenarios driven live against the product
Scenario Result Live Evidence
An idle Claude primary receives and acknowledges an attended main-only close while its successor watcher stays alive ⏸️ untested no The fixture-backed host-and-hook regression passed, but it is not an interactive Claude primary. A disposable worktree-local Claude configuration reported loggedIn=false. The supplied live guard requi…
After re-arming, a second main-only close and the successor’s own close reach the idle primary ⏸️ untested no The fixture-backed successor-close regression passed, but repeated interactive turns require an authenticated Claude primary. The isolated configuration had no login, and the supplied guard cannot run…
A close that becomes main-only at its turn rewakes main rather than remaining marked as handled ⏸️ untested no The host-and-hook regression passed with a fake harness, not a live Claude turn. The worktree-local Claude configuration was unauthenticated; a disposable login is needed for an interactive proof.
A remote-reply listener remains owned while repeated hand-backs are delivered ⏸️ untested no The listener’s repeated-delivery scenario depends on the supplied interactive guard. Its isolated Claude configuration had no login, while the guard as written would write outside the worktree.
The pre-fix host leaves the same idle Claude primary unwoken as a negative control ⏸️ untested no The control requires a real authenticated idle Claude primary. The disposable configuration reported loggedIn=false, and using the guard’s operator-login transcript store is outside the permitted boun…
  • bash tests/fm-supervision-host.test.sh (timed out after 500 seconds; relevant host-and-hook cases passed before timeout)
  • bash tests/fm-claude-stop-autoarm.test.sh (completed)
  • CLAUDE_CONFIG_DIR=<disposable worktree directory> claude auth status (reported loggedIn=false)
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

An attended main-only pass-through confirmed a handling handoff for the
successor it leaves running, which flipped the recovery marker to handling.
The Claude Stop hook only rewakes main while that marker reads downtime, so
the close reached no one and an idle primary slept with wakes queued.

The pass-through now leaves the marker at downtime, and a close that turns
main-only at its turn hands the consumed handoff back to downtime before it
reaches main. Regression tests drive the real Stop hook around the real host
on both paths and for the successor's own later close, and a new opt-in live
guard proves it against an idle interactive Claude primary with a pre-fix
negative control.
@greptile-apps

greptile-apps Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[High risk] Changes supervision host handoff and recovery logic.

The PR appears safe to merge on the reviewed code, although the opt-in live guard remains scheduling-sensitive and interactive validation was not completed.

Reviews (3) · Last reviewed commit: "no-mistakes(ci): Fixed the Stop hook’s r..."

Comment thread bin/fm-supervision-host.sh
Comment thread tests/fm-supervision-host-attended-live-e2e.test.sh Outdated
…k notifies main instead of silently dropping the close. Corrected the live guard’s tracked-hook check and added the requested at-turn main-only scenario. The new regression failed before the fix and passed after it; the host suite, syntax checks, and diff check pass. The credentialed live guard was not run in this CI phase because it writes outside the worktree
Comment thread bin/fm-claude-stop-autoarm.sh Outdated
Comment on lines +382 to +386
decide_at_handoff "$lab" "$lab/fm/state/demo2.status" \
"needs-decision [at=$(date +%s)] [key=lab-e4]: pick a rollout window" > "$lab/handoff.out" &
injector=$!
e4=$(date +%s)
printf 'working [at=%s]: rollout prep started\n' "$e4" >> "$lab/fm/state/demo2.status"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Decision injection races host. The live test waits to see the handling marker before appending the decision, but the host can complete its second offer before the polling process runs. The close is then handled instead of handed back, making this opt-in validation fail depending on scheduling rather than reliably testing the at-turn hand-back.

… gets its bounded retry before a non-crash hand-back failure is reported. The Stop-hook suite passes, including the crash regression. The host suite passed the failed-marker-write regression but timed out before completing; syntax and diff checks pass
@kunchenguid
kunchenguid merged commit 1b82b77 into main Sep 28, 2026
38 of 39 checks passed
@kunchenguid
kunchenguid deleted the fm/fm-watcher-down-recurrence-s1 branch September 28, 2026 05:40
neel-mishra pushed a commit to neel-mishra/firstmate that referenced this pull request Sep 28, 2026
…guid#5961)

* fix: wake an idle Claude primary for attended main-only hand-backs

An attended main-only pass-through confirmed a handling handoff for the
successor it leaves running, which flipped the recovery marker to handling.
The Claude Stop hook only rewakes main while that marker reads downtime, so
the close reached no one and an idle primary slept with wakes queued.

The pass-through now leaves the marker at downtime, and a close that turns
main-only at its turn hands the consumed handoff back to downtime before it
reaches main. Regression tests drive the real Stop hook around the real host
on both paths and for the successor's own later close, and a new opt-in live
guard proves it against an idle interactive Claude primary with a pre-fix
negative control.

* no-mistakes(review): Assert live lab Stop-hook registration via parsed settings JSON

* no-mistakes(document): Correct supervision hand-back documentation

* no-mistakes(ci): Fixed the failed downtime-write path so the Stop hook notifies main instead of silently dropping the close. Corrected the live guard’s tracked-hook check and added the requested at-turn main-only scenario. The new regression failed before the fix and passed after it; the host suite, syntax checks, and diff check pass. The credentialed live guard was not run in this CI phase because it writes outside the worktree

* no-mistakes(ci): Fixed the Stop hook’s retry ordering: a crashed host gets its bounded retry before a non-crash hand-back failure is reported. The Stop-hook suite passes, including the crash regression. The host suite passed the failed-marker-write regression but timed out before completing; syntax and diff checks pass
neel-mishra added a commit to neel-mishra/firstmate that referenced this pull request Sep 28, 2026
#3)

* fix: grant Claude workers access to Firstmate task channels (kunchenguid#5884)

* fix(bin): grant Claude workers their task-channel dirs via --add-dir

Since Claude Code 2.1.257, a file-tool read (Read/Glob/Grep, and an
Edit's mandatory prior read) of a path outside the working directories
parks --permission-mode auto panes on a one-time interactive question,
and a "Block" answer lands permissions.blockReadsOutsideWorkingDirectories
in user settings, refusing the same reads even under bypass. Firstmate
launches Claude with no --add-dir, so a secondmate's parent-home steering
inbox and a ship or scout worker's launch record, steering inbox, brief
dir, and code-root .agents/skills were all outside: workers wedged on
the question the first time they read a steer.

Every Claude launch, spawn and relaunch, in both permission modes, now
grants exactly the task's channel directories: state/<id>.inbox for a
secondmate (in the parent home), or state/operational-inbox,
state/<id>.inbox, data/<id>, and the code root's .agents/skills for a
ship or scout. Paths resolve to real paths and lazily created channel
dirs are made before launch so the grant never names a not-yet-existing
directory; the whole state/ is deliberately never granted.

The grant keeps the bypass-mode launch argv changed on purpose: it also
protects bypass workers against a machine-recorded Block answer.

* no-mistakes(document): Consolidate Claude launch guidance in configuration reference

* no-mistakes(document): Clarify Claude permission documentation reference

* fix(bin): prevent idle recovery loops without stranding wakes (kunchenguid#4819)

* fix(supervision): prevent idle recovery loops without stranding wakes

* no-mistakes(review): Remove unused wake-append rollback helper

* fix: reduce remote worker and polling helper process churn (kunchenguid#5889)

* fix(bin): stop the remote-job worker busy-polling an idle queue

The serving loop slept 50ms between passes and re-ran state preparation
(chmod on every queue directory), the heartbeat publish, and the stale sweep
on every pass. It now blocks on a worker.wake FIFO that staging,
cancellation, and lane exit nudge, keeps a short fast-poll window after
activity, refreshes the heartbeat at most once a second, and runs the sweep
(which re-applies the queue directories' 0700 modes) at startup and then on
a bounded interval. Lane-owned records are no longer re-read every pass.

Measured with a fork/execve-interposing counter on a --serve worker in a
disposable HOME and queue, bash 3.2, 20-second windows (the counter slows
the old loop to about 5 passes a second, so real-host rates were higher):
  idle worker             146 forks/s, 61 execs/s -> 4.8 forks/s, 3.1 execs/s
  one running long job    232 forks/s, 100 execs/s -> 15 forks/s, 11 execs/s
Stage-to-result latency for a no-op job, idle and back to back, stayed at
about 0.8-1.2s in both versions (dominated by job execution, not pickup).

* perf(bin): drop per-cycle forks from watcher, drain, and lock helpers

The watcher, drain, inactive-reconcile scan, and branch-outcome reads forked
small external commands on every cycle where bash can do the same work.

- fm-wake-lib.sh gains fm_dirname_to, fm_basename_to, and
  fm_epoch_seconds_to, exact stand-ins for $(dirname --), $(basename --),
  and $(date +%s); the clock uses printf %(%s)T on bash 4.2+ and still forks
  date exactly once on stock macOS bash 3.2.
- fm_lock_abs_path, fm_wake_signal_seen_path, fm_path_age, the watcher's
  age_of and wedge timer, and the recovery-marker line count use them or
  plain reads instead of dirname/basename/tr/date/wc.
- window_to_task reads a meta file once instead of two
  grep | tail -1 | cut -d= -f2- pipelines per file per call.
- fm-classify-lib.sh reads uname -s once at source time instead of in every
  status stat helper.
- Libraries sourced every cycle derive their own directory without forking
  dirname, including the backend adapter siblings a subshell re-sources on
  each probe.

tests/fm-fork-free-helpers.test.sh pins each replacement against the command
it replaces on edge-case inputs, under every available bash and both the C
and a UTF-8 locale; CI's stock macOS bash lane runs it under /bin/bash 3.2.

Measured with a fork/execve-interposing counter in a disposable home, one
tmux crew task, FM_POLL=1 (forks and execs per watcher cycle, per run
otherwise):
  watcher cycle      bash 5.3  299/138 -> 199/66   bash 3.2  341/146 -> 224/80
  drain              bash 5.3  492/238 -> 430/200  bash 3.2  567/250 -> 491/212
  inactive scan      bash 5.3   27/14  ->  17/4    bash 3.2   37/14  ->  17/4
  branch-outcome     bash 5.3   40/21  ->  35/16   bash 3.2   48/24  ->  38/19

* test: note the interpreter-expanded version probe for shellcheck

* no-mistakes(review): Fix worker idle bounds and fork-free contributions snapshot

* no-mistakes(review): Coalesce buffered worker wake nudges into one wake

* no-mistakes(review): Coalesce wake nudges via pending marker so publishers never block

* no-mistakes(review): Claim wake nudges atomically via noclobber pending marker

* no-mistakes(review): Release abandoned wake claims only after a 30-second bound

* no-mistakes(review): Drop worker wake FIFO; load path helpers side-effect free

* no-mistakes(document): Document remote worker polling and preemption cadence

* no-mistakes(ci): Fixed both failing CI shards: isolated remote and teardown test fixtures now include fm-path-lib.sh, which fm-wake-lib.sh requires. The three affected tests, fm-lint.sh, and git diff --check pass locally

* fix: keep remote reply listeners and watcher cycles running (kunchenguid#5941)

* fix: keep a successor watcher and remote-reply listeners across the gaps that dropped them

A main-only supervision pass-through exited without leaving a watcher, and each remote-reply poll released its claim until the next cycle, so short-lived listeners stayed down.

* no-mistakes(document): Clarify listener and supervision continuity documentation

* no-mistakes(ci): Fixed the three Greptile findings: failed ingestion leaves one durable capture, failed reads exit instead of relistening, and the disposable-checkout guard rejects state paths outside the marked lab. Added behavioral tests; the remote-reply and watcher-lock suites, shell syntax checks, and git diff checks passed

* no-mistakes(ci): Fixed a race in the wake-queue interruption test: it now waits for the drain to own the lock and enter handling before signaling it. The wake-queue suite, shell syntax check, and diff check pass

* fix: make attended cutover outcome re-presentation check-first (kunchenguid#5925)

* fix: date replayed branch outcomes and ask main to check current state first

A captain outcome main never acknowledged is presented again, which after a
harness or posture switch, or the first drain after the upgrade whose earlier
presenter never advanced the read cursor, can be days after its situation
settled. The replay read as fresh news, so a PR since merged looked ready.

bin/fm-branch-outcome.sh now adds a "recordedAgo" age (minutes, hours, then
days) to present and unprocessed rows, one owner of that wording for both
presenters. The drain's BRANCH OUTCOMES captain lines and the Pi branch's
processing request name that age and ask main to check the task's current
state first; an outcome already settled needs only the acknowledgement, with
nothing relayed to the captain. Nothing is adopted as processed, so a fresh
home's first outcome is still presented until acknowledged.

* no-mistakes(review): Absent processed marker reads 0; never adopt read cursor

* no-mistakes(review): Require recordedAgo in Pi requests; report undated rows to main

* no-mistakes(review): Keep recordedAgo on captain rows only in present output

* no-mistakes(document): Correct cutover documentation and retire stale migration guidance

* fix: keep settled branch outcomes out of main's reply to the captain

A live Pi primary that took over a host-drain home received the carried-over
outcomes dated and check-first, but its processing reply still told the
captain about an outcome whose decision had since been answered. The request
also claimed every outcome was already shown as an anchor entry in this
transcript, which is false for an outcome carried over from before a restart
or a switch of primary.

The Pi processing request now says each outcome was recorded earlier and may
already have been seen or handled, and that a settled outcome gets no
captain-facing mention at all in the reply or any recap, not even that it is
settled. The drain's BRANCH OUTCOMES header and the supervision docs state the
same rule, and the tests check both delivered texts.

* fix: scope main's outcome reply to what is still open

Telling main what not to say about a settled outcome was not enough: in two
live Pi trials the processing reply still told the captain that an answered
decision was settled. Main now sorts the outcomes by current state first, and
its reply to the captain covers only the still-open ones, written as if the
settled ones had never been listed. With that framing three live Pi trials
kept the settled outcome out of the reply and relayed the open one each time.

The drain's BRANCH OUTCOMES header and the supervision docs use the same
framing, and the tests check both delivered texts.

* no-mistakes(review): Clarify that main acknowledges every presented captain outcome

* no-mistakes(document): Clarify outcome cursor ownership across Pi and host

* no-mistakes(ci): Fixed Pi replay by batching unprocessed captain outcomes oldest-first and acknowledging only through each batch. Verified a marker-less backlog over 1 MiB replays through all batches. A real-drain regression confirms an older keyed decision remains under OPEN DECISIONS after a newer branch row is acknowledged; the check-first instruction now names those decisions. Relevant targeted tests and branch-supervision tests passed; the full host suite timed out

* no-mistakes(ci): Fixed the host drain’s check-first wording in bin/fm-wake-drain.sh; the CI fixture now passes. The full host suite passed the affected fixtures but timed out later. Syntax and diff checks passed

* no-mistakes(ci): Fixed ci-1: abbreviated Pi outcome summaries now stay within 1,024 characters and include a row-specific full-outcome lookup command. The delivered instruction requires reading the full outcome before acting, relaying, or acknowledging it. The new extension-driver regression failed before the fix and passes now; the Pi and supervision-host suites pass

* no-mistakes(ci): Corrected the batching sentence in docs/pi-supervision-branch.md. The cancelled CI check needs no code fix; its clean rerun passed. The Pi branch extension suite and git diff check passed

* fix: restore primary rewakes after attended main-only closes (kunchenguid#5961)

* fix: wake an idle Claude primary for attended main-only hand-backs

An attended main-only pass-through confirmed a handling handoff for the
successor it leaves running, which flipped the recovery marker to handling.
The Claude Stop hook only rewakes main while that marker reads downtime, so
the close reached no one and an idle primary slept with wakes queued.

The pass-through now leaves the marker at downtime, and a close that turns
main-only at its turn hands the consumed handoff back to downtime before it
reaches main. Regression tests drive the real Stop hook around the real host
on both paths and for the successor's own later close, and a new opt-in live
guard proves it against an idle interactive Claude primary with a pre-fix
negative control.

* no-mistakes(review): Assert live lab Stop-hook registration via parsed settings JSON

* no-mistakes(document): Correct supervision hand-back documentation

* no-mistakes(ci): Fixed the failed downtime-write path so the Stop hook notifies main instead of silently dropping the close. Corrected the live guard’s tracked-hook check and added the requested at-turn main-only scenario. The new regression failed before the fix and passed after it; the host suite, syntax checks, and diff check pass. The credentialed live guard was not run in this CI phase because it writes outside the worktree

* no-mistakes(ci): Fixed the Stop hook’s retry ordering: a crashed host gets its bounded retry before a non-crash hand-back failure is reported. The Stop-hook suite passes, including the crash regression. The host suite passed the failed-marker-write regression but timed out before completing; syntax and diff checks pass

* Clarify live Claude login for opted-in tests (kunchenguid#5975)

* fix(bin): ensure resumed worker launches enter their recorded worktree (kunchenguid#5916)

* Fix worker launches to enter recorded worktrees

* no-mistakes(review): placeholder

* no-mistakes(document): Update agent-control.md worktree-refusal note to match new universal cd+assert

* no-mistakes(review): Add regression tests for Orca spawn/relaunch worktree carve-outs

* Fix PR relaunch and prelaunch cwd verification

* no-mistakes(document): Fix docs/agent-control.md: worktree cwd check is pre-launch, not post-launch

* fix: slim worktree launch change onto upstream main

* fix(bin): make a Herdr task pane render before its launch is delivered

A Herdr pane created with --no-focus is not rendered until its tab has
been the active tab of a focused workspace once. Until then the launch
still executes but `pane read` stays empty and Herdr's screen-based agent
state never observes the worker, so a spawned worker's terminal reads
blank and `agent prompt` stalls (Herdr issue kunchenguid#2449).

The spawn now activates the task endpoint immediately before delivering
the launch command and restores the exact prior focused workspace and
tab right after the Enter, matching the tmux backend where
`capture-pane` reads the live pane terminal directly.

Adds bin/backends/herdr.sh task-rendering activation primitives, the
fm-spawn wiring, an updated focus note in docs/herdr-backend.md, and a
portable fake-CLI regression in tests/fm-backend-herdr.test.sh.

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: Joseph Kim <jokim1@gmail.com>
Co-authored-by: Mehul Bhagwani <mehulbhagwani@gmail.com>
Co-authored-by: Neel Mishra <neelmishra@Neels-Mac-Mini.local>
knowttl pushed a commit to knowttl/firstmate that referenced this pull request Sep 29, 2026
…guid#5961)

* fix: wake an idle Claude primary for attended main-only hand-backs

An attended main-only pass-through confirmed a handling handoff for the
successor it leaves running, which flipped the recovery marker to handling.
The Claude Stop hook only rewakes main while that marker reads downtime, so
the close reached no one and an idle primary slept with wakes queued.

The pass-through now leaves the marker at downtime, and a close that turns
main-only at its turn hands the consumed handoff back to downtime before it
reaches main. Regression tests drive the real Stop hook around the real host
on both paths and for the successor's own later close, and a new opt-in live
guard proves it against an idle interactive Claude primary with a pre-fix
negative control.

* no-mistakes(review): Assert live lab Stop-hook registration via parsed settings JSON

* no-mistakes(document): Correct supervision hand-back documentation

* no-mistakes(ci): Fixed the failed downtime-write path so the Stop hook notifies main instead of silently dropping the close. Corrected the live guard’s tracked-hook check and added the requested at-turn main-only scenario. The new regression failed before the fix and passed after it; the host suite, syntax checks, and diff check pass. The credentialed live guard was not run in this CI phase because it writes outside the worktree

* no-mistakes(ci): Fixed the Stop hook’s retry ordering: a crashed host gets its bounded retry before a non-crash hand-back failure is reported. The Stop-hook suite passes, including the crash regression. The host suite passed the failed-marker-write regression but timed out before completing; syntax and diff checks pass
RooseveltAdvisors pushed a commit to RooseveltAdvisors/firstmate that referenced this pull request Sep 29, 2026
…guid#5961)

* fix: wake an idle Claude primary for attended main-only hand-backs

An attended main-only pass-through confirmed a handling handoff for the
successor it leaves running, which flipped the recovery marker to handling.
The Claude Stop hook only rewakes main while that marker reads downtime, so
the close reached no one and an idle primary slept with wakes queued.

The pass-through now leaves the marker at downtime, and a close that turns
main-only at its turn hands the consumed handoff back to downtime before it
reaches main. Regression tests drive the real Stop hook around the real host
on both paths and for the successor's own later close, and a new opt-in live
guard proves it against an idle interactive Claude primary with a pre-fix
negative control.

* no-mistakes(review): Assert live lab Stop-hook registration via parsed settings JSON

* no-mistakes(document): Correct supervision hand-back documentation

* no-mistakes(ci): Fixed the failed downtime-write path so the Stop hook notifies main instead of silently dropping the close. Corrected the live guard’s tracked-hook check and added the requested at-turn main-only scenario. The new regression failed before the fix and passed after it; the host suite, syntax checks, and diff check pass. The credentialed live guard was not run in this CI phase because it writes outside the worktree

* no-mistakes(ci): Fixed the Stop hook’s retry ordering: a crashed host gets its bounded retry before a non-crash hand-back failure is reported. The Stop-hook suite passes, including the crash regression. The host suite passed the failed-marker-write regression but timed out before completing; syntax and diff checks pass
andrewesweet pushed a commit to andrewesweet/firstmate that referenced this pull request Sep 30, 2026
…guid#5961)

* fix: wake an idle Claude primary for attended main-only hand-backs

An attended main-only pass-through confirmed a handling handoff for the
successor it leaves running, which flipped the recovery marker to handling.
The Claude Stop hook only rewakes main while that marker reads downtime, so
the close reached no one and an idle primary slept with wakes queued.

The pass-through now leaves the marker at downtime, and a close that turns
main-only at its turn hands the consumed handoff back to downtime before it
reaches main. Regression tests drive the real Stop hook around the real host
on both paths and for the successor's own later close, and a new opt-in live
guard proves it against an idle interactive Claude primary with a pre-fix
negative control.

* no-mistakes(review): Assert live lab Stop-hook registration via parsed settings JSON

* no-mistakes(document): Correct supervision hand-back documentation

* no-mistakes(ci): Fixed the failed downtime-write path so the Stop hook notifies main instead of silently dropping the close. Corrected the live guard’s tracked-hook check and added the requested at-turn main-only scenario. The new regression failed before the fix and passed after it; the host suite, syntax checks, and diff check pass. The credentialed live guard was not run in this CI phase because it writes outside the worktree

* no-mistakes(ci): Fixed the Stop hook’s retry ordering: a crashed host gets its bounded retry before a non-crash hand-back failure is reported. The Stop-hook suite passes, including the crash regression. The host suite passed the failed-marker-write regression but timed out before completing; syntax and diff checks pass
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant