Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
14 commits
Select commit Hold shift + click to select a range
688edbf
test(fixtures): preserve the 499-record watcher cycle-exit capture
mattadams-dev Jul 31, 2026
a038a5f
fix(watch): arm the successor before handling a wake, and give parked…
mattadams-dev Jul 31, 2026
26a63d5
test(watch): cover successor arming, attached-cycle delivery, and the…
mattadams-dev Jul 31, 2026
8d5ccd3
docs: record the arm-layer continuity floor and the parked check-in c…
mattadams-dev Jul 31, 2026
67a58d7
fix(tests): register temp roots so a suite cannot leak its own processes
mattadams-dev Jul 31, 2026
397ed82
docs(verification): record the captured defect and the guard-class mu…
mattadams-dev Jul 31, 2026
d6d8fd6
test(fixtures): key the cycle-exit capture on its consumer directory
mattadams-dev Jul 31, 2026
55a5c2d
docs(scripts): note that the arm wrapper arms the next cycle before d…
mattadams-dev Jul 31, 2026
f3cdbcc
fix(tests): create the cleanup registry only when a suite takes a tem…
mattadams-dev Jul 31, 2026
d2152a7
no-mistakes(review): fix parked check-in batching and successor outpu…
mattadams-dev Jul 31, 2026
0921b11
no-mistakes(test): wait for TERM-resistant peer handler before restar…
mattadams-dev Jul 31, 2026
58dae2b
no-mistakes(document): dedupe cycle-exit capture summary and fix stal…
mattadams-dev Jul 31, 2026
076cbbf
no-mistakes(document): re-measure guard-class mutation matrix and add…
mattadams-dev Jul 31, 2026
32b26f1
no-mistakes(document): record surviving mutant at redundant pause-rel…
mattadams-dev Jul 31, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -111,7 +111,7 @@ state/ volatile runtime signals; gitignored
.afk durable away-mode flag; present = sub-supervisor may inject escalations (set by /afk, cleared on user return)
.watch.lock .wake-queue.lock watcher singleton and queue serialization locks
.claude-autoarm.lock .claude-autoarm-epoch .turnend-claude-blocks Claude Stop auto-arm single-flight, epoch, and guard-budget records; never touch
.hash-* .count-* .stale-* .stale-since-* .paused-* .wedge-escalations-* .seen-* .hb-surfaced-* .last-* .heartbeat-streak watcher internals; never touch
.hash-* .count-* .stale-* .stale-since-* .paused-* .wedge-escalations-* .seen-* .hb-surfaced-* .last-* .heartbeat-streak .watch-successor-output.* watcher internals; never touch
.watch-triage.log watcher's absorbed-wake debug log (size-capped); never relied on, safe to delete
.last-watcher-beat watcher liveness beacon, touched every poll (including while absorbing benign wakes); guard scripts read it
.subsuper-* .supervise-daemon.* sub-supervisor internals; never touch
Expand Down
149 changes: 146 additions & 3 deletions bin/fm-watch-arm.sh
Original file line number Diff line number Diff line change
Expand Up @@ -42,6 +42,23 @@
# loud. A live cycle already present means re-arm attaches - do not start a second
# watcher.
#
# ARMING THE SUCCESSOR IS THE FIRST ACT OF CONSUMING A WAKE, before the wake is
# handled. The watcher is one-shot by design: it exits to deliver one actionable
# reason. Nothing here used to start the next cycle, so supervision ENDED on
# every delivered wake and resumed only when some adapter above this layer
# happened to re-arm - a blind window as wide as the whole handling turn, and
# unbounded whenever that adapter path was inert (a forked Claude session whose
# Stop hook can never claim the home is the observed case). Handling first and
# arming afterwards cannot close that window; only this ordering can.
# The successor is deliberately NOT this arm's child: it must outlive this
# process, which exits as soon as it delivers the wake. Its stdout goes to a
# pid-keyed state/.watch-successor-output.<pid> file, so whichever arm later
# attaches to it can still deliver that cycle's real reason instead of reporting
# a genuine wake as an unexplained empty close.
# A successor is armed only where supervision would otherwise end: a delivered
# actionable wake, in a home that still needs supervision and is not away. Typed
# failures stay loud and unarmed so the adapter and the model still repair them.
#
# Every observed watcher cycle appends one tab-separated lifecycle record to
# state/.watch-cycle-exits.log. The arm layer owns that bounded ledger; it records
# arm/watcher identities, timestamps, exit/signal classification, beacon age,
Expand All @@ -61,6 +78,8 @@ set -u
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
# shellcheck source=bin/fm-wake-lib.sh
. "$SCRIPT_DIR/fm-wake-lib.sh"
# shellcheck source=bin/fm-supervision-lib.sh
. "$SCRIPT_DIR/fm-supervision-lib.sh"

WATCH="$SCRIPT_DIR/fm-watch.sh"
WATCH_LOCK="$STATE/.watch.lock"
Expand All @@ -79,6 +98,11 @@ CONFIRM_TIMEOUT=${FM_ARM_CONFIRM_TIMEOUT:-$ARM_CONFIRM_DEFAULT}
ATTACH_POLL=${FM_ARM_ATTACH_POLL:-0.5}
CYCLE_LOG="$STATE/.watch-cycle-exits.log"
CYCLE_LOG_LOCK="$STATE/.watch-cycle-exits.lock"
# Where a detached successor's one printed reason lands until an arm claims it.
SUCCESSOR_OUT_PREFIX="$STATE/.watch-successor-output"
# How long an unclaimed successor output survives before it is treated as debris.
SUCCESSOR_OUT_TTL=${FM_WATCH_SUCCESSOR_OUT_TTL:-3600}
case "$SUCCESSOR_OUT_TTL" in ''|*[!0-9]*) SUCCESSOR_OUT_TTL=3600 ;; esac
CYCLE_LOG_MAX_BYTES=${FM_WATCH_CYCLE_LOG_MAX_BYTES:-262144}
CYCLE_LOG_KEEP_LINES=${FM_WATCH_CYCLE_LOG_KEEP_LINES:-1000}
ARM_PID=${BASHPID:-$$}
Expand Down Expand Up @@ -241,6 +265,101 @@ report_attached() {
echo "watcher: attached pid=$HEALTHY_PID (beacon ${age}s)"
}

successor_output_path() { # <watcher-pid>
printf '%s.%s' "$SUCCESSOR_OUT_PREFIX" "$1"
}

# Drop only successor outputs no live watcher can still be writing to AND that
# no arm plausibly still owes a delivery for. Age is the second condition on
# purpose: pruning purely on a dead pid would race an attached arm that has just
# observed the close and is about to claim that exact file.
prune_successor_outputs() {
local f pid
for f in "$SUCCESSOR_OUT_PREFIX".*; do
[ -e "$f" ] || continue
pid=${f##*.}
case "$pid" in
''|*[!0-9]*) ;;
*) fm_pid_alive "$pid" && continue ;;
esac
[ "$(fm_path_age "$f")" -ge "$SUCCESSOR_OUT_TTL" ] || continue
rm -f "$f" 2>/dev/null || true
done
}

# A successor is armed only where supervision would otherwise end. An away home
# is skipped because the daemon owns the cycle there, and a home with no
# in-flight work and no relay poll is skipped so a detached watcher can never
# outlive the work it was supervising. FM_WATCH_ARM_SUCCESSOR=0 is the explicit
# opt-out for a caller that owns continuity itself.
successor_wanted() {
[ "${FM_WATCH_ARM_SUCCESSOR:-1}" != 0 ] || return 1
[ -e "$STATE/.afk" ] && return 1
fm_supervision_needed "$STATE" "$GRACE"
}

# Start one detached watcher and confirm it through the same honesty gate an
# owned child passes. Prints "started:<pid>" once confirmed, "attached:<pid>"
# when another watcher won the singleton instead, and "none" otherwise. It NEVER
# fails the wake delivery it precedes: a home left without a successor is
# recorded as such in the ledger and still gets its reason.
arm_successor() { # <just-closed-watcher-pid>
local closed_pid=${1:-none} tmp pidfile forked deadline pid target
successor_wanted || { printf 'none'; return 0; }
prune_successor_outputs
tmp=$(mktemp "$SUCCESSOR_OUT_PREFIX.pending.XXXXXX" 2>/dev/null) || { printf 'none'; return 0; }
pidfile=$(mktemp "$SUCCESSOR_OUT_PREFIX.forkpid.XXXXXX" 2>/dev/null) || {
rm -f "$tmp" 2>/dev/null || true
printf 'none'
return 0
}
# Double-fork so the watcher is reparented away from this arm: it must survive
# the exit that delivers the wake. setsid additionally detaches the controlling
# terminal where it exists; the bare subshell fallback covers hosts without it.
# The double fork hides the watcher's pid from $!, so the forked shell records
# its own pid and then execs the watcher into it: the pid it wrote IS the
# watcher's, whether or not setsid inserted a fork of its own.
# shellcheck disable=SC2016 # $$ and $1/$2 must expand in the forked shell, not here.
if command -v setsid >/dev/null 2>&1; then
( setsid "${BASH:-bash}" -c 'printf "%s" "$$" > "$1"; exec "$2"' fm-successor "$pidfile" "$WATCH" \
>"$tmp" 2>/dev/null </dev/null & ) 2>/dev/null
else
( "${BASH:-bash}" -c 'printf "%s" "$$" > "$1"; exec "$2"' fm-successor "$pidfile" "$WATCH" \
>"$tmp" 2>/dev/null </dev/null & ) 2>/dev/null
fi
deadline=$(( $(date +%s) + CONFIRM_TIMEOUT + 1 ))
while :; do
if healthy_watcher && [ "$HEALTHY_PID" != "$closed_pid" ]; then
pid=$HEALTHY_PID
forked=$(cat "$pidfile" 2>/dev/null || true)
if [ -n "$forked" ] && [ "$forked" = "$pid" ]; then
target=$(successor_output_path "$pid")
# Another arm may already own this watcher and its output file. Never
# clobber a reason someone else is still owed.
if [ -e "$target" ]; then
rm -f "$tmp" 2>/dev/null || true
else
mv -f "$tmp" "$target" 2>/dev/null || rm -f "$tmp" 2>/dev/null || true
fi
rm -f "$pidfile" 2>/dev/null || true
printf 'started:%s' "$pid"
return 0
fi
# Someone else's watcher holds the singleton, so our fork stood down and
# its stdout is a "already running" note, not that cycle's reason. The
# winner's reason lands under its own arm's file; publishing ours there
# would make a later attached arm read a genuine wake as an empty close.
rm -f "$tmp" "$pidfile" 2>/dev/null || true
printf 'attached:%s' "$pid"
return 0
fi
[ "$(date +%s)" -ge "$deadline" ] && break
sleep 0.2
done
rm -f "$tmp" "$pidfile" 2>/dev/null || true
printf 'none'
}

# Give a successor the same bounded confirmation window used for a fresh child.
# Adapter-owned continuations normally win immediately, but the bound avoids a
# false failure when process-close delivery and lock publication cross briefly.
Expand All @@ -265,7 +384,7 @@ fail_unexplained_cycle() {
# to a verified successor. With no successor, fail loudly instead of returning a
# clean empty completion that an adapter could mistake for a no-op.
attach_and_wait() {
local attached_pid=$1
local attached_pid=$1 out claimed reason_type successor
while :; do
if healthy_watcher; then
if [ "$HEALTHY_PID" != "$attached_pid" ]; then
Expand All @@ -277,6 +396,27 @@ attach_and_wait() {
sleep "$ATTACH_POLL"
continue
fi
# The cycle this arm was following has closed. A watcher writes its one
# reason and only then releases the lock, so if that cycle woke for
# something real its reason is already complete on disk. Deliver it - an
# attached close carrying a genuine wake is a wake, not the unexplained
# empty completion this used to report.
# The rename is the claim: with two arms attached to the same watcher only
# one can win it. The loser did NOT get the reason, so it must not report a
# delivery - it falls through to the successor path and either follows the
# cycle the winner armed or fails loudly, never exits clean and empty.
out=$(successor_output_path "$attached_pid")
if watch_output_has_wake "$out"; then
reason_type=$(watch_output_reason_type "$out")
claimed="$out.claimed.$ARM_PID"
if mv -f "$out" "$claimed" 2>/dev/null; then
successor=$(arm_successor "$attached_pid")
cycle_log_append 0 none "$reason_type" "$successor"
print_watch_output "$claimed"
rm -f "$claimed" 2>/dev/null || true
return 0
fi
fi
if wait_for_healthy_successor; then
cycle_log_append unknown unknown attached-cycle-ended "attached:$HEALTHY_PID"
attached_pid=$HEALTHY_PID
Expand Down Expand Up @@ -405,11 +545,14 @@ cycle_begin "$child" started
child_done=0

owned_child_finished() {
local rc=$1 signal reason_type status
local rc=$1 signal reason_type status successor
signal=$(cycle_signal_name "$rc")
if [ "$rc" -eq 0 ] && watch_output_has_wake "$child_out"; then
reason_type=$(watch_output_reason_type "$child_out")
cycle_log_append "$rc" "$signal" "$reason_type" none
# Arm first, deliver second. Nothing below this line may run before the
# next cycle exists, or the home is blind for the whole handling turn.
successor=$(arm_successor "$cycle_watcher_pid")
cycle_log_append "$rc" "$signal" "$reason_type" "$successor"
print_watch_output "$child_out"
rm -f "$child_out" 2>/dev/null || true
child=
Expand Down
77 changes: 70 additions & 7 deletions bin/fm-watch.sh
Original file line number Diff line number Diff line change
Expand Up @@ -153,6 +153,18 @@ BUSY_TURN_MAX_SECS=${FM_BUSY_TURN_MAX_SECS:-3600}
# These cases re-surface once for a recheck every PAUSE_RESURFACE_SECS - far
# longer than the wedge threshold, but finite so a forgotten hold cannot rot invisibly.
PAUSE_RESURFACE_SECS=${FM_PAUSE_RESURFACE_SECS:-$FM_PAUSE_RESURFACE_SECS_DEFAULT}
# A parked lane - a declared external wait, or a captain hold whose agent has
# confidently exited - cannot clear until something outside this fleet acts, so
# it gets a CHECK-IN CADENCE rather than a standing per-cycle trigger. Each lane
# keeps its own PAUSE_RESURFACE_SECS recheck age, but every lane due in one poll
# is BATCHED into a single wake, and that batch is rate-limited home-wide by the
# .last-parked-checkin mtime - the same restart-surviving, time-based shape the
# slow per-task checks use. Without the batch, parked lanes at independent
# phases each ended a supervision cycle of their own, and a fresh watcher died
# within seconds of every start on whichever lane happened to be next overdue.
# This is cadence, never silence: a parked lane is still re-checked, because a
# declared wait can stop holding.
PARKED_CHECKIN_SECS=${FM_PARKED_CHECKIN_SECS:-$PAUSE_RESURFACE_SECS}
# Consecutive event-path failures (fm_backend_wait_transition returning 2 -
# connect/subscribe failure) before the push fast-path is disabled for the rest
# of this watcher process and the loop reverts to pure polling (report section
Expand Down Expand Up @@ -310,15 +322,18 @@ busy_turn_over_age() { # <task>
}

# Absorb a stale pane under a declared external-wait pause (paused:) or a
# dead-agent captain-held transfer, and re-surface it once every
# PAUSE_RESURFACE_SECS for a recheck so it cannot rot invisibly. Called on any
# dead-agent captain-held transfer, and mark it due for a recheck once every
# PAUSE_RESURFACE_SECS so it cannot rot invisibly. Called on any
# stale poll once pause_state_class permits the bounded cadence, so it must be
# cheap: it NEVER re-reads crew state. The re-surface age is anchored on the
# status file mtime, not a per-hash marker, so a churny idle pane (a ticking
# clock, a token counter) cannot keep resetting the cadence the way a hash-tied
# timer would. A .paused-resurfaced-<key> throttle marker records the last
# re-surface epoch so, once past the window, it fires once per window rather than
# every poll. Advances the stale suppressor to <hash> and flags the key paused.
# re-surface epoch so, once past the window, a lane comes due once per window
# rather than every poll. A due lane is surfaced here only in away mode; in
# normal mode it is collected for flush_parked_checkin, which owns the home-wide
# batch and stamps that marker for exactly the lanes it carries.
# Advances the stale suppressor to <hash> and flags the key paused.
handle_paused_stale() { # <window> <task> <hash>
local win=$1 task=$2 h=$3 key statusf mtime age rf rf_age reason
key=$(printf '%s' "$win" | tr ':/.' '___')
Expand All @@ -333,13 +348,53 @@ handle_paused_stale() { # <window> <task> <hash>
rf_age=$(age_of "$rf") # 999999 when no prior re-surface
if [ "$age" -ge "$PAUSE_RESURFACE_SECS" ] && [ "$rf_age" -ge "$PAUSE_RESURFACE_SECS" ]; then
reason="stale: $win (paused ${age}s, awaiting external - declared pause, rechecked on a long cadence not a wedge; confirm the wait still holds)"
fm_wake_append stale "$win" "$reason" || exit 1
date +%s > "$rf"
wake "$reason"
if afk_present; then
# The away-mode daemon owns triage and classifies one window per printed
# reason, so it keeps the unbatched one-shot form it was built to read.
fm_wake_append stale "$win" "$reason" || exit 1
date +%s > "$rf"
wake "$reason"
fi
# Due, but not surfaced here: collect it and let flush_parked_checkin decide
# once, for the whole fleet, whether this poll is a check-in. One line per
# lane: flush_parked_checkin reads this back a record at a time.
parked_due="${parked_due}${win}"$'\t'"${reason}"$'\n'
triage_log "parked check-in due (paused ${age}s): $win"
return 0
fi
triage_log "absorbed stale (paused, awaiting external, age ${age}s): $win"
}

# Emit at most ONE parked check-in per PARKED_CHECKIN_SECS, carrying every lane
# that came due in this poll. Only the lanes actually included in an emitted
# check-in have their per-lane cadence stamped, so a deferred lane stays due and
# is picked up by the check-in that does fire.
flush_parked_checkin() {
local win reason combined='' first=1 key
[ -n "$parked_due" ] || return 0
if [ "$(age_of "$STATE/.last-parked-checkin")" -lt "$PARKED_CHECKIN_SECS" ]; then
triage_log "deferred parked check-in to its cadence"
return 0
fi
while IFS=$(printf '\t') read -r win reason; do
[ -n "$win" ] || continue
fm_wake_append stale "$win" "$reason" || exit 1
key=$(printf '%s' "$win" | tr ':/.' '___')
date +%s > "$STATE/.paused-resurfaced-$key"
if [ "$first" -eq 1 ]; then
combined="$reason"
first=0
else
combined="$combined; $reason"
fi
done <<EOF
$parked_due
EOF
[ -n "$combined" ] || return 0
touch "$STATE/.last-parked-checkin"
wake "$combined"
}

clear_pause_state() { # <window>
local win=$1 key
key=${win//:/_}
Expand Down Expand Up @@ -856,6 +911,9 @@ EOF
# signature means the crewmate finished, is waiting, or is wedged. Each distinct
# stale hash is surfaced, absorbed, or timed toward escalation once (.stale-*
# remembers the hash already classified).
# Parked lanes found in this pass accumulate here instead of each ending the
# cycle on its own; flush_parked_checkin below decides the batch once.
parked_due=''
while IFS= read -r w; do
kind=$(window_kind "$w")
task=$(window_to_task "$w" "$STATE")
Expand Down Expand Up @@ -1020,6 +1078,11 @@ EOF
fi
done < <(recorded_windows)

# One check-in for every parked lane that came due in this pass, at most once
# per PARKED_CHECKIN_SECS. Runs after the whole window loop so a check-in
# never truncates the pass that produced it.
flush_parked_checkin

# Heartbeat: the watcher runs a cheap fleet-scan at a regular cadence no matter
# what. Time-based via .last-heartbeat mtime; interval doubles per consecutive
# no-change heartbeat (idle fleet) up to HEARTBEAT_MAX, and resets on any
Expand Down
Loading
Loading