Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
19 commits
Select commit Hold shift + click to select a range
5153288
fix(bin): bound recovery-episode reopen so a stuck watcher can stay up
marcus-vsc-moreira Sep 23, 2026
69495b0
docs: make watcher-continuity easier to read (#5608)
tmchow Sep 25, 2026
ebdfc36
fix(bin): refuse watchers from disposable checkouts and exit when the…
karotkriss Sep 25, 2026
31aac60
fix(bin): bound recovery-episode reopen so a stuck watcher can stay up
marcus-vsc-moreira Sep 23, 2026
59a9d16
no-mistakes(review): Keep settled recovery episodes from re-announcin…
marcus-vsc-moreira Sep 25, 2026
ce15ad9
no-mistakes(review): Skip queued-row re-announce only for bound-settl…
marcus-vsc-moreira Sep 25, 2026
9976424
no-mistakes(document): Document reopen limit env var and regression c…
marcus-vsc-moreira Sep 25, 2026
82d8173
no-mistakes(review): Remove duplicated watcher tests and clear settle…
marcus-vsc-moreira Sep 25, 2026
36eec6e
no-mistakes(test): Clear bound-settle sentinel on genuine ack with qu…
marcus-vsc-moreira Sep 25, 2026
754bba3
fix(bin): repair a missing session lock while a harness is still active
marcus-vsc-moreira Sep 28, 2026
713aaf9
no-mistakes(review): Limit guard session-lock repair to the main actor
marcus-vsc-moreira Sep 28, 2026
3bf81b7
no-mistakes(review): Move missing-lock repair out of guard to owner path
marcus-vsc-moreira Sep 28, 2026
7298de8
no-mistakes(test): Fix empty-queue bounded-reopen test to expect watc…
marcus-vsc-moreira Sep 28, 2026
e7f88af
no-mistakes(test): Keep bound-settled recovery episode across watcher…
marcus-vsc-moreira Sep 28, 2026
7d0d915
no-mistakes(ci): I fixed the four Greptile findings. All the tests I …
marcus-vsc-moreira Sep 28, 2026
82e22fb
no-mistakes(review): Lower recovery reopen limit default from 3 to 1
marcus-vsc-moreira Sep 30, 2026
0475f89
no-mistakes(review): Reset reopen budget when append mints fresh episode
marcus-vsc-moreira Sep 30, 2026
52271b1
no-mistakes(review): Clear reopen count only after append succeeds
marcus-vsc-moreira Sep 30, 2026
6780209
no-mistakes(document): Note bounded reopen also settles recovery epis…
marcus-vsc-moreira Sep 30, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions .agents/skills/operational-home-layout/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -102,6 +102,8 @@ state/ runtime records and signals; gitignored
.startup-network.* status, report, per-step elapsed timings, inline-print claim, and lock for the deferred startup stage that runs network checks and the inactive-outcome scan off the digest's blocking path; bin/fm-startup-network.sh
.wake-queue durable queued wakes retained until post-handling acknowledgement: epoch<TAB>seq<TAB>kind<TAB>key<TAB>payload
.watcher-down private generation-bound recovery state coupling watcher downtime, durable wake presentation, and post-handling acknowledgement; never touch
.watcher-down.reopen-count private per-episode reopen counter bounding how many unacknowledged plain restarts reopen the same stuck episode before it settles (docs/watcher-continuity.md "Recovery episode acknowledgement"); never touch
.watcher-down.reopen-settled private generation of the recovery episode the reopen bound last settled, so a watcher start does not re-announce it for queued rows (docs/watcher-continuity.md "Recovery episode acknowledgement"); never touch
.<id>.open-decisions-cursor per-task byte cursor and folded open-decision set bounding the OPEN DECISIONS scan's cost to new status-log appends; written only by fm-classify-lib.sh's status_open_decisions_incremental, removed by teardown, safe to delete (forces one full re-fold)
.<id>.home-appends per-task ledger of byte ranges this home itself appended as bookkeeping closes, so a wake scan can tell its own growth from a foreign write instead of waking on it; presentation is unaffected, so both the signal annotation and UNREAD STATUS still print those lines; written only by fm-classify-lib.sh's status_home_appends_record; its sibling .<id>.home-appends.lock serializes that ledger's read-merge-write; both removed by teardown, safe to delete
.status-presentation-cursor .status-presentation-lock fleet-wide per-task status identity plus independent annotation and outcome-backstop byte offsets, with a serialization lock preventing already-presented lines from replaying while preserving delayed signal annotations; owned by fm-classify-lib.sh, with each task's row retired by teardown
Expand Down
6 changes: 5 additions & 1 deletion bin/fm-wake-drain.sh
Original file line number Diff line number Diff line change
Expand Up @@ -927,7 +927,11 @@ if [ -n "$ACK_THROUGH" ]; then
;;
esac
else
fm_recovery_marker_snapshot "$RECOVERY_MARKER" || exit 1
# Only an ack that consumed rows retires the reopen bound's settle
# distinction; a retried stale ack that consumed nothing must not.
SNAPSHOT_ACK_GENERATION=
[ "$ACK_REMOVED" -gt 0 ] && SNAPSHOT_ACK_GENERATION=$ACK_GENERATION
fm_recovery_marker_snapshot "$RECOVERY_MARKER" "$SNAPSHOT_ACK_GENERATION" || exit 1
RECOVERY_MARKER_TOKEN=$FM_RECOVERY_MARKER_TOKEN
if [ "${RECOVERY_MARKER_TOKEN##*:}" != "$ACK_GENERATION" ]; then
RECOVERY_ACK_MOVED=true
Expand Down
108 changes: 93 additions & 15 deletions bin/fm-wake-lib.sh
Original file line number Diff line number Diff line change
Expand Up @@ -709,12 +709,14 @@ _fm_recovery_marker_write_locked() {
}

# Apply the downtime republication states owned by docs/watcher-continuity.md
# while preserving an outstanding generation-bound acknowledgement.
# while preserving an outstanding generation-bound acknowledgement. A watcher
# lock ending (source close) keeps an episode the reopen bound settled, so the
# next restart does not reopen that same unwatched stretch all over again.
_fm_recovery_marker_publish() {
local marker=$1 kind=${2:-downtime} bound=${3:-} source=${4:-watcher}
local lock saved_token generation='' status=pending previous_append_token=''
case "$kind" in handling|downtime) ;; *) return 1 ;; esac
case "$source" in watcher|append) ;; *) return 1 ;; esac
case "$source" in watcher|close|append) ;; *) return 1 ;; esac
if [ "$source" = append ]; then
FM_WAKE_APPEND_RECOVERY_PREVIOUS_TOKEN=
FM_WAKE_APPEND_RECOVERY_PUBLISHED_TOKEN=
Expand Down Expand Up @@ -748,11 +750,18 @@ _fm_recovery_marker_publish() {
status=pending
;;
announced:downtime:*)
if [ "$source" = watcher ]; then
if [ "$source" != append ]; then
generation=${FM_RECOVERY_MARKER_TOKEN##*:}
status=announced
fi
;;
acked:downtime:*)
if [ "$source" = close ] \
&& [ "$(cat "${marker}.reopen-settled" 2>/dev/null || true)" = "${FM_RECOVERY_MARKER_TOKEN##*:}" ]; then
generation=${FM_RECOVERY_MARKER_TOKEN##*:}
status=acked
fi
;;
esac
fi
FM_RECOVERY_MARKER_TOKEN=$saved_token
Expand Down Expand Up @@ -832,12 +841,19 @@ _fm_recovery_marker_begin_handling() {
fm_lock_release "$lock"
}

# With an acknowledged generation, a snapshot that still names it also retires
# the reopen bound's settle distinction: a genuine ack that leaves newer rows
# queued must keep the acked-plus-queue re-announce on the next arm.
fm_recovery_marker_snapshot() {
local marker=$1 lock
local marker=$1 acked_generation=${2:-} lock
FM_RECOVERY_MARKER_TOKEN=
lock="${marker}.lock"
fm_lock_acquire_wait "$lock" || return 1
fm_recovery_marker_read "$marker" || true
if [ -n "$acked_generation" ] && [ -n "$FM_RECOVERY_MARKER_TOKEN" ] \
&& [ "${FM_RECOVERY_MARKER_TOKEN##*:}" = "$acked_generation" ]; then
rm -f -- "${marker}.reopen-count" "${marker}.reopen-settled" 2>/dev/null || true
Comment on lines +853 to +855

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Empty ack restarts recovery If an older --ack-through command is retried for a settled generation while newer rows remain queued, it can consume zero rows and still reach this snapshot. The matching generation clears the settlement sentinel, so the next arm re-announces the episode and exits instead of keeping the watcher up until the next drain.

Knowledge Base Used: Watch and wake workflows

fi
fm_lock_release "$lock"
}

Expand All @@ -854,7 +870,11 @@ _fm_recovery_marker_ack() {
line=$FM_RECOVERY_MARKER_TOKEN
case "$line" in
pending:*|announced:*) line="acked:${line#*:}" ;;
acked:*) fm_lock_release "$lock"; return 0 ;;
acked:*)
rm -f -- "${marker}.reopen-count" "${marker}.reopen-settled" 2>/dev/null || true
fm_lock_release "$lock"
return 0
;;
*) fm_lock_release "$lock"; return 1 ;;
esac
tmp=$(mktemp "${marker}.tmp.XXXXXX") || { fm_lock_release "$lock"; return 1; }
Expand All @@ -865,6 +885,10 @@ _fm_recovery_marker_ack() {
fm_lock_release "$lock"
return 1
fi
# A genuine acknowledgement proves this episode was actually seen, so a
# later down stretch's reopen starts with a fresh FM_RECOVERY_REOPEN_LIMIT
# budget rather than inheriting this one's count.
rm -f -- "${marker}.reopen-count" "${marker}.reopen-settled" 2>/dev/null || true
fm_lock_release "$lock"
}

Expand All @@ -884,6 +908,7 @@ _fm_recovery_marker_arm_check() {
fm_lock_release "$FM_WAKE_QUEUE_LOCK"
return 1
fi
rm -f -- "${marker}.reopen-count" 2>/dev/null || true
FM_RECOVERY_MARKER_ACTION='recover'
fi
fm_lock_release "$lock"
Expand All @@ -904,6 +929,7 @@ _fm_recovery_marker_arm_check() {
fm_lock_release "$FM_WAKE_QUEUE_LOCK"
return 1
fi
rm -f -- "${marker}.reopen-count" 2>/dev/null || true
FM_RECOVERY_MARKER_ACTION='recover'
fm_lock_release "$lock"
fm_lock_release "$FM_WAKE_QUEUE_LOCK"
Expand All @@ -927,7 +953,8 @@ _fm_recovery_marker_arm_check() {
FM_RECOVERY_MARKER_ACTION='recover'
;;
acked:*)
if [ -s "$FM_WAKE_QUEUE" ]; then
if [ -s "$FM_WAKE_QUEUE" ] \
&& [ "$(cat "${marker}.reopen-settled" 2>/dev/null || true)" != "${line##*:}" ]; then
if ! _fm_recovery_marker_write_locked "$marker" downtime "" announced; then
fm_lock_release "$lock"
fm_lock_release "$FM_WAKE_QUEUE_LOCK"
Expand All @@ -944,9 +971,31 @@ _fm_recovery_marker_arm_check() {

# Apply the owner-documented announced-episode arm transition atomically with
# the queue read. Handling successors must not call this transition.
#
# Nothing else ever retires that generation when no live session runs the
# printed acknowledgement, so a plain restart with no re-arm loop and no
# session reopens the SAME stuck episode into a fresh generation every single
# time, forever: each generation is used for exactly one resurface and then
# abandoned still announced, and the very next restart reopens it again. Bound
# that: past FM_RECOVERY_REOPEN_LIMIT consecutive reopens of one episode with
# no intervening explicit acknowledgement, settle it to acked directly instead
# of minting yet another generation nobody is watching, so a watcher can start
# and stay up. The settle records its generation in .watcher-down.reopen-settled
# so arm-check does not re-announce that bound-settled episode just because the
# queue is non-empty: its queued rows stay durable for the next session's
# drain, while a genuinely acked episode with queued rows still resurfaces.
# A real acknowledgement
# (_fm_recovery_marker_ack), arm-check minting a fresh episode from a missing
# or invalid marker, and a durable append minting a fresh episode from an
# announced one all clear the counter, so this bound never shortens the
# once-per-genuine-generation resurface a live, attentive session relies on.
FM_RECOVERY_REOPEN_LIMIT=${FM_RECOVERY_REOPEN_LIMIT:-1}
case "$FM_RECOVERY_REOPEN_LIMIT" in ''|*[!0-9]*) FM_RECOVERY_REOPEN_LIMIT=1 ;; esac

_fm_recovery_marker_reopen_announced() {
local marker=$1 lock
local marker=$1 lock counter_file count generation
lock="${marker}.lock"
counter_file="${marker}.reopen-count"
fm_lock_acquire_wait "$FM_WAKE_QUEUE_LOCK" || return 1
if ! fm_lock_acquire_wait "$lock"; then
fm_lock_release "$FM_WAKE_QUEUE_LOCK"
Expand All @@ -959,11 +1008,37 @@ _fm_recovery_marker_reopen_announced() {
fi
case "$FM_RECOVERY_MARKER_TOKEN" in
announced:*)
if [ -s "$FM_WAKE_QUEUE" ] \
&& ! _fm_recovery_marker_write_locked "$marker" downtime ""; then
fm_lock_release "$lock"
fm_lock_release "$FM_WAKE_QUEUE_LOCK"
return 1
if [ -s "$FM_WAKE_QUEUE" ]; then
count=$(cat "$counter_file" 2>/dev/null || true)
case "$count" in ''|*[!0-9]*) count=0 ;; esac
count=$((count + 1))
if [ "$count" -gt "$FM_RECOVERY_REOPEN_LIMIT" ]; then
generation=${FM_RECOVERY_MARKER_TOKEN##*:}
if ! _fm_recovery_marker_write_locked "$marker" downtime "$generation" acked; then
fm_lock_release "$lock"
fm_lock_release "$FM_WAKE_QUEUE_LOCK"
return 1
fi
if ! printf '%s\n' "$generation" > "${marker}.reopen-settled" 2>/dev/null; then
fm_lock_release "$lock"
fm_lock_release "$FM_WAKE_QUEUE_LOCK"
return 1
fi
rm -f -- "$counter_file" 2>/dev/null || true
fm_lock_release "$lock"
fm_lock_release "$FM_WAKE_QUEUE_LOCK"
return 0
fi
if ! printf '%s\n' "$count" > "$counter_file" 2>/dev/null; then
fm_lock_release "$lock"
fm_lock_release "$FM_WAKE_QUEUE_LOCK"
return 1
fi
if ! _fm_recovery_marker_write_locked "$marker" downtime ""; then
fm_lock_release "$lock"
fm_lock_release "$FM_WAKE_QUEUE_LOCK"
return 1
fi
fi
;;
esac
Expand All @@ -988,7 +1063,7 @@ fm_recovery_transition() {
;;
release-lock)
[ -n "$target" ] || return 1
_fm_recovery_marker_publish "$marker" "${value:-downtime}" "$bound" || return 1
_fm_recovery_marker_publish "$marker" "${value:-downtime}" "$bound" close || return 1
fm_lock_release "$target"
;;
release-lock-existing)
Expand All @@ -1008,7 +1083,7 @@ fm_recovery_transition() {
;;
clear-stale-lock)
[ -n "$target" ] || return 1
_fm_recovery_marker_publish "$marker" "${value:-downtime}" "$bound" || return 1
_fm_recovery_marker_publish "$marker" "${value:-downtime}" "$bound" close || return 1
fm_lock_remove_path "$target"
;;
*) return 2 ;;
Expand Down Expand Up @@ -1169,7 +1244,7 @@ fm_lock_try_acquire() {
fi

if [ "$lockdir" = "$STATE/.watch.lock" ] \
&& ! _fm_recovery_marker_publish "$STATE/.watcher-down" downtime; then
&& ! _fm_recovery_marker_publish "$STATE/.watcher-down" downtime "" close; then
fm_lock_release "$steal"
FM_LOCK_HELD_PID=$cur
FM_LOCK_OWNER_DIR=
Expand Down Expand Up @@ -2038,6 +2113,9 @@ fm_wake_append_locked() {
if [ "$status" -ne 0 ]; then
_fm_wake_append_recovery_restore_locked || true
else
case "$FM_WAKE_APPEND_RECOVERY_PREVIOUS_TOKEN" in
announced:downtime:*) rm -f -- "${recovery_marker}.reopen-count" 2>/dev/null || true ;;
esac
FM_WAKE_APPEND_RECOVERY_PREVIOUS_TOKEN=
FM_WAKE_APPEND_RECOVERY_PUBLISHED_TOKEN=
fi
Expand Down
1 change: 1 addition & 0 deletions docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -2333,6 +2333,7 @@ FM_WATCH_ARM_RETIRE_TIMEOUT_MS=1000 # milliseconds Pi/OpenCode wait for an unr
FM_WATCH_REARM_RETRY_BASE_MS=250 # Pi/OpenCode adapter base delay for continuity restoration retries
FM_WATCH_REARM_RETRY_MAX_MS=4000 # Pi/OpenCode adapter cap for exponential continuity retry delay
FM_WATCH_REARM_RETRY_LIMIT=5 # Pi/OpenCode adapter launch-failure retries before surfacing restoration failure
FM_RECOVERY_REOPEN_LIMIT=1 # consecutive unacknowledged watcher restarts that reopen one stuck recovery episode before it settles, so the watcher stays up (docs/watcher-continuity.md "Bounded reopen")
FM_WATCH_CYCLE_LOG_MAX_BYTES=262144 # size cap for the arm-owned watcher lifecycle ledger
FM_WATCH_CYCLE_LOG_KEEP_LINES=1000 # newest complete lifecycle rows considered when the ledger is capped
FM_WATCHER_STALE_GRACE=300 # defaults to FM_GUARD_GRACE if set, else the poll-derived grace (docs/turnend-guard.md "Guard grace and the poll cadence"); seconds before a fresh arm refuses a live holder's stale beacon (attached arms: FM_WATCHER_STALL_BOUND)
Expand Down
15 changes: 14 additions & 1 deletion docs/watcher-continuity.md
Original file line number Diff line number Diff line change
Expand Up @@ -206,7 +206,7 @@ In its `--claude` mode it cooperates with the auto-arm.
## Recovery episode acknowledgement

A recovery episode is one generation of the `state/.watcher-down` marker.
It is retired only by the generation-bound acknowledgement the drain prints as `WAKE_ACK_REQUIRED`.
It is retired by the generation-bound acknowledgement the drain prints as `WAKE_ACK_REQUIRED`, or settled by the bounded reopen below when nobody runs that acknowledgement.
The away return brief treats a still-open handling episode as a wake in progress, not watcher downtime; an open downtime episode remains a gap.

### Announcement
Expand All @@ -217,6 +217,18 @@ A non-successor watcher start checks the durable queue and recovery marker under
If an announced-but-unacknowledged episode has an empty queue, the arm leaves that generation announced, making repeated empty-queue arms idempotent while a long-poll source is merely alive.
If a durable row arrived after the announcement, the arm opens a fresh pending downtime generation so buried work still resurfaces once.

### Bounded reopen

Nothing else ever retires that generation when no live session runs the printed acknowledgement.
So a plain restart with no re-arm loop and no session would otherwise reopen the same stuck episode into a fresh generation forever, one resurface-then-exit cycle per restart.
`state/.watcher-down.reopen-count` bounds that.
Past `FM_RECOVERY_REOPEN_LIMIT` (default 1) consecutive reopens of one episode with no intervening explicit acknowledgement, the next reopen settles the episode to acked directly instead of minting another generation, so the watcher can finally start and stay up.
The settle records that generation in `state/.watcher-down.reopen-settled`.
A watcher start does not re-announce that bound-settled episode just because the queue is non-empty, so its queued rows cannot make the watcher exit; they stay durable and the next session's drain presents them.
When a watcher's lock later ends (a `--restart` retiring it, or a stale lock cleared), that close keeps the bound-settled episode instead of publishing a fresh downtime generation, so the next restart stays up too.
A genuinely acknowledged episode with queued rows still re-announces and resurfaces them on the next arm.
A real acknowledgement, a watcher start that mints a fresh episode from a missing or invalid marker, or a successful durable append that mints a fresh episode from an announced one clears the counter, so this bound never shortens the once-per-genuine-generation resurface a live, attentive session relies on.

### Generation reuse

An ordinary watcher close attempts to publish downtime, and every durable queue append publishes it.
Expand Down Expand Up @@ -435,6 +447,7 @@ They also prove that a legacy or handoff-phase watcher marker from an absent rep
- Decision-only OPEN DECISIONS recovery.
- Interrupted handling replay.
- Generation-bound acknowledgement.
- The bounded reopen of a stuck unacknowledged episode, with and without queued rows, and a genuine acknowledgement with a queued row that still resurfaces.
- A persistent live successor after recovery.
- An idle live Lavish source that stays quiet until its real result wakes promptly.
- An append that reopens an announced empty recovery.
Expand Down
Loading