Skip to content

fix(pi): restore watcher continuity across successor gaps and make extension log opt-in - #5489

Merged
kunchenguid merged 11 commits into
kunchenguid:mainfrom
RibatTRW:fm/firstmate-watcher-successor-gap-fix
Oct 3, 2026
Merged

kunchenguid merged 11 commits into
kunchenguid:mainfrom
RibatTRW:fm/firstmate-watcher-successor-gap-fix

Conversation

@RibatTRW

@RibatTRW RibatTRW commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Intent

Work on the merge conflict for #5489.

Context needed to read the ask: that pull request, "fix(pi): restore watcher continuity across successor gaps and make extension log opt-in", comes from the fork RibatTRW/firstmate, head branch fm/firstmate-watcher-successor-gap-fix (head 3b237f6), base main, and closes #5492.
It treats a dead-but-unclosed arm child as an empty slot via liveArmChild so repair and retry start a fresh arm, pins handling confirmation to the restoration's own recovery token so a superseded generation is delivered without a failure appendix and nothing is retired unless the failed token names the exact current pid and generation, treats an already-acknowledged handling or downtime episode as a confirming no-op when the caller names its generation, and makes the Pi extension diagnostic log state/.watch-extension.log opt-in and off by default (only a positive FM_WATCH_EXTENSION_LOG_KEEP_LINES enables it).
GitHub now reports the pull request as conflicting with main, so it cannot merge.
The pull request's existing behavior must be preserved while bringing it up to date with current main.

Yes lets do the recommended fixes

Those recommended fixes are items 1-4 of a code review of that pull request:

  1. A superseded handling delivery (the confirmation reports a generation mismatch) bypasses the Pi supervision branch and lands on main, while a confirmed delivery is offered to the branch first. A superseded delivery should route like a confirmed one - offered to the Pi supervision branch, with no failure appendix, retiring nothing - with a test that uses an accepting branch bus to assert the superseded wake is offered to the branch.
  2. docs/watcher-continuity.md now says every adapter (Pi, omp, and OpenCode) confirms against the restoration's own recovery token, retries against that same token, and retires a successor only when the failed token names that exact pid and generation, but only Pi changed; omp and OpenCode still confirm against the current successor and retire it whenever the watcher pid is dead. Scope those sentences to Pi (or port the fix to omp and OpenCode).
  3. Two claimed protections are untested: the narrowed retire guard (a failed confirmation must not retire a newer healthy arm that replaced the successor), and dead-but-unclosed slot handling in the scheduled retry and deferred-close paths (a retry must start a fresh arm over a dead-but-unclosed child). Add the missing tests, each failing when its fix is reverted.
  4. test_handling_delivered_rejects_a_superseded_generation asserts something other than its label (its closing step is a plain arm check, not the handling-successor path, and it pins behavior that already exists on main): fix it or relabel it and its docs bullet. Its dead-pid rejection case uses pid 1, which is always alive; use a genuinely dead, reaped pid instead.

What Changed

  • Pi extension (.pi/extensions/fm-primary-pi-watch.ts) now treats a dead-but-unclosed arm child as an empty slot (liveArmChild) in repair, scheduled retry and deferred close. Handling confirmation and its single retry use the restoration's own recovery token. A generation mismatch (exit status 3) is reported as superseded and is offered to the Pi supervision branch like a confirmed delivery, with no failure appendix and nothing retired. A failed confirmation retires the current arm only when the failed token names its exact pid and generation and that pid is dead.
  • fm-wake-lib.sh treats an already-acknowledged handling or downtime episode as a confirming no-op when the caller names its generation. Without a named generation it is still rejected. The Pi extension diagnostic log state/.watch-extension.log is now opt-in and off by default. Only a positive FM_WATCH_EXTENSION_LOG_KEEP_LINES enables it, and the new knob is documented in docs/configuration.md.
  • docs/watcher-continuity.md scopes the recovery-token confirm, retry and narrowed retire guard to Pi, since omp and OpenCode still confirm against the current successor. The extension and arm test suites cover:
    • the superseded wake being offered to an accepting branch bus;
    • the narrowed retire guard sparing a newer healthy arm;
    • a repair, retry and deferred close starting a fresh arm over a dead-but-unclosed child;
    • the already-acknowledged no-op, with a genuinely dead, reaped pid used for the dead-pid rejection case.

Risk Assessment

✅ Low: The change is well-bounded to the Pi extension and the handling-confirmation marker, and I found no concrete defect. Superseded deliveries now route like confirmed ones, the retire guard matches the recorded token including generation, and docs are scoped to Pi. Tests were added for each of the four requested items, but I did not run them.

Testing

I ran the two targeted test files that exercise this change, and both pass with exit 0. They cover superseded delivery being offered to the branch, dead-child fresh-arm restart on repair/retry/deferred close, the opt-in log, and handling confirmation. I did not run a real Pi or Herdr lab session, so no scenario is recorded as a live pass. I also did not revert each fix to confirm its test fails.

  • Live validation: ⚠️ inconclusive - 0 of 4 scenarios driven live against the product
Scenario Result Live Evidence
Superseded handling delivery is offered to the Pi supervision branch like a confirmed one, with no failure appendix ⏸️ untested no The prior payload only ran this through a test harness with a fake bus, not a live Pi session, so it did not establish a live result. Run a real Pi session under Herdr to drive it.
A dead-but-unclosed arm child is treated as an empty slot, so repair, scheduled retry and deferred close each start a fresh arm ⏸️ untested no The prior payload only ran this through a test harness, not a live Pi session, so it did not establish a live result. Run a real Pi session under Herdr to drive it.
Handling confirmation: an already-acknowledged confirmation is a no-op, and a churned generation reports a mismatch while an arm check keeps it ⏸️ untested no The prior payload ran this against fm-watch-arm.sh in a temp home, not a live fleet, so it did not establish a live result. Run against a live fleet to drive it.
The Pi extension diagnostic log stays off unless FM_WATCH_EXTENSION_LOG_KEEP_LINES is positive ⏸️ untested no The prior payload only ran this through a test harness, not a live Pi session, so it did not establish a live result. Run a real Pi session under Herdr to drive it.
Evidence: Pi extension test log

Source: Pi extension test log

ok - Pi extension reports external healthy watcher output
ok - Pi custom tool exposes repair-only metadata and returns automatic-continuation guidance
ok - Pi redundant tool call returns ownership guidance and spawns no second child
ok - Pi scheduled retry remains extension-owned after another tool call
ok - Pi actionable close starts one successor before wake delivery settles
ok - Pi actionable output waits for predecessor close before successor restoration
ok - Pi dispatcher branch offer owns accepted wakes and falls back to main
ok - Pi dispatcher flags a fleet-wide heartbeat offer as branch-eligible
ok - a co-present check row neither vetoes nor rides a heartbeat into main
ok - every main-only check class still reaches main, never the supervision branch
ok - a captain-held signal trigger reaches main with routine rows present
ok - an unread pending-reply escalation keeps later stale aliases on main
ok - a mixed batch of two distinct files - one routine, one needs-decision - routes wholly to main
ok - a co-present needs-decision row neither vetoes nor rides a heartbeat into main
ok - heartbeat restoration failure stays on main
ok - watcher-failure repair stays with main even with a live, accepting branch listener
ok - under the away-posture record every actionable row is offered to the branch while broken-queue wakes and watcher-failure alarms still reach main
ok - Pi refused handling handshake is classified and not swallowed
ok - Pi confirm failure retires the named arm with distinct watcher pid
ok - Pi confirm failure for a stale successor spares the newer arm
ok - Pi superseded handling delivery carries no rejection appendix and is logged
ok - Pi superseded handling delivery is offered to the branch like a confirmed one
ok - Pi extension diagnostic log stays off unless opted in
ok - Pi repair starts a fresh arm instead of no-opping on a dead child
ok - Pi scheduled retry starts a fresh arm instead of stalling on a dead child
ok - Pi deferred close starts a fresh arm instead of stalling on a dead child
ok - Pi hung successor falls back to one typed actionable wake
ok - Pi unretired successor falls back without an overlapping retry
ok - Pi late unretired closes resume classified supervision
ok - Pi clean empty close triggers a bounded continuity retry
ok - Pi established clean closes stop at the configured retry limit
ok - Pi close handler verifies session-lock ownership before successor launch
ok - Pi watcher arm distinguishes all session lock ownership states
ok - Pi session transitions auto-arm through a generation owner across /new /resume /fork/reload, stale callbacks, and quit
ok - Pi session replacement auto-arms and carries its in-flight actionable close
ok - Pi replacement replays a streaming follow-up before consumption
ok - Pi streaming-time wake delivery keeps the successor chain and replays only unconsumed wakes
ok - Pi retries a verified successor that failed during wake delivery once that delivery settles
ok - Pi replacement receives actionable closes after retirement timeout
ok - Pi replacement handoff tokens stay unique across fresh modules
ok - Pi replacement persistence failure keeps its predecessor until a successor commits
ok - Pi process-exit cleanup listener remains singular across session replacement
ok - Pi process-exit cleanup stops the attached arm child
ok - OpenCode plugins have an explicit ESM boundary even under a typeless parent package
ok - OpenCode watcher plugin uses the effective FM_HOME state
ok - OpenCode watcher plugin sources the effective config
ok - OpenCode watcher plugin requires session lock ownership
ok - OpenCode watcher coordinator respects primary scope
ok - OpenCode watcher plugin starts one successor before wake prompt delivery settles
ok - OpenCode watcher plugin runs the supervision host on an opted-in home and relays every host line (away record)
ok - OpenCode watcher plugin runs the supervision host on an opted-in home and relays every host line (quiet record)
ok - OpenCode pre-ready actionable close preserves its successor
ok - OpenCode hung successor falls back to one typed actionable wake
ok - OpenCode unretired successor falls back without an overlapping retry
ok - OpenCode late unretired closes resume classified supervision
ok - OpenCode clean empty close triggers a bounded continuity retry
ok - OpenCode established clean closes stop at the configured retry limit
ok - OpenCode close handler verifies session-lock ownership before successor launch
ok - OpenCode watcher plugin coordinates with the turn-end guard
ok - OpenCode healthy arm output does not suppress the turn-end guard
rc=0
Evidence: watch-arm test log

Source: watch-arm test log

ok - watch-arm: an attached arm reports the wake its cycle delivered instead of a false failure
ok - watch-arm: a delivered wake consumed by the handling turn still closes the attached arm cleanly
ok - watch-arm: an unusable launch confirm window refuses to arm by name
ok - watch-arm: a disposable validation checkout refuses to arm
ok - watch-arm: a watcher exits when its state directory is removed
ok - watch-arm: a watcher exits when its home is removed
ok - watch-arm: the test reaper stops a watcher armed for a tracked temporary home
ok - watch-arm: a cycle that delivered no wake of its own still fails loudly
ok - watch-arm: an attached arm keeps following a slow live holder and reports its wake
ok - watch-arm: an attached arm hands a holder stalled past the bound to its owner's replacement
~/.no-mistakes/worktrees/f070a5b5bfb8/01M3ZKCH5EQGVYDFA762D0AA7M/bin/fm-watch-arm.sh: line 766: 1670150 Killed                     "$WATCH" > "$child_out"
WAKE_ACK_REQUIRED: after handling completes run bin/fm-wake-drain.sh --ack-through 2 --recovery-generation 1670710.1790988443.MrYnb5
ok - watch-arm: re-arm surfaces every queued wake and an open remote decision after downtime
~/.no-mistakes/worktrees/f070a5b5bfb8/01M3ZKCH5EQGVYDFA762D0AA7M/bin/fm-watch-arm.sh: line 766: 1677678 Killed                     "$WATCH" > "$child_out"
WAKE_ACK_REQUIRED: after handling completes run bin/fm-wake-drain.sh --ack-through 1 --recovery-generation 1678271.1790988448.knOAoN
ok - watch-arm: a re-arm whose recovery cycle runs slowly still surfaces it
watcher: recovery state could not be persisted; retaining stale lock evidence
ok - watch-arm: marker publication failure retains stale-lock recovery evidence
WAKE_ACK_REQUIRED: after handling completes run bin/fm-wake-drain.sh --ack-through 2 --recovery-generation 1685671.1790988464.8ULsOT
WAKE_ACK_REQUIRED: after handling completes run bin/fm-wake-drain.sh --ack-through 3 --recovery-generation 1688670.1790988465.tXD6vI
ok - watch-arm: a wake queued after handling drain is recovered once at successor arm
ok - watch-arm: interrupted handling leaves its wake durable for successor re-drain
WAKE_ACK_REQUIRED: after handling completes run bin/fm-wake-drain.sh --ack-through 0 --recovery-generation 1697028.1790988471.cKj7tL
ok - watch-arm: malformed recovery state is quarantined without a successor loop
WAKE_ACK_REQUIRED: after handling completes run bin/fm-wake-drain.sh --ack-through 1 --recovery-generation 1699528.1790988472.7nmsPk
ok - watch-arm: publication after recovery handoff is surfaced
ok - watch-arm: restart publishes recovery before clearing a reused-pid watcher lock
WAKE_ACK_REQUIRED: after handling completes run bin/fm-wake-drain.sh --ack-through 7 --recovery-generation 1702103.1790988475.79plo4
ok - watch-arm: markerless legacy queues are adopted and recovered
ok - watch-arm: an idle Lavish source stays quiet and its real result wakes promptly
ok - watch-arm: appending work reopens an announced empty recovery
ok - watch-arm: a watcher close during handling keeps the printed acknowledgement valid
ok - watch-arm: a moved recovery generation consumes handled rows and names its remedy
ok - watch-arm: downtime marker publication does not follow symlinks
ok - watch-arm: --stop ends only this home's watcher, publishes downtime, and reports when none runs
ok - watch-arm: an already-acknowledged handling confirmation succeeds as a no-op
ok - watch-arm: a churned generation's handling confirmation reports a mismatch and an arm check keeps it
ok - watch-arm: --take-over attaches to a cycle the named arm does not own and leaves it running
ok - watch-arm: --take-over owns a fresh cycle without a recovery wake and still surfaces queued work
watcher: secondmate liveness check failed
ok - watch-arm: takeover preserves self-exit downtime and surfaces a recovery wake
rc=0
- Outcome: ⚠️ 1 warning across 1 run (5m58s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

⚠️ **Test** - 1 warning
  • ⚠️ live validation verdict: inconclusive (0 of 4 scenarios were driven live against the product); untested: Superseded handling delivery is offered to the Pi supervision branch like a confirmed one, with no failure appendix, A dead-but-unclosed arm child is treated as an empty slot, so repair, scheduled retry and deferred close each start a fresh arm, Handling confirmation: an already-acknowledged confirmation is a no-op, and a churned generation reports a mismatch while an arm check keeps it, The Pi extension diagnostic log stays off unless FM_WATCH_EXTENSION_LOG_KEEP_LINES is positive
  • Live validation: ⚠️ inconclusive - 0 of 4 scenarios driven live against the product
Scenario Result Live Evidence
Superseded handling delivery is offered to the Pi supervision branch like a confirmed one, with no failure appendix ⏸️ untested no The prior payload only ran this through a test harness with a fake bus, not a live Pi session, so it did not establish a live result. Run a real Pi session under Herdr to drive it.
A dead-but-unclosed arm child is treated as an empty slot, so repair, scheduled retry and deferred close each start a fresh arm ⏸️ untested no The prior payload only ran this through a test harness, not a live Pi session, so it did not establish a live result. Run a real Pi session under Herdr to drive it.
Handling confirmation: an already-acknowledged confirmation is a no-op, and a churned generation reports a mismatch while an arm check keeps it ⏸️ untested no The prior payload ran this against fm-watch-arm.sh in a temp home, not a live fleet, so it did not establish a live result. Run against a live fleet to drive it.
The Pi extension diagnostic log stays off unless FM_WATCH_EXTENSION_LOG_KEEP_LINES is positive ⏸️ untested no The prior payload only ran this through a test harness, not a live Pi session, so it did not establish a live result. Run a real Pi session under Herdr to drive it.
  • bash tests/fm-pi-watch-extension.test.sh (60 ok, exit 0)
  • bash tests/fm-watch-arm.test.sh (30 ok, exit 0)
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

@RibatTRW
RibatTRW force-pushed the fm/firstmate-watcher-successor-gap-fix branch from 34ebdaa to 7d0005b Compare September 24, 2026 05:00
@RibatTRW RibatTRW changed the title fix: repair Pi watcher successor-gap handling confirmations and add bounded extension log fix: restore Pi watcher continuity across successor gaps and dead arm slots Sep 24, 2026

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate: stamped waiting-ci.

HEAD 7d0005b441498c0c34395eb0011b3d52ec5dae04. Attestation MATCH. Require no-mistakes SUCCESS. CI 35958032376 approved and in progress. MERGEABLE/UNSTABLE. Closes #5492 verified in body (issue open; defects match tip). Fork workflows approved after safe diff review.

contract-class: new-default (split: repair paths are restore; always-on log is new-default).

  • Restore parts (tip vs main): dead-but-unclosed arm child gated via liveArmChild so repair/retry recover; handling confirmation pinned to restoration's own recovery token (superseded ≠ rejected; retire only on exact pid+generation match); already-acked confirmation is a no-op when generation is named — these repair a broken continuity path from Pi watcher continuity: successor-gap confirmations and a dead-but-unclosed arm child block repair/retry recovery #5492.
  • New-default part: appendExtensionLog writes always-on to state/.watch-extension.log; FM_WATCH_EXTENSION_LOG_KEEP_LINES defaults to 200, and positiveInteger treats <=0 as fallback 200 so logging cannot be turned off — new default-on observer surface (FM-LEARN-4627). Not opt-in.

Overall class for auto-merge: new-default — will not auto-merge even when green. Firstmate-flag deferred until otherwise ready (CI still running). Help the PR; no competing PR.

VISION: One captain/interface — repair aligns; always-on log adds below-deck noise surface (partial); Authority aligns (no new autonomy grant); Scripts/judgment aligns; Restart — continuity repair aligns; Delegation n/a; Fleet/vendor — Pi adapter repair aligns; Scope — watcher continuity aligns, new diagnostic file is scope creep. Overall: aligns on repair, does not auto-admit for the new default-on log.

Accept an already-acknowledged handling confirmation as a no-op when the
generation matches, confirm the restoration's own recovery token with a
superseded (not rejected) outcome on generation mismatch, retire an arm on
confirm failure only when the failed token names that exact pid, and record
restore attempts, readiness timeouts, and confirm results in the bounded
state/.watch-extension.log. Regression tests: already-acked no-op plus
mismatch/dead-pid/lock-mismatch rejections and the manual-restart churn
contract in fm-watch-arm.test.sh, and a mid-restore marker advance with no
rejection appendix in fm-pi-watch-extension.test.sh.
startArm and scheduleRetry answered unchanged while holding a ChildProcess
whose OS process was already gone but whose close had not fired, so neither
the repair tool nor the retry timer started anything until that close fired.
Gate slot occupancy on a liveness check (exit/signal codes plus pid probe)
and start a fresh arm instead, with a regression test driving the repair
tool against a dead-but-unclosed child.
Only a positive FM_WATCH_EXTENSION_LOG_KEEP_LINES enables
state/.watch-extension.log. Unset, empty, non-numeric, zero, and
negative values disable logging entirely, so the default run writes
nothing and never creates the file. The shared positiveInteger
fallback semantics stay untouched for the retry and timeout knobs.
docs/configuration.md owns the knob contract and
docs/watcher-continuity.md points at it. Tests: the superseded-delivery
case runs opted in, and a new case proves unset, zero, and non-numeric
values create no log file while delivery still succeeds.
@RibatTRW
RibatTRW force-pushed the fm/firstmate-watcher-successor-gap-fix branch from 7d0005b to 0037e9e Compare September 28, 2026 03:00
@RibatTRW RibatTRW changed the title fix: restore Pi watcher continuity across successor gaps and dead arm slots fix(pi): restore watcher continuity across successor gaps and make extension log opt-in Sep 28, 2026
@greptile-apps

greptile-apps Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[High risk] Modifies watcher continuity and recovery logic for process lifecycle.

The PR appears safe to merge; the latest changes preserve durable wake delivery and watcher ownership while closing the identified successor-gap cases.

Reviews (4) · Last reviewed commit: "Route superseded Pi deliveries like conf..."

… no-mistakes run 36372069208) show conclusion action_required with 0 jobs and no logs because this is a fork PR (RibatTRW/firstmate) and GitHub is holding the workflow runs awaiting maintainer approval; that gate is external to the code and needs a maintainer to approve the runs. While verifying the change locally I found a real defect the approved CI run would hit: the PR's new test test_handling_delivered_rejects_a_superseded_generation failed deterministically. Invariant violated: reopen-announced (a non-successor/manual arm start) mints a fresh recovery generation only when the durable wake queue holds unrecovered work; an announced episode with an empty queue must be left untouched so idle arm starts never churn generations (the [ -s queue ] guard from kunchenguid#4819, relied on by bin/fm-watch.sh:2413 and covered by the append-reopens and announcement-bound sibling tests). The test called reopen with an empty queue and expected churn, so the fix establishes the queued-work precondition (append one wake, re-announce, re-read the generation) before asserting the reopen mints and the old confirmation mismatches. Test-only change, 14 lines in tests/fm-watch-arm.test.sh. Verified: fm-watch-arm.test.sh 25/25 ok on two consecutive runs, fm-pi-watch-extension.test.sh 55/55 ok, and shellcheck reports only one pre-existing warning outside the edited region
@RibatTRW

Copy link
Copy Markdown
Contributor Author

Replying to the scope review on the diagnostic log: it is now opt-in and default-off. Setting FM_WATCH_EXTENSION_LOG_KEEP_LINES to a positive line count enables state/.watch-extension.log; unset, empty, non-numeric, zero, or negative values disable it entirely with no fallback, and a disabled log returns before touching the filesystem, so the file is never created. The watcher-continuity repair behavior (dead-slot gating, recovery-token confirmation, already-acked no-op) is unchanged. This comment was written with AI assistance; the change is covered by the repo regression suites and the pipeline's recorded live validation.

@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate: whole thread re-read (prior waiting-ci new-default stamp on 7d0005b4; author reply that log is now opt-in default-off; tip rebound).

HEAD 3b237f6a284f98fd785489ae6028f79c4d4096d9. Attestation MATCH. Tip NM 36374833255 / 36373277908 SUCCESS. Tip CI 36373277882 in_progress after this-pass approve. Author RibatTRW not blocked. MERGEABLE pending checks.

Closes #5492: body Fixes #5492 — accurate (dead-but-unclosed arm, wrong-token confirm, already-acked rejection).

Contract-class: restore (own tip-vs-main; FM-LEARN-CLAIMS / FM-LEARN-4627). Continuity repair restores the already-specified watcher recovery path: liveArmChild empty-slot gating, confirmation pinned to restoration's own recovery token (superseded ≠ rejected; retire only on exact pid+generation), already-acked confirmation no-op when generation is named. The diagnostic log is now opt-in and default-off (FM_WATCH_EXTENSION_LOG_KEEP_LINES positive enables; unset/0/negative disable and never touch the filesystem) — no longer the always-on observer that previously forced new-default. Overall auto-merge class once green: restore (opt-in log does not block).

VISION.md (each rule)

  1. One captain, one interface — aligns (continuity holds; opt-in log stays below deck when off).
  2. Authority explicit — aligns (log gated; no new autonomy grant).
  3. Scripts own mechanics — aligns (arm/confirm/retire stay scripted).
  4. Restart non-event — aligns (repair restores durable watcher continuity).
  5. Delegation with a spine — aligns (n/a beyond supervision continuity).
  6. Fleet outlives vendor — aligns (Pi adapter repair + version-pinned reading debt unchanged).
  7. Scope — aligns (watcher continuity; diagnostic file only when opted in).

Outcome: waiting-ci. Firstmate flag no. Workflow approvals this pass: CI 36373277882, NM 36373277908, NM 36374833255. If tip goes fully green + MATCH + CLEAN, eligible for auto-merge (restore).

…successor-gap-fix

# Conflicts:
#	docs/configuration.md
#	docs/watcher-continuity.md
#	tests/fm-watch-arm.test.sh
…re guard

A superseded handling confirmation now falls through to the normal delivery
path, so an accepting supervision branch owns the wake instead of main.
Scope the watcher-continuity token-pinned confirmation and narrowed retire
rule to Pi, since omp and OpenCode still confirm against the current
successor. Add tests that fail when the retire guard, the scheduled-retry
gate, or the deferred-close gate is reverted, relabel the churned-generation
characterization test, and use a reaped pid for the dead-pid rejection.
@kunchenguid
kunchenguid merged commit e31bc6e into kunchenguid:main Oct 3, 2026
20 checks passed
@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate: this is merged. Thank you @RibatTRW — really appreciate you taking the time on this.

bramdokman pushed a commit to bramdokman/firstmate that referenced this pull request Oct 3, 2026
Upstream advanced by one commit (kunchenguid#5489, Pi watcher continuity across
successor gaps and an opt-in extension log) while this sync was being
verified. It merges cleanly: the only fork-touched path it changes is
docs/configuration.md, where both sides are kept, and the fork delta
against upstream is unchanged.
yehezkieled added a commit to yehezkieled/firstmate that referenced this pull request Oct 5, 2026
* fix: reclaim orphaned watcher arms on the next park (kunchenguid#6335)

* fix(bin): take over the watcher cycle a main-only pass-through leaves

An attended main-only pass-through leaves a successor watcher cycle
running through main's handling turn. The session's next park attached
to that cycle instead of owning it, so the successor's arm, orphaned by
its host's exit, kept owning the watcher while the new park's arm polled
it twice a second until the next close or the park boundary, hours later
in a quiet second mate. A second-mate restart hit this every time, since
its persist request is a main-only close.

The host now records the successor it leaves for main, and the next
host's first cycle runs bin/fm-watch-arm.sh --take-over on it: when that
arm still owns the healthy watcher, the new arm stops it, reports a
reason the cycle delivered first, and otherwise owns a fresh cycle. The
stop's own downtime publication is undone over an acknowledged episode
when no wake was appended in between, so the handover wakes nobody.

* no-mistakes(review): Keep left-arm record until the orphaned arm is gone

* no-mistakes(review): Relinquish successor arm only after durably recording it

* no-mistakes(review): Relinquish successor only after its record reads back

* no-mistakes(document): Clarify successor takeover guarantees and authoritative documentation

* no-mistakes(ci): Fixed ci-1: acknowledgement restore now requires the exact taken-over arm/watcher ledger row with signal=TERM, awaited within a short bound. Otherwise takeover proceeds without erasing downtime. Added a self-exit regression confirmed failing before the fix and passing afterward; ordinary takeover tests and the full watcher-arm suite pass. Updated Generation reuse documentation. Syntax, diff checks, and ShellCheck pass with existing SC1091/SC2034 warnings excluded. ci-2 remains unchanged

* no-mistakes(document): Clarify watcher take-over recovery and restart limits

* fix(bin): restore downtime on supervision-host hand-back when the successor already closed (kunchenguid#6355)

* fix(bin): restore supervision host hand-back continuity

* no-mistakes(review): Scope host hand-back failure fallback to lost pending:handling

* no-mistakes(review): Remove stray scratch test copy tests/.tmp-rest.test.sh

* no-mistakes(test): Initialise successor globals so early hand-back survives set -u

* no-mistakes(document): Document host hand-back downtime failure and Claude lost-handback notice

* no-mistakes(ci): I fixed the Greptile finding. The rule that must hold: when the supervision host hands back an actionable wake, its rewake is refused, and no watcher is healthy, the hand-back still has to reach main as a delivered notice. That must be true whether the recovery marker is `pending:handling` or `announced:handling`. Only one place applies this check: the lost hand-back fallback in `bin/fm-claude-stop-autoarm.sh`. **Fix:** that check now accepts both `pending:handling:*` and `announced:handling:*` tokens (a one-line change). Nothing else in the fallback changed: - Refusals on any other marker, such as an already acknowledged one, still exit 0 silently and open no failure episode. - The notice is still sent once per episode, and repeats are recorded as `failed-suppressed`. **Tests:** - `tests/fm-claude-stop-autoarm.test.sh`: the lost hand-back test now runs as a shared helper with two variants, one writing a `pending:handling` marker and a new one writing `announced:handling` (`test_host_lost_announced_handback_notifies_once_per_episode`). - `tests/fm-supervision-host.test.sh`: the end-to-end test where downtime restoration fails is now a shared helper too, with a new `announced` variant (`test_claude_stop_hook_notifies_when_closed_announced_successor_downtime_restore_fails`). It moves the handling episode to `announced` before the host hands back. Without the fix the hook would exit 0 here; the test requires exit 2, `outcome=failed` and a delivered failure notice. **Verification (all under nice -n 10):** - The full `tests/fm-claude-stop-autoarm.test.sh` suite passed (rc=0), including both lost hand-back variants and the benign-refusal test. - In `tests/fm-supervision-host.test.sh` I ran only the four hand-back test functions, all passing (rc=0). The suite can't run single functions, so I used a temporary copy with a trimmed test list and deleted it afterwards; `git status` shows only the 3 intended files changed. - shellcheck is clean on all three changed files. I did not run `bin/fm-lint.sh`. - I did not run the new tests against the unfixed code; the claim that they fail without the fix comes from reading the old check

* fix: reduce remote-job polling process churn (kunchenguid#6363)

* perf: cut remote-job idle process creation in the three hot loops

Post-update host measurement still attributes most idle churn to three
per-sample loops: result-consumer state reads and date calls, the delta
reader's capture/hash pass on every poll, and the lane preemption scan's
per-field pipelines. This drops each to its minimum without touching the
contracts around them.

* fm_remote_job_read_state gains an optional result-variable form backed
  by fm_remote_job_read_line, a builtin-only bounded record read (regular
  non-symlink file, byte bound, one newline-terminated line, tolerated
  unterminated tail, no carriage returns). fm_remote_job_wait samples
  state and the SECONDS clock with no per-sample children; one date call
  converts the epoch deadline once.
* fm-remote-delta-read stats the log each poll and re-runs the bounded
  capture and hashing only when size, mtime, ctime, inode, or device
  change. The snapshot's own stat writes the comparison key, so a log
  that moves between the gate and the capture is never read as stable.
* worker_preempting_waiter_exists reads state, home, and the staged argv
  head with builtins only. The now-unused worker_job_command goes away.

The bounded reads use -d '' -n, which behaves identically on the macOS
stock bash 3.2 and current bash; -N does not exist on 3.2. Tests cover
the malformed-record corpus, delta identity gating, fork-free lane
scanning through counting PATH shims, and same-home versus cross-home
preemption. No signal traps or sleep contracts change.

* no-mistakes(review): Restore subsecond delta keys and byte-bounded builtin record reads

* no-mistakes(document): Clarify delta snapshot caching and coarse-timestamp fallback

* no-mistakes(lint): Scope UTF-8 regression locales to individual function calls

* no-mistakes(ci): Fixed both lint failures by applying the documented production-library analysis boundary at the two affected test imports. Runtime behavior is unchanged; the library remains independently linted. Canonical full-analysis lint passed for the library and both suites, as did bash syntax checks and git diff --check

* fix(bin): recognize titled Claude top rules and preserve grey slash commands (kunchenguid#5963)

* fix(composer): read a titled Claude top rule as the composer's edge

A named Claude Code session draws its title into the composer's top rule.
The strict separator predicate rejected that row, so the closing rule read
as a lower unmatched separator and an idle, empty composer classified
unknown on every cursorless backend, refusing fm-send, exit, and relaunch.

Spare a bare agent-glyph row sandwiched between a width-proven titled rule
and the screen's only unmatched separator directly below it. The strict
separator predicate, dead-shell rule, and blank-row posture are unchanged.

Fixes kunchenguid#5601
Fixes kunchenguid#5558

* no-mistakes(test): Keep Claude's grey slash command in Herdr payload proof

* no-mistakes(test): Make missing-herdr version check hermetic to installed herdr

* fix: prevent Pi trust prompts in seeded secondmate homes (kunchenguid#6387)

* fix: pre-approve Pi trust for seeded secondmate homes

Unattended first launches of Firstmate-seeded Pi secondmate homes stalled on
"Trust project folder?" until Enter. Probe --approve like --tui-mode and pass
it only for --secondmate when help advertises it (.fm-secondmate-home signal),
leaving ordinary workers and older Pi unchanged.

* no-mistakes(document): Consolidate Pi seeded-home trust documentation ownership

* no-mistakes(ci): Fixed Lint 1’s unused polling counter. Diagnosed Behavior portable serial 4 as a pre-existing delta-reader test clock race; replaced timing-dependent rewrite and deletion with deterministic executable-boundary synchronization. ShellCheck, Bash syntax checks, and git diff --check passed. Delta-reader tests passed three consecutive runs; all three live Pi trust cases passed. Production behavior unchanged

* fix(bin): retry ShellCheck roots that hit the memory ceiling without --external-sources (kunchenguid#6443)

* fix(lint): retry memory-bound roots without external sources

* no-mistakes(review): Make fallback tests portable and correct source-following telemetry

* no-mistakes(review): Remove committed parity fixtures and use disposable test roots

* no-mistakes(document): Document ShellCheck memory fallback and telemetry

* no-mistakes(review): Cover bounded and unbounded fallback RSS behavior

* no-mistakes(document): Correct stale lint fallback documentation

* no-mistakes(document): Correct stale lint test documentation

* no-mistakes(ci): The memory fallback (the retry without --external-sources) now gets only the time left in its root's original deadline, so it can no longer outlast the CI job. Invariant: one root's first attempt plus its fallback must fit inside a single FM_LINT_ROOT_SECONDS deadline, plus the cleanup grace. Only one site started a new deadline: the fallback call in fm_lint_run_root. The deadline is the only budget involved, because the memory limit already applies to each process separately. Changes in bin/fm-lint.sh: - fm_lint_exec_root now takes a <seconds> argument instead of always reading FM_LINT_INTERNAL_ROOT_SECS. - The first attempt passes the full deadline. - The fallback passes floor((start + deadline - now) / 1000) seconds. - When bounds are enforced and less than 1 second is left, no retry starts. fm_exec_timed rejects 0 seconds, so the retry cannot run with no time. The root keeps reason=memory, and the shard output says "no time left in its Ns deadline to retry without it". - Unbounded local runs have no deadline and behave as before. - The header comment now describes the shared deadline. Changes in tests/fm-lint.test.sh: a new test, test_memory_fallback_spends_only_the_remaining_root_deadline, runs only on hosts that can enforce bounds. It uses a 6 s deadline and 1 s grace. - Case 1: the first attempt runs 3 s and then fails with memory status 251. The test asserts one fallback ran, reported reason=timeout, and the root's recorded duration is under 7000 ms. - Case 2: the first attempt runs 5.2 s. The test asserts no fallback starts, the skip is explained, and the sidecar records memory with source-following 1. Verification: - Full `nice -n 10 bash tests/fm-lint.test.sh` passed, including the new test, in about 5 minutes. - Case 1 run against the HEAD script: the root took 9168 ms, so the under-7000 ms check fails before the fix. - `bin/fm-lint.sh bin/fm-lint.sh tests/fm-lint.test.sh` reported no findings. - The CI workflow is unchanged, so the Test step still runs only tests/fm-lint.test.sh with nice -n 10 and the 12 GiB ShellCheck limit

* fix(pi): restore watcher continuity across successor gaps and make extension log opt-in (kunchenguid#5489)

* Fix Pi watcher successor-gap confirmations and add extension log

Accept an already-acknowledged handling confirmation as a no-op when the
generation matches, confirm the restoration's own recovery token with a
superseded (not rejected) outcome on generation mismatch, retire an arm on
confirm failure only when the failed token names that exact pid, and record
restore attempts, readiness timeouts, and confirm results in the bounded
state/.watch-extension.log. Regression tests: already-acked no-op plus
mismatch/dead-pid/lock-mismatch rejections and the manual-restart churn
contract in fm-watch-arm.test.sh, and a mid-restore marker advance with no
rejection appendix in fm-pi-watch-extension.test.sh.

* Treat a dead arm child as an empty slot so repair and retry recover

startArm and scheduleRetry answered unchanged while holding a ChildProcess
whose OS process was already gone but whose close had not fired, so neither
the repair tool nor the retry timer started anything until that close fired.
Gate slot occupancy on a liveness check (exit/signal codes plus pid probe)
and start a fresh arm instead, with a regression test driving the repair
tool against a dead-but-unclosed child.

* no-mistakes(document): Document new Pi extension log knob

* no-mistakes(review): Fix confirm-failure retire token match, add distinct-pid test

* no-mistakes(document): Clarify retire guard needs pid and generation

* Make the Pi extension diagnostic log opt-in and default-off

Only a positive FM_WATCH_EXTENSION_LOG_KEEP_LINES enables
state/.watch-extension.log. Unset, empty, non-numeric, zero, and
negative values disable logging entirely, so the default run writes
nothing and never creates the file. The shared positiveInteger
fallback semantics stay untouched for the retry and timeout knobs.
docs/configuration.md owns the knob contract and
docs/watcher-continuity.md points at it. Tests: the superseded-delivery
case runs opted in, and a new case proves unset, zero, and non-numeric
values create no log file while delivery still succeeds.

* no-mistakes(document): Qualify extension-log coverage bullet as opt-in

* no-mistakes(ci): The two reported checks (CI run 36372002913, Require no-mistakes run 36372069208) show conclusion action_required with 0 jobs and no logs because this is a fork PR (RibatTRW/firstmate) and GitHub is holding the workflow runs awaiting maintainer approval; that gate is external to the code and needs a maintainer to approve the runs. While verifying the change locally I found a real defect the approved CI run would hit: the PR's new test test_handling_delivered_rejects_a_superseded_generation failed deterministically. Invariant violated: reopen-announced (a non-successor/manual arm start) mints a fresh recovery generation only when the durable wake queue holds unrecovered work; an announced episode with an empty queue must be left untouched so idle arm starts never churn generations (the [ -s queue ] guard from kunchenguid#4819, relied on by bin/fm-watch.sh:2413 and covered by the append-reopens and announcement-bound sibling tests). The test called reopen with an empty queue and expected churn, so the fix establishes the queued-work precondition (append one wake, re-announce, re-read the generation) before asserting the reopen mints and the old confirmation mismatches. Test-only change, 14 lines in tests/fm-watch-arm.test.sh. Verified: fm-watch-arm.test.sh 25/25 ok on two consecutive runs, fm-pi-watch-extension.test.sh 55/55 ok, and shellcheck reports only one pre-existing warning outside the edited region

* no-mistakes(document): Restore blank line in watcher-continuity docs

* Route superseded Pi deliveries like confirmed ones and cover the retire guard

A superseded handling confirmation now falls through to the normal delivery
path, so an accepting supervision branch owns the wake instead of main.
Scope the watcher-continuity token-pinned confirmation and narrowed retire
rule to Pi, since omp and OpenCode still confirm against the current
successor. Add tests that fail when the retire guard, the scheduled-retry
gate, or the deferred-close gate is reverted, relabel the churned-generation
characterization test, and use a reaped pid for the dead-pid rejection.

* fix(pi): hide duplicate assistant finals from hidden processing retries (kunchenguid#5863)

* fix(pi): silence unacknowledged processing retry replies

Suppress autonomous processing prose before persistence and during streaming while retaining tool calls, signed reasoning, and retryable outcomes. Restore ordinary output after acknowledgement or a user message.

Fixes kunchenguid#4954

* no-mistakes(review): Silence only processing retries, keep first presentation visible

* fix(pi): preserve differing processing retry replies

* test(pi): accept Pi 1.0.1's renamed HTML export renderer lookup (kunchenguid#6530)

Pi 1.0.1's createToolHtmlRenderer reads getToolRenderers and ignores
getToolDefinition. Calm /export still includes stock grep HTML; the
fixture has to pass the lookup key the installed Pi actually reads.

* fix(bin): tolerate transient quota read failures (kunchenguid#6490)

* fix(procevent-quota): tolerate consecutive slow quota-axi reads

The quota allowance poll treated any quota_json failure as terminal, so one
slow quota-axi --json (measured max ~29s under a 48s derived bound) shut the
watch down until someone re-armed it, and the detail always said
"missing/incompatible". Tolerate three consecutive failed or timed-out reads
before going terminal, reset the streak on any good read, and report a
timeout distinctly from a missing or incompatible tool. Each timed poll runs
exactly one bounded --version and one bounded --json: validate the captured
version text through fm_quota_axi_version_compatible rather than launching a
second probe, and describe a mixed failure streak by count plus last cause.

* no-mistakes(document): Document quota polling failure tolerance

* no-mistakes(ci): Fixed ci-2 and ci-3. Permanent quota read failures (rc 2 missing, rc 3 incompatible) now report on the first poll, while transient rc 1/4 failures retain the existing three-failure retry behavior. The missing-binary test now uses an isolated PATH without quota-axi and asserts both permanent failures stop at condition_polls: 1. Verification passed: tests/fm-procevent-quota.test.sh, canonical fast lint for both changed files, bash syntax checks, and git diff --check

* no-mistakes(review): Classify untimed quota version failures as transient

* no-mistakes(document): Clarify quota polling failure budget

* fix(bin): clarify scratch guidance and dirty teardown refusals (kunchenguid#6505)

* fix(teardown): clarify scratch guidance and dirty worktree refusals

Keep ship proof material outside the task worktree and distinguish untracked-only leftovers from tracked edits without changing cleanup guards.

Fixes kunchenguid#6319

* fix(ci): Fixed ci-3 only. Both promotion outputs now replace the scout restriction and require external scratch storage and a clean worktree before done. Verification: 27 delivery tests passed, five mutations caught, restored test passed, pinned ShellCheck and syntax/whitespace checks passed. ci-1, ci-2, and ci-4 remain untouched

* fix(bin): escalate inbox instructions blocked by busy workers (kunchenguid#6518)

* Escalate inbox instructions stuck behind a busy worker

Count consecutive busy-deferred due doorbells durably and escalate at the configured bound without typing into the worker pane.

Fixes kunchenguid#6445

* fix(review): Fix inbox escalation deduplication and busy streak resets

* fix(review): Preserve busy inbox escalations through daemon supervision

* fix(document): Correct busy-inbox escalation documentation

* fix(ci): Fixed SC2034 in tests/fm-task-inbox.test.sh by including the loop counter in the failure diagnostic. Source-aware lint, all 34 inbox tests, and git diff --check pass. Behavior portable serial 6 reproduces identically on base 1f3e769 and target 78156b8 with Pi 1.0.1: an unrelated renderer API change breaks the unchanged Calm test. No Calm changes made; that failure is addressed separately by kunchenguid#6516. Logs retained in scratchpad-ci/

* fix(ci): Fixed ci-2, ci-3, and ci-4: successor failures surface, reset alerts deduplicate, and oversized busy limits fall back to two. Passed 47 inbox tests, 7 focused daemon checks, all 13 mutation checks, lint, documentation checks, and diff checks. Evidence: scratchpad-ci-selected/summary.json. ci-1 remains unchanged and unwaived. Fresh live Herdr proof remains with the outer driver

* fix: prevent unnecessary remote worker turnover (kunchenguid#6431)

* fix: prevent healthy remote job worker turnover

* no-mistakes(review): Serialize full LaunchAgent repair and verify launchd-tracked owners

* no-mistakes(review): Let launchd-tracked unpublished spawns start before reloading

* no-mistakes(document): Document remote worker heartbeat and serialized LaunchAgent recovery

* no-mistakes(ci): Full CI log showed the idle-worker regression exceeded its outdated command budget (82 versus 80) after independent heartbeat ownership checks were added. Raised the budget to 120 while retaining the separate busy-poll sleep limit. Remote-job and LaunchAgent executable tests passed, as did bash syntax validation and git diff --check. No production behavior changed

* no-mistakes(ci): Fixed missing readiness recovery under verified live ownership, preserving the serving PID and private file mode. Added executable regressions for deletion during a blocked sweep and stale readiness diagnostics without LaunchAgent reload. Deletion regression failed before the fix. Both remote-job test suites, bash syntax validation, and git diff --check passed

* no-mistakes(ci): Full CI log identified a flaky ownership-loss test racing an already-authorized heartbeat refresh. Replaced backdating and a fixed sleep with bounded observation of readiness expiry through the public probe. Production behavior unchanged. Remote-job and LaunchAgent executable suites passed; bash syntax validation and git diff --check passed

* fix: wait for launchd bootout cleanup

* no-mistakes(document): Document remote worker recovery and read-only turnover verification

* no-mistakes(review): Publish worker identity before lock owner records

* no-mistakes(document): Document worker identity publication safety invariant

* no-mistakes(ci): Fixed ci-1: replacement workers discard predecessor readiness before publishing identity and roll back identity if lock-owner recording fails. Added executable regressions reproducing both defects. LaunchAgent and orphan-reap suites, shell syntax checks, and git diff --check passed. No live service state was modified

* no-mistakes(ci): Restored lock-owner-before-identity publication and removed the identity rollback and reordering-only tests. Retained an executable regression proving predecessor readiness is rejected until replacement startup completes. LaunchAgent and orphan-reap suites, shell syntax checks, and git diff --check passed. Other changes remain intact; the retained-identity interrupted-repair edge remains out of scope. No live service state was modified

* fix: speed up remote-job sequence claim cleanup (kunchenguid#6575)

* fix(remote-job): reap expired seq claims with one directory walk

The hourly claim sweep forked uname+stat per .seq-claims entry and blocked
serving for ~85s at ~17k dirs. Delete expired empty claim dirs with a single
find -exec rmdir batch and cache the host uname for remaining mtime reads.

* no-mistakes(review): Restore original path mtime helper and drop uname cache

* no-mistakes(test): Restore claim retention eligibility; focused retention and serving tests pass

* no-mistakes(document): Document single-walk claim cleanup and regression entrypoints

* no-mistakes(ci): Fixed Lint 2’s reproduced SC1091 by adding the tests/lib.sh ShellCheck source directive to the retention test. Runtime behavior is unchanged. ShellCheck passed for both new claim tests; Bash syntax, the retention behavior test, and git diff --check passed

* fix(tests): keep fixture registries out of git worktree roots

A TMPDIR pointed at a repository root placed live .fm-test-* registries
beside tracked files, and a concurrent git add during the claim-walk CI
fix round committed three of them. Route registries and fixture roots
through a TMPDIR that refuses git worktree roots, remove the stray files,
and pin the escape with a behavioral cleanup test.

* no-mistakes(review): Preserve whole-second claim expiry in single-walk sweep

* no-mistakes(document): Correct temporary-directory resolution documentation

* no-mistakes(ci): Fixed ci-3 by changing only the stale, fresh, and read-only orphan fixture paths in tests/fm-test-fixture-cleanup.test.sh to use $FM_TEST_TMPDIR. All seven tests passed both normally and with TMPDIR set to the worktree root. Shell syntax and git diff --check passed

* no-mistakes(document): Align merged docs with upstream hand-back and policy owners

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: Tiago <tiagop@hey.com>
Co-authored-by: RibatTRW <aydinhrrs@gmail.com>
Co-authored-by: Ian Brown <742554+zestysoft@users.noreply.github.com>
Co-authored-by: Joseph Kim <jokim1@gmail.com>
Co-authored-by: Mickaël Rémond <mremond@process-one.net>
Co-authored-by: Hezki <hezki@users.noreply.github.com>
max-metaplanet pushed a commit to max-metaplanet/firstmate that referenced this pull request Oct 6, 2026
…tension log opt-in (kunchenguid#5489)

* Fix Pi watcher successor-gap confirmations and add extension log

Accept an already-acknowledged handling confirmation as a no-op when the
generation matches, confirm the restoration's own recovery token with a
superseded (not rejected) outcome on generation mismatch, retire an arm on
confirm failure only when the failed token names that exact pid, and record
restore attempts, readiness timeouts, and confirm results in the bounded
state/.watch-extension.log. Regression tests: already-acked no-op plus
mismatch/dead-pid/lock-mismatch rejections and the manual-restart churn
contract in fm-watch-arm.test.sh, and a mid-restore marker advance with no
rejection appendix in fm-pi-watch-extension.test.sh.

* Treat a dead arm child as an empty slot so repair and retry recover

startArm and scheduleRetry answered unchanged while holding a ChildProcess
whose OS process was already gone but whose close had not fired, so neither
the repair tool nor the retry timer started anything until that close fired.
Gate slot occupancy on a liveness check (exit/signal codes plus pid probe)
and start a fresh arm instead, with a regression test driving the repair
tool against a dead-but-unclosed child.

* no-mistakes(document): Document new Pi extension log knob

* no-mistakes(review): Fix confirm-failure retire token match, add distinct-pid test

* no-mistakes(document): Clarify retire guard needs pid and generation

* Make the Pi extension diagnostic log opt-in and default-off

Only a positive FM_WATCH_EXTENSION_LOG_KEEP_LINES enables
state/.watch-extension.log. Unset, empty, non-numeric, zero, and
negative values disable logging entirely, so the default run writes
nothing and never creates the file. The shared positiveInteger
fallback semantics stay untouched for the retry and timeout knobs.
docs/configuration.md owns the knob contract and
docs/watcher-continuity.md points at it. Tests: the superseded-delivery
case runs opted in, and a new case proves unset, zero, and non-numeric
values create no log file while delivery still succeeds.

* no-mistakes(document): Qualify extension-log coverage bullet as opt-in

* no-mistakes(ci): The two reported checks (CI run 36372002913, Require no-mistakes run 36372069208) show conclusion action_required with 0 jobs and no logs because this is a fork PR (RibatTRW/firstmate) and GitHub is holding the workflow runs awaiting maintainer approval; that gate is external to the code and needs a maintainer to approve the runs. While verifying the change locally I found a real defect the approved CI run would hit: the PR's new test test_handling_delivered_rejects_a_superseded_generation failed deterministically. Invariant violated: reopen-announced (a non-successor/manual arm start) mints a fresh recovery generation only when the durable wake queue holds unrecovered work; an announced episode with an empty queue must be left untouched so idle arm starts never churn generations (the [ -s queue ] guard from kunchenguid#4819, relied on by bin/fm-watch.sh:2413 and covered by the append-reopens and announcement-bound sibling tests). The test called reopen with an empty queue and expected churn, so the fix establishes the queued-work precondition (append one wake, re-announce, re-read the generation) before asserting the reopen mints and the old confirmation mismatches. Test-only change, 14 lines in tests/fm-watch-arm.test.sh. Verified: fm-watch-arm.test.sh 25/25 ok on two consecutive runs, fm-pi-watch-extension.test.sh 55/55 ok, and shellcheck reports only one pre-existing warning outside the edited region

* no-mistakes(document): Restore blank line in watcher-continuity docs

* Route superseded Pi deliveries like confirmed ones and cover the retire guard

A superseded handling confirmation now falls through to the normal delivery
path, so an accepting supervision branch owns the wake instead of main.
Scope the watcher-continuity token-pinned confirmation and narrowed retire
rule to Pi, since omp and OpenCode still confirm against the current
successor. Add tests that fail when the retire guard, the scheduled-retry
gate, or the deferred-close gate is reverted, relabel the churned-generation
characterization test, and use a reaped pid for the dead-pid rejection.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Pi watcher continuity: successor-gap confirmations and a dead-but-unclosed arm child block repair/retry recovery

2 participants