Skip to content
This repository was archived by the owner on Aug 25, 2026. It is now read-only.

feat: add durable secondmate report recovery - #81

Merged
JTInventory merged 26 commits into
mainfrom
fm/firstmate-adopt-phase2-secondmate-0726
Jul 27, 2026
Merged

JTInventory merged 26 commits into
mainfrom
fm/firstmate-adopt-phase2-secondmate-0726

Conversation

@JTInventory

Copy link
Copy Markdown
Owner

Intent

Adopt Phase 2 secondmate resilience onto JTInventory/firstmate from owner PRs kunchenguid#950 then kunchenguid#834. Preserve the JTInventory fork contracts already landed in Phase 1, including strict fm- send routing, Herdr lifecycle and presentation behavior, PR #79's merged-poll retirement path, and fork identity checks. Add parent-owned correlation, bounded recovery repost, one-time escalation, and durable pending-reply handling for missed secondmate reports. Keep the surface limited to secondmate spawn, recovery, handoff, reporting, and watcher integration; exclude AFK kunchenguid#758, broader concurrency work, quota changes, and Herdr rewrites. Ship by PR to JTInventory/firstmate without merging.

What Changed

  • Add parent-owned correlation for routed secondmate requests, with one bounded recovery repost, one-time escalation, and durable resolution history.
  • Make secondmate teardown and late-report handoffs crash-safe so unresolved replies cannot be silently orphaned, including forced and nested retirement paths.
  • Gate sends, bootstrap, and self-updates on verified watcher-protocol migration while preserving replayable instruction reread and secondmate nudge obligations.

Risk Assessment

✅ Low: The narrow change correctly preserves future-only obligations through the first retry without making them prematurely acknowledgeable or altering fork contracts.

Testing

The supplied full behavior baseline and all focused recovery, watcher, update, teardown, strict-routing, Herdr, and lifecycle checks passed. Manual evidence demonstrates one correlated repost, no duplicate repost, one parent-visible escalation, restart durability, and late-report resolution. This is a CLI/runtime change with no rendered UI, so visual screenshots were not applicable.

Evidence: End-to-end missed-report lifecycle transcript
SCENARIO: parent sends a marked request to fm-hibit
correlation=0a79759de8f9fe11
phase=awaiting_report
parent_status=/tmp/no-mistakes-evidence/01KYG36X848KZV0AAW00ARAEMG/missed-report-state/state/hibit.status

SCENARIO: secondmate finishes the request turn but its report never reaches the parent
phase_after_first_miss=recovery_sent
recovery_transport:
  target=hibit
  message=[fm-from-firstmate]⁣corr=0a79759de8f9fe11 REPOST REQUIRED: previous marked request had no correlated parent report. Reply on the parent status channel including corr=0a79759de8f9fe11. Original request: Give me the Phase 7 status
recovery_send_count=1

SCENARIO: repeated scan does not repost again
recovery_send_count_after_repeat_scan=1

SCENARIO: recovery turn also finishes without a correlated report
phase_after_second_miss=escalated
parent_visible_status:
  blocked: pending-reply-missed: task=hibit pending-reply-id=0a79759de8f9fe11 request=Give me the Phase 7 status
escalation_count=1

SCENARIO: a fresh process reads the same durable obligation after restart
restart_phase=escalated
restart_parent_status=/tmp/no-mistakes-evidence/01KYG36X848KZV0AAW00ARAEMG/missed-report-state/state/hibit.status

SCENARIO: a late correlated secondmate report resolves the original obligation
final_phase=resolved
resolved_via=helper
final_parent_status:
  blocked: pending-reply-missed: task=hibit pending-reply-id=0a79759de8f9fe11 request=Give me the Phase 7 status
  done [corr=0a79759de8f9fe11]: Phase 7 is complete (via-helper)
Evidence: Focused secondmate resilience tests
ok - normal correlated reply resolves once (idempotent)
ok - completed turn with no report triggers exactly one recovery
ok - recovery send preserves the fork's strict fm-<id> target contract
ok - recovery attempts reconcile without reinjection
ok - recovery reply resolves the original expectation
ok - second missed turn escalates once and remains durable
ok - concurrent escalation publishes one blocked status
ok - incomplete transaction locks are reclaimed
ok - stale generation reclaim preserves the live owner
ok - failed owner promotion removes its waiter generation
ok - live legacy waiter remains fail-closed pending watcher restart
ok - failed escalation publication remains retryable and publishes once
ok - transport success cannot masquerade as reply success
ok - undelivered records remain immutable across scan paths
ok - delivery confirmation fallback reconciles durably
ok - unrelated events and stale correlation ids cannot resolve
ok - restart preserves expectation and exact parent destination
ok - wrong-home reports are detected but do not silently acknowledge
ok - direct unmarked captain input creates no expectation
ok - fm-send marked secondmate path creates pending and embeds corr
ok - marked sends fail closed behind watcher protocol fence
ok - status-pointed document resolves the expectation
ok - optional helper report resolves without being required for correctness
ok - backend busy/idle observation covers Pi/Claude paths without conversation scrape
ok - tmux and zellij unknown states use bounded capture fallback
ok - tick skips terminal records and reuses target observations
ok - partial resolution reconciles and archives after restart
ok - correlations are reused only for matching open task records
ok - tick end-to-end: miss -> one recovery -> escalate -> durable
ok - failed transport discards undelivered expectation only
ok - forced-retirement receipts remain bound to their source state
ok - staged resolution commits retained history before receipt cleanup
ok - staging reconciles a resolution that won the lock race
ok - wrong-home diagnostics serialize with retirement staging
ok - finalization rechecks late resolution under correlation lock
ok - finalization recovers cleared history markers
ok - finalization preserves reports when resolution archival fails
ok - all pending-reply tests passed
ok - legacy watcher creates a durable pending-reply fence
ok - protocol restart stays fenced until tracked re-arm
ok - plain arm performs verified legacy primary takeover
ok - nested pending-reply gate validates child owner scope
ok - protocol restart refuses AFK and preserves X cadence
ok - verified AFK daemon ownership satisfies protocol gate
# all watcher protocol tests passed
ok - local-only worktree with HEAD on a fork remote is torn down (fix holds)
ok - teardown prompts tasks-axi backlog refresh when compatible
ok - teardown honors config/backlog-backend=manual even when tasks-axi is compatible
ok - teardown refuses arbitrary tasktmp cleanup targets from meta
ok - local-only worktree with truly unpushed work is refused (safety preserved)
ok - local-only worktree with work merged into local main is torn down (no regression)
ok - no-mistakes worktree with HEAD on origin is torn down (no regression)
ok - no-mistakes worktree with genuinely unlanded work is refused (safety preserved)
ok - local-only worktree with unpushed work is torn down under --force (escape hatch)
ok - squash-merged + deleted-branch worktree (PR merged) is torn down (the fix)
ok - squash-merged PR accepts a local HEAD that is an ancestor of the final PR head
ok - squash-merged PR accepts replayed unpushed local patches contained in the PR head
ok - merged PR does not allow teardown after a later local commit
ok - fm-pr-check does not refresh PR head after HEAD moves
ok - fm-pr-check records the remote PR head when the local worktree lags
ok - worktree whose content already landed in the default branch is torn down (content fallback)
ok - content fallback refreshes origin default before comparing trees
ok - dirty worktree is refused even when its committed work has landed (dirty always wins)
ok - gh lookup error with content not in default refuses (fail-safe)
ok - teardown retries a transient index lock without weakening landed-work checks
ok - forced secondmate teardown retries a transient child worktree index lock
ok - forced and recursive secondmate teardown propagate child close failures
ok - secondmate teardown preserves routing while a reply remains open
ok - forced teardown durably hands off an escalated reply
ok - failed forced teardown retries with its original staged phase
ok - recursive teardown preserves an un-escalated nested reply
ok - recursive teardown hands off nested reply history to durable parent state
ok - late nested reports migrate resolved history before teardown
ok - recursive teardown migrates already archived nested history
ok - failed nested teardown keeps the staged reply active
ok - Herdr teardown serializes every exact endpoint close through focus-safe verification
ok - confirmed projection journals retire before return failure and child teardown shares safe closure
ok - T1 main + secondmate fast-forward (single-parent), reread + nudge signalled
ok - T3 reread gates on instruction surface, nudge on advancement
ok - T4 dirty secondmate skipped, local edit preserved
ok - T5 diverged secondmate skipped, local commit preserved
ok - T6 idempotent: a second run is a no-op
ok - T7 registry backstop resolves, dedups meta+registry, excludes the firstmate repo
ok - T9 firstmate off its default branch is skipped, not forced
ok - T10 firstmate detached HEAD is skipped
ok - T11 unsafe secondmate home is not fast-forwarded
ok - T12 interrupted update obligations persist until acknowledged
ok - T13 real predecessor requires the installed updater pass
ok - T14 stale acknowledgements cannot clear newer generations
ok - T15 Herdr acknowledgements resolve exact live metadata
ok - T16 ancestor obligations remain acknowledgeable and retries are durable
ok - T17 skipped updates retain acknowledgement generations
ok - T18 future legacy generations survive concurrent acknowledgements
ok - T19 future-only legacy generations recover on the first retry
# all fm-update tests passed
ok - fm-send: a kind=secondmate target gets the from-firstmate marker and corr prepended
ok - fm-send: a kind=ship (crewmate) target is sent unmarked
ok - fm-send: an explicit session:window target is never marked
ok - fm-send: the --key path carries no marker (no literal text is typed)
ok - fm-send: the marker is exactly '[fm-from-firstmate]' + U+2063, detector keys on it
ok - fm-send: marked secondmate payload preserves trailing newline bytes
ok - fm-marker-lib: marker transformation is idempotent and newline-safe
Evidence: JT routing, Herdr, and lifecycle regression tests
ok - fm-send preserves the secondmate marker through the Herdr backend
ok - fm-send keeps journal-only recovery on the Herdr backend
ok - seed: registry scope+projects, charter copied, clones+origins, no-mistakes init in subhome only
ok - spawn: launches in the subhome with persistent charter, records routing meta
ok - send: a bare fm-<id> secondmate routes to the meta window with the from-firstmate marker
ok - handoff: in-scope item blocks move verbatim, out-of-scope stays, idempotent
ok - recovery: respawns from the durable registry and persistent home
ok - teardown: removes the home, then clears meta and the registry route
ok - backend selection precedence, metadata default, and known/unknown validation
ok - fm-spawn refuses unknown backends from FM_BACKEND, config/backend, and --backend
ok - selector resolution and capture/key/readability/kill dispatch use tmux adapter
ok - selector recovery retains Herdr routing and persisted workspace identity
ok - tmux adapter propagates kill and container-creation failures
ok - teardown preserves task state when backend kill fails
- Outcome: 🔧 1 issue found → auto-fixed ✅ across 2 runs (1h2m19s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 4 issues found → auto-fixed (21) ✅
  • 🚨 bin/fm-pending-reply-lib.sh:404 - Required criterion: “durable pending-reply handling for missed secondmate reports.” This function treats any status line containing the correlation token as a completed reply, including a nonterminal working [corr=...] update. If the final report is then missed, recovery and escalation remain permanently disabled. Restrict resolution to answer-bearing states such as done, needs-decision, blocked, or failed.
  • 🚨 bin/fm-pending-reply-lib.sh:1018 - Required criterion: “durable pending-reply handling for missed secondmate reports.” The new open-reply query has no call sites, so normal secondmate teardown can remove the task metadata, status route, and endpoint while a reply remains open. Without completion evidence, the retained record can then never repost or escalate. Integrate this guard into teardown or define an explicit durable handoff before cleanup.
  • ⚠️ bin/fm-pending-reply-lib.sh:494 - Resolution commits phase=resolved before delivered_epoch, resolved_epoch, and resolved_via. A failure between these writes leaves a partial record that every later watcher tick skips. Persist the bookkeeping first and commit the terminal phase last, with restart reconciliation for partial records.
  • ⚠️ bin/fm-pending-reply-lib.sh:933 - Resolved records remain permanently in the active directory, and every 15-second watcher poll still launches repeated grep/tail/cut pipelines to read three fields before skipping each one. Long-lived homes therefore accumulate unbounded process work. Preserve any desired history outside the active scan or maintain an active-record index.

🔧 Fix: Harden secondmate pending-reply lifecycle
3 issues (2 errors, 1 warning) still open:

  • 🚨 bin/fm-teardown.sh:785 - The required “durable pending-reply handling” now blocks every non-resolved phase before checking --force. Because escalated, recovery_failed, and recovery_unknown records are retained indefinitely, a dead secondmate that cannot send a late report can never be retired. Define a durable post-escalation handoff/archive that permits explicitly approved forced retirement, or confirm this permanent lockout is intended.
  • 🚨 bin/fm-teardown.sh:785 - Required criterion: cleanup must not orphan an open reply. This guard covers only the top-level secondmate. Forced recursive cleanup can remove a nested secondmate home without checking that owning home’s pending replies. Apply the same guard or durable handoff for each nested kind=secondmate before its endpoint and home are removed.
  • ⚠️ bin/fm-pending-reply-lib.sh:150 - Correlation matching uses substring containment, so a terminal line containing corr=0123456789abcdef0 incorrectly acknowledges 0123456789abcdef. Require a token boundary after the 16 hexadecimal characters or extract and compare the complete token.

🔧 Fix: Preserve pending replies through forced teardown
4 issues (3 errors, 1 warning) still open:

  • 🚨 bin/fm-teardown.sh:797 - Required criterion: “durable pending-reply handling.” Forced retirement removes the expectation before recursive validation and endpoint cleanup succeed. A later refusal or close failure leaves a live secondmate with monitoring disabled and history falsely marked forced-teardown; nested cleanup has the same ordering at line 747. Stage the handoff and finalize it only after successful teardown, or restore the active record on failure.
  • 🚨 bin/fm-pending-reply-lib.sh:1153 - Required criterion: “one-time escalation.” Forced retirement accepts recovery_failed and recovery_unknown directly, bypassing fm_pending_reply_maybe_escalate; a force request in this window archives the record without writing the required blocked escalation. Publish and durably record the escalation before retiring these phases.
  • 🚨 bin/fm-pending-reply-lib.sh:1157 - Nested retirement commits phase=retired before moving the record to parent history. If execution stops or mv fails, the nested watcher can archive it inside the soon-to-be-deleted nested home, losing the durable handoff. Persist the destination and reconcile to it, or publish the parent-history record atomically before committing retirement.
  • ⚠️ bin/fm-pending-reply-lib.sh:76 - The correlation regex still lacks a trailing token boundary: corr=0123456789abcdefg yields the valid 16-hex prefix and can falsely resolve 0123456789abcdef. Require end-of-string or a non-alphanumeric/non-underscore delimiter after the ID.

🔧 Fix: Make forced retirement failure-safe
3 issues (2 errors, 1 warning) still open:

  • 🚨 bin/fm-pending-reply-lib.sh:1175 - After staging a recovery_failed or recovery_unknown record, an endpoint failure leaves a receipt whose retired_from is the original recovery phase. On retry, this line overwrites retirement_staged_from with escalated, so the existing receipt fails validation and approved forced retirement remains blocked. Preserve the original staged source phase across retries or clear the staged transaction when teardown fails.
  • 🚨 bin/fm-pending-reply-lib.sh:1168 - fm_pending_reply_maybe_escalate can discover a late report, resolve and archive the record, then this branch returns staging success without creating a retirement receipt. Teardown performs destructive cleanup before finalization fails with found=0; for nested homes, retry can subsequently delete the resolved history. Treat late resolution as a completed durable handoff and migrate its history to the retained parent state.
  • ⚠️ bin/fm-pending-reply-lib.sh:1252 - Finalization claims every shared .retire-* receipt matching only task_id, although nested task IDs are scoped to separate Firstmate homes. Concurrent same-named nested teardowns can publish another home's retirement before its cleanup succeeds. Bind receipts to their canonical source state or a unique teardown transaction.

🔧 Fix: Bind retirement handoffs to source state
1 error still open:

  • 🚨 bin/fm-pending-reply-lib.sh:521 - Required criterion: “durable pending-reply handling for missed secondmate reports.” If a late report resolves after retirement staging, this line deletes the source-bound handoff receipt before the resolved record is moved from the nested home's local history to retained parent history. Finalization then fails, and retry can delete the home containing the only history. The resolved-at-stage branch has the same crash window between lines 1243-1245. Move resolved history to retirement_history_state atomically before deleting the receipt or active record.

🔧 Fix: Promote resolved history before receipt cleanup
3 errors still open:

  • 🚨 bin/fm-pending-reply-lib.sh:1286 - Required criterion: “durable pending-reply handling for missed secondmate reports.” Lines 1286-1292 persist retirement metadata through four separate rewrites while watcher resolution remains active. If resolution occurs before both destination and source are recorded, the record is archived inside the nested home; staging then fails, and a retry sees no active reply before deleting that home and its only history. Persist the staging transaction atomically or serialize/reconcile resolution against partial staging.
  • 🚨 bin/fm-pending-reply-lib.sh:986 - Required criterion: “one-time escalation.” The watcher and forced-teardown staging can concurrently pass this non-atomic status check and both append the same blocked: escalation before either commits phase=escalated. Use a narrow per-correlation transaction lock around the resolve recheck, publication, and phase commit.
  • 🚨 bin/fm-pending-reply-lib.sh:596 - Required criterion: “durable pending-reply handling for missed secondmate reports.” Resolved promotion deletes the forced-retirement receipt before creating its resolved replacement. Failure in that gap leaves retained history but no receipt; the already-running finalizer then reaches found=0 after destructive teardown. Atomically replace the receipt or let finalization recognize source-bound resolved history without one.

🔧 Fix: Serialize pending-reply handoff transactions
5 issues (4 errors, 1 warning) still open:

  • 🚨 bin/fm-pending-reply-lib.sh:170 - Required criterion: “bounded recovery repost, one-time escalation, and durable pending-reply handling.” The lock directory becomes visible before its three ownership files are written, while stale recovery requires all three fields. A crash during acquisition or release therefore leaves an incomplete directory that is never reclaimed, permanently blocking resolution, escalation, and teardown for that correlation. Make incomplete-lock recovery safe and bounded.
  • 🚨 bin/fm-teardown.sh:752 - Required criterion: “durable pending-reply handling for missed secondmate reports.” Recursive teardown calls the resolved-history handoff only when task_has_open is true. An already resolved reply has left the active directory, so this condition is false and the nested home containing its only history is deleted. Invoke the handoff for every nested secondmate before removing its home.
  • 🚨 bin/fm-pending-reply-lib.sh:1617 - Required criterion: “durable pending-reply handling for missed secondmate reports.” Finalization consumes the retirement receipt outside the correlation lock. A concurrent late resolution can commit phase=resolved while finalization moves the older retired receipt; resolved promotion then rejects that destination, leaving conflicting active and historical records after teardown removes the report route. Finalize each correlation under the same lock with a final state recheck.
  • 🚨 bin/fm-pending-reply-lib.sh:1619 - Required criterion: “durable pending-reply handling for missed secondmate reports.” Finalization clears the history’s staging fields before deleting the active record. A crash between those operations leaves an escalated active record and receipt-free history; every retry skips that history because retirement_staged_epoch is empty and returns found=0. Retain the transaction marker until active cleanup commits or recognize this validated recovery state on retry.
  • ⚠️ bin/fm-pending-reply-lib.sh:1154 - The wrong-home observer rewrites the correlation record without the new lock. A stale read can land after forced-retirement staging and erase its retirement fields, allowing endpoint/home deletion before finalization rejects the malformed active record. Protect this writer with the same correlation lock or store its diagnostic counters separately.

🔧 Fix: Harden pending-reply transaction recovery
3 issues (2 errors, 1 warning) still open:

  • 🚨 bin/fm-pending-reply-lib.sh:202 - Required criterion: “one-time escalation” and “durable pending-reply handling.” Stale-lock validation and mv are not atomic: after one reclaimer removes the inspected lock and a new owner publishes its lock, another paused reclaimer can move and delete that new live lock. Two writers can then enter the correlation transaction concurrently. Use a takeover protocol that cannot remove a newer lock generation.
  • 🚨 bin/fm-pending-reply-lib.sh:1614 - Required criterion: “durable pending-reply handling for missed secondmate reports.” This suppresses both “no correlated report” and persistence failures after a real report is found. If saving resolution evidence fails, finalization can consume the retired receipt and delete the active record, losing the report. Return a distinct no-report result and abort finalization on persistence or archival failure.
  • ⚠️ bin/fm-teardown.sh:779 - An already archived nested resolution is migrated through a .handoff-* receipt, but child_retire_staged remains false, so finalization never consumes that receipt. Each such teardown leaves a permanent duplicate record outside the active scan. Finalize migrated resolved handoffs even when no active reply was staged, or migrate them without creating a receipt.

🔧 Fix: Harden pending-reply takeover and finalization
3 issues (2 errors, 1 warning) still open:

  • 🚨 bin/fm-pending-reply-lib.sh:275 - Required criteria: “one-time escalation” and “durable pending-reply handling.” Election and promotion are not atomic: waiter A can select itself, waiter B can then publish a lexicographically earlier record and become owned, after which A also becomes owned. Both processes can enter the same correlation transaction and duplicate escalation or conflict with resolution/teardown. Replace snapshot-based promotion with one atomic, generation-bound ownership claim.
  • 🚨 bin/fm-pending-reply-lib.sh:276 - Required criterion: “durable pending-reply handling.” If the waiting→owned rewrite fails, this immediate return leaves a waiting generation owned by the still-live watcher process. When it sorts before later contenders, future acquisitions time out and resolution, escalation, and teardown remain blocked for the watcher's lifetime. Remove the caller's generation on every failed or early acquisition exit.
  • ⚠️ bin/fm-teardown.sh:766 - After resolved history is migrated, an endpoint-close failure skips finalization. On retry, the source history is already gone, so child_resolved_handoff remains zero and finalization is skipped again, leaving the source-bound .handoff-* receipt permanently. Treat an existing validated destination history or receipt as pending finalization on retry.

🔧 Fix: Harden pending-reply ownership and handoff retries
2 errors still open:

  • 🚨 bin/fm-pending-reply-lib.sh:260 - Required criteria: “bounded recovery repost, one-time escalation, and durable pending-reply handling.” Live watchers from the starting revision can leave waiting/owned lock generations without ticket= during an in-place update. This validation rejects that generation but removes only the new contender, so all correlation transactions remain blocked until the old watcher exits. Add legacy-generation migration/draining or enforce a watcher-restart barrier.
  • 🚨 bin/fm-send.sh:120 - Required criterion: “Add parent-owned correlation ... and durable pending-reply handling for missed secondmate reports.” This fallback extracts corr= from ordinary request text and reuses its open record. A new follow-up quoting an earlier correlation therefore gets no expectation of its own and can resolve or recover under the previous request. Require explicit reuse through FM_PENDING_REPLY_EXISTING_CORR; do not infer it from free text.

🔧 Fix: Drain legacy locks and require explicit correlation reuse
1 error still open:

  • 🚨 bin/fm-pending-reply-lib.sh:293 - Required criteria: “one-time escalation” and “durable pending-reply handling.” Stable polling cannot prove a live ticketless waiter is abandoned, and the separate checks before this rm are racy. The legacy process can promote waiting to owned between validation and deletion, or resume a delayed promotion after the new process acquires, producing two concurrent owners. Use a watcher-restart barrier or a takeover protocol understood by both versions; do not time-reclaim live legacy generations.

🔧 Fix: Enforce watcher restart barrier for legacy locks
3 errors still open:

  • 🚨 bin/fm-update.sh:58 - Required criterion: “bounded recovery repost, one-time escalation, and durable pending-reply handling.” The first update installing this barrier still runs the previous updater and already-loaded skill, which know nothing about watcher restarts. A retry is already current, so this line never requests the missed restart. Persist a protocol-migration marker or otherwise enforce the restart across the first upgrade.
  • 🚨 .agents/skills/updatefirstmate/SKILL.md:42 - Required criterion: “one-time escalation” and “durable pending-reply handling.” Secondmate restart is only a conversational instruction; neither the updater nor the skill verifies that it ran. A delayed, missed, or prematurely completed nudge can leave the legacy watcher active indefinitely. Require a verified home-scoped restart or a durable fence for every listed home.
  • 🚨 bin/fm-pending-reply-lib.sh:260 - Required criterion: “durable pending-reply handling.” Session-start local-head sync remains a separate bypass: bin/fm-bootstrap.sh:215 fast-forwards live secondmates with a reread-only nudge and no watcher restart. Their old watcher can keep producing live ticketless generations that this new fail-closed branch rejects. Apply the restart barrier to bootstrap convergence too.

🔧 Fix: Enforce verified watcher protocol migration
6 issues (5 errors, 1 warning) still open:

  • 🚨 bin/fm-pending-reply-lib.sh:183 - Required criterion: “durable pending-reply handling for missed secondmate reports.” Nested teardown supplies a child-owned state directory, but this gate validates the parent’s ambient FM_HOME and watcher path. A healthy child watcher therefore fails proof, permanently fencing staged retirement and history handoff. Pass the state owner’s canonical home and watcher path explicitly.
  • 🚨 bin/fm-watch-arm.sh:361 - Required watcher integration is broken: --restart-verify exits immediately after starting the detached watcher, and its EXIT trap releases the only follower claim. The replacement can later exit on an actionable wake without notifying the harness. Complete a durable runner/follower handoff before reporting restart success.
  • 🚨 bin/fm-watch-arm.sh:101 - The sanctioned direct git pull path can leave a live legacy primary watcher. This new protocol check makes normal arm reject it, but normal arm never performs the required restart; it repeatedly fails or reports the old follower while pending-reply work remains fenced. Enforce primary migration during startup or make normal arm perform the verified home-scoped takeover.
  • 🚨 bin/fm-send.sh:125 - The requested fail-closed migration does not cover new sends. Correlation creation, delivery preparation, and confirmation bypass the protocol gate, so a marked request can be delivered while a legacy watcher or durable migration fence remains active. Gate correlation reuse/creation before delivery.
  • 🚨 bin/fm-watcher-protocol-lib.sh:70 - Forbidden criterion: “exclude AFK fix: prevent AFK idle stalls and stale run attribution kunchenguid/firstmate#758.” The restart helper has no .afk guard and invokes a standalone watcher restart even though AFK requires the daemon to own watcher lifecycle. Refuse migration while AFK is active or route replacement through the daemon-owned lifecycle.
  • ⚠️ bin/fm-watcher-protocol-lib.sh:72 - The verified restart exports only home/root/state and never sources the target’s config/x-mode.env. An X-mode watcher therefore silently changes from its documented 30-second check interval to the default 300 seconds after migration. Preserve the target watcher’s runtime configuration during restart.

🔧 Fix: Harden watcher migration and pending-reply protocol gates
5 errors still open:

  • 🚨 bin/fm-watcher-protocol-lib.sh:55 - Forbidden criterion: “exclude AFK fix: prevent AFK idle stalls and stale run attribution kunchenguid/firstmate#758.” This gate requires a normal fm-watch-arm.sh follower, but AFK starts fm-watch.sh directly from fm-supervise-daemon.sh:1111 and forbids separately arming it. Consequently marked sends and pending-reply resolution, recovery, and escalation remain fenced throughout AFK. Accept verified daemon ownership without adding a normal follower.
  • 🚨 bin/fm-watcher-protocol-lib.sh:56 - Follower ownership is now part of the protocol proof, but the protocol remains pending-reply-ticket-v1, identical to the starting revision. An already-running watcher therefore remains indistinguishable from the repaired generation and can continue using its previously sourced, follower-unaware pending-reply code after this update. Introduce a new durable generation or re-exec migration that cannot accept the old proof.
  • 🚨 bin/fm-watch-arm.sh:116 - Automatic takeover recognizes a live legacy watcher only while its heartbeat is fresh. A live, identity-pinned legacy process with a stale heartbeat is not stopped; a replacement loses the singleton race or the command reports an existing follower, leaving the protocol fence active indefinitely. Apply verified takeover to every safely pinned live legacy watcher.
  • 🚨 bin/fm-update.sh:64 - The updater fast-forwards firstmate before verified re-arm. If migration then exits here, the retry sees an already-current checkout and prints reread-firstmate: no, permanently losing the mandatory instruction reread. Persist and replay the post-update reread obligation across the stop/re-arm/retry transaction.
  • 🚨 bin/fm-ff-lib.sh:405 - The secondmate has already fast-forwarded when the post-update callback can fail its watcher migration. This return occurs before FF_NUDGE_WINDOWS is updated and, during bootstrap, before its durable nudge marker is created. Retry sees the home as current, so the required AGENTS.md reread nudge is never sent. Record the obligation before migration or replay updated-but-unnudged homes.

🔧 Fix: Harden watcher migration and replay update obligations
5 issues (4 errors, 1 warning) still open:

  • 🚨 bin/fm-update.sh:40 - Required criteria: “one-time escalation” and “durable pending-reply handling.” The updater sources the v1 protocol library before fast-forwarding itself, so the first v1→v2 installation keeps using v1 in memory. A live v1 watcher can pass the old gate and avoid restart, leaving newly installed v2 sends fenced. Re-exec or reload the installed updater after self-update before validating watcher generation.
  • 🚨 bin/fm-update.sh:122 - Required criterion: “durable pending-reply handling.” The firstmate reread marker is deleted before the updater returns and before the agent actually rereads AGENTS.md. An interruption after deletion loses the obligation permanently. Clear it only after an explicit reread acknowledgement.
  • 🚨 bin/fm-update.sh:129 - Required criterion: “durable pending-reply handling.” Secondmate nudge markers are deleted before the skill sends or confirms the reread nudges. Output loss or a failed send therefore makes retries report no pending work. Retain each marker until its nudge is acknowledged.
  • 🚨 bin/fm-ff-lib.sh:412 - Required criterion: “durable pending-reply handling.” The secondmate is fast-forwarded before its reread marker is created; a crash in that gap leaves an already-current home with no replayable nudge obligation. Publish the obligation before committing the fast-forward, or reconcile updated-but-unmarked homes on retry. The firstmate path has the same ordering.
  • ⚠️ tests/fm-update.test.sh:92 - The migration tests always execute the target worktree's already-v2 updater while only the fixture repository is behind. They cannot reproduce the real first-install failure where a running v1 updater installs v2 but retains v1 helpers in memory. Run the updater from the behind fixture checkout and seed a live v1 watcher/follower.

🔧 Fix: Make update obligations durable and explicitly acknowledged
5 issues (3 errors, 2 warnings) still open:

  • 🚨 bin/fm-update.sh:119 - Required criterion: “durable pending-reply handling.” This re-exec branch cannot bootstrap the actual 384ed2b→3894711 upgrade because the already-running 384ed2b updater does not contain it. That process can install the new files but finish with its previously loaded helpers and legacy watcher proof. Use a migration entry point already understood by the predecessor or require a durable second invocation.
  • 🚨 bin/fm-update.sh:57 - Required criterion: “Preserve ... Herdr lifecycle and presentation behavior.” Herdr window targets are opaque values such as default:w1:p2; taking the final colon component produces p2, which this branch rejects because it is not fm-&lt;id&gt;. A delivered Herdr nudge can therefore never be acknowledged and repeats indefinitely. Resolve the exact target through live metadata instead of deriving the ID from tmux naming.
  • 🚨 bin/fm-ff-lib.sh:361 - Existing obligation markers are left unchanged, while acknowledgement later performs an unconditional delete. If update B advances a home while update A remains unacknowledged, A's delayed acknowledgement can remove B's newer reread/nudge obligation. Store the target commit or generation in the marker and compare-and-clear only the acknowledged generation.
  • ⚠️ bin/fm-bootstrap.sh:325 - After a failed bootstrap nudge succeeds on retry, this path removes only the parent retry receipt. The child reread marker remains, so the immediately following sweep sends the same nudge again. Clear the validated child marker after successful retry delivery, matching the direct-send path.
  • ⚠️ tests/fm-update.test.sh:82 - The cross-version fixture copies the target's new bin/ and changes only the protocol constant, so its alleged v1 updater already contains the new re-exec branch. It cannot expose the real predecessor bootstrap failure. Seed the fixture from the actual 384ed2b updater and helpers.

🔧 Fix: Make update obligations generation-safe across protocol migration
3 issues (2 errors, 1 warning) still open:

  • 🚨 .agents/skills/updatefirstmate/SKILL.md:40 - Required criterion: “durable pending-reply handling for missed secondmate reports.” The mandatory installed-updater second pass is documented only in the newly installed skill, but a base 9c6a4d8 invocation is still following the predecessor skill already loaded in memory; that workflow rereads only AGENTS.md, which contains no second-pass rule. The old updater can therefore finish without applying the v3 watcher fence. Put the migration trigger on a surface the base workflow already rereads, or require a genuinely staged prerequisite release.
  • 🚨 bin/fm-ff-lib.sh:47 - Generation acknowledgement is still a check-then-delete race: after the cat matches generation A, another update can atomically publish generation B before rm, causing A's delayed acknowledgement to delete B. Failed-merge rollback has the same problem at lines 399-403, where it can restore or remove a marker over a newer transaction. Serialize marker write/rollback/ack per home or use an atomic generation-bound claim.
  • ⚠️ bin/fm-ff-lib.sh:307 - An existing reread marker is decoded only after fetch, branch, and dirty-tree guards. Any early skip leaves FF_OBLIGATION_GENERATION empty while callers still report the pending action, producing reread-firstmate: yes with generation none, or a secondmate nudge without its generation line. Read and validate the existing marker before early returns so the documented acknowledgement remains possible.

🔧 Fix: Make update obligation claims atomic and replayable
2 errors still open:

  • 🚨 bin/fm-ff-lib.sh:75 - The atomic claim renames the only canonical obligation marker to a temporary path. A crash after this mv leaves only an orphan .update-obligation-claim.*, which later loads never reconcile, so the reread/nudge obligation disappears. Rollback and legacy normalization have the same window at lines 92 and 126. Recover orphan claims or retain a canonical durable transaction record until completion.
  • 🚨 bin/fm-ff-lib.sh:505 - Marker production remains a stale snapshot followed by unconditional replacement. Another updater or bootstrap can publish a newer generation between the snapshot and this write; the older writer then overwrites it, and rollback may restore an even older value. Serialize the complete publish/fast-forward/rollback transaction or use monotonic generation-aware replacement.

🔧 Fix: Make update obligations immutable and crash-safe
1 error still open:

  • 🚨 bin/fm-ff-lib.sh:100 - Required criterion: “durable pending-reply handling for missed secondmate reports.” Acknowledgement requires the obligation generation to equal exact HEAD. If instruction update B remains unacknowledged and a later non-instruction update advances to C, B is still selected as pending because it is an ancestor of C, but it can never be acknowledged. A crash during the multi-record deletion loop can create the same dead-end by deleting the newest record first. Validate the token against the currently selected pending generation and make partial acknowledgement replayable.

🔧 Fix: Keep ancestor update acknowledgements replayable
1 error still open:

  • 🚨 bin/fm-ff-lib.sh:84 - Required criterion: “durable pending-reply handling for missed secondmate reports.” This conversion preserves generation=C only when C is already an ancestor of HEAD. A predecessor updater writes that marker before fast-forwarding, so a concurrent acknowledgement at HEAD B can replace C with B and delete C; if the predecessor later advances to C and exits, the C reread/nudge obligation is lost. Preserve every valid, repository-resolvable generation as an immutable record; selection already ignores future generations until HEAD reaches them.

🔧 Fix: Preserve future legacy update obligations
2 issues (1 error, 1 warning) still open:

  • 🚨 bin/fm-ff-lib.sh:91 - Required criterion: “durable pending-reply handling for missed secondmate reports.” Conversion reads the legacy marker and later deletes its pathname without claiming that exact file. A predecessor updater can replace generation=C with newer generation=D between those operations; this loader persists C and then deletes D, losing the newer reread/nudge obligation. Atomically claim and recover the exact marker before conversion, or serialize legacy producers and converters.
  • ⚠️ bin/fm-ff-lib.sh:93 - When the crash state contains only a future generation=C marker while HEAD is B, conversion preserves C but then requires an installed ancestor record. Selection fails, so ff_target falsely reports the valid obligation as invalid and skips the fast-forward until another invocation. Allow a preserved future-only record to load without a current acknowledgement generation so recovery proceeds immediately.

🔧 Fix: Recover future-only obligations on first retry
✅ Re-checked - no issues remain.

🔧 **Test** - 1 issue found → auto-fixed ✅
  • 🚨 tests failed with exit code 1
  • bash bin/fm-run-behavior-tests.sh

🔧 Fix: Fix pending-reply teardown test fixtures
✅ Re-checked - no issues remain.

  • bash bin/fm-run-behavior-tests.sh
  • Supplied successful baseline: bash bin/fm-run-behavior-tests.sh
  • bash tests/fm-pending-reply.test.sh (initial direct attempt identified the expected gate-worktree environment refusal)
  • In a clean normal clone using the same gate-refusal test shim as the baseline runner: bash tests/fm-pending-reply.test.sh &amp;&amp; bash tests/fm-watcher-protocol.test.sh &amp;&amp; bash tests/fm-teardown.test.sh &amp;&amp; bash tests/fm-update.test.sh &amp;&amp; bash tests/fm-send-secondmate-marker.test.sh
  • In the same normal-clone setup: bash tests/fm-send-herdr-secondmate-marker.test.sh &amp;&amp; bash tests/fm-secondmate-lifecycle-e2e.test.sh &amp;&amp; bash tests/fm-backend.test.sh
  • Manual lifecycle scenario: deliver a correlated request, simulate a missed turn, verify one recovery repost, repeat the scan, simulate a second miss, verify one escalation, reload state in a fresh process, then submit a late correlated report and verify resolution
  • bash bin/fm-no-mistakes-pr-target-guard.sh
  • Inspected the base-to-target diff for forbidden AFK, quota, broad concurrency, and Herdr rewrite scope
  • git status --short after testing
✅ **Document** - passed

✅ No issues found.

🔧 **Lint** - 1 issue found → auto-fixed ✅
  • ⚠️ linter found issues (exit code 1)

🔧 Fix: Fix ShellCheck warnings in secondmate resilience scripts
✅ Re-checked - no issues remain.

✅ **Push** - passed

✅ No issues found.

@JTInventory
JTInventory merged commit 52a66df into main Jul 27, 2026
6 checks passed
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant