This repository was archived by the owner on Aug 25, 2026. It is now read-only.
feat: add durable secondmate report recovery - #81
Merged
Merged
Conversation
…rotocol migration
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Adopt Phase 2 secondmate resilience onto JTInventory/firstmate from owner PRs kunchenguid#950 then kunchenguid#834. Preserve the JTInventory fork contracts already landed in Phase 1, including strict fm- send routing, Herdr lifecycle and presentation behavior, PR #79's merged-poll retirement path, and fork identity checks. Add parent-owned correlation, bounded recovery repost, one-time escalation, and durable pending-reply handling for missed secondmate reports. Keep the surface limited to secondmate spawn, recovery, handoff, reporting, and watcher integration; exclude AFK kunchenguid#758, broader concurrency work, quota changes, and Herdr rewrites. Ship by PR to JTInventory/firstmate without merging.
What Changed
Risk Assessment
✅ Low: The narrow change correctly preserves future-only obligations through the first retry without making them prematurely acknowledgeable or altering fork contracts.
Testing
The supplied full behavior baseline and all focused recovery, watcher, update, teardown, strict-routing, Herdr, and lifecycle checks passed. Manual evidence demonstrates one correlated repost, no duplicate repost, one parent-visible escalation, restart durability, and late-report resolution. This is a CLI/runtime change with no rendered UI, so visual screenshots were not applicable.
Evidence: End-to-end missed-report lifecycle transcript
Evidence: Focused secondmate resilience tests
Evidence: JT routing, Herdr, and lifecycle regression tests
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
🔧 **Review** - 4 issues found → auto-fixed (21) ✅
bin/fm-pending-reply-lib.sh:404- Required criterion: “durable pending-reply handling for missed secondmate reports.” This function treats any status line containing the correlation token as a completed reply, including a nonterminalworking [corr=...]update. If the final report is then missed, recovery and escalation remain permanently disabled. Restrict resolution to answer-bearing states such asdone,needs-decision,blocked, orfailed.bin/fm-pending-reply-lib.sh:1018- Required criterion: “durable pending-reply handling for missed secondmate reports.” The new open-reply query has no call sites, so normal secondmate teardown can remove the task metadata, status route, and endpoint while a reply remains open. Without completion evidence, the retained record can then never repost or escalate. Integrate this guard into teardown or define an explicit durable handoff before cleanup.bin/fm-pending-reply-lib.sh:494- Resolution commitsphase=resolvedbeforedelivered_epoch,resolved_epoch, andresolved_via. A failure between these writes leaves a partial record that every later watcher tick skips. Persist the bookkeeping first and commit the terminal phase last, with restart reconciliation for partial records.bin/fm-pending-reply-lib.sh:933- Resolved records remain permanently in the active directory, and every 15-second watcher poll still launches repeatedgrep/tail/cutpipelines to read three fields before skipping each one. Long-lived homes therefore accumulate unbounded process work. Preserve any desired history outside the active scan or maintain an active-record index.🔧 Fix: Harden secondmate pending-reply lifecycle
3 issues (2 errors, 1 warning) still open:
bin/fm-teardown.sh:785- The required “durable pending-reply handling” now blocks every non-resolved phase before checking--force. Becauseescalated,recovery_failed, andrecovery_unknownrecords are retained indefinitely, a dead secondmate that cannot send a late report can never be retired. Define a durable post-escalation handoff/archive that permits explicitly approved forced retirement, or confirm this permanent lockout is intended.bin/fm-teardown.sh:785- Required criterion: cleanup must not orphan an open reply. This guard covers only the top-level secondmate. Forced recursive cleanup can remove a nested secondmate home without checking that owning home’s pending replies. Apply the same guard or durable handoff for each nestedkind=secondmatebefore its endpoint and home are removed.bin/fm-pending-reply-lib.sh:150- Correlation matching uses substring containment, so a terminal line containingcorr=0123456789abcdef0incorrectly acknowledges0123456789abcdef. Require a token boundary after the 16 hexadecimal characters or extract and compare the complete token.🔧 Fix: Preserve pending replies through forced teardown
4 issues (3 errors, 1 warning) still open:
bin/fm-teardown.sh:797- Required criterion: “durable pending-reply handling.” Forced retirement removes the expectation before recursive validation and endpoint cleanup succeed. A later refusal or close failure leaves a live secondmate with monitoring disabled and history falsely markedforced-teardown; nested cleanup has the same ordering at line 747. Stage the handoff and finalize it only after successful teardown, or restore the active record on failure.bin/fm-pending-reply-lib.sh:1153- Required criterion: “one-time escalation.” Forced retirement acceptsrecovery_failedandrecovery_unknowndirectly, bypassingfm_pending_reply_maybe_escalate; a force request in this window archives the record without writing the required blocked escalation. Publish and durably record the escalation before retiring these phases.bin/fm-pending-reply-lib.sh:1157- Nested retirement commitsphase=retiredbefore moving the record to parent history. If execution stops ormvfails, the nested watcher can archive it inside the soon-to-be-deleted nested home, losing the durable handoff. Persist the destination and reconcile to it, or publish the parent-history record atomically before committing retirement.bin/fm-pending-reply-lib.sh:76- The correlation regex still lacks a trailing token boundary:corr=0123456789abcdefgyields the valid 16-hex prefix and can falsely resolve0123456789abcdef. Require end-of-string or a non-alphanumeric/non-underscore delimiter after the ID.🔧 Fix: Make forced retirement failure-safe
3 issues (2 errors, 1 warning) still open:
bin/fm-pending-reply-lib.sh:1175- After staging arecovery_failedorrecovery_unknownrecord, an endpoint failure leaves a receipt whoseretired_fromis the original recovery phase. On retry, this line overwritesretirement_staged_fromwithescalated, so the existing receipt fails validation and approved forced retirement remains blocked. Preserve the original staged source phase across retries or clear the staged transaction when teardown fails.bin/fm-pending-reply-lib.sh:1168-fm_pending_reply_maybe_escalatecan discover a late report, resolve and archive the record, then this branch returns staging success without creating a retirement receipt. Teardown performs destructive cleanup before finalization fails withfound=0; for nested homes, retry can subsequently delete the resolved history. Treat late resolution as a completed durable handoff and migrate its history to the retained parent state.bin/fm-pending-reply-lib.sh:1252- Finalization claims every shared.retire-*receipt matching onlytask_id, although nested task IDs are scoped to separate Firstmate homes. Concurrent same-named nested teardowns can publish another home's retirement before its cleanup succeeds. Bind receipts to their canonical source state or a unique teardown transaction.🔧 Fix: Bind retirement handoffs to source state
1 error still open:
bin/fm-pending-reply-lib.sh:521- Required criterion: “durable pending-reply handling for missed secondmate reports.” If a late report resolves after retirement staging, this line deletes the source-bound handoff receipt before the resolved record is moved from the nested home's local history to retained parent history. Finalization then fails, and retry can delete the home containing the only history. The resolved-at-stage branch has the same crash window between lines 1243-1245. Move resolved history toretirement_history_stateatomically before deleting the receipt or active record.🔧 Fix: Promote resolved history before receipt cleanup
3 errors still open:
bin/fm-pending-reply-lib.sh:1286- Required criterion: “durable pending-reply handling for missed secondmate reports.” Lines 1286-1292 persist retirement metadata through four separate rewrites while watcher resolution remains active. If resolution occurs before both destination and source are recorded, the record is archived inside the nested home; staging then fails, and a retry sees no active reply before deleting that home and its only history. Persist the staging transaction atomically or serialize/reconcile resolution against partial staging.bin/fm-pending-reply-lib.sh:986- Required criterion: “one-time escalation.” The watcher and forced-teardown staging can concurrently pass this non-atomic status check and both append the sameblocked:escalation before either commitsphase=escalated. Use a narrow per-correlation transaction lock around the resolve recheck, publication, and phase commit.bin/fm-pending-reply-lib.sh:596- Required criterion: “durable pending-reply handling for missed secondmate reports.” Resolved promotion deletes the forced-retirement receipt before creating its resolved replacement. Failure in that gap leaves retained history but no receipt; the already-running finalizer then reachesfound=0after destructive teardown. Atomically replace the receipt or let finalization recognize source-bound resolved history without one.🔧 Fix: Serialize pending-reply handoff transactions
5 issues (4 errors, 1 warning) still open:
bin/fm-pending-reply-lib.sh:170- Required criterion: “bounded recovery repost, one-time escalation, and durable pending-reply handling.” The lock directory becomes visible before its three ownership files are written, while stale recovery requires all three fields. A crash during acquisition or release therefore leaves an incomplete directory that is never reclaimed, permanently blocking resolution, escalation, and teardown for that correlation. Make incomplete-lock recovery safe and bounded.bin/fm-teardown.sh:752- Required criterion: “durable pending-reply handling for missed secondmate reports.” Recursive teardown calls the resolved-history handoff only whentask_has_openis true. An already resolved reply has left the active directory, so this condition is false and the nested home containing its only history is deleted. Invoke the handoff for every nested secondmate before removing its home.bin/fm-pending-reply-lib.sh:1617- Required criterion: “durable pending-reply handling for missed secondmate reports.” Finalization consumes the retirement receipt outside the correlation lock. A concurrent late resolution can commitphase=resolvedwhile finalization moves the older retired receipt; resolved promotion then rejects that destination, leaving conflicting active and historical records after teardown removes the report route. Finalize each correlation under the same lock with a final state recheck.bin/fm-pending-reply-lib.sh:1619- Required criterion: “durable pending-reply handling for missed secondmate reports.” Finalization clears the history’s staging fields before deleting the active record. A crash between those operations leaves an escalated active record and receipt-free history; every retry skips that history becauseretirement_staged_epochis empty and returnsfound=0. Retain the transaction marker until active cleanup commits or recognize this validated recovery state on retry.bin/fm-pending-reply-lib.sh:1154- The wrong-home observer rewrites the correlation record without the new lock. A stale read can land after forced-retirement staging and erase its retirement fields, allowing endpoint/home deletion before finalization rejects the malformed active record. Protect this writer with the same correlation lock or store its diagnostic counters separately.🔧 Fix: Harden pending-reply transaction recovery
3 issues (2 errors, 1 warning) still open:
bin/fm-pending-reply-lib.sh:202- Required criterion: “one-time escalation” and “durable pending-reply handling.” Stale-lock validation andmvare not atomic: after one reclaimer removes the inspected lock and a new owner publishes its lock, another paused reclaimer can move and delete that new live lock. Two writers can then enter the correlation transaction concurrently. Use a takeover protocol that cannot remove a newer lock generation.bin/fm-pending-reply-lib.sh:1614- Required criterion: “durable pending-reply handling for missed secondmate reports.” This suppresses both “no correlated report” and persistence failures after a real report is found. If saving resolution evidence fails, finalization can consume the retired receipt and delete the active record, losing the report. Return a distinct no-report result and abort finalization on persistence or archival failure.bin/fm-teardown.sh:779- An already archived nested resolution is migrated through a.handoff-*receipt, butchild_retire_stagedremains false, so finalization never consumes that receipt. Each such teardown leaves a permanent duplicate record outside the active scan. Finalize migrated resolved handoffs even when no active reply was staged, or migrate them without creating a receipt.🔧 Fix: Harden pending-reply takeover and finalization
3 issues (2 errors, 1 warning) still open:
bin/fm-pending-reply-lib.sh:275- Required criteria: “one-time escalation” and “durable pending-reply handling.” Election and promotion are not atomic: waiter A can select itself, waiter B can then publish a lexicographically earlier record and becomeowned, after which A also becomesowned. Both processes can enter the same correlation transaction and duplicate escalation or conflict with resolution/teardown. Replace snapshot-based promotion with one atomic, generation-bound ownership claim.bin/fm-pending-reply-lib.sh:276- Required criterion: “durable pending-reply handling.” If thewaiting→ownedrewrite fails, this immediate return leaves a waiting generation owned by the still-live watcher process. When it sorts before later contenders, future acquisitions time out and resolution, escalation, and teardown remain blocked for the watcher's lifetime. Remove the caller's generation on every failed or early acquisition exit.bin/fm-teardown.sh:766- After resolved history is migrated, an endpoint-close failure skips finalization. On retry, the source history is already gone, sochild_resolved_handoffremains zero and finalization is skipped again, leaving the source-bound.handoff-*receipt permanently. Treat an existing validated destination history or receipt as pending finalization on retry.🔧 Fix: Harden pending-reply ownership and handoff retries
2 errors still open:
bin/fm-pending-reply-lib.sh:260- Required criteria: “bounded recovery repost, one-time escalation, and durable pending-reply handling.” Live watchers from the starting revision can leavewaiting/ownedlock generations withoutticket=during an in-place update. This validation rejects that generation but removes only the new contender, so all correlation transactions remain blocked until the old watcher exits. Add legacy-generation migration/draining or enforce a watcher-restart barrier.bin/fm-send.sh:120- Required criterion: “Add parent-owned correlation ... and durable pending-reply handling for missed secondmate reports.” This fallback extractscorr=from ordinary request text and reuses its open record. A new follow-up quoting an earlier correlation therefore gets no expectation of its own and can resolve or recover under the previous request. Require explicit reuse throughFM_PENDING_REPLY_EXISTING_CORR; do not infer it from free text.🔧 Fix: Drain legacy locks and require explicit correlation reuse
1 error still open:
bin/fm-pending-reply-lib.sh:293- Required criteria: “one-time escalation” and “durable pending-reply handling.” Stable polling cannot prove a live ticketless waiter is abandoned, and the separate checks before thisrmare racy. The legacy process can promotewaitingtoownedbetween validation and deletion, or resume a delayed promotion after the new process acquires, producing two concurrent owners. Use a watcher-restart barrier or a takeover protocol understood by both versions; do not time-reclaim live legacy generations.🔧 Fix: Enforce watcher restart barrier for legacy locks
3 errors still open:
bin/fm-update.sh:58- Required criterion: “bounded recovery repost, one-time escalation, and durable pending-reply handling.” The first update installing this barrier still runs the previous updater and already-loaded skill, which know nothing about watcher restarts. A retry is already current, so this line never requests the missed restart. Persist a protocol-migration marker or otherwise enforce the restart across the first upgrade..agents/skills/updatefirstmate/SKILL.md:42- Required criterion: “one-time escalation” and “durable pending-reply handling.” Secondmate restart is only a conversational instruction; neither the updater nor the skill verifies that it ran. A delayed, missed, or prematurely completed nudge can leave the legacy watcher active indefinitely. Require a verified home-scoped restart or a durable fence for every listed home.bin/fm-pending-reply-lib.sh:260- Required criterion: “durable pending-reply handling.” Session-start local-head sync remains a separate bypass:bin/fm-bootstrap.sh:215fast-forwards live secondmates with a reread-only nudge and no watcher restart. Their old watcher can keep producing live ticketless generations that this new fail-closed branch rejects. Apply the restart barrier to bootstrap convergence too.🔧 Fix: Enforce verified watcher protocol migration
6 issues (5 errors, 1 warning) still open:
bin/fm-pending-reply-lib.sh:183- Required criterion: “durable pending-reply handling for missed secondmate reports.” Nested teardown supplies a child-owned state directory, but this gate validates the parent’s ambientFM_HOMEand watcher path. A healthy child watcher therefore fails proof, permanently fencing staged retirement and history handoff. Pass the state owner’s canonical home and watcher path explicitly.bin/fm-watch-arm.sh:361- Required watcher integration is broken:--restart-verifyexits immediately after starting the detached watcher, and its EXIT trap releases the only follower claim. The replacement can later exit on an actionable wake without notifying the harness. Complete a durable runner/follower handoff before reporting restart success.bin/fm-watch-arm.sh:101- The sanctioned directgit pullpath can leave a live legacy primary watcher. This new protocol check makes normal arm reject it, but normal arm never performs the required restart; it repeatedly fails or reports the old follower while pending-reply work remains fenced. Enforce primary migration during startup or make normal arm perform the verified home-scoped takeover.bin/fm-send.sh:125- The requested fail-closed migration does not cover new sends. Correlation creation, delivery preparation, and confirmation bypass the protocol gate, so a marked request can be delivered while a legacy watcher or durable migration fence remains active. Gate correlation reuse/creation before delivery.bin/fm-watcher-protocol-lib.sh:70- Forbidden criterion: “exclude AFK fix: prevent AFK idle stalls and stale run attribution kunchenguid/firstmate#758.” The restart helper has no.afkguard and invokes a standalone watcher restart even though AFK requires the daemon to own watcher lifecycle. Refuse migration while AFK is active or route replacement through the daemon-owned lifecycle.bin/fm-watcher-protocol-lib.sh:72- The verified restart exports only home/root/state and never sources the target’sconfig/x-mode.env. An X-mode watcher therefore silently changes from its documented 30-second check interval to the default 300 seconds after migration. Preserve the target watcher’s runtime configuration during restart.🔧 Fix: Harden watcher migration and pending-reply protocol gates
5 errors still open:
bin/fm-watcher-protocol-lib.sh:55- Forbidden criterion: “exclude AFK fix: prevent AFK idle stalls and stale run attribution kunchenguid/firstmate#758.” This gate requires a normalfm-watch-arm.shfollower, but AFK startsfm-watch.shdirectly fromfm-supervise-daemon.sh:1111and forbids separately arming it. Consequently marked sends and pending-reply resolution, recovery, and escalation remain fenced throughout AFK. Accept verified daemon ownership without adding a normal follower.bin/fm-watcher-protocol-lib.sh:56- Follower ownership is now part of the protocol proof, but the protocol remainspending-reply-ticket-v1, identical to the starting revision. An already-running watcher therefore remains indistinguishable from the repaired generation and can continue using its previously sourced, follower-unaware pending-reply code after this update. Introduce a new durable generation or re-exec migration that cannot accept the old proof.bin/fm-watch-arm.sh:116- Automatic takeover recognizes a live legacy watcher only while its heartbeat is fresh. A live, identity-pinned legacy process with a stale heartbeat is not stopped; a replacement loses the singleton race or the command reports an existing follower, leaving the protocol fence active indefinitely. Apply verified takeover to every safely pinned live legacy watcher.bin/fm-update.sh:64- The updater fast-forwards firstmate before verified re-arm. If migration then exits here, the retry sees an already-current checkout and printsreread-firstmate: no, permanently losing the mandatory instruction reread. Persist and replay the post-update reread obligation across the stop/re-arm/retry transaction.bin/fm-ff-lib.sh:405- The secondmate has already fast-forwarded when the post-update callback can fail its watcher migration. This return occurs beforeFF_NUDGE_WINDOWSis updated and, during bootstrap, before its durable nudge marker is created. Retry sees the home as current, so the required AGENTS.md reread nudge is never sent. Record the obligation before migration or replay updated-but-unnudged homes.🔧 Fix: Harden watcher migration and replay update obligations
5 issues (4 errors, 1 warning) still open:
bin/fm-update.sh:40- Required criteria: “one-time escalation” and “durable pending-reply handling.” The updater sources the v1 protocol library before fast-forwarding itself, so the first v1→v2 installation keeps using v1 in memory. A live v1 watcher can pass the old gate and avoid restart, leaving newly installed v2 sends fenced. Re-exec or reload the installed updater after self-update before validating watcher generation.bin/fm-update.sh:122- Required criterion: “durable pending-reply handling.” The firstmate reread marker is deleted before the updater returns and before the agent actually rereads AGENTS.md. An interruption after deletion loses the obligation permanently. Clear it only after an explicit reread acknowledgement.bin/fm-update.sh:129- Required criterion: “durable pending-reply handling.” Secondmate nudge markers are deleted before the skill sends or confirms the reread nudges. Output loss or a failed send therefore makes retries report no pending work. Retain each marker until its nudge is acknowledged.bin/fm-ff-lib.sh:412- Required criterion: “durable pending-reply handling.” The secondmate is fast-forwarded before its reread marker is created; a crash in that gap leaves an already-current home with no replayable nudge obligation. Publish the obligation before committing the fast-forward, or reconcile updated-but-unmarked homes on retry. The firstmate path has the same ordering.tests/fm-update.test.sh:92- The migration tests always execute the target worktree's already-v2 updater while only the fixture repository is behind. They cannot reproduce the real first-install failure where a running v1 updater installs v2 but retains v1 helpers in memory. Run the updater from the behind fixture checkout and seed a live v1 watcher/follower.🔧 Fix: Make update obligations durable and explicitly acknowledged
5 issues (3 errors, 2 warnings) still open:
bin/fm-update.sh:119- Required criterion: “durable pending-reply handling.” This re-exec branch cannot bootstrap the actual 384ed2b→3894711 upgrade because the already-running 384ed2b updater does not contain it. That process can install the new files but finish with its previously loaded helpers and legacy watcher proof. Use a migration entry point already understood by the predecessor or require a durable second invocation.bin/fm-update.sh:57- Required criterion: “Preserve ... Herdr lifecycle and presentation behavior.” Herdr window targets are opaque values such asdefault:w1:p2; taking the final colon component producesp2, which this branch rejects because it is notfm-<id>. A delivered Herdr nudge can therefore never be acknowledged and repeats indefinitely. Resolve the exact target through live metadata instead of deriving the ID from tmux naming.bin/fm-ff-lib.sh:361- Existing obligation markers are left unchanged, while acknowledgement later performs an unconditional delete. If update B advances a home while update A remains unacknowledged, A's delayed acknowledgement can remove B's newer reread/nudge obligation. Store the target commit or generation in the marker and compare-and-clear only the acknowledged generation.bin/fm-bootstrap.sh:325- After a failed bootstrap nudge succeeds on retry, this path removes only the parent retry receipt. The child reread marker remains, so the immediately following sweep sends the same nudge again. Clear the validated child marker after successful retry delivery, matching the direct-send path.tests/fm-update.test.sh:82- The cross-version fixture copies the target's newbin/and changes only the protocol constant, so its alleged v1 updater already contains the new re-exec branch. It cannot expose the real predecessor bootstrap failure. Seed the fixture from the actual 384ed2b updater and helpers.🔧 Fix: Make update obligations generation-safe across protocol migration
3 issues (2 errors, 1 warning) still open:
.agents/skills/updatefirstmate/SKILL.md:40- Required criterion: “durable pending-reply handling for missed secondmate reports.” The mandatory installed-updater second pass is documented only in the newly installed skill, but a base9c6a4d8invocation is still following the predecessor skill already loaded in memory; that workflow rereads onlyAGENTS.md, which contains no second-pass rule. The old updater can therefore finish without applying the v3 watcher fence. Put the migration trigger on a surface the base workflow already rereads, or require a genuinely staged prerequisite release.bin/fm-ff-lib.sh:47- Generation acknowledgement is still a check-then-delete race: after thecatmatches generation A, another update can atomically publish generation B beforerm, causing A's delayed acknowledgement to delete B. Failed-merge rollback has the same problem at lines 399-403, where it can restore or remove a marker over a newer transaction. Serialize marker write/rollback/ack per home or use an atomic generation-bound claim.bin/fm-ff-lib.sh:307- An existing reread marker is decoded only after fetch, branch, and dirty-tree guards. Any early skip leavesFF_OBLIGATION_GENERATIONempty while callers still report the pending action, producingreread-firstmate: yeswith generationnone, or a secondmate nudge without its generation line. Read and validate the existing marker before early returns so the documented acknowledgement remains possible.🔧 Fix: Make update obligation claims atomic and replayable
2 errors still open:
bin/fm-ff-lib.sh:75- The atomic claim renames the only canonical obligation marker to a temporary path. A crash after thismvleaves only an orphan.update-obligation-claim.*, which later loads never reconcile, so the reread/nudge obligation disappears. Rollback and legacy normalization have the same window at lines 92 and 126. Recover orphan claims or retain a canonical durable transaction record until completion.bin/fm-ff-lib.sh:505- Marker production remains a stale snapshot followed by unconditional replacement. Another updater or bootstrap can publish a newer generation between the snapshot and this write; the older writer then overwrites it, and rollback may restore an even older value. Serialize the complete publish/fast-forward/rollback transaction or use monotonic generation-aware replacement.🔧 Fix: Make update obligations immutable and crash-safe
1 error still open:
bin/fm-ff-lib.sh:100- Required criterion: “durable pending-reply handling for missed secondmate reports.” Acknowledgement requires the obligation generation to equal exactHEAD. If instruction update B remains unacknowledged and a later non-instruction update advances to C, B is still selected as pending because it is an ancestor of C, but it can never be acknowledged. A crash during the multi-record deletion loop can create the same dead-end by deleting the newest record first. Validate the token against the currently selected pending generation and make partial acknowledgement replayable.🔧 Fix: Keep ancestor update acknowledgements replayable
1 error still open:
bin/fm-ff-lib.sh:84- Required criterion: “durable pending-reply handling for missed secondmate reports.” This conversion preservesgeneration=Conly when C is already an ancestor of HEAD. A predecessor updater writes that marker before fast-forwarding, so a concurrent acknowledgement at HEAD B can replace C with B and delete C; if the predecessor later advances to C and exits, the C reread/nudge obligation is lost. Preserve every valid, repository-resolvable generation as an immutable record; selection already ignores future generations until HEAD reaches them.🔧 Fix: Preserve future legacy update obligations
2 issues (1 error, 1 warning) still open:
bin/fm-ff-lib.sh:91- Required criterion: “durable pending-reply handling for missed secondmate reports.” Conversion reads the legacy marker and later deletes its pathname without claiming that exact file. A predecessor updater can replacegeneration=Cwith newergeneration=Dbetween those operations; this loader persists C and then deletes D, losing the newer reread/nudge obligation. Atomically claim and recover the exact marker before conversion, or serialize legacy producers and converters.bin/fm-ff-lib.sh:93- When the crash state contains only a futuregeneration=Cmarker while HEAD is B, conversion preserves C but then requires an installed ancestor record. Selection fails, soff_targetfalsely reports the valid obligation as invalid and skips the fast-forward until another invocation. Allow a preserved future-only record to load without a current acknowledgement generation so recovery proceeds immediately.🔧 Fix: Recover future-only obligations on first retry
✅ Re-checked - no issues remain.
🔧 **Test** - 1 issue found → auto-fixed ✅
bash bin/fm-run-behavior-tests.sh🔧 Fix: Fix pending-reply teardown test fixtures
✅ Re-checked - no issues remain.
bash bin/fm-run-behavior-tests.shSupplied successful baseline:bash bin/fm-run-behavior-tests.shbash tests/fm-pending-reply.test.sh(initial direct attempt identified the expected gate-worktree environment refusal)In a clean normal clone using the same gate-refusal test shim as the baseline runner:bash tests/fm-pending-reply.test.sh && bash tests/fm-watcher-protocol.test.sh && bash tests/fm-teardown.test.sh && bash tests/fm-update.test.sh && bash tests/fm-send-secondmate-marker.test.shIn the same normal-clone setup:bash tests/fm-send-herdr-secondmate-marker.test.sh && bash tests/fm-secondmate-lifecycle-e2e.test.sh && bash tests/fm-backend.test.shManual lifecycle scenario: deliver a correlated request, simulate a missed turn, verify one recovery repost, repeat the scan, simulate a second miss, verify one escalation, reload state in a fresh process, then submit a late correlated report and verify resolutionbash bin/fm-no-mistakes-pr-target-guard.shInspected the base-to-target diff for forbidden AFK, quota, broad concurrency, and Herdr rewrite scopegit status --shortafter testing✅ **Document** - passed
✅ No issues found.
🔧 **Lint** - 1 issue found → auto-fixed ✅
🔧 Fix: Fix ShellCheck warnings in secondmate resilience scripts
✅ Re-checked - no issues remain.
✅ **Push** - passed
✅ No issues found.