fix(bin): retire a pending backlog close whose row left the backlog - #32
Merged
Merged
Conversation
A row that closed and then aged out of done_keep retention before its own teardown's close attempt ran reads back from tasks-axi as code: NOT_FOUND, indistinguishable from a row that never existed. fm_backlog_close_marker_replay already retires that pending close as stale instead of retrying forever, but the library's own CRASH RECOVERY contract never said so. Document the behavior in the one place that owns it, and add a regression test proving the record retires so this stays fixed.
…s with absent-row close
…ILE; correct overclaims
rub-a-dub-dub
force-pushed
the
fm/firstmate-unreplayable-backlog-close-v4
branch
from
September 27, 2026 04:45
efc1c4f to
5977a32
Compare
rub-a-dub-dub
added a commit
that referenced
this pull request
Sep 27, 2026
Base advanced by three commits (#31, #30, #32) since this branch last took origin/main at 929a366. Only 4e5c8ee (#31, pin Codex workers to the standard service tier) overlapped the reconciliation, in three files: - .agents/skills/harness-adapters/references/harness/codex.md: keep upstream's Effort flag row, which now advertises `max` for gpt-5.6-luna, and add the base's new Service tier row alongside it. - bin/fm-spawn.sh: carry `-c 'service_tier="default"'` on both codex launch templates, keeping upstream's `--disable hooks` on the crewmate template, and take the base's service-tier rationale into the launch contract header. - tests/fm-spawn-dispatch-profile.test.sh: register the base's new test_codex_scout_uses_standard_service_tier next to upstream's restructured max-effort and hook-layer cases, and thread the service-tier flag through upstream's own Luna max-effort launch assertion. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Observed live in the main home on 2026-09-19, and filed by firstmate rather than
hand-fixed because bin/fm-backlog-transition-lib.sh owns the record's format and replay.
A teardown ran for a task whose backlog row had already closed and aged out of done_keep
retention. The endpoint and worktree cleaned up successfully (cleanup_incomplete=0), but
the backlog close could not be recorded:
error: ... its backlog item could not be closed atomically (error: Task "..." not found
in this backlog); the pending close is recorded and the next session start retries it
That retry can never succeed. state/.backlog-close retires only when its transition
lands, and this transition can never land against a task that no longer exists.
What Changed
fm_backlog_close_transitionnow probes the row after a failed close: a confirmedNOT_FOUNDremoves the pending-close record and raises the newFM_BACKLOG_CLOSE_ROW_ABSENTflag, while a row that is still present, or a lookup error, keeps the original close failure (annotated with the read error) and preserves the record for retry.absent/absent_incompleteresults via a newfm_backlog_close_replay_resulthelper plusFM_BACKLOG_CLOSE_REPLAY_DELIVERABLE, so a record retired against a vanished row is reported bybin/fm-bootstrap.shunderBACKLOG_RECONCILEnaming the completion link it could not confirm as applied (and any surviving endpoint or local copy) instead of being silently labelledstale; the superseded-incarnation case now prints its ownBOOTSTRAP_INFOline, andbin/fm-teardown.shprints an absent-row reminder rather than claiming a close it did not make.tests/fm-backlog-atomicity.test.shcovering retention-archived and outright-removed rows, unconfirmed absences from a broken row probe, and interrupted-cleanup warnings; updateddocs/configuration.md,AGENTS.md, and the bootstrap-diagnostics skill to document the retirement path and its diagnostic line.Risk Assessment
✅ Low: The change is well-bounded and intent-conformant: absence is accepted only on a positively confirmed NOT_FOUND from the row probe, every other failure still keeps the record and refuses loudly, the new absent/absent_incomplete labels and their wording agree across teardown, session start, docs and the skill contract, and the new paths are covered by six behavioral tests; the only surviving findings are comment/doc accuracy nits and one already-acknowledged durability tradeoff.
Testing
Installed the CI-pinned tasks-axi 0.2.5 into a throwaway temp prefix (it is absent on this host, which would otherwise make the whole owning suite self-skip), then reproduced the filed incident end to end against the real bin/fm-teardown.sh and bin/fm-bootstrap.sh with a real data/backlog.md: at base 5719dac completion prints the incident text verbatim, exits 1, and leaves an unlandable pending-close record, while at target 1889a5f the identical fixture exits 0, retires the record, and reports that the row had already left the backlog with its completion link named as unconfirmed rather than never applied; session start went from printing nothing and silently dropping the recorded PR link to reporting the retirement under BACKLOG_RECONCILE and naming it. I also captured both loud-refusal paths verbatim, confirming they still exit 1, keep the record, name the backlog read failure, and make no absence claim when the close never reported one. Each of the change's seven new tests fails at base with the incident message or the missing wording and passes at target; the owning suite came back 120 ok / 0 not ok and the two other suites that assert this operator wording are green. This change is terminal copy rather than a rendered UI surface, so CLI transcripts are the end-user artifact and no screenshot applies. Worktree left clean and all temp artifacts removed.
Evidence: Incident reproduced at base and fixed at target (headline before/after CLI transcript)
Source: Incident reproduced at base and fixed at target (headline before/after CLI transcript)
Evidence: Round summary: what changed, transcripts index, fail-before/pass-after table, suites run
Source: Round summary: what changed, transcripts index, fail-before/pass-after table, suites run
Evidence: Completion against an aged-out row — target 1889a5f (raw)
Source: Completion against an aged-out row — target 1889a5f (raw)
$ bin/fm-teardown.sh fm-incident-aged-out teardown fm-incident-aged-out complete (window firstmate:fm-fm-incident-aged-out, worktree .../evidence-completion/absent-worktree) Backlog: fm-incident-aged-out had already left .../home/data/backlog.md, so cleanup recorded no close there and its completion link (local main) could not be confirmed as applied and should be checked. Run bin/fm-tasks-axi.sh ready for dependency-cleared candidates, check date gates, and dispatch only work whose blockers are gone and date is due. $ echo $? 0 $ ls home/state/ | grep -c backlog-close # pending-close records left behind 0Evidence: Completion against an aged-out row — base 5719dac (the incident, raw)
Source: Completion against an aged-out row — base 5719dac (the incident, raw)
$ bin/fm-teardown.sh fm-incident-aged-out error: fm-incident-aged-out's endpoint and local copy are cleaned up, but its backlog item could not be closed atomically (error: "Task "fm-incident-aged-out" not found in this backlog"); the pending close is recorded and the next session start retries it $ echo $? 1 $ ls home/state/ | grep -c backlog-close # pending-close records left behind 1Evidence: Session start replaying an unlandable recorded close — target 1889a5f
Source: Session start replaying an unlandable recorded close — target 1889a5f
$ bin/fm-bootstrap.sh # next session start BACKLOG_RECONCILE: fm-incident-replay: the recorded backlog close was retired because its backlog row had already left this backlog, so no close was left to land; its recorded completion link (PR https://example.test/pr/7) could not be confirmed as applied and should be checked $ test -e home/state/fm-incident-replay.backlog-close && echo PRESENT || echo RETIRED RETIREDEvidence: Session start replay — base 5719dac (printed nothing, retired silently)
Source: Session start replay — base 5719dac (printed nothing, retired silently)
Evidence: Both loud-refusal paths, verbatim: unconfirmed absence, and a close that never reported one
Source: Both loud-refusal paths, verbatim: unconfirmed absence, and a close that never reported one
=== (1) close reported NOT_FOUND but the confirming row read failed === error: ... could not be closed atomically (error: Task "fm-incident-unconfirmed" not found in this backlog; this home's backlog row could not be read to confirm whether the item still exists (error: "backlog is unreadable"), so the next session start retries this close); the pending close is recorded and the next session start retries it exit 1 | state/fm-incident-unconfirmed.backlog-close: PRESENT === (2) close failed for a reason that is NOT an absence; row read also failed === error: ... could not be closed atomically (error: "backlog is unwritable"; this home's backlog row could not be read to confirm whether the item still exists (error: "backlog is unreadable"), so the next session start retries this close); the pending close is recorded and the next session start retries it exit 1 | state/fm-incident-no-absence.backlog-close: PRESENT (no absence is claimed: neither "absence" nor "had already left" appears)Evidence: Full owning suite at target: tests/fm-backlog-atomicity.test.sh — 120 ok, 0 not ok
Source: Full owning suite at target: tests/fm-backlog-atomicity.test.sh — 120 ok, 0 not ok
Pipeline
Updates from git push no-mistakes
... (9 earlier update rounds omitted to keep the PR body within GitHub's 65536-char limit; full history is in the run log.)
🔧 Fix: report retired absent close under BACKLOG_RECONCILE; correct overclaims
1 warning still open:
bin/fm-bootstrap.sh:1293- Theabsentdisposition asserts "its recorded completion link (<link>) was never applied and should be reconciled", butabsentis reached from two states and only one of them proves that. Path A (bin/fm-backlog-transition-lib.sh:1300) is the probe finding no row — there the claim is sound. Path B isfm_backlog_close_transitionsetting FM_BACKLOG_CLOSE_ROW_ABSENT=1 at bin/fm-backlog-transition-lib.sh:917 after the close itself returned NOT_FOUND, which replay reaches from thedone \*branch (bin/fm-backlog-transition-lib.sh:1281) and the fall-through (:1334) — i.e. after the probe positively saw the row. Concrete sequence: teardown runstasks-axi done <id> --pr <url>; it lands (row becomes done, PR recorded); the process is killed beforefm_backlog_record_remove "$marker"at bin/fm-backlog-transition-lib.sh:919, or that removal fails (the state tests/fm-backlog-atomicity.test.sh already covers as test_completion_fails_when_its_close_marker_cannot_be_removed); done_keep retention then prunes the row before the next session start. Replay probes, gets NOT_FOUND, retires the record — correct — and prints that the merged PR "was never applied and should be reconciled", which is the opposite of what happened. The change's own test test_recovery_reports_a_row_that_left_the_backlog_mid_close (tests/fm-backlog-atomicity.test.sh:2247) enshrines this for the sharper case: it asserts the row isdonebefore replay, so the link had almost certainly already been backfilled, yet the asserted output still claims it was never applied. Per AGENTS.md:564 this is an actionable line, so the agent is directed to re-file an artifact that is already recorded against a row that no longer exists. The retirement itself is right in every one of these cases; only the claim about the link is wrong. The code cannot distinguish the two states without re-introducing the archive-consultation component deliberately removed in 65e82d2, so the smallest honest remedies are both user-visible: soften the shared wording to say the link could not be confirmed applied, or carry the pre-close row state so thedone \*race reports its own disposition. Because that reverses wording the author deliberately chose in round 4 ("the recorded completion link is named in the message"), the remedy needs authorization rather than an auto-fix. Note the teardown-side twin at bin/fm-teardown.sh:1447 is not affected: therefm_backlog_donedemonstrably failed with NOT_FOUND, so "was never applied" is accurate.🔧 Fix: soften replay wording to unconfirmed completion link
3 issues (2 warnings, 1 info) still open:
bin/fm-bootstrap.sh:103- The sweep's header contract still asserts the opposite of what the code now reports. bin/fm-bootstrap.sh:102-104 says "Replayed transitions and restored In-flight rows print BOOTSTRAP_INFO facts; a record retired without landing its close leaves an unapplied completion link, so it reports through BACKLOG_RECONCILE instead." Round 5 deliberately removed that assertion from the emitted line (bin/fm-bootstrap.sh:1293 now says the link "could not be confirmed as applied and should be checked") precisely becauseabsentis reachable from a state where the link WAS already applied: fm_backlog_close_transition sets FM_BACKLOG_CLOSE_ROW_ABSENT=1 at bin/fm-backlog-transition-lib.sh:918 aftertasks-axi donelanded the link and the process died before fm_backlog_record_remove, with done_keep then pruning the row - the path the change's own test test_recovery_reports_a_row_that_left_the_backlog_mid_close (tests/fm-backlog-atomicity.test.sh:2287) exercises. docs/configuration.md:102 now states this explicitly ("session start says only that it could not be confirmed as applied, because a row that left the backlog before replay read it may already have carried the link from an earlier close"), so the header is the last surviving copy of the retracted claim and it is the file-level contract a maintainer reads before touching backlog_record_reconcile. The stale text was introduced by round 4 (4ab947d) and round 5 (231b167) corrected the message and docs/configuration.md without following it into this comment. Remedy is mechanical and changes no behavior or emitted output: make line 103-104 read that the retirement leaves a completion link whose application could not be confirmed.bin/fm-backlog-transition-lib.sh:915- When the close reports NOT_FOUND but the confirming probe fails, bin/fm-backlog-transition-lib.sh:915 restores FM_BACKLOG_TRANSITION_ERROR to the close error and discards FM_BACKLOG_ROW_ERROR entirely, so the only reason the record was preserved is never surfaced. Concrete sequence: a Beads or markdown home wheretasks-axi done <id>fails witherror: Task "<id>" not found in this backlog(a real absence) and the followingtasks-axi show <id>then exceeds FM_TASKS_AXI_TIMEOUT or hits an unreadable backlog - the state the change's own test break_row_probe_after_absent_close (tests/fm-backlog-atomicity.test.sh:245) fabricates. fm_backlog_row_probe sets FM_BACKLOG_ROW_ERROR to the read failure, line 913 sees FM_BACKLOG_ROW_RESULT=error, and line 915 overwrites it. bin/fm-teardown.sh:1447 then prints exactlyerror: <id>'s endpoint and local copy are cleaned up, but its backlog item could not be closed atomically (error: Task "<id>" not found in this backlog); the pending close is recorded and the next session start retries it- byte-for-byte the incident text quoted in the User intent, whose stated conclusion is "That retry can never succeed." Post-change that conclusion is wrong for this sub-case (the retry will confirm the absence and retire the record), but the operator or agent reading the message has no way to tell it apart from the pre-fix incident, and none of the actionable information - that the backlog read timed out or was unreadable, which is what actually needs fixing - reaches them. The wrong result is not the retirement decision, which is correct; it is that the refusal reports a cause it did not act on. Smallest honest remedy: when the probe fails, preserve both, e.g. set FM_BACKLOG_TRANSITION_ERROR to the close error plus the probe's FM_BACKLOG_ROW_ERROR, so the message names the unconfirmed absence. The test at tests/fm-backlog-atomicity.test.sh:1876 currently pins the lossy message (assert_contains "could not be closed") and would need the new wording. Flagged ask-user rather than auto-fix because the library header at bin/fm-backlog-transition-lib.sh:905 states the discard as a deliberate rule ("a probe that errors, times out, or still finds the row keeps the original close failure") and the fix changes operator-visible error text..agents/skills/bootstrap-diagnostics/SKILL.md:56- The new skill entry explains the unconfirmed link with "The row left the backlog before this replay could read it", but that is false for one of the two paths that reach theabsentresult. In the mid-close path the replay DID read the row - fm_backlog_close_marker_replay probes it, matches thedone *branch at bin/fm-backlog-transition-lib.sh:1280, and the row only leaves during the close that follows, which is exactly the state test_recovery_reports_a_row_that_left_the_backlog_mid_close (tests/fm-backlog-atomicity.test.sh:2287) constructs and the state that motivated round 5's softened wording in the first place. The prescribed action that follows ("Check where the closed row would have carried it, file it only if it is missing") is correct for both paths, so an agent following the entry still does the right thing; only the stated reason is wrong, and it is wrong in the direction that understates how likely the link already landed. Remedy is a one-clause comment fix with no behavior change: say the row left the backlog before the close could be applied, rather than before the replay could read it.🔧 Fix: name the failed backlog read in unconfirmed-absence refusal
2 issues (1 warning, 1 info) still open:
bin/fm-backlog-transition-lib.sh:919- The round-7 wording is applied to every close failure followed by a probe failure, not only to the NOT_FOUND case it was written for, so a close that never reported an absence is reported as an unconfirmed absence. bin/fm-backlog-transition-lib.sh:918 gates only on FM_BACKLOG_ROW_RESULT=error; nothing checks thatclose_errorwas itself a not-found. Concrete reachable sequence in a markdown home: TEARDOWN_BACKLOG_APPLIES is decided early (fm_backlog_transition_applies, bin/fm-backlog-transition-lib.sh:335 requires backlog.md), and the close runs much later at bin/fm-teardown.sh:3488. Remove or truncate data/backlog.md in that window. fm_backlog_done -> fm_backlog_mutate -> fm_backlog_source_present -> fm_backlog_record_present fails with "backlog file is not a regular file at <path>", so close_error is that text. fm_backlog_row_probe then fails on the identical check (bin/fm-backlog-transition-lib.sh:516-519) and sets FM_BACKLOG_ROW_ERROR to the same string with FM_BACKLOG_ROW_RESULT=error. Line 919 then produces: "backlog file is not a regular file at <path>; that absence could not be confirmed because this home's backlog row could not be read (backlog file is not a regular file at <path>), so the next session start retries the confirmation", which bin/fm-teardown.sh:1447 wraps as "...could not be closed atomically (<that>); the pending close is recorded and the next session start retries it". The result is wrong on three counts with no error raised: it asserts an absence that was never reported, repeats the same cause verbatim twice, and says the surviving record exists to settle a confirmation when what the next session start will actually retry is the close (fm_backlog_close_marker_replay probes, finds the row, and re-runs it). Before 8a1e23e this same case printed just "(backlog file is not a regular file at <path>)", so the fix round made it strictly worse. The same composed string is also what bin/fm-bootstrap.sh:1315 prints as "recorded backlog close could not be replayed: <error>". The test that pins this wording, tests/fm-backlog-atomicity.test.sh:1893, fakesdoneas NOT_FOUND, so it exercises only the good case and cannot catch this. Smallest honest remedy that adds no state: make the clause neutral about absence, e.g. "$close_error; this home's backlog row could not be read to confirm whether the item still exists ($FM_BACKLOG_ROW_ERROR), so the next session start retries this close", and follow the header rule at bin/fm-backlog-transition-lib.sh:901-908 into the same neutral form. Flagged ask-user rather than auto-fix because the round-6 decision dictated this exact wording ("worded so the message says the absence could not be confirmed because the backlog read failed, and that the next session start will retry the confirmation"); the defect is the unconditional scope of application, and correcting it necessarily reopens that wording. The precise alternative - only composing the absence clause when the mutation itself reported NOT_FOUND - would require a new output global from fm_backlog_mutate, which extends the change rather than correcting it.bin/fm-backlog-transition-lib.sh:1305- Context for judging scope, no action needed. The User intent's premise that "state/<id>.backlog-close retires only when its transition lands" is only true of teardown's own close. At base commit 6ad419d the replay's absent-row branch already removed the marker and labelled itstale, and bin/fm-bootstrap.sh's case had nostalearm, so session start retired the record silently and printed nothing. The substantive replay-side change here is therefore the honest reporting (the newabsent/absent_incompletelabels and the BACKLOG_RECONCILE line at bin/fm-bootstrap.sh:1300), not the retirement itself; the retirement that the intent asks for is the new teardown-side acceptance at bin/fm-backlog-transition-lib.sh:916-925, which is what turns the incident's loud exit-1 into an accepted completion. The change is intent-conformant on that reading, and I found no component in it that the intent or the prior rounds' accepted findings do not require, so the simplification pass yields nothing to remove.🔧 Fix: claim no absence when close reported none
2 warnings still open:
bin/fm-teardown.sh:1445- Teardown's absent-row reminder makes a hard claim ("its completion link ($deliverable) was never applied - reconcile that artifact by hand.") that has a concrete reachable counterexample, so completion reports a wrong label while exiting 0. docs/configuration.md:102 justifies the hard claim as "because its own close demonstrably failed", but a nonzero exit fromtasks-axi donedoes not demonstrate that no write landed. Concrete sequence in adone_keep = 0home - a configuration this repo supports and tests (tests/fm-captain-hold-lifecycle.test.sh:3148 test_merge_approval_releases_before_zero_done_retention proves a singledonethere removes the row from data/backlog.md and writes it plus its PR into data/done-archive.md): teardown reaches bin/fm-teardown.sh:3488, fm_backlog_done -> fm_backlog_mutate -> fm_tasks_axi wraps the child intimeout -k $bound $bound(bin/fm-backlog-transition-lib.sh:383). The bound expires after tasks-axi has written both backlog.md and done-archive.md but before it exits; the child is TERM/KILLed and fm_backlog_mutate returns 124 with FM_BACKLOG_TRANSITION_ERROR="tasks-axi done <id> did not finish within <n>s". fm_backlog_row_probe then reads the active backlog, finds no row (the close removed it), and reports not_found, so bin/fm-backlog-transition-lib.sh:925 sets FM_BACKLOG_CLOSE_ROW_ABSENT=1, the marker is retired, teardown exits 0, and line 1445 prints "Backlog: <id> had already left <backlog>, so cleanup recorded no close there and its completion link (PR <url>) was never applied - reconcile that artifact by hand." Every clause after "had already left" is false: the close did land and the PR link is in done-archive.md, so the operator is sent to re-apply an artifact that is already applied. Under the defaultdone_keep = 10this is unreachable (the just-closed row is the newest and survives the prune, so the probe finds itdoneand teardown refuses with the close error alone), which is why the hazard is confined to zero-retention homes. Reported as ask-user, not auto-fix, because the remedy is not a repair of a mechanical slip: the smallest honest fix is to align completion's wording with session start's ("could not be confirmed as applied", bin/fm-bootstrap.sh:1292), which directly contradicts the sentence the author wrote deliberately in round 5 at docs/configuration.md:102 and changes user-visible output. Do not re-propose proving the close from the archive - round 2 of this run already removed that archive-proof component on purpose.tests/fm-backlog-atomicity.test.sh:244- The fix round that addedbreak_close_and_row_probe(06425cc) inserted the new helper between the existing comment block and the function that comment documents, so tests/fm-backlog-atomicity.test.sh:244-250 now stacks two mutually contradictory descriptions on one function and leaves the other helper undocumented. Lines 244-247 ("a close reports the NOT_FOUND a vanished row reads as while the row lookup that would confirm that absence fails outright ... exactly the case a recorded close must survive rather than be retired on") describebreak_row_probe_after_absent_closeat line 274, whose fakedoneemitscode: NOT_FOUND. Lines 248-250 ("the close fails for a reason that is not an absence") describebreak_close_and_row_probeat line 251, whose fakedoneemitserror: "backlog is unwritable"and never a NOT_FOUND. A maintainer reading line 251's header is told the helper produces a confirmed absence, which is the precise opposite of the state test_completion_refusal_claims_no_absence_when_the_close_never_reported_one (line 1932) was added to pin, andbreak_row_probe_after_absent_close- the helper the comment actually belongs to, and the one whose NOT_FOUND-plus-read-error pairing is the subtle case - is now bare. Remedy is mechanical with no behavior change: move lines 244-247 down to immediately above line 274 and leave lines 248-250 on line 251.🔧 Fix: align teardown link wording with session start
✅ Re-checked - no issues remain.
bin/fm-bootstrap.sh:103- The script header still asserts the claim rounds 8-9 deliberately removed from the output: "a record retired without landing its close leaves an unapplied completion link, so it reports through BACKLOG_RECONCILE instead." The line the code actually emits (bin/fm-bootstrap.sh:1341 with the disposition from :1334) says only that the link "could not be confirmed as applied", precisely because a close can be killed after its write reached the backlog (docs/configuration.md:102, commit 0a28998). The header also misstates the routing rule: when the record carried no args, FM_BACKLOG_CLOSE_REPLAY_DELIVERABLE is empty, the line reads "it recorded no completion link to reconcile", and the BACKLOG_RECONCILE prefix is still used - the routing is unconditional for absent/absent_incomplete, not conditioned on an unapplied link. Remedy is a comment edit with no behavior change: say a retired close leaves a completion link that could not be confirmed as applied, and that the retirement always reports through BACKLOG_RECONCILE rather than BOOTSTRAP_INFO..agents/skills/bootstrap-diagnostics/SKILL.md:58- The new agent-facing contract line explains the BACKLOG_RECONCILE retirement with "The row left the backlog before this replay could read it", but the change itself adds a test for the reachable subcase where that is false. In test_recovery_reports_a_row_that_left_the_backlog_mid_close (tests/fm-backlog-atomicity.test.sh:2891) the replay's probe reads the row asdone(bin/fm-backlog-transition-lib.sh:1518), then the close inside fm_backlog_close_transition hits NOT_FOUND and sets FM_BACKLOG_CLOSE_ROW_ABSENT=1, so the identical line is emitted for a row the replay demonstrably did read. Line 57's lead-in "the row was already gone" is wrong for the same subcase. The remediation guidance that follows (check where the closed row would have carried the link, file it only if missing) stays correct, so the impact is a diagnostic an agent may reason about with the wrong sequence in mind. Remedy is a wording edit: say the row is no longer in the backlog by the time the close ran, rather than before the replay could read it.bin/fm-bootstrap.sh:1344- The change adds a new operator-facing line for thestalereplay result ("discarded a pending close for <id> recorded by a superseded incarnation; the incarnation now on record still owes its own close"), introduced in 17932d0 whenabsentwas split out ofstale. Nothing asserts it:grepfor "superseded incarnation" or "discarded a pending close" across tests/ returns no hit in fm-backlog-atomicity.test.sh, while every sibling outcome in the same case statement (closed, closed_incomplete, retained, answered, absent) is pinned by an assert_contains. The existing test_recovery_drops_a_close_for_a_newer_meta_incarnation (tests/fm-backlog-atomicity.test.sh:3289) already drives exactly this path and captures bootstrap's output inout, but only checks the marker and row, so a regression in the label or the message would pass silently. Remedy is one assert_contains on$outin that existing test - no new fixture.bin/fm-backlog-transition-lib.sh:1584- Noting an acknowledged tradeoff, no action needed. The absent-close retirement deletes the marker (fm_backlog_close_marker_remove) and reports its completion link exactly once, from a single BACKLOG_RECONCILE line at session start; after that the only durable copy of the link is gone. The sibling retain path deliberately does the opposite: at bin/fm-backlog-transition-lib.sh:1577 it calls fm_backlog_reconcile_marker_write specifically "so the deliverable it carried survives being read only once (or never)", and bin/fm-bootstrap.sh re-reports that record every session start until acked. The author is aware of the asymmetry - the helper's own header at :1445-1448 says "Retiring that record discards the only durable copy of the completion link it carried, so name that link too" - and the two cases differ materially (an absent row leaves nothing outstanding in the backlog to ack, whereas an unresolved captain call does). Flagged only so the one-shot reporting is a visible choice rather than an oversight; closing the gap would mean a new durable record and an ack path, which extends the change beyond its intent and would need authorization, so nothing here should block the merge.✅ **Test** - passed
✅ No issues found.
bash tests/fm-backlog-atomicity.test.sh— the owning suite, 109 ok / 0 not ok at targetbash tests/fm-backlog-read-bound.test.sh— 10 ok / 0 not ok (sibling suite touching the same library globals)bash tests/fm-tasks-axi.test.sh— 8 ok / 0 not okPer-test regression proof, each new test run in isolation with onlybin/rolled back to base6ad419dathen again at target:test_completion_accepts_a_row_already_archived_by_retention,test_completion_keeps_a_close_whose_row_lookup_fails,test_completion_refusal_claims_no_absence_when_the_close_never_reported_one,test_interrupted_cleanup_of_an_archived_row_still_warns_about_its_endpoint,test_recovery_retires_a_close_for_a_row_archived_by_retention,test_recovery_retires_a_close_for_a_row_removed_without_closing,test_recovery_reports_a_row_that_left_the_backlog_mid_close— all rc=1 at base, rc=0 at targetManual end-to-end operator transcript: built a real home with a realdata/backlog.md, rantasks-axi done <id>thentasks-axi prune --keep 0 --state done, confirmedtasks-axi show <id>answerscode: NOT_FOUND, then drove the realbin/fm-teardown.shand captured its verbatim completion output, exit status, and the presence/absence ofstate/<id>.backlog-closeManual end-to-end session-start transcript: hand-wrote astate/<id>.backlog-closerecord carryingarg=--pr https://github.com/example/repo/pull/42for a pruned row, ran the realbin/fm-bootstrap.sh, and captured theBACKLOG_RECONCILEline and the marker's fateManual unconfirmed-absence transcript: shadowedtasks-axisodonereports NOT_FOUND and the followingshowfails, then ran real teardown and captured the refusal text, exit 1, and the preserved recordManual control transcript: ordinary in-flight row through real teardown — exit 0,is closed in <backlog>, row readsdoneSetup fix: installedtasks-axi@0.2.6(repo floorFM_TASKS_AXI_MIN=0.2.4) into a throwaway/tmpprefix, since the host had none and the suite self-skips without it; removed the prefix afterward✅ No issues found.
Installed the CI-pinnedtasks-axi@0.2.5into a throwaway /tmp npm prefix (not installed on this host, which otherwise makes the whole suite self-skip); removed afterwardsbash tests/fm-backlog-atomicity.test.sh— full owning suite at target: 120 ok, 0 not oktests/fm-backlog-atomicity.test.sh::test_completion_accepts_a_row_already_archived_by_retention(the filed incident)tests/fm-backlog-atomicity.test.sh::test_completion_keeps_a_close_whose_row_lookup_failstests/fm-backlog-atomicity.test.sh::test_completion_refusal_claims_no_absence_when_the_close_never_reported_onetests/fm-backlog-atomicity.test.sh::test_interrupted_cleanup_of_an_archived_row_still_warns_about_its_endpointtests/fm-backlog-atomicity.test.sh::test_recovery_retires_a_close_for_a_row_archived_by_retentiontests/fm-backlog-atomicity.test.sh::test_recovery_retires_a_close_for_a_row_removed_without_closingtests/fm-backlog-atomicity.test.sh::test_recovery_reports_a_row_that_left_the_backlog_mid_closeNeighbour regression guards re-run at target:test_completion_fails_loudly_and_records_the_close_it_still_owes,test_interrupted_destructive_cleanup_leaves_a_recoverable_close,test_recovery_preserves_a_close_when_the_backlog_cannot_be_read,test_recovery_reconcile_record_preserves_incomplete_cleanup_warningbash tests/fm-backlog-read-bound.test.sh— 10 ok, 0 not okbash tests/fm-tasks-axi.test.sh— 8 ok, 0 not okFail-before proof:git archive 5719dacinto a temp tree, copied in only the newtests/fm-backlog-atomicity.test.sh, ran each of the 7 new tests individually — all 7 rc=1 with the incident message / missing-wording assertionsManual evidence capture: drove the realbin/fm-teardown.shagainst a row closed then pruned withtasks-axi prune --keep 0 --state done(sotasks-axi showanswerscode: NOT_FOUND) at base and at target, recording stdout, exit status, and whetherstate/<id>.backlog-closesurvivedManual evidence capture: drove the realbin/fm-bootstrap.shover a hand-writtenstate/<id>.backlog-closecarryingarg=--pr https://example.test/pr/7for a row removed withtasks-axi rm, at base and at targetManual evidence capture: both loud-refusal paths (close reports NOT_FOUND with a failing confirming row read; close fails for a non-absence reason with a failing row read) recorded verbatim viabin/fm-teardown.shgit status --porcelain— worktree clean; temp npm prefix, temp base checkout, and scratch drivers removed✅ **Document** - passed
✅ No issues found.
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ No issues found.
✅ **Push** - passed
✅ No issues found.
✅ No issues found.